الگوریتم‌های یادگیری ماشین در پیش‌بینی عملکرد غلات: یک مرور تحلیلی

نوع مقاله : مروری

نویسندگان

بخش تحقیقات زراعی و باغی، مرکز تحقیقات و آموزش کشاورزی و منابع طبیعی فارس، سازمان تحقیقات، آموزش و ترویج کشاورزی، شیراز، ایران

10.22126/cbb.2026.13903.1145

چکیده

مقدمه: پیش‌بینی دقیق عملکرد محصولات زراعی، به‌ویژه غلات به‌عنوان اصلی‌ترین منبع تأمین انرژی غذایی جهان، نقشی کلیدی در مدیریت امنیت غذایی، برنامه‌ریزی تولید، تجارت محصولات کشاورزی و سازگاری با تغییرات اقلیمی ایفا می‌کند. روش‌های سنتی مبتنی بر مدل‌های رگرسیون خطی و آمار کلاسیک، اگرچه در بسیاری از مطالعات مورد استفاده قرار گرفته‌اند، اما در مواجهه با داده‌های پیچیده، غیرخطی، چندبعدی و ناهمگن با محدودیت‌هایی مواجه هستند. در سال‌های اخیر، پیشرفت فناوری‌های سنجش از دور، حسگرهای مزرعه‌ای، سامانه‌های اطلاعات جغرافیایی و افزایش حجم داده‌های کشاورزی، زمینه استفاده گسترده از روش‌های یادگیری ماشین را فراهم کرده است. این الگوریتم‌ها با قابلیت شناسایی الگوهای پنهان و روابط غیرخطی میان عوامل اقلیمی، مدیریتی، خاکی و ژنتیکی، به ابزارهای مؤثری برای پیش‌بینی عملکرد محصولات تبدیل شده‌اند. هدف این مقاله مروری، بررسی کاربرد رایج‌ترین الگوریتم‌های یادگیری ماشین در پیش‌بینی عملکرد گندم، جو، برنج و ذرت، تحلیل مراحل توسعه مدل‌های پیش‌بینی و تبیین چالش‌ها و فرصت‌های پژوهشی موجود در این حوزه است.
مواد و روش‌ها: در این مطالعه، پژوهش‌های منتشرشده در زمینه پیش‌بینی عملکرد محصولات زراعی با استفاده از یادگیری ماشین مورد بررسی قرار گرفتند. الگوریتم‌هایی نظیر رگرسیون چندمتغیره، درخت رگرسیون، جنگل تصادفی، ماشین بردار پشتیبان، پرسپترون چندلایه، شبکه‌های عصبی مصنوعی و مدل‌های یادگیری عمیق از جمله شبکه‌های عصبی عمیق و حافظه بلندمدت کوتاه‌مدت از پرکاربردترین روش‌های مورد استفاده در مطالعات مختلف هستند. توسعه مدل‌های پیش‌بینی شامل مراحل جمع‌آوری داده، پیش‌پردازش، انتخاب ویژگی، آموزش مدل، اعتبارسنجی و ارزیابی عملکرد است. در این فرآیند، داده‌ها به مجموعه‌های آموزشی و آزمایشی تقسیم شده و عملکرد مدل‌ها با استفاده از معیارهای آماری مناسب مورد ارزیابی قرار می‌گیرد.
یافته‌ها: نتایج مطالعات نشان می‌دهد که کیفیت داده‌های ورودی، انتخاب ویژگی‌های مناسب و روش ارزیابی مدل از مهم‌ترین عوامل مؤثر بر دقت پیش‌بینی هستند. متغیرهایی نظیر دما، بارش، رطوبت خاک، خصوصیات فیزیکی و شیمیایی خاک، شاخص‌های پوشش گیاهی حاصل از سنجش از دور، عملیات مدیریتی مزرعه و اطلاعات ژنتیکی گیاهان از مهم‌ترین متغیرهای مورد استفاده در مدل‌های پیش‌بینی عملکرد به شمار می‌روند. پیش‌پردازش داده‌ها شامل حذف داده‌های گمشده، شناسایی مقادیر پرت، نرمال‌سازی و کاهش ابعاد داده، نقش مهمی در بهبود کارایی مدل‌ها دارد. همچنین، انتخاب ویژگی با استفاده از روش‌های فیلتر، بسته‌بند و تعبیه‌شده موجب حذف متغیرهای غیرمؤثر و افزایش دقت پیش‌بینی می‌شود. بررسی مطالعات مختلف نشان داد که الگوریتم‌های مبتنی بر درخت مانند جنگل تصادفی و تقویت گرادیان، در بسیاری از شرایط عملکرد مطلوبی داشته‌اند، در حالی که مدل‌های یادگیری عمیق در صورت دسترسی به مجموعه داده‌های بزرگ، توانایی بیشتری در استخراج روابط پیچیده و بهبود دقت پیش‌بینی نشان داده‌اند. ارزیابی عملکرد مدل‌ها عمدتاً با استفاده از معیارهایی نظیر ضریب تبیین، ریشه میانگین مربعات خطا، میانگین خطای مطلق، دقت، فراخوانی و امتیاز F1 انجام می‌شود. همچنین، اعتبارسنجی متقابل، به‌ویژه روش K-fold، یکی از متداول‌ترین رویکردها برای ارزیابی قابلیت تعمیم مدل‌ها است.
نتیجه‌گیری: یادگیری ماشین ظرفیت قابل‌توجهی برای افزایش دقت پیش‌بینی عملکرد محصولات زراعی و پشتیبانی از تصمیم‌گیری در کشاورزی دقیق دارد. با این حال، محدودیت‌هایی نظیر کمبود داده‌های استاندارد و بلندمدت، ناهمگنی داده‌ها، هزینه بالای فناوری‌های جمع‌آوری داده و تفسیرپذیری محدود برخی مدل‌های پیچیده همچنان از چالش‌های اصلی این حوزه محسوب می‌شوند. توسعه پایگاه‌های داده چندمنبعی، ادغام داده‌های اقلیمی، خاکی، ژنتیکی و سنجش از دور، بهره‌گیری از روش‌های یادگیری عمیق و هوش مصنوعی توضیح‌پذیر و همچنین اعتبارسنجی مدل‌ها در شرایط اقلیمی و مدیریتی متنوع، از مهم‌ترین اولویت‌های پژوهشی آینده برای افزایش قابلیت کاربرد و تعمیم‌پذیری مدل‌های پیش‌بینی عملکرد خواهند بود.

کلیدواژه‌ها


عنوان مقاله [English]

Machine Learning Algorithms in Cereal Yield Prediction: An Analytical Review

نویسندگان [English]

  • Leyla Nazari
  • Khadijeh Alijani
Crop and Horticultural Science Research Department, Fars Agricultural and Natural Resources Research and Education Center, Agricultural Research, Education and Extension Organization (AREEO), Shiraz, Iran
چکیده [English]

Introduction: Accurate crop yield prediction, particularly for cereal crops as the primary source of global food energy, plays a critical role in food security management, production planning, agricultural trade, and climate change adaptation. Although traditional approaches based on linear regression and classical statistical methods have been widely applied, they often face limitations when dealing with complex, nonlinear, high-dimensional, and heterogeneous datasets. Recent advances in remote sensing technologies, field-based sensors, geographic information systems (GIS), and the rapid growth of agricultural data have accelerated the adoption of machine learning techniques in crop yield prediction. By capturing hidden patterns and nonlinear interactions among climatic, management, soil, and genetic factors, machine learning algorithms have emerged as powerful tools for predictive modeling in agriculture. This review aims to examine the application of the most widely used machine learning algorithms for predicting the yield of major cereal crops, including wheat, barley, rice, and maize, while analyzing the key stages of model development and highlighting current challenges and future research opportunities.
Materials and Methods: This review synthesizes published studies on crop yield prediction using machine learning approaches. Commonly applied algorithms include multiple regression, regression trees, random forests, support vector machines (SVM), multilayer perceptrons (MLP), artificial neural networks (ANN), and deep learning architectures such as deep neural networks (DNN) and long short-term memory (LSTM) networks. The development of predictive models typically involves several sequential stages, including data acquisition, preprocessing, feature selection, model training, validation, and performance evaluation. In most studies, datasets are divided into training and testing subsets, and model performance is assessed using appropriate statistical and machine learning evaluation metrics.
Results: The reviewed studies indicate that input data quality, feature selection strategies, and validation procedures are among the most influential factors affecting prediction accuracy. Frequently used predictors include temperature, precipitation, soil moisture, soil physicochemical properties, remotely sensed vegetation indices, crop management practices, and plant genetic information. Data preprocessing procedures, including missing-value treatment, outlier detection, normalization, and dimensionality reduction, substantially improves model robustness and predictive performance. Furthermore, feature selection methods based on filter, wrapper, and embedded approaches help eliminate redundant variables and enhance model efficiency. Comparative analyses across studies revealed that tree-based algorithms, particularly Random Forest and Gradient Boosting methods, consistently achieve strong predictive performance under diverse environmental conditions. Deep learning models demonstrate superior capability in capturing complex nonlinear relationships when large-scale datasets are available. Model performance is commonly evaluated using metrics such as the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), accuracy, recall, and F1-score. In addition, cross-validation techniques, especially K-fold cross-validation, are widely employed to assess model stability and generalizability.
Conclusion: Machine learning offers substantial potential for improving crop yield prediction accuracy and supporting data-driven decision-making in precision agriculture. However, several challenges remain, including the limited availability of standardized long-term datasets, data heterogeneity, the high cost of data acquisition technologies, and the limited interpretability of complex predictive models. Future research should focus on the development of large-scale multi-source databases, the integration of climatic, soil, genetic, and remote sensing information, the adoption of advanced deep learning and explainable artificial intelligence (XAI) frameworks, and the validation of predictive models across diverse environmental and management conditions. Addressing these challenges will enhance the robustness, scalability, and practical applicability of machine learning-based crop yield prediction systems.

کلیدواژه‌ها [English]

  • Cereal Crops
  • Cross-validation
  • Deep Learning
  • Feature Selection
  • Machine Learning
  • Yield Prediction