نوع مقاله : مروری
نویسندگان
بخش تحقیقات زراعی و باغی، مرکز تحقیقات و آموزش کشاورزی و منابع طبیعی فارس، سازمان تحقیقات، آموزش و ترویج کشاورزی، شیراز، ایران
چکیده
کلیدواژهها
عنوان مقاله [English]
نویسندگان [English]
Introduction: Accurate crop yield prediction, particularly for cereal crops as the primary source of global food energy, plays a critical role in food security management, production planning, agricultural trade, and climate change adaptation. Although traditional approaches based on linear regression and classical statistical methods have been widely applied, they often face limitations when dealing with complex, nonlinear, high-dimensional, and heterogeneous datasets. Recent advances in remote sensing technologies, field-based sensors, geographic information systems (GIS), and the rapid growth of agricultural data have accelerated the adoption of machine learning techniques in crop yield prediction. By capturing hidden patterns and nonlinear interactions among climatic, management, soil, and genetic factors, machine learning algorithms have emerged as powerful tools for predictive modeling in agriculture. This review aims to examine the application of the most widely used machine learning algorithms for predicting the yield of major cereal crops, including wheat, barley, rice, and maize, while analyzing the key stages of model development and highlighting current challenges and future research opportunities.
Materials and Methods: This review synthesizes published studies on crop yield prediction using machine learning approaches. Commonly applied algorithms include multiple regression, regression trees, random forests, support vector machines (SVM), multilayer perceptrons (MLP), artificial neural networks (ANN), and deep learning architectures such as deep neural networks (DNN) and long short-term memory (LSTM) networks. The development of predictive models typically involves several sequential stages, including data acquisition, preprocessing, feature selection, model training, validation, and performance evaluation. In most studies, datasets are divided into training and testing subsets, and model performance is assessed using appropriate statistical and machine learning evaluation metrics.
Results: The reviewed studies indicate that input data quality, feature selection strategies, and validation procedures are among the most influential factors affecting prediction accuracy. Frequently used predictors include temperature, precipitation, soil moisture, soil physicochemical properties, remotely sensed vegetation indices, crop management practices, and plant genetic information. Data preprocessing procedures, including missing-value treatment, outlier detection, normalization, and dimensionality reduction, substantially improves model robustness and predictive performance. Furthermore, feature selection methods based on filter, wrapper, and embedded approaches help eliminate redundant variables and enhance model efficiency. Comparative analyses across studies revealed that tree-based algorithms, particularly Random Forest and Gradient Boosting methods, consistently achieve strong predictive performance under diverse environmental conditions. Deep learning models demonstrate superior capability in capturing complex nonlinear relationships when large-scale datasets are available. Model performance is commonly evaluated using metrics such as the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), accuracy, recall, and F1-score. In addition, cross-validation techniques, especially K-fold cross-validation, are widely employed to assess model stability and generalizability.
Conclusion: Machine learning offers substantial potential for improving crop yield prediction accuracy and supporting data-driven decision-making in precision agriculture. However, several challenges remain, including the limited availability of standardized long-term datasets, data heterogeneity, the high cost of data acquisition technologies, and the limited interpretability of complex predictive models. Future research should focus on the development of large-scale multi-source databases, the integration of climatic, soil, genetic, and remote sensing information, the adoption of advanced deep learning and explainable artificial intelligence (XAI) frameworks, and the validation of predictive models across diverse environmental and management conditions. Addressing these challenges will enhance the robustness, scalability, and practical applicability of machine learning-based crop yield prediction systems.
کلیدواژهها [English]