Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

Crop Yield Prediction in India using Machine Learning: An Exploratory Data Analysis Approach

Authors

Manjesh R, Prasad M R, Parinitha P Rao, Vineeta Koppalkar

Abstract

Agriculture is one of the most important sectors of Indian economy and thus needs reliable yield forecasts to provide food security and better management of resources as well as formulate informed policy decisions. This paper proposes to utilize a systematic way of forecasting crop yields by integrating Exploratory Data Analysis (EDA) with machine learning techniques. The dataset with state, district, crop year, season, crop type, cultivated area, and production as its attributes allowed revealing the important patterns in the form of seasonal changes, regional variations, and yield trends of the past. Predictive models were then developed by using these observations. Four algorithms, which are Linear Regression, Decision Tree Regressor, Random Forest Regressor, and XGBoost, were used and compared. Random Forest has shown the best performance with the highest R2 of 0.96 when compared to the Decision Tree (0.95), XGBoost (0.93), and Linear Regression (0.89). This finding is an indication of the capability of Random Forest to capture the complex and nonlinearities that exist in agricultural data. The key findings of the work are as follows (i) comprehensive data preprocessing, (ii) the application of EDA knowledge to assist in model development, and (iii) the use of seasonal factors as predictors. The combination of these steps enhanced the models in terms of accuracy and interpretability. The results of this study can be of good advice to the policy makers and other stakeholders in agriculture in India to make plans and encourage sustainable agricultural practices.