Small-Sample Tree-Ensemble Models for Climate-Smart Crop Selection in Low-Resource Ethiopian Agriculture
DOI:
https://doi.org/10.69987/JACS.2024.40608Keywords:
climate-smart agriculture, crop recommendation, small-sample learning, tree ensembles, XGBoost, CatBoost, RandomForest, EthiopiaAbstract
Low-resource agricultural decision support requires models that can learn from small, heterogeneous, and locally collected tabular data. This study evaluates a crop-selection workflow for an Ethiopian crop recommendation table containing soil color, soil pH, macro- and micronutrients, seasonal humidity, maximum and minimum temperature, precipitation, wind, cloud amount, surface pressure, and crop labels. Because the main supervised table contains crop labels but no row-level yield target, the experiment is framed as multiclass crop-label recommendation rather than yield regression. To clarify the data scope, a regional cereal yield table and the EthCT2020 georeferenced crop-type reference are used as contextual datasets, not as merged row-level training targets. CatBoost, RandomForest, and XGBoost were fitted on a fixed stratified 75/25 split with low-resource training sizes of 120, 240, 480, 960, and 2,900 records. On the full training pool, RandomForest achieved the highest top-1 accuracy at 0.501, followed by XGBoost at 0.492 and CatBoost at 0.481. XGBoost produced the strongest top-3 shortlist accuracy at 0.838, while CatBoost reached 0.825 and RandomForest reached 0.810. Macro-F1 remained low for all models, showing that rare crops were not reliably recovered under the observed class imbalance. The feature ablation and RandomForest importance analysis show that soil color, pH, and nutrient fields carry most of the separative signal, while climate variables add context but are weaker when used alone. The results support fast tabular crop recommendation for data-scarce settings, while also showing that rare-crop coverage and row-level yield prediction require additional field-level data.







