{
  "id": 614776,
  "title": "FutureCrop Challenge — 2nd Place Solution",
  "url": "/competitions/the-future-crop-challenge/writeups/futurecrop-challenge-2nd-place-solution",
  "author_name": "",
  "post_date": "2025-11-06T12:12:24.277Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>TL;DR</h2>\n<ul>\n<li><strong>Task:</strong> Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.</li>\n<li><strong>Data:</strong> Train 1982–2020 → Test 2021–2098 (high-emissions, rainfed).</li>\n<li><strong>Core idea:</strong> Summarize 240 daily features into global and 8×30-day chunk stats, plus train-only per-location yield priors (lat, lon).</li>\n<li><strong>Model:</strong> LightGBM (GBDT) — ~100k trees, lr≈0.01, early stopping≈50.</li>\n<li><strong>Performance:</strong> sMAPE ≈ 14.7 %, R² ≈ 0.96, MAE ≈ 0.31, RMSE ≈ 0.49 (train); RMSE ≈ 1.07 (public validation, 2nd place).</li>\n<li><strong>Why it works:</strong> Captures climate patterns, seasonality, and site productivity while avoiding overfit and enabling fast and robust training.</li>\n</ul>\n<hr>\n<h1>Problem &amp; Data</h1>\n<ul>\n<li><strong>Task:</strong> Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.</li>\n<li><strong>Time horizon:</strong> Forecast <strong>2021–2098</strong> using models trained on <strong>1982–2020</strong> (train : test ≈ 1 : 2).</li>\n<li><strong>Scenario:</strong> High-emissions (rising CO₂), rainfed conditions with time-invariant soils and fixed per-country nitrogen rates.</li>\n<li><strong>Inputs (per datapoint):</strong><ul>\n<li><strong>240 daily climate variables:</strong><ul>\n<li>Precipitation (PR)</li>\n<li>Shortwave radiation (RSDS)</li>\n<li>Mean, minimum, and maximum temperature (TAS, TMIN, TMAX)</li></ul></li>\n<li><strong>Auxiliaries:</strong><ul>\n<li><strong>Soil texture class</strong> (USDA 1–13)</li>\n<li><strong>Nitrogen rate</strong> (country-level, fixed)</li>\n<li><strong>CO₂ concentration</strong> (spatially constant, rising over time)</li>\n<li><strong>Planting date</strong> and real_year</li></ul></li></ul></li>\n<li><strong>Evaluation metrics:</strong><ul>\n<li><strong>Primary:</strong> RMSE on 2021–2098 (private leaderboard)</li>\n<li><strong>Secondary:</strong> Median per-cell R² and regional production R² (Iowa maize, Germany wheat).</li></ul></li>\n</ul>\n<hr>\n<h2>Design Principles</h2>\n<ol>\n<li><strong>Respect the shift:</strong> Future climate and CO₂ drift ⇒ prefer features that remain meaningful under shift (nonparametric summaries, chunk seasonality) over high-capacity sequence models that can overfit training climatology.</li>\n<li><strong>Keep it lean:</strong> Prioritize fast iteration and low memory; extensive aggregation &gt; raw 240-day sequences.</li>\n<li><strong>Add geographic priors (but no leakage):</strong> Per-location mean yields computed only on training add ststrong signal without touching future labels.</li>\n<li><strong>Determinism:</strong> Fixed seeds, version-locked libs, explicit preprocessing.</li>\n</ol>\n<hr>\n<h2>Feature Engineering</h2>\n<ul>\n<li><strong>Daily → Global Summaries (per variable, per sample):</strong> <code>mean</code>, <code>median</code>, <code>sum</code>, <code>min</code>, <code>max</code> for each of: <code>pr</code>, <code>rsds</code>, <code>tas</code> (and analogously for <code>tmin</code>, <code>tmax</code> when helpful).</li>\n<li><strong>8×30-day “Chunk” Statistics:</strong><ul>\n<li>For each variable, split the 240 days into 8 contiguous 30-day windows and compute the same statistics per chunk.</li>\n<li>Captures intra-season timing (e.g., early-season heat, mid-season radiation, late-season water stress) without high-dimensional sequences.</li></ul></li>\n<li><strong>Per-Location Yield Priors (Train-only):</strong><ul>\n<li>Compute mean yield at each <code>(latitude, longitude)</code> separately for maize and wheat using only training years; join these two features into both crops at inference.</li>\n<li>Intuition: encodes site-level productivity (soils/management beyond observable features) <strong>without</strong> using future labels.</li></ul></li>\n<li><strong>Simple Interactions &amp; Calendar:</strong><ul>\n<li>Mild interactions such as <code>nitrogen × texture_class</code>, <code>nitrogen / CO₂</code>, and a coarse decade feature from <code>real_year</code>.</li>\n<li>Avoid heavy cross terms to keep modemodel simple and stable across shifts.</li></ul></li>\n<li><strong>Pruning:</strong> After creating summaries, drop the raw 240-day columns to reduce memory and prevent tree splits on noisy day-level artifacts.</li>\n</ul>\n<hr>\n<h2>Model</h2>\n<ul>\n<li><strong>Algorithm:</strong> LightGBM (GBDT)</li>\n<li><strong>Key params:</strong><ul>\n<li><code>objective=regression</code></li>\n<li><code>metric=rmse</code></li>\n<li><code>learning_rate=0.01</code></li>\n<li><code>n_estimators=100000</code> with early stopping ≈ 50 rounds</li>\n<li>Default depth/leaf settings performed well; minor tuning had a marginal impact</li></ul></li>\n<li><strong>Why LightGBM:</strong> strong tabular baseline, fast, robust to mixed scales, and tolerant of many moderately informative features.</li>\n</ul>\n<hr>\n<h2>Validation &amp; Leakage Guards</h2>\n<ul>\n<li><strong>Primary split</strong> for model selection: <strong>random 80/20</strong> (to iterate quickly on FE/model).</li>\n<li><strong>Temporal sanity checks:</strong> additional <strong>forward holdouts</strong> (e.g., last 5 training years) to verify decade-to-decade stability.</li>\n<li><strong>Leakage checks:</strong><ul>\n<li>Per-location means computed <strong>only from training</strong> and merged by <code>(latitude, longitude)</code>; <strong>no</strong> target leakage from test.</li>\n<li>Avoid using the internal <code>year</code> simulation index as a proxy for time on test.</li>\n<li>Confirm units and scales (e.g., <code>pr</code> in kg m⁻² s⁻¹).</li></ul></li>\n<li><strong>Seeds:</strong> fixed (<code>random_state=42</code>) for reproducibility.</li>\n</ul>\n<hr>\n<h2>Results Summary</h2>\n<table>\n<thead>\n<tr>\n<th><strong>Metric</strong></th>\n<th><strong>Training (1982–2020)</strong></th>\n<th><strong>Public Validation (2021–2050, Kaggle LB)</strong></th>\n<th><strong>Remarks</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>RMSE</strong></td>\n<td><strong>≈ 0.49</strong></td>\n<td><strong>≈ 1.07 (2nd place)</strong></td>\n<td>Primary leaderboard metric</td>\n</tr>\n<tr>\n<td><strong>MAE</strong></td>\n<td>≈ 0.31</td>\n<td>—</td>\n<td>Low absolute deviation across grids</td>\n</tr>\n<tr>\n<td><strong>MAPE</strong></td>\n<td>≈ 14.7 %</td>\n<td>—</td>\n<td>Stable relative accuracy across crops</td>\n</tr>\n<tr>\n<td><strong>R²</strong></td>\n<td>≈ 0.96</td>\n<td>—</td>\n<td>Strong fit and generalization on past decades</td>\n</tr>\n</tbody>\n</table>\n<p>Comprehensive evaluation results will be published by the AgML competition team following the official review.</p>\n<hr>\n<h2>Ablations (What Moved the Needle)</h2>\n<ul>\n<li><strong>+ Per-location mean yields (2 features):</strong> Large gain; strong prior for site productivity.</li>\n<li><strong>+ 8×30-day chunks:</strong> Medium gain; preserves phenology/seasonality patterns.</li>\n<li><strong>+ Interactions (nitrogen×texture, nitrogen/CO₂):</strong> Small but consistent.</li>\n<li><strong>– Removing raw 240-day columns:</strong> Neutral→positive (less overfit, lower memory).</li>\n<li><strong>± Heavier models (XGBoost/CatBoost/NNs):</strong> No consistent win vs LGBM after robust FE; training slower.</li>\n<li><strong>± Extensive hyper-tuning:</strong> marginal improvements (&lt; LB jitter) once core FE is in place.</li>\n</ul>\n<hr>\n<h2>Ethics &amp; Responsible Use</h2>\n<p>This solution emulates simulated yields under a specific high-emissions scenario and rainfed assumptions. Real-world decisions should consider local agronomy, uncertainty, and the gap between simulated and observed systems (irrigation, policy, technology change).</p>\n<hr>\n<h2>Acknowledgments</h2>\n<p>Special thanks to the AgML / AgMIP team for providing the FutureCrop benchmark and fostering open agricultural AI research. Gratitude to the broader community and contributors whose datasets, feedback, and discussions helped refine this work.</p>",
  "messages": [
    {
      "id": "3312083",
      "postDate": "11/06/2025 12:12:03",
      "content": "<h2>TL;DR</h2>\n<ul>\n<li><strong>Task:</strong> Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.</li>\n<li><strong>Data:</strong> Train 1982–2020 → Test 2021–2098 (high-emissions, rainfed).</li>\n<li><strong>Core idea:</strong> Summarize 240 daily features into global and 8×30-day chunk stats, plus train-only per-location yield priors (lat, lon).</li>\n<li><strong>Model:</strong> LightGBM (GBDT) — ~100k trees, lr≈0.01, early stopping≈50.</li>\n<li><strong>Performance:</strong> sMAPE ≈ 14.7 %, R² ≈ 0.96, MAE ≈ 0.31, RMSE ≈ 0.49 (train); RMSE ≈ 1.07 (public validation, 2nd place).</li>\n<li><strong>Why it works:</strong> Captures climate patterns, seasonality, and site productivity while avoiding overfit and enabling fast and robust training.</li>\n</ul>\n<hr>\n<h1>Problem &amp; Data</h1>\n<ul>\n<li><strong>Task:</strong> Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.</li>\n<li><strong>Time horizon:</strong> Forecast <strong>2021–2098</strong> using models trained on <strong>1982–2020</strong> (train : test ≈ 1 : 2).</li>\n<li><strong>Scenario:</strong> High-emissions (rising CO₂), rainfed conditions with time-invariant soils and fixed per-country nitrogen rates.</li>\n<li><strong>Inputs (per datapoint):</strong><ul>\n<li><strong>240 daily climate variables:</strong><ul>\n<li>Precipitation (PR)</li>\n<li>Shortwave radiation (RSDS)</li>\n<li>Mean, minimum, and maximum temperature (TAS, TMIN, TMAX)</li></ul></li>\n<li><strong>Auxiliaries:</strong><ul>\n<li><strong>Soil texture class</strong> (USDA 1–13)</li>\n<li><strong>Nitrogen rate</strong> (country-level, fixed)</li>\n<li><strong>CO₂ concentration</strong> (spatially constant, rising over time)</li>\n<li><strong>Planting date</strong> and real_year</li></ul></li></ul></li>\n<li><strong>Evaluation metrics:</strong><ul>\n<li><strong>Primary:</strong> RMSE on 2021–2098 (private leaderboard)</li>\n<li><strong>Secondary:</strong> Median per-cell R² and regional production R² (Iowa maize, Germany wheat).</li></ul></li>\n</ul>\n<hr>\n<h2>Design Principles</h2>\n<ol>\n<li><strong>Respect the shift:</strong> Future climate and CO₂ drift ⇒ prefer features that remain meaningful under shift (nonparametric summaries, chunk seasonality) over high-capacity sequence models that can overfit training climatology.</li>\n<li><strong>Keep it lean:</strong> Prioritize fast iteration and low memory; extensive aggregation &gt; raw 240-day sequences.</li>\n<li><strong>Add geographic priors (but no leakage):</strong> Per-location mean yields computed only on training add ststrong signal without touching future labels.</li>\n<li><strong>Determinism:</strong> Fixed seeds, version-locked libs, explicit preprocessing.</li>\n</ol>\n<hr>\n<h2>Feature Engineering</h2>\n<ul>\n<li><strong>Daily → Global Summaries (per variable, per sample):</strong> <code>mean</code>, <code>median</code>, <code>sum</code>, <code>min</code>, <code>max</code> for each of: <code>pr</code>, <code>rsds</code>, <code>tas</code> (and analogously for <code>tmin</code>, <code>tmax</code> when helpful).</li>\n<li><strong>8×30-day “Chunk” Statistics:</strong><ul>\n<li>For each variable, split the 240 days into 8 contiguous 30-day windows and compute the same statistics per chunk.</li>\n<li>Captures intra-season timing (e.g., early-season heat, mid-season radiation, late-season water stress) without high-dimensional sequences.</li></ul></li>\n<li><strong>Per-Location Yield Priors (Train-only):</strong><ul>\n<li>Compute mean yield at each <code>(latitude, longitude)</code> separately for maize and wheat using only training years; join these two features into both crops at inference.</li>\n<li>Intuition: encodes site-level productivity (soils/management beyond observable features) <strong>without</strong> using future labels.</li></ul></li>\n<li><strong>Simple Interactions &amp; Calendar:</strong><ul>\n<li>Mild interactions such as <code>nitrogen × texture_class</code>, <code>nitrogen / CO₂</code>, and a coarse decade feature from <code>real_year</code>.</li>\n<li>Avoid heavy cross terms to keep modemodel simple and stable across shifts.</li></ul></li>\n<li><strong>Pruning:</strong> After creating summaries, drop the raw 240-day columns to reduce memory and prevent tree splits on noisy day-level artifacts.</li>\n</ul>\n<hr>\n<h2>Model</h2>\n<ul>\n<li><strong>Algorithm:</strong> LightGBM (GBDT)</li>\n<li><strong>Key params:</strong><ul>\n<li><code>objective=regression</code></li>\n<li><code>metric=rmse</code></li>\n<li><code>learning_rate=0.01</code></li>\n<li><code>n_estimators=100000</code> with early stopping ≈ 50 rounds</li>\n<li>Default depth/leaf settings performed well; minor tuning had a marginal impact</li></ul></li>\n<li><strong>Why LightGBM:</strong> strong tabular baseline, fast, robust to mixed scales, and tolerant of many moderately informative features.</li>\n</ul>\n<hr>\n<h2>Validation &amp; Leakage Guards</h2>\n<ul>\n<li><strong>Primary split</strong> for model selection: <strong>random 80/20</strong> (to iterate quickly on FE/model).</li>\n<li><strong>Temporal sanity checks:</strong> additional <strong>forward holdouts</strong> (e.g., last 5 training years) to verify decade-to-decade stability.</li>\n<li><strong>Leakage checks:</strong><ul>\n<li>Per-location means computed <strong>only from training</strong> and merged by <code>(latitude, longitude)</code>; <strong>no</strong> target leakage from test.</li>\n<li>Avoid using the internal <code>year</code> simulation index as a proxy for time on test.</li>\n<li>Confirm units and scales (e.g., <code>pr</code> in kg m⁻² s⁻¹).</li></ul></li>\n<li><strong>Seeds:</strong> fixed (<code>random_state=42</code>) for reproducibility.</li>\n</ul>\n<hr>\n<h2>Results Summary</h2>\n<table>\n<thead>\n<tr>\n<th><strong>Metric</strong></th>\n<th><strong>Training (1982–2020)</strong></th>\n<th><strong>Public Validation (2021–2050, Kaggle LB)</strong></th>\n<th><strong>Remarks</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>RMSE</strong></td>\n<td><strong>≈ 0.49</strong></td>\n<td><strong>≈ 1.07 (2nd place)</strong></td>\n<td>Primary leaderboard metric</td>\n</tr>\n<tr>\n<td><strong>MAE</strong></td>\n<td>≈ 0.31</td>\n<td>—</td>\n<td>Low absolute deviation across grids</td>\n</tr>\n<tr>\n<td><strong>MAPE</strong></td>\n<td>≈ 14.7 %</td>\n<td>—</td>\n<td>Stable relative accuracy across crops</td>\n</tr>\n<tr>\n<td><strong>R²</strong></td>\n<td>≈ 0.96</td>\n<td>—</td>\n<td>Strong fit and generalization on past decades</td>\n</tr>\n</tbody>\n</table>\n<p>Comprehensive evaluation results will be published by the AgML competition team following the official review.</p>\n<hr>\n<h2>Ablations (What Moved the Needle)</h2>\n<ul>\n<li><strong>+ Per-location mean yields (2 features):</strong> Large gain; strong prior for site productivity.</li>\n<li><strong>+ 8×30-day chunks:</strong> Medium gain; preserves phenology/seasonality patterns.</li>\n<li><strong>+ Interactions (nitrogen×texture, nitrogen/CO₂):</strong> Small but consistent.</li>\n<li><strong>– Removing raw 240-day columns:</strong> Neutral→positive (less overfit, lower memory).</li>\n<li><strong>± Heavier models (XGBoost/CatBoost/NNs):</strong> No consistent win vs LGBM after robust FE; training slower.</li>\n<li><strong>± Extensive hyper-tuning:</strong> marginal improvements (&lt; LB jitter) once core FE is in place.</li>\n</ul>\n<hr>\n<h2>Ethics &amp; Responsible Use</h2>\n<p>This solution emulates simulated yields under a specific high-emissions scenario and rainfed assumptions. Real-world decisions should consider local agronomy, uncertainty, and the gap between simulated and observed systems (irrigation, policy, technology change).</p>\n<hr>\n<h2>Acknowledgments</h2>\n<p>Special thanks to the AgML / AgMIP team for providing the FutureCrop benchmark and fostering open agricultural AI research. Gratitude to the broader community and contributors whose datasets, feedback, and discussions helped refine this work.</p>",
      "rawMarkdown": "## TL;DR\n* **Task:** Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.\n* **Data:** Train 1982–2020 → Test 2021–2098 (high-emissions, rainfed).\n* **Core idea:** Summarize 240 daily features into global and 8×30-day chunk stats, plus train-only per-location yield priors (lat, lon).\n* **Model:** LightGBM (GBDT) — ~100k trees, lr≈0.01, early stopping≈50.\n* **Performance:** sMAPE ≈ 14.7 %, R² ≈ 0.96, MAE ≈ 0.31, RMSE ≈ 0.49 (train); RMSE ≈ 1.07 (public validation, 2nd place).\n* **Why it works:** Captures climate patterns, seasonality, and site productivity while avoiding overfit and enabling fast and robust training.\n\n---\n\n# Problem & Data\n* **Task:** Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.\n* **Time horizon:** Forecast **2021–2098** using models trained on **1982–2020** (train : test ≈ 1 : 2).\n* **Scenario:** High-emissions (rising CO₂), rainfed conditions with time-invariant soils and fixed per-country nitrogen rates.\n* **Inputs (per datapoint):**\n  * **240 daily climate variables:**\n      * Precipitation (PR)\n      * Shortwave radiation (RSDS)\n      * Mean, minimum, and maximum temperature (TAS, TMIN, TMAX)\n  * **Auxiliaries:**\n      * **Soil texture class** (USDA 1–13)\n      * **Nitrogen rate** (country-level, fixed)\n      * **CO₂ concentration** (spatially constant, rising over time)\n      * **Planting date** and real_year\n* **Evaluation metrics:**\n  * **Primary:** RMSE on 2021–2098 (private leaderboard)\n  * **Secondary:** Median per-cell R² and regional production R² (Iowa maize, Germany wheat).\n\n---\n\n## Design Principles\n\n1. **Respect the shift:** Future climate and CO₂ drift ⇒ prefer features that remain meaningful under shift (nonparametric summaries, chunk seasonality) over high-capacity sequence models that can overfit training climatology.\n2. **Keep it lean:** Prioritize fast iteration and low memory; extensive aggregation > raw 240-day sequences.\n3. **Add geographic priors (but no leakage):** Per-location mean yields computed only on training add ststrong signal without touching future labels.\n4. **Determinism:** Fixed seeds, version-locked libs, explicit preprocessing.\n\n---\n\n## Feature Engineering\n\n* **Daily → Global Summaries (per variable, per sample):** `mean`, `median`, `sum`, `min`, `max` for each of: `pr`, `rsds`, `tas` (and analogously for `tmin`, `tmax` when helpful).\n* **8×30-day “Chunk” Statistics:**\n  * For each variable, split the 240 days into 8 contiguous 30-day windows and compute the same statistics per chunk.\n  * Captures intra-season timing (e.g., early-season heat, mid-season radiation, late-season water stress) without high-dimensional sequences.\n* **Per-Location Yield Priors (Train-only):**\n  * Compute mean yield at each `(latitude, longitude)` separately for maize and wheat using only training years; join these two features into both crops at inference.\n  * Intuition: encodes site-level productivity (soils/management beyond observable features) **without** using future labels.\n* **Simple Interactions & Calendar:**\n  * Mild interactions such as `nitrogen × texture_class`, `nitrogen / CO₂`, and a coarse decade feature from `real_year`.\n  * Avoid heavy cross terms to keep modemodel simple and stable across shifts.\n* **Pruning:** After creating summaries, drop the raw 240-day columns to reduce memory and prevent tree splits on noisy day-level artifacts.\n\n---\n\n## Model\n\n* **Algorithm:** LightGBM (GBDT)\n* **Key params:**\n  * `objective=regression`\n  * `metric=rmse`\n  * `learning_rate=0.01`\n  * `n_estimators=100000` with early stopping ≈ 50 rounds\n  * Default depth/leaf settings performed well; minor tuning had a marginal impact\n* **Why LightGBM:** strong tabular baseline, fast, robust to mixed scales, and tolerant of many moderately informative features.\n\n---\n\n## Validation & Leakage Guards\n\n* **Primary split** for model selection: **random 80/20** (to iterate quickly on FE/model).\n* **Temporal sanity checks:** additional **forward holdouts** (e.g., last 5 training years) to verify decade-to-decade stability.\n* **Leakage checks:**\n    * Per-location means computed **only from training** and merged by `(latitude, longitude)`; **no** target leakage from test.\n    * Avoid using the internal `year` simulation index as a proxy for time on test.\n    * Confirm units and scales (e.g., `pr` in kg m⁻² s⁻¹).\n* **Seeds:** fixed (`random_state=42`) for reproducibility.\n\n---\n\n## Results Summary\n\n| **Metric** | **Training (1982–2020)** | **Public Validation (2021–2050, Kaggle LB)** | **Remarks**|\n| ----------------------------- | ------------- | ------------------------------ | --------------------------- |\n| **RMSE**| **≈ 0.49**   | **≈ 1.07 (2nd place)**  | Primary leaderboard metric   |\n| **MAE** | ≈ 0.31  | —  | Low absolute deviation across grids   |\n| **MAPE**   | ≈ 14.7 % | — | Stable relative accuracy across crops  |\n| **R²**| ≈ 0.96  | — | Strong fit and generalization on past decades |\n\n\n\n\nComprehensive evaluation results will be published by the AgML competition team following the official review.\n\n---\n\n## Ablations (What Moved the Needle)\n\n* **+ Per-location mean yields (2 features):** Large gain; strong prior for site productivity.\n* **+ 8×30-day chunks:** Medium gain; preserves phenology/seasonality patterns.\n* **+ Interactions (nitrogen×texture, nitrogen/CO₂):** Small but consistent.\n* **– Removing raw 240-day columns:** Neutral→positive (less overfit, lower memory).\n* **± Heavier models (XGBoost/CatBoost/NNs):** No consistent win vs LGBM after robust FE; training slower.\n* **± Extensive hyper-tuning:** marginal improvements (< LB jitter) once core FE is in place.\n\n---\n\n## Ethics & Responsible Use\n\nThis solution emulates simulated yields under a specific high-emissions scenario and rainfed assumptions. Real-world decisions should consider local agronomy, uncertainty, and the gap between simulated and observed systems (irrigation, policy, technology change).\n\n---\n\n## Acknowledgments\n\nSpecial thanks to the AgML / AgMIP team for providing the FutureCrop benchmark and fostering open agricultural AI research. Gratitude to the broader community and contributors whose datasets, feedback, and discussions helped refine this work.",
      "votes": null
    },
    {
      "id": "3312213",
      "postDate": "11/06/2025 17:46:24",
      "content": "<p>Great write-up, thanks a lot! And congrats on the amazing work!</p>",
      "rawMarkdown": "Great write-up, thanks a lot! And congrats on the amazing work!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3312213,
      "author_name": "lilybellesweet",
      "author_url": "",
      "post_date": "11/06/2025 17:46:24",
      "content": "<p>Great write-up, thanks a lot! And congrats on the amazing work!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3312083": "## TL;DR\n* **Task:** Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.\n* **Data:** Train 1982–2020 → Test 2021–2098 (high-emissions, rainfed).\n* **Core idea:** Summarize 240 daily features into global and 8×30-day chunk stats, plus train-only per-location yield priors (lat, lon).\n* **Model:** LightGBM (GBDT) — ~100k trees, lr≈0.01, early stopping≈50.\n* **Performance:** sMAPE ≈ 14.7 %, R² ≈ 0.96, MAE ≈ 0.31, RMSE ≈ 0.49 (train); RMSE ≈ 1.07 (public validation, 2nd place).\n* **Why it works:** Captures climate patterns, seasonality, and site productivity while avoiding overfit and enabling fast and robust training.\n\n---\n\n# Problem & Data\n* **Task:** Predict end-of-season yield (t ha⁻¹) for maize and wheat at each 0.5° grid cell × year.\n* **Time horizon:** Forecast **2021–2098** using models trained on **1982–2020** (train : test ≈ 1 : 2).\n* **Scenario:** High-emissions (rising CO₂), rainfed conditions with time-invariant soils and fixed per-country nitrogen rates.\n* **Inputs (per datapoint):**\n  * **240 daily climate variables:**\n      * Precipitation (PR)\n      * Shortwave radiation (RSDS)\n      * Mean, minimum, and maximum temperature (TAS, TMIN, TMAX)\n  * **Auxiliaries:**\n      * **Soil texture class** (USDA 1–13)\n      * **Nitrogen rate** (country-level, fixed)\n      * **CO₂ concentration** (spatially constant, rising over time)\n      * **Planting date** and real_year\n* **Evaluation metrics:**\n  * **Primary:** RMSE on 2021–2098 (private leaderboard)\n  * **Secondary:** Median per-cell R² and regional production R² (Iowa maize, Germany wheat).\n\n---\n\n## Design Principles\n\n1. **Respect the shift:** Future climate and CO₂ drift ⇒ prefer features that remain meaningful under shift (nonparametric summaries, chunk seasonality) over high-capacity sequence models that can overfit training climatology.\n2. **Keep it lean:** Prioritize fast iteration and low memory; extensive aggregation > raw 240-day sequences.\n3. **Add geographic priors (but no leakage):** Per-location mean yields computed only on training add ststrong signal without touching future labels.\n4. **Determinism:** Fixed seeds, version-locked libs, explicit preprocessing.\n\n---\n\n## Feature Engineering\n\n* **Daily → Global Summaries (per variable, per sample):** `mean`, `median`, `sum`, `min`, `max` for each of: `pr`, `rsds`, `tas` (and analogously for `tmin`, `tmax` when helpful).\n* **8×30-day “Chunk” Statistics:**\n  * For each variable, split the 240 days into 8 contiguous 30-day windows and compute the same statistics per chunk.\n  * Captures intra-season timing (e.g., early-season heat, mid-season radiation, late-season water stress) without high-dimensional sequences.\n* **Per-Location Yield Priors (Train-only):**\n  * Compute mean yield at each `(latitude, longitude)` separately for maize and wheat using only training years; join these two features into both crops at inference.\n  * Intuition: encodes site-level productivity (soils/management beyond observable features) **without** using future labels.\n* **Simple Interactions & Calendar:**\n  * Mild interactions such as `nitrogen × texture_class`, `nitrogen / CO₂`, and a coarse decade feature from `real_year`.\n  * Avoid heavy cross terms to keep modemodel simple and stable across shifts.\n* **Pruning:** After creating summaries, drop the raw 240-day columns to reduce memory and prevent tree splits on noisy day-level artifacts.\n\n---\n\n## Model\n\n* **Algorithm:** LightGBM (GBDT)\n* **Key params:**\n  * `objective=regression`\n  * `metric=rmse`\n  * `learning_rate=0.01`\n  * `n_estimators=100000` with early stopping ≈ 50 rounds\n  * Default depth/leaf settings performed well; minor tuning had a marginal impact\n* **Why LightGBM:** strong tabular baseline, fast, robust to mixed scales, and tolerant of many moderately informative features.\n\n---\n\n## Validation & Leakage Guards\n\n* **Primary split** for model selection: **random 80/20** (to iterate quickly on FE/model).\n* **Temporal sanity checks:** additional **forward holdouts** (e.g., last 5 training years) to verify decade-to-decade stability.\n* **Leakage checks:**\n    * Per-location means computed **only from training** and merged by `(latitude, longitude)`; **no** target leakage from test.\n    * Avoid using the internal `year` simulation index as a proxy for time on test.\n    * Confirm units and scales (e.g., `pr` in kg m⁻² s⁻¹).\n* **Seeds:** fixed (`random_state=42`) for reproducibility.\n\n---\n\n## Results Summary\n\n| **Metric** | **Training (1982–2020)** | **Public Validation (2021–2050, Kaggle LB)** | **Remarks**|\n| ----------------------------- | ------------- | ------------------------------ | --------------------------- |\n| **RMSE**| **≈ 0.49**   | **≈ 1.07 (2nd place)**  | Primary leaderboard metric   |\n| **MAE** | ≈ 0.31  | —  | Low absolute deviation across grids   |\n| **MAPE**   | ≈ 14.7 % | — | Stable relative accuracy across crops  |\n| **R²**| ≈ 0.96  | — | Strong fit and generalization on past decades |\n\n\n\n\nComprehensive evaluation results will be published by the AgML competition team following the official review.\n\n---\n\n## Ablations (What Moved the Needle)\n\n* **+ Per-location mean yields (2 features):** Large gain; strong prior for site productivity.\n* **+ 8×30-day chunks:** Medium gain; preserves phenology/seasonality patterns.\n* **+ Interactions (nitrogen×texture, nitrogen/CO₂):** Small but consistent.\n* **– Removing raw 240-day columns:** Neutral→positive (less overfit, lower memory).\n* **± Heavier models (XGBoost/CatBoost/NNs):** No consistent win vs LGBM after robust FE; training slower.\n* **± Extensive hyper-tuning:** marginal improvements (< LB jitter) once core FE is in place.\n\n---\n\n## Ethics & Responsible Use\n\nThis solution emulates simulated yields under a specific high-emissions scenario and rainfed assumptions. Real-world decisions should consider local agronomy, uncertainty, and the gap between simulated and observed systems (irrigation, policy, technology change).\n\n---\n\n## Acknowledgments\n\nSpecial thanks to the AgML / AgMIP team for providing the FutureCrop benchmark and fostering open agricultural AI research. Gratitude to the broader community and contributors whose datasets, feedback, and discussions helped refine this work.",
    "3312213": "Great write-up, thanks a lot! And congrats on the amazing work!"
  },
  "source": "meta"
}