{
  "id": 743921,
  "title": "Plot identity is recoverable from GrainWeight, and leaf-random CV inflates accuracy by ~0.14",
  "url": "/competitions/HyperLeaf2024/discussion/743921",
  "author_name": "Resham Raj Shivwanshi",
  "post_date": "2026-09-27T18:51:20.293000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Leaves in HyperLeaf2024 are nested inside field plots, and each plot carries exactly one cultivar. If your validation split is random over leaves, leaves from the same plot land on both sides and the model can learn plot, not genotype.</p>\n<p>There is no plot column in the release, but you do not need one. <strong>GrainWeight is measured once per plot</strong>, so grouping by it gives 36 plots with purity 1.000 (9 per cultivar, 3 per cultivar x fertilizer cell). The north-western block is 4 cultivars x 3 fertilizer x 3 density, so every plot is a unique design triple.</p>\n<p>What we measured on all 2,410 labelled leaves:</p>\n<ul>\n<li>Across 15 model x feature combinations, leaf-random CV was higher than plot-grouped CV every time, by 0.105 to 0.195 (median 0.142).</li>\n<li>Best leaf-random SVM on per-band quantiles: 0.918. Same model, plot-grouped: 0.789.</li>\n<li>A permutation null that shuffles cultivar across whole plots gives 0.239 against an observed 0.618 (P = 0.005), so the genotype signal is real. It is just smaller than leaf-level scores suggest.</li>\n<li>Intra-plot correlation is high enough (design effect about 21) that the effective sample size is closer to 113 than to 2,410.</li>\n</ul>\n<p>Practical suggestions: use <code>GroupKFold</code> with GrainWeight as the group, and mask the zero background before averaging bands (unmasked means are biased by roughly two thirds). Per-band quantiles beat the mean spectrum in our runs.</p>\n<p>Happy to hear if anyone has found a cleaner grouping key.</p>",
  "messages": [
    {
      "id": 3529260,
      "postDate": "2026-09-27T18:51:20.293Z",
      "content": "<p>Leaves in HyperLeaf2024 are nested inside field plots, and each plot carries exactly one cultivar. If your validation split is random over leaves, leaves from the same plot land on both sides and the model can learn plot, not genotype.</p>\n<p>There is no plot column in the release, but you do not need one. <strong>GrainWeight is measured once per plot</strong>, so grouping by it gives 36 plots with purity 1.000 (9 per cultivar, 3 per cultivar x fertilizer cell). The north-western block is 4 cultivars x 3 fertilizer x 3 density, so every plot is a unique design triple.</p>\n<p>What we measured on all 2,410 labelled leaves:</p>\n<ul>\n<li>Across 15 model x feature combinations, leaf-random CV was higher than plot-grouped CV every time, by 0.105 to 0.195 (median 0.142).</li>\n<li>Best leaf-random SVM on per-band quantiles: 0.918. Same model, plot-grouped: 0.789.</li>\n<li>A permutation null that shuffles cultivar across whole plots gives 0.239 against an observed 0.618 (P = 0.005), so the genotype signal is real. It is just smaller than leaf-level scores suggest.</li>\n<li>Intra-plot correlation is high enough (design effect about 21) that the effective sample size is closer to 113 than to 2,410.</li>\n</ul>\n<p>Practical suggestions: use <code>GroupKFold</code> with GrainWeight as the group, and mask the zero background before averaging bands (unmasked means are biased by roughly two thirds). Per-band quantiles beat the mean spectrum in our runs.</p>\n<p>Happy to hear if anyone has found a cleaner grouping key.</p>",
      "rawMarkdown": "Leaves in HyperLeaf2024 are nested inside field plots, and each plot carries exactly one cultivar. If your validation split is random over leaves, leaves from the same plot land on both sides and the model can learn plot, not genotype.\n\nThere is no plot column in the release, but you do not need one. **GrainWeight is measured once per plot**, so grouping by it gives 36 plots with purity 1.000 (9 per cultivar, 3 per cultivar x fertilizer cell). The north-western block is 4 cultivars x 3 fertilizer x 3 density, so every plot is a unique design triple.\n\nWhat we measured on all 2,410 labelled leaves:\n\n- Across 15 model x feature combinations, leaf-random CV was higher than plot-grouped CV every time, by 0.105 to 0.195 (median 0.142).\n- Best leaf-random SVM on per-band quantiles: 0.918. Same model, plot-grouped: 0.789.\n- A permutation null that shuffles cultivar across whole plots gives 0.239 against an observed 0.618 (P = 0.005), so the genotype signal is real. It is just smaller than leaf-level scores suggest.\n- Intra-plot correlation is high enough (design effect about 21) that the effective sample size is closer to 113 than to 2,410.\n\nPractical suggestions: use `GroupKFold` with GrainWeight as the group, and mask the zero background before averaging bands (unmasked means are biased by roughly two thirds). Per-band quantiles beat the mean spectrum in our runs.\n\nHappy to hear if anyone has found a cleaner grouping key."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3529260": "Leaves in HyperLeaf2024 are nested inside field plots, and each plot carries exactly one cultivar. If your validation split is random over leaves, leaves from the same plot land on both sides and the model can learn plot, not genotype.\n\nThere is no plot column in the release, but you do not need one. **GrainWeight is measured once per plot**, so grouping by it gives 36 plots with purity 1.000 (9 per cultivar, 3 per cultivar x fertilizer cell). The north-western block is 4 cultivars x 3 fertilizer x 3 density, so every plot is a unique design triple.\n\nWhat we measured on all 2,410 labelled leaves:\n\n- Across 15 model x feature combinations, leaf-random CV was higher than plot-grouped CV every time, by 0.105 to 0.195 (median 0.142).\n- Best leaf-random SVM on per-band quantiles: 0.918. Same model, plot-grouped: 0.789.\n- A permutation null that shuffles cultivar across whole plots gives 0.239 against an observed 0.618 (P = 0.005), so the genotype signal is real. It is just smaller than leaf-level scores suggest.\n- Intra-plot correlation is high enough (design effect about 21) that the effective sample size is closer to 113 than to 2,410.\n\nPractical suggestions: use `GroupKFold` with GrainWeight as the group, and mask the zero background before averaging bands (unmasked means are biased by roughly two thirds). Per-band quantiles beat the mean spectrum in our runs.\n\nHappy to hear if anyone has found a cleaner grouping key."
  }
}