{
  "id": 244653,
  "title": "Observations about the data and suggestions for beating my score 🌱️🤔️",
  "url": "/competitions/sorghum-biomass-prediction/discussion/244653",
  "author_name": "",
  "post_date": "2021-06-07T18:07:44.391991400Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was excited when I saw this competition at CVPPA, because I'm a huge fan of TERRA-REF. I used to go to the Phenome conferences in AZ and one of my big regrets is that I never went out to see the facility while I was out there. So when I saw this competition I thought I would do my part and try to help folks out with this.</p>\n<p>I've burned about two weeks of combined GPU time searching over models, and I should probably get back to research now, so here's what I've learned!</p>\n<h3>The signal in the data is quite small.</h3>\n<p>There's a lot of what ML people would think of as label noise - and I hypothesize that it varies with maturity. For example, here are two late-season images, both at 140 DAP, but one is the extreme high end of biomass and the other is at the extreme low end.</p>\n<p><img src=\"https://i.imgur.com/88r3JxJ.png\" alt=\"\"><br>\n<img src=\"https://i.imgur.com/2ZTf2ID.png\" alt=\"\"></p>\n<p>Can you tell which is which? Me neither m8. <em>This is why I think that most of the signal is in the early-season data - specifically in early season vigour and growth rate</em>.</p>\n<h3>About my baseline.</h3>\n<p>For the baseline I submitted, I'm just extracting image features with a pretrained resnet50 and then concatenating on the metadata, shoving it through a simple two-layer MLP, and training with SGD. That's it. No fancy optimizers, no state-of-the-art vision models, nothing. I didn't even tune any hyperparameters.</p>\n<p>Comparing the distribution of predictions against the distribution of labels in the training set (which I assume should be <em>i.i.d</em>), the model doesn't predict a lot of the variance. I'm too lazy to write the code to break it down by maturity, but I'm pretty sure there will be more variance for earlier timepoints.</p>\n<p><img src=\"https://i.imgur.com/0RgOmxm.png\" alt=\"\"></p>\n<h3>The model capacity doesn't really matter.</h3>\n<blockquote>\n  <p>It don't matter. None of this matters. <br>\n  - Carl Brutananadilewski</p>\n</blockquote>\n<p>I trained a few iterations of my baseline model and I can tell you that basically none of it matters. You can use a resnet50 or a resnet18, or probably something really tiny. You can fine-tune a pretrained model, train one from scratch, or just freeze the whole thing and use it as a feature extractor. It all yields basically the same validation error. This is probably because they are all capable of extracting vegetation area, which is where I think almost all of the signal comes from. </p>\n<p>I've also tried training with and without the heat units data, and it doesn't change validation error either.</p>\n<h3>How to beat my submission.</h3>\n<p>As I've alluded to, I'm pretty sure that the key to recovering more signal is to model the timeseries, which will allow you to get early-season growth rate. The data does include timeseries of plots over a growing season, which is great! However, the times are varied and, in the test set, the test points are divided into discrete early, mid, and late stages only. </p>\n<p>And of course, as with all Kaggle competitions, the other secret is to ride that submission button like it owes you money. The <a href=\"https://lukeoakdenrayner.wordpress.com/2019/09/19/ai-competitions-dont-produce-useful-models/\" target=\"_blank\">more times you sample the same model</a> with a different seed, the better your best submission will probably be (unless the data held out from the public leaderboard is from a different domain).</p>\n<p>Anyway, thanks to the organizers for putting this together 🤗️.  I look forward to seeing the final models, and I hope that this has been useful to people who are thinking of participating.</p>",
  "messages": [
    {
      "id": "1340220",
      "postDate": "06/07/2021 18:07:44",
      "content": "<p>I was excited when I saw this competition at CVPPA, because I'm a huge fan of TERRA-REF. I used to go to the Phenome conferences in AZ and one of my big regrets is that I never went out to see the facility while I was out there. So when I saw this competition I thought I would do my part and try to help folks out with this.</p>\n<p>I've burned about two weeks of combined GPU time searching over models, and I should probably get back to research now, so here's what I've learned!</p>\n<h3>The signal in the data is quite small.</h3>\n<p>There's a lot of what ML people would think of as label noise - and I hypothesize that it varies with maturity. For example, here are two late-season images, both at 140 DAP, but one is the extreme high end of biomass and the other is at the extreme low end.</p>\n<p><img src=\"https://i.imgur.com/88r3JxJ.png\" alt=\"\"><br>\n<img src=\"https://i.imgur.com/2ZTf2ID.png\" alt=\"\"></p>\n<p>Can you tell which is which? Me neither m8. <em>This is why I think that most of the signal is in the early-season data - specifically in early season vigour and growth rate</em>.</p>\n<h3>About my baseline.</h3>\n<p>For the baseline I submitted, I'm just extracting image features with a pretrained resnet50 and then concatenating on the metadata, shoving it through a simple two-layer MLP, and training with SGD. That's it. No fancy optimizers, no state-of-the-art vision models, nothing. I didn't even tune any hyperparameters.</p>\n<p>Comparing the distribution of predictions against the distribution of labels in the training set (which I assume should be <em>i.i.d</em>), the model doesn't predict a lot of the variance. I'm too lazy to write the code to break it down by maturity, but I'm pretty sure there will be more variance for earlier timepoints.</p>\n<p><img src=\"https://i.imgur.com/0RgOmxm.png\" alt=\"\"></p>\n<h3>The model capacity doesn't really matter.</h3>\n<blockquote>\n  <p>It don't matter. None of this matters. <br>\n  - Carl Brutananadilewski</p>\n</blockquote>\n<p>I trained a few iterations of my baseline model and I can tell you that basically none of it matters. You can use a resnet50 or a resnet18, or probably something really tiny. You can fine-tune a pretrained model, train one from scratch, or just freeze the whole thing and use it as a feature extractor. It all yields basically the same validation error. This is probably because they are all capable of extracting vegetation area, which is where I think almost all of the signal comes from. </p>\n<p>I've also tried training with and without the heat units data, and it doesn't change validation error either.</p>\n<h3>How to beat my submission.</h3>\n<p>As I've alluded to, I'm pretty sure that the key to recovering more signal is to model the timeseries, which will allow you to get early-season growth rate. The data does include timeseries of plots over a growing season, which is great! However, the times are varied and, in the test set, the test points are divided into discrete early, mid, and late stages only. </p>\n<p>And of course, as with all Kaggle competitions, the other secret is to ride that submission button like it owes you money. The <a href=\"https://lukeoakdenrayner.wordpress.com/2019/09/19/ai-competitions-dont-produce-useful-models/\" target=\"_blank\">more times you sample the same model</a> with a different seed, the better your best submission will probably be (unless the data held out from the public leaderboard is from a different domain).</p>\n<p>Anyway, thanks to the organizers for putting this together 🤗️.  I look forward to seeing the final models, and I hope that this has been useful to people who are thinking of participating.</p>",
      "rawMarkdown": "I was excited when I saw this competition at CVPPA, because I'm a huge fan of TERRA-REF. I used to go to the Phenome conferences in AZ and one of my big regrets is that I never went out to see the facility while I was out there. So when I saw this competition I thought I would do my part and try to help folks out with this.\n\nI've burned about two weeks of combined GPU time searching over models, and I should probably get back to research now, so here's what I've learned!\n\n### The signal in the data is quite small.\n\nThere's a lot of what ML people would think of as label noise - and I hypothesize that it varies with maturity. For example, here are two late-season images, both at 140 DAP, but one is the extreme high end of biomass and the other is at the extreme low end.\n\n![](https://i.imgur.com/88r3JxJ.png)\n![](https://i.imgur.com/2ZTf2ID.png)\n\nCan you tell which is which? Me neither m8. *This is why I think that most of the signal is in the early-season data - specifically in early season vigour and growth rate*.\n\n### About my baseline. \n\nFor the baseline I submitted, I'm just extracting image features with a pretrained resnet50 and then concatenating on the metadata, shoving it through a simple two-layer MLP, and training with SGD. That's it. No fancy optimizers, no state-of-the-art vision models, nothing. I didn't even tune any hyperparameters.\n\nComparing the distribution of predictions against the distribution of labels in the training set (which I assume should be *i.i.d*), the model doesn't predict a lot of the variance. I'm too lazy to write the code to break it down by maturity, but I'm pretty sure there will be more variance for earlier timepoints.\n\n![](https://i.imgur.com/0RgOmxm.png)\n\n### The model capacity doesn't really matter.\n\n> It don't matter. None of this matters. \n> \\- Carl Brutananadilewski\n\nI trained a few iterations of my baseline model and I can tell you that basically none of it matters. You can use a resnet50 or a resnet18, or probably something really tiny. You can fine-tune a pretrained model, train one from scratch, or just freeze the whole thing and use it as a feature extractor. It all yields basically the same validation error. This is probably because they are all capable of extracting vegetation area, which is where I think almost all of the signal comes from. \n\nI've also tried training with and without the heat units data, and it doesn't change validation error either.\n\n### How to beat my submission.\n\nAs I've alluded to, I'm pretty sure that the key to recovering more signal is to model the timeseries, which will allow you to get early-season growth rate. The data does include timeseries of plots over a growing season, which is great! However, the times are varied and, in the test set, the test points are divided into discrete early, mid, and late stages only. \n\nAnd of course, as with all Kaggle competitions, the other secret is to ride that submission button like it owes you money. The [more times you sample the same model](https://lukeoakdenrayner.wordpress.com/2019/09/19/ai-competitions-dont-produce-useful-models/) with a different seed, the better your best submission will probably be (unless the data held out from the public leaderboard is from a different domain).\n\nAnyway, thanks to the organizers for putting this together 🤗️.  I look forward to seeing the final models, and I hope that this has been useful to people who are thinking of participating.",
      "votes": null
    },
    {
      "id": "1366392",
      "postDate": "06/26/2021 17:57:27",
      "content": "<p>Hi Jordan this is great work. I’d second the idea that using date would be essential. I am excited about the competition but won’t enter myself (since I’m on the TERRA-REF team). But if I did I would create a plant / soil mask and calculate %cover. Then can fit a logistic growth curve from which the max rate of increase could be a good predictor.</p>\n<p>So I started running the masking if that would be of help … it is taking a few days to process. And that made me realize that any analysis must be very computationally intensive!</p>\n<p>I’d be happy to share masks and % cover with anyone interested when the processing is done.</p>\n<p>I’d also be happy to help folks get computing time - would first suggest they get an account at cyverse.org and search the discovery environment for the vice app of their choice. I can put the images, masks, and %cover data in a shared folder and otherwise help.</p>",
      "rawMarkdown": "Hi Jordan this is great work. I’d second the idea that using date would be essential. I am excited about the competition but won’t enter myself (since I’m on the TERRA-REF team). But if I did I would create a plant / soil mask and calculate %cover. Then can fit a logistic growth curve from which the max rate of increase could be a good predictor.\n\nSo I started running the masking if that would be of help … it is taking a few days to process. And that made me realize that any analysis must be very computationally intensive!\n\nI’d be happy to share masks and % cover with anyone interested when the processing is done.\n\nI’d also be happy to help folks get computing time - would first suggest they get an account at cyverse.org and search the discovery environment for the vice app of their choice. I can put the images, masks, and %cover data in a shared folder and otherwise help.",
      "votes": null
    },
    {
      "id": "1366540",
      "postDate": "06/26/2021 22:27:28",
      "content": "<p>Hi David,</p>\n<p>Yes that was almost exactly what I thought of doing! I gave the vegetation masks a try with excess green but couldn't find a reliable threshold and I ran out of time.  I'm sure those masks would be immensely helpful to participants. :)</p>",
      "rawMarkdown": "Hi David,\n\nYes that was almost exactly what I thought of doing! I gave the vegetation masks a try with excess green but couldn't find a reliable threshold and I ran out of time.  I'm sure those masks would be immensely helpful to participants. :)",
      "votes": null
    },
    {
      "id": "1368545",
      "postDate": "06/28/2021 17:04:56",
      "content": "<p>Hi Jordan … a lot of the images are overexposed, so our algorithm skips those. Without doing a lot of qa/qc or cleaning up. Code is in <a href=\"https://github.com/az-digitalag/sorghum_biomass_prediction\" target=\"_blank\">github.com/az-digitalag/sorghum_biomass_prediction</a> and there is a list of files and their calculated canopy cover values in <a href=\"https://github.com/az-digitalag/sorghum_biomass_prediction/blob/main/output/canopy_cover.txt?raw=true\" target=\"_blank\">output/canopy_cover.txt</a></p>\n<p>The canopy cover value goes from 0 = no plants to 1 = all plants / no soil. The file is structured with the file name followed by the canopy cover value. Sometimes it appears there are two file names in a row (in which case I suspect the first file threw an error but haven't dug in). </p>\n<p>I just wanted to create masks since I figured that would be the most difficult part and find interested parties to pick up the ball. I'll create a separate post as soon as I can with a link to both the masks and these data once I tar up the masks and put them somewhere.</p>",
      "rawMarkdown": "Hi Jordan ... a lot of the images are overexposed, so our algorithm skips those. Without doing a lot of qa/qc or cleaning up. Code is in [github.com/az-digitalag/sorghum_biomass_prediction](https://github.com/az-digitalag/sorghum_biomass_prediction) and there is a list of files and their calculated canopy cover values in [output/canopy_cover.txt](https://github.com/az-digitalag/sorghum_biomass_prediction/blob/main/output/canopy_cover.txt?raw=true)\n\nThe canopy cover value goes from 0 = no plants to 1 = all plants / no soil. The file is structured with the file name followed by the canopy cover value. Sometimes it appears there are two file names in a row (in which case I suspect the first file threw an error but haven't dug in). \n\nI just wanted to create masks since I figured that would be the most difficult part and find interested parties to pick up the ball. I'll create a separate post as soon as I can with a link to both the masks and these data once I tar up the masks and put them somewhere.",
      "votes": null
    },
    {
      "id": "1384905",
      "postDate": "07/12/2021 10:09:10",
      "content": "<p>Thanks for these tips - I my approach was largely based on this. I found that my models quickly overfit the training set.<br>\nI'd be interested to know if other competitors had private LB scores better than mine from their earlier submissions, and whether anything more sophisticated worked.</p>\n<p>I'll try to write up a brief summary of what I did later.</p>",
      "rawMarkdown": "Thanks for these tips - I my approach was largely based on this. I found that my models quickly overfit the training set.\nI'd be interested to know if other competitors had private LB scores better than mine from their earlier submissions, and whether anything more sophisticated worked.\n\nI'll try to write up a brief summary of what I did later.",
      "votes": null
    },
    {
      "id": "1405446",
      "postDate": "07/30/2021 21:04:52",
      "content": "<p>Hi everyone. I just realize about this interesting competition and want to contribute to the discussion here. you may find our work at UIUC useful <a href=\"https://www.mdpi.com/2072-4292/13/9/1763\" target=\"_blank\">https://www.mdpi.com/2072-4292/13/9/1763</a>  :) </p>",
      "rawMarkdown": "Hi everyone. I just realize about this interesting competition and want to contribute to the discussion here. you may find our work at UIUC useful https://www.mdpi.com/2072-4292/13/9/1763  :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1366392,
      "author_name": "dlebauer",
      "author_url": "",
      "post_date": "06/26/2021 17:57:27",
      "content": "<p>Hi Jordan this is great work. I’d second the idea that using date would be essential. I am excited about the competition but won’t enter myself (since I’m on the TERRA-REF team). But if I did I would create a plant / soil mask and calculate %cover. Then can fit a logistic growth curve from which the max rate of increase could be a good predictor.</p>\n<p>So I started running the masking if that would be of help … it is taking a few days to process. And that made me realize that any analysis must be very computationally intensive!</p>\n<p>I’d be happy to share masks and % cover with anyone interested when the processing is done.</p>\n<p>I’d also be happy to help folks get computing time - would first suggest they get an account at cyverse.org and search the discovery environment for the vice app of their choice. I can put the images, masks, and %cover data in a shared folder and otherwise help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1366540,
          "author_name": "jubbens",
          "author_url": "",
          "post_date": "06/26/2021 22:27:28",
          "content": "<p>Hi David,</p>\n<p>Yes that was almost exactly what I thought of doing! I gave the vegetation masks a try with excess green but couldn't find a reliable threshold and I ran out of time.  I'm sure those masks would be immensely helpful to participants. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1368545,
          "author_name": "dlebauer",
          "author_url": "",
          "post_date": "06/28/2021 17:04:56",
          "content": "<p>Hi Jordan … a lot of the images are overexposed, so our algorithm skips those. Without doing a lot of qa/qc or cleaning up. Code is in <a href=\"https://github.com/az-digitalag/sorghum_biomass_prediction\" target=\"_blank\">github.com/az-digitalag/sorghum_biomass_prediction</a> and there is a list of files and their calculated canopy cover values in <a href=\"https://github.com/az-digitalag/sorghum_biomass_prediction/blob/main/output/canopy_cover.txt?raw=true\" target=\"_blank\">output/canopy_cover.txt</a></p>\n<p>The canopy cover value goes from 0 = no plants to 1 = all plants / no soil. The file is structured with the file name followed by the canopy cover value. Sometimes it appears there are two file names in a row (in which case I suspect the first file threw an error but haven't dug in). </p>\n<p>I just wanted to create masks since I figured that would be the most difficult part and find interested parties to pick up the ball. I'll create a separate post as soon as I can with a link to both the masks and these data once I tar up the masks and put them somewhere.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1384905,
      "author_name": "johnfozard",
      "author_url": "",
      "post_date": "07/12/2021 10:09:10",
      "content": "<p>Thanks for these tips - I my approach was largely based on this. I found that my models quickly overfit the training set.<br>\nI'd be interested to know if other competitors had private LB scores better than mine from their earlier submissions, and whether anything more sophisticated worked.</p>\n<p>I'll try to write up a brief summary of what I did later.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1405446,
      "author_name": "pixelvar",
      "author_url": "",
      "post_date": "07/30/2021 21:04:52",
      "content": "<p>Hi everyone. I just realize about this interesting competition and want to contribute to the discussion here. you may find our work at UIUC useful <a href=\"https://www.mdpi.com/2072-4292/13/9/1763\" target=\"_blank\">https://www.mdpi.com/2072-4292/13/9/1763</a>  :) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1340220": "I was excited when I saw this competition at CVPPA, because I'm a huge fan of TERRA-REF. I used to go to the Phenome conferences in AZ and one of my big regrets is that I never went out to see the facility while I was out there. So when I saw this competition I thought I would do my part and try to help folks out with this.\n\nI've burned about two weeks of combined GPU time searching over models, and I should probably get back to research now, so here's what I've learned!\n\n### The signal in the data is quite small.\n\nThere's a lot of what ML people would think of as label noise - and I hypothesize that it varies with maturity. For example, here are two late-season images, both at 140 DAP, but one is the extreme high end of biomass and the other is at the extreme low end.\n\n![](https://i.imgur.com/88r3JxJ.png)\n![](https://i.imgur.com/2ZTf2ID.png)\n\nCan you tell which is which? Me neither m8. *This is why I think that most of the signal is in the early-season data - specifically in early season vigour and growth rate*.\n\n### About my baseline. \n\nFor the baseline I submitted, I'm just extracting image features with a pretrained resnet50 and then concatenating on the metadata, shoving it through a simple two-layer MLP, and training with SGD. That's it. No fancy optimizers, no state-of-the-art vision models, nothing. I didn't even tune any hyperparameters.\n\nComparing the distribution of predictions against the distribution of labels in the training set (which I assume should be *i.i.d*), the model doesn't predict a lot of the variance. I'm too lazy to write the code to break it down by maturity, but I'm pretty sure there will be more variance for earlier timepoints.\n\n![](https://i.imgur.com/0RgOmxm.png)\n\n### The model capacity doesn't really matter.\n\n> It don't matter. None of this matters. \n> \\- Carl Brutananadilewski\n\nI trained a few iterations of my baseline model and I can tell you that basically none of it matters. You can use a resnet50 or a resnet18, or probably something really tiny. You can fine-tune a pretrained model, train one from scratch, or just freeze the whole thing and use it as a feature extractor. It all yields basically the same validation error. This is probably because they are all capable of extracting vegetation area, which is where I think almost all of the signal comes from. \n\nI've also tried training with and without the heat units data, and it doesn't change validation error either.\n\n### How to beat my submission.\n\nAs I've alluded to, I'm pretty sure that the key to recovering more signal is to model the timeseries, which will allow you to get early-season growth rate. The data does include timeseries of plots over a growing season, which is great! However, the times are varied and, in the test set, the test points are divided into discrete early, mid, and late stages only. \n\nAnd of course, as with all Kaggle competitions, the other secret is to ride that submission button like it owes you money. The [more times you sample the same model](https://lukeoakdenrayner.wordpress.com/2019/09/19/ai-competitions-dont-produce-useful-models/) with a different seed, the better your best submission will probably be (unless the data held out from the public leaderboard is from a different domain).\n\nAnyway, thanks to the organizers for putting this together 🤗️.  I look forward to seeing the final models, and I hope that this has been useful to people who are thinking of participating.",
    "1366392": "Hi Jordan this is great work. I’d second the idea that using date would be essential. I am excited about the competition but won’t enter myself (since I’m on the TERRA-REF team). But if I did I would create a plant / soil mask and calculate %cover. Then can fit a logistic growth curve from which the max rate of increase could be a good predictor.\n\nSo I started running the masking if that would be of help … it is taking a few days to process. And that made me realize that any analysis must be very computationally intensive!\n\nI’d be happy to share masks and % cover with anyone interested when the processing is done.\n\nI’d also be happy to help folks get computing time - would first suggest they get an account at cyverse.org and search the discovery environment for the vice app of their choice. I can put the images, masks, and %cover data in a shared folder and otherwise help.",
    "1366540": "Hi David,\n\nYes that was almost exactly what I thought of doing! I gave the vegetation masks a try with excess green but couldn't find a reliable threshold and I ran out of time.  I'm sure those masks would be immensely helpful to participants. :)",
    "1368545": "Hi Jordan ... a lot of the images are overexposed, so our algorithm skips those. Without doing a lot of qa/qc or cleaning up. Code is in [github.com/az-digitalag/sorghum_biomass_prediction](https://github.com/az-digitalag/sorghum_biomass_prediction) and there is a list of files and their calculated canopy cover values in [output/canopy_cover.txt](https://github.com/az-digitalag/sorghum_biomass_prediction/blob/main/output/canopy_cover.txt?raw=true)\n\nThe canopy cover value goes from 0 = no plants to 1 = all plants / no soil. The file is structured with the file name followed by the canopy cover value. Sometimes it appears there are two file names in a row (in which case I suspect the first file threw an error but haven't dug in). \n\nI just wanted to create masks since I figured that would be the most difficult part and find interested parties to pick up the ball. I'll create a separate post as soon as I can with a link to both the masks and these data once I tar up the masks and put them somewhere.",
    "1384905": "Thanks for these tips - I my approach was largely based on this. I found that my models quickly overfit the training set.\nI'd be interested to know if other competitors had private LB scores better than mine from their earlier submissions, and whether anything more sophisticated worked.\n\nI'll try to write up a brief summary of what I did later.",
    "1405446": "Hi everyone. I just realize about this interesting competition and want to contribute to the discussion here. you may find our work at UIUC useful https://www.mdpi.com/2072-4292/13/9/1763  :)"
  },
  "source": "meta"
}