{
  "id": 520894,
  "title": "Predictors and leakage (don't look if you don't want spoilers from my submissions)",
  "url": "/competitions/the-future-crop-challenge/discussion/520894",
  "author_name": "",
  "post_date": "2024-07-17T19:47:49.039761400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In my more recent attempts I made two changes: I subtracted the average of the grid from the target, re-adding it after modeling, and I included this average as a predictor. From a score of 0.95680, just the subtraction led to a performance of 0.82203 and including it as a predictor, to an even better performance of 0.79163.</p>\n<p>I'm thinking of removing the average as a predictor because it feels like a partial leakage. However, it did not seem to be a problem in my hold-out set and it does not seem to be a problem with the evaluation in the leaderboard. My idea when I included was to convey the yield level of the location (if it is a high-yield or low-yield and well, why not use an actual number if I have it).</p>\n<p>So, right now, I'm like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F231819%2Fb6d3391857713ef93d7b10146128edff%2Fmeme.jpeg?generation=1721233081037926&amp;alt=media\"></p>\n<p>Any thoughts?</p>",
  "messages": [
    {
      "id": "2926571",
      "postDate": "07/17/2024 19:47:49",
      "content": "<p>In my more recent attempts I made two changes: I subtracted the average of the grid from the target, re-adding it after modeling, and I included this average as a predictor. From a score of 0.95680, just the subtraction led to a performance of 0.82203 and including it as a predictor, to an even better performance of 0.79163.</p>\n<p>I'm thinking of removing the average as a predictor because it feels like a partial leakage. However, it did not seem to be a problem in my hold-out set and it does not seem to be a problem with the evaluation in the leaderboard. My idea when I included was to convey the yield level of the location (if it is a high-yield or low-yield and well, why not use an actual number if I have it).</p>\n<p>So, right now, I'm like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F231819%2Fb6d3391857713ef93d7b10146128edff%2Fmeme.jpeg?generation=1721233081037926&amp;alt=media\"></p>\n<p>Any thoughts?</p>",
      "rawMarkdown": "In my more recent attempts I made two changes: I subtracted the average of the grid from the target, re-adding it after modeling, and I included this average as a predictor. From a score of 0.95680, just the subtraction led to a performance of 0.82203 and including it as a predictor, to an even better performance of 0.79163.\n\nI'm thinking of removing the average as a predictor because it feels like a partial leakage. However, it did not seem to be a problem in my hold-out set and it does not seem to be a problem with the evaluation in the leaderboard. My idea when I included was to convey the yield level of the location (if it is a high-yield or low-yield and well, why not use an actual number if I have it).\n\nSo, right now, I'm like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F231819%2Fb6d3391857713ef93d7b10146128edff%2Fmeme.jpeg?generation=1721233081037926&alt=media)\n\nAny thoughts?",
      "votes": null
    },
    {
      "id": "2927116",
      "postDate": "07/18/2024 09:18:15",
      "content": "<p>it's not leakage in my opinion! you're not using any information from the test set. totally reasonable approach in my opinion. of course, it might not help you get the trend in the yields…</p>",
      "rawMarkdown": "it's not leakage in my opinion! you're not using any information from the test set. totally reasonable approach in my opinion. of course, it might not help you get the trend in the yields...",
      "votes": null
    },
    {
      "id": "2927177",
      "postDate": "07/18/2024 10:10:48",
      "content": "<p>Good points, thank you!</p>",
      "rawMarkdown": "Good points, thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2927116,
      "author_name": "lilybellesweet",
      "author_url": "",
      "post_date": "07/18/2024 09:18:15",
      "content": "<p>it's not leakage in my opinion! you're not using any information from the test set. totally reasonable approach in my opinion. of course, it might not help you get the trend in the yields…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2927177,
          "author_name": "picsoflily",
          "author_url": "",
          "post_date": "07/18/2024 10:10:48",
          "content": "<p>Good points, thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2926571": "In my more recent attempts I made two changes: I subtracted the average of the grid from the target, re-adding it after modeling, and I included this average as a predictor. From a score of 0.95680, just the subtraction led to a performance of 0.82203 and including it as a predictor, to an even better performance of 0.79163.\n\nI'm thinking of removing the average as a predictor because it feels like a partial leakage. However, it did not seem to be a problem in my hold-out set and it does not seem to be a problem with the evaluation in the leaderboard. My idea when I included was to convey the yield level of the location (if it is a high-yield or low-yield and well, why not use an actual number if I have it).\n\nSo, right now, I'm like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F231819%2Fb6d3391857713ef93d7b10146128edff%2Fmeme.jpeg?generation=1721233081037926&alt=media)\n\nAny thoughts?",
    "2927116": "it's not leakage in my opinion! you're not using any information from the test set. totally reasonable approach in my opinion. of course, it might not help you get the trend in the yields...",
    "2927177": "Good points, thank you!"
  },
  "source": "meta"
}