{
  "id": 304784,
  "title": "Why not make it a code only competition?( To kaggle team and Organizers)",
  "url": "/competitions/happy-whale-and-dolphin/discussion/304784",
  "author_name": "DeepUnderstanding",
  "post_date": "2022-02-02T13:20:44.415000",
  "votes": 19,
  "comment_count": 4,
  "views": 0,
  "content": "<p>To Kaggle team and Organisers,</p>\n<p>In the Data section, it said that there are some issues with the data and I guess there might be some leakage(maybe).<br>\nI know some mistakes can happen but why take any chance, I think making the test data hidden and making it a code only competition would be better.</p>",
  "messages": [
    {
      "id": 1673051,
      "postDate": "2022-02-02T13:20:44.417Z",
      "content": "<p>To Kaggle team and Organisers,</p>\n<p>In the Data section, it said that there are some issues with the data and I guess there might be some leakage(maybe).<br>\nI know some mistakes can happen but why take any chance, I think making the test data hidden and making it a code only competition would be better.</p>",
      "rawMarkdown": "To Kaggle team and Organisers,\n\nIn the Data section, it said that there are some issues with the data and I guess there might be some leakage(maybe).\nI know some mistakes can happen but why take any chance, I think making the test data hidden and making it a code only competition would be better.",
      "votes": 19
    },
    {
      "id": 1674595,
      "postDate": "2022-02-03T16:11:06.923Z",
      "content": "<p>Great question! Theoretically, we could make everything a Code competition, since that better simulates \"real life\" machine learning where you don't have access to the future unseen test data. But Code competitions are an extra friction point for some people who prefer to have the entire workflow on their own system, so we generally only make a Code competition when it really is necessary for technical reasons.</p>\n<p>Why not this competition because of the (possible) data issues? The main reason is that the possible data issues wouldn't be solved with a Code competition. The most likely issue is an individual who is only photographed during a single encounter with similar images. We've tried to be smart with how theses are split <em>when we know they are from the same encounter</em>, but for some of the species we did not have photograph dates, so there may be unintentional, e.g., background leakage, that a classifier learns but that is not solved by having the leaky image \"hidden\" from the competitors.</p>\n<p>Another consideration is that it actually helps marine mammal researchers when the Kaggle community discovers and reports weaknesses in the data setup. The scrutiny helps improve the dataset for future iterations.</p>",
      "rawMarkdown": "Great question! Theoretically, we could make everything a Code competition, since that better simulates \"real life\" machine learning where you don't have access to the future unseen test data. But Code competitions are an extra friction point for some people who prefer to have the entire workflow on their own system, so we generally only make a Code competition when it really is necessary for technical reasons.\n\nWhy not this competition because of the (possible) data issues? The main reason is that the possible data issues wouldn't be solved with a Code competition. The most likely issue is an individual who is only photographed during a single encounter with similar images. We've tried to be smart with how theses are split *when we know they are from the same encounter*, but for some of the species we did not have photograph dates, so there may be unintentional, e.g., background leakage, that a classifier learns but that is not solved by having the leaky image \"hidden\" from the competitors.\n\nAnother consideration is that it actually helps marine mammal researchers when the Kaggle community discovers and reports weaknesses in the data setup. The scrutiny helps improve the dataset for future iterations.",
      "votes": 14,
      "replies": [
        {
          "id": 1675185,
          "postDate": "2022-02-04T04:58:15.407Z",
          "content": "<p>Thanks for explaining it very thoughtfully,<br>\n<code>background leakage,</code> in my initial model I can see this, images with the almost the same background are clustered together irrespective of their species. Anyway that's something I have to solve<br>\nThanks for the reply</p>",
          "rawMarkdown": "Thanks for explaining it very thoughtfully,\n`background leakage,` in my initial model I can see this, images with the almost the same background are clustered together irrespective of their species. Anyway that's something I have to solve\nThanks for the reply",
          "votes": 4
        }
      ]
    },
    {
      "id": 1673335,
      "postDate": "2022-02-02T16:56:34.967Z",
      "content": "<p>Yes, it is a good сomment, but I DON'T WANT TO WAIT FOR N HOURS! 😄 just kidding…</p>",
      "rawMarkdown": "Yes, it is a good сomment, but I DON'T WANT TO WAIT FOR N HOURS! 😄 just kidding...",
      "votes": 4
    },
    {
      "id": 1673107,
      "postDate": "2022-02-02T14:12:50.913Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1674595,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2022-02-03T16:11:06.923000",
      "content": "<p>Great question! Theoretically, we could make everything a Code competition, since that better simulates \"real life\" machine learning where you don't have access to the future unseen test data. But Code competitions are an extra friction point for some people who prefer to have the entire workflow on their own system, so we generally only make a Code competition when it really is necessary for technical reasons.</p>\n<p>Why not this competition because of the (possible) data issues? The main reason is that the possible data issues wouldn't be solved with a Code competition. The most likely issue is an individual who is only photographed during a single encounter with similar images. We've tried to be smart with how theses are split <em>when we know they are from the same encounter</em>, but for some of the species we did not have photograph dates, so there may be unintentional, e.g., background leakage, that a classifier learns but that is not solved by having the leaky image \"hidden\" from the competitors.</p>\n<p>Another consideration is that it actually helps marine mammal researchers when the Kaggle community discovers and reports weaknesses in the data setup. The scrutiny helps improve the dataset for future iterations.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 1675185,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-04T04:58:15.407000",
          "content": "<p>Thanks for explaining it very thoughtfully,<br>\n<code>background leakage,</code> in my initial model I can see this, images with the almost the same background are clustered together irrespective of their species. Anyway that's something I have to solve<br>\nThanks for the reply</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1673335,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2022-02-02T16:56:34.967000",
      "content": "<p>Yes, it is a good сomment, but I DON'T WANT TO WAIT FOR N HOURS! 😄 just kidding…</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1673107,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-02T14:12:50.913000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1673051": "To Kaggle team and Organisers,\n\nIn the Data section, it said that there are some issues with the data and I guess there might be some leakage(maybe).\nI know some mistakes can happen but why take any chance, I think making the test data hidden and making it a code only competition would be better.",
    "1674595": "Great question! Theoretically, we could make everything a Code competition, since that better simulates \"real life\" machine learning where you don't have access to the future unseen test data. But Code competitions are an extra friction point for some people who prefer to have the entire workflow on their own system, so we generally only make a Code competition when it really is necessary for technical reasons.\n\nWhy not this competition because of the (possible) data issues? The main reason is that the possible data issues wouldn't be solved with a Code competition. The most likely issue is an individual who is only photographed during a single encounter with similar images. We've tried to be smart with how theses are split *when we know they are from the same encounter*, but for some of the species we did not have photograph dates, so there may be unintentional, e.g., background leakage, that a classifier learns but that is not solved by having the leaky image \"hidden\" from the competitors.\n\nAnother consideration is that it actually helps marine mammal researchers when the Kaggle community discovers and reports weaknesses in the data setup. The scrutiny helps improve the dataset for future iterations.",
    "1673335": "Yes, it is a good сomment, but I DON'T WANT TO WAIT FOR N HOURS! 😄 just kidding...",
    "1673107": ""
  }
}