{
  "id": 206018,
  "title": "Duplicates across 2019 and 2020 data, and potential leaks.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206018",
  "author_name": "",
  "post_date": "2020-12-22T20:32:40.810395700Z",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>By now everybody knows that there are duplicates between this competition's train set and the labelled part of the <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">2019 in-class competiton</a>. However, I have not seen anything about duplicates between the current data and the unlabelled test set and extra images of the 2019 dataset. </p>\n<p>I used <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">this great notebook</a> to find duplicates and here are my findings.</p>\n<p>I analyzed duplicates between all of the following: train 2020, train 2019, test 2019, extra 2019.<br>\nI found a total of <strong>5570</strong> duplicate pairs.</p>\n<p>This is the distribution of them across datasets:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F2cb5fa008b5993e40a37ae2146d8ee75%2FScreenshot%20from%202020-12-22%2023-16-38.png?generation=1608668217600380&amp;alt=media\" alt=\"\"></p>\n<p>There are a few things that should catch your attention.</p>\n<ol>\n<li>The train data in 2020 contains 4258 images that were in extra 2019. Could they have been labelled using the winning solutions of 2019 competition? Maybe the labels on those 4258 images are noisier than others? </li>\n<li>There were duplicates within the test of 2019. There were also images with labels in the train set of 2019, and the same images without labels in extra 2019.  </li>\n<li>There were leaks in the 2019 competition: there were images from the test set in the extra data in 2019.</li>\n</ol>\n<p>If this competition was prepared using the same tools and approaches as the previous one, similar leaks and issues could be present. I hope the organizers will investigate.</p>\n<p>Here are some examples.<br>\nImages 656 and 1443 from 2019 test set. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F5a723898f4d2e0dce98633b1490a1452%2Fdownload%20(1).png?generation=1608668584328693&amp;alt=media\" alt=\"\"></p>\n<p>Image 2937 from test and its duplicate extra 8731.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F68e72deac62f39c713df940e8a797f2d%2Fdownload%20(2).png?generation=1608668668036273&amp;alt=media\" alt=\"\"></p>\n<p><strong>There are images where labels from 2019 and 2020 do not agree.</strong><br>\nImage 20755 from train 2020 and 327 from train 2019.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F28a4c12e3b46115895f78b9bae748741%2Fdownload%20(3).png?generation=1608668883228234&amp;alt=media\" alt=\"\"><br>\nSubjectively it seems like new labels are cleaner.</p>\n<p>The distribution of duplicates across labels seems to be as expected, I don't know what to make of it:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F8546e1d537e9111a3c1608b45591c502%2FScreenshot%20from%202020-12-22%2023-31-10.png?generation=1608669100604370&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1122976",
      "postDate": "12/22/2020 20:32:40",
      "content": "<p>By now everybody knows that there are duplicates between this competition's train set and the labelled part of the <a href=\"https://www.kaggle.com/c/cassava-disease\" target=\"_blank\">2019 in-class competiton</a>. However, I have not seen anything about duplicates between the current data and the unlabelled test set and extra images of the 2019 dataset. </p>\n<p>I used <a href=\"https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions\" target=\"_blank\">this great notebook</a> to find duplicates and here are my findings.</p>\n<p>I analyzed duplicates between all of the following: train 2020, train 2019, test 2019, extra 2019.<br>\nI found a total of <strong>5570</strong> duplicate pairs.</p>\n<p>This is the distribution of them across datasets:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F2cb5fa008b5993e40a37ae2146d8ee75%2FScreenshot%20from%202020-12-22%2023-16-38.png?generation=1608668217600380&amp;alt=media\" alt=\"\"></p>\n<p>There are a few things that should catch your attention.</p>\n<ol>\n<li>The train data in 2020 contains 4258 images that were in extra 2019. Could they have been labelled using the winning solutions of 2019 competition? Maybe the labels on those 4258 images are noisier than others? </li>\n<li>There were duplicates within the test of 2019. There were also images with labels in the train set of 2019, and the same images without labels in extra 2019.  </li>\n<li>There were leaks in the 2019 competition: there were images from the test set in the extra data in 2019.</li>\n</ol>\n<p>If this competition was prepared using the same tools and approaches as the previous one, similar leaks and issues could be present. I hope the organizers will investigate.</p>\n<p>Here are some examples.<br>\nImages 656 and 1443 from 2019 test set. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F5a723898f4d2e0dce98633b1490a1452%2Fdownload%20(1).png?generation=1608668584328693&amp;alt=media\" alt=\"\"></p>\n<p>Image 2937 from test and its duplicate extra 8731.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F68e72deac62f39c713df940e8a797f2d%2Fdownload%20(2).png?generation=1608668668036273&amp;alt=media\" alt=\"\"></p>\n<p><strong>There are images where labels from 2019 and 2020 do not agree.</strong><br>\nImage 20755 from train 2020 and 327 from train 2019.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F28a4c12e3b46115895f78b9bae748741%2Fdownload%20(3).png?generation=1608668883228234&amp;alt=media\" alt=\"\"><br>\nSubjectively it seems like new labels are cleaner.</p>\n<p>The distribution of duplicates across labels seems to be as expected, I don't know what to make of it:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F8546e1d537e9111a3c1608b45591c502%2FScreenshot%20from%202020-12-22%2023-31-10.png?generation=1608669100604370&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "By now everybody knows that there are duplicates between this competition's train set and the labelled part of the [2019 in-class competiton](https://www.kaggle.com/c/cassava-disease). However, I have not seen anything about duplicates between the current data and the unlabelled test set and extra images of the 2019 dataset. \n\nI used [this great notebook](https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions) to find duplicates and here are my findings.\n\nI analyzed duplicates between all of the following: train 2020, train 2019, test 2019, extra 2019.\nI found a total of **5570** duplicate pairs.\n\nThis is the distribution of them across datasets:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F2cb5fa008b5993e40a37ae2146d8ee75%2FScreenshot%20from%202020-12-22%2023-16-38.png?generation=1608668217600380&alt=media)\n\nThere are a few things that should catch your attention.\n1. The train data in 2020 contains 4258 images that were in extra 2019. Could they have been labelled using the winning solutions of 2019 competition? Maybe the labels on those 4258 images are noisier than others? \n2. There were duplicates within the test of 2019. There were also images with labels in the train set of 2019, and the same images without labels in extra 2019.  \n3. There were leaks in the 2019 competition: there were images from the test set in the extra data in 2019.\n\nIf this competition was prepared using the same tools and approaches as the previous one, similar leaks and issues could be present. I hope the organizers will investigate.\n\nHere are some examples.\nImages 656 and 1443 from 2019 test set. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F5a723898f4d2e0dce98633b1490a1452%2Fdownload%20(1).png?generation=1608668584328693&alt=media)\n\n\nImage 2937 from test and its duplicate extra 8731.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F68e72deac62f39c713df940e8a797f2d%2Fdownload%20(2).png?generation=1608668668036273&alt=media)\n\n**There are images where labels from 2019 and 2020 do not agree.**\nImage 20755 from train 2020 and 327 from train 2019.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F28a4c12e3b46115895f78b9bae748741%2Fdownload%20(3).png?generation=1608668883228234&alt=media)\nSubjectively it seems like new labels are cleaner.\n\nThe distribution of duplicates across labels seems to be as expected, I don't know what to make of it:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F8546e1d537e9111a3c1608b45591c502%2FScreenshot%20from%202020-12-22%2023-31-10.png?generation=1608669100604370&alt=media)",
      "votes": null
    },
    {
      "id": "1122991",
      "postDate": "12/22/2020 20:48:06",
      "content": "<p>The preparation pipeline for this competition was built from scratch relative to the 2019 competition. We are aware of the situation and have taken measures to ensure the integrity of the leaderboard.</p>",
      "rawMarkdown": "The preparation pipeline for this competition was built from scratch relative to the 2019 competition. We are aware of the situation and have taken measures to ensure the integrity of the leaderboard.",
      "votes": null
    },
    {
      "id": "1123544",
      "postDate": "12/23/2020 10:11:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/btseytlin\" target=\"_blank\">@btseytlin</a> interesting discussion.<br>\nDo you mind to share your notebook to find those duplicates? (the link in your post does not point to a notebook)</p>",
      "rawMarkdown": "Hi @btseytlin interesting discussion.\nDo you mind to share your notebook to find those duplicates? (the link in your post does not point to a notebook)",
      "votes": null
    },
    {
      "id": "1123574",
      "postDate": "12/23/2020 10:40:37",
      "content": "<p>Sorry, I fixed the link!</p>\n<p>I will consider making a public notebook (currently the code is not tidy enough :) ), but this original notebook should be enough to reproduce the results.</p>",
      "rawMarkdown": "Sorry, I fixed the link!\n\nI will consider making a public notebook (currently the code is not tidy enough :) ), but this original notebook should be enough to reproduce the results.",
      "votes": null
    },
    {
      "id": "1123587",
      "postDate": "12/23/2020 10:57:59",
      "content": "<p><code>The train data in 2020 contains 4258 images that were in extra 2019. They have been labelled using the winning solutions of 2019 competition.</code></p>\n<p>Has been this officially stated by competition host or kaggle staff? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "`The train data in 2020 contains 4258 images that were in extra 2019. They have been labelled using the winning solutions of 2019 competition.`\n\nHas been this officially stated by competition host or kaggle staff? @sohier",
      "votes": null
    },
    {
      "id": "1123615",
      "postDate": "12/23/2020 11:22:54",
      "content": "<p>It was not stated, sorry for confusing you, I fixed the wording. Supposed to say \"maybe they have been labelled\" :)</p>",
      "rawMarkdown": "It was not stated, sorry for confusing you, I fixed the wording. Supposed to say \"maybe they have been labelled\" :)",
      "votes": null
    },
    {
      "id": "1123620",
      "postDate": "12/23/2020 11:29:17",
      "content": "<p>Don't worry, they  have already removed all potential leaks during LB re-score. </p>\n<p>Clearly, Unlabelled 2019 dataset was present in the previous LB. </p>",
      "rawMarkdown": "Don't worry, they  have already removed all potential leaks during LB re-score. \n\nClearly, Unlabelled 2019 dataset was present in the previous LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1122991,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "12/22/2020 20:48:06",
      "content": "<p>The preparation pipeline for this competition was built from scratch relative to the 2019 competition. We are aware of the situation and have taken measures to ensure the integrity of the leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1123544,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "12/23/2020 10:11:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/btseytlin\" target=\"_blank\">@btseytlin</a> interesting discussion.<br>\nDo you mind to share your notebook to find those duplicates? (the link in your post does not point to a notebook)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123574,
          "author_name": "btseytlin",
          "author_url": "",
          "post_date": "12/23/2020 10:40:37",
          "content": "<p>Sorry, I fixed the link!</p>\n<p>I will consider making a public notebook (currently the code is not tidy enough :) ), but this original notebook should be enough to reproduce the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123587,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "12/23/2020 10:57:59",
      "content": "<p><code>The train data in 2020 contains 4258 images that were in extra 2019. They have been labelled using the winning solutions of 2019 competition.</code></p>\n<p>Has been this officially stated by competition host or kaggle staff? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1123615,
          "author_name": "btseytlin",
          "author_url": "",
          "post_date": "12/23/2020 11:22:54",
          "content": "<p>It was not stated, sorry for confusing you, I fixed the wording. Supposed to say \"maybe they have been labelled\" :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123620,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/23/2020 11:29:17",
      "content": "<p>Don't worry, they  have already removed all potential leaks during LB re-score. </p>\n<p>Clearly, Unlabelled 2019 dataset was present in the previous LB. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1122976": "By now everybody knows that there are duplicates between this competition's train set and the labelled part of the [2019 in-class competiton](https://www.kaggle.com/c/cassava-disease). However, I have not seen anything about duplicates between the current data and the unlabelled test set and extra images of the 2019 dataset. \n\nI used [this great notebook](https://www.kaggle.com/zzy990106/duplicate-images-in-two-competitions) to find duplicates and here are my findings.\n\nI analyzed duplicates between all of the following: train 2020, train 2019, test 2019, extra 2019.\nI found a total of **5570** duplicate pairs.\n\nThis is the distribution of them across datasets:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F2cb5fa008b5993e40a37ae2146d8ee75%2FScreenshot%20from%202020-12-22%2023-16-38.png?generation=1608668217600380&alt=media)\n\nThere are a few things that should catch your attention.\n1. The train data in 2020 contains 4258 images that were in extra 2019. Could they have been labelled using the winning solutions of 2019 competition? Maybe the labels on those 4258 images are noisier than others? \n2. There were duplicates within the test of 2019. There were also images with labels in the train set of 2019, and the same images without labels in extra 2019.  \n3. There were leaks in the 2019 competition: there were images from the test set in the extra data in 2019.\n\nIf this competition was prepared using the same tools and approaches as the previous one, similar leaks and issues could be present. I hope the organizers will investigate.\n\nHere are some examples.\nImages 656 and 1443 from 2019 test set. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F5a723898f4d2e0dce98633b1490a1452%2Fdownload%20(1).png?generation=1608668584328693&alt=media)\n\n\nImage 2937 from test and its duplicate extra 8731.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F68e72deac62f39c713df940e8a797f2d%2Fdownload%20(2).png?generation=1608668668036273&alt=media)\n\n**There are images where labels from 2019 and 2020 do not agree.**\nImage 20755 from train 2020 and 327 from train 2019.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F28a4c12e3b46115895f78b9bae748741%2Fdownload%20(3).png?generation=1608668883228234&alt=media)\nSubjectively it seems like new labels are cleaner.\n\nThe distribution of duplicates across labels seems to be as expected, I don't know what to make of it:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1423624%2F8546e1d537e9111a3c1608b45591c502%2FScreenshot%20from%202020-12-22%2023-31-10.png?generation=1608669100604370&alt=media)",
    "1122991": "The preparation pipeline for this competition was built from scratch relative to the 2019 competition. We are aware of the situation and have taken measures to ensure the integrity of the leaderboard.",
    "1123544": "Hi @btseytlin interesting discussion.\nDo you mind to share your notebook to find those duplicates? (the link in your post does not point to a notebook)",
    "1123574": "Sorry, I fixed the link!\n\nI will consider making a public notebook (currently the code is not tidy enough :) ), but this original notebook should be enough to reproduce the results.",
    "1123587": "`The train data in 2020 contains 4258 images that were in extra 2019. They have been labelled using the winning solutions of 2019 competition.`\n\nHas been this officially stated by competition host or kaggle staff? @sohier",
    "1123615": "It was not stated, sorry for confusing you, I fixed the wording. Supposed to say \"maybe they have been labelled\" :)",
    "1123620": "Don't worry, they  have already removed all potential leaks during LB re-score. \n\nClearly, Unlabelled 2019 dataset was present in the previous LB."
  },
  "source": "meta"
}