{
  "id": 104754,
  "title": "Images with different labels in train and test",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104754",
  "author_name": "",
  "post_date": "2019-08-19T02:58:02.081687800Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We now know that the same image exists for train and test. This means that the labels of some of the images in the test are known.\nI have submitted a submission.csv that the label of Leakage Images replaced with that of train.csv. Surprisingly, the score dropped from 0.821 to 0.820. (Most results were the same as in train.csv, but there were slight changes in about 3 images.)\nThis means that test and train have the same image with different labels !!\nWhich labels should we trust?</p>",
  "messages": [
    {
      "id": "602384",
      "postDate": "08/19/2019 02:58:02",
      "content": "<p>We now know that the same image exists for train and test. This means that the labels of some of the images in the test are known.\nI have submitted a submission.csv that the label of Leakage Images replaced with that of train.csv. Surprisingly, the score dropped from 0.821 to 0.820. (Most results were the same as in train.csv, but there were slight changes in about 3 images.)\nThis means that test and train have the same image with different labels !!\nWhich labels should we trust?</p>",
      "rawMarkdown": "We now know that the same image exists for train and test. This means that the labels of some of the images in the test are known.\nI have submitted a submission.csv that the label of Leakage Images replaced with that of train.csv. Surprisingly, the score dropped from 0.821 to 0.820. (Most results were the same as in train.csv, but there were slight changes in about 3 images.)\nThis means that test and train have the same image with different labels !!\nWhich labels should we trust?",
      "votes": null
    },
    {
      "id": "602405",
      "postDate": "08/19/2019 03:42:56",
      "content": "<p>Can confirm this behaviour. Score dropped from 0.802 to 0.800. I have 4 submissions with 0.801 score with totally different target distributions(Checked on test.csv id_code). I stopped believing in Public LB, expecting a huge shake-up!</p>",
      "rawMarkdown": "Can confirm this behaviour. Score dropped from 0.802 to 0.800. I have 4 submissions with 0.801 score with totally different target distributions(Checked on test.csv id_code). I stopped believing in Public LB, expecting a huge shake-up!",
      "votes": null
    },
    {
      "id": "602410",
      "postDate": "08/19/2019 03:52:38",
      "content": "<p>It's a really noisy dataset. I wouldn't be surprised if private test data were annotated in a different way  with that of train data. Good Luck !!</p>",
      "rawMarkdown": "It's a really noisy dataset. I wouldn't be surprised if private test data were annotated in a different way  with that of train data. Good Luck !!",
      "votes": null
    },
    {
      "id": "602422",
      "postDate": "08/19/2019 04:05:28",
      "content": "<p>maybe the same images are labeled by difference doctor (this is normal). This behaviour just prove that training data is noise. I believe in the organizer and test set will be clean. And another important thing is that the metrics of this competition are kappa score, you just need to make sure the predicted labels +-1 compared with GT. I don't expect a huge shake-up :))) </p>",
      "rawMarkdown": "maybe the same images are labeled by difference doctor (this is normal). This behaviour just prove that training data is noise. I believe in the organizer and test set will be clean. And another important thing is that the metrics of this competition are kappa score, you just need to make sure the predicted labels +-1 compared with GT. I don't expect a huge shake-up :)))",
      "votes": null
    },
    {
      "id": "602467",
      "postDate": "08/19/2019 04:54:54",
      "content": "<p>in train dataset, there are images having very different labels. I hope there are not many such patterns in private test datasets....</p>\n\n<p>for example)\n3ee4841936ef: 4\n7005be54cab1: 1</p>\n\n<p>46cdc8b685bd: 4\ne4151feb8443: 2</p>\n\n<p>8273fdb4405e: 1\nf0098e9d4aee: 4</p>",
      "rawMarkdown": "in train dataset, there are images having very different labels. I hope there are not many such patterns in private test datasets....\n\nfor example)\n3ee4841936ef: 4\n7005be54cab1: 1\n\n46cdc8b685bd: 4\ne4151feb8443: 2\n\n8273fdb4405e: 1\nf0098e9d4aee: 4",
      "votes": null
    },
    {
      "id": "602492",
      "postDate": "08/19/2019 05:59:15",
      "content": "<p>There are quite a few wrongly labeled images also in test, and not only +-1 wrongly labeled.</p>",
      "rawMarkdown": "There are quite a few wrongly labeled images also in test, and not only +-1 wrongly labeled.",
      "votes": null
    },
    {
      "id": "602498",
      "postDate": "08/19/2019 06:13:26",
      "content": "<p>Do you have clean data train?. My clean data set is now about 3500 samples. Train with clean dataset, my LB decrease about 0.5 - 1% compared full train data set :)</p>",
      "rawMarkdown": "Do you have clean data train?. My clean data set is now about 3500 samples. Train with clean dataset, my LB decrease about 0.5 - 1% compared full train data set :)",
      "votes": null
    },
    {
      "id": "602499",
      "postDate": "08/19/2019 06:14:03",
      "content": "<p>just remove duplicate image with different label :). Hope private dataset is clean :)) </p>",
      "rawMarkdown": "just remove duplicate image with different label :). Hope private dataset is clean :))",
      "votes": null
    },
    {
      "id": "602543",
      "postDate": "08/19/2019 07:10:55",
      "content": "<p>Of course it decreases if test also contains those errors/inconsistencies.</p>",
      "rawMarkdown": "Of course it decreases if test also contains those errors/inconsistencies.",
      "votes": null
    },
    {
      "id": "602879",
      "postDate": "08/19/2019 15:53:53",
      "content": "<p>train in test, diff label in train roughly about 150 images</p>",
      "rawMarkdown": "train in test, diff label in train roughly about 150 images",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 602405,
      "author_name": "kirankunapuli",
      "author_url": "",
      "post_date": "08/19/2019 03:42:56",
      "content": "<p>Can confirm this behaviour. Score dropped from 0.802 to 0.800. I have 4 submissions with 0.801 score with totally different target distributions(Checked on test.csv id_code). I stopped believing in Public LB, expecting a huge shake-up!</p>",
      "votes": null,
      "replies": [
        {
          "id": 602410,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/19/2019 03:52:38",
          "content": "<p>It's a really noisy dataset. I wouldn't be surprised if private test data were annotated in a different way  with that of train data. Good Luck !!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 602422,
      "author_name": "dathudeptrai",
      "author_url": "",
      "post_date": "08/19/2019 04:05:28",
      "content": "<p>maybe the same images are labeled by difference doctor (this is normal). This behaviour just prove that training data is noise. I believe in the organizer and test set will be clean. And another important thing is that the metrics of this competition are kappa score, you just need to make sure the predicted labels +-1 compared with GT. I don't expect a huge shake-up :))) </p>",
      "votes": null,
      "replies": [
        {
          "id": 602467,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/19/2019 04:54:54",
          "content": "<p>in train dataset, there are images having very different labels. I hope there are not many such patterns in private test datasets....</p>\n\n<p>for example)\n3ee4841936ef: 4\n7005be54cab1: 1</p>\n\n<p>46cdc8b685bd: 4\ne4151feb8443: 2</p>\n\n<p>8273fdb4405e: 1\nf0098e9d4aee: 4</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602492,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/19/2019 05:59:15",
          "content": "<p>There are quite a few wrongly labeled images also in test, and not only +-1 wrongly labeled.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602498,
          "author_name": "dathudeptrai",
          "author_url": "",
          "post_date": "08/19/2019 06:13:26",
          "content": "<p>Do you have clean data train?. My clean data set is now about 3500 samples. Train with clean dataset, my LB decrease about 0.5 - 1% compared full train data set :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602499,
          "author_name": "dathudeptrai",
          "author_url": "",
          "post_date": "08/19/2019 06:14:03",
          "content": "<p>just remove duplicate image with different label :). Hope private dataset is clean :)) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602543,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/19/2019 07:10:55",
          "content": "<p>Of course it decreases if test also contains those errors/inconsistencies.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602879,
          "author_name": "ccjoshua",
          "author_url": "",
          "post_date": "08/19/2019 15:53:53",
          "content": "<p>train in test, diff label in train roughly about 150 images</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "602384": "We now know that the same image exists for train and test. This means that the labels of some of the images in the test are known.\nI have submitted a submission.csv that the label of Leakage Images replaced with that of train.csv. Surprisingly, the score dropped from 0.821 to 0.820. (Most results were the same as in train.csv, but there were slight changes in about 3 images.)\nThis means that test and train have the same image with different labels !!\nWhich labels should we trust?",
    "602405": "Can confirm this behaviour. Score dropped from 0.802 to 0.800. I have 4 submissions with 0.801 score with totally different target distributions(Checked on test.csv id_code). I stopped believing in Public LB, expecting a huge shake-up!",
    "602410": "It's a really noisy dataset. I wouldn't be surprised if private test data were annotated in a different way  with that of train data. Good Luck !!",
    "602422": "maybe the same images are labeled by difference doctor (this is normal). This behaviour just prove that training data is noise. I believe in the organizer and test set will be clean. And another important thing is that the metrics of this competition are kappa score, you just need to make sure the predicted labels +-1 compared with GT. I don't expect a huge shake-up :)))",
    "602467": "in train dataset, there are images having very different labels. I hope there are not many such patterns in private test datasets....\n\nfor example)\n3ee4841936ef: 4\n7005be54cab1: 1\n\n46cdc8b685bd: 4\ne4151feb8443: 2\n\n8273fdb4405e: 1\nf0098e9d4aee: 4",
    "602492": "There are quite a few wrongly labeled images also in test, and not only +-1 wrongly labeled.",
    "602498": "Do you have clean data train?. My clean data set is now about 3500 samples. Train with clean dataset, my LB decrease about 0.5 - 1% compared full train data set :)",
    "602499": "just remove duplicate image with different label :). Hope private dataset is clean :))",
    "602543": "Of course it decreases if test also contains those errors/inconsistencies.",
    "602879": "train in test, diff label in train roughly about 150 images"
  },
  "source": "meta"
}