{
  "id": 103797,
  "title": "Why to remove duplicates? ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/103797",
  "author_name": "",
  "post_date": "2019-08-12T02:53:52.367462Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello guys,  I am only using the present competition data. Before removing duplicates my highest LB was 79.4 but after removing duplicates my LB is not going beyond 78. If this is the case, why to remove duplicates?  </p>",
  "messages": [
    {
      "id": "597220",
      "postDate": "08/12/2019 02:53:52",
      "content": "<p>Hello guys,  I am only using the present competition data. Before removing duplicates my highest LB was 79.4 but after removing duplicates my LB is not going beyond 78. If this is the case, why to remove duplicates?  </p>",
      "rawMarkdown": "Hello guys,  I am only using the present competition data. Before removing duplicates my highest LB was 79.4 but after removing duplicates my LB is not going beyond 78. If this is the case, why to remove duplicates?",
      "votes": null
    },
    {
      "id": "597378",
      "postDate": "08/12/2019 08:49:54",
      "content": "<p>the duplicate could make label inconsistent, however in this competition the private test data may have the same label inconsistent issue like train data, may be this is the reason why you remove the duplicates and get worse LB</p>",
      "rawMarkdown": "the duplicate could make label inconsistent, however in this competition the private test data may have the same label inconsistent issue like train data, may be this is the reason why you remove the duplicates and get worse LB",
      "votes": null
    },
    {
      "id": "597424",
      "postDate": "08/12/2019 10:35:32",
      "content": "<p>IMO, the duplicated images means diagnosis from dfferent doctors(for some images, the identical image has different diagnosises). so, there's no reason to remove them.\nbut when you split the dataset into train and validation, some duplicated images may fall into both splits. thus there's leakage from train to validation.  i think we should do something to avoid the leakage. my method is: for the same images, i just keep one image, and the label is average of those diagnosises. </p>\n\n<p>BTW,  how you get LB 79.4 without other datasets? i only got 7.0 with densenet so far. \nso what's your preprocessing, model, or other skills like ensemble, TTA, etc. ?</p>",
      "rawMarkdown": "IMO, the duplicated images means diagnosis from dfferent doctors(for some images, the identical image has different diagnosises). so, there's no reason to remove them.\nbut when you split the dataset into train and validation, some duplicated images may fall into both splits. thus there's leakage from train to validation.  i think we should do something to avoid the leakage. my method is: for the same images, i just keep one image, and the label is average of those diagnosises. \n\nBTW,  how you get LB 79.4 without other datasets? i only got 7.0 with densenet so far. \nso what's your preprocessing, model, or other skills like ensemble, TTA, etc. ?",
      "votes": null
    },
    {
      "id": "597432",
      "postDate": "08/12/2019 10:50:11",
      "content": "<p>I used EffifcientNetB5, cropping + ben's processing, augmentations like flips and zoom, and treated as multi-labeled classification. I didn't get good results with multiclass classification and regression. I didn't get good results with TTA aswell. I haven't yet tried ensembling.   </p>",
      "rawMarkdown": "I used EffifcientNetB5, cropping + ben's processing, augmentations like flips and zoom, and treated as multi-labeled classification. I didn't get good results with multiclass classification and regression. I didn't get good results with TTA aswell. I haven't yet tried ensembling.",
      "votes": null
    },
    {
      "id": "597458",
      "postDate": "08/12/2019 11:23:10",
      "content": "<p>As leixiang@seedsmed mentioned, the public test has duplicated images which are same as some training dataset, LB itself is not that reliable.  And reomoving them, the Public LB decreases. there are some possbile reasons. From your train validate data set split. to your stop at loss or kappa. and you would have a check after the competition whether your private LB drop.</p>",
      "rawMarkdown": "As leixiang@seedsmed mentioned, the public test has duplicated images which are same as some training dataset, LB itself is not that reliable.  And reomoving them, the Public LB decreases. there are some possbile reasons. From your train validate data set split. to your stop at loss or kappa. and you would have a check after the competition whether your private LB drop.",
      "votes": null
    },
    {
      "id": "597603",
      "postDate": "08/12/2019 14:59:51",
      "content": "<p>that's amazing!\ni have tried efficientnet-b4, train it as you said, but still got approximate 0.7 LB score☹️ </p>",
      "rawMarkdown": "that's amazing!\ni have tried efficientnet-b4, train it as you said, but still got approximate 0.7 LB score☹️",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 597378,
      "author_name": "leixiang",
      "author_url": "",
      "post_date": "08/12/2019 08:49:54",
      "content": "<p>the duplicate could make label inconsistent, however in this competition the private test data may have the same label inconsistent issue like train data, may be this is the reason why you remove the duplicates and get worse LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 597424,
      "author_name": "frank518",
      "author_url": "",
      "post_date": "08/12/2019 10:35:32",
      "content": "<p>IMO, the duplicated images means diagnosis from dfferent doctors(for some images, the identical image has different diagnosises). so, there's no reason to remove them.\nbut when you split the dataset into train and validation, some duplicated images may fall into both splits. thus there's leakage from train to validation.  i think we should do something to avoid the leakage. my method is: for the same images, i just keep one image, and the label is average of those diagnosises. </p>\n\n<p>BTW,  how you get LB 79.4 without other datasets? i only got 7.0 with densenet so far. \nso what's your preprocessing, model, or other skills like ensemble, TTA, etc. ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 597432,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/12/2019 10:50:11",
          "content": "<p>I used EffifcientNetB5, cropping + ben's processing, augmentations like flips and zoom, and treated as multi-labeled classification. I didn't get good results with multiclass classification and regression. I didn't get good results with TTA aswell. I haven't yet tried ensembling.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 597603,
          "author_name": "frank518",
          "author_url": "",
          "post_date": "08/12/2019 14:59:51",
          "content": "<p>that's amazing!\ni have tried efficientnet-b4, train it as you said, but still got approximate 0.7 LB score☹️ </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 597458,
      "author_name": "abnerzhang",
      "author_url": "",
      "post_date": "08/12/2019 11:23:10",
      "content": "<p>As leixiang@seedsmed mentioned, the public test has duplicated images which are same as some training dataset, LB itself is not that reliable.  And reomoving them, the Public LB decreases. there are some possbile reasons. From your train validate data set split. to your stop at loss or kappa. and you would have a check after the competition whether your private LB drop.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "597220": "Hello guys,  I am only using the present competition data. Before removing duplicates my highest LB was 79.4 but after removing duplicates my LB is not going beyond 78. If this is the case, why to remove duplicates?",
    "597378": "the duplicate could make label inconsistent, however in this competition the private test data may have the same label inconsistent issue like train data, may be this is the reason why you remove the duplicates and get worse LB",
    "597424": "IMO, the duplicated images means diagnosis from dfferent doctors(for some images, the identical image has different diagnosises). so, there's no reason to remove them.\nbut when you split the dataset into train and validation, some duplicated images may fall into both splits. thus there's leakage from train to validation.  i think we should do something to avoid the leakage. my method is: for the same images, i just keep one image, and the label is average of those diagnosises. \n\nBTW,  how you get LB 79.4 without other datasets? i only got 7.0 with densenet so far. \nso what's your preprocessing, model, or other skills like ensemble, TTA, etc. ?",
    "597432": "I used EffifcientNetB5, cropping + ben's processing, augmentations like flips and zoom, and treated as multi-labeled classification. I didn't get good results with multiclass classification and regression. I didn't get good results with TTA aswell. I haven't yet tried ensembling.",
    "597458": "As leixiang@seedsmed mentioned, the public test has duplicated images which are same as some training dataset, LB itself is not that reliable.  And reomoving them, the Public LB decreases. there are some possbile reasons. From your train validate data set split. to your stop at loss or kappa. and you would have a check after the competition whether your private LB drop.",
    "597603": "that's amazing!\ni have tried efficientnet-b4, train it as you said, but still got approximate 0.7 LB score☹️"
  },
  "source": "meta"
}