{
  "id": 107945,
  "title": "Things I don't like about this competition",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107945",
  "author_name": "",
  "post_date": "2019-09-08T04:09:18.006638200Z",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The competition has ended. Compared to the 2015 Diabetic Retinopathy, there are several things I don't like about this competition. I think the future competitions could have a better experience if these things can be solved.</p>\n\n<ol>\n<li><p><strong>Aspect Ratio Leakage</strong>\nAlthough the images are of different aspect ratios/resolutions in the data of this competition, you will find that there are actually only 2 aspect ratios (1:1 and 4:3) after removing the black exterior region. The class distribution for these two kinds of images are following,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fdf3f800544f6081c6887f202f963a243%2FWeChat%20Image_20190907174956.png?generation=1567913720075761&amp;alt=media\" alt=\"\">\nYou can easily find the severe leakage in the 1:1 images for this competition, which does not exist in the 2015 data. The 4:3 images seem to be the zoomed in images and the 4:3 2019 images are enlarged at different levels. </p></li>\n<li><p><strong>Small Public Test Size</strong>\nThe public vs private test size is around 1:6.5. I have showed the significant impacts brought by small sample size a few days ago. You can find the details <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#617240\">here</a>. In summary, assuming the samples are drawn randomly (which is not true in this competition), small sampling size may lead to +/- 0.04 in the QWK. If we are using the ratio of 2015 (~1:4), the shake would shrink to less than +/- 0.01. </p></li>\n<li><p><strong>Inconsistent Public/Private Distribution</strong>\nWith the private score revealed, it is certain the the private test set has a distribution similar to train set, rather than the public test set. I believe this is another reason that leads to the huge variation between the public and private score. </p></li>\n</ol>\n\n<p>QWK is a metrics that heavily relies on distribution. Small sampling size + inconsistent distributions are nightmares for this kind of metrics. I believe this is the reason for over 0.1 variation in public/private scores. I understand that Item #1 and #2 are situations we may face in the real world, but for Item #3 I have no clue why the competition is designed in this way. Perhaps you guys could enlighten me a little bit. </p>",
  "messages": [
    {
      "id": "620913",
      "postDate": "09/08/2019 04:09:18",
      "content": "<p>The competition has ended. Compared to the 2015 Diabetic Retinopathy, there are several things I don't like about this competition. I think the future competitions could have a better experience if these things can be solved.</p>\n\n<ol>\n<li><p><strong>Aspect Ratio Leakage</strong>\nAlthough the images are of different aspect ratios/resolutions in the data of this competition, you will find that there are actually only 2 aspect ratios (1:1 and 4:3) after removing the black exterior region. The class distribution for these two kinds of images are following,\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fdf3f800544f6081c6887f202f963a243%2FWeChat%20Image_20190907174956.png?generation=1567913720075761&amp;alt=media\" alt=\"\">\nYou can easily find the severe leakage in the 1:1 images for this competition, which does not exist in the 2015 data. The 4:3 images seem to be the zoomed in images and the 4:3 2019 images are enlarged at different levels. </p></li>\n<li><p><strong>Small Public Test Size</strong>\nThe public vs private test size is around 1:6.5. I have showed the significant impacts brought by small sample size a few days ago. You can find the details <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#617240\">here</a>. In summary, assuming the samples are drawn randomly (which is not true in this competition), small sampling size may lead to +/- 0.04 in the QWK. If we are using the ratio of 2015 (~1:4), the shake would shrink to less than +/- 0.01. </p></li>\n<li><p><strong>Inconsistent Public/Private Distribution</strong>\nWith the private score revealed, it is certain the the private test set has a distribution similar to train set, rather than the public test set. I believe this is another reason that leads to the huge variation between the public and private score. </p></li>\n</ol>\n\n<p>QWK is a metrics that heavily relies on distribution. Small sampling size + inconsistent distributions are nightmares for this kind of metrics. I believe this is the reason for over 0.1 variation in public/private scores. I understand that Item #1 and #2 are situations we may face in the real world, but for Item #3 I have no clue why the competition is designed in this way. Perhaps you guys could enlighten me a little bit. </p>",
      "rawMarkdown": "The competition has ended. Compared to the 2015 Diabetic Retinopathy, there are several things I don't like about this competition. I think the future competitions could have a better experience if these things can be solved.\n\n1. **Aspect Ratio Leakage**\nAlthough the images are of different aspect ratios/resolutions in the data of this competition, you will find that there are actually only 2 aspect ratios (1:1 and 4:3) after removing the black exterior region. The class distribution for these two kinds of images are following,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fdf3f800544f6081c6887f202f963a243%2FWeChat%20Image_20190907174956.png?generation=1567913720075761&amp;alt=media)\nYou can easily find the severe leakage in the 1:1 images for this competition, which does not exist in the 2015 data. The 4:3 images seem to be the zoomed in images and the 4:3 2019 images are enlarged at different levels. \n\n2. **Small Public Test Size**\nThe public vs private test size is around 1:6.5. I have showed the significant impacts brought by small sample size a few days ago. You can find the details [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#617240). In summary, assuming the samples are drawn randomly (which is not true in this competition), small sampling size may lead to +/- 0.04 in the QWK. If we are using the ratio of 2015 (~1:4), the shake would shrink to less than +/- 0.01. \n\n3. **Inconsistent Public/Private Distribution**\nWith the private score revealed, it is certain the the private test set has a distribution similar to train set, rather than the public test set. I believe this is another reason that leads to the huge variation between the public and private score. \n\nQWK is a metrics that heavily relies on distribution. Small sampling size + inconsistent distributions are nightmares for this kind of metrics. I believe this is the reason for over 0.1 variation in public/private scores. I understand that Item #1 and #2 are situations we may face in the real world, but for Item #3 I have no clue why the competition is designed in this way. Perhaps you guys could enlighten me a little bit.",
      "votes": null
    },
    {
      "id": "620961",
      "postDate": "09/08/2019 05:25:59",
      "content": "<p>I agree. I also found the inconsistent public/private distribution to be very frustrating. To trust your CV in this competition required a massive leap of faith, since the public test set seemed to be so different from the train set provided. I also think it resulted in a lot more submissions than otherwise would be necessary which appears to have resulted in problems with Kaggle's infrastructure.</p>\n\n<p>Congrats on your final score <a href=\"/naivelamb\">@naivelamb</a>!</p>",
      "rawMarkdown": "I agree. I also found the inconsistent public/private distribution to be very frustrating. To trust your CV in this competition required a massive leap of faith, since the public test set seemed to be so different from the train set provided. I also think it resulted in a lot more submissions than otherwise would be necessary which appears to have resulted in problems with Kaggle's infrastructure.\n\nCongrats on your final score @naivelamb!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 620961,
      "author_name": "lextoumbourou",
      "author_url": "",
      "post_date": "09/08/2019 05:25:59",
      "content": "<p>I agree. I also found the inconsistent public/private distribution to be very frustrating. To trust your CV in this competition required a massive leap of faith, since the public test set seemed to be so different from the train set provided. I also think it resulted in a lot more submissions than otherwise would be necessary which appears to have resulted in problems with Kaggle's infrastructure.</p>\n\n<p>Congrats on your final score <a href=\"/naivelamb\">@naivelamb</a>!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "620913": "The competition has ended. Compared to the 2015 Diabetic Retinopathy, there are several things I don't like about this competition. I think the future competitions could have a better experience if these things can be solved.\n\n1. **Aspect Ratio Leakage**\nAlthough the images are of different aspect ratios/resolutions in the data of this competition, you will find that there are actually only 2 aspect ratios (1:1 and 4:3) after removing the black exterior region. The class distribution for these two kinds of images are following,\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fdf3f800544f6081c6887f202f963a243%2FWeChat%20Image_20190907174956.png?generation=1567913720075761&amp;alt=media)\nYou can easily find the severe leakage in the 1:1 images for this competition, which does not exist in the 2015 data. The 4:3 images seem to be the zoomed in images and the 4:3 2019 images are enlarged at different levels. \n\n2. **Small Public Test Size**\nThe public vs private test size is around 1:6.5. I have showed the significant impacts brought by small sample size a few days ago. You can find the details [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#617240). In summary, assuming the samples are drawn randomly (which is not true in this competition), small sampling size may lead to +/- 0.04 in the QWK. If we are using the ratio of 2015 (~1:4), the shake would shrink to less than +/- 0.01. \n\n3. **Inconsistent Public/Private Distribution**\nWith the private score revealed, it is certain the the private test set has a distribution similar to train set, rather than the public test set. I believe this is another reason that leads to the huge variation between the public and private score. \n\nQWK is a metrics that heavily relies on distribution. Small sampling size + inconsistent distributions are nightmares for this kind of metrics. I believe this is the reason for over 0.1 variation in public/private scores. I understand that Item #1 and #2 are situations we may face in the real world, but for Item #3 I have no clue why the competition is designed in this way. Perhaps you guys could enlighten me a little bit.",
    "620961": "I agree. I also found the inconsistent public/private distribution to be very frustrating. To trust your CV in this competition required a massive leap of faith, since the public test set seemed to be so different from the train set provided. I also think it resulted in a lot more submissions than otherwise would be necessary which appears to have resulted in problems with Kaggle's infrastructure.\n\nCongrats on your final score @naivelamb!"
  },
  "source": "meta"
}