{
  "id": 98992,
  "title": "Is test set analysis ethical?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98992",
  "author_name": "",
  "post_date": "2019-07-08T05:13:38.447264900Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In many kaggle competitions, competitors try to analyse test set (leakage, distribution, etc.). In careercon 2019, LB score of 1 was achieved due to leakage. It is sometimes really helpful to win competition and doesn't matter much in competitions like careerCon.</p>\n\n<p>But, I have seen discussions around image size, distribution, etc in this competition also. It may be exploited and can be beneficial in improving LB score. In previous DR(Diabetic Retinopathy) competition, 4 years ago, image size was provided as metadata by top scoring team.</p>\n\n<p><em>But should we do that?</em> </p>\n\n<p>Where our algorithms might be used for detecting DR in real world, does it make sense to judge DR presence based on image size, for example? Imagine algorithm not predicting DR, because it was taken from different equipment and hence different variance and image size? It might ruin someone's life.</p>\n\n<p>So, IMHO, we should really avoid introducing such biases and exploiting test set, <em>for greater good of society.</em></p>",
  "messages": [
    {
      "id": "570289",
      "postDate": "07/08/2019 05:13:38",
      "content": "<p>In many kaggle competitions, competitors try to analyse test set (leakage, distribution, etc.). In careercon 2019, LB score of 1 was achieved due to leakage. It is sometimes really helpful to win competition and doesn't matter much in competitions like careerCon.</p>\n\n<p>But, I have seen discussions around image size, distribution, etc in this competition also. It may be exploited and can be beneficial in improving LB score. In previous DR(Diabetic Retinopathy) competition, 4 years ago, image size was provided as metadata by top scoring team.</p>\n\n<p><em>But should we do that?</em> </p>\n\n<p>Where our algorithms might be used for detecting DR in real world, does it make sense to judge DR presence based on image size, for example? Imagine algorithm not predicting DR, because it was taken from different equipment and hence different variance and image size? It might ruin someone's life.</p>\n\n<p>So, IMHO, we should really avoid introducing such biases and exploiting test set, <em>for greater good of society.</em></p>",
      "rawMarkdown": "In many kaggle competitions, competitors try to analyse test set (leakage, distribution, etc.). In careercon 2019, LB score of 1 was achieved due to leakage. It is sometimes really helpful to win competition and doesn't matter much in competitions like careerCon.\n\nBut, I have seen discussions around image size, distribution, etc in this competition also. It may be exploited and can be beneficial in improving LB score. In previous DR(Diabetic Retinopathy) competition, 4 years ago, image size was provided as metadata by top scoring team.\n\n*But should we do that?* \n\nWhere our algorithms might be used for detecting DR in real world, does it make sense to judge DR presence based on image size, for example? Imagine algorithm not predicting DR, because it was taken from different equipment and hence different variance and image size? It might ruin someone's life.\n\nSo, IMHO, we should really avoid introducing such biases and exploiting test set, *for greater good of society.*",
      "votes": null
    },
    {
      "id": "570383",
      "postDate": "07/08/2019 07:57:57",
      "content": "<p>Real world is different from Competitions...</p>",
      "rawMarkdown": "Real world is different from Competitions...",
      "votes": null
    },
    {
      "id": "579507",
      "postDate": "07/18/2019 22:38:54",
      "content": "<p>Hi <a href=\"/ashwan1\">@ashwan1</a> !\nI believe that analyzing test sets is insightful at least in Kaggle world. By understanding how test data are constructed and which kinds of leakage they have, we can discuss better ways to avoid overfitting to such meaningless features, and so machine learning techniques in real world will improve. Of course, I totally agree that private test data should not be improved using such meaningless features, and am interested in what private test dataset is like..</p>",
      "rawMarkdown": "Hi @ashwan1 !\nI believe that analyzing test sets is insightful at least in Kaggle world. By understanding how test data are constructed and which kinds of leakage they have, we can discuss better ways to avoid overfitting to such meaningless features, and so machine learning techniques in real world will improve. Of course, I totally agree that private test data should not be improved using such meaningless features, and am interested in what private test dataset is like..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 570383,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "07/08/2019 07:57:57",
      "content": "<p>Real world is different from Competitions...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 579507,
      "author_name": "kkondo",
      "author_url": "",
      "post_date": "07/18/2019 22:38:54",
      "content": "<p>Hi <a href=\"/ashwan1\">@ashwan1</a> !\nI believe that analyzing test sets is insightful at least in Kaggle world. By understanding how test data are constructed and which kinds of leakage they have, we can discuss better ways to avoid overfitting to such meaningless features, and so machine learning techniques in real world will improve. Of course, I totally agree that private test data should not be improved using such meaningless features, and am interested in what private test dataset is like..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "570289": "In many kaggle competitions, competitors try to analyse test set (leakage, distribution, etc.). In careercon 2019, LB score of 1 was achieved due to leakage. It is sometimes really helpful to win competition and doesn't matter much in competitions like careerCon.\n\nBut, I have seen discussions around image size, distribution, etc in this competition also. It may be exploited and can be beneficial in improving LB score. In previous DR(Diabetic Retinopathy) competition, 4 years ago, image size was provided as metadata by top scoring team.\n\n*But should we do that?* \n\nWhere our algorithms might be used for detecting DR in real world, does it make sense to judge DR presence based on image size, for example? Imagine algorithm not predicting DR, because it was taken from different equipment and hence different variance and image size? It might ruin someone's life.\n\nSo, IMHO, we should really avoid introducing such biases and exploiting test set, *for greater good of society.*",
    "570383": "Real world is different from Competitions...",
    "579507": "Hi @ashwan1 !\nI believe that analyzing test sets is insightful at least in Kaggle world. By understanding how test data are constructed and which kinds of leakage they have, we can discuss better ways to avoid overfitting to such meaningless features, and so machine learning techniques in real world will improve. Of course, I totally agree that private test data should not be improved using such meaningless features, and am interested in what private test dataset is like.."
  },
  "source": "meta"
}