{
  "id": 13789,
  "title": "Were the Training and Testing sets randomly picked from same image pool?",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/13789",
  "author_name": "",
  "post_date": "2015-04-26T19:24:43.890Z",
  "votes": 1,
  "comment_count": 4,
  "views": 1207,
  "content": "<p>That is, can it be assumed that training and testing sets have similar distributions with respect to classes as well as other characteristics?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "74898",
      "postDate": "04/26/2015 19:24:43",
      "content": "<p>That is, can it be assumed that training and testing sets have similar distributions with respect to classes as well as other characteristics?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74941",
      "postDate": "04/27/2015 01:06:49",
      "content": "<p>I would be interested in an answer to this as well. &nbsp;Are these sets chosen using stratified sampling? &nbsp;How about the 80/20 split for the leaderboard vs final score?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75059",
      "postDate": "04/27/2015 16:27:54",
      "content": "<p>[quote=Dames;74898]</p>\n<p>can it be assumed that training and testing sets have similar distributions with respect to classes as well as other characteristics?</p>\n<p>[/quote]</p>\n<p>Yes (up to random variations). Splitting was random, not stratified.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75136",
      "postDate": "04/27/2015 22:29:07",
      "content": "<p>&nbsp;Thanks!&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75517",
      "postDate": "04/29/2015 22:16:24",
      "content": "<p>In the same line of enquiry, I was wondering if the 20% of the test images that you use to pre-score for the leaderboard is a fixed set of images from the testing lot or is it a random sample that is generated with every submission?&nbsp;</p>\n<p>Dames</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 74941,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/27/2015 01:06:49",
      "content": "<p>I would be interested in an answer to this as well. &nbsp;Are these sets chosen using stratified sampling? &nbsp;How about the 80/20 split for the leaderboard vs final score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75059,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "04/27/2015 16:27:54",
      "content": "<p>[quote=Dames;74898]</p>\n<p>can it be assumed that training and testing sets have similar distributions with respect to classes as well as other characteristics?</p>\n<p>[/quote]</p>\n<p>Yes (up to random variations). Splitting was random, not stratified.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75136,
      "author_name": "damianfondevila",
      "author_url": "",
      "post_date": "04/27/2015 22:29:07",
      "content": "<p>&nbsp;Thanks!&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75517,
      "author_name": "damianfondevila",
      "author_url": "",
      "post_date": "04/29/2015 22:16:24",
      "content": "<p>In the same line of enquiry, I was wondering if the 20% of the test images that you use to pre-score for the leaderboard is a fixed set of images from the testing lot or is it a random sample that is generated with every submission?&nbsp;</p>\n<p>Dames</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "74898": "",
    "74941": "",
    "75059": "",
    "75136": "",
    "75517": ""
  },
  "source": "meta"
}