{
  "id": 393693,
  "title": "Test Set Distribution [Handedness, etc.]",
  "url": "/competitions/asl-signs/discussion/393693",
  "author_name": "",
  "post_date": "2023-03-10T12:31:21.729179Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there, I'm curious to know whether the test set distribution can be expected to be similar to the train set w.r.t. handedness and any other secondary population defining characteristics.</p>\n<ul>\n<li>This would mean that the test set is controlled to ensure a similar distribution.</li>\n</ul>\n<p>Alternatively, the data could be stratified and grouped by participant_id and then both train and test could have been sampled from some super set.</p>\n<ul>\n<li>This could potentially result in differing train/test secondary distributions for things like handedness.</li>\n</ul>\n<p>Any insights into this? <a href=\"https://www.kaggle.com/samsepah\" target=\"_blank\">@samsepah</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
  "messages": [
    {
      "id": "2176151",
      "postDate": "03/10/2023 12:31:21",
      "content": "<p>Hi there, I'm curious to know whether the test set distribution can be expected to be similar to the train set w.r.t. handedness and any other secondary population defining characteristics.</p>\n<ul>\n<li>This would mean that the test set is controlled to ensure a similar distribution.</li>\n</ul>\n<p>Alternatively, the data could be stratified and grouped by participant_id and then both train and test could have been sampled from some super set.</p>\n<ul>\n<li>This could potentially result in differing train/test secondary distributions for things like handedness.</li>\n</ul>\n<p>Any insights into this? <a href=\"https://www.kaggle.com/samsepah\" target=\"_blank\">@samsepah</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "Hi there, I'm curious to know whether the test set distribution can be expected to be similar to the train set w.r.t. handedness and any other secondary population defining characteristics.\n* This would mean that the test set is controlled to ensure a similar distribution.\n\nAlternatively, the data could be stratified and grouped by participant_id and then both train and test could have been sampled from some super set.\n* This could potentially result in differing train/test secondary distributions for things like handedness.\n\nAny insights into this? @samsepah @sohier",
      "votes": null
    },
    {
      "id": "2176450",
      "postDate": "03/10/2023 16:57:26",
      "content": "<p>My default assumption and question for organizers is whether we can expect the test set to be far more diverse (much less than 4000 rows per participant in test set)? Or was the test set generated similarly to the train set (\"Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. \")?</p>",
      "rawMarkdown": "My default assumption and question for organizers is whether we can expect the test set to be far more diverse (much less than 4000 rows per participant in test set)? Or was the test set generated similarly to the train set (\"Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. \")?",
      "votes": null
    },
    {
      "id": "2177210",
      "postDate": "03/11/2023 09:29:12",
      "content": "<p>The safe assumption I generally make is always that the test set will be independent from and contain similar variability to the train set (IID assumption).</p>",
      "rawMarkdown": "The safe assumption I generally make is always that the test set will be independent from and contain similar variability to the train set (IID assumption).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2176450,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "03/10/2023 16:57:26",
      "content": "<p>My default assumption and question for organizers is whether we can expect the test set to be far more diverse (much less than 4000 rows per participant in test set)? Or was the test set generated similarly to the train set (\"Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. \")?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2177210,
          "author_name": "wonderingalice",
          "author_url": "",
          "post_date": "03/11/2023 09:29:12",
          "content": "<p>The safe assumption I generally make is always that the test set will be independent from and contain similar variability to the train set (IID assumption).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2176151": "Hi there, I'm curious to know whether the test set distribution can be expected to be similar to the train set w.r.t. handedness and any other secondary population defining characteristics.\n* This would mean that the test set is controlled to ensure a similar distribution.\n\nAlternatively, the data could be stratified and grouped by participant_id and then both train and test could have been sampled from some super set.\n* This could potentially result in differing train/test secondary distributions for things like handedness.\n\nAny insights into this? @samsepah @sohier",
    "2176450": "My default assumption and question for organizers is whether we can expect the test set to be far more diverse (much less than 4000 rows per participant in test set)? Or was the test set generated similarly to the train set (\"Signers who communicate using American Sign Language as their primary language were recruited from across the United States. They were shipped a Pixel 4a smartphone with an installed collection app. \")?",
    "2177210": "The safe assumption I generally make is always that the test set will be independent from and contain similar variability to the train set (IID assumption)."
  },
  "source": "meta"
}