{
  "id": 468592,
  "title": "A pitfall in public best notebook",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/468592",
  "author_name": "",
  "post_date": "2024-01-17T07:27:10.233992700Z",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I want to thank <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for sharing his amazing pipeline <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/#Training\" target=\"_blank\">here</a>. However, there is a minor issue on that notebook that can cause inconsistent results.</p>\n<p>The problem arises from this part<br>\n<code>train = train.groupby(\"spectrogram_id\").head(1).reset_index(drop=True)</code></p>\n<p>This line is executed before creating the folds. When someone wants to increase sampling per spectrogram_id, they will change the validation folds too.</p>\n<p>Folds should be created first on the entire dataset and then sampling should be done within each folds training portion. This way ensures the validation sets are fixed so you can get comparable results.</p>",
  "messages": [
    {
      "id": "2605695",
      "postDate": "01/17/2024 07:27:10",
      "content": "<p>I want to thank <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for sharing his amazing pipeline <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/#Training\" target=\"_blank\">here</a>. However, there is a minor issue on that notebook that can cause inconsistent results.</p>\n<p>The problem arises from this part<br>\n<code>train = train.groupby(\"spectrogram_id\").head(1).reset_index(drop=True)</code></p>\n<p>This line is executed before creating the folds. When someone wants to increase sampling per spectrogram_id, they will change the validation folds too.</p>\n<p>Folds should be created first on the entire dataset and then sampling should be done within each folds training portion. This way ensures the validation sets are fixed so you can get comparable results.</p>",
      "rawMarkdown": "I want to thank @ttahara for sharing his amazing pipeline [here](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/#Training). However, there is a minor issue on that notebook that can cause inconsistent results.\n\nThe problem arises from this part\n`train = train.groupby(\"spectrogram_id\").head(1).reset_index(drop=True)`\n\nThis line is executed before creating the folds. When someone wants to increase sampling per spectrogram_id, they will change the validation folds too.\n\nFolds should be created first on the entire dataset and then sampling should be done within each folds training portion. This way ensures the validation sets are fixed so you can get comparable results.",
      "votes": null
    },
    {
      "id": "2605809",
      "postDate": "01/17/2024 08:55:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> .. I was wondering why it is like this ..</p>",
      "rawMarkdown": "Thanks @gunesevitan .. I was wondering why it is like this ..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2605809,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "01/17/2024 08:55:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> .. I was wondering why it is like this ..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2605695": "I want to thank @ttahara for sharing his amazing pipeline [here](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/#Training). However, there is a minor issue on that notebook that can cause inconsistent results.\n\nThe problem arises from this part\n`train = train.groupby(\"spectrogram_id\").head(1).reset_index(drop=True)`\n\nThis line is executed before creating the folds. When someone wants to increase sampling per spectrogram_id, they will change the validation folds too.\n\nFolds should be created first on the entire dataset and then sampling should be done within each folds training portion. This way ensures the validation sets are fixed so you can get comparable results.",
    "2605809": "Thanks @gunesevitan .. I was wondering why it is like this .."
  },
  "source": "meta"
}