{
  "id": 467612,
  "title": "Train Split CV",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467612",
  "author_name": "",
  "post_date": "2024-01-13T06:56:01.578690Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F23b13419f9702d79ffc55b93014df8c0%2F1.png?generation=1705128282771189&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F99d93f74896c304de2cac5f8f7af26fd%2F2.png?generation=1705128295000598&amp;alt=media\"></p>\n<p><a href=\"https://www.kaggle.com/code/seshurajup/eegs-train-split-cv\" target=\"_blank\">Notebook - Eegs Train Split (CV)</a><br>\n<a href=\"https://www.kaggle.com/datasets/seshurajup/eegs-train-split/data\" target=\"_blank\">Dataset - Eegs Train Split</a></p>\n<h2>Other than swapping patient_id or using different seeds, is any good techniques to balance the folds? I used swap approach</h2>\n<h2>with GroupKFold, it is better</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Ff8150501f014970990bb78b727bc25f7%2FScreenshot%202024-01-14%20at%208.24.13AM.png?generation=1705200901861322&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2599831",
      "postDate": "01/13/2024 06:56:01",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F23b13419f9702d79ffc55b93014df8c0%2F1.png?generation=1705128282771189&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F99d93f74896c304de2cac5f8f7af26fd%2F2.png?generation=1705128295000598&amp;alt=media\"></p>\n<p><a href=\"https://www.kaggle.com/code/seshurajup/eegs-train-split-cv\" target=\"_blank\">Notebook - Eegs Train Split (CV)</a><br>\n<a href=\"https://www.kaggle.com/datasets/seshurajup/eegs-train-split/data\" target=\"_blank\">Dataset - Eegs Train Split</a></p>\n<h2>Other than swapping patient_id or using different seeds, is any good techniques to balance the folds? I used swap approach</h2>\n<h2>with GroupKFold, it is better</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Ff8150501f014970990bb78b727bc25f7%2FScreenshot%202024-01-14%20at%208.24.13AM.png?generation=1705200901861322&amp;alt=media\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F23b13419f9702d79ffc55b93014df8c0%2F1.png?generation=1705128282771189&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F99d93f74896c304de2cac5f8f7af26fd%2F2.png?generation=1705128295000598&alt=media)\n\n[Notebook - Eegs Train Split (CV)](https://www.kaggle.com/code/seshurajup/eegs-train-split-cv)\n[Dataset - Eegs Train Split](https://www.kaggle.com/datasets/seshurajup/eegs-train-split/data)\n\n## Other than swapping patient_id or using different seeds, is any good techniques to balance the folds? I used swap approach\n\n## with GroupKFold, it is better \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Ff8150501f014970990bb78b727bc25f7%2FScreenshot%202024-01-14%20at%208.24.13AM.png?generation=1705200901861322&alt=media)",
      "votes": null
    },
    {
      "id": "2626846",
      "postDate": "01/30/2024 10:05:53",
      "content": "<p>Nice visualization, I see that the ratio of each category is basically consistent using gkf</p>",
      "rawMarkdown": "Nice visualization, I see that the ratio of each category is basically consistent using gkf",
      "votes": null
    },
    {
      "id": "2626978",
      "postDate": "01/30/2024 12:06:38",
      "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> yes compare to sgkf </p>",
      "rawMarkdown": "gentlezdh yes compare to sgkf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2626846,
      "author_name": "gentlezdh",
      "author_url": "",
      "post_date": "01/30/2024 10:05:53",
      "content": "<p>Nice visualization, I see that the ratio of each category is basically consistent using gkf</p>",
      "votes": null,
      "replies": [
        {
          "id": 2626978,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/30/2024 12:06:38",
          "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> yes compare to sgkf </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2599831": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F23b13419f9702d79ffc55b93014df8c0%2F1.png?generation=1705128282771189&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F99d93f74896c304de2cac5f8f7af26fd%2F2.png?generation=1705128295000598&alt=media)\n\n[Notebook - Eegs Train Split (CV)](https://www.kaggle.com/code/seshurajup/eegs-train-split-cv)\n[Dataset - Eegs Train Split](https://www.kaggle.com/datasets/seshurajup/eegs-train-split/data)\n\n## Other than swapping patient_id or using different seeds, is any good techniques to balance the folds? I used swap approach\n\n## with GroupKFold, it is better \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Ff8150501f014970990bb78b727bc25f7%2FScreenshot%202024-01-14%20at%208.24.13AM.png?generation=1705200901861322&alt=media)",
    "2626846": "Nice visualization, I see that the ratio of each category is basically consistent using gkf",
    "2626978": "gentlezdh yes compare to sgkf"
  },
  "source": "meta"
}