{
  "id": 294084,
  "title": "Balanced Fold Splitting",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/294084",
  "author_name": "",
  "post_date": "2021-12-08T10:57:29.842135500Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the current public notebooks, most of people use GroupKFold to split the train/valid data.<br>\nIt assures no leak happens in validation dataset. But you also have to consider sufficient #COTSs appears in both train/validation sets for each fold. Naively applying GroupKFold may cause unbalanced folds. </p>\n<p>So I came up with the idea splitting each sequence to subsequence, to get more balanced fold.<br>\nThe algorithms is simple, but it mitigates the unbalance -- decreasing standard deviation of #COTSs and #COTSs/frame between folds.</p>\n<p>The algorithm is naive, but the idea should be applied to other more sophisticated fold splitting algorithms.</p>\n<p>The code is here:<br>\n<a href=\"https://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118\" target=\"_blank\">https://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118</a></p>",
  "messages": [
    {
      "id": "1611880",
      "postDate": "12/08/2021 10:57:29",
      "content": "<p>In the current public notebooks, most of people use GroupKFold to split the train/valid data.<br>\nIt assures no leak happens in validation dataset. But you also have to consider sufficient #COTSs appears in both train/validation sets for each fold. Naively applying GroupKFold may cause unbalanced folds. </p>\n<p>So I came up with the idea splitting each sequence to subsequence, to get more balanced fold.<br>\nThe algorithms is simple, but it mitigates the unbalance -- decreasing standard deviation of #COTSs and #COTSs/frame between folds.</p>\n<p>The algorithm is naive, but the idea should be applied to other more sophisticated fold splitting algorithms.</p>\n<p>The code is here:<br>\n<a href=\"https://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118\" target=\"_blank\">https://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118</a></p>",
      "rawMarkdown": "In the current public notebooks, most of people use GroupKFold to split the train/valid data.\nIt assures no leak happens in validation dataset. But you also have to consider sufficient #COTSs appears in both train/validation sets for each fold. Naively applying GroupKFold may cause unbalanced folds. \n\nSo I came up with the idea splitting each sequence to subsequence, to get more balanced fold.\nThe algorithms is simple, but it mitigates the unbalance -- decreasing standard deviation of #COTSs and #COTSs/frame between folds.\n\nThe algorithm is naive, but the idea should be applied to other more sophisticated fold splitting algorithms.\n\nThe code is here:\nhttps://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118",
      "votes": null
    },
    {
      "id": "1617262",
      "postDate": "12/14/2021 01:58:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, Do you have a good result with new split algorithm than groupkfold ? <br>\nHas the gap between CV and LB been reduced with new algorithm ? </p>",
      "rawMarkdown": "Hi @tatamikenn, Do you have a good result with new split algorithm than groupkfold ? \nHas the gap between CV and LB been reduced with new algorithm ?",
      "votes": null
    },
    {
      "id": "1617360",
      "postDate": "12/14/2021 03:46:49",
      "content": "<p><a href=\"https://www.kaggle.com/seongwook93\" target=\"_blank\">@seongwook93</a> </p>\n<p>Personally I didn't test. But someone reported CV/LB was correlated well using a similar split [1] (the one which <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> proposed[2]. It is almost identical to the approach except for the fold allocation algorithm).</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences</a></li>\n</ul>",
      "rawMarkdown": "seongwook93 \n\nPersonally I didn't test. But someone reported CV/LB was correlated well using a similar split [1] (the one which @julian3833 proposed[2]. It is almost identical to the approach except for the fold allocation algorithm).\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788\n* [2] https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
      "votes": null
    },
    {
      "id": "1617427",
      "postDate": "12/14/2021 04:39:35",
      "content": "<p>Thank you! </p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1617262,
      "author_name": "seongwook93",
      "author_url": "",
      "post_date": "12/14/2021 01:58:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, Do you have a good result with new split algorithm than groupkfold ? <br>\nHas the gap between CV and LB been reduced with new algorithm ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1617360,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "12/14/2021 03:46:49",
          "content": "<p><a href=\"https://www.kaggle.com/seongwook93\" target=\"_blank\">@seongwook93</a> </p>\n<p>Personally I didn't test. But someone reported CV/LB was correlated well using a similar split [1] (the one which <a href=\"https://www.kaggle.com/julian3833\" target=\"_blank\">@julian3833</a> proposed[2]. It is almost identical to the approach except for the fold allocation algorithm).</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1617427,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "12/14/2021 04:39:35",
          "content": "<p>Thank you! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1611880": "In the current public notebooks, most of people use GroupKFold to split the train/valid data.\nIt assures no leak happens in validation dataset. But you also have to consider sufficient #COTSs appears in both train/validation sets for each fold. Naively applying GroupKFold may cause unbalanced folds. \n\nSo I came up with the idea splitting each sequence to subsequence, to get more balanced fold.\nThe algorithms is simple, but it mitigates the unbalance -- decreasing standard deviation of #COTSs and #COTSs/frame between folds.\n\nThe algorithm is naive, but the idea should be applied to other more sophisticated fold splitting algorithms.\n\nThe code is here:\nhttps://www.kaggle.com/tatamikenn/a-balanced-fold-splitting-algorithm?scriptVersionId=81854118",
    "1617262": "Hi @tatamikenn, Do you have a good result with new split algorithm than groupkfold ? \nHas the gap between CV and LB been reduced with new algorithm ?",
    "1617360": "seongwook93 \n\nPersonally I didn't test. But someone reported CV/LB was correlated well using a similar split [1] (the one which @julian3833 proposed[2]. It is almost identical to the approach except for the fold allocation algorithm).\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/290757#1612788\n* [2] https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
    "1617427": "Thank you!"
  },
  "source": "meta"
}