{
  "id": 360070,
  "title": "Cross-Validation, GroupKFold (Implementation)",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/360070",
  "author_name": "",
  "post_date": "2022-10-14T20:58:05.365284900Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear kagglers,</p>\n<p><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714\" target=\"_blank\">It has been posted by Alexandr here</a> that we are to avoid overfitting and data leakage. For this purpose I implemented this code a while ago. There is no reason to keep it private. Hope you find it useful, happy training!</p>\n<pre><code>from sklearn.model_selection import GroupKFold\nimport string\n\ndef create_kf_groups(game_num, n_folds=10):\n    \"\"\"\n    Creates groups for the GroupKFold based on the game_num.\n    We don't want to let train data spill over the validation set.\n    :param game_num: pandas.core.series.Series\n    :param n_folds: int (default 10)\n    :return: pandas.core.series.Series (alpha list)\n    \"\"\"\n    count_unique_games = game_num.unique().shape[0]\n    group_size = count_unique_games // n_folds\n    bins = [group_size * i for i in range(n_folds)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:n_folds]]\n    kf_groups = pd.cut(game_num, bins=bins, labels=labels)\n    return kf_groups\n\n\nn_folds = 10\ngame_num = train.game_num\ngroups = create_kf_groups(game_num, n_folds=n_folds)\ngkf = GroupKFold(n_splits=n_folds)\nfor fold, (idx_train, idx_val) in enumerate(gkf.split(X, y, groups=groups))\n    X_train, y_train = X[idx_train, :], y[idx_train, :]\n    X_valid, y_valid = X[idx_val, :], y[idx_val, :]\n</code></pre>",
  "messages": [
    {
      "id": "1987832",
      "postDate": "10/14/2022 20:58:05",
      "content": "<p>Dear kagglers,</p>\n<p><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714\" target=\"_blank\">It has been posted by Alexandr here</a> that we are to avoid overfitting and data leakage. For this purpose I implemented this code a while ago. There is no reason to keep it private. Hope you find it useful, happy training!</p>\n<pre><code>from sklearn.model_selection import GroupKFold\nimport string\n\ndef create_kf_groups(game_num, n_folds=10):\n    \"\"\"\n    Creates groups for the GroupKFold based on the game_num.\n    We don't want to let train data spill over the validation set.\n    :param game_num: pandas.core.series.Series\n    :param n_folds: int (default 10)\n    :return: pandas.core.series.Series (alpha list)\n    \"\"\"\n    count_unique_games = game_num.unique().shape[0]\n    group_size = count_unique_games // n_folds\n    bins = [group_size * i for i in range(n_folds)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:n_folds]]\n    kf_groups = pd.cut(game_num, bins=bins, labels=labels)\n    return kf_groups\n\n\nn_folds = 10\ngame_num = train.game_num\ngroups = create_kf_groups(game_num, n_folds=n_folds)\ngkf = GroupKFold(n_splits=n_folds)\nfor fold, (idx_train, idx_val) in enumerate(gkf.split(X, y, groups=groups))\n    X_train, y_train = X[idx_train, :], y[idx_train, :]\n    X_valid, y_valid = X[idx_val, :], y[idx_val, :]\n</code></pre>",
      "rawMarkdown": "Dear kagglers,\n\n[It has been posted by Alexandr here](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714) that we are to avoid overfitting and data leakage. For this purpose I implemented this code a while ago. There is no reason to keep it private. Hope you find it useful, happy training!\n\n```\nfrom sklearn.model_selection import GroupKFold\nimport string\n\ndef create_kf_groups(game_num, n_folds=10):\n    \"\"\"\n    Creates groups for the GroupKFold based on the game_num.\n    We don't want to let train data spill over the validation set.\n    :param game_num: pandas.core.series.Series\n    :param n_folds: int (default 10)\n    :return: pandas.core.series.Series (alpha list)\n    \"\"\"\n    count_unique_games = game_num.unique().shape[0]\n    group_size = count_unique_games // n_folds\n    bins = [group_size * i for i in range(n_folds)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:n_folds]]\n    kf_groups = pd.cut(game_num, bins=bins, labels=labels)\n    return kf_groups\n\n\nn_folds = 10\ngame_num = train.game_num\ngroups = create_kf_groups(game_num, n_folds=n_folds)\ngkf = GroupKFold(n_splits=n_folds)\nfor fold, (idx_train, idx_val) in enumerate(gkf.split(X, y, groups=groups))\n    X_train, y_train = X[idx_train, :], y[idx_train, :]\n    X_valid, y_valid = X[idx_val, :], y[idx_val, :]\n```",
      "votes": null
    },
    {
      "id": "1987885",
      "postDate": "10/14/2022 21:47:04",
      "content": "<p>Thanks for sharing this piece of code, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. Very instructive!</p>\n<p>For beginners (<em>like me</em>) in the early stage of modeling. I'm using this <a href=\"https://stackoverflow.com/a/61337958\" target=\"_blank\">implementation</a> (kudos to <code>yatu</code>) to split data between train and test:</p>\n<pre><code>from sklearn.model_selection import GroupShuffleSplit\n\ngsp = GroupShuffleSplit(n_splits=2, test_size=0.3, random_state=777)\ntrain_index, test_index = next(gsp.split(df, groups=df.index.get_level_values(\"game_num\")))\n\n# since GroupShuffleSplit allows only 2 or more splits, I'm using next() to get the second train/test split\n\nX_train = df[FEATURES].iloc[train_index]\ny_train = df[TARGET].iloc[train_index]\n\nX_test = df[FEATURES].iloc[test_index]\ny_test = df[TARGET].iloc[test_index]\n</code></pre>",
      "rawMarkdown": "Thanks for sharing this piece of code, @sergiosaharovskiy. Very instructive!\n\nFor beginners (_like me_) in the early stage of modeling. I'm using this [implementation](https://stackoverflow.com/a/61337958) (kudos to `yatu`) to split data between train and test:\n\n```python\nfrom sklearn.model_selection import GroupShuffleSplit\n\ngsp = GroupShuffleSplit(n_splits=2, test_size=0.3, random_state=777)\ntrain_index, test_index = next(gsp.split(df, groups=df.index.get_level_values(\"game_num\")))\n\n# since GroupShuffleSplit allows only 2 or more splits, I'm using next() to get the second train/test split\n\nX_train = df[FEATURES].iloc[train_index]\ny_train = df[TARGET].iloc[train_index]\n\nX_test = df[FEATURES].iloc[test_index]\ny_test = df[TARGET].iloc[test_index]\n```",
      "votes": null
    },
    {
      "id": "1987952",
      "postDate": "10/14/2022 23:55:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "rawMarkdown": "Thanks for sharing @sergiosaharovskiy",
      "votes": null
    },
    {
      "id": "1994865",
      "postDate": "10/19/2022 10:09:54",
      "content": "<p>Great Efforts, appreciated <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "rawMarkdown": "Great Efforts, appreciated @sergiosaharovskiy",
      "votes": null
    },
    {
      "id": "2000448",
      "postDate": "10/23/2022 10:38:12",
      "content": "<p>Thank you for your code. I've written a notebook explaining the problem, citing your code and adding my implementation that makes it parallelizable.<br>\nYou can find it here: <a href=\"https://www.kaggle.com/code/donatoriccio/why-you-need-groupkfold-on-this-dataset/notebook\" target=\"_blank\">Understanding data + why you need GroupKFold</a></p>",
      "rawMarkdown": "Thank you for your code. I've written a notebook explaining the problem, citing your code and adding my implementation that makes it parallelizable.\nYou can find it here: [Understanding data + why you need GroupKFold](https://www.kaggle.com/code/donatoriccio/why-you-need-groupkfold-on-this-dataset/notebook)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1987885,
      "author_name": "iancastillo",
      "author_url": "",
      "post_date": "10/14/2022 21:47:04",
      "content": "<p>Thanks for sharing this piece of code, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. Very instructive!</p>\n<p>For beginners (<em>like me</em>) in the early stage of modeling. I'm using this <a href=\"https://stackoverflow.com/a/61337958\" target=\"_blank\">implementation</a> (kudos to <code>yatu</code>) to split data between train and test:</p>\n<pre><code>from sklearn.model_selection import GroupShuffleSplit\n\ngsp = GroupShuffleSplit(n_splits=2, test_size=0.3, random_state=777)\ntrain_index, test_index = next(gsp.split(df, groups=df.index.get_level_values(\"game_num\")))\n\n# since GroupShuffleSplit allows only 2 or more splits, I'm using next() to get the second train/test split\n\nX_train = df[FEATURES].iloc[train_index]\ny_train = df[TARGET].iloc[train_index]\n\nX_test = df[FEATURES].iloc[test_index]\ny_test = df[TARGET].iloc[test_index]\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1987952,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "10/14/2022 23:55:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1994865,
      "author_name": "pranayteaches",
      "author_url": "",
      "post_date": "10/19/2022 10:09:54",
      "content": "<p>Great Efforts, appreciated <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2000448,
      "author_name": "donatoriccio",
      "author_url": "",
      "post_date": "10/23/2022 10:38:12",
      "content": "<p>Thank you for your code. I've written a notebook explaining the problem, citing your code and adding my implementation that makes it parallelizable.<br>\nYou can find it here: <a href=\"https://www.kaggle.com/code/donatoriccio/why-you-need-groupkfold-on-this-dataset/notebook\" target=\"_blank\">Understanding data + why you need GroupKFold</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1987832": "Dear kagglers,\n\n[It has been posted by Alexandr here](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714) that we are to avoid overfitting and data leakage. For this purpose I implemented this code a while ago. There is no reason to keep it private. Hope you find it useful, happy training!\n\n```\nfrom sklearn.model_selection import GroupKFold\nimport string\n\ndef create_kf_groups(game_num, n_folds=10):\n    \"\"\"\n    Creates groups for the GroupKFold based on the game_num.\n    We don't want to let train data spill over the validation set.\n    :param game_num: pandas.core.series.Series\n    :param n_folds: int (default 10)\n    :return: pandas.core.series.Series (alpha list)\n    \"\"\"\n    count_unique_games = game_num.unique().shape[0]\n    group_size = count_unique_games // n_folds\n    bins = [group_size * i for i in range(n_folds)]\n    bins += [count_unique_games]\n    labels = [letter for letter in string.ascii_letters[:n_folds]]\n    kf_groups = pd.cut(game_num, bins=bins, labels=labels)\n    return kf_groups\n\n\nn_folds = 10\ngame_num = train.game_num\ngroups = create_kf_groups(game_num, n_folds=n_folds)\ngkf = GroupKFold(n_splits=n_folds)\nfor fold, (idx_train, idx_val) in enumerate(gkf.split(X, y, groups=groups))\n    X_train, y_train = X[idx_train, :], y[idx_train, :]\n    X_valid, y_valid = X[idx_val, :], y[idx_val, :]\n```",
    "1987885": "Thanks for sharing this piece of code, @sergiosaharovskiy. Very instructive!\n\nFor beginners (_like me_) in the early stage of modeling. I'm using this [implementation](https://stackoverflow.com/a/61337958) (kudos to `yatu`) to split data between train and test:\n\n```python\nfrom sklearn.model_selection import GroupShuffleSplit\n\ngsp = GroupShuffleSplit(n_splits=2, test_size=0.3, random_state=777)\ntrain_index, test_index = next(gsp.split(df, groups=df.index.get_level_values(\"game_num\")))\n\n# since GroupShuffleSplit allows only 2 or more splits, I'm using next() to get the second train/test split\n\nX_train = df[FEATURES].iloc[train_index]\ny_train = df[TARGET].iloc[train_index]\n\nX_test = df[FEATURES].iloc[test_index]\ny_test = df[TARGET].iloc[test_index]\n```",
    "1987952": "Thanks for sharing @sergiosaharovskiy",
    "1994865": "Great Efforts, appreciated @sergiosaharovskiy",
    "2000448": "Thank you for your code. I've written a notebook explaining the problem, citing your code and adding my implementation that makes it parallelizable.\nYou can find it here: [Understanding data + why you need GroupKFold](https://www.kaggle.com/code/donatoriccio/why-you-need-groupkfold-on-this-dataset/notebook)"
  },
  "source": "meta"
}