{
  "id": 204638,
  "title": "A simple way to split folds",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/204638",
  "author_name": "Gunes Evitan",
  "post_date": "2020-12-16T05:26:58.848000",
  "votes": 44,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> made two good points on how to properly split data in this notebook.</p>\n<p><a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">https://www.kaggle.com/underwearfitting/how-to-properly-split-folds</a></p>\n<p>The same result can be achieved with simple <code>sklearn.model_selection.GroupKFold</code> as well. It has stratification functionality even though it is not explicitly said. It stratifies the features when y is passed to split method. y can be a single feature or an array-like object with multiple features like in this case.</p>\n<p><code>gkf.split(df_train, df_train[targets], df_train['PatientID'])</code></p>\n<p>After the folds are created, they have equal number of samples.</p>\n<p><img src=\"https://i.ibb.co/hL2PG38/samples.jpg\" alt=\"samples\"></p>\n<p>Targets are almost equally represented in every fold.</p>\n<p><img src=\"https://i.ibb.co/Q8hHVGg/folds.jpg\" alt=\"targets\"></p>\n<p>and finally, intersection of patient ids in every fold is an empty set.</p>\n<pre><code>set.intersection(set(df_train[df_train['fold'] == 1]['PatientID']),\n                           set(df_train[df_train['fold'] == 2]['PatientID']),\n                           set(df_train[df_train['fold'] == 3]['PatientID']),\n                           set(df_train[df_train['fold'] == 4]['PatientID']),\n                           set(df_train[df_train['fold'] == 5]['PatientID']))\n</code></pre>\n<p>returns <code>set()</code>.</p>",
  "messages": [
    {
      "id": 1115225,
      "postDate": "2020-12-16T05:26:58.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> made two good points on how to properly split data in this notebook.</p>\n<p><a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">https://www.kaggle.com/underwearfitting/how-to-properly-split-folds</a></p>\n<p>The same result can be achieved with simple <code>sklearn.model_selection.GroupKFold</code> as well. It has stratification functionality even though it is not explicitly said. It stratifies the features when y is passed to split method. y can be a single feature or an array-like object with multiple features like in this case.</p>\n<p><code>gkf.split(df_train, df_train[targets], df_train['PatientID'])</code></p>\n<p>After the folds are created, they have equal number of samples.</p>\n<p><img src=\"https://i.ibb.co/hL2PG38/samples.jpg\" alt=\"samples\"></p>\n<p>Targets are almost equally represented in every fold.</p>\n<p><img src=\"https://i.ibb.co/Q8hHVGg/folds.jpg\" alt=\"targets\"></p>\n<p>and finally, intersection of patient ids in every fold is an empty set.</p>\n<pre><code>set.intersection(set(df_train[df_train['fold'] == 1]['PatientID']),\n                           set(df_train[df_train['fold'] == 2]['PatientID']),\n                           set(df_train[df_train['fold'] == 3]['PatientID']),\n                           set(df_train[df_train['fold'] == 4]['PatientID']),\n                           set(df_train[df_train['fold'] == 5]['PatientID']))\n</code></pre>\n<p>returns <code>set()</code>.</p>",
      "rawMarkdown": "@underwearfitting made two good points on how to properly split data in this notebook.\n\nhttps://www.kaggle.com/underwearfitting/how-to-properly-split-folds\n\nThe same result can be achieved with simple `sklearn.model_selection.GroupKFold` as well. It has stratification functionality even though it is not explicitly said. It stratifies the features when y is passed to split method. y can be a single feature or an array-like object with multiple features like in this case.\n\n`gkf.split(df_train, df_train[targets], df_train['PatientID'])`\n\nAfter the folds are created, they have equal number of samples.\n\n![samples](https://i.ibb.co/hL2PG38/samples.jpg)\n\nTargets are almost equally represented in every fold.\n\n![targets](https://i.ibb.co/Q8hHVGg/folds.jpg)\n\nand finally, intersection of patient ids in every fold is an empty set.\n\n```\nset.intersection(set(df_train[df_train['fold'] == 1]['PatientID']),\n                           set(df_train[df_train['fold'] == 2]['PatientID']),\n                           set(df_train[df_train['fold'] == 3]['PatientID']),\n                           set(df_train[df_train['fold'] == 4]['PatientID']),\n                           set(df_train[df_train['fold'] == 5]['PatientID']))\n```\nreturns `set()`.\n\n",
      "votes": 44
    },
    {
      "id": 1138156,
      "postDate": "2021-01-04T12:59:21.027Z",
      "content": "<p>I think the <code>y</code> argument is only there for API completeness and doesn't do anything . The stratification we see is because of the large number of samples (law of large numbers), not because <code>GroupKFold</code> is deciding to balance the targets. You could probably test this by performing splits on a small sample of the dataset (or even simpler set <code>y=None</code> and see if you get the same result)</p>\n<p><code>GroupKFold</code> is fully deterministic, which is why it doesn't require a <code>random_state</code> in the same way <code>StratifiedKFold</code> does. In the CHAMPS Molecules competition, this came in handy for teaming up as most people were already using the same folds 😀</p>\n<p>I only just started looking at the data, but I agree with you, <code>GroupKFold</code> looks like it should be good enough to start with</p>\n<p>Edit: I tried <code>y=None</code> and <code>y=df[targets]</code> and the splits are identical</p>",
      "rawMarkdown": "I think the `y` argument is only there for API completeness and doesn't do anything . The stratification we see is because of the large number of samples (law of large numbers), not because `GroupKFold` is deciding to balance the targets. You could probably test this by performing splits on a small sample of the dataset (or even simpler set `y=None` and see if you get the same result)\n\n`GroupKFold` is fully deterministic, which is why it doesn't require a `random_state` in the same way `StratifiedKFold` does. In the CHAMPS Molecules competition, this came in handy for teaming up as most people were already using the same folds 😀\n\nI only just started looking at the data, but I agree with you, `GroupKFold` looks like it should be good enough to start with\n\nEdit: I tried `y=None` and `y=df[targets]` and the splits are identical",
      "votes": 7
    },
    {
      "id": 1116368,
      "postDate": "2020-12-17T05:37:36.620Z",
      "content": "<p>\"sklearn.model_selection.GroupKFold\" is not random.<br>\ni have a feeling that most kagglers will be using this and most of us will have exactly the same split.</p>\n<p>(making our CV results very comparable?)</p>",
      "rawMarkdown": "\"sklearn.model_selection.GroupKFold\" is not random.\ni have a feeling that most kagglers will be using this and most of us will have exactly the same split.\n\n(making our CV results very comparable?)",
      "votes": 8,
      "replies": [
        {
          "id": 1117811,
          "postDate": "2020-12-18T13:22:17.180Z",
          "content": "<p>I wonder is that ever happened in the history of Kaggle.</p>",
          "rawMarkdown": "I wonder is that ever happened in the history of Kaggle.",
          "votes": 4
        },
        {
          "id": 1137049,
          "postDate": "2021-01-03T16:21:59.267Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1115240,
      "postDate": "2020-12-16T05:49:53.403Z",
      "content": "<p>Thanks a lot for the plots. Much better than my plain notebook😃. And thanks for introducing the <code>sklearn.model_selection.GroupKFold</code>, I never knew this.</p>",
      "rawMarkdown": "Thanks a lot for the plots. Much better than my plain notebook😃. And thanks for introducing the `sklearn.model_selection.GroupKFold`, I never knew this.",
      "votes": 4,
      "replies": [
        {
          "id": 1115401,
          "postDate": "2020-12-16T09:07:43.800Z",
          "content": "<p>Here are the codes for generating the plots in case you want to add them to your notebook.</p>\n<pre><code>fig = plt.figure(figsize=(16, 6))\n\nax = sns.countplot(df_train['fold'])\n\nax.tick_params(axis='x', labelsize=20)\nax.tick_params(axis='y', labelsize=20)\nax.set_xticklabels([f'{value} ({count:,})' for value, count in df_train['fold'].value_counts().sort_index().to_dict().items()])\nax.set_xlabel('Folds', size=20, labelpad=20)\nax.set_ylabel('Samples', size=20, labelpad=20)\n\nplt.title(f'Training Set Number of Samples in Folds', size=20, pad=20)\n\nplt.show()\n</code></pre>\n<pre><code>splits = df_train.groupby('fold').sum()[targets] \\\n        .reset_index(drop=True) \\\n        .T \\\n        .rename(columns={fold - 1: fold for fold in sorted(df_train['fold'].unique())}) \\\n        .reset_index() \\\n        .rename(columns={'index': 'Target'})\n\nsplits = pd.melt(splits, id_vars=['Target'], value_name='Count')\nsplits['Total'] = splits.groupby('Target')['Count'].transform('sum')\nsplits = splits.sort_values(by=['Total', 'Target'], ascending=False).reset_index(drop=True)\nsplits['variable'] = 'Fold ' + splits['variable'].astype(str)\n\nfig = plt.figure(figsize=(16, 8), dpi=100)\n\nsns.barplot(x=splits['Count'],\n            y=splits['Target'],\n            hue=splits['variable'])\n\nplt.xlabel('')\nplt.ylabel('')\nplt.tick_params(axis='x', labelsize=15)\nplt.tick_params(axis='y', labelsize=15)\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0, prop={'size': 20})\nplt.title('Multi Label Stratified GroupKFold Target Counts', size=18, pad=18)\n\nplt.show()\n</code></pre>",
          "rawMarkdown": "Here are the codes for generating the plots in case you want to add them to your notebook.\n\n```\nfig = plt.figure(figsize=(16, 6))\n\nax = sns.countplot(df_train['fold'])\n\nax.tick_params(axis='x', labelsize=20)\nax.tick_params(axis='y', labelsize=20)\nax.set_xticklabels([f'{value} ({count:,})' for value, count in df_train['fold'].value_counts().sort_index().to_dict().items()])\nax.set_xlabel('Folds', size=20, labelpad=20)\nax.set_ylabel('Samples', size=20, labelpad=20)\n    \nplt.title(f'Training Set Number of Samples in Folds', size=20, pad=20)\n\nplt.show()\n```\n\n```\nsplits = df_train.groupby('fold').sum()[targets] \\\n        .reset_index(drop=True) \\\n        .T \\\n        .rename(columns={fold - 1: fold for fold in sorted(df_train['fold'].unique())}) \\\n        .reset_index() \\\n        .rename(columns={'index': 'Target'})\n\nsplits = pd.melt(splits, id_vars=['Target'], value_name='Count')\nsplits['Total'] = splits.groupby('Target')['Count'].transform('sum')\nsplits = splits.sort_values(by=['Total', 'Target'], ascending=False).reset_index(drop=True)\nsplits['variable'] = 'Fold ' + splits['variable'].astype(str)\n\nfig = plt.figure(figsize=(16, 8), dpi=100)\n\nsns.barplot(x=splits['Count'],\n            y=splits['Target'],\n            hue=splits['variable'])\n\nplt.xlabel('')\nplt.ylabel('')\nplt.tick_params(axis='x', labelsize=15)\nplt.tick_params(axis='y', labelsize=15)\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0, prop={'size': 20})\nplt.title('Multi Label Stratified GroupKFold Target Counts', size=18, pad=18)\n\nplt.show()\n```",
          "votes": 4
        },
        {
          "id": 1115795,
          "postDate": "2020-12-16T15:25:45.887Z",
          "content": "<p>Same could have been done in the MoA competition…</p>",
          "rawMarkdown": "Same could have been done in the MoA competition...",
          "votes": 1
        }
      ]
    },
    {
      "id": 1213521,
      "postDate": "2021-02-22T07:10:57.940Z",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  dint we try this <br>\n<a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">https://github.com/trent-b/iterative-stratification</a></p>",
      "rawMarkdown": "@gunesevitan  dint we try this \nhttps://github.com/trent-b/iterative-stratification"
    },
    {
      "id": 1137046,
      "postDate": "2021-01-03T16:17:02.850Z",
      "content": "<p>Thanks for this topic. I wasn't knowing that GroupKfold stratifies the multilabel targets without having any explicit parameter for it. </p>",
      "rawMarkdown": "Thanks for this topic. I wasn't knowing that GroupKfold stratifies the multilabel targets without having any explicit parameter for it. "
    }
  ],
  "comments": [
    {
      "id": 1138156,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-01-04T12:59:21.027000",
      "content": "<p>I think the <code>y</code> argument is only there for API completeness and doesn't do anything . The stratification we see is because of the large number of samples (law of large numbers), not because <code>GroupKFold</code> is deciding to balance the targets. You could probably test this by performing splits on a small sample of the dataset (or even simpler set <code>y=None</code> and see if you get the same result)</p>\n<p><code>GroupKFold</code> is fully deterministic, which is why it doesn't require a <code>random_state</code> in the same way <code>StratifiedKFold</code> does. In the CHAMPS Molecules competition, this came in handy for teaming up as most people were already using the same folds 😀</p>\n<p>I only just started looking at the data, but I agree with you, <code>GroupKFold</code> looks like it should be good enough to start with</p>\n<p>Edit: I tried <code>y=None</code> and <code>y=df[targets]</code> and the splits are identical</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1116368,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-17T05:37:36.620000",
      "content": "<p>\"sklearn.model_selection.GroupKFold\" is not random.<br>\ni have a feeling that most kagglers will be using this and most of us will have exactly the same split.</p>\n<p>(making our CV results very comparable?)</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1117811,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-12-18T13:22:17.180000",
          "content": "<p>I wonder is that ever happened in the history of Kaggle.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1137049,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-03T16:21:59.267000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1115240,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-12-16T05:49:53.403000",
      "content": "<p>Thanks a lot for the plots. Much better than my plain notebook😃. And thanks for introducing the <code>sklearn.model_selection.GroupKFold</code>, I never knew this.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1115401,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-12-16T09:07:43.800000",
          "content": "<p>Here are the codes for generating the plots in case you want to add them to your notebook.</p>\n<pre><code>fig = plt.figure(figsize=(16, 6))\n\nax = sns.countplot(df_train['fold'])\n\nax.tick_params(axis='x', labelsize=20)\nax.tick_params(axis='y', labelsize=20)\nax.set_xticklabels([f'{value} ({count:,})' for value, count in df_train['fold'].value_counts().sort_index().to_dict().items()])\nax.set_xlabel('Folds', size=20, labelpad=20)\nax.set_ylabel('Samples', size=20, labelpad=20)\n\nplt.title(f'Training Set Number of Samples in Folds', size=20, pad=20)\n\nplt.show()\n</code></pre>\n<pre><code>splits = df_train.groupby('fold').sum()[targets] \\\n        .reset_index(drop=True) \\\n        .T \\\n        .rename(columns={fold - 1: fold for fold in sorted(df_train['fold'].unique())}) \\\n        .reset_index() \\\n        .rename(columns={'index': 'Target'})\n\nsplits = pd.melt(splits, id_vars=['Target'], value_name='Count')\nsplits['Total'] = splits.groupby('Target')['Count'].transform('sum')\nsplits = splits.sort_values(by=['Total', 'Target'], ascending=False).reset_index(drop=True)\nsplits['variable'] = 'Fold ' + splits['variable'].astype(str)\n\nfig = plt.figure(figsize=(16, 8), dpi=100)\n\nsns.barplot(x=splits['Count'],\n            y=splits['Target'],\n            hue=splits['variable'])\n\nplt.xlabel('')\nplt.ylabel('')\nplt.tick_params(axis='x', labelsize=15)\nplt.tick_params(axis='y', labelsize=15)\nplt.legend(bbox_to_anchor=(1.05, 1), loc=2, borderaxespad=0, prop={'size': 20})\nplt.title('Multi Label Stratified GroupKFold Target Counts', size=18, pad=18)\n\nplt.show()\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1115795,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-16T15:25:45.887000",
          "content": "<p>Same could have been done in the MoA competition…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1213521,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2021-02-22T07:10:57.940000",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>  dint we try this <br>\n<a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">https://github.com/trent-b/iterative-stratification</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1137046,
      "author_name": "Prateek Mishra",
      "author_url": "",
      "post_date": "2021-01-03T16:17:02.850000",
      "content": "<p>Thanks for this topic. I wasn't knowing that GroupKfold stratifies the multilabel targets without having any explicit parameter for it. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1115225": "@underwearfitting made two good points on how to properly split data in this notebook.\n\nhttps://www.kaggle.com/underwearfitting/how-to-properly-split-folds\n\nThe same result can be achieved with simple `sklearn.model_selection.GroupKFold` as well. It has stratification functionality even though it is not explicitly said. It stratifies the features when y is passed to split method. y can be a single feature or an array-like object with multiple features like in this case.\n\n`gkf.split(df_train, df_train[targets], df_train['PatientID'])`\n\nAfter the folds are created, they have equal number of samples.\n\n![samples](https://i.ibb.co/hL2PG38/samples.jpg)\n\nTargets are almost equally represented in every fold.\n\n![targets](https://i.ibb.co/Q8hHVGg/folds.jpg)\n\nand finally, intersection of patient ids in every fold is an empty set.\n\n```\nset.intersection(set(df_train[df_train['fold'] == 1]['PatientID']),\n                           set(df_train[df_train['fold'] == 2]['PatientID']),\n                           set(df_train[df_train['fold'] == 3]['PatientID']),\n                           set(df_train[df_train['fold'] == 4]['PatientID']),\n                           set(df_train[df_train['fold'] == 5]['PatientID']))\n```\nreturns `set()`.\n\n",
    "1138156": "I think the `y` argument is only there for API completeness and doesn't do anything . The stratification we see is because of the large number of samples (law of large numbers), not because `GroupKFold` is deciding to balance the targets. You could probably test this by performing splits on a small sample of the dataset (or even simpler set `y=None` and see if you get the same result)\n\n`GroupKFold` is fully deterministic, which is why it doesn't require a `random_state` in the same way `StratifiedKFold` does. In the CHAMPS Molecules competition, this came in handy for teaming up as most people were already using the same folds 😀\n\nI only just started looking at the data, but I agree with you, `GroupKFold` looks like it should be good enough to start with\n\nEdit: I tried `y=None` and `y=df[targets]` and the splits are identical",
    "1116368": "\"sklearn.model_selection.GroupKFold\" is not random.\ni have a feeling that most kagglers will be using this and most of us will have exactly the same split.\n\n(making our CV results very comparable?)",
    "1115240": "Thanks a lot for the plots. Much better than my plain notebook😃. And thanks for introducing the `sklearn.model_selection.GroupKFold`, I never knew this.",
    "1213521": "@gunesevitan  dint we try this \nhttps://github.com/trent-b/iterative-stratification",
    "1137046": "Thanks for this topic. I wasn't knowing that GroupKfold stratifies the multilabel targets without having any explicit parameter for it. "
  }
}