{
  "id": 171514,
  "title": "Scikit-learn StratifiedGroupKFold",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171514",
  "author_name": "",
  "post_date": "2020-08-01T07:11:41.719106600Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This is the official code (work in progress, not yet in production as far as I know), of Scikit's <code>StratifiedGroupKFold</code>.  So some of you have been using <code>StratifiedKFold</code>, or <code>GroupKFold</code>, well now you have both!</p>\n\n<p>Save file as <code>split.py</code> and then add to your code:</p>\n\n<p>```\nfrom split import StratifiedGroupKFold, RepeatedStratifiedGroupKFold</p>\n\n<p>```</p>\n\n<p>Share how you are using it.  In my code, I use <code>patient_id</code> as a group, so that way I do not have leaks between folds of the same patient.</p>\n\n<p><code>cv = StratifiedGroupKFold(n_splits=param['num_splits'], shuffle = True, random_state=867)</code></p>\n\n<p>I exclude indexes past <code>33125</code> in <code>val</code> since those are not part of the original data.</p>\n\n<p>```\nfor fold, (train_idx, val_idx) in enumerate(cv.split(X=np.zeros(len(train_df)), \n                                                       y=train_df['target'], \n                                                       groups=train_df['patient_id'].tolist()), 1):\n    val_idx = val_idx[val_idx &lt;= 33125] # only original 2020 data in validation</p>\n\n<p>```</p>\n\n<p>Here is the class</p>\n\n<p>```\nfrom collections import Counter, defaultdict</p>\n\n<p>import numpy as np</p>\n\n<p>from sklearn.model_selection._split import _BaseKFold, _RepeatedSplits\nfrom sklearn.utils.validation import check_random_state</p>\n\n<p>class StratifiedGroupKFold(_BaseKFold):\n    \"\"\"Stratified K-Folds iterator variant with non-overlapping groups.</p>\n\n<pre><code>This cross-validation object is a variation of StratifiedKFold that returns\nstratified folds with non-overlapping groups. The folds are made by\npreserving the percentage of samples for each class.\n\nThe same group will not appear in two different folds (the number of\ndistinct groups has to be at least equal to the number of folds).\n\nThe difference between GroupKFold and StratifiedGroupKFold is that\nthe former attempts to create balanced folds such that the number of\ndistinct groups is approximately the same in each fold, whereas\nStratifiedGroupKFold attempts to create folds which preserve the\npercentage of samples for each class.\n\nRead more in the :ref:`User Guide &lt;cross_validation&gt;`.\n\nParameters\n----------\nn_splits : int, default=5\n    Number of folds. Must be at least 2.\n\nshuffle : bool, default=False\n    Whether to shuffle each class's samples before splitting into batches.\n    Note that the samples within each split will not be shuffled.\n\nrandom_state : int or RandomState instance, default=None\n    When `shuffle` is True, `random_state` affects the ordering of the\n    indices, which controls the randomness of each fold for each class.\n    Otherwise, leave `random_state` as `None`.\n    Pass an int for reproducible output across multiple function calls.\n    See :term:`Glossary &lt;random_state&gt;`.\n\nExamples\n--------\n&gt;&gt;&gt; import numpy as np\n&gt;&gt;&gt; from sklearn.model_selection import StratifiedGroupKFold\n&gt;&gt;&gt; X = np.ones((17, 2))\n&gt;&gt;&gt; y = np.array([0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0])\n&gt;&gt;&gt; groups = np.array([1, 1, 2, 2, 3, 3, 3, 4, 5, 5, 5, 5, 6, 6, 7, 8, 8])\n&gt;&gt;&gt; cv = StratifiedGroupKFold(n_splits=3)\n&gt;&gt;&gt; for train_idxs, test_idxs in cv.split(X, y, groups):\n...     print(\"TRAIN:\", groups[train_idxs])\n...     print(\"      \", y[train_idxs])\n...     print(\" TEST:\", groups[test_idxs])\n...     print(\"      \", y[test_idxs])\nTRAIN: [2 2 4 5 5 5 5 6 6 7]\n       [1 1 1 0 0 0 0 0 0 0]\n TEST: [1 1 3 3 3 8 8]\n       [0 0 1 1 1 0 0]\nTRAIN: [1 1 3 3 3 4 5 5 5 5 8 8]\n       [0 0 1 1 1 1 0 0 0 0 0 0]\n TEST: [2 2 6 6 7]\n       [1 1 0 0 0]\nTRAIN: [1 1 2 2 3 3 3 6 6 7 8 8]\n       [0 0 1 1 1 1 1 0 0 0 0 0]\n TEST: [4 5 5 5 5]\n       [1 0 0 0 0]\n\nSee also\n--------\nStratifiedKFold: Takes class information into account to build folds which\n    retain class distributions (for binary or multiclass classification\n    tasks).\n\nGroupKFold: K-fold iterator variant with non-overlapping groups.\n\"\"\"\n\ndef __init__(self, n_splits=5, shuffle=False, random_state=None):\n    super().__init__(n_splits=n_splits, shuffle=shuffle,\n                     random_state=random_state)\n\n# Implementation based on this kaggle kernel:\n# https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\ndef _iter_test_indices(self, X, y, groups):\n    labels_num = np.max(y) + 1\n    y_counts_per_group = defaultdict(lambda: np.zeros(labels_num))\n    y_distr = Counter()\n    for label, group in zip(y, groups):\n        y_counts_per_group[group][label] += 1\n        y_distr[label] += 1\n\n    y_counts_per_fold = defaultdict(lambda: np.zeros(labels_num))\n    groups_per_fold = defaultdict(set)\n\n    groups_and_y_counts = list(y_counts_per_group.items())\n    rng = check_random_state(self.random_state)\n    if self.shuffle:\n        rng.shuffle(groups_and_y_counts)\n\n    for group, y_counts in sorted(groups_and_y_counts,\n                                  key=lambda x: -np.std(x[1])):\n        best_fold = None\n        min_eval = None\n        for i in range(self.n_splits):\n            y_counts_per_fold[i] += y_counts\n            std_per_label = []\n            for label in range(labels_num):\n                std_per_label.append(np.std(\n                    [y_counts_per_fold[j][label] / y_distr[label]\n                     for j in range(self.n_splits)]))\n            y_counts_per_fold[i] -= y_counts\n            fold_eval = np.mean(std_per_label)\n            if min_eval is None or fold_eval &lt; min_eval:\n                min_eval = fold_eval\n                best_fold = i\n        y_counts_per_fold[best_fold] += y_counts\n        groups_per_fold[best_fold].add(group)\n\n    for i in range(self.n_splits):\n        test_indices = [idx for idx, group in enumerate(groups)\n                        if group in groups_per_fold[i]]\n        yield test_indices\n</code></pre>\n\n<p>class RepeatedStratifiedGroupKFold(_RepeatedSplits):\n    \"\"\"Repeated Stratified K-Fold cross validator.</p>\n\n<pre><code>Repeats Stratified K-Fold with non-overlapping groups n times with\ndifferent randomization in each repetition.\n\nRead more in the :ref:`User Guide &lt;cross_validation&gt;`.\n\nParameters\n----------\nn_splits : int, default=5\n    Number of folds. Must be at least 2.\n\nn_repeats : int, default=10\n    Number of times cross-validator needs to be repeated.\n\nrandom_state : int or RandomState instance, default=None\n    Controls the generation of the random states for each repetition.\n    Pass an int for reproducible output across multiple function calls.\n    See :term:`Glossary &lt;random_state&gt;`.\n\nExamples\n--------\n&gt;&gt;&gt; import numpy as np\n&gt;&gt;&gt; from sklearn.model_selection import RepeatedStratifiedGroupKFold\n&gt;&gt;&gt; X = np.ones((17, 2))\n&gt;&gt;&gt; y = np.array([0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0])\n&gt;&gt;&gt; groups = np.array([1, 1, 2, 2, 3, 3, 3, 4, 5, 5, 5, 5, 6, 6, 7, 8, 8])\n&gt;&gt;&gt; cv = RepeatedStratifiedGroupKFold(n_splits=2, n_repeats=2,\n...                                   random_state=36851234)\n&gt;&gt;&gt; for train_index, test_index in cv.split(X, y, groups):\n...     print(\"TRAIN:\", groups[train_idxs])\n...     print(\"      \", y[train_idxs])\n...     print(\" TEST:\", groups[test_idxs])\n...     print(\"      \", y[test_idxs])\nTRAIN: [2 2 4 5 5 5 5 8 8]\n       [1 1 1 0 0 0 0 0 0]\n TEST: [1 1 3 3 3 6 6 7]\n       [0 0 1 1 1 0 0 0]\nTRAIN: [1 1 3 3 3 6 6 7]\n       [0 0 1 1 1 0 0 0]\n TEST: [2 2 4 5 5 5 5 8 8]\n       [1 1 1 0 0 0 0 0 0]\nTRAIN: [3 3 3 4 7 8 8]\n       [1 1 1 1 0 0 0]\n TEST: [1 1 2 2 5 5 5 5 6 6]\n       [0 0 1 1 0 0 0 0 0 0]\nTRAIN: [1 1 2 2 5 5 5 5 6 6]\n       [0 0 1 1 0 0 0 0 0 0]\n TEST: [3 3 3 4 7 8 8]\n       [1 1 1 1 0 0 0]\n\nNotes\n-----\nRandomized CV splitters may return different results for each call of\nsplit. You can make the results identical by setting `random_state`\nto an integer.\n\nSee also\n--------\nRepeatedStratifiedKFold: Repeats Stratified K-Fold n times.\n\"\"\"\n\ndef __init__(self, n_splits=5, n_repeats=10, random_state=None):\n    super().__init__(StratifiedGroupKFold, n_splits=n_splits,\n                     n_repeats=n_repeats, random_state=random_state)\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "953846",
      "postDate": "08/01/2020 07:11:41",
      "content": "<p>This is the official code (work in progress, not yet in production as far as I know), of Scikit's <code>StratifiedGroupKFold</code>.  So some of you have been using <code>StratifiedKFold</code>, or <code>GroupKFold</code>, well now you have both!</p>\n\n<p>Save file as <code>split.py</code> and then add to your code:</p>\n\n<p>```\nfrom split import StratifiedGroupKFold, RepeatedStratifiedGroupKFold</p>\n\n<p>```</p>\n\n<p>Share how you are using it.  In my code, I use <code>patient_id</code> as a group, so that way I do not have leaks between folds of the same patient.</p>\n\n<p><code>cv = StratifiedGroupKFold(n_splits=param['num_splits'], shuffle = True, random_state=867)</code></p>\n\n<p>I exclude indexes past <code>33125</code> in <code>val</code> since those are not part of the original data.</p>\n\n<p>```\nfor fold, (train_idx, val_idx) in enumerate(cv.split(X=np.zeros(len(train_df)), \n                                                       y=train_df['target'], \n                                                       groups=train_df['patient_id'].tolist()), 1):\n    val_idx = val_idx[val_idx &lt;= 33125] # only original 2020 data in validation</p>\n\n<p>```</p>\n\n<p>Here is the class</p>\n\n<p>```\nfrom collections import Counter, defaultdict</p>\n\n<p>import numpy as np</p>\n\n<p>from sklearn.model_selection._split import _BaseKFold, _RepeatedSplits\nfrom sklearn.utils.validation import check_random_state</p>\n\n<p>class StratifiedGroupKFold(_BaseKFold):\n    \"\"\"Stratified K-Folds iterator variant with non-overlapping groups.</p>\n\n<pre><code>This cross-validation object is a variation of StratifiedKFold that returns\nstratified folds with non-overlapping groups. The folds are made by\npreserving the percentage of samples for each class.\n\nThe same group will not appear in two different folds (the number of\ndistinct groups has to be at least equal to the number of folds).\n\nThe difference between GroupKFold and StratifiedGroupKFold is that\nthe former attempts to create balanced folds such that the number of\ndistinct groups is approximately the same in each fold, whereas\nStratifiedGroupKFold attempts to create folds which preserve the\npercentage of samples for each class.\n\nRead more in the :ref:`User Guide &lt;cross_validation&gt;`.\n\nParameters\n----------\nn_splits : int, default=5\n    Number of folds. Must be at least 2.\n\nshuffle : bool, default=False\n    Whether to shuffle each class's samples before splitting into batches.\n    Note that the samples within each split will not be shuffled.\n\nrandom_state : int or RandomState instance, default=None\n    When `shuffle` is True, `random_state` affects the ordering of the\n    indices, which controls the randomness of each fold for each class.\n    Otherwise, leave `random_state` as `None`.\n    Pass an int for reproducible output across multiple function calls.\n    See :term:`Glossary &lt;random_state&gt;`.\n\nExamples\n--------\n&gt;&gt;&gt; import numpy as np\n&gt;&gt;&gt; from sklearn.model_selection import StratifiedGroupKFold\n&gt;&gt;&gt; X = np.ones((17, 2))\n&gt;&gt;&gt; y = np.array([0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0])\n&gt;&gt;&gt; groups = np.array([1, 1, 2, 2, 3, 3, 3, 4, 5, 5, 5, 5, 6, 6, 7, 8, 8])\n&gt;&gt;&gt; cv = StratifiedGroupKFold(n_splits=3)\n&gt;&gt;&gt; for train_idxs, test_idxs in cv.split(X, y, groups):\n...     print(\"TRAIN:\", groups[train_idxs])\n...     print(\"      \", y[train_idxs])\n...     print(\" TEST:\", groups[test_idxs])\n...     print(\"      \", y[test_idxs])\nTRAIN: [2 2 4 5 5 5 5 6 6 7]\n       [1 1 1 0 0 0 0 0 0 0]\n TEST: [1 1 3 3 3 8 8]\n       [0 0 1 1 1 0 0]\nTRAIN: [1 1 3 3 3 4 5 5 5 5 8 8]\n       [0 0 1 1 1 1 0 0 0 0 0 0]\n TEST: [2 2 6 6 7]\n       [1 1 0 0 0]\nTRAIN: [1 1 2 2 3 3 3 6 6 7 8 8]\n       [0 0 1 1 1 1 1 0 0 0 0 0]\n TEST: [4 5 5 5 5]\n       [1 0 0 0 0]\n\nSee also\n--------\nStratifiedKFold: Takes class information into account to build folds which\n    retain class distributions (for binary or multiclass classification\n    tasks).\n\nGroupKFold: K-fold iterator variant with non-overlapping groups.\n\"\"\"\n\ndef __init__(self, n_splits=5, shuffle=False, random_state=None):\n    super().__init__(n_splits=n_splits, shuffle=shuffle,\n                     random_state=random_state)\n\n# Implementation based on this kaggle kernel:\n# https://www.kaggle.com/jakubwasikowski/stratified-group-k-fold-cross-validation\ndef _iter_test_indices(self, X, y, groups):\n    labels_num = np.max(y) + 1\n    y_counts_per_group = defaultdict(lambda: np.zeros(labels_num))\n    y_distr = Counter()\n    for label, group in zip(y, groups):\n        y_counts_per_group[group][label] += 1\n        y_distr[label] += 1\n\n    y_counts_per_fold = defaultdict(lambda: np.zeros(labels_num))\n    groups_per_fold = defaultdict(set)\n\n    groups_and_y_counts = list(y_counts_per_group.items())\n    rng = check_random_state(self.random_state)\n    if self.shuffle:\n        rng.shuffle(groups_and_y_counts)\n\n    for group, y_counts in sorted(groups_and_y_counts,\n                                  key=lambda x: -np.std(x[1])):\n        best_fold = None\n        min_eval = None\n        for i in range(self.n_splits):\n            y_counts_per_fold[i] += y_counts\n            std_per_label = []\n            for label in range(labels_num):\n                std_per_label.append(np.std(\n                    [y_counts_per_fold[j][label] / y_distr[label]\n                     for j in range(self.n_splits)]))\n            y_counts_per_fold[i] -= y_counts\n            fold_eval = np.mean(std_per_label)\n            if min_eval is None or fold_eval &lt; min_eval:\n                min_eval = fold_eval\n                best_fold = i\n        y_counts_per_fold[best_fold] += y_counts\n        groups_per_fold[best_fold].add(group)\n\n    for i in range(self.n_splits):\n        test_indices = [idx for idx, group in enumerate(groups)\n                        if group in groups_per_fold[i]]\n        yield test_indices\n</code></pre>\n\n<p>class RepeatedStratifiedGroupKFold(_RepeatedSplits):\n    \"\"\"Repeated Stratified K-Fold cross validator.</p>\n\n<pre><code>Repeats Stratified K-Fold with non-overlapping groups n times with\ndifferent randomization in each repetition.\n\nRead more in the :ref:`User Guide &lt;cross_validation&gt;`.\n\nParameters\n----------\nn_splits : int, default=5\n    Number of folds. Must be at least 2.\n\nn_repeats : int, default=10\n    Number of times cross-validator needs to be repeated.\n\nrandom_state : int or RandomState instance, default=None\n    Controls the generation of the random states for each repetition.\n    Pass an int for reproducible output across multiple function calls.\n    See :term:`Glossary &lt;random_state&gt;`.\n\nExamples\n--------\n&gt;&gt;&gt; import numpy as np\n&gt;&gt;&gt; from sklearn.model_selection import RepeatedStratifiedGroupKFold\n&gt;&gt;&gt; X = np.ones((17, 2))\n&gt;&gt;&gt; y = np.array([0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0])\n&gt;&gt;&gt; groups = np.array([1, 1, 2, 2, 3, 3, 3, 4, 5, 5, 5, 5, 6, 6, 7, 8, 8])\n&gt;&gt;&gt; cv = RepeatedStratifiedGroupKFold(n_splits=2, n_repeats=2,\n...                                   random_state=36851234)\n&gt;&gt;&gt; for train_index, test_index in cv.split(X, y, groups):\n...     print(\"TRAIN:\", groups[train_idxs])\n...     print(\"      \", y[train_idxs])\n...     print(\" TEST:\", groups[test_idxs])\n...     print(\"      \", y[test_idxs])\nTRAIN: [2 2 4 5 5 5 5 8 8]\n       [1 1 1 0 0 0 0 0 0]\n TEST: [1 1 3 3 3 6 6 7]\n       [0 0 1 1 1 0 0 0]\nTRAIN: [1 1 3 3 3 6 6 7]\n       [0 0 1 1 1 0 0 0]\n TEST: [2 2 4 5 5 5 5 8 8]\n       [1 1 1 0 0 0 0 0 0]\nTRAIN: [3 3 3 4 7 8 8]\n       [1 1 1 1 0 0 0]\n TEST: [1 1 2 2 5 5 5 5 6 6]\n       [0 0 1 1 0 0 0 0 0 0]\nTRAIN: [1 1 2 2 5 5 5 5 6 6]\n       [0 0 1 1 0 0 0 0 0 0]\n TEST: [3 3 3 4 7 8 8]\n       [1 1 1 1 0 0 0]\n\nNotes\n-----\nRandomized CV splitters may return different results for each call of\nsplit. You can make the results identical by setting `random_state`\nto an integer.\n\nSee also\n--------\nRepeatedStratifiedKFold: Repeats Stratified K-Fold n times.\n\"\"\"\n\ndef __init__(self, n_splits=5, n_repeats=10, random_state=None):\n    super().__init__(StratifiedGroupKFold, n_splits=n_splits,\n                     n_repeats=n_repeats, random_state=random_state)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "This is the official code (work in progress, not yet in production as far as I know), of Scikit's `StratifiedGroupKFold`.  So some of you have been using `StratifiedKFold`, or `GroupKFold`, well now you have both!\n\nSave file as `split.py` and then add to your code:\n\n```\nfrom split import StratifiedGroupKFold, RepeatedStratifiedGroupKFold\n\n```\n\nShare how you are using it.  In my code, I use `patient_id` as a group, so that way I do not have leaks between folds of the same patient.\n\n`cv = StratifiedGroupKFold(n_splits=param['num_splits'], shuffle = True, random_state=867)`\n\nI exclude indexes past `33125` in `val` since those are not part of the original data.\n\n```\nfor fold, (train_idx, val_idx) in enumerate(cv.split(X=np.zeros(len(train_df)), \n                                                       y=train_df['target'], \n                                                       groups=train_df['patient_id'].tolist()), 1):\n    val_idx = val_idx[val_idx &lt;= 33125] # only original 2020 data in validation\n\n```\n\nHere is the class\n\n\n```\nfrom collections import Counter, defaultdict\n\nimport numpy as np\n\nfrom sklearn.model_selection._split import _BaseKFold, _RepeatedSplits\nfrom sklearn.utils.validation import check_random_state\n\nclass StratifiedGroupKFold(_BaseKFold):\n    \"\"\"Stratified K-Folds iterator variant with non-overlapping groups.\n\n\n    This cross-validation object is a variation of StratifiedKFold that returns\n    stratified folds with non-overlapping groups. The folds are made by\n    preserving the percentage of samples for each class.\n\n    The same group will not appear in two different folds (the number of\n    distinct groups has to be at least equal to the number of folds).\n\n    The difference between GroupKFold and StratifiedGroupKFold is that\n    the former attempts to create balanced folds such that the number of\n    distinct groups is approximately the same in each fold, whereas\n    StratifiedGroupKFold attempts to create folds which preserve the\n    percentage of samples for each class.\n\n    Read more in the :ref:`User Guide",
      "votes": null
    },
    {
      "id": "954674",
      "postDate": "08/02/2020 01:51:37",
      "content": "<p>Thanks for the info. This function makes double stratified KFolds. I provide code to make triple stratifed KFolds <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526#923147\">here</a>. (1) stratify by patient (2) stratify malignant (3) stratify patient count</p>",
      "rawMarkdown": "Thanks for the info. This function makes double stratified KFolds. I provide code to make triple stratifed KFolds [here][1]. (1) stratify by patient (2) stratify malignant (3) stratify patient count\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526#923147",
      "votes": null
    },
    {
      "id": "955744",
      "postDate": "08/02/2020 22:09:18",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> so for someone like me, who is using JPEGs instead of TFRecords, how could I take advantage of the work you have done? Would I simply tell my CV to use TFRecord field and use that as a Group?  So in other words use GroupKFold where group=TFRecord?</p>",
      "rawMarkdown": "cdeotte so for someone like me, who is using JPEGs instead of TFRecords, how could I take advantage of the work you have done? Would I simply tell my CV to use TFRecord field and use that as a Group?  So in other words use GroupKFold where group=TFRecord?",
      "votes": null
    },
    {
      "id": "956029",
      "postDate": "08/03/2020 07:05:36",
      "content": "<p>Yes GroupKFold with <code>group = tfrecord</code> will work. But first remove all rows with <code>tfrecord= -1</code>. Or do, what i do in the notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> which is basic KFold on the numbers 0 thru 14.</p>",
      "rawMarkdown": "Yes GroupKFold with `group = tfrecord` will work. But first remove all rows with `tfrecord= -1`. Or do, what i do in the notebook [here][1] which is basic KFold on the numbers 0 thru 14.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 954674,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/02/2020 01:51:37",
      "content": "<p>Thanks for the info. This function makes double stratified KFolds. I provide code to make triple stratifed KFolds <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526#923147\">here</a>. (1) stratify by patient (2) stratify malignant (3) stratify patient count</p>",
      "votes": null,
      "replies": [
        {
          "id": 955744,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/02/2020 22:09:18",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> so for someone like me, who is using JPEGs instead of TFRecords, how could I take advantage of the work you have done? Would I simply tell my CV to use TFRecord field and use that as a Group?  So in other words use GroupKFold where group=TFRecord?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956029,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/03/2020 07:05:36",
          "content": "<p>Yes GroupKFold with <code>group = tfrecord</code> will work. But first remove all rows with <code>tfrecord= -1</code>. Or do, what i do in the notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> which is basic KFold on the numbers 0 thru 14.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "953846": "This is the official code (work in progress, not yet in production as far as I know), of Scikit's `StratifiedGroupKFold`.  So some of you have been using `StratifiedKFold`, or `GroupKFold`, well now you have both!\n\nSave file as `split.py` and then add to your code:\n\n```\nfrom split import StratifiedGroupKFold, RepeatedStratifiedGroupKFold\n\n```\n\nShare how you are using it.  In my code, I use `patient_id` as a group, so that way I do not have leaks between folds of the same patient.\n\n`cv = StratifiedGroupKFold(n_splits=param['num_splits'], shuffle = True, random_state=867)`\n\nI exclude indexes past `33125` in `val` since those are not part of the original data.\n\n```\nfor fold, (train_idx, val_idx) in enumerate(cv.split(X=np.zeros(len(train_df)), \n                                                       y=train_df['target'], \n                                                       groups=train_df['patient_id'].tolist()), 1):\n    val_idx = val_idx[val_idx &lt;= 33125] # only original 2020 data in validation\n\n```\n\nHere is the class\n\n\n```\nfrom collections import Counter, defaultdict\n\nimport numpy as np\n\nfrom sklearn.model_selection._split import _BaseKFold, _RepeatedSplits\nfrom sklearn.utils.validation import check_random_state\n\nclass StratifiedGroupKFold(_BaseKFold):\n    \"\"\"Stratified K-Folds iterator variant with non-overlapping groups.\n\n\n    This cross-validation object is a variation of StratifiedKFold that returns\n    stratified folds with non-overlapping groups. The folds are made by\n    preserving the percentage of samples for each class.\n\n    The same group will not appear in two different folds (the number of\n    distinct groups has to be at least equal to the number of folds).\n\n    The difference between GroupKFold and StratifiedGroupKFold is that\n    the former attempts to create balanced folds such that the number of\n    distinct groups is approximately the same in each fold, whereas\n    StratifiedGroupKFold attempts to create folds which preserve the\n    percentage of samples for each class.\n\n    Read more in the :ref:`User Guide",
    "954674": "Thanks for the info. This function makes double stratified KFolds. I provide code to make triple stratifed KFolds [here][1]. (1) stratify by patient (2) stratify malignant (3) stratify patient count\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526#923147",
    "955744": "cdeotte so for someone like me, who is using JPEGs instead of TFRecords, how could I take advantage of the work you have done? Would I simply tell my CV to use TFRecord field and use that as a Group?  So in other words use GroupKFold where group=TFRecord?",
    "956029": "Yes GroupKFold with `group = tfrecord` will work. But first remove all rows with `tfrecord= -1`. Or do, what i do in the notebook [here][1] which is basic KFold on the numbers 0 thru 14.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates"
  },
  "source": "meta"
}