{
  "id": 90444,
  "title": "Stratified KFolds",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90444",
  "author_name": "",
  "post_date": "2019-04-23T21:17:39.401190100Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Inspired by this discussion <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-521819\">Interesting insight from shuffling ?</a>, I tried to find a way to construct a solid cross validation method.</p>\n\n<p>Looking at the internet, I found this scikit-learn github issue <a href=\"https://github.com/scikit-learn/scikit-learn/issues/4757\">https://github.com/scikit-learn/scikit-learn/issues/4757</a>, talking about stratified kfold for regression.</p>\n\n<p>In particular, someone wrote an extension of the scikit-learn StratifiedKFold, in order to have a continuous target distributed the same way in each fold <a href=\"https://github.com/biocore/calour/blob/master/calour/training.py#L144\">https://github.com/biocore/calour/blob/master/calour/training.py#L144</a></p>\n\n<p>I pasted it here, for those who doesn't want to click on the links ;-) </p>\n\n<p>```\nclass SortedStratifiedKFold(StratifiedKFold):\n    def <strong>init</strong>(self, n_splits=3, shuffle=False, random_state=None):\n        super().<strong>init</strong>(n_splits, shuffle, random_state)</p>\n\n<pre><code>def _sort_partition(self, y):\n    n = len(y)\n    cats = np.empty(n, dtype='u4')\n    div, mod = divmod(n, self.n_splits)\n    cats[:n-mod] = np.repeat(range(div), self.n_splits)\n    cats[n-mod:] = div + 1\n    # run argsort twice to get the rank of each y value\n    return cats[np.argsort(np.argsort(y))]\n\ndef split(self, X, y, groups=None):\n    y_cat = self._sort_partition(y)\n    return super().split(X, y_cat, groups)\n</code></pre>\n\n<p>```</p>\n\n<p>In this competition, it allows me to have a standard deviation between folds of 0.015 to 0.025 which is quite lower that what I had with standard KFold (shuffling or not) with a average CV of 1.98 to 2.1\n(I managed to get an even lower CV but with features that where not distributed the same way in train and test... I don't  know why yet but that could be another discussion to have ;-) )</p>\n\n<p>So for me this is the way to go but I didn't manage to make it works for LB (I probably don't have the good features to do so yet). </p>\n\n<p>What's your expert opinions about this?</p>",
  "messages": [
    {
      "id": "522093",
      "postDate": "04/23/2019 21:17:39",
      "content": "<p>Inspired by this discussion <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-521819\">Interesting insight from shuffling ?</a>, I tried to find a way to construct a solid cross validation method.</p>\n\n<p>Looking at the internet, I found this scikit-learn github issue <a href=\"https://github.com/scikit-learn/scikit-learn/issues/4757\">https://github.com/scikit-learn/scikit-learn/issues/4757</a>, talking about stratified kfold for regression.</p>\n\n<p>In particular, someone wrote an extension of the scikit-learn StratifiedKFold, in order to have a continuous target distributed the same way in each fold <a href=\"https://github.com/biocore/calour/blob/master/calour/training.py#L144\">https://github.com/biocore/calour/blob/master/calour/training.py#L144</a></p>\n\n<p>I pasted it here, for those who doesn't want to click on the links ;-) </p>\n\n<p>```\nclass SortedStratifiedKFold(StratifiedKFold):\n    def <strong>init</strong>(self, n_splits=3, shuffle=False, random_state=None):\n        super().<strong>init</strong>(n_splits, shuffle, random_state)</p>\n\n<pre><code>def _sort_partition(self, y):\n    n = len(y)\n    cats = np.empty(n, dtype='u4')\n    div, mod = divmod(n, self.n_splits)\n    cats[:n-mod] = np.repeat(range(div), self.n_splits)\n    cats[n-mod:] = div + 1\n    # run argsort twice to get the rank of each y value\n    return cats[np.argsort(np.argsort(y))]\n\ndef split(self, X, y, groups=None):\n    y_cat = self._sort_partition(y)\n    return super().split(X, y_cat, groups)\n</code></pre>\n\n<p>```</p>\n\n<p>In this competition, it allows me to have a standard deviation between folds of 0.015 to 0.025 which is quite lower that what I had with standard KFold (shuffling or not) with a average CV of 1.98 to 2.1\n(I managed to get an even lower CV but with features that where not distributed the same way in train and test... I don't  know why yet but that could be another discussion to have ;-) )</p>\n\n<p>So for me this is the way to go but I didn't manage to make it works for LB (I probably don't have the good features to do so yet). </p>\n\n<p>What's your expert opinions about this?</p>",
      "rawMarkdown": "Inspired by this discussion [Interesting insight from shuffling ?](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-521819), I tried to find a way to construct a solid cross validation method.\n\nLooking at the internet, I found this scikit-learn github issue https://github.com/scikit-learn/scikit-learn/issues/4757, talking about stratified kfold for regression.\n\nIn particular, someone wrote an extension of the scikit-learn StratifiedKFold, in order to have a continuous target distributed the same way in each fold https://github.com/biocore/calour/blob/master/calour/training.py#L144\n\nI pasted it here, for those who doesn't want to click on the links ;-) \n\n```\nclass SortedStratifiedKFold(StratifiedKFold):\n    def __init__(self, n_splits=3, shuffle=False, random_state=None):\n        super().__init__(n_splits, shuffle, random_state)\n\n    def _sort_partition(self, y):\n        n = len(y)\n        cats = np.empty(n, dtype='u4')\n        div, mod = divmod(n, self.n_splits)\n        cats[:n-mod] = np.repeat(range(div), self.n_splits)\n        cats[n-mod:] = div + 1\n        # run argsort twice to get the rank of each y value\n        return cats[np.argsort(np.argsort(y))]\n\n    def split(self, X, y, groups=None):\n        y_cat = self._sort_partition(y)\n        return super().split(X, y_cat, groups)\n```\n\nIn this competition, it allows me to have a standard deviation between folds of 0.015 to 0.025 which is quite lower that what I had with standard KFold (shuffling or not) with a average CV of 1.98 to 2.1\n(I managed to get an even lower CV but with features that where not distributed the same way in train and test... I don't  know why yet but that could be another discussion to have ;-) )\n\nSo for me this is the way to go but I didn't manage to make it works for LB (I probably don't have the good features to do so yet). \n\nWhat's your expert opinions about this?",
      "votes": null
    },
    {
      "id": "522104",
      "postDate": "04/23/2019 21:33:21",
      "content": "<p>I think finding the proper way to do CV this competition will be one of the main keys for sure so thank you for sharing this class. CPMP brought up the fact that in reality there is only 16 observations (16 different times to labquake) in the training set, so I'm not fully convinced that KFolds on segments divided every 150,000 observations is the answer. However, it is what I am sticking with for now, so I will definitely try this out.</p>",
      "rawMarkdown": "I think finding the proper way to do CV this competition will be one of the main keys for sure so thank you for sharing this class. CPMP brought up the fact that in reality there is only 16 observations (16 different times to labquake) in the training set, so I'm not fully convinced that KFolds on segments divided every 150,000 observations is the answer. However, it is what I am sticking with for now, so I will definitely try this out.",
      "votes": null
    },
    {
      "id": "522108",
      "postDate": "04/23/2019 21:36:19",
      "content": "<p>I just check this validation. CV rises from 1.907 to 1.9750, std amomg folds 0.053 -&gt; 0.147.  LB 1.393 -&gt; 1.383.  I do not really think that the model improves. It is more like equlizing prediction`s median among folds. This validation may be better for this particular public data set due to closer median of the predictions to public median. Or just bagging gives extra boost from bigger std of predictions (0.352 -&gt; 0.386). Woud it be better in private? Is this method better for training model or validation? I do not know.</p>",
      "rawMarkdown": "I just check this validation. CV rises from 1.907 to 1.9750, std amomg folds 0.053 -&gt; 0.147.  LB 1.393 -&gt; 1.383.  I do not really think that the model improves. It is more like equlizing prediction`s median among folds. This validation may be better for this particular public data set due to closer median of the predictions to public median. Or just bagging gives extra boost from bigger std of predictions (0.352 -&gt; 0.386). Woud it be better in private? Is this method better for training model or validation? I do not know.",
      "votes": null
    },
    {
      "id": "522114",
      "postDate": "04/23/2019 21:54:44",
      "content": "<p>Sure, I didn't try the folds based on labquake yet, but I definitely will</p>",
      "rawMarkdown": "Sure, I didn't try the folds based on labquake yet, but I definitely will",
      "votes": null
    },
    {
      "id": "522117",
      "postDate": "04/23/2019 22:01:45",
      "content": "<p>I do not know either :-) But I'm glad it made you climb the LB a bit ;-)</p>",
      "rawMarkdown": "I do not know either :-) But I'm glad it made you climb the LB a bit ;-)",
      "votes": null
    },
    {
      "id": "522121",
      "postDate": "04/23/2019 22:08:57",
      "content": "<p>Thanks, but lb means nothing in this competition, especially with such minor improvements IMO. I am waiting for huge shake-up.</p>",
      "rawMarkdown": "Thanks, but lb means nothing in this competition, especially with such minor improvements IMO. I am waiting for huge shake-up.",
      "votes": null
    },
    {
      "id": "522123",
      "postDate": "04/23/2019 22:11:00",
      "content": "<p>Yeah I know, it was just a joke. And this LB earthquake is quite easy to predict :-D</p>",
      "rawMarkdown": "Yeah I know, it was just a joke. And this LB earthquake is quite easy to predict :-D",
      "votes": null
    },
    {
      "id": "522320",
      "postDate": "04/24/2019 08:26:59",
      "content": "<p>Didn't work for me either, CV remained almost the same</p>",
      "rawMarkdown": "Didn't work for me either, CV remained almost the same",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 522104,
      "author_name": "halldalton94",
      "author_url": "",
      "post_date": "04/23/2019 21:33:21",
      "content": "<p>I think finding the proper way to do CV this competition will be one of the main keys for sure so thank you for sharing this class. CPMP brought up the fact that in reality there is only 16 observations (16 different times to labquake) in the training set, so I'm not fully convinced that KFolds on segments divided every 150,000 observations is the answer. However, it is what I am sticking with for now, so I will definitely try this out.</p>",
      "votes": null,
      "replies": [
        {
          "id": 522114,
          "author_name": "daijin12",
          "author_url": "",
          "post_date": "04/23/2019 21:54:44",
          "content": "<p>Sure, I didn't try the folds based on labquake yet, but I definitely will</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522108,
      "author_name": "simakov",
      "author_url": "",
      "post_date": "04/23/2019 21:36:19",
      "content": "<p>I just check this validation. CV rises from 1.907 to 1.9750, std amomg folds 0.053 -&gt; 0.147.  LB 1.393 -&gt; 1.383.  I do not really think that the model improves. It is more like equlizing prediction`s median among folds. This validation may be better for this particular public data set due to closer median of the predictions to public median. Or just bagging gives extra boost from bigger std of predictions (0.352 -&gt; 0.386). Woud it be better in private? Is this method better for training model or validation? I do not know.</p>",
      "votes": null,
      "replies": [
        {
          "id": 522117,
          "author_name": "daijin12",
          "author_url": "",
          "post_date": "04/23/2019 22:01:45",
          "content": "<p>I do not know either :-) But I'm glad it made you climb the LB a bit ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522121,
          "author_name": "simakov",
          "author_url": "",
          "post_date": "04/23/2019 22:08:57",
          "content": "<p>Thanks, but lb means nothing in this competition, especially with such minor improvements IMO. I am waiting for huge shake-up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522123,
          "author_name": "daijin12",
          "author_url": "",
          "post_date": "04/23/2019 22:11:00",
          "content": "<p>Yeah I know, it was just a joke. And this LB earthquake is quite easy to predict :-D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522320,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "04/24/2019 08:26:59",
      "content": "<p>Didn't work for me either, CV remained almost the same</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "522093": "Inspired by this discussion [Interesting insight from shuffling ?](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-521819), I tried to find a way to construct a solid cross validation method.\n\nLooking at the internet, I found this scikit-learn github issue https://github.com/scikit-learn/scikit-learn/issues/4757, talking about stratified kfold for regression.\n\nIn particular, someone wrote an extension of the scikit-learn StratifiedKFold, in order to have a continuous target distributed the same way in each fold https://github.com/biocore/calour/blob/master/calour/training.py#L144\n\nI pasted it here, for those who doesn't want to click on the links ;-) \n\n```\nclass SortedStratifiedKFold(StratifiedKFold):\n    def __init__(self, n_splits=3, shuffle=False, random_state=None):\n        super().__init__(n_splits, shuffle, random_state)\n\n    def _sort_partition(self, y):\n        n = len(y)\n        cats = np.empty(n, dtype='u4')\n        div, mod = divmod(n, self.n_splits)\n        cats[:n-mod] = np.repeat(range(div), self.n_splits)\n        cats[n-mod:] = div + 1\n        # run argsort twice to get the rank of each y value\n        return cats[np.argsort(np.argsort(y))]\n\n    def split(self, X, y, groups=None):\n        y_cat = self._sort_partition(y)\n        return super().split(X, y_cat, groups)\n```\n\nIn this competition, it allows me to have a standard deviation between folds of 0.015 to 0.025 which is quite lower that what I had with standard KFold (shuffling or not) with a average CV of 1.98 to 2.1\n(I managed to get an even lower CV but with features that where not distributed the same way in train and test... I don't  know why yet but that could be another discussion to have ;-) )\n\nSo for me this is the way to go but I didn't manage to make it works for LB (I probably don't have the good features to do so yet). \n\nWhat's your expert opinions about this?",
    "522104": "I think finding the proper way to do CV this competition will be one of the main keys for sure so thank you for sharing this class. CPMP brought up the fact that in reality there is only 16 observations (16 different times to labquake) in the training set, so I'm not fully convinced that KFolds on segments divided every 150,000 observations is the answer. However, it is what I am sticking with for now, so I will definitely try this out.",
    "522108": "I just check this validation. CV rises from 1.907 to 1.9750, std amomg folds 0.053 -&gt; 0.147.  LB 1.393 -&gt; 1.383.  I do not really think that the model improves. It is more like equlizing prediction`s median among folds. This validation may be better for this particular public data set due to closer median of the predictions to public median. Or just bagging gives extra boost from bigger std of predictions (0.352 -&gt; 0.386). Woud it be better in private? Is this method better for training model or validation? I do not know.",
    "522114": "Sure, I didn't try the folds based on labquake yet, but I definitely will",
    "522117": "I do not know either :-) But I'm glad it made you climb the LB a bit ;-)",
    "522121": "Thanks, but lb means nothing in this competition, especially with such minor improvements IMO. I am waiting for huge shake-up.",
    "522123": "Yeah I know, it was just a joke. And this LB earthquake is quite easy to predict :-D",
    "522320": "Didn't work for me either, CV remained almost the same"
  },
  "source": "meta"
}