{
  "id": 56346,
  "title": "Validation strategy",
  "url": "/competitions/avito-demand-prediction/discussion/56346",
  "author_name": "Konrad Banachewicz",
  "post_date": "2018-05-08T20:44:21.853000",
  "votes": 29,
  "comment_count": 37,
  "views": 0,
  "content": "<p>I'm curious if anybody is willing to discuss their validation approach - I tried random split (loose relationship with LB), time based split based on similar time length (i.e. creating a validation set spanning a period of similar length as test set - too few observations left), split based on observation count (okayish, but nothing to write home about - local validation score 0.2231, lb 0.2291).</p>",
  "messages": [
    {
      "id": 325763,
      "postDate": "2018-05-08T20:44:21.853Z",
      "content": "<p>I'm curious if anybody is willing to discuss their validation approach - I tried random split (loose relationship with LB), time based split based on similar time length (i.e. creating a validation set spanning a period of similar length as test set - too few observations left), split based on observation count (okayish, but nothing to write home about - local validation score 0.2231, lb 0.2291).</p>",
      "rawMarkdown": "I'm curious if anybody is willing to discuss their validation approach - I tried random split (loose relationship with LB), time based split based on similar time length (i.e. creating a validation set spanning a period of similar length as test set - too few observations left), split based on observation count (okayish, but nothing to write home about - local validation score 0.2231, lb 0.2291).",
      "votes": 29
    },
    {
      "id": 325775,
      "postDate": "2018-05-08T21:03:55.713Z",
      "content": "<p>We have 10 random KFold and constant difference between CV and LB near 0.005.</p>",
      "rawMarkdown": "We have 10 random KFold and constant difference between CV and LB near 0.005.",
      "votes": 15,
      "replies": [
        {
          "id": 325777,
          "postDate": "2018-05-08T21:05:39.373Z",
          "content": "<p>Interesting - I will give it a try, then. Thanks!</p>",
          "rawMarkdown": "Interesting - I will give it a try, then. Thanks!",
          "votes": 1
        },
        {
          "id": 325879,
          "postDate": "2018-05-09T02:13:51.677Z",
          "content": "<p>Seems perfect</p>",
          "rawMarkdown": "Seems perfect"
        },
        {
          "id": 326610,
          "postDate": "2018-05-10T02:39:03.840Z",
          "content": "<p>same here, 5 random KFold and difference between CV and LB about 0.004-0.005, but when I performed stacking or using optim function, the difference is up to 0.006-0.008, weird.</p>",
          "rawMarkdown": "same here, 5 random KFold and difference between CV and LB about 0.004-0.005, but when I performed stacking or using optim function, the difference is up to 0.006-0.008, weird."
        },
        {
          "id": 326615,
          "postDate": "2018-05-10T03:15:57.927Z",
          "content": "<p>@Andrew-Chang: That's pretty normal in my experience as usually stacking overfits a little bit more.</p>",
          "rawMarkdown": "@Andrew-Chang: That's pretty normal in my experience as usually stacking overfits a little bit more."
        },
        {
          "id": 326619,
          "postDate": "2018-05-10T03:20:56.703Z",
          "content": "<p>You're right, but when I use k-folds and optim() function to determine the weight of every model, it also leads to big difference. It really confused me.</p>",
          "rawMarkdown": "You're right, but when I use k-folds and optim() function to determine the weight of every model, it also leads to big difference. It really confused me."
        },
        {
          "id": 326820,
          "postDate": "2018-05-10T11:40:40.640Z",
          "content": "<p>Update: so a 5-fold cv seems to give a nice and stable 0.005 difference between public and private LB. Thanks for the tip @Sergei.</p>",
          "rawMarkdown": "Update: so a 5-fold cv seems to give a nice and stable 0.005 difference between public and private LB. Thanks for the tip @Sergei.",
          "votes": 9
        },
        {
          "id": 348635,
          "postDate": "2018-06-27T03:10:23.683Z",
          "content": "<p>@Peter Hurfold, does it mean for the stacking model using 10 fold cross validation might lead to more precise estimation than 5 fold? <br>\nAnd is it comparable for model using 5 fold CV and 10 fold?  Just curious on how to find a best way to make the best use of CV. Thanks!</p>",
          "rawMarkdown": "@Peter Hurfold, does it mean for the stacking model using 10 fold cross validation might lead to more precise estimation than 5 fold?  \nAnd is it comparable for model using 5 fold CV and 10 fold?  Just curious on how to find a best way to make the best use of CV. Thanks!"
        },
        {
          "id": 348637,
          "postDate": "2018-06-27T03:11:48.193Z",
          "content": "<p>10 fold would be more precise than 5 fold, but maybe not by much. And it would be twice as time consuming. I personally stick with 5 fold, but I know some purists who do otherwise. As long as you're consistent throughout, you should be fine. :)</p>",
          "rawMarkdown": "10 fold would be more precise than 5 fold, but maybe not by much. And it would be twice as time consuming. I personally stick with 5 fold, but I know some purists who do otherwise. As long as you're consistent throughout, you should be fine. :)",
          "votes": 1
        },
        {
          "id": 348642,
          "postDate": "2018-06-27T03:33:45.167Z",
          "content": "<p>Maybe another question from some purists might ask: usually how much the difference between CV and LB can be tolerated for the private LB? </p>",
          "rawMarkdown": "Maybe another question from some purists might ask: usually how much the difference between CV and LB can be tolerated for the private LB? "
        },
        {
          "id": 348651,
          "postDate": "2018-06-27T03:53:27.997Z",
          "content": "<p>It depends on the competition, which is why it is great to ask other people what they are experiencing and see if yours is similar.</p>",
          "rawMarkdown": "It depends on the competition, which is why it is great to ask other people what they are experiencing and see if yours is similar.",
          "votes": 1
        },
        {
          "id": 348652,
          "postDate": "2018-06-27T03:53:58.873Z",
          "content": "<p>Another important aspect is how much variability there is in your gap between your submissions. The more consistency, the safer you can feel.</p>",
          "rawMarkdown": "Another important aspect is how much variability there is in your gap between your submissions. The more consistency, the safer you can feel.",
          "votes": 1
        },
        {
          "id": 348653,
          "postDate": "2018-06-27T03:54:23.890Z",
          "content": "<p>Generally, you should trust your CV score most because it's based on the most data. Having a time-ordered LB can undermine that trust some, though.</p>",
          "rawMarkdown": "Generally, you should trust your CV score most because it's based on the most data. Having a time-ordered LB can undermine that trust some, though."
        }
      ]
    },
    {
      "id": 325778,
      "postDate": "2018-05-08T21:06:14.847Z",
      "content": "<p>I really like seeing this thread so I can see if I'm overfitting. :) Right now, I'm doing 5-fold CV and have a consistent gap between CV and LB between 0.0031 and 0.0038. I tried a time-based validation strategy once but it made my score worse.</p>",
      "rawMarkdown": "I really like seeing this thread so I can see if I'm overfitting. :) Right now, I'm doing 5-fold CV and have a consistent gap between CV and LB between 0.0031 and 0.0038. I tried a time-based validation strategy once but it made my score worse.",
      "votes": 9
    },
    {
      "id": 326945,
      "postDate": "2018-05-10T14:54:17.447Z",
      "content": "<p>Be careful with the last days of training data set:</p>\n\n<p>training[(training['activation_date']&gt;=pd.to_datetime('2017-03-29'))].shape</p>\n\n<p>(100, 17)</p>\n\n<p>training[(training['activation_date']&gt;=pd.to_datetime('2017-03-27'))].shape</p>\n\n<p>(227848, 17)</p>\n\n<p>Only 100 samples after 2017-03-29.</p>",
      "rawMarkdown": "Be careful with the last days of training data set:\n\ntraining[(training['activation_date']&gt;=pd.to_datetime('2017-03-29'))].shape\n\n(100, 17)\n\ntraining[(training['activation_date']&gt;=pd.to_datetime('2017-03-27'))].shape\n\n(227848, 17)\n\nOnly 100 samples after 2017-03-29.",
      "votes": 3,
      "replies": [
        {
          "id": 327002,
          "postDate": "2018-05-10T16:37:30.413Z",
          "content": "<p>similar pattern can be seen in last days of test set.</p>\n\n<pre><code>2017-04-12    81824\n2017-04-13    77176\n2017-04-14    70366\n2017-04-15    58793\n2017-04-16    58909\n2017-04-17    80191\n2017-04-18    81114\n2017-04-19       64\n2017-04-20        1\n</code></pre>",
          "rawMarkdown": "similar pattern can be seen in last days of test set.\n\n    2017-04-12    81824\n    2017-04-13    77176\n    2017-04-14    70366\n    2017-04-15    58793\n    2017-04-16    58909\n    2017-04-17    80191\n    2017-04-18    81114\n    2017-04-19       64\n    2017-04-20        1\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 330580,
      "postDate": "2018-05-19T06:40:00.020Z",
      "content": "<p>I am using this idea of stratified KFold on binned target:</p>\n\n<pre><code>import numpy as np\nfrom sklearn.model_selection import train_test_split\n\ndef stratified_train_valid_split(meta_train,  target_column,  target_bins, valid_size, random_state=1234):\n\n    y = meta_train[target_column].values\n    bins = np.linspace(0, y.shape[0], target_bins)\n    y_binned = np.digitize(y, bins)\n\n   return train_test_split(meta_train, test_size=valid_size, stratify=y_binned, random_state=random_state)\n</code></pre>\n\n<p>It seems that the target distribution in train/valid is similiar. So far only tested it on bad models but it seems to correlate with the LB improvements. </p>\n\n<p>What do you think of this approach?</p>",
      "rawMarkdown": "I am using this idea of stratified KFold on binned target:\n\n\n    import numpy as np\n    from sklearn.model_selection import train_test_split\n \n    def stratified_train_valid_split(meta_train,  target_column,  target_bins, valid_size, random_state=1234):\n\n        y = meta_train[target_column].values\n        bins = np.linspace(0, y.shape[0], target_bins)\n        y_binned = np.digitize(y, bins)\n\n       return train_test_split(meta_train, test_size=valid_size, stratify=y_binned, random_state=random_state)\n\nIt seems that the target distribution in train/valid is similiar. So far only tested it on bad models but it seems to correlate with the LB improvements. \n\nWhat do you think of this approach?",
      "votes": 4,
      "replies": [
        {
          "id": 330811,
          "postDate": "2018-05-19T19:06:12.543Z",
          "content": "<p>I'm doing the same thing, but I'm getting very different LB results. How similar are yours?</p>",
          "rawMarkdown": "I'm doing the same thing, but I'm getting very different LB results. How similar are yours?"
        },
        {
          "id": 332461,
          "postDate": "2018-05-23T07:23:37.900Z",
          "content": "<p>Finally ran new models, and it seems that 0.005 delta is what this approach gets me. </p>",
          "rawMarkdown": "Finally ran new models, and it seems that 0.005 delta is what this approach gets me. "
        }
      ]
    },
    {
      "id": 336491,
      "postDate": "2018-05-31T20:37:14.747Z",
      "content": "<p>I'm using March 26 - 28 as my validation set. Like most people here, I'm also getting ~0.005 delta. I'm still starting out though so not sure how well this will hold.</p>\n\n<p><img src=\"https://i.imgur.com/rm9b2AL.png\" alt=\"validation thingy\"></p>",
      "rawMarkdown": "I'm using March 26 - 28 as my validation set. Like most people here, I'm also getting ~0.005 delta. I'm still starting out though so not sure how well this will hold.\n\n![validation thingy][1]\n\n\n  [1]: https://i.imgur.com/rm9b2AL.png",
      "votes": 2,
      "replies": [
        {
          "id": 336554,
          "postDate": "2018-05-31T23:43:56.093Z",
          "content": "<p>That's a cool sheet. Thanks for sharing.</p>",
          "rawMarkdown": "That's a cool sheet. Thanks for sharing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 336350,
      "postDate": "2018-05-31T13:43:33.810Z",
      "content": "<p>I am on about 0.004 CV/LB delta, but I am only starting out - will be interesting to see how this gap will develop as I add more magic/non-magic into the mix :) </p>",
      "rawMarkdown": "I am on about 0.004 CV/LB delta, but I am only starting out - will be interesting to see how this gap will develop as I add more magic/non-magic into the mix :) ",
      "replies": [
        {
          "id": 336361,
          "postDate": "2018-05-31T14:09:34.777Z",
          "content": "<p>Now that I've added features from train_active / test_active and the images, I've found my gap has widened to ~0.0043.</p>",
          "rawMarkdown": "Now that I've added features from train_active / test_active and the images, I've found my gap has widened to ~0.0043."
        }
      ]
    },
    {
      "id": 331960,
      "postDate": "2018-05-22T08:50:05.693Z",
      "content": "<p>I use valid set as 10% random split from train  set. It gives me  ~0.0034 CV-LB delta</p>",
      "rawMarkdown": "I use valid set as 10% random split from train  set. It gives me  ~0.0034 CV-LB delta"
    },
    {
      "id": 331033,
      "postDate": "2018-05-20T08:11:21.160Z",
      "content": "<p>I confirm that using 5-fold random split results in a ~0.005 separation wrt LB. I tried decreasing and increasing the number of folds but 5 seems to be the optimal choice (also considering time). <br>\nI am also using the last 20% of the train dataset as validation, which results in a slightly smaller shift, around ~0.004</p>",
      "rawMarkdown": "I confirm that using 5-fold random split results in a ~0.005 separation wrt LB. I tried decreasing and increasing the number of folds but 5 seems to be the optimal choice (also considering time).  \nI am also using the last 20% of the train dataset as validation, which results in a slightly smaller shift, around ~0.004",
      "replies": [
        {
          "id": 331040,
          "postDate": "2018-05-20T08:30:41.333Z",
          "content": "<blockquote>\n  <p><strong>jtrane87 wrote</strong></p>\n  \n  <blockquote>\n    <p>I confirm that using 5-fold random split results in a ~0.005</p>\n  </blockquote>\n</blockquote>\n\n<p>It may depend on the model </p>\n\n<p>5-fold random split results in a  0.002x-0.003 CV-LB delta  for me </p>",
          "rawMarkdown": "\n&gt; **jtrane87 wrote**\n&gt; \n&gt; &gt; I confirm that using 5-fold random split results in a ~0.005\n\nIt may depend on the model \n\n5-fold random split results in a  0.002x-0.003 CV-LB delta  for me \n",
          "votes": 1
        },
        {
          "id": 331044,
          "postDate": "2018-05-20T08:34:54.263Z",
          "content": "<p>Sure, it may depend on the model. Let's say: these are my observations so far. </p>",
          "rawMarkdown": "Sure, it may depend on the model. Let's say: these are my observations so far. "
        }
      ]
    },
    {
      "id": 328702,
      "postDate": "2018-05-14T23:40:17.410Z",
      "content": "<p>I don't think random spliting (or CV) would be a good approach, because you would be using future data to predict the past, which most of the time gives a more optimistic accuracy.</p>\n\n<p><a href=\"http://www.fast.ai/2017/11/13/validation-sets/\">http://www.fast.ai/2017/11/13/validation-sets/</a></p>",
      "rawMarkdown": "I don't think random spliting (or CV) would be a good approach, because you would be using future data to predict the past, which most of the time gives a more optimistic accuracy.\n\nhttp://www.fast.ai/2017/11/13/validation-sets/",
      "replies": [
        {
          "id": 328712,
          "postDate": "2018-05-15T00:08:55.970Z",
          "content": "<p>In principle I agree, but: 1. this particular series seems to be pretty stationary 2. k-fold cv gives a more consistent estimate (smaller gap with public lb) then the time-based split.</p>",
          "rawMarkdown": "In principle I agree, but: 1. this particular series seems to be pretty stationary 2. k-fold cv gives a more consistent estimate (smaller gap with public lb) then the time-based split.",
          "votes": 2
        },
        {
          "id": 328713,
          "postDate": "2018-05-15T00:09:01.223Z",
          "content": "<p>That's not always the case... Think about it this way. Every Kaggle problem is technically a time series problem, because the data was collected over time ...we just don't usually get to see the timestamps. What matters is not whether there is a timestamp or time-order to the data, but whether the target actually changes over time in a meaningful way. In this competition, it does not look like there is a reason so far to think that the time order matters. (But I'm still pretty scared of this and will keep hunting!)</p>",
          "rawMarkdown": "That's not always the case... Think about it this way. Every Kaggle problem is technically a time series problem, because the data was collected over time ...we just don't usually get to see the timestamps. What matters is not whether there is a timestamp or time-order to the data, but whether the target actually changes over time in a meaningful way. In this competition, it does not look like there is a reason so far to think that the time order matters. (But I'm still pretty scared of this and will keep hunting!)",
          "votes": 11
        },
        {
          "id": 329099,
          "postDate": "2018-05-15T17:57:23.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 348647,
          "postDate": "2018-06-27T03:48:53.650Z",
          "content": "<p>@Konrad Banachewicz, treating time-series doesn't mean you cannot apply n-fold strategy. In fact, a more accurate estimation of time-series (or any sequence data where samples are dependent either on time or other order)  could be \"leave-one-out \" or \"leave-n-out\"where you could leave n future samples for testing intentionally or incrementally if historical data give power of prediction. Because it is really time consuming, most people won't do such LOO or LnO CV with such big data.</p>\n\n<p>Maybe some statics guru can shed some light about conducting time series CV effectively. </p>\n\n<p>Please see this post for more explanation: <a href=\"https://robjhyndman.com/hyndsight/tscv/\">https://robjhyndman.com/hyndsight/tscv/</a></p>",
          "rawMarkdown": "@Konrad Banachewicz, treating time-series doesn't mean you cannot apply n-fold strategy. In fact, a more accurate estimation of time-series (or any sequence data where samples are dependent either on time or other order)  could be \"leave-one-out \" or \"leave-n-out\"where you could leave n future samples for testing intentionally or incrementally if historical data give power of prediction. Because it is really time consuming, most people won't do such LOO or LnO CV with such big data.\n\nMaybe some statics guru can shed some light about conducting time series CV effectively. \n\nPlease see this post for more explanation: https://robjhyndman.com/hyndsight/tscv/"
        }
      ]
    },
    {
      "id": 327057,
      "postDate": "2018-05-10T18:53:14.510Z",
      "content": "<p>I tried 8:2 random sample many times with different seed, but Validation loss didn't change more than 0.0001.\nIt seems that the train data well distributed, not same as testset.\nWhat's more, as the loss going to be lower, the gap seems to be lower. but still shows difference about 0.005.(bcuz my LB is still high)\n<br>\n| model || Val_loss | LB      | <br>\n| lgb      | | 0.2351 | 0.2412 | <br>\n| NN      | | 0.2588 | 0.3032 |  <br>\n| NN      | | 0.2475 | 0.2537 | <br>\n| NN      | |  0.2430 |  0.2488 | <br> \n| NN      | | 0.2407 |  0.2459 |  <br></p>\n\n<p><br>\nref : <a href=\"https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code\">https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code</a></p>",
      "rawMarkdown": "I tried 8:2 random sample many times with different seed, but Validation loss didn't change more than 0.0001.\nIt seems that the train data well distributed, not same as testset.\nWhat's more, as the loss going to be lower, the gap seems to be lower. but still shows difference about 0.005.(bcuz my LB is still high)\n<br>\n| model || Val_loss | LB      | <br>\n| lgb      | | 0.2351 | 0.2412 | <br>\n| NN      | | 0.2588 | 0.3032 |  <br>\n| NN      | | 0.2475 | 0.2537 | <br>\n| NN      | |  0.2430 |  0.2488 | <br> \n| NN      | | 0.2407 |  0.2459 |  <br>\n\n<br>\nref : https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code"
    },
    {
      "id": 326188,
      "postDate": "2018-05-09T12:31:40.240Z",
      "content": "<p>Anyone treating it as a time series problem?</p>",
      "rawMarkdown": "Anyone treating it as a time series problem?",
      "replies": [
        {
          "id": 326192,
          "postDate": "2018-05-09T12:35:18.880Z",
          "content": "<p>Branden Murray mentioned he used activation_date to create train-validation split. Could be worth trying, athough I believe cross-validation would be a better approach this competition.</p>\n\n<p>See: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/55587\">https://www.kaggle.com/c/avito-demand-prediction/discussion/55587</a></p>",
          "rawMarkdown": "Branden Murray mentioned he used activation_date to create train-validation split. Could be worth trying, athough I believe cross-validation would be a better approach this competition.\n\nSee: https://www.kaggle.com/c/avito-demand-prediction/discussion/55587",
          "votes": 2
        },
        {
          "id": 326193,
          "postDate": "2018-05-09T12:35:54.933Z",
          "content": "<p>I tried, but the lousy result is what prompted me to ask the original question...</p>",
          "rawMarkdown": "I tried, but the lousy result is what prompted me to ask the original question...",
          "votes": 3
        },
        {
          "id": 326398,
          "postDate": "2018-05-09T16:30:00.007Z",
          "content": "<p>I'm more worried about this now., as we now know that <a href=\"https://www.kaggle.com/cczaixian/test-on-lb-split\">the leaderboard is split by time</a> with the public leaderboard coming from April 12th and 13th, and the private leaderboard coming from April 14-18th. But we also know that <code>deal_probability</code> seems to be quite stable across time... if this is a time series problem, it's a very stationary one. So it's unclear to me if we need to do time series or not.</p>",
          "rawMarkdown": "I'm more worried about this now., as we now know that [the leaderboard is split by time](https://www.kaggle.com/cczaixian/test-on-lb-split) with the public leaderboard coming from April 12th and 13th, and the private leaderboard coming from April 14-18th. But we also know that `deal_probability` seems to be quite stable across time... if this is a time series problem, it's a very stationary one. So it's unclear to me if we need to do time series or not.",
          "votes": 16
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 325775,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2018-05-08T21:03:55.713000",
      "content": "<p>We have 10 random KFold and constant difference between CV and LB near 0.005.</p>",
      "votes": 15,
      "replies": [
        {
          "id": 325777,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-08T21:05:39.373000",
          "content": "<p>Interesting - I will give it a try, then. Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325879,
          "author_name": "Victor An",
          "author_url": "",
          "post_date": "2018-05-09T02:13:51.677000",
          "content": "<p>Seems perfect</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326610,
          "author_name": "Angus Chang",
          "author_url": "",
          "post_date": "2018-05-10T02:39:03.840000",
          "content": "<p>same here, 5 random KFold and difference between CV and LB about 0.004-0.005, but when I performed stacking or using optim function, the difference is up to 0.006-0.008, weird.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326615,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-10T03:15:57.927000",
          "content": "<p>@Andrew-Chang: That's pretty normal in my experience as usually stacking overfits a little bit more.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326619,
          "author_name": "Angus Chang",
          "author_url": "",
          "post_date": "2018-05-10T03:20:56.703000",
          "content": "<p>You're right, but when I use k-folds and optim() function to determine the weight of every model, it also leads to big difference. It really confused me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326820,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-10T11:40:40.640000",
          "content": "<p>Update: so a 5-fold cv seems to give a nice and stable 0.005 difference between public and private LB. Thanks for the tip @Sergei.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 348635,
          "author_name": "Rene Wang",
          "author_url": "",
          "post_date": "2018-06-27T03:10:23.683000",
          "content": "<p>@Peter Hurfold, does it mean for the stacking model using 10 fold cross validation might lead to more precise estimation than 5 fold? <br>\nAnd is it comparable for model using 5 fold CV and 10 fold?  Just curious on how to find a best way to make the best use of CV. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348637,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-27T03:11:48.193000",
          "content": "<p>10 fold would be more precise than 5 fold, but maybe not by much. And it would be twice as time consuming. I personally stick with 5 fold, but I know some purists who do otherwise. As long as you're consistent throughout, you should be fine. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348642,
          "author_name": "Rene Wang",
          "author_url": "",
          "post_date": "2018-06-27T03:33:45.167000",
          "content": "<p>Maybe another question from some purists might ask: usually how much the difference between CV and LB can be tolerated for the private LB? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348651,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-27T03:53:27.997000",
          "content": "<p>It depends on the competition, which is why it is great to ask other people what they are experiencing and see if yours is similar.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348652,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-27T03:53:58.873000",
          "content": "<p>Another important aspect is how much variability there is in your gap between your submissions. The more consistency, the safer you can feel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348653,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-06-27T03:54:23.890000",
          "content": "<p>Generally, you should trust your CV score most because it's based on the most data. Having a time-ordered LB can undermine that trust some, though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325778,
      "author_name": "Peter Hurford",
      "author_url": "",
      "post_date": "2018-05-08T21:06:14.847000",
      "content": "<p>I really like seeing this thread so I can see if I'm overfitting. :) Right now, I'm doing 5-fold CV and have a consistent gap between CV and LB between 0.0031 and 0.0038. I tried a time-based validation strategy once but it made my score worse.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 326945,
      "author_name": "Dimitris Leventis",
      "author_url": "",
      "post_date": "2018-05-10T14:54:17.447000",
      "content": "<p>Be careful with the last days of training data set:</p>\n\n<p>training[(training['activation_date']&gt;=pd.to_datetime('2017-03-29'))].shape</p>\n\n<p>(100, 17)</p>\n\n<p>training[(training['activation_date']&gt;=pd.to_datetime('2017-03-27'))].shape</p>\n\n<p>(227848, 17)</p>\n\n<p>Only 100 samples after 2017-03-29.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 327002,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-05-10T16:37:30.413000",
          "content": "<p>similar pattern can be seen in last days of test set.</p>\n\n<pre><code>2017-04-12    81824\n2017-04-13    77176\n2017-04-14    70366\n2017-04-15    58793\n2017-04-16    58909\n2017-04-17    80191\n2017-04-18    81114\n2017-04-19       64\n2017-04-20        1\n</code></pre>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 330580,
      "author_name": "Jakub Czakon",
      "author_url": "",
      "post_date": "2018-05-19T06:40:00.020000",
      "content": "<p>I am using this idea of stratified KFold on binned target:</p>\n\n<pre><code>import numpy as np\nfrom sklearn.model_selection import train_test_split\n\ndef stratified_train_valid_split(meta_train,  target_column,  target_bins, valid_size, random_state=1234):\n\n    y = meta_train[target_column].values\n    bins = np.linspace(0, y.shape[0], target_bins)\n    y_binned = np.digitize(y, bins)\n\n   return train_test_split(meta_train, test_size=valid_size, stratify=y_binned, random_state=random_state)\n</code></pre>\n\n<p>It seems that the target distribution in train/valid is similiar. So far only tested it on bad models but it seems to correlate with the LB improvements. </p>\n\n<p>What do you think of this approach?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 330811,
          "author_name": "ianchute",
          "author_url": "",
          "post_date": "2018-05-19T19:06:12.543000",
          "content": "<p>I'm doing the same thing, but I'm getting very different LB results. How similar are yours?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 332461,
          "author_name": "Jakub Czakon",
          "author_url": "",
          "post_date": "2018-05-23T07:23:37.900000",
          "content": "<p>Finally ran new models, and it seems that 0.005 delta is what this approach gets me. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 336491,
      "author_name": "Lem Ko",
      "author_url": "",
      "post_date": "2018-05-31T20:37:14.747000",
      "content": "<p>I'm using March 26 - 28 as my validation set. Like most people here, I'm also getting ~0.005 delta. I'm still starting out though so not sure how well this will hold.</p>\n\n<p><img src=\"https://i.imgur.com/rm9b2AL.png\" alt=\"validation thingy\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 336554,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-31T23:43:56.093000",
          "content": "<p>That's a cool sheet. Thanks for sharing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 336350,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2018-05-31T13:43:33.810000",
      "content": "<p>I am on about 0.004 CV/LB delta, but I am only starting out - will be interesting to see how this gap will develop as I add more magic/non-magic into the mix :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 336361,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-31T14:09:34.777000",
          "content": "<p>Now that I've added features from train_active / test_active and the images, I've found my gap has widened to ~0.0043.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 331960,
      "author_name": "AlexTru",
      "author_url": "",
      "post_date": "2018-05-22T08:50:05.693000",
      "content": "<p>I use valid set as 10% random split from train  set. It gives me  ~0.0034 CV-LB delta</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 331033,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2018-05-20T08:11:21.160000",
      "content": "<p>I confirm that using 5-fold random split results in a ~0.005 separation wrt LB. I tried decreasing and increasing the number of folds but 5 seems to be the optimal choice (also considering time). <br>\nI am also using the last 20% of the train dataset as validation, which results in a slightly smaller shift, around ~0.004</p>",
      "votes": 0,
      "replies": [
        {
          "id": 331040,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-05-20T08:30:41.333000",
          "content": "<blockquote>\n  <p><strong>jtrane87 wrote</strong></p>\n  \n  <blockquote>\n    <p>I confirm that using 5-fold random split results in a ~0.005</p>\n  </blockquote>\n</blockquote>\n\n<p>It may depend on the model </p>\n\n<p>5-fold random split results in a  0.002x-0.003 CV-LB delta  for me </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 331044,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-05-20T08:34:54.263000",
          "content": "<p>Sure, it may depend on the model. Let's say: these are my observations so far. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 328702,
      "author_name": "Axel Straminsky",
      "author_url": "",
      "post_date": "2018-05-14T23:40:17.410000",
      "content": "<p>I don't think random spliting (or CV) would be a good approach, because you would be using future data to predict the past, which most of the time gives a more optimistic accuracy.</p>\n\n<p><a href=\"http://www.fast.ai/2017/11/13/validation-sets/\">http://www.fast.ai/2017/11/13/validation-sets/</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 328712,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-15T00:08:55.970000",
          "content": "<p>In principle I agree, but: 1. this particular series seems to be pretty stationary 2. k-fold cv gives a more consistent estimate (smaller gap with public lb) then the time-based split.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 328713,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-15T00:09:01.223000",
          "content": "<p>That's not always the case... Think about it this way. Every Kaggle problem is technically a time series problem, because the data was collected over time ...we just don't usually get to see the timestamps. What matters is not whether there is a timestamp or time-order to the data, but whether the target actually changes over time in a meaningful way. In this competition, it does not look like there is a reason so far to think that the time order matters. (But I'm still pretty scared of this and will keep hunting!)</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 329099,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-15T17:57:23.560000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348647,
          "author_name": "Rene Wang",
          "author_url": "",
          "post_date": "2018-06-27T03:48:53.650000",
          "content": "<p>@Konrad Banachewicz, treating time-series doesn't mean you cannot apply n-fold strategy. In fact, a more accurate estimation of time-series (or any sequence data where samples are dependent either on time or other order)  could be \"leave-one-out \" or \"leave-n-out\"where you could leave n future samples for testing intentionally or incrementally if historical data give power of prediction. Because it is really time consuming, most people won't do such LOO or LnO CV with such big data.</p>\n\n<p>Maybe some statics guru can shed some light about conducting time series CV effectively. </p>\n\n<p>Please see this post for more explanation: <a href=\"https://robjhyndman.com/hyndsight/tscv/\">https://robjhyndman.com/hyndsight/tscv/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327057,
      "author_name": "YeongTaek Oh",
      "author_url": "",
      "post_date": "2018-05-10T18:53:14.510000",
      "content": "<p>I tried 8:2 random sample many times with different seed, but Validation loss didn't change more than 0.0001.\nIt seems that the train data well distributed, not same as testset.\nWhat's more, as the loss going to be lower, the gap seems to be lower. but still shows difference about 0.005.(bcuz my LB is still high)\n<br>\n| model || Val_loss | LB      | <br>\n| lgb      | | 0.2351 | 0.2412 | <br>\n| NN      | | 0.2588 | 0.3032 |  <br>\n| NN      | | 0.2475 | 0.2537 | <br>\n| NN      | |  0.2430 |  0.2488 | <br> \n| NN      | | 0.2407 |  0.2459 |  <br></p>\n\n<p><br>\nref : <a href=\"https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code\">https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326188,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-05-09T12:31:40.240000",
      "content": "<p>Anyone treating it as a time series problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326192,
          "author_name": "Kevin",
          "author_url": "",
          "post_date": "2018-05-09T12:35:18.880000",
          "content": "<p>Branden Murray mentioned he used activation_date to create train-validation split. Could be worth trying, athough I believe cross-validation would be a better approach this competition.</p>\n\n<p>See: <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/55587\">https://www.kaggle.com/c/avito-demand-prediction/discussion/55587</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326193,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-05-09T12:35:54.933000",
          "content": "<p>I tried, but the lousy result is what prompted me to ask the original question...</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 326398,
          "author_name": "Peter Hurford",
          "author_url": "",
          "post_date": "2018-05-09T16:30:00.007000",
          "content": "<p>I'm more worried about this now., as we now know that <a href=\"https://www.kaggle.com/cczaixian/test-on-lb-split\">the leaderboard is split by time</a> with the public leaderboard coming from April 12th and 13th, and the private leaderboard coming from April 14-18th. But we also know that <code>deal_probability</code> seems to be quite stable across time... if this is a time series problem, it's a very stationary one. So it's unclear to me if we need to do time series or not.</p>",
          "votes": 16,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325763": "I'm curious if anybody is willing to discuss their validation approach - I tried random split (loose relationship with LB), time based split based on similar time length (i.e. creating a validation set spanning a period of similar length as test set - too few observations left), split based on observation count (okayish, but nothing to write home about - local validation score 0.2231, lb 0.2291).",
    "325775": "We have 10 random KFold and constant difference between CV and LB near 0.005.",
    "325778": "I really like seeing this thread so I can see if I'm overfitting. :) Right now, I'm doing 5-fold CV and have a consistent gap between CV and LB between 0.0031 and 0.0038. I tried a time-based validation strategy once but it made my score worse.",
    "326945": "Be careful with the last days of training data set:\n\ntraining[(training['activation_date']&gt;=pd.to_datetime('2017-03-29'))].shape\n\n(100, 17)\n\ntraining[(training['activation_date']&gt;=pd.to_datetime('2017-03-27'))].shape\n\n(227848, 17)\n\nOnly 100 samples after 2017-03-29.",
    "330580": "I am using this idea of stratified KFold on binned target:\n\n\n    import numpy as np\n    from sklearn.model_selection import train_test_split\n \n    def stratified_train_valid_split(meta_train,  target_column,  target_bins, valid_size, random_state=1234):\n\n        y = meta_train[target_column].values\n        bins = np.linspace(0, y.shape[0], target_bins)\n        y_binned = np.digitize(y, bins)\n\n       return train_test_split(meta_train, test_size=valid_size, stratify=y_binned, random_state=random_state)\n\nIt seems that the target distribution in train/valid is similiar. So far only tested it on bad models but it seems to correlate with the LB improvements. \n\nWhat do you think of this approach?",
    "336491": "I'm using March 26 - 28 as my validation set. Like most people here, I'm also getting ~0.005 delta. I'm still starting out though so not sure how well this will hold.\n\n![validation thingy][1]\n\n\n  [1]: https://i.imgur.com/rm9b2AL.png",
    "336350": "I am on about 0.004 CV/LB delta, but I am only starting out - will be interesting to see how this gap will develop as I add more magic/non-magic into the mix :) ",
    "331960": "I use valid set as 10% random split from train  set. It gives me  ~0.0034 CV-LB delta",
    "331033": "I confirm that using 5-fold random split results in a ~0.005 separation wrt LB. I tried decreasing and increasing the number of folds but 5 seems to be the optimal choice (also considering time).  \nI am also using the last 20% of the train dataset as validation, which results in a slightly smaller shift, around ~0.004",
    "328702": "I don't think random spliting (or CV) would be a good approach, because you would be using future data to predict the past, which most of the time gives a more optimistic accuracy.\n\nhttp://www.fast.ai/2017/11/13/validation-sets/",
    "327057": "I tried 8:2 random sample many times with different seed, but Validation loss didn't change more than 0.0001.\nIt seems that the train data well distributed, not same as testset.\nWhat's more, as the loss going to be lower, the gap seems to be lower. but still shows difference about 0.005.(bcuz my LB is still high)\n<br>\n| model || Val_loss | LB      | <br>\n| lgb      | | 0.2351 | 0.2412 | <br>\n| NN      | | 0.2588 | 0.3032 |  <br>\n| NN      | | 0.2475 | 0.2537 | <br>\n| NN      | |  0.2430 |  0.2488 | <br> \n| NN      | | 0.2407 |  0.2459 |  <br>\n\n<br>\nref : https://www.kaggle.com/dhznsdl/nn-model-adding-variables-step-by-step/code",
    "326188": "Anyone treating it as a time series problem?"
  }
}