{
  "id": 53126,
  "title": "Shuffled split significantly improves the score on LB",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53126",
  "author_name": "",
  "post_date": "2018-03-27T10:03:56.856498900Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello fellow Kagglers,</p>\n\n<p>I have done some performance bench-marking with different strategies to split validation data and would like to share my observation. </p>\n\n<p>Single lightGBM can go as far as <a href=\"https://www.kaggle.com/pranav84/single-lightgbm-in-r-with-75-mln-rows-lb-0-9690/code\">0.9690</a>  with <strong>shuffled split</strong>  compared to time based split. Also XGBoost with <a href=\"https://www.kaggle.com/pranav84/single-hist-xgboost-hitting-0-9684-on-lb\">histogram optimized</a>  version goes further and the gap between local cv and LB score is significantly reduced. </p>\n\n<p>Below is the code used in R for split:</p>\n\n<pre><code>#shuffled split\nsample =  sample.split(train$is_attributed, SplitRatio = 0.95)\ndtrain =  subset(train, sample == TRUE)\nvalid  =  subset(train, sample == FALSE)\n# LIghtGBM local cv:  0.9773 | LB: 0.9690\n# XGBoost  local cv:  0.9735 | LB: 0.9684\n\n#time based split\ntr_index &lt;- nrow(train)\ndtrain &lt;-  train %&gt;% head(0.95 * tr_index)\nvalid  &lt;-  train %&gt;% tail(0.05 * tr_index)\n# LIghtGBM local cv:  0.9841 | LB: 0.9685\n</code></pre>\n\n<p>Since we still have almost a month to go, please share your observations/ tips and let's try to hit 0.98 on public LB with single model with shared learning. I know 0.98 may be too much but may not be impossible with wisdom of crowd. I hope this is helpful. </p>",
  "messages": [
    {
      "id": "304291",
      "postDate": "03/27/2018 10:03:56",
      "content": "<p>Hello fellow Kagglers,</p>\n\n<p>I have done some performance bench-marking with different strategies to split validation data and would like to share my observation. </p>\n\n<p>Single lightGBM can go as far as <a href=\"https://www.kaggle.com/pranav84/single-lightgbm-in-r-with-75-mln-rows-lb-0-9690/code\">0.9690</a>  with <strong>shuffled split</strong>  compared to time based split. Also XGBoost with <a href=\"https://www.kaggle.com/pranav84/single-hist-xgboost-hitting-0-9684-on-lb\">histogram optimized</a>  version goes further and the gap between local cv and LB score is significantly reduced. </p>\n\n<p>Below is the code used in R for split:</p>\n\n<pre><code>#shuffled split\nsample =  sample.split(train$is_attributed, SplitRatio = 0.95)\ndtrain =  subset(train, sample == TRUE)\nvalid  =  subset(train, sample == FALSE)\n# LIghtGBM local cv:  0.9773 | LB: 0.9690\n# XGBoost  local cv:  0.9735 | LB: 0.9684\n\n#time based split\ntr_index &lt;- nrow(train)\ndtrain &lt;-  train %&gt;% head(0.95 * tr_index)\nvalid  &lt;-  train %&gt;% tail(0.05 * tr_index)\n# LIghtGBM local cv:  0.9841 | LB: 0.9685\n</code></pre>\n\n<p>Since we still have almost a month to go, please share your observations/ tips and let's try to hit 0.98 on public LB with single model with shared learning. I know 0.98 may be too much but may not be impossible with wisdom of crowd. I hope this is helpful. </p>",
      "rawMarkdown": "Hello fellow Kagglers,\n\nI have done some performance bench-marking with different strategies to split validation data and would like to share my observation. \n\nSingle lightGBM can go as far as [0.9690][1]  with **shuffled split**  compared to time based split. Also XGBoost with [histogram optimized][2]  version goes further and the gap between local cv and LB score is significantly reduced. \n\nBelow is the code used in R for split:\n    \n    #shuffled split\n    sample =  sample.split(train$is_attributed, SplitRatio = 0.95)\n    dtrain =  subset(train, sample == TRUE)\n    valid  =  subset(train, sample == FALSE)\n    # LIghtGBM local cv:  0.9773 | LB: 0.9690\n    # XGBoost  local cv:  0.9735 | LB: 0.9684\n\n    #time based split\n    tr_index &lt;- nrow(train)\n    dtrain &lt;-  train %&gt;% head(0.95 * tr_index)\n    valid  &lt;-  train %&gt;% tail(0.05 * tr_index)\n    # LIghtGBM local cv:  0.9841 | LB: 0.9685\n\nSince we still have almost a month to go, please share your observations/ tips and let's try to hit 0.98 on public LB with single model with shared learning. I know 0.98 may be too much but may not be impossible with wisdom of crowd. I hope this is helpful. \n\n  [1]: https://www.kaggle.com/pranav84/single-lightgbm-in-r-with-75-mln-rows-lb-0-9690/code\n  [2]: https://www.kaggle.com/pranav84/single-hist-xgboost-hitting-0-9684-on-lb",
      "votes": null
    },
    {
      "id": "304294",
      "postDate": "03/27/2018 10:10:16",
      "content": "<p>Use data between <code>2017-11-09 04:00:00</code> and <code>2017-11-09 05:00:00</code> as val data.\nGap between local cv and public lb should less than 0.001</p>",
      "rawMarkdown": "Use data between `2017-11-09 04:00:00` and `2017-11-09 05:00:00` as val data.\nGap between local cv and public lb should less than 0.001",
      "votes": null
    },
    {
      "id": "304297",
      "postDate": "03/27/2018 10:15:00",
      "content": "<p>That will overfit on public LB, because public LB split is most probably from same split as your validation split.</p>",
      "rawMarkdown": "That will overfit on public LB, because public LB split is most probably from same split as your validation split.",
      "votes": null
    },
    {
      "id": "304303",
      "postDate": "03/27/2018 10:29:51",
      "content": "<p>Yes, you are certainly correct. We could use two val data. </p>\n\n<ul>\n<li>val for public lb: between <code>2017-11-09 04:00:00</code> and <code>2017-11-09 05:00:00</code></li>\n<li>val for private lb: other corresponding time in test data.</li>\n</ul>\n\n<p>Since click events vary a lot at different hours. These two val scores should have some gap.</p>",
      "rawMarkdown": "Yes, you are certainly correct. We could use two val data. \n\n* val for public lb: between `2017-11-09 04:00:00` and `2017-11-09 05:00:00`\n* val for private lb: other corresponding time in test data.\n\nSince click events vary a lot at different hours. These two val scores should have some gap.",
      "votes": null
    },
    {
      "id": "304404",
      "postDate": "03/27/2018 14:35:49",
      "content": "<p>Bear in mind that, if you use the same training for prediction as you do for validation, then your choice of validation approach will affect the training set. I think that may be part of what is going on here: if you hold out the final samples for validation, you're excluding from training the data that are closest in time to the test data. (See my comment on your kernel.) A better test of validation procedures would be to validate both ways and do separate final predictions without validation, using parameters based on each of the validation results.</p>",
      "rawMarkdown": "Bear in mind that, if you use the same training for prediction as you do for validation, then your choice of validation approach will affect the training set. I think that may be part of what is going on here: if you hold out the final samples for validation, you're excluding from training the data that are closest in time to the test data. (See my comment on your kernel.) A better test of validation procedures would be to validate both ways and do separate final predictions without validation, using parameters based on each of the validation results.",
      "votes": null
    },
    {
      "id": "304415",
      "postDate": "03/27/2018 15:11:51",
      "content": "<p>Thank you so much for the detailed response and explanation about proper validation strategy. This will certainly help me concentrate in the right direction. Thanks again. </p>",
      "rawMarkdown": "Thank you so much for the detailed response and explanation about proper validation strategy. This will certainly help me concentrate in the right direction. Thanks again.",
      "votes": null
    },
    {
      "id": "304419",
      "postDate": "03/27/2018 15:21:24",
      "content": "<p>for me this is working good so far</p>\n\n<pre><code>fit the model on x0, evaluate on x1 - re-fit on complete train, predict on test\n</code></pre>\n\n<p>as  by @konard validation setting <a href=\"https://www.kaggle.com/konradb/validation-set#295343\">here</a></p>",
      "rawMarkdown": "for me this is working good so far\n\n    fit the model on x0, evaluate on x1 - re-fit on complete train, predict on test\n\n\nas  by @konard validation setting [here][1]\n\n\n  [1]: https://www.kaggle.com/konradb/validation-set#295343",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 304294,
      "author_name": "liujilong",
      "author_url": "",
      "post_date": "03/27/2018 10:10:16",
      "content": "<p>Use data between <code>2017-11-09 04:00:00</code> and <code>2017-11-09 05:00:00</code> as val data.\nGap between local cv and public lb should less than 0.001</p>",
      "votes": null,
      "replies": [
        {
          "id": 304297,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/27/2018 10:15:00",
          "content": "<p>That will overfit on public LB, because public LB split is most probably from same split as your validation split.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304303,
          "author_name": "liujilong",
          "author_url": "",
          "post_date": "03/27/2018 10:29:51",
          "content": "<p>Yes, you are certainly correct. We could use two val data. </p>\n\n<ul>\n<li>val for public lb: between <code>2017-11-09 04:00:00</code> and <code>2017-11-09 05:00:00</code></li>\n<li>val for private lb: other corresponding time in test data.</li>\n</ul>\n\n<p>Since click events vary a lot at different hours. These two val scores should have some gap.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 304404,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/27/2018 14:35:49",
      "content": "<p>Bear in mind that, if you use the same training for prediction as you do for validation, then your choice of validation approach will affect the training set. I think that may be part of what is going on here: if you hold out the final samples for validation, you're excluding from training the data that are closest in time to the test data. (See my comment on your kernel.) A better test of validation procedures would be to validate both ways and do separate final predictions without validation, using parameters based on each of the validation results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 304415,
          "author_name": "pranav84",
          "author_url": "",
          "post_date": "03/27/2018 15:11:51",
          "content": "<p>Thank you so much for the detailed response and explanation about proper validation strategy. This will certainly help me concentrate in the right direction. Thanks again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304419,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/27/2018 15:21:24",
          "content": "<p>for me this is working good so far</p>\n\n<pre><code>fit the model on x0, evaluate on x1 - re-fit on complete train, predict on test\n</code></pre>\n\n<p>as  by @konard validation setting <a href=\"https://www.kaggle.com/konradb/validation-set#295343\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "304291": "Hello fellow Kagglers,\n\nI have done some performance bench-marking with different strategies to split validation data and would like to share my observation. \n\nSingle lightGBM can go as far as [0.9690][1]  with **shuffled split**  compared to time based split. Also XGBoost with [histogram optimized][2]  version goes further and the gap between local cv and LB score is significantly reduced. \n\nBelow is the code used in R for split:\n    \n    #shuffled split\n    sample =  sample.split(train$is_attributed, SplitRatio = 0.95)\n    dtrain =  subset(train, sample == TRUE)\n    valid  =  subset(train, sample == FALSE)\n    # LIghtGBM local cv:  0.9773 | LB: 0.9690\n    # XGBoost  local cv:  0.9735 | LB: 0.9684\n\n    #time based split\n    tr_index &lt;- nrow(train)\n    dtrain &lt;-  train %&gt;% head(0.95 * tr_index)\n    valid  &lt;-  train %&gt;% tail(0.05 * tr_index)\n    # LIghtGBM local cv:  0.9841 | LB: 0.9685\n\nSince we still have almost a month to go, please share your observations/ tips and let's try to hit 0.98 on public LB with single model with shared learning. I know 0.98 may be too much but may not be impossible with wisdom of crowd. I hope this is helpful. \n\n  [1]: https://www.kaggle.com/pranav84/single-lightgbm-in-r-with-75-mln-rows-lb-0-9690/code\n  [2]: https://www.kaggle.com/pranav84/single-hist-xgboost-hitting-0-9684-on-lb",
    "304294": "Use data between `2017-11-09 04:00:00` and `2017-11-09 05:00:00` as val data.\nGap between local cv and public lb should less than 0.001",
    "304297": "That will overfit on public LB, because public LB split is most probably from same split as your validation split.",
    "304303": "Yes, you are certainly correct. We could use two val data. \n\n* val for public lb: between `2017-11-09 04:00:00` and `2017-11-09 05:00:00`\n* val for private lb: other corresponding time in test data.\n\nSince click events vary a lot at different hours. These two val scores should have some gap.",
    "304404": "Bear in mind that, if you use the same training for prediction as you do for validation, then your choice of validation approach will affect the training set. I think that may be part of what is going on here: if you hold out the final samples for validation, you're excluding from training the data that are closest in time to the test data. (See my comment on your kernel.) A better test of validation procedures would be to validate both ways and do separate final predictions without validation, using parameters based on each of the validation results.",
    "304415": "Thank you so much for the detailed response and explanation about proper validation strategy. This will certainly help me concentrate in the right direction. Thanks again.",
    "304419": "for me this is working good so far\n\n    fit the model on x0, evaluate on x1 - re-fit on complete train, predict on test\n\n\nas  by @konard validation setting [here][1]\n\n\n  [1]: https://www.kaggle.com/konradb/validation-set#295343"
  },
  "source": "meta"
}