{
  "id": 189669,
  "title": "Validation ideas to avoid leaks",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189669",
  "author_name": "",
  "post_date": "2020-10-08T08:53:17.243902600Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This competition uses time-series data.<br>\nAlso, a non-leaking target encoding is proposed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437\" target=\"_blank\">here</a> on discussion.  </p>\n<p>I think it is also <strong>important to create good validations</strong>.<br>\nI had two ideas. Let me explain.</p>\n<p>The vertical axis is the user's ID and the horizontal axis is the time.<br>\n(light blue: training set, red: test set and assume there are no leaks in target encoding.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F67e2642c2a25bfcdf320edd228a881e5%2F2020-10-08%2017-40-04.png?generation=1602146418700915&amp;alt=media\" alt=\"\"><br>\n<strong>Separate for each user</strong></p>\n<p>GOOD POINT</p>\n<ul>\n<li>easy to implement</li>\n</ul>\n<p>BAD POINT</p>\n<ul>\n<li>sometimes the tests don't match the predictions<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F1c831a8886c77840a5f2474c458d3ecc%2F2020-10-08%2017-38-49.png?generation=1602146362483094&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>GOOD POINT  </p>\n<ul>\n<li>close to the test data</li>\n</ul>\n<p>BAD POINT</p>\n<ul>\n<li>difficult to implement</li>\n</ul>\n<p>What is your opinion on these?<br>\n<strong>Please share your thoughts!</strong></p>",
  "messages": [
    {
      "id": "1042479",
      "postDate": "10/08/2020 08:53:17",
      "content": "<p>This competition uses time-series data.<br>\nAlso, a non-leaking target encoding is proposed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437\" target=\"_blank\">here</a> on discussion.  </p>\n<p>I think it is also <strong>important to create good validations</strong>.<br>\nI had two ideas. Let me explain.</p>\n<p>The vertical axis is the user's ID and the horizontal axis is the time.<br>\n(light blue: training set, red: test set and assume there are no leaks in target encoding.)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F67e2642c2a25bfcdf320edd228a881e5%2F2020-10-08%2017-40-04.png?generation=1602146418700915&amp;alt=media\" alt=\"\"><br>\n<strong>Separate for each user</strong></p>\n<p>GOOD POINT</p>\n<ul>\n<li>easy to implement</li>\n</ul>\n<p>BAD POINT</p>\n<ul>\n<li>sometimes the tests don't match the predictions<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F1c831a8886c77840a5f2474c458d3ecc%2F2020-10-08%2017-38-49.png?generation=1602146362483094&amp;alt=media\" alt=\"\"></li>\n</ul>\n<p>GOOD POINT  </p>\n<ul>\n<li>close to the test data</li>\n</ul>\n<p>BAD POINT</p>\n<ul>\n<li>difficult to implement</li>\n</ul>\n<p>What is your opinion on these?<br>\n<strong>Please share your thoughts!</strong></p>",
      "rawMarkdown": "This competition uses time-series data.\n\nAlso, a non-leaking target encoding is proposed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437) on discussion.  \n  \nI think it is also **important to create good validations**.\nI had two ideas. Let me explain.\n  \nThe vertical axis is the user's ID and the horizontal axis is the time.\n(light blue: training set, red: test set and assume there are no leaks in target encoding.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F67e2642c2a25bfcdf320edd228a881e5%2F2020-10-08%2017-40-04.png?generation=1602146418700915&alt=media)\n\n**Separate for each user**\n  \nGOOD POINT\n  \n- easy to implement\n  \nBAD POINT\n- sometimes the tests don't match the predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F1c831a8886c77840a5f2474c458d3ecc%2F2020-10-08%2017-38-49.png?generation=1602146362483094&alt=media)\n  \nGOOD POINT  \n- close to the test data\n  \nBAD POINT\n- difficult to implement\n  \n \nWhat is your opinion on these?\n**Please share your thoughts!**",
      "votes": null
    },
    {
      "id": "1045432",
      "postDate": "10/10/2020 16:17:38",
      "content": "<p>I think it depends on how to split train/public/private, and I don't know the split information now…<br>\nDo you know that?</p>",
      "rawMarkdown": "I think it depends on how to split train/public/private, and I don't know the split information now...\nDo you know that?",
      "votes": null
    },
    {
      "id": "1045694",
      "postDate": "10/10/2020 22:58:34",
      "content": "<p>Thanks for sharing with me.<br>\nNo, I don't. And you are right.</p>\n<p>I'm going to avoid using future data when creating features.</p>",
      "rawMarkdown": "Thanks for sharing with me.\nNo, I don't. And you are right.\n\nI'm going to avoid using future data when creating features.",
      "votes": null
    },
    {
      "id": "1046568",
      "postDate": "10/11/2020 19:35:52",
      "content": "<p>Time-series (or other intrinsically ordered data) can be problematic for cross-validation. If some pattern emerges in year 3 and stays for years 4-6, then your model can pick up on it, even though it wasn't part of years 1 &amp; 2.</p>\n<p>An approach that's sometimes more principled for time series is forward chaining, where your procedure would be something like this:</p>\n<p>fold 1 : training [1], test [2]<br>\nfold 2 : training [1 2], test [3]<br>\nfold 3 : training [1 2 3], test [4]<br>\nfold 4 : training [1 2 3 4], test [5]<br>\nfold 5 : training [1 2 3 4 5], test [6]<br>\nThat more accurately models the situation you'll see at prediction time, where you'll model on past data and predict on forward-looking data. It also will give you a sense of the dependence of your modeling on data size.</p>",
      "rawMarkdown": "Time-series (or other intrinsically ordered data) can be problematic for cross-validation. If some pattern emerges in year 3 and stays for years 4-6, then your model can pick up on it, even though it wasn't part of years 1 & 2.\n\nAn approach that's sometimes more principled for time series is forward chaining, where your procedure would be something like this:\n\nfold 1 : training [1], test [2]\nfold 2 : training [1 2], test [3]\nfold 3 : training [1 2 3], test [4]\nfold 4 : training [1 2 3 4], test [5]\nfold 5 : training [1 2 3 4 5], test [6]\nThat more accurately models the situation you'll see at prediction time, where you'll model on past data and predict on forward-looking data. It also will give you a sense of the dependence of your modeling on data size.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1045432,
      "author_name": "go5kuramubon",
      "author_url": "",
      "post_date": "10/10/2020 16:17:38",
      "content": "<p>I think it depends on how to split train/public/private, and I don't know the split information now…<br>\nDo you know that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045694,
          "author_name": "chizuchizu",
          "author_url": "",
          "post_date": "10/10/2020 22:58:34",
          "content": "<p>Thanks for sharing with me.<br>\nNo, I don't. And you are right.</p>\n<p>I'm going to avoid using future data when creating features.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046568,
      "author_name": "elvinagammed",
      "author_url": "",
      "post_date": "10/11/2020 19:35:52",
      "content": "<p>Time-series (or other intrinsically ordered data) can be problematic for cross-validation. If some pattern emerges in year 3 and stays for years 4-6, then your model can pick up on it, even though it wasn't part of years 1 &amp; 2.</p>\n<p>An approach that's sometimes more principled for time series is forward chaining, where your procedure would be something like this:</p>\n<p>fold 1 : training [1], test [2]<br>\nfold 2 : training [1 2], test [3]<br>\nfold 3 : training [1 2 3], test [4]<br>\nfold 4 : training [1 2 3 4], test [5]<br>\nfold 5 : training [1 2 3 4 5], test [6]<br>\nThat more accurately models the situation you'll see at prediction time, where you'll model on past data and predict on forward-looking data. It also will give you a sense of the dependence of your modeling on data size.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1042479": "This competition uses time-series data.\n\nAlso, a non-leaking target encoding is proposed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437) on discussion.  \n  \nI think it is also **important to create good validations**.\nI had two ideas. Let me explain.\n  \nThe vertical axis is the user's ID and the horizontal axis is the time.\n(light blue: training set, red: test set and assume there are no leaks in target encoding.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F67e2642c2a25bfcdf320edd228a881e5%2F2020-10-08%2017-40-04.png?generation=1602146418700915&alt=media)\n\n**Separate for each user**\n  \nGOOD POINT\n  \n- easy to implement\n  \nBAD POINT\n- sometimes the tests don't match the predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2523201%2F1c831a8886c77840a5f2474c458d3ecc%2F2020-10-08%2017-38-49.png?generation=1602146362483094&alt=media)\n  \nGOOD POINT  \n- close to the test data\n  \nBAD POINT\n- difficult to implement\n  \n \nWhat is your opinion on these?\n**Please share your thoughts!**",
    "1045432": "I think it depends on how to split train/public/private, and I don't know the split information now...\nDo you know that?",
    "1045694": "Thanks for sharing with me.\nNo, I don't. And you are right.\n\nI'm going to avoid using future data when creating features.",
    "1046568": "Time-series (or other intrinsically ordered data) can be problematic for cross-validation. If some pattern emerges in year 3 and stays for years 4-6, then your model can pick up on it, even though it wasn't part of years 1 & 2.\n\nAn approach that's sometimes more principled for time series is forward chaining, where your procedure would be something like this:\n\nfold 1 : training [1], test [2]\nfold 2 : training [1 2], test [3]\nfold 3 : training [1 2 3], test [4]\nfold 4 : training [1 2 3 4], test [5]\nfold 5 : training [1 2 3 4 5], test [6]\nThat more accurately models the situation you'll see at prediction time, where you'll model on past data and predict on forward-looking data. It also will give you a sense of the dependence of your modeling on data size."
  },
  "source": "meta"
}