{
  "id": 207460,
  "title": "Task_container_id and data leakage.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207460",
  "author_name": "",
  "post_date": "2020-12-29T19:30:17.624070800Z",
  "votes": null,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Right now, my single LGBM model achieves 0.813 CV with 12m training and 3m validation sets and 24 features. My question is:</p>\n<p>I have not taken task_container_id into consideration till now, therefore there is a bit of data leakage. Do you think it's realistic to get 0.813 on a single LGBM model with 24 features? I'm having difficulties with inference therefore i have not been able to submit my model yet. A previous model with 0.806 CV is running for submission for the last 3 hours. My notebooks usually get a scoring error after 5-6 hours of running, this is why i'm asking here if this data leakage is a serious problem, it will be long before i figure it out myself.</p>",
  "messages": [
    {
      "id": "1131519",
      "postDate": "12/29/2020 19:30:17",
      "content": "<p>Right now, my single LGBM model achieves 0.813 CV with 12m training and 3m validation sets and 24 features. My question is:</p>\n<p>I have not taken task_container_id into consideration till now, therefore there is a bit of data leakage. Do you think it's realistic to get 0.813 on a single LGBM model with 24 features? I'm having difficulties with inference therefore i have not been able to submit my model yet. A previous model with 0.806 CV is running for submission for the last 3 hours. My notebooks usually get a scoring error after 5-6 hours of running, this is why i'm asking here if this data leakage is a serious problem, it will be long before i figure it out myself.</p>",
      "rawMarkdown": "Right now, my single LGBM model achieves 0.813 CV with 12m training and 3m validation sets and 24 features. My question is:\n\nI have not taken task_container_id into consideration till now, therefore there is a bit of data leakage. Do you think it's realistic to get 0.813 on a single LGBM model with 24 features? I'm having difficulties with inference therefore i have not been able to submit my model yet. A previous model with 0.806 CV is running for submission for the last 3 hours. My notebooks usually get a scoring error after 5-6 hours of running, this is why i'm asking here if this data leakage is a serious problem, it will be long before i figure it out myself.",
      "votes": null
    },
    {
      "id": "1131547",
      "postDate": "12/29/2020 19:57:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bilgicbe\" target=\"_blank\">@bilgicbe</a> , I think it's not realistic, unless you found a magic feature. Did you remove the column \"user_answer\" from the data before training?</p>",
      "rawMarkdown": "Hi @bilgicbe , I think it's not realistic, unless you found a magic feature. Did you remove the column \"user_answer\" from the data before training?",
      "votes": null
    },
    {
      "id": "1131564",
      "postDate": "12/29/2020 20:09:18",
      "content": "<p><a href=\"https://www.kaggle.com/stmandl\" target=\"_blank\">@stmandl</a> You comment made me laugh :) yes i did.</p>",
      "rawMarkdown": "stmandl You comment made me laugh :) yes i did.",
      "votes": null
    },
    {
      "id": "1131566",
      "postDate": "12/29/2020 20:12:07",
      "content": "<p>How do you pick your validation data?</p>",
      "rawMarkdown": "How do you pick your validation data?",
      "votes": null
    },
    {
      "id": "1131571",
      "postDate": "12/29/2020 20:15:39",
      "content": "<p>Shuffle and sample 15m, train_test_split with 0.2 test size. I use random_state to be consistent so i can keep track of my progress.</p>",
      "rawMarkdown": "Shuffle and sample 15m, train_test_split with 0.2 test size. I use random_state to be consistent so i can keep track of my progress.",
      "votes": null
    },
    {
      "id": "1131576",
      "postDate": "12/29/2020 20:21:39",
      "content": "<p>It is bad idea to shuffle timeseries</p>",
      "rawMarkdown": "It is bad idea to shuffle timeseries",
      "votes": null
    },
    {
      "id": "1131586",
      "postDate": "12/29/2020 20:27:07",
      "content": "<p>That's good 👍😄 </p>",
      "rawMarkdown": "That's good 👍😄",
      "votes": null
    },
    {
      "id": "1131589",
      "postDate": "12/29/2020 20:29:10",
      "content": "<p><a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> Would you explain why? I have the assumption that LGBM treats every row independently.</p>\n<p>I only shuffle after feature engineering, before training.</p>",
      "rawMarkdown": "fredegrec Would you explain why? I have the assumption that LGBM treats every row independently.\n\nI only shuffle after feature engineering, before training.",
      "votes": null
    },
    {
      "id": "1131592",
      "postDate": "12/29/2020 20:31:34",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> is right. I personally make the test/val split by on the distinct values of the user_id column.</p>",
      "rawMarkdown": "Yes @fredegrec is right. I personally make the test/val split by on the distinct values of the user_id column.",
      "votes": null
    },
    {
      "id": "1131596",
      "postDate": "12/29/2020 20:34:03",
      "content": "<p>It does, but still you will have future samples in your training set for at least part of your validation data. Not sure how much of an impact that has here though.</p>",
      "rawMarkdown": "It does, but still you will have future samples in your training set for at least part of your validation data. Not sure how much of an impact that has here though.",
      "votes": null
    },
    {
      "id": "1131597",
      "postDate": "12/29/2020 20:38:18",
      "content": "<p><a href=\"https://www.kaggle.com/bilgicbe\" target=\"_blank\">@bilgicbe</a> <br>\nSome lines in the train will have information from the future </p>",
      "rawMarkdown": "bilgicbe \nSome lines in the train will have information from the future",
      "votes": null
    },
    {
      "id": "1131625",
      "postDate": "12/29/2020 21:25:25",
      "content": "<p><a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> No row in my train dataset has information from the future, except for the same task containers.</p>",
      "rawMarkdown": "fredegrec No row in my train dataset has information from the future, except for the same task containers.",
      "votes": null
    },
    {
      "id": "1131704",
      "postDate": "12/29/2020 22:40:40",
      "content": "<p>Did you shift your cumulative features forward after computing them? That might a source of data leakage. 0.800 + in LGBM requires serious feature engineering.</p>",
      "rawMarkdown": "Did you shift your cumulative features forward after computing them? That might a source of data leakage. 0.800 + in LGBM requires serious feature engineering.",
      "votes": null
    },
    {
      "id": "1131721",
      "postDate": "12/29/2020 22:58:23",
      "content": "<p>I used loops, i checked manually if my features are using information from the future, they don't, except for task_container_id, which can be a problem. My last submission failed again, i'm trying to debug the pipeline, it seems i can never be sure without a proper submission.</p>",
      "rawMarkdown": "I used loops, i checked manually if my features are using information from the future, they don't, except for task_container_id, which can be a problem. My last submission failed again, i'm trying to debug the pipeline, it seems i can never be sure without a proper submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1131547,
      "author_name": "stmandl",
      "author_url": "",
      "post_date": "12/29/2020 19:57:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bilgicbe\" target=\"_blank\">@bilgicbe</a> , I think it's not realistic, unless you found a magic feature. Did you remove the column \"user_answer\" from the data before training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131564,
          "author_name": "bilgicbe",
          "author_url": "",
          "post_date": "12/29/2020 20:09:18",
          "content": "<p><a href=\"https://www.kaggle.com/stmandl\" target=\"_blank\">@stmandl</a> You comment made me laugh :) yes i did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131586,
          "author_name": "stmandl",
          "author_url": "",
          "post_date": "12/29/2020 20:27:07",
          "content": "<p>That's good 👍😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131566,
      "author_name": "nicohrubec",
      "author_url": "",
      "post_date": "12/29/2020 20:12:07",
      "content": "<p>How do you pick your validation data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131571,
          "author_name": "bilgicbe",
          "author_url": "",
          "post_date": "12/29/2020 20:15:39",
          "content": "<p>Shuffle and sample 15m, train_test_split with 0.2 test size. I use random_state to be consistent so i can keep track of my progress.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131576,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "12/29/2020 20:21:39",
          "content": "<p>It is bad idea to shuffle timeseries</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131589,
          "author_name": "bilgicbe",
          "author_url": "",
          "post_date": "12/29/2020 20:29:10",
          "content": "<p><a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> Would you explain why? I have the assumption that LGBM treats every row independently.</p>\n<p>I only shuffle after feature engineering, before training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131592,
          "author_name": "stmandl",
          "author_url": "",
          "post_date": "12/29/2020 20:31:34",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> is right. I personally make the test/val split by on the distinct values of the user_id column.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131596,
          "author_name": "nicohrubec",
          "author_url": "",
          "post_date": "12/29/2020 20:34:03",
          "content": "<p>It does, but still you will have future samples in your training set for at least part of your validation data. Not sure how much of an impact that has here though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131597,
          "author_name": "fredegrec",
          "author_url": "",
          "post_date": "12/29/2020 20:38:18",
          "content": "<p><a href=\"https://www.kaggle.com/bilgicbe\" target=\"_blank\">@bilgicbe</a> <br>\nSome lines in the train will have information from the future </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131625,
          "author_name": "bilgicbe",
          "author_url": "",
          "post_date": "12/29/2020 21:25:25",
          "content": "<p><a href=\"https://www.kaggle.com/fredegrec\" target=\"_blank\">@fredegrec</a> No row in my train dataset has information from the future, except for the same task containers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131704,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "12/29/2020 22:40:40",
      "content": "<p>Did you shift your cumulative features forward after computing them? That might a source of data leakage. 0.800 + in LGBM requires serious feature engineering.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131721,
          "author_name": "bilgicbe",
          "author_url": "",
          "post_date": "12/29/2020 22:58:23",
          "content": "<p>I used loops, i checked manually if my features are using information from the future, they don't, except for task_container_id, which can be a problem. My last submission failed again, i'm trying to debug the pipeline, it seems i can never be sure without a proper submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1131519": "Right now, my single LGBM model achieves 0.813 CV with 12m training and 3m validation sets and 24 features. My question is:\n\nI have not taken task_container_id into consideration till now, therefore there is a bit of data leakage. Do you think it's realistic to get 0.813 on a single LGBM model with 24 features? I'm having difficulties with inference therefore i have not been able to submit my model yet. A previous model with 0.806 CV is running for submission for the last 3 hours. My notebooks usually get a scoring error after 5-6 hours of running, this is why i'm asking here if this data leakage is a serious problem, it will be long before i figure it out myself.",
    "1131547": "Hi @bilgicbe , I think it's not realistic, unless you found a magic feature. Did you remove the column \"user_answer\" from the data before training?",
    "1131564": "stmandl You comment made me laugh :) yes i did.",
    "1131566": "How do you pick your validation data?",
    "1131571": "Shuffle and sample 15m, train_test_split with 0.2 test size. I use random_state to be consistent so i can keep track of my progress.",
    "1131576": "It is bad idea to shuffle timeseries",
    "1131586": "That's good 👍😄",
    "1131589": "fredegrec Would you explain why? I have the assumption that LGBM treats every row independently.\n\nI only shuffle after feature engineering, before training.",
    "1131592": "Yes @fredegrec is right. I personally make the test/val split by on the distinct values of the user_id column.",
    "1131596": "It does, but still you will have future samples in your training set for at least part of your validation data. Not sure how much of an impact that has here though.",
    "1131597": "bilgicbe \nSome lines in the train will have information from the future",
    "1131625": "fredegrec No row in my train dataset has information from the future, except for the same task containers.",
    "1131704": "Did you shift your cumulative features forward after computing them? That might a source of data leakage. 0.800 + in LGBM requires serious feature engineering.",
    "1131721": "I used loops, i checked manually if my features are using information from the future, they don't, except for task_container_id, which can be a problem. My last submission failed again, i'm trying to debug the pipeline, it seems i can never be sure without a proper submission."
  },
  "source": "meta"
}