{
  "id": 189437,
  "title": "Beware of target leakage",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189437",
  "author_name": "",
  "post_date": "2020-10-07T15:35:51.282318200Z",
  "votes": 151,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I see that a growing amount of people are using target encoding to extract features. Although this <em>might</em> seem to be work, you have to be careful about target leakage.</p>\n<p>As an example, if you calculate the average <code>answered_correctly</code> per <code>user_id</code>, then you'll be leaking information from the future into your average. This will make your model overly confident, which will be detrimental on the test set.</p>\n<p>In my book, there are two things you can do in order to improve your target encoding procedure:</p>\n<ol>\n<li>Calculate rolling averages that only include past values. I've written a blog post with an example that you can find <a href=\"https://maxhalford.github.io/blog/pandas-tricks/#target-encoding-for-time-series\" target=\"_blank\">here</a>.</li>\n<li>Use a Bayesian average to counterbalance the fact that some groups have too few instances. I've written a blog post which you read <a href=\"https://maxhalford.github.io/blog/target-encoding/\" target=\"_blank\">here</a>.</li>\n</ol>\n<p>I know it seems like I'm doing shameless publicity for my blog, but that's just a coincidence. I'm very interested in time series. In particular, I find it fascinating how they challenge the ways in which practictioners are used to doing machine learning.</p>\n<p>I think we can all agree that early kernels have a lot of influence on the competition because they get copy/pasted all over the place. This is my attempt to avoid target leakage inserting itself into everyone's code!</p>\n<p>Peace out ✌️</p>",
  "messages": [
    {
      "id": "1041175",
      "postDate": "10/07/2020 15:35:51",
      "content": "<p>I see that a growing amount of people are using target encoding to extract features. Although this <em>might</em> seem to be work, you have to be careful about target leakage.</p>\n<p>As an example, if you calculate the average <code>answered_correctly</code> per <code>user_id</code>, then you'll be leaking information from the future into your average. This will make your model overly confident, which will be detrimental on the test set.</p>\n<p>In my book, there are two things you can do in order to improve your target encoding procedure:</p>\n<ol>\n<li>Calculate rolling averages that only include past values. I've written a blog post with an example that you can find <a href=\"https://maxhalford.github.io/blog/pandas-tricks/#target-encoding-for-time-series\" target=\"_blank\">here</a>.</li>\n<li>Use a Bayesian average to counterbalance the fact that some groups have too few instances. I've written a blog post which you read <a href=\"https://maxhalford.github.io/blog/target-encoding/\" target=\"_blank\">here</a>.</li>\n</ol>\n<p>I know it seems like I'm doing shameless publicity for my blog, but that's just a coincidence. I'm very interested in time series. In particular, I find it fascinating how they challenge the ways in which practictioners are used to doing machine learning.</p>\n<p>I think we can all agree that early kernels have a lot of influence on the competition because they get copy/pasted all over the place. This is my attempt to avoid target leakage inserting itself into everyone's code!</p>\n<p>Peace out ✌️</p>",
      "rawMarkdown": "I see that a growing amount of people are using target encoding to extract features. Although this *might* seem to be work, you have to be careful about target leakage.\n\nAs an example, if you calculate the average `answered_correctly` per `user_id`, then you'll be leaking information from the future into your average. This will make your model overly confident, which will be detrimental on the test set.\n\nIn my book, there are two things you can do in order to improve your target encoding procedure:\n\n1. Calculate rolling averages that only include past values. I've written a blog post with an example that you can find [here](https://maxhalford.github.io/blog/pandas-tricks/#target-encoding-for-time-series).\n2. Use a Bayesian average to counterbalance the fact that some groups have too few instances. I've written a blog post which you read [here](https://maxhalford.github.io/blog/target-encoding/).\n\nI know it seems like I'm doing shameless publicity for my blog, but that's just a coincidence. I'm very interested in time series. In particular, I find it fascinating how they challenge the ways in which practictioners are used to doing machine learning.\n\nI think we can all agree that early kernels have a lot of influence on the competition because they get copy/pasted all over the place. This is my attempt to avoid target leakage inserting itself into everyone's code!\n\nPeace out ✌️",
      "votes": null
    },
    {
      "id": "1041409",
      "postDate": "10/07/2020 18:12:18",
      "content": "<p>Totally agreed. Relatedly, I think that traditional cross validation might be a mistake -- it may work well enough, but at least on paper it's definitely also future leaking. If students' early progress influences the questions they get routed to, I can easily see how a model might learn some overfit ways to infer the past with the future. I plan to start by holding out the end of users' question histories for validation instead of using random splits. </p>",
      "rawMarkdown": "Totally agreed. Relatedly, I think that traditional cross validation might be a mistake -- it may work well enough, but at least on paper it's definitely also future leaking. If students' early progress influences the questions they get routed to, I can easily see how a model might learn some overfit ways to infer the past with the future. I plan to start by holding out the end of users' question histories for validation instead of using random splits.",
      "votes": null
    },
    {
      "id": "1042683",
      "postDate": "10/08/2020 11:52:08",
      "content": "<p>Thanks for the advice :)</p>",
      "rawMarkdown": "Thanks for the advice :)",
      "votes": null
    },
    {
      "id": "1043145",
      "postDate": "10/08/2020 18:09:08",
      "content": "<p>I agree too. Thought to post similar ideas about leakage. Thanks for sharing. </p>",
      "rawMarkdown": "I agree too. Thought to post similar ideas about leakage. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1043356",
      "postDate": "10/08/2020 22:38:26",
      "content": "<p>I already have read your very useful blogs and discussed in IEEE fraud detection competition  here: <br>\n<a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/108311\" target=\"_blank\">IEEE Discussion</a></p>\n<p>Another point i have to mention is that, LGB uses \"Fisher Algorithm\" to encode categorical features, If you specify cats for the algorithm:</p>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/Features.html#optimal-split-for-categorical-features\" target=\"_blank\">Advanced Topics in LGB</a><br>\n<a href=\"https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html\" target=\"_blank\">Optimal Split for Categorical Features</a></p>\n<p>But when you send cat features directly to LGB you have better to use ('mindatapergroup' and 'catsmooth') in your LGB parameters( seems smoothing also has been implemented by LGB). By the way,  high cardinality of the cats in this comp will consume lots of memory during LGB training runtime ! </p>",
      "rawMarkdown": "I already have read your very useful blogs and discussed in IEEE fraud detection competition  here: \n[IEEE Discussion](https://www.kaggle.com/c/ieee-fraud-detection/discussion/108311)\n\nAnother point i have to mention is that, LGB uses \"Fisher Algorithm\" to encode categorical features, If you specify cats for the algorithm:\n\n[Advanced Topics in LGB]( https://lightgbm.readthedocs.io/en/latest/Features.html#optimal-split-for-categorical-features)\n[Optimal Split for Categorical Features](https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html)\n\n\n But when you send cat features directly to LGB you have better to use ('mindatapergroup' and 'catsmooth') in your LGB parameters( seems smoothing also has been implemented by LGB). By the way,  high cardinality of the cats in this comp will consume lots of memory during LGB training runtime !",
      "votes": null
    },
    {
      "id": "1045512",
      "postDate": "10/10/2020 17:57:49",
      "content": "<p>can u elaborate this a bit more if possible? <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> </p>",
      "rawMarkdown": "can u elaborate this a bit more if possible? @aquatic",
      "votes": null
    },
    {
      "id": "1046184",
      "postDate": "10/11/2020 12:22:27",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> you are correct! When I use standard CV my public score decreases even it gives 0.76 AUC validation score. But I didn't get how to do validation is this situation? Can you please elaborate more on that? Thanks!</p>",
      "rawMarkdown": "aquatic you are correct! When I use standard CV my public score decreases even it gives 0.76 AUC validation score. But I didn't get how to do validation is this situation? Can you please elaborate more on that? Thanks!",
      "votes": null
    },
    {
      "id": "1047446",
      "postDate": "10/12/2020 16:02:48",
      "content": "<p>Basically you should pre mark a validation set that is chronologically separate from your train set on a per user basis. ie. for any given user, highest timestamp in the train set &lt; lowest timestamp in the val set.</p>",
      "rawMarkdown": "Basically you should pre mark a validation set that is chronologically separate from your train set on a per user basis. ie. for any given user, highest timestamp in the train set < lowest timestamp in the val set.",
      "votes": null
    },
    {
      "id": "1048807",
      "postDate": "10/13/2020 19:46:49",
      "content": "<p>nice blog, congrats!</p>",
      "rawMarkdown": "nice blog, congrats!",
      "votes": null
    },
    {
      "id": "1048829",
      "postDate": "10/13/2020 20:19:56",
      "content": "<p>But then you wouldn't have new users on your validation set.<br>\nInstead of doing this, I've been creating a \"simulated timeline\", where each user is assigned a random starting point at the timeline, and each of their events is corrected by that starting point.<br>\nThen I can get my validation set by getting the last few weeks/months worth of events in this simulated timeline</p>",
      "rawMarkdown": "But then you wouldn't have new users on your validation set.\nInstead of doing this, I've been creating a \"simulated timeline\", where each user is assigned a random starting point at the timeline, and each of their events is corrected by that starting point.\nThen I can get my validation set by getting the last few weeks/months worth of events in this simulated timeline",
      "votes": null
    },
    {
      "id": "1053661",
      "postDate": "10/19/2020 07:57:20",
      "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/maxhalford\" target=\"_blank\">@maxhalford</a>  , I have just begin with this competition and also thought using mean of 'answered_correctly' based on 'user_id' groups , Thanks for Sharing Your blogs will check them out !</p>",
      "rawMarkdown": "Thanks alot @maxhalford  , I have just begin with this competition and also thought using mean of 'answered_correctly' based on 'user_id' groups , Thanks for Sharing Your blogs will check them out !",
      "votes": null
    },
    {
      "id": "1055605",
      "postDate": "10/21/2020 00:33:42",
      "content": "<p>Thanks for the good post!</p>",
      "rawMarkdown": "Thanks for the good post!",
      "votes": null
    },
    {
      "id": "1104618",
      "postDate": "12/07/2020 06:03:17",
      "content": "<p>This is very pertinent. Thanks for sharing the 2 notebooks! </p>",
      "rawMarkdown": "This is very pertinent. Thanks for sharing the 2 notebooks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041409,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/07/2020 18:12:18",
      "content": "<p>Totally agreed. Relatedly, I think that traditional cross validation might be a mistake -- it may work well enough, but at least on paper it's definitely also future leaking. If students' early progress influences the questions they get routed to, I can easily see how a model might learn some overfit ways to infer the past with the future. I plan to start by holding out the end of users' question histories for validation instead of using random splits. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1045512,
          "author_name": "elvinagammed",
          "author_url": "",
          "post_date": "10/10/2020 17:57:49",
          "content": "<p>can u elaborate this a bit more if possible? <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046184,
          "author_name": "kaushal2896",
          "author_url": "",
          "post_date": "10/11/2020 12:22:27",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> you are correct! When I use standard CV my public score decreases even it gives 0.76 AUC validation score. But I didn't get how to do validation is this situation? Can you please elaborate more on that? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047446,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/12/2020 16:02:48",
          "content": "<p>Basically you should pre mark a validation set that is chronologically separate from your train set on a per user basis. ie. for any given user, highest timestamp in the train set &lt; lowest timestamp in the val set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048829,
          "author_name": "fredcaroli",
          "author_url": "",
          "post_date": "10/13/2020 20:19:56",
          "content": "<p>But then you wouldn't have new users on your validation set.<br>\nInstead of doing this, I've been creating a \"simulated timeline\", where each user is assigned a random starting point at the timeline, and each of their events is corrected by that starting point.<br>\nThen I can get my validation set by getting the last few weeks/months worth of events in this simulated timeline</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042683,
      "author_name": "domizianostingi",
      "author_url": "",
      "post_date": "10/08/2020 11:52:08",
      "content": "<p>Thanks for the advice :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043145,
      "author_name": "vmk058",
      "author_url": "",
      "post_date": "10/08/2020 18:09:08",
      "content": "<p>I agree too. Thought to post similar ideas about leakage. Thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1043356,
      "author_name": "arashnic",
      "author_url": "",
      "post_date": "10/08/2020 22:38:26",
      "content": "<p>I already have read your very useful blogs and discussed in IEEE fraud detection competition  here: <br>\n<a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/108311\" target=\"_blank\">IEEE Discussion</a></p>\n<p>Another point i have to mention is that, LGB uses \"Fisher Algorithm\" to encode categorical features, If you specify cats for the algorithm:</p>\n<p><a href=\"https://lightgbm.readthedocs.io/en/latest/Features.html#optimal-split-for-categorical-features\" target=\"_blank\">Advanced Topics in LGB</a><br>\n<a href=\"https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html\" target=\"_blank\">Optimal Split for Categorical Features</a></p>\n<p>But when you send cat features directly to LGB you have better to use ('mindatapergroup' and 'catsmooth') in your LGB parameters( seems smoothing also has been implemented by LGB). By the way,  high cardinality of the cats in this comp will consume lots of memory during LGB training runtime ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1048807,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "10/13/2020 19:46:49",
      "content": "<p>nice blog, congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1053661,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "10/19/2020 07:57:20",
      "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/maxhalford\" target=\"_blank\">@maxhalford</a>  , I have just begin with this competition and also thought using mean of 'answered_correctly' based on 'user_id' groups , Thanks for Sharing Your blogs will check them out !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1055605,
      "author_name": "dongkyunkim",
      "author_url": "",
      "post_date": "10/21/2020 00:33:42",
      "content": "<p>Thanks for the good post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104618,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "12/07/2020 06:03:17",
      "content": "<p>This is very pertinent. Thanks for sharing the 2 notebooks! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1041175": "I see that a growing amount of people are using target encoding to extract features. Although this *might* seem to be work, you have to be careful about target leakage.\n\nAs an example, if you calculate the average `answered_correctly` per `user_id`, then you'll be leaking information from the future into your average. This will make your model overly confident, which will be detrimental on the test set.\n\nIn my book, there are two things you can do in order to improve your target encoding procedure:\n\n1. Calculate rolling averages that only include past values. I've written a blog post with an example that you can find [here](https://maxhalford.github.io/blog/pandas-tricks/#target-encoding-for-time-series).\n2. Use a Bayesian average to counterbalance the fact that some groups have too few instances. I've written a blog post which you read [here](https://maxhalford.github.io/blog/target-encoding/).\n\nI know it seems like I'm doing shameless publicity for my blog, but that's just a coincidence. I'm very interested in time series. In particular, I find it fascinating how they challenge the ways in which practictioners are used to doing machine learning.\n\nI think we can all agree that early kernels have a lot of influence on the competition because they get copy/pasted all over the place. This is my attempt to avoid target leakage inserting itself into everyone's code!\n\nPeace out ✌️",
    "1041409": "Totally agreed. Relatedly, I think that traditional cross validation might be a mistake -- it may work well enough, but at least on paper it's definitely also future leaking. If students' early progress influences the questions they get routed to, I can easily see how a model might learn some overfit ways to infer the past with the future. I plan to start by holding out the end of users' question histories for validation instead of using random splits.",
    "1042683": "Thanks for the advice :)",
    "1043145": "I agree too. Thought to post similar ideas about leakage. Thanks for sharing.",
    "1043356": "I already have read your very useful blogs and discussed in IEEE fraud detection competition  here: \n[IEEE Discussion](https://www.kaggle.com/c/ieee-fraud-detection/discussion/108311)\n\nAnother point i have to mention is that, LGB uses \"Fisher Algorithm\" to encode categorical features, If you specify cats for the algorithm:\n\n[Advanced Topics in LGB]( https://lightgbm.readthedocs.io/en/latest/Features.html#optimal-split-for-categorical-features)\n[Optimal Split for Categorical Features](https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html)\n\n\n But when you send cat features directly to LGB you have better to use ('mindatapergroup' and 'catsmooth') in your LGB parameters( seems smoothing also has been implemented by LGB). By the way,  high cardinality of the cats in this comp will consume lots of memory during LGB training runtime !",
    "1045512": "can u elaborate this a bit more if possible? @aquatic",
    "1046184": "aquatic you are correct! When I use standard CV my public score decreases even it gives 0.76 AUC validation score. But I didn't get how to do validation is this situation? Can you please elaborate more on that? Thanks!",
    "1047446": "Basically you should pre mark a validation set that is chronologically separate from your train set on a per user basis. ie. for any given user, highest timestamp in the train set < lowest timestamp in the val set.",
    "1048807": "nice blog, congrats!",
    "1048829": "But then you wouldn't have new users on your validation set.\nInstead of doing this, I've been creating a \"simulated timeline\", where each user is assigned a random starting point at the timeline, and each of their events is corrected by that starting point.\nThen I can get my validation set by getting the last few weeks/months worth of events in this simulated timeline",
    "1053661": "Thanks alot @maxhalford  , I have just begin with this competition and also thought using mean of 'answered_correctly' based on 'user_id' groups , Thanks for Sharing Your blogs will check them out !",
    "1055605": "Thanks for the good post!",
    "1104618": "This is very pertinent. Thanks for sharing the 2 notebooks!"
  },
  "source": "meta"
}