{
  "id": 55147,
  "title": "attributed frequencies overfitting",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55147",
  "author_name": "Edward Chen",
  "post_date": "2018-04-22T19:05:07.442000",
  "votes": 11,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I've been testing a few features and noticed that features that calculate the frequency of is_attributed all seem to overfit the training data. My thinking behind this is that, since most of the apps are not download, the model is trained to simply predict 0 for any frequencies of 0 and 1 for any frequencies &gt; 0, and since there are so many 0s the model overfits. </p>\n\n<p>Is my reasoning behind that correct or am I misunderstanding something? Also, would any feature that uses is_attributed as a frequency be considered target encoding? Is it possible for target encoding to not overfit?</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 317889,
      "postDate": "2018-04-22T19:05:07.443Z",
      "content": "<p>I've been testing a few features and noticed that features that calculate the frequency of is_attributed all seem to overfit the training data. My thinking behind this is that, since most of the apps are not download, the model is trained to simply predict 0 for any frequencies of 0 and 1 for any frequencies &gt; 0, and since there are so many 0s the model overfits. </p>\n\n<p>Is my reasoning behind that correct or am I misunderstanding something? Also, would any feature that uses is_attributed as a frequency be considered target encoding? Is it possible for target encoding to not overfit?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I've been testing a few features and noticed that features that calculate the frequency of is_attributed all seem to overfit the training data. My thinking behind this is that, since most of the apps are not download, the model is trained to simply predict 0 for any frequencies of 0 and 1 for any frequencies &gt; 0, and since there are so many 0s the model overfits. \n\nIs my reasoning behind that correct or am I misunderstanding something? Also, would any feature that uses is_attributed as a frequency be considered target encoding? Is it possible for target encoding to not overfit?\n\nThanks!",
      "votes": 11
    },
    {
      "id": 317893,
      "postDate": "2018-04-22T19:23:07.973Z",
      "content": "<p>Yes, this is called 'leaking target information'.   One way to avoid this is to compute target average on some part of the data, and use that to create feature on another part of the data.  </p>",
      "rawMarkdown": "Yes, this is called 'leaking target information'.   One way to avoid this is to compute target average on some part of the data, and use that to create feature on another part of the data.  ",
      "votes": 10,
      "replies": [
        {
          "id": 317894,
          "postDate": "2018-04-22T19:39:32.490Z",
          "content": "<p>Thank you very much for your reply! What you mentioned sounds interesting though. To clarify my understanding, could this be an example of that?</p>\n\n<p>Calculate frequencies of is_attributed for each ip - let's call it \"ip_freq\"\nUse \"ip_freq\" to create some other feature x. When training, only use feature x and not feature ip_freq.</p>\n\n<p>Is that what you mean by \"create feature on another part of the data\"? I can't seem to think of how you could create the feature x, according to my example, and not leak anything about the target either</p>",
          "rawMarkdown": "Thank you very much for your reply! What you mentioned sounds interesting though. To clarify my understanding, could this be an example of that?\n\nCalculate frequencies of is_attributed for each ip - let's call it \"ip_freq\"\nUse \"ip_freq\" to create some other feature x. When training, only use feature x and not feature ip_freq.\n\nIs that what you mean by \"create feature on another part of the data\"? I can't seem to think of how you could create the feature x, according to my example, and not leak anything about the target either"
        },
        {
          "id": 317925,
          "postDate": "2018-04-22T21:11:13.297Z",
          "content": "<p>Another technique would be to make your target encoding a bit more noisy. Two example methods to accomplish this:</p>\n\n<ol>\n<li>rather than using ALL your training data, use a subset of the training data to do target encoding. Your different folds can use different bags if you want to capture all the signal.</li>\n<li>or alternatively, add some gauss noise to it. figuring out the optimal level of noise should be a function of your validation scheme</li>\n</ol>",
          "rawMarkdown": "Another technique would be to make your target encoding a bit more noisy. Two example methods to accomplish this:\n\n1. rather than using ALL your training data, use a subset of the training data to do target encoding. Your different folds can use different bags if you want to capture all the signal.\n2. or alternatively, add some gauss noise to it. figuring out the optimal level of noise should be a function of your validation scheme"
        },
        {
          "id": 318195,
          "postDate": "2018-04-23T11:24:49.223Z",
          "content": "<p>@Edward, not, this is not what I mean.  By 'part of the data' I mean some split of the data by rows, not by columns.  If you compute target frequency for a row using the rwo (and other rows), then you leak the target information of the row into your new feature.</p>",
          "rawMarkdown": "@Edward, not, this is not what I mean.  By 'part of the data' I mean some split of the data by rows, not by columns.  If you compute target frequency for a row using the rwo (and other rows), then you leak the target information of the row into your new feature.",
          "votes": 2
        },
        {
          "id": 318400,
          "postDate": "2018-04-23T18:52:07.177Z",
          "content": "<p><a href=\"/authman\">@authman</a> Thank you for your ideas! If I'm understanding correctly, your 1st point is very similar to what CPMP is mentioning right?</p>\n\n<p>@CPMP Thanks for the clarification! To clarify my understanding, so, for example, I could compute the attributed frequency for some feature 'x' only for day 'y' in the training data but then merge those frequencies into the other days in the training set. So in that case, would it still be okay to use day 'y' as part of training? I can see that the overfitting would be a lot less since there would technically be more noise added right?</p>",
          "rawMarkdown": "@authman Thank you for your ideas! If I'm understanding correctly, your 1st point is very similar to what CPMP is mentioning right?\n\n@CPMP Thanks for the clarification! To clarify my understanding, so, for example, I could compute the attributed frequency for some feature 'x' only for day 'y' in the training data but then merge those frequencies into the other days in the training set. So in that case, would it still be okay to use day 'y' as part of training? I can see that the overfitting would be a lot less since there would technically be more noise added right?"
        },
        {
          "id": 318432,
          "postDate": "2018-04-23T20:28:19.380Z",
          "content": "<p>@Edward, you got it.</p>",
          "rawMarkdown": "@Edward, you got it."
        },
        {
          "id": 318471,
          "postDate": "2018-04-23T22:21:25.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 318809,
          "postDate": "2018-04-24T14:09:43.993Z",
          "content": "<p>Here is a (so far under the radar but likely about to blow up) compact, runnable kernel of N.2 above: <a href=\"https://www.kaggle.com/scirpus/a-bit-more-target-encoding-with-gp\">https://www.kaggle.com/scirpus/a-bit-more-target-encoding-with-gp</a></p>",
          "rawMarkdown": "Here is a (so far under the radar but likely about to blow up) compact, runnable kernel of N.2 above: https://www.kaggle.com/scirpus/a-bit-more-target-encoding-with-gp"
        },
        {
          "id": 318914,
          "postDate": "2018-04-24T18:22:46.623Z",
          "content": "<p>Hi, CPMP. I am also try to use this attributed frequency feature. My trainset is about 100millionrows. How many rows of the subset I should choose is appropriate?</p>",
          "rawMarkdown": "Hi, CPMP. I am also try to use this attributed frequency feature. My trainset is about 100millionrows. How many rows of the subset I should choose is appropriate?"
        }
      ]
    },
    {
      "id": 318673,
      "postDate": "2018-04-24T08:51:03.563Z",
      "content": "<p>That was my experience too.\nWhat @CPMP suggested is what i did up to a few days ago, I did target encoding for each day separately, using the previous days: day 8 gets the ratios of day 7, day 9 gets ratios of day 8+ 7  and test data day 10 gets ratios of all previous days (all training). this decreased somewhat the amount of leakage and overfitting, but took a long time to preprocess and removed the ability to train on day 7, which is pretty costly.</p>\n\n<p>Play with it and see how it goes for you =)\nwill be happy to know how it goes.</p>",
      "rawMarkdown": "That was my experience too.\nWhat @CPMP suggested is what i did up to a few days ago, I did target encoding for each day separately, using the previous days: day 8 gets the ratios of day 7, day 9 gets ratios of day 8+ 7  and test data day 10 gets ratios of all previous days (all training). this decreased somewhat the amount of leakage and overfitting, but took a long time to preprocess and removed the ability to train on day 7, which is pretty costly.\n\nPlay with it and see how it goes for you =)\nwill be happy to know how it goes.",
      "votes": 2,
      "replies": [
        {
          "id": 319393,
          "postDate": "2018-04-26T00:41:31.437Z",
          "content": "<p>DO THESE FEATURES WORK WELL IN YOUR MODEL ？</p>",
          "rawMarkdown": "DO THESE FEATURES WORK WELL IN YOUR MODEL ？",
          "votes": -1
        },
        {
          "id": 319666,
          "postDate": "2018-04-26T15:02:19.830Z",
          "content": "<p>hard question, i'll tell you this - They were in top in feature importance, but i ended up removing them for different reasons. took a long time to process and i suspect they were overfitting my model</p>",
          "rawMarkdown": "hard question, i'll tell you this - They were in top in feature importance, but i ended up removing them for different reasons. took a long time to process and i suspect they were overfitting my model",
          "votes": 1
        },
        {
          "id": 319670,
          "postDate": "2018-04-26T15:10:03.873Z",
          "content": "<p>Same as you.Especially ip related features overfit model .</p>",
          "rawMarkdown": "Same as you.Especially ip related features overfit model .",
          "votes": 1
        }
      ]
    },
    {
      "id": 318530,
      "postDate": "2018-04-24T02:35:45.553Z",
      "content": "<p>Hey,</p>\n\n<p>I am also thinking of using target encoding as one of the features. </p>\n\n<p>To solve your problems you can maybe try using cumulative count instead of summing throughout the entire training set. This has less amount of leakage.</p>\n\n<p>However, one problem I am trying to tackle at the moment is: How do you create that feature for the test set? Do you just shift the feature down by number of test set and discard the start of training set?</p>",
      "rawMarkdown": "Hey,\n\nI am also thinking of using target encoding as one of the features. \n\nTo solve your problems you can maybe try using cumulative count instead of summing throughout the entire training set. This has less amount of leakage.\n\nHowever, one problem I am trying to tackle at the moment is: How do you create that feature for the test set? Do you just shift the feature down by number of test set and discard the start of training set?\n\n",
      "replies": [
        {
          "id": 318556,
          "postDate": "2018-04-24T03:38:04.453Z",
          "content": "<p>Hey Budi, I've thought of that exact same dilemma. Honestly I'm not too sure how to approach it and have yet to find a solution. I'm still considering trying different ways in the near future though and will update if I do. Please update as well if you do end up finding a way!</p>",
          "rawMarkdown": "Hey Budi, I've thought of that exact same dilemma. Honestly I'm not too sure how to approach it and have yet to find a solution. I'm still considering trying different ways in the near future though and will update if I do. Please update as well if you do end up finding a way!"
        },
        {
          "id": 318561,
          "postDate": "2018-04-24T03:54:04.917Z",
          "content": "<p>I am going to these 2 approaches when I have time:</p>\n\n<ol>\n<li>The one I mentioned above: Append test set to training, then shift everything down by the number of test set. I think this method is pretty stupid and will not help anyway, but at least it is a start xD.</li>\n<li>Get the cumulative mean of target variable in training --&gt; append test to training --&gt; fill in the target encoding feature in test data by getting it from 'day - 1' (thinking of doing this by <code>pandas.groupby</code>). I think this method will have many <code>NaN</code> values because there are many feature combinations which will not match the ones in training.</li>\n</ol>\n\n<p>However, take this with a grain of salt. I am by no means an expert and this is my first featured competition on Kaggle xD.  I'm really eager to get my first medal though.</p>\n\n<p>Maybe gurus from silver rank or above can help us? </p>",
          "rawMarkdown": "I am going to these 2 approaches when I have time:\n\n 1. The one I mentioned above: Append test set to training, then shift everything down by the number of test set. I think this method is pretty stupid and will not help anyway, but at least it is a start xD.\n 2. Get the cumulative mean of target variable in training --&gt; append test to training --&gt; fill in the target encoding feature in test data by getting it from 'day - 1' (thinking of doing this by `pandas.groupby`). I think this method will have many `NaN` values because there are many feature combinations which will not match the ones in training.\n\nHowever, take this with a grain of salt. I am by no means an expert and this is my first featured competition on Kaggle xD.  I'm really eager to get my first medal though.\n\nMaybe gurus from silver rank or above can help us? "
        },
        {
          "id": 318581,
          "postDate": "2018-04-24T05:17:18.917Z",
          "content": "<p>There's essentially 2 ways you can do this.</p>\n\n<ol>\n<li>Calculate metrics for train. For example, maybe you get the mean is_attributed encoding for each app in the train set. You then do simple left merge of those app-xmean values from train against apps in test. You might have some apps in test set that do not appear in train; no big deal, those will have nan/missing xmean values.</li>\n<li>Alternatively, since you're already doing pre-processing on the entire dataset, you can use your best lvl-1 model to predict the is_attributed target of your test set. Armed with is_attributed values for both train+test, you can now calculate xmean (target encoding) for the entire dataset.</li>\n</ol>\n\n<p>The second method can be preferred when |train| &gt;&gt; |test|, especially since has the added 2birds 1stone benefit of adding some needed noise which can assist with fighting overfitting.</p>\n\n<p>Other techniques that can be used to fill the nans from method #1 above would be like calculating the target encoded values that appear in train but not in test, and test.fillna() with the mean of them all (the mean of the mean encoded values).</p>\n\n<p>There's also the cumulative, rolling mean method . . .</p>\n\n<p>I don't think there's any right or best answer here, you really just have to see what jives best with your models.</p>",
          "rawMarkdown": "There's essentially 2 ways you can do this.\n\n1. Calculate metrics for train. For example, maybe you get the mean is_attributed encoding for each app in the train set. You then do simple left merge of those app-xmean values from train against apps in test. You might have some apps in test set that do not appear in train; no big deal, those will have nan/missing xmean values.\n2. Alternatively, since you're already doing pre-processing on the entire dataset, you can use your best lvl-1 model to predict the is_attributed target of your test set. Armed with is_attributed values for both train+test, you can now calculate xmean (target encoding) for the entire dataset.\n\nThe second method can be preferred when |train| &gt;&gt; |test|, especially since has the added 2birds 1stone benefit of adding some needed noise which can assist with fighting overfitting.\n\nOther techniques that can be used to fill the nans from method #1 above would be like calculating the target encoded values that appear in train but not in test, and test.fillna() with the mean of them all (the mean of the mean encoded values).\n\nThere's also the cumulative, rolling mean method . . .\n\nI don't think there's any right or best answer here, you really just have to see what jives best with your models.",
          "votes": 2
        },
        {
          "id": 318592,
          "postDate": "2018-04-24T05:27:33.007Z",
          "content": "<p>Number 2 is brilliant! Thanks, I will certainly try it out.</p>\n\n<p>Do you recommend using different model for both scenarios? Oh well I can try all possibilities..</p>",
          "rawMarkdown": "Number 2 is brilliant! Thanks, I will certainly try it out.\n\nDo you recommend using different model for both scenarios? Oh well I can try all possibilities.."
        },
        {
          "id": 318923,
          "postDate": "2018-04-24T18:57:04.613Z",
          "content": "<p>Hi, Budi. I a newbee here. Could you illustrate the meaning of the cumulative mean of target variable in training.</p>",
          "rawMarkdown": "Hi, Budi. I a newbee here. Could you illustrate the meaning of the cumulative mean of target variable in training."
        },
        {
          "id": 319007,
          "postDate": "2018-04-25T02:54:40.630Z",
          "content": "<p>You can calculate the mean for that category so far considering the timestamp. So it does not leak future information.</p>",
          "rawMarkdown": "You can calculate the mean for that category so far considering the timestamp. So it does not leak future information."
        },
        {
          "id": 319022,
          "postDate": "2018-04-25T04:09:52.347Z",
          "content": "<p>Given some columns that you have decided to groupby, you aggregate the target variable by using cumulative sum divided by its cumulative count.</p>\n\n<p>I learnt this method in a Coursera course, proven to be relatively leak free and there is no need for any hyperparameter tuning.</p>",
          "rawMarkdown": "Given some columns that you have decided to groupby, you aggregate the target variable by using cumulative sum divided by its cumulative count.\n\nI learnt this method in a Coursera course, proven to be relatively leak free and there is no need for any hyperparameter tuning.",
          "votes": 1
        },
        {
          "id": 319026,
          "postDate": "2018-04-25T04:17:14.977Z",
          "content": "<p>@Budi, that sounds like a pretty good idea. I have yet to try it but it may work. You mentioned previously that the groups in the test set would then obtain the previous day's values right? So they would essentially be taking the overall frequency then</p>",
          "rawMarkdown": "@Budi, that sounds like a pretty good idea. I have yet to try it but it may work. You mentioned previously that the groups in the test set would then obtain the previous day's values right? So they would essentially be taking the overall frequency then"
        },
        {
          "id": 319033,
          "postDate": "2018-04-25T04:39:02.653Z",
          "content": "<p>@Budi Ryan, could you tell me the name of the coursera class you mention? </p>",
          "rawMarkdown": "@Budi Ryan, could you tell me the name of the coursera class you mention? "
        },
        {
          "id": 319034,
          "postDate": "2018-04-25T04:40:34.580Z",
          "content": "<p>It's most likely this one: <a href=\"https://www.coursera.org/learn/competitive-data-science\">https://www.coursera.org/learn/competitive-data-science</a> It's a fun course. I wouldn't recommend taking it mid-competition though.</p>",
          "rawMarkdown": "It's most likely this one: https://www.coursera.org/learn/competitive-data-science It's a fun course. I wouldn't recommend taking it mid-competition though."
        },
        {
          "id": 319039,
          "postDate": "2018-04-25T04:51:25.247Z",
          "content": "<p><a href=\"/authman\">@authman</a> it is a dope course!</p>",
          "rawMarkdown": "@authman it is a dope course!",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 317893,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-22T19:23:07.973000",
      "content": "<p>Yes, this is called 'leaking target information'.   One way to avoid this is to compute target average on some part of the data, and use that to create feature on another part of the data.  </p>",
      "votes": 10,
      "replies": [
        {
          "id": 317894,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-22T19:39:32.490000",
          "content": "<p>Thank you very much for your reply! What you mentioned sounds interesting though. To clarify my understanding, could this be an example of that?</p>\n\n<p>Calculate frequencies of is_attributed for each ip - let's call it \"ip_freq\"\nUse \"ip_freq\" to create some other feature x. When training, only use feature x and not feature ip_freq.</p>\n\n<p>Is that what you mean by \"create feature on another part of the data\"? I can't seem to think of how you could create the feature x, according to my example, and not leak anything about the target either</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317925,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-22T21:11:13.297000",
          "content": "<p>Another technique would be to make your target encoding a bit more noisy. Two example methods to accomplish this:</p>\n\n<ol>\n<li>rather than using ALL your training data, use a subset of the training data to do target encoding. Your different folds can use different bags if you want to capture all the signal.</li>\n<li>or alternatively, add some gauss noise to it. figuring out the optimal level of noise should be a function of your validation scheme</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318195,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-23T11:24:49.223000",
          "content": "<p>@Edward, not, this is not what I mean.  By 'part of the data' I mean some split of the data by rows, not by columns.  If you compute target frequency for a row using the rwo (and other rows), then you leak the target information of the row into your new feature.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 318400,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-23T18:52:07.177000",
          "content": "<p><a href=\"/authman\">@authman</a> Thank you for your ideas! If I'm understanding correctly, your 1st point is very similar to what CPMP is mentioning right?</p>\n\n<p>@CPMP Thanks for the clarification! To clarify my understanding, so, for example, I could compute the attributed frequency for some feature 'x' only for day 'y' in the training data but then merge those frequencies into the other days in the training set. So in that case, would it still be okay to use day 'y' as part of training? I can see that the overfitting would be a lot less since there would technically be more noise added right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318432,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-23T20:28:19.380000",
          "content": "<p>@Edward, you got it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318471,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T22:21:25.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318809,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-24T14:09:43.993000",
          "content": "<p>Here is a (so far under the radar but likely about to blow up) compact, runnable kernel of N.2 above: <a href=\"https://www.kaggle.com/scirpus/a-bit-more-target-encoding-with-gp\">https://www.kaggle.com/scirpus/a-bit-more-target-encoding-with-gp</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318914,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-04-24T18:22:46.623000",
          "content": "<p>Hi, CPMP. I am also try to use this attributed frequency feature. My trainset is about 100millionrows. How many rows of the subset I should choose is appropriate?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 318673,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-04-24T08:51:03.563000",
      "content": "<p>That was my experience too.\nWhat @CPMP suggested is what i did up to a few days ago, I did target encoding for each day separately, using the previous days: day 8 gets the ratios of day 7, day 9 gets ratios of day 8+ 7  and test data day 10 gets ratios of all previous days (all training). this decreased somewhat the amount of leakage and overfitting, but took a long time to preprocess and removed the ability to train on day 7, which is pretty costly.</p>\n\n<p>Play with it and see how it goes for you =)\nwill be happy to know how it goes.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 319393,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-26T00:41:31.437000",
          "content": "<p>DO THESE FEATURES WORK WELL IN YOUR MODEL ？</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 319666,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-04-26T15:02:19.830000",
          "content": "<p>hard question, i'll tell you this - They were in top in feature importance, but i ended up removing them for different reasons. took a long time to process and i suspect they were overfitting my model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319670,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-04-26T15:10:03.873000",
          "content": "<p>Same as you.Especially ip related features overfit model .</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 318530,
      "author_name": "Budi Ryan",
      "author_url": "",
      "post_date": "2018-04-24T02:35:45.553000",
      "content": "<p>Hey,</p>\n\n<p>I am also thinking of using target encoding as one of the features. </p>\n\n<p>To solve your problems you can maybe try using cumulative count instead of summing throughout the entire training set. This has less amount of leakage.</p>\n\n<p>However, one problem I am trying to tackle at the moment is: How do you create that feature for the test set? Do you just shift the feature down by number of test set and discard the start of training set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 318556,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-24T03:38:04.453000",
          "content": "<p>Hey Budi, I've thought of that exact same dilemma. Honestly I'm not too sure how to approach it and have yet to find a solution. I'm still considering trying different ways in the near future though and will update if I do. Please update as well if you do end up finding a way!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318561,
          "author_name": "Budi Ryan",
          "author_url": "",
          "post_date": "2018-04-24T03:54:04.917000",
          "content": "<p>I am going to these 2 approaches when I have time:</p>\n\n<ol>\n<li>The one I mentioned above: Append test set to training, then shift everything down by the number of test set. I think this method is pretty stupid and will not help anyway, but at least it is a start xD.</li>\n<li>Get the cumulative mean of target variable in training --&gt; append test to training --&gt; fill in the target encoding feature in test data by getting it from 'day - 1' (thinking of doing this by <code>pandas.groupby</code>). I think this method will have many <code>NaN</code> values because there are many feature combinations which will not match the ones in training.</li>\n</ol>\n\n<p>However, take this with a grain of salt. I am by no means an expert and this is my first featured competition on Kaggle xD.  I'm really eager to get my first medal though.</p>\n\n<p>Maybe gurus from silver rank or above can help us? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318581,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-24T05:17:18.917000",
          "content": "<p>There's essentially 2 ways you can do this.</p>\n\n<ol>\n<li>Calculate metrics for train. For example, maybe you get the mean is_attributed encoding for each app in the train set. You then do simple left merge of those app-xmean values from train against apps in test. You might have some apps in test set that do not appear in train; no big deal, those will have nan/missing xmean values.</li>\n<li>Alternatively, since you're already doing pre-processing on the entire dataset, you can use your best lvl-1 model to predict the is_attributed target of your test set. Armed with is_attributed values for both train+test, you can now calculate xmean (target encoding) for the entire dataset.</li>\n</ol>\n\n<p>The second method can be preferred when |train| &gt;&gt; |test|, especially since has the added 2birds 1stone benefit of adding some needed noise which can assist with fighting overfitting.</p>\n\n<p>Other techniques that can be used to fill the nans from method #1 above would be like calculating the target encoded values that appear in train but not in test, and test.fillna() with the mean of them all (the mean of the mean encoded values).</p>\n\n<p>There's also the cumulative, rolling mean method . . .</p>\n\n<p>I don't think there's any right or best answer here, you really just have to see what jives best with your models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 318592,
          "author_name": "Budi Ryan",
          "author_url": "",
          "post_date": "2018-04-24T05:27:33.007000",
          "content": "<p>Number 2 is brilliant! Thanks, I will certainly try it out.</p>\n\n<p>Do you recommend using different model for both scenarios? Oh well I can try all possibilities..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318923,
          "author_name": "Siyuan Dang",
          "author_url": "",
          "post_date": "2018-04-24T18:57:04.613000",
          "content": "<p>Hi, Budi. I a newbee here. Could you illustrate the meaning of the cumulative mean of target variable in training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319007,
          "author_name": "Luis Moneda",
          "author_url": "",
          "post_date": "2018-04-25T02:54:40.630000",
          "content": "<p>You can calculate the mean for that category so far considering the timestamp. So it does not leak future information.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319022,
          "author_name": "Budi Ryan",
          "author_url": "",
          "post_date": "2018-04-25T04:09:52.347000",
          "content": "<p>Given some columns that you have decided to groupby, you aggregate the target variable by using cumulative sum divided by its cumulative count.</p>\n\n<p>I learnt this method in a Coursera course, proven to be relatively leak free and there is no need for any hyperparameter tuning.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 319026,
          "author_name": "Edward Chen",
          "author_url": "",
          "post_date": "2018-04-25T04:17:14.977000",
          "content": "<p>@Budi, that sounds like a pretty good idea. I have yet to try it but it may work. You mentioned previously that the groups in the test set would then obtain the previous day's values right? So they would essentially be taking the overall frequency then</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319033,
          "author_name": "MengYe",
          "author_url": "",
          "post_date": "2018-04-25T04:39:02.653000",
          "content": "<p>@Budi Ryan, could you tell me the name of the coursera class you mention? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319034,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-25T04:40:34.580000",
          "content": "<p>It's most likely this one: <a href=\"https://www.coursera.org/learn/competitive-data-science\">https://www.coursera.org/learn/competitive-data-science</a> It's a fun course. I wouldn't recommend taking it mid-competition though.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 319039,
          "author_name": "Budi Ryan",
          "author_url": "",
          "post_date": "2018-04-25T04:51:25.247000",
          "content": "<p><a href=\"/authman\">@authman</a> it is a dope course!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "317889": "I've been testing a few features and noticed that features that calculate the frequency of is_attributed all seem to overfit the training data. My thinking behind this is that, since most of the apps are not download, the model is trained to simply predict 0 for any frequencies of 0 and 1 for any frequencies &gt; 0, and since there are so many 0s the model overfits. \n\nIs my reasoning behind that correct or am I misunderstanding something? Also, would any feature that uses is_attributed as a frequency be considered target encoding? Is it possible for target encoding to not overfit?\n\nThanks!",
    "317893": "Yes, this is called 'leaking target information'.   One way to avoid this is to compute target average on some part of the data, and use that to create feature on another part of the data.  ",
    "318673": "That was my experience too.\nWhat @CPMP suggested is what i did up to a few days ago, I did target encoding for each day separately, using the previous days: day 8 gets the ratios of day 7, day 9 gets ratios of day 8+ 7  and test data day 10 gets ratios of all previous days (all training). this decreased somewhat the amount of leakage and overfitting, but took a long time to preprocess and removed the ability to train on day 7, which is pretty costly.\n\nPlay with it and see how it goes for you =)\nwill be happy to know how it goes.",
    "318530": "Hey,\n\nI am also thinking of using target encoding as one of the features. \n\nTo solve your problems you can maybe try using cumulative count instead of summing throughout the entire training set. This has less amount of leakage.\n\nHowever, one problem I am trying to tackle at the moment is: How do you create that feature for the test set? Do you just shift the feature down by number of test set and discard the start of training set?\n\n"
  }
}