{
  "id": 51492,
  "title": "Validation",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51492",
  "author_name": "",
  "post_date": "2018-03-09T15:06:43.771388700Z",
  "votes": 28,
  "comment_count": 42,
  "views": 0,
  "content": "<p>When I first looked at this competition, I thought, \"Oh, good. This is simple. We can just take a random sample for validation, and maybe use a k-fold procedure, just like in the good old days.\"  But after reading discussion topics and thinking about it, I realize I'm completely wrong about that.  That validation procedure has a big leak.  One of the most important predictors is going to be frequency: if one entity (however identified) does an unusually high number of clicks, that's an indication that something weird is happening (and, conditional perhaps on various other factors, probably means they are unlikely to download).  But frequency can change over time.  To the extent that we use frequency as a predictor, it needs to be an indication of <em>future</em> frequency.  If you choose validation cases at random, so that, timewise, they are intermingled with your training cases, you are getting an unfair glimpse into the future.</p>\n\n<p>Fortunately, one good thing about having a huge amount of data is that you can get a timewise holdout that's large enough to use for validation.  So I guess that's what I'll try.</p>",
  "messages": [
    {
      "id": "293275",
      "postDate": "03/09/2018 15:06:43",
      "content": "<p>When I first looked at this competition, I thought, \"Oh, good. This is simple. We can just take a random sample for validation, and maybe use a k-fold procedure, just like in the good old days.\"  But after reading discussion topics and thinking about it, I realize I'm completely wrong about that.  That validation procedure has a big leak.  One of the most important predictors is going to be frequency: if one entity (however identified) does an unusually high number of clicks, that's an indication that something weird is happening (and, conditional perhaps on various other factors, probably means they are unlikely to download).  But frequency can change over time.  To the extent that we use frequency as a predictor, it needs to be an indication of <em>future</em> frequency.  If you choose validation cases at random, so that, timewise, they are intermingled with your training cases, you are getting an unfair glimpse into the future.</p>\n\n<p>Fortunately, one good thing about having a huge amount of data is that you can get a timewise holdout that's large enough to use for validation.  So I guess that's what I'll try.</p>",
      "rawMarkdown": "When I first looked at this competition, I thought, \"Oh, good. This is simple. We can just take a random sample for validation, and maybe use a k-fold procedure, just like in the good old days.\"  But after reading discussion topics and thinking about it, I realize I'm completely wrong about that.  That validation procedure has a big leak.  One of the most important predictors is going to be frequency: if one entity (however identified) does an unusually high number of clicks, that's an indication that something weird is happening (and, conditional perhaps on various other factors, probably means they are unlikely to download).  But frequency can change over time.  To the extent that we use frequency as a predictor, it needs to be an indication of *future* frequency.  If you choose validation cases at random, so that, timewise, they are intermingled with your training cases, you are getting an unfair glimpse into the future.\n\nFortunately, one good thing about having a huge amount of data is that you can get a timewise holdout that's large enough to use for validation.  So I guess that's what I'll try.",
      "votes": null
    },
    {
      "id": "293314",
      "postDate": "03/09/2018 16:10:43",
      "content": "<p>This is a very good observation Andy, I have trying to use random sample and this not work very well.\nMaybe using the time for validation as you mentioned will be the best!</p>",
      "rawMarkdown": "This is a very good observation Andy, I have trying to use random sample and this not work very well.\nMaybe using the time for validation as you mentioned will be the best!",
      "votes": null
    },
    {
      "id": "293587",
      "postDate": "03/10/2018 07:27:58",
      "content": "<p>Exactly! It's better to use, say, 11.6-11.8 as training, and 11.9 as validation.</p>",
      "rawMarkdown": "Exactly! It's better to use, say, 11.6-11.8 as training, and 11.9 as validation.",
      "votes": null
    },
    {
      "id": "293837",
      "postDate": "03/10/2018 18:25:33",
      "content": "<p>I think setting <code>shuffle=False</code> in <code>sklearn.model_selection.train_test_split</code> will work, since the data appear to be in chronological order.</p>",
      "rawMarkdown": "I think setting `shuffle=False` in `sklearn.model_selection.train_test_split` will work, since the data appear to be in chronological order.",
      "votes": null
    },
    {
      "id": "293919",
      "postDate": "03/10/2018 21:19:03",
      "content": "<p>...or maybe not.  One problem with doing a simple split is that the time gap between train and validation may not match that between train and test.  And the times of day for validation typically won't match the test set.  So I suppose I will try taking time into account explicitly.</p>",
      "rawMarkdown": "...or maybe not.  One problem with doing a simple split is that the time gap between train and validation may not match that between train and test.  And the times of day for validation typically won't match the test set.  So I suppose I will try taking time into account explicitly.",
      "votes": null
    },
    {
      "id": "294081",
      "postDate": "03/11/2018 08:26:28",
      "content": "<p>Indeed, I believe the best would be to take the latest say 20% of the data as a validation set. Seems fine to me to still feature engineer on the whole dataset (so deriving the frequency for the whole set). Does somebody knows if that would be reasonable?</p>",
      "rawMarkdown": "Indeed, I believe the best would be to take the latest say 20% of the data as a validation set. Seems fine to me to still feature engineer on the whole dataset (so deriving the frequency for the whole set). Does somebody knows if that would be reasonable?",
      "votes": null
    },
    {
      "id": "295068",
      "postDate": "03/13/2018 02:50:57",
      "content": "<p>I'm trying a moving validation set: I train 3 times, first with 70% for train and 30% for validation( without shuffling), then 80% train and 20% validation, and 90% train 10% validation. I think this scheme is better than doing only one split for validation.</p>",
      "rawMarkdown": "I'm trying a moving validation set: I train 3 times, first with 70% for train and 30% for validation( without shuffling), then 80% train and 20% validation, and 90% train 10% validation. I think this scheme is better than doing only one split for validation.",
      "votes": null
    },
    {
      "id": "295084",
      "postDate": "03/13/2018 03:26:08",
      "content": "<p>Something like frequency on the whole training set will leak future frequency into the past beyond the time window available at prediction time for the test data. So for example, a training label on day 1 would be predicted with frequency characteristics partially drawn from day 3. For the test data, you can only look backward from later in day 4, not multiple days later. </p>\n\n<p>I'd guess it might still work ok to do total frequency, but a safer bet could be to use trailing frequency + forward frequency over only the remainder of that day to mimic what the test data gets to see. You'd basically end up computing 1 day, 2 day, and 3 day frequencies and using those features for clicks on each of the respective days.  </p>",
      "rawMarkdown": "Something like frequency on the whole training set will leak future frequency into the past beyond the time window available at prediction time for the test data. So for example, a training label on day 1 would be predicted with frequency characteristics partially drawn from day 3. For the test data, you can only look backward from later in day 4, not multiple days later. \n\nI'd guess it might still work ok to do total frequency, but a safer bet could be to use trailing frequency + forward frequency over only the remainder of that day to mimic what the test data gets to see. You'd basically end up computing 1 day, 2 day, and 3 day frequencies and using those features for clicks on each of the respective days.",
      "votes": null
    },
    {
      "id": "295416",
      "postDate": "03/13/2018 16:13:20",
      "content": "<p>Konrad Banachewicz implements a time-aware validation set <a href=\"https://www.kaggle.com/konradb/validation-set\">here</a>.</p>",
      "rawMarkdown": "Konrad Banachewicz implements a time-aware validation set [here][1].\n\n [1]: https://www.kaggle.com/konradb/validation-set",
      "votes": null
    },
    {
      "id": "295419",
      "postDate": "03/13/2018 16:15:52",
      "content": "<p>Thanks for the shout-out.</p>",
      "rawMarkdown": "Thanks for the shout-out.",
      "votes": null
    },
    {
      "id": "296733",
      "postDate": "03/15/2018 17:24:08",
      "content": "<p>Given that we can also get frequencies from test data, wouldn't getting counts on whole train set (prior to splitting to train/val) be the same as getting frequencies from train and test for the final run?  Would it make sense?</p>",
      "rawMarkdown": "Given that we can also get frequencies from test data, wouldn't getting counts on whole train set (prior to splitting to train/val) be the same as getting frequencies from train and test for the final run?  Would it make sense?",
      "votes": null
    },
    {
      "id": "296760",
      "postDate": "03/15/2018 18:12:41",
      "content": "<p>The problem lies in what the model would learn if you pull in counts from the future. If you count across all 4 days and use that count information to make predictions for day 1 clicks, your model gets a look into the future that you don't have when you predict the test set. The total counts across train/test would be the same but the <strong>meaning</strong> of those counts would be different. The misalignment in meaning might not be a big problem, but it also might cause your model to generalize to the test data less well.</p>",
      "rawMarkdown": "The problem lies in what the model would learn if you pull in counts from the future. If you count across all 4 days and use that count information to make predictions for day 1 clicks, your model gets a look into the future that you don't have when you predict the test set. The total counts across train/test would be the same but the **meaning** of those counts would be different. The misalignment in meaning might not be a big problem, but it also might cause your model to generalize to the test data less well.",
      "votes": null
    },
    {
      "id": "296844",
      "postDate": "03/15/2018 20:50:23",
      "content": "<p>Essentially, using counts more than a day in the future is a problem in the training phase rather than the validation phase.  If you do what @yulia suggests, you would get an (almost) legitimate validation score, but you might not be training the optimal model to predict the test data.  OTOH the extra count data might improve the model to the extent that it's worth including despite being unavailable at test time.  In principle, you can test this by getting separate validation scores for models that do and don't see counts from future days.  The test isn't perfect, because, when training for the validation score, you have one fewer day to train on than when the training for submission, so the early data are seeing further into the future.</p>",
      "rawMarkdown": "Essentially, using counts more than a day in the future is a problem in the training phase rather than the validation phase.  If you do what @yulia suggests, you would get an (almost) legitimate validation score, but you might not be training the optimal model to predict the test data.  OTOH the extra count data might improve the model to the extent that it's worth including despite being unavailable at test time.  In principle, you can test this by getting separate validation scores for models that do and don't see counts from future days.  The test isn't perfect, because, when training for the validation score, you have one fewer day to train on than when the training for submission, so the early data are seeing further into the future.",
      "votes": null
    },
    {
      "id": "296851",
      "postDate": "03/15/2018 21:01:40",
      "content": "<p>^+1, agree there's a tradeoff where any leakiness may be offset by signal gains from more count data. From the kernels it seems clear that the full count data approach works very well, but it's not clear if leakless counts would work better and would be difficult to truly validate for the reason you mentioned.</p>\n\n<p>At the end of the day leaking only a couple days back in time may just not be a big deal.</p>",
      "rawMarkdown": "^+1, agree there's a tradeoff where any leakiness may be offset by signal gains from more count data. From the kernels it seems clear that the full count data approach works very well, but it's not clear if leakless counts would work better and would be difficult to truly validate for the reason you mentioned.\n\nAt the end of the day leaking only a couple days back in time may just not be a big deal.",
      "votes": null
    },
    {
      "id": "296927",
      "postDate": "03/16/2018 01:32:48",
      "content": "<p>@Joe Eddy, I like the way you explain things! Very clear to me, thanks!</p>",
      "rawMarkdown": "Joe Eddy, I like the way you explain things! Very clear to me, thanks!",
      "votes": null
    },
    {
      "id": "297013",
      "postDate": "03/16/2018 05:02:35",
      "content": "<p>Thanks! Glad if it's helpful :)</p>",
      "rawMarkdown": "Thanks! Glad if it's helpful :)",
      "votes": null
    },
    {
      "id": "302835",
      "postDate": "03/24/2018 20:04:24",
      "content": "<p>Are you able to build CV - LB constant in your validation setting? So far I have been trying different validation splits as discussed <a href=\"https://www.kaggle.com/konradb/validation-set\">here</a> by @Konrad and <a href=\"https://www.kaggle.com/aharless/training-and-validation-data-pickle/code\">here</a> by @Andy, and my CV is consistent with my LB but I am not able to build a CV - LB constant and that's why I think that I might be overfitting to public LB.  Validation is tough here.</p>",
      "rawMarkdown": "Are you able to build CV - LB constant in your validation setting? So far I have been trying different validation splits as discussed [here][1] by @Konrad and [here][2] by @Andy, and my CV is consistent with my LB but I am not able to build a CV - LB constant and that's why I think that I might be overfitting to public LB.  Validation is tough here.\n\n\n  [1]: https://www.kaggle.com/konradb/validation-set\n  [2]: https://www.kaggle.com/aharless/training-and-validation-data-pickle/code",
      "votes": null
    },
    {
      "id": "302852",
      "postDate": "03/24/2018 21:07:07",
      "content": "<p>Perhaps we should exclude from validation data the range of IP codes that is not represented in the test data and seems to have different characteristics, as found, e.g., by CPMP in <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">this kernel</a>.</p>",
      "rawMarkdown": "Perhaps we should exclude from validation data the range of IP codes that is not represented in the test data and seems to have different characteristics, as found, e.g., by CPMP in [this kernel][1].\n\n [1]: https://www.kaggle.com/cpmpml/ip-download-rates",
      "votes": null
    },
    {
      "id": "302865",
      "postDate": "03/24/2018 21:34:45",
      "content": "<p>As discussed <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">here</a> by @yulia, </p>\n\n<pre><code>The pattern of IP assignment in Train does not mimic the one in test.\n</code></pre>\n\n<p>So we are not sure that IP's of same range have same characteristics in train and test and even if we remove the IP no.s from validation split which are not present in test set, how can we be sure that our all val set IP's are present in test set as they do not follow the mapping rule of train data.</p>",
      "rawMarkdown": "As discussed [here][1] by @yulia, \n\n    The pattern of IP assignment in Train does not mimic the one in test.\n\nSo we are not sure that IP's of same range have same characteristics in train and test and even if we remove the IP no.s from validation split which are not present in test set, how can we be sure that our all val set IP's are present in test set as they do not follow the mapping rule of train data.\n\n\n  [1]: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal",
      "votes": null
    },
    {
      "id": "302941",
      "postDate": "03/25/2018 03:04:04",
      "content": "<p>The issue isn't about particular IP addresses but about ranges of codes.  The mapping between codes and IP addresses should be one-to-one.  In general we don't know much about how the codes were assigned, but we've at least identified a range of codes that is present in the training data and not the test data.  And we also know that range is associated with more downloads.  It's not clear what this means, but it seems like these might be unusual in some way and perhaps should be excluded from validation data.  Certainly if one were doing adversarial validation, the algorithm would identify those records as unlikely to be from the test data and therefore not use them for validation.</p>",
      "rawMarkdown": "The issue isn't about particular IP addresses but about ranges of codes.  The mapping between codes and IP addresses should be one-to-one.  In general we don't know much about how the codes were assigned, but we've at least identified a range of codes that is present in the training data and not the test data.  And we also know that range is associated with more downloads.  It's not clear what this means, but it seems like these might be unusual in some way and perhaps should be excluded from validation data.  Certainly if one were doing adversarial validation, the algorithm would identify those records as unlikely to be from the test data and therefore not use them for validation.",
      "votes": null
    },
    {
      "id": "303536",
      "postDate": "03/26/2018 11:15:26",
      "content": "<p>So here's a small update: the validation setup I'm running at the moment is along the lines of my kernel </p>\n\n<p><a href=\"https://www.kaggle.com/konradb/validation-set/code\">https://www.kaggle.com/konradb/validation-set/code</a></p>\n\n<p>It turns out that once you limit your evaluation set to the first hour (November 9th between 4 and 5 am) there is a match with LB at 4 significant digits. Obvious, I know - for some reason did not occur to me before, so I thought I'd share it here to save others some time :-)</p>",
      "rawMarkdown": "So here's a small update: the validation setup I'm running at the moment is along the lines of my kernel \n\nhttps://www.kaggle.com/konradb/validation-set/code\n\nIt turns out that once you limit your evaluation set to the first hour (November 9th between 4 and 5 am) there is a match with LB at 4 significant digits. Obvious, I know - for some reason did not occur to me before, so I thought I'd share it here to save others some time :-)",
      "votes": null
    },
    {
      "id": "303668",
      "postDate": "03/26/2018 14:23:26",
      "content": "<p>Maybe because public LB is based on the same hour in test, see: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060</a></p>",
      "rawMarkdown": "Maybe because public LB is based on the same hour in test, see: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060",
      "votes": null
    },
    {
      "id": "303683",
      "postDate": "03/26/2018 14:39:18",
      "content": "<p>I haven't been getting consistent results from the first hour, but maybe I need to tweak some things.</p>",
      "rawMarkdown": "I haven't been getting consistent results from the first hour, but maybe I need to tweak some things.",
      "votes": null
    },
    {
      "id": "303833",
      "postDate": "03/26/2018 17:44:18",
      "content": "<p>@CPMP - that's what i meant by 'obvious' :-)</p>",
      "rawMarkdown": "CPMP - that's what i meant by 'obvious' :-)",
      "votes": null
    },
    {
      "id": "303859",
      "postDate": "03/26/2018 18:21:14",
      "content": "<p>Konrad, it was not obvious to me so far ;)</p>",
      "rawMarkdown": "Konrad, it was not obvious to me so far ;)",
      "votes": null
    },
    {
      "id": "303872",
      "postDate": "03/26/2018 18:29:37",
      "content": "<p>For me neither - that's why my sarcastic autopilot kicked in :-)</p>",
      "rawMarkdown": "For me neither - that's why my sarcastic autopilot kicked in :-)",
      "votes": null
    },
    {
      "id": "304028",
      "postDate": "03/26/2018 22:35:18",
      "content": "<p>Thanks for your findings, I am not sure if I am getting this right but since there is a match for the hour in validation set and LB, validating on it can get you closer to CV - LB ~ 0, would that not mean we might tend to overfit the leaderboard though indirectly by optimizing on a validation which is similar to LB data and not the test data.</p>",
      "rawMarkdown": "Thanks for your findings, I am not sure if I am getting this right but since there is a match for the hour in validation set and LB, validating on it can get you closer to CV - LB ~ 0, would that not mean we might tend to overfit the leaderboard though indirectly by optimizing on a validation which is similar to LB data and not the test data.",
      "votes": null
    },
    {
      "id": "304255",
      "postDate": "03/27/2018 09:04:28",
      "content": "<p>Has any one tried Kfold CV here?</p>",
      "rawMarkdown": "Has any one tried Kfold CV here?",
      "votes": null
    },
    {
      "id": "304260",
      "postDate": "03/27/2018 09:17:28",
      "content": "<p>I am too trying to get my validation sorted, kfold seems to be a go-to option but I would assume that the  time series nature of the data would make random data allocation to folds, not a great cv strategy, \nI am looking something along the lines of </p>\n\n<p>Splitting training data into time-series1,time-series2, time-series3, time-series4</p>\n\n<p>fold 1 : training [time-series1], test [time-series2]</p>\n\n<p>fold 2 : training [time-series1 + time-series2], test [time-series3] </p>\n\n<p>fold 3 : training [time-series1 + time-series2 + time-series3], test [time-series4] </p>",
      "rawMarkdown": "I am too trying to get my validation sorted, kfold seems to be a go-to option but I would assume that the  time series nature of the data would make random data allocation to folds, not a great cv strategy, \nI am looking something along the lines of \n\nSplitting training data into time-series1,time-series2, time-series3, time-series4\n\nfold 1 : training [time-series1], test [time-series2]\n\nfold 2 : training [time-series1 + time-series2], test [time-series3] \n\nfold 3 : training [time-series1 + time-series2 + time-series3], test [time-series4]",
      "votes": null
    },
    {
      "id": "304620",
      "postDate": "03/27/2018 19:55:32",
      "content": "<p>Yes, this is what it means. Nevertheless, you can use the corresponding hours from 9th as validation also for the other hours from the test set (which are \"reserved\" for the private, final evaluation). Basically, you would like your validation set to resemble as closely as possible the test data.</p>\n\n<p>Note the last sentance applies only for Kaggle competitions :)</p>",
      "rawMarkdown": "Yes, this is what it means. Nevertheless, you can use the corresponding hours from 9th as validation also for the other hours from the test set (which are \"reserved\" for the private, final evaluation). Basically, you would like your validation set to resemble as closely as possible the test data.\n\nNote the last sentance applies only for Kaggle competitions :)",
      "votes": null
    },
    {
      "id": "304623",
      "postDate": "03/27/2018 19:57:25",
      "content": "<p>It is an interesting idea worth to be verified! My personal intuition tells me that those IPs carry also other valuable information like conversion rate per os, device, many possible composed features, etc.</p>",
      "rawMarkdown": "It is an interesting idea worth to be verified! My personal intuition tells me that those IPs carry also other valuable information like conversion rate per os, device, many possible composed features, etc.",
      "votes": null
    },
    {
      "id": "304624",
      "postDate": "03/27/2018 19:59:02",
      "content": "<p>Exactly my thought. Well, almost, I would use only those hours from 11.9 that correspond to the hours from the test data for 11.10. Note, I still wouldn't use the other hours from 11.9 as training.</p>",
      "rawMarkdown": "Exactly my thought. Well, almost, I would use only those hours from 11.9 that correspond to the hours from the test data for 11.10. Note, I still wouldn't use the other hours from 11.9 as training.",
      "votes": null
    },
    {
      "id": "304631",
      "postDate": "03/27/2018 20:09:54",
      "content": "<p>In the light of <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">Cute Chibiko's kernel</a>, I don't think we should exclude those IPs from validation.  It seems the sponsors started with the test data when assigning IP codes, so the higher codes are just for IPs that didn't happen to appear in the test data, not ones that are qualitatively different from the those in the test data.  </p>\n\n<p>The important thing is not to use the numeric values of the IP codes in training (so don't feed them directly into a GBT algorithm that will try to split its trees according to those numeric values), or it will overfit (both in training and in validation).  As long as you treat IP as a pure categorical feature and are careful to make your model ignore the implicit ordering of the codes, I expect IP has useful information.</p>",
      "rawMarkdown": "In the light of [Cute Chibiko's kernel][1], I don't think we should exclude those IPs from validation.  It seems the sponsors started with the test data when assigning IP codes, so the higher codes are just for IPs that didn't happen to appear in the test data, not ones that are qualitatively different from the those in the test data.  \n\nThe important thing is not to use the numeric values of the IP codes in training (so don't feed them directly into a GBT algorithm that will try to split its trees according to those numeric values), or it will overfit (both in training and in validation).  As long as you treat IP as a pure categorical feature and are careful to make your model ignore the implicit ordering of the codes, I expect IP has useful information.\n\n [1]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
      "votes": null
    },
    {
      "id": "304701",
      "postDate": "03/27/2018 22:39:34",
      "content": "<p>Are you guys training with all the data? If not, how will you choose the subsample?</p>\n\n<p>Moreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?</p>",
      "rawMarkdown": "Are you guys training with all the data? If not, how will you choose the subsample?\n\nMoreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?",
      "votes": null
    },
    {
      "id": "304738",
      "postDate": "03/28/2018 00:40:08",
      "content": "<p>I'm only using 40m rows per run on my macbook, but its best to separate subsamples by time. The test set only uses 4 hours of data so you could separate those out into your validation set.</p>",
      "rawMarkdown": "I'm only using 40m rows per run on my macbook, but its best to separate subsamples by time. The test set only uses 4 hours of data so you could separate those out into your validation set.",
      "votes": null
    },
    {
      "id": "304740",
      "postDate": "03/28/2018 00:50:23",
      "content": "<p>Thank you very much for your answer. Imagining we would have 1 full day for test, how would you construct your validation set?</p>",
      "rawMarkdown": "Thank you very much for your answer. Imagining we would have 1 full day for test, how would you construct your validation set?",
      "votes": null
    },
    {
      "id": "304747",
      "postDate": "03/28/2018 01:02:48",
      "content": "<p>I'd probably use the last day as validation. </p>",
      "rawMarkdown": "I'd probably use the last day as validation.",
      "votes": null
    },
    {
      "id": "304757",
      "postDate": "03/28/2018 01:26:07",
      "content": "<p>One more, if it is not asking a lot :)</p>\n\n<p>How would I tune the hyperparameters of my model if I'd use the last day as validation? I'd need another subset to confirm that I am not overfitting that validation set, correct?</p>",
      "rawMarkdown": "One more, if it is not asking a lot :)\n\nHow would I tune the hyperparameters of my model if I'd use the last day as validation? I'd need another subset to confirm that I am not overfitting that validation set, correct?",
      "votes": null
    },
    {
      "id": "304759",
      "postDate": "03/28/2018 01:31:52",
      "content": "<p>In theory you should test it over many subsets. If you haven't done much FE then the super conservative parameters that are popular in the kernels should be a good place to start.</p>",
      "rawMarkdown": "In theory you should test it over many subsets. If you haven't done much FE then the super conservative parameters that are popular in the kernels should be a good place to start.",
      "votes": null
    },
    {
      "id": "304764",
      "postDate": "03/28/2018 01:45:52",
      "content": "<p>Thanks a lot for your advices!</p>",
      "rawMarkdown": "Thanks a lot for your advices!",
      "votes": null
    },
    {
      "id": "305586",
      "postDate": "03/29/2018 05:29:03",
      "content": "<p>what is trailing and forward frequency?</p>",
      "rawMarkdown": "what is trailing and forward frequency?",
      "votes": null
    },
    {
      "id": "306065",
      "postDate": "03/29/2018 20:15:24",
      "content": "<p>@shivraj I believe the idea is that the close match between day 9 hour 4 validation score and day 10 hour 4 score (LB) provides evidence that using the day 9 test hours for validation could be a good strategy. It's definitely not a good idea to only use hour 4 for validation, for the reason you describe (overfitting to public LB). However, since hour 4 shows very consistent validation/public LB results, we can hope that the other validation hours will also show behavior consistent with private LB results. </p>",
      "rawMarkdown": "shivraj I believe the idea is that the close match between day 9 hour 4 validation score and day 10 hour 4 score (LB) provides evidence that using the day 9 test hours for validation could be a good strategy. It's definitely not a good idea to only use hour 4 for validation, for the reason you describe (overfitting to public LB). However, since hour 4 shows very consistent validation/public LB results, we can hope that the other validation hours will also show behavior consistent with private LB results.",
      "votes": null
    },
    {
      "id": "306202",
      "postDate": "03/30/2018 02:57:14",
      "content": "<p>Yep! That's the right way.</p>",
      "rawMarkdown": "Yep! That's the right way.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 293314,
      "author_name": "joaopmpeinado",
      "author_url": "",
      "post_date": "03/09/2018 16:10:43",
      "content": "<p>This is a very good observation Andy, I have trying to use random sample and this not work very well.\nMaybe using the time for validation as you mentioned will be the best!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 293587,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "03/10/2018 07:27:58",
      "content": "<p>Exactly! It's better to use, say, 11.6-11.8 as training, and 11.9 as validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 304624,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/27/2018 19:59:02",
          "content": "<p>Exactly my thought. Well, almost, I would use only those hours from 11.9 that correspond to the hours from the test data for 11.10. Note, I still wouldn't use the other hours from 11.9 as training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306202,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "03/30/2018 02:57:14",
          "content": "<p>Yep! That's the right way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 293837,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/10/2018 18:25:33",
      "content": "<p>I think setting <code>shuffle=False</code> in <code>sklearn.model_selection.train_test_split</code> will work, since the data appear to be in chronological order.</p>",
      "votes": null,
      "replies": [
        {
          "id": 293919,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/10/2018 21:19:03",
          "content": "<p>...or maybe not.  One problem with doing a simple split is that the time gap between train and validation may not match that between train and test.  And the times of day for validation typically won't match the test set.  So I suppose I will try taking time into account explicitly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295416,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/13/2018 16:13:20",
          "content": "<p>Konrad Banachewicz implements a time-aware validation set <a href=\"https://www.kaggle.com/konradb/validation-set\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295419,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/13/2018 16:15:52",
          "content": "<p>Thanks for the shout-out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 294081,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/11/2018 08:26:28",
      "content": "<p>Indeed, I believe the best would be to take the latest say 20% of the data as a validation set. Seems fine to me to still feature engineer on the whole dataset (so deriving the frequency for the whole set). Does somebody knows if that would be reasonable?</p>",
      "votes": null,
      "replies": [
        {
          "id": 295084,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/13/2018 03:26:08",
          "content": "<p>Something like frequency on the whole training set will leak future frequency into the past beyond the time window available at prediction time for the test data. So for example, a training label on day 1 would be predicted with frequency characteristics partially drawn from day 3. For the test data, you can only look backward from later in day 4, not multiple days later. </p>\n\n<p>I'd guess it might still work ok to do total frequency, but a safer bet could be to use trailing frequency + forward frequency over only the remainder of that day to mimic what the test data gets to see. You'd basically end up computing 1 day, 2 day, and 3 day frequencies and using those features for clicks on each of the respective days.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296733,
          "author_name": "yuliagm",
          "author_url": "",
          "post_date": "03/15/2018 17:24:08",
          "content": "<p>Given that we can also get frequencies from test data, wouldn't getting counts on whole train set (prior to splitting to train/val) be the same as getting frequencies from train and test for the final run?  Would it make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296760,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/15/2018 18:12:41",
          "content": "<p>The problem lies in what the model would learn if you pull in counts from the future. If you count across all 4 days and use that count information to make predictions for day 1 clicks, your model gets a look into the future that you don't have when you predict the test set. The total counts across train/test would be the same but the <strong>meaning</strong> of those counts would be different. The misalignment in meaning might not be a big problem, but it also might cause your model to generalize to the test data less well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296844,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/15/2018 20:50:23",
          "content": "<p>Essentially, using counts more than a day in the future is a problem in the training phase rather than the validation phase.  If you do what @yulia suggests, you would get an (almost) legitimate validation score, but you might not be training the optimal model to predict the test data.  OTOH the extra count data might improve the model to the extent that it's worth including despite being unavailable at test time.  In principle, you can test this by getting separate validation scores for models that do and don't see counts from future days.  The test isn't perfect, because, when training for the validation score, you have one fewer day to train on than when the training for submission, so the early data are seeing further into the future.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296851,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/15/2018 21:01:40",
          "content": "<p>^+1, agree there's a tradeoff where any leakiness may be offset by signal gains from more count data. From the kernels it seems clear that the full count data approach works very well, but it's not clear if leakless counts would work better and would be difficult to truly validate for the reason you mentioned.</p>\n\n<p>At the end of the day leaking only a couple days back in time may just not be a big deal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296927,
          "author_name": "muhammadalfiansyah",
          "author_url": "",
          "post_date": "03/16/2018 01:32:48",
          "content": "<p>@Joe Eddy, I like the way you explain things! Very clear to me, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297013,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/16/2018 05:02:35",
          "content": "<p>Thanks! Glad if it's helpful :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 305586,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/29/2018 05:29:03",
          "content": "<p>what is trailing and forward frequency?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 295068,
      "author_name": "pedroschoen",
      "author_url": "",
      "post_date": "03/13/2018 02:50:57",
      "content": "<p>I'm trying a moving validation set: I train 3 times, first with 70% for train and 30% for validation( without shuffling), then 80% train and 20% validation, and 90% train 10% validation. I think this scheme is better than doing only one split for validation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302835,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/24/2018 20:04:24",
      "content": "<p>Are you able to build CV - LB constant in your validation setting? So far I have been trying different validation splits as discussed <a href=\"https://www.kaggle.com/konradb/validation-set\">here</a> by @Konrad and <a href=\"https://www.kaggle.com/aharless/training-and-validation-data-pickle/code\">here</a> by @Andy, and my CV is consistent with my LB but I am not able to build a CV - LB constant and that's why I think that I might be overfitting to public LB.  Validation is tough here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302852,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/24/2018 21:07:07",
      "content": "<p>Perhaps we should exclude from validation data the range of IP codes that is not represented in the test data and seems to have different characteristics, as found, e.g., by CPMP in <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">this kernel</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 302865,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/24/2018 21:34:45",
          "content": "<p>As discussed <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">here</a> by @yulia, </p>\n\n<pre><code>The pattern of IP assignment in Train does not mimic the one in test.\n</code></pre>\n\n<p>So we are not sure that IP's of same range have same characteristics in train and test and even if we remove the IP no.s from validation split which are not present in test set, how can we be sure that our all val set IP's are present in test set as they do not follow the mapping rule of train data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302941,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/25/2018 03:04:04",
          "content": "<p>The issue isn't about particular IP addresses but about ranges of codes.  The mapping between codes and IP addresses should be one-to-one.  In general we don't know much about how the codes were assigned, but we've at least identified a range of codes that is present in the training data and not the test data.  And we also know that range is associated with more downloads.  It's not clear what this means, but it seems like these might be unusual in some way and perhaps should be excluded from validation data.  Certainly if one were doing adversarial validation, the algorithm would identify those records as unlikely to be from the test data and therefore not use them for validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304623,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/27/2018 19:57:25",
          "content": "<p>It is an interesting idea worth to be verified! My personal intuition tells me that those IPs carry also other valuable information like conversion rate per os, device, many possible composed features, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304631,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/27/2018 20:09:54",
          "content": "<p>In the light of <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">Cute Chibiko's kernel</a>, I don't think we should exclude those IPs from validation.  It seems the sponsors started with the test data when assigning IP codes, so the higher codes are just for IPs that didn't happen to appear in the test data, not ones that are qualitatively different from the those in the test data.  </p>\n\n<p>The important thing is not to use the numeric values of the IP codes in training (so don't feed them directly into a GBT algorithm that will try to split its trees according to those numeric values), or it will overfit (both in training and in validation).  As long as you treat IP as a pure categorical feature and are careful to make your model ignore the implicit ordering of the codes, I expect IP has useful information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303536,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/26/2018 11:15:26",
      "content": "<p>So here's a small update: the validation setup I'm running at the moment is along the lines of my kernel </p>\n\n<p><a href=\"https://www.kaggle.com/konradb/validation-set/code\">https://www.kaggle.com/konradb/validation-set/code</a></p>\n\n<p>It turns out that once you limit your evaluation set to the first hour (November 9th between 4 and 5 am) there is a match with LB at 4 significant digits. Obvious, I know - for some reason did not occur to me before, so I thought I'd share it here to save others some time :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 303668,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 14:23:26",
          "content": "<p>Maybe because public LB is based on the same hour in test, see: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303683,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/26/2018 14:39:18",
          "content": "<p>I haven't been getting consistent results from the first hour, but maybe I need to tweak some things.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303833,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/26/2018 17:44:18",
          "content": "<p>@CPMP - that's what i meant by 'obvious' :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303859,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 18:21:14",
          "content": "<p>Konrad, it was not obvious to me so far ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303872,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/26/2018 18:29:37",
          "content": "<p>For me neither - that's why my sarcastic autopilot kicked in :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304028,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "03/26/2018 22:35:18",
          "content": "<p>Thanks for your findings, I am not sure if I am getting this right but since there is a match for the hour in validation set and LB, validating on it can get you closer to CV - LB ~ 0, would that not mean we might tend to overfit the leaderboard though indirectly by optimizing on a validation which is similar to LB data and not the test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304620,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/27/2018 19:55:32",
          "content": "<p>Yes, this is what it means. Nevertheless, you can use the corresponding hours from 9th as validation also for the other hours from the test set (which are \"reserved\" for the private, final evaluation). Basically, you would like your validation set to resemble as closely as possible the test data.</p>\n\n<p>Note the last sentance applies only for Kaggle competitions :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306065,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/29/2018 20:15:24",
          "content": "<p>@shivraj I believe the idea is that the close match between day 9 hour 4 validation score and day 10 hour 4 score (LB) provides evidence that using the day 9 test hours for validation could be a good strategy. It's definitely not a good idea to only use hour 4 for validation, for the reason you describe (overfitting to public LB). However, since hour 4 shows very consistent validation/public LB results, we can hope that the other validation hours will also show behavior consistent with private LB results. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 304255,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/27/2018 09:04:28",
      "content": "<p>Has any one tried Kfold CV here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 304260,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "03/27/2018 09:17:28",
          "content": "<p>I am too trying to get my validation sorted, kfold seems to be a go-to option but I would assume that the  time series nature of the data would make random data allocation to folds, not a great cv strategy, \nI am looking something along the lines of </p>\n\n<p>Splitting training data into time-series1,time-series2, time-series3, time-series4</p>\n\n<p>fold 1 : training [time-series1], test [time-series2]</p>\n\n<p>fold 2 : training [time-series1 + time-series2], test [time-series3] </p>\n\n<p>fold 3 : training [time-series1 + time-series2 + time-series3], test [time-series4] </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 304701,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "03/27/2018 22:39:34",
      "content": "<p>Are you guys training with all the data? If not, how will you choose the subsample?</p>\n\n<p>Moreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 304738,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "03/28/2018 00:40:08",
          "content": "<p>I'm only using 40m rows per run on my macbook, but its best to separate subsamples by time. The test set only uses 4 hours of data so you could separate those out into your validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304740,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "03/28/2018 00:50:23",
          "content": "<p>Thank you very much for your answer. Imagining we would have 1 full day for test, how would you construct your validation set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304747,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "03/28/2018 01:02:48",
          "content": "<p>I'd probably use the last day as validation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304757,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "03/28/2018 01:26:07",
          "content": "<p>One more, if it is not asking a lot :)</p>\n\n<p>How would I tune the hyperparameters of my model if I'd use the last day as validation? I'd need another subset to confirm that I am not overfitting that validation set, correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304759,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "03/28/2018 01:31:52",
          "content": "<p>In theory you should test it over many subsets. If you haven't done much FE then the super conservative parameters that are popular in the kernels should be a good place to start.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304764,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "03/28/2018 01:45:52",
          "content": "<p>Thanks a lot for your advices!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "293275": "When I first looked at this competition, I thought, \"Oh, good. This is simple. We can just take a random sample for validation, and maybe use a k-fold procedure, just like in the good old days.\"  But after reading discussion topics and thinking about it, I realize I'm completely wrong about that.  That validation procedure has a big leak.  One of the most important predictors is going to be frequency: if one entity (however identified) does an unusually high number of clicks, that's an indication that something weird is happening (and, conditional perhaps on various other factors, probably means they are unlikely to download).  But frequency can change over time.  To the extent that we use frequency as a predictor, it needs to be an indication of *future* frequency.  If you choose validation cases at random, so that, timewise, they are intermingled with your training cases, you are getting an unfair glimpse into the future.\n\nFortunately, one good thing about having a huge amount of data is that you can get a timewise holdout that's large enough to use for validation.  So I guess that's what I'll try.",
    "293314": "This is a very good observation Andy, I have trying to use random sample and this not work very well.\nMaybe using the time for validation as you mentioned will be the best!",
    "293587": "Exactly! It's better to use, say, 11.6-11.8 as training, and 11.9 as validation.",
    "293837": "I think setting `shuffle=False` in `sklearn.model_selection.train_test_split` will work, since the data appear to be in chronological order.",
    "293919": "...or maybe not.  One problem with doing a simple split is that the time gap between train and validation may not match that between train and test.  And the times of day for validation typically won't match the test set.  So I suppose I will try taking time into account explicitly.",
    "294081": "Indeed, I believe the best would be to take the latest say 20% of the data as a validation set. Seems fine to me to still feature engineer on the whole dataset (so deriving the frequency for the whole set). Does somebody knows if that would be reasonable?",
    "295068": "I'm trying a moving validation set: I train 3 times, first with 70% for train and 30% for validation( without shuffling), then 80% train and 20% validation, and 90% train 10% validation. I think this scheme is better than doing only one split for validation.",
    "295084": "Something like frequency on the whole training set will leak future frequency into the past beyond the time window available at prediction time for the test data. So for example, a training label on day 1 would be predicted with frequency characteristics partially drawn from day 3. For the test data, you can only look backward from later in day 4, not multiple days later. \n\nI'd guess it might still work ok to do total frequency, but a safer bet could be to use trailing frequency + forward frequency over only the remainder of that day to mimic what the test data gets to see. You'd basically end up computing 1 day, 2 day, and 3 day frequencies and using those features for clicks on each of the respective days.",
    "295416": "Konrad Banachewicz implements a time-aware validation set [here][1].\n\n [1]: https://www.kaggle.com/konradb/validation-set",
    "295419": "Thanks for the shout-out.",
    "296733": "Given that we can also get frequencies from test data, wouldn't getting counts on whole train set (prior to splitting to train/val) be the same as getting frequencies from train and test for the final run?  Would it make sense?",
    "296760": "The problem lies in what the model would learn if you pull in counts from the future. If you count across all 4 days and use that count information to make predictions for day 1 clicks, your model gets a look into the future that you don't have when you predict the test set. The total counts across train/test would be the same but the **meaning** of those counts would be different. The misalignment in meaning might not be a big problem, but it also might cause your model to generalize to the test data less well.",
    "296844": "Essentially, using counts more than a day in the future is a problem in the training phase rather than the validation phase.  If you do what @yulia suggests, you would get an (almost) legitimate validation score, but you might not be training the optimal model to predict the test data.  OTOH the extra count data might improve the model to the extent that it's worth including despite being unavailable at test time.  In principle, you can test this by getting separate validation scores for models that do and don't see counts from future days.  The test isn't perfect, because, when training for the validation score, you have one fewer day to train on than when the training for submission, so the early data are seeing further into the future.",
    "296851": "^+1, agree there's a tradeoff where any leakiness may be offset by signal gains from more count data. From the kernels it seems clear that the full count data approach works very well, but it's not clear if leakless counts would work better and would be difficult to truly validate for the reason you mentioned.\n\nAt the end of the day leaking only a couple days back in time may just not be a big deal.",
    "296927": "Joe Eddy, I like the way you explain things! Very clear to me, thanks!",
    "297013": "Thanks! Glad if it's helpful :)",
    "302835": "Are you able to build CV - LB constant in your validation setting? So far I have been trying different validation splits as discussed [here][1] by @Konrad and [here][2] by @Andy, and my CV is consistent with my LB but I am not able to build a CV - LB constant and that's why I think that I might be overfitting to public LB.  Validation is tough here.\n\n\n  [1]: https://www.kaggle.com/konradb/validation-set\n  [2]: https://www.kaggle.com/aharless/training-and-validation-data-pickle/code",
    "302852": "Perhaps we should exclude from validation data the range of IP codes that is not represented in the test data and seems to have different characteristics, as found, e.g., by CPMP in [this kernel][1].\n\n [1]: https://www.kaggle.com/cpmpml/ip-download-rates",
    "302865": "As discussed [here][1] by @yulia, \n\n    The pattern of IP assignment in Train does not mimic the one in test.\n\nSo we are not sure that IP's of same range have same characteristics in train and test and even if we remove the IP no.s from validation split which are not present in test set, how can we be sure that our all val set IP's are present in test set as they do not follow the mapping rule of train data.\n\n\n  [1]: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal",
    "302941": "The issue isn't about particular IP addresses but about ranges of codes.  The mapping between codes and IP addresses should be one-to-one.  In general we don't know much about how the codes were assigned, but we've at least identified a range of codes that is present in the training data and not the test data.  And we also know that range is associated with more downloads.  It's not clear what this means, but it seems like these might be unusual in some way and perhaps should be excluded from validation data.  Certainly if one were doing adversarial validation, the algorithm would identify those records as unlikely to be from the test data and therefore not use them for validation.",
    "303536": "So here's a small update: the validation setup I'm running at the moment is along the lines of my kernel \n\nhttps://www.kaggle.com/konradb/validation-set/code\n\nIt turns out that once you limit your evaluation set to the first hour (November 9th between 4 and 5 am) there is a match with LB at 4 significant digits. Obvious, I know - for some reason did not occur to me before, so I thought I'd share it here to save others some time :-)",
    "303668": "Maybe because public LB is based on the same hour in test, see: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833#302060",
    "303683": "I haven't been getting consistent results from the first hour, but maybe I need to tweak some things.",
    "303833": "CPMP - that's what i meant by 'obvious' :-)",
    "303859": "Konrad, it was not obvious to me so far ;)",
    "303872": "For me neither - that's why my sarcastic autopilot kicked in :-)",
    "304028": "Thanks for your findings, I am not sure if I am getting this right but since there is a match for the hour in validation set and LB, validating on it can get you closer to CV - LB ~ 0, would that not mean we might tend to overfit the leaderboard though indirectly by optimizing on a validation which is similar to LB data and not the test data.",
    "304255": "Has any one tried Kfold CV here?",
    "304260": "I am too trying to get my validation sorted, kfold seems to be a go-to option but I would assume that the  time series nature of the data would make random data allocation to folds, not a great cv strategy, \nI am looking something along the lines of \n\nSplitting training data into time-series1,time-series2, time-series3, time-series4\n\nfold 1 : training [time-series1], test [time-series2]\n\nfold 2 : training [time-series1 + time-series2], test [time-series3] \n\nfold 3 : training [time-series1 + time-series2 + time-series3], test [time-series4]",
    "304620": "Yes, this is what it means. Nevertheless, you can use the corresponding hours from 9th as validation also for the other hours from the test set (which are \"reserved\" for the private, final evaluation). Basically, you would like your validation set to resemble as closely as possible the test data.\n\nNote the last sentance applies only for Kaggle competitions :)",
    "304623": "It is an interesting idea worth to be verified! My personal intuition tells me that those IPs carry also other valuable information like conversion rate per os, device, many possible composed features, etc.",
    "304624": "Exactly my thought. Well, almost, I would use only those hours from 11.9 that correspond to the hours from the test data for 11.10. Note, I still wouldn't use the other hours from 11.9 as training.",
    "304631": "In the light of [Cute Chibiko's kernel][1], I don't think we should exclude those IPs from validation.  It seems the sponsors started with the test data when assigning IP codes, so the higher codes are just for IPs that didn't happen to appear in the test data, not ones that are qualitatively different from the those in the test data.  \n\nThe important thing is not to use the numeric values of the IP codes in training (so don't feed them directly into a GBT algorithm that will try to split its trees according to those numeric values), or it will overfit (both in training and in validation).  As long as you treat IP as a pure categorical feature and are careful to make your model ignore the implicit ordering of the codes, I expect IP has useful information.\n\n [1]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
    "304701": "Are you guys training with all the data? If not, how will you choose the subsample?\n\nMoreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?",
    "304738": "I'm only using 40m rows per run on my macbook, but its best to separate subsamples by time. The test set only uses 4 hours of data so you could separate those out into your validation set.",
    "304740": "Thank you very much for your answer. Imagining we would have 1 full day for test, how would you construct your validation set?",
    "304747": "I'd probably use the last day as validation.",
    "304757": "One more, if it is not asking a lot :)\n\nHow would I tune the hyperparameters of my model if I'd use the last day as validation? I'd need another subset to confirm that I am not overfitting that validation set, correct?",
    "304759": "In theory you should test it over many subsets. If you haven't done much FE then the super conservative parameters that are popular in the kernels should be a good place to start.",
    "304764": "Thanks a lot for your advices!",
    "305586": "what is trailing and forward frequency?",
    "306065": "shivraj I believe the idea is that the close match between day 9 hour 4 validation score and day 10 hour 4 score (LB) provides evidence that using the day 9 test hours for validation could be a good strategy. It's definitely not a good idea to only use hour 4 for validation, for the reason you describe (overfitting to public LB). However, since hour 4 shows very consistent validation/public LB results, we can hope that the other validation hours will also show behavior consistent with private LB results.",
    "306202": "Yep! That's the right way."
  },
  "source": "meta"
}