{
  "id": 53325,
  "title": "How do you test your solution?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53325",
  "author_name": "little_snail",
  "post_date": "2018-03-29T12:17:04.612000",
  "votes": 5,
  "comment_count": 28,
  "views": 0,
  "content": "<p>I use day8 to train my model, and use day9 to test my model. </p>\n\n<p>I use day9 to train my model, and predict day10</p>",
  "messages": [
    {
      "id": 305792,
      "postDate": "2018-03-29T12:17:04.613Z",
      "content": "<p>I use day8 to train my model, and use day9 to test my model. </p>\n\n<p>I use day9 to train my model, and predict day10</p>",
      "rawMarkdown": "I use day8 to train my model, and use day9 to test my model. \n\nI use day9 to train my model, and predict day10",
      "votes": 5
    },
    {
      "id": 305847,
      "postDate": "2018-03-29T14:01:34.937Z",
      "content": "<p>right now i validate on the significant hours of the test set so 4,5,9,10,14,15 from day 9. Seems to work fine for me atleast.</p>",
      "rawMarkdown": "right now i validate on the significant hours of the test set so 4,5,9,10,14,15 from day 9. Seems to work fine for me atleast.",
      "votes": 1
    },
    {
      "id": 305822,
      "postDate": "2018-03-29T13:00:55.160Z",
      "content": "<p>I train my model with day 8, validate on day 9, retrain on day 8 + 9 and predict test set. So far this setting is working good for me.</p>",
      "rawMarkdown": "I train my model with day 8, validate on day 9, retrain on day 8 + 9 and predict test set. So far this setting is working good for me.",
      "votes": 1,
      "replies": [
        {
          "id": 306041,
          "postDate": "2018-03-29T19:26:52.520Z",
          "content": "<p>so are you considering the whole 9th day for validation or specific hours similar to test set?</p>",
          "rawMarkdown": "so are you considering the whole 9th day for validation or specific hours similar to test set?"
        },
        {
          "id": 306052,
          "postDate": "2018-03-29T19:58:17.230Z",
          "content": "<p>Whole day.</p>",
          "rawMarkdown": "Whole day."
        },
        {
          "id": 306233,
          "postDate": "2018-03-30T04:54:52.060Z",
          "content": "<p>Why so? Shouldn't the validation set be parallel to test set?</p>",
          "rawMarkdown": "Why so? Shouldn't the validation set be parallel to test set?",
          "votes": 1
        },
        {
          "id": 306377,
          "postDate": "2018-03-30T10:31:40.013Z",
          "content": "<p>Because I need my model to learn from the closest past day. As  I judge my validation score based on the the training till last day. <br></p>",
          "rawMarkdown": "Because I need my model to learn from the closest past day. As  I judge my validation score based on the the training till last day. <br>"
        },
        {
          "id": 306378,
          "postDate": "2018-03-30T10:34:23.057Z",
          "content": "<p>Because training on day 0, 1 and validating on day 3 does not make sense to me.  </p>",
          "rawMarkdown": "Because training on day 0, 1 and validating on day 3 does not make sense to me.  "
        }
      ]
    },
    {
      "id": 307047,
      "postDate": "2018-03-31T17:48:41.560Z",
      "content": "<p>@Joe Eddy: I am using day 9 for validation and total for submission. I cannot get consistent local validation - LB score. I am using Lgb. What model are you using? Did you have to change the parameters many times to get the consistency? I even tried with no feature engineering, without the IP, still the validation and LB score are not consistent. My local CV score is .98xx and the LB is .96xx</p>",
      "rawMarkdown": "@Joe Eddy: I am using day 9 for validation and total for submission. I cannot get consistent local validation - LB score. I am using Lgb. What model are you using? Did you have to change the parameters many times to get the consistency? I even tried with no feature engineering, without the IP, still the validation and LB score are not consistent. My local CV score is .98xx and the LB is .96xx",
      "replies": [
        {
          "id": 307116,
          "postDate": "2018-03-31T20:54:17.030Z",
          "content": "<p>I'm also using lgb, and no, so far the consistency does not seem parameter sensitive. I've only tried this once so far, but a small change in parameters that increased day 9 hour 4 validation by ~.0002 had the exact same PLB effect. But like I said before, it's impossible to know if the consistency will hold up as I try different things, and I might just be getting lucky right now.</p>\n\n<p>.98xx to .96xx seems like a very large gap, I would double check your setup. Are you using features derived from future days when training?</p>",
          "rawMarkdown": "I'm also using lgb, and no, so far the consistency does not seem parameter sensitive. I've only tried this once so far, but a small change in parameters that increased day 9 hour 4 validation by ~.0002 had the exact same PLB effect. But like I said before, it's impossible to know if the consistency will hold up as I try different things, and I might just be getting lucky right now.\n\n.98xx to .96xx seems like a very large gap, I would double check your setup. Are you using features derived from future days when training?",
          "votes": 1
        },
        {
          "id": 307138,
          "postDate": "2018-03-31T22:23:12.267Z",
          "content": "<p>Thank you so much for the reply. I don't think i am leaking any future day information. I combine my train and test to do feature engineering but the features i engineered so far all grouped by first <strong>IP</strong> then by <strong>day</strong>. Then for validation i am using day 9 so i do not think i am leaking any future day information into my model. Right now i am running a model without any feature engineering to see the difference.</p>",
          "rawMarkdown": "Thank you so much for the reply. I don't think i am leaking any future day information. I combine my train and test to do feature engineering but the features i engineered so far all grouped by first **IP** then by **day**. Then for validation i am using day 9 so i do not think i am leaking any future day information into my model. Right now i am running a model without any feature engineering to see the difference."
        },
        {
          "id": 307838,
          "postDate": "2018-04-02T15:24:29.240Z",
          "content": "<p>Deleting this post. Sorry Eddy</p>",
          "rawMarkdown": "Deleting this post. Sorry Eddy"
        },
        {
          "id": 307843,
          "postDate": "2018-04-02T15:27:59.500Z",
          "content": "<p>@Joe Eddy, haha Joe you are such a nice person. But I would recommend you to post your paypal address here, even we couldn't win the prize of this competition, we still have the income of debugging.</p>",
          "rawMarkdown": "@Joe Eddy, haha Joe you are such a nice person. But I would recommend you to post your paypal address here, even we couldn't win the prize of this competition, we still have the income of debugging.",
          "votes": 2
        }
      ]
    },
    {
      "id": 306048,
      "postDate": "2018-03-29T19:46:29.247Z",
      "content": "<p>I think using day 9 test hours for validation is the way to go. Right now my validation AUC on day 9 hour 4 is within .0002 of my public LB score.</p>",
      "rawMarkdown": "I think using day 9 test hours for validation is the way to go. Right now my validation AUC on day 9 hour 4 is within .0002 of my public LB score.",
      "replies": [
        {
          "id": 306072,
          "postDate": "2018-03-29T20:23:53.330Z",
          "content": "<p>I'm getting validation results for day 9 hour 4 that are consistently lower than public LB scores by an inconsistent amount.  But this is looking at a variety of different kinds of models, most of which come from public kernels and are sloppy about some aggregation features.  I guess I should try doing something with perfectly clean aggregations and see if it works.</p>",
          "rawMarkdown": "I'm getting validation results for day 9 hour 4 that are consistently lower than public LB scores by an inconsistent amount.  But this is looking at a variety of different kinds of models, most of which come from public kernels and are sloppy about some aggregation features.  I guess I should try doing something with perfectly clean aggregations and see if it works."
        },
        {
          "id": 306076,
          "postDate": "2018-03-29T20:28:46.537Z",
          "content": "<p>Yeah, for what it's worth I'm using leakless features. Also, the consistency I'm seeing is based on only a few submissions so hard to say that it will hold up.</p>",
          "rawMarkdown": "Yeah, for what it's worth I'm using leakless features. Also, the consistency I'm seeing is based on only a few submissions so hard to say that it will hold up.",
          "votes": 2
        },
        {
          "id": 306089,
          "postDate": "2018-03-29T20:48:51.963Z",
          "content": "<p>To minimize leakage I aggregate features by day, so that today aggregation is not influenced by the next day. What more can be done?</p>",
          "rawMarkdown": "To minimize leakage I aggregate features by day, so that today aggregation is not influenced by the next day. What more can be done?"
        },
        {
          "id": 306098,
          "postDate": "2018-03-29T21:12:58.400Z",
          "content": "<p>That's the right idea. Just think carefully about what information is available at prediction time and make sure that your features are not skewed by the time range you compute them on. For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.</p>",
          "rawMarkdown": "That's the right idea. Just think carefully about what information is available at prediction time and make sure that your features are not skewed by the time range you compute them on. For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.",
          "votes": 3
        },
        {
          "id": 306104,
          "postDate": "2018-03-29T21:24:45.733Z",
          "content": "<p>Thanks for the great explanation :) </p>",
          "rawMarkdown": "Thanks for the great explanation :) "
        },
        {
          "id": 306151,
          "postDate": "2018-03-30T00:23:03.240Z",
          "content": "<p>My validation score is about 0.0004 lower than my public lb score. Could I ask what is your estimation of your private lb score for your current model?</p>",
          "rawMarkdown": "My validation score is about 0.0004 lower than my public lb score. Could I ask what is your estimation of your private lb score for your current model?"
        },
        {
          "id": 306158,
          "postDate": "2018-03-30T00:58:55.230Z",
          "content": "<p>Sure, it's about .977 (day 9 all of the test hours).</p>\n\n<p>Now that I think of it, it might make sense to drop hour 4 from that validation set to get a true match with the private LB hours. It's funny how this competition is encouraging us to overfit to specific times of day.</p>",
          "rawMarkdown": "Sure, it's about .977 (day 9 all of the test hours).\n\nNow that I think of it, it might make sense to drop hour 4 from that validation set to get a true match with the private LB hours. It's funny how this competition is encouraging us to overfit to specific times of day.\n",
          "votes": 2
        },
        {
          "id": 306173,
          "postDate": "2018-03-30T01:50:38.457Z",
          "content": "<p>Thank you very much for your reply. I got several thoughts regarding your result and my result: 1. is there a method to specifically overfit specific times of day? 2. If not, those top rankers are monsters coz their private lb score should be around 0.985... , really want to hear from some top rankers of what they did 3. some kernels have very high validation auc coz they choose last rows which is around hour 14 and 15, and based on our validation score we could find our models perform differently on different hours, which is very interesting</p>",
          "rawMarkdown": "Thank you very much for your reply. I got several thoughts regarding your result and my result: 1. is there a method to specifically overfit specific times of day? 2. If not, those top rankers are monsters coz their private lb score should be around 0.985... , really want to hear from some top rankers of what they did 3. some kernels have very high validation auc coz they choose last rows which is around hour 14 and 15, and based on our validation score we could find our models perform differently on different hours, which is very interesting"
        },
        {
          "id": 306185,
          "postDate": "2018-03-30T02:08:32.743Z",
          "content": "<p>HI, Joe Eddy, after the validation, to get a model, using the whole data is a good idea, or just the first two days?Thanks.</p>",
          "rawMarkdown": "HI, Joe Eddy, after the validation, to get a model, using the whole data is a good idea, or just the first two days?Thanks."
        },
        {
          "id": 306189,
          "postDate": "2018-03-30T02:22:52.807Z",
          "content": "<p>I think it depends on the features you're using and what your RAM can accommodate. You should expect that more (relevant) data will tend to improve your model, but increasing the sample may have diminishing returns. If you're computationally restrained, I would prioritize including more (good) features over including more rows. </p>",
          "rawMarkdown": "I think it depends on the features you're using and what your RAM can accommodate. You should expect that more (relevant) data will tend to improve your model, but increasing the sample may have diminishing returns. If you're computationally restrained, I would prioritize including more (good) features over including more rows. ",
          "votes": 1
        },
        {
          "id": 306205,
          "postDate": "2018-03-30T03:07:15.260Z",
          "content": "<p>Thanks for your advices. :D</p>",
          "rawMarkdown": "Thanks for your advices. :D"
        },
        {
          "id": 306206,
          "postDate": "2018-03-30T03:07:18.317Z",
          "content": "<p>@Snorlax I don't know of a specific way you would overfit to certain hours of the day beyond just only including those hours in your validation set. I'd suspect a model that does well on those hours will tend to do well generally, but it's not impossible to imagine a world where a model that works really well for certain times might not fare as well for other times.</p>\n\n<p>Based on my own experience so far I would guess that top scorers have higher full test hours validation than their public LB scores, but it's definitely hard to know. It might depend a lot on the model used and/or features.</p>\n\n<p>It's not surprising to me that kernel validation scores for hours later in the same day as training look artificially good. Both because it's easier for the model to understand behaviors on that exact day than to generalize them to the next day, and because some of these kernels are using leaky features that pull in data from the future. I think that validation without a day gap and future day leaking features are both a mistake.</p>",
          "rawMarkdown": "@Snorlax I don't know of a specific way you would overfit to certain hours of the day beyond just only including those hours in your validation set. I'd suspect a model that does well on those hours will tend to do well generally, but it's not impossible to imagine a world where a model that works really well for certain times might not fare as well for other times.\n\nBased on my own experience so far I would guess that top scorers have higher full test hours validation than their public LB scores, but it's definitely hard to know. It might depend a lot on the model used and/or features.\n\nIt's not surprising to me that kernel validation scores for hours later in the same day as training look artificially good. Both because it's easier for the model to understand behaviors on that exact day than to generalize them to the next day, and because some of these kernels are using leaky features that pull in data from the future. I think that validation without a day gap and future day leaking features are both a mistake."
        },
        {
          "id": 306234,
          "postDate": "2018-03-30T04:56:05.853Z",
          "content": "<p>what's a leakless feature?</p>",
          "rawMarkdown": "what's a leakless feature?\n"
        },
        {
          "id": 306445,
          "postDate": "2018-03-30T12:40:51.990Z",
          "content": "<p>Influence of future on present. @Joe explained it like this <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098\">here</a></p>\n\n<pre><code>For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n</code></pre>",
          "rawMarkdown": "Influence of future on present. @Joe explained it like this [here][1]\n\n    For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n\n\n\n\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098"
        },
        {
          "id": 306769,
          "postDate": "2018-03-31T01:08:02.080Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 305847,
      "author_name": "ms",
      "author_url": "",
      "post_date": "2018-03-29T14:01:34.937000",
      "content": "<p>right now i validate on the significant hours of the test set so 4,5,9,10,14,15 from day 9. Seems to work fine for me atleast.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 305822,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-03-29T13:00:55.160000",
      "content": "<p>I train my model with day 8, validate on day 9, retrain on day 8 + 9 and predict test set. So far this setting is working good for me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 306041,
          "author_name": "Nitish Kumar",
          "author_url": "",
          "post_date": "2018-03-29T19:26:52.520000",
          "content": "<p>so are you considering the whole 9th day for validation or specific hours similar to test set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306052,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-29T19:58:17.230000",
          "content": "<p>Whole day.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306233,
          "author_name": "Nitish Kumar",
          "author_url": "",
          "post_date": "2018-03-30T04:54:52.060000",
          "content": "<p>Why so? Shouldn't the validation set be parallel to test set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 306377,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-30T10:31:40.013000",
          "content": "<p>Because I need my model to learn from the closest past day. As  I judge my validation score based on the the training till last day. <br></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306378,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-30T10:34:23.057000",
          "content": "<p>Because training on day 0, 1 and validating on day 3 does not make sense to me.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 307047,
      "author_name": "Anon",
      "author_url": "",
      "post_date": "2018-03-31T17:48:41.560000",
      "content": "<p>@Joe Eddy: I am using day 9 for validation and total for submission. I cannot get consistent local validation - LB score. I am using Lgb. What model are you using? Did you have to change the parameters many times to get the consistency? I even tried with no feature engineering, without the IP, still the validation and LB score are not consistent. My local CV score is .98xx and the LB is .96xx</p>",
      "votes": 0,
      "replies": [
        {
          "id": 307116,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-31T20:54:17.030000",
          "content": "<p>I'm also using lgb, and no, so far the consistency does not seem parameter sensitive. I've only tried this once so far, but a small change in parameters that increased day 9 hour 4 validation by ~.0002 had the exact same PLB effect. But like I said before, it's impossible to know if the consistency will hold up as I try different things, and I might just be getting lucky right now.</p>\n\n<p>.98xx to .96xx seems like a very large gap, I would double check your setup. Are you using features derived from future days when training?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 307138,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2018-03-31T22:23:12.267000",
          "content": "<p>Thank you so much for the reply. I don't think i am leaking any future day information. I combine my train and test to do feature engineering but the features i engineered so far all grouped by first <strong>IP</strong> then by <strong>day</strong>. Then for validation i am using day 9 so i do not think i am leaking any future day information into my model. Right now i am running a model without any feature engineering to see the difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 307838,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2018-04-02T15:24:29.240000",
          "content": "<p>Deleting this post. Sorry Eddy</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 307843,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-02T15:27:59.500000",
          "content": "<p>@Joe Eddy, haha Joe you are such a nice person. But I would recommend you to post your paypal address here, even we couldn't win the prize of this competition, we still have the income of debugging.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 306048,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-03-29T19:46:29.247000",
      "content": "<p>I think using day 9 test hours for validation is the way to go. Right now my validation AUC on day 9 hour 4 is within .0002 of my public LB score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 306072,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-29T20:23:53.330000",
          "content": "<p>I'm getting validation results for day 9 hour 4 that are consistently lower than public LB scores by an inconsistent amount.  But this is looking at a variety of different kinds of models, most of which come from public kernels and are sloppy about some aggregation features.  I guess I should try doing something with perfectly clean aggregations and see if it works.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306076,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-29T20:28:46.537000",
          "content": "<p>Yeah, for what it's worth I'm using leakless features. Also, the consistency I'm seeing is based on only a few submissions so hard to say that it will hold up.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 306089,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-29T20:48:51.963000",
          "content": "<p>To minimize leakage I aggregate features by day, so that today aggregation is not influenced by the next day. What more can be done?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306098,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-29T21:12:58.400000",
          "content": "<p>That's the right idea. Just think carefully about what information is available at prediction time and make sure that your features are not skewed by the time range you compute them on. For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 306104,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-29T21:24:45.733000",
          "content": "<p>Thanks for the great explanation :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306151,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-03-30T00:23:03.240000",
          "content": "<p>My validation score is about 0.0004 lower than my public lb score. Could I ask what is your estimation of your private lb score for your current model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306158,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-30T00:58:55.230000",
          "content": "<p>Sure, it's about .977 (day 9 all of the test hours).</p>\n\n<p>Now that I think of it, it might make sense to drop hour 4 from that validation set to get a true match with the private LB hours. It's funny how this competition is encouraging us to overfit to specific times of day.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 306173,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-03-30T01:50:38.457000",
          "content": "<p>Thank you very much for your reply. I got several thoughts regarding your result and my result: 1. is there a method to specifically overfit specific times of day? 2. If not, those top rankers are monsters coz their private lb score should be around 0.985... , really want to hear from some top rankers of what they did 3. some kernels have very high validation auc coz they choose last rows which is around hour 14 and 15, and based on our validation score we could find our models perform differently on different hours, which is very interesting</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306185,
          "author_name": "YulinGUO",
          "author_url": "",
          "post_date": "2018-03-30T02:08:32.743000",
          "content": "<p>HI, Joe Eddy, after the validation, to get a model, using the whole data is a good idea, or just the first two days?Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306189,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-30T02:22:52.807000",
          "content": "<p>I think it depends on the features you're using and what your RAM can accommodate. You should expect that more (relevant) data will tend to improve your model, but increasing the sample may have diminishing returns. If you're computationally restrained, I would prioritize including more (good) features over including more rows. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 306205,
          "author_name": "YulinGUO",
          "author_url": "",
          "post_date": "2018-03-30T03:07:15.260000",
          "content": "<p>Thanks for your advices. :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306206,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-30T03:07:18.317000",
          "content": "<p>@Snorlax I don't know of a specific way you would overfit to certain hours of the day beyond just only including those hours in your validation set. I'd suspect a model that does well on those hours will tend to do well generally, but it's not impossible to imagine a world where a model that works really well for certain times might not fare as well for other times.</p>\n\n<p>Based on my own experience so far I would guess that top scorers have higher full test hours validation than their public LB scores, but it's definitely hard to know. It might depend a lot on the model used and/or features.</p>\n\n<p>It's not surprising to me that kernel validation scores for hours later in the same day as training look artificially good. Both because it's easier for the model to understand behaviors on that exact day than to generalize them to the next day, and because some of these kernels are using leaky features that pull in data from the future. I think that validation without a day gap and future day leaking features are both a mistake.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306234,
          "author_name": "Nitish Kumar",
          "author_url": "",
          "post_date": "2018-03-30T04:56:05.853000",
          "content": "<p>what's a leakless feature?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306445,
          "author_name": "Sohaib Omar",
          "author_url": "",
          "post_date": "2018-03-30T12:40:51.990000",
          "content": "<p>Influence of future on present. @Joe explained it like this <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098\">here</a></p>\n\n<pre><code>For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306769,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-03-31T01:08:02.080000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "305792": "I use day8 to train my model, and use day9 to test my model. \n\nI use day9 to train my model, and predict day10",
    "305847": "right now i validate on the significant hours of the test set so 4,5,9,10,14,15 from day 9. Seems to work fine for me atleast.",
    "305822": "I train my model with day 8, validate on day 9, retrain on day 8 + 9 and predict test set. So far this setting is working good for me.",
    "307047": "@Joe Eddy: I am using day 9 for validation and total for submission. I cannot get consistent local validation - LB score. I am using Lgb. What model are you using? Did you have to change the parameters many times to get the consistency? I even tried with no feature engineering, without the IP, still the validation and LB score are not consistent. My local CV score is .98xx and the LB is .96xx",
    "306048": "I think using day 9 test hours for validation is the way to go. Right now my validation AUC on day 9 hour 4 is within .0002 of my public LB score."
  }
}