{
  "id": 53987,
  "title": "What's your estimated private lb score vs your public lb score?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53987",
  "author_name": "",
  "post_date": "2018-04-08T01:10:56.956870Z",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi Guys,</p>\n\n<p>Thank you for stopping by and I'm really curious about your estimated private lb score vs your public lb score.</p>\n\n<p>For me, I use data before day 8 16:00 as training data, use data between day 9 4:00 to 5:00 as validation data set 1(mimic test data on public lb) and data after day 9 5:00 as validation data set 2(mimic test data on private lb). My best single model(trained on data before day 8 16:00) has auc = 0.972 on validation set 1 and auc = 0.977 on validation set 2. Accordingly, my public lb score is 0.9718 and I guess my private lb score would be around 0.9765 to 0.9768.</p>\n\n<p>I'm really curious about your your estimated private lb score vs your public lb score and I really want to discuss with your guys here that:</p>\n\n<ol>\n<li><p>Does my validation strategy has a big flaw or could I improve it? (Till now, this is the best strategy I have that my validation auc is near to public lb score)</p></li>\n<li><p>I've discussed this issue a bit with @Joe Eddy before, and both of our validation auc on data after day 9 5:00 is much higher then our validation auc on data within day 9 4:00-5:00. Why is it the case? I mean why our model has a much lower auc within 4:00 to 5:00.?</p></li>\n<li><p>I really want to hear something from top rankers that given you have public lb score = 0.980~0.982, do you still estimate that your private lb score would be much higher than your public lb score?</p></li>\n<li><p>I understand that this question is very private and you may invest so much to get a better validation strategy than me. Thus if you don't want to share your insights, I'm totally fine I still appreciate your stopping. </p></li>\n</ol>\n\n<p>Thank you again and looking forward to hearing from you.</p>",
  "messages": [
    {
      "id": "310561",
      "postDate": "04/08/2018 01:10:56",
      "content": "<p>Hi Guys,</p>\n\n<p>Thank you for stopping by and I'm really curious about your estimated private lb score vs your public lb score.</p>\n\n<p>For me, I use data before day 8 16:00 as training data, use data between day 9 4:00 to 5:00 as validation data set 1(mimic test data on public lb) and data after day 9 5:00 as validation data set 2(mimic test data on private lb). My best single model(trained on data before day 8 16:00) has auc = 0.972 on validation set 1 and auc = 0.977 on validation set 2. Accordingly, my public lb score is 0.9718 and I guess my private lb score would be around 0.9765 to 0.9768.</p>\n\n<p>I'm really curious about your your estimated private lb score vs your public lb score and I really want to discuss with your guys here that:</p>\n\n<ol>\n<li><p>Does my validation strategy has a big flaw or could I improve it? (Till now, this is the best strategy I have that my validation auc is near to public lb score)</p></li>\n<li><p>I've discussed this issue a bit with @Joe Eddy before, and both of our validation auc on data after day 9 5:00 is much higher then our validation auc on data within day 9 4:00-5:00. Why is it the case? I mean why our model has a much lower auc within 4:00 to 5:00.?</p></li>\n<li><p>I really want to hear something from top rankers that given you have public lb score = 0.980~0.982, do you still estimate that your private lb score would be much higher than your public lb score?</p></li>\n<li><p>I understand that this question is very private and you may invest so much to get a better validation strategy than me. Thus if you don't want to share your insights, I'm totally fine I still appreciate your stopping. </p></li>\n</ol>\n\n<p>Thank you again and looking forward to hearing from you.</p>",
      "rawMarkdown": "Hi Guys,\n\nThank you for stopping by and I'm really curious about your estimated private lb score vs your public lb score.\n\nFor me, I use data before day 8 16:00 as training data, use data between day 9 4:00 to 5:00 as validation data set 1(mimic test data on public lb) and data after day 9 5:00 as validation data set 2(mimic test data on private lb). My best single model(trained on data before day 8 16:00) has auc = 0.972 on validation set 1 and auc = 0.977 on validation set 2. Accordingly, my public lb score is 0.9718 and I guess my private lb score would be around 0.9765 to 0.9768.\n\nI'm really curious about your your estimated private lb score vs your public lb score and I really want to discuss with your guys here that:\n\n1. Does my validation strategy has a big flaw or could I improve it? (Till now, this is the best strategy I have that my validation auc is near to public lb score)\n\n2. I've discussed this issue a bit with @Joe Eddy before, and both of our validation auc on data after day 9 5:00 is much higher then our validation auc on data within day 9 4:00-5:00. Why is it the case? I mean why our model has a much lower auc within 4:00 to 5:00.?\n\n3. I really want to hear something from top rankers that given you have public lb score = 0.980~0.982, do you still estimate that your private lb score would be much higher than your public lb score?\n\n4. I understand that this question is very private and you may invest so much to get a better validation strategy than me. Thus if you don't want to share your insights, I'm totally fine I still appreciate your stopping. \n\nThank you again and looking forward to hearing from you.",
      "votes": null
    },
    {
      "id": "310576",
      "postDate": "04/08/2018 02:35:18",
      "content": "<p>I also wonder what tricks do top rankers did. From my experience, maybe we should kick off all other hours which are not in the test. I suppose there could be big differences in users behavior in different hours, maybe we should train them sperately I guess? I'm not sure. Newbie's idea.</p>",
      "rawMarkdown": "I also wonder what tricks do top rankers did. From my experience, maybe we should kick off all other hours which are not in the test. I suppose there could be big differences in users behavior in different hours, maybe we should train them sperately I guess? I'm not sure. Newbie's idea.",
      "votes": null
    },
    {
      "id": "310578",
      "postDate": "04/08/2018 02:40:42",
      "content": "<p>Thank you for your reply! Actually I've tried only using data in hours that only appear in test. The score doesn't improve</p>",
      "rawMarkdown": "Thank you for your reply! Actually I've tried only using data in hours that only appear in test. The score doesn't improve",
      "votes": null
    },
    {
      "id": "310621",
      "postDate": "04/08/2018 05:32:17",
      "content": "<p>Pity, I can't figure out how to improve. ToT</p>",
      "rawMarkdown": "Pity, I can't figure out how to improve. ToT",
      "votes": null
    },
    {
      "id": "310663",
      "postDate": "04/08/2018 07:42:18",
      "content": "<p>There is always the risk that, with a high position on public LB, you will end up with a lower position on private LB; as well as you can overfit your training data (thing that you can avoid by using cross-validation), you can \"overfit\" the public LB i.e. you align your solution to the reduced (and unknown) sample from total testing data represented by the subset used for public LB score. What top competitors are doing typically is not very complicate, the principles can be boiled down to few points:</p>\n\n<ol>\n<li><p>do not trust any improvement in public LB score if it is not the\nresult of an improvement in the cross-validation score;  </p></li>\n<li><p>validate the best set of algorithm parameters + set of features\nfor their approach by doing cross-validation; usually they eliminate\nfeatures and see if this gives an improvement in the\ncross-validation score;  </p></li>\n<li><p>do a lot of trial and error;  </p></li>\n<li><p>do not look for quick wins.  </p></li>\n</ol>\n\n<p>Many of the top performers are exposing their strategy after a competition so reading their stories can help you understand about their methods.</p>",
      "rawMarkdown": "There is always the risk that, with a high position on public LB, you will end up with a lower position on private LB; as well as you can overfit your training data (thing that you can avoid by using cross-validation), you can \"overfit\" the public LB i.e. you align your solution to the reduced (and unknown) sample from total testing data represented by the subset used for public LB score. What top competitors are doing typically is not very complicate, the principles can be boiled down to few points:\n\n 1. do not trust any improvement in public LB score if it is not the\n   result of an improvement in the cross-validation score;  \n\n 2. validate the best set of algorithm parameters + set of features\n    for their approach by doing cross-validation; usually they eliminate\n    features and see if this gives an improvement in the\n    cross-validation score;  \n\n 3. do a lot of trial and error;  \n\n 4. do not look for quick wins.  \n\nMany of the top performers are exposing their strategy after a competition so reading their stories can help you understand about their methods.",
      "votes": null
    },
    {
      "id": "310671",
      "postDate": "04/08/2018 08:09:58",
      "content": "<p>I am playing around with just small portion of data but looking forward to touch full data soon. I am using a feature that classifies observations into most frequent <code>c(\"4\",\"5\",\"9\",\"10\",\"13\",\"14\")</code> , least frequent <code>c(\"6\",\"11\",\"15\")</code>  and remaining hours <code>c(\"7\",\"8\",\"12\")</code> in test data. Experiment shows improvement in score but just on small portion of training data. </p>\n\n<p>I one experiment I tried maximum 80 million observations (close to test set) and used hms (which has enormous factor levels) and score was not good. Dropping hms and keeping features based on hours leads to ~8-10 % duplicates. I would very much appreciate thoughts/general suggestions on making use of time features (specifically mins and secs) with more data. </p>",
      "rawMarkdown": "I am playing around with just small portion of data but looking forward to touch full data soon. I am using a feature that classifies observations into most frequent `c(\"4\",\"5\",\"9\",\"10\",\"13\",\"14\")` , least frequent `c(\"6\",\"11\",\"15\")`  and remaining hours `c(\"7\",\"8\",\"12\")` in test data. Experiment shows improvement in score but just on small portion of training data. \n\nI one experiment I tried maximum 80 million observations (close to test set) and used hms (which has enormous factor levels) and score was not good. Dropping hms and keeping features based on hours leads to ~8-10 % duplicates. I would very much appreciate thoughts/general suggestions on making use of time features (specifically mins and secs) with more data.",
      "votes": null
    },
    {
      "id": "310729",
      "postDate": "04/08/2018 12:13:15",
      "content": "<p>Pranav, why hms 'has enormous factor levels' is a problem? Is it categorical in your model? I convert day time to [0.0, 1.0] range, so I suppose this is similar to what you mean by hms.</p>\n\n<p>Regarding small portion of data vs size close to test, I observed decrease in score for the same model trained on smaller portion of data. Perhaps it makes sense to check your ideas on data size similar to test (I mean test.csv, which is ~18M lines).</p>",
      "rawMarkdown": "Pranav, why hms 'has enormous factor levels' is a problem? Is it categorical in your model? I convert day time to [0.0, 1.0] range, so I suppose this is similar to what you mean by hms.\n\nRegarding small portion of data vs size close to test, I observed decrease in score for the same model trained on smaller portion of data. Perhaps it makes sense to check your ideas on data size similar to test (I mean test.csv, which is ~18M lines).",
      "votes": null
    },
    {
      "id": "310735",
      "postDate": "04/08/2018 12:29:35",
      "content": "<p>Thanks Alex, \nAfter extracting hh:mm:ss from click time, I converted them to numeric seconds but used it as categorical. :( Scaling them definitely makes sense. And thanks a lot for sharing your observation and suggestion regarding data size. I very much appreciate that. </p>",
      "rawMarkdown": "Thanks Alex, \nAfter extracting hh:mm:ss from click time, I converted them to numeric seconds but used it as categorical. :( Scaling them definitely makes sense. And thanks a lot for sharing your observation and suggestion regarding data size. I very much appreciate that.",
      "votes": null
    },
    {
      "id": "312869",
      "postDate": "04/12/2018 13:38:27",
      "content": "<p>That's the questions I wanted to ask!</p>\n\n<p>I also make this kind of validation split and I also regularly observe this 0.005 gap between public valid and private valid AUC. I think this is something interesting to explore further.</p>",
      "rawMarkdown": "That's the questions I wanted to ask!\n\nI also make this kind of validation split and I also regularly observe this 0.005 gap between public valid and private valid AUC. I think this is something interesting to explore further.",
      "votes": null
    },
    {
      "id": "312913",
      "postDate": "04/12/2018 14:26:12",
      "content": "<p>Thank you for your reply. So your public lb is 0.9767 and your estimated private lb score should be around 0.981 to 0.982, right?</p>",
      "rawMarkdown": "Thank you for your reply. So your public lb is 0.9767 and your estimated private lb score should be around 0.981 to 0.982, right?",
      "votes": null
    },
    {
      "id": "313000",
      "postDate": "04/12/2018 16:43:36",
      "content": "<p>More like 0.980-0.981 but yeah ;)</p>",
      "rawMarkdown": "More like 0.980-0.981 but yeah ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 310576,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "04/08/2018 02:35:18",
      "content": "<p>I also wonder what tricks do top rankers did. From my experience, maybe we should kick off all other hours which are not in the test. I suppose there could be big differences in users behavior in different hours, maybe we should train them sperately I guess? I'm not sure. Newbie's idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 310578,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/08/2018 02:40:42",
          "content": "<p>Thank you for your reply! Actually I've tried only using data in hours that only appear in test. The score doesn't improve</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310621,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "04/08/2018 05:32:17",
          "content": "<p>Pity, I can't figure out how to improve. ToT</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310663,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "04/08/2018 07:42:18",
      "content": "<p>There is always the risk that, with a high position on public LB, you will end up with a lower position on private LB; as well as you can overfit your training data (thing that you can avoid by using cross-validation), you can \"overfit\" the public LB i.e. you align your solution to the reduced (and unknown) sample from total testing data represented by the subset used for public LB score. What top competitors are doing typically is not very complicate, the principles can be boiled down to few points:</p>\n\n<ol>\n<li><p>do not trust any improvement in public LB score if it is not the\nresult of an improvement in the cross-validation score;  </p></li>\n<li><p>validate the best set of algorithm parameters + set of features\nfor their approach by doing cross-validation; usually they eliminate\nfeatures and see if this gives an improvement in the\ncross-validation score;  </p></li>\n<li><p>do a lot of trial and error;  </p></li>\n<li><p>do not look for quick wins.  </p></li>\n</ol>\n\n<p>Many of the top performers are exposing their strategy after a competition so reading their stories can help you understand about their methods.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 310671,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/08/2018 08:09:58",
      "content": "<p>I am playing around with just small portion of data but looking forward to touch full data soon. I am using a feature that classifies observations into most frequent <code>c(\"4\",\"5\",\"9\",\"10\",\"13\",\"14\")</code> , least frequent <code>c(\"6\",\"11\",\"15\")</code>  and remaining hours <code>c(\"7\",\"8\",\"12\")</code> in test data. Experiment shows improvement in score but just on small portion of training data. </p>\n\n<p>I one experiment I tried maximum 80 million observations (close to test set) and used hms (which has enormous factor levels) and score was not good. Dropping hms and keeping features based on hours leads to ~8-10 % duplicates. I would very much appreciate thoughts/general suggestions on making use of time features (specifically mins and secs) with more data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 310729,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/08/2018 12:13:15",
          "content": "<p>Pranav, why hms 'has enormous factor levels' is a problem? Is it categorical in your model? I convert day time to [0.0, 1.0] range, so I suppose this is similar to what you mean by hms.</p>\n\n<p>Regarding small portion of data vs size close to test, I observed decrease in score for the same model trained on smaller portion of data. Perhaps it makes sense to check your ideas on data size similar to test (I mean test.csv, which is ~18M lines).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310735,
          "author_name": "pranav84",
          "author_url": "",
          "post_date": "04/08/2018 12:29:35",
          "content": "<p>Thanks Alex, \nAfter extracting hh:mm:ss from click time, I converted them to numeric seconds but used it as categorical. :( Scaling them definitely makes sense. And thanks a lot for sharing your observation and suggestion regarding data size. I very much appreciate that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 312869,
      "author_name": "kamilkk",
      "author_url": "",
      "post_date": "04/12/2018 13:38:27",
      "content": "<p>That's the questions I wanted to ask!</p>\n\n<p>I also make this kind of validation split and I also regularly observe this 0.005 gap between public valid and private valid AUC. I think this is something interesting to explore further.</p>",
      "votes": null,
      "replies": [
        {
          "id": 312913,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/12/2018 14:26:12",
          "content": "<p>Thank you for your reply. So your public lb is 0.9767 and your estimated private lb score should be around 0.981 to 0.982, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313000,
          "author_name": "kamilkk",
          "author_url": "",
          "post_date": "04/12/2018 16:43:36",
          "content": "<p>More like 0.980-0.981 but yeah ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "310561": "Hi Guys,\n\nThank you for stopping by and I'm really curious about your estimated private lb score vs your public lb score.\n\nFor me, I use data before day 8 16:00 as training data, use data between day 9 4:00 to 5:00 as validation data set 1(mimic test data on public lb) and data after day 9 5:00 as validation data set 2(mimic test data on private lb). My best single model(trained on data before day 8 16:00) has auc = 0.972 on validation set 1 and auc = 0.977 on validation set 2. Accordingly, my public lb score is 0.9718 and I guess my private lb score would be around 0.9765 to 0.9768.\n\nI'm really curious about your your estimated private lb score vs your public lb score and I really want to discuss with your guys here that:\n\n1. Does my validation strategy has a big flaw or could I improve it? (Till now, this is the best strategy I have that my validation auc is near to public lb score)\n\n2. I've discussed this issue a bit with @Joe Eddy before, and both of our validation auc on data after day 9 5:00 is much higher then our validation auc on data within day 9 4:00-5:00. Why is it the case? I mean why our model has a much lower auc within 4:00 to 5:00.?\n\n3. I really want to hear something from top rankers that given you have public lb score = 0.980~0.982, do you still estimate that your private lb score would be much higher than your public lb score?\n\n4. I understand that this question is very private and you may invest so much to get a better validation strategy than me. Thus if you don't want to share your insights, I'm totally fine I still appreciate your stopping. \n\nThank you again and looking forward to hearing from you.",
    "310576": "I also wonder what tricks do top rankers did. From my experience, maybe we should kick off all other hours which are not in the test. I suppose there could be big differences in users behavior in different hours, maybe we should train them sperately I guess? I'm not sure. Newbie's idea.",
    "310578": "Thank you for your reply! Actually I've tried only using data in hours that only appear in test. The score doesn't improve",
    "310621": "Pity, I can't figure out how to improve. ToT",
    "310663": "There is always the risk that, with a high position on public LB, you will end up with a lower position on private LB; as well as you can overfit your training data (thing that you can avoid by using cross-validation), you can \"overfit\" the public LB i.e. you align your solution to the reduced (and unknown) sample from total testing data represented by the subset used for public LB score. What top competitors are doing typically is not very complicate, the principles can be boiled down to few points:\n\n 1. do not trust any improvement in public LB score if it is not the\n   result of an improvement in the cross-validation score;  \n\n 2. validate the best set of algorithm parameters + set of features\n    for their approach by doing cross-validation; usually they eliminate\n    features and see if this gives an improvement in the\n    cross-validation score;  \n\n 3. do a lot of trial and error;  \n\n 4. do not look for quick wins.  \n\nMany of the top performers are exposing their strategy after a competition so reading their stories can help you understand about their methods.",
    "310671": "I am playing around with just small portion of data but looking forward to touch full data soon. I am using a feature that classifies observations into most frequent `c(\"4\",\"5\",\"9\",\"10\",\"13\",\"14\")` , least frequent `c(\"6\",\"11\",\"15\")`  and remaining hours `c(\"7\",\"8\",\"12\")` in test data. Experiment shows improvement in score but just on small portion of training data. \n\nI one experiment I tried maximum 80 million observations (close to test set) and used hms (which has enormous factor levels) and score was not good. Dropping hms and keeping features based on hours leads to ~8-10 % duplicates. I would very much appreciate thoughts/general suggestions on making use of time features (specifically mins and secs) with more data.",
    "310729": "Pranav, why hms 'has enormous factor levels' is a problem? Is it categorical in your model? I convert day time to [0.0, 1.0] range, so I suppose this is similar to what you mean by hms.\n\nRegarding small portion of data vs size close to test, I observed decrease in score for the same model trained on smaller portion of data. Perhaps it makes sense to check your ideas on data size similar to test (I mean test.csv, which is ~18M lines).",
    "310735": "Thanks Alex, \nAfter extracting hh:mm:ss from click time, I converted them to numeric seconds but used it as categorical. :( Scaling them definitely makes sense. And thanks a lot for sharing your observation and suggestion regarding data size. I very much appreciate that.",
    "312869": "That's the questions I wanted to ask!\n\nI also make this kind of validation split and I also regularly observe this 0.005 gap between public valid and private valid AUC. I think this is something interesting to explore further.",
    "312913": "Thank you for your reply. So your public lb is 0.9767 and your estimated private lb score should be around 0.981 to 0.982, right?",
    "313000": "More like 0.980-0.981 but yeah ;)"
  },
  "source": "meta"
}