{
  "id": 192919,
  "title": "CV vs LB scores",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192919",
  "author_name": "Vopani",
  "post_date": "2020-10-24T06:51:36.512000",
  "votes": 89,
  "comment_count": 162,
  "views": 0,
  "content": "<p>With quite a large dataset at hand and a unique submission format there can be a lot of interesting ideas, models and validation approaches that can be used.</p>\n<p>Large datasets often give a stable performance and so far my experience has been positive.</p>\n<p>Opening this thread for anyone willing to share their validation / leaderboard approaches and / or scores.</p>",
  "messages": [
    {
      "id": 1058728,
      "postDate": "2020-10-24T06:51:36.513Z",
      "content": "<p>With quite a large dataset at hand and a unique submission format there can be a lot of interesting ideas, models and validation approaches that can be used.</p>\n<p>Large datasets often give a stable performance and so far my experience has been positive.</p>\n<p>Opening this thread for anyone willing to share their validation / leaderboard approaches and / or scores.</p>",
      "rawMarkdown": "With quite a large dataset at hand and a unique submission format there can be a lot of interesting ideas, models and validation approaches that can be used.\n\nLarge datasets often give a stable performance and so far my experience has been positive.\n\nOpening this thread for anyone willing to share their validation / leaderboard approaches and / or scores.",
      "votes": 88
    },
    {
      "id": 1112090,
      "postDate": "2020-12-14T09:16:54.680Z",
      "content": "<ul>\n<li>CV: 0.80664, LB: 0.806</li>\n<li>Single SAINT-based model</li>\n<li>without lectures</li>\n</ul>",
      "rawMarkdown": "- CV: 0.80664, LB: 0.806\n- Single SAINT-based model\n- without lectures",
      "votes": 18,
      "replies": [
        {
          "id": 1112092,
          "postDate": "2020-12-14T09:20:43.613Z",
          "content": "<p>Can you share the model size? I've reach around 0.792 using SAINT+ now but still using 128 d_model and 4 num_layers. Would increasing d_model help much?</p>",
          "rawMarkdown": "Can you share the model size? I've reach around 0.792 using SAINT+ now but still using 128 d_model and 4 num_layers. Would increasing d_model help much?"
        },
        {
          "id": 1112097,
          "postDate": "2020-12-14T09:23:38.140Z",
          "content": "<p>same as the paper: d_model = 512, num_layers = 4</p>",
          "rawMarkdown": "same as the paper: d_model = 512, num_layers = 4",
          "votes": 3
        },
        {
          "id": 1112109,
          "postDate": "2020-12-14T09:31:12.723Z",
          "content": "<p>Thanks, I'll see how the scaled up version works for me as well. Issue is, one model training takes up almost all of the week's GPU quota :/</p>\n<p>Have to do iterations once a week only…</p>\n<p>Anyway, best of luck. Include Lectures btw, they increased my score by ~0.01 i think</p>",
          "rawMarkdown": "Thanks, I'll see how the scaled up version works for me as well. Issue is, one model training takes up almost all of the week's GPU quota :/\n\nHave to do iterations once a week only...\n\nAnyway, best of luck. Include Lectures btw, they increased my score by ~0.01 i think",
          "votes": 3
        },
        {
          "id": 1112131,
          "postDate": "2020-12-14T10:07:05.447Z",
          "content": "<p>legend 🙏🙏🙏🙏🙏</p>",
          "rawMarkdown": "legend 🙏🙏🙏🙏🙏",
          "votes": 1
        },
        {
          "id": 1112456,
          "postDate": "2020-12-14T16:11:45.440Z",
          "content": "<p>thanks, I will also try saint++ model !</p>",
          "rawMarkdown": "thanks, I will also try saint++ model !",
          "votes": 4
        },
        {
          "id": 1112913,
          "postDate": "2020-12-15T02:27:40.230Z",
          "content": "<p>Nice socre, do you mind share how you split train and valid ?</p>",
          "rawMarkdown": "Nice socre, do you mind share how you split train and valid ?"
        },
        {
          "id": 1113487,
          "postDate": "2020-12-15T13:42:04.707Z",
          "content": "<p>same as <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> split, dropping last 2.5M samples.</p>",
          "rawMarkdown": "same as [this](https://www.kaggle.com/its7171/cv-strategy) split, dropping last 2.5M samples.",
          "votes": 2
        },
        {
          "id": 1115459,
          "postDate": "2020-12-16T10:01:27.203Z",
          "content": "<p>This is what looks like when you surrond by all genious people 😊</p>",
          "rawMarkdown": "This is what looks like when you surrond by all genious people 😊",
          "votes": 2
        },
        {
          "id": 1117434,
          "postDate": "2020-12-18T03:50:02.353Z",
          "content": "<p><a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> ; Just curious, If you won't mind, Can you share how are you handling larger sequences? Splitting them up of picking randomly (in sort of time-sorted fashion)? <br>\nTy!</p>",
          "rawMarkdown": "@sakami ; Just curious, If you won't mind, Can you share how are you handling larger sequences? Splitting them up of picking randomly (in sort of time-sorted fashion)? \nTy!"
        },
        {
          "id": 1123749,
          "postDate": "2020-12-23T13:27:21.937Z",
          "content": "<p>Hi Sakami, did you get that score using only the features used in the paper or did you use some additional features. I also tried the the saint approach, but it cant beat my single lstm (0.797)</p>",
          "rawMarkdown": "Hi Sakami, did you get that score using only the features used in the paper or did you use some additional features. I also tried the the saint approach, but it cant beat my single lstm (0.797)",
          "votes": 3
        },
        {
          "id": 1124164,
          "postDate": "2020-12-23T17:52:23.783Z",
          "content": "<p>Wow! Single LSTM scoring .797 is mind blowing!</p>",
          "rawMarkdown": "Wow! Single LSTM scoring .797 is mind blowing!"
        },
        {
          "id": 1125306,
          "postDate": "2020-12-24T16:00:37.480Z",
          "content": "<p>Sorry for late reply.</p>\n<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I'm using the same sampling method used in many public notebooks.<br>\n<a href=\"https://www.kaggle.com/mchahhou\" target=\"_blank\">@mchahhou</a> I'm using completely same features as the paper.</p>",
          "rawMarkdown": "Sorry for late reply.\n\n@adityaecdrid I'm using the same sampling method used in many public notebooks.\n@mchahhou I'm using completely same features as the paper.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1059325,
      "postDate": "2020-10-24T23:42:47.820Z",
      "content": "<p>Last LB score increases:</p>\n<p>Local: <strong>0.7609</strong> LB: <strong>0.759</strong><br>\nLocal: <strong>0.7757</strong> LB: <strong>0.772</strong><br>\nLocal: <strong>0.7817</strong> LB: <strong>0.780</strong><br>\nLocal: <strong>0.7866</strong> LB: <strong>0.786</strong></p>\n<p>I'm using 10m rows to train and 2.5m to validate per fold.</p>",
      "rawMarkdown": "Last LB score increases:\n\nLocal: **0.7609** LB: **0.759**\nLocal: **0.7757** LB: **0.772**\nLocal: **0.7817** LB: **0.780**\nLocal: **0.7866** LB: **0.786**\n\nI'm using 10m rows to train and 2.5m to validate per fold.",
      "votes": 19,
      "replies": [
        {
          "id": 1072868,
          "postDate": "2020-11-08T19:28:41.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/mingpan07\" target=\"_blank\">@mingpan07</a> it would be great if you can share your approach how did you get this strong CV/LB correlation ?</p>",
          "rawMarkdown": "@mingpan07 it would be great if you can share your approach how did you get this strong CV/LB correlation ?"
        }
      ]
    },
    {
      "id": 1058744,
      "postDate": "2020-10-24T07:29:02.003Z",
      "content": "<p>Regarding CV/LB scores, my last three jumps on the LB were the following:</p>\n<p>CV: <strong>0.725487</strong>, LB: <strong>0.761</strong><br>\nCV: <strong>0.740906</strong>, LB: <strong>0.767</strong><br>\nCV: <strong>0.747638</strong>, LB: <strong>0.772</strong></p>\n<p>My approach to validation is pretty simple, just the following to create the training/validation set. I use the remaining 90+% of the data for feature aggregates and the like. Even though the CV score isn't too similar to the LB, an increase in CV typically comes with an increase in LB:</p>\n<pre><code>print(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n</code></pre>\n<p>Update:</p>\n<p>CV: <strong>0.790365</strong>, LB: <strong>0.787</strong></p>\n<p>I switched up my validation strategy to instead use the one described in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a>. Seems to be relatively accurate from what I've noticed throughout the thread, typically &lt;= 0.003 of a difference beween CV/LB. This score is calculated using 10M rows for training, 2.5M rows for validation. I'm not using any lecture features yet in the implementation.</p>\n<p>Update #2:</p>\n<p>CV: <strong>0.793102</strong>, LB: <strong>0.789</strong></p>\n<p>Added a few more features, changed the framework of my code throughout to make it much easier to add/test new features. Still using only 10M rows, trying to see if I can hit ~0.795+ before I really look into training with the entire training set, given the estimated 0.01 boost from other people.</p>",
      "rawMarkdown": "Regarding CV/LB scores, my last three jumps on the LB were the following:\n\nCV: **0.725487**, LB: **0.761**\nCV: **0.740906**, LB: **0.767**\nCV: **0.747638**, LB: **0.772**\n\nMy approach to validation is pretty simple, just the following to create the training/validation set. I use the remaining 90+% of the data for feature aggregates and the like. Even though the CV score isn't too similar to the LB, an increase in CV typically comes with an increase in LB:\n\n```\nprint(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n```\n\nUpdate:\n\nCV: **0.790365**, LB: **0.787**\n\nI switched up my validation strategy to instead use the one described in https://www.kaggle.com/its7171/cv-strategy. Seems to be relatively accurate from what I've noticed throughout the thread, typically <= 0.003 of a difference beween CV/LB. This score is calculated using 10M rows for training, 2.5M rows for validation. I'm not using any lecture features yet in the implementation.\n\nUpdate #2:\n\nCV: **0.793102**, LB: **0.789**\n\nAdded a few more features, changed the framework of my code throughout to make it much easier to add/test new features. Still using only 10M rows, trying to see if I can hit ~0.795+ before I really look into training with the entire training set, given the estimated 0.01 boost from other people.",
      "votes": 13
    },
    {
      "id": 1058962,
      "postDate": "2020-10-24T13:00:36.780Z",
      "content": "<p>Local Validation: <strong>0.770</strong>, LB: <strong>0.766</strong>.<br>\nUpdate: Local <strong>0.777</strong>, LB: <strong>0.772</strong>. </p>\n<p>That's with training on ~4% of the data, and validating on ~2%. My validation scheme is set up to mirror the way the test predictions need to be made, so I'm fairly confident in it. I think a bit of a gap is normal especially given the time element.</p>\n<p>I strongly agree with what Alex said below about prioritizing finding the right features before using more of the data. 2% is already millions of samples and likely large enough to give reliable model feedback IMO (and even close to the same order as the actual test set). The key thing early in a competition is to be able to iterate on feature/model ideas quickly, and I'm pretty sure training time scales above linear with # of training rows for the typical gradient boosting model.</p>\n<p>Edit: see update </p>",
      "rawMarkdown": "Local Validation: **0.770**, LB: **0.766**.\nUpdate: Local **0.777**, LB: **0.772**. \n\nThat's with training on ~4% of the data, and validating on ~2%. My validation scheme is set up to mirror the way the test predictions need to be made, so I'm fairly confident in it. I think a bit of a gap is normal especially given the time element.\n\nI strongly agree with what Alex said below about prioritizing finding the right features before using more of the data. 2% is already millions of samples and likely large enough to give reliable model feedback IMO (and even close to the same order as the actual test set). The key thing early in a competition is to be able to iterate on feature/model ideas quickly, and I'm pretty sure training time scales above linear with # of training rows for the typical gradient boosting model.\n\nEdit: see update ",
      "votes": 11,
      "replies": [
        {
          "id": 1058966,
          "postDate": "2020-10-24T13:11:33.600Z",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Thanks for the info<br>\nAre you are now generating features for the 4% of data or for the whole data and train it with 4% of the data.</p>",
          "rawMarkdown": "@aquatic Thanks for the info\nAre you are now generating features for the 4% of data or for the whole data and train it with 4% of the data.",
          "votes": 1
        },
        {
          "id": 1058969,
          "postDate": "2020-10-24T13:17:45.893Z",
          "content": "<p>You definitely need to use all the training data for features to handle test well, so I do that. In particular, question level features come from all the data. One distinction though - my 6% train/val is a self-contained set of users, so I only really need that 6% of full user features to train/validate. For the rest of the data I have summarized user features.</p>",
          "rawMarkdown": "You definitely need to use all the training data for features to handle test well, so I do that. In particular, question level features come from all the data. One distinction though - my 6% train/val is a self-contained set of users, so I only really need that 6% of full user features to train/validate. For the rest of the data I have summarized user features.",
          "votes": 2
        },
        {
          "id": 1058975,
          "postDate": "2020-10-24T13:32:53.970Z",
          "content": "<p>Would you mind to share how to handle the features for unseen user during test time? <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>  <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> ? Or point me to a notebook that handle this?</p>",
          "rawMarkdown": "Would you mind to share how to handle the features for unseen user during test time? @spacelx  @aquatic ? Or point me to a notebook that handle this?"
        },
        {
          "id": 1058986,
          "postDate": "2020-10-24T13:50:28.873Z",
          "content": "<p>I essentially have a table of users where I keep some useful attributes for each of them… it starts out empty and gets filled by processing train and test data, simply adding a new row/entry for each previously unseen user and updating information as needed for already existing users.</p>",
          "rawMarkdown": "I essentially have a table of users where I keep some useful attributes for each of them... it starts out empty and gets filled by processing train and test data, simply adding a new row/entry for each previously unseen user and updating information as needed for already existing users.",
          "votes": 8
        },
        {
          "id": 1059004,
          "postDate": "2020-10-24T14:11:39.940Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I actually have prepared a notebook like that, but unfortunately I don't feel comfortable sharing it yet because it scores fairly highly (top 50ish). I haven't seen any public examples as of yet. I'm hoping to share a little later on, at a point when it wouldn't disrupt the leaderboard as much.</p>\n<p>My method is similar to what Alex describes, where a table gets updated on a running basis as the test data comes in. </p>",
          "rawMarkdown": "@yihdarshieh I actually have prepared a notebook like that, but unfortunately I don't feel comfortable sharing it yet because it scores fairly highly (top 50ish). I haven't seen any public examples as of yet. I'm hoping to share a little later on, at a point when it wouldn't disrupt the leaderboard as much.\n\nMy method is similar to what Alex describes, where a table gets updated on a running basis as the test data comes in. ",
          "votes": 5
        },
        {
          "id": 1059011,
          "postDate": "2020-10-24T14:22:17.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> No problem. I am just curious how this is done in FE. As I mentioned, I focus on sequential model, which should just use sequences as input, so there is no actual problem for unseen users. Of course, for unseen users, there is no history to use for the first prediction during test time - but the model, during training, will learn the distribution of targets in this case anyway. (You can think it as predicting the first word in a translation problem).</p>\n<p>Being said so, you could try to interoperate features among all users, like <code>what's the ratio of answer correctly a question if this question is the first question in a user interaction sequences.</code></p>",
          "rawMarkdown": "@aquatic No problem. I am just curious how this is done in FE. As I mentioned, I focus on sequential model, which should just use sequences as input, so there is no actual problem for unseen users. Of course, for unseen users, there is no history to use for the first prediction during test time - but the model, during training, will learn the distribution of targets in this case anyway. (You can think it as predicting the first word in a translation problem).\n\nBeing said so, you could try to interoperate features among all users, like `what's the ratio of answer correctly a question if this question is the first question in a user interaction sequences.`",
          "votes": 2
        },
        {
          "id": 1059193,
          "postDate": "2020-10-24T18:56:32.740Z",
          "content": "<blockquote>\n  <p>Features are entirely non-future-leaking</p>\n</blockquote>\n<p>Little confused here, let's say a user which you saw in train is also there in your val, so shall you use training stats and update that or use fresh stats? In case we don't use fresh, then we are leaking the user's past performance in val data, right? Am i missing something here? Ty!</p>\n<blockquote>\n  <p>Would you mind to share how to handle the features for unseen user during test time? Or point me to a notebook that handle this?</p>\n</blockquote>\n<p>One way to do this is what's already mentioned in this thread by Alex and Joe, we can index the user_stats_df's with the user_id (as the uid) for e.g.</p>",
          "rawMarkdown": ">Features are entirely non-future-leaking\n\nLittle confused here, let's say a user which you saw in train is also there in your val, so shall you use training stats and update that or use fresh stats? In case we don't use fresh, then we are leaking the user's past performance in val data, right? Am i missing something here? Ty!\n\n>Would you mind to share how to handle the features for unseen user during test time? Or point me to a notebook that handle this?\n\nOne way to do this is what's already mentioned in this thread by Alex and Joe, we can index the user_stats_df's with the user_id (as the uid) for e.g."
        },
        {
          "id": 1059198,
          "postDate": "2020-10-24T19:10:41.137Z",
          "content": "<blockquote>\n  <p>In case we don't use fresh, then we are leaking the user's past performance in val data, right?</p>\n</blockquote>\n<p>Using past performance info to predict the future is not a leak. Leaks refer to cases where the model has access to information that would not actually be available at prediction time, like future performance. In this problem, when we process and predict on the test data, we have access to past performance for users that aren't brand new. So long as the features you use are derived entirely from what the user has done in the past, you will be leak-free.</p>\n<p>The bigger picture way of thinking about this is that it's not really about the idea of an \"information leak\", but more a question of what's known at training time vs. what's known at prediction time. In most cases those two should mirror each other as much as possible for optimal results, both to prevent overfitting and underfitting. In particular, if you avoid using past information you're not protecting yourself from the harm of a leak, but instead denying your model relevant information that should almost certainly improve its performance.</p>\n<p>Another extension of this idea is that in a kaggle competition setting, it can be correct to use unrealistic, future-leaking features (talkingdata click prediction from a few years ago is a good example), so long as all parts of the timeline are accessible. Luckily, this competition gives you a much more realistic paradigm via the API submission process.</p>",
          "rawMarkdown": "> In case we don't use fresh, then we are leaking the user's past performance in val data, right?\n\nUsing past performance info to predict the future is not a leak. Leaks refer to cases where the model has access to information that would not actually be available at prediction time, like future performance. In this problem, when we process and predict on the test data, we have access to past performance for users that aren't brand new. So long as the features you use are derived entirely from what the user has done in the past, you will be leak-free.\n\nThe bigger picture way of thinking about this is that it's not really about the idea of an \"information leak\", but more a question of what's known at training time vs. what's known at prediction time. In most cases those two should mirror each other as much as possible for optimal results, both to prevent overfitting and underfitting. In particular, if you avoid using past information you're not protecting yourself from the harm of a leak, but instead denying your model relevant information that should almost certainly improve its performance.\n\nAnother extension of this idea is that in a kaggle competition setting, it can be correct to use unrealistic, future-leaking features (talkingdata click prediction from a few years ago is a good example), so long as all parts of the timeline are accessible. Luckily, this competition gives you a much more realistic paradigm via the API submission process.\n\n",
          "votes": 7
        },
        {
          "id": 1059205,
          "postDate": "2020-10-24T19:15:07.287Z",
          "content": "<p>For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -&gt; mean correctness etc.). The features in each row then only depend on past information.</p>\n<p>So I end up having a very long list of some features (historic mean user correctness, mean question correctness, etc.) and the information whether this question was answered correctly and that's what I'm training on. User IDs are dropped so I couldn't even tell anymore which user is which. The goal is to find a general model that predicts the chance of a random user getting the question right based on that random user's past performance, pretty much.</p>\n<p>So of course the user's past performance will \"leak\" into the validation data if you want to call it that, but then which features would we even predict on if we didn't want to take into account past performance?</p>\n<p>My validation is a random split, so it might well be (and will undoubtedly happen) that I train on the 500th question answered by a user, using knowledge of the outcome of all their previous questions, and validate the result using, amongst other data, the 10th question answered by this user given features calculated from the outcome of their previous 9 questions.</p>\n<p>Does this make sense? I'm not sure whether that answers your question.</p>",
          "rawMarkdown": "For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -> mean correctness etc.). The features in each row then only depend on past information.\n\nSo I end up having a very long list of some features (historic mean user correctness, mean question correctness, etc.) and the information whether this question was answered correctly and that's what I'm training on. User IDs are dropped so I couldn't even tell anymore which user is which. The goal is to find a general model that predicts the chance of a random user getting the question right based on that random user's past performance, pretty much.\n\nSo of course the user's past performance will \"leak\" into the validation data if you want to call it that, but then which features would we even predict on if we didn't want to take into account past performance?\n\nMy validation is a random split, so it might well be (and will undoubtedly happen) that I train on the 500th question answered by a user, using knowledge of the outcome of all their previous questions, and validate the result using, amongst other data, the 10th question answered by this user given features calculated from the outcome of their previous 9 questions.\n\nDoes this make sense? I'm not sure whether that answers your question.",
          "votes": 14
        },
        {
          "id": 1059400,
          "postDate": "2020-10-25T03:39:37.357Z",
          "content": "<p>Thanks a lot Joe and Alex!  It's much clear now!👍</p>",
          "rawMarkdown": "Thanks a lot Joe and Alex!  It's much clear now!👍"
        },
        {
          "id": 1059911,
          "postDate": "2020-10-25T15:24:17.317Z",
          "content": "<blockquote>\n  <p>For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -&gt; mean correctness etc.). The features in each row then only depend on past information.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> How are you able to do all this for 393,656 users without surpassing the 9h time limit? You only have 0.08 seconds available per user, right?</p>",
          "rawMarkdown": "> For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -> mean correctness etc.). The features in each row then only depend on past information.\n\n@spacelx How are you able to do all this for 393,656 users without surpassing the 9h time limit? You only have 0.08 seconds available per user, right?",
          "votes": 1
        },
        {
          "id": 1059937,
          "postDate": "2020-10-25T15:51:21.710Z",
          "content": "<p>My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.</p>\n<p>I don't process the data literally row by row, just row by row <em>per user</em>. In practical terms this means I keep a list of test set rows for each user, sorted by timestamp, and batch the processing by that. I.e., first batch is the first entry of each user, second batch the second entry of each user and so on.</p>\n<p>The kernel split is done on a user-basis, so each kernel handles about 80k users. So the first batch in the preprocessing loop will be some 80k rows (takes some 5 secs maybe), and the batches get smaller, in a somewhat exponential fashion. The most active users have some &gt;15k rows (quite impressive dedication of them actually) so by the end I'm forced to go row by row, but that's not a big issue either.</p>",
          "rawMarkdown": "My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.\n\nI don't process the data literally row by row, just row by row *per user*. In practical terms this means I keep a list of test set rows for each user, sorted by timestamp, and batch the processing by that. I.e., first batch is the first entry of each user, second batch the second entry of each user and so on.\n\nThe kernel split is done on a user-basis, so each kernel handles about 80k users. So the first batch in the preprocessing loop will be some 80k rows (takes some 5 secs maybe), and the batches get smaller, in a somewhat exponential fashion. The most active users have some >15k rows (quite impressive dedication of them actually) so by the end I'm forced to go row by row, but that's not a big issue either.",
          "votes": 4,
          "replies": [
            {
              "id": 1060026,
              "postDate": "2020-10-25T17:32:37.023Z",
              "content": "<blockquote>\n  <p>My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> During test time, how do you get to load these 20GB of features into memory and use them for the prediction? No OOM issue?</p>",
              "rawMarkdown": "> My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.\n> \n\n@spacelx During test time, how do you get to load these 20GB of features into memory and use them for the prediction? No OOM issue?"
            },
            {
              "id": 1060047,
              "postDate": "2020-10-25T18:03:54.940Z",
              "content": "<p>Currently I'm simply not using all the data to train my model, just a fraction of it.</p>",
              "rawMarkdown": "Currently I'm simply not using all the data to train my model, just a fraction of it."
            },
            {
              "id": 1060052,
              "postDate": "2020-10-25T18:09:59.013Z",
              "content": "<p>I mean the prediction for submission. During inference, you don't need to use the features that are prepared for training? </p>",
              "rawMarkdown": "I mean the prediction for submission. During inference, you don't need to use the features that are prepared for training? "
            },
            {
              "id": 1060102,
              "postDate": "2020-10-25T19:18:29.130Z",
              "content": "<p>Once the model is trained the train features aren't needed anymore, so in the inference kernel I just load the pre-trained model. Test data needs to be preprocessed though before sticking it through the model.</p>",
              "rawMarkdown": "Once the model is trained the train features aren't needed anymore, so in the inference kernel I just load the pre-trained model. Test data needs to be preprocessed though before sticking it through the model.",
              "votes": 3
            }
          ]
        },
        {
          "id": 1059943,
          "postDate": "2020-10-25T15:58:03.757Z",
          "content": "<blockquote>\n  <p>The most active users have some &gt;15k rows (quite impressive dedication of them actually)</p>\n</blockquote>\n<p>They certainly have a bright future 😁</p>",
          "rawMarkdown": "> The most active users have some >15k rows (quite impressive dedication of them actually)\n\nThey certainly have a bright future 😁",
          "votes": 3
        },
        {
          "id": 1059967,
          "postDate": "2020-10-25T16:19:29.773Z",
          "content": "<p>20 GB+ in features really means a hell lot of features 🙏🙏! Nice idea to split the work load's in multiple kernels for faster turn-around time as well!</p>",
          "rawMarkdown": "20 GB+ in features really means a hell lot of features 🙏🙏! Nice idea to split the work load's in multiple kernels for faster turn-around time as well!"
        },
        {
          "id": 1059970,
          "postDate": "2020-10-25T16:23:33.240Z",
          "content": "<p>Meh, it's not that many actually… I should probably just switch to saving in a more efficient format than floats in plain csv…</p>",
          "rawMarkdown": "Meh, it's not that many actually... I should probably just switch to saving in a more efficient format than floats in plain csv..."
        },
        {
          "id": 1059991,
          "postDate": "2020-10-25T16:49:36.100Z",
          "content": "<p>Why not store them as string's when saving them and while loading back, load as float point numbers?</p>",
          "rawMarkdown": "Why not store them as string's when saving them and while loading back, load as float point numbers?"
        },
        {
          "id": 1059994,
          "postDate": "2020-10-25T16:55:47.160Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 1059999,
          "postDate": "2020-10-25T17:07:46.900Z",
          "content": "<blockquote>\n  <p>Why not store them as string's when saving them and while loading back, load as float point numbers?</p>\n</blockquote>\n<p>In a .csv it saves the string representation of the floats (which can be pretty long), but it would be more space-efficient to save the floats themselves in binary format. I should really go and save in feather or something 😛</p>",
          "rawMarkdown": "> Why not store them as string's when saving them and while loading back, load as float point numbers?\n\nIn a .csv it saves the string representation of the floats (which can be pretty long), but it would be more space-efficient to save the floats themselves in binary format. I should really go and save in feather or something 😛",
          "votes": 1
        },
        {
          "id": 1060004,
          "postDate": "2020-10-25T17:09:53.990Z",
          "content": "<p><code>They certainly have a bright future 😁</code></p>\n<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> , if they come to Kaggle to compete, I won't be able to beat them 😂</p>",
          "rawMarkdown": "`They certainly have a bright future 😁`\n\n@rohanrao , if they come to Kaggle to compete, I won't be able to beat them 😂",
          "votes": 2
        },
        {
          "id": 1060024,
          "postDate": "2020-10-25T17:29:06.723Z",
          "content": "<p>Yep, we can go for binary format for sure as numpy will help us there (never did it myself until today;) or array in Py but it's not used much in day-to-day activities. But do we need more than 4-5 decimals of precision?</p>",
          "rawMarkdown": "Yep, we can go for binary format for sure as numpy will help us there (never did it myself until today;) or array in Py but it's not used much in day-to-day activities. But do we need more than 4-5 decimals of precision?"
        },
        {
          "id": 1060046,
          "postDate": "2020-10-25T18:00:44.913Z",
          "content": "<blockquote>\n  <p>But do we need more than 4-5 decimals of precision?</p>\n</blockquote>\n<p>That's what I thought too, but a run with more decimals gave me a better score, and not a negligible increase at that. Or maybe I added a feature in between these runs and forgot about it. <br>\nAnyway, better safe than sorry is my current approach…</p>\n<p>Pandas has plenty of saving options so binary is not a problem at all from a practical point of view. Upside of .csv is that in Kaggle's dataset viewer you can directly see column histograms and check whether everything went right with scaling and such.</p>",
          "rawMarkdown": "> But do we need more than 4-5 decimals of precision?\n\nThat's what I thought too, but a run with more decimals gave me a better score, and not a negligible increase at that. Or maybe I added a feature in between these runs and forgot about it. \nAnyway, better safe than sorry is my current approach...\n\nPandas has plenty of saving options so binary is not a problem at all from a practical point of view. Upside of .csv is that in Kaggle's dataset viewer you can directly see column histograms and check whether everything went right with scaling and such.",
          "votes": 1
        },
        {
          "id": 1060092,
          "postDate": "2020-10-25T18:55:13.240Z",
          "content": "<blockquote>\n  <p>The most active users have some &gt;15k rows (quite impressive dedication of them actually)</p>\n</blockquote>\n<p>How do we know that they are not bots who were scrapping riddi? (out of topic, just kidding)</p>",
          "rawMarkdown": "> The most active users have some >15k rows (quite impressive dedication of them actually)\n\nHow do we know that they are not bots who were scrapping riddi? (out of topic, just kidding)"
        },
        {
          "id": 1060926,
          "postDate": "2020-10-26T16:16:15.073Z",
          "content": "<p>Is it possible to do ensembling and stacking in this challenge<br>\nbecause test data can only be called once right??</p>",
          "rawMarkdown": "Is it possible to do ensembling and stacking in this challenge\nbecause test data can only be called once right??"
        },
        {
          "id": 1060995,
          "postDate": "2020-10-26T17:21:03.983Z",
          "content": "<p>I would simply suggest, do something simpler first as that itself is quite complicated and challenging many times. And then one can build on top that!</p>",
          "rawMarkdown": "I would simply suggest, do something simpler first as that itself is quite complicated and challenging many times. And then one can build on top that!"
        },
        {
          "id": 1105473,
          "postDate": "2020-12-07T22:58:00.777Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Of course they have a bright future. If one of them is already here on kaggle, she just check the competion: <br>\n\"hmm, I created the training data for years, let's win this :P\"</p>",
          "rawMarkdown": "@yihdarshieh @rohanrao Of course they have a bright future. If one of them is already here on kaggle, she just check the competion: \n\"hmm, I created the training data for years, let's win this :P\""
        }
      ]
    },
    {
      "id": 1093890,
      "postDate": "2020-11-28T05:24:03.260Z",
      "content": "<p>single NN, inference for ~3h<br>\nI use <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> strategy.<br>\nCV:  0.772, LB: 0.776</p>",
      "rawMarkdown": "single NN, inference for ~3h\nI use [this](https://www.kaggle.com/its7171/cv-strategy) strategy.\nCV:  0.772, LB: 0.776",
      "votes": 7,
      "replies": [
        {
          "id": 1093913,
          "postDate": "2020-11-28T05:49:23.743Z",
          "content": "<p>Nice CV! Can you write few words about your NN! Is it transformer's based? Ty!</p>",
          "rawMarkdown": "Nice CV! Can you write few words about your NN! Is it transformer's based? Ty!"
        },
        {
          "id": 1125940,
          "postDate": "2020-12-25T07:32:00.760Z",
          "content": "<p>NN based, implemented with tensorflow, inference for 7h<br>\n<a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">same cv strategy</a><br>\nCV: 0.803, LB: 0.799</p>",
          "rawMarkdown": "NN based, implemented with tensorflow, inference for 7h\n[same cv strategy](https://www.kaggle.com/its7171/cv-strategy)\nCV: 0.803, LB: 0.799",
          "votes": 1
        },
        {
          "id": 1125986,
          "postDate": "2020-12-25T08:12:04.197Z",
          "content": "<p>Awesome Score! Good Luck!</p>",
          "rawMarkdown": "Awesome Score! Good Luck!"
        },
        {
          "id": 1130563,
          "postDate": "2020-12-29T05:59:50.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> we have strong pipeline with single saint score of 78.6 we work on saint plus. let me know your thought about teamup</p>",
          "rawMarkdown": "@nadare we have strong pipeline with single saint score of 78.6 we work on saint plus. let me know your thought about teamup"
        }
      ]
    },
    {
      "id": 1064170,
      "postDate": "2020-10-29T19:54:33.937Z",
      "content": "<p>So far CV-LB correlation looks very stable - it's constantly +0.003 on LB compared to CV for all my submissions in the last 2 weeks.</p>",
      "rawMarkdown": "So far CV-LB correlation looks very stable - it's constantly +0.003 on LB compared to CV for all my submissions in the last 2 weeks.",
      "votes": 8
    },
    {
      "id": 1071447,
      "postDate": "2020-11-06T22:30:38.373Z",
      "content": "<p>CV: .792 <br>\nLB: .791</p>",
      "rawMarkdown": "CV: .792 \nLB: .791",
      "votes": 7,
      "replies": [
        {
          "id": 1073958,
          "postDate": "2020-11-10T04:27:57.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Could you please explain the validation strategy?😊</p>",
          "rawMarkdown": "@mamasinkgs Could you please explain the validation strategy?😊",
          "votes": 1
        },
        {
          "id": 1074479,
          "postDate": "2020-11-10T17:37:32.283Z",
          "content": "<p>Okay, I will explain the details in the winning solution.👌</p>",
          "rawMarkdown": "Okay, I will explain the details in the winning solution.👌",
          "votes": 5
        },
        {
          "id": 1101258,
          "postDate": "2020-12-03T18:54:21.157Z",
          "content": "<p>UPDATES:<br>\nCV: 0.8090<br>\nLB: 0.808X</p>",
          "rawMarkdown": "UPDATES:\nCV: 0.8090\nLB: 0.808X",
          "votes": 3
        },
        {
          "id": 1101446,
          "postDate": "2020-12-03T22:52:01.147Z",
          "content": "<p>stop plz! 😂</p>",
          "rawMarkdown": "stop plz! 😂",
          "votes": 6
        },
        {
          "id": 1102368,
          "postDate": "2020-12-04T21:09:07.627Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Did you achieve this score using LGBM?  I am asking because I am quite curious where the limit of GBDT model in this competition is.  Thanks in advance! </p>",
          "rawMarkdown": "@mamasinkgs Did you achieve this score using LGBM?  I am asking because I am quite curious where the limit of GBDT model in this competition is.  Thanks in advance! "
        },
        {
          "id": 1102597,
          "postDate": "2020-12-05T05:18:49.610Z",
          "content": "<p>It's a top secret! But, please remember nyanp says his LGBM is .801, and Nikola Bacic (.806 in LB) says he is using Transformer. From their comments, I think Transformer may perform a bit better, but LGBM is competitive.</p>",
          "rawMarkdown": "It's a top secret! But, please remember nyanp says his LGBM is .801, and Nikola Bacic (.806 in LB) says he is using Transformer. From their comments, I think Transformer may perform a bit better, but LGBM is competitive.",
          "votes": 7
        },
        {
          "id": 1103287,
          "postDate": "2020-12-05T19:39:31.283Z",
          "content": "<p>LGBM specifically, or GBM generally, including xgboost?</p>",
          "rawMarkdown": "LGBM specifically, or GBM generally, including xgboost?"
        },
        {
          "id": 1103955,
          "postDate": "2020-12-06T13:00:45.217Z",
          "content": "<p>I think 100M rows are too computationally expensive for XGBoost/CatBoost.</p>",
          "rawMarkdown": "I think 100M rows are too computationally expensive for XGBoost/CatBoost."
        },
        {
          "id": 1111106,
          "postDate": "2020-12-13T12:12:52.303Z",
          "content": "<p>UPDATES: <br>\nCV: 0.8125<br>\nLB: 0.811x</p>\n<p>seems I'm overfitting :(</p>",
          "rawMarkdown": "UPDATES: \nCV: 0.8125\nLB: 0.811x\n\nseems I'm overfitting :(",
          "votes": 4
        },
        {
          "id": 1111116,
          "postDate": "2020-12-13T12:33:42.893Z",
          "content": "<p>poor you :)</p>",
          "rawMarkdown": "poor you :)",
          "votes": 1
        },
        {
          "id": 1118037,
          "postDate": "2020-12-18T17:13:51.960Z",
          "content": "<p>thus, nn might win in such case. I strongly believe.</p>",
          "rawMarkdown": "thus, nn might win in such case. I strongly believe."
        }
      ]
    },
    {
      "id": 1117386,
      "postDate": "2020-12-18T02:11:25.650Z",
      "content": "<p>CV: 0.802, LB: 0.802</p>\n<p>single lightgbm model with 100+ features, nothing found in lectures😓. Anyone boost score from lectures features?</p>",
      "rawMarkdown": "CV: 0.802, LB: 0.802\n\nsingle lightgbm model with 100+ features, nothing found in lectures😓. Anyone boost score from lectures features?",
      "votes": 6,
      "replies": [
        {
          "id": 1118049,
          "postDate": "2020-12-18T17:21:25.227Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> you're not alone, we're not using lectures, every attempt with lectures failed for us. Interested too with feeback on lectures features.</p>",
          "rawMarkdown": "Hey @juzqyxs you're not alone, we're not using lectures, every attempt with lectures failed for us. Interested too with feeback on lectures features.",
          "votes": 1
        },
        {
          "id": 1129008,
          "postDate": "2020-12-28T01:53:30.073Z",
          "content": "<p>update, single lgb CV:0.806, LB:0.806, still without lectures features.</p>",
          "rawMarkdown": "update, single lgb CV:0.806, LB:0.806, still without lectures features.",
          "votes": 1
        },
        {
          "id": 1130744,
          "postDate": "2020-12-29T09:38:57.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Do you guys mean lectures don't improve your validation score or LB?</p>",
          "rawMarkdown": "@juzqyxs @mpware Do you guys mean lectures don't improve your validation score or LB?",
          "votes": 1
        },
        {
          "id": 1130784,
          "postDate": "2020-12-29T09:55:33.813Z",
          "content": "<p>Your guess is right. So lectures improve your score?</p>",
          "rawMarkdown": "Your guess is right. So lectures improve your score?"
        },
        {
          "id": 1130788,
          "postDate": "2020-12-29T09:57:05.403Z",
          "content": "<p>It improves CV (a little bit) but not LB. I will give another try this week, I may have a bug somewhere.<br>\nBut lectures are only 2% of data so it might not be important.</p>",
          "rawMarkdown": "It improves CV (a little bit) but not LB. I will give another try this week, I may have a bug somewhere.\nBut lectures are only 2% of data so it might not be important."
        },
        {
          "id": 1130806,
          "postDate": "2020-12-29T10:12:54.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> I asked an OR question. So which one?:) </p>\n<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> 2% of data is actually important since we all compete for 3rd decimal on metric. I and several other people reported that it improves CV but not LB. Maybe we all have the same inference bug.</p>",
          "rawMarkdown": "@juzqyxs I asked an OR question. So which one?:) \n\n@mpware 2% of data is actually important since we all compete for 3rd decimal on metric. I and several other people reported that it improves CV but not LB. Maybe we all have the same inference bug.",
          "votes": 1
        },
        {
          "id": 1130864,
          "postDate": "2020-12-29T10:59:13.777Z",
          "content": "<p>Because lectures features did not improve our local cv, we did not submit. If you generated features from lectures after user's answer records, it might improve your local cv, but the online lectures info maybe in next group, this could be the reason that LB not improve.</p>",
          "rawMarkdown": "Because lectures features did not improve our local cv, we did not submit. If you generated features from lectures after user's answer records, it might improve your local cv, but the online lectures info maybe in next group, this could be the reason that LB not improve."
        }
      ]
    },
    {
      "id": 1122135,
      "postDate": "2020-12-22T08:09:12.127Z",
      "content": "<p>CV: 0.808, LB: 0.805</p>\n<p>Looking at everyone has better CV-LB gap, I either have a bug in inference or something else is wrong.</p>",
      "rawMarkdown": "CV: 0.808, LB: 0.805\n\nLooking at everyone has better CV-LB gap, I either have a bug in inference or something else is wrong.",
      "votes": 3,
      "replies": [
        {
          "id": 1122149,
          "postDate": "2020-12-22T08:26:17.310Z",
          "content": "<p>Can I form a team with you to study.</p>",
          "rawMarkdown": "Can I form a team with you to study."
        },
        {
          "id": 1122263,
          "postDate": "2020-12-22T10:08:54.930Z",
          "content": "<p>No.</p>\n<p></p>",
          "rawMarkdown": "No.\n\n~~Your message must have at least 10 characters.~~",
          "votes": 1
        },
        {
          "id": 1122272,
          "postDate": "2020-12-22T10:18:49.190Z",
          "content": "<p>During inference, you can have multiple questions for the same user within the same batch. If you've some kind of expanding features then are you updating those features in that case or do you wait for the next batch? I'm facing this issue with simular gap on my local simulator.</p>",
          "rawMarkdown": "During inference, you can have multiple questions for the same user within the same batch. If you've some kind of expanding features then are you updating those features in that case or do you wait for the next batch? I'm facing this issue with simular gap on my local simulator.",
          "votes": 1
        },
        {
          "id": 1122279,
          "postDate": "2020-12-22T10:27:48.510Z",
          "content": "<p>It depends if multiple questions have different timestamps in the same batch. I assume they have the same timestamp?</p>",
          "rawMarkdown": "It depends if multiple questions have different timestamps in the same batch. I assume they have the same timestamp?",
          "votes": 1
        },
        {
          "id": 1122282,
          "postDate": "2020-12-22T10:29:51.690Z",
          "content": "<p>Yes, same timestamp, all within the same task_container_id.</p>",
          "rawMarkdown": "Yes, same timestamp, all within the same task_container_id.\n",
          "votes": 1
        },
        {
          "id": 1122286,
          "postDate": "2020-12-22T10:31:18.420Z",
          "content": "<p>I can confirm from my personal probe.</p>",
          "rawMarkdown": "I can confirm from my personal probe."
        },
        {
          "id": 1122564,
          "postDate": "2020-12-22T14:29:16.927Z",
          "content": "<p>update features ，and update dict next batch.may be userful </p>",
          "rawMarkdown": "update features ，and update dict next batch.may be userful ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1071389,
      "postDate": "2020-11-06T20:18:28.260Z",
      "content": "<p>Local: <strong>0.7602</strong>   LB: <strong>0.773</strong><br>\nLocal: <strong>0.7699</strong>   LB: <strong>0.782</strong></p>",
      "rawMarkdown": "Local: **0.7602**   LB: **0.773**\nLocal: **0.7699**   LB: **0.782**",
      "votes": 3
    },
    {
      "id": 1071175,
      "postDate": "2020-11-06T15:28:07.570Z",
      "content": "<p>Local:766<br>\nLB:785</p>",
      "rawMarkdown": "Local:766\nLB:785",
      "votes": 3
    },
    {
      "id": 1066774,
      "postDate": "2020-11-02T02:57:24.543Z",
      "content": "<p>Local:763,LB:780<br>\nUpdate: Local:771,LB:790</p>",
      "rawMarkdown": "Local:763,LB:780\nUpdate: Local:771,LB:790",
      "votes": 3
    },
    {
      "id": 1066232,
      "postDate": "2020-11-01T14:02:54.863Z",
      "content": "<p>CV:0.798<br>\nLB:0.771<br>\nValidating method: GroupKFold by user_id.</p>",
      "rawMarkdown": "CV:0.798\nLB:0.771\nValidating method: GroupKFold by user_id.",
      "votes": 3
    },
    {
      "id": 1058825,
      "postDate": "2020-10-24T10:33:37.097Z",
      "content": "<p>My last two submissions were</p>\n<p>Local CV: <strong>0.7637</strong>, LB: <strong>0.767</strong><br>\nLocal CV: <strong>0.7567</strong>, LB: <strong>0.760</strong></p>\n<p>So both times my LB was about 0.033 higher than my local CV score. Not sure why but I'll take it.</p>",
      "rawMarkdown": "My last two submissions were\n\nLocal CV: **0.7637**, LB: **0.767**\nLocal CV: **0.7567**, LB: **0.760**\n\nSo both times my LB was about 0.033 higher than my local CV score. Not sure why but I'll take it.",
      "votes": 3,
      "replies": [
        {
          "id": 1058828,
          "postDate": "2020-10-24T10:37:19.400Z",
          "content": "<p>As for approaches, I'm still working with only small random selections of the train data - my last model was a quick-build on about 10% or maybe even less of the training data, validated with some 2% of it. Features are entirely non-future-leaking so I can happily train and validate on whichever rows I feel like without having to pay attention to that.</p>",
          "rawMarkdown": "As for approaches, I'm still working with only small random selections of the train data - my last model was a quick-build on about 10% or maybe even less of the training data, validated with some 2% of it. Features are entirely non-future-leaking so I can happily train and validate on whichever rows I feel like without having to pay attention to that.",
          "votes": 2
        },
        {
          "id": 1058867,
          "postDate": "2020-10-24T11:28:05.680Z",
          "content": "<p>18th with 10% 👍</p>",
          "rawMarkdown": "18th with 10% 👍"
        },
        {
          "id": 1058874,
          "postDate": "2020-10-24T11:34:28.970Z",
          "content": "<p>Finding good features makes all the difference really. Sure I can get some increases using all the set (and I'll eventually do that), but finding good features is much more important in the beginning and will give you quite some improvements to work with.</p>",
          "rawMarkdown": "Finding good features makes all the difference really. Sure I can get some increases using all the set (and I'll eventually do that), but finding good features is much more important in the beginning and will give you quite some improvements to work with.",
          "votes": 4
        },
        {
          "id": 1058878,
          "postDate": "2020-10-24T11:36:35.203Z",
          "content": "<p>Yeah, I know. I focus on sequential models though - haven't done any FE yet.</p>",
          "rawMarkdown": "Yeah, I know. I focus on sequential models though - haven't done any FE yet."
        },
        {
          "id": 1058887,
          "postDate": "2020-10-24T11:42:39.193Z",
          "content": "<p>I guess 10 % to train the model, but FE is on the full data , right ? </p>",
          "rawMarkdown": "I guess 10 % to train the model, but FE is on the full data , right ? "
        },
        {
          "id": 1058899,
          "postDate": "2020-10-24T11:53:30.653Z",
          "content": "<p>Ok so in more detail, I get features for the entire train set, going for each user by timestep essentially to get non-future-leaked features (like mean user accuracy and such, will be very noisy for the first few questions of each user but getting better over time).<br>\nThis gives a me a full train set of features as well as a post-train status of some parameters which are used to calculate the features of the test set and updated with each test row (imagine, number of answers given by user x or so).</p>\n<p>So my model is currently only trained on like 10% of the train data, but when evaluating the test data with it I have the full cumulative history, so to say, of all previous events in order to calculate test features.</p>",
          "rawMarkdown": "Ok so in more detail, I get features for the entire train set, going for each user by timestep essentially to get non-future-leaked features (like mean user accuracy and such, will be very noisy for the first few questions of each user but getting better over time).\nThis gives a me a full train set of features as well as a post-train status of some parameters which are used to calculate the features of the test set and updated with each test row (imagine, number of answers given by user x or so).\n\nSo my model is currently only trained on like 10% of the train data, but when evaluating the test data with it I have the full cumulative history, so to say, of all previous events in order to calculate test features.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1082908,
      "postDate": "2020-11-18T11:15:57.953Z",
      "content": "<p>Local: 0.798 LB: not submitted (with time-series split)</p>\n<p>The road to 1st submission is a long way for me 😂</p>",
      "rawMarkdown": "Local: 0.798 LB: not submitted (with time-series split)\n\nThe road to 1st submission is a long way for me 😂",
      "votes": 4,
      "replies": [
        {
          "id": 1082922,
          "postDate": "2020-11-18T11:35:39.917Z",
          "content": "<p>STOP PLS!!!</p>",
          "rawMarkdown": "STOP PLS!!!",
          "votes": -8
        },
        {
          "id": 1101447,
          "postDate": "2020-12-03T22:53:18.930Z",
          "content": "<p>updated:</p>\n<p>CV: 0.803 LB: 0.801 with single LightGBM</p>",
          "rawMarkdown": "updated:\n\nCV: 0.803 LB: 0.801 with single LightGBM",
          "votes": 5
        },
        {
          "id": 1101478,
          "postDate": "2020-12-03T23:51:44.653Z",
          "content": "<p>Nice result - how much do you think is related to:</p>\n<ul>\n<li>training/validation split</li>\n<li>model parameters</li>\n<li>features</li>\n</ul>\n<p>Thank you</p>",
          "rawMarkdown": "Nice result - how much do you think is related to:\n- training/validation split\n- model parameters\n- features\n\nThank you"
        },
        {
          "id": 1101536,
          "postDate": "2020-12-04T02:03:27.743Z",
          "content": "<p>How many rows are you training with? I'm currently at a CV/LB of 0.793/0.789 with a single LGB and 10M rows, and trying to figure out how much further I can push it before upping the number of rows.</p>",
          "rawMarkdown": "How many rows are you training with? I'm currently at a CV/LB of 0.793/0.789 with a single LGB and 10M rows, and trying to figure out how much further I can push it before upping the number of rows.",
          "votes": 1
        },
        {
          "id": 1101543,
          "postDate": "2020-12-04T02:13:10.890Z",
          "content": "<p>amazing work</p>",
          "rawMarkdown": "amazing work"
        },
        {
          "id": 1101577,
          "postDate": "2020-12-04T03:25:19.270Z",
          "content": "<p>I always use full dataset for submission. In my case, using full datasets increases CV by +0.006 compared to 6M rows. All hyperparameters except for num_leaves and learning_rate (for speedup training) remain the default. I think feature engineering on the solid validation is much more important than model parameters in case of GBDT.</p>",
          "rawMarkdown": "I always use full dataset for submission. In my case, using full datasets increases CV by +0.006 compared to 6M rows. All hyperparameters except for num_leaves and learning_rate (for speedup training) remain the default. I think feature engineering on the solid validation is much more important than model parameters in case of GBDT.",
          "votes": 14
        },
        {
          "id": 1101580,
          "postDate": "2020-12-04T03:29:35.460Z",
          "content": "<p>Perfect Feature Engineering👍👍👍</p>",
          "rawMarkdown": "Perfect Feature Engineering👍👍👍",
          "votes": 1
        },
        {
          "id": 1102849,
          "postDate": "2020-12-05T11:48:14.457Z",
          "content": "<p>nyanp, did you really change <code>num_leaves</code> for full data training? or do you mean <code>nround</code>?</p>",
          "rawMarkdown": "nyanp, did you really change `num_leaves` for full data training? or do you mean `nround`?"
        },
        {
          "id": 1102861,
          "postDate": "2020-12-05T12:08:05.047Z",
          "content": "<p>I forgot about the nround, it's based on early stopping, so of course it's different for full rows and 6M rows. As for the other parameters, I adjusted the num_leaves and learning_rate in 6M rows and used the same parameters in the full data as well.</p>",
          "rawMarkdown": "I forgot about the nround, it's based on early stopping, so of course it's different for full rows and 6M rows. As for the other parameters, I adjusted the num_leaves and learning_rate in 6M rows and used the same parameters in the full data as well.",
          "votes": 4
        },
        {
          "id": 1102865,
          "postDate": "2020-12-05T12:12:55.537Z",
          "content": "<p>Ah ok thanks, I got it!</p>",
          "rawMarkdown": "Ah ok thanks, I got it!"
        }
      ]
    },
    {
      "id": 1071733,
      "postDate": "2020-11-07T10:47:21.740Z",
      "content": "<p>CV: 0.777<br>\nLB: 0.784</p>",
      "rawMarkdown": "CV: 0.777\nLB: 0.784",
      "votes": 4
    },
    {
      "id": 1065768,
      "postDate": "2020-10-31T18:08:31.307Z",
      "content": "<p>I have at the moment 0.7579 on validation and 0.763 on LB. I'm validating only against unseen users, and that's probably why it is quite higher on LB.</p>\n<p>EDIT: I'm using 5 M rows, 33% of them as validation.</p>",
      "rawMarkdown": "I have at the moment 0.7579 on validation and 0.763 on LB. I'm validating only against unseen users, and that's probably why it is quite higher on LB.\n\nEDIT: I'm using 5 M rows, 33% of them as validation.",
      "votes": 4
    },
    {
      "id": 1127999,
      "postDate": "2020-12-27T05:08:31.843Z",
      "content": "<p>Same as <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> cv strategy.<br>\nLocal: 0.8027, LB: 0.805x</p>",
      "rawMarkdown": "Same as [this](https://www.kaggle.com/its7171/cv-strategy) cv strategy.\nLocal: 0.8027, LB: 0.805x",
      "votes": 1
    },
    {
      "id": 1058729,
      "postDate": "2020-10-24T06:56:11.427Z",
      "content": "<p>My last 3 submissions (using holdout validation set):</p>\n<p>Local: <strong>0.756</strong> LB: <strong>0.758</strong><br>\nLocal: <strong>0.748</strong> LB: <strong>0.751</strong><br>\nLocal: <strong>0.741</strong> LB: <strong>0.742</strong></p>",
      "rawMarkdown": "My last 3 submissions (using holdout validation set):\n\nLocal: **0.756** LB: **0.758**\nLocal: **0.748** LB: **0.751**\nLocal: **0.741** LB: **0.742**",
      "votes": 2,
      "replies": [
        {
          "id": 1058735,
          "postDate": "2020-10-24T07:11:54.130Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1103563,
      "postDate": "2020-12-06T03:15:47.543Z",
      "content": "<p>CV: 0.778 - LB: 0.779</p>",
      "rawMarkdown": "CV: 0.778 - LB: 0.779",
      "votes": 1
    },
    {
      "id": 1090124,
      "postDate": "2020-11-25T05:11:11.037Z",
      "content": "<p>some last models:<br>\nCV 0.756//LB0.759<br>\nCV0.770//LB0.766<br>\nCV0.771//LB0.769<br>\nCV0.774//LB0.772<br>\nCV0.777//LB0.774<br>\nit Looks like the LB score is fluctuating around .003.<br>\nsingle  LGBM LB0.801 sounds so magic,need more features..💪💪💪</p>",
      "rawMarkdown": "some last models:\nCV 0.756//LB0.759\nCV0.770//LB0.766\nCV0.771//LB0.769\nCV0.774//LB0.772\nCV0.777//LB0.774\nit Looks like the LB score is fluctuating around .003.\nsingle  LGBM LB0.801 sounds so magic,need more features..💪💪💪",
      "votes": 1,
      "replies": [
        {
          "id": 1102687,
          "postDate": "2020-12-05T07:33:01.243Z",
          "content": "<p>but, seems you are lucky that you received a consistent improvement.</p>",
          "rawMarkdown": "but, seems you are lucky that you received a consistent improvement."
        },
        {
          "id": 1125319,
          "postDate": "2020-12-24T16:13:37.280Z",
          "content": "<p>update my features style ,now CV 0.781,LB0.781</p>",
          "rawMarkdown": "update my features style ,now CV 0.781,LB0.781"
        }
      ]
    },
    {
      "id": 1079477,
      "postDate": "2020-11-16T05:37:12.520Z",
      "content": "<p>Local: <strong>0.7638</strong> LB <strong>0.765</strong></p>\n<p>Using all training data and follow the cross-validation strategy in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>",
      "rawMarkdown": "Local: **0.7638** LB **0.765**\n\nUsing all training data and follow the cross-validation strategy in https://www.kaggle.com/its7171/cv-strategy",
      "votes": 1,
      "replies": [
        {
          "id": 1089909,
          "postDate": "2020-11-24T22:16:51.303Z",
          "content": "<p>I'm using the same cross-validation strategy, however I'm experiencing a gap in LB vs Local (0.757 vs 0.773). I noticed this gap exists only when I used cumulative user features. Do you experience the same ? Thanks!</p>",
          "rawMarkdown": "I'm using the same cross-validation strategy, however I'm experiencing a gap in LB vs Local (0.757 vs 0.773). I noticed this gap exists only when I used cumulative user features. Do you experience the same ? Thanks!"
        }
      ]
    },
    {
      "id": 1078828,
      "postDate": "2020-11-15T11:02:56.707Z",
      "content": "<p>Local CV: 0.740, LB: 0.727</p>",
      "rawMarkdown": "Local CV: 0.740, LB: 0.727",
      "votes": 1
    },
    {
      "id": 1075135,
      "postDate": "2020-11-11T12:23:35.723Z",
      "content": "<p>CV: 0.763 LB: 0.771 with simple split by user_id</p>",
      "rawMarkdown": "CV: 0.763 LB: 0.771 with simple split by user_id",
      "votes": 1,
      "replies": [
        {
          "id": 1089910,
          "postDate": "2020-11-24T22:17:27.473Z",
          "content": "<p>Can you tell me how did you split by user_id ? Thanks!</p>",
          "rawMarkdown": "Can you tell me how did you split by user_id ? Thanks!"
        },
        {
          "id": 1104145,
          "postDate": "2020-12-06T16:51:39.337Z",
          "content": "<p>Update:<br>\nvalidation: 0.8064 LB: 0.808x</p>\n<p>but at what cost…</p>",
          "rawMarkdown": "Update:\nvalidation: 0.8064 LB: 0.808x\n\nbut at what cost...",
          "votes": 2
        }
      ]
    },
    {
      "id": 1064367,
      "postDate": "2020-10-30T03:42:16.150Z",
      "content": "<p>Local: 0.79305<br>\nLB:     0.792</p>",
      "rawMarkdown": "Local: 0.79305\nLB:     0.792",
      "votes": 1
    },
    {
      "id": 1062828,
      "postDate": "2020-10-28T08:24:33.617Z",
      "content": "<p>Hi,<br>\nLocal : 0.733<br>\nLB: 0.758</p>\n<p>Pretty classic validation scheme:<br>\nMy training set contains about 10M rows, 3M for validation. Features are aggretated on the remaining rows</p>",
      "rawMarkdown": "Hi,\nLocal : 0.733\nLB: 0.758\n\nPretty classic validation scheme:\nMy training set contains about 10M rows, 3M for validation. Features are aggretated on the remaining rows",
      "votes": 1,
      "replies": [
        {
          "id": 1063029,
          "postDate": "2020-10-28T12:52:48.650Z",
          "content": "<p>The gap is similar to what <a href=\"https://www.kaggle.com/misfyre\" target=\"_blank\">@misfyre</a> reported.</p>",
          "rawMarkdown": "The gap is similar to what @misfyre reported."
        },
        {
          "id": 1063145,
          "postDate": "2020-10-28T14:55:56.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a> I'm impressed by the strong correlation between your local validation and public LB. Are you using something similar to <a href=\"https://www.kaggle.com/misfyre\" target=\"_blank\">@misfyre</a> ?</p>\n<p>I've built my hold-out set in the same way so it makes sense that there is a similar gap. However I cannot reprocude his jump from 0.725 to 0.747 😔</p>",
          "rawMarkdown": "@vopani I'm impressed by the strong correlation between your local validation and public LB. Are you using something similar to @misfyre ?\n\nI've built my hold-out set in the same way so it makes sense that there is a similar gap. However I cannot reprocude his jump from 0.725 to 0.747 😔"
        },
        {
          "id": 1063174,
          "postDate": "2020-10-28T15:38:20.943Z",
          "content": "<p>Many others in this thread have reported much closer (and higher) correlation.</p>\n<p>I just used a small random sample to start with (only made 5 successful submissions), so yet to explore how different validation setups work.</p>",
          "rawMarkdown": "Many others in this thread have reported much closer (and higher) correlation.\n\nI just used a small random sample to start with (only made 5 successful submissions), so yet to explore how different validation setups work.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1083238,
      "postDate": "2020-11-18T18:35:58.437Z",
      "content": "<p>I'm seeing very strong correlation between CV and LB so far, by using out-of-time validation and validating on the last 2M rows. <br>\nSome of my last models:<br>\nCV: 75.9 // LB: 76.0<br>\nCV: 77.0 // LB 76.9<br>\nCV: 77.6 // LB: 77.4 </p>",
      "rawMarkdown": "I'm seeing very strong correlation between CV and LB so far, by using out-of-time validation and validating on the last 2M rows. \nSome of my last models:\nCV: 75.9 // LB: 76.0\nCV: 77.0 // LB 76.9\nCV: 77.6 // LB: 77.4 ",
      "votes": 2
    },
    {
      "id": 1074330,
      "postDate": "2020-11-10T14:22:18.030Z",
      "content": "<p>Haven't yet made the sub but the below stats is using 1M/8M rows for training and 2.5M rows for validation with 16 feats. Thanks a lot everyone for the inspirations!</p>\n<pre><code>training's auc: 0.769886    valid's auc: 0.765254 # 1M train, 2.5M rows val\n\ntraining's auc: 0.770383    valid's auc: 0.766083 # 8M train, 2.5M rows val\n</code></pre>",
      "rawMarkdown": "Haven't yet made the sub but the below stats is using 1M/8M rows for training and 2.5M rows for validation with 16 feats. Thanks a lot everyone for the inspirations!\n\n```\ntraining's auc: 0.769886    valid's auc: 0.765254 # 1M train, 2.5M rows val\n\ntraining's auc: 0.770383    valid's auc: 0.766083 # 8M train, 2.5M rows val\n```",
      "votes": 2,
      "replies": [
        {
          "id": 1092929,
          "postDate": "2020-11-27T09:56:08.257Z",
          "content": "<p>So i have finally made a sub for a lgbm trained with 8M rows,  ~LB .762/.763 now.</p>",
          "rawMarkdown": "So i have finally made a sub for a lgbm trained with 8M rows,  ~LB .762/.763 now."
        }
      ]
    },
    {
      "id": 1071976,
      "postDate": "2020-11-07T17:09:38.927Z",
      "content": "<p>Someone has a perfect 1.0 score on the LB (as it seems for now). [it's a glitch i hope, no point downvoting me] <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">bug</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3fea56de822a51982159f43811503050%2FScreenshot%202020-11-07%20at%2010.37.31%20PM.png?generation=1604768976117675&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Someone has a perfect 1.0 score on the LB (as it seems for now). [it's a glitch i hope, no point downvoting me] [bug](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3fea56de822a51982159f43811503050%2FScreenshot%202020-11-07%20at%2010.37.31%20PM.png?generation=1604768976117675&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1071984,
          "postDate": "2020-11-07T17:23:04.803Z",
          "content": "<p>Very suspicious…</p>",
          "rawMarkdown": "Very suspicious..."
        },
        {
          "id": 1071988,
          "postDate": "2020-11-07T17:27:01.507Z",
          "content": "<p>I hope so, not sure why people would downvote, i just shared the information and TBH was amazed to see. Refreshed the page multiple times as well 😅</p>",
          "rawMarkdown": "I hope so, not sure why people would downvote, i just shared the information and TBH was amazed to see. Refreshed the page multiple times as well 😅",
          "votes": 1
        },
        {
          "id": 1071989,
          "postDate": "2020-11-07T17:28:04.103Z",
          "content": "<p>This person has kindly shared a bug here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815</a>. Presumably it will just be fixed. It's not foul play, more like a white hat hacking sort of situation from what I can tell.</p>",
          "rawMarkdown": "This person has kindly shared a bug here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815. Presumably it will just be fixed. It's not foul play, more like a white hat hacking sort of situation from what I can tell.",
          "votes": 2
        },
        {
          "id": 1071998,
          "postDate": "2020-11-07T17:35:46.010Z",
          "content": "<p>edit-&gt; placeholder [deleted]</p>",
          "rawMarkdown": "edit-> placeholder [deleted]"
        }
      ]
    },
    {
      "id": 1061720,
      "postDate": "2020-10-27T08:58:21.313Z",
      "content": "<p>nice implementation</p>",
      "rawMarkdown": "nice implementation",
      "votes": -11
    },
    {
      "id": 1061719,
      "postDate": "2020-10-27T08:58:00.180Z",
      "content": "<p>Nice implementation</p>",
      "rawMarkdown": "Nice implementation",
      "votes": -11
    },
    {
      "id": 1109706,
      "postDate": "2020-12-12T01:30:00.810Z",
      "content": "<p>CV 783 ,  LB : 783 , training with 7.5 million rows. However, my last big boost in cv was not a big boost in public leaderboard. I am working in geting a better data representation and looking for bugs in my last submissions.</p>",
      "rawMarkdown": "CV 783 ,  LB : 783 , training with 7.5 million rows. However, my last big boost in cv was not a big boost in public leaderboard. I am working in geting a better data representation and looking for bugs in my last submissions.",
      "replies": [
        {
          "id": 1109745,
          "postDate": "2020-12-12T02:59:22.517Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1109772,
          "postDate": "2020-12-12T03:37:27.487Z",
          "content": "<p>Lightgbm, no lecture features used</p>",
          "rawMarkdown": "Lightgbm, no lecture features used"
        }
      ]
    },
    {
      "id": 1105943,
      "postDate": "2020-12-08T11:03:03.023Z",
      "content": "<p>1.CV 0.767 / LB 0.766<br>\n2.CV 0.771 / LB 0.771</p>",
      "rawMarkdown": "1.CV 0.767 / LB 0.766\n2.CV 0.771 / LB 0.771",
      "replies": [
        {
          "id": 1115760,
          "postDate": "2020-12-16T14:57:09.357Z",
          "content": "<p>updated:</p>\n<p>CV: 0.785 LB: 0.786 with single model</p>",
          "rawMarkdown": "updated:\n\nCV: 0.785 LB: 0.786 with single model",
          "votes": 1
        },
        {
          "id": 1120677,
          "postDate": "2020-12-21T03:41:29.830Z",
          "content": "<p>How many features did you use? Did you train on kaggle's kernel？thank you.</p>",
          "rawMarkdown": "How many features did you use? Did you train on kaggle's kernel？thank you."
        },
        {
          "id": 1120915,
          "postDate": "2020-12-21T07:41:33.670Z",
          "content": "<p>It's around thirty features. Our team train the all model on local machine. Only use kaggle's kernel in inference.</p>",
          "rawMarkdown": "It's around thirty features. Our team train the all model on local machine. Only use kaggle's kernel in inference."
        }
      ]
    },
    {
      "id": 1105028,
      "postDate": "2020-12-07T13:21:53.483Z",
      "content": "<p>My CV is 0.756, but LB is 0.553. I can't find the reason.<br>\nWhen I drop variables, which is min/max/std/mean/skew of answer data of each user ID, CV is 0.701 and LB is 0.682.<br>\nI think maybe the additional variables are related with the problem, but I still don't know the exact reason.<br>\nAnyone having the similar experience?</p>",
      "rawMarkdown": "My CV is 0.756, but LB is 0.553. I can't find the reason.\nWhen I drop variables, which is min/max/std/mean/skew of answer data of each user ID, CV is 0.701 and LB is 0.682.\nI think maybe the additional variables are related with the problem, but I still don't know the exact reason.\nAnyone having the similar experience?",
      "replies": [
        {
          "id": 1105526,
          "postDate": "2020-12-08T01:09:38.147Z",
          "content": "<p>Maybe you had some errors in feature engineering during the testing phase.<br>\nTherefore, your LB and CV are very different</p>",
          "rawMarkdown": "Maybe you had some errors in feature engineering during the testing phase.\nTherefore, your LB and CV are very different"
        },
        {
          "id": 1105673,
          "postDate": "2020-12-08T04:36:46.043Z",
          "content": "<p>User aggregated information may result in a data leakage.</p>",
          "rawMarkdown": "User aggregated information may result in a data leakage.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1073485,
      "postDate": "2020-11-09T15:08:36.427Z",
      "content": "<p>CV: 0.76  LB: 0.737<br>\nCV: 0.92  LB: 0.702</p>\n<p>I don't get it, it makes no sense. My CV preparation is simple. I get the first 60% of the whole training set, for my training set, then do [ validation = X.groupby('user_id').tail(100) ]. I end up with 13 milion CV and 46 milion Train rows.</p>\n<p>I did switch from Pandas to cuDF in my submission notebook to do faster merging and all, so unless cuDF is doing some wild messed up merges, I have no idea why such a wild difference. Almost giving up</p>",
      "rawMarkdown": "CV: 0.76  LB: 0.737\nCV: 0.92  LB: 0.702\n\nI don't get it, it makes no sense. My CV preparation is simple. I get the first 60% of the whole training set, for my training set, then do [ validation = X.groupby('user_id').tail(100) ]. I end up with 13 milion CV and 46 milion Train rows.\n\nI did switch from Pandas to cuDF in my submission notebook to do faster merging and all, so unless cuDF is doing some wild messed up merges, I have no idea why such a wild difference. Almost giving up",
      "replies": [
        {
          "id": 1073562,
          "postDate": "2020-11-09T16:20:30.600Z",
          "content": "<p>That second CV score looks extremely high and suggests target-leaking features or overlap between your train and validation to me. These sorts of discrepancies almost always end up coming back to data preparation IMO -- can't emphasize enough how important it is to carefully double and triple check basically everything you do with features (especially at inference time). </p>\n<p>Also, I'd suggest using a different validation setup - taking the last 100 records is going to remove all of the relatively light users from your training data, so your model might be badly biased. One possible approach is to do it based on percentage instead of raw number: e.g. take the last 10% of each user for val so that you get user-record distributions in train and val that mirror user frequency.</p>",
          "rawMarkdown": "That second CV score looks extremely high and suggests target-leaking features or overlap between your train and validation to me. These sorts of discrepancies almost always end up coming back to data preparation IMO -- can't emphasize enough how important it is to carefully double and triple check basically everything you do with features (especially at inference time). \n\nAlso, I'd suggest using a different validation setup - taking the last 100 records is going to remove all of the relatively light users from your training data, so your model might be badly biased. One possible approach is to do it based on percentage instead of raw number: e.g. take the last 10% of each user for val so that you get user-record distributions in train and val that mirror user frequency.",
          "votes": 7
        },
        {
          "id": 1073574,
          "postDate": "2020-11-09T16:33:35.817Z",
          "content": "<p>It's not overlap nor leaking, I made sure to drop validation index and all that. I thought the score was really high because maybe 13 milion wasn't alot or maybe the examples were easy. So I wasn't expecting 0.92 on leaderboard, but not a drop from 0.737 to 0.702, since I only added 2 more features that do look really promising on CV and Feature Importance (gain and split for lgbm).</p>\n<p>I'll try what you said on the last 10% of each user, for now I switched validation split to [ validation = X.tail(int(X.shape[0]*0.33)) ]. CV and Training Loss are really close now, and the mean correctness of both are very close also. Got 40 milion for training and 20 for CV. Maybe it does better now.</p>",
          "rawMarkdown": "It's not overlap nor leaking, I made sure to drop validation index and all that. I thought the score was really high because maybe 13 milion wasn't alot or maybe the examples were easy. So I wasn't expecting 0.92 on leaderboard, but not a drop from 0.737 to 0.702, since I only added 2 more features that do look really promising on CV and Feature Importance (gain and split for lgbm).\n\nI'll try what you said on the last 10% of each user, for now I switched validation split to [ validation = X.tail(int(X.shape[0]*0.33)) ]. CV and Training Loss are really close now, and the mean correctness of both are very close also. Got 40 milion for training and 20 for CV. Maybe it does better now."
        },
        {
          "id": 1073578,
          "postDate": "2020-11-09T16:41:26.523Z",
          "content": "<blockquote>\n  <p>IMO -- can't emphasize enough how important it is to carefully double and triple check basically everything you do with features (especially at inference time).</p>\n</blockquote>\n<p>This happened yesterday on my local tests. I was doing something wrong and i wasn't aware of the same and i leaked something which resulted in a huge diff b/w train and val scores. (completely un intentionally) It took me around an hour to figure it out what went wrong.</p>\n<p>The idea here is that for each row, ensure you don't create any feature that you cannot compute at that point of time or you know about the same during that row. Plus try creating features which you can compute during inference as well easily first (to get a pipeline up) and then complicate stuffs.</p>",
          "rawMarkdown": ">IMO -- can't emphasize enough how important it is to carefully double and triple check basically everything you do with features (especially at inference time).\n\nThis happened yesterday on my local tests. I was doing something wrong and i wasn't aware of the same and i leaked something which resulted in a huge diff b/w train and val scores. (completely un intentionally) It took me around an hour to figure it out what went wrong.\n\nThe idea here is that for each row, ensure you don't create any feature that you cannot compute at that point of time or you know about the same during that row. Plus try creating features which you can compute during inference as well easily first (to get a pipeline up) and then complicate stuffs.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1064345,
      "postDate": "2020-10-30T03:00:01.670Z",
      "content": "<p>Im puzzled that with 10M trainset and 2.5M vset,  Local: 0.7683 and LB: 0.69. I think I have got a overfitting dilemma.</p>",
      "rawMarkdown": "Im puzzled that with 10M trainset and 2.5M vset,  Local: 0.7683 and LB: 0.69. I think I have got a overfitting dilemma."
    },
    {
      "id": 1063886,
      "postDate": "2020-10-29T13:06:14.153Z",
      "content": "<p>My LB and Local are radically different, the local one is 0.752 and on the LB is 0.680. Any idea ideas on why is that ? </p>",
      "rawMarkdown": "My LB and Local are radically different, the local one is 0.752 and on the LB is 0.680. Any idea ideas on why is that ? ",
      "replies": [
        {
          "id": 1063896,
          "postDate": "2020-10-29T13:18:04.617Z",
          "content": "<p>A gap this large points to data preparation issues, in my view. Check your code carefully to make sure that any feature preparation/transformation steps that you do on train are exactly the same as what do you on test (order of the columns when predicting, any label encoding/dummy variables/scaling including reusing the transformer fit on train instead of refitting it, null handling, etc.). Also might be worth checking that your prediction is formatted as expected as an actual probability, not a 0/1.  </p>",
          "rawMarkdown": "A gap this large points to data preparation issues, in my view. Check your code carefully to make sure that any feature preparation/transformation steps that you do on train are exactly the same as what do you on test (order of the columns when predicting, any label encoding/dummy variables/scaling including reusing the transformer fit on train instead of refitting it, null handling, etc.). Also might be worth checking that your prediction is formatted as expected as an actual probability, not a 0/1.  ",
          "votes": 4
        },
        {
          "id": 1064553,
          "postDate": "2020-10-30T08:41:15.093Z",
          "content": "<p>I have checked that all the preparation/transformation steps I have the same as env_test do.  But my local one is 0.768 with LB 0.690, what's wrong with such case? Will that be caused by wrong chosen features or wrong methods of data pretreatment? </p>",
          "rawMarkdown": "I have checked that all the preparation/transformation steps I have the same as env_test do.  But my local one is 0.768 with LB 0.690, what's wrong with such case? Will that be caused by wrong chosen features or wrong methods of data pretreatment? "
        },
        {
          "id": 1064680,
          "postDate": "2020-10-30T12:05:39.127Z",
          "content": "<p>What closed the gap a bit for my case (not completely though), is that I extracted few features from the main CSV file, then I merged them during training. Since I didn't use the entire 100M row to extract these features, some content_ids were missing, thus resulting in a nan, leading the the decline of accuracy during scoring later on.<br>\nWhat I've done is simply use the entire dataset to extract features, and a portion of the dataset for training, which led to a bump in score 0.680 -&gt; 0.724 vs local score 0.760</p>",
          "rawMarkdown": "What closed the gap a bit for my case (not completely though), is that I extracted few features from the main CSV file, then I merged them during training. Since I didn't use the entire 100M row to extract these features, some content_ids were missing, thus resulting in a nan, leading the the decline of accuracy during scoring later on.\nWhat I've done is simply use the entire dataset to extract features, and a portion of the dataset for training, which led to a bump in score 0.680 -> 0.724 vs local score 0.760",
          "votes": 2
        }
      ]
    },
    {
      "id": 1058805,
      "postDate": "2020-10-24T10:03:28.310Z",
      "content": "<p>There is another thread <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191413\" target=\"_blank\">191413</a> on validation</p>",
      "rawMarkdown": "There is another thread [191413](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191413) on validation",
      "replies": [
        {
          "id": 1066778,
          "postDate": "2020-11-02T03:00:53.600Z",
          "content": "<p>But no one shared scores there 😐</p>",
          "rawMarkdown": "But no one shared scores there 😐",
          "votes": 1
        }
      ]
    },
    {
      "id": 1076266,
      "postDate": "2020-11-12T11:38:37.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1062749,
      "postDate": "2020-10-28T06:43:10.823Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1063014,
          "postDate": "2020-10-28T12:33:03.190Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true,
          "replies": [
            {
              "id": 1063086,
              "postDate": "2020-10-28T13:49:50.130Z",
              "content": "<p><a href=\"https://www.kaggle.com/dhyeypatel1234\" target=\"_blank\">@dhyeypatel1234</a> :</p>\n<blockquote>\n  <p>Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.</p>\n</blockquote>\n<p>^From the competition's evaluation info [<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/overview/evaluation\" target=\"_blank\">link</a>]</p>",
              "rawMarkdown": "@dhyeypatel1234 :\n> Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n\n^From the competition's evaluation info [[link](https://www.kaggle.com/c/riiid-test-answer-prediction/overview/evaluation)]",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 1060810,
      "postDate": "2020-10-26T14:51:45.027Z",
      "rawMarkdown": "",
      "votes": -5,
      "isDeleted": true,
      "replies": [
        {
          "id": 1060857,
          "postDate": "2020-10-26T15:17:49.560Z",
          "content": "<p>…why plagiarize so transparently? </p>",
          "rawMarkdown": "...why plagiarize so transparently? "
        },
        {
          "id": 1060863,
          "postDate": "2020-10-26T15:23:29.893Z",
          "content": "<p>Did you just……………..👌😆😂</p>",
          "rawMarkdown": "Did you just.................👌😆😂"
        },
        {
          "id": 1060868,
          "postDate": "2020-10-26T15:30:48.447Z",
          "content": "<p>Or maybe he didn't realise he was logged in from the wrong id 😂</p>",
          "rawMarkdown": "Or maybe he didn't realise he was logged in from the wrong id 😂",
          "votes": 2
        },
        {
          "id": 1061170,
          "postDate": "2020-10-26T19:33:40.467Z",
          "content": "<p>Looks like he has a nice armada of downvoting accounts 😮</p>",
          "rawMarkdown": "Looks like he has a nice armada of downvoting accounts 😮",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1112090,
      "author_name": "sakami",
      "author_url": "",
      "post_date": "2020-12-14T09:16:54.680000",
      "content": "<ul>\n<li>CV: 0.80664, LB: 0.806</li>\n<li>Single SAINT-based model</li>\n<li>without lectures</li>\n</ul>",
      "votes": 18,
      "replies": [
        {
          "id": 1112092,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-14T09:20:43.613000",
          "content": "<p>Can you share the model size? I've reach around 0.792 using SAINT+ now but still using 128 d_model and 4 num_layers. Would increasing d_model help much?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112097,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "2020-12-14T09:23:38.140000",
          "content": "<p>same as the paper: d_model = 512, num_layers = 4</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1112109,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-14T09:31:12.723000",
          "content": "<p>Thanks, I'll see how the scaled up version works for me as well. Issue is, one model training takes up almost all of the week's GPU quota :/</p>\n<p>Have to do iterations once a week only…</p>\n<p>Anyway, best of luck. Include Lectures btw, they increased my score by ~0.01 i think</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1112131,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-14T10:07:05.447000",
          "content": "<p>legend 🙏🙏🙏🙏🙏</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1112456,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-14T16:11:45.440000",
          "content": "<p>thanks, I will also try saint++ model !</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1112913,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-12-15T02:27:40.230000",
          "content": "<p>Nice socre, do you mind share how you split train and valid ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1113487,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "2020-12-15T13:42:04.707000",
          "content": "<p>same as <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> split, dropping last 2.5M samples.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1115459,
          "author_name": "AK",
          "author_url": "",
          "post_date": "2020-12-16T10:01:27.203000",
          "content": "<p>This is what looks like when you surrond by all genious people 😊</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1117434,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-18T03:50:02.353000",
          "content": "<p><a href=\"https://www.kaggle.com/sakami\" target=\"_blank\">@sakami</a> ; Just curious, If you won't mind, Can you share how are you handling larger sequences? Splitting them up of picking randomly (in sort of time-sorted fashion)? <br>\nTy!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1123749,
          "author_name": "Chahhou Mohamed",
          "author_url": "",
          "post_date": "2020-12-23T13:27:21.937000",
          "content": "<p>Hi Sakami, did you get that score using only the features used in the paper or did you use some additional features. I also tried the the saint approach, but it cant beat my single lstm (0.797)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1124164,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-23T17:52:23.783000",
          "content": "<p>Wow! Single LSTM scoring .797 is mind blowing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1125306,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "2020-12-24T16:00:37.480000",
          "content": "<p>Sorry for late reply.</p>\n<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I'm using the same sampling method used in many public notebooks.<br>\n<a href=\"https://www.kaggle.com/mchahhou\" target=\"_blank\">@mchahhou</a> I'm using completely same features as the paper.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1059325,
      "author_name": "Ming Pan",
      "author_url": "",
      "post_date": "2020-10-24T23:42:47.820000",
      "content": "<p>Last LB score increases:</p>\n<p>Local: <strong>0.7609</strong> LB: <strong>0.759</strong><br>\nLocal: <strong>0.7757</strong> LB: <strong>0.772</strong><br>\nLocal: <strong>0.7817</strong> LB: <strong>0.780</strong><br>\nLocal: <strong>0.7866</strong> LB: <strong>0.786</strong></p>\n<p>I'm using 10m rows to train and 2.5m to validate per fold.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 1072868,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-11-08T19:28:41.137000",
          "content": "<p><a href=\"https://www.kaggle.com/mingpan07\" target=\"_blank\">@mingpan07</a> it would be great if you can share your approach how did you get this strong CV/LB correlation ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1058744,
      "author_name": "Nick Sarris",
      "author_url": "",
      "post_date": "2020-10-24T07:29:02.003000",
      "content": "<p>Regarding CV/LB scores, my last three jumps on the LB were the following:</p>\n<p>CV: <strong>0.725487</strong>, LB: <strong>0.761</strong><br>\nCV: <strong>0.740906</strong>, LB: <strong>0.767</strong><br>\nCV: <strong>0.747638</strong>, LB: <strong>0.772</strong></p>\n<p>My approach to validation is pretty simple, just the following to create the training/validation set. I use the remaining 90+% of the data for feature aggregates and the like. Even though the CV score isn't too similar to the LB, an increase in CV typically comes with an increase in LB:</p>\n<pre><code>print(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n</code></pre>\n<p>Update:</p>\n<p>CV: <strong>0.790365</strong>, LB: <strong>0.787</strong></p>\n<p>I switched up my validation strategy to instead use the one described in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a>. Seems to be relatively accurate from what I've noticed throughout the thread, typically &lt;= 0.003 of a difference beween CV/LB. This score is calculated using 10M rows for training, 2.5M rows for validation. I'm not using any lecture features yet in the implementation.</p>\n<p>Update #2:</p>\n<p>CV: <strong>0.793102</strong>, LB: <strong>0.789</strong></p>\n<p>Added a few more features, changed the framework of my code throughout to make it much easier to add/test new features. Still using only 10M rows, trying to see if I can hit ~0.795+ before I really look into training with the entire training set, given the estimated 0.01 boost from other people.</p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 1058962,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2020-10-24T13:00:36.780000",
      "content": "<p>Local Validation: <strong>0.770</strong>, LB: <strong>0.766</strong>.<br>\nUpdate: Local <strong>0.777</strong>, LB: <strong>0.772</strong>. </p>\n<p>That's with training on ~4% of the data, and validating on ~2%. My validation scheme is set up to mirror the way the test predictions need to be made, so I'm fairly confident in it. I think a bit of a gap is normal especially given the time element.</p>\n<p>I strongly agree with what Alex said below about prioritizing finding the right features before using more of the data. 2% is already millions of samples and likely large enough to give reliable model feedback IMO (and even close to the same order as the actual test set). The key thing early in a competition is to be able to iterate on feature/model ideas quickly, and I'm pretty sure training time scales above linear with # of training rows for the typical gradient boosting model.</p>\n<p>Edit: see update </p>",
      "votes": 11,
      "replies": [
        {
          "id": 1058966,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-10-24T13:11:33.600000",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> Thanks for the info<br>\nAre you are now generating features for the 4% of data or for the whole data and train it with 4% of the data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1058969,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2020-10-24T13:17:45.893000",
          "content": "<p>You definitely need to use all the training data for features to handle test well, so I do that. In particular, question level features come from all the data. One distinction though - my 6% train/val is a self-contained set of users, so I only really need that 6% of full user features to train/validate. For the rest of the data I have summarized user features.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1058975,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-24T13:32:53.970000",
          "content": "<p>Would you mind to share how to handle the features for unseen user during test time? <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a>  <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> ? Or point me to a notebook that handle this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058986,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T13:50:28.873000",
          "content": "<p>I essentially have a table of users where I keep some useful attributes for each of them… it starts out empty and gets filled by processing train and test data, simply adding a new row/entry for each previously unseen user and updating information as needed for already existing users.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1059004,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2020-10-24T14:11:39.940000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I actually have prepared a notebook like that, but unfortunately I don't feel comfortable sharing it yet because it scores fairly highly (top 50ish). I haven't seen any public examples as of yet. I'm hoping to share a little later on, at a point when it wouldn't disrupt the leaderboard as much.</p>\n<p>My method is similar to what Alex describes, where a table gets updated on a running basis as the test data comes in. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1059011,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-24T14:22:17.867000",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> No problem. I am just curious how this is done in FE. As I mentioned, I focus on sequential model, which should just use sequences as input, so there is no actual problem for unseen users. Of course, for unseen users, there is no history to use for the first prediction during test time - but the model, during training, will learn the distribution of targets in this case anyway. (You can think it as predicting the first word in a translation problem).</p>\n<p>Being said so, you could try to interoperate features among all users, like <code>what's the ratio of answer correctly a question if this question is the first question in a user interaction sequences.</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1059193,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-24T18:56:32.740000",
          "content": "<blockquote>\n  <p>Features are entirely non-future-leaking</p>\n</blockquote>\n<p>Little confused here, let's say a user which you saw in train is also there in your val, so shall you use training stats and update that or use fresh stats? In case we don't use fresh, then we are leaking the user's past performance in val data, right? Am i missing something here? Ty!</p>\n<blockquote>\n  <p>Would you mind to share how to handle the features for unseen user during test time? Or point me to a notebook that handle this?</p>\n</blockquote>\n<p>One way to do this is what's already mentioned in this thread by Alex and Joe, we can index the user_stats_df's with the user_id (as the uid) for e.g.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059198,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2020-10-24T19:10:41.137000",
          "content": "<blockquote>\n  <p>In case we don't use fresh, then we are leaking the user's past performance in val data, right?</p>\n</blockquote>\n<p>Using past performance info to predict the future is not a leak. Leaks refer to cases where the model has access to information that would not actually be available at prediction time, like future performance. In this problem, when we process and predict on the test data, we have access to past performance for users that aren't brand new. So long as the features you use are derived entirely from what the user has done in the past, you will be leak-free.</p>\n<p>The bigger picture way of thinking about this is that it's not really about the idea of an \"information leak\", but more a question of what's known at training time vs. what's known at prediction time. In most cases those two should mirror each other as much as possible for optimal results, both to prevent overfitting and underfitting. In particular, if you avoid using past information you're not protecting yourself from the harm of a leak, but instead denying your model relevant information that should almost certainly improve its performance.</p>\n<p>Another extension of this idea is that in a kaggle competition setting, it can be correct to use unrealistic, future-leaking features (talkingdata click prediction from a few years ago is a good example), so long as all parts of the timeline are accessible. Luckily, this competition gives you a much more realistic paradigm via the API submission process.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1059205,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T19:15:07.287000",
          "content": "<p>For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -&gt; mean correctness etc.). The features in each row then only depend on past information.</p>\n<p>So I end up having a very long list of some features (historic mean user correctness, mean question correctness, etc.) and the information whether this question was answered correctly and that's what I'm training on. User IDs are dropped so I couldn't even tell anymore which user is which. The goal is to find a general model that predicts the chance of a random user getting the question right based on that random user's past performance, pretty much.</p>\n<p>So of course the user's past performance will \"leak\" into the validation data if you want to call it that, but then which features would we even predict on if we didn't want to take into account past performance?</p>\n<p>My validation is a random split, so it might well be (and will undoubtedly happen) that I train on the 500th question answered by a user, using knowledge of the outcome of all their previous questions, and validate the result using, amongst other data, the 10th question answered by this user given features calculated from the outcome of their previous 9 questions.</p>\n<p>Does this make sense? I'm not sure whether that answers your question.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1059400,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-25T03:39:37.357000",
          "content": "<p>Thanks a lot Joe and Alex!  It's much clear now!👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059911,
          "author_name": "Nya 🚀",
          "author_url": "",
          "post_date": "2020-10-25T15:24:17.317000",
          "content": "<blockquote>\n  <p>For a single user, I go row by row in order of time - for each row calculating all features based on the data available from previous rows (number of answers, number of correct answers -&gt; mean correctness etc.). The features in each row then only depend on past information.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> How are you able to do all this for 393,656 users without surpassing the 9h time limit? You only have 0.08 seconds available per user, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1059937,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-25T15:51:21.710000",
          "content": "<p>My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.</p>\n<p>I don't process the data literally row by row, just row by row <em>per user</em>. In practical terms this means I keep a list of test set rows for each user, sorted by timestamp, and batch the processing by that. I.e., first batch is the first entry of each user, second batch the second entry of each user and so on.</p>\n<p>The kernel split is done on a user-basis, so each kernel handles about 80k users. So the first batch in the preprocessing loop will be some 80k rows (takes some 5 secs maybe), and the batches get smaller, in a somewhat exponential fashion. The most active users have some &gt;15k rows (quite impressive dedication of them actually) so by the end I'm forced to go row by row, but that's not a big issue either.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 1060026,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-10-25T17:32:37.023000",
              "content": "<blockquote>\n  <p>My test preprocessing is currently split across 5 kernels, but all finish within about 1.5 hours - the split is to accommodate the large size of the output feature files which get close to 20 GB or so I guess. It could definitely finish in one kernel, maybe in 4-5 hours at a guess.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> During test time, how do you get to load these 20GB of features into memory and use them for the prediction? No OOM issue?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1060047,
              "author_name": "Alex Bader",
              "author_url": "",
              "post_date": "2020-10-25T18:03:54.940000",
              "content": "<p>Currently I'm simply not using all the data to train my model, just a fraction of it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1060052,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-10-25T18:09:59.013000",
              "content": "<p>I mean the prediction for submission. During inference, you don't need to use the features that are prepared for training? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1060102,
              "author_name": "Alex Bader",
              "author_url": "",
              "post_date": "2020-10-25T19:18:29.130000",
              "content": "<p>Once the model is trained the train features aren't needed anymore, so in the inference kernel I just load the pre-trained model. Test data needs to be preprocessed though before sticking it through the model.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 1059943,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-25T15:58:03.757000",
          "content": "<blockquote>\n  <p>The most active users have some &gt;15k rows (quite impressive dedication of them actually)</p>\n</blockquote>\n<p>They certainly have a bright future 😁</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1059967,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-25T16:19:29.773000",
          "content": "<p>20 GB+ in features really means a hell lot of features 🙏🙏! Nice idea to split the work load's in multiple kernels for faster turn-around time as well!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059970,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-25T16:23:33.240000",
          "content": "<p>Meh, it's not that many actually… I should probably just switch to saving in a more efficient format than floats in plain csv…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059991,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-25T16:49:36.100000",
          "content": "<p>Why not store them as string's when saving them and while loading back, load as float point numbers?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059994,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-25T16:55:47.160000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1059999,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-25T17:07:46.900000",
          "content": "<blockquote>\n  <p>Why not store them as string's when saving them and while loading back, load as float point numbers?</p>\n</blockquote>\n<p>In a .csv it saves the string representation of the floats (which can be pretty long), but it would be more space-efficient to save the floats themselves in binary format. I should really go and save in feather or something 😛</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1060004,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-25T17:09:53.990000",
          "content": "<p><code>They certainly have a bright future 😁</code></p>\n<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> , if they come to Kaggle to compete, I won't be able to beat them 😂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1060024,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-25T17:29:06.723000",
          "content": "<p>Yep, we can go for binary format for sure as numpy will help us there (never did it myself until today;) or array in Py but it's not used much in day-to-day activities. But do we need more than 4-5 decimals of precision?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060046,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-25T18:00:44.913000",
          "content": "<blockquote>\n  <p>But do we need more than 4-5 decimals of precision?</p>\n</blockquote>\n<p>That's what I thought too, but a run with more decimals gave me a better score, and not a negligible increase at that. Or maybe I added a feature in between these runs and forgot about it. <br>\nAnyway, better safe than sorry is my current approach…</p>\n<p>Pandas has plenty of saving options so binary is not a problem at all from a practical point of view. Upside of .csv is that in Kaggle's dataset viewer you can directly see column histograms and check whether everything went right with scaling and such.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1060092,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-25T18:55:13.240000",
          "content": "<blockquote>\n  <p>The most active users have some &gt;15k rows (quite impressive dedication of them actually)</p>\n</blockquote>\n<p>How do we know that they are not bots who were scrapping riddi? (out of topic, just kidding)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060926,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-10-26T16:16:15.073000",
          "content": "<p>Is it possible to do ensembling and stacking in this challenge<br>\nbecause test data can only be called once right??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060995,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-26T17:21:03.983000",
          "content": "<p>I would simply suggest, do something simpler first as that itself is quite complicated and challenging many times. And then one can build on top that!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1105473,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2020-12-07T22:58:00.777000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Of course they have a bright future. If one of them is already here on kaggle, she just check the competion: <br>\n\"hmm, I created the training data for years, let's win this :P\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1093890,
      "author_name": "nadare",
      "author_url": "",
      "post_date": "2020-11-28T05:24:03.260000",
      "content": "<p>single NN, inference for ~3h<br>\nI use <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> strategy.<br>\nCV:  0.772, LB: 0.776</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1093913,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-28T05:49:23.743000",
          "content": "<p>Nice CV! Can you write few words about your NN! Is it transformer's based? Ty!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1125940,
          "author_name": "nadare",
          "author_url": "",
          "post_date": "2020-12-25T07:32:00.760000",
          "content": "<p>NN based, implemented with tensorflow, inference for 7h<br>\n<a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">same cv strategy</a><br>\nCV: 0.803, LB: 0.799</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1125986,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-25T08:12:04.197000",
          "content": "<p>Awesome Score! Good Luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130563,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-29T05:59:50.383000",
          "content": "<p><a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> we have strong pipeline with single saint score of 78.6 we work on saint plus. let me know your thought about teamup</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1064170,
      "author_name": "alijs",
      "author_url": "",
      "post_date": "2020-10-29T19:54:33.937000",
      "content": "<p>So far CV-LB correlation looks very stable - it's constantly +0.003 on LB compared to CV for all my submissions in the last 2 weeks.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1071447,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2020-11-06T22:30:38.373000",
      "content": "<p>CV: .792 <br>\nLB: .791</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1073958,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-11-10T04:27:57.037000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Could you please explain the validation strategy?😊</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1074479,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-11-10T17:37:32.283000",
          "content": "<p>Okay, I will explain the details in the winning solution.👌</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1101258,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-03T18:54:21.157000",
          "content": "<p>UPDATES:<br>\nCV: 0.8090<br>\nLB: 0.808X</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1101446,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2020-12-03T22:52:01.147000",
          "content": "<p>stop plz! 😂</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1102368,
          "author_name": "Tonghui Li",
          "author_url": "",
          "post_date": "2020-12-04T21:09:07.627000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> Did you achieve this score using LGBM?  I am asking because I am quite curious where the limit of GBDT model in this competition is.  Thanks in advance! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1102597,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-05T05:18:49.610000",
          "content": "<p>It's a top secret! But, please remember nyanp says his LGBM is .801, and Nikola Bacic (.806 in LB) says he is using Transformer. From their comments, I think Transformer may perform a bit better, but LGBM is competitive.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1103287,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-12-05T19:39:31.283000",
          "content": "<p>LGBM specifically, or GBM generally, including xgboost?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1103955,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-06T13:00:45.217000",
          "content": "<p>I think 100M rows are too computationally expensive for XGBoost/CatBoost.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1111106,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-13T12:12:52.303000",
          "content": "<p>UPDATES: <br>\nCV: 0.8125<br>\nLB: 0.811x</p>\n<p>seems I'm overfitting :(</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1111116,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2020-12-13T12:33:42.893000",
          "content": "<p>poor you :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1118037,
          "author_name": "Zhenghan Chen",
          "author_url": "",
          "post_date": "2020-12-18T17:13:51.960000",
          "content": "<p>thus, nn might win in such case. I strongly believe.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117386,
      "author_name": "qyxs",
      "author_url": "",
      "post_date": "2020-12-18T02:11:25.650000",
      "content": "<p>CV: 0.802, LB: 0.802</p>\n<p>single lightgbm model with 100+ features, nothing found in lectures😓. Anyone boost score from lectures features?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1118049,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-18T17:21:25.227000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> you're not alone, we're not using lectures, every attempt with lectures failed for us. Interested too with feeback on lectures features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129008,
          "author_name": "qyxs",
          "author_url": "",
          "post_date": "2020-12-28T01:53:30.073000",
          "content": "<p>update, single lgb CV:0.806, LB:0.806, still without lectures features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1130744,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-12-29T09:38:57.793000",
          "content": "<p><a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Do you guys mean lectures don't improve your validation score or LB?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1130784,
          "author_name": "qyxs",
          "author_url": "",
          "post_date": "2020-12-29T09:55:33.813000",
          "content": "<p>Your guess is right. So lectures improve your score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130788,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-29T09:57:05.403000",
          "content": "<p>It improves CV (a little bit) but not LB. I will give another try this week, I may have a bug somewhere.<br>\nBut lectures are only 2% of data so it might not be important.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1130806,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-12-29T10:12:54.343000",
          "content": "<p><a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> I asked an OR question. So which one?:) </p>\n<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> 2% of data is actually important since we all compete for 3rd decimal on metric. I and several other people reported that it improves CV but not LB. Maybe we all have the same inference bug.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1130864,
          "author_name": "qyxs",
          "author_url": "",
          "post_date": "2020-12-29T10:59:13.777000",
          "content": "<p>Because lectures features did not improve our local cv, we did not submit. If you generated features from lectures after user's answer records, it might improve your local cv, but the online lectures info maybe in next group, this could be the reason that LB not improve.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1122135,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2020-12-22T08:09:12.127000",
      "content": "<p>CV: 0.808, LB: 0.805</p>\n<p>Looking at everyone has better CV-LB gap, I either have a bug in inference or something else is wrong.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1122149,
          "author_name": "ML_and_DL",
          "author_url": "",
          "post_date": "2020-12-22T08:26:17.310000",
          "content": "<p>Can I form a team with you to study.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1122263,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-12-22T10:08:54.930000",
          "content": "<p>No.</p>\n<p></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122272,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-22T10:18:49.190000",
          "content": "<p>During inference, you can have multiple questions for the same user within the same batch. If you've some kind of expanding features then are you updating those features in that case or do you wait for the next batch? I'm facing this issue with simular gap on my local simulator.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122279,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-12-22T10:27:48.510000",
          "content": "<p>It depends if multiple questions have different timestamps in the same batch. I assume they have the same timestamp?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122282,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-22T10:29:51.690000",
          "content": "<p>Yes, same timestamp, all within the same task_container_id.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122286,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-22T10:31:18.420000",
          "content": "<p>I can confirm from my personal probe.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1122564,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-22T14:29:16.927000",
          "content": "<p>update features ，and update dict next batch.may be userful </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1071389,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2020-11-06T20:18:28.260000",
      "content": "<p>Local: <strong>0.7602</strong>   LB: <strong>0.773</strong><br>\nLocal: <strong>0.7699</strong>   LB: <strong>0.782</strong></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1071175,
      "author_name": "大顺",
      "author_url": "",
      "post_date": "2020-11-06T15:28:07.570000",
      "content": "<p>Local:766<br>\nLB:785</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1066774,
      "author_name": "Young for you",
      "author_url": "",
      "post_date": "2020-11-02T02:57:24.543000",
      "content": "<p>Local:763,LB:780<br>\nUpdate: Local:771,LB:790</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1066232,
      "author_name": "toxu",
      "author_url": "",
      "post_date": "2020-11-01T14:02:54.863000",
      "content": "<p>CV:0.798<br>\nLB:0.771<br>\nValidating method: GroupKFold by user_id.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1058825,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-10-24T10:33:37.097000",
      "content": "<p>My last two submissions were</p>\n<p>Local CV: <strong>0.7637</strong>, LB: <strong>0.767</strong><br>\nLocal CV: <strong>0.7567</strong>, LB: <strong>0.760</strong></p>\n<p>So both times my LB was about 0.033 higher than my local CV score. Not sure why but I'll take it.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1058828,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T10:37:19.400000",
          "content": "<p>As for approaches, I'm still working with only small random selections of the train data - my last model was a quick-build on about 10% or maybe even less of the training data, validated with some 2% of it. Features are entirely non-future-leaking so I can happily train and validate on whichever rows I feel like without having to pay attention to that.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1058867,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-24T11:28:05.680000",
          "content": "<p>18th with 10% 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058874,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T11:34:28.970000",
          "content": "<p>Finding good features makes all the difference really. Sure I can get some increases using all the set (and I'll eventually do that), but finding good features is much more important in the beginning and will give you quite some improvements to work with.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1058878,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-24T11:36:35.203000",
          "content": "<p>Yeah, I know. I focus on sequential models though - haven't done any FE yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058887,
          "author_name": "Ala Eddine Ayadi",
          "author_url": "",
          "post_date": "2020-10-24T11:42:39.193000",
          "content": "<p>I guess 10 % to train the model, but FE is on the full data , right ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058899,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T11:53:30.653000",
          "content": "<p>Ok so in more detail, I get features for the entire train set, going for each user by timestep essentially to get non-future-leaked features (like mean user accuracy and such, will be very noisy for the first few questions of each user but getting better over time).<br>\nThis gives a me a full train set of features as well as a post-train status of some parameters which are used to calculate the features of the test set and updated with each test row (imagine, number of answers given by user x or so).</p>\n<p>So my model is currently only trained on like 10% of the train data, but when evaluating the test data with it I have the full cumulative history, so to say, of all previous events in order to calculate test features.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1082908,
      "author_name": "nyanp",
      "author_url": "",
      "post_date": "2020-11-18T11:15:57.953000",
      "content": "<p>Local: 0.798 LB: not submitted (with time-series split)</p>\n<p>The road to 1st submission is a long way for me 😂</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1082922,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-11-18T11:35:39.917000",
          "content": "<p>STOP PLS!!!</p>",
          "votes": -8,
          "replies": []
        },
        {
          "id": 1101447,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2020-12-03T22:53:18.930000",
          "content": "<p>updated:</p>\n<p>CV: 0.803 LB: 0.801 with single LightGBM</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1101478,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-12-03T23:51:44.653000",
          "content": "<p>Nice result - how much do you think is related to:</p>\n<ul>\n<li>training/validation split</li>\n<li>model parameters</li>\n<li>features</li>\n</ul>\n<p>Thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101536,
          "author_name": "Nick Sarris",
          "author_url": "",
          "post_date": "2020-12-04T02:03:27.743000",
          "content": "<p>How many rows are you training with? I'm currently at a CV/LB of 0.793/0.789 with a single LGB and 10M rows, and trying to figure out how much further I can push it before upping the number of rows.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1101543,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-04T02:13:10.890000",
          "content": "<p>amazing work</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101577,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2020-12-04T03:25:19.270000",
          "content": "<p>I always use full dataset for submission. In my case, using full datasets increases CV by +0.006 compared to 6M rows. All hyperparameters except for num_leaves and learning_rate (for speedup training) remain the default. I think feature engineering on the solid validation is much more important than model parameters in case of GBDT.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1101580,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-04T03:29:35.460000",
          "content": "<p>Perfect Feature Engineering👍👍👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1102849,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-05T11:48:14.457000",
          "content": "<p>nyanp, did you really change <code>num_leaves</code> for full data training? or do you mean <code>nround</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1102861,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2020-12-05T12:08:05.047000",
          "content": "<p>I forgot about the nround, it's based on early stopping, so of course it's different for full rows and 6M rows. As for the other parameters, I adjusted the num_leaves and learning_rate in 6M rows and used the same parameters in the full data as well.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1102865,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-05T12:12:55.537000",
          "content": "<p>Ah ok thanks, I got it!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1071733,
      "author_name": "kurupical",
      "author_url": "",
      "post_date": "2020-11-07T10:47:21.740000",
      "content": "<p>CV: 0.777<br>\nLB: 0.784</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1065768,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-10-31T18:08:31.307000",
      "content": "<p>I have at the moment 0.7579 on validation and 0.763 on LB. I'm validating only against unseen users, and that's probably why it is quite higher on LB.</p>\n<p>EDIT: I'm using 5 M rows, 33% of them as validation.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1127999,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2020-12-27T05:08:31.843000",
      "content": "<p>Same as <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this</a> cv strategy.<br>\nLocal: 0.8027, LB: 0.805x</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1058729,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-10-24T06:56:11.427000",
      "content": "<p>My last 3 submissions (using holdout validation set):</p>\n<p>Local: <strong>0.756</strong> LB: <strong>0.758</strong><br>\nLocal: <strong>0.748</strong> LB: <strong>0.751</strong><br>\nLocal: <strong>0.741</strong> LB: <strong>0.742</strong></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1058735,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-24T07:11:54.130000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1103563,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-12-06T03:15:47.543000",
      "content": "<p>CV: 0.778 - LB: 0.779</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1090124,
      "author_name": "qiaqia",
      "author_url": "",
      "post_date": "2020-11-25T05:11:11.037000",
      "content": "<p>some last models:<br>\nCV 0.756//LB0.759<br>\nCV0.770//LB0.766<br>\nCV0.771//LB0.769<br>\nCV0.774//LB0.772<br>\nCV0.777//LB0.774<br>\nit Looks like the LB score is fluctuating around .003.<br>\nsingle  LGBM LB0.801 sounds so magic,need more features..💪💪💪</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1102687,
          "author_name": "biubiuG",
          "author_url": "",
          "post_date": "2020-12-05T07:33:01.243000",
          "content": "<p>but, seems you are lucky that you received a consistent improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1125319,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-24T16:13:37.280000",
          "content": "<p>update my features style ,now CV 0.781,LB0.781</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1079477,
      "author_name": "Ken Ho",
      "author_url": "",
      "post_date": "2020-11-16T05:37:12.520000",
      "content": "<p>Local: <strong>0.7638</strong> LB <strong>0.765</strong></p>\n<p>Using all training data and follow the cross-validation strategy in <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1089909,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-11-24T22:16:51.303000",
          "content": "<p>I'm using the same cross-validation strategy, however I'm experiencing a gap in LB vs Local (0.757 vs 0.773). I noticed this gap exists only when I used cumulative user features. Do you experience the same ? Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1078828,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-11-15T11:02:56.707000",
      "content": "<p>Local CV: 0.740, LB: 0.727</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1075135,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2020-11-11T12:23:35.723000",
      "content": "<p>CV: 0.763 LB: 0.771 with simple split by user_id</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1089910,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-11-24T22:17:27.473000",
          "content": "<p>Can you tell me how did you split by user_id ? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1104145,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2020-12-06T16:51:39.337000",
          "content": "<p>Update:<br>\nvalidation: 0.8064 LB: 0.808x</p>\n<p>but at what cost…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1064367,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2020-10-30T03:42:16.150000",
      "content": "<p>Local: 0.79305<br>\nLB:     0.792</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062828,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-10-28T08:24:33.617000",
      "content": "<p>Hi,<br>\nLocal : 0.733<br>\nLB: 0.758</p>\n<p>Pretty classic validation scheme:<br>\nMy training set contains about 10M rows, 3M for validation. Features are aggretated on the remaining rows</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1063029,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-28T12:52:48.650000",
          "content": "<p>The gap is similar to what <a href=\"https://www.kaggle.com/misfyre\" target=\"_blank\">@misfyre</a> reported.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1063145,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-10-28T14:55:56.747000",
          "content": "<p><a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a> I'm impressed by the strong correlation between your local validation and public LB. Are you using something similar to <a href=\"https://www.kaggle.com/misfyre\" target=\"_blank\">@misfyre</a> ?</p>\n<p>I've built my hold-out set in the same way so it makes sense that there is a similar gap. However I cannot reprocude his jump from 0.725 to 0.747 😔</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1063174,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-28T15:38:20.943000",
          "content": "<p>Many others in this thread have reported much closer (and higher) correlation.</p>\n<p>I just used a small random sample to start with (only made 5 successful submissions), so yet to explore how different validation setups work.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1083238,
      "author_name": "Ravox",
      "author_url": "",
      "post_date": "2020-11-18T18:35:58.437000",
      "content": "<p>I'm seeing very strong correlation between CV and LB so far, by using out-of-time validation and validating on the last 2M rows. <br>\nSome of my last models:<br>\nCV: 75.9 // LB: 76.0<br>\nCV: 77.0 // LB 76.9<br>\nCV: 77.6 // LB: 77.4 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1074330,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-10T14:22:18.030000",
      "content": "<p>Haven't yet made the sub but the below stats is using 1M/8M rows for training and 2.5M rows for validation with 16 feats. Thanks a lot everyone for the inspirations!</p>\n<pre><code>training's auc: 0.769886    valid's auc: 0.765254 # 1M train, 2.5M rows val\n\ntraining's auc: 0.770383    valid's auc: 0.766083 # 8M train, 2.5M rows val\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1092929,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-27T09:56:08.257000",
          "content": "<p>So i have finally made a sub for a lgbm trained with 8M rows,  ~LB .762/.763 now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1071976,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-07T17:09:38.927000",
      "content": "<p>Someone has a perfect 1.0 score on the LB (as it seems for now). [it's a glitch i hope, no point downvoting me] <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">bug</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3fea56de822a51982159f43811503050%2FScreenshot%202020-11-07%20at%2010.37.31%20PM.png?generation=1604768976117675&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1071984,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-11-07T17:23:04.803000",
          "content": "<p>Very suspicious…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1071988,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-07T17:27:01.507000",
          "content": "<p>I hope so, not sure why people would downvote, i just shared the information and TBH was amazed to see. Refreshed the page multiple times as well 😅</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1071989,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2020-11-07T17:28:04.103000",
          "content": "<p>This person has kindly shared a bug here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815</a>. Presumably it will just be fixed. It's not foul play, more like a white hat hacking sort of situation from what I can tell.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1071998,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-07T17:35:46.010000",
          "content": "<p>edit-&gt; placeholder [deleted]</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061720,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T08:58:21.313000",
      "content": "",
      "votes": -11,
      "replies": []
    },
    {
      "id": 1061719,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T08:58:00.180000",
      "content": "",
      "votes": -11,
      "replies": []
    },
    {
      "id": 1109706,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-12T01:30:00.810000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1109745,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-12T02:59:22.517000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109772,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-12T03:37:27.487000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1105943,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-08T11:03:03.023000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1115760,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-16T14:57:09.357000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120677,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-21T03:41:29.830000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1120915,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-21T07:41:33.670000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1105028,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-07T13:21:53.483000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1105526,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-08T01:09:38.147000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1105673,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-08T04:36:46.043000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1073485,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-09T15:08:36.427000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1073562,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-09T16:20:30.600000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1073574,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-09T16:33:35.817000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1073578,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-09T16:41:26.523000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1064345,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-30T03:00:01.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1063886,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-29T13:06:14.153000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1063896,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-29T13:18:04.617000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1064553,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-30T08:41:15.093000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1064680,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-30T12:05:39.127000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1058805,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-24T10:03:28.310000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1066778,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-02T03:00:53.600000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1076266,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-12T11:38:37.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1062749,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-28T06:43:10.823000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 1063014,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-28T12:33:03.190000",
          "content": "",
          "votes": -1,
          "replies": [
            {
              "id": 1063086,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-10-28T13:49:50.130000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1060810,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-26T14:51:45.027000",
      "content": "",
      "votes": -5,
      "replies": [
        {
          "id": 1060857,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T15:17:49.560000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060863,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T15:23:29.893000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060868,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T15:30:48.447000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061170,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-26T19:33:40.467000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1058728": "With quite a large dataset at hand and a unique submission format there can be a lot of interesting ideas, models and validation approaches that can be used.\n\nLarge datasets often give a stable performance and so far my experience has been positive.\n\nOpening this thread for anyone willing to share their validation / leaderboard approaches and / or scores.",
    "1112090": "- CV: 0.80664, LB: 0.806\n- Single SAINT-based model\n- without lectures",
    "1059325": "Last LB score increases:\n\nLocal: **0.7609** LB: **0.759**\nLocal: **0.7757** LB: **0.772**\nLocal: **0.7817** LB: **0.780**\nLocal: **0.7866** LB: **0.786**\n\nI'm using 10m rows to train and 2.5m to validate per fold.",
    "1058744": "Regarding CV/LB scores, my last three jumps on the LB were the following:\n\nCV: **0.725487**, LB: **0.761**\nCV: **0.740906**, LB: **0.767**\nCV: **0.747638**, LB: **0.772**\n\nMy approach to validation is pretty simple, just the following to create the training/validation set. I use the remaining 90+% of the data for feature aggregates and the like. Even though the CV score isn't too similar to the LB, an increase in CV typically comes with an increase in LB:\n\n```\nprint(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n```\n\nUpdate:\n\nCV: **0.790365**, LB: **0.787**\n\nI switched up my validation strategy to instead use the one described in https://www.kaggle.com/its7171/cv-strategy. Seems to be relatively accurate from what I've noticed throughout the thread, typically <= 0.003 of a difference beween CV/LB. This score is calculated using 10M rows for training, 2.5M rows for validation. I'm not using any lecture features yet in the implementation.\n\nUpdate #2:\n\nCV: **0.793102**, LB: **0.789**\n\nAdded a few more features, changed the framework of my code throughout to make it much easier to add/test new features. Still using only 10M rows, trying to see if I can hit ~0.795+ before I really look into training with the entire training set, given the estimated 0.01 boost from other people.",
    "1058962": "Local Validation: **0.770**, LB: **0.766**.\nUpdate: Local **0.777**, LB: **0.772**. \n\nThat's with training on ~4% of the data, and validating on ~2%. My validation scheme is set up to mirror the way the test predictions need to be made, so I'm fairly confident in it. I think a bit of a gap is normal especially given the time element.\n\nI strongly agree with what Alex said below about prioritizing finding the right features before using more of the data. 2% is already millions of samples and likely large enough to give reliable model feedback IMO (and even close to the same order as the actual test set). The key thing early in a competition is to be able to iterate on feature/model ideas quickly, and I'm pretty sure training time scales above linear with # of training rows for the typical gradient boosting model.\n\nEdit: see update ",
    "1093890": "single NN, inference for ~3h\nI use [this](https://www.kaggle.com/its7171/cv-strategy) strategy.\nCV:  0.772, LB: 0.776",
    "1064170": "So far CV-LB correlation looks very stable - it's constantly +0.003 on LB compared to CV for all my submissions in the last 2 weeks.",
    "1071447": "CV: .792 \nLB: .791",
    "1117386": "CV: 0.802, LB: 0.802\n\nsingle lightgbm model with 100+ features, nothing found in lectures😓. Anyone boost score from lectures features?",
    "1122135": "CV: 0.808, LB: 0.805\n\nLooking at everyone has better CV-LB gap, I either have a bug in inference or something else is wrong.",
    "1071389": "Local: **0.7602**   LB: **0.773**\nLocal: **0.7699**   LB: **0.782**",
    "1071175": "Local:766\nLB:785",
    "1066774": "Local:763,LB:780\nUpdate: Local:771,LB:790",
    "1066232": "CV:0.798\nLB:0.771\nValidating method: GroupKFold by user_id.",
    "1058825": "My last two submissions were\n\nLocal CV: **0.7637**, LB: **0.767**\nLocal CV: **0.7567**, LB: **0.760**\n\nSo both times my LB was about 0.033 higher than my local CV score. Not sure why but I'll take it.",
    "1082908": "Local: 0.798 LB: not submitted (with time-series split)\n\nThe road to 1st submission is a long way for me 😂",
    "1071733": "CV: 0.777\nLB: 0.784",
    "1065768": "I have at the moment 0.7579 on validation and 0.763 on LB. I'm validating only against unseen users, and that's probably why it is quite higher on LB.\n\nEDIT: I'm using 5 M rows, 33% of them as validation.",
    "1127999": "Same as [this](https://www.kaggle.com/its7171/cv-strategy) cv strategy.\nLocal: 0.8027, LB: 0.805x",
    "1058729": "My last 3 submissions (using holdout validation set):\n\nLocal: **0.756** LB: **0.758**\nLocal: **0.748** LB: **0.751**\nLocal: **0.741** LB: **0.742**",
    "1103563": "CV: 0.778 - LB: 0.779",
    "1090124": "some last models:\nCV 0.756//LB0.759\nCV0.770//LB0.766\nCV0.771//LB0.769\nCV0.774//LB0.772\nCV0.777//LB0.774\nit Looks like the LB score is fluctuating around .003.\nsingle  LGBM LB0.801 sounds so magic,need more features..💪💪💪",
    "1079477": "Local: **0.7638** LB **0.765**\n\nUsing all training data and follow the cross-validation strategy in https://www.kaggle.com/its7171/cv-strategy",
    "1078828": "Local CV: 0.740, LB: 0.727",
    "1075135": "CV: 0.763 LB: 0.771 with simple split by user_id",
    "1064367": "Local: 0.79305\nLB:     0.792",
    "1062828": "Hi,\nLocal : 0.733\nLB: 0.758\n\nPretty classic validation scheme:\nMy training set contains about 10M rows, 3M for validation. Features are aggretated on the remaining rows",
    "1083238": "I'm seeing very strong correlation between CV and LB so far, by using out-of-time validation and validating on the last 2M rows. \nSome of my last models:\nCV: 75.9 // LB: 76.0\nCV: 77.0 // LB 76.9\nCV: 77.6 // LB: 77.4 ",
    "1074330": "Haven't yet made the sub but the below stats is using 1M/8M rows for training and 2.5M rows for validation with 16 feats. Thanks a lot everyone for the inspirations!\n\n```\ntraining's auc: 0.769886    valid's auc: 0.765254 # 1M train, 2.5M rows val\n\ntraining's auc: 0.770383    valid's auc: 0.766083 # 8M train, 2.5M rows val\n```",
    "1071976": "Someone has a perfect 1.0 score on the LB (as it seems for now). [it's a glitch i hope, no point downvoting me] [bug](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3fea56de822a51982159f43811503050%2FScreenshot%202020-11-07%20at%2010.37.31%20PM.png?generation=1604768976117675&alt=media)",
    "1061720": "nice implementation",
    "1061719": "Nice implementation",
    "1109706": "CV 783 ,  LB : 783 , training with 7.5 million rows. However, my last big boost in cv was not a big boost in public leaderboard. I am working in geting a better data representation and looking for bugs in my last submissions.",
    "1105943": "1.CV 0.767 / LB 0.766\n2.CV 0.771 / LB 0.771",
    "1105028": "My CV is 0.756, but LB is 0.553. I can't find the reason.\nWhen I drop variables, which is min/max/std/mean/skew of answer data of each user ID, CV is 0.701 and LB is 0.682.\nI think maybe the additional variables are related with the problem, but I still don't know the exact reason.\nAnyone having the similar experience?",
    "1073485": "CV: 0.76  LB: 0.737\nCV: 0.92  LB: 0.702\n\nI don't get it, it makes no sense. My CV preparation is simple. I get the first 60% of the whole training set, for my training set, then do [ validation = X.groupby('user_id').tail(100) ]. I end up with 13 milion CV and 46 milion Train rows.\n\nI did switch from Pandas to cuDF in my submission notebook to do faster merging and all, so unless cuDF is doing some wild messed up merges, I have no idea why such a wild difference. Almost giving up",
    "1064345": "Im puzzled that with 10M trainset and 2.5M vset,  Local: 0.7683 and LB: 0.69. I think I have got a overfitting dilemma.",
    "1063886": "My LB and Local are radically different, the local one is 0.752 and on the LB is 0.680. Any idea ideas on why is that ? ",
    "1058805": "There is another thread [191413](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191413) on validation",
    "1076266": "",
    "1062749": "",
    "1060810": ""
  }
}