{
  "id": 196942,
  "title": "A few notes...",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196942",
  "author_name": "Silogram",
  "post_date": "2020-11-13T13:41:29.646000",
  "votes": 261,
  "comment_count": 41,
  "views": 0,
  "content": "<p>Since it’s still relatively early in the competition, I’m sharing a few general observations that may help others:</p>\n<ol>\n<li><p>This appears to be a very well-designed competition with an interesting data set. So far, I haven’t found any leaks or ‘magic’. Kudos to the Kaggle and Riiid teams.</p></li>\n<li><p>If you construct your train/valid split to mimic the train/test separation, you should be able to rely on your CV scores. For my last 10 submissions, the LB score has consistently been 0.000-0.001 higher than my CV score.</p></li>\n<li><p>As others have noted, ditch Pandas. Pandas is great if you only need to generate features once, but the overhead is way too big for processing many small dataframes as we need to do during inference. Using Pandas, I was constantly getting submission errors due to using too much time. Since switching to NumPy, the inference time dropped to about 1 hour.</p></li>\n<li><p>Although NN models definitely have an edge with this type of data, GBMs can be competitive. My current solution (LB 0.789) is based on a single LGBM model. The winning solution will almost certainly be an NN/GBM blend.</p></li>\n<li><p>Due to the inference API and the time/memory constraints, this competition is really as much about software engineering as it is about ML. When I started, I had separate feature-generation functions for the training and inference parts. This led to all sorts of irritating and difficult bugs due to small, unintended differences in the two functions. Since switching to a single function used during both training and inference, the pipeline has become much simpler and the bugs have disappeared.</p></li>\n<li><p>One of the trickiest parts about this competition is deciding how to deal with lectures when generating features. Sometimes it makes sense to drop them and sometimes it’s better to include them. </p></li>\n</ol>",
  "messages": [
    {
      "id": 1077337,
      "postDate": "2020-11-13T13:41:29.647Z",
      "content": "<p>Since it’s still relatively early in the competition, I’m sharing a few general observations that may help others:</p>\n<ol>\n<li><p>This appears to be a very well-designed competition with an interesting data set. So far, I haven’t found any leaks or ‘magic’. Kudos to the Kaggle and Riiid teams.</p></li>\n<li><p>If you construct your train/valid split to mimic the train/test separation, you should be able to rely on your CV scores. For my last 10 submissions, the LB score has consistently been 0.000-0.001 higher than my CV score.</p></li>\n<li><p>As others have noted, ditch Pandas. Pandas is great if you only need to generate features once, but the overhead is way too big for processing many small dataframes as we need to do during inference. Using Pandas, I was constantly getting submission errors due to using too much time. Since switching to NumPy, the inference time dropped to about 1 hour.</p></li>\n<li><p>Although NN models definitely have an edge with this type of data, GBMs can be competitive. My current solution (LB 0.789) is based on a single LGBM model. The winning solution will almost certainly be an NN/GBM blend.</p></li>\n<li><p>Due to the inference API and the time/memory constraints, this competition is really as much about software engineering as it is about ML. When I started, I had separate feature-generation functions for the training and inference parts. This led to all sorts of irritating and difficult bugs due to small, unintended differences in the two functions. Since switching to a single function used during both training and inference, the pipeline has become much simpler and the bugs have disappeared.</p></li>\n<li><p>One of the trickiest parts about this competition is deciding how to deal with lectures when generating features. Sometimes it makes sense to drop them and sometimes it’s better to include them. </p></li>\n</ol>",
      "rawMarkdown": "Since it’s still relatively early in the competition, I’m sharing a few general observations that may help others:\n\n1. This appears to be a very well-designed competition with an interesting data set. So far, I haven’t found any leaks or ‘magic’. Kudos to the Kaggle and Riiid teams.\n\n2. If you construct your train/valid split to mimic the train/test separation, you should be able to rely on your CV scores. For my last 10 submissions, the LB score has consistently been 0.000-0.001 higher than my CV score.\n\n3. As others have noted, ditch Pandas. Pandas is great if you only need to generate features once, but the overhead is way too big for processing many small dataframes as we need to do during inference. Using Pandas, I was constantly getting submission errors due to using too much time. Since switching to NumPy, the inference time dropped to about 1 hour.\n\n4. Although NN models definitely have an edge with this type of data, GBMs can be competitive. My current solution (LB 0.789) is based on a single LGBM model. The winning solution will almost certainly be an NN/GBM blend.\n\n5. Due to the inference API and the time/memory constraints, this competition is really as much about software engineering as it is about ML. When I started, I had separate feature-generation functions for the training and inference parts. This led to all sorts of irritating and difficult bugs due to small, unintended differences in the two functions. Since switching to a single function used during both training and inference, the pipeline has become much simpler and the bugs have disappeared.\n\n6. One of the trickiest parts about this competition is deciding how to deal with lectures when generating features. Sometimes it makes sense to drop them and sometimes it’s better to include them. ",
      "votes": 260
    },
    {
      "id": 1077363,
      "postDate": "2020-11-13T14:14:34.870Z",
      "content": "<p>+1 for well-designed competition. I really like the API part on inference, even if it's not easy to tune/troubleshoot, it really simulates real life cycle with unseen data and student progress. Also, no full <code>submission.csv</code> file prevent us from crazy/blindly blend kernels.</p>\n<p>LGB gives better results than NN transformer so far for me (CV 0.77 vs 0.75). But I'm working to make my transformer model better (like SAINT).</p>",
      "rawMarkdown": "+1 for well-designed competition. I really like the API part on inference, even if it's not easy to tune/troubleshoot, it really simulates real life cycle with unseen data and student progress. Also, no full `submission.csv` file prevent us from crazy/blindly blend kernels.\n\nLGB gives better results than NN transformer so far for me (CV 0.77 vs 0.75). But I'm working to make my transformer model better (like SAINT).",
      "votes": 13
    },
    {
      "id": 1083256,
      "postDate": "2020-11-18T19:09:13.530Z",
      "content": "<p>Thanks for sharing! I have some questions,</p>\n<p>2) currently my cv and lb correlates but not as close as yours ~ 0.005. I was wondering do you do the cv on the complete train data? i found processing and training using the whole data takes quite some time at least for my pipeline so the validation time is long, when i use less e.g. 5% data, the performance is much lower than using all (i guess not surprising since more data gives a more precise target encoding), if you are using less data for fast validation, how are sure if certain features is improving your model performance or the scale of improvement when it translates to complete train data.</p>\n<p>3) i'd like to ditch pandas too but didnt find a good way to start. One thing conceptually bothers me is that, how do you deal with groupby operations or others (e.g. fillna etc that pandas offers out of box) when using numpy? I was imagining if i switched to numpy and used the same function processing test to process train, then i would probably need using for loop and keep a dictionary around (and numba might help with the for loop), it feels like its gonna be really messy to me without using these out-of-box functionalities pandas  provides (i.e. perhaps i need to write a lot customized function in numpy?)</p>\n<p>4) this is probably a bit sensitive, but if you dont mind, around how many features are you using to achieve such strong single model performance? and is this due to one or two really strong features or a set of features?</p>\n<p>6) dealing with lectures rows in my pipeline often complicates my code, it would have been easier without lectures. there are only ~ 2% lectures, maybe completely dropping them would not matter too much for feature generation, are you currently using them in you feature generations?</p>",
      "rawMarkdown": "Thanks for sharing! I have some questions,\n\n2) currently my cv and lb correlates but not as close as yours ~ 0.005. I was wondering do you do the cv on the complete train data? i found processing and training using the whole data takes quite some time at least for my pipeline so the validation time is long, when i use less e.g. 5% data, the performance is much lower than using all (i guess not surprising since more data gives a more precise target encoding), if you are using less data for fast validation, how are sure if certain features is improving your model performance or the scale of improvement when it translates to complete train data.\n\n3) i'd like to ditch pandas too but didnt find a good way to start. One thing conceptually bothers me is that, how do you deal with groupby operations or others (e.g. fillna etc that pandas offers out of box) when using numpy? I was imagining if i switched to numpy and used the same function processing test to process train, then i would probably need using for loop and keep a dictionary around (and numba might help with the for loop), it feels like its gonna be really messy to me without using these out-of-box functionalities pandas  provides (i.e. perhaps i need to write a lot customized function in numpy?)\n\n4) this is probably a bit sensitive, but if you dont mind, around how many features are you using to achieve such strong single model performance? and is this due to one or two really strong features or a set of features?\n\n6) dealing with lectures rows in my pipeline often complicates my code, it would have been easier without lectures. there are only ~ 2% lectures, maybe completely dropping them would not matter too much for feature generation, are you currently using them in you feature generations?",
      "votes": 8,
      "replies": [
        {
          "id": 1083751,
          "postDate": "2020-11-19T10:37:34.077Z",
          "content": "<p>Those are interesting questions:</p>\n<p>2) This is certainly an issue as feature importance will vary depending on how much of the data is being used during training. If you're training with 5% data and have a feature you're not sure about, you might try training a model with 10% and see if the feature still helps. Also, permutation feature importance is a good way to eliminate questionable features. One aspect of this competition that's especially tricky is that we don't know what percentage of users in the test set overlap with users in the train set. This also affects validation accuracy and feature importance.</p>\n<p>3) Yes, pandas is very convenient. And in fact, I still do use pandas for some feature generation. I would recommend finding the parts of your pipeline that are slowest and try converting them to numpy one at a time. Lots of little errors can emerge during this process so I find it helpful to make a submission after each change even if I know it won't help my score, just to make sure that I haven't introduced any bugs.</p>\n<p>4) Currently, my model has about 45 features.</p>\n<p>6) I do use lectures for some features. It's hard to say how much lower my score would be if I ignored lectures. </p>",
          "rawMarkdown": "Those are interesting questions:\n\n2) This is certainly an issue as feature importance will vary depending on how much of the data is being used during training. If you're training with 5% data and have a feature you're not sure about, you might try training a model with 10% and see if the feature still helps. Also, permutation feature importance is a good way to eliminate questionable features. One aspect of this competition that's especially tricky is that we don't know what percentage of users in the test set overlap with users in the train set. This also affects validation accuracy and feature importance.\n\n3) Yes, pandas is very convenient. And in fact, I still do use pandas for some feature generation. I would recommend finding the parts of your pipeline that are slowest and try converting them to numpy one at a time. Lots of little errors can emerge during this process so I find it helpful to make a submission after each change even if I know it won't help my score, just to make sure that I haven't introduced any bugs.\n\n4) Currently, my model has about 45 features.\n\n6) I do use lectures for some features. It's hard to say how much lower my score would be if I ignored lectures. ",
          "votes": 14
        },
        {
          "id": 1086918,
          "postDate": "2020-11-22T07:15:27.393Z",
          "content": "<p>Is your inference time is still about 1 hour with 45 features and new API ?</p>",
          "rawMarkdown": "Is your inference time is still about 1 hour with 45 features and new API ?"
        },
        {
          "id": 1088212,
          "postDate": "2020-11-23T12:40:29.870Z",
          "content": "<p>No, it's longer now.</p>",
          "rawMarkdown": "No, it's longer now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1081449,
      "postDate": "2020-11-17T05:42:55.447Z",
      "content": "<p>Thank You 🙏.This is my very 1st comment on Kaggle.😊</p>",
      "rawMarkdown": "Thank You 🙏.This is my very 1st comment on Kaggle.😊",
      "votes": 6
    },
    {
      "id": 1080958,
      "postDate": "2020-11-16T17:26:31.253Z",
      "content": "<p>Thanks for your very interesting remarks !<br>\nAs the column <em>timestamp</em> is actually a \"relative\" timestamp for each user (it should be named <em>user_timestamp</em>), we cannot order all the interactions on a single absolute timeline, and I don't think the <em>row_id</em> can help for this.<br>\nDo you confirm this ?<br>\nThis is problematic if we want to compute features involving multiple users because <strong>we risk using data from the future</strong> (future data leakage).</p>",
      "rawMarkdown": "Thanks for your very interesting remarks !\nAs the column *timestamp* is actually a \"relative\" timestamp for each user (it should be named *user_timestamp*), we cannot order all the interactions on a single absolute timeline, and I don't think the *row_id* can help for this.\nDo you confirm this ?\nThis is problematic if we want to compute features involving multiple users because **we risk using data from the future** (future data leakage).",
      "votes": 1,
      "replies": [
        {
          "id": 1081560,
          "postDate": "2020-11-17T07:48:35.667Z",
          "content": "<p>Yes, this is the way I understand the data, that there is no way to know whether user A answering a question came before or after user B answering a question. This prevents us from building certain kinds of features, such as question accuracy 'drift' (e.g., dis a question become easier or harder over time) or identifying clusters of users who may be sharing answers. But I don't think this leads to any problem of future data leakage as long as the training data is always constrained to only include information about past questions per user.</p>",
          "rawMarkdown": "Yes, this is the way I understand the data, that there is no way to know whether user A answering a question came before or after user B answering a question. This prevents us from building certain kinds of features, such as question accuracy 'drift' (e.g., dis a question become easier or harder over time) or identifying clusters of users who may be sharing answers. But I don't think this leads to any problem of future data leakage as long as the training data is always constrained to only include information about past questions per user.",
          "votes": 6
        }
      ]
    },
    {
      "id": 1078894,
      "postDate": "2020-11-15T12:11:30.970Z",
      "content": "<p>It really hurts me that i have to wait for 9 hours to see whether i have a successful submission or not. And i am not even sure why it's so slow either. It's really frustrating :(</p>",
      "rawMarkdown": "It really hurts me that i have to wait for 9 hours to see whether i have a successful submission or not. And i am not even sure why it's so slow either. It's really frustrating :(",
      "votes": 1,
      "replies": [
        {
          "id": 1079236,
          "postDate": "2020-11-15T19:32:57.523Z",
          "content": "<p>try to measure:</p>\n<pre><code>- your pipeline timing, but replace the model prediction by a dummy prediction (=0.5)\n- without your pipeline, but only your model prediction timing per batch (about 30 ~ 40 rows)\n</code></pre>\n<p>you might get a better idea where to optimize</p>",
          "rawMarkdown": "try to measure:\n\n    - your pipeline timing, but replace the model prediction by a dummy prediction (=0.5)\n    - without your pipeline, but only your model prediction timing per batch (about 30 ~ 40 rows)\n\nyou might get a better idea where to optimize",
          "votes": 2
        },
        {
          "id": 1079426,
          "postDate": "2020-11-16T03:51:57.423Z",
          "content": "<p>Thanks for the suggestions, I will try them. But the sad part for me is that my inference code passes <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> iter_test emulator pretty much yet it <strong>still</strong> fails on the submission and i am having a hard time since last 3 days trying to see why, It fails pretty quick though, like just after 1-1.5 mins pretty much on the submission.</p>\n<p>Will re-check everything once more and then have to make the nbs public to seek help;</p>\n<pre><code># on making preds with model on sample test_df\nCPU times: user 7.1 s, sys: 101 ms, total: 7.2 s\nWall time: 4.92 s\n</code></pre>",
          "rawMarkdown": "Thanks for the suggestions, I will try them. But the sad part for me is that my inference code passes @its7171 iter_test emulator pretty much yet it **still** fails on the submission and i am having a hard time since last 3 days trying to see why, It fails pretty quick though, like just after 1-1.5 mins pretty much on the submission.\n\nWill re-check everything once more and then have to make the nbs public to seek help;\n\n```\n# on making preds with model on sample test_df\nCPU times: user 7.1 s, sys: 101 ms, total: 7.2 s\nWall time: 4.92 s\n```"
        },
        {
          "id": 1079451,
          "postDate": "2020-11-16T04:37:23.980Z",
          "content": "<p>That sounds like an error related to something in the actual iter test vs. the mock up. Have you modified your state to exclude at least some of the user and content ids included in the mock?</p>",
          "rawMarkdown": "That sounds like an error related to something in the actual iter test vs. the mock up. Have you modified your state to exclude at least some of the user and content ids included in the mock?"
        },
        {
          "id": 1079462,
          "postDate": "2020-11-16T04:58:53.600Z",
          "content": "<p>Yep, that shouldn't break the snip i believe. Would be helpful if someone can just spend 5 mins on <a href=\"https://www.kaggle.com/adityaecdrid/lgbm-pipeline-test\" target=\"_blank\">this</a>, I have open sourced the snip and my <a href=\"https://www.kaggle.com/adityaecdrid/cv-2-cache-files-10th-nov-2020\" target=\"_blank\">cache files</a>. It does has a feature which i haven't yet seen on public kernels at-least and it gave me a nice boost as well.</p>\n<p>Ty! looking forward to your comments.</p>",
          "rawMarkdown": "Yep, that shouldn't break the snip i believe. Would be helpful if someone can just spend 5 mins on [this](https://www.kaggle.com/adityaecdrid/lgbm-pipeline-test), I have open sourced the snip and my [cache files](https://www.kaggle.com/adityaecdrid/cv-2-cache-files-10th-nov-2020). It does has a feature which i haven't yet seen on public kernels at-least and it gave me a nice boost as well.\n\nTy! looking forward to your comments."
        },
        {
          "id": 1079565,
          "postDate": "2020-11-16T08:11:57.540Z",
          "content": "<p>Okay, so I guess probably my bug was having non-integral keys of a dictionary. Notebook's running now, will update if it succeeds.</p>",
          "rawMarkdown": "Okay, so I guess probably my bug was having non-integral keys of a dictionary. Notebook's running now, will update if it succeeds."
        },
        {
          "id": 1079648,
          "postDate": "2020-11-16T10:52:16.330Z",
          "content": "<p>There's a lot going on there and looks like time is tight/over. If you wanted to try something different, <a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">this</a> is working in under two hours and should accommodate  most of your features. You basically load your state tables into an sqllite database and let it do its thing, sidestepping pandas for everything except reading and submitting batches.</p>",
          "rawMarkdown": "There's a lot going on there and looks like time is tight/over. If you wanted to try something different, [this](https://www.kaggle.com/calebeverett/riiid-submit) is working in under two hours and should accommodate  most of your features. You basically load your state tables into an sqllite database and let it do its thing, sidestepping pandas for everything except reading and submitting batches.",
          "votes": 1
        },
        {
          "id": 1079680,
          "postDate": "2020-11-16T11:40:36.303Z",
          "content": "<p>Thanks a lot for the heads up! I m looking for other options for sure. I am not that strong in SQL unfortunately 🤐and this comp suites a dB like thing.</p>\n<p>It's been 4 hours now since i submitted and it's still running 🤐</p>",
          "rawMarkdown": "Thanks a lot for the heads up! I m looking for other options for sure. I am not that strong in SQL unfortunately 🤐and this comp suites a dB like thing.\n\nIt's been 4 hours now since i submitted and it's still running 🤐"
        }
      ]
    },
    {
      "id": 1077577,
      "postDate": "2020-11-13T17:58:14.310Z",
      "content": "<p>Thanks for sharing. I am currious about how to use Numpy instead of pd.merge?</p>",
      "rawMarkdown": "Thanks for sharing. I am currious about how to use Numpy instead of pd.merge?",
      "votes": 1,
      "replies": [
        {
          "id": 1077610,
          "postDate": "2020-11-13T18:25:16.153Z",
          "content": "<p>We can iteratively compute stats and treat numpy vector's column as a column of the matrix and just re-assign that particular column back to the dataframe after doing all the ops at that index?</p>",
          "rawMarkdown": "We can iteratively compute stats and treat numpy vector's column as a column of the matrix and just re-assign that particular column back to the dataframe after doing all the ops at that index?",
          "votes": 1
        },
        {
          "id": 1077797,
          "postDate": "2020-11-13T23:52:59.603Z",
          "content": "<p>In my opinion, pd.merge is super slow/memory hungry and better ditch it. some dict + list comprehension can speed up significantly.</p>",
          "rawMarkdown": "In my opinion, pd.merge is super slow/memory hungry and better ditch it. some dict + list comprehension can speed up significantly.",
          "votes": 1
        },
        {
          "id": 1077900,
          "postDate": "2020-11-14T04:34:15.843Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<p>I applied it this way, is that correct or do you mean a different way?</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
          "rawMarkdown": "@adityaecdrid \n\nI applied it this way, is that correct or do you mean a different way?\n\n```\nenv = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n```"
        },
        {
          "id": 1078303,
          "postDate": "2020-11-14T15:54:44.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a> Thanks for your suggestions. I have try transform the some features from dataframe to dict, but they also cause merrory issue. </p>",
          "rawMarkdown": "@yl1202 Thanks for your suggestions. I have try transform the some features from dataframe to dict, but they also cause merrory issue. "
        },
        {
          "id": 1078331,
          "postDate": "2020-11-14T16:36:15.913Z",
          "content": "<p>Fully agree with <a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a>, from best to worse performances: </p>\n<ul>\n<li>Python list</li>\n<li>Numpy</li>\n<li>Pandas</li>\n</ul>",
          "rawMarkdown": "Fully agree with @yl1202, from best to worse performances: \n- Python list\n- Numpy\n- Pandas",
          "votes": 2
        }
      ]
    },
    {
      "id": 1077556,
      "postDate": "2020-11-13T17:41:42.887Z",
      "content": "<p>Thanks for the high level comments and it is very helpful!</p>",
      "rawMarkdown": "Thanks for the high level comments and it is very helpful!\n\n",
      "votes": 1
    },
    {
      "id": 1077459,
      "postDate": "2020-11-13T16:20:26.313Z",
      "content": "<p>How do you currently split your train/test data to get such reliable CV scores? I could not get a reliable CV score so far…</p>",
      "rawMarkdown": "How do you currently split your train/test data to get such reliable CV scores? I could not get a reliable CV score so far...",
      "votes": 1,
      "replies": [
        {
          "id": 1077507,
          "postDate": "2020-11-13T16:49:00.340Z",
          "content": "<p>I don't want to give away my precise methodology, but I'll say that it follows these guidelines:</p>\n<ol>\n<li>I want to train and validate on complete histories of sample users.</li>\n<li>The validation set should have a mix of new users and users that appear in the train set.</li>\n<li>For common users, all questions in the valid set should follow questions in the train set.</li>\n</ol>",
          "rawMarkdown": "I don't want to give away my precise methodology, but I'll say that it follows these guidelines:\n1. I want to train and validate on complete histories of sample users.\n2. The validation set should have a mix of new users and users that appear in the train set.\n3. For common users, all questions in the valid set should follow questions in the train set.",
          "votes": 28
        },
        {
          "id": 1077603,
          "postDate": "2020-11-13T18:19:24.590Z",
          "content": "<p>I mostly spent my time in improving feature engineering as quick as possible so my validation scheme is pretty basic; there is always a gap between validation and LB. However the changes are consistent (i.e. +0.005 in val leads to +0.005 in LB)<br>\nI did not manage yet to get a good validation scheme too</p>",
          "rawMarkdown": "I mostly spent my time in improving feature engineering as quick as possible so my validation scheme is pretty basic; there is always a gap between validation and LB. However the changes are consistent (i.e. +0.005 in val leads to +0.005 in LB)\nI did not manage yet to get a good validation scheme too",
          "votes": 2
        },
        {
          "id": 1077801,
          "postDate": "2020-11-13T23:57:45.817Z",
          "content": "<p>I assume that in the test data api, new lecture/question entries to an existing user are fed in chronological order. is this the correct assumption? or test entry can take place anytime in an user's history?</p>",
          "rawMarkdown": "\nI assume that in the test data api, new lecture/question entries to an existing user are fed in chronological order. is this the correct assumption? or test entry can take place anytime in an user's history?"
        }
      ]
    },
    {
      "id": 1077371,
      "postDate": "2020-11-13T14:25:34.250Z",
      "content": "<p>Yep! This comp is gonna be a lot of fun as people are gonna bring in the best of SDE and write optimised code to the best possible extent as well unlike in past comps (few are exceptions though)! Plus CV/LB is in complete sync as well!</p>",
      "rawMarkdown": "Yep! This comp is gonna be a lot of fun as people are gonna bring in the best of SDE and write optimised code to the best possible extent as well unlike in past comps (few are exceptions though)! Plus CV/LB is in complete sync as well!",
      "votes": 1
    },
    {
      "id": 1090647,
      "postDate": "2020-11-25T14:07:16.213Z",
      "content": "<p>Thanks for sharing , it's very helpful <br>\nI have some more questions : </p>\n<p>1-how we can deal  with prior elpsed time <br>\n2-do you think it's better to shift them and deal with alpsed time  instead of dealing with the prior.</p>",
      "rawMarkdown": "Thanks for sharing , it's very helpful \nI have some more questions : \n\n1-how we can deal  with prior elpsed time \n2-do you think it's better to shift them and deal with alpsed time  instead of dealing with the prior."
    },
    {
      "id": 1085532,
      "postDate": "2020-11-21T00:37:44.270Z",
      "content": "<p>Thanks for sharing. one questio though:<br>\nwhen you say ditch Pandas do you mean for the submission part or also for exploring and visualizing the data?</p>",
      "rawMarkdown": "Thanks for sharing. one questio though:\nwhen you say ditch Pandas do you mean for the submission part or also for exploring and visualizing the data?",
      "replies": [
        {
          "id": 1085770,
          "postDate": "2020-11-21T07:34:39.867Z",
          "content": "<p>I mean especially during inference. For exploring and visualizing data, pandas is invaluable. And even in the inference section, I do use pandas here and there, but I've replaced it where it's much slower than numpy.</p>",
          "rawMarkdown": "I mean especially during inference. For exploring and visualizing data, pandas is invaluable. And even in the inference section, I do use pandas here and there, but I've replaced it where it's much slower than numpy."
        }
      ]
    },
    {
      "id": 1083939,
      "postDate": "2020-11-19T14:57:34.953Z",
      "content": "<p>I think it is a legit resource of information! Thank you for your insights! </p>",
      "rawMarkdown": "I think it is a legit resource of information! Thank you for your insights! "
    },
    {
      "id": 1079568,
      "postDate": "2020-11-16T08:16:23.977Z",
      "content": "<p>Sir, I have got into trouble on submission time/space error. I have used pandas merge and np.array comparation.<br>\nI have not noticed any space restrict.  Should that only be executing time error but not space insufficient?<br>\nwhich could be supplement for pandas merge?</p>",
      "rawMarkdown": "Sir, I have got into trouble on submission time/space error. I have used pandas merge and np.array comparation.\nI have not noticed any space restrict.  Should that only be executing time error but not space insufficient?\nwhich could be supplement for pandas merge?",
      "replies": [
        {
          "id": 1081562,
          "postDate": "2020-11-17T07:49:49.140Z",
          "content": "<p>I'm pretty sure our kernels are limited to 16GB, so you may be bumping up against that limit.</p>",
          "rawMarkdown": "I'm pretty sure our kernels are limited to 16GB, so you may be bumping up against that limit.",
          "votes": 1
        },
        {
          "id": 1082928,
          "postDate": "2020-11-18T11:45:13.160Z",
          "content": "<p>Thank you for your valuable advice!</p>",
          "rawMarkdown": "Thank you for your valuable advice!"
        }
      ]
    },
    {
      "id": 1079072,
      "postDate": "2020-11-15T15:27:14.980Z",
      "content": "<p>This is very helpful for me (I am about a week into the competition). I am really curious about point 6 if you would be willing to elaborate any more. I am assuming something around time of lecture needs to be considered, currently I am treating all features time agnostic which needs to be corrected.</p>",
      "rawMarkdown": "This is very helpful for me (I am about a week into the competition). I am really curious about point 6 if you would be willing to elaborate any more. I am assuming something around time of lecture needs to be considered, currently I am treating all features time agnostic which needs to be corrected.",
      "replies": [
        {
          "id": 1079556,
          "postDate": "2020-11-16T07:46:41.713Z",
          "content": "<p>Time is important ;-)</p>",
          "rawMarkdown": "Time is important ;-)",
          "votes": 5
        },
        {
          "id": 1081299,
          "postDate": "2020-11-17T01:01:32.260Z",
          "content": "<p>Sounds good! 👍 Thanks again for an awesome post.</p>",
          "rawMarkdown": "Sounds good! 👍 Thanks again for an awesome post."
        }
      ]
    },
    {
      "id": 1083378,
      "postDate": "2020-11-18T22:51:40.827Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1097257,
      "postDate": "2020-12-01T01:54:17.493Z",
      "content": "<p>Thank you for sharing! </p>",
      "rawMarkdown": "Thank you for sharing! "
    },
    {
      "id": 1083683,
      "postDate": "2020-11-19T08:35:25.443Z",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing "
    }
  ],
  "comments": [
    {
      "id": 1077363,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-13T14:14:34.870000",
      "content": "<p>+1 for well-designed competition. I really like the API part on inference, even if it's not easy to tune/troubleshoot, it really simulates real life cycle with unseen data and student progress. Also, no full <code>submission.csv</code> file prevent us from crazy/blindly blend kernels.</p>\n<p>LGB gives better results than NN transformer so far for me (CV 0.77 vs 0.75). But I'm working to make my transformer model better (like SAINT).</p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 1083256,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "2020-11-18T19:09:13.530000",
      "content": "<p>Thanks for sharing! I have some questions,</p>\n<p>2) currently my cv and lb correlates but not as close as yours ~ 0.005. I was wondering do you do the cv on the complete train data? i found processing and training using the whole data takes quite some time at least for my pipeline so the validation time is long, when i use less e.g. 5% data, the performance is much lower than using all (i guess not surprising since more data gives a more precise target encoding), if you are using less data for fast validation, how are sure if certain features is improving your model performance or the scale of improvement when it translates to complete train data.</p>\n<p>3) i'd like to ditch pandas too but didnt find a good way to start. One thing conceptually bothers me is that, how do you deal with groupby operations or others (e.g. fillna etc that pandas offers out of box) when using numpy? I was imagining if i switched to numpy and used the same function processing test to process train, then i would probably need using for loop and keep a dictionary around (and numba might help with the for loop), it feels like its gonna be really messy to me without using these out-of-box functionalities pandas  provides (i.e. perhaps i need to write a lot customized function in numpy?)</p>\n<p>4) this is probably a bit sensitive, but if you dont mind, around how many features are you using to achieve such strong single model performance? and is this due to one or two really strong features or a set of features?</p>\n<p>6) dealing with lectures rows in my pipeline often complicates my code, it would have been easier without lectures. there are only ~ 2% lectures, maybe completely dropping them would not matter too much for feature generation, are you currently using them in you feature generations?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1083751,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-19T10:37:34.077000",
          "content": "<p>Those are interesting questions:</p>\n<p>2) This is certainly an issue as feature importance will vary depending on how much of the data is being used during training. If you're training with 5% data and have a feature you're not sure about, you might try training a model with 10% and see if the feature still helps. Also, permutation feature importance is a good way to eliminate questionable features. One aspect of this competition that's especially tricky is that we don't know what percentage of users in the test set overlap with users in the train set. This also affects validation accuracy and feature importance.</p>\n<p>3) Yes, pandas is very convenient. And in fact, I still do use pandas for some feature generation. I would recommend finding the parts of your pipeline that are slowest and try converting them to numpy one at a time. Lots of little errors can emerge during this process so I find it helpful to make a submission after each change even if I know it won't help my score, just to make sure that I haven't introduced any bugs.</p>\n<p>4) Currently, my model has about 45 features.</p>\n<p>6) I do use lectures for some features. It's hard to say how much lower my score would be if I ignored lectures. </p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1086918,
          "author_name": "ln",
          "author_url": "",
          "post_date": "2020-11-22T07:15:27.393000",
          "content": "<p>Is your inference time is still about 1 hour with 45 features and new API ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1088212,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-23T12:40:29.870000",
          "content": "<p>No, it's longer now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1081449,
      "author_name": "Beast",
      "author_url": "",
      "post_date": "2020-11-17T05:42:55.447000",
      "content": "<p>Thank You 🙏.This is my very 1st comment on Kaggle.😊</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1080958,
      "author_name": "ISMAX",
      "author_url": "",
      "post_date": "2020-11-16T17:26:31.253000",
      "content": "<p>Thanks for your very interesting remarks !<br>\nAs the column <em>timestamp</em> is actually a \"relative\" timestamp for each user (it should be named <em>user_timestamp</em>), we cannot order all the interactions on a single absolute timeline, and I don't think the <em>row_id</em> can help for this.<br>\nDo you confirm this ?<br>\nThis is problematic if we want to compute features involving multiple users because <strong>we risk using data from the future</strong> (future data leakage).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1081560,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-17T07:48:35.667000",
          "content": "<p>Yes, this is the way I understand the data, that there is no way to know whether user A answering a question came before or after user B answering a question. This prevents us from building certain kinds of features, such as question accuracy 'drift' (e.g., dis a question become easier or harder over time) or identifying clusters of users who may be sharing answers. But I don't think this leads to any problem of future data leakage as long as the training data is always constrained to only include information about past questions per user.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1078894,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-15T12:11:30.970000",
      "content": "<p>It really hurts me that i have to wait for 9 hours to see whether i have a successful submission or not. And i am not even sure why it's so slow either. It's really frustrating :(</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1079236,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-15T19:32:57.523000",
          "content": "<p>try to measure:</p>\n<pre><code>- your pipeline timing, but replace the model prediction by a dummy prediction (=0.5)\n- without your pipeline, but only your model prediction timing per batch (about 30 ~ 40 rows)\n</code></pre>\n<p>you might get a better idea where to optimize</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1079426,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-16T03:51:57.423000",
          "content": "<p>Thanks for the suggestions, I will try them. But the sad part for me is that my inference code passes <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> iter_test emulator pretty much yet it <strong>still</strong> fails on the submission and i am having a hard time since last 3 days trying to see why, It fails pretty quick though, like just after 1-1.5 mins pretty much on the submission.</p>\n<p>Will re-check everything once more and then have to make the nbs public to seek help;</p>\n<pre><code># on making preds with model on sample test_df\nCPU times: user 7.1 s, sys: 101 ms, total: 7.2 s\nWall time: 4.92 s\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1079451,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-16T04:37:23.980000",
          "content": "<p>That sounds like an error related to something in the actual iter test vs. the mock up. Have you modified your state to exclude at least some of the user and content ids included in the mock?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1079462,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-16T04:58:53.600000",
          "content": "<p>Yep, that shouldn't break the snip i believe. Would be helpful if someone can just spend 5 mins on <a href=\"https://www.kaggle.com/adityaecdrid/lgbm-pipeline-test\" target=\"_blank\">this</a>, I have open sourced the snip and my <a href=\"https://www.kaggle.com/adityaecdrid/cv-2-cache-files-10th-nov-2020\" target=\"_blank\">cache files</a>. It does has a feature which i haven't yet seen on public kernels at-least and it gave me a nice boost as well.</p>\n<p>Ty! looking forward to your comments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1079565,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-16T08:11:57.540000",
          "content": "<p>Okay, so I guess probably my bug was having non-integral keys of a dictionary. Notebook's running now, will update if it succeeds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1079648,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-11-16T10:52:16.330000",
          "content": "<p>There's a lot going on there and looks like time is tight/over. If you wanted to try something different, <a href=\"https://www.kaggle.com/calebeverett/riiid-submit\" target=\"_blank\">this</a> is working in under two hours and should accommodate  most of your features. You basically load your state tables into an sqllite database and let it do its thing, sidestepping pandas for everything except reading and submitting batches.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1079680,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-16T11:40:36.303000",
          "content": "<p>Thanks a lot for the heads up! I m looking for other options for sure. I am not that strong in SQL unfortunately 🤐and this comp suites a dB like thing.</p>\n<p>It's been 4 hours now since i submitted and it's still running 🤐</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1077577,
      "author_name": "Ethan",
      "author_url": "",
      "post_date": "2020-11-13T17:58:14.310000",
      "content": "<p>Thanks for sharing. I am currious about how to use Numpy instead of pd.merge?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1077610,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-13T18:25:16.153000",
          "content": "<p>We can iteratively compute stats and treat numpy vector's column as a column of the matrix and just re-assign that particular column back to the dataframe after doing all the ops at that index?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1077797,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-11-13T23:52:59.603000",
          "content": "<p>In my opinion, pd.merge is super slow/memory hungry and better ditch it. some dict + list comprehension can speed up significantly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1077900,
          "author_name": "Mnka",
          "author_url": "",
          "post_date": "2020-11-14T04:34:15.843000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<p>I applied it this way, is that correct or do you mean a different way?</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078303,
          "author_name": "Ethan",
          "author_url": "",
          "post_date": "2020-11-14T15:54:44.350000",
          "content": "<p><a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a> Thanks for your suggestions. I have try transform the some features from dataframe to dict, but they also cause merrory issue. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1078331,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-14T16:36:15.913000",
          "content": "<p>Fully agree with <a href=\"https://www.kaggle.com/yl1202\" target=\"_blank\">@yl1202</a>, from best to worse performances: </p>\n<ul>\n<li>Python list</li>\n<li>Numpy</li>\n<li>Pandas</li>\n</ul>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1077556,
      "author_name": "cxue34",
      "author_url": "",
      "post_date": "2020-11-13T17:41:42.887000",
      "content": "<p>Thanks for the high level comments and it is very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1077459,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2020-11-13T16:20:26.313000",
      "content": "<p>How do you currently split your train/test data to get such reliable CV scores? I could not get a reliable CV score so far…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1077507,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-13T16:49:00.340000",
          "content": "<p>I don't want to give away my precise methodology, but I'll say that it follows these guidelines:</p>\n<ol>\n<li>I want to train and validate on complete histories of sample users.</li>\n<li>The validation set should have a mix of new users and users that appear in the train set.</li>\n<li>For common users, all questions in the valid set should follow questions in the train set.</li>\n</ol>",
          "votes": 28,
          "replies": []
        },
        {
          "id": 1077603,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-11-13T18:19:24.590000",
          "content": "<p>I mostly spent my time in improving feature engineering as quick as possible so my validation scheme is pretty basic; there is always a gap between validation and LB. However the changes are consistent (i.e. +0.005 in val leads to +0.005 in LB)<br>\nI did not manage yet to get a good validation scheme too</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1077801,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-11-13T23:57:45.817000",
          "content": "<p>I assume that in the test data api, new lecture/question entries to an existing user are fed in chronological order. is this the correct assumption? or test entry can take place anytime in an user's history?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1077371,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-13T14:25:34.250000",
      "content": "<p>Yep! This comp is gonna be a lot of fun as people are gonna bring in the best of SDE and write optimised code to the best possible extent as well unlike in past comps (few are exceptions though)! Plus CV/LB is in complete sync as well!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1090647,
      "author_name": "AIMEN REFOUFI",
      "author_url": "",
      "post_date": "2020-11-25T14:07:16.213000",
      "content": "<p>Thanks for sharing , it's very helpful <br>\nI have some more questions : </p>\n<p>1-how we can deal  with prior elpsed time <br>\n2-do you think it's better to shift them and deal with alpsed time  instead of dealing with the prior.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1085532,
      "author_name": "Omar Bassam",
      "author_url": "",
      "post_date": "2020-11-21T00:37:44.270000",
      "content": "<p>Thanks for sharing. one questio though:<br>\nwhen you say ditch Pandas do you mean for the submission part or also for exploring and visualizing the data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1085770,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-21T07:34:39.867000",
          "content": "<p>I mean especially during inference. For exploring and visualizing data, pandas is invaluable. And even in the inference section, I do use pandas here and there, but I've replaced it where it's much slower than numpy.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1083939,
      "author_name": "Albert Nikanorov",
      "author_url": "",
      "post_date": "2020-11-19T14:57:34.953000",
      "content": "<p>I think it is a legit resource of information! Thank you for your insights! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1079568,
      "author_name": "Zhenghan Chen",
      "author_url": "",
      "post_date": "2020-11-16T08:16:23.977000",
      "content": "<p>Sir, I have got into trouble on submission time/space error. I have used pandas merge and np.array comparation.<br>\nI have not noticed any space restrict.  Should that only be executing time error but not space insufficient?<br>\nwhich could be supplement for pandas merge?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1081562,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-17T07:49:49.140000",
          "content": "<p>I'm pretty sure our kernels are limited to 16GB, so you may be bumping up against that limit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082928,
          "author_name": "Zhenghan Chen",
          "author_url": "",
          "post_date": "2020-11-18T11:45:13.160000",
          "content": "<p>Thank you for your valuable advice!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1079072,
      "author_name": "Chicago JC Kaggle",
      "author_url": "",
      "post_date": "2020-11-15T15:27:14.980000",
      "content": "<p>This is very helpful for me (I am about a week into the competition). I am really curious about point 6 if you would be willing to elaborate any more. I am assuming something around time of lecture needs to be considered, currently I am treating all features time agnostic which needs to be corrected.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1079556,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2020-11-16T07:46:41.713000",
          "content": "<p>Time is important ;-)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1081299,
          "author_name": "Chicago JC Kaggle",
          "author_url": "",
          "post_date": "2020-11-17T01:01:32.260000",
          "content": "<p>Sounds good! 👍 Thanks again for an awesome post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1083378,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-18T22:51:40.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1097257,
      "author_name": "Joao Pinto",
      "author_url": "",
      "post_date": "2020-12-01T01:54:17.493000",
      "content": "<p>Thank you for sharing! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1083683,
      "author_name": "Houssem Ayed",
      "author_url": "",
      "post_date": "2020-11-19T08:35:25.443000",
      "content": "<p>Thanks for sharing </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1077337": "Since it’s still relatively early in the competition, I’m sharing a few general observations that may help others:\n\n1. This appears to be a very well-designed competition with an interesting data set. So far, I haven’t found any leaks or ‘magic’. Kudos to the Kaggle and Riiid teams.\n\n2. If you construct your train/valid split to mimic the train/test separation, you should be able to rely on your CV scores. For my last 10 submissions, the LB score has consistently been 0.000-0.001 higher than my CV score.\n\n3. As others have noted, ditch Pandas. Pandas is great if you only need to generate features once, but the overhead is way too big for processing many small dataframes as we need to do during inference. Using Pandas, I was constantly getting submission errors due to using too much time. Since switching to NumPy, the inference time dropped to about 1 hour.\n\n4. Although NN models definitely have an edge with this type of data, GBMs can be competitive. My current solution (LB 0.789) is based on a single LGBM model. The winning solution will almost certainly be an NN/GBM blend.\n\n5. Due to the inference API and the time/memory constraints, this competition is really as much about software engineering as it is about ML. When I started, I had separate feature-generation functions for the training and inference parts. This led to all sorts of irritating and difficult bugs due to small, unintended differences in the two functions. Since switching to a single function used during both training and inference, the pipeline has become much simpler and the bugs have disappeared.\n\n6. One of the trickiest parts about this competition is deciding how to deal with lectures when generating features. Sometimes it makes sense to drop them and sometimes it’s better to include them. ",
    "1077363": "+1 for well-designed competition. I really like the API part on inference, even if it's not easy to tune/troubleshoot, it really simulates real life cycle with unseen data and student progress. Also, no full `submission.csv` file prevent us from crazy/blindly blend kernels.\n\nLGB gives better results than NN transformer so far for me (CV 0.77 vs 0.75). But I'm working to make my transformer model better (like SAINT).",
    "1083256": "Thanks for sharing! I have some questions,\n\n2) currently my cv and lb correlates but not as close as yours ~ 0.005. I was wondering do you do the cv on the complete train data? i found processing and training using the whole data takes quite some time at least for my pipeline so the validation time is long, when i use less e.g. 5% data, the performance is much lower than using all (i guess not surprising since more data gives a more precise target encoding), if you are using less data for fast validation, how are sure if certain features is improving your model performance or the scale of improvement when it translates to complete train data.\n\n3) i'd like to ditch pandas too but didnt find a good way to start. One thing conceptually bothers me is that, how do you deal with groupby operations or others (e.g. fillna etc that pandas offers out of box) when using numpy? I was imagining if i switched to numpy and used the same function processing test to process train, then i would probably need using for loop and keep a dictionary around (and numba might help with the for loop), it feels like its gonna be really messy to me without using these out-of-box functionalities pandas  provides (i.e. perhaps i need to write a lot customized function in numpy?)\n\n4) this is probably a bit sensitive, but if you dont mind, around how many features are you using to achieve such strong single model performance? and is this due to one or two really strong features or a set of features?\n\n6) dealing with lectures rows in my pipeline often complicates my code, it would have been easier without lectures. there are only ~ 2% lectures, maybe completely dropping them would not matter too much for feature generation, are you currently using them in you feature generations?",
    "1081449": "Thank You 🙏.This is my very 1st comment on Kaggle.😊",
    "1080958": "Thanks for your very interesting remarks !\nAs the column *timestamp* is actually a \"relative\" timestamp for each user (it should be named *user_timestamp*), we cannot order all the interactions on a single absolute timeline, and I don't think the *row_id* can help for this.\nDo you confirm this ?\nThis is problematic if we want to compute features involving multiple users because **we risk using data from the future** (future data leakage).",
    "1078894": "It really hurts me that i have to wait for 9 hours to see whether i have a successful submission or not. And i am not even sure why it's so slow either. It's really frustrating :(",
    "1077577": "Thanks for sharing. I am currious about how to use Numpy instead of pd.merge?",
    "1077556": "Thanks for the high level comments and it is very helpful!\n\n",
    "1077459": "How do you currently split your train/test data to get such reliable CV scores? I could not get a reliable CV score so far...",
    "1077371": "Yep! This comp is gonna be a lot of fun as people are gonna bring in the best of SDE and write optimised code to the best possible extent as well unlike in past comps (few are exceptions though)! Plus CV/LB is in complete sync as well!",
    "1090647": "Thanks for sharing , it's very helpful \nI have some more questions : \n\n1-how we can deal  with prior elpsed time \n2-do you think it's better to shift them and deal with alpsed time  instead of dealing with the prior.",
    "1085532": "Thanks for sharing. one questio though:\nwhen you say ditch Pandas do you mean for the submission part or also for exploring and visualizing the data?",
    "1083939": "I think it is a legit resource of information! Thank you for your insights! ",
    "1079568": "Sir, I have got into trouble on submission time/space error. I have used pandas merge and np.array comparation.\nI have not noticed any space restrict.  Should that only be executing time error but not space insufficient?\nwhich could be supplement for pandas merge?",
    "1079072": "This is very helpful for me (I am about a week into the competition). I am really curious about point 6 if you would be willing to elaborate any more. I am assuming something around time of lecture needs to be considered, currently I am treating all features time agnostic which needs to be corrected.",
    "1083378": "",
    "1097257": "Thank you for sharing! ",
    "1083683": "Thanks for sharing "
  }
}