{
  "id": 210354,
  "title": "9th place solution : 6 Transformers + 2 LightGBMs",
  "url": "/competitions/riiid-test-answer-prediction/writeups/tito-nyanp-9th-place-solution-6-transformers-2-lig",
  "author_name": "",
  "post_date": "2021-01-11T14:39:11.417Z",
  "votes": 74,
  "comment_count": 16,
  "views": 0,
  "content": "<p>First of all, thanks to the hosts for a great competition! This was one of the toughest competitions I have ever entered, but well worth the effort.</p>\n<p>Prediction of our team (tito <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and nyanp <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>)consists of following models.</p>\n<ul>\n<li>tito's transformer: LB 0.813</li>\n<li>nyanp's SAINT+ transformer: LB 0.808</li>\n<li>nyanp's LightGBM: LB 0.806</li>\n</ul>\n<p>Simple blending of these models scored LB 0.814 / private 0.816.</p>\n<h2>Pipeline</h2>\n<p>We converted the entire train.csv data into hdf5, and loaded only the user_id that appeared during inference into a np array (97% RAM savings compared to hold entire training data in RAM). We estimate that the overhead due to I/O in hdf is ~45 minutes. This some overhead allowed us to combine tito's large transformer with nyanp's feature engineering pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532198%2F21ed1bbf93a37f6219fb7dc369761f0e%2Friiid%20pipeline.png?generation=1610287506256559&amp;alt=media\" alt=\"\"></p>\n<h2>Transformer (tito, LB 0.814)</h2>\n<p>This is transformer model with encoder only, based on <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>'s <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">nice kernel</a></p>\n<h3>summary</h3>\n<ul>\n<li>only trained and predicted for answered_correctly of last question of the sequence.</li>\n<li>all features are concatenated (only position encoding is added).</li>\n<li>used lectures, in timestamp order as it is</li>\n<li>Window size 300-600</li>\n<li>batch size 1000</li>\n<li>drop_out 0</li>\n<li>n_encoder_layers 3-5</li>\n<li>augmentation to replace content_ids with dummy ids at a certain rate</li>\n<li>kept only one question in the last task to avoid leaks</li>\n</ul>\n<h3>features</h3>\n<p>embedded or dense was decided by CV.</p>\n<ul>\n<li>(embedded) content id</li>\n<li>(embedded) part id</li>\n<li>(embedded) same task question size</li>\n<li>(dense) answered_correctly</li>\n<li>(dense) had_explanation</li>\n<li>(dense) elapsed time</li>\n<li>(dense) lag time</li>\n<li>(dense) diff of timestamp from the last question</li>\n</ul>\n<h3>combining models</h3>\n<p>To avoid the overhead of calling model.predict() multiple times for ensemble, I made a combined model that links four models.</p>\n<pre><code>inputs = tf.keras.Input(shape=(input_shape, n_features))\nout1 = model1(inputs[:,-window_size1:,:])\nout2 = model2(inputs[:,-window_size2:,:])\nout3 = model3(inputs[:,-window_size3:,:])\nout4 = model4(inputs[:,-window_size4:,:])\ncombo_model = tf.keras.Model(inputs, [out1,out2,out3,out4])\n</code></pre>\n<h2>SAINT+ (nyanp, LB 0.808)</h2>\n<ul>\n<li>d_model = 256</li>\n<li>window_size = 200</li>\n<li>n_layers = 3</li>\n<li>attention dropout = 0.03</li>\n<li>question, part, lag are embed to encoder</li>\n<li>response, elapsed time, has_explanation are embed to decoder</li>\n</ul>\n<p>To prevent leakage, in addition to upper triangular attention mask, I masked the loss in questions other than the beginning of each task_container_id. Questions with the same task_container_id were shuffled in each batch during training, and the loss weights were adjusted to reduce the effect of masks. This mask improved LB by 0.0003.</p>\n<p>(Note: I believe that indirect leaks still exist, but I've spent 80% of my time on the LightGBM implementation and data pipeline, so I couldn't improve it any further)</p>\n<p>Other than that, there is nothing special about this NN. It scored .806 in single, .808 by averaging 2 models.</p>\n<h2>LightGBM (nyanp, LB 0.806)</h2>\n<p>LightGBM models are trained on 264 features. To speed up inference, I fixed the number of trees to 3000 and ensembled two models with different seeds (This is better than single large LGBM in terms of both speed and accuracy). By compiling this model with <a href=\"https://github.com/dmlc/treelite\" target=\"_blank\">treelite</a>, inference time became 3x faster (~10ms/batch, ~10min in total).</p>\n<h3>features</h3>\n<p>I mapped prior_question_* rows with their respective rows by following code and utilized them in some features.</p>\n<pre><code>df['elapsed_time'] = df.groupby('user_id')['prior_question_elapsed_time'].shift(-1)\ndf['elapsed_time'] = df.groupby(['user_id', 'timestamp'])['question_elapsed_time'].transform('last')\n\ndf['has_explanation'] = df.groupby('user_id')['prior_question_had_explanation'].shift(-1)\ndf['has_explanation'] = df.groupby(['user_id', 'timestamp'])['has_explanation'].transform('last')\n</code></pre>\n<p>Here is the list of my features. There is no magic here; no single feature boost CV more than 0.0002. I repeated feature engineering based on well-known techniques and a little bit of domain knowledge.</p>\n<h4>question features</h4>\n<ul>\n<li>count encoding</li>\n<li>target encoding</li>\n<li>number of tags</li>\n<li>one-hot encoding of tag (top-10 frequent tags)</li>\n<li>SVD, LDA, item2vec using user_id x content_id matrix</li>\n<li>LDA, item2vec using user_id x content_id matrix (filtered by answered_correctly == 0)<ul>\n<li>Typical word2vec model is trained on next-word prediction task. By constructing word2vec model over incorrectly answered questions, <br>\nthe latent vectors extracted from the model can be used to capture which incorrect question are likely to co-occur with each other.</li></ul></li>\n<li>10%, 20%, 50%, 80% elapsed time of all users response, correct response, wrong response</li>\n<li>SAINT embedding vector + PCA</li>\n</ul>\n<h4>user features</h4>\n<ul>\n<li>avg/median/max/std elapsed_time</li>\n<li>avg/median/max/std elapsed_time by part</li>\n<li>avg has_explanation flag</li>\n<li>avg has_explanation flag by part</li>\n<li>nunique of question, part, lecture</li>\n<li>cumcount / timestamp</li>\n<li>avg answered_correctly with recent 10/30/100/300 questions, recent 10min/7 days</li>\n<li>avg answered_correctly by part, question, bundle, order of response, question difficulty, question difficulty x part<ul>\n<li>question difficulty: discretize the avg answered_correctly of all users for each question into 10 levels</li></ul></li>\n<li>cumcount by part, question, bundle, question difficulty, question difficulty x part</li>\n<li>lag from 1/2/3/4 step before</li>\n<li>correctness, lag, elapsed_time, has_explanation in the same question last time</li>\n<li>tag-level aggregation features<ul>\n<li>calculate tag-level feature for each user x tag, then aggregate them by min/avg/max</li>\n<li>cumcount of wrong answer, cumcount of correct answer, avg target, lag</li></ul></li>\n<li>(estimated elapsed time) - (avg/median/min/max elapsed time within same part)<ul>\n<li>estimated elapsed time = (timestamp - prev timestamp) / (# of questions within the bundle)</li></ul></li>\n<li>(estimated elapsed time) - (10%, 20%, 50%, 80% elapsed time of all users correct/wrong response)<ul>\n<li>Because TOEIC part1-4 questions are usually answered after listening to the conversation, the correct answers tend to be concentrated immediately after the conversation ends</li></ul></li>\n<li>task_container_id - previous task_container_id<ul>\n<li>I'm not sure why this worked. There might be a difference in the correct rate if people answered from different devices than usual (multi-user?).</li></ul></li>\n<li>part of last lecture</li>\n<li>lag from last lecture</li>\n<li>whether the question contains the same tag as the last lecture</li>\n<li>whether the part is the same as the previous question</li>\n<li>median lag - median elapsed_time</li>\n<li>part of the first problem the user solved</li>\n<li>lag - median lag</li>\n<li>inner-product of user-{correct|incorrect}-question-vector and question-vector<ul>\n<li>user-correct-question-vector: average of LDA vectors for each question that the user answered correctly.</li></ul></li>\n<li>rank of lag compared to the same user's past lag (filtered by answered_correctly == 0, 1 respectively)</li>\n</ul>\n<h3>feature calculation</h3>\n<p>Instead of updating the user dictionary, I calculate user features from scratch for each bacth.</p>\n<p>The numpy array of historical data was loaded from the hdf storage and then split into questions and lectures, which were then wrapped in a pandas-like API and passed to their respective feature functions.</p>\n<pre><code>@feature('lag1.user_lecture')\ndef lag1_user_lecture(df: RiiidData, pool: DataPool) -&gt; np.ndarray:\n    \"\"\"\n    time elapsed from last lectures for each user\n\n    :param df: data in current batch\n    :param pool: cached data storage\n    \"\"\"\n    lag1 = {}\n    for u, t in set(zip(df.questions['user_id'], df.questions['timestamp'])):\n        past_lect = pool.users[u].lectures  # history of lecture\n        if len(past_lect) &gt; 0:\n            lag1[u] = t - past_lect['timestamp'][-1]\n\n    return np.array(list(map(lambda x: lag1.get(x, np.nan), df.questions['user_id'])), dtype=np.float32)\n</code></pre>\n<p>For training, I use the same functions as inference time. This way, there were almost no restrictions on feature creation, and I did not have to worry about bugs of train-test difference.</p>\n<p>The feature functions were frequently benchmarked by a dedicated benchmark script, and functions with high overhead were optimized by various ways (numba, bottleneck and various algorithm improvement) .</p>\n<p>The question features were precomputed and made into a global numpy array of shape (13523, *) and merged into the feature data using fancy index.</p>\n<pre><code>@feature(tuple(f'question_item2vec_{i}' for i in range(20)))\ndef question_item2vec(df: RiiidData, _: DataPool):\n    \"\"\"\n    question embedding vector using item2vec\n    \"\"\"\n    qid = df.questions['content_id']\n\n    ret = QUESTION_ITEM2VEC[qid, :]  # mere fancy index, faster than pd.merge\n\n    return ret\n</code></pre>\n<h2>Other ideas</h2>\n<ul>\n<li>TTA on SAINT+ by np.roll (it did improve SAINT+ by 0.001+, but we couldn't include it because of submission timeout)</li>\n<li>linear blending based on public LB labels (timeout, too)</li>\n</ul>\n<h2>Feedback on Time-Series API Competition</h2>\n<p>Although the competition was well-designed, our team still found that the time-series API allowed us to obtain private score information with 1-bit probing even after the known vulnerability was fixed.</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\n\nground_truth = []\nprediction = []\nthreshold = 0.810\n\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    ground_truth.extend(list(test_df['answered_correctly'].values))\n\n    predicted = model.predict(...)\n\n    prediction.extend(list(predicted))\n\n    if len(prediction) &gt;= 2500000:\n        private_auc = roc_auc_score(prediction[500000:], ground_truth[500000:])\n        if private_auc &lt; threshold:\n            raise RuntimeError()\n\n# private auc of this submission is guaranteed to exceed .810 if submission is succeeded\n</code></pre>\n<p>By using this probing, we can select the final submission with the highest private score, or determine the best<br>\n ensemble weight for private score by using the hill-climbing method.</p>\n<p>We contacted the Kaggle Team and asked them if this is legal, and they said it was a \"low signal matter\". We still think this is a gray-area and decided not to use this probing.</p>\n<p>We think it wasn't critical in this competition, but if there is a future competition with the same format but without AUC metrics, the hack with the \"magic coefficient\" could improve the ranking significantly. If a competition with the same format is held on Kaggle in the future, we suggest the Kaggle team to fix this problem (e.g. put a dummy value in the private label during the competition).</p>",
  "messages": [
    {
      "id": "1147440",
      "postDate": "01/10/2021 14:06:57",
      "content": "<p>First of all, thanks to the hosts for a great competition! This was one of the toughest competitions I have ever entered, but well worth the effort.</p>\n<p>Prediction of our team (tito <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and nyanp <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>)consists of following models.</p>\n<ul>\n<li>tito's transformer: LB 0.813</li>\n<li>nyanp's SAINT+ transformer: LB 0.808</li>\n<li>nyanp's LightGBM: LB 0.806</li>\n</ul>\n<p>Simple blending of these models scored LB 0.814 / private 0.816.</p>\n<h2>Pipeline</h2>\n<p>We converted the entire train.csv data into hdf5, and loaded only the user_id that appeared during inference into a np array (97% RAM savings compared to hold entire training data in RAM). We estimate that the overhead due to I/O in hdf is ~45 minutes. This some overhead allowed us to combine tito's large transformer with nyanp's feature engineering pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532198%2F21ed1bbf93a37f6219fb7dc369761f0e%2Friiid%20pipeline.png?generation=1610287506256559&amp;alt=media\" alt=\"\"></p>\n<h2>Transformer (tito, LB 0.814)</h2>\n<p>This is transformer model with encoder only, based on <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>'s <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">nice kernel</a></p>\n<h3>summary</h3>\n<ul>\n<li>only trained and predicted for answered_correctly of last question of the sequence.</li>\n<li>all features are concatenated (only position encoding is added).</li>\n<li>used lectures, in timestamp order as it is</li>\n<li>Window size 300-600</li>\n<li>batch size 1000</li>\n<li>drop_out 0</li>\n<li>n_encoder_layers 3-5</li>\n<li>augmentation to replace content_ids with dummy ids at a certain rate</li>\n<li>kept only one question in the last task to avoid leaks</li>\n</ul>\n<h3>features</h3>\n<p>embedded or dense was decided by CV.</p>\n<ul>\n<li>(embedded) content id</li>\n<li>(embedded) part id</li>\n<li>(embedded) same task question size</li>\n<li>(dense) answered_correctly</li>\n<li>(dense) had_explanation</li>\n<li>(dense) elapsed time</li>\n<li>(dense) lag time</li>\n<li>(dense) diff of timestamp from the last question</li>\n</ul>\n<h3>combining models</h3>\n<p>To avoid the overhead of calling model.predict() multiple times for ensemble, I made a combined model that links four models.</p>\n<pre><code>inputs = tf.keras.Input(shape=(input_shape, n_features))\nout1 = model1(inputs[:,-window_size1:,:])\nout2 = model2(inputs[:,-window_size2:,:])\nout3 = model3(inputs[:,-window_size3:,:])\nout4 = model4(inputs[:,-window_size4:,:])\ncombo_model = tf.keras.Model(inputs, [out1,out2,out3,out4])\n</code></pre>\n<h2>SAINT+ (nyanp, LB 0.808)</h2>\n<ul>\n<li>d_model = 256</li>\n<li>window_size = 200</li>\n<li>n_layers = 3</li>\n<li>attention dropout = 0.03</li>\n<li>question, part, lag are embed to encoder</li>\n<li>response, elapsed time, has_explanation are embed to decoder</li>\n</ul>\n<p>To prevent leakage, in addition to upper triangular attention mask, I masked the loss in questions other than the beginning of each task_container_id. Questions with the same task_container_id were shuffled in each batch during training, and the loss weights were adjusted to reduce the effect of masks. This mask improved LB by 0.0003.</p>\n<p>(Note: I believe that indirect leaks still exist, but I've spent 80% of my time on the LightGBM implementation and data pipeline, so I couldn't improve it any further)</p>\n<p>Other than that, there is nothing special about this NN. It scored .806 in single, .808 by averaging 2 models.</p>\n<h2>LightGBM (nyanp, LB 0.806)</h2>\n<p>LightGBM models are trained on 264 features. To speed up inference, I fixed the number of trees to 3000 and ensembled two models with different seeds (This is better than single large LGBM in terms of both speed and accuracy). By compiling this model with <a href=\"https://github.com/dmlc/treelite\" target=\"_blank\">treelite</a>, inference time became 3x faster (~10ms/batch, ~10min in total).</p>\n<h3>features</h3>\n<p>I mapped prior_question_* rows with their respective rows by following code and utilized them in some features.</p>\n<pre><code>df['elapsed_time'] = df.groupby('user_id')['prior_question_elapsed_time'].shift(-1)\ndf['elapsed_time'] = df.groupby(['user_id', 'timestamp'])['question_elapsed_time'].transform('last')\n\ndf['has_explanation'] = df.groupby('user_id')['prior_question_had_explanation'].shift(-1)\ndf['has_explanation'] = df.groupby(['user_id', 'timestamp'])['has_explanation'].transform('last')\n</code></pre>\n<p>Here is the list of my features. There is no magic here; no single feature boost CV more than 0.0002. I repeated feature engineering based on well-known techniques and a little bit of domain knowledge.</p>\n<h4>question features</h4>\n<ul>\n<li>count encoding</li>\n<li>target encoding</li>\n<li>number of tags</li>\n<li>one-hot encoding of tag (top-10 frequent tags)</li>\n<li>SVD, LDA, item2vec using user_id x content_id matrix</li>\n<li>LDA, item2vec using user_id x content_id matrix (filtered by answered_correctly == 0)<ul>\n<li>Typical word2vec model is trained on next-word prediction task. By constructing word2vec model over incorrectly answered questions, <br>\nthe latent vectors extracted from the model can be used to capture which incorrect question are likely to co-occur with each other.</li></ul></li>\n<li>10%, 20%, 50%, 80% elapsed time of all users response, correct response, wrong response</li>\n<li>SAINT embedding vector + PCA</li>\n</ul>\n<h4>user features</h4>\n<ul>\n<li>avg/median/max/std elapsed_time</li>\n<li>avg/median/max/std elapsed_time by part</li>\n<li>avg has_explanation flag</li>\n<li>avg has_explanation flag by part</li>\n<li>nunique of question, part, lecture</li>\n<li>cumcount / timestamp</li>\n<li>avg answered_correctly with recent 10/30/100/300 questions, recent 10min/7 days</li>\n<li>avg answered_correctly by part, question, bundle, order of response, question difficulty, question difficulty x part<ul>\n<li>question difficulty: discretize the avg answered_correctly of all users for each question into 10 levels</li></ul></li>\n<li>cumcount by part, question, bundle, question difficulty, question difficulty x part</li>\n<li>lag from 1/2/3/4 step before</li>\n<li>correctness, lag, elapsed_time, has_explanation in the same question last time</li>\n<li>tag-level aggregation features<ul>\n<li>calculate tag-level feature for each user x tag, then aggregate them by min/avg/max</li>\n<li>cumcount of wrong answer, cumcount of correct answer, avg target, lag</li></ul></li>\n<li>(estimated elapsed time) - (avg/median/min/max elapsed time within same part)<ul>\n<li>estimated elapsed time = (timestamp - prev timestamp) / (# of questions within the bundle)</li></ul></li>\n<li>(estimated elapsed time) - (10%, 20%, 50%, 80% elapsed time of all users correct/wrong response)<ul>\n<li>Because TOEIC part1-4 questions are usually answered after listening to the conversation, the correct answers tend to be concentrated immediately after the conversation ends</li></ul></li>\n<li>task_container_id - previous task_container_id<ul>\n<li>I'm not sure why this worked. There might be a difference in the correct rate if people answered from different devices than usual (multi-user?).</li></ul></li>\n<li>part of last lecture</li>\n<li>lag from last lecture</li>\n<li>whether the question contains the same tag as the last lecture</li>\n<li>whether the part is the same as the previous question</li>\n<li>median lag - median elapsed_time</li>\n<li>part of the first problem the user solved</li>\n<li>lag - median lag</li>\n<li>inner-product of user-{correct|incorrect}-question-vector and question-vector<ul>\n<li>user-correct-question-vector: average of LDA vectors for each question that the user answered correctly.</li></ul></li>\n<li>rank of lag compared to the same user's past lag (filtered by answered_correctly == 0, 1 respectively)</li>\n</ul>\n<h3>feature calculation</h3>\n<p>Instead of updating the user dictionary, I calculate user features from scratch for each bacth.</p>\n<p>The numpy array of historical data was loaded from the hdf storage and then split into questions and lectures, which were then wrapped in a pandas-like API and passed to their respective feature functions.</p>\n<pre><code>@feature('lag1.user_lecture')\ndef lag1_user_lecture(df: RiiidData, pool: DataPool) -&gt; np.ndarray:\n    \"\"\"\n    time elapsed from last lectures for each user\n\n    :param df: data in current batch\n    :param pool: cached data storage\n    \"\"\"\n    lag1 = {}\n    for u, t in set(zip(df.questions['user_id'], df.questions['timestamp'])):\n        past_lect = pool.users[u].lectures  # history of lecture\n        if len(past_lect) &gt; 0:\n            lag1[u] = t - past_lect['timestamp'][-1]\n\n    return np.array(list(map(lambda x: lag1.get(x, np.nan), df.questions['user_id'])), dtype=np.float32)\n</code></pre>\n<p>For training, I use the same functions as inference time. This way, there were almost no restrictions on feature creation, and I did not have to worry about bugs of train-test difference.</p>\n<p>The feature functions were frequently benchmarked by a dedicated benchmark script, and functions with high overhead were optimized by various ways (numba, bottleneck and various algorithm improvement) .</p>\n<p>The question features were precomputed and made into a global numpy array of shape (13523, *) and merged into the feature data using fancy index.</p>\n<pre><code>@feature(tuple(f'question_item2vec_{i}' for i in range(20)))\ndef question_item2vec(df: RiiidData, _: DataPool):\n    \"\"\"\n    question embedding vector using item2vec\n    \"\"\"\n    qid = df.questions['content_id']\n\n    ret = QUESTION_ITEM2VEC[qid, :]  # mere fancy index, faster than pd.merge\n\n    return ret\n</code></pre>\n<h2>Other ideas</h2>\n<ul>\n<li>TTA on SAINT+ by np.roll (it did improve SAINT+ by 0.001+, but we couldn't include it because of submission timeout)</li>\n<li>linear blending based on public LB labels (timeout, too)</li>\n</ul>\n<h2>Feedback on Time-Series API Competition</h2>\n<p>Although the competition was well-designed, our team still found that the time-series API allowed us to obtain private score information with 1-bit probing even after the known vulnerability was fixed.</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\n\nground_truth = []\nprediction = []\nthreshold = 0.810\n\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    ground_truth.extend(list(test_df['answered_correctly'].values))\n\n    predicted = model.predict(...)\n\n    prediction.extend(list(predicted))\n\n    if len(prediction) &gt;= 2500000:\n        private_auc = roc_auc_score(prediction[500000:], ground_truth[500000:])\n        if private_auc &lt; threshold:\n            raise RuntimeError()\n\n# private auc of this submission is guaranteed to exceed .810 if submission is succeeded\n</code></pre>\n<p>By using this probing, we can select the final submission with the highest private score, or determine the best<br>\n ensemble weight for private score by using the hill-climbing method.</p>\n<p>We contacted the Kaggle Team and asked them if this is legal, and they said it was a \"low signal matter\". We still think this is a gray-area and decided not to use this probing.</p>\n<p>We think it wasn't critical in this competition, but if there is a future competition with the same format but without AUC metrics, the hack with the \"magic coefficient\" could improve the ranking significantly. If a competition with the same format is held on Kaggle in the future, we suggest the Kaggle team to fix this problem (e.g. put a dummy value in the private label during the competition).</p>",
      "rawMarkdown": "First of all, thanks to the hosts for a great competition! This was one of the toughest competitions I have ever entered, but well worth the effort.\n\nPrediction of our team (tito @its7171 and nyanp @nyanpn)consists of following models.\n\n- tito's transformer: LB 0.813\n- nyanp's SAINT+ transformer: LB 0.808\n- nyanp's LightGBM: LB 0.806\n\nSimple blending of these models scored LB 0.814 / private 0.816.\n\n## Pipeline\n\nWe converted the entire train.csv data into hdf5, and loaded only the user_id that appeared during inference into a np array (97% RAM savings compared to hold entire training data in RAM). We estimate that the overhead due to I/O in hdf is ~45 minutes. This some overhead allowed us to combine tito's large transformer with nyanp's feature engineering pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532198%2F21ed1bbf93a37f6219fb7dc369761f0e%2Friiid%20pipeline.png?generation=1610287506256559&alt=media)\n\n## Transformer (tito, LB 0.814)\nThis is transformer model with encoder only, based on @claverru's [nice kernel](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)\n\n### summary\n* only trained and predicted for answered_correctly of last question of the sequence.\n* all features are concatenated (only position encoding is added).\n* used lectures, in timestamp order as it is\n* Window size 300-600\n* batch size 1000\n* drop_out 0\n* n_encoder_layers 3-5\n* augmentation to replace content_ids with dummy ids at a certain rate\n* kept only one question in the last task to avoid leaks\n\n### features\nembedded or dense was decided by CV.\n* (embedded) content id\n* (embedded) part id\n* (embedded) same task question size\n* (dense) answered_correctly\n* (dense) had_explanation\n* (dense) elapsed time\n* (dense) lag time\n* (dense) diff of timestamp from the last question\n\n### combining models\nTo avoid the overhead of calling model.predict() multiple times for ensemble, I made a combined model that links four models.\n\n```python\ninputs = tf.keras.Input(shape=(input_shape, n_features))\nout1 = model1(inputs[:,-window_size1:,:])\nout2 = model2(inputs[:,-window_size2:,:])\nout3 = model3(inputs[:,-window_size3:,:])\nout4 = model4(inputs[:,-window_size4:,:])\ncombo_model = tf.keras.Model(inputs, [out1,out2,out3,out4])\n```\n\n\n## SAINT+ (nyanp, LB 0.808)\n- d_model = 256\n- window_size = 200\n- n_layers = 3\n- attention dropout = 0.03\n- question, part, lag are embed to encoder\n- response, elapsed time, has_explanation are embed to decoder\n\nTo prevent leakage, in addition to upper triangular attention mask, I masked the loss in questions other than the beginning of each task_container_id. Questions with the same task_container_id were shuffled in each batch during training, and the loss weights were adjusted to reduce the effect of masks. This mask improved LB by 0.0003.\n\n(Note: I believe that indirect leaks still exist, but I've spent 80% of my time on the LightGBM implementation and data pipeline, so I couldn't improve it any further)\n\nOther than that, there is nothing special about this NN. It scored .806 in single, .808 by averaging 2 models.\n\n## LightGBM (nyanp, LB 0.806)\nLightGBM models are trained on 264 features. To speed up inference, I fixed the number of trees to 3000 and ensembled two models with different seeds (This is better than single large LGBM in terms of both speed and accuracy). By compiling this model with [treelite](https://github.com/dmlc/treelite), inference time became 3x faster (~10ms/batch, ~10min in total).\n\n### features\nI mapped prior_question_* rows with their respective rows by following code and utilized them in some features.\n\n```python\ndf['elapsed_time'] = df.groupby('user_id')['prior_question_elapsed_time'].shift(-1)\ndf['elapsed_time'] = df.groupby(['user_id', 'timestamp'])['question_elapsed_time'].transform('last')\n\ndf['has_explanation'] = df.groupby('user_id')['prior_question_had_explanation'].shift(-1)\ndf['has_explanation'] = df.groupby(['user_id', 'timestamp'])['has_explanation'].transform('last')\n```\n\nHere is the list of my features. There is no magic here; no single feature boost CV more than 0.0002. I repeated feature engineering based on well-known techniques and a little bit of domain knowledge.\n\n#### question features\n- count encoding\n- target encoding\n- number of tags\n- one-hot encoding of tag (top-10 frequent tags)\n- SVD, LDA, item2vec using user_id x content_id matrix\n- LDA, item2vec using user_id x content_id matrix (filtered by answered_correctly == 0)\n    - Typical word2vec model is trained on next-word prediction task. By constructing word2vec model over incorrectly answered questions, \nthe latent vectors extracted from the model can be used to capture which incorrect question are likely to co-occur with each other.\n- 10%, 20%, 50%, 80% elapsed time of all users response, correct response, wrong response\n- SAINT embedding vector + PCA\n\n#### user features\n- avg/median/max/std elapsed_time\n- avg/median/max/std elapsed_time by part\n- avg has_explanation flag\n- avg has_explanation flag by part\n- nunique of question, part, lecture\n- cumcount / timestamp\n- avg answered_correctly with recent 10/30/100/300 questions, recent 10min/7 days\n- avg answered_correctly by part, question, bundle, order of response, question difficulty, question difficulty x part\n    - question difficulty: discretize the avg answered_correctly of all users for each question into 10 levels\n- cumcount by part, question, bundle, question difficulty, question difficulty x part\n- lag from 1/2/3/4 step before\n- correctness, lag, elapsed_time, has_explanation in the same question last time\n- tag-level aggregation features\n    - calculate tag-level feature for each user x tag, then aggregate them by min/avg/max\n    - cumcount of wrong answer, cumcount of correct answer, avg target, lag\n- (estimated elapsed time) - (avg/median/min/max elapsed time within same part)\n    - estimated elapsed time = (timestamp - prev timestamp) / (# of questions within the bundle)\n- (estimated elapsed time) - (10%, 20%, 50%, 80% elapsed time of all users correct/wrong response)\n    - Because TOEIC part1-4 questions are usually answered after listening to the conversation, the correct answers tend to be concentrated immediately after the conversation ends\n- task_container_id - previous task_container_id\n    - I'm not sure why this worked. There might be a difference in the correct rate if people answered from different devices than usual (multi-user?).\n- part of last lecture\n- lag from last lecture\n- whether the question contains the same tag as the last lecture\n- whether the part is the same as the previous question\n- median lag - median elapsed_time\n- part of the first problem the user solved\n- lag - median lag\n- inner-product of user-{correct|incorrect}-question-vector and question-vector\n    - user-correct-question-vector: average of LDA vectors for each question that the user answered correctly.\n- rank of lag compared to the same user's past lag (filtered by answered_correctly == 0, 1 respectively)\n\n\n### feature calculation\nInstead of updating the user dictionary, I calculate user features from scratch for each bacth.\n\nThe numpy array of historical data was loaded from the hdf storage and then split into questions and lectures, which were then wrapped in a pandas-like API and passed to their respective feature functions.\n\n```python\n@feature('lag1.user_lecture')\ndef lag1_user_lecture(df: RiiidData, pool: DataPool) -> np.ndarray:\n    \"\"\"\n    time elapsed from last lectures for each user\n\n    :param df: data in current batch\n    :param pool: cached data storage\n    \"\"\"\n    lag1 = {}\n    for u, t in set(zip(df.questions['user_id'], df.questions['timestamp'])):\n        past_lect = pool.users[u].lectures  # history of lecture\n        if len(past_lect) > 0:\n            lag1[u] = t - past_lect['timestamp'][-1]\n\n    return np.array(list(map(lambda x: lag1.get(x, np.nan), df.questions['user_id'])), dtype=np.float32)\n```\n\nFor training, I use the same functions as inference time. This way, there were almost no restrictions on feature creation, and I did not have to worry about bugs of train-test difference.\n\nThe feature functions were frequently benchmarked by a dedicated benchmark script, and functions with high overhead were optimized by various ways (numba, bottleneck and various algorithm improvement) .\n\nThe question features were precomputed and made into a global numpy array of shape (13523, *) and merged into the feature data using fancy index.\n\n```python\n@feature(tuple(f'question_item2vec_{i}' for i in range(20)))\ndef question_item2vec(df: RiiidData, _: DataPool):\n    \"\"\"\n    question embedding vector using item2vec\n    \"\"\"\n    qid = df.questions['content_id']\n\n    ret = QUESTION_ITEM2VEC[qid, :]  # mere fancy index, faster than pd.merge\n\n    return ret\n```\n\n\n## Other ideas\n- TTA on SAINT+ by np.roll (it did improve SAINT+ by 0.001+, but we couldn't include it because of submission timeout)\n- linear blending based on public LB labels (timeout, too)\n\n## Feedback on Time-Series API Competition\nAlthough the competition was well-designed, our team still found that the time-series API allowed us to obtain private score information with 1-bit probing even after the known vulnerability was fixed.\n\n```python\nenv = riiideducation.make_env()\niter_test = env.iter_test()\n\nground_truth = []\nprediction = []\nthreshold = 0.810\n\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    ground_truth.extend(list(test_df['answered_correctly'].values))\n\n    predicted = model.predict(...)\n\n    prediction.extend(list(predicted))\n\n    if len(prediction) >= 2500000:\n        private_auc = roc_auc_score(prediction[500000:], ground_truth[500000:])\n        if private_auc < threshold:\n            raise RuntimeError()\n\n# private auc of this submission is guaranteed to exceed .810 if submission is succeeded\n```\n\nBy using this probing, we can select the final submission with the highest private score, or determine the best\n ensemble weight for private score by using the hill-climbing method.\n\nWe contacted the Kaggle Team and asked them if this is legal, and they said it was a \"low signal matter\". We still think this is a gray-area and decided not to use this probing.\n\nWe think it wasn't critical in this competition, but if there is a future competition with the same format but without AUC metrics, the hack with the \"magic coefficient\" could improve the ranking significantly. If a competition with the same format is held on Kaggle in the future, we suggest the Kaggle team to fix this problem (e.g. put a dummy value in the private label during the competition).",
      "votes": null
    },
    {
      "id": "1147474",
      "postDate": "01/10/2021 14:31:29",
      "content": "<p>Congrats nyanp, your feature is really great :)! Wow… seems you found the security hole of Time-Series API. I think kaggle team has to deal with this problem in future competitions.</p>",
      "rawMarkdown": "Congrats nyanp, your feature is really great :)! Wow... seems you found the security hole of Time-Series API. I think kaggle team has to deal with this problem in future competitions.",
      "votes": null
    },
    {
      "id": "1147477",
      "postDate": "01/10/2021 14:33:01",
      "content": "<p>there is also an ongoing competition that relates to time series API, </p>\n<p><a href=\"https://www.kaggle.com/c/jane-street-market-prediction\" target=\"_blank\">Jane Street Market Prediction</a></p>\n<p>Not sure if there is a hole there, but better for Kaggle to checkout.</p>",
      "rawMarkdown": "there is also an ongoing competition that relates to time series API, \n\n[Jane Street Market Prediction](https://www.kaggle.com/c/jane-street-market-prediction)\n\nNot sure if there is a hole there, but better for Kaggle to checkout.",
      "votes": null
    },
    {
      "id": "1147491",
      "postDate": "01/10/2021 14:42:41",
      "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> Congrats!</p>\n<p>About</p>\n<pre><code>Instead of updating the user dictionary, I calculate user features from scratch for each bacth.\n</code></pre>\n<p>Do you mean, even during the inference (submission time), once you get the <code>prior_user_answer</code> and <code>prior_answer_correctness</code>, you use them along with other standard information from <code>test_df</code> in the previous inference batch, and from these raw data to calculate all the features you created?</p>",
      "rawMarkdown": "nyanpn Congrats!\n\nAbout\n\n```\nInstead of updating the user dictionary, I calculate user features from scratch for each bacth.\n```\n\nDo you mean, even during the inference (submission time), once you get the `prior_user_answer` and `prior_answer_correctness`, you use them along with other standard information from `test_df` in the previous inference batch, and from these raw data to calculate all the features you created?",
      "votes": null
    },
    {
      "id": "1147572",
      "postDate": "01/10/2021 15:34:39",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> and tito on 9th place and thanks for sharing your team solution. Thanks tito for cross validation strategy. </p>",
      "rawMarkdown": "Congrats @nyanpn and tito on 9th place and thanks for sharing your team solution. Thanks tito for cross validation strategy.",
      "votes": null
    },
    {
      "id": "1147648",
      "postDate": "01/10/2021 16:19:19",
      "content": "<p>I guess there is no hole in Jane, because it has <code>Forecasting Timeline</code>.</p>",
      "rawMarkdown": "I guess there is no hole in Jane, because it has `Forecasting Timeline`.",
      "votes": null
    },
    {
      "id": "1148574",
      "postDate": "01/11/2021 08:19:48",
      "content": "<p>Congrats nyanp!</p>\n<p>Could you show the sample code about how to  get item2vec?<br>\nThank you.</p>",
      "rawMarkdown": "Congrats nyanp!\n\nCould you show the sample code about how to  get item2vec?\nThank you.",
      "votes": null
    },
    {
      "id": "1149019",
      "postDate": "01/11/2021 14:33:24",
      "content": "<p>Exactly.<br>\nI kept the raw historical data for each user in a numpy array, and <code>prior_user_answer</code> and <code>prior_answer_correctness</code> were also merged into the historical data once and then used for features in the same fashion as other standard information.</p>",
      "rawMarkdown": "Exactly.\nI kept the raw historical data for each user in a numpy array, and `prior_user_answer` and `prior_answer_correctness` were also merged into the historical data once and then used for features in the same fashion as other standard information.",
      "votes": null
    },
    {
      "id": "1149020",
      "postDate": "01/11/2021 14:33:24",
      "content": "<p>Exactly.<br>\nI kept the raw historical data for each user in a numpy array, and <code>prior_user_answer</code> and <code>prior_answer_correctness</code> were also merged into the historical data once and then used for features in the same fashion as other standard information.</p>",
      "rawMarkdown": "Exactly.\nI kept the raw historical data for each user in a numpy array, and `prior_user_answer` and `prior_answer_correctness` were also merged into the historical data once and then used for features in the same fashion as other standard information.",
      "votes": null
    },
    {
      "id": "1149027",
      "postDate": "01/11/2021 14:37:17",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and congrats too! I personally don't think this issue is a \"low signal matter\" at all. I hope it will be dealt with appropriately.</p>\n<blockquote>\n  <p>I guess there is no hole in Jane, because it has Forecasting Timeline.</p>\n</blockquote>\n<p>Yes. <strong>the problem occurs when the Time-Series API is combined with the Synchronous Code Competition format</strong>; a competition where recalculation is performed for the Private Leaderboard would not have the problem.</p>",
      "rawMarkdown": "Thanks @mamasinkgs and congrats too! I personally don't think this issue is a \"low signal matter\" at all. I hope it will be dealt with appropriately.\n\n> I guess there is no hole in Jane, because it has Forecasting Timeline.\n\nYes. **the problem occurs when the Time-Series API is combined with the Synchronous Code Competition format**; a competition where recalculation is performed for the Private Leaderboard would not have the problem.",
      "votes": null
    },
    {
      "id": "1149030",
      "postDate": "01/11/2021 14:40:58",
      "content": "<p>Thanks! I also agree that <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> 's validation is awesome.</p>",
      "rawMarkdown": "Thanks! I also agree that @its7171 's validation is awesome.",
      "votes": null
    },
    {
      "id": "1149036",
      "postDate": "01/11/2021 14:42:51",
      "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> Thank you for confirmation. That's indeed a good way to avoid the inconsistency between the training and inference time, which I introduced (these errors) at the end of the competition.</p>",
      "rawMarkdown": "nyanpn Thank you for confirmation. That's indeed a good way to avoid the inconsistency between the training and inference time, which I introduced (these errors) at the end of the competition.",
      "votes": null
    },
    {
      "id": "1149049",
      "postDate": "01/11/2021 14:52:58",
      "content": "<p>item2vec is a simple idea to run word2vec with a sequence of question IDs as sentences, and turn latent expressions into features.<br>\nPlease note that it cannot be run in kaggle notebook due to lack of RAM.</p>\n<pre><code>import pandas as pd\nimport numpy as np\nfrom gensim.models import Word2Vec\n\nEMBEDDING_DIM = 20\n\ntrain = pd.read_csv('train.csv')\nquestions = pd.read_csv('questions.csv')\n\ntrain = train[(train.content_type_id == 0) &amp; (train.answered_correctly == 0)]\n\nsentences = train.groupby('user_id')['content_id'].apply(lambda x: [str(t) for t in x]).values\n\nmodel = Word2Vec(sentences, size=EMBEDDING_DIM , window=100, seed=0, workers=16)\n\nresult_vector = np.zeros((len(q), EMBEDDING_DIM ))\n\nfor i in range(len(q)):\n    try:\n        result_vector[i, :] = model.wv[str(i)]\n    except:\n        pass\n</code></pre>",
      "rawMarkdown": "item2vec is a simple idea to run word2vec with a sequence of question IDs as sentences, and turn latent expressions into features.\nPlease note that it cannot be run in kaggle notebook due to lack of RAM.\n\n```python\nimport pandas as pd\nimport numpy as np\nfrom gensim.models import Word2Vec\n\nEMBEDDING_DIM = 20\n\ntrain = pd.read_csv('train.csv')\nquestions = pd.read_csv('questions.csv')\n\ntrain = train[(train.content_type_id == 0) & (train.answered_correctly == 0)]\n\nsentences = train.groupby('user_id')['content_id'].apply(lambda x: [str(t) for t in x]).values\n\nmodel = Word2Vec(sentences, size=EMBEDDING_DIM , window=100, seed=0, workers=16)\n\nresult_vector = np.zeros((len(q), EMBEDDING_DIM ))\n\nfor i in range(len(q)):\n    try:\n        result_vector[i, :] = model.wv[str(i)]\n    except:\n        pass\n```",
      "votes": null
    },
    {
      "id": "1150420",
      "postDate": "01/12/2021 15:15:10",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>, many thanks for sharing ! </p>\n<p>Can give more details about LDA on <code>user_id x content_id matrix</code> ? Is the LDA trained on the number of occurrences ? How did you choose a well-fitted n_components ? </p>\n<p>Thanks again <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> for your great kernels. It seems almost everyone relied on your cv strategy …</p>",
      "rawMarkdown": "Congratz @nyanpn, many thanks for sharing ! \n\nCan give more details about LDA on `user_id x content_id matrix` ? Is the LDA trained on the number of occurrences ? How did you choose a well-fitted n_components ? \n\nThanks again @its7171 for your great kernels. It seems almost everyone relied on your cv strategy ...",
      "votes": null
    },
    {
      "id": "1151720",
      "postDate": "01/13/2021 14:12:23",
      "content": "<p>Thanks, I've made my FE code public. Hope this helps ;)<br>\n<a href=\"https://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm\" target=\"_blank\">https://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm</a></p>",
      "rawMarkdown": "Thanks, I've made my FE code public. Hope this helps ;)\nhttps://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm",
      "votes": null
    },
    {
      "id": "1151729",
      "postDate": "01/13/2021 14:18:08",
      "content": "<blockquote>\n  <p>How did you choose a well-fitted n_components ?</p>\n</blockquote>\n<p>The <code>n_component</code> for LDA is set to 10, but I don't know if this is the best because I didn't have time to explore other values. Since the number of questions is only ~15,000, setting the n_component larger than 30 will probably not have much effect.</p>",
      "rawMarkdown": "> How did you choose a well-fitted n_components ?\n\nThe `n_component` for LDA is set to 10, but I don't know if this is the best because I didn't have time to explore other values. Since the number of questions is only ~15,000, setting the n_component larger than 30 will probably not have much effect.",
      "votes": null
    },
    {
      "id": "1153931",
      "postDate": "01/15/2021 09:01:31",
      "content": "<p>Awesome, thanks !!</p>",
      "rawMarkdown": "Awesome, thanks !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1147474,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "01/10/2021 14:31:29",
      "content": "<p>Congrats nyanp, your feature is really great :)! Wow… seems you found the security hole of Time-Series API. I think kaggle team has to deal with this problem in future competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147477,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/10/2021 14:33:01",
          "content": "<p>there is also an ongoing competition that relates to time series API, </p>\n<p><a href=\"https://www.kaggle.com/c/jane-street-market-prediction\" target=\"_blank\">Jane Street Market Prediction</a></p>\n<p>Not sure if there is a hole there, but better for Kaggle to checkout.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147648,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "01/10/2021 16:19:19",
          "content": "<p>I guess there is no hole in Jane, because it has <code>Forecasting Timeline</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149027,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/11/2021 14:37:17",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and congrats too! I personally don't think this issue is a \"low signal matter\" at all. I hope it will be dealt with appropriately.</p>\n<blockquote>\n  <p>I guess there is no hole in Jane, because it has Forecasting Timeline.</p>\n</blockquote>\n<p>Yes. <strong>the problem occurs when the Time-Series API is combined with the Synchronous Code Competition format</strong>; a competition where recalculation is performed for the Private Leaderboard would not have the problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147491,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/10/2021 14:42:41",
      "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> Congrats!</p>\n<p>About</p>\n<pre><code>Instead of updating the user dictionary, I calculate user features from scratch for each bacth.\n</code></pre>\n<p>Do you mean, even during the inference (submission time), once you get the <code>prior_user_answer</code> and <code>prior_answer_correctness</code>, you use them along with other standard information from <code>test_df</code> in the previous inference batch, and from these raw data to calculate all the features you created?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149019,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/11/2021 14:33:24",
          "content": "<p>Exactly.<br>\nI kept the raw historical data for each user in a numpy array, and <code>prior_user_answer</code> and <code>prior_answer_correctness</code> were also merged into the historical data once and then used for features in the same fashion as other standard information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149020,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/11/2021 14:33:24",
          "content": "<p>Exactly.<br>\nI kept the raw historical data for each user in a numpy array, and <code>prior_user_answer</code> and <code>prior_answer_correctness</code> were also merged into the historical data once and then used for features in the same fashion as other standard information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149036,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/11/2021 14:42:51",
          "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> Thank you for confirmation. That's indeed a good way to avoid the inconsistency between the training and inference time, which I introduced (these errors) at the end of the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147572,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/10/2021 15:34:39",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> and tito on 9th place and thanks for sharing your team solution. Thanks tito for cross validation strategy. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1149030,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/11/2021 14:40:58",
          "content": "<p>Thanks! I also agree that <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> 's validation is awesome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148574,
      "author_name": "m10515009",
      "author_url": "",
      "post_date": "01/11/2021 08:19:48",
      "content": "<p>Congrats nyanp!</p>\n<p>Could you show the sample code about how to  get item2vec?<br>\nThank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149049,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/11/2021 14:52:58",
          "content": "<p>item2vec is a simple idea to run word2vec with a sequence of question IDs as sentences, and turn latent expressions into features.<br>\nPlease note that it cannot be run in kaggle notebook due to lack of RAM.</p>\n<pre><code>import pandas as pd\nimport numpy as np\nfrom gensim.models import Word2Vec\n\nEMBEDDING_DIM = 20\n\ntrain = pd.read_csv('train.csv')\nquestions = pd.read_csv('questions.csv')\n\ntrain = train[(train.content_type_id == 0) &amp; (train.answered_correctly == 0)]\n\nsentences = train.groupby('user_id')['content_id'].apply(lambda x: [str(t) for t in x]).values\n\nmodel = Word2Vec(sentences, size=EMBEDDING_DIM , window=100, seed=0, workers=16)\n\nresult_vector = np.zeros((len(q), EMBEDDING_DIM ))\n\nfor i in range(len(q)):\n    try:\n        result_vector[i, :] = model.wv[str(i)]\n    except:\n        pass\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1150420,
      "author_name": "johannhuber",
      "author_url": "",
      "post_date": "01/12/2021 15:15:10",
      "content": "<p>Congratz <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>, many thanks for sharing ! </p>\n<p>Can give more details about LDA on <code>user_id x content_id matrix</code> ? Is the LDA trained on the number of occurrences ? How did you choose a well-fitted n_components ? </p>\n<p>Thanks again <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> for your great kernels. It seems almost everyone relied on your cv strategy …</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151720,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/13/2021 14:12:23",
          "content": "<p>Thanks, I've made my FE code public. Hope this helps ;)<br>\n<a href=\"https://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm\" target=\"_blank\">https://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151729,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "01/13/2021 14:18:08",
          "content": "<blockquote>\n  <p>How did you choose a well-fitted n_components ?</p>\n</blockquote>\n<p>The <code>n_component</code> for LDA is set to 10, but I don't know if this is the best because I didn't have time to explore other values. Since the number of questions is only ~15,000, setting the n_component larger than 30 will probably not have much effect.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1153931,
          "author_name": "johannhuber",
          "author_url": "",
          "post_date": "01/15/2021 09:01:31",
          "content": "<p>Awesome, thanks !!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1147440": "First of all, thanks to the hosts for a great competition! This was one of the toughest competitions I have ever entered, but well worth the effort.\n\nPrediction of our team (tito @its7171 and nyanp @nyanpn)consists of following models.\n\n- tito's transformer: LB 0.813\n- nyanp's SAINT+ transformer: LB 0.808\n- nyanp's LightGBM: LB 0.806\n\nSimple blending of these models scored LB 0.814 / private 0.816.\n\n## Pipeline\n\nWe converted the entire train.csv data into hdf5, and loaded only the user_id that appeared during inference into a np array (97% RAM savings compared to hold entire training data in RAM). We estimate that the overhead due to I/O in hdf is ~45 minutes. This some overhead allowed us to combine tito's large transformer with nyanp's feature engineering pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1532198%2F21ed1bbf93a37f6219fb7dc369761f0e%2Friiid%20pipeline.png?generation=1610287506256559&alt=media)\n\n## Transformer (tito, LB 0.814)\nThis is transformer model with encoder only, based on @claverru's [nice kernel](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)\n\n### summary\n* only trained and predicted for answered_correctly of last question of the sequence.\n* all features are concatenated (only position encoding is added).\n* used lectures, in timestamp order as it is\n* Window size 300-600\n* batch size 1000\n* drop_out 0\n* n_encoder_layers 3-5\n* augmentation to replace content_ids with dummy ids at a certain rate\n* kept only one question in the last task to avoid leaks\n\n### features\nembedded or dense was decided by CV.\n* (embedded) content id\n* (embedded) part id\n* (embedded) same task question size\n* (dense) answered_correctly\n* (dense) had_explanation\n* (dense) elapsed time\n* (dense) lag time\n* (dense) diff of timestamp from the last question\n\n### combining models\nTo avoid the overhead of calling model.predict() multiple times for ensemble, I made a combined model that links four models.\n\n```python\ninputs = tf.keras.Input(shape=(input_shape, n_features))\nout1 = model1(inputs[:,-window_size1:,:])\nout2 = model2(inputs[:,-window_size2:,:])\nout3 = model3(inputs[:,-window_size3:,:])\nout4 = model4(inputs[:,-window_size4:,:])\ncombo_model = tf.keras.Model(inputs, [out1,out2,out3,out4])\n```\n\n\n## SAINT+ (nyanp, LB 0.808)\n- d_model = 256\n- window_size = 200\n- n_layers = 3\n- attention dropout = 0.03\n- question, part, lag are embed to encoder\n- response, elapsed time, has_explanation are embed to decoder\n\nTo prevent leakage, in addition to upper triangular attention mask, I masked the loss in questions other than the beginning of each task_container_id. Questions with the same task_container_id were shuffled in each batch during training, and the loss weights were adjusted to reduce the effect of masks. This mask improved LB by 0.0003.\n\n(Note: I believe that indirect leaks still exist, but I've spent 80% of my time on the LightGBM implementation and data pipeline, so I couldn't improve it any further)\n\nOther than that, there is nothing special about this NN. It scored .806 in single, .808 by averaging 2 models.\n\n## LightGBM (nyanp, LB 0.806)\nLightGBM models are trained on 264 features. To speed up inference, I fixed the number of trees to 3000 and ensembled two models with different seeds (This is better than single large LGBM in terms of both speed and accuracy). By compiling this model with [treelite](https://github.com/dmlc/treelite), inference time became 3x faster (~10ms/batch, ~10min in total).\n\n### features\nI mapped prior_question_* rows with their respective rows by following code and utilized them in some features.\n\n```python\ndf['elapsed_time'] = df.groupby('user_id')['prior_question_elapsed_time'].shift(-1)\ndf['elapsed_time'] = df.groupby(['user_id', 'timestamp'])['question_elapsed_time'].transform('last')\n\ndf['has_explanation'] = df.groupby('user_id')['prior_question_had_explanation'].shift(-1)\ndf['has_explanation'] = df.groupby(['user_id', 'timestamp'])['has_explanation'].transform('last')\n```\n\nHere is the list of my features. There is no magic here; no single feature boost CV more than 0.0002. I repeated feature engineering based on well-known techniques and a little bit of domain knowledge.\n\n#### question features\n- count encoding\n- target encoding\n- number of tags\n- one-hot encoding of tag (top-10 frequent tags)\n- SVD, LDA, item2vec using user_id x content_id matrix\n- LDA, item2vec using user_id x content_id matrix (filtered by answered_correctly == 0)\n    - Typical word2vec model is trained on next-word prediction task. By constructing word2vec model over incorrectly answered questions, \nthe latent vectors extracted from the model can be used to capture which incorrect question are likely to co-occur with each other.\n- 10%, 20%, 50%, 80% elapsed time of all users response, correct response, wrong response\n- SAINT embedding vector + PCA\n\n#### user features\n- avg/median/max/std elapsed_time\n- avg/median/max/std elapsed_time by part\n- avg has_explanation flag\n- avg has_explanation flag by part\n- nunique of question, part, lecture\n- cumcount / timestamp\n- avg answered_correctly with recent 10/30/100/300 questions, recent 10min/7 days\n- avg answered_correctly by part, question, bundle, order of response, question difficulty, question difficulty x part\n    - question difficulty: discretize the avg answered_correctly of all users for each question into 10 levels\n- cumcount by part, question, bundle, question difficulty, question difficulty x part\n- lag from 1/2/3/4 step before\n- correctness, lag, elapsed_time, has_explanation in the same question last time\n- tag-level aggregation features\n    - calculate tag-level feature for each user x tag, then aggregate them by min/avg/max\n    - cumcount of wrong answer, cumcount of correct answer, avg target, lag\n- (estimated elapsed time) - (avg/median/min/max elapsed time within same part)\n    - estimated elapsed time = (timestamp - prev timestamp) / (# of questions within the bundle)\n- (estimated elapsed time) - (10%, 20%, 50%, 80% elapsed time of all users correct/wrong response)\n    - Because TOEIC part1-4 questions are usually answered after listening to the conversation, the correct answers tend to be concentrated immediately after the conversation ends\n- task_container_id - previous task_container_id\n    - I'm not sure why this worked. There might be a difference in the correct rate if people answered from different devices than usual (multi-user?).\n- part of last lecture\n- lag from last lecture\n- whether the question contains the same tag as the last lecture\n- whether the part is the same as the previous question\n- median lag - median elapsed_time\n- part of the first problem the user solved\n- lag - median lag\n- inner-product of user-{correct|incorrect}-question-vector and question-vector\n    - user-correct-question-vector: average of LDA vectors for each question that the user answered correctly.\n- rank of lag compared to the same user's past lag (filtered by answered_correctly == 0, 1 respectively)\n\n\n### feature calculation\nInstead of updating the user dictionary, I calculate user features from scratch for each bacth.\n\nThe numpy array of historical data was loaded from the hdf storage and then split into questions and lectures, which were then wrapped in a pandas-like API and passed to their respective feature functions.\n\n```python\n@feature('lag1.user_lecture')\ndef lag1_user_lecture(df: RiiidData, pool: DataPool) -> np.ndarray:\n    \"\"\"\n    time elapsed from last lectures for each user\n\n    :param df: data in current batch\n    :param pool: cached data storage\n    \"\"\"\n    lag1 = {}\n    for u, t in set(zip(df.questions['user_id'], df.questions['timestamp'])):\n        past_lect = pool.users[u].lectures  # history of lecture\n        if len(past_lect) > 0:\n            lag1[u] = t - past_lect['timestamp'][-1]\n\n    return np.array(list(map(lambda x: lag1.get(x, np.nan), df.questions['user_id'])), dtype=np.float32)\n```\n\nFor training, I use the same functions as inference time. This way, there were almost no restrictions on feature creation, and I did not have to worry about bugs of train-test difference.\n\nThe feature functions were frequently benchmarked by a dedicated benchmark script, and functions with high overhead were optimized by various ways (numba, bottleneck and various algorithm improvement) .\n\nThe question features were precomputed and made into a global numpy array of shape (13523, *) and merged into the feature data using fancy index.\n\n```python\n@feature(tuple(f'question_item2vec_{i}' for i in range(20)))\ndef question_item2vec(df: RiiidData, _: DataPool):\n    \"\"\"\n    question embedding vector using item2vec\n    \"\"\"\n    qid = df.questions['content_id']\n\n    ret = QUESTION_ITEM2VEC[qid, :]  # mere fancy index, faster than pd.merge\n\n    return ret\n```\n\n\n## Other ideas\n- TTA on SAINT+ by np.roll (it did improve SAINT+ by 0.001+, but we couldn't include it because of submission timeout)\n- linear blending based on public LB labels (timeout, too)\n\n## Feedback on Time-Series API Competition\nAlthough the competition was well-designed, our team still found that the time-series API allowed us to obtain private score information with 1-bit probing even after the known vulnerability was fixed.\n\n```python\nenv = riiideducation.make_env()\niter_test = env.iter_test()\n\nground_truth = []\nprediction = []\nthreshold = 0.810\n\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    ground_truth.extend(list(test_df['answered_correctly'].values))\n\n    predicted = model.predict(...)\n\n    prediction.extend(list(predicted))\n\n    if len(prediction) >= 2500000:\n        private_auc = roc_auc_score(prediction[500000:], ground_truth[500000:])\n        if private_auc < threshold:\n            raise RuntimeError()\n\n# private auc of this submission is guaranteed to exceed .810 if submission is succeeded\n```\n\nBy using this probing, we can select the final submission with the highest private score, or determine the best\n ensemble weight for private score by using the hill-climbing method.\n\nWe contacted the Kaggle Team and asked them if this is legal, and they said it was a \"low signal matter\". We still think this is a gray-area and decided not to use this probing.\n\nWe think it wasn't critical in this competition, but if there is a future competition with the same format but without AUC metrics, the hack with the \"magic coefficient\" could improve the ranking significantly. If a competition with the same format is held on Kaggle in the future, we suggest the Kaggle team to fix this problem (e.g. put a dummy value in the private label during the competition).",
    "1147474": "Congrats nyanp, your feature is really great :)! Wow... seems you found the security hole of Time-Series API. I think kaggle team has to deal with this problem in future competitions.",
    "1147477": "there is also an ongoing competition that relates to time series API, \n\n[Jane Street Market Prediction](https://www.kaggle.com/c/jane-street-market-prediction)\n\nNot sure if there is a hole there, but better for Kaggle to checkout.",
    "1147491": "nyanpn Congrats!\n\nAbout\n\n```\nInstead of updating the user dictionary, I calculate user features from scratch for each bacth.\n```\n\nDo you mean, even during the inference (submission time), once you get the `prior_user_answer` and `prior_answer_correctness`, you use them along with other standard information from `test_df` in the previous inference batch, and from these raw data to calculate all the features you created?",
    "1147572": "Congrats @nyanpn and tito on 9th place and thanks for sharing your team solution. Thanks tito for cross validation strategy.",
    "1147648": "I guess there is no hole in Jane, because it has `Forecasting Timeline`.",
    "1148574": "Congrats nyanp!\n\nCould you show the sample code about how to  get item2vec?\nThank you.",
    "1149019": "Exactly.\nI kept the raw historical data for each user in a numpy array, and `prior_user_answer` and `prior_answer_correctness` were also merged into the historical data once and then used for features in the same fashion as other standard information.",
    "1149020": "Exactly.\nI kept the raw historical data for each user in a numpy array, and `prior_user_answer` and `prior_answer_correctness` were also merged into the historical data once and then used for features in the same fashion as other standard information.",
    "1149027": "Thanks @mamasinkgs and congrats too! I personally don't think this issue is a \"low signal matter\" at all. I hope it will be dealt with appropriately.\n\n> I guess there is no hole in Jane, because it has Forecasting Timeline.\n\nYes. **the problem occurs when the Time-Series API is combined with the Synchronous Code Competition format**; a competition where recalculation is performed for the Private Leaderboard would not have the problem.",
    "1149030": "Thanks! I also agree that @its7171 's validation is awesome.",
    "1149036": "nyanpn Thank you for confirmation. That's indeed a good way to avoid the inconsistency between the training and inference time, which I introduced (these errors) at the end of the competition.",
    "1149049": "item2vec is a simple idea to run word2vec with a sequence of question IDs as sentences, and turn latent expressions into features.\nPlease note that it cannot be run in kaggle notebook due to lack of RAM.\n\n```python\nimport pandas as pd\nimport numpy as np\nfrom gensim.models import Word2Vec\n\nEMBEDDING_DIM = 20\n\ntrain = pd.read_csv('train.csv')\nquestions = pd.read_csv('questions.csv')\n\ntrain = train[(train.content_type_id == 0) & (train.answered_correctly == 0)]\n\nsentences = train.groupby('user_id')['content_id'].apply(lambda x: [str(t) for t in x]).values\n\nmodel = Word2Vec(sentences, size=EMBEDDING_DIM , window=100, seed=0, workers=16)\n\nresult_vector = np.zeros((len(q), EMBEDDING_DIM ))\n\nfor i in range(len(q)):\n    try:\n        result_vector[i, :] = model.wv[str(i)]\n    except:\n        pass\n```",
    "1150420": "Congratz @nyanpn, many thanks for sharing ! \n\nCan give more details about LDA on `user_id x content_id matrix` ? Is the LDA trained on the number of occurrences ? How did you choose a well-fitted n_components ? \n\nThanks again @its7171 for your great kernels. It seems almost everyone relied on your cv strategy ...",
    "1151720": "Thanks, I've made my FE code public. Hope this helps ;)\nhttps://www.kaggle.com/nyanpn/lda-feature-for-riiid-part-of-9th-place-lightgbm",
    "1151729": "> How did you choose a well-fitted n_components ?\n\nThe `n_component` for LDA is set to 10, but I don't know if this is the best because I didn't have time to explore other values. Since the number of questions is only ~15,000, setting the n_component larger than 30 will probably not have much effect.",
    "1153931": "Awesome, thanks !!"
  },
  "source": "meta"
}