{
  "id": 210025,
  "title": "21st Place Solution - Ensemble of 3 Saint++ Models",
  "url": "/competitions/riiid-test-answer-prediction/writeups/ancient-magics-21st-place-solution-ensemble-of-3-s",
  "author_name": "",
  "post_date": "2021-01-09T21:28:14.680Z",
  "votes": 26,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi all!<br>\nAfter seeing everyone share their solution we have decided to also share ours in the same spirit.</p>\n<p>We ran a modified SAINT+ that <strong>included lectures, tags and a couple of aggregate features.</strong></p>\n<p>All our code is available on github ( <a href=\"https://github.com/gautierdag/riiid\" target=\"_blank\">https://github.com/gautierdag/riiid</a> ). We used Pytorch/Pytorch Lightning and Hydra.</p>\n<p>Our single model public LB was 0.808, and we were able to push that to LB 0.810 (Final LB 0.813) through ensembling.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F583548%2Fa19776338ca1ba57bd1260bf8e0ec678%2Friiid_model.png?generation=1610188697948445&amp;alt=media\" alt=\"\"></p>\n<h2>Modifications to original SAINT:</h2>\n<ul>\n<li><p><strong>Lectures</strong>: lectures were included and had a special trainable embedding vector instead of the answer embedding. The loss was not masked for lecture steps. We tested both with and without, and the difference wasn't crazy. The big benefit they offered was that they made everything much simpler to handle since we did not need to filter anything.</p></li>\n<li><p><strong>Tags</strong>: Since questions had a maximum of six tags, we passed sequences of length 6 to a tag embedding layer and summed the output. Therefore our tag embeddings were learned for each tag but then sum across to obtain an embedding for all the tags in a question/lecture. The padding simply returned a 0 vector.</p></li>\n<li><p><strong>Agg Feats</strong>: This is what gave us a healthy boost of 0.002 on public LB. By including 12 aggregate features in the decoder we were able to use some of our learnings in previous LGBM implementations. </p></li>\n</ul>\n<h4>Agg feats</h4>\n<p>We used the following aggregate features:</p>\n<ul>\n<li>attempts: number of previous attempts on current question (clamped to 5 and normalized to 1)</li>\n<li>mean_content_mean: average of all the average question accuracy (population wise) seen until this step by the user</li>\n<li>mean_user_mean: average of user's accuracy until this step </li>\n<li>parts_mean: seven dimensional (the average of the user's accuracy on each part)</li>\n<li>m_user_session_mean: the average accuracy in the current user's session </li>\n<li>session_number: the current session number (how many sessions up till now)</li>\n</ul>\n<p>Note: session is just a way of dividing up the questions if a difference in time between questions is greater than two hours.</p>\n<h4>Float vs Categorical</h4>\n<p>We switched to use a float representation of answers in the last few days. This seemed to have little effect on our LB auc unfortunately, but had an effect locally. The idea was that since we auto-regress on answers we would propagate the uncertainty that the model displayed into future predictions.</p>\n<p>All aggs and time features are floats that are ran through a linear layer (without bias). They were all normalized to either [-1,1] or [0, 1].</p>\n<p>All other embeddings are categorical.</p>\n<h2>Training:</h2>\n<p>What made a <strong>big</strong> difference early on was our sampling methodology. Unlike most approaches in public kernels, we did not take a user-centric approach to sample our training examples. Instead, we simply sampled based on row_id and would load in the previous window_size history for every row.</p>\n<p>So for instance if we sampled row 55. Then we would find the user id that row 55 corresponds to, and load in their history up until that row. The size of the window would then be the min(window_size, len(history_until_row_55)).</p>\n<p>We used Adam with 0.001 learning rate, early stopping and lr decay based on val_auc during training.</p>\n<h2>Validation</h2>\n<p>For validation, we kept a holdout set of 2.5mil (like the test set) and generated randomly with a certain proportion guaranteed of new users. We used only 250,000 rows to actually validate on during training.</p>\n<p>For every row in validation, we would pick a random number of inference steps between 1 and 10. We would then infer over these steps, not allowing the model to see the true answers and having to use its own.</p>\n<h2>Inference</h2>\n<p>During inference, when predicting multiple questions for a single user, we fed back the previous prediction of our model as answers in the model. This auto-regression was helpful in propagating uncertainty and helped the auc.</p>\n<p>I saw a writeup that said that this was not possible to do because it constrains you to batch size = 1. That is wrong, you can actually do this for a whole batch in parallel and it's a little tricky with indexes but it is doable. You simply have to keep track of each sequence lengths and how many steps you are predicting for each.</p>\n<p>Unfortunately since some of our aggs are also based on user answers, these do not get updated until the next batch because they are calculated on CPU and not updated in that inference loop.</p>\n<h2>Parameters</h2>\n<p>We ensemble three models.</p>\n<p>First model:</p>\n<ul>\n<li>64 emb_dim</li>\n<li>4/6 encoder/decoder layers</li>\n<li>256 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>100 window size</li>\n</ul>\n<p>Second model:</p>\n<ul>\n<li>64 emb_dim</li>\n<li>4/6 encoder/decoder layers</li>\n<li>256 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>200 window size (expanded by finetuning on 200)</li>\n</ul>\n<p>Third model:</p>\n<ul>\n<li>256 emb_dim</li>\n<li>2/2 encoder/decoder layers</li>\n<li>512 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>256 window size</li>\n</ul>\n<h2>Hardware</h2>\n<p>We rented three machines but we could have probably gotten farther with using a larger single machine:</p>\n<ul>\n<li>3 X  1 Tesla V100(16gb VRAM)</li>\n</ul>\n<h2>Other things we tried</h2>\n<ul>\n<li>Noam lr schedule (super slow to converge)</li>\n<li>Linear attention / Reformer / Linformer (spent way to much time on this)</li>\n<li>Local Attention</li>\n<li>Additional Aggs in output layer and a myriad of other aggs</li>\n<li>Concatenating Embeds instead of summing</li>\n<li>A lot of different configurations of num_layers / dims / .. etc</li>\n<li>Custom attention based on bundle </li>\n<li>Ensembling with LGBM</li>\n<li>Predicting time taken to answer question as well as answer in Loss</li>\n<li>K Beam decoding (predicting in parallel K possible paths of answers and taking the one that maximized joint probability of sequence)</li>\n<li>Running average of agg features (different windows or averaging over time)</li>\n<li>Causal 1D convolutions before the Transformer</li>\n<li>Increasing window size during training</li>\n<li>…</li>\n</ul>\n<p>Finally, thanks to <a href=\"https://www.kaggle.com/eeeedev\" target=\"_blank\">@eeeedev</a> for the adventure!</p>",
  "messages": [
    {
      "id": "1145876",
      "postDate": "01/09/2021 11:58:53",
      "content": "<p>Hi all!<br>\nAfter seeing everyone share their solution we have decided to also share ours in the same spirit.</p>\n<p>We ran a modified SAINT+ that <strong>included lectures, tags and a couple of aggregate features.</strong></p>\n<p>All our code is available on github ( <a href=\"https://github.com/gautierdag/riiid\" target=\"_blank\">https://github.com/gautierdag/riiid</a> ). We used Pytorch/Pytorch Lightning and Hydra.</p>\n<p>Our single model public LB was 0.808, and we were able to push that to LB 0.810 (Final LB 0.813) through ensembling.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F583548%2Fa19776338ca1ba57bd1260bf8e0ec678%2Friiid_model.png?generation=1610188697948445&amp;alt=media\" alt=\"\"></p>\n<h2>Modifications to original SAINT:</h2>\n<ul>\n<li><p><strong>Lectures</strong>: lectures were included and had a special trainable embedding vector instead of the answer embedding. The loss was not masked for lecture steps. We tested both with and without, and the difference wasn't crazy. The big benefit they offered was that they made everything much simpler to handle since we did not need to filter anything.</p></li>\n<li><p><strong>Tags</strong>: Since questions had a maximum of six tags, we passed sequences of length 6 to a tag embedding layer and summed the output. Therefore our tag embeddings were learned for each tag but then sum across to obtain an embedding for all the tags in a question/lecture. The padding simply returned a 0 vector.</p></li>\n<li><p><strong>Agg Feats</strong>: This is what gave us a healthy boost of 0.002 on public LB. By including 12 aggregate features in the decoder we were able to use some of our learnings in previous LGBM implementations. </p></li>\n</ul>\n<h4>Agg feats</h4>\n<p>We used the following aggregate features:</p>\n<ul>\n<li>attempts: number of previous attempts on current question (clamped to 5 and normalized to 1)</li>\n<li>mean_content_mean: average of all the average question accuracy (population wise) seen until this step by the user</li>\n<li>mean_user_mean: average of user's accuracy until this step </li>\n<li>parts_mean: seven dimensional (the average of the user's accuracy on each part)</li>\n<li>m_user_session_mean: the average accuracy in the current user's session </li>\n<li>session_number: the current session number (how many sessions up till now)</li>\n</ul>\n<p>Note: session is just a way of dividing up the questions if a difference in time between questions is greater than two hours.</p>\n<h4>Float vs Categorical</h4>\n<p>We switched to use a float representation of answers in the last few days. This seemed to have little effect on our LB auc unfortunately, but had an effect locally. The idea was that since we auto-regress on answers we would propagate the uncertainty that the model displayed into future predictions.</p>\n<p>All aggs and time features are floats that are ran through a linear layer (without bias). They were all normalized to either [-1,1] or [0, 1].</p>\n<p>All other embeddings are categorical.</p>\n<h2>Training:</h2>\n<p>What made a <strong>big</strong> difference early on was our sampling methodology. Unlike most approaches in public kernels, we did not take a user-centric approach to sample our training examples. Instead, we simply sampled based on row_id and would load in the previous window_size history for every row.</p>\n<p>So for instance if we sampled row 55. Then we would find the user id that row 55 corresponds to, and load in their history up until that row. The size of the window would then be the min(window_size, len(history_until_row_55)).</p>\n<p>We used Adam with 0.001 learning rate, early stopping and lr decay based on val_auc during training.</p>\n<h2>Validation</h2>\n<p>For validation, we kept a holdout set of 2.5mil (like the test set) and generated randomly with a certain proportion guaranteed of new users. We used only 250,000 rows to actually validate on during training.</p>\n<p>For every row in validation, we would pick a random number of inference steps between 1 and 10. We would then infer over these steps, not allowing the model to see the true answers and having to use its own.</p>\n<h2>Inference</h2>\n<p>During inference, when predicting multiple questions for a single user, we fed back the previous prediction of our model as answers in the model. This auto-regression was helpful in propagating uncertainty and helped the auc.</p>\n<p>I saw a writeup that said that this was not possible to do because it constrains you to batch size = 1. That is wrong, you can actually do this for a whole batch in parallel and it's a little tricky with indexes but it is doable. You simply have to keep track of each sequence lengths and how many steps you are predicting for each.</p>\n<p>Unfortunately since some of our aggs are also based on user answers, these do not get updated until the next batch because they are calculated on CPU and not updated in that inference loop.</p>\n<h2>Parameters</h2>\n<p>We ensemble three models.</p>\n<p>First model:</p>\n<ul>\n<li>64 emb_dim</li>\n<li>4/6 encoder/decoder layers</li>\n<li>256 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>100 window size</li>\n</ul>\n<p>Second model:</p>\n<ul>\n<li>64 emb_dim</li>\n<li>4/6 encoder/decoder layers</li>\n<li>256 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>200 window size (expanded by finetuning on 200)</li>\n</ul>\n<p>Third model:</p>\n<ul>\n<li>256 emb_dim</li>\n<li>2/2 encoder/decoder layers</li>\n<li>512 feed foward in Transformer</li>\n<li>4 Heads</li>\n<li>256 window size</li>\n</ul>\n<h2>Hardware</h2>\n<p>We rented three machines but we could have probably gotten farther with using a larger single machine:</p>\n<ul>\n<li>3 X  1 Tesla V100(16gb VRAM)</li>\n</ul>\n<h2>Other things we tried</h2>\n<ul>\n<li>Noam lr schedule (super slow to converge)</li>\n<li>Linear attention / Reformer / Linformer (spent way to much time on this)</li>\n<li>Local Attention</li>\n<li>Additional Aggs in output layer and a myriad of other aggs</li>\n<li>Concatenating Embeds instead of summing</li>\n<li>A lot of different configurations of num_layers / dims / .. etc</li>\n<li>Custom attention based on bundle </li>\n<li>Ensembling with LGBM</li>\n<li>Predicting time taken to answer question as well as answer in Loss</li>\n<li>K Beam decoding (predicting in parallel K possible paths of answers and taking the one that maximized joint probability of sequence)</li>\n<li>Running average of agg features (different windows or averaging over time)</li>\n<li>Causal 1D convolutions before the Transformer</li>\n<li>Increasing window size during training</li>\n<li>…</li>\n</ul>\n<p>Finally, thanks to <a href=\"https://www.kaggle.com/eeeedev\" target=\"_blank\">@eeeedev</a> for the adventure!</p>",
      "rawMarkdown": "Hi all!\nAfter seeing everyone share their solution we have decided to also share ours in the same spirit.\n\nWe ran a modified SAINT+ that **included lectures, tags and a couple of aggregate features.**\n\nAll our code is available on github ( https://github.com/gautierdag/riiid ). We used Pytorch/Pytorch Lightning and Hydra.\n\nOur single model public LB was 0.808, and we were able to push that to LB 0.810 (Final LB 0.813) through ensembling.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F583548%2Fa19776338ca1ba57bd1260bf8e0ec678%2Friiid_model.png?generation=1610188697948445&alt=media)\n\n## Modifications to original SAINT:\n\n- **Lectures**: lectures were included and had a special trainable embedding vector instead of the answer embedding. The loss was not masked for lecture steps. We tested both with and without, and the difference wasn't crazy. The big benefit they offered was that they made everything much simpler to handle since we did not need to filter anything.\n\n- **Tags**: Since questions had a maximum of six tags, we passed sequences of length 6 to a tag embedding layer and summed the output. Therefore our tag embeddings were learned for each tag but then sum across to obtain an embedding for all the tags in a question/lecture. The padding simply returned a 0 vector.\n\n- **Agg Feats**: This is what gave us a healthy boost of 0.002 on public LB. By including 12 aggregate features in the decoder we were able to use some of our learnings in previous LGBM implementations. \n\n\n#### Agg feats\n\nWe used the following aggregate features:\n- attempts: number of previous attempts on current question (clamped to 5 and normalized to 1)\n- mean_content_mean: average of all the average question accuracy (population wise) seen until this step by the user\n- mean_user_mean: average of user's accuracy until this step \n- parts_mean: seven dimensional (the average of the user's accuracy on each part)\n- m_user_session_mean: the average accuracy in the current user's session \n- session_number: the current session number (how many sessions up till now)\n\nNote: session is just a way of dividing up the questions if a difference in time between questions is greater than two hours.\n\n#### Float vs Categorical\n\nWe switched to use a float representation of answers in the last few days. This seemed to have little effect on our LB auc unfortunately, but had an effect locally. The idea was that since we auto-regress on answers we would propagate the uncertainty that the model displayed into future predictions.\n\nAll aggs and time features are floats that are ran through a linear layer (without bias). They were all normalized to either [-1,1] or [0, 1].\n\nAll other embeddings are categorical.\n\n## Training:\n\nWhat made a **big** difference early on was our sampling methodology. Unlike most approaches in public kernels, we did not take a user-centric approach to sample our training examples. Instead, we simply sampled based on row_id and would load in the previous window_size history for every row.\n\nSo for instance if we sampled row 55. Then we would find the user id that row 55 corresponds to, and load in their history up until that row. The size of the window would then be the min(window_size, len(history_until_row_55)).\n\nWe used Adam with 0.001 learning rate, early stopping and lr decay based on val_auc during training.\n\n## Validation\n\nFor validation, we kept a holdout set of 2.5mil (like the test set) and generated randomly with a certain proportion guaranteed of new users. We used only 250,000 rows to actually validate on during training.\n\nFor every row in validation, we would pick a random number of inference steps between 1 and 10. We would then infer over these steps, not allowing the model to see the true answers and having to use its own.\n\n## Inference\n\nDuring inference, when predicting multiple questions for a single user, we fed back the previous prediction of our model as answers in the model. This auto-regression was helpful in propagating uncertainty and helped the auc.\n\nI saw a writeup that said that this was not possible to do because it constrains you to batch size = 1. That is wrong, you can actually do this for a whole batch in parallel and it's a little tricky with indexes but it is doable. You simply have to keep track of each sequence lengths and how many steps you are predicting for each.\n \nUnfortunately since some of our aggs are also based on user answers, these do not get updated until the next batch because they are calculated on CPU and not updated in that inference loop.\n\n## Parameters\n\nWe ensemble three models.\n\nFirst model:\n- 64 emb_dim\n- 4/6 encoder/decoder layers\n- 256 feed foward in Transformer\n- 4 Heads\n- 100 window size\n\nSecond model:\n- 64 emb_dim\n- 4/6 encoder/decoder layers\n- 256 feed foward in Transformer\n- 4 Heads\n- 200 window size (expanded by finetuning on 200)\n\nThird model:\n- 256 emb_dim\n- 2/2 encoder/decoder layers\n- 512 feed foward in Transformer\n- 4 Heads\n- 256 window size\n\n\n## Hardware\nWe rented three machines but we could have probably gotten farther with using a larger single machine:\n\n- 3 X  1 Tesla V100(16gb VRAM)\n\n## Other things we tried\n- Noam lr schedule (super slow to converge)\n- Linear attention / Reformer / Linformer (spent way to much time on this)\n- Local Attention\n- Additional Aggs in output layer and a myriad of other aggs\n- Concatenating Embeds instead of summing\n- A lot of different configurations of num_layers / dims / .. etc\n- Custom attention based on bundle \n- Ensembling with LGBM\n- Predicting time taken to answer question as well as answer in Loss\n- K Beam decoding (predicting in parallel K possible paths of answers and taking the one that maximized joint probability of sequence)\n- Running average of agg features (different windows or averaging over time)\n- Causal 1D convolutions before the Transformer\n- Increasing window size during training\n- ...\n\nFinally, thanks to @eeeedev for the adventure!",
      "votes": null
    },
    {
      "id": "1146399",
      "postDate": "01/09/2021 18:40:40",
      "content": "<p>Congratulations and thanks for sharing, interesting sampling method. Was your ensemble just an average of the 3 models? When you compare Adam to Noam was the difference like half the epochs required to converge?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing, interesting sampling method. Was your ensemble just an average of the 3 models? When you compare Adam to Noam was the difference like half the epochs required to converge?",
      "votes": null
    },
    {
      "id": "1146499",
      "postDate": "01/09/2021 20:01:44",
      "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> github link is not working.</p>",
      "rawMarkdown": "brotye github link is not working.",
      "votes": null
    },
    {
      "id": "1146567",
      "postDate": "01/09/2021 21:29:28",
      "content": "<p>Thanks for noticing! The parenthesis was wrongly included in the url - fixed it! </p>",
      "rawMarkdown": "Thanks for noticing! The parenthesis was wrongly included in the url - fixed it!",
      "votes": null
    },
    {
      "id": "1146582",
      "postDate": "01/09/2021 21:40:33",
      "content": "<p>Thank you!</p>\n<p>Yes the ensemble was a simple average of all models, all models were all too close individually to weight one over another.</p>\n<p>The problem with Noam was that because it's a linear increase that needs to be parametrised (you have to decide number of warmup steps) - the network really learns nothing at the beginning. Then the exponential decay hits immediately after reaching your target lr. While this might be good to warmup very deep Transformers, we found that using a simple lr and starting directly at the target lr was much faster. </p>\n<p>I don't remember the exact difference, but probably something like 1.5x faster to <strong>not</strong> use Noam and no loss/auc difference. It would obviously depend on how you parametrize the Noam scheme. The default and advertised Noam scheme is 4k warm up steps but that should also depend on your batch size as well.</p>",
      "rawMarkdown": "Thank you!\n\nYes the ensemble was a simple average of all models, all models were all too close individually to weight one over another.\n\nThe problem with Noam was that because it's a linear increase that needs to be parametrised (you have to decide number of warmup steps) - the network really learns nothing at the beginning. Then the exponential decay hits immediately after reaching your target lr. While this might be good to warmup very deep Transformers, we found that using a simple lr and starting directly at the target lr was much faster. \n\nI don't remember the exact difference, but probably something like 1.5x faster to **not** use Noam and no loss/auc difference. It would obviously depend on how you parametrize the Noam scheme. The default and advertised Noam scheme is 4k warm up steps but that should also depend on your batch size as well.",
      "votes": null
    },
    {
      "id": "1146619",
      "postDate": "01/09/2021 22:16:01",
      "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> I was not able to find the code for preparation of <code>.npy files</code>. Can you point out the file which does that or is it not present in repo ?</p>\n<p>Also, can you please explain a bit about your model fine tuning ?</p>",
      "rawMarkdown": "brotye I was not able to find the code for preparation of `.npy files`. Can you point out the file which does that or is it not present in repo ?\n\nAlso, can you please explain a bit about your model fine tuning ?",
      "votes": null
    },
    {
      "id": "1146672",
      "postDate": "01/09/2021 23:53:22",
      "content": "<p>Congrats and thank you for sharing. I felt that all of the efforts are excellent, considering the characteristic of this competition!</p>",
      "rawMarkdown": "Congrats and thank you for sharing. I felt that all of the efforts are excellent, considering the characteristic of this competition!",
      "votes": null
    },
    {
      "id": "1146893",
      "postDate": "01/10/2021 06:27:47",
      "content": "<p>Nice solution, I also had linear transformers in mind although due to time constraint did not try it. Can you give more details (lb/cv) on the linear transformers? Also, did you happen to try performers?</p>",
      "rawMarkdown": "Nice solution, I also had linear transformers in mind although due to time constraint did not try it. Can you give more details (lb/cv) on the linear transformers? Also, did you happen to try performers?",
      "votes": null
    },
    {
      "id": "1147091",
      "postDate": "01/10/2021 09:29:29",
      "content": "<p>Thanks! </p>\n<p>Well the problem is that I never got any of the alternative attentions to even converge. I used public implementations of them so I'm quite certain it wasn't from a bug. I spent about a week on testing a lot of them (I didn't try the Performer architecture). I think they are much harder to train for some reason on this dataset, this could be due to me not being patient enough in training or having some wrong hyperparameter/initialization scheme. </p>",
      "rawMarkdown": "Thanks! \n\nWell the problem is that I never got any of the alternative attentions to even converge. I used public implementations of them so I'm quite certain it wasn't from a bug. I spent about a week on testing a lot of them (I didn't try the Performer architecture). I think they are much harder to train for some reason on this dataset, this could be due to me not being patient enough in training or having some wrong hyperparameter/initialization scheme.",
      "votes": null
    },
    {
      "id": "1147098",
      "postDate": "01/10/2021 09:37:12",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> you can find the generation of some of the <code>.npy</code> files in the <code>preprocessing.py</code> file. Some are also included in the repo itself in the <code>data</code> folder. </p>\n<p>Apologies if some are missing, if you let me know in particular which one you are referring to I can rectify it. </p>\n<p>There is one I know is missing: <code>questions_lectures_pct.npy</code>, which is a content based feature one, but it is not actually needed to reproduce our results since we did not use it in the end.</p>\n<p>Finetuning was done by simply resuming our best checkpoint after training on (90%) to a window size double of what it had seen (100 -&gt; 200) and the remainder of the training set (10%). This made training faster (since we could train on smaller windows) and models with bigger windows tended to be a little bit more robust, but it did not greatly improve our AUC. </p>",
      "rawMarkdown": "abdurrehman245 you can find the generation of some of the `.npy` files in the `preprocessing.py` file. Some are also included in the repo itself in the `data` folder. \n\nApologies if some are missing, if you let me know in particular which one you are referring to I can rectify it. \n\nThere is one I know is missing: `questions_lectures_pct.npy`, which is a content based feature one, but it is not actually needed to reproduce our results since we did not use it in the end.\n\n\nFinetuning was done by simply resuming our best checkpoint after training on (90%) to a window size double of what it had seen (100 -> 200) and the remainder of the training set (10%). This made training faster (since we could train on smaller windows) and models with bigger windows tended to be a little bit more robust, but it did not greatly improve our AUC.",
      "votes": null
    },
    {
      "id": "1147251",
      "postDate": "01/10/2021 11:39:34",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> 👍</p>\n<p>Nice solution</p>",
      "rawMarkdown": "Congrats @brotye 👍\n\nNice solution",
      "votes": null
    },
    {
      "id": "1147265",
      "postDate": "01/10/2021 11:52:01",
      "content": "<p>Thanks for the detailed response much appreciated</p>",
      "rawMarkdown": "Thanks for the detailed response much appreciated",
      "votes": null
    },
    {
      "id": "1147452",
      "postDate": "01/10/2021 14:15:05",
      "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a>  I think these files preparation code is not present in repo. Can you explain a bit about how you generate these files? </p>\n<pre><code>val_2500000_lec_True.npy\nval_2500000_lec_False.npy\nuser_weights.npy\n</code></pre>\n<p>Also, can you confirm by using <code>neural.ipynb</code> notebook in repo, I can reproduce one of your best single model by specifying the one of the configurations mentioned in description?</p>\n<pre><code>self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)\nself.embed_parts = nn.Embedding(n_part, emb_dim, padding_idx=0)\nself.embed_tags = nn.Embedding(n_tags, emb_dim, padding_idx=188)\n</code></pre>\n<p>I am newbie in transformers so asking some silly questions as well. You used <code>padding_idx</code>in embeddings because we can have some content_id or parts in test data that we don't have in train data ? </p>",
      "rawMarkdown": "brotye  I think these files preparation code is not present in repo. Can you explain a bit about how you generate these files? \n\n```\nval_2500000_lec_True.npy\nval_2500000_lec_False.npy\nuser_weights.npy\n```\nAlso, can you confirm by using `neural.ipynb` notebook in repo, I can reproduce one of your best single model by specifying the one of the configurations mentioned in description?\n\n```\nself.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)\nself.embed_parts = nn.Embedding(n_part, emb_dim, padding_idx=0)\nself.embed_tags = nn.Embedding(n_tags, emb_dim, padding_idx=188)\n```\nI am newbie in transformers so asking some silly questions as well. You used `padding_idx `in embeddings because we can have some content_id or parts in test data that we don't have in train data ?",
      "votes": null
    },
    {
      "id": "1147665",
      "postDate": "01/10/2021 16:30:34",
      "content": "<p><code>val_2500000_lec_True.npy</code> and <code>val_2500000_lec_False.npy</code> are just numpy arrays that describe the generated val set - its just an array of length 2.5mil of row ids that belong to val. </p>\n<p>You don't actually need them since it will get generated randomly when you run the code.</p>\n<p><code>user_weights.npy</code> is not used anymore.</p>\n<p>If you want to reproduce, it will be a little trickier than just running the notebook. You'll have to use the python <code>main.py</code> script. The notebook is unfortunately out of date.</p>\n<p><code>padding_idx</code> is used because we pass batches of variable lengths sequences to the network, so for instance, if we have two sequences (one length 6 and one length 3):</p>\n<p>[[1, 2, 3, 4, 5, 6],<br>\n[1, 2, 3, 0, 0, 0]]</p>\n<p>Then when we pass them through embedding we don't care about the embedding of the padding token so we can just set the output vector to 0 directly.</p>",
      "rawMarkdown": "`val_2500000_lec_True.npy` and `val_2500000_lec_False.npy` are just numpy arrays that describe the generated val set - its just an array of length 2.5mil of row ids that belong to val. \n\nYou don't actually need them since it will get generated randomly when you run the code.\n\n`user_weights.npy` is not used anymore.\n\nIf you want to reproduce, it will be a little trickier than just running the notebook. You'll have to use the python `main.py` script. The notebook is unfortunately out of date.\n\n`padding_idx` is used because we pass batches of variable lengths sequences to the network, so for instance, if we have two sequences (one length 6 and one length 3):\n\n[[1, 2, 3, 4, 5, 6],\n[1, 2, 3, 0, 0, 0]]\n\nThen when we pass them through embedding we don't care about the embedding of the padding token so we can just set the output vector to 0 directly.",
      "votes": null
    },
    {
      "id": "1147724",
      "postDate": "01/10/2021 16:50:01",
      "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> thanks for all the replies.</p>\n<p>So I have to go through the repo files to reproduce them in kaggle notebook.</p>\n<p>As you mentioned in above, you append 0's to sequence to make them of equal length and then we get embedding vector of 0's corresponding to <code>padding_idx</code> so this means in case of <code>content_id</code> sequnence whose length is less than <code>MAX_SEQ_LEN</code> you pad them with <code>13942</code> to get vector of 0's for those positions. Why you did not use the same <code>padding_idx</code> in each embedding instead of <code>0, 13942,188</code>. Is there any specific reason of it ?</p>\n<p><code>self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)</code></p>\n<p>Did you try to feed the <code>aggregated features</code> to encoder as well? I am just trying to understand the intuition of feeding them to decoder layer.</p>\n<p>How did you come up with the hyperparameters( by some hyperparameter turning or some other method) ? I am asking this because most of the solution used 8 heads as mentioned in Attention paper but you used 4 heads in your solution.</p>",
      "rawMarkdown": "brotye thanks for all the replies.\n\nSo I have to go through the repo files to reproduce them in kaggle notebook.\n\nAs you mentioned in above, you append 0's to sequence to make them of equal length and then we get embedding vector of 0's corresponding to `padding_idx ` so this means in case of `content_id ` sequnence whose length is less than `MAX_SEQ_LEN` you pad them with `13942` to get vector of 0's for those positions. Why you did not use the same `padding_idx ` in each embedding instead of `0, 13942,188`. Is there any specific reason of it ?\n\n`self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)`\n\nDid you try to feed the `aggregated features` to encoder as well? I am just trying to understand the intuition of feeding them to decoder layer.\n\nHow did you come up with the hyperparameters( by some hyperparameter turning or some other method) ? I am asking this because most of the solution used 8 heads as mentioned in Attention paper but you used 4 heads in your solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1146399,
      "author_name": "darren1515",
      "author_url": "",
      "post_date": "01/09/2021 18:40:40",
      "content": "<p>Congratulations and thanks for sharing, interesting sampling method. Was your ensemble just an average of the 3 models? When you compare Adam to Noam was the difference like half the epochs required to converge?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146582,
          "author_name": "brotye",
          "author_url": "",
          "post_date": "01/09/2021 21:40:33",
          "content": "<p>Thank you!</p>\n<p>Yes the ensemble was a simple average of all models, all models were all too close individually to weight one over another.</p>\n<p>The problem with Noam was that because it's a linear increase that needs to be parametrised (you have to decide number of warmup steps) - the network really learns nothing at the beginning. Then the exponential decay hits immediately after reaching your target lr. While this might be good to warmup very deep Transformers, we found that using a simple lr and starting directly at the target lr was much faster. </p>\n<p>I don't remember the exact difference, but probably something like 1.5x faster to <strong>not</strong> use Noam and no loss/auc difference. It would obviously depend on how you parametrize the Noam scheme. The default and advertised Noam scheme is 4k warm up steps but that should also depend on your batch size as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147265,
          "author_name": "darren1515",
          "author_url": "",
          "post_date": "01/10/2021 11:52:01",
          "content": "<p>Thanks for the detailed response much appreciated</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146499,
      "author_name": "abdurrehman245",
      "author_url": "",
      "post_date": "01/09/2021 20:01:44",
      "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> github link is not working.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146567,
          "author_name": "brotye",
          "author_url": "",
          "post_date": "01/09/2021 21:29:28",
          "content": "<p>Thanks for noticing! The parenthesis was wrongly included in the url - fixed it! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146619,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "01/09/2021 22:16:01",
          "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> I was not able to find the code for preparation of <code>.npy files</code>. Can you point out the file which does that or is it not present in repo ?</p>\n<p>Also, can you please explain a bit about your model fine tuning ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147098,
          "author_name": "brotye",
          "author_url": "",
          "post_date": "01/10/2021 09:37:12",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> you can find the generation of some of the <code>.npy</code> files in the <code>preprocessing.py</code> file. Some are also included in the repo itself in the <code>data</code> folder. </p>\n<p>Apologies if some are missing, if you let me know in particular which one you are referring to I can rectify it. </p>\n<p>There is one I know is missing: <code>questions_lectures_pct.npy</code>, which is a content based feature one, but it is not actually needed to reproduce our results since we did not use it in the end.</p>\n<p>Finetuning was done by simply resuming our best checkpoint after training on (90%) to a window size double of what it had seen (100 -&gt; 200) and the remainder of the training set (10%). This made training faster (since we could train on smaller windows) and models with bigger windows tended to be a little bit more robust, but it did not greatly improve our AUC. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147452,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "01/10/2021 14:15:05",
          "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a>  I think these files preparation code is not present in repo. Can you explain a bit about how you generate these files? </p>\n<pre><code>val_2500000_lec_True.npy\nval_2500000_lec_False.npy\nuser_weights.npy\n</code></pre>\n<p>Also, can you confirm by using <code>neural.ipynb</code> notebook in repo, I can reproduce one of your best single model by specifying the one of the configurations mentioned in description?</p>\n<pre><code>self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)\nself.embed_parts = nn.Embedding(n_part, emb_dim, padding_idx=0)\nself.embed_tags = nn.Embedding(n_tags, emb_dim, padding_idx=188)\n</code></pre>\n<p>I am newbie in transformers so asking some silly questions as well. You used <code>padding_idx</code>in embeddings because we can have some content_id or parts in test data that we don't have in train data ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147665,
          "author_name": "brotye",
          "author_url": "",
          "post_date": "01/10/2021 16:30:34",
          "content": "<p><code>val_2500000_lec_True.npy</code> and <code>val_2500000_lec_False.npy</code> are just numpy arrays that describe the generated val set - its just an array of length 2.5mil of row ids that belong to val. </p>\n<p>You don't actually need them since it will get generated randomly when you run the code.</p>\n<p><code>user_weights.npy</code> is not used anymore.</p>\n<p>If you want to reproduce, it will be a little trickier than just running the notebook. You'll have to use the python <code>main.py</code> script. The notebook is unfortunately out of date.</p>\n<p><code>padding_idx</code> is used because we pass batches of variable lengths sequences to the network, so for instance, if we have two sequences (one length 6 and one length 3):</p>\n<p>[[1, 2, 3, 4, 5, 6],<br>\n[1, 2, 3, 0, 0, 0]]</p>\n<p>Then when we pass them through embedding we don't care about the embedding of the padding token so we can just set the output vector to 0 directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1147724,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "01/10/2021 16:50:01",
          "content": "<p><a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> thanks for all the replies.</p>\n<p>So I have to go through the repo files to reproduce them in kaggle notebook.</p>\n<p>As you mentioned in above, you append 0's to sequence to make them of equal length and then we get embedding vector of 0's corresponding to <code>padding_idx</code> so this means in case of <code>content_id</code> sequnence whose length is less than <code>MAX_SEQ_LEN</code> you pad them with <code>13942</code> to get vector of 0's for those positions. Why you did not use the same <code>padding_idx</code> in each embedding instead of <code>0, 13942,188</code>. Is there any specific reason of it ?</p>\n<p><code>self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)</code></p>\n<p>Did you try to feed the <code>aggregated features</code> to encoder as well? I am just trying to understand the intuition of feeding them to decoder layer.</p>\n<p>How did you come up with the hyperparameters( by some hyperparameter turning or some other method) ? I am asking this because most of the solution used 8 heads as mentioned in Attention paper but you used 4 heads in your solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146672,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/09/2021 23:53:22",
      "content": "<p>Congrats and thank you for sharing. I felt that all of the efforts are excellent, considering the characteristic of this competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146893,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "01/10/2021 06:27:47",
      "content": "<p>Nice solution, I also had linear transformers in mind although due to time constraint did not try it. Can you give more details (lb/cv) on the linear transformers? Also, did you happen to try performers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147091,
          "author_name": "brotye",
          "author_url": "",
          "post_date": "01/10/2021 09:29:29",
          "content": "<p>Thanks! </p>\n<p>Well the problem is that I never got any of the alternative attentions to even converge. I used public implementations of them so I'm quite certain it wasn't from a bug. I spent about a week on testing a lot of them (I didn't try the Performer architecture). I think they are much harder to train for some reason on this dataset, this could be due to me not being patient enough in training or having some wrong hyperparameter/initialization scheme. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147251,
      "author_name": "jagadish13",
      "author_url": "",
      "post_date": "01/10/2021 11:39:34",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/brotye\" target=\"_blank\">@brotye</a> 👍</p>\n<p>Nice solution</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1145876": "Hi all!\nAfter seeing everyone share their solution we have decided to also share ours in the same spirit.\n\nWe ran a modified SAINT+ that **included lectures, tags and a couple of aggregate features.**\n\nAll our code is available on github ( https://github.com/gautierdag/riiid ). We used Pytorch/Pytorch Lightning and Hydra.\n\nOur single model public LB was 0.808, and we were able to push that to LB 0.810 (Final LB 0.813) through ensembling.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F583548%2Fa19776338ca1ba57bd1260bf8e0ec678%2Friiid_model.png?generation=1610188697948445&alt=media)\n\n## Modifications to original SAINT:\n\n- **Lectures**: lectures were included and had a special trainable embedding vector instead of the answer embedding. The loss was not masked for lecture steps. We tested both with and without, and the difference wasn't crazy. The big benefit they offered was that they made everything much simpler to handle since we did not need to filter anything.\n\n- **Tags**: Since questions had a maximum of six tags, we passed sequences of length 6 to a tag embedding layer and summed the output. Therefore our tag embeddings were learned for each tag but then sum across to obtain an embedding for all the tags in a question/lecture. The padding simply returned a 0 vector.\n\n- **Agg Feats**: This is what gave us a healthy boost of 0.002 on public LB. By including 12 aggregate features in the decoder we were able to use some of our learnings in previous LGBM implementations. \n\n\n#### Agg feats\n\nWe used the following aggregate features:\n- attempts: number of previous attempts on current question (clamped to 5 and normalized to 1)\n- mean_content_mean: average of all the average question accuracy (population wise) seen until this step by the user\n- mean_user_mean: average of user's accuracy until this step \n- parts_mean: seven dimensional (the average of the user's accuracy on each part)\n- m_user_session_mean: the average accuracy in the current user's session \n- session_number: the current session number (how many sessions up till now)\n\nNote: session is just a way of dividing up the questions if a difference in time between questions is greater than two hours.\n\n#### Float vs Categorical\n\nWe switched to use a float representation of answers in the last few days. This seemed to have little effect on our LB auc unfortunately, but had an effect locally. The idea was that since we auto-regress on answers we would propagate the uncertainty that the model displayed into future predictions.\n\nAll aggs and time features are floats that are ran through a linear layer (without bias). They were all normalized to either [-1,1] or [0, 1].\n\nAll other embeddings are categorical.\n\n## Training:\n\nWhat made a **big** difference early on was our sampling methodology. Unlike most approaches in public kernels, we did not take a user-centric approach to sample our training examples. Instead, we simply sampled based on row_id and would load in the previous window_size history for every row.\n\nSo for instance if we sampled row 55. Then we would find the user id that row 55 corresponds to, and load in their history up until that row. The size of the window would then be the min(window_size, len(history_until_row_55)).\n\nWe used Adam with 0.001 learning rate, early stopping and lr decay based on val_auc during training.\n\n## Validation\n\nFor validation, we kept a holdout set of 2.5mil (like the test set) and generated randomly with a certain proportion guaranteed of new users. We used only 250,000 rows to actually validate on during training.\n\nFor every row in validation, we would pick a random number of inference steps between 1 and 10. We would then infer over these steps, not allowing the model to see the true answers and having to use its own.\n\n## Inference\n\nDuring inference, when predicting multiple questions for a single user, we fed back the previous prediction of our model as answers in the model. This auto-regression was helpful in propagating uncertainty and helped the auc.\n\nI saw a writeup that said that this was not possible to do because it constrains you to batch size = 1. That is wrong, you can actually do this for a whole batch in parallel and it's a little tricky with indexes but it is doable. You simply have to keep track of each sequence lengths and how many steps you are predicting for each.\n \nUnfortunately since some of our aggs are also based on user answers, these do not get updated until the next batch because they are calculated on CPU and not updated in that inference loop.\n\n## Parameters\n\nWe ensemble three models.\n\nFirst model:\n- 64 emb_dim\n- 4/6 encoder/decoder layers\n- 256 feed foward in Transformer\n- 4 Heads\n- 100 window size\n\nSecond model:\n- 64 emb_dim\n- 4/6 encoder/decoder layers\n- 256 feed foward in Transformer\n- 4 Heads\n- 200 window size (expanded by finetuning on 200)\n\nThird model:\n- 256 emb_dim\n- 2/2 encoder/decoder layers\n- 512 feed foward in Transformer\n- 4 Heads\n- 256 window size\n\n\n## Hardware\nWe rented three machines but we could have probably gotten farther with using a larger single machine:\n\n- 3 X  1 Tesla V100(16gb VRAM)\n\n## Other things we tried\n- Noam lr schedule (super slow to converge)\n- Linear attention / Reformer / Linformer (spent way to much time on this)\n- Local Attention\n- Additional Aggs in output layer and a myriad of other aggs\n- Concatenating Embeds instead of summing\n- A lot of different configurations of num_layers / dims / .. etc\n- Custom attention based on bundle \n- Ensembling with LGBM\n- Predicting time taken to answer question as well as answer in Loss\n- K Beam decoding (predicting in parallel K possible paths of answers and taking the one that maximized joint probability of sequence)\n- Running average of agg features (different windows or averaging over time)\n- Causal 1D convolutions before the Transformer\n- Increasing window size during training\n- ...\n\nFinally, thanks to @eeeedev for the adventure!",
    "1146399": "Congratulations and thanks for sharing, interesting sampling method. Was your ensemble just an average of the 3 models? When you compare Adam to Noam was the difference like half the epochs required to converge?",
    "1146499": "brotye github link is not working.",
    "1146567": "Thanks for noticing! The parenthesis was wrongly included in the url - fixed it!",
    "1146582": "Thank you!\n\nYes the ensemble was a simple average of all models, all models were all too close individually to weight one over another.\n\nThe problem with Noam was that because it's a linear increase that needs to be parametrised (you have to decide number of warmup steps) - the network really learns nothing at the beginning. Then the exponential decay hits immediately after reaching your target lr. While this might be good to warmup very deep Transformers, we found that using a simple lr and starting directly at the target lr was much faster. \n\nI don't remember the exact difference, but probably something like 1.5x faster to **not** use Noam and no loss/auc difference. It would obviously depend on how you parametrize the Noam scheme. The default and advertised Noam scheme is 4k warm up steps but that should also depend on your batch size as well.",
    "1146619": "brotye I was not able to find the code for preparation of `.npy files`. Can you point out the file which does that or is it not present in repo ?\n\nAlso, can you please explain a bit about your model fine tuning ?",
    "1146672": "Congrats and thank you for sharing. I felt that all of the efforts are excellent, considering the characteristic of this competition!",
    "1146893": "Nice solution, I also had linear transformers in mind although due to time constraint did not try it. Can you give more details (lb/cv) on the linear transformers? Also, did you happen to try performers?",
    "1147091": "Thanks! \n\nWell the problem is that I never got any of the alternative attentions to even converge. I used public implementations of them so I'm quite certain it wasn't from a bug. I spent about a week on testing a lot of them (I didn't try the Performer architecture). I think they are much harder to train for some reason on this dataset, this could be due to me not being patient enough in training or having some wrong hyperparameter/initialization scheme.",
    "1147098": "abdurrehman245 you can find the generation of some of the `.npy` files in the `preprocessing.py` file. Some are also included in the repo itself in the `data` folder. \n\nApologies if some are missing, if you let me know in particular which one you are referring to I can rectify it. \n\nThere is one I know is missing: `questions_lectures_pct.npy`, which is a content based feature one, but it is not actually needed to reproduce our results since we did not use it in the end.\n\n\nFinetuning was done by simply resuming our best checkpoint after training on (90%) to a window size double of what it had seen (100 -> 200) and the remainder of the training set (10%). This made training faster (since we could train on smaller windows) and models with bigger windows tended to be a little bit more robust, but it did not greatly improve our AUC.",
    "1147251": "Congrats @brotye 👍\n\nNice solution",
    "1147265": "Thanks for the detailed response much appreciated",
    "1147452": "brotye  I think these files preparation code is not present in repo. Can you explain a bit about how you generate these files? \n\n```\nval_2500000_lec_True.npy\nval_2500000_lec_False.npy\nuser_weights.npy\n```\nAlso, can you confirm by using `neural.ipynb` notebook in repo, I can reproduce one of your best single model by specifying the one of the configurations mentioned in description?\n\n```\nself.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)\nself.embed_parts = nn.Embedding(n_part, emb_dim, padding_idx=0)\nself.embed_tags = nn.Embedding(n_tags, emb_dim, padding_idx=188)\n```\nI am newbie in transformers so asking some silly questions as well. You used `padding_idx `in embeddings because we can have some content_id or parts in test data that we don't have in train data ?",
    "1147665": "`val_2500000_lec_True.npy` and `val_2500000_lec_False.npy` are just numpy arrays that describe the generated val set - its just an array of length 2.5mil of row ids that belong to val. \n\nYou don't actually need them since it will get generated randomly when you run the code.\n\n`user_weights.npy` is not used anymore.\n\nIf you want to reproduce, it will be a little trickier than just running the notebook. You'll have to use the python `main.py` script. The notebook is unfortunately out of date.\n\n`padding_idx` is used because we pass batches of variable lengths sequences to the network, so for instance, if we have two sequences (one length 6 and one length 3):\n\n[[1, 2, 3, 4, 5, 6],\n[1, 2, 3, 0, 0, 0]]\n\nThen when we pass them through embedding we don't care about the embedding of the padding token so we can just set the output vector to 0 directly.",
    "1147724": "brotye thanks for all the replies.\n\nSo I have to go through the repo files to reproduce them in kaggle notebook.\n\nAs you mentioned in above, you append 0's to sequence to make them of equal length and then we get embedding vector of 0's corresponding to `padding_idx ` so this means in case of `content_id ` sequnence whose length is less than `MAX_SEQ_LEN` you pad them with `13942` to get vector of 0's for those positions. Why you did not use the same `padding_idx ` in each embedding instead of `0, 13942,188`. Is there any specific reason of it ?\n\n`self.embed_content_id = nn.Embedding(n_content_id, emb_dim, padding_idx=13942)`\n\nDid you try to feed the `aggregated features` to encoder as well? I am just trying to understand the intuition of feeding them to decoder layer.\n\nHow did you come up with the hyperparameters( by some hyperparameter turning or some other method) ? I am asking this because most of the solution used 8 heads as mentioned in Attention paper but you used 4 heads in your solution."
  },
  "source": "meta"
}