{
  "id": 209599,
  "title": "76th Place Solution | Minimal SAINT+ Model!",
  "url": "/competitions/riiid-test-answer-prediction/writeups/a105-76th-place-solution-minimal-saint-model",
  "author_name": "",
  "post_date": "2021-07-28T14:08:19.700Z",
  "votes": 22,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the organizers of this very interesting and difficult competition. Congratulations to all the winners. I have learned a lot in this competition. I am not a novice myself but not an expert either. This competition had some of the best discussion threads I could ask for. 2 days back I wasn't sure if my team is even gonna get a medal let alone in the top 100. I should say it is all thanks to the simple SAINT+ model we have build and the efforts of my teammates <a href=\"https://www.kaggle.com/vishwajeet993511\" target=\"_blank\">@vishwajeet993511</a> and <a href=\"https://www.kaggle.com/padma3\" target=\"_blank\">@padma3</a>. We are calling our model minimal because this is the first version we could come up with. Didn't get to try any hyperparameter tuning due to the constraint of time and computation power.</p>\n<h2>SAINT+</h2>\n<h3>Encoder</h3>\n<p>We used the following features in the encoder model:</p>\n<ul>\n<li><code>question_id</code>: Embedding: 0 for padding</li>\n<li><code>part_id</code>: Embedding: 0 for padding</li>\n<li><code>correct_answer</code>: Embedding: 0 for padding. we thought that by giving both questions actual answers and the user's previous answers, the model should be able to come with a pattern and improve the prediction accuracy.</li>\n<li><code>lsi_id</code>: Embedding: 0 for padding. I took lsi tags from <a href=\"https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated\" target=\"_blank\">this popular notebook</a>. </li>\n<li><code>question_accuracy</code>: Linear: This and the below features are adding value in LightGBM. Giving the transformer model this info might have also helped. This is the mean value of accuracy for the question_id.</li>\n<li><code>question_elapsed</code>: Linear: Quantile Transformed (Normal, 100 quantiles). Large Inputs to Neural Networks make it hard to converge. From my previous competition (MoA), I understood this much, so used it here. I can confirm that the convergence was faster. Could have tried bucketing as specified in the paper but didn't get time. This is the mean value of elapsed time for the question_id.</li>\n<li><code>question_lagtime</code>: Linear. Quantile Transformed (Normal, 100 quantiles). I have discussed this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>. This is the mean value of lagtime for the question_id.</li>\n<li><code>part_accuracy</code>: Linear</li>\n<li><code>lsi_accuracy</code>: Linear</li>\n</ul>\n<p>We take to add the embedding above and send it to the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">transformer model</a> as src input.</p>\n<h3>Decoder</h3>\n<p>We used the following features in the decoder model:</p>\n<ul>\n<li><code>responses</code>: Embedding: 0/1 feature. Shifted back by 1 interaction. I had also tried the way SAKT encodes responses but it wasn't optimal for me.</li>\n<li><code>user_answer</code>: Embedding. Shifted back by 1 interaction.</li>\n<li><code>prior_question_had_explanation</code>: Embedding. </li>\n<li><code>elapsed_time</code>: Linear: Quantile Transformed (Normal, 100 quantiles).</li>\n<li><code>lagtime</code>: Linear: Quantile Transformed (Normal, 100 quantiles).</li>\n</ul>\n<p>We take to add the embedding above and send it to the transformer model as tgt input.</p>\n<h3>Masking</h3>\n<ul>\n<li>Masking is very important for transformer models. A lot of people have posted an AUC of 1.0. This is due to bad masking. I have discussed masking <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>.</li>\n<li>Should have used future masking for encoder and task_container_id based masking for the decoder.</li>\n<li>1 thing I would like to highlight is that, look for the dtype of masking. Changing from bool to float makes inverses the mask. This created a lot of pain for me. </li>\n</ul>\n<h3>Parameters</h3>\n<p>MAX_SEQ = 180<br>\nNUM_ENCODER_LAYERS = 2<br>\nNUM_DECODER_LAYERS = 2<br>\nEMBED_SIZE = D_MODEL = 512<br>\nBATCH_SIZE = 128<br>\nDROPOUT = 0.1<br>\nTEST_SIZE = 0.1<br>\nLEARNING_RATE = 2e-4<br>\nEPOCHS = 20<br>\nEARLY_STOPPING_ROUNDS = 3</p>\n<p>We tried out embedding sizes of 128 and 256. They were good but increasing the size helped us. Initially, we were using a higher learning rate, but the results were not optimal.</p>\n<h3>Training</h3>\n<ul>\n<li>We understood that training a transformer model is not easy. The learning rate, optimizer, and scheduler are very important. Try to use a small learning rate with Adam optimizer or a variant of it. Initially we used <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR\" target=\"_blank\">OneCycleLR</a> but we shifted to <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\" target=\"_blank\">CosineAnnealingWarmRestarts</a></li>\n<li>We were constrained by the Kaggle GPUs. Just training the model 4 times was enough to complete a week of our GPU quota.</li>\n<li>Training each model took around 20-30 minutes per epoch. 5 to 6 hours in total.</li>\n<li>We weren't able to reach early stopping in our best model. </li>\n<li>I discussed debugging my model <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>.</li>\n<li>The single model gives CV 0.792, Public 0.793 and <strong>Private 0.795</strong></li>\n<li>If we had time, we wanted to do <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209633\" target=\"_blank\">pretraining of SAINT+</a> with the next question prediction task. </li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>We had our LightGBM pipeline ready but weren't able to spend training LightGBM models. Our features are very basic. In the end, I used the public popular LightGBM model. 1 change we did is that instead of using the first x million data as input, we used the last 20/30 million rows. Since the test data immediately start from the end of the train data, it makes sense to use the last part instead of the first. </li>\n<li>It gave us a boost of 0.004 to the SAINT+ model we had.</li>\n<li>We also tried to use ELO and FTRL models but they were getting 0 weights in our bagging experiments we did on the last day.</li>\n<li>One thing we missed is that we should have used my SAKT model in the ensemble. But we weren't sure if it could have added value as it is very similar to SAINT.</li>\n</ul>\n<h2>Final Words</h2>\n<ul>\n<li>Thanks to all the people who were sharing good notebooks and sharing their ideas in the discussion forum. Even if I hadn't got a medal, I would have been happy with the things I learned in this competition. </li>\n<li>For the next tabular competition, I would like to team up with other competitors who are good at building Boosting Models. I am pretty bad at feature engineering for LightGBM kind of models. I want to learn it. </li>\n<li>I will be uploading my code to GitHub. Once I have done this, I will add the link.</li>\n<li>I will be becoming a <strong>Kaggle Expert</strong>. Would like to try for a GOLD medal in the next one 😆. Kaggle Master, here I come.</li>\n<li>Thanks a lot to everyone who participated. This is a good experience. </li>\n</ul>\n<p>(edited)</p>",
  "messages": [
    {
      "id": "1143623",
      "postDate": "01/08/2021 01:17:54",
      "content": "<p>Thanks to the organizers of this very interesting and difficult competition. Congratulations to all the winners. I have learned a lot in this competition. I am not a novice myself but not an expert either. This competition had some of the best discussion threads I could ask for. 2 days back I wasn't sure if my team is even gonna get a medal let alone in the top 100. I should say it is all thanks to the simple SAINT+ model we have build and the efforts of my teammates <a href=\"https://www.kaggle.com/vishwajeet993511\" target=\"_blank\">@vishwajeet993511</a> and <a href=\"https://www.kaggle.com/padma3\" target=\"_blank\">@padma3</a>. We are calling our model minimal because this is the first version we could come up with. Didn't get to try any hyperparameter tuning due to the constraint of time and computation power.</p>\n<h2>SAINT+</h2>\n<h3>Encoder</h3>\n<p>We used the following features in the encoder model:</p>\n<ul>\n<li><code>question_id</code>: Embedding: 0 for padding</li>\n<li><code>part_id</code>: Embedding: 0 for padding</li>\n<li><code>correct_answer</code>: Embedding: 0 for padding. we thought that by giving both questions actual answers and the user's previous answers, the model should be able to come with a pattern and improve the prediction accuracy.</li>\n<li><code>lsi_id</code>: Embedding: 0 for padding. I took lsi tags from <a href=\"https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated\" target=\"_blank\">this popular notebook</a>. </li>\n<li><code>question_accuracy</code>: Linear: This and the below features are adding value in LightGBM. Giving the transformer model this info might have also helped. This is the mean value of accuracy for the question_id.</li>\n<li><code>question_elapsed</code>: Linear: Quantile Transformed (Normal, 100 quantiles). Large Inputs to Neural Networks make it hard to converge. From my previous competition (MoA), I understood this much, so used it here. I can confirm that the convergence was faster. Could have tried bucketing as specified in the paper but didn't get time. This is the mean value of elapsed time for the question_id.</li>\n<li><code>question_lagtime</code>: Linear. Quantile Transformed (Normal, 100 quantiles). I have discussed this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>. This is the mean value of lagtime for the question_id.</li>\n<li><code>part_accuracy</code>: Linear</li>\n<li><code>lsi_accuracy</code>: Linear</li>\n</ul>\n<p>We take to add the embedding above and send it to the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">transformer model</a> as src input.</p>\n<h3>Decoder</h3>\n<p>We used the following features in the decoder model:</p>\n<ul>\n<li><code>responses</code>: Embedding: 0/1 feature. Shifted back by 1 interaction. I had also tried the way SAKT encodes responses but it wasn't optimal for me.</li>\n<li><code>user_answer</code>: Embedding. Shifted back by 1 interaction.</li>\n<li><code>prior_question_had_explanation</code>: Embedding. </li>\n<li><code>elapsed_time</code>: Linear: Quantile Transformed (Normal, 100 quantiles).</li>\n<li><code>lagtime</code>: Linear: Quantile Transformed (Normal, 100 quantiles).</li>\n</ul>\n<p>We take to add the embedding above and send it to the transformer model as tgt input.</p>\n<h3>Masking</h3>\n<ul>\n<li>Masking is very important for transformer models. A lot of people have posted an AUC of 1.0. This is due to bad masking. I have discussed masking <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>.</li>\n<li>Should have used future masking for encoder and task_container_id based masking for the decoder.</li>\n<li>1 thing I would like to highlight is that, look for the dtype of masking. Changing from bool to float makes inverses the mask. This created a lot of pain for me. </li>\n</ul>\n<h3>Parameters</h3>\n<p>MAX_SEQ = 180<br>\nNUM_ENCODER_LAYERS = 2<br>\nNUM_DECODER_LAYERS = 2<br>\nEMBED_SIZE = D_MODEL = 512<br>\nBATCH_SIZE = 128<br>\nDROPOUT = 0.1<br>\nTEST_SIZE = 0.1<br>\nLEARNING_RATE = 2e-4<br>\nEPOCHS = 20<br>\nEARLY_STOPPING_ROUNDS = 3</p>\n<p>We tried out embedding sizes of 128 and 256. They were good but increasing the size helped us. Initially, we were using a higher learning rate, but the results were not optimal.</p>\n<h3>Training</h3>\n<ul>\n<li>We understood that training a transformer model is not easy. The learning rate, optimizer, and scheduler are very important. Try to use a small learning rate with Adam optimizer or a variant of it. Initially we used <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR\" target=\"_blank\">OneCycleLR</a> but we shifted to <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\" target=\"_blank\">CosineAnnealingWarmRestarts</a></li>\n<li>We were constrained by the Kaggle GPUs. Just training the model 4 times was enough to complete a week of our GPU quota.</li>\n<li>Training each model took around 20-30 minutes per epoch. 5 to 6 hours in total.</li>\n<li>We weren't able to reach early stopping in our best model. </li>\n<li>I discussed debugging my model <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303\" target=\"_blank\">in this thread</a>.</li>\n<li>The single model gives CV 0.792, Public 0.793 and <strong>Private 0.795</strong></li>\n<li>If we had time, we wanted to do <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209633\" target=\"_blank\">pretraining of SAINT+</a> with the next question prediction task. </li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>We had our LightGBM pipeline ready but weren't able to spend training LightGBM models. Our features are very basic. In the end, I used the public popular LightGBM model. 1 change we did is that instead of using the first x million data as input, we used the last 20/30 million rows. Since the test data immediately start from the end of the train data, it makes sense to use the last part instead of the first. </li>\n<li>It gave us a boost of 0.004 to the SAINT+ model we had.</li>\n<li>We also tried to use ELO and FTRL models but they were getting 0 weights in our bagging experiments we did on the last day.</li>\n<li>One thing we missed is that we should have used my SAKT model in the ensemble. But we weren't sure if it could have added value as it is very similar to SAINT.</li>\n</ul>\n<h2>Final Words</h2>\n<ul>\n<li>Thanks to all the people who were sharing good notebooks and sharing their ideas in the discussion forum. Even if I hadn't got a medal, I would have been happy with the things I learned in this competition. </li>\n<li>For the next tabular competition, I would like to team up with other competitors who are good at building Boosting Models. I am pretty bad at feature engineering for LightGBM kind of models. I want to learn it. </li>\n<li>I will be uploading my code to GitHub. Once I have done this, I will add the link.</li>\n<li>I will be becoming a <strong>Kaggle Expert</strong>. Would like to try for a GOLD medal in the next one 😆. Kaggle Master, here I come.</li>\n<li>Thanks a lot to everyone who participated. This is a good experience. </li>\n</ul>\n<p>(edited)</p>",
      "rawMarkdown": "Thanks to the organizers of this very interesting and difficult competition. Congratulations to all the winners. I have learned a lot in this competition. I am not a novice myself but not an expert either. This competition had some of the best discussion threads I could ask for. 2 days back I wasn't sure if my team is even gonna get a medal let alone in the top 100. I should say it is all thanks to the simple SAINT+ model we have build and the efforts of my teammates @vishwajeet993511 and @padma3. We are calling our model minimal because this is the first version we could come up with. Didn't get to try any hyperparameter tuning due to the constraint of time and computation power.\n\n## SAINT+\n### Encoder\nWe used the following features in the encoder model:\n- `question_id`: Embedding: 0 for padding\n- `part_id`: Embedding: 0 for padding\n- `correct_answer`: Embedding: 0 for padding. we thought that by giving both questions actual answers and the user's previous answers, the model should be able to come with a pattern and improve the prediction accuracy.\n- `lsi_id`: Embedding: 0 for padding. I took lsi tags from [this popular notebook](https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated). \n- `question_accuracy`: Linear: This and the below features are adding value in LightGBM. Giving the transformer model this info might have also helped. This is the mean value of accuracy for the question_id.\n- `question_elapsed`: Linear: Quantile Transformed (Normal, 100 quantiles). Large Inputs to Neural Networks make it hard to converge. From my previous competition (MoA), I understood this much, so used it here. I can confirm that the convergence was faster. Could have tried bucketing as specified in the paper but didn't get time. This is the mean value of elapsed time for the question_id.\n- `question_lagtime`: Linear. Quantile Transformed (Normal, 100 quantiles). I have discussed this [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303). This is the mean value of lagtime for the question_id.\n- `part_accuracy`: Linear\n- `lsi_accuracy`: Linear\n\nWe take to add the embedding above and send it to the [transformer model](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html) as src input.\n\n### Decoder\nWe used the following features in the decoder model:\n- `responses`: Embedding: 0/1 feature. Shifted back by 1 interaction. I had also tried the way SAKT encodes responses but it wasn't optimal for me.\n- `user_answer`: Embedding. Shifted back by 1 interaction.\n- `prior_question_had_explanation`: Embedding. \n- `elapsed_time`: Linear: Quantile Transformed (Normal, 100 quantiles).\n- `lagtime`: Linear: Quantile Transformed (Normal, 100 quantiles).\n\nWe take to add the embedding above and send it to the transformer model as tgt input.\n\n### Masking\n- Masking is very important for transformer models. A lot of people have posted an AUC of 1.0. This is due to bad masking. I have discussed masking [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303).\n- Should have used future masking for encoder and task_container_id based masking for the decoder.\n- 1 thing I would like to highlight is that, look for the dtype of masking. Changing from bool to float makes inverses the mask. This created a lot of pain for me. \n\n### Parameters\nMAX_SEQ = 180\nNUM_ENCODER_LAYERS = 2\nNUM_DECODER_LAYERS = 2\nEMBED_SIZE = D_MODEL = 512\nBATCH_SIZE = 128\nDROPOUT = 0.1\nTEST_SIZE = 0.1\nLEARNING_RATE = 2e-4\nEPOCHS = 20\nEARLY_STOPPING_ROUNDS = 3\n\nWe tried out embedding sizes of 128 and 256. They were good but increasing the size helped us. Initially, we were using a higher learning rate, but the results were not optimal.\n\n### Training\n- We understood that training a transformer model is not easy. The learning rate, optimizer, and scheduler are very important. Try to use a small learning rate with Adam optimizer or a variant of it. Initially we used [OneCycleLR](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR) but we shifted to [CosineAnnealingWarmRestarts](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.CosineAnnealingWarmRestarts)\n- We were constrained by the Kaggle GPUs. Just training the model 4 times was enough to complete a week of our GPU quota.\n- Training each model took around 20-30 minutes per epoch. 5 to 6 hours in total.\n- We weren't able to reach early stopping in our best model. \n- I discussed debugging my model [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303).\n- The single model gives CV 0.792, Public 0.793 and **Private 0.795**\n- If we had time, we wanted to do [pretraining of SAINT+](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209633) with the next question prediction task. \n\n## Ensemble\n- We had our LightGBM pipeline ready but weren't able to spend training LightGBM models. Our features are very basic. In the end, I used the public popular LightGBM model. 1 change we did is that instead of using the first x million data as input, we used the last 20/30 million rows. Since the test data immediately start from the end of the train data, it makes sense to use the last part instead of the first. \n- It gave us a boost of 0.004 to the SAINT+ model we had.\n- We also tried to use ELO and FTRL models but they were getting 0 weights in our bagging experiments we did on the last day.\n- One thing we missed is that we should have used my SAKT model in the ensemble. But we weren't sure if it could have added value as it is very similar to SAINT.\n\n## Final Words\n- Thanks to all the people who were sharing good notebooks and sharing their ideas in the discussion forum. Even if I hadn't got a medal, I would have been happy with the things I learned in this competition. \n- For the next tabular competition, I would like to team up with other competitors who are good at building Boosting Models. I am pretty bad at feature engineering for LightGBM kind of models. I want to learn it. \n- I will be uploading my code to GitHub. Once I have done this, I will add the link.\n- I will be becoming a **Kaggle Expert**. Would like to try for a GOLD medal in the next one 😆. Kaggle Master, here I come.\n- Thanks a lot to everyone who participated. This is a good experience. \n\n(edited)",
      "votes": null
    },
    {
      "id": "1145642",
      "postDate": "01/09/2021 09:20:43",
      "content": "<p>congratulations！</p>",
      "rawMarkdown": "congratulations！",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1145642,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/09/2021 09:20:43",
      "content": "<p>congratulations！</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143623": "Thanks to the organizers of this very interesting and difficult competition. Congratulations to all the winners. I have learned a lot in this competition. I am not a novice myself but not an expert either. This competition had some of the best discussion threads I could ask for. 2 days back I wasn't sure if my team is even gonna get a medal let alone in the top 100. I should say it is all thanks to the simple SAINT+ model we have build and the efforts of my teammates @vishwajeet993511 and @padma3. We are calling our model minimal because this is the first version we could come up with. Didn't get to try any hyperparameter tuning due to the constraint of time and computation power.\n\n## SAINT+\n### Encoder\nWe used the following features in the encoder model:\n- `question_id`: Embedding: 0 for padding\n- `part_id`: Embedding: 0 for padding\n- `correct_answer`: Embedding: 0 for padding. we thought that by giving both questions actual answers and the user's previous answers, the model should be able to come with a pattern and improve the prediction accuracy.\n- `lsi_id`: Embedding: 0 for padding. I took lsi tags from [this popular notebook](https://www.kaggle.com/gaozhanfire/riiid-lgbm-val-0-788-offline-and-lsimodel-updated). \n- `question_accuracy`: Linear: This and the below features are adding value in LightGBM. Giving the transformer model this info might have also helped. This is the mean value of accuracy for the question_id.\n- `question_elapsed`: Linear: Quantile Transformed (Normal, 100 quantiles). Large Inputs to Neural Networks make it hard to converge. From my previous competition (MoA), I understood this much, so used it here. I can confirm that the convergence was faster. Could have tried bucketing as specified in the paper but didn't get time. This is the mean value of elapsed time for the question_id.\n- `question_lagtime`: Linear. Quantile Transformed (Normal, 100 quantiles). I have discussed this [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303). This is the mean value of lagtime for the question_id.\n- `part_accuracy`: Linear\n- `lsi_accuracy`: Linear\n\nWe take to add the embedding above and send it to the [transformer model](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html) as src input.\n\n### Decoder\nWe used the following features in the decoder model:\n- `responses`: Embedding: 0/1 feature. Shifted back by 1 interaction. I had also tried the way SAKT encodes responses but it wasn't optimal for me.\n- `user_answer`: Embedding. Shifted back by 1 interaction.\n- `prior_question_had_explanation`: Embedding. \n- `elapsed_time`: Linear: Quantile Transformed (Normal, 100 quantiles).\n- `lagtime`: Linear: Quantile Transformed (Normal, 100 quantiles).\n\nWe take to add the embedding above and send it to the transformer model as tgt input.\n\n### Masking\n- Masking is very important for transformer models. A lot of people have posted an AUC of 1.0. This is due to bad masking. I have discussed masking [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303).\n- Should have used future masking for encoder and task_container_id based masking for the decoder.\n- 1 thing I would like to highlight is that, look for the dtype of masking. Changing from bool to float makes inverses the mask. This created a lot of pain for me. \n\n### Parameters\nMAX_SEQ = 180\nNUM_ENCODER_LAYERS = 2\nNUM_DECODER_LAYERS = 2\nEMBED_SIZE = D_MODEL = 512\nBATCH_SIZE = 128\nDROPOUT = 0.1\nTEST_SIZE = 0.1\nLEARNING_RATE = 2e-4\nEPOCHS = 20\nEARLY_STOPPING_ROUNDS = 3\n\nWe tried out embedding sizes of 128 and 256. They were good but increasing the size helped us. Initially, we were using a higher learning rate, but the results were not optimal.\n\n### Training\n- We understood that training a transformer model is not easy. The learning rate, optimizer, and scheduler are very important. Try to use a small learning rate with Adam optimizer or a variant of it. Initially we used [OneCycleLR](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR) but we shifted to [CosineAnnealingWarmRestarts](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.CosineAnnealingWarmRestarts)\n- We were constrained by the Kaggle GPUs. Just training the model 4 times was enough to complete a week of our GPU quota.\n- Training each model took around 20-30 minutes per epoch. 5 to 6 hours in total.\n- We weren't able to reach early stopping in our best model. \n- I discussed debugging my model [in this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209303).\n- The single model gives CV 0.792, Public 0.793 and **Private 0.795**\n- If we had time, we wanted to do [pretraining of SAINT+](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209633) with the next question prediction task. \n\n## Ensemble\n- We had our LightGBM pipeline ready but weren't able to spend training LightGBM models. Our features are very basic. In the end, I used the public popular LightGBM model. 1 change we did is that instead of using the first x million data as input, we used the last 20/30 million rows. Since the test data immediately start from the end of the train data, it makes sense to use the last part instead of the first. \n- It gave us a boost of 0.004 to the SAINT+ model we had.\n- We also tried to use ELO and FTRL models but they were getting 0 weights in our bagging experiments we did on the last day.\n- One thing we missed is that we should have used my SAKT model in the ensemble. But we weren't sure if it could have added value as it is very similar to SAINT.\n\n## Final Words\n- Thanks to all the people who were sharing good notebooks and sharing their ideas in the discussion forum. Even if I hadn't got a medal, I would have been happy with the things I learned in this competition. \n- For the next tabular competition, I would like to team up with other competitors who are good at building Boosting Models. I am pretty bad at feature engineering for LightGBM kind of models. I want to learn it. \n- I will be uploading my code to GitHub. Once I have done this, I will add the link.\n- I will be becoming a **Kaggle Expert**. Would like to try for a GOLD medal in the next one 😆. Kaggle Master, here I come.\n- Thanks a lot to everyone who participated. This is a good experience. \n\n(edited)",
    "1145642": "congratulations！"
  },
  "source": "meta"
}