{
  "id": 209624,
  "title": "37th place solution (Transformer part)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209624",
  "author_name": "Dean",
  "post_date": "2021-01-08T03:53:14.441000",
  "votes": 25,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank oganizer and Kaggle for hosting the competition. In addition, thanks for <a href=\"https://www.kaggle.com/vk00st\" target=\"_blank\">@vk00st</a> inviting us to merge as a team, and other teammates <a href=\"https://www.kaggle.com/unkoff\" target=\"_blank\">@unkoff</a> <a href=\"https://www.kaggle.com/arvissu\" target=\"_blank\">@arvissu</a> <a href=\"https://www.kaggle.com/syxuming\" target=\"_blank\">@syxuming</a>, we have great team work.</p>\n<p>I also want to thank <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a> <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> great SAKT starter notebook. I can’t implement the Transformer model by myself without these.</p>\n<p>I summarize the detail of our SAINT+ like model as follows.</p>\n<h2>Input features</h2>\n<ol>\n<li>content_id</li>\n<li>response (answered_correctly)</li>\n<li>part</li>\n<li>prior_question_elapsed_time</li>\n<li>prior_question_had_explanation</li>\n<li>lag_time1 - convert lag time to seconds. if lag_time1 &gt;= 300 than 300.</li>\n<li>lag_time2 - convert lag time to minutes. if lag_time2 &gt;= 1440 than 300 (one day).</li>\n<li>lag_time3 - convert lag time to days. if lag_time3 &gt;= 365 than 365 (one year).</li>\n</ol>\n<p>Lag time split to different time format boosting score around 0.003.</p>\n<h2>Distinguish zero padding</h2>\n<ul>\n<li>In order to distinguish zero padding and zero of <strong>content_id</strong>, I add one to <strong>content_id</strong> and <strong>prior_question_had_explanation</strong> and <strong>response</strong>. It will help model distinguish zero padding and features. It boost score around 0.003.</li>\n</ul>\n<pre><code>q_ = q_+1\npri_exp_ = pri_exp_+1\nres_ = qa_+1\n</code></pre>\n<h3>Transformer</h3>\n<h5>Encoder Input</h5>\n<ul>\n<li>question embedding</li>\n<li>part embedding</li>\n<li>position embedding</li>\n<li>prior question had explanation embedding</li>\n</ul>\n<h5>Decoder Input</h5>\n<ul>\n<li>position embedding</li>\n<li>reponse embedding</li>\n<li>prior elapsed time embedding</li>\n<li>lag_time1 categorical embedding</li>\n<li>lag_time2 categorical embedding</li>\n<li>lag_time3 categorical embedding</li>\n<li>Note that I tried categorical and continuous embedding in prior elapsed time and lag time. The performance of categorical embedding is better than continuous embedding.</li>\n</ul>\n<h5>Parameter of  Transformer</h5>\n<ul>\n<li>max sequence: 100</li>\n<li>d model: 256 </li>\n<li>number of layer of encoder: 2</li>\n<li>number of layer of decoder: 2</li>\n<li>batch size: 256</li>\n<li>dropout: 0.1</li>\n<li>learning rate: 5e-4 with AdamW</li>\n</ul>\n<h3>Blending</h3>\n<ul>\n<li>LGBM + Catboost + three Transformer</li>\n<li>LGBM + two Transformer</li>\n</ul>\n<h3>Result</h3>\n<ul>\n<li>Finally, our team get LGBM public LB 0.794 and Transformer public LB 0.799.</li>\n<li>Get public LB 0.804 with blending.</li>\n</ul>\n<h3>Sample code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/m10515009/saint-is-all-you-need-training-private-0-801\" target=\"_blank\">Training</a></li>\n<li><a href=\"https://www.kaggle.com/m10515009/saint-is-all-you-need-inference-private-0-801\" target=\"_blank\">Inference</a></li>\n</ul>",
  "messages": [
    {
      "id": 1143765,
      "postDate": "2021-01-08T03:53:14.443Z",
      "content": "<p>First of all, I would like to thank oganizer and Kaggle for hosting the competition. In addition, thanks for <a href=\"https://www.kaggle.com/vk00st\" target=\"_blank\">@vk00st</a> inviting us to merge as a team, and other teammates <a href=\"https://www.kaggle.com/unkoff\" target=\"_blank\">@unkoff</a> <a href=\"https://www.kaggle.com/arvissu\" target=\"_blank\">@arvissu</a> <a href=\"https://www.kaggle.com/syxuming\" target=\"_blank\">@syxuming</a>, we have great team work.</p>\n<p>I also want to thank <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a> <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> great SAKT starter notebook. I can’t implement the Transformer model by myself without these.</p>\n<p>I summarize the detail of our SAINT+ like model as follows.</p>\n<h2>Input features</h2>\n<ol>\n<li>content_id</li>\n<li>response (answered_correctly)</li>\n<li>part</li>\n<li>prior_question_elapsed_time</li>\n<li>prior_question_had_explanation</li>\n<li>lag_time1 - convert lag time to seconds. if lag_time1 &gt;= 300 than 300.</li>\n<li>lag_time2 - convert lag time to minutes. if lag_time2 &gt;= 1440 than 300 (one day).</li>\n<li>lag_time3 - convert lag time to days. if lag_time3 &gt;= 365 than 365 (one year).</li>\n</ol>\n<p>Lag time split to different time format boosting score around 0.003.</p>\n<h2>Distinguish zero padding</h2>\n<ul>\n<li>In order to distinguish zero padding and zero of <strong>content_id</strong>, I add one to <strong>content_id</strong> and <strong>prior_question_had_explanation</strong> and <strong>response</strong>. It will help model distinguish zero padding and features. It boost score around 0.003.</li>\n</ul>\n<pre><code>q_ = q_+1\npri_exp_ = pri_exp_+1\nres_ = qa_+1\n</code></pre>\n<h3>Transformer</h3>\n<h5>Encoder Input</h5>\n<ul>\n<li>question embedding</li>\n<li>part embedding</li>\n<li>position embedding</li>\n<li>prior question had explanation embedding</li>\n</ul>\n<h5>Decoder Input</h5>\n<ul>\n<li>position embedding</li>\n<li>reponse embedding</li>\n<li>prior elapsed time embedding</li>\n<li>lag_time1 categorical embedding</li>\n<li>lag_time2 categorical embedding</li>\n<li>lag_time3 categorical embedding</li>\n<li>Note that I tried categorical and continuous embedding in prior elapsed time and lag time. The performance of categorical embedding is better than continuous embedding.</li>\n</ul>\n<h5>Parameter of  Transformer</h5>\n<ul>\n<li>max sequence: 100</li>\n<li>d model: 256 </li>\n<li>number of layer of encoder: 2</li>\n<li>number of layer of decoder: 2</li>\n<li>batch size: 256</li>\n<li>dropout: 0.1</li>\n<li>learning rate: 5e-4 with AdamW</li>\n</ul>\n<h3>Blending</h3>\n<ul>\n<li>LGBM + Catboost + three Transformer</li>\n<li>LGBM + two Transformer</li>\n</ul>\n<h3>Result</h3>\n<ul>\n<li>Finally, our team get LGBM public LB 0.794 and Transformer public LB 0.799.</li>\n<li>Get public LB 0.804 with blending.</li>\n</ul>\n<h3>Sample code</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/m10515009/saint-is-all-you-need-training-private-0-801\" target=\"_blank\">Training</a></li>\n<li><a href=\"https://www.kaggle.com/m10515009/saint-is-all-you-need-inference-private-0-801\" target=\"_blank\">Inference</a></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank oganizer and Kaggle for hosting the competition. In addition, thanks for @vk00st inviting us to merge as a team, and other teammates @unkoff @arvissu @syxuming, we have great team work.\n\nI also want to thank @manikanthr5 @wangsg @leadbest great SAKT starter notebook. I can’t implement the Transformer model by myself without these.\n\nI summarize the detail of our SAINT+ like model as follows.\n\n\n## Input features\n\n1. content_id\n2. response (answered_correctly)\n3. part\n4. prior_question_elapsed_time\n5. prior_question_had_explanation\n6. lag_time1 - convert lag time to seconds. if lag_time1 >= 300 than 300.\n7. lag_time2 - convert lag time to minutes. if lag_time2 >= 1440 than 300 (one day).\n8. lag_time3 - convert lag time to days. if lag_time3 >= 365 than 365 (one year).\n\nLag time split to different time format boosting score around 0.003.\n\n## Distinguish zero padding\n* In order to distinguish zero padding and zero of **content_id**, I add one to **content_id** and **prior_question_had_explanation** and **response**. It will help model distinguish zero padding and features. It boost score around 0.003.\n\n```\nq_ = q_+1\npri_exp_ = pri_exp_+1\nres_ = qa_+1\n```\n### Transformer\n\n##### Encoder Input\n* question embedding\n* part embedding\n* position embedding\n* prior question had explanation embedding\n\n##### Decoder Input\n\n* position embedding\n* reponse embedding\n* prior elapsed time embedding\n* lag_time1 categorical embedding\n* lag_time2 categorical embedding\n* lag_time3 categorical embedding\n* Note that I tried categorical and continuous embedding in prior elapsed time and lag time. The performance of categorical embedding is better than continuous embedding.\n\n##### Parameter of  Transformer\n* max sequence: 100\n* d model: 256 \n* number of layer of encoder: 2\n* number of layer of decoder: 2\n* batch size: 256\n* dropout: 0.1\n* learning rate: 5e-4 with AdamW\n\n### Blending\n\n* LGBM + Catboost + three Transformer\n* LGBM + two Transformer\n\n### Result\n\n* Finally, our team get LGBM public LB 0.794 and Transformer public LB 0.799.\n* Get public LB 0.804 with blending.\n\n### Sample code\n* [Training](https://www.kaggle.com/m10515009/saint-is-all-you-need-training-private-0-801)\n* [Inference](https://www.kaggle.com/m10515009/saint-is-all-you-need-inference-private-0-801)",
      "votes": 24
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1143765": "First of all, I would like to thank oganizer and Kaggle for hosting the competition. In addition, thanks for @vk00st inviting us to merge as a team, and other teammates @unkoff @arvissu @syxuming, we have great team work.\n\nI also want to thank @manikanthr5 @wangsg @leadbest great SAKT starter notebook. I can’t implement the Transformer model by myself without these.\n\nI summarize the detail of our SAINT+ like model as follows.\n\n\n## Input features\n\n1. content_id\n2. response (answered_correctly)\n3. part\n4. prior_question_elapsed_time\n5. prior_question_had_explanation\n6. lag_time1 - convert lag time to seconds. if lag_time1 >= 300 than 300.\n7. lag_time2 - convert lag time to minutes. if lag_time2 >= 1440 than 300 (one day).\n8. lag_time3 - convert lag time to days. if lag_time3 >= 365 than 365 (one year).\n\nLag time split to different time format boosting score around 0.003.\n\n## Distinguish zero padding\n* In order to distinguish zero padding and zero of **content_id**, I add one to **content_id** and **prior_question_had_explanation** and **response**. It will help model distinguish zero padding and features. It boost score around 0.003.\n\n```\nq_ = q_+1\npri_exp_ = pri_exp_+1\nres_ = qa_+1\n```\n### Transformer\n\n##### Encoder Input\n* question embedding\n* part embedding\n* position embedding\n* prior question had explanation embedding\n\n##### Decoder Input\n\n* position embedding\n* reponse embedding\n* prior elapsed time embedding\n* lag_time1 categorical embedding\n* lag_time2 categorical embedding\n* lag_time3 categorical embedding\n* Note that I tried categorical and continuous embedding in prior elapsed time and lag time. The performance of categorical embedding is better than continuous embedding.\n\n##### Parameter of  Transformer\n* max sequence: 100\n* d model: 256 \n* number of layer of encoder: 2\n* number of layer of decoder: 2\n* batch size: 256\n* dropout: 0.1\n* learning rate: 5e-4 with AdamW\n\n### Blending\n\n* LGBM + Catboost + three Transformer\n* LGBM + two Transformer\n\n### Result\n\n* Finally, our team get LGBM public LB 0.794 and Transformer public LB 0.799.\n* Get public LB 0.804 with blending.\n\n### Sample code\n* [Training](https://www.kaggle.com/m10515009/saint-is-all-you-need-training-private-0-801)\n* [Inference](https://www.kaggle.com/m10515009/saint-is-all-you-need-inference-private-0-801)"
  }
}