{
  "id": 209689,
  "title": "38th place solution: Single SAINT+",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209689",
  "author_name": "Furkan Ömerustaoğlu",
  "post_date": "2021-01-08T08:40:43.118000",
  "votes": 36,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, thank you to the host and Kaggle team for the competition. I think it is a very useful challenge for those who want to try transformer models outside of the nlp. I would also like to thank everyone who shared their ideas for this competition. I can say that it contributed to my learning journey in many ways.</p>\n<p>Thank you to Tito and his great <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">notebook</a> about loop for feature engineering<br>\nThank you to Mark and his great <a href=\"https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state\" target=\"_blank\">notebook</a> for the harmonic mean idea. </p>\n<h2>Results</h2>\n<p>My local train/val AUC: <strong>0.804</strong> / <strong>0.806</strong><br>\nI already knew the difference between my public and private lb score, because of the private leak.<br>\nAs a CV strategy, I used a similar but not the same as tito.</p>\n<h2>Model details</h2>\n<p>The model structure is very much the same as <a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a> paper, except below parameters.</p>\n<ul>\n<li>2 encoder - 2 decoder layers</li>\n<li>256 embed dim</li>\n<li>0.1 dropout for attention layers, 0.2 dropout for fully connected heads at the end of the encoder-decoder layers</li>\n<li>100 seq length</li>\n</ul>\n<p><strong>Encoder embeddings (Known features)</strong></p>\n<ul>\n<li>Pos embedding (Categorical)</li>\n<li>Question id embedding (Categorical)</li>\n<li>Bundle id embedding (Categorical)</li>\n<li>Part embedding (Categorical)</li>\n<li>Question encoded tag embedding (New number for each unique value) (Categorical)</li>\n<li>Question attempt embedding (Categorical)</li>\n<li>Lag time embedding (current task timestamp - prev task timestamp) (Categorical)</li>\n<li>Harmonic mean embedding (Continuous)</li>\n</ul>\n<p><strong>Decoder embeddings (Future features)</strong><br>\n<em>All features shifted/adjusted</em></p>\n<ul>\n<li>Pos embedding - same embedding layer for encoder (Categorical)</li>\n<li>Answer embedding (Categorical)</li>\n<li>Elapsed time (Categorical)</li>\n<li>Answer bundle id - same embedding layer for bundle id (Categorical)</li>\n</ul>\n<h2>Training:</h2>\n<ul>\n<li>25 epoch, train almost took 7-8 hours on GPU, inference 3-4 hours</li>\n<li>Adam with 1e-3 lr for first 20 epoch, then 1e-4 last 5 epoch</li>\n<li>128 batch size</li>\n</ul>\n<p>Notebook <a href=\"https://www.kaggle.com/fomdata/38th-solution-saint-model\" target=\"_blank\">link</a></p>",
  "messages": [
    {
      "id": 1144115,
      "postDate": "2021-01-08T08:40:43.120Z",
      "content": "<p>First of all, thank you to the host and Kaggle team for the competition. I think it is a very useful challenge for those who want to try transformer models outside of the nlp. I would also like to thank everyone who shared their ideas for this competition. I can say that it contributed to my learning journey in many ways.</p>\n<p>Thank you to Tito and his great <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">notebook</a> about loop for feature engineering<br>\nThank you to Mark and his great <a href=\"https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state\" target=\"_blank\">notebook</a> for the harmonic mean idea. </p>\n<h2>Results</h2>\n<p>My local train/val AUC: <strong>0.804</strong> / <strong>0.806</strong><br>\nI already knew the difference between my public and private lb score, because of the private leak.<br>\nAs a CV strategy, I used a similar but not the same as tito.</p>\n<h2>Model details</h2>\n<p>The model structure is very much the same as <a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a> paper, except below parameters.</p>\n<ul>\n<li>2 encoder - 2 decoder layers</li>\n<li>256 embed dim</li>\n<li>0.1 dropout for attention layers, 0.2 dropout for fully connected heads at the end of the encoder-decoder layers</li>\n<li>100 seq length</li>\n</ul>\n<p><strong>Encoder embeddings (Known features)</strong></p>\n<ul>\n<li>Pos embedding (Categorical)</li>\n<li>Question id embedding (Categorical)</li>\n<li>Bundle id embedding (Categorical)</li>\n<li>Part embedding (Categorical)</li>\n<li>Question encoded tag embedding (New number for each unique value) (Categorical)</li>\n<li>Question attempt embedding (Categorical)</li>\n<li>Lag time embedding (current task timestamp - prev task timestamp) (Categorical)</li>\n<li>Harmonic mean embedding (Continuous)</li>\n</ul>\n<p><strong>Decoder embeddings (Future features)</strong><br>\n<em>All features shifted/adjusted</em></p>\n<ul>\n<li>Pos embedding - same embedding layer for encoder (Categorical)</li>\n<li>Answer embedding (Categorical)</li>\n<li>Elapsed time (Categorical)</li>\n<li>Answer bundle id - same embedding layer for bundle id (Categorical)</li>\n</ul>\n<h2>Training:</h2>\n<ul>\n<li>25 epoch, train almost took 7-8 hours on GPU, inference 3-4 hours</li>\n<li>Adam with 1e-3 lr for first 20 epoch, then 1e-4 last 5 epoch</li>\n<li>128 batch size</li>\n</ul>\n<p>Notebook <a href=\"https://www.kaggle.com/fomdata/38th-solution-saint-model\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "First of all, thank you to the host and Kaggle team for the competition. I think it is a very useful challenge for those who want to try transformer models outside of the nlp. I would also like to thank everyone who shared their ideas for this competition. I can say that it contributed to my learning journey in many ways.\n\nThank you to Tito and his great [notebook](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) about loop for feature engineering\nThank you to Mark and his great [notebook](https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state) for the harmonic mean idea. \n\nResults\n- \nMy local train/val AUC: **0.804** / **0.806**\nI already knew the difference between my public and private lb score, because of the private leak.\nAs a CV strategy, I used a similar but not the same as tito.\nModel details\n- \nThe model structure is very much the same as [SAINT+](https://arxiv.org/abs/2010.12042) paper, except below parameters.\n- 2 encoder - 2 decoder layers\n- 256 embed dim\n- 0.1 dropout for attention layers, 0.2 dropout for fully connected heads at the end of the encoder-decoder layers\n- 100 seq length\n\n**Encoder embeddings (Known features)**\n- Pos embedding (Categorical)\n- Question id embedding (Categorical)\n- Bundle id embedding (Categorical)\n- Part embedding (Categorical)\n- Question encoded tag embedding (New number for each unique value) (Categorical)\n- Question attempt embedding (Categorical)\n- Lag time embedding (current task timestamp - prev task timestamp) (Categorical)\n- Harmonic mean embedding (Continuous)\n\n**Decoder embeddings (Future features)**\n*All features shifted/adjusted*\n- Pos embedding - same embedding layer for encoder (Categorical)\n- Answer embedding (Categorical)\n- Elapsed time (Categorical)\n- Answer bundle id - same embedding layer for bundle id (Categorical)\n\nTraining:\n- \n- 25 epoch, train almost took 7-8 hours on GPU, inference 3-4 hours\n- Adam with 1e-3 lr for first 20 epoch, then 1e-4 last 5 epoch\n- 128 batch size\n\nNotebook [link](https://www.kaggle.com/fomdata/38th-solution-saint-model)",
      "votes": 36
    },
    {
      "id": 1144229,
      "postDate": "2021-01-08T10:35:04.693Z",
      "content": "<p><a href=\"https://www.kaggle.com/fomdata\" target=\"_blank\">@fomdata</a> congrats on achieving the medal. Would you like to share your code ?</p>",
      "rawMarkdown": "@fomdata congrats on achieving the medal. Would you like to share your code ?",
      "replies": [
        {
          "id": 1144447,
          "postDate": "2021-01-08T13:19:21.437Z",
          "content": "<p>Notebook <a href=\"https://www.kaggle.com/fomdata/38th-solution-saint-model\" target=\"_blank\">link</a> added</p>",
          "rawMarkdown": "Notebook [link](https://www.kaggle.com/fomdata/38th-solution-saint-model) added",
          "votes": 1
        }
      ]
    },
    {
      "id": 1144247,
      "postDate": "2021-01-08T10:49:44.420Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1144229,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2021-01-08T10:35:04.693000",
      "content": "<p><a href=\"https://www.kaggle.com/fomdata\" target=\"_blank\">@fomdata</a> congrats on achieving the medal. Would you like to share your code ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1144447,
          "author_name": "Furkan Ömerustaoğlu",
          "author_url": "",
          "post_date": "2021-01-08T13:19:21.437000",
          "content": "<p>Notebook <a href=\"https://www.kaggle.com/fomdata/38th-solution-saint-model\" target=\"_blank\">link</a> added</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1144247,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-08T10:49:44.420000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144115": "First of all, thank you to the host and Kaggle team for the competition. I think it is a very useful challenge for those who want to try transformer models outside of the nlp. I would also like to thank everyone who shared their ideas for this competition. I can say that it contributed to my learning journey in many ways.\n\nThank you to Tito and his great [notebook](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering) about loop for feature engineering\nThank you to Mark and his great [notebook](https://www.kaggle.com/markwijkhuizen/riiid-training-and-prediction-using-a-state) for the harmonic mean idea. \n\nResults\n- \nMy local train/val AUC: **0.804** / **0.806**\nI already knew the difference between my public and private lb score, because of the private leak.\nAs a CV strategy, I used a similar but not the same as tito.\nModel details\n- \nThe model structure is very much the same as [SAINT+](https://arxiv.org/abs/2010.12042) paper, except below parameters.\n- 2 encoder - 2 decoder layers\n- 256 embed dim\n- 0.1 dropout for attention layers, 0.2 dropout for fully connected heads at the end of the encoder-decoder layers\n- 100 seq length\n\n**Encoder embeddings (Known features)**\n- Pos embedding (Categorical)\n- Question id embedding (Categorical)\n- Bundle id embedding (Categorical)\n- Part embedding (Categorical)\n- Question encoded tag embedding (New number for each unique value) (Categorical)\n- Question attempt embedding (Categorical)\n- Lag time embedding (current task timestamp - prev task timestamp) (Categorical)\n- Harmonic mean embedding (Continuous)\n\n**Decoder embeddings (Future features)**\n*All features shifted/adjusted*\n- Pos embedding - same embedding layer for encoder (Categorical)\n- Answer embedding (Categorical)\n- Elapsed time (Categorical)\n- Answer bundle id - same embedding layer for bundle id (Categorical)\n\nTraining:\n- \n- 25 epoch, train almost took 7-8 hours on GPU, inference 3-4 hours\n- Adam with 1e-3 lr for first 20 epoch, then 1e-4 last 5 epoch\n- 128 batch size\n\nNotebook [link](https://www.kaggle.com/fomdata/38th-solution-saint-model)",
    "1144229": "@fomdata congrats on achieving the medal. Would you like to share your code ?",
    "1144247": ""
  }
}