{
  "id": 209581,
  "title": "6th Place Solution: Very Custom GRU",
  "url": "/competitions/riiid-test-answer-prediction/writeups/ahmet-erdem-6th-place-solution-very-custom-gru",
  "author_name": "",
  "post_date": "2021-01-08T12:08:09.110Z",
  "votes": 146,
  "comment_count": 24,
  "views": 0,
  "content": "<p>First of all congrats to <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> and <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and other top teams. I will shortly explain my solution. It is late here, so I may be missing some points.</p>\n<p><strong>Some General Details</strong></p>\n<ul>\n<li>Used questions only. Lectures improve validation score but increase train-val gap and don’t improve LB.</li>\n<li>Didn't use dataframes. Used numpy arrays partitioned by user ids.</li>\n<li>Sequence length 256</li>\n<li>15 epochs. 1/3 of the data each time with reducing LR.</li>\n<li>8192 batch size</li>\n<li>Ensemble of the same model with 7 different seeds trained on whole data</li>\n<li>0.8136 single model validation score, 0.813 LB. Ensemble: 0.815.</li>\n<li>8 hours training on 4-GPU machine</li>\n<li>Used Github and committed any improvement with a message like: Add one more GRU layer (Val: 0.8136, LB: 0.813)</li>\n</ul>\n<p><strong>Inputs</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F4e161bd44dd25c8329c5ced0fa6c9ab3%2Fimg1.png?generation=1610064408784528&amp;alt=media\" alt=\"\"></p>\n<p><strong>Engineered Features</strong></p>\n<p>Assume current question’s correct answer is X. Logarithm of:</p>\n<ul>\n<li>Number of questions since last X.</li>\n<li>Length of current X streak. (can be zero)</li>\n<li>Length of current streak on any non-X answer. (can be zero)</li>\n</ul>\n<p>This helps with users who always pick A as answer etc.</p>\n<p><strong>Embeddings</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F8256c5fe356ce1bedeedaf9bc91e9884%2Fimg2.png?generation=1610064559098164&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F73d8b182375f185231d35f8b141b8cbf%2Fimg3.png?generation=1610064576810743&amp;alt=media\" alt=\"\"></p>\n<p>Content Cosine Similarity</p>\n<ul>\n<li>A bit similar to attention with 16 heads</li>\n<li>Linear transformation and l2 norm applied on content vectors</li>\n<li>For 16 different transformation, cosine similarity between current content and history contents are calculated.</li>\n<li>Transformation is symmetric for content and history contents.</li>\n</ul>\n<p>U-GRU:</p>\n<ul>\n<li>GRU with 2 directions but not BiGRU</li>\n<li>First does reverse pass, concatenates the output and then does forward pass</li>\n</ul>\n<p>MLP:</p>\n<ul>\n<li>2 layers of [Linear, BatchNorm, Relu]</li>\n</ul>\n<p>Edit: Part embedding is trainable. There is actually sigmoid x tanh layer before GRUs.</p>",
  "messages": [
    {
      "id": "1143534",
      "postDate": "01/08/2021 00:13:28",
      "content": "<p>First of all congrats to <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> and <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and other top teams. I will shortly explain my solution. It is late here, so I may be missing some points.</p>\n<p><strong>Some General Details</strong></p>\n<ul>\n<li>Used questions only. Lectures improve validation score but increase train-val gap and don’t improve LB.</li>\n<li>Didn't use dataframes. Used numpy arrays partitioned by user ids.</li>\n<li>Sequence length 256</li>\n<li>15 epochs. 1/3 of the data each time with reducing LR.</li>\n<li>8192 batch size</li>\n<li>Ensemble of the same model with 7 different seeds trained on whole data</li>\n<li>0.8136 single model validation score, 0.813 LB. Ensemble: 0.815.</li>\n<li>8 hours training on 4-GPU machine</li>\n<li>Used Github and committed any improvement with a message like: Add one more GRU layer (Val: 0.8136, LB: 0.813)</li>\n</ul>\n<p><strong>Inputs</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F4e161bd44dd25c8329c5ced0fa6c9ab3%2Fimg1.png?generation=1610064408784528&amp;alt=media\" alt=\"\"></p>\n<p><strong>Engineered Features</strong></p>\n<p>Assume current question’s correct answer is X. Logarithm of:</p>\n<ul>\n<li>Number of questions since last X.</li>\n<li>Length of current X streak. (can be zero)</li>\n<li>Length of current streak on any non-X answer. (can be zero)</li>\n</ul>\n<p>This helps with users who always pick A as answer etc.</p>\n<p><strong>Embeddings</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F8256c5fe356ce1bedeedaf9bc91e9884%2Fimg2.png?generation=1610064559098164&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F73d8b182375f185231d35f8b141b8cbf%2Fimg3.png?generation=1610064576810743&amp;alt=media\" alt=\"\"></p>\n<p>Content Cosine Similarity</p>\n<ul>\n<li>A bit similar to attention with 16 heads</li>\n<li>Linear transformation and l2 norm applied on content vectors</li>\n<li>For 16 different transformation, cosine similarity between current content and history contents are calculated.</li>\n<li>Transformation is symmetric for content and history contents.</li>\n</ul>\n<p>U-GRU:</p>\n<ul>\n<li>GRU with 2 directions but not BiGRU</li>\n<li>First does reverse pass, concatenates the output and then does forward pass</li>\n</ul>\n<p>MLP:</p>\n<ul>\n<li>2 layers of [Linear, BatchNorm, Relu]</li>\n</ul>\n<p>Edit: Part embedding is trainable. There is actually sigmoid x tanh layer before GRUs.</p>",
      "rawMarkdown": "First of all congrats to @keetar and @mamasinkgs and other top teams. I will shortly explain my solution. It is late here, so I may be missing some points.\n\n**Some General Details**\n\n- Used questions only. Lectures improve validation score but increase train-val gap and don’t improve LB.\n- Didn't use dataframes. Used numpy arrays partitioned by user ids.\n- Sequence length 256\n- 15 epochs. 1/3 of the data each time with reducing LR.\n- 8192 batch size\n- Ensemble of the same model with 7 different seeds trained on whole data\n- 0.8136 single model validation score, 0.813 LB. Ensemble: 0.815.\n- 8 hours training on 4-GPU machine\n- Used Github and committed any improvement with a message like: Add one more GRU layer (Val: 0.8136, LB: 0.813)\n\n**Inputs**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F4e161bd44dd25c8329c5ced0fa6c9ab3%2Fimg1.png?generation=1610064408784528&alt=media)\n\n**Engineered Features**\n\nAssume current question’s correct answer is X. Logarithm of:\n- Number of questions since last X.\n- Length of current X streak. (can be zero)\n- Length of current streak on any non-X answer. (can be zero)\n\nThis helps with users who always pick A as answer etc.\n\n**Embeddings**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F8256c5fe356ce1bedeedaf9bc91e9884%2Fimg2.png?generation=1610064559098164&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F73d8b182375f185231d35f8b141b8cbf%2Fimg3.png?generation=1610064576810743&alt=media)\n\nContent Cosine Similarity\n- A bit similar to attention with 16 heads\n- Linear transformation and l2 norm applied on content vectors\n- For 16 different transformation, cosine similarity between current content and history contents are calculated.\n- Transformation is symmetric for content and history contents.\n\nU-GRU:\n- GRU with 2 directions but not BiGRU\n- First does reverse pass, concatenates the output and then does forward pass\n\nMLP:\n- 2 layers of [Linear, BatchNorm, Relu]\n\n\nEdit: Part embedding is trainable. There is actually sigmoid x tanh layer before GRUs.",
      "votes": null
    },
    {
      "id": "1143538",
      "postDate": "01/08/2021 00:16:46",
      "content": "<p>Nice solution!</p>",
      "rawMarkdown": "Nice solution!",
      "votes": null
    },
    {
      "id": "1143540",
      "postDate": "01/08/2021 00:17:45",
      "content": "<p>Wow. Extremelly innovative. Congratulations.</p>",
      "rawMarkdown": "Wow. Extremelly innovative. Congratulations.",
      "votes": null
    },
    {
      "id": "1143546",
      "postDate": "01/08/2021 00:20:32",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> Congrats to you and your teammate with the 3rd place!</p>",
      "rawMarkdown": "Thanks @antorsae Congrats to you and your teammate with the 3rd place!",
      "votes": null
    },
    {
      "id": "1143550",
      "postDate": "01/08/2021 00:22:16",
      "content": "<p>Impressive solution. Thank you for sharing.</p>",
      "rawMarkdown": "Impressive solution. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1143553",
      "postDate": "01/08/2021 00:23:27",
      "content": "<p>Great solution, congrats on 6th place <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>!</p>",
      "rawMarkdown": "Great solution, congrats on 6th place @aerdem4!",
      "votes": null
    },
    {
      "id": "1143569",
      "postDate": "01/08/2021 00:33:11",
      "content": "<p>Brilliant approach. Congratulations!</p>",
      "rawMarkdown": "Brilliant approach. Congratulations!",
      "votes": null
    },
    {
      "id": "1143600",
      "postDate": "01/08/2021 00:57:06",
      "content": "<p>Suspect for special award ^^</p>",
      "rawMarkdown": "Suspect for special award ^^",
      "votes": null
    },
    {
      "id": "1143712",
      "postDate": "01/08/2021 02:54:20",
      "content": "<p>Nice solution！Congratulations!</p>",
      "rawMarkdown": "Nice solution！Congratulations!",
      "votes": null
    },
    {
      "id": "1143869",
      "postDate": "01/08/2021 05:37:52",
      "content": "<p>Congrats, simple and efficient model👍</p>",
      "rawMarkdown": "Congrats, simple and efficient model👍",
      "votes": null
    },
    {
      "id": "1143971",
      "postDate": "01/08/2021 06:49:54",
      "content": "<p>Congrats, very nice approach, thanks for sharing.</p>",
      "rawMarkdown": "Congrats, very nice approach, thanks for sharing.",
      "votes": null
    },
    {
      "id": "1144030",
      "postDate": "01/08/2021 07:34:04",
      "content": "<p>Congrats, thank you for sharing, the U-GRU part is very interesting. I see that the dimensions of your embedding are different, how do they blend together?</p>",
      "rawMarkdown": "Congrats, thank you for sharing, the U-GRU part is very interesting. I see that the dimensions of your embedding are different, how do they blend together?",
      "votes": null
    },
    {
      "id": "1144090",
      "postDate": "01/08/2021 08:20:43",
      "content": "<p>Very interesting thanks !</p>",
      "rawMarkdown": "Very interesting thanks !",
      "votes": null
    },
    {
      "id": "1144156",
      "postDate": "01/08/2021 09:12:45",
      "content": "<p>Well done Ahmet, I like the idea of tracking if all <code>A's</code> are chosen - like run length encoding. <br>\nWhat hidden dimension did you use in the GRU to fit 8192 in batch ? </p>",
      "rawMarkdown": "Well done Ahmet, I like the idea of tracking if all `A's` are chosen - like run length encoding. \nWhat hidden dimension did you use in the GRU to fit 8192 in batch ?",
      "votes": null
    },
    {
      "id": "1144260",
      "postDate": "01/08/2021 10:58:02",
      "content": "<p>GRU hidden dimension was 128. Larger GRUs were overfitting. Model size is around 13 MB, which is mostly content embedding weights.</p>",
      "rawMarkdown": "GRU hidden dimension was 128. Larger GRUs were overfitting. Model size is around 13 MB, which is mostly content embedding weights.",
      "votes": null
    },
    {
      "id": "1144262",
      "postDate": "01/08/2021 10:58:31",
      "content": "<p>I concatenate them.</p>",
      "rawMarkdown": "I concatenate them.",
      "votes": null
    },
    {
      "id": "1144351",
      "postDate": "01/08/2021 12:15:10",
      "content": "<p>Congratz, and thanks for the very nice write-up !</p>",
      "rawMarkdown": "Congratz, and thanks for the very nice write-up !",
      "votes": null
    },
    {
      "id": "1144577",
      "postDate": "01/08/2021 14:43:10",
      "content": "<p>Very interesting, thanks for sharing!</p>",
      "rawMarkdown": "Very interesting, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1146437",
      "postDate": "01/09/2021 19:18:04",
      "content": "<p>It's really terrific to get to such a high score without Transformer. Congrats!</p>",
      "rawMarkdown": "It's really terrific to get to such a high score without Transformer. Congrats!",
      "votes": null
    },
    {
      "id": "1146663",
      "postDate": "01/09/2021 23:29:07",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> !! Very innovative sol!! especially your training scheme with the U GRU and the CSS transformation.. are you planning to share any code (full or key parts) ? </p>",
      "rawMarkdown": "Well done @aerdem4 !! Very innovative sol!! especially your training scheme with the U GRU and the CSS transformation.. are you planning to share any code (full or key parts) ?",
      "votes": null
    },
    {
      "id": "1146954",
      "postDate": "01/10/2021 07:57:06",
      "content": "<p>Congratulations !!! 🔥🔥. I have few questions(May be not related to your solution) about getting started with competition as a beginner. </p>\n<ul>\n<li>What are the resources(other than hardware) you have referred during the competition.</li>\n<li>As I am new to the kaggle competitions… Do I need industry level knowledge about the kind of problems given in the competition.<br>\nWhat are some tips to beginners like me.  </li>\n</ul>",
      "rawMarkdown": "Congratulations !!! 🔥🔥. I have few questions(May be not related to your solution) about getting started with competition as a beginner. \n- What are the resources(other than hardware) you have referred during the competition.\n- As I am new to the kaggle competitions... Do I need industry level knowledge about the kind of problems given in the competition.\nWhat are some tips to beginners like me.",
      "votes": null
    },
    {
      "id": "1153925",
      "postDate": "01/15/2021 08:51:58",
      "content": "<p>Brilliant, thanks for sharing and congrats!</p>",
      "rawMarkdown": "Brilliant, thanks for sharing and congrats!",
      "votes": null
    },
    {
      "id": "1153960",
      "postDate": "01/15/2021 09:33:38",
      "content": "<p>Thank you. Normally I share my code but this time it is in notebooks rather than well-organized scripts. I don't know if it will be readable.</p>",
      "rawMarkdown": "Thank you. Normally I share my code but this time it is in notebooks rather than well-organized scripts. I don't know if it will be readable.",
      "votes": null
    },
    {
      "id": "1154270",
      "postDate": "01/15/2021 14:29:45",
      "content": "<p>Thanks, maybe just the model/dataset definition would be helpful and nice learning task for anyone that is interested </p>",
      "rawMarkdown": "Thanks, maybe just the model/dataset definition would be helpful and nice learning task for anyone that is interested",
      "votes": null
    },
    {
      "id": "1156563",
      "postDate": "01/17/2021 08:55:16",
      "content": "<p>I hope my inference kernel can be enough for that, therefore I made it public now: <a href=\"https://www.kaggle.com/aerdem4/riiid-starter\" target=\"_blank\">https://www.kaggle.com/aerdem4/riiid-starter</a></p>",
      "rawMarkdown": "I hope my inference kernel can be enough for that, therefore I made it public now: https://www.kaggle.com/aerdem4/riiid-starter",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143538,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "01/08/2021 00:16:46",
      "content": "<p>Nice solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143540,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/08/2021 00:17:45",
      "content": "<p>Wow. Extremelly innovative. Congratulations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143546,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/08/2021 00:20:32",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> Congrats to you and your teammate with the 3rd place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143550,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 00:22:16",
      "content": "<p>Impressive solution. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143553,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 00:23:27",
      "content": "<p>Great solution, congrats on 6th place <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143569,
      "author_name": "maherelouahabi",
      "author_url": "",
      "post_date": "01/08/2021 00:33:11",
      "content": "<p>Brilliant approach. Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143600,
      "author_name": "killimi",
      "author_url": "",
      "post_date": "01/08/2021 00:57:06",
      "content": "<p>Suspect for special award ^^</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143712,
      "author_name": "a763337092",
      "author_url": "",
      "post_date": "01/08/2021 02:54:20",
      "content": "<p>Nice solution！Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143869,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "01/08/2021 05:37:52",
      "content": "<p>Congrats, simple and efficient model👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143971,
      "author_name": "fomdata",
      "author_url": "",
      "post_date": "01/08/2021 06:49:54",
      "content": "<p>Congrats, very nice approach, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144030,
      "author_name": "hardmilkcake",
      "author_url": "",
      "post_date": "01/08/2021 07:34:04",
      "content": "<p>Congrats, thank you for sharing, the U-GRU part is very interesting. I see that the dimensions of your embedding are different, how do they blend together?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144262,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/08/2021 10:58:31",
          "content": "<p>I concatenate them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144090,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "01/08/2021 08:20:43",
      "content": "<p>Very interesting thanks !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144156,
      "author_name": "darraghdog",
      "author_url": "",
      "post_date": "01/08/2021 09:12:45",
      "content": "<p>Well done Ahmet, I like the idea of tracking if all <code>A's</code> are chosen - like run length encoding. <br>\nWhat hidden dimension did you use in the GRU to fit 8192 in batch ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1144260,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/08/2021 10:58:02",
          "content": "<p>GRU hidden dimension was 128. Larger GRUs were overfitting. Model size is around 13 MB, which is mostly content embedding weights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144351,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:15:10",
      "content": "<p>Congratz, and thanks for the very nice write-up !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144577,
      "author_name": "levonian",
      "author_url": "",
      "post_date": "01/08/2021 14:43:10",
      "content": "<p>Very interesting, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146437,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "01/09/2021 19:18:04",
      "content": "<p>It's really terrific to get to such a high score without Transformer. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146663,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/09/2021 23:29:07",
      "content": "<p>Well done <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> !! Very innovative sol!! especially your training scheme with the U GRU and the CSS transformation.. are you planning to share any code (full or key parts) ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1153960,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/15/2021 09:33:38",
          "content": "<p>Thank you. Normally I share my code but this time it is in notebooks rather than well-organized scripts. I don't know if it will be readable.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1154270,
              "author_name": "imeintanis",
              "author_url": "",
              "post_date": "01/15/2021 14:29:45",
              "content": "<p>Thanks, maybe just the model/dataset definition would be helpful and nice learning task for anyone that is interested </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1156563,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "01/17/2021 08:55:16",
          "content": "<p>I hope my inference kernel can be enough for that, therefore I made it public now: <a href=\"https://www.kaggle.com/aerdem4/riiid-starter\" target=\"_blank\">https://www.kaggle.com/aerdem4/riiid-starter</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146954,
      "author_name": "jarupula",
      "author_url": "",
      "post_date": "01/10/2021 07:57:06",
      "content": "<p>Congratulations !!! 🔥🔥. I have few questions(May be not related to your solution) about getting started with competition as a beginner. </p>\n<ul>\n<li>What are the resources(other than hardware) you have referred during the competition.</li>\n<li>As I am new to the kaggle competitions… Do I need industry level knowledge about the kind of problems given in the competition.<br>\nWhat are some tips to beginners like me.  </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1153925,
      "author_name": "neilgibbons",
      "author_url": "",
      "post_date": "01/15/2021 08:51:58",
      "content": "<p>Brilliant, thanks for sharing and congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143534": "First of all congrats to @keetar and @mamasinkgs and other top teams. I will shortly explain my solution. It is late here, so I may be missing some points.\n\n**Some General Details**\n\n- Used questions only. Lectures improve validation score but increase train-val gap and don’t improve LB.\n- Didn't use dataframes. Used numpy arrays partitioned by user ids.\n- Sequence length 256\n- 15 epochs. 1/3 of the data each time with reducing LR.\n- 8192 batch size\n- Ensemble of the same model with 7 different seeds trained on whole data\n- 0.8136 single model validation score, 0.813 LB. Ensemble: 0.815.\n- 8 hours training on 4-GPU machine\n- Used Github and committed any improvement with a message like: Add one more GRU layer (Val: 0.8136, LB: 0.813)\n\n**Inputs**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F4e161bd44dd25c8329c5ced0fa6c9ab3%2Fimg1.png?generation=1610064408784528&alt=media)\n\n**Engineered Features**\n\nAssume current question’s correct answer is X. Logarithm of:\n- Number of questions since last X.\n- Length of current X streak. (can be zero)\n- Length of current streak on any non-X answer. (can be zero)\n\nThis helps with users who always pick A as answer etc.\n\n**Embeddings**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F8256c5fe356ce1bedeedaf9bc91e9884%2Fimg2.png?generation=1610064559098164&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471945%2F73d8b182375f185231d35f8b141b8cbf%2Fimg3.png?generation=1610064576810743&alt=media)\n\nContent Cosine Similarity\n- A bit similar to attention with 16 heads\n- Linear transformation and l2 norm applied on content vectors\n- For 16 different transformation, cosine similarity between current content and history contents are calculated.\n- Transformation is symmetric for content and history contents.\n\nU-GRU:\n- GRU with 2 directions but not BiGRU\n- First does reverse pass, concatenates the output and then does forward pass\n\nMLP:\n- 2 layers of [Linear, BatchNorm, Relu]\n\n\nEdit: Part embedding is trainable. There is actually sigmoid x tanh layer before GRUs.",
    "1143538": "Nice solution!",
    "1143540": "Wow. Extremelly innovative. Congratulations.",
    "1143546": "Thanks @antorsae Congrats to you and your teammate with the 3rd place!",
    "1143550": "Impressive solution. Thank you for sharing.",
    "1143553": "Great solution, congrats on 6th place @aerdem4!",
    "1143569": "Brilliant approach. Congratulations!",
    "1143600": "Suspect for special award ^^",
    "1143712": "Nice solution！Congratulations!",
    "1143869": "Congrats, simple and efficient model👍",
    "1143971": "Congrats, very nice approach, thanks for sharing.",
    "1144030": "Congrats, thank you for sharing, the U-GRU part is very interesting. I see that the dimensions of your embedding are different, how do they blend together?",
    "1144090": "Very interesting thanks !",
    "1144156": "Well done Ahmet, I like the idea of tracking if all `A's` are chosen - like run length encoding. \nWhat hidden dimension did you use in the GRU to fit 8192 in batch ?",
    "1144260": "GRU hidden dimension was 128. Larger GRUs were overfitting. Model size is around 13 MB, which is mostly content embedding weights.",
    "1144262": "I concatenate them.",
    "1144351": "Congratz, and thanks for the very nice write-up !",
    "1144577": "Very interesting, thanks for sharing!",
    "1146437": "It's really terrific to get to such a high score without Transformer. Congrats!",
    "1146663": "Well done @aerdem4 !! Very innovative sol!! especially your training scheme with the U GRU and the CSS transformation.. are you planning to share any code (full or key parts) ?",
    "1146954": "Congratulations !!! 🔥🔥. I have few questions(May be not related to your solution) about getting started with competition as a beginner. \n- What are the resources(other than hardware) you have referred during the competition.\n- As I am new to the kaggle competitions... Do I need industry level knowledge about the kind of problems given in the competition.\nWhat are some tips to beginners like me.",
    "1153925": "Brilliant, thanks for sharing and congrats!",
    "1153960": "Thank you. Normally I share my code but this time it is in notebooks rather than well-organized scripts. I don't know if it will be readable.",
    "1154270": "Thanks, maybe just the model/dataset definition would be helpful and nice learning task for anyone that is interested",
    "1156563": "I hope my inference kernel can be enough for that, therefore I made it public now: https://www.kaggle.com/aerdem4/riiid-starter"
  },
  "source": "meta"
}