{
  "id": 209587,
  "title": "20th solution, transformer encoder only model (SERT) with code and notebook",
  "url": "/competitions/riiid-test-answer-prediction/writeups/sert-20th-solution-transformer-encoder-only-model-",
  "author_name": "",
  "post_date": "2021-01-09T03:00:10.603Z",
  "votes": 48,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting a comp with very little shakeup. This is always nice to see. I wanna thank <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> and <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> for the starter kernels. Without those I would not have known what to do in this difficult comp.</p>\n<h1>Architecture</h1>\n<p>I use the transformer encoder only SERT (SIngle-directional Encoder Representation from Transformers), to make predictions just one linear layer after the last encoder layer. This is probably a mistake</p>\n<h1>Embeddings</h1>\n<ol>\n<li>Question_id</li>\n<li>Prior question correctness</li>\n<li>Timestamp difference between bundles</li>\n<li>Prior question elapsed time</li>\n<li>Prior question explanation</li>\n<li>Tag cluster thanks to <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> </li>\n<li>Tag vector</li>\n<li>Fixed pos encoding, same as in Attention is All You Need</li>\n</ol>\n<h1>Key modifications</h1>\n<ol>\n<li>encoder only, this is kind of stupid and a mistake. I think this cost me a few places</li>\n<li>fixed pos encoding, which allows retraining the model with longer sequences</li>\n<li>layer norm and dropout after embedding layers</li>\n<li>Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss</li>\n</ol>\n<h1>Mistakes I couldn't resolve</h1>\n<ol>\n<li>Intra-bundle leakage, I made a nice task mask implementation, but it only made my score worse, probably because of the lack of a decoder. I ended using just an autoregressive mask</li>\n<li>First few positions during inference and training may have wrong timestamp difference since i simply do t[1:]-t[:-1]</li>\n</ol>\n<h1>Code and submission kernel</h1>\n<p><a href=\"https://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution\" target=\"_blank\">https://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution</a></p>\n<p><a href=\"https://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583\" target=\"_blank\">https://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583</a></p>\n<h2>Feel free to ask questions. This is a very brief write-up</h2>",
  "messages": [
    {
      "id": "1143562",
      "postDate": "01/08/2021 00:27:39",
      "content": "<p>Thanks to the organizers for hosting a comp with very little shakeup. This is always nice to see. I wanna thank <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> and <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> for the starter kernels. Without those I would not have known what to do in this difficult comp.</p>\n<h1>Architecture</h1>\n<p>I use the transformer encoder only SERT (SIngle-directional Encoder Representation from Transformers), to make predictions just one linear layer after the last encoder layer. This is probably a mistake</p>\n<h1>Embeddings</h1>\n<ol>\n<li>Question_id</li>\n<li>Prior question correctness</li>\n<li>Timestamp difference between bundles</li>\n<li>Prior question elapsed time</li>\n<li>Prior question explanation</li>\n<li>Tag cluster thanks to <a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> </li>\n<li>Tag vector</li>\n<li>Fixed pos encoding, same as in Attention is All You Need</li>\n</ol>\n<h1>Key modifications</h1>\n<ol>\n<li>encoder only, this is kind of stupid and a mistake. I think this cost me a few places</li>\n<li>fixed pos encoding, which allows retraining the model with longer sequences</li>\n<li>layer norm and dropout after embedding layers</li>\n<li>Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss</li>\n</ol>\n<h1>Mistakes I couldn't resolve</h1>\n<ol>\n<li>Intra-bundle leakage, I made a nice task mask implementation, but it only made my score worse, probably because of the lack of a decoder. I ended using just an autoregressive mask</li>\n<li>First few positions during inference and training may have wrong timestamp difference since i simply do t[1:]-t[:-1]</li>\n</ol>\n<h1>Code and submission kernel</h1>\n<p><a href=\"https://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution\" target=\"_blank\">https://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution</a></p>\n<p><a href=\"https://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583\" target=\"_blank\">https://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583</a></p>\n<h2>Feel free to ask questions. This is a very brief write-up</h2>",
      "rawMarkdown": "Thanks to the organizers for hosting a comp with very little shakeup. This is always nice to see. I wanna thank @wangsg and @leadbest for the starter kernels. Without those I would not have known what to do in this difficult comp.\n\n\n# Architecture\n\nI use the transformer encoder only SERT (SIngle-directional Encoder Representation from Transformers), to make predictions just one linear layer after the last encoder layer. This is probably a mistake\n\n# Embeddings\n\n1. Question_id\n2. Prior question correctness\n3. Timestamp difference between bundles\n4. Prior question elapsed time\n5. Prior question explanation\n6. Tag cluster thanks to @spacelx \n7.  Tag vector\n8. Fixed pos encoding, same as in Attention is All You Need\n\n# Key modifications\n1. encoder only, this is kind of stupid and a mistake. I think this cost me a few places\n2. fixed pos encoding, which allows retraining the model with longer sequences\n3. layer norm and dropout after embedding layers\n4. Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss\n\n# Mistakes I couldn't resolve\n1.\tIntra-bundle leakage, I made a nice task mask implementation, but it only made my score worse, probably because of the lack of a decoder. I ended using just an autoregressive mask\n2.\tFirst few positions during inference and training may have wrong timestamp difference since i simply do t[1:]-t[:-1]\n\n# Code and submission kernel\nhttps://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution\n\nhttps://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583\n\n## Feel free to ask questions. This is a very brief write-up",
      "votes": null
    },
    {
      "id": "1143568",
      "postDate": "01/08/2021 00:33:05",
      "content": "<p>Strongly model. Congrats and thanks for sharing solution and code <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>",
      "rawMarkdown": "Strongly model. Congrats and thanks for sharing solution and code @shujun717",
      "votes": null
    },
    {
      "id": "1143576",
      "postDate": "01/08/2021 00:37:42",
      "content": "<p>Thanks gold was within reach but alas, mistakes were made</p>",
      "rawMarkdown": "Thanks gold was within reach but alas, mistakes were made",
      "votes": null
    },
    {
      "id": "1143584",
      "postDate": "01/08/2021 00:44:11",
      "content": "<p>Wow! This is cool work! Can you share more details about that tag vector of yours?</p>",
      "rawMarkdown": "Wow! This is cool work! Can you share more details about that tag vector of yours?",
      "votes": null
    },
    {
      "id": "1143586",
      "postDate": "01/08/2021 00:47:19",
      "content": "<p>Yes, so there are 188 tags in total, so for each question I use a tag vector of size 188 of 0's and 1's. For instance, it a question has tags 1 and 140, the 1th and 140th element would be 1's and the rest would be 0's, and then a linear layer to transform it to embedding dim</p>",
      "rawMarkdown": "Yes, so there are 188 tags in total, so for each question I use a tag vector of size 188 of 0's and 1's. For instance, it a question has tags 1 and 140, the 1th and 140th element would be 1's and the rest would be 0's, and then a linear layer to transform it to embedding dim",
      "votes": null
    },
    {
      "id": "1143609",
      "postDate": "01/08/2021 01:07:11",
      "content": "<p>Excellent work <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> !! Thank you for the write up and the code.. I will study it more carefully in the weekend</p>\n<p>Edit: it would be nice to add the milestones wrt to LB scores, e.g. starting from the public nb: LB 07xx, add lag info: LB 078xx, …, up to 0.810 - if you have and you like to share such info</p>",
      "rawMarkdown": "Excellent work @shujun717 !! Thank you for the write up and the code.. I will study it more carefully in the weekend\n\nEdit: it would be nice to add the milestones wrt to LB scores, e.g. starting from the public nb: LB 07xx, add lag info: LB 078xx, ..., up to 0.810 - if you have and you like to share such info",
      "votes": null
    },
    {
      "id": "1143662",
      "postDate": "01/08/2021 01:50:40",
      "content": "<p>Thanks for your sharing.<br>\nBy the way, why did you use encoder only? </p>",
      "rawMarkdown": "Thanks for your sharing.\nBy the way, why did you use encoder only?",
      "votes": null
    },
    {
      "id": "1143674",
      "postDate": "01/08/2021 02:01:46",
      "content": "<p>Thx and nice suggestion. I will try to add that information</p>",
      "rawMarkdown": "Thx and nice suggestion. I will try to add that information",
      "votes": null
    },
    {
      "id": "1143676",
      "postDate": "01/08/2021 02:03:53",
      "content": "<p>I always prefer less complexity. To me I though the decoder was not necessary, so I got rid of it. Most of my changes was also just to reduce complexity (e.g. learnable pos encoding -&gt; fixed)</p>",
      "rawMarkdown": "I always prefer less complexity. To me I though the decoder was not necessary, so I got rid of it. Most of my changes was also just to reduce complexity (e.g. learnable pos encoding -> fixed)",
      "votes": null
    },
    {
      "id": "1143835",
      "postDate": "01/08/2021 05:00:51",
      "content": "<p>Great ….. Thank you for sharing 💯</p>",
      "rawMarkdown": "Great ..... Thank you for sharing 💯",
      "votes": null
    },
    {
      "id": "1143878",
      "postDate": "01/08/2021 05:43:24",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> , nice solution!</p>",
      "rawMarkdown": "Congrats @shujun717 , nice solution!",
      "votes": null
    },
    {
      "id": "1146522",
      "postDate": "01/09/2021 20:26:38",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> congrats and thanks for sharing.</p>\n<p><code>Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss</code></p>\n<p>Why you are giving more weight to loss for later position ? Are they more difficult to predict for your model and how did you figure out this ?</p>\n<p>Also, in your code, you have <code>question_cmnts.csv</code>. Can you tell a bit about it or can you provide some reference?</p>",
      "rawMarkdown": "shujun717 congrats and thanks for sharing.\n\n`Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss`\n\nWhy you are giving more weight to loss for later position ? Are they more difficult to predict for your model and how did you figure out this ?\n\nAlso, in your code, you have `question_cmnts.csv`. Can you tell a bit about it or can you provide some reference?",
      "votes": null
    },
    {
      "id": "1146667",
      "postDate": "01/09/2021 23:36:57",
      "content": "<p>That's a good question  <a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> . Because of the autoregressive mask, the items in the sequence can only see items that come before. For instance, the 1st question can only see itself, and the 2nd can only see itself and the 1st. Therefore, the early questions have less context compared to the later questions. If I let all positions have equal weight, the model will overfit to earlier positions with little context. Another consideration is that during inference, I'm always using the last position to make predictions, so it doesnt make sense to give equal weighting to everything. Using this loss weight gave me 0.002 boost.</p>\n<p>As for question_cmnts.csv, here's the notebook <a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags</a></p>",
      "rawMarkdown": "That's a good question  @abdurrehman245 . Because of the autoregressive mask, the items in the sequence can only see items that come before. For instance, the 1st question can only see itself, and the 2nd can only see itself and the 1st. Therefore, the early questions have less context compared to the later questions. If I let all positions have equal weight, the model will overfit to earlier positions with little context. Another consideration is that during inference, I'm always using the last position to make predictions, so it doesnt make sense to give equal weighting to everything. Using this loss weight gave me 0.002 boost.\n\nAs for question_cmnts.csv, here's the notebook https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags",
      "votes": null
    },
    {
      "id": "1154019",
      "postDate": "01/15/2021 10:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> thanks for sharing, I enjoyed reading this and it was super helpful to look through your code</p>",
      "rawMarkdown": "shujun717 thanks for sharing, I enjoyed reading this and it was super helpful to look through your code",
      "votes": null
    },
    {
      "id": "1157612",
      "postDate": "01/18/2021 03:47:05",
      "content": "<p><a href=\"https://www.kaggle.com/neilgibbons\" target=\"_blank\">@neilgibbons</a> no problem! Glad it was of help.</p>",
      "rawMarkdown": "neilgibbons no problem! Glad it was of help.",
      "votes": null
    },
    {
      "id": "1176989",
      "postDate": "01/30/2021 00:14:59",
      "content": "<p>This is great. What hardware was it run on?</p>",
      "rawMarkdown": "This is great. What hardware was it run on?",
      "votes": null
    },
    {
      "id": "1177239",
      "postDate": "01/30/2021 06:36:04",
      "content": "<p>Thanks. Mostly a 2x3090 machine with a 24 core threadripper</p>",
      "rawMarkdown": "Thanks. Mostly a 2x3090 machine with a 24 core threadripper",
      "votes": null
    },
    {
      "id": "1179906",
      "postDate": "02/01/2021 00:45:49",
      "content": "<p>Thank you for sharing and congratulations!<br>\nI'm learning a lot from your code at github :)</p>",
      "rawMarkdown": "Thank you for sharing and congratulations!\nI'm learning a lot from your code at github :)",
      "votes": null
    },
    {
      "id": "1181326",
      "postDate": "02/01/2021 19:34:17",
      "content": "<p>You're welcome! Glad it is useful </p>",
      "rawMarkdown": "You're welcome! Glad it is useful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143568,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 00:33:05",
      "content": "<p>Strongly model. Congrats and thanks for sharing solution and code <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1143576,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/08/2021 00:37:42",
          "content": "<p>Thanks gold was within reach but alas, mistakes were made</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143584,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/08/2021 00:44:11",
      "content": "<p>Wow! This is cool work! Can you share more details about that tag vector of yours?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143586,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/08/2021 00:47:19",
          "content": "<p>Yes, so there are 188 tags in total, so for each question I use a tag vector of size 188 of 0's and 1's. For instance, it a question has tags 1 and 140, the 1th and 140th element would be 1's and the rest would be 0's, and then a linear layer to transform it to embedding dim</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143609,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/08/2021 01:07:11",
      "content": "<p>Excellent work <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> !! Thank you for the write up and the code.. I will study it more carefully in the weekend</p>\n<p>Edit: it would be nice to add the milestones wrt to LB scores, e.g. starting from the public nb: LB 07xx, add lag info: LB 078xx, …, up to 0.810 - if you have and you like to share such info</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143674,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/08/2021 02:01:46",
          "content": "<p>Thx and nice suggestion. I will try to add that information</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143662,
      "author_name": "m10515009",
      "author_url": "",
      "post_date": "01/08/2021 01:50:40",
      "content": "<p>Thanks for your sharing.<br>\nBy the way, why did you use encoder only? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1143676,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/08/2021 02:03:53",
          "content": "<p>I always prefer less complexity. To me I though the decoder was not necessary, so I got rid of it. Most of my changes was also just to reduce complexity (e.g. learnable pos encoding -&gt; fixed)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143835,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "01/08/2021 05:00:51",
      "content": "<p>Great ….. Thank you for sharing 💯</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143878,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "01/08/2021 05:43:24",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> , nice solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146522,
      "author_name": "abdurrehman245",
      "author_url": "",
      "post_date": "01/09/2021 20:26:38",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> congrats and thanks for sharing.</p>\n<p><code>Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss</code></p>\n<p>Why you are giving more weight to loss for later position ? Are they more difficult to predict for your model and how did you figure out this ?</p>\n<p>Also, in your code, you have <code>question_cmnts.csv</code>. Can you tell a bit about it or can you provide some reference?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146667,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/09/2021 23:36:57",
          "content": "<p>That's a good question  <a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> . Because of the autoregressive mask, the items in the sequence can only see items that come before. For instance, the 1st question can only see itself, and the 2nd can only see itself and the 1st. Therefore, the early questions have less context compared to the later questions. If I let all positions have equal weight, the model will overfit to earlier positions with little context. Another consideration is that during inference, I'm always using the last position to make predictions, so it doesnt make sense to give equal weighting to everything. Using this loss weight gave me 0.002 boost.</p>\n<p>As for question_cmnts.csv, here's the notebook <a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1154019,
      "author_name": "neilgibbons",
      "author_url": "",
      "post_date": "01/15/2021 10:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> thanks for sharing, I enjoyed reading this and it was super helpful to look through your code</p>",
      "votes": null,
      "replies": [
        {
          "id": 1157612,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/18/2021 03:47:05",
          "content": "<p><a href=\"https://www.kaggle.com/neilgibbons\" target=\"_blank\">@neilgibbons</a> no problem! Glad it was of help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1176989,
      "author_name": "philipkd",
      "author_url": "",
      "post_date": "01/30/2021 00:14:59",
      "content": "<p>This is great. What hardware was it run on?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1177239,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/30/2021 06:36:04",
          "content": "<p>Thanks. Mostly a 2x3090 machine with a 24 core threadripper</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1179906,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "02/01/2021 00:45:49",
      "content": "<p>Thank you for sharing and congratulations!<br>\nI'm learning a lot from your code at github :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1181326,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "02/01/2021 19:34:17",
          "content": "<p>You're welcome! Glad it is useful </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143562": "Thanks to the organizers for hosting a comp with very little shakeup. This is always nice to see. I wanna thank @wangsg and @leadbest for the starter kernels. Without those I would not have known what to do in this difficult comp.\n\n\n# Architecture\n\nI use the transformer encoder only SERT (SIngle-directional Encoder Representation from Transformers), to make predictions just one linear layer after the last encoder layer. This is probably a mistake\n\n# Embeddings\n\n1. Question_id\n2. Prior question correctness\n3. Timestamp difference between bundles\n4. Prior question elapsed time\n5. Prior question explanation\n6. Tag cluster thanks to @spacelx \n7.  Tag vector\n8. Fixed pos encoding, same as in Attention is All You Need\n\n# Key modifications\n1. encoder only, this is kind of stupid and a mistake. I think this cost me a few places\n2. fixed pos encoding, which allows retraining the model with longer sequences\n3. layer norm and dropout after embedding layers\n4. Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss\n\n# Mistakes I couldn't resolve\n1.\tIntra-bundle leakage, I made a nice task mask implementation, but it only made my score worse, probably because of the lack of a decoder. I ended using just an autoregressive mask\n2.\tFirst few positions during inference and training may have wrong timestamp difference since i simply do t[1:]-t[:-1]\n\n# Code and submission kernel\nhttps://github.com/Shujun-He/Riiid-Answer-Correctness-Prediction-20th-solution\n\nhttps://www.kaggle.com/shujun717/fork-of-tag-encoding-with-loss-weight?scriptVersionId=51272583\n\n## Feel free to ask questions. This is a very brief write-up",
    "1143568": "Strongly model. Congrats and thanks for sharing solution and code @shujun717",
    "1143576": "Thanks gold was within reach but alas, mistakes were made",
    "1143584": "Wow! This is cool work! Can you share more details about that tag vector of yours?",
    "1143586": "Yes, so there are 188 tags in total, so for each question I use a tag vector of size 188 of 0's and 1's. For instance, it a question has tags 1 and 140, the 1th and 140th element would be 1's and the rest would be 0's, and then a linear layer to transform it to embedding dim",
    "1143609": "Excellent work @shujun717 !! Thank you for the write up and the code.. I will study it more carefully in the weekend\n\nEdit: it would be nice to add the milestones wrt to LB scores, e.g. starting from the public nb: LB 07xx, add lag info: LB 078xx, ..., up to 0.810 - if you have and you like to share such info",
    "1143662": "Thanks for your sharing.\nBy the way, why did you use encoder only?",
    "1143674": "Thx and nice suggestion. I will try to add that information",
    "1143676": "I always prefer less complexity. To me I though the decoder was not necessary, so I got rid of it. Most of my changes was also just to reduce complexity (e.g. learnable pos encoding -> fixed)",
    "1143835": "Great ..... Thank you for sharing 💯",
    "1143878": "Congrats @shujun717 , nice solution!",
    "1146522": "shujun717 congrats and thanks for sharing.\n\n`Loss weight favoring later positions, np.arange(0,1,1/seq_length)*loss`\n\nWhy you are giving more weight to loss for later position ? Are they more difficult to predict for your model and how did you figure out this ?\n\nAlso, in your code, you have `question_cmnts.csv`. Can you tell a bit about it or can you provide some reference?",
    "1146667": "That's a good question  @abdurrehman245 . Because of the autoregressive mask, the items in the sequence can only see items that come before. For instance, the 1st question can only see itself, and the 2nd can only see itself and the 1st. Therefore, the early questions have less context compared to the later questions. If I let all positions have equal weight, the model will overfit to earlier positions with little context. Another consideration is that during inference, I'm always using the last position to make predictions, so it doesnt make sense to give equal weighting to everything. Using this loss weight gave me 0.002 boost.\n\nAs for question_cmnts.csv, here's the notebook https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags",
    "1154019": "shujun717 thanks for sharing, I enjoyed reading this and it was super helpful to look through your code",
    "1157612": "neilgibbons no problem! Glad it was of help.",
    "1176989": "This is great. What hardware was it run on?",
    "1177239": "Thanks. Mostly a 2x3090 machine with a 24 core threadripper",
    "1179906": "Thank you for sharing and congratulations!\nI'm learning a lot from your code at github :)",
    "1181326": "You're welcome! Glad it is useful"
  },
  "source": "meta"
}