{
  "id": 209615,
  "title": "SAINT+ Model with code [Private 0.800 / 68th place]",
  "url": "/competitions/riiid-test-answer-prediction/writeups/marisaka-shuntaro-saint-model-with-code-private-0-",
  "author_name": "",
  "post_date": "2021-01-08T03:10:49.336843600Z",
  "votes": 24,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I'd like to share our solution. All codes are available on <a href=\"https://github.com/marisakamozz/riiid\" target=\"_blank\">GitHub</a>. You can reproduce our submission with <a href=\"https://www.kaggle.com/marisakamozz/riiid-saint-solution\" target=\"_blank\">this notebook</a>.</p>\n<p>I hope it will be helpful to you. Thank you.</p>\n<h1>Model Summary</h1>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a> based model</li>\n<li>used train.csv only.</li>\n<li>used questions only.</li>\n</ul>\n<p>We didn't use lectures, because lectures didn't improve validation score.</p>\n<h2>Input</h2>\n<p>Encoder input:</p>\n<ul>\n<li>positional embedding</li>\n<li>question_id embedding</li>\n<li>lag embedding (= timestamp - last timestamp)</li>\n<li>user_count: number of questions the user has solved</li>\n</ul>\n<p>lag feature is categorized by 100 quantiles.<br>\nuser_count is categorized by <a href=\"https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test\" target=\"_blank\">proficiency test</a>.</p>\n<p>Decoder input:</p>\n<ul>\n<li>positional embedding</li>\n<li>answered_correctly</li>\n<li>prior_question_elapsed_time</li>\n<li>prior_question_had_explanation</li>\n</ul>\n<p>prior_question_elapsed_time and prior_question_had_explanation are shifted -1.<br>\nprior_question_elapsed_time is categorized by 100 quantiles.</p>\n<h2>Model Architecture</h2>\n<p>Embedding Layer + Tranformer + Feed Forward Layer (with residual connection)</p>\n<p>Weights of embedding layers are sampled from a normal distribution with a standard deviation of 0.01.</p>\n<p>Transformer architecture:</p>\n<ul>\n<li>sequence length: 150</li>\n<li>number of dimensions: 128</li>\n<li>number of encoder layers: 3</li>\n<li>number of decoder layers: 3 (self-attention) + 3 (encoder-attention)</li>\n<li>number of heads: 8</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>batch size: 256</li>\n<li>optimizer: Adam</li>\n<li>training time is almost 20 hours on 1 K80 GPU machine.</li>\n<li>early stopping by 1% unseen user's auc score. (but it didn't work well.)</li>\n</ul>",
  "messages": [
    {
      "id": "1143730",
      "postDate": "01/08/2021 03:10:49",
      "content": "<p>I'd like to share our solution. All codes are available on <a href=\"https://github.com/marisakamozz/riiid\" target=\"_blank\">GitHub</a>. You can reproduce our submission with <a href=\"https://www.kaggle.com/marisakamozz/riiid-saint-solution\" target=\"_blank\">this notebook</a>.</p>\n<p>I hope it will be helpful to you. Thank you.</p>\n<h1>Model Summary</h1>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a> based model</li>\n<li>used train.csv only.</li>\n<li>used questions only.</li>\n</ul>\n<p>We didn't use lectures, because lectures didn't improve validation score.</p>\n<h2>Input</h2>\n<p>Encoder input:</p>\n<ul>\n<li>positional embedding</li>\n<li>question_id embedding</li>\n<li>lag embedding (= timestamp - last timestamp)</li>\n<li>user_count: number of questions the user has solved</li>\n</ul>\n<p>lag feature is categorized by 100 quantiles.<br>\nuser_count is categorized by <a href=\"https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test\" target=\"_blank\">proficiency test</a>.</p>\n<p>Decoder input:</p>\n<ul>\n<li>positional embedding</li>\n<li>answered_correctly</li>\n<li>prior_question_elapsed_time</li>\n<li>prior_question_had_explanation</li>\n</ul>\n<p>prior_question_elapsed_time and prior_question_had_explanation are shifted -1.<br>\nprior_question_elapsed_time is categorized by 100 quantiles.</p>\n<h2>Model Architecture</h2>\n<p>Embedding Layer + Tranformer + Feed Forward Layer (with residual connection)</p>\n<p>Weights of embedding layers are sampled from a normal distribution with a standard deviation of 0.01.</p>\n<p>Transformer architecture:</p>\n<ul>\n<li>sequence length: 150</li>\n<li>number of dimensions: 128</li>\n<li>number of encoder layers: 3</li>\n<li>number of decoder layers: 3 (self-attention) + 3 (encoder-attention)</li>\n<li>number of heads: 8</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>batch size: 256</li>\n<li>optimizer: Adam</li>\n<li>training time is almost 20 hours on 1 K80 GPU machine.</li>\n<li>early stopping by 1% unseen user's auc score. (but it didn't work well.)</li>\n</ul>",
      "rawMarkdown": "I'd like to share our solution. All codes are available on [GitHub](https://github.com/marisakamozz/riiid). You can reproduce our submission with [this notebook](https://www.kaggle.com/marisakamozz/riiid-saint-solution).\n\nI hope it will be helpful to you. Thank you.\n\n# Model Summary\n\n* [SAINT+](https://arxiv.org/abs/2010.12042) based model\n* used train.csv only.\n* used questions only.\n\nWe didn't use lectures, because lectures didn't improve validation score.\n\n## Input\n\nEncoder input:\n\n* positional embedding\n* question_id embedding\n* lag embedding (= timestamp - last timestamp)\n* user_count: number of questions the user has solved\n\nlag feature is categorized by 100 quantiles.\nuser_count is categorized by [proficiency test](https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test).\n\nDecoder input:\n\n* positional embedding\n* answered_correctly\n* prior_question_elapsed_time\n* prior_question_had_explanation\n\nprior_question_elapsed_time and prior_question_had_explanation are shifted -1.\nprior_question_elapsed_time is categorized by 100 quantiles.\n\n## Model Architecture\n\nEmbedding Layer + Tranformer + Feed Forward Layer (with residual connection)\n\nWeights of embedding layers are sampled from a normal distribution with a standard deviation of 0.01.\n\nTransformer architecture:\n\n* sequence length: 150\n* number of dimensions: 128\n* number of encoder layers: 3\n* number of decoder layers: 3 (self-attention) + 3 (encoder-attention)\n* number of heads: 8\n\n## Training\n\n* batch size: 256\n* optimizer: Adam\n* training time is almost 20 hours on 1 K80 GPU machine.\n* early stopping by 1% unseen user's auc score. (but it didn't work well.)",
      "votes": null
    },
    {
      "id": "1143746",
      "postDate": "01/08/2021 03:29:20",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a>! 🎉</p>",
      "rawMarkdown": "Congratulations @marisakamozz! 🎉",
      "votes": null
    },
    {
      "id": "1143777",
      "postDate": "01/08/2021 04:04:14",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>!</p>",
      "rawMarkdown": "Thank you, @doctorkael!",
      "votes": null
    },
    {
      "id": "1143857",
      "postDate": "01/08/2021 05:20:20",
      "content": "<p>Thanks for your sharing.</p>",
      "rawMarkdown": "Thanks for your sharing.",
      "votes": null
    },
    {
      "id": "1143886",
      "postDate": "01/08/2021 05:46:54",
      "content": "<p>Thanks for sharing ! Can you help me with understanding. Is the input ‘answered_correctly’ to decoder shifted left by 1 (i.e. lag)? Because if not, won’t it be target leak?</p>",
      "rawMarkdown": "Thanks for sharing ! Can you help me with understanding. Is the input ‘answered_correctly’ to decoder shifted left by 1 (i.e. lag)? Because if not, won’t it be target leak?",
      "votes": null
    },
    {
      "id": "1143928",
      "postDate": "01/08/2021 06:20:49",
      "content": "<p>Thank you for your response.</p>\n<p>I omitted the explanation, the input of the decoder is shifted by +1 in SAINT+ to avoid target leak.</p>\n<p>prior_question_elapsed_time and prior_question_had_explanation are shifted -1. As a result, it becomes 0. Because prior_question_elapsed_time and prior_question_had_explanation are information about the user's response to the \"prior\" question.</p>",
      "rawMarkdown": "Thank you for your response.\n\nI omitted the explanation, the input of the decoder is shifted by +1 in SAINT+ to avoid target leak.\n\nprior_question_elapsed_time and prior_question_had_explanation are shifted -1. As a result, it becomes 0. Because prior_question_elapsed_time and prior_question_had_explanation are information about the user's response to the \"prior\" question.",
      "votes": null
    },
    {
      "id": "1144243",
      "postDate": "01/08/2021 10:44:38",
      "content": "<p>Congrats on winning medal.</p>\n<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> In SAINT+ paper, they used <code>lag_time</code> embedding in decoder instead of encoder and also they used <code>category</code> embedding in encoder. Did you guys experiment with these embeddings and what were the results ?</p>",
      "rawMarkdown": "Congrats on winning medal.\n\n@marisakamozz In SAINT+ paper, they used `lag_time ` embedding in decoder instead of encoder and also they used `category ` embedding in encoder. Did you guys experiment with these embeddings and what were the results ?",
      "votes": null
    },
    {
      "id": "1144327",
      "postDate": "01/08/2021 11:53:37",
      "content": "<p>Thanks again and congratulations !! </p>",
      "rawMarkdown": "Thanks again and congratulations !!",
      "votes": null
    },
    {
      "id": "1144380",
      "postDate": "01/08/2021 12:34:31",
      "content": "<p>Thank you for your response.</p>\n<p>I tried <code>lag_time</code> in decoder, but validation score is lower than in encoder. Since <code>lag_time</code> is information that is known before answering the question, I think that it is better to enter it from the encoder side.</p>",
      "rawMarkdown": "Thank you for your response.\n\nI tried `lag_time` in decoder, but validation score is lower than in encoder. Since `lag_time` is information that is known before answering the question, I think that it is better to enter it from the encoder side.",
      "votes": null
    },
    {
      "id": "1144506",
      "postDate": "01/08/2021 13:51:13",
      "content": "<p>You are welcome.</p>",
      "rawMarkdown": "You are welcome.",
      "votes": null
    },
    {
      "id": "1145095",
      "postDate": "01/08/2021 22:05:46",
      "content": "<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> I could not understand how did you use <code>user_count</code> embedding. It would be nice if you can explain it a bit.</p>",
      "rawMarkdown": "marisakamozz I could not understand how did you use `user_count` embedding. It would be nice if you can explain it a bit.",
      "votes": null
    },
    {
      "id": "1145808",
      "postDate": "01/09/2021 10:59:29",
      "content": "<p><code>user_count</code> is categorized as follows.</p>\n<p>1 - 3 (part 1) : 1<br>\n4 (part 2) : 2<br>\n5 - 7 (part 3) : 3<br>\n8 - 10 (part 4) : 4<br>\n11 - 13 (part 4) : 5<br>\n14 - 16 (part 4) : 6<br>\n17 - 22 (part 5) : 7<br>\n23 - 26 (part 6) : 8<br>\n27 - 30 (part 7) : 9<br>\n30 - : 10</p>\n<p>We found that our model did not score very well for the first 30 questions. Therefore, We decided to add <code>user_count</code>. This increased the Public score by 0.03.</p>\n<p>This is one of important contributions of my partner, <a href=\"https://www.kaggle.com/shuntarotanaka\" target=\"_blank\">@shuntarotanaka</a> . Please see <a href=\"https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test\" target=\"_blank\">this notebook</a> for details.</p>",
      "rawMarkdown": "`user_count` is categorized as follows.\n\n1 - 3 (part 1) : 1\n4 (part 2) : 2\n5 - 7 (part 3) : 3\n8 - 10 (part 4) : 4\n11 - 13 (part 4) : 5\n14 - 16 (part 4) : 6\n17 - 22 (part 5) : 7\n23 - 26 (part 6) : 8\n27 - 30 (part 7) : 9\n30 - : 10\n\nWe found that our model did not score very well for the first 30 questions. Therefore, We decided to add `user_count`. This increased the Public score by 0.03.\n\nThis is one of important contributions of my partner, @shuntarotanaka . Please see [this notebook](https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test) for details.",
      "votes": null
    },
    {
      "id": "1145864",
      "postDate": "01/09/2021 11:39:35",
      "content": "<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> as you mentioned you used <code>lag_time</code> in decoder because that is known before answering the question so I think same goes for <code>prior_question_had_explanation</code> that is also know before answering the question but you fed it into the decoder instead of encoder. Any specific reason of that ?</p>",
      "rawMarkdown": "marisakamozz as you mentioned you used `lag_time` in decoder because that is known before answering the question so I think same goes for `prior_question_had_explanation` that is also know before answering the question but you fed it into the decoder instead of encoder. Any specific reason of that ?",
      "votes": null
    },
    {
      "id": "1145943",
      "postDate": "01/09/2021 12:44:48",
      "content": "<p><code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> are information about \"prior\" question. In other words, that information is additional information about the user's response. For example, even if the same correct answer is given, the meaning of the correct answer differs depending on whether the correct answer is taken over a long period of time or in a short time. Therefore, enter it in the decoder with <code>answered_correctly</code>. Please see <a href=\"https://www.kaggle.com/marisakamozz/prior-question-is-info-about-prior-question\" target=\"_blank\">this notebook</a> for details.</p>",
      "rawMarkdown": "`prior_question_elapsed_time` and `prior_question_had_explanation` are information about \"prior\" question. In other words, that information is additional information about the user's response. For example, even if the same correct answer is given, the meaning of the correct answer differs depending on whether the correct answer is taken over a long period of time or in a short time. Therefore, enter it in the decoder with `answered_correctly`. Please see [this notebook](https://www.kaggle.com/marisakamozz/prior-question-is-info-about-prior-question) for details.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143746,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "01/08/2021 03:29:20",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a>! 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143777,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/08/2021 04:04:14",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/doctorkael\" target=\"_blank\">@doctorkael</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143857,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "01/08/2021 05:20:20",
      "content": "<p>Thanks for your sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144506,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/08/2021 13:51:13",
          "content": "<p>You are welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143886,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "01/08/2021 05:46:54",
      "content": "<p>Thanks for sharing ! Can you help me with understanding. Is the input ‘answered_correctly’ to decoder shifted left by 1 (i.e. lag)? Because if not, won’t it be target leak?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143928,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/08/2021 06:20:49",
          "content": "<p>Thank you for your response.</p>\n<p>I omitted the explanation, the input of the decoder is shifted by +1 in SAINT+ to avoid target leak.</p>\n<p>prior_question_elapsed_time and prior_question_had_explanation are shifted -1. As a result, it becomes 0. Because prior_question_elapsed_time and prior_question_had_explanation are information about the user's response to the \"prior\" question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144327,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "01/08/2021 11:53:37",
          "content": "<p>Thanks again and congratulations !! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144243,
      "author_name": "abdurrehman245",
      "author_url": "",
      "post_date": "01/08/2021 10:44:38",
      "content": "<p>Congrats on winning medal.</p>\n<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> In SAINT+ paper, they used <code>lag_time</code> embedding in decoder instead of encoder and also they used <code>category</code> embedding in encoder. Did you guys experiment with these embeddings and what were the results ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144380,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/08/2021 12:34:31",
          "content": "<p>Thank you for your response.</p>\n<p>I tried <code>lag_time</code> in decoder, but validation score is lower than in encoder. Since <code>lag_time</code> is information that is known before answering the question, I think that it is better to enter it from the encoder side.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145095,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "01/08/2021 22:05:46",
          "content": "<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> I could not understand how did you use <code>user_count</code> embedding. It would be nice if you can explain it a bit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145808,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/09/2021 10:59:29",
          "content": "<p><code>user_count</code> is categorized as follows.</p>\n<p>1 - 3 (part 1) : 1<br>\n4 (part 2) : 2<br>\n5 - 7 (part 3) : 3<br>\n8 - 10 (part 4) : 4<br>\n11 - 13 (part 4) : 5<br>\n14 - 16 (part 4) : 6<br>\n17 - 22 (part 5) : 7<br>\n23 - 26 (part 6) : 8<br>\n27 - 30 (part 7) : 9<br>\n30 - : 10</p>\n<p>We found that our model did not score very well for the first 30 questions. Therefore, We decided to add <code>user_count</code>. This increased the Public score by 0.03.</p>\n<p>This is one of important contributions of my partner, <a href=\"https://www.kaggle.com/shuntarotanaka\" target=\"_blank\">@shuntarotanaka</a> . Please see <a href=\"https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test\" target=\"_blank\">this notebook</a> for details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145864,
          "author_name": "abdurrehman245",
          "author_url": "",
          "post_date": "01/09/2021 11:39:35",
          "content": "<p><a href=\"https://www.kaggle.com/marisakamozz\" target=\"_blank\">@marisakamozz</a> as you mentioned you used <code>lag_time</code> in decoder because that is known before answering the question so I think same goes for <code>prior_question_had_explanation</code> that is also know before answering the question but you fed it into the decoder instead of encoder. Any specific reason of that ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145943,
          "author_name": "marisakamozz",
          "author_url": "",
          "post_date": "01/09/2021 12:44:48",
          "content": "<p><code>prior_question_elapsed_time</code> and <code>prior_question_had_explanation</code> are information about \"prior\" question. In other words, that information is additional information about the user's response. For example, even if the same correct answer is given, the meaning of the correct answer differs depending on whether the correct answer is taken over a long period of time or in a short time. Therefore, enter it in the decoder with <code>answered_correctly</code>. Please see <a href=\"https://www.kaggle.com/marisakamozz/prior-question-is-info-about-prior-question\" target=\"_blank\">this notebook</a> for details.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143730": "I'd like to share our solution. All codes are available on [GitHub](https://github.com/marisakamozz/riiid). You can reproduce our submission with [this notebook](https://www.kaggle.com/marisakamozz/riiid-saint-solution).\n\nI hope it will be helpful to you. Thank you.\n\n# Model Summary\n\n* [SAINT+](https://arxiv.org/abs/2010.12042) based model\n* used train.csv only.\n* used questions only.\n\nWe didn't use lectures, because lectures didn't improve validation score.\n\n## Input\n\nEncoder input:\n\n* positional embedding\n* question_id embedding\n* lag embedding (= timestamp - last timestamp)\n* user_count: number of questions the user has solved\n\nlag feature is categorized by 100 quantiles.\nuser_count is categorized by [proficiency test](https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test).\n\nDecoder input:\n\n* positional embedding\n* answered_correctly\n* prior_question_elapsed_time\n* prior_question_had_explanation\n\nprior_question_elapsed_time and prior_question_had_explanation are shifted -1.\nprior_question_elapsed_time is categorized by 100 quantiles.\n\n## Model Architecture\n\nEmbedding Layer + Tranformer + Feed Forward Layer (with residual connection)\n\nWeights of embedding layers are sampled from a normal distribution with a standard deviation of 0.01.\n\nTransformer architecture:\n\n* sequence length: 150\n* number of dimensions: 128\n* number of encoder layers: 3\n* number of decoder layers: 3 (self-attention) + 3 (encoder-attention)\n* number of heads: 8\n\n## Training\n\n* batch size: 256\n* optimizer: Adam\n* training time is almost 20 hours on 1 K80 GPU machine.\n* early stopping by 1% unseen user's auc score. (but it didn't work well.)",
    "1143746": "Congratulations @marisakamozz! 🎉",
    "1143777": "Thank you, @doctorkael!",
    "1143857": "Thanks for your sharing.",
    "1143886": "Thanks for sharing ! Can you help me with understanding. Is the input ‘answered_correctly’ to decoder shifted left by 1 (i.e. lag)? Because if not, won’t it be target leak?",
    "1143928": "Thank you for your response.\n\nI omitted the explanation, the input of the decoder is shifted by +1 in SAINT+ to avoid target leak.\n\nprior_question_elapsed_time and prior_question_had_explanation are shifted -1. As a result, it becomes 0. Because prior_question_elapsed_time and prior_question_had_explanation are information about the user's response to the \"prior\" question.",
    "1144243": "Congrats on winning medal.\n\n@marisakamozz In SAINT+ paper, they used `lag_time ` embedding in decoder instead of encoder and also they used `category ` embedding in encoder. Did you guys experiment with these embeddings and what were the results ?",
    "1144327": "Thanks again and congratulations !!",
    "1144380": "Thank you for your response.\n\nI tried `lag_time` in decoder, but validation score is lower than in encoder. Since `lag_time` is information that is known before answering the question, I think that it is better to enter it from the encoder side.",
    "1144506": "You are welcome.",
    "1145095": "marisakamozz I could not understand how did you use `user_count` embedding. It would be nice if you can explain it a bit.",
    "1145808": "`user_count` is categorized as follows.\n\n1 - 3 (part 1) : 1\n4 (part 2) : 2\n5 - 7 (part 3) : 3\n8 - 10 (part 4) : 4\n11 - 13 (part 4) : 5\n14 - 16 (part 4) : 6\n17 - 22 (part 5) : 7\n23 - 26 (part 6) : 8\n27 - 30 (part 7) : 9\n30 - : 10\n\nWe found that our model did not score very well for the first 30 questions. Therefore, We decided to add `user_count`. This increased the Public score by 0.03.\n\nThis is one of important contributions of my partner, @shuntarotanaka . Please see [this notebook](https://www.kaggle.com/shuntarotanaka/riiid-the-first-30-questions-are-proficiency-test) for details.",
    "1145864": "marisakamozz as you mentioned you used `lag_time` in decoder because that is known before answering the question so I think same goes for `prior_question_had_explanation` that is also know before answering the question but you fed it into the decoder instead of encoder. Any specific reason of that ?",
    "1145943": "`prior_question_elapsed_time` and `prior_question_had_explanation` are information about \"prior\" question. In other words, that information is additional information about the user's response. For example, even if the same correct answer is given, the meaning of the correct answer differs depending on whether the correct answer is taken over a long period of time or in a short time. Therefore, enter it in the decoder with `answered_correctly`. Please see [this notebook](https://www.kaggle.com/marisakamozz/prior-question-is-info-about-prior-question) for details."
  },
  "source": "meta"
}