{
  "id": 209793,
  "title": "58th solution: SAINT+ based predicts user_answer and then answered_correctly",
  "url": "/competitions/riiid-test-answer-prediction/writeups/claverru-58th-solution-saint-based-predicts-user-a",
  "author_name": "",
  "post_date": "2021-01-10T12:00:11.657Z",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi folks!</p>\n<p>My solution has something different from what I've seen until now. I have two outputs in a serial fashion:</p>\n<ul>\n<li>First the user answer prediction (A, B, C, D).</li>\n<li>Then, inside the NN graph I use the correct answer id to gather the correctness probability.</li>\n</ul>\n<p>This way I have two different backwards signals that helped me to grow a little bit in LB.</p>\n<p>Other things to notice:</p>\n<ul>\n<li>Log-normalization applied to time features (they have a pareto-like distribution). At the end of the competition I realized I didn't <em>need</em> to clip these features with such normalization.</li>\n<li>All tags usage.</li>\n<li>Continuous task_container_id (adds user expertise information).</li>\n<li>I add a mask to make sure sequence elements in the same container don't attend to each other.</li>\n<li>Pad tokens and lectures are masked in the loss functions.</li>\n</ul>\n<p>The rule I used to choose whether a feature should go on encoder or decoder is: static information on encoder, variable information on decoder.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F71579dff863e98b9cdec6a2d0dfd48ee%2FUntitled%20Diagram%20(1).png?generation=1610226098358123&amp;alt=media\" alt=\"My solution\"></p>",
  "messages": [
    {
      "id": "1144665",
      "postDate": "01/08/2021 15:46:56",
      "content": "<p>Hi folks!</p>\n<p>My solution has something different from what I've seen until now. I have two outputs in a serial fashion:</p>\n<ul>\n<li>First the user answer prediction (A, B, C, D).</li>\n<li>Then, inside the NN graph I use the correct answer id to gather the correctness probability.</li>\n</ul>\n<p>This way I have two different backwards signals that helped me to grow a little bit in LB.</p>\n<p>Other things to notice:</p>\n<ul>\n<li>Log-normalization applied to time features (they have a pareto-like distribution). At the end of the competition I realized I didn't <em>need</em> to clip these features with such normalization.</li>\n<li>All tags usage.</li>\n<li>Continuous task_container_id (adds user expertise information).</li>\n<li>I add a mask to make sure sequence elements in the same container don't attend to each other.</li>\n<li>Pad tokens and lectures are masked in the loss functions.</li>\n</ul>\n<p>The rule I used to choose whether a feature should go on encoder or decoder is: static information on encoder, variable information on decoder.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F71579dff863e98b9cdec6a2d0dfd48ee%2FUntitled%20Diagram%20(1).png?generation=1610226098358123&amp;alt=media\" alt=\"My solution\"></p>",
      "rawMarkdown": "Hi folks!\n\nMy solution has something different from what I've seen until now. I have two outputs in a serial fashion:\n\n- First the user answer prediction (A, B, C, D).\n- Then, inside the NN graph I use the correct answer id to gather the correctness probability.\n\nThis way I have two different backwards signals that helped me to grow a little bit in LB.\n\nOther things to notice:\n- Log-normalization applied to time features (they have a pareto-like distribution). At the end of the competition I realized I didn't _need_ to clip these features with such normalization.\n- All tags usage.\n- Continuous task_container_id (adds user expertise information).\n- I add a mask to make sure sequence elements in the same container don't attend to each other.\n- Pad tokens and lectures are masked in the loss functions.\n\nThe rule I used to choose whether a feature should go on encoder or decoder is: static information on encoder, variable information on decoder.\n\n![My solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F71579dff863e98b9cdec6a2d0dfd48ee%2FUntitled%20Diagram%20(1).png?generation=1610226098358123&alt=media)",
      "votes": null
    },
    {
      "id": "1144688",
      "postDate": "01/08/2021 15:59:24",
      "content": "<p>Congratulation, <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>! I also predict user answer, but just a joint prediction (i.e. there are 2 losses) in my encoder-decoder SAINT + like model. It helps me to avoid overfitting, and gain some boost.</p>",
      "rawMarkdown": "Congratulation, @claverru! I also predict user answer, but just a joint prediction (i.e. there are 2 losses) in my encoder-decoder SAINT + like model. It helps me to avoid overfitting, and gain some boost.",
      "votes": null
    },
    {
      "id": "1144694",
      "postDate": "01/08/2021 16:04:48",
      "content": "<p>Yeah nice! I tried that too and indeed it helped to reduce overfitting. Hope to see you around next comp, mate!</p>",
      "rawMarkdown": "Yeah nice! I tried that too and indeed it helped to reduce overfitting. Hope to see you around next comp, mate!",
      "votes": null
    },
    {
      "id": "1145877",
      "postDate": "01/09/2021 12:00:31",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>! I learned a lot from your transformer notebook and discussions. Are you planning to share your final solution code?</p>",
      "rawMarkdown": "Congratulations @claverru! I learned a lot from your transformer notebook and discussions. Are you planning to share your final solution code?",
      "votes": null
    },
    {
      "id": "1146066",
      "postDate": "01/09/2021 14:15:42",
      "content": "<p>I didn't think about it, I have everything in my local now and it would be somehow a lot of work. Maybe will do it within the next weeks. Thank you mate, I apreciate your appreciation :D</p>",
      "rawMarkdown": "I didn't think about it, I have everything in my local now and it would be somehow a lot of work. Maybe will do it within the next weeks. Thank you mate, I apreciate your appreciation :D",
      "votes": null
    },
    {
      "id": "1146139",
      "postDate": "01/09/2021 15:01:47",
      "content": "<p>Congrats on solo silver medal <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> and thanks for sharing your solution </p>",
      "rawMarkdown": "Congrats on solo silver medal @claverru and thanks for sharing your solution",
      "votes": null
    },
    {
      "id": "1146670",
      "postDate": "01/09/2021 23:44:46",
      "content": "<p>Thank you for sharing and congrats on your 58th place. Well visualized figure describing your approach!</p>",
      "rawMarkdown": "Thank you for sharing and congrats on your 58th place. Well visualized figure describing your approach!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144688,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/08/2021 15:59:24",
      "content": "<p>Congratulation, <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>! I also predict user answer, but just a joint prediction (i.e. there are 2 losses) in my encoder-decoder SAINT + like model. It helps me to avoid overfitting, and gain some boost.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144694,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "01/08/2021 16:04:48",
          "content": "<p>Yeah nice! I tried that too and indeed it helped to reduce overfitting. Hope to see you around next comp, mate!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145877,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "01/09/2021 12:00:31",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>! I learned a lot from your transformer notebook and discussions. Are you planning to share your final solution code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146066,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "01/09/2021 14:15:42",
          "content": "<p>I didn't think about it, I have everything in my local now and it would be somehow a lot of work. Maybe will do it within the next weeks. Thank you mate, I apreciate your appreciation :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146139,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/09/2021 15:01:47",
      "content": "<p>Congrats on solo silver medal <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> and thanks for sharing your solution </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146670,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/09/2021 23:44:46",
      "content": "<p>Thank you for sharing and congrats on your 58th place. Well visualized figure describing your approach!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144665": "Hi folks!\n\nMy solution has something different from what I've seen until now. I have two outputs in a serial fashion:\n\n- First the user answer prediction (A, B, C, D).\n- Then, inside the NN graph I use the correct answer id to gather the correctness probability.\n\nThis way I have two different backwards signals that helped me to grow a little bit in LB.\n\nOther things to notice:\n- Log-normalization applied to time features (they have a pareto-like distribution). At the end of the competition I realized I didn't _need_ to clip these features with such normalization.\n- All tags usage.\n- Continuous task_container_id (adds user expertise information).\n- I add a mask to make sure sequence elements in the same container don't attend to each other.\n- Pad tokens and lectures are masked in the loss functions.\n\nThe rule I used to choose whether a feature should go on encoder or decoder is: static information on encoder, variable information on decoder.\n\n![My solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1820636%2F71579dff863e98b9cdec6a2d0dfd48ee%2FUntitled%20Diagram%20(1).png?generation=1610226098358123&alt=media)",
    "1144688": "Congratulation, @claverru! I also predict user answer, but just a joint prediction (i.e. there are 2 losses) in my encoder-decoder SAINT + like model. It helps me to avoid overfitting, and gain some boost.",
    "1144694": "Yeah nice! I tried that too and indeed it helped to reduce overfitting. Hope to see you around next comp, mate!",
    "1145877": "Congratulations @claverru! I learned a lot from your transformer notebook and discussions. Are you planning to share your final solution code?",
    "1146066": "I didn't think about it, I have everything in my local now and it would be somehow a lot of work. Maybe will do it within the next weeks. Thank you mate, I apreciate your appreciation :D",
    "1146139": "Congrats on solo silver medal @claverru and thanks for sharing your solution",
    "1146670": "Thank you for sharing and congrats on your 58th place. Well visualized figure describing your approach!"
  },
  "source": "meta"
}