{
  "id": 209800,
  "title": "encoder self-attention : pay attention to questions in the same bunlde",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209800",
  "author_name": "",
  "post_date": "2021-01-08T16:18:23.727494500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, I want to share an observation and want to see if any of you have use the same idea.</p>\n<p>For transformer like model, even for encoder, we should use the causal masking to avoid seeing the future questions.</p>\n<p>However, in this dataset and API, the questions in the same bundle are given at the same time, so I thought I can pay attention to future questions as long as if they are in the same bundle.</p>\n<p>I therefore modified the attention mask, so a question can see other questions in the same bundle. It does help me to gain a bit, but I don't remember the exact number (not very much though).</p>",
  "messages": [
    {
      "id": "1144717",
      "postDate": "01/08/2021 16:18:23",
      "content": "<p>Hi, I want to share an observation and want to see if any of you have use the same idea.</p>\n<p>For transformer like model, even for encoder, we should use the causal masking to avoid seeing the future questions.</p>\n<p>However, in this dataset and API, the questions in the same bundle are given at the same time, so I thought I can pay attention to future questions as long as if they are in the same bundle.</p>\n<p>I therefore modified the attention mask, so a question can see other questions in the same bundle. It does help me to gain a bit, but I don't remember the exact number (not very much though).</p>",
      "rawMarkdown": "Hi, I want to share an observation and want to see if any of you have use the same idea.\n\nFor transformer like model, even for encoder, we should use the causal masking to avoid seeing the future questions.\n\nHowever, in this dataset and API, the questions in the same bundle are given at the same time, so I thought I can pay attention to future questions as long as if they are in the same bundle.\n\nI therefore modified the attention mask, so a question can see other questions in the same bundle. It does help me to gain a bit, but I don't remember the exact number (not very much though).",
      "votes": null
    },
    {
      "id": "1144720",
      "postDate": "01/08/2021 16:19:40",
      "content": "<p>That's clever !</p>",
      "rawMarkdown": "That's clever !",
      "votes": null
    },
    {
      "id": "1144730",
      "postDate": "01/08/2021 16:28:35",
      "content": "<p>I think this was very important to squeeze the last, both for the model to be prevented to look \"back\" at Qs of same bundle (b/c they are not really \"back), but also for the decoder to decode all Qs from the same bundle at once.</p>",
      "rawMarkdown": "I think this was very important to squeeze the last, both for the model to be prevented to look \"back\" at Qs of same bundle (b/c they are not really \"back), but also for the decoder to decode all Qs from the same bundle at once.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144720,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "01/08/2021 16:19:40",
      "content": "<p>That's clever !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144730,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/08/2021 16:28:35",
      "content": "<p>I think this was very important to squeeze the last, both for the model to be prevented to look \"back\" at Qs of same bundle (b/c they are not really \"back), but also for the decoder to decode all Qs from the same bundle at once.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144717": "Hi, I want to share an observation and want to see if any of you have use the same idea.\n\nFor transformer like model, even for encoder, we should use the causal masking to avoid seeing the future questions.\n\nHowever, in this dataset and API, the questions in the same bundle are given at the same time, so I thought I can pay attention to future questions as long as if they are in the same bundle.\n\nI therefore modified the attention mask, so a question can see other questions in the same bundle. It does help me to gain a bit, but I don't remember the exact number (not very much though).",
    "1144720": "That's clever !",
    "1144730": "I think this was very important to squeeze the last, both for the model to be prevented to look \"back\" at Qs of same bundle (b/c they are not really \"back), but also for the decoder to decode all Qs from the same bundle at once."
  },
  "source": "meta"
}