{
  "id": 76583,
  "title": "[Question] Was the popular attention script correct?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76583",
  "author_name": "",
  "post_date": "2019-01-04T12:57:20.606341100Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I just had a look at the attention in the kernel. I previously looked through it in some other source but they are a bit different so I thought maybe the attention in the kernel is wrong?</p>\n\n<p>I got the attention output tensor shaped [batch_size,feature_size]. If I'm not wrong it means the attention was applied on each sample, weighting different features. Wasn't attention supposed to weight different samples on a training sequence? Correct me if I'm wrong thanks.</p>\n\n<p>Update : Turns out I was wrong. The kernel sums up the weighted features instead of weighting features.</p>",
  "messages": [
    {
      "id": "450175",
      "postDate": "01/04/2019 12:57:20",
      "content": "<p>I just had a look at the attention in the kernel. I previously looked through it in some other source but they are a bit different so I thought maybe the attention in the kernel is wrong?</p>\n\n<p>I got the attention output tensor shaped [batch_size,feature_size]. If I'm not wrong it means the attention was applied on each sample, weighting different features. Wasn't attention supposed to weight different samples on a training sequence? Correct me if I'm wrong thanks.</p>\n\n<p>Update : Turns out I was wrong. The kernel sums up the weighted features instead of weighting features.</p>",
      "rawMarkdown": "I just had a look at the attention in the kernel. I previously looked through it in some other source but they are a bit different so I thought maybe the attention in the kernel is wrong?\n\nI got the attention output tensor shaped [batch_size,feature_size]. If I'm not wrong it means the attention was applied on each sample, weighting different features. Wasn't attention supposed to weight different samples on a training sequence? Correct me if I'm wrong thanks.\n\nUpdate : Turns out I was wrong. The kernel sums up the weighted features instead of weighting features.",
      "votes": null
    },
    {
      "id": "450177",
      "postDate": "01/04/2019 13:02:59",
      "content": "<p>Attention just move focus of model to specific part of sequence.It doesn't attend different samples.</p>",
      "rawMarkdown": "Attention just move focus of model to specific part of sequence.It doesn't attend different samples.",
      "votes": null
    },
    {
      "id": "450193",
      "postDate": "01/04/2019 13:28:42",
      "content": "<p>Yes, that's why I’m thinking the attention in the kernel maybe wrong.</p>",
      "rawMarkdown": "Yes, that's why I’m thinking the attention in the kernel maybe wrong.",
      "votes": null
    },
    {
      "id": "450210",
      "postDate": "01/04/2019 13:48:56",
      "content": "<p>Not completely wrong. He/she did one step extra to take weighted average of entire sentence.</p>",
      "rawMarkdown": "Not completely wrong. He/she did one step extra to take weighted average of entire sentence.",
      "votes": null
    },
    {
      "id": "450235",
      "postDate": "01/04/2019 14:29:27",
      "content": "<p>OK thanks I just got how it works. He took the sum of it tho. I still don't know why doing this would help.</p>",
      "rawMarkdown": "OK thanks I just got how it works. He took the sum of it tho. I still don't know why doing this would help.",
      "votes": null
    },
    {
      "id": "450795",
      "postDate": "01/05/2019 19:36:37",
      "content": "<p>Most of the attention ( atleast inside kaggle ) is wrong. Because of not taking into account of  tokens while taking softmax. I implemented all these, but attention makes my model very bad. I will publish my attention kernel after the competition. NB: I am not using attention in my models. </p>",
      "rawMarkdown": "Most of the attention ( atleast inside kaggle ) is wrong. Because of not taking into account of",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450177,
      "author_name": "mayurnewase",
      "author_url": "",
      "post_date": "01/04/2019 13:02:59",
      "content": "<p>Attention just move focus of model to specific part of sequence.It doesn't attend different samples.</p>",
      "votes": null,
      "replies": [
        {
          "id": 450193,
          "author_name": "icemekaveli",
          "author_url": "",
          "post_date": "01/04/2019 13:28:42",
          "content": "<p>Yes, that's why I’m thinking the attention in the kernel maybe wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450795,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "01/05/2019 19:36:37",
          "content": "<p>Most of the attention ( atleast inside kaggle ) is wrong. Because of not taking into account of  tokens while taking softmax. I implemented all these, but attention makes my model very bad. I will publish my attention kernel after the competition. NB: I am not using attention in my models. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 450210,
      "author_name": "mayurnewase",
      "author_url": "",
      "post_date": "01/04/2019 13:48:56",
      "content": "<p>Not completely wrong. He/she did one step extra to take weighted average of entire sentence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 450235,
          "author_name": "icemekaveli",
          "author_url": "",
          "post_date": "01/04/2019 14:29:27",
          "content": "<p>OK thanks I just got how it works. He took the sum of it tho. I still don't know why doing this would help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "450175": "I just had a look at the attention in the kernel. I previously looked through it in some other source but they are a bit different so I thought maybe the attention in the kernel is wrong?\n\nI got the attention output tensor shaped [batch_size,feature_size]. If I'm not wrong it means the attention was applied on each sample, weighting different features. Wasn't attention supposed to weight different samples on a training sequence? Correct me if I'm wrong thanks.\n\nUpdate : Turns out I was wrong. The kernel sums up the weighted features instead of weighting features.",
    "450177": "Attention just move focus of model to specific part of sequence.It doesn't attend different samples.",
    "450193": "Yes, that's why I’m thinking the attention in the kernel maybe wrong.",
    "450210": "Not completely wrong. He/she did one step extra to take weighted average of entire sentence.",
    "450235": "OK thanks I just got how it works. He took the sum of it tho. I still don't know why doing this would help.",
    "450795": "Most of the attention ( atleast inside kaggle ) is wrong. Because of not taking into account of"
  },
  "source": "meta"
}