{
  "id": 76148,
  "title": "Very confused on what an attention layer is",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76148",
  "author_name": "",
  "post_date": "2018-12-29T19:03:07.145243900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So from the research I have done here is my current (and likely incorrect) description of an attention layer:</p>\n\n<p>It is a fully connected layer on top of a recurrent neural network. The input nodes of the attention layer are the outputs of the recurrent neural network that that sequential point in time. </p>\n\n<p>Is this correct? If so, why are bidirectional RNNs used with attention layers? Doesn't an attention layer sort of have a global perspective on the sequence of inputs, which is basically what a bidirectional rnn accomplishes? </p>",
  "messages": [
    {
      "id": "447404",
      "postDate": "12/29/2018 19:03:07",
      "content": "<p>So from the research I have done here is my current (and likely incorrect) description of an attention layer:</p>\n\n<p>It is a fully connected layer on top of a recurrent neural network. The input nodes of the attention layer are the outputs of the recurrent neural network that that sequential point in time. </p>\n\n<p>Is this correct? If so, why are bidirectional RNNs used with attention layers? Doesn't an attention layer sort of have a global perspective on the sequence of inputs, which is basically what a bidirectional rnn accomplishes? </p>",
      "rawMarkdown": "So from the research I have done here is my current (and likely incorrect) description of an attention layer:\n\nIt is a fully connected layer on top of a recurrent neural network. The input nodes of the attention layer are the outputs of the recurrent neural network that that sequential point in time. \n\nIs this correct? If so, why are bidirectional RNNs used with attention layers? Doesn't an attention layer sort of have a global perspective on the sequence of inputs, which is basically what a bidirectional rnn accomplishes?",
      "votes": null
    },
    {
      "id": "447595",
      "postDate": "12/30/2018 06:20:44",
      "content": "<p>According to my understanding : Attention Layer calculates a weighted sum of the different outputs(at different timestamps) of the Bidirectional RNN. As this is a classification task, the model needs to learn the importance of certain words with respect to others to correctly classify the sentence as sincere/insincere. The Attention layer weights are trained to identify these important words, the words that are weighted more then decide the final result of classification.</p>\n\n<p>Also as discussed in other discussion topics, a MaxPool layer does the same job as the Attention Layer in a classification task and also an Attention layer is more useful for language translation tasks. </p>",
      "rawMarkdown": "According to my understanding : Attention Layer calculates a weighted sum of the different outputs(at different timestamps) of the Bidirectional RNN. As this is a classification task, the model needs to learn the importance of certain words with respect to others to correctly classify the sentence as sincere/insincere. The Attention layer weights are trained to identify these important words, the words that are weighted more then decide the final result of classification.\n\nAlso as discussed in other discussion topics, a MaxPool layer does the same job as the Attention Layer in a classification task and also an Attention layer is more useful for language translation tasks.",
      "votes": null
    },
    {
      "id": "447829",
      "postDate": "12/30/2018 17:00:45",
      "content": "<p>Wrote this to explain attention for this competition here. Do have a look \n<a href=\"https://mlwhiz.com/blog/2018/12/17/text_classification/\">https://mlwhiz.com/blog/2018/12/17/text_classification/</a></p>",
      "rawMarkdown": "Wrote this to explain attention for this competition here. Do have a look \nhttps://mlwhiz.com/blog/2018/12/17/text_classification/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 447595,
      "author_name": "ashish2123",
      "author_url": "",
      "post_date": "12/30/2018 06:20:44",
      "content": "<p>According to my understanding : Attention Layer calculates a weighted sum of the different outputs(at different timestamps) of the Bidirectional RNN. As this is a classification task, the model needs to learn the importance of certain words with respect to others to correctly classify the sentence as sincere/insincere. The Attention layer weights are trained to identify these important words, the words that are weighted more then decide the final result of classification.</p>\n\n<p>Also as discussed in other discussion topics, a MaxPool layer does the same job as the Attention Layer in a classification task and also an Attention layer is more useful for language translation tasks. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447829,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "12/30/2018 17:00:45",
      "content": "<p>Wrote this to explain attention for this competition here. Do have a look \n<a href=\"https://mlwhiz.com/blog/2018/12/17/text_classification/\">https://mlwhiz.com/blog/2018/12/17/text_classification/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "447404": "So from the research I have done here is my current (and likely incorrect) description of an attention layer:\n\nIt is a fully connected layer on top of a recurrent neural network. The input nodes of the attention layer are the outputs of the recurrent neural network that that sequential point in time. \n\nIs this correct? If so, why are bidirectional RNNs used with attention layers? Doesn't an attention layer sort of have a global perspective on the sequence of inputs, which is basically what a bidirectional rnn accomplishes?",
    "447595": "According to my understanding : Attention Layer calculates a weighted sum of the different outputs(at different timestamps) of the Bidirectional RNN. As this is a classification task, the model needs to learn the importance of certain words with respect to others to correctly classify the sentence as sincere/insincere. The Attention layer weights are trained to identify these important words, the words that are weighted more then decide the final result of classification.\n\nAlso as discussed in other discussion topics, a MaxPool layer does the same job as the Attention Layer in a classification task and also an Attention layer is more useful for language translation tasks.",
    "447829": "Wrote this to explain attention for this competition here. Do have a look \nhttps://mlwhiz.com/blog/2018/12/17/text_classification/"
  },
  "source": "meta"
}