{
  "id": 76133,
  "title": "Attention layer",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76133",
  "author_name": "",
  "post_date": "2018-12-29T17:13:22.261463900Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have gone through many notebooks and found an attention layer. Could anybody explain me about it. I am newbie and could not grasp that.\nThanks in advance</p>",
  "messages": [
    {
      "id": "447339",
      "postDate": "12/29/2018 17:13:22",
      "content": "<p>I have gone through many notebooks and found an attention layer. Could anybody explain me about it. I am newbie and could not grasp that.\nThanks in advance</p>",
      "rawMarkdown": "I have gone through many notebooks and found an attention layer. Could anybody explain me about it. I am newbie and could not grasp that.\nThanks in advance",
      "votes": null
    },
    {
      "id": "447348",
      "postDate": "12/29/2018 17:36:04",
      "content": "<p>Here is explanation of different types of attention:\n<a href=\"https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html\">https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html</a></p>",
      "rawMarkdown": "Here is explanation of different types of attention:\nhttps://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html",
      "votes": null
    },
    {
      "id": "447350",
      "postDate": "12/29/2018 17:40:30",
      "content": "<p>Try this as a fairly high coverage starting point: <a href=\"https://medium.com/@joealato/attention-in-nlp-734c6fa9d983\">https://medium.com/@joealato/attention-in-nlp-734c6fa9d983</a></p>\n\n<p>In terms of its use in this challenge most people are using as an alternative to simpler pooling layers over the sequence (usually an RNN output).  In this regard think of the common usage that seems to be in many notebooks here as key-query attention where the keys are the sequence outputs at each step, and we just train a single common query vector.</p>",
      "rawMarkdown": "Try this as a fairly high coverage starting point: https://medium.com/@joealato/attention-in-nlp-734c6fa9d983\n\nIn terms of its use in this challenge most people are using as an alternative to simpler pooling layers over the sequence (usually an RNN output).  In this regard think of the common usage that seems to be in many notebooks here as key-query attention where the keys are the sequence outputs at each step, and we just train a single common query vector.",
      "votes": null
    },
    {
      "id": "447497",
      "postDate": "12/29/2018 23:44:24",
      "content": "<p>thanks :)</p>",
      "rawMarkdown": "thanks :)",
      "votes": null
    },
    {
      "id": "447498",
      "postDate": "12/29/2018 23:44:43",
      "content": "<p>thanks Steve :)</p>",
      "rawMarkdown": "thanks Steve :)",
      "votes": null
    },
    {
      "id": "447514",
      "postDate": "12/30/2018 01:40:08",
      "content": "<p>In a nutshell, for every word you have in a sentence, you are trying to find which word is more significant to that sentence. This is done with the help of a \"network\", which is nothing but a simple neural network again. It operated after you  got your \"RNN\" outputs for the sentence. Otherway to think is, take TF-IDF of words, multiply those values with RNN output vectors, you get same (\"intuitively\")  kind of thing.</p>",
      "rawMarkdown": "In a nutshell, for every word you have in a sentence, you are trying to find which word is more significant to that sentence. This is done with the help of a \"network\", which is nothing but a simple neural network again. It operated after you  got your \"RNN\" outputs for the sentence. Otherway to think is, take TF-IDF of words, multiply those values with RNN output vectors, you get same (\"intuitively\")  kind of thing.",
      "votes": null
    },
    {
      "id": "447828",
      "postDate": "12/30/2018 17:00:22",
      "content": "<p>Wrote this to explain attention for this competition here. Do have a look \n<a href=\"https://mlwhiz.com/blog/2018/12/17/text_classification/\">https://mlwhiz.com/blog/2018/12/17/text_classification/</a></p>",
      "rawMarkdown": "Wrote this to explain attention for this competition here. Do have a look \nhttps://mlwhiz.com/blog/2018/12/17/text_classification/",
      "votes": null
    },
    {
      "id": "447877",
      "postDate": "12/30/2018 18:41:53",
      "content": "<p>thank you that was an interesting article.</p>",
      "rawMarkdown": "thank you that was an interesting article.",
      "votes": null
    },
    {
      "id": "447878",
      "postDate": "12/30/2018 18:42:27",
      "content": "<p>thanks for the answer :)</p>",
      "rawMarkdown": "thanks for the answer :)",
      "votes": null
    },
    {
      "id": "457103",
      "postDate": "01/17/2019 00:24:46",
      "content": "<p>Hi Rahul, great blog post thank you it helps me to understand attention.\nJust one question though - in your diagram for attention it shows calculating vt as exp(dot product(u, ut)). Is u the matrix composed of u1, u2, ..., uT (where T is seq length) and therefore vt should be the matrix multiplication of u with ut (not the dot product)? Thanks</p>",
      "rawMarkdown": "Hi Rahul, great blog post thank you it helps me to understand attention.\nJust one question though - in your diagram for attention it shows calculating vt as exp(dot product(u, ut)). Is u the matrix composed of u1, u2, ..., uT (where T is seq length) and therefore vt should be the matrix multiplication of u with ut (not the dot product)? Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 447348,
      "author_name": "vikas15",
      "author_url": "",
      "post_date": "12/29/2018 17:36:04",
      "content": "<p>Here is explanation of different types of attention:\n<a href=\"https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html\">https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 447497,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "12/29/2018 23:44:24",
          "content": "<p>thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447350,
      "author_name": "stevedraper",
      "author_url": "",
      "post_date": "12/29/2018 17:40:30",
      "content": "<p>Try this as a fairly high coverage starting point: <a href=\"https://medium.com/@joealato/attention-in-nlp-734c6fa9d983\">https://medium.com/@joealato/attention-in-nlp-734c6fa9d983</a></p>\n\n<p>In terms of its use in this challenge most people are using as an alternative to simpler pooling layers over the sequence (usually an RNN output).  In this regard think of the common usage that seems to be in many notebooks here as key-query attention where the keys are the sequence outputs at each step, and we just train a single common query vector.</p>",
      "votes": null,
      "replies": [
        {
          "id": 447498,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "12/29/2018 23:44:43",
          "content": "<p>thanks Steve :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447514,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "12/30/2018 01:40:08",
      "content": "<p>In a nutshell, for every word you have in a sentence, you are trying to find which word is more significant to that sentence. This is done with the help of a \"network\", which is nothing but a simple neural network again. It operated after you  got your \"RNN\" outputs for the sentence. Otherway to think is, take TF-IDF of words, multiply those values with RNN output vectors, you get same (\"intuitively\")  kind of thing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 447878,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "12/30/2018 18:42:27",
          "content": "<p>thanks for the answer :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447828,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "12/30/2018 17:00:22",
      "content": "<p>Wrote this to explain attention for this competition here. Do have a look \n<a href=\"https://mlwhiz.com/blog/2018/12/17/text_classification/\">https://mlwhiz.com/blog/2018/12/17/text_classification/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 447877,
          "author_name": "kalyankkr",
          "author_url": "",
          "post_date": "12/30/2018 18:41:53",
          "content": "<p>thank you that was an interesting article.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457103,
          "author_name": "hedemann",
          "author_url": "",
          "post_date": "01/17/2019 00:24:46",
          "content": "<p>Hi Rahul, great blog post thank you it helps me to understand attention.\nJust one question though - in your diagram for attention it shows calculating vt as exp(dot product(u, ut)). Is u the matrix composed of u1, u2, ..., uT (where T is seq length) and therefore vt should be the matrix multiplication of u with ut (not the dot product)? Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "447339": "I have gone through many notebooks and found an attention layer. Could anybody explain me about it. I am newbie and could not grasp that.\nThanks in advance",
    "447348": "Here is explanation of different types of attention:\nhttps://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html",
    "447350": "Try this as a fairly high coverage starting point: https://medium.com/@joealato/attention-in-nlp-734c6fa9d983\n\nIn terms of its use in this challenge most people are using as an alternative to simpler pooling layers over the sequence (usually an RNN output).  In this regard think of the common usage that seems to be in many notebooks here as key-query attention where the keys are the sequence outputs at each step, and we just train a single common query vector.",
    "447497": "thanks :)",
    "447498": "thanks Steve :)",
    "447514": "In a nutshell, for every word you have in a sentence, you are trying to find which word is more significant to that sentence. This is done with the help of a \"network\", which is nothing but a simple neural network again. It operated after you  got your \"RNN\" outputs for the sentence. Otherway to think is, take TF-IDF of words, multiply those values with RNN output vectors, you get same (\"intuitively\")  kind of thing.",
    "447828": "Wrote this to explain attention for this competition here. Do have a look \nhttps://mlwhiz.com/blog/2018/12/17/text_classification/",
    "447877": "thank you that was an interesting article.",
    "447878": "thanks for the answer :)",
    "457103": "Hi Rahul, great blog post thank you it helps me to understand attention.\nJust one question though - in your diagram for attention it shows calculating vt as exp(dot product(u, ut)). Is u the matrix composed of u1, u2, ..., uT (where T is seq length) and therefore vt should be the matrix multiplication of u with ut (not the dot product)? Thanks"
  },
  "source": "meta"
}