{
  "id": 209634,
  "title": "Positional Embedding Done Right",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209634",
  "author_name": "عثمان",
  "post_date": "2021-01-08T04:49:51.812000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> and a few of us started <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">discussing mid competition</a> the saint architecture generally and specifically the decision to add positional encoding vectors as oppose to concatenate them. SAINT's entire architecture is about adding all vectors together, only separating between exercise and response embedding with an implicit interaction embedding being built by the model itself.</p>\n<p>Well, our NLP sisters and brothers over at Microsoft <a href=\"https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/\" target=\"_blank\">just released DeBERTa yesterday</a>, which puts their language model at the helm in a myriad of evaluations such as SuperGLUE and surpasses human annotator performance, just as we've seen CV algorithms do in the years prior.</p>\n<p>Their secret sauce?</p>\n<p>Disentangling the positional embeddings:</p>\n<blockquote>\n  <p>DEBERTa proposes a disentangled self-attention mechanism. Unlike BERT where each word in the input layer is represented using a vector which is the sum of its word (content) embedding and position embedding, each word in DeBERTa is represented using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangled matrices based on their contents and relative positions, respectively. This is motivated by the observation that the attention weight of a word pair depends on not only their contents but their relative positions. For example, the dependency between the words “deep” and “learning” is much stronger when they occur next to each other than when they occur in different sentences.</p>\n</blockquote>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure-2_DeBERTa_high-res.jpg\" alt=\"DeBERTa\"></p>\n<ul>\n<li>Check out <a href=\"https://arxiv.org/abs/2006.03654\" target=\"_blank\">the paper</a> on arxiv</li>\n<li><a href=\"https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/\" target=\"_blank\">Microsoft Research</a> website for the publication</li>\n</ul>\n<p>It would be interesting to see some of the top scoring saint-based models converted over to this style architecture, as we have observed different target behavior for newer accounts vs seasons users vs answer spammers on this dataset.</p>",
  "messages": [
    {
      "id": 1143825,
      "postDate": "2021-01-08T04:49:51.813Z",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> and a few of us started <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">discussing mid competition</a> the saint architecture generally and specifically the decision to add positional encoding vectors as oppose to concatenate them. SAINT's entire architecture is about adding all vectors together, only separating between exercise and response embedding with an implicit interaction embedding being built by the model itself.</p>\n<p>Well, our NLP sisters and brothers over at Microsoft <a href=\"https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/\" target=\"_blank\">just released DeBERTa yesterday</a>, which puts their language model at the helm in a myriad of evaluations such as SuperGLUE and surpasses human annotator performance, just as we've seen CV algorithms do in the years prior.</p>\n<p>Their secret sauce?</p>\n<p>Disentangling the positional embeddings:</p>\n<blockquote>\n  <p>DEBERTa proposes a disentangled self-attention mechanism. Unlike BERT where each word in the input layer is represented using a vector which is the sum of its word (content) embedding and position embedding, each word in DeBERTa is represented using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangled matrices based on their contents and relative positions, respectively. This is motivated by the observation that the attention weight of a word pair depends on not only their contents but their relative positions. For example, the dependency between the words “deep” and “learning” is much stronger when they occur next to each other than when they occur in different sentences.</p>\n</blockquote>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure-2_DeBERTa_high-res.jpg\" alt=\"DeBERTa\"></p>\n<ul>\n<li>Check out <a href=\"https://arxiv.org/abs/2006.03654\" target=\"_blank\">the paper</a> on arxiv</li>\n<li><a href=\"https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/\" target=\"_blank\">Microsoft Research</a> website for the publication</li>\n</ul>\n<p>It would be interesting to see some of the top scoring saint-based models converted over to this style architecture, as we have observed different target behavior for newer accounts vs seasons users vs answer spammers on this dataset.</p>",
      "rawMarkdown": "@allohvk and a few of us started [discussing mid competition](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584) the saint architecture generally and specifically the decision to add positional encoding vectors as oppose to concatenate them. SAINT's entire architecture is about adding all vectors together, only separating between exercise and response embedding with an implicit interaction embedding being built by the model itself.\n\nWell, our NLP sisters and brothers over at Microsoft [just released DeBERTa yesterday](https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/), which puts their language model at the helm in a myriad of evaluations such as SuperGLUE and surpasses human annotator performance, just as we've seen CV algorithms do in the years prior.\n\nTheir secret sauce?\n\nDisentangling the positional embeddings:\n\n> DEBERTa proposes a disentangled self-attention mechanism. Unlike BERT where each word in the input layer is represented using a vector which is the sum of its word (content) embedding and position embedding, each word in DeBERTa is represented using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangled matrices based on their contents and relative positions, respectively. This is motivated by the observation that the attention weight of a word pair depends on not only their contents but their relative positions. For example, the dependency between the words “deep” and “learning” is much stronger when they occur next to each other than when they occur in different sentences.\n\n![DeBERTa](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure-2_DeBERTa_high-res.jpg)\n\n- Check out [the paper](https://arxiv.org/abs/2006.03654) on arxiv\n- [Microsoft Research](https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/) website for the publication\n\nIt would be interesting to see some of the top scoring saint-based models converted over to this style architecture, as we have observed different target behavior for newer accounts vs seasons users vs answer spammers on this dataset.",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1143825": "@allohvk and a few of us started [discussing mid competition](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584) the saint architecture generally and specifically the decision to add positional encoding vectors as oppose to concatenate them. SAINT's entire architecture is about adding all vectors together, only separating between exercise and response embedding with an implicit interaction embedding being built by the model itself.\n\nWell, our NLP sisters and brothers over at Microsoft [just released DeBERTa yesterday](https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/), which puts their language model at the helm in a myriad of evaluations such as SuperGLUE and surpasses human annotator performance, just as we've seen CV algorithms do in the years prior.\n\nTheir secret sauce?\n\nDisentangling the positional embeddings:\n\n> DEBERTa proposes a disentangled self-attention mechanism. Unlike BERT where each word in the input layer is represented using a vector which is the sum of its word (content) embedding and position embedding, each word in DeBERTa is represented using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangled matrices based on their contents and relative positions, respectively. This is motivated by the observation that the attention weight of a word pair depends on not only their contents but their relative positions. For example, the dependency between the words “deep” and “learning” is much stronger when they occur next to each other than when they occur in different sentences.\n\n![DeBERTa](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure-2_DeBERTa_high-res.jpg)\n\n- Check out [the paper](https://arxiv.org/abs/2006.03654) on arxiv\n- [Microsoft Research](https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/) website for the publication\n\nIt would be interesting to see some of the top scoring saint-based models converted over to this style architecture, as we have observed different target behavior for newer accounts vs seasons users vs answer spammers on this dataset."
  }
}