{
  "id": 566910,
  "title": "Need to Innovate on Position Encoding Techniques",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566910",
  "author_name": "Bhargav Borah",
  "post_date": "2025-03-07T10:56:10.299000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As we can already observe from the top public notebooks, the winning solutions of Stanford Ribonanza RNA Folding, and <a href=\"https://www.nature.com/articles/s41586-021-03819-2\" target=\"_blank\">AlphaFold</a>,  we need to work with some variant of the transformer architecture as it is a very strong candidate for dealing with sequential data.</p>\n<p>However, transformers are permutation invariant, that is, they are not able to account for the inherent sequential nature of the input embeddings. Therefore, position encodings are needed which can feed information to the transformer about which part of the sequence a given token (represented by a token embedding) belongs to.</p>\n<p>Some of the well-known position encoding techniques are:</p>\n<ol>\n<li><a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">Rotary Position Embedding (RoPE)</a> (SoTA for NLP tasks)</li>\n<li><a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Sinusoidal Position Encoding </a></li>\n<li><a href=\"https://arxiv.org/abs/2108.12409\" target=\"_blank\">Attention with Linear Biases (ALiBi)</a></li>\n</ol>\n<p>Note that the position encodings can be fixed or learnable.</p>\n<p>The aforementioned techniques are mainly used for NLP tasks, and may not work well for this problem as is. In my opinion, there is a serious need to work with domain knowledge to develop a position encoding technique for a competitive transformer-based solution.</p>\n<p>Thoughts?</p>",
  "messages": [
    {
      "id": 3143573,
      "postDate": "2025-03-07T10:56:10.300Z",
      "content": "<p>As we can already observe from the top public notebooks, the winning solutions of Stanford Ribonanza RNA Folding, and <a href=\"https://www.nature.com/articles/s41586-021-03819-2\" target=\"_blank\">AlphaFold</a>,  we need to work with some variant of the transformer architecture as it is a very strong candidate for dealing with sequential data.</p>\n<p>However, transformers are permutation invariant, that is, they are not able to account for the inherent sequential nature of the input embeddings. Therefore, position encodings are needed which can feed information to the transformer about which part of the sequence a given token (represented by a token embedding) belongs to.</p>\n<p>Some of the well-known position encoding techniques are:</p>\n<ol>\n<li><a href=\"https://arxiv.org/abs/2104.09864\" target=\"_blank\">Rotary Position Embedding (RoPE)</a> (SoTA for NLP tasks)</li>\n<li><a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">Sinusoidal Position Encoding </a></li>\n<li><a href=\"https://arxiv.org/abs/2108.12409\" target=\"_blank\">Attention with Linear Biases (ALiBi)</a></li>\n</ol>\n<p>Note that the position encodings can be fixed or learnable.</p>\n<p>The aforementioned techniques are mainly used for NLP tasks, and may not work well for this problem as is. In my opinion, there is a serious need to work with domain knowledge to develop a position encoding technique for a competitive transformer-based solution.</p>\n<p>Thoughts?</p>",
      "rawMarkdown": "As we can already observe from the top public notebooks, the winning solutions of Stanford Ribonanza RNA Folding, and [AlphaFold](https://www.nature.com/articles/s41586-021-03819-2),  we need to work with some variant of the transformer architecture as it is a very strong candidate for dealing with sequential data.\n\nHowever, transformers are permutation invariant, that is, they are not able to account for the inherent sequential nature of the input embeddings. Therefore, position encodings are needed which can feed information to the transformer about which part of the sequence a given token (represented by a token embedding) belongs to.\n\nSome of the well-known position encoding techniques are:\n1. [Rotary Position Embedding (RoPE)](https://arxiv.org/abs/2104.09864) (SoTA for NLP tasks)\n2. [Sinusoidal Position Encoding ](https://arxiv.org/abs/1706.03762)\n3. [Attention with Linear Biases (ALiBi)](https://arxiv.org/abs/2108.12409)\n\nNote that the position encodings can be fixed or learnable.\n\nThe aforementioned techniques are mainly used for NLP tasks, and may not work well for this problem as is. In my opinion, there is a serious need to work with domain knowledge to develop a position encoding technique for a competitive transformer-based solution.\n\nThoughts?",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3143573": "As we can already observe from the top public notebooks, the winning solutions of Stanford Ribonanza RNA Folding, and [AlphaFold](https://www.nature.com/articles/s41586-021-03819-2),  we need to work with some variant of the transformer architecture as it is a very strong candidate for dealing with sequential data.\n\nHowever, transformers are permutation invariant, that is, they are not able to account for the inherent sequential nature of the input embeddings. Therefore, position encodings are needed which can feed information to the transformer about which part of the sequence a given token (represented by a token embedding) belongs to.\n\nSome of the well-known position encoding techniques are:\n1. [Rotary Position Embedding (RoPE)](https://arxiv.org/abs/2104.09864) (SoTA for NLP tasks)\n2. [Sinusoidal Position Encoding ](https://arxiv.org/abs/1706.03762)\n3. [Attention with Linear Biases (ALiBi)](https://arxiv.org/abs/2108.12409)\n\nNote that the position encodings can be fixed or learnable.\n\nThe aforementioned techniques are mainly used for NLP tasks, and may not work well for this problem as is. In my opinion, there is a serious need to work with domain knowledge to develop a position encoding technique for a competitive transformer-based solution.\n\nThoughts?"
  }
}