{
  "id": 575448,
  "title": "Best way to deal with long sequences",
  "url": "/competitions/stanford-rna-3d-folding/discussion/575448",
  "author_name": "",
  "post_date": "2025-04-28T17:24:20.441482300Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What is the best way to deal with long sequences? </p>\n<p>So far I've seen that most algorithms create at some point pairwise representations of shape <code>(seq_length, seq_length, embedding_dim)</code> which is a great memory bottleneck since <code>seq_length</code> can be very large, in the order of thousands. For example, if <code>embedding_dim</code> is quite short, e.g. 32 and the sequence is mid-length, e.g. 512 this would result in a tensor of shape (512,512,32) which is quite large for only 16 GB of VRAM as in P100 GPU. Truncating sequences and then filling with random or 0 values seems the usual way in this competition.</p>\n<p>I wonder if there is a way to reduce this bottleneck:</p>\n<ol>\n<li>For instance, RibonanzaNet seems to first project the embedding dimension to a lower dimensional space in the pairwise representations to somewhat manage this problem.</li>\n<li>Apply a moving window. I haven't tried this yet, but I think two main problems may rise. The first one being the alignment of the resulting predictions and the second one is that it breaks the inherent sequential information of the sequence.</li>\n</ol>\n<p>I quite unsuccessfully tried adding a tokenizer to RibonanzaNet that encodes the sequence and thus results in shortening the sequence length about 4x times. I pretrained this network and the score was a bit poorer than the original RibonanzaNet. This may be due to the fact that the tokenizer embeddings must be trained from scratch and the available data may not be enough to learn good representations.</p>\n<p>I look forward to hearing from you :)</p>",
  "messages": [
    {
      "id": "3189071",
      "postDate": "04/28/2025 17:24:20",
      "content": "<p>What is the best way to deal with long sequences? </p>\n<p>So far I've seen that most algorithms create at some point pairwise representations of shape <code>(seq_length, seq_length, embedding_dim)</code> which is a great memory bottleneck since <code>seq_length</code> can be very large, in the order of thousands. For example, if <code>embedding_dim</code> is quite short, e.g. 32 and the sequence is mid-length, e.g. 512 this would result in a tensor of shape (512,512,32) which is quite large for only 16 GB of VRAM as in P100 GPU. Truncating sequences and then filling with random or 0 values seems the usual way in this competition.</p>\n<p>I wonder if there is a way to reduce this bottleneck:</p>\n<ol>\n<li>For instance, RibonanzaNet seems to first project the embedding dimension to a lower dimensional space in the pairwise representations to somewhat manage this problem.</li>\n<li>Apply a moving window. I haven't tried this yet, but I think two main problems may rise. The first one being the alignment of the resulting predictions and the second one is that it breaks the inherent sequential information of the sequence.</li>\n</ol>\n<p>I quite unsuccessfully tried adding a tokenizer to RibonanzaNet that encodes the sequence and thus results in shortening the sequence length about 4x times. I pretrained this network and the score was a bit poorer than the original RibonanzaNet. This may be due to the fact that the tokenizer embeddings must be trained from scratch and the available data may not be enough to learn good representations.</p>\n<p>I look forward to hearing from you :)</p>",
      "rawMarkdown": "What is the best way to deal with long sequences? \n\nSo far I've seen that most algorithms create at some point pairwise representations of shape `(seq_length, seq_length, embedding_dim)` which is a great memory bottleneck since `seq_length` can be very large, in the order of thousands. For example, if `embedding_dim` is quite short, e.g. 32 and the sequence is mid-length, e.g. 512 this would result in a tensor of shape (512,512,32) which is quite large for only 16 GB of VRAM as in P100 GPU. Truncating sequences and then filling with random or 0 values seems the usual way in this competition.\n\nI wonder if there is a way to reduce this bottleneck:\n1. For instance, RibonanzaNet seems to first project the embedding dimension to a lower dimensional space in the pairwise representations to somewhat manage this problem.\n2. Apply a moving window. I haven't tried this yet, but I think two main problems may rise. The first one being the alignment of the resulting predictions and the second one is that it breaks the inherent sequential information of the sequence.\n\nI quite unsuccessfully tried adding a tokenizer to RibonanzaNet that encodes the sequence and thus results in shortening the sequence length about 4x times. I pretrained this network and the score was a bit poorer than the original RibonanzaNet. This may be due to the fact that the tokenizer embeddings must be trained from scratch and the available data may not be enough to learn good representations.\n\nI look forward to hearing from you :)",
      "votes": null
    },
    {
      "id": "3189072",
      "postDate": "04/28/2025 17:30:27",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> has your team tried tokenizing the sequence? I do not know from a domain perspective if this makes any sense. To me the only benefit would be allowing to deal with longer sequences.</p>",
      "rawMarkdown": "rhijudas has your team tried tokenizing the sequence? I do not know from a domain perspective if this makes any sense. To me the only benefit would be allowing to deal with longer sequences.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3189072,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "04/28/2025 17:30:27",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> has your team tried tokenizing the sequence? I do not know from a domain perspective if this makes any sense. To me the only benefit would be allowing to deal with longer sequences.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3189071": "What is the best way to deal with long sequences? \n\nSo far I've seen that most algorithms create at some point pairwise representations of shape `(seq_length, seq_length, embedding_dim)` which is a great memory bottleneck since `seq_length` can be very large, in the order of thousands. For example, if `embedding_dim` is quite short, e.g. 32 and the sequence is mid-length, e.g. 512 this would result in a tensor of shape (512,512,32) which is quite large for only 16 GB of VRAM as in P100 GPU. Truncating sequences and then filling with random or 0 values seems the usual way in this competition.\n\nI wonder if there is a way to reduce this bottleneck:\n1. For instance, RibonanzaNet seems to first project the embedding dimension to a lower dimensional space in the pairwise representations to somewhat manage this problem.\n2. Apply a moving window. I haven't tried this yet, but I think two main problems may rise. The first one being the alignment of the resulting predictions and the second one is that it breaks the inherent sequential information of the sequence.\n\nI quite unsuccessfully tried adding a tokenizer to RibonanzaNet that encodes the sequence and thus results in shortening the sequence length about 4x times. I pretrained this network and the score was a bit poorer than the original RibonanzaNet. This may be due to the fact that the tokenizer embeddings must be trained from scratch and the available data may not be enough to learn good representations.\n\nI look forward to hearing from you :)",
    "3189072": "rhijudas has your team tried tokenizing the sequence? I do not know from a domain perspective if this makes any sense. To me the only benefit would be allowing to deal with longer sequences."
  },
  "source": "meta"
}