{
  "id": 580433,
  "title": "Strategies for inference on long sequences",
  "url": "/competitions/stanford-rna-3d-folding/discussion/580433",
  "author_name": "Shujun",
  "post_date": "2025-05-23T21:39:22.554000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Kagglers, there has been some questions about long sequences on future test. While you can assume the sequence length distribution will be similar on future test data to public/private, it will be important to make sure your notebook doesn't OOM so it actually generates a submission.csv. Here are some simple strategies to make sure your notebook actually generates a submission.csv even on very long sequences that would cause OOM:</p>\n<ol>\n<li>first preset a reasonable MAX_LEN for inference and only do inference on the first MAX_LEN residues</li>\n<li>even if your code works for first MAX_LEN residues on public/private, it could still OOM for some unforeseen reason on future test. So you can add a simple try/except loop that tries to do inference on progressively shorter crops following OOMs</li>\n<li>if you'd like to make predictions for the entire sequence still, you can make predictions for overlapping windows. Since predictions are in 3D space in this competition, it is not reasonable to simply concat predictions for different windows. Instead, for every 2 adjacent windows, you can svd align overlapping residues and then concat, which should put the 2 windows roughly in the same frame</li>\n</ol>\n<p>Enjoy the last week of the competition and let me know you have any questions!</p>",
  "messages": [
    {
      "id": 3208318,
      "postDate": "2025-05-23T21:39:22.553Z",
      "content": "<p>Hi Kagglers, there has been some questions about long sequences on future test. While you can assume the sequence length distribution will be similar on future test data to public/private, it will be important to make sure your notebook doesn't OOM so it actually generates a submission.csv. Here are some simple strategies to make sure your notebook actually generates a submission.csv even on very long sequences that would cause OOM:</p>\n<ol>\n<li>first preset a reasonable MAX_LEN for inference and only do inference on the first MAX_LEN residues</li>\n<li>even if your code works for first MAX_LEN residues on public/private, it could still OOM for some unforeseen reason on future test. So you can add a simple try/except loop that tries to do inference on progressively shorter crops following OOMs</li>\n<li>if you'd like to make predictions for the entire sequence still, you can make predictions for overlapping windows. Since predictions are in 3D space in this competition, it is not reasonable to simply concat predictions for different windows. Instead, for every 2 adjacent windows, you can svd align overlapping residues and then concat, which should put the 2 windows roughly in the same frame</li>\n</ol>\n<p>Enjoy the last week of the competition and let me know you have any questions!</p>",
      "rawMarkdown": "Hi Kagglers, there has been some questions about long sequences on future test. While you can assume the sequence length distribution will be similar on future test data to public/private, it will be important to make sure your notebook doesn't OOM so it actually generates a submission.csv. Here are some simple strategies to make sure your notebook actually generates a submission.csv even on very long sequences that would cause OOM:\n\n1. first preset a reasonable MAX_LEN for inference and only do inference on the first MAX_LEN residues\n2. even if your code works for first MAX_LEN residues on public/private, it could still OOM for some unforeseen reason on future test. So you can add a simple try/except loop that tries to do inference on progressively shorter crops following OOMs\n3. if you'd like to make predictions for the entire sequence still, you can make predictions for overlapping windows. Since predictions are in 3D space in this competition, it is not reasonable to simply concat predictions for different windows. Instead, for every 2 adjacent windows, you can svd align overlapping residues and then concat, which should put the 2 windows roughly in the same frame\n\nEnjoy the last week of the competition and let me know you have any questions!\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3208318": "Hi Kagglers, there has been some questions about long sequences on future test. While you can assume the sequence length distribution will be similar on future test data to public/private, it will be important to make sure your notebook doesn't OOM so it actually generates a submission.csv. Here are some simple strategies to make sure your notebook actually generates a submission.csv even on very long sequences that would cause OOM:\n\n1. first preset a reasonable MAX_LEN for inference and only do inference on the first MAX_LEN residues\n2. even if your code works for first MAX_LEN residues on public/private, it could still OOM for some unforeseen reason on future test. So you can add a simple try/except loop that tries to do inference on progressively shorter crops following OOMs\n3. if you'd like to make predictions for the entire sequence still, you can make predictions for overlapping windows. Since predictions are in 3D space in this competition, it is not reasonable to simply concat predictions for different windows. Instead, for every 2 adjacent windows, you can svd align overlapping residues and then concat, which should put the 2 windows roughly in the same frame\n\nEnjoy the last week of the competition and let me know you have any questions!\n"
  }
}