{
  "id": 460277,
  "title": "81st Place Solution for the Stanford Ribonanza RNA Folding Competition",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460277",
  "author_name": "neilus",
  "post_date": "2023-12-08T14:12:44.919000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>I referenced the model used in the <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">ASL Fingerspelling Competition 2nd place solution</a>.<br>\nI employed a model consisting of Conv1D (x3) + Transformer.</p>\n<h1>Details of the Submission</h1>\n<h2>Model</h2>\n<p>I initially used a model with 12 Transformer layers, based on the <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">excellent baseline provided by Iafoss</a>.<br>\nHowever, due to stagnation in LB performance, I investigated other models. <br>\nInfluenced by the model used by the runner-up in a ASL Fingerspelling Competition, which incorporated Conv1D,<br>\nI saw improved performance and thus switched to the following structure:</p>\n<p>pseudo model code:</p>\n<pre><code> _  range(5):\n     _  range(3):\n        x = Conv1DBlock(=192, =17, =4, =0.2)(x)\n    x = TransformerBlock(=192, =2, =4, =0.2, =0.2)(x)\n</code></pre>\n<h2>Dataset</h2>\n<p>Modifications applied to <code>train_data.csv</code> included:</p>\n<ul>\n<li>Filtering for reads &gt;= 100.</li>\n<li>Adjusting for large errors in reactivity and reactivity_error by setting them to NaN, excluding them from loss calculations.\nThis involved computing signal-to-noise for each reactivity value and replacing reactivity with NaN under the following conditions:<ul>\n<li>reactivity_error &gt; 1.5</li>\n<li>reactivity / reactivity_error &gt; 1.5</li></ul></li>\n</ul>\n<p>This approach was inspired by the <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">Stanford COVID Vaccine Competition winner's solution</a></p>\n<h2>My Mistake</h2>\n<p>For the final submission, I used an ensemble approach different from my main score.<br>\nWhile it improved my Public LB score to 0.14381, it detrimentally lowered my Private LB score to 0.18093.</p>\n<p>The mistake was including a model with an extremely poor Private LB score of 0.32353 in the ensemble.<br>\n(comparable to the performance of sample_submission.csv at Private 0.33290)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F2ba2f5941a271b75c9111a6675669677%2F.PNG?generation=1702044697974444&amp;alt=media\" alt=\"\"></p>\n<p>This flawed model was created by merely changing the seed during training, with almost no predictive performance for lengths not present in the private dataset (207, 357, 457). <br>\nThe reason is unclear, but I suspect overfitting to a length of 177. Attempting to address potential leakage in the public LB by setting the MaP of leaked sequences to 0 made no difference.</p>\n<p>I regret not validating with test data lengths like 206, which could have potentially averted this issue.</p>\n<h2>Source</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620</a></li>\n<li><a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb</a></li>\n</ul>",
  "messages": [
    {
      "id": 2553793,
      "postDate": "2023-12-08T14:12:44.920Z",
      "content": "<p>Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>I referenced the model used in the <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">ASL Fingerspelling Competition 2nd place solution</a>.<br>\nI employed a model consisting of Conv1D (x3) + Transformer.</p>\n<h1>Details of the Submission</h1>\n<h2>Model</h2>\n<p>I initially used a model with 12 Transformer layers, based on the <a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">excellent baseline provided by Iafoss</a>.<br>\nHowever, due to stagnation in LB performance, I investigated other models. <br>\nInfluenced by the model used by the runner-up in a ASL Fingerspelling Competition, which incorporated Conv1D,<br>\nI saw improved performance and thus switched to the following structure:</p>\n<p>pseudo model code:</p>\n<pre><code> _  range(5):\n     _  range(3):\n        x = Conv1DBlock(=192, =17, =4, =0.2)(x)\n    x = TransformerBlock(=192, =2, =4, =0.2, =0.2)(x)\n</code></pre>\n<h2>Dataset</h2>\n<p>Modifications applied to <code>train_data.csv</code> included:</p>\n<ul>\n<li>Filtering for reads &gt;= 100.</li>\n<li>Adjusting for large errors in reactivity and reactivity_error by setting them to NaN, excluding them from loss calculations.\nThis involved computing signal-to-noise for each reactivity value and replacing reactivity with NaN under the following conditions:<ul>\n<li>reactivity_error &gt; 1.5</li>\n<li>reactivity / reactivity_error &gt; 1.5</li></ul></li>\n</ul>\n<p>This approach was inspired by the <a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">Stanford COVID Vaccine Competition winner's solution</a></p>\n<h2>My Mistake</h2>\n<p>For the final submission, I used an ensemble approach different from my main score.<br>\nWhile it improved my Public LB score to 0.14381, it detrimentally lowered my Private LB score to 0.18093.</p>\n<p>The mistake was including a model with an extremely poor Private LB score of 0.32353 in the ensemble.<br>\n(comparable to the performance of sample_submission.csv at Private 0.33290)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F2ba2f5941a271b75c9111a6675669677%2F.PNG?generation=1702044697974444&amp;alt=media\" alt=\"\"></p>\n<p>This flawed model was created by merely changing the seed during training, with almost no predictive performance for lengths not present in the private dataset (207, 357, 457). <br>\nThe reason is unclear, but I suspect overfitting to a length of 177. Attempting to address potential leakage in the public LB by setting the MaP of leaked sequences to 0 made no difference.</p>\n<p>I regret not validating with test data lengths like 206, which could have potentially averted this issue.</p>\n<h2>Source</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620</a></li>\n<li><a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb</a></li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!\n\n# Context\n- Business context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\n- Data context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\n\n# Overview of the Approach\nI referenced the model used in the [ASL Fingerspelling Competition 2nd place solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588).\nI employed a model consisting of Conv1D (x3) + Transformer.\n\n# Details of the Submission\n\n## Model\nI initially used a model with 12 Transformer layers, based on the [excellent baseline provided by Iafoss](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb).\nHowever, due to stagnation in LB performance, I investigated other models. \nInfluenced by the model used by the runner-up in a ASL Fingerspelling Competition, which incorporated Conv1D,\nI saw improved performance and thus switched to the following structure:\n\npseudo model code:\n\n```\nfor _ in range(5):\n    for _ in range(3):\n        x = Conv1DBlock(channel_size=192, kernel_size=17, expand_ratio=4, drop_rate=0.2)(x)\n    x = TransformerBlock(channel_size=192, expand=2, num_heads=4, drop_rate=0.2, attn_dropout=0.2)(x)\n```\n\n\n\n## Dataset\nModifications applied to `train_data.csv` included:\n\n- Filtering for reads >= 100.\n- Adjusting for large errors in reactivity and reactivity_error by setting them to NaN, excluding them from loss calculations.\nThis involved computing signal-to-noise for each reactivity value and replacing reactivity with NaN under the following conditions:\n  - reactivity_error > 1.5\n  - reactivity / reactivity_error > 1.5\n\nThis approach was inspired by the [Stanford COVID Vaccine Competition winner's solution](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620)\n\n## My Mistake\nFor the final submission, I used an ensemble approach different from my main score.\nWhile it improved my Public LB score to 0.14381, it detrimentally lowered my Private LB score to 0.18093.\n\nThe mistake was including a model with an extremely poor Private LB score of 0.32353 in the ensemble.\n(comparable to the performance of sample_submission.csv at Private 0.33290)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F2ba2f5941a271b75c9111a6675669677%2F.PNG?generation=1702044697974444&alt=media)\n\nThis flawed model was created by merely changing the seed during training, with almost no predictive performance for lengths not present in the private dataset (207, 357, 457). \nThe reason is unclear, but I suspect overfitting to a length of 177. Attempting to address potential leakage in the public LB by setting the MaP of leaked sequences to 0 made no difference.\n\nI regret not validating with test data lengths like 206, which could have potentially averted this issue.\n\n## Source\n- https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\n- https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\n- https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2553793": "Thanks to the organizers and Kaggle staff for holding the competition, and congratulations to the winners!\n\n# Context\n- Business context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/overview\n- Data context: https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\n\n# Overview of the Approach\nI referenced the model used in the [ASL Fingerspelling Competition 2nd place solution](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588).\nI employed a model consisting of Conv1D (x3) + Transformer.\n\n# Details of the Submission\n\n## Model\nI initially used a model with 12 Transformer layers, based on the [excellent baseline provided by Iafoss](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb).\nHowever, due to stagnation in LB performance, I investigated other models. \nInfluenced by the model used by the runner-up in a ASL Fingerspelling Competition, which incorporated Conv1D,\nI saw improved performance and thus switched to the following structure:\n\npseudo model code:\n\n```\nfor _ in range(5):\n    for _ in range(3):\n        x = Conv1DBlock(channel_size=192, kernel_size=17, expand_ratio=4, drop_rate=0.2)(x)\n    x = TransformerBlock(channel_size=192, expand=2, num_heads=4, drop_rate=0.2, attn_dropout=0.2)(x)\n```\n\n\n\n## Dataset\nModifications applied to `train_data.csv` included:\n\n- Filtering for reads >= 100.\n- Adjusting for large errors in reactivity and reactivity_error by setting them to NaN, excluding them from loss calculations.\nThis involved computing signal-to-noise for each reactivity value and replacing reactivity with NaN under the following conditions:\n  - reactivity_error > 1.5\n  - reactivity / reactivity_error > 1.5\n\nThis approach was inspired by the [Stanford COVID Vaccine Competition winner's solution](https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620)\n\n## My Mistake\nFor the final submission, I used an ensemble approach different from my main score.\nWhile it improved my Public LB score to 0.14381, it detrimentally lowered my Private LB score to 0.18093.\n\nThe mistake was including a model with an extremely poor Private LB score of 0.32353 in the ensemble.\n(comparable to the performance of sample_submission.csv at Private 0.33290)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8577629%2F2ba2f5941a271b75c9111a6675669677%2F.PNG?generation=1702044697974444&alt=media)\n\nThis flawed model was created by merely changing the seed during training, with almost no predictive performance for lengths not present in the private dataset (207, 357, 457). \nThe reason is unclear, but I suspect overfitting to a length of 177. Attempting to address potential leakage in the public LB by setting the MaP of leaked sequences to 0 made no difference.\n\nI regret not validating with test data lengths like 206, which could have potentially averted this issue.\n\n## Source\n- https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\n- https://www.kaggle.com/competitions/stanford-covid-vaccine/discussion/189620\n- https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb"
  }
}