{
  "id": 460130,
  "title": "15th place solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460130",
  "author_name": "Javier Martín",
  "post_date": "2023-12-08T01:17:28.019000",
  "votes": 21,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First and foremost, I'd like to express my gratitude to Kaggle for hosting this competition, and to the organizers for their active involvement and responsiveness in the forums. I gained valuable insights from participating in this competition. Special thanks to my teammates, Anton <a href=\"https://www.kaggle.com/ant0nch\" target=\"_blank\">@ant0nch</a> and María <a href=\"https://www.kaggle.com/manaves\" target=\"_blank\">@manaves</a>, and a shout-out to María for her extensive domain knowledge as a biotechnologist. Lastly, I want to extend my thanks to my employer, Freepik, for providing additional computing resources that proved crucial in the final weeks.</p>\n<p>Our solution is based on a single transformer encoder module. We used the standard pytorch implementation with <code>d_model: 256</code>, <code>dim_feedforward: 768</code>, <code>num_layers: 16</code>, <code>dropout: 0.1</code>, GELU activation and normalization first.</p>\n<p>The model was trained on a 5-fold split grouped by clusters, and clusters were obtained using KMeans of A, C, G, U -&gt; 1, 2, 3, 4 mapped vectors.</p>\n<p>We used AdamW with weight decay 1e-2 and OneCycleLR schedule with <code>pct_start: 0.02</code> and <code>max_lr: 1e-3</code>.</p>\n<p>The encoder input is:</p>\n<ul>\n<li>A 252 dimensional vector encoding the sequence as the sum of nucleotide embeddings:</li>\n</ul>\n<pre><code>    self.embs = nn.\n</code></pre>\n<ul>\n<li>We also encoded the secondary structures from the 47 provided algorithms using an embedding bag where only present data counted towards the mean. Data from the structures were encoded like this: . -&gt; 1, ( -&gt; 2, ) -&gt; 3, [ or &lt; or { -&gt; 4, ] or &gt; or } -&gt; 5.</li>\n</ul>\n<pre><code>    self.se_embs = nn.\n</code></pre>\n<ul>\n<li><p>The two embeddings were added together, and then the following features from the BPP files were concatenated for each nt: max, median, mean and std.</p></li>\n<li><p>Additionally, both the eternafold BPP data and adjacency matrices computed from the secondary structures were used as attention biases via a learnable linear layer. We used 48 matrices for regular pair adjacency and another 48 matrices for pseudoknot adjacency, plus the BPP matrix (so total = 48 * 2 + 1). We augmented the given data by generating additional secondary structures with mxfold2.</p></li>\n</ul>\n<pre><code>    self.algo_conv = nn.Conv2d(48 * 2 + 1, nhead, =1, =0, =)\n</code></pre>\n<ul>\n<li>Positional encoding: we used the standard sinusoidal positional encoding from the original transformer paper. We attempt to achieve longer sequence generalization by assigning each nucleotide in the sequence random correlative positions sampled uniformly from <code>[0, 512[</code>. This hopefully forces the network to learn from the relative order of the nucleotides rather than their absolute position.</li>\n</ul>\n<pre><code>    \n    rnd_pos = torch.empty((n, s), dtype=torch.int64, device=pred.device)\n     row  \n        row[] = torch.randperm(.max_seq_len, device=pred.device)[].sort().values\n    pos = .pos(rnd_pos) \n</code></pre>\n<p>We trained several models for 100 epochs with a batch size of 32 and different ablations of the above features, different SNR filtering strategies (&gt;1 and &gt;0.5) and different GroupKFold shuffle seeds. In the end, we averaged the predictions of all 5 folds of our 7 best models.</p>\n<h2>What didn't work:</h2>\n<ul>\n<li>training in bfloat16 gave substantially poorer results</li>\n<li>pseudolabels, or we just didn't know how to implement this properly</li>\n</ul>",
  "messages": [
    {
      "id": 2553038,
      "postDate": "2023-12-08T01:17:28.020Z",
      "content": "<p>First and foremost, I'd like to express my gratitude to Kaggle for hosting this competition, and to the organizers for their active involvement and responsiveness in the forums. I gained valuable insights from participating in this competition. Special thanks to my teammates, Anton <a href=\"https://www.kaggle.com/ant0nch\" target=\"_blank\">@ant0nch</a> and María <a href=\"https://www.kaggle.com/manaves\" target=\"_blank\">@manaves</a>, and a shout-out to María for her extensive domain knowledge as a biotechnologist. Lastly, I want to extend my thanks to my employer, Freepik, for providing additional computing resources that proved crucial in the final weeks.</p>\n<p>Our solution is based on a single transformer encoder module. We used the standard pytorch implementation with <code>d_model: 256</code>, <code>dim_feedforward: 768</code>, <code>num_layers: 16</code>, <code>dropout: 0.1</code>, GELU activation and normalization first.</p>\n<p>The model was trained on a 5-fold split grouped by clusters, and clusters were obtained using KMeans of A, C, G, U -&gt; 1, 2, 3, 4 mapped vectors.</p>\n<p>We used AdamW with weight decay 1e-2 and OneCycleLR schedule with <code>pct_start: 0.02</code> and <code>max_lr: 1e-3</code>.</p>\n<p>The encoder input is:</p>\n<ul>\n<li>A 252 dimensional vector encoding the sequence as the sum of nucleotide embeddings:</li>\n</ul>\n<pre><code>    self.embs = nn.\n</code></pre>\n<ul>\n<li>We also encoded the secondary structures from the 47 provided algorithms using an embedding bag where only present data counted towards the mean. Data from the structures were encoded like this: . -&gt; 1, ( -&gt; 2, ) -&gt; 3, [ or &lt; or { -&gt; 4, ] or &gt; or } -&gt; 5.</li>\n</ul>\n<pre><code>    self.se_embs = nn.\n</code></pre>\n<ul>\n<li><p>The two embeddings were added together, and then the following features from the BPP files were concatenated for each nt: max, median, mean and std.</p></li>\n<li><p>Additionally, both the eternafold BPP data and adjacency matrices computed from the secondary structures were used as attention biases via a learnable linear layer. We used 48 matrices for regular pair adjacency and another 48 matrices for pseudoknot adjacency, plus the BPP matrix (so total = 48 * 2 + 1). We augmented the given data by generating additional secondary structures with mxfold2.</p></li>\n</ul>\n<pre><code>    self.algo_conv = nn.Conv2d(48 * 2 + 1, nhead, =1, =0, =)\n</code></pre>\n<ul>\n<li>Positional encoding: we used the standard sinusoidal positional encoding from the original transformer paper. We attempt to achieve longer sequence generalization by assigning each nucleotide in the sequence random correlative positions sampled uniformly from <code>[0, 512[</code>. This hopefully forces the network to learn from the relative order of the nucleotides rather than their absolute position.</li>\n</ul>\n<pre><code>    \n    rnd_pos = torch.empty((n, s), dtype=torch.int64, device=pred.device)\n     row  \n        row[] = torch.randperm(.max_seq_len, device=pred.device)[].sort().values\n    pos = .pos(rnd_pos) \n</code></pre>\n<p>We trained several models for 100 epochs with a batch size of 32 and different ablations of the above features, different SNR filtering strategies (&gt;1 and &gt;0.5) and different GroupKFold shuffle seeds. In the end, we averaged the predictions of all 5 folds of our 7 best models.</p>\n<h2>What didn't work:</h2>\n<ul>\n<li>training in bfloat16 gave substantially poorer results</li>\n<li>pseudolabels, or we just didn't know how to implement this properly</li>\n</ul>",
      "rawMarkdown": "First and foremost, I'd like to express my gratitude to Kaggle for hosting this competition, and to the organizers for their active involvement and responsiveness in the forums. I gained valuable insights from participating in this competition. Special thanks to my teammates, Anton @ant0nch and María @manaves, and a shout-out to María for her extensive domain knowledge as a biotechnologist. Lastly, I want to extend my thanks to my employer, Freepik, for providing additional computing resources that proved crucial in the final weeks.\n\nOur solution is based on a single transformer encoder module. We used the standard pytorch implementation with `d_model: 256`, `dim_feedforward: 768`, `num_layers: 16`, `dropout: 0.1`, GELU activation and normalization first.\n\nThe model was trained on a 5-fold split grouped by clusters, and clusters were obtained using KMeans of A, C, G, U -> 1, 2, 3, 4 mapped vectors.\n\nWe used AdamW with weight decay 1e-2 and OneCycleLR schedule with `pct_start: 0.02` and `max_lr: 1e-3`.\n\nThe encoder input is:\n\n- A 252 dimensional vector encoding the sequence as the sum of nucleotide embeddings:\n\n```\n    self.embs = nn.Embedding(5, d_model - extra_features)\n```\n\n- We also encoded the secondary structures from the 47 provided algorithms using an embedding bag where only present data counted towards the mean. Data from the structures were encoded like this: . -> 1, ( -> 2, ) -> 3, [ or < or { -> 4, ] or > or } -> 5.\n\n```\n    self.se_embs = nn.EmbeddingBag(6, d_model - extra_features, padding_idx=0)\n```\n\n- The two embeddings were added together, and then the following features from the BPP files were concatenated for each nt: max, median, mean and std.\n\n- Additionally, both the eternafold BPP data and adjacency matrices computed from the secondary structures were used as attention biases via a learnable linear layer. We used 48 matrices for regular pair adjacency and another 48 matrices for pseudoknot adjacency, plus the BPP matrix (so total = 48 * 2 + 1). We augmented the given data by generating additional secondary structures with mxfold2.\n\n```\n    self.algo_conv = nn.Conv2d(48 * 2 + 1, nhead, kernel_size=1, padding=0, bias=True)\n```\n\n- Positional encoding: we used the standard sinusoidal positional encoding from the original transformer paper. We attempt to achieve longer sequence generalization by assigning each nucleotide in the sequence random correlative positions sampled uniformly from `[0, 512[`. This hopefully forces the network to learn from the relative order of the nucleotides rather than their absolute position.\n\n```\n    # Make a sorted list of positions randomly sampled from [0, max_seq_len)\n    rnd_pos = torch.empty((n, s), dtype=torch.int64, device=pred.device)\n    for row in rnd_pos:\n        row[:] = torch.randperm(self.max_seq_len, device=pred.device)[:s].sort().values\n    pos = self.pos(rnd_pos) # N, S, E-4\n```\n\nWe trained several models for 100 epochs with a batch size of 32 and different ablations of the above features, different SNR filtering strategies (>1 and >0.5) and different GroupKFold shuffle seeds. In the end, we averaged the predictions of all 5 folds of our 7 best models.\n\n## What didn't work:\n\n- training in bfloat16 gave substantially poorer results\n- pseudolabels, or we just didn't know how to implement this properly\n",
      "votes": 21
    },
    {
      "id": 2553673,
      "postDate": "2023-12-08T12:35:03.110Z",
      "content": "<p>Congrats, thanks for sharing the solution.</p>",
      "rawMarkdown": "Congrats, thanks for sharing the solution.",
      "votes": 1
    },
    {
      "id": 2553238,
      "postDate": "2023-12-08T05:31:08.733Z",
      "content": "<p>Congratulations on the topping in the LB as the problem is much computationally rich. </p>",
      "rawMarkdown": "Congratulations on the topping in the LB as the problem is much computationally rich. ",
      "votes": 1
    },
    {
      "id": 2553175,
      "postDate": "2023-12-08T03:57:29.683Z",
      "content": "<p>Congratulations. Thanks for sharing the details of your approach. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details of your approach. ",
      "votes": 1
    },
    {
      "id": 2553048,
      "postDate": "2023-12-08T01:29:45.153Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2553673,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-12-08T12:35:03.110000",
      "content": "<p>Congrats, thanks for sharing the solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553238,
      "author_name": "Sarun P M",
      "author_url": "",
      "post_date": "2023-12-08T05:31:08.733000",
      "content": "<p>Congratulations on the topping in the LB as the problem is much computationally rich. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553175,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-08T03:57:29.683000",
      "content": "<p>Congratulations. Thanks for sharing the details of your approach. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2553048,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T01:29:45.153000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553038": "First and foremost, I'd like to express my gratitude to Kaggle for hosting this competition, and to the organizers for their active involvement and responsiveness in the forums. I gained valuable insights from participating in this competition. Special thanks to my teammates, Anton @ant0nch and María @manaves, and a shout-out to María for her extensive domain knowledge as a biotechnologist. Lastly, I want to extend my thanks to my employer, Freepik, for providing additional computing resources that proved crucial in the final weeks.\n\nOur solution is based on a single transformer encoder module. We used the standard pytorch implementation with `d_model: 256`, `dim_feedforward: 768`, `num_layers: 16`, `dropout: 0.1`, GELU activation and normalization first.\n\nThe model was trained on a 5-fold split grouped by clusters, and clusters were obtained using KMeans of A, C, G, U -> 1, 2, 3, 4 mapped vectors.\n\nWe used AdamW with weight decay 1e-2 and OneCycleLR schedule with `pct_start: 0.02` and `max_lr: 1e-3`.\n\nThe encoder input is:\n\n- A 252 dimensional vector encoding the sequence as the sum of nucleotide embeddings:\n\n```\n    self.embs = nn.Embedding(5, d_model - extra_features)\n```\n\n- We also encoded the secondary structures from the 47 provided algorithms using an embedding bag where only present data counted towards the mean. Data from the structures were encoded like this: . -> 1, ( -> 2, ) -> 3, [ or < or { -> 4, ] or > or } -> 5.\n\n```\n    self.se_embs = nn.EmbeddingBag(6, d_model - extra_features, padding_idx=0)\n```\n\n- The two embeddings were added together, and then the following features from the BPP files were concatenated for each nt: max, median, mean and std.\n\n- Additionally, both the eternafold BPP data and adjacency matrices computed from the secondary structures were used as attention biases via a learnable linear layer. We used 48 matrices for regular pair adjacency and another 48 matrices for pseudoknot adjacency, plus the BPP matrix (so total = 48 * 2 + 1). We augmented the given data by generating additional secondary structures with mxfold2.\n\n```\n    self.algo_conv = nn.Conv2d(48 * 2 + 1, nhead, kernel_size=1, padding=0, bias=True)\n```\n\n- Positional encoding: we used the standard sinusoidal positional encoding from the original transformer paper. We attempt to achieve longer sequence generalization by assigning each nucleotide in the sequence random correlative positions sampled uniformly from `[0, 512[`. This hopefully forces the network to learn from the relative order of the nucleotides rather than their absolute position.\n\n```\n    # Make a sorted list of positions randomly sampled from [0, max_seq_len)\n    rnd_pos = torch.empty((n, s), dtype=torch.int64, device=pred.device)\n    for row in rnd_pos:\n        row[:] = torch.randperm(self.max_seq_len, device=pred.device)[:s].sort().values\n    pos = self.pos(rnd_pos) # N, S, E-4\n```\n\nWe trained several models for 100 epochs with a batch size of 32 and different ablations of the above features, different SNR filtering strategies (>1 and >0.5) and different GroupKFold shuffle seeds. In the end, we averaged the predictions of all 5 folds of our 7 best models.\n\n## What didn't work:\n\n- training in bfloat16 gave substantially poorer results\n- pseudolabels, or we just didn't know how to implement this properly\n",
    "2553673": "Congrats, thanks for sharing the solution.",
    "2553238": "Congratulations on the topping in the LB as the problem is much computationally rich. ",
    "2553175": "Congratulations. Thanks for sharing the details of your approach. ",
    "2553048": ""
  }
}