{
  "id": 461545,
  "title": "20th place solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/461545",
  "author_name": "",
  "post_date": "2023-12-15T02:55:56.794104500Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This was a fascinating challenge and the top solutions are really inspiring! My solution is pretty simple, so I am happy that it achieved the 20th place.</p>\n<h1>Model.</h1>\n<p>I use the Transformer model (namely, BertForTokenClassification from huggingface, and I use its final \"logits\" for regression). I set its hyperparameters to <code>hidden_size=768, dropout=0.1, num_layers=8, intermediate_size=1024, heads=16</code>. I did not do any hyperparameter search.<br>\nI represent the task as regression, where per each token I produce N outputs. I use the same model for predicting all the mapping values (no separate models for DMS and 2A3). During training, I mask out the loss for <code>None</code> target values.<br>\nFor both stages, I separate a random 5% valid set. I use it for early stopping and LR scheduling - I use ReduceLROnPlateau.<br>\nI don't do any augmentation (some discussion suggested that randomly reversing sequences does not help)</p>\n<h1>Stage 1. Pretraining</h1>\n<p>I start with pretraining the model on predicting BPP values. Instead of predicting the full matrix, I predict per token 4 values: 1) the sum of all BPPs between this token and all other tokens from the sequence and 2) the top 1, 2, and 3 BPPs between this token and all other tokens from the sequence.<br>\nFor the pretraining dataset, I concatenate training and test sets provided for the competition.<br>\nI run only one pretraining, for a relatively long time.</p>\n<h1>Stage 2. Finetuning</h1>\n<p>For finetuning, I add to the pretraining model a randomly initialized \"probing\" block. It consists of two linear layers, interleaved with ReLU and with internal hidden size of 256. The input to the first layer is a concatenation of activations from the 8 attention layers from the backbone model. The output of the second layer is the final regression. From my experiments, such an approach works better than simply replacing the last layer of the pre-trained Transformer with a new regression layer and finetuning it.<br>\nI start finetuning with 1 epoch of \"warmup\", during which I only train the new \"probing\" block, while other weights are frozen.<br>\nFor the finetuning dataset, I concatenate the provided training dataset (only data with SN_filter == 1) with an external dataset provided in <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397\" target=\"_blank\">this discussion post</a>. I finetune for \"multi-output\" regression - I predict 12 values per token (DMS, 2A3, and all the values from the external dataset that appeared at least 1000 times).<br>\nI noticed that clipping values during training to &lt; 1 works worse than not doing it during training and clipping after inference on the test set. I clip to &gt; 0 during both training and inference.</p>\n<h1>Final solution</h1>\n<p>My final solution is an ensemble of 16 finetuning runs, all starting from the same pretraining run, but at different checkpoints. For some of them, I use stochastic weight averaging of the 10 best checkpoints. I average the predictions of these models and clip the final prediction to &lt; 1.</p>",
  "messages": [
    {
      "id": "2562002",
      "postDate": "12/15/2023 02:55:56",
      "content": "<p>This was a fascinating challenge and the top solutions are really inspiring! My solution is pretty simple, so I am happy that it achieved the 20th place.</p>\n<h1>Model.</h1>\n<p>I use the Transformer model (namely, BertForTokenClassification from huggingface, and I use its final \"logits\" for regression). I set its hyperparameters to <code>hidden_size=768, dropout=0.1, num_layers=8, intermediate_size=1024, heads=16</code>. I did not do any hyperparameter search.<br>\nI represent the task as regression, where per each token I produce N outputs. I use the same model for predicting all the mapping values (no separate models for DMS and 2A3). During training, I mask out the loss for <code>None</code> target values.<br>\nFor both stages, I separate a random 5% valid set. I use it for early stopping and LR scheduling - I use ReduceLROnPlateau.<br>\nI don't do any augmentation (some discussion suggested that randomly reversing sequences does not help)</p>\n<h1>Stage 1. Pretraining</h1>\n<p>I start with pretraining the model on predicting BPP values. Instead of predicting the full matrix, I predict per token 4 values: 1) the sum of all BPPs between this token and all other tokens from the sequence and 2) the top 1, 2, and 3 BPPs between this token and all other tokens from the sequence.<br>\nFor the pretraining dataset, I concatenate training and test sets provided for the competition.<br>\nI run only one pretraining, for a relatively long time.</p>\n<h1>Stage 2. Finetuning</h1>\n<p>For finetuning, I add to the pretraining model a randomly initialized \"probing\" block. It consists of two linear layers, interleaved with ReLU and with internal hidden size of 256. The input to the first layer is a concatenation of activations from the 8 attention layers from the backbone model. The output of the second layer is the final regression. From my experiments, such an approach works better than simply replacing the last layer of the pre-trained Transformer with a new regression layer and finetuning it.<br>\nI start finetuning with 1 epoch of \"warmup\", during which I only train the new \"probing\" block, while other weights are frozen.<br>\nFor the finetuning dataset, I concatenate the provided training dataset (only data with SN_filter == 1) with an external dataset provided in <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397\" target=\"_blank\">this discussion post</a>. I finetune for \"multi-output\" regression - I predict 12 values per token (DMS, 2A3, and all the values from the external dataset that appeared at least 1000 times).<br>\nI noticed that clipping values during training to &lt; 1 works worse than not doing it during training and clipping after inference on the test set. I clip to &gt; 0 during both training and inference.</p>\n<h1>Final solution</h1>\n<p>My final solution is an ensemble of 16 finetuning runs, all starting from the same pretraining run, but at different checkpoints. For some of them, I use stochastic weight averaging of the 10 best checkpoints. I average the predictions of these models and clip the final prediction to &lt; 1.</p>",
      "rawMarkdown": "This was a fascinating challenge and the top solutions are really inspiring! My solution is pretty simple, so I am happy that it achieved the 20th place.\n\n# Model.\nI use the Transformer model (namely, BertForTokenClassification from huggingface, and I use its final \"logits\" for regression). I set its hyperparameters to `hidden_size=768, dropout=0.1, num_layers=8, intermediate_size=1024, heads=16`. I did not do any hyperparameter search.\nI represent the task as regression, where per each token I produce N outputs. I use the same model for predicting all the mapping values (no separate models for DMS and 2A3). During training, I mask out the loss for `None` target values.\nFor both stages, I separate a random 5% valid set. I use it for early stopping and LR scheduling - I use ReduceLROnPlateau.\nI don't do any augmentation (some discussion suggested that randomly reversing sequences does not help)\n\n# Stage 1. Pretraining\nI start with pretraining the model on predicting BPP values. Instead of predicting the full matrix, I predict per token 4 values: 1) the sum of all BPPs between this token and all other tokens from the sequence and 2) the top 1, 2, and 3 BPPs between this token and all other tokens from the sequence.\nFor the pretraining dataset, I concatenate training and test sets provided for the competition.\nI run only one pretraining, for a relatively long time.\n\n# Stage 2. Finetuning\nFor finetuning, I add to the pretraining model a randomly initialized \"probing\" block. It consists of two linear layers, interleaved with ReLU and with internal hidden size of 256. The input to the first layer is a concatenation of activations from the 8 attention layers from the backbone model. The output of the second layer is the final regression. From my experiments, such an approach works better than simply replacing the last layer of the pre-trained Transformer with a new regression layer and finetuning it.\nI start finetuning with 1 epoch of \"warmup\", during which I only train the new \"probing\" block, while other weights are frozen.\nFor the finetuning dataset, I concatenate the provided training dataset (only data with SN_filter == 1) with an external dataset provided in [this discussion post](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397). I finetune for \"multi-output\" regression - I predict 12 values per token (DMS, 2A3, and all the values from the external dataset that appeared at least 1000 times).\nI noticed that clipping values during training to < 1 works worse than not doing it during training and clipping after inference on the test set. I clip to > 0 during both training and inference.\n\n# Final solution\nMy final solution is an ensemble of 16 finetuning runs, all starting from the same pretraining run, but at different checkpoints. For some of them, I use stochastic weight averaging of the 10 best checkpoints. I average the predictions of these models and clip the final prediction to < 1.",
      "votes": null
    },
    {
      "id": "2562957",
      "postDate": "12/15/2023 20:27:21",
      "content": "<p>Congratulations for your meaningful 20th place and thanks for sharing your solution Sacha.</p>",
      "rawMarkdown": "Congratulations for your meaningful 20th place and thanks for sharing your solution Sacha.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2562957,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "12/15/2023 20:27:21",
      "content": "<p>Congratulations for your meaningful 20th place and thanks for sharing your solution Sacha.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2562002": "This was a fascinating challenge and the top solutions are really inspiring! My solution is pretty simple, so I am happy that it achieved the 20th place.\n\n# Model.\nI use the Transformer model (namely, BertForTokenClassification from huggingface, and I use its final \"logits\" for regression). I set its hyperparameters to `hidden_size=768, dropout=0.1, num_layers=8, intermediate_size=1024, heads=16`. I did not do any hyperparameter search.\nI represent the task as regression, where per each token I produce N outputs. I use the same model for predicting all the mapping values (no separate models for DMS and 2A3). During training, I mask out the loss for `None` target values.\nFor both stages, I separate a random 5% valid set. I use it for early stopping and LR scheduling - I use ReduceLROnPlateau.\nI don't do any augmentation (some discussion suggested that randomly reversing sequences does not help)\n\n# Stage 1. Pretraining\nI start with pretraining the model on predicting BPP values. Instead of predicting the full matrix, I predict per token 4 values: 1) the sum of all BPPs between this token and all other tokens from the sequence and 2) the top 1, 2, and 3 BPPs between this token and all other tokens from the sequence.\nFor the pretraining dataset, I concatenate training and test sets provided for the competition.\nI run only one pretraining, for a relatively long time.\n\n# Stage 2. Finetuning\nFor finetuning, I add to the pretraining model a randomly initialized \"probing\" block. It consists of two linear layers, interleaved with ReLU and with internal hidden size of 256. The input to the first layer is a concatenation of activations from the 8 attention layers from the backbone model. The output of the second layer is the final regression. From my experiments, such an approach works better than simply replacing the last layer of the pre-trained Transformer with a new regression layer and finetuning it.\nI start finetuning with 1 epoch of \"warmup\", during which I only train the new \"probing\" block, while other weights are frozen.\nFor the finetuning dataset, I concatenate the provided training dataset (only data with SN_filter == 1) with an external dataset provided in [this discussion post](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/454397). I finetune for \"multi-output\" regression - I predict 12 values per token (DMS, 2A3, and all the values from the external dataset that appeared at least 1000 times).\nI noticed that clipping values during training to < 1 works worse than not doing it during training and clipping after inference on the test set. I clip to > 0 during both training and inference.\n\n# Final solution\nMy final solution is an ensemble of 16 finetuning runs, all starting from the same pretraining run, but at different checkpoints. For some of them, I use stochastic weight averaging of the 10 best checkpoints. I average the predictions of these models and clip the final prediction to < 1.",
    "2562957": "Congratulations for your meaningful 20th place and thanks for sharing your solution Sacha."
  },
  "source": "meta"
}