{
  "id": 460904,
  "title": "Approximately 22nd Place Solution",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/rna-2-approximately-22nd-place-solution",
  "author_name": "",
  "post_date": "2023-12-11T17:35:08.645141700Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>\"Approximately\" in the Post Title: Lessons from a Near-Miss Blend and the Power of a Top Model</strong></p>\n<p>In this post, I dive into an intriguing aspect of our team's journey (only my part) in the recent Kaggle competition. Our final submission was a blend that, unfortunately, didn't perform as expected. If we had solely relied on our best model, we would have secured the 22nd place. This post focuses on that model - what it got right, where it faltered, and how it ultimately impacted our team's standing.</p>\n<p>Despite the setbacks, I am grateful to the competition organizers and my teammate for an enriching experience. This competition was not just about the rankings but also about the invaluable learning and the chance to delve deeper into the subject in a competitive format. Here, I'll share insights and takeaways from our journey, hoping they can be useful for fellow Kagglers.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F586539%2F8b21536b8858508549d6e8db4642dd13%2FScreenshot%202023-12-11%20at%2016.48.46.png?generation=1702309752894013&amp;alt=media\" alt=\"\"></p>\n<p><strong>Exploring Model Backbone: From Scratch to Pre-trained Efficiency</strong></p>\n<p>In the begging, I initially contemplated building a model from scratch. The idea was to create a lighter model and significantly reducing the vocabulary size, thus downsizing the model itself. I started with a model based on BERT and DeBERTa, but training was painfully slow.</p>\n<p>I then experimented with pre-training on the task of subsequence masking. Although this approach slightly improved the model's initial performance, the gains were minimal, and the training remained sluggish. It's worth noting that my setup included only a 1080TI GPU, which restricted me to smaller batch sizes, contributing to the slower training speed.</p>\n<p>Eventually, I shifted my strategy to leveraging a pre-trained, smaller language model. That's how I chose 'deberta-v3-xsmall' as my base model. This choice was also motivated by the potential benefits of relative positional embeddings, especially for sequences of varying lengths. This post details this journey, highlighting the challenges and learnings from each step.</p>\n<p><strong>Implementing Auxiliary Losses: A Strategic Approach to Model Improvement</strong></p>\n<p>In my solution, I considered incorporating auxiliary losses, given the plethora of potential candidates. My approach involved adding a dedicated 'head' for each candidate and training the model for five epochs. In the end, I decided to keep only those heads where the validation loss was decreasing, adjusting their weights accordingly. However, I faced time constraints due to personal commitments like work and family:)</p>\n<p>The final model included heads for such as targets, reactivity errors, structure, dataset name, and signal-to-noise ratios. Unfortunately, I couldn't explore the 'bpp' aspect as much as I would have liked. I selected the weights for the losses based on intuition.</p>\n<p><strong>Utilizing the Full Dataset</strong></p>\n<p>In my approach, I made a strategic decision to assign weights to each sample in the loss calculation. These weights were based on the signal-to-noise ratio and the number of reads. Initially, due to my limited computational resources, I incrementally added parts of the dataset to the training process. However, I soon realized the necessity of utilizing the entire dataset.</p>\n<p>This approach significantly extended the training duration. As a result, I couldn't complete the computation for all folds - for instance, in the image above, only 2 out of the 5 folds were processed. This post will explore the challenges and insights gained from this method, particularly how integrating the entire dataset and weighted samples impacted the model's training and performance.</p>\n<p><strong>Blending Checkpoints</strong></p>\n<p>In the concluding phase of my solution for the Kaggle competition, I focused on blending models at different checkpoints. Initially, I was conserving the top 3 checkpoints. As an experiment, I expanded this to include the best 5, and eventually, I blended the top 10 checkpoints. Intriguingly, each time I increased the number of checkpoints in the blend, I observed a corresponding improvement in my leaderboard score.</p>",
  "messages": [
    {
      "id": "2557823",
      "postDate": "12/11/2023 17:35:08",
      "content": "<p><strong>\"Approximately\" in the Post Title: Lessons from a Near-Miss Blend and the Power of a Top Model</strong></p>\n<p>In this post, I dive into an intriguing aspect of our team's journey (only my part) in the recent Kaggle competition. Our final submission was a blend that, unfortunately, didn't perform as expected. If we had solely relied on our best model, we would have secured the 22nd place. This post focuses on that model - what it got right, where it faltered, and how it ultimately impacted our team's standing.</p>\n<p>Despite the setbacks, I am grateful to the competition organizers and my teammate for an enriching experience. This competition was not just about the rankings but also about the invaluable learning and the chance to delve deeper into the subject in a competitive format. Here, I'll share insights and takeaways from our journey, hoping they can be useful for fellow Kagglers.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F586539%2F8b21536b8858508549d6e8db4642dd13%2FScreenshot%202023-12-11%20at%2016.48.46.png?generation=1702309752894013&amp;alt=media\" alt=\"\"></p>\n<p><strong>Exploring Model Backbone: From Scratch to Pre-trained Efficiency</strong></p>\n<p>In the begging, I initially contemplated building a model from scratch. The idea was to create a lighter model and significantly reducing the vocabulary size, thus downsizing the model itself. I started with a model based on BERT and DeBERTa, but training was painfully slow.</p>\n<p>I then experimented with pre-training on the task of subsequence masking. Although this approach slightly improved the model's initial performance, the gains were minimal, and the training remained sluggish. It's worth noting that my setup included only a 1080TI GPU, which restricted me to smaller batch sizes, contributing to the slower training speed.</p>\n<p>Eventually, I shifted my strategy to leveraging a pre-trained, smaller language model. That's how I chose 'deberta-v3-xsmall' as my base model. This choice was also motivated by the potential benefits of relative positional embeddings, especially for sequences of varying lengths. This post details this journey, highlighting the challenges and learnings from each step.</p>\n<p><strong>Implementing Auxiliary Losses: A Strategic Approach to Model Improvement</strong></p>\n<p>In my solution, I considered incorporating auxiliary losses, given the plethora of potential candidates. My approach involved adding a dedicated 'head' for each candidate and training the model for five epochs. In the end, I decided to keep only those heads where the validation loss was decreasing, adjusting their weights accordingly. However, I faced time constraints due to personal commitments like work and family:)</p>\n<p>The final model included heads for such as targets, reactivity errors, structure, dataset name, and signal-to-noise ratios. Unfortunately, I couldn't explore the 'bpp' aspect as much as I would have liked. I selected the weights for the losses based on intuition.</p>\n<p><strong>Utilizing the Full Dataset</strong></p>\n<p>In my approach, I made a strategic decision to assign weights to each sample in the loss calculation. These weights were based on the signal-to-noise ratio and the number of reads. Initially, due to my limited computational resources, I incrementally added parts of the dataset to the training process. However, I soon realized the necessity of utilizing the entire dataset.</p>\n<p>This approach significantly extended the training duration. As a result, I couldn't complete the computation for all folds - for instance, in the image above, only 2 out of the 5 folds were processed. This post will explore the challenges and insights gained from this method, particularly how integrating the entire dataset and weighted samples impacted the model's training and performance.</p>\n<p><strong>Blending Checkpoints</strong></p>\n<p>In the concluding phase of my solution for the Kaggle competition, I focused on blending models at different checkpoints. Initially, I was conserving the top 3 checkpoints. As an experiment, I expanded this to include the best 5, and eventually, I blended the top 10 checkpoints. Intriguingly, each time I increased the number of checkpoints in the blend, I observed a corresponding improvement in my leaderboard score.</p>",
      "rawMarkdown": "**\"Approximately\" in the Post Title: Lessons from a Near-Miss Blend and the Power of a Top Model**\n\nIn this post, I dive into an intriguing aspect of our team's journey (only my part) in the recent Kaggle competition. Our final submission was a blend that, unfortunately, didn't perform as expected. If we had solely relied on our best model, we would have secured the 22nd place. This post focuses on that model - what it got right, where it faltered, and how it ultimately impacted our team's standing.\n\nDespite the setbacks, I am grateful to the competition organizers and my teammate for an enriching experience. This competition was not just about the rankings but also about the invaluable learning and the chance to delve deeper into the subject in a competitive format. Here, I'll share insights and takeaways from our journey, hoping they can be useful for fellow Kagglers.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F586539%2F8b21536b8858508549d6e8db4642dd13%2FScreenshot%202023-12-11%20at%2016.48.46.png?generation=1702309752894013&alt=media)\n\n**Exploring Model Backbone: From Scratch to Pre-trained Efficiency**\n\nIn the begging, I initially contemplated building a model from scratch. The idea was to create a lighter model and significantly reducing the vocabulary size, thus downsizing the model itself. I started with a model based on BERT and DeBERTa, but training was painfully slow.\n\nI then experimented with pre-training on the task of subsequence masking. Although this approach slightly improved the model's initial performance, the gains were minimal, and the training remained sluggish. It's worth noting that my setup included only a 1080TI GPU, which restricted me to smaller batch sizes, contributing to the slower training speed.\n\nEventually, I shifted my strategy to leveraging a pre-trained, smaller language model. That's how I chose 'deberta-v3-xsmall' as my base model. This choice was also motivated by the potential benefits of relative positional embeddings, especially for sequences of varying lengths. This post details this journey, highlighting the challenges and learnings from each step.\n\n**Implementing Auxiliary Losses: A Strategic Approach to Model Improvement**\n\nIn my solution, I considered incorporating auxiliary losses, given the plethora of potential candidates. My approach involved adding a dedicated 'head' for each candidate and training the model for five epochs. In the end, I decided to keep only those heads where the validation loss was decreasing, adjusting their weights accordingly. However, I faced time constraints due to personal commitments like work and family:)\n\nThe final model included heads for such as targets, reactivity errors, structure, dataset name, and signal-to-noise ratios. Unfortunately, I couldn't explore the 'bpp' aspect as much as I would have liked. I selected the weights for the losses based on intuition.\n\n**Utilizing the Full Dataset**\n\nIn my approach, I made a strategic decision to assign weights to each sample in the loss calculation. These weights were based on the signal-to-noise ratio and the number of reads. Initially, due to my limited computational resources, I incrementally added parts of the dataset to the training process. However, I soon realized the necessity of utilizing the entire dataset.\n\nThis approach significantly extended the training duration. As a result, I couldn't complete the computation for all folds - for instance, in the image above, only 2 out of the 5 folds were processed. This post will explore the challenges and insights gained from this method, particularly how integrating the entire dataset and weighted samples impacted the model's training and performance.\n\n**Blending Checkpoints**\n\nIn the concluding phase of my solution for the Kaggle competition, I focused on blending models at different checkpoints. Initially, I was conserving the top 3 checkpoints. As an experiment, I expanded this to include the best 5, and eventually, I blended the top 10 checkpoints. Intriguingly, each time I increased the number of checkpoints in the blend, I observed a corresponding improvement in my leaderboard score.",
      "votes": null
    },
    {
      "id": "2558556",
      "postDate": "12/12/2023 08:04:02",
      "content": "<p>Have you only used sequences, but not used BPP? It is not very clear from the description.</p>",
      "rawMarkdown": "Have you only used sequences, but not used BPP? It is not very clear from the description.",
      "votes": null
    },
    {
      "id": "2558665",
      "postDate": "12/12/2023 09:12:11",
      "content": "<p>Sorry for this <br>\nyes, I didn't use BPP at all</p>",
      "rawMarkdown": "Sorry for this \nyes, I didn't use BPP at all",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2558556,
      "author_name": "alekseytrepetsky",
      "author_url": "",
      "post_date": "12/12/2023 08:04:02",
      "content": "<p>Have you only used sequences, but not used BPP? It is not very clear from the description.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2558665,
          "author_name": "suvalex",
          "author_url": "",
          "post_date": "12/12/2023 09:12:11",
          "content": "<p>Sorry for this <br>\nyes, I didn't use BPP at all</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2557823": "**\"Approximately\" in the Post Title: Lessons from a Near-Miss Blend and the Power of a Top Model**\n\nIn this post, I dive into an intriguing aspect of our team's journey (only my part) in the recent Kaggle competition. Our final submission was a blend that, unfortunately, didn't perform as expected. If we had solely relied on our best model, we would have secured the 22nd place. This post focuses on that model - what it got right, where it faltered, and how it ultimately impacted our team's standing.\n\nDespite the setbacks, I am grateful to the competition organizers and my teammate for an enriching experience. This competition was not just about the rankings but also about the invaluable learning and the chance to delve deeper into the subject in a competitive format. Here, I'll share insights and takeaways from our journey, hoping they can be useful for fellow Kagglers.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F586539%2F8b21536b8858508549d6e8db4642dd13%2FScreenshot%202023-12-11%20at%2016.48.46.png?generation=1702309752894013&alt=media)\n\n**Exploring Model Backbone: From Scratch to Pre-trained Efficiency**\n\nIn the begging, I initially contemplated building a model from scratch. The idea was to create a lighter model and significantly reducing the vocabulary size, thus downsizing the model itself. I started with a model based on BERT and DeBERTa, but training was painfully slow.\n\nI then experimented with pre-training on the task of subsequence masking. Although this approach slightly improved the model's initial performance, the gains were minimal, and the training remained sluggish. It's worth noting that my setup included only a 1080TI GPU, which restricted me to smaller batch sizes, contributing to the slower training speed.\n\nEventually, I shifted my strategy to leveraging a pre-trained, smaller language model. That's how I chose 'deberta-v3-xsmall' as my base model. This choice was also motivated by the potential benefits of relative positional embeddings, especially for sequences of varying lengths. This post details this journey, highlighting the challenges and learnings from each step.\n\n**Implementing Auxiliary Losses: A Strategic Approach to Model Improvement**\n\nIn my solution, I considered incorporating auxiliary losses, given the plethora of potential candidates. My approach involved adding a dedicated 'head' for each candidate and training the model for five epochs. In the end, I decided to keep only those heads where the validation loss was decreasing, adjusting their weights accordingly. However, I faced time constraints due to personal commitments like work and family:)\n\nThe final model included heads for such as targets, reactivity errors, structure, dataset name, and signal-to-noise ratios. Unfortunately, I couldn't explore the 'bpp' aspect as much as I would have liked. I selected the weights for the losses based on intuition.\n\n**Utilizing the Full Dataset**\n\nIn my approach, I made a strategic decision to assign weights to each sample in the loss calculation. These weights were based on the signal-to-noise ratio and the number of reads. Initially, due to my limited computational resources, I incrementally added parts of the dataset to the training process. However, I soon realized the necessity of utilizing the entire dataset.\n\nThis approach significantly extended the training duration. As a result, I couldn't complete the computation for all folds - for instance, in the image above, only 2 out of the 5 folds were processed. This post will explore the challenges and insights gained from this method, particularly how integrating the entire dataset and weighted samples impacted the model's training and performance.\n\n**Blending Checkpoints**\n\nIn the concluding phase of my solution for the Kaggle competition, I focused on blending models at different checkpoints. Initially, I was conserving the top 3 checkpoints. As an experiment, I expanded this to include the best 5, and eventually, I blended the top 10 checkpoints. Intriguingly, each time I increased the number of checkpoints in the blend, I observed a corresponding improvement in my leaderboard score.",
    "2558556": "Have you only used sequences, but not used BPP? It is not very clear from the description.",
    "2558665": "Sorry for this \nyes, I didn't use BPP at all"
  },
  "source": "meta"
}