{
  "id": 460401,
  "title": "58th solution: simple Roformer model",
  "url": "/competitions/stanford-ribonanza-rna-folding/writeups/darek-k-eczek-58th-solution-simple-roformer-model",
  "author_name": "",
  "post_date": "2023-12-09T03:57:00.415792100Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to the host and Kaggle for this amazing competition! </p>\n<p>My goal was to try some simple experiments so that I can better appreciate reading the top solutions, so all-in-all I am pretty happy about the result. </p>\n<p>Thank you for really great public notebooks that helped me get started quickly: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">RNA Starter</a> by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a></li>\n<li><a href=\"https://www.kaggle.com/code/barteksadlej123/rna-transformer-with-rotary-embedding\" target=\"_blank\">RNA Transformer with Rotary Embedding</a> by <a href=\"https://www.kaggle.com/barteksadlej123\" target=\"_blank\">@barteksadlej123</a> </li>\n</ul>\n<p>Also thanks a lot for super insightful discussions by the hosts, especially the generalization test shared by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>. </p>\n<p>I used the fastai pipeline from the first notebook and the dataset/encoding (sequence+structure) from the second one. Models tested:</p>\n<ul>\n<li>Debertav2 - worked pretty well on cv/lb but didn't generalize well on the test visualization. I hoped the relative position embeddings would work well, but it didn't work out over longer distances. If anyone was able to overcome that, I'd really love to learn about it!</li>\n<li>Longformer - sliding window attention worked much better on the test visualization, but didn't get good cv/lb scores. </li>\n<li>Roformer - this model finally started looking well on the test visualization, but underperformed Deberta on cv/lb. I decided to go with it anyways because it seemed like it should generalize well to private lb. </li>\n</ul>\n<p>Training scheme: </p>\n<ul>\n<li>80 epochs with high lr (5e-4), one cycle schedule, and high regularization (dropout 0.3, eps 1e-7) on noisier data (signal_to_noise &gt; 0.6, reads &gt; 100)</li>\n<li>10 epochs with lower lr (1e-5), one cycle schedule, removing dropout (0.0) and lower eps (1e-12) on clean data (SN_filter == 1)</li>\n</ul>\n<p>Final submission is a single Roformer model, blend of 4 folds. </p>",
  "messages": [
    {
      "id": "2554321",
      "postDate": "12/09/2023 03:57:00",
      "content": "<p>Thanks to the host and Kaggle for this amazing competition! </p>\n<p>My goal was to try some simple experiments so that I can better appreciate reading the top solutions, so all-in-all I am pretty happy about the result. </p>\n<p>Thank you for really great public notebooks that helped me get started quickly: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb\" target=\"_blank\">RNA Starter</a> by <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a></li>\n<li><a href=\"https://www.kaggle.com/code/barteksadlej123/rna-transformer-with-rotary-embedding\" target=\"_blank\">RNA Transformer with Rotary Embedding</a> by <a href=\"https://www.kaggle.com/barteksadlej123\" target=\"_blank\">@barteksadlej123</a> </li>\n</ul>\n<p>Also thanks a lot for super insightful discussions by the hosts, especially the generalization test shared by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>. </p>\n<p>I used the fastai pipeline from the first notebook and the dataset/encoding (sequence+structure) from the second one. Models tested:</p>\n<ul>\n<li>Debertav2 - worked pretty well on cv/lb but didn't generalize well on the test visualization. I hoped the relative position embeddings would work well, but it didn't work out over longer distances. If anyone was able to overcome that, I'd really love to learn about it!</li>\n<li>Longformer - sliding window attention worked much better on the test visualization, but didn't get good cv/lb scores. </li>\n<li>Roformer - this model finally started looking well on the test visualization, but underperformed Deberta on cv/lb. I decided to go with it anyways because it seemed like it should generalize well to private lb. </li>\n</ul>\n<p>Training scheme: </p>\n<ul>\n<li>80 epochs with high lr (5e-4), one cycle schedule, and high regularization (dropout 0.3, eps 1e-7) on noisier data (signal_to_noise &gt; 0.6, reads &gt; 100)</li>\n<li>10 epochs with lower lr (1e-5), one cycle schedule, removing dropout (0.0) and lower eps (1e-12) on clean data (SN_filter == 1)</li>\n</ul>\n<p>Final submission is a single Roformer model, blend of 4 folds. </p>",
      "rawMarkdown": "Thanks to the host and Kaggle for this amazing competition! \n\nMy goal was to try some simple experiments so that I can better appreciate reading the top solutions, so all-in-all I am pretty happy about the result. \n\nThank you for really great public notebooks that helped me get started quickly: \n- [RNA Starter](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb) by @iafoss\n- [RNA Transformer with Rotary Embedding](https://www.kaggle.com/code/barteksadlej123/rna-transformer-with-rotary-embedding) by @barteksadlej123 \n\nAlso thanks a lot for super insightful discussions by the hosts, especially the generalization test shared by @shujun717. \n\nI used the fastai pipeline from the first notebook and the dataset/encoding (sequence+structure) from the second one. Models tested:\n- Debertav2 - worked pretty well on cv/lb but didn't generalize well on the test visualization. I hoped the relative position embeddings would work well, but it didn't work out over longer distances. If anyone was able to overcome that, I'd really love to learn about it!\n- Longformer - sliding window attention worked much better on the test visualization, but didn't get good cv/lb scores. \n- Roformer - this model finally started looking well on the test visualization, but underperformed Deberta on cv/lb. I decided to go with it anyways because it seemed like it should generalize well to private lb. \n\nTraining scheme: \n- 80 epochs with high lr (5e-4), one cycle schedule, and high regularization (dropout 0.3, eps 1e-7) on noisier data (signal_to_noise > 0.6, reads > 100)\n- 10 epochs with lower lr (1e-5), one cycle schedule, removing dropout (0.0) and lower eps (1e-12) on clean data (SN_filter == 1)\n\nFinal submission is a single Roformer model, blend of 4 folds.",
      "votes": null
    },
    {
      "id": "2558989",
      "postDate": "12/12/2023 14:15:03",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> is it possible to train a HuggingFace model (as you mentioned <code>deberta</code>) on a very different vocabulary to the one which was pretrained? I personally didn't joined this comp, but the RNA sequences are series of letters, much different to the vocabulary tokens used in the pretrained models. I wonder how you can overcome this, like how can you train the architecture from scratch?</p>",
      "rawMarkdown": "thedrcat is it possible to train a HuggingFace model (as you mentioned `deberta`) on a very different vocabulary to the one which was pretrained? I personally didn't joined this comp, but the RNA sequences are series of letters, much different to the vocabulary tokens used in the pretrained models. I wonder how you can overcome this, like how can you train the architecture from scratch?",
      "votes": null
    },
    {
      "id": "2559019",
      "postDate": "12/12/2023 14:40:47",
      "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> yes, you need to create a new model from config, and in the config you need to adjust the vocab size. You can either create/use a corresponding tokenizer, or handle tokenization yourself. Here's an example: </p>\n<pre><code> = DebertaConfig()\n =  \n = DebertaModel(config)\n</code></pre>",
      "rawMarkdown": "alejopaullier yes, you need to create a new model from config, and in the config you need to adjust the vocab size. You can either create/use a corresponding tokenizer, or handle tokenization yourself. Here's an example: \n```\nconfig = DebertaConfig()\nconfig.vocab_size = 4 # or more, depending on how you tokenize, if you use cls/eos/pad tokens etc.\nmodel = DebertaModel(config)\n```",
      "votes": null
    },
    {
      "id": "2559049",
      "postDate": "12/12/2023 15:05:15",
      "content": "<p>I used pretrained deberta-v3 from HuggingFace and it works. </p>",
      "rawMarkdown": "I used pretrained deberta-v3 from HuggingFace and it works.",
      "votes": null
    },
    {
      "id": "2559162",
      "postDate": "12/12/2023 16:25:33",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> <a href=\"https://www.kaggle.com/suvalex\" target=\"_blank\">@suvalex</a> thanks, does anyone of you know a code where this is done? Sorry for the trouble</p>",
      "rawMarkdown": "thedrcat @suvalex thanks, does anyone of you know a code where this is done? Sorry for the trouble",
      "votes": null
    },
    {
      "id": "2559199",
      "postDate": "12/12/2023 16:42:44",
      "content": "<p>Split a sequence and works like a regular text transformer.</p>",
      "rawMarkdown": "Split a sequence and works like a regular text transformer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2558989,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "12/12/2023 14:15:03",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> is it possible to train a HuggingFace model (as you mentioned <code>deberta</code>) on a very different vocabulary to the one which was pretrained? I personally didn't joined this comp, but the RNA sequences are series of letters, much different to the vocabulary tokens used in the pretrained models. I wonder how you can overcome this, like how can you train the architecture from scratch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2559019,
          "author_name": "thedrcat",
          "author_url": "",
          "post_date": "12/12/2023 14:40:47",
          "content": "<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> yes, you need to create a new model from config, and in the config you need to adjust the vocab size. You can either create/use a corresponding tokenizer, or handle tokenization yourself. Here's an example: </p>\n<pre><code> = DebertaConfig()\n =  \n = DebertaModel(config)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2559049,
          "author_name": "suvalex",
          "author_url": "",
          "post_date": "12/12/2023 15:05:15",
          "content": "<p>I used pretrained deberta-v3 from HuggingFace and it works. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2559162,
              "author_name": "alejopaullier",
              "author_url": "",
              "post_date": "12/12/2023 16:25:33",
              "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> <a href=\"https://www.kaggle.com/suvalex\" target=\"_blank\">@suvalex</a> thanks, does anyone of you know a code where this is done? Sorry for the trouble</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2559199,
                  "author_name": "suvalex",
                  "author_url": "",
                  "post_date": "12/12/2023 16:42:44",
                  "content": "<p>Split a sequence and works like a regular text transformer.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2554321": "Thanks to the host and Kaggle for this amazing competition! \n\nMy goal was to try some simple experiments so that I can better appreciate reading the top solutions, so all-in-all I am pretty happy about the result. \n\nThank you for really great public notebooks that helped me get started quickly: \n- [RNA Starter](https://www.kaggle.com/code/iafoss/rna-starter-0-186-lb) by @iafoss\n- [RNA Transformer with Rotary Embedding](https://www.kaggle.com/code/barteksadlej123/rna-transformer-with-rotary-embedding) by @barteksadlej123 \n\nAlso thanks a lot for super insightful discussions by the hosts, especially the generalization test shared by @shujun717. \n\nI used the fastai pipeline from the first notebook and the dataset/encoding (sequence+structure) from the second one. Models tested:\n- Debertav2 - worked pretty well on cv/lb but didn't generalize well on the test visualization. I hoped the relative position embeddings would work well, but it didn't work out over longer distances. If anyone was able to overcome that, I'd really love to learn about it!\n- Longformer - sliding window attention worked much better on the test visualization, but didn't get good cv/lb scores. \n- Roformer - this model finally started looking well on the test visualization, but underperformed Deberta on cv/lb. I decided to go with it anyways because it seemed like it should generalize well to private lb. \n\nTraining scheme: \n- 80 epochs with high lr (5e-4), one cycle schedule, and high regularization (dropout 0.3, eps 1e-7) on noisier data (signal_to_noise > 0.6, reads > 100)\n- 10 epochs with lower lr (1e-5), one cycle schedule, removing dropout (0.0) and lower eps (1e-12) on clean data (SN_filter == 1)\n\nFinal submission is a single Roformer model, blend of 4 folds.",
    "2558989": "thedrcat is it possible to train a HuggingFace model (as you mentioned `deberta`) on a very different vocabulary to the one which was pretrained? I personally didn't joined this comp, but the RNA sequences are series of letters, much different to the vocabulary tokens used in the pretrained models. I wonder how you can overcome this, like how can you train the architecture from scratch?",
    "2559019": "alejopaullier yes, you need to create a new model from config, and in the config you need to adjust the vocab size. You can either create/use a corresponding tokenizer, or handle tokenization yourself. Here's an example: \n```\nconfig = DebertaConfig()\nconfig.vocab_size = 4 # or more, depending on how you tokenize, if you use cls/eos/pad tokens etc.\nmodel = DebertaModel(config)\n```",
    "2559049": "I used pretrained deberta-v3 from HuggingFace and it works.",
    "2559162": "thedrcat @suvalex thanks, does anyone of you know a code where this is done? Sorry for the trouble",
    "2559199": "Split a sequence and works like a regular text transformer."
  },
  "source": "meta"
}