{
  "id": 458121,
  "title": "Improving Generalization ",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/458121",
  "author_name": "",
  "post_date": "2023-11-28T11:02:41.208890900Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Generalization seems to be the most important aspect of this contest </p>\n<p>I'll share : Sliding window quite surprisingly didn't help for generalization, using rotary embeddings helped a bit but still far away… <br>\nThe biggest catch for generalization is how we're dealing with the data/embeddings though trying various augmentation techniques shifting the sequences randomly eg GAB -&gt; ABG -&gt;BGA , etc seemingly helps.<br>\nFor now i'm going to try and do extremally aggressive augmentation and see how it generalizes…</p>\n<p>My cv is at 0.67x and i'm first trying to improve generalization before the score</p>\n<p>haven't implemented bpps as features etc yet, not sure if it helps with generalization, i know that it should help with the score based on this discussions</p>\n<p>What is everyone doing to improve generalization? Anything anyone willing to share?</p>",
  "messages": [
    {
      "id": "2541281",
      "postDate": "11/28/2023 11:02:41",
      "content": "<p>Generalization seems to be the most important aspect of this contest </p>\n<p>I'll share : Sliding window quite surprisingly didn't help for generalization, using rotary embeddings helped a bit but still far away… <br>\nThe biggest catch for generalization is how we're dealing with the data/embeddings though trying various augmentation techniques shifting the sequences randomly eg GAB -&gt; ABG -&gt;BGA , etc seemingly helps.<br>\nFor now i'm going to try and do extremally aggressive augmentation and see how it generalizes…</p>\n<p>My cv is at 0.67x and i'm first trying to improve generalization before the score</p>\n<p>haven't implemented bpps as features etc yet, not sure if it helps with generalization, i know that it should help with the score based on this discussions</p>\n<p>What is everyone doing to improve generalization? Anything anyone willing to share?</p>",
      "rawMarkdown": "Generalization seems to be the most important aspect of this contest \n\nI'll share : Sliding window quite surprisingly didn't help for generalization, using rotary embeddings helped a bit but still far away… \nThe biggest catch for generalization is how we're dealing with the data/embeddings though trying various augmentation techniques shifting the sequences randomly eg GAB -> ABG ->BGA , etc seemingly helps.\nFor now i'm going to try and do extremally aggressive augmentation and see how it generalizes...\n\nMy cv is at 0.67x and i'm first trying to improve generalization before the score\n\nhaven't implemented bpps as features etc yet, not sure if it helps with generalization, i know that it should help with the score based on this discussions\n\nWhat is everyone doing to improve generalization? Anything anyone willing to share?",
      "votes": null
    },
    {
      "id": "2541863",
      "postDate": "11/28/2023 21:25:36",
      "content": "<p>Same issue here, we initially trained on a subset of the highest quality sequences and had great validation score, but super low testing score. It seems the best way for generalization is to use as much of the data as possible, but the training gets much worse. <br>\nI haven't tried data augmentation technics, what do you mean with shifting sequences randomly ? Do you cut sequences into pieces and rearrange them ?</p>",
      "rawMarkdown": "Same issue here, we initially trained on a subset of the highest quality sequences and had great validation score, but super low testing score. It seems the best way for generalization is to use as much of the data as possible, but the training gets much worse. \nI haven't tried data augmentation technics, what do you mean with shifting sequences randomly ? Do you cut sequences into pieces and rearrange them ?",
      "votes": null
    },
    {
      "id": "2541883",
      "postDate": "11/28/2023 22:03:32",
      "content": "<p>what i mean is if a sequence is N units long start reading it from K (a random no)<br>\nread it like K-&gt;N+1-&gt;K<br>\nSo if you have GABAAAA<br>\nit can become BAAAAGA</p>",
      "rawMarkdown": "what i mean is if a sequence is N units long start reading it from K (a random no)\nread it like K->N+1->K\nSo if you have GABAAAA\nit can become BAAAAGA",
      "votes": null
    },
    {
      "id": "2541889",
      "postDate": "11/28/2023 22:15:32",
      "content": "<p>it's interesting that it works, because in theory the new sequence could have a very different structure when rearranged this way</p>",
      "rawMarkdown": "it's interesting that it works, because in theory the new sequence could have a very different structure when rearranged this way",
      "votes": null
    },
    {
      "id": "2541895",
      "postDate": "11/28/2023 22:23:55",
      "content": "<p>turns out just info on the average reactivity of g/a/c/u etc is somewhat correlated<br>\nSee: <a href=\"https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu\" target=\"_blank\">https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu</a></p>\n<p>So if doing this gets rids of certain signals maybe that's what we're learning and that is something which should generalize quite well. </p>\n<p>Please don't use this for final submissions since the correlation may not be accurate enough(or it may😅)</p>",
      "rawMarkdown": "turns out just info on the average reactivity of g/a/c/u etc is somewhat correlated\nSee: https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu\n\nSo if doing this gets rids of certain signals maybe that's what we're learning and that is something which should generalize quite well. \n\nPlease don't use this for final submissions since the correlation may not be accurate enough(or it may😅)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2541863,
      "author_name": "albericlajarte",
      "author_url": "",
      "post_date": "11/28/2023 21:25:36",
      "content": "<p>Same issue here, we initially trained on a subset of the highest quality sequences and had great validation score, but super low testing score. It seems the best way for generalization is to use as much of the data as possible, but the training gets much worse. <br>\nI haven't tried data augmentation technics, what do you mean with shifting sequences randomly ? Do you cut sequences into pieces and rearrange them ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2541883,
          "author_name": "dhruvdhilla",
          "author_url": "",
          "post_date": "11/28/2023 22:03:32",
          "content": "<p>what i mean is if a sequence is N units long start reading it from K (a random no)<br>\nread it like K-&gt;N+1-&gt;K<br>\nSo if you have GABAAAA<br>\nit can become BAAAAGA</p>",
          "votes": null,
          "replies": [
            {
              "id": 2541889,
              "author_name": "albericlajarte",
              "author_url": "",
              "post_date": "11/28/2023 22:15:32",
              "content": "<p>it's interesting that it works, because in theory the new sequence could have a very different structure when rearranged this way</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2541895,
                  "author_name": "dhruvdhilla",
                  "author_url": "",
                  "post_date": "11/28/2023 22:23:55",
                  "content": "<p>turns out just info on the average reactivity of g/a/c/u etc is somewhat correlated<br>\nSee: <a href=\"https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu\" target=\"_blank\">https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu</a></p>\n<p>So if doing this gets rids of certain signals maybe that's what we're learning and that is something which should generalize quite well. </p>\n<p>Please don't use this for final submissions since the correlation may not be accurate enough(or it may😅)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2541281": "Generalization seems to be the most important aspect of this contest \n\nI'll share : Sliding window quite surprisingly didn't help for generalization, using rotary embeddings helped a bit but still far away… \nThe biggest catch for generalization is how we're dealing with the data/embeddings though trying various augmentation techniques shifting the sequences randomly eg GAB -> ABG ->BGA , etc seemingly helps.\nFor now i'm going to try and do extremally aggressive augmentation and see how it generalizes...\n\nMy cv is at 0.67x and i'm first trying to improve generalization before the score\n\nhaven't implemented bpps as features etc yet, not sure if it helps with generalization, i know that it should help with the score based on this discussions\n\nWhat is everyone doing to improve generalization? Anything anyone willing to share?",
    "2541863": "Same issue here, we initially trained on a subset of the highest quality sequences and had great validation score, but super low testing score. It seems the best way for generalization is to use as much of the data as possible, but the training gets much worse. \nI haven't tried data augmentation technics, what do you mean with shifting sequences randomly ? Do you cut sequences into pieces and rearrange them ?",
    "2541883": "what i mean is if a sequence is N units long start reading it from K (a random no)\nread it like K->N+1->K\nSo if you have GABAAAA\nit can become BAAAAGA",
    "2541889": "it's interesting that it works, because in theory the new sequence could have a very different structure when rearranged this way",
    "2541895": "turns out just info on the average reactivity of g/a/c/u etc is somewhat correlated\nSee: https://www.kaggle.com/code/rgieseking/ribonanza-baseline-average-per-nucleotide-acgu\n\nSo if doing this gets rids of certain signals maybe that's what we're learning and that is something which should generalize quite well. \n\nPlease don't use this for final submissions since the correlation may not be accurate enough(or it may😅)"
  },
  "source": "meta"
}