{
  "id": 454999,
  "title": "about data augumentation",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/454999",
  "author_name": "Koshiro",
  "post_date": "2023-11-12T23:37:52.231000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>What do you think about data augumentation? Do you have any good ideas?</p>",
  "messages": [
    {
      "id": 2522694,
      "postDate": "2023-11-12T23:37:52.230Z",
      "content": "<p>What do you think about data augumentation? Do you have any good ideas?</p>",
      "rawMarkdown": "What do you think about data augumentation? Do you have any good ideas?",
      "votes": 6
    },
    {
      "id": 2523082,
      "postDate": "2023-11-13T08:28:30.507Z",
      "content": "<p>The obvious one is flipping the sequence around (i.e. ACGG -&gt; GGCA). </p>\n<p>I implemented it (recently), but have seen no uplift (no degradation either, just slightly slower convergence.)</p>\n<p>I don't know if anyone has had a similar/different experience. </p>",
      "rawMarkdown": "The obvious one is flipping the sequence around (i.e. ACGG -> GGCA). \n\nI implemented it (recently), but have seen no uplift (no degradation either, just slightly slower convergence.)\n\nI don't know if anyone has had a similar/different experience. ",
      "votes": 2,
      "replies": [
        {
          "id": 2523148,
          "postDate": "2023-11-13T10:06:46.113Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2523334,
          "postDate": "2023-11-13T12:22:23.413Z",
          "content": "<p>I have the same result, flipping doesn't help</p>",
          "rawMarkdown": "I have the same result, flipping doesn't help",
          "votes": 5,
          "replies": [
            {
              "id": 2523550,
              "postDate": "2023-11-13T15:07:50.173Z",
              "content": "<p>I also got kinda the same result, but I found out that at least in my case it made the models worse, local validation was similar to other models without reversing the sequences, but if I compare the prediction for the longer sequence that the organizers posted, is way worse than a model trained without reversing the sequences.</p>\n<p>What I think is happening is that it breaks the locality.</p>\n<p>original sequence, inverse sequence:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F561026143dfb4c0aae5f611a7cfd972c%2FScreenshot%20from%202023-11-13%2015-54-08.png?generation=1699887655326634&amp;alt=media\" alt=\"\"></p>\n<p>So in the end just add 70-90% (depends on val size) more tokens that are at best noise, the model ends up discerning the noise from the real distribution but did way more steps that overfit to shorter sequences.</p>\n<p>Could be wrong and maybe someone made it work.</p>",
              "rawMarkdown": "I also got kinda the same result, but I found out that at least in my case it made the models worse, local validation was similar to other models without reversing the sequences, but if I compare the prediction for the longer sequence that the organizers posted, is way worse than a model trained without reversing the sequences.\n\nWhat I think is happening is that it breaks the locality.\n\noriginal sequence, inverse sequence:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F561026143dfb4c0aae5f611a7cfd972c%2FScreenshot%20from%202023-11-13%2015-54-08.png?generation=1699887655326634&alt=media)\n\nSo in the end just add 70-90% (depends on val size) more tokens that are at best noise, the model ends up discerning the noise from the real distribution but did way more steps that overfit to shorter sequences.\n\nCould be wrong and maybe someone made it work.",
              "votes": 4
            }
          ]
        },
        {
          "id": 2523818,
          "postDate": "2023-11-13T19:15:56.087Z",
          "content": "<p>Well there are few years from my biology classes. But as I remember sequences are not flippeable since there is a direction in the structure depending if it goes from 5' to 3' carbons or viceversa. So 5'AUG3' is not the same structure than 5'GUA3'.</p>",
          "rawMarkdown": "Well there are few years from my biology classes. But as I remember sequences are not flippeable since there is a direction in the structure depending if it goes from 5' to 3' carbons or viceversa. So 5'AUG3' is not the same structure than 5'GUA3'.",
          "votes": 5,
          "replies": [
            {
              "id": 2523934,
              "postDate": "2023-11-13T21:18:20.287Z",
              "content": "<p>Interesting. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> , is this the case? <br>\nIs there a canonical ordering to RNA sequences? </p>",
              "rawMarkdown": "Interesting. @rhijudas , is this the case? \nIs there a canonical ordering to RNA sequences? ",
              "votes": 1
            },
            {
              "id": 2524019,
              "postDate": "2023-11-13T23:35:20.153Z",
              "content": "<p>Yes, <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> is right -- there is a canonical ordering to sequences, so that in principle a flipped RNA sequence will not have the same structures as the original. But in practice, such data augmentation might help -- it's difficult to know until you've tried it with your model.</p>",
              "rawMarkdown": "Yes, @sacuscreed is right -- there is a canonical ordering to sequences, so that in principle a flipped RNA sequence will not have the same structures as the original. But in practice, such data augmentation might help -- it's difficult to know until you've tried it with your model.",
              "votes": 3
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2523082,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2023-11-13T08:28:30.507000",
      "content": "<p>The obvious one is flipping the sequence around (i.e. ACGG -&gt; GGCA). </p>\n<p>I implemented it (recently), but have seen no uplift (no degradation either, just slightly slower convergence.)</p>\n<p>I don't know if anyone has had a similar/different experience. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2523148,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-11-13T10:06:46.113000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2523334,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2023-11-13T12:22:23.413000",
          "content": "<p>I have the same result, flipping doesn't help</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2523550,
              "author_name": "Antonio Félix",
              "author_url": "",
              "post_date": "2023-11-13T15:07:50.173000",
              "content": "<p>I also got kinda the same result, but I found out that at least in my case it made the models worse, local validation was similar to other models without reversing the sequences, but if I compare the prediction for the longer sequence that the organizers posted, is way worse than a model trained without reversing the sequences.</p>\n<p>What I think is happening is that it breaks the locality.</p>\n<p>original sequence, inverse sequence:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F561026143dfb4c0aae5f611a7cfd972c%2FScreenshot%20from%202023-11-13%2015-54-08.png?generation=1699887655326634&amp;alt=media\" alt=\"\"></p>\n<p>So in the end just add 70-90% (depends on val size) more tokens that are at best noise, the model ends up discerning the noise from the real distribution but did way more steps that overfit to shorter sequences.</p>\n<p>Could be wrong and maybe someone made it work.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 2523818,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2023-11-13T19:15:56.087000",
          "content": "<p>Well there are few years from my biology classes. But as I remember sequences are not flippeable since there is a direction in the structure depending if it goes from 5' to 3' carbons or viceversa. So 5'AUG3' is not the same structure than 5'GUA3'.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2523934,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2023-11-13T21:18:20.287000",
              "content": "<p>Interesting. <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> , is this the case? <br>\nIs there a canonical ordering to RNA sequences? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2524019,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2023-11-13T23:35:20.153000",
              "content": "<p>Yes, <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> is right -- there is a canonical ordering to sequences, so that in principle a flipped RNA sequence will not have the same structures as the original. But in practice, such data augmentation might help -- it's difficult to know until you've tried it with your model.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2522694": "What do you think about data augumentation? Do you have any good ideas?",
    "2523082": "The obvious one is flipping the sequence around (i.e. ACGG -> GGCA). \n\nI implemented it (recently), but have seen no uplift (no degradation either, just slightly slower convergence.)\n\nI don't know if anyone has had a similar/different experience. "
  }
}