{
  "id": 442561,
  "title": "First 34 nucleotides are dominated by sequence GGGAACGACUCGAGUAGAGUCGAAAAGAUAUGGA - why ?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/442561",
  "author_name": "",
  "post_date": "2023-09-23T07:54:13.657075100Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First 26 nucleotides are the same for all sequences: GGGAACGACUCGAGUAGAGUCGAAAA <br>\nas explained in info - that is  for some technical reasons, and targets are absent (NAN) for these positions. </p>\n<p>If look further - then up to position 34 majority of sequences are: GGGAACGACUCGAGUAGAGUCGAAAAGAUAUGGA, and it fails from the next position 35 -  see:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&amp;cellId=6\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&amp;cellId=6</a></p>\n<p>Why is that ? </p>\n<p>Positions 26-33 are included in train targets - do we expect some affect on target for these particular reason ?</p>\n<p>PS</p>\n<p>In the test set there is no such phenomena:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&amp;cellId=18\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&amp;cellId=18</a></p>",
  "messages": [
    {
      "id": "2452298",
      "postDate": "09/23/2023 07:54:13",
      "content": "<p>First 26 nucleotides are the same for all sequences: GGGAACGACUCGAGUAGAGUCGAAAA <br>\nas explained in info - that is  for some technical reasons, and targets are absent (NAN) for these positions. </p>\n<p>If look further - then up to position 34 majority of sequences are: GGGAACGACUCGAGUAGAGUCGAAAAGAUAUGGA, and it fails from the next position 35 -  see:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&amp;cellId=6\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&amp;cellId=6</a></p>\n<p>Why is that ? </p>\n<p>Positions 26-33 are included in train targets - do we expect some affect on target for these particular reason ?</p>\n<p>PS</p>\n<p>In the test set there is no such phenomena:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&amp;cellId=18\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&amp;cellId=18</a></p>",
      "rawMarkdown": "First 26 nucleotides are the same for all sequences: GGGAACGACUCGAGUAGAGUCGAAAA \nas explained in info - that is  for some technical reasons, and targets are absent (NAN) for these positions. \n\nIf look further - then up to position 34 majority of sequences are: GGGAACGACUCGAGUAGAGUCGAAAAGAUAUGGA, and it fails from the next position 35 -  see:\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&cellId=6\n\nWhy is that ? \n\nPositions 26-33 are included in train targets - do we expect some affect on target for these particular reason ?\n\nPS\n\nIn the test set there is no such phenomena:\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&cellId=18",
      "votes": null
    },
    {
      "id": "2457053",
      "postDate": "09/26/2023 15:20:37",
      "content": "<p>I think RNA sequences in train data groups with simple order, same lengths groups together for example. It possible than similar sequences with equal fragments (up to position 34) may appear in train data</p>",
      "rawMarkdown": "I think RNA sequences in train data groups with simple order, same lengths groups together for example. It possible than similar sequences with equal fragments (up to position 34) may appear in train data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2457053,
      "author_name": "alexg5",
      "author_url": "",
      "post_date": "09/26/2023 15:20:37",
      "content": "<p>I think RNA sequences in train data groups with simple order, same lengths groups together for example. It possible than similar sequences with equal fragments (up to position 34) may appear in train data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2452298": "First 26 nucleotides are the same for all sequences: GGGAACGACUCGAGUAGAGUCGAAAA \nas explained in info - that is  for some technical reasons, and targets are absent (NAN) for these positions. \n\nIf look further - then up to position 34 majority of sequences are: GGGAACGACUCGAGUAGAGUCGAAAAGAUAUGGA, and it fails from the next position 35 -  see:\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143909103&cellId=6\n\nWhy is that ? \n\nPositions 26-33 are included in train targets - do we expect some affect on target for these particular reason ?\n\nPS\n\nIn the test set there is no such phenomena:\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=143986175&cellId=18",
    "2457053": "I think RNA sequences in train data groups with simple order, same lengths groups together for example. It possible than similar sequences with equal fragments (up to position 34) may appear in train data"
  },
  "source": "meta"
}