{
  "id": 444306,
  "title": "Train of length 206 - only 3 core sequences - CORONOVIRUS (?) -  with only small variations around - seems NOT similar to private test",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/444306",
  "author_name": "Alexander Chervov",
  "post_date": "2023-10-01T10:29:11.184000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In train most sequences are of length 177 (95%) and very few of length 206. <br>\nThe private test contain sequences of length 207 mostly (99.9% ? ). </p>\n<p>So it may seems that train of length 206 might serve as \"proxy\" to private test.<br>\nBut it seems NOT the case, because: </p>\n<p>Train sequences of length 206 seems contain only 3 \"core\" sequences - at positions 0 - 100, with small variations around at later positions.<br>\n Here are most frequent sequence in train length 206, taking only first 26+90 positions:<br>\nCode is here: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=17\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=17</a></p>\n<pre><code>First +N  \nUCUACGGACACGAGUAACUCGUCUAUCUUCUGCAGGCUGCUUACGGUUUCGUCCGUGUUGCAGC    \nGGAGCGUCGUGUCUCUUGUACGUCUCGGUCACAAUACACGGUUUCGUCCGGUGCGUGGCAAUUC    \nGGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUC    \n</code></pre>\n<p>Going to blast <a href=\"https://blast.ncbi.nlm.nih.gov/Blast.cgi\" target=\"_blank\">https://blast.ncbi.nlm.nih.gov/Blast.cgi</a><br>\nWith any of 3 of these sequences we will see coronovirus related genomes. </p>\n<p>E.g. blast for the third sequence: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F886c8a597c2cd19254161a641275bc03%2FScreenshot%202023-10-01%20122255.png?generation=1696155796466824&amp;alt=media\" alt=\"\"></p>\n<p>The other way to visualize that sequence of length 206 are very special (i.e. mostly do not have variation in nucleotides) is  the following:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fcb9903b5ee7468425dc64284e8a9c690%2FScreenshot%202023-10-01%20122504.png?generation=1696155932992173&amp;alt=media\" alt=\"\"><br>\nSee: <a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=13\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=13</a></p>\n<hr>\n<p>Another way:<br>\nThree clusters can be seen on PCA and UMAP for one-hot of these sequences:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Facc28a1a79c15aac110d4349e8e0d5d6%2FScreenshot%202023-10-01%20125521.png?generation=1696157768170236&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=10</a></p>\n<p>Actually with ordinal encoding we also see these clusters: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=22\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=22</a></p>",
  "messages": [
    {
      "id": 2463440,
      "postDate": "2023-10-01T10:29:11.183Z",
      "content": "<p>In train most sequences are of length 177 (95%) and very few of length 206. <br>\nThe private test contain sequences of length 207 mostly (99.9% ? ). </p>\n<p>So it may seems that train of length 206 might serve as \"proxy\" to private test.<br>\nBut it seems NOT the case, because: </p>\n<p>Train sequences of length 206 seems contain only 3 \"core\" sequences - at positions 0 - 100, with small variations around at later positions.<br>\n Here are most frequent sequence in train length 206, taking only first 26+90 positions:<br>\nCode is here: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=17\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=17</a></p>\n<pre><code>First +N  \nUCUACGGACACGAGUAACUCGUCUAUCUUCUGCAGGCUGCUUACGGUUUCGUCCGUGUUGCAGC    \nGGAGCGUCGUGUCUCUUGUACGUCUCGGUCACAAUACACGGUUUCGUCCGGUGCGUGGCAAUUC    \nGGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUC    \n</code></pre>\n<p>Going to blast <a href=\"https://blast.ncbi.nlm.nih.gov/Blast.cgi\" target=\"_blank\">https://blast.ncbi.nlm.nih.gov/Blast.cgi</a><br>\nWith any of 3 of these sequences we will see coronovirus related genomes. </p>\n<p>E.g. blast for the third sequence: <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F886c8a597c2cd19254161a641275bc03%2FScreenshot%202023-10-01%20122255.png?generation=1696155796466824&amp;alt=media\" alt=\"\"></p>\n<p>The other way to visualize that sequence of length 206 are very special (i.e. mostly do not have variation in nucleotides) is  the following:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fcb9903b5ee7468425dc64284e8a9c690%2FScreenshot%202023-10-01%20122504.png?generation=1696155932992173&amp;alt=media\" alt=\"\"><br>\nSee: <a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=13\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=13</a></p>\n<hr>\n<p>Another way:<br>\nThree clusters can be seen on PCA and UMAP for one-hot of these sequences:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Facc28a1a79c15aac110d4349e8e0d5d6%2FScreenshot%202023-10-01%20125521.png?generation=1696157768170236&amp;alt=media\" alt=\"\"><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=10\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=10</a></p>\n<p>Actually with ordinal encoding we also see these clusters: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=22\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&amp;cellId=22</a></p>",
      "rawMarkdown": "In train most sequences are of length 177 (95%) and very few of length 206. \nThe private test contain sequences of length 207 mostly (99.9% ? ). \n\nSo it may seems that train of length 206 might serve as \"proxy\" to private test.\nBut it seems NOT the case, because: \n\nTrain sequences of length 206 seems contain only 3 \"core\" sequences - at positions 0 - 100, with small variations around at later positions.\n Here are most frequent sequence in train length 206, taking only first 26+90 positions:\nCode is here: \nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=17\n\n```python\nFirst 26+N  90\nUCUACGGACACGAGUAACUCGUCUAUCUUCUGCAGGCUGCUUACGGUUUCGUCCGUGUUGCAGC    904\nGGAGCGUCGUGUCUCUUGUACGUCUCGGUCACAAUACACGGUUUCGUCCGGUGCGUGGCAAUUC    686\nGGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUC    540\n\n```\n\nGoing to blast https://blast.ncbi.nlm.nih.gov/Blast.cgi\nWith any of 3 of these sequences we will see coronovirus related genomes. \n\nE.g. blast for the third sequence: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F886c8a597c2cd19254161a641275bc03%2FScreenshot%202023-10-01%20122255.png?generation=1696155796466824&alt=media)\n\nThe other way to visualize that sequence of length 206 are very special (i.e. mostly do not have variation in nucleotides) is  the following:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fcb9903b5ee7468425dc64284e8a9c690%2FScreenshot%202023-10-01%20122504.png?generation=1696155932992173&alt=media)\nSee: https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=13\n\n------------------------------------------\n\nAnother way:\nThree clusters can be seen on PCA and UMAP for one-hot of these sequences:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Facc28a1a79c15aac110d4349e8e0d5d6%2FScreenshot%202023-10-01%20125521.png?generation=1696157768170236&alt=media)\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=10\n\nActually with ordinal encoding we also see these clusters: \nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=22\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2463440": "In train most sequences are of length 177 (95%) and very few of length 206. \nThe private test contain sequences of length 207 mostly (99.9% ? ). \n\nSo it may seems that train of length 206 might serve as \"proxy\" to private test.\nBut it seems NOT the case, because: \n\nTrain sequences of length 206 seems contain only 3 \"core\" sequences - at positions 0 - 100, with small variations around at later positions.\n Here are most frequent sequence in train length 206, taking only first 26+90 positions:\nCode is here: \nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=17\n\n```python\nFirst 26+N  90\nUCUACGGACACGAGUAACUCGUCUAUCUUCUGCAGGCUGCUUACGGUUUCGUCCGUGUUGCAGC    904\nGGAGCGUCGUGUCUCUUGUACGUCUCGGUCACAAUACACGGUUUCGUCCGGUGCGUGGCAAUUC    686\nGGAGCAUCGUGUCUCAAGUGCUUCACGGUCACAAUAUACCGUUUCGUCGGGUGCGUGGCAAUUC    540\n\n```\n\nGoing to blast https://blast.ncbi.nlm.nih.gov/Blast.cgi\nWith any of 3 of these sequences we will see coronovirus related genomes. \n\nE.g. blast for the third sequence: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F886c8a597c2cd19254161a641275bc03%2FScreenshot%202023-10-01%20122255.png?generation=1696155796466824&alt=media)\n\nThe other way to visualize that sequence of length 206 are very special (i.e. mostly do not have variation in nucleotides) is  the following:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Fcb9903b5ee7468425dc64284e8a9c690%2FScreenshot%202023-10-01%20122504.png?generation=1696155932992173&alt=media)\nSee: https://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=13\n\n------------------------------------------\n\nAnother way:\nThree clusters can be seen on PCA and UMAP for one-hot of these sequences:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2Facc28a1a79c15aac110d4349e8e0d5d6%2FScreenshot%202023-10-01%20125521.png?generation=1696157768170236&alt=media)\nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=10\n\nActually with ordinal encoding we also see these clusters: \nhttps://www.kaggle.com/code/alexandervc/ribonanza-1-eda?scriptVersionId=144787111&cellId=22\n"
  }
}