{
  "id": 447895,
  "title": "Make Use of Structure",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/447895",
  "author_name": "",
  "post_date": "2023-10-17T17:50:44.082683900Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have some ideas to utilize the structure.</p>\n<ol>\n<li>Sequence x Structure</li>\n</ol>\n<p>Sequences have 4 letters, and structures have 3 simbols (, ., ).</p>\n<p>G * ( = 0<br>\nG * . = 1<br>\nG * ) = 2<br>\nand so on.</p>\n<p>In this way, we can represent the structure with sequences.</p>\n<ol>\n<li>Sequence + Structure</li>\n</ol>\n<p>When there are brackets, they are usually grouped. Therefore, it would be enough to indicate the start and the end. 0~3 are for sequences. 4 and 5 are for open, and 6 and 7 are for close. If the paired RNAs have useful information, the sequential models will find out some relationships btw RNAs in 4 and 5 and ones in 6 and 7.</p>\n<ol>\n<li>Auxiliary task</li>\n</ol>\n<p>As an auxiliary task, models should predict structure from sequence as well. By doing so, the model will generate structure information dynamically.</p>\n<ul>\n<li>Further discussion</li>\n</ul>\n<ol>\n<li><p>I tried this idea, and then, I wanted to submit the model. However, it was not possible because, it takes 48 hours to extract structures from the test sequences. If you are still interested in the implementation, I will share the note.</p></li>\n<li><p>I didn't tried this idea, and it needs the structures of test sets. Therefore, we'd better solve the problem before trying out this idea.</p></li>\n<li><p>Although it requires lots of time to extract structures from train_data, it is still feasible.</p></li>\n</ol>\n<p><a href=\"https://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings\" target=\"_blank\">https://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings</a></p>\n<p>This is the structure data of sequences whose SN_filter values are 1.<br>\nUsing the data, we train our models. We don't need structures of test_data.</p>\n<ul>\n<li>Additional discussion.</li>\n</ul>\n<p>If we already generated structures using arnie, we can generate additional structures. Read the code below.</p>\n<p><a href=\"https://www.kaggle.com/code/crimson206/dataaugmentation\" target=\"_blank\">https://www.kaggle.com/code/crimson206/dataaugmentation</a></p>\n<p>I just make her outdated code runable.<br>\n<a href=\"https://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook</a><br>\n(from the champion of OpenVaccine competition)</p>\n<p>If we solve the runtime issue, we can generate two or more inputs from sequences using the idea 1 or 2. </p>",
  "messages": [
    {
      "id": "2486141",
      "postDate": "10/17/2023 17:50:44",
      "content": "<p>I have some ideas to utilize the structure.</p>\n<ol>\n<li>Sequence x Structure</li>\n</ol>\n<p>Sequences have 4 letters, and structures have 3 simbols (, ., ).</p>\n<p>G * ( = 0<br>\nG * . = 1<br>\nG * ) = 2<br>\nand so on.</p>\n<p>In this way, we can represent the structure with sequences.</p>\n<ol>\n<li>Sequence + Structure</li>\n</ol>\n<p>When there are brackets, they are usually grouped. Therefore, it would be enough to indicate the start and the end. 0~3 are for sequences. 4 and 5 are for open, and 6 and 7 are for close. If the paired RNAs have useful information, the sequential models will find out some relationships btw RNAs in 4 and 5 and ones in 6 and 7.</p>\n<ol>\n<li>Auxiliary task</li>\n</ol>\n<p>As an auxiliary task, models should predict structure from sequence as well. By doing so, the model will generate structure information dynamically.</p>\n<ul>\n<li>Further discussion</li>\n</ul>\n<ol>\n<li><p>I tried this idea, and then, I wanted to submit the model. However, it was not possible because, it takes 48 hours to extract structures from the test sequences. If you are still interested in the implementation, I will share the note.</p></li>\n<li><p>I didn't tried this idea, and it needs the structures of test sets. Therefore, we'd better solve the problem before trying out this idea.</p></li>\n<li><p>Although it requires lots of time to extract structures from train_data, it is still feasible.</p></li>\n</ol>\n<p><a href=\"https://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings\" target=\"_blank\">https://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings</a></p>\n<p>This is the structure data of sequences whose SN_filter values are 1.<br>\nUsing the data, we train our models. We don't need structures of test_data.</p>\n<ul>\n<li>Additional discussion.</li>\n</ul>\n<p>If we already generated structures using arnie, we can generate additional structures. Read the code below.</p>\n<p><a href=\"https://www.kaggle.com/code/crimson206/dataaugmentation\" target=\"_blank\">https://www.kaggle.com/code/crimson206/dataaugmentation</a></p>\n<p>I just make her outdated code runable.<br>\n<a href=\"https://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook</a><br>\n(from the champion of OpenVaccine competition)</p>\n<p>If we solve the runtime issue, we can generate two or more inputs from sequences using the idea 1 or 2. </p>",
      "rawMarkdown": "I have some ideas to utilize the structure.\n\n1. Sequence x Structure\n\nSequences have 4 letters, and structures have 3 simbols (, ., ).\n\nG * ( = 0\nG * . = 1\nG * ) = 2\nand so on.\n\nIn this way, we can represent the structure with sequences.\n\n2. Sequence + Structure\n\nWhen there are brackets, they are usually grouped. Therefore, it would be enough to indicate the start and the end. 0~3 are for sequences. 4 and 5 are for open, and 6 and 7 are for close. If the paired RNAs have useful information, the sequential models will find out some relationships btw RNAs in 4 and 5 and ones in 6 and 7.\n\n3. Auxiliary task\n\nAs an auxiliary task, models should predict structure from sequence as well. By doing so, the model will generate structure information dynamically.\n\n- Further discussion\n\n1. I tried this idea, and then, I wanted to submit the model. However, it was not possible because, it takes 48 hours to extract structures from the test sequences. If you are still interested in the implementation, I will share the note.\n\n2. I didn't tried this idea, and it needs the structures of test sets. Therefore, we'd better solve the problem before trying out this idea.\n\n3. Although it requires lots of time to extract structures from train_data, it is still feasible.\n\nhttps://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings\n\nThis is the structure data of sequences whose SN_filter values are 1.\nUsing the data, we train our models. We don't need structures of test_data.\n\n\n- Additional discussion.\n\nIf we already generated structures using arnie, we can generate additional structures. Read the code below.\n\nhttps://www.kaggle.com/code/crimson206/dataaugmentation\n\nI just make her outdated code runable.\nhttps://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook\n(from the champion of OpenVaccine competition)\n\nIf we solve the runtime issue, we can generate two or more inputs from sequences using the idea 1 or 2.",
      "votes": null
    },
    {
      "id": "2486688",
      "postDate": "10/18/2023 05:37:32",
      "content": "<p>I thought the idea here was that it might be possible to infer the reactivity from the sequence without deciding the whole structure. Since there are many methods for extracting the structure, then not much is added if we go this path? I might be misunderstanding the objective, but that was my understanding. </p>",
      "rawMarkdown": "I thought the idea here was that it might be possible to infer the reactivity from the sequence without deciding the whole structure. Since there are many methods for extracting the structure, then not much is added if we go this path? I might be misunderstanding the objective, but that was my understanding.",
      "votes": null
    },
    {
      "id": "2489136",
      "postDate": "10/19/2023 17:46:59",
      "content": "<blockquote>\n  <p>Since there are many methods for extracting the structure, then not much is added if we go this path?</p>\n</blockquote>\n<p>Keep in mind that while there are many existing methods, they have limited accuracy. Predicting a structure and predicting reactivity from a sequence are (while not identical) analogous problems. Using existing methods that predict structures as part of your solution may be useful (eg because it is using thermodynamic properties that aren't apparent from the sequence itself), but it certainly is not expected to be the entire solution!</p>",
      "rawMarkdown": "> Since there are many methods for extracting the structure, then not much is added if we go this path?\n\nKeep in mind that while there are many existing methods, they have limited accuracy. Predicting a structure and predicting reactivity from a sequence are (while not identical) analogous problems. Using existing methods that predict structures as part of your solution may be useful (eg because it is using thermodynamic properties that aren't apparent from the sequence itself), but it certainly is not expected to be the entire solution!",
      "votes": null
    },
    {
      "id": "2502863",
      "postDate": "10/28/2023 14:32:39",
      "content": "<p>Nice work. I am trying to use your dataset for structure! I am wondering what is column sequence_ext and how did you generate that.</p>",
      "rawMarkdown": "Nice work. I am trying to use your dataset for structure! I am wondering what is column sequence_ext and how did you generate that.",
      "votes": null
    },
    {
      "id": "2503430",
      "postDate": "10/29/2023 04:52:56",
      "content": "<p>Please read the note.</p>\n<p><a href=\"https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer\" target=\"_blank\">https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer</a></p>",
      "rawMarkdown": "Please read the note.\n\nhttps://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer",
      "votes": null
    },
    {
      "id": "2535580",
      "postDate": "11/23/2023 12:53:52",
      "content": "<p>Nice work, could you publish stucture data you generated? the linker <a href=\"https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer\" target=\"_blank\">https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer</a> is not reachable?</p>",
      "rawMarkdown": "Nice work, could you publish stucture data you generated? the linker https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer is not reachable?",
      "votes": null
    },
    {
      "id": "2535881",
      "postDate": "11/23/2023 17:53:38",
      "content": "<p>Nice!</p>\n<p>I don't know if anyone has had a different experience, but I trained a model on secondary structure prediction, and then fine-tuned it on the main task. </p>\n<p>While I saw faster convergence, it converged to the same MAE as without pre-training. </p>",
      "rawMarkdown": "Nice!\n\nI don't know if anyone has had a different experience, but I trained a model on secondary structure prediction, and then fine-tuned it on the main task. \n\nWhile I saw faster convergence, it converged to the same MAE as without pre-training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2486688,
      "author_name": "ajenningsfrankston",
      "author_url": "",
      "post_date": "10/18/2023 05:37:32",
      "content": "<p>I thought the idea here was that it might be possible to infer the reactivity from the sequence without deciding the whole structure. Since there are many methods for extracting the structure, then not much is added if we go this path? I might be misunderstanding the objective, but that was my understanding. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2489136,
          "author_name": "jonathanromano",
          "author_url": "",
          "post_date": "10/19/2023 17:46:59",
          "content": "<blockquote>\n  <p>Since there are many methods for extracting the structure, then not much is added if we go this path?</p>\n</blockquote>\n<p>Keep in mind that while there are many existing methods, they have limited accuracy. Predicting a structure and predicting reactivity from a sequence are (while not identical) analogous problems. Using existing methods that predict structures as part of your solution may be useful (eg because it is using thermodynamic properties that aren't apparent from the sequence itself), but it certainly is not expected to be the entire solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2502863,
      "author_name": "venkatapadavala",
      "author_url": "",
      "post_date": "10/28/2023 14:32:39",
      "content": "<p>Nice work. I am trying to use your dataset for structure! I am wondering what is column sequence_ext and how did you generate that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503430,
          "author_name": "",
          "author_url": "",
          "post_date": "10/29/2023 04:52:56",
          "content": "<p>Please read the note.</p>\n<p><a href=\"https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer\" target=\"_blank\">https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2535580,
      "author_name": "qianren007",
      "author_url": "",
      "post_date": "11/23/2023 12:53:52",
      "content": "<p>Nice work, could you publish stucture data you generated? the linker <a href=\"https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer\" target=\"_blank\">https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer</a> is not reachable?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2535881,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "11/23/2023 17:53:38",
      "content": "<p>Nice!</p>\n<p>I don't know if anyone has had a different experience, but I trained a model on secondary structure prediction, and then fine-tuned it on the main task. </p>\n<p>While I saw faster convergence, it converged to the same MAE as without pre-training. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486141": "I have some ideas to utilize the structure.\n\n1. Sequence x Structure\n\nSequences have 4 letters, and structures have 3 simbols (, ., ).\n\nG * ( = 0\nG * . = 1\nG * ) = 2\nand so on.\n\nIn this way, we can represent the structure with sequences.\n\n2. Sequence + Structure\n\nWhen there are brackets, they are usually grouped. Therefore, it would be enough to indicate the start and the end. 0~3 are for sequences. 4 and 5 are for open, and 6 and 7 are for close. If the paired RNAs have useful information, the sequential models will find out some relationships btw RNAs in 4 and 5 and ones in 6 and 7.\n\n3. Auxiliary task\n\nAs an auxiliary task, models should predict structure from sequence as well. By doing so, the model will generate structure information dynamically.\n\n- Further discussion\n\n1. I tried this idea, and then, I wanted to submit the model. However, it was not possible because, it takes 48 hours to extract structures from the test sequences. If you are still interested in the implementation, I will share the note.\n\n2. I didn't tried this idea, and it needs the structures of test sets. Therefore, we'd better solve the problem before trying out this idea.\n\n3. Although it requires lots of time to extract structures from train_data, it is still feasible.\n\nhttps://www.kaggle.com/datasets/crimson206/rna-smalldatawithstructure/settings\n\nThis is the structure data of sequences whose SN_filter values are 1.\nUsing the data, we train our models. We don't need structures of test_data.\n\n\n- Additional discussion.\n\nIf we already generated structures using arnie, we can generate additional structures. Read the code below.\n\nhttps://www.kaggle.com/code/crimson206/dataaugmentation\n\nI just make her outdated code runable.\nhttps://www.kaggle.com/code/its7171/how-to-generate-augmentation-data/notebook\n(from the champion of OpenVaccine competition)\n\nIf we solve the runtime issue, we can generate two or more inputs from sequences using the idea 1 or 2.",
    "2486688": "I thought the idea here was that it might be possible to infer the reactivity from the sequence without deciding the whole structure. Since there are many methods for extracting the structure, then not much is added if we go this path? I might be misunderstanding the objective, but that was my understanding.",
    "2489136": "> Since there are many methods for extracting the structure, then not much is added if we go this path?\n\nKeep in mind that while there are many existing methods, they have limited accuracy. Predicting a structure and predicting reactivity from a sequence are (while not identical) analogous problems. Using existing methods that predict structures as part of your solution may be useful (eg because it is using thermodynamic properties that aren't apparent from the sequence itself), but it certainly is not expected to be the entire solution!",
    "2502863": "Nice work. I am trying to use your dataset for structure! I am wondering what is column sequence_ext and how did you generate that.",
    "2503430": "Please read the note.\n\nhttps://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer",
    "2535580": "Nice work, could you publish stucture data you generated? the linker https://www.kaggle.com/code/crimson206/rna-seq-struct-flexibletransformer is not reachable?",
    "2535881": "Nice!\n\nI don't know if anyone has had a different experience, but I trained a model on secondary structure prediction, and then fine-tuned it on the main task. \n\nWhile I saw faster convergence, it converged to the same MAE as without pre-training."
  },
  "source": "meta"
}