{
  "id": 460252,
  "title": "GraphAttention solution approach",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460252",
  "author_name": "VITALIY",
  "post_date": "2023-12-08T12:12:59.053000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi there,thanks Stanford University for this competition!<br>\nMight be someone will found my work is interesting.<br>\nNote enough of time was my biggest problem,but after all this is my first medal,even though its silver, i am really happy.</p>\n<ol>\n<li>Preprocessing RNA sequence to graph:<br>\n1.1. Used Eternafold pkg for extracting secondary structure.<br>\n1.2. OHE nucleotids -&gt; node features<br>\n1.3. Used as edge features -&gt; [phosphodiester_bond,base_pairing(canonical or wobble),BPPS]</li>\n<li>My final solution was:<br>\n2.1. Random walk as positional encoding.<br>\n2.2. Architecture: <br>\nCombination of:<br>\nlocal attention -&gt; GraphTransformer.<br>\nglobal attention -&gt; Attention encoder.<br>\n20 layers depth nn(192 hidden dim), for me expands hidden dimension doesn't help.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10350612%2F29ca6292363042edcf0923d231234396%2Fd.drawio.png?generation=1702059627165899&amp;alt=media\" alt=\"\"></p>\n<p>This works helps me a lot:<br>\n<a href=\"https://arxiv.org/abs/2009.03509\" target=\"_blank\">https://arxiv.org/abs/2009.03509</a><br>\n<a href=\"https://arxiv.org/abs/2205.12454\" target=\"_blank\">https://arxiv.org/abs/2205.12454</a></p>\n<p>3.<br>\nDefault splitting data on train and validation(10 % of SN filter =1) <br>\nDuring train process using rna sequence with signal to noise &gt;= 0.8</p>\n<p>For final submission used only one model prediction no stucks, that might increase quality of prediction as i said before no time was problem.</p>\n<p>github:<a href=\"https://github.com/cerenov94/ribonanzaRNA\" target=\"_blank\">https://github.com/cerenov94/ribonanzaRNA</a><br>\nInstruments:<br>\nPytorch Geometric,Graphein</p>",
  "messages": [
    {
      "id": 2553640,
      "postDate": "2023-12-08T12:12:59.053Z",
      "content": "<p>Hi there,thanks Stanford University for this competition!<br>\nMight be someone will found my work is interesting.<br>\nNote enough of time was my biggest problem,but after all this is my first medal,even though its silver, i am really happy.</p>\n<ol>\n<li>Preprocessing RNA sequence to graph:<br>\n1.1. Used Eternafold pkg for extracting secondary structure.<br>\n1.2. OHE nucleotids -&gt; node features<br>\n1.3. Used as edge features -&gt; [phosphodiester_bond,base_pairing(canonical or wobble),BPPS]</li>\n<li>My final solution was:<br>\n2.1. Random walk as positional encoding.<br>\n2.2. Architecture: <br>\nCombination of:<br>\nlocal attention -&gt; GraphTransformer.<br>\nglobal attention -&gt; Attention encoder.<br>\n20 layers depth nn(192 hidden dim), for me expands hidden dimension doesn't help.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10350612%2F29ca6292363042edcf0923d231234396%2Fd.drawio.png?generation=1702059627165899&amp;alt=media\" alt=\"\"></p>\n<p>This works helps me a lot:<br>\n<a href=\"https://arxiv.org/abs/2009.03509\" target=\"_blank\">https://arxiv.org/abs/2009.03509</a><br>\n<a href=\"https://arxiv.org/abs/2205.12454\" target=\"_blank\">https://arxiv.org/abs/2205.12454</a></p>\n<p>3.<br>\nDefault splitting data on train and validation(10 % of SN filter =1) <br>\nDuring train process using rna sequence with signal to noise &gt;= 0.8</p>\n<p>For final submission used only one model prediction no stucks, that might increase quality of prediction as i said before no time was problem.</p>\n<p>github:<a href=\"https://github.com/cerenov94/ribonanzaRNA\" target=\"_blank\">https://github.com/cerenov94/ribonanzaRNA</a><br>\nInstruments:<br>\nPytorch Geometric,Graphein</p>",
      "rawMarkdown": "Hi there,thanks Stanford University for this competition!\nMight be someone will found my work is interesting.\nNote enough of time was my biggest problem,but after all this is my first medal,even though its silver, i am really happy.\n\n\n1. Preprocessing RNA sequence to graph:\n1.1. Used Eternafold pkg for extracting secondary structure.\n1.2. OHE nucleotids -> node features\n1.3. Used as edge features -> [phosphodiester_bond,base_pairing(canonical or wobble),BPPS]\n2. My final solution was:\n2.1. Random walk as positional encoding.\n2.2. Architecture: \nCombination of:\nlocal attention -> GraphTransformer.\nglobal attention -> Attention encoder.\n20 layers depth nn(192 hidden dim), for me expands hidden dimension doesn't help.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10350612%2F29ca6292363042edcf0923d231234396%2Fd.drawio.png?generation=1702059627165899&alt=media)\n\nThis works helps me a lot:\nhttps://arxiv.org/abs/2009.03509\nhttps://arxiv.org/abs/2205.12454\n\n3.\nDefault splitting data on train and validation(10 % of SN filter =1) \nDuring train process using rna sequence with signal to noise >= 0.8\n\nFor final submission used only one model prediction no stucks, that might increase quality of prediction as i said before no time was problem.\n\n\ngithub:https://github.com/cerenov94/ribonanzaRNA\nInstruments:\nPytorch Geometric,Graphein",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2553640": "Hi there,thanks Stanford University for this competition!\nMight be someone will found my work is interesting.\nNote enough of time was my biggest problem,but after all this is my first medal,even though its silver, i am really happy.\n\n\n1. Preprocessing RNA sequence to graph:\n1.1. Used Eternafold pkg for extracting secondary structure.\n1.2. OHE nucleotids -> node features\n1.3. Used as edge features -> [phosphodiester_bond,base_pairing(canonical or wobble),BPPS]\n2. My final solution was:\n2.1. Random walk as positional encoding.\n2.2. Architecture: \nCombination of:\nlocal attention -> GraphTransformer.\nglobal attention -> Attention encoder.\n20 layers depth nn(192 hidden dim), for me expands hidden dimension doesn't help.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10350612%2F29ca6292363042edcf0923d231234396%2Fd.drawio.png?generation=1702059627165899&alt=media)\n\nThis works helps me a lot:\nhttps://arxiv.org/abs/2009.03509\nhttps://arxiv.org/abs/2205.12454\n\n3.\nDefault splitting data on train and validation(10 % of SN filter =1) \nDuring train process using rna sequence with signal to noise >= 0.8\n\nFor final submission used only one model prediction no stucks, that might increase quality of prediction as i said before no time was problem.\n\n\ngithub:https://github.com/cerenov94/ribonanzaRNA\nInstruments:\nPytorch Geometric,Graphein"
  }
}