{
  "id": 463253,
  "title": "[37th place solution🥈] Single Graph Transformer + Gated GCN Model",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/463253",
  "author_name": "imaginist",
  "post_date": "2023-12-24T08:32:22.043000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>1. Leadboard score</h1>\n<p>Public score: 0.15187; Private Score: 0.15349</p>\n<h1>2. Data preprocess</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Fd294a27c2af60aaccf72f496078dd7b1%2F1.png?generation=1703406681187161&amp;alt=media\" alt=\"\"></p>\n<h2>2.1 Graph node feature</h2>\n<p>input node feature:</p>\n<pre><code>RNA_ONE_HOT = {\n    : [, , , ],\n    : [, , , ],\n    : [, , , ],\n    : [, , , ]\n}\n</code></pre>\n<p>output node feature: DMS_MaP and 2A3_MaP score, which were clipped to [0, 1] when training.</p>\n<h2>2.2 Edge feature</h2>\n<p>For neighbor edge in original RNA sequence, the edge feature = 1.0;<br>\nFor Ribonanza position pairs edge, the edge feature = 1.0 + 10 * Watson-Crick base pair probabilities, which are recorded in the supply Ribonanza_bpp_files.</p>\n<h1>3. Model</h1>\n<p>The model I chose was GraphGPS: github:<a href=\"https://github.com/rampasek/GraphGPS\" target=\"_blank\">https://github.com/rampasek/GraphGPS</a></p>\n<h2>3.1 Basic layer</h2>\n<p>The basic GraphGPS layer is composed of Graph Transformer and Gated GCN.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Ff441e8b5ec70ff3a31f1e3c657f66d0c%2FSnipaste_2023-12-24_15-31-40.png?generation=1703406706736766&amp;alt=media\" alt=\"\"></p>\n<h2>3.2 Positional encoding</h2>\n<p>I tried <strong>Laplacian Positional Encoding</strong> at the begining, but it didn't improve the performance of GraphGPS, which may means the Laplacian Positional Encoding is not useful when facing with topology diversity.<br>\nHence, I removed the positional encoding in GraphGPS.</p>\n<h2>3.3 Model hyper-parameters</h2>\n<pre><code>\n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n</code></pre>\n<h2>3.4 Data augmentation</h2>\n<p>I tried two data augmentation methods: <strong>sub graph sampling</strong> and <strong>RNA sequence reversing</strong>.<br>\nTo be specific, the sub graph sampling is choosing a random center node from original graph, and sampling its k-steps neighbors to generate a new graph; RNA sequence reversing means random reversing the input sequence when training.<br>\nExperimentally, I found the <strong>rna sequence reversing</strong> is more work for this mission.</p>",
  "messages": [
    {
      "id": 2572437,
      "postDate": "2023-12-24T08:32:22.043Z",
      "content": "<h1>1. Leadboard score</h1>\n<p>Public score: 0.15187; Private Score: 0.15349</p>\n<h1>2. Data preprocess</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Fd294a27c2af60aaccf72f496078dd7b1%2F1.png?generation=1703406681187161&amp;alt=media\" alt=\"\"></p>\n<h2>2.1 Graph node feature</h2>\n<p>input node feature:</p>\n<pre><code>RNA_ONE_HOT = {\n    : [, , , ],\n    : [, , , ],\n    : [, , , ],\n    : [, , , ]\n}\n</code></pre>\n<p>output node feature: DMS_MaP and 2A3_MaP score, which were clipped to [0, 1] when training.</p>\n<h2>2.2 Edge feature</h2>\n<p>For neighbor edge in original RNA sequence, the edge feature = 1.0;<br>\nFor Ribonanza position pairs edge, the edge feature = 1.0 + 10 * Watson-Crick base pair probabilities, which are recorded in the supply Ribonanza_bpp_files.</p>\n<h1>3. Model</h1>\n<p>The model I chose was GraphGPS: github:<a href=\"https://github.com/rampasek/GraphGPS\" target=\"_blank\">https://github.com/rampasek/GraphGPS</a></p>\n<h2>3.1 Basic layer</h2>\n<p>The basic GraphGPS layer is composed of Graph Transformer and Gated GCN.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Ff441e8b5ec70ff3a31f1e3c657f66d0c%2FSnipaste_2023-12-24_15-31-40.png?generation=1703406706736766&amp;alt=media\" alt=\"\"></p>\n<h2>3.2 Positional encoding</h2>\n<p>I tried <strong>Laplacian Positional Encoding</strong> at the begining, but it didn't improve the performance of GraphGPS, which may means the Laplacian Positional Encoding is not useful when facing with topology diversity.<br>\nHence, I removed the positional encoding in GraphGPS.</p>\n<h2>3.3 Model hyper-parameters</h2>\n<pre><code>\n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n   \n   \n   \n\n   \n   \n   \n   \n</code></pre>\n<h2>3.4 Data augmentation</h2>\n<p>I tried two data augmentation methods: <strong>sub graph sampling</strong> and <strong>RNA sequence reversing</strong>.<br>\nTo be specific, the sub graph sampling is choosing a random center node from original graph, and sampling its k-steps neighbors to generate a new graph; RNA sequence reversing means random reversing the input sequence when training.<br>\nExperimentally, I found the <strong>rna sequence reversing</strong> is more work for this mission.</p>",
      "rawMarkdown": "# 1. Leadboard score\nPublic score: 0.15187; Private Score: 0.15349\n# 2. Data preprocess\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Fd294a27c2af60aaccf72f496078dd7b1%2F1.png?generation=1703406681187161&alt=media)\n## 2.1 Graph node feature\ninput node feature:\n```python\nRNA_ONE_HOT = {\n    'A': [1, 0, 0, 0],\n    'U': [0, 1, 0, 0],\n    'G': [0, 0, 1, 0],\n    'C': [0, 0, 0, 1]\n}\n```\noutput node feature: DMS_MaP and 2A3_MaP score, which were clipped to [0, 1] when training.\n## 2.2 Edge feature\nFor neighbor edge in original RNA sequence, the edge feature = 1.0;\nFor Ribonanza position pairs edge, the edge feature = 1.0 + 10 * Watson-Crick base pair probabilities, which are recorded in the supply Ribonanza_bpp_files.\n# 3. Model\nThe model I chose was GraphGPS: github:[https://github.com/rampasek/GraphGPS](https://github.com/rampasek/GraphGPS)\n## 3.1 Basic layer\nThe basic GraphGPS layer is composed of Graph Transformer and Gated GCN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Ff441e8b5ec70ff3a31f1e3c657f66d0c%2FSnipaste_2023-12-24_15-31-40.png?generation=1703406706736766&alt=media)\n## 3.2 Positional encoding\nI tried **Laplacian Positional Encoding** at the begining, but it didn't improve the performance of GraphGPS, which may means the Laplacian Positional Encoding is not useful when facing with topology diversity.\nHence, I removed the positional encoding in GraphGPS.\n## 3.3 Model hyper-parameters\n```yaml\nmodel:\n  type: GPSModel\n  loss_fun: l1\ngt:\n  layer_type: CustomGatedGCN+Transformer\n  layers: 16\n  n_heads: 8\n  dim_hidden: 256\n  dropout: 0.1\n  attn_dropout: 0.1\n  layer_norm: False\n  batch_norm: True\ngnn:\n  head: inductive_node\n  layers_pre_mp: 0\n  layers_post_mp: 3\n  dim_inner: 256\n  batchnorm: True\n  act: relu\n  dropout: 0.0\n  agg: mean\n  normalize_adj: False\noptim:\n  clip_grad_norm: True\n  optimizer: adamW\n  weight_decay: 1e-5\n  base_lr: 0.001\n  max_epoch: 50\n  scheduler: cosine_with_warmup\n  num_warmup_epochs: 3\nshare:\n  dim_in: 4\n  dim_out: 2\n  num_splits: 3\n  edge_dim_in: 1\n```\n## 3.4 Data augmentation\nI tried two data augmentation methods: **sub graph sampling** and **RNA sequence reversing**.\nTo be specific, the sub graph sampling is choosing a random center node from original graph, and sampling its k-steps neighbors to generate a new graph; RNA sequence reversing means random reversing the input sequence when training.\nExperimentally, I found the **rna sequence reversing** is more work for this mission.\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2572437": "# 1. Leadboard score\nPublic score: 0.15187; Private Score: 0.15349\n# 2. Data preprocess\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Fd294a27c2af60aaccf72f496078dd7b1%2F1.png?generation=1703406681187161&alt=media)\n## 2.1 Graph node feature\ninput node feature:\n```python\nRNA_ONE_HOT = {\n    'A': [1, 0, 0, 0],\n    'U': [0, 1, 0, 0],\n    'G': [0, 0, 1, 0],\n    'C': [0, 0, 0, 1]\n}\n```\noutput node feature: DMS_MaP and 2A3_MaP score, which were clipped to [0, 1] when training.\n## 2.2 Edge feature\nFor neighbor edge in original RNA sequence, the edge feature = 1.0;\nFor Ribonanza position pairs edge, the edge feature = 1.0 + 10 * Watson-Crick base pair probabilities, which are recorded in the supply Ribonanza_bpp_files.\n# 3. Model\nThe model I chose was GraphGPS: github:[https://github.com/rampasek/GraphGPS](https://github.com/rampasek/GraphGPS)\n## 3.1 Basic layer\nThe basic GraphGPS layer is composed of Graph Transformer and Gated GCN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3160035%2Ff441e8b5ec70ff3a31f1e3c657f66d0c%2FSnipaste_2023-12-24_15-31-40.png?generation=1703406706736766&alt=media)\n## 3.2 Positional encoding\nI tried **Laplacian Positional Encoding** at the begining, but it didn't improve the performance of GraphGPS, which may means the Laplacian Positional Encoding is not useful when facing with topology diversity.\nHence, I removed the positional encoding in GraphGPS.\n## 3.3 Model hyper-parameters\n```yaml\nmodel:\n  type: GPSModel\n  loss_fun: l1\ngt:\n  layer_type: CustomGatedGCN+Transformer\n  layers: 16\n  n_heads: 8\n  dim_hidden: 256\n  dropout: 0.1\n  attn_dropout: 0.1\n  layer_norm: False\n  batch_norm: True\ngnn:\n  head: inductive_node\n  layers_pre_mp: 0\n  layers_post_mp: 3\n  dim_inner: 256\n  batchnorm: True\n  act: relu\n  dropout: 0.0\n  agg: mean\n  normalize_adj: False\noptim:\n  clip_grad_norm: True\n  optimizer: adamW\n  weight_decay: 1e-5\n  base_lr: 0.001\n  max_epoch: 50\n  scheduler: cosine_with_warmup\n  num_warmup_epochs: 3\nshare:\n  dim_in: 4\n  dim_out: 2\n  num_splits: 3\n  edge_dim_in: 1\n```\n## 3.4 Data augmentation\nI tried two data augmentation methods: **sub graph sampling** and **RNA sequence reversing**.\nTo be specific, the sub graph sampling is choosing a random center node from original graph, and sampling its k-steps neighbors to generate a new graph; RNA sequence reversing means random reversing the input sequence when training.\nExperimentally, I found the **rna sequence reversing** is more work for this mission.\n"
  }
}