{
  "id": 443168,
  "title": "🕸️ Graph neural network approaches - with starter notebook 🕸️",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/443168",
  "author_name": "",
  "post_date": "2023-09-25T19:22:31.639942Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, so I'm pretty interested in GNNs, so I threw together a <a href=\"https://www.kaggle.com/code/fnands/a-quick-gnn-baseline\" target=\"_blank\">starter notebook </a> yesterday afternoon that others might find interesting. </p>\n<p>Most of my time was spent on the dataset, and I just plugged in a basic GNN.</p>\n<p>I am not 100% sure if the best approach to this challenge is sequence based or graph based, but there is a pretty good argument to be made that <a href=\"https://thegradient.pub/transformers-are-graph-neural-networks/\" target=\"_blank\">transformers are graph neural networks</a> (if you squint hard enough anyway).  </p>\n<p>I think the \"trick\" for GNNs here will be deciding how to define adjacency. <br>\nIn my notebook, I simply connected nodes to the <code>n</code> nodes to the left and right, but this is overly simplistic. You could just wire it densely (connect every node to every other node, this worked well in a <a href=\"https://www.kaggle.com/code/fnands/1-mpnn\" target=\"_blank\">previous challenge</a>) as is often done in cases where GNNs are used on molecules, although memory usage and compute goes as <code>O(n^2)</code>, so be careful with that. </p>\n<p>A good start might be using the pairing information in the additional data folders, as (as far as I understand) being paired greatly affects the reactivity of a base.   </p>\n<p>If you have any good ideas, I'd be happy to discuss them below. </p>",
  "messages": [
    {
      "id": "2455929",
      "postDate": "09/25/2023 19:22:31",
      "content": "<p>Hi, so I'm pretty interested in GNNs, so I threw together a <a href=\"https://www.kaggle.com/code/fnands/a-quick-gnn-baseline\" target=\"_blank\">starter notebook </a> yesterday afternoon that others might find interesting. </p>\n<p>Most of my time was spent on the dataset, and I just plugged in a basic GNN.</p>\n<p>I am not 100% sure if the best approach to this challenge is sequence based or graph based, but there is a pretty good argument to be made that <a href=\"https://thegradient.pub/transformers-are-graph-neural-networks/\" target=\"_blank\">transformers are graph neural networks</a> (if you squint hard enough anyway).  </p>\n<p>I think the \"trick\" for GNNs here will be deciding how to define adjacency. <br>\nIn my notebook, I simply connected nodes to the <code>n</code> nodes to the left and right, but this is overly simplistic. You could just wire it densely (connect every node to every other node, this worked well in a <a href=\"https://www.kaggle.com/code/fnands/1-mpnn\" target=\"_blank\">previous challenge</a>) as is often done in cases where GNNs are used on molecules, although memory usage and compute goes as <code>O(n^2)</code>, so be careful with that. </p>\n<p>A good start might be using the pairing information in the additional data folders, as (as far as I understand) being paired greatly affects the reactivity of a base.   </p>\n<p>If you have any good ideas, I'd be happy to discuss them below. </p>",
      "rawMarkdown": "Hi, so I'm pretty interested in GNNs, so I threw together a [starter notebook ](https://www.kaggle.com/code/fnands/a-quick-gnn-baseline) yesterday afternoon that others might find interesting. \n\nMost of my time was spent on the dataset, and I just plugged in a basic GNN.\n\nI am not 100% sure if the best approach to this challenge is sequence based or graph based, but there is a pretty good argument to be made that [transformers are graph neural networks](https://thegradient.pub/transformers-are-graph-neural-networks/) (if you squint hard enough anyway).  \n\nI think the \"trick\" for GNNs here will be deciding how to define adjacency. \nIn my notebook, I simply connected nodes to the `n` nodes to the left and right, but this is overly simplistic. You could just wire it densely (connect every node to every other node, this worked well in a [previous challenge](https://www.kaggle.com/code/fnands/1-mpnn)) as is often done in cases where GNNs are used on molecules, although memory usage and compute goes as `O(n^2)`, so be careful with that. \n\nA good start might be using the pairing information in the additional data folders, as (as far as I understand) being paired greatly affects the reactivity of a base.   \n\nIf you have any good ideas, I'd be happy to discuss them below.",
      "votes": null
    },
    {
      "id": "2461277",
      "postDate": "09/29/2023 11:58:53",
      "content": "<p>Hello, I have done a similar thing, also incorporated positional encoding into my next model (am yet to submit a submission for this model). Something you may consider (which I'm working to integrate right now) is stochastically defining some connections, as well as the more structured definition for the neighbors.</p>",
      "rawMarkdown": "Hello, I have done a similar thing, also incorporated positional encoding into my next model (am yet to submit a submission for this model). Something you may consider (which I'm working to integrate right now) is stochastically defining some connections, as well as the more structured definition for the neighbors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2461277,
      "author_name": "marcusbrady",
      "author_url": "",
      "post_date": "09/29/2023 11:58:53",
      "content": "<p>Hello, I have done a similar thing, also incorporated positional encoding into my next model (am yet to submit a submission for this model). Something you may consider (which I'm working to integrate right now) is stochastically defining some connections, as well as the more structured definition for the neighbors.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2455929": "Hi, so I'm pretty interested in GNNs, so I threw together a [starter notebook ](https://www.kaggle.com/code/fnands/a-quick-gnn-baseline) yesterday afternoon that others might find interesting. \n\nMost of my time was spent on the dataset, and I just plugged in a basic GNN.\n\nI am not 100% sure if the best approach to this challenge is sequence based or graph based, but there is a pretty good argument to be made that [transformers are graph neural networks](https://thegradient.pub/transformers-are-graph-neural-networks/) (if you squint hard enough anyway).  \n\nI think the \"trick\" for GNNs here will be deciding how to define adjacency. \nIn my notebook, I simply connected nodes to the `n` nodes to the left and right, but this is overly simplistic. You could just wire it densely (connect every node to every other node, this worked well in a [previous challenge](https://www.kaggle.com/code/fnands/1-mpnn)) as is often done in cases where GNNs are used on molecules, although memory usage and compute goes as `O(n^2)`, so be careful with that. \n\nA good start might be using the pairing information in the additional data folders, as (as far as I understand) being paired greatly affects the reactivity of a base.   \n\nIf you have any good ideas, I'd be happy to discuss them below.",
    "2461277": "Hello, I have done a similar thing, also incorporated positional encoding into my next model (am yet to submit a submission for this model). Something you may consider (which I'm working to integrate right now) is stochastically defining some connections, as well as the more structured definition for the neighbors."
  },
  "source": "meta"
}