{
  "id": 402926,
  "title": "33rd Place Solution - Geometry Aware Transformer",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402926",
  "author_name": "",
  "post_date": "2023-04-20T09:00:53.324453400Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Here's my competition writeup! Unfortunately I didn't make it into my goal of the top 30, but I am happy with the result nonetheless.</p>\n<p>I used a single-model geometry-aware transformer. After seeing the performance of the GraphNet baseline, I tried other point-cloud inspired models. I landed on the geometry-aware transformer (from <a href=\"https://arxiv.org/abs/2108.08839\" target=\"_blank\">this paper</a>) as the backbone for the model.</p>\n<p>The geometry-aware transformer block I used is modified from the PoinTr paper, and uses a radius-graph query (much like the edge-conv blocks in GraphNet), along side multi-head attention. I found that a radius graph performed significantly better than a KNN graph for the edges.<br>\nThe particular structure I got the best performance out of is below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8094969%2F590706064952d89072a4f8e44c0b91c3%2Fgeometry-aware-transformer-diagram.drawio.png?generation=1681980865962692&amp;alt=media\" alt=\"\"></p>\n<p>It is interesting to me to see that a few of the top solutions were also transformers, with some differences and many more blocks: I only used 4 transformer blocks for my final model, with an embedding dimension of 256 - by the end of the competition I had limited time to train this model - I would have liked to try larger models as well. I think in this respect I was also held back by using the edge-conv structure, which seems to not scale as well as only self-attention.</p>",
  "messages": [
    {
      "id": "2228085",
      "postDate": "04/20/2023 09:00:53",
      "content": "<p>Here's my competition writeup! Unfortunately I didn't make it into my goal of the top 30, but I am happy with the result nonetheless.</p>\n<p>I used a single-model geometry-aware transformer. After seeing the performance of the GraphNet baseline, I tried other point-cloud inspired models. I landed on the geometry-aware transformer (from <a href=\"https://arxiv.org/abs/2108.08839\" target=\"_blank\">this paper</a>) as the backbone for the model.</p>\n<p>The geometry-aware transformer block I used is modified from the PoinTr paper, and uses a radius-graph query (much like the edge-conv blocks in GraphNet), along side multi-head attention. I found that a radius graph performed significantly better than a KNN graph for the edges.<br>\nThe particular structure I got the best performance out of is below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8094969%2F590706064952d89072a4f8e44c0b91c3%2Fgeometry-aware-transformer-diagram.drawio.png?generation=1681980865962692&amp;alt=media\" alt=\"\"></p>\n<p>It is interesting to me to see that a few of the top solutions were also transformers, with some differences and many more blocks: I only used 4 transformer blocks for my final model, with an embedding dimension of 256 - by the end of the competition I had limited time to train this model - I would have liked to try larger models as well. I think in this respect I was also held back by using the edge-conv structure, which seems to not scale as well as only self-attention.</p>",
      "rawMarkdown": "Here's my competition writeup! Unfortunately I didn't make it into my goal of the top 30, but I am happy with the result nonetheless.\n\nI used a single-model geometry-aware transformer. After seeing the performance of the GraphNet baseline, I tried other point-cloud inspired models. I landed on the geometry-aware transformer (from [this paper](https://arxiv.org/abs/2108.08839)) as the backbone for the model.\n\nThe geometry-aware transformer block I used is modified from the PoinTr paper, and uses a radius-graph query (much like the edge-conv blocks in GraphNet), along side multi-head attention. I found that a radius graph performed significantly better than a KNN graph for the edges.\nThe particular structure I got the best performance out of is below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8094969%2F590706064952d89072a4f8e44c0b91c3%2Fgeometry-aware-transformer-diagram.drawio.png?generation=1681980865962692&alt=media)\n\nIt is interesting to me to see that a few of the top solutions were also transformers, with some differences and many more blocks: I only used 4 transformer blocks for my final model, with an embedding dimension of 256 - by the end of the competition I had limited time to train this model - I would have liked to try larger models as well. I think in this respect I was also held back by using the edge-conv structure, which seems to not scale as well as only self-attention.",
      "votes": null
    },
    {
      "id": "2228475",
      "postDate": "04/20/2023 15:17:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/sebastianvangerwen\" target=\"_blank\">@sebastianvangerwen</a>! Thanks for sharing your solution; it’s interesting to see this transformer network!</p>",
      "rawMarkdown": "Congrats @sebastianvangerwen! Thanks for sharing your solution; it’s interesting to see this transformer network!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228475,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "04/20/2023 15:17:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/sebastianvangerwen\" target=\"_blank\">@sebastianvangerwen</a>! Thanks for sharing your solution; it’s interesting to see this transformer network!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228085": "Here's my competition writeup! Unfortunately I didn't make it into my goal of the top 30, but I am happy with the result nonetheless.\n\nI used a single-model geometry-aware transformer. After seeing the performance of the GraphNet baseline, I tried other point-cloud inspired models. I landed on the geometry-aware transformer (from [this paper](https://arxiv.org/abs/2108.08839)) as the backbone for the model.\n\nThe geometry-aware transformer block I used is modified from the PoinTr paper, and uses a radius-graph query (much like the edge-conv blocks in GraphNet), along side multi-head attention. I found that a radius graph performed significantly better than a KNN graph for the edges.\nThe particular structure I got the best performance out of is below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8094969%2F590706064952d89072a4f8e44c0b91c3%2Fgeometry-aware-transformer-diagram.drawio.png?generation=1681980865962692&alt=media)\n\nIt is interesting to me to see that a few of the top solutions were also transformers, with some differences and many more blocks: I only used 4 transformer blocks for my final model, with an embedding dimension of 256 - by the end of the competition I had limited time to train this model - I would have liked to try larger models as well. I think in this respect I was also held back by using the edge-conv structure, which seems to not scale as well as only self-attention.",
    "2228475": "Congrats @sebastianvangerwen! Thanks for sharing your solution; it’s interesting to see this transformer network!"
  },
  "source": "meta"
}