{
  "id": 402976,
  "title": "1st Place Solution",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/tito-1st-place-solution",
  "author_name": "",
  "post_date": "2023-04-23T05:20:26.330Z",
  "votes": 77,
  "comment_count": 34,
  "views": 0,
  "content": "<p>First and foremost, I would like to express my gratitude to the hosts and the Kaggle team for organizing such an engaging competition. <br>\nAdditionally, I want to extend my appreciation to those who shared their valuable insights by publishing notebooks and participating in the discussions.<br>\nIn particular, the following notebooks and posts have been incredibly helpful.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">graphnet_baseline_submission</a> by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a></li>\n<li><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">graphnet_example</a> by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> and <a href=\"https://www.kaggle.com/pellerphys\" target=\"_blank\">@pellerphys</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392096\" target=\"_blank\">∞ Explanation and Improvement von Mises-Fisher Loss</a> by <a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a></li>\n<li><a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">Early Sharing Prize - DynEdge - 1.046</a> by <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a></li>\n<li><a href=\"https://www.kaggle.com/code/seungmoklee/3-lstms-with-data-picking-and-shifting\" target=\"_blank\">3 LSTMs; with Data Picking and Shifting</a> by <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a></li>\n<li><a href=\"https://www.kaggle.com/code/dschettler8845/ndi-let-s-learn-together-eli5-and-eda\" target=\"_blank\">NDI – Let's Learn Together – ELI5 and EDA</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">Tensorflow LSTM Model Training TPU</a> by <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a></li>\n</ul>\n<h3>Architecture</h3>\n<p>My model has a simple structure with EdgeConv and Transformer connected.<br>\nI aimed to combine EdgeConv and Transformers in my model to efficiently gather information. GNNs were employed to capture local neighborhood information, while Transformers were used to collect global context, ensuring a comprehensive understanding of the data.<br>\nThis simple model achieves a public score of 0.9628 points and a private score of 0.9633 points as a single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2Ff28760ca31a4cedaebaad8cac80352ac%2Ficecube2.png?generation=1682000618971016&amp;alt=media\" alt=\"model\"></p>\n<h3>Static edge selection for GNN</h3>\n<p>In the <a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">original paper</a>, edges are calculated in each layer dynamically, but this edge selection is not differentiable and cannot be trained. This still works well for the segmentation task in the original paper, as points in the same segment are trained to be close in the latent space. However, the situation is different in this task, so I did not think it would make sense to dynamically select edges. Therefore, the edge selection used in EdgeConv is calculated at the input.</p>\n<h3>Simple modification of EdgeConv</h3>\n<p>Using <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html\" target=\"_blank\">EdgeConv</a>, the latent parameter \\({x}_i\\) of \\({dom}_i\\) is updated using the difference: \\({x}_j-{x}_i\\), where \\({x}_j\\) is the latent parameter of the \\({dom}_i\\)'s k-nearest neighbour \\({dom}_j\\).<br>\nThis is fine for x, y, z and time, but for charge and auxiliary, the absolute values are meaningful, so I make \\({x}_i\\) be updated using not only \\({x}_j-{x}_i\\) but also \\({x}_j\\).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F99153c1b48d05472673c6dcd1827d187%2Fedge4.png?generation=1682093716754213&amp;alt=media\" alt=\"edgeConv\"></p>\n<h3>Loss function</h3>\n<p>VMFLoss is a good and stable loss function, but θ is in the form of cos:</p>\n<pre><code>VMFLoss = -κ*cos(θ) + C(κ)\n</code></pre>\n<p>(where θ is the angle between the ground truth and the prediction and κ is the length of the 3D prediction).<br>\nOn the other hand, the metric of this competition is the angle θ itself.<br>\nTo minimise θ itself, I defined the loss function as follows:</p>\n<pre><code>MyLoss = -θ - κ*cos(θ) + C(κ)\n</code></pre>\n<p>This simple modification resulted in a 0.005👍 gain compared to VMFLoss.</p>\n<h3>Sequence Bucketing</h3>\n<p>In transformers, computational complexity is proportional to the square of the sequence length, making it essential to reduce the sequence length. As a result, it is effective to group data with similar sequence lengths and create mini-batches. This approach speeds up computation and further decreases the GPU memory usage, allowing for the utilization of larger mini-batche size.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F3674d2af7f7552e176a208149f3ba2ec%2Fmini_batch.png?generation=1682128087385204&amp;alt=media\" alt=\"\"></p>\n<p>It is easy to implement by changing collate_fn in the DataLoader.</p>\n<pre><code> ():\n    graphs = [g  g  graphs  g.n_pulses &gt; ]\n    graphs.sort(key= x: x.n_pulses)\n    batch_list = []\n     minp, maxp  ([]+split_list[:-], split_list):\n        min_idx = (minp*(graphs))\n        max_idx = (maxp*(graphs))\n        this_graphs = graphs[min_idx:max_idx]\n        this_batch = Batch.from_data_list(this_graphs)\n        batch_list.append(this_batch)\n     batch_list\n</code></pre>\n<p>Note that even if gradients are calculated individually for each length-biased mini-mini-batch in the batch_list, the negative effects can be mitigated by updating the weights collectively afterwards.</p>\n<h3>Data loading</h3>\n<p>Because the data was huge and it was difficult to put everything in memory, I tried to load one batch per epoch.<br>\nThe pseudo code is shown below.</p>\n<pre><code> ():\n     ():\n        \n        self.this_batch_id = self.get_next_batch_id()\n        \n        self.this_batch = pd.read_parquet()\n        \n        self.this_meta = self.meta[self.meta.batch_id == self.this_batch_id]\n\n     () -&gt; :\n         (self.this_meta)\n\n    ・・・\n\n ():\n     ():\n        ・・・\n        self._dataset = dataset\n\n     ():\n        ・・・\n        \n        \n        self._dataset.reset_batch()\n</code></pre>\n<h3>Model parameters</h3>\n<p>This is some information about base models which are used for stacking. <br>\n4 Layters, 6M parameters seems to be smaller than other top solutions.</p>\n<ul>\n<li>Depth of EdgeConv+Transformer block layer: 3 or 4</li>\n<li>Embedding dimension: 256</li>\n<li>Used sequence length (train): 200 to 500</li>\n<li>Used sequence length (inference): 6000</li>\n<li>With/without global features</li>\n<li>Number of parameters: 6M</li>\n<li>Parameters used in kNN for GNN: x,y,z or x,y,z,time</li>\n<li>Training time: 10 to 14 hours/epoch (on 2x titan RTX)</li>\n<li>Inference time: 30 minutes/5 batches ( on Kaggle notebook)</li>\n</ul>\n<p>Inference time is relatively slow due to the use of 6000 sequences.</p>\n<h3>Stacking</h3>\n<p>The framework for stacking was primarily based on the original base model. <br>\nThe main difference was that the model was replaced with a 3-layer MLP. <br>\nThrough stacking, an improvement of +0.003 was achieved.</p>\n<h3>Code</h3>\n<p>Inference notebook is available <a href=\"https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "2228403",
      "postDate": "04/20/2023 14:14:47",
      "content": "<p>First and foremost, I would like to express my gratitude to the hosts and the Kaggle team for organizing such an engaging competition. <br>\nAdditionally, I want to extend my appreciation to those who shared their valuable insights by publishing notebooks and participating in the discussions.<br>\nIn particular, the following notebooks and posts have been incredibly helpful.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">graphnet_baseline_submission</a> by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a></li>\n<li><a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">graphnet_example</a> by <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> and <a href=\"https://www.kaggle.com/pellerphys\" target=\"_blank\">@pellerphys</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392096\" target=\"_blank\">∞ Explanation and Improvement von Mises-Fisher Loss</a> by <a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a></li>\n<li><a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">Early Sharing Prize - DynEdge - 1.046</a> by <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a></li>\n<li><a href=\"https://www.kaggle.com/code/seungmoklee/3-lstms-with-data-picking-and-shifting\" target=\"_blank\">3 LSTMs; with Data Picking and Shifting</a> by <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a></li>\n<li><a href=\"https://www.kaggle.com/code/dschettler8845/ndi-let-s-learn-together-eli5-and-eda\" target=\"_blank\">NDI – Let's Learn Together – ELI5 and EDA</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">Tensorflow LSTM Model Training TPU</a> by <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a></li>\n</ul>\n<h3>Architecture</h3>\n<p>My model has a simple structure with EdgeConv and Transformer connected.<br>\nI aimed to combine EdgeConv and Transformers in my model to efficiently gather information. GNNs were employed to capture local neighborhood information, while Transformers were used to collect global context, ensuring a comprehensive understanding of the data.<br>\nThis simple model achieves a public score of 0.9628 points and a private score of 0.9633 points as a single model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2Ff28760ca31a4cedaebaad8cac80352ac%2Ficecube2.png?generation=1682000618971016&amp;alt=media\" alt=\"model\"></p>\n<h3>Static edge selection for GNN</h3>\n<p>In the <a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">original paper</a>, edges are calculated in each layer dynamically, but this edge selection is not differentiable and cannot be trained. This still works well for the segmentation task in the original paper, as points in the same segment are trained to be close in the latent space. However, the situation is different in this task, so I did not think it would make sense to dynamically select edges. Therefore, the edge selection used in EdgeConv is calculated at the input.</p>\n<h3>Simple modification of EdgeConv</h3>\n<p>Using <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html\" target=\"_blank\">EdgeConv</a>, the latent parameter \\({x}_i\\) of \\({dom}_i\\) is updated using the difference: \\({x}_j-{x}_i\\), where \\({x}_j\\) is the latent parameter of the \\({dom}_i\\)'s k-nearest neighbour \\({dom}_j\\).<br>\nThis is fine for x, y, z and time, but for charge and auxiliary, the absolute values are meaningful, so I make \\({x}_i\\) be updated using not only \\({x}_j-{x}_i\\) but also \\({x}_j\\).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F99153c1b48d05472673c6dcd1827d187%2Fedge4.png?generation=1682093716754213&amp;alt=media\" alt=\"edgeConv\"></p>\n<h3>Loss function</h3>\n<p>VMFLoss is a good and stable loss function, but θ is in the form of cos:</p>\n<pre><code>VMFLoss = -κ*cos(θ) + C(κ)\n</code></pre>\n<p>(where θ is the angle between the ground truth and the prediction and κ is the length of the 3D prediction).<br>\nOn the other hand, the metric of this competition is the angle θ itself.<br>\nTo minimise θ itself, I defined the loss function as follows:</p>\n<pre><code>MyLoss = -θ - κ*cos(θ) + C(κ)\n</code></pre>\n<p>This simple modification resulted in a 0.005👍 gain compared to VMFLoss.</p>\n<h3>Sequence Bucketing</h3>\n<p>In transformers, computational complexity is proportional to the square of the sequence length, making it essential to reduce the sequence length. As a result, it is effective to group data with similar sequence lengths and create mini-batches. This approach speeds up computation and further decreases the GPU memory usage, allowing for the utilization of larger mini-batche size.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F3674d2af7f7552e176a208149f3ba2ec%2Fmini_batch.png?generation=1682128087385204&amp;alt=media\" alt=\"\"></p>\n<p>It is easy to implement by changing collate_fn in the DataLoader.</p>\n<pre><code> ():\n    graphs = [g  g  graphs  g.n_pulses &gt; ]\n    graphs.sort(key= x: x.n_pulses)\n    batch_list = []\n     minp, maxp  ([]+split_list[:-], split_list):\n        min_idx = (minp*(graphs))\n        max_idx = (maxp*(graphs))\n        this_graphs = graphs[min_idx:max_idx]\n        this_batch = Batch.from_data_list(this_graphs)\n        batch_list.append(this_batch)\n     batch_list\n</code></pre>\n<p>Note that even if gradients are calculated individually for each length-biased mini-mini-batch in the batch_list, the negative effects can be mitigated by updating the weights collectively afterwards.</p>\n<h3>Data loading</h3>\n<p>Because the data was huge and it was difficult to put everything in memory, I tried to load one batch per epoch.<br>\nThe pseudo code is shown below.</p>\n<pre><code> ():\n     ():\n        \n        self.this_batch_id = self.get_next_batch_id()\n        \n        self.this_batch = pd.read_parquet()\n        \n        self.this_meta = self.meta[self.meta.batch_id == self.this_batch_id]\n\n     () -&gt; :\n         (self.this_meta)\n\n    ・・・\n\n ():\n     ():\n        ・・・\n        self._dataset = dataset\n\n     ():\n        ・・・\n        \n        \n        self._dataset.reset_batch()\n</code></pre>\n<h3>Model parameters</h3>\n<p>This is some information about base models which are used for stacking. <br>\n4 Layters, 6M parameters seems to be smaller than other top solutions.</p>\n<ul>\n<li>Depth of EdgeConv+Transformer block layer: 3 or 4</li>\n<li>Embedding dimension: 256</li>\n<li>Used sequence length (train): 200 to 500</li>\n<li>Used sequence length (inference): 6000</li>\n<li>With/without global features</li>\n<li>Number of parameters: 6M</li>\n<li>Parameters used in kNN for GNN: x,y,z or x,y,z,time</li>\n<li>Training time: 10 to 14 hours/epoch (on 2x titan RTX)</li>\n<li>Inference time: 30 minutes/5 batches ( on Kaggle notebook)</li>\n</ul>\n<p>Inference time is relatively slow due to the use of 6000 sequences.</p>\n<h3>Stacking</h3>\n<p>The framework for stacking was primarily based on the original base model. <br>\nThe main difference was that the model was replaced with a 3-layer MLP. <br>\nThrough stacking, an improvement of +0.003 was achieved.</p>\n<h3>Code</h3>\n<p>Inference notebook is available <a href=\"https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "First and foremost, I would like to express my gratitude to the hosts and the Kaggle team for organizing such an engaging competition. \nAdditionally, I want to extend my appreciation to those who shared their valuable insights by publishing notebooks and participating in the discussions.\nIn particular, the following notebooks and posts have been incredibly helpful.\n\n- [graphnet_baseline_submission](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) by @rasmusrse\n- [graphnet_example](https://www.kaggle.com/code/rasmusrse/graphnet-example) by @rasmusrse and @pellerphys\n- [∞ Explanation and Improvement von Mises-Fisher Loss](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392096) by @synset\n- [Early Sharing Prize - DynEdge - 1.046](https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046) by @anjum48\n- [3 LSTMs; with Data Picking and Shifting](https://www.kaggle.com/code/seungmoklee/3-lstms-with-data-picking-and-shifting) by @seungmoklee\n- [NDI – Let's Learn Together – ELI5 and EDA](https://www.kaggle.com/code/dschettler8845/ndi-let-s-learn-together-eli5-and-eda) by @dschettler8845\n- [Tensorflow LSTM Model Training TPU](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) by @rsmits\n\n\n\n### Architecture\nMy model has a simple structure with EdgeConv and Transformer connected.\nI aimed to combine EdgeConv and Transformers in my model to efficiently gather information. GNNs were employed to capture local neighborhood information, while Transformers were used to collect global context, ensuring a comprehensive understanding of the data.\nThis simple model achieves a public score of 0.9628 points and a private score of 0.9633 points as a single model.\n\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2Ff28760ca31a4cedaebaad8cac80352ac%2Ficecube2.png?generation=1682000618971016&alt=media)\n\n### Static edge selection for GNN\nIn the [original paper](https://arxiv.org/abs/1801.07829), edges are calculated in each layer dynamically, but this edge selection is not differentiable and cannot be trained. This still works well for the segmentation task in the original paper, as points in the same segment are trained to be close in the latent space. However, the situation is different in this task, so I did not think it would make sense to dynamically select edges. Therefore, the edge selection used in EdgeConv is calculated at the input.\n\n### Simple modification of EdgeConv\nUsing [EdgeConv](https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html), the latent parameter \\\\({x}_i\\\\) of \\\\({dom}_i\\\\) is updated using the difference: \\\\({x}_j-{x}_i\\\\), where \\\\({x}_j\\\\) is the latent parameter of the \\\\({dom}_i\\\\)'s k-nearest neighbour \\\\({dom}_j\\\\).\nThis is fine for x, y, z and time, but for charge and auxiliary, the absolute values are meaningful, so I make \\\\({x}_i\\\\) be updated using not only \\\\({x}_j-{x}_i\\\\) but also \\\\({x}_j\\\\).\n\n![edgeConv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F99153c1b48d05472673c6dcd1827d187%2Fedge4.png?generation=1682093716754213&alt=media)\n\n### Loss function\nVMFLoss is a good and stable loss function, but θ is in the form of cos:\n```python\nVMFLoss = -κ*cos(θ) + C(κ)\n\n```\n(where θ is the angle between the ground truth and the prediction and κ is the length of the 3D prediction).\nOn the other hand, the metric of this competition is the angle θ itself.\nTo minimise θ itself, I defined the loss function as follows:\n```python\nMyLoss = -θ - κ*cos(θ) + C(κ)\n\n```\nThis simple modification resulted in a 0.005👍 gain compared to VMFLoss.\n\n### Sequence Bucketing\nIn transformers, computational complexity is proportional to the square of the sequence length, making it essential to reduce the sequence length. As a result, it is effective to group data with similar sequence lengths and create mini-batches. This approach speeds up computation and further decreases the GPU memory usage, allowing for the utilization of larger mini-batche size.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F3674d2af7f7552e176a208149f3ba2ec%2Fmini_batch.png?generation=1682128087385204&alt=media)\n\nIt is easy to implement by changing collate_fn in the DataLoader.\n```python\ndef collate_fn(graphs, split_list=[0.8, 1.0]):\n    graphs = [g for g in graphs if g.n_pulses > 1]\n    graphs.sort(key=lambda x: x.n_pulses)\n    batch_list = []\n    for minp, maxp in zip([0]+split_list[:-1], split_list):\n        min_idx = int(minp*len(graphs))\n        max_idx = int(maxp*len(graphs))\n        this_graphs = graphs[min_idx:max_idx]\n        this_batch = Batch.from_data_list(this_graphs)\n        batch_list.append(this_batch)\n    return batch_list\n```\nNote that even if gradients are calculated individually for each length-biased mini-mini-batch in the batch_list, the negative effects can be mitigated by updating the weights collectively afterwards.\n\n### Data loading\nBecause the data was huge and it was difficult to put everything in memory, I tried to load one batch per epoch.\nThe pseudo code is shown below.\n\n```python\nclass IceCubeDataset(Dataset):\n    def reset_batch(self):\n        # get next batch id\n        self.this_batch_id = self.get_next_batch_id()\n        # read batch for next batch id\n        self.this_batch = pd.read_parquet(f\"{BATCH_DIR}/batch_{self.this_batch_id}.parquet\")\n        # filter meta with next batch id\n        self.this_meta = self.meta[self.meta.batch_id == self.this_batch_id]\n        \n    def __len__(self) -> int:\n        return len(self.this_meta)\n    \n    ・・・\n    \nclass IceCubeModel(Model):\n    def __init__(\n        self,\n        dataset,\n        ・・・\n    ):\n        ・・・\n        self._dataset = dataset\n        \n    def training_epoch_end(self, outputs):\n        ・・・\n        # reset batch every epoch\n        # (reload_dataloaders_every_n_epochs must be set to 1)\n        self._dataset.reset_batch()\n\n```   \n\n### Model parameters \nThis is some information about base models which are used for stacking. \n4 Layters, 6M parameters seems to be smaller than other top solutions.\n\n- Depth of EdgeConv+Transformer block layer: 3 or 4\n- Embedding dimension: 256\n- Used sequence length (train): 200 to 500\n- Used sequence length (inference): 6000\n- With/without global features\n- Number of parameters: 6M\n- Parameters used in kNN for GNN: x,y,z or x,y,z,time\n- Training time: 10 to 14 hours/epoch (on 2x titan RTX)\n- Inference time: 30 minutes/5 batches ( on Kaggle notebook)\n\nInference time is relatively slow due to the use of 6000 sequences.\n\n### Stacking\nThe framework for stacking was primarily based on the original base model. \nThe main difference was that the model was replaced with a 3-layer MLP. \nThrough stacking, an improvement of +0.003 was achieved.\n\n### Code \n\nInference notebook is available [here](https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference).",
      "votes": null
    },
    {
      "id": "2228412",
      "postDate": "04/20/2023 14:27:22",
      "content": "<p>Great write-up and congrats on your position <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, very well deserved!</p>",
      "rawMarkdown": "Great write-up and congrats on your position @its7171, very well deserved!",
      "votes": null
    },
    {
      "id": "2228753",
      "postDate": "04/20/2023 19:22:12",
      "content": "<p>Great work and thank you for posting your method. Edge detection is new to me, so a little confused, but looks pretty interesting.</p>",
      "rawMarkdown": "Great work and thank you for posting your method. Edge detection is new to me, so a little confused, but looks pretty interesting.",
      "votes": null
    },
    {
      "id": "2228795",
      "postDate": "04/20/2023 20:26:20",
      "content": "<p>Really impressive work <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> Congratulations with the achieved 1st place.</p>",
      "rawMarkdown": "Really impressive work @its7171 Congratulations with the achieved 1st place.",
      "votes": null
    },
    {
      "id": "2229270",
      "postDate": "04/21/2023 07:55:11",
      "content": "<p>Thank you for participating in this competition and sharing your solution! It's amazing to see that your very straightforward single model alone ranks gold. Looking forward to reading the complete version of this write-up as well as your code if possible! Thanks again.</p>",
      "rawMarkdown": "Thank you for participating in this competition and sharing your solution! It's amazing to see that your very straightforward single model alone ranks gold. Looking forward to reading the complete version of this write-up as well as your code if possible! Thanks again.",
      "votes": null
    },
    {
      "id": "2229622",
      "postDate": "04/21/2023 14:29:09",
      "content": "<p>Thank you for sharing the solution. All of us will take a look at it. You gave an enormous impact! </p>",
      "rawMarkdown": "Thank you for sharing the solution. All of us will take a look at it. You gave an enormous impact!",
      "votes": null
    },
    {
      "id": "2229804",
      "postDate": "04/21/2023 17:39:32",
      "content": "<p>Awesome work <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>! Really love your solution. Well deserved first place and top 15 overall!</p>",
      "rawMarkdown": "Awesome work @its7171! Really love your solution. Well deserved first place and top 15 overall!",
      "votes": null
    },
    {
      "id": "2230121",
      "postDate": "04/22/2023 03:43:11",
      "content": "<p>Thank you! And congratulations to new GM, Well deserved!🎉</p>",
      "rawMarkdown": "Thank you! And congratulations to new GM, Well deserved!🎉",
      "votes": null
    },
    {
      "id": "2230122",
      "postDate": "04/22/2023 03:43:39",
      "content": "<p>Thank you. And thank you for publishing the notebook. It was very helpful.</p>",
      "rawMarkdown": "Thank you. And thank you for publishing the notebook. It was very helpful.",
      "votes": null
    },
    {
      "id": "2230123",
      "postDate": "04/22/2023 03:44:05",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2230125",
      "postDate": "04/22/2023 03:48:28",
      "content": "<p>Thank you!<br>\nThe following documentation on EdgeConv may be helpful.</p>\n<p><a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">https://arxiv.org/abs/1801.07829</a><br>\n<a href=\"https://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf\" target=\"_blank\">https://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf</a><br>\n<a href=\"https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html\" target=\"_blank\">https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html</a></p>",
      "rawMarkdown": "Thank you!\nThe following documentation on EdgeConv may be helpful.\n\nhttps://arxiv.org/abs/1801.07829\nhttps://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf\nhttps://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html",
      "votes": null
    },
    {
      "id": "2230128",
      "postDate": "04/22/2023 03:51:21",
      "content": "<p>Thank you. And thank you for publishing the notebook. It was insightful.</p>",
      "rawMarkdown": "Thank you. And thank you for publishing the notebook. It was insightful.",
      "votes": null
    },
    {
      "id": "2230135",
      "postDate": "04/22/2023 03:53:44",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "2230388",
      "postDate": "04/22/2023 10:24:17",
      "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> would you be willing to share the code of your model part?</p>",
      "rawMarkdown": "its7171 would you be willing to share the code of your model part?",
      "votes": null
    },
    {
      "id": "2230960",
      "postDate": "04/23/2023 00:05:14",
      "content": "<p><a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> I made my inference notebook pulblic.<br>\n<a href=\"https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference\" target=\"_blank\">https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference</a></p>",
      "rawMarkdown": "crodoc I made my inference notebook pulblic.\nhttps://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference",
      "votes": null
    },
    {
      "id": "2230987",
      "postDate": "04/23/2023 01:32:09",
      "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> awesome solution, thanks for sharing it. Congratulations on your results</p>",
      "rawMarkdown": "its7171 awesome solution, thanks for sharing it. Congratulations on your results",
      "votes": null
    },
    {
      "id": "2231341",
      "postDate": "04/23/2023 08:38:50",
      "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>!</p>",
      "rawMarkdown": "Thanks so much @its7171!",
      "votes": null
    },
    {
      "id": "2231992",
      "postDate": "04/23/2023 22:07:00",
      "content": "<p>This is a really great solution, thanks for sharing! And congratulations on getting 1st place!</p>",
      "rawMarkdown": "This is a really great solution, thanks for sharing! And congratulations on getting 1st place!",
      "votes": null
    },
    {
      "id": "2232004",
      "postDate": "04/23/2023 22:26:51",
      "content": "<p>Great work and thank you for posting your method!</p>",
      "rawMarkdown": "Great work and thank you for posting your method!",
      "votes": null
    },
    {
      "id": "2232807",
      "postDate": "04/24/2023 16:34:25",
      "content": "<p>it was nice to go through your solution</p>",
      "rawMarkdown": "it was nice to go through your solution",
      "votes": null
    },
    {
      "id": "2234856",
      "postDate": "04/25/2023 14:30:35",
      "content": "<p>Congratulations on 1st place!<br>\nYour approach to speed up the training is really great!<br>\nI think that speeding up the training was very important in this competition.</p>",
      "rawMarkdown": "Congratulations on 1st place!\nYour approach to speed up the training is really great!\nI think that speeding up the training was very important in this competition.",
      "votes": null
    },
    {
      "id": "2235439",
      "postDate": "04/26/2023 03:55:47",
      "content": "<p>Congratulations. Very nicely explained. </p>",
      "rawMarkdown": "Congratulations. Very nicely explained.",
      "votes": null
    },
    {
      "id": "2236130",
      "postDate": "04/26/2023 14:56:58",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "2236231",
      "postDate": "04/26/2023 16:30:37",
      "content": "<p>Despite the high complexity of the task, the author clearly presented the solution process in the form of a simple flowchart. In addition, he explained how to efficiently use the PC's memory for such a huge amount of data. The article contains links to solutions on individual topics, which makes it very useful and relevant. I wish the author continued success!</p>",
      "rawMarkdown": "Despite the high complexity of the task, the author clearly presented the solution process in the form of a simple flowchart. In addition, he explained how to efficiently use the PC's memory for such a huge amount of data. The article contains links to solutions on individual topics, which makes it very useful and relevant. I wish the author continued success!",
      "votes": null
    },
    {
      "id": "2236434",
      "postDate": "04/26/2023 21:16:09",
      "content": "<p>Thank you for posting your method. Great work.</p>",
      "rawMarkdown": "Thank you for posting your method. Great work.",
      "votes": null
    },
    {
      "id": "2237622",
      "postDate": "04/27/2023 19:38:27",
      "content": "<p>Thanks for sharing!<br>\nCould I know why did you think EdgeConv is suitable here? it is a little new for me, so just want to know when it is best to use it?</p>",
      "rawMarkdown": "Thanks for sharing!\nCould I know why did you think EdgeConv is suitable here? it is a little new for me, so just want to know when it is best to use it?",
      "votes": null
    },
    {
      "id": "2238188",
      "postDate": "04/28/2023 09:20:11",
      "content": "<p>I just thought that it was the state-of-the-art for this task by this paper from hosts.<br>\n<a href=\"https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf\" target=\"_blank\">https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf</a></p>\n<p>It is typically used in Point Clouds segmentation tasks.<br>\n<a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">https://arxiv.org/abs/1801.07829</a></p>\n<p>The data x,y,z and time contained of this competition are only meaningful for their differences.<br>\nIn this respect it is similar to Point Clouds, which I believe is why it was effective in this competition task.</p>",
      "rawMarkdown": "I just thought that it was the state-of-the-art for this task by this paper from hosts.\nhttps://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf\n\nIt is typically used in Point Clouds segmentation tasks.\nhttps://arxiv.org/abs/1801.07829\n\nThe data x,y,z and time contained of this competition are only meaningful for their differences.\nIn this respect it is similar to Point Clouds, which I believe is why it was effective in this competition task.",
      "votes": null
    },
    {
      "id": "2238390",
      "postDate": "04/28/2023 13:26:40",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "2241869",
      "postDate": "05/01/2023 23:04:11",
      "content": "<p>great work\"</p>",
      "rawMarkdown": "great work\"",
      "votes": null
    },
    {
      "id": "2243982",
      "postDate": "05/03/2023 10:56:36",
      "content": "<p>Thanks for the solution and congradulations!!</p>",
      "rawMarkdown": "Thanks for the solution and congradulations!!",
      "votes": null
    },
    {
      "id": "2247334",
      "postDate": "05/05/2023 22:05:37",
      "content": "<p>great solution</p>",
      "rawMarkdown": "great solution",
      "votes": null
    },
    {
      "id": "2247848",
      "postDate": "05/06/2023 10:09:45",
      "content": "<p>Good job, congratulations <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>",
      "rawMarkdown": "Good job, congratulations @its7171",
      "votes": null
    },
    {
      "id": "2253363",
      "postDate": "05/10/2023 05:54:49",
      "content": "<p>Great work!! Love your work!! Thanks for sharing with us👌.</p>",
      "rawMarkdown": "Great work!! Love your work!! Thanks for sharing with us👌.",
      "votes": null
    },
    {
      "id": "2422170",
      "postDate": "09/03/2023 18:39:37",
      "content": "<p>Very nice information! Outstanding work!</p>",
      "rawMarkdown": "Very nice information! Outstanding work!",
      "votes": null
    },
    {
      "id": "2456005",
      "postDate": "09/25/2023 21:31:43",
      "content": "<p>Thank you for the excellent learning materials! </p>",
      "rawMarkdown": "Thank you for the excellent learning materials!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228412,
      "author_name": "hozaifakhalid",
      "author_url": "",
      "post_date": "04/20/2023 14:27:22",
      "content": "<p>Great write-up and congrats on your position <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, very well deserved!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230135,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:53:44",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228753,
      "author_name": "edguy99",
      "author_url": "",
      "post_date": "04/20/2023 19:22:12",
      "content": "<p>Great work and thank you for posting your method. Edge detection is new to me, so a little confused, but looks pretty interesting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230125,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:48:28",
          "content": "<p>Thank you!<br>\nThe following documentation on EdgeConv may be helpful.</p>\n<p><a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">https://arxiv.org/abs/1801.07829</a><br>\n<a href=\"https://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf\" target=\"_blank\">https://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf</a><br>\n<a href=\"https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html\" target=\"_blank\">https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228795,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/20/2023 20:26:20",
      "content": "<p>Really impressive work <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> Congratulations with the achieved 1st place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230128,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:51:21",
          "content": "<p>Thank you. And thank you for publishing the notebook. It was insightful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229270,
      "author_name": "riow1983",
      "author_url": "",
      "post_date": "04/21/2023 07:55:11",
      "content": "<p>Thank you for participating in this competition and sharing your solution! It's amazing to see that your very straightforward single model alone ranks gold. Looking forward to reading the complete version of this write-up as well as your code if possible! Thanks again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230123,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:44:05",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229622,
      "author_name": "seungmoklee",
      "author_url": "",
      "post_date": "04/21/2023 14:29:09",
      "content": "<p>Thank you for sharing the solution. All of us will take a look at it. You gave an enormous impact! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2230122,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:43:39",
          "content": "<p>Thank you. And thank you for publishing the notebook. It was very helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229804,
      "author_name": "crodoc",
      "author_url": "",
      "post_date": "04/21/2023 17:39:32",
      "content": "<p>Awesome work <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>! Really love your solution. Well deserved first place and top 15 overall!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230121,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/22/2023 03:43:11",
          "content": "<p>Thank you! And congratulations to new GM, Well deserved!🎉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2230388,
      "author_name": "crodoc",
      "author_url": "",
      "post_date": "04/22/2023 10:24:17",
      "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> would you be willing to share the code of your model part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230960,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/23/2023 00:05:14",
          "content": "<p><a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> I made my inference notebook pulblic.<br>\n<a href=\"https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference\" target=\"_blank\">https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2231341,
              "author_name": "crodoc",
              "author_url": "",
              "post_date": "04/23/2023 08:38:50",
              "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2230987,
      "author_name": "chrisaqm",
      "author_url": "",
      "post_date": "04/23/2023 01:32:09",
      "content": "<p><a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> awesome solution, thanks for sharing it. Congratulations on your results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2231992,
      "author_name": "nickrod068",
      "author_url": "",
      "post_date": "04/23/2023 22:07:00",
      "content": "<p>This is a really great solution, thanks for sharing! And congratulations on getting 1st place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232004,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "04/23/2023 22:26:51",
      "content": "<p>Great work and thank you for posting your method!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232807,
      "author_name": "ursmaheshj",
      "author_url": "",
      "post_date": "04/24/2023 16:34:25",
      "content": "<p>it was nice to go through your solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234856,
      "author_name": "yamashitamotokazu",
      "author_url": "",
      "post_date": "04/25/2023 14:30:35",
      "content": "<p>Congratulations on 1st place!<br>\nYour approach to speed up the training is really great!<br>\nI think that speeding up the training was very important in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2235439,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "04/26/2023 03:55:47",
      "content": "<p>Congratulations. Very nicely explained. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2236130,
      "author_name": "adithyakolluru",
      "author_url": "",
      "post_date": "04/26/2023 14:56:58",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2236231,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "04/26/2023 16:30:37",
      "content": "<p>Despite the high complexity of the task, the author clearly presented the solution process in the form of a simple flowchart. In addition, he explained how to efficiently use the PC's memory for such a huge amount of data. The article contains links to solutions on individual topics, which makes it very useful and relevant. I wish the author continued success!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2236434,
      "author_name": "maciekgomka",
      "author_url": "",
      "post_date": "04/26/2023 21:16:09",
      "content": "<p>Thank you for posting your method. Great work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2237622,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "04/27/2023 19:38:27",
      "content": "<p>Thanks for sharing!<br>\nCould I know why did you think EdgeConv is suitable here? it is a little new for me, so just want to know when it is best to use it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2238188,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/28/2023 09:20:11",
          "content": "<p>I just thought that it was the state-of-the-art for this task by this paper from hosts.<br>\n<a href=\"https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf\" target=\"_blank\">https://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf</a></p>\n<p>It is typically used in Point Clouds segmentation tasks.<br>\n<a href=\"https://arxiv.org/abs/1801.07829\" target=\"_blank\">https://arxiv.org/abs/1801.07829</a></p>\n<p>The data x,y,z and time contained of this competition are only meaningful for their differences.<br>\nIn this respect it is similar to Point Clouds, which I believe is why it was effective in this competition task.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2238390,
      "author_name": "nicolanicassio",
      "author_url": "",
      "post_date": "04/28/2023 13:26:40",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2241869,
      "author_name": "thomazbarrospires",
      "author_url": "",
      "post_date": "05/01/2023 23:04:11",
      "content": "<p>great work\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2243982,
      "author_name": "hmenezes97",
      "author_url": "",
      "post_date": "05/03/2023 10:56:36",
      "content": "<p>Thanks for the solution and congradulations!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247334,
      "author_name": "asierae",
      "author_url": "",
      "post_date": "05/05/2023 22:05:37",
      "content": "<p>great solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247848,
      "author_name": "szebiniso",
      "author_url": "",
      "post_date": "05/06/2023 10:09:45",
      "content": "<p>Good job, congratulations <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2253363,
      "author_name": "samsamurai",
      "author_url": "",
      "post_date": "05/10/2023 05:54:49",
      "content": "<p>Great work!! Love your work!! Thanks for sharing with us👌.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2422170,
      "author_name": "sathyanarayanrao89",
      "author_url": "",
      "post_date": "09/03/2023 18:39:37",
      "content": "<p>Very nice information! Outstanding work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2456005,
      "author_name": "jhynes",
      "author_url": "",
      "post_date": "09/25/2023 21:31:43",
      "content": "<p>Thank you for the excellent learning materials! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228403": "First and foremost, I would like to express my gratitude to the hosts and the Kaggle team for organizing such an engaging competition. \nAdditionally, I want to extend my appreciation to those who shared their valuable insights by publishing notebooks and participating in the discussions.\nIn particular, the following notebooks and posts have been incredibly helpful.\n\n- [graphnet_baseline_submission](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) by @rasmusrse\n- [graphnet_example](https://www.kaggle.com/code/rasmusrse/graphnet-example) by @rasmusrse and @pellerphys\n- [∞ Explanation and Improvement von Mises-Fisher Loss](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392096) by @synset\n- [Early Sharing Prize - DynEdge - 1.046](https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046) by @anjum48\n- [3 LSTMs; with Data Picking and Shifting](https://www.kaggle.com/code/seungmoklee/3-lstms-with-data-picking-and-shifting) by @seungmoklee\n- [NDI – Let's Learn Together – ELI5 and EDA](https://www.kaggle.com/code/dschettler8845/ndi-let-s-learn-together-eli5-and-eda) by @dschettler8845\n- [Tensorflow LSTM Model Training TPU](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) by @rsmits\n\n\n\n### Architecture\nMy model has a simple structure with EdgeConv and Transformer connected.\nI aimed to combine EdgeConv and Transformers in my model to efficiently gather information. GNNs were employed to capture local neighborhood information, while Transformers were used to collect global context, ensuring a comprehensive understanding of the data.\nThis simple model achieves a public score of 0.9628 points and a private score of 0.9633 points as a single model.\n\n![model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2Ff28760ca31a4cedaebaad8cac80352ac%2Ficecube2.png?generation=1682000618971016&alt=media)\n\n### Static edge selection for GNN\nIn the [original paper](https://arxiv.org/abs/1801.07829), edges are calculated in each layer dynamically, but this edge selection is not differentiable and cannot be trained. This still works well for the segmentation task in the original paper, as points in the same segment are trained to be close in the latent space. However, the situation is different in this task, so I did not think it would make sense to dynamically select edges. Therefore, the edge selection used in EdgeConv is calculated at the input.\n\n### Simple modification of EdgeConv\nUsing [EdgeConv](https://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html), the latent parameter \\\\({x}_i\\\\) of \\\\({dom}_i\\\\) is updated using the difference: \\\\({x}_j-{x}_i\\\\), where \\\\({x}_j\\\\) is the latent parameter of the \\\\({dom}_i\\\\)'s k-nearest neighbour \\\\({dom}_j\\\\).\nThis is fine for x, y, z and time, but for charge and auxiliary, the absolute values are meaningful, so I make \\\\({x}_i\\\\) be updated using not only \\\\({x}_j-{x}_i\\\\) but also \\\\({x}_j\\\\).\n\n![edgeConv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F99153c1b48d05472673c6dcd1827d187%2Fedge4.png?generation=1682093716754213&alt=media)\n\n### Loss function\nVMFLoss is a good and stable loss function, but θ is in the form of cos:\n```python\nVMFLoss = -κ*cos(θ) + C(κ)\n\n```\n(where θ is the angle between the ground truth and the prediction and κ is the length of the 3D prediction).\nOn the other hand, the metric of this competition is the angle θ itself.\nTo minimise θ itself, I defined the loss function as follows:\n```python\nMyLoss = -θ - κ*cos(θ) + C(κ)\n\n```\nThis simple modification resulted in a 0.005👍 gain compared to VMFLoss.\n\n### Sequence Bucketing\nIn transformers, computational complexity is proportional to the square of the sequence length, making it essential to reduce the sequence length. As a result, it is effective to group data with similar sequence lengths and create mini-batches. This approach speeds up computation and further decreases the GPU memory usage, allowing for the utilization of larger mini-batche size.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F548996%2F3674d2af7f7552e176a208149f3ba2ec%2Fmini_batch.png?generation=1682128087385204&alt=media)\n\nIt is easy to implement by changing collate_fn in the DataLoader.\n```python\ndef collate_fn(graphs, split_list=[0.8, 1.0]):\n    graphs = [g for g in graphs if g.n_pulses > 1]\n    graphs.sort(key=lambda x: x.n_pulses)\n    batch_list = []\n    for minp, maxp in zip([0]+split_list[:-1], split_list):\n        min_idx = int(minp*len(graphs))\n        max_idx = int(maxp*len(graphs))\n        this_graphs = graphs[min_idx:max_idx]\n        this_batch = Batch.from_data_list(this_graphs)\n        batch_list.append(this_batch)\n    return batch_list\n```\nNote that even if gradients are calculated individually for each length-biased mini-mini-batch in the batch_list, the negative effects can be mitigated by updating the weights collectively afterwards.\n\n### Data loading\nBecause the data was huge and it was difficult to put everything in memory, I tried to load one batch per epoch.\nThe pseudo code is shown below.\n\n```python\nclass IceCubeDataset(Dataset):\n    def reset_batch(self):\n        # get next batch id\n        self.this_batch_id = self.get_next_batch_id()\n        # read batch for next batch id\n        self.this_batch = pd.read_parquet(f\"{BATCH_DIR}/batch_{self.this_batch_id}.parquet\")\n        # filter meta with next batch id\n        self.this_meta = self.meta[self.meta.batch_id == self.this_batch_id]\n        \n    def __len__(self) -> int:\n        return len(self.this_meta)\n    \n    ・・・\n    \nclass IceCubeModel(Model):\n    def __init__(\n        self,\n        dataset,\n        ・・・\n    ):\n        ・・・\n        self._dataset = dataset\n        \n    def training_epoch_end(self, outputs):\n        ・・・\n        # reset batch every epoch\n        # (reload_dataloaders_every_n_epochs must be set to 1)\n        self._dataset.reset_batch()\n\n```   \n\n### Model parameters \nThis is some information about base models which are used for stacking. \n4 Layters, 6M parameters seems to be smaller than other top solutions.\n\n- Depth of EdgeConv+Transformer block layer: 3 or 4\n- Embedding dimension: 256\n- Used sequence length (train): 200 to 500\n- Used sequence length (inference): 6000\n- With/without global features\n- Number of parameters: 6M\n- Parameters used in kNN for GNN: x,y,z or x,y,z,time\n- Training time: 10 to 14 hours/epoch (on 2x titan RTX)\n- Inference time: 30 minutes/5 batches ( on Kaggle notebook)\n\nInference time is relatively slow due to the use of 6000 sequences.\n\n### Stacking\nThe framework for stacking was primarily based on the original base model. \nThe main difference was that the model was replaced with a 3-layer MLP. \nThrough stacking, an improvement of +0.003 was achieved.\n\n### Code \n\nInference notebook is available [here](https://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference).",
    "2228412": "Great write-up and congrats on your position @its7171, very well deserved!",
    "2228753": "Great work and thank you for posting your method. Edge detection is new to me, so a little confused, but looks pretty interesting.",
    "2228795": "Really impressive work @its7171 Congratulations with the achieved 1st place.",
    "2229270": "Thank you for participating in this competition and sharing your solution! It's amazing to see that your very straightforward single model alone ranks gold. Looking forward to reading the complete version of this write-up as well as your code if possible! Thanks again.",
    "2229622": "Thank you for sharing the solution. All of us will take a look at it. You gave an enormous impact!",
    "2229804": "Awesome work @its7171! Really love your solution. Well deserved first place and top 15 overall!",
    "2230121": "Thank you! And congratulations to new GM, Well deserved!🎉",
    "2230122": "Thank you. And thank you for publishing the notebook. It was very helpful.",
    "2230123": "Thank you!",
    "2230125": "Thank you!\nThe following documentation on EdgeConv may be helpful.\n\nhttps://arxiv.org/abs/1801.07829\nhttps://www.nbi.dk/~petersen/Teaching/ML2022/Week5/IntroductionToGNNs_RasmusO.pdf\nhttps://pytorch-geometric.readthedocs.io/en/latest/generated/torch_geometric.nn.conv.EdgeConv.html",
    "2230128": "Thank you. And thank you for publishing the notebook. It was insightful.",
    "2230135": "Thank you very much!",
    "2230388": "its7171 would you be willing to share the code of your model part?",
    "2230960": "crodoc I made my inference notebook pulblic.\nhttps://www.kaggle.com/code/its7171/icecube-edgeconv-transformer-inference",
    "2230987": "its7171 awesome solution, thanks for sharing it. Congratulations on your results",
    "2231341": "Thanks so much @its7171!",
    "2231992": "This is a really great solution, thanks for sharing! And congratulations on getting 1st place!",
    "2232004": "Great work and thank you for posting your method!",
    "2232807": "it was nice to go through your solution",
    "2234856": "Congratulations on 1st place!\nYour approach to speed up the training is really great!\nI think that speeding up the training was very important in this competition.",
    "2235439": "Congratulations. Very nicely explained.",
    "2236130": "Congratulations!",
    "2236231": "Despite the high complexity of the task, the author clearly presented the solution process in the form of a simple flowchart. In addition, he explained how to efficiently use the PC's memory for such a huge amount of data. The article contains links to solutions on individual topics, which makes it very useful and relevant. I wish the author continued success!",
    "2236434": "Thank you for posting your method. Great work.",
    "2237622": "Thanks for sharing!\nCould I know why did you think EdgeConv is suitable here? it is a little new for me, so just want to know when it is best to use it?",
    "2238188": "I just thought that it was the state-of-the-art for this task by this paper from hosts.\nhttps://iopscience.iop.org/article/10.1088/1748-0221/17/11/P11003/pdf\n\nIt is typically used in Point Clouds segmentation tasks.\nhttps://arxiv.org/abs/1801.07829\n\nThe data x,y,z and time contained of this competition are only meaningful for their differences.\nIn this respect it is similar to Point Clouds, which I believe is why it was effective in this competition task.",
    "2238390": "Congratulations!",
    "2241869": "great work\"",
    "2243982": "Thanks for the solution and congradulations!!",
    "2247334": "great solution",
    "2247848": "Good job, congratulations @its7171",
    "2253363": "Great work!! Love your work!! Thanks for sharing with us👌.",
    "2422170": "Very nice information! Outstanding work!",
    "2456005": "Thank you for the excellent learning materials!"
  },
  "source": "meta"
}