{
  "id": 402849,
  "title": "9th Place Solution: GNNs Ensemble and MLP Stacking",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/hearts-of-ice-9th-place-solution-gnns-ensemble-and",
  "author_name": "",
  "post_date": "2024-10-06T13:54:07.760Z",
  "votes": 41,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, we want to thank Kaggle and IceCube Collaboration for providing this spectacular competition!</p>\n<p>Our solution is based on an MLP stack with shared weights of 6 DynEdge models + a blend of 4 GNNs inside the segments by pulse number.</p>\n<h2>1. Model Zoo</h2>\n<p>We took the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">GraphNeT</a> model with DynEdge architecture as a baseline and modified it in different ways. Main variations and their nicknames:</p>\n<ul>\n<li>v1 <strong>Vanilla</strong>. Changed the KNN value and included additional pooling schemes. Replaced Adam optimizer with <a href=\"https://arxiv.org/abs/1810.06801\" target=\"_blank\">QHAdam</a>.</li>\n<li>v2 <strong>Bigman</strong>. Added more layers and made them wider. Changed activation functions to <a href=\"https://arxiv.org/abs/1702.03118\" target=\"_blank\">SiLU</a> (they proved to be very powerful in various domains). Added Batch Normalization layers.</li>\n<li>v3 <strong>DynEmb</strong>. To account for ice heterogeneity, we integrated sensor embeddings into the model (embeddings are initialized by XYZ coordinates). This approach can be generalized and used with any type of Observatory that has sensors distributed in ice or water. We also added Weight Normalization and switched to FP16 training with the new <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> optimizer. For this model, we published a submission <a href=\"https://www.kaggle.com/code/ukrainskiydv/icecube-submission-dynemb-example/notebook\" target=\"_blank\">example</a> that gives the top 20 LB in solo for 1 hour 20 minutes.</li>\n</ul>\n<h2>2. Training Protocol</h2>\n<p>Every model had longrun training from scratch with von Mises-Fisher Loss on events up to 600 pulses. Then a couple of additional epochs were trained with Euclidean Distance Loss and up to 200 pulses. As <a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> noted below, this loss works much better than the original one for second-stage training. The details on the optimizer choice and second-stage training strategy are presented by <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a> in the comment section.</p>\n<h2>3. Ensembling</h2>\n<p>Finally, SWA was applied to the best checkpoints of each model for better robustness. As a result, we had v1, v2, v3 models with performance &lt; 1.000 LB and v1, v2, v3 models with performance &lt; 0.990 LB. These 6 models were blended and stacked. Stack showed better performance ~ 0.977 LB in the very final version, <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a> added the details below.</p>\n<h2>4. Alternative Approach</h2>\n<p>Another part of our solution consists of several GNN models for different neutrino events. The dataset was segmented by ten uniformly bins on pulse number (that represents energy levels of neutrino), and 4 new models were trained on their compositions. Eventually, the models were blended segment by segment, <a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> added the details on this part of the solution in the post <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402995\" target=\"_blank\">here</a>.</p>\n<h2>5. Results</h2>\n<p>The final model is a 60 / 40 blend of 6-stacked / 4-segmented GNNs and shows 0.975 LB (0.976 Private). We used a blend of stacked models to improve the overall stability of the ensemble and reduce the variance of individual models. We also hope this approach leads to better resilience of the model to possible noise in the data.</p>\n<p>For sure, Transformer-based models have great performance on the task, but a diversified ensemble of properly tuned GNNs can also show a keen score on the considered reconstruction problem.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743648%2F829df4033d8353d91b74f0559259506c%2FModels%20Zoo%20RESIZED.png?generation=1682034404975611&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2227682",
      "postDate": "04/20/2023 00:07:28",
      "content": "<p>First of all, we want to thank Kaggle and IceCube Collaboration for providing this spectacular competition!</p>\n<p>Our solution is based on an MLP stack with shared weights of 6 DynEdge models + a blend of 4 GNNs inside the segments by pulse number.</p>\n<h2>1. Model Zoo</h2>\n<p>We took the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">GraphNeT</a> model with DynEdge architecture as a baseline and modified it in different ways. Main variations and their nicknames:</p>\n<ul>\n<li>v1 <strong>Vanilla</strong>. Changed the KNN value and included additional pooling schemes. Replaced Adam optimizer with <a href=\"https://arxiv.org/abs/1810.06801\" target=\"_blank\">QHAdam</a>.</li>\n<li>v2 <strong>Bigman</strong>. Added more layers and made them wider. Changed activation functions to <a href=\"https://arxiv.org/abs/1702.03118\" target=\"_blank\">SiLU</a> (they proved to be very powerful in various domains). Added Batch Normalization layers.</li>\n<li>v3 <strong>DynEmb</strong>. To account for ice heterogeneity, we integrated sensor embeddings into the model (embeddings are initialized by XYZ coordinates). This approach can be generalized and used with any type of Observatory that has sensors distributed in ice or water. We also added Weight Normalization and switched to FP16 training with the new <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> optimizer. For this model, we published a submission <a href=\"https://www.kaggle.com/code/ukrainskiydv/icecube-submission-dynemb-example/notebook\" target=\"_blank\">example</a> that gives the top 20 LB in solo for 1 hour 20 minutes.</li>\n</ul>\n<h2>2. Training Protocol</h2>\n<p>Every model had longrun training from scratch with von Mises-Fisher Loss on events up to 600 pulses. Then a couple of additional epochs were trained with Euclidean Distance Loss and up to 200 pulses. As <a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> noted below, this loss works much better than the original one for second-stage training. The details on the optimizer choice and second-stage training strategy are presented by <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a> in the comment section.</p>\n<h2>3. Ensembling</h2>\n<p>Finally, SWA was applied to the best checkpoints of each model for better robustness. As a result, we had v1, v2, v3 models with performance &lt; 1.000 LB and v1, v2, v3 models with performance &lt; 0.990 LB. These 6 models were blended and stacked. Stack showed better performance ~ 0.977 LB in the very final version, <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a> added the details below.</p>\n<h2>4. Alternative Approach</h2>\n<p>Another part of our solution consists of several GNN models for different neutrino events. The dataset was segmented by ten uniformly bins on pulse number (that represents energy levels of neutrino), and 4 new models were trained on their compositions. Eventually, the models were blended segment by segment, <a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> added the details on this part of the solution in the post <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402995\" target=\"_blank\">here</a>.</p>\n<h2>5. Results</h2>\n<p>The final model is a 60 / 40 blend of 6-stacked / 4-segmented GNNs and shows 0.975 LB (0.976 Private). We used a blend of stacked models to improve the overall stability of the ensemble and reduce the variance of individual models. We also hope this approach leads to better resilience of the model to possible noise in the data.</p>\n<p>For sure, Transformer-based models have great performance on the task, but a diversified ensemble of properly tuned GNNs can also show a keen score on the considered reconstruction problem.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743648%2F829df4033d8353d91b74f0559259506c%2FModels%20Zoo%20RESIZED.png?generation=1682034404975611&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, we want to thank Kaggle and IceCube Collaboration for providing this spectacular competition!\n\nOur solution is based on an MLP stack with shared weights of 6 DynEdge models + a blend of 4 GNNs inside the segments by pulse number.\n\n## 1. Model Zoo\n\nWe took the [GraphNeT](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524) model with DynEdge architecture as a baseline and modified it in different ways. Main variations and their nicknames:\n- v1 **Vanilla**. Changed the KNN value and included additional pooling schemes. Replaced Adam optimizer with [QHAdam](https://arxiv.org/abs/1810.06801).\n- v2 **Bigman**. Added more layers and made them wider. Changed activation functions to [SiLU](https://arxiv.org/abs/1702.03118) (they proved to be very powerful in various domains). Added Batch Normalization layers.\n- v3 **DynEmb**. To account for ice heterogeneity, we integrated sensor embeddings into the model (embeddings are initialized by XYZ coordinates). This approach can be generalized and used with any type of Observatory that has sensors distributed in ice or water. We also added Weight Normalization and switched to FP16 training with the new [Lion](https://arxiv.org/abs/2302.06675) optimizer. For this model, we published a submission [example](https://www.kaggle.com/code/ukrainskiydv/icecube-submission-dynemb-example/notebook) that gives the top 20 LB in solo for 1 hour 20 minutes.\n\n## 2. Training Protocol\n\nEvery model had longrun training from scratch with von Mises-Fisher Loss on events up to 600 pulses. Then a couple of additional epochs were trained with Euclidean Distance Loss and up to 200 pulses. As @rusg77 noted below, this loss works much better than the original one for second-stage training. The details on the optimizer choice and second-stage training strategy are presented by @churkinnikita in the comment section.\n\n## 3. Ensembling\n\nFinally, SWA was applied to the best checkpoints of each model for better robustness. As a result, we had v1, v2, v3 models with performance < 1.000 LB and v1, v2, v3 models with performance < 0.990 LB. These 6 models were blended and stacked. Stack showed better performance ~ 0.977 LB in the very final version, @simakov added the details below.\n\n## 4. Alternative Approach\n\nAnother part of our solution consists of several GNN models for different neutrino events. The dataset was segmented by ten uniformly bins on pulse number (that represents energy levels of neutrino), and 4 new models were trained on their compositions. Eventually, the models were blended segment by segment, @rusg77 added the details on this part of the solution in the post [here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402995).\n\n## 5. Results \n\nThe final model is a 60 / 40 blend of 6-stacked / 4-segmented GNNs and shows 0.975 LB (0.976 Private). We used a blend of stacked models to improve the overall stability of the ensemble and reduce the variance of individual models. We also hope this approach leads to better resilience of the model to possible noise in the data.\n\nFor sure, Transformer-based models have great performance on the task, but a diversified ensemble of properly tuned GNNs can also show a keen score on the considered reconstruction problem.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743648%2F829df4033d8353d91b74f0559259506c%2FModels%20Zoo%20RESIZED.png?generation=1682034404975611&alt=media)",
      "votes": null
    },
    {
      "id": "2227685",
      "postDate": "04/20/2023 00:13:56",
      "content": "<p>Stack Model:</p>\n<pre><code> (nn.Module):\n     ():\n        (MLPShared, self).__init__()\n\n        self.fc1 = nn.Linear(, ) \n        self.fc2 = nn.Linear(+, )\n        self.fc3 = nn.Linear(+, )\n\n        self.act = nn.SiLU()\n\n        self.bn1 = nn.BatchNorm1d()\n        self.bn2 = nn.BatchNorm1d()\n\n     ():\n        x, meta = x[:, :-], x[:, -:]\n        meta = torch.repeat_interleave(meta, , dim=)\n        meta = meta.view(-, , )\n\n        x = x.view(-, , )\n        x = torch.cat([x, meta], dim=)\n        x = torch.permute(x, (, , ))\n        x1 = self.fc1(x)\n        x1 = torch.permute(x1, (, , ))\n        x1 = self.bn1(x1)\n        x1 = torch.permute(x1, (, , ))\n        x1 = self.act(x1)\n\n        \n        x1 = torch.cat((x, x1), dim=-) \n        \n        x2 = self.fc2(x1)\n        x2 = torch.permute(x2, (, , ))\n        x2 = self.bn2(x2)\n        x2 = torch.permute(x2, (, , ))\n        x2 = self.act(x2)\n        \n        x2 = torch.cat((x, x2), dim=-)\n        \n\n        x3 = self.fc3(x2)\n        x3 = x3.view(-, )\n         x3\n</code></pre>",
      "rawMarkdown": "Stack Model:\n\n```python\nclass MLPShared(nn.Module):\n    def __init__(self,):\n        super(MLPShared, self).__init__()\n\n        self.fc1 = nn.Linear(11, 32) # 6 models + 5 meta features \n        self.fc2 = nn.Linear(32+11, 16)\n        self.fc3 = nn.Linear(16+11, 1)\n\n        self.act = nn.SiLU()\n\n        self.bn1 = nn.BatchNorm1d(32)\n        self.bn2 = nn.BatchNorm1d(16)\n\n    def forward(self, x):\n        x, meta = x[:, :-5], x[:, -5:]\n        meta = torch.repeat_interleave(meta, 3, dim=1)\n        meta = meta.view(-1, 5, 3)\n\n        x = x.view(-1, 6, 3)\n        x = torch.cat([x, meta], dim=1)\n        x = torch.permute(x, (0, 2, 1))\n        x1 = self.fc1(x)\n        x1 = torch.permute(x1, (0, 2, 1))\n        x1 = self.bn1(x1)\n        x1 = torch.permute(x1, (0, 2, 1))\n        x1 = self.act(x1)\n\n        #\n        x1 = torch.cat((x, x1), dim=-1) # dence-light connection\n        #\n        x2 = self.fc2(x1)\n        x2 = torch.permute(x2, (0, 2, 1))\n        x2 = self.bn2(x2)\n        x2 = torch.permute(x2, (0, 2, 1))\n        x2 = self.act(x2)\n        #\n        x2 = torch.cat((x, x2), dim=-1)\n        #\n\n        x3 = self.fc3(x2)\n        x3 = x3.view(-1, 3)\n        return x3\n```",
      "votes": null
    },
    {
      "id": "2227688",
      "postDate": "04/20/2023 00:22:36",
      "content": "<p><code>EuclideanDistanceLoss</code> shows better performance for pretrained models </p>",
      "rawMarkdown": "`EuclideanDistanceLoss` shows better performance for pretrained models",
      "votes": null
    },
    {
      "id": "2228474",
      "postDate": "04/20/2023 15:17:11",
      "content": "<p>Congratulations on your gold medals! Solid purely-GNN score.</p>",
      "rawMarkdown": "Congratulations on your gold medals! Solid purely-GNN score.",
      "votes": null
    },
    {
      "id": "2228618",
      "postDate": "04/20/2023 16:49:24",
      "content": "<p><strong>GNN Training Process.</strong></p>\n<p>Inital loss is von Mises-Fisher 3D Loss. For initial network learning we used <a href=\"https://arxiv.org/abs/1810.06801\" target=\"_blank\">QHAdam</a>, <a href=\"https://arxiv.org/abs/1711.05101\" target=\"_blank\">AdamW</a> and <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> optimizers.</p>\n<p>I have already touched <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> in my <a href=\"https://www.kaggle.com/competitions/learning-equality-curriculum-recommendations/discussion/394827\" target=\"_blank\">previous competition</a>. Learning process often gets whimsical when I try to train using Lion: usually extremely fast convergence to low loss values (but not enough low) followed by bizzare behaviour or complete divergence. But after some tuning and tinkering (<a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">there</a> are some tips for working with Lion) it was possible to get nice checkpoint.</p>\n<p>We continued training our models using L2 loss and different training protocol.<br>\nOur hypothesis on learning our GNNs: loss surface often looked like <a href=\"https://www.sfu.ca/~ssurjano/rastr.html\" target=\"_blank\">Rastrigin function</a>, sophisticated surface with dozens of local minima.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F578229%2F7f4f0049a0ec841238a88275ead39f8f%2Frastr.png?generation=1682008443614288&amp;alt=media\" alt=\"\"><br>\nSource: <a href=\"https://www.sfu.ca/~ssurjano/rastr.html\" target=\"_blank\">https://www.sfu.ca/~ssurjano/rastr.html</a><br>\nLion optimizer led to extremely bad results on that stage; we grid searched through optimizers from <a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">pytorch optimizer lib</a> that have the very best performance in the \"Rastrigin scenario\". In our setup <a href=\"https://openreview.net/forum?id=Bkg3g2R9FX\" target=\"_blank\">AdaBound</a> &amp; <a href=\"https://arxiv.org/abs/1908.03265\" target=\"_blank\">RAdam</a> provided splendid performance.</p>\n<p>We used short warmup (5%), non-default betas for optimizers and lower learning rate during 2nd stage training.<br>\nFor my models (\"DynEmb\" variations) only 30m-1h of training time was enough to receive impressive loss drop.</p>",
      "rawMarkdown": "**GNN Training Process.**\n\nInital loss is von Mises-Fisher 3D Loss. For initial network learning we used [QHAdam](https://arxiv.org/abs/1810.06801), [AdamW](https://arxiv.org/abs/1711.05101) and [Lion](https://arxiv.org/abs/2302.06675) optimizers.\n\nI have already touched [Lion](https://arxiv.org/abs/2302.06675) in my [previous competition](https://www.kaggle.com/competitions/learning-equality-curriculum-recommendations/discussion/394827). Learning process often gets whimsical when I try to train using Lion: usually extremely fast convergence to low loss values (but not enough low) followed by bizzare behaviour or complete divergence. But after some tuning and tinkering ([there](https://github.com/lucidrains/lion-pytorch) are some tips for working with Lion) it was possible to get nice checkpoint.\n\nWe continued training our models using L2 loss and different training protocol.\nOur hypothesis on learning our GNNs: loss surface often looked like [Rastrigin function](https://www.sfu.ca/~ssurjano/rastr.html), sophisticated surface with dozens of local minima.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F578229%2F7f4f0049a0ec841238a88275ead39f8f%2Frastr.png?generation=1682008443614288&alt=media)\nSource: https://www.sfu.ca/~ssurjano/rastr.html\nLion optimizer led to extremely bad results on that stage; we grid searched through optimizers from [pytorch optimizer lib](https://github.com/jettify/pytorch-optimizer) that have the very best performance in the \"Rastrigin scenario\". In our setup [AdaBound](https://openreview.net/forum?id=Bkg3g2R9FX) & [RAdam](https://arxiv.org/abs/1908.03265) provided splendid performance.\n\nWe used short warmup (5%), non-default betas for optimizers and lower learning rate during 2nd stage training.\nFor my models (\"DynEmb\" variations) only 30m-1h of training time was enough to receive impressive loss drop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227685,
      "author_name": "simakov",
      "author_url": "",
      "post_date": "04/20/2023 00:13:56",
      "content": "<p>Stack Model:</p>\n<pre><code> (nn.Module):\n     ():\n        (MLPShared, self).__init__()\n\n        self.fc1 = nn.Linear(, ) \n        self.fc2 = nn.Linear(+, )\n        self.fc3 = nn.Linear(+, )\n\n        self.act = nn.SiLU()\n\n        self.bn1 = nn.BatchNorm1d()\n        self.bn2 = nn.BatchNorm1d()\n\n     ():\n        x, meta = x[:, :-], x[:, -:]\n        meta = torch.repeat_interleave(meta, , dim=)\n        meta = meta.view(-, , )\n\n        x = x.view(-, , )\n        x = torch.cat([x, meta], dim=)\n        x = torch.permute(x, (, , ))\n        x1 = self.fc1(x)\n        x1 = torch.permute(x1, (, , ))\n        x1 = self.bn1(x1)\n        x1 = torch.permute(x1, (, , ))\n        x1 = self.act(x1)\n\n        \n        x1 = torch.cat((x, x1), dim=-) \n        \n        x2 = self.fc2(x1)\n        x2 = torch.permute(x2, (, , ))\n        x2 = self.bn2(x2)\n        x2 = torch.permute(x2, (, , ))\n        x2 = self.act(x2)\n        \n        x2 = torch.cat((x, x2), dim=-)\n        \n\n        x3 = self.fc3(x2)\n        x3 = x3.view(-, )\n         x3\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2227688,
      "author_name": "rusg77",
      "author_url": "",
      "post_date": "04/20/2023 00:22:36",
      "content": "<p><code>EuclideanDistanceLoss</code> shows better performance for pretrained models </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228474,
      "author_name": "sofaflyyoufools",
      "author_url": "",
      "post_date": "04/20/2023 15:17:11",
      "content": "<p>Congratulations on your gold medals! Solid purely-GNN score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228618,
      "author_name": "churkinnikita",
      "author_url": "",
      "post_date": "04/20/2023 16:49:24",
      "content": "<p><strong>GNN Training Process.</strong></p>\n<p>Inital loss is von Mises-Fisher 3D Loss. For initial network learning we used <a href=\"https://arxiv.org/abs/1810.06801\" target=\"_blank\">QHAdam</a>, <a href=\"https://arxiv.org/abs/1711.05101\" target=\"_blank\">AdamW</a> and <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> optimizers.</p>\n<p>I have already touched <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">Lion</a> in my <a href=\"https://www.kaggle.com/competitions/learning-equality-curriculum-recommendations/discussion/394827\" target=\"_blank\">previous competition</a>. Learning process often gets whimsical when I try to train using Lion: usually extremely fast convergence to low loss values (but not enough low) followed by bizzare behaviour or complete divergence. But after some tuning and tinkering (<a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">there</a> are some tips for working with Lion) it was possible to get nice checkpoint.</p>\n<p>We continued training our models using L2 loss and different training protocol.<br>\nOur hypothesis on learning our GNNs: loss surface often looked like <a href=\"https://www.sfu.ca/~ssurjano/rastr.html\" target=\"_blank\">Rastrigin function</a>, sophisticated surface with dozens of local minima.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F578229%2F7f4f0049a0ec841238a88275ead39f8f%2Frastr.png?generation=1682008443614288&amp;alt=media\" alt=\"\"><br>\nSource: <a href=\"https://www.sfu.ca/~ssurjano/rastr.html\" target=\"_blank\">https://www.sfu.ca/~ssurjano/rastr.html</a><br>\nLion optimizer led to extremely bad results on that stage; we grid searched through optimizers from <a href=\"https://github.com/jettify/pytorch-optimizer\" target=\"_blank\">pytorch optimizer lib</a> that have the very best performance in the \"Rastrigin scenario\". In our setup <a href=\"https://openreview.net/forum?id=Bkg3g2R9FX\" target=\"_blank\">AdaBound</a> &amp; <a href=\"https://arxiv.org/abs/1908.03265\" target=\"_blank\">RAdam</a> provided splendid performance.</p>\n<p>We used short warmup (5%), non-default betas for optimizers and lower learning rate during 2nd stage training.<br>\nFor my models (\"DynEmb\" variations) only 30m-1h of training time was enough to receive impressive loss drop.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2227682": "First of all, we want to thank Kaggle and IceCube Collaboration for providing this spectacular competition!\n\nOur solution is based on an MLP stack with shared weights of 6 DynEdge models + a blend of 4 GNNs inside the segments by pulse number.\n\n## 1. Model Zoo\n\nWe took the [GraphNeT](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524) model with DynEdge architecture as a baseline and modified it in different ways. Main variations and their nicknames:\n- v1 **Vanilla**. Changed the KNN value and included additional pooling schemes. Replaced Adam optimizer with [QHAdam](https://arxiv.org/abs/1810.06801).\n- v2 **Bigman**. Added more layers and made them wider. Changed activation functions to [SiLU](https://arxiv.org/abs/1702.03118) (they proved to be very powerful in various domains). Added Batch Normalization layers.\n- v3 **DynEmb**. To account for ice heterogeneity, we integrated sensor embeddings into the model (embeddings are initialized by XYZ coordinates). This approach can be generalized and used with any type of Observatory that has sensors distributed in ice or water. We also added Weight Normalization and switched to FP16 training with the new [Lion](https://arxiv.org/abs/2302.06675) optimizer. For this model, we published a submission [example](https://www.kaggle.com/code/ukrainskiydv/icecube-submission-dynemb-example/notebook) that gives the top 20 LB in solo for 1 hour 20 minutes.\n\n## 2. Training Protocol\n\nEvery model had longrun training from scratch with von Mises-Fisher Loss on events up to 600 pulses. Then a couple of additional epochs were trained with Euclidean Distance Loss and up to 200 pulses. As @rusg77 noted below, this loss works much better than the original one for second-stage training. The details on the optimizer choice and second-stage training strategy are presented by @churkinnikita in the comment section.\n\n## 3. Ensembling\n\nFinally, SWA was applied to the best checkpoints of each model for better robustness. As a result, we had v1, v2, v3 models with performance < 1.000 LB and v1, v2, v3 models with performance < 0.990 LB. These 6 models were blended and stacked. Stack showed better performance ~ 0.977 LB in the very final version, @simakov added the details below.\n\n## 4. Alternative Approach\n\nAnother part of our solution consists of several GNN models for different neutrino events. The dataset was segmented by ten uniformly bins on pulse number (that represents energy levels of neutrino), and 4 new models were trained on their compositions. Eventually, the models were blended segment by segment, @rusg77 added the details on this part of the solution in the post [here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402995).\n\n## 5. Results \n\nThe final model is a 60 / 40 blend of 6-stacked / 4-segmented GNNs and shows 0.975 LB (0.976 Private). We used a blend of stacked models to improve the overall stability of the ensemble and reduce the variance of individual models. We also hope this approach leads to better resilience of the model to possible noise in the data.\n\nFor sure, Transformer-based models have great performance on the task, but a diversified ensemble of properly tuned GNNs can also show a keen score on the considered reconstruction problem.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743648%2F829df4033d8353d91b74f0559259506c%2FModels%20Zoo%20RESIZED.png?generation=1682034404975611&alt=media)",
    "2227685": "Stack Model:\n\n```python\nclass MLPShared(nn.Module):\n    def __init__(self,):\n        super(MLPShared, self).__init__()\n\n        self.fc1 = nn.Linear(11, 32) # 6 models + 5 meta features \n        self.fc2 = nn.Linear(32+11, 16)\n        self.fc3 = nn.Linear(16+11, 1)\n\n        self.act = nn.SiLU()\n\n        self.bn1 = nn.BatchNorm1d(32)\n        self.bn2 = nn.BatchNorm1d(16)\n\n    def forward(self, x):\n        x, meta = x[:, :-5], x[:, -5:]\n        meta = torch.repeat_interleave(meta, 3, dim=1)\n        meta = meta.view(-1, 5, 3)\n\n        x = x.view(-1, 6, 3)\n        x = torch.cat([x, meta], dim=1)\n        x = torch.permute(x, (0, 2, 1))\n        x1 = self.fc1(x)\n        x1 = torch.permute(x1, (0, 2, 1))\n        x1 = self.bn1(x1)\n        x1 = torch.permute(x1, (0, 2, 1))\n        x1 = self.act(x1)\n\n        #\n        x1 = torch.cat((x, x1), dim=-1) # dence-light connection\n        #\n        x2 = self.fc2(x1)\n        x2 = torch.permute(x2, (0, 2, 1))\n        x2 = self.bn2(x2)\n        x2 = torch.permute(x2, (0, 2, 1))\n        x2 = self.act(x2)\n        #\n        x2 = torch.cat((x, x2), dim=-1)\n        #\n\n        x3 = self.fc3(x2)\n        x3 = x3.view(-1, 3)\n        return x3\n```",
    "2227688": "`EuclideanDistanceLoss` shows better performance for pretrained models",
    "2228474": "Congratulations on your gold medals! Solid purely-GNN score.",
    "2228618": "**GNN Training Process.**\n\nInital loss is von Mises-Fisher 3D Loss. For initial network learning we used [QHAdam](https://arxiv.org/abs/1810.06801), [AdamW](https://arxiv.org/abs/1711.05101) and [Lion](https://arxiv.org/abs/2302.06675) optimizers.\n\nI have already touched [Lion](https://arxiv.org/abs/2302.06675) in my [previous competition](https://www.kaggle.com/competitions/learning-equality-curriculum-recommendations/discussion/394827). Learning process often gets whimsical when I try to train using Lion: usually extremely fast convergence to low loss values (but not enough low) followed by bizzare behaviour or complete divergence. But after some tuning and tinkering ([there](https://github.com/lucidrains/lion-pytorch) are some tips for working with Lion) it was possible to get nice checkpoint.\n\nWe continued training our models using L2 loss and different training protocol.\nOur hypothesis on learning our GNNs: loss surface often looked like [Rastrigin function](https://www.sfu.ca/~ssurjano/rastr.html), sophisticated surface with dozens of local minima.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F578229%2F7f4f0049a0ec841238a88275ead39f8f%2Frastr.png?generation=1682008443614288&alt=media)\nSource: https://www.sfu.ca/~ssurjano/rastr.html\nLion optimizer led to extremely bad results on that stage; we grid searched through optimizers from [pytorch optimizer lib](https://github.com/jettify/pytorch-optimizer) that have the very best performance in the \"Rastrigin scenario\". In our setup [AdaBound](https://openreview.net/forum?id=Bkg3g2R9FX) & [RAdam](https://arxiv.org/abs/1908.03265) provided splendid performance.\n\nWe used short warmup (5%), non-default betas for optimizers and lower learning rate during 2nd stage training.\nFor my models (\"DynEmb\" variations) only 30m-1h of training time was enough to receive impressive loss drop."
  },
  "source": "meta"
}