{
  "id": 402995,
  "title": "9th place solution. 0.983/0.982 GraphNet only (rusg77 part)",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402995",
  "author_name": "",
  "post_date": "2023-04-20T15:47:07.782692Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I'd like to thank organizers and congratulations to all the winners!<br>\nSpecial thanks to:</p>\n<ul>\n<li>my teammates: Nikita Churkin <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a>, Dmitry Simakov <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a>, Dmitriy Ukrainskiy <a href=\"https://www.kaggle.com/ukrainskiydv\" target=\"_blank\">@ukrainskiydv</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for per batch iteration approach (<a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching</a>)</li>\n<li><a href=\"https://www.kaggle.com/antonsevostianov\" target=\"_blank\">@antonsevostianov</a> for sensors feature engineering (<a href=\"https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook\" target=\"_blank\">https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook</a>)</li>\n<li><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> for clean code and tidy code (<a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">https://github.com/graphnet-team/graphnet</a>)</li>\n</ul>\n<p>Our entire solution is available here: <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849</a> </p>\n<p>Here is my part. I'll skip all things I tried and focus on the most significant improvements.</p>\n<h2>Reproduce baseline score (1.017) from scratch</h2>\n<p>For me it wasn't easy. But when tried to play with learning rate I achieved <strong>1.007/1.005</strong></p>\n<h2>Add more features</h2>\n<p><code>additional_features = ['is_abovedust', 'is_deepcore', 'is_dustlayer', 'is_underdust', 'scattering', 'absorption', 'relative_qe']</code><br>\nTrain on first 200 batches, 3x3090 rented host, 1d 6h. <strong>0.999/0.998</strong></p>\n<h2>Two models</h2>\n<p>I noticed that model performance is poor for events with a low amount of pulses, especially less than 50. I tuned a model for such events with increased nb_neighbours. <strong>0.992/0.990</strong></p>\n<p>Two parts inference:</p>\n<ul>\n<li>model_1 for events with low amount of pulses</li>\n<li>model_2 for the rest</li>\n</ul>\n<h2>Train pipeline optimization</h2>\n<p>SQlite approach required a lot of disk space, long processing. I tried to switch to Posgresql but the result was almost the same. CPU, disk space, memory for each connection are bottlenecks. Then <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> published his amazing approach with per batch iteration and shuffling batches every epoch. With little modification I make it work with multiple GPUs.</p>\n<h2>EuclidianDistanceLoss</h2>\n<p>When I was about to despair, I suddenly discovered that EuclidianDistanceLoss shows much better results.</p>\n<h2>Split events by n_pulses in 10 buckets</h2>\n<p>So now I have a good train pipeline. I continued to train models but this time I split data into ~equal 10 buckets by number of pulses. <br>\nThe most dramatic change I achieved was for the last bucket with 145+ pulses. Train on events with 145-2000 pulses. Rented machine with 4x3090 ~2 days. <strong>0.983/0.982</strong><br>\n<a href=\"https://postimg.cc/JsVKVvds\" target=\"_blank\"><img src=\"https://i.postimg.cc/K84WLyPD/Screen-Shot-2023-04-20-at-17-21-45.png\" alt=\"Screen-Shot-2023-04-20-at-17-21-45.png\"></a></p>\n<h2>Team merge</h2>\n<p>Weights 60/40 (my part) for x,y,z coordinates + norm to 1<br>\nFirst day: 0.6 * 0.984 + 0.4 * 0.987 -&gt; <strong>0.981/0.979</strong><br>\nLast day: 0.6 * 0.978 + 0.4 * 0.982 -&gt; <strong>0.976/0.975</strong></p>\n<h2>Lessons</h2>\n<ul>\n<li>spend a little more time on code quality, experiments logging</li>\n<li>GPUs is not all you need, but it matters</li>\n<li>Try well known architectures instead of trying paperswithcode graph convolutions</li>\n</ul>\n<p>Good luck and happy kaggling!</p>",
  "messages": [
    {
      "id": "2228523",
      "postDate": "04/20/2023 15:47:07",
      "content": "<p>First of all, I'd like to thank organizers and congratulations to all the winners!<br>\nSpecial thanks to:</p>\n<ul>\n<li>my teammates: Nikita Churkin <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a>, Dmitry Simakov <a href=\"https://www.kaggle.com/simakov\" target=\"_blank\">@simakov</a>, Dmitriy Ukrainskiy <a href=\"https://www.kaggle.com/ukrainskiydv\" target=\"_blank\">@ukrainskiydv</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for per batch iteration approach (<a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching\" target=\"_blank\">https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching</a>)</li>\n<li><a href=\"https://www.kaggle.com/antonsevostianov\" target=\"_blank\">@antonsevostianov</a> for sensors feature engineering (<a href=\"https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook\" target=\"_blank\">https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook</a>)</li>\n<li><a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> for clean code and tidy code (<a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">https://github.com/graphnet-team/graphnet</a>)</li>\n</ul>\n<p>Our entire solution is available here: <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849</a> </p>\n<p>Here is my part. I'll skip all things I tried and focus on the most significant improvements.</p>\n<h2>Reproduce baseline score (1.017) from scratch</h2>\n<p>For me it wasn't easy. But when tried to play with learning rate I achieved <strong>1.007/1.005</strong></p>\n<h2>Add more features</h2>\n<p><code>additional_features = ['is_abovedust', 'is_deepcore', 'is_dustlayer', 'is_underdust', 'scattering', 'absorption', 'relative_qe']</code><br>\nTrain on first 200 batches, 3x3090 rented host, 1d 6h. <strong>0.999/0.998</strong></p>\n<h2>Two models</h2>\n<p>I noticed that model performance is poor for events with a low amount of pulses, especially less than 50. I tuned a model for such events with increased nb_neighbours. <strong>0.992/0.990</strong></p>\n<p>Two parts inference:</p>\n<ul>\n<li>model_1 for events with low amount of pulses</li>\n<li>model_2 for the rest</li>\n</ul>\n<h2>Train pipeline optimization</h2>\n<p>SQlite approach required a lot of disk space, long processing. I tried to switch to Posgresql but the result was almost the same. CPU, disk space, memory for each connection are bottlenecks. Then <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> published his amazing approach with per batch iteration and shuffling batches every epoch. With little modification I make it work with multiple GPUs.</p>\n<h2>EuclidianDistanceLoss</h2>\n<p>When I was about to despair, I suddenly discovered that EuclidianDistanceLoss shows much better results.</p>\n<h2>Split events by n_pulses in 10 buckets</h2>\n<p>So now I have a good train pipeline. I continued to train models but this time I split data into ~equal 10 buckets by number of pulses. <br>\nThe most dramatic change I achieved was for the last bucket with 145+ pulses. Train on events with 145-2000 pulses. Rented machine with 4x3090 ~2 days. <strong>0.983/0.982</strong><br>\n<a href=\"https://postimg.cc/JsVKVvds\" target=\"_blank\"><img src=\"https://i.postimg.cc/K84WLyPD/Screen-Shot-2023-04-20-at-17-21-45.png\" alt=\"Screen-Shot-2023-04-20-at-17-21-45.png\"></a></p>\n<h2>Team merge</h2>\n<p>Weights 60/40 (my part) for x,y,z coordinates + norm to 1<br>\nFirst day: 0.6 * 0.984 + 0.4 * 0.987 -&gt; <strong>0.981/0.979</strong><br>\nLast day: 0.6 * 0.978 + 0.4 * 0.982 -&gt; <strong>0.976/0.975</strong></p>\n<h2>Lessons</h2>\n<ul>\n<li>spend a little more time on code quality, experiments logging</li>\n<li>GPUs is not all you need, but it matters</li>\n<li>Try well known architectures instead of trying paperswithcode graph convolutions</li>\n</ul>\n<p>Good luck and happy kaggling!</p>",
      "rawMarkdown": "First of all, I'd like to thank organizers and congratulations to all the winners!\nSpecial thanks to:\n- my teammates: Nikita Churkin @churkinnikita, Dmitry Simakov @simakov, Dmitriy Ukrainskiy @ukrainskiydv\n- @iafoss for per batch iteration approach (https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching)\n- @antonsevostianov for sensors feature engineering (https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook)\n- @rasmusrse for clean code and tidy code (https://github.com/graphnet-team/graphnet)\n\nOur entire solution is available here: https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849 \n\nHere is my part. I'll skip all things I tried and focus on the most significant improvements.\n\n## Reproduce baseline score (1.017) from scratch \nFor me it wasn't easy. But when tried to play with learning rate I achieved **1.007/1.005**\n\n## Add more features\n`additional_features = ['is_abovedust', 'is_deepcore', 'is_dustlayer', 'is_underdust', 'scattering', 'absorption', 'relative_qe']`\nTrain on first 200 batches, 3x3090 rented host, 1d 6h. **0.999/0.998**\n\n## Two models\nI noticed that model performance is poor for events with a low amount of pulses, especially less than 50. I tuned a model for such events with increased nb_neighbours. **0.992/0.990**\n\nTwo parts inference:\n- model_1 for events with low amount of pulses\n- model_2 for the rest\n\n## Train pipeline optimization\nSQlite approach required a lot of disk space, long processing. I tried to switch to Posgresql but the result was almost the same. CPU, disk space, memory for each connection are bottlenecks. Then @iafoss published his amazing approach with per batch iteration and shuffling batches every epoch. With little modification I make it work with multiple GPUs.\n\n## EuclidianDistanceLoss\nWhen I was about to despair, I suddenly discovered that EuclidianDistanceLoss shows much better results.\n\n## Split events by n_pulses in 10 buckets\nSo now I have a good train pipeline. I continued to train models but this time I split data into ~equal 10 buckets by number of pulses. \nThe most dramatic change I achieved was for the last bucket with 145+ pulses. Train on events with 145-2000 pulses. Rented machine with 4x3090 ~2 days. **0.983/0.982**\n[![Screen-Shot-2023-04-20-at-17-21-45.png](https://i.postimg.cc/K84WLyPD/Screen-Shot-2023-04-20-at-17-21-45.png)](https://postimg.cc/JsVKVvds)\n\n## Team merge\nWeights 60/40 (my part) for x,y,z coordinates + norm to 1\nFirst day: 0.6 * 0.984 + 0.4 * 0.987 -> **0.981/0.979**\nLast day: 0.6 * 0.978 + 0.4 * 0.982 -> **0.976/0.975**\n\n## Lessons\n- spend a little more time on code quality, experiments logging\n- GPUs is not all you need, but it matters\n- Try well known architectures instead of trying paperswithcode graph convolutions\n\nGood luck and happy kaggling!",
      "votes": null
    },
    {
      "id": "2230910",
      "postDate": "04/22/2023 21:57:09",
      "content": "<p>Congratulations, Ruslan! Glad that my messy code (found a few mislabeled sensors later on) still helped you! Probably could've squeezed a few more 0.00001s if not for me being lazy bones, haha.</p>",
      "rawMarkdown": "Congratulations, Ruslan! Glad that my messy code (found a few mislabeled sensors later on) still helped you! Probably could've squeezed a few more 0.00001s if not for me being lazy bones, haha.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2230910,
      "author_name": "antonsevostianov",
      "author_url": "",
      "post_date": "04/22/2023 21:57:09",
      "content": "<p>Congratulations, Ruslan! Glad that my messy code (found a few mislabeled sensors later on) still helped you! Probably could've squeezed a few more 0.00001s if not for me being lazy bones, haha.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228523": "First of all, I'd like to thank organizers and congratulations to all the winners!\nSpecial thanks to:\n- my teammates: Nikita Churkin @churkinnikita, Dmitry Simakov @simakov, Dmitriy Ukrainskiy @ukrainskiydv\n- @iafoss for per batch iteration approach (https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching)\n- @antonsevostianov for sensors feature engineering (https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook)\n- @rasmusrse for clean code and tidy code (https://github.com/graphnet-team/graphnet)\n\nOur entire solution is available here: https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/402849 \n\nHere is my part. I'll skip all things I tried and focus on the most significant improvements.\n\n## Reproduce baseline score (1.017) from scratch \nFor me it wasn't easy. But when tried to play with learning rate I achieved **1.007/1.005**\n\n## Add more features\n`additional_features = ['is_abovedust', 'is_deepcore', 'is_dustlayer', 'is_underdust', 'scattering', 'absorption', 'relative_qe']`\nTrain on first 200 batches, 3x3090 rented host, 1d 6h. **0.999/0.998**\n\n## Two models\nI noticed that model performance is poor for events with a low amount of pulses, especially less than 50. I tuned a model for such events with increased nb_neighbours. **0.992/0.990**\n\nTwo parts inference:\n- model_1 for events with low amount of pulses\n- model_2 for the rest\n\n## Train pipeline optimization\nSQlite approach required a lot of disk space, long processing. I tried to switch to Posgresql but the result was almost the same. CPU, disk space, memory for each connection are bottlenecks. Then @iafoss published his amazing approach with per batch iteration and shuffling batches every epoch. With little modification I make it work with multiple GPUs.\n\n## EuclidianDistanceLoss\nWhen I was about to despair, I suddenly discovered that EuclidianDistanceLoss shows much better results.\n\n## Split events by n_pulses in 10 buckets\nSo now I have a good train pipeline. I continued to train models but this time I split data into ~equal 10 buckets by number of pulses. \nThe most dramatic change I achieved was for the last bucket with 145+ pulses. Train on events with 145-2000 pulses. Rented machine with 4x3090 ~2 days. **0.983/0.982**\n[![Screen-Shot-2023-04-20-at-17-21-45.png](https://i.postimg.cc/K84WLyPD/Screen-Shot-2023-04-20-at-17-21-45.png)](https://postimg.cc/JsVKVvds)\n\n## Team merge\nWeights 60/40 (my part) for x,y,z coordinates + norm to 1\nFirst day: 0.6 * 0.984 + 0.4 * 0.987 -> **0.981/0.979**\nLast day: 0.6 * 0.978 + 0.4 * 0.982 -> **0.976/0.975**\n\n## Lessons\n- spend a little more time on code quality, experiments logging\n- GPUs is not all you need, but it matters\n- Try well known architectures instead of trying paperswithcode graph convolutions\n\nGood luck and happy kaggling!",
    "2230910": "Congratulations, Ruslan! Glad that my messy code (found a few mislabeled sensors later on) still helped you! Probably could've squeezed a few more 0.00001s if not for me being lazy bones, haha."
  },
  "source": "meta"
}