{
  "id": 402861,
  "title": "17th place solution (silver)",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/men-of-the-year-17th-place-solution-silver",
  "author_name": "",
  "post_date": "2023-04-20T14:55:36.943Z",
  "votes": 24,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many thanks to everyone, it was an exciting competition!</p>\n<p>I also want to give thanks to <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> and <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> for your notebooks with LSTM solutions to this problem. They really helped me to understand the possible approaches for this task. </p>\n<h3>Our solution is an ensemble of 2 types of models:</h3>\n<h4>1. Ensemble of 5 LSTM models, that are trained on classification problem to predict bins of angles</h4>\n<p>Tried different architectures. It appeared, that actually LSTM here worked better, than GRU. However, in one model we used 3 LSTM + 2 GRU layers and it also worked very good. We also used exponential LR, which gave slight boost. Cosine annealing LR didn't work good in our case. Models very trained on 200-300 files. I once tried 400-600 files, but training was much slower and the score increased also slowly. There was a deadline soon, so it was more \"profitable\" to use 200-300 files. We also tried to predict just a XYZ vector and the convert it to angle, tried MSE and VmF loss, didn't work well. It was ok, but not better, than predicting bins of angles.</p>\n<h4>2. DynEdge graphnets (5-7 in best submissions), just trained on VmF3d loss to predict XYZ vectors (as in the baseline)</h4>\n<p>Didn't really change anything, just trained this model on 200 files. Tried different lengths of pulses: from 96 to 200. More pulses gave better results. We managed to train about 20 epochs before the deadline, but it can definitely give better results with more training.</p>\n<p>While ensembling, we used different ensemble weights to predict azimuth and zenith. They were chosen on a validation set.</p>\n<p>Each graphnet takes around 30 minutes on inference. LSTM ensemble takes around 2 hours and has 0.999 on the private LB. We didn't check the score of pure graphnet ensemble on the LB, but the best single model is around 1.006. Mixing these 2 types of models gave us a nice boost (0.987 on the private).</p>\n<h3>Additional thoughts:</h3>\n<ul>\n<li>GraphNet is a cool thing, however, it felt like it can be pushed further in our case. It trains well, but after some epochs loss starts to increase slowly. Probably should tried different learning rate schedulers, but didn't have time at the end. </li>\n<li>We also tried a transformer model, trained on XYZ vector with MSE/VmF loss. Moreover, I believed, that in this competition it should win top places. However, for some reason it didn't overperform the LSTMs. Maybe I chose bad params in its architecture (was my first time, yeah). Maybe we needed some additional techniques and other losses to make it work.</li>\n<li>Ensembling gives nice boosts. I think that models are still not robust and in many cases they are not confident in predictions.</li>\n</ul>",
  "messages": [
    {
      "id": "2227738",
      "postDate": "04/20/2023 01:49:16",
      "content": "<p>Many thanks to everyone, it was an exciting competition!</p>\n<p>I also want to give thanks to <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> and <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">@seungmoklee</a> for your notebooks with LSTM solutions to this problem. They really helped me to understand the possible approaches for this task. </p>\n<h3>Our solution is an ensemble of 2 types of models:</h3>\n<h4>1. Ensemble of 5 LSTM models, that are trained on classification problem to predict bins of angles</h4>\n<p>Tried different architectures. It appeared, that actually LSTM here worked better, than GRU. However, in one model we used 3 LSTM + 2 GRU layers and it also worked very good. We also used exponential LR, which gave slight boost. Cosine annealing LR didn't work good in our case. Models very trained on 200-300 files. I once tried 400-600 files, but training was much slower and the score increased also slowly. There was a deadline soon, so it was more \"profitable\" to use 200-300 files. We also tried to predict just a XYZ vector and the convert it to angle, tried MSE and VmF loss, didn't work well. It was ok, but not better, than predicting bins of angles.</p>\n<h4>2. DynEdge graphnets (5-7 in best submissions), just trained on VmF3d loss to predict XYZ vectors (as in the baseline)</h4>\n<p>Didn't really change anything, just trained this model on 200 files. Tried different lengths of pulses: from 96 to 200. More pulses gave better results. We managed to train about 20 epochs before the deadline, but it can definitely give better results with more training.</p>\n<p>While ensembling, we used different ensemble weights to predict azimuth and zenith. They were chosen on a validation set.</p>\n<p>Each graphnet takes around 30 minutes on inference. LSTM ensemble takes around 2 hours and has 0.999 on the private LB. We didn't check the score of pure graphnet ensemble on the LB, but the best single model is around 1.006. Mixing these 2 types of models gave us a nice boost (0.987 on the private).</p>\n<h3>Additional thoughts:</h3>\n<ul>\n<li>GraphNet is a cool thing, however, it felt like it can be pushed further in our case. It trains well, but after some epochs loss starts to increase slowly. Probably should tried different learning rate schedulers, but didn't have time at the end. </li>\n<li>We also tried a transformer model, trained on XYZ vector with MSE/VmF loss. Moreover, I believed, that in this competition it should win top places. However, for some reason it didn't overperform the LSTMs. Maybe I chose bad params in its architecture (was my first time, yeah). Maybe we needed some additional techniques and other losses to make it work.</li>\n<li>Ensembling gives nice boosts. I think that models are still not robust and in many cases they are not confident in predictions.</li>\n</ul>",
      "rawMarkdown": "Many thanks to everyone, it was an exciting competition!\n\nI also want to give thanks to @rsmits and @seungmoklee for your notebooks with LSTM solutions to this problem. They really helped me to understand the possible approaches for this task. \n\n### Our solution is an ensemble of 2 types of models:\n#### 1. Ensemble of 5 LSTM models, that are trained on classification problem to predict bins of angles\nTried different architectures. It appeared, that actually LSTM here worked better, than GRU. However, in one model we used 3 LSTM + 2 GRU layers and it also worked very good. We also used exponential LR, which gave slight boost. Cosine annealing LR didn't work good in our case. Models very trained on 200-300 files. I once tried 400-600 files, but training was much slower and the score increased also slowly. There was a deadline soon, so it was more \"profitable\" to use 200-300 files. We also tried to predict just a XYZ vector and the convert it to angle, tried MSE and VmF loss, didn't work well. It was ok, but not better, than predicting bins of angles.\n\n#### 2. DynEdge graphnets (5-7 in best submissions), just trained on VmF3d loss to predict XYZ vectors (as in the baseline)\nDidn't really change anything, just trained this model on 200 files. Tried different lengths of pulses: from 96 to 200. More pulses gave better results. We managed to train about 20 epochs before the deadline, but it can definitely give better results with more training.\n\nWhile ensembling, we used different ensemble weights to predict azimuth and zenith. They were chosen on a validation set.\n\nEach graphnet takes around 30 minutes on inference. LSTM ensemble takes around 2 hours and has 0.999 on the private LB. We didn't check the score of pure graphnet ensemble on the LB, but the best single model is around 1.006. Mixing these 2 types of models gave us a nice boost (0.987 on the private).\n\n### Additional thoughts:\n- GraphNet is a cool thing, however, it felt like it can be pushed further in our case. It trains well, but after some epochs loss starts to increase slowly. Probably should tried different learning rate schedulers, but didn't have time at the end. \n- We also tried a transformer model, trained on XYZ vector with MSE/VmF loss. Moreover, I believed, that in this competition it should win top places. However, for some reason it didn't overperform the LSTMs. Maybe I chose bad params in its architecture (was my first time, yeah). Maybe we needed some additional techniques and other losses to make it work.\n- Ensembling gives nice boosts. I think that models are still not robust and in many cases they are not confident in predictions.",
      "votes": null
    },
    {
      "id": "2227872",
      "postDate": "04/20/2023 04:47:42",
      "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> interesting solution. Thanks for sharing!</p>",
      "rawMarkdown": "manwithaflower interesting solution. Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227872,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "04/20/2023 04:47:42",
      "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> interesting solution. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2227738": "Many thanks to everyone, it was an exciting competition!\n\nI also want to give thanks to @rsmits and @seungmoklee for your notebooks with LSTM solutions to this problem. They really helped me to understand the possible approaches for this task. \n\n### Our solution is an ensemble of 2 types of models:\n#### 1. Ensemble of 5 LSTM models, that are trained on classification problem to predict bins of angles\nTried different architectures. It appeared, that actually LSTM here worked better, than GRU. However, in one model we used 3 LSTM + 2 GRU layers and it also worked very good. We also used exponential LR, which gave slight boost. Cosine annealing LR didn't work good in our case. Models very trained on 200-300 files. I once tried 400-600 files, but training was much slower and the score increased also slowly. There was a deadline soon, so it was more \"profitable\" to use 200-300 files. We also tried to predict just a XYZ vector and the convert it to angle, tried MSE and VmF loss, didn't work well. It was ok, but not better, than predicting bins of angles.\n\n#### 2. DynEdge graphnets (5-7 in best submissions), just trained on VmF3d loss to predict XYZ vectors (as in the baseline)\nDidn't really change anything, just trained this model on 200 files. Tried different lengths of pulses: from 96 to 200. More pulses gave better results. We managed to train about 20 epochs before the deadline, but it can definitely give better results with more training.\n\nWhile ensembling, we used different ensemble weights to predict azimuth and zenith. They were chosen on a validation set.\n\nEach graphnet takes around 30 minutes on inference. LSTM ensemble takes around 2 hours and has 0.999 on the private LB. We didn't check the score of pure graphnet ensemble on the LB, but the best single model is around 1.006. Mixing these 2 types of models gave us a nice boost (0.987 on the private).\n\n### Additional thoughts:\n- GraphNet is a cool thing, however, it felt like it can be pushed further in our case. It trains well, but after some epochs loss starts to increase slowly. Probably should tried different learning rate schedulers, but didn't have time at the end. \n- We also tried a transformer model, trained on XYZ vector with MSE/VmF loss. Moreover, I believed, that in this competition it should win top places. However, for some reason it didn't overperform the LSTMs. Maybe I chose bad params in its architecture (was my first time, yeah). Maybe we needed some additional techniques and other losses to make it work.\n- Ensembling gives nice boosts. I think that models are still not robust and in many cases they are not confident in predictions.",
    "2227872": "manwithaflower interesting solution. Thanks for sharing!"
  },
  "source": "meta"
}