{
  "id": 403398,
  "title": "5th place solution: combining GraphNet and Transformer with LSTM meta-model",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/big-bento-box-of-science-5th-place-solution-combin",
  "author_name": "",
  "post_date": "2023-04-27T17:47:19.043Z",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We'd like to thank the organizers and Kaggle staff for putting together this enthralling challenge. As we share the scientific feast we've cooked up in our \"Big Bento Box of Science,\" please accept our apologies for the slightly delayed write-up. The competition was very intense, with our team initially lagging over 400 spots away from the medal zone just two weeks before the end. Through hard work and persistence, we managed to improve our standing, ultimately finishing in 5th place and achieving our first gold medal.</p>\n<h2>Summary:</h2>\n<ul>\n<li>Two classification heads with 256 bins each for transformer, smart label smoothing -&gt; <strong>1.003</strong></li>\n<li>Increasing the number of features GraphNet uses in KNN graph builder, changing activation functions to LayerNorm, overall increasing model size, predicting <code>(x, y, z)</code> direction instead of azimuth and zenith directly -&gt; <strong>0.990</strong></li>\n<li>LSTM model selector which took the input data as well as predictions from Transformer and GraphNet, and returned prediction which of these models to use -&gt; <strong>0.974</strong></li>\n</ul>\n<h2>Transformer</h2>\n<p>We used standard encoder architecture with 8 layers, 512 model size and 8 attention heads. The input consisted of <code>(x, y, z, t, charge, auxiliary, is_core, rank)</code>. <code>is_core</code> was a 0/1 value indicating whether the sensor was part of a deep core or not. <code>rank</code> was computed the same way as in the <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">following great notebook with LSTM</a>. For features like <code>(x, y, z, t, charge)</code> we used a few linear layers to preprocess them before passing them to the Transformer. While <code>(auxiliary, is_core, rank)</code> due to their discrete nature were embedded.</p>\n<p>We had two output heads for both azimuth and zenith. Each head predicted 256 values that represented bins. We used equally spaced bins and trained them using cross entropy with smoothed labels. The labels were smoothed so that neighboring bins also have some probability (not like in standard label smoothing where remaining probability is equally assigned to all classes). For azimuth this smoothing was wrapped around in circular fashion, so that predicting \\(0\\) and \\(2 \\pi \\) was equivalent. Prediction was just an argmax over the bins. We experimented with a different number of bins, ranging from 16 to 256, and more bins resulted in better performance. However, the difference between 128 and 256 was not significant and we chose to use 256. </p>\n<p>We trained the model for 760k steps on TPU with batch size of 512 and sequence length of 256, so around 3 epochs (which is not much). We used a cosine schedule for LR with peak LR=3e-5 and repeated the cycle every 10k steps (looking back at it the cycle should probably be longer, but at the beginning we weren’t sure what to expect from the training). For the last 160k steps we changed peak LR to 3e-6. Whole training took around 26h. After 760k steps we could still see some slight improvements as the training progressed but decided to focus more on other ideas. This led to a model with a score of <strong>1.003</strong> on public LB.</p>\n<h2>GraphNet</h2>\n<p>As a starting point we’ve used <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">standard GraphNet model</a> with <a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">additional input features from early sharing price notebook</a>.</p>\n<p>Model input was composed of <code>(x, y, z, t, charge, auxiliary, quantum_effeciency, wrong_charge_dom, ice_absorption, ice_scatter)</code> where <code>wrong_charge_dom</code> was a marking of one sensor (num. 1229) because our analysis shown there might be something wrong with it. Output of the model was a direction. </p>\n<p>To achieve the final LB score we've made a few changes to the model.</p>\n<p>First of all, after looking through the code of <code>DynEdge</code> we saw that <code>features_subset</code> argument (latent feature subset for computing nearest neighbors in <code>DynEdge</code>) is applied not only on the first layer, but also on inputs to other layers. Intuitively it felt limiting, we changed it so that <code>features_subset</code> for the non-first layer was variable. In the end we set it to 48 (bigger seems to work better). With this change our model achieved <strong>0.998</strong> on LB.</p>\n<p>Next big change was moving from <code>BatchNorm1d</code> to <code>LayerNorm</code>. This decision was guided by a big difference in MAE metrics we observed between training and validation. Even the standard model (with <code>BatchNorm1d</code>) got better results without switching to eval mode, so we decided to move to <code>LayerNorm</code> that does not differ between training and evaluation.</p>\n<p>Last but not least, we experimented and tuned the whole training procedure. We used AdaBelief optimizer. Unlike the Transformer, with GraphNet model we took the approach of changing the learning rate in steps, starting from 1e-3 and reducing the learning rate by a factor of 2 every 200k steps. Total number of steps was 700k with 512 batch size. It took around 3 days on RTX 4090.</p>\n<p>Finally, with all changes and scaling of the model we were able to achieve a <strong>0.990</strong> score on LB. </p>\n<h2>Model stacking ensemble</h2>\n<p>In the last few days of competition we had two models and no time to train them any further. We looked previously at ensembling them but results were not promising. We made some analysis and found that the predictions are sometimes very different from both models. If we could pick the correct model perfectly the lower bound was around 0.810. </p>\n<p>Two days before the end of competition we dumped the embeddings from our best Transformer (last common layer before two categorical heads) and GraphNet (layer before predicting direction) for the first 10 data batches. At first we fed both of those embeddings along with model predictions into XGBoost. It was able to achieve 57% accuracy which resulted in <strong>0.976</strong> score on LB (and 0.978 on our internal validation).</p>\n<p>Later on we coded up a simple 3 layer Bi-directional LSTM as a model selector. Our thinking was that having a different architecture (not Transformer or GraphNet) as a selector will be beneficial in learning some different features from data resulting in better selections.</p>\n<p>LSTMs hidden states were initialized with concatenated Transformer and GraphNet embeddings and subset of pulses were given as input sequence*. We thought that training it on event data will help it make better decisions. We trained the model for around 9 epochs (keep in mind it’s only on the first 10 batches) which took about 8h. Doing this improved accuracy to 58% and resulted in a score of <strong>0.974</strong> which was our final result.</p>\n<p>(*) We also experimented with concatenating the embeddings to the input sequence or just before the last classification layer, but both of those approaches turned out to be worse than just initializing hidden states with embeddings.</p>\n<p>Overall architecture of our final submission looked like this:</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/2230884/19025/model-idea.png\" alt=\"model diagram\"></p>\n<h2>Things that we tried</h2>\n<ul>\n<li>We trained two Transformer models on different subset of dataset and tried simple ensembling (adding logits/softmaxes and argmaxing) but this resulted only in marginal improvements of 0.002</li>\n<li>Training a separate categorical heads (like in case of transformer) on GraphNet embeddings and then ensemble both models by simply adding logits/softmaxes and picking argmax - this slightly reduced the error but was not spectacular</li>\n<li>Classification task for GraphNet didn’t work that great, neither have azimuth and zenith regression. Direction regression is what performed best.</li>\n<li>At the beginning of the competition we found a paper “A Convolutional Neural Network based Cascade Reconstruction for the IceCube Neutrino Observatory” and implemented a variant of a 3d convolutional network with so called hexagonal convolutions, which are supposed to be localized in the DOMs space. However, we couldn’t get it to work and get the mean angular error below 1.2. A likely reason is that the data in the competition was quite different from the data used in the paper. The latter had great resolution in the time dimension, because the pulses were collected every 2ns, while our data was much more sparse.</li>\n</ul>\n<h2>Inference code</h2>\n<p><a href=\"https://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer\" target=\"_blank\">https://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer</a></p>",
  "messages": [
    {
      "id": "2230884",
      "postDate": "04/22/2023 20:56:54",
      "content": "<p>We'd like to thank the organizers and Kaggle staff for putting together this enthralling challenge. As we share the scientific feast we've cooked up in our \"Big Bento Box of Science,\" please accept our apologies for the slightly delayed write-up. The competition was very intense, with our team initially lagging over 400 spots away from the medal zone just two weeks before the end. Through hard work and persistence, we managed to improve our standing, ultimately finishing in 5th place and achieving our first gold medal.</p>\n<h2>Summary:</h2>\n<ul>\n<li>Two classification heads with 256 bins each for transformer, smart label smoothing -&gt; <strong>1.003</strong></li>\n<li>Increasing the number of features GraphNet uses in KNN graph builder, changing activation functions to LayerNorm, overall increasing model size, predicting <code>(x, y, z)</code> direction instead of azimuth and zenith directly -&gt; <strong>0.990</strong></li>\n<li>LSTM model selector which took the input data as well as predictions from Transformer and GraphNet, and returned prediction which of these models to use -&gt; <strong>0.974</strong></li>\n</ul>\n<h2>Transformer</h2>\n<p>We used standard encoder architecture with 8 layers, 512 model size and 8 attention heads. The input consisted of <code>(x, y, z, t, charge, auxiliary, is_core, rank)</code>. <code>is_core</code> was a 0/1 value indicating whether the sensor was part of a deep core or not. <code>rank</code> was computed the same way as in the <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">following great notebook with LSTM</a>. For features like <code>(x, y, z, t, charge)</code> we used a few linear layers to preprocess them before passing them to the Transformer. While <code>(auxiliary, is_core, rank)</code> due to their discrete nature were embedded.</p>\n<p>We had two output heads for both azimuth and zenith. Each head predicted 256 values that represented bins. We used equally spaced bins and trained them using cross entropy with smoothed labels. The labels were smoothed so that neighboring bins also have some probability (not like in standard label smoothing where remaining probability is equally assigned to all classes). For azimuth this smoothing was wrapped around in circular fashion, so that predicting \\(0\\) and \\(2 \\pi \\) was equivalent. Prediction was just an argmax over the bins. We experimented with a different number of bins, ranging from 16 to 256, and more bins resulted in better performance. However, the difference between 128 and 256 was not significant and we chose to use 256. </p>\n<p>We trained the model for 760k steps on TPU with batch size of 512 and sequence length of 256, so around 3 epochs (which is not much). We used a cosine schedule for LR with peak LR=3e-5 and repeated the cycle every 10k steps (looking back at it the cycle should probably be longer, but at the beginning we weren’t sure what to expect from the training). For the last 160k steps we changed peak LR to 3e-6. Whole training took around 26h. After 760k steps we could still see some slight improvements as the training progressed but decided to focus more on other ideas. This led to a model with a score of <strong>1.003</strong> on public LB.</p>\n<h2>GraphNet</h2>\n<p>As a starting point we’ve used <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission\" target=\"_blank\">standard GraphNet model</a> with <a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">additional input features from early sharing price notebook</a>.</p>\n<p>Model input was composed of <code>(x, y, z, t, charge, auxiliary, quantum_effeciency, wrong_charge_dom, ice_absorption, ice_scatter)</code> where <code>wrong_charge_dom</code> was a marking of one sensor (num. 1229) because our analysis shown there might be something wrong with it. Output of the model was a direction. </p>\n<p>To achieve the final LB score we've made a few changes to the model.</p>\n<p>First of all, after looking through the code of <code>DynEdge</code> we saw that <code>features_subset</code> argument (latent feature subset for computing nearest neighbors in <code>DynEdge</code>) is applied not only on the first layer, but also on inputs to other layers. Intuitively it felt limiting, we changed it so that <code>features_subset</code> for the non-first layer was variable. In the end we set it to 48 (bigger seems to work better). With this change our model achieved <strong>0.998</strong> on LB.</p>\n<p>Next big change was moving from <code>BatchNorm1d</code> to <code>LayerNorm</code>. This decision was guided by a big difference in MAE metrics we observed between training and validation. Even the standard model (with <code>BatchNorm1d</code>) got better results without switching to eval mode, so we decided to move to <code>LayerNorm</code> that does not differ between training and evaluation.</p>\n<p>Last but not least, we experimented and tuned the whole training procedure. We used AdaBelief optimizer. Unlike the Transformer, with GraphNet model we took the approach of changing the learning rate in steps, starting from 1e-3 and reducing the learning rate by a factor of 2 every 200k steps. Total number of steps was 700k with 512 batch size. It took around 3 days on RTX 4090.</p>\n<p>Finally, with all changes and scaling of the model we were able to achieve a <strong>0.990</strong> score on LB. </p>\n<h2>Model stacking ensemble</h2>\n<p>In the last few days of competition we had two models and no time to train them any further. We looked previously at ensembling them but results were not promising. We made some analysis and found that the predictions are sometimes very different from both models. If we could pick the correct model perfectly the lower bound was around 0.810. </p>\n<p>Two days before the end of competition we dumped the embeddings from our best Transformer (last common layer before two categorical heads) and GraphNet (layer before predicting direction) for the first 10 data batches. At first we fed both of those embeddings along with model predictions into XGBoost. It was able to achieve 57% accuracy which resulted in <strong>0.976</strong> score on LB (and 0.978 on our internal validation).</p>\n<p>Later on we coded up a simple 3 layer Bi-directional LSTM as a model selector. Our thinking was that having a different architecture (not Transformer or GraphNet) as a selector will be beneficial in learning some different features from data resulting in better selections.</p>\n<p>LSTMs hidden states were initialized with concatenated Transformer and GraphNet embeddings and subset of pulses were given as input sequence*. We thought that training it on event data will help it make better decisions. We trained the model for around 9 epochs (keep in mind it’s only on the first 10 batches) which took about 8h. Doing this improved accuracy to 58% and resulted in a score of <strong>0.974</strong> which was our final result.</p>\n<p>(*) We also experimented with concatenating the embeddings to the input sequence or just before the last classification layer, but both of those approaches turned out to be worse than just initializing hidden states with embeddings.</p>\n<p>Overall architecture of our final submission looked like this:</p>\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/2230884/19025/model-idea.png\" alt=\"model diagram\"></p>\n<h2>Things that we tried</h2>\n<ul>\n<li>We trained two Transformer models on different subset of dataset and tried simple ensembling (adding logits/softmaxes and argmaxing) but this resulted only in marginal improvements of 0.002</li>\n<li>Training a separate categorical heads (like in case of transformer) on GraphNet embeddings and then ensemble both models by simply adding logits/softmaxes and picking argmax - this slightly reduced the error but was not spectacular</li>\n<li>Classification task for GraphNet didn’t work that great, neither have azimuth and zenith regression. Direction regression is what performed best.</li>\n<li>At the beginning of the competition we found a paper “A Convolutional Neural Network based Cascade Reconstruction for the IceCube Neutrino Observatory” and implemented a variant of a 3d convolutional network with so called hexagonal convolutions, which are supposed to be localized in the DOMs space. However, we couldn’t get it to work and get the mean angular error below 1.2. A likely reason is that the data in the competition was quite different from the data used in the paper. The latter had great resolution in the time dimension, because the pulses were collected every 2ns, while our data was much more sparse.</li>\n</ul>\n<h2>Inference code</h2>\n<p><a href=\"https://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer\" target=\"_blank\">https://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer</a></p>",
      "rawMarkdown": "We'd like to thank the organizers and Kaggle staff for putting together this enthralling challenge. As we share the scientific feast we've cooked up in our \"Big Bento Box of Science,\" please accept our apologies for the slightly delayed write-up. The competition was very intense, with our team initially lagging over 400 spots away from the medal zone just two weeks before the end. Through hard work and persistence, we managed to improve our standing, ultimately finishing in 5th place and achieving our first gold medal.\n\n\nSummary:\n-----\n- Two classification heads with 256 bins each for transformer, smart label smoothing -> **1.003**\n- Increasing the number of features GraphNet uses in KNN graph builder, changing activation functions to LayerNorm, overall increasing model size, predicting `(x, y, z)` direction instead of azimuth and zenith directly -> **0.990**\n- LSTM model selector which took the input data as well as predictions from Transformer and GraphNet, and returned prediction which of these models to use -> **0.974**\n\n\nTransformer\n-----\nWe used standard encoder architecture with 8 layers, 512 model size and 8 attention heads. The input consisted of `(x, y, z, t, charge, auxiliary, is_core, rank)`. `is_core` was a 0/1 value indicating whether the sensor was part of a deep core or not. `rank` was computed the same way as in the [following great notebook with LSTM](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference). For features like `(x, y, z, t, charge)` we used a few linear layers to preprocess them before passing them to the Transformer. While `(auxiliary, is_core, rank)` due to their discrete nature were embedded.\n\nWe had two output heads for both azimuth and zenith. Each head predicted 256 values that represented bins. We used equally spaced bins and trained them using cross entropy with smoothed labels. The labels were smoothed so that neighboring bins also have some probability (not like in standard label smoothing where remaining probability is equally assigned to all classes). For azimuth this smoothing was wrapped around in circular fashion, so that predicting \\\\(0\\\\) and \\\\(2 \\pi \\\\) was equivalent. Prediction was just an argmax over the bins. We experimented with a different number of bins, ranging from 16 to 256, and more bins resulted in better performance. However, the difference between 128 and 256 was not significant and we chose to use 256. \n\nWe trained the model for 760k steps on TPU with batch size of 512 and sequence length of 256, so around 3 epochs (which is not much). We used a cosine schedule for LR with peak LR=3e-5 and repeated the cycle every 10k steps (looking back at it the cycle should probably be longer, but at the beginning we weren’t sure what to expect from the training). For the last 160k steps we changed peak LR to 3e-6. Whole training took around 26h. After 760k steps we could still see some slight improvements as the training progressed but decided to focus more on other ideas. This led to a model with a score of **1.003** on public LB.\n\n\nGraphNet\n-----\nAs a starting point we’ve used [standard GraphNet model](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) with [additional input features from early sharing price notebook](https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046).\n\nModel input was composed of `(x, y, z, t, charge, auxiliary, quantum_effeciency, wrong_charge_dom, ice_absorption, ice_scatter)` where `wrong_charge_dom` was a marking of one sensor (num. 1229) because our analysis shown there might be something wrong with it. Output of the model was a direction. \n\nTo achieve the final LB score we've made a few changes to the model.\n\nFirst of all, after looking through the code of `DynEdge` we saw that `features_subset` argument (latent feature subset for computing nearest neighbors in `DynEdge`) is applied not only on the first layer, but also on inputs to other layers. Intuitively it felt limiting, we changed it so that `features_subset` for the non-first layer was variable. In the end we set it to 48 (bigger seems to work better). With this change our model achieved **0.998** on LB.\n\nNext big change was moving from `BatchNorm1d` to `LayerNorm`. This decision was guided by a big difference in MAE metrics we observed between training and validation. Even the standard model (with `BatchNorm1d`) got better results without switching to eval mode, so we decided to move to `LayerNorm` that does not differ between training and evaluation.\n \nLast but not least, we experimented and tuned the whole training procedure. We used AdaBelief optimizer. Unlike the Transformer, with GraphNet model we took the approach of changing the learning rate in steps, starting from 1e-3 and reducing the learning rate by a factor of 2 every 200k steps. Total number of steps was 700k with 512 batch size. It took around 3 days on RTX 4090.\n\nFinally, with all changes and scaling of the model we were able to achieve a **0.990** score on LB. \n\n\nModel stacking ensemble\n-----\nIn the last few days of competition we had two models and no time to train them any further. We looked previously at ensembling them but results were not promising. We made some analysis and found that the predictions are sometimes very different from both models. If we could pick the correct model perfectly the lower bound was around 0.810. \n\nTwo days before the end of competition we dumped the embeddings from our best Transformer (last common layer before two categorical heads) and GraphNet (layer before predicting direction) for the first 10 data batches. At first we fed both of those embeddings along with model predictions into XGBoost. It was able to achieve 57% accuracy which resulted in **0.976** score on LB (and 0.978 on our internal validation).\n\nLater on we coded up a simple 3 layer Bi-directional LSTM as a model selector. Our thinking was that having a different architecture (not Transformer or GraphNet) as a selector will be beneficial in learning some different features from data resulting in better selections.\n\nLSTMs hidden states were initialized with concatenated Transformer and GraphNet embeddings and subset of pulses were given as input sequence*. We thought that training it on event data will help it make better decisions. We trained the model for around 9 epochs (keep in mind it’s only on the first 10 batches) which took about 8h. Doing this improved accuracy to 58% and resulted in a score of **0.974** which was our final result.\n\n(*) We also experimented with concatenating the embeddings to the input sequence or just before the last classification layer, but both of those approaches turned out to be worse than just initializing hidden states with embeddings.\n\nOverall architecture of our final submission looked like this:\n\n![model diagram](https://storage.googleapis.com/kaggle-forum-message-attachments/2230884/19025/model-idea.png)\n\nThings that we tried\n-----\n- We trained two Transformer models on different subset of dataset and tried simple ensembling (adding logits/softmaxes and argmaxing) but this resulted only in marginal improvements of 0.002\n- Training a separate categorical heads (like in case of transformer) on GraphNet embeddings and then ensemble both models by simply adding logits/softmaxes and picking argmax - this slightly reduced the error but was not spectacular\n- Classification task for GraphNet didn’t work that great, neither have azimuth and zenith regression. Direction regression is what performed best.\n- At the beginning of the competition we found a paper “A Convolutional Neural Network based Cascade Reconstruction for the IceCube Neutrino Observatory” and implemented a variant of a 3d convolutional network with so called hexagonal convolutions, which are supposed to be localized in the DOMs space. However, we couldn’t get it to work and get the mean angular error below 1.2. A likely reason is that the data in the competition was quite different from the data used in the paper. The latter had great resolution in the time dimension, because the pulses were collected every 2ns, while our data was much more sparse.\n\n\nInference code\n-----\nhttps://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer",
      "votes": null
    },
    {
      "id": "2231726",
      "postDate": "04/23/2023 15:26:17",
      "content": "<p>What a great result, combining those two models effectively.  When you trained the LSTM selector, was it simply a sigmoid output, trained against a target of 0=Graphnet better, 1=Transformer better?</p>",
      "rawMarkdown": "What a great result, combining those two models effectively.  When you trained the LSTM selector, was it simply a sigmoid output, trained against a target of 0=Graphnet better, 1=Transformer better?",
      "votes": null
    },
    {
      "id": "2231770",
      "postDate": "04/23/2023 16:29:24",
      "content": "<p>Yes, it was just a single output with sigmoid activation. For the loss we used binary cross entropy. Our experiments with softening the loss, based on the difference in MAE between two models, turned out worse than just the basic approach.</p>",
      "rawMarkdown": "Yes, it was just a single output with sigmoid activation. For the loss we used binary cross entropy. Our experiments with softening the loss, based on the difference in MAE between two models, turned out worse than just the basic approach.",
      "votes": null
    },
    {
      "id": "2232119",
      "postDate": "04/24/2023 03:36:08",
      "content": "<p>Thanks for sharing, the lstm selector is really a neat idea. </p>\n<p>Interesting that you have a decently large transformer yet the score is less than what I would expect. Notably, the layers are less, I think that might have made a difference. For this data more layers was more important than embedding dim.</p>",
      "rawMarkdown": "Thanks for sharing, the lstm selector is really a neat idea. \n\nInteresting that you have a decently large transformer yet the score is less than what I would expect. Notably, the layers are less, I think that might have made a difference. For this data more layers was more important than embedding dim.",
      "votes": null
    },
    {
      "id": "2232771",
      "postDate": "04/24/2023 15:41:55",
      "content": "<p>Using LSTM to build the ensemble is very interesting approach, and it seems nobody else did this. </p>\n<p>Regarding the CNN paper, we also tried this way at the beginning with no good results. I think one of the reasons is indeed the dataset difference, the paper dealt specifically with cascade events, which means higher energies and more pulses than what we worked with. </p>",
      "rawMarkdown": "Using LSTM to build the ensemble is very interesting approach, and it seems nobody else did this. \n\nRegarding the CNN paper, we also tried this way at the beginning with no good results. I think one of the reasons is indeed the dataset difference, the paper dealt specifically with cascade events, which means higher energies and more pulses than what we worked with.",
      "votes": null
    },
    {
      "id": "2232843",
      "postDate": "04/24/2023 16:59:11",
      "content": "<p>Good to know that we were not the only ones who didn't make the CNN work 😅</p>",
      "rawMarkdown": "Good to know that we were not the only ones who didn't make the CNN work 😅",
      "votes": null
    },
    {
      "id": "2246289",
      "postDate": "05/05/2023 03:42:08",
      "content": "<p>Congratulations on the gold medal and thanks for sharing!</p>\n<p>I also tried CNNs without much success, though my attempts were limited to 100 batches (the Kaggle weekly GPU quota runs out very quickly, and I focused on using it to try new ideas).</p>\n<p>Using the embeddings of all L1 models in the L2 model was a great idea. My L1 models had good diversity and by picking the best one for each event would lead to a 0.6x score. I tried to leverage that using various architectures (LGBM, MLP, and LSTM), feature combinations (oof predictions, kappa, raw data, extra features, event statistics, embeddings of best model), and targets (classification for best model, regression for direction, etc.), but I was unable to get results significantly better than a simple optimization of the oofs. I ended up using only the latter for an improvement of 0.01 versus my best model, because it was simpler and just slightly worse (1e-5, easily within the margin of error).</p>\n<p>I expected that the embeddings would provide valuable features for the L2 classifier/regressor, but assumed that using just those of the best model would be enough. It makes sense that the L2 model benefits from having all embeddings. </p>",
      "rawMarkdown": "Congratulations on the gold medal and thanks for sharing!\n\nI also tried CNNs without much success, though my attempts were limited to 100 batches (the Kaggle weekly GPU quota runs out very quickly, and I focused on using it to try new ideas).\n\nUsing the embeddings of all L1 models in the L2 model was a great idea. My L1 models had good diversity and by picking the best one for each event would lead to a 0.6x score. I tried to leverage that using various architectures (LGBM, MLP, and LSTM), feature combinations (oof predictions, kappa, raw data, extra features, event statistics, embeddings of best model), and targets (classification for best model, regression for direction, etc.), but I was unable to get results significantly better than a simple optimization of the oofs. I ended up using only the latter for an improvement of 0.01 versus my best model, because it was simpler and just slightly worse (1e-5, easily within the margin of error).\n\nI expected that the embeddings would provide valuable features for the L2 classifier/regressor, but assumed that using just those of the best model would be enough. It makes sense that the L2 model benefits from having all embeddings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2231726,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "04/23/2023 15:26:17",
      "content": "<p>What a great result, combining those two models effectively.  When you trained the LSTM selector, was it simply a sigmoid output, trained against a target of 0=Graphnet better, 1=Transformer better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2231770,
          "author_name": "grzego",
          "author_url": "",
          "post_date": "04/23/2023 16:29:24",
          "content": "<p>Yes, it was just a single output with sigmoid activation. For the loss we used binary cross entropy. Our experiments with softening the loss, based on the difference in MAE between two models, turned out worse than just the basic approach.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2232119,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "04/24/2023 03:36:08",
      "content": "<p>Thanks for sharing, the lstm selector is really a neat idea. </p>\n<p>Interesting that you have a decently large transformer yet the score is less than what I would expect. Notably, the layers are less, I think that might have made a difference. For this data more layers was more important than embedding dim.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232771,
      "author_name": "alexz0",
      "author_url": "",
      "post_date": "04/24/2023 15:41:55",
      "content": "<p>Using LSTM to build the ensemble is very interesting approach, and it seems nobody else did this. </p>\n<p>Regarding the CNN paper, we also tried this way at the beginning with no good results. I think one of the reasons is indeed the dataset difference, the paper dealt specifically with cascade events, which means higher energies and more pulses than what we worked with. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2232843,
          "author_name": "gabfil",
          "author_url": "",
          "post_date": "04/24/2023 16:59:11",
          "content": "<p>Good to know that we were not the only ones who didn't make the CNN work 😅</p>",
          "votes": null,
          "replies": [
            {
              "id": 2246289,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "05/05/2023 03:42:08",
              "content": "<p>Congratulations on the gold medal and thanks for sharing!</p>\n<p>I also tried CNNs without much success, though my attempts were limited to 100 batches (the Kaggle weekly GPU quota runs out very quickly, and I focused on using it to try new ideas).</p>\n<p>Using the embeddings of all L1 models in the L2 model was a great idea. My L1 models had good diversity and by picking the best one for each event would lead to a 0.6x score. I tried to leverage that using various architectures (LGBM, MLP, and LSTM), feature combinations (oof predictions, kappa, raw data, extra features, event statistics, embeddings of best model), and targets (classification for best model, regression for direction, etc.), but I was unable to get results significantly better than a simple optimization of the oofs. I ended up using only the latter for an improvement of 0.01 versus my best model, because it was simpler and just slightly worse (1e-5, easily within the margin of error).</p>\n<p>I expected that the embeddings would provide valuable features for the L2 classifier/regressor, but assumed that using just those of the best model would be enough. It makes sense that the L2 model benefits from having all embeddings. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2230884": "We'd like to thank the organizers and Kaggle staff for putting together this enthralling challenge. As we share the scientific feast we've cooked up in our \"Big Bento Box of Science,\" please accept our apologies for the slightly delayed write-up. The competition was very intense, with our team initially lagging over 400 spots away from the medal zone just two weeks before the end. Through hard work and persistence, we managed to improve our standing, ultimately finishing in 5th place and achieving our first gold medal.\n\n\nSummary:\n-----\n- Two classification heads with 256 bins each for transformer, smart label smoothing -> **1.003**\n- Increasing the number of features GraphNet uses in KNN graph builder, changing activation functions to LayerNorm, overall increasing model size, predicting `(x, y, z)` direction instead of azimuth and zenith directly -> **0.990**\n- LSTM model selector which took the input data as well as predictions from Transformer and GraphNet, and returned prediction which of these models to use -> **0.974**\n\n\nTransformer\n-----\nWe used standard encoder architecture with 8 layers, 512 model size and 8 attention heads. The input consisted of `(x, y, z, t, charge, auxiliary, is_core, rank)`. `is_core` was a 0/1 value indicating whether the sensor was part of a deep core or not. `rank` was computed the same way as in the [following great notebook with LSTM](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference). For features like `(x, y, z, t, charge)` we used a few linear layers to preprocess them before passing them to the Transformer. While `(auxiliary, is_core, rank)` due to their discrete nature were embedded.\n\nWe had two output heads for both azimuth and zenith. Each head predicted 256 values that represented bins. We used equally spaced bins and trained them using cross entropy with smoothed labels. The labels were smoothed so that neighboring bins also have some probability (not like in standard label smoothing where remaining probability is equally assigned to all classes). For azimuth this smoothing was wrapped around in circular fashion, so that predicting \\\\(0\\\\) and \\\\(2 \\pi \\\\) was equivalent. Prediction was just an argmax over the bins. We experimented with a different number of bins, ranging from 16 to 256, and more bins resulted in better performance. However, the difference between 128 and 256 was not significant and we chose to use 256. \n\nWe trained the model for 760k steps on TPU with batch size of 512 and sequence length of 256, so around 3 epochs (which is not much). We used a cosine schedule for LR with peak LR=3e-5 and repeated the cycle every 10k steps (looking back at it the cycle should probably be longer, but at the beginning we weren’t sure what to expect from the training). For the last 160k steps we changed peak LR to 3e-6. Whole training took around 26h. After 760k steps we could still see some slight improvements as the training progressed but decided to focus more on other ideas. This led to a model with a score of **1.003** on public LB.\n\n\nGraphNet\n-----\nAs a starting point we’ve used [standard GraphNet model](https://www.kaggle.com/code/rasmusrse/graphnet-baseline-submission) with [additional input features from early sharing price notebook](https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046).\n\nModel input was composed of `(x, y, z, t, charge, auxiliary, quantum_effeciency, wrong_charge_dom, ice_absorption, ice_scatter)` where `wrong_charge_dom` was a marking of one sensor (num. 1229) because our analysis shown there might be something wrong with it. Output of the model was a direction. \n\nTo achieve the final LB score we've made a few changes to the model.\n\nFirst of all, after looking through the code of `DynEdge` we saw that `features_subset` argument (latent feature subset for computing nearest neighbors in `DynEdge`) is applied not only on the first layer, but also on inputs to other layers. Intuitively it felt limiting, we changed it so that `features_subset` for the non-first layer was variable. In the end we set it to 48 (bigger seems to work better). With this change our model achieved **0.998** on LB.\n\nNext big change was moving from `BatchNorm1d` to `LayerNorm`. This decision was guided by a big difference in MAE metrics we observed between training and validation. Even the standard model (with `BatchNorm1d`) got better results without switching to eval mode, so we decided to move to `LayerNorm` that does not differ between training and evaluation.\n \nLast but not least, we experimented and tuned the whole training procedure. We used AdaBelief optimizer. Unlike the Transformer, with GraphNet model we took the approach of changing the learning rate in steps, starting from 1e-3 and reducing the learning rate by a factor of 2 every 200k steps. Total number of steps was 700k with 512 batch size. It took around 3 days on RTX 4090.\n\nFinally, with all changes and scaling of the model we were able to achieve a **0.990** score on LB. \n\n\nModel stacking ensemble\n-----\nIn the last few days of competition we had two models and no time to train them any further. We looked previously at ensembling them but results were not promising. We made some analysis and found that the predictions are sometimes very different from both models. If we could pick the correct model perfectly the lower bound was around 0.810. \n\nTwo days before the end of competition we dumped the embeddings from our best Transformer (last common layer before two categorical heads) and GraphNet (layer before predicting direction) for the first 10 data batches. At first we fed both of those embeddings along with model predictions into XGBoost. It was able to achieve 57% accuracy which resulted in **0.976** score on LB (and 0.978 on our internal validation).\n\nLater on we coded up a simple 3 layer Bi-directional LSTM as a model selector. Our thinking was that having a different architecture (not Transformer or GraphNet) as a selector will be beneficial in learning some different features from data resulting in better selections.\n\nLSTMs hidden states were initialized with concatenated Transformer and GraphNet embeddings and subset of pulses were given as input sequence*. We thought that training it on event data will help it make better decisions. We trained the model for around 9 epochs (keep in mind it’s only on the first 10 batches) which took about 8h. Doing this improved accuracy to 58% and resulted in a score of **0.974** which was our final result.\n\n(*) We also experimented with concatenating the embeddings to the input sequence or just before the last classification layer, but both of those approaches turned out to be worse than just initializing hidden states with embeddings.\n\nOverall architecture of our final submission looked like this:\n\n![model diagram](https://storage.googleapis.com/kaggle-forum-message-attachments/2230884/19025/model-idea.png)\n\nThings that we tried\n-----\n- We trained two Transformer models on different subset of dataset and tried simple ensembling (adding logits/softmaxes and argmaxing) but this resulted only in marginal improvements of 0.002\n- Training a separate categorical heads (like in case of transformer) on GraphNet embeddings and then ensemble both models by simply adding logits/softmaxes and picking argmax - this slightly reduced the error but was not spectacular\n- Classification task for GraphNet didn’t work that great, neither have azimuth and zenith regression. Direction regression is what performed best.\n- At the beginning of the competition we found a paper “A Convolutional Neural Network based Cascade Reconstruction for the IceCube Neutrino Observatory” and implemented a variant of a 3d convolutional network with so called hexagonal convolutions, which are supposed to be localized in the DOMs space. However, we couldn’t get it to work and get the mean angular error below 1.2. A likely reason is that the data in the competition was quite different from the data used in the paper. The latter had great resolution in the time dimension, because the pulses were collected every 2ns, while our data was much more sparse.\n\n\nInference code\n-----\nhttps://www.kaggle.com/grzego/inference-lstm-ensembler-graphnet-transformer",
    "2231726": "What a great result, combining those two models effectively.  When you trained the LSTM selector, was it simply a sigmoid output, trained against a target of 0=Graphnet better, 1=Transformer better?",
    "2231770": "Yes, it was just a single output with sigmoid activation. For the loss we used binary cross entropy. Our experiments with softening the loss, based on the difference in MAE between two models, turned out worse than just the basic approach.",
    "2232119": "Thanks for sharing, the lstm selector is really a neat idea. \n\nInteresting that you have a decently large transformer yet the score is less than what I would expect. Notably, the layers are less, I think that might have made a difference. For this data more layers was more important than embedding dim.",
    "2232771": "Using LSTM to build the ensemble is very interesting approach, and it seems nobody else did this. \n\nRegarding the CNN paper, we also tried this way at the beginning with no good results. I think one of the reasons is indeed the dataset difference, the paper dealt specifically with cascade events, which means higher energies and more pulses than what we worked with.",
    "2232843": "Good to know that we were not the only ones who didn't make the CNN work 😅",
    "2246289": "Congratulations on the gold medal and thanks for sharing!\n\nI also tried CNNs without much success, though my attempts were limited to 100 batches (the Kaggle weekly GPU quota runs out very quickly, and I focused on using it to try new ideas).\n\nUsing the embeddings of all L1 models in the L2 model was a great idea. My L1 models had good diversity and by picking the best one for each event would lead to a 0.6x score. I tried to leverage that using various architectures (LGBM, MLP, and LSTM), feature combinations (oof predictions, kappa, raw data, extra features, event statistics, embeddings of best model), and targets (classification for best model, regression for direction, etc.), but I was unable to get results significantly better than a simple optimization of the oofs. I ended up using only the latter for an improvement of 0.01 versus my best model, because it was simpler and just slightly worse (1e-5, easily within the margin of error).\n\nI expected that the embeddings would provide valuable features for the L2 classifier/regressor, but assumed that using just those of the best model would be enough. It makes sense that the L2 model benefits from having all embeddings."
  },
  "source": "meta"
}