{
  "id": 402888,
  "title": "3rd Place - Attention + XGBoost Ensembler",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/gpus-on-ice-3rd-place-attention-xgboost-ensembler",
  "author_name": "",
  "post_date": "2023-04-21T10:03:50.257Z",
  "votes": 56,
  "comment_count": 25,
  "views": 0,
  "content": "<h1>3rd place Solution writeup - Attention + XGBoost</h1>\n<p>Participated on Kaggle after 3 years, first gold for me, had a really exciting time. 😃</p>\n<p><strong>Check the TL;Dr at the end for a short summary</strong> - The post is long with a mixture of model information and personal lessons.</p>\n<h2>The bitter lesson - Attention is all you need</h2>\n<p>This data is not a natural graph, I'm confused by all the graph models being trained on it. The graph edges are artificially selected based on criteria such as time or distance, which I didn't think was the right approach. Hence I spent my time on LSTM and later moved to Self-Attention.</p>\n<h3>Base Model</h3>\n<p>The base model is a simple <strong>self-attention model</strong>. Most of the model code is from Andrej Karpathy's NanoGPT <a href=\"https://github.com/karpathy/nanoGPT.git\" target=\"_blank\">https://github.com/karpathy/nanoGPT.git</a></p>\n<p>The inputs are normalized pulse data with transparency information from Datasaurus' Dynedge baseline <a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046</a> - Sequences are sampled upto a max length of 256, prioritizing sampling of the pulses where <code>auxiliary=False</code></p>\n<p>The outputs are <strong>two classification heads</strong> for each angle with 128 bins each. - I train it with a custom loss that locally smoothens the one hot target vectors, more on that below.</p>\n<p>A small version of the model can be trained on an RTX 3060 to give a score close to ~1.02 in 1 hour. (Embedding size 128 - 6 layers).</p>\n<h3>The bitter lesson - Scale beats everything</h3>\n<p>The competition was undoubtedly compute-heavy, but I only had access to my home PC with an RTX 3060. Naturally, I spent a lot of time hand-designing features, optimizing the code, and trying ideas that felt like they should help. Only to have to throw out most of the experiments which help a smaller model but get washed out when training a big model with more data.</p>\n<p>Notable ideas that helped but later got discarded:</p>\n<ol>\n<li>Simple augmentation like moving the points slightly, changing the charges etc.</li>\n<li>Supervised contrastive loss style training to predict the angle between two events.</li>\n<li>Mutliple pooling instead of just average pooling at the end. (in fact this caused overfitting)</li>\n<li>Angle classifier on sphere coorindates instead of two classifiers.</li>\n</ol>\n<p>All of the above individually helped get better performance on a small model trained for 2-3 hours, but reach the same performance with bigger models. There is no overfitting in the models even with 20 batches of training data.</p>\n<p>The real improvements in score came from scaling the model. Training on more data was crucial as well, the final models I use train on 650 batches of data. They finally start to overfit after 4 cycles through the entire data. </p>\n<p>All the wasted days on experimentation reminds me of Richard Sutton's blog on \"The Bitter Lesson\". <a href=\"http://www.incompleteideas.net/IncIdeas/BitterLesson.html\" target=\"_blank\">http://www.incompleteideas.net/IncIdeas/BitterLesson.html</a> - Haha GPUs go bitterrrrrrr.</p>\n<h3>Scaling and engineering tricks</h3>\n<p>At the start of April, I decided to buy a new GPU just seeing how much compute is needed. Upgraded to RTX 4080, which is 3x faster. (No this is not the only scaling I did 😅, but more compute was crucial)</p>\n<p>Final submitted model has 18 self attention layers, embedding size 512. This reaches a LB score of 0.982, and I'm sure training a bigger model will get even better. I trained the same model with 15 layers, but had worse performance due to mixed precision instability.</p>\n<p>Engineering tricks</p>\n<ol>\n<li>Mixed precision training is much faster, I train with <strong>FP16 with FlashAttention</strong>. However this started getting unstable after reaching a score of 0.990, so I switched to FP32 after that. Only the bigger models showed this instability and I didn't get to investigate how to continue on FP16, probably the scales of the inputs or the loss can be managed.</li>\n<li>The sequence lengths in the dataset has high variance. After sampling the data upto 256 pulses, the average length of each event is close of 100. This means a lot of attention processing (when not using FlashAttention) is wasted on padding. To speed up training I <strong>group the batches by event length</strong>. This gives a 1.5x speedup over FlashAttention and <strong>3.5x speedup</strong> over normal attention implementation. Packing also speeds up inference in the same ratios.</li>\n<li>Loading a portion of the training data for each \"epoch\". This is straightforward, switch out the bathces in the dataloader after each epoch, or do it as a background process for prefetch.</li>\n</ol>\n<p>Scaling worked incredibly well:</p>\n<ol>\n<li>128 Embedding - 9 Layers - 200 batches - Score ~1.001</li>\n<li>256 Embedding - 12 Layers - 400 batches - Score ~0.990</li>\n<li>512 Embedding - 18 Layers - 650 batches - Score ~0.982 (Some overfitting at the end)</li>\n</ol>\n<p>Further, I fine tuned the model for sequence lengths upto 3072, which gives a slight imporvement of 0.002.</p>\n<h3>Smoothed cross entropy loss</h3>\n<p>The Von Mises-Fisher Loss proved unstable with Attention model, and I never got it to go beyond 1.04 when training end to end.</p>\n<p>For the classifier I digitize the angles to 128 bins. Cross entropy loss works well, but I wanted to introduce inductive bias in the model that the clases are ordered, and not independent. To do so, I apply <strong>1D convolution with a gaussian kernel</strong> to the target one hot vectors. For azimuth this covolution has wraparound, for zenith I extend the bins based on the length of the kernel. (Notebook explaining the loss will be released later). This gives slightly more stable training, though with good hyperparam tuning even cross entropy might work just as well.</p>\n<h2>The sweet ending - Attention is not all you need</h2>\n<p>Obviously attention was not enough to win, 0.982 isn't even gold zone (cries in corner). However, the base model was cruicial for the next ideas to work.</p>\n<h3>Stack vMF model</h3>\n<p>Training an MLP with vMF loss on top of the encoder embeddings works. I dump 100 batches of average pooled encoder embeddings and train the stack model with it, achieves a score of ~0.978. Simple average ensemble of vMF and base model gives a score ~0.976.</p>\n<h3>The magic - Ensembling using XGBoost 🔥</h3>\n<p>I know the standard ensembling tricks, but this is the first time I have ensembled using another model. If you are aware of this used in practice, <strong>please let me know if there is literature on this</strong>. I got the idea just 3 days before the end of the competition, barely got to test it properly, yet it worked like magic.</p>\n<p>My hypothesis on why this worked - The classifier and vMF models are far from perfect scores, and have high disagreement. Here's a histogram of the zenith angles.</p>\n<p>Both the model have its own biases, and its not obvious how to combine them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3468635%2Fd4807872db2ad427f9484a1f6ac8ccbd%2Fzenith_histogram.png?generation=1681967255551488&amp;alt=media\" alt=\"\"></p>\n<p>First I tried a simple voting rule - If they're too far then don't average:<br>\nThis already bumped the score from 0.976 -&gt; 0.973, already a top 5 solution.</p>\n<pre><code>avg_zenith = (zn_stack + zn_base) / \nvote_zenith = avg_zenith.clone()\nvoter = torch.(zn_stack - zn_base) &gt; \nvote_zenith[voter] = zn_stack[voter].clone()\n</code></pre>\n<p>Both the models have some way of indicating \"confidence\" in prediction. Probabilty of bin for classifier, and kappa for vMF. I played around with some thresholding for those, along with event lengths and z coordinates, and it bumped score to 0.9715, so I decided to try XGBoost. On the first try the score went straight to 0.969. 🚀</p>\n<p>Further, I tried another boosting classifier to \"predict which predictions are bad\". And used it to flip the angles, this also gives a slight improvement in score.</p>\n<p>Adding some more simple features the final score for the 18 layer model was 0.965 (I made a submission with the 15 layer model which got 0.967).</p>\n<p>I used sklearn's HistGradientBoostingClassifier for the final setup, but XGBoost has similar results.</p>\n<h3>Inference timing stat</h3>\n<p>The inference time for the base model is around 4 minutes per batch, and another 4 minutes for the long sequence model, which only gives a small boost. So its possible to get a score of <strong>0.967 within 25 minutes</strong>.</p>\n<h3>Final submission</h3>\n<p>The final submission has 3 models. The 18 layer, 15 layer, and another checkpoint from same 18 layer run.</p>\n<p>Another boosting classifier to merge 18 layer and 15 layer. And then simple voting to merge the 3.</p>\n<p>Here's the inference notebooks for final submission and single attention model.</p>\n<ol>\n<li>Final - <a href=\"https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler</a></li>\n<li>Attention only - <a href=\"https://www.kaggle.com/code/dipamc77/single-attention-model-inference\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/single-attention-model-inference</a></li>\n</ol>\n<h2>Summary of learnings</h2>\n<ol>\n<li>Test the model for scaling, especially if its underfitting, before trying feature engineering/other ideas.</li>\n<li>Spend the initial engineering effort to optimize the pipeline, to test ideas fast.</li>\n<li>If models have high disagreement and not close to max score, try to train a classifier to ensemble to models.</li>\n<li>Even though the final improvement came from ensembling tricks, scaling was the most important. Bigger models will likely give even better results.</li>\n<li>Don't give up, I really threw everything in the kitchen sink to improve scores at the end!</li>\n<li>Buy more GPUs. 💸</li>\n</ol>\n<p>In hindsight, I should have tried some of the graph models as they might have disagreed with the classifier, and ensembling them might give a good improvement.</p>\n<h2>Notable Failed ideas</h2>\n<p>Didn't expect these to fail, and didn't abandon prematurely, but couldn't make them work.</p>\n<ol>\n<li><p>Finding noise in the dataset to ignore during training.</p></li>\n<li><p>vMF loss training with attention model.</p>\n<ul>\n<li>Edit - From reading the 2nd place solution, this might be because I didn't encode the inputs properly. Perhaps it was lack of understanding on my part. I thought encoding continuous variables with a normal FC layer is good enough. Clearly it did work with the classification loss though.</li></ul></li>\n<li><p>Reducing bias in zenith predictions by class weightage and oversampling.</p></li>\n</ol>\n<h2>TL;Dr - Final Solution summary</h2>\n<p>Pulses as tokens, sequence length upto 256, priority sampling for \"real\" pulses.</p>\n<ul>\n<li>(Self Attention (512 embedding, 18 layer, Avg pool, 2048 FF) + 2 Angle classifier heads) -&gt; ~0.982</li>\n<li>(Fine tuned attention model for sequence length upto 3072) -&gt; 0.980</li>\n<li>(MLP with vMF loss on attention model embeddings) -&gt; ~0.976</li>\n<li>(Combine classifier and vMF angles with boosting classifier) -&gt; ~0.965 🔥</li>\n<li>(15 layer + 18 layer + 1 checkpoint from same 18 layer run - combine with another boosting classifier) -&gt; ~0.962</li>\n</ul>\n<hr>\n<p>P.S - Doesn't matter what your rank is, please share your learnings and also failed ideas. It helps you organize what you learnt and get better insights, I'll make sure to read all of them.</p>\n<p><strong>Training Code - Please share your feedback.</strong></p>\n<p><a href=\"https://github.com/dipamc/kaggle-icecube-neutrinos\" target=\"_blank\">https://github.com/dipamc/kaggle-icecube-neutrinos</a></p>",
  "messages": [
    {
      "id": "2227888",
      "postDate": "04/20/2023 05:08:18",
      "content": "<h1>3rd place Solution writeup - Attention + XGBoost</h1>\n<p>Participated on Kaggle after 3 years, first gold for me, had a really exciting time. 😃</p>\n<p><strong>Check the TL;Dr at the end for a short summary</strong> - The post is long with a mixture of model information and personal lessons.</p>\n<h2>The bitter lesson - Attention is all you need</h2>\n<p>This data is not a natural graph, I'm confused by all the graph models being trained on it. The graph edges are artificially selected based on criteria such as time or distance, which I didn't think was the right approach. Hence I spent my time on LSTM and later moved to Self-Attention.</p>\n<h3>Base Model</h3>\n<p>The base model is a simple <strong>self-attention model</strong>. Most of the model code is from Andrej Karpathy's NanoGPT <a href=\"https://github.com/karpathy/nanoGPT.git\" target=\"_blank\">https://github.com/karpathy/nanoGPT.git</a></p>\n<p>The inputs are normalized pulse data with transparency information from Datasaurus' Dynedge baseline <a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046\" target=\"_blank\">https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046</a> - Sequences are sampled upto a max length of 256, prioritizing sampling of the pulses where <code>auxiliary=False</code></p>\n<p>The outputs are <strong>two classification heads</strong> for each angle with 128 bins each. - I train it with a custom loss that locally smoothens the one hot target vectors, more on that below.</p>\n<p>A small version of the model can be trained on an RTX 3060 to give a score close to ~1.02 in 1 hour. (Embedding size 128 - 6 layers).</p>\n<h3>The bitter lesson - Scale beats everything</h3>\n<p>The competition was undoubtedly compute-heavy, but I only had access to my home PC with an RTX 3060. Naturally, I spent a lot of time hand-designing features, optimizing the code, and trying ideas that felt like they should help. Only to have to throw out most of the experiments which help a smaller model but get washed out when training a big model with more data.</p>\n<p>Notable ideas that helped but later got discarded:</p>\n<ol>\n<li>Simple augmentation like moving the points slightly, changing the charges etc.</li>\n<li>Supervised contrastive loss style training to predict the angle between two events.</li>\n<li>Mutliple pooling instead of just average pooling at the end. (in fact this caused overfitting)</li>\n<li>Angle classifier on sphere coorindates instead of two classifiers.</li>\n</ol>\n<p>All of the above individually helped get better performance on a small model trained for 2-3 hours, but reach the same performance with bigger models. There is no overfitting in the models even with 20 batches of training data.</p>\n<p>The real improvements in score came from scaling the model. Training on more data was crucial as well, the final models I use train on 650 batches of data. They finally start to overfit after 4 cycles through the entire data. </p>\n<p>All the wasted days on experimentation reminds me of Richard Sutton's blog on \"The Bitter Lesson\". <a href=\"http://www.incompleteideas.net/IncIdeas/BitterLesson.html\" target=\"_blank\">http://www.incompleteideas.net/IncIdeas/BitterLesson.html</a> - Haha GPUs go bitterrrrrrr.</p>\n<h3>Scaling and engineering tricks</h3>\n<p>At the start of April, I decided to buy a new GPU just seeing how much compute is needed. Upgraded to RTX 4080, which is 3x faster. (No this is not the only scaling I did 😅, but more compute was crucial)</p>\n<p>Final submitted model has 18 self attention layers, embedding size 512. This reaches a LB score of 0.982, and I'm sure training a bigger model will get even better. I trained the same model with 15 layers, but had worse performance due to mixed precision instability.</p>\n<p>Engineering tricks</p>\n<ol>\n<li>Mixed precision training is much faster, I train with <strong>FP16 with FlashAttention</strong>. However this started getting unstable after reaching a score of 0.990, so I switched to FP32 after that. Only the bigger models showed this instability and I didn't get to investigate how to continue on FP16, probably the scales of the inputs or the loss can be managed.</li>\n<li>The sequence lengths in the dataset has high variance. After sampling the data upto 256 pulses, the average length of each event is close of 100. This means a lot of attention processing (when not using FlashAttention) is wasted on padding. To speed up training I <strong>group the batches by event length</strong>. This gives a 1.5x speedup over FlashAttention and <strong>3.5x speedup</strong> over normal attention implementation. Packing also speeds up inference in the same ratios.</li>\n<li>Loading a portion of the training data for each \"epoch\". This is straightforward, switch out the bathces in the dataloader after each epoch, or do it as a background process for prefetch.</li>\n</ol>\n<p>Scaling worked incredibly well:</p>\n<ol>\n<li>128 Embedding - 9 Layers - 200 batches - Score ~1.001</li>\n<li>256 Embedding - 12 Layers - 400 batches - Score ~0.990</li>\n<li>512 Embedding - 18 Layers - 650 batches - Score ~0.982 (Some overfitting at the end)</li>\n</ol>\n<p>Further, I fine tuned the model for sequence lengths upto 3072, which gives a slight imporvement of 0.002.</p>\n<h3>Smoothed cross entropy loss</h3>\n<p>The Von Mises-Fisher Loss proved unstable with Attention model, and I never got it to go beyond 1.04 when training end to end.</p>\n<p>For the classifier I digitize the angles to 128 bins. Cross entropy loss works well, but I wanted to introduce inductive bias in the model that the clases are ordered, and not independent. To do so, I apply <strong>1D convolution with a gaussian kernel</strong> to the target one hot vectors. For azimuth this covolution has wraparound, for zenith I extend the bins based on the length of the kernel. (Notebook explaining the loss will be released later). This gives slightly more stable training, though with good hyperparam tuning even cross entropy might work just as well.</p>\n<h2>The sweet ending - Attention is not all you need</h2>\n<p>Obviously attention was not enough to win, 0.982 isn't even gold zone (cries in corner). However, the base model was cruicial for the next ideas to work.</p>\n<h3>Stack vMF model</h3>\n<p>Training an MLP with vMF loss on top of the encoder embeddings works. I dump 100 batches of average pooled encoder embeddings and train the stack model with it, achieves a score of ~0.978. Simple average ensemble of vMF and base model gives a score ~0.976.</p>\n<h3>The magic - Ensembling using XGBoost 🔥</h3>\n<p>I know the standard ensembling tricks, but this is the first time I have ensembled using another model. If you are aware of this used in practice, <strong>please let me know if there is literature on this</strong>. I got the idea just 3 days before the end of the competition, barely got to test it properly, yet it worked like magic.</p>\n<p>My hypothesis on why this worked - The classifier and vMF models are far from perfect scores, and have high disagreement. Here's a histogram of the zenith angles.</p>\n<p>Both the model have its own biases, and its not obvious how to combine them.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3468635%2Fd4807872db2ad427f9484a1f6ac8ccbd%2Fzenith_histogram.png?generation=1681967255551488&amp;alt=media\" alt=\"\"></p>\n<p>First I tried a simple voting rule - If they're too far then don't average:<br>\nThis already bumped the score from 0.976 -&gt; 0.973, already a top 5 solution.</p>\n<pre><code>avg_zenith = (zn_stack + zn_base) / \nvote_zenith = avg_zenith.clone()\nvoter = torch.(zn_stack - zn_base) &gt; \nvote_zenith[voter] = zn_stack[voter].clone()\n</code></pre>\n<p>Both the models have some way of indicating \"confidence\" in prediction. Probabilty of bin for classifier, and kappa for vMF. I played around with some thresholding for those, along with event lengths and z coordinates, and it bumped score to 0.9715, so I decided to try XGBoost. On the first try the score went straight to 0.969. 🚀</p>\n<p>Further, I tried another boosting classifier to \"predict which predictions are bad\". And used it to flip the angles, this also gives a slight improvement in score.</p>\n<p>Adding some more simple features the final score for the 18 layer model was 0.965 (I made a submission with the 15 layer model which got 0.967).</p>\n<p>I used sklearn's HistGradientBoostingClassifier for the final setup, but XGBoost has similar results.</p>\n<h3>Inference timing stat</h3>\n<p>The inference time for the base model is around 4 minutes per batch, and another 4 minutes for the long sequence model, which only gives a small boost. So its possible to get a score of <strong>0.967 within 25 minutes</strong>.</p>\n<h3>Final submission</h3>\n<p>The final submission has 3 models. The 18 layer, 15 layer, and another checkpoint from same 18 layer run.</p>\n<p>Another boosting classifier to merge 18 layer and 15 layer. And then simple voting to merge the 3.</p>\n<p>Here's the inference notebooks for final submission and single attention model.</p>\n<ol>\n<li>Final - <a href=\"https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler</a></li>\n<li>Attention only - <a href=\"https://www.kaggle.com/code/dipamc77/single-attention-model-inference\" target=\"_blank\">https://www.kaggle.com/code/dipamc77/single-attention-model-inference</a></li>\n</ol>\n<h2>Summary of learnings</h2>\n<ol>\n<li>Test the model for scaling, especially if its underfitting, before trying feature engineering/other ideas.</li>\n<li>Spend the initial engineering effort to optimize the pipeline, to test ideas fast.</li>\n<li>If models have high disagreement and not close to max score, try to train a classifier to ensemble to models.</li>\n<li>Even though the final improvement came from ensembling tricks, scaling was the most important. Bigger models will likely give even better results.</li>\n<li>Don't give up, I really threw everything in the kitchen sink to improve scores at the end!</li>\n<li>Buy more GPUs. 💸</li>\n</ol>\n<p>In hindsight, I should have tried some of the graph models as they might have disagreed with the classifier, and ensembling them might give a good improvement.</p>\n<h2>Notable Failed ideas</h2>\n<p>Didn't expect these to fail, and didn't abandon prematurely, but couldn't make them work.</p>\n<ol>\n<li><p>Finding noise in the dataset to ignore during training.</p></li>\n<li><p>vMF loss training with attention model.</p>\n<ul>\n<li>Edit - From reading the 2nd place solution, this might be because I didn't encode the inputs properly. Perhaps it was lack of understanding on my part. I thought encoding continuous variables with a normal FC layer is good enough. Clearly it did work with the classification loss though.</li></ul></li>\n<li><p>Reducing bias in zenith predictions by class weightage and oversampling.</p></li>\n</ol>\n<h2>TL;Dr - Final Solution summary</h2>\n<p>Pulses as tokens, sequence length upto 256, priority sampling for \"real\" pulses.</p>\n<ul>\n<li>(Self Attention (512 embedding, 18 layer, Avg pool, 2048 FF) + 2 Angle classifier heads) -&gt; ~0.982</li>\n<li>(Fine tuned attention model for sequence length upto 3072) -&gt; 0.980</li>\n<li>(MLP with vMF loss on attention model embeddings) -&gt; ~0.976</li>\n<li>(Combine classifier and vMF angles with boosting classifier) -&gt; ~0.965 🔥</li>\n<li>(15 layer + 18 layer + 1 checkpoint from same 18 layer run - combine with another boosting classifier) -&gt; ~0.962</li>\n</ul>\n<hr>\n<p>P.S - Doesn't matter what your rank is, please share your learnings and also failed ideas. It helps you organize what you learnt and get better insights, I'll make sure to read all of them.</p>\n<p><strong>Training Code - Please share your feedback.</strong></p>\n<p><a href=\"https://github.com/dipamc/kaggle-icecube-neutrinos\" target=\"_blank\">https://github.com/dipamc/kaggle-icecube-neutrinos</a></p>",
      "rawMarkdown": "# 3rd place Solution writeup - Attention + XGBoost \n\nParticipated on Kaggle after 3 years, first gold for me, had a really exciting time. 😃\n\n**Check the TL;Dr at the end for a short summary** - The post is long with a mixture of model information and personal lessons.\n\n## The bitter lesson - Attention is all you need\n\nThis data is not a natural graph, I'm confused by all the graph models being trained on it. The graph edges are artificially selected based on criteria such as time or distance, which I didn't think was the right approach. Hence I spent my time on LSTM and later moved to Self-Attention.\n\n### Base Model\n\nThe base model is a simple **self-attention model**. Most of the model code is from Andrej Karpathy's NanoGPT https://github.com/karpathy/nanoGPT.git\n\nThe inputs are normalized pulse data with transparency information from Datasaurus' Dynedge baseline https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046 - Sequences are sampled upto a max length of 256, prioritizing sampling of the pulses where `auxiliary=False`\n\nThe outputs are **two classification heads** for each angle with 128 bins each. - I train it with a custom loss that locally smoothens the one hot target vectors, more on that below.\n\nA small version of the model can be trained on an RTX 3060 to give a score close to ~1.02 in 1 hour. (Embedding size 128 - 6 layers).\n\n### The bitter lesson - Scale beats everything\n\nThe competition was undoubtedly compute-heavy, but I only had access to my home PC with an RTX 3060. Naturally, I spent a lot of time hand-designing features, optimizing the code, and trying ideas that felt like they should help. Only to have to throw out most of the experiments which help a smaller model but get washed out when training a big model with more data.\n\nNotable ideas that helped but later got discarded:\n\n1. Simple augmentation like moving the points slightly, changing the charges etc.\n2. Supervised contrastive loss style training to predict the angle between two events.\n3. Mutliple pooling instead of just average pooling at the end. (in fact this caused overfitting)\n4. Angle classifier on sphere coorindates instead of two classifiers.\n\nAll of the above individually helped get better performance on a small model trained for 2-3 hours, but reach the same performance with bigger models. There is no overfitting in the models even with 20 batches of training data.\n\nThe real improvements in score came from scaling the model. Training on more data was crucial as well, the final models I use train on 650 batches of data. They finally start to overfit after 4 cycles through the entire data. \n\nAll the wasted days on experimentation reminds me of Richard Sutton's blog on \"The Bitter Lesson\". http://www.incompleteideas.net/IncIdeas/BitterLesson.html - Haha GPUs go bitterrrrrrr.\n\n### Scaling and engineering tricks\n\nAt the start of April, I decided to buy a new GPU just seeing how much compute is needed. Upgraded to RTX 4080, which is 3x faster. (No this is not the only scaling I did 😅, but more compute was crucial)\n\nFinal submitted model has 18 self attention layers, embedding size 512. This reaches a LB score of 0.982, and I'm sure training a bigger model will get even better. I trained the same model with 15 layers, but had worse performance due to mixed precision instability.\n\nEngineering tricks\n\n1. Mixed precision training is much faster, I train with **FP16 with FlashAttention**. However this started getting unstable after reaching a score of 0.990, so I switched to FP32 after that. Only the bigger models showed this instability and I didn't get to investigate how to continue on FP16, probably the scales of the inputs or the loss can be managed.\n2. The sequence lengths in the dataset has high variance. After sampling the data upto 256 pulses, the average length of each event is close of 100. This means a lot of attention processing (when not using FlashAttention) is wasted on padding. To speed up training I **group the batches by event length**. This gives a 1.5x speedup over FlashAttention and **3.5x speedup** over normal attention implementation. Packing also speeds up inference in the same ratios.\n3. Loading a portion of the training data for each \"epoch\". This is straightforward, switch out the bathces in the dataloader after each epoch, or do it as a background process for prefetch.\n\nScaling worked incredibly well:\n\n1. 128 Embedding - 9 Layers - 200 batches - Score ~1.001\n2. 256 Embedding - 12 Layers - 400 batches - Score ~0.990\n3. 512 Embedding - 18 Layers - 650 batches - Score ~0.982 (Some overfitting at the end)\n\nFurther, I fine tuned the model for sequence lengths upto 3072, which gives a slight imporvement of 0.002.\n\n### Smoothed cross entropy loss\n\nThe Von Mises-Fisher Loss proved unstable with Attention model, and I never got it to go beyond 1.04 when training end to end.\n\nFor the classifier I digitize the angles to 128 bins. Cross entropy loss works well, but I wanted to introduce inductive bias in the model that the clases are ordered, and not independent. To do so, I apply **1D convolution with a gaussian kernel** to the target one hot vectors. For azimuth this covolution has wraparound, for zenith I extend the bins based on the length of the kernel. (Notebook explaining the loss will be released later). This gives slightly more stable training, though with good hyperparam tuning even cross entropy might work just as well.\n\n## The sweet ending - Attention is not all you need\n\nObviously attention was not enough to win, 0.982 isn't even gold zone (cries in corner). However, the base model was cruicial for the next ideas to work.\n\n### Stack vMF model\n\nTraining an MLP with vMF loss on top of the encoder embeddings works. I dump 100 batches of average pooled encoder embeddings and train the stack model with it, achieves a score of ~0.978. Simple average ensemble of vMF and base model gives a score ~0.976.\n\n### The magic - Ensembling using XGBoost 🔥\n\nI know the standard ensembling tricks, but this is the first time I have ensembled using another model. If you are aware of this used in practice, **please let me know if there is literature on this**. I got the idea just 3 days before the end of the competition, barely got to test it properly, yet it worked like magic.\n\nMy hypothesis on why this worked - The classifier and vMF models are far from perfect scores, and have high disagreement. Here's a histogram of the zenith angles.\n\nBoth the model have its own biases, and its not obvious how to combine them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3468635%2Fd4807872db2ad427f9484a1f6ac8ccbd%2Fzenith_histogram.png?generation=1681967255551488&alt=media)\n\nFirst I tried a simple voting rule - If they're too far then don't average:\nThis already bumped the score from 0.976 -> 0.973, already a top 5 solution.\n\n```python\navg_zenith = (zn_stack + zn_base) / 2\nvote_zenith = avg_zenith.clone()\nvoter = torch.abs(zn_stack - zn_base) > 0.4\nvote_zenith[voter] = zn_stack[voter].clone()\n```\n\nBoth the models have some way of indicating \"confidence\" in prediction. Probabilty of bin for classifier, and kappa for vMF. I played around with some thresholding for those, along with event lengths and z coordinates, and it bumped score to 0.9715, so I decided to try XGBoost. On the first try the score went straight to 0.969. 🚀\n\nFurther, I tried another boosting classifier to \"predict which predictions are bad\". And used it to flip the angles, this also gives a slight improvement in score.\n\nAdding some more simple features the final score for the 18 layer model was 0.965 (I made a submission with the 15 layer model which got 0.967).\n\nI used sklearn's HistGradientBoostingClassifier for the final setup, but XGBoost has similar results.\n\n### Inference timing stat\n\nThe inference time for the base model is around 4 minutes per batch, and another 4 minutes for the long sequence model, which only gives a small boost. So its possible to get a score of **0.967 within 25 minutes**.\n\n### Final submission\n\nThe final submission has 3 models. The 18 layer, 15 layer, and another checkpoint from same 18 layer run.\n\nAnother boosting classifier to merge 18 layer and 15 layer. And then simple voting to merge the 3.\n\nHere's the inference notebooks for final submission and single attention model.\n\n1. Final - https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler\n2. Attention only - https://www.kaggle.com/code/dipamc77/single-attention-model-inference\n\n## Summary of learnings\n\n1. Test the model for scaling, especially if its underfitting, before trying feature engineering/other ideas.\n2. Spend the initial engineering effort to optimize the pipeline, to test ideas fast.\n3. If models have high disagreement and not close to max score, try to train a classifier to ensemble to models.\n4. Even though the final improvement came from ensembling tricks, scaling was the most important. Bigger models will likely give even better results.\n5. Don't give up, I really threw everything in the kitchen sink to improve scores at the end!\n6. Buy more GPUs. 💸\n\nIn hindsight, I should have tried some of the graph models as they might have disagreed with the classifier, and ensembling them might give a good improvement.\n\n## Notable Failed ideas\n\nDidn't expect these to fail, and didn't abandon prematurely, but couldn't make them work.\n\n1. Finding noise in the dataset to ignore during training.\n2. vMF loss training with attention model.\n     - Edit - From reading the 2nd place solution, this might be because I didn't encode the inputs properly. Perhaps it was lack of understanding on my part. I thought encoding continuous variables with a normal FC layer is good enough. Clearly it did work with the classification loss though.\n\n3. Reducing bias in zenith predictions by class weightage and oversampling.\n\n## TL;Dr - Final Solution summary\n\nPulses as tokens, sequence length upto 256, priority sampling for \"real\" pulses.\n\n- (Self Attention (512 embedding, 18 layer, Avg pool, 2048 FF) + 2 Angle classifier heads) -> ~0.982\n- (Fine tuned attention model for sequence length upto 3072) -> 0.980\n- (MLP with vMF loss on attention model embeddings) -> ~0.976\n- (Combine classifier and vMF angles with boosting classifier) -> ~0.965 🔥\n- (15 layer + 18 layer + 1 checkpoint from same 18 layer run - combine with another boosting classifier) -> ~0.962\n\n-----------------\n\nP.S - Doesn't matter what your rank is, please share your learnings and also failed ideas. It helps you organize what you learnt and get better insights, I'll make sure to read all of them.\n\n**Training Code - Please share your feedback.**\n\nhttps://github.com/dipamc/kaggle-icecube-neutrinos",
      "votes": null
    },
    {
      "id": "2227950",
      "postDate": "04/20/2023 06:21:45",
      "content": "<p>Thanks for sharing! I heard the magic of Ensembling using XGBoost for the first time! This is very valuable to learn<br>\nLooking forward to your train code.</p>\n<p>(going to sleep before exploring the code</p>",
      "rawMarkdown": "Thanks for sharing! I heard the magic of Ensembling using XGBoost for the first time! This is very valuable to learn\nLooking forward to your train code.\n\n(going to sleep before exploring the code",
      "votes": null
    },
    {
      "id": "2227963",
      "postDate": "04/20/2023 06:33:31",
      "content": "<p>This is awesome, self attention is one of the first things I went to try but could not get it to work so nice to see you solved it, thanks for the write up!</p>\n<p>Congratulations on solo gold and (soon) upgrading to Kaggle Master 😄</p>",
      "rawMarkdown": "This is awesome, self attention is one of the first things I went to try but could not get it to work so nice to see you solved it, thanks for the write up!\n\nCongratulations on solo gold and (soon) upgrading to Kaggle Master 😄",
      "votes": null
    },
    {
      "id": "2227977",
      "postDate": "04/20/2023 07:02:22",
      "content": "<p>Congratulations on the solo gold! We were all amazed by your progress and hoping you would get the win!</p>",
      "rawMarkdown": "Congratulations on the solo gold! We were all amazed by your progress and hoping you would get the win!",
      "votes": null
    },
    {
      "id": "2227989",
      "postDate": "04/20/2023 07:26:02",
      "content": "<p>Thank you so much for the kind words.</p>\n<p>A week ago I thought the max score I can get is probably around 0.973, the boosting model ensembler really did the magic!   </p>",
      "rawMarkdown": "Thank you so much for the kind words.\n\nA week ago I thought the max score I can get is probably around 0.973, the boosting model ensembler really did the magic!",
      "votes": null
    },
    {
      "id": "2228002",
      "postDate": "04/20/2023 07:36:30",
      "content": "<p>Congratulations on the solo gold! Thank you so much for providing an in-depth report and sharing the notebooks! There's a wealth of valuable information here.</p>",
      "rawMarkdown": "Congratulations on the solo gold! Thank you so much for providing an in-depth report and sharing the notebooks! There's a wealth of valuable information here.",
      "votes": null
    },
    {
      "id": "2228010",
      "postDate": "04/20/2023 07:41:40",
      "content": "<p>Truly it was quite unexpected, really want to if generally ensembling using decision trees is a thing.</p>",
      "rawMarkdown": "Truly it was quite unexpected, really want to if generally ensembling using decision trees is a thing.",
      "votes": null
    },
    {
      "id": "2228025",
      "postDate": "04/20/2023 07:49:44",
      "content": "<p>Congratulations and thanks for sharing!<br>\nI realized that scaling was the key to advance. That's why I honestly almost gave up the competition. I used Kaggle TPU with 330 GB RAM, but I found the training to be quite slow and couldn't complete in time (9 hours). But anyway it was fun.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!\nI realized that scaling was the key to advance. That's why I honestly almost gave up the competition. I used Kaggle TPU with 330 GB RAM, but I found the training to be quite slow and couldn't complete in time (9 hours). But anyway it was fun.",
      "votes": null
    },
    {
      "id": "2228041",
      "postDate": "04/20/2023 08:07:45",
      "content": "<p>Pays off to make a machine at home, I'm definitely buying another GPU before the next competition. 😅</p>",
      "rawMarkdown": "Pays off to make a machine at home, I'm definitely buying another GPU before the next competition. 😅",
      "votes": null
    },
    {
      "id": "2228045",
      "postDate": "04/20/2023 08:13:33",
      "content": "<p>Well, can't say no after getting such a prize.😂 You deserve it👏</p>",
      "rawMarkdown": "Well, can't say no after getting such a prize.😂 You deserve it👏",
      "votes": null
    },
    {
      "id": "2228098",
      "postDate": "04/20/2023 09:15:21",
      "content": "<p>Congratulations. From what I read, you trained the final model on 650 batches of data for 4 epochs, is that correct? How much time does it take on a 4080? What about the learning rate, did you use a constant learning rate or schedule it to decrease.  </p>",
      "rawMarkdown": "Congratulations. From what I read, you trained the final model on 650 batches of data for 4 epochs, is that correct? How much time does it take on a 4080? What about the learning rate, did you use a constant learning rate or schedule it to decrease.",
      "votes": null
    },
    {
      "id": "2228164",
      "postDate": "04/20/2023 10:17:36",
      "content": "<p>I use OneCycle learning rate scheduler, max rate 5e-5 for the largest model, final learning rate is 2e-6. The scheduler is certainly important.</p>\n<p>The final model I trained on a 4080 for 5 days, switching out from FP16 to FP32 after around ~1.3 epochs.</p>",
      "rawMarkdown": "I use OneCycle learning rate scheduler, max rate 5e-5 for the largest model, final learning rate is 2e-6. The scheduler is certainly important.\n\nThe final model I trained on a 4080 for 5 days, switching out from FP16 to FP32 after around ~1.3 epochs.",
      "votes": null
    },
    {
      "id": "2228168",
      "postDate": "04/20/2023 10:19:33",
      "content": "<p>Thanks, attention definitely, seemed like a more promising idea than using a graphnet, for the reason I described above. Making it work took some time. </p>",
      "rawMarkdown": "Thanks, attention definitely, seemed like a more promising idea than using a graphnet, for the reason I described above. Making it work took some time.",
      "votes": null
    },
    {
      "id": "2228202",
      "postDate": "04/20/2023 10:49:41",
      "content": "<p>Impressive, learned a lot from this competition. Interested to see the data pipeline. I got lazy and just stuck with 200 batches of data loaded into the memory. In the past, when training rnn based models, we usually have lr at 1e-3, which is still the default for pytorch. I tried the transformer-based model a bit with large lr and not enough layers, quickly abandoned that and back to rnn. Didn't realize you need that many parameters and a small learning rate to train. Thanks for sharing.</p>",
      "rawMarkdown": "Impressive, learned a lot from this competition. Interested to see the data pipeline. I got lazy and just stuck with 200 batches of data loaded into the memory. In the past, when training rnn based models, we usually have lr at 1e-3, which is still the default for pytorch. I tried the transformer-based model a bit with large lr and not enough layers, quickly abandoned that and back to rnn. Didn't realize you need that many parameters and a small learning rate to train. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2228285",
      "postDate": "04/20/2023 12:23:34",
      "content": "<p>Also, I just realized, congratulations on soon becoming a Competitions Grandmaster.</p>",
      "rawMarkdown": "Also, I just realized, congratulations on soon becoming a Competitions Grandmaster.",
      "votes": null
    },
    {
      "id": "2228574",
      "postDate": "04/20/2023 16:17:33",
      "content": "<p>Congrats for winning! I knew that GNN only approach is not all I need :)</p>",
      "rawMarkdown": "Congrats for winning! I knew that GNN only approach is not all I need :)",
      "votes": null
    },
    {
      "id": "2228808",
      "postDate": "04/20/2023 20:40:50",
      "content": "<p>First of all - congratulations! You made important step - solo gold. This is really amazing achievement. <br>\nYour solution is true magic. I saw your submission time during the last day of competition - 54 minutes 0.966, 105 minutes 0.964. Transformer + XGBoost -&gt; outstanding. This is really great example how to think out of the box - combine different ML techniques - deep learning (attention), boosting … respect 👍</p>",
      "rawMarkdown": "First of all - congratulations! You made important step - solo gold. This is really amazing achievement. \nYour solution is true magic. I saw your submission time during the last day of competition - 54 minutes 0.966, 105 minutes 0.964. Transformer + XGBoost -> outstanding. This is really great example how to think out of the box - combine different ML techniques - deep learning (attention), boosting … respect 👍",
      "votes": null
    },
    {
      "id": "2228823",
      "postDate": "04/20/2023 20:58:49",
      "content": "<p>If you look for Enseble books I recommend you books I have:</p>\n<ul>\n<li><a href=\"https://www.manning.com/books/ensemble-methods-for-machine-learning\" target=\"_blank\">https://www.manning.com/books/ensemble-methods-for-machine-learning</a></li>\n<li><a href=\"https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/\" target=\"_blank\">https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/</a></li>\n</ul>",
      "rawMarkdown": "If you look for Enseble books I recommend you books I have:\n- https://www.manning.com/books/ensemble-methods-for-machine-learning\n- https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/",
      "votes": null
    },
    {
      "id": "2228853",
      "postDate": "04/20/2023 21:47:06",
      "content": "<p>Thanks.</p>\n<p>I remember going like, wrote some for hacky loops for test \"if kappa &gt; k and zenith_prob &lt; p and zenith_base - zenith_stack &gt; d and z_first &gt; z\" etc, and score were improving. Then I realized, hey that's a cute little decision tree. </p>\n<p>Also, noob question, how to get these exact time values for submission runs?</p>",
      "rawMarkdown": "Thanks.\n\nI remember going like, wrote some for hacky loops for test \"if kappa > k and zenith_prob < p and zenith_base - zenith_stack > d and z_first > z\" etc, and score were improving. Then I realized, hey that's a cute little decision tree. \n\nAlso, noob question, how to get these exact time values for submission runs?",
      "votes": null
    },
    {
      "id": "2228854",
      "postDate": "04/20/2023 21:47:41",
      "content": "<p>Thank you so much for these.</p>",
      "rawMarkdown": "Thank you so much for these.",
      "votes": null
    },
    {
      "id": "2228935",
      "postDate": "04/21/2023 00:19:49",
      "content": "<p>Big congratulations on solo gold! It is quite interesting to read your writeup and follow the steps you made to get the improvement. Really great work!</p>",
      "rawMarkdown": "Big congratulations on solo gold! It is quite interesting to read your writeup and follow the steps you made to get the improvement. Really great work!",
      "votes": null
    },
    {
      "id": "2229285",
      "postDate": "04/21/2023 08:09:40",
      "content": "<p>Congrats, guys! Well deserved!</p>",
      "rawMarkdown": "Congrats, guys! Well deserved!",
      "votes": null
    },
    {
      "id": "2230400",
      "postDate": "04/22/2023 10:38:45",
      "content": "<p><a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> well done! Thanks for the detailed write-up and sharing your code on git.</p>",
      "rawMarkdown": "dipamc77 well done! Thanks for the detailed write-up and sharing your code on git.",
      "votes": null
    },
    {
      "id": "2231223",
      "postDate": "04/23/2023 06:58:35",
      "content": "<p>Do you recommend any material to know more about boosting model ensembler  🙌 </p>",
      "rawMarkdown": "Do you recommend any material to know more about boosting model ensembler  🙌",
      "votes": null
    },
    {
      "id": "2231886",
      "postDate": "04/23/2023 19:26:31",
      "content": "<p>Very good that you share this.</p>",
      "rawMarkdown": "Very good that you share this.",
      "votes": null
    },
    {
      "id": "2232005",
      "postDate": "04/23/2023 22:27:56",
      "content": "<p>Congratulations)</p>",
      "rawMarkdown": "Congratulations)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227950,
      "author_name": "horikitasaku",
      "author_url": "",
      "post_date": "04/20/2023 06:21:45",
      "content": "<p>Thanks for sharing! I heard the magic of Ensembling using XGBoost for the first time! This is very valuable to learn<br>\nLooking forward to your train code.</p>\n<p>(going to sleep before exploring the code</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228010,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 07:41:40",
          "content": "<p>Truly it was quite unexpected, really want to if generally ensembling using decision trees is a thing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2227963,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "04/20/2023 06:33:31",
      "content": "<p>This is awesome, self attention is one of the first things I went to try but could not get it to work so nice to see you solved it, thanks for the write up!</p>\n<p>Congratulations on solo gold and (soon) upgrading to Kaggle Master 😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228168,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 10:19:33",
          "content": "<p>Thanks, attention definitely, seemed like a more promising idea than using a graphnet, for the reason I described above. Making it work took some time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2227977,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "04/20/2023 07:02:22",
      "content": "<p>Congratulations on the solo gold! We were all amazed by your progress and hoping you would get the win!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2227989,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 07:26:02",
          "content": "<p>Thank you so much for the kind words.</p>\n<p>A week ago I thought the max score I can get is probably around 0.973, the boosting model ensembler really did the magic!   </p>",
          "votes": null,
          "replies": [
            {
              "id": 2231223,
              "author_name": "wbenitezdava",
              "author_url": "",
              "post_date": "04/23/2023 06:58:35",
              "content": "<p>Do you recommend any material to know more about boosting model ensembler  🙌 </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2228285,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 12:23:34",
          "content": "<p>Also, I just realized, congratulations on soon becoming a Competitions Grandmaster.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228002,
      "author_name": "inartimiryasov",
      "author_url": "",
      "post_date": "04/20/2023 07:36:30",
      "content": "<p>Congratulations on the solo gold! Thank you so much for providing an in-depth report and sharing the notebooks! There's a wealth of valuable information here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228025,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "04/20/2023 07:49:44",
      "content": "<p>Congratulations and thanks for sharing!<br>\nI realized that scaling was the key to advance. That's why I honestly almost gave up the competition. I used Kaggle TPU with 330 GB RAM, but I found the training to be quite slow and couldn't complete in time (9 hours). But anyway it was fun.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228041,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 08:07:45",
          "content": "<p>Pays off to make a machine at home, I'm definitely buying another GPU before the next competition. 😅</p>",
          "votes": null,
          "replies": [
            {
              "id": 2228045,
              "author_name": "mohammad2012191",
              "author_url": "",
              "post_date": "04/20/2023 08:13:33",
              "content": "<p>Well, can't say no after getting such a prize.😂 You deserve it👏</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2228098,
      "author_name": "wenrui29",
      "author_url": "",
      "post_date": "04/20/2023 09:15:21",
      "content": "<p>Congratulations. From what I read, you trained the final model on 650 batches of data for 4 epochs, is that correct? How much time does it take on a 4080? What about the learning rate, did you use a constant learning rate or schedule it to decrease.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2228164,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 10:17:36",
          "content": "<p>I use OneCycle learning rate scheduler, max rate 5e-5 for the largest model, final learning rate is 2e-6. The scheduler is certainly important.</p>\n<p>The final model I trained on a 4080 for 5 days, switching out from FP16 to FP32 after around ~1.3 epochs.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2228202,
              "author_name": "wenrui29",
              "author_url": "",
              "post_date": "04/20/2023 10:49:41",
              "content": "<p>Impressive, learned a lot from this competition. Interested to see the data pipeline. I got lazy and just stuck with 200 batches of data loaded into the memory. In the past, when training rnn based models, we usually have lr at 1e-3, which is still the default for pytorch. I tried the transformer-based model a bit with large lr and not enough layers, quickly abandoned that and back to rnn. Didn't realize you need that many parameters and a small learning rate to train. Thanks for sharing.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2228574,
      "author_name": "rusg77",
      "author_url": "",
      "post_date": "04/20/2023 16:17:33",
      "content": "<p>Congrats for winning! I knew that GNN only approach is not all I need :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228808,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/20/2023 20:40:50",
      "content": "<p>First of all - congratulations! You made important step - solo gold. This is really amazing achievement. <br>\nYour solution is true magic. I saw your submission time during the last day of competition - 54 minutes 0.966, 105 minutes 0.964. Transformer + XGBoost -&gt; outstanding. This is really great example how to think out of the box - combine different ML techniques - deep learning (attention), boosting … respect 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228823,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/20/2023 20:58:49",
          "content": "<p>If you look for Enseble books I recommend you books I have:</p>\n<ul>\n<li><a href=\"https://www.manning.com/books/ensemble-methods-for-machine-learning\" target=\"_blank\">https://www.manning.com/books/ensemble-methods-for-machine-learning</a></li>\n<li><a href=\"https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/\" target=\"_blank\">https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/</a></li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2228854,
              "author_name": "dipamc77",
              "author_url": "",
              "post_date": "04/20/2023 21:47:41",
              "content": "<p>Thank you so much for these.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2228853,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "04/20/2023 21:47:06",
          "content": "<p>Thanks.</p>\n<p>I remember going like, wrote some for hacky loops for test \"if kappa &gt; k and zenith_prob &lt; p and zenith_base - zenith_stack &gt; d and z_first &gt; z\" etc, and score were improving. Then I realized, hey that's a cute little decision tree. </p>\n<p>Also, noob question, how to get these exact time values for submission runs?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228935,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "04/21/2023 00:19:49",
      "content": "<p>Big congratulations on solo gold! It is quite interesting to read your writeup and follow the steps you made to get the improvement. Really great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2229285,
      "author_name": "wenlihong",
      "author_url": "",
      "post_date": "04/21/2023 08:09:40",
      "content": "<p>Congrats, guys! Well deserved!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2230400,
      "author_name": "crodoc",
      "author_url": "",
      "post_date": "04/22/2023 10:38:45",
      "content": "<p><a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> well done! Thanks for the detailed write-up and sharing your code on git.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2231886,
      "author_name": "jaholm",
      "author_url": "",
      "post_date": "04/23/2023 19:26:31",
      "content": "<p>Very good that you share this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232005,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "04/23/2023 22:27:56",
      "content": "<p>Congratulations)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2227888": "# 3rd place Solution writeup - Attention + XGBoost \n\nParticipated on Kaggle after 3 years, first gold for me, had a really exciting time. 😃\n\n**Check the TL;Dr at the end for a short summary** - The post is long with a mixture of model information and personal lessons.\n\n## The bitter lesson - Attention is all you need\n\nThis data is not a natural graph, I'm confused by all the graph models being trained on it. The graph edges are artificially selected based on criteria such as time or distance, which I didn't think was the right approach. Hence I spent my time on LSTM and later moved to Self-Attention.\n\n### Base Model\n\nThe base model is a simple **self-attention model**. Most of the model code is from Andrej Karpathy's NanoGPT https://github.com/karpathy/nanoGPT.git\n\nThe inputs are normalized pulse data with transparency information from Datasaurus' Dynedge baseline https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046 - Sequences are sampled upto a max length of 256, prioritizing sampling of the pulses where `auxiliary=False`\n\nThe outputs are **two classification heads** for each angle with 128 bins each. - I train it with a custom loss that locally smoothens the one hot target vectors, more on that below.\n\nA small version of the model can be trained on an RTX 3060 to give a score close to ~1.02 in 1 hour. (Embedding size 128 - 6 layers).\n\n### The bitter lesson - Scale beats everything\n\nThe competition was undoubtedly compute-heavy, but I only had access to my home PC with an RTX 3060. Naturally, I spent a lot of time hand-designing features, optimizing the code, and trying ideas that felt like they should help. Only to have to throw out most of the experiments which help a smaller model but get washed out when training a big model with more data.\n\nNotable ideas that helped but later got discarded:\n\n1. Simple augmentation like moving the points slightly, changing the charges etc.\n2. Supervised contrastive loss style training to predict the angle between two events.\n3. Mutliple pooling instead of just average pooling at the end. (in fact this caused overfitting)\n4. Angle classifier on sphere coorindates instead of two classifiers.\n\nAll of the above individually helped get better performance on a small model trained for 2-3 hours, but reach the same performance with bigger models. There is no overfitting in the models even with 20 batches of training data.\n\nThe real improvements in score came from scaling the model. Training on more data was crucial as well, the final models I use train on 650 batches of data. They finally start to overfit after 4 cycles through the entire data. \n\nAll the wasted days on experimentation reminds me of Richard Sutton's blog on \"The Bitter Lesson\". http://www.incompleteideas.net/IncIdeas/BitterLesson.html - Haha GPUs go bitterrrrrrr.\n\n### Scaling and engineering tricks\n\nAt the start of April, I decided to buy a new GPU just seeing how much compute is needed. Upgraded to RTX 4080, which is 3x faster. (No this is not the only scaling I did 😅, but more compute was crucial)\n\nFinal submitted model has 18 self attention layers, embedding size 512. This reaches a LB score of 0.982, and I'm sure training a bigger model will get even better. I trained the same model with 15 layers, but had worse performance due to mixed precision instability.\n\nEngineering tricks\n\n1. Mixed precision training is much faster, I train with **FP16 with FlashAttention**. However this started getting unstable after reaching a score of 0.990, so I switched to FP32 after that. Only the bigger models showed this instability and I didn't get to investigate how to continue on FP16, probably the scales of the inputs or the loss can be managed.\n2. The sequence lengths in the dataset has high variance. After sampling the data upto 256 pulses, the average length of each event is close of 100. This means a lot of attention processing (when not using FlashAttention) is wasted on padding. To speed up training I **group the batches by event length**. This gives a 1.5x speedup over FlashAttention and **3.5x speedup** over normal attention implementation. Packing also speeds up inference in the same ratios.\n3. Loading a portion of the training data for each \"epoch\". This is straightforward, switch out the bathces in the dataloader after each epoch, or do it as a background process for prefetch.\n\nScaling worked incredibly well:\n\n1. 128 Embedding - 9 Layers - 200 batches - Score ~1.001\n2. 256 Embedding - 12 Layers - 400 batches - Score ~0.990\n3. 512 Embedding - 18 Layers - 650 batches - Score ~0.982 (Some overfitting at the end)\n\nFurther, I fine tuned the model for sequence lengths upto 3072, which gives a slight imporvement of 0.002.\n\n### Smoothed cross entropy loss\n\nThe Von Mises-Fisher Loss proved unstable with Attention model, and I never got it to go beyond 1.04 when training end to end.\n\nFor the classifier I digitize the angles to 128 bins. Cross entropy loss works well, but I wanted to introduce inductive bias in the model that the clases are ordered, and not independent. To do so, I apply **1D convolution with a gaussian kernel** to the target one hot vectors. For azimuth this covolution has wraparound, for zenith I extend the bins based on the length of the kernel. (Notebook explaining the loss will be released later). This gives slightly more stable training, though with good hyperparam tuning even cross entropy might work just as well.\n\n## The sweet ending - Attention is not all you need\n\nObviously attention was not enough to win, 0.982 isn't even gold zone (cries in corner). However, the base model was cruicial for the next ideas to work.\n\n### Stack vMF model\n\nTraining an MLP with vMF loss on top of the encoder embeddings works. I dump 100 batches of average pooled encoder embeddings and train the stack model with it, achieves a score of ~0.978. Simple average ensemble of vMF and base model gives a score ~0.976.\n\n### The magic - Ensembling using XGBoost 🔥\n\nI know the standard ensembling tricks, but this is the first time I have ensembled using another model. If you are aware of this used in practice, **please let me know if there is literature on this**. I got the idea just 3 days before the end of the competition, barely got to test it properly, yet it worked like magic.\n\nMy hypothesis on why this worked - The classifier and vMF models are far from perfect scores, and have high disagreement. Here's a histogram of the zenith angles.\n\nBoth the model have its own biases, and its not obvious how to combine them.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3468635%2Fd4807872db2ad427f9484a1f6ac8ccbd%2Fzenith_histogram.png?generation=1681967255551488&alt=media)\n\nFirst I tried a simple voting rule - If they're too far then don't average:\nThis already bumped the score from 0.976 -> 0.973, already a top 5 solution.\n\n```python\navg_zenith = (zn_stack + zn_base) / 2\nvote_zenith = avg_zenith.clone()\nvoter = torch.abs(zn_stack - zn_base) > 0.4\nvote_zenith[voter] = zn_stack[voter].clone()\n```\n\nBoth the models have some way of indicating \"confidence\" in prediction. Probabilty of bin for classifier, and kappa for vMF. I played around with some thresholding for those, along with event lengths and z coordinates, and it bumped score to 0.9715, so I decided to try XGBoost. On the first try the score went straight to 0.969. 🚀\n\nFurther, I tried another boosting classifier to \"predict which predictions are bad\". And used it to flip the angles, this also gives a slight improvement in score.\n\nAdding some more simple features the final score for the 18 layer model was 0.965 (I made a submission with the 15 layer model which got 0.967).\n\nI used sklearn's HistGradientBoostingClassifier for the final setup, but XGBoost has similar results.\n\n### Inference timing stat\n\nThe inference time for the base model is around 4 minutes per batch, and another 4 minutes for the long sequence model, which only gives a small boost. So its possible to get a score of **0.967 within 25 minutes**.\n\n### Final submission\n\nThe final submission has 3 models. The 18 layer, 15 layer, and another checkpoint from same 18 layer run.\n\nAnother boosting classifier to merge 18 layer and 15 layer. And then simple voting to merge the 3.\n\nHere's the inference notebooks for final submission and single attention model.\n\n1. Final - https://www.kaggle.com/code/dipamc77/3rd-place-attention-xgboost-ensembler\n2. Attention only - https://www.kaggle.com/code/dipamc77/single-attention-model-inference\n\n## Summary of learnings\n\n1. Test the model for scaling, especially if its underfitting, before trying feature engineering/other ideas.\n2. Spend the initial engineering effort to optimize the pipeline, to test ideas fast.\n3. If models have high disagreement and not close to max score, try to train a classifier to ensemble to models.\n4. Even though the final improvement came from ensembling tricks, scaling was the most important. Bigger models will likely give even better results.\n5. Don't give up, I really threw everything in the kitchen sink to improve scores at the end!\n6. Buy more GPUs. 💸\n\nIn hindsight, I should have tried some of the graph models as they might have disagreed with the classifier, and ensembling them might give a good improvement.\n\n## Notable Failed ideas\n\nDidn't expect these to fail, and didn't abandon prematurely, but couldn't make them work.\n\n1. Finding noise in the dataset to ignore during training.\n2. vMF loss training with attention model.\n     - Edit - From reading the 2nd place solution, this might be because I didn't encode the inputs properly. Perhaps it was lack of understanding on my part. I thought encoding continuous variables with a normal FC layer is good enough. Clearly it did work with the classification loss though.\n\n3. Reducing bias in zenith predictions by class weightage and oversampling.\n\n## TL;Dr - Final Solution summary\n\nPulses as tokens, sequence length upto 256, priority sampling for \"real\" pulses.\n\n- (Self Attention (512 embedding, 18 layer, Avg pool, 2048 FF) + 2 Angle classifier heads) -> ~0.982\n- (Fine tuned attention model for sequence length upto 3072) -> 0.980\n- (MLP with vMF loss on attention model embeddings) -> ~0.976\n- (Combine classifier and vMF angles with boosting classifier) -> ~0.965 🔥\n- (15 layer + 18 layer + 1 checkpoint from same 18 layer run - combine with another boosting classifier) -> ~0.962\n\n-----------------\n\nP.S - Doesn't matter what your rank is, please share your learnings and also failed ideas. It helps you organize what you learnt and get better insights, I'll make sure to read all of them.\n\n**Training Code - Please share your feedback.**\n\nhttps://github.com/dipamc/kaggle-icecube-neutrinos",
    "2227950": "Thanks for sharing! I heard the magic of Ensembling using XGBoost for the first time! This is very valuable to learn\nLooking forward to your train code.\n\n(going to sleep before exploring the code",
    "2227963": "This is awesome, self attention is one of the first things I went to try but could not get it to work so nice to see you solved it, thanks for the write up!\n\nCongratulations on solo gold and (soon) upgrading to Kaggle Master 😄",
    "2227977": "Congratulations on the solo gold! We were all amazed by your progress and hoping you would get the win!",
    "2227989": "Thank you so much for the kind words.\n\nA week ago I thought the max score I can get is probably around 0.973, the boosting model ensembler really did the magic!",
    "2228002": "Congratulations on the solo gold! Thank you so much for providing an in-depth report and sharing the notebooks! There's a wealth of valuable information here.",
    "2228010": "Truly it was quite unexpected, really want to if generally ensembling using decision trees is a thing.",
    "2228025": "Congratulations and thanks for sharing!\nI realized that scaling was the key to advance. That's why I honestly almost gave up the competition. I used Kaggle TPU with 330 GB RAM, but I found the training to be quite slow and couldn't complete in time (9 hours). But anyway it was fun.",
    "2228041": "Pays off to make a machine at home, I'm definitely buying another GPU before the next competition. 😅",
    "2228045": "Well, can't say no after getting such a prize.😂 You deserve it👏",
    "2228098": "Congratulations. From what I read, you trained the final model on 650 batches of data for 4 epochs, is that correct? How much time does it take on a 4080? What about the learning rate, did you use a constant learning rate or schedule it to decrease.",
    "2228164": "I use OneCycle learning rate scheduler, max rate 5e-5 for the largest model, final learning rate is 2e-6. The scheduler is certainly important.\n\nThe final model I trained on a 4080 for 5 days, switching out from FP16 to FP32 after around ~1.3 epochs.",
    "2228168": "Thanks, attention definitely, seemed like a more promising idea than using a graphnet, for the reason I described above. Making it work took some time.",
    "2228202": "Impressive, learned a lot from this competition. Interested to see the data pipeline. I got lazy and just stuck with 200 batches of data loaded into the memory. In the past, when training rnn based models, we usually have lr at 1e-3, which is still the default for pytorch. I tried the transformer-based model a bit with large lr and not enough layers, quickly abandoned that and back to rnn. Didn't realize you need that many parameters and a small learning rate to train. Thanks for sharing.",
    "2228285": "Also, I just realized, congratulations on soon becoming a Competitions Grandmaster.",
    "2228574": "Congrats for winning! I knew that GNN only approach is not all I need :)",
    "2228808": "First of all - congratulations! You made important step - solo gold. This is really amazing achievement. \nYour solution is true magic. I saw your submission time during the last day of competition - 54 minutes 0.966, 105 minutes 0.964. Transformer + XGBoost -> outstanding. This is really great example how to think out of the box - combine different ML techniques - deep learning (attention), boosting … respect 👍",
    "2228823": "If you look for Enseble books I recommend you books I have:\n- https://www.manning.com/books/ensemble-methods-for-machine-learning\n- https://machinelearningmastery.com/ensemble-learning-algorithms-with-python/",
    "2228853": "Thanks.\n\nI remember going like, wrote some for hacky loops for test \"if kappa > k and zenith_prob < p and zenith_base - zenith_stack > d and z_first > z\" etc, and score were improving. Then I realized, hey that's a cute little decision tree. \n\nAlso, noob question, how to get these exact time values for submission runs?",
    "2228854": "Thank you so much for these.",
    "2228935": "Big congratulations on solo gold! It is quite interesting to read your writeup and follow the steps you made to get the improvement. Really great work!",
    "2229285": "Congrats, guys! Well deserved!",
    "2230400": "dipamc77 well done! Thanks for the detailed write-up and sharing your code on git.",
    "2231223": "Do you recommend any material to know more about boosting model ensembler  🙌",
    "2231886": "Very good that you share this.",
    "2232005": "Congratulations)"
  },
  "source": "meta"
}