{
  "id": 403713,
  "title": "8th Place Solution",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/403713",
  "author_name": "Isamu",
  "post_date": "2023-04-24T13:46:15.221000",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers and staff for hosting such a wonderful competition. And thank you to all my wonderful teammates! <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a>. </p>\n<p>Congratulations <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>!  your promotion to GM!</p>\n<p>Our code will be available <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\" target=\"_blank\">here</a>.</p>\n<h1>Datasaurus Part</h1>\n<h2>Preprocessing</h2>\n<h3>Raw data</h3>\n<p>This dataset was really big, and repeatedly doing pandas operations every epoch would have been a waste of CPU cycles (polars does not appear to work with PyTorch data loaders with many workers yet). To address this, I made PyTorch Geometric <code>Data</code> objects for each event and saved them as <code>.pt</code> files which could be loaded during training. This took about 8 hours to create using 32 threads, and required about 1TB of space. </p>\n<p>The issue with this was that that 1TB was across 130+ millon tiny files. A Linux partition has a finite number of “index nodes” or <code>inodes</code>, i.e. an index to a certain file. Since these files are so small I ran into my inode limit before I ran out of space on my 2TB drive. As a workaround, I had to spread some of these files across two drives, so heads up for anyone trying to reproduce this method or run my code.</p>\n<p>A more efficient way could be to store <code>.pt</code> files that have already been pre-batched which would require fewer files, but then you lose the ability to shuffle every epoch which may/may not make a difference with this much data. I didn’t try the sqlite method suggested by the GraphNet team.</p>\n<h2>Features</h2>\n<p>In the context of GNNs, each DOM is considered as a node. Each node was given the following 11 features:</p>\n<ul>\n<li>X: X location from sensor_geometry.csv / 500</li>\n<li>Y: Y location from sensor_geometry.csv / 500</li>\n<li>Z: Z location from sensor_geometry.csv / 500</li>\n<li>T: (Time from batch_[n].parquet - 1e4) / 3e3</li>\n<li>Charge: log10(Charge from batch_[n].parquet) / 3.0</li>\n<li>QE: See below</li>\n<li>Aux: (False = -0.5, True = 0.5)</li>\n<li>Scattering Length: See below</li>\n<li>Distance to previous hit: See below</li>\n<li>Time delta since previous hit: See below</li>\n<li>Scattering flag (False = -0.5, True = 0.5). See below</li>\n</ul>\n<p>Many of the normalisation methods were taken from GraphNet as a <a href=\"https://github.com/graphnet-team/graphnet/blob/4df8f396400da3cfca4ff1e0593a0c7d1b5b5195/src/graphnet/models/detector/icecube.py#L64-L69\" target=\"_blank\">starting point</a>, but I altered the scale for time, since time is a very important feature here.</p>\n<p>For events with large numbers of hits, to prevent OOM errors, I sampled 256 hits. This can make the process slightly non-deterministic.</p>\n<h2>Quantum efficiency</h2>\n<p>QE is the quantum efficiency of the photomutipliers in the DOMs. The DeepCore DOMs are quoted to have 35% higher QE than the regular DOMs (Figure 1 of this <a href=\"https://arxiv.org/pdf/2209.03042.pdf\" target=\"_blank\">paper</a>), so QE was set to 1 everywhere, and 1.35 for the lower 50 DOMs in DeepCore. The final QE feature was scaled using (QE - 1.25) / 0.25.</p>\n<h2>Scattering length</h2>\n<p>Scattering and absorption lengths are important to characterise differences in the clarity of the ice. This data is published on page 31 of this <a href=\"https://arxiv.org/abs/1301.5361\" target=\"_blank\">paper</a>. A datum depth of 1920 metres was used so that z = (depth - 1920) / 500. The data was resampled to the z values using <code>scipy.interpolate.interp1d</code>. I found that after passing the data though <code>RobustScaler</code>, the scattering and absorption data was near identical, so I only used scattering length.</p>\n<h2>Previous hit features</h2>\n<p>The two main types of events are track and cascade events. Looking at some of the amazing <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388858\" target=\"_blank\">visualisation tools</a> for example from edguy99, I got the idea that if a node had some understanding where and when the nearest previous hit was, it might help the model differentiate between these two groups. To calculate this for each event, sorted the hits by time, calculated the pairwise distances of all hits, masked any hits from the future and calculated the distance, d, to the nearest previous hit. This was scaled using (d - 0.5) / 0.5. The time delta from the previous hit was also calculated using the same method and scaled using (t - 0.1) / 0.1.</p>\n<h2>Scattering flag</h2>\n<p>I tried to create a flag that could discern whether a hit was caused directly from a track, or some secondary scattering, inspired by section 2.1 of this <a href=\"https://arxiv.org/pdf/2203.02303.pdf\" target=\"_blank\">paper</a>. A side effect of adding this flag was that training was much more stable. The flag is generated as follows:</p>\n<ol>\n<li>Identify the hit with the largest charge</li>\n<li>From this DOM location, calculate the distances &amp; time delta to every other hit</li>\n<li>If the time taken to travel that distance is &gt; speed of light in ice, assume that the photon is a result of scattering</li>\n</ol>\n<h2>Validation</h2>\n<p>I used 90% - 10% train-validation split, and no cross validation due to the size of the dataset. The split was done by creating 10 bins of log10(n_hits), and then using <code>StratifiedKFold</code> using 10 splits.</p>\n<h2>Models</h2>\n<p>I used the <code>DirectionReconstructionWithKappa</code> task directly from GraphNet, meaning that an embedding (e.g. shape of 128) will be projected to a shape of 4 (x, y, z, kappa)</p>\n<h2>Architectures</h2>\n<p>I used the following 3 architectures. All validation scores are with 6x TTA applied</p>\n<p><a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">GraphNet/DynEdge</a> -  Val = 0.98501<br>\n<a href=\"https://arxiv.org/abs/2205.12454\" target=\"_blank\">GPS</a> - Val = 0.98945*<br>\n<a href=\"https://arxiv.org/abs/1902.07987\" target=\"_blank\">GravNet</a> - Val = 0.98519</p>\n<p>The average of these 3 models gave a 0.982 LB score.</p>\n<p>The GraphNet/DynEdge model had very little modfification, other than changing to GELU activations.</p>\n<p>GPS &amp; GravNet used 8 blocks and you can find the exact architectures for both in the code <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice/blob/main/src/modules.py\" target=\"_blank\">here</a>.</p>\n<p>*GPS was the most powerful model, but also slowest to train being a transformer type model (roughly 11 hours/epoch on my machine). I managed to train a model which achieved a validation score of 0.98XX but was too late to include in our final submission.</p>\n<h2>Loss</h2>\n<p>I used VonMisesFisher3DLoss + (1 - CosineSimilarity) as the final loss function, since cosine similarity is a nice proxy for mean angular error. For CosineSimilarity I transformed the target azimuth &amp; zenith values to cartesian coordinates.</p>\n<p>This performed much better than separate losses for azimuth (VMF2D) &amp; zenith (MSE) which I the route I <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383546\" target=\"_blank\">initially went down</a>.</p>\n<h2>Augmentation</h2>\n<p>I centered the data on string 35 (the DeepCore string) and rotated about the z-axis in 60 degree steps. This didn’t actually improve validation performance but did have the benefit of making the models rotationally invariant so that I could take advantage of the detector symmetry and apply a 6x test time augmentation (TTA). This often improved scores by 0.002-0.003. </p>\n<h2>Training parameters</h2>\n<ul>\n<li>AdamW optimiser</li>\n<li>Epochs = 6</li>\n<li>Cosine schedule (no warmup)</li>\n<li>Learning rate = 0.0002</li>\n<li>Batch size = 1024</li>\n<li>Weight decay = 0.001 - 0.1 depending on model</li>\n<li>FP16 training</li>\n<li>Hardware: 2x RTX 3090, 128 GB RAM</li>\n</ul>\n<h2>Final submissions</h2>\n<p>Circular mean</p>\n<h2>Robustness to perturbation</h2>\n<p>TBC</p>\n<h2>Lessons learned/stuff that didn’t work</h2>\n<ul>\n<li>The GraphNet DynEdge baseline is extremely strong and tough to improve on - kudos to the team! It is also the fastest/efficient model, and what I used for the majority of experimentation</li>\n<li>More data = more better. The issue with this though is that I found that some conclusions drawn from experiments on 1% or 5% of the data were no longer applicable on the full dataset. This made experimentation slow and expensive</li>\n<li>Batch normalisation made things unstable and didn’t show improvements</li>\n<li>Lion optimiser didn’t generalise as well as AdamW</li>\n<li>Weight decay was important for some models. As a result I assumed changing the epsilon value in Adam would have an effect, but I didn’t see anything significant</li>\n<li>In my experiments, GNNs seem to benefit from leaky activations, e.g. GELU</li>\n<li>For MPNN aggregation, it seems that [min, max, mean, sum] is sufficient. Adding more didn’t appear to make significant gains</li>\n<li>Realigning all of the times to the time of the first hit of each event deteriorates performance, possibly due to noise in the data/false triggers etc.</li>\n<li>Radial nearest neighbours didn’t work any better than KNN when defining graph edges</li>\n<li>Only using 1 - CosineSimilarity as a loss function wasn’t very stable. Adding VMF3D helped a lot</li>\n</ul>\n<h2>Code</h2>\n<p>All my code will be available here soon: <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\" target=\"_blank\">https://github.com/Anjum48/icecube-neutrinos-in-deep-ice</a></p>\n<h1>Isamu Part</h1>\n<p>Before the team merge, I created LSTM and GraphNet models. After the team merge, I focused on GraphNet because Remek's LSTM model was superior to mine. I used graphnet (<a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">https://github.com/graphnet-team/graphnet</a>) as a baseline and made several changes to improve its accuracy. I list below some of the experiments I performed that worked(There are tons of things that didn't work)</p>\n<ul>\n<li>random sampling (random sampling from DB if the specified data length exceeds 800)</li>\n<li>Increasing nearest neighbors of KNN layer(8-&gt;16)</li>\n<li>Addition of features<ul>\n<li>x, y, z</li>\n<li>time</li>\n<li>charge</li>\n<li>auxiliary</li>\n<li>ice_transparency feature</li></ul></li>\n<li>2-stage model with kappa(sigma) of vonMisesFisher distribution <ul>\n<li>Train 1st stage model to predict x, y, z, kappa<ul>\n<li>restart from the weights of GraphNet from public baseline notebook</li>\n<li>About 1-250 batches were used</li>\n<li>batch size 512</li>\n<li>epoch 20</li>\n<li>DirectionReconstructionWithKappa</li></ul></li>\n<li>Split data into easy and hard parts according to 1st stage kappa value<ul>\n<li>Inference was performed using the 1st model and classified into two sets of data(easy part and hard part) according to their predicted kappa value</li></ul></li>\n<li>Train expert models for easy and hard parts and combine their predictions<ul>\n<li>About 250-350 batches were used</li>\n<li>batch size 512</li>\n<li>epoch 20</li>\n<li>DirectionReconstructionWithKappa</li></ul></li></ul></li>\n<li>TTA <ul>\n<li>rotation 180-degree TTA about the z-axis</li></ul></li>\n<li>Loss<ul>\n<li>DirectionReconstructionWithKappa</li></ul></li>\n<li>Hardware: RTX 3090, 64 GB RAM, 8TB HDD, 2TB SSD, Google Colab Pro</li>\n</ul>\n<p>The ensemble of models(1st and 2nd) created above gave public LB 0.995669, private LB 0.996550 </p>\n<h1>Remek Part - LSTM</h1>\n<p>For LSTM training we used the attitude proposed by Robin Smits (@rsmits) with some improvements.</p>\n<ul>\n<li><p>We added more LSTM/GRU layers (4 GRU/LSTM) - more and less than 4 layers was not better in our experiments.</p></li>\n<li><p>We added ice transparency as additional features (we used both features – transparency and absorption).</p></li>\n<li><p>We redesigned the training loop to train models on all batches – we can train models on different parts of DS (from one batch to all DS in one epoch).</p></li>\n<li><p>We checked many hypotheses (for part of them weight and biases Swipe tool was used):</p>\n<ul>\n<li>Different size of LSTM units – finally 196 was the best in our model.</li>\n<li>Different bin size – 24 was our final choice.</li>\n<li>Different amount of features and impulse selection - finally we use strategy - first select non_aux events then add random aux events (it gave us score boost as well - this was kind of an augmentation technique).</li>\n<li>Different model architectures – finally for our score blend we used pure LSTM/GRU setup. We tested transformer architecture but as it appeared we gave up too early (first scores were way worse then pure LSTM).</li>\n<li>Different optimizers (AdamW, NAdam) – final choice was Adam.</li>\n<li>Different schedulers – CosineDecay, OneCycle but then we use step LR scheduling described below.</li></ul></li>\n</ul>\n<p>Training was divided  into three parts scheduling LR:</p>\n<ul>\n<li>Step 1 - Train baseline models (two models) – batches 4-330 (m1) and 331-660 (m2) using LR = 0.005, for 6-8 epochs, sparse categorical crossentropy loss. Both models were validated on batches 1-3.</li>\n<li>Step 2 - Fine tuning models m1 and m2 using LR/10 = 0.0005 for 3 epochs, sparse categorical crossentropy loss.</li>\n<li>Step 3 – Fine tuning models m1 and m2 using LR = 0.00025 for 2 epochs with different loss function – categorical crossentropy with label smoothing 0.05</li>\n<li>Then we took 4 best models (according to MAE metrics) for both model m1 and m2 (8 models in total) and produced one model using SWA (Stochastic Weight Averaging). We simply averaged model weights. As it appeared the model had the same performance compared for 8 ensembled models but we significantly decreased LSTM inference time. Single model score (public LB) (no TTA): 1.0024.</li>\n</ul>\n<p>Things did not improve our score:</p>\n<ul>\n<li>Dropout in LSTM/GRU layers and linear head.</li>\n<li>GaussianNoise Layer after Masking layer or after LSTM/GRU layer.</li>\n<li>Adam + Lookahead optimizer.</li>\n<li>Gradient Accumulation to simulate TPU big batch size – I (Remek) had problem to implement it properly in TF/Keras (it is easy in Pytorch, as it appeared not exactly easy to implement in TF/Keras – my daily choice is Pytorch)</li>\n</ul>\n<p>Additional tools:</p>\n<ul>\n<li>Weights and biases for two tasks – logging and monitoring training process and swipe for hyperparameter tuning (number of LSTM units, LR scheduler, Optimizer and LR).</li>\n</ul>\n<p>For my (Remek) experiment I use (in first part of competition) ZbyHP Z4 with 2xA5000 but then HP sent me  ZbyHP Z8 workstation with 2x Intel Xeon CPU and Nvidia A6000 GPU. Personally I can say that this help me to establish fast experimentation pipeline. I was able to process dataset files very fast. Having A6000 gave me possibility to train models on bigger batch size. Thank HP for supporting my work.</p>\n<p>My final words - we set up a great team in my opinion - each of us was responsible for part of the solution, we discussed a lot but there was no “my is better”. Although we worked together for the first time, I had the impression that we had known each other forever. Great team, great result! Thank you guys for having the opportunity to learn from great AI guys.</p>\n<h1>Alvor part - blending</h1>\n<p>My main contribution to this competition was the development of methods for blending my teammates’ solutions. Analysis of the different types of solutions (GraphNet, LSTM etc.) showed that they have different efficiency at different predicted zeniths and azimuths values (zenith mainly). So I changed my initial \"constant weight\" approach to a \"bins\" approach.<br>\nThe method consists in splitting the predicted zenith values into 10 bins of equal width. Thus, when combining two solutions, we get 100 bins in total. The value of 10 is configurable, but experiments showed it to be close to optimal.</p>\n<p>After that, the blending weight for the zenith was found in each bin and the blending weight for the azimuth was found. The blending of predicted zenith values was done by a simple linear combination. Blending the azimuth values was a little more complicated, due to the possible transition through 2*pi. Therefore, to begin with, the difference in the predicted azimuths in the two solutions was calculated and the direction of the second value relative to the first value.<br>\nWeights fitting was carried out on a sample of training data close in size to the test data (5 batches with 1M events).</p>\n<p>Several approaches to improve this blending method were also tested. In particular, the use of GBDT.</p>\n<p>One of these approaches was an attempt to classify events into “simple” and “complex” ones (inspired by the kernel <a href=\"https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93)\" target=\"_blank\">https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93)</a>.</p>\n<p>Experiments have shown that some types of neural networks are better at handling simple events, while other types are better at handling complex events. If we could accurately determine the type of each event, this would greatly improve the competition metric. A model was built, the efficiency of which in events classifying was at the level of the public kernel (AUС 0.93), but this was not enough to get a noticeable improvement in the blend quality.</p>\n<p>Several other models of classification (binary variable - which of the two neural networks better predicts a given event) and regression (target variable - the difference in competition metrics from two neural networks for a given event) were also trained. Unfortunately, all these approaches showed only a minor improvement in the result (fourth decimal place), but they were very overfitting-sensitive.</p>\n<p>Later, my method was significantly improved by Wojtek Rosa. Therefore, it was his approach that was used in the final submissions, which he describes in more detail in his part. In particular, I would like to note his brilliant idea of using a Decision Tree model to build blending bins.</p>\n<h1>Wojtek part - more blending</h1>\n<p>Thank you competition host for exciting competition IceCube!<br>\nAlso thank you my Teammates - once again congrats for their prizes, GM titles and great solutions.<br>\nMy part was about blending.<br>\nI used batch_ids 1-5 for evaluation and adjusting parameters, later I used batch_ids 655+ only for evaluation purposes.<br>\nWhen I joined Remek&amp;Alvor team, we have great LSTM solution, public Graphnet, and amazing blend technique:</p>\n<pre><code> ():\n   s[] = np.(s[] - s[])\n   s[] = np.where(\n   s[] &lt; np.pi,\n   np.sign(s[] - s[]),\n   -np.sign(s[] - s[])\n   )\n   N = \n   s[] = (N*np.floor(N*s[]/np.pi) + np.floor(N*s[]/np.pi)).astype()\n</code></pre>\n<p>and minimize score for each bin, finding best qu, alpha such as:</p>\n<pre><code>s0[] = s0[] + alpha * s0[] * s0[]\ns0[] = (-qu) * s0[] + qu * s0[]\n</code></pre>\n<p>I realize, that compared to simple/public vector weight ensembling this method is much better.<br>\nI tried to improve score with greater N values but this leads me to overfit.<br>\nAfter this, I managed to improve the score, by creating bins using Regression Trees with target = score_1 - score_2 and features:</p>\n<pre><code>cls = [,,,,,, ,,,,,,\n      ,,,]\n</code></pre>\n<p>where direction_% are from Isamu submission, zenith_3 is from crazy quick and clean 1.183 Robert linear solution. I tried many different features directly from event, but<br>\nThe simplicity of Decision Tree Regressor gave us (almost) full control of N bins for adjustment allvors parameters qu and alpha:</p>\n<pre><code>regr_1 = DecisionTreeRegressor(max_depth=, min_samples_leaf=, min_samples_split=)\ns[] = regr_1.predict(s[cls])\ns[] = (s[],).astype()\n</code></pre>\n<p>Using DecissionTrees gave us 0.0027 boost both with local score and LB. I tried other methods of blending such as MLP or LGBM without success.<br>\nLater on, I also slightly modified <code>adjusting</code> loop, for finding full linear combination of zenith_1 and zenith_2:</p>\n<pre><code>s0[] = s0[] + alpha * s0[] * s0[]\ns0[] = ru * s0[] + qu * s0[]\n</code></pre>\n<p>It takes much more computation to find ru and qu, but that gives us additional 0.001 improvement.<br>\nOur final blend consist:<br>\nIsamu solution .995 -&gt;  Remek LSTM 1.002 -&gt;  datasaurus Graphnet .982 -&gt; Robert 1.183<br>\nlocal score 0.9774692   Public LB 0.975545, Private LB 0.97658</p>\n<h2>Findings:</h2>\n<ul>\n<li>we found interesting type of error, ‘model give up’ and predicts zenith close to 0.<br>\nOur final solution zenith prediction vs score (x-axis zenith predicted, y-axis zenith ground truth):</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F4a8b117472efd756290083dcf0b787c1%2Fimage1.png?generation=1682343676129581&amp;alt=media\" alt=\"\"></p>\n<p>Ground truth zenith (x-axis) vs score (y-axis) - ‘error triangle’ visible:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F2ee6d7032191782027e45522daff9081%2Fimage2.png?generation=1682343631946316&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>we found that predictions from LSTM classification goes incredibly good at the center of bins, chart: zenith_pred_lstm*1000 (x-axis), mean score (y-axis), zenith edges (red lines)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F7fde090be1100c83a09f045fbf57f5fd%2Fimage3.png?generation=1682343415933098&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2232604,
      "postDate": "2023-04-24T13:46:15.223Z",
      "content": "<p>First of all, I would like to thank the organizers and staff for hosting such a wonderful competition. And thank you to all my wonderful teammates! <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>, <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a>. </p>\n<p>Congratulations <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>!  your promotion to GM!</p>\n<p>Our code will be available <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\" target=\"_blank\">here</a>.</p>\n<h1>Datasaurus Part</h1>\n<h2>Preprocessing</h2>\n<h3>Raw data</h3>\n<p>This dataset was really big, and repeatedly doing pandas operations every epoch would have been a waste of CPU cycles (polars does not appear to work with PyTorch data loaders with many workers yet). To address this, I made PyTorch Geometric <code>Data</code> objects for each event and saved them as <code>.pt</code> files which could be loaded during training. This took about 8 hours to create using 32 threads, and required about 1TB of space. </p>\n<p>The issue with this was that that 1TB was across 130+ millon tiny files. A Linux partition has a finite number of “index nodes” or <code>inodes</code>, i.e. an index to a certain file. Since these files are so small I ran into my inode limit before I ran out of space on my 2TB drive. As a workaround, I had to spread some of these files across two drives, so heads up for anyone trying to reproduce this method or run my code.</p>\n<p>A more efficient way could be to store <code>.pt</code> files that have already been pre-batched which would require fewer files, but then you lose the ability to shuffle every epoch which may/may not make a difference with this much data. I didn’t try the sqlite method suggested by the GraphNet team.</p>\n<h2>Features</h2>\n<p>In the context of GNNs, each DOM is considered as a node. Each node was given the following 11 features:</p>\n<ul>\n<li>X: X location from sensor_geometry.csv / 500</li>\n<li>Y: Y location from sensor_geometry.csv / 500</li>\n<li>Z: Z location from sensor_geometry.csv / 500</li>\n<li>T: (Time from batch_[n].parquet - 1e4) / 3e3</li>\n<li>Charge: log10(Charge from batch_[n].parquet) / 3.0</li>\n<li>QE: See below</li>\n<li>Aux: (False = -0.5, True = 0.5)</li>\n<li>Scattering Length: See below</li>\n<li>Distance to previous hit: See below</li>\n<li>Time delta since previous hit: See below</li>\n<li>Scattering flag (False = -0.5, True = 0.5). See below</li>\n</ul>\n<p>Many of the normalisation methods were taken from GraphNet as a <a href=\"https://github.com/graphnet-team/graphnet/blob/4df8f396400da3cfca4ff1e0593a0c7d1b5b5195/src/graphnet/models/detector/icecube.py#L64-L69\" target=\"_blank\">starting point</a>, but I altered the scale for time, since time is a very important feature here.</p>\n<p>For events with large numbers of hits, to prevent OOM errors, I sampled 256 hits. This can make the process slightly non-deterministic.</p>\n<h2>Quantum efficiency</h2>\n<p>QE is the quantum efficiency of the photomutipliers in the DOMs. The DeepCore DOMs are quoted to have 35% higher QE than the regular DOMs (Figure 1 of this <a href=\"https://arxiv.org/pdf/2209.03042.pdf\" target=\"_blank\">paper</a>), so QE was set to 1 everywhere, and 1.35 for the lower 50 DOMs in DeepCore. The final QE feature was scaled using (QE - 1.25) / 0.25.</p>\n<h2>Scattering length</h2>\n<p>Scattering and absorption lengths are important to characterise differences in the clarity of the ice. This data is published on page 31 of this <a href=\"https://arxiv.org/abs/1301.5361\" target=\"_blank\">paper</a>. A datum depth of 1920 metres was used so that z = (depth - 1920) / 500. The data was resampled to the z values using <code>scipy.interpolate.interp1d</code>. I found that after passing the data though <code>RobustScaler</code>, the scattering and absorption data was near identical, so I only used scattering length.</p>\n<h2>Previous hit features</h2>\n<p>The two main types of events are track and cascade events. Looking at some of the amazing <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388858\" target=\"_blank\">visualisation tools</a> for example from edguy99, I got the idea that if a node had some understanding where and when the nearest previous hit was, it might help the model differentiate between these two groups. To calculate this for each event, sorted the hits by time, calculated the pairwise distances of all hits, masked any hits from the future and calculated the distance, d, to the nearest previous hit. This was scaled using (d - 0.5) / 0.5. The time delta from the previous hit was also calculated using the same method and scaled using (t - 0.1) / 0.1.</p>\n<h2>Scattering flag</h2>\n<p>I tried to create a flag that could discern whether a hit was caused directly from a track, or some secondary scattering, inspired by section 2.1 of this <a href=\"https://arxiv.org/pdf/2203.02303.pdf\" target=\"_blank\">paper</a>. A side effect of adding this flag was that training was much more stable. The flag is generated as follows:</p>\n<ol>\n<li>Identify the hit with the largest charge</li>\n<li>From this DOM location, calculate the distances &amp; time delta to every other hit</li>\n<li>If the time taken to travel that distance is &gt; speed of light in ice, assume that the photon is a result of scattering</li>\n</ol>\n<h2>Validation</h2>\n<p>I used 90% - 10% train-validation split, and no cross validation due to the size of the dataset. The split was done by creating 10 bins of log10(n_hits), and then using <code>StratifiedKFold</code> using 10 splits.</p>\n<h2>Models</h2>\n<p>I used the <code>DirectionReconstructionWithKappa</code> task directly from GraphNet, meaning that an embedding (e.g. shape of 128) will be projected to a shape of 4 (x, y, z, kappa)</p>\n<h2>Architectures</h2>\n<p>I used the following 3 architectures. All validation scores are with 6x TTA applied</p>\n<p><a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">GraphNet/DynEdge</a> -  Val = 0.98501<br>\n<a href=\"https://arxiv.org/abs/2205.12454\" target=\"_blank\">GPS</a> - Val = 0.98945*<br>\n<a href=\"https://arxiv.org/abs/1902.07987\" target=\"_blank\">GravNet</a> - Val = 0.98519</p>\n<p>The average of these 3 models gave a 0.982 LB score.</p>\n<p>The GraphNet/DynEdge model had very little modfification, other than changing to GELU activations.</p>\n<p>GPS &amp; GravNet used 8 blocks and you can find the exact architectures for both in the code <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice/blob/main/src/modules.py\" target=\"_blank\">here</a>.</p>\n<p>*GPS was the most powerful model, but also slowest to train being a transformer type model (roughly 11 hours/epoch on my machine). I managed to train a model which achieved a validation score of 0.98XX but was too late to include in our final submission.</p>\n<h2>Loss</h2>\n<p>I used VonMisesFisher3DLoss + (1 - CosineSimilarity) as the final loss function, since cosine similarity is a nice proxy for mean angular error. For CosineSimilarity I transformed the target azimuth &amp; zenith values to cartesian coordinates.</p>\n<p>This performed much better than separate losses for azimuth (VMF2D) &amp; zenith (MSE) which I the route I <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383546\" target=\"_blank\">initially went down</a>.</p>\n<h2>Augmentation</h2>\n<p>I centered the data on string 35 (the DeepCore string) and rotated about the z-axis in 60 degree steps. This didn’t actually improve validation performance but did have the benefit of making the models rotationally invariant so that I could take advantage of the detector symmetry and apply a 6x test time augmentation (TTA). This often improved scores by 0.002-0.003. </p>\n<h2>Training parameters</h2>\n<ul>\n<li>AdamW optimiser</li>\n<li>Epochs = 6</li>\n<li>Cosine schedule (no warmup)</li>\n<li>Learning rate = 0.0002</li>\n<li>Batch size = 1024</li>\n<li>Weight decay = 0.001 - 0.1 depending on model</li>\n<li>FP16 training</li>\n<li>Hardware: 2x RTX 3090, 128 GB RAM</li>\n</ul>\n<h2>Final submissions</h2>\n<p>Circular mean</p>\n<h2>Robustness to perturbation</h2>\n<p>TBC</p>\n<h2>Lessons learned/stuff that didn’t work</h2>\n<ul>\n<li>The GraphNet DynEdge baseline is extremely strong and tough to improve on - kudos to the team! It is also the fastest/efficient model, and what I used for the majority of experimentation</li>\n<li>More data = more better. The issue with this though is that I found that some conclusions drawn from experiments on 1% or 5% of the data were no longer applicable on the full dataset. This made experimentation slow and expensive</li>\n<li>Batch normalisation made things unstable and didn’t show improvements</li>\n<li>Lion optimiser didn’t generalise as well as AdamW</li>\n<li>Weight decay was important for some models. As a result I assumed changing the epsilon value in Adam would have an effect, but I didn’t see anything significant</li>\n<li>In my experiments, GNNs seem to benefit from leaky activations, e.g. GELU</li>\n<li>For MPNN aggregation, it seems that [min, max, mean, sum] is sufficient. Adding more didn’t appear to make significant gains</li>\n<li>Realigning all of the times to the time of the first hit of each event deteriorates performance, possibly due to noise in the data/false triggers etc.</li>\n<li>Radial nearest neighbours didn’t work any better than KNN when defining graph edges</li>\n<li>Only using 1 - CosineSimilarity as a loss function wasn’t very stable. Adding VMF3D helped a lot</li>\n</ul>\n<h2>Code</h2>\n<p>All my code will be available here soon: <a href=\"https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\" target=\"_blank\">https://github.com/Anjum48/icecube-neutrinos-in-deep-ice</a></p>\n<h1>Isamu Part</h1>\n<p>Before the team merge, I created LSTM and GraphNet models. After the team merge, I focused on GraphNet because Remek's LSTM model was superior to mine. I used graphnet (<a href=\"https://github.com/graphnet-team/graphnet\" target=\"_blank\">https://github.com/graphnet-team/graphnet</a>) as a baseline and made several changes to improve its accuracy. I list below some of the experiments I performed that worked(There are tons of things that didn't work)</p>\n<ul>\n<li>random sampling (random sampling from DB if the specified data length exceeds 800)</li>\n<li>Increasing nearest neighbors of KNN layer(8-&gt;16)</li>\n<li>Addition of features<ul>\n<li>x, y, z</li>\n<li>time</li>\n<li>charge</li>\n<li>auxiliary</li>\n<li>ice_transparency feature</li></ul></li>\n<li>2-stage model with kappa(sigma) of vonMisesFisher distribution <ul>\n<li>Train 1st stage model to predict x, y, z, kappa<ul>\n<li>restart from the weights of GraphNet from public baseline notebook</li>\n<li>About 1-250 batches were used</li>\n<li>batch size 512</li>\n<li>epoch 20</li>\n<li>DirectionReconstructionWithKappa</li></ul></li>\n<li>Split data into easy and hard parts according to 1st stage kappa value<ul>\n<li>Inference was performed using the 1st model and classified into two sets of data(easy part and hard part) according to their predicted kappa value</li></ul></li>\n<li>Train expert models for easy and hard parts and combine their predictions<ul>\n<li>About 250-350 batches were used</li>\n<li>batch size 512</li>\n<li>epoch 20</li>\n<li>DirectionReconstructionWithKappa</li></ul></li></ul></li>\n<li>TTA <ul>\n<li>rotation 180-degree TTA about the z-axis</li></ul></li>\n<li>Loss<ul>\n<li>DirectionReconstructionWithKappa</li></ul></li>\n<li>Hardware: RTX 3090, 64 GB RAM, 8TB HDD, 2TB SSD, Google Colab Pro</li>\n</ul>\n<p>The ensemble of models(1st and 2nd) created above gave public LB 0.995669, private LB 0.996550 </p>\n<h1>Remek Part - LSTM</h1>\n<p>For LSTM training we used the attitude proposed by Robin Smits (@rsmits) with some improvements.</p>\n<ul>\n<li><p>We added more LSTM/GRU layers (4 GRU/LSTM) - more and less than 4 layers was not better in our experiments.</p></li>\n<li><p>We added ice transparency as additional features (we used both features – transparency and absorption).</p></li>\n<li><p>We redesigned the training loop to train models on all batches – we can train models on different parts of DS (from one batch to all DS in one epoch).</p></li>\n<li><p>We checked many hypotheses (for part of them weight and biases Swipe tool was used):</p>\n<ul>\n<li>Different size of LSTM units – finally 196 was the best in our model.</li>\n<li>Different bin size – 24 was our final choice.</li>\n<li>Different amount of features and impulse selection - finally we use strategy - first select non_aux events then add random aux events (it gave us score boost as well - this was kind of an augmentation technique).</li>\n<li>Different model architectures – finally for our score blend we used pure LSTM/GRU setup. We tested transformer architecture but as it appeared we gave up too early (first scores were way worse then pure LSTM).</li>\n<li>Different optimizers (AdamW, NAdam) – final choice was Adam.</li>\n<li>Different schedulers – CosineDecay, OneCycle but then we use step LR scheduling described below.</li></ul></li>\n</ul>\n<p>Training was divided  into three parts scheduling LR:</p>\n<ul>\n<li>Step 1 - Train baseline models (two models) – batches 4-330 (m1) and 331-660 (m2) using LR = 0.005, for 6-8 epochs, sparse categorical crossentropy loss. Both models were validated on batches 1-3.</li>\n<li>Step 2 - Fine tuning models m1 and m2 using LR/10 = 0.0005 for 3 epochs, sparse categorical crossentropy loss.</li>\n<li>Step 3 – Fine tuning models m1 and m2 using LR = 0.00025 for 2 epochs with different loss function – categorical crossentropy with label smoothing 0.05</li>\n<li>Then we took 4 best models (according to MAE metrics) for both model m1 and m2 (8 models in total) and produced one model using SWA (Stochastic Weight Averaging). We simply averaged model weights. As it appeared the model had the same performance compared for 8 ensembled models but we significantly decreased LSTM inference time. Single model score (public LB) (no TTA): 1.0024.</li>\n</ul>\n<p>Things did not improve our score:</p>\n<ul>\n<li>Dropout in LSTM/GRU layers and linear head.</li>\n<li>GaussianNoise Layer after Masking layer or after LSTM/GRU layer.</li>\n<li>Adam + Lookahead optimizer.</li>\n<li>Gradient Accumulation to simulate TPU big batch size – I (Remek) had problem to implement it properly in TF/Keras (it is easy in Pytorch, as it appeared not exactly easy to implement in TF/Keras – my daily choice is Pytorch)</li>\n</ul>\n<p>Additional tools:</p>\n<ul>\n<li>Weights and biases for two tasks – logging and monitoring training process and swipe for hyperparameter tuning (number of LSTM units, LR scheduler, Optimizer and LR).</li>\n</ul>\n<p>For my (Remek) experiment I use (in first part of competition) ZbyHP Z4 with 2xA5000 but then HP sent me  ZbyHP Z8 workstation with 2x Intel Xeon CPU and Nvidia A6000 GPU. Personally I can say that this help me to establish fast experimentation pipeline. I was able to process dataset files very fast. Having A6000 gave me possibility to train models on bigger batch size. Thank HP for supporting my work.</p>\n<p>My final words - we set up a great team in my opinion - each of us was responsible for part of the solution, we discussed a lot but there was no “my is better”. Although we worked together for the first time, I had the impression that we had known each other forever. Great team, great result! Thank you guys for having the opportunity to learn from great AI guys.</p>\n<h1>Alvor part - blending</h1>\n<p>My main contribution to this competition was the development of methods for blending my teammates’ solutions. Analysis of the different types of solutions (GraphNet, LSTM etc.) showed that they have different efficiency at different predicted zeniths and azimuths values (zenith mainly). So I changed my initial \"constant weight\" approach to a \"bins\" approach.<br>\nThe method consists in splitting the predicted zenith values into 10 bins of equal width. Thus, when combining two solutions, we get 100 bins in total. The value of 10 is configurable, but experiments showed it to be close to optimal.</p>\n<p>After that, the blending weight for the zenith was found in each bin and the blending weight for the azimuth was found. The blending of predicted zenith values was done by a simple linear combination. Blending the azimuth values was a little more complicated, due to the possible transition through 2*pi. Therefore, to begin with, the difference in the predicted azimuths in the two solutions was calculated and the direction of the second value relative to the first value.<br>\nWeights fitting was carried out on a sample of training data close in size to the test data (5 batches with 1M events).</p>\n<p>Several approaches to improve this blending method were also tested. In particular, the use of GBDT.</p>\n<p>One of these approaches was an attempt to classify events into “simple” and “complex” ones (inspired by the kernel <a href=\"https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93)\" target=\"_blank\">https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93)</a>.</p>\n<p>Experiments have shown that some types of neural networks are better at handling simple events, while other types are better at handling complex events. If we could accurately determine the type of each event, this would greatly improve the competition metric. A model was built, the efficiency of which in events classifying was at the level of the public kernel (AUС 0.93), but this was not enough to get a noticeable improvement in the blend quality.</p>\n<p>Several other models of classification (binary variable - which of the two neural networks better predicts a given event) and regression (target variable - the difference in competition metrics from two neural networks for a given event) were also trained. Unfortunately, all these approaches showed only a minor improvement in the result (fourth decimal place), but they were very overfitting-sensitive.</p>\n<p>Later, my method was significantly improved by Wojtek Rosa. Therefore, it was his approach that was used in the final submissions, which he describes in more detail in his part. In particular, I would like to note his brilliant idea of using a Decision Tree model to build blending bins.</p>\n<h1>Wojtek part - more blending</h1>\n<p>Thank you competition host for exciting competition IceCube!<br>\nAlso thank you my Teammates - once again congrats for their prizes, GM titles and great solutions.<br>\nMy part was about blending.<br>\nI used batch_ids 1-5 for evaluation and adjusting parameters, later I used batch_ids 655+ only for evaluation purposes.<br>\nWhen I joined Remek&amp;Alvor team, we have great LSTM solution, public Graphnet, and amazing blend technique:</p>\n<pre><code> ():\n   s[] = np.(s[] - s[])\n   s[] = np.where(\n   s[] &lt; np.pi,\n   np.sign(s[] - s[]),\n   -np.sign(s[] - s[])\n   )\n   N = \n   s[] = (N*np.floor(N*s[]/np.pi) + np.floor(N*s[]/np.pi)).astype()\n</code></pre>\n<p>and minimize score for each bin, finding best qu, alpha such as:</p>\n<pre><code>s0[] = s0[] + alpha * s0[] * s0[]\ns0[] = (-qu) * s0[] + qu * s0[]\n</code></pre>\n<p>I realize, that compared to simple/public vector weight ensembling this method is much better.<br>\nI tried to improve score with greater N values but this leads me to overfit.<br>\nAfter this, I managed to improve the score, by creating bins using Regression Trees with target = score_1 - score_2 and features:</p>\n<pre><code>cls = [,,,,,, ,,,,,,\n      ,,,]\n</code></pre>\n<p>where direction_% are from Isamu submission, zenith_3 is from crazy quick and clean 1.183 Robert linear solution. I tried many different features directly from event, but<br>\nThe simplicity of Decision Tree Regressor gave us (almost) full control of N bins for adjustment allvors parameters qu and alpha:</p>\n<pre><code>regr_1 = DecisionTreeRegressor(max_depth=, min_samples_leaf=, min_samples_split=)\ns[] = regr_1.predict(s[cls])\ns[] = (s[],).astype()\n</code></pre>\n<p>Using DecissionTrees gave us 0.0027 boost both with local score and LB. I tried other methods of blending such as MLP or LGBM without success.<br>\nLater on, I also slightly modified <code>adjusting</code> loop, for finding full linear combination of zenith_1 and zenith_2:</p>\n<pre><code>s0[] = s0[] + alpha * s0[] * s0[]\ns0[] = ru * s0[] + qu * s0[]\n</code></pre>\n<p>It takes much more computation to find ru and qu, but that gives us additional 0.001 improvement.<br>\nOur final blend consist:<br>\nIsamu solution .995 -&gt;  Remek LSTM 1.002 -&gt;  datasaurus Graphnet .982 -&gt; Robert 1.183<br>\nlocal score 0.9774692   Public LB 0.975545, Private LB 0.97658</p>\n<h2>Findings:</h2>\n<ul>\n<li>we found interesting type of error, ‘model give up’ and predicts zenith close to 0.<br>\nOur final solution zenith prediction vs score (x-axis zenith predicted, y-axis zenith ground truth):</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F4a8b117472efd756290083dcf0b787c1%2Fimage1.png?generation=1682343676129581&amp;alt=media\" alt=\"\"></p>\n<p>Ground truth zenith (x-axis) vs score (y-axis) - ‘error triangle’ visible:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F2ee6d7032191782027e45522daff9081%2Fimage2.png?generation=1682343631946316&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>we found that predictions from LSTM classification goes incredibly good at the center of bins, chart: zenith_pred_lstm*1000 (x-axis), mean score (y-axis), zenith edges (red lines)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F7fde090be1100c83a09f045fbf57f5fd%2Fimage3.png?generation=1682343415933098&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, I would like to thank the organizers and staff for hosting such a wonderful competition. And thank you to all my wonderful teammates! @anjum48, @allvor, @remekkinas, @wrrosa. \n\nCongratulations @anjum48!  your promotion to GM!\n\nOur code will be available [here](https://github.com/Anjum48/icecube-neutrinos-in-deep-ice).\n\n\n# Datasaurus Part\n## Preprocessing\n### Raw data\n\nThis dataset was really big, and repeatedly doing pandas operations every epoch would have been a waste of CPU cycles (polars does not appear to work with PyTorch data loaders with many workers yet). To address this, I made PyTorch Geometric `Data` objects for each event and saved them as `.pt` files which could be loaded during training. This took about 8 hours to create using 32 threads, and required about 1TB of space. \n\nThe issue with this was that that 1TB was across 130+ millon tiny files. A Linux partition has a finite number of “index nodes” or `inodes`, i.e. an index to a certain file. Since these files are so small I ran into my inode limit before I ran out of space on my 2TB drive. As a workaround, I had to spread some of these files across two drives, so heads up for anyone trying to reproduce this method or run my code.\n\nA more efficient way could be to store `.pt` files that have already been pre-batched which would require fewer files, but then you lose the ability to shuffle every epoch which may/may not make a difference with this much data. I didn’t try the sqlite method suggested by the GraphNet team.\n\n## Features\nIn the context of GNNs, each DOM is considered as a node. Each node was given the following 11 features:\n\n- X: X location from sensor_geometry.csv / 500\n- Y: Y location from sensor_geometry.csv / 500\n- Z: Z location from sensor_geometry.csv / 500\n- T: (Time from batch_[n].parquet - 1e4) / 3e3\n- Charge: log10(Charge from batch_[n].parquet) / 3.0\n- QE: See below\n- Aux: (False = -0.5, True = 0.5)\n- Scattering Length: See below\n- Distance to previous hit: See below\n- Time delta since previous hit: See below\n- Scattering flag (False = -0.5, True = 0.5). See below\n\nMany of the normalisation methods were taken from GraphNet as a [starting point] (https://github.com/graphnet-team/graphnet/blob/4df8f396400da3cfca4ff1e0593a0c7d1b5b5195/src/graphnet/models/detector/icecube.py#L64-L69), but I altered the scale for time, since time is a very important feature here.\n\nFor events with large numbers of hits, to prevent OOM errors, I sampled 256 hits. This can make the process slightly non-deterministic.\n\n## Quantum efficiency\nQE is the quantum efficiency of the photomutipliers in the DOMs. The DeepCore DOMs are quoted to have 35% higher QE than the regular DOMs (Figure 1 of this [paper](https://arxiv.org/pdf/2209.03042.pdf)), so QE was set to 1 everywhere, and 1.35 for the lower 50 DOMs in DeepCore. The final QE feature was scaled using (QE - 1.25) / 0.25.\n\n## Scattering length\nScattering and absorption lengths are important to characterise differences in the clarity of the ice. This data is published on page 31 of this [paper](https://arxiv.org/abs/1301.5361). A datum depth of 1920 metres was used so that z = (depth - 1920) / 500. The data was resampled to the z values using `scipy.interpolate.interp1d`. I found that after passing the data though `RobustScaler`, the scattering and absorption data was near identical, so I only used scattering length.\n\n## Previous hit features\nThe two main types of events are track and cascade events. Looking at some of the amazing [visualisation tools](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388858) for example from edguy99, I got the idea that if a node had some understanding where and when the nearest previous hit was, it might help the model differentiate between these two groups. To calculate this for each event, sorted the hits by time, calculated the pairwise distances of all hits, masked any hits from the future and calculated the distance, d, to the nearest previous hit. This was scaled using (d - 0.5) / 0.5. The time delta from the previous hit was also calculated using the same method and scaled using (t - 0.1) / 0.1.\n\n## Scattering flag\nI tried to create a flag that could discern whether a hit was caused directly from a track, or some secondary scattering, inspired by section 2.1 of this [paper](https://arxiv.org/pdf/2203.02303.pdf). A side effect of adding this flag was that training was much more stable. The flag is generated as follows:\n1. Identify the hit with the largest charge\n2. From this DOM location, calculate the distances & time delta to every other hit\n3. If the time taken to travel that distance is > speed of light in ice, assume that the photon is a result of scattering\n\n## Validation\nI used 90% - 10% train-validation split, and no cross validation due to the size of the dataset. The split was done by creating 10 bins of log10(n_hits), and then using `StratifiedKFold` using 10 splits.\n\n## Models\nI used the `DirectionReconstructionWithKappa` task directly from GraphNet, meaning that an embedding (e.g. shape of 128) will be projected to a shape of 4 (x, y, z, kappa)\n\n## Architectures\nI used the following 3 architectures. All validation scores are with 6x TTA applied\n\n[GraphNet/DynEdge](https://github.com/graphnet-team/graphnet) -  Val = 0.98501\n[GPS](https://arxiv.org/abs/2205.12454) - Val = 0.98945*\n[GravNet](https://arxiv.org/abs/1902.07987) - Val = 0.98519\n\nThe average of these 3 models gave a 0.982 LB score.\n\nThe GraphNet/DynEdge model had very little modfification, other than changing to GELU activations.\n\nGPS & GravNet used 8 blocks and you can find the exact architectures for both in the code [here](https://github.com/Anjum48/icecube-neutrinos-in-deep-ice/blob/main/src/modules.py).\n\n*GPS was the most powerful model, but also slowest to train being a transformer type model (roughly 11 hours/epoch on my machine). I managed to train a model which achieved a validation score of 0.98XX but was too late to include in our final submission.\n\n## Loss\nI used VonMisesFisher3DLoss + (1 - CosineSimilarity) as the final loss function, since cosine similarity is a nice proxy for mean angular error. For CosineSimilarity I transformed the target azimuth & zenith values to cartesian coordinates.\n\nThis performed much better than separate losses for azimuth (VMF2D) & zenith (MSE) which I the route I [initially went down](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383546).\n\n## Augmentation\nI centered the data on string 35 (the DeepCore string) and rotated about the z-axis in 60 degree steps. This didn’t actually improve validation performance but did have the benefit of making the models rotationally invariant so that I could take advantage of the detector symmetry and apply a 6x test time augmentation (TTA). This often improved scores by 0.002-0.003. \n\n## Training parameters\n- AdamW optimiser\n- Epochs = 6\n- Cosine schedule (no warmup)\n- Learning rate = 0.0002\n- Batch size = 1024\n- Weight decay = 0.001 - 0.1 depending on model\n- FP16 training\n- Hardware: 2x RTX 3090, 128 GB RAM\n\n## Final submissions\nCircular mean\n\n## Robustness to perturbation\nTBC\n\n## Lessons learned/stuff that didn’t work\n- The GraphNet DynEdge baseline is extremely strong and tough to improve on - kudos to the team! It is also the fastest/efficient model, and what I used for the majority of experimentation\n- More data = more better. The issue with this though is that I found that some conclusions drawn from experiments on 1% or 5% of the data were no longer applicable on the full dataset. This made experimentation slow and expensive\n- Batch normalisation made things unstable and didn’t show improvements\n- Lion optimiser didn’t generalise as well as AdamW\n- Weight decay was important for some models. As a result I assumed changing the epsilon value in Adam would have an effect, but I didn’t see anything significant\n- In my experiments, GNNs seem to benefit from leaky activations, e.g. GELU\n- For MPNN aggregation, it seems that [min, max, mean, sum] is sufficient. Adding more didn’t appear to make significant gains\n- Realigning all of the times to the time of the first hit of each event deteriorates performance, possibly due to noise in the data/false triggers etc.\n- Radial nearest neighbours didn’t work any better than KNN when defining graph edges\n- Only using 1 - CosineSimilarity as a loss function wasn’t very stable. Adding VMF3D helped a lot\n\n## Code\nAll my code will be available here soon: https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\n\n\n# Isamu Part\n\nBefore the team merge, I created LSTM and GraphNet models. After the team merge, I focused on GraphNet because Remek's LSTM model was superior to mine. I used graphnet (https://github.com/graphnet-team/graphnet) as a baseline and made several changes to improve its accuracy. I list below some of the experiments I performed that worked(There are tons of things that didn't work)\n\n- random sampling (random sampling from DB if the specified data length exceeds 800)\n- Increasing nearest neighbors of KNN layer(8->16)\n- Addition of features\n  - x, y, z\n  - time\n  - charge\n  - auxiliary\n  - ice_transparency feature\n- 2-stage model with kappa(sigma) of vonMisesFisher distribution \n  - Train 1st stage model to predict x, y, z, kappa\n      - restart from the weights of GraphNet from public baseline notebook\n      - About 1-250 batches were used\n      - batch size 512\n      - epoch 20\n      - DirectionReconstructionWithKappa\n  - Split data into easy and hard parts according to 1st stage kappa value\n       - Inference was performed using the 1st model and classified into two sets of data(easy part and hard part) according to their predicted kappa value\n  - Train expert models for easy and hard parts and combine their predictions\n      - About 250-350 batches were used\n      - batch size 512\n      - epoch 20\n      - DirectionReconstructionWithKappa\n- TTA \n  - rotation 180-degree TTA about the z-axis\n- Loss\n  - DirectionReconstructionWithKappa\n- Hardware: RTX 3090, 64 GB RAM, 8TB HDD, 2TB SSD, Google Colab Pro\n\nThe ensemble of models(1st and 2nd) created above gave public LB 0.995669, private LB 0.996550 \n\n# Remek Part - LSTM\nFor LSTM training we used the attitude proposed by Robin Smits (@rsmits) with some improvements.\n- We added more LSTM/GRU layers (4 GRU/LSTM) - more and less than 4 layers was not better in our experiments.\n- We added ice transparency as additional features (we used both features – transparency and absorption).\n- We redesigned the training loop to train models on all batches – we can train models on different parts of DS (from one batch to all DS in one epoch).\n- We checked many hypotheses (for part of them weight and biases Swipe tool was used):\n\n\t- Different size of LSTM units – finally 196 was the best in our model.\n\t- Different bin size – 24 was our final choice.\n\t- Different amount of features and impulse selection - finally we use strategy - first select non_aux events then add random aux events (it gave us score boost as well - this was kind of an augmentation technique).\n\t- Different model architectures – finally for our score blend we used pure LSTM/GRU setup. We tested transformer architecture but as it appeared we gave up too early (first scores were way worse then pure LSTM).\n\t- Different optimizers (AdamW, NAdam) – final choice was Adam.\n\t- Different schedulers – CosineDecay, OneCycle but then we use step LR scheduling described below.\n \nTraining was divided  into three parts scheduling LR:\n- Step 1 - Train baseline models (two models) – batches 4-330 (m1) and 331-660 (m2) using LR = 0.005, for 6-8 epochs, sparse categorical crossentropy loss. Both models were validated on batches 1-3.\n- Step 2 - Fine tuning models m1 and m2 using LR/10 = 0.0005 for 3 epochs, sparse categorical crossentropy loss.\n- Step 3 – Fine tuning models m1 and m2 using LR = 0.00025 for 2 epochs with different loss function – categorical crossentropy with label smoothing 0.05\n- Then we took 4 best models (according to MAE metrics) for both model m1 and m2 (8 models in total) and produced one model using SWA (Stochastic Weight Averaging). We simply averaged model weights. As it appeared the model had the same performance compared for 8 ensembled models but we significantly decreased LSTM inference time. Single model score (public LB) (no TTA): 1.0024.\n \nThings did not improve our score:\n- Dropout in LSTM/GRU layers and linear head.\n- GaussianNoise Layer after Masking layer or after LSTM/GRU layer.\n- Adam + Lookahead optimizer.\n- Gradient Accumulation to simulate TPU big batch size – I (Remek) had problem to implement it properly in TF/Keras (it is easy in Pytorch, as it appeared not exactly easy to implement in TF/Keras – my daily choice is Pytorch)\n \nAdditional tools:\n- Weights and biases for two tasks – logging and monitoring training process and swipe for hyperparameter tuning (number of LSTM units, LR scheduler, Optimizer and LR).\n \nFor my (Remek) experiment I use (in first part of competition) ZbyHP Z4 with 2xA5000 but then HP sent me  ZbyHP Z8 workstation with 2x Intel Xeon CPU and Nvidia A6000 GPU. Personally I can say that this help me to establish fast experimentation pipeline. I was able to process dataset files very fast. Having A6000 gave me possibility to train models on bigger batch size. Thank HP for supporting my work.\n\nMy final words - we set up a great team in my opinion - each of us was responsible for part of the solution, we discussed a lot but there was no “my is better”. Although we worked together for the first time, I had the impression that we had known each other forever. Great team, great result! Thank you guys for having the opportunity to learn from great AI guys.\n\n# Alvor part - blending\nMy main contribution to this competition was the development of methods for blending my teammates’ solutions. Analysis of the different types of solutions (GraphNet, LSTM etc.) showed that they have different efficiency at different predicted zeniths and azimuths values (zenith mainly). So I changed my initial \"constant weight\" approach to a \"bins\" approach.\nThe method consists in splitting the predicted zenith values into 10 bins of equal width. Thus, when combining two solutions, we get 100 bins in total. The value of 10 is configurable, but experiments showed it to be close to optimal.\n\nAfter that, the blending weight for the zenith was found in each bin and the blending weight for the azimuth was found. The blending of predicted zenith values was done by a simple linear combination. Blending the azimuth values was a little more complicated, due to the possible transition through 2*pi. Therefore, to begin with, the difference in the predicted azimuths in the two solutions was calculated and the direction of the second value relative to the first value.\nWeights fitting was carried out on a sample of training data close in size to the test data (5 batches with 1M events).\n\nSeveral approaches to improve this blending method were also tested. In particular, the use of GBDT.\n\nOne of these approaches was an attempt to classify events into “simple” and “complex” ones (inspired by the kernel https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93).\n\nExperiments have shown that some types of neural networks are better at handling simple events, while other types are better at handling complex events. If we could accurately determine the type of each event, this would greatly improve the competition metric. A model was built, the efficiency of which in events classifying was at the level of the public kernel (AUС 0.93), but this was not enough to get a noticeable improvement in the blend quality.\n\nSeveral other models of classification (binary variable - which of the two neural networks better predicts a given event) and regression (target variable - the difference in competition metrics from two neural networks for a given event) were also trained. Unfortunately, all these approaches showed only a minor improvement in the result (fourth decimal place), but they were very overfitting-sensitive.\n\nLater, my method was significantly improved by Wojtek Rosa. Therefore, it was his approach that was used in the final submissions, which he describes in more detail in his part. In particular, I would like to note his brilliant idea of using a Decision Tree model to build blending bins.\n\n\n# Wojtek part - more blending\nThank you competition host for exciting competition IceCube!\nAlso thank you my Teammates - once again congrats for their prizes, GM titles and great solutions.\nMy part was about blending.\nI used batch_ids 1-5 for evaluation and adjusting parameters, later I used batch_ids 655+ only for evaluation purposes.\nWhen I joined Remek&Alvor team, we have great LSTM solution, public Graphnet, and amazing blend technique:\n```python\ndef alvor_blend(s):\n   s['diff'] = np.abs(s['azimuth_2'] - s['azimuth_1'])\n   s['direction'] = np.where(\n   s['diff'] < np.pi,\n   np.sign(s['azimuth_2'] - s['azimuth_1']),\n   -np.sign(s['azimuth_2'] - s['azimuth_1'])\n   )\n   N = 10\n   s['bin'] = (N*np.floor(N*s['zenith_1']/np.pi) + np.floor(N*s['zenith_2']/np.pi)).astype(int)\n```\nand minimize score for each bin, finding best qu, alpha such as:\n```python\ns0['azimuth_pred'] = s0['azimuth_1'] + alpha * s0['diff'] * s0['direction']\ns0['zenith_pred'] = (1-qu) * s0['zenith_1'] + qu * s0['zenith_2']\n```\nI realize, that compared to simple/public vector weight ensembling this method is much better.\nI tried to improve score with greater N values but this leads me to overfit.\nAfter this, I managed to improve the score, by creating bins using Regression Trees with target = score_1 - score_2 and features:\n```python\ncls = ['azimuth_1','zenith_1','azimuth_2','zenith_2','diff','direction', 'n_pulses','zenith_1_2','direction_x','direction_y','direction_z','direction_kappa',\n  \t'zenith_3','azimuth_3','score_1_3','score_2_3']\n```\nwhere direction_% are from Isamu submission, zenith_3 is from crazy quick and clean 1.183 Robert linear solution. I tried many different features directly from event, but\nThe simplicity of Decision Tree Regressor gave us (almost) full control of N bins for adjustment allvors parameters qu and alpha:\n```python\nregr_1 = DecisionTreeRegressor(max_depth=11, min_samples_leaf=500, min_samples_split=500)\ns['bin_raw'] = regr_1.predict(s[cls])\ns['bin'] = round(s['bin_raw'],5).astype(str)\n```\nUsing DecissionTrees gave us 0.0027 boost both with local score and LB. I tried other methods of blending such as MLP or LGBM without success.\nLater on, I also slightly modified `adjusting` loop, for finding full linear combination of zenith_1 and zenith_2:\n```python\ns0['azimuth_pred'] = s0['azimuth_1'] + alpha * s0['diff'] * s0['direction']\ns0['zenith_pred'] = ru * s0['zenith_1'] + qu * s0['zenith_2']\n```\nIt takes much more computation to find ru and qu, but that gives us additional 0.001 improvement.\nOur final blend consist:\nIsamu solution .995 ->  Remek LSTM 1.002 ->  datasaurus Graphnet .982 -> Robert 1.183\nlocal score 0.9774692   Public LB 0.975545, Private LB 0.97658\n## Findings:\n- we found interesting type of error, ‘model give up’ and predicts zenith close to 0.\nOur final solution zenith prediction vs score (x-axis zenith predicted, y-axis zenith ground truth):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F4a8b117472efd756290083dcf0b787c1%2Fimage1.png?generation=1682343676129581&alt=media)\n\nGround truth zenith (x-axis) vs score (y-axis) - ‘error triangle’ visible:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F2ee6d7032191782027e45522daff9081%2Fimage2.png?generation=1682343631946316&alt=media)\n\n- we found that predictions from LSTM classification goes incredibly good at the center of bins, chart: zenith_pred_lstm*1000 (x-axis), mean score (y-axis), zenith edges (red lines)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F7fde090be1100c83a09f045fbf57f5fd%2Fimage3.png?generation=1682343415933098&alt=media)\n",
      "votes": 23
    },
    {
      "id": 2234115,
      "postDate": "2023-04-24T22:21:45.877Z",
      "content": "<p>This topic is very informative and clearly explained. I learned a lot from reading this post. You have done a great job of presenting the concepts in a simple and engaging way. I appreciate your effort and knowledge.</p>\n<p>I am going to bookmark this post for future reference. This is a valuable resource that I want to revisit and review. Thank you for sharing this with us. You and your teammates are awesome <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>! 😊</p>",
      "rawMarkdown": "This topic is very informative and clearly explained. I learned a lot from reading this post. You have done a great job of presenting the concepts in a simple and engaging way. I appreciate your effort and knowledge.\n\nI am going to bookmark this post for future reference. This is a valuable resource that I want to revisit and review. Thank you for sharing this with us. You and your teammates are awesome @yamsam! 😊",
      "votes": 1
    },
    {
      "id": 2234866,
      "postDate": "2023-04-25T14:37:50.237Z",
      "content": "<p>Congratulations!<br>\nI really wonder how you could find good models such as GPS and GravNet.<br>\nI also searched about models but could not find good models.</p>\n<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>'s notebook on GNN was very helpful. Thank you very much!</p>",
      "rawMarkdown": "Congratulations!\nI really wonder how you could find good models such as GPS and GravNet.\nI also searched about models but could not find good models.\n\n@anjum48's notebook on GNN was very helpful. Thank you very much!",
      "votes": 2
    },
    {
      "id": 2233975,
      "postDate": "2023-04-24T18:30:23.113Z",
      "content": "<p><a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> thanks for the detailed topic! Enjoyed reading</p>",
      "rawMarkdown": "@yamsam thanks for the detailed topic! Enjoyed reading",
      "votes": 2
    },
    {
      "id": 2941059,
      "postDate": "2024-07-30T16:15:08.530Z",
      "content": "<p>Congrats to <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> and the team for the impressive 8th place finish in the IceCube Neutrino Observatory competition! Your approach to preprocessing, model architectures, and blending is truly commendable. The detailed insights on handling large datasets and optimizing model performance are incredibly valuable. A special shoutout to <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> for the GM promotion and <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, and <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> for their contributions. Looking forward to exploring the code and learning more from your solutions. <br>\nThanks for sharing! </p>",
      "rawMarkdown": "Congrats to @yamsam and the team for the impressive 8th place finish in the IceCube Neutrino Observatory competition! Your approach to preprocessing, model architectures, and blending is truly commendable. The detailed insights on handling large datasets and optimizing model performance are incredibly valuable. A special shoutout to @anjum48 for the GM promotion and @allvor, @remekkinas, and @wrrosa for their contributions. Looking forward to exploring the code and learning more from your solutions. \nThanks for sharing! "
    },
    {
      "id": 2234993,
      "postDate": "2023-04-25T16:11:24.557Z",
      "content": "<p>Congratulations on the achievement!</p>",
      "rawMarkdown": "Congratulations on the achievement!"
    },
    {
      "id": 2363820,
      "postDate": "2023-07-29T00:51:03.970Z",
      "content": "<p>Congratulations！Amazing work with informative figure and clear explaination!<br>\nThanks for your sharing! <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> </p>",
      "rawMarkdown": "Congratulations！Amazing work with informative figure and clear explaination!\nThanks for your sharing! @yamsam ",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2234115,
      "author_name": "Hozaifa Ahmed",
      "author_url": "",
      "post_date": "2023-04-24T22:21:45.877000",
      "content": "<p>This topic is very informative and clearly explained. I learned a lot from reading this post. You have done a great job of presenting the concepts in a simple and engaging way. I appreciate your effort and knowledge.</p>\n<p>I am going to bookmark this post for future reference. This is a valuable resource that I want to revisit and review. Thank you for sharing this with us. You and your teammates are awesome <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a>! 😊</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2234866,
      "author_name": "moto",
      "author_url": "",
      "post_date": "2023-04-25T14:37:50.237000",
      "content": "<p>Congratulations!<br>\nI really wonder how you could find good models such as GPS and GravNet.<br>\nI also searched about models but could not find good models.</p>\n<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>'s notebook on GNN was very helpful. Thank you very much!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2233975,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "2023-04-24T18:30:23.113000",
      "content": "<p><a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> thanks for the detailed topic! Enjoyed reading</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2941059,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T16:15:08.530000",
      "content": "<p>Congrats to <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> and the team for the impressive 8th place finish in the IceCube Neutrino Observatory competition! Your approach to preprocessing, model architectures, and blending is truly commendable. The detailed insights on handling large datasets and optimizing model performance are incredibly valuable. A special shoutout to <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> for the GM promotion and <a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a>, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, and <a href=\"https://www.kaggle.com/wrrosa\" target=\"_blank\">@wrrosa</a> for their contributions. Looking forward to exploring the code and learning more from your solutions. <br>\nThanks for sharing! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2234993,
      "author_name": "Ericka42",
      "author_url": "",
      "post_date": "2023-04-25T16:11:24.557000",
      "content": "<p>Congratulations on the achievement!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2363820,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-29T00:51:03.970000",
      "content": "<p>Congratulations！Amazing work with informative figure and clear explaination!<br>\nThanks for your sharing! <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2232604": "First of all, I would like to thank the organizers and staff for hosting such a wonderful competition. And thank you to all my wonderful teammates! @anjum48, @allvor, @remekkinas, @wrrosa. \n\nCongratulations @anjum48!  your promotion to GM!\n\nOur code will be available [here](https://github.com/Anjum48/icecube-neutrinos-in-deep-ice).\n\n\n# Datasaurus Part\n## Preprocessing\n### Raw data\n\nThis dataset was really big, and repeatedly doing pandas operations every epoch would have been a waste of CPU cycles (polars does not appear to work with PyTorch data loaders with many workers yet). To address this, I made PyTorch Geometric `Data` objects for each event and saved them as `.pt` files which could be loaded during training. This took about 8 hours to create using 32 threads, and required about 1TB of space. \n\nThe issue with this was that that 1TB was across 130+ millon tiny files. A Linux partition has a finite number of “index nodes” or `inodes`, i.e. an index to a certain file. Since these files are so small I ran into my inode limit before I ran out of space on my 2TB drive. As a workaround, I had to spread some of these files across two drives, so heads up for anyone trying to reproduce this method or run my code.\n\nA more efficient way could be to store `.pt` files that have already been pre-batched which would require fewer files, but then you lose the ability to shuffle every epoch which may/may not make a difference with this much data. I didn’t try the sqlite method suggested by the GraphNet team.\n\n## Features\nIn the context of GNNs, each DOM is considered as a node. Each node was given the following 11 features:\n\n- X: X location from sensor_geometry.csv / 500\n- Y: Y location from sensor_geometry.csv / 500\n- Z: Z location from sensor_geometry.csv / 500\n- T: (Time from batch_[n].parquet - 1e4) / 3e3\n- Charge: log10(Charge from batch_[n].parquet) / 3.0\n- QE: See below\n- Aux: (False = -0.5, True = 0.5)\n- Scattering Length: See below\n- Distance to previous hit: See below\n- Time delta since previous hit: See below\n- Scattering flag (False = -0.5, True = 0.5). See below\n\nMany of the normalisation methods were taken from GraphNet as a [starting point] (https://github.com/graphnet-team/graphnet/blob/4df8f396400da3cfca4ff1e0593a0c7d1b5b5195/src/graphnet/models/detector/icecube.py#L64-L69), but I altered the scale for time, since time is a very important feature here.\n\nFor events with large numbers of hits, to prevent OOM errors, I sampled 256 hits. This can make the process slightly non-deterministic.\n\n## Quantum efficiency\nQE is the quantum efficiency of the photomutipliers in the DOMs. The DeepCore DOMs are quoted to have 35% higher QE than the regular DOMs (Figure 1 of this [paper](https://arxiv.org/pdf/2209.03042.pdf)), so QE was set to 1 everywhere, and 1.35 for the lower 50 DOMs in DeepCore. The final QE feature was scaled using (QE - 1.25) / 0.25.\n\n## Scattering length\nScattering and absorption lengths are important to characterise differences in the clarity of the ice. This data is published on page 31 of this [paper](https://arxiv.org/abs/1301.5361). A datum depth of 1920 metres was used so that z = (depth - 1920) / 500. The data was resampled to the z values using `scipy.interpolate.interp1d`. I found that after passing the data though `RobustScaler`, the scattering and absorption data was near identical, so I only used scattering length.\n\n## Previous hit features\nThe two main types of events are track and cascade events. Looking at some of the amazing [visualisation tools](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388858) for example from edguy99, I got the idea that if a node had some understanding where and when the nearest previous hit was, it might help the model differentiate between these two groups. To calculate this for each event, sorted the hits by time, calculated the pairwise distances of all hits, masked any hits from the future and calculated the distance, d, to the nearest previous hit. This was scaled using (d - 0.5) / 0.5. The time delta from the previous hit was also calculated using the same method and scaled using (t - 0.1) / 0.1.\n\n## Scattering flag\nI tried to create a flag that could discern whether a hit was caused directly from a track, or some secondary scattering, inspired by section 2.1 of this [paper](https://arxiv.org/pdf/2203.02303.pdf). A side effect of adding this flag was that training was much more stable. The flag is generated as follows:\n1. Identify the hit with the largest charge\n2. From this DOM location, calculate the distances & time delta to every other hit\n3. If the time taken to travel that distance is > speed of light in ice, assume that the photon is a result of scattering\n\n## Validation\nI used 90% - 10% train-validation split, and no cross validation due to the size of the dataset. The split was done by creating 10 bins of log10(n_hits), and then using `StratifiedKFold` using 10 splits.\n\n## Models\nI used the `DirectionReconstructionWithKappa` task directly from GraphNet, meaning that an embedding (e.g. shape of 128) will be projected to a shape of 4 (x, y, z, kappa)\n\n## Architectures\nI used the following 3 architectures. All validation scores are with 6x TTA applied\n\n[GraphNet/DynEdge](https://github.com/graphnet-team/graphnet) -  Val = 0.98501\n[GPS](https://arxiv.org/abs/2205.12454) - Val = 0.98945*\n[GravNet](https://arxiv.org/abs/1902.07987) - Val = 0.98519\n\nThe average of these 3 models gave a 0.982 LB score.\n\nThe GraphNet/DynEdge model had very little modfification, other than changing to GELU activations.\n\nGPS & GravNet used 8 blocks and you can find the exact architectures for both in the code [here](https://github.com/Anjum48/icecube-neutrinos-in-deep-ice/blob/main/src/modules.py).\n\n*GPS was the most powerful model, but also slowest to train being a transformer type model (roughly 11 hours/epoch on my machine). I managed to train a model which achieved a validation score of 0.98XX but was too late to include in our final submission.\n\n## Loss\nI used VonMisesFisher3DLoss + (1 - CosineSimilarity) as the final loss function, since cosine similarity is a nice proxy for mean angular error. For CosineSimilarity I transformed the target azimuth & zenith values to cartesian coordinates.\n\nThis performed much better than separate losses for azimuth (VMF2D) & zenith (MSE) which I the route I [initially went down](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383546).\n\n## Augmentation\nI centered the data on string 35 (the DeepCore string) and rotated about the z-axis in 60 degree steps. This didn’t actually improve validation performance but did have the benefit of making the models rotationally invariant so that I could take advantage of the detector symmetry and apply a 6x test time augmentation (TTA). This often improved scores by 0.002-0.003. \n\n## Training parameters\n- AdamW optimiser\n- Epochs = 6\n- Cosine schedule (no warmup)\n- Learning rate = 0.0002\n- Batch size = 1024\n- Weight decay = 0.001 - 0.1 depending on model\n- FP16 training\n- Hardware: 2x RTX 3090, 128 GB RAM\n\n## Final submissions\nCircular mean\n\n## Robustness to perturbation\nTBC\n\n## Lessons learned/stuff that didn’t work\n- The GraphNet DynEdge baseline is extremely strong and tough to improve on - kudos to the team! It is also the fastest/efficient model, and what I used for the majority of experimentation\n- More data = more better. The issue with this though is that I found that some conclusions drawn from experiments on 1% or 5% of the data were no longer applicable on the full dataset. This made experimentation slow and expensive\n- Batch normalisation made things unstable and didn’t show improvements\n- Lion optimiser didn’t generalise as well as AdamW\n- Weight decay was important for some models. As a result I assumed changing the epsilon value in Adam would have an effect, but I didn’t see anything significant\n- In my experiments, GNNs seem to benefit from leaky activations, e.g. GELU\n- For MPNN aggregation, it seems that [min, max, mean, sum] is sufficient. Adding more didn’t appear to make significant gains\n- Realigning all of the times to the time of the first hit of each event deteriorates performance, possibly due to noise in the data/false triggers etc.\n- Radial nearest neighbours didn’t work any better than KNN when defining graph edges\n- Only using 1 - CosineSimilarity as a loss function wasn’t very stable. Adding VMF3D helped a lot\n\n## Code\nAll my code will be available here soon: https://github.com/Anjum48/icecube-neutrinos-in-deep-ice\n\n\n# Isamu Part\n\nBefore the team merge, I created LSTM and GraphNet models. After the team merge, I focused on GraphNet because Remek's LSTM model was superior to mine. I used graphnet (https://github.com/graphnet-team/graphnet) as a baseline and made several changes to improve its accuracy. I list below some of the experiments I performed that worked(There are tons of things that didn't work)\n\n- random sampling (random sampling from DB if the specified data length exceeds 800)\n- Increasing nearest neighbors of KNN layer(8->16)\n- Addition of features\n  - x, y, z\n  - time\n  - charge\n  - auxiliary\n  - ice_transparency feature\n- 2-stage model with kappa(sigma) of vonMisesFisher distribution \n  - Train 1st stage model to predict x, y, z, kappa\n      - restart from the weights of GraphNet from public baseline notebook\n      - About 1-250 batches were used\n      - batch size 512\n      - epoch 20\n      - DirectionReconstructionWithKappa\n  - Split data into easy and hard parts according to 1st stage kappa value\n       - Inference was performed using the 1st model and classified into two sets of data(easy part and hard part) according to their predicted kappa value\n  - Train expert models for easy and hard parts and combine their predictions\n      - About 250-350 batches were used\n      - batch size 512\n      - epoch 20\n      - DirectionReconstructionWithKappa\n- TTA \n  - rotation 180-degree TTA about the z-axis\n- Loss\n  - DirectionReconstructionWithKappa\n- Hardware: RTX 3090, 64 GB RAM, 8TB HDD, 2TB SSD, Google Colab Pro\n\nThe ensemble of models(1st and 2nd) created above gave public LB 0.995669, private LB 0.996550 \n\n# Remek Part - LSTM\nFor LSTM training we used the attitude proposed by Robin Smits (@rsmits) with some improvements.\n- We added more LSTM/GRU layers (4 GRU/LSTM) - more and less than 4 layers was not better in our experiments.\n- We added ice transparency as additional features (we used both features – transparency and absorption).\n- We redesigned the training loop to train models on all batches – we can train models on different parts of DS (from one batch to all DS in one epoch).\n- We checked many hypotheses (for part of them weight and biases Swipe tool was used):\n\n\t- Different size of LSTM units – finally 196 was the best in our model.\n\t- Different bin size – 24 was our final choice.\n\t- Different amount of features and impulse selection - finally we use strategy - first select non_aux events then add random aux events (it gave us score boost as well - this was kind of an augmentation technique).\n\t- Different model architectures – finally for our score blend we used pure LSTM/GRU setup. We tested transformer architecture but as it appeared we gave up too early (first scores were way worse then pure LSTM).\n\t- Different optimizers (AdamW, NAdam) – final choice was Adam.\n\t- Different schedulers – CosineDecay, OneCycle but then we use step LR scheduling described below.\n \nTraining was divided  into three parts scheduling LR:\n- Step 1 - Train baseline models (two models) – batches 4-330 (m1) and 331-660 (m2) using LR = 0.005, for 6-8 epochs, sparse categorical crossentropy loss. Both models were validated on batches 1-3.\n- Step 2 - Fine tuning models m1 and m2 using LR/10 = 0.0005 for 3 epochs, sparse categorical crossentropy loss.\n- Step 3 – Fine tuning models m1 and m2 using LR = 0.00025 for 2 epochs with different loss function – categorical crossentropy with label smoothing 0.05\n- Then we took 4 best models (according to MAE metrics) for both model m1 and m2 (8 models in total) and produced one model using SWA (Stochastic Weight Averaging). We simply averaged model weights. As it appeared the model had the same performance compared for 8 ensembled models but we significantly decreased LSTM inference time. Single model score (public LB) (no TTA): 1.0024.\n \nThings did not improve our score:\n- Dropout in LSTM/GRU layers and linear head.\n- GaussianNoise Layer after Masking layer or after LSTM/GRU layer.\n- Adam + Lookahead optimizer.\n- Gradient Accumulation to simulate TPU big batch size – I (Remek) had problem to implement it properly in TF/Keras (it is easy in Pytorch, as it appeared not exactly easy to implement in TF/Keras – my daily choice is Pytorch)\n \nAdditional tools:\n- Weights and biases for two tasks – logging and monitoring training process and swipe for hyperparameter tuning (number of LSTM units, LR scheduler, Optimizer and LR).\n \nFor my (Remek) experiment I use (in first part of competition) ZbyHP Z4 with 2xA5000 but then HP sent me  ZbyHP Z8 workstation with 2x Intel Xeon CPU and Nvidia A6000 GPU. Personally I can say that this help me to establish fast experimentation pipeline. I was able to process dataset files very fast. Having A6000 gave me possibility to train models on bigger batch size. Thank HP for supporting my work.\n\nMy final words - we set up a great team in my opinion - each of us was responsible for part of the solution, we discussed a lot but there was no “my is better”. Although we worked together for the first time, I had the impression that we had known each other forever. Great team, great result! Thank you guys for having the opportunity to learn from great AI guys.\n\n# Alvor part - blending\nMy main contribution to this competition was the development of methods for blending my teammates’ solutions. Analysis of the different types of solutions (GraphNet, LSTM etc.) showed that they have different efficiency at different predicted zeniths and azimuths values (zenith mainly). So I changed my initial \"constant weight\" approach to a \"bins\" approach.\nThe method consists in splitting the predicted zenith values into 10 bins of equal width. Thus, when combining two solutions, we get 100 bins in total. The value of 10 is configurable, but experiments showed it to be close to optimal.\n\nAfter that, the blending weight for the zenith was found in each bin and the blending weight for the azimuth was found. The blending of predicted zenith values was done by a simple linear combination. Blending the azimuth values was a little more complicated, due to the possible transition through 2*pi. Therefore, to begin with, the difference in the predicted azimuths in the two solutions was calculated and the direction of the second value relative to the first value.\nWeights fitting was carried out on a sample of training data close in size to the test data (5 batches with 1M events).\n\nSeveral approaches to improve this blending method were also tested. In particular, the use of GBDT.\n\nOne of these approaches was an attempt to classify events into “simple” and “complex” ones (inspired by the kernel https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93).\n\nExperiments have shown that some types of neural networks are better at handling simple events, while other types are better at handling complex events. If we could accurately determine the type of each event, this would greatly improve the competition metric. A model was built, the efficiency of which in events classifying was at the level of the public kernel (AUС 0.93), but this was not enough to get a noticeable improvement in the blend quality.\n\nSeveral other models of classification (binary variable - which of the two neural networks better predicts a given event) and regression (target variable - the difference in competition metrics from two neural networks for a given event) were also trained. Unfortunately, all these approaches showed only a minor improvement in the result (fourth decimal place), but they were very overfitting-sensitive.\n\nLater, my method was significantly improved by Wojtek Rosa. Therefore, it was his approach that was used in the final submissions, which he describes in more detail in his part. In particular, I would like to note his brilliant idea of using a Decision Tree model to build blending bins.\n\n\n# Wojtek part - more blending\nThank you competition host for exciting competition IceCube!\nAlso thank you my Teammates - once again congrats for their prizes, GM titles and great solutions.\nMy part was about blending.\nI used batch_ids 1-5 for evaluation and adjusting parameters, later I used batch_ids 655+ only for evaluation purposes.\nWhen I joined Remek&Alvor team, we have great LSTM solution, public Graphnet, and amazing blend technique:\n```python\ndef alvor_blend(s):\n   s['diff'] = np.abs(s['azimuth_2'] - s['azimuth_1'])\n   s['direction'] = np.where(\n   s['diff'] < np.pi,\n   np.sign(s['azimuth_2'] - s['azimuth_1']),\n   -np.sign(s['azimuth_2'] - s['azimuth_1'])\n   )\n   N = 10\n   s['bin'] = (N*np.floor(N*s['zenith_1']/np.pi) + np.floor(N*s['zenith_2']/np.pi)).astype(int)\n```\nand minimize score for each bin, finding best qu, alpha such as:\n```python\ns0['azimuth_pred'] = s0['azimuth_1'] + alpha * s0['diff'] * s0['direction']\ns0['zenith_pred'] = (1-qu) * s0['zenith_1'] + qu * s0['zenith_2']\n```\nI realize, that compared to simple/public vector weight ensembling this method is much better.\nI tried to improve score with greater N values but this leads me to overfit.\nAfter this, I managed to improve the score, by creating bins using Regression Trees with target = score_1 - score_2 and features:\n```python\ncls = ['azimuth_1','zenith_1','azimuth_2','zenith_2','diff','direction', 'n_pulses','zenith_1_2','direction_x','direction_y','direction_z','direction_kappa',\n  \t'zenith_3','azimuth_3','score_1_3','score_2_3']\n```\nwhere direction_% are from Isamu submission, zenith_3 is from crazy quick and clean 1.183 Robert linear solution. I tried many different features directly from event, but\nThe simplicity of Decision Tree Regressor gave us (almost) full control of N bins for adjustment allvors parameters qu and alpha:\n```python\nregr_1 = DecisionTreeRegressor(max_depth=11, min_samples_leaf=500, min_samples_split=500)\ns['bin_raw'] = regr_1.predict(s[cls])\ns['bin'] = round(s['bin_raw'],5).astype(str)\n```\nUsing DecissionTrees gave us 0.0027 boost both with local score and LB. I tried other methods of blending such as MLP or LGBM without success.\nLater on, I also slightly modified `adjusting` loop, for finding full linear combination of zenith_1 and zenith_2:\n```python\ns0['azimuth_pred'] = s0['azimuth_1'] + alpha * s0['diff'] * s0['direction']\ns0['zenith_pred'] = ru * s0['zenith_1'] + qu * s0['zenith_2']\n```\nIt takes much more computation to find ru and qu, but that gives us additional 0.001 improvement.\nOur final blend consist:\nIsamu solution .995 ->  Remek LSTM 1.002 ->  datasaurus Graphnet .982 -> Robert 1.183\nlocal score 0.9774692   Public LB 0.975545, Private LB 0.97658\n## Findings:\n- we found interesting type of error, ‘model give up’ and predicts zenith close to 0.\nOur final solution zenith prediction vs score (x-axis zenith predicted, y-axis zenith ground truth):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F4a8b117472efd756290083dcf0b787c1%2Fimage1.png?generation=1682343676129581&alt=media)\n\nGround truth zenith (x-axis) vs score (y-axis) - ‘error triangle’ visible:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F2ee6d7032191782027e45522daff9081%2Fimage2.png?generation=1682343631946316&alt=media)\n\n- we found that predictions from LSTM classification goes incredibly good at the center of bins, chart: zenith_pred_lstm*1000 (x-axis), mean score (y-axis), zenith edges (red lines)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3974868%2F7fde090be1100c83a09f045fbf57f5fd%2Fimage3.png?generation=1682343415933098&alt=media)\n",
    "2234115": "This topic is very informative and clearly explained. I learned a lot from reading this post. You have done a great job of presenting the concepts in a simple and engaging way. I appreciate your effort and knowledge.\n\nI am going to bookmark this post for future reference. This is a valuable resource that I want to revisit and review. Thank you for sharing this with us. You and your teammates are awesome @yamsam! 😊",
    "2234866": "Congratulations!\nI really wonder how you could find good models such as GPS and GravNet.\nI also searched about models but could not find good models.\n\n@anjum48's notebook on GNN was very helpful. Thank you very much!",
    "2233975": "@yamsam thanks for the detailed topic! Enjoyed reading",
    "2941059": "Congrats to @yamsam and the team for the impressive 8th place finish in the IceCube Neutrino Observatory competition! Your approach to preprocessing, model architectures, and blending is truly commendable. The detailed insights on handling large datasets and optimizing model performance are incredibly valuable. A special shoutout to @anjum48 for the GM promotion and @allvor, @remekkinas, and @wrrosa for their contributions. Looking forward to exploring the code and learning more from your solutions. \nThanks for sharing! ",
    "2234993": "Congratulations on the achievement!",
    "2363820": "Congratulations！Amazing work with informative figure and clear explaination!\nThanks for your sharing! @yamsam "
  }
}