{
  "id": 402882,
  "title": "2nd place solution: Neutrino direction prediction with transformers",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/icemix-2nd-place-solution-neutrino-direction-predi",
  "author_name": "",
  "post_date": "2023-04-27T03:00:44.187Z",
  "votes": 103,
  "comment_count": 34,
  "views": 0,
  "content": "<h1>Summary</h1>\n<ul>\n<li><strong>Transformer-based solution</strong></li>\n<li>Chunk-based data loading with caching and length matching</li>\n<li>Fourier encoder</li>\n<li>Relative spacetime interval bias</li>\n<li>von Mises-Fisher Loss + the competition metric</li>\n</ul>\n<p><strong>Our code is available at <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">this repo</a></strong>.</p>\n<h1>Introduction</h1>\n<p>While neutrinos are one the most abundant particles in the universe that may bring information about violent astrophysical events, their fundamental properties make them difficult to detect and requires a huge volume of material to interact with. The IceCube Neutrino Observatory is the first detector of its kind, which consists of a cubic kilometer of ice and is designed to search for nearly massless neutrinos. The rare interaction of neutrinos with the matter may produce high-energy particles emitting Cherenkov radiation, which is recorded with photo-detectors immersed in ice. The objective of this challenge was a reconstruction of the neutrino trace direction based on the photon detections. <br>\nTo begin with, I would really like to express my gratitude to the organizers and Kaggle team for making this competition possible. Reading the description of the competition I immediately realized that transformers may be a really good approach for it. Since I wanted to experiment with transformers in depth, I decided to join this competition. Meanwhile, my outstanding teammate <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> has joined this competition <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381481#2117294\" target=\"_blank\">hoping to visit Antarctica</a>. </p>\n<h1>Details</h1>\n<p>Our solution is based on transformers, which appeared to be ideal for the considered task. Below we provide the key components of our approach.</p>\n<h2>Why transformers</h2>\n<p>Graph Neural Networks (GNNs) exhibit limitations in certain areas:<br>\n(1)    GNNs function primarily on a local level, as updates are performed based on neighboring nodes. Meanwhile, these models are expected to predict global quantities, such as the direction of a track. Organizers have attempted to address this issue through dynamic edge building, but this method falls short in comparison to the multi-head self-attention mechanism found in transformers. Moreover, dynamic edge building does not provide gradient feedback for neighbor selection.<br>\n(2)    Upon initial examination, GNNs might appear faster than transformers due to their consideration of only a fixed number of neighbors. Nevertheless, extensive sparse operations and data rearrangements in GNNs can be significantly slower than the highly optimized dense tensor operations in transformers for a reasonable sequence length and considering all possible interactions. In our experiments, the organizers’ baseline has the same computational cost as our reference T model (discussed below) while having only 1.4M parameters vs. 7.4M for T model.<br>\n<strong>Transformer is a GNN on steroids</strong>. It may be considered as GNN on a fully connected graph using attention to estimate the edge weights dynamically. But this power comes with the need for huge training data.</p>\n<h2>Data</h2>\n<p>The organizers generously provided a sizable training dataset, which may be somewhat challenging to manage. Since transformers require massive data for training, it was vital to set up a proper data pipeline at the very beginning. Our approach consists of three primary components: (1) caching the selected data chunks to minimize the substantial cost of data loading and preprocessing, (2) employing chunk-based random sampling to effectively utilize caching, and (3) <strong>implementing length-matched data sampling</strong> (sapling batches with approximately equal lengths) to reduce the computational overhead associated with padding tokens when truncating the batch at the longest sequence. Components (1) and (2) offer a computationally efficient and RAM-friendly handling of the data and were shared through the <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching/notebook\" target=\"_blank\">\"Chunk-based Data Loading with Caching\" notebook</a> during the competition to facilitate the setup of a comprehensive training data pipeline for participants. <strong>Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time)</strong>. Inference at a 768 -length led to a 25 bps improvement over the 192-length. Initially, we also attempted to use Hugging Face datasets, which offered the ability to utilize chunked data but required additional file conversion; ultimately, we settled on a chunk caching pipeline.<br>\nIn our study, we observed no significant difference between utilizing only basic features (time, sensor position, charge, auxiliary, and the total number of detections) and incorporating additional features related to ice transparency and absorption. This outcome is expected, as for sufficiently large training datasets, the extra features merely offer an alternate representation of the z-coordinate, with the corresponding transparency and absorption distributions being explicitly learned. To ensure a diverse range of models, we employed both the base and extended feature sets in our experiments. The features are normalized in a similar way as in the organizers' baseline. <strong>It is important to consider detections with both auxiliary true and false.</strong> For events exceeding the maximum considered sequence length, we implemented the following selection process: Initially, we randomly selected detections from the auxiliary false subset, and if this subset proved insufficient to create a sequence of the required length, we then randomly sampled noisy auxiliary true detections. It is interesting that this random selection approach outperformed more complex methods, such as selecting based on charge or considering only events within a specific temporal boundary (IceCube size divided by the speed of light) surrounding the detection with the maximum charge.</p>\n<h2>Model</h2>\n<p>The model is schematically illustrated in the plot below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F552b4994e2931207aa2267c170d6044d%2FModel2.png?generation=1681961703319090&amp;alt=media\" alt=\"\"><br>\n<em>Fourier encoder</em>. Our solution is based on a transformer model, which considers each event as a sequence of detections. It is crucial to process the provided continuous input, such as time or charge, into a form suitable for transformers. We use <a href=\"https://arxiv.org/pdf/1706.03762.pdf\" target=\"_blank\">Fourier encoding representation</a>, often used to describe the position in the sequence in language models. This method can be viewed as a soft digitization of the continuous input signal into a set of codes determined by Fourier frequencies. This procedure is applied for all continuous input variables, while the discrete auxiliary flag is encoded with learnable embedding. We multiply the normalized input variables by 1024-4096 to have a sufficient resolution after digitizing. For example, after multiplying the normalized time by 4096, the temporal resolution becomes 7.3 ns, and the model may understand smaller time variations because of the continuous nature of the Fourier encoding. This multiplication is critical, and even a change of the coefficient from 128 to 4096 gives 20 bps boost.<br>\nThe encoding of all variables is concatenated, passed through one GELU layer, and reprojected into the transformer dimension. An alternative to Fourier encoding may be a multilayer network that learns the corresponding representation of continuous variables, but it may require many more layers than we consider in our setup. <br>\n<em>GraphNet encoder</em>. In addition to the use of Fourier encoding, in several of our models, we incorporated GraphNet feature extractor. In addition to learning the representation of input variables, it also constructs features accounting for the neighbors. <br>\n<em>Relative spacetime interval bias</em>. In the special theory of relativity there is a quantity called spacetime interval: ds^2=c^2 dt^2-dx^2-dy^2-dz^2. For particles moving with a speed close to the speed of light, it is close to zero. What is more useful is that all particles and photons produced in the reactions caused by a single neutrino should have ds^2 close to zero too (it is not fully precise because the refractive index of ice is about 1.3, and the speed of light in it is lower than c).  Therefore, by <strong>computing ds^2 between all pairs of detections in the given event, it is relatively straightforward to distinguish between detections originating from the same neutrino and those that are merely noise</strong>, as demonstrated in the figure below. This criterion can be naturally introduced to a transformer as a <a href=\"https://arxiv.org/pdf/1803.02155v2.pdf\" target=\"_blank\">relative bias</a>. Therefore, during the construction of the attention matrix, the transformer automatically groups detections based on the source event effectively filtering out noise. We use Fourier encoding representation of ds^2 to build the bias term. <strong>Use of relative bias boosts the performance by 40 bps</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fa96019c14603b2c54be402d9d960464f%2Fds.png?generation=1681961957521492&amp;alt=media\" alt=\"\"><br>\n<em>Model</em>. We use the standard <a href=\"https://arxiv.org/pdf/2208.06366.pdf\" target=\"_blank\">BEiT building blocks</a>  (transformer with learnable shortcuts). To incorporate relative bias, the first transformer blocks are modified according to <a href=\"https://arxiv.org/pdf/1803.02155v2.pdf\" target=\"_blank\">this paper</a>. The model size is characterized with a standard ViT notation: T – tiny (dim 192), S – small (dim 384), B – base (dim 768). We use 16 transformer blocks in total: 4 blocks with relative bias + 12 blocks with cls token. The use of an additional cls token, included in the input sequence, enables the gradual collection of information about the track direction throughout the entire model. The attention head size is varied between 32 and 64: the smaller head size gives better performance, especially for the T model (10 bps). The details of the considered models are summarized below.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>dimension</th>\n<th>number of heads</th>\n<th>depth</th>\n<th>number of parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>T</td>\n<td>192</td>\n<td>3-6</td>\n<td>4+12</td>\n<td>7.57M</td>\n</tr>\n<tr>\n<td>S</td>\n<td>384</td>\n<td>6-12</td>\n<td>4+12</td>\n<td>29.3M</td>\n</tr>\n<tr>\n<td>B</td>\n<td>768</td>\n<td>12-24</td>\n<td>4+12</td>\n<td>115.6M</td>\n</tr>\n</tbody>\n</table>\n<p><em>Training setup</em>. The models are trained with AdamW optimizer (weight decay 0.05) for 4-5 epochs (all data is considered). The learning rate is changed with a cosine annealing scheduler with a warmup. The maximum learning rate at the first epoch is 5e-4 for T and S models, and 1e-4 for B models. The following epochs are trained with the maximum learning rate of 0.5e-5 - 2e-5. The effective batch size is 4096 in all experiments (with the use of gradient accumulation). Starting from the second epoch Stochastic Weight Average is used. <br>\nWe perform training for the first 2-3 epochs with von Mises-Fisher Loss (the model predicts 3d vectors, and kappa is computed based on the length of the produced vector), and the remaining epochs are performed with using the competition metric as the objective function + 0.05 von Mises-Fisher Loss (to give the model feedback on the vector length vs. confidence relation and simplify ensembling step). Use of the competition metric as the loss gives 55 bps boost. </p>\n<h1>Results</h1>\n<p>The table below summarizes the performance of our finalized models. CV is evaluated at L=512. (*) refers to models trained on 2×A6000, which in contrast to rtx4090 enables larger actual batch size. Further finetuning with L=256 may result in about 5 bps boost.</p>\n<table>\n<thead>\n<tr>\n<th>Model setup</th>\n<th>CV512</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>T d32</td>\n<td>0.9704</td>\n<td>0.9693</td>\n<td>0.9698</td>\n</tr>\n<tr>\n<td>S d32</td>\n<td>0.9671</td>\n<td>0.9654</td>\n<td>0.9659</td>\n</tr>\n<tr>\n<td>B d32</td>\n<td>0.9642</td>\n<td>0.9623</td>\n<td>0.9632</td>\n</tr>\n<tr>\n<td>B d64</td>\n<td>0.9645</td>\n<td>0.9635</td>\n<td>0.9629</td>\n</tr>\n<tr>\n<td>B+gnn d48</td>\n<td>0.9643</td>\n<td>0.9624</td>\n<td>0.9627</td>\n</tr>\n<tr>\n<td>* S+gnn d32</td>\n<td>0.9639</td>\n<td>0.9620</td>\n<td>0.9628</td>\n</tr>\n<tr>\n<td>* B+gnn d32</td>\n<td>0.9633</td>\n<td>0.9609</td>\n<td>0.9621</td>\n</tr>\n</tbody>\n</table>\n<p>Each increase of the model to the next size T-&gt;S-&gt;B gives approximately 30 bps but is accompanied by an ×4 increase in the number of model parameters. However, the provided data may be insufficient for training models larger than B. Meanwhile, the results of the use GNN in addition to the Fourier extractor are rather contradictive and may be affected by the difference in the hardware configuration.  <br>\nSo, even <strong>the smallest considered T model with only 7.53M parameters and fast training (under 12 hours on rtx4090 per epoch) can take the top-5 in this competition</strong> by itself. Meanwhile, with additional finetuning at L=256, <strong>our best model, B+gnn d32, reaches 0.9628 CV512, 0.9608 public, and 0.9618 private LB as a single model</strong>. </p>\n<h2>Model ensemble</h2>\n<p>With the use of von Mises-Fisher Loss, the vector length corresponds to the model confidence. Therefore, a weighted average of the predicted vectors automatically incorporates the confidence and biases the prediction towards the most confident direction. <br>\nOur best final submission in the linear ensemble of 6 models gives 0.9594 at public and 0.9602  at private LB (the contribution of each model is evaluated as an average of weights fitted in 5 random 5-fold splits). Ensembling gives only an insignificant improvement over our best single model submission of 0.9608 public and 0.9618 at private LB. We also tried to consider more elaborated ensembling by utilizing additional features such as sequence length, first detection time, total charge, etc., but without any improvement over the simple linear ensemble</p>\n<h2>Training and inference time</h2>\n<p><strong>On rtx4090 the inference time can be as short as 3 min per 1M samples for the T model</strong> (without inference optimization) and may increase to 8-10 min for the B model. On outdated GPUs several generations behind, such as P100 at Kaggle, the inference time may increase by x10. The training time varies from 12 hours per epoch for the T model to 56 hours per epoch for the B+gnn model. On a machine with 2×A6000 the speed is approximately the same (~10% faster) than running the model on a single rtx4090.</p>\n<h2>Robustness</h2>\n<p>Regarding the robustness of model predictions with respect to perturbation of the input parameters, we have considered the following cases: (1) <strong>temporal noise</strong>: add normally distributed random variable to detection time of each event; (2) <strong>charge noise</strong>: multiply charge by 10^ normally distributed randomly with variance reported at the horizontal axis; (3) <strong>auxiliary noise</strong>: randomly flip auxiliary flag with a given probability; and (4) <strong>length reduction</strong>: remove a given percentage of detections. The plot below reports results for our reference T model (noise-free CV512 is 0.9704).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ffdbfc788586698d1965af8a024868dd7%2FNoise.png?generation=1681963294066000&amp;alt=media\" alt=\"\"><br>\nAlthough the model was not explicitly trained to consider the abovementioned types of noise, the predictions in most cases do not degrade by more than 15% at the largest magnitude of the considered noise, e.g. adding 200 ns random time error or ×2 reduction of the number of detected events. Training the model with specific data augmentation during training may improve the robustness of the model. However, there may be a limit when the error starts rapidly increasing, e.g. 10 ns temporal noise or a change in the value of detected charge by 40%.<br>\nThe considered model is versatile enough to apply to any neutrino telescope with a similar detector concept as IceCube, regardless of the number of detectors and their arrangement. However, since during the training of our models, the positions of the detectors are fixed, the model may not have a complete understanding of the geometry and experience the degradation of the performance if the detectors are moved. For example, the performance of T model drops from 0.9704 to 1.004 under flip of all detector positions along X or Y axes. To avoid such performance degradation and improve the understanding of geometry by the model, the training could be performed at randomly generated configurations of detectors at each detection event. Incorporation of geometry augmentation, such as shifts, flips, and rotation, would also improve the model robustness to the change of the detector geometry. The less optimal solution is the generation of new training data for a new neutrino telescope. Consideration of <a href=\"https://arxiv.org/pdf/2006.10503.pdf\" target=\"_blank\">3D Roto-Translation Equivariant Attention</a> models, in which detector positions are incorporated through attention rather than direct input, could potentially address the issue of detector geometry. However, this improvement may come with both the computational cost and VRAM requirement increase, and, therefore, a more practical solution may be training a simple model with spatial augmentation and randomized detector positions.<br>\nThe plot below visualizes the distribution of error for confident and unconfident prediction.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd1fe483e3972b5ae6aafa7c4db3f4ea9%2Fdist.png?generation=1681963384438419&amp;alt=media\" alt=\"\"></p>\n<h2>Interesting findings</h2>\n<p>It appeared that <strong>the majority of participants of the competition are unconsciously using the following leak</strong>. The provided data is the result of a simulation, and 0 time has a particular meaning revealing some details of the generation process (the time of the first detection with respect to 0 time may probably give such information as the energy and approximate direction limitations based on the traveling time of the neutrino from the box boundary to the detector). The models learn how to utilize this leak and improve its performance. For example, one of our experiments with the T model trained (1) on data with keeping 0-time-reference and (2) subtracting the time of the first detection from all detections gives the performance gap of 145 bps at the end of the first epoch. <br>\nIn the real detector data, the 0-time-reference is not available, and the time might be considered with the reference to the first detected event only. For models trained using 0-time-reference, such experimental data should be shifted by the average first detection time. However, in this case, the performance may drop by 150+ bps, comparable to the results of the model trained without the 0-time leak. </p>\n<h2>Things that did not work</h2>\n<p>In this competition, we have experimented with a number of things that, unfortunately, did not give an improvement:<br>\n<strong>Graph Neural Networks</strong>. After building the first transformer-based pipeline it becomes apparent that GNNs are well behind (0.98-0.99 vs. 1.0 LB for first models). One interesting comparison we ran at that moment is dividing the predictions based on the kappa by 30% of confident and 70% of unconfident (noisy input and very small number of detections for track reconstruction). The performance of both models on unconfident predictions was comparable, while the improvement on confident predictions was the thing giving the performance gap between transformers and GNNs. It is fully in line with our initial expectation that transformers with global consideration of the input data are more appropriate for the prediction of the global quantity as the track direction in contrast to local GNNs. So further efforts were spent on enhancing the transformer setup. So, in the end, we used GNN as an optional addition to our Fourier feature extractor for diversity.<br>\n<strong>Ice properties</strong>. We did not see that use of additional input, such as ice properties, helps to improve model performance.<br>\n<strong>Local attention</strong>. We spent several weeks experimenting with local attention. The basic idea is simple: before giving a sequence of detections to the transformer model, which considers all-vs-all interactions between detections, process the features within the local neighbors. In theory, local attention is expected to be fast and memory efficient since it examines only a small number of k=8 neighbors instead of requiring L×L attention. However, in practice, local attention appears to be quite slow because of extensive data shuffling (similar to GNN), and in some naïve implementations, it also may consume significantly more VRAM than highly optimized transformers. Surprisingly, the fastest local attention is just regular transformer attention with masking attention for elements beyond k neighbors. Such masking helps at the beginning of training, but at the later stage, full attention outperforms it or gives comparable results. We also tried to consider configurations with alternating global-local attention (sandwich) or parallel blocks of local and global attention in the part of the model preceding the regular 12-layer transformer. However, a simple configuration with consideration of 4 layers with full attention and rel bias + 12-layer regular transformer outperformed the setups with local attention.<br>\n<strong>Classification loss</strong>. We tried to subdivide the angle space into bins and utilize classification loss to fight against angle uncertainty. The ground truth is represented as a 2D Gaussian around the provided label, as illustrated in the figure below. We introduced 8 cls tokens to predict the probability for a 64×128 angular bin map (32×32 per token). Unfortunately, this model ended up at approximately 1.0 CV, and we did not perform any further checks on this approach. However, this model is quite interesting to mention, and it may provide a good visualization of the predictions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F1a2b2b993bb6a79ad13babb00ea689dc%2Fcls.png?generation=1681963515503826&amp;alt=media\" alt=\"\"><br>\n<strong>L and H models</strong> (400 and 800M parameters). Because of  Nvidia virtual address space driver bug, 2×rtx4090 could not be used together until 1-2 weeks before the end of the competition. And even after this bug was fixed, disabled p2p makes multi-GPU training with these cards to be quite inefficient. Therefore, we were unable to scale up our runs and stopped at relatively small B models. The expected boost at each model size increase is about 30 bps, and L and H single models could likely reach 0.958 and 0.955-0.956 LB as a single model but with a massive increase of the required compute for both training and inference.</p>",
  "messages": [
    {
      "id": "2227847",
      "postDate": "04/20/2023 04:09:27",
      "content": "<h1>Summary</h1>\n<ul>\n<li><strong>Transformer-based solution</strong></li>\n<li>Chunk-based data loading with caching and length matching</li>\n<li>Fourier encoder</li>\n<li>Relative spacetime interval bias</li>\n<li>von Mises-Fisher Loss + the competition metric</li>\n</ul>\n<p><strong>Our code is available at <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">this repo</a></strong>.</p>\n<h1>Introduction</h1>\n<p>While neutrinos are one the most abundant particles in the universe that may bring information about violent astrophysical events, their fundamental properties make them difficult to detect and requires a huge volume of material to interact with. The IceCube Neutrino Observatory is the first detector of its kind, which consists of a cubic kilometer of ice and is designed to search for nearly massless neutrinos. The rare interaction of neutrinos with the matter may produce high-energy particles emitting Cherenkov radiation, which is recorded with photo-detectors immersed in ice. The objective of this challenge was a reconstruction of the neutrino trace direction based on the photon detections. <br>\nTo begin with, I would really like to express my gratitude to the organizers and Kaggle team for making this competition possible. Reading the description of the competition I immediately realized that transformers may be a really good approach for it. Since I wanted to experiment with transformers in depth, I decided to join this competition. Meanwhile, my outstanding teammate <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> has joined this competition <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381481#2117294\" target=\"_blank\">hoping to visit Antarctica</a>. </p>\n<h1>Details</h1>\n<p>Our solution is based on transformers, which appeared to be ideal for the considered task. Below we provide the key components of our approach.</p>\n<h2>Why transformers</h2>\n<p>Graph Neural Networks (GNNs) exhibit limitations in certain areas:<br>\n(1)    GNNs function primarily on a local level, as updates are performed based on neighboring nodes. Meanwhile, these models are expected to predict global quantities, such as the direction of a track. Organizers have attempted to address this issue through dynamic edge building, but this method falls short in comparison to the multi-head self-attention mechanism found in transformers. Moreover, dynamic edge building does not provide gradient feedback for neighbor selection.<br>\n(2)    Upon initial examination, GNNs might appear faster than transformers due to their consideration of only a fixed number of neighbors. Nevertheless, extensive sparse operations and data rearrangements in GNNs can be significantly slower than the highly optimized dense tensor operations in transformers for a reasonable sequence length and considering all possible interactions. In our experiments, the organizers’ baseline has the same computational cost as our reference T model (discussed below) while having only 1.4M parameters vs. 7.4M for T model.<br>\n<strong>Transformer is a GNN on steroids</strong>. It may be considered as GNN on a fully connected graph using attention to estimate the edge weights dynamically. But this power comes with the need for huge training data.</p>\n<h2>Data</h2>\n<p>The organizers generously provided a sizable training dataset, which may be somewhat challenging to manage. Since transformers require massive data for training, it was vital to set up a proper data pipeline at the very beginning. Our approach consists of three primary components: (1) caching the selected data chunks to minimize the substantial cost of data loading and preprocessing, (2) employing chunk-based random sampling to effectively utilize caching, and (3) <strong>implementing length-matched data sampling</strong> (sapling batches with approximately equal lengths) to reduce the computational overhead associated with padding tokens when truncating the batch at the longest sequence. Components (1) and (2) offer a computationally efficient and RAM-friendly handling of the data and were shared through the <a href=\"https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching/notebook\" target=\"_blank\">\"Chunk-based Data Loading with Caching\" notebook</a> during the competition to facilitate the setup of a comprehensive training data pipeline for participants. <strong>Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time)</strong>. Inference at a 768 -length led to a 25 bps improvement over the 192-length. Initially, we also attempted to use Hugging Face datasets, which offered the ability to utilize chunked data but required additional file conversion; ultimately, we settled on a chunk caching pipeline.<br>\nIn our study, we observed no significant difference between utilizing only basic features (time, sensor position, charge, auxiliary, and the total number of detections) and incorporating additional features related to ice transparency and absorption. This outcome is expected, as for sufficiently large training datasets, the extra features merely offer an alternate representation of the z-coordinate, with the corresponding transparency and absorption distributions being explicitly learned. To ensure a diverse range of models, we employed both the base and extended feature sets in our experiments. The features are normalized in a similar way as in the organizers' baseline. <strong>It is important to consider detections with both auxiliary true and false.</strong> For events exceeding the maximum considered sequence length, we implemented the following selection process: Initially, we randomly selected detections from the auxiliary false subset, and if this subset proved insufficient to create a sequence of the required length, we then randomly sampled noisy auxiliary true detections. It is interesting that this random selection approach outperformed more complex methods, such as selecting based on charge or considering only events within a specific temporal boundary (IceCube size divided by the speed of light) surrounding the detection with the maximum charge.</p>\n<h2>Model</h2>\n<p>The model is schematically illustrated in the plot below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F552b4994e2931207aa2267c170d6044d%2FModel2.png?generation=1681961703319090&amp;alt=media\" alt=\"\"><br>\n<em>Fourier encoder</em>. Our solution is based on a transformer model, which considers each event as a sequence of detections. It is crucial to process the provided continuous input, such as time or charge, into a form suitable for transformers. We use <a href=\"https://arxiv.org/pdf/1706.03762.pdf\" target=\"_blank\">Fourier encoding representation</a>, often used to describe the position in the sequence in language models. This method can be viewed as a soft digitization of the continuous input signal into a set of codes determined by Fourier frequencies. This procedure is applied for all continuous input variables, while the discrete auxiliary flag is encoded with learnable embedding. We multiply the normalized input variables by 1024-4096 to have a sufficient resolution after digitizing. For example, after multiplying the normalized time by 4096, the temporal resolution becomes 7.3 ns, and the model may understand smaller time variations because of the continuous nature of the Fourier encoding. This multiplication is critical, and even a change of the coefficient from 128 to 4096 gives 20 bps boost.<br>\nThe encoding of all variables is concatenated, passed through one GELU layer, and reprojected into the transformer dimension. An alternative to Fourier encoding may be a multilayer network that learns the corresponding representation of continuous variables, but it may require many more layers than we consider in our setup. <br>\n<em>GraphNet encoder</em>. In addition to the use of Fourier encoding, in several of our models, we incorporated GraphNet feature extractor. In addition to learning the representation of input variables, it also constructs features accounting for the neighbors. <br>\n<em>Relative spacetime interval bias</em>. In the special theory of relativity there is a quantity called spacetime interval: ds^2=c^2 dt^2-dx^2-dy^2-dz^2. For particles moving with a speed close to the speed of light, it is close to zero. What is more useful is that all particles and photons produced in the reactions caused by a single neutrino should have ds^2 close to zero too (it is not fully precise because the refractive index of ice is about 1.3, and the speed of light in it is lower than c).  Therefore, by <strong>computing ds^2 between all pairs of detections in the given event, it is relatively straightforward to distinguish between detections originating from the same neutrino and those that are merely noise</strong>, as demonstrated in the figure below. This criterion can be naturally introduced to a transformer as a <a href=\"https://arxiv.org/pdf/1803.02155v2.pdf\" target=\"_blank\">relative bias</a>. Therefore, during the construction of the attention matrix, the transformer automatically groups detections based on the source event effectively filtering out noise. We use Fourier encoding representation of ds^2 to build the bias term. <strong>Use of relative bias boosts the performance by 40 bps</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fa96019c14603b2c54be402d9d960464f%2Fds.png?generation=1681961957521492&amp;alt=media\" alt=\"\"><br>\n<em>Model</em>. We use the standard <a href=\"https://arxiv.org/pdf/2208.06366.pdf\" target=\"_blank\">BEiT building blocks</a>  (transformer with learnable shortcuts). To incorporate relative bias, the first transformer blocks are modified according to <a href=\"https://arxiv.org/pdf/1803.02155v2.pdf\" target=\"_blank\">this paper</a>. The model size is characterized with a standard ViT notation: T – tiny (dim 192), S – small (dim 384), B – base (dim 768). We use 16 transformer blocks in total: 4 blocks with relative bias + 12 blocks with cls token. The use of an additional cls token, included in the input sequence, enables the gradual collection of information about the track direction throughout the entire model. The attention head size is varied between 32 and 64: the smaller head size gives better performance, especially for the T model (10 bps). The details of the considered models are summarized below.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>dimension</th>\n<th>number of heads</th>\n<th>depth</th>\n<th>number of parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>T</td>\n<td>192</td>\n<td>3-6</td>\n<td>4+12</td>\n<td>7.57M</td>\n</tr>\n<tr>\n<td>S</td>\n<td>384</td>\n<td>6-12</td>\n<td>4+12</td>\n<td>29.3M</td>\n</tr>\n<tr>\n<td>B</td>\n<td>768</td>\n<td>12-24</td>\n<td>4+12</td>\n<td>115.6M</td>\n</tr>\n</tbody>\n</table>\n<p><em>Training setup</em>. The models are trained with AdamW optimizer (weight decay 0.05) for 4-5 epochs (all data is considered). The learning rate is changed with a cosine annealing scheduler with a warmup. The maximum learning rate at the first epoch is 5e-4 for T and S models, and 1e-4 for B models. The following epochs are trained with the maximum learning rate of 0.5e-5 - 2e-5. The effective batch size is 4096 in all experiments (with the use of gradient accumulation). Starting from the second epoch Stochastic Weight Average is used. <br>\nWe perform training for the first 2-3 epochs with von Mises-Fisher Loss (the model predicts 3d vectors, and kappa is computed based on the length of the produced vector), and the remaining epochs are performed with using the competition metric as the objective function + 0.05 von Mises-Fisher Loss (to give the model feedback on the vector length vs. confidence relation and simplify ensembling step). Use of the competition metric as the loss gives 55 bps boost. </p>\n<h1>Results</h1>\n<p>The table below summarizes the performance of our finalized models. CV is evaluated at L=512. (*) refers to models trained on 2×A6000, which in contrast to rtx4090 enables larger actual batch size. Further finetuning with L=256 may result in about 5 bps boost.</p>\n<table>\n<thead>\n<tr>\n<th>Model setup</th>\n<th>CV512</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>T d32</td>\n<td>0.9704</td>\n<td>0.9693</td>\n<td>0.9698</td>\n</tr>\n<tr>\n<td>S d32</td>\n<td>0.9671</td>\n<td>0.9654</td>\n<td>0.9659</td>\n</tr>\n<tr>\n<td>B d32</td>\n<td>0.9642</td>\n<td>0.9623</td>\n<td>0.9632</td>\n</tr>\n<tr>\n<td>B d64</td>\n<td>0.9645</td>\n<td>0.9635</td>\n<td>0.9629</td>\n</tr>\n<tr>\n<td>B+gnn d48</td>\n<td>0.9643</td>\n<td>0.9624</td>\n<td>0.9627</td>\n</tr>\n<tr>\n<td>* S+gnn d32</td>\n<td>0.9639</td>\n<td>0.9620</td>\n<td>0.9628</td>\n</tr>\n<tr>\n<td>* B+gnn d32</td>\n<td>0.9633</td>\n<td>0.9609</td>\n<td>0.9621</td>\n</tr>\n</tbody>\n</table>\n<p>Each increase of the model to the next size T-&gt;S-&gt;B gives approximately 30 bps but is accompanied by an ×4 increase in the number of model parameters. However, the provided data may be insufficient for training models larger than B. Meanwhile, the results of the use GNN in addition to the Fourier extractor are rather contradictive and may be affected by the difference in the hardware configuration.  <br>\nSo, even <strong>the smallest considered T model with only 7.53M parameters and fast training (under 12 hours on rtx4090 per epoch) can take the top-5 in this competition</strong> by itself. Meanwhile, with additional finetuning at L=256, <strong>our best model, B+gnn d32, reaches 0.9628 CV512, 0.9608 public, and 0.9618 private LB as a single model</strong>. </p>\n<h2>Model ensemble</h2>\n<p>With the use of von Mises-Fisher Loss, the vector length corresponds to the model confidence. Therefore, a weighted average of the predicted vectors automatically incorporates the confidence and biases the prediction towards the most confident direction. <br>\nOur best final submission in the linear ensemble of 6 models gives 0.9594 at public and 0.9602  at private LB (the contribution of each model is evaluated as an average of weights fitted in 5 random 5-fold splits). Ensembling gives only an insignificant improvement over our best single model submission of 0.9608 public and 0.9618 at private LB. We also tried to consider more elaborated ensembling by utilizing additional features such as sequence length, first detection time, total charge, etc., but without any improvement over the simple linear ensemble</p>\n<h2>Training and inference time</h2>\n<p><strong>On rtx4090 the inference time can be as short as 3 min per 1M samples for the T model</strong> (without inference optimization) and may increase to 8-10 min for the B model. On outdated GPUs several generations behind, such as P100 at Kaggle, the inference time may increase by x10. The training time varies from 12 hours per epoch for the T model to 56 hours per epoch for the B+gnn model. On a machine with 2×A6000 the speed is approximately the same (~10% faster) than running the model on a single rtx4090.</p>\n<h2>Robustness</h2>\n<p>Regarding the robustness of model predictions with respect to perturbation of the input parameters, we have considered the following cases: (1) <strong>temporal noise</strong>: add normally distributed random variable to detection time of each event; (2) <strong>charge noise</strong>: multiply charge by 10^ normally distributed randomly with variance reported at the horizontal axis; (3) <strong>auxiliary noise</strong>: randomly flip auxiliary flag with a given probability; and (4) <strong>length reduction</strong>: remove a given percentage of detections. The plot below reports results for our reference T model (noise-free CV512 is 0.9704).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ffdbfc788586698d1965af8a024868dd7%2FNoise.png?generation=1681963294066000&amp;alt=media\" alt=\"\"><br>\nAlthough the model was not explicitly trained to consider the abovementioned types of noise, the predictions in most cases do not degrade by more than 15% at the largest magnitude of the considered noise, e.g. adding 200 ns random time error or ×2 reduction of the number of detected events. Training the model with specific data augmentation during training may improve the robustness of the model. However, there may be a limit when the error starts rapidly increasing, e.g. 10 ns temporal noise or a change in the value of detected charge by 40%.<br>\nThe considered model is versatile enough to apply to any neutrino telescope with a similar detector concept as IceCube, regardless of the number of detectors and their arrangement. However, since during the training of our models, the positions of the detectors are fixed, the model may not have a complete understanding of the geometry and experience the degradation of the performance if the detectors are moved. For example, the performance of T model drops from 0.9704 to 1.004 under flip of all detector positions along X or Y axes. To avoid such performance degradation and improve the understanding of geometry by the model, the training could be performed at randomly generated configurations of detectors at each detection event. Incorporation of geometry augmentation, such as shifts, flips, and rotation, would also improve the model robustness to the change of the detector geometry. The less optimal solution is the generation of new training data for a new neutrino telescope. Consideration of <a href=\"https://arxiv.org/pdf/2006.10503.pdf\" target=\"_blank\">3D Roto-Translation Equivariant Attention</a> models, in which detector positions are incorporated through attention rather than direct input, could potentially address the issue of detector geometry. However, this improvement may come with both the computational cost and VRAM requirement increase, and, therefore, a more practical solution may be training a simple model with spatial augmentation and randomized detector positions.<br>\nThe plot below visualizes the distribution of error for confident and unconfident prediction.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd1fe483e3972b5ae6aafa7c4db3f4ea9%2Fdist.png?generation=1681963384438419&amp;alt=media\" alt=\"\"></p>\n<h2>Interesting findings</h2>\n<p>It appeared that <strong>the majority of participants of the competition are unconsciously using the following leak</strong>. The provided data is the result of a simulation, and 0 time has a particular meaning revealing some details of the generation process (the time of the first detection with respect to 0 time may probably give such information as the energy and approximate direction limitations based on the traveling time of the neutrino from the box boundary to the detector). The models learn how to utilize this leak and improve its performance. For example, one of our experiments with the T model trained (1) on data with keeping 0-time-reference and (2) subtracting the time of the first detection from all detections gives the performance gap of 145 bps at the end of the first epoch. <br>\nIn the real detector data, the 0-time-reference is not available, and the time might be considered with the reference to the first detected event only. For models trained using 0-time-reference, such experimental data should be shifted by the average first detection time. However, in this case, the performance may drop by 150+ bps, comparable to the results of the model trained without the 0-time leak. </p>\n<h2>Things that did not work</h2>\n<p>In this competition, we have experimented with a number of things that, unfortunately, did not give an improvement:<br>\n<strong>Graph Neural Networks</strong>. After building the first transformer-based pipeline it becomes apparent that GNNs are well behind (0.98-0.99 vs. 1.0 LB for first models). One interesting comparison we ran at that moment is dividing the predictions based on the kappa by 30% of confident and 70% of unconfident (noisy input and very small number of detections for track reconstruction). The performance of both models on unconfident predictions was comparable, while the improvement on confident predictions was the thing giving the performance gap between transformers and GNNs. It is fully in line with our initial expectation that transformers with global consideration of the input data are more appropriate for the prediction of the global quantity as the track direction in contrast to local GNNs. So further efforts were spent on enhancing the transformer setup. So, in the end, we used GNN as an optional addition to our Fourier feature extractor for diversity.<br>\n<strong>Ice properties</strong>. We did not see that use of additional input, such as ice properties, helps to improve model performance.<br>\n<strong>Local attention</strong>. We spent several weeks experimenting with local attention. The basic idea is simple: before giving a sequence of detections to the transformer model, which considers all-vs-all interactions between detections, process the features within the local neighbors. In theory, local attention is expected to be fast and memory efficient since it examines only a small number of k=8 neighbors instead of requiring L×L attention. However, in practice, local attention appears to be quite slow because of extensive data shuffling (similar to GNN), and in some naïve implementations, it also may consume significantly more VRAM than highly optimized transformers. Surprisingly, the fastest local attention is just regular transformer attention with masking attention for elements beyond k neighbors. Such masking helps at the beginning of training, but at the later stage, full attention outperforms it or gives comparable results. We also tried to consider configurations with alternating global-local attention (sandwich) or parallel blocks of local and global attention in the part of the model preceding the regular 12-layer transformer. However, a simple configuration with consideration of 4 layers with full attention and rel bias + 12-layer regular transformer outperformed the setups with local attention.<br>\n<strong>Classification loss</strong>. We tried to subdivide the angle space into bins and utilize classification loss to fight against angle uncertainty. The ground truth is represented as a 2D Gaussian around the provided label, as illustrated in the figure below. We introduced 8 cls tokens to predict the probability for a 64×128 angular bin map (32×32 per token). Unfortunately, this model ended up at approximately 1.0 CV, and we did not perform any further checks on this approach. However, this model is quite interesting to mention, and it may provide a good visualization of the predictions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F1a2b2b993bb6a79ad13babb00ea689dc%2Fcls.png?generation=1681963515503826&amp;alt=media\" alt=\"\"><br>\n<strong>L and H models</strong> (400 and 800M parameters). Because of  Nvidia virtual address space driver bug, 2×rtx4090 could not be used together until 1-2 weeks before the end of the competition. And even after this bug was fixed, disabled p2p makes multi-GPU training with these cards to be quite inefficient. Therefore, we were unable to scale up our runs and stopped at relatively small B models. The expected boost at each model size increase is about 30 bps, and L and H single models could likely reach 0.958 and 0.955-0.956 LB as a single model but with a massive increase of the required compute for both training and inference.</p>",
      "rawMarkdown": "# Summary\n-\t**Transformer-based solution**\n-\tChunk-based data loading with caching and length matching\n-\tFourier encoder\n-\tRelative spacetime interval bias\n-\tvon Mises-Fisher Loss + the competition metric\n\n**Our code is available at [this repo](https://github.com/DrHB/icecube-2nd-place)**.\n\n# Introduction\nWhile neutrinos are one the most abundant particles in the universe that may bring information about violent astrophysical events, their fundamental properties make them difficult to detect and requires a huge volume of material to interact with. The IceCube Neutrino Observatory is the first detector of its kind, which consists of a cubic kilometer of ice and is designed to search for nearly massless neutrinos. The rare interaction of neutrinos with the matter may produce high-energy particles emitting Cherenkov radiation, which is recorded with photo-detectors immersed in ice. The objective of this challenge was a reconstruction of the neutrino trace direction based on the photon detections. \nTo begin with, I would really like to express my gratitude to the organizers and Kaggle team for making this competition possible. Reading the description of the competition I immediately realized that transformers may be a really good approach for it. Since I wanted to experiment with transformers in depth, I decided to join this competition. Meanwhile, my outstanding teammate @drhabib has joined this competition [hoping to visit Antarctica](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381481#2117294). \n\n# Details\nOur solution is based on transformers, which appeared to be ideal for the considered task. Below we provide the key components of our approach.\n## Why transformers\nGraph Neural Networks (GNNs) exhibit limitations in certain areas:\n(1)\tGNNs function primarily on a local level, as updates are performed based on neighboring nodes. Meanwhile, these models are expected to predict global quantities, such as the direction of a track. Organizers have attempted to address this issue through dynamic edge building, but this method falls short in comparison to the multi-head self-attention mechanism found in transformers. Moreover, dynamic edge building does not provide gradient feedback for neighbor selection.\n(2)\tUpon initial examination, GNNs might appear faster than transformers due to their consideration of only a fixed number of neighbors. Nevertheless, extensive sparse operations and data rearrangements in GNNs can be significantly slower than the highly optimized dense tensor operations in transformers for a reasonable sequence length and considering all possible interactions. In our experiments, the organizers’ baseline has the same computational cost as our reference T model (discussed below) while having only 1.4M parameters vs. 7.4M for T model.\n**Transformer is a GNN on steroids**. It may be considered as GNN on a fully connected graph using attention to estimate the edge weights dynamically. But this power comes with the need for huge training data.\n## Data\nThe organizers generously provided a sizable training dataset, which may be somewhat challenging to manage. Since transformers require massive data for training, it was vital to set up a proper data pipeline at the very beginning. Our approach consists of three primary components: (1) caching the selected data chunks to minimize the substantial cost of data loading and preprocessing, (2) employing chunk-based random sampling to effectively utilize caching, and (3) **implementing length-matched data sampling** (sapling batches with approximately equal lengths) to reduce the computational overhead associated with padding tokens when truncating the batch at the longest sequence. Components (1) and (2) offer a computationally efficient and RAM-friendly handling of the data and were shared through the [\"Chunk-based Data Loading with Caching\" notebook](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching/notebook) during the competition to facilitate the setup of a comprehensive training data pipeline for participants. **Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time)**. Inference at a 768 -length led to a 25 bps improvement over the 192-length. Initially, we also attempted to use Hugging Face datasets, which offered the ability to utilize chunked data but required additional file conversion; ultimately, we settled on a chunk caching pipeline.\nIn our study, we observed no significant difference between utilizing only basic features (time, sensor position, charge, auxiliary, and the total number of detections) and incorporating additional features related to ice transparency and absorption. This outcome is expected, as for sufficiently large training datasets, the extra features merely offer an alternate representation of the z-coordinate, with the corresponding transparency and absorption distributions being explicitly learned. To ensure a diverse range of models, we employed both the base and extended feature sets in our experiments. The features are normalized in a similar way as in the organizers' baseline. **It is important to consider detections with both auxiliary true and false.** For events exceeding the maximum considered sequence length, we implemented the following selection process: Initially, we randomly selected detections from the auxiliary false subset, and if this subset proved insufficient to create a sequence of the required length, we then randomly sampled noisy auxiliary true detections. It is interesting that this random selection approach outperformed more complex methods, such as selecting based on charge or considering only events within a specific temporal boundary (IceCube size divided by the speed of light) surrounding the detection with the maximum charge.\n## Model\nThe model is schematically illustrated in the plot below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F552b4994e2931207aa2267c170d6044d%2FModel2.png?generation=1681961703319090&alt=media)\n*Fourier encoder*. Our solution is based on a transformer model, which considers each event as a sequence of detections. It is crucial to process the provided continuous input, such as time or charge, into a form suitable for transformers. We use [Fourier encoding representation](https://arxiv.org/pdf/1706.03762.pdf), often used to describe the position in the sequence in language models. This method can be viewed as a soft digitization of the continuous input signal into a set of codes determined by Fourier frequencies. This procedure is applied for all continuous input variables, while the discrete auxiliary flag is encoded with learnable embedding. We multiply the normalized input variables by 1024-4096 to have a sufficient resolution after digitizing. For example, after multiplying the normalized time by 4096, the temporal resolution becomes 7.3 ns, and the model may understand smaller time variations because of the continuous nature of the Fourier encoding. This multiplication is critical, and even a change of the coefficient from 128 to 4096 gives 20 bps boost.\nThe encoding of all variables is concatenated, passed through one GELU layer, and reprojected into the transformer dimension. An alternative to Fourier encoding may be a multilayer network that learns the corresponding representation of continuous variables, but it may require many more layers than we consider in our setup. \n*GraphNet encoder*. In addition to the use of Fourier encoding, in several of our models, we incorporated GraphNet feature extractor. In addition to learning the representation of input variables, it also constructs features accounting for the neighbors. \n*Relative spacetime interval bias*. In the special theory of relativity there is a quantity called spacetime interval: ds^2=c^2 dt^2-dx^2-dy^2-dz^2. For particles moving with a speed close to the speed of light, it is close to zero. What is more useful is that all particles and photons produced in the reactions caused by a single neutrino should have ds^2 close to zero too (it is not fully precise because the refractive index of ice is about 1.3, and the speed of light in it is lower than c).  Therefore, by **computing ds^2 between all pairs of detections in the given event, it is relatively straightforward to distinguish between detections originating from the same neutrino and those that are merely noise**, as demonstrated in the figure below. This criterion can be naturally introduced to a transformer as a [relative bias](https://arxiv.org/pdf/1803.02155v2.pdf). Therefore, during the construction of the attention matrix, the transformer automatically groups detections based on the source event effectively filtering out noise. We use Fourier encoding representation of ds^2 to build the bias term. **Use of relative bias boosts the performance by 40 bps**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fa96019c14603b2c54be402d9d960464f%2Fds.png?generation=1681961957521492&alt=media)\n*Model*. We use the standard [BEiT building blocks](https://arxiv.org/pdf/2208.06366.pdf)  (transformer with learnable shortcuts). To incorporate relative bias, the first transformer blocks are modified according to [this paper](https://arxiv.org/pdf/1803.02155v2.pdf). The model size is characterized with a standard ViT notation: T – tiny (dim 192), S – small (dim 384), B – base (dim 768). We use 16 transformer blocks in total: 4 blocks with relative bias + 12 blocks with cls token. The use of an additional cls token, included in the input sequence, enables the gradual collection of information about the track direction throughout the entire model. The attention head size is varied between 32 and 64: the smaller head size gives better performance, especially for the T model (10 bps). The details of the considered models are summarized below.\n\n| Model | dimension | number of heads | depth | number of parameters |\n| --- | --- | --- | --- | --- |\n| T | 192 | 3-6 | 4+12 | 7.57M |\n| S | 384 | 6-12 | 4+12 | 29.3M |\n| B | 768\t| 12-24 | 4+12 | 115.6M |\n\n\n*Training setup*. The models are trained with AdamW optimizer (weight decay 0.05) for 4-5 epochs (all data is considered). The learning rate is changed with a cosine annealing scheduler with a warmup. The maximum learning rate at the first epoch is 5e-4 for T and S models, and 1e-4 for B models. The following epochs are trained with the maximum learning rate of 0.5e-5 - 2e-5. The effective batch size is 4096 in all experiments (with the use of gradient accumulation). Starting from the second epoch Stochastic Weight Average is used. \nWe perform training for the first 2-3 epochs with von Mises-Fisher Loss (the model predicts 3d vectors, and kappa is computed based on the length of the produced vector), and the remaining epochs are performed with using the competition metric as the objective function + 0.05 von Mises-Fisher Loss (to give the model feedback on the vector length vs. confidence relation and simplify ensembling step). Use of the competition metric as the loss gives 55 bps boost. \n\n# Results\nThe table below summarizes the performance of our finalized models. CV is evaluated at L=512. (*) refers to models trained on 2×A6000, which in contrast to rtx4090 enables larger actual batch size. Further finetuning with L=256 may result in about 5 bps boost.\n\t\n| Model setup | CV512 | Public LB | Private LB| \n| --- | --- | --- | --- |\n| T d32 | 0.9704 | 0.9693 | 0.9698 |\n| S d32 | 0.9671 | 0.9654 | 0.9659 | \n| B d32 | 0.9642 | 0.9623 | 0.9632 |\n| B d64 | 0.9645 | 0.9635 | 0.9629 |\n| B+gnn d48 | 0.9643 | 0.9624 | 0.9627 |\n| * S+gnn d32 | 0.9639 | 0.9620 | 0.9628 |\n| * B+gnn d32 | 0.9633 | 0.9609 | 0.9621 |\n\nEach increase of the model to the next size T->S->B gives approximately 30 bps but is accompanied by an ×4 increase in the number of model parameters. However, the provided data may be insufficient for training models larger than B. Meanwhile, the results of the use GNN in addition to the Fourier extractor are rather contradictive and may be affected by the difference in the hardware configuration.  \nSo, even **the smallest considered T model with only 7.53M parameters and fast training (under 12 hours on rtx4090 per epoch) can take the top-5 in this competition** by itself. Meanwhile, with additional finetuning at L=256, **our best model, B+gnn d32, reaches 0.9628 CV512, 0.9608 public, and 0.9618 private LB as a single model**. \n\n## Model ensemble\nWith the use of von Mises-Fisher Loss, the vector length corresponds to the model confidence. Therefore, a weighted average of the predicted vectors automatically incorporates the confidence and biases the prediction towards the most confident direction. \nOur best final submission in the linear ensemble of 6 models gives 0.9594 at public and 0.9602  at private LB (the contribution of each model is evaluated as an average of weights fitted in 5 random 5-fold splits). Ensembling gives only an insignificant improvement over our best single model submission of 0.9608 public and 0.9618 at private LB. We also tried to consider more elaborated ensembling by utilizing additional features such as sequence length, first detection time, total charge, etc., but without any improvement over the simple linear ensemble\n\n## Training and inference time\n**On rtx4090 the inference time can be as short as 3 min per 1M samples for the T model** (without inference optimization) and may increase to 8-10 min for the B model. On outdated GPUs several generations behind, such as P100 at Kaggle, the inference time may increase by x10. The training time varies from 12 hours per epoch for the T model to 56 hours per epoch for the B+gnn model. On a machine with 2×A6000 the speed is approximately the same (~10% faster) than running the model on a single rtx4090.\n\n## Robustness\nRegarding the robustness of model predictions with respect to perturbation of the input parameters, we have considered the following cases: (1) **temporal noise**: add normally distributed random variable to detection time of each event; (2) **charge noise**: multiply charge by 10^ normally distributed randomly with variance reported at the horizontal axis; (3) **auxiliary noise**: randomly flip auxiliary flag with a given probability; and (4) **length reduction**: remove a given percentage of detections. The plot below reports results for our reference T model (noise-free CV512 is 0.9704).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ffdbfc788586698d1965af8a024868dd7%2FNoise.png?generation=1681963294066000&alt=media)\nAlthough the model was not explicitly trained to consider the abovementioned types of noise, the predictions in most cases do not degrade by more than 15% at the largest magnitude of the considered noise, e.g. adding 200 ns random time error or ×2 reduction of the number of detected events. Training the model with specific data augmentation during training may improve the robustness of the model. However, there may be a limit when the error starts rapidly increasing, e.g. 10 ns temporal noise or a change in the value of detected charge by 40%.\nThe considered model is versatile enough to apply to any neutrino telescope with a similar detector concept as IceCube, regardless of the number of detectors and their arrangement. However, since during the training of our models, the positions of the detectors are fixed, the model may not have a complete understanding of the geometry and experience the degradation of the performance if the detectors are moved. For example, the performance of T model drops from 0.9704 to 1.004 under flip of all detector positions along X or Y axes. To avoid such performance degradation and improve the understanding of geometry by the model, the training could be performed at randomly generated configurations of detectors at each detection event. Incorporation of geometry augmentation, such as shifts, flips, and rotation, would also improve the model robustness to the change of the detector geometry. The less optimal solution is the generation of new training data for a new neutrino telescope. Consideration of [3D Roto-Translation Equivariant Attention](https://arxiv.org/pdf/2006.10503.pdf) models, in which detector positions are incorporated through attention rather than direct input, could potentially address the issue of detector geometry. However, this improvement may come with both the computational cost and VRAM requirement increase, and, therefore, a more practical solution may be training a simple model with spatial augmentation and randomized detector positions.\nThe plot below visualizes the distribution of error for confident and unconfident prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd1fe483e3972b5ae6aafa7c4db3f4ea9%2Fdist.png?generation=1681963384438419&alt=media)\n## Interesting findings\nIt appeared that **the majority of participants of the competition are unconsciously using the following leak**. The provided data is the result of a simulation, and 0 time has a particular meaning revealing some details of the generation process (the time of the first detection with respect to 0 time may probably give such information as the energy and approximate direction limitations based on the traveling time of the neutrino from the box boundary to the detector). The models learn how to utilize this leak and improve its performance. For example, one of our experiments with the T model trained (1) on data with keeping 0-time-reference and (2) subtracting the time of the first detection from all detections gives the performance gap of 145 bps at the end of the first epoch. \nIn the real detector data, the 0-time-reference is not available, and the time might be considered with the reference to the first detected event only. For models trained using 0-time-reference, such experimental data should be shifted by the average first detection time. However, in this case, the performance may drop by 150+ bps, comparable to the results of the model trained without the 0-time leak. \n\n## Things that did not work\nIn this competition, we have experimented with a number of things that, unfortunately, did not give an improvement:\n**Graph Neural Networks**. After building the first transformer-based pipeline it becomes apparent that GNNs are well behind (0.98-0.99 vs. 1.0 LB for first models). One interesting comparison we ran at that moment is dividing the predictions based on the kappa by 30% of confident and 70% of unconfident (noisy input and very small number of detections for track reconstruction). The performance of both models on unconfident predictions was comparable, while the improvement on confident predictions was the thing giving the performance gap between transformers and GNNs. It is fully in line with our initial expectation that transformers with global consideration of the input data are more appropriate for the prediction of the global quantity as the track direction in contrast to local GNNs. So further efforts were spent on enhancing the transformer setup. So, in the end, we used GNN as an optional addition to our Fourier feature extractor for diversity.\n**Ice properties**. We did not see that use of additional input, such as ice properties, helps to improve model performance.\n**Local attention**. We spent several weeks experimenting with local attention. The basic idea is simple: before giving a sequence of detections to the transformer model, which considers all-vs-all interactions between detections, process the features within the local neighbors. In theory, local attention is expected to be fast and memory efficient since it examines only a small number of k=8 neighbors instead of requiring L×L attention. However, in practice, local attention appears to be quite slow because of extensive data shuffling (similar to GNN), and in some naïve implementations, it also may consume significantly more VRAM than highly optimized transformers. Surprisingly, the fastest local attention is just regular transformer attention with masking attention for elements beyond k neighbors. Such masking helps at the beginning of training, but at the later stage, full attention outperforms it or gives comparable results. We also tried to consider configurations with alternating global-local attention (sandwich) or parallel blocks of local and global attention in the part of the model preceding the regular 12-layer transformer. However, a simple configuration with consideration of 4 layers with full attention and rel bias + 12-layer regular transformer outperformed the setups with local attention.\n**Classification loss**. We tried to subdivide the angle space into bins and utilize classification loss to fight against angle uncertainty. The ground truth is represented as a 2D Gaussian around the provided label, as illustrated in the figure below. We introduced 8 cls tokens to predict the probability for a 64×128 angular bin map (32×32 per token). Unfortunately, this model ended up at approximately 1.0 CV, and we did not perform any further checks on this approach. However, this model is quite interesting to mention, and it may provide a good visualization of the predictions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F1a2b2b993bb6a79ad13babb00ea689dc%2Fcls.png?generation=1681963515503826&alt=media)\n**L and H models** (400 and 800M parameters). Because of  Nvidia virtual address space driver bug, 2×rtx4090 could not be used together until 1-2 weeks before the end of the competition. And even after this bug was fixed, disabled p2p makes multi-GPU training with these cards to be quite inefficient. Therefore, we were unable to scale up our runs and stopped at relatively small B models. The expected boost at each model size increase is about 30 bps, and L and H single models could likely reach 0.958 and 0.955-0.956 LB as a single model but with a massive increase of the required compute for both training and inference.",
      "votes": null
    },
    {
      "id": "2227866",
      "postDate": "04/20/2023 04:43:36",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks. Very detailed and interesting solution! Congrats with 2 place!</p>",
      "rawMarkdown": "iafoss thanks. Very detailed and interesting solution! Congrats with 2 place!",
      "votes": null
    },
    {
      "id": "2227877",
      "postDate": "04/20/2023 04:51:47",
      "content": "<p>Wow, this is truly impressive! Congratulations, and thank you for providing such a comprehensive solution!</p>\n<p>We had nearly all the necessary components (Fourier encoder, Relative spacetime interval bias, and a speedy dataloader built on a memmap file), but were training our model on a GTX 1080. Consequently, we never observed any significant improvements compared to a GRU model 😂</p>\n<p>Additionally, the awareness of data leakage is crucial for physics applications. It's fantastic that you have drawn attention to this issue.</p>",
      "rawMarkdown": "Wow, this is truly impressive! Congratulations, and thank you for providing such a comprehensive solution!\n\nWe had nearly all the necessary components (Fourier encoder, Relative spacetime interval bias, and a speedy dataloader built on a memmap file), but were training our model on a GTX 1080. Consequently, we never observed any significant improvements compared to a GRU model 😂\n\nAdditionally, the awareness of data leakage is crucial for physics applications. It's fantastic that you have drawn attention to this issue.",
      "votes": null
    },
    {
      "id": "2227882",
      "postDate": "04/20/2023 04:59:38",
      "content": "<p>Thanks     </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "2227914",
      "postDate": "04/20/2023 05:30:30",
      "content": "<p>Wow I never thought of that 0 time leak, always trained with the first detection time subtracted. Wonder how much difference that would have made.</p>\n<p>The classifier you mention is similar to mine, instead of joint map I have 2 heads and it did work, it was rather much more stable for me.</p>\n<p>Your use of fourier encoder and relative bias is amazing, learnt a lot from reading this, thank you for sharing.</p>",
      "rawMarkdown": "Wow I never thought of that 0 time leak, always trained with the first detection time subtracted. Wonder how much difference that would have made.\n\nThe classifier you mention is similar to mine, instead of joint map I have 2 heads and it did work, it was rather much more stable for me.\n\nYour use of fourier encoder and relative bias is amazing, learnt a lot from reading this, thank you for sharing.",
      "votes": null
    },
    {
      "id": "2228063",
      "postDate": "04/20/2023 08:38:22",
      "content": "<p>I am New to Neural networks hope you guys will help me out 😁 </p>",
      "rawMarkdown": "I am New to Neural networks hope you guys will help me out 😁",
      "votes": null
    },
    {
      "id": "2228240",
      "postDate": "04/20/2023 11:30:56",
      "content": "<p>Congratulations on the second place. This write-up is quite interesting and the architecture is impressive. Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations on the second place. This write-up is quite interesting and the architecture is impressive. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2228289",
      "postDate": "04/20/2023 12:28:11",
      "content": "<p>Congratulations on your finish and the very thorough writeup.  What tool did you use to draw your model diagram?</p>",
      "rawMarkdown": "Congratulations on your finish and the very thorough writeup.  What tool did you use to draw your model diagram?",
      "votes": null
    },
    {
      "id": "2228301",
      "postDate": "04/20/2023 12:43:02",
      "content": "<p>Sorry for a nitpick, but your link in \"<em>We use Fourier encoding representation, often used</em>\" seems to be the Attention paper.  Is that where Fourier Encoding is discussed?</p>",
      "rawMarkdown": "Sorry for a nitpick, but your link in \"*We use Fourier encoding representation, often used*\" seems to be the Attention paper.  Is that where Fourier Encoding is discussed?",
      "votes": null
    },
    {
      "id": "2228353",
      "postDate": "04/20/2023 13:40:44",
      "content": "<p>Yes, it is the place where it was introduced for the first time to my knowledge for the description of the position of tokens in the sequence, but I may be wrong.</p>",
      "rawMarkdown": "Yes, it is the place where it was introduced for the first time to my knowledge for the description of the position of tokens in the sequence, but I may be wrong.",
      "votes": null
    },
    {
      "id": "2228370",
      "postDate": "04/20/2023 13:50:29",
      "content": "<p>Thank you. It was combination of Figma and a little bit of midjourneyv5:) </p>",
      "rawMarkdown": "Thank you. It was combination of Figma and a little bit of midjourneyv5:)",
      "votes": null
    },
    {
      "id": "2228383",
      "postDate": "04/20/2023 14:01:01",
      "content": "<p>Congratulations Dream Team! I am reading your solution description and it is amazing. I will spend a lot of time learning from this description and from you. 🙏🙏🙏</p>",
      "rawMarkdown": "Congratulations Dream Team! I am reading your solution description and it is amazing. I will spend a lot of time learning from this description and from you. 🙏🙏🙏",
      "votes": null
    },
    {
      "id": "2228460",
      "postDate": "04/20/2023 15:07:45",
      "content": "<p>Really interesting solution! Thanks for sharing your write up, and congrats on 2nd place <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>!</p>",
      "rawMarkdown": "Really interesting solution! Thanks for sharing your write up, and congrats on 2nd place @iafoss!",
      "votes": null
    },
    {
      "id": "2228513",
      "postDate": "04/20/2023 15:45:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team. Really details solution, thanks for sharing!</p>",
      "rawMarkdown": "Congrats @iafoss and team. Really details solution, thanks for sharing!",
      "votes": null
    },
    {
      "id": "2228563",
      "postDate": "04/20/2023 16:12:27",
      "content": "<p>Brilliant work! Thank you.<br>\nAs a physicist, I especially liked the use of \"Relative spacetime interval bias\".<br>\nWe will study your solution in more depth. It has a lot of great ideas.</p>",
      "rawMarkdown": "Brilliant work! Thank you.\nAs a physicist, I especially liked the use of \"Relative spacetime interval bias\".\nWe will study your solution in more depth. It has a lot of great ideas.",
      "votes": null
    },
    {
      "id": "2228567",
      "postDate": "04/20/2023 16:15:09",
      "content": "<p>Congrats! And thanks for sharing \"Chunk-based Data Loading with Caching\". </p>",
      "rawMarkdown": "Congrats! And thanks for sharing \"Chunk-based Data Loading with Caching\".",
      "votes": null
    },
    {
      "id": "2228812",
      "postDate": "04/20/2023 20:44:41",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Really impressive solution and write-up. congratulations!</p>",
      "rawMarkdown": "iafoss Really impressive solution and write-up. congratulations!",
      "votes": null
    },
    {
      "id": "2228921",
      "postDate": "04/20/2023 23:32:48",
      "content": "<p>Congratulations on winning second place! The sharing scheme is great and interesting, and deserves further thought.</p>",
      "rawMarkdown": "Congratulations on winning second place! The sharing scheme is great and interesting, and deserves further thought.",
      "votes": null
    },
    {
      "id": "2230174",
      "postDate": "04/22/2023 04:53:18",
      "content": "<p>congrats man</p>",
      "rawMarkdown": "congrats man",
      "votes": null
    },
    {
      "id": "2230230",
      "postDate": "04/22/2023 05:57:48",
      "content": "<p>congratulations and thank you for sharing this solution with us</p>",
      "rawMarkdown": "congratulations and thank you for sharing this solution with us",
      "votes": null
    },
    {
      "id": "2230390",
      "postDate": "04/22/2023 10:27:27",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> awesome achievement. Would you be willing to share your model code?</p>",
      "rawMarkdown": "iafoss awesome achievement. Would you be willing to share your model code?",
      "votes": null
    },
    {
      "id": "2230478",
      "postDate": "04/22/2023 11:54:04",
      "content": "<p>we are in process making repo and making inference notebook public! we will update the post with the links</p>",
      "rawMarkdown": "we are in process making repo and making inference notebook public! we will update the post with the links",
      "votes": null
    },
    {
      "id": "2230533",
      "postDate": "04/22/2023 13:12:58",
      "content": "<p>Awesome! I had issues creating a transformer architecture that is better than LSTM. Didn't have much time to experiment too much with it though.</p>",
      "rawMarkdown": "Awesome! I had issues creating a transformer architecture that is better than LSTM. Didn't have much time to experiment too much with it though.",
      "votes": null
    },
    {
      "id": "2231268",
      "postDate": "04/23/2023 07:37:47",
      "content": "<p>Congratulations on securing the 2nd position, and thank you for providing a detailed explanation of your solution!</p>\n<p>I have a question regarding your statement: \"Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time).\" How were you able to use different sequence lengths for training and inference? My understanding was that the Transformer architecture has a fixed input sequence length.</p>",
      "rawMarkdown": "Congratulations on securing the 2nd position, and thank you for providing a detailed explanation of your solution!\n\nI have a question regarding your statement: \"Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time).\" How were you able to use different sequence lengths for training and inference? My understanding was that the Transformer architecture has a fixed input sequence length.",
      "votes": null
    },
    {
      "id": "2231915",
      "postDate": "04/23/2023 20:01:57",
      "content": "<p>Transformer can use any length from the architecture point of view, the limitation you refer is coming from the predefined positional encoding in some transformer models, like ViT. In our case we do not need anything like that. </p>",
      "rawMarkdown": "Transformer can use any length from the architecture point of view, the limitation you refer is coming from the predefined positional encoding in some transformer models, like ViT. In our case we do not need anything like that.",
      "votes": null
    },
    {
      "id": "2231939",
      "postDate": "04/23/2023 20:50:22",
      "content": "<p>Oh, I see! Thanks for clearing that up.</p>",
      "rawMarkdown": "Oh, I see! Thanks for clearing that up.",
      "votes": null
    },
    {
      "id": "2232003",
      "postDate": "04/23/2023 22:26:12",
      "content": "<p>Congrats!)))</p>",
      "rawMarkdown": "Congrats!)))",
      "votes": null
    },
    {
      "id": "2232046",
      "postDate": "04/24/2023 00:22:07",
      "content": "<p>Congratulations on your win!!</p>",
      "rawMarkdown": "Congratulations on your win!!",
      "votes": null
    },
    {
      "id": "2232320",
      "postDate": "04/24/2023 07:42:20",
      "content": "<p>Congratulations on your win!</p>",
      "rawMarkdown": "Congratulations on your win!",
      "votes": null
    },
    {
      "id": "2232820",
      "postDate": "04/24/2023 16:46:55",
      "content": "<p>Great solution and great write-up!  Several insights which will take some thinking over.</p>\n<p>Your random selection rule outperforming more involved algos is interesting, and I could think of an intuitive reason.  For low-energy (few-pulses) events there is no need to drop them. For the high-energy (many-pulses) events, to predict the direction it is more important to preserve the general topology of the pulse cloud, than to retain earlier or more charged pulses.  When you drop randomly from a dense cloud you get a thinner cloud of the same form. When you drop by charge or time, you may well eliminate  half of the track and bring accuracy down.  </p>",
      "rawMarkdown": "Great solution and great write-up!  Several insights which will take some thinking over.\n\nYour random selection rule outperforming more involved algos is interesting, and I could think of an intuitive reason.  For low-energy (few-pulses) events there is no need to drop them. For the high-energy (many-pulses) events, to predict the direction it is more important to preserve the general topology of the pulse cloud, than to retain earlier or more charged pulses.  When you drop randomly from a dense cloud you get a thinner cloud of the same form. When you drop by charge or time, you may well eliminate  half of the track and bring accuracy down.",
      "votes": null
    },
    {
      "id": "2234114",
      "postDate": "04/24/2023 22:17:10",
      "content": "<p>Congrats :) </p>",
      "rawMarkdown": "Congrats :)",
      "votes": null
    },
    {
      "id": "2234844",
      "postDate": "04/25/2023 14:21:49",
      "content": "<p>Congratulations on the 2nd place!<br>\nYour ideas are excellent and I was amazed very much.<br>\nI could understand well thanks to your great write-up.</p>",
      "rawMarkdown": "Congratulations on the 2nd place!\nYour ideas are excellent and I was amazed very much.\nI could understand well thanks to your great write-up.",
      "votes": null
    },
    {
      "id": "2234978",
      "postDate": "04/25/2023 16:04:30",
      "content": "<p>Congratulations on your win!!!!!</p>",
      "rawMarkdown": "Congratulations on your win!!!!!",
      "votes": null
    },
    {
      "id": "2235955",
      "postDate": "04/26/2023 12:53:00",
      "content": "<p>Thanks for the great work! The given article is not only interesting, but may also be useful to other participants in solving and presenting the results of other similar tasks.</p>",
      "rawMarkdown": "Thanks for the great work! The given article is not only interesting, but may also be useful to other participants in solving and presenting the results of other similar tasks.",
      "votes": null
    },
    {
      "id": "2236649",
      "postDate": "04/27/2023 03:01:30",
      "content": "<p>The code is available now <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">https://github.com/DrHB/icecube-2nd-place</a> </p>",
      "rawMarkdown": "The code is available now https://github.com/DrHB/icecube-2nd-place",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227866,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "04/20/2023 04:43:36",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks. Very detailed and interesting solution! Congrats with 2 place!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2227882,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/20/2023 04:59:38",
          "content": "<p>Thanks     </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2227877,
      "author_name": "inartimiryasov",
      "author_url": "",
      "post_date": "04/20/2023 04:51:47",
      "content": "<p>Wow, this is truly impressive! Congratulations, and thank you for providing such a comprehensive solution!</p>\n<p>We had nearly all the necessary components (Fourier encoder, Relative spacetime interval bias, and a speedy dataloader built on a memmap file), but were training our model on a GTX 1080. Consequently, we never observed any significant improvements compared to a GRU model 😂</p>\n<p>Additionally, the awareness of data leakage is crucial for physics applications. It's fantastic that you have drawn attention to this issue.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2227914,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "04/20/2023 05:30:30",
      "content": "<p>Wow I never thought of that 0 time leak, always trained with the first detection time subtracted. Wonder how much difference that would have made.</p>\n<p>The classifier you mention is similar to mine, instead of joint map I have 2 heads and it did work, it was rather much more stable for me.</p>\n<p>Your use of fourier encoder and relative bias is amazing, learnt a lot from reading this, thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228063,
      "author_name": "johin714",
      "author_url": "",
      "post_date": "04/20/2023 08:38:22",
      "content": "<p>I am New to Neural networks hope you guys will help me out 😁 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228240,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/20/2023 11:30:56",
      "content": "<p>Congratulations on the second place. This write-up is quite interesting and the architecture is impressive. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228289,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "04/20/2023 12:28:11",
      "content": "<p>Congratulations on your finish and the very thorough writeup.  What tool did you use to draw your model diagram?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228370,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/20/2023 13:50:29",
          "content": "<p>Thank you. It was combination of Figma and a little bit of midjourneyv5:) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228301,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "04/20/2023 12:43:02",
      "content": "<p>Sorry for a nitpick, but your link in \"<em>We use Fourier encoding representation, often used</em>\" seems to be the Attention paper.  Is that where Fourier Encoding is discussed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228353,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/20/2023 13:40:44",
          "content": "<p>Yes, it is the place where it was introduced for the first time to my knowledge for the description of the position of tokens in the sequence, but I may be wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228383,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/20/2023 14:01:01",
      "content": "<p>Congratulations Dream Team! I am reading your solution description and it is amazing. I will spend a lot of time learning from this description and from you. 🙏🙏🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228460,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "04/20/2023 15:07:45",
      "content": "<p>Really interesting solution! Thanks for sharing your write up, and congrats on 2nd place <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228513,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "04/20/2023 15:45:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and team. Really details solution, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228563,
      "author_name": "synset",
      "author_url": "",
      "post_date": "04/20/2023 16:12:27",
      "content": "<p>Brilliant work! Thank you.<br>\nAs a physicist, I especially liked the use of \"Relative spacetime interval bias\".<br>\nWe will study your solution in more depth. It has a lot of great ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228567,
      "author_name": "rusg77",
      "author_url": "",
      "post_date": "04/20/2023 16:15:09",
      "content": "<p>Congrats! And thanks for sharing \"Chunk-based Data Loading with Caching\". </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228812,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/20/2023 20:44:41",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Really impressive solution and write-up. congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228921,
      "author_name": "mewmlelswm",
      "author_url": "",
      "post_date": "04/20/2023 23:32:48",
      "content": "<p>Congratulations on winning second place! The sharing scheme is great and interesting, and deserves further thought.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2230174,
      "author_name": "shreyashkawalkar",
      "author_url": "",
      "post_date": "04/22/2023 04:53:18",
      "content": "<p>congrats man</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2230230,
      "author_name": "thilakshangayan",
      "author_url": "",
      "post_date": "04/22/2023 05:57:48",
      "content": "<p>congratulations and thank you for sharing this solution with us</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2230390,
      "author_name": "crodoc",
      "author_url": "",
      "post_date": "04/22/2023 10:27:27",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> awesome achievement. Would you be willing to share your model code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230478,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/22/2023 11:54:04",
          "content": "<p>we are in process making repo and making inference notebook public! we will update the post with the links</p>",
          "votes": null,
          "replies": [
            {
              "id": 2230533,
              "author_name": "crodoc",
              "author_url": "",
              "post_date": "04/22/2023 13:12:58",
              "content": "<p>Awesome! I had issues creating a transformer architecture that is better than LSTM. Didn't have much time to experiment too much with it though.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2236649,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/27/2023 03:01:30",
          "content": "<p>The code is available now <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">https://github.com/DrHB/icecube-2nd-place</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2231268,
      "author_name": "viktorcikojevic",
      "author_url": "",
      "post_date": "04/23/2023 07:37:47",
      "content": "<p>Congratulations on securing the 2nd position, and thank you for providing a detailed explanation of your solution!</p>\n<p>I have a question regarding your statement: \"Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time).\" How were you able to use different sequence lengths for training and inference? My understanding was that the Transformer architecture has a fixed input sequence length.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2231915,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/23/2023 20:01:57",
          "content": "<p>Transformer can use any length from the architecture point of view, the limitation you refer is coming from the predefined positional encoding in some transformer models, like ViT. In our case we do not need anything like that. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2231939,
              "author_name": "viktorcikojevic",
              "author_url": "",
              "post_date": "04/23/2023 20:50:22",
              "content": "<p>Oh, I see! Thanks for clearing that up.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2232003,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "04/23/2023 22:26:12",
      "content": "<p>Congrats!)))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232046,
      "author_name": "ifeanyichukwunwobodo",
      "author_url": "",
      "post_date": "04/24/2023 00:22:07",
      "content": "<p>Congratulations on your win!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232320,
      "author_name": "yoshikihiragaki",
      "author_url": "",
      "post_date": "04/24/2023 07:42:20",
      "content": "<p>Congratulations on your win!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2232820,
      "author_name": "alexz0",
      "author_url": "",
      "post_date": "04/24/2023 16:46:55",
      "content": "<p>Great solution and great write-up!  Several insights which will take some thinking over.</p>\n<p>Your random selection rule outperforming more involved algos is interesting, and I could think of an intuitive reason.  For low-energy (few-pulses) events there is no need to drop them. For the high-energy (many-pulses) events, to predict the direction it is more important to preserve the general topology of the pulse cloud, than to retain earlier or more charged pulses.  When you drop randomly from a dense cloud you get a thinner cloud of the same form. When you drop by charge or time, you may well eliminate  half of the track and bring accuracy down.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234114,
      "author_name": "branislavdoubek",
      "author_url": "",
      "post_date": "04/24/2023 22:17:10",
      "content": "<p>Congrats :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234844,
      "author_name": "yamashitamotokazu",
      "author_url": "",
      "post_date": "04/25/2023 14:21:49",
      "content": "<p>Congratulations on the 2nd place!<br>\nYour ideas are excellent and I was amazed very much.<br>\nI could understand well thanks to your great write-up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2234978,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "04/25/2023 16:04:30",
      "content": "<p>Congratulations on your win!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2235955,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "04/26/2023 12:53:00",
      "content": "<p>Thanks for the great work! The given article is not only interesting, but may also be useful to other participants in solving and presenting the results of other similar tasks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2227847": "# Summary\n-\t**Transformer-based solution**\n-\tChunk-based data loading with caching and length matching\n-\tFourier encoder\n-\tRelative spacetime interval bias\n-\tvon Mises-Fisher Loss + the competition metric\n\n**Our code is available at [this repo](https://github.com/DrHB/icecube-2nd-place)**.\n\n# Introduction\nWhile neutrinos are one the most abundant particles in the universe that may bring information about violent astrophysical events, their fundamental properties make them difficult to detect and requires a huge volume of material to interact with. The IceCube Neutrino Observatory is the first detector of its kind, which consists of a cubic kilometer of ice and is designed to search for nearly massless neutrinos. The rare interaction of neutrinos with the matter may produce high-energy particles emitting Cherenkov radiation, which is recorded with photo-detectors immersed in ice. The objective of this challenge was a reconstruction of the neutrino trace direction based on the photon detections. \nTo begin with, I would really like to express my gratitude to the organizers and Kaggle team for making this competition possible. Reading the description of the competition I immediately realized that transformers may be a really good approach for it. Since I wanted to experiment with transformers in depth, I decided to join this competition. Meanwhile, my outstanding teammate @drhabib has joined this competition [hoping to visit Antarctica](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381481#2117294). \n\n# Details\nOur solution is based on transformers, which appeared to be ideal for the considered task. Below we provide the key components of our approach.\n## Why transformers\nGraph Neural Networks (GNNs) exhibit limitations in certain areas:\n(1)\tGNNs function primarily on a local level, as updates are performed based on neighboring nodes. Meanwhile, these models are expected to predict global quantities, such as the direction of a track. Organizers have attempted to address this issue through dynamic edge building, but this method falls short in comparison to the multi-head self-attention mechanism found in transformers. Moreover, dynamic edge building does not provide gradient feedback for neighbor selection.\n(2)\tUpon initial examination, GNNs might appear faster than transformers due to their consideration of only a fixed number of neighbors. Nevertheless, extensive sparse operations and data rearrangements in GNNs can be significantly slower than the highly optimized dense tensor operations in transformers for a reasonable sequence length and considering all possible interactions. In our experiments, the organizers’ baseline has the same computational cost as our reference T model (discussed below) while having only 1.4M parameters vs. 7.4M for T model.\n**Transformer is a GNN on steroids**. It may be considered as GNN on a fully connected graph using attention to estimate the edge weights dynamically. But this power comes with the need for huge training data.\n## Data\nThe organizers generously provided a sizable training dataset, which may be somewhat challenging to manage. Since transformers require massive data for training, it was vital to set up a proper data pipeline at the very beginning. Our approach consists of three primary components: (1) caching the selected data chunks to minimize the substantial cost of data loading and preprocessing, (2) employing chunk-based random sampling to effectively utilize caching, and (3) **implementing length-matched data sampling** (sapling batches with approximately equal lengths) to reduce the computational overhead associated with padding tokens when truncating the batch at the longest sequence. Components (1) and (2) offer a computationally efficient and RAM-friendly handling of the data and were shared through the [\"Chunk-based Data Loading with Caching\" notebook](https://www.kaggle.com/code/iafoss/chunk-based-data-loading-with-caching/notebook) during the competition to facilitate the setup of a comprehensive training data pipeline for participants. **Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time)**. Inference at a 768 -length led to a 25 bps improvement over the 192-length. Initially, we also attempted to use Hugging Face datasets, which offered the ability to utilize chunked data but required additional file conversion; ultimately, we settled on a chunk caching pipeline.\nIn our study, we observed no significant difference between utilizing only basic features (time, sensor position, charge, auxiliary, and the total number of detections) and incorporating additional features related to ice transparency and absorption. This outcome is expected, as for sufficiently large training datasets, the extra features merely offer an alternate representation of the z-coordinate, with the corresponding transparency and absorption distributions being explicitly learned. To ensure a diverse range of models, we employed both the base and extended feature sets in our experiments. The features are normalized in a similar way as in the organizers' baseline. **It is important to consider detections with both auxiliary true and false.** For events exceeding the maximum considered sequence length, we implemented the following selection process: Initially, we randomly selected detections from the auxiliary false subset, and if this subset proved insufficient to create a sequence of the required length, we then randomly sampled noisy auxiliary true detections. It is interesting that this random selection approach outperformed more complex methods, such as selecting based on charge or considering only events within a specific temporal boundary (IceCube size divided by the speed of light) surrounding the detection with the maximum charge.\n## Model\nThe model is schematically illustrated in the plot below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F552b4994e2931207aa2267c170d6044d%2FModel2.png?generation=1681961703319090&alt=media)\n*Fourier encoder*. Our solution is based on a transformer model, which considers each event as a sequence of detections. It is crucial to process the provided continuous input, such as time or charge, into a form suitable for transformers. We use [Fourier encoding representation](https://arxiv.org/pdf/1706.03762.pdf), often used to describe the position in the sequence in language models. This method can be viewed as a soft digitization of the continuous input signal into a set of codes determined by Fourier frequencies. This procedure is applied for all continuous input variables, while the discrete auxiliary flag is encoded with learnable embedding. We multiply the normalized input variables by 1024-4096 to have a sufficient resolution after digitizing. For example, after multiplying the normalized time by 4096, the temporal resolution becomes 7.3 ns, and the model may understand smaller time variations because of the continuous nature of the Fourier encoding. This multiplication is critical, and even a change of the coefficient from 128 to 4096 gives 20 bps boost.\nThe encoding of all variables is concatenated, passed through one GELU layer, and reprojected into the transformer dimension. An alternative to Fourier encoding may be a multilayer network that learns the corresponding representation of continuous variables, but it may require many more layers than we consider in our setup. \n*GraphNet encoder*. In addition to the use of Fourier encoding, in several of our models, we incorporated GraphNet feature extractor. In addition to learning the representation of input variables, it also constructs features accounting for the neighbors. \n*Relative spacetime interval bias*. In the special theory of relativity there is a quantity called spacetime interval: ds^2=c^2 dt^2-dx^2-dy^2-dz^2. For particles moving with a speed close to the speed of light, it is close to zero. What is more useful is that all particles and photons produced in the reactions caused by a single neutrino should have ds^2 close to zero too (it is not fully precise because the refractive index of ice is about 1.3, and the speed of light in it is lower than c).  Therefore, by **computing ds^2 between all pairs of detections in the given event, it is relatively straightforward to distinguish between detections originating from the same neutrino and those that are merely noise**, as demonstrated in the figure below. This criterion can be naturally introduced to a transformer as a [relative bias](https://arxiv.org/pdf/1803.02155v2.pdf). Therefore, during the construction of the attention matrix, the transformer automatically groups detections based on the source event effectively filtering out noise. We use Fourier encoding representation of ds^2 to build the bias term. **Use of relative bias boosts the performance by 40 bps**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fa96019c14603b2c54be402d9d960464f%2Fds.png?generation=1681961957521492&alt=media)\n*Model*. We use the standard [BEiT building blocks](https://arxiv.org/pdf/2208.06366.pdf)  (transformer with learnable shortcuts). To incorporate relative bias, the first transformer blocks are modified according to [this paper](https://arxiv.org/pdf/1803.02155v2.pdf). The model size is characterized with a standard ViT notation: T – tiny (dim 192), S – small (dim 384), B – base (dim 768). We use 16 transformer blocks in total: 4 blocks with relative bias + 12 blocks with cls token. The use of an additional cls token, included in the input sequence, enables the gradual collection of information about the track direction throughout the entire model. The attention head size is varied between 32 and 64: the smaller head size gives better performance, especially for the T model (10 bps). The details of the considered models are summarized below.\n\n| Model | dimension | number of heads | depth | number of parameters |\n| --- | --- | --- | --- | --- |\n| T | 192 | 3-6 | 4+12 | 7.57M |\n| S | 384 | 6-12 | 4+12 | 29.3M |\n| B | 768\t| 12-24 | 4+12 | 115.6M |\n\n\n*Training setup*. The models are trained with AdamW optimizer (weight decay 0.05) for 4-5 epochs (all data is considered). The learning rate is changed with a cosine annealing scheduler with a warmup. The maximum learning rate at the first epoch is 5e-4 for T and S models, and 1e-4 for B models. The following epochs are trained with the maximum learning rate of 0.5e-5 - 2e-5. The effective batch size is 4096 in all experiments (with the use of gradient accumulation). Starting from the second epoch Stochastic Weight Average is used. \nWe perform training for the first 2-3 epochs with von Mises-Fisher Loss (the model predicts 3d vectors, and kappa is computed based on the length of the produced vector), and the remaining epochs are performed with using the competition metric as the objective function + 0.05 von Mises-Fisher Loss (to give the model feedback on the vector length vs. confidence relation and simplify ensembling step). Use of the competition metric as the loss gives 55 bps boost. \n\n# Results\nThe table below summarizes the performance of our finalized models. CV is evaluated at L=512. (*) refers to models trained on 2×A6000, which in contrast to rtx4090 enables larger actual batch size. Further finetuning with L=256 may result in about 5 bps boost.\n\t\n| Model setup | CV512 | Public LB | Private LB| \n| --- | --- | --- | --- |\n| T d32 | 0.9704 | 0.9693 | 0.9698 |\n| S d32 | 0.9671 | 0.9654 | 0.9659 | \n| B d32 | 0.9642 | 0.9623 | 0.9632 |\n| B d64 | 0.9645 | 0.9635 | 0.9629 |\n| B+gnn d48 | 0.9643 | 0.9624 | 0.9627 |\n| * S+gnn d32 | 0.9639 | 0.9620 | 0.9628 |\n| * B+gnn d32 | 0.9633 | 0.9609 | 0.9621 |\n\nEach increase of the model to the next size T->S->B gives approximately 30 bps but is accompanied by an ×4 increase in the number of model parameters. However, the provided data may be insufficient for training models larger than B. Meanwhile, the results of the use GNN in addition to the Fourier extractor are rather contradictive and may be affected by the difference in the hardware configuration.  \nSo, even **the smallest considered T model with only 7.53M parameters and fast training (under 12 hours on rtx4090 per epoch) can take the top-5 in this competition** by itself. Meanwhile, with additional finetuning at L=256, **our best model, B+gnn d32, reaches 0.9628 CV512, 0.9608 public, and 0.9618 private LB as a single model**. \n\n## Model ensemble\nWith the use of von Mises-Fisher Loss, the vector length corresponds to the model confidence. Therefore, a weighted average of the predicted vectors automatically incorporates the confidence and biases the prediction towards the most confident direction. \nOur best final submission in the linear ensemble of 6 models gives 0.9594 at public and 0.9602  at private LB (the contribution of each model is evaluated as an average of weights fitted in 5 random 5-fold splits). Ensembling gives only an insignificant improvement over our best single model submission of 0.9608 public and 0.9618 at private LB. We also tried to consider more elaborated ensembling by utilizing additional features such as sequence length, first detection time, total charge, etc., but without any improvement over the simple linear ensemble\n\n## Training and inference time\n**On rtx4090 the inference time can be as short as 3 min per 1M samples for the T model** (without inference optimization) and may increase to 8-10 min for the B model. On outdated GPUs several generations behind, such as P100 at Kaggle, the inference time may increase by x10. The training time varies from 12 hours per epoch for the T model to 56 hours per epoch for the B+gnn model. On a machine with 2×A6000 the speed is approximately the same (~10% faster) than running the model on a single rtx4090.\n\n## Robustness\nRegarding the robustness of model predictions with respect to perturbation of the input parameters, we have considered the following cases: (1) **temporal noise**: add normally distributed random variable to detection time of each event; (2) **charge noise**: multiply charge by 10^ normally distributed randomly with variance reported at the horizontal axis; (3) **auxiliary noise**: randomly flip auxiliary flag with a given probability; and (4) **length reduction**: remove a given percentage of detections. The plot below reports results for our reference T model (noise-free CV512 is 0.9704).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Ffdbfc788586698d1965af8a024868dd7%2FNoise.png?generation=1681963294066000&alt=media)\nAlthough the model was not explicitly trained to consider the abovementioned types of noise, the predictions in most cases do not degrade by more than 15% at the largest magnitude of the considered noise, e.g. adding 200 ns random time error or ×2 reduction of the number of detected events. Training the model with specific data augmentation during training may improve the robustness of the model. However, there may be a limit when the error starts rapidly increasing, e.g. 10 ns temporal noise or a change in the value of detected charge by 40%.\nThe considered model is versatile enough to apply to any neutrino telescope with a similar detector concept as IceCube, regardless of the number of detectors and their arrangement. However, since during the training of our models, the positions of the detectors are fixed, the model may not have a complete understanding of the geometry and experience the degradation of the performance if the detectors are moved. For example, the performance of T model drops from 0.9704 to 1.004 under flip of all detector positions along X or Y axes. To avoid such performance degradation and improve the understanding of geometry by the model, the training could be performed at randomly generated configurations of detectors at each detection event. Incorporation of geometry augmentation, such as shifts, flips, and rotation, would also improve the model robustness to the change of the detector geometry. The less optimal solution is the generation of new training data for a new neutrino telescope. Consideration of [3D Roto-Translation Equivariant Attention](https://arxiv.org/pdf/2006.10503.pdf) models, in which detector positions are incorporated through attention rather than direct input, could potentially address the issue of detector geometry. However, this improvement may come with both the computational cost and VRAM requirement increase, and, therefore, a more practical solution may be training a simple model with spatial augmentation and randomized detector positions.\nThe plot below visualizes the distribution of error for confident and unconfident prediction.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fd1fe483e3972b5ae6aafa7c4db3f4ea9%2Fdist.png?generation=1681963384438419&alt=media)\n## Interesting findings\nIt appeared that **the majority of participants of the competition are unconsciously using the following leak**. The provided data is the result of a simulation, and 0 time has a particular meaning revealing some details of the generation process (the time of the first detection with respect to 0 time may probably give such information as the energy and approximate direction limitations based on the traveling time of the neutrino from the box boundary to the detector). The models learn how to utilize this leak and improve its performance. For example, one of our experiments with the T model trained (1) on data with keeping 0-time-reference and (2) subtracting the time of the first detection from all detections gives the performance gap of 145 bps at the end of the first epoch. \nIn the real detector data, the 0-time-reference is not available, and the time might be considered with the reference to the first detected event only. For models trained using 0-time-reference, such experimental data should be shifted by the average first detection time. However, in this case, the performance may drop by 150+ bps, comparable to the results of the model trained without the 0-time leak. \n\n## Things that did not work\nIn this competition, we have experimented with a number of things that, unfortunately, did not give an improvement:\n**Graph Neural Networks**. After building the first transformer-based pipeline it becomes apparent that GNNs are well behind (0.98-0.99 vs. 1.0 LB for first models). One interesting comparison we ran at that moment is dividing the predictions based on the kappa by 30% of confident and 70% of unconfident (noisy input and very small number of detections for track reconstruction). The performance of both models on unconfident predictions was comparable, while the improvement on confident predictions was the thing giving the performance gap between transformers and GNNs. It is fully in line with our initial expectation that transformers with global consideration of the input data are more appropriate for the prediction of the global quantity as the track direction in contrast to local GNNs. So further efforts were spent on enhancing the transformer setup. So, in the end, we used GNN as an optional addition to our Fourier feature extractor for diversity.\n**Ice properties**. We did not see that use of additional input, such as ice properties, helps to improve model performance.\n**Local attention**. We spent several weeks experimenting with local attention. The basic idea is simple: before giving a sequence of detections to the transformer model, which considers all-vs-all interactions between detections, process the features within the local neighbors. In theory, local attention is expected to be fast and memory efficient since it examines only a small number of k=8 neighbors instead of requiring L×L attention. However, in practice, local attention appears to be quite slow because of extensive data shuffling (similar to GNN), and in some naïve implementations, it also may consume significantly more VRAM than highly optimized transformers. Surprisingly, the fastest local attention is just regular transformer attention with masking attention for elements beyond k neighbors. Such masking helps at the beginning of training, but at the later stage, full attention outperforms it or gives comparable results. We also tried to consider configurations with alternating global-local attention (sandwich) or parallel blocks of local and global attention in the part of the model preceding the regular 12-layer transformer. However, a simple configuration with consideration of 4 layers with full attention and rel bias + 12-layer regular transformer outperformed the setups with local attention.\n**Classification loss**. We tried to subdivide the angle space into bins and utilize classification loss to fight against angle uncertainty. The ground truth is represented as a 2D Gaussian around the provided label, as illustrated in the figure below. We introduced 8 cls tokens to predict the probability for a 64×128 angular bin map (32×32 per token). Unfortunately, this model ended up at approximately 1.0 CV, and we did not perform any further checks on this approach. However, this model is quite interesting to mention, and it may provide a good visualization of the predictions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F1a2b2b993bb6a79ad13babb00ea689dc%2Fcls.png?generation=1681963515503826&alt=media)\n**L and H models** (400 and 800M parameters). Because of  Nvidia virtual address space driver bug, 2×rtx4090 could not be used together until 1-2 weeks before the end of the competition. And even after this bug was fixed, disabled p2p makes multi-GPU training with these cards to be quite inefficient. Therefore, we were unable to scale up our runs and stopped at relatively small B models. The expected boost at each model size increase is about 30 bps, and L and H single models could likely reach 0.958 and 0.955-0.956 LB as a single model but with a massive increase of the required compute for both training and inference.",
    "2227866": "iafoss thanks. Very detailed and interesting solution! Congrats with 2 place!",
    "2227877": "Wow, this is truly impressive! Congratulations, and thank you for providing such a comprehensive solution!\n\nWe had nearly all the necessary components (Fourier encoder, Relative spacetime interval bias, and a speedy dataloader built on a memmap file), but were training our model on a GTX 1080. Consequently, we never observed any significant improvements compared to a GRU model 😂\n\nAdditionally, the awareness of data leakage is crucial for physics applications. It's fantastic that you have drawn attention to this issue.",
    "2227882": "Thanks",
    "2227914": "Wow I never thought of that 0 time leak, always trained with the first detection time subtracted. Wonder how much difference that would have made.\n\nThe classifier you mention is similar to mine, instead of joint map I have 2 heads and it did work, it was rather much more stable for me.\n\nYour use of fourier encoder and relative bias is amazing, learnt a lot from reading this, thank you for sharing.",
    "2228063": "I am New to Neural networks hope you guys will help me out 😁",
    "2228240": "Congratulations on the second place. This write-up is quite interesting and the architecture is impressive. Thanks for sharing!",
    "2228289": "Congratulations on your finish and the very thorough writeup.  What tool did you use to draw your model diagram?",
    "2228301": "Sorry for a nitpick, but your link in \"*We use Fourier encoding representation, often used*\" seems to be the Attention paper.  Is that where Fourier Encoding is discussed?",
    "2228353": "Yes, it is the place where it was introduced for the first time to my knowledge for the description of the position of tokens in the sequence, but I may be wrong.",
    "2228370": "Thank you. It was combination of Figma and a little bit of midjourneyv5:)",
    "2228383": "Congratulations Dream Team! I am reading your solution description and it is amazing. I will spend a lot of time learning from this description and from you. 🙏🙏🙏",
    "2228460": "Really interesting solution! Thanks for sharing your write up, and congrats on 2nd place @iafoss!",
    "2228513": "Congrats @iafoss and team. Really details solution, thanks for sharing!",
    "2228563": "Brilliant work! Thank you.\nAs a physicist, I especially liked the use of \"Relative spacetime interval bias\".\nWe will study your solution in more depth. It has a lot of great ideas.",
    "2228567": "Congrats! And thanks for sharing \"Chunk-based Data Loading with Caching\".",
    "2228812": "iafoss Really impressive solution and write-up. congratulations!",
    "2228921": "Congratulations on winning second place! The sharing scheme is great and interesting, and deserves further thought.",
    "2230174": "congrats man",
    "2230230": "congratulations and thank you for sharing this solution with us",
    "2230390": "iafoss awesome achievement. Would you be willing to share your model code?",
    "2230478": "we are in process making repo and making inference notebook public! we will update the post with the links",
    "2230533": "Awesome! I had issues creating a transformer architecture that is better than LSTM. Didn't have much time to experiment too much with it though.",
    "2231268": "Congratulations on securing the 2nd position, and thank you for providing a detailed explanation of your solution!\n\nI have a question regarding your statement: \"Component (3) considerably expedited training and allowed us to use the maximum sequence length of 192 for training and 768 for inference (without a noticeable increase in inference time).\" How were you able to use different sequence lengths for training and inference? My understanding was that the Transformer architecture has a fixed input sequence length.",
    "2231915": "Transformer can use any length from the architecture point of view, the limitation you refer is coming from the predefined positional encoding in some transformer models, like ViT. In our case we do not need anything like that.",
    "2231939": "Oh, I see! Thanks for clearing that up.",
    "2232003": "Congrats!)))",
    "2232046": "Congratulations on your win!!",
    "2232320": "Congratulations on your win!",
    "2232820": "Great solution and great write-up!  Several insights which will take some thinking over.\n\nYour random selection rule outperforming more involved algos is interesting, and I could think of an intuitive reason.  For low-energy (few-pulses) events there is no need to drop them. For the high-energy (many-pulses) events, to predict the direction it is more important to preserve the general topology of the pulse cloud, than to retain earlier or more charged pulses.  When you drop randomly from a dense cloud you get a thinner cloud of the same form. When you drop by charge or time, you may well eliminate  half of the track and bring accuracy down.",
    "2234114": "Congrats :)",
    "2234844": "Congratulations on the 2nd place!\nYour ideas are excellent and I was amazed very much.\nI could understand well thanks to your great write-up.",
    "2234978": "Congratulations on your win!!!!!",
    "2235955": "Thanks for the great work! The given article is not only interesting, but may also be useful to other participants in solving and presenting the results of other similar tasks.",
    "2236649": "The code is available now https://github.com/DrHB/icecube-2nd-place"
  },
  "source": "meta"
}