{
  "id": 403000,
  "title": "12th Place Solution Details - IceCube",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/solverworld-12th-place-solution-details-icecube",
  "author_name": "",
  "post_date": "2023-04-27T11:28:53.257Z",
  "votes": 25,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Summary</h1>\n<p>My final solution used an ensemble with a simple weighted average of 6 models that were generalizations of publicly posted models: (1) the DynEdge model <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> , <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> , and (2) LSTM model by <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>.  Each model predicted directly a 3-dimensional vector of the target direction, using a loss that was the norm of the difference between the predicted vector and the target vector.  The predicted vector was not normalized first.</p>\n<h1>TLDR:</h1>\n<ol>\n<li>Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.</li>\n<li>Pool of all models included 30 different models, best combination of 6 selected through optimization method.<br>\nModels included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks</li>\n<li>Models included different number of pulses selected for training and different number of pulses used for inference (not always the same). Pulses were selected randomly from entire set of pulses.</li>\n<li>Each model was used (effectively) 4 times and averaged (because of the random selection of pulse set) - a simple observation was that smaller events did not require multiple evaluations.</li>\n<li>For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling). Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.</li>\n<li>Each model used as loss the norm of predicted vector minus true unit direction vector.</li>\n</ol>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">Robin Smits LSTM models</a><br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">RASMUS ØRSØE DynEdge Post</a><br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">My Derivation of Least Squares Fit</a><br>\n<a href=\"https://www.kaggle.com/code/solverworld/12th-place-ensemble-models\" target=\"_blank\">My Ensembling Notebook</a> [The attached dataset contains the pytorch models]<br>\nMirco Hunnefield Masters Thesis, Online Reconstruction of Muon-Neutrino Events in IceCube using Deep Learning Techniques<br>\nKai Schatto, PhD Thesis, Stacked searches for high-energy neutrinos from blazars with IceCube</p>\n<h1>Early Days</h1>\n<p>I started by reading as many of the neutrino detection papers as I could, trying to understand the nature of the problem.  There were many excellent posts with graphics (taken from the papers) showing how neutrinos generated a muon that then formed spinoff collisions that illuminated a large set of sensors embedded in the Antarctic Ice.  I won't repeat all that excellent information, but I will try to outline my thought process as my solution evolved.</p>\n<p>My first submissions were the (<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">line-fit solutions</a> that got 1.214 and later improved to 1.174 with more clever picking of pulses through neighbor counting (thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> !).</p>\n<p>I then decided to move on to neural net models.  Since there was a lot of excitement around the DynEdge models, I though maybe I should zig while others were zagging, and looked into other graph neural networks.  I implemented PointNet and then PointNet++, which are basically DynEdge models where you do not dynamically alter the graph at each iteration.  That is, the graph is fixed by neighbors within a certain radius, and those neighbors are then operated on much as a Convolutional NN (CNN) does for the entire solution.  The geometric pytorch  <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/cheatsheet/gnn_cheatsheet.html\" target=\"_blank\">cheatsheet</a> gives a nice summary of all the GNN models it supports with linked papers.  </p>\n<p>After looking at all the GNN models, and failing to have PointNet++ do much, I thought I should dive into the DynEdge models.  Trying to install the graphnet module required backing up too many python modules (and using Python 3.7 instead of my local 3.10).  So, since I hate to go backwards, I pulled out the dynedge code from the entire graphnet module, and used py-geometric to build my DynEdge models.  This also freed me from the SQlite and other framework difficulties so I could build my own pipeline.</p>\n<p>I solved the data size problem by preprocessing each parquet batch file into 780 individual hdf5 files which contained 256x2000xK matrices which were the data for 256 events.  The pulses were limited to 2000 by Aux and Charge sorting.  The rest were zero-padded, but because hdf5 compresses the data, this did not result in larger storage.  The K features were x,y,z, time, charge, aux, num-of-neighbors, and some pre-computed time window information.</p>\n<p>I then was able to shuffle both batches and the 256-groups of events within each batch without massive amounts of memory.  I was able to feed my GPU efficiently using the preprocessed files.</p>\n<h1>Progression</h1>\n<p>I decided to use batches 1-100 for training, and batch 601 and 602 for validation.  I typically trained for 600-800 batches, thus used each batch only 6-8 times.  I see now from the top finishers that people trained on much more data and for longer.  My models were probably not complex enough to need more data.</p>\n<p>I got the DynEdge models working close to the reported public notebooks, and eventually tried a few variations, including adding more layers (up to 6) and a larger output head.  The default was just a 128 size FC layer, but I eventually used [384,256,128,64,3].  My best DynEdge models got around 1.001 on Batch 602 (my test batch).</p>\n<h1>Choice of Output</h1>\n<p>I tried outputting bins (thanks again, <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>) but I found better results with directly predicting the 3-dimension direction and using the norm ||v-pred|| as the loss function.  The gradient of this function for v close to to pred looks much like the Angular Error function, so I thought it should work.  I could not get the Angular Error to work directly as a loss function, because it does not control for the vector size and allows the predicted vector to shrink to zero size.  There also some inaccuracies involved in making sure you don't normalize a length 0 vector.</p>\n<h1>Choice of Input</h1>\n<p>I tried various ways of selecting the input pulses to use in my models.  Every sorting or choosing method seemed to work worse than random selection.  Because of the random selection nature, model accuracy could be improved by simply averaging 4 predictions <em>from the same model</em>.  One simple trick was to realize that events with less pulses than were required for a model did not have a random element and thus those did not need to be repeated 4 times.  This meant I could do inference 4x with only a 15% increase in computation effort on submissions.</p>\n<h1>Back to LSTM Models</h1>\n<p>I decided to go back the LSTM models and taking a hint from the 2nd place finish of the RSNA Intracranial Hemorrhage challenge ( <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> )  <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/117228\" target=\"_blank\">2nd Place</a> I used multiple LSTM layers, with the outputs of each layer concatenated together into a final multi-layer fully connected prediction head.<br>\nThe best models were 512 wide LSTMs in 3 layers.  I used prediction heads with layers [256,64,64,3] and variations thereof.  I also tried an Attention modification on the LSTM output which helped a little bit (this was simply a learned weight that is dotted with the LSTM output to produce a signal proportional the weight that should be given to that element in the sequence).</p>\n<h1>Trying all the Models</h1>\n<p>Here are the scores of 35 models, just to show the variation and improvement from averaging.  The s&lt;i&gt; indicates a different random selection of the pulses from each event, the p&lt;n&gt; indicates how many pulses were used in the inference (prediction) and the last two columns indicate the change in score by using an average of the 4 predictions and the standard deviation, respectively.  The score was done on Batch 602.<br>\nYou can see the average helps about 35 bps on models using 128 pulses, about 6-10 bps for models using 350 or 384 pulses.  The description includes the type of model as well as the number of pulses used for training.  E.g. LSTM512x3-384 was a 3 layer LSTM model trained on 384 pulses.  ATT means attention was used.<br>\nThe best GNN model gets 1.001, the best LSTM model gets 0.9859 on batch 602.  My scores on the leaderboard were about .002 lower than my cross validation scores on batch 602.  I did not submit any individual models to the leaderboard, as I did want to contaminate my procedure with fitting to the leaderboard.</p>\n<pre><code>000\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s0  0.98869\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s1  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s2  0.98886\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s3  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 --  0.98810 -0.00069  0.00006\n001\n01 dog-ust-state-752-s9 LSTM512x3        p128 s0  0.99525\n01 dog-ust-state-752-s9 LSTM512x3        p128 s1  0.99547\n01 dog-ust-state-752-s9 LSTM512x3        p128 s2  0.99538\n01 dog-ust-state-752-s9 LSTM512x3        p128 s3  0.99533\n01 dog-ust-state-752-s9 LSTM512x3        p128 --  0.99184 -0.00352  0.00008\n002\n02 dog-ust-state-752-s9 LSTM512x3        p200 s0  0.99056\n02 dog-ust-state-752-s9 LSTM512x3        p200 s1  0.99073\n02 dog-ust-state-752-s9 LSTM512x3        p200 s2  0.99069\n02 dog-ust-state-752-s9 LSTM512x3        p200 s3  0.99061\n02 dog-ust-state-752-s9 LSTM512x3        p200 --  0.98900 -0.00165  0.00007\n003\n03 dog-ust-state-752-s9 LSTM512x3        p256 s0  0.98949\n03 dog-ust-state-752-s9 LSTM512x3        p256 s1  0.98929\n03 dog-ust-state-752-s9 LSTM512x3        p256 s2  0.98943\n03 dog-ust-state-752-s9 LSTM512x3        p256 s3  0.98945\n03 dog-ust-state-752-s9 LSTM512x3        p256 --  0.98831 -0.00110  0.00008\n004\n04 dog-ust-state-752-s9 LSTM512x3        p320 s0  0.98896\n04 dog-ust-state-752-s9 LSTM512x3        p320 s1  0.98897\n04 dog-ust-state-752-s9 LSTM512x3        p320 s2  0.98885\n04 dog-ust-state-752-s9 LSTM512x3        p320 s3  0.98891\n04 dog-ust-state-752-s9 LSTM512x3        p320 --  0.98813 -0.00079  0.00005\n005\n05 dog-ust-state-752-s9 LSTM512x3        p384 s0  0.98868\n05 dog-ust-state-752-s9 LSTM512x3        p384 s1  0.98881\n05 dog-ust-state-752-s9 LSTM512x3        p384 s2  0.98874\n05 dog-ust-state-752-s9 LSTM512x3        p384 s3  0.98872\n05 dog-ust-state-752-s9 LSTM512x3        p384 --  0.98815 -0.00058  0.00005\n006\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s0  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s1  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s2  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s3  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 --  0.99047  0.00000  0.00000\n007\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s0  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s1  0.99138\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s2  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s3  0.99140\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 --  0.98955 -0.00175  0.00010\n008\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s0  0.98947\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s1  0.98931\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s2  0.98939\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s3  0.98936\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 --  0.98819 -0.00119  0.00006\n009\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s0  0.98846\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s1  0.98843\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s2  0.98848\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s3  0.98835\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 --  0.98756 -0.00087  0.00005\n010\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s0  0.98793\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s1  0.98807\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s2  0.98810\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s3  0.98798\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 --  0.98741 -0.00061  0.00007\n011\n11 dyno-dxu-state-456-s GNN5-RO          p200 s0  1.00377\n11 dyno-dxu-state-456-s GNN5-RO          p200 s1  1.00385\n11 dyno-dxu-state-456-s GNN5-RO          p200 s2  1.00370\n11 dyno-dxu-state-456-s GNN5-RO          p200 s3  1.00397\n11 dyno-dxu-state-456-s GNN5-RO          p200 --  1.00193 -0.00189  0.00010\n012\n12 dyno-dxu-state-456-s GNN5-RO          p256 s0  1.00346\n12 dyno-dxu-state-456-s GNN5-RO          p256 s1  1.00343\n12 dyno-dxu-state-456-s GNN5-RO          p256 s2  1.00341\n12 dyno-dxu-state-456-s GNN5-RO          p256 s3  1.00340\n12 dyno-dxu-state-456-s GNN5-RO          p256 --  1.00213 -0.00129  0.00002\n013\n13 dyno-dxu-state-456-s GNN5-RO          p350 s0  1.00300\n13 dyno-dxu-state-456-s GNN5-RO          p350 s1  1.00308\n13 dyno-dxu-state-456-s GNN5-RO          p350 s2  1.00316\n13 dyno-dxu-state-456-s GNN5-RO          p350 s3  1.00299\n13 dyno-dxu-state-456-s GNN5-RO          p350 --  1.00213 -0.00093  0.00007\n014\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s0  1.00250\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s1  1.00255\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s2  1.00268\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s3  1.00245\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 --  1.00158 -0.00096  0.00009\n015\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s0  1.00200\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s1  1.00209\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s2  1.00218\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s3  1.00194\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 --  1.00107 -0.00098  0.00009\n016\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s0  1.00572\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s1  1.00582\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s2  1.00562\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s3  1.00587\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 --  1.00374 -0.00202  0.00010\n017\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s0  1.00503\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s1  1.00487\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s2  1.00496\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s3  1.00504\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 --  1.00352 -0.00146  0.00007\n018\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s0  1.00432\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s1  1.00434\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s2  1.00439\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s3  1.00445\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 --  1.00337 -0.00101  0.00005\n019\n19 dyno-wnq-state-476-s GNN4             p200 s0  1.00640\n19 dyno-wnq-state-476-s GNN4             p200 s1  1.00645\n19 dyno-wnq-state-476-s GNN4             p200 s2  1.00646\n19 dyno-wnq-state-476-s GNN4             p200 s3  1.00660\n19 dyno-wnq-state-476-s GNN4             p200 --  1.00457 -0.00191  0.00007\n020\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s0  0.99903\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s1  0.99885\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s2  0.99879\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s3  0.99877\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 --  0.99553 -0.00333  0.00010\n021\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s0  0.99426\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s1  0.99452\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s2  0.99433\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s3  0.99448\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 --  0.99282 -0.00158  0.00011\n022\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s0  0.99456\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s1  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s2  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s3  0.99453\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 --  0.99396 -0.00060  0.00002\n023\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s0  0.99215\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s1  0.99226\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s2  0.99228\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s3  0.99238\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 --  0.99057 -0.00170  0.00008\n024\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s0  0.99058\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s1  0.99062\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s2  0.99072\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s3  0.99065\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 --  0.99003 -0.00061  0.00005\n025\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s0  0.99316\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s1  0.99324\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s2  0.99298\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s3  0.99322\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 --  0.99145 -0.00170  0.00010\n026\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s0  0.99138\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s1  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s2  0.99129\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s3  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 --  0.99050 -0.00087  0.00005\n027\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s0  0.99141\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s1  0.99143\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s2  0.99151\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s3  0.99137\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 --  0.99071 -0.00072  0.00005\n028\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s0  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s1  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s2  0.98837\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s3  0.98838\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 --  0.98771 -0.00063  0.00003\n029\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s0  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s1  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s2  0.98784\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s3  0.98781\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 --  0.98717 -0.00063  0.00003\n030\n30 goo-uel-state-652-s9 GOO-200          p200 s0  0.99403\n30 goo-uel-state-652-s9 GOO-200          p200 s1  0.99415\n30 goo-uel-state-652-s9 GOO-200          p200 s2  0.99398\n30 goo-uel-state-652-s9 GOO-200          p200 s3  0.99416\n30 goo-uel-state-652-s9 GOO-200          p200 --  0.99241 -0.00167  0.00008\n031\n31 goo-uel-state-652-s9 GOO-200          p350 s0  0.99091\n31 goo-uel-state-652-s9 GOO-200          p350 s1  0.99087\n31 goo-uel-state-652-s9 GOO-200          p350 s2  0.99092\n31 goo-uel-state-652-s9 GOO-200          p350 s3  0.99093\n31 goo-uel-state-652-s9 GOO-200          p350 --  0.99013 -0.00078  0.00002\n032\n32 goo-wqz-state-664-s9 GOO-384          p384 s0  0.99036\n32 goo-wqz-state-664-s9 GOO-384          p384 s1  0.99049\n32 goo-wqz-state-664-s9 GOO-384          p384 s2  0.99043\n32 goo-wqz-state-664-s9 GOO-384          p384 s3  0.99039\n32 goo-wqz-state-664-s9 GOO-384          p384 --  0.98976 -0.00066  0.00005\n033\n33 dyno-fdk-state-396-s GNN6-200         p200 s0  1.01130\n33 dyno-fdk-state-396-s GNN6-200         p200 s1  1.01137\n33 dyno-fdk-state-396-s GNN6-200         p200 s2  1.01136\n33 dyno-fdk-state-396-s GNN6-200         p200 s3  1.01143\n33 dyno-fdk-state-396-s GNN6-200         p200 --  1.00933 -0.00204  0.00005\n034\n34 dyno-rqy-state-408-s GNN6-350         p350 s0  1.00987\n34 dyno-rqy-state-408-s GNN6-350         p350 s1  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s2  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s3  1.01000\n34 dyno-rqy-state-408-s GNN6-350         p350 --  1.00892 -0.00104  0.00005\n035\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s0  0.98744\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s1  0.98755\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s2  0.98756\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s3  0.98758\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 --  0.98687 -0.00066  0.00006\n036\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s0  0.98654\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s1  0.98662\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s2  0.98664\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s3  0.98666\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 --  0.98595 -0.00066  0.00005\n\nTotal models present 37\n</code></pre>\n<h1>Combining the Models</h1>\n<p>I used a linear combination of a subset of these models.  For a baseline, I simple calculated the optimal weights on all 37 models using scipy.optimization.minimize with the Powell method.  Using least squares gives about 20 bps worse score because the true metric is the mean angular error.  The baseline score was .97948 on Batch 601 (used for weight calculation) and .97964 on Batch 602 (for validation).  This is a 63bp improvement over the best single model.  Since I didn't want to  calculate all 37 models for submission, I choose a subset of models.  You cannot simply use the models with the highest weights to pick a subset because the models are correlated.  For example, if model A and model B are the best and yet effectively the same, then their weights might be approximately equal and the highest of the bunch.  A subset should only include one of them.  So I took an iterative approach, removing a single model at a time and removing the one that had the least negative impact.  These are the results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651278%2Fa906b327fa59d5259a491c6098ce82eb%2Fscore_by_models.png?generation=1682004978180418&amp;alt=media\" alt=\"\"></p>\n<p>You can see that too many models is bad on cross-validation due to overfitting.  I choose 6 models as a compromise of inference time and best score.</p>\n<h1>Lessons Learned</h1>\n<ol>\n<li>Learn about Transformers!</li>\n<li>Have a better framework in place to evaluate different models and parameters without pulling my hair out when changing stuff.</li>\n<li>Spend less time on searching for small tweaks when there is a large gap to the leaders</li>\n<li>To save time, do more fine tuning of models when adding things rather than training from scratch</li>\n</ol>\n<h1>Things that didn't work for me</h1>\n<ol>\n<li>I could not get vonMises-Fisher loss to work</li>\n<li>Adding additional features, like number of neighbors, sensor type, line-fit estimated direction </li>\n<li>Dropout or weight decay.  My models tended not to overfit, probably because they were not complex enough</li>\n<li>Pulse selection.  I could never improve on random selection, which still mystifies me</li>\n<li>More complex ensembles.  I see <a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> did well with this, but I could not get XGBOOST or a neural network to combine predictions or even the embeddings into anything better than weighted averaging.</li>\n</ol>\n<h1>Conclusion</h1>\n<p>Congratulations to all the top finishers!  And a big thank you to all the people posting information to help everyone, to everyone for participating, and to the contest organizers for putting on such a great competition.</p>",
  "messages": [
    {
      "id": "2228540",
      "postDate": "04/20/2023 15:57:33",
      "content": "<h1>Summary</h1>\n<p>My final solution used an ensemble with a simple weighted average of 6 models that were generalizations of publicly posted models: (1) the DynEdge model <a href=\"https://www.kaggle.com/rasmusrse\" target=\"_blank\">@rasmusrse</a> , <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> , and (2) LSTM model by <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>.  Each model predicted directly a 3-dimensional vector of the target direction, using a loss that was the norm of the difference between the predicted vector and the target vector.  The predicted vector was not normalized first.</p>\n<h1>TLDR:</h1>\n<ol>\n<li>Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.</li>\n<li>Pool of all models included 30 different models, best combination of 6 selected through optimization method.<br>\nModels included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks</li>\n<li>Models included different number of pulses selected for training and different number of pulses used for inference (not always the same). Pulses were selected randomly from entire set of pulses.</li>\n<li>Each model was used (effectively) 4 times and averaged (because of the random selection of pulse set) - a simple observation was that smaller events did not require multiple evaluations.</li>\n<li>For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling). Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.</li>\n<li>Each model used as loss the norm of predicted vector minus true unit direction vector.</li>\n</ol>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">Robin Smits LSTM models</a><br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">RASMUS ØRSØE DynEdge Post</a><br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">My Derivation of Least Squares Fit</a><br>\n<a href=\"https://www.kaggle.com/code/solverworld/12th-place-ensemble-models\" target=\"_blank\">My Ensembling Notebook</a> [The attached dataset contains the pytorch models]<br>\nMirco Hunnefield Masters Thesis, Online Reconstruction of Muon-Neutrino Events in IceCube using Deep Learning Techniques<br>\nKai Schatto, PhD Thesis, Stacked searches for high-energy neutrinos from blazars with IceCube</p>\n<h1>Early Days</h1>\n<p>I started by reading as many of the neutrino detection papers as I could, trying to understand the nature of the problem.  There were many excellent posts with graphics (taken from the papers) showing how neutrinos generated a muon that then formed spinoff collisions that illuminated a large set of sensors embedded in the Antarctic Ice.  I won't repeat all that excellent information, but I will try to outline my thought process as my solution evolved.</p>\n<p>My first submissions were the (<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">line-fit solutions</a> that got 1.214 and later improved to 1.174 with more clever picking of pulses through neighbor counting (thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> !).</p>\n<p>I then decided to move on to neural net models.  Since there was a lot of excitement around the DynEdge models, I though maybe I should zig while others were zagging, and looked into other graph neural networks.  I implemented PointNet and then PointNet++, which are basically DynEdge models where you do not dynamically alter the graph at each iteration.  That is, the graph is fixed by neighbors within a certain radius, and those neighbors are then operated on much as a Convolutional NN (CNN) does for the entire solution.  The geometric pytorch  <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/cheatsheet/gnn_cheatsheet.html\" target=\"_blank\">cheatsheet</a> gives a nice summary of all the GNN models it supports with linked papers.  </p>\n<p>After looking at all the GNN models, and failing to have PointNet++ do much, I thought I should dive into the DynEdge models.  Trying to install the graphnet module required backing up too many python modules (and using Python 3.7 instead of my local 3.10).  So, since I hate to go backwards, I pulled out the dynedge code from the entire graphnet module, and used py-geometric to build my DynEdge models.  This also freed me from the SQlite and other framework difficulties so I could build my own pipeline.</p>\n<p>I solved the data size problem by preprocessing each parquet batch file into 780 individual hdf5 files which contained 256x2000xK matrices which were the data for 256 events.  The pulses were limited to 2000 by Aux and Charge sorting.  The rest were zero-padded, but because hdf5 compresses the data, this did not result in larger storage.  The K features were x,y,z, time, charge, aux, num-of-neighbors, and some pre-computed time window information.</p>\n<p>I then was able to shuffle both batches and the 256-groups of events within each batch without massive amounts of memory.  I was able to feed my GPU efficiently using the preprocessed files.</p>\n<h1>Progression</h1>\n<p>I decided to use batches 1-100 for training, and batch 601 and 602 for validation.  I typically trained for 600-800 batches, thus used each batch only 6-8 times.  I see now from the top finishers that people trained on much more data and for longer.  My models were probably not complex enough to need more data.</p>\n<p>I got the DynEdge models working close to the reported public notebooks, and eventually tried a few variations, including adding more layers (up to 6) and a larger output head.  The default was just a 128 size FC layer, but I eventually used [384,256,128,64,3].  My best DynEdge models got around 1.001 on Batch 602 (my test batch).</p>\n<h1>Choice of Output</h1>\n<p>I tried outputting bins (thanks again, <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>) but I found better results with directly predicting the 3-dimension direction and using the norm ||v-pred|| as the loss function.  The gradient of this function for v close to to pred looks much like the Angular Error function, so I thought it should work.  I could not get the Angular Error to work directly as a loss function, because it does not control for the vector size and allows the predicted vector to shrink to zero size.  There also some inaccuracies involved in making sure you don't normalize a length 0 vector.</p>\n<h1>Choice of Input</h1>\n<p>I tried various ways of selecting the input pulses to use in my models.  Every sorting or choosing method seemed to work worse than random selection.  Because of the random selection nature, model accuracy could be improved by simply averaging 4 predictions <em>from the same model</em>.  One simple trick was to realize that events with less pulses than were required for a model did not have a random element and thus those did not need to be repeated 4 times.  This meant I could do inference 4x with only a 15% increase in computation effort on submissions.</p>\n<h1>Back to LSTM Models</h1>\n<p>I decided to go back the LSTM models and taking a hint from the 2nd place finish of the RSNA Intracranial Hemorrhage challenge ( <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> )  <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/117228\" target=\"_blank\">2nd Place</a> I used multiple LSTM layers, with the outputs of each layer concatenated together into a final multi-layer fully connected prediction head.<br>\nThe best models were 512 wide LSTMs in 3 layers.  I used prediction heads with layers [256,64,64,3] and variations thereof.  I also tried an Attention modification on the LSTM output which helped a little bit (this was simply a learned weight that is dotted with the LSTM output to produce a signal proportional the weight that should be given to that element in the sequence).</p>\n<h1>Trying all the Models</h1>\n<p>Here are the scores of 35 models, just to show the variation and improvement from averaging.  The s&lt;i&gt; indicates a different random selection of the pulses from each event, the p&lt;n&gt; indicates how many pulses were used in the inference (prediction) and the last two columns indicate the change in score by using an average of the 4 predictions and the standard deviation, respectively.  The score was done on Batch 602.<br>\nYou can see the average helps about 35 bps on models using 128 pulses, about 6-10 bps for models using 350 or 384 pulses.  The description includes the type of model as well as the number of pulses used for training.  E.g. LSTM512x3-384 was a 3 layer LSTM model trained on 384 pulses.  ATT means attention was used.<br>\nThe best GNN model gets 1.001, the best LSTM model gets 0.9859 on batch 602.  My scores on the leaderboard were about .002 lower than my cross validation scores on batch 602.  I did not submit any individual models to the leaderboard, as I did want to contaminate my procedure with fitting to the leaderboard.</p>\n<pre><code>000\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s0  0.98869\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s1  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s2  0.98886\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s3  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 --  0.98810 -0.00069  0.00006\n001\n01 dog-ust-state-752-s9 LSTM512x3        p128 s0  0.99525\n01 dog-ust-state-752-s9 LSTM512x3        p128 s1  0.99547\n01 dog-ust-state-752-s9 LSTM512x3        p128 s2  0.99538\n01 dog-ust-state-752-s9 LSTM512x3        p128 s3  0.99533\n01 dog-ust-state-752-s9 LSTM512x3        p128 --  0.99184 -0.00352  0.00008\n002\n02 dog-ust-state-752-s9 LSTM512x3        p200 s0  0.99056\n02 dog-ust-state-752-s9 LSTM512x3        p200 s1  0.99073\n02 dog-ust-state-752-s9 LSTM512x3        p200 s2  0.99069\n02 dog-ust-state-752-s9 LSTM512x3        p200 s3  0.99061\n02 dog-ust-state-752-s9 LSTM512x3        p200 --  0.98900 -0.00165  0.00007\n003\n03 dog-ust-state-752-s9 LSTM512x3        p256 s0  0.98949\n03 dog-ust-state-752-s9 LSTM512x3        p256 s1  0.98929\n03 dog-ust-state-752-s9 LSTM512x3        p256 s2  0.98943\n03 dog-ust-state-752-s9 LSTM512x3        p256 s3  0.98945\n03 dog-ust-state-752-s9 LSTM512x3        p256 --  0.98831 -0.00110  0.00008\n004\n04 dog-ust-state-752-s9 LSTM512x3        p320 s0  0.98896\n04 dog-ust-state-752-s9 LSTM512x3        p320 s1  0.98897\n04 dog-ust-state-752-s9 LSTM512x3        p320 s2  0.98885\n04 dog-ust-state-752-s9 LSTM512x3        p320 s3  0.98891\n04 dog-ust-state-752-s9 LSTM512x3        p320 --  0.98813 -0.00079  0.00005\n005\n05 dog-ust-state-752-s9 LSTM512x3        p384 s0  0.98868\n05 dog-ust-state-752-s9 LSTM512x3        p384 s1  0.98881\n05 dog-ust-state-752-s9 LSTM512x3        p384 s2  0.98874\n05 dog-ust-state-752-s9 LSTM512x3        p384 s3  0.98872\n05 dog-ust-state-752-s9 LSTM512x3        p384 --  0.98815 -0.00058  0.00005\n006\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s0  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s1  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s2  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s3  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 --  0.99047  0.00000  0.00000\n007\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s0  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s1  0.99138\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s2  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s3  0.99140\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 --  0.98955 -0.00175  0.00010\n008\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s0  0.98947\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s1  0.98931\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s2  0.98939\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s3  0.98936\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 --  0.98819 -0.00119  0.00006\n009\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s0  0.98846\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s1  0.98843\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s2  0.98848\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s3  0.98835\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 --  0.98756 -0.00087  0.00005\n010\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s0  0.98793\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s1  0.98807\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s2  0.98810\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s3  0.98798\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 --  0.98741 -0.00061  0.00007\n011\n11 dyno-dxu-state-456-s GNN5-RO          p200 s0  1.00377\n11 dyno-dxu-state-456-s GNN5-RO          p200 s1  1.00385\n11 dyno-dxu-state-456-s GNN5-RO          p200 s2  1.00370\n11 dyno-dxu-state-456-s GNN5-RO          p200 s3  1.00397\n11 dyno-dxu-state-456-s GNN5-RO          p200 --  1.00193 -0.00189  0.00010\n012\n12 dyno-dxu-state-456-s GNN5-RO          p256 s0  1.00346\n12 dyno-dxu-state-456-s GNN5-RO          p256 s1  1.00343\n12 dyno-dxu-state-456-s GNN5-RO          p256 s2  1.00341\n12 dyno-dxu-state-456-s GNN5-RO          p256 s3  1.00340\n12 dyno-dxu-state-456-s GNN5-RO          p256 --  1.00213 -0.00129  0.00002\n013\n13 dyno-dxu-state-456-s GNN5-RO          p350 s0  1.00300\n13 dyno-dxu-state-456-s GNN5-RO          p350 s1  1.00308\n13 dyno-dxu-state-456-s GNN5-RO          p350 s2  1.00316\n13 dyno-dxu-state-456-s GNN5-RO          p350 s3  1.00299\n13 dyno-dxu-state-456-s GNN5-RO          p350 --  1.00213 -0.00093  0.00007\n014\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s0  1.00250\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s1  1.00255\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s2  1.00268\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s3  1.00245\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 --  1.00158 -0.00096  0.00009\n015\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s0  1.00200\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s1  1.00209\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s2  1.00218\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s3  1.00194\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 --  1.00107 -0.00098  0.00009\n016\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s0  1.00572\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s1  1.00582\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s2  1.00562\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s3  1.00587\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 --  1.00374 -0.00202  0.00010\n017\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s0  1.00503\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s1  1.00487\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s2  1.00496\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s3  1.00504\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 --  1.00352 -0.00146  0.00007\n018\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s0  1.00432\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s1  1.00434\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s2  1.00439\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s3  1.00445\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 --  1.00337 -0.00101  0.00005\n019\n19 dyno-wnq-state-476-s GNN4             p200 s0  1.00640\n19 dyno-wnq-state-476-s GNN4             p200 s1  1.00645\n19 dyno-wnq-state-476-s GNN4             p200 s2  1.00646\n19 dyno-wnq-state-476-s GNN4             p200 s3  1.00660\n19 dyno-wnq-state-476-s GNN4             p200 --  1.00457 -0.00191  0.00007\n020\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s0  0.99903\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s1  0.99885\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s2  0.99879\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s3  0.99877\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 --  0.99553 -0.00333  0.00010\n021\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s0  0.99426\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s1  0.99452\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s2  0.99433\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s3  0.99448\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 --  0.99282 -0.00158  0.00011\n022\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s0  0.99456\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s1  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s2  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s3  0.99453\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 --  0.99396 -0.00060  0.00002\n023\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s0  0.99215\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s1  0.99226\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s2  0.99228\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s3  0.99238\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 --  0.99057 -0.00170  0.00008\n024\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s0  0.99058\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s1  0.99062\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s2  0.99072\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s3  0.99065\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 --  0.99003 -0.00061  0.00005\n025\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s0  0.99316\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s1  0.99324\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s2  0.99298\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s3  0.99322\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 --  0.99145 -0.00170  0.00010\n026\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s0  0.99138\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s1  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s2  0.99129\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s3  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 --  0.99050 -0.00087  0.00005\n027\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s0  0.99141\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s1  0.99143\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s2  0.99151\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s3  0.99137\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 --  0.99071 -0.00072  0.00005\n028\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s0  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s1  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s2  0.98837\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s3  0.98838\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 --  0.98771 -0.00063  0.00003\n029\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s0  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s1  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s2  0.98784\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s3  0.98781\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 --  0.98717 -0.00063  0.00003\n030\n30 goo-uel-state-652-s9 GOO-200          p200 s0  0.99403\n30 goo-uel-state-652-s9 GOO-200          p200 s1  0.99415\n30 goo-uel-state-652-s9 GOO-200          p200 s2  0.99398\n30 goo-uel-state-652-s9 GOO-200          p200 s3  0.99416\n30 goo-uel-state-652-s9 GOO-200          p200 --  0.99241 -0.00167  0.00008\n031\n31 goo-uel-state-652-s9 GOO-200          p350 s0  0.99091\n31 goo-uel-state-652-s9 GOO-200          p350 s1  0.99087\n31 goo-uel-state-652-s9 GOO-200          p350 s2  0.99092\n31 goo-uel-state-652-s9 GOO-200          p350 s3  0.99093\n31 goo-uel-state-652-s9 GOO-200          p350 --  0.99013 -0.00078  0.00002\n032\n32 goo-wqz-state-664-s9 GOO-384          p384 s0  0.99036\n32 goo-wqz-state-664-s9 GOO-384          p384 s1  0.99049\n32 goo-wqz-state-664-s9 GOO-384          p384 s2  0.99043\n32 goo-wqz-state-664-s9 GOO-384          p384 s3  0.99039\n32 goo-wqz-state-664-s9 GOO-384          p384 --  0.98976 -0.00066  0.00005\n033\n33 dyno-fdk-state-396-s GNN6-200         p200 s0  1.01130\n33 dyno-fdk-state-396-s GNN6-200         p200 s1  1.01137\n33 dyno-fdk-state-396-s GNN6-200         p200 s2  1.01136\n33 dyno-fdk-state-396-s GNN6-200         p200 s3  1.01143\n33 dyno-fdk-state-396-s GNN6-200         p200 --  1.00933 -0.00204  0.00005\n034\n34 dyno-rqy-state-408-s GNN6-350         p350 s0  1.00987\n34 dyno-rqy-state-408-s GNN6-350         p350 s1  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s2  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s3  1.01000\n34 dyno-rqy-state-408-s GNN6-350         p350 --  1.00892 -0.00104  0.00005\n035\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s0  0.98744\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s1  0.98755\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s2  0.98756\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s3  0.98758\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 --  0.98687 -0.00066  0.00006\n036\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s0  0.98654\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s1  0.98662\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s2  0.98664\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s3  0.98666\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 --  0.98595 -0.00066  0.00005\n\nTotal models present 37\n</code></pre>\n<h1>Combining the Models</h1>\n<p>I used a linear combination of a subset of these models.  For a baseline, I simple calculated the optimal weights on all 37 models using scipy.optimization.minimize with the Powell method.  Using least squares gives about 20 bps worse score because the true metric is the mean angular error.  The baseline score was .97948 on Batch 601 (used for weight calculation) and .97964 on Batch 602 (for validation).  This is a 63bp improvement over the best single model.  Since I didn't want to  calculate all 37 models for submission, I choose a subset of models.  You cannot simply use the models with the highest weights to pick a subset because the models are correlated.  For example, if model A and model B are the best and yet effectively the same, then their weights might be approximately equal and the highest of the bunch.  A subset should only include one of them.  So I took an iterative approach, removing a single model at a time and removing the one that had the least negative impact.  These are the results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651278%2Fa906b327fa59d5259a491c6098ce82eb%2Fscore_by_models.png?generation=1682004978180418&amp;alt=media\" alt=\"\"></p>\n<p>You can see that too many models is bad on cross-validation due to overfitting.  I choose 6 models as a compromise of inference time and best score.</p>\n<h1>Lessons Learned</h1>\n<ol>\n<li>Learn about Transformers!</li>\n<li>Have a better framework in place to evaluate different models and parameters without pulling my hair out when changing stuff.</li>\n<li>Spend less time on searching for small tweaks when there is a large gap to the leaders</li>\n<li>To save time, do more fine tuning of models when adding things rather than training from scratch</li>\n</ol>\n<h1>Things that didn't work for me</h1>\n<ol>\n<li>I could not get vonMises-Fisher loss to work</li>\n<li>Adding additional features, like number of neighbors, sensor type, line-fit estimated direction </li>\n<li>Dropout or weight decay.  My models tended not to overfit, probably because they were not complex enough</li>\n<li>Pulse selection.  I could never improve on random selection, which still mystifies me</li>\n<li>More complex ensembles.  I see <a href=\"https://www.kaggle.com/dipamc77\" target=\"_blank\">@dipamc77</a> did well with this, but I could not get XGBOOST or a neural network to combine predictions or even the embeddings into anything better than weighted averaging.</li>\n</ol>\n<h1>Conclusion</h1>\n<p>Congratulations to all the top finishers!  And a big thank you to all the people posting information to help everyone, to everyone for participating, and to the contest organizers for putting on such a great competition.</p>",
      "rawMarkdown": "#Summary\nMy final solution used an ensemble with a simple weighted average of 6 models that were generalizations of publicly posted models: (1) the DynEdge model @rasmusrse , @anjum48 , and (2) LSTM model by @rsmits.  Each model predicted directly a 3-dimensional vector of the target direction, using a loss that was the norm of the difference between the predicted vector and the target vector.  The predicted vector was not normalized first.\n\n#TLDR:\n1. Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.\n1. Pool of all models included 30 different models, best combination of 6 selected through optimization method.\nModels included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks\n1. Models included different number of pulses selected for training and different number of pulses used for inference (not always the same). Pulses were selected randomly from entire set of pulses.\n1. Each model was used (effectively) 4 times and averaged (because of the random selection of pulse set) - a simple observation was that smaller events did not require multiple evaluations.\n1. For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling). Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.\n1. Each model used as loss the norm of predicted vector minus true unit direction vector.\n\n#References\n[Robin Smits LSTM models](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu)\n[RASMUS ØRSØE DynEdge Post](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524)\n[My Derivation of Least Squares Fit](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747)\n[My Ensembling Notebook](https://www.kaggle.com/code/solverworld/12th-place-ensemble-models) [The attached dataset contains the pytorch models]\nMirco Hunnefield Masters Thesis, Online Reconstruction of Muon-Neutrino Events in IceCube using Deep Learning Techniques\nKai Schatto, PhD Thesis, Stacked searches for high-energy neutrinos from blazars with IceCube\n#Early Days\nI started by reading as many of the neutrino detection papers as I could, trying to understand the nature of the problem.  There were many excellent posts with graphics (taken from the papers) showing how neutrinos generated a muon that then formed spinoff collisions that illuminated a large set of sensors embedded in the Antarctic Ice.  I won't repeat all that excellent information, but I will try to outline my thought process as my solution evolved.\n\nMy first submissions were the ([line-fit solutions](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747) that got 1.214 and later improved to 1.174 with more clever picking of pulses through neighbor counting (thanks @roberthatch !).\n\nI then decided to move on to neural net models.  Since there was a lot of excitement around the DynEdge models, I though maybe I should zig while others were zagging, and looked into other graph neural networks.  I implemented PointNet and then PointNet++, which are basically DynEdge models where you do not dynamically alter the graph at each iteration.  That is, the graph is fixed by neighbors within a certain radius, and those neighbors are then operated on much as a Convolutional NN (CNN) does for the entire solution.  The geometric pytorch  [cheatsheet](https://pytorch-geometric.readthedocs.io/en/latest/cheatsheet/gnn_cheatsheet.html) gives a nice summary of all the GNN models it supports with linked papers.  \n\nAfter looking at all the GNN models, and failing to have PointNet++ do much, I thought I should dive into the DynEdge models.  Trying to install the graphnet module required backing up too many python modules (and using Python 3.7 instead of my local 3.10).  So, since I hate to go backwards, I pulled out the dynedge code from the entire graphnet module, and used py-geometric to build my DynEdge models.  This also freed me from the SQlite and other framework difficulties so I could build my own pipeline.\n\nI solved the data size problem by preprocessing each parquet batch file into 780 individual hdf5 files which contained 256x2000xK matrices which were the data for 256 events.  The pulses were limited to 2000 by Aux and Charge sorting.  The rest were zero-padded, but because hdf5 compresses the data, this did not result in larger storage.  The K features were x,y,z, time, charge, aux, num-of-neighbors, and some pre-computed time window information.\n\nI then was able to shuffle both batches and the 256-groups of events within each batch without massive amounts of memory.  I was able to feed my GPU efficiently using the preprocessed files.\n\n#Progression\nI decided to use batches 1-100 for training, and batch 601 and 602 for validation.  I typically trained for 600-800 batches, thus used each batch only 6-8 times.  I see now from the top finishers that people trained on much more data and for longer.  My models were probably not complex enough to need more data.\n\nI got the DynEdge models working close to the reported public notebooks, and eventually tried a few variations, including adding more layers (up to 6) and a larger output head.  The default was just a 128 size FC layer, but I eventually used [384,256,128,64,3].  My best DynEdge models got around 1.001 on Batch 602 (my test batch).\n\n#Choice of Output\nI tried outputting bins (thanks again, @rsmits) but I found better results with directly predicting the 3-dimension direction and using the norm ||v-pred|| as the loss function.  The gradient of this function for v close to to pred looks much like the Angular Error function, so I thought it should work.  I could not get the Angular Error to work directly as a loss function, because it does not control for the vector size and allows the predicted vector to shrink to zero size.  There also some inaccuracies involved in making sure you don't normalize a length 0 vector.\n\n#Choice of Input\nI tried various ways of selecting the input pulses to use in my models.  Every sorting or choosing method seemed to work worse than random selection.  Because of the random selection nature, model accuracy could be improved by simply averaging 4 predictions *from the same model*.  One simple trick was to realize that events with less pulses than were required for a model did not have a random element and thus those did not need to be repeated 4 times.  This meant I could do inference 4x with only a 15% increase in computation effort on submissions.\n\n#Back to LSTM Models\nI decided to go back the LSTM models and taking a hint from the 2nd place finish of the RSNA Intracranial Hemorrhage challenge ( @darraghdog )  [2nd Place](https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/117228) I used multiple LSTM layers, with the outputs of each layer concatenated together into a final multi-layer fully connected prediction head.\nThe best models were 512 wide LSTMs in 3 layers.  I used prediction heads with layers [256,64,64,3] and variations thereof.  I also tried an Attention modification on the LSTM output which helped a little bit (this was simply a learned weight that is dotted with the LSTM output to produce a signal proportional the weight that should be given to that element in the sequence).\n\n#Trying all the Models\nHere are the scores of 35 models, just to show the variation and improvement from averaging.  The s<i\\> indicates a different random selection of the pulses from each event, the p<n\\> indicates how many pulses were used in the inference (prediction) and the last two columns indicate the change in score by using an average of the 4 predictions and the standard deviation, respectively.  The score was done on Batch 602.\nYou can see the average helps about 35 bps on models using 128 pulses, about 6-10 bps for models using 350 or 384 pulses.  The description includes the type of model as well as the number of pulses used for training.  E.g. LSTM512x3-384 was a 3 layer LSTM model trained on 384 pulses.  ATT means attention was used.\nThe best GNN model gets 1.001, the best LSTM model gets 0.9859 on batch 602.  My scores on the leaderboard were about .002 lower than my cross validation scores on batch 602.  I did not submit any individual models to the leaderboard, as I did want to contaminate my procedure with fitting to the leaderboard.\n```\n000\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s0  0.98869\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s1  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s2  0.98886\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s3  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 --  0.98810 -0.00069  0.00006\n001\n01 dog-ust-state-752-s9 LSTM512x3        p128 s0  0.99525\n01 dog-ust-state-752-s9 LSTM512x3        p128 s1  0.99547\n01 dog-ust-state-752-s9 LSTM512x3        p128 s2  0.99538\n01 dog-ust-state-752-s9 LSTM512x3        p128 s3  0.99533\n01 dog-ust-state-752-s9 LSTM512x3        p128 --  0.99184 -0.00352  0.00008\n002\n02 dog-ust-state-752-s9 LSTM512x3        p200 s0  0.99056\n02 dog-ust-state-752-s9 LSTM512x3        p200 s1  0.99073\n02 dog-ust-state-752-s9 LSTM512x3        p200 s2  0.99069\n02 dog-ust-state-752-s9 LSTM512x3        p200 s3  0.99061\n02 dog-ust-state-752-s9 LSTM512x3        p200 --  0.98900 -0.00165  0.00007\n003\n03 dog-ust-state-752-s9 LSTM512x3        p256 s0  0.98949\n03 dog-ust-state-752-s9 LSTM512x3        p256 s1  0.98929\n03 dog-ust-state-752-s9 LSTM512x3        p256 s2  0.98943\n03 dog-ust-state-752-s9 LSTM512x3        p256 s3  0.98945\n03 dog-ust-state-752-s9 LSTM512x3        p256 --  0.98831 -0.00110  0.00008\n004\n04 dog-ust-state-752-s9 LSTM512x3        p320 s0  0.98896\n04 dog-ust-state-752-s9 LSTM512x3        p320 s1  0.98897\n04 dog-ust-state-752-s9 LSTM512x3        p320 s2  0.98885\n04 dog-ust-state-752-s9 LSTM512x3        p320 s3  0.98891\n04 dog-ust-state-752-s9 LSTM512x3        p320 --  0.98813 -0.00079  0.00005\n005\n05 dog-ust-state-752-s9 LSTM512x3        p384 s0  0.98868\n05 dog-ust-state-752-s9 LSTM512x3        p384 s1  0.98881\n05 dog-ust-state-752-s9 LSTM512x3        p384 s2  0.98874\n05 dog-ust-state-752-s9 LSTM512x3        p384 s3  0.98872\n05 dog-ust-state-752-s9 LSTM512x3        p384 --  0.98815 -0.00058  0.00005\n006\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s0  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s1  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s2  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s3  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 --  0.99047  0.00000  0.00000\n007\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s0  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s1  0.99138\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s2  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s3  0.99140\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 --  0.98955 -0.00175  0.00010\n008\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s0  0.98947\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s1  0.98931\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s2  0.98939\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s3  0.98936\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 --  0.98819 -0.00119  0.00006\n009\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s0  0.98846\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s1  0.98843\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s2  0.98848\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s3  0.98835\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 --  0.98756 -0.00087  0.00005\n010\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s0  0.98793\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s1  0.98807\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s2  0.98810\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s3  0.98798\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 --  0.98741 -0.00061  0.00007\n011\n11 dyno-dxu-state-456-s GNN5-RO          p200 s0  1.00377\n11 dyno-dxu-state-456-s GNN5-RO          p200 s1  1.00385\n11 dyno-dxu-state-456-s GNN5-RO          p200 s2  1.00370\n11 dyno-dxu-state-456-s GNN5-RO          p200 s3  1.00397\n11 dyno-dxu-state-456-s GNN5-RO          p200 --  1.00193 -0.00189  0.00010\n012\n12 dyno-dxu-state-456-s GNN5-RO          p256 s0  1.00346\n12 dyno-dxu-state-456-s GNN5-RO          p256 s1  1.00343\n12 dyno-dxu-state-456-s GNN5-RO          p256 s2  1.00341\n12 dyno-dxu-state-456-s GNN5-RO          p256 s3  1.00340\n12 dyno-dxu-state-456-s GNN5-RO          p256 --  1.00213 -0.00129  0.00002\n013\n13 dyno-dxu-state-456-s GNN5-RO          p350 s0  1.00300\n13 dyno-dxu-state-456-s GNN5-RO          p350 s1  1.00308\n13 dyno-dxu-state-456-s GNN5-RO          p350 s2  1.00316\n13 dyno-dxu-state-456-s GNN5-RO          p350 s3  1.00299\n13 dyno-dxu-state-456-s GNN5-RO          p350 --  1.00213 -0.00093  0.00007\n014\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s0  1.00250\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s1  1.00255\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s2  1.00268\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s3  1.00245\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 --  1.00158 -0.00096  0.00009\n015\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s0  1.00200\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s1  1.00209\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s2  1.00218\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s3  1.00194\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 --  1.00107 -0.00098  0.00009\n016\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s0  1.00572\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s1  1.00582\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s2  1.00562\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s3  1.00587\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 --  1.00374 -0.00202  0.00010\n017\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s0  1.00503\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s1  1.00487\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s2  1.00496\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s3  1.00504\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 --  1.00352 -0.00146  0.00007\n018\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s0  1.00432\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s1  1.00434\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s2  1.00439\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s3  1.00445\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 --  1.00337 -0.00101  0.00005\n019\n19 dyno-wnq-state-476-s GNN4             p200 s0  1.00640\n19 dyno-wnq-state-476-s GNN4             p200 s1  1.00645\n19 dyno-wnq-state-476-s GNN4             p200 s2  1.00646\n19 dyno-wnq-state-476-s GNN4             p200 s3  1.00660\n19 dyno-wnq-state-476-s GNN4             p200 --  1.00457 -0.00191  0.00007\n020\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s0  0.99903\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s1  0.99885\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s2  0.99879\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s3  0.99877\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 --  0.99553 -0.00333  0.00010\n021\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s0  0.99426\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s1  0.99452\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s2  0.99433\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s3  0.99448\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 --  0.99282 -0.00158  0.00011\n022\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s0  0.99456\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s1  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s2  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s3  0.99453\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 --  0.99396 -0.00060  0.00002\n023\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s0  0.99215\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s1  0.99226\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s2  0.99228\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s3  0.99238\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 --  0.99057 -0.00170  0.00008\n024\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s0  0.99058\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s1  0.99062\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s2  0.99072\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s3  0.99065\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 --  0.99003 -0.00061  0.00005\n025\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s0  0.99316\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s1  0.99324\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s2  0.99298\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s3  0.99322\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 --  0.99145 -0.00170  0.00010\n026\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s0  0.99138\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s1  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s2  0.99129\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s3  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 --  0.99050 -0.00087  0.00005\n027\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s0  0.99141\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s1  0.99143\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s2  0.99151\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s3  0.99137\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 --  0.99071 -0.00072  0.00005\n028\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s0  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s1  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s2  0.98837\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s3  0.98838\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 --  0.98771 -0.00063  0.00003\n029\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s0  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s1  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s2  0.98784\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s3  0.98781\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 --  0.98717 -0.00063  0.00003\n030\n30 goo-uel-state-652-s9 GOO-200          p200 s0  0.99403\n30 goo-uel-state-652-s9 GOO-200          p200 s1  0.99415\n30 goo-uel-state-652-s9 GOO-200          p200 s2  0.99398\n30 goo-uel-state-652-s9 GOO-200          p200 s3  0.99416\n30 goo-uel-state-652-s9 GOO-200          p200 --  0.99241 -0.00167  0.00008\n031\n31 goo-uel-state-652-s9 GOO-200          p350 s0  0.99091\n31 goo-uel-state-652-s9 GOO-200          p350 s1  0.99087\n31 goo-uel-state-652-s9 GOO-200          p350 s2  0.99092\n31 goo-uel-state-652-s9 GOO-200          p350 s3  0.99093\n31 goo-uel-state-652-s9 GOO-200          p350 --  0.99013 -0.00078  0.00002\n032\n32 goo-wqz-state-664-s9 GOO-384          p384 s0  0.99036\n32 goo-wqz-state-664-s9 GOO-384          p384 s1  0.99049\n32 goo-wqz-state-664-s9 GOO-384          p384 s2  0.99043\n32 goo-wqz-state-664-s9 GOO-384          p384 s3  0.99039\n32 goo-wqz-state-664-s9 GOO-384          p384 --  0.98976 -0.00066  0.00005\n033\n33 dyno-fdk-state-396-s GNN6-200         p200 s0  1.01130\n33 dyno-fdk-state-396-s GNN6-200         p200 s1  1.01137\n33 dyno-fdk-state-396-s GNN6-200         p200 s2  1.01136\n33 dyno-fdk-state-396-s GNN6-200         p200 s3  1.01143\n33 dyno-fdk-state-396-s GNN6-200         p200 --  1.00933 -0.00204  0.00005\n034\n34 dyno-rqy-state-408-s GNN6-350         p350 s0  1.00987\n34 dyno-rqy-state-408-s GNN6-350         p350 s1  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s2  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s3  1.01000\n34 dyno-rqy-state-408-s GNN6-350         p350 --  1.00892 -0.00104  0.00005\n035\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s0  0.98744\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s1  0.98755\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s2  0.98756\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s3  0.98758\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 --  0.98687 -0.00066  0.00006\n036\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s0  0.98654\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s1  0.98662\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s2  0.98664\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s3  0.98666\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 --  0.98595 -0.00066  0.00005\n\nTotal models present 37\n\n```\n# Combining the Models\nI used a linear combination of a subset of these models.  For a baseline, I simple calculated the optimal weights on all 37 models using scipy.optimization.minimize with the Powell method.  Using least squares gives about 20 bps worse score because the true metric is the mean angular error.  The baseline score was .97948 on Batch 601 (used for weight calculation) and .97964 on Batch 602 (for validation).  This is a 63bp improvement over the best single model.  Since I didn't want to  calculate all 37 models for submission, I choose a subset of models.  You cannot simply use the models with the highest weights to pick a subset because the models are correlated.  For example, if model A and model B are the best and yet effectively the same, then their weights might be approximately equal and the highest of the bunch.  A subset should only include one of them.  So I took an iterative approach, removing a single model at a time and removing the one that had the least negative impact.  These are the results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651278%2Fa906b327fa59d5259a491c6098ce82eb%2Fscore_by_models.png?generation=1682004978180418&alt=media)\n\nYou can see that too many models is bad on cross-validation due to overfitting.  I choose 6 models as a compromise of inference time and best score.\n\n# Lessons Learned\n1. Learn about Transformers!\n1. Have a better framework in place to evaluate different models and parameters without pulling my hair out when changing stuff.\n1. Spend less time on searching for small tweaks when there is a large gap to the leaders\n1. To save time, do more fine tuning of models when adding things rather than training from scratch\n#Things that didn't work for me\n1. I could not get vonMises-Fisher loss to work\n1. Adding additional features, like number of neighbors, sensor type, line-fit estimated direction \n1. Dropout or weight decay.  My models tended not to overfit, probably because they were not complex enough\n1. Pulse selection.  I could never improve on random selection, which still mystifies me\n1. More complex ensembles.  I see @dipamc77 did well with this, but I could not get XGBOOST or a neural network to combine predictions or even the embeddings into anything better than weighted averaging.\n#Conclusion\nCongratulations to all the top finishers!  And a big thank you to all the people posting information to help everyone, to everyone for participating, and to the contest organizers for putting on such a great competition.",
      "votes": null
    },
    {
      "id": "2228614",
      "postDate": "04/20/2023 16:45:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> Really nice and interresting solutions. Congratulations! </p>\n<p>I fully recognize that more and longer LSTM layers can help significantly in increasing score.<br>\nI'am glad my work was usefull to you.</p>",
      "rawMarkdown": "Hi @solverworld Really nice and interresting solutions. Congratulations! \n\nI fully recognize that more and longer LSTM layers can help significantly in increasing score.\nI'am glad my work was usefull to you.",
      "votes": null
    },
    {
      "id": "2228620",
      "postDate": "04/20/2023 16:51:12",
      "content": "<p>Am I the only one who stuggled so much with vMF/vector training end to end? I must have been doing something seriously wrong. 😯</p>\n<p>I guess small unseen details somehow change results so much, xgboost ensemble worked just for me on the first shot. </p>",
      "rawMarkdown": "Am I the only one who stuggled so much with vMF/vector training end to end? I must have been doing something seriously wrong. 😯\n\nI guess small unseen details somehow change results so much, xgboost ensemble worked just for me on the first shot.",
      "votes": null
    },
    {
      "id": "2228644",
      "postDate": "04/20/2023 17:28:54",
      "content": "<p>Yes, many thanks.  I was surprised to see LSTM did better than DynEdge in the end.</p>",
      "rawMarkdown": "Yes, many thanks.  I was surprised to see LSTM did better than DynEdge in the end.",
      "votes": null
    },
    {
      "id": "2228646",
      "postDate": "04/20/2023 17:31:21",
      "content": "<p>vMF just didn't work for me.  By using norm of vector error, you get some benefit of vMF - when the model is totally uncertain, it picks small length prediction vector, because that minimizes ||v-pred|| over a (longer) wrong direction.  This helps in the ensembling, because those predictions get automatically weighted less.  That's something I forgot to mention in the writeup.  Normalizing the predictions before averaging gives a worse solution because of this effect.</p>",
      "rawMarkdown": "vMF just didn't work for me.  By using norm of vector error, you get some benefit of vMF - when the model is totally uncertain, it picks small length prediction vector, because that minimizes ||v-pred|| over a (longer) wrong direction.  This helps in the ensembling, because those predictions get automatically weighted less.  That's something I forgot to mention in the writeup.  Normalizing the predictions before averaging gives a worse solution because of this effect.",
      "votes": null
    },
    {
      "id": "2229201",
      "postDate": "04/21/2023 06:42:01",
      "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> thanks for a topic with cool details. It is very important to share experiences with the community. I read the topic and get new knowledge!</p>",
      "rawMarkdown": "solverworld thanks for a topic with cool details. It is very important to share experiences with the community. I read the topic and get new knowledge!",
      "votes": null
    },
    {
      "id": "2231447",
      "postDate": "04/23/2023 10:14:33",
      "content": "<p>That's really shocking for me, too. I've tried to attach many things on LSTM… But it seems larger LSTM and attention was all we needed…</p>",
      "rawMarkdown": "That's really shocking for me, too. I've tried to attach many things on LSTM... But it seems larger LSTM and attention was all we needed...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228614,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/20/2023 16:45:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> Really nice and interresting solutions. Congratulations! </p>\n<p>I fully recognize that more and longer LSTM layers can help significantly in increasing score.<br>\nI'am glad my work was usefull to you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228644,
          "author_name": "solverworld",
          "author_url": "",
          "post_date": "04/20/2023 17:28:54",
          "content": "<p>Yes, many thanks.  I was surprised to see LSTM did better than DynEdge in the end.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2231447,
              "author_name": "seungmoklee",
              "author_url": "",
              "post_date": "04/23/2023 10:14:33",
              "content": "<p>That's really shocking for me, too. I've tried to attach many things on LSTM… But it seems larger LSTM and attention was all we needed…</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2228620,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "04/20/2023 16:51:12",
      "content": "<p>Am I the only one who stuggled so much with vMF/vector training end to end? I must have been doing something seriously wrong. 😯</p>\n<p>I guess small unseen details somehow change results so much, xgboost ensemble worked just for me on the first shot. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2228646,
          "author_name": "solverworld",
          "author_url": "",
          "post_date": "04/20/2023 17:31:21",
          "content": "<p>vMF just didn't work for me.  By using norm of vector error, you get some benefit of vMF - when the model is totally uncertain, it picks small length prediction vector, because that minimizes ||v-pred|| over a (longer) wrong direction.  This helps in the ensembling, because those predictions get automatically weighted less.  That's something I forgot to mention in the writeup.  Normalizing the predictions before averaging gives a worse solution because of this effect.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229201,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "04/21/2023 06:42:01",
      "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> thanks for a topic with cool details. It is very important to share experiences with the community. I read the topic and get new knowledge!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228540": "#Summary\nMy final solution used an ensemble with a simple weighted average of 6 models that were generalizations of publicly posted models: (1) the DynEdge model @rasmusrse , @anjum48 , and (2) LSTM model by @rsmits.  Each model predicted directly a 3-dimensional vector of the target direction, using a loss that was the norm of the difference between the predicted vector and the target vector.  The predicted vector was not normalized first.\n\n#TLDR:\n1. Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.\n1. Pool of all models included 30 different models, best combination of 6 selected through optimization method.\nModels included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks\n1. Models included different number of pulses selected for training and different number of pulses used for inference (not always the same). Pulses were selected randomly from entire set of pulses.\n1. Each model was used (effectively) 4 times and averaged (because of the random selection of pulse set) - a simple observation was that smaller events did not require multiple evaluations.\n1. For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling). Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.\n1. Each model used as loss the norm of predicted vector minus true unit direction vector.\n\n#References\n[Robin Smits LSTM models](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu)\n[RASMUS ØRSØE DynEdge Post](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524)\n[My Derivation of Least Squares Fit](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747)\n[My Ensembling Notebook](https://www.kaggle.com/code/solverworld/12th-place-ensemble-models) [The attached dataset contains the pytorch models]\nMirco Hunnefield Masters Thesis, Online Reconstruction of Muon-Neutrino Events in IceCube using Deep Learning Techniques\nKai Schatto, PhD Thesis, Stacked searches for high-energy neutrinos from blazars with IceCube\n#Early Days\nI started by reading as many of the neutrino detection papers as I could, trying to understand the nature of the problem.  There were many excellent posts with graphics (taken from the papers) showing how neutrinos generated a muon that then formed spinoff collisions that illuminated a large set of sensors embedded in the Antarctic Ice.  I won't repeat all that excellent information, but I will try to outline my thought process as my solution evolved.\n\nMy first submissions were the ([line-fit solutions](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747) that got 1.214 and later improved to 1.174 with more clever picking of pulses through neighbor counting (thanks @roberthatch !).\n\nI then decided to move on to neural net models.  Since there was a lot of excitement around the DynEdge models, I though maybe I should zig while others were zagging, and looked into other graph neural networks.  I implemented PointNet and then PointNet++, which are basically DynEdge models where you do not dynamically alter the graph at each iteration.  That is, the graph is fixed by neighbors within a certain radius, and those neighbors are then operated on much as a Convolutional NN (CNN) does for the entire solution.  The geometric pytorch  [cheatsheet](https://pytorch-geometric.readthedocs.io/en/latest/cheatsheet/gnn_cheatsheet.html) gives a nice summary of all the GNN models it supports with linked papers.  \n\nAfter looking at all the GNN models, and failing to have PointNet++ do much, I thought I should dive into the DynEdge models.  Trying to install the graphnet module required backing up too many python modules (and using Python 3.7 instead of my local 3.10).  So, since I hate to go backwards, I pulled out the dynedge code from the entire graphnet module, and used py-geometric to build my DynEdge models.  This also freed me from the SQlite and other framework difficulties so I could build my own pipeline.\n\nI solved the data size problem by preprocessing each parquet batch file into 780 individual hdf5 files which contained 256x2000xK matrices which were the data for 256 events.  The pulses were limited to 2000 by Aux and Charge sorting.  The rest were zero-padded, but because hdf5 compresses the data, this did not result in larger storage.  The K features were x,y,z, time, charge, aux, num-of-neighbors, and some pre-computed time window information.\n\nI then was able to shuffle both batches and the 256-groups of events within each batch without massive amounts of memory.  I was able to feed my GPU efficiently using the preprocessed files.\n\n#Progression\nI decided to use batches 1-100 for training, and batch 601 and 602 for validation.  I typically trained for 600-800 batches, thus used each batch only 6-8 times.  I see now from the top finishers that people trained on much more data and for longer.  My models were probably not complex enough to need more data.\n\nI got the DynEdge models working close to the reported public notebooks, and eventually tried a few variations, including adding more layers (up to 6) and a larger output head.  The default was just a 128 size FC layer, but I eventually used [384,256,128,64,3].  My best DynEdge models got around 1.001 on Batch 602 (my test batch).\n\n#Choice of Output\nI tried outputting bins (thanks again, @rsmits) but I found better results with directly predicting the 3-dimension direction and using the norm ||v-pred|| as the loss function.  The gradient of this function for v close to to pred looks much like the Angular Error function, so I thought it should work.  I could not get the Angular Error to work directly as a loss function, because it does not control for the vector size and allows the predicted vector to shrink to zero size.  There also some inaccuracies involved in making sure you don't normalize a length 0 vector.\n\n#Choice of Input\nI tried various ways of selecting the input pulses to use in my models.  Every sorting or choosing method seemed to work worse than random selection.  Because of the random selection nature, model accuracy could be improved by simply averaging 4 predictions *from the same model*.  One simple trick was to realize that events with less pulses than were required for a model did not have a random element and thus those did not need to be repeated 4 times.  This meant I could do inference 4x with only a 15% increase in computation effort on submissions.\n\n#Back to LSTM Models\nI decided to go back the LSTM models and taking a hint from the 2nd place finish of the RSNA Intracranial Hemorrhage challenge ( @darraghdog )  [2nd Place](https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/117228) I used multiple LSTM layers, with the outputs of each layer concatenated together into a final multi-layer fully connected prediction head.\nThe best models were 512 wide LSTMs in 3 layers.  I used prediction heads with layers [256,64,64,3] and variations thereof.  I also tried an Attention modification on the LSTM output which helped a little bit (this was simply a learned weight that is dotted with the LSTM output to produce a signal proportional the weight that should be given to that element in the sequence).\n\n#Trying all the Models\nHere are the scores of 35 models, just to show the variation and improvement from averaging.  The s<i\\> indicates a different random selection of the pulses from each event, the p<n\\> indicates how many pulses were used in the inference (prediction) and the last two columns indicate the change in score by using an average of the 4 predictions and the standard deviation, respectively.  The score was done on Batch 602.\nYou can see the average helps about 35 bps on models using 128 pulses, about 6-10 bps for models using 350 or 384 pulses.  The description includes the type of model as well as the number of pulses used for training.  E.g. LSTM512x3-384 was a 3 layer LSTM model trained on 384 pulses.  ATT means attention was used.\nThe best GNN model gets 1.001, the best LSTM model gets 0.9859 on batch 602.  My scores on the leaderboard were about .002 lower than my cross validation scores on batch 602.  I did not submit any individual models to the leaderboard, as I did want to contaminate my procedure with fitting to the leaderboard.\n```\n000\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s0  0.98869\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s1  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s2  0.98886\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 s3  0.98881\n00 dog-qnr-state-664-s9 LSTM512x3-384    p384 --  0.98810 -0.00069  0.00006\n001\n01 dog-ust-state-752-s9 LSTM512x3        p128 s0  0.99525\n01 dog-ust-state-752-s9 LSTM512x3        p128 s1  0.99547\n01 dog-ust-state-752-s9 LSTM512x3        p128 s2  0.99538\n01 dog-ust-state-752-s9 LSTM512x3        p128 s3  0.99533\n01 dog-ust-state-752-s9 LSTM512x3        p128 --  0.99184 -0.00352  0.00008\n002\n02 dog-ust-state-752-s9 LSTM512x3        p200 s0  0.99056\n02 dog-ust-state-752-s9 LSTM512x3        p200 s1  0.99073\n02 dog-ust-state-752-s9 LSTM512x3        p200 s2  0.99069\n02 dog-ust-state-752-s9 LSTM512x3        p200 s3  0.99061\n02 dog-ust-state-752-s9 LSTM512x3        p200 --  0.98900 -0.00165  0.00007\n003\n03 dog-ust-state-752-s9 LSTM512x3        p256 s0  0.98949\n03 dog-ust-state-752-s9 LSTM512x3        p256 s1  0.98929\n03 dog-ust-state-752-s9 LSTM512x3        p256 s2  0.98943\n03 dog-ust-state-752-s9 LSTM512x3        p256 s3  0.98945\n03 dog-ust-state-752-s9 LSTM512x3        p256 --  0.98831 -0.00110  0.00008\n004\n04 dog-ust-state-752-s9 LSTM512x3        p320 s0  0.98896\n04 dog-ust-state-752-s9 LSTM512x3        p320 s1  0.98897\n04 dog-ust-state-752-s9 LSTM512x3        p320 s2  0.98885\n04 dog-ust-state-752-s9 LSTM512x3        p320 s3  0.98891\n04 dog-ust-state-752-s9 LSTM512x3        p320 --  0.98813 -0.00079  0.00005\n005\n05 dog-ust-state-752-s9 LSTM512x3        p384 s0  0.98868\n05 dog-ust-state-752-s9 LSTM512x3        p384 s1  0.98881\n05 dog-ust-state-752-s9 LSTM512x3        p384 s2  0.98874\n05 dog-ust-state-752-s9 LSTM512x3        p384 s3  0.98872\n05 dog-ust-state-752-s9 LSTM512x3        p384 --  0.98815 -0.00058  0.00005\n006\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s0  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s1  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s2  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 s3  0.99047\n06 dog-uva-state-648-s9 LSTM512x3-200    p320 --  0.99047  0.00000  0.00000\n007\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s0  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s1  0.99138\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s2  0.99120\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 s3  0.99140\n07 dog-uva-state-768-s9 LSTM512x3-200    p200 --  0.98955 -0.00175  0.00010\n008\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s0  0.98947\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s1  0.98931\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s2  0.98939\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 s3  0.98936\n08 dog-uva-state-768-s9 LSTM512x3-200    p256 --  0.98819 -0.00119  0.00006\n009\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s0  0.98846\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s1  0.98843\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s2  0.98848\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 s3  0.98835\n09 dog-uva-state-768-s9 LSTM512x3-200    p320 --  0.98756 -0.00087  0.00005\n010\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s0  0.98793\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s1  0.98807\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s2  0.98810\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 s3  0.98798\n10 dog-uva-state-768-s9 LSTM512x3-200    p384 --  0.98741 -0.00061  0.00007\n011\n11 dyno-dxu-state-456-s GNN5-RO          p200 s0  1.00377\n11 dyno-dxu-state-456-s GNN5-RO          p200 s1  1.00385\n11 dyno-dxu-state-456-s GNN5-RO          p200 s2  1.00370\n11 dyno-dxu-state-456-s GNN5-RO          p200 s3  1.00397\n11 dyno-dxu-state-456-s GNN5-RO          p200 --  1.00193 -0.00189  0.00010\n012\n12 dyno-dxu-state-456-s GNN5-RO          p256 s0  1.00346\n12 dyno-dxu-state-456-s GNN5-RO          p256 s1  1.00343\n12 dyno-dxu-state-456-s GNN5-RO          p256 s2  1.00341\n12 dyno-dxu-state-456-s GNN5-RO          p256 s3  1.00340\n12 dyno-dxu-state-456-s GNN5-RO          p256 --  1.00213 -0.00129  0.00002\n013\n13 dyno-dxu-state-456-s GNN5-RO          p350 s0  1.00300\n13 dyno-dxu-state-456-s GNN5-RO          p350 s1  1.00308\n13 dyno-dxu-state-456-s GNN5-RO          p350 s2  1.00316\n13 dyno-dxu-state-456-s GNN5-RO          p350 s3  1.00299\n13 dyno-dxu-state-456-s GNN5-RO          p350 --  1.00213 -0.00093  0.00007\n014\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s0  1.00250\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s1  1.00255\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s2  1.00268\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 s3  1.00245\n14 dyno-kxc-state-524-s GNN5-RO-350      p350 --  1.00158 -0.00096  0.00009\n015\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s0  1.00200\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s1  1.00209\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s2  1.00218\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 s3  1.00194\n15 dyno-kxc-state-580-s GNN5-RO-350      p350 --  1.00107 -0.00098  0.00009\n016\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s0  1.00572\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s1  1.00582\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s2  1.00562\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 s3  1.00587\n16 dyno-rjl-state-708-s GNN4-RO-350      p200 --  1.00374 -0.00202  0.00010\n017\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s0  1.00503\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s1  1.00487\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s2  1.00496\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 s3  1.00504\n17 dyno-rjl-state-708-s GNN4-RO-350      p256 --  1.00352 -0.00146  0.00007\n018\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s0  1.00432\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s1  1.00434\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s2  1.00439\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 s3  1.00445\n18 dyno-rjl-state-708-s GNN4-RO-350      p350 --  1.00337 -0.00101  0.00005\n019\n19 dyno-wnq-state-476-s GNN4             p200 s0  1.00640\n19 dyno-wnq-state-476-s GNN4             p200 s1  1.00645\n19 dyno-wnq-state-476-s GNN4             p200 s2  1.00646\n19 dyno-wnq-state-476-s GNN4             p200 s3  1.00660\n19 dyno-wnq-state-476-s GNN4             p200 --  1.00457 -0.00191  0.00007\n020\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s0  0.99903\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s1  0.99885\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s2  0.99879\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 s3  0.99877\n20 egg-jhn-state-604-s9 LSTM384x3-ATT    p128 --  0.99553 -0.00333  0.00010\n021\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s0  0.99426\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s1  0.99452\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s2  0.99433\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 s3  0.99448\n21 egg-jhn-state-604-s9 LSTM384x3-ATT    p200 --  0.99282 -0.00158  0.00011\n022\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s0  0.99456\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s1  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s2  0.99457\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 s3  0.99453\n22 egg-jhn-state-604-s9 LSTM384x3-ATT    p350 --  0.99396 -0.00060  0.00002\n023\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s0  0.99215\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s1  0.99226\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s2  0.99228\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 s3  0.99238\n23 egg-jqn-state-572-s9 LSTM512x3-ATT200 p200 --  0.99057 -0.00170  0.00008\n024\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s0  0.99058\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s1  0.99062\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s2  0.99072\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 s3  0.99065\n24 egg-jqn-state-572-s9 LSTM512x3-ATT200 p384 --  0.99003 -0.00061  0.00005\n025\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s0  0.99316\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s1  0.99324\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s2  0.99298\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 s3  0.99322\n25 fir-irl-state-456-s9 LSTM512x3-200-mod p200 --  0.99145 -0.00170  0.00010\n026\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s0  0.99138\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s1  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s2  0.99129\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 s3  0.99142\n26 fir-irl-state-456-s9 LSTM512x3-200-mod p300 --  0.99050 -0.00087  0.00005\n027\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s0  0.99141\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s1  0.99143\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s2  0.99151\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 s3  0.99137\n27 fir-irl-state-456-s9 LSTM512x3-200-mod p350 --  0.99071 -0.00072  0.00005\n028\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s0  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s1  0.98831\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s2  0.98837\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 s3  0.98838\n28 egg-kvu-state-516-s9 LSTM512x3-ATT384 p384 --  0.98771 -0.00063  0.00003\n029\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s0  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s1  0.98777\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s2  0.98784\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 s3  0.98781\n29 egg-kvu-state-548-s9 LSTM512x3-ATT384 p384 --  0.98717 -0.00063  0.00003\n030\n30 goo-uel-state-652-s9 GOO-200          p200 s0  0.99403\n30 goo-uel-state-652-s9 GOO-200          p200 s1  0.99415\n30 goo-uel-state-652-s9 GOO-200          p200 s2  0.99398\n30 goo-uel-state-652-s9 GOO-200          p200 s3  0.99416\n30 goo-uel-state-652-s9 GOO-200          p200 --  0.99241 -0.00167  0.00008\n031\n31 goo-uel-state-652-s9 GOO-200          p350 s0  0.99091\n31 goo-uel-state-652-s9 GOO-200          p350 s1  0.99087\n31 goo-uel-state-652-s9 GOO-200          p350 s2  0.99092\n31 goo-uel-state-652-s9 GOO-200          p350 s3  0.99093\n31 goo-uel-state-652-s9 GOO-200          p350 --  0.99013 -0.00078  0.00002\n032\n32 goo-wqz-state-664-s9 GOO-384          p384 s0  0.99036\n32 goo-wqz-state-664-s9 GOO-384          p384 s1  0.99049\n32 goo-wqz-state-664-s9 GOO-384          p384 s2  0.99043\n32 goo-wqz-state-664-s9 GOO-384          p384 s3  0.99039\n32 goo-wqz-state-664-s9 GOO-384          p384 --  0.98976 -0.00066  0.00005\n033\n33 dyno-fdk-state-396-s GNN6-200         p200 s0  1.01130\n33 dyno-fdk-state-396-s GNN6-200         p200 s1  1.01137\n33 dyno-fdk-state-396-s GNN6-200         p200 s2  1.01136\n33 dyno-fdk-state-396-s GNN6-200         p200 s3  1.01143\n33 dyno-fdk-state-396-s GNN6-200         p200 --  1.00933 -0.00204  0.00005\n034\n34 dyno-rqy-state-408-s GNN6-350         p350 s0  1.00987\n34 dyno-rqy-state-408-s GNN6-350         p350 s1  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s2  1.00999\n34 dyno-rqy-state-408-s GNN6-350         p350 s3  1.01000\n34 dyno-rqy-state-408-s GNN6-350         p350 --  1.00892 -0.00104  0.00005\n035\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s0  0.98744\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s1  0.98755\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s2  0.98756\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 s3  0.98758\n35 egg-fsh-state-620-s9 LSTM512x3-ATT384 p384 --  0.98687 -0.00066  0.00006\n036\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s0  0.98654\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s1  0.98662\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s2  0.98664\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 s3  0.98666\n36 egg-fsh-state-640-s9 LSTM512x3-ATT384 p384 --  0.98595 -0.00066  0.00005\n\nTotal models present 37\n\n```\n# Combining the Models\nI used a linear combination of a subset of these models.  For a baseline, I simple calculated the optimal weights on all 37 models using scipy.optimization.minimize with the Powell method.  Using least squares gives about 20 bps worse score because the true metric is the mean angular error.  The baseline score was .97948 on Batch 601 (used for weight calculation) and .97964 on Batch 602 (for validation).  This is a 63bp improvement over the best single model.  Since I didn't want to  calculate all 37 models for submission, I choose a subset of models.  You cannot simply use the models with the highest weights to pick a subset because the models are correlated.  For example, if model A and model B are the best and yet effectively the same, then their weights might be approximately equal and the highest of the bunch.  A subset should only include one of them.  So I took an iterative approach, removing a single model at a time and removing the one that had the least negative impact.  These are the results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651278%2Fa906b327fa59d5259a491c6098ce82eb%2Fscore_by_models.png?generation=1682004978180418&alt=media)\n\nYou can see that too many models is bad on cross-validation due to overfitting.  I choose 6 models as a compromise of inference time and best score.\n\n# Lessons Learned\n1. Learn about Transformers!\n1. Have a better framework in place to evaluate different models and parameters without pulling my hair out when changing stuff.\n1. Spend less time on searching for small tweaks when there is a large gap to the leaders\n1. To save time, do more fine tuning of models when adding things rather than training from scratch\n#Things that didn't work for me\n1. I could not get vonMises-Fisher loss to work\n1. Adding additional features, like number of neighbors, sensor type, line-fit estimated direction \n1. Dropout or weight decay.  My models tended not to overfit, probably because they were not complex enough\n1. Pulse selection.  I could never improve on random selection, which still mystifies me\n1. More complex ensembles.  I see @dipamc77 did well with this, but I could not get XGBOOST or a neural network to combine predictions or even the embeddings into anything better than weighted averaging.\n#Conclusion\nCongratulations to all the top finishers!  And a big thank you to all the people posting information to help everyone, to everyone for participating, and to the contest organizers for putting on such a great competition.",
    "2228614": "Hi @solverworld Really nice and interresting solutions. Congratulations! \n\nI fully recognize that more and longer LSTM layers can help significantly in increasing score.\nI'am glad my work was usefull to you.",
    "2228620": "Am I the only one who stuggled so much with vMF/vector training end to end? I must have been doing something seriously wrong. 😯\n\nI guess small unseen details somehow change results so much, xgboost ensemble worked just for me on the first shot.",
    "2228644": "Yes, many thanks.  I was surprised to see LSTM did better than DynEdge in the end.",
    "2228646": "vMF just didn't work for me.  By using norm of vector error, you get some benefit of vMF - when the model is totally uncertain, it picks small length prediction vector, because that minimizes ||v-pred|| over a (longer) wrong direction.  This helps in the ensembling, because those predictions get automatically weighted less.  That's something I forgot to mention in the writeup.  Normalizing the predictions before averaging gives a worse solution because of this effect.",
    "2229201": "solverworld thanks for a topic with cool details. It is very important to share experiences with the community. I read the topic and get new knowledge!",
    "2231447": "That's really shocking for me, too. I've tried to attach many things on LSTM... But it seems larger LSTM and attention was all we needed..."
  },
  "source": "meta"
}