{
  "id": 402860,
  "title": "12th Place Solution IceCube",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402860",
  "author_name": "",
  "post_date": "2023-04-20T01:45:13.080249700Z",
  "votes": 22,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Quick summary, details to come:  <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403000\" target=\"_blank\">details here</a></p>\n<ol>\n<li>Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.</li>\n<li>Pool of all models included 30 different models, best combination of 6 selected through optimization method.</li>\n<li>Models included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks</li>\n<li>Models included different number of pulses selected for training and different number of pulses used for inference (not always the same).  Pulses were selected randomly from entire set of pulses.  </li>\n<li>Each model was used 4 times and averaged (because of the random selection of pulse set) - although a simple observation was that smaller events did not require multiple evaluations.</li>\n<li>For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling).  Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.</li>\n<li>Each model used as loss the norm of predicted vector minus true unit direction vector.</li>\n</ol>",
  "messages": [
    {
      "id": "2227733",
      "postDate": "04/20/2023 01:45:13",
      "content": "<p>Quick summary, details to come:  <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403000\" target=\"_blank\">details here</a></p>\n<ol>\n<li>Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.</li>\n<li>Pool of all models included 30 different models, best combination of 6 selected through optimization method.</li>\n<li>Models included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks</li>\n<li>Models included different number of pulses selected for training and different number of pulses used for inference (not always the same).  Pulses were selected randomly from entire set of pulses.  </li>\n<li>Each model was used 4 times and averaged (because of the random selection of pulse set) - although a simple observation was that smaller events did not require multiple evaluations.</li>\n<li>For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling).  Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.</li>\n<li>Each model used as loss the norm of predicted vector minus true unit direction vector.</li>\n</ol>",
      "rawMarkdown": "Quick summary, details to come:  [details here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403000)\n1. Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.\n2. Pool of all models included 30 different models, best combination of 6 selected through optimization method.\n3. Models included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks\n4. Models included different number of pulses selected for training and different number of pulses used for inference (not always the same).  Pulses were selected randomly from entire set of pulses.  \n5. Each model was used 4 times and averaged (because of the random selection of pulse set) - although a simple observation was that smaller events did not require multiple evaluations.\n6. For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling).  Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.\n7. Each model used as loss the norm of predicted vector minus true unit direction vector.",
      "votes": null
    },
    {
      "id": "2227874",
      "postDate": "04/20/2023 04:48:33",
      "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> thanks, for sharing</p>",
      "rawMarkdown": "solverworld thanks, for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227874,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "04/20/2023 04:48:33",
      "content": "<p><a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> thanks, for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2227733": "Quick summary, details to come:  [details here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403000)\n1. Final Ensemble of 6 models, linear weighting determined by nonlinear optimization minimizing MAE (Mean Angular Error) on entire held-out batch.\n2. Pool of all models included 30 different models, best combination of 6 selected through optimization method.\n3. Models included Dynamic Graph Networks of different sizes, and LSTM networks of different sizes and head networks\n4. Models included different number of pulses selected for training and different number of pulses used for inference (not always the same).  Pulses were selected randomly from entire set of pulses.  \n5. Each model was used 4 times and averaged (because of the random selection of pulse set) - although a simple observation was that smaller events did not require multiple evaluations.\n6. For prediction, each batch in the test set was read into memory and cached for use in evaluating all 6 models (and their 4x sampling).  Because of the distribution of number of pulses per event, only about 15% additional inferences were required to produce a 4x average for each model.\n7. Each model used as loss the norm of predicted vector minus true unit direction vector.",
    "2227874": "solverworld thanks, for sharing"
  },
  "source": "meta"
}