{
  "id": 401163,
  "title": "Intergarate infoes of data analysis and how to use them?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/401163",
  "author_name": "RogerOcean",
  "post_date": "2023-04-12T04:08:07.115000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>hi everyone:</p>\n<p>After reading so many discussions and notebooks, <strong>in data analysis parts</strong>, I think people concentrate on below three aspects:</p>\n<ol>\n<li>Different kind of sensors may preform differently<ol>\n<li><a href=\"https://www.kaggle.com/code/suhancho/eda-inspection-of-noisy-sensors/notebook\" target=\"_blank\">EDA - Inspection of noisy sensors</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381702\" target=\"_blank\">Sensor IDs For 3 Sections of the IceCube Detector</a></li>\n<li><a href=\"https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook\" target=\"_blank\">IceCube Sensor Efficiency: Feature Engineering</a></li></ol></li>\n<li>Divide eventid into different categories, you can easily find <strong>angular error distrbution</strong>, maybe in <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">graphnet_example</a>, many analysis base on or use it, <ol>\n<li>Use <strong>kappa</strong> to group events into different categories. not many people discussed. According to <a href=\"https://github.com/graphnet-team/graphnet/blob/main/src/graphnet/models/task/reconstruction.py\" target=\"_blank\">source code</a>, it's a kind of \"regularization coefficient\". </li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/390873\" target=\"_blank\">How many \"bad samples\" there are in the dataset?</a>. In which writer shared to distinguish bad sample(event_id) by calculating unknown <strong>w</strong>, which is similar to use <strong>kappa</strong>.(writer also attached many relevant links you may check)</li>\n<li>Many discussions are around \"line-fitting\" (most of them do data filter auxiliary=False firstly),like <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392098\" target=\"_blank\">Neutrino classification of events</a>( in which writer use linear model's performance to tell different events ), <a href=\"https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93\" target=\"_blank\">🚀Neutrino Event Type Classifier🚀 AUC Score 0.93</a> (use classifier to tell if eventid locates in area of \"liner model\" performs good),<a href=\"https://www.kaggle.com/code/bowaka/icecube-simple-detector-of-muon-neutrino-type/notebook\" target=\"_blank\">IceCube - Simple Detector of \"Muon Neutrino\" type</a>(, which focus on \"muon event\", )</li>\n<li><a href=\"https://www.kaggle.com/code/mmakhyanov/classification-of-the-events\" target=\"_blank\">Classification of the events</a>, whick also tidy up many infoes from different papers</li></ol></li>\n<li>Analysis on y_label: a/z (not many people discussed, and I can't understand😅)<ol>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/394733\" target=\"_blank\">Data anisotropy along the z-axis</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">Non-uniform (non-isotropic) distribution of incident neutrinos in training data</a></li></ol></li>\n</ol>\n<p>In my thoughs, one way to work on the compitation is modify the GNN model host provided directly(it preform so good it's worth to work on it, maybe add more data)(and there are indeed many other models like lstm I didn't try a lot), but another way is to take advantage of infoes mentioned aboved(and more I haven't mentioned yet), but how to <strong>integrate them to our soluation</strong>, honestly I have no idea.😅</p>\n<p>So how do you use those infoes, I'm quiet curious👀</p>",
  "messages": [
    {
      "id": 2218815,
      "postDate": "2023-04-12T04:08:07.117Z",
      "content": "<p>hi everyone:</p>\n<p>After reading so many discussions and notebooks, <strong>in data analysis parts</strong>, I think people concentrate on below three aspects:</p>\n<ol>\n<li>Different kind of sensors may preform differently<ol>\n<li><a href=\"https://www.kaggle.com/code/suhancho/eda-inspection-of-noisy-sensors/notebook\" target=\"_blank\">EDA - Inspection of noisy sensors</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381702\" target=\"_blank\">Sensor IDs For 3 Sections of the IceCube Detector</a></li>\n<li><a href=\"https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook\" target=\"_blank\">IceCube Sensor Efficiency: Feature Engineering</a></li></ol></li>\n<li>Divide eventid into different categories, you can easily find <strong>angular error distrbution</strong>, maybe in <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">graphnet_example</a>, many analysis base on or use it, <ol>\n<li>Use <strong>kappa</strong> to group events into different categories. not many people discussed. According to <a href=\"https://github.com/graphnet-team/graphnet/blob/main/src/graphnet/models/task/reconstruction.py\" target=\"_blank\">source code</a>, it's a kind of \"regularization coefficient\". </li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/390873\" target=\"_blank\">How many \"bad samples\" there are in the dataset?</a>. In which writer shared to distinguish bad sample(event_id) by calculating unknown <strong>w</strong>, which is similar to use <strong>kappa</strong>.(writer also attached many relevant links you may check)</li>\n<li>Many discussions are around \"line-fitting\" (most of them do data filter auxiliary=False firstly),like <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392098\" target=\"_blank\">Neutrino classification of events</a>( in which writer use linear model's performance to tell different events ), <a href=\"https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93\" target=\"_blank\">🚀Neutrino Event Type Classifier🚀 AUC Score 0.93</a> (use classifier to tell if eventid locates in area of \"liner model\" performs good),<a href=\"https://www.kaggle.com/code/bowaka/icecube-simple-detector-of-muon-neutrino-type/notebook\" target=\"_blank\">IceCube - Simple Detector of \"Muon Neutrino\" type</a>(, which focus on \"muon event\", )</li>\n<li><a href=\"https://www.kaggle.com/code/mmakhyanov/classification-of-the-events\" target=\"_blank\">Classification of the events</a>, whick also tidy up many infoes from different papers</li></ol></li>\n<li>Analysis on y_label: a/z (not many people discussed, and I can't understand😅)<ol>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/394733\" target=\"_blank\">Data anisotropy along the z-axis</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">Non-uniform (non-isotropic) distribution of incident neutrinos in training data</a></li></ol></li>\n</ol>\n<p>In my thoughs, one way to work on the compitation is modify the GNN model host provided directly(it preform so good it's worth to work on it, maybe add more data)(and there are indeed many other models like lstm I didn't try a lot), but another way is to take advantage of infoes mentioned aboved(and more I haven't mentioned yet), but how to <strong>integrate them to our soluation</strong>, honestly I have no idea.😅</p>\n<p>So how do you use those infoes, I'm quiet curious👀</p>",
      "rawMarkdown": "hi everyone:\n\nAfter reading so many discussions and notebooks, **in data analysis parts**, I think people concentrate on below three aspects:\n1. Different kind of sensors may preform differently\n    1. [EDA - Inspection of noisy sensors](https://www.kaggle.com/code/suhancho/eda-inspection-of-noisy-sensors/notebook)\n    2. [Sensor IDs For 3 Sections of the IceCube Detector](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381702)\n    3. [IceCube Sensor Efficiency: Feature Engineering](https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook)\n2. Divide eventid into different categories, you can easily find **angular error distrbution**, maybe in [graphnet_example](https://www.kaggle.com/code/rasmusrse/graphnet-example), many analysis base on or use it, \n    1. Use **kappa** to group events into different categories. not many people discussed. According to [source code](https://github.com/graphnet-team/graphnet/blob/main/src/graphnet/models/task/reconstruction.py), it's a kind of \"regularization coefficient\". \n    2. [How many \"bad samples\" there are in the dataset?](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/390873). In which writer shared to distinguish bad sample(event_id) by calculating unknown **w**, which is similar to use **kappa**.(writer also attached many relevant links you may check)\n    3. Many discussions are around \"line-fitting\" (most of them do data filter auxiliary=False firstly),like [Neutrino classification of events](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392098)( in which writer use linear model's performance to tell different events ), [🚀Neutrino Event Type Classifier🚀 AUC Score 0.93](https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93) (use classifier to tell if eventid locates in area of \"liner model\" performs good),[IceCube - Simple Detector of \"Muon Neutrino\" type](https://www.kaggle.com/code/bowaka/icecube-simple-detector-of-muon-neutrino-type/notebook)(, which focus on \"muon event\", )\n    4. [Classification of the events](https://www.kaggle.com/code/mmakhyanov/classification-of-the-events), whick also tidy up many infoes from different papers\n3. Analysis on y_label: a/z (not many people discussed, and I can't understand😅)\n    1. [Data anisotropy along the z-axis](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/394733)\n    2. [Non-uniform (non-isotropic) distribution of incident neutrinos in training data](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537)\n\nIn my thoughs, one way to work on the compitation is modify the GNN model host provided directly(it preform so good it's worth to work on it, maybe add more data)(and there are indeed many other models like lstm I didn't try a lot), but another way is to take advantage of infoes mentioned aboved(and more I haven't mentioned yet), but how to **integrate them to our soluation**, honestly I have no idea.😅\n\nSo how do you use those infoes, I'm quiet curious👀\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2218815": "hi everyone:\n\nAfter reading so many discussions and notebooks, **in data analysis parts**, I think people concentrate on below three aspects:\n1. Different kind of sensors may preform differently\n    1. [EDA - Inspection of noisy sensors](https://www.kaggle.com/code/suhancho/eda-inspection-of-noisy-sensors/notebook)\n    2. [Sensor IDs For 3 Sections of the IceCube Detector](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381702)\n    3. [IceCube Sensor Efficiency: Feature Engineering](https://www.kaggle.com/code/antonsevostianov/icecube-sensor-efficiency-feature-engineering/notebook)\n2. Divide eventid into different categories, you can easily find **angular error distrbution**, maybe in [graphnet_example](https://www.kaggle.com/code/rasmusrse/graphnet-example), many analysis base on or use it, \n    1. Use **kappa** to group events into different categories. not many people discussed. According to [source code](https://github.com/graphnet-team/graphnet/blob/main/src/graphnet/models/task/reconstruction.py), it's a kind of \"regularization coefficient\". \n    2. [How many \"bad samples\" there are in the dataset?](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/390873). In which writer shared to distinguish bad sample(event_id) by calculating unknown **w**, which is similar to use **kappa**.(writer also attached many relevant links you may check)\n    3. Many discussions are around \"line-fitting\" (most of them do data filter auxiliary=False firstly),like [Neutrino classification of events](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/392098)( in which writer use linear model's performance to tell different events ), [🚀Neutrino Event Type Classifier🚀 AUC Score 0.93](https://www.kaggle.com/code/tatelarkin/neutrino-event-type-classifier-auc-score-0-93) (use classifier to tell if eventid locates in area of \"liner model\" performs good),[IceCube - Simple Detector of \"Muon Neutrino\" type](https://www.kaggle.com/code/bowaka/icecube-simple-detector-of-muon-neutrino-type/notebook)(, which focus on \"muon event\", )\n    4. [Classification of the events](https://www.kaggle.com/code/mmakhyanov/classification-of-the-events), whick also tidy up many infoes from different papers\n3. Analysis on y_label: a/z (not many people discussed, and I can't understand😅)\n    1. [Data anisotropy along the z-axis](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/394733)\n    2. [Non-uniform (non-isotropic) distribution of incident neutrinos in training data](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537)\n\nIn my thoughs, one way to work on the compitation is modify the GNN model host provided directly(it preform so good it's worth to work on it, maybe add more data)(and there are indeed many other models like lstm I didn't try a lot), but another way is to take advantage of infoes mentioned aboved(and more I haven't mentioned yet), but how to **integrate them to our soluation**, honestly I have no idea.😅\n\nSo how do you use those infoes, I'm quiet curious👀\n"
  }
}