{
  "id": 382038,
  "title": "Understanding the direction vector",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/382038",
  "author_name": "",
  "post_date": "2023-01-29T10:11:01.830148800Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I was doing my preliminary EDA when I came across a bunch of events where there is a very obvious path that the detections occur along. I am attaching one example here.<br><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9836294%2F4c687cd4dcbef1e35876c3dcded0335d%2Fevent_244748_trace_issue.PNG?generation=1674986903743018&amp;alt=media\" alt=\"\"></p>\n<p>I understand the azimuth and zenith angles define only the direction and not the position. But I expect that by viewing the plot from certain vantage points would make that vector and the detections line up. This however isn't happening.<br>\nIs there something wrong in my understanding? Perhaps its a case of the neutrinos and muons going off at different directions? </p>",
  "messages": [
    {
      "id": "2120036",
      "postDate": "01/29/2023 10:11:01",
      "content": "<p>I was doing my preliminary EDA when I came across a bunch of events where there is a very obvious path that the detections occur along. I am attaching one example here.<br><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9836294%2F4c687cd4dcbef1e35876c3dcded0335d%2Fevent_244748_trace_issue.PNG?generation=1674986903743018&amp;alt=media\" alt=\"\"></p>\n<p>I understand the azimuth and zenith angles define only the direction and not the position. But I expect that by viewing the plot from certain vantage points would make that vector and the detections line up. This however isn't happening.<br>\nIs there something wrong in my understanding? Perhaps its a case of the neutrinos and muons going off at different directions? </p>",
      "rawMarkdown": "I was doing my preliminary EDA when I came across a bunch of events where there is a very obvious path that the detections occur along. I am attaching one example here.<br>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9836294%2F4c687cd4dcbef1e35876c3dcded0335d%2Fevent_244748_trace_issue.PNG?generation=1674986903743018&alt=media)\n\nI understand the azimuth and zenith angles define only the direction and not the position. But I expect that by viewing the plot from certain vantage points would make that vector and the detections line up. This however isn't happening.\nIs there something wrong in my understanding? Perhaps its a case of the neutrinos and muons going off at different directions?",
      "votes": null
    },
    {
      "id": "2120983",
      "postDate": "01/30/2023 00:09:36",
      "content": "<p>Well, I'll take a stab at it.<br>\nI think that the conspicuous line of sensor reports that we see comes from a particle, but it isn't the particle we're interested in.</p>\n<p>TLDR: There are 3 types of signatures in this contest: signatures from targets, signatures from noise, and signatures from decoys. The decoys represent nuisance particles from the atmosphere and we need to figure out how to reject them - that's probably the crux of the contest.</p>\n<p>When creating an episode, first they add a target signature, but it might not result in many sensor reports, and they won't necessarily show up in a nice line. Then they add noise. Then they roll the dice to decide if they add any decoys.</p>\n<p>In the episode above, it's hard to see which, if any, sensor reports come from the target. The ones we notice come from a decoy. (This is all just my opinion, and it's worth what you paid for it.) </p>\n<p>Suppose we just run the non-auxiliary reports from an episode like this through a PCA (or line fit, or similar). It will give us the direction of the decoy. That's not the desired answer, although it's certainly pertinent.</p>\n<p>The zeniths for the target particles have a Gaussian-looking distribution centered around 90 degrees. The zeniths for the decoy particles are centered around 45 degrees.</p>\n<p>If we use PCA/line fitting to get our answers (without some fancy filtering of reports or some such), sometimes it will find the target, and sometimes a decoy. If we look at the resulting distribution of zeniths, we will see a mixture of the 2 distributions. This notebook, by  <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> shows it nicely. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ymatioun/icecube-plots-of-predictions</a> </p>",
      "rawMarkdown": "Well, I'll take a stab at it.\nI think that the conspicuous line of sensor reports that we see comes from a particle, but it isn't the particle we're interested in.\n\nTLDR: There are 3 types of signatures in this contest: signatures from targets, signatures from noise, and signatures from decoys. The decoys represent nuisance particles from the atmosphere and we need to figure out how to reject them - that's probably the crux of the contest.\n\nWhen creating an episode, first they add a target signature, but it might not result in many sensor reports, and they won't necessarily show up in a nice line. Then they add noise. Then they roll the dice to decide if they add any decoys.\n\nIn the episode above, it's hard to see which, if any, sensor reports come from the target. The ones we notice come from a decoy. (This is all just my opinion, and it's worth what you paid for it.) \n\nSuppose we just run the non-auxiliary reports from an episode like this through a PCA (or line fit, or similar). It will give us the direction of the decoy. That's not the desired answer, although it's certainly pertinent.\n\nThe zeniths for the target particles have a Gaussian-looking distribution centered around 90 degrees. The zeniths for the decoy particles are centered around 45 degrees.\n\nIf we use PCA/line fitting to get our answers (without some fancy filtering of reports or some such), sometimes it will find the target, and sometimes a decoy. If we look at the resulting distribution of zeniths, we will see a mixture of the 2 distributions. This notebook, by  @ymatioun shows it nicely. [https://www.kaggle.com/code/ymatioun/icecube-plots-of-predictions](url)",
      "votes": null
    },
    {
      "id": "2121523",
      "postDate": "01/30/2023 11:39:08",
      "content": "<p>I agree, one major part of the challenge is to assess if the nice line of particules is or is not coming from a neutrino.</p>\n<p>I would also add that a crucial challenge is to assess if we are able to detect the neutrinos from auxiliary sensors, which would be kind of an unsupervised task.</p>",
      "rawMarkdown": "I agree, one major part of the challenge is to assess if the nice line of particules is or is not coming from a neutrino.\n\nI would also add that a crucial challenge is to assess if we are able to detect the neutrinos from auxiliary sensors, which would be kind of an unsupervised task.",
      "votes": null
    },
    {
      "id": "2121846",
      "postDate": "01/30/2023 15:32:53",
      "content": "<p>Interesting take on the presence of two distributions. I read elsewhere that the dataset is indeed synthetic so a decoy is definitely possible. I wonder what's the criteria for calling a detection auxiliary in that case. Because the detection is always going to be due to some other particle giving off Cherenkov radiation. So far my take as been aux==False means muon/other-particle from the target neutrino interaction, aux==True means it's from the secondary particle shower + instrument noise.</p>",
      "rawMarkdown": "Interesting take on the presence of two distributions. I read elsewhere that the dataset is indeed synthetic so a decoy is definitely possible. I wonder what's the criteria for calling a detection auxiliary in that case. Because the detection is always going to be due to some other particle giving off Cherenkov radiation. So far my take as been aux==False means muon/other-particle from the target neutrino interaction, aux==True means it's from the secondary particle shower + instrument noise.",
      "votes": null
    },
    {
      "id": "2122140",
      "postDate": "01/30/2023 18:14:17",
      "content": "<p>Unfortunately, auxiliary label is not so nice. It's just a very simple logic more relevant for real-time data gathering (of the waveforms we weren't given for ANY points!) for the sensors.</p>\n<p>See <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/379677\" target=\"_blank\">discussion here</a>.</p>\n<p>And my <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">notebook here</a> shows clear improvement when tuning my own aux=False labels (from 1.213 to 1.183, so a small but significant 0.030 avg improvement)</p>",
      "rawMarkdown": "Unfortunately, auxiliary label is not so nice. It's just a very simple logic more relevant for real-time data gathering (of the waveforms we weren't given for ANY points!) for the sensors.\n\nSee [discussion here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/379677).\n\nAnd my [notebook here](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) shows clear improvement when tuning my own aux=False labels (from 1.213 to 1.183, so a small but significant 0.030 avg improvement)",
      "votes": null
    },
    {
      "id": "2122261",
      "postDate": "01/30/2023 19:57:34",
      "content": "<p>I missed this one before I made a similar post myself.  I was expecting to see an alignment with the direction of the Auxiliary == False data points when I plotted there solutions.</p>",
      "rawMarkdown": "I missed this one before I made a similar post myself.  I was expecting to see an alignment with the direction of the Auxiliary == False data points when I plotted there solutions.",
      "votes": null
    },
    {
      "id": "2123649",
      "postDate": "01/31/2023 16:10:13",
      "content": "<p>Yes, I was expecting the same. But perhaps as Robert pointed out, the auxiliary label isn't to be interpreted this way.</p>",
      "rawMarkdown": "Yes, I was expecting the same. But perhaps as Robert pointed out, the auxiliary label isn't to be interpreted this way.",
      "votes": null
    },
    {
      "id": "2123659",
      "postDate": "01/31/2023 16:17:22",
      "content": "<p>Hey that's an interesting find! Upvoted. So clearly the auxiliary label isn't to be taken at face value. It does make me wonder what that'll do to model explicability.</p>",
      "rawMarkdown": "Hey that's an interesting find! Upvoted. So clearly the auxiliary label isn't to be taken at face value. It does make me wonder what that'll do to model explicability.",
      "votes": null
    },
    {
      "id": "2123672",
      "postDate": "01/31/2023 16:21:13",
      "content": "<p><a href=\"https://www.kaggle.com/glazed\" target=\"_blank\">@glazed</a> Thank you so much for this explanation. </p>",
      "rawMarkdown": "glazed Thank you so much for this explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2120983,
      "author_name": "glazed",
      "author_url": "",
      "post_date": "01/30/2023 00:09:36",
      "content": "<p>Well, I'll take a stab at it.<br>\nI think that the conspicuous line of sensor reports that we see comes from a particle, but it isn't the particle we're interested in.</p>\n<p>TLDR: There are 3 types of signatures in this contest: signatures from targets, signatures from noise, and signatures from decoys. The decoys represent nuisance particles from the atmosphere and we need to figure out how to reject them - that's probably the crux of the contest.</p>\n<p>When creating an episode, first they add a target signature, but it might not result in many sensor reports, and they won't necessarily show up in a nice line. Then they add noise. Then they roll the dice to decide if they add any decoys.</p>\n<p>In the episode above, it's hard to see which, if any, sensor reports come from the target. The ones we notice come from a decoy. (This is all just my opinion, and it's worth what you paid for it.) </p>\n<p>Suppose we just run the non-auxiliary reports from an episode like this through a PCA (or line fit, or similar). It will give us the direction of the decoy. That's not the desired answer, although it's certainly pertinent.</p>\n<p>The zeniths for the target particles have a Gaussian-looking distribution centered around 90 degrees. The zeniths for the decoy particles are centered around 45 degrees.</p>\n<p>If we use PCA/line fitting to get our answers (without some fancy filtering of reports or some such), sometimes it will find the target, and sometimes a decoy. If we look at the resulting distribution of zeniths, we will see a mixture of the 2 distributions. This notebook, by  <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> shows it nicely. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ymatioun/icecube-plots-of-predictions</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2121523,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "01/30/2023 11:39:08",
          "content": "<p>I agree, one major part of the challenge is to assess if the nice line of particules is or is not coming from a neutrino.</p>\n<p>I would also add that a crucial challenge is to assess if we are able to detect the neutrinos from auxiliary sensors, which would be kind of an unsupervised task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2121846,
          "author_name": "chahbazaman",
          "author_url": "",
          "post_date": "01/30/2023 15:32:53",
          "content": "<p>Interesting take on the presence of two distributions. I read elsewhere that the dataset is indeed synthetic so a decoy is definitely possible. I wonder what's the criteria for calling a detection auxiliary in that case. Because the detection is always going to be due to some other particle giving off Cherenkov radiation. So far my take as been aux==False means muon/other-particle from the target neutrino interaction, aux==True means it's from the secondary particle shower + instrument noise.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2122140,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "01/30/2023 18:14:17",
              "content": "<p>Unfortunately, auxiliary label is not so nice. It's just a very simple logic more relevant for real-time data gathering (of the waveforms we weren't given for ANY points!) for the sensors.</p>\n<p>See <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/379677\" target=\"_blank\">discussion here</a>.</p>\n<p>And my <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">notebook here</a> shows clear improvement when tuning my own aux=False labels (from 1.213 to 1.183, so a small but significant 0.030 avg improvement)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2123659,
                  "author_name": "chahbazaman",
                  "author_url": "",
                  "post_date": "01/31/2023 16:17:22",
                  "content": "<p>Hey that's an interesting find! Upvoted. So clearly the auxiliary label isn't to be taken at face value. It does make me wonder what that'll do to model explicability.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2123672,
          "author_name": "microtus",
          "author_url": "",
          "post_date": "01/31/2023 16:21:13",
          "content": "<p><a href=\"https://www.kaggle.com/glazed\" target=\"_blank\">@glazed</a> Thank you so much for this explanation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2122261,
      "author_name": "kpkent",
      "author_url": "",
      "post_date": "01/30/2023 19:57:34",
      "content": "<p>I missed this one before I made a similar post myself.  I was expecting to see an alignment with the direction of the Auxiliary == False data points when I plotted there solutions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2123649,
          "author_name": "chahbazaman",
          "author_url": "",
          "post_date": "01/31/2023 16:10:13",
          "content": "<p>Yes, I was expecting the same. But perhaps as Robert pointed out, the auxiliary label isn't to be interpreted this way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2120036": "I was doing my preliminary EDA when I came across a bunch of events where there is a very obvious path that the detections occur along. I am attaching one example here.<br>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9836294%2F4c687cd4dcbef1e35876c3dcded0335d%2Fevent_244748_trace_issue.PNG?generation=1674986903743018&alt=media)\n\nI understand the azimuth and zenith angles define only the direction and not the position. But I expect that by viewing the plot from certain vantage points would make that vector and the detections line up. This however isn't happening.\nIs there something wrong in my understanding? Perhaps its a case of the neutrinos and muons going off at different directions?",
    "2120983": "Well, I'll take a stab at it.\nI think that the conspicuous line of sensor reports that we see comes from a particle, but it isn't the particle we're interested in.\n\nTLDR: There are 3 types of signatures in this contest: signatures from targets, signatures from noise, and signatures from decoys. The decoys represent nuisance particles from the atmosphere and we need to figure out how to reject them - that's probably the crux of the contest.\n\nWhen creating an episode, first they add a target signature, but it might not result in many sensor reports, and they won't necessarily show up in a nice line. Then they add noise. Then they roll the dice to decide if they add any decoys.\n\nIn the episode above, it's hard to see which, if any, sensor reports come from the target. The ones we notice come from a decoy. (This is all just my opinion, and it's worth what you paid for it.) \n\nSuppose we just run the non-auxiliary reports from an episode like this through a PCA (or line fit, or similar). It will give us the direction of the decoy. That's not the desired answer, although it's certainly pertinent.\n\nThe zeniths for the target particles have a Gaussian-looking distribution centered around 90 degrees. The zeniths for the decoy particles are centered around 45 degrees.\n\nIf we use PCA/line fitting to get our answers (without some fancy filtering of reports or some such), sometimes it will find the target, and sometimes a decoy. If we look at the resulting distribution of zeniths, we will see a mixture of the 2 distributions. This notebook, by  @ymatioun shows it nicely. [https://www.kaggle.com/code/ymatioun/icecube-plots-of-predictions](url)",
    "2121523": "I agree, one major part of the challenge is to assess if the nice line of particules is or is not coming from a neutrino.\n\nI would also add that a crucial challenge is to assess if we are able to detect the neutrinos from auxiliary sensors, which would be kind of an unsupervised task.",
    "2121846": "Interesting take on the presence of two distributions. I read elsewhere that the dataset is indeed synthetic so a decoy is definitely possible. I wonder what's the criteria for calling a detection auxiliary in that case. Because the detection is always going to be due to some other particle giving off Cherenkov radiation. So far my take as been aux==False means muon/other-particle from the target neutrino interaction, aux==True means it's from the secondary particle shower + instrument noise.",
    "2122140": "Unfortunately, auxiliary label is not so nice. It's just a very simple logic more relevant for real-time data gathering (of the waveforms we weren't given for ANY points!) for the sensors.\n\nSee [discussion here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/379677).\n\nAnd my [notebook here](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) shows clear improvement when tuning my own aux=False labels (from 1.213 to 1.183, so a small but significant 0.030 avg improvement)",
    "2122261": "I missed this one before I made a similar post myself.  I was expecting to see an alignment with the direction of the Auxiliary == False data points when I plotted there solutions.",
    "2123649": "Yes, I was expecting the same. But perhaps as Robert pointed out, the auxiliary label isn't to be interpreted this way.",
    "2123659": "Hey that's an interesting find! Upvoted. So clearly the auxiliary label isn't to be taken at face value. It does make me wonder what that'll do to model explicability.",
    "2123672": "glazed Thank you so much for this explanation."
  },
  "source": "meta"
}