{
  "id": 390873,
  "title": "How many \"bad samples\" there are in the dataset?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/390873",
  "author_name": "",
  "post_date": "2023-02-27T16:30:09.088201Z",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I want to share some techniques that we (QuData team) use to analyze model errors.</p>\n<p>Our models predict not angles but a direction vector. This vector is compared with the target vector. The angle 𝛥ψ between them is the model error (angular error). Another angle (azimuthal error 𝛥α) is calculated between the projections of the target and predicted vectors onto the (x,y) plane. Zenith angle errors did not yet give us any interesting information. For the 𝛥ψ and  𝛥α angles, the histograms are plotted, given below by three algorithms: Weighted Line-fit method, RNN and GNN:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fab8c6f3e850f5bae7c02d329bdfb4500%2F01.png?generation=1677514529586631&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fc30e72a2e75fad4cac46f228ad076b67%2F02.png?generation=1677514571533717&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F0ec159762a7b2e9a399905685558cb40%2F03.png?generation=1677514629707122&amp;alt=media\" alt=\"\"></p>\n<p>The average angular error is the metric for our competition. In addition to the mean, we calculate the median, which is usually used in physical publications when comparing different methods of trajectory reconstruction. In all our models, the median is lower than the mean (and the better the model, the greater this difference).</p>\n<p>In addition to these two characteristics, we evaluate the proportion of examples that turned out to be “bad” (that is, not predicted accurately) for a given method.</p>\n<p>Let’s assume that  all examples for angular errors can be divided into two classes - “good” and “bad”. The proportion of bad examples is w, and of good examples, 1-w. Then the distribution of errors of both classes has the form:<br>\n$$<br>\nP(ψ) = (1-w)  P_0(ψ) + w \\, \\sin(ψ) / 2<br>\n$$<br>\nThe second term is the distribution of the angle ψ = [0, π] between a fixed (target) vector and a vector with a random isotropic direction. For more or less working prediction methods, the first term P(ψ) decreases rapidly. Therefore, to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms. Knowing w, we can construct the error distribution of “good examples” (green line) and “bad” ones (red line). The dotted line is the cumulative probability distribution from the cumulative histogram.</p>\n<p>In the case of azimuth errors, the cumulative distribution is approximated as follows:<br>\n$$<br>\nP(α) = (1-w)  P_0(α)   +    w / π,<br>\n$$<br>\nThe distribution of “bad events” in this case is a uniform distribution.</p>\n<p>As can be seen from the figures above, the proportion of “bad events” by different methods gives close enough results: 60% - 70%. Although as the method improves, this value gradually decreases. :)<br>\nNote that when calculating the projections of vectors on the plane (x, y), it is necessary to control their length in this plane. If the length turns out to be equal to zero, the angle should be chosen randomly from the uniform distribution [0…π]. If this is not done, a sharp outlier will be observed on the distribution in the Line-fit method at α=π/2. It corresponds to single string events in which the sensors are triggered on only one string. In the dataset, with the auxiliary==False filter (only “reliable” pulses), there are 17.7% of single-string events:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F872fe2a73ca48be571d2c3d08b6cbb42%2F04.png?generation=1677514759040284&amp;alt=media\" alt=\"\"></p>\n<p>Clearly, for single-string events, Line-fit predicts the direction vector along the string. Therefore, its projection on the plane (x,y) will be equal to zero, and the value of the azimuth angle is not defined.</p>\n<p>The distributions of the predicted zenith and azimuth angles also look quite interesting:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F6a6004f4a1bb91efe9bb4ce5d346cadb%2F05.png?generation=1677514820098371&amp;alt=media\" alt=\"\"></p>\n<p>Two large values for bins at the zenith angles 0 and π are associated with single-string events (17.7%). On the rest of the distribution, one can also distinguish a random part of sin /2 and a “non-random” one. The six peaks in the distribution of azimuthal angles are related to the “discrete hexagonality” of space in the (x,y) plane - <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">see this discussion\n</a></p>\n<p>The implementation of the function that builds the error distribution for the Line-fit method can be found in the notebook: <a href=\"https://www.kaggle.com/synset/icecube-qudata-line-fit-method\" target=\"_blank\">https://www.kaggle.com/synset/icecube-qudata-line-fit-method</a><br>\nWe use weighted averages, which allows us to reduce the error from 1.21 to 1.18.</p>\n<p>I note that this is only part of the “error hunting” that we are doing. We will post here other aspects of the analysis a bit later in another message.</p>\n<p>Happy kaggling!</p>\n<p>P.S. Useful links:</p>\n<ul>\n<li><p>Graphnet-example <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example</a> - the distribution of the angular error for the Graphnet model is given, the need for segmentation of examples is discussed</p></li>\n<li><p>What is accuracy of current state of the art?<br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191</a> angular, azimuth and zenith errors for Graphnet</p></li>\n<li><p>Derivation of Line-fit least squares <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747</a> </p></li>\n<li><p>Non-uniform (non-isotropic) distribution of incident neutrinos in training data <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537</a> </p></li>\n</ul>",
  "messages": [
    {
      "id": "2161659",
      "postDate": "02/27/2023 16:30:09",
      "content": "<p>Hi all!</p>\n<p>I want to share some techniques that we (QuData team) use to analyze model errors.</p>\n<p>Our models predict not angles but a direction vector. This vector is compared with the target vector. The angle 𝛥ψ between them is the model error (angular error). Another angle (azimuthal error 𝛥α) is calculated between the projections of the target and predicted vectors onto the (x,y) plane. Zenith angle errors did not yet give us any interesting information. For the 𝛥ψ and  𝛥α angles, the histograms are plotted, given below by three algorithms: Weighted Line-fit method, RNN and GNN:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fab8c6f3e850f5bae7c02d329bdfb4500%2F01.png?generation=1677514529586631&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fc30e72a2e75fad4cac46f228ad076b67%2F02.png?generation=1677514571533717&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F0ec159762a7b2e9a399905685558cb40%2F03.png?generation=1677514629707122&amp;alt=media\" alt=\"\"></p>\n<p>The average angular error is the metric for our competition. In addition to the mean, we calculate the median, which is usually used in physical publications when comparing different methods of trajectory reconstruction. In all our models, the median is lower than the mean (and the better the model, the greater this difference).</p>\n<p>In addition to these two characteristics, we evaluate the proportion of examples that turned out to be “bad” (that is, not predicted accurately) for a given method.</p>\n<p>Let’s assume that  all examples for angular errors can be divided into two classes - “good” and “bad”. The proportion of bad examples is w, and of good examples, 1-w. Then the distribution of errors of both classes has the form:<br>\n$$<br>\nP(ψ) = (1-w)  P_0(ψ) + w \\, \\sin(ψ) / 2<br>\n$$<br>\nThe second term is the distribution of the angle ψ = [0, π] between a fixed (target) vector and a vector with a random isotropic direction. For more or less working prediction methods, the first term P(ψ) decreases rapidly. Therefore, to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms. Knowing w, we can construct the error distribution of “good examples” (green line) and “bad” ones (red line). The dotted line is the cumulative probability distribution from the cumulative histogram.</p>\n<p>In the case of azimuth errors, the cumulative distribution is approximated as follows:<br>\n$$<br>\nP(α) = (1-w)  P_0(α)   +    w / π,<br>\n$$<br>\nThe distribution of “bad events” in this case is a uniform distribution.</p>\n<p>As can be seen from the figures above, the proportion of “bad events” by different methods gives close enough results: 60% - 70%. Although as the method improves, this value gradually decreases. :)<br>\nNote that when calculating the projections of vectors on the plane (x, y), it is necessary to control their length in this plane. If the length turns out to be equal to zero, the angle should be chosen randomly from the uniform distribution [0…π]. If this is not done, a sharp outlier will be observed on the distribution in the Line-fit method at α=π/2. It corresponds to single string events in which the sensors are triggered on only one string. In the dataset, with the auxiliary==False filter (only “reliable” pulses), there are 17.7% of single-string events:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F872fe2a73ca48be571d2c3d08b6cbb42%2F04.png?generation=1677514759040284&amp;alt=media\" alt=\"\"></p>\n<p>Clearly, for single-string events, Line-fit predicts the direction vector along the string. Therefore, its projection on the plane (x,y) will be equal to zero, and the value of the azimuth angle is not defined.</p>\n<p>The distributions of the predicted zenith and azimuth angles also look quite interesting:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F6a6004f4a1bb91efe9bb4ce5d346cadb%2F05.png?generation=1677514820098371&amp;alt=media\" alt=\"\"></p>\n<p>Two large values for bins at the zenith angles 0 and π are associated with single-string events (17.7%). On the rest of the distribution, one can also distinguish a random part of sin /2 and a “non-random” one. The six peaks in the distribution of azimuthal angles are related to the “discrete hexagonality” of space in the (x,y) plane - <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">see this discussion\n</a></p>\n<p>The implementation of the function that builds the error distribution for the Line-fit method can be found in the notebook: <a href=\"https://www.kaggle.com/synset/icecube-qudata-line-fit-method\" target=\"_blank\">https://www.kaggle.com/synset/icecube-qudata-line-fit-method</a><br>\nWe use weighted averages, which allows us to reduce the error from 1.21 to 1.18.</p>\n<p>I note that this is only part of the “error hunting” that we are doing. We will post here other aspects of the analysis a bit later in another message.</p>\n<p>Happy kaggling!</p>\n<p>P.S. Useful links:</p>\n<ul>\n<li><p>Graphnet-example <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\" target=\"_blank\">https://www.kaggle.com/code/rasmusrse/graphnet-example</a> - the distribution of the angular error for the Graphnet model is given, the need for segmentation of examples is discussed</p></li>\n<li><p>What is accuracy of current state of the art?<br>\n<a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191</a> angular, azimuth and zenith errors for Graphnet</p></li>\n<li><p>Derivation of Line-fit least squares <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747</a> </p></li>\n<li><p>Non-uniform (non-isotropic) distribution of incident neutrinos in training data <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537</a> </p></li>\n</ul>",
      "rawMarkdown": "Hi all!\n\nI want to share some techniques that we (QuData team) use to analyze model errors.\n\nOur models predict not angles but a direction vector. This vector is compared with the target vector. The angle 𝛥ψ between them is the model error (angular error). Another angle (azimuthal error 𝛥α) is calculated between the projections of the target and predicted vectors onto the (x,y) plane. Zenith angle errors did not yet give us any interesting information. For the 𝛥ψ and  𝛥α angles, the histograms are plotted, given below by three algorithms: Weighted Line-fit method, RNN and GNN:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fab8c6f3e850f5bae7c02d329bdfb4500%2F01.png?generation=1677514529586631&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fc30e72a2e75fad4cac46f228ad076b67%2F02.png?generation=1677514571533717&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F0ec159762a7b2e9a399905685558cb40%2F03.png?generation=1677514629707122&alt=media)\n\nThe average angular error is the metric for our competition. In addition to the mean, we calculate the median, which is usually used in physical publications when comparing different methods of trajectory reconstruction. In all our models, the median is lower than the mean (and the better the model, the greater this difference).\n\nIn addition to these two characteristics, we evaluate the proportion of examples that turned out to be “bad” (that is, not predicted accurately) for a given method.\n\n\nLet’s assume that  all examples for angular errors can be divided into two classes - “good” and “bad”. The proportion of bad examples is w, and of good examples, 1-w. Then the distribution of errors of both classes has the form:\n$$\nP(ψ) = (1-w)  P_0(ψ) + w \\, \\sin(ψ) / 2\n$$\nThe second term is the distribution of the angle ψ = [0, π] between a fixed (target) vector and a vector with a random isotropic direction. For more or less working prediction methods, the first term P<sub>0</sub>(ψ) decreases rapidly. Therefore, to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms. Knowing w, we can construct the error distribution of “good examples” (green line) and “bad” ones (red line). The dotted line is the cumulative probability distribution from the cumulative histogram.\n\nIn the case of azimuth errors, the cumulative distribution is approximated as follows:\n$$\nP(α) = (1-w)  P_0(α)   +    w / π,\n$$\nThe distribution of “bad events” in this case is a uniform distribution.\n\n\n\nAs can be seen from the figures above, the proportion of “bad events” by different methods gives close enough results: 60% - 70%. Although as the method improves, this value gradually decreases. :)\nNote that when calculating the projections of vectors on the plane (x, y), it is necessary to control their length in this plane. If the length turns out to be equal to zero, the angle should be chosen randomly from the uniform distribution [0…π]. If this is not done, a sharp outlier will be observed on the distribution in the Line-fit method at α=π/2. It corresponds to single string events in which the sensors are triggered on only one string. In the dataset, with the auxiliary==False filter (only “reliable” pulses), there are 17.7% of single-string events:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F872fe2a73ca48be571d2c3d08b6cbb42%2F04.png?generation=1677514759040284&alt=media)\n\nClearly, for single-string events, Line-fit predicts the direction vector along the string. Therefore, its projection on the plane (x,y) will be equal to zero, and the value of the azimuth angle is not defined.\n\nThe distributions of the predicted zenith and azimuth angles also look quite interesting:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F6a6004f4a1bb91efe9bb4ce5d346cadb%2F05.png?generation=1677514820098371&alt=media)\n\nTwo large values for bins at the zenith angles 0 and π are associated with single-string events (17.7%). On the rest of the distribution, one can also distinguish a random part of sin /2 and a “non-random” one. The six peaks in the distribution of azimuthal angles are related to the “discrete hexagonality” of space in the (x,y) plane - [see this discussion\n](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537 )\n\nThe implementation of the function that builds the error distribution for the Line-fit method can be found in the notebook: https://www.kaggle.com/synset/icecube-qudata-line-fit-method\nWe use weighted averages, which allows us to reduce the error from 1.21 to 1.18.\n\nI note that this is only part of the “error hunting” that we are doing. We will post here other aspects of the analysis a bit later in another message.\n\nHappy kaggling!\n\nP.S. Useful links:\n\n* Graphnet-example https://www.kaggle.com/code/rasmusrse/graphnet-example - the distribution of the angular error for the Graphnet model is given, the need for segmentation of examples is discussed\n\n* What is accuracy of current state of the art?\nhttps://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191 angular, azimuth and zenith errors for Graphnet\n\n* Derivation of Line-fit least squares https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747 \n\n* Non-uniform (non-isotropic) distribution of incident neutrinos in training data https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537",
      "votes": null
    },
    {
      "id": "2212617",
      "postDate": "04/06/2023 23:19:00",
      "content": "<p>You have done very systematic analysis. Thank you! </p>",
      "rawMarkdown": "You have done very systematic analysis. Thank you!",
      "votes": null
    },
    {
      "id": "2218728",
      "postDate": "04/12/2023 01:15:16",
      "content": "<p>thanks for sharing, after reading I have a question👀, <br>\nI understand how you calculate the unknow <strong>w</strong>, but based on what logic, I know you said</p>\n<blockquote>\n  <p>to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms</p>\n</blockquote>\n<p>but I can't understand why it's approprate</p>\n<p>please answer that if you have time😄</p>",
      "rawMarkdown": "thanks for sharing, after reading I have a question👀, \nI understand how you calculate the unknow **w**, but based on what logic, I know you said\n> to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms\n\nbut I can't understand why it's approprate\n\nplease answer that if you have time😄",
      "votes": null
    },
    {
      "id": "2220322",
      "postDate": "04/13/2023 09:36:18",
      "content": "<p>Good question :)</p>\n<p>Generally speaking, one should specify an analytic expression for each component of the distribution. We know the expression for \"noise\". The expression for the \"correct predictions\" error must be specified in some way (this requires a further study). In any case, it will depend on the parameters.</p>\n<p>Then all components are summed up with weights and the problem of minimizing the deviation of the resulting sum from the empirical distribution is solved. The parameters in this problem are both the weights and the parameters of the unknown distribution component. Unfortunately, there is no time for all this…</p>\n<p>However, in our case, the component responsible for \"correct predictions\" decreases very quickly. Therefore, in our algorithm, we have significantly simplified our lives, while receiving a quite adequate estimate.</p>\n<p>In general, the usual ML-magic :)</p>",
      "rawMarkdown": "Good question :)\n\nGenerally speaking, one should specify an analytic expression for each component of the distribution. We know the expression for \"noise\". The expression for the \"correct predictions\" error must be specified in some way (this requires a further study). In any case, it will depend on the parameters.\n\nThen all components are summed up with weights and the problem of minimizing the deviation of the resulting sum from the empirical distribution is solved. The parameters in this problem are both the weights and the parameters of the unknown distribution component. Unfortunately, there is no time for all this...\n\nHowever, in our case, the component responsible for \"correct predictions\" decreases very quickly. Therefore, in our algorithm, we have significantly simplified our lives, while receiving a quite adequate estimate.\n\nIn general, the usual ML-magic :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2212617,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "04/06/2023 23:19:00",
      "content": "<p>You have done very systematic analysis. Thank you! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2218728,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "04/12/2023 01:15:16",
      "content": "<p>thanks for sharing, after reading I have a question👀, <br>\nI understand how you calculate the unknow <strong>w</strong>, but based on what logic, I know you said</p>\n<blockquote>\n  <p>to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms</p>\n</blockquote>\n<p>but I can't understand why it's approprate</p>\n<p>please answer that if you have time😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 2220322,
          "author_name": "synset",
          "author_url": "",
          "post_date": "04/13/2023 09:36:18",
          "content": "<p>Good question :)</p>\n<p>Generally speaking, one should specify an analytic expression for each component of the distribution. We know the expression for \"noise\". The expression for the \"correct predictions\" error must be specified in some way (this requires a further study). In any case, it will depend on the parameters.</p>\n<p>Then all components are summed up with weights and the problem of minimizing the deviation of the resulting sum from the empirical distribution is solved. The parameters in this problem are both the weights and the parameters of the unknown distribution component. Unfortunately, there is no time for all this…</p>\n<p>However, in our case, the component responsible for \"correct predictions\" decreases very quickly. Therefore, in our algorithm, we have significantly simplified our lives, while receiving a quite adequate estimate.</p>\n<p>In general, the usual ML-magic :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2161659": "Hi all!\n\nI want to share some techniques that we (QuData team) use to analyze model errors.\n\nOur models predict not angles but a direction vector. This vector is compared with the target vector. The angle 𝛥ψ between them is the model error (angular error). Another angle (azimuthal error 𝛥α) is calculated between the projections of the target and predicted vectors onto the (x,y) plane. Zenith angle errors did not yet give us any interesting information. For the 𝛥ψ and  𝛥α angles, the histograms are plotted, given below by three algorithms: Weighted Line-fit method, RNN and GNN:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fab8c6f3e850f5bae7c02d329bdfb4500%2F01.png?generation=1677514529586631&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2Fc30e72a2e75fad4cac46f228ad076b67%2F02.png?generation=1677514571533717&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F0ec159762a7b2e9a399905685558cb40%2F03.png?generation=1677514629707122&alt=media)\n\nThe average angular error is the metric for our competition. In addition to the mean, we calculate the median, which is usually used in physical publications when comparing different methods of trajectory reconstruction. In all our models, the median is lower than the mean (and the better the model, the greater this difference).\n\nIn addition to these two characteristics, we evaluate the proportion of examples that turned out to be “bad” (that is, not predicted accurately) for a given method.\n\n\nLet’s assume that  all examples for angular errors can be divided into two classes - “good” and “bad”. The proportion of bad examples is w, and of good examples, 1-w. Then the distribution of errors of both classes has the form:\n$$\nP(ψ) = (1-w)  P_0(ψ) + w \\, \\sin(ψ) / 2\n$$\nThe second term is the distribution of the angle ψ = [0, π] between a fixed (target) vector and a vector with a random isotropic direction. For more or less working prediction methods, the first term P<sub>0</sub>(ψ) decreases rapidly. Therefore, to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms. Knowing w, we can construct the error distribution of “good examples” (green line) and “bad” ones (red line). The dotted line is the cumulative probability distribution from the cumulative histogram.\n\nIn the case of azimuth errors, the cumulative distribution is approximated as follows:\n$$\nP(α) = (1-w)  P_0(α)   +    w / π,\n$$\nThe distribution of “bad events” in this case is a uniform distribution.\n\n\n\nAs can be seen from the figures above, the proportion of “bad events” by different methods gives close enough results: 60% - 70%. Although as the method improves, this value gradually decreases. :)\nNote that when calculating the projections of vectors on the plane (x, y), it is necessary to control their length in this plane. If the length turns out to be equal to zero, the angle should be chosen randomly from the uniform distribution [0…π]. If this is not done, a sharp outlier will be observed on the distribution in the Line-fit method at α=π/2. It corresponds to single string events in which the sensors are triggered on only one string. In the dataset, with the auxiliary==False filter (only “reliable” pulses), there are 17.7% of single-string events:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F872fe2a73ca48be571d2c3d08b6cbb42%2F04.png?generation=1677514759040284&alt=media)\n\nClearly, for single-string events, Line-fit predicts the direction vector along the string. Therefore, its projection on the plane (x,y) will be equal to zero, and the value of the azimuth angle is not defined.\n\nThe distributions of the predicted zenith and azimuth angles also look quite interesting:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F6a6004f4a1bb91efe9bb4ce5d346cadb%2F05.png?generation=1677514820098371&alt=media)\n\nTwo large values for bins at the zenith angles 0 and π are associated with single-string events (17.7%). On the rest of the distribution, one can also distinguish a random part of sin /2 and a “non-random” one. The six peaks in the distribution of azimuthal angles are related to the “discrete hexagonality” of space in the (x,y) plane - [see this discussion\n](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537 )\n\nThe implementation of the function that builds the error distribution for the Line-fit method can be found in the notebook: https://www.kaggle.com/synset/icecube-qudata-line-fit-method\nWe use weighted averages, which allows us to reduce the error from 1.21 to 1.18.\n\nI note that this is only part of the “error hunting” that we are doing. We will post here other aspects of the analysis a bit later in another message.\n\nHappy kaggling!\n\nP.S. Useful links:\n\n* Graphnet-example https://www.kaggle.com/code/rasmusrse/graphnet-example - the distribution of the angular error for the Graphnet model is given, the need for segmentation of examples is discussed\n\n* What is accuracy of current state of the art?\nhttps://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384191 angular, azimuth and zenith errors for Graphnet\n\n* Derivation of Line-fit least squares https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/381747 \n\n* Non-uniform (non-isotropic) distribution of incident neutrinos in training data https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/386537",
    "2212617": "You have done very systematic analysis. Thank you!",
    "2218728": "thanks for sharing, after reading I have a question👀, \nI understand how you calculate the unknow **w**, but based on what logic, I know you said\n> to estimate the weight w, the distribution integral is used in the range [π/2, π] in which the distribution of random errors predominates. The calculated values of this weight are given above in the headings of the histograms\n\nbut I can't understand why it's approprate\n\nplease answer that if you have time😄",
    "2220322": "Good question :)\n\nGenerally speaking, one should specify an analytic expression for each component of the distribution. We know the expression for \"noise\". The expression for the \"correct predictions\" error must be specified in some way (this requires a further study). In any case, it will depend on the parameters.\n\nThen all components are summed up with weights and the problem of minimizing the deviation of the resulting sum from the empirical distribution is solved. The parameters in this problem are both the weights and the parameters of the unknown distribution component. Unfortunately, there is no time for all this...\n\nHowever, in our case, the component responsible for \"correct predictions\" decreases very quickly. Therefore, in our algorithm, we have significantly simplified our lives, while receiving a quite adequate estimate.\n\nIn general, the usual ML-magic :)"
  },
  "source": "meta"
}