{
  "id": 380319,
  "title": "Do we try to predict an algorithm?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/380319",
  "author_name": "",
  "post_date": "2023-01-22T19:57:08.471872900Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In my understanding, IceCube doesn't have absolute knowledge about the neutrino source position- the angles we are given in the train data are calculated from the same (?) data we are given, and the problem is that they need to calculate it faster. Thus, our predictions try to predict their algorithm, not the 'real' angles. It would be very interesting to know what algorithm they used…maybe we can just numba/jit it and reproduce their results  with 100% accuracy in the required time (just kidding…probably? haha)</p>",
  "messages": [
    {
      "id": "2111340",
      "postDate": "01/22/2023 19:57:08",
      "content": "<p>In my understanding, IceCube doesn't have absolute knowledge about the neutrino source position- the angles we are given in the train data are calculated from the same (?) data we are given, and the problem is that they need to calculate it faster. Thus, our predictions try to predict their algorithm, not the 'real' angles. It would be very interesting to know what algorithm they used…maybe we can just numba/jit it and reproduce their results  with 100% accuracy in the required time (just kidding…probably? haha)</p>",
      "rawMarkdown": "In my understanding, IceCube doesn't have absolute knowledge about the neutrino source position- the angles we are given in the train data are calculated from the same (?) data we are given, and the problem is that they need to calculate it faster. Thus, our predictions try to predict their algorithm, not the 'real' angles. It would be very interesting to know what algorithm they used...maybe we can just numba/jit it and reproduce their results  with 100% accuracy in the required time (just kidding...probably? haha)",
      "votes": null
    },
    {
      "id": "2111369",
      "postDate": "01/22/2023 20:52:20",
      "content": "<p>Great point, but I'd be rather surprised if that was the case.   It's probably either synthetic data with noise around  known travel paths with fixed angles, or they have a pretty good way to calculate the angles accurately but with a great deal of more horsepower required and is not realtime.  </p>\n<p>As Kaggle comps go, this one is pretty sweet, though probably less about traditional ML and more about performance engineering skills + strong math.</p>\n<p>One way to test your theory out is do some deep analysis on a subset of points and see if the labeled angles are always as accurate as they should be given unlimited time and space.  If they aren't, then yeah, I guess the goal is to overfit their bias..</p>",
      "rawMarkdown": "Great point, but I'd be rather surprised if that was the case.   It's probably either synthetic data with noise around  known travel paths with fixed angles, or they have a pretty good way to calculate the angles accurately but with a great deal of more horsepower required and is not realtime.  \n\nAs Kaggle comps go, this one is pretty sweet, though probably less about traditional ML and more about performance engineering skills + strong math.\n\nOne way to test your theory out is do some deep analysis on a subset of points and see if the labeled angles are always as accurate as they should be given unlimited time and space.  If they aren't, then yeah, I guess the goal is to overfit their bias..",
      "votes": null
    },
    {
      "id": "2111419",
      "postDate": "01/22/2023 22:09:59",
      "content": "<p>It's simulated data, so it's possible some of the best feature engineering and/or algorithms will be optimizing against the assumptions, constraints, etc of the simulation. But the better the physics of the simulation, the more applicable our solutions will be to the real world. It's a good question, though.</p>",
      "rawMarkdown": "It's simulated data, so it's possible some of the best feature engineering and/or algorithms will be optimizing against the assumptions, constraints, etc of the simulation. But the better the physics of the simulation, the more applicable our solutions will be to the real world. It's a good question, though.",
      "votes": null
    },
    {
      "id": "2111421",
      "postDate": "01/22/2023 22:14:30",
      "content": "<p>They did disclose it is simulated data, both train and test, based on a comment from organizer in <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/355653\" target=\"_blank\">this thread</a>. Additionally, <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> shared <a href=\"https://docushare.icecube.wisc.edu/dsweb/Get/Version-104920/09%20-%20IceCube%20Simulation%20Software.pdf\" target=\"_blank\">this IceCube Simulation PDF</a>.</p>",
      "rawMarkdown": "They did disclose it is simulated data, both train and test, based on a comment from organizer in [this thread](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/355653). Additionally, @nofreewill shared [this IceCube Simulation PDF](https://docushare.icecube.wisc.edu/dsweb/Get/Version-104920/09%20-%20IceCube%20Simulation%20Software.pdf).",
      "votes": null
    },
    {
      "id": "2111456",
      "postDate": "01/22/2023 22:43:31",
      "content": "<p>I missed it! Thanks</p>",
      "rawMarkdown": "I missed it! Thanks",
      "votes": null
    },
    {
      "id": "2111461",
      "postDate": "01/22/2023 22:47:43",
      "content": "<p>I wonder if they ran an adversarial validation against their data.  I didn't see it in the deck.  maybe there's a data leak in there somewhere.  there was a comp for synthetic seti that found something that helped them - <a href=\"https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385</a></p>",
      "rawMarkdown": "I wonder if they ran an adversarial validation against their data.  I didn't see it in the deck.  maybe there's a data leak in there somewhere.  there was a comp for synthetic seti that found something that helped them - https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385",
      "votes": null
    },
    {
      "id": "2147488",
      "postDate": "02/16/2023 16:52:04",
      "content": "<p>You are correct. And  I would say they don't really know the incoming angle of the neutrino. Here is a simplified version of my guess of how it works : The entire IcuCube detector is modeled in the simulator. Given any incoming neutrino the simulator will give a sensor response (the training data they give us). So it is a matter of finding the \"inverse function of the simulator\": sensor data --&gt; incoming neutrino,  so that given any sensor response, you can know the incoming neutrino direction. The accuracy of the results relys on the accuracy of the simulator (which encodes the physics of the whole experiment and assumptions of neutrino interaction) and the \"inverse function\" you find. Traditionally the \"inverse function\" is done by a likelihood approach as illustrated in their talk, and they are looking for some ML approach to do it faster and more accurate. </p>\n<p>That being said. If you know more detail about the simulator, and more physics involved in the simulation, it will definitely help you to build your model. Since somehow your ML model should encode the same physics.</p>",
      "rawMarkdown": "You are correct. And  I would say they don't really know the incoming angle of the neutrino. Here is a simplified version of my guess of how it works : The entire IcuCube detector is modeled in the simulator. Given any incoming neutrino the simulator will give a sensor response (the training data they give us). So it is a matter of finding the \"inverse function of the simulator\": sensor data --> incoming neutrino,  so that given any sensor response, you can know the incoming neutrino direction. The accuracy of the results relys on the accuracy of the simulator (which encodes the physics of the whole experiment and assumptions of neutrino interaction) and the \"inverse function\" you find. Traditionally the \"inverse function\" is done by a likelihood approach as illustrated in their talk, and they are looking for some ML approach to do it faster and more accurate. \n\nThat being said. If you know more detail about the simulator, and more physics involved in the simulation, it will definitely help you to build your model. Since somehow your ML model should encode the same physics.",
      "votes": null
    },
    {
      "id": "2147619",
      "postDate": "02/16/2023 18:36:24",
      "content": "<p>Reading many of the publications and considering the number of folks involved and the years of activity it would seem they have a pretty decent way to simulate data.  Since only a few dozen of the really interesting particles have been detected over the decade most of the work they do starts with simulated data.</p>\n<p>I have seen (somewhere in the pubs and docs) the work flow for generating the data and like most simulation data it adds noise.</p>\n<p>On past kaggle competitions using simulated data it did seem that learning the algorithm was more important to a good score.  </p>\n<p>So far my impression is that we need to find how to ignore the noise.  Many of the events seem to be fairly noise free and the math solutions generate pretty good scores on those events.  </p>",
      "rawMarkdown": "Reading many of the publications and considering the number of folks involved and the years of activity it would seem they have a pretty decent way to simulate data.  Since only a few dozen of the really interesting particles have been detected over the decade most of the work they do starts with simulated data.\n\nI have seen (somewhere in the pubs and docs) the work flow for generating the data and like most simulation data it adds noise.\n\nOn past kaggle competitions using simulated data it did seem that learning the algorithm was more important to a good score.  \n\nSo far my impression is that we need to find how to ignore the noise.  Many of the events seem to be fairly noise free and the math solutions generate pretty good scores on those events.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2111369,
      "author_name": "kaggleqrdl",
      "author_url": "",
      "post_date": "01/22/2023 20:52:20",
      "content": "<p>Great point, but I'd be rather surprised if that was the case.   It's probably either synthetic data with noise around  known travel paths with fixed angles, or they have a pretty good way to calculate the angles accurately but with a great deal of more horsepower required and is not realtime.  </p>\n<p>As Kaggle comps go, this one is pretty sweet, though probably less about traditional ML and more about performance engineering skills + strong math.</p>\n<p>One way to test your theory out is do some deep analysis on a subset of points and see if the labeled angles are always as accurate as they should be given unlimited time and space.  If they aren't, then yeah, I guess the goal is to overfit their bias..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2111421,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "01/22/2023 22:14:30",
          "content": "<p>They did disclose it is simulated data, both train and test, based on a comment from organizer in <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/355653\" target=\"_blank\">this thread</a>. Additionally, <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> shared <a href=\"https://docushare.icecube.wisc.edu/dsweb/Get/Version-104920/09%20-%20IceCube%20Simulation%20Software.pdf\" target=\"_blank\">this IceCube Simulation PDF</a>.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2111456,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "01/22/2023 22:43:31",
              "content": "<p>I missed it! Thanks</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2111461,
              "author_name": "kaggleqrdl",
              "author_url": "",
              "post_date": "01/22/2023 22:47:43",
              "content": "<p>I wonder if they ran an adversarial validation against their data.  I didn't see it in the deck.  maybe there's a data leak in there somewhere.  there was a comp for synthetic seti that found something that helped them - <a href=\"https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2111419,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "01/22/2023 22:09:59",
      "content": "<p>It's simulated data, so it's possible some of the best feature engineering and/or algorithms will be optimizing against the assumptions, constraints, etc of the simulation. But the better the physics of the simulation, the more applicable our solutions will be to the real world. It's a good question, though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2147488,
      "author_name": "xingchenxu28",
      "author_url": "",
      "post_date": "02/16/2023 16:52:04",
      "content": "<p>You are correct. And  I would say they don't really know the incoming angle of the neutrino. Here is a simplified version of my guess of how it works : The entire IcuCube detector is modeled in the simulator. Given any incoming neutrino the simulator will give a sensor response (the training data they give us). So it is a matter of finding the \"inverse function of the simulator\": sensor data --&gt; incoming neutrino,  so that given any sensor response, you can know the incoming neutrino direction. The accuracy of the results relys on the accuracy of the simulator (which encodes the physics of the whole experiment and assumptions of neutrino interaction) and the \"inverse function\" you find. Traditionally the \"inverse function\" is done by a likelihood approach as illustrated in their talk, and they are looking for some ML approach to do it faster and more accurate. </p>\n<p>That being said. If you know more detail about the simulator, and more physics involved in the simulation, it will definitely help you to build your model. Since somehow your ML model should encode the same physics.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2147619,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/16/2023 18:36:24",
      "content": "<p>Reading many of the publications and considering the number of folks involved and the years of activity it would seem they have a pretty decent way to simulate data.  Since only a few dozen of the really interesting particles have been detected over the decade most of the work they do starts with simulated data.</p>\n<p>I have seen (somewhere in the pubs and docs) the work flow for generating the data and like most simulation data it adds noise.</p>\n<p>On past kaggle competitions using simulated data it did seem that learning the algorithm was more important to a good score.  </p>\n<p>So far my impression is that we need to find how to ignore the noise.  Many of the events seem to be fairly noise free and the math solutions generate pretty good scores on those events.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2111340": "In my understanding, IceCube doesn't have absolute knowledge about the neutrino source position- the angles we are given in the train data are calculated from the same (?) data we are given, and the problem is that they need to calculate it faster. Thus, our predictions try to predict their algorithm, not the 'real' angles. It would be very interesting to know what algorithm they used...maybe we can just numba/jit it and reproduce their results  with 100% accuracy in the required time (just kidding...probably? haha)",
    "2111369": "Great point, but I'd be rather surprised if that was the case.   It's probably either synthetic data with noise around  known travel paths with fixed angles, or they have a pretty good way to calculate the angles accurately but with a great deal of more horsepower required and is not realtime.  \n\nAs Kaggle comps go, this one is pretty sweet, though probably less about traditional ML and more about performance engineering skills + strong math.\n\nOne way to test your theory out is do some deep analysis on a subset of points and see if the labeled angles are always as accurate as they should be given unlimited time and space.  If they aren't, then yeah, I guess the goal is to overfit their bias..",
    "2111419": "It's simulated data, so it's possible some of the best feature engineering and/or algorithms will be optimizing against the assumptions, constraints, etc of the simulation. But the better the physics of the simulation, the more applicable our solutions will be to the real world. It's a good question, though.",
    "2111421": "They did disclose it is simulated data, both train and test, based on a comment from organizer in [this thread](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/355653). Additionally, @nofreewill shared [this IceCube Simulation PDF](https://docushare.icecube.wisc.edu/dsweb/Get/Version-104920/09%20-%20IceCube%20Simulation%20Software.pdf).",
    "2111456": "I missed it! Thanks",
    "2111461": "I wonder if they ran an adversarial validation against their data.  I didn't see it in the deck.  maybe there's a data leak in there somewhere.  there was a comp for synthetic seti that found something that helped them - https://www.kaggle.com/competitions/seti-breakthrough-listen/discussion/266385",
    "2147488": "You are correct. And  I would say they don't really know the incoming angle of the neutrino. Here is a simplified version of my guess of how it works : The entire IcuCube detector is modeled in the simulator. Given any incoming neutrino the simulator will give a sensor response (the training data they give us). So it is a matter of finding the \"inverse function of the simulator\": sensor data --> incoming neutrino,  so that given any sensor response, you can know the incoming neutrino direction. The accuracy of the results relys on the accuracy of the simulator (which encodes the physics of the whole experiment and assumptions of neutrino interaction) and the \"inverse function\" you find. Traditionally the \"inverse function\" is done by a likelihood approach as illustrated in their talk, and they are looking for some ML approach to do it faster and more accurate. \n\nThat being said. If you know more detail about the simulator, and more physics involved in the simulation, it will definitely help you to build your model. Since somehow your ML model should encode the same physics.",
    "2147619": "Reading many of the publications and considering the number of folks involved and the years of activity it would seem they have a pretty decent way to simulate data.  Since only a few dozen of the really interesting particles have been detected over the decade most of the work they do starts with simulated data.\n\nI have seen (somewhere in the pubs and docs) the work flow for generating the data and like most simulation data it adds noise.\n\nOn past kaggle competitions using simulated data it did seem that learning the algorithm was more important to a good score.  \n\nSo far my impression is that we need to find how to ignore the noise.  Many of the events seem to be fairly noise free and the math solutions generate pretty good scores on those events."
  },
  "source": "meta"
}