{
  "id": 393339,
  "title": "Seemingly bimodal distribution for the events start times",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/393339",
  "author_name": "",
  "post_date": "2023-03-09T00:22:02.272360400Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was exploring the first batch data and came across this distribution:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1004598%2F67187533df84535d55f13fd7d7dbbe58%2Fbimodal.png?generation=1678320486827583&amp;alt=media\" alt=\"\"></p>\n<p>Does anyone know why could that be?</p>\n<pre><code># code to reproduce the plot\n# note that ../sample/train_meta_batch_1.pq is a slice of our train_meta\n\nimport pandas as pd\nimport seaborn as sns\nsns.set_style('whitegrid')\n\nbase_dir = '../'\n\npulses_df = pd.read_parquet(base_dir + 'sample/train_batch_1.pq').reset_index()\nevents_df = pd.read_parquet(base_dir + 'sample/train_meta_batch_1.pq').set_index('event_id')\nevents_df['first_pulse_time'] = pulses_df.iloc[events_df.first_pulse_index].time.set_axis(events_df.index)\n\nax = sns.histplot(events_df.first_pulse_time, log_scale=10)\nax.set_yscale('log')\nax.set_xticks([1e4, 1e5])\nax.tick_params(which=\"both\", bottom=True, color=\".8\")\n</code></pre>",
  "messages": [
    {
      "id": "2174215",
      "postDate": "03/09/2023 00:22:02",
      "content": "<p>I was exploring the first batch data and came across this distribution:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1004598%2F67187533df84535d55f13fd7d7dbbe58%2Fbimodal.png?generation=1678320486827583&amp;alt=media\" alt=\"\"></p>\n<p>Does anyone know why could that be?</p>\n<pre><code># code to reproduce the plot\n# note that ../sample/train_meta_batch_1.pq is a slice of our train_meta\n\nimport pandas as pd\nimport seaborn as sns\nsns.set_style('whitegrid')\n\nbase_dir = '../'\n\npulses_df = pd.read_parquet(base_dir + 'sample/train_batch_1.pq').reset_index()\nevents_df = pd.read_parquet(base_dir + 'sample/train_meta_batch_1.pq').set_index('event_id')\nevents_df['first_pulse_time'] = pulses_df.iloc[events_df.first_pulse_index].time.set_axis(events_df.index)\n\nax = sns.histplot(events_df.first_pulse_time, log_scale=10)\nax.set_yscale('log')\nax.set_xticks([1e4, 1e5])\nax.tick_params(which=\"both\", bottom=True, color=\".8\")\n</code></pre>",
      "rawMarkdown": "I was exploring the first batch data and came across this distribution:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1004598%2F67187533df84535d55f13fd7d7dbbe58%2Fbimodal.png?generation=1678320486827583&alt=media)\n\nDoes anyone know why could that be?\n\n```\n# code to reproduce the plot\n# note that ../sample/train_meta_batch_1.pq is a slice of our train_meta\n\nimport pandas as pd\nimport seaborn as sns\nsns.set_style('whitegrid')\n\nbase_dir = '../'\n\npulses_df = pd.read_parquet(base_dir + 'sample/train_batch_1.pq').reset_index()\nevents_df = pd.read_parquet(base_dir + 'sample/train_meta_batch_1.pq').set_index('event_id')\nevents_df['first_pulse_time'] = pulses_df.iloc[events_df.first_pulse_index].time.set_axis(events_df.index)\n\nax = sns.histplot(events_df.first_pulse_time, log_scale=10)\nax.set_yscale('log')\nax.set_xticks([1e4, 1e5])\nax.tick_params(which=\"both\", bottom=True, color=\".8\")\n```",
      "votes": null
    },
    {
      "id": "2174566",
      "postDate": "03/09/2023 08:21:53",
      "content": "<p>The two distinct clusters of events you found are interesting.</p>\n<p>However, since you are using a logarithmic scale, events with a late start time contribute very little (less than 3%)</p>\n<p>The organizers emphasized that the initial moment of time does not play an important role. It occurs in the \"cutting algorithm\" of an event from a continuous stream of time.</p>",
      "rawMarkdown": "The two distinct clusters of events you found are interesting.\n\nHowever, since you are using a logarithmic scale, events with a late start time contribute very little (less than 3%)\n\nThe organizers emphasized that the initial moment of time does not play an important role. It occurs in the \"cutting algorithm\" of an event from a continuous stream of time.",
      "votes": null
    },
    {
      "id": "2174573",
      "postDate": "03/09/2023 08:34:05",
      "content": "<p>The left cluster pretty much matches what you will find in the documentation scattered across the web for how a cut is made from the data flow and an event created.  There are some 'rules' that trigger the event - event data starts with some head space added.  The head space would likely be there to look for veto conditions - did the event start outside the cube, etc.</p>\n<p>The right as noted by Sergey is a small percentage - in the web cast it was noted that there's some junk thrown in - this might be junk that's typical of a bad trigger condition that they often see - or its junk or as noted by Sergey its just a strange cut from the data flow.</p>\n<p>Some of the events I have looked at don't seem to have met the 'trigger' rules - not sure if those were part of your right hand distribution.</p>",
      "rawMarkdown": "The left cluster pretty much matches what you will find in the documentation scattered across the web for how a cut is made from the data flow and an event created.  There are some 'rules' that trigger the event - event data starts with some head space added.  The head space would likely be there to look for veto conditions - did the event start outside the cube, etc.\n\nThe right as noted by Sergey is a small percentage - in the web cast it was noted that there's some junk thrown in - this might be junk that's typical of a bad trigger condition that they often see - or its junk or as noted by Sergey its just a strange cut from the data flow.\n\nSome of the events I have looked at don't seem to have met the 'trigger' rules - not sure if those were part of your right hand distribution.",
      "votes": null
    },
    {
      "id": "2175428",
      "postDate": "03/09/2023 20:47:54",
      "content": "<p>I used the log scale because the right distribution would be too small to see otherwise haha. Hm, I recalculated your number (less than 3%) and it happens to be exactly 2.75% for this batch (batch 1), a number that seems highly suspicious for me.</p>\n<p>I might look into this a little closer later but don't wanna waste too much time on something that won't add too much value (I mean, the - basic- solution is just to add this information to the model and hope that the it captures something if there's some info there, and that it ignores if there isn't).</p>\n<p>In any case, thanks for answering :)</p>",
      "rawMarkdown": "I used the log scale because the right distribution would be too small to see otherwise haha. Hm, I recalculated your number (less than 3%) and it happens to be exactly 2.75% for this batch (batch 1), a number that seems highly suspicious for me.\n\nI might look into this a little closer later but don't wanna waste too much time on something that won't add too much value (I mean, the - basic- solution is just to add this information to the model and hope that the it captures something if there's some info there, and that it ignores if there isn't).\n\nIn any case, thanks for answering :)",
      "votes": null
    },
    {
      "id": "2175430",
      "postDate": "03/09/2023 20:49:26",
      "content": "<p>A bad trigger you mean like the algorithm detected that there was a neutrino interaction and actually there was none or like, something else?</p>\n<p>I'm not sure if I remember correctly, but I remember seeing somewhere that we are guaranteed that all events must have a least a neutrino.</p>\n<p>Edit: Thanks for the answer BTW :)</p>",
      "rawMarkdown": "A bad trigger you mean like the algorithm detected that there was a neutrino interaction and actually there was none or like, something else?\n\nI'm not sure if I remember correctly, but I remember seeing somewhere that we are guaranteed that all events must have a least a neutrino.\n\nEdit: Thanks for the answer BTW :)",
      "votes": null
    },
    {
      "id": "2175523",
      "postDate": "03/09/2023 23:04:34",
      "content": "<p>I can't find the document which described a typical trigger rule.  As I recall a first piece was that a pair set of sensors would be see light within a very short time frame (sensors apparently wired in pairs).  Seems like a second part required several sensors would see a pulse.  Once that set of rules was met than the software creates the event data from the continuous flow of data.  The creation added some head space.  I assume the head space used to look at some veto conditions.</p>\n<p>So the trigger rules describe a sequence of pulses that might be a neutrino interaction but could also be an atmospheric shower ,which is source of noise to the system.  As I understand it, noise can met the trigger rules and start an event.</p>\n<p>In the web cast it did seem very clear that every event in our data would have a neutrino but also seemed very clear that noise of one form or another could have been the trigger and that our neutrino is hidden in the weeds for us to find.  (In the real world data set every event does not have a neutrino).</p>\n<p>Looking at the output from several of the shared notebooks and my own results there are a large number of events where the prediction is pretty darn good.  My assumption would be that these events had a \"good\" trigger, very little noise, no coincident event and were track like.</p>\n<p>Other events seem to have decent agreement, probably had a good trigger, but not track like.</p>\n<p>The really bad predictions would seem like they have a trigger from noise and a very sneaky neutrino who's pulses are aux = True.  </p>\n<p>Will take a bit of looking to understand if the right hand side cluster that you see in your analysis is a clue to neutrino's in the weeds, or as stated by both Sergey and myself, just a late starter with no abnormal mix to the type of events.</p>\n<p>UPDATE:  As part of other work I was doing I created a plot of angular score using <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">Robert Hatch's</a> notebook vs earliest time when aux = false and see no visual suggestion that the late bloomers are bad actors, and have what looks to be the same distribution of angular scores.  </p>",
      "rawMarkdown": "I can't find the document which described a typical trigger rule.  As I recall a first piece was that a pair set of sensors would be see light within a very short time frame (sensors apparently wired in pairs).  Seems like a second part required several sensors would see a pulse.  Once that set of rules was met than the software creates the event data from the continuous flow of data.  The creation added some head space.  I assume the head space used to look at some veto conditions.\n\nSo the trigger rules describe a sequence of pulses that might be a neutrino interaction but could also be an atmospheric shower ,which is source of noise to the system.  As I understand it, noise can met the trigger rules and start an event.\n\nIn the web cast it did seem very clear that every event in our data would have a neutrino but also seemed very clear that noise of one form or another could have been the trigger and that our neutrino is hidden in the weeds for us to find.  (In the real world data set every event does not have a neutrino).\n\nLooking at the output from several of the shared notebooks and my own results there are a large number of events where the prediction is pretty darn good.  My assumption would be that these events had a \"good\" trigger, very little noise, no coincident event and were track like.\n\nOther events seem to have decent agreement, probably had a good trigger, but not track like.\n\nThe really bad predictions would seem like they have a trigger from noise and a very sneaky neutrino who's pulses are aux = True.  \n\nWill take a bit of looking to understand if the right hand side cluster that you see in your analysis is a clue to neutrino's in the weeds, or as stated by both Sergey and myself, just a late starter with no abnormal mix to the type of events.\n\nUPDATE:  As part of other work I was doing I created a plot of angular score using [Robert Hatch's](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) notebook vs earliest time when aux = false and see no visual suggestion that the late bloomers are bad actors, and have what looks to be the same distribution of angular scores.",
      "votes": null
    },
    {
      "id": "2175530",
      "postDate": "03/09/2023 23:19:33",
      "content": "<blockquote>\n  <p>a number that seems highly suspicious for me</p>\n</blockquote>\n<p>The data is simulated.  Reading lots of documentation on how it's simulated you can see that folks requesting data can describe they nature of the data in a number of parameters.  Depending on what problem your tackling, you can get data descriptive of that problem.</p>\n<p>IF the 2.75% represents a specific problem than you might find that same level of distribution in every batch of the training (and test??) depending on how the host built the data set for our use.  </p>\n<p>I lean towards the view that \"highly suspicious\" are the right words and that 2.75% +/- x does represent a type of problem that occurs in the real world data.  My experience on kaggle for simulated data has been that you can often find clues to improving the model when you guess some of the logic that built the dataset.  That is to say, that some good models in the past have arrived as a result of solving \"how was the data created\" rather than solving the problem. </p>",
      "rawMarkdown": ">a number that seems highly suspicious for me\n\nThe data is simulated.  Reading lots of documentation on how it's simulated you can see that folks requesting data can describe they nature of the data in a number of parameters.  Depending on what problem your tackling, you can get data descriptive of that problem.\n\nIF the 2.75% represents a specific problem than you might find that same level of distribution in every batch of the training (and test??) depending on how the host built the data set for our use.  \n\nI lean towards the view that \"highly suspicious\" are the right words and that 2.75% +/- x does represent a type of problem that occurs in the real world data.  My experience on kaggle for simulated data has been that you can often find clues to improving the model when you guess some of the logic that built the dataset.  That is to say, that some good models in the past have arrived as a result of solving \"how was the data created\" rather than solving the problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2174566,
      "author_name": "synset",
      "author_url": "",
      "post_date": "03/09/2023 08:21:53",
      "content": "<p>The two distinct clusters of events you found are interesting.</p>\n<p>However, since you are using a logarithmic scale, events with a late start time contribute very little (less than 3%)</p>\n<p>The organizers emphasized that the initial moment of time does not play an important role. It occurs in the \"cutting algorithm\" of an event from a continuous stream of time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2175428,
          "author_name": "araraonline",
          "author_url": "",
          "post_date": "03/09/2023 20:47:54",
          "content": "<p>I used the log scale because the right distribution would be too small to see otherwise haha. Hm, I recalculated your number (less than 3%) and it happens to be exactly 2.75% for this batch (batch 1), a number that seems highly suspicious for me.</p>\n<p>I might look into this a little closer later but don't wanna waste too much time on something that won't add too much value (I mean, the - basic- solution is just to add this information to the model and hope that the it captures something if there's some info there, and that it ignores if there isn't).</p>\n<p>In any case, thanks for answering :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2175530,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "03/09/2023 23:19:33",
              "content": "<blockquote>\n  <p>a number that seems highly suspicious for me</p>\n</blockquote>\n<p>The data is simulated.  Reading lots of documentation on how it's simulated you can see that folks requesting data can describe they nature of the data in a number of parameters.  Depending on what problem your tackling, you can get data descriptive of that problem.</p>\n<p>IF the 2.75% represents a specific problem than you might find that same level of distribution in every batch of the training (and test??) depending on how the host built the data set for our use.  </p>\n<p>I lean towards the view that \"highly suspicious\" are the right words and that 2.75% +/- x does represent a type of problem that occurs in the real world data.  My experience on kaggle for simulated data has been that you can often find clues to improving the model when you guess some of the logic that built the dataset.  That is to say, that some good models in the past have arrived as a result of solving \"how was the data created\" rather than solving the problem. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2174573,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "03/09/2023 08:34:05",
      "content": "<p>The left cluster pretty much matches what you will find in the documentation scattered across the web for how a cut is made from the data flow and an event created.  There are some 'rules' that trigger the event - event data starts with some head space added.  The head space would likely be there to look for veto conditions - did the event start outside the cube, etc.</p>\n<p>The right as noted by Sergey is a small percentage - in the web cast it was noted that there's some junk thrown in - this might be junk that's typical of a bad trigger condition that they often see - or its junk or as noted by Sergey its just a strange cut from the data flow.</p>\n<p>Some of the events I have looked at don't seem to have met the 'trigger' rules - not sure if those were part of your right hand distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2175430,
          "author_name": "araraonline",
          "author_url": "",
          "post_date": "03/09/2023 20:49:26",
          "content": "<p>A bad trigger you mean like the algorithm detected that there was a neutrino interaction and actually there was none or like, something else?</p>\n<p>I'm not sure if I remember correctly, but I remember seeing somewhere that we are guaranteed that all events must have a least a neutrino.</p>\n<p>Edit: Thanks for the answer BTW :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2175523,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "03/09/2023 23:04:34",
              "content": "<p>I can't find the document which described a typical trigger rule.  As I recall a first piece was that a pair set of sensors would be see light within a very short time frame (sensors apparently wired in pairs).  Seems like a second part required several sensors would see a pulse.  Once that set of rules was met than the software creates the event data from the continuous flow of data.  The creation added some head space.  I assume the head space used to look at some veto conditions.</p>\n<p>So the trigger rules describe a sequence of pulses that might be a neutrino interaction but could also be an atmospheric shower ,which is source of noise to the system.  As I understand it, noise can met the trigger rules and start an event.</p>\n<p>In the web cast it did seem very clear that every event in our data would have a neutrino but also seemed very clear that noise of one form or another could have been the trigger and that our neutrino is hidden in the weeds for us to find.  (In the real world data set every event does not have a neutrino).</p>\n<p>Looking at the output from several of the shared notebooks and my own results there are a large number of events where the prediction is pretty darn good.  My assumption would be that these events had a \"good\" trigger, very little noise, no coincident event and were track like.</p>\n<p>Other events seem to have decent agreement, probably had a good trigger, but not track like.</p>\n<p>The really bad predictions would seem like they have a trigger from noise and a very sneaky neutrino who's pulses are aux = True.  </p>\n<p>Will take a bit of looking to understand if the right hand side cluster that you see in your analysis is a clue to neutrino's in the weeds, or as stated by both Sergey and myself, just a late starter with no abnormal mix to the type of events.</p>\n<p>UPDATE:  As part of other work I was doing I created a plot of angular score using <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">Robert Hatch's</a> notebook vs earliest time when aux = false and see no visual suggestion that the late bloomers are bad actors, and have what looks to be the same distribution of angular scores.  </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2174215": "I was exploring the first batch data and came across this distribution:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1004598%2F67187533df84535d55f13fd7d7dbbe58%2Fbimodal.png?generation=1678320486827583&alt=media)\n\nDoes anyone know why could that be?\n\n```\n# code to reproduce the plot\n# note that ../sample/train_meta_batch_1.pq is a slice of our train_meta\n\nimport pandas as pd\nimport seaborn as sns\nsns.set_style('whitegrid')\n\nbase_dir = '../'\n\npulses_df = pd.read_parquet(base_dir + 'sample/train_batch_1.pq').reset_index()\nevents_df = pd.read_parquet(base_dir + 'sample/train_meta_batch_1.pq').set_index('event_id')\nevents_df['first_pulse_time'] = pulses_df.iloc[events_df.first_pulse_index].time.set_axis(events_df.index)\n\nax = sns.histplot(events_df.first_pulse_time, log_scale=10)\nax.set_yscale('log')\nax.set_xticks([1e4, 1e5])\nax.tick_params(which=\"both\", bottom=True, color=\".8\")\n```",
    "2174566": "The two distinct clusters of events you found are interesting.\n\nHowever, since you are using a logarithmic scale, events with a late start time contribute very little (less than 3%)\n\nThe organizers emphasized that the initial moment of time does not play an important role. It occurs in the \"cutting algorithm\" of an event from a continuous stream of time.",
    "2174573": "The left cluster pretty much matches what you will find in the documentation scattered across the web for how a cut is made from the data flow and an event created.  There are some 'rules' that trigger the event - event data starts with some head space added.  The head space would likely be there to look for veto conditions - did the event start outside the cube, etc.\n\nThe right as noted by Sergey is a small percentage - in the web cast it was noted that there's some junk thrown in - this might be junk that's typical of a bad trigger condition that they often see - or its junk or as noted by Sergey its just a strange cut from the data flow.\n\nSome of the events I have looked at don't seem to have met the 'trigger' rules - not sure if those were part of your right hand distribution.",
    "2175428": "I used the log scale because the right distribution would be too small to see otherwise haha. Hm, I recalculated your number (less than 3%) and it happens to be exactly 2.75% for this batch (batch 1), a number that seems highly suspicious for me.\n\nI might look into this a little closer later but don't wanna waste too much time on something that won't add too much value (I mean, the - basic- solution is just to add this information to the model and hope that the it captures something if there's some info there, and that it ignores if there isn't).\n\nIn any case, thanks for answering :)",
    "2175430": "A bad trigger you mean like the algorithm detected that there was a neutrino interaction and actually there was none or like, something else?\n\nI'm not sure if I remember correctly, but I remember seeing somewhere that we are guaranteed that all events must have a least a neutrino.\n\nEdit: Thanks for the answer BTW :)",
    "2175523": "I can't find the document which described a typical trigger rule.  As I recall a first piece was that a pair set of sensors would be see light within a very short time frame (sensors apparently wired in pairs).  Seems like a second part required several sensors would see a pulse.  Once that set of rules was met than the software creates the event data from the continuous flow of data.  The creation added some head space.  I assume the head space used to look at some veto conditions.\n\nSo the trigger rules describe a sequence of pulses that might be a neutrino interaction but could also be an atmospheric shower ,which is source of noise to the system.  As I understand it, noise can met the trigger rules and start an event.\n\nIn the web cast it did seem very clear that every event in our data would have a neutrino but also seemed very clear that noise of one form or another could have been the trigger and that our neutrino is hidden in the weeds for us to find.  (In the real world data set every event does not have a neutrino).\n\nLooking at the output from several of the shared notebooks and my own results there are a large number of events where the prediction is pretty darn good.  My assumption would be that these events had a \"good\" trigger, very little noise, no coincident event and were track like.\n\nOther events seem to have decent agreement, probably had a good trigger, but not track like.\n\nThe really bad predictions would seem like they have a trigger from noise and a very sneaky neutrino who's pulses are aux = True.  \n\nWill take a bit of looking to understand if the right hand side cluster that you see in your analysis is a clue to neutrino's in the weeds, or as stated by both Sergey and myself, just a late starter with no abnormal mix to the type of events.\n\nUPDATE:  As part of other work I was doing I created a plot of angular score using [Robert Hatch's](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) notebook vs earliest time when aux = false and see no visual suggestion that the late bloomers are bad actors, and have what looks to be the same distribution of angular scores.",
    "2175530": ">a number that seems highly suspicious for me\n\nThe data is simulated.  Reading lots of documentation on how it's simulated you can see that folks requesting data can describe they nature of the data in a number of parameters.  Depending on what problem your tackling, you can get data descriptive of that problem.\n\nIF the 2.75% represents a specific problem than you might find that same level of distribution in every batch of the training (and test??) depending on how the host built the data set for our use.  \n\nI lean towards the view that \"highly suspicious\" are the right words and that 2.75% +/- x does represent a type of problem that occurs in the real world data.  My experience on kaggle for simulated data has been that you can often find clues to improving the model when you guess some of the logic that built the dataset.  That is to say, that some good models in the past have arrived as a result of solving \"how was the data created\" rather than solving the problem."
  },
  "source": "meta"
}