{
  "id": 379677,
  "title": "The logic used for 'auxiliary' column True vs False",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/379677",
  "author_name": "Robert Hatch",
  "post_date": "2023-01-20T15:43:37.647000",
  "votes": 39,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The data we're given doesn't have many fields, and clever features replacing/using/manipulating the 'auxiliary' field might end up holding major value. The picture on <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">the data tab</a> highlights the significance of that field. So let's spend 5 minutes researching to find out the initial algorithm that labeled this field for us!</p>\n<p>As expected, the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/overview/additional-resources\" target=\"_blank\">additional resources</a> has a wealth of information. There's a \"Detector &amp; Ice\" section. Digging into the first link: <a href=\"https://arxiv.org/abs/1612.05093\" target=\"_blank\">The IceCube Neutrino Observatory: Instrumentation and Online Systems</a></p>\n<p>What I found is:<br>\n\"Every digitizer launch results in a “hit” record. Hits are transferred from the FPGA to SDRAM lookback memory (LBM) via Direct Memory Access (DMA), and the Main Board CPU bundles them and sends them on request to the surface computers. The amount of information included in a hit depends on whether a signal was also detected in one of the neighboring DOMs. In case of an isolated signal (no coincidence), only a time stamp and brief charge summary are sent, and the digitization process is aborted. Conversely, when a nearest or next-to-nearest neighbor DOM also signals a launch within ±1 µs (local coincidence), the full waveform is compressed and included in the hit record. \"</p>\n<p>That sounds pretty straightforward! And if willing to put in the coding effort, it's very testable/verifiable, to try to define nearest neighbors (presumably a grid layout, probably diagonals aren't included as 'nearest'?) and look at timestamps, and confirm that all 'auxiliary' False have at least one matching 'local coincidence' hit within the allowed timeframe and within the 'next-to-nearest-neighbor' distance, and that all 'auxiliary' True do NOT have any 'local coincidence' matches.</p>\n<p>The fun starts when you start testing alternate features using variations of this logic, or even novel dissimilar approaches! :D What if only including nearest neighbor (aka '1' away)? Or allowing +/- 2 microseconds for neighbors up to '3' away? What will help the next stage algorithm get the most accurate results?</p>\n<p>Good luck!</p>\n<p>Oh, and if you come across more quotes related to this 'auxiliary' label, please update this thread! This was just the quick 5 minute search. :)</p>",
  "messages": [
    {
      "id": 2108528,
      "postDate": "2023-01-20T15:43:37.647Z",
      "content": "<p>The data we're given doesn't have many fields, and clever features replacing/using/manipulating the 'auxiliary' field might end up holding major value. The picture on <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">the data tab</a> highlights the significance of that field. So let's spend 5 minutes researching to find out the initial algorithm that labeled this field for us!</p>\n<p>As expected, the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/overview/additional-resources\" target=\"_blank\">additional resources</a> has a wealth of information. There's a \"Detector &amp; Ice\" section. Digging into the first link: <a href=\"https://arxiv.org/abs/1612.05093\" target=\"_blank\">The IceCube Neutrino Observatory: Instrumentation and Online Systems</a></p>\n<p>What I found is:<br>\n\"Every digitizer launch results in a “hit” record. Hits are transferred from the FPGA to SDRAM lookback memory (LBM) via Direct Memory Access (DMA), and the Main Board CPU bundles them and sends them on request to the surface computers. The amount of information included in a hit depends on whether a signal was also detected in one of the neighboring DOMs. In case of an isolated signal (no coincidence), only a time stamp and brief charge summary are sent, and the digitization process is aborted. Conversely, when a nearest or next-to-nearest neighbor DOM also signals a launch within ±1 µs (local coincidence), the full waveform is compressed and included in the hit record. \"</p>\n<p>That sounds pretty straightforward! And if willing to put in the coding effort, it's very testable/verifiable, to try to define nearest neighbors (presumably a grid layout, probably diagonals aren't included as 'nearest'?) and look at timestamps, and confirm that all 'auxiliary' False have at least one matching 'local coincidence' hit within the allowed timeframe and within the 'next-to-nearest-neighbor' distance, and that all 'auxiliary' True do NOT have any 'local coincidence' matches.</p>\n<p>The fun starts when you start testing alternate features using variations of this logic, or even novel dissimilar approaches! :D What if only including nearest neighbor (aka '1' away)? Or allowing +/- 2 microseconds for neighbors up to '3' away? What will help the next stage algorithm get the most accurate results?</p>\n<p>Good luck!</p>\n<p>Oh, and if you come across more quotes related to this 'auxiliary' label, please update this thread! This was just the quick 5 minute search. :)</p>",
      "rawMarkdown": "The data we're given doesn't have many fields, and clever features replacing/using/manipulating the 'auxiliary' field might end up holding major value. The picture on [the data tab](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) highlights the significance of that field. So let's spend 5 minutes researching to find out the initial algorithm that labeled this field for us!\n\nAs expected, the [additional resources](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/overview/additional-resources) has a wealth of information. There's a \"Detector & Ice\" section. Digging into the first link: [The IceCube Neutrino Observatory: Instrumentation and Online Systems](https://arxiv.org/abs/1612.05093)\n\nWhat I found is:\n\"Every digitizer launch results in a “hit” record. Hits are transferred from the FPGA to SDRAM lookback memory (LBM) via Direct Memory Access (DMA), and the Main Board CPU bundles them and sends them on request to the surface computers. The amount of information included in a hit depends on whether a signal was also detected in one of the neighboring DOMs. In case of an isolated signal (no coincidence), only a time stamp and brief charge summary are sent, and the digitization process is aborted. Conversely, when a nearest or next-to-nearest neighbor DOM also signals a launch within ±1 µs (local coincidence), the full waveform is compressed and included in the hit record. \"\n\nThat sounds pretty straightforward! And if willing to put in the coding effort, it's very testable/verifiable, to try to define nearest neighbors (presumably a grid layout, probably diagonals aren't included as 'nearest'?) and look at timestamps, and confirm that all 'auxiliary' False have at least one matching 'local coincidence' hit within the allowed timeframe and within the 'next-to-nearest-neighbor' distance, and that all 'auxiliary' True do NOT have any 'local coincidence' matches.\n\nThe fun starts when you start testing alternate features using variations of this logic, or even novel dissimilar approaches! :D What if only including nearest neighbor (aka '1' away)? Or allowing +/- 2 microseconds for neighbors up to '3' away? What will help the next stage algorithm get the most accurate results?\n\nGood luck!\n\nOh, and if you come across more quotes related to this 'auxiliary' label, please update this thread! This was just the quick 5 minute search. :)",
      "votes": 39
    },
    {
      "id": 2112754,
      "postDate": "2023-01-23T20:00:44.900Z",
      "content": "<p>UPDATE: The big unknown is how and why a solo auxiliary = False can exist in the data?</p>\n<p>Snippet example from the first batch, first event (batch 1, event 24):</p>\n<p><strong>sensor_id   time  charge  auxiliary       x       y       z</strong><br>\n8        3609   8572   1.025       True -313.60  237.44  348.01<br>\n9        <strong>5057</strong>   8680   3.975       <strong>True</strong>   -9.68  -79.50 -205.47<br>\n10       <strong>5057</strong>   8723   0.775       <strong>True</strong>   -9.68  -79.50 -205.47<br>\n11       2977   8747   1.025       True  576.37  170.92 -135.72<br>\n12       <strong>5059</strong>   9868   1.375      <strong>False</strong>   -9.68  -79.50 -219.49<br>\n13       3496   9976   0.825       True  505.27  257.88  233.90<br>\n14       3161  10259   0.775       True -234.95  140.44 -197.79</p>\n<p>Why is hit #12 aux=False, but hits 9 and 10 (just OVER a microsecond earlier) are aux=True?</p>\n<p>Most of the data, in very little inspecting, did seem to be in (at least) pairs and line up pretty well with the described algo, but the very first aux=False in the entire train dataset does not. :)</p>",
      "rawMarkdown": "UPDATE: The big unknown is how and why a solo auxiliary = False can exist in the data?\n\nSnippet example from the first batch, first event (batch 1, event 24):\n\n**sensor_id   time  charge  auxiliary       x       y       z**\n8        3609   8572   1.025       True -313.60  237.44  348.01\n9        **5057**   8680   3.975       **True**   -9.68  -79.50 -205.47\n10       **5057**   8723   0.775       **True**   -9.68  -79.50 -205.47\n11       2977   8747   1.025       True  576.37  170.92 -135.72\n12       **5059**   9868   1.375      **False**   -9.68  -79.50 -219.49\n13       3496   9976   0.825       True  505.27  257.88  233.90\n14       3161  10259   0.775       True -234.95  140.44 -197.79\n\nWhy is hit #12 aux=False, but hits 9 and 10 (just OVER a microsecond earlier) are aux=True?\n\nMost of the data, in very little inspecting, did seem to be in (at least) pairs and line up pretty well with the described algo, but the very first aux=False in the entire train dataset does not. :)",
      "votes": 4,
      "replies": [
        {
          "id": 2112802,
          "postDate": "2023-01-23T20:58:10.907Z",
          "content": "<p>Yes, I wondered about that too. Really strange.</p>",
          "rawMarkdown": "Yes, I wondered about that too. Really strange."
        },
        {
          "id": 2113763,
          "postDate": "2023-01-24T13:43:33.530Z",
          "content": "<p>The detector is quite a complex machinery, and what a recorded signal is depends on many factors. Just to give one example, per module (sensor) there are different waveform digitizers on board with multiple channels each - a fast digitizer and a slow one for instance. Depending on the current state of a module, some channels may be busy / rearming. Modules per string are linked together in groups, within such a group local coincidences can be formed. A few modules are also \"dead\", meaning that we lost them at some point during the last 10+ years of operation. Different trigger decisions and readout windows can happen and overlap, multiple particle interactions take place, etc.</p>\n<p>So long story short, in the end what you see in one \"event\" may not map 1:1 to the ideal case of what one could expect. I think most of the instances that you checked are indeed pairs of optical modules when auxiliary = False, and most auxiliary hits are isolated. ✔️</p>",
          "rawMarkdown": "The detector is quite a complex machinery, and what a recorded signal is depends on many factors. Just to give one example, per module (sensor) there are different waveform digitizers on board with multiple channels each - a fast digitizer and a slow one for instance. Depending on the current state of a module, some channels may be busy / rearming. Modules per string are linked together in groups, within such a group local coincidences can be formed. A few modules are also \"dead\", meaning that we lost them at some point during the last 10+ years of operation. Different trigger decisions and readout windows can happen and overlap, multiple particle interactions take place, etc.\n\nSo long story short, in the end what you see in one \"event\" may not map 1:1 to the ideal case of what one could expect. I think most of the instances that you checked are indeed pairs of optical modules when auxiliary = False, and most auxiliary hits are isolated. ✔️",
          "votes": 11,
          "replies": [
            {
              "id": 2147727,
              "postDate": "2023-02-16T20:33:08.623Z",
              "content": "<p>Thank you for this information. Do you have insights on how to think about \"Depending on the current state of a module, some channels may be busy / rearming.\". <br>\nIe. It seems we are getting signals from the sensors in the nano-second range and light takes maybe 20 or 30 nano-seconds to get from sensor to sensor. Are these sensors resetting in the nano-second range? or are they maybe building charge over time, then going down over maybe a longer time range? is there a simple way to think of them?</p>",
              "rawMarkdown": "Thank you for this information. Do you have insights on how to think about \"Depending on the current state of a module, some channels may be busy / rearming.\". \nIe. It seems we are getting signals from the sensors in the nano-second range and light takes maybe 20 or 30 nano-seconds to get from sensor to sensor. Are these sensors resetting in the nano-second range? or are they maybe building charge over time, then going down over maybe a longer time range? is there a simple way to think of them?"
            }
          ]
        }
      ]
    },
    {
      "id": 2135915,
      "postDate": "2023-02-09T00:40:37.663Z",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> would the ideal scenario have no <code>True</code> cases? E.g: have a classifier which detects with 100% accuracy <code>True</code> vs <code>False</code>, train on only <code>False</code> cases?</p>",
      "rawMarkdown": "@roberthatch would the ideal scenario have no `True` cases? E.g: have a classifier which detects with 100% accuracy `True` vs `False`, train on only `False` cases?",
      "replies": [
        {
          "id": 2136012,
          "postDate": "2023-02-09T03:22:50.697Z",
          "content": "<p>Hmm, for something simplistic like line fitting, definitely. </p>\n<p>And if you had a way to 100% detect, then eliminating noise is a good thing. </p>\n<p>But for something like a NN, my guess is that it learns to detect noise at the same time as it learns how to interpret signal. So probably don't need to teach it separately, and that could hurt the performance. </p>\n<p>There still might be ways to use more complex geometry truths in conjunction with more trust worthy signals. Or even to somehow ensemble NN on everything with NN on only the most trustworthy signals, or maybe try a GNN with \"edges\" that only connect to or from auxiliary=False?</p>\n<p>But one of the potential benefits of NNs is not necessarily needing to figure details like this out, and just getting good results regardless</p>",
          "rawMarkdown": "Hmm, for something simplistic like line fitting, definitely. \n\nAnd if you had a way to 100% detect, then eliminating noise is a good thing. \n\nBut for something like a NN, my guess is that it learns to detect noise at the same time as it learns how to interpret signal. So probably don't need to teach it separately, and that could hurt the performance. \n\nThere still might be ways to use more complex geometry truths in conjunction with more trust worthy signals. Or even to somehow ensemble NN on everything with NN on only the most trustworthy signals, or maybe try a GNN with \"edges\" that only connect to or from auxiliary=False?\n\nBut one of the potential benefits of NNs is not necessarily needing to figure details like this out, and just getting good results regardless",
          "votes": 1,
          "replies": [
            {
              "id": 2136695,
              "postDate": "2023-02-09T13:48:48.500Z",
              "content": "<p>It might be worth checking training a GNN (the GraphNet for example) with False only vs with both and check if it performs significantly better. I'll probably try this and come back to you. Thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>",
              "rawMarkdown": "It might be worth checking training a GNN (the GraphNet for example) with False only vs with both and check if it performs significantly better. I'll probably try this and come back to you. Thanks @roberthatch "
            },
            {
              "id": 2137054,
              "postDate": "2023-02-09T17:42:57.437Z",
              "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> Robert, in case you are interested, I trained 2 GraphNets models with and without auxiliary=<code>True</code> and saw no significant differences in model performance. Perhaps training on a single batch may not be representative to show there is no significance. Here is the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384833\" target=\"_blank\">discussion</a></p>",
              "rawMarkdown": "@roberthatch Robert, in case you are interested, I trained 2 GraphNets models with and without auxiliary=`True` and saw no significant differences in model performance. Perhaps training on a single batch may not be representative to show there is no significance. Here is the [discussion](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384833)",
              "votes": 1
            },
            {
              "id": 2137057,
              "postDate": "2023-02-09T17:49:06.457Z",
              "content": "<p>Interesting. \"no difference\" could be a decent case for aux=False only. As a way of slightly reducing memory footprint and training runtime. </p>",
              "rawMarkdown": "Interesting. \"no difference\" could be a decent case for aux=False only. As a way of slightly reducing memory footprint and training runtime. "
            }
          ]
        }
      ]
    },
    {
      "id": 2108581,
      "postDate": "2023-01-20T16:35:18.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2112754,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-01-23T20:00:44.900000",
      "content": "<p>UPDATE: The big unknown is how and why a solo auxiliary = False can exist in the data?</p>\n<p>Snippet example from the first batch, first event (batch 1, event 24):</p>\n<p><strong>sensor_id   time  charge  auxiliary       x       y       z</strong><br>\n8        3609   8572   1.025       True -313.60  237.44  348.01<br>\n9        <strong>5057</strong>   8680   3.975       <strong>True</strong>   -9.68  -79.50 -205.47<br>\n10       <strong>5057</strong>   8723   0.775       <strong>True</strong>   -9.68  -79.50 -205.47<br>\n11       2977   8747   1.025       True  576.37  170.92 -135.72<br>\n12       <strong>5059</strong>   9868   1.375      <strong>False</strong>   -9.68  -79.50 -219.49<br>\n13       3496   9976   0.825       True  505.27  257.88  233.90<br>\n14       3161  10259   0.775       True -234.95  140.44 -197.79</p>\n<p>Why is hit #12 aux=False, but hits 9 and 10 (just OVER a microsecond earlier) are aux=True?</p>\n<p>Most of the data, in very little inspecting, did seem to be in (at least) pairs and line up pretty well with the described algo, but the very first aux=False in the entire train dataset does not. :)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2112802,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-01-23T20:58:10.907000",
          "content": "<p>Yes, I wondered about that too. Really strange.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2113763,
          "author_name": "Philipp Eller",
          "author_url": "",
          "post_date": "2023-01-24T13:43:33.530000",
          "content": "<p>The detector is quite a complex machinery, and what a recorded signal is depends on many factors. Just to give one example, per module (sensor) there are different waveform digitizers on board with multiple channels each - a fast digitizer and a slow one for instance. Depending on the current state of a module, some channels may be busy / rearming. Modules per string are linked together in groups, within such a group local coincidences can be formed. A few modules are also \"dead\", meaning that we lost them at some point during the last 10+ years of operation. Different trigger decisions and readout windows can happen and overlap, multiple particle interactions take place, etc.</p>\n<p>So long story short, in the end what you see in one \"event\" may not map 1:1 to the ideal case of what one could expect. I think most of the instances that you checked are indeed pairs of optical modules when auxiliary = False, and most auxiliary hits are isolated. ✔️</p>",
          "votes": 11,
          "replies": [
            {
              "id": 2147727,
              "author_name": "edguy99",
              "author_url": "",
              "post_date": "2023-02-16T20:33:08.623000",
              "content": "<p>Thank you for this information. Do you have insights on how to think about \"Depending on the current state of a module, some channels may be busy / rearming.\". <br>\nIe. It seems we are getting signals from the sensors in the nano-second range and light takes maybe 20 or 30 nano-seconds to get from sensor to sensor. Are these sensors resetting in the nano-second range? or are they maybe building charge over time, then going down over maybe a longer time range? is there a simple way to think of them?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2135915,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2023-02-09T00:40:37.663000",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> would the ideal scenario have no <code>True</code> cases? E.g: have a classifier which detects with 100% accuracy <code>True</code> vs <code>False</code>, train on only <code>False</code> cases?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2136012,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2023-02-09T03:22:50.697000",
          "content": "<p>Hmm, for something simplistic like line fitting, definitely. </p>\n<p>And if you had a way to 100% detect, then eliminating noise is a good thing. </p>\n<p>But for something like a NN, my guess is that it learns to detect noise at the same time as it learns how to interpret signal. So probably don't need to teach it separately, and that could hurt the performance. </p>\n<p>There still might be ways to use more complex geometry truths in conjunction with more trust worthy signals. Or even to somehow ensemble NN on everything with NN on only the most trustworthy signals, or maybe try a GNN with \"edges\" that only connect to or from auxiliary=False?</p>\n<p>But one of the potential benefits of NNs is not necessarily needing to figure details like this out, and just getting good results regardless</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2136695,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2023-02-09T13:48:48.500000",
              "content": "<p>It might be worth checking training a GNN (the GraphNet for example) with False only vs with both and check if it performs significantly better. I'll probably try this and come back to you. Thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2137054,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2023-02-09T17:42:57.437000",
              "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> Robert, in case you are interested, I trained 2 GraphNets models with and without auxiliary=<code>True</code> and saw no significant differences in model performance. Perhaps training on a single batch may not be representative to show there is no significance. Here is the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/384833\" target=\"_blank\">discussion</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2137057,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-09T17:49:06.457000",
              "content": "<p>Interesting. \"no difference\" could be a decent case for aux=False only. As a way of slightly reducing memory footprint and training runtime. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2108581,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-20T16:35:18.087000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2108528": "The data we're given doesn't have many fields, and clever features replacing/using/manipulating the 'auxiliary' field might end up holding major value. The picture on [the data tab](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) highlights the significance of that field. So let's spend 5 minutes researching to find out the initial algorithm that labeled this field for us!\n\nAs expected, the [additional resources](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/overview/additional-resources) has a wealth of information. There's a \"Detector & Ice\" section. Digging into the first link: [The IceCube Neutrino Observatory: Instrumentation and Online Systems](https://arxiv.org/abs/1612.05093)\n\nWhat I found is:\n\"Every digitizer launch results in a “hit” record. Hits are transferred from the FPGA to SDRAM lookback memory (LBM) via Direct Memory Access (DMA), and the Main Board CPU bundles them and sends them on request to the surface computers. The amount of information included in a hit depends on whether a signal was also detected in one of the neighboring DOMs. In case of an isolated signal (no coincidence), only a time stamp and brief charge summary are sent, and the digitization process is aborted. Conversely, when a nearest or next-to-nearest neighbor DOM also signals a launch within ±1 µs (local coincidence), the full waveform is compressed and included in the hit record. \"\n\nThat sounds pretty straightforward! And if willing to put in the coding effort, it's very testable/verifiable, to try to define nearest neighbors (presumably a grid layout, probably diagonals aren't included as 'nearest'?) and look at timestamps, and confirm that all 'auxiliary' False have at least one matching 'local coincidence' hit within the allowed timeframe and within the 'next-to-nearest-neighbor' distance, and that all 'auxiliary' True do NOT have any 'local coincidence' matches.\n\nThe fun starts when you start testing alternate features using variations of this logic, or even novel dissimilar approaches! :D What if only including nearest neighbor (aka '1' away)? Or allowing +/- 2 microseconds for neighbors up to '3' away? What will help the next stage algorithm get the most accurate results?\n\nGood luck!\n\nOh, and if you come across more quotes related to this 'auxiliary' label, please update this thread! This was just the quick 5 minute search. :)",
    "2112754": "UPDATE: The big unknown is how and why a solo auxiliary = False can exist in the data?\n\nSnippet example from the first batch, first event (batch 1, event 24):\n\n**sensor_id   time  charge  auxiliary       x       y       z**\n8        3609   8572   1.025       True -313.60  237.44  348.01\n9        **5057**   8680   3.975       **True**   -9.68  -79.50 -205.47\n10       **5057**   8723   0.775       **True**   -9.68  -79.50 -205.47\n11       2977   8747   1.025       True  576.37  170.92 -135.72\n12       **5059**   9868   1.375      **False**   -9.68  -79.50 -219.49\n13       3496   9976   0.825       True  505.27  257.88  233.90\n14       3161  10259   0.775       True -234.95  140.44 -197.79\n\nWhy is hit #12 aux=False, but hits 9 and 10 (just OVER a microsecond earlier) are aux=True?\n\nMost of the data, in very little inspecting, did seem to be in (at least) pairs and line up pretty well with the described algo, but the very first aux=False in the entire train dataset does not. :)",
    "2135915": "@roberthatch would the ideal scenario have no `True` cases? E.g: have a classifier which detects with 100% accuracy `True` vs `False`, train on only `False` cases?",
    "2108581": ""
  }
}