{
  "id": 57004,
  "title": "a small leak?",
  "url": "/competitions/trackml-particle-identification/discussion/57004",
  "author_name": "",
  "post_date": "2018-05-17T23:34:39.795349600Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>About 0.14-0.19 of all particle-id is equal 0. However, there is a small leak (it is hard to call it a magic feature) which enables to obtain a subset with 2.5 times higher probability that particle-id /track-id is equal 0. Summarize \"value\" by \"hit-id\" in cells file, merge it with truth file and look at the hits of the sum equal 5,6,7 or 8. Probability that particle-id is equal 0 reaches 0.39-0.49 for this subset. I hope there is no more serious leaks in the datasets.</p>\n\n<p>Q&amp;D illustration:</p>\n\n<p><a href=\"https://www.kaggle.com/sionek/a-small-leak\">https://www.kaggle.com/sionek/a-small-leak</a></p>",
  "messages": [
    {
      "id": "330058",
      "postDate": "05/17/2018 23:34:39",
      "content": "<p>About 0.14-0.19 of all particle-id is equal 0. However, there is a small leak (it is hard to call it a magic feature) which enables to obtain a subset with 2.5 times higher probability that particle-id /track-id is equal 0. Summarize \"value\" by \"hit-id\" in cells file, merge it with truth file and look at the hits of the sum equal 5,6,7 or 8. Probability that particle-id is equal 0 reaches 0.39-0.49 for this subset. I hope there is no more serious leaks in the datasets.</p>\n\n<p>Q&amp;D illustration:</p>\n\n<p><a href=\"https://www.kaggle.com/sionek/a-small-leak\">https://www.kaggle.com/sionek/a-small-leak</a></p>",
      "rawMarkdown": "About 0.14-0.19 of all particle-id is equal 0. However, there is a small leak (it is hard to call it a magic feature) which enables to obtain a subset with 2.5 times higher probability that particle-id /track-id is equal 0. Summarize \"value\" by \"hit-id\" in cells file, merge it with truth file and look at the hits of the sum equal 5,6,7 or 8. Probability that particle-id is equal 0 reaches 0.39-0.49 for this subset. I hope there is no more serious leaks in the datasets.\n\nQ&amp;D illustration:\n\nhttps://www.kaggle.com/sionek/a-small-leak",
      "votes": null
    },
    {
      "id": "330062",
      "postDate": "05/17/2018 23:44:32",
      "content": "<p>nice finding :)</p>",
      "rawMarkdown": "nice finding :)",
      "votes": null
    },
    {
      "id": "330068",
      "postDate": "05/18/2018 00:37:48",
      "content": "<p>Thanks for this. I was starting to wonder when and why trackml decides to categorize a hit as unclassified (track = 0).\nI think high sum(value) means the hit was quite transversal and crossed many detectors in the same module (following one of the organizers kernel, it should be possible to estimate the direction of the hit) and the value being a whole number means it was an outter layer, which might help us estimate when a hit should be unclassified. </p>",
      "rawMarkdown": "Thanks for this. I was starting to wonder when and why trackml decides to categorize a hit as unclassified (track = 0).\nI think high sum(value) means the hit was quite transversal and crossed many detectors in the same module (following one of the organizers kernel, it should be possible to estimate the direction of the hit) and the value being a whole number means it was an outter layer, which might help us estimate when a hit should be unclassified.",
      "votes": null
    },
    {
      "id": "330071",
      "postDate": "05/18/2018 01:06:12",
      "content": "<p>@Riad, Probably you are right. Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.</p>",
      "rawMarkdown": "Riad, Probably you are right. Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.",
      "votes": null
    },
    {
      "id": "330164",
      "postDate": "05/18/2018 07:16:28",
      "content": "<blockquote>\n  <p>Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.</p>\n</blockquote>\n\n<p>That's an interesting thought. Thanks for sharing the find.</p>",
      "rawMarkdown": "&gt; Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.\n\nThat's an interesting thought. Thanks for sharing the find.",
      "votes": null
    },
    {
      "id": "331758",
      "postDate": "05/21/2018 20:16:26",
      "content": "<p>Not a leak but a feature. Random hits do not have the same signature as genuine hits. Now this is not a yes or no, as you're seeing.</p>",
      "rawMarkdown": "Not a leak but a feature. Random hits do not have the same signature as genuine hits. Now this is not a yes or no, as you're seeing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330062,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "05/17/2018 23:44:32",
      "content": "<p>nice finding :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330068,
      "author_name": "riadsouissi",
      "author_url": "",
      "post_date": "05/18/2018 00:37:48",
      "content": "<p>Thanks for this. I was starting to wonder when and why trackml decides to categorize a hit as unclassified (track = 0).\nI think high sum(value) means the hit was quite transversal and crossed many detectors in the same module (following one of the organizers kernel, it should be possible to estimate the direction of the hit) and the value being a whole number means it was an outter layer, which might help us estimate when a hit should be unclassified. </p>",
      "votes": null,
      "replies": [
        {
          "id": 330071,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/18/2018 01:06:12",
          "content": "<p>@Riad, Probably you are right. Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330164,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "05/18/2018 07:16:28",
          "content": "<blockquote>\n  <p>Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.</p>\n</blockquote>\n\n<p>That's an interesting thought. Thanks for sharing the find.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331758,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/21/2018 20:16:26",
      "content": "<p>Not a leak but a feature. Random hits do not have the same signature as genuine hits. Now this is not a yes or no, as you're seeing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "330058": "About 0.14-0.19 of all particle-id is equal 0. However, there is a small leak (it is hard to call it a magic feature) which enables to obtain a subset with 2.5 times higher probability that particle-id /track-id is equal 0. Summarize \"value\" by \"hit-id\" in cells file, merge it with truth file and look at the hits of the sum equal 5,6,7 or 8. Probability that particle-id is equal 0 reaches 0.39-0.49 for this subset. I hope there is no more serious leaks in the datasets.\n\nQ&amp;D illustration:\n\nhttps://www.kaggle.com/sionek/a-small-leak",
    "330062": "nice finding :)",
    "330068": "Thanks for this. I was starting to wonder when and why trackml decides to categorize a hit as unclassified (track = 0).\nI think high sum(value) means the hit was quite transversal and crossed many detectors in the same module (following one of the organizers kernel, it should be possible to estimate the direction of the hit) and the value being a whole number means it was an outter layer, which might help us estimate when a hit should be unclassified.",
    "330071": "Riad, Probably you are right. Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.",
    "330164": "&gt; Maybe some modules are oriented in such a way that some numbers of measured values (direction) make no sense as origin from the collision region. In such a case the above is not a leak but a part of very important feature.\n\nThat's an interesting thought. Thanks for sharing the find.",
    "331758": "Not a leak but a feature. Random hits do not have the same signature as genuine hits. Now this is not a yes or no, as you're seeing."
  },
  "source": "meta"
}