{
  "id": 57527,
  "title": "Throw out Garbage (ID = 0)",
  "url": "/competitions/trackml-particle-identification/discussion/57527",
  "author_name": "",
  "post_date": "2018-05-24T22:53:43.723935900Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As we've all already seen, there's a lot of trash hits in the dataset that aren't part of any particle track, with Particle ID 0 used as the garbage collection for hits that can't be classified. My question is: has anyone had any success with classification tools either labeling or clustering these out before doing the helix+dbscan tricks? </p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "333334",
      "postDate": "05/24/2018 22:53:43",
      "content": "<p>As we've all already seen, there's a lot of trash hits in the dataset that aren't part of any particle track, with Particle ID 0 used as the garbage collection for hits that can't be classified. My question is: has anyone had any success with classification tools either labeling or clustering these out before doing the helix+dbscan tricks? </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "As we've all already seen, there's a lot of trash hits in the dataset that aren't part of any particle track, with Particle ID 0 used as the garbage collection for hits that can't be classified. My question is: has anyone had any success with classification tools either labeling or clustering these out before doing the helix+dbscan tricks? \n\nThanks!",
      "votes": null
    },
    {
      "id": "333574",
      "postDate": "05/25/2018 12:40:08",
      "content": "<p>I have tried without much success.  </p>\n\n<p>The amount of charge deposited in the material is proportional to the length of the material traversed by the particle.  Once the track parameters are known, the path length in material can be calculated and the charge deposition “value” can be used to discard bad hits.  It should be the best feature.  Unfortunately this requires you to have the track parameters, and to calculate the translation and rotation into local u, v, w coordinates.  </p>",
      "rawMarkdown": "I have tried without much success.  \n\nThe amount of charge deposited in the material is proportional to the length of the material traversed by the particle.  Once the track parameters are known, the path length in material can be calculated and the charge deposition “value” can be used to discard bad hits.  It should be the best feature.  Unfortunately this requires you to have the track parameters, and to calculate the translation and rotation into local u, v, w coordinates.",
      "votes": null
    },
    {
      "id": "334233",
      "postDate": "05/26/2018 18:47:22",
      "content": "<p>I believe false positives aren't penalized, so throwing them out would only be useful in terms of speeding up downstream analysis or if your process doesn't have a means of overwriting previous, incorrect labels. </p>\n\n<p>For example, if your script iterates through multiple attempts to label tracks, and will assign a label to a hit based on the number of hits in the new track, then that hit should score the same at the end of the process, even if you remove the spurious hit at the start. It is worth noting that incorrectly removing a hit should hurt your score. </p>\n\n<p>In practice, by modifying the script by Grzegorz Sionkowski to discard 'trash hits' I've hurt my score slightly a few times, but usually it stays the same. I saw the same results if I removed any track IDs with only one hit at the end of all iterations. I believe this would only happen if an early iteration labeled the hit with other hits into a track, and then later the other hits were reassigned to tracks with more hits. </p>",
      "rawMarkdown": "I believe false positives aren't penalized, so throwing them out would only be useful in terms of speeding up downstream analysis or if your process doesn't have a means of overwriting previous, incorrect labels. \n\nFor example, if your script iterates through multiple attempts to label tracks, and will assign a label to a hit based on the number of hits in the new track, then that hit should score the same at the end of the process, even if you remove the spurious hit at the start. It is worth noting that incorrectly removing a hit should hurt your score. \n\nIn practice, by modifying the script by Grzegorz Sionkowski to discard 'trash hits' I've hurt my score slightly a few times, but usually it stays the same. I saw the same results if I removed any track IDs with only one hit at the end of all iterations. I believe this would only happen if an early iteration labeled the hit with other hits into a track, and then later the other hits were reassigned to tracks with more hits.",
      "votes": null
    },
    {
      "id": "334263",
      "postDate": "05/26/2018 20:02:35",
      "content": "<p>It depends on your strategy.  I think the bad hits would hurt sequential based alogorithms by falsely linking hits together.  </p>",
      "rawMarkdown": "It depends on your strategy.  I think the bad hits would hurt sequential based alogorithms by falsely linking hits together.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 333574,
      "author_name": "dmriser",
      "author_url": "",
      "post_date": "05/25/2018 12:40:08",
      "content": "<p>I have tried without much success.  </p>\n\n<p>The amount of charge deposited in the material is proportional to the length of the material traversed by the particle.  Once the track parameters are known, the path length in material can be calculated and the charge deposition “value” can be used to discard bad hits.  It should be the best feature.  Unfortunately this requires you to have the track parameters, and to calculate the translation and rotation into local u, v, w coordinates.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 334233,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "05/26/2018 18:47:22",
      "content": "<p>I believe false positives aren't penalized, so throwing them out would only be useful in terms of speeding up downstream analysis or if your process doesn't have a means of overwriting previous, incorrect labels. </p>\n\n<p>For example, if your script iterates through multiple attempts to label tracks, and will assign a label to a hit based on the number of hits in the new track, then that hit should score the same at the end of the process, even if you remove the spurious hit at the start. It is worth noting that incorrectly removing a hit should hurt your score. </p>\n\n<p>In practice, by modifying the script by Grzegorz Sionkowski to discard 'trash hits' I've hurt my score slightly a few times, but usually it stays the same. I saw the same results if I removed any track IDs with only one hit at the end of all iterations. I believe this would only happen if an early iteration labeled the hit with other hits into a track, and then later the other hits were reassigned to tracks with more hits. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 334263,
      "author_name": "dmriser",
      "author_url": "",
      "post_date": "05/26/2018 20:02:35",
      "content": "<p>It depends on your strategy.  I think the bad hits would hurt sequential based alogorithms by falsely linking hits together.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "333334": "As we've all already seen, there's a lot of trash hits in the dataset that aren't part of any particle track, with Particle ID 0 used as the garbage collection for hits that can't be classified. My question is: has anyone had any success with classification tools either labeling or clustering these out before doing the helix+dbscan tricks? \n\nThanks!",
    "333574": "I have tried without much success.  \n\nThe amount of charge deposited in the material is proportional to the length of the material traversed by the particle.  Once the track parameters are known, the path length in material can be calculated and the charge deposition “value” can be used to discard bad hits.  It should be the best feature.  Unfortunately this requires you to have the track parameters, and to calculate the translation and rotation into local u, v, w coordinates.",
    "334233": "I believe false positives aren't penalized, so throwing them out would only be useful in terms of speeding up downstream analysis or if your process doesn't have a means of overwriting previous, incorrect labels. \n\nFor example, if your script iterates through multiple attempts to label tracks, and will assign a label to a hit based on the number of hits in the new track, then that hit should score the same at the end of the process, even if you remove the spurious hit at the start. It is worth noting that incorrectly removing a hit should hurt your score. \n\nIn practice, by modifying the script by Grzegorz Sionkowski to discard 'trash hits' I've hurt my score slightly a few times, but usually it stays the same. I saw the same results if I removed any track IDs with only one hit at the end of all iterations. I believe this would only happen if an early iteration labeled the hit with other hits into a track, and then later the other hits were reassigned to tracks with more hits.",
    "334263": "It depends on your strategy.  I think the bad hits would hurt sequential based alogorithms by falsely linking hits together."
  },
  "source": "meta"
}