{
  "id": 63400,
  "title": "Assign track_id trick from #19 solution",
  "url": "/competitions/trackml-particle-identification/writeups/steins-gate-assign-track-id-trick-from-19-solution",
  "author_name": "",
  "post_date": "2018-08-15T19:35:09.389922300Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The full solution is not worth reading given there are already many brilliant solutions shared in the forum. But I'd like to share a small trick that greatly helps speed up the algorithm: how to assign track_id properly to remove duplicate tracks fast. This is useful if you're using primitive clustering method.</p>\n\n<p>Duplicate tracks (track with exactly the same hits) may be discovered in two ways:</p>\n\n<ul>\n<li><p>The same track is re-discovered under a different scanning parameter</p></li>\n<li><p>Two tracks initially are different, but after outliers and two-hits-on-the-same-detector are removed, they end up with being the same track.</p></li>\n</ul>\n\n<p>If you use the track_id generating scheme from the public kernel, you may have noticed a lot of tracks (with different track_id) are actually the same track. </p>\n\n<p>My way to generate track_id is to use the hash function to hash a list of hit_id that belongs to the same track. In Python it looks like</p>\n\n<pre><code>track_id = hash(frozenset(hits.hit_id.values))\n</code></pre>\n\n<p>Then by just looking at the track_id one can immediately tell if two tracks are actually the same. Then you can remove duplicates tracks from a pool of track candidates in O(n) time. Just a quick example, in my first stage of the clustering with track length &gt;= 14, there are 68960 tracks found by DBSCAN, but after duplicate tracks are removed there are only 4737 tracks left. (I probably scanned too much.....). Remove this many duplicates greatly helps speed up the curve fitting, outlier removal and merge processes later on. Merge is never a running time bottleneck for my case.</p>\n\n<p>I hope this is useful to someone.</p>",
  "messages": [
    {
      "id": "371004",
      "postDate": "08/15/2018 19:35:09",
      "content": "<p>The full solution is not worth reading given there are already many brilliant solutions shared in the forum. But I'd like to share a small trick that greatly helps speed up the algorithm: how to assign track_id properly to remove duplicate tracks fast. This is useful if you're using primitive clustering method.</p>\n\n<p>Duplicate tracks (track with exactly the same hits) may be discovered in two ways:</p>\n\n<ul>\n<li><p>The same track is re-discovered under a different scanning parameter</p></li>\n<li><p>Two tracks initially are different, but after outliers and two-hits-on-the-same-detector are removed, they end up with being the same track.</p></li>\n</ul>\n\n<p>If you use the track_id generating scheme from the public kernel, you may have noticed a lot of tracks (with different track_id) are actually the same track. </p>\n\n<p>My way to generate track_id is to use the hash function to hash a list of hit_id that belongs to the same track. In Python it looks like</p>\n\n<pre><code>track_id = hash(frozenset(hits.hit_id.values))\n</code></pre>\n\n<p>Then by just looking at the track_id one can immediately tell if two tracks are actually the same. Then you can remove duplicates tracks from a pool of track candidates in O(n) time. Just a quick example, in my first stage of the clustering with track length &gt;= 14, there are 68960 tracks found by DBSCAN, but after duplicate tracks are removed there are only 4737 tracks left. (I probably scanned too much.....). Remove this many duplicates greatly helps speed up the curve fitting, outlier removal and merge processes later on. Merge is never a running time bottleneck for my case.</p>\n\n<p>I hope this is useful to someone.</p>",
      "rawMarkdown": "The full solution is not worth reading given there are already many brilliant solutions shared in the forum. But I'd like to share a small trick that greatly helps speed up the algorithm: how to assign track_id properly to remove duplicate tracks fast. This is useful if you're using primitive clustering method.\n\nDuplicate tracks (track with exactly the same hits) may be discovered in two ways:\n\n - The same track is re-discovered under a different scanning parameter\n\n - Two tracks initially are different, but after outliers and two-hits-on-the-same-detector are removed, they end up with being the same track.\n\nIf you use the track_id generating scheme from the public kernel, you may have noticed a lot of tracks (with different track_id) are actually the same track. \n\nMy way to generate track_id is to use the hash function to hash a list of hit_id that belongs to the same track. In Python it looks like\n\n    track_id = hash(frozenset(hits.hit_id.values))\n\nThen by just looking at the track_id one can immediately tell if two tracks are actually the same. Then you can remove duplicates tracks from a pool of track candidates in O(n) time. Just a quick example, in my first stage of the clustering with track length &gt;= 14, there are 68960 tracks found by DBSCAN, but after duplicate tracks are removed there are only 4737 tracks left. (I probably scanned too much.....). Remove this many duplicates greatly helps speed up the curve fitting, outlier removal and merge processes later on. Merge is never a running time bottleneck for my case.\n\nI hope this is useful to someone.",
      "votes": null
    },
    {
      "id": "371021",
      "postDate": "08/15/2018 20:36:15",
      "content": "<p>Interesting, I'll try it.</p>",
      "rawMarkdown": "Interesting, I'll try it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 371021,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/15/2018 20:36:15",
      "content": "<p>Interesting, I'll try it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "371004": "The full solution is not worth reading given there are already many brilliant solutions shared in the forum. But I'd like to share a small trick that greatly helps speed up the algorithm: how to assign track_id properly to remove duplicate tracks fast. This is useful if you're using primitive clustering method.\n\nDuplicate tracks (track with exactly the same hits) may be discovered in two ways:\n\n - The same track is re-discovered under a different scanning parameter\n\n - Two tracks initially are different, but after outliers and two-hits-on-the-same-detector are removed, they end up with being the same track.\n\nIf you use the track_id generating scheme from the public kernel, you may have noticed a lot of tracks (with different track_id) are actually the same track. \n\nMy way to generate track_id is to use the hash function to hash a list of hit_id that belongs to the same track. In Python it looks like\n\n    track_id = hash(frozenset(hits.hit_id.values))\n\nThen by just looking at the track_id one can immediately tell if two tracks are actually the same. Then you can remove duplicates tracks from a pool of track candidates in O(n) time. Just a quick example, in my first stage of the clustering with track length &gt;= 14, there are 68960 tracks found by DBSCAN, but after duplicate tracks are removed there are only 4737 tracks left. (I probably scanned too much.....). Remove this many duplicates greatly helps speed up the curve fitting, outlier removal and merge processes later on. Merge is never a running time bottleneck for my case.\n\nI hope this is useful to someone.",
    "371021": "Interesting, I'll try it."
  },
  "source": "meta"
}