{
  "id": 63259,
  "title": "54th Place Solution for bkKaggle Team",
  "url": "/competitions/trackml-particle-identification/writeups/bkkaggle-54th-place-solution-for-bkkaggle-team",
  "author_name": "",
  "post_date": "2018-08-14T03:30:05.228707500Z",
  "votes": 17,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>54th Place Solution for bkKaggle team; 0.59229 private LB</h1>\n\n<h2>Third Kaggle competition and a second bronze medal!</h2>\n\n<p>Hi, I’m a 15 year old Kaggle beginner and this is my third Kaggle competition and second bronze medal! This competition was different from others since it wasn’t a straightforward supervised learning problem which made it challenging, and I wouldn't have achieved this score without all the ideas shared on the discussion forum. I know that my score doesn't compare to the winners's solutions, but I wanted to share it anyway.</p>\n\n<p>For most the first two months of the competition, I mostly focused on deep learning based solutions but found them hard to train and they took a long time. It was only in the last month of the competition that I focused on implementing an unsupervised clustering based solution.</p>\n\n<p>Some of my deep learning based ideas that I partially completed or didn’t attempt included: Using an mlp to find triplets, implementing a PointNet to cluster tracks, and creating word2vec style embeddings for each track and then clustering the embeddings. Unfortunately, I wasn’t able to get any of these approaches to work really well in the 2 months. After that, I shifted my focus to clustering based approaches.</p>\n\n<p>My final solution was a DBSCAN clustering and helix unrolling with z shifting and track extension.</p>\n\n<h3>Features</h3>\n\n<p>The features I ended up using were:</p>\n\n<p><code>cos(a), sin(a), z/rt, z/r, x/r, y/r</code> </p>\n\n<p>where </p>\n\n<p><code>a = arctan2(y, x) - arccos(mm * ii * rt)</code></p>\n\n<p><code>r = sqrt(x^2 + y^2 + z^2)</code></p>\n\n<p><code>rt = sqrt(x^2 + y^2)</code></p>\n\n<p>These features were from <a href=\"https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel\">Grzegorz's</a> and <a href=\"https://www.kaggle.com/khahuras/0-53x-clustering-using-hough-features-basic\">Kha A Vo's</a> kernels. Unfortunately, I didn’t have the time to find better features with the ideas shared in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/61590\">this</a> discussion post. I originally used <code>cos a, sin a, z/r, and z/rt</code> as my features; When merging based on track length, these features give you a score of 0.35 with no weights, and I couldn’t take it beyond 0.5 even after extensive Bayesian optimization. The addition of <code>x/r and y/r</code> and the default weights from the second kernel gives me a score of 0.54 without any z-shifting or track extension.</p>\n\n<h3>Merging</h3>\n\n<p>My merging strategy is very simple since I didn’t have time to improve it. I simply assign hits to the longest track with no more than 25 hits. </p>\n\n<h3>Z-shifting</h3>\n\n<p>I use 11 z-shifts uniformly distributed around the origin. I search +/- 5mm with steps of 1mm and the origin itself for a total of 11 z-shifts. The drawback of using so many z-shifts is that my running time increases by a factor of 11.</p>\n\n<h3>Track Extension</h3>\n\n<p>I use a modified version of track extension from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58194\">Heng's track extension post</a> . I found that for me, the optimum number of track extension rounds is 6. I wasn't able to use more because each further track extension gave diminishing increases in score and took an extra 1-2 minutes.</p>\n\n<h3>Hyperparameters</h3>\n\n<p>I used Bayesian optimization on top of the weights given in the second kernel. Also optimizing the number of DBSCAN iterations and it's epsilon hyperparameter further increased the score. My final hyperparameters are:</p>\n\n<pre><code>cos(aa) and sin(aa): 1.7\nz/rt: 0.8\nz/r: 0.2\nx/r: 0.015\ny/r: 0.015\n</code></pre>\n\n<h3>Compute and Parallel Processing</h3>\n\n<p>For this competition, I used a preemptible 96 core 86 Gb RAM virtual machine from GCP. All together, I'm running 300 iterations * 11 z-shifts * 2 directions = 6600 iterations of DBSCAN and 6 rounds of track extension. Each event takes about 9-10 min when I use Python's multiprocessing.Pool to parallelize the iterations; After reading <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/62883\">cpmp's post</a> , I see that most people parallelize over the events of the test set to reduce overhead.</p>\n\n<h3>Things I Didn't Do</h3>\n\n<p>Look for better features; I didn't have the math expertise of the higher ranking competitors and didn't have the time at the end of the competition when the \" Criteria for Good Features\" discussion post was posted.</p>\n\n<p>Develop more sophisticated track merging; Focusing more on merging than z-shifting at first could have let me get a higher score out of my earlier features.</p>\n\n<p>Optimizing more hyperparameters; I didn't optimize the helix unrolling or track extension hyperparameters.</p>\n\n<p>My code is available at <a href=\"https://github.com/bkahn-github/TrackML\">this</a> GitHub repository</p>",
  "messages": [
    {
      "id": "369984",
      "postDate": "08/14/2018 03:30:05",
      "content": "<h1>54th Place Solution for bkKaggle team; 0.59229 private LB</h1>\n\n<h2>Third Kaggle competition and a second bronze medal!</h2>\n\n<p>Hi, I’m a 15 year old Kaggle beginner and this is my third Kaggle competition and second bronze medal! This competition was different from others since it wasn’t a straightforward supervised learning problem which made it challenging, and I wouldn't have achieved this score without all the ideas shared on the discussion forum. I know that my score doesn't compare to the winners's solutions, but I wanted to share it anyway.</p>\n\n<p>For most the first two months of the competition, I mostly focused on deep learning based solutions but found them hard to train and they took a long time. It was only in the last month of the competition that I focused on implementing an unsupervised clustering based solution.</p>\n\n<p>Some of my deep learning based ideas that I partially completed or didn’t attempt included: Using an mlp to find triplets, implementing a PointNet to cluster tracks, and creating word2vec style embeddings for each track and then clustering the embeddings. Unfortunately, I wasn’t able to get any of these approaches to work really well in the 2 months. After that, I shifted my focus to clustering based approaches.</p>\n\n<p>My final solution was a DBSCAN clustering and helix unrolling with z shifting and track extension.</p>\n\n<h3>Features</h3>\n\n<p>The features I ended up using were:</p>\n\n<p><code>cos(a), sin(a), z/rt, z/r, x/r, y/r</code> </p>\n\n<p>where </p>\n\n<p><code>a = arctan2(y, x) - arccos(mm * ii * rt)</code></p>\n\n<p><code>r = sqrt(x^2 + y^2 + z^2)</code></p>\n\n<p><code>rt = sqrt(x^2 + y^2)</code></p>\n\n<p>These features were from <a href=\"https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel\">Grzegorz's</a> and <a href=\"https://www.kaggle.com/khahuras/0-53x-clustering-using-hough-features-basic\">Kha A Vo's</a> kernels. Unfortunately, I didn’t have the time to find better features with the ideas shared in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/61590\">this</a> discussion post. I originally used <code>cos a, sin a, z/r, and z/rt</code> as my features; When merging based on track length, these features give you a score of 0.35 with no weights, and I couldn’t take it beyond 0.5 even after extensive Bayesian optimization. The addition of <code>x/r and y/r</code> and the default weights from the second kernel gives me a score of 0.54 without any z-shifting or track extension.</p>\n\n<h3>Merging</h3>\n\n<p>My merging strategy is very simple since I didn’t have time to improve it. I simply assign hits to the longest track with no more than 25 hits. </p>\n\n<h3>Z-shifting</h3>\n\n<p>I use 11 z-shifts uniformly distributed around the origin. I search +/- 5mm with steps of 1mm and the origin itself for a total of 11 z-shifts. The drawback of using so many z-shifts is that my running time increases by a factor of 11.</p>\n\n<h3>Track Extension</h3>\n\n<p>I use a modified version of track extension from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58194\">Heng's track extension post</a> . I found that for me, the optimum number of track extension rounds is 6. I wasn't able to use more because each further track extension gave diminishing increases in score and took an extra 1-2 minutes.</p>\n\n<h3>Hyperparameters</h3>\n\n<p>I used Bayesian optimization on top of the weights given in the second kernel. Also optimizing the number of DBSCAN iterations and it's epsilon hyperparameter further increased the score. My final hyperparameters are:</p>\n\n<pre><code>cos(aa) and sin(aa): 1.7\nz/rt: 0.8\nz/r: 0.2\nx/r: 0.015\ny/r: 0.015\n</code></pre>\n\n<h3>Compute and Parallel Processing</h3>\n\n<p>For this competition, I used a preemptible 96 core 86 Gb RAM virtual machine from GCP. All together, I'm running 300 iterations * 11 z-shifts * 2 directions = 6600 iterations of DBSCAN and 6 rounds of track extension. Each event takes about 9-10 min when I use Python's multiprocessing.Pool to parallelize the iterations; After reading <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/62883\">cpmp's post</a> , I see that most people parallelize over the events of the test set to reduce overhead.</p>\n\n<h3>Things I Didn't Do</h3>\n\n<p>Look for better features; I didn't have the math expertise of the higher ranking competitors and didn't have the time at the end of the competition when the \" Criteria for Good Features\" discussion post was posted.</p>\n\n<p>Develop more sophisticated track merging; Focusing more on merging than z-shifting at first could have let me get a higher score out of my earlier features.</p>\n\n<p>Optimizing more hyperparameters; I didn't optimize the helix unrolling or track extension hyperparameters.</p>\n\n<p>My code is available at <a href=\"https://github.com/bkahn-github/TrackML\">this</a> GitHub repository</p>",
      "rawMarkdown": "# 54th Place Solution for bkKaggle team; 0.59229 private LB \n\n## Third Kaggle competition and a second bronze medal!  \n\nHi, I’m a 15 year old Kaggle beginner and this is my third Kaggle competition and second bronze medal! This competition was different from others since it wasn’t a straightforward supervised learning problem which made it challenging, and I wouldn't have achieved this score without all the ideas shared on the discussion forum. I know that my score doesn't compare to the winners's solutions, but I wanted to share it anyway.\n\nFor most the first two months of the competition, I mostly focused on deep learning based solutions but found them hard to train and they took a long time. It was only in the last month of the competition that I focused on implementing an unsupervised clustering based solution.\n\nSome of my deep learning based ideas that I partially completed or didn’t attempt included: Using an mlp to find triplets, implementing a PointNet to cluster tracks, and creating word2vec style embeddings for each track and then clustering the embeddings. Unfortunately, I wasn’t able to get any of these approaches to work really well in the 2 months. After that, I shifted my focus to clustering based approaches.\n\nMy final solution was a DBSCAN clustering and helix unrolling with z shifting and track extension.\n\n### Features\n\nThe features I ended up using were:\n\n`cos(a), sin(a), z/rt, z/r, x/r, y/r` \n\nwhere \n\n`a = arctan2(y, x) - arccos(mm * ii * rt)`\n\n`r = sqrt(x^2 + y^2 + z^2)`\n\n`rt = sqrt(x^2 + y^2)`\n\nThese features were from [Grzegorz's] [1] and [Kha A Vo's] [2] kernels. Unfortunately, I didn’t have the time to find better features with the ideas shared in [this] [3] discussion post. I originally used `cos a, sin a, z/r, and z/rt` as my features; When merging based on track length, these features give you a score of 0.35 with no weights, and I couldn’t take it beyond 0.5 even after extensive Bayesian optimization. The addition of `x/r and y/r` and the default weights from the second kernel gives me a score of 0.54 without any z-shifting or track extension.\n\n### Merging\n\nMy merging strategy is very simple since I didn’t have time to improve it. I simply assign hits to the longest track with no more than 25 hits. \n\n### Z-shifting\n\nI use 11 z-shifts uniformly distributed around the origin. I search +/- 5mm with steps of 1mm and the origin itself for a total of 11 z-shifts. The drawback of using so many z-shifts is that my running time increases by a factor of 11.\n\n### Track Extension\n\nI use a modified version of track extension from [Heng's track extension post] [4] . I found that for me, the optimum number of track extension rounds is 6. I wasn't able to use more because each further track extension gave diminishing increases in score and took an extra 1-2 minutes.\n\n### Hyperparameters\n\nI used Bayesian optimization on top of the weights given in the second kernel. Also optimizing the number of DBSCAN iterations and it's epsilon hyperparameter further increased the score. My final hyperparameters are:\n\n    cos(aa) and sin(aa): 1.7\n    z/rt: 0.8\n    z/r: 0.2\n    x/r: 0.015\n    y/r: 0.015\n\n### Compute and Parallel Processing\n\nFor this competition, I used a preemptible 96 core 86 Gb RAM virtual machine from GCP. All together, I'm running 300 iterations * 11 z-shifts * 2 directions = 6600 iterations of DBSCAN and 6 rounds of track extension. Each event takes about 9-10 min when I use Python's multiprocessing.Pool to parallelize the iterations; After reading [cpmp's post] [5] , I see that most people parallelize over the events of the test set to reduce overhead.\n\n### Things I Didn't Do\n\nLook for better features; I didn't have the math expertise of the higher ranking competitors and didn't have the time at the end of the competition when the \" Criteria for Good Features\" discussion post was posted.\n\nDevelop more sophisticated track merging; Focusing more on merging than z-shifting at first could have let me get a higher score out of my earlier features.\n\nOptimizing more hyperparameters; I didn't optimize the helix unrolling or track extension hyperparameters.\n\nMy code is available at [this] [6] GitHub repository\n\n[1]: https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel\n[2]: https://www.kaggle.com/khahuras/0-53x-clustering-using-hough-features-basic\n[3]: https://www.kaggle.com/c/trackml-particle-identification/discussion/61590\n[4]: https://www.kaggle.com/c/trackml-particle-identification/discussion/58194\n[5]: https://www.kaggle.com/c/trackml-particle-identification/discussion/62883\n[6]: https://github.com/bkahn-github/TrackML",
      "votes": null
    },
    {
      "id": "369987",
      "postDate": "08/14/2018 03:50:25",
      "content": "<p>@Bilal ML  15 years old??? and already 2 bronze medals out of 3 competitions, that's awesome!! I don't know what I was doing at 15, and you managed to do well in a research competition many scientists haven been working on at the PhD-thesis level. Kudos to you! That's very inspiring, thank you for sharing.</p>",
      "rawMarkdown": "Bilal ML  15 years old??? and already 2 bronze medals out of 3 competitions, that's awesome!! I don't know what I was doing at 15, and you managed to do well in a research competition many scientists haven been working on at the PhD-thesis level. Kudos to you! That's very inspiring, thank you for sharing.",
      "votes": null
    },
    {
      "id": "370000",
      "postDate": "08/14/2018 04:25:14",
      "content": "<p>What do we got here? A talented teen Kaggler! At your age I also used the computer to research so much, almost every day and night. But what I researched is to how to beat the next boss of a computer game.</p>",
      "rawMarkdown": "What do we got here? A talented teen Kaggler! At your age I also used the computer to research so much, almost every day and night. But what I researched is to how to beat the next boss of a computer game.",
      "votes": null
    },
    {
      "id": "370004",
      "postDate": "08/14/2018 04:37:50",
      "content": "<p>Congrats on your Bronze..... Wish you all the best for your future competitions too...</p>",
      "rawMarkdown": "Congrats on your Bronze..... Wish you all the best for your future competitions too...",
      "votes": null
    },
    {
      "id": "370024",
      "postDate": "08/14/2018 05:28:29",
      "content": "<p>Congratulations the bkKaggle team and thanks for sharing your solution.</p>",
      "rawMarkdown": "Congratulations the bkKaggle team and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "370126",
      "postDate": "08/14/2018 10:23:15",
      "content": "<p>Congrats on the result, and thanks for sharing!</p>\n\n<blockquote>\n  <p>Hi, I’m a 15 year old Kaggle beginner</p>\n</blockquote>\n\n<p>Seems <a href=\"/anokas\">@anokas</a> has some serious competition coming ;)</p>",
      "rawMarkdown": "Congrats on the result, and thanks for sharing!\n\n&gt; Hi, I’m a 15 year old Kaggle beginner\n\nSeems @anokas has some serious competition coming ;)",
      "votes": null
    },
    {
      "id": "370131",
      "postDate": "08/14/2018 10:27:49",
      "content": "<p>Haha, congrats Bilal! Elegant solution :)</p>",
      "rawMarkdown": "Haha, congrats Bilal! Elegant solution :)",
      "votes": null
    },
    {
      "id": "370264",
      "postDate": "08/14/2018 14:16:45",
      "content": "<p>Well done! With regards to track extension, we ended up using 5 rounds, with a different limit for each round - we found the 0.02/0.04/0.06/0.08/0.10 combination worked best (achieved a better track extension score than using the same limit 5 times in a row). Did you change the limit for each round as well? Another thing we tried that proved to be beneficial (but not as beneficial as changing the limit) was to split a single larger angle delta up into smaller units when there were too many hits found - in our case, when there were over 2000 hits found in one angular slice, we split the results up into 4 separate slices. And, another simple tweak from Heng's base version is to look at the two track lengths before extending a track when the extension would end up taking that hit from another track - letting the longest-track win worked reasonably well there (and you could improve performance even more if you got more sophisticated than just longest-track wins, of course).</p>",
      "rawMarkdown": "Well done! With regards to track extension, we ended up using 5 rounds, with a different limit for each round - we found the 0.02/0.04/0.06/0.08/0.10 combination worked best (achieved a better track extension score than using the same limit 5 times in a row). Did you change the limit for each round as well? Another thing we tried that proved to be beneficial (but not as beneficial as changing the limit) was to split a single larger angle delta up into smaller units when there were too many hits found - in our case, when there were over 2000 hits found in one angular slice, we split the results up into 4 separate slices. And, another simple tweak from Heng's base version is to look at the two track lengths before extending a track when the extension would end up taking that hit from another track - letting the longest-track win worked reasonably well there (and you could improve performance even more if you got more sophisticated than just longest-track wins, of course).",
      "votes": null
    },
    {
      "id": "370268",
      "postDate": "08/14/2018 14:24:48",
      "content": "<p><strong>For track extension code</strong> Liam mentioned\n<a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py\">https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py</a></p>",
      "rawMarkdown": "**For track extension code** Liam mentioned\nhttps://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py",
      "votes": null
    },
    {
      "id": "378776",
      "postDate": "08/30/2018 14:47:20",
      "content": "<p>Hi,</p>\n\n<p>I used the same threshold of 0.1 for all my track extension rounds. Your way of gradually increasing the limit for extending tracks to first extend higher quality tracks then extend lower quality tracks makes intuitive sense, and is something I should have done. It seems like using different thresholds for track extension is a way to determine the quality of a track candidate which could then be used to do more sophisticated merging techniques. </p>\n\n<p>Would combining the first and third improvements in your comment by only extending a track if the extended track can pass through a lower threshold than the original track be a good way to prevent low quality extended tracks?</p>",
      "rawMarkdown": "Hi,\n\nI used the same threshold of 0.1 for all my track extension rounds. Your way of gradually increasing the limit for extending tracks to first extend higher quality tracks then extend lower quality tracks makes intuitive sense, and is something I should have done. It seems like using different thresholds for track extension is a way to determine the quality of a track candidate which could then be used to do more sophisticated merging techniques. \n\nWould combining the first and third improvements in your comment by only extending a track if the extended track can pass through a lower threshold than the original track be a good way to prevent low quality extended tracks?",
      "votes": null
    },
    {
      "id": "379343",
      "postDate": "08/31/2018 07:55:13",
      "content": "<p>Hmm.... I hadn't thought of that, but yeah, that could work well too - makes sense that ones with a lower threshold are more likely to be correct. I guess the hard part is that the track extension code only extends from each end of the target track, but when taking a hit away from another track, that hit could be in the middle. In which case, I guess you could look at the threshold coming from both sides and average them.</p>",
      "rawMarkdown": "Hmm.... I hadn't thought of that, but yeah, that could work well too - makes sense that ones with a lower threshold are more likely to be correct. I guess the hard part is that the track extension code only extends from each end of the target track, but when taking a hit away from another track, that hit could be in the middle. In which case, I guess you could look at the threshold coming from both sides and average them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 369987,
      "author_name": "nicolefinnie",
      "author_url": "",
      "post_date": "08/14/2018 03:50:25",
      "content": "<p>@Bilal ML  15 years old??? and already 2 bronze medals out of 3 competitions, that's awesome!! I don't know what I was doing at 15, and you managed to do well in a research competition many scientists haven been working on at the PhD-thesis level. Kudos to you! That's very inspiring, thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 370000,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "08/14/2018 04:25:14",
      "content": "<p>What do we got here? A talented teen Kaggler! At your age I also used the computer to research so much, almost every day and night. But what I researched is to how to beat the next boss of a computer game.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 370004,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "08/14/2018 04:37:50",
      "content": "<p>Congrats on your Bronze..... Wish you all the best for your future competitions too...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 370024,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "08/14/2018 05:28:29",
      "content": "<p>Congratulations the bkKaggle team and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 370126,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/14/2018 10:23:15",
      "content": "<p>Congrats on the result, and thanks for sharing!</p>\n\n<blockquote>\n  <p>Hi, I’m a 15 year old Kaggle beginner</p>\n</blockquote>\n\n<p>Seems <a href=\"/anokas\">@anokas</a> has some serious competition coming ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 370131,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "08/14/2018 10:27:49",
          "content": "<p>Haha, congrats Bilal! Elegant solution :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 370264,
      "author_name": "jliamfinnie",
      "author_url": "",
      "post_date": "08/14/2018 14:16:45",
      "content": "<p>Well done! With regards to track extension, we ended up using 5 rounds, with a different limit for each round - we found the 0.02/0.04/0.06/0.08/0.10 combination worked best (achieved a better track extension score than using the same limit 5 times in a row). Did you change the limit for each round as well? Another thing we tried that proved to be beneficial (but not as beneficial as changing the limit) was to split a single larger angle delta up into smaller units when there were too many hits found - in our case, when there were over 2000 hits found in one angular slice, we split the results up into 4 separate slices. And, another simple tweak from Heng's base version is to look at the two track lengths before extending a track when the extension would end up taking that hit from another track - letting the longest-track win worked reasonably well there (and you could improve performance even more if you got more sophisticated than just longest-track wins, of course).</p>",
      "votes": null,
      "replies": [
        {
          "id": 370268,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/14/2018 14:24:48",
          "content": "<p><strong>For track extension code</strong> Liam mentioned\n<a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py\">https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 378776,
          "author_name": "bkkaggle",
          "author_url": "",
          "post_date": "08/30/2018 14:47:20",
          "content": "<p>Hi,</p>\n\n<p>I used the same threshold of 0.1 for all my track extension rounds. Your way of gradually increasing the limit for extending tracks to first extend higher quality tracks then extend lower quality tracks makes intuitive sense, and is something I should have done. It seems like using different thresholds for track extension is a way to determine the quality of a track candidate which could then be used to do more sophisticated merging techniques. </p>\n\n<p>Would combining the first and third improvements in your comment by only extending a track if the extended track can pass through a lower threshold than the original track be a good way to prevent low quality extended tracks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 379343,
          "author_name": "jliamfinnie",
          "author_url": "",
          "post_date": "08/31/2018 07:55:13",
          "content": "<p>Hmm.... I hadn't thought of that, but yeah, that could work well too - makes sense that ones with a lower threshold are more likely to be correct. I guess the hard part is that the track extension code only extends from each end of the target track, but when taking a hit away from another track, that hit could be in the middle. In which case, I guess you could look at the threshold coming from both sides and average them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "369984": "# 54th Place Solution for bkKaggle team; 0.59229 private LB \n\n## Third Kaggle competition and a second bronze medal!  \n\nHi, I’m a 15 year old Kaggle beginner and this is my third Kaggle competition and second bronze medal! This competition was different from others since it wasn’t a straightforward supervised learning problem which made it challenging, and I wouldn't have achieved this score without all the ideas shared on the discussion forum. I know that my score doesn't compare to the winners's solutions, but I wanted to share it anyway.\n\nFor most the first two months of the competition, I mostly focused on deep learning based solutions but found them hard to train and they took a long time. It was only in the last month of the competition that I focused on implementing an unsupervised clustering based solution.\n\nSome of my deep learning based ideas that I partially completed or didn’t attempt included: Using an mlp to find triplets, implementing a PointNet to cluster tracks, and creating word2vec style embeddings for each track and then clustering the embeddings. Unfortunately, I wasn’t able to get any of these approaches to work really well in the 2 months. After that, I shifted my focus to clustering based approaches.\n\nMy final solution was a DBSCAN clustering and helix unrolling with z shifting and track extension.\n\n### Features\n\nThe features I ended up using were:\n\n`cos(a), sin(a), z/rt, z/r, x/r, y/r` \n\nwhere \n\n`a = arctan2(y, x) - arccos(mm * ii * rt)`\n\n`r = sqrt(x^2 + y^2 + z^2)`\n\n`rt = sqrt(x^2 + y^2)`\n\nThese features were from [Grzegorz's] [1] and [Kha A Vo's] [2] kernels. Unfortunately, I didn’t have the time to find better features with the ideas shared in [this] [3] discussion post. I originally used `cos a, sin a, z/r, and z/rt` as my features; When merging based on track length, these features give you a score of 0.35 with no weights, and I couldn’t take it beyond 0.5 even after extensive Bayesian optimization. The addition of `x/r and y/r` and the default weights from the second kernel gives me a score of 0.54 without any z-shifting or track extension.\n\n### Merging\n\nMy merging strategy is very simple since I didn’t have time to improve it. I simply assign hits to the longest track with no more than 25 hits. \n\n### Z-shifting\n\nI use 11 z-shifts uniformly distributed around the origin. I search +/- 5mm with steps of 1mm and the origin itself for a total of 11 z-shifts. The drawback of using so many z-shifts is that my running time increases by a factor of 11.\n\n### Track Extension\n\nI use a modified version of track extension from [Heng's track extension post] [4] . I found that for me, the optimum number of track extension rounds is 6. I wasn't able to use more because each further track extension gave diminishing increases in score and took an extra 1-2 minutes.\n\n### Hyperparameters\n\nI used Bayesian optimization on top of the weights given in the second kernel. Also optimizing the number of DBSCAN iterations and it's epsilon hyperparameter further increased the score. My final hyperparameters are:\n\n    cos(aa) and sin(aa): 1.7\n    z/rt: 0.8\n    z/r: 0.2\n    x/r: 0.015\n    y/r: 0.015\n\n### Compute and Parallel Processing\n\nFor this competition, I used a preemptible 96 core 86 Gb RAM virtual machine from GCP. All together, I'm running 300 iterations * 11 z-shifts * 2 directions = 6600 iterations of DBSCAN and 6 rounds of track extension. Each event takes about 9-10 min when I use Python's multiprocessing.Pool to parallelize the iterations; After reading [cpmp's post] [5] , I see that most people parallelize over the events of the test set to reduce overhead.\n\n### Things I Didn't Do\n\nLook for better features; I didn't have the math expertise of the higher ranking competitors and didn't have the time at the end of the competition when the \" Criteria for Good Features\" discussion post was posted.\n\nDevelop more sophisticated track merging; Focusing more on merging than z-shifting at first could have let me get a higher score out of my earlier features.\n\nOptimizing more hyperparameters; I didn't optimize the helix unrolling or track extension hyperparameters.\n\nMy code is available at [this] [6] GitHub repository\n\n[1]: https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel\n[2]: https://www.kaggle.com/khahuras/0-53x-clustering-using-hough-features-basic\n[3]: https://www.kaggle.com/c/trackml-particle-identification/discussion/61590\n[4]: https://www.kaggle.com/c/trackml-particle-identification/discussion/58194\n[5]: https://www.kaggle.com/c/trackml-particle-identification/discussion/62883\n[6]: https://github.com/bkahn-github/TrackML",
    "369987": "Bilal ML  15 years old??? and already 2 bronze medals out of 3 competitions, that's awesome!! I don't know what I was doing at 15, and you managed to do well in a research competition many scientists haven been working on at the PhD-thesis level. Kudos to you! That's very inspiring, thank you for sharing.",
    "370000": "What do we got here? A talented teen Kaggler! At your age I also used the computer to research so much, almost every day and night. But what I researched is to how to beat the next boss of a computer game.",
    "370004": "Congrats on your Bronze..... Wish you all the best for your future competitions too...",
    "370024": "Congratulations the bkKaggle team and thanks for sharing your solution.",
    "370126": "Congrats on the result, and thanks for sharing!\n\n&gt; Hi, I’m a 15 year old Kaggle beginner\n\nSeems @anokas has some serious competition coming ;)",
    "370131": "Haha, congrats Bilal! Elegant solution :)",
    "370264": "Well done! With regards to track extension, we ended up using 5 rounds, with a different limit for each round - we found the 0.02/0.04/0.06/0.08/0.10 combination worked best (achieved a better track extension score than using the same limit 5 times in a row). Did you change the limit for each round as well? Another thing we tried that proved to be beneficial (but not as beneficial as changing the limit) was to split a single larger angle delta up into smaller units when there were too many hits found - in our case, when there were over 2000 hits found in one angular slice, we split the results up into 4 separate slices. And, another simple tweak from Heng's base version is to look at the two track lengths before extending a track when the extension would end up taking that hit from another track - letting the longest-track win worked reasonably well there (and you could improve performance even more if you got more sophisticated than just longest-track wins, of course).",
    "370268": "**For track extension code** Liam mentioned\nhttps://github.com/jliamfinnie/kaggle-trackml/blob/master/src/extension.py",
    "378776": "Hi,\n\nI used the same threshold of 0.1 for all my track extension rounds. Your way of gradually increasing the limit for extending tracks to first extend higher quality tracks then extend lower quality tracks makes intuitive sense, and is something I should have done. It seems like using different thresholds for track extension is a way to determine the quality of a track candidate which could then be used to do more sophisticated merging techniques. \n\nWould combining the first and third improvements in your comment by only extending a track if the extended track can pass through a lower threshold than the original track be a good way to prevent low quality extended tracks?",
    "379343": "Hmm.... I hadn't thought of that, but yeah, that could work well too - makes sense that ones with a lower threshold are more likely to be correct. I guess the hard part is that the track extension code only extends from each end of the target track, but when taking a hit away from another track, that hit could be in the middle. In which case, I guess you could look at the threshold coming from both sides and average them."
  },
  "source": "meta"
}