{
  "id": 63244,
  "title": "Ensembling Helix 42 - #12 Solution",
  "url": "/competitions/trackml-particle-identification/discussion/63244",
  "author_name": "Liam Finnie",
  "post_date": "2018-08-14T00:10:20.030000",
  "votes": 25,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Why 42? That's the largest internal DBScan helix cluster ID we merged. If you include each z-shift as a separate model, we actually merge a total of 45 models. That's a lot of merging!</p>\n\n<p><a href=\"https://github.com/jliamfinnie/kaggle-trackml.git\">[You can find all our code in this github repository]</a>(<a href=\"https://github.com/jliamfinnie/kaggle-trackml.git\">https://github.com/jliamfinnie/kaggle-trackml.git</a>)</p>\n\n<p><strong>Non-mathematicians solution from second-time Kagglers</strong></p>\n\n<p>Nicole and I (Liam Finnie) started this Kaggle competition because it sounded pretty cool, however without a strong math or physics background, we quickly found ourselves at a disadvantage. So, we did what we know - write lots of code! Hopefully at least some of this will prove useful to someone, even as an example of 'what not to do!'.</p>\n\n<p>Our solution consists of many DBScan variants with different features, z-shifts, etc. For post-processing, we use heavy outlier removal (both hits and entire tracks) and track extension individually on each of the DBScan results. We then split each of the results into 3 categories - strong, medium, and weak - before merging them.</p>\n\n<p><strong>DBScan results</strong></p>\n\n<p>Many thanks to @Luis for providing us our base DBScan kernel. Nicole did most of the math work on our team to develop our clustering features. I can't do the math justice, so if you understand advanced helix math, check out <code>hits_clustering.py</code>, class <code>Clusterer</code>, method <code>dbscan()</code>. We used several of the features discussed in the forum such as z-shifts and sampled <code>r0</code> values, as well as some of our own tweaks. Our raw DBScan scores tended to mostly be in the range of 0.35 to 0.55.</p>\n\n<p><strong>Outlier removal</strong></p>\n\n<p>Outlier removal is tricky - it lowers your LB, however allows for much better merging later on. Approaches we used for outlier removal:</p>\n\n<ul>\n<li>use <code>z/r</code> to eliminate hits that are out-of-place</li>\n<li>look for hits with the exact same <code>z</code> value from the same <code>volume_id</code> and <code>layer_id</code>, remove one of them.</li>\n<li>calculate the slope between each pair of adjacent hits, remove hits whose slopes are very different. </li>\n</ul>\n\n<p>The outlier removal code entry point is in <code>merge.py</code> in function <code>remove_outliers()</code>.</p>\n\n<p><strong>Helix Track extension</strong></p>\n\n<p>Many thanks to @Heng who provided an initial track extension prototype. From this base, we added:</p>\n\n<ul>\n<li><code>z/r</code> to improve the KDTree clustering</li>\n<li>track scoring (length + quality of track) to determine when to steal hits from another track</li>\n<li>different number of KDTree neighbours, angle slices, etc. </li>\n</ul>\n\n<p>The track extension code can be found in the <code>extension.py</code> file, function <code>do_all_track_extensions()</code>. This type of track extension typically gave us a boost of between 0.05 and 0.15 for a single DBScan model.</p>\n\n<p><strong>Straight Track extension</strong></p>\n\n<p>Some tracks are more 'straight' than 'helix-like' - we do straight-track extension for track fragments from volumes 7 or 9. To extend straight tracks, we:</p>\n\n<ul>\n<li>compute <code>z/r</code> for each hit</li>\n<li>if our track does not have an entry in the adjacent <code>layer_id</code>, calculate the expected <code>z/r</code> for that adjacent <code>layer_id</code>, and assign any found hits to our track</li>\n<li>try to merge with track fragments from an adjacent <code>volume_id</code>.</li>\n</ul>\n\n<p>This type of track extension typically gave us a boost of between 0.01 and 0.02 for a single DBScan model. Code is in <code>straight_tracks.py</code>, function <code>extend_straight_tracks()</code>.</p>\n\n<p><strong>Merging</strong></p>\n\n<p>When merging clusters with different z-shifts, we found the order mattered a lot - for example, we could merge better with bigger jumps between successive z-shifts, i.e.the order (-6, 3, -3, 6) works better than (-6, -3, 3, 6).</p>\n\n<p>For our final merge at the end, we split each DBScan cluster into <strong>strong</strong>, <strong>medium</strong> and <strong>weak</strong> components based on the consistency of the <strong>helix curvature</strong>. Strong tracks are merged first, then medium, and finally weak ones at the end, getting more conservative at each step.</p>\n\n<p>The main problem with merging is how to tell whether two tracks are really the same, or should be separate? We tend to favour extending existing tracks when possible, but will create a new track if there is too little overlap with any existing track. Some rough pseudo-code for our merging heuristics:</p>\n\n<pre><code>foreach new_track in new_tracks:\n   if (no overlap with existing tracks)\n     assign new_track to merged results\n   elif (existing track is longer than new_track and includes all hits)\n     do nothing\n   else\n     determine longest overlapping track\n     if (longest overlapping track is track '0', i.e. unassigned hits)\n       consider second longest track for extension\n     if (too little overlap with existing longest overlapping track)\n       assign non-outlier hits from new_track to merged results\n     else\n       extend longest track to include non-outlier hits from new_track\n</code></pre>\n\n<p>Our merging code is in <code>merge.py</code>, function <code>heuristic_merge_tracks()</code>. We found simple merging ('longest-track-wins') hurts scores when there are more than 2 or 3 models, our current merging code was able to merge about 45 different sets of DBScan cluster results well.</p>\n\n<p><strong>Acknowledgement</strong></p>\n\n<p>Thanks to all Kagglers sharing in this competition, notably @Luis for the initial DBScan kernel, @Heng for the track-extension code, and @Yuval, @CPMP, @johnhsweeney, the chemist @Grzegorz, and many others for good discussions and DBScan feature suggestions.</p>",
  "messages": [
    {
      "id": 369896,
      "postDate": "2018-08-14T00:10:20.030Z",
      "content": "<p>Why 42? That's the largest internal DBScan helix cluster ID we merged. If you include each z-shift as a separate model, we actually merge a total of 45 models. That's a lot of merging!</p>\n\n<p><a href=\"https://github.com/jliamfinnie/kaggle-trackml.git\">[You can find all our code in this github repository]</a>(<a href=\"https://github.com/jliamfinnie/kaggle-trackml.git\">https://github.com/jliamfinnie/kaggle-trackml.git</a>)</p>\n\n<p><strong>Non-mathematicians solution from second-time Kagglers</strong></p>\n\n<p>Nicole and I (Liam Finnie) started this Kaggle competition because it sounded pretty cool, however without a strong math or physics background, we quickly found ourselves at a disadvantage. So, we did what we know - write lots of code! Hopefully at least some of this will prove useful to someone, even as an example of 'what not to do!'.</p>\n\n<p>Our solution consists of many DBScan variants with different features, z-shifts, etc. For post-processing, we use heavy outlier removal (both hits and entire tracks) and track extension individually on each of the DBScan results. We then split each of the results into 3 categories - strong, medium, and weak - before merging them.</p>\n\n<p><strong>DBScan results</strong></p>\n\n<p>Many thanks to @Luis for providing us our base DBScan kernel. Nicole did most of the math work on our team to develop our clustering features. I can't do the math justice, so if you understand advanced helix math, check out <code>hits_clustering.py</code>, class <code>Clusterer</code>, method <code>dbscan()</code>. We used several of the features discussed in the forum such as z-shifts and sampled <code>r0</code> values, as well as some of our own tweaks. Our raw DBScan scores tended to mostly be in the range of 0.35 to 0.55.</p>\n\n<p><strong>Outlier removal</strong></p>\n\n<p>Outlier removal is tricky - it lowers your LB, however allows for much better merging later on. Approaches we used for outlier removal:</p>\n\n<ul>\n<li>use <code>z/r</code> to eliminate hits that are out-of-place</li>\n<li>look for hits with the exact same <code>z</code> value from the same <code>volume_id</code> and <code>layer_id</code>, remove one of them.</li>\n<li>calculate the slope between each pair of adjacent hits, remove hits whose slopes are very different. </li>\n</ul>\n\n<p>The outlier removal code entry point is in <code>merge.py</code> in function <code>remove_outliers()</code>.</p>\n\n<p><strong>Helix Track extension</strong></p>\n\n<p>Many thanks to @Heng who provided an initial track extension prototype. From this base, we added:</p>\n\n<ul>\n<li><code>z/r</code> to improve the KDTree clustering</li>\n<li>track scoring (length + quality of track) to determine when to steal hits from another track</li>\n<li>different number of KDTree neighbours, angle slices, etc. </li>\n</ul>\n\n<p>The track extension code can be found in the <code>extension.py</code> file, function <code>do_all_track_extensions()</code>. This type of track extension typically gave us a boost of between 0.05 and 0.15 for a single DBScan model.</p>\n\n<p><strong>Straight Track extension</strong></p>\n\n<p>Some tracks are more 'straight' than 'helix-like' - we do straight-track extension for track fragments from volumes 7 or 9. To extend straight tracks, we:</p>\n\n<ul>\n<li>compute <code>z/r</code> for each hit</li>\n<li>if our track does not have an entry in the adjacent <code>layer_id</code>, calculate the expected <code>z/r</code> for that adjacent <code>layer_id</code>, and assign any found hits to our track</li>\n<li>try to merge with track fragments from an adjacent <code>volume_id</code>.</li>\n</ul>\n\n<p>This type of track extension typically gave us a boost of between 0.01 and 0.02 for a single DBScan model. Code is in <code>straight_tracks.py</code>, function <code>extend_straight_tracks()</code>.</p>\n\n<p><strong>Merging</strong></p>\n\n<p>When merging clusters with different z-shifts, we found the order mattered a lot - for example, we could merge better with bigger jumps between successive z-shifts, i.e.the order (-6, 3, -3, 6) works better than (-6, -3, 3, 6).</p>\n\n<p>For our final merge at the end, we split each DBScan cluster into <strong>strong</strong>, <strong>medium</strong> and <strong>weak</strong> components based on the consistency of the <strong>helix curvature</strong>. Strong tracks are merged first, then medium, and finally weak ones at the end, getting more conservative at each step.</p>\n\n<p>The main problem with merging is how to tell whether two tracks are really the same, or should be separate? We tend to favour extending existing tracks when possible, but will create a new track if there is too little overlap with any existing track. Some rough pseudo-code for our merging heuristics:</p>\n\n<pre><code>foreach new_track in new_tracks:\n   if (no overlap with existing tracks)\n     assign new_track to merged results\n   elif (existing track is longer than new_track and includes all hits)\n     do nothing\n   else\n     determine longest overlapping track\n     if (longest overlapping track is track '0', i.e. unassigned hits)\n       consider second longest track for extension\n     if (too little overlap with existing longest overlapping track)\n       assign non-outlier hits from new_track to merged results\n     else\n       extend longest track to include non-outlier hits from new_track\n</code></pre>\n\n<p>Our merging code is in <code>merge.py</code>, function <code>heuristic_merge_tracks()</code>. We found simple merging ('longest-track-wins') hurts scores when there are more than 2 or 3 models, our current merging code was able to merge about 45 different sets of DBScan cluster results well.</p>\n\n<p><strong>Acknowledgement</strong></p>\n\n<p>Thanks to all Kagglers sharing in this competition, notably @Luis for the initial DBScan kernel, @Heng for the track-extension code, and @Yuval, @CPMP, @johnhsweeney, the chemist @Grzegorz, and many others for good discussions and DBScan feature suggestions.</p>",
      "rawMarkdown": "\n\nWhy 42? That's the largest internal DBScan helix cluster ID we merged. If you include each z-shift as a separate model, we actually merge a total of 45 models. That's a lot of merging!\n\n[\\[You can find all our code in this github repository\\]][1](https://github.com/jliamfinnie/kaggle-trackml.git)\n\n**Non-mathematicians solution from second-time Kagglers**\n\nNicole and I (Liam Finnie) started this Kaggle competition because it sounded pretty cool, however without a strong math or physics background, we quickly found ourselves at a disadvantage. So, we did what we know - write lots of code! Hopefully at least some of this will prove useful to someone, even as an example of 'what not to do!'.\n\nOur solution consists of many DBScan variants with different features, z-shifts, etc. For post-processing, we use heavy outlier removal (both hits and entire tracks) and track extension individually on each of the DBScan results. We then split each of the results into 3 categories - strong, medium, and weak - before merging them.\n\n**DBScan results**\n\nMany thanks to @Luis for providing us our base DBScan kernel. Nicole did most of the math work on our team to develop our clustering features. I can't do the math justice, so if you understand advanced helix math, check out `hits_clustering.py`, class `Clusterer`, method `dbscan()`. We used several of the features discussed in the forum such as z-shifts and sampled `r0` values, as well as some of our own tweaks. Our raw DBScan scores tended to mostly be in the range of 0.35 to 0.55.\n\n**Outlier removal**\n\nOutlier removal is tricky - it lowers your LB, however allows for much better merging later on. Approaches we used for outlier removal:\n\n- use `z/r` to eliminate hits that are out-of-place\n- look for hits with the exact same `z` value from the same `volume_id` and `layer_id`, remove one of them.\n- calculate the slope between each pair of adjacent hits, remove hits whose slopes are very different. \n\nThe outlier removal code entry point is in `merge.py` in function `remove_outliers()`.\n\n**Helix Track extension**\n\nMany thanks to @Heng who provided an initial track extension prototype. From this base, we added:\n\n- `z/r` to improve the KDTree clustering\n- track scoring (length + quality of track) to determine when to steal hits from another track\n- different number of KDTree neighbours, angle slices, etc. \n\nThe track extension code can be found in the `extension.py` file, function `do_all_track_extensions()`. This type of track extension typically gave us a boost of between 0.05 and 0.15 for a single DBScan model.\n\n**Straight Track extension**\n\nSome tracks are more 'straight' than 'helix-like' - we do straight-track extension for track fragments from volumes 7 or 9. To extend straight tracks, we:\n\n- compute `z/r` for each hit\n- if our track does not have an entry in the adjacent `layer_id`, calculate the expected `z/r` for that adjacent `layer_id`, and assign any found hits to our track\n- try to merge with track fragments from an adjacent `volume_id`.\n\n    \nThis type of track extension typically gave us a boost of between 0.01 and 0.02 for a single DBScan model. Code is in `straight_tracks.py`, function `extend_straight_tracks()`.\n\n**Merging**\n\nWhen merging clusters with different z-shifts, we found the order mattered a lot - for example, we could merge better with bigger jumps between successive z-shifts, i.e.the order (-6, 3, -3, 6) works better than (-6, -3, 3, 6).\n\nFor our final merge at the end, we split each DBScan cluster into **strong**, **medium** and **weak** components based on the consistency of the **helix curvature**. Strong tracks are merged first, then medium, and finally weak ones at the end, getting more conservative at each step.\n\nThe main problem with merging is how to tell whether two tracks are really the same, or should be separate? We tend to favour extending existing tracks when possible, but will create a new track if there is too little overlap with any existing track. Some rough pseudo-code for our merging heuristics:\n\n    foreach new_track in new_tracks:\n       if (no overlap with existing tracks)\n         assign new_track to merged results\n       elif (existing track is longer than new_track and includes all hits)\n         do nothing\n       else\n         determine longest overlapping track\n         if (longest overlapping track is track '0', i.e. unassigned hits)\n           consider second longest track for extension\n         if (too little overlap with existing longest overlapping track)\n           assign non-outlier hits from new_track to merged results\n         else\n           extend longest track to include non-outlier hits from new_track\n\nOur merging code is in `merge.py`, function `heuristic_merge_tracks()`. We found simple merging ('longest-track-wins') hurts scores when there are more than 2 or 3 models, our current merging code was able to merge about 45 different sets of DBScan cluster results well.\n\n\n**Acknowledgement**\n\nThanks to all Kagglers sharing in this competition, notably @Luis for the initial DBScan kernel, @Heng for the track-extension code, and @Yuval, @CPMP, @johnhsweeney, the chemist @Grzegorz, and many others for good discussions and DBScan feature suggestions.\n\n\n  [1]: https://github.com/jliamfinnie/kaggle-trackml.git",
      "votes": 25
    },
    {
      "id": 369938,
      "postDate": "2018-08-14T00:56:11.040Z",
      "content": "<p>And I'm sharing the main DBSCAN helix approach and all the references we used as promised. :)</p>\n\n<h2>Generate the hidden feature - helix radii as input</h2>\n\n<ul>\n<li><p>I find the most tricky part doing clustering in this competition is simulating the range of helix radii. We've tried different distributions, such as linear distribution, Gaussian distribution, but the most effective way is to generate real helix radii from the train data. I stole @Heng's code from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">this post</a>.</p></li>\n<li><p>You can find <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/generate_radii_samples.ipynb\">my notebook</a> which generates helix radii from the train data or you can download pre-generated radii named after their event id <a href=\"https://github.com/nicolefinnie/kaggle-trackml/tree/master/input/r0_list\">here</a></p></li>\n</ul>\n\n<h2>Helix unrolling function</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369901/10051/helix_unrolling.png\" alt=\"helix unrolling\"></p>\n\n<h3>z-axis centered tracks, tracks crossing (x,y) = (0,0)</h3>\n\n<ul>\n<li><p>The closet approach <code>D0 = 0</code>, the tracks start from close to <code>(0,0,z)</code>.</p></li>\n<li><p>Accurate version - reach <code>score=0.5</code> within 1 minute with 40 radius samples. To get a higher accuracy, you need to run more radius samples and it can take much longer. </p></li>\n</ul>\n\nThe track can go in either  direction, and <code>theta0</code> should be constant for all hits on the same  track in a perfect helix form.\n\n<pre><code>dfh['cos_theta'] = dfh.r/2/r0\nif ii &lt; r0_list.shape[0]/2:\n    dfh['theta0'] = dfh['phi'] - np.arccos(dfh['cos_theta'])\nelse:\n    dfh['theta0'] = dfh['phi'] + np.arccos(dfh['cos_theta'])\n</code></pre>\n\n<ul>\n<li>Self-made version - reach <code>score=0.5</code> in 2 minutes but it rarely go above 0.5 since the unrolling function is not accurate. The reason why we use this self-made version is that it can find different tracks for later merging, which is good.</li>\n</ul>\n\nThis tries to find possible <code>theta0</code> using an approximation function <code>ii</code> from -120 to 120\n\n<pre><code>STEPRR = 0.03\nrr = dfh.r/1000\ndfh['theta0'] = dfh['phi'] + (rr + STEPRR*rr**2)*ii/180*np.pi + (0.00001*ii)*dfh.z*np.sign(dfh.z)/180*np.pi\n</code></pre>\n\n<h3>Main features that are constant in a perfect helix form</h3>\n\n<ul>\n<li><code>sin(theta0)</code></li>\n<li><code>cos(theta0)</code></li>\n<li><p><code>(z-z0)/arc</code></p></li>\n<li><p>The problem is <code>z/arc</code> is still uneven in the magnetic field, so I've been trying to improve this problem by using following features in different models. Other Kagglers definitely have a more accurate equation. </p></li>\n<li><code>log(1 + abs((z-z0)/arc))*sign(z)</code></li>\n<li><code>(z-z0)/arc*sqrt(sin(arctan2(r,z-z0)))</code> I use this square root sine function as a correction term for the azimuthal angle on the x-y plane projection</li>\n</ul>\n\n<h3>Main features that are often constant</h3>\n\n<ul>\n<li><code>(z-z0)/r</code> where <code>r</code> is the Euclidean distance from the hit to the origin.</li>\n<li><code>log(1 + abs((z-z0)/r))*sign(z)</code> is an approach to get z values closer to the origin to improve the problem with uneven <code>z/r</code> values</li>\n</ul>\n\n<h3>Side features</h3>\n\n<ul>\n<li>Those are often not constant but we can find different tracks using them with small weights when we cluster</li>\n<li><code>x/d</code> where <code>d</code> is the Eucliean distance from the hit to the origin in 3D <code>(x**2+y**2+z**2)**0.5</code></li>\n<li><code>y/d</code></li>\n<li><code>arctan2(z-r0,r)</code></li>\n<li><code>px, py</code> in my code: <code>-r*cos(theta0)*cos(phi)-r*sin(theta0)*sin(phi)</code> and <code>-r*cos(theta0)*sin(phi)+r*sin(theta0)*cos(phi)</code>. I happened to find this feature that can find the seeds of non-z-axis centered tracks and we can extend the tracks using the found seeds. </li>\n</ul>\n\n<h3>Non-z-axis centered tracks</h3>\n\n<ul>\n<li>Skip this part since we didn't have time to implement it. Add the closest approach <code>D0</code> to your equations, you can find full equations with <code>D0</code> and <code>D1</code> from <a href=\"http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf\">Full helix equations for ATLAS</a></li>\n</ul>\n\n<h2>LSTM approach for track fitting</h2>\n\n<ul>\n<li>My <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/train_LSTM.ipynb\">notebook</a> including visualization</li>\n<li>We take first five hits as a seeded track and predict next 5 hits. I believe this has its potential as a track validation tool if right features were trained. I shared the detail in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60455#352645\">this post</a></li>\n</ul>\n\n<h2>pointNet approach</h2>\n\n<ul>\n<li>PointNet is a lightweight CNN that can be used for pixel level classification(segementation) or classification. I put experimental code <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/train_pointnet.py\">here</a>, the challenge is to generate the right train data.</li>\n</ul>\n\n<h2>Background knowledge</h2>\n\n<ul>\n<li><p><a href=\"http://ific.uv.es/~nebot/IDPASC/Material/Tracking-Vertexing/Tracking-Vertexing-Slides.pdf\">Very good slides for beginners</a></p></li>\n<li><p><a href=\"http://www.physics.iitm.ac.in/~sercehep2013/track2_Gagan_Mohanty.pdf\">Lecture of particles tracking</a></p></li>\n<li><p><a href=\"http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf\">Full helix equations for ATLAS</a> - All equations you need!</p></li>\n<li><p><a href=\"http://physik.uibk.ac.at/hephy/theses/dipl_as.pdf\">Diplom thesis</a> of Andreas Salzburger (Wow, he started in this field as a CERN student already in 2001 :p )</p></li>\n<li><p><a href=\"http://physik.uibk.ac.at/hephy/theses/diss_as.pdf\">Doctor thesis</a> of Andreas Salzburger</p></li>\n<li><p><a href=\"https://gitlab.cern.ch/acts/acts-core\">CERN tracking software Acts</a> - Sadly, we didn't have time to explore it :) </p></li>\n</ul>",
      "rawMarkdown": "And I'm sharing the main DBSCAN helix approach and all the references we used as promised. :)\n\n## Generate the hidden feature - helix radii as input \n\n* I find the most tricky part doing clustering in this competition is simulating the range of helix radii. We've tried different distributions, such as linear distribution, Gaussian distribution, but the most effective way is to generate real helix radii from the train data. I stole @Heng's code from [this post](https://www.kaggle.com/c/trackml-particle-identification/discussion/57643).\n \n* You can find [my notebook](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/generate_radii_samples.ipynb) which generates helix radii from the train data or you can download pre-generated radii named after their event id [here](https://github.com/nicolefinnie/kaggle-trackml/tree/master/input/r0_list)\n\n## Helix unrolling function\n\n![helix unrolling][1]\n[1]: https://storage.googleapis.com/kaggle-forum-message-attachments/369901/10051/helix_unrolling.png\n\n\n### z-axis centered tracks, tracks crossing (x,y) = (0,0)\n* The closet approach `D0 = 0`, the tracks start from close to `(0,0,z)`.\n\n* Accurate version - reach `score=0.5` within 1 minute with 40 radius samples. To get a higher accuracy, you need to run more radius samples and it can take much longer. \n\n\n#### The track can go in either  direction, and `theta0` should be constant for all hits on the same  track in a perfect helix form.\n\n\n    dfh['cos_theta'] = dfh.r/2/r0\n    if ii &lt; r0_list.shape[0]/2:\n        dfh['theta0'] = dfh['phi'] - np.arccos(dfh['cos_theta'])\n    else:\n        dfh['theta0'] = dfh['phi'] + np.arccos(dfh['cos_theta'])\n\n\n* Self-made version - reach `score=0.5` in 2 minutes but it rarely go above 0.5 since the unrolling function is not accurate. The reason why we use this self-made version is that it can find different tracks for later merging, which is good.\n\n####This tries to find possible `theta0` using an approximation function `ii` from -120 to 120\n\n\n    STEPRR = 0.03\n    rr = dfh.r/1000\n    dfh['theta0'] = dfh['phi'] + (rr + STEPRR*rr**2)*ii/180*np.pi + (0.00001*ii)*dfh.z*np.sign(dfh.z)/180*np.pi\n\n\n###Main features that are constant in a perfect helix form\n* `sin(theta0)`\n* `cos(theta0)`\n* `(z-z0)/arc`\n\n* The problem is `z/arc` is still uneven in the magnetic field, so I've been trying to improve this problem by using following features in different models. Other Kagglers definitely have a more accurate equation. \n* `log(1 + abs((z-z0)/arc))*sign(z)`\n* `(z-z0)/arc*sqrt(sin(arctan2(r,z-z0)))` I use this square root sine function as a correction term for the azimuthal angle on the x-y plane projection\n\n###Main features that are often constant\n* `(z-z0)/r` where `r` is the Euclidean distance from the hit to the origin.\n* `log(1 + abs((z-z0)/r))*sign(z)` is an approach to get z values closer to the origin to improve the problem with uneven `z/r` values\n\n###Side features\n* Those are often not constant but we can find different tracks using them with small weights when we cluster\n* `x/d` where `d` is the Eucliean distance from the hit to the origin in 3D `(x**2+y**2+z**2)**0.5`\n* `y/d`\n* `arctan2(z-r0,r)`\n* `px, py` in my code: `-r*cos(theta0)*cos(phi)-r*sin(theta0)*sin(phi)` and `-r*cos(theta0)*sin(phi)+r*sin(theta0)*cos(phi)`. I happened to find this feature that can find the seeds of non-z-axis centered tracks and we can extend the tracks using the found seeds. \n\n###Non-z-axis centered tracks\n* Skip this part since we didn't have time to implement it. Add the closest approach `D0` to your equations, you can find full equations with `D0` and `D1` from [Full helix equations for ATLAS](http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf)\n\n\n## LSTM approach for track fitting\n* My [notebook](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/train_LSTM.ipynb) including visualization\n* We take first five hits as a seeded track and predict next 5 hits. I believe this has its potential as a track validation tool if right features were trained. I shared the detail in [this post](https://www.kaggle.com/c/trackml-particle-identification/discussion/60455#352645)\n\n## pointNet approach \n* PointNet is a lightweight CNN that can be used for pixel level classification(segementation) or classification. I put experimental code [here](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/train_pointnet.py), the challenge is to generate the right train data.\n\n\n## Background knowledge \n\n* [Very good slides for beginners](http://ific.uv.es/~nebot/IDPASC/Material/Tracking-Vertexing/Tracking-Vertexing-Slides.pdf)\n\n* [Lecture of particles tracking](http://www.physics.iitm.ac.in/~sercehep2013/track2_Gagan_Mohanty.pdf)\n\n\n* [Full helix equations for ATLAS](http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf) - All equations you need!\n\n\n* [Diplom thesis](http://physik.uibk.ac.at/hephy/theses/dipl_as.pdf) of Andreas Salzburger (Wow, he started in this field as a CERN student already in 2001 :p )\n\n* [Doctor thesis](http://physik.uibk.ac.at/hephy/theses/diss_as.pdf) of Andreas Salzburger\n\n* [CERN tracking software Acts](https://gitlab.cern.ch/acts/acts-core) - Sadly, we didn't have time to explore it :) \n\n\n",
      "votes": 7,
      "replies": [
        {
          "id": 370284,
          "postDate": "2018-08-14T15:06:09.250Z",
          "content": "<p>Oh - you made me feel old now ... :-) </p>\n\n<p>Thanks for participating and I hope you had fun in the challenge!!</p>",
          "rawMarkdown": "Oh - you made me feel old now ... :-) \n\nThanks for participating and I hope you had fun in the challenge!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 370015,
      "postDate": "2018-08-14T05:04:59.670Z",
      "content": "<p>Congratulations and thanks so much @Nicole and @Liam for sharing your code and all that you have shared in the discussion threads. It is an amazing approach to this problem. </p>\n\n<p>I will need time to go through your code later as now I have to run over to what is a messed up Santander competition to try my best in the remaining 6 days.</p>",
      "rawMarkdown": "Congratulations and thanks so much @Nicole and @Liam for sharing your code and all that you have shared in the discussion threads. It is an amazing approach to this problem. \n\nI will need time to go through your code later as now I have to run over to what is a messed up Santander competition to try my best in the remaining 6 days.",
      "votes": 1
    },
    {
      "id": 369909,
      "postDate": "2018-08-14T00:27:37.440Z",
      "content": "<p>Congrats. Sad to see you guys missed the gold. Thank you for the solution.</p>",
      "rawMarkdown": "Congrats. Sad to see you guys missed the gold. Thank you for the solution.",
      "votes": 1
    },
    {
      "id": 369908,
      "postDate": "2018-08-14T00:27:11.710Z",
      "content": "<p>Congrats and thanks for sharing your approach and code! Will be helpful in figuring out where I went wrong trying to find better features.</p>",
      "rawMarkdown": "Congrats and thanks for sharing your approach and code! Will be helpful in figuring out where I went wrong trying to find better features.",
      "votes": 1
    },
    {
      "id": 369900,
      "postDate": "2018-08-14T00:15:40.857Z",
      "content": "<p>Congrats and sorry for your miss of gold by one.  I won't say I hope someone cheated in front of you, as I don't think it happened, but I thought of it still ;)</p>\n\n<p>Thanks for the writeup.</p>",
      "rawMarkdown": "Congrats and sorry for your miss of gold by one.  I won't say I hope someone cheated in front of you, as I don't think it happened, but I thought of it still ;)\n\nThanks for the writeup.\n",
      "votes": 1
    },
    {
      "id": 370008,
      "postDate": "2018-08-14T04:51:27.407Z",
      "content": "<p>How can we forget to thank @Kha too, he spawned the discussion thread with @yuval together and made us realise we did it all wrong. 😉 I wish you shared it earlier. </p>",
      "rawMarkdown": "How can we forget to thank @Kha too, he spawned the discussion thread with @yuval together and made us realise we did it all wrong. 😉 I wish you shared it earlier. ",
      "votes": 2,
      "replies": [
        {
          "id": 370137,
          "postDate": "2018-08-14T10:32:44.170Z",
          "content": "<p>Thanks Nicole.</p>",
          "rawMarkdown": "Thanks Nicole."
        }
      ]
    },
    {
      "id": 369911,
      "postDate": "2018-08-14T00:29:28.710Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 369901,
      "postDate": "2018-08-14T00:16:18.800Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 369938,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-08-14T00:56:11.040000",
      "content": "<p>And I'm sharing the main DBSCAN helix approach and all the references we used as promised. :)</p>\n\n<h2>Generate the hidden feature - helix radii as input</h2>\n\n<ul>\n<li><p>I find the most tricky part doing clustering in this competition is simulating the range of helix radii. We've tried different distributions, such as linear distribution, Gaussian distribution, but the most effective way is to generate real helix radii from the train data. I stole @Heng's code from <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">this post</a>.</p></li>\n<li><p>You can find <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/generate_radii_samples.ipynb\">my notebook</a> which generates helix radii from the train data or you can download pre-generated radii named after their event id <a href=\"https://github.com/nicolefinnie/kaggle-trackml/tree/master/input/r0_list\">here</a></p></li>\n</ul>\n\n<h2>Helix unrolling function</h2>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369901/10051/helix_unrolling.png\" alt=\"helix unrolling\"></p>\n\n<h3>z-axis centered tracks, tracks crossing (x,y) = (0,0)</h3>\n\n<ul>\n<li><p>The closet approach <code>D0 = 0</code>, the tracks start from close to <code>(0,0,z)</code>.</p></li>\n<li><p>Accurate version - reach <code>score=0.5</code> within 1 minute with 40 radius samples. To get a higher accuracy, you need to run more radius samples and it can take much longer. </p></li>\n</ul>\n\nThe track can go in either  direction, and <code>theta0</code> should be constant for all hits on the same  track in a perfect helix form.\n\n<pre><code>dfh['cos_theta'] = dfh.r/2/r0\nif ii &lt; r0_list.shape[0]/2:\n    dfh['theta0'] = dfh['phi'] - np.arccos(dfh['cos_theta'])\nelse:\n    dfh['theta0'] = dfh['phi'] + np.arccos(dfh['cos_theta'])\n</code></pre>\n\n<ul>\n<li>Self-made version - reach <code>score=0.5</code> in 2 minutes but it rarely go above 0.5 since the unrolling function is not accurate. The reason why we use this self-made version is that it can find different tracks for later merging, which is good.</li>\n</ul>\n\nThis tries to find possible <code>theta0</code> using an approximation function <code>ii</code> from -120 to 120\n\n<pre><code>STEPRR = 0.03\nrr = dfh.r/1000\ndfh['theta0'] = dfh['phi'] + (rr + STEPRR*rr**2)*ii/180*np.pi + (0.00001*ii)*dfh.z*np.sign(dfh.z)/180*np.pi\n</code></pre>\n\n<h3>Main features that are constant in a perfect helix form</h3>\n\n<ul>\n<li><code>sin(theta0)</code></li>\n<li><code>cos(theta0)</code></li>\n<li><p><code>(z-z0)/arc</code></p></li>\n<li><p>The problem is <code>z/arc</code> is still uneven in the magnetic field, so I've been trying to improve this problem by using following features in different models. Other Kagglers definitely have a more accurate equation. </p></li>\n<li><code>log(1 + abs((z-z0)/arc))*sign(z)</code></li>\n<li><code>(z-z0)/arc*sqrt(sin(arctan2(r,z-z0)))</code> I use this square root sine function as a correction term for the azimuthal angle on the x-y plane projection</li>\n</ul>\n\n<h3>Main features that are often constant</h3>\n\n<ul>\n<li><code>(z-z0)/r</code> where <code>r</code> is the Euclidean distance from the hit to the origin.</li>\n<li><code>log(1 + abs((z-z0)/r))*sign(z)</code> is an approach to get z values closer to the origin to improve the problem with uneven <code>z/r</code> values</li>\n</ul>\n\n<h3>Side features</h3>\n\n<ul>\n<li>Those are often not constant but we can find different tracks using them with small weights when we cluster</li>\n<li><code>x/d</code> where <code>d</code> is the Eucliean distance from the hit to the origin in 3D <code>(x**2+y**2+z**2)**0.5</code></li>\n<li><code>y/d</code></li>\n<li><code>arctan2(z-r0,r)</code></li>\n<li><code>px, py</code> in my code: <code>-r*cos(theta0)*cos(phi)-r*sin(theta0)*sin(phi)</code> and <code>-r*cos(theta0)*sin(phi)+r*sin(theta0)*cos(phi)</code>. I happened to find this feature that can find the seeds of non-z-axis centered tracks and we can extend the tracks using the found seeds. </li>\n</ul>\n\n<h3>Non-z-axis centered tracks</h3>\n\n<ul>\n<li>Skip this part since we didn't have time to implement it. Add the closest approach <code>D0</code> to your equations, you can find full equations with <code>D0</code> and <code>D1</code> from <a href=\"http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf\">Full helix equations for ATLAS</a></li>\n</ul>\n\n<h2>LSTM approach for track fitting</h2>\n\n<ul>\n<li>My <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/train_LSTM.ipynb\">notebook</a> including visualization</li>\n<li>We take first five hits as a seeded track and predict next 5 hits. I believe this has its potential as a track validation tool if right features were trained. I shared the detail in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60455#352645\">this post</a></li>\n</ul>\n\n<h2>pointNet approach</h2>\n\n<ul>\n<li>PointNet is a lightweight CNN that can be used for pixel level classification(segementation) or classification. I put experimental code <a href=\"https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/train_pointnet.py\">here</a>, the challenge is to generate the right train data.</li>\n</ul>\n\n<h2>Background knowledge</h2>\n\n<ul>\n<li><p><a href=\"http://ific.uv.es/~nebot/IDPASC/Material/Tracking-Vertexing/Tracking-Vertexing-Slides.pdf\">Very good slides for beginners</a></p></li>\n<li><p><a href=\"http://www.physics.iitm.ac.in/~sercehep2013/track2_Gagan_Mohanty.pdf\">Lecture of particles tracking</a></p></li>\n<li><p><a href=\"http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf\">Full helix equations for ATLAS</a> - All equations you need!</p></li>\n<li><p><a href=\"http://physik.uibk.ac.at/hephy/theses/dipl_as.pdf\">Diplom thesis</a> of Andreas Salzburger (Wow, he started in this field as a CERN student already in 2001 :p )</p></li>\n<li><p><a href=\"http://physik.uibk.ac.at/hephy/theses/diss_as.pdf\">Doctor thesis</a> of Andreas Salzburger</p></li>\n<li><p><a href=\"https://gitlab.cern.ch/acts/acts-core\">CERN tracking software Acts</a> - Sadly, we didn't have time to explore it :) </p></li>\n</ul>",
      "votes": 7,
      "replies": [
        {
          "id": 370284,
          "author_name": "Andreas Salzburger",
          "author_url": "",
          "post_date": "2018-08-14T15:06:09.250000",
          "content": "<p>Oh - you made me feel old now ... :-) </p>\n\n<p>Thanks for participating and I hope you had fun in the challenge!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 370015,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-08-14T05:04:59.670000",
      "content": "<p>Congratulations and thanks so much @Nicole and @Liam for sharing your code and all that you have shared in the discussion threads. It is an amazing approach to this problem. </p>\n\n<p>I will need time to go through your code later as now I have to run over to what is a messed up Santander competition to try my best in the remaining 6 days.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 369909,
      "author_name": "Akila Wajirasena",
      "author_url": "",
      "post_date": "2018-08-14T00:27:37.440000",
      "content": "<p>Congrats. Sad to see you guys missed the gold. Thank you for the solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 369908,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2018-08-14T00:27:11.710000",
      "content": "<p>Congrats and thanks for sharing your approach and code! Will be helpful in figuring out where I went wrong trying to find better features.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 369900,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-14T00:15:40.857000",
      "content": "<p>Congrats and sorry for your miss of gold by one.  I won't say I hope someone cheated in front of you, as I don't think it happened, but I thought of it still ;)</p>\n\n<p>Thanks for the writeup.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 370008,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-08-14T04:51:27.407000",
      "content": "<p>How can we forget to thank @Kha too, he spawned the discussion thread with @yuval together and made us realise we did it all wrong. 😉 I wish you shared it earlier. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 370137,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-08-14T10:32:44.170000",
          "content": "<p>Thanks Nicole.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 369911,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-14T00:29:28.710000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 369901,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-14T00:16:18.800000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "369896": "\n\nWhy 42? That's the largest internal DBScan helix cluster ID we merged. If you include each z-shift as a separate model, we actually merge a total of 45 models. That's a lot of merging!\n\n[\\[You can find all our code in this github repository\\]][1](https://github.com/jliamfinnie/kaggle-trackml.git)\n\n**Non-mathematicians solution from second-time Kagglers**\n\nNicole and I (Liam Finnie) started this Kaggle competition because it sounded pretty cool, however without a strong math or physics background, we quickly found ourselves at a disadvantage. So, we did what we know - write lots of code! Hopefully at least some of this will prove useful to someone, even as an example of 'what not to do!'.\n\nOur solution consists of many DBScan variants with different features, z-shifts, etc. For post-processing, we use heavy outlier removal (both hits and entire tracks) and track extension individually on each of the DBScan results. We then split each of the results into 3 categories - strong, medium, and weak - before merging them.\n\n**DBScan results**\n\nMany thanks to @Luis for providing us our base DBScan kernel. Nicole did most of the math work on our team to develop our clustering features. I can't do the math justice, so if you understand advanced helix math, check out `hits_clustering.py`, class `Clusterer`, method `dbscan()`. We used several of the features discussed in the forum such as z-shifts and sampled `r0` values, as well as some of our own tweaks. Our raw DBScan scores tended to mostly be in the range of 0.35 to 0.55.\n\n**Outlier removal**\n\nOutlier removal is tricky - it lowers your LB, however allows for much better merging later on. Approaches we used for outlier removal:\n\n- use `z/r` to eliminate hits that are out-of-place\n- look for hits with the exact same `z` value from the same `volume_id` and `layer_id`, remove one of them.\n- calculate the slope between each pair of adjacent hits, remove hits whose slopes are very different. \n\nThe outlier removal code entry point is in `merge.py` in function `remove_outliers()`.\n\n**Helix Track extension**\n\nMany thanks to @Heng who provided an initial track extension prototype. From this base, we added:\n\n- `z/r` to improve the KDTree clustering\n- track scoring (length + quality of track) to determine when to steal hits from another track\n- different number of KDTree neighbours, angle slices, etc. \n\nThe track extension code can be found in the `extension.py` file, function `do_all_track_extensions()`. This type of track extension typically gave us a boost of between 0.05 and 0.15 for a single DBScan model.\n\n**Straight Track extension**\n\nSome tracks are more 'straight' than 'helix-like' - we do straight-track extension for track fragments from volumes 7 or 9. To extend straight tracks, we:\n\n- compute `z/r` for each hit\n- if our track does not have an entry in the adjacent `layer_id`, calculate the expected `z/r` for that adjacent `layer_id`, and assign any found hits to our track\n- try to merge with track fragments from an adjacent `volume_id`.\n\n    \nThis type of track extension typically gave us a boost of between 0.01 and 0.02 for a single DBScan model. Code is in `straight_tracks.py`, function `extend_straight_tracks()`.\n\n**Merging**\n\nWhen merging clusters with different z-shifts, we found the order mattered a lot - for example, we could merge better with bigger jumps between successive z-shifts, i.e.the order (-6, 3, -3, 6) works better than (-6, -3, 3, 6).\n\nFor our final merge at the end, we split each DBScan cluster into **strong**, **medium** and **weak** components based on the consistency of the **helix curvature**. Strong tracks are merged first, then medium, and finally weak ones at the end, getting more conservative at each step.\n\nThe main problem with merging is how to tell whether two tracks are really the same, or should be separate? We tend to favour extending existing tracks when possible, but will create a new track if there is too little overlap with any existing track. Some rough pseudo-code for our merging heuristics:\n\n    foreach new_track in new_tracks:\n       if (no overlap with existing tracks)\n         assign new_track to merged results\n       elif (existing track is longer than new_track and includes all hits)\n         do nothing\n       else\n         determine longest overlapping track\n         if (longest overlapping track is track '0', i.e. unassigned hits)\n           consider second longest track for extension\n         if (too little overlap with existing longest overlapping track)\n           assign non-outlier hits from new_track to merged results\n         else\n           extend longest track to include non-outlier hits from new_track\n\nOur merging code is in `merge.py`, function `heuristic_merge_tracks()`. We found simple merging ('longest-track-wins') hurts scores when there are more than 2 or 3 models, our current merging code was able to merge about 45 different sets of DBScan cluster results well.\n\n\n**Acknowledgement**\n\nThanks to all Kagglers sharing in this competition, notably @Luis for the initial DBScan kernel, @Heng for the track-extension code, and @Yuval, @CPMP, @johnhsweeney, the chemist @Grzegorz, and many others for good discussions and DBScan feature suggestions.\n\n\n  [1]: https://github.com/jliamfinnie/kaggle-trackml.git",
    "369938": "And I'm sharing the main DBSCAN helix approach and all the references we used as promised. :)\n\n## Generate the hidden feature - helix radii as input \n\n* I find the most tricky part doing clustering in this competition is simulating the range of helix radii. We've tried different distributions, such as linear distribution, Gaussian distribution, but the most effective way is to generate real helix radii from the train data. I stole @Heng's code from [this post](https://www.kaggle.com/c/trackml-particle-identification/discussion/57643).\n \n* You can find [my notebook](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/generate_radii_samples.ipynb) which generates helix radii from the train data or you can download pre-generated radii named after their event id [here](https://github.com/nicolefinnie/kaggle-trackml/tree/master/input/r0_list)\n\n## Helix unrolling function\n\n![helix unrolling][1]\n[1]: https://storage.googleapis.com/kaggle-forum-message-attachments/369901/10051/helix_unrolling.png\n\n\n### z-axis centered tracks, tracks crossing (x,y) = (0,0)\n* The closet approach `D0 = 0`, the tracks start from close to `(0,0,z)`.\n\n* Accurate version - reach `score=0.5` within 1 minute with 40 radius samples. To get a higher accuracy, you need to run more radius samples and it can take much longer. \n\n\n#### The track can go in either  direction, and `theta0` should be constant for all hits on the same  track in a perfect helix form.\n\n\n    dfh['cos_theta'] = dfh.r/2/r0\n    if ii &lt; r0_list.shape[0]/2:\n        dfh['theta0'] = dfh['phi'] - np.arccos(dfh['cos_theta'])\n    else:\n        dfh['theta0'] = dfh['phi'] + np.arccos(dfh['cos_theta'])\n\n\n* Self-made version - reach `score=0.5` in 2 minutes but it rarely go above 0.5 since the unrolling function is not accurate. The reason why we use this self-made version is that it can find different tracks for later merging, which is good.\n\n####This tries to find possible `theta0` using an approximation function `ii` from -120 to 120\n\n\n    STEPRR = 0.03\n    rr = dfh.r/1000\n    dfh['theta0'] = dfh['phi'] + (rr + STEPRR*rr**2)*ii/180*np.pi + (0.00001*ii)*dfh.z*np.sign(dfh.z)/180*np.pi\n\n\n###Main features that are constant in a perfect helix form\n* `sin(theta0)`\n* `cos(theta0)`\n* `(z-z0)/arc`\n\n* The problem is `z/arc` is still uneven in the magnetic field, so I've been trying to improve this problem by using following features in different models. Other Kagglers definitely have a more accurate equation. \n* `log(1 + abs((z-z0)/arc))*sign(z)`\n* `(z-z0)/arc*sqrt(sin(arctan2(r,z-z0)))` I use this square root sine function as a correction term for the azimuthal angle on the x-y plane projection\n\n###Main features that are often constant\n* `(z-z0)/r` where `r` is the Euclidean distance from the hit to the origin.\n* `log(1 + abs((z-z0)/r))*sign(z)` is an approach to get z values closer to the origin to improve the problem with uneven `z/r` values\n\n###Side features\n* Those are often not constant but we can find different tracks using them with small weights when we cluster\n* `x/d` where `d` is the Eucliean distance from the hit to the origin in 3D `(x**2+y**2+z**2)**0.5`\n* `y/d`\n* `arctan2(z-r0,r)`\n* `px, py` in my code: `-r*cos(theta0)*cos(phi)-r*sin(theta0)*sin(phi)` and `-r*cos(theta0)*sin(phi)+r*sin(theta0)*cos(phi)`. I happened to find this feature that can find the seeds of non-z-axis centered tracks and we can extend the tracks using the found seeds. \n\n###Non-z-axis centered tracks\n* Skip this part since we didn't have time to implement it. Add the closest approach `D0` to your equations, you can find full equations with `D0` and `D1` from [Full helix equations for ATLAS](http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf)\n\n\n## LSTM approach for track fitting\n* My [notebook](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/notebooks/train_LSTM.ipynb) including visualization\n* We take first five hits as a seeded track and predict next 5 hits. I believe this has its potential as a track validation tool if right features were trained. I shared the detail in [this post](https://www.kaggle.com/c/trackml-particle-identification/discussion/60455#352645)\n\n## pointNet approach \n* PointNet is a lightweight CNN that can be used for pixel level classification(segementation) or classification. I put experimental code [here](https://github.com/jliamfinnie/kaggle-trackml/blob/master/src/train_pointnet.py), the challenge is to generate the right train data.\n\n\n## Background knowledge \n\n* [Very good slides for beginners](http://ific.uv.es/~nebot/IDPASC/Material/Tracking-Vertexing/Tracking-Vertexing-Slides.pdf)\n\n* [Lecture of particles tracking](http://www.physics.iitm.ac.in/~sercehep2013/track2_Gagan_Mohanty.pdf)\n\n\n* [Full helix equations for ATLAS](http://www.hep.ucl.ac.uk/atlas/atlantis/files/helix_equations_1.pdf) - All equations you need!\n\n\n* [Diplom thesis](http://physik.uibk.ac.at/hephy/theses/dipl_as.pdf) of Andreas Salzburger (Wow, he started in this field as a CERN student already in 2001 :p )\n\n* [Doctor thesis](http://physik.uibk.ac.at/hephy/theses/diss_as.pdf) of Andreas Salzburger\n\n* [CERN tracking software Acts](https://gitlab.cern.ch/acts/acts-core) - Sadly, we didn't have time to explore it :) \n\n\n",
    "370015": "Congratulations and thanks so much @Nicole and @Liam for sharing your code and all that you have shared in the discussion threads. It is an amazing approach to this problem. \n\nI will need time to go through your code later as now I have to run over to what is a messed up Santander competition to try my best in the remaining 6 days.",
    "369909": "Congrats. Sad to see you guys missed the gold. Thank you for the solution.",
    "369908": "Congrats and thanks for sharing your approach and code! Will be helpful in figuring out where I went wrong trying to find better features.",
    "369900": "Congrats and sorry for your miss of gold by one.  I won't say I hope someone cheated in front of you, as I don't think it happened, but I thought of it still ;)\n\nThanks for the writeup.\n",
    "370008": "How can we forget to thank @Kha too, he spawned the discussion thread with @yuval together and made us realise we did it all wrong. 😉 I wish you shared it earlier. ",
    "369911": "",
    "369901": ""
  }
}