{
  "id": 57180,
  "title": "Idea for 0.3x solution: DBSCAN on unrolled helixes",
  "url": "/competitions/trackml-particle-identification/discussion/57180",
  "author_name": "",
  "post_date": "2018-05-20T15:41:32.312419600Z",
  "votes": 45,
  "comment_count": 43,
  "views": 0,
  "content": "<p>I'd like to share an approach I used in the last few submissions, getting 0.28 -- 0.38 depending on some details. This is quite far even from the current leaders, and I'm not sure how viable is this approach in the long run, but maybe it will be helpful. At least it's fast to run and does not require any training data.</p>\n\n<p>Idea is based on DBSCAN baseline: <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark</a>. It seems that it works best on relatively straight tracks which are more or less aligned with the z axis. It's possible to unroll helixes (make them straight) by rotating each hit around the z axis by the angle that is proportional to the distance from the z axis. The angle that unrolls the helix is different for each helix, so the idea is to:</p>\n\n<pre><code>for angle in ~20 evenly spaced angles:\n     x2, y2 &lt;- rotate x, y by angle * (distance to z), unrolling some helixes\n     do helix transform and dbscan clustering, as in the original kernel\nmerge clusters from all angles\n</code></pre>\n\n<p>Probably this can be extended by doing other transformations, besides just rotations to unroll helixes, or post-processing of discovered clusters. But this has no chance to obtain tracks that start far from origin, it seems.</p>",
  "messages": [
    {
      "id": "331193",
      "postDate": "05/20/2018 15:41:32",
      "content": "<p>I'd like to share an approach I used in the last few submissions, getting 0.28 -- 0.38 depending on some details. This is quite far even from the current leaders, and I'm not sure how viable is this approach in the long run, but maybe it will be helpful. At least it's fast to run and does not require any training data.</p>\n\n<p>Idea is based on DBSCAN baseline: <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark</a>. It seems that it works best on relatively straight tracks which are more or less aligned with the z axis. It's possible to unroll helixes (make them straight) by rotating each hit around the z axis by the angle that is proportional to the distance from the z axis. The angle that unrolls the helix is different for each helix, so the idea is to:</p>\n\n<pre><code>for angle in ~20 evenly spaced angles:\n     x2, y2 &lt;- rotate x, y by angle * (distance to z), unrolling some helixes\n     do helix transform and dbscan clustering, as in the original kernel\nmerge clusters from all angles\n</code></pre>\n\n<p>Probably this can be extended by doing other transformations, besides just rotations to unroll helixes, or post-processing of discovered clusters. But this has no chance to obtain tracks that start far from origin, it seems.</p>",
      "rawMarkdown": "I'd like to share an approach I used in the last few submissions, getting 0.28 -- 0.38 depending on some details. This is quite far even from the current leaders, and I'm not sure how viable is this approach in the long run, but maybe it will be helpful. At least it's fast to run and does not require any training data.\n\nIdea is based on DBSCAN baseline: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark. It seems that it works best on relatively straight tracks which are more or less aligned with the z axis. It's possible to unroll helixes (make them straight) by rotating each hit around the z axis by the angle that is proportional to the distance from the z axis. The angle that unrolls the helix is different for each helix, so the idea is to:\n\n    for angle in ~20 evenly spaced angles:\n         x2, y2 &lt;- rotate x, y by angle * (distance to z), unrolling some helixes\n         do helix transform and dbscan clustering, as in the original kernel\n    merge clusters from all angles\n\nProbably this can be extended by doing other transformations, besides just rotations to unroll helixes, or post-processing of discovered clusters. But this has no chance to obtain tracks that start far from origin, it seems.",
      "votes": null
    },
    {
      "id": "331201",
      "postDate": "05/20/2018 15:50:34",
      "content": "<p>@Konstantin, It was my first approach. Fitting the parameters enables to obtain at least 0.42. Best wishes.</p>\n\n<p>EDIT: I use over 50 angles (nonlinearly distributed), eps being a function of angle and a little bit different approach to z.</p>",
      "rawMarkdown": "Konstantin, It was my first approach. Fitting the parameters enables to obtain at least 0.42. Best wishes.\n\nEDIT: I use over 50 angles (nonlinearly distributed), eps being a function of angle and a little bit different approach to z.",
      "votes": null
    },
    {
      "id": "331220",
      "postDate": "05/20/2018 16:54:15",
      "content": "<p>Glad others thought the same, I was actually working on that yesterday, but wanted to find the best way to unroll a helix to straight line (not just proportional to z or radius) </p>",
      "rawMarkdown": "Glad others thought the same, I was actually working on that yesterday, but wanted to find the best way to unroll a helix to straight line (not just proportional to z or radius)",
      "votes": null
    },
    {
      "id": "331290",
      "postDate": "05/20/2018 21:38:59",
      "content": "<p>Have you tried using HDBSCAN? I'm just a beginner, but from my experience so far on the 0.2x range, this algorithm works better than DBSCAN. Does this situation change once some helices are straighten up?</p>",
      "rawMarkdown": "Have you tried using HDBSCAN? I'm just a beginner, but from my experience so far on the 0.2x range, this algorithm works better than DBSCAN. Does this situation change once some helices are straighten up?",
      "votes": null
    },
    {
      "id": "331338",
      "postDate": "05/21/2018 03:37:04",
      "content": "<p>@Konstantin, thanks for the solution!</p>",
      "rawMarkdown": "Konstantin, thanks for the solution!",
      "votes": null
    },
    {
      "id": "331359",
      "postDate": "05/21/2018 04:59:50",
      "content": "<p>Here's an idea I've been toying with (but haven't fully implemented):</p>\n\n<ul>\n<li>project to the $XY$-plane (so now helices look like circles and lines --well-- still look like lines),</li>\n<li>try to fit lines to points using RANSAC, Hough, DBSCAN, etc.,</li>\n<li>try to fit circles to points,</li>\n<li>try those last two steps in different orders (after filtering out previously classified lines/circles).</li>\n</ul>\n\n<p>Have you tried using projections? I've had some moderate success using projections (and, in general, different transformations of the data).</p>",
      "rawMarkdown": "Here's an idea I've been toying with (but haven't fully implemented):\n\n * project to the $XY$-plane (so now helices look like circles and lines --well-- still look like lines),\n * try to fit lines to points using RANSAC, Hough, DBSCAN, etc.,\n * try to fit circles to points,\n * try those last two steps in different orders (after filtering out previously classified lines/circles).\n\nHave you tried using projections? I've had some moderate success using projections (and, in general, different transformations of the data).",
      "votes": null
    },
    {
      "id": "331387",
      "postDate": "05/21/2018 06:17:36",
      "content": "<p>I briefly tried HDBSCAN with this approach and it worked worse, I think the reason is that it gives more false matches, which do not allow to do proper cluster merging. But maybe I did something wrong, please share your results if you try it.</p>",
      "rawMarkdown": "I briefly tried HDBSCAN with this approach and it worked worse, I think the reason is that it gives more false matches, which do not allow to do proper cluster merging. But maybe I did something wrong, please share your results if you try it.",
      "votes": null
    },
    {
      "id": "331446",
      "postDate": "05/21/2018 08:51:33",
      "content": "<p>you may want to check this projection. hope that i did not make any mistake.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331446/9475/projection.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "you may want to check this projection. hope that i did not make any mistake.\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331446/9475/projection.png",
      "votes": null
    },
    {
      "id": "331495",
      "postDate": "05/21/2018 12:01:54",
      "content": "<p>oops. i am sorry that there is a mistake. </p>\n\n<p>phi = np.arctan2(y,x)/r is effectively zero</p>",
      "rawMarkdown": "oops. i am sorry that there is a mistake. \n\nphi = np.arctan2(y,x)/r is effectively zero",
      "votes": null
    },
    {
      "id": "331563",
      "postDate": "05/21/2018 13:44:56",
      "content": "<p>hey, Heng! Shouldn't the rotation be something like \nhx = x*np.cos(phi*r) - y*np.sin(phi*r)\nand\nhy = x*np.sin(phi*r) + y*np.cos(phi*r) ?\nI mean, from your plot it clearly works, but I'm' not understanding why it works if the rotation matrix is written differently from the above.</p>",
      "rawMarkdown": "hey, Heng! Shouldn't the rotation be something like \nhx = x*np.cos(phi*r) - y*np.sin(phi*r)\nand\nhy = x*np.sin(phi*r) + y*np.cos(phi*r) ?\nI mean, from your plot it clearly works, but I'm' not understanding why it works if the rotation matrix is written differently from the above.",
      "votes": null
    },
    {
      "id": "331617",
      "postDate": "05/21/2018 15:20:53",
      "content": "<p>your equation is correct. </p>\n\n<p>Mine is a mistake. It simply project z to a plane. x and y are not used at all.</p>",
      "rawMarkdown": "your equation is correct. \n\nMine is a mistake. It simply project z to a plane. x and y are not used at all.",
      "votes": null
    },
    {
      "id": "331620",
      "postDate": "05/21/2018 15:31:13",
      "content": "<p>@Henrique Mello</p>\n\n<p>corrected code and results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331620/9481/corrected.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Henrique Mello\n \ncorrected code and results\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331620/9481/corrected.png",
      "votes": null
    },
    {
      "id": "331635",
      "postDate": "05/21/2018 16:01:38",
      "content": "<p>In the x-y plane tracks which pass through the origin (it seems most of them) can be made linear by making a conformal mapping.\n$$\nu = \\frac{x}{x^2 + y^2},  \\\nv = \\frac{y}{x^2 + y^2} <br>\n$$\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331635/9482/conformal.png\" alt=\"Figure\"></p>\n\n<p>Unlike shown above, this unrolling procedure <em>might</em> work because it straightens only a subset of tracks out at a time.  This increases the density for only a subset in the clustered space which can be iteratively removed.  If this is the case, it might be interesting to try and learn the distribution of this unrolling angle from the truth dataset, and concentrate more iterations of clustering in those regions with more tracks (ie: choose a non-uniform grid of angles to unroll).  </p>",
      "rawMarkdown": "In the x-y plane tracks which pass through the origin (it seems most of them) can be made linear by making a conformal mapping.\n$$\nu = \\frac{x}{x^2 + y^2},  \\\\\nv = \\frac{y}{x^2 + y^2}  \n$$\n![Figure][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331635/9482/conformal.png\n\nUnlike shown above, this unrolling procedure *might* work because it straightens only a subset of tracks out at a time.  This increases the density for only a subset in the clustered space which can be iteratively removed.  If this is the case, it might be interesting to try and learn the distribution of this unrolling angle from the truth dataset, and concentrate more iterations of clustering in those regions with more tracks (ie: choose a non-uniform grid of angles to unroll).",
      "votes": null
    },
    {
      "id": "331667",
      "postDate": "05/21/2018 17:09:14",
      "content": "<p>@Konstantin, thanks a lot for sharing your findings, this is triggering lots of thoughts for me.  Hope they'll translate into something effective.</p>",
      "rawMarkdown": "Konstantin, thanks a lot for sharing your findings, this is triggering lots of thoughts for me.  Hope they'll translate into something effective.",
      "votes": null
    },
    {
      "id": "332021",
      "postDate": "05/22/2018 11:17:14",
      "content": "<p>Great stuff. Any basic ideas with regards to merging clusters? Say you run clustering 10 times with different angles. How to choose which of the clusterings to use for a particular hit?</p>",
      "rawMarkdown": "Great stuff. Any basic ideas with regards to merging clusters? Say you run clustering 10 times with different angles. How to choose which of the clusterings to use for a particular hit?",
      "votes": null
    },
    {
      "id": "332026",
      "postDate": "05/22/2018 11:29:15",
      "content": "<p>@maka, I use similar approach as Konstantin. Because more energetic particles and bigger tracks are more prefered by scoring function, I start from bigger angles and finish at 0 angle changing assigment of the hit if the present cluster is bigger than the previous one.</p>\n\n<p>Edit: (Cluster/track) bigger in terms of number of hits. It is a very simple model. Just to start and find problems.</p>",
      "rawMarkdown": "maka, I use similar approach as Konstantin. Because more energetic particles and bigger tracks are more prefered by scoring function, I start from bigger angles and finish at 0 angle changing assigment of the hit if the present cluster is bigger than the previous one.\n\nEdit: (Cluster/track) bigger in terms of number of hits. It is a very simple model. Just to start and find problems.",
      "votes": null
    },
    {
      "id": "332029",
      "postDate": "05/22/2018 11:35:24",
      "content": "<p><a href=\"/sionek\">@sionek</a> Bigger in terms of number of hits in cluster or in terms of cluster std.dev in helix coordinates? Would assume one would like the first high and the latter low?</p>",
      "rawMarkdown": "sionek Bigger in terms of number of hits in cluster or in terms of cluster std.dev in helix coordinates? Would assume one would like the first high and the latter low?",
      "votes": null
    },
    {
      "id": "332090",
      "postDate": "05/22/2018 13:56:26",
      "content": "<p>My merging algorithm is a bit involved. First I calculate the size of all clusters, and discard clusters bigger than 20 hits. After that I eagerly assign each hit to the biggest cluster it belongs too (this also assigns all other hits belonging to this cluster), starting from the points with highest number of clusters.</p>",
      "rawMarkdown": "My merging algorithm is a bit involved. First I calculate the size of all clusters, and discard clusters bigger than 20 hits. After that I eagerly assign each hit to the biggest cluster it belongs too (this also assigns all other hits belonging to this cluster), starting from the points with highest number of clusters.",
      "votes": null
    },
    {
      "id": "332102",
      "postDate": "05/22/2018 14:36:23",
      "content": "<p>@Konstantin, It would be nice to compare someday our models and dates of their creation. Just to study telepathy. A piece of my code regarding discarding clusters bigger than 20 (hm, rather 19 ;)</p>\n\n<blockquote>\n  <p>dfh[,s1:=ifelse(N2&gt;N1 &amp; N2&lt;20,s2+maxs1,s1)]</p>\n</blockquote>\n\n<p>(N2 and N1 - number of hits in present and previous cluster. s1 and s2 - previous and present cluster number). The difference - in my model after removing a hit from the previous cluster, it becomes smaller (less important) and then it is easier to remove other hits from it, but the final effect is probably similar. Best wishes.</p>",
      "rawMarkdown": "Konstantin, It would be nice to compare someday our models and dates of their creation. Just to study telepathy. A piece of my code regarding discarding clusters bigger than 20 (hm, rather 19 ;)\n\n&gt; dfh[,s1:=ifelse(N2&gt;N1 &amp; N2&lt;20,s2+maxs1,s1)]\n\n(N2 and N1 - number of hits in present and previous cluster. s1 and s2 - previous and present cluster number). The difference - in my model after removing a hit from the previous cluster, it becomes smaller (less important) and then it is easier to remove other hits from it, but the final effect is probably similar. Best wishes.",
      "votes": null
    },
    {
      "id": "332116",
      "postDate": "05/22/2018 15:02:10",
      "content": "<p>Thanks @Konstantin <a href=\"/sionek\">@sionek</a> . </p>\n\n<p>I need to just get a basic solution up and running myself. But here is one idea you might consider:</p>\n\n<p>From what I can understand from the equations(see my kernel), the angle that needs to be unrolled is $qB(z-z0)/p_{0,z}$. q is charge, B is magnetic field( -0.00057 is a good value for first event in train_1). Since we know what the distribution of p_{0,z} should be approximately, maybe we can use this to choose angles or merge clusters better. </p>",
      "rawMarkdown": "Thanks @Konstantin @sionek . \n\nI need to just get a basic solution up and running myself. But here is one idea you might consider:\n\nFrom what I can understand from the equations(see my kernel), the angle that needs to be unrolled is $qB(z-z0)/p_{0,z}$. q is charge, B is magnetic field( -0.00057 is a good value for first event in train_1). Since we know what the distribution of p_{0,z} should be approximately, maybe we can use this to choose angles or merge clusters better.",
      "votes": null
    },
    {
      "id": "332134",
      "postDate": "05/22/2018 15:47:01",
      "content": "<p>@maka, thanks for your kernel.  I have been looking at the exact same approach in the past few days.</p>",
      "rawMarkdown": "maka, thanks for your kernel.  I have been looking at the exact same approach in the past few days.",
      "votes": null
    },
    {
      "id": "332157",
      "postDate": "05/22/2018 16:26:20",
      "content": "<p>@Grzegorz &amp; @Konstantin <br>\nIt seems telepathic waves are flowing around...  In my case, 20 is working better  ;)</p>\n\n<pre><code>sel=function(x,Nx,y,Ny,N){\n    I=x\n    I[Ny&gt;Nx &amp; Ny&lt;N]=max(x)+y[Ny&gt;Nx &amp; Ny &lt;N]\n    I[x==-1]=max(x)+y[x==-1]\n    I\n  }\n mm[,I:=sel(I.x,Nx,I.y,Ny,21)]\n</code></pre>",
      "rawMarkdown": "Grzegorz &amp; @Konstantin  \nIt seems telepathic waves are flowing around...  In my case, 20 is working better  ;)\n\n    sel=function(x,Nx,y,Ny,N){\n        I=x\n        I[Ny&gt;Nx &amp; Ny",
      "votes": null
    },
    {
      "id": "332213",
      "postDate": "05/22/2018 19:23:23",
      "content": "<p>@Vicens, I'm afraid, these waves flow around Europe only or Mickey (Japan) uses \"skupiacz myśli\" (~thought concentrator) - a kind of deflector disabling thoughts to go away from the head, usually made of steel or kevlar, used in the army. I can't hear even a short haiku from him.</p>",
      "rawMarkdown": "Vicens, I'm afraid, these waves flow around Europe only or Mickey (Japan) uses \"skupiacz myśli\" (~thought concentrator) - a kind of deflector disabling thoughts to go away from the head, usually made of steel or kevlar, used in the army. I can't hear even a short haiku from him.",
      "votes": null
    },
    {
      "id": "332232",
      "postDate": "05/22/2018 20:35:26",
      "content": "<p>Hey <a href=\"/hengck23\">@hengck23</a>, FYI, we tried out your original code (the one you said now has the wrong equation), as well as the 'fixed' version. For us, the original code produced better results - i.e. after implementing for all combinations (x&gt;0, x&lt;0, y&gt;0, y&lt;0, z&gt;0, z&lt;0), the 'basic' DBSCAN kernel score improved from 0.204 to 0.209 (for your selected event 1029), and our HDBSCAN score improved from 0.250 to 0.258. The 'fixed' equations produced only a very small benefit on top of the base DBSCAN/HDBSCAN models. Of course, now the only problem is how to re-create without having the ground-truth tracks and momentum :-) I guess this is where <a href=\"/lopuhin\">@lopuhin</a> and <a href=\"/sionek\">@sionek</a>'s suggestions can be used....</p>\n\n<p>However, from reading some of the other research material (<a href=\"https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf\">https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf</a>), it sounds like this approach (conformal mapping and histogramming) seems to be discouraged, or at least is seen to have minimal benefit. I guess maybe this can still be useful as a 'first step', used in combination with other approaches....</p>",
      "rawMarkdown": "Hey @hengck23, FYI, we tried out your original code (the one you said now has the wrong equation), as well as the 'fixed' version. For us, the original code produced better results - i.e. after implementing for all combinations (x&gt;0, x&lt;0, y&gt;0, y&lt;0, z&gt;0, z&lt;0), the 'basic' DBSCAN kernel score improved from 0.204 to 0.209 (for your selected event 1029), and our HDBSCAN score improved from 0.250 to 0.258. The 'fixed' equations produced only a very small benefit on top of the base DBSCAN/HDBSCAN models. Of course, now the only problem is how to re-create without having the ground-truth tracks and momentum :-) I guess this is where @lopuhin and @sionek's suggestions can be used....\n\nHowever, from reading some of the other research material ([https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf][1]), it sounds like this approach (conformal mapping and histogramming) seems to be discouraged, or at least is seen to have minimal benefit. I guess maybe this can still be useful as a 'first step', used in combination with other approaches....\n\n  [1]: https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf",
      "votes": null
    },
    {
      "id": "332451",
      "postDate": "05/23/2018 07:13:04",
      "content": "<p><a href=\"/sionek\">@sionek</a> haha, yeah I think we have a curious case indeed ;) Here is my code, to be fair I checked a few different values (12 -- 24):</p>\n\n<pre><code>_, inverse, counts = np.unique(labels, return_inverse=True, return_counts=True)\ncounts = counts[inverse]\ncounts[labels == -1] = 0\ncounts[counts &gt; 20] = 0  # important: if the cluster is large, it's bad\n</code></pre>",
      "rawMarkdown": "sionek haha, yeah I think we have a curious case indeed ;) Here is my code, to be fair I checked a few different values (12 -- 24):\n \n    _, inverse, counts = np.unique(labels, return_inverse=True, return_counts=True)\n    counts = counts[inverse]\n    counts[labels == -1] = 0\n    counts[counts &gt; 20] = 0  # important: if the cluster is large, it's bad",
      "votes": null
    },
    {
      "id": "332714",
      "postDate": "05/23/2018 16:00:21",
      "content": "<p>Neat1</p>",
      "rawMarkdown": "Neat1",
      "votes": null
    },
    {
      "id": "332899",
      "postDate": "05/24/2018 02:34:53",
      "content": "<p>I also tried this implementation with HDBSCAN and saw worse results too</p>",
      "rawMarkdown": "I also tried this implementation with HDBSCAN and saw worse results too",
      "votes": null
    },
    {
      "id": "333005",
      "postDate": "05/24/2018 07:57:35",
      "content": "<p>Analyzing @Grzegorz Sionkowski kernel (0.3472 DBSCAN).</p>\n\n<p>conclusion:</p>\n\n<ul>\n<li><p>the tracks are rather long (good for back-fitting lines)</p></li>\n<li><p>the clustered hits are not completed,  missing the first or last hits of the ground truth tracks. This is where the weights are high.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/333005/9514/back%20fitting.png\" alt=\"enter image description here\"></p></li>\n</ul>",
      "rawMarkdown": "Analyzing @Grzegorz Sionkowski kernel (0.3472 DBSCAN).\n\nconclusion:\n\n  - the tracks are rather long (good for back-fitting lines)\n  \n  - the clustered hits are not completed,  missing the first or last hits of the ground truth tracks. This is where the weights are high.\n  \n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/333005/9514/back%20fitting.png",
      "votes": null
    },
    {
      "id": "333087",
      "postDate": "05/24/2018 11:30:30",
      "content": "<p>I see...Thanks a lot, this works awesome but question... Am I dumb or is it Ok to have no idea whats happening here? :)</p>",
      "rawMarkdown": "I see...Thanks a lot, this works awesome but question... Am I dumb or is it Ok to have no idea whats happening here? :)",
      "votes": null
    },
    {
      "id": "333104",
      "postDate": "05/24/2018 12:10:30",
      "content": "<p>@NxGTR, European matters. CERN has to prepare to the changes of EU regulations. Officially, it is called General Data Protection Regulation but there is a hidden directive: every curve must be straight. </p>",
      "rawMarkdown": "NxGTR, European matters. CERN has to prepare to the changes of EU regulations. Officially, it is called General Data Protection Regulation but there is a hidden directive: every curve must be straight.",
      "votes": null
    },
    {
      "id": "333181",
      "postDate": "05/24/2018 15:16:20",
      "content": "<blockquote>\n  <p>is it Ok to have no idea whats happening here?</p>\n</blockquote>\n\n<p>Are you saying you use a public kernel as is? ;)</p>\n\n<p>I feel free to add comments given I did not submit yet.  I may shut up for a while after getting a miserable score ;)</p>",
      "rawMarkdown": "&gt; is it Ok to have no idea whats happening here?\n\nAre you saying you use a public kernel as is? ;)\n\nI feel free to add comments given I did not submit yet.  I may shut up for a while after getting a miserable score ;)",
      "votes": null
    },
    {
      "id": "333186",
      "postDate": "05/24/2018 15:21:59",
      "content": "<p>Amm... let's say I was busy with other stuff and accidentally happened, proof that that is totally possible <a href=\"https://en.wikipedia.org/wiki/Infinite_monkey_theorem\">here</a>.</p>",
      "rawMarkdown": "Amm... let's say I was busy with other stuff and accidentally happened, proof that that is totally possible [here][1].\n\n\n  [1]: https://en.wikipedia.org/wiki/Infinite_monkey_theorem",
      "votes": null
    },
    {
      "id": "333683",
      "postDate": "05/25/2018 17:05:29",
      "content": "<p>I've done some experiments with this map. I'll return the favor --here is another one for you to try:</p>\n\n<p>$$\nx,y \\quad \\mapsto \\quad \\theta, \\log(r),\n$$</p>\n\n<p>where </p>\n\n<p>$$\nr := \\sqrt{x^2+y^2} \\quad , \\quad \\theta := \\arctan \\left( \\frac{y}{x} \\right) .\n$$</p>\n\n<p>This maps points in the punctured plane lying on a ray to vertical lines in a plane.</p>\n\n<p>One can then run a quicker version of the Hough transform (or other clustering/line-identifying method) which only captures vertical lines.</p>",
      "rawMarkdown": "I've done some experiments with this map. I'll return the favor --here is another one for you to try:\n\n$$\nx,y \\quad \\mapsto \\quad \\theta, \\log(r),\n$$\n\nwhere \n\n$$\nr := \\sqrt{x^2+y^2} \\quad , \\quad \\theta := \\arctan \\left( \\frac{y}{x} \\right) .\n$$\n\nThis maps points in the punctured plane lying on a ray to vertical lines in a plane.\n\nOne can then run a quicker version of the Hough transform (or other clustering/line-identifying method) which only captures vertical lines.",
      "votes": null
    },
    {
      "id": "333717",
      "postDate": "05/25/2018 18:16:11",
      "content": "<p>Side note: sometimes you need to be careful about where the \"cut\" occurs (use a different theta). \nThe noise can cause what should be one vertical line to map to two different vertical lines.</p>",
      "rawMarkdown": "Side note: sometimes you need to be careful about where the \"cut\" occurs (use a different theta). \nThe noise can cause what should be one vertical line to map to two different vertical lines.",
      "votes": null
    },
    {
      "id": "334631",
      "postDate": "05/28/2018 03:55:04",
      "content": "<p>Did you find that making eps as a function of angle increased the score? It seemed to be detrimental in my version.\nAlso, I tried to fork your kernel but could import dbscan.  Did you simply add a custom github repo?</p>",
      "rawMarkdown": "Did you find that making eps as a function of angle increased the score? It seemed to be detrimental in my version.\nAlso, I tried to fork your kernel but could import dbscan.  Did you simply add a custom github repo?",
      "votes": null
    },
    {
      "id": "334636",
      "postDate": "05/28/2018 04:13:37",
      "content": "<p>You can increment EPS and angle at the same time and get better results, but incrementing them both by small steps and using different step sizes for each. </p>\n\n<p>See the high scoring public kernels based on dbscan for reference.</p>",
      "rawMarkdown": "You can increment EPS and angle at the same time and get better results, but incrementing them both by small steps and using different step sizes for each. \n\nSee the high scoring public kernels based on dbscan for reference.",
      "votes": null
    },
    {
      "id": "334643",
      "postDate": "05/28/2018 04:27:01",
      "content": "<p>Yeah, I am using a python implementation similar to Heng's of Grzegorz's kernel found here: <a href=\"https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843\">https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843</a></p>\n\n<p>However, when I add the optimizations to eps like,</p>\n\n<pre><code>eps = self.eps + (i * 0.000005)\n</code></pre>\n\n<p>it made the model perform worse</p>",
      "rawMarkdown": "Yeah, I am using a python implementation similar to Heng's of Grzegorz's kernel found here: [https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843][1]\n\nHowever, when I add the optimizations to eps like,\n\n    eps = self.eps + (i * 0.000005)\n\nit made the model perform worse\n\n  [1]: https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843",
      "votes": null
    },
    {
      "id": "334650",
      "postDate": "05/28/2018 05:02:27",
      "content": "<p>I tried tweaking the similar line in the notebook by Grzegorz Sionkowski</p>\n\n<p>'res=dbscan(dfs,eps=0.0035+ii*stepeps'...\nwhere stepeps = 0.000005</p>\n\n<p>and most of the tweaks I tried hurt performance, making me think he already did a decent job of tuning the scaling of eps with angle, and started looking elsewhere. Are you sure you're passing a reasonable value in with self.eps? </p>\n\n<p>...strange... sorry I'm not of more help. </p>",
      "rawMarkdown": "I tried tweaking the similar line in the notebook by Grzegorz Sionkowski\n\n'res=dbscan(dfs,eps=0.0035+ii*stepeps'...\nwhere stepeps = 0.000005\n\nand most of the tweaks I tried hurt performance, making me think he already did a decent job of tuning the scaling of eps with angle, and started looking elsewhere. Are you sure you're passing a reasonable value in with self.eps? \n\n...strange... sorry I'm not of more help.",
      "votes": null
    },
    {
      "id": "334661",
      "postDate": "05/28/2018 05:42:35",
      "content": "<p>@MathewMasters, <a href=\"/macfarll\">@macfarll</a>, Talking about parameters of my model, I was talking about the model of the score 0.42 (there are some ideas I want to keep for myself) . Using that model I have obtained better results decreasing eps a little bit in each iteration. In the model 0.34-0.37 I just show how you can additionally fit the parameters and looking at the leaderboard I can see many Kagglers not only have found better parameters but also they have implemented their own ideas basing on the DBSCAN benchmark skeleton.</p>\n\n<p>My advice. Be more open for ideas. Do not concentrate on my parameters and features. Take only the idea of iterations. Read about the model by @Konstantin - in his model (and in mine of the score 0.42) the angle is not the most important feature for DBSCAN, it is just a way to obtain new values of x and y. </p>",
      "rawMarkdown": "MathewMasters, @macfarll, Talking about parameters of my model, I was talking about the model of the score 0.42 (there are some ideas I want to keep for myself) . Using that model I have obtained better results decreasing eps a little bit in each iteration. In the model 0.34-0.37 I just show how you can additionally fit the parameters and looking at the leaderboard I can see many Kagglers not only have found better parameters but also they have implemented their own ideas basing on the DBSCAN benchmark skeleton.\n\nMy advice. Be more open for ideas. Do not concentrate on my parameters and features. Take only the idea of iterations. Read about the model by @Konstantin - in his model (and in mine of the score 0.42) the angle is not the most important feature for DBSCAN, it is just a way to obtain new values of x and y.",
      "votes": null
    },
    {
      "id": "334904",
      "postDate": "05/28/2018 17:24:41",
      "content": "<p>@Grzegorz, thanks for your reply and advice, I appreciate it. Just one question, how were you able to import dbscan into the kaggle kernel?  A github repo?</p>",
      "rawMarkdown": "Grzegorz, thanks for your reply and advice, I appreciate it. Just one question, how were you able to import dbscan into the kaggle kernel?  A github repo?",
      "votes": null
    },
    {
      "id": "334911",
      "postDate": "05/28/2018 17:35:18",
      "content": "<p>Sorry if I misspoke about your model. I was trying to refer to the kernel you published, and thought that since I put some time into understanding it I could try to help answer some questions. My way of understanding how it worked was mainly through tweaking and tinkering things and seeing what happened.</p>\n\n<p>That being said, I obviously have a lot to learn. </p>\n\n<p>Since then I have been working on different ideas instead of just tweaking parameters, such as extending tracks where the end points are missing. </p>\n\n<p>Sorry if I caused any trouble / was misleading.</p>",
      "rawMarkdown": "Sorry if I misspoke about your model. I was trying to refer to the kernel you published, and thought that since I put some time into understanding it I could try to help answer some questions. My way of understanding how it worked was mainly through tweaking and tinkering things and seeing what happened.\n\nThat being said, I obviously have a lot to learn. \n\nSince then I have been working on different ideas instead of just tweaking parameters, such as extending tracks where the end points are missing. \n\nSorry if I caused any trouble / was misleading.",
      "votes": null
    },
    {
      "id": "371829",
      "postDate": "08/17/2018 16:49:54",
      "content": "<p>Hi, No one ever thanked you for this idea, but I think it is what made DBSCAN a good option here.  Therefore a big THANK YOU.</p>",
      "rawMarkdown": "Hi, No one ever thanked you for this idea, but I think it is what made DBSCAN a good option here.  Therefore a big THANK YOU.",
      "votes": null
    },
    {
      "id": "371879",
      "postDate": "08/17/2018 18:59:31",
      "content": "<p>I thanked in my kernel:</p>\n\n<p><strong>to Konstantin Lopuhin for pulishing the idea in</strong></p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57180\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57180</a></p>",
      "rawMarkdown": "I thanked in my kernel:\n\n**to Konstantin Lopuhin for pulishing the idea in**\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/57180",
      "votes": null
    },
    {
      "id": "371898",
      "postDate": "08/17/2018 19:48:25",
      "content": "<blockquote>\n  <p>I thanked in my kernel</p>\n</blockquote>\n\n<p>Great, you're better than me!</p>",
      "rawMarkdown": "&gt; I thanked in my kernel\n\nGreat, you're better than me!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 331201,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "05/20/2018 15:50:34",
      "content": "<p>@Konstantin, It was my first approach. Fitting the parameters enables to obtain at least 0.42. Best wishes.</p>\n\n<p>EDIT: I use over 50 angles (nonlinearly distributed), eps being a function of angle and a little bit different approach to z.</p>",
      "votes": null,
      "replies": [
        {
          "id": 334631,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "05/28/2018 03:55:04",
          "content": "<p>Did you find that making eps as a function of angle increased the score? It seemed to be detrimental in my version.\nAlso, I tried to fork your kernel but could import dbscan.  Did you simply add a custom github repo?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334636,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "05/28/2018 04:13:37",
          "content": "<p>You can increment EPS and angle at the same time and get better results, but incrementing them both by small steps and using different step sizes for each. </p>\n\n<p>See the high scoring public kernels based on dbscan for reference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334643,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "05/28/2018 04:27:01",
          "content": "<p>Yeah, I am using a python implementation similar to Heng's of Grzegorz's kernel found here: <a href=\"https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843\">https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843</a></p>\n\n<p>However, when I add the optimizations to eps like,</p>\n\n<pre><code>eps = self.eps + (i * 0.000005)\n</code></pre>\n\n<p>it made the model perform worse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334650,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "05/28/2018 05:02:27",
          "content": "<p>I tried tweaking the similar line in the notebook by Grzegorz Sionkowski</p>\n\n<p>'res=dbscan(dfs,eps=0.0035+ii*stepeps'...\nwhere stepeps = 0.000005</p>\n\n<p>and most of the tweaks I tried hurt performance, making me think he already did a decent job of tuning the scaling of eps with angle, and started looking elsewhere. Are you sure you're passing a reasonable value in with self.eps? </p>\n\n<p>...strange... sorry I'm not of more help. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334661,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/28/2018 05:42:35",
          "content": "<p>@MathewMasters, <a href=\"/macfarll\">@macfarll</a>, Talking about parameters of my model, I was talking about the model of the score 0.42 (there are some ideas I want to keep for myself) . Using that model I have obtained better results decreasing eps a little bit in each iteration. In the model 0.34-0.37 I just show how you can additionally fit the parameters and looking at the leaderboard I can see many Kagglers not only have found better parameters but also they have implemented their own ideas basing on the DBSCAN benchmark skeleton.</p>\n\n<p>My advice. Be more open for ideas. Do not concentrate on my parameters and features. Take only the idea of iterations. Read about the model by @Konstantin - in his model (and in mine of the score 0.42) the angle is not the most important feature for DBSCAN, it is just a way to obtain new values of x and y. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334904,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "05/28/2018 17:24:41",
          "content": "<p>@Grzegorz, thanks for your reply and advice, I appreciate it. Just one question, how were you able to import dbscan into the kaggle kernel?  A github repo?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334911,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "05/28/2018 17:35:18",
          "content": "<p>Sorry if I misspoke about your model. I was trying to refer to the kernel you published, and thought that since I put some time into understanding it I could try to help answer some questions. My way of understanding how it worked was mainly through tweaking and tinkering things and seeing what happened.</p>\n\n<p>That being said, I obviously have a lot to learn. </p>\n\n<p>Since then I have been working on different ideas instead of just tweaking parameters, such as extending tracks where the end points are missing. </p>\n\n<p>Sorry if I caused any trouble / was misleading.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331220,
      "author_name": "riadsouissi",
      "author_url": "",
      "post_date": "05/20/2018 16:54:15",
      "content": "<p>Glad others thought the same, I was actually working on that yesterday, but wanted to find the best way to unroll a helix to straight line (not just proportional to z or radius) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331290,
      "author_name": "hrmello",
      "author_url": "",
      "post_date": "05/20/2018 21:38:59",
      "content": "<p>Have you tried using HDBSCAN? I'm just a beginner, but from my experience so far on the 0.2x range, this algorithm works better than DBSCAN. Does this situation change once some helices are straighten up?</p>",
      "votes": null,
      "replies": [
        {
          "id": 331387,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/21/2018 06:17:36",
          "content": "<p>I briefly tried HDBSCAN with this approach and it worked worse, I think the reason is that it gives more false matches, which do not allow to do proper cluster merging. But maybe I did something wrong, please share your results if you try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332899,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "05/24/2018 02:34:53",
          "content": "<p>I also tried this implementation with HDBSCAN and saw worse results too</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331338,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/21/2018 03:37:04",
      "content": "<p>@Konstantin, thanks for the solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331359,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "05/21/2018 04:59:50",
      "content": "<p>Here's an idea I've been toying with (but haven't fully implemented):</p>\n\n<ul>\n<li>project to the $XY$-plane (so now helices look like circles and lines --well-- still look like lines),</li>\n<li>try to fit lines to points using RANSAC, Hough, DBSCAN, etc.,</li>\n<li>try to fit circles to points,</li>\n<li>try those last two steps in different orders (after filtering out previously classified lines/circles).</li>\n</ul>\n\n<p>Have you tried using projections? I've had some moderate success using projections (and, in general, different transformations of the data).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331446,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/21/2018 08:51:33",
      "content": "<p>you may want to check this projection. hope that i did not make any mistake.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331446/9475/projection.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 331495,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/21/2018 12:01:54",
          "content": "<p>oops. i am sorry that there is a mistake. </p>\n\n<p>phi = np.arctan2(y,x)/r is effectively zero</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331563,
          "author_name": "hrmello",
          "author_url": "",
          "post_date": "05/21/2018 13:44:56",
          "content": "<p>hey, Heng! Shouldn't the rotation be something like \nhx = x*np.cos(phi*r) - y*np.sin(phi*r)\nand\nhy = x*np.sin(phi*r) + y*np.cos(phi*r) ?\nI mean, from your plot it clearly works, but I'm' not understanding why it works if the rotation matrix is written differently from the above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331617,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/21/2018 15:20:53",
          "content": "<p>your equation is correct. </p>\n\n<p>Mine is a mistake. It simply project z to a plane. x and y are not used at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332232,
          "author_name": "jliamfinnie",
          "author_url": "",
          "post_date": "05/22/2018 20:35:26",
          "content": "<p>Hey <a href=\"/hengck23\">@hengck23</a>, FYI, we tried out your original code (the one you said now has the wrong equation), as well as the 'fixed' version. For us, the original code produced better results - i.e. after implementing for all combinations (x&gt;0, x&lt;0, y&gt;0, y&lt;0, z&gt;0, z&lt;0), the 'basic' DBSCAN kernel score improved from 0.204 to 0.209 (for your selected event 1029), and our HDBSCAN score improved from 0.250 to 0.258. The 'fixed' equations produced only a very small benefit on top of the base DBSCAN/HDBSCAN models. Of course, now the only problem is how to re-create without having the ground-truth tracks and momentum :-) I guess this is where <a href=\"/lopuhin\">@lopuhin</a> and <a href=\"/sionek\">@sionek</a>'s suggestions can be used....</p>\n\n<p>However, from reading some of the other research material (<a href=\"https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf\">https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf</a>), it sounds like this approach (conformal mapping and histogramming) seems to be discouraged, or at least is seen to have minimal benefit. I guess maybe this can still be useful as a 'first step', used in combination with other approaches....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331620,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/21/2018 15:31:13",
      "content": "<p>@Henrique Mello</p>\n\n<p>corrected code and results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331620/9481/corrected.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331635,
      "author_name": "dmriser",
      "author_url": "",
      "post_date": "05/21/2018 16:01:38",
      "content": "<p>In the x-y plane tracks which pass through the origin (it seems most of them) can be made linear by making a conformal mapping.\n$$\nu = \\frac{x}{x^2 + y^2},  \\\nv = \\frac{y}{x^2 + y^2} <br>\n$$\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331635/9482/conformal.png\" alt=\"Figure\"></p>\n\n<p>Unlike shown above, this unrolling procedure <em>might</em> work because it straightens only a subset of tracks out at a time.  This increases the density for only a subset in the clustered space which can be iteratively removed.  If this is the case, it might be interesting to try and learn the distribution of this unrolling angle from the truth dataset, and concentrate more iterations of clustering in those regions with more tracks (ie: choose a non-uniform grid of angles to unroll).  </p>",
      "votes": null,
      "replies": [
        {
          "id": 333683,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "05/25/2018 17:05:29",
          "content": "<p>I've done some experiments with this map. I'll return the favor --here is another one for you to try:</p>\n\n<p>$$\nx,y \\quad \\mapsto \\quad \\theta, \\log(r),\n$$</p>\n\n<p>where </p>\n\n<p>$$\nr := \\sqrt{x^2+y^2} \\quad , \\quad \\theta := \\arctan \\left( \\frac{y}{x} \\right) .\n$$</p>\n\n<p>This maps points in the punctured plane lying on a ray to vertical lines in a plane.</p>\n\n<p>One can then run a quicker version of the Hough transform (or other clustering/line-identifying method) which only captures vertical lines.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333717,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "05/25/2018 18:16:11",
          "content": "<p>Side note: sometimes you need to be careful about where the \"cut\" occurs (use a different theta). \nThe noise can cause what should be one vertical line to map to two different vertical lines.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331667,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/21/2018 17:09:14",
      "content": "<p>@Konstantin, thanks a lot for sharing your findings, this is triggering lots of thoughts for me.  Hope they'll translate into something effective.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332021,
      "author_name": "makahana",
      "author_url": "",
      "post_date": "05/22/2018 11:17:14",
      "content": "<p>Great stuff. Any basic ideas with regards to merging clusters? Say you run clustering 10 times with different angles. How to choose which of the clusterings to use for a particular hit?</p>",
      "votes": null,
      "replies": [
        {
          "id": 332026,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/22/2018 11:29:15",
          "content": "<p>@maka, I use similar approach as Konstantin. Because more energetic particles and bigger tracks are more prefered by scoring function, I start from bigger angles and finish at 0 angle changing assigment of the hit if the present cluster is bigger than the previous one.</p>\n\n<p>Edit: (Cluster/track) bigger in terms of number of hits. It is a very simple model. Just to start and find problems.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332029,
          "author_name": "makahana",
          "author_url": "",
          "post_date": "05/22/2018 11:35:24",
          "content": "<p><a href=\"/sionek\">@sionek</a> Bigger in terms of number of hits in cluster or in terms of cluster std.dev in helix coordinates? Would assume one would like the first high and the latter low?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332090,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/22/2018 13:56:26",
          "content": "<p>My merging algorithm is a bit involved. First I calculate the size of all clusters, and discard clusters bigger than 20 hits. After that I eagerly assign each hit to the biggest cluster it belongs too (this also assigns all other hits belonging to this cluster), starting from the points with highest number of clusters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332102,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/22/2018 14:36:23",
          "content": "<p>@Konstantin, It would be nice to compare someday our models and dates of their creation. Just to study telepathy. A piece of my code regarding discarding clusters bigger than 20 (hm, rather 19 ;)</p>\n\n<blockquote>\n  <p>dfh[,s1:=ifelse(N2&gt;N1 &amp; N2&lt;20,s2+maxs1,s1)]</p>\n</blockquote>\n\n<p>(N2 and N1 - number of hits in present and previous cluster. s1 and s2 - previous and present cluster number). The difference - in my model after removing a hit from the previous cluster, it becomes smaller (less important) and then it is easier to remove other hits from it, but the final effect is probably similar. Best wishes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332116,
          "author_name": "makahana",
          "author_url": "",
          "post_date": "05/22/2018 15:02:10",
          "content": "<p>Thanks @Konstantin <a href=\"/sionek\">@sionek</a> . </p>\n\n<p>I need to just get a basic solution up and running myself. But here is one idea you might consider:</p>\n\n<p>From what I can understand from the equations(see my kernel), the angle that needs to be unrolled is $qB(z-z0)/p_{0,z}$. q is charge, B is magnetic field( -0.00057 is a good value for first event in train_1). Since we know what the distribution of p_{0,z} should be approximately, maybe we can use this to choose angles or merge clusters better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332134,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/22/2018 15:47:01",
          "content": "<p>@maka, thanks for your kernel.  I have been looking at the exact same approach in the past few days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332451,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/23/2018 07:13:04",
          "content": "<p><a href=\"/sionek\">@sionek</a> haha, yeah I think we have a curious case indeed ;) Here is my code, to be fair I checked a few different values (12 -- 24):</p>\n\n<pre><code>_, inverse, counts = np.unique(labels, return_inverse=True, return_counts=True)\ncounts = counts[inverse]\ncounts[labels == -1] = 0\ncounts[counts &gt; 20] = 0  # important: if the cluster is large, it's bad\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332157,
      "author_name": "vicensgaitan",
      "author_url": "",
      "post_date": "05/22/2018 16:26:20",
      "content": "<p>@Grzegorz &amp; @Konstantin <br>\nIt seems telepathic waves are flowing around...  In my case, 20 is working better  ;)</p>\n\n<pre><code>sel=function(x,Nx,y,Ny,N){\n    I=x\n    I[Ny&gt;Nx &amp; Ny&lt;N]=max(x)+y[Ny&gt;Nx &amp; Ny &lt;N]\n    I[x==-1]=max(x)+y[x==-1]\n    I\n  }\n mm[,I:=sel(I.x,Nx,I.y,Ny,21)]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 332213,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/22/2018 19:23:23",
          "content": "<p>@Vicens, I'm afraid, these waves flow around Europe only or Mickey (Japan) uses \"skupiacz myśli\" (~thought concentrator) - a kind of deflector disabling thoughts to go away from the head, usually made of steel or kevlar, used in the army. I can't hear even a short haiku from him.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332714,
      "author_name": "xmaayy",
      "author_url": "",
      "post_date": "05/23/2018 16:00:21",
      "content": "<p>Neat1</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 333005,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/24/2018 07:57:35",
      "content": "<p>Analyzing @Grzegorz Sionkowski kernel (0.3472 DBSCAN).</p>\n\n<p>conclusion:</p>\n\n<ul>\n<li><p>the tracks are rather long (good for back-fitting lines)</p></li>\n<li><p>the clustered hits are not completed,  missing the first or last hits of the ground truth tracks. This is where the weights are high.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/333005/9514/back%20fitting.png\" alt=\"enter image description here\"></p></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 333087,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "05/24/2018 11:30:30",
      "content": "<p>I see...Thanks a lot, this works awesome but question... Am I dumb or is it Ok to have no idea whats happening here? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 333104,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "05/24/2018 12:10:30",
          "content": "<p>@NxGTR, European matters. CERN has to prepare to the changes of EU regulations. Officially, it is called General Data Protection Regulation but there is a hidden directive: every curve must be straight. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333181,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/24/2018 15:16:20",
          "content": "<blockquote>\n  <p>is it Ok to have no idea whats happening here?</p>\n</blockquote>\n\n<p>Are you saying you use a public kernel as is? ;)</p>\n\n<p>I feel free to add comments given I did not submit yet.  I may shut up for a while after getting a miserable score ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333186,
          "author_name": "carloshuertas",
          "author_url": "",
          "post_date": "05/24/2018 15:21:59",
          "content": "<p>Amm... let's say I was busy with other stuff and accidentally happened, proof that that is totally possible <a href=\"https://en.wikipedia.org/wiki/Infinite_monkey_theorem\">here</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 371829,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/17/2018 16:49:54",
      "content": "<p>Hi, No one ever thanked you for this idea, but I think it is what made DBSCAN a good option here.  Therefore a big THANK YOU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 371879,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/17/2018 18:59:31",
          "content": "<p>I thanked in my kernel:</p>\n\n<p><strong>to Konstantin Lopuhin for pulishing the idea in</strong></p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57180\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57180</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371898,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/17/2018 19:48:25",
          "content": "<blockquote>\n  <p>I thanked in my kernel</p>\n</blockquote>\n\n<p>Great, you're better than me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "331193": "I'd like to share an approach I used in the last few submissions, getting 0.28 -- 0.38 depending on some details. This is quite far even from the current leaders, and I'm not sure how viable is this approach in the long run, but maybe it will be helpful. At least it's fast to run and does not require any training data.\n\nIdea is based on DBSCAN baseline: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark. It seems that it works best on relatively straight tracks which are more or less aligned with the z axis. It's possible to unroll helixes (make them straight) by rotating each hit around the z axis by the angle that is proportional to the distance from the z axis. The angle that unrolls the helix is different for each helix, so the idea is to:\n\n    for angle in ~20 evenly spaced angles:\n         x2, y2 &lt;- rotate x, y by angle * (distance to z), unrolling some helixes\n         do helix transform and dbscan clustering, as in the original kernel\n    merge clusters from all angles\n\nProbably this can be extended by doing other transformations, besides just rotations to unroll helixes, or post-processing of discovered clusters. But this has no chance to obtain tracks that start far from origin, it seems.",
    "331201": "Konstantin, It was my first approach. Fitting the parameters enables to obtain at least 0.42. Best wishes.\n\nEDIT: I use over 50 angles (nonlinearly distributed), eps being a function of angle and a little bit different approach to z.",
    "331220": "Glad others thought the same, I was actually working on that yesterday, but wanted to find the best way to unroll a helix to straight line (not just proportional to z or radius)",
    "331290": "Have you tried using HDBSCAN? I'm just a beginner, but from my experience so far on the 0.2x range, this algorithm works better than DBSCAN. Does this situation change once some helices are straighten up?",
    "331338": "Konstantin, thanks for the solution!",
    "331359": "Here's an idea I've been toying with (but haven't fully implemented):\n\n * project to the $XY$-plane (so now helices look like circles and lines --well-- still look like lines),\n * try to fit lines to points using RANSAC, Hough, DBSCAN, etc.,\n * try to fit circles to points,\n * try those last two steps in different orders (after filtering out previously classified lines/circles).\n\nHave you tried using projections? I've had some moderate success using projections (and, in general, different transformations of the data).",
    "331387": "I briefly tried HDBSCAN with this approach and it worked worse, I think the reason is that it gives more false matches, which do not allow to do proper cluster merging. But maybe I did something wrong, please share your results if you try it.",
    "331446": "you may want to check this projection. hope that i did not make any mistake.\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331446/9475/projection.png",
    "331495": "oops. i am sorry that there is a mistake. \n\nphi = np.arctan2(y,x)/r is effectively zero",
    "331563": "hey, Heng! Shouldn't the rotation be something like \nhx = x*np.cos(phi*r) - y*np.sin(phi*r)\nand\nhy = x*np.sin(phi*r) + y*np.cos(phi*r) ?\nI mean, from your plot it clearly works, but I'm' not understanding why it works if the rotation matrix is written differently from the above.",
    "331617": "your equation is correct. \n\nMine is a mistake. It simply project z to a plane. x and y are not used at all.",
    "331620": "Henrique Mello\n \ncorrected code and results\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331620/9481/corrected.png",
    "331635": "In the x-y plane tracks which pass through the origin (it seems most of them) can be made linear by making a conformal mapping.\n$$\nu = \\frac{x}{x^2 + y^2},  \\\\\nv = \\frac{y}{x^2 + y^2}  \n$$\n![Figure][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/331635/9482/conformal.png\n\nUnlike shown above, this unrolling procedure *might* work because it straightens only a subset of tracks out at a time.  This increases the density for only a subset in the clustered space which can be iteratively removed.  If this is the case, it might be interesting to try and learn the distribution of this unrolling angle from the truth dataset, and concentrate more iterations of clustering in those regions with more tracks (ie: choose a non-uniform grid of angles to unroll).",
    "331667": "Konstantin, thanks a lot for sharing your findings, this is triggering lots of thoughts for me.  Hope they'll translate into something effective.",
    "332021": "Great stuff. Any basic ideas with regards to merging clusters? Say you run clustering 10 times with different angles. How to choose which of the clusterings to use for a particular hit?",
    "332026": "maka, I use similar approach as Konstantin. Because more energetic particles and bigger tracks are more prefered by scoring function, I start from bigger angles and finish at 0 angle changing assigment of the hit if the present cluster is bigger than the previous one.\n\nEdit: (Cluster/track) bigger in terms of number of hits. It is a very simple model. Just to start and find problems.",
    "332029": "sionek Bigger in terms of number of hits in cluster or in terms of cluster std.dev in helix coordinates? Would assume one would like the first high and the latter low?",
    "332090": "My merging algorithm is a bit involved. First I calculate the size of all clusters, and discard clusters bigger than 20 hits. After that I eagerly assign each hit to the biggest cluster it belongs too (this also assigns all other hits belonging to this cluster), starting from the points with highest number of clusters.",
    "332102": "Konstantin, It would be nice to compare someday our models and dates of their creation. Just to study telepathy. A piece of my code regarding discarding clusters bigger than 20 (hm, rather 19 ;)\n\n&gt; dfh[,s1:=ifelse(N2&gt;N1 &amp; N2&lt;20,s2+maxs1,s1)]\n\n(N2 and N1 - number of hits in present and previous cluster. s1 and s2 - previous and present cluster number). The difference - in my model after removing a hit from the previous cluster, it becomes smaller (less important) and then it is easier to remove other hits from it, but the final effect is probably similar. Best wishes.",
    "332116": "Thanks @Konstantin @sionek . \n\nI need to just get a basic solution up and running myself. But here is one idea you might consider:\n\nFrom what I can understand from the equations(see my kernel), the angle that needs to be unrolled is $qB(z-z0)/p_{0,z}$. q is charge, B is magnetic field( -0.00057 is a good value for first event in train_1). Since we know what the distribution of p_{0,z} should be approximately, maybe we can use this to choose angles or merge clusters better.",
    "332134": "maka, thanks for your kernel.  I have been looking at the exact same approach in the past few days.",
    "332157": "Grzegorz &amp; @Konstantin  \nIt seems telepathic waves are flowing around...  In my case, 20 is working better  ;)\n\n    sel=function(x,Nx,y,Ny,N){\n        I=x\n        I[Ny&gt;Nx &amp; Ny",
    "332213": "Vicens, I'm afraid, these waves flow around Europe only or Mickey (Japan) uses \"skupiacz myśli\" (~thought concentrator) - a kind of deflector disabling thoughts to go away from the head, usually made of steel or kevlar, used in the army. I can't hear even a short haiku from him.",
    "332232": "Hey @hengck23, FYI, we tried out your original code (the one you said now has the wrong equation), as well as the 'fixed' version. For us, the original code produced better results - i.e. after implementing for all combinations (x&gt;0, x&lt;0, y&gt;0, y&lt;0, z&gt;0, z&lt;0), the 'basic' DBSCAN kernel score improved from 0.204 to 0.209 (for your selected event 1029), and our HDBSCAN score improved from 0.250 to 0.258. The 'fixed' equations produced only a very small benefit on top of the base DBSCAN/HDBSCAN models. Of course, now the only problem is how to re-create without having the ground-truth tracks and momentum :-) I guess this is where @lopuhin and @sionek's suggestions can be used....\n\nHowever, from reading some of the other research material ([https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf][1]), it sounds like this approach (conformal mapping and histogramming) seems to be discouraged, or at least is seen to have minimal benefit. I guess maybe this can still be useful as a 'first step', used in combination with other approaches....\n\n  [1]: https://www.physik.uni-heidelberg.de/c/image/exp/f/highrr/Kisel_HD_12.04.2016.pdf",
    "332451": "sionek haha, yeah I think we have a curious case indeed ;) Here is my code, to be fair I checked a few different values (12 -- 24):\n \n    _, inverse, counts = np.unique(labels, return_inverse=True, return_counts=True)\n    counts = counts[inverse]\n    counts[labels == -1] = 0\n    counts[counts &gt; 20] = 0  # important: if the cluster is large, it's bad",
    "332714": "Neat1",
    "332899": "I also tried this implementation with HDBSCAN and saw worse results too",
    "333005": "Analyzing @Grzegorz Sionkowski kernel (0.3472 DBSCAN).\n\nconclusion:\n\n  - the tracks are rather long (good for back-fitting lines)\n  \n  - the clustered hits are not completed,  missing the first or last hits of the ground truth tracks. This is where the weights are high.\n  \n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/333005/9514/back%20fitting.png",
    "333087": "I see...Thanks a lot, this works awesome but question... Am I dumb or is it Ok to have no idea whats happening here? :)",
    "333104": "NxGTR, European matters. CERN has to prepare to the changes of EU regulations. Officially, it is called General Data Protection Regulation but there is a hidden directive: every curve must be straight.",
    "333181": "&gt; is it Ok to have no idea whats happening here?\n\nAre you saying you use a public kernel as is? ;)\n\nI feel free to add comments given I did not submit yet.  I may shut up for a while after getting a miserable score ;)",
    "333186": "Amm... let's say I was busy with other stuff and accidentally happened, proof that that is totally possible [here][1].\n\n\n  [1]: https://en.wikipedia.org/wiki/Infinite_monkey_theorem",
    "333683": "I've done some experiments with this map. I'll return the favor --here is another one for you to try:\n\n$$\nx,y \\quad \\mapsto \\quad \\theta, \\log(r),\n$$\n\nwhere \n\n$$\nr := \\sqrt{x^2+y^2} \\quad , \\quad \\theta := \\arctan \\left( \\frac{y}{x} \\right) .\n$$\n\nThis maps points in the punctured plane lying on a ray to vertical lines in a plane.\n\nOne can then run a quicker version of the Hough transform (or other clustering/line-identifying method) which only captures vertical lines.",
    "333717": "Side note: sometimes you need to be careful about where the \"cut\" occurs (use a different theta). \nThe noise can cause what should be one vertical line to map to two different vertical lines.",
    "334631": "Did you find that making eps as a function of angle increased the score? It seemed to be detrimental in my version.\nAlso, I tried to fork your kernel but could import dbscan.  Did you simply add a custom github repo?",
    "334636": "You can increment EPS and angle at the same time and get better results, but incrementing them both by small steps and using different step sizes for each. \n\nSee the high scoring public kernels based on dbscan for reference.",
    "334643": "Yeah, I am using a python implementation similar to Heng's of Grzegorz's kernel found here: [https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843][1]\n\nHowever, when I add the optimizations to eps like,\n\n    eps = self.eps + (i * 0.000005)\n\nit made the model perform worse\n\n  [1]: https://www.kaggle.com/sionek/mod-dbscan-x-100-parallel/comments#332843",
    "334650": "I tried tweaking the similar line in the notebook by Grzegorz Sionkowski\n\n'res=dbscan(dfs,eps=0.0035+ii*stepeps'...\nwhere stepeps = 0.000005\n\nand most of the tweaks I tried hurt performance, making me think he already did a decent job of tuning the scaling of eps with angle, and started looking elsewhere. Are you sure you're passing a reasonable value in with self.eps? \n\n...strange... sorry I'm not of more help.",
    "334661": "MathewMasters, @macfarll, Talking about parameters of my model, I was talking about the model of the score 0.42 (there are some ideas I want to keep for myself) . Using that model I have obtained better results decreasing eps a little bit in each iteration. In the model 0.34-0.37 I just show how you can additionally fit the parameters and looking at the leaderboard I can see many Kagglers not only have found better parameters but also they have implemented their own ideas basing on the DBSCAN benchmark skeleton.\n\nMy advice. Be more open for ideas. Do not concentrate on my parameters and features. Take only the idea of iterations. Read about the model by @Konstantin - in his model (and in mine of the score 0.42) the angle is not the most important feature for DBSCAN, it is just a way to obtain new values of x and y.",
    "334904": "Grzegorz, thanks for your reply and advice, I appreciate it. Just one question, how were you able to import dbscan into the kaggle kernel?  A github repo?",
    "334911": "Sorry if I misspoke about your model. I was trying to refer to the kernel you published, and thought that since I put some time into understanding it I could try to help answer some questions. My way of understanding how it worked was mainly through tweaking and tinkering things and seeing what happened.\n\nThat being said, I obviously have a lot to learn. \n\nSince then I have been working on different ideas instead of just tweaking parameters, such as extending tracks where the end points are missing. \n\nSorry if I caused any trouble / was misleading.",
    "371829": "Hi, No one ever thanked you for this idea, but I think it is what made DBSCAN a good option here.  Therefore a big THANK YOU.",
    "371879": "I thanked in my kernel:\n\n**to Konstantin Lopuhin for pulishing the idea in**\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/57180",
    "371898": "&gt; I thanked in my kernel\n\nGreat, you're better than me!"
  },
  "source": "meta"
}