{
  "id": 61081,
  "title": "Finding features for DBSCAN",
  "url": "/competitions/trackml-particle-identification/discussion/61081",
  "author_name": "bilal2vec",
  "post_date": "2018-07-14T01:57:20.026000",
  "votes": 11,
  "comment_count": 67,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I was trying to find good features for DBSCAN and was wondering at what score is it a good idea to stop looking for better features and optimizing weights on your features and instead focus on track extension, z shifting and finding better ways to merge tracks?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": 356609,
      "postDate": "2018-07-14T01:57:20.027Z",
      "content": "<p>Hi,</p>\n\n<p>I was trying to find good features for DBSCAN and was wondering at what score is it a good idea to stop looking for better features and optimizing weights on your features and instead focus on track extension, z shifting and finding better ways to merge tracks?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\n\nI was trying to find good features for DBSCAN and was wondering at what score is it a good idea to stop looking for better features and optimizing weights on your features and instead focus on track extension, z shifting and finding better ways to merge tracks?\n\nThanks",
      "votes": 11
    },
    {
      "id": 356653,
      "postDate": "2018-07-14T06:06:52.160Z",
      "content": "<p>If I remember correcty what I did, you can get significantly above 0.6 without z shifting, using DBSCAN, helix unrolling, and Heng's track extension code.</p>",
      "rawMarkdown": "If I remember correcty what I did, you can get significantly above 0.6 without z shifting, using DBSCAN, helix unrolling, and Heng's track extension code.",
      "votes": 1,
      "replies": [
        {
          "id": 356894,
          "postDate": "2018-07-14T18:11:32.420Z",
          "content": "<p>z shifting on top can get you above 0.7, but I doubt this can get above 0.8.</p>",
          "rawMarkdown": "z shifting on top can get you above 0.7, but I doubt this can get above 0.8."
        },
        {
          "id": 356906,
          "postDate": "2018-07-14T18:35:12.667Z",
          "content": "<p>you need supervised learning to get beyond 0.80</p>",
          "rawMarkdown": "you need supervised learning to get beyond 0.80\n\n\n"
        },
        {
          "id": 356916,
          "postDate": "2018-07-14T18:52:53.200Z",
          "content": "<p>I know, working on it since I'm back to this competition ;)</p>",
          "rawMarkdown": "I know, working on it since I'm back to this competition ;)"
        },
        {
          "id": 356925,
          "postDate": "2018-07-14T19:13:36.587Z",
          "content": "<p>What I just submitted was a reaction to Grzegorz saying we cannot get over 0.7 with merging based on track length only.  I just checked if this was true or not.</p>",
          "rawMarkdown": "What I just submitted was a reaction to Grzegorz saying we cannot get over 0.7 with merging based on track length only.  I just checked if this was true or not."
        },
        {
          "id": 357027,
          "postDate": "2018-07-15T05:42:07.463Z",
          "content": "<p>I can get beyond 0.75 selecting and extending tracks on length only.\nAnd the only other filtering I add is not allowing two hits on the exact same sensor.</p>",
          "rawMarkdown": "I can get beyond 0.75 selecting and extending tracks on length only.\nAnd the only other filtering I add is not allowing two hits on the exact same sensor.",
          "votes": 8
        },
        {
          "id": 359036,
          "postDate": "2018-07-19T11:12:22.957Z",
          "content": "<p>@yuval it seems your clustering approach is way more accurate than dbscan if you only use 2-3 features without heavy track fitting. </p>",
          "rawMarkdown": "@yuval it seems your clustering approach is way more accurate than dbscan if you only use 2-3 features without heavy track fitting. "
        },
        {
          "id": 359053,
          "postDate": "2018-07-19T11:57:38.170Z",
          "content": "<p>Actually I don't really do clustering, I use  bins (similar to the bins used in the Hough transform but implemented very efficiently). What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7.\nPlatu - can't get further with a clustering technique (neither bins nor dbscan) and need to use some track extending technique.</p>",
          "rawMarkdown": "Actually I don't really do clustering, I use  bins (similar to the bins used in the Hough transform but implemented very efficiently). What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7.\nPlatu - can't get further with a clustering technique (neither bins nor dbscan) and need to use some track extending technique.",
          "votes": 10
        },
        {
          "id": 359073,
          "postDate": "2018-07-19T12:51:55.820Z",
          "content": "<p>Interesting, I'm now at 0.7 without track extension, I guess you just confirm track extension should be my next focus.  And I'm now using 2 features as well if cos and sin counts for 1.  Maybe the same as yours actually.</p>",
          "rawMarkdown": "Interesting, I'm now at 0.7 without track extension, I guess you just confirm track extension should be my next focus.  And I'm now using 2 features as well if cos and sin counts for 1.  Maybe the same as yours actually."
        },
        {
          "id": 359139,
          "postDate": "2018-07-19T14:18:54.197Z",
          "content": "<p>0.7 using only 2 features and no track extension? Impressive! congrats</p>",
          "rawMarkdown": "0.7 using only 2 features and no track extension? Impressive! congrats",
          "votes": 1
        },
        {
          "id": 359204,
          "postDate": "2018-07-19T16:39:12.737Z",
          "content": "<p>@Giba, thanks, but I am sure the guys above 0.8 have way better tracks than me before they extend them.</p>",
          "rawMarkdown": "@Giba, thanks, but I am sure the guys above 0.8 have way better tracks than me before they extend them.",
          "votes": 1
        },
        {
          "id": 359354,
          "postDate": "2018-07-20T00:06:53.293Z",
          "content": "<p>with dbscan? that's still impresive</p>",
          "rawMarkdown": "with dbscan? that's still impresive"
        },
        {
          "id": 359365,
          "postDate": "2018-07-20T01:08:23.863Z",
          "content": "<p>How is z shifting implemented? Is it somewhat like the for loop for helix unrolling?</p>",
          "rawMarkdown": "How is z shifting implemented? Is it somewhat like the for loop for helix unrolling?"
        },
        {
          "id": 359368,
          "postDate": "2018-07-20T01:36:18.970Z",
          "content": "<p>@yuval r that is very impressive and congratulations. I am currently stuck below 0.6, would you care to elaborate on the following?</p>\n\n<p>&gt;  ...try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7 ...</p>",
          "rawMarkdown": "@yuval r that is very impressive and congratulations. I am currently stuck below 0.6, would you care to elaborate on the following?\n\n&gt;  ...try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7 ..."
        },
        {
          "id": 359383,
          "postDate": "2018-07-20T02:47:40.663Z",
          "content": "<p>The idea it to unroll the helix correctly. As I do not know the radius of the Helix (and the direction = the Q of the particle) I try different random values for the radius  (actually not for the radius but for 1/Radius).\nAlso the main collision area is +-5.5 mm around the origin in the Z axis =&gt; I try different random values for z0\nBecause I use binning and not dbscan, I need very accurate initial unrolling =&gt; I try a lot of different values.</p>\n\n<p>The way to do a very fast binning in python is using numpy.unique:</p>\n\n<p>un,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)</p>\n\n<p>hits['track_id']=inv</p>\n\n<p>hits['track_length']=count[inv]</p>\n\n<p>where hits['cat'] - is the binning variable.\nIn my main loop I do as little as possible - just adjust my features according to the random radius and z0, do the above binning and for every hit decide if to use this track_id according to track length.\nthis main loop takes less the 1min/1500 random values on my laptop = &gt;3 min to get 0.6 and about 40 to get to the plateau (get very little progress even after doubling the amount of values) </p>",
          "rawMarkdown": "The idea it to unroll the helix correctly. As I do not know the radius of the Helix (and the direction = the Q of the particle) I try different random values for the radius  (actually not for the radius but for 1/Radius).\nAlso the main collision area is +-5.5 mm around the origin in the Z axis =&gt; I try different random values for z0\nBecause I use binning and not dbscan, I need very accurate initial unrolling =&gt; I try a lot of different values.\n\nThe way to do a very fast binning in python is using numpy.unique:\n            \nun,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)\n\nhits['track_id']=inv\n\nhits['track_length']=count[inv]\n\nwhere hits['cat'] - is the binning variable.\nIn my main loop I do as little as possible - just adjust my features according to the random radius and z0, do the above binning and for every hit decide if to use this track_id according to track length.\nthis main loop takes less the 1min/1500 random values on my laptop = &gt;3 min to get 0.6 and about 40 to get to the plateau (get very little progress even after doubling the amount of values) \n ",
          "votes": 8
        },
        {
          "id": 359505,
          "postDate": "2018-07-20T08:22:55.960Z",
          "content": "<p><a href=\"/martin\">@martin</a>, yes, using dbscan.  My approach is very similar to what @yuval describes, including what he says about his features, but I'm using dbscan instead of binning.  It is taking way more time than what he does, at least 10x more time.  I plan to see if I can switch to binning, but thats not my priority. Priority is to improve score first, running time second ;)</p>\n\n<p>@btKaggle , we discussed z shifting a while back in the forum.  The idea is simple.  If you construct features for dbscan, you construct features as a function of hit coordinates, right?  So if your features are f(x, y, z), then shifting z means you compute features as f(x, y, z - z0), where z0 is the new origin for your tracks.  You then need to find ways to combine the tracks coming from varying z0.</p>",
          "rawMarkdown": "@martin, yes, using dbscan.  My approach is very similar to what @yuval describes, including what he says about his features, but I'm using dbscan instead of binning.  It is taking way more time than what he does, at least 10x more time.  I plan to see if I can switch to binning, but thats not my priority. Priority is to improve score first, running time second ;)\n\n@btKaggle , we discussed z shifting a while back in the forum.  The idea is simple.  If you construct features for dbscan, you construct features as a function of hit coordinates, right?  So if your features are f(x, y, z), then shifting z means you compute features as f(x, y, z - z0), where z0 is the new origin for your tracks.  You then need to find ways to combine the tracks coming from varying z0."
        },
        {
          "id": 359511,
          "postDate": "2018-07-20T08:39:07.420Z",
          "content": "<p>Thanks @yuval r, that is very clear and since it is faster than dbscan, I will try binning and see if I can improve my model.</p>",
          "rawMarkdown": "Thanks @yuval r, that is very clear and since it is faster than dbscan, I will try binning and see if I can improve my model."
        },
        {
          "id": 359774,
          "postDate": "2018-07-20T18:34:19.720Z",
          "content": "<p>@yuval r what do you mean by \"binning variable\"?</p>",
          "rawMarkdown": "@yuval r what do you mean by \"binning variable\"?"
        },
        {
          "id": 359827,
          "postDate": "2018-07-20T21:36:41.057Z",
          "content": "<p>@Martin I am doing a one dimensional binning (for efficiency). Let's say we have two features f1 and f2, We need to define F=f1+K*f2. Where K&gt;max(f1). Now we can do binning on F which I call the binning variable.\nAnother thing you must do is make F an integer (every bin is a different integer) to do this you can multiply by a large number and use .astype('int'). The large number is very important because it actually determine the size of the bins.</p>\n\n<p>To put it all together:</p>\n\n<p>f1,f2 are the features in the range [-1,1]</p>\n\n<p>L1,L2 are large integers</p>\n\n<p>K - an integer where K&gt;L1</p>\n\n<p>We define:</p>\n\n<p>F= (L1*f1). astype ('int')+K*(L2*f2).astype('int')</p>",
          "rawMarkdown": "@Martin I am doing a one dimensional binning (for efficiency). Let's say we have two features f1 and f2, We need to define F=f1+K*f2. Where K&gt;max(f1). Now we can do binning on F which I call the binning variable.\nAnother thing you must do is make F an integer (every bin is a different integer) to do this you can multiply by a large number and use .astype('int'). The large number is very important because it actually determine the size of the bins.\n\nTo put it all together:\n\nf1,f2 are the features in the range [-1,1]\n\nL1,L2 are large integers\n\nK - an integer where K&gt;L1\n\nWe define:\n\nF= (L1*f1). astype ('int')+K*(L2*f2).astype('int')",
          "votes": 2
        },
        {
          "id": 359889,
          "postDate": "2018-07-21T01:52:39.363Z",
          "content": "<p>Thanks for your comment, @yuval6769. Your approach using np.unique is so smart! May I ask you a small question? Do you use Hough transform by varying 1/r0 and calculate theta by theta = phi - arccos(r/2r0)? If so, your large integer L1, L2 and K be the same across different 1/r0 or you use adaptive L1, L2, K with respect to r0?  I am using that approach but face some difficulties by selecting the range for 1/r0. When doing that way, my 2 features will by z/r and theta, which can be plugged into your suggested equation for 1-D binning. I am suspecting the feasibility of this approach, because we need to be careful of the range of arccos and arctan2 functions, because they are used in one equation together ( theta = phi - arccos(r/2r0) = arctan2(y/x) - arccos(r/2r0) ). </p>",
          "rawMarkdown": "Thanks for your comment, @yuval6769. Your approach using np.unique is so smart! May I ask you a small question? Do you use Hough transform by varying 1/r0 and calculate theta by theta = phi - arccos(r/2r0)? If so, your large integer L1, L2 and K be the same across different 1/r0 or you use adaptive L1, L2, K with respect to r0?  I am using that approach but face some difficulties by selecting the range for 1/r0. When doing that way, my 2 features will by z/r and theta, which can be plugged into your suggested equation for 1-D binning. I am suspecting the feasibility of this approach, because we need to be careful of the range of arccos and arctan2 functions, because they are used in one equation together ( theta = phi - arccos(r/2r0) = arctan2(y/x) - arccos(r/2r0) ). "
        },
        {
          "id": 359989,
          "postDate": "2018-07-21T07:53:33.683Z",
          "content": "<p>@Kha A. Vo</p>\n\n<p>Although my approach is close to yours, I’m not using Hough transform but clustering (or binning) because Hough transform isn’t a suitable solution for this challenge. Let me start by explaining this statement and later I’ll address your issues.</p>\n\n<p>Why Hough transform isn’t suitable:</p>\n\n<p>•   Hough transform does not assume the tracks starts at the origin which leaves you with too many free variables. The Hough transform kernel decrease this number by assuming that in the Z axis it does start from the origin, also, it does not scan for different possible positions for the center of the helix in the XY plain.</p>\n\n<p>•   If you don’t assume the tracks starts at the origin, you’ll get a very large number of wrong tracks. The sensors are arranged in circular structures centered at the origin, hence you have an endless number of optional Helixes centered at the origin.</p>\n\n<p><strong>One remark about the origin assumption – about 18% of the tracks don’t start at the origin, 80% of these tracks couldn’t be found while using the origin assumption (I believe this is the reason the leaders are stuck around 0.82 - in some discussions it seems <a href=\"/outrunner\">@outrunner</a> solved this issue and we will soon se a 0.9 from him)</strong></p>\n\n<p>Now Let’s address your issues.</p>\n\n<p>Arccos(r/2r0) is a good and accurate choice, I believe it is better then: <code>theta = Phi + k*z</code>, you find in some kernals. The problem with arccos is that arccos(t) isn’t defined for |t|&gt;1. And for some choices of 1/r0 you will get r/2r0&gt;1 for some hits. You need to find a way to overcome this issue.</p>\n\n<p>On the other hand, z/r is not a good feature, as it has two major lacks. First it is unbounded and unevenly spread, and 2nd it does not take into account the curve of the track between (0,0) and (x,y).\nThe first issue can be easily solved using arctan() and for the 2nd issue you’ll need a little bit of geometry.  </p>\n\n<p>As for your question about changing L1, L2, K, I don’t change them for different values of 1/r0. I don’t change K at all, I do use more than one value for L1, L2 – which is equivalent to using different bin sizes but I do it regardless of the value of 1/r0.</p>",
          "rawMarkdown": "@Kha A. Vo\n\nAlthough my approach is close to yours, I’m not using Hough transform but clustering (or binning) because Hough transform isn’t a suitable solution for this challenge. Let me start by explaining this statement and later I’ll address your issues.\n\nWhy Hough transform isn’t suitable:\n\n•\tHough transform does not assume the tracks starts at the origin which leaves you with too many free variables. The Hough transform kernel decrease this number by assuming that in the Z axis it does start from the origin, also, it does not scan for different possible positions for the center of the helix in the XY plain.\n\n•\tIf you don’t assume the tracks starts at the origin, you’ll get a very large number of wrong tracks. The sensors are arranged in circular structures centered at the origin, hence you have an endless number of optional Helixes centered at the origin.\n\n**One remark about the origin assumption – about 18% of the tracks don’t start at the origin, 80% of these tracks couldn’t be found while using the origin assumption (I believe this is the reason the leaders are stuck around 0.82 - in some discussions it seems @outrunner solved this issue and we will soon se a 0.9 from him)**\n\nNow Let’s address your issues.\n\nArccos(r/2r0) is a good and accurate choice, I believe it is better then: `theta = Phi + k*z`, you find in some kernals. The problem with arccos is that arccos(t) isn’t defined for |t|&gt;1. And for some choices of 1/r0 you will get r/2r0&gt;1 for some hits. You need to find a way to overcome this issue.\n\nOn the other hand, z/r is not a good feature, as it has two major lacks. First it is unbounded and unevenly spread, and 2nd it does not take into account the curve of the track between (0,0) and (x,y).\nThe first issue can be easily solved using arctan() and for the 2nd issue you’ll need a little bit of geometry.  \n\nAs for your question about changing L1, L2, K, I don’t change them for different values of 1/r0. I don’t change K at all, I do use more than one value for L1, L2 – which is equivalent to using different bin sizes but I do it regardless of the value of 1/r0.\n",
          "votes": 3
        },
        {
          "id": 360025,
          "postDate": "2018-07-21T09:50:39.083Z",
          "content": "<p>Thanks @yuval r.  I read your comment and believe you may miss some points I made. First, you said \"I’m not using Hough transform but clustering (or binning)\". Indeed, I also used clustering, but by Hough features. Now I would like to shift it to the binning to avoid DBSCAN. Let me explain it further.</p>\n\n<p>1) The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0. This assumes that, with different i, all hits which belong to a true track, will have near Euclidean distance using features [sin(phi), cos(phi), ...some other features...]. With i = 0, phi is kept untouched, so straight tracks will be detected. With i large, phi is distorted, and hits lying in the curve tracks will have the same distorted phi.</p>\n\n<p>However, this is not extremely accurate, because phi is not exactly a linear function of r, or quadratic... Indeed, Hough transform did provide an exact function relating phi and r, that is r/r0 = 2cos(phi - theta).</p>\n\n<p>2) That equation is exact, and ALREADY assumes that the track starts at origin (0,0,z) in the XY plane!  It is because given any true track starts from origin (0,0,z), the track will draw a curve in the XY plane starting from (0,0).  This curve is an arc of a circle, whose diameter is 2r0 and the origin of this circle will have a specific and unique theta.  </p>\n\n<p>3) So, using the Hough equation and doing clustering with it, is surely a better method than using handcrafted equation. Now returning to point number 1) I made above. Varying i and calculating the new phi as mentioned in 1), will be equivalent to varying theta and calculating 1/r0 for each hit, by the Hough equation: 1/r0 = 2cos(phi-theta)/r. Then, using 1/r0 as a feature, all hits lying on the same true tracks with the provided theta, will have the same 1/r0. This will be clustered by DBSCAN. Indeed, the </p>\n\n<p>4) However, the approach 3) is good, but not smart. As you said, varying and sampling 1/r0 is better because we can use a Gaussian distribution for 1/r0 centered at 0. With varying 1/r0, now we switch the role of 1/r0 and theta. Now 1/r0 to be varied, and theta is to be calculated as a feature. With DBSCAN, this would be impossible to scan all possible values for 1/r0. Therefore, I want to shift to binning. Indeed, the binning with Hough features was also done by a public kernel and produce 0.1x on LB.</p>\n\n<p>5) With binning, however using K, L1, L2 as you said is very sensitive. I tried it, with just one value of 1/r0 = 0, and get the score of 0.002. This means that it is feasible to run for different 1/r0. But there are wrong tracks as you said.  Moreover, the hits binning results are very sensitive to large K, L1, L2(for example, hits with binning variables 23431, 23429, 23464 should be clustered together because the values are nearly equal. Now I need to find K, L1, L2 so that all of them must equal the unique integer (that is hard!!).</p>\n\n<p>Anyway, thanks for your answer. </p>",
          "rawMarkdown": "Thanks @yuval r.  I read your comment and believe you may miss some points I made. First, you said \"I’m not using Hough transform but clustering (or binning)\". Indeed, I also used clustering, but by Hough features. Now I would like to shift it to the binning to avoid DBSCAN. Let me explain it further.\n\n1) The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0. This assumes that, with different i, all hits which belong to a true track, will have near Euclidean distance using features [sin(phi), cos(phi), ...some other features...]. With i = 0, phi is kept untouched, so straight tracks will be detected. With i large, phi is distorted, and hits lying in the curve tracks will have the same distorted phi.\n\nHowever, this is not extremely accurate, because phi is not exactly a linear function of r, or quadratic... Indeed, Hough transform did provide an exact function relating phi and r, that is r/r0 = 2cos(phi - theta).\n\n2) That equation is exact, and ALREADY assumes that the track starts at origin (0,0,z) in the XY plane!  It is because given any true track starts from origin (0,0,z), the track will draw a curve in the XY plane starting from (0,0).  This curve is an arc of a circle, whose diameter is 2r0 and the origin of this circle will have a specific and unique theta.  \n\n3) So, using the Hough equation and doing clustering with it, is surely a better method than using handcrafted equation. Now returning to point number 1) I made above. Varying i and calculating the new phi as mentioned in 1), will be equivalent to varying theta and calculating 1/r0 for each hit, by the Hough equation: 1/r0 = 2cos(phi-theta)/r. Then, using 1/r0 as a feature, all hits lying on the same true tracks with the provided theta, will have the same 1/r0. This will be clustered by DBSCAN. Indeed, the \n\n4) However, the approach 3) is good, but not smart. As you said, varying and sampling 1/r0 is better because we can use a Gaussian distribution for 1/r0 centered at 0. With varying 1/r0, now we switch the role of 1/r0 and theta. Now 1/r0 to be varied, and theta is to be calculated as a feature. With DBSCAN, this would be impossible to scan all possible values for 1/r0. Therefore, I want to shift to binning. Indeed, the binning with Hough features was also done by a public kernel and produce 0.1x on LB.\n\n5) With binning, however using K, L1, L2 as you said is very sensitive. I tried it, with just one value of 1/r0 = 0, and get the score of 0.002. This means that it is feasible to run for different 1/r0. But there are wrong tracks as you said.  Moreover, the hits binning results are very sensitive to large K, L1, L2(for example, hits with binning variables 23431, 23429, 23464 should be clustered together because the values are nearly equal. Now I need to find K, L1, L2 so that all of them must equal the unique integer (that is hard!!).\n\nAnyway, thanks for your answer. \n",
          "votes": 3
        },
        {
          "id": 360027,
          "postDate": "2018-07-21T09:51:26.553Z",
          "content": "<p>Well well, if anyone is not above 0.7 by end of competition with this then something is wrong ;)</p>",
          "rawMarkdown": "Well well, if anyone is not above 0.7 by end of competition with this then something is wrong ;)"
        },
        {
          "id": 360050,
          "postDate": "2018-07-21T11:49:03.737Z",
          "content": "<p>@Kha I agree with 1) - 4) but not with 5). binning is not sensitive to the values of K, L1, L2. You should be careful when choosing K because if K is too small the different feature would blend together, but the values of L1, L2 can be changed by 10% without getting almost any difference in the score.</p>\n\n<p>There is one sensitivity in binning which could be easily explain by the following example: lets say we have only one feature: f1 and for a certain true track for some hits we get f1=1e-5 and others f1=-1e-5, although the values are very close,  these hits will never fall in the same bin if you do binning by <code>bin=(L1*f1).astype(int)</code>.  This issue has a few elegant solutions.</p>\n\n<p>Another thing to understand it that binning is much more sensitive to the values of 1/r0 and z0 because you need all hits to fall in the same bin, this is why I try so many different values.</p>\n\n<p>One last comment - the distribution of 1/r0 is definitely not Gaussian - 1/r0=0 means the particle is not rotating which means it has no charge - I think uncharged particles are invisible for the detectors. I don't know what is the real distribution, but it looks like gamma distribution of type 2, mirrored to the left because of the particles charge. <strong>The nice thing it you can use Gaussian distribution with no score penalty, and a small time penalty =&gt; I use a Gaussian distribution</strong>.</p>",
          "rawMarkdown": "@Kha I agree with 1) - 4) but not with 5). binning is not sensitive to the values of K, L1, L2. You should be careful when choosing K because if K is too small the different feature would blend together, but the values of L1, L2 can be changed by 10% without getting almost any difference in the score.\n\nThere is one sensitivity in binning which could be easily explain by the following example: lets say we have only one feature: f1 and for a certain true track for some hits we get f1=1e-5 and others f1=-1e-5, although the values are very close,  these hits will never fall in the same bin if you do binning by `bin=(L1*f1).astype(int)`.  This issue has a few elegant solutions.\n\nAnother thing to understand it that binning is much more sensitive to the values of 1/r0 and z0 because you need all hits to fall in the same bin, this is why I try so many different values.\n\n\nOne last comment - the distribution of 1/r0 is definitely not Gaussian - 1/r0=0 means the particle is not rotating which means it has no charge - I think uncharged particles are invisible for the detectors. I don't know what is the real distribution, but it looks like gamma distribution of type 2, mirrored to the left because of the particles charge. **The nice thing it you can use Gaussian distribution with no score penalty, and a small time penalty =&gt; I use a Gaussian distribution**.",
          "votes": 2
        },
        {
          "id": 360052,
          "postDate": "2018-07-21T11:53:13Z",
          "content": "<p>@Kha, you seem to miss that hough transform maps each hit to a curve, not just a point in another feature space.  You use the equation of a circle using polar coordinates, which is indeed an ingredient of hough transform, but not all of it.  See <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60824#355799\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60824#355799</a></p>",
          "rawMarkdown": "@Kha, you seem to miss that hough transform maps each hit to a curve, not just a point in another feature space.  You use the equation of a circle using polar coordinates, which is indeed an ingredient of hough transform, but not all of it.  See https://www.kaggle.com/c/trackml-particle-identification/discussion/60824#355799",
          "votes": 1
        },
        {
          "id": 360054,
          "postDate": "2018-07-21T11:57:35.373Z",
          "content": "<p>@CPMP you are correct, but I believe it's time this competition will start to heat up a little as it seems to be stuck. More Kagglers  above 0.7 means more ideas and more knowledge and this is what it is all about. :)</p>\n\n<p>(But to get to 0.7 ,even with all the hints around  one really needs to work. If I was  one of the organizers I would have started with an example kernel which get 0.7 because here is where all the fun begins).   </p>",
          "rawMarkdown": "@CPMP you are correct, but I believe it's time this competition will start to heat up a little as it seems to be stuck. More Kagglers  above 0.7 means more ideas and more knowledge and this is what it is all about. :)\n\n(But to get to 0.7 ,even with all the hints around  one really needs to work. If I was  one of the organizers I would have started with an example kernel which get 0.7 because here is where all the fun begins).   ",
          "votes": 4
        },
        {
          "id": 360058,
          "postDate": "2018-07-21T12:18:10.263Z",
          "content": "<p>I don't think competition is stuck by any mean, at least if I look at my local score.  But thing is people aren't submitting much because you can use local score quite effectively, and because it is computationally intensive. They may need the compute resource to try new ideas instead of creating a submission.</p>\n\n<p>I am not against sharing, and I do my share of it, but I am against people getting medals just be replicating what has been shared.  I'm afraid you are offering medals on a plate here.  </p>\n\n<p>This said, this is my opinion, and just my opinion.  You can safely ignore it.  And there are probably a bunch of people who liked your sharing.</p>",
          "rawMarkdown": "I don't think competition is stuck by any mean, at least if I look at my local score.  But thing is people aren't submitting much because you can use local score quite effectively, and because it is computationally intensive. They may need the compute resource to try new ideas instead of creating a submission.\n\nI am not against sharing, and I do my share of it, but I am against people getting medals just be replicating what has been shared.  I'm afraid you are offering medals on a plate here.  \n\nThis said, this is my opinion, and just my opinion.  You can safely ignore it.  And there are probably a bunch of people who liked your sharing.",
          "votes": 1
        },
        {
          "id": 360059,
          "postDate": "2018-07-21T12:21:01.137Z",
          "content": "<p>Thanks @yuval. I believe my statement stating that choosing K, L1, L2 is sensitive, is because I haven't tried enough, and also I may use the inappropriate features. I will try the old features of DBSCAN clustering with sin, cos, and do it with binning, to see its efficiency. That is totally worth it. Before your share about this, I never believe one could do 10K iterations for merging tracks. </p>\n\n<p>@CPMP. Thanks. Indeed I read your thread there all. In there, <a href=\"/trian2018\">@trian2018</a> said </p>\n\n<p>\"  <em>At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).</em> \"</p>\n\n<p>That is, to my knowledge, equivalent to the binning or clustering method of @Grzegorz. I agree to your statement that \"hough transform maps each hit to a curve, not just a point in another feature space\". However we did the clustering by detecting the nearby transformed hits at each iteration (the intersection point).</p>",
          "rawMarkdown": "Thanks @yuval. I believe my statement stating that choosing K, L1, L2 is sensitive, is because I haven't tried enough, and also I may use the inappropriate features. I will try the old features of DBSCAN clustering with sin, cos, and do it with binning, to see its efficiency. That is totally worth it. Before your share about this, I never believe one could do 10K iterations for merging tracks. \n\n@CPMP. Thanks. Indeed I read your thread there all. In there, @trian2018 said \n\n\"  *At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).* \"\n\nThat is, to my knowledge, equivalent to the binning or clustering method of @Grzegorz. I agree to your statement that \"hough transform maps each hit to a curve, not just a point in another feature space\". However we did the clustering by detecting the nearby transformed hits at each iteration (the intersection point)."
        },
        {
          "id": 360060,
          "postDate": "2018-07-21T12:21:18.050Z",
          "content": "<p>&gt; I would have started with an example kernel which get 0.7 *</p>\n\n<p>This assumes they could produce one at the start of the competition, which I doubt.  This data is way more dense and complex than what the research community has been working with so far.  The fact that it took 2 month for a specialist like Edwin Steiner to move to 0.8 shows it is hard, even for specialists.</p>",
          "rawMarkdown": "&gt; I would have started with an example kernel which get 0.7 *\n\nThis assumes they could produce one at the start of the competition, which I doubt.  This data is way more dense and complex than what the research community has been working with so far.  The fact that it took 2 month for a specialist like Edwin Steiner to move to 0.8 shows it is hard, even for specialists.",
          "votes": 1
        },
        {
          "id": 360062,
          "postDate": "2018-07-21T12:27:50.560Z",
          "content": "<p>@CPMP What do you mean by \"specialist\"?</p>",
          "rawMarkdown": "@CPMP What do you mean by \"specialist\"?"
        },
        {
          "id": 360065,
          "postDate": "2018-07-21T12:29:41.753Z",
          "content": "<p>Someone working on track reconstruction for high energy physics</p>",
          "rawMarkdown": "Someone working on track reconstruction for high energy physics"
        },
        {
          "id": 360068,
          "postDate": "2018-07-21T12:35:54.507Z",
          "content": "<p>@CPMP  - Don't worry about it. We didn't use any of these approaches, simply DBSCAN with the chemist's features <code>sin, cos, z1, z2</code> + z-shifting + handcrafted merging code to pass the 0.7 plateau. Right now we're working on track extension and exploring supervised learning approaches. </p>\n\n<p>A couple months ago, when the chemist wanted to share a kernel with z-shifting, we kind of agreed that we wouldn't share critical approaches in the final month and we still keep that consensus.  I think it's more fair to others who have worked really hard in this competition.</p>\n\n<p><strong>UPDATE on July 31st</strong>\nI finally figured out the secret sauce after having read this, pretty clever. I used to have trouble calculating the 1/r0 before and only ran 100 iterations of 1/r0 and the score was pretty bad and gave up on it, so I handcrafted 2 quadratic functions to approximate the new unrolling angles. Thanks to <strong>@yuval, @Kha and @CPMP</strong> sharing your thoughts here.</p>",
          "rawMarkdown": "@CPMP  - Don't worry about it. We didn't use any of these approaches, simply DBSCAN with the chemist's features `sin, cos, z1, z2 ` + z-shifting + handcrafted merging code to pass the 0.7 plateau. Right now we're working on track extension and exploring supervised learning approaches. \n\nA couple months ago, when the chemist wanted to share a kernel with z-shifting, we kind of agreed that we wouldn't share critical approaches in the final month and we still keep that consensus.  I think it's more fair to others who have worked really hard in this competition.\n\n**UPDATE on July 31st**\nI finally figured out the secret sauce after having read this, pretty clever. I used to have trouble calculating the 1/r0 before and only ran 100 iterations of 1/r0 and the score was pretty bad and gave up on it, so I handcrafted 2 quadratic functions to approximate the new unrolling angles. Thanks to **@yuval, @Kha and @CPMP** sharing your thoughts here.",
          "votes": 2
        },
        {
          "id": 360074,
          "postDate": "2018-07-21T13:30:07.597Z",
          "content": "<p>@Nicole, I don't worry ;)  Other than that I agree with you.</p>",
          "rawMarkdown": "@Nicole, I don't worry ;)  Other than that I agree with you."
        },
        {
          "id": 360522,
          "postDate": "2018-07-22T16:37:12.083Z",
          "content": "<blockquote>\n  <p>What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7. </p>\n</blockquote>\n\n<p>@yuval\nDo you mean that the total number of radiuses and Z0 is 4500? (i.e. the sum, not every object).</p>\n\n<p>I used only about 200 radiuses, and five z0. It takes 20 minutes for an event, and I consider this slow. Maybe I should move further for better score. I'm also trying ML, but no success so far.</p>",
          "rawMarkdown": "&gt; What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7. \n\n@yuval\nDo you mean that the total number of radiuses and Z0 is 4500? (i.e. the sum, not every object).\n\nI used only about 200 radiuses, and five z0. It takes 20 minutes for an event, and I consider this slow. Maybe I should move further for better score. I'm also trying ML, but no success so far."
        },
        {
          "id": 360535,
          "postDate": "2018-07-22T16:56:38.063Z",
          "content": "<p>I mean 4500 combinations. my combinations are random pairs (I use Gaussian distribution for each).\nI understand you are trying 200*5=1000 pairs. \nAs I wrote I am using binning and not dbscan =&gt; it takes only 3 min for an event to get about 0.6, It takes about 40 min to check enough pairs to get to about 0.7. \nI don't think any ML is needed until you get to 0.7 (as @CPMP also confirmed) going beyond 0.7 is a different story.  </p>",
          "rawMarkdown": "I mean 4500 combinations. my combinations are random pairs (I use Gaussian distribution for each).\nI understand you are trying 200*5=1000 pairs. \nAs I wrote I am using binning and not dbscan =&gt; it takes only 3 min for an event to get about 0.6, It takes about 40 min to check enough pairs to get to about 0.7. \nI don't think any ML is needed until you get to 0.7 (as @CPMP also confirmed) going beyond 0.7 is a different story.  ",
          "votes": 2
        },
        {
          "id": 360546,
          "postDate": "2018-07-22T17:13:46.490Z",
          "content": "<p>@yuval Thanks, now it's clear to me. Yes, I used &lt; 1000 pairs. I will try more.\nYour approach is quite fast! But I'm afraid there's no time to change to your variant (I'm using DBSCAN).</p>",
          "rawMarkdown": "@yuval Thanks, now it's clear to me. Yes, I used &lt; 1000 pairs. I will try more.\nYour approach is quite fast! But I'm afraid there's no time to change to your variant (I'm using DBSCAN)."
        },
        {
          "id": 360561,
          "postDate": "2018-07-22T17:45:25.953Z",
          "content": "<p>@Serguey, using 20,000 pairs with DBSCAN moves me over 0.71 locally (my LB score is usually higher than the local score) before any track extension.  But running times are way longer than @yuval's, about 20 minutes to get to 0.6 on a single core.  This is to say that there is hope with DBSCAN, but fast binning as @yuval uses is tempting.</p>",
          "rawMarkdown": "@Serguey, using 20,000 pairs with DBSCAN moves me over 0.71 locally (my LB score is usually higher than the local score) before any track extension.  But running times are way longer than @yuval's, about 20 minutes to get to 0.6 on a single core.  This is to say that there is hope with DBSCAN, but fast binning as @yuval uses is tempting.",
          "votes": 2
        },
        {
          "id": 360615,
          "postDate": "2018-07-22T20:58:36.483Z",
          "content": "<p>@CPMP Thank you for info! What's about time for 20,000 pairs? Outrunner wrote about 10-20 hours for an event. This shocked me! :) </p>",
          "rawMarkdown": "@CPMP Thank you for info! What's about time for 20,000 pairs? Outrunner wrote about 10-20 hours for an event. This shocked me! :) "
        },
        {
          "id": 360619,
          "postDate": "2018-07-22T21:06:32.490Z",
          "content": "<p>About 9-10 hours or so on a single core.</p>",
          "rawMarkdown": "About 9-10 hours or so on a single core.",
          "votes": 1
        },
        {
          "id": 360621,
          "postDate": "2018-07-22T21:14:33.053Z",
          "content": "<p>I’m lucky, I have access to some wicked machinery, but only have an intermediate level of knowledge, so I have a .65 local score but I’ve been struggling to implement a supervised learning approach for weeks 😭</p>",
          "rawMarkdown": "I’m lucky, I have access to some wicked machinery, but only have an intermediate level of knowledge, so I have a .65 local score but I’ve been struggling to implement a supervised learning approach for weeks 😭"
        },
        {
          "id": 360689,
          "postDate": "2018-07-23T02:49:58.907Z",
          "content": "<p>20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. The bad news is the upcoming update data. :(</p>",
          "rawMarkdown": "20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. The bad news is the upcoming update data. :(",
          "votes": 1
        },
        {
          "id": 360739,
          "postDate": "2018-07-23T06:16:52.087Z",
          "content": "<blockquote>\n  <p>I’ve been struggling to implement a supervised learning approach for weeks</p>\n</blockquote>\n\n<p>That's one of the challenges here.  I am not sure one cannot get to 0.8 without supervised learning, but I think using it effectively can help. I'm in the middle of my attempt here, we'll see if it works as expected.</p>",
          "rawMarkdown": "&gt; I’ve been struggling to implement a supervised learning approach for weeks\n\nThat's one of the challenges here.  I am not sure one cannot get to 0.8 without supervised learning, but I think using it effectively can help. I'm in the middle of my attempt here, we'll see if it works as expected."
        },
        {
          "id": 360740,
          "postDate": "2018-07-23T06:18:26.950Z",
          "content": "<blockquote>\n  <p>20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. </p>\n</blockquote>\n\n<p>Sure, I'd be happy to double my running time to get your results!</p>",
          "rawMarkdown": "&gt; 20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. \n\nSure, I'd be happy to double my running time to get your results!"
        },
        {
          "id": 360916,
          "postDate": "2018-07-23T13:53:00.357Z",
          "content": "<p>I tried to estimate the maximum scrore I'll be able to get with my current clustering and an ideal track extending and refining.\nI took a local 0.705 solution achieved only by clustering, assumed I would somehow be able to collect all the score of every hit related to a particle if I  clustered at least 5 of this particle's hits (even if currently I don't score on this track)\nThe hypothetical score I got was about 0.825. </p>\n\n<p>From this result I understand, I wouldn't be able to improve my score above 0.8, unless I'll change my clustering strategy, probably start clustering tracks that start far from the origin.</p>",
          "rawMarkdown": "I tried to estimate the maximum scrore I'll be able to get with my current clustering and an ideal track extending and refining.\nI took a local 0.705 solution achieved only by clustering, assumed I would somehow be able to collect all the score of every hit related to a particle if I  clustered at least 5 of this particle's hits (even if currently I don't score on this track)\nThe hypothetical score I got was about 0.825. \n\nFrom this result I understand, I wouldn't be able to improve my score above 0.8, unless I'll change my clustering strategy, probably start clustering tracks that start far from the origin."
        },
        {
          "id": 361309,
          "postDate": "2018-07-24T08:09:45.160Z",
          "content": "<p>I've made good progress yesterday, with better features and some tuning.  I'm now getting to 0.70 local in about 30 minutes per core and per event with DBSCAN only, assigning hits to longest track, on the fly, using about 3,500 pairs.  Letting it run to plateau with about 40,000 pairs gets me to 0.744 in about 10 hours per core and per event.  I'm surprised that a simple clustering works that well actually.  However, I don't think it will be easy to improve it significantly,  and I'm now looking into supervised learning, and into better track merge/extension.</p>\n\n<p>edited: it is 30 min to get to 0.7, not 15 min.</p>",
          "rawMarkdown": "I've made good progress yesterday, with better features and some tuning.  I'm now getting to 0.70 local in about 30 minutes per core and per event with DBSCAN only, assigning hits to longest track, on the fly, using about 3,500 pairs.  Letting it run to plateau with about 40,000 pairs gets me to 0.744 in about 10 hours per core and per event.  I'm surprised that a simple clustering works that well actually.  However, I don't think it will be easy to improve it significantly,  and I'm now looking into supervised learning, and into better track merge/extension.\n\nedited: it is 30 min to get to 0.7, not 15 min.",
          "votes": 3
        },
        {
          "id": 361324,
          "postDate": "2018-07-24T08:34:18.540Z",
          "content": "<blockquote>\n  <p>I'm now getting to 0.70 local in about 15 minutes per core and per event with DBSCAN only</p>\n</blockquote>\n\n<p>Wow!!! It's hard to believe that. Congrats!</p>",
          "rawMarkdown": "&gt; I'm now getting to 0.70 local in about 15 minutes per core and per event with DBSCAN only\n\nWow!!! It's hard to believe that. Congrats!",
          "votes": 1
        },
        {
          "id": 361329,
          "postDate": "2018-07-24T08:39:06.823Z",
          "content": "<p>Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?</p>",
          "rawMarkdown": "Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?"
        },
        {
          "id": 361338,
          "postDate": "2018-07-24T09:02:23.840Z",
          "content": "<blockquote>\n  <p>It's hard to believe that. </p>\n</blockquote>\n\n<p>I find it hard to believe too ;)</p>\n\n<p>I'll make a sub to be sure, it will take few days.</p>\n\n<p>edited: yes, it was wrong, it takes 30 min to get to 0.7</p>",
          "rawMarkdown": "&gt; It's hard to believe that. \n\nI find it hard to believe too ;)\n\nI'll make a sub to be sure, it will take few days.\n\nedited: yes, it was wrong, it takes 30 min to get to 0.7"
        },
        {
          "id": 361339,
          "postDate": "2018-07-24T09:03:10.283Z",
          "content": "<blockquote>\n  <p>Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?</p>\n</blockquote>\n\n<p>What do you mean?  A report like outrunner's?</p>",
          "rawMarkdown": "&gt; Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?\n\nWhat do you mean?  A report like outrunner's?"
        },
        {
          "id": 361345,
          "postDate": "2018-07-24T09:18:25.070Z",
          "content": "<p>just two plots... (mostly to understand if you find easier to find tracks in the endcaps than in the barrel..)</p>",
          "rawMarkdown": "just two plots... (mostly to understand if you find easier to find tracks in the endcaps than in the barrel..)"
        },
        {
          "id": 361356,
          "postDate": "2018-07-24T09:42:12.530Z",
          "content": "<p>Here is a report similar to what <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60054\">outrunner shared</a> 2 weeks ago for event 1001.  This shows the very large gap between what I have and what he had.  Sure he uses clever track extension/merging, but I think his initial tracks are better than what I have.</p>\n\n<p>I have not done a theta distribution analysis but that's a good idea probably.</p>\n\n<pre><code>    vr range    GT weights  reconstruct     rate    loss\n0   0 - 0.01    0.218332    0.192069    0.879712    0.026263\n1   0.01 - 0.1  0.612320    0.528006    0.862304    0.084314\n2    0.1 - 1    0.011578    0.009683    0.836329    0.001895\n3     1 - 10    0.018677    0.009844    0.527069    0.008833\n4    10 -100    0.083535    0.009656    0.115595    0.073879\n5   100 - inf   0.055558    0.003833    0.068993    0.051725\n</code></pre>\n\n<p>​</p>\n\n<p>Local score for that event is 0.753</p>",
          "rawMarkdown": "Here is a report similar to what [outrunner shared][1] 2 weeks ago for event 1001.  This shows the very large gap between what I have and what he had.  Sure he uses clever track extension/merging, but I think his initial tracks are better than what I have.\n\nI have not done a theta distribution analysis but that's a good idea probably.\n\n     \tvr range \tGT weights \treconstruct \trate \tloss\n    0 \t0 - 0.01 \t0.218332 \t0.192069 \t0.879712 \t0.026263\n    1 \t0.01 - 0.1 \t0.612320 \t0.528006 \t0.862304 \t0.084314\n    2 \t 0.1 - 1   \t0.011578 \t0.009683 \t0.836329 \t0.001895\n    3 \t  1 - 10  \t0.018677 \t0.009844 \t0.527069 \t0.008833\n    4 \t 10 -100 \t0.083535 \t0.009656 \t0.115595 \t0.073879\n    5 \t100 - inf \t0.055558 \t0.003833 \t0.068993 \t0.051725\n\n​\n\nLocal score for that event is 0.753\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/60054"
        },
        {
          "id": 361358,
          "postDate": "2018-07-24T09:49:28.360Z",
          "content": "<blockquote>\n  <p>mostly to understand if you find easier to find tracks in the endcaps than in the barrel..</p>\n</blockquote>\n\n<p>I did look at rate per volume, best volumes are 7 and 9, worse are the 3 outer volumes.</p>",
          "rawMarkdown": "&gt; mostly to understand if you find easier to find tracks in the endcaps than in the barrel..\n\nI did look at rate per volume, best volumes are 7 and 9, worse are the 3 outer volumes."
        },
        {
          "id": 361419,
          "postDate": "2018-07-24T12:40:56.543Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 361424,
          "postDate": "2018-07-24T13:05:45.703Z",
          "content": "<p>I use numpy.unique for binning. Please look at one of my previous post in this discussion, I posted a full description of my python binning code (4 lines of code)</p>",
          "rawMarkdown": "I use numpy.unique for binning. Please look at one of my previous post in this discussion, I posted a full description of my python binning code (4 lines of code)",
          "votes": 1
        },
        {
          "id": 361461,
          "postDate": "2018-07-24T14:33:05.040Z",
          "content": "<p>@CPMP congratulation on your great achievement and news (The 0.744) and <strong>thanks</strong>, I was about to give up  and decide I can't go any further but now you gave me new hope.</p>",
          "rawMarkdown": "@CPMP congratulation on your great achievement and news (The 0.744) and **thanks**, I was about to give up  and decide I can't go any further but now you gave me new hope.\n",
          "votes": 1
        },
        {
          "id": 361734,
          "postDate": "2018-07-25T02:40:16.213Z",
          "content": "<p>I realized today that I made a stupid and was only z-shifting in 1 direction haha instant improvement in score smh</p>",
          "rawMarkdown": "I realized today that I made a stupid and was only z-shifting in 1 direction haha instant improvement in score smh",
          "votes": 1
        },
        {
          "id": 361819,
          "postDate": "2018-07-25T06:15:46.380Z",
          "content": "<p>@yuval,  thanks as well!  One can always make progress until the last minute.   Well, sort of 'last minute' given the time it takes to build a submission here ;)</p>",
          "rawMarkdown": "@yuval,  thanks as well!  One can always make progress until the last minute.   Well, sort of 'last minute' given the time it takes to build a submission here ;)"
        },
        {
          "id": 362346,
          "postDate": "2018-07-26T07:15:29.680Z",
          "content": "<p>I just submitted a 10k run (40k run underway), it takes about 2-2.5 hour per i7 core per event.  Local score is 0.7294, and LB is 0.7303.  This is just DBSCAN with assiging hits to longest track on the fly.  No postprocessing at all.</p>",
          "rawMarkdown": "I just submitted a 10k run (40k run underway), it takes about 2-2.5 hour per i7 core per event.  Local score is 0.7294, and LB is 0.7303.  This is just DBSCAN with assiging hits to longest track on the fly.  No postprocessing at all.",
          "votes": 1
        },
        {
          "id": 362348,
          "postDate": "2018-07-26T07:20:49.550Z",
          "content": "<p>@CPMP, may I ask, when using different z-shifting, your merging method stays the same? I mean what scheme do you use between the below 2:</p>\n\n<ul>\n<li><p>Find different clusters of different z-shifting, then merging those clusters. (each cluster is performed with loops of changing angle).</p></li>\n<li><p>Find one cluster of a specific z, change z shifting, after each new angle loop of the new z-shifting, immediate merge the new cluster of that loop to the existing cluster. Speaking differently, this scheme do not know about z-shifting at all (use the same merge function for all loop cluster).</p></li>\n</ul>\n\n<p>When using the latter scheme, I found the result sometimes worsen.</p>",
          "rawMarkdown": "@CPMP, may I ask, when using different z-shifting, your merging method stays the same? I mean what scheme do you use between the below 2:\n\n+ Find different clusters of different z-shifting, then merging those clusters. (each cluster is performed with loops of changing angle).\n\n+ Find one cluster of a specific z, change z shifting, after each new angle loop of the new z-shifting, immediate merge the new cluster of that loop to the existing cluster. Speaking differently, this scheme do not know about z-shifting at all (use the same merge function for all loop cluster).\n\nWhen using the latter scheme, I found the result sometimes worsen."
        },
        {
          "id": 362371,
          "postDate": "2018-07-26T08:10:51.880Z",
          "content": "<p>I'm using the latter.  Several people said the former is better, but it uses way too much memory.  </p>",
          "rawMarkdown": "I'm using the latter.  Several people said the former is better, but it uses way too much memory.  ",
          "votes": 1
        },
        {
          "id": 363514,
          "postDate": "2018-07-29T11:03:10.590Z",
          "content": "<blockquote>\n  <p>I'll make a sub to be sure, it will take few days.</p>\n</blockquote>\n\n<p>I just submitted my DBSCAN run with 40k pairs.  Local score was 0.744, LB score is 0.7449, no post processing, just DBSCAN.</p>\n\n<p>Working now on track extension and other ideas.  </p>",
          "rawMarkdown": "&gt; I'll make a sub to be sure, it will take few days.\n\nI just submitted my DBSCAN run with 40k pairs.  Local score was 0.744, LB score is 0.7449, no post processing, just DBSCAN.\n\nWorking now on track extension and other ideas.  ",
          "votes": 2
        },
        {
          "id": 363713,
          "postDate": "2018-07-29T23:44:21.147Z",
          "content": "<p>Kudos @CPMP, the 2 numbers are really close. This challenge is very interesting but it takes forever to run and create a submission. Make a change, run it for a day on my local machine and submit is all I am doing at the moment. I do not use my cloud service for this contest.</p>\n\n<p>Good luck to us all and happy Kaggling!!!</p>",
          "rawMarkdown": "Kudos @CPMP, the 2 numbers are really close. This challenge is very interesting but it takes forever to run and create a submission. Make a change, run it for a day on my local machine and submit is all I am doing at the moment. I do not use my cloud service for this contest.\n\nGood luck to us all and happy Kaggling!!!"
        },
        {
          "id": 367414,
          "postDate": "2018-08-07T17:41:46.033Z",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>Someone working on track reconstruction for high energy physics</p>\n</blockquote>\n\n<p>I'm not a specialist in this sense, except to the degree I have become one in this competition (not really).\nThe nickname <a href=\"/icecuber\">@icecuber</a> could suggest that he or she is a specialist, but I'm just guessing. It will be interesting to read the names in the copyright notices that will hopefully become visible. :-)</p>\n\n<p>I'm a physics enthusiast privately, but as you correctly stated, not much advanced physics/math is needed to compete in this competition.</p>",
          "rawMarkdown": "@CPMP:\n\n&gt; Someone working on track reconstruction for high energy physics\n\nI'm not a specialist in this sense, except to the degree I have become one in this competition (not really).\nThe nickname @icecuber could suggest that he or she is a specialist, but I'm just guessing. It will be interesting to read the names in the copyright notices that will hopefully become visible. :-)\n\nI'm a physics enthusiast privately, but as you correctly stated, not much advanced physics/math is needed to compete in this competition.",
          "votes": 1
        },
        {
          "id": 367422,
          "postDate": "2018-08-07T18:08:49.750Z",
          "content": "<p>Edwin, fair enough!  Your score is a motivation for me ;)</p>",
          "rawMarkdown": "Edwin, fair enough!  Your score is a motivation for me ;)",
          "votes": 1
        },
        {
          "id": 369094,
          "postDate": "2018-08-11T23:20:27.067Z",
          "content": "<p>Thanks @Kha A. Vo. \nCould you tell me where I can find the public kernel mentioned in your post above? I'd like to try both your transformation as well as this kernel.\n\"The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0\"</p>",
          "rawMarkdown": "Thanks @Kha A. Vo. \nCould you tell me where I can find the public kernel mentioned in your post above? I'd like to try both your transformation as well as this kernel.\n\"The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0\""
        },
        {
          "id": 369406,
          "postDate": "2018-08-13T01:20:07.323Z",
          "content": "<p>@HaiJiang: You can go and look at the \"Kernel\" section of this competition. I wrote a simple kernel which can produce 0.53 on LB. But that kernel didn't use the good second feature.</p>",
          "rawMarkdown": "@HaiJiang: You can go and look at the \"Kernel\" section of this competition. I wrote a simple kernel which can produce 0.53 on LB. But that kernel didn't use the good second feature."
        },
        {
          "id": 369897,
          "postDate": "2018-08-14T00:12:06.510Z",
          "content": "<p>Many thanks! @Kha A. Vo. I have learned a lot from your kernel. ; )</p>",
          "rawMarkdown": "Many thanks! @Kha A. Vo. I have learned a lot from your kernel. ; )",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 356653,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-07-14T06:06:52.160000",
      "content": "<p>If I remember correcty what I did, you can get significantly above 0.6 without z shifting, using DBSCAN, helix unrolling, and Heng's track extension code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 356894,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-14T18:11:32.420000",
          "content": "<p>z shifting on top can get you above 0.7, but I doubt this can get above 0.8.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356906,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-07-14T18:35:12.667000",
          "content": "<p>you need supervised learning to get beyond 0.80</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356916,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-14T18:52:53.200000",
          "content": "<p>I know, working on it since I'm back to this competition ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 356925,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-14T19:13:36.587000",
          "content": "<p>What I just submitted was a reaction to Grzegorz saying we cannot get over 0.7 with merging based on track length only.  I just checked if this was true or not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 357027,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-15T05:42:07.463000",
          "content": "<p>I can get beyond 0.75 selecting and extending tracks on length only.\nAnd the only other filtering I add is not allowing two hits on the exact same sensor.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 359036,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-19T11:12:22.957000",
          "content": "<p>@yuval it seems your clustering approach is way more accurate than dbscan if you only use 2-3 features without heavy track fitting. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359053,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-19T11:57:38.170000",
          "content": "<p>Actually I don't really do clustering, I use  bins (similar to the bins used in the Hough transform but implemented very efficiently). What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7.\nPlatu - can't get further with a clustering technique (neither bins nor dbscan) and need to use some track extending technique.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 359073,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-19T12:51:55.820000",
          "content": "<p>Interesting, I'm now at 0.7 without track extension, I guess you just confirm track extension should be my next focus.  And I'm now using 2 features as well if cos and sin counts for 1.  Maybe the same as yours actually.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359139,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2018-07-19T14:18:54.197000",
          "content": "<p>0.7 using only 2 features and no track extension? Impressive! congrats</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 359204,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-19T16:39:12.737000",
          "content": "<p>@Giba, thanks, but I am sure the guys above 0.8 have way better tracks than me before they extend them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 359354,
          "author_name": "Poseidon",
          "author_url": "",
          "post_date": "2018-07-20T00:06:53.293000",
          "content": "<p>with dbscan? that's still impresive</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359365,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-07-20T01:08:23.863000",
          "content": "<p>How is z shifting implemented? Is it somewhat like the for loop for helix unrolling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359368,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-20T01:36:18.970000",
          "content": "<p>@yuval r that is very impressive and congratulations. I am currently stuck below 0.6, would you care to elaborate on the following?</p>\n\n<p>&gt;  ...try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7 ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359383,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-20T02:47:40.663000",
          "content": "<p>The idea it to unroll the helix correctly. As I do not know the radius of the Helix (and the direction = the Q of the particle) I try different random values for the radius  (actually not for the radius but for 1/Radius).\nAlso the main collision area is +-5.5 mm around the origin in the Z axis =&gt; I try different random values for z0\nBecause I use binning and not dbscan, I need very accurate initial unrolling =&gt; I try a lot of different values.</p>\n\n<p>The way to do a very fast binning in python is using numpy.unique:</p>\n\n<p>un,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)</p>\n\n<p>hits['track_id']=inv</p>\n\n<p>hits['track_length']=count[inv]</p>\n\n<p>where hits['cat'] - is the binning variable.\nIn my main loop I do as little as possible - just adjust my features according to the random radius and z0, do the above binning and for every hit decide if to use this track_id according to track length.\nthis main loop takes less the 1min/1500 random values on my laptop = &gt;3 min to get 0.6 and about 40 to get to the plateau (get very little progress even after doubling the amount of values) </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 359505,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-20T08:22:55.960000",
          "content": "<p><a href=\"/martin\">@martin</a>, yes, using dbscan.  My approach is very similar to what @yuval describes, including what he says about his features, but I'm using dbscan instead of binning.  It is taking way more time than what he does, at least 10x more time.  I plan to see if I can switch to binning, but thats not my priority. Priority is to improve score first, running time second ;)</p>\n\n<p>@btKaggle , we discussed z shifting a while back in the forum.  The idea is simple.  If you construct features for dbscan, you construct features as a function of hit coordinates, right?  So if your features are f(x, y, z), then shifting z means you compute features as f(x, y, z - z0), where z0 is the new origin for your tracks.  You then need to find ways to combine the tracks coming from varying z0.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359511,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-20T08:39:07.420000",
          "content": "<p>Thanks @yuval r, that is very clear and since it is faster than dbscan, I will try binning and see if I can improve my model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359774,
          "author_name": "Poseidon",
          "author_url": "",
          "post_date": "2018-07-20T18:34:19.720000",
          "content": "<p>@yuval r what do you mean by \"binning variable\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359827,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-20T21:36:41.057000",
          "content": "<p>@Martin I am doing a one dimensional binning (for efficiency). Let's say we have two features f1 and f2, We need to define F=f1+K*f2. Where K&gt;max(f1). Now we can do binning on F which I call the binning variable.\nAnother thing you must do is make F an integer (every bin is a different integer) to do this you can multiply by a large number and use .astype('int'). The large number is very important because it actually determine the size of the bins.</p>\n\n<p>To put it all together:</p>\n\n<p>f1,f2 are the features in the range [-1,1]</p>\n\n<p>L1,L2 are large integers</p>\n\n<p>K - an integer where K&gt;L1</p>\n\n<p>We define:</p>\n\n<p>F= (L1*f1). astype ('int')+K*(L2*f2).astype('int')</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 359889,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-07-21T01:52:39.363000",
          "content": "<p>Thanks for your comment, @yuval6769. Your approach using np.unique is so smart! May I ask you a small question? Do you use Hough transform by varying 1/r0 and calculate theta by theta = phi - arccos(r/2r0)? If so, your large integer L1, L2 and K be the same across different 1/r0 or you use adaptive L1, L2, K with respect to r0?  I am using that approach but face some difficulties by selecting the range for 1/r0. When doing that way, my 2 features will by z/r and theta, which can be plugged into your suggested equation for 1-D binning. I am suspecting the feasibility of this approach, because we need to be careful of the range of arccos and arctan2 functions, because they are used in one equation together ( theta = phi - arccos(r/2r0) = arctan2(y/x) - arccos(r/2r0) ). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359989,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-21T07:53:33.683000",
          "content": "<p>@Kha A. Vo</p>\n\n<p>Although my approach is close to yours, I’m not using Hough transform but clustering (or binning) because Hough transform isn’t a suitable solution for this challenge. Let me start by explaining this statement and later I’ll address your issues.</p>\n\n<p>Why Hough transform isn’t suitable:</p>\n\n<p>•   Hough transform does not assume the tracks starts at the origin which leaves you with too many free variables. The Hough transform kernel decrease this number by assuming that in the Z axis it does start from the origin, also, it does not scan for different possible positions for the center of the helix in the XY plain.</p>\n\n<p>•   If you don’t assume the tracks starts at the origin, you’ll get a very large number of wrong tracks. The sensors are arranged in circular structures centered at the origin, hence you have an endless number of optional Helixes centered at the origin.</p>\n\n<p><strong>One remark about the origin assumption – about 18% of the tracks don’t start at the origin, 80% of these tracks couldn’t be found while using the origin assumption (I believe this is the reason the leaders are stuck around 0.82 - in some discussions it seems <a href=\"/outrunner\">@outrunner</a> solved this issue and we will soon se a 0.9 from him)</strong></p>\n\n<p>Now Let’s address your issues.</p>\n\n<p>Arccos(r/2r0) is a good and accurate choice, I believe it is better then: <code>theta = Phi + k*z</code>, you find in some kernals. The problem with arccos is that arccos(t) isn’t defined for |t|&gt;1. And for some choices of 1/r0 you will get r/2r0&gt;1 for some hits. You need to find a way to overcome this issue.</p>\n\n<p>On the other hand, z/r is not a good feature, as it has two major lacks. First it is unbounded and unevenly spread, and 2nd it does not take into account the curve of the track between (0,0) and (x,y).\nThe first issue can be easily solved using arctan() and for the 2nd issue you’ll need a little bit of geometry.  </p>\n\n<p>As for your question about changing L1, L2, K, I don’t change them for different values of 1/r0. I don’t change K at all, I do use more than one value for L1, L2 – which is equivalent to using different bin sizes but I do it regardless of the value of 1/r0.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 360025,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-07-21T09:50:39.083000",
          "content": "<p>Thanks @yuval r.  I read your comment and believe you may miss some points I made. First, you said \"I’m not using Hough transform but clustering (or binning)\". Indeed, I also used clustering, but by Hough features. Now I would like to shift it to the binning to avoid DBSCAN. Let me explain it further.</p>\n\n<p>1) The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0. This assumes that, with different i, all hits which belong to a true track, will have near Euclidean distance using features [sin(phi), cos(phi), ...some other features...]. With i = 0, phi is kept untouched, so straight tracks will be detected. With i large, phi is distorted, and hits lying in the curve tracks will have the same distorted phi.</p>\n\n<p>However, this is not extremely accurate, because phi is not exactly a linear function of r, or quadratic... Indeed, Hough transform did provide an exact function relating phi and r, that is r/r0 = 2cos(phi - theta).</p>\n\n<p>2) That equation is exact, and ALREADY assumes that the track starts at origin (0,0,z) in the XY plane!  It is because given any true track starts from origin (0,0,z), the track will draw a curve in the XY plane starting from (0,0).  This curve is an arc of a circle, whose diameter is 2r0 and the origin of this circle will have a specific and unique theta.  </p>\n\n<p>3) So, using the Hough equation and doing clustering with it, is surely a better method than using handcrafted equation. Now returning to point number 1) I made above. Varying i and calculating the new phi as mentioned in 1), will be equivalent to varying theta and calculating 1/r0 for each hit, by the Hough equation: 1/r0 = 2cos(phi-theta)/r. Then, using 1/r0 as a feature, all hits lying on the same true tracks with the provided theta, will have the same 1/r0. This will be clustered by DBSCAN. Indeed, the </p>\n\n<p>4) However, the approach 3) is good, but not smart. As you said, varying and sampling 1/r0 is better because we can use a Gaussian distribution for 1/r0 centered at 0. With varying 1/r0, now we switch the role of 1/r0 and theta. Now 1/r0 to be varied, and theta is to be calculated as a feature. With DBSCAN, this would be impossible to scan all possible values for 1/r0. Therefore, I want to shift to binning. Indeed, the binning with Hough features was also done by a public kernel and produce 0.1x on LB.</p>\n\n<p>5) With binning, however using K, L1, L2 as you said is very sensitive. I tried it, with just one value of 1/r0 = 0, and get the score of 0.002. This means that it is feasible to run for different 1/r0. But there are wrong tracks as you said.  Moreover, the hits binning results are very sensitive to large K, L1, L2(for example, hits with binning variables 23431, 23429, 23464 should be clustered together because the values are nearly equal. Now I need to find K, L1, L2 so that all of them must equal the unique integer (that is hard!!).</p>\n\n<p>Anyway, thanks for your answer. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 360027,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T09:51:26.553000",
          "content": "<p>Well well, if anyone is not above 0.7 by end of competition with this then something is wrong ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360050,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-21T11:49:03.737000",
          "content": "<p>@Kha I agree with 1) - 4) but not with 5). binning is not sensitive to the values of K, L1, L2. You should be careful when choosing K because if K is too small the different feature would blend together, but the values of L1, L2 can be changed by 10% without getting almost any difference in the score.</p>\n\n<p>There is one sensitivity in binning which could be easily explain by the following example: lets say we have only one feature: f1 and for a certain true track for some hits we get f1=1e-5 and others f1=-1e-5, although the values are very close,  these hits will never fall in the same bin if you do binning by <code>bin=(L1*f1).astype(int)</code>.  This issue has a few elegant solutions.</p>\n\n<p>Another thing to understand it that binning is much more sensitive to the values of 1/r0 and z0 because you need all hits to fall in the same bin, this is why I try so many different values.</p>\n\n<p>One last comment - the distribution of 1/r0 is definitely not Gaussian - 1/r0=0 means the particle is not rotating which means it has no charge - I think uncharged particles are invisible for the detectors. I don't know what is the real distribution, but it looks like gamma distribution of type 2, mirrored to the left because of the particles charge. <strong>The nice thing it you can use Gaussian distribution with no score penalty, and a small time penalty =&gt; I use a Gaussian distribution</strong>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 360052,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T11:53:13",
          "content": "<p>@Kha, you seem to miss that hough transform maps each hit to a curve, not just a point in another feature space.  You use the equation of a circle using polar coordinates, which is indeed an ingredient of hough transform, but not all of it.  See <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60824#355799\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60824#355799</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360054,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-21T11:57:35.373000",
          "content": "<p>@CPMP you are correct, but I believe it's time this competition will start to heat up a little as it seems to be stuck. More Kagglers  above 0.7 means more ideas and more knowledge and this is what it is all about. :)</p>\n\n<p>(But to get to 0.7 ,even with all the hints around  one really needs to work. If I was  one of the organizers I would have started with an example kernel which get 0.7 because here is where all the fun begins).   </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 360058,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T12:18:10.263000",
          "content": "<p>I don't think competition is stuck by any mean, at least if I look at my local score.  But thing is people aren't submitting much because you can use local score quite effectively, and because it is computationally intensive. They may need the compute resource to try new ideas instead of creating a submission.</p>\n\n<p>I am not against sharing, and I do my share of it, but I am against people getting medals just be replicating what has been shared.  I'm afraid you are offering medals on a plate here.  </p>\n\n<p>This said, this is my opinion, and just my opinion.  You can safely ignore it.  And there are probably a bunch of people who liked your sharing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360059,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-07-21T12:21:01.137000",
          "content": "<p>Thanks @yuval. I believe my statement stating that choosing K, L1, L2 is sensitive, is because I haven't tried enough, and also I may use the inappropriate features. I will try the old features of DBSCAN clustering with sin, cos, and do it with binning, to see its efficiency. That is totally worth it. Before your share about this, I never believe one could do 10K iterations for merging tracks. </p>\n\n<p>@CPMP. Thanks. Indeed I read your thread there all. In there, <a href=\"/trian2018\">@trian2018</a> said </p>\n\n<p>\"  <em>At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).</em> \"</p>\n\n<p>That is, to my knowledge, equivalent to the binning or clustering method of @Grzegorz. I agree to your statement that \"hough transform maps each hit to a curve, not just a point in another feature space\". However we did the clustering by detecting the nearby transformed hits at each iteration (the intersection point).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360060,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T12:21:18.050000",
          "content": "<p>&gt; I would have started with an example kernel which get 0.7 *</p>\n\n<p>This assumes they could produce one at the start of the competition, which I doubt.  This data is way more dense and complex than what the research community has been working with so far.  The fact that it took 2 month for a specialist like Edwin Steiner to move to 0.8 shows it is hard, even for specialists.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360062,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-07-21T12:27:50.560000",
          "content": "<p>@CPMP What do you mean by \"specialist\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360065,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T12:29:41.753000",
          "content": "<p>Someone working on track reconstruction for high energy physics</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360068,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-21T12:35:54.507000",
          "content": "<p>@CPMP  - Don't worry about it. We didn't use any of these approaches, simply DBSCAN with the chemist's features <code>sin, cos, z1, z2</code> + z-shifting + handcrafted merging code to pass the 0.7 plateau. Right now we're working on track extension and exploring supervised learning approaches. </p>\n\n<p>A couple months ago, when the chemist wanted to share a kernel with z-shifting, we kind of agreed that we wouldn't share critical approaches in the final month and we still keep that consensus.  I think it's more fair to others who have worked really hard in this competition.</p>\n\n<p><strong>UPDATE on July 31st</strong>\nI finally figured out the secret sauce after having read this, pretty clever. I used to have trouble calculating the 1/r0 before and only ran 100 iterations of 1/r0 and the score was pretty bad and gave up on it, so I handcrafted 2 quadratic functions to approximate the new unrolling angles. Thanks to <strong>@yuval, @Kha and @CPMP</strong> sharing your thoughts here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 360074,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-21T13:30:07.597000",
          "content": "<p>@Nicole, I don't worry ;)  Other than that I agree with you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360522,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-22T16:37:12.083000",
          "content": "<blockquote>\n  <p>What I do is unroll the helix very accurately and try a lot of helix radiuses and Z0 (origin of the partical) - about 4500 to get 0.6 and 60000 to the platu at about 0.7. </p>\n</blockquote>\n\n<p>@yuval\nDo you mean that the total number of radiuses and Z0 is 4500? (i.e. the sum, not every object).</p>\n\n<p>I used only about 200 radiuses, and five z0. It takes 20 minutes for an event, and I consider this slow. Maybe I should move further for better score. I'm also trying ML, but no success so far.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360535,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-22T16:56:38.063000",
          "content": "<p>I mean 4500 combinations. my combinations are random pairs (I use Gaussian distribution for each).\nI understand you are trying 200*5=1000 pairs. \nAs I wrote I am using binning and not dbscan =&gt; it takes only 3 min for an event to get about 0.6, It takes about 40 min to check enough pairs to get to about 0.7. \nI don't think any ML is needed until you get to 0.7 (as @CPMP also confirmed) going beyond 0.7 is a different story.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 360546,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-22T17:13:46.490000",
          "content": "<p>@yuval Thanks, now it's clear to me. Yes, I used &lt; 1000 pairs. I will try more.\nYour approach is quite fast! But I'm afraid there's no time to change to your variant (I'm using DBSCAN).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360561,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-22T17:45:25.953000",
          "content": "<p>@Serguey, using 20,000 pairs with DBSCAN moves me over 0.71 locally (my LB score is usually higher than the local score) before any track extension.  But running times are way longer than @yuval's, about 20 minutes to get to 0.6 on a single core.  This is to say that there is hope with DBSCAN, but fast binning as @yuval uses is tempting.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 360615,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-22T20:58:36.483000",
          "content": "<p>@CPMP Thank you for info! What's about time for 20,000 pairs? Outrunner wrote about 10-20 hours for an event. This shocked me! :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360619,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-22T21:06:32.490000",
          "content": "<p>About 9-10 hours or so on a single core.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360621,
          "author_name": "Poseidon",
          "author_url": "",
          "post_date": "2018-07-22T21:14:33.053000",
          "content": "<p>I’m lucky, I have access to some wicked machinery, but only have an intermediate level of knowledge, so I have a .65 local score but I’ve been struggling to implement a supervised learning approach for weeks 😭</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360689,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-07-23T02:49:58.907000",
          "content": "<p>20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. The bad news is the upcoming update data. :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 360739,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-23T06:16:52.087000",
          "content": "<blockquote>\n  <p>I’ve been struggling to implement a supervised learning approach for weeks</p>\n</blockquote>\n\n<p>That's one of the challenges here.  I am not sure one cannot get to 0.8 without supervised learning, but I think using it effectively can help. I'm in the middle of my attempt here, we'll see if it works as expected.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360740,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-23T06:18:26.950000",
          "content": "<blockquote>\n  <p>20 hours is a reasonable upper bound for a core i7, and 10 hours gives more flexible. </p>\n</blockquote>\n\n<p>Sure, I'd be happy to double my running time to get your results!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360916,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-23T13:53:00.357000",
          "content": "<p>I tried to estimate the maximum scrore I'll be able to get with my current clustering and an ideal track extending and refining.\nI took a local 0.705 solution achieved only by clustering, assumed I would somehow be able to collect all the score of every hit related to a particle if I  clustered at least 5 of this particle's hits (even if currently I don't score on this track)\nThe hypothetical score I got was about 0.825. </p>\n\n<p>From this result I understand, I wouldn't be able to improve my score above 0.8, unless I'll change my clustering strategy, probably start clustering tracks that start far from the origin.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361309,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-24T08:09:45.160000",
          "content": "<p>I've made good progress yesterday, with better features and some tuning.  I'm now getting to 0.70 local in about 30 minutes per core and per event with DBSCAN only, assigning hits to longest track, on the fly, using about 3,500 pairs.  Letting it run to plateau with about 40,000 pairs gets me to 0.744 in about 10 hours per core and per event.  I'm surprised that a simple clustering works that well actually.  However, I don't think it will be easy to improve it significantly,  and I'm now looking into supervised learning, and into better track merge/extension.</p>\n\n<p>edited: it is 30 min to get to 0.7, not 15 min.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 361324,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-24T08:34:18.540000",
          "content": "<blockquote>\n  <p>I'm now getting to 0.70 local in about 15 minutes per core and per event with DBSCAN only</p>\n</blockquote>\n\n<p>Wow!!! It's hard to believe that. Congrats!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361329,
          "author_name": "VinInn",
          "author_url": "",
          "post_date": "2018-07-24T08:39:06.823000",
          "content": "<p>Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361338,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-24T09:02:23.840000",
          "content": "<blockquote>\n  <p>It's hard to believe that. </p>\n</blockquote>\n\n<p>I find it hard to believe too ;)</p>\n\n<p>I'll make a sub to be sure, it will take few days.</p>\n\n<p>edited: yes, it was wrong, it takes 30 min to get to 0.7</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361339,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-24T09:03:10.283000",
          "content": "<blockquote>\n  <p>Do you have a theta (eta would be better) track distribution for the 15 minutes and 10 hours cases?</p>\n</blockquote>\n\n<p>What do you mean?  A report like outrunner's?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361345,
          "author_name": "VinInn",
          "author_url": "",
          "post_date": "2018-07-24T09:18:25.070000",
          "content": "<p>just two plots... (mostly to understand if you find easier to find tracks in the endcaps than in the barrel..)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361356,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-24T09:42:12.530000",
          "content": "<p>Here is a report similar to what <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60054\">outrunner shared</a> 2 weeks ago for event 1001.  This shows the very large gap between what I have and what he had.  Sure he uses clever track extension/merging, but I think his initial tracks are better than what I have.</p>\n\n<p>I have not done a theta distribution analysis but that's a good idea probably.</p>\n\n<pre><code>    vr range    GT weights  reconstruct     rate    loss\n0   0 - 0.01    0.218332    0.192069    0.879712    0.026263\n1   0.01 - 0.1  0.612320    0.528006    0.862304    0.084314\n2    0.1 - 1    0.011578    0.009683    0.836329    0.001895\n3     1 - 10    0.018677    0.009844    0.527069    0.008833\n4    10 -100    0.083535    0.009656    0.115595    0.073879\n5   100 - inf   0.055558    0.003833    0.068993    0.051725\n</code></pre>\n\n<p>​</p>\n\n<p>Local score for that event is 0.753</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361358,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-24T09:49:28.360000",
          "content": "<blockquote>\n  <p>mostly to understand if you find easier to find tracks in the endcaps than in the barrel..</p>\n</blockquote>\n\n<p>I did look at rate per volume, best volumes are 7 and 9, worse are the 3 outer volumes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361419,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-24T12:40:56.543000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361424,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-24T13:05:45.703000",
          "content": "<p>I use numpy.unique for binning. Please look at one of my previous post in this discussion, I posted a full description of my python binning code (4 lines of code)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361461,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-24T14:33:05.040000",
          "content": "<p>@CPMP congratulation on your great achievement and news (The 0.744) and <strong>thanks</strong>, I was about to give up  and decide I can't go any further but now you gave me new hope.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361734,
          "author_name": "Poseidon",
          "author_url": "",
          "post_date": "2018-07-25T02:40:16.213000",
          "content": "<p>I realized today that I made a stupid and was only z-shifting in 1 direction haha instant improvement in score smh</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361819,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-25T06:15:46.380000",
          "content": "<p>@yuval,  thanks as well!  One can always make progress until the last minute.   Well, sort of 'last minute' given the time it takes to build a submission here ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362346,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-26T07:15:29.680000",
          "content": "<p>I just submitted a 10k run (40k run underway), it takes about 2-2.5 hour per i7 core per event.  Local score is 0.7294, and LB is 0.7303.  This is just DBSCAN with assiging hits to longest track on the fly.  No postprocessing at all.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 362348,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-07-26T07:20:49.550000",
          "content": "<p>@CPMP, may I ask, when using different z-shifting, your merging method stays the same? I mean what scheme do you use between the below 2:</p>\n\n<ul>\n<li><p>Find different clusters of different z-shifting, then merging those clusters. (each cluster is performed with loops of changing angle).</p></li>\n<li><p>Find one cluster of a specific z, change z shifting, after each new angle loop of the new z-shifting, immediate merge the new cluster of that loop to the existing cluster. Speaking differently, this scheme do not know about z-shifting at all (use the same merge function for all loop cluster).</p></li>\n</ul>\n\n<p>When using the latter scheme, I found the result sometimes worsen.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362371,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-26T08:10:51.880000",
          "content": "<p>I'm using the latter.  Several people said the former is better, but it uses way too much memory.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 363514,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-29T11:03:10.590000",
          "content": "<blockquote>\n  <p>I'll make a sub to be sure, it will take few days.</p>\n</blockquote>\n\n<p>I just submitted my DBSCAN run with 40k pairs.  Local score was 0.744, LB score is 0.7449, no post processing, just DBSCAN.</p>\n\n<p>Working now on track extension and other ideas.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 363713,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-29T23:44:21.147000",
          "content": "<p>Kudos @CPMP, the 2 numbers are really close. This challenge is very interesting but it takes forever to run and create a submission. Make a change, run it for a day on my local machine and submit is all I am doing at the moment. I do not use my cloud service for this contest.</p>\n\n<p>Good luck to us all and happy Kaggling!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367414,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-07T17:41:46.033000",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>Someone working on track reconstruction for high energy physics</p>\n</blockquote>\n\n<p>I'm not a specialist in this sense, except to the degree I have become one in this competition (not really).\nThe nickname <a href=\"/icecuber\">@icecuber</a> could suggest that he or she is a specialist, but I'm just guessing. It will be interesting to read the names in the copyright notices that will hopefully become visible. :-)</p>\n\n<p>I'm a physics enthusiast privately, but as you correctly stated, not much advanced physics/math is needed to compete in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367422,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-07T18:08:49.750000",
          "content": "<p>Edwin, fair enough!  Your score is a motivation for me ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 369094,
          "author_name": "HaiJiang",
          "author_url": "",
          "post_date": "2018-08-11T23:20:27.067000",
          "content": "<p>Thanks @Kha A. Vo. \nCould you tell me where I can find the public kernel mentioned in your post above? I'd like to try both your transformation as well as this kernel.\n\"The public kernel (which produced 0.49) works by using DBSCAN. In the main loop, they vary phi by phi = phi_0 + i*f(r), where f is a quadratic function of r, and i is incremented a little by each loop from 0\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 369406,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-08-13T01:20:07.323000",
          "content": "<p>@HaiJiang: You can go and look at the \"Kernel\" section of this competition. I wrote a simple kernel which can produce 0.53 on LB. But that kernel didn't use the good second feature.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 369897,
          "author_name": "HaiJiang",
          "author_url": "",
          "post_date": "2018-08-14T00:12:06.510000",
          "content": "<p>Many thanks! @Kha A. Vo. I have learned a lot from your kernel. ; )</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "356609": "Hi,\n\nI was trying to find good features for DBSCAN and was wondering at what score is it a good idea to stop looking for better features and optimizing weights on your features and instead focus on track extension, z shifting and finding better ways to merge tracks?\n\nThanks",
    "356653": "If I remember correcty what I did, you can get significantly above 0.6 without z shifting, using DBSCAN, helix unrolling, and Heng's track extension code."
  }
}