{
  "id": 58194,
  "title": "track-extension code : +0.04LB",
  "url": "/competitions/trackml-particle-identification/discussion/58194",
  "author_name": "hengck23",
  "post_date": "2018-06-04T08:41:23.365000",
  "votes": 35,
  "comment_count": 53,
  "views": 0,
  "content": "<p>this is example of track extension. It should gives improvement of LB +0.04 What you should do:</p>\n\n<ol>\n<li><p>use this your extension submission track_id. you can run it over a few times (submission and hits are the dataframe of csv files):</p>\n\n<pre><code>for i in range(8): \n       submission = extend(submission, hits)\n</code></pre></li>\n<li><p>tune the parameters to get better results</p></li>\n<li><p>make some visualizations. It is easy to debug, fault find and further improve results.</p></li>\n</ol>\n\n<hr>\n\n<p>if you improve it by multi-threading, code efficiency, etc ... please share your improved version back here. thanks!</p>\n\n<p>here is the code</p>\n\n<pre><code>def extend(submission,hits):\n\n    df = submission.merge(hits,  on=['hit_id'], how='left')\n    df = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n    df = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n    df = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n    ... see extension.py as attached ... \n</code></pre>\n\n<p>example results</p>\n\n<p>pink: extended tracks</p>\n\n<p>red: submission tracks</p>\n\n<p>black: truth tracks</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338013/9565/after.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": 338013,
      "postDate": "2018-06-04T08:41:23.367Z",
      "content": "<p>this is example of track extension. It should gives improvement of LB +0.04 What you should do:</p>\n\n<ol>\n<li><p>use this your extension submission track_id. you can run it over a few times (submission and hits are the dataframe of csv files):</p>\n\n<pre><code>for i in range(8): \n       submission = extend(submission, hits)\n</code></pre></li>\n<li><p>tune the parameters to get better results</p></li>\n<li><p>make some visualizations. It is easy to debug, fault find and further improve results.</p></li>\n</ol>\n\n<hr>\n\n<p>if you improve it by multi-threading, code efficiency, etc ... please share your improved version back here. thanks!</p>\n\n<p>here is the code</p>\n\n<pre><code>def extend(submission,hits):\n\n    df = submission.merge(hits,  on=['hit_id'], how='left')\n    df = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n    df = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n    df = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n    ... see extension.py as attached ... \n</code></pre>\n\n<p>example results</p>\n\n<p>pink: extended tracks</p>\n\n<p>red: submission tracks</p>\n\n<p>black: truth tracks</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338013/9565/after.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "this is example of track extension. It should gives improvement of LB +0.04 What you should do:\n\n1. use this your extension submission track_id. you can run it over a few times (submission and hits are the dataframe of csv files):\n\n        for i in range(8): \n               submission = extend(submission, hits)\n\n2. tune the parameters to get better results\n\n3. make some visualizations. It is easy to debug, fault find and further improve results.\n\n---\nif you improve it by multi-threading, code efficiency, etc ... please share your improved version back here. thanks!\n\nhere is the code\n\n    def extend(submission,hits):\n\n\t\tdf = submission.merge(hits,  on=['hit_id'], how='left')\n\t\tdf = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n\t\tdf = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n\t\tdf = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n\t\t... see extension.py as attached ... \n\n\nexample results\n\npink: extended tracks\n\nred: submission tracks\n\nblack: truth tracks\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/338013/9565/after.png",
      "votes": 34
    },
    {
      "id": 340763,
      "postDate": "2018-06-10T08:41:46.270Z",
      "content": "<p>my updated code</p>\n\n<pre><code>def extend(submission,hits,limit=0.04, num_neighbours=18):\n\ndf = submission.merge(hits,  on=['hit_id'], how='left')\ndf = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\ndf = df.assign(r = np.sqrt( df.x**2 + df.y**2))\ndf = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\nfor angle in range(-90,90,1):\n\n    print ('\\r %f'%angle, end='',flush=True)\n    #df1 = df.loc[(df.arctan2&gt;(angle-0.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+0.5)/180*np.pi)]\n    df1 = df.loc[(df.arctan2&gt;(angle-1.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+1.5)/180*np.pi)]\n\n    min_num_neighbours = len(df1)\n    if min_num_neighbours&lt;3: continue\n\n    hit_ids = df1.hit_id.values\n    x,y,z = df1[['x', 'y', 'z']].values.T\n    r  = (x**2 + y**2)**0.5\n    r  = r/1000\n    a  = np.arctan2(y,x)\n    c = np.cos(a)\n    s = np.sin(a)\n    #tree = KDTree(np.column_stack([a,r]), metric='euclidean')\n    tree = KDTree(np.column_stack([c, s, r]), metric='euclidean')\n\n\n    track_ids = list(df1.track_id.unique())\n    num_track_ids = len(track_ids)\n    min_length=3\n\n    for i in range(num_track_ids):\n        p = track_ids[i]\n        if p==0: continue\n\n        idx = np.where(df1.track_id==p)[0]\n        if len(idx)&lt;min_length: continue\n\n        if angle&gt;0:\n            idx = idx[np.argsort( z[idx])]\n        else:\n            idx = idx[np.argsort(-z[idx])]\n\n\n        ## start and end points  ##\n        idx0,idx1 = idx[0],idx[-1]\n        a0 = a[idx0]\n        a1 = a[idx1]\n        r0 = r[idx0]\n        r1 = r[idx1]\n        c0 = c[idx0]\n        c1 = c[idx1]\n        s0 = s[idx0]\n        s1 = s[idx1]\n\n        da0 = a[idx[1]] - a[idx[0]]  #direction\n        dr0 = r[idx[1]] - r[idx[0]]\n        direction0 = np.arctan2(dr0,da0)\n\n        da1 = a[idx[-1]] - a[idx[-2]]\n        dr1 = r[idx[-1]] - r[idx[-2]]\n        direction1 = np.arctan2(dr1,da1)\n\n\n\n        ## extend start point\n        ns = tree.query([[c0, s0, r0]], k=min(num_neighbours, min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n\n        direction = np.arctan2(r0 - r[ns], a0 - a[ns])\n        diff = 1 - np.cos(direction - direction0)\n        ns = ns[(r0 - r[ns] &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]\n        for n in ns: df.loc[df.hit_id == hit_ids[n], 'track_id'] = p\n\n        ## extend end point\n        ns = tree.query([[c1, s1, r1]], k=min(num_neighbours, min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n\n        direction = np.arctan2(r[ns] - r1, a[ns] - a1)\n        diff = 1 - np.cos(direction - direction1)\n        ns = ns[(r[ns] - r1 &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]\n        for n in ns:  df.loc[df.hit_id == hit_ids[n], 'track_id'] = p\n\n#print ('\\r')\ndf = df[['event_id', 'hit_id', 'track_id']]\nreturn df\n</code></pre>",
      "rawMarkdown": "my updated code\n\n\n    def extend(submission,hits,limit=0.04, num_neighbours=18):\n\n    df = submission.merge(hits,  on=['hit_id'], how='left')\n    df = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n    df = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n    df = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n    for angle in range(-90,90,1):\n\n        print ('\\r %f'%angle, end='',flush=True)\n        #df1 = df.loc[(df.arctan2&gt;(angle-0.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+0.5)/180*np.pi)]\n        df1 = df.loc[(df.arctan2&gt;(angle-1.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+1.5)/180*np.pi)]\n\n        min_num_neighbours = len(df1)\n        if min_num_neighbours&lt;3: continue\n\n        hit_ids = df1.hit_id.values\n        x,y,z = df1[['x', 'y', 'z']].values.T\n        r  = (x**2 + y**2)**0.5\n        r  = r/1000\n        a  = np.arctan2(y,x)\n        c = np.cos(a)\n        s = np.sin(a)\n        #tree = KDTree(np.column_stack([a,r]), metric='euclidean')\n        tree = KDTree(np.column_stack([c, s, r]), metric='euclidean')\n\n\n        track_ids = list(df1.track_id.unique())\n        num_track_ids = len(track_ids)\n        min_length=3\n\n        for i in range(num_track_ids):\n            p = track_ids[i]\n            if p==0: continue\n\n            idx = np.where(df1.track_id==p)[0]\n            if len(idx)",
      "votes": 10,
      "replies": [
        {
          "id": 361863,
          "postDate": "2018-07-25T08:05:08.233Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 361867,
          "postDate": "2018-07-25T08:09:21.693Z",
          "content": "<p><a href=\"/starhao\">@starhao</a> This is a bug. One round takes a few minutes.</p>",
          "rawMarkdown": "@starhao This is a bug. One round takes a few minutes.",
          "votes": 1
        },
        {
          "id": 361868,
          "postDate": "2018-07-25T08:11:47.417Z",
          "content": "<p>Hello.. So submission is the complete dataframe and hits is the hit_id column in the dataframe?</p>",
          "rawMarkdown": "Hello.. So submission is the complete dataframe and hits is the hit_id column in the dataframe?"
        },
        {
          "id": 361871,
          "postDate": "2018-07-25T08:19:42.797Z",
          "content": "<p>@Samrat P</p>\n\n<p>submission is a submission (dataframe) for an event (columns: event_id,hit_id,track_id)</p>\n\n<p>hits is a hits dataframe, obtained via load_dataset function</p>",
          "rawMarkdown": "@Samrat P\n\nsubmission is a submission (dataframe) for an event (columns: event_id,hit_id,track_id)\n\nhits is a hits dataframe, obtained via load_dataset function",
          "votes": 1
        },
        {
          "id": 361872,
          "postDate": "2018-07-25T08:24:33.387Z",
          "content": "<p>Thanks <a href=\"/sergeyzlobin\">@sergeyzlobin</a> ... Let me try that out....</p>",
          "rawMarkdown": "Thanks @sergeyzlobin ... Let me try that out...."
        },
        {
          "id": 361903,
          "postDate": "2018-07-25T09:31:45.857Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 361930,
          "postDate": "2018-07-25T10:22:56.153Z",
          "content": "<p><a href=\"/starhao\">@starhao</a></p>\n\n<p>I didn't have that problem. So I don't know. I just confirm that this code should be fast.\nI suggest you look where it hangs.</p>",
          "rawMarkdown": "@starhao\n\nI didn't have that problem. So I don't know. I just confirm that this code should be fast.\nI suggest you look where it hangs.",
          "votes": 1
        },
        {
          "id": 361939,
          "postDate": "2018-07-25T10:26:01.273Z",
          "content": "<p><a href=\"/starhao\">@starhao</a>\nDo you see a progress? It is in this line:\nprint ('\\r %f'%angle, end='',flush=True)</p>",
          "rawMarkdown": "@starhao\nDo you see a progress? It is in this line:\nprint ('\\r %f'%angle, end='',flush=True)",
          "votes": 1
        },
        {
          "id": 361977,
          "postDate": "2018-07-25T12:23:47.900Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 361980,
          "postDate": "2018-07-25T12:28:41.947Z",
          "content": "<p>You have wrong hits. It should goes from the Kaggle dataset, not from sumbission. For example,</p>\n\n<p><strong>hits</strong>, cells, particles, truth = load_event(data_dir + '/event' + event_id)</p>\n\n<p>or using load_dataset function</p>",
          "rawMarkdown": "You have wrong hits. It should goes from the Kaggle dataset, not from sumbission. For example,\n\n   **hits**, cells, particles, truth = load_event(data_dir + '/event' + event_id)\n\nor using load_dataset function",
          "votes": 2
        },
        {
          "id": 362002,
          "postDate": "2018-07-25T13:18:27.063Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 362039,
          "postDate": "2018-07-25T14:43:59.587Z",
          "content": "<p><a href=\"/sergeyzlobin\">@sergeyzlobin</a> I still couldn't get it working.. Could you please let me know if I'm missing something..</p>\n\n<pre><code>sub_df = pd.read_csv(\"../input/trackml-best/sub_jul_23.csv\")\npath_to_test = \"../input/trackml-particle-identification/test\"\n\n all_hits = pd.DataFrame()\n for event_id, hits in load_dataset(path_to_test, parts=['hits']):\n     all_hits=all_hits.append(hits)\n\n for i in range(8):\n      sub = extend(sub_df, all_hits)\nsub.to_csv('submission_final.csv', index=False)\n</code></pre>\n\n<p>sub_df.shape =&gt; (13741466, 3)\nall_hits.shape =&gt; (13741466, 7)</p>\n\n<p>I'm getting below error:</p>\n\n<p>ValueError: operands could not be broadcast together with shapes (18,) (108117,7) </p>\n\n<p>---&gt; 67             ns = ns[(r0 - r[ns] &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]</p>",
          "rawMarkdown": "@sergeyzlobin I still couldn't get it working.. Could you please let me know if I'm missing something..\n\n    sub_df = pd.read_csv(\"../input/trackml-best/sub_jul_23.csv\")\n    path_to_test = \"../input/trackml-particle-identification/test\"\n\n     all_hits = pd.DataFrame()\n     for event_id, hits in load_dataset(path_to_test, parts=['hits']):\n         all_hits=all_hits.append(hits)\n\n     for i in range(8):\n          sub = extend(sub_df, all_hits)\n    sub.to_csv('submission_final.csv', index=False)\n\nsub_df.shape =&gt; (13741466, 3)\nall_hits.shape =&gt; (13741466, 7)\n\nI'm getting below error:\n\nValueError: operands could not be broadcast together with shapes (18,) (108117,7) \n\n---&gt; 67             ns = ns[(r0 - r[ns] &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]\n"
        },
        {
          "id": 362164,
          "postDate": "2018-07-25T20:02:13.157Z",
          "content": "<p>The function 'extend' should be used for every event independently.</p>",
          "rawMarkdown": "The function 'extend' should be used for every event independently.",
          "votes": 1
        },
        {
          "id": 362247,
          "postDate": "2018-07-26T02:10:05.420Z",
          "content": "<p>Thank You.. <a href=\"/sergeyzlobin\">@sergeyzlobin</a></p>",
          "rawMarkdown": "Thank You.. @sergeyzlobin"
        },
        {
          "id": 362277,
          "postDate": "2018-07-26T04:25:55.520Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 362412,
          "postDate": "2018-07-26T09:51:44.263Z",
          "content": "<blockquote>\n  <p>Am I missing on something?</p>\n</blockquote>\n\n<p>A bit of work maybe?</p>\n\n<p>Sorry to be blunt, but you need to add you twist to what is shared here.  A lot have been shared in the forum, here and in other topics, and using all that has been shared is enough to get over 0.7.</p>",
          "rawMarkdown": "&gt; Am I missing on something?\n\nA bit of work maybe?\n\nSorry to be blunt, but you need to add you twist to what is shared here.  A lot have been shared in the forum, here and in other topics, and using all that has been shared is enough to get over 0.7."
        },
        {
          "id": 362471,
          "postDate": "2018-07-26T12:50:53.227Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 362498,
          "postDate": "2018-07-26T13:45:41.077Z",
          "content": "<p>Read the forum.  Ask questions when you don't understand what people say.  Here Heng said his code needs parameter tuning for instance.  Have you tried that ?</p>",
          "rawMarkdown": "Read the forum.  Ask questions when you don't understand what people say.  Here Heng said his code needs parameter tuning for instance.  Have you tried that ?",
          "votes": 1
        },
        {
          "id": 363369,
          "postDate": "2018-07-28T19:04:06.160Z",
          "content": "<p>I responded that way because I think you should spend time reading the forum rather than asking me (or others) to spend time re reading it for you. This said, what I use is pretty much described by what yuval, Nicole Finnie, Grzegorz Sionkowski and I disclosed in the forum.  Just read our posts.  For other types of approaches read Heng Cher Keng posts, he lists a number of deep learning approaches, and often provides some code.  My current approach targets tracks originating near the z axis, hence cannot yield a score above 082 or so.</p>",
          "rawMarkdown": "I responded that way because I think you should spend time reading the forum rather than asking me (or others) to spend time re reading it for you. This said, what I use is pretty much described by what yuval, Nicole Finnie, Grzegorz Sionkowski and I disclosed in the forum.  Just read our posts.  For other types of approaches read Heng Cher Keng posts, he lists a number of deep learning approaches, and often provides some code.  My current approach targets tracks originating near the z axis, hence cannot yield a score above 082 or so.",
          "votes": 1
        },
        {
          "id": 363397,
          "postDate": "2018-07-28T20:58:41.847Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 363555,
          "postDate": "2018-07-29T13:41:46.423Z",
          "content": "<p>@Zijun Yao</p>\n\n<p>Everything looks good. I do almost the same.\nI can only suggest you try to load hits with the TrackML library like this</p>\n\n<p>hits, cells, particles, truth = load_event('Data/train_100_events/event000001000')</p>",
          "rawMarkdown": "@Zijun Yao\n\nEverything looks good. I do almost the same.\nI can only suggest you try to load hits with the TrackML library like this\n\nhits, cells, particles, truth = load_event('Data/train_100_events/event000001000')\n\n"
        }
      ]
    },
    {
      "id": 358942,
      "postDate": "2018-07-19T08:07:02.403Z",
      "content": "<p>Thank you very much for sharing your code, it greatly worked for me ! </p>\n\n<p>I re-write it in R for those of you who may be interested.</p>",
      "rawMarkdown": "Thank you very much for sharing your code, it greatly worked for me ! \n\nI re-write it in R for those of you who may be interested.",
      "votes": 1
    },
    {
      "id": 341433,
      "postDate": "2018-06-11T15:24:07.330Z",
      "content": "<p>I wonder if any kaggler has time to try this experiment:</p>\n\n<p>1) DBSCAN on point set:</p>\n\n<ul>\n<li><p>make a cone slice.</p></li>\n<li><p>if cone slice = x1,x2,x3,x4 ...x10, it means each data point has 3 dim={x,yz} and there are 10 data points. </p></li>\n<li><p>apply dbscan.</p></li>\n</ul>\n\n<p>2) DBSCAN on pair set:</p>\n\n<ul>\n<li><p>make a cone slice.</p></li>\n<li><p>if cone slice = x1,x2,x3,x4, ... x10. make pairs. one pair has 4 dim: pij ={ midpoint(xi,xj), direction(xi to xj) } = { x,y,z, theta }. there are 10x10=100 pairs.</p></li>\n<li><p>apply dbscan on pairs </p></li>\n<li><p>decode clustered pairs into points. e,g, if a clustered pairs ={p12, p23,p34}, then decoded results={x1,x2,x3,x4}</p></li>\n</ul>",
      "rawMarkdown": "I wonder if any kaggler has time to try this experiment:\n\n1) DBSCAN on point set:\n\n - make a cone slice.\n\n -  if cone slice = x1,x2,x3,x4 ...x10, it means each data point has 3 dim={x,yz} and there are 10 data points. \n\n - apply dbscan.\n\n2) DBSCAN on pair set:\n\n - make a cone slice.\n\n -  if cone slice = x1,x2,x3,x4, ... x10. make pairs. one pair has 4 dim: pij ={ midpoint(xi,xj), direction(xi to xj) } = { x,y,z, theta }. there are 10x10=100 pairs.\n\n -  apply dbscan on pairs \n\n - decode clustered pairs into points. e,g, if a clustered pairs ={p12, p23,p34}, then decoded results={x1,x2,x3,x4}",
      "votes": 1,
      "replies": [
        {
          "id": 341484,
          "postDate": "2018-06-11T16:55:35.923Z",
          "content": "<p>For such a series of point pairs, to reduce the total input you could try to limit your input pairs to potentially adjacent pairs. For example, a track containing hits x1, x2, x3, x4, ... x10 would only want to consider pairs {x1, x2}, {x2, x3}, etc. </p>\n\n<p>To accomplish this, only include pairs where the pair distance is less than a some threshold. You could also include exclude pairs where the angle with respect to z changes more than some other threshold. Both of these thresholds would probably scale with z. The reason I think these criteria make sense is because adjacent pairs on a track wouldn't be too far apart (very close to origin and on the edge of detectors shouldn't be adjacent on a track), or vary too greatly in angle(pairs of hits won't be on the opposite 'edge' of a cone), but hits tend to have more variance in their angles and occur less closely at more extreme values of z. </p>\n\n<p>The downside of trying to limit to only pairs is that you potentially lose some confidence in your clusters. For the above hypothetical track, {x1, x2} and {x2, x3} would cluster together, but {x1, x3} would also cluster together and would be a midpoint between the two. </p>\n\n<p>Possibly an iterative approach where early on in the iterations, nearby pairs are considered, such as  {x1, x2} and {x2, x3}, and then for and given potential cluster, {x1, x3} would also be evaluated. Eventually a cluster of all combinations could be evaluated for that track, and would hopefully reduce the total amount of pairs evaluated.</p>",
          "rawMarkdown": "For such a series of point pairs, to reduce the total input you could try to limit your input pairs to potentially adjacent pairs. For example, a track containing hits x1, x2, x3, x4, ... x10 would only want to consider pairs {x1, x2}, {x2, x3}, etc. \n\nTo accomplish this, only include pairs where the pair distance is less than a some threshold. You could also include exclude pairs where the angle with respect to z changes more than some other threshold. Both of these thresholds would probably scale with z. The reason I think these criteria make sense is because adjacent pairs on a track wouldn't be too far apart (very close to origin and on the edge of detectors shouldn't be adjacent on a track), or vary too greatly in angle(pairs of hits won't be on the opposite 'edge' of a cone), but hits tend to have more variance in their angles and occur less closely at more extreme values of z. \n\nThe downside of trying to limit to only pairs is that you potentially lose some confidence in your clusters. For the above hypothetical track, {x1, x2} and {x2, x3} would cluster together, but {x1, x3} would also cluster together and would be a midpoint between the two. \n\nPossibly an iterative approach where early on in the iterations, nearby pairs are considered, such as  {x1, x2} and {x2, x3}, and then for and given potential cluster, {x1, x3} would also be evaluated. Eventually a cluster of all combinations could be evaluated for that track, and would hopefully reduce the total amount of pairs evaluated."
        }
      ]
    },
    {
      "id": 340349,
      "postDate": "2018-06-09T00:59:16.570Z",
      "content": "<p>Wow! </p>",
      "rawMarkdown": "Wow! ",
      "votes": 1
    },
    {
      "id": 339412,
      "postDate": "2018-06-06T22:58:47.583Z",
      "content": "<p>Thank you for this! I have also been trying to extend lines but with a much slower, less efficient algorithm.</p>",
      "rawMarkdown": "Thank you for this! I have also been trying to extend lines but with a much slower, less efficient algorithm.",
      "votes": 1
    },
    {
      "id": 338333,
      "postDate": "2018-06-04T20:56:54.117Z",
      "content": "<p>I've been attempting to extend tracks as well, but with a different approach. I haven't submitted anything yet since I've not had any promising results... I either get a very small amount of true positives or I start getting too many false positives. </p>\n\n<p>Any critique on my approach would be appreciated:</p>\n\n<p>Using the dbscan tracks as inputs, I attempt to fit a regression to the tracks and then finding points that are very close to the projected track. Normally this would be very inefficient, so I added in some filter criteria... \n1. Only include hits in a similar 'cone' around the z axis\n2. Only include hits where the z is greater than the max z in the track, or less than the minimum (it seems to be the edges that get missed with the dbscan approach)</p>\n\n<p>My model worked decently when I only tried to find linear tracks, but trying to fit helices has been hard, I'm currently trying to fit with the method by Luis Andre Dutra e Silva...</p>",
      "rawMarkdown": "I've been attempting to extend tracks as well, but with a different approach. I haven't submitted anything yet since I've not had any promising results... I either get a very small amount of true positives or I start getting too many false positives. \n\nAny critique on my approach would be appreciated:\n\nUsing the dbscan tracks as inputs, I attempt to fit a regression to the tracks and then finding points that are very close to the projected track. Normally this would be very inefficient, so I added in some filter criteria... \n1. Only include hits in a similar 'cone' around the z axis\n2. Only include hits where the z is greater than the max z in the track, or less than the minimum (it seems to be the edges that get missed with the dbscan approach)\n\nMy model worked decently when I only tried to find linear tracks, but trying to fit helices has been hard, I'm currently trying to fit with the method by Luis Andre Dutra e Silva...",
      "votes": 1
    },
    {
      "id": 338841,
      "postDate": "2018-06-05T20:42:03.060Z",
      "content": "<p>Thanks for that code, I shamelessly reused it ;)</p>",
      "rawMarkdown": "Thanks for that code, I shamelessly reused it ;)",
      "replies": [
        {
          "id": 339567,
          "postDate": "2018-06-07T06:25:42.807Z",
          "content": "<p>The boost I get from it goes down as my score got better.  Now it is about 0.02 (in the current run I'll submit tomorrow).</p>",
          "rawMarkdown": "The boost I get from it goes down as my score got better.  Now it is about 0.02 (in the current run I'll submit tomorrow)."
        },
        {
          "id": 340038,
          "postDate": "2018-06-08T07:46:14.350Z",
          "content": "<p>@CPMP</p>\n\n<p>the code is a sample code. you can make some visualization and improve on it. I think such post processing  for linking can get LB around 0.64.</p>\n\n<p>there is a bug in the code: discontinuity for 0 and 2pi when comparing angular distance</p>",
          "rawMarkdown": "@CPMP\n\nthe code is a sample code. you can make some visualization and improve on it. I think such post processing  for linking can get LB around 0.64.\n\n\nthere is a bug in the code: discontinuity for 0 and 2pi when comparing angular distance",
          "votes": 1
        },
        {
          "id": 340042,
          "postDate": "2018-06-08T08:05:34.767Z",
          "content": "<p>@Heng, I had the same problem (0 and 2*pi), when I had to use the angle and converting it into two features (sin and cos) was not acceptable. I just turned the XY plane by pi, performed all operations once more and then ensembled both results.  Turning the XY plane by pi is very easy: x = -x,  y= -y.</p>",
          "rawMarkdown": "@Heng, I had the same problem (0 and 2*pi), when I had to use the angle and converting it into two features (sin and cos) was not acceptable. I just turned the XY plane by pi, performed all operations once more and then ensembled both results.  Turning the XY plane by pi is very easy: x = -x,  y= -y.",
          "votes": 3
        },
        {
          "id": 340043,
          "postDate": "2018-06-08T08:11:23.923Z",
          "content": "<p>@Heng,  I indeed plan to implement something more complex.  I also think one can reach 0.64 with better back fitting.  We'll see if I am right this week end ;)</p>\n\n<p>@Grzegorz, I turn xy by pi/2 in other parts of my code with x = y, y = -x</p>",
          "rawMarkdown": "@Heng,  I indeed plan to implement something more complex.  I also think one can reach 0.64 with better back fitting.  We'll see if I am right this week end ;)\n\n@Grzegorz, I turn xy by pi/2 in other parts of my code with x = y, y = -x",
          "votes": 2
        },
        {
          "id": 340044,
          "postDate": "2018-06-08T08:13:52.660Z",
          "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>thanks for the suggestion. An alternative is to change a to (cos(a), sin(a)), i.e. unit direction vector. Then angular distance can measured by dot product (i.e. cosine distance)</p>",
          "rawMarkdown": "@Grzegorz Sionkowski\n\nthanks for the suggestion. An alternative is to change a to (cos(a), sin(a)), i.e. unit direction vector. Then angular distance can measured by dot product (i.e. cosine distance)",
          "votes": 2
        },
        {
          "id": 340223,
          "postDate": "2018-06-08T16:47:29.910Z",
          "content": "<p>I tried both approaches of @CPMP and @Grzegorz by shifting <code>pi</code> or <code>pi/2</code>, unsurprisingly, they yielded exactly the same local scores since they're meant to solve the same problem. The improvement is only marginal and it took twice as long for clustering since I had to run the same operations on the shifted plane again, so I guess it's not the secret of your current public LB scores, too bad :p</p>",
          "rawMarkdown": "I tried both approaches of @CPMP and @Grzegorz by shifting `pi` or `pi/2`, unsurprisingly, they yielded exactly the same local scores since they're meant to solve the same problem. The improvement is only marginal and it took twice as long for clustering since I had to run the same operations on the shifted plane again, so I guess it's not the secret of your current public LB scores, too bad :p"
        },
        {
          "id": 340234,
          "postDate": "2018-06-08T17:12:40.377Z",
          "content": "<p>@Nicole, I was thinking the same thing driving to work today (small improvement for twice the computation).</p>\n\n<p>However, since you are only trying to avoid lost opportunity from the discontinuity at 0, 2*pi, you don't need to scan the entire angular range again...</p>\n\n<p>Maybe you can get the improvement for free by scanning half the range, rotate by pi and scan again to cover the rest of the range? </p>\n\n<p>BTW - I also want to think about @Heng's comment (always a good idea) to measure angular distance using cosine similarity or some similar technique.  This would naturally eliminate the discontinuity. </p>",
          "rawMarkdown": "@Nicole, I was thinking the same thing driving to work today (small improvement for twice the computation).\n\nHowever, since you are only trying to avoid lost opportunity from the discontinuity at 0, 2*pi, you don't need to scan the entire angular range again...\n\nMaybe you can get the improvement for free by scanning half the range, rotate by pi and scan again to cover the rest of the range? \n\nBTW - I also want to think about @Heng's comment (always a good idea) to measure angular distance using cosine similarity or some similar technique.  This would naturally eliminate the discontinuity. ",
          "votes": 1
        },
        {
          "id": 340243,
          "postDate": "2018-06-08T17:53:18.373Z",
          "content": "<p>Hey @John, thanks for the idea, that was my first attempt cutting down the steps  by half but I forgot to change the length of angular displacement intervals, and that hurt the accuracy but after having read your comment I tried again and fixed the bug. Thanks! There's almost no improvement in the score though. It's good to see more discussions here, somehow not too many teams have entered this competition. </p>",
          "rawMarkdown": "Hey @John, thanks for the idea, that was my first attempt cutting down the steps  by half but I forgot to change the length of angular displacement intervals, and that hurt the accuracy but after having read your comment I tried again and fixed the bug. Thanks! There's almost no improvement in the score though. It's good to see more discussions here, somehow not too many teams have entered this competition. "
        },
        {
          "id": 340302,
          "postDate": "2018-06-08T20:58:53.697Z",
          "content": "<p>What is the issue with working with the sin/cos of the angle? It seems to work effectively for the clustering portion of the challenge. Is it just because the sin/cos of the angle isn't as easy to implement in your track extension framework?</p>\n\n<p>I think the smaller entrant count is a combination of the smaller prize pool, the far away closing date, and the challenge seeming more intimidating at first glance. I imagine several people see physics and step away, even if the competition is mainly based on data.</p>",
          "rawMarkdown": "What is the issue with working with the sin/cos of the angle? It seems to work effectively for the clustering portion of the challenge. Is it just because the sin/cos of the angle isn't as easy to implement in your track extension framework?\n\nI think the smaller entrant count is a combination of the smaller prize pool, the far away closing date, and the challenge seeming more intimidating at first glance. I imagine several people see physics and step away, even if the competition is mainly based on data."
        },
        {
          "id": 340566,
          "postDate": "2018-06-09T16:01:44.643Z",
          "content": "<p>I used cosine similarity to avoid the 0 vs 2*pi discontinuity.  I used the head and tail points to create directional unit vectors, and then used a dot product to measure the angular distance between unit vectors.  Maybe I did something wrong, but it only helped a very little bit (+0.0002).  </p>\n\n<p>Here is a snippet of the change from @Heng's original code:</p>\n\n<p>`           da0 = a[idx[1]] - a[idx[0]]  #direction\n            dr0 = r[idx[1]] - r[idx[0]]\n            divisor0 = (da0*<em>2+dr0</em>*2)**0.5\n            if divisor0 == 0 : divisor0 = 1\n            direction0 = np.array([da0/divisor0,dr0/divisor0])</p>\n\n<pre><code>        da1 = a[idx[-1]] - a[idx[-2]]\n        dr1 = r[idx[-1]] - r[idx[-2]]\n        divisor1 = (da1**2+dr1**2)**0.5\n        if divisor1 == 0: divisor1 = 1\n        direction1 = np.array([da1/divisor1,dr1/divisor1]) \n\n\n\n        ## extend start point\n        ns = tree.query([[a0,r0]], k=min(20,min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n        da0ns = a0-a[ns]\n        dr0ns = r0-r[ns]\n        divisor0ns = (da0ns**2+dr0ns**2)**0.5\n        divisor0ns[np.where(divisor0ns==0)]=1\n\n        direction = np.array([da0ns/divisor0ns,dr0ns/divisor0ns]) \n        ns = ns[(r0-r[ns]&gt;0.01) &amp;(np.matmul(direction.T,direction0)&gt;0.9991)]\n\n        for n in ns:\n            df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p \n\n        ## extend end point\n        ns = tree.query([[a1,r1]], k=min(20,min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n        da1ns = a[ns]-a1\n        dr1ns = r[ns]-r1\n        divisor1ns = (da1ns**2+dr1ns**2)**0.5\n        divisor1ns[np.where(divisor1ns==0)]=1\n\n        direction = np.array([da1ns/divisor1ns,dr1ns/divisor1ns]) \n        ns = ns[(r[ns]-r1&gt;0.01) &amp;(np.matmul(direction.T,direction1)&gt;0.9991)] \n\n        for n in ns:\n            df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p`\n</code></pre>",
          "rawMarkdown": "I used cosine similarity to avoid the 0 vs 2*pi discontinuity.  I used the head and tail points to create directional unit vectors, and then used a dot product to measure the angular distance between unit vectors.  Maybe I did something wrong, but it only helped a very little bit (+0.0002).  \n\nHere is a snippet of the change from @Heng's original code:\n\n`           da0 = a[idx[1]] - a[idx[0]]  #direction\n            dr0 = r[idx[1]] - r[idx[0]]\n            divisor0 = (da0**2+dr0**2)**0.5\n            if divisor0 == 0 : divisor0 = 1\n            direction0 = np.array([da0/divisor0,dr0/divisor0])\n\n\n            da1 = a[idx[-1]] - a[idx[-2]]\n            dr1 = r[idx[-1]] - r[idx[-2]]\n            divisor1 = (da1**2+dr1**2)**0.5\n            if divisor1 == 0: divisor1 = 1\n            direction1 = np.array([da1/divisor1,dr1/divisor1]) \n\n\n\t\n\t\t\t## extend start point\n            ns = tree.query([[a0,r0]], k=min(20,min_num_neighbours), return_distance=False)\n            ns = np.concatenate(ns)\n            da0ns = a0-a[ns]\n            dr0ns = r0-r[ns]\n            divisor0ns = (da0ns**2+dr0ns**2)**0.5\n            divisor0ns[np.where(divisor0ns==0)]=1\n            \n            direction = np.array([da0ns/divisor0ns,dr0ns/divisor0ns]) \n            ns = ns[(r0-r[ns]&gt;0.01) &amp;(np.matmul(direction.T,direction0)&gt;0.9991)]\n\t\n            for n in ns:\n                df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p \n\n\t\t\t## extend end point\n            ns = tree.query([[a1,r1]], k=min(20,min_num_neighbours), return_distance=False)\n            ns = np.concatenate(ns)\n            da1ns = a[ns]-a1\n            dr1ns = r[ns]-r1\n            divisor1ns = (da1ns**2+dr1ns**2)**0.5\n            divisor1ns[np.where(divisor1ns==0)]=1\n            \n            direction = np.array([da1ns/divisor1ns,dr1ns/divisor1ns]) \n            ns = ns[(r[ns]-r1&gt;0.01) &amp;(np.matmul(direction.T,direction1)&gt;0.9991)] \n\t\t\t\n            for n in ns:\n                df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p`\n\n",
          "votes": 3
        },
        {
          "id": 340648,
          "postDate": "2018-06-09T22:15:42.200Z",
          "content": "<p>@John, thanks for sharing your code, highly appreciated. As said I shifted the plane by <code>pi/2</code> as suggested by @CPMP in the clustering code (not just the backfitting code) and the improvement was marginal, around 0.002 for clustering (backfitting too). I forwardfit on the original XY plane and backwardfit on the shifted plane, however I removed the code for now since the gain is marginal. There are only about 1500 hits between -1 and 1 degrees, so I <em>guess</em> that's why the discontinuity doesn't hurt the local score too much. As mentioned by @Heng and @CPMP we can still improve the backfitting code to get an accuracy above 0.6. I'm working on it.</p>",
          "rawMarkdown": "@John, thanks for sharing your code, highly appreciated. As said I shifted the plane by `pi/2` as suggested by @CPMP in the clustering code (not just the backfitting code) and the improvement was marginal, around 0.002 for clustering (backfitting too). I forwardfit on the original XY plane and backwardfit on the shifted plane, however I removed the code for now since the gain is marginal. There are only about 1500 hits between -1 and 1 degrees, so I *guess* that's why the discontinuity doesn't hurt the local score too much. As mentioned by @Heng and @CPMP we can still improve the backfitting code to get an accuracy above 0.6. I'm working on it."
        },
        {
          "id": 340686,
          "postDate": "2018-06-10T02:36:04.287Z",
          "content": "<p>@John, I made a similar modification yesterday, this gave me a 0.005 boost.</p>\n\n<p>The main difference with yours is I replace</p>\n\n<pre><code>    divisor1 = (da1**2+dr1**2)**0.5\n    if divisor1 == 0: divisor1 = 1\n</code></pre>\n\n<p>with</p>\n\n<pre><code>divisor1 = (da1**2+dr1**2)**0.5 + 1e-6\n</code></pre>",
          "rawMarkdown": "@John, I made a similar modification yesterday, this gave me a 0.005 boost.\n\nThe main difference with yours is I replace\n\n        divisor1 = (da1**2+dr1**2)**0.5\n        if divisor1 == 0: divisor1 = 1\n\nwith\n\n    divisor1 = (da1**2+dr1**2)**0.5 + 1e-6",
          "votes": 3
        },
        {
          "id": 340744,
          "postDate": "2018-06-10T07:35:50.140Z",
          "content": "<p>@John, your code seems right to me, I changed the threshold to <code>0.999</code> to find more nearest neighbours, the improvement on all cone slices for event 1000 is around <code>0.002</code> for the event 1000. I guess you put one zero too much.</p>\n\n<p>I scored on the found tracks using this cone slicing approach. </p>\n\n<pre><code>Cone slice score for event 1000: 0.25068574\nCone slice score for event 1000: 0.25284123\n</code></pre>",
          "rawMarkdown": "@John, your code seems right to me, I changed the threshold to `0.999` to find more nearest neighbours, the improvement on all cone slices for event 1000 is around `0.002` for the event 1000. I guess you put one zero too much.\n\nI scored on the found tracks using this cone slicing approach. \n\n    Cone slice score for event 1000: 0.25068574\n    Cone slice score for event 1000: 0.25284123\n\n"
        },
        {
          "id": 340761,
          "postDate": "2018-06-10T08:37:12.723Z",
          "content": "<p>some notes:</p>\n\n<ol>\n<li><p>create a set of different models. The objective is that the models must detect different tracks. Hence score of each model may not be necessary high.  the only way to tell is to draw out the tracks and visually inspect. </p></li>\n<li><p>then run extend() to extend the tracks. you can use decrease thresholds when you run multiple runs. This detect more curvy tracks.</p></li>\n</ol>",
          "rawMarkdown": "some notes:\n\n1. create a set of different models. The objective is that the models must detect different tracks. Hence score of each model may not be necessary high.  the only way to tell is to draw out the tracks and visually inspect. \n\n2. then run extend() to extend the tracks. you can use decrease thresholds when you run multiple runs. This detect more curvy tracks.",
          "votes": 2
        }
      ]
    },
    {
      "id": 361515,
      "postDate": "2018-07-24T16:23:28.733Z",
      "content": "<p>Hello <a href=\"/hengck23\">@hengck23</a> / <a href=\"/crysis\">@crysis</a> ... I'm using the python script for track extension and I have a small doubt.. So, in the extend routine the submission variable takes the whole DataFrame of the submission file and the hits variable takes the hits values - submission[[hit_id]].. Am I correct? Coz even with around 60GB RAM I'm getting a MemoryError in the first line of the code where the merge is done.... Thanks in Adv...</p>",
      "rawMarkdown": "Hello @hengck23 / @crysis ... I'm using the python script for track extension and I have a small doubt.. So, in the extend routine the submission variable takes the whole DataFrame of the submission file and the hits variable takes the hits values - submission[[hit_id]].. Am I correct? Coz even with around 60GB RAM I'm getting a MemoryError in the first line of the code where the merge is done.... Thanks in Adv..."
    },
    {
      "id": 340777,
      "postDate": "2018-06-10T09:31:14.270Z",
      "content": "<p>for i in range(8): \n       submission = extend(submission, hits)</p>\n\n<p>I didn't get it. What‘s the meaning of the i？</p>",
      "rawMarkdown": "for i in range(8): \n       submission = extend(submission, hits)\n\n\nI didn't get it. What‘s the meaning of the i？",
      "replies": [
        {
          "id": 340810,
          "postDate": "2018-06-10T11:41:05.923Z",
          "content": "<p>You can change <code>i</code> to an underscore if it bothers you, but I guess you were not asking a programmatic question. I use <code>i</code> to do different post processing (or extend) and change values passed in each time, e.g. <code>extend(submission, hits, shift_plane=i%2 == 1)</code> and the labels get improved after each post-processing run. If you use @Heng's code directly, <code>8</code> is the optimal value as a trade-off between time and accuracy, at least for me.</p>",
          "rawMarkdown": "You can change `i` to an underscore if it bothers you, but I guess you were not asking a programmatic question. I use `i` to do different post processing (or extend) and change values passed in each time, e.g. `extend(submission, hits, shift_plane=i%2 == 1)` and the labels get improved after each post-processing run. If you use @Heng's code directly, `8` is the optimal value as a trade-off between time and accuracy, at least for me.",
          "votes": 1
        },
        {
          "id": 340833,
          "postDate": "2018-06-10T12:37:42.257Z",
          "content": "<p>It means extend is applied 8 times.</p>",
          "rawMarkdown": "It means extend is applied 8 times.",
          "votes": 1
        }
      ]
    },
    {
      "id": 340473,
      "postDate": "2018-06-09T09:06:26.480Z",
      "content": "<p>@Heng, In your code df.arctan2 ranges from -pi/2 to pi/2, but you scan the interval [-pi, pi].  </p>\n\n<p>Not a big deal but I thought I should report it.</p>",
      "rawMarkdown": "@Heng, In your code df.arctan2 ranges from -pi/2 to pi/2, but you scan the interval [-pi, pi].  \n\nNot a big deal but I thought I should report it.",
      "replies": [
        {
          "id": 340535,
          "postDate": "2018-06-09T13:22:05.787Z",
          "content": "<p>Thanks. I confirm that scanning from -pi/2 to pi/2 is enough</p>",
          "rawMarkdown": "Thanks. I confirm that scanning from -pi/2 to pi/2 is enough",
          "votes": 2
        }
      ]
    },
    {
      "id": 338499,
      "postDate": "2018-06-05T07:50:33.453Z",
      "content": "<p>Thank you for sharing your approach!</p>\n\n<p>I've not investigated extending tracks yet. I understand why we can lose last hits (due to fluctuations within every hit with a detector). However I don't understand why we lose a first hit.</p>",
      "rawMarkdown": "Thank you for sharing your approach!\n\nI've not investigated extending tracks yet. I understand why we can lose last hits (due to fluctuations within every hit with a detector). However I don't understand why we lose a first hit."
    },
    {
      "id": 338300,
      "postDate": "2018-06-04T19:21:09.457Z",
      "content": "<p>Thank-you for all of the information you have been posting!  Much appreciated!  in your code above you define arctan2 = np.arctan2(df.z, df.r).  Should this be np.arctan2(df.r, df.z) ??</p>",
      "rawMarkdown": "Thank-you for all of the information you have been posting!  Much appreciated!  in your code above you define arctan2 = np.arctan2(df.z, df.r).  Should this be np.arctan2(df.r, df.z) ??"
    },
    {
      "id": 338042,
      "postDate": "2018-06-04T09:58:48.770Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 340577,
      "postDate": "2018-06-09T17:10:26.590Z",
      "content": "<p>Thank you, that's really useful!</p>",
      "rawMarkdown": "Thank you, that's really useful!"
    }
  ],
  "comments": [
    {
      "id": 340763,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-10T08:41:46.270000",
      "content": "<p>my updated code</p>\n\n<pre><code>def extend(submission,hits,limit=0.04, num_neighbours=18):\n\ndf = submission.merge(hits,  on=['hit_id'], how='left')\ndf = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\ndf = df.assign(r = np.sqrt( df.x**2 + df.y**2))\ndf = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\nfor angle in range(-90,90,1):\n\n    print ('\\r %f'%angle, end='',flush=True)\n    #df1 = df.loc[(df.arctan2&gt;(angle-0.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+0.5)/180*np.pi)]\n    df1 = df.loc[(df.arctan2&gt;(angle-1.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+1.5)/180*np.pi)]\n\n    min_num_neighbours = len(df1)\n    if min_num_neighbours&lt;3: continue\n\n    hit_ids = df1.hit_id.values\n    x,y,z = df1[['x', 'y', 'z']].values.T\n    r  = (x**2 + y**2)**0.5\n    r  = r/1000\n    a  = np.arctan2(y,x)\n    c = np.cos(a)\n    s = np.sin(a)\n    #tree = KDTree(np.column_stack([a,r]), metric='euclidean')\n    tree = KDTree(np.column_stack([c, s, r]), metric='euclidean')\n\n\n    track_ids = list(df1.track_id.unique())\n    num_track_ids = len(track_ids)\n    min_length=3\n\n    for i in range(num_track_ids):\n        p = track_ids[i]\n        if p==0: continue\n\n        idx = np.where(df1.track_id==p)[0]\n        if len(idx)&lt;min_length: continue\n\n        if angle&gt;0:\n            idx = idx[np.argsort( z[idx])]\n        else:\n            idx = idx[np.argsort(-z[idx])]\n\n\n        ## start and end points  ##\n        idx0,idx1 = idx[0],idx[-1]\n        a0 = a[idx0]\n        a1 = a[idx1]\n        r0 = r[idx0]\n        r1 = r[idx1]\n        c0 = c[idx0]\n        c1 = c[idx1]\n        s0 = s[idx0]\n        s1 = s[idx1]\n\n        da0 = a[idx[1]] - a[idx[0]]  #direction\n        dr0 = r[idx[1]] - r[idx[0]]\n        direction0 = np.arctan2(dr0,da0)\n\n        da1 = a[idx[-1]] - a[idx[-2]]\n        dr1 = r[idx[-1]] - r[idx[-2]]\n        direction1 = np.arctan2(dr1,da1)\n\n\n\n        ## extend start point\n        ns = tree.query([[c0, s0, r0]], k=min(num_neighbours, min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n\n        direction = np.arctan2(r0 - r[ns], a0 - a[ns])\n        diff = 1 - np.cos(direction - direction0)\n        ns = ns[(r0 - r[ns] &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]\n        for n in ns: df.loc[df.hit_id == hit_ids[n], 'track_id'] = p\n\n        ## extend end point\n        ns = tree.query([[c1, s1, r1]], k=min(num_neighbours, min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n\n        direction = np.arctan2(r[ns] - r1, a[ns] - a1)\n        diff = 1 - np.cos(direction - direction1)\n        ns = ns[(r[ns] - r1 &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]\n        for n in ns:  df.loc[df.hit_id == hit_ids[n], 'track_id'] = p\n\n#print ('\\r')\ndf = df[['event_id', 'hit_id', 'track_id']]\nreturn df\n</code></pre>",
      "votes": 10,
      "replies": [
        {
          "id": 361863,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-25T08:05:08.233000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361867,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T08:09:21.693000",
          "content": "<p><a href=\"/starhao\">@starhao</a> This is a bug. One round takes a few minutes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361868,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-07-25T08:11:47.417000",
          "content": "<p>Hello.. So submission is the complete dataframe and hits is the hit_id column in the dataframe?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361871,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T08:19:42.797000",
          "content": "<p>@Samrat P</p>\n\n<p>submission is a submission (dataframe) for an event (columns: event_id,hit_id,track_id)</p>\n\n<p>hits is a hits dataframe, obtained via load_dataset function</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361872,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-07-25T08:24:33.387000",
          "content": "<p>Thanks <a href=\"/sergeyzlobin\">@sergeyzlobin</a> ... Let me try that out....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361903,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-25T09:31:45.857000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361930,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T10:22:56.153000",
          "content": "<p><a href=\"/starhao\">@starhao</a></p>\n\n<p>I didn't have that problem. So I don't know. I just confirm that this code should be fast.\nI suggest you look where it hangs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361939,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T10:26:01.273000",
          "content": "<p><a href=\"/starhao\">@starhao</a>\nDo you see a progress? It is in this line:\nprint ('\\r %f'%angle, end='',flush=True)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 361977,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-25T12:23:47.900000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361980,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T12:28:41.947000",
          "content": "<p>You have wrong hits. It should goes from the Kaggle dataset, not from sumbission. For example,</p>\n\n<p><strong>hits</strong>, cells, particles, truth = load_event(data_dir + '/event' + event_id)</p>\n\n<p>or using load_dataset function</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 362002,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-25T13:18:27.063000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362039,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-07-25T14:43:59.587000",
          "content": "<p><a href=\"/sergeyzlobin\">@sergeyzlobin</a> I still couldn't get it working.. Could you please let me know if I'm missing something..</p>\n\n<pre><code>sub_df = pd.read_csv(\"../input/trackml-best/sub_jul_23.csv\")\npath_to_test = \"../input/trackml-particle-identification/test\"\n\n all_hits = pd.DataFrame()\n for event_id, hits in load_dataset(path_to_test, parts=['hits']):\n     all_hits=all_hits.append(hits)\n\n for i in range(8):\n      sub = extend(sub_df, all_hits)\nsub.to_csv('submission_final.csv', index=False)\n</code></pre>\n\n<p>sub_df.shape =&gt; (13741466, 3)\nall_hits.shape =&gt; (13741466, 7)</p>\n\n<p>I'm getting below error:</p>\n\n<p>ValueError: operands could not be broadcast together with shapes (18,) (108117,7) </p>\n\n<p>---&gt; 67             ns = ns[(r0 - r[ns] &gt; 0.01) &amp; (diff &lt; (1 - np.cos(limit)))]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362164,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-25T20:02:13.157000",
          "content": "<p>The function 'extend' should be used for every event independently.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 362247,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-07-26T02:10:05.420000",
          "content": "<p>Thank You.. <a href=\"/sergeyzlobin\">@sergeyzlobin</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362277,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-26T04:25:55.520000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362412,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-26T09:51:44.263000",
          "content": "<blockquote>\n  <p>Am I missing on something?</p>\n</blockquote>\n\n<p>A bit of work maybe?</p>\n\n<p>Sorry to be blunt, but you need to add you twist to what is shared here.  A lot have been shared in the forum, here and in other topics, and using all that has been shared is enough to get over 0.7.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362471,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-26T12:50:53.227000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 362498,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-26T13:45:41.077000",
          "content": "<p>Read the forum.  Ask questions when you don't understand what people say.  Here Heng said his code needs parameter tuning for instance.  Have you tried that ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 363369,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-28T19:04:06.160000",
          "content": "<p>I responded that way because I think you should spend time reading the forum rather than asking me (or others) to spend time re reading it for you. This said, what I use is pretty much described by what yuval, Nicole Finnie, Grzegorz Sionkowski and I disclosed in the forum.  Just read our posts.  For other types of approaches read Heng Cher Keng posts, he lists a number of deep learning approaches, and often provides some code.  My current approach targets tracks originating near the z axis, hence cannot yield a score above 082 or so.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 363397,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-28T20:58:41.847000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 363555,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-07-29T13:41:46.423000",
          "content": "<p>@Zijun Yao</p>\n\n<p>Everything looks good. I do almost the same.\nI can only suggest you try to load hits with the TrackML library like this</p>\n\n<p>hits, cells, particles, truth = load_event('Data/train_100_events/event000001000')</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 358942,
      "author_name": "Crysis",
      "author_url": "",
      "post_date": "2018-07-19T08:07:02.403000",
      "content": "<p>Thank you very much for sharing your code, it greatly worked for me ! </p>\n\n<p>I re-write it in R for those of you who may be interested.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 341433,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-11T15:24:07.330000",
      "content": "<p>I wonder if any kaggler has time to try this experiment:</p>\n\n<p>1) DBSCAN on point set:</p>\n\n<ul>\n<li><p>make a cone slice.</p></li>\n<li><p>if cone slice = x1,x2,x3,x4 ...x10, it means each data point has 3 dim={x,yz} and there are 10 data points. </p></li>\n<li><p>apply dbscan.</p></li>\n</ul>\n\n<p>2) DBSCAN on pair set:</p>\n\n<ul>\n<li><p>make a cone slice.</p></li>\n<li><p>if cone slice = x1,x2,x3,x4, ... x10. make pairs. one pair has 4 dim: pij ={ midpoint(xi,xj), direction(xi to xj) } = { x,y,z, theta }. there are 10x10=100 pairs.</p></li>\n<li><p>apply dbscan on pairs </p></li>\n<li><p>decode clustered pairs into points. e,g, if a clustered pairs ={p12, p23,p34}, then decoded results={x1,x2,x3,x4}</p></li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 341484,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-11T16:55:35.923000",
          "content": "<p>For such a series of point pairs, to reduce the total input you could try to limit your input pairs to potentially adjacent pairs. For example, a track containing hits x1, x2, x3, x4, ... x10 would only want to consider pairs {x1, x2}, {x2, x3}, etc. </p>\n\n<p>To accomplish this, only include pairs where the pair distance is less than a some threshold. You could also include exclude pairs where the angle with respect to z changes more than some other threshold. Both of these thresholds would probably scale with z. The reason I think these criteria make sense is because adjacent pairs on a track wouldn't be too far apart (very close to origin and on the edge of detectors shouldn't be adjacent on a track), or vary too greatly in angle(pairs of hits won't be on the opposite 'edge' of a cone), but hits tend to have more variance in their angles and occur less closely at more extreme values of z. </p>\n\n<p>The downside of trying to limit to only pairs is that you potentially lose some confidence in your clusters. For the above hypothetical track, {x1, x2} and {x2, x3} would cluster together, but {x1, x3} would also cluster together and would be a midpoint between the two. </p>\n\n<p>Possibly an iterative approach where early on in the iterations, nearby pairs are considered, such as  {x1, x2} and {x2, x3}, and then for and given potential cluster, {x1, x3} would also be evaluated. Eventually a cluster of all combinations could be evaluated for that track, and would hopefully reduce the total amount of pairs evaluated.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340349,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2018-06-09T00:59:16.570000",
      "content": "<p>Wow! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 339412,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2018-06-06T22:58:47.583000",
      "content": "<p>Thank you for this! I have also been trying to extend lines but with a much slower, less efficient algorithm.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 338333,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "2018-06-04T20:56:54.117000",
      "content": "<p>I've been attempting to extend tracks as well, but with a different approach. I haven't submitted anything yet since I've not had any promising results... I either get a very small amount of true positives or I start getting too many false positives. </p>\n\n<p>Any critique on my approach would be appreciated:</p>\n\n<p>Using the dbscan tracks as inputs, I attempt to fit a regression to the tracks and then finding points that are very close to the projected track. Normally this would be very inefficient, so I added in some filter criteria... \n1. Only include hits in a similar 'cone' around the z axis\n2. Only include hits where the z is greater than the max z in the track, or less than the minimum (it seems to be the edges that get missed with the dbscan approach)</p>\n\n<p>My model worked decently when I only tried to find linear tracks, but trying to fit helices has been hard, I'm currently trying to fit with the method by Luis Andre Dutra e Silva...</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 338841,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-06-05T20:42:03.060000",
      "content": "<p>Thanks for that code, I shamelessly reused it ;)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 339567,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-07T06:25:42.807000",
          "content": "<p>The boost I get from it goes down as my score got better.  Now it is about 0.02 (in the current run I'll submit tomorrow).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340038,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-08T07:46:14.350000",
          "content": "<p>@CPMP</p>\n\n<p>the code is a sample code. you can make some visualization and improve on it. I think such post processing  for linking can get LB around 0.64.</p>\n\n<p>there is a bug in the code: discontinuity for 0 and 2pi when comparing angular distance</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340042,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-08T08:05:34.767000",
          "content": "<p>@Heng, I had the same problem (0 and 2*pi), when I had to use the angle and converting it into two features (sin and cos) was not acceptable. I just turned the XY plane by pi, performed all operations once more and then ensembled both results.  Turning the XY plane by pi is very easy: x = -x,  y= -y.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 340043,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-08T08:11:23.923000",
          "content": "<p>@Heng,  I indeed plan to implement something more complex.  I also think one can reach 0.64 with better back fitting.  We'll see if I am right this week end ;)</p>\n\n<p>@Grzegorz, I turn xy by pi/2 in other parts of my code with x = y, y = -x</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 340044,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-08T08:13:52.660000",
          "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>thanks for the suggestion. An alternative is to change a to (cos(a), sin(a)), i.e. unit direction vector. Then angular distance can measured by dot product (i.e. cosine distance)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 340223,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-08T16:47:29.910000",
          "content": "<p>I tried both approaches of @CPMP and @Grzegorz by shifting <code>pi</code> or <code>pi/2</code>, unsurprisingly, they yielded exactly the same local scores since they're meant to solve the same problem. The improvement is only marginal and it took twice as long for clustering since I had to run the same operations on the shifted plane again, so I guess it's not the secret of your current public LB scores, too bad :p</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340234,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-08T17:12:40.377000",
          "content": "<p>@Nicole, I was thinking the same thing driving to work today (small improvement for twice the computation).</p>\n\n<p>However, since you are only trying to avoid lost opportunity from the discontinuity at 0, 2*pi, you don't need to scan the entire angular range again...</p>\n\n<p>Maybe you can get the improvement for free by scanning half the range, rotate by pi and scan again to cover the rest of the range? </p>\n\n<p>BTW - I also want to think about @Heng's comment (always a good idea) to measure angular distance using cosine similarity or some similar technique.  This would naturally eliminate the discontinuity. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340243,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-08T17:53:18.373000",
          "content": "<p>Hey @John, thanks for the idea, that was my first attempt cutting down the steps  by half but I forgot to change the length of angular displacement intervals, and that hurt the accuracy but after having read your comment I tried again and fixed the bug. Thanks! There's almost no improvement in the score though. It's good to see more discussions here, somehow not too many teams have entered this competition. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340302,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-08T20:58:53.697000",
          "content": "<p>What is the issue with working with the sin/cos of the angle? It seems to work effectively for the clustering portion of the challenge. Is it just because the sin/cos of the angle isn't as easy to implement in your track extension framework?</p>\n\n<p>I think the smaller entrant count is a combination of the smaller prize pool, the far away closing date, and the challenge seeming more intimidating at first glance. I imagine several people see physics and step away, even if the competition is mainly based on data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340566,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-09T16:01:44.643000",
          "content": "<p>I used cosine similarity to avoid the 0 vs 2*pi discontinuity.  I used the head and tail points to create directional unit vectors, and then used a dot product to measure the angular distance between unit vectors.  Maybe I did something wrong, but it only helped a very little bit (+0.0002).  </p>\n\n<p>Here is a snippet of the change from @Heng's original code:</p>\n\n<p>`           da0 = a[idx[1]] - a[idx[0]]  #direction\n            dr0 = r[idx[1]] - r[idx[0]]\n            divisor0 = (da0*<em>2+dr0</em>*2)**0.5\n            if divisor0 == 0 : divisor0 = 1\n            direction0 = np.array([da0/divisor0,dr0/divisor0])</p>\n\n<pre><code>        da1 = a[idx[-1]] - a[idx[-2]]\n        dr1 = r[idx[-1]] - r[idx[-2]]\n        divisor1 = (da1**2+dr1**2)**0.5\n        if divisor1 == 0: divisor1 = 1\n        direction1 = np.array([da1/divisor1,dr1/divisor1]) \n\n\n\n        ## extend start point\n        ns = tree.query([[a0,r0]], k=min(20,min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n        da0ns = a0-a[ns]\n        dr0ns = r0-r[ns]\n        divisor0ns = (da0ns**2+dr0ns**2)**0.5\n        divisor0ns[np.where(divisor0ns==0)]=1\n\n        direction = np.array([da0ns/divisor0ns,dr0ns/divisor0ns]) \n        ns = ns[(r0-r[ns]&gt;0.01) &amp;(np.matmul(direction.T,direction0)&gt;0.9991)]\n\n        for n in ns:\n            df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p \n\n        ## extend end point\n        ns = tree.query([[a1,r1]], k=min(20,min_num_neighbours), return_distance=False)\n        ns = np.concatenate(ns)\n        da1ns = a[ns]-a1\n        dr1ns = r[ns]-r1\n        divisor1ns = (da1ns**2+dr1ns**2)**0.5\n        divisor1ns[np.where(divisor1ns==0)]=1\n\n        direction = np.array([da1ns/divisor1ns,dr1ns/divisor1ns]) \n        ns = ns[(r[ns]-r1&gt;0.01) &amp;(np.matmul(direction.T,direction1)&gt;0.9991)] \n\n        for n in ns:\n            df.loc[ df.hit_id==hit_ids[n],'track_id' ] = p`\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 340648,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-09T22:15:42.200000",
          "content": "<p>@John, thanks for sharing your code, highly appreciated. As said I shifted the plane by <code>pi/2</code> as suggested by @CPMP in the clustering code (not just the backfitting code) and the improvement was marginal, around 0.002 for clustering (backfitting too). I forwardfit on the original XY plane and backwardfit on the shifted plane, however I removed the code for now since the gain is marginal. There are only about 1500 hits between -1 and 1 degrees, so I <em>guess</em> that's why the discontinuity doesn't hurt the local score too much. As mentioned by @Heng and @CPMP we can still improve the backfitting code to get an accuracy above 0.6. I'm working on it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340686,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-10T02:36:04.287000",
          "content": "<p>@John, I made a similar modification yesterday, this gave me a 0.005 boost.</p>\n\n<p>The main difference with yours is I replace</p>\n\n<pre><code>    divisor1 = (da1**2+dr1**2)**0.5\n    if divisor1 == 0: divisor1 = 1\n</code></pre>\n\n<p>with</p>\n\n<pre><code>divisor1 = (da1**2+dr1**2)**0.5 + 1e-6\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 340744,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-10T07:35:50.140000",
          "content": "<p>@John, your code seems right to me, I changed the threshold to <code>0.999</code> to find more nearest neighbours, the improvement on all cone slices for event 1000 is around <code>0.002</code> for the event 1000. I guess you put one zero too much.</p>\n\n<p>I scored on the found tracks using this cone slicing approach. </p>\n\n<pre><code>Cone slice score for event 1000: 0.25068574\nCone slice score for event 1000: 0.25284123\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340761,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-10T08:37:12.723000",
          "content": "<p>some notes:</p>\n\n<ol>\n<li><p>create a set of different models. The objective is that the models must detect different tracks. Hence score of each model may not be necessary high.  the only way to tell is to draw out the tracks and visually inspect. </p></li>\n<li><p>then run extend() to extend the tracks. you can use decrease thresholds when you run multiple runs. This detect more curvy tracks.</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 361515,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-07-24T16:23:28.733000",
      "content": "<p>Hello <a href=\"/hengck23\">@hengck23</a> / <a href=\"/crysis\">@crysis</a> ... I'm using the python script for track extension and I have a small doubt.. So, in the extend routine the submission variable takes the whole DataFrame of the submission file and the hits variable takes the hits values - submission[[hit_id]].. Am I correct? Coz even with around 60GB RAM I'm getting a MemoryError in the first line of the code where the merge is done.... Thanks in Adv...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 340777,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-06-10T09:31:14.270000",
      "content": "<p>for i in range(8): \n       submission = extend(submission, hits)</p>\n\n<p>I didn't get it. What‘s the meaning of the i？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 340810,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-10T11:41:05.923000",
          "content": "<p>You can change <code>i</code> to an underscore if it bothers you, but I guess you were not asking a programmatic question. I use <code>i</code> to do different post processing (or extend) and change values passed in each time, e.g. <code>extend(submission, hits, shift_plane=i%2 == 1)</code> and the labels get improved after each post-processing run. If you use @Heng's code directly, <code>8</code> is the optimal value as a trade-off between time and accuracy, at least for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 340833,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-10T12:37:42.257000",
          "content": "<p>It means extend is applied 8 times.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 340473,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-06-09T09:06:26.480000",
      "content": "<p>@Heng, In your code df.arctan2 ranges from -pi/2 to pi/2, but you scan the interval [-pi, pi].  </p>\n\n<p>Not a big deal but I thought I should report it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 340535,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-09T13:22:05.787000",
          "content": "<p>Thanks. I confirm that scanning from -pi/2 to pi/2 is enough</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 338499,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-06-05T07:50:33.453000",
      "content": "<p>Thank you for sharing your approach!</p>\n\n<p>I've not investigated extending tracks yet. I understand why we can lose last hits (due to fluctuations within every hit with a detector). However I don't understand why we lose a first hit.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 338300,
      "author_name": "Michael Maguire",
      "author_url": "",
      "post_date": "2018-06-04T19:21:09.457000",
      "content": "<p>Thank-you for all of the information you have been posting!  Much appreciated!  in your code above you define arctan2 = np.arctan2(df.z, df.r).  Should this be np.arctan2(df.r, df.z) ??</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 338042,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-04T09:58:48.770000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 340577,
      "author_name": "Alice B",
      "author_url": "",
      "post_date": "2018-06-09T17:10:26.590000",
      "content": "<p>Thank you, that's really useful!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "338013": "this is example of track extension. It should gives improvement of LB +0.04 What you should do:\n\n1. use this your extension submission track_id. you can run it over a few times (submission and hits are the dataframe of csv files):\n\n        for i in range(8): \n               submission = extend(submission, hits)\n\n2. tune the parameters to get better results\n\n3. make some visualizations. It is easy to debug, fault find and further improve results.\n\n---\nif you improve it by multi-threading, code efficiency, etc ... please share your improved version back here. thanks!\n\nhere is the code\n\n    def extend(submission,hits):\n\n\t\tdf = submission.merge(hits,  on=['hit_id'], how='left')\n\t\tdf = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n\t\tdf = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n\t\tdf = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n\t\t... see extension.py as attached ... \n\n\nexample results\n\npink: extended tracks\n\nred: submission tracks\n\nblack: truth tracks\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/338013/9565/after.png",
    "340763": "my updated code\n\n\n    def extend(submission,hits,limit=0.04, num_neighbours=18):\n\n    df = submission.merge(hits,  on=['hit_id'], how='left')\n    df = df.assign(d = np.sqrt( df.x**2 + df.y**2 + df.z**2 ))\n    df = df.assign(r = np.sqrt( df.x**2 + df.y**2))\n    df = df.assign(arctan2 = np.arctan2(df.z, df.r))\n\n    for angle in range(-90,90,1):\n\n        print ('\\r %f'%angle, end='',flush=True)\n        #df1 = df.loc[(df.arctan2&gt;(angle-0.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+0.5)/180*np.pi)]\n        df1 = df.loc[(df.arctan2&gt;(angle-1.5)/180*np.pi) &amp; (df.arctan2&lt;(angle+1.5)/180*np.pi)]\n\n        min_num_neighbours = len(df1)\n        if min_num_neighbours&lt;3: continue\n\n        hit_ids = df1.hit_id.values\n        x,y,z = df1[['x', 'y', 'z']].values.T\n        r  = (x**2 + y**2)**0.5\n        r  = r/1000\n        a  = np.arctan2(y,x)\n        c = np.cos(a)\n        s = np.sin(a)\n        #tree = KDTree(np.column_stack([a,r]), metric='euclidean')\n        tree = KDTree(np.column_stack([c, s, r]), metric='euclidean')\n\n\n        track_ids = list(df1.track_id.unique())\n        num_track_ids = len(track_ids)\n        min_length=3\n\n        for i in range(num_track_ids):\n            p = track_ids[i]\n            if p==0: continue\n\n            idx = np.where(df1.track_id==p)[0]\n            if len(idx)",
    "358942": "Thank you very much for sharing your code, it greatly worked for me ! \n\nI re-write it in R for those of you who may be interested.",
    "341433": "I wonder if any kaggler has time to try this experiment:\n\n1) DBSCAN on point set:\n\n - make a cone slice.\n\n -  if cone slice = x1,x2,x3,x4 ...x10, it means each data point has 3 dim={x,yz} and there are 10 data points. \n\n - apply dbscan.\n\n2) DBSCAN on pair set:\n\n - make a cone slice.\n\n -  if cone slice = x1,x2,x3,x4, ... x10. make pairs. one pair has 4 dim: pij ={ midpoint(xi,xj), direction(xi to xj) } = { x,y,z, theta }. there are 10x10=100 pairs.\n\n -  apply dbscan on pairs \n\n - decode clustered pairs into points. e,g, if a clustered pairs ={p12, p23,p34}, then decoded results={x1,x2,x3,x4}",
    "340349": "Wow! ",
    "339412": "Thank you for this! I have also been trying to extend lines but with a much slower, less efficient algorithm.",
    "338333": "I've been attempting to extend tracks as well, but with a different approach. I haven't submitted anything yet since I've not had any promising results... I either get a very small amount of true positives or I start getting too many false positives. \n\nAny critique on my approach would be appreciated:\n\nUsing the dbscan tracks as inputs, I attempt to fit a regression to the tracks and then finding points that are very close to the projected track. Normally this would be very inefficient, so I added in some filter criteria... \n1. Only include hits in a similar 'cone' around the z axis\n2. Only include hits where the z is greater than the max z in the track, or less than the minimum (it seems to be the edges that get missed with the dbscan approach)\n\nMy model worked decently when I only tried to find linear tracks, but trying to fit helices has been hard, I'm currently trying to fit with the method by Luis Andre Dutra e Silva...",
    "338841": "Thanks for that code, I shamelessly reused it ;)",
    "361515": "Hello @hengck23 / @crysis ... I'm using the python script for track extension and I have a small doubt.. So, in the extend routine the submission variable takes the whole DataFrame of the submission file and the hits variable takes the hits values - submission[[hit_id]].. Am I correct? Coz even with around 60GB RAM I'm getting a MemoryError in the first line of the code where the merge is done.... Thanks in Adv...",
    "340777": "for i in range(8): \n       submission = extend(submission, hits)\n\n\nI didn't get it. What‘s the meaning of the i？",
    "340473": "@Heng, In your code df.arctan2 ranges from -pi/2 to pi/2, but you scan the interval [-pi, pi].  \n\nNot a big deal but I thought I should report it.",
    "338499": "Thank you for sharing your approach!\n\nI've not investigated extending tracks yet. I understand why we can lose last hits (due to fluctuations within every hit with a detector). However I don't understand why we lose a first hit.",
    "338300": "Thank-you for all of the information you have been posting!  Much appreciated!  in your code above you define arctan2 = np.arctan2(df.z, df.r).  Should this be np.arctan2(df.r, df.z) ??",
    "338042": "",
    "340577": "Thank you, that's really useful!"
  }
}