{
  "id": 63256,
  "title": "2nd place solution",
  "url": "/competitions/trackml-particle-identification/discussion/63256",
  "author_name": "outrunner",
  "post_date": "2018-08-14T03:03:16.114000",
  "votes": 47,
  "comment_count": 26,
  "views": 0,
  "content": "<p>In the beginning, I want to build a model which can input all hits and output all tracks. But after simple calculation, it could not be done. So I split it to minimum unit: input two hits. output 1 if two hits are in the same track, 0 otherwise.</p>\n\n<p>The difference with most other DL approaches is that they only do \"connect the dots\", if some dots lost, the connection break.</p>\n\n<p><strong>I connect all the dots</strong>. <a href=\"https://www.kaggle.com/outrunner/trackml-2-solution-example\">here is a example of kernel (update 08/16)</a></p>\n\n<p>In my real case, the difference is: (read the kernel in detail)</p>\n\n<ul>\n<li>input size: 27 (x, y, z etc. and use cells to get hit's direction)</li>\n<li>model size: 5 hidden layers with 4k-2k-2k-2k-1k neurons</li>\n</ul>\n\n<p>The well trained model can get 0.8 by only use the predictions to reconstruct tracks, just like the kernel. Add simple curve fitting (I use scipy.optimize.leastsq to fit circle in xy plane) can get 0.9, and add z-axis constrain (dr/dz) improve 0.003 in the end. I don't spend much time on curve fitting since I think CERN do it better, and someone can get much improvement from better curve fitting.</p>\n\n<ul>\n<li>attached is a prediction of event1001 and you may give it a try.</li>\n<li>fig 01 shows the seed(large circle) and it's corresponding candidates(the same color)</li>\n<li>fig 02 shows the sum of predict prob. of hits in direct ratio to diameter\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10058/TrackML_01.png\" alt=\"enter image description here\"></li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10059/TrackML_02.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": 369970,
      "postDate": "2018-08-14T03:03:16.113Z",
      "content": "<p>In the beginning, I want to build a model which can input all hits and output all tracks. But after simple calculation, it could not be done. So I split it to minimum unit: input two hits. output 1 if two hits are in the same track, 0 otherwise.</p>\n\n<p>The difference with most other DL approaches is that they only do \"connect the dots\", if some dots lost, the connection break.</p>\n\n<p><strong>I connect all the dots</strong>. <a href=\"https://www.kaggle.com/outrunner/trackml-2-solution-example\">here is a example of kernel (update 08/16)</a></p>\n\n<p>In my real case, the difference is: (read the kernel in detail)</p>\n\n<ul>\n<li>input size: 27 (x, y, z etc. and use cells to get hit's direction)</li>\n<li>model size: 5 hidden layers with 4k-2k-2k-2k-1k neurons</li>\n</ul>\n\n<p>The well trained model can get 0.8 by only use the predictions to reconstruct tracks, just like the kernel. Add simple curve fitting (I use scipy.optimize.leastsq to fit circle in xy plane) can get 0.9, and add z-axis constrain (dr/dz) improve 0.003 in the end. I don't spend much time on curve fitting since I think CERN do it better, and someone can get much improvement from better curve fitting.</p>\n\n<ul>\n<li>attached is a prediction of event1001 and you may give it a try.</li>\n<li>fig 01 shows the seed(large circle) and it's corresponding candidates(the same color)</li>\n<li>fig 02 shows the sum of predict prob. of hits in direct ratio to diameter\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10058/TrackML_01.png\" alt=\"enter image description here\"></li>\n</ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10059/TrackML_02.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "In the beginning, I want to build a model which can input all hits and output all tracks. But after simple calculation, it could not be done. So I split it to minimum unit: input two hits. output 1 if two hits are in the same track, 0 otherwise.\n\n\nThe difference with most other DL approaches is that they only do \"connect the dots\", if some dots lost, the connection break.\n\n\n**I connect all the dots**. [here is a example of kernel (update 08/16)][1]\n\nIn my real case, the difference is: (read the kernel in detail)\n\n - input size: 27 (x, y, z etc. and use cells to get hit's direction)\n - model size: 5 hidden layers with 4k-2k-2k-2k-1k neurons\n\n\nThe well trained model can get 0.8 by only use the predictions to reconstruct tracks, just like the kernel. Add simple curve fitting (I use scipy.optimize.leastsq to fit circle in xy plane) can get 0.9, and add z-axis constrain (dr/dz) improve 0.003 in the end. I don't spend much time on curve fitting since I think CERN do it better, and someone can get much improvement from better curve fitting.\n\n - attached is a prediction of event1001 and you may give it a try.\n - fig 01 shows the seed(large circle) and it's corresponding candidates(the same color)\n - fig 02 shows the sum of predict prob. of hits in direct ratio to diameter\n![enter image description here][2]\n\n![enter image description here][3]\n\n\n\n  [1]: https://www.kaggle.com/outrunner/trackml-2-solution-example\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10058/TrackML_01.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10059/TrackML_02.png",
      "votes": 47
    },
    {
      "id": 370216,
      "postDate": "2018-08-14T12:57:30.187Z",
      "content": "<p>Congratulations outrunner on a great solution! It looks like I was right to be worried right up to the last minute, given how much potential your method has. It looks like you spent most of your time on candidates, while I spent my time extension / selection. I can only imagine how well a solution using your near-optimal pair scoring and my track extension / selection could score.</p>",
      "rawMarkdown": "Congratulations outrunner on a great solution! It looks like I was right to be worried right up to the last minute, given how much potential your method has. It looks like you spent most of your time on candidates, while I spent my time extension / selection. I can only imagine how well a solution using your near-optimal pair scoring and my track extension / selection could score.",
      "votes": 5,
      "replies": [
        {
          "id": 370235,
          "postDate": "2018-08-14T13:31:20.710Z",
          "content": "<p>Congratulations and thank you for open the gap so I can sleep well in the last two days. :) You did a great job! </p>",
          "rawMarkdown": "Congratulations and thank you for open the gap so I can sleep well in the last two days. :) You did a great job! ",
          "votes": 4
        }
      ]
    },
    {
      "id": 370113,
      "postDate": "2018-08-14T09:52:41.327Z",
      "content": "<p>Well done, congrats on the result, even if you lost 1st place in the last few days.  I hope you're not too disappointed by this. </p>\n\n<p>Your sharing of your good results throughout the competition was very motivating, thanks for that.</p>\n\n<p>Since the start of the competition I was thinking that it would be silly to use machine learning to discover the laws of physics, but your solution proves it is not silly at all!  To me it looks like your approach and icuber/ersol approach are similar, except you seed your tracks with a neural net whereas they use another way.  Is your way of building tracks from likely hit pairs significantly different from icecuber's?  </p>",
      "rawMarkdown": "Well done, congrats on the result, even if you lost 1st place in the last few days.  I hope you're not too disappointed by this. \n\nYour sharing of your good results throughout the competition was very motivating, thanks for that.\n\nSince the start of the competition I was thinking that it would be silly to use machine learning to discover the laws of physics, but your solution proves it is not silly at all!  To me it looks like your approach and icuber/ersol approach are similar, except you seed your tracks with a neural net whereas they use another way.  Is your way of building tracks from likely hit pairs significantly different from icecuber's?  ",
      "votes": 1,
      "replies": [
        {
          "id": 370172,
          "postDate": "2018-08-14T11:46:23.800Z",
          "content": "<p>In most kaggle competition I think I will win, so it is fine. :) I will explain my approach in detail as below.</p>",
          "rawMarkdown": "In most kaggle competition I think I will win, so it is fine. :) I will explain my approach in detail as below.",
          "votes": 1
        },
        {
          "id": 370189,
          "postDate": "2018-08-14T12:19:07.243Z",
          "content": "<p>For a event with N hits, I predict NxN probability matrix. For example, a event with 5 hits, the prediction is:</p>\n\n<pre><code>    h1   h2   h3   h4   h5\nh1   -   0.8  0.2  0.9  0.4\nh2  0.8   -   0.5  0.7  0.7\nh3  0.2  0.5   -   0.3  0.4\nh4  0.9  0.7  0.3   -   0.4\nh5  0.4  0.7  0.4  0.4   -\n</code></pre>\n\n<p>let us assume the threshold is 0.65, so:</p>\n\n<pre><code>    h1   h2   h3   h4   h5\nh1   -   0.8  0.   0.9  0. \nh2  0.8   -   0.   0.7  0.7\nh3  0.   0.    -   0.   0. \nh4  0.9  0.7  0.    -   0. \nh5  0.   0.7  0.   0.    -\n</code></pre>\n\n<p>pick <strong>h1</strong> as seed, then <strong>h4</strong> is the next most likely hit, and p(h1,h4)=0.9&gt;0.65, so let's go on:</p>\n\n<pre><code>        h1   h2   h3   h4   h5\n    h1   -   0.8  0.   0.9  0. \n    h4  0.9  0.7  0.    -   0. \n   ---------------------------\n         -   1.5  0.    -   0. \n</code></pre>\n\n<p>0.8 and 0.7 are all large than threshold, so the next hit is <strong>h2</strong>, then:</p>\n\n<pre><code>        h1   h2   h3   h4   h5\n    h1   -   0.8  0.   0.9  0. \n    h4  0.9  0.7  0.    -   0. \n    h2  0.8   -   0.   0.7  0.7\n</code></pre>\n\n<p><strong>h3</strong> and <strong>h5</strong> are not qualify (all prod. large than threshold), so we stop here.</p>\n\n<p>And the track we find is <strong>h1-h2-h4</strong></p>\n\n<p>The next seed is h2 and so on, I reconstruct N tracks by N hits in one event.</p>",
          "rawMarkdown": "For a event with N hits, I predict NxN probability matrix. For example, a event with 5 hits, the prediction is:\n\n        h1   h2   h3   h4   h5\n    h1   -   0.8  0.2  0.9  0.4\n    h2  0.8   -   0.5  0.7  0.7\n    h3  0.2  0.5   -   0.3  0.4\n    h4  0.9  0.7  0.3   -   0.4\n    h5  0.4  0.7  0.4  0.4   -\n\nlet us assume the threshold is 0.65, so:\n\n        h1   h2   h3   h4   h5\n    h1   -   0.8  0.   0.9  0. \n    h2  0.8   -   0.   0.7  0.7\n    h3  0.   0.    -   0.   0. \n    h4  0.9  0.7  0.    -   0. \n    h5  0.   0.7  0.   0.    -\n\npick **h1** as seed, then **h4** is the next most likely hit, and p(h1,h4)=0.9&gt;0.65, so let's go on:\n\n            h1   h2   h3   h4   h5\n        h1   -   0.8  0.   0.9  0. \n        h4  0.9  0.7  0.    -   0. \n       ---------------------------\n             -   1.5  0.    -   0. \n\n0.8 and 0.7 are all large than threshold, so the next hit is **h2**, then:\n\n            h1   h2   h3   h4   h5\n        h1   -   0.8  0.   0.9  0. \n        h4  0.9  0.7  0.    -   0. \n        h2  0.8   -   0.   0.7  0.7\n\n**h3** and **h5** are not qualify (all prod. large than threshold), so we stop here.\n\nAnd the track we find is **h1-h2-h4**\n\nThe next seed is h2 and so on, I reconstruct N tracks by N hits in one event.\n",
          "votes": 3
        },
        {
          "id": 370197,
          "postDate": "2018-08-14T12:25:39.207Z",
          "content": "<p>Thank you, makes lots of sense.  Very nice way to use supervised ML here.  I'll look at your code to see how you handle a 100k wide objective!</p>",
          "rawMarkdown": "Thank you, makes lots of sense.  Very nice way to use supervised ML here.  I'll look at your code to see how you handle a 100k wide objective!"
        },
        {
          "id": 370204,
          "postDate": "2018-08-14T12:33:00.030Z",
          "content": "<p>as the example, all the tracks are:</p>\n\n<ul>\n<li>1-2-4</li>\n<li>2-1-4</li>\n<li>3</li>\n<li>4-1-2</li>\n<li>5-2</li>\n</ul>\n\n<p>then I calculate the similarity of all tracks as merge priority, so the final submission is:</p>\n\n<pre><code>hit_id,track_id\n1,1\n2,1\n3,0\n4,1\n5,0\n</code></pre>",
          "rawMarkdown": "as the example, all the tracks are:\n\n - 1-2-4\n - 2-1-4\n - 3\n - 4-1-2\n - 5-2\n\nthen I calculate the similarity of all tracks as merge priority, so the final submission is:\n\n    hit_id,track_id\n    1,1\n    2,1\n    3,0\n    4,1\n    5,0\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 370314,
      "postDate": "2018-08-14T15:50:13.420Z",
      "content": "<p>Thanks for your great solution.</p>",
      "rawMarkdown": "Thanks for your great solution.",
      "votes": 1,
      "replies": [
        {
          "id": 375544,
          "postDate": "2018-08-25T12:24:49.140Z",
          "content": "<p><a href=\"/bestfitting\">@bestfitting</a></p>\n\n<p>would you like to share your solution? I am interested to know how you can come up with top solution with only few submissions.</p>",
          "rawMarkdown": "@bestfitting\n \nwould you like to share your solution? I am interested to know how you can come up with top solution with only few submissions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 458122,
      "postDate": "2019-01-18T20:37:11.237Z",
      "content": "<p>Great solution ! I was wondering if you had posted the full code anywhere with the full sized network ?</p>",
      "rawMarkdown": "Great solution ! I was wondering if you had posted the full code anywhere with the full sized network ?"
    },
    {
      "id": 441566,
      "postDate": "2018-12-18T19:43:06.583Z",
      "content": "<p>Hello <a href=\"/outrunner\">@outrunner</a>,  I was reading into your solution and in the reconstruction of the tracks, in the step 3 you mention \"Test the new hit to see whether it fits the circle in x-y plane by existing hits after the track has two or three hits. (Without this step I can only get to an 0.8 score.)\" and I was wondering if you could clear my confusion on this.</p>\n\n<p>From my understanding, this hit 'k' has already been added to the track, so I don't quite understand why you then test it. Do you remove the hit 'k' from the track if it doesn't fit on the circle from the existing points in the track ? And is there some sort of threshold in place to decide what classifies as being fitted to the circle?</p>\n\n<p>Thank you for the solution and any explanation with this would be much appreciated!</p>",
      "rawMarkdown": "Hello @outrunner,  I was reading into your solution and in the reconstruction of the tracks, in the step 3 you mention \"Test the new hit to see whether it fits the circle in x-y plane by existing hits after the track has two or three hits. (Without this step I can only get to an 0.8 score.)\" and I was wondering if you could clear my confusion on this.\n\nFrom my understanding, this hit 'k' has already been added to the track, so I don't quite understand why you then test it. Do you remove the hit 'k' from the track if it doesn't fit on the circle from the existing points in the track ? And is there some sort of threshold in place to decide what classifies as being fitted to the circle?\n\nThank you for the solution and any explanation with this would be much appreciated!"
    },
    {
      "id": 386891,
      "postDate": "2018-09-13T20:47:37.307Z",
      "content": "<p>Second, \"Throughput\" phase is online, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/65525\">this post</a></p>",
      "rawMarkdown": "Second, \"Throughput\" phase is online, see [this post][1]\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/65525"
    },
    {
      "id": 382316,
      "postDate": "2018-09-06T05:30:48.783Z",
      "content": "<p>Very cool. I assumed pairwise hit comparison would be too sparse but this is fantastic. I added your dense network to my approach and it's looking a lot better.  Crossing my fingers that the red and blue distributions separate with further training.  :)</p>",
      "rawMarkdown": "Very cool. I assumed pairwise hit comparison would be too sparse but this is fantastic. I added your dense network to my approach and it's looking a lot better.  Crossing my fingers that the red and blue distributions separate with further training.  :)"
    },
    {
      "id": 375536,
      "postDate": "2018-08-25T11:50:40.920Z",
      "content": "<p><a href=\"/outrunner\">@outrunner</a> </p>\n\n<p>Congrats Nice application of deep learning method! Any plan to speed up your method and apply for the speed phase challenge as well?</p>",
      "rawMarkdown": "@outrunner \n\nCongrats Nice application of deep learning method! Any plan to speed up your method and apply for the speed phase challenge as well?",
      "replies": [
        {
          "id": 375565,
          "postDate": "2018-08-25T13:48:34.977Z",
          "content": "<p>I have no plan to join the second phase, thanks.</p>",
          "rawMarkdown": "I have no plan to join the second phase, thanks."
        }
      ]
    },
    {
      "id": 370027,
      "postDate": "2018-08-14T05:33:10.787Z",
      "content": "<p>Congratulations <a href=\"/outrunner\">@outrunner</a> on coming 2nd and sharing your solution. Thanks also for pushing the rest of us to believe scores above 0.3 and beyond is possible.</p>\n\n<p>I also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.</p>",
      "rawMarkdown": "Congratulations @outrunner on coming 2nd and sharing your solution. Thanks also for pushing the rest of us to believe scores above 0.3 and beyond is possible.\n\nI also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition."
    },
    {
      "id": 370816,
      "postDate": "2018-08-15T13:43:59.653Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 370885,
          "postDate": "2018-08-15T15:50:22.643Z",
          "content": "<p>It is trail-and-error. In the beginning, I do some experiments to make sure it can do some kind of curve fitting in proper precision. Then I start training with 4 layers, and it never overfit even when I add more neurons and one layer. Every time I extend the model size, the accuracy improve a little. I think there is a way to make the model more efficient but need time to do more experiments.</p>",
          "rawMarkdown": "It is trail-and-error. In the beginning, I do some experiments to make sure it can do some kind of curve fitting in proper precision. Then I start training with 4 layers, and it never overfit even when I add more neurons and one layer. Every time I extend the model size, the accuracy improve a little. I think there is a way to make the model more efficient but need time to do more experiments.",
          "votes": 1
        },
        {
          "id": 370938,
          "postDate": "2018-08-15T17:42:43.197Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 371110,
          "postDate": "2018-08-16T02:22:22.567Z",
          "content": "<p>You are welcome.</p>\n\n<ol>\n<li>It would be sure</li>\n<li>selu. About input, 5 as the kernel [x, y, z, count(cells), sum(cells.value)], and two unit vector come from cells to estimate the hit's direction, random reverse when training. Now we have (5+6)x2 = 22. The rest 4 are assumed that the two hits are linear or helix with (0,0,z0), calculate the abs(cos()) with previous two estimated unit vector, and the last is z0.</li>\n<li>I fixed it, thanks.</li>\n</ol>",
          "rawMarkdown": "You are welcome.\n\n 1. It would be sure\n 2. selu. About input, 5 as the kernel [x, y, z, count(cells), sum(cells.value)], and two unit vector come from cells to estimate the hit's direction, random reverse when training. Now we have (5+6)x2 = 22. The rest 4 are assumed that the two hits are linear or helix with (0,0,z0), calculate the abs(cos()) with previous two estimated unit vector, and the last is z0.\n 3. I fixed it, thanks.\n"
        },
        {
          "id": 381015,
          "postDate": "2018-09-03T23:00:24.113Z",
          "content": "<p><a href=\"/outrunner\">@outrunner</a></p>\n\n<p>Thank you for sharing your solution. Would you please elaborate further on \"two unit vector come from cells to estimate the hit's direction, random reverse when training\" part? </p>\n\n<p>Thanks in advance.</p>",
          "rawMarkdown": "@outrunner\n\nThank you for sharing your solution. Would you please elaborate further on \"two unit vector come from cells to estimate the hit's direction, random reverse when training\" part? \n\nThanks in advance."
        },
        {
          "id": 381045,
          "postDate": "2018-09-04T01:33:38.407Z",
          "content": "<p><a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">reference this kernel</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10251/TrackML_Fig2.png\" alt=\"enter image description here\"></p>\n\n<p>I assumed that the particle crosses the center of two end point cells.</p>",
          "rawMarkdown": "[reference this kernel][2]\n\n![enter image description here][1]\n\nI assumed that the particle crosses the center of two end point cells.\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10251/TrackML_Fig2.png\n  [2]: https://www.kaggle.com/asalzburger/pixel-detector-cells"
        },
        {
          "id": 382067,
          "postDate": "2018-09-05T16:30:11.773Z",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>\nThanks :)</p>",
          "rawMarkdown": "@outrunner\nThanks :)"
        },
        {
          "id": 386125,
          "postDate": "2018-09-12T07:36:34.417Z",
          "content": "<p>Hi <a href=\"/outrunner\">@outrunner</a>, might be a stupid question, could you explain also a bit how do you compute z0?</p>",
          "rawMarkdown": "Hi @outrunner, might be a stupid question, could you explain also a bit how do you compute z0?"
        },
        {
          "id": 386473,
          "postDate": "2018-09-13T01:18:53.070Z",
          "content": "<p>find the circle on x-y plane by two hits and (0,0), then find the delta z by arc length.</p>",
          "rawMarkdown": "find the circle on x-y plane by two hits and (0,0), then find the delta z by arc length.",
          "votes": 1
        },
        {
          "id": 386607,
          "postDate": "2018-09-13T08:14:49.337Z",
          "content": "<p>I see! many thanks! </p>",
          "rawMarkdown": "I see! many thanks! "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 370216,
      "author_name": "icecuber",
      "author_url": "",
      "post_date": "2018-08-14T12:57:30.187000",
      "content": "<p>Congratulations outrunner on a great solution! It looks like I was right to be worried right up to the last minute, given how much potential your method has. It looks like you spent most of your time on candidates, while I spent my time extension / selection. I can only imagine how well a solution using your near-optimal pair scoring and my track extension / selection could score.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 370235,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-14T13:31:20.710000",
          "content": "<p>Congratulations and thank you for open the gap so I can sleep well in the last two days. :) You did a great job! </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 370113,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-14T09:52:41.327000",
      "content": "<p>Well done, congrats on the result, even if you lost 1st place in the last few days.  I hope you're not too disappointed by this. </p>\n\n<p>Your sharing of your good results throughout the competition was very motivating, thanks for that.</p>\n\n<p>Since the start of the competition I was thinking that it would be silly to use machine learning to discover the laws of physics, but your solution proves it is not silly at all!  To me it looks like your approach and icuber/ersol approach are similar, except you seed your tracks with a neural net whereas they use another way.  Is your way of building tracks from likely hit pairs significantly different from icecuber's?  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 370172,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-14T11:46:23.800000",
          "content": "<p>In most kaggle competition I think I will win, so it is fine. :) I will explain my approach in detail as below.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 370189,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-14T12:19:07.243000",
          "content": "<p>For a event with N hits, I predict NxN probability matrix. For example, a event with 5 hits, the prediction is:</p>\n\n<pre><code>    h1   h2   h3   h4   h5\nh1   -   0.8  0.2  0.9  0.4\nh2  0.8   -   0.5  0.7  0.7\nh3  0.2  0.5   -   0.3  0.4\nh4  0.9  0.7  0.3   -   0.4\nh5  0.4  0.7  0.4  0.4   -\n</code></pre>\n\n<p>let us assume the threshold is 0.65, so:</p>\n\n<pre><code>    h1   h2   h3   h4   h5\nh1   -   0.8  0.   0.9  0. \nh2  0.8   -   0.   0.7  0.7\nh3  0.   0.    -   0.   0. \nh4  0.9  0.7  0.    -   0. \nh5  0.   0.7  0.   0.    -\n</code></pre>\n\n<p>pick <strong>h1</strong> as seed, then <strong>h4</strong> is the next most likely hit, and p(h1,h4)=0.9&gt;0.65, so let's go on:</p>\n\n<pre><code>        h1   h2   h3   h4   h5\n    h1   -   0.8  0.   0.9  0. \n    h4  0.9  0.7  0.    -   0. \n   ---------------------------\n         -   1.5  0.    -   0. \n</code></pre>\n\n<p>0.8 and 0.7 are all large than threshold, so the next hit is <strong>h2</strong>, then:</p>\n\n<pre><code>        h1   h2   h3   h4   h5\n    h1   -   0.8  0.   0.9  0. \n    h4  0.9  0.7  0.    -   0. \n    h2  0.8   -   0.   0.7  0.7\n</code></pre>\n\n<p><strong>h3</strong> and <strong>h5</strong> are not qualify (all prod. large than threshold), so we stop here.</p>\n\n<p>And the track we find is <strong>h1-h2-h4</strong></p>\n\n<p>The next seed is h2 and so on, I reconstruct N tracks by N hits in one event.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 370197,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-14T12:25:39.207000",
          "content": "<p>Thank you, makes lots of sense.  Very nice way to use supervised ML here.  I'll look at your code to see how you handle a 100k wide objective!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 370204,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-14T12:33:00.030000",
          "content": "<p>as the example, all the tracks are:</p>\n\n<ul>\n<li>1-2-4</li>\n<li>2-1-4</li>\n<li>3</li>\n<li>4-1-2</li>\n<li>5-2</li>\n</ul>\n\n<p>then I calculate the similarity of all tracks as merge priority, so the final submission is:</p>\n\n<pre><code>hit_id,track_id\n1,1\n2,1\n3,0\n4,1\n5,0\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 370314,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "2018-08-14T15:50:13.420000",
      "content": "<p>Thanks for your great solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 375544,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-08-25T12:24:49.140000",
          "content": "<p><a href=\"/bestfitting\">@bestfitting</a></p>\n\n<p>would you like to share your solution? I am interested to know how you can come up with top solution with only few submissions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 458122,
      "author_name": "Ewan",
      "author_url": "",
      "post_date": "2019-01-18T20:37:11.237000",
      "content": "<p>Great solution ! I was wondering if you had posted the full code anywhere with the full sized network ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441566,
      "author_name": "Colm",
      "author_url": "",
      "post_date": "2018-12-18T19:43:06.583000",
      "content": "<p>Hello <a href=\"/outrunner\">@outrunner</a>,  I was reading into your solution and in the reconstruction of the tracks, in the step 3 you mention \"Test the new hit to see whether it fits the circle in x-y plane by existing hits after the track has two or three hits. (Without this step I can only get to an 0.8 score.)\" and I was wondering if you could clear my confusion on this.</p>\n\n<p>From my understanding, this hit 'k' has already been added to the track, so I don't quite understand why you then test it. Do you remove the hit 'k' from the track if it doesn't fit on the circle from the existing points in the track ? And is there some sort of threshold in place to decide what classifies as being fitted to the circle?</p>\n\n<p>Thank you for the solution and any explanation with this would be much appreciated!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 386891,
      "author_name": "David Rousseau",
      "author_url": "",
      "post_date": "2018-09-13T20:47:37.307000",
      "content": "<p>Second, \"Throughput\" phase is online, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/65525\">this post</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 382316,
      "author_name": "Physicsman",
      "author_url": "",
      "post_date": "2018-09-06T05:30:48.783000",
      "content": "<p>Very cool. I assumed pairwise hit comparison would be too sparse but this is fantastic. I added your dense network to my approach and it's looking a lot better.  Crossing my fingers that the red and blue distributions separate with further training.  :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 375536,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-08-25T11:50:40.920000",
      "content": "<p><a href=\"/outrunner\">@outrunner</a> </p>\n\n<p>Congrats Nice application of deep learning method! Any plan to speed up your method and apply for the speed phase challenge as well?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 375565,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-25T13:48:34.977000",
          "content": "<p>I have no plan to join the second phase, thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 370027,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-08-14T05:33:10.787000",
      "content": "<p>Congratulations <a href=\"/outrunner\">@outrunner</a> on coming 2nd and sharing your solution. Thanks also for pushing the rest of us to believe scores above 0.3 and beyond is possible.</p>\n\n<p>I also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 370816,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-15T13:43:59.653000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 370885,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-15T15:50:22.643000",
          "content": "<p>It is trail-and-error. In the beginning, I do some experiments to make sure it can do some kind of curve fitting in proper precision. Then I start training with 4 layers, and it never overfit even when I add more neurons and one layer. Every time I extend the model size, the accuracy improve a little. I think there is a way to make the model more efficient but need time to do more experiments.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 370938,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-15T17:42:43.197000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371110,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-08-16T02:22:22.567000",
          "content": "<p>You are welcome.</p>\n\n<ol>\n<li>It would be sure</li>\n<li>selu. About input, 5 as the kernel [x, y, z, count(cells), sum(cells.value)], and two unit vector come from cells to estimate the hit's direction, random reverse when training. Now we have (5+6)x2 = 22. The rest 4 are assumed that the two hits are linear or helix with (0,0,z0), calculate the abs(cos()) with previous two estimated unit vector, and the last is z0.</li>\n<li>I fixed it, thanks.</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381015,
          "author_name": "bhartha",
          "author_url": "",
          "post_date": "2018-09-03T23:00:24.113000",
          "content": "<p><a href=\"/outrunner\">@outrunner</a></p>\n\n<p>Thank you for sharing your solution. Would you please elaborate further on \"two unit vector come from cells to estimate the hit's direction, random reverse when training\" part? </p>\n\n<p>Thanks in advance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381045,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-09-04T01:33:38.407000",
          "content": "<p><a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">reference this kernel</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10251/TrackML_Fig2.png\" alt=\"enter image description here\"></p>\n\n<p>I assumed that the particle crosses the center of two end point cells.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382067,
          "author_name": "bhartha",
          "author_url": "",
          "post_date": "2018-09-05T16:30:11.773000",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>\nThanks :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 386125,
          "author_name": "Heisenberg",
          "author_url": "",
          "post_date": "2018-09-12T07:36:34.417000",
          "content": "<p>Hi <a href=\"/outrunner\">@outrunner</a>, might be a stupid question, could you explain also a bit how do you compute z0?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 386473,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-09-13T01:18:53.070000",
          "content": "<p>find the circle on x-y plane by two hits and (0,0), then find the delta z by arc length.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 386607,
          "author_name": "Heisenberg",
          "author_url": "",
          "post_date": "2018-09-13T08:14:49.337000",
          "content": "<p>I see! many thanks! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "369970": "In the beginning, I want to build a model which can input all hits and output all tracks. But after simple calculation, it could not be done. So I split it to minimum unit: input two hits. output 1 if two hits are in the same track, 0 otherwise.\n\n\nThe difference with most other DL approaches is that they only do \"connect the dots\", if some dots lost, the connection break.\n\n\n**I connect all the dots**. [here is a example of kernel (update 08/16)][1]\n\nIn my real case, the difference is: (read the kernel in detail)\n\n - input size: 27 (x, y, z etc. and use cells to get hit's direction)\n - model size: 5 hidden layers with 4k-2k-2k-2k-1k neurons\n\n\nThe well trained model can get 0.8 by only use the predictions to reconstruct tracks, just like the kernel. Add simple curve fitting (I use scipy.optimize.leastsq to fit circle in xy plane) can get 0.9, and add z-axis constrain (dr/dz) improve 0.003 in the end. I don't spend much time on curve fitting since I think CERN do it better, and someone can get much improvement from better curve fitting.\n\n - attached is a prediction of event1001 and you may give it a try.\n - fig 01 shows the seed(large circle) and it's corresponding candidates(the same color)\n - fig 02 shows the sum of predict prob. of hits in direct ratio to diameter\n![enter image description here][2]\n\n![enter image description here][3]\n\n\n\n  [1]: https://www.kaggle.com/outrunner/trackml-2-solution-example\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10058/TrackML_01.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/369970/10059/TrackML_02.png",
    "370216": "Congratulations outrunner on a great solution! It looks like I was right to be worried right up to the last minute, given how much potential your method has. It looks like you spent most of your time on candidates, while I spent my time extension / selection. I can only imagine how well a solution using your near-optimal pair scoring and my track extension / selection could score.",
    "370113": "Well done, congrats on the result, even if you lost 1st place in the last few days.  I hope you're not too disappointed by this. \n\nYour sharing of your good results throughout the competition was very motivating, thanks for that.\n\nSince the start of the competition I was thinking that it would be silly to use machine learning to discover the laws of physics, but your solution proves it is not silly at all!  To me it looks like your approach and icuber/ersol approach are similar, except you seed your tracks with a neural net whereas they use another way.  Is your way of building tracks from likely hit pairs significantly different from icecuber's?  ",
    "370314": "Thanks for your great solution.",
    "458122": "Great solution ! I was wondering if you had posted the full code anywhere with the full sized network ?",
    "441566": "Hello @outrunner,  I was reading into your solution and in the reconstruction of the tracks, in the step 3 you mention \"Test the new hit to see whether it fits the circle in x-y plane by existing hits after the track has two or three hits. (Without this step I can only get to an 0.8 score.)\" and I was wondering if you could clear my confusion on this.\n\nFrom my understanding, this hit 'k' has already been added to the track, so I don't quite understand why you then test it. Do you remove the hit 'k' from the track if it doesn't fit on the circle from the existing points in the track ? And is there some sort of threshold in place to decide what classifies as being fitted to the circle?\n\nThank you for the solution and any explanation with this would be much appreciated!",
    "386891": "Second, \"Throughput\" phase is online, see [this post][1]\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/65525",
    "382316": "Very cool. I assumed pairwise hit comparison would be too sparse but this is fantastic. I added your dense network to my approach and it's looking a lot better.  Crossing my fingers that the red and blue distributions separate with further training.  :)",
    "375536": "@outrunner \n\nCongrats Nice application of deep learning method! Any plan to speed up your method and apply for the speed phase challenge as well?",
    "370027": "Congratulations @outrunner on coming 2nd and sharing your solution. Thanks also for pushing the rest of us to believe scores above 0.3 and beyond is possible.\n\nI also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.",
    "370816": ""
  }
}