{
  "id": 63249,
  "title": "1st place solution - with code and official documentation",
  "url": "/competitions/trackml-particle-identification/discussion/63249",
  "author_name": "icecuber",
  "post_date": "2018-08-14T00:34:49.698000",
  "votes": 106,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Hello everyone, thank you for a great competition! This was my first serious Kaggle competition, and I must say I'm impressed with how much fun the competition has been to me. I think the organizers have done a great job in making the scope of competition task large enough to be interesting, while not requiring much background knowledge from the field.</p>\n\n<p>Edit: official documentation and code are now available at:\n<a href=\"https://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf\">https://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf</a>\n<a href=\"https://github.com/top-quarks/top-quarks\">https://github.com/top-quarks/top-quarks</a></p>\n\n<p>I'm sorry, but I did not use much machine learning (only some logistic regression for candidate pruning), but rather classical mathematical modeling with statistics and 3d geometry. This, combined with the fact that I wrote everything in C++ with no dependencies, made the final code quite fast: about 8 minutes per event per cpu core for my final submission. So I believe my code could be a good starting point for the throughput phase.</p>\n\n<p><strong>Now for my approach</strong></p>\n\n<p>I divided my algorithm into several steps, and created a scoring metric after each step, so that I could easily tell at which step I could earn the most score. I also made load / score function after each step for rapid debugging and tuning.</p>\n\n<p>There were 48 layers in the detector, each either an annulus or cylinder (approximately). I sorted these approximately so that each track would pass the layers in increasing order. I considered multiple hits of one particle on a single detector to be duplicate measurements, and only looked for a single hit per detector per track until step 4.</p>\n\n<p><strong>1. Select promising pairs of hits.</strong></p>\n\n<p>This was done by considering all pairs of hits on 50 pairs of adjacent layers that covered most of the tracks. These candidates were pruned heavily by a logistic regression model of several heuristics. Some of the heuristics were how far the line passing through the two hits passes from the origin, and the angle between the direction between hits and the direction given by the cells data for each of the hits.\nThis gave about 7 million candidate pairs covering about 99% of the score (meaning for tracks worth 0.99 had at least one pair on that track).</p>\n\n<p><strong>2. Extend the pairs to triples</strong></p>\n\n<p>This was done by extending the line passing through a pair, and looking where it hits the next adjacent detector layers using 3d geometry. I set the 10 closest hits to the intersection as triple candidates. Then I did another pass of pruning by logistic regression to get about 12 million candidate triples. In this step we had three points, so we could fit a helix through them, and we even had one degree of freedom left as a feature for the logistic regression. Other features were (the logarithm of) the radius of the helix, and again the deviation from the direction given by the cell data. The triples covered about 97% of the score (meaning for tracks worth 0.99 I had at least one triple on that track). And the remaining tracks were short, crooked (low momentum), and started far from the z axis.</p>\n\n<p><strong>3. Extend triples to tracks</strong></p>\n\n<p>We fitted a helix through the three hits, and extended it to the adjacent layers using 3d geometry. I always used the helix fitted by the 3 nearest hits on the track to the layer in question. Also here I added the closest hit to the intersection. The resulting (still about 12 million) tracks now contained about 60 million hits, and about 95% of the score (meaning if we optimally assigned tracks using the ground truth data, added all duplicate hits to each track, and ignored &gt;50% coverage constraints, we could get score 0.95).</p>\n\n<p><strong>4. Add duplicate hits</strong></p>\n\n<p>For each track we added the hits closest to it on each layer it passed through. I'm not exactly sure how, but now we covered about 96% of the score :) and I'm not complaining.</p>\n\n<p><strong>5. Assign hits to tracks</strong></p>\n\n<p>Until now all tracks had been processed completely separately, so they were massively overlapping. The goal here was to pick the best paths, and resolve any conflicts between them. My algorithm for this step was based on taking the \"best\" track (I will come back to the metric), removing all hits contained in it from all conflicting paths, and then repeating until there was nothing more to do. This was done efficiently using a data-structure based on a priority queue and dynamic updating of track scores.</p>\n\n<p>The scoring metric to determine the \"best\" tracks was originally based on a random forest and distance from helixes, but I later found something much better. I didn't manage to model the perturbed helix noise. At least, I didn't feel like I had enough quantitative information to do this properly. This meant modeling the probabilities accurately as needed f.ex in a Kalman filter was infeasible. So instead of modeling the inliers (actual helix track), I modeled the probability of outliers (that we would find this track by chance). This was based on the assumption that we could model outliers by the density of hits on a layer, which I assumed was independent of the angle around the z-axis. This outlier density idea was also used for thresholding in all previous steps, so f.ex. saying \"I want 0.1 outlier duplicates on average from each hit\" for making the thresholding distance for duplicates.</p>\n\n<p>The full algorithm gave the final score of about 0.92, using about 90% of the hits.</p>\n\n<p><strong>More important considerations</strong></p>\n\n<p>Of course there were several very important implementation details, note that the above explanation is a simplification down to the most important parts. A crucial technique considering performance, was that I used an acceleration data-structure to quickly access points to close to the helix intersection with a layer. This data-structure based on quad-trees was highly efficient, supported elliptic queries, and took into consideration imprefectness of the layers (they are not exactly annuluses and cylinders), and used polar coordinates to make the maths tractable. I also made a O(1) lookup for close to analytic outlier probability densities in any elliptic region on a detector. A crude model of the magnetic field strength as function of z position of the detector ( \"1.002-z'<em>3e-2-z'^2</em>(0.55-0.3*(1-z'^2))\", where z' = z/2750) gave a 0.003 score boost. On top of that there were a lot of parameters to tune, which were what gave me the last 0.01, and I'm sure there is more to gain if I had the patience.</p>\n\n<p><strong>My takeaways from the competition:</strong></p>\n\n<ul>\n<li>Kaggle has some really interesting competitions.</li>\n<li>Loading bars are really cool! I used them everywhere :)</li>\n<li>It's fun to submit to the leaderboard, even when it isn't strictly strategical considering winning chances.</li>\n<li>Computational resources aren't everything. I got access to a supercomputer, but was unable to improve my score by increasing computational load.</li>\n</ul>\n\n<p>Edit: I added @ersol to the team, as he had experience with cloud computing services. However, in practice I didn't need that, so he didn't end up helping me.</p>",
  "messages": [
    {
      "id": 369917,
      "postDate": "2018-08-14T00:34:49.700Z",
      "content": "<p>Hello everyone, thank you for a great competition! This was my first serious Kaggle competition, and I must say I'm impressed with how much fun the competition has been to me. I think the organizers have done a great job in making the scope of competition task large enough to be interesting, while not requiring much background knowledge from the field.</p>\n\n<p>Edit: official documentation and code are now available at:\n<a href=\"https://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf\">https://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf</a>\n<a href=\"https://github.com/top-quarks/top-quarks\">https://github.com/top-quarks/top-quarks</a></p>\n\n<p>I'm sorry, but I did not use much machine learning (only some logistic regression for candidate pruning), but rather classical mathematical modeling with statistics and 3d geometry. This, combined with the fact that I wrote everything in C++ with no dependencies, made the final code quite fast: about 8 minutes per event per cpu core for my final submission. So I believe my code could be a good starting point for the throughput phase.</p>\n\n<p><strong>Now for my approach</strong></p>\n\n<p>I divided my algorithm into several steps, and created a scoring metric after each step, so that I could easily tell at which step I could earn the most score. I also made load / score function after each step for rapid debugging and tuning.</p>\n\n<p>There were 48 layers in the detector, each either an annulus or cylinder (approximately). I sorted these approximately so that each track would pass the layers in increasing order. I considered multiple hits of one particle on a single detector to be duplicate measurements, and only looked for a single hit per detector per track until step 4.</p>\n\n<p><strong>1. Select promising pairs of hits.</strong></p>\n\n<p>This was done by considering all pairs of hits on 50 pairs of adjacent layers that covered most of the tracks. These candidates were pruned heavily by a logistic regression model of several heuristics. Some of the heuristics were how far the line passing through the two hits passes from the origin, and the angle between the direction between hits and the direction given by the cells data for each of the hits.\nThis gave about 7 million candidate pairs covering about 99% of the score (meaning for tracks worth 0.99 had at least one pair on that track).</p>\n\n<p><strong>2. Extend the pairs to triples</strong></p>\n\n<p>This was done by extending the line passing through a pair, and looking where it hits the next adjacent detector layers using 3d geometry. I set the 10 closest hits to the intersection as triple candidates. Then I did another pass of pruning by logistic regression to get about 12 million candidate triples. In this step we had three points, so we could fit a helix through them, and we even had one degree of freedom left as a feature for the logistic regression. Other features were (the logarithm of) the radius of the helix, and again the deviation from the direction given by the cell data. The triples covered about 97% of the score (meaning for tracks worth 0.99 I had at least one triple on that track). And the remaining tracks were short, crooked (low momentum), and started far from the z axis.</p>\n\n<p><strong>3. Extend triples to tracks</strong></p>\n\n<p>We fitted a helix through the three hits, and extended it to the adjacent layers using 3d geometry. I always used the helix fitted by the 3 nearest hits on the track to the layer in question. Also here I added the closest hit to the intersection. The resulting (still about 12 million) tracks now contained about 60 million hits, and about 95% of the score (meaning if we optimally assigned tracks using the ground truth data, added all duplicate hits to each track, and ignored &gt;50% coverage constraints, we could get score 0.95).</p>\n\n<p><strong>4. Add duplicate hits</strong></p>\n\n<p>For each track we added the hits closest to it on each layer it passed through. I'm not exactly sure how, but now we covered about 96% of the score :) and I'm not complaining.</p>\n\n<p><strong>5. Assign hits to tracks</strong></p>\n\n<p>Until now all tracks had been processed completely separately, so they were massively overlapping. The goal here was to pick the best paths, and resolve any conflicts between them. My algorithm for this step was based on taking the \"best\" track (I will come back to the metric), removing all hits contained in it from all conflicting paths, and then repeating until there was nothing more to do. This was done efficiently using a data-structure based on a priority queue and dynamic updating of track scores.</p>\n\n<p>The scoring metric to determine the \"best\" tracks was originally based on a random forest and distance from helixes, but I later found something much better. I didn't manage to model the perturbed helix noise. At least, I didn't feel like I had enough quantitative information to do this properly. This meant modeling the probabilities accurately as needed f.ex in a Kalman filter was infeasible. So instead of modeling the inliers (actual helix track), I modeled the probability of outliers (that we would find this track by chance). This was based on the assumption that we could model outliers by the density of hits on a layer, which I assumed was independent of the angle around the z-axis. This outlier density idea was also used for thresholding in all previous steps, so f.ex. saying \"I want 0.1 outlier duplicates on average from each hit\" for making the thresholding distance for duplicates.</p>\n\n<p>The full algorithm gave the final score of about 0.92, using about 90% of the hits.</p>\n\n<p><strong>More important considerations</strong></p>\n\n<p>Of course there were several very important implementation details, note that the above explanation is a simplification down to the most important parts. A crucial technique considering performance, was that I used an acceleration data-structure to quickly access points to close to the helix intersection with a layer. This data-structure based on quad-trees was highly efficient, supported elliptic queries, and took into consideration imprefectness of the layers (they are not exactly annuluses and cylinders), and used polar coordinates to make the maths tractable. I also made a O(1) lookup for close to analytic outlier probability densities in any elliptic region on a detector. A crude model of the magnetic field strength as function of z position of the detector ( \"1.002-z'<em>3e-2-z'^2</em>(0.55-0.3*(1-z'^2))\", where z' = z/2750) gave a 0.003 score boost. On top of that there were a lot of parameters to tune, which were what gave me the last 0.01, and I'm sure there is more to gain if I had the patience.</p>\n\n<p><strong>My takeaways from the competition:</strong></p>\n\n<ul>\n<li>Kaggle has some really interesting competitions.</li>\n<li>Loading bars are really cool! I used them everywhere :)</li>\n<li>It's fun to submit to the leaderboard, even when it isn't strictly strategical considering winning chances.</li>\n<li>Computational resources aren't everything. I got access to a supercomputer, but was unable to improve my score by increasing computational load.</li>\n</ul>\n\n<p>Edit: I added @ersol to the team, as he had experience with cloud computing services. However, in practice I didn't need that, so he didn't end up helping me.</p>",
      "rawMarkdown": "Hello everyone, thank you for a great competition! This was my first serious Kaggle competition, and I must say I'm impressed with how much fun the competition has been to me. I think the organizers have done a great job in making the scope of competition task large enough to be interesting, while not requiring much background knowledge from the field.\n\nEdit: official documentation and code are now available at:\nhttps://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf\nhttps://github.com/top-quarks/top-quarks\n\nI'm sorry, but I did not use much machine learning (only some logistic regression for candidate pruning), but rather classical mathematical modeling with statistics and 3d geometry. This, combined with the fact that I wrote everything in C++ with no dependencies, made the final code quite fast: about 8 minutes per event per cpu core for my final submission. So I believe my code could be a good starting point for the throughput phase.\n\n**Now for my approach**\n\nI divided my algorithm into several steps, and created a scoring metric after each step, so that I could easily tell at which step I could earn the most score. I also made load / score function after each step for rapid debugging and tuning.\n\nThere were 48 layers in the detector, each either an annulus or cylinder (approximately). I sorted these approximately so that each track would pass the layers in increasing order. I considered multiple hits of one particle on a single detector to be duplicate measurements, and only looked for a single hit per detector per track until step 4.\n\n**1. Select promising pairs of hits.**\n\nThis was done by considering all pairs of hits on 50 pairs of adjacent layers that covered most of the tracks. These candidates were pruned heavily by a logistic regression model of several heuristics. Some of the heuristics were how far the line passing through the two hits passes from the origin, and the angle between the direction between hits and the direction given by the cells data for each of the hits.\nThis gave about 7 million candidate pairs covering about 99% of the score (meaning for tracks worth 0.99 had at least one pair on that track).\n\n**2. Extend the pairs to triples**\n\nThis was done by extending the line passing through a pair, and looking where it hits the next adjacent detector layers using 3d geometry. I set the 10 closest hits to the intersection as triple candidates. Then I did another pass of pruning by logistic regression to get about 12 million candidate triples. In this step we had three points, so we could fit a helix through them, and we even had one degree of freedom left as a feature for the logistic regression. Other features were (the logarithm of) the radius of the helix, and again the deviation from the direction given by the cell data. The triples covered about 97% of the score (meaning for tracks worth 0.99 I had at least one triple on that track). And the remaining tracks were short, crooked (low momentum), and started far from the z axis.\n\n**3. Extend triples to tracks**\n\nWe fitted a helix through the three hits, and extended it to the adjacent layers using 3d geometry. I always used the helix fitted by the 3 nearest hits on the track to the layer in question. Also here I added the closest hit to the intersection. The resulting (still about 12 million) tracks now contained about 60 million hits, and about 95% of the score (meaning if we optimally assigned tracks using the ground truth data, added all duplicate hits to each track, and ignored &gt;50% coverage constraints, we could get score 0.95).\n\n**4. Add duplicate hits**\n\nFor each track we added the hits closest to it on each layer it passed through. I'm not exactly sure how, but now we covered about 96% of the score :) and I'm not complaining.\n\n**5. Assign hits to tracks**\n\nUntil now all tracks had been processed completely separately, so they were massively overlapping. The goal here was to pick the best paths, and resolve any conflicts between them. My algorithm for this step was based on taking the \"best\" track (I will come back to the metric), removing all hits contained in it from all conflicting paths, and then repeating until there was nothing more to do. This was done efficiently using a data-structure based on a priority queue and dynamic updating of track scores.\n\nThe scoring metric to determine the \"best\" tracks was originally based on a random forest and distance from helixes, but I later found something much better. I didn't manage to model the perturbed helix noise. At least, I didn't feel like I had enough quantitative information to do this properly. This meant modeling the probabilities accurately as needed f.ex in a Kalman filter was infeasible. So instead of modeling the inliers (actual helix track), I modeled the probability of outliers (that we would find this track by chance). This was based on the assumption that we could model outliers by the density of hits on a layer, which I assumed was independent of the angle around the z-axis. This outlier density idea was also used for thresholding in all previous steps, so f.ex. saying \"I want 0.1 outlier duplicates on average from each hit\" for making the thresholding distance for duplicates.\n\nThe full algorithm gave the final score of about 0.92, using about 90% of the hits.\n\n**More important considerations**\n\nOf course there were several very important implementation details, note that the above explanation is a simplification down to the most important parts. A crucial technique considering performance, was that I used an acceleration data-structure to quickly access points to close to the helix intersection with a layer. This data-structure based on quad-trees was highly efficient, supported elliptic queries, and took into consideration imprefectness of the layers (they are not exactly annuluses and cylinders), and used polar coordinates to make the maths tractable. I also made a O(1) lookup for close to analytic outlier probability densities in any elliptic region on a detector. A crude model of the magnetic field strength as function of z position of the detector ( \"1.002-z'*3e-2-z'^2*(0.55-0.3*(1-z'^2))\", where z' = z/2750) gave a 0.003 score boost. On top of that there were a lot of parameters to tune, which were what gave me the last 0.01, and I'm sure there is more to gain if I had the patience.\n\n**My takeaways from the competition:**\n\n - Kaggle has some really interesting competitions.\n - Loading bars are really cool! I used them everywhere :)\n - It's fun to submit to the leaderboard, even when it isn't strictly strategical considering winning chances.\n - Computational resources aren't everything. I got access to a supercomputer, but was unable to improve my score by increasing computational load.\n\nEdit: I added @ersol to the team, as he had experience with cloud computing services. However, in practice I didn't need that, so he didn't end up helping me.",
      "votes": 106
    },
    {
      "id": 369959,
      "postDate": "2018-08-14T02:06:07.547Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a>, I'm in awe and speechless at the same time after having read your approach and thinking \"what have we been doing in the past 3 months\"? You made us feel really stupid. :p  It proved machine learning can't outperform laws of physics. (Science is the king!!) I don't know what people can do in the throughput competition, probably just optimizing your code. :)  Your solution is absolutely extraordinary and I wonder if CERN will still bother to explore any DL approaches after having read this post. May I ask if you're a physicist or a mathematician?</p>",
      "rawMarkdown": "@icecuber, I'm in awe and speechless at the same time after having read your approach and thinking \"what have we been doing in the past 3 months\"? You made us feel really stupid. :p  It proved machine learning can't outperform laws of physics. (Science is the king!!) I don't know what people can do in the throughput competition, probably just optimizing your code. :)  Your solution is absolutely extraordinary and I wonder if CERN will still bother to explore any DL approaches after having read this post. May I ask if you're a physicist or a mathematician?",
      "votes": 7,
      "replies": [
        {
          "id": 370049,
          "postDate": "2018-08-14T06:33:23.573Z",
          "content": "<p>I agree with you! The very hard work. And I think the throughput competition is somewhat useless now. </p>",
          "rawMarkdown": "I agree with you! The very hard work. And I think the throughput competition is somewhat useless now. "
        },
        {
          "id": 370199,
          "postDate": "2018-08-14T12:28:20.520Z",
          "content": "<p>I'm studying a program called \"physics and maths\", with specialization in applied mathematics. Next year I will be writing my master's thesis on deep learning.</p>",
          "rawMarkdown": "I'm studying a program called \"physics and maths\", with specialization in applied mathematics. Next year I will be writing my master's thesis on deep learning.",
          "votes": 4
        },
        {
          "id": 370226,
          "postDate": "2018-08-14T13:15:40.290Z",
          "content": "<p><a href=\"/icecuber\">@icecuber</a>, thanks for letting us know you're \"just\" a master student in physics and math. :)  I was thinking you were a PhD student or a post doc in quantum physics from your profile picture. :D  I thought you named yourself icecuber because you come from Trondheim. We love Norway though we didn't go further north and stopped in Balestrand. Everything in Norway is way too expensive for us, so we ran out of our money before reaching Trondheim... :p  Have to save up more for the next trip to Norway. My previous Kaggle profile picture was actually taken at Jostedalsbreen when @Liam and myself were on a glacier walk. Good luck with your future journey and you'll nail your DL thesis I'm sure. See you next time at Kaggle.</p>",
          "rawMarkdown": "@icecuber, thanks for letting us know you're \"just\" a master student in physics and math. :)  I was thinking you were a PhD student or a post doc in quantum physics from your profile picture. :D  I thought you named yourself icecuber because you come from Trondheim. We love Norway though we didn't go further north and stopped in Balestrand. Everything in Norway is way too expensive for us, so we ran out of our money before reaching Trondheim... :p  Have to save up more for the next trip to Norway. My previous Kaggle profile picture was actually taken at Jostedalsbreen when @Liam and myself were on a glacier walk. Good luck with your future journey and you'll nail your DL thesis I'm sure. See you next time at Kaggle.",
          "votes": 1
        }
      ]
    },
    {
      "id": 369941,
      "postDate": "2018-08-14T01:03:35Z",
      "content": "<p>Amazing! LR is the way, Thanks for sharing.</p>",
      "rawMarkdown": "Amazing! LR is the way, Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 370665,
      "postDate": "2018-08-15T08:45:47.903Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> Well done. Your work is a good example that just ML is not the holy grail of everything. It´s just a tool we should know how to use. </p>",
      "rawMarkdown": "@icecuber Well done. Your work is a good example that just ML is not the holy grail of everything. It´s just a tool we should know how to use. ",
      "votes": 2
    },
    {
      "id": 370072,
      "postDate": "2018-08-14T07:40:16.343Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a>, @ersol, first congratulations to your wonderful result and thank you for explaining your approach!</p>\n\n<p><a href=\"/icecuber\">@icecuber</a>, I am curious due to your nickname: Are you associated with the IceCube experiment? Are the two of you professional physicists?</p>\n\n<p>Reading your description I was taken aback, because it is almost 1-to-1 a description of my own approach. You clearly did decisively better in some places, most importantly in the finding of the track seeds (first two, then three hits per track). Judging by the numbers you give, that is the step were I lose most of the difference, and then probably also some in the assigning step. I had expected to see quite a different approach to emerge as the number one!</p>\n\n<p>Apart from that there are surprisingly many similarities. I also used a geometric/combinatorial approach with nearest neighbor search in each detector layer (I reused scikit's NearestNeighbors' Ball/KD tree and did poor-man's elliptic neighborhoods by scaling the cylinder for cylinder layers). I also used only one hit per layer for the helix fitting, using always the latest three layer crossings, and then in a separate step I find the closest additional hits per layer, etc. Even the background modeling: I model the background hits (everything but the track) as a Poisson point process with density learned from the training data and try to apply Bayesian cuts (I say \"try\" because it worked with some success, but I did not really get it as clean and well-motivated as I'd like to have it.) I do some learning of systematic deviations from helix tracks, but that gives me only a small score improvement which is dwarfed by what you gain compared to me in the other steps.</p>\n\n<p>My approach is pure Python currently and takes about 3 minutes per event, but I guess yours would be faster than mine if you'd scale  down the candidate lists to go down to my score. So I think you have the best basis for the throughput phase, too.</p>",
      "rawMarkdown": "@icecuber, @ersol, first congratulations to your wonderful result and thank you for explaining your approach!\n\n@icecuber, I am curious due to your nickname: Are you associated with the IceCube experiment? Are the two of you professional physicists?\n\nReading your description I was taken aback, because it is almost 1-to-1 a description of my own approach. You clearly did decisively better in some places, most importantly in the finding of the track seeds (first two, then three hits per track). Judging by the numbers you give, that is the step were I lose most of the difference, and then probably also some in the assigning step. I had expected to see quite a different approach to emerge as the number one!\n\nApart from that there are surprisingly many similarities. I also used a geometric/combinatorial approach with nearest neighbor search in each detector layer (I reused scikit's NearestNeighbors' Ball/KD tree and did poor-man's elliptic neighborhoods by scaling the cylinder for cylinder layers). I also used only one hit per layer for the helix fitting, using always the latest three layer crossings, and then in a separate step I find the closest additional hits per layer, etc. Even the background modeling: I model the background hits (everything but the track) as a Poisson point process with density learned from the training data and try to apply Bayesian cuts (I say \"try\" because it worked with some success, but I did not really get it as clean and well-motivated as I'd like to have it.) I do some learning of systematic deviations from helix tracks, but that gives me only a small score improvement which is dwarfed by what you gain compared to me in the other steps.\n\nMy approach is pure Python currently and takes about 3 minutes per event, but I guess yours would be faster than mine if you'd scale  down the candidate lists to go down to my score. So I think you have the best basis for the throughput phase, too.",
      "votes": 2,
      "replies": [
        {
          "id": 370112,
          "postDate": "2018-08-14T09:47:02.390Z",
          "content": "<p>Edwin, congrats on your strong result.  Your solution and icecuber solution confirm what I was thinking from the start: it is somewhat silly to use machine learning to learn physics laws, we'd rather use them directly.  Yet outrunner result show machine learning can do a very good job at finding these laws!</p>",
          "rawMarkdown": "Edwin, congrats on your strong result.  Your solution and icecuber solution confirm what I was thinking from the start: it is somewhat silly to use machine learning to learn physics laws, we'd rather use them directly.  Yet outrunner result show machine learning can do a very good job at finding these laws!",
          "votes": 1
        },
        {
          "id": 370208,
          "postDate": "2018-08-14T12:44:21.267Z",
          "content": "<p>Thank you! The nickname has nothing to do with the IceCube project, it was simply a nickname I took from the speedsolving forum :) . Combining \"ice\" from my cold home country, and \"cuber\" from my interest to solve the Rubik's cube quickly. Regarding profession, I'm just a student.</p>\n\n<p>It is really interesting that you found so many of the same ideas! I haven't heard about Bayesian cuts. It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed. However, I'm sure it makes more sense to you to keep that to yourself until after the throughput phase.</p>",
          "rawMarkdown": "Thank you! The nickname has nothing to do with the IceCube project, it was simply a nickname I took from the speedsolving forum :) . Combining \"ice\" from my cold home country, and \"cuber\" from my interest to solve the Rubik's cube quickly. Regarding profession, I'm just a student.\n\nIt is really interesting that you found so many of the same ideas! I haven't heard about Bayesian cuts. It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed. However, I'm sure it makes more sense to you to keep that to yourself until after the throughput phase.",
          "votes": 3
        },
        {
          "id": 370386,
          "postDate": "2018-08-14T18:18:10.483Z",
          "content": "<blockquote>\n  <p>Regarding profession, I'm just a student.</p>\n</blockquote>\n\n<p>Impressive. With that kind of problem solving skills you have a bright future before you, whether you choose to go into research or industry. Pick wisely and follow your deepest interests, because anything less can get dreadfully boring after a few years!</p>\n\n<blockquote>\n  <p>It is really interesting that you found so many of the same ideas!</p>\n</blockquote>\n\n<p>Yes, it was surprising for me, too.  I guess we are not the only ones and we partially re-invented the conventional approach used in the HEP community.</p>\n\n<blockquote>\n  <p>I haven't heard about Bayesian cuts.</p>\n</blockquote>\n\n<p>It may be a bit pompous on my part to call them such. It's just that I used Bayes' theorem to derive a formula for evaluating how much a hit raises or decreases the probability that a candidate track corresponds to a true particle.</p>\n\n<p>The idea is the following: Let's say you have a set of helix parameters H and an estimated prior probability P0 that the parameters do not correspond to a true track (null hypothesis). Now you look at an intersection of the helix with a detector layer. Let's assume you also have a formal estimate of the error for predicting this intersection (including the estimated error of hit measurement). The intersection point will have some nearest neighbor hit at distance d from the helix. Question: Depending on the value d, what is your posterior probability to believe the null hypothesis?</p>\n\n<p>I got the formula\n<code>P0_post = P0 / (P0 + x * (1 - P0))</code> where 'x' is an update factor which is multiplicative in a sequence of such probability updates and is calculated as</p>\n\n<p><code>x = P(true_d &gt; d|H is real) + p(true_d = d|H is real) / p(background_d = d) * P(background_d &gt; d)</code></p>\n\n<p>where</p>\n\n<ul>\n<li><code>true_d</code> is the helix distance of the nearest hit in the layer that is actually caused by the particle assuming</li>\n<li><code>H is real</code> (the complement of the null hypothesis)</li>\n<li><code>background_d</code> is the helix distance of the nearest \"background hit\", i.e. anything not corresponding to a particle matching the helix parameters H</li>\n<li>the upper case <code>P</code>s are probabilities and</li>\n<li>the lower case <code>p</code>s are probability densities corresponding to the (negative) first derivatives of the <code>P</code>s.</li>\n</ul>\n\n<p>Assuming some models about the distribution of true hits and the distribution of background hits, one can then actually calculate <code>x</code> as a function of <code>d</code>, the prediction errors, etc., and background hit density. <code>x</code> is not a probability, but an unbounded non-negative real with the interpretations (x=0...H is certainly garbage, 0&lt;x&lt;1...H is more likely garbage than we thought before, x=1...the hit we found tells us nothing new, 1&lt;x&lt;+inf...H is more likely a true trajectory than we thought before, x=+inf...H certainly corresponds to a true trajectory).</p>\n\n<p>For the distribution of <code>true_d</code> I assumed that there is a non-zero probability (hyperopt gave 0.63%) that a true particle does not leave any hit on the layer (<code>true_d = inf</code>), otherwise the prediction error is normally distributed in the two dimensions transverse to the helix. For modeling <code>background_d</code> I assumed that the background hits follow a Poisson process with a local density learned from training events.</p>\n\n<p>Well, it kind of worked, at least better than the crude heuristics I had before, but not as well as I hoped. I also did not implement it very cleanly and used it both in the neighbor selection and in track evaluation, which is certainly fishy. One problem that I did not fully resolve is that it generally makes sense to treat the azimuthal direction and the polar/radial direction differently due to the very different measurement and prediction errors. The formulas get quite complicated if one wants to take that into account. At least I did not find sufficient simplifications to make it clean, and in the end I only implemented the parts that actually improved my score.</p>\n\n<blockquote>\n  <p>It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed.</p>\n</blockquote>\n\n<p>I will certainly make my code public, but maybe only after the throughput phase. Maybe earlier, because I think the throughput phase will mostly be about optimizing your code, or similar code from the HEP community, and even if I incorporated the parts were you did significantly better, I could probably compete only in accuracy, not in throughput. My code is fast too, but a C++ solution has probably much more potential for throughput optimization.</p>\n\n<p>For now, let me just timestamp my code in any case:</p>\n\n<p>e6966c5dc3d12eaf82781a3c9a516760f6cfab52 *main.py\ne5f03a6d5bdbbe63f2edf6304358384d83c29fb4 *trackml_solution/algorithm.py\ne9301b282e6a558667b934355f8eb2960bfca670 *trackml_solution/candidates.py\n161f2844d0031f815f078037b6c5301ca223c5d2 *trackml_solution/cells.py\nc88933f20fd79e7fb20211992a4272459ec29f64 *trackml_solution/corrections.py\n42f6fb57555c8bc539799b364852ef2890c14b57 *trackml_solution/data.py\n2fbb74d936ecc4616db47c14e7f0117ef2c427a2 *trackml_solution/geometry.py\n5fa4836b2c3ee7f649c8ad45853ce910e95a474d *trackml_solution/logging.py\n2300e56d798722caeb18670f9b46fd65f44136ff *trackml_solution/neighbors.py\n8e99ebd2e4b2fedb25e7adfa6697ed929f2dc787 *trackml_solution/supervised.py</p>",
          "rawMarkdown": "&gt; Regarding profession, I'm just a student.\n\nImpressive. With that kind of problem solving skills you have a bright future before you, whether you choose to go into research or industry. Pick wisely and follow your deepest interests, because anything less can get dreadfully boring after a few years!\n\n&gt; It is really interesting that you found so many of the same ideas!\n\nYes, it was surprising for me, too.  I guess we are not the only ones and we partially re-invented the conventional approach used in the HEP community.\n\n&gt; I haven't heard about Bayesian cuts.\n\nIt may be a bit pompous on my part to call them such. It's just that I used Bayes' theorem to derive a formula for evaluating how much a hit raises or decreases the probability that a candidate track corresponds to a true particle.\n\nThe idea is the following: Let's say you have a set of helix parameters H and an estimated prior probability P0 that the parameters do not correspond to a true track (null hypothesis). Now you look at an intersection of the helix with a detector layer. Let's assume you also have a formal estimate of the error for predicting this intersection (including the estimated error of hit measurement). The intersection point will have some nearest neighbor hit at distance d from the helix. Question: Depending on the value d, what is your posterior probability to believe the null hypothesis?\n\nI got the formula\n`P0_post = P0 / (P0 + x * (1 - P0))` where 'x' is an update factor which is multiplicative in a sequence of such probability updates and is calculated as\n\n`x = P(true_d &gt; d|H is real) + p(true_d = d|H is real) / p(background_d = d) * P(background_d &gt; d)`\n\nwhere\n\n - `true_d` is the helix distance of the nearest hit in the layer that is actually caused by the particle assuming\n - `H is real` (the complement of the null hypothesis)\n - `background_d` is the helix distance of the nearest \"background hit\", i.e. anything not corresponding to a particle matching the helix parameters H\n - the upper case `P`s are probabilities and\n - the lower case ` p`s are probability densities corresponding to the (negative) first derivatives of the `P`s.\n\nAssuming some models about the distribution of true hits and the distribution of background hits, one can then actually calculate `x` as a function of `d`, the prediction errors, etc., and background hit density. `x` is not a probability, but an unbounded non-negative real with the interpretations (x=0...H is certainly garbage, 0&lt;x&lt;1...H is more likely garbage than we thought before, x=1...the hit we found tells us nothing new, 1&lt;x&lt;+inf...H is more likely a true trajectory than we thought before, x=+inf...H certainly corresponds to a true trajectory).\n\nFor the distribution of `true_d` I assumed that there is a non-zero probability (hyperopt gave 0.63%) that a true particle does not leave any hit on the layer (`true_d = inf`), otherwise the prediction error is normally distributed in the two dimensions transverse to the helix. For modeling `background_d` I assumed that the background hits follow a Poisson process with a local density learned from training events.\n\nWell, it kind of worked, at least better than the crude heuristics I had before, but not as well as I hoped. I also did not implement it very cleanly and used it both in the neighbor selection and in track evaluation, which is certainly fishy. One problem that I did not fully resolve is that it generally makes sense to treat the azimuthal direction and the polar/radial direction differently due to the very different measurement and prediction errors. The formulas get quite complicated if one wants to take that into account. At least I did not find sufficient simplifications to make it clean, and in the end I only implemented the parts that actually improved my score.\n\n&gt; It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed.\n\nI will certainly make my code public, but maybe only after the throughput phase. Maybe earlier, because I think the throughput phase will mostly be about optimizing your code, or similar code from the HEP community, and even if I incorporated the parts were you did significantly better, I could probably compete only in accuracy, not in throughput. My code is fast too, but a C++ solution has probably much more potential for throughput optimization.\n\nFor now, let me just timestamp my code in any case:\n\ne6966c5dc3d12eaf82781a3c9a516760f6cfab52 *main.py\ne5f03a6d5bdbbe63f2edf6304358384d83c29fb4 *trackml_solution/algorithm.py\ne9301b282e6a558667b934355f8eb2960bfca670 *trackml_solution/candidates.py\n161f2844d0031f815f078037b6c5301ca223c5d2 *trackml_solution/cells.py\nc88933f20fd79e7fb20211992a4272459ec29f64 *trackml_solution/corrections.py\n42f6fb57555c8bc539799b364852ef2890c14b57 *trackml_solution/data.py\n2fbb74d936ecc4616db47c14e7f0117ef2c427a2 *trackml_solution/geometry.py\n5fa4836b2c3ee7f649c8ad45853ce910e95a474d *trackml_solution/logging.py\n2300e56d798722caeb18670f9b46fd65f44136ff *trackml_solution/neighbors.py\n8e99ebd2e4b2fedb25e7adfa6697ed929f2dc787 *trackml_solution/supervised.py",
          "votes": 4
        }
      ]
    },
    {
      "id": 370315,
      "postDate": "2018-08-14T15:51:19.670Z",
      "content": "<p>Excellent Solution ! Congratulations and Thanks.</p>",
      "rawMarkdown": "Excellent Solution ! Congratulations and Thanks."
    },
    {
      "id": 370205,
      "postDate": "2018-08-14T12:33:12.347Z",
      "content": "<p>Hi, how many threads do you use per event?  </p>",
      "rawMarkdown": "Hi, how many threads do you use per event?  ",
      "replies": [
        {
          "id": 370211,
          "postDate": "2018-08-14T12:45:51.437Z",
          "content": "<p>only one. I just run multiple events in parallel.</p>",
          "rawMarkdown": "only one. I just run multiple events in parallel.",
          "votes": 3
        }
      ]
    },
    {
      "id": 370334,
      "postDate": "2018-08-14T16:36:13.660Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a></p>\n\n<p>Great idea. Thanks for sharing it and describing step by step. \nCongratulations! </p>\n\n<p>I thought you are working in IceCube collaboration ))) </p>",
      "rawMarkdown": "@icecuber\n\nGreat idea. Thanks for sharing it and describing step by step. \nCongratulations! \n\nI thought you are working in IceCube collaboration ))) ",
      "votes": 1
    },
    {
      "id": 376771,
      "postDate": "2018-08-28T03:26:47.397Z",
      "content": "<p>Great work! Congratulations</p>",
      "rawMarkdown": "Great work! Congratulations"
    },
    {
      "id": 375534,
      "postDate": "2018-08-25T11:45:45.197Z",
      "content": "<p>Congratulations and thank you for the solution!</p>",
      "rawMarkdown": "Congratulations and thank you for the solution!"
    },
    {
      "id": 372970,
      "postDate": "2018-08-20T19:19:31.767Z",
      "content": "<p>Excellent Solution ! Congratulations and Thanks.</p>",
      "rawMarkdown": "Excellent Solution ! Congratulations and Thanks."
    },
    {
      "id": 372699,
      "postDate": "2018-08-20T06:16:39.957Z",
      "content": "<p>(y)</p>",
      "rawMarkdown": "(y)"
    },
    {
      "id": 372563,
      "postDate": "2018-08-19T17:55:42.933Z",
      "content": "<p>Your solution shows how hard you worked for this win. Congratulations!</p>",
      "rawMarkdown": "Your solution shows how hard you worked for this win. Congratulations!"
    },
    {
      "id": 372368,
      "postDate": "2018-08-19T05:43:03.333Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> Well Done. Congratulations </p>",
      "rawMarkdown": "@icecuber Well Done. Congratulations "
    },
    {
      "id": 370289,
      "postDate": "2018-08-14T15:17:35.260Z",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> did you find any discrepancies in the cell vs truth data e.g. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/63027\">event000001000/19932</a>? They are going to regenerate the dataset for the throughput phase so if there is a bug in cell data it would be good to flag this to them before they do it.</p>",
      "rawMarkdown": "@icecuber did you find any discrepancies in the cell vs truth data e.g. [event000001000/19932][1]? They are going to regenerate the dataset for the throughput phase so if there is a bug in cell data it would be good to flag this to them before they do it.\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/63027",
      "replies": [
        {
          "id": 370403,
          "postDate": "2018-08-14T19:09:25.703Z",
          "content": "<p>I also used the cells data in the end and did weighted linear regression to find the direction vectors, probably similar to what <a href=\"/icecuber\">@icecuber</a> did. I found some cases where the implied direction vector was at a quite large angle to the true momentum vector. Maybe I'll try to reproduce such cases. However, it does not necessarily mean that those are bugs. It could be some rare large-angle scattering effects happening between the activation of the cells and the storage of the true momentum vector.</p>",
          "rawMarkdown": "I also used the cells data in the end and did weighted linear regression to find the direction vectors, probably similar to what @icecuber did. I found some cases where the implied direction vector was at a quite large angle to the true momentum vector. Maybe I'll try to reproduce such cases. However, it does not necessarily mean that those are bugs. It could be some rare large-angle scattering effects happening between the activation of the cells and the storage of the true momentum vector."
        },
        {
          "id": 371001,
          "postDate": "2018-08-15T19:34:04.173Z",
          "content": "<p><a href=\"/edwinst\">@edwinst</a> I have double checked event000001000/19932 and there is no scattering in this case. If I compare the momentum with the lines connecting the hits there is virtually no difference. If you would like to investigate and report it, as your word here is worth more than mine, I would recommend looking at this hit. The organisers ignored my post and I am moving on to another project so I am leaving this with you guys. All the best in the 2nd round.</p>",
          "rawMarkdown": "@edwinst I have double checked event000001000/19932 and there is no scattering in this case. If I compare the momentum with the lines connecting the hits there is virtually no difference. If you would like to investigate and report it, as your word here is worth more than mine, I would recommend looking at this hit. The organisers ignored my post and I am moving on to another project so I am leaving this with you guys. All the best in the 2nd round."
        },
        {
          "id": 371461,
          "postDate": "2018-08-16T21:01:59.007Z",
          "content": "<p>&gt; The organisers ignored my post</p>\n\n<p>I noticed that the organizers read more of the comments than is immediately apparent from their responses. If I can get around to it, I will post some examples of the cells data including strange ones and also one looking at the hit you specify.</p>",
          "rawMarkdown": "&gt; The organisers ignored my post\n\nI noticed that the organizers read more of the comments than is immediately apparent from their responses. If I can get around to it, I will post some examples of the cells data including strange ones and also one looking at the hit you specify."
        },
        {
          "id": 371483,
          "postDate": "2018-08-16T22:17:55.213Z",
          "content": "<p>That would be useful. We ce investigated and understood the funny trajectories seen by some people (see pinned topic(. I don t recall something funny with the cells, but this could have escaped us. </p>",
          "rawMarkdown": "That would be useful. We ce investigated and understood the funny trajectories seen by some people (see pinned topic(. I don t recall something funny with the cells, but this could have escaped us. "
        }
      ]
    },
    {
      "id": 370124,
      "postDate": "2018-08-14T10:22:36.027Z",
      "content": "<p>Great work!\nCongratulations <a href=\"/icecuber\">@icecuber</a>.</p>",
      "rawMarkdown": "Great work!\nCongratulations @icecuber."
    },
    {
      "id": 370019,
      "postDate": "2018-08-14T05:17:57.740Z",
      "content": "<p>Congratulations <a href=\"/icecuber\">@icecuber</a> and team. Thanks for sharing your so elegant solution.  Wow  \"... 8 minutes per event per cpu core...\", how impressive. </p>\n\n<p>I also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.</p>",
      "rawMarkdown": "Congratulations @icecuber and team. Thanks for sharing your so elegant solution.  Wow  \"... 8 minutes per event per cpu core...\", how impressive. \n\nI also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.",
      "replies": [
        {
          "id": 370155,
          "postDate": "2018-08-14T11:05:11.470Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 370157,
          "postDate": "2018-08-14T11:05:43.737Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 369973,
      "postDate": "2018-08-14T03:10:41.857Z",
      "content": "<p>Thoughtful approach. Well done!!</p>",
      "rawMarkdown": "Thoughtful approach. Well done!!"
    },
    {
      "id": 369963,
      "postDate": "2018-08-14T02:15:09.790Z",
      "content": "<p>Hi there, congratulations on your finish! Can you share your source code?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi there, congratulations on your finish! Can you share your source code?\n\nThanks.",
      "replies": [
        {
          "id": 370191,
          "postDate": "2018-08-14T12:20:42.207Z",
          "content": "<p>Thank you! I will share my source code soon, after I've set up a license and written some instruction on how to run it.</p>",
          "rawMarkdown": "Thank you! I will share my source code soon, after I've set up a license and written some instruction on how to run it.",
          "votes": 1,
          "replies": [
            {
              "id": 3339451,
              "postDate": "2025-11-19T04:46:51.670Z",
              "content": "<p>Can you share source code Today?</p>",
              "rawMarkdown": "Can you share source code Today?"
            }
          ]
        }
      ]
    },
    {
      "id": 369939,
      "postDate": "2018-08-14T00:56:13.983Z",
      "content": "<p>Congrats! Really amazing work, thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Really amazing work, thanks for sharing."
    },
    {
      "id": 369936,
      "postDate": "2018-08-14T00:55:21.580Z",
      "content": "<p>At some stage of the competition, I thought that if someone could think of the way to grow tracks from hits in the global scope (like bacteria grows and duplicate themselves in the air). Now, the first solution is similar like this.  That's truly a wonder to behold. Thanks for sharing.</p>",
      "rawMarkdown": "At some stage of the competition, I thought that if someone could think of the way to grow tracks from hits in the global scope (like bacteria grows and duplicate themselves in the air). Now, the first solution is similar like this.  That's truly a wonder to behold. Thanks for sharing."
    },
    {
      "id": 369932,
      "postDate": "2018-08-14T00:52:56.860Z",
      "content": "<p>This is great. Congratulations. </p>\n\n<p>Regarding using direction from cell's data: Was this crucial in your steps 1 and 2, or could you have done without it, at a cost of a few percentage points to the score? </p>\n\n<p>Calculating direction from the cells file seemed very interesting, but there was also quite some noise (direction not available for all hits, and also some real inaccuracies involved). You surely mastered to work with that data. Great work.</p>",
      "rawMarkdown": "This is great. Congratulations. \n\nRegarding using direction from cell's data: Was this crucial in your steps 1 and 2, or could you have done without it, at a cost of a few percentage points to the score? \n\nCalculating direction from the cells file seemed very interesting, but there was also quite some noise (direction not available for all hits, and also some real inaccuracies involved). You surely mastered to work with that data. Great work.",
      "replies": [
        {
          "id": 369942,
          "postDate": "2018-08-14T01:03:47.717Z",
          "content": "<p>Thank you. I don't think the cell's data was very important, if I remember correctly I added it after I was at 0.9-something and it didn't give too much improvement.\nThe cells code itself wasn't too hard after I figured out what the data meant. I used regression on the plane of the cell, and then an analytical formula for the angle with the plane based on number of cells intersected. It's about 40 lines of code.</p>",
          "rawMarkdown": "Thank you. I don't think the cell's data was very important, if I remember correctly I added it after I was at 0.9-something and it didn't give too much improvement.\nThe cells code itself wasn't too hard after I figured out what the data meant. I used regression on the plane of the cell, and then an analytical formula for the angle with the plane based on number of cells intersected. It's about 40 lines of code.",
          "votes": 2
        }
      ]
    },
    {
      "id": 771299,
      "postDate": "2020-03-14T01:42:23.103Z",
      "content": "<p>Thank u for sharing</p>",
      "rawMarkdown": "Thank u for sharing"
    },
    {
      "id": 374695,
      "postDate": "2018-08-23T15:31:01.253Z",
      "content": "<p>Excellent! thank you for sharing</p>",
      "rawMarkdown": "Excellent! thank you for sharing"
    },
    {
      "id": 372093,
      "postDate": "2018-08-18T09:35:11.283Z",
      "content": "<p>Grt work! Thanks</p>",
      "rawMarkdown": "Grt work! Thanks"
    },
    {
      "id": 370059,
      "postDate": "2018-08-14T07:01:10.430Z",
      "content": "<p>Amazing, thank you !</p>",
      "rawMarkdown": "Amazing, thank you !"
    },
    {
      "id": 369923,
      "postDate": "2018-08-14T00:39:57.940Z",
      "content": "<p>Wow, thanks for sharing, and congrats.</p>",
      "rawMarkdown": "Wow, thanks for sharing, and congrats."
    }
  ],
  "comments": [
    {
      "id": 369959,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-08-14T02:06:07.547000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a>, I'm in awe and speechless at the same time after having read your approach and thinking \"what have we been doing in the past 3 months\"? You made us feel really stupid. :p  It proved machine learning can't outperform laws of physics. (Science is the king!!) I don't know what people can do in the throughput competition, probably just optimizing your code. :)  Your solution is absolutely extraordinary and I wonder if CERN will still bother to explore any DL approaches after having read this post. May I ask if you're a physicist or a mathematician?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 370049,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-14T06:33:23.573000",
          "content": "<p>I agree with you! The very hard work. And I think the throughput competition is somewhat useless now. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 370199,
          "author_name": "icecuber",
          "author_url": "",
          "post_date": "2018-08-14T12:28:20.520000",
          "content": "<p>I'm studying a program called \"physics and maths\", with specialization in applied mathematics. Next year I will be writing my master's thesis on deep learning.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 370226,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-14T13:15:40.290000",
          "content": "<p><a href=\"/icecuber\">@icecuber</a>, thanks for letting us know you're \"just\" a master student in physics and math. :)  I was thinking you were a PhD student or a post doc in quantum physics from your profile picture. :D  I thought you named yourself icecuber because you come from Trondheim. We love Norway though we didn't go further north and stopped in Balestrand. Everything in Norway is way too expensive for us, so we ran out of our money before reaching Trondheim... :p  Have to save up more for the next trip to Norway. My previous Kaggle profile picture was actually taken at Jostedalsbreen when @Liam and myself were on a glacier walk. Good luck with your future journey and you'll nail your DL thesis I'm sure. See you next time at Kaggle.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 369941,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-08-14T01:03:35",
      "content": "<p>Amazing! LR is the way, Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 370665,
      "author_name": "Thomas Loock",
      "author_url": "",
      "post_date": "2018-08-15T08:45:47.903000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> Well done. Your work is a good example that just ML is not the holy grail of everything. It´s just a tool we should know how to use. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 370072,
      "author_name": "Edwin Steiner",
      "author_url": "",
      "post_date": "2018-08-14T07:40:16.343000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a>, @ersol, first congratulations to your wonderful result and thank you for explaining your approach!</p>\n\n<p><a href=\"/icecuber\">@icecuber</a>, I am curious due to your nickname: Are you associated with the IceCube experiment? Are the two of you professional physicists?</p>\n\n<p>Reading your description I was taken aback, because it is almost 1-to-1 a description of my own approach. You clearly did decisively better in some places, most importantly in the finding of the track seeds (first two, then three hits per track). Judging by the numbers you give, that is the step were I lose most of the difference, and then probably also some in the assigning step. I had expected to see quite a different approach to emerge as the number one!</p>\n\n<p>Apart from that there are surprisingly many similarities. I also used a geometric/combinatorial approach with nearest neighbor search in each detector layer (I reused scikit's NearestNeighbors' Ball/KD tree and did poor-man's elliptic neighborhoods by scaling the cylinder for cylinder layers). I also used only one hit per layer for the helix fitting, using always the latest three layer crossings, and then in a separate step I find the closest additional hits per layer, etc. Even the background modeling: I model the background hits (everything but the track) as a Poisson point process with density learned from the training data and try to apply Bayesian cuts (I say \"try\" because it worked with some success, but I did not really get it as clean and well-motivated as I'd like to have it.) I do some learning of systematic deviations from helix tracks, but that gives me only a small score improvement which is dwarfed by what you gain compared to me in the other steps.</p>\n\n<p>My approach is pure Python currently and takes about 3 minutes per event, but I guess yours would be faster than mine if you'd scale  down the candidate lists to go down to my score. So I think you have the best basis for the throughput phase, too.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 370112,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-14T09:47:02.390000",
          "content": "<p>Edwin, congrats on your strong result.  Your solution and icecuber solution confirm what I was thinking from the start: it is somewhat silly to use machine learning to learn physics laws, we'd rather use them directly.  Yet outrunner result show machine learning can do a very good job at finding these laws!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 370208,
          "author_name": "icecuber",
          "author_url": "",
          "post_date": "2018-08-14T12:44:21.267000",
          "content": "<p>Thank you! The nickname has nothing to do with the IceCube project, it was simply a nickname I took from the speedsolving forum :) . Combining \"ice\" from my cold home country, and \"cuber\" from my interest to solve the Rubik's cube quickly. Regarding profession, I'm just a student.</p>\n\n<p>It is really interesting that you found so many of the same ideas! I haven't heard about Bayesian cuts. It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed. However, I'm sure it makes more sense to you to keep that to yourself until after the throughput phase.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 370386,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-14T18:18:10.483000",
          "content": "<blockquote>\n  <p>Regarding profession, I'm just a student.</p>\n</blockquote>\n\n<p>Impressive. With that kind of problem solving skills you have a bright future before you, whether you choose to go into research or industry. Pick wisely and follow your deepest interests, because anything less can get dreadfully boring after a few years!</p>\n\n<blockquote>\n  <p>It is really interesting that you found so many of the same ideas!</p>\n</blockquote>\n\n<p>Yes, it was surprising for me, too.  I guess we are not the only ones and we partially re-invented the conventional approach used in the HEP community.</p>\n\n<blockquote>\n  <p>I haven't heard about Bayesian cuts.</p>\n</blockquote>\n\n<p>It may be a bit pompous on my part to call them such. It's just that I used Bayes' theorem to derive a formula for evaluating how much a hit raises or decreases the probability that a candidate track corresponds to a true particle.</p>\n\n<p>The idea is the following: Let's say you have a set of helix parameters H and an estimated prior probability P0 that the parameters do not correspond to a true track (null hypothesis). Now you look at an intersection of the helix with a detector layer. Let's assume you also have a formal estimate of the error for predicting this intersection (including the estimated error of hit measurement). The intersection point will have some nearest neighbor hit at distance d from the helix. Question: Depending on the value d, what is your posterior probability to believe the null hypothesis?</p>\n\n<p>I got the formula\n<code>P0_post = P0 / (P0 + x * (1 - P0))</code> where 'x' is an update factor which is multiplicative in a sequence of such probability updates and is calculated as</p>\n\n<p><code>x = P(true_d &gt; d|H is real) + p(true_d = d|H is real) / p(background_d = d) * P(background_d &gt; d)</code></p>\n\n<p>where</p>\n\n<ul>\n<li><code>true_d</code> is the helix distance of the nearest hit in the layer that is actually caused by the particle assuming</li>\n<li><code>H is real</code> (the complement of the null hypothesis)</li>\n<li><code>background_d</code> is the helix distance of the nearest \"background hit\", i.e. anything not corresponding to a particle matching the helix parameters H</li>\n<li>the upper case <code>P</code>s are probabilities and</li>\n<li>the lower case <code>p</code>s are probability densities corresponding to the (negative) first derivatives of the <code>P</code>s.</li>\n</ul>\n\n<p>Assuming some models about the distribution of true hits and the distribution of background hits, one can then actually calculate <code>x</code> as a function of <code>d</code>, the prediction errors, etc., and background hit density. <code>x</code> is not a probability, but an unbounded non-negative real with the interpretations (x=0...H is certainly garbage, 0&lt;x&lt;1...H is more likely garbage than we thought before, x=1...the hit we found tells us nothing new, 1&lt;x&lt;+inf...H is more likely a true trajectory than we thought before, x=+inf...H certainly corresponds to a true trajectory).</p>\n\n<p>For the distribution of <code>true_d</code> I assumed that there is a non-zero probability (hyperopt gave 0.63%) that a true particle does not leave any hit on the layer (<code>true_d = inf</code>), otherwise the prediction error is normally distributed in the two dimensions transverse to the helix. For modeling <code>background_d</code> I assumed that the background hits follow a Poisson process with a local density learned from training events.</p>\n\n<p>Well, it kind of worked, at least better than the crude heuristics I had before, but not as well as I hoped. I also did not implement it very cleanly and used it both in the neighbor selection and in track evaluation, which is certainly fishy. One problem that I did not fully resolve is that it generally makes sense to treat the azimuthal direction and the polar/radial direction differently due to the very different measurement and prediction errors. The formulas get quite complicated if one wants to take that into account. At least I did not find sufficient simplifications to make it clean, and in the end I only implemented the parts that actually improved my score.</p>\n\n<blockquote>\n  <p>It would be really interesting to hear about the more detailed ideas you found, which would probably cover some I missed.</p>\n</blockquote>\n\n<p>I will certainly make my code public, but maybe only after the throughput phase. Maybe earlier, because I think the throughput phase will mostly be about optimizing your code, or similar code from the HEP community, and even if I incorporated the parts were you did significantly better, I could probably compete only in accuracy, not in throughput. My code is fast too, but a C++ solution has probably much more potential for throughput optimization.</p>\n\n<p>For now, let me just timestamp my code in any case:</p>\n\n<p>e6966c5dc3d12eaf82781a3c9a516760f6cfab52 *main.py\ne5f03a6d5bdbbe63f2edf6304358384d83c29fb4 *trackml_solution/algorithm.py\ne9301b282e6a558667b934355f8eb2960bfca670 *trackml_solution/candidates.py\n161f2844d0031f815f078037b6c5301ca223c5d2 *trackml_solution/cells.py\nc88933f20fd79e7fb20211992a4272459ec29f64 *trackml_solution/corrections.py\n42f6fb57555c8bc539799b364852ef2890c14b57 *trackml_solution/data.py\n2fbb74d936ecc4616db47c14e7f0117ef2c427a2 *trackml_solution/geometry.py\n5fa4836b2c3ee7f649c8ad45853ce910e95a474d *trackml_solution/logging.py\n2300e56d798722caeb18670f9b46fd65f44136ff *trackml_solution/neighbors.py\n8e99ebd2e4b2fedb25e7adfa6697ed929f2dc787 *trackml_solution/supervised.py</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 370315,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "2018-08-14T15:51:19.670000",
      "content": "<p>Excellent Solution ! Congratulations and Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 370205,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-14T12:33:12.347000",
      "content": "<p>Hi, how many threads do you use per event?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 370211,
          "author_name": "icecuber",
          "author_url": "",
          "post_date": "2018-08-14T12:45:51.437000",
          "content": "<p>only one. I just run multiple events in parallel.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 370334,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2018-08-14T16:36:13.660000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a></p>\n\n<p>Great idea. Thanks for sharing it and describing step by step. \nCongratulations! </p>\n\n<p>I thought you are working in IceCube collaboration ))) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 376771,
      "author_name": "raghav",
      "author_url": "",
      "post_date": "2018-08-28T03:26:47.397000",
      "content": "<p>Great work! Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 375534,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-08-25T11:45:45.197000",
      "content": "<p>Congratulations and thank you for the solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372970,
      "author_name": "Leandro Ferreira",
      "author_url": "",
      "post_date": "2018-08-20T19:19:31.767000",
      "content": "<p>Excellent Solution ! Congratulations and Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372699,
      "author_name": "Abhishek Jain",
      "author_url": "",
      "post_date": "2018-08-20T06:16:39.957000",
      "content": "<p>(y)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372563,
      "author_name": "Siddharth Yadav",
      "author_url": "",
      "post_date": "2018-08-19T17:55:42.933000",
      "content": "<p>Your solution shows how hard you worked for this win. Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372368,
      "author_name": "Abhijeet Vichare",
      "author_url": "",
      "post_date": "2018-08-19T05:43:03.333000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> Well Done. Congratulations </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 370289,
      "author_name": "Dthrone",
      "author_url": "",
      "post_date": "2018-08-14T15:17:35.260000",
      "content": "<p><a href=\"/icecuber\">@icecuber</a> did you find any discrepancies in the cell vs truth data e.g. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/63027\">event000001000/19932</a>? They are going to regenerate the dataset for the throughput phase so if there is a bug in cell data it would be good to flag this to them before they do it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 370403,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-14T19:09:25.703000",
          "content": "<p>I also used the cells data in the end and did weighted linear regression to find the direction vectors, probably similar to what <a href=\"/icecuber\">@icecuber</a> did. I found some cases where the implied direction vector was at a quite large angle to the true momentum vector. Maybe I'll try to reproduce such cases. However, it does not necessarily mean that those are bugs. It could be some rare large-angle scattering effects happening between the activation of the cells and the storage of the true momentum vector.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371001,
          "author_name": "Dthrone",
          "author_url": "",
          "post_date": "2018-08-15T19:34:04.173000",
          "content": "<p><a href=\"/edwinst\">@edwinst</a> I have double checked event000001000/19932 and there is no scattering in this case. If I compare the momentum with the lines connecting the hits there is virtually no difference. If you would like to investigate and report it, as your word here is worth more than mine, I would recommend looking at this hit. The organisers ignored my post and I am moving on to another project so I am leaving this with you guys. All the best in the 2nd round.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371461,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-16T21:01:59.007000",
          "content": "<p>&gt; The organisers ignored my post</p>\n\n<p>I noticed that the organizers read more of the comments than is immediately apparent from their responses. If I can get around to it, I will post some examples of the cells data including strange ones and also one looking at the hit you specify.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371483,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-08-16T22:17:55.213000",
          "content": "<p>That would be useful. We ce investigated and understood the funny trajectories seen by some people (see pinned topic(. I don t recall something funny with the cells, but this could have escaped us. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 370124,
      "author_name": "Amit Kumar Jaiswal",
      "author_url": "",
      "post_date": "2018-08-14T10:22:36.027000",
      "content": "<p>Great work!\nCongratulations <a href=\"/icecuber\">@icecuber</a>.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 370019,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-08-14T05:17:57.740000",
      "content": "<p>Congratulations <a href=\"/icecuber\">@icecuber</a> and team. Thanks for sharing your so elegant solution.  Wow  \"... 8 minutes per event per cpu core...\", how impressive. </p>\n\n<p>I also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 370155,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-14T11:05:11.470000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 370157,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-14T11:05:43.737000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 369973,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2018-08-14T03:10:41.857000",
      "content": "<p>Thoughtful approach. Well done!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 369963,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2018-08-14T02:15:09.790000",
      "content": "<p>Hi there, congratulations on your finish! Can you share your source code?</p>\n\n<p>Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 370191,
          "author_name": "icecuber",
          "author_url": "",
          "post_date": "2018-08-14T12:20:42.207000",
          "content": "<p>Thank you! I will share my source code soon, after I've set up a license and written some instruction on how to run it.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3339451,
              "author_name": "kanishkbakshi2004",
              "author_url": "",
              "post_date": "2025-11-19T04:46:51.670000",
              "content": "<p>Can you share source code Today?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 369939,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2018-08-14T00:56:13.983000",
      "content": "<p>Congrats! Really amazing work, thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 369936,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2018-08-14T00:55:21.580000",
      "content": "<p>At some stage of the competition, I thought that if someone could think of the way to grow tracks from hits in the global scope (like bacteria grows and duplicate themselves in the air). Now, the first solution is similar like this.  That's truly a wonder to behold. Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 369932,
      "author_name": "Trian",
      "author_url": "",
      "post_date": "2018-08-14T00:52:56.860000",
      "content": "<p>This is great. Congratulations. </p>\n\n<p>Regarding using direction from cell's data: Was this crucial in your steps 1 and 2, or could you have done without it, at a cost of a few percentage points to the score? </p>\n\n<p>Calculating direction from the cells file seemed very interesting, but there was also quite some noise (direction not available for all hits, and also some real inaccuracies involved). You surely mastered to work with that data. Great work.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 369942,
          "author_name": "icecuber",
          "author_url": "",
          "post_date": "2018-08-14T01:03:47.717000",
          "content": "<p>Thank you. I don't think the cell's data was very important, if I remember correctly I added it after I was at 0.9-something and it didn't give too much improvement.\nThe cells code itself wasn't too hard after I figured out what the data meant. I used regression on the plane of the cell, and then an analytical formula for the angle with the plane based on number of cells intersected. It's about 40 lines of code.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 771299,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-14T01:42:23.103000",
      "content": "<p>Thank u for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 374695,
      "author_name": "Shailendra Tomar",
      "author_url": "",
      "post_date": "2018-08-23T15:31:01.253000",
      "content": "<p>Excellent! thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372093,
      "author_name": "leo022",
      "author_url": "",
      "post_date": "2018-08-18T09:35:11.283000",
      "content": "<p>Grt work! Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 370059,
      "author_name": "Ulysse31",
      "author_url": "",
      "post_date": "2018-08-14T07:01:10.430000",
      "content": "<p>Amazing, thank you !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 369923,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-14T00:39:57.940000",
      "content": "<p>Wow, thanks for sharing, and congrats.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "369917": "Hello everyone, thank you for a great competition! This was my first serious Kaggle competition, and I must say I'm impressed with how much fun the competition has been to me. I think the organizers have done a great job in making the scope of competition task large enough to be interesting, while not requiring much background knowledge from the field.\n\nEdit: official documentation and code are now available at:\nhttps://github.com/top-quarks/top-quarks/blob/master/top-quarks_documentation.pdf\nhttps://github.com/top-quarks/top-quarks\n\nI'm sorry, but I did not use much machine learning (only some logistic regression for candidate pruning), but rather classical mathematical modeling with statistics and 3d geometry. This, combined with the fact that I wrote everything in C++ with no dependencies, made the final code quite fast: about 8 minutes per event per cpu core for my final submission. So I believe my code could be a good starting point for the throughput phase.\n\n**Now for my approach**\n\nI divided my algorithm into several steps, and created a scoring metric after each step, so that I could easily tell at which step I could earn the most score. I also made load / score function after each step for rapid debugging and tuning.\n\nThere were 48 layers in the detector, each either an annulus or cylinder (approximately). I sorted these approximately so that each track would pass the layers in increasing order. I considered multiple hits of one particle on a single detector to be duplicate measurements, and only looked for a single hit per detector per track until step 4.\n\n**1. Select promising pairs of hits.**\n\nThis was done by considering all pairs of hits on 50 pairs of adjacent layers that covered most of the tracks. These candidates were pruned heavily by a logistic regression model of several heuristics. Some of the heuristics were how far the line passing through the two hits passes from the origin, and the angle between the direction between hits and the direction given by the cells data for each of the hits.\nThis gave about 7 million candidate pairs covering about 99% of the score (meaning for tracks worth 0.99 had at least one pair on that track).\n\n**2. Extend the pairs to triples**\n\nThis was done by extending the line passing through a pair, and looking where it hits the next adjacent detector layers using 3d geometry. I set the 10 closest hits to the intersection as triple candidates. Then I did another pass of pruning by logistic regression to get about 12 million candidate triples. In this step we had three points, so we could fit a helix through them, and we even had one degree of freedom left as a feature for the logistic regression. Other features were (the logarithm of) the radius of the helix, and again the deviation from the direction given by the cell data. The triples covered about 97% of the score (meaning for tracks worth 0.99 I had at least one triple on that track). And the remaining tracks were short, crooked (low momentum), and started far from the z axis.\n\n**3. Extend triples to tracks**\n\nWe fitted a helix through the three hits, and extended it to the adjacent layers using 3d geometry. I always used the helix fitted by the 3 nearest hits on the track to the layer in question. Also here I added the closest hit to the intersection. The resulting (still about 12 million) tracks now contained about 60 million hits, and about 95% of the score (meaning if we optimally assigned tracks using the ground truth data, added all duplicate hits to each track, and ignored &gt;50% coverage constraints, we could get score 0.95).\n\n**4. Add duplicate hits**\n\nFor each track we added the hits closest to it on each layer it passed through. I'm not exactly sure how, but now we covered about 96% of the score :) and I'm not complaining.\n\n**5. Assign hits to tracks**\n\nUntil now all tracks had been processed completely separately, so they were massively overlapping. The goal here was to pick the best paths, and resolve any conflicts between them. My algorithm for this step was based on taking the \"best\" track (I will come back to the metric), removing all hits contained in it from all conflicting paths, and then repeating until there was nothing more to do. This was done efficiently using a data-structure based on a priority queue and dynamic updating of track scores.\n\nThe scoring metric to determine the \"best\" tracks was originally based on a random forest and distance from helixes, but I later found something much better. I didn't manage to model the perturbed helix noise. At least, I didn't feel like I had enough quantitative information to do this properly. This meant modeling the probabilities accurately as needed f.ex in a Kalman filter was infeasible. So instead of modeling the inliers (actual helix track), I modeled the probability of outliers (that we would find this track by chance). This was based on the assumption that we could model outliers by the density of hits on a layer, which I assumed was independent of the angle around the z-axis. This outlier density idea was also used for thresholding in all previous steps, so f.ex. saying \"I want 0.1 outlier duplicates on average from each hit\" for making the thresholding distance for duplicates.\n\nThe full algorithm gave the final score of about 0.92, using about 90% of the hits.\n\n**More important considerations**\n\nOf course there were several very important implementation details, note that the above explanation is a simplification down to the most important parts. A crucial technique considering performance, was that I used an acceleration data-structure to quickly access points to close to the helix intersection with a layer. This data-structure based on quad-trees was highly efficient, supported elliptic queries, and took into consideration imprefectness of the layers (they are not exactly annuluses and cylinders), and used polar coordinates to make the maths tractable. I also made a O(1) lookup for close to analytic outlier probability densities in any elliptic region on a detector. A crude model of the magnetic field strength as function of z position of the detector ( \"1.002-z'*3e-2-z'^2*(0.55-0.3*(1-z'^2))\", where z' = z/2750) gave a 0.003 score boost. On top of that there were a lot of parameters to tune, which were what gave me the last 0.01, and I'm sure there is more to gain if I had the patience.\n\n**My takeaways from the competition:**\n\n - Kaggle has some really interesting competitions.\n - Loading bars are really cool! I used them everywhere :)\n - It's fun to submit to the leaderboard, even when it isn't strictly strategical considering winning chances.\n - Computational resources aren't everything. I got access to a supercomputer, but was unable to improve my score by increasing computational load.\n\nEdit: I added @ersol to the team, as he had experience with cloud computing services. However, in practice I didn't need that, so he didn't end up helping me.",
    "369959": "@icecuber, I'm in awe and speechless at the same time after having read your approach and thinking \"what have we been doing in the past 3 months\"? You made us feel really stupid. :p  It proved machine learning can't outperform laws of physics. (Science is the king!!) I don't know what people can do in the throughput competition, probably just optimizing your code. :)  Your solution is absolutely extraordinary and I wonder if CERN will still bother to explore any DL approaches after having read this post. May I ask if you're a physicist or a mathematician?",
    "369941": "Amazing! LR is the way, Thanks for sharing.",
    "370665": "@icecuber Well done. Your work is a good example that just ML is not the holy grail of everything. It´s just a tool we should know how to use. ",
    "370072": "@icecuber, @ersol, first congratulations to your wonderful result and thank you for explaining your approach!\n\n@icecuber, I am curious due to your nickname: Are you associated with the IceCube experiment? Are the two of you professional physicists?\n\nReading your description I was taken aback, because it is almost 1-to-1 a description of my own approach. You clearly did decisively better in some places, most importantly in the finding of the track seeds (first two, then three hits per track). Judging by the numbers you give, that is the step were I lose most of the difference, and then probably also some in the assigning step. I had expected to see quite a different approach to emerge as the number one!\n\nApart from that there are surprisingly many similarities. I also used a geometric/combinatorial approach with nearest neighbor search in each detector layer (I reused scikit's NearestNeighbors' Ball/KD tree and did poor-man's elliptic neighborhoods by scaling the cylinder for cylinder layers). I also used only one hit per layer for the helix fitting, using always the latest three layer crossings, and then in a separate step I find the closest additional hits per layer, etc. Even the background modeling: I model the background hits (everything but the track) as a Poisson point process with density learned from the training data and try to apply Bayesian cuts (I say \"try\" because it worked with some success, but I did not really get it as clean and well-motivated as I'd like to have it.) I do some learning of systematic deviations from helix tracks, but that gives me only a small score improvement which is dwarfed by what you gain compared to me in the other steps.\n\nMy approach is pure Python currently and takes about 3 minutes per event, but I guess yours would be faster than mine if you'd scale  down the candidate lists to go down to my score. So I think you have the best basis for the throughput phase, too.",
    "370315": "Excellent Solution ! Congratulations and Thanks.",
    "370205": "Hi, how many threads do you use per event?  ",
    "370334": "@icecuber\n\nGreat idea. Thanks for sharing it and describing step by step. \nCongratulations! \n\nI thought you are working in IceCube collaboration ))) ",
    "376771": "Great work! Congratulations",
    "375534": "Congratulations and thank you for the solution!",
    "372970": "Excellent Solution ! Congratulations and Thanks.",
    "372699": "(y)",
    "372563": "Your solution shows how hard you worked for this win. Congratulations!",
    "372368": "@icecuber Well Done. Congratulations ",
    "370289": "@icecuber did you find any discrepancies in the cell vs truth data e.g. [event000001000/19932][1]? They are going to regenerate the dataset for the throughput phase so if there is a bug in cell data it would be good to flag this to them before they do it.\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/63027",
    "370124": "Great work!\nCongratulations @icecuber.",
    "370019": "Congratulations @icecuber and team. Thanks for sharing your so elegant solution.  Wow  \"... 8 minutes per event per cpu core...\", how impressive. \n\nI also want to give shout out to @yuval r and @heng for all they have shared in the discussions sections of this competition.",
    "369973": "Thoughtful approach. Well done!!",
    "369963": "Hi there, congratulations on your finish! Can you share your source code?\n\nThanks.",
    "369939": "Congrats! Really amazing work, thanks for sharing.",
    "369936": "At some stage of the competition, I thought that if someone could think of the way to grow tracks from hits in the global scope (like bacteria grows and duplicate themselves in the air). Now, the first solution is similar like this.  That's truly a wonder to behold. Thanks for sharing.",
    "369932": "This is great. Congratulations. \n\nRegarding using direction from cell's data: Was this crucial in your steps 1 and 2, or could you have done without it, at a cost of a few percentage points to the score? \n\nCalculating direction from the cells file seemed very interesting, but there was also quite some noise (direction not available for all hits, and also some real inaccuracies involved). You surely mastered to work with that data. Great work.",
    "771299": "Thank u for sharing",
    "374695": "Excellent! thank you for sharing",
    "372093": "Grt work! Thanks",
    "370059": "Amazing, thank you !",
    "369923": "Wow, thanks for sharing, and congrats."
  }
}