{
  "id": 63313,
  "title": "7# solution",
  "url": "/competitions/trackml-particle-identification/writeups/yuval-trian-7-solution",
  "author_name": "",
  "post_date": "2018-08-14T18:54:01.376526100Z",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Although I shared many of  our algorithm's basic ideas, now is the time to share the full solution\nI have posted a kernel with the major part of this solution. <a href=\"https://www.kaggle.com/yuval6967/7th-place-clustering-and-extending\">#7 place solution</a>\nRunning this kernel on training event 1000 will score ~0.635 after the clustering stage, and 0.735 after expending stage.\nEvery stage takes about 8-10 min on Kaggle, and about half the time on my laptop.\nIn the clustering part of the kernel the algorithm 5500 pairs of z0, 1/2R (More of it below) by increasing the number to about 100,000 the score will plateau at about 0.765 (after expending).\nHow does it work:\nIn each clustering loop the algorithm try to find all tracks originating from <code>(0,0,z0)</code> and with a radius of <code>1/(2*kt).</code>\nIf a hit (x,y,z) is on a track the helix can be fully defined by the following features (1), (2)</p>\n\n<pre><code>rr=(x**2+y**2)**0.5\ntheta_=arctan(y/x)\ndtheta = arcsin(kt*rr)\n(1) Theta=theta_+dtheta\n(2) (z-z0)*kt/dtheta\n</code></pre>\n\n<p>To solve the +pi,-pi problem we use sin, cos for theta.\n To make (2) more uniform, we use <code>arctan((z-z0)/(3.3*dtheta/kt))</code>\nAfter calculating the features, the algorithm tries to cluster all the hits with the same features. This is done by sparse binning – using np.unique.\nThe disadvantage of sparse binning over dbscan is it’s sensitivity, the advantages are its speed and its sensitivity (almost no outliners).\nAfter clustering every hit choose if his cluster is good according to the clusters length.\nEvery 500 loops all hits belonging to tracks which are long enough are removed from the dataset\nIf two hits from the same detector are on the same track, the one which is closest to the track’s center of mass is chosen.\nThe z0, kt pairs a chosen randomly\nWhile running, the algorithm changes the bin width and the length of the minimum track to be extracted from the dataset.</p>\n\n<p>Expending is done by selecting the un-clustered hits which are close to the center of mass of the track.</p>\n\n<p>To get better then 0.765, we merged a few long runs together, this was done by scoring the tracks with a ML algorithm Trian wrote (please share below).\nWe also gained sum points by clustering from outside of the origin, starting the track from a hit (it was very efficient and slow)</p>",
  "messages": [
    {
      "id": "370398",
      "postDate": "08/14/2018 18:54:01",
      "content": "<p>Although I shared many of  our algorithm's basic ideas, now is the time to share the full solution\nI have posted a kernel with the major part of this solution. <a href=\"https://www.kaggle.com/yuval6967/7th-place-clustering-and-extending\">#7 place solution</a>\nRunning this kernel on training event 1000 will score ~0.635 after the clustering stage, and 0.735 after expending stage.\nEvery stage takes about 8-10 min on Kaggle, and about half the time on my laptop.\nIn the clustering part of the kernel the algorithm 5500 pairs of z0, 1/2R (More of it below) by increasing the number to about 100,000 the score will plateau at about 0.765 (after expending).\nHow does it work:\nIn each clustering loop the algorithm try to find all tracks originating from <code>(0,0,z0)</code> and with a radius of <code>1/(2*kt).</code>\nIf a hit (x,y,z) is on a track the helix can be fully defined by the following features (1), (2)</p>\n\n<pre><code>rr=(x**2+y**2)**0.5\ntheta_=arctan(y/x)\ndtheta = arcsin(kt*rr)\n(1) Theta=theta_+dtheta\n(2) (z-z0)*kt/dtheta\n</code></pre>\n\n<p>To solve the +pi,-pi problem we use sin, cos for theta.\n To make (2) more uniform, we use <code>arctan((z-z0)/(3.3*dtheta/kt))</code>\nAfter calculating the features, the algorithm tries to cluster all the hits with the same features. This is done by sparse binning – using np.unique.\nThe disadvantage of sparse binning over dbscan is it’s sensitivity, the advantages are its speed and its sensitivity (almost no outliners).\nAfter clustering every hit choose if his cluster is good according to the clusters length.\nEvery 500 loops all hits belonging to tracks which are long enough are removed from the dataset\nIf two hits from the same detector are on the same track, the one which is closest to the track’s center of mass is chosen.\nThe z0, kt pairs a chosen randomly\nWhile running, the algorithm changes the bin width and the length of the minimum track to be extracted from the dataset.</p>\n\n<p>Expending is done by selecting the un-clustered hits which are close to the center of mass of the track.</p>\n\n<p>To get better then 0.765, we merged a few long runs together, this was done by scoring the tracks with a ML algorithm Trian wrote (please share below).\nWe also gained sum points by clustering from outside of the origin, starting the track from a hit (it was very efficient and slow)</p>",
      "rawMarkdown": "Although I shared many of  our algorithm's basic ideas, now is the time to share the full solution\nI have posted a kernel with the major part of this solution. [#7 place solution][1]\nRunning this kernel on training event 1000 will score ~0.635 after the clustering stage, and 0.735 after expending stage.\nEvery stage takes about 8-10 min on Kaggle, and about half the time on my laptop.\nIn the clustering part of the kernel the algorithm 5500 pairs of z0, 1/2R (More of it below) by increasing the number to about 100,000 the score will plateau at about 0.765 (after expending).\nHow does it work:\nIn each clustering loop the algorithm try to find all tracks originating from `(0,0,z0)` and with a radius of `1/(2*kt).`\nIf a hit (x,y,z) is on a track the helix can be fully defined by the following features (1), (2)\n\n    rr=(x**2+y**2)**0.5\n    theta_=arctan(y/x)\n    dtheta = arcsin(kt*rr)\n    (1)\tTheta=theta_+dtheta\n    (2)\t(z-z0)*kt/dtheta\n\nTo solve the +pi,-pi problem we use sin, cos for theta.\n To make (2) more uniform, we use `arctan((z-z0)/(3.3*dtheta/kt))`\nAfter calculating the features, the algorithm tries to cluster all the hits with the same features. This is done by sparse binning – using np.unique.\nThe disadvantage of sparse binning over dbscan is it’s sensitivity, the advantages are its speed and its sensitivity (almost no outliners).\nAfter clustering every hit choose if his cluster is good according to the clusters length.\nEvery 500 loops all hits belonging to tracks which are long enough are removed from the dataset\nIf two hits from the same detector are on the same track, the one which is closest to the track’s center of mass is chosen.\nThe z0, kt pairs a chosen randomly\nWhile running, the algorithm changes the bin width and the length of the minimum track to be extracted from the dataset.\n\nExpending is done by selecting the un-clustered hits which are close to the center of mass of the track.\n\nTo get better then 0.765, we merged a few long runs together, this was done by scoring the tracks with a ML algorithm Trian wrote (please share below).\nWe also gained sum points by clustering from outside of the origin, starting the track from a hit (it was very efficient and slow)\n\n\n  [1]: https://www.kaggle.com/yuval6967/7th-place-clustering-and-extending",
      "votes": null
    },
    {
      "id": "370504",
      "postDate": "08/14/2018 23:47:27",
      "content": "<p>Let me say a couple of things on the ML algorithm we employed. Spoiler alert: Nothing fancy happened here. </p>\n\n<p>I estimate the net gain of this ML algorithm to our <strong>final score to be maybe around +0.01</strong> (with potential for more if time allowed + it can speed up the submission generation). The \"problem\" is that it initially may give higher gains, but part of those gains can also be achieved using other post-processing steps, that we had in place.</p>\n\n<p>1) So, we started with <strong>8 submissions</strong> worth around 0.7x (obtained via clustering + post-processing). The concept is to merge these successively into 1 submission: Merge sub1 &amp; sub2 into sub12, then merge sub12 &amp; sub3 into sub123.... until sub12345678.</p>\n\n<p>2) We merge by giving <strong>probabilities to each track_id</strong>, and choosing for each hit the track_id, which has the highest probability to be correct. Those probabilities are calculated using <strong>LightGBM</strong> (several tests suggest that similar results would have been achieved using logistic regression, or a deep neural net). Actually, we added to the probabilities a fraction of the track-length, and after a couple of merges, we also asked the new probability of subX to be at least c higher then sub123...(X-1)</p>\n\n<p>Note: We also tried neural networks with a few hidden layers (and embeddings) and LSTM (where the input is a sequence of one value for each of the up to 20 hits, where the value was e.g. volume-layer-module id). The first model gave similar performance to the LightGBM and the latter model slightly worse, though still better than I thought would be possible by just using volume-layer-module id!</p>\n\n<p>3) <strong>13 features</strong> were used for each track:\n- The most important were: variance of x, variance of y\n- We also used z_var, z_mean, x_min, y_min, z_min, x_max, y_max, z_max, \n- volume_id of first hit (if we sort by absolute z-value with ascending order)\n- number_of_clusters of  track (calculated using dbscan with eps=10; 1 cluster = a small area where the track has many hits, think distance &lt;5 in each of x,y,z)\n- number_of_hits / number_of_clusters\nNote: We tried a lot more features, but they seemed to give only minor additional gains.</p>\n\n<p>4) Training data from roughly <strong>250 training events with roughly 5 million tracks</strong> were used:\n- Correct tracks (target=1): all true tracks, taken from truth files\n- Wrong tracks (target=0): all tracks from simulated 0.7x submissions (for those 100 events), which got score 0\nNote: Good results were already achieved with 50.000 training tracks, with diminishing returns after that. Training was a matter of minutes (some parameters: LightGBM with 3000 steps, learning_rate=0.05, 128 leafs). We also tried giving a \"purity score\" to each track (and \"objective=xentropy\" instead of \"binary\"), depending on how many percent of the full suggested track belong to the same particle (as suggested by Heng). This did not change things enough, to be further considered.</p>\n\n<p>5) Eventually, on the validation events (around 5-10, which were not used for training, but were created similarly), we achieved rates of <strong>roughly 95% for precision, accuracy and recall</strong>. This sounds quite good, BUT the big problem was, that <strong>the used correct/wrong tracks of our training/validation data did NOT RESEMBLE the final situation well</strong>, at which the model was employed. One step to improve this, was to take the \"purity score\" as described above, but this did not help our implementation.</p>\n\n<p>Final thoughts: <br>\nFor ML to be helpful in our situation, it needs to distinguish correct from wrong tracks with very high precision (tp/(fp+tp)) AND do so for various sets of track candidates (especially such, which are generated if one tries to find tracks which originate far away from the origin; in those situations, often a lot of bad candidates are produced).</p>\n\n<p>Eventually, besides speeding up submission generation, the best thing ML may be able to do here (as @CPMP said somewhere else I think) is to find out physics laws (or 3d detector architecture) just by feeding in data. So, on the one hand it may be better to use those physics rules directly. On the other hand, sometimes a lot of forces interact, and combining all those formulas and 3d detector architecture may be not (100%) feasible.</p>",
      "rawMarkdown": "Let me say a couple of things on the ML algorithm we employed. Spoiler alert: Nothing fancy happened here. \n\nI estimate the net gain of this ML algorithm to our **final score to be maybe around +0.01** (with potential for more if time allowed + it can speed up the submission generation). The \"problem\" is that it initially may give higher gains, but part of those gains can also be achieved using other post-processing steps, that we had in place.\n\n1) So, we started with **8 submissions** worth around 0.7x (obtained via clustering + post-processing). The concept is to merge these successively into 1 submission: Merge sub1 &amp; sub2 into sub12, then merge sub12 &amp; sub3 into sub123.... until sub12345678.\n\n2) We merge by giving **probabilities to each track_id**, and choosing for each hit the track_id, which has the highest probability to be correct. Those probabilities are calculated using **LightGBM** (several tests suggest that similar results would have been achieved using logistic regression, or a deep neural net). Actually, we added to the probabilities a fraction of the track-length, and after a couple of merges, we also asked the new probability of subX to be at least c higher then sub123...(X-1)\n\nNote: We also tried neural networks with a few hidden layers (and embeddings) and LSTM (where the input is a sequence of one value for each of the up to 20 hits, where the value was e.g. volume-layer-module id). The first model gave similar performance to the LightGBM and the latter model slightly worse, though still better than I thought would be possible by just using volume-layer-module id!\n\n3) **13 features** were used for each track:\n- The most important were: variance of x, variance of y\n- We also used z_var, z_mean, x_min, y_min, z_min, x_max, y_max, z_max, \n- volume_id of first hit (if we sort by absolute z-value with ascending order)\n- number_of_clusters of  track (calculated using dbscan with eps=10; 1 cluster = a small area where the track has many hits, think distance &lt;5 in each of x,y,z)\n- number_of_hits / number_of_clusters\nNote: We tried a lot more features, but they seemed to give only minor additional gains.\n\n4) Training data from roughly **250 training events with roughly 5 million tracks** were used:\n- Correct tracks (target=1): all true tracks, taken from truth files\n- Wrong tracks (target=0): all tracks from simulated 0.7x submissions (for those 100 events), which got score 0\nNote: Good results were already achieved with 50.000 training tracks, with diminishing returns after that. Training was a matter of minutes (some parameters: LightGBM with 3000 steps, learning_rate=0.05, 128 leafs). We also tried giving a \"purity score\" to each track (and \"objective=xentropy\" instead of \"binary\"), depending on how many percent of the full suggested track belong to the same particle (as suggested by Heng). This did not change things enough, to be further considered.\n\n5) Eventually, on the validation events (around 5-10, which were not used for training, but were created similarly), we achieved rates of **roughly 95% for precision, accuracy and recall**. This sounds quite good, BUT the big problem was, that **the used correct/wrong tracks of our training/validation data did NOT RESEMBLE the final situation well**, at which the model was employed. One step to improve this, was to take the \"purity score\" as described above, but this did not help our implementation.\n\n\nFinal thoughts:  \nFor ML to be helpful in our situation, it needs to distinguish correct from wrong tracks with very high precision (tp/(fp+tp)) AND do so for various sets of track candidates (especially such, which are generated if one tries to find tracks which originate far away from the origin; in those situations, often a lot of bad candidates are produced).\n\nEventually, besides speeding up submission generation, the best thing ML may be able to do here (as @CPMP said somewhere else I think) is to find out physics laws (or 3d detector architecture) just by feeding in data. So, on the one hand it may be better to use those physics rules directly. On the other hand, sometimes a lot of forces interact, and combining all those formulas and 3d detector architecture may be not (100%) feasible.",
      "votes": null
    },
    {
      "id": "370576",
      "postDate": "08/15/2018 04:10:51",
      "content": "<p>Thank you both for sharing, and congfrats on the result.  I see few things I could use to improve my own solution.  I thought our solutions were very similar, based on what yuval shared, but they are rather different actually.</p>",
      "rawMarkdown": "Thank you both for sharing, and congfrats on the result.  I see few things I could use to improve my own solution.  I thought our solutions were very similar, based on what yuval shared, but they are rather different actually.",
      "votes": null
    },
    {
      "id": "375812",
      "postDate": "08/26/2018 05:47:42",
      "content": "<p>We published our full detailed solution <a href=\"https://github.com/tx1985/kaggle-trackML/tree/master\">here</a>\nAnd a detailed description of the solution in PDF format <a href=\"https://github.com/tx1985/kaggle-trackML/blob/master/trackML_solution.pdf\">here</a></p>",
      "rawMarkdown": "We published our full detailed solution [here][1]\nAnd a detailed description of the solution in PDF format [here][2]\n\n\n  [1]: https://github.com/tx1985/kaggle-trackML/tree/master\n  [2]: https://github.com/tx1985/kaggle-trackML/blob/master/trackML_solution.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 370504,
      "author_name": "trian2018",
      "author_url": "",
      "post_date": "08/14/2018 23:47:27",
      "content": "<p>Let me say a couple of things on the ML algorithm we employed. Spoiler alert: Nothing fancy happened here. </p>\n\n<p>I estimate the net gain of this ML algorithm to our <strong>final score to be maybe around +0.01</strong> (with potential for more if time allowed + it can speed up the submission generation). The \"problem\" is that it initially may give higher gains, but part of those gains can also be achieved using other post-processing steps, that we had in place.</p>\n\n<p>1) So, we started with <strong>8 submissions</strong> worth around 0.7x (obtained via clustering + post-processing). The concept is to merge these successively into 1 submission: Merge sub1 &amp; sub2 into sub12, then merge sub12 &amp; sub3 into sub123.... until sub12345678.</p>\n\n<p>2) We merge by giving <strong>probabilities to each track_id</strong>, and choosing for each hit the track_id, which has the highest probability to be correct. Those probabilities are calculated using <strong>LightGBM</strong> (several tests suggest that similar results would have been achieved using logistic regression, or a deep neural net). Actually, we added to the probabilities a fraction of the track-length, and after a couple of merges, we also asked the new probability of subX to be at least c higher then sub123...(X-1)</p>\n\n<p>Note: We also tried neural networks with a few hidden layers (and embeddings) and LSTM (where the input is a sequence of one value for each of the up to 20 hits, where the value was e.g. volume-layer-module id). The first model gave similar performance to the LightGBM and the latter model slightly worse, though still better than I thought would be possible by just using volume-layer-module id!</p>\n\n<p>3) <strong>13 features</strong> were used for each track:\n- The most important were: variance of x, variance of y\n- We also used z_var, z_mean, x_min, y_min, z_min, x_max, y_max, z_max, \n- volume_id of first hit (if we sort by absolute z-value with ascending order)\n- number_of_clusters of  track (calculated using dbscan with eps=10; 1 cluster = a small area where the track has many hits, think distance &lt;5 in each of x,y,z)\n- number_of_hits / number_of_clusters\nNote: We tried a lot more features, but they seemed to give only minor additional gains.</p>\n\n<p>4) Training data from roughly <strong>250 training events with roughly 5 million tracks</strong> were used:\n- Correct tracks (target=1): all true tracks, taken from truth files\n- Wrong tracks (target=0): all tracks from simulated 0.7x submissions (for those 100 events), which got score 0\nNote: Good results were already achieved with 50.000 training tracks, with diminishing returns after that. Training was a matter of minutes (some parameters: LightGBM with 3000 steps, learning_rate=0.05, 128 leafs). We also tried giving a \"purity score\" to each track (and \"objective=xentropy\" instead of \"binary\"), depending on how many percent of the full suggested track belong to the same particle (as suggested by Heng). This did not change things enough, to be further considered.</p>\n\n<p>5) Eventually, on the validation events (around 5-10, which were not used for training, but were created similarly), we achieved rates of <strong>roughly 95% for precision, accuracy and recall</strong>. This sounds quite good, BUT the big problem was, that <strong>the used correct/wrong tracks of our training/validation data did NOT RESEMBLE the final situation well</strong>, at which the model was employed. One step to improve this, was to take the \"purity score\" as described above, but this did not help our implementation.</p>\n\n<p>Final thoughts: <br>\nFor ML to be helpful in our situation, it needs to distinguish correct from wrong tracks with very high precision (tp/(fp+tp)) AND do so for various sets of track candidates (especially such, which are generated if one tries to find tracks which originate far away from the origin; in those situations, often a lot of bad candidates are produced).</p>\n\n<p>Eventually, besides speeding up submission generation, the best thing ML may be able to do here (as @CPMP said somewhere else I think) is to find out physics laws (or 3d detector architecture) just by feeding in data. So, on the one hand it may be better to use those physics rules directly. On the other hand, sometimes a lot of forces interact, and combining all those formulas and 3d detector architecture may be not (100%) feasible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 370576,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 04:10:51",
          "content": "<p>Thank you both for sharing, and congfrats on the result.  I see few things I could use to improve my own solution.  I thought our solutions were very similar, based on what yuval shared, but they are rather different actually.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 375812,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "08/26/2018 05:47:42",
      "content": "<p>We published our full detailed solution <a href=\"https://github.com/tx1985/kaggle-trackML/tree/master\">here</a>\nAnd a detailed description of the solution in PDF format <a href=\"https://github.com/tx1985/kaggle-trackML/blob/master/trackML_solution.pdf\">here</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "370398": "Although I shared many of  our algorithm's basic ideas, now is the time to share the full solution\nI have posted a kernel with the major part of this solution. [#7 place solution][1]\nRunning this kernel on training event 1000 will score ~0.635 after the clustering stage, and 0.735 after expending stage.\nEvery stage takes about 8-10 min on Kaggle, and about half the time on my laptop.\nIn the clustering part of the kernel the algorithm 5500 pairs of z0, 1/2R (More of it below) by increasing the number to about 100,000 the score will plateau at about 0.765 (after expending).\nHow does it work:\nIn each clustering loop the algorithm try to find all tracks originating from `(0,0,z0)` and with a radius of `1/(2*kt).`\nIf a hit (x,y,z) is on a track the helix can be fully defined by the following features (1), (2)\n\n    rr=(x**2+y**2)**0.5\n    theta_=arctan(y/x)\n    dtheta = arcsin(kt*rr)\n    (1)\tTheta=theta_+dtheta\n    (2)\t(z-z0)*kt/dtheta\n\nTo solve the +pi,-pi problem we use sin, cos for theta.\n To make (2) more uniform, we use `arctan((z-z0)/(3.3*dtheta/kt))`\nAfter calculating the features, the algorithm tries to cluster all the hits with the same features. This is done by sparse binning – using np.unique.\nThe disadvantage of sparse binning over dbscan is it’s sensitivity, the advantages are its speed and its sensitivity (almost no outliners).\nAfter clustering every hit choose if his cluster is good according to the clusters length.\nEvery 500 loops all hits belonging to tracks which are long enough are removed from the dataset\nIf two hits from the same detector are on the same track, the one which is closest to the track’s center of mass is chosen.\nThe z0, kt pairs a chosen randomly\nWhile running, the algorithm changes the bin width and the length of the minimum track to be extracted from the dataset.\n\nExpending is done by selecting the un-clustered hits which are close to the center of mass of the track.\n\nTo get better then 0.765, we merged a few long runs together, this was done by scoring the tracks with a ML algorithm Trian wrote (please share below).\nWe also gained sum points by clustering from outside of the origin, starting the track from a hit (it was very efficient and slow)\n\n\n  [1]: https://www.kaggle.com/yuval6967/7th-place-clustering-and-extending",
    "370504": "Let me say a couple of things on the ML algorithm we employed. Spoiler alert: Nothing fancy happened here. \n\nI estimate the net gain of this ML algorithm to our **final score to be maybe around +0.01** (with potential for more if time allowed + it can speed up the submission generation). The \"problem\" is that it initially may give higher gains, but part of those gains can also be achieved using other post-processing steps, that we had in place.\n\n1) So, we started with **8 submissions** worth around 0.7x (obtained via clustering + post-processing). The concept is to merge these successively into 1 submission: Merge sub1 &amp; sub2 into sub12, then merge sub12 &amp; sub3 into sub123.... until sub12345678.\n\n2) We merge by giving **probabilities to each track_id**, and choosing for each hit the track_id, which has the highest probability to be correct. Those probabilities are calculated using **LightGBM** (several tests suggest that similar results would have been achieved using logistic regression, or a deep neural net). Actually, we added to the probabilities a fraction of the track-length, and after a couple of merges, we also asked the new probability of subX to be at least c higher then sub123...(X-1)\n\nNote: We also tried neural networks with a few hidden layers (and embeddings) and LSTM (where the input is a sequence of one value for each of the up to 20 hits, where the value was e.g. volume-layer-module id). The first model gave similar performance to the LightGBM and the latter model slightly worse, though still better than I thought would be possible by just using volume-layer-module id!\n\n3) **13 features** were used for each track:\n- The most important were: variance of x, variance of y\n- We also used z_var, z_mean, x_min, y_min, z_min, x_max, y_max, z_max, \n- volume_id of first hit (if we sort by absolute z-value with ascending order)\n- number_of_clusters of  track (calculated using dbscan with eps=10; 1 cluster = a small area where the track has many hits, think distance &lt;5 in each of x,y,z)\n- number_of_hits / number_of_clusters\nNote: We tried a lot more features, but they seemed to give only minor additional gains.\n\n4) Training data from roughly **250 training events with roughly 5 million tracks** were used:\n- Correct tracks (target=1): all true tracks, taken from truth files\n- Wrong tracks (target=0): all tracks from simulated 0.7x submissions (for those 100 events), which got score 0\nNote: Good results were already achieved with 50.000 training tracks, with diminishing returns after that. Training was a matter of minutes (some parameters: LightGBM with 3000 steps, learning_rate=0.05, 128 leafs). We also tried giving a \"purity score\" to each track (and \"objective=xentropy\" instead of \"binary\"), depending on how many percent of the full suggested track belong to the same particle (as suggested by Heng). This did not change things enough, to be further considered.\n\n5) Eventually, on the validation events (around 5-10, which were not used for training, but were created similarly), we achieved rates of **roughly 95% for precision, accuracy and recall**. This sounds quite good, BUT the big problem was, that **the used correct/wrong tracks of our training/validation data did NOT RESEMBLE the final situation well**, at which the model was employed. One step to improve this, was to take the \"purity score\" as described above, but this did not help our implementation.\n\n\nFinal thoughts:  \nFor ML to be helpful in our situation, it needs to distinguish correct from wrong tracks with very high precision (tp/(fp+tp)) AND do so for various sets of track candidates (especially such, which are generated if one tries to find tracks which originate far away from the origin; in those situations, often a lot of bad candidates are produced).\n\nEventually, besides speeding up submission generation, the best thing ML may be able to do here (as @CPMP said somewhere else I think) is to find out physics laws (or 3d detector architecture) just by feeding in data. So, on the one hand it may be better to use those physics rules directly. On the other hand, sometimes a lot of forces interact, and combining all those formulas and 3d detector architecture may be not (100%) feasible.",
    "370576": "Thank you both for sharing, and congfrats on the result.  I see few things I could use to improve my own solution.  I thought our solutions were very similar, based on what yuval shared, but they are rather different actually.",
    "375812": "We published our full detailed solution [here][1]\nAnd a detailed description of the solution in PDF format [here][2]\n\n\n  [1]: https://github.com/tx1985/kaggle-trackML/tree/master\n  [2]: https://github.com/tx1985/kaggle-trackML/blob/master/trackML_solution.pdf"
  },
  "source": "meta"
}