{
  "id": 63330,
  "title": "3rd place solution",
  "url": "/competitions/trackml-particle-identification/writeups/sergey-gorbunov-3rd-place-solution",
  "author_name": "",
  "post_date": "2018-08-19T19:54:33.830Z",
  "votes": 14,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello everyone and thank you for the nice competition!</p>\n\n<p>Unfortunately I joined it late and didn't follow discussions on the forum, I had to concentrate on solving the task.</p>\n\n<p>I developed a combinatorial algorithm, which is very similar to the 1-st place algorithm from icecuber. </p>\n\n<p>In the amount of time I had I only managed to finish a combinatorial \"engine\" of the algorithm a day before the deadline. During the last day I just tried to run this engine many times more or less chaotically  everywhere in the detector. The resulting code looks therefore horrible and is extremely slow. But the engine itself seems to work, the rest still needs to be developed.</p>\n\n<p>Ok. The engine consist of the two parts: </p>\n\n<p>-the first part converts hits to short tracklets in some local part of the detector. It's goal is to kill the hit-to-hit combinatorics in order to work later on on some structured data.</p>\n\n<p>-the second part prolongates the tracklets through the detector and collects their hits. </p>\n\n<p>I run this engine several times in different parts of the detector with different constraints.\nThen I sort all the found track candidates in some simple way and choose the best tracks. Hits which belong to the found tracks I remove from the plane. Then the next round of the tracklet search starts in other detector parts or with other constraints.\nThis way I clean up the data slice by slice, so to say, until nothing is left on the plane. </p>\n\n<p>The engine. </p>\n\n<ol>\n<li>The tracklet constructor.</li>\n</ol>\n\n<p>There are 2 options here. </p>\n\n<p>option a) \nIt creates two-hit tracklets which are constrained to the event vertex (which is pretty much (0,0) point in XY and +-2.5 cm in Z )\nUsing the vertex constraint is the trick here. It significantly reduces amount of possible tracklet candidates. And, as 75% of the tracks are coming from the vertex, one can really clean up the data by removing the vertex tracks before doing anything else.</p>\n\n<p>option b) \nIt creates 3-hit tracklets without vertex constraint.  </p>\n\n<p>The layers where the tracklets are created are predefined by the main program. Unlike icecuber I decided to not create all possible tracklets everywhere but rather save the compute time by developing some smart seeding sequence strategy, which is still need to be developed:) </p>\n\n<ol>\n<li>The tracklet prolongation. Similar to the winner algorithm, I don't have a global trajectory. </li>\n</ol>\n\n<p>To prolongate a track to the next layer, I use a local helix created by last 3 hits of the track. </p>\n\n<p>Amazingly, it works pretty well. I think this is due to very precise measurements in silicon. One can follow all the local features of magnetic field and trajectory scattering in the material, and complicated energy losses, and god knows what else without even knowing the value of the magnetic field!  And it doesn't cost any cpu time.</p>\n\n<p>But what I found, the magnetic field is varying  dramatically in the detector, from 20 kGaus to -10!!. Especially between the detector volumes. I realised it when I was checking - how much the local curvature changes along a track. Oh yes, it changes. </p>\n\n<p>To investigate this, I have fitted the magnetic field value using neighbouring truth points and truth momentum vectors. </p>\n\n<p>Once I realised that the field is non-constant in many regions, I decided to modify my track model. </p>\n\n<p>It is still a helix, but it is parameterized not with its geometrical radius, but with a physical parameter Pt (transverse momentum). \nThese are just proportional: Pt = B*r.   Now, having 1) the physical parameterisation and 2) the magnetic field values, I can fit the trajectory with one field value (i.e. inside a radial volume), but then prolongate it using another field value (i.e the value between radial and forward volumes). </p>\n\n<p>Or, in the other words, I scale the helix radius  according to the field change. (New Radius=OldRadius*NewField/OldField)</p>\n\n<p>The magnetic field I parametrised for each detector layer individually using some polynoms. \nHere is my formula for the field: \nB(z,phi) = (c0+c1*z) + (c2+c3*z)*sin(phi) + (c4+c5*z)*cos(phi).\nHere z, phi are  hit angular and z coordinates on a layer, coefficients I have fitted with old good LSM method. The approximation is  not very accurate, maybe one can replace it with just an average field value on a layer.  To save the time during track search, I calculate the filed value for every hit and store it in the hit structure before the search starts. </p>\n\n<p>As I remember, use of the physical model improved my accuracy of prolongation, but I can't tell now how big was the improvement. May be at the end it is not needed at all.</p>\n\n<p>One big problem I have here - I have to manually set cuts for picking up hits. For the proof-of-concept it was fine, but then I wind up with a huge file with copy-pasted and slightly modified hardcoded numbers.  My plan was to have this hits picking-up cuts to be set automatically from the test data. But due to the naive trajectory model, hit deviations from  trajectories are not nicely distributed. At the end I had to look to every distribution and decide where to cut it. One should do something about that.</p>\n\n<p>The search strategy. </p>\n\n<p>First, I search tracks which are coming from the vertex, then the other tracks. First I find high-momentum tracks (applying angular and momentum cuts on the tracklets), then the low-momentum tracks. </p>\n\n<p>Well, that is pretty much the algorithm. </p>\n\n<p>Thank again for the nice competition!</p>\n\n<p>Here is a link to the code: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker\">https://github.com/sgorbuno/TrackML_CombinatorialTracker</a>\nHere is a description: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf\">https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf</a></p>",
  "messages": [
    {
      "id": "370522",
      "postDate": "08/15/2018 00:49:07",
      "content": "<p>Hello everyone and thank you for the nice competition!</p>\n\n<p>Unfortunately I joined it late and didn't follow discussions on the forum, I had to concentrate on solving the task.</p>\n\n<p>I developed a combinatorial algorithm, which is very similar to the 1-st place algorithm from icecuber. </p>\n\n<p>In the amount of time I had I only managed to finish a combinatorial \"engine\" of the algorithm a day before the deadline. During the last day I just tried to run this engine many times more or less chaotically  everywhere in the detector. The resulting code looks therefore horrible and is extremely slow. But the engine itself seems to work, the rest still needs to be developed.</p>\n\n<p>Ok. The engine consist of the two parts: </p>\n\n<p>-the first part converts hits to short tracklets in some local part of the detector. It's goal is to kill the hit-to-hit combinatorics in order to work later on on some structured data.</p>\n\n<p>-the second part prolongates the tracklets through the detector and collects their hits. </p>\n\n<p>I run this engine several times in different parts of the detector with different constraints.\nThen I sort all the found track candidates in some simple way and choose the best tracks. Hits which belong to the found tracks I remove from the plane. Then the next round of the tracklet search starts in other detector parts or with other constraints.\nThis way I clean up the data slice by slice, so to say, until nothing is left on the plane. </p>\n\n<p>The engine. </p>\n\n<ol>\n<li>The tracklet constructor.</li>\n</ol>\n\n<p>There are 2 options here. </p>\n\n<p>option a) \nIt creates two-hit tracklets which are constrained to the event vertex (which is pretty much (0,0) point in XY and +-2.5 cm in Z )\nUsing the vertex constraint is the trick here. It significantly reduces amount of possible tracklet candidates. And, as 75% of the tracks are coming from the vertex, one can really clean up the data by removing the vertex tracks before doing anything else.</p>\n\n<p>option b) \nIt creates 3-hit tracklets without vertex constraint.  </p>\n\n<p>The layers where the tracklets are created are predefined by the main program. Unlike icecuber I decided to not create all possible tracklets everywhere but rather save the compute time by developing some smart seeding sequence strategy, which is still need to be developed:) </p>\n\n<ol>\n<li>The tracklet prolongation. Similar to the winner algorithm, I don't have a global trajectory. </li>\n</ol>\n\n<p>To prolongate a track to the next layer, I use a local helix created by last 3 hits of the track. </p>\n\n<p>Amazingly, it works pretty well. I think this is due to very precise measurements in silicon. One can follow all the local features of magnetic field and trajectory scattering in the material, and complicated energy losses, and god knows what else without even knowing the value of the magnetic field!  And it doesn't cost any cpu time.</p>\n\n<p>But what I found, the magnetic field is varying  dramatically in the detector, from 20 kGaus to -10!!. Especially between the detector volumes. I realised it when I was checking - how much the local curvature changes along a track. Oh yes, it changes. </p>\n\n<p>To investigate this, I have fitted the magnetic field value using neighbouring truth points and truth momentum vectors. </p>\n\n<p>Once I realised that the field is non-constant in many regions, I decided to modify my track model. </p>\n\n<p>It is still a helix, but it is parameterized not with its geometrical radius, but with a physical parameter Pt (transverse momentum). \nThese are just proportional: Pt = B*r.   Now, having 1) the physical parameterisation and 2) the magnetic field values, I can fit the trajectory with one field value (i.e. inside a radial volume), but then prolongate it using another field value (i.e the value between radial and forward volumes). </p>\n\n<p>Or, in the other words, I scale the helix radius  according to the field change. (New Radius=OldRadius*NewField/OldField)</p>\n\n<p>The magnetic field I parametrised for each detector layer individually using some polynoms. \nHere is my formula for the field: \nB(z,phi) = (c0+c1*z) + (c2+c3*z)*sin(phi) + (c4+c5*z)*cos(phi).\nHere z, phi are  hit angular and z coordinates on a layer, coefficients I have fitted with old good LSM method. The approximation is  not very accurate, maybe one can replace it with just an average field value on a layer.  To save the time during track search, I calculate the filed value for every hit and store it in the hit structure before the search starts. </p>\n\n<p>As I remember, use of the physical model improved my accuracy of prolongation, but I can't tell now how big was the improvement. May be at the end it is not needed at all.</p>\n\n<p>One big problem I have here - I have to manually set cuts for picking up hits. For the proof-of-concept it was fine, but then I wind up with a huge file with copy-pasted and slightly modified hardcoded numbers.  My plan was to have this hits picking-up cuts to be set automatically from the test data. But due to the naive trajectory model, hit deviations from  trajectories are not nicely distributed. At the end I had to look to every distribution and decide where to cut it. One should do something about that.</p>\n\n<p>The search strategy. </p>\n\n<p>First, I search tracks which are coming from the vertex, then the other tracks. First I find high-momentum tracks (applying angular and momentum cuts on the tracklets), then the low-momentum tracks. </p>\n\n<p>Well, that is pretty much the algorithm. </p>\n\n<p>Thank again for the nice competition!</p>\n\n<p>Here is a link to the code: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker\">https://github.com/sgorbuno/TrackML_CombinatorialTracker</a>\nHere is a description: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf\">https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf</a></p>",
      "rawMarkdown": "Hello everyone and thank you for the nice competition!\n\n\nUnfortunately I joined it late and didn't follow discussions on the forum, I had to concentrate on solving the task.\n\nI developed a combinatorial algorithm, which is very similar to the 1-st place algorithm from icecuber. \n\nIn the amount of time I had I only managed to finish a combinatorial \"engine\" of the algorithm a day before the deadline. During the last day I just tried to run this engine many times more or less chaotically  everywhere in the detector. The resulting code looks therefore horrible and is extremely slow. But the engine itself seems to work, the rest still needs to be developed.\n\n\nOk. The engine consist of the two parts: \n\n-the first part converts hits to short tracklets in some local part of the detector. It's goal is to kill the hit-to-hit combinatorics in order to work later on on some structured data.\n\n-the second part prolongates the tracklets through the detector and collects their hits. \n\n\nI run this engine several times in different parts of the detector with different constraints.\nThen I sort all the found track candidates in some simple way and choose the best tracks. Hits which belong to the found tracks I remove from the plane. Then the next round of the tracklet search starts in other detector parts or with other constraints.\nThis way I clean up the data slice by slice, so to say, until nothing is left on the plane. \n\n\nThe engine. \n\n\n1. The tracklet constructor.\n\nThere are 2 options here. \n\noption a) \nIt creates two-hit tracklets which are constrained to the event vertex (which is pretty much (0,0) point in XY and +-2.5 cm in Z )\nUsing the vertex constraint is the trick here. It significantly reduces amount of possible tracklet candidates. And, as 75% of the tracks are coming from the vertex, one can really clean up the data by removing the vertex tracks before doing anything else.\n\noption b) \nIt creates 3-hit tracklets without vertex constraint.  \n\nThe layers where the tracklets are created are predefined by the main program. Unlike icecuber I decided to not create all possible tracklets everywhere but rather save the compute time by developing some smart seeding sequence strategy, which is still need to be developed:) \n\n\n2. The tracklet prolongation. Similar to the winner algorithm, I don't have a global trajectory. \n\nTo prolongate a track to the next layer, I use a local helix created by last 3 hits of the track. \n\nAmazingly, it works pretty well. I think this is due to very precise measurements in silicon. One can follow all the local features of magnetic field and trajectory scattering in the material, and complicated energy losses, and god knows what else without even knowing the value of the magnetic field!  And it doesn't cost any cpu time.\n\nBut what I found, the magnetic field is varying  dramatically in the detector, from 20 kGaus to -10!!. Especially between the detector volumes. I realised it when I was checking - how much the local curvature changes along a track. Oh yes, it changes. \n\nTo investigate this, I have fitted the magnetic field value using neighbouring truth points and truth momentum vectors. \n\nOnce I realised that the field is non-constant in many regions, I decided to modify my track model. \n\nIt is still a helix, but it is parameterized not with its geometrical radius, but with a physical parameter Pt (transverse momentum). \nThese are just proportional: Pt = B*r.   Now, having 1) the physical parameterisation and 2) the magnetic field values, I can fit the trajectory with one field value (i.e. inside a radial volume), but then prolongate it using another field value (i.e the value between radial and forward volumes). \n\nOr, in the other words, I scale the helix radius  according to the field change. (New Radius=OldRadius*NewField/OldField)\n\nThe magnetic field I parametrised for each detector layer individually using some polynoms. \nHere is my formula for the field: \nB(z,phi) = (c0+c1*z) + (c2+c3*z)*sin(phi) + (c4+c5*z)*cos(phi).\nHere z, phi are  hit angular and z coordinates on a layer, coefficients I have fitted with old good LSM method. The approximation is  not very accurate, maybe one can replace it with just an average field value on a layer.  To save the time during track search, I calculate the filed value for every hit and store it in the hit structure before the search starts. \n\nAs I remember, use of the physical model improved my accuracy of prolongation, but I can't tell now how big was the improvement. May be at the end it is not needed at all.\n\nOne big problem I have here - I have to manually set cuts for picking up hits. For the proof-of-concept it was fine, but then I wind up with a huge file with copy-pasted and slightly modified hardcoded numbers.  My plan was to have this hits picking-up cuts to be set automatically from the test data. But due to the naive trajectory model, hit deviations from  trajectories are not nicely distributed. At the end I had to look to every distribution and decide where to cut it. One should do something about that.\n\n\nThe search strategy. \n\nFirst, I search tracks which are coming from the vertex, then the other tracks. First I find high-momentum tracks (applying angular and momentum cuts on the tracklets), then the low-momentum tracks. \n\n\nWell, that is pretty much the algorithm. \n\nThank again for the nice competition!\n\nHere is a link to the code: https://github.com/sgorbuno/TrackML_CombinatorialTracker\nHere is a description: https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf",
      "votes": null
    },
    {
      "id": "370581",
      "postDate": "08/15/2018 04:16:18",
      "content": "<p>Wow, congrats.  I wish we were given the magnetic field, because, as you noticed, it varies much more than what one could expect from data description.  But I didn't know how to estimate it from the data as you did.  You clearly seem to have some physics background.  May I ask what you do for a living?</p>",
      "rawMarkdown": "Wow, congrats.  I wish we were given the magnetic field, because, as you noticed, it varies much more than what one could expect from data description.  But I didn't know how to estimate it from the data as you did.  You clearly seem to have some physics background.  May I ask what you do for a living?",
      "votes": null
    },
    {
      "id": "370587",
      "postDate": "08/15/2018 04:38:31",
      "content": "<p>Afterthought: the variation in magnetic field may be due to a change in direction, as hinted at in the CERN document shared at the start of competition, rather than intensity.  Have you taken this into account when computing magnetic field intensity?</p>",
      "rawMarkdown": "Afterthought: the variation in magnetic field may be due to a change in direction, as hinted at in the CERN document shared at the start of competition, rather than intensity.  Have you taken this into account when computing magnetic field intensity?",
      "votes": null
    },
    {
      "id": "370626",
      "postDate": "08/15/2018 06:57:27",
      "content": "<p>Thanks for sharing. Congrats on your last minute dash to the 3rd place.\nreading about all the deviation  and correction to the magnetic field, I'm surprised my solution was able to work, as I assumed the helix is perfect. </p>",
      "rawMarkdown": "Thanks for sharing. Congrats on your last minute dash to the 3rd place.\nreading about all the deviation  and correction to the magnetic field, I'm surprised my solution was able to work, as I assumed the helix is perfect.",
      "votes": null
    },
    {
      "id": "370637",
      "postDate": "08/15/2018 07:30:33",
      "content": "<blockquote>\n  <p>I'm surprised my solution was able to work, as I assumed the helix is perfect. </p>\n</blockquote>\n\n<p>That's why your binning output is not that high ;)  You had to compensate with rather smart post processing.</p>",
      "rawMarkdown": "&gt; I'm surprised my solution was able to work, as I assumed the helix is perfect. \n\nThat's why your binning output is not that high ;)  You had to compensate with rather smart post processing.",
      "votes": null
    },
    {
      "id": "370645",
      "postDate": "08/15/2018 07:39:51",
      "content": "<p>I do get to 0.725 from the binning alone and my post processing also assums perfect helix. It does allow some more distance in the distant volumes</p>",
      "rawMarkdown": "I do get to 0.725 from the binning alone and my post processing also assums perfect helix. It does allow some more distance in the distant volumes",
      "votes": null
    },
    {
      "id": "370651",
      "postDate": "08/15/2018 07:51:43",
      "content": "<blockquote>\n  <p>I use a local helix created by last 3 hits of the track.\n  Amazingly, it works pretty well. I think this is due to very precise measurements in silicon.</p>\n</blockquote>\n\n<p>Indeed. I also only ever used the latest 3 layer crossings to fit the helix parameters for track extension. Any attempt to include older hits lowered my score, from which I concluded that the loss from mixing in the older pre-scattering information is larger than the gain from averaging out random measurement errors (the latter are quite small, and extremely small in the innermost layers).</p>",
      "rawMarkdown": "&gt; I use a local helix created by last 3 hits of the track.\n&gt; Amazingly, it works pretty well. I think this is due to very precise measurements in silicon.\n\nIndeed. I also only ever used the latest 3 layer crossings to fit the helix parameters for track extension. Any attempt to include older hits lowered my score, from which I concluded that the loss from mixing in the older pre-scattering information is larger than the gain from averaging out random measurement errors (the latter are quite small, and extremely small in the innermost layers).",
      "votes": null
    },
    {
      "id": "370697",
      "postDate": "08/15/2018 10:03:06",
      "content": "<p>Hi, thanks!</p>\n\n<p>Yes, I'm working on the online tracking in ALICE experiment. It has  completely different detector - it is one huge gas volume without any material inside. We have relatively unprecise measurements, but up to 160 of them per track  and almost constant field and no scattering. We need to collect all the measurements along the trajectory in order to get good track position/momentum estimation. Here the situation is different. I get better knowledge about local trajectory when I ignore all other the other measurements.  Interesting. </p>\n\n<p>Concerning the missing field description - I think the organisers wanted to give ML algorithms  a head start, because they can just learn the missing field by training (I'm not an expert here),  and limit others in looking for some simple approaches - tricks and hacks. Perhaps they don't want to see here the Kalman Filter monsters as they already have them :)</p>\n\n<p>But finding back the field is easy. You fit a circle, get its radius. Then you divide the radius by the truth momentum and you get the field. </p>",
      "rawMarkdown": "Hi, thanks!\n\n  Yes, I'm working on the online tracking in ALICE experiment. It has  completely different detector - it is one huge gas volume without any material inside. We have relatively unprecise measurements, but up to 160 of them per track  and almost constant field and no scattering. We need to collect all the measurements along the trajectory in order to get good track position/momentum estimation. Here the situation is different. I get better knowledge about local trajectory when I ignore all other the other measurements.  Interesting. \n\n Concerning the missing field description - I think the organisers wanted to give ML algorithms  a head start, because they can just learn the missing field by training (I'm not an expert here),  and limit others in looking for some simple approaches - tricks and hacks. Perhaps they don't want to see here the Kalman Filter monsters as they already have them :)\n\nBut finding back the field is easy. You fit a circle, get its radius. Then you divide the radius by the truth momentum and you get the field.",
      "votes": null
    },
    {
      "id": "370699",
      "postDate": "08/15/2018 10:08:26",
      "content": "<p>Changing the field direction - yes, I also thought about this.  It shouldn't be difficult  to implement - one just rotates the track before the prolongation. But it will slow down everything on the other side.. I think, may be one should do it in some certain regions.. </p>",
      "rawMarkdown": "Changing the field direction - yes, I also thought about this.  It shouldn't be difficult  to implement - one just rotates the track before the prolongation. But it will slow down everything on the other side.. I think, may be one should do it in some certain regions..",
      "votes": null
    },
    {
      "id": "370705",
      "postDate": "08/15/2018 10:14:29",
      "content": "<p>Exactly. I think, if the measurement errors would be 0, it won't not change much. Tracks are measured with the infinite precision but then they scattered like hell right after the measurement. </p>",
      "rawMarkdown": "Exactly. I think, if the measurement errors would be 0, it won't not change much. Tracks are measured with the infinite precision but then they scattered like hell right after the measurement.",
      "votes": null
    },
    {
      "id": "370731",
      "postDate": "08/15/2018 11:03:52",
      "content": "<p>Indeed, we discussed this lengthly and we decided not to give the field description, because we didn't want to get a solution that re-implements what we are doing including the 1001st numerical integration algorithm for a magnetic field integration.</p>",
      "rawMarkdown": "Indeed, we discussed this lengthly and we decided not to give the field description, because we didn't want to get a solution that re-implements what we are doing including the 1001st numerical integration algorithm for a magnetic field integration.",
      "votes": null
    },
    {
      "id": "370756",
      "postDate": "08/15/2018 12:14:53",
      "content": "<p>The ML model Trian developed does not assume perfect helix.  You need to relax that assumption at a point if you want to move over 0.8.</p>",
      "rawMarkdown": "The ML model Trian developed does not assume perfect helix.  You need to relax that assumption at a point if you want to move over 0.8.",
      "votes": null
    },
    {
      "id": "370759",
      "postDate": "08/15/2018 12:17:38",
      "content": "<p>@Andrea, makes sense.</p>\n\n<p>@Serguey, my point about filed direction was different.  You say its intensity could be half of what it is at the vertex, but I think it rather is more aligned with particle track.  Now that I see how you compute intensity you compute the projection of the field on the z axis, which is not the same as the overall intensity.</p>",
      "rawMarkdown": "Andrea, makes sense.\n\n@Serguey, my point about filed direction was different.  You say its intensity could be half of what it is at the vertex, but I think it rather is more aligned with particle track.  Now that I see how you compute intensity you compute the projection of the field on the z axis, which is not the same as the overall intensity.",
      "votes": null
    },
    {
      "id": "370762",
      "postDate": "08/15/2018 12:21:58",
      "content": "<p>@CPMP\nBy the way, why do you write 'Serguey' (it's not the first time). I've never seen writing my name this way. :) \nThe French variant is 'Serge'  I think.</p>",
      "rawMarkdown": "CPMP\nBy the way, why do you write 'Serguey' (it's not the first time). I've never seen writing my name this way. :) \nThe French variant is 'Serge'  I think.",
      "votes": null
    },
    {
      "id": "370789",
      "postDate": "08/15/2018 13:10:30",
      "content": "<p>Yes, I fit the z-component of the field only.  </p>\n\n<p>Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation. It is kind of a parameter in the track model, which is calibrated on the truth data. Every hit has this calibrated parameter. Actually, it has three different parameters - one  for forward prolongation, one for backward prolongation, and one for helix construction. I forgot to mention that.  </p>\n\n<p>This parameters are calibrated i advance. I think it is very similar to what ML should do. ( but as i said, i'm not an expert here. I  first need to learn the machine learning before making such statements )</p>",
      "rawMarkdown": "Yes, I fit the z-component of the field only.  \n\nBasically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation. It is kind of a parameter in the track model, which is calibrated on the truth data. Every hit has this calibrated parameter. Actually, it has three different parameters - one  for forward prolongation, one for backward prolongation, and one for helix construction. I forgot to mention that.  \n\nThis parameters are calibrated i advance. I think it is very similar to what ML should do. ( but as i said, i'm not an expert here. I  first need to learn the machine learning before making such statements )",
      "votes": null
    },
    {
      "id": "370800",
      "postDate": "08/15/2018 13:28:38",
      "content": "<blockquote>\n  <p>why do you write 'Serguey'</p>\n</blockquote>\n\n<p>I imagine that a native French speaker could need the 'u' to approximate the non-French pronunciation of the 'g'. :-)</p>",
      "rawMarkdown": "&gt; why do you write 'Serguey'\n\nI imagine that a native French speaker could need the 'u' to approximate the non-French pronunciation of the 'g'. :-)",
      "votes": null
    },
    {
      "id": "370804",
      "postDate": "08/15/2018 13:34:48",
      "content": "<blockquote>\n  <p>why do you write 'Serguey'</p>\n</blockquote>\n\n<p>I'm just bad at name spelling, sorry for that.  Thanks for the reminder, I'll try to write it correctly next time.</p>\n\n<p>Edwin is right though, in French a 'g' is always followed by a 'u' when the 'g' is pronounced the hard way (as opposed to be pronounced as a 'j')</p>",
      "rawMarkdown": "&gt; why do you write 'Serguey'\n\nI'm just bad at name spelling, sorry for that.  Thanks for the reminder, I'll try to write it correctly next time.\n\nEdwin is right though, in French a 'g' is always followed by a 'u' when the 'g' is pronounced the hard way (as opposed to be pronounced as a 'j')",
      "votes": null
    },
    {
      "id": "370808",
      "postDate": "08/15/2018 13:38:01",
      "content": "<blockquote>\n  <p>Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation</p>\n</blockquote>\n\n<p>I agree, and I treated it as such as well.  </p>",
      "rawMarkdown": "&gt; Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation\n\nI agree, and I treated it as such as well.",
      "votes": null
    },
    {
      "id": "372591",
      "postDate": "08/19/2018 19:54:48",
      "content": "<p>Here is a link to the code: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker\">https://github.com/sgorbuno/TrackML_CombinatorialTracker</a>\nHere is a description: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf\">https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf</a></p>",
      "rawMarkdown": "Here is a link to the code: https://github.com/sgorbuno/TrackML_CombinatorialTracker\nHere is a description: https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 370581,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/15/2018 04:16:18",
      "content": "<p>Wow, congrats.  I wish we were given the magnetic field, because, as you noticed, it varies much more than what one could expect from data description.  But I didn't know how to estimate it from the data as you did.  You clearly seem to have some physics background.  May I ask what you do for a living?</p>",
      "votes": null,
      "replies": [
        {
          "id": 370587,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 04:38:31",
          "content": "<p>Afterthought: the variation in magnetic field may be due to a change in direction, as hinted at in the CERN document shared at the start of competition, rather than intensity.  Have you taken this into account when computing magnetic field intensity?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370697,
          "author_name": "sgorbuno",
          "author_url": "",
          "post_date": "08/15/2018 10:03:06",
          "content": "<p>Hi, thanks!</p>\n\n<p>Yes, I'm working on the online tracking in ALICE experiment. It has  completely different detector - it is one huge gas volume without any material inside. We have relatively unprecise measurements, but up to 160 of them per track  and almost constant field and no scattering. We need to collect all the measurements along the trajectory in order to get good track position/momentum estimation. Here the situation is different. I get better knowledge about local trajectory when I ignore all other the other measurements.  Interesting. </p>\n\n<p>Concerning the missing field description - I think the organisers wanted to give ML algorithms  a head start, because they can just learn the missing field by training (I'm not an expert here),  and limit others in looking for some simple approaches - tricks and hacks. Perhaps they don't want to see here the Kalman Filter monsters as they already have them :)</p>\n\n<p>But finding back the field is easy. You fit a circle, get its radius. Then you divide the radius by the truth momentum and you get the field. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370699,
          "author_name": "sgorbuno",
          "author_url": "",
          "post_date": "08/15/2018 10:08:26",
          "content": "<p>Changing the field direction - yes, I also thought about this.  It shouldn't be difficult  to implement - one just rotates the track before the prolongation. But it will slow down everything on the other side.. I think, may be one should do it in some certain regions.. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370731,
          "author_name": "asalzburger",
          "author_url": "",
          "post_date": "08/15/2018 11:03:52",
          "content": "<p>Indeed, we discussed this lengthly and we decided not to give the field description, because we didn't want to get a solution that re-implements what we are doing including the 1001st numerical integration algorithm for a magnetic field integration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370759,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 12:17:38",
          "content": "<p>@Andrea, makes sense.</p>\n\n<p>@Serguey, my point about filed direction was different.  You say its intensity could be half of what it is at the vertex, but I think it rather is more aligned with particle track.  Now that I see how you compute intensity you compute the projection of the field on the z axis, which is not the same as the overall intensity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370762,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/15/2018 12:21:58",
          "content": "<p>@CPMP\nBy the way, why do you write 'Serguey' (it's not the first time). I've never seen writing my name this way. :) \nThe French variant is 'Serge'  I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370789,
          "author_name": "sgorbuno",
          "author_url": "",
          "post_date": "08/15/2018 13:10:30",
          "content": "<p>Yes, I fit the z-component of the field only.  </p>\n\n<p>Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation. It is kind of a parameter in the track model, which is calibrated on the truth data. Every hit has this calibrated parameter. Actually, it has three different parameters - one  for forward prolongation, one for backward prolongation, and one for helix construction. I forgot to mention that.  </p>\n\n<p>This parameters are calibrated i advance. I think it is very similar to what ML should do. ( but as i said, i'm not an expert here. I  first need to learn the machine learning before making such statements )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370800,
          "author_name": "edwinst",
          "author_url": "",
          "post_date": "08/15/2018 13:28:38",
          "content": "<blockquote>\n  <p>why do you write 'Serguey'</p>\n</blockquote>\n\n<p>I imagine that a native French speaker could need the 'u' to approximate the non-French pronunciation of the 'g'. :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370804,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 13:34:48",
          "content": "<blockquote>\n  <p>why do you write 'Serguey'</p>\n</blockquote>\n\n<p>I'm just bad at name spelling, sorry for that.  Thanks for the reminder, I'll try to write it correctly next time.</p>\n\n<p>Edwin is right though, in French a 'g' is always followed by a 'u' when the 'g' is pronounced the hard way (as opposed to be pronounced as a 'j')</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370808,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 13:38:01",
          "content": "<blockquote>\n  <p>Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation</p>\n</blockquote>\n\n<p>I agree, and I treated it as such as well.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 370626,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "08/15/2018 06:57:27",
      "content": "<p>Thanks for sharing. Congrats on your last minute dash to the 3rd place.\nreading about all the deviation  and correction to the magnetic field, I'm surprised my solution was able to work, as I assumed the helix is perfect. </p>",
      "votes": null,
      "replies": [
        {
          "id": 370637,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 07:30:33",
          "content": "<blockquote>\n  <p>I'm surprised my solution was able to work, as I assumed the helix is perfect. </p>\n</blockquote>\n\n<p>That's why your binning output is not that high ;)  You had to compensate with rather smart post processing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370645,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "08/15/2018 07:39:51",
          "content": "<p>I do get to 0.725 from the binning alone and my post processing also assums perfect helix. It does allow some more distance in the distant volumes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 370756,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2018 12:14:53",
          "content": "<p>The ML model Trian developed does not assume perfect helix.  You need to relax that assumption at a point if you want to move over 0.8.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 370651,
      "author_name": "edwinst",
      "author_url": "",
      "post_date": "08/15/2018 07:51:43",
      "content": "<blockquote>\n  <p>I use a local helix created by last 3 hits of the track.\n  Amazingly, it works pretty well. I think this is due to very precise measurements in silicon.</p>\n</blockquote>\n\n<p>Indeed. I also only ever used the latest 3 layer crossings to fit the helix parameters for track extension. Any attempt to include older hits lowered my score, from which I concluded that the loss from mixing in the older pre-scattering information is larger than the gain from averaging out random measurement errors (the latter are quite small, and extremely small in the innermost layers).</p>",
      "votes": null,
      "replies": [
        {
          "id": 370705,
          "author_name": "sgorbuno",
          "author_url": "",
          "post_date": "08/15/2018 10:14:29",
          "content": "<p>Exactly. I think, if the measurement errors would be 0, it won't not change much. Tracks are measured with the infinite precision but then they scattered like hell right after the measurement. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372591,
      "author_name": "sgorbuno",
      "author_url": "",
      "post_date": "08/19/2018 19:54:48",
      "content": "<p>Here is a link to the code: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker\">https://github.com/sgorbuno/TrackML_CombinatorialTracker</a>\nHere is a description: <a href=\"https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf\">https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "370522": "Hello everyone and thank you for the nice competition!\n\n\nUnfortunately I joined it late and didn't follow discussions on the forum, I had to concentrate on solving the task.\n\nI developed a combinatorial algorithm, which is very similar to the 1-st place algorithm from icecuber. \n\nIn the amount of time I had I only managed to finish a combinatorial \"engine\" of the algorithm a day before the deadline. During the last day I just tried to run this engine many times more or less chaotically  everywhere in the detector. The resulting code looks therefore horrible and is extremely slow. But the engine itself seems to work, the rest still needs to be developed.\n\n\nOk. The engine consist of the two parts: \n\n-the first part converts hits to short tracklets in some local part of the detector. It's goal is to kill the hit-to-hit combinatorics in order to work later on on some structured data.\n\n-the second part prolongates the tracklets through the detector and collects their hits. \n\n\nI run this engine several times in different parts of the detector with different constraints.\nThen I sort all the found track candidates in some simple way and choose the best tracks. Hits which belong to the found tracks I remove from the plane. Then the next round of the tracklet search starts in other detector parts or with other constraints.\nThis way I clean up the data slice by slice, so to say, until nothing is left on the plane. \n\n\nThe engine. \n\n\n1. The tracklet constructor.\n\nThere are 2 options here. \n\noption a) \nIt creates two-hit tracklets which are constrained to the event vertex (which is pretty much (0,0) point in XY and +-2.5 cm in Z )\nUsing the vertex constraint is the trick here. It significantly reduces amount of possible tracklet candidates. And, as 75% of the tracks are coming from the vertex, one can really clean up the data by removing the vertex tracks before doing anything else.\n\noption b) \nIt creates 3-hit tracklets without vertex constraint.  \n\nThe layers where the tracklets are created are predefined by the main program. Unlike icecuber I decided to not create all possible tracklets everywhere but rather save the compute time by developing some smart seeding sequence strategy, which is still need to be developed:) \n\n\n2. The tracklet prolongation. Similar to the winner algorithm, I don't have a global trajectory. \n\nTo prolongate a track to the next layer, I use a local helix created by last 3 hits of the track. \n\nAmazingly, it works pretty well. I think this is due to very precise measurements in silicon. One can follow all the local features of magnetic field and trajectory scattering in the material, and complicated energy losses, and god knows what else without even knowing the value of the magnetic field!  And it doesn't cost any cpu time.\n\nBut what I found, the magnetic field is varying  dramatically in the detector, from 20 kGaus to -10!!. Especially between the detector volumes. I realised it when I was checking - how much the local curvature changes along a track. Oh yes, it changes. \n\nTo investigate this, I have fitted the magnetic field value using neighbouring truth points and truth momentum vectors. \n\nOnce I realised that the field is non-constant in many regions, I decided to modify my track model. \n\nIt is still a helix, but it is parameterized not with its geometrical radius, but with a physical parameter Pt (transverse momentum). \nThese are just proportional: Pt = B*r.   Now, having 1) the physical parameterisation and 2) the magnetic field values, I can fit the trajectory with one field value (i.e. inside a radial volume), but then prolongate it using another field value (i.e the value between radial and forward volumes). \n\nOr, in the other words, I scale the helix radius  according to the field change. (New Radius=OldRadius*NewField/OldField)\n\nThe magnetic field I parametrised for each detector layer individually using some polynoms. \nHere is my formula for the field: \nB(z,phi) = (c0+c1*z) + (c2+c3*z)*sin(phi) + (c4+c5*z)*cos(phi).\nHere z, phi are  hit angular and z coordinates on a layer, coefficients I have fitted with old good LSM method. The approximation is  not very accurate, maybe one can replace it with just an average field value on a layer.  To save the time during track search, I calculate the filed value for every hit and store it in the hit structure before the search starts. \n\nAs I remember, use of the physical model improved my accuracy of prolongation, but I can't tell now how big was the improvement. May be at the end it is not needed at all.\n\nOne big problem I have here - I have to manually set cuts for picking up hits. For the proof-of-concept it was fine, but then I wind up with a huge file with copy-pasted and slightly modified hardcoded numbers.  My plan was to have this hits picking-up cuts to be set automatically from the test data. But due to the naive trajectory model, hit deviations from  trajectories are not nicely distributed. At the end I had to look to every distribution and decide where to cut it. One should do something about that.\n\n\nThe search strategy. \n\nFirst, I search tracks which are coming from the vertex, then the other tracks. First I find high-momentum tracks (applying angular and momentum cuts on the tracklets), then the low-momentum tracks. \n\n\nWell, that is pretty much the algorithm. \n\nThank again for the nice competition!\n\nHere is a link to the code: https://github.com/sgorbuno/TrackML_CombinatorialTracker\nHere is a description: https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf",
    "370581": "Wow, congrats.  I wish we were given the magnetic field, because, as you noticed, it varies much more than what one could expect from data description.  But I didn't know how to estimate it from the data as you did.  You clearly seem to have some physics background.  May I ask what you do for a living?",
    "370587": "Afterthought: the variation in magnetic field may be due to a change in direction, as hinted at in the CERN document shared at the start of competition, rather than intensity.  Have you taken this into account when computing magnetic field intensity?",
    "370626": "Thanks for sharing. Congrats on your last minute dash to the 3rd place.\nreading about all the deviation  and correction to the magnetic field, I'm surprised my solution was able to work, as I assumed the helix is perfect.",
    "370637": "&gt; I'm surprised my solution was able to work, as I assumed the helix is perfect. \n\nThat's why your binning output is not that high ;)  You had to compensate with rather smart post processing.",
    "370645": "I do get to 0.725 from the binning alone and my post processing also assums perfect helix. It does allow some more distance in the distant volumes",
    "370651": "&gt; I use a local helix created by last 3 hits of the track.\n&gt; Amazingly, it works pretty well. I think this is due to very precise measurements in silicon.\n\nIndeed. I also only ever used the latest 3 layer crossings to fit the helix parameters for track extension. Any attempt to include older hits lowered my score, from which I concluded that the loss from mixing in the older pre-scattering information is larger than the gain from averaging out random measurement errors (the latter are quite small, and extremely small in the innermost layers).",
    "370697": "Hi, thanks!\n\n  Yes, I'm working on the online tracking in ALICE experiment. It has  completely different detector - it is one huge gas volume without any material inside. We have relatively unprecise measurements, but up to 160 of them per track  and almost constant field and no scattering. We need to collect all the measurements along the trajectory in order to get good track position/momentum estimation. Here the situation is different. I get better knowledge about local trajectory when I ignore all other the other measurements.  Interesting. \n\n Concerning the missing field description - I think the organisers wanted to give ML algorithms  a head start, because they can just learn the missing field by training (I'm not an expert here),  and limit others in looking for some simple approaches - tricks and hacks. Perhaps they don't want to see here the Kalman Filter monsters as they already have them :)\n\nBut finding back the field is easy. You fit a circle, get its radius. Then you divide the radius by the truth momentum and you get the field.",
    "370699": "Changing the field direction - yes, I also thought about this.  It shouldn't be difficult  to implement - one just rotates the track before the prolongation. But it will slow down everything on the other side.. I think, may be one should do it in some certain regions..",
    "370705": "Exactly. I think, if the measurement errors would be 0, it won't not change much. Tracks are measured with the infinite precision but then they scattered like hell right after the measurement.",
    "370731": "Indeed, we discussed this lengthly and we decided not to give the field description, because we didn't want to get a solution that re-implements what we are doing including the 1001st numerical integration algorithm for a magnetic field integration.",
    "370756": "The ML model Trian developed does not assume perfect helix.  You need to relax that assumption at a point if you want to move over 0.8.",
    "370759": "Andrea, makes sense.\n\n@Serguey, my point about filed direction was different.  You say its intensity could be half of what it is at the vertex, but I think it rather is more aligned with particle track.  Now that I see how you compute intensity you compute the projection of the field on the z axis, which is not the same as the overall intensity.",
    "370762": "CPMP\nBy the way, why do you write 'Serguey' (it's not the first time). I've never seen writing my name this way. :) \nThe French variant is 'Serge'  I think.",
    "370789": "Yes, I fit the z-component of the field only.  \n\nBasically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation. It is kind of a parameter in the track model, which is calibrated on the truth data. Every hit has this calibrated parameter. Actually, it has three different parameters - one  for forward prolongation, one for backward prolongation, and one for helix construction. I forgot to mention that.  \n\nThis parameters are calibrated i advance. I think it is very similar to what ML should do. ( but as i said, i'm not an expert here. I  first need to learn the machine learning before making such statements )",
    "370800": "&gt; why do you write 'Serguey'\n\nI imagine that a native French speaker could need the 'u' to approximate the non-French pronunciation of the 'g'. :-)",
    "370804": "&gt; why do you write 'Serguey'\n\nI'm just bad at name spelling, sorry for that.  Thanks for the reminder, I'll try to write it correctly next time.\n\nEdwin is right though, in French a 'g' is always followed by a 'u' when the 'g' is pronounced the hard way (as opposed to be pronounced as a 'j')",
    "370808": "&gt; Basically, one can forget about the magnetic field and consider all this as some fixed scaling factor for helix radius during the prolongation\n\nI agree, and I treated it as such as well.",
    "372591": "Here is a link to the code: https://github.com/sgorbuno/TrackML_CombinatorialTracker\nHere is a description: https://github.com/sgorbuno/TrackML_CombinatorialTracker/blob/master/doc/TrackML_AlgorithmDescription.pdf"
  },
  "source": "meta"
}