{
  "id": 59635,
  "title": "A better ensembling method",
  "url": "/competitions/trackml-particle-identification/discussion/59635",
  "author_name": "hengck23",
  "post_date": "2018-06-25T09:45:04.722000",
  "votes": 10,
  "comment_count": 44,
  "views": 0,
  "content": "<p>See the code for details. There is no LB submission for this code yet.</p>\n\n<ol>\n<li><p>run dbscan for different z offset, angle offset: z(r)= (z+delta)/r;   a(r) = a + delta*r</p></li>\n<li><p>store all clusters from different dbscan results. Sort them increasing length (cluster count)</p></li>\n<li><p>assign hits according to the filtering criterion in the code.</p></li>\n</ol>\n\n<hr>\n\n<p>example results:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/347757/9687/ensmble.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li>this does not use track extension yet. </li>\n<li>this does not fix angle discontinuity yet</li>\n<li>this code can be used for finding seeds of straight tracks</li>\n</ul>",
  "messages": [
    {
      "id": 347757,
      "postDate": "2018-06-25T09:45:04.723Z",
      "content": "<p>See the code for details. There is no LB submission for this code yet.</p>\n\n<ol>\n<li><p>run dbscan for different z offset, angle offset: z(r)= (z+delta)/r;   a(r) = a + delta*r</p></li>\n<li><p>store all clusters from different dbscan results. Sort them increasing length (cluster count)</p></li>\n<li><p>assign hits according to the filtering criterion in the code.</p></li>\n</ol>\n\n<hr>\n\n<p>example results:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/347757/9687/ensmble.png\" alt=\"enter image description here\"></p>\n\n<ul>\n<li>this does not use track extension yet. </li>\n<li>this does not fix angle discontinuity yet</li>\n<li>this code can be used for finding seeds of straight tracks</li>\n</ul>",
      "rawMarkdown": "See the code for details. There is no LB submission for this code yet.\n\n1. run dbscan for different z offset, angle offset: z(r)= (z+delta)/r;   a(r) = a + delta*r\n\n2. store all clusters from different dbscan results. Sort them increasing length (cluster count)\n\n3. assign hits according to the filtering criterion in the code.\n\n\n---\n\nexample results:\n\n\n\n  ![enter image description here][1]\n\n\n- this does not use track extension yet. \n- this does not fix angle discontinuity yet\n- this code can be used for finding seeds of straight tracks\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/347757/9687/ensmble.png",
      "votes": 10
    },
    {
      "id": 347950,
      "postDate": "2018-06-25T18:51:38.580Z",
      "content": "<p>@cpmp</p>\n\n<p>You can imagine that i am doing a ransac for line fitting. The cluster count is the number of inliners. Hence, assign hits to the line with the largest number of inliners </p>",
      "rawMarkdown": "@cpmp\n\nYou can imagine that i am doing a ransac for line fitting. The cluster count is the number of inliners. Hence, assign hits to the line with the largest number of inliners ",
      "votes": 1,
      "replies": [
        {
          "id": 347995,
          "postDate": "2018-06-25T21:07:34.490Z",
          "content": "<p>For those (like me) who aren't in the know, ransac is a method of determining a best line of fit by excluding some outlier points. </p>\n\n<p>For a track that might be very well fit for 9 hits, and a final hit that is questionably on the track, that track will effectively be considered of length 9, so a new track with 10 hits that for the track well will cause the hits to be reassigned to the new track. </p>\n\n<p><a href=\"https://en.m.wikipedia.org/wiki/Random_sample_consensus\">https://en.m.wikipedia.org/wiki/Random_sample_consensus</a></p>\n\n<p>That being said, from reading the code and reading how ransac works (from a theoretical perspective), I think there is something I'm missing... Looks like I have some homework to do. </p>",
          "rawMarkdown": "For those (like me) who aren't in the know, ransac is a method of determining a best line of fit by excluding some outlier points. \n\nFor a track that might be very well fit for 9 hits, and a final hit that is questionably on the track, that track will effectively be considered of length 9, so a new track with 10 hits that for the track well will cause the hits to be reassigned to the new track. \n\nhttps://en.m.wikipedia.org/wiki/Random_sample_consensus\n\nThat being said, from reading the code and reading how ransac works (from a theoretical perspective), I think there is something I'm missing... Looks like I have some homework to do. ",
          "votes": 2
        },
        {
          "id": 348050,
          "postDate": "2018-06-26T01:31:08.677Z",
          "content": "<blockquote>\n  <p>Hence, assign hits to the line with the largest number of inliners </p>\n</blockquote>\n\n<p>Your code does something different: for a candidate, if a point is assigned to an existing cluster larger than it then you throw away the whole candidate.  I don't.  </p>\n\n<p>I've tried your filtering and it degrades a bit my local score.</p>",
          "rawMarkdown": "&gt; Hence, assign hits to the line with the largest number of inliners \n\nYour code does something different: for a candidate, if a point is assigned to an existing cluster larger than it then you throw away the whole candidate.  I don't.  \n\nI've tried your filtering and it degrades a bit my local score."
        }
      ]
    },
    {
      "id": 347775,
      "postDate": "2018-06-25T10:41:45.690Z",
      "content": "<p>@CPMP</p>\n\n<p>this is for short tracks only (e.g. across 4 layer_ids).</p>\n\n<p>E.g you have divide all hits into different volume_id, and then choose a few layer_ids, to ensure tracklets are straight.</p>\n\n<p>Different volume_id may use different parameters</p>",
      "rawMarkdown": "@CPMP\n\n\nthis is for short tracks only (e.g. across 4 layer_ids).\n\nE.g you have divide all hits into different volume_id, and then choose a few layer_ids, to ensure tracklets are straight.\n\nDifferent volume_id may use different parameters",
      "votes": 1
    },
    {
      "id": 355240,
      "postDate": "2018-07-11T09:11:03.450Z",
      "content": "<p>@Heng, Looking at the values of dj, I guess your explore.py was written before you realized, that the beam spot σz  is not equal 55 mm but much less. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315</a></p>",
      "rawMarkdown": "@Heng, Looking at the values of dj, I guess your explore.py was written before you realized, that the beam spot σz  is not equal 55 mm but much less. https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315",
      "votes": 2,
      "replies": [
        {
          "id": 355246,
          "postDate": "2018-07-11T09:33:10.223Z",
          "content": "<p>Grzegorz, your remark about it in another thread made me revisit my code too.  It yields a significant improvement at minimal cost.  Thanks a lot for that.  A sub is underway from me (won't be close to top, but encouraging when resuming after 15 days of absence from this competition).</p>",
          "rawMarkdown": "Grzegorz, your remark about it in another thread made me revisit my code too.  It yields a significant improvement at minimal cost.  Thanks a lot for that.  A sub is underway from me (won't be close to top, but encouraging when resuming after 15 days of absence from this competition)."
        },
        {
          "id": 355275,
          "postDate": "2018-07-11T11:01:57.377Z",
          "content": "<p>@Grzegorz, @CPMP thank you so much!!! I got it wrong too. </p>",
          "rawMarkdown": "@Grzegorz, @CPMP thank you so much!!! I got it wrong too. "
        },
        {
          "id": 355277,
          "postDate": "2018-07-11T11:15:49.760Z",
          "content": "<p>This is probably another daft question, but where is beam spot σz defined? I read the correction - but I couldn't find where you saw the original 55 mm?</p>",
          "rawMarkdown": "This is probably another daft question, but where is beam spot σz defined? I read the correction - but I couldn't find where you saw the original 55 mm?"
        },
        {
          "id": 355280,
          "postDate": "2018-07-11T11:22:47.407Z",
          "content": "<p>There is a link in \"Welcome from the organizers! read this first!\" to:</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a>\npage 6</p>",
          "rawMarkdown": "There is a link in \"Welcome from the organizers! read this first!\" to:\n\nhttps://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\npage 6",
          "votes": 2
        },
        {
          "id": 355283,
          "postDate": "2018-07-11T11:30:51.880Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        },
        {
          "id": 355427,
          "postDate": "2018-07-11T16:59:59.023Z",
          "content": "<blockquote>\n  <p>Grzegorz, your remark about it in another thread made me revisit my code too</p>\n</blockquote>\n\n<p>@CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.</p>",
          "rawMarkdown": "&gt;Grzegorz, your remark about it in another thread made me revisit my code too\n\n@CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.",
          "votes": 1
        },
        {
          "id": 355447,
          "postDate": "2018-07-11T17:46:52.917Z",
          "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>Thank you for the comment. you are right. The dj values are determined by trial and errors. </p>\n\n<p>Currently, I am putting DBSCA aside and focus on deep learning to learn pairwise links.\n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60447\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60447</a></p>\n\n<p>I can understand now why <a href=\"/outrunner\">@outrunner</a> tracks  are of good quality. </p>",
          "rawMarkdown": "@Grzegorz Sionkowski\n \nThank you for the comment. you are right. The dj values are determined by trial and errors. \n\nCurrently, I am putting DBSCA aside and focus on deep learning to learn pairwise links.\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/60447\n\nI can understand now why @outrunner tracks  are of good quality. \n"
        },
        {
          "id": 355574,
          "postDate": "2018-07-12T02:31:42.673Z",
          "content": "<blockquote>\n  <p>@CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.</p>\n</blockquote>\n\n<p>Well, I also forgot about it.  It was good to move away from this for 2 weeks as I now have a fresh look at it.  And your posts contain many thought provocative insights, thanks again for that.</p>",
          "rawMarkdown": "&gt; @CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.\n\nWell, I also forgot about it.  It was good to move away from this for 2 weeks as I now have a fresh look at it.  And your posts contain many thought provocative insights, thanks again for that."
        },
        {
          "id": 355748,
          "postDate": "2018-07-12T09:34:40.953Z",
          "content": "<p>@Grzegorz, @CPMP,</p>\n\n<blockquote>\n  <p>that the beam spot σz is not equal 55 mm but much less.</p>\n</blockquote>\n\n<p>I read it first in the slides linked by @CPMP. It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.</p>\n\n<p>I don't understand why the organizers downplay this difference so much (except for personal reasons, of course), since it significantly decreases the geometric \"mixing\" of trajectories in the inner layers, doesn't it? The difference between 55mm and 5.5mm is comparable to or even larger than dimensions of the innermost detector structures. This is highly significant geometrically, I'd say. Even worse, it will affect different solution approaches differently. E.g. I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.</p>",
          "rawMarkdown": "@Grzegorz, @CPMP,\n\n&gt; that the beam spot σz is not equal 55 mm but much less.\n\nI read it first in the slides linked by @CPMP. It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.\n\nI don't understand why the organizers downplay this difference so much (except for personal reasons, of course), since it significantly decreases the geometric \"mixing\" of trajectories in the inner layers, doesn't it? The difference between 55mm and 5.5mm is comparable to or even larger than dimensions of the innermost detector structures. This is highly significant geometrically, I'd say. Even worse, it will affect different solution approaches differently. E.g. I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.",
          "votes": 2
        },
        {
          "id": 355791,
          "postDate": "2018-07-12T11:19:00.350Z",
          "content": "<blockquote>\n  <p>It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.</p>\n</blockquote>\n\n<p>I had the same observation, and did spent time trying to find out why this was not working as it should but I didn't check the validity of the description provided to us.  That's why i thank Grzegorz here.</p>\n\n<blockquote>\n  <p>I don't understand why the organizers downplay this difference so much</p>\n</blockquote>\n\n<p>They are not the ones sweating while reverse engineering the simulator.</p>\n\n<blockquote>\n  <p>I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.</p>\n</blockquote>\n\n<p>Indeed.  At the width of the beam in x,y plane is also 10x smaller than described, which explains why approaches focusing on tracks originating from x,y = 0,0 work surprisingly well.</p>",
          "rawMarkdown": "&gt; It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.\n\nI had the same observation, and did spent time trying to find out why this was not working as it should but I didn't check the validity of the description provided to us.  That's why i thank Grzegorz here.\n\n&gt; I don't understand why the organizers downplay this difference so much\n\nThey are not the ones sweating while reverse engineering the simulator.\n\n&gt; I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.\n\nIndeed.  At the width of the beam in x,y plane is also 10x smaller than described, which explains why approaches focusing on tracks originating from x,y = 0,0 work surprisingly well."
        }
      ]
    },
    {
      "id": 351495,
      "postDate": "2018-07-02T10:06:59.763Z",
      "content": "<p>@Heng I analyzed your code last night and the highest score I could get is still lower than doing dbscan on the full event hits, this is the score on the event 000001003 in the volume 7 and 9 using your code (with some of my hyper parameters to maximize your score)</p>\n\n<pre><code>max_score = df.weight.sum() = 0.28114\nscore1= 0.24930  (0.88674)\nscore2= 0.24930  (0.88674)\n</code></pre>\n\n<p>Using the full labels found by the current best dbscan model  (with the score 0.675)</p>\n\n<pre><code>max_score = df.weight.sum() = 0.28114\nscore1= 0.25874  (0.92033)\nscore2= 0.25874  (0.92033)\n</code></pre>\n\n<p>It indicates that this seeding approach may only find a subset of the straight tracks that the full-scan dbscan model can find. My takeaway is that if your dbscan model's score is over <code>0.6</code>, your model may outperform this approach in terms of finding seeds from the volume 7 and 9, the challenge is to find seeds of the tracks in the volume 8 and outside the pixel detectors, regardless of being straight or not. Of course, seeding is not necessary for this competition, we can do without, just seeding is a standard approach used by CERN. </p>",
      "rawMarkdown": "@Heng I analyzed your code last night and the highest score I could get is still lower than doing dbscan on the full event hits, this is the score on the event 000001003 in the volume 7 and 9 using your code (with some of my hyper parameters to maximize your score)\n\n    max_score = df.weight.sum() = 0.28114\n    score1= 0.24930  (0.88674)\n    score2= 0.24930  (0.88674)\n\n\nUsing the full labels found by the current best dbscan model  (with the score 0.675)\n\n    max_score = df.weight.sum() = 0.28114\n    score1= 0.25874  (0.92033)\n    score2= 0.25874  (0.92033)\n\n\nIt indicates that this seeding approach may only find a subset of the straight tracks that the full-scan dbscan model can find. My takeaway is that if your dbscan model's score is over `0.6`, your model may outperform this approach in terms of finding seeds from the volume 7 and 9, the challenge is to find seeds of the tracks in the volume 8 and outside the pixel detectors, regardless of being straight or not. Of course, seeding is not necessary for this competition, we can do without, just seeding is a standard approach used by CERN. ",
      "votes": 2,
      "replies": [
        {
          "id": 351652,
          "postDate": "2018-07-02T18:23:59.407Z",
          "content": "<p>When working with shifting on the z axis, I've found that limiting the max length of tracks gives me a worse score than allowing full tracks, which I figured would be counter intuitive because of tracks that originate further from the origin tend to be shorter. </p>\n\n<p>I have a guess as to what the limitation is, but I'm not sure how to test it. With the extension code adding hits one at a time, a short track being extended iteratively will only work until the first hit that differs too greatly from the previous pair. </p>\n\n<p>So if hits 1, 2,... 8 are all in a track, and you only seed hits 2, 3, and 4, but 5 is noisy, you might not be able to extend to 6, while all 8 might be detected with dbscan. </p>\n\n<p>It is worth noting that my local scores still aren't at .6x yet. </p>\n\n<p>Now that I'm writing all this I think I need to compare two submissions and find tracks where the dbscan can outperform an extended short track, and look at only that track... </p>",
          "rawMarkdown": "When working with shifting on the z axis, I've found that limiting the max length of tracks gives me a worse score than allowing full tracks, which I figured would be counter intuitive because of tracks that originate further from the origin tend to be shorter. \n\nI have a guess as to what the limitation is, but I'm not sure how to test it. With the extension code adding hits one at a time, a short track being extended iteratively will only work until the first hit that differs too greatly from the previous pair. \n\nSo if hits 1, 2,... 8 are all in a track, and you only seed hits 2, 3, and 4, but 5 is noisy, you might not be able to extend to 6, while all 8 might be detected with dbscan. \n\nIt is worth noting that my local scores still aren't at .6x yet. \n\nNow that I'm writing all this I think I need to compare two submissions and find tracks where the dbscan can outperform an extended short track, and look at only that track... ",
          "votes": 1
        },
        {
          "id": 355069,
          "postDate": "2018-07-10T21:32:53.647Z",
          "content": "<p>I am getting an error, when I try to use @Heng's explore.py script. Analyzing it I am not sure what the error is about. I think it is in the scoring function. See below:-</p>\n\n<p>&gt; TypeError: merge() got an unexpected keyword argument 'validate'</p>\n\n<p>@Nicole, @macfarII can you help?</p>",
          "rawMarkdown": "I am getting an error, when I try to use @Heng's explore.py script. Analyzing it I am not sure what the error is about. I think it is in the scoring function. See below:-\n\n&gt; TypeError: merge() got an unexpected keyword argument 'validate'\n\n@Nicole, @macfarII can you help?"
        },
        {
          "id": 355074,
          "postDate": "2018-07-10T21:48:10.460Z",
          "content": "<p>We've seen this before, my guess is that you should  update to pandas latest version.</p>",
          "rawMarkdown": "We've seen this before, my guess is that you should  update to pandas latest version.",
          "votes": 2
        },
        {
          "id": 355078,
          "postDate": "2018-07-10T21:53:10.527Z",
          "content": "<p>I can confirm that fixed it for me...</p>",
          "rawMarkdown": "I can confirm that fixed it for me...",
          "votes": 1
        },
        {
          "id": 355081,
          "postDate": "2018-07-10T22:05:36.563Z",
          "content": "<p>Thanks @David and @John, that worked.</p>",
          "rawMarkdown": "Thanks @David and @John, that worked."
        }
      ]
    },
    {
      "id": 348308,
      "postDate": "2018-06-26T13:55:28.853Z",
      "content": "<p>here are some ideas for your reference:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9703/e5.png\" alt=\"enter image description here\">\n   typo error:  ... supervised (or unsupervised)  learning ... coordinates should be a,r,z/r and not x,y,z</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9701/e2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9702/e3.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "here are some ideas for your reference:\n\n\n  ![enter image description here][1]\n   typo error:  ... supervised (or unsupervised)  learning ... coordinates should be a,r,z/r and not x,y,z\n\n  ![enter image description here][2]\n\n  ![enter image description here][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9703/e5.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9701/e2.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9702/e3.png",
      "replies": [
        {
          "id": 348365,
          "postDate": "2018-06-26T15:25:27.357Z",
          "content": "<p>You're saying that for linear tracks, the track in transformed coordinates is directional, and all hits are in sequence with respect to layer id. </p>\n\n<p>Shouldn't a helix in transformed space should also be directional and form a curve,  each hit is in sequence with respect to layer id?</p>\n\n<p>From looking at your secret sauce, you're using z as thevway of determining if the layers are in the correct order... </p>\n\n<p>Would it make more sense to use the straightened track, and then evaluate that z is in the correct order, in that x, y are a function of z/layer id?</p>\n\n<p>I'm not sure why a track candidate would have hits that have an order of layer id that is different than the order of layer id...</p>",
          "rawMarkdown": "You're saying that for linear tracks, the track in transformed coordinates is directional, and all hits are in sequence with respect to layer id. \n\nShouldn't a helix in transformed space should also be directional and form a curve,  each hit is in sequence with respect to layer id?\n\nFrom looking at your secret sauce, you're using z as thevway of determining if the layers are in the correct order... \n\nWould it make more sense to use the straightened track, and then evaluate that z is in the correct order, in that x, y are a function of z/layer id?\n\nI'm not sure why a track candidate would have hits that have an order of layer id that is different than the order of layer id..."
        },
        {
          "id": 348366,
          "postDate": "2018-06-26T15:29:06.087Z",
          "content": "<p>I am using a, r, z/r to represent the hits. In this coordinate system, tracks are almost piecewise linear</p>",
          "rawMarkdown": "I am using a, r, z/r to represent the hits. In this coordinate system, tracks are almost piecewise linear"
        },
        {
          "id": 348369,
          "postDate": "2018-06-26T15:35:50.217Z",
          "content": "<p>Would a track have r and z in the same order, which would remove the need for layer information?</p>\n\n<p>I guess my question was what does the layer id information contribute that the coordinates don't already have?</p>",
          "rawMarkdown": "Would a track have r and z in the same order, which would remove the need for layer information?\n\nI guess my question was what does the layer id information contribute that the coordinates don't already have?"
        },
        {
          "id": 348371,
          "postDate": "2018-06-26T15:38:02.803Z",
          "content": "<p>Layer infromation is useful to remove false positives. Eg a track that moves from layer 2 to 6 without 4 is rare</p>",
          "rawMarkdown": "Layer infromation is useful to remove false positives. Eg a track that moves from layer 2 to 6 without 4 is rare"
        },
        {
          "id": 348559,
          "postDate": "2018-06-26T23:16:17.893Z",
          "content": "<p>the projection plane can be treated like \"convolutional weights\"  and to be learned. Since only one of the planes should work, some kind of \"view pooling\" can be added:</p>\n\n<p>(a,r, z/r) ---&gt; projected to a set of 2d planes (normal of planes learned via back propagation) ---&gt; pool (choose most compact) ---&gt; ...</p>",
          "rawMarkdown": "the projection plane can be treated like \"convolutional weights\"  and to be learned. Since only one of the planes should work, some kind of \"view pooling\" can be added:\n\n\n(a,r, z/r) ---&gt; projected to a set of 2d planes (normal of planes learned via back propagation) ---&gt; pool (choose most compact) ---&gt; ...",
          "votes": 2
        }
      ]
    },
    {
      "id": 348239,
      "postDate": "2018-06-26T10:52:34.427Z",
      "content": "<p>@Heng, the aim of the scoring functions written by @Vicens and @CPMP was to be fast. That is why there is a line:</p>\n\n<pre><code>sum(df[r1&gt;.5 &amp; r2&gt;.5,weight])\n</code></pre>\n\n<p>instead of slower:</p>\n\n<pre><code>sum(df[r1&gt;.5 &amp; r2&gt;.5 &amp; Nt&gt;3 &amp; Np&gt;3,weight])/sum(df[Np&gt;3,weight])\n</code></pre>\n\n<p>They work excellent with full event files, but the results obtained for a part of them are not exactly what we want. What do I mean? Imagine a particle which starts from (0,0,0) and in the region z&gt;500 &amp; r&lt;200 makes 1 hit only. If your clustering method creates a cluster for this hit only, its weight is added to the score. I do not think, you want to call a one-hit-cluster a tracklet. In the case above, discrimination of really small clusters (1-3 hits) gives a bit different results. I am not familiar with Python, so I do not know if there is such discrimination in your code.</p>",
      "rawMarkdown": "@Heng, the aim of the scoring functions written by @Vicens and @CPMP was to be fast. That is why there is a line:\n\n    sum(df[r1&gt;.5 &amp; r2&gt;.5,weight])\n\ninstead of slower:\n\n    sum(df[r1&gt;.5 &amp; r2&gt;.5 &amp; Nt&gt;3 &amp; Np&gt;3,weight])/sum(df[Np&gt;3,weight])\n\nThey work excellent with full event files, but the results obtained for a part of them are not exactly what we want. What do I mean? Imagine a particle which starts from (0,0,0) and in the region z&gt;500 &amp; r&lt;200 makes 1 hit only. If your clustering method creates a cluster for this hit only, its weight is added to the score. I do not think, you want to call a one-hit-cluster a tracklet. In the case above, discrimination of really small clusters (1-3 hits) gives a bit different results. I am not familiar with Python, so I do not know if there is such discrimination in your code."
    },
    {
      "id": 347774,
      "postDate": "2018-06-25T10:38:53.393Z",
      "content": "<p>@Heng, thanks for sharing.  I noticed that layer_id is not unique across volumes, for instance there are hits with layer_id  2 in both volume_id 8 and volume_id 9.  Did you pre process your data to get unique layer_ids?  If not then the test in your code should be refined I'm afraid.  Anyway, the idea of looking for missing layers is interesting.  </p>\n\n<p>Your idea to not use a track if a point on it is already assigned a longer one is puzzling (line 213). I'm not sure this is right though, but I'll try it and report if it helps or not here.  </p>\n\n<p>Last point, what about count ties?  The order in which you generate the candidates is probably quite important.  </p>",
      "rawMarkdown": "@Heng, thanks for sharing.  I noticed that layer_id is not unique across volumes, for instance there are hits with layer_id  2 in both volume_id 8 and volume_id 9.  Did you pre process your data to get unique layer_ids?  If not then the test in your code should be refined I'm afraid.  Anyway, the idea of looking for missing layers is interesting.  \n\nYour idea to not use a track if a point on it is already assigned a longer one is puzzling (line 213). I'm not sure this is right though, but I'll try it and report if it helps or not here.  \n\nLast point, what about count ties?  The order in which you generate the candidates is probably quite important.  ",
      "replies": [
        {
          "id": 347920,
          "postDate": "2018-06-25T17:19:27.940Z",
          "content": "<p>My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. </p>\n\n<p>I doubt the above assumption is a good assumption, but a 'good enough' assumption, in that it is better than favoring the later track. </p>\n\n<p>I've tried methods that would try to evaluate which track is more probable, for example I'd think an actual track will scale r as a function of z, but I think the logic I was trying to use was simply a more simple implementation of the logic behind dbscan... It reduced score and increased runtime. </p>\n\n<p>Possibly will need a more advanced solution to evaluate probable tracks, especially with shifting origins </p>",
          "rawMarkdown": "My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. \n\nI doubt the above assumption is a good assumption, but a 'good enough' assumption, in that it is better than favoring the later track. \n\nI've tried methods that would try to evaluate which track is more probable, for example I'd think an actual track will scale r as a function of z, but I think the logic I was trying to use was simply a more simple implementation of the logic behind dbscan... It reduced score and increased runtime. \n\nPossibly will need a more advanced solution to evaluate probable tracks, especially with shifting origins ",
          "votes": 1
        },
        {
          "id": 347927,
          "postDate": "2018-06-25T17:42:08.823Z",
          "content": "<p>Interesting, I was thinking of doing something similar, checking linearity of proposed tracks as a measure of confidence (z-z0)/r ~ const...</p>\n\n<p>I'll let you know if I end up with the same conclusion (hurts both score and execution time...)</p>",
          "rawMarkdown": "Interesting, I was thinking of doing something similar, checking linearity of proposed tracks as a measure of confidence (z-z0)/r ~ const...\n\nI'll let you know if I end up with the same conclusion (hurts both score and execution time...)"
        },
        {
          "id": 347930,
          "postDate": "2018-06-25T17:59:04.500Z",
          "content": "<p>I also tried linear fitting tracks for track extension... It performed way more poorly than hkc's extension method. </p>\n\n<p>I think linear models are very comfortable, so a tempting direction to go, but not robust enough for many of the more advanced tracks. Possibly modeling the transformed tracks, including the modified origin and modified 'unrolled' angle used in the creation of the cluster?</p>",
          "rawMarkdown": "I also tried linear fitting tracks for track extension... It performed way more poorly than hkc's extension method. \n\nI think linear models are very comfortable, so a tempting direction to go, but not robust enough for many of the more advanced tracks. Possibly modeling the transformed tracks, including the modified origin and modified 'unrolled' angle used in the creation of the cluster?"
        },
        {
          "id": 347939,
          "postDate": "2018-06-25T18:11:51.353Z",
          "content": "<blockquote>\n  <p>My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (</p>\n</blockquote>\n\n<p>I don't ensemble models outside my helix unrolling loop, but I do merge clusters obtained with various unrolling angles via that exact method.  I don't claim it is mine either, I found it independently but it has been shared in kernels by Grzegorz and others.</p>\n\n<p>The main difference with Heng's proposed method above is:\n 1. No secret sauce filtering at this point\n 2. Assign points to the largest cluster, point by point.  </p>\n\n<p>My secret sauce is more on the helix unrolling, but that sauce is far from being as tasty as outrunner's  sauce!  </p>\n\n<p>While I stull think my approach can yield few additional percents, I think the next step for me is to deal with the fact that particles are not following helix trajectories exactly.  And high weights hits are often associated with significant deviations from helix.  Again, I'm not the first one to make that observation, Heng made it loud and clear when presenting his track extension code.  Reading Heng carefully is a must in this competition ;)</p>",
          "rawMarkdown": "&gt; My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (\n\nI don't ensemble models outside my helix unrolling loop, but I do merge clusters obtained with various unrolling angles via that exact method.  I don't claim it is mine either, I found it independently but it has been shared in kernels by Grzegorz and others.\n\nThe main difference with Heng's proposed method above is:\n 1. No secret sauce filtering at this point\n 2. Assign points to the largest cluster, point by point.  \n\nMy secret sauce is more on the helix unrolling, but that sauce is far from being as tasty as outrunner's  sauce!  \n\nWhile I stull think my approach can yield few additional percents, I think the next step for me is to deal with the fact that particles are not following helix trajectories exactly.  And high weights hits are often associated with significant deviations from helix.  Again, I'm not the first one to make that observation, Heng made it loud and clear when presenting his track extension code.  Reading Heng carefully is a must in this competition ;)",
          "votes": 3
        },
        {
          "id": 348190,
          "postDate": "2018-06-26T08:18:30.663Z",
          "content": "<p>For me, until now, the only \"secret sauce\" is in helix unrolling - if it's accurate enough, clustering (and extending) is easy and you don't need any fancy algorithms, but when it is not accurate, nothing helps (I'm struggling now with tracks which start very far from the origin).  </p>",
          "rawMarkdown": "For me, until now, the only \"secret sauce\" is in helix unrolling - if it's accurate enough, clustering (and extending) is easy and you don't need any fancy algorithms, but when it is not accurate, nothing helps (I'm struggling now with tracks which start very far from the origin).  ",
          "votes": 1
        },
        {
          "id": 348196,
          "postDate": "2018-06-26T08:35:13.367Z",
          "content": "<p>@yuval r</p>\n\n<p>\"... the only \"secret sauce\" is in helix unrolling\". This is quite true. For the code i attached, it can get about 90% of the score for z&gt;500, r&lt;200. This is the volume region of straight tracks. For the  rest of the rest of the volume, if we just need to find a way to make the track straight.</p>",
          "rawMarkdown": "@yuval r\n \n \"... the only \"secret sauce\" is in helix unrolling\". This is quite true. For the code i attached, it can get about 90% of the score for z&gt;500, r&lt;200. This is the volume region of straight tracks. For the  rest of the rest of the volume, if we just need to find a way to make the track straight.",
          "votes": 1
        },
        {
          "id": 348200,
          "postDate": "2018-06-26T08:47:18.547Z",
          "content": "<p>@Heng, in your code I see you scan over different  angle offsets linearly, I found it much better to do it in a stochastic way, using normal distribution. (same for z0 scanning) </p>",
          "rawMarkdown": "@Heng, in your code I see you scan over different  angle offsets linearly, I found it much better to do it in a stochastic way, using normal distribution. (same for z0 scanning) ",
          "votes": 1
        },
        {
          "id": 348280,
          "postDate": "2018-06-26T12:39:58.693Z",
          "content": "<p>@yuval, interesting, I was experimenting with non linear scan</p>",
          "rawMarkdown": "@yuval, interesting, I was experimenting with non linear scan",
          "votes": 1
        },
        {
          "id": 348754,
          "postDate": "2018-06-27T08:12:18.383Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 348755,
          "postDate": "2018-06-27T08:14:11.460Z",
          "content": "<p>It is what @macfaril wrote:</p>\n\n<blockquote>\n  <p>starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. </p>\n</blockquote>",
          "rawMarkdown": "It is what @macfaril wrote:\n\n&gt; starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. ",
          "votes": 1
        },
        {
          "id": 349276,
          "postDate": "2018-06-28T01:22:50.267Z",
          "content": "<p>I do merge more complex, someone can do something like this:</p>\n\n<ol>\n<li><p>try to find the maximum score in your clusters according to GT.</p></li>\n<li><p>if your pipeline is merge -&gt; extend, try to add more stage like: merge high confidence -&gt; extend -&gt; merge others -&gt; extend ...</p></li>\n</ol>",
          "rawMarkdown": "I do merge more complex, someone can do something like this:\n\n1. try to find the maximum score in your clusters according to GT.\n\n2. if your pipeline is merge -&gt; extend, try to add more stage like: merge high confidence -&gt; extend -&gt; merge others -&gt; extend ...\n",
          "votes": 8
        },
        {
          "id": 349334,
          "postDate": "2018-06-28T02:36:46.350Z",
          "content": "<p>&gt; I do merge more complex, someone can do something like this:</p>\n\n<p>I'm sure you do more complex than me, the difference in score is there to tell us ;)  Thanks for sharing a bit of your secret sauce, you made me think about many new things.  I'll share some of them if they prove to be effective.</p>",
          "rawMarkdown": "&gt; I do merge more complex, someone can do something like this:\n\nI'm sure you do more complex than me, the difference in score is there to tell us ;)  Thanks for sharing a bit of your secret sauce, you made me think about many new things.  I'll share some of them if they prove to be effective.",
          "votes": 1
        },
        {
          "id": 349616,
          "postDate": "2018-06-28T11:29:37.273Z",
          "content": "<p><a href=\"/outrunner\">@outrunner</a> this might be a daft question, but what do you mean by the 'GT' in 'try to find the maximum score in your clusters according to GT'</p>",
          "rawMarkdown": "@outrunner this might be a daft question, but what do you mean by the 'GT' in 'try to find the maximum score in your clusters according to GT'",
          "votes": 1
        },
        {
          "id": 349640,
          "postDate": "2018-06-28T12:27:53.377Z",
          "content": "<p>Ground Truth</p>",
          "rawMarkdown": "Ground Truth"
        },
        {
          "id": 349647,
          "postDate": "2018-06-28T12:42:54.457Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 347950,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-25T18:51:38.580000",
      "content": "<p>@cpmp</p>\n\n<p>You can imagine that i am doing a ransac for line fitting. The cluster count is the number of inliners. Hence, assign hits to the line with the largest number of inliners </p>",
      "votes": 1,
      "replies": [
        {
          "id": 347995,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-25T21:07:34.490000",
          "content": "<p>For those (like me) who aren't in the know, ransac is a method of determining a best line of fit by excluding some outlier points. </p>\n\n<p>For a track that might be very well fit for 9 hits, and a final hit that is questionably on the track, that track will effectively be considered of length 9, so a new track with 10 hits that for the track well will cause the hits to be reassigned to the new track. </p>\n\n<p><a href=\"https://en.m.wikipedia.org/wiki/Random_sample_consensus\">https://en.m.wikipedia.org/wiki/Random_sample_consensus</a></p>\n\n<p>That being said, from reading the code and reading how ransac works (from a theoretical perspective), I think there is something I'm missing... Looks like I have some homework to do. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 348050,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-26T01:31:08.677000",
          "content": "<blockquote>\n  <p>Hence, assign hits to the line with the largest number of inliners </p>\n</blockquote>\n\n<p>Your code does something different: for a candidate, if a point is assigned to an existing cluster larger than it then you throw away the whole candidate.  I don't.  </p>\n\n<p>I've tried your filtering and it degrades a bit my local score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 347775,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-25T10:41:45.690000",
      "content": "<p>@CPMP</p>\n\n<p>this is for short tracks only (e.g. across 4 layer_ids).</p>\n\n<p>E.g you have divide all hits into different volume_id, and then choose a few layer_ids, to ensure tracklets are straight.</p>\n\n<p>Different volume_id may use different parameters</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 355240,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-07-11T09:11:03.450000",
      "content": "<p>@Heng, Looking at the values of dj, I guess your explore.py was written before you realized, that the beam spot σz  is not equal 55 mm but much less. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 355246,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-11T09:33:10.223000",
          "content": "<p>Grzegorz, your remark about it in another thread made me revisit my code too.  It yields a significant improvement at minimal cost.  Thanks a lot for that.  A sub is underway from me (won't be close to top, but encouraging when resuming after 15 days of absence from this competition).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355275,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-11T11:01:57.377000",
          "content": "<p>@Grzegorz, @CPMP thank you so much!!! I got it wrong too. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355277,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-07-11T11:15:49.760000",
          "content": "<p>This is probably another daft question, but where is beam spot σz defined? I read the correction - but I couldn't find where you saw the original 55 mm?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355280,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-07-11T11:22:47.407000",
          "content": "<p>There is a link in \"Welcome from the organizers! read this first!\" to:</p>\n\n<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a>\npage 6</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 355283,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-07-11T11:30:51.880000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355427,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-07-11T16:59:59.023000",
          "content": "<blockquote>\n  <p>Grzegorz, your remark about it in another thread made me revisit my code too</p>\n</blockquote>\n\n<p>@CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 355447,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-07-11T17:46:52.917000",
          "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>Thank you for the comment. you are right. The dj values are determined by trial and errors. </p>\n\n<p>Currently, I am putting DBSCA aside and focus on deep learning to learn pairwise links.\n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60447\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60447</a></p>\n\n<p>I can understand now why <a href=\"/outrunner\">@outrunner</a> tracks  are of good quality. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355574,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-12T02:31:42.673000",
          "content": "<blockquote>\n  <p>@CPMP, You mean my \"Even Z is in most cases much below 0+/-55 mm declared by the organizers\" in one of your early threads? I forgot it and had to discover once more. A lack of magnesium? I have to change the ratio of coffee/beer.</p>\n</blockquote>\n\n<p>Well, I also forgot about it.  It was good to move away from this for 2 weeks as I now have a fresh look at it.  And your posts contain many thought provocative insights, thanks again for that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355748,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-07-12T09:34:40.953000",
          "content": "<p>@Grzegorz, @CPMP,</p>\n\n<blockquote>\n  <p>that the beam spot σz is not equal 55 mm but much less.</p>\n</blockquote>\n\n<p>I read it first in the slides linked by @CPMP. It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.</p>\n\n<p>I don't understand why the organizers downplay this difference so much (except for personal reasons, of course), since it significantly decreases the geometric \"mixing\" of trajectories in the inner layers, doesn't it? The difference between 55mm and 5.5mm is comparable to or even larger than dimensions of the innermost detector structures. This is highly significant geometrically, I'd say. Even worse, it will affect different solution approaches differently. E.g. I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 355791,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-12T11:19:00.350000",
          "content": "<blockquote>\n  <p>It explains a puzzling observation I made: My algorithm prefers a much less elongated primary origin area then would be expected from the organizers' documents.</p>\n</blockquote>\n\n<p>I had the same observation, and did spent time trying to find out why this was not working as it should but I didn't check the validity of the description provided to us.  That's why i thank Grzegorz here.</p>\n\n<blockquote>\n  <p>I don't understand why the organizers downplay this difference so much</p>\n</blockquote>\n\n<p>They are not the ones sweating while reverse engineering the simulator.</p>\n\n<blockquote>\n  <p>I'd expect clustering approaches to benefit strongly from the reduced spot size, while other approaches might not benefit that much.</p>\n</blockquote>\n\n<p>Indeed.  At the width of the beam in x,y plane is also 10x smaller than described, which explains why approaches focusing on tracks originating from x,y = 0,0 work surprisingly well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 351495,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-07-02T10:06:59.763000",
      "content": "<p>@Heng I analyzed your code last night and the highest score I could get is still lower than doing dbscan on the full event hits, this is the score on the event 000001003 in the volume 7 and 9 using your code (with some of my hyper parameters to maximize your score)</p>\n\n<pre><code>max_score = df.weight.sum() = 0.28114\nscore1= 0.24930  (0.88674)\nscore2= 0.24930  (0.88674)\n</code></pre>\n\n<p>Using the full labels found by the current best dbscan model  (with the score 0.675)</p>\n\n<pre><code>max_score = df.weight.sum() = 0.28114\nscore1= 0.25874  (0.92033)\nscore2= 0.25874  (0.92033)\n</code></pre>\n\n<p>It indicates that this seeding approach may only find a subset of the straight tracks that the full-scan dbscan model can find. My takeaway is that if your dbscan model's score is over <code>0.6</code>, your model may outperform this approach in terms of finding seeds from the volume 7 and 9, the challenge is to find seeds of the tracks in the volume 8 and outside the pixel detectors, regardless of being straight or not. Of course, seeding is not necessary for this competition, we can do without, just seeding is a standard approach used by CERN. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 351652,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-07-02T18:23:59.407000",
          "content": "<p>When working with shifting on the z axis, I've found that limiting the max length of tracks gives me a worse score than allowing full tracks, which I figured would be counter intuitive because of tracks that originate further from the origin tend to be shorter. </p>\n\n<p>I have a guess as to what the limitation is, but I'm not sure how to test it. With the extension code adding hits one at a time, a short track being extended iteratively will only work until the first hit that differs too greatly from the previous pair. </p>\n\n<p>So if hits 1, 2,... 8 are all in a track, and you only seed hits 2, 3, and 4, but 5 is noisy, you might not be able to extend to 6, while all 8 might be detected with dbscan. </p>\n\n<p>It is worth noting that my local scores still aren't at .6x yet. </p>\n\n<p>Now that I'm writing all this I think I need to compare two submissions and find tracks where the dbscan can outperform an extended short track, and look at only that track... </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 355069,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-10T21:32:53.647000",
          "content": "<p>I am getting an error, when I try to use @Heng's explore.py script. Analyzing it I am not sure what the error is about. I think it is in the scoring function. See below:-</p>\n\n<p>&gt; TypeError: merge() got an unexpected keyword argument 'validate'</p>\n\n<p>@Nicole, @macfarII can you help?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355074,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-07-10T21:48:10.460000",
          "content": "<p>We've seen this before, my guess is that you should  update to pandas latest version.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 355078,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-07-10T21:53:10.527000",
          "content": "<p>I can confirm that fixed it for me...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 355081,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-10T22:05:36.563000",
          "content": "<p>Thanks @David and @John, that worked.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 348308,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-26T13:55:28.853000",
      "content": "<p>here are some ideas for your reference:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9703/e5.png\" alt=\"enter image description here\">\n   typo error:  ... supervised (or unsupervised)  learning ... coordinates should be a,r,z/r and not x,y,z</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9701/e2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9702/e3.png\" alt=\"enter image description here\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 348365,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-26T15:25:27.357000",
          "content": "<p>You're saying that for linear tracks, the track in transformed coordinates is directional, and all hits are in sequence with respect to layer id. </p>\n\n<p>Shouldn't a helix in transformed space should also be directional and form a curve,  each hit is in sequence with respect to layer id?</p>\n\n<p>From looking at your secret sauce, you're using z as thevway of determining if the layers are in the correct order... </p>\n\n<p>Would it make more sense to use the straightened track, and then evaluate that z is in the correct order, in that x, y are a function of z/layer id?</p>\n\n<p>I'm not sure why a track candidate would have hits that have an order of layer id that is different than the order of layer id...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348366,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-26T15:29:06.087000",
          "content": "<p>I am using a, r, z/r to represent the hits. In this coordinate system, tracks are almost piecewise linear</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348369,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-26T15:35:50.217000",
          "content": "<p>Would a track have r and z in the same order, which would remove the need for layer information?</p>\n\n<p>I guess my question was what does the layer id information contribute that the coordinates don't already have?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348371,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-26T15:38:02.803000",
          "content": "<p>Layer infromation is useful to remove false positives. Eg a track that moves from layer 2 to 6 without 4 is rare</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348559,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-26T23:16:17.893000",
          "content": "<p>the projection plane can be treated like \"convolutional weights\"  and to be learned. Since only one of the planes should work, some kind of \"view pooling\" can be added:</p>\n\n<p>(a,r, z/r) ---&gt; projected to a set of 2d planes (normal of planes learned via back propagation) ---&gt; pool (choose most compact) ---&gt; ...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 348239,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-06-26T10:52:34.427000",
      "content": "<p>@Heng, the aim of the scoring functions written by @Vicens and @CPMP was to be fast. That is why there is a line:</p>\n\n<pre><code>sum(df[r1&gt;.5 &amp; r2&gt;.5,weight])\n</code></pre>\n\n<p>instead of slower:</p>\n\n<pre><code>sum(df[r1&gt;.5 &amp; r2&gt;.5 &amp; Nt&gt;3 &amp; Np&gt;3,weight])/sum(df[Np&gt;3,weight])\n</code></pre>\n\n<p>They work excellent with full event files, but the results obtained for a part of them are not exactly what we want. What do I mean? Imagine a particle which starts from (0,0,0) and in the region z&gt;500 &amp; r&lt;200 makes 1 hit only. If your clustering method creates a cluster for this hit only, its weight is added to the score. I do not think, you want to call a one-hit-cluster a tracklet. In the case above, discrimination of really small clusters (1-3 hits) gives a bit different results. I am not familiar with Python, so I do not know if there is such discrimination in your code.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 347774,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-06-25T10:38:53.393000",
      "content": "<p>@Heng, thanks for sharing.  I noticed that layer_id is not unique across volumes, for instance there are hits with layer_id  2 in both volume_id 8 and volume_id 9.  Did you pre process your data to get unique layer_ids?  If not then the test in your code should be refined I'm afraid.  Anyway, the idea of looking for missing layers is interesting.  </p>\n\n<p>Your idea to not use a track if a point on it is already assigned a longer one is puzzling (line 213). I'm not sure this is right though, but I'll try it and report if it helps or not here.  </p>\n\n<p>Last point, what about count ties?  The order in which you generate the candidates is probably quite important.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 347920,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-25T17:19:27.940000",
          "content": "<p>My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. </p>\n\n<p>I doubt the above assumption is a good assumption, but a 'good enough' assumption, in that it is better than favoring the later track. </p>\n\n<p>I've tried methods that would try to evaluate which track is more probable, for example I'd think an actual track will scale r as a function of z, but I think the logic I was trying to use was simply a more simple implementation of the logic behind dbscan... It reduced score and increased runtime. </p>\n\n<p>Possibly will need a more advanced solution to evaluate probable tracks, especially with shifting origins </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 347927,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-25T17:42:08.823000",
          "content": "<p>Interesting, I was thinking of doing something similar, checking linearity of proposed tracks as a measure of confidence (z-z0)/r ~ const...</p>\n\n<p>I'll let you know if I end up with the same conclusion (hurts both score and execution time...)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 347930,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-25T17:59:04.500000",
          "content": "<p>I also tried linear fitting tracks for track extension... It performed way more poorly than hkc's extension method. </p>\n\n<p>I think linear models are very comfortable, so a tempting direction to go, but not robust enough for many of the more advanced tracks. Possibly modeling the transformed tracks, including the modified origin and modified 'unrolled' angle used in the creation of the cluster?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 347939,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-25T18:11:51.353000",
          "content": "<blockquote>\n  <p>My understanding is that starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (</p>\n</blockquote>\n\n<p>I don't ensemble models outside my helix unrolling loop, but I do merge clusters obtained with various unrolling angles via that exact method.  I don't claim it is mine either, I found it independently but it has been shared in kernels by Grzegorz and others.</p>\n\n<p>The main difference with Heng's proposed method above is:\n 1. No secret sauce filtering at this point\n 2. Assign points to the largest cluster, point by point.  </p>\n\n<p>My secret sauce is more on the helix unrolling, but that sauce is far from being as tasty as outrunner's  sauce!  </p>\n\n<p>While I stull think my approach can yield few additional percents, I think the next step for me is to deal with the fact that particles are not following helix trajectories exactly.  And high weights hits are often associated with significant deviations from helix.  Again, I'm not the first one to make that observation, Heng made it loud and clear when presenting his track extension code.  Reading Heng carefully is a must in this competition ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 348190,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-06-26T08:18:30.663000",
          "content": "<p>For me, until now, the only \"secret sauce\" is in helix unrolling - if it's accurate enough, clustering (and extending) is easy and you don't need any fancy algorithms, but when it is not accurate, nothing helps (I'm struggling now with tracks which start very far from the origin).  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348196,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-26T08:35:13.367000",
          "content": "<p>@yuval r</p>\n\n<p>\"... the only \"secret sauce\" is in helix unrolling\". This is quite true. For the code i attached, it can get about 90% of the score for z&gt;500, r&lt;200. This is the volume region of straight tracks. For the  rest of the rest of the volume, if we just need to find a way to make the track straight.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348200,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-06-26T08:47:18.547000",
          "content": "<p>@Heng, in your code I see you scan over different  angle offsets linearly, I found it much better to do it in a stochastic way, using normal distribution. (same for z0 scanning) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348280,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-26T12:39:58.693000",
          "content": "<p>@yuval, interesting, I was experimenting with non linear scan</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 348754,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-27T08:12:18.383000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 348755,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-27T08:14:11.460000",
          "content": "<p>It is what @macfaril wrote:</p>\n\n<blockquote>\n  <p>starting with smaller scaling of the angle, dbscan will find straighter tracks first, and then you'll move towards more helical tracks in the later iterations. Since you're starting with the more likely (straighter), you want to keep the prior track in the case of a tie. </p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349276,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-06-28T01:22:50.267000",
          "content": "<p>I do merge more complex, someone can do something like this:</p>\n\n<ol>\n<li><p>try to find the maximum score in your clusters according to GT.</p></li>\n<li><p>if your pipeline is merge -&gt; extend, try to add more stage like: merge high confidence -&gt; extend -&gt; merge others -&gt; extend ...</p></li>\n</ol>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 349334,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-28T02:36:46.350000",
          "content": "<p>&gt; I do merge more complex, someone can do something like this:</p>\n\n<p>I'm sure you do more complex than me, the difference in score is there to tell us ;)  Thanks for sharing a bit of your secret sauce, you made me think about many new things.  I'll share some of them if they prove to be effective.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349616,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-06-28T11:29:37.273000",
          "content": "<p><a href=\"/outrunner\">@outrunner</a> this might be a daft question, but what do you mean by the 'GT' in 'try to find the maximum score in your clusters according to GT'</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349640,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-06-28T12:27:53.377000",
          "content": "<p>Ground Truth</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349647,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-06-28T12:42:54.457000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "347757": "See the code for details. There is no LB submission for this code yet.\n\n1. run dbscan for different z offset, angle offset: z(r)= (z+delta)/r;   a(r) = a + delta*r\n\n2. store all clusters from different dbscan results. Sort them increasing length (cluster count)\n\n3. assign hits according to the filtering criterion in the code.\n\n\n---\n\nexample results:\n\n\n\n  ![enter image description here][1]\n\n\n- this does not use track extension yet. \n- this does not fix angle discontinuity yet\n- this code can be used for finding seeds of straight tracks\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/347757/9687/ensmble.png",
    "347950": "@cpmp\n\nYou can imagine that i am doing a ransac for line fitting. The cluster count is the number of inliners. Hence, assign hits to the line with the largest number of inliners ",
    "347775": "@CPMP\n\n\nthis is for short tracks only (e.g. across 4 layer_ids).\n\nE.g you have divide all hits into different volume_id, and then choose a few layer_ids, to ensure tracklets are straight.\n\nDifferent volume_id may use different parameters",
    "355240": "@Heng, Looking at the values of dj, I guess your explore.py was written before you realized, that the beam spot σz  is not equal 55 mm but much less. https://www.kaggle.com/c/trackml-particle-identification/discussion/60735#354315",
    "351495": "@Heng I analyzed your code last night and the highest score I could get is still lower than doing dbscan on the full event hits, this is the score on the event 000001003 in the volume 7 and 9 using your code (with some of my hyper parameters to maximize your score)\n\n    max_score = df.weight.sum() = 0.28114\n    score1= 0.24930  (0.88674)\n    score2= 0.24930  (0.88674)\n\n\nUsing the full labels found by the current best dbscan model  (with the score 0.675)\n\n    max_score = df.weight.sum() = 0.28114\n    score1= 0.25874  (0.92033)\n    score2= 0.25874  (0.92033)\n\n\nIt indicates that this seeding approach may only find a subset of the straight tracks that the full-scan dbscan model can find. My takeaway is that if your dbscan model's score is over `0.6`, your model may outperform this approach in terms of finding seeds from the volume 7 and 9, the challenge is to find seeds of the tracks in the volume 8 and outside the pixel detectors, regardless of being straight or not. Of course, seeding is not necessary for this competition, we can do without, just seeding is a standard approach used by CERN. ",
    "348308": "here are some ideas for your reference:\n\n\n  ![enter image description here][1]\n   typo error:  ... supervised (or unsupervised)  learning ... coordinates should be a,r,z/r and not x,y,z\n\n  ![enter image description here][2]\n\n  ![enter image description here][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9703/e5.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9701/e2.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/348308/9702/e3.png",
    "348239": "@Heng, the aim of the scoring functions written by @Vicens and @CPMP was to be fast. That is why there is a line:\n\n    sum(df[r1&gt;.5 &amp; r2&gt;.5,weight])\n\ninstead of slower:\n\n    sum(df[r1&gt;.5 &amp; r2&gt;.5 &amp; Nt&gt;3 &amp; Np&gt;3,weight])/sum(df[Np&gt;3,weight])\n\nThey work excellent with full event files, but the results obtained for a part of them are not exactly what we want. What do I mean? Imagine a particle which starts from (0,0,0) and in the region z&gt;500 &amp; r&lt;200 makes 1 hit only. If your clustering method creates a cluster for this hit only, its weight is added to the score. I do not think, you want to call a one-hit-cluster a tracklet. In the case above, discrimination of really small clusters (1-3 hits) gives a bit different results. I am not familiar with Python, so I do not know if there is such discrimination in your code.",
    "347774": "@Heng, thanks for sharing.  I noticed that layer_id is not unique across volumes, for instance there are hits with layer_id  2 in both volume_id 8 and volume_id 9.  Did you pre process your data to get unique layer_ids?  If not then the test in your code should be refined I'm afraid.  Anyway, the idea of looking for missing layers is interesting.  \n\nYour idea to not use a track if a point on it is already assigned a longer one is puzzling (line 213). I'm not sure this is right though, but I'll try it and report if it helps or not here.  \n\nLast point, what about count ties?  The order in which you generate the candidates is probably quite important.  "
  }
}