{
  "id": 60638,
  "title": "how to score a track is good or not",
  "url": "/competitions/trackml-particle-identification/discussion/60638",
  "author_name": "hengck23",
  "post_date": "2018-07-07T18:19:35.209000",
  "votes": 5,
  "comment_count": 22,
  "views": 0,
  "content": "<p>let a track candidate be given as a sequence of T= { (i0,p0), (i1,p1) .... (iN,pN) }</p>\n\n<p>Here p can be (x,y,z) or other representation like (a,z/r,z) or an embedded feature that encode a hit and its neighborhood. Let assume p =(x,yz)</p>\n\n<p>And i is an unique int that identify the i=(volume_id,layer_id). Assume i range from 0 to L.</p>\n\n<p>Create a zero Nx3 array. Set  array[i]=p. Use this array as input feature to train a classifier or ranking function.</p>\n\n<p>Note:</p>\n\n<ol>\n<li><p>the temporal information is encoded in i.</p></li>\n<li><p>In case where a track has multiple hit in a single i=(volume_id,layer_id), represent this track in multiple sequence, each has only one hit in a single i.</p></li>\n</ol>\n\n<p>to be updated:  code and results .</p>\n\n<p>e.g. label is 1 or 0</p>\n\n<p>{ (i0,p0), (i1,p1) .... (iN,pN) }  --&gt; fixed length array --&gt;  [linear network] --&gt; {label0, label1 ....labelN}</p>",
  "messages": [
    {
      "id": 353772,
      "postDate": "2018-07-07T18:19:35.210Z",
      "content": "<p>let a track candidate be given as a sequence of T= { (i0,p0), (i1,p1) .... (iN,pN) }</p>\n\n<p>Here p can be (x,y,z) or other representation like (a,z/r,z) or an embedded feature that encode a hit and its neighborhood. Let assume p =(x,yz)</p>\n\n<p>And i is an unique int that identify the i=(volume_id,layer_id). Assume i range from 0 to L.</p>\n\n<p>Create a zero Nx3 array. Set  array[i]=p. Use this array as input feature to train a classifier or ranking function.</p>\n\n<p>Note:</p>\n\n<ol>\n<li><p>the temporal information is encoded in i.</p></li>\n<li><p>In case where a track has multiple hit in a single i=(volume_id,layer_id), represent this track in multiple sequence, each has only one hit in a single i.</p></li>\n</ol>\n\n<p>to be updated:  code and results .</p>\n\n<p>e.g. label is 1 or 0</p>\n\n<p>{ (i0,p0), (i1,p1) .... (iN,pN) }  --&gt; fixed length array --&gt;  [linear network] --&gt; {label0, label1 ....labelN}</p>",
      "rawMarkdown": "let a track candidate be given as a sequence of T= { (i0,p0), (i1,p1) .... (iN,pN) }\n\nHere p can be (x,y,z) or other representation like (a,z/r,z) or an embedded feature that encode a hit and its neighborhood. Let assume p =(x,yz)\n\nAnd i is an unique int that identify the i=(volume_id,layer_id). Assume i range from 0 to L.\n\nCreate a zero Nx3 array. Set  array[i]=p. Use this array as input feature to train a classifier or ranking function.\n\nNote:\n\n1.  the temporal information is encoded in i.\n\n2. In case where a track has multiple hit in a single i=(volume_id,layer_id), represent this track in multiple sequence, each has only one hit in a single i.\n\nto be updated:  code and results .\n\ne.g. label is 1 or 0\n\n{ (i0,p0), (i1,p1) .... (iN,pN) }  --&gt; fixed length array --&gt;  [linear network] --&gt; {label0, label1 ....labelN}\n\n\n\n\n\n",
      "votes": 5
    },
    {
      "id": 353913,
      "postDate": "2018-07-08T09:26:11.413Z",
      "content": "<p>＠Heng, interesting, I just tried a very similar approach yesterday but my simple linear NN trained with 100 events doesn't work too well on this, the accuracy rate was only about <code>58%</code> , a random guess would outperform my NN :) that's why I don't bother to share my not-working code here.  Incidentally, <code>module_id</code> can be useful for double hits but most of time very noisy, have you thought about how to add <code>module_id</code> as train data too?</p>\n\n<p>( My human random guess would outperform this prediction I'm sure.....) </p>\n\n<pre><code>Groundtruth\n[ [ 1.] [ 0.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 0.] [ 0.] [ 1.] [ 1.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 1.] ]\n\nPrediction \n[[ 0.48636526]\n [ 0.55879432]\n [ 0.50961298]\n [ 0.48604107]\n [ 0.61274189]\n [ 0.49084449]\n [ 0.92523319]\n [ 0.49379969]\n [ 0.56074548]\n [ 0.50928968]\n [ 0.50403577]\n [ 0.48605412]\n [ 0.88126868]\n [ 0.65616256]\n [ 0.49116713]\n [ 0.92684376]\n [ 0.61099911]\n [ 0.88783711]\n [ 0.50963062]\n [ 0.60075444]]\n</code></pre>",
      "rawMarkdown": "＠Heng, interesting, I just tried a very similar approach yesterday but my simple linear NN trained with 100 events doesn't work too well on this, the accuracy rate was only about `58%` , a random guess would outperform my NN :) that's why I don't bother to share my not-working code here.  Incidentally, `module_id` can be useful for double hits but most of time very noisy, have you thought about how to add `module_id` as train data too?\n\n( My human random guess would outperform this prediction I'm sure.....) \n\n    Groundtruth\n    [ [ 1.] [ 0.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 0.] [ 0.] [ 1.] [ 1.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 1.] ]\n    \n    Prediction \n    [[ 0.48636526]\n     [ 0.55879432]\n     [ 0.50961298]\n     [ 0.48604107]\n     [ 0.61274189]\n     [ 0.49084449]\n     [ 0.92523319]\n     [ 0.49379969]\n     [ 0.56074548]\n     [ 0.50928968]\n     [ 0.50403577]\n     [ 0.48605412]\n     [ 0.88126868]\n     [ 0.65616256]\n     [ 0.49116713]\n     [ 0.92684376]\n     [ 0.61099911]\n     [ 0.88783711]\n     [ 0.50963062]\n     [ 0.60075444]]\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 353942,
          "postDate": "2018-07-08T10:30:17.827Z",
          "content": "<p>the input representation is important.  i will post my code soon.</p>",
          "rawMarkdown": "the input representation is important.  i will post my code soon.\n\n",
          "votes": 2
        },
        {
          "id": 354051,
          "postDate": "2018-07-08T15:33:06.653Z",
          "content": "<p>@Heng  my input was</p>\n\n<pre><code>input  = np.column_stack((x/1000, y/1000, z/3000))\nor \ninput  = np.column_stack((a, r/1000, x/1000, y/1000, z/3000))\n</code></pre>\n\n<p>each sample has valid hits (labelled 1) mixed with noise and the hits from other tracks (labelled 0) without losing the majority. The valid hits were sorted based on z (kind of like your encoded detector information) or shuffled, both accuracy were not good, and my losses didn't go down as much. I think the problem could be that the hits of other tracks were very similar and the NN had difficulty learning the pattern, since if I only mixed my tracks with noise, the NN yielded around 80% of accuracy. My loss function was <code>binary_crossentropy</code> since it's kind of similar to a multi-label classification problem. I was wondering if this idea would work with a simple NN at all or we have to use a PointNet and I'm glad you pointed out this idea. </p>\n\n<p>I look forward to your input and results. </p>",
          "rawMarkdown": "@Heng  my input was\n\n    input  = np.column_stack((x/1000, y/1000, z/3000))\n    or \n    input  = np.column_stack((a, r/1000, x/1000, y/1000, z/3000))\n\neach sample has valid hits (labelled 1) mixed with noise and the hits from other tracks (labelled 0) without losing the majority. The valid hits were sorted based on z (kind of like your encoded detector information) or shuffled, both accuracy were not good, and my losses didn't go down as much. I think the problem could be that the hits of other tracks were very similar and the NN had difficulty learning the pattern, since if I only mixed my tracks with noise, the NN yielded around 80% of accuracy. My loss function was `binary_crossentropy` since it's kind of similar to a multi-label classification problem. I was wondering if this idea would work with a simple NN at all or we have to use a PointNet and I'm glad you pointed out this idea. \n\nI look forward to your input and results. "
        },
        {
          "id": 354071,
          "postDate": "2018-07-08T16:51:51.577Z",
          "content": "<p>I think your problem is hit's order, and pointnet will help since it is order free.</p>",
          "rawMarkdown": "I think your problem is hit's order, and pointnet will help since it is order free.",
          "votes": 1
        },
        {
          "id": 354141,
          "postDate": "2018-07-08T21:15:31.013Z",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>, I think my feedforward NN was not good for this classification, regardless of hits being sorted or not, it yielded the same accuracy. I'll try other models. Thanks!</p>",
          "rawMarkdown": "@outrunner, I think my feedforward NN was not good for this classification, regardless of hits being sorted or not, it yielded the same accuracy. I'll try other models. Thanks!"
        },
        {
          "id": 354354,
          "postDate": "2018-07-09T11:12:27.247Z",
          "content": "<p>the order is important, because it help to reduce outliers.</p>\n\n<p>for me, the issue is alignment. how to align different tracks to \"remove variation\".  e,g, </p>\n\n<p>(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; (hit0, hit1, hit2 ...) - hit0</p>\n\n<p>or</p>\n\n<p>(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; project_in_direction ( (hit0, hit1, hit2 ...),  (hit1-hitLast) ) </p>\n\n<p>Note that original poinetnet has a spatial transform network.  </p>",
          "rawMarkdown": "the order is important, because it help to reduce outliers.\n\nfor me, the issue is alignment. how to align different tracks to \"remove variation\".  e,g, \n\n(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; (hit0, hit1, hit2 ...) - hit0\n\nor\n\n(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; project_in_direction ( (hit0, hit1, hit2 ...),  (hit1-hitLast) ) \n\n\nNote that original poinetnet has a spatial transform network.  \n\n"
        },
        {
          "id": 354557,
          "postDate": "2018-07-09T19:06:09.557Z",
          "content": "<p>@Heng, you're right, the original pointnet intends to make the model invariant to the input permutation but the order does help in practice.  \"The problem with the alignment\" -&gt; do you mean you have trouble making the input more invariant to geometric transformation? pointnet comes with a joint alignment network (T-net) or I guess you mean you'd have this problem in a NN in general.</p>",
          "rawMarkdown": "@Heng, you're right, the original pointnet intends to make the model invariant to the input permutation but the order does help in practice.  \"The problem with the alignment\" -&gt; do you mean you have trouble making the input more invariant to geometric transformation? pointnet comes with a joint alignment network (T-net) or I guess you mean you'd have this problem in a NN in general."
        }
      ]
    },
    {
      "id": 354020,
      "postDate": "2018-07-08T14:26:28.030Z",
      "content": "<p>A track with higher probability of occurrence should a higher score. This is important in ensembling results (or merging track), where we suppress low score tracks with high score one if they overlap (aka. non max suppression)</p>\n\n<p>Here are some rules , based on my observation. Given 2 tracks (in a,z/r,z coordinates) that are well fitted by some parametric equations:</p>\n\n<p>\"1\" is highest score</p>\n\n<ol>\n<li><p>longest length, i.e. high hit count</p></li>\n<li><p>start from lowest z, transverse sequentially in layers (not missing any layers)</p></li>\n<li><p>straight track  </p></li>\n<li><p>Low error from curve fitting</p></li>\n<li><p>does not cross other track</p></li>\n</ol>\n\n<p>Other observations:</p>\n\n<ol>\n<li><p>Do not treat all tracks the same.  Grow and extend high confident track first</p></li>\n<li><p>if you have difficult in using supervised learning, apply training and testing at for \"potential high confident tracks\" tracks only. E.g. take care of those that are straight, long tracks, assume they transverse sequentially in layers , etc ...</p></li>\n<li><p>To increase LB score, always try to link the hits that re regarded as background noise (labelled 0) in the \"last step\",.This is false positive FP don't give lower LB score because of the metric. We just need to ensure FB don't break up the inlier tracks. </p></li>\n<li><p>the curvature, length, etc of track is related  a, z/r. use this to improve outlier removal.  </p></li>\n</ol>\n\n<p>... to be updated ...</p>",
      "rawMarkdown": "A track with higher probability of occurrence should a higher score. This is important in ensembling results (or merging track), where we suppress low score tracks with high score one if they overlap (aka. non max suppression)\n\nHere are some rules , based on my observation. Given 2 tracks (in a,z/r,z coordinates) that are well fitted by some parametric equations:\n\n\"1\" is highest score\n\n1.  longest length, i.e. high hit count\n\n2. start from lowest z, transverse sequentially in layers (not missing any layers)\n\n3. straight track  \n\n4. Low error from curve fitting\n\n5. does not cross other track\n\n\nOther observations:\n\n1. Do not treat all tracks the same.  Grow and extend high confident track first\n\n2. if you have difficult in using supervised learning, apply training and testing at for \"potential high confident tracks\" tracks only. E.g. take care of those that are straight, long tracks, assume they transverse sequentially in layers , etc ...\n\n3. To increase LB score, always try to link the hits that re regarded as background noise (labelled 0) in the \"last step\",.This is false positive FP don't give lower LB score because of the metric. We just need to ensure FB don't break up the inlier tracks. \n\n4. the curvature, length, etc of track is related  a, z/r. use this to improve outlier removal.  \n \n \n... to be updated ...",
      "votes": 2,
      "replies": [
        {
          "id": 354032,
          "postDate": "2018-07-08T14:49:19.603Z",
          "content": "<p>Step 3 can lower LB score because of the condition that a track must contain more than 50% of hits from same particle.</p>",
          "rawMarkdown": "Step 3 can lower LB score because of the condition that a track must contain more than 50% of hits from same particle."
        },
        {
          "id": 354035,
          "postDate": "2018-07-08T14:52:13.360Z",
          "content": "<p>We can just cluster the background noise. We do not merge the clustered noise with the already determined tracks</p>",
          "rawMarkdown": "We can just cluster the background noise. We do not merge the clustered noise with the already determined tracks"
        },
        {
          "id": 354048,
          "postDate": "2018-07-08T15:25:47.323Z",
          "content": "<p>@Heng, thanks for your thought. Indeed, the idea of the confidence of tracks has brought up by <a href=\"/outrunner\">@outrunner</a> in one of his posts here. \n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60330\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60330</a></p>\n\n<p>I think dbscan clustering did find enough tracks but we merged them badly without ranking the tracks and giving them confidence/quality score. This is an interesting idea to work on. </p>",
          "rawMarkdown": "@Heng, thanks for your thought. Indeed, the idea of the confidence of tracks has brought up by @outrunner in one of his posts here. \nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/60330\n\nI think dbscan clustering did find enough tracks but we merged them badly without ranking the tracks and giving them confidence/quality score. This is an interesting idea to work on. \n",
          "votes": 1
        },
        {
          "id": 354053,
          "postDate": "2018-07-08T15:35:59.390Z",
          "content": "<p>Ad. point 3. We can attach noise hits to the tracks of even number of hits without compromising the score. For example, when the track has 8 hits, and 5 good ones (&gt; 50%), you can add one more hit, because 5/9 is still &gt; 50%.</p>",
          "rawMarkdown": "Ad. point 3. We can attach noise hits to the tracks of even number of hits without compromising the score. For example, when the track has 8 hits, and 5 good ones (&gt; 50%), you can add one more hit, because 5/9 is still &gt; 50%.",
          "votes": 5
        },
        {
          "id": 354146,
          "postDate": "2018-07-08T21:42:41.143Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 355052,
          "postDate": "2018-07-10T20:24:55.467Z",
          "content": "<p>Yes, and very honorable of Grzegorz to share it. I'm focused on improvements that are more or less physically motivated, so I would most likely have missed it. It's a bit disappointing that the scoring function has such a non-physical loophole, though.</p>",
          "rawMarkdown": "Yes, and very honorable of Grzegorz to share it. I'm focused on improvements that are more or less physically motivated, so I would most likely have missed it. It's a bit disappointing that the scoring function has such a non-physical loophole, though.",
          "votes": 2
        },
        {
          "id": 355072,
          "postDate": "2018-07-10T21:45:19.993Z",
          "content": "<p>I would not say this particular trick was anticipated, but we've deliberately put emphasis on finding tracks candidates with many hits, rather than having clean tracks. Because cleaning a track from outliers we consider relatively easy to do in a finalizing stage, while missed tracks cannot be easily recovered.</p>",
          "rawMarkdown": "I would not say this particular trick was anticipated, but we've deliberately put emphasis on finding tracks candidates with many hits, rather than having clean tracks. Because cleaning a track from outliers we consider relatively easy to do in a finalizing stage, while missed tracks cannot be easily recovered.",
          "votes": 3
        }
      ]
    },
    {
      "id": 353820,
      "postDate": "2018-07-07T22:54:41.307Z",
      "content": "<p>How would the track candidate be selected at evaluation time? Would it be some random set of hits of a random length between 1 and about 20 hits?</p>",
      "rawMarkdown": "How would the track candidate be selected at evaluation time? Would it be some random set of hits of a random length between 1 and about 20 hits?",
      "replies": [
        {
          "id": 354561,
          "postDate": "2018-07-09T19:14:28.993Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 354681,
          "postDate": "2018-07-10T03:56:28.653Z",
          "content": "<p>That makes sense. Using dbscan or a similar preprocessing step to find good candidates sounds like a much better approach, that would actually work, than random selection. </p>",
          "rawMarkdown": "That makes sense. Using dbscan or a similar preprocessing step to find good candidates sounds like a much better approach, that would actually work, than random selection. "
        },
        {
          "id": 354823,
          "postDate": "2018-07-10T10:16:15.580Z",
          "content": "<p>@Jack,   @Heng shared quite a few approaches of how to find initial tracks. I shared our results too, e.g. with dbscan, we can find over 92% of initial tracks in the volume 7 and 9 (seeding area) but only 68% in the volume 8 since it's very dense. \n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59635\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59635</a></p>\n\n<p>dbscan or clustering in general is actually a very powerful tool to find your candidates and then you can extend your tracks from there.</p>",
          "rawMarkdown": "@Jack,   @Heng shared quite a few approaches of how to find initial tracks. I shared our results too, e.g. with dbscan, we can find over 92% of initial tracks in the volume 7 and 9 (seeding area) but only 68% in the volume 8 since it's very dense. \nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/59635\n\ndbscan or clustering in general is actually a very powerful tool to find your candidates and then you can extend your tracks from there.",
          "votes": 1
        },
        {
          "id": 355131,
          "postDate": "2018-07-11T03:00:52.723Z",
          "content": "<p>Thanks Nicole, very helpful information. Is volume 13 much the same case as volume 8, lower accuracy using dbscan to find seed tracks since the hit count is much denser than volume 7 and 9? </p>",
          "rawMarkdown": "Thanks Nicole, very helpful information. Is volume 13 much the same case as volume 8, lower accuracy using dbscan to find seed tracks since the hit count is much denser than volume 7 and 9? "
        }
      ]
    },
    {
      "id": 353816,
      "postDate": "2018-07-07T22:15:45.267Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 353861,
          "postDate": "2018-07-08T05:27:54.803Z",
          "content": "<p>I am using something similar, to remove points that are too far from the best helix fit.</p>",
          "rawMarkdown": "I am using something similar, to remove points that are too far from the best helix fit.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 353913,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-07-08T09:26:11.413000",
      "content": "<p>＠Heng, interesting, I just tried a very similar approach yesterday but my simple linear NN trained with 100 events doesn't work too well on this, the accuracy rate was only about <code>58%</code> , a random guess would outperform my NN :) that's why I don't bother to share my not-working code here.  Incidentally, <code>module_id</code> can be useful for double hits but most of time very noisy, have you thought about how to add <code>module_id</code> as train data too?</p>\n\n<p>( My human random guess would outperform this prediction I'm sure.....) </p>\n\n<pre><code>Groundtruth\n[ [ 1.] [ 0.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 0.] [ 0.] [ 1.] [ 1.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 1.] ]\n\nPrediction \n[[ 0.48636526]\n [ 0.55879432]\n [ 0.50961298]\n [ 0.48604107]\n [ 0.61274189]\n [ 0.49084449]\n [ 0.92523319]\n [ 0.49379969]\n [ 0.56074548]\n [ 0.50928968]\n [ 0.50403577]\n [ 0.48605412]\n [ 0.88126868]\n [ 0.65616256]\n [ 0.49116713]\n [ 0.92684376]\n [ 0.61099911]\n [ 0.88783711]\n [ 0.50963062]\n [ 0.60075444]]\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 353942,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-07-08T10:30:17.827000",
          "content": "<p>the input representation is important.  i will post my code soon.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 354051,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-08T15:33:06.653000",
          "content": "<p>@Heng  my input was</p>\n\n<pre><code>input  = np.column_stack((x/1000, y/1000, z/3000))\nor \ninput  = np.column_stack((a, r/1000, x/1000, y/1000, z/3000))\n</code></pre>\n\n<p>each sample has valid hits (labelled 1) mixed with noise and the hits from other tracks (labelled 0) without losing the majority. The valid hits were sorted based on z (kind of like your encoded detector information) or shuffled, both accuracy were not good, and my losses didn't go down as much. I think the problem could be that the hits of other tracks were very similar and the NN had difficulty learning the pattern, since if I only mixed my tracks with noise, the NN yielded around 80% of accuracy. My loss function was <code>binary_crossentropy</code> since it's kind of similar to a multi-label classification problem. I was wondering if this idea would work with a simple NN at all or we have to use a PointNet and I'm glad you pointed out this idea. </p>\n\n<p>I look forward to your input and results. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354071,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-07-08T16:51:51.577000",
          "content": "<p>I think your problem is hit's order, and pointnet will help since it is order free.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 354141,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-08T21:15:31.013000",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>, I think my feedforward NN was not good for this classification, regardless of hits being sorted or not, it yielded the same accuracy. I'll try other models. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354354,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-07-09T11:12:27.247000",
          "content": "<p>the order is important, because it help to reduce outliers.</p>\n\n<p>for me, the issue is alignment. how to align different tracks to \"remove variation\".  e,g, </p>\n\n<p>(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; (hit0, hit1, hit2 ...) - hit0</p>\n\n<p>or</p>\n\n<p>(hit0, hit1, hit2 ...)  --&gt; [alignment]  --&gt; project_in_direction ( (hit0, hit1, hit2 ...),  (hit1-hitLast) ) </p>\n\n<p>Note that original poinetnet has a spatial transform network.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354557,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-09T19:06:09.557000",
          "content": "<p>@Heng, you're right, the original pointnet intends to make the model invariant to the input permutation but the order does help in practice.  \"The problem with the alignment\" -&gt; do you mean you have trouble making the input more invariant to geometric transformation? pointnet comes with a joint alignment network (T-net) or I guess you mean you'd have this problem in a NN in general.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 354020,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-07-08T14:26:28.030000",
      "content": "<p>A track with higher probability of occurrence should a higher score. This is important in ensembling results (or merging track), where we suppress low score tracks with high score one if they overlap (aka. non max suppression)</p>\n\n<p>Here are some rules , based on my observation. Given 2 tracks (in a,z/r,z coordinates) that are well fitted by some parametric equations:</p>\n\n<p>\"1\" is highest score</p>\n\n<ol>\n<li><p>longest length, i.e. high hit count</p></li>\n<li><p>start from lowest z, transverse sequentially in layers (not missing any layers)</p></li>\n<li><p>straight track  </p></li>\n<li><p>Low error from curve fitting</p></li>\n<li><p>does not cross other track</p></li>\n</ol>\n\n<p>Other observations:</p>\n\n<ol>\n<li><p>Do not treat all tracks the same.  Grow and extend high confident track first</p></li>\n<li><p>if you have difficult in using supervised learning, apply training and testing at for \"potential high confident tracks\" tracks only. E.g. take care of those that are straight, long tracks, assume they transverse sequentially in layers , etc ...</p></li>\n<li><p>To increase LB score, always try to link the hits that re regarded as background noise (labelled 0) in the \"last step\",.This is false positive FP don't give lower LB score because of the metric. We just need to ensure FB don't break up the inlier tracks. </p></li>\n<li><p>the curvature, length, etc of track is related  a, z/r. use this to improve outlier removal.  </p></li>\n</ol>\n\n<p>... to be updated ...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 354032,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-08T14:49:19.603000",
          "content": "<p>Step 3 can lower LB score because of the condition that a track must contain more than 50% of hits from same particle.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354035,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-07-08T14:52:13.360000",
          "content": "<p>We can just cluster the background noise. We do not merge the clustered noise with the already determined tracks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354048,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-08T15:25:47.323000",
          "content": "<p>@Heng, thanks for your thought. Indeed, the idea of the confidence of tracks has brought up by <a href=\"/outrunner\">@outrunner</a> in one of his posts here. \n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60330\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60330</a></p>\n\n<p>I think dbscan clustering did find enough tracks but we merged them badly without ranking the tracks and giving them confidence/quality score. This is an interesting idea to work on. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 354053,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-07-08T15:35:59.390000",
          "content": "<p>Ad. point 3. We can attach noise hits to the tracks of even number of hits without compromising the score. For example, when the track has 8 hits, and 5 good ones (&gt; 50%), you can add one more hit, because 5/9 is still &gt; 50%.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 354146,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-08T21:42:41.143000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 355052,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-07-10T20:24:55.467000",
          "content": "<p>Yes, and very honorable of Grzegorz to share it. I'm focused on improvements that are more or less physically motivated, so I would most likely have missed it. It's a bit disappointing that the scoring function has such a non-physical loophole, though.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 355072,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-07-10T21:45:19.993000",
          "content": "<p>I would not say this particular trick was anticipated, but we've deliberately put emphasis on finding tracks candidates with many hits, rather than having clean tracks. Because cleaning a track from outliers we consider relatively easy to do in a finalizing stage, while missed tracks cannot be easily recovered.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 353820,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2018-07-07T22:54:41.307000",
      "content": "<p>How would the track candidate be selected at evaluation time? Would it be some random set of hits of a random length between 1 and about 20 hits?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 354561,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-09T19:14:28.993000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 354681,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2018-07-10T03:56:28.653000",
          "content": "<p>That makes sense. Using dbscan or a similar preprocessing step to find good candidates sounds like a much better approach, that would actually work, than random selection. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354823,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-10T10:16:15.580000",
          "content": "<p>@Jack,   @Heng shared quite a few approaches of how to find initial tracks. I shared our results too, e.g. with dbscan, we can find over 92% of initial tracks in the volume 7 and 9 (seeding area) but only 68% in the volume 8 since it's very dense. \n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59635\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59635</a></p>\n\n<p>dbscan or clustering in general is actually a very powerful tool to find your candidates and then you can extend your tracks from there.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 355131,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2018-07-11T03:00:52.723000",
          "content": "<p>Thanks Nicole, very helpful information. Is volume 13 much the same case as volume 8, lower accuracy using dbscan to find seed tracks since the hit count is much denser than volume 7 and 9? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 353816,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-07T22:15:45.267000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 353861,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-08T05:27:54.803000",
          "content": "<p>I am using something similar, to remove points that are too far from the best helix fit.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "353772": "let a track candidate be given as a sequence of T= { (i0,p0), (i1,p1) .... (iN,pN) }\n\nHere p can be (x,y,z) or other representation like (a,z/r,z) or an embedded feature that encode a hit and its neighborhood. Let assume p =(x,yz)\n\nAnd i is an unique int that identify the i=(volume_id,layer_id). Assume i range from 0 to L.\n\nCreate a zero Nx3 array. Set  array[i]=p. Use this array as input feature to train a classifier or ranking function.\n\nNote:\n\n1.  the temporal information is encoded in i.\n\n2. In case where a track has multiple hit in a single i=(volume_id,layer_id), represent this track in multiple sequence, each has only one hit in a single i.\n\nto be updated:  code and results .\n\ne.g. label is 1 or 0\n\n{ (i0,p0), (i1,p1) .... (iN,pN) }  --&gt; fixed length array --&gt;  [linear network] --&gt; {label0, label1 ....labelN}\n\n\n\n\n\n",
    "353913": "＠Heng, interesting, I just tried a very similar approach yesterday but my simple linear NN trained with 100 events doesn't work too well on this, the accuracy rate was only about `58%` , a random guess would outperform my NN :) that's why I don't bother to share my not-working code here.  Incidentally, `module_id` can be useful for double hits but most of time very noisy, have you thought about how to add `module_id` as train data too?\n\n( My human random guess would outperform this prediction I'm sure.....) \n\n    Groundtruth\n    [ [ 1.] [ 0.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 0.] [ 0.] [ 1.] [ 1.] [ 0.] [ 1.] [ 1.] [ 1.] [ 1.] [ 0.] [ 1.] ]\n    \n    Prediction \n    [[ 0.48636526]\n     [ 0.55879432]\n     [ 0.50961298]\n     [ 0.48604107]\n     [ 0.61274189]\n     [ 0.49084449]\n     [ 0.92523319]\n     [ 0.49379969]\n     [ 0.56074548]\n     [ 0.50928968]\n     [ 0.50403577]\n     [ 0.48605412]\n     [ 0.88126868]\n     [ 0.65616256]\n     [ 0.49116713]\n     [ 0.92684376]\n     [ 0.61099911]\n     [ 0.88783711]\n     [ 0.50963062]\n     [ 0.60075444]]\n\n",
    "354020": "A track with higher probability of occurrence should a higher score. This is important in ensembling results (or merging track), where we suppress low score tracks with high score one if they overlap (aka. non max suppression)\n\nHere are some rules , based on my observation. Given 2 tracks (in a,z/r,z coordinates) that are well fitted by some parametric equations:\n\n\"1\" is highest score\n\n1.  longest length, i.e. high hit count\n\n2. start from lowest z, transverse sequentially in layers (not missing any layers)\n\n3. straight track  \n\n4. Low error from curve fitting\n\n5. does not cross other track\n\n\nOther observations:\n\n1. Do not treat all tracks the same.  Grow and extend high confident track first\n\n2. if you have difficult in using supervised learning, apply training and testing at for \"potential high confident tracks\" tracks only. E.g. take care of those that are straight, long tracks, assume they transverse sequentially in layers , etc ...\n\n3. To increase LB score, always try to link the hits that re regarded as background noise (labelled 0) in the \"last step\",.This is false positive FP don't give lower LB score because of the metric. We just need to ensure FB don't break up the inlier tracks. \n\n4. the curvature, length, etc of track is related  a, z/r. use this to improve outlier removal.  \n \n \n... to be updated ...",
    "353820": "How would the track candidate be selected at evaluation time? Would it be some random set of hits of a random length between 1 and about 20 hits?",
    "353816": ""
  }
}