{
  "id": 59053,
  "title": "How many features are you using?",
  "url": "/competitions/trackml-particle-identification/discussion/59053",
  "author_name": "",
  "post_date": "2018-06-18T01:19:41.069921900Z",
  "votes": 5,
  "comment_count": 41,
  "views": 0,
  "content": "<p>Hi, everyone!</p>\n\n<p>How are you guys dealing with dimensionality? From your experience, increasing the number of features used are good or bad to correct labeling? The best score I've got until now had 5 features being used, while my teammate got a similar score using 7-9 features that were not necessarily the same ones I had used. What are your opinions (and experiences, if you want to talk about it)? </p>",
  "messages": [
    {
      "id": "344428",
      "postDate": "06/18/2018 01:19:41",
      "content": "<p>Hi, everyone!</p>\n\n<p>How are you guys dealing with dimensionality? From your experience, increasing the number of features used are good or bad to correct labeling? The best score I've got until now had 5 features being used, while my teammate got a similar score using 7-9 features that were not necessarily the same ones I had used. What are your opinions (and experiences, if you want to talk about it)? </p>",
      "rawMarkdown": "Hi, everyone!\n\nHow are you guys dealing with dimensionality? From your experience, increasing the number of features used are good or bad to correct labeling? The best score I've got until now had 5 features being used, while my teammate got a similar score using 7-9 features that were not necessarily the same ones I had used. What are your opinions (and experiences, if you want to talk about it)?",
      "votes": null
    },
    {
      "id": "344431",
      "postDate": "06/18/2018 01:29:53",
      "content": "<p>I use 6 features. I believe obtaining the same score with fewer features is advantageous because it will run faster, right?</p>",
      "rawMarkdown": "I use 6 features. I believe obtaining the same score with fewer features is advantageous because it will run faster, right?",
      "votes": null
    },
    {
      "id": "344434",
      "postDate": "06/18/2018 01:50:55",
      "content": "<p>My best model, which gets a score of 0.575 uses 15 features!  Most of the other models I am ensembling have 5 features.</p>",
      "rawMarkdown": "My best model, which gets a score of 0.575 uses 15 features!  Most of the other models I am ensembling have 5 features.",
      "votes": null
    },
    {
      "id": "344438",
      "postDate": "06/18/2018 02:19:05",
      "content": "<p>Referring to DBSCAN:\nIf you consider the sin&amp;cos of an angle a single feature, then I'm using 9 features. I've found in total ~15 that will occasionally give some lift, but many of them would only improve score very slightly on average and would actually reduce score on certain local evaluations. </p>\n\n<p>I'm not currently looking into more features, I'm following a hunch that there will be big improvements outside of DBSCAN features, although I might circle back to features in the future. </p>",
      "rawMarkdown": "Referring to DBSCAN:\nIf you consider the sin&amp;cos of an angle a single feature, then I'm using 9 features. I've found in total ~15 that will occasionally give some lift, but many of them would only improve score very slightly on average and would actually reduce score on certain local evaluations. \n\nI'm not currently looking into more features, I'm following a hunch that there will be big improvements outside of DBSCAN features, although I might circle back to features in the future.",
      "votes": null
    },
    {
      "id": "344519",
      "postDate": "06/18/2018 06:15:29",
      "content": "<p>I use only 2 features to get 0.62 (I count sin&amp;cos as one feature)</p>",
      "rawMarkdown": "I use only 2 features to get 0.62 (I count sin&amp;cos as one feature)",
      "votes": null
    },
    {
      "id": "344522",
      "postDate": "06/18/2018 06:21:39",
      "content": "<p>Wow - very interesting!  How many models are you ensembling?</p>",
      "rawMarkdown": "Wow - very interesting!  How many models are you ensembling?",
      "votes": null
    },
    {
      "id": "344546",
      "postDate": "06/18/2018 07:38:28",
      "content": "<p>None, yet.  </p>",
      "rawMarkdown": "None, yet.",
      "votes": null
    },
    {
      "id": "344555",
      "postDate": "06/18/2018 07:57:59",
      "content": "<p>Ensembling may mean different things here.  The best public kernel iterates over a number of angle corrections.  It ensembles tracks found for each angle prediction.  In that way most of us are using some ensembling.  I guess you are discussing another form of ensembling, which would be to, say, create a new submission out of two or more submissions.</p>",
      "rawMarkdown": "Ensembling may mean different things here.  The best public kernel iterates over a number of angle corrections.  It ensembles tracks found for each angle prediction.  In that way most of us are using some ensembling.  I guess you are discussing another form of ensembling, which would be to, say, create a new submission out of two or more submissions.",
      "votes": null
    },
    {
      "id": "344558",
      "postDate": "06/18/2018 08:13:18",
      "content": "<p>Our strongest single model of the latest submission gives us around <code>0.58</code> using <code>sin and cos</code> + basic features in the chemist's kernel.  And then we ensembled other 4-5 weaker models to give us over  <code>0.632</code>. However, if I made each model stronger, says the strongest model with <code>0.59</code>, then the local score goes down to <code>0.625</code> when we merge them together. Ensembling is a very hard thing to do in this competition. Using more features doesn't necessarily help if you look at the tracks dbscan finds, some features are contradicting to each other. However, as @CPMP pointed out, the chemist's kernel is already \"ensembling\" tracks for each scan in some sense.  </p>",
      "rawMarkdown": "Our strongest single model of the latest submission gives us around `0.58` using `sin and cos` + basic features in the chemist's kernel.  And then we ensembled other 4-5 weaker models to give us over  `0.632`. However, if I made each model stronger, says the strongest model with `0.59`, then the local score goes down to `0.625` when we merge them together. Ensembling is a very hard thing to do in this competition. Using more features doesn't necessarily help if you look at the tracks dbscan finds, some features are contradicting to each other. However, as @CPMP pointed out, the chemist's kernel is already \"ensembling\" tracks for each scan in some sense.",
      "votes": null
    },
    {
      "id": "344575",
      "postDate": "06/18/2018 08:46:13",
      "content": "<p>You are correct. I don't call these runtime iterations ensembles (I do &gt;10K iterations).</p>",
      "rawMarkdown": "You are correct. I don't call these runtime iterations ensembles (I do &gt;10K iterations).",
      "votes": null
    },
    {
      "id": "344812",
      "postDate": "06/18/2018 17:59:28",
      "content": "<p>I use 3 features if cos/sin counts for 1 in the DBSCAN part of my code to get 0.66.  But my code does a bit more than just running DBSCAN in a loop ;)</p>\n\n<p>I am not ensembling models outside the DBSCAN loop.  This means there is hope when working on a single model ;)</p>",
      "rawMarkdown": "I use 3 features if cos/sin counts for 1 in the DBSCAN part of my code to get 0.66.  But my code does a bit more than just running DBSCAN in a loop ;)\n\nI am not ensembling models outside the DBSCAN loop.  This means there is hope when working on a single model ;)",
      "votes": null
    },
    {
      "id": "344880",
      "postDate": "06/18/2018 20:55:57",
      "content": "<p>Wow, 3 features, 1 model, 0.66 for you, 2 features, 1 model, 0.62 for yuval...</p>\n\n<p>I didn't expect that!  Makes this challenge fun!</p>",
      "rawMarkdown": "Wow, 3 features, 1 model, 0.66 for you, 2 features, 1 model, 0.62 for yuval...\n\nI didn't expect that!  Makes this challenge fun!",
      "votes": null
    },
    {
      "id": "344886",
      "postDate": "06/18/2018 21:07:47",
      "content": "<p>Indeed, ensembling has its limit, merging two models can only give you a benefit of 0.01 (naive merging, what we're using) ~0.04 (if all the hits belong to the right track, very unlikely to achieve, we haven't achieved this yet), so I have to revisit the single strong model approach again. </p>",
      "rawMarkdown": "Indeed, ensembling has its limit, merging two models can only give you a benefit of 0.01 (naive merging, what we're using) ~0.04 (if all the hits belong to the right track, very unlikely to achieve, we haven't achieved this yet), so I have to revisit the single strong model approach again.",
      "votes": null
    },
    {
      "id": "344894",
      "postDate": "06/18/2018 21:26:20",
      "content": "<p>On the other hand, an improvement of 0.01 is nice to find!  Of course, for me I think they cost about 0.00025 improvement per hour of lost sleep!</p>",
      "rawMarkdown": "On the other hand, an improvement of 0.01 is nice to find!  Of course, for me I think they cost about 0.00025 improvement per hour of lost sleep!",
      "votes": null
    },
    {
      "id": "344917",
      "postDate": "06/18/2018 23:07:45",
      "content": "<p>@yuval r, 10k iterations, wow! We've run 240-480 iterations for each model, however, each model is weaker than 0.6. I guess from what you and @CPMP posted, this is something we can still work on.</p>",
      "rawMarkdown": "yuval r, 10k iterations, wow! We've run 240-480 iterations for each model, however, each model is weaker than 0.6. I guess from what you and @CPMP posted, this is something we can still work on.",
      "votes": null
    },
    {
      "id": "344934",
      "postDate": "06/19/2018 00:19:58",
      "content": "<p>@Nicole Finnie, yes definitely a ways to go from the public model! My current idea of improving the tracks with more iterations is by shifting the origin, mostly in the z-direction since the luminous region has a longitudinal width of 55mm and I've noticed ground truth tracks that I thought this would capture. However, I think I'm either not doing something right or my idea is flawed because it is pretty detrimental to the score. Thoughts?</p>",
      "rawMarkdown": "Nicole Finnie, yes definitely a ways to go from the public model! My current idea of improving the tracks with more iterations is by shifting the origin, mostly in the z-direction since the luminous region has a longitudinal width of 55mm and I've noticed ground truth tracks that I thought this would capture. However, I think I'm either not doing something right or my idea is flawed because it is pretty detrimental to the score. Thoughts?",
      "votes": null
    },
    {
      "id": "344943",
      "postDate": "06/19/2018 01:03:12",
      "content": "<p>@Matthew, I'm not sure if shifting the origin along the z-axis would help since our features are cylindrical and those are constant traits approximately shared among the hits that belong to the same track, e.g. constant radius space. Shifting hits along the Z-axis probably wouldn't make dbscan(or the clustering approach you use) to find more tracks since their relation to the origin wouldn't change, such as <code>r, sin(a) cos(a)</code> would stay the same. @CPMP you're a mathematician/physicist, any thoughts? </p>",
      "rawMarkdown": "Matthew, I'm not sure if shifting the origin along the z-axis would help since our features are cylindrical and those are constant traits approximately shared among the hits that belong to the same track, e.g. constant radius space. Shifting hits along the Z-axis probably wouldn't make dbscan(or the clustering approach you use) to find more tracks since their relation to the origin wouldn't change, such as `r, sin(a) cos(a)` would stay the same. @CPMP you're a mathematician/physicist, any thoughts?",
      "votes": null
    },
    {
      "id": "344948",
      "postDate": "06/19/2018 01:15:42",
      "content": "<p>How I imagined it working was by shifting the hits in their xyz coordinates and then recomputing the features to use for each shift. I agree it wouldn't change features such as sin,cos,and r but my other features that use z and d are changed from the shift. I'm really not sure this is even a good idea so I'm glad to hear feedback. :)</p>",
      "rawMarkdown": "How I imagined it working was by shifting the hits in their xyz coordinates and then recomputing the features to use for each shift. I agree it wouldn't change features such as sin,cos,and r but my other features that use z and d are changed from the shift. I'm really not sure this is even a good idea so I'm glad to hear feedback. :)",
      "votes": null
    },
    {
      "id": "344950",
      "postDate": "06/19/2018 01:21:29",
      "content": "<p>@Nicole, not sure I agree with you.  Many features shared publicly assume tracks start with z = 0, be it in the helix unrolling kernels, or Heng's track extension.  Finding tracks that do not start at (0,0,0) is certainly key to win this.  I'm not very good at this for now either.</p>",
      "rawMarkdown": "Nicole, not sure I agree with you.  Many features shared publicly assume tracks start with z = 0, be it in the helix unrolling kernels, or Heng's track extension.  Finding tracks that do not start at (0,0,0) is certainly key to win this.  I'm not very good at this for now either.",
      "votes": null
    },
    {
      "id": "344951",
      "postDate": "06/19/2018 01:22:43",
      "content": "<p>@CPMP Ooo, interesting. Thank you for your input!</p>",
      "rawMarkdown": "CPMP Ooo, interesting. Thank you for your input!",
      "votes": null
    },
    {
      "id": "344955",
      "postDate": "06/19/2018 01:29:19",
      "content": "<p>@CPMP, true, I didn't look at that aspect (starting particles). The starting hits of those tracks may not be detected by the detectors in the first place. The tracks of &lt; 4 hits (57% of them don't start from the origin) don't have any weights. Almost all tracks made up by more than 9 hits can be detected by the innermost detectors. The ones we could tackle here would be the tracks of between 4 and 9 hits.</p>\n\n<p>*<em>EDIT on June 23rd *</em></p>\n\n<pre><code>For hits = 9, train event=1000\n42 out of 551 tracks - hits cannot be detected by the innermost detectors\n0.07622504537205081% of the tracks -  hits cannot be detected in the innermost detectors\n</code></pre>\n\n<p><a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing</a></p>",
      "rawMarkdown": "CPMP, true, I didn't look at that aspect (starting particles). The starting hits of those tracks may not be detected by the detectors in the first place. The tracks of &lt; 4 hits (57% of them don't start from the origin) don't have any weights. Almost all tracks made up by more than 9 hits can be detected by the innermost detectors. The ones we could tackle here would be the tracks of between 4 and 9 hits.\n\n**EDIT on June 23rd **\n\n    For hits = 9, train event=1000\n    42 out of 551 tracks - hits cannot be detected by the innermost detectors\n    0.07622504537205081% of the tracks -  hits cannot be detected in the innermost detectors\n\nhttps://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing",
      "votes": null
    },
    {
      "id": "344959",
      "postDate": "06/19/2018 01:41:44",
      "content": "<blockquote>\n  <p>All tracks made up by more than 8 hits start from the origin</p>\n</blockquote>\n\n<p>I didn't realize this, thanks for sharing.  </p>",
      "rawMarkdown": "&gt; All tracks made up by more than 8 hits start from the origin\n\nI didn't realize this, thanks for sharing.",
      "votes": null
    },
    {
      "id": "344977",
      "postDate": "06/19/2018 02:12:34",
      "content": "<p>The unrolling strategy considers a new track valid if it has more hits in the track and the new track has less than 20 hits... </p>\n\n<p>Similarly an unrolling strategy with an alternative z origin could only assign a new track if the total hits in the new track are 8 or less. This would reduce the amount of false positives... </p>",
      "rawMarkdown": "The unrolling strategy considers a new track valid if it has more hits in the track and the new track has less than 20 hits... \n\nSimilarly an unrolling strategy with an alternative z origin could only assign a new track if the total hits in the new track are 8 or less. This would reduce the amount of false positives...",
      "votes": null
    },
    {
      "id": "344984",
      "postDate": "06/19/2018 02:28:44",
      "content": "<p>@Nicole - Is your data about track length vs location of origin based on input from the organizers or on analysis of the \"truth\"???  Thanks</p>",
      "rawMarkdown": "Nicole - Is your data about track length vs location of origin based on input from the organizers or on analysis of the \"truth\"???  Thanks",
      "votes": null
    },
    {
      "id": "344990",
      "postDate": "06/19/2018 02:44:15",
      "content": "<p><a href=\"/macfarll\">@macfarll</a>, Right. I am just experimenting with this now. It still seems to hurt the model though. I feel like it is because the interval I use for shifting z is too large. However, making this interval smaller means a combinatoric explosion since I run around 240 dbscan iterations per z-shift.</p>\n\n<p>@John @Nicole, I am also curious about this. Is this statistic true for all events?</p>",
      "rawMarkdown": "macfarll, Right. I am just experimenting with this now. It still seems to hurt the model though. I feel like it is because the interval I use for shifting z is too large. However, making this interval smaller means a combinatoric explosion since I run around 240 dbscan iterations per z-shift.\n\n@John @Nicole, I am also curious about this. Is this statistic true for all events?",
      "votes": null
    },
    {
      "id": "345002",
      "postDate": "06/19/2018 03:45:30",
      "content": "<p>Hmm...well then the question becomes what is going to help your solution more, more iterations per origin, or more origins? I'd imagine both will give diminishing returns so you'll just need to find the right sweetspot.</p>\n\n<p>240 iterations seems like a ton to me... </p>",
      "rawMarkdown": "Hmm...well then the question becomes what is going to help your solution more, more iterations per origin, or more origins? I'd imagine both will give diminishing returns so you'll just need to find the right sweetspot.\n\n240 iterations seems like a ton to me...",
      "votes": null
    },
    {
      "id": "345045",
      "postDate": "06/19/2018 05:26:39",
      "content": "<p>@Nicole, It is probably strange to many others of the score above 0.6 (at least to me), that you have obtained 0.63 without shifting z. Ensembling 7 shifted models can give boost 0.08.</p>",
      "rawMarkdown": "Nicole, It is probably strange to many others of the score above 0.6 (at least to me), that you have obtained 0.63 without shifting z. Ensembling 7 shifted models can give boost 0.08.",
      "votes": null
    },
    {
      "id": "345050",
      "postDate": "06/19/2018 05:39:13",
      "content": "<p>@John, EDA of the ground truth, I only looked at one event, but I assume every event is representative. </p>",
      "rawMarkdown": "John, EDA of the ground truth, I only looked at one event, but I assume every event is representative.",
      "votes": null
    },
    {
      "id": "345052",
      "postDate": "06/19/2018 05:40:24",
      "content": "<p>@Grzegorz, sounds like the right way to go then. I'll give it a try. <code>0.63</code> is nothing when there's <code>0.8</code> :)</p>",
      "rawMarkdown": "Grzegorz, sounds like the right way to go then. I'll give it a try. `0.63` is nothing when there's `0.8` :)",
      "votes": null
    },
    {
      "id": "345057",
      "postDate": "06/19/2018 05:52:12",
      "content": "<p>0.8 it is almost all tracks of the origin (0,0,z) making less than 0.5 turn inside the detector. It will be a dream and the limit for many of us. The jump from 0.8 to 0.9 would be a quality jump, not simply quantity one. </p>\n\n<p>EDIT: <a href=\"https://www.kaggle.com/sionek/score-limit-for-0-0-z\">https://www.kaggle.com/sionek/score-limit-for-0-0-z</a></p>",
      "rawMarkdown": "0.8 it is almost all tracks of the origin (0,0,z) making less than 0.5 turn inside the detector. It will be a dream and the limit for many of us. The jump from 0.8 to 0.9 would be a quality jump, not simply quantity one. \n\nEDIT: https://www.kaggle.com/sionek/score-limit-for-0-0-z",
      "votes": null
    },
    {
      "id": "345059",
      "postDate": "06/19/2018 05:57:43",
      "content": "<p>I can do a large number of iterations because I don't use  DBSCAN  but a very simple and extremely fast clustering (using bins). \nUsing different assumptions for the origin in the Z axis is very important (X,Y are a waste of time). \nNow we have a new challenge - the  0.8 ;)</p>",
      "rawMarkdown": "I can do a large number of iterations because I don't use  DBSCAN  but a very simple and extremely fast clustering (using bins). \nUsing different assumptions for the origin in the Z axis is very important (X,Y are a waste of time). \nNow we have a new challenge - the  0.8 ;)",
      "votes": null
    },
    {
      "id": "345101",
      "postDate": "06/19/2018 07:50:36",
      "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>thanks for the hint. I confirm changing from z =z +dz improves my results on my validation set. Each dz gives potentially a boost of 0.01 ... but the run time exploded as well.</p>",
      "rawMarkdown": "Grzegorz Sionkowski\n\nthanks for the hint. I confirm changing from z =z +dz improves my results on my validation set. Each dz gives potentially a boost of 0.01 ... but the run time exploded as well.",
      "votes": null
    },
    {
      "id": "345142",
      "postDate": "06/19/2018 09:05:06",
      "content": "<p>@CPMP, just checked out my EDA result and found out almost all tracks made up of more than <code>hits = 9</code> are from the origin. I corrected my comment. However, they're still in the helix-like form. </p>\n\n<p>I added visualization here\n<a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing</a></p>",
      "rawMarkdown": "CPMP, just checked out my EDA result and found out almost all tracks made up of more than `hits = 9` are from the origin. I corrected my comment. However, they're still in the helix-like form. \n\nI added visualization here\nhttps://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing",
      "votes": null
    },
    {
      "id": "345993",
      "postDate": "06/20/2018 21:42:24",
      "content": "<p>@Nicole - how did you determine which tracks were from the origin? \nI took a look in the particles file at the vx, vy, vz fields which are the 'initial position or vertex (in millimeters) in global coordinates' - if this means the origin of the particles then there are many away from the origin. I haven't analysed the length of these tracks but I took a quick look at the effect on the score and it seemed to be significant</p>",
      "rawMarkdown": "Nicole - how did you determine which tracks were from the origin? \nI took a look in the particles file at the vx, vy, vz fields which are the 'initial position or vertex (in millimeters) in global coordinates' - if this means the origin of the particles then there are many away from the origin. I haven't analysed the length of these tracks but I took a quick look at the effect on the score and it seemed to be significant",
      "votes": null
    },
    {
      "id": "346003",
      "postDate": "06/20/2018 21:59:26",
      "content": "<p>@Seb, we could have been wrong with our approach, but we were looking at the geometry information provided with the hits to give us hints as to which hits are considered close to the origin. Our approach is very rough, we likely missed many. No hits start at exactly to 0,0,0 - but we were not looking at vx,vy,vz directly.</p>",
      "rawMarkdown": "Seb, we could have been wrong with our approach, but we were looking at the geometry information provided with the hits to give us hints as to which hits are considered close to the origin. Our approach is very rough, we likely missed many. No hits start at exactly to 0,0,0 - but we were not looking at vx,vy,vz directly.",
      "votes": null
    },
    {
      "id": "346012",
      "postDate": "06/20/2018 22:55:50",
      "content": "<p>I remember you saying you haven't used z shifted models yet, which would be my first guess at what you meant by a weak model... </p>\n\n<p>Looking back at this comment it seems like it might have been a much more valuable hint than I previously realized. Thanks for sharing. </p>\n\n<p>Also, I like the term 'the chemist'.</p>",
      "rawMarkdown": "I remember you saying you haven't used z shifted models yet, which would be my first guess at what you meant by a weak model... \n\nLooking back at this comment it seems like it might have been a much more valuable hint than I previously realized. Thanks for sharing. \n\nAlso, I like the term 'the chemist'.",
      "votes": null
    },
    {
      "id": "346015",
      "postDate": "06/20/2018 23:14:10",
      "content": "<p><a href=\"/macfarll\">@macfarll</a>,  Yes, we haven't used z shifted models yet, work keeps us busy. For the weak models I mean the models that yield a lower mean score. They may find different tracks and you can ensemble them, the downside is that a weak model can be very noisy and tends to \"overwrite\" the good tracks from a strong model and that makes ensembling pretty challenging. We're moving away from the merging approach tho, since @CPMP's / @yuval r's LB scores proved a single model is much stronger than a merged one.</p>\n\n<p>P.S. Yes, I like @Grzegorz's title \"the chemist\" since I don't know how to pronounce his name in Polish. :) I guess it's probably similar to \"Gregor\" in German. </p>",
      "rawMarkdown": "macfarll,  Yes, we haven't used z shifted models yet, work keeps us busy. For the weak models I mean the models that yield a lower mean score. They may find different tracks and you can ensemble them, the downside is that a weak model can be very noisy and tends to \"overwrite\" the good tracks from a strong model and that makes ensembling pretty challenging. We're moving away from the merging approach tho, since @CPMP's / @yuval r's LB scores proved a single model is much stronger than a merged one.\n\nP.S. Yes, I like @Grzegorz's title \"the chemist\" since I don't know how to pronounce his name in Polish. :) I guess it's probably similar to \"Gregor\" in German.",
      "votes": null
    },
    {
      "id": "347450",
      "postDate": "06/24/2018 11:26:18",
      "content": "<p>@Seb just found out a couple days ago our original EDA was to find out \"the hits that don't pass the inner most detectors with the volume ID 7, 8, 9\" sorry for the misleading information. I updated my comment.</p>\n\n<p>Using detector information can be helpful for classifying the hits (before dbscan) and removing outliers, since the current dbscan is only good for predicting helix-like tracks, I remember the <a href=\"/chemist\">@chemist</a> or @CPMP said we missed approximately 18% of the tracks only using dbscan. </p>\n\n<p>our approximate approach to filter on hits from the origin is  </p>\n\n<pre><code>       if np.absolute(particle_list[i].vx) &lt; 0.01 and np.absolute(particle_list[i].vy) &lt; 0.01 and np.absolute(particle_list[i].vz) &lt; 0.01:\n</code></pre>\n\n<p>For example, for the train event 0001054 with all tracks of the length of 12 hits:</p>\n\n<pre><code> 251 tracks out of 1338  come from the origin\n Hits counts: 12\n</code></pre>",
      "rawMarkdown": "Seb just found out a couple days ago our original EDA was to find out \"the hits that don't pass the inner most detectors with the volume ID 7, 8, 9\" sorry for the misleading information. I updated my comment.\n\nUsing detector information can be helpful for classifying the hits (before dbscan) and removing outliers, since the current dbscan is only good for predicting helix-like tracks, I remember the @chemist or @CPMP said we missed approximately 18% of the tracks only using dbscan. \n\n our approximate approach to filter on hits from the origin is  \n\n           if np.absolute(particle_list[i].vx) &lt; 0.01 and np.absolute(particle_list[i].vy) &lt; 0.01 and np.absolute(particle_list[i].vz) &lt; 0.01:\n\nFor example, for the train event 0001054 with all tracks of the length of 12 hits:\n\n  \n     251 tracks out of 1338  come from the origin\n     Hits counts: 12",
      "votes": null
    },
    {
      "id": "347491",
      "postDate": "06/24/2018 13:56:48",
      "content": "<p>@Finnies Thanks for your responses. I haven't looked using the detector information much yet, it's an interesting idea.\nMy concern with dbscan at the moment is the dependency on the origin - sure you can shift it and get better scores, but with small increases in score for large increases in runtime, I feel like I should try another way... I just haven't figured out what the other way should be yet!\nGiven that the current leader seems to have a good track record in image based competitions, I'd speculate that the same sort of techniques might be used in this competition. Unfortunately for me image recognition is not something I've had that much experience of so I've been doing more reading/research/thinking this week than coding</p>",
      "rawMarkdown": "Finnies Thanks for your responses. I haven't looked using the detector information much yet, it's an interesting idea.\nMy concern with dbscan at the moment is the dependency on the origin - sure you can shift it and get better scores, but with small increases in score for large increases in runtime, I feel like I should try another way... I just haven't figured out what the other way should be yet!\nGiven that the current leader seems to have a good track record in image based competitions, I'd speculate that the same sort of techniques might be used in this competition. Unfortunately for me image recognition is not something I've had that much experience of so I've been doing more reading/research/thinking this week than coding",
      "votes": null
    },
    {
      "id": "347540",
      "postDate": "06/24/2018 16:19:25",
      "content": "<p>@Seb  hey, z-shifting has its limits, without z-shifting, our best single model was around 0.625 with the chemist's kernel's features, with z-shifting, it gave us up to 0.035 with the same features for a single model, that's the limit of the z-shifting I can see, since we still missed tons of tracks outside the (x,y)=(0,0) . And you're right, a small increase in score in dbscan does increase runtime in lots of cases. And you're very observant with the current leader's past experience ;)  I will move away from the unsupervised learning approach as well. However, everything you've done for dbscan (clustering, post processing, outlier removal, track fitting) will be useful for your next step. Learning/reading/researching is the most important and fun part of kaggling isn't it? :)</p>",
      "rawMarkdown": "Seb  hey, z-shifting has its limits, without z-shifting, our best single model was around 0.625 with the chemist's kernel's features, with z-shifting, it gave us up to 0.035 with the same features for a single model, that's the limit of the z-shifting I can see, since we still missed tons of tracks outside the (x,y)=(0,0) . And you're right, a small increase in score in dbscan does increase runtime in lots of cases. And you're very observant with the current leader's past experience ;)  I will move away from the unsupervised learning approach as well. However, everything you've done for dbscan (clustering, post processing, outlier removal, track fitting) will be useful for your next step. Learning/reading/researching is the most important and fun part of kaggling isn't it? :)",
      "votes": null
    },
    {
      "id": "350413",
      "postDate": "06/29/2018 18:07:34",
      "content": "<p>@Nicole you're right, learning stuff / improving is the best part - it's why I'm doing this. It seems like there's a lot to learn on this one, especially with <a href=\"/outrunner\">@outrunner</a>'s 0.9 score not too far away now! Everyday life has an annoying habit of getting in the way of time spent on Kaggle! :)</p>",
      "rawMarkdown": "Nicole you're right, learning stuff / improving is the best part - it's why I'm doing this. It seems like there's a lot to learn on this one, especially with @outrunner's 0.9 score not too far away now! Everyday life has an annoying habit of getting in the way of time spent on Kaggle! :)",
      "votes": null
    },
    {
      "id": "350416",
      "postDate": "06/29/2018 18:12:49",
      "content": "<p>@Seb, exactly, <a href=\"/outrunner\">@outrunner</a> proved how far we can still get, we may not be able to replicate his results but we can learn his approach after the end of this competition.  I've been struggling implementing a new model, no luck so far.  (Just like the German National Football Team, no luck, ouch ouch ouch!! )</p>",
      "rawMarkdown": "Seb, exactly, @outrunner proved how far we can still get, we may not be able to replicate his results but we can learn his approach after the end of this competition.  I've been struggling implementing a new model, no luck so far.  (Just like the German National Football Team, no luck, ouch ouch ouch!! )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 344431,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "06/18/2018 01:29:53",
      "content": "<p>I use 6 features. I believe obtaining the same score with fewer features is advantageous because it will run faster, right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 344434,
      "author_name": "johnhsweeney",
      "author_url": "",
      "post_date": "06/18/2018 01:50:55",
      "content": "<p>My best model, which gets a score of 0.575 uses 15 features!  Most of the other models I am ensembling have 5 features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 344438,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "06/18/2018 02:19:05",
      "content": "<p>Referring to DBSCAN:\nIf you consider the sin&amp;cos of an angle a single feature, then I'm using 9 features. I've found in total ~15 that will occasionally give some lift, but many of them would only improve score very slightly on average and would actually reduce score on certain local evaluations. </p>\n\n<p>I'm not currently looking into more features, I'm following a hunch that there will be big improvements outside of DBSCAN features, although I might circle back to features in the future. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 344519,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "06/18/2018 06:15:29",
      "content": "<p>I use only 2 features to get 0.62 (I count sin&amp;cos as one feature)</p>",
      "votes": null,
      "replies": [
        {
          "id": 344522,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "06/18/2018 06:21:39",
          "content": "<p>Wow - very interesting!  How many models are you ensembling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344546,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "06/18/2018 07:38:28",
          "content": "<p>None, yet.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344555,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/18/2018 07:57:59",
          "content": "<p>Ensembling may mean different things here.  The best public kernel iterates over a number of angle corrections.  It ensembles tracks found for each angle prediction.  In that way most of us are using some ensembling.  I guess you are discussing another form of ensembling, which would be to, say, create a new submission out of two or more submissions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344575,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "06/18/2018 08:46:13",
          "content": "<p>You are correct. I don't call these runtime iterations ensembles (I do &gt;10K iterations).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344917,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/18/2018 23:07:45",
          "content": "<p>@yuval r, 10k iterations, wow! We've run 240-480 iterations for each model, however, each model is weaker than 0.6. I guess from what you and @CPMP posted, this is something we can still work on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344934,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "06/19/2018 00:19:58",
          "content": "<p>@Nicole Finnie, yes definitely a ways to go from the public model! My current idea of improving the tracks with more iterations is by shifting the origin, mostly in the z-direction since the luminous region has a longitudinal width of 55mm and I've noticed ground truth tracks that I thought this would capture. However, I think I'm either not doing something right or my idea is flawed because it is pretty detrimental to the score. Thoughts?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344943,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/19/2018 01:03:12",
          "content": "<p>@Matthew, I'm not sure if shifting the origin along the z-axis would help since our features are cylindrical and those are constant traits approximately shared among the hits that belong to the same track, e.g. constant radius space. Shifting hits along the Z-axis probably wouldn't make dbscan(or the clustering approach you use) to find more tracks since their relation to the origin wouldn't change, such as <code>r, sin(a) cos(a)</code> would stay the same. @CPMP you're a mathematician/physicist, any thoughts? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344948,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "06/19/2018 01:15:42",
          "content": "<p>How I imagined it working was by shifting the hits in their xyz coordinates and then recomputing the features to use for each shift. I agree it wouldn't change features such as sin,cos,and r but my other features that use z and d are changed from the shift. I'm really not sure this is even a good idea so I'm glad to hear feedback. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344950,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/19/2018 01:21:29",
          "content": "<p>@Nicole, not sure I agree with you.  Many features shared publicly assume tracks start with z = 0, be it in the helix unrolling kernels, or Heng's track extension.  Finding tracks that do not start at (0,0,0) is certainly key to win this.  I'm not very good at this for now either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344951,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "06/19/2018 01:22:43",
          "content": "<p>@CPMP Ooo, interesting. Thank you for your input!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344955,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/19/2018 01:29:19",
          "content": "<p>@CPMP, true, I didn't look at that aspect (starting particles). The starting hits of those tracks may not be detected by the detectors in the first place. The tracks of &lt; 4 hits (57% of them don't start from the origin) don't have any weights. Almost all tracks made up by more than 9 hits can be detected by the innermost detectors. The ones we could tackle here would be the tracks of between 4 and 9 hits.</p>\n\n<p>*<em>EDIT on June 23rd *</em></p>\n\n<pre><code>For hits = 9, train event=1000\n42 out of 551 tracks - hits cannot be detected by the innermost detectors\n0.07622504537205081% of the tracks -  hits cannot be detected in the innermost detectors\n</code></pre>\n\n<p><a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344959,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/19/2018 01:41:44",
          "content": "<blockquote>\n  <p>All tracks made up by more than 8 hits start from the origin</p>\n</blockquote>\n\n<p>I didn't realize this, thanks for sharing.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344977,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "06/19/2018 02:12:34",
          "content": "<p>The unrolling strategy considers a new track valid if it has more hits in the track and the new track has less than 20 hits... </p>\n\n<p>Similarly an unrolling strategy with an alternative z origin could only assign a new track if the total hits in the new track are 8 or less. This would reduce the amount of false positives... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344984,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "06/19/2018 02:28:44",
          "content": "<p>@Nicole - Is your data about track length vs location of origin based on input from the organizers or on analysis of the \"truth\"???  Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344990,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "06/19/2018 02:44:15",
          "content": "<p><a href=\"/macfarll\">@macfarll</a>, Right. I am just experimenting with this now. It still seems to hurt the model though. I feel like it is because the interval I use for shifting z is too large. However, making this interval smaller means a combinatoric explosion since I run around 240 dbscan iterations per z-shift.</p>\n\n<p>@John @Nicole, I am also curious about this. Is this statistic true for all events?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345002,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "06/19/2018 03:45:30",
          "content": "<p>Hmm...well then the question becomes what is going to help your solution more, more iterations per origin, or more origins? I'd imagine both will give diminishing returns so you'll just need to find the right sweetspot.</p>\n\n<p>240 iterations seems like a ton to me... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345045,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "06/19/2018 05:26:39",
          "content": "<p>@Nicole, It is probably strange to many others of the score above 0.6 (at least to me), that you have obtained 0.63 without shifting z. Ensembling 7 shifted models can give boost 0.08.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345050,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/19/2018 05:39:13",
          "content": "<p>@John, EDA of the ground truth, I only looked at one event, but I assume every event is representative. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345052,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/19/2018 05:40:24",
          "content": "<p>@Grzegorz, sounds like the right way to go then. I'll give it a try. <code>0.63</code> is nothing when there's <code>0.8</code> :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345057,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "06/19/2018 05:52:12",
          "content": "<p>0.8 it is almost all tracks of the origin (0,0,z) making less than 0.5 turn inside the detector. It will be a dream and the limit for many of us. The jump from 0.8 to 0.9 would be a quality jump, not simply quantity one. </p>\n\n<p>EDIT: <a href=\"https://www.kaggle.com/sionek/score-limit-for-0-0-z\">https://www.kaggle.com/sionek/score-limit-for-0-0-z</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345059,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "06/19/2018 05:57:43",
          "content": "<p>I can do a large number of iterations because I don't use  DBSCAN  but a very simple and extremely fast clustering (using bins). \nUsing different assumptions for the origin in the Z axis is very important (X,Y are a waste of time). \nNow we have a new challenge - the  0.8 ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345101,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/19/2018 07:50:36",
          "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>thanks for the hint. I confirm changing from z =z +dz improves my results on my validation set. Each dz gives potentially a boost of 0.01 ... but the run time exploded as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345142,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/19/2018 09:05:06",
          "content": "<p>@CPMP, just checked out my EDA result and found out almost all tracks made up of more than <code>hits = 9</code> are from the origin. I corrected my comment. However, they're still in the helix-like form. </p>\n\n<p>I added visualization here\n<a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345993,
          "author_name": "sjb1988",
          "author_url": "",
          "post_date": "06/20/2018 21:42:24",
          "content": "<p>@Nicole - how did you determine which tracks were from the origin? \nI took a look in the particles file at the vx, vy, vz fields which are the 'initial position or vertex (in millimeters) in global coordinates' - if this means the origin of the particles then there are many away from the origin. I haven't analysed the length of these tracks but I took a quick look at the effect on the score and it seemed to be significant</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346003,
          "author_name": "jliamfinnie",
          "author_url": "",
          "post_date": "06/20/2018 21:59:26",
          "content": "<p>@Seb, we could have been wrong with our approach, but we were looking at the geometry information provided with the hits to give us hints as to which hits are considered close to the origin. Our approach is very rough, we likely missed many. No hits start at exactly to 0,0,0 - but we were not looking at vx,vy,vz directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347450,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/24/2018 11:26:18",
          "content": "<p>@Seb just found out a couple days ago our original EDA was to find out \"the hits that don't pass the inner most detectors with the volume ID 7, 8, 9\" sorry for the misleading information. I updated my comment.</p>\n\n<p>Using detector information can be helpful for classifying the hits (before dbscan) and removing outliers, since the current dbscan is only good for predicting helix-like tracks, I remember the <a href=\"/chemist\">@chemist</a> or @CPMP said we missed approximately 18% of the tracks only using dbscan. </p>\n\n<p>our approximate approach to filter on hits from the origin is  </p>\n\n<pre><code>       if np.absolute(particle_list[i].vx) &lt; 0.01 and np.absolute(particle_list[i].vy) &lt; 0.01 and np.absolute(particle_list[i].vz) &lt; 0.01:\n</code></pre>\n\n<p>For example, for the train event 0001054 with all tracks of the length of 12 hits:</p>\n\n<pre><code> 251 tracks out of 1338  come from the origin\n Hits counts: 12\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347491,
          "author_name": "sjb1988",
          "author_url": "",
          "post_date": "06/24/2018 13:56:48",
          "content": "<p>@Finnies Thanks for your responses. I haven't looked using the detector information much yet, it's an interesting idea.\nMy concern with dbscan at the moment is the dependency on the origin - sure you can shift it and get better scores, but with small increases in score for large increases in runtime, I feel like I should try another way... I just haven't figured out what the other way should be yet!\nGiven that the current leader seems to have a good track record in image based competitions, I'd speculate that the same sort of techniques might be used in this competition. Unfortunately for me image recognition is not something I've had that much experience of so I've been doing more reading/research/thinking this week than coding</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347540,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/24/2018 16:19:25",
          "content": "<p>@Seb  hey, z-shifting has its limits, without z-shifting, our best single model was around 0.625 with the chemist's kernel's features, with z-shifting, it gave us up to 0.035 with the same features for a single model, that's the limit of the z-shifting I can see, since we still missed tons of tracks outside the (x,y)=(0,0) . And you're right, a small increase in score in dbscan does increase runtime in lots of cases. And you're very observant with the current leader's past experience ;)  I will move away from the unsupervised learning approach as well. However, everything you've done for dbscan (clustering, post processing, outlier removal, track fitting) will be useful for your next step. Learning/reading/researching is the most important and fun part of kaggling isn't it? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350413,
          "author_name": "sjb1988",
          "author_url": "",
          "post_date": "06/29/2018 18:07:34",
          "content": "<p>@Nicole you're right, learning stuff / improving is the best part - it's why I'm doing this. It seems like there's a lot to learn on this one, especially with <a href=\"/outrunner\">@outrunner</a>'s 0.9 score not too far away now! Everyday life has an annoying habit of getting in the way of time spent on Kaggle! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 350416,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/29/2018 18:12:49",
          "content": "<p>@Seb, exactly, <a href=\"/outrunner\">@outrunner</a> proved how far we can still get, we may not be able to replicate his results but we can learn his approach after the end of this competition.  I've been struggling implementing a new model, no luck so far.  (Just like the German National Football Team, no luck, ouch ouch ouch!! )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 344558,
      "author_name": "nicolefinnie",
      "author_url": "",
      "post_date": "06/18/2018 08:13:18",
      "content": "<p>Our strongest single model of the latest submission gives us around <code>0.58</code> using <code>sin and cos</code> + basic features in the chemist's kernel.  And then we ensembled other 4-5 weaker models to give us over  <code>0.632</code>. However, if I made each model stronger, says the strongest model with <code>0.59</code>, then the local score goes down to <code>0.625</code> when we merge them together. Ensembling is a very hard thing to do in this competition. Using more features doesn't necessarily help if you look at the tracks dbscan finds, some features are contradicting to each other. However, as @CPMP pointed out, the chemist's kernel is already \"ensembling\" tracks for each scan in some sense.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 346012,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "06/20/2018 22:55:50",
          "content": "<p>I remember you saying you haven't used z shifted models yet, which would be my first guess at what you meant by a weak model... </p>\n\n<p>Looking back at this comment it seems like it might have been a much more valuable hint than I previously realized. Thanks for sharing. </p>\n\n<p>Also, I like the term 'the chemist'.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346015,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/20/2018 23:14:10",
          "content": "<p><a href=\"/macfarll\">@macfarll</a>,  Yes, we haven't used z shifted models yet, work keeps us busy. For the weak models I mean the models that yield a lower mean score. They may find different tracks and you can ensemble them, the downside is that a weak model can be very noisy and tends to \"overwrite\" the good tracks from a strong model and that makes ensembling pretty challenging. We're moving away from the merging approach tho, since @CPMP's / @yuval r's LB scores proved a single model is much stronger than a merged one.</p>\n\n<p>P.S. Yes, I like @Grzegorz's title \"the chemist\" since I don't know how to pronounce his name in Polish. :) I guess it's probably similar to \"Gregor\" in German. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 344812,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/18/2018 17:59:28",
      "content": "<p>I use 3 features if cos/sin counts for 1 in the DBSCAN part of my code to get 0.66.  But my code does a bit more than just running DBSCAN in a loop ;)</p>\n\n<p>I am not ensembling models outside the DBSCAN loop.  This means there is hope when working on a single model ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 344880,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "06/18/2018 20:55:57",
          "content": "<p>Wow, 3 features, 1 model, 0.66 for you, 2 features, 1 model, 0.62 for yuval...</p>\n\n<p>I didn't expect that!  Makes this challenge fun!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344886,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/18/2018 21:07:47",
          "content": "<p>Indeed, ensembling has its limit, merging two models can only give you a benefit of 0.01 (naive merging, what we're using) ~0.04 (if all the hits belong to the right track, very unlikely to achieve, we haven't achieved this yet), so I have to revisit the single strong model approach again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344894,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "06/18/2018 21:26:20",
          "content": "<p>On the other hand, an improvement of 0.01 is nice to find!  Of course, for me I think they cost about 0.00025 improvement per hour of lost sleep!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "344428": "Hi, everyone!\n\nHow are you guys dealing with dimensionality? From your experience, increasing the number of features used are good or bad to correct labeling? The best score I've got until now had 5 features being used, while my teammate got a similar score using 7-9 features that were not necessarily the same ones I had used. What are your opinions (and experiences, if you want to talk about it)?",
    "344431": "I use 6 features. I believe obtaining the same score with fewer features is advantageous because it will run faster, right?",
    "344434": "My best model, which gets a score of 0.575 uses 15 features!  Most of the other models I am ensembling have 5 features.",
    "344438": "Referring to DBSCAN:\nIf you consider the sin&amp;cos of an angle a single feature, then I'm using 9 features. I've found in total ~15 that will occasionally give some lift, but many of them would only improve score very slightly on average and would actually reduce score on certain local evaluations. \n\nI'm not currently looking into more features, I'm following a hunch that there will be big improvements outside of DBSCAN features, although I might circle back to features in the future.",
    "344519": "I use only 2 features to get 0.62 (I count sin&amp;cos as one feature)",
    "344522": "Wow - very interesting!  How many models are you ensembling?",
    "344546": "None, yet.",
    "344555": "Ensembling may mean different things here.  The best public kernel iterates over a number of angle corrections.  It ensembles tracks found for each angle prediction.  In that way most of us are using some ensembling.  I guess you are discussing another form of ensembling, which would be to, say, create a new submission out of two or more submissions.",
    "344558": "Our strongest single model of the latest submission gives us around `0.58` using `sin and cos` + basic features in the chemist's kernel.  And then we ensembled other 4-5 weaker models to give us over  `0.632`. However, if I made each model stronger, says the strongest model with `0.59`, then the local score goes down to `0.625` when we merge them together. Ensembling is a very hard thing to do in this competition. Using more features doesn't necessarily help if you look at the tracks dbscan finds, some features are contradicting to each other. However, as @CPMP pointed out, the chemist's kernel is already \"ensembling\" tracks for each scan in some sense.",
    "344575": "You are correct. I don't call these runtime iterations ensembles (I do &gt;10K iterations).",
    "344812": "I use 3 features if cos/sin counts for 1 in the DBSCAN part of my code to get 0.66.  But my code does a bit more than just running DBSCAN in a loop ;)\n\nI am not ensembling models outside the DBSCAN loop.  This means there is hope when working on a single model ;)",
    "344880": "Wow, 3 features, 1 model, 0.66 for you, 2 features, 1 model, 0.62 for yuval...\n\nI didn't expect that!  Makes this challenge fun!",
    "344886": "Indeed, ensembling has its limit, merging two models can only give you a benefit of 0.01 (naive merging, what we're using) ~0.04 (if all the hits belong to the right track, very unlikely to achieve, we haven't achieved this yet), so I have to revisit the single strong model approach again.",
    "344894": "On the other hand, an improvement of 0.01 is nice to find!  Of course, for me I think they cost about 0.00025 improvement per hour of lost sleep!",
    "344917": "yuval r, 10k iterations, wow! We've run 240-480 iterations for each model, however, each model is weaker than 0.6. I guess from what you and @CPMP posted, this is something we can still work on.",
    "344934": "Nicole Finnie, yes definitely a ways to go from the public model! My current idea of improving the tracks with more iterations is by shifting the origin, mostly in the z-direction since the luminous region has a longitudinal width of 55mm and I've noticed ground truth tracks that I thought this would capture. However, I think I'm either not doing something right or my idea is flawed because it is pretty detrimental to the score. Thoughts?",
    "344943": "Matthew, I'm not sure if shifting the origin along the z-axis would help since our features are cylindrical and those are constant traits approximately shared among the hits that belong to the same track, e.g. constant radius space. Shifting hits along the Z-axis probably wouldn't make dbscan(or the clustering approach you use) to find more tracks since their relation to the origin wouldn't change, such as `r, sin(a) cos(a)` would stay the same. @CPMP you're a mathematician/physicist, any thoughts?",
    "344948": "How I imagined it working was by shifting the hits in their xyz coordinates and then recomputing the features to use for each shift. I agree it wouldn't change features such as sin,cos,and r but my other features that use z and d are changed from the shift. I'm really not sure this is even a good idea so I'm glad to hear feedback. :)",
    "344950": "Nicole, not sure I agree with you.  Many features shared publicly assume tracks start with z = 0, be it in the helix unrolling kernels, or Heng's track extension.  Finding tracks that do not start at (0,0,0) is certainly key to win this.  I'm not very good at this for now either.",
    "344951": "CPMP Ooo, interesting. Thank you for your input!",
    "344955": "CPMP, true, I didn't look at that aspect (starting particles). The starting hits of those tracks may not be detected by the detectors in the first place. The tracks of &lt; 4 hits (57% of them don't start from the origin) don't have any weights. Almost all tracks made up by more than 9 hits can be detected by the innermost detectors. The ones we could tackle here would be the tracks of between 4 and 9 hits.\n\n**EDIT on June 23rd **\n\n    For hits = 9, train event=1000\n    42 out of 551 tracks - hits cannot be detected by the innermost detectors\n    0.07622504537205081% of the tracks -  hits cannot be detected in the innermost detectors\n\nhttps://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing",
    "344959": "&gt; All tracks made up by more than 8 hits start from the origin\n\nI didn't realize this, thanks for sharing.",
    "344977": "The unrolling strategy considers a new track valid if it has more hits in the track and the new track has less than 20 hits... \n\nSimilarly an unrolling strategy with an alternative z origin could only assign a new track if the total hits in the new track are 8 or less. This would reduce the amount of false positives...",
    "344984": "Nicole - Is your data about track length vs location of origin based on input from the organizers or on analysis of the \"truth\"???  Thanks",
    "344990": "macfarll, Right. I am just experimenting with this now. It still seems to hurt the model though. I feel like it is because the interval I use for shifting z is too large. However, making this interval smaller means a combinatoric explosion since I run around 240 dbscan iterations per z-shift.\n\n@John @Nicole, I am also curious about this. Is this statistic true for all events?",
    "345002": "Hmm...well then the question becomes what is going to help your solution more, more iterations per origin, or more origins? I'd imagine both will give diminishing returns so you'll just need to find the right sweetspot.\n\n240 iterations seems like a ton to me...",
    "345045": "Nicole, It is probably strange to many others of the score above 0.6 (at least to me), that you have obtained 0.63 without shifting z. Ensembling 7 shifted models can give boost 0.08.",
    "345050": "John, EDA of the ground truth, I only looked at one event, but I assume every event is representative.",
    "345052": "Grzegorz, sounds like the right way to go then. I'll give it a try. `0.63` is nothing when there's `0.8` :)",
    "345057": "0.8 it is almost all tracks of the origin (0,0,z) making less than 0.5 turn inside the detector. It will be a dream and the limit for many of us. The jump from 0.8 to 0.9 would be a quality jump, not simply quantity one. \n\nEDIT: https://www.kaggle.com/sionek/score-limit-for-0-0-z",
    "345059": "I can do a large number of iterations because I don't use  DBSCAN  but a very simple and extremely fast clustering (using bins). \nUsing different assumptions for the origin in the Z axis is very important (X,Y are a waste of time). \nNow we have a new challenge - the  0.8 ;)",
    "345101": "Grzegorz Sionkowski\n\nthanks for the hint. I confirm changing from z =z +dz improves my results on my validation set. Each dz gives potentially a boost of 0.01 ... but the run time exploded as well.",
    "345142": "CPMP, just checked out my EDA result and found out almost all tracks made up of more than `hits = 9` are from the origin. I corrected my comment. However, they're still in the helix-like form. \n\nI added visualization here\nhttps://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing",
    "345993": "Nicole - how did you determine which tracks were from the origin? \nI took a look in the particles file at the vx, vy, vz fields which are the 'initial position or vertex (in millimeters) in global coordinates' - if this means the origin of the particles then there are many away from the origin. I haven't analysed the length of these tracks but I took a quick look at the effect on the score and it seemed to be significant",
    "346003": "Seb, we could have been wrong with our approach, but we were looking at the geometry information provided with the hits to give us hints as to which hits are considered close to the origin. Our approach is very rough, we likely missed many. No hits start at exactly to 0,0,0 - but we were not looking at vx,vy,vz directly.",
    "346012": "I remember you saying you haven't used z shifted models yet, which would be my first guess at what you meant by a weak model... \n\nLooking back at this comment it seems like it might have been a much more valuable hint than I previously realized. Thanks for sharing. \n\nAlso, I like the term 'the chemist'.",
    "346015": "macfarll,  Yes, we haven't used z shifted models yet, work keeps us busy. For the weak models I mean the models that yield a lower mean score. They may find different tracks and you can ensemble them, the downside is that a weak model can be very noisy and tends to \"overwrite\" the good tracks from a strong model and that makes ensembling pretty challenging. We're moving away from the merging approach tho, since @CPMP's / @yuval r's LB scores proved a single model is much stronger than a merged one.\n\nP.S. Yes, I like @Grzegorz's title \"the chemist\" since I don't know how to pronounce his name in Polish. :) I guess it's probably similar to \"Gregor\" in German.",
    "347450": "Seb just found out a couple days ago our original EDA was to find out \"the hits that don't pass the inner most detectors with the volume ID 7, 8, 9\" sorry for the misleading information. I updated my comment.\n\nUsing detector information can be helpful for classifying the hits (before dbscan) and removing outliers, since the current dbscan is only good for predicting helix-like tracks, I remember the @chemist or @CPMP said we missed approximately 18% of the tracks only using dbscan. \n\n our approximate approach to filter on hits from the origin is  \n\n           if np.absolute(particle_list[i].vx) &lt; 0.01 and np.absolute(particle_list[i].vy) &lt; 0.01 and np.absolute(particle_list[i].vz) &lt; 0.01:\n\nFor example, for the train event 0001054 with all tracks of the length of 12 hits:\n\n  \n     251 tracks out of 1338  come from the origin\n     Hits counts: 12",
    "347491": "Finnies Thanks for your responses. I haven't looked using the detector information much yet, it's an interesting idea.\nMy concern with dbscan at the moment is the dependency on the origin - sure you can shift it and get better scores, but with small increases in score for large increases in runtime, I feel like I should try another way... I just haven't figured out what the other way should be yet!\nGiven that the current leader seems to have a good track record in image based competitions, I'd speculate that the same sort of techniques might be used in this competition. Unfortunately for me image recognition is not something I've had that much experience of so I've been doing more reading/research/thinking this week than coding",
    "347540": "Seb  hey, z-shifting has its limits, without z-shifting, our best single model was around 0.625 with the chemist's kernel's features, with z-shifting, it gave us up to 0.035 with the same features for a single model, that's the limit of the z-shifting I can see, since we still missed tons of tracks outside the (x,y)=(0,0) . And you're right, a small increase in score in dbscan does increase runtime in lots of cases. And you're very observant with the current leader's past experience ;)  I will move away from the unsupervised learning approach as well. However, everything you've done for dbscan (clustering, post processing, outlier removal, track fitting) will be useful for your next step. Learning/reading/researching is the most important and fun part of kaggling isn't it? :)",
    "350413": "Nicole you're right, learning stuff / improving is the best part - it's why I'm doing this. It seems like there's a lot to learn on this one, especially with @outrunner's 0.9 score not too far away now! Everyday life has an annoying habit of getting in the way of time spent on Kaggle! :)",
    "350416": "Seb, exactly, @outrunner proved how far we can still get, we may not be able to replicate his results but we can learn his approach after the end of this competition.  I've been struggling implementing a new model, no luck so far.  (Just like the German National Football Team, no luck, ouch ouch ouch!! )"
  },
  "source": "meta"
}