{
  "id": 62163,
  "title": "Shakeup?",
  "url": "/competitions/trackml-particle-identification/discussion/62163",
  "author_name": "CPMP",
  "post_date": "2018-07-28T19:08:42.615000",
  "votes": 8,
  "comment_count": 115,
  "views": 0,
  "content": "<p>How many people will suddenly appear in the last week with very good LB scores?</p>\n\n<p>Will a score above 0.8 (which is my goal) be sufficient to get a gold medal?</p>\n\n<p>These questions are really open as one can rely on local score to assess progress.  it remind me of the <a href=\"https://www.kaggle.com/c/web-traffic-time-series-forecasting\">wikipedia forecasting competition</a> where a number of people (including #1) submitted good solutions without submitting much during the competition.  I was submitting my best there,  and here too.</p>\n\n<p>I really appreciate that demelian is open to show us his current results.  I wish other good contenders were doing the same.  </p>",
  "messages": [
    {
      "id": 363371,
      "postDate": "2018-07-28T19:08:42.617Z",
      "content": "<p>How many people will suddenly appear in the last week with very good LB scores?</p>\n\n<p>Will a score above 0.8 (which is my goal) be sufficient to get a gold medal?</p>\n\n<p>These questions are really open as one can rely on local score to assess progress.  it remind me of the <a href=\"https://www.kaggle.com/c/web-traffic-time-series-forecasting\">wikipedia forecasting competition</a> where a number of people (including #1) submitted good solutions without submitting much during the competition.  I was submitting my best there,  and here too.</p>\n\n<p>I really appreciate that demelian is open to show us his current results.  I wish other good contenders were doing the same.  </p>",
      "rawMarkdown": "How many people will suddenly appear in the last week with very good LB scores?\n\nWill a score above 0.8 (which is my goal) be sufficient to get a gold medal?\n\nThese questions are really open as one can rely on local score to assess progress.  it remind me of the [wikipedia forecasting competition][1] where a number of people (including #1) submitted good solutions without submitting much during the competition.  I was submitting my best there,  and here too.\n\nI really appreciate that demelian is open to show us his current results.  I wish other good contenders were doing the same.  \n\n\n  [1]: https://www.kaggle.com/c/web-traffic-time-series-forecasting",
      "votes": 8
    },
    {
      "id": 365129,
      "postDate": "2018-08-02T00:09:36.007Z",
      "content": "<p>Now me and @Zidmie are experiencing a feeling of being watched and chased by someone right behind us... :-) @Finnies</p>",
      "rawMarkdown": "Now me and @Zidmie are experiencing a feeling of being watched and chased by someone right behind us... :-) @Finnies",
      "votes": 1
    },
    {
      "id": 364310,
      "postDate": "2018-07-31T09:26:59.250Z",
      "content": "<p>@CPMP, I am still struggling around low score, but I would like to get more higher score of course. So I am not only trying algorithms but also reviewing various kernels and discussions until now (\"standing on the shoulders of giants\").  However, I have a concern that I must not steal the giants' ideas/solutions. I think that I must add my originality on it.\nWhat are \"open\", ideal competitions for you ?\n* I am not good at English, sorry if this is difficult to read.</p>",
      "rawMarkdown": "@CPMP, I am still struggling around low score, but I would like to get more higher score of course. So I am not only trying algorithms but also reviewing various kernels and discussions until now (\"standing on the shoulders of giants\").  However, I have a concern that I must not steal the giants' ideas/solutions. I think that I must add my originality on it.\nWhat are \"open\", ideal competitions for you ?\n* I am not good at English, sorry if this is difficult to read.",
      "votes": 1,
      "replies": [
        {
          "id": 366143,
          "postDate": "2018-08-04T05:02:52.833Z",
          "content": "<p>Sorry, I only read this now, I was away most of the week.  This competition is unusual for Kaggle as it is not clear it is a supervised machine learning competition.  And it is clearly a research competition.</p>\n\n<p>Have you tried the playground competitions?  If not then I would start there.  If you have, then maybe the home credit default risk competition is a good one.  The Santander one is weird because of  a massive leak which makes it also rather unusual.  I have not looked at other ongoing ones like airbus or tgs, but they seem interesting too.</p>",
          "rawMarkdown": "Sorry, I only read this now, I was away most of the week.  This competition is unusual for Kaggle as it is not clear it is a supervised machine learning competition.  And it is clearly a research competition.\n\nHave you tried the playground competitions?  If not then I would start there.  If you have, then maybe the home credit default risk competition is a good one.  The Santander one is weird because of  a massive leak which makes it also rather unusual.  I have not looked at other ongoing ones like airbus or tgs, but they seem interesting too.\n\n",
          "votes": 1
        },
        {
          "id": 366628,
          "postDate": "2018-08-06T04:30:44.807Z",
          "content": "<p>Thank you for your reply. I have not tried the playground competitions, but I have joined this TrackML just because of my interest in physics/astrophysics (I am also interested in data science, of course). </p>\n\n<p>I found a suitable playground competition <a href=\"https://www.kaggle.com/c/flavours-of-physics-kernels-only\">\"Flavours of Physics\"</a>. I have not read details yet, but I will join it. I will also check competitions that you mentioned.</p>\n\n<p>Thank you very much !</p>",
          "rawMarkdown": "Thank you for your reply. I have not tried the playground competitions, but I have joined this TrackML just because of my interest in physics/astrophysics (I am also interested in data science, of course). \n\nI found a suitable playground competition [\"Flavours of Physics\"][1]. I have not read details yet, but I will join it. I will also check competitions that you mentioned.\n\nThank you very much !\n  [1]: https://www.kaggle.com/c/flavours-of-physics-kernels-only"
        },
        {
          "id": 366635,
          "postDate": "2018-08-06T05:12:26.070Z",
          "content": "<p>@ykit if you don't mind me adding my 2 cents here. I suggest you take a look at the completed playground competitions because they are geared towards teaching new comers to Kaggle and its full of great kernels to learn from. I would suggest the <a href=\"https://www.kaggle.com/c/nyc-taxi-trip-duration\">\"New York City Taxi Trip Duration\"</a> as a good one because I took part in it as a way of giving back yet I learnt a lot about visualization tools in python.</p>\n\n<p>By the way, I noticed you thanked @CPMP but did not up-vote his answer. Since you said you are new I thought I should mention that up-voting is how we show our appreciations on Kaggle :-)</p>",
          "rawMarkdown": "@ykit if you don't mind me adding my 2 cents here. I suggest you take a look at the completed playground competitions because they are geared towards teaching new comers to Kaggle and its full of great kernels to learn from. I would suggest the [\"New York City Taxi Trip Duration\"][1] as a good one because I took part in it as a way of giving back yet I learnt a lot about visualization tools in python.\n\nBy the way, I noticed you thanked @CPMP but did not up-vote his answer. Since you said you are new I thought I should mention that up-voting is how we show our appreciations on Kaggle :-)\n\n\n  [1]: https://www.kaggle.com/c/nyc-taxi-trip-duration",
          "votes": 2
        },
        {
          "id": 366701,
          "postDate": "2018-08-06T08:51:13.417Z",
          "content": "<p>@YaGana thank you for your suggestion. I upvote comments. I am enjoying and learning a lot in this competition :D</p>",
          "rawMarkdown": "@YaGana thank you for your suggestion. I upvote comments. I am enjoying and learning a lot in this competition :D"
        }
      ]
    },
    {
      "id": 363425,
      "postDate": "2018-07-28T23:56:02.727Z",
      "content": "<p>You got me. I decide to submit in the last week. But I think there is not much people will above 0.8 since it is hard for a clustering only approach, and DL only is harder.</p>",
      "rawMarkdown": "You got me. I decide to submit in the last week. But I think there is not much people will above 0.8 since it is hard for a clustering only approach, and DL only is harder.",
      "votes": 1,
      "replies": [
        {
          "id": 363493,
          "postDate": "2018-07-29T08:58:10.447Z",
          "content": "<p>Hi, I'm not rally targeting you here.  indeed, you disclosed your progress on a regular basis, even without submitting.  I guess you're now running your final submission code...</p>",
          "rawMarkdown": "Hi, I'm not rally targeting you here.  indeed, you disclosed your progress on a regular basis, even without submitting.  I guess you're now running your final submission code...",
          "votes": 1
        },
        {
          "id": 363500,
          "postDate": "2018-07-29T09:38:48.857Z",
          "content": "<p>I get you. I found some big event takes 3 days to process, haha.</p>",
          "rawMarkdown": "I get you. I found some big event takes 3 days to process, haha.",
          "votes": 3
        },
        {
          "id": 363501,
          "postDate": "2018-07-29T09:47:45.570Z",
          "content": "<p>I'm sure that time is well spent and that you will blew us all on the LB!</p>",
          "rawMarkdown": "I'm sure that time is well spent and that you will blew us all on the LB!",
          "votes": 1
        },
        {
          "id": 364152,
          "postDate": "2018-07-30T21:43:54.937Z",
          "content": "<p>No wonder why I can't improve my score. It only takes me 1 hour per event...</p>",
          "rawMarkdown": "No wonder why I can't improve my score. It only takes me 1 hour per event..."
        },
        {
          "id": 364163,
          "postDate": "2018-07-30T22:31:08.180Z",
          "content": "<p>We have a quick solution too, I think it takes 30-60 minutes or so per event to get 0.66 with only one z-shifting + track fitting, merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks).  @CPMP said he could get above 0.7 as posted in another discussion thread. But for us, the score above 0.7 does require 2-3 more hours for dbscan clustering, binning is definitely much faster, like what @yuval and @trian and probably also @Kha and @Sergey are doing.</p>",
          "rawMarkdown": "We have a quick solution too, I think it takes 30-60 minutes or so per event to get 0.66 with only one z-shifting + track fitting, merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks).  @CPMP said he could get above 0.7 as posted in another discussion thread. But for us, the score above 0.7 does require 2-3 more hours for dbscan clustering, binning is definitely much faster, like what @yuval and @trian and probably also @Kha and @Sergey are doing.",
          "votes": 4
        },
        {
          "id": 364194,
          "postDate": "2018-07-31T01:06:08.170Z",
          "content": "<blockquote>\n  <p>No wonder why I can't improve my score. It only takes me 1 hour per event...</p>\n</blockquote>\n\n<p>I get to 0.7 in less than one our per event, and yuval too.  Keep trying ;)</p>",
          "rawMarkdown": "&gt; No wonder why I can't improve my score. It only takes me 1 hour per event...\n\nI get to 0.7 in less than one our per event, and yuval too.  Keep trying ;)"
        },
        {
          "id": 364249,
          "postDate": "2018-07-31T05:55:50.193Z",
          "content": "<blockquote>\n  <p>0.66 with only one z-shifting </p>\n</blockquote>\n\n<p>Are you really only using one alternative origin for tracks? </p>\n\n<p>Or do you use one alternative origin in each direction (+/- z-axis)?</p>",
          "rawMarkdown": "&gt; 0.66 with only one z-shifting \n\nAre you really only using one alternative origin for tracks? \n\nOr do you use one alternative origin in each direction (+/- z-axis)?"
        },
        {
          "id": 364269,
          "postDate": "2018-07-31T07:00:35.330Z",
          "content": "<p>@bkKaggle, no just one, 3 or 2 (-3 or -2 should work well too) I remember, I think from 0, 0, +/-3 you get the highest dbscan score. from (0,0,0) it's about 0.02 lower I recall.  The power is merging different tracks using different models of different features + diffe, for example, we used different z over r alternatives shared in other posts (the chemist shared some too), z/r alternatives are not as sensitive to dbscan as to binning so it helps us a lot.   </p>",
          "rawMarkdown": "@bkKaggle, no just one, 3 or 2 (-3 or -2 should work well too) I remember, I think from 0, 0, +/-3 you get the highest dbscan score. from (0,0,0) it's about 0.02 lower I recall.  The power is merging different tracks using different models of different features + diffe, for example, we used different z over r alternatives shared in other posts (the chemist shared some too), z/r alternatives are not as sensitive to dbscan as to binning so it helps us a lot.   "
        },
        {
          "id": 364317,
          "postDate": "2018-07-31T09:58:12.403Z",
          "content": "<p>I  wonder if you all do some special track select mechanism or just use track length.</p>",
          "rawMarkdown": "I  wonder if you all do some special track select mechanism or just use track length.\n",
          "votes": 1
        },
        {
          "id": 364452,
          "postDate": "2018-07-31T14:46:09.193Z",
          "content": "<blockquote>\n  <p>from (0,0,0) it's about 0.02 lower I recall. </p>\n</blockquote>\n\n<p>@Nicole I am getting 0.02 lower score with z = 3 shifted, you observed the opposite and LB is higher? Did I get you right?</p>",
          "rawMarkdown": "&gt; from (0,0,0) it's about 0.02 lower I recall. \n\n@Nicole I am getting 0.02 lower score with z = 3 shifted, you observed the opposite and LB is higher? Did I get you right?\n\n"
        },
        {
          "id": 364465,
          "postDate": "2018-07-31T15:05:27.543Z",
          "content": "<p>Really? It can be event dependent, I'll check when I get home, I don't use the origin 0,0,0 at all, I only use shifted values,   try to run multiple z0 to find more tracks and merge them, that's what we do. </p>",
          "rawMarkdown": "Really? It can be event dependent, I'll check when I get home, I don't use the origin 0,0,0 at all, I only use shifted values,   try to run multiple z0 to find more tracks and merge them, that's what we do. ",
          "votes": 3
        },
        {
          "id": 364488,
          "postDate": "2018-07-31T16:03:30.527Z",
          "content": "<p>Thanks @Nicole, I will try that when I get back home later today.</p>",
          "rawMarkdown": "Thanks @Nicole, I will try that when I get back home later today."
        },
        {
          "id": 364566,
          "postDate": "2018-07-31T19:19:42.690Z",
          "content": "<p><a href=\"/atom1231\">@atom1231</a> We do 'special' merging :-) Longest-track-wins is what is used within the provided DBScan kernels, however we found that does not work too well when merging different models, different z-shifts, etc. We spent a lot of time coming up with better heuristics when merging. In a post a while ago, <a href=\"/outrunner\">@outrunner</a> mentioned his merging makes use of 'track quality', we do something similar as well. One particular area of concern when merging is that you can have one ground truth track that is partially covered by a track from one model (along with possibly some other 'noise' hits not related to that track), and partially covered by a track from a different model (again, along with possibly some other 'noise'). The challenge when merging is to identify they are part of the same track, and to 'merge' these two tracks together, rather than just selecting one or the other as a winner (but how to tell they are actually the same track, rather than two distinct tracks that should not be combined? an ongoing challenge for us....)</p>",
          "rawMarkdown": "@atom1231 We do 'special' merging :-) Longest-track-wins is what is used within the provided DBScan kernels, however we found that does not work too well when merging different models, different z-shifts, etc. We spent a lot of time coming up with better heuristics when merging. In a post a while ago, @outrunner mentioned his merging makes use of 'track quality', we do something similar as well. One particular area of concern when merging is that you can have one ground truth track that is partially covered by a track from one model (along with possibly some other 'noise' hits not related to that track), and partially covered by a track from a different model (again, along with possibly some other 'noise'). The challenge when merging is to identify they are part of the same track, and to 'merge' these two tracks together, rather than just selecting one or the other as a winner (but how to tell they are actually the same track, rather than two distinct tracks that should not be combined? an ongoing challenge for us....)",
          "votes": 7
        },
        {
          "id": 364680,
          "postDate": "2018-08-01T03:19:10.123Z",
          "content": "<p>Thanks for all your share.</p>\n\n<p>I refer @outruuner @yuval merging flow to hierarchy check , reserve best track\n( In a track,I try to distinguish which one is noise/error point but fail so finally drop all candidates)\nbut it's time consuming.</p>\n\n<p>with dbscan  it takes 1.5 hours (with complete flow/z-shifting)  to build a event to local 0.62.\nNow I change my code to binning , and merging time is bottleneck.\n(@yuval  @CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :(   )</p>\n\n<p>That 's why I ask the question.</p>\n\n<p>In brief ,good feature/different feature combination is always the key. <br>\nI did not do more fail case analysis  and the goal is still far away.</p>",
          "rawMarkdown": "Thanks for all your share.\n\nI refer @outruuner @yuval merging flow to hierarchy check , reserve best track\n( In a track,I try to distinguish which one is noise/error point but fail so finally drop all candidates)\nbut it's time consuming.\n\nwith dbscan  it takes 1.5 hours (with complete flow/z-shifting)  to build a event to local 0.62.\nNow I change my code to binning , and merging time is bottleneck.\n(@yuval  @CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :(   )\n\nThat 's why I ask the question.\n\nIn brief ,good feature/different feature combination is always the key.  \nI did not do more fail case analysis  and the goal is still far away.\n\n\n"
        },
        {
          "id": 364736,
          "postDate": "2018-08-01T06:19:56.667Z",
          "content": "<blockquote>\n  <p>@CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :( )</p>\n</blockquote>\n\n<p>Im a using DBSCAN, not binning.  I get to 0.7 with DBSCAN in less than one hour, about 30 min actually.  </p>\n\n<p>I tried binning recently, and while I could get it run, I could not get over 0.72 with it, whereas DBSCAN gives me nearly 0.75 now.  DBSCAN is slower but better, at least for me.</p>",
          "rawMarkdown": "&gt; @CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :( )\n\nIm a using DBSCAN, not binning.  I get to 0.7 with DBSCAN in less than one hour, about 30 min actually.  \n\nI tried binning recently, and while I could get it run, I could not get over 0.72 with it, whereas DBSCAN gives me nearly 0.75 now.  DBSCAN is slower but better, at least for me."
        },
        {
          "id": 364758,
          "postDate": "2018-08-01T07:21:33.430Z",
          "content": "<p>@CPMP you are right that the maximum initial score with binning is lower then with dbsacn (we also get only to about 0.72) but this can be easily  compensated by simple tack extending =&gt; at the end results are the same.\n(a 0.63 binning is extended to 0.72 in less then 3 min, and a 0.72 is extended to 0.78 also in 3 min)</p>\n\n<p>@ atom1231\nmy 3 min binning also include the merging which is done very simply - select the longest track.</p>\n\n<p>This is a code for binning and merging with previous tracks.</p>\n\n<pre><code>        hits['cat']=(K1*F1).astype('int64')+(K2*F2).astype('int64')*10000            \n        un,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)\n        hits['new_track_id']=inv+10000000    #you need to offset new ids\n        hits['new_track_size']=count[inv]\n        better = (hits.new_track_size&gt;hits.track_size)\n        hits['track_id']=hits['new_track_id'].where(better,hits.track_id)\n</code></pre>\n\n<p>This code if for 2 features F1, F2.  K1,K2 define the bin size</p>\n\n<p>I don't do any sophisticated merging. \nJust some post clustering  track extension  </p>",
          "rawMarkdown": "@CPMP you are right that the maximum initial score with binning is lower then with dbsacn (we also get only to about 0.72) but this can be easily  compensated by simple tack extending =&gt; at the end results are the same.\n(a 0.63 binning is extended to 0.72 in less then 3 min, and a 0.72 is extended to 0.78 also in 3 min)\n\n@ atom1231\nmy 3 min binning also include the merging which is done very simply - select the longest track.\n\nThis is a code for binning and merging with previous tracks.\n \n            hits['cat']=(K1*F1).astype('int64')+(K2*F2).astype('int64')*10000            \n            un,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)\n            hits['new_track_id']=inv+10000000    #you need to offset new ids\n            hits['new_track_size']=count[inv]\n            better = (hits.new_track_size&gt;hits.track_size)\n            hits['track_id']=hits['new_track_id'].where(better,hits.track_id)\n\nThis code if for 2 features F1, F2.  K1,K2 define the bin size\n\nI don't do any sophisticated merging. \nJust some post clustering  track extension  \n\n ",
          "votes": 5
        },
        {
          "id": 364769,
          "postDate": "2018-08-01T07:49:18.270Z",
          "content": "<p>@yuval\nThanks for your share , I did similar in my code.\nI will keep on finding good features~</p>",
          "rawMarkdown": "@yuval\nThanks for your share , I did similar in my code.\nI will keep on finding good features~"
        },
        {
          "id": 364883,
          "postDate": "2018-08-01T12:57:20.200Z",
          "content": "<p>@YaGana, you're right, I didn't use (0,0,0) so I didn't know the score of the origin was higher, I just tried it, this is my raw score for two different z0s.  Cool, maybe I should add this model to our jumbo model too.  :D  </p>\n\n<pre><code>(0,0,-3)\nfor event 1000: 0.50525238\n\n(0,0,0)\nfor event 1000: 0.51460913\n</code></pre>",
          "rawMarkdown": "@YaGana, you're right, I didn't use (0,0,0) so I didn't know the score of the origin was higher, I just tried it, this is my raw score for two different z0s.  Cool, maybe I should add this model to our jumbo model too.  :D  \n\n\n    (0,0,-3)\n    for event 1000: 0.50525238\n    \n    (0,0,0)\n    for event 1000: 0.51460913\n\n\n",
          "votes": 2
        },
        {
          "id": 365097,
          "postDate": "2018-08-01T22:27:26.927Z",
          "content": "<p>@Nicole, thanks for confirming my observation. When I read your comment about it, I started to think that there is something wrong with my implementation. Because of the time it takes per event, I tested my changes on a small number of training events and observed :</p>\n\n<p>&gt; <strong>(0,0,0) Average of 5 events: 0.5338</strong></p>\n\n<p>&gt; <strong>(0,0,3) Average of 5 events: 0.5327</strong></p>\n\n<p>I am seeing a much better performance on (0,0,0) with the current version of my model, the average of 5 events : 0.57+ . I will know how this model does on the LB tomorrow as it takes almost 2 days to run on my weak HW.</p>",
          "rawMarkdown": "@Nicole, thanks for confirming my observation. When I read your comment about it, I started to think that there is something wrong with my implementation. Because of the time it takes per event, I tested my changes on a small number of training events and observed :\n\n&gt; **(0,0,0) Average of 5 events: 0.5338**\n\n&gt; **(0,0,3) Average of 5 events: 0.5327**\n\nI am seeing a much better performance on (0,0,0) with the current version of my model, the average of 5 events : 0.57+ . I will know how this model does on the LB tomorrow as it takes almost 2 days to run on my weak HW."
        },
        {
          "id": 365102,
          "postDate": "2018-08-01T22:47:20.310Z",
          "content": "<p>@YaGana Are those raw scores right from the dbscan? Looking good, don't think anything wrong with your implementation, your scores are higher than our raw scores :) from some z shifts we only got 0.35 or so and merge them together, the scores don't mean much themselves, very different tracks can be found with different weights and iterations and features, and they often yield low scores. </p>",
          "rawMarkdown": "@YaGana Are those raw scores right from the dbscan? Looking good, don't think anything wrong with your implementation, your scores are higher than our raw scores :) from some z shifts we only got 0.35 or so and merge them together, the scores don't mean much themselves, very different tracks can be found with different weights and iterations and features, and they often yield low scores. "
        },
        {
          "id": 365133,
          "postDate": "2018-08-02T00:18:55.423Z",
          "content": "<p>@Nicole thanks.  I am not sure what you mean by raw scores but my dbscan clusterer  implementation include pre-processing, outliers removal etc. then do prediction all in one go. Perhaps not the fastest way but that is what I have at the moment as I do not have that much time to work on this going forward.</p>\n\n<p>I just realized I swapped my results posted above between (0,0,0) and (0,0,3). I have edited and fixed it as (0,0,0) is higher.</p>",
          "rawMarkdown": "@Nicole thanks.  I am not sure what you mean by raw scores but my dbscan clusterer  implementation include pre-processing, outliers removal etc. then do prediction all in one go. Perhaps not the fastest way but that is what I have at the moment as I do not have that much time to work on this going forward.\n\nI just realized I swapped my results posted above between (0,0,0) and (0,0,3). I have edited and fixed it as (0,0,0) is higher."
        },
        {
          "id": 365185,
          "postDate": "2018-08-02T04:35:17.560Z",
          "content": "<p>Why use +/- 3mm as the zshift? I found <a href=\"https://www.kaggle.com/bkkaggle/distribution-of-vz\">here</a> that most of the particles near the origin are distributed in peaks around -2mm, 0mm, and +2mm</p>",
          "rawMarkdown": "Why use +/- 3mm as the zshift? I found [here] [1] that most of the particles near the origin are distributed in peaks around -2mm, 0mm, and +2mm\n\n[1]: https://www.kaggle.com/bkkaggle/distribution-of-vz"
        },
        {
          "id": 365196,
          "postDate": "2018-08-02T05:20:30.727Z",
          "content": "<p>@bkKaggle Yeah +-2.75mm after CERN corrected it in the forum a month ago, however, exploring different shifts gets different tracks though they yield lower scores since there are less tracks starting outside the beam beam collision region, so we merge them together.\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a> (the doc still says 55 mm though)</p>",
          "rawMarkdown": "@bkKaggle Yeah +-2.75mm after CERN corrected it in the forum a month ago, however, exploring different shifts gets different tracks though they yield lower scores since there are less tracks starting outside the beam beam collision region, so we merge them together.\nhttps://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf (the doc still says 55 mm though)",
          "votes": 2
        },
        {
          "id": 365201,
          "postDate": "2018-08-02T05:26:57.973Z",
          "content": "<p>@YaGana, I see, the raw score I mean is the score without preprocessing/postprocessing that comes right out of dbscan (except the longest track winning code between each iteration of dbscan, which is kind of postprocessing too) </p>",
          "rawMarkdown": "@YaGana, I see, the raw score I mean is the score without preprocessing/postprocessing that comes right out of dbscan (except the longest track winning code between each iteration of dbscan, which is kind of postprocessing too) "
        },
        {
          "id": 365527,
          "postDate": "2018-08-02T20:38:07.670Z",
          "content": "<p>I am amazed how much mileage people get out of clustering and binning approaches. After my EDA and after playing a bit with the public kernels, I discarded both ideas. My estimation was that clustering approaches would get stuck at about 0.65 with weight coming mostly from the forward regions. @CPMP's numbers show that this was an underestimation.</p>\n\n<p>I also discarded binning quite early since I concluded that one needed high precision in helix parameter space for separating the tracks and I assumed the required number of bins would make it intractable. It seems I underestimated this approach even more, when I read @yuval's comments.</p>\n\n<p>It will be very interesting to see how much variety of approaches one will find in the winner's  solutions.</p>\n\n<p>P.S. Thinking twice about @yuval's comments, I realize that he found an efficient way to do sparse binning of hits. Back then I thought more into the direction of binning not only hits but associated sub-manifolds in an image space as is done in the Hough transformation. I still think that would be infeasible with the required precision, but who knows what will turn up when the best solutions are opened.</p>",
          "rawMarkdown": "I am amazed how much mileage people get out of clustering and binning approaches. After my EDA and after playing a bit with the public kernels, I discarded both ideas. My estimation was that clustering approaches would get stuck at about 0.65 with weight coming mostly from the forward regions. @CPMP's numbers show that this was an underestimation.\n\nI also discarded binning quite early since I concluded that one needed high precision in helix parameter space for separating the tracks and I assumed the required number of bins would make it intractable. It seems I underestimated this approach even more, when I read @yuval's comments.\n\nIt will be very interesting to see how much variety of approaches one will find in the winner's  solutions.\n\nP.S. Thinking twice about @yuval's comments, I realize that he found an efficient way to do sparse binning of hits. Back then I thought more into the direction of binning not only hits but associated sub-manifolds in an image space as is done in the Hough transformation. I still think that would be infeasible with the required precision, but who knows what will turn up when the best solutions are opened.",
          "votes": 1
        },
        {
          "id": 365744,
          "postDate": "2018-08-03T09:47:38.900Z",
          "content": "<blockquote>\n  <p>I am amazed how much mileage people get out of clustering and binning approaches. </p>\n</blockquote>\n\n<p>I think it can go way higher ;)</p>",
          "rawMarkdown": "&gt; I am amazed how much mileage people get out of clustering and binning approaches. \n\nI think it can go way higher ;)"
        },
        {
          "id": 365794,
          "postDate": "2018-08-03T12:37:47.603Z",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>I think it can go way higher ;)</p>\n</blockquote>\n\n<p>Do you mean just binning/clustering alone or plus track extension?</p>",
          "rawMarkdown": "@CPMP:\n&gt; I think it can go way higher ;)\n\nDo you mean just binning/clustering alone or plus track extension?"
        },
        {
          "id": 365803,
          "postDate": "2018-08-03T12:51:45.743Z",
          "content": "<p>I mean so far I only applied it to centered tracks.  I think I found how to also apply it to out of center tracks.  But it will take days before it shows (if it shows...)</p>",
          "rawMarkdown": "I mean so far I only applied it to centered tracks.  I think I found how to also apply it to out of center tracks.  But it will take days before it shows (if it shows...)\n\n"
        },
        {
          "id": 365816,
          "postDate": "2018-08-03T13:15:43.163Z",
          "content": "<p>I think you are correct. It could be applied to out of center tracks. The issues are:</p>\n\n<ol>\n<li>Where to start?</li>\n<li>How to select good tracks. (This time only length is not enough)</li>\n<li>It is much slower then centered tracks</li>\n</ol>\n\n<p>Not sure we'll have enough time to really implement our solution.</p>",
          "rawMarkdown": "I think you are correct. It could be applied to out of center tracks. The issues are:\n\n1. Where to start?\n2. How to select good tracks. (This time only length is not enough)\n3. It is much slower then centered tracks\n\nNot sure we'll have enough time to really implement our solution.",
          "votes": 1
        },
        {
          "id": 365821,
          "postDate": "2018-08-03T13:18:27.903Z",
          "content": "<blockquote>\n  <p>It is much slower then centered tracks</p>\n</blockquote>\n\n<p>Definitely.</p>",
          "rawMarkdown": "&gt; It is much slower then centered tracks\n\nDefinitely."
        },
        {
          "id": 365825,
          "postDate": "2018-08-03T13:36:44.363Z",
          "content": "<p>@yuval, @CPMP, agree, the very same equations can be applied to the non-centred tracks using exhaustive search, you need multiple physical cores to do so to meet the deadline. :)     I haven't even finished writing the code of finding all tracks from the origin, I'll pass. :) </p>",
          "rawMarkdown": "@yuval, @CPMP, agree, the very same equations can be applied to the non-centred tracks using exhaustive search, you need multiple physical cores to do so to meet the deadline. :)     I haven't even finished writing the code of finding all tracks from the origin, I'll pass. :) "
        },
        {
          "id": 366023,
          "postDate": "2018-08-03T20:35:30.063Z",
          "content": "<p>Nicole, compute is nearly free...\nYou can lease a 48 core server for $15/day from Google.</p>\n\n<p>Not that it did me much good...  0.66 seems to be the best I can do on my own.</p>\n\n<p>Now I need to decode the clues recently posted to \"steal\" a score &gt; 0.7 😀</p>",
          "rawMarkdown": "Nicole, compute is nearly free...\nYou can lease a 48 core server for $15/day from Google.\n\nNot that it did me much good...  0.66 seems to be the best I can do on my own.\n\nNow I need to decode the clues recently posted to \"steal\" a score &gt; 0.7 😀\n\n",
          "votes": 1
        },
        {
          "id": 366029,
          "postDate": "2018-08-03T20:56:10.660Z",
          "content": "<p>@John, welcome back!!! Thanks for the info :D  we do have access to a CPU / GPU farm for Kaggle and for free!  lol   Just I don't think we'll have time to implement it anymore since we're still working on a new solution, so I'll pass. :p  Hang in there, only 10 days are left, then we can take a break from Kaggle, YAH! </p>",
          "rawMarkdown": "@John, welcome back!!! Thanks for the info :D  we do have access to a CPU / GPU farm for Kaggle and for free!  lol   Just I don't think we'll have time to implement it anymore since we're still working on a new solution, so I'll pass. :p  Hang in there, only 10 days are left, then we can take a break from Kaggle, YAH! ",
          "votes": 1
        },
        {
          "id": 366036,
          "postDate": "2018-08-03T21:18:07.433Z",
          "content": "<p>@yuval r, I tried to test a partial implementation and found it to be super slow and gave up on it. I had to focus on my relatively faster modifications of improving my results on the centered tracks. There is just not enough time to implement a laborious solution at this stage. Oh time :-)</p>\n\n<p>@Nicole, did you end up adding the (0,0,0) model to your ensemble? Seen any improvement?</p>",
          "rawMarkdown": "@yuval r, I tried to test a partial implementation and found it to be super slow and gave up on it. I had to focus on my relatively faster modifications of improving my results on the centered tracks. There is just not enough time to implement a laborious solution at this stage. Oh time :-)\n\n@Nicole, did you end up adding the (0,0,0) model to your ensemble? Seen any improvement?"
        },
        {
          "id": 366057,
          "postDate": "2018-08-03T22:11:54.370Z",
          "content": "<p>@YaGana  There is no reason to try and find off centered tracks before you get to at least 0.77. We moved only now to this area. Our current score use only centered model (with a small Z shift of course)</p>",
          "rawMarkdown": "@YaGana  There is no reason to try and find off centered tracks before you get to at least 0.77. We moved only now to this area. Our current score use only centered model (with a small Z shift of course)",
          "votes": 2
        },
        {
          "id": 366058,
          "postDate": "2018-08-03T22:13:47.707Z",
          "content": "<p>@YaGana, I haven't added the (0,0,0) model to my existing jumbo model yet, but I think I'll add it to the new model if it works well. </p>",
          "rawMarkdown": "@YaGana, I haven't added the (0,0,0) model to my existing jumbo model yet, but I think I'll add it to the new model if it works well. "
        },
        {
          "id": 366060,
          "postDate": "2018-08-03T22:26:28.887Z",
          "content": "<p>Thanks for the advise @yuval r and for all the tips you have shared in this contest.</p>",
          "rawMarkdown": "Thanks for the advise @yuval r and for all the tips you have shared in this contest."
        },
        {
          "id": 366147,
          "postDate": "2018-08-04T05:14:42.853Z",
          "content": "<p>I guess off center tracks are basically tracks with a large zshift.</p>\n\n<p>It seems that my problem is not finding more tracks, but finding higher quality tracks and merging them together to generate longer tracks. The average length of my tracks is 2-5 while the average length of the ground truth particles is around 12.</p>\n\n<p>Somewhere else on the discussion forum, I saw that the organizers said that the emphasis is more on finding high quality track candidates than on perfecting the length of the tracks. If this is true, does this mean that CERN has track merging approaches that are a lot more sophisticated than the length and quality based approaches that are being used here?</p>",
          "rawMarkdown": "I guess off center tracks are basically tracks with a large zshift.\n\nIt seems that my problem is not finding more tracks, but finding higher quality tracks and merging them together to generate longer tracks. The average length of my tracks is 2-5 while the average length of the ground truth particles is around 12.\n\nSomewhere else on the discussion forum, I saw that the organizers said that the emphasis is more on finding high quality track candidates than on perfecting the length of the tracks. If this is true, does this mean that CERN has track merging approaches that are a lot more sophisticated than the length and quality based approaches that are being used here?",
          "votes": 1
        },
        {
          "id": 366148,
          "postDate": "2018-08-04T05:25:28.180Z",
          "content": "<blockquote>\n  <p>I guess off center tracks are basically tracks with a large zshift.</p>\n</blockquote>\n\n<p>No, they are tracks that do not get near z axis.  </p>",
          "rawMarkdown": "&gt; I guess off center tracks are basically tracks with a large zshift.\n\nNo, they are tracks that do not get near z axis.  "
        },
        {
          "id": 366184,
          "postDate": "2018-08-04T08:21:04.157Z",
          "content": "<p>@bkKaggle -   my 2 cents, I think those are the tracks that do not originate from (x,y) = (0,0) (the initial position of the particles)  I shared this image a couple months ago. I think the threshold for this EDA I set at the time was x&gt;1 or y&gt;1. You can see some clearly don't start from (x,y) = (0,0)  or get close to the z-axis by CPMP's definition. </p>\n\n<p><a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">tracks of 9 hits that don't start from near the (0,0,z) </a></p>",
          "rawMarkdown": "@bkKaggle -   my 2 cents, I think those are the tracks that do not originate from (x,y) = (0,0) (the initial position of the particles)  I shared this image a couple months ago. I think the threshold for this EDA I set at the time was x&gt;1 or y&gt;1. You can see some clearly don't start from (x,y) = (0,0)  or get close to the z-axis by CPMP's definition. \n\n[tracks of 9 hits that don't start from near the (0,0,z) ][1]\n\n\n\n  [1]: https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing"
        },
        {
          "id": 366204,
          "postDate": "2018-08-04T10:37:02.663Z",
          "content": "<p>A trajectory may start far from z axis, yet be on an helix that goes through the z axis.  Such tracks can be caught by code that look for helix going through the z axis.  An out of center helix is an helix that never is close to the z axis.</p>",
          "rawMarkdown": "A trajectory may start far from z axis, yet be on an helix that goes through the z axis.  Such tracks can be caught by code that look for helix going through the z axis.  An out of center helix is an helix that never is close to the z axis."
        },
        {
          "id": 366542,
          "postDate": "2018-08-05T18:27:12.273Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 363924,
      "postDate": "2018-07-30T10:44:51.217Z",
      "content": "<p>Good to see icecuber move in the open.  How many others to come?</p>",
      "rawMarkdown": "Good to see icecuber move in the open.  How many others to come?",
      "votes": 2,
      "replies": [
        {
          "id": 363929,
          "postDate": "2018-07-30T10:52:20.330Z",
          "content": "<p>I thought the same too :)  Mickey, Heng, the chemist, and others who haven't made submissions yet :D  </p>",
          "rawMarkdown": "I thought the same too :)  Mickey, Heng, the chemist, and others who haven't made submissions yet :D  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 385948,
      "postDate": "2018-09-11T20:57:58.517Z",
      "content": "<p>The second \"Throughput\" phase of this competition is online, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/65525\">this post</a></p>",
      "rawMarkdown": "The second \"Throughput\" phase of this competition is online, see [this post][1]\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/65525"
    },
    {
      "id": 366794,
      "postDate": "2018-08-06T14:25:42.993Z",
      "content": "<p>Oh, the new participant (bestfitting) has come with the good score. </p>",
      "rawMarkdown": "Oh, the new participant (bestfitting) has come with the good score. \n",
      "replies": [
        {
          "id": 366823,
          "postDate": "2018-08-06T15:29:39.450Z",
          "content": "<p>I think we should expect a lot of that this week...</p>",
          "rawMarkdown": "I think we should expect a lot of that this week..."
        },
        {
          "id": 366830,
          "postDate": "2018-08-06T15:45:29.503Z",
          "content": "<p>Indeed, I guess at the end of the competition, you may need 0.8 to get a gold medal. </p>",
          "rawMarkdown": "Indeed, I guess at the end of the competition, you may need 0.8 to get a gold medal. "
        },
        {
          "id": 366878,
          "postDate": "2018-08-06T17:11:03.170Z",
          "content": "<p>I hope I can get a silver with 0.66, but that is literally what is keeping me up at night 😁</p>",
          "rawMarkdown": "I hope I can get a silver with 0.66, but that is literally what is keeping me up at night 😁"
        },
        {
          "id": 366891,
          "postDate": "2018-08-06T17:34:09.920Z",
          "content": "<p>@John, that's certain ;)  today is the final day of making the first submission if I understood the rules correctly. With your current score, you'll get a high silver medal for sure if you do nothing until the end of the competition. ;) </p>",
          "rawMarkdown": "@John, that's certain ;)  today is the final day of making the first submission if I understood the rules correctly. With your current score, you'll get a high silver medal for sure if you do nothing until the end of the competition. ;) ",
          "votes": 1
        },
        {
          "id": 366896,
          "postDate": "2018-08-06T17:45:25.630Z",
          "content": "<p>Nicole,  Are you sure? I thought we had to accept the rules in order to download the data? </p>",
          "rawMarkdown": "Nicole,  Are you sure? I thought we had to accept the rules in order to download the data? ",
          "votes": 1
        },
        {
          "id": 366902,
          "postDate": "2018-08-06T17:57:56.383Z",
          "content": "<p>No I'm not sure , under the time line , the first rule indicates today is only accepting the rules to compete, the second rule says , August 6, 2018 - Team Merger deadline. This is the last day participants may join or merge teams. So its not too clear what this means, maybe it only applies to team merge, i thought you have to make the first sub to be able to merge with another team. A related discussion posted by <a href=\"/inversion\">@inversion</a> 2 years ago\n<a href=\"https://www.kaggle.com/product-feedback/20451\">https://www.kaggle.com/product-feedback/20451</a></p>",
          "rawMarkdown": "No I'm not sure , under the time line , the first rule indicates today is only accepting the rules to compete, the second rule says , August 6, 2018 - Team Merger deadline. This is the last day participants may join or merge teams. So its not too clear what this means, maybe it only applies to team merge, i thought you have to make the first sub to be able to merge with another team. A related discussion posted by @inversion 2 years ago\nhttps://www.kaggle.com/product-feedback/20451",
          "votes": 1
        },
        {
          "id": 366904,
          "postDate": "2018-08-06T18:05:16.053Z",
          "content": "<p>Interesting, totally not clear, but that could be exactly what it means. Now the question do I make a throwaway submission today in order to keep myself in the competition, when my otherwise promising solution seems prone to every bug under the sun? Nuts.</p>\n\n<p>thanks for the info. </p>",
          "rawMarkdown": "Interesting, totally not clear, but that could be exactly what it means. Now the question do I make a throwaway submission today in order to keep myself in the competition, when my otherwise promising solution seems prone to every bug under the sun? Nuts.\n\nthanks for the info. "
        },
        {
          "id": 366907,
          "postDate": "2018-08-06T18:09:25.290Z",
          "content": "<p>@Matthew, I'd definitely do so to be safe. There was a turmoil after the data science bowl this year because the rules between Kaggle and the competition were conflicting each other. </p>",
          "rawMarkdown": "@Matthew, I'd definitely do so to be safe. There was a turmoil after the data science bowl this year because the rules between Kaggle and the competition were conflicting each other. ",
          "votes": 1
        },
        {
          "id": 366908,
          "postDate": "2018-08-06T18:18:20.820Z",
          "content": "<p>@Matthew Wander: I just read the rules again and honestly it is not clear to me what exactly constitutes \"entry\" in the competition. It is clear that \"entry\" implies the acceptance of the competition rules, but it is not clear to me whether there is an implication in the other direction.</p>\n\n<p>To be on the safe side, I'd definitely make a submission before the upcoming merger &amp; entry deadline.</p>\n\n<p>P.S.: If you're very short on time, remember that your submission only needs to be technically valid, it does not need to achieve a high score, so you could immediately submit a submission created by one of the public kernels, if you need one, or even a synthetic trivial one, I guess. (Although the latter could strictly be said to be a case of \"hand-labeling\", which is forbidden, if one interprets it in a very inflexible way.)</p>",
          "rawMarkdown": "@Matthew Wander: I just read the rules again and honestly it is not clear to me what exactly constitutes \"entry\" in the competition. It is clear that \"entry\" implies the acceptance of the competition rules, but it is not clear to me whether there is an implication in the other direction.\n\nTo be on the safe side, I'd definitely make a submission before the upcoming merger &amp; entry deadline.\n\nP.S.: If you're very short on time, remember that your submission only needs to be technically valid, it does not need to achieve a high score, so you could immediately submit a submission created by one of the public kernels, if you need one, or even a synthetic trivial one, I guess. (Although the latter could strictly be said to be a case of \"hand-labeling\", which is forbidden, if one interprets it in a very inflexible way.)",
          "votes": 1
        },
        {
          "id": 366939,
          "postDate": "2018-08-06T19:18:48.783Z",
          "content": "<p>@Matthew, In the past you had to have made at least one submission prior to the last week of the competition, but then I believe they changed it so that you only had to accept the rules by the last week and could make your first submission after that.  But I haven't tested this and don't want to be blamed if you get in trouble, so it might be best to make a throwaway submission like others have suggested.</p>",
          "rawMarkdown": "@Matthew, In the past you had to have made at least one submission prior to the last week of the competition, but then I believe they changed it so that you only had to accept the rules by the last week and could make your first submission after that.  But I haven't tested this and don't want to be blamed if you get in trouble, so it might be best to make a throwaway submission like others have suggested.",
          "votes": 2
        },
        {
          "id": 366941,
          "postDate": "2018-08-06T19:19:50.283Z",
          "content": "<p>@John Sweeney, that's funny.</p>",
          "rawMarkdown": "@John Sweeney, that's funny."
        },
        {
          "id": 366947,
          "postDate": "2018-08-06T19:28:23.003Z",
          "content": "<p>thanks all, but I think I am going to pull out. there were just too many hurdles to the solution and a submission would take 3-4 days to work out. I think I am going to take it as a sign that the clock just ran out and work on other projects.  </p>",
          "rawMarkdown": "thanks all, but I think I am going to pull out. there were just too many hurdles to the solution and a submission would take 3-4 days to work out. I think I am going to take it as a sign that the clock just ran out and work on other projects.  "
        },
        {
          "id": 366954,
          "postDate": "2018-08-06T19:41:59.777Z",
          "content": "<p>@Matthew, no no, didn't mean to discourage you, you can make a real submission before the deadline and tell us if it works ;)  and I agree with \"too many hurdles to the solution part\", it was a very tough battle for someone like me who sucks at math. </p>",
          "rawMarkdown": "@Matthew, no no, didn't mean to discourage you, you can make a real submission before the deadline and tell us if it works ;)  and I agree with \"too many hurdles to the solution part\", it was a very tough battle for someone like me who sucks at math. "
        },
        {
          "id": 366958,
          "postDate": "2018-08-06T19:47:09.493Z",
          "content": "<blockquote>\n  <p>it was a very tough battle</p>\n</blockquote>\n\n<p>@Nicole And unfortunately I have to say that you are now in 11th place, the last for a gold medal. I hope no one will come, and you will get it.</p>\n\n<p>I think I can relax now and wait for a silver medal. :)</p>",
          "rawMarkdown": "&gt;  it was a very tough battle\n\n@Nicole And unfortunately I have to say that you are now in 11th place, the last for a gold medal. I hope no one will come, and you will get it.\n\nI think I can relax now and wait for a silver medal. :)"
        },
        {
          "id": 366962,
          "postDate": "2018-08-06T19:54:00.187Z",
          "content": "<p>Thanks @Sergey, we are the slowest runners who will be caught by the bear ;) (From the joke @Robert made some days ago) - @Heng shows an amazing result in DL and I expect him to make a submission over 0.8 :)   we expect to get a silver medal as well. We've been working in the wrong direction in the past 2 months completely (building rules and writing lots of code) and justly realized it's all about math equations in the last week after I've read @yuval's post.  Hopefully our code will still be useful after we share our github repository. </p>",
          "rawMarkdown": "Thanks @Sergey, we are the slowest runners who will be caught by the bear ;) (From the joke @Robert made some days ago) - @Heng shows an amazing result in DL and I expect him to make a submission over 0.8 :)   we expect to get a silver medal as well. We've been working in the wrong direction in the past 2 months completely (building rules and writing lots of code) and justly realized it's all about math equations in the last week after I've read @yuval's post.  Hopefully our code will still be useful after we share our github repository. "
        },
        {
          "id": 366963,
          "postDate": "2018-08-06T19:54:04.577Z",
          "content": "<p>@Nicole. No discouragement, just realism. I'm too slow a python programmer to make it work in time. The math was sound: you use pairs of points to determine the features: helix radii, helix xy center, and screw axis. I would probably need to go to triplets of points to get past 85%. The challenge is getting from pairs of points back to singlets.  I had some ideas but the results are just not where they need to be. Too bad I wasted a week trying to get the cells file to yield derivatives. Shrug. Now to go and try my hand at salt deposits.</p>",
          "rawMarkdown": "@Nicole. No discouragement, just realism. I'm too slow a python programmer to make it work in time. The math was sound: you use pairs of points to determine the features: helix radii, helix xy center, and screw axis. I would probably need to go to triplets of points to get past 85%. The challenge is getting from pairs of points back to singlets.  I had some ideas but the results are just not where they need to be. Too bad I wasted a week trying to get the cells file to yield derivatives. Shrug. Now to go and try my hand at salt deposits.",
          "votes": 1
        },
        {
          "id": 366974,
          "postDate": "2018-08-06T20:37:08.670Z",
          "content": "<p>@Matthew Wander: Honestly, I think you are underestimating the difficulty of the problem.</p>\n\n<blockquote>\n  <p>I would probably need to go to triplets of points to get past 85%.</p>\n</blockquote>\n\n<p>If you have good reasons to think that you can get 85% with rather simple code, you must definitely make a submission for the sake of science! (I say that even though you most likely would kick me down the LB :-) Such a solution would be really interesting to the CERN people, I think.</p>\n\n<p>I have no idea what the leaders are doing, but I'd be surprised if their code is quite simple.\nMy own solution is currently at about 3.8K pure code lines. Not all of it is necessary and I will remove some dead code before the conclusion, but still it's a lot.</p>\n\n<p>I somehow expect the winners' solutions to be much more concise and elegant than mine. Let's see.</p>",
          "rawMarkdown": "@Matthew Wander: Honestly, I think you are underestimating the difficulty of the problem.\n\n&gt; I would probably need to go to triplets of points to get past 85%.\n\nIf you have good reasons to think that you can get 85% with rather simple code, you must definitely make a submission for the sake of science! (I say that even though you most likely would kick me down the LB :-) Such a solution would be really interesting to the CERN people, I think.\n\nI have no idea what the leaders are doing, but I'd be surprised if their code is quite simple.\nMy own solution is currently at about 3.8K pure code lines. Not all of it is necessary and I will remove some dead code before the conclusion, but still it's a lot.\n\nI somehow expect the winners' solutions to be much more concise and elegant than mine. Let's see.",
          "votes": 3
        },
        {
          "id": 366982,
          "postDate": "2018-08-06T20:53:22.450Z",
          "content": "<p>@edwin Good to know our team is not the only one writing lots of code - we have about 4K lines of code, mostly trying to deal with all the complications as a result of EDA (i.e. how to extend different types of tracks, how to detect outliers, etc.). Definitely a lot more complicated than we had expected at the beginning of the competition. However, even with lots of code, we are far from the top, and far below your score (guess your 3.8K LOC are more powerful than ours, darn). I think those that understand advanced math have a big edge though, finding the right features can increase the score by a huge amount, writing lots of LOC mostly just nudges the score a little, from our experience.</p>",
          "rawMarkdown": "@edwin Good to know our team is not the only one writing lots of code - we have about 4K lines of code, mostly trying to deal with all the complications as a result of EDA (i.e. how to extend different types of tracks, how to detect outliers, etc.). Definitely a lot more complicated than we had expected at the beginning of the competition. However, even with lots of code, we are far from the top, and far below your score (guess your 3.8K LOC are more powerful than ours, darn). I think those that understand advanced math have a big edge though, finding the right features can increase the score by a huge amount, writing lots of LOC mostly just nudges the score a little, from our experience.",
          "votes": 3
        },
        {
          "id": 367002,
          "postDate": "2018-08-06T21:50:16.963Z",
          "content": "<blockquote>\n  <p>writing lots of LOC mostly just nudges the score a little, from our experience.</p>\n</blockquote>\n\n<p>True. My typical \"great idea\" gives me a +0.0020 increment after an implementation time between 5 minutes and two weeks, with no discernible relation between time invested and score improvement ;-/.</p>\n\n<p>There were some very satisfying +0.0200 ideas, but I seem to have run out of them.</p>\n\n<p>My guess would be that <a href=\"/icecuber\">@icecuber</a> &amp; co have found a good way to train a classifier for evaluation and subsequent merging of track candidates. That is currently one of the biggest weaknesses of my algorithm: the evaluation of track candidates. I tried some ML approaches but with little success. I am also completely new to ML, which doesn't help.</p>",
          "rawMarkdown": "&gt; writing lots of LOC mostly just nudges the score a little, from our experience.\n\nTrue. My typical \"great idea\" gives me a +0.0020 increment after an implementation time between 5 minutes and two weeks, with no discernible relation between time invested and score improvement ;-/.\n\nThere were some very satisfying +0.0200 ideas, but I seem to have run out of them.\n\nMy guess would be that @icecuber &amp; co have found a good way to train a classifier for evaluation and subsequent merging of track candidates. That is currently one of the biggest weaknesses of my algorithm: the evaluation of track candidates. I tried some ML approaches but with little success. I am also completely new to ML, which doesn't help.",
          "votes": 2
        },
        {
          "id": 367006,
          "postDate": "2018-08-06T22:15:59.160Z",
          "content": "<p>@Edwin That is simply based on my estimation of the number of points on the helix that start with X and Y both ~=0. Yes I am much further away probably at least a month to complete the internal transformation accurately enough to achieve even that level of quality. Mathematical limits start hitting hard you only need four points to confirm a helix originating at 0,0 but 7 for a helix centered anywhere else. I don't know the count but I know the very best solutions are close to that limit. </p>",
          "rawMarkdown": "@Edwin That is simply based on my estimation of the number of points on the helix that start with X and Y both ~=0. Yes I am much further away probably at least a month to complete the internal transformation accurately enough to achieve even that level of quality. Mathematical limits start hitting hard you only need four points to confirm a helix originating at 0,0 but 7 for a helix centered anywhere else. I don't know the count but I know the very best solutions are close to that limit. ",
          "votes": 1
        },
        {
          "id": 367011,
          "postDate": "2018-08-06T22:33:01.353Z",
          "content": "<p>@Matthew Wander: You are correct in that the helices which do not pass through the line (0, 0, z) are much harder to find. In general, a helix with axis parallel to the z-axis has 5 degrees of freedom and every 3-d point gives you 2 (because you need to subtract one for the curve parameter). So in general you need three points to pin down the helix (with one degree to spare for checks), but if you know the helix passes through (0,0,z) you only need two points, which makes a huge difference.</p>\n\n<p>HOWEVER, the particle trajectories are NOT helices! The problem really wouldn't be that hard if they were. For many reasons, which are mostly mentioned in the introductory document by the organizers, the trajectories deviate from perfect helices, sometimes quite strongly.</p>\n\n<p>What really makes the problem so hard is that <em>the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks</em>!</p>",
          "rawMarkdown": "@Matthew Wander: You are correct in that the helices which do not pass through the line (0, 0, z) are much harder to find. In general, a helix with axis parallel to the z-axis has 5 degrees of freedom and every 3-d point gives you 2 (because you need to subtract one for the curve parameter). So in general you need three points to pin down the helix (with one degree to spare for checks), but if you know the helix passes through (0,0,z) you only need two points, which makes a huge difference.\n\nHOWEVER, the particle trajectories are NOT helices! The problem really wouldn't be that hard if they were. For many reasons, which are mostly mentioned in the introductory document by the organizers, the trajectories deviate from perfect helices, sometimes quite strongly.\n\nWhat really makes the problem so hard is that *the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks*!",
          "votes": 3
        },
        {
          "id": 367050,
          "postDate": "2018-08-07T01:05:38.653Z",
          "content": "<blockquote>\n  <p>What really makes the problem so hard is that the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks!</p>\n</blockquote>\n\n<p>Yep @Edwin Steiner you have succinctly defined the issue that made accurate clustering a challenge.</p>",
          "rawMarkdown": "&gt; What really makes the problem so hard is that the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks!\n\nYep @Edwin Steiner you have succinctly defined the issue that made accurate clustering a challenge."
        },
        {
          "id": 367055,
          "postDate": "2018-08-07T01:36:15.397Z",
          "content": "<blockquote>\n  <p>think those that understand advanced math have a big edge though</p>\n</blockquote>\n\n<p>How much math do you need to know to understand everything? What type of math is it anyway?</p>",
          "rawMarkdown": "&gt;  think those that understand advanced math have a big edge though\n\nHow much math do you need to know to understand everything? What type of math is it anyway?"
        },
        {
          "id": 367073,
          "postDate": "2018-08-07T02:41:47.913Z",
          "content": "<blockquote>\n  <p>I think those that understand advanced math have a big edge though,</p>\n</blockquote>\n\n<p>Elementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.</p>\n\n<p>I really don't think math is the issue here, all necessary equations are provided in material shared on the forum.  </p>\n\n<p>This is down to usual ML: proper EDA, proper feature engineering, choice of algorithm and model, and some tuning.</p>\n\n<p>The real split between top entries and the rest (I'm part of the rest), is to find effective ways for identifying tracks not originating near z axis.  I'm working on it but I doubt I'll get something significant by end of competition.</p>",
          "rawMarkdown": "&gt; I think those that understand advanced math have a big edge though,\n\nElementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.\n\nI really don't think math is the issue here, all necessary equations are provided in material shared on the forum.  \n\nThis is down to usual ML: proper EDA, proper feature engineering, choice of algorithm and model, and some tuning.\n\nThe real split between top entries and the rest (I'm part of the rest), is to find effective ways for identifying tracks not originating near z axis.  I'm working on it but I doubt I'll get something significant by end of competition.",
          "votes": 1
        },
        {
          "id": 367152,
          "postDate": "2018-08-07T07:45:03.017Z",
          "content": "<p>I am very much looking forward to reading your solution when the competition is finished!</p>\n\n<p>240 LOC ==&gt; 0.778 sounds like a beautiful and elegant solution!</p>\n\n<p>Like the Finnies mentioned, I spent most of the competition trying to hand craft exotic features and merging methods when the solution was in the complete opposite direction!</p>\n\n<p>After the recent board discussions I think I know the math but am obviously missing some property of symmetry that is cutting my score in half!!!</p>\n\n<p>I'm guessing that once the competition is finished, I'll kick myself for missing something rather obvious!</p>",
          "rawMarkdown": "I am very much looking forward to reading your solution when the competition is finished!\n\n240 LOC ==&gt; 0.778 sounds like a beautiful and elegant solution!\n\nLike the Finnies mentioned, I spent most of the competition trying to hand craft exotic features and merging methods when the solution was in the complete opposite direction!\n\nAfter the recent board discussions I think I know the math but am obviously missing some property of symmetry that is cutting my score in half!!!\n\nI'm guessing that once the competition is finished, I'll kick myself for missing something rather obvious!",
          "votes": 3
        },
        {
          "id": 367168,
          "postDate": "2018-08-07T08:18:34.970Z",
          "content": "<p>Thanks, but I am not using any magic feature as we could see in other competitions sometimes. It really is along the line that yuval shared <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/61081#356653\">in this topic</a>, except I'm using DBSCAN instead of binning.  </p>",
          "rawMarkdown": "Thanks, but I am not using any magic feature as we could see in other competitions sometimes. It really is along the line that yuval shared [in this topic][1], except I'm using DBSCAN instead of binning.  \n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/61081#356653"
        },
        {
          "id": 367182,
          "postDate": "2018-08-07T08:57:52.863Z",
          "content": "<p><strong>@John</strong>, Yes, the problem we have is we couldn't find the accurate math equations to solve the problem with the uneven magnetic field, and I didn't even realize this was a problem until 2 weeks ago.  Thanks to <strong>@yuval,</strong> it was very generous of him sharing that important information. That's why we have been working in the trial and error mode without knowing what we're doing. And <strong>@Edwin</strong> also pointed out that the biggest challenge is the deviations of the trajectories from perfect helices, if the equation is not accurate enough, you get neighbour tracks' hits, and that's the problem we have. And once you get the tracks wrong in the first place, you have to spend a crazy amount of time removing the wrong hits and you will only get half of them right. Not knowing math leads to 4k LOC which still yields a low score, and I don't know if the exact equations of solving the uneven/unbounded problem have been shared in the forum, if it is, I must have overlooked some discussions. I think this is something that will set people's score apart from others and people wouldn't easily share, just my 2 cents. </p>",
          "rawMarkdown": "**@John**, Yes, the problem we have is we couldn't find the accurate math equations to solve the problem with the uneven magnetic field, and I didn't even realize this was a problem until 2 weeks ago.  Thanks to **@yuval,** it was very generous of him sharing that important information. That's why we have been working in the trial and error mode without knowing what we're doing. And **@Edwin** also pointed out that the biggest challenge is the deviations of the trajectories from perfect helices, if the equation is not accurate enough, you get neighbour tracks' hits, and that's the problem we have. And once you get the tracks wrong in the first place, you have to spend a crazy amount of time removing the wrong hits and you will only get half of them right. Not knowing math leads to 4k LOC which still yields a low score, and I don't know if the exact equations of solving the uneven/unbounded problem have been shared in the forum, if it is, I must have overlooked some discussions. I think this is something that will set people's score apart from others and people wouldn't easily share, just my 2 cents. ",
          "votes": 3
        },
        {
          "id": 367435,
          "postDate": "2018-08-07T18:32:11.653Z",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>Elementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.</p>\n</blockquote>\n\n<p>That's very neat, indeed.</p>\n\n<p>I did a fun experiment to see how much code I can remove while staying above 0.80 (extrapolated). I got to  about 1500 code lines that give 0.7999. I ripped out one line too much, it seems. :) It's not minimized, but without obfuscation the minimum number will not be too far below it, I guess.</p>\n\n<p>It clearly shows that my approach is fundamentally more complicated than clustering. It also shows that I need about 60% of my code for getting an additional few percent of score. Brutal diminishing returns.</p>",
          "rawMarkdown": "@CPMP:\n\n&gt; Elementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.\n\nThat's very neat, indeed.\n\nI did a fun experiment to see how much code I can remove while staying above 0.80 (extrapolated). I got to  about 1500 code lines that give 0.7999. I ripped out one line too much, it seems. :) It's not minimized, but without obfuscation the minimum number will not be too far below it, I guess.\n\nIt clearly shows that my approach is fundamentally more complicated than clustering. It also shows that I need about 60% of my code for getting an additional few percent of score. Brutal diminishing returns.",
          "votes": 2
        },
        {
          "id": 367579,
          "postDate": "2018-08-08T03:47:54.430Z",
          "content": "<p>Edwin, I'd be happy to bloat my code to get your score ;)</p>",
          "rawMarkdown": "Edwin, I'd be happy to bloat my code to get your score ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 366743,
      "postDate": "2018-08-06T11:38:20.040Z",
      "content": "<p>I added something new to my track extension code and now I see negative values. I wonder if anyone have seen something similar?</p>",
      "rawMarkdown": "I added something new to my track extension code and now I see negative values. I wonder if anyone have seen something similar?\n\n",
      "replies": [
        {
          "id": 367078,
          "postDate": "2018-08-07T03:00:52.207Z",
          "content": "<p>you probably overflow integers.  </p>",
          "rawMarkdown": "you probably overflow integers.  ",
          "votes": 1
        },
        {
          "id": 367148,
          "postDate": "2018-08-07T07:23:54.670Z",
          "content": "<p>Thanks I figured it out but did not delete my post so that it does not look suspicious. It was a silly mistake in my code I could not see with my tired eyes when I posted my question :-)</p>",
          "rawMarkdown": "Thanks I figured it out but did not delete my post so that it does not look suspicious. It was a silly mistake in my code I could not see with my tired eyes when I posted my question :-)"
        }
      ]
    },
    {
      "id": 364707,
      "postDate": "2018-08-01T05:01:51.240Z",
      "content": "<p>&gt; How many people will suddenly appear in the last week with very good LB scores?</p>\n\n<p>The new participiant (Robert) came into top-11, just 1 submission. How???</p>\n\n<p><em>*</em> Dreaming of the gold medal. Wanna be 11-th. :))))</p>",
      "rawMarkdown": "&gt; How many people will suddenly appear in the last week with very good LB scores?\n\nThe new participiant (Robert) came into top-11, just 1 submission. How???\n\n*** Dreaming of the gold medal. Wanna be 11-th. :))))",
      "replies": [
        {
          "id": 364737,
          "postDate": "2018-08-01T06:21:52.240Z",
          "content": "<p>I was serious when I said yuval shared enough to get above 0.7.  Just read carefully what he shared, implement it, et voilà!</p>",
          "rawMarkdown": "I was serious when I said yuval shared enough to get above 0.7.  Just read carefully what he shared, implement it, et voilà!",
          "votes": 2
        },
        {
          "id": 364895,
          "postDate": "2018-08-01T13:32:23.987Z",
          "content": "<p>The \"just 1 submission\" is because I was depending on my local validation for making improvements.  It's safer to do it in this type of competition since there test data would be expected to be similar to train data.  Anyway 11th is not the best position to be in, it reminds me of a saying \"You don't have to be able to run faster than the bear, as long as you are not the slowest\".  I'm the slowest.</p>\n\n<p>@CPMP I think you shouldn't have any problem getting a gold medal.  I think most people like myself who delay submitting were probably just trying to improve their solution and couldn't dedicate the processor time to making a submission from the huge test set.</p>",
          "rawMarkdown": "The \"just 1 submission\" is because I was depending on my local validation for making improvements.  It's safer to do it in this type of competition since there test data would be expected to be similar to train data.  Anyway 11th is not the best position to be in, it reminds me of a saying \"You don't have to be able to run faster than the bear, as long as you are not the slowest\".  I'm the slowest.\n\n@CPMP I think you shouldn't have any problem getting a gold medal.  I think most people like myself who delay submitting were probably just trying to improve their solution and couldn't dedicate the processor time to making a submission from the huge test set.",
          "votes": 4
        },
        {
          "id": 364905,
          "postDate": "2018-08-01T13:46:03.297Z",
          "content": "<p>@Robert\nI wonder how long have you solve the problem? Or did you decide to join a few days ago?</p>\n\n<p>What I find surprising is that many new Kaggle contestants appear in the last 1 or 2 weeks (I'm not about this problem, but more generally). Perhaps the reason is that there are already many discussions about a problem.</p>",
          "rawMarkdown": "@Robert\nI wonder how long have you solve the problem? Or did you decide to join a few days ago?\n\nWhat I find surprising is that many new Kaggle contestants appear in the last 1 or 2 weeks (I'm not about this problem, but more generally). Perhaps the reason is that there are already many discussions about a problem.\n"
        },
        {
          "id": 364956,
          "postDate": "2018-08-01T16:02:11.530Z",
          "content": "<p>Don't worry, it wasn't a few days ago :)</p>\n\n<p>I've been at it for a while but I didn't think it was that necessary to make a submission, since I believed the feedback from local validation was good enough.  Plus it takes forever to create a submission.  I think this high cost (processor time) of making submissions is what has caused these late contestants to appear.</p>",
          "rawMarkdown": "Don't worry, it wasn't a few days ago :)\n\nI've been at it for a while but I didn't think it was that necessary to make a submission, since I believed the feedback from local validation was good enough.  Plus it takes forever to create a submission.  I think this high cost (processor time) of making submissions is what has caused these late contestants to appear.",
          "votes": 1
        },
        {
          "id": 365101,
          "postDate": "2018-08-01T22:43:39.367Z",
          "content": "<p>@Robert well done and I agree with your assessment. </p>\n\n<p>I was doing the same thing when I initially joined but read people talking about finding good events, event000001001 being one of the best. So I decided to run my simple and fastest model with a different event randomly then submit to see the score. My simple model's local validation score is not very good but I am able to find good events at the cost of racking up my submission numbers :-)</p>",
          "rawMarkdown": "@Robert well done and I agree with your assessment. \n\nI was doing the same thing when I initially joined but read people talking about finding good events, event000001001 being one of the best. So I decided to run my simple and fastest model with a different event randomly then submit to see the score. My simple model's local validation score is not very good but I am able to find good events at the cost of racking up my submission numbers :-)"
        },
        {
          "id": 365132,
          "postDate": "2018-08-02T00:16:26.313Z",
          "content": "<p>Thanks YaGana.  I don't understand the need to find good events.  You still have to submit solutions for all events.</p>",
          "rawMarkdown": "Thanks YaGana.  I don't understand the need to find good events.  You still have to submit solutions for all events.",
          "votes": 1
        },
        {
          "id": 365134,
          "postDate": "2018-08-02T00:24:18.540Z",
          "content": "<p>@Robert, I was talking about good training events especially with Bay optimization model. Train with one event and predict on test data. Initially it was my best model :-)</p>",
          "rawMarkdown": "@Robert, I was talking about good training events especially with Bay optimization model. Train with one event and predict on test data. Initially it was my best model :-)"
        },
        {
          "id": 365961,
          "postDate": "2018-08-03T18:36:24.263Z",
          "content": "<p>Robert, don't you want to make a team? \nI'll submit solution in 2 days for 0.66x score (I think). That's near you but still no good and the deadline is near. :(\nMaybe together we can break through 0.7.</p>",
          "rawMarkdown": "Robert, don't you want to make a team? \nI'll submit solution in 2 days for 0.66x score (I think). That's near you but still no good and the deadline is near. :(\nMaybe together we can break through 0.7."
        },
        {
          "id": 365981,
          "postDate": "2018-08-03T19:22:45.500Z",
          "content": "<p>Sergey thanks for the offer but I'm not planning on making a team.  Also it wouldn't be possible for us to combine solutions since my solution is a bit different from the approaches mentioned in the forums.  Most of them do a lot of dbscans then merge the track candidates, I don't have a merge stage.  Thanks again for the offer.</p>",
          "rawMarkdown": "Sergey thanks for the offer but I'm not planning on making a team.  Also it wouldn't be possible for us to combine solutions since my solution is a bit different from the approaches mentioned in the forums.  Most of them do a lot of dbscans then merge the track candidates, I don't have a merge stage.  Thanks again for the offer."
        },
        {
          "id": 365992,
          "postDate": "2018-08-03T19:45:58.767Z",
          "content": "<p>Robert, ok, as you wish. \nBy the way, it is possible to merge final submissions (by the longest track). It gives a boost. I have already merged my own submissions.</p>",
          "rawMarkdown": "Robert, ok, as you wish. \nBy the way, it is possible to merge final submissions (by the longest track). It gives a boost. I have already merged my own submissions."
        },
        {
          "id": 366031,
          "postDate": "2018-08-03T21:08:44.640Z",
          "content": "<p>Sergey, are you talking about Heng's track extension code?  I've already tried that as is, and it doesn't improve the score for me.</p>",
          "rawMarkdown": "Sergey, are you talking about Heng's track extension code?  I've already tried that as is, and it doesn't improve the score for me."
        },
        {
          "id": 366037,
          "postDate": "2018-08-03T21:18:28.260Z",
          "content": "<p>No, I mean merging different models like Nicole wrote here: \"merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks)\". It is the same as merging thousands of mini-submissions for each pair (z0, R).\nThe simple algorithm assigns a hit to the longest track in all of  (mini-)submissions. Nicole used the more complicated algorithm.</p>",
          "rawMarkdown": "No, I mean merging different models like Nicole wrote here: \"merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks)\". It is the same as merging thousands of mini-submissions for each pair (z0, R).\nThe simple algorithm assigns a hit to the longest track in all of  (mini-)submissions. Nicole used the more complicated algorithm.\n"
        },
        {
          "id": 366044,
          "postDate": "2018-08-03T21:31:56.143Z",
          "content": "<p>Ah.. I see.  My algorithm goes a different route.  I don't have multiple tracks to merge.  Each track is finalized as it is created so it's not possible to do any merging since I don't have different versions of each track.</p>",
          "rawMarkdown": "Ah.. I see.  My algorithm goes a different route.  I don't have multiple tracks to merge.  Each track is finalized as it is created so it's not possible to do any merging since I don't have different versions of each track."
        },
        {
          "id": 366144,
          "postDate": "2018-08-04T05:08:05.953Z",
          "content": "<p>Robert, the point Serguey is trying to make is that it may be beneficial to merge your tracks with his.  This is called ensembling.  To your point, ensembling is unusual here as we have to ensemble clusters, but there are ways to do it.  Heng in particular shared existing litterature and packages in the forum.  </p>\n\n<p>It does not mean you have to team of course, but I just wanted to clarify a point you seem to have missed.</p>",
          "rawMarkdown": "Robert, the point Serguey is trying to make is that it may be beneficial to merge your tracks with his.  This is called ensembling.  To your point, ensembling is unusual here as we have to ensemble clusters, but there are ways to do it.  Heng in particular shared existing litterature and packages in the forum.  \n\nIt does not mean you have to team of course, but I just wanted to clarify a point you seem to have missed."
        },
        {
          "id": 366152,
          "postDate": "2018-08-04T05:33:49.077Z",
          "content": "<p>Thanks, CPMP! That's what I mean. You described it better than me. :)</p>",
          "rawMarkdown": "Thanks, CPMP! That's what I mean. You described it better than me. :)",
          "votes": 1
        },
        {
          "id": 366153,
          "postDate": "2018-08-04T05:35:27.907Z",
          "content": "<p>&gt;  I don't have different versions of each track</p>\n\n<p>Robert, you can also try to ensemble your own models with different params/features/etc.</p>",
          "rawMarkdown": "&gt;  I don't have different versions of each track\n\nRobert, you can also try to ensemble your own models with different params/features/etc."
        },
        {
          "id": 366174,
          "postDate": "2018-08-04T07:45:41.133Z",
          "content": "<blockquote>\n  <p>we have to ensemble clusters, but there are ways to do it. Heng in particular shared existing litterature and packages in the forum. </p>\n</blockquote>\n\n<p>Do you mean this topic?\n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59635\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59635</a></p>",
          "rawMarkdown": "&gt; we have to ensemble clusters, but there are ways to do it. Heng in particular shared existing litterature and packages in the forum. \n\nDo you mean this topic?\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/59635\n"
        },
        {
          "id": 366201,
          "postDate": "2018-08-04T10:25:30.707Z",
          "content": "<p>Heng shared a number of references on cluster ensembling, in another topic I think.</p>",
          "rawMarkdown": "Heng shared a number of references on cluster ensembling, in another topic I think."
        },
        {
          "id": 366207,
          "postDate": "2018-08-04T10:46:11.277Z",
          "content": "<p>I've found it! <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58323\">https://www.kaggle.com/c/trackml-particle-identification/discussion/58323</a></p>",
          "rawMarkdown": "I've found it! https://www.kaggle.com/c/trackml-particle-identification/discussion/58323\n",
          "votes": 3
        },
        {
          "id": 366208,
          "postDate": "2018-08-04T10:57:07.867Z",
          "content": "<p>Right, that's the one!</p>",
          "rawMarkdown": "Right, that's the one!"
        },
        {
          "id": 366283,
          "postDate": "2018-08-04T16:33:12.410Z",
          "content": "<p>Sergey and CPMP, thanks for the info.  If there was more time I would try that approach.  Currently my setup does not provide me with different cluster candidates for each point, so I can't do any ensembling at the moment.  I would have to abandon what I have so far and basically start from scratch to go that route, which is much too risky since I am still working on improving my current score.  I'm curious to know if you all are using your own custom ensembling or are you using any packages.   I found this one for python\n<a href=\"https://pypi.org/project/Cluster_Ensembles/\">https://pypi.org/project/Cluster_Ensembles/</a></p>",
          "rawMarkdown": "Sergey and CPMP, thanks for the info.  If there was more time I would try that approach.  Currently my setup does not provide me with different cluster candidates for each point, so I can't do any ensembling at the moment.  I would have to abandon what I have so far and basically start from scratch to go that route, which is much too risky since I am still working on improving my current score.  I'm curious to know if you all are using your own custom ensembling or are you using any packages.   I found this one for python\nhttps://pypi.org/project/Cluster_Ensembles/\n"
        },
        {
          "id": 366293,
          "postDate": "2018-08-04T17:22:25.923Z",
          "content": "<p>@Sergey, small correction, I found the log of my short run (30- 60 minutes including post-processing). I randomly chose z-shifts, the reason why I tried this was that @CPMP posted something saying he could get 0.7 within 30 minutes, so I wanted to try if it's possible to get 0.7 within 30 minutes too, so I removed most z-shifts and only chose one for each model. All the features we use were mentioned in this forum, just this combination works well for our code. However, this is our old approach, we're still working on a new one. </p>\n\n<pre><code>0.62101987  (0,0, -3) \n0.57350157  (0,0,-3) \n0.53517226 (0,0,1)\n0.51663778 (0,0,-2)\n 0.54309654 (0,0,-2)\n\nmerged score: 0.66708409\n</code></pre>",
          "rawMarkdown": "@Sergey, small correction, I found the log of my short run (30- 60 minutes including post-processing). I randomly chose z-shifts, the reason why I tried this was that @CPMP posted something saying he could get 0.7 within 30 minutes, so I wanted to try if it's possible to get 0.7 within 30 minutes too, so I removed most z-shifts and only chose one for each model. All the features we use were mentioned in this forum, just this combination works well for our code. However, this is our old approach, we're still working on a new one. \n\n\n    0.62101987  (0,0, -3) \n    0.57350157  (0,0,-3) \n    0.53517226 (0,0,1)\n    0.51663778 (0,0,-2)\n     0.54309654 (0,0,-2)\n    \n    merged score: 0.66708409\n\n"
        },
        {
          "id": 366312,
          "postDate": "2018-08-04T18:29:06.567Z",
          "content": "<p>@Robert,  you don't get what I want to say, it is absolutely not about you changing anything in what you do so far.  Given I can't find another way to say what I said I'll leave it here.</p>",
          "rawMarkdown": "@Robert,  you don't get what I want to say, it is absolutely not about you changing anything in what you do so far.  Given I can't find another way to say what I said I'll leave it here."
        },
        {
          "id": 366341,
          "postDate": "2018-08-04T21:03:20.767Z",
          "content": "<p>@Nicole, I took 30 minutes mentioned by @CPMP to mean - 30 mins per event. Is that right if not I am doing something wrong.</p>",
          "rawMarkdown": "@Nicole, I took 30 minutes mentioned by @CPMP to mean - 30 mins per event. Is that right if not I am doing something wrong."
        },
        {
          "id": 366344,
          "postDate": "2018-08-04T21:11:40.660Z",
          "content": "<p>@YaGana, it depends on how you cluster your hits and which features you use, there's no fixed amount of time. The more representative features you use, the faster you can get a high score, you are not doing anything wrong, just using a different approach.</p>",
          "rawMarkdown": "@YaGana, it depends on how you cluster your hits and which features you use, there's no fixed amount of time. The more representative features you use, the faster you can get a high score, you are not doing anything wrong, just using a different approach.",
          "votes": 2
        },
        {
          "id": 366360,
          "postDate": "2018-08-04T22:27:26.070Z",
          "content": "<p>Thanks @Nicole, I am still at 26 minutes per event but working on tuning the model to improve scores. I will re-evaluate my features on the side to see if I have something that is not very useful there as well.</p>",
          "rawMarkdown": "Thanks @Nicole, I am still at 26 minutes per event but working on tuning the model to improve scores. I will re-evaluate my features on the side to see if I have something that is not very useful there as well."
        },
        {
          "id": 366386,
          "postDate": "2018-08-05T01:23:36.007Z",
          "content": "<blockquote>\n  <p>I took 30 minutes mentioned by @CPMP to mean - 30 mins per event</p>\n</blockquote>\n\n<p>Yes, that's what I meant.</p>",
          "rawMarkdown": "&gt; I took 30 minutes mentioned by @CPMP to mean - 30 mins per event\n\nYes, that's what I meant.",
          "votes": 1
        },
        {
          "id": 366393,
          "postDate": "2018-08-05T02:03:49.930Z",
          "content": "<p>Thanks @CPMP.</p>",
          "rawMarkdown": "Thanks @CPMP."
        },
        {
          "id": 366560,
          "postDate": "2018-08-05T20:17:59.867Z",
          "content": "<blockquote>\n  <p>I found this one for python <a href=\"https://pypi.org/project/Cluster_Ensembles/\">https://pypi.org/project/Cluster_Ensembles/</a></p>\n</blockquote>\n\n<p>I've tried it, but without success. I wrote about it here: <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58323\">https://www.kaggle.com/c/trackml-particle-identification/discussion/58323</a></p>",
          "rawMarkdown": "&gt; I found this one for python https://pypi.org/project/Cluster_Ensembles/\n\nI've tried it, but without success. I wrote about it here: https://www.kaggle.com/c/trackml-particle-identification/discussion/58323"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 365129,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2018-08-02T00:09:36.007000",
      "content": "<p>Now me and @Zidmie are experiencing a feeling of being watched and chased by someone right behind us... :-) @Finnies</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 364310,
      "author_name": "ykit",
      "author_url": "",
      "post_date": "2018-07-31T09:26:59.250000",
      "content": "<p>@CPMP, I am still struggling around low score, but I would like to get more higher score of course. So I am not only trying algorithms but also reviewing various kernels and discussions until now (\"standing on the shoulders of giants\").  However, I have a concern that I must not steal the giants' ideas/solutions. I think that I must add my originality on it.\nWhat are \"open\", ideal competitions for you ?\n* I am not good at English, sorry if this is difficult to read.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 366143,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T05:02:52.833000",
          "content": "<p>Sorry, I only read this now, I was away most of the week.  This competition is unusual for Kaggle as it is not clear it is a supervised machine learning competition.  And it is clearly a research competition.</p>\n\n<p>Have you tried the playground competitions?  If not then I would start there.  If you have, then maybe the home credit default risk competition is a good one.  The Santander one is weird because of  a massive leak which makes it also rather unusual.  I have not looked at other ongoing ones like airbus or tgs, but they seem interesting too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366628,
          "author_name": "ykit",
          "author_url": "",
          "post_date": "2018-08-06T04:30:44.807000",
          "content": "<p>Thank you for your reply. I have not tried the playground competitions, but I have joined this TrackML just because of my interest in physics/astrophysics (I am also interested in data science, of course). </p>\n\n<p>I found a suitable playground competition <a href=\"https://www.kaggle.com/c/flavours-of-physics-kernels-only\">\"Flavours of Physics\"</a>. I have not read details yet, but I will join it. I will also check competitions that you mentioned.</p>\n\n<p>Thank you very much !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366635,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-06T05:12:26.070000",
          "content": "<p>@ykit if you don't mind me adding my 2 cents here. I suggest you take a look at the completed playground competitions because they are geared towards teaching new comers to Kaggle and its full of great kernels to learn from. I would suggest the <a href=\"https://www.kaggle.com/c/nyc-taxi-trip-duration\">\"New York City Taxi Trip Duration\"</a> as a good one because I took part in it as a way of giving back yet I learnt a lot about visualization tools in python.</p>\n\n<p>By the way, I noticed you thanked @CPMP but did not up-vote his answer. Since you said you are new I thought I should mention that up-voting is how we show our appreciations on Kaggle :-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 366701,
          "author_name": "ykit",
          "author_url": "",
          "post_date": "2018-08-06T08:51:13.417000",
          "content": "<p>@YaGana thank you for your suggestion. I upvote comments. I am enjoying and learning a lot in this competition :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 363425,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "2018-07-28T23:56:02.727000",
      "content": "<p>You got me. I decide to submit in the last week. But I think there is not much people will above 0.8 since it is hard for a clustering only approach, and DL only is harder.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 363493,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-29T08:58:10.447000",
          "content": "<p>Hi, I'm not rally targeting you here.  indeed, you disclosed your progress on a regular basis, even without submitting.  I guess you're now running your final submission code...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 363500,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-07-29T09:38:48.857000",
          "content": "<p>I get you. I found some big event takes 3 days to process, haha.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 363501,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-29T09:47:45.570000",
          "content": "<p>I'm sure that time is well spent and that you will blew us all on the LB!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 364152,
          "author_name": "Yang Wang",
          "author_url": "",
          "post_date": "2018-07-30T21:43:54.937000",
          "content": "<p>No wonder why I can't improve my score. It only takes me 1 hour per event...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364163,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-30T22:31:08.180000",
          "content": "<p>We have a quick solution too, I think it takes 30-60 minutes or so per event to get 0.66 with only one z-shifting + track fitting, merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks).  @CPMP said he could get above 0.7 as posted in another discussion thread. But for us, the score above 0.7 does require 2-3 more hours for dbscan clustering, binning is definitely much faster, like what @yuval and @trian and probably also @Kha and @Sergey are doing.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 364194,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-07-31T01:06:08.170000",
          "content": "<blockquote>\n  <p>No wonder why I can't improve my score. It only takes me 1 hour per event...</p>\n</blockquote>\n\n<p>I get to 0.7 in less than one our per event, and yuval too.  Keep trying ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364249,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-07-31T05:55:50.193000",
          "content": "<blockquote>\n  <p>0.66 with only one z-shifting </p>\n</blockquote>\n\n<p>Are you really only using one alternative origin for tracks? </p>\n\n<p>Or do you use one alternative origin in each direction (+/- z-axis)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364269,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-31T07:00:35.330000",
          "content": "<p>@bkKaggle, no just one, 3 or 2 (-3 or -2 should work well too) I remember, I think from 0, 0, +/-3 you get the highest dbscan score. from (0,0,0) it's about 0.02 lower I recall.  The power is merging different tracks using different models of different features + diffe, for example, we used different z over r alternatives shared in other posts (the chemist shared some too), z/r alternatives are not as sensitive to dbscan as to binning so it helps us a lot.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364317,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "2018-07-31T09:58:12.403000",
          "content": "<p>I  wonder if you all do some special track select mechanism or just use track length.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 364452,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-31T14:46:09.193000",
          "content": "<blockquote>\n  <p>from (0,0,0) it's about 0.02 lower I recall. </p>\n</blockquote>\n\n<p>@Nicole I am getting 0.02 lower score with z = 3 shifted, you observed the opposite and LB is higher? Did I get you right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364465,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-31T15:05:27.543000",
          "content": "<p>Really? It can be event dependent, I'll check when I get home, I don't use the origin 0,0,0 at all, I only use shifted values,   try to run multiple z0 to find more tracks and merge them, that's what we do. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 364488,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-07-31T16:03:30.527000",
          "content": "<p>Thanks @Nicole, I will try that when I get back home later today.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364566,
          "author_name": "Liam Finnie",
          "author_url": "",
          "post_date": "2018-07-31T19:19:42.690000",
          "content": "<p><a href=\"/atom1231\">@atom1231</a> We do 'special' merging :-) Longest-track-wins is what is used within the provided DBScan kernels, however we found that does not work too well when merging different models, different z-shifts, etc. We spent a lot of time coming up with better heuristics when merging. In a post a while ago, <a href=\"/outrunner\">@outrunner</a> mentioned his merging makes use of 'track quality', we do something similar as well. One particular area of concern when merging is that you can have one ground truth track that is partially covered by a track from one model (along with possibly some other 'noise' hits not related to that track), and partially covered by a track from a different model (again, along with possibly some other 'noise'). The challenge when merging is to identify they are part of the same track, and to 'merge' these two tracks together, rather than just selecting one or the other as a winner (but how to tell they are actually the same track, rather than two distinct tracks that should not be combined? an ongoing challenge for us....)</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 364680,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "2018-08-01T03:19:10.123000",
          "content": "<p>Thanks for all your share.</p>\n\n<p>I refer @outruuner @yuval merging flow to hierarchy check , reserve best track\n( In a track,I try to distinguish which one is noise/error point but fail so finally drop all candidates)\nbut it's time consuming.</p>\n\n<p>with dbscan  it takes 1.5 hours (with complete flow/z-shifting)  to build a event to local 0.62.\nNow I change my code to binning , and merging time is bottleneck.\n(@yuval  @CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :(   )</p>\n\n<p>That 's why I ask the question.</p>\n\n<p>In brief ,good feature/different feature combination is always the key. <br>\nI did not do more fail case analysis  and the goal is still far away.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364736,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-01T06:19:56.667000",
          "content": "<blockquote>\n  <p>@CPMP said using binning 3+ min to get 0.6 , but my merging take 1 min+ at the time :( )</p>\n</blockquote>\n\n<p>Im a using DBSCAN, not binning.  I get to 0.7 with DBSCAN in less than one hour, about 30 min actually.  </p>\n\n<p>I tried binning recently, and while I could get it run, I could not get over 0.72 with it, whereas DBSCAN gives me nearly 0.75 now.  DBSCAN is slower but better, at least for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364758,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-01T07:21:33.430000",
          "content": "<p>@CPMP you are right that the maximum initial score with binning is lower then with dbsacn (we also get only to about 0.72) but this can be easily  compensated by simple tack extending =&gt; at the end results are the same.\n(a 0.63 binning is extended to 0.72 in less then 3 min, and a 0.72 is extended to 0.78 also in 3 min)</p>\n\n<p>@ atom1231\nmy 3 min binning also include the merging which is done very simply - select the longest track.</p>\n\n<p>This is a code for binning and merging with previous tracks.</p>\n\n<pre><code>        hits['cat']=(K1*F1).astype('int64')+(K2*F2).astype('int64')*10000            \n        un,inv,count = np.unique(hits['cat'],return_inverse=True, return_counts=True)\n        hits['new_track_id']=inv+10000000    #you need to offset new ids\n        hits['new_track_size']=count[inv]\n        better = (hits.new_track_size&gt;hits.track_size)\n        hits['track_id']=hits['new_track_id'].where(better,hits.track_id)\n</code></pre>\n\n<p>This code if for 2 features F1, F2.  K1,K2 define the bin size</p>\n\n<p>I don't do any sophisticated merging. \nJust some post clustering  track extension  </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 364769,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "2018-08-01T07:49:18.270000",
          "content": "<p>@yuval\nThanks for your share , I did similar in my code.\nI will keep on finding good features~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364883,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-01T12:57:20.200000",
          "content": "<p>@YaGana, you're right, I didn't use (0,0,0) so I didn't know the score of the origin was higher, I just tried it, this is my raw score for two different z0s.  Cool, maybe I should add this model to our jumbo model too.  :D  </p>\n\n<pre><code>(0,0,-3)\nfor event 1000: 0.50525238\n\n(0,0,0)\nfor event 1000: 0.51460913\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 365097,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-01T22:27:26.927000",
          "content": "<p>@Nicole, thanks for confirming my observation. When I read your comment about it, I started to think that there is something wrong with my implementation. Because of the time it takes per event, I tested my changes on a small number of training events and observed :</p>\n\n<p>&gt; <strong>(0,0,0) Average of 5 events: 0.5338</strong></p>\n\n<p>&gt; <strong>(0,0,3) Average of 5 events: 0.5327</strong></p>\n\n<p>I am seeing a much better performance on (0,0,0) with the current version of my model, the average of 5 events : 0.57+ . I will know how this model does on the LB tomorrow as it takes almost 2 days to run on my weak HW.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365102,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-01T22:47:20.310000",
          "content": "<p>@YaGana Are those raw scores right from the dbscan? Looking good, don't think anything wrong with your implementation, your scores are higher than our raw scores :) from some z shifts we only got 0.35 or so and merge them together, the scores don't mean much themselves, very different tracks can be found with different weights and iterations and features, and they often yield low scores. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365133,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-02T00:18:55.423000",
          "content": "<p>@Nicole thanks.  I am not sure what you mean by raw scores but my dbscan clusterer  implementation include pre-processing, outliers removal etc. then do prediction all in one go. Perhaps not the fastest way but that is what I have at the moment as I do not have that much time to work on this going forward.</p>\n\n<p>I just realized I swapped my results posted above between (0,0,0) and (0,0,3). I have edited and fixed it as (0,0,0) is higher.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365185,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-08-02T04:35:17.560000",
          "content": "<p>Why use +/- 3mm as the zshift? I found <a href=\"https://www.kaggle.com/bkkaggle/distribution-of-vz\">here</a> that most of the particles near the origin are distributed in peaks around -2mm, 0mm, and +2mm</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365196,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-02T05:20:30.727000",
          "content": "<p>@bkKaggle Yeah +-2.75mm after CERN corrected it in the forum a month ago, however, exploring different shifts gets different tracks though they yield lower scores since there are less tracks starting outside the beam beam collision region, so we merge them together.\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a> (the doc still says 55 mm though)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 365201,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-02T05:26:57.973000",
          "content": "<p>@YaGana, I see, the raw score I mean is the score without preprocessing/postprocessing that comes right out of dbscan (except the longest track winning code between each iteration of dbscan, which is kind of postprocessing too) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365527,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-02T20:38:07.670000",
          "content": "<p>I am amazed how much mileage people get out of clustering and binning approaches. After my EDA and after playing a bit with the public kernels, I discarded both ideas. My estimation was that clustering approaches would get stuck at about 0.65 with weight coming mostly from the forward regions. @CPMP's numbers show that this was an underestimation.</p>\n\n<p>I also discarded binning quite early since I concluded that one needed high precision in helix parameter space for separating the tracks and I assumed the required number of bins would make it intractable. It seems I underestimated this approach even more, when I read @yuval's comments.</p>\n\n<p>It will be very interesting to see how much variety of approaches one will find in the winner's  solutions.</p>\n\n<p>P.S. Thinking twice about @yuval's comments, I realize that he found an efficient way to do sparse binning of hits. Back then I thought more into the direction of binning not only hits but associated sub-manifolds in an image space as is done in the Hough transformation. I still think that would be infeasible with the required precision, but who knows what will turn up when the best solutions are opened.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365744,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-03T09:47:38.900000",
          "content": "<blockquote>\n  <p>I am amazed how much mileage people get out of clustering and binning approaches. </p>\n</blockquote>\n\n<p>I think it can go way higher ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365794,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-03T12:37:47.603000",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>I think it can go way higher ;)</p>\n</blockquote>\n\n<p>Do you mean just binning/clustering alone or plus track extension?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365803,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-03T12:51:45.743000",
          "content": "<p>I mean so far I only applied it to centered tracks.  I think I found how to also apply it to out of center tracks.  But it will take days before it shows (if it shows...)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365816,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-03T13:15:43.163000",
          "content": "<p>I think you are correct. It could be applied to out of center tracks. The issues are:</p>\n\n<ol>\n<li>Where to start?</li>\n<li>How to select good tracks. (This time only length is not enough)</li>\n<li>It is much slower then centered tracks</li>\n</ol>\n\n<p>Not sure we'll have enough time to really implement our solution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365821,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-03T13:18:27.903000",
          "content": "<blockquote>\n  <p>It is much slower then centered tracks</p>\n</blockquote>\n\n<p>Definitely.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365825,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-03T13:36:44.363000",
          "content": "<p>@yuval, @CPMP, agree, the very same equations can be applied to the non-centred tracks using exhaustive search, you need multiple physical cores to do so to meet the deadline. :)     I haven't even finished writing the code of finding all tracks from the origin, I'll pass. :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366023,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-08-03T20:35:30.063000",
          "content": "<p>Nicole, compute is nearly free...\nYou can lease a 48 core server for $15/day from Google.</p>\n\n<p>Not that it did me much good...  0.66 seems to be the best I can do on my own.</p>\n\n<p>Now I need to decode the clues recently posted to \"steal\" a score &gt; 0.7 😀</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366029,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-03T20:56:10.660000",
          "content": "<p>@John, welcome back!!! Thanks for the info :D  we do have access to a CPU / GPU farm for Kaggle and for free!  lol   Just I don't think we'll have time to implement it anymore since we're still working on a new solution, so I'll pass. :p  Hang in there, only 10 days are left, then we can take a break from Kaggle, YAH! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366036,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-03T21:18:07.433000",
          "content": "<p>@yuval r, I tried to test a partial implementation and found it to be super slow and gave up on it. I had to focus on my relatively faster modifications of improving my results on the centered tracks. There is just not enough time to implement a laborious solution at this stage. Oh time :-)</p>\n\n<p>@Nicole, did you end up adding the (0,0,0) model to your ensemble? Seen any improvement?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366057,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-03T22:11:54.370000",
          "content": "<p>@YaGana  There is no reason to try and find off centered tracks before you get to at least 0.77. We moved only now to this area. Our current score use only centered model (with a small Z shift of course)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 366058,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-03T22:13:47.707000",
          "content": "<p>@YaGana, I haven't added the (0,0,0) model to my existing jumbo model yet, but I think I'll add it to the new model if it works well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366060,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-03T22:26:28.887000",
          "content": "<p>Thanks for the advise @yuval r and for all the tips you have shared in this contest.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366147,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-08-04T05:14:42.853000",
          "content": "<p>I guess off center tracks are basically tracks with a large zshift.</p>\n\n<p>It seems that my problem is not finding more tracks, but finding higher quality tracks and merging them together to generate longer tracks. The average length of my tracks is 2-5 while the average length of the ground truth particles is around 12.</p>\n\n<p>Somewhere else on the discussion forum, I saw that the organizers said that the emphasis is more on finding high quality track candidates than on perfecting the length of the tracks. If this is true, does this mean that CERN has track merging approaches that are a lot more sophisticated than the length and quality based approaches that are being used here?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366148,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T05:25:28.180000",
          "content": "<blockquote>\n  <p>I guess off center tracks are basically tracks with a large zshift.</p>\n</blockquote>\n\n<p>No, they are tracks that do not get near z axis.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366184,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-04T08:21:04.157000",
          "content": "<p>@bkKaggle -   my 2 cents, I think those are the tracks that do not originate from (x,y) = (0,0) (the initial position of the particles)  I shared this image a couple months ago. I think the threshold for this EDA I set at the time was x&gt;1 or y&gt;1. You can see some clearly don't start from (x,y) = (0,0)  or get close to the z-axis by CPMP's definition. </p>\n\n<p><a href=\"https://drive.google.com/file/d/1-nsvtrkDWtnXHO1C-5C9uwHBkceqZhhe/view?usp=sharing\">tracks of 9 hits that don't start from near the (0,0,z) </a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366204,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T10:37:02.663000",
          "content": "<p>A trajectory may start far from z axis, yet be on an helix that goes through the z axis.  Such tracks can be caught by code that look for helix going through the z axis.  An out of center helix is an helix that never is close to the z axis.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366542,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-05T18:27:12.273000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 363924,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-07-30T10:44:51.217000",
      "content": "<p>Good to see icecuber move in the open.  How many others to come?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 363929,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-07-30T10:52:20.330000",
          "content": "<p>I thought the same too :)  Mickey, Heng, the chemist, and others who haven't made submissions yet :D  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 385948,
      "author_name": "David Rousseau",
      "author_url": "",
      "post_date": "2018-09-11T20:57:58.517000",
      "content": "<p>The second \"Throughput\" phase of this competition is online, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/65525\">this post</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 366794,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-08-06T14:25:42.993000",
      "content": "<p>Oh, the new participant (bestfitting) has come with the good score. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 366823,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-08-06T15:29:39.450000",
          "content": "<p>I think we should expect a lot of that this week...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366830,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T15:45:29.503000",
          "content": "<p>Indeed, I guess at the end of the competition, you may need 0.8 to get a gold medal. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366878,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-08-06T17:11:03.170000",
          "content": "<p>I hope I can get a silver with 0.66, but that is literally what is keeping me up at night 😁</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366891,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T17:34:09.920000",
          "content": "<p>@John, that's certain ;)  today is the final day of making the first submission if I understood the rules correctly. With your current score, you'll get a high silver medal for sure if you do nothing until the end of the competition. ;) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366896,
          "author_name": "Matthew Wander",
          "author_url": "",
          "post_date": "2018-08-06T17:45:25.630000",
          "content": "<p>Nicole,  Are you sure? I thought we had to accept the rules in order to download the data? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366902,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T17:57:56.383000",
          "content": "<p>No I'm not sure , under the time line , the first rule indicates today is only accepting the rules to compete, the second rule says , August 6, 2018 - Team Merger deadline. This is the last day participants may join or merge teams. So its not too clear what this means, maybe it only applies to team merge, i thought you have to make the first sub to be able to merge with another team. A related discussion posted by <a href=\"/inversion\">@inversion</a> 2 years ago\n<a href=\"https://www.kaggle.com/product-feedback/20451\">https://www.kaggle.com/product-feedback/20451</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366904,
          "author_name": "Matthew Wander",
          "author_url": "",
          "post_date": "2018-08-06T18:05:16.053000",
          "content": "<p>Interesting, totally not clear, but that could be exactly what it means. Now the question do I make a throwaway submission today in order to keep myself in the competition, when my otherwise promising solution seems prone to every bug under the sun? Nuts.</p>\n\n<p>thanks for the info. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366907,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T18:09:25.290000",
          "content": "<p>@Matthew, I'd definitely do so to be safe. There was a turmoil after the data science bowl this year because the rules between Kaggle and the competition were conflicting each other. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366908,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-06T18:18:20.820000",
          "content": "<p>@Matthew Wander: I just read the rules again and honestly it is not clear to me what exactly constitutes \"entry\" in the competition. It is clear that \"entry\" implies the acceptance of the competition rules, but it is not clear to me whether there is an implication in the other direction.</p>\n\n<p>To be on the safe side, I'd definitely make a submission before the upcoming merger &amp; entry deadline.</p>\n\n<p>P.S.: If you're very short on time, remember that your submission only needs to be technically valid, it does not need to achieve a high score, so you could immediately submit a submission created by one of the public kernels, if you need one, or even a synthetic trivial one, I guess. (Although the latter could strictly be said to be a case of \"hand-labeling\", which is forbidden, if one interprets it in a very inflexible way.)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366939,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-06T19:18:48.783000",
          "content": "<p>@Matthew, In the past you had to have made at least one submission prior to the last week of the competition, but then I believe they changed it so that you only had to accept the rules by the last week and could make your first submission after that.  But I haven't tested this and don't want to be blamed if you get in trouble, so it might be best to make a throwaway submission like others have suggested.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 366941,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-06T19:19:50.283000",
          "content": "<p>@John Sweeney, that's funny.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366947,
          "author_name": "Matthew Wander",
          "author_url": "",
          "post_date": "2018-08-06T19:28:23.003000",
          "content": "<p>thanks all, but I think I am going to pull out. there were just too many hurdles to the solution and a submission would take 3-4 days to work out. I think I am going to take it as a sign that the clock just ran out and work on other projects.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366954,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T19:41:59.777000",
          "content": "<p>@Matthew, no no, didn't mean to discourage you, you can make a real submission before the deadline and tell us if it works ;)  and I agree with \"too many hurdles to the solution part\", it was a very tough battle for someone like me who sucks at math. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366958,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-06T19:47:09.493000",
          "content": "<blockquote>\n  <p>it was a very tough battle</p>\n</blockquote>\n\n<p>@Nicole And unfortunately I have to say that you are now in 11th place, the last for a gold medal. I hope no one will come, and you will get it.</p>\n\n<p>I think I can relax now and wait for a silver medal. :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366962,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-06T19:54:00.187000",
          "content": "<p>Thanks @Sergey, we are the slowest runners who will be caught by the bear ;) (From the joke @Robert made some days ago) - @Heng shows an amazing result in DL and I expect him to make a submission over 0.8 :)   we expect to get a silver medal as well. We've been working in the wrong direction in the past 2 months completely (building rules and writing lots of code) and justly realized it's all about math equations in the last week after I've read @yuval's post.  Hopefully our code will still be useful after we share our github repository. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366963,
          "author_name": "Matthew Wander",
          "author_url": "",
          "post_date": "2018-08-06T19:54:04.577000",
          "content": "<p>@Nicole. No discouragement, just realism. I'm too slow a python programmer to make it work in time. The math was sound: you use pairs of points to determine the features: helix radii, helix xy center, and screw axis. I would probably need to go to triplets of points to get past 85%. The challenge is getting from pairs of points back to singlets.  I had some ideas but the results are just not where they need to be. Too bad I wasted a week trying to get the cells file to yield derivatives. Shrug. Now to go and try my hand at salt deposits.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366974,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-06T20:37:08.670000",
          "content": "<p>@Matthew Wander: Honestly, I think you are underestimating the difficulty of the problem.</p>\n\n<blockquote>\n  <p>I would probably need to go to triplets of points to get past 85%.</p>\n</blockquote>\n\n<p>If you have good reasons to think that you can get 85% with rather simple code, you must definitely make a submission for the sake of science! (I say that even though you most likely would kick me down the LB :-) Such a solution would be really interesting to the CERN people, I think.</p>\n\n<p>I have no idea what the leaders are doing, but I'd be surprised if their code is quite simple.\nMy own solution is currently at about 3.8K pure code lines. Not all of it is necessary and I will remove some dead code before the conclusion, but still it's a lot.</p>\n\n<p>I somehow expect the winners' solutions to be much more concise and elegant than mine. Let's see.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 366982,
          "author_name": "Liam Finnie",
          "author_url": "",
          "post_date": "2018-08-06T20:53:22.450000",
          "content": "<p>@edwin Good to know our team is not the only one writing lots of code - we have about 4K lines of code, mostly trying to deal with all the complications as a result of EDA (i.e. how to extend different types of tracks, how to detect outliers, etc.). Definitely a lot more complicated than we had expected at the beginning of the competition. However, even with lots of code, we are far from the top, and far below your score (guess your 3.8K LOC are more powerful than ours, darn). I think those that understand advanced math have a big edge though, finding the right features can increase the score by a huge amount, writing lots of LOC mostly just nudges the score a little, from our experience.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 367002,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-06T21:50:16.963000",
          "content": "<blockquote>\n  <p>writing lots of LOC mostly just nudges the score a little, from our experience.</p>\n</blockquote>\n\n<p>True. My typical \"great idea\" gives me a +0.0020 increment after an implementation time between 5 minutes and two weeks, with no discernible relation between time invested and score improvement ;-/.</p>\n\n<p>There were some very satisfying +0.0200 ideas, but I seem to have run out of them.</p>\n\n<p>My guess would be that <a href=\"/icecuber\">@icecuber</a> &amp; co have found a good way to train a classifier for evaluation and subsequent merging of track candidates. That is currently one of the biggest weaknesses of my algorithm: the evaluation of track candidates. I tried some ML approaches but with little success. I am also completely new to ML, which doesn't help.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 367006,
          "author_name": "Matthew Wander",
          "author_url": "",
          "post_date": "2018-08-06T22:15:59.160000",
          "content": "<p>@Edwin That is simply based on my estimation of the number of points on the helix that start with X and Y both ~=0. Yes I am much further away probably at least a month to complete the internal transformation accurately enough to achieve even that level of quality. Mathematical limits start hitting hard you only need four points to confirm a helix originating at 0,0 but 7 for a helix centered anywhere else. I don't know the count but I know the very best solutions are close to that limit. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367011,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-06T22:33:01.353000",
          "content": "<p>@Matthew Wander: You are correct in that the helices which do not pass through the line (0, 0, z) are much harder to find. In general, a helix with axis parallel to the z-axis has 5 degrees of freedom and every 3-d point gives you 2 (because you need to subtract one for the curve parameter). So in general you need three points to pin down the helix (with one degree to spare for checks), but if you know the helix passes through (0,0,z) you only need two points, which makes a huge difference.</p>\n\n<p>HOWEVER, the particle trajectories are NOT helices! The problem really wouldn't be that hard if they were. For many reasons, which are mostly mentioned in the introductory document by the organizers, the trajectories deviate from perfect helices, sometimes quite strongly.</p>\n\n<p>What really makes the problem so hard is that <em>the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks</em>!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 367050,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-07T01:05:38.653000",
          "content": "<blockquote>\n  <p>What really makes the problem so hard is that the deviations of the trajectories from perfect helices are comparable to the differences between neighboring tracks!</p>\n</blockquote>\n\n<p>Yep @Edwin Steiner you have succinctly defined the issue that made accurate clustering a challenge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367055,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-08-07T01:36:15.397000",
          "content": "<blockquote>\n  <p>think those that understand advanced math have a big edge though</p>\n</blockquote>\n\n<p>How much math do you need to know to understand everything? What type of math is it anyway?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367073,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-07T02:41:47.913000",
          "content": "<blockquote>\n  <p>I think those that understand advanced math have a big edge though,</p>\n</blockquote>\n\n<p>Elementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.</p>\n\n<p>I really don't think math is the issue here, all necessary equations are provided in material shared on the forum.  </p>\n\n<p>This is down to usual ML: proper EDA, proper feature engineering, choice of algorithm and model, and some tuning.</p>\n\n<p>The real split between top entries and the rest (I'm part of the rest), is to find effective ways for identifying tracks not originating near z axis.  I'm working on it but I doubt I'll get something significant by end of competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367152,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-08-07T07:45:03.017000",
          "content": "<p>I am very much looking forward to reading your solution when the competition is finished!</p>\n\n<p>240 LOC ==&gt; 0.778 sounds like a beautiful and elegant solution!</p>\n\n<p>Like the Finnies mentioned, I spent most of the competition trying to hand craft exotic features and merging methods when the solution was in the complete opposite direction!</p>\n\n<p>After the recent board discussions I think I know the math but am obviously missing some property of symmetry that is cutting my score in half!!!</p>\n\n<p>I'm guessing that once the competition is finished, I'll kick myself for missing something rather obvious!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 367168,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-07T08:18:34.970000",
          "content": "<p>Thanks, but I am not using any magic feature as we could see in other competitions sometimes. It really is along the line that yuval shared <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/61081#356653\">in this topic</a>, except I'm using DBSCAN instead of binning.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367182,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-07T08:57:52.863000",
          "content": "<p><strong>@John</strong>, Yes, the problem we have is we couldn't find the accurate math equations to solve the problem with the uneven magnetic field, and I didn't even realize this was a problem until 2 weeks ago.  Thanks to <strong>@yuval,</strong> it was very generous of him sharing that important information. That's why we have been working in the trial and error mode without knowing what we're doing. And <strong>@Edwin</strong> also pointed out that the biggest challenge is the deviations of the trajectories from perfect helices, if the equation is not accurate enough, you get neighbour tracks' hits, and that's the problem we have. And once you get the tracks wrong in the first place, you have to spend a crazy amount of time removing the wrong hits and you will only get half of them right. Not knowing math leads to 4k LOC which still yields a low score, and I don't know if the exact equations of solving the uneven/unbounded problem have been shared in the forum, if it is, I must have overlooked some discussions. I think this is something that will set people's score apart from others and people wouldn't easily share, just my 2 cents. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 367435,
          "author_name": "Edwin Steiner",
          "author_url": "",
          "post_date": "2018-08-07T18:32:11.653000",
          "content": "<p>@CPMP:</p>\n\n<blockquote>\n  <p>Elementary math and 240 lines of code give me local score of 0.778 out of DBSCAN now without any post processing.</p>\n</blockquote>\n\n<p>That's very neat, indeed.</p>\n\n<p>I did a fun experiment to see how much code I can remove while staying above 0.80 (extrapolated). I got to  about 1500 code lines that give 0.7999. I ripped out one line too much, it seems. :) It's not minimized, but without obfuscation the minimum number will not be too far below it, I guess.</p>\n\n<p>It clearly shows that my approach is fundamentally more complicated than clustering. It also shows that I need about 60% of my code for getting an additional few percent of score. Brutal diminishing returns.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 367579,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-08T03:47:54.430000",
          "content": "<p>Edwin, I'd be happy to bloat my code to get your score ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 366743,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-08-06T11:38:20.040000",
      "content": "<p>I added something new to my track extension code and now I see negative values. I wonder if anyone have seen something similar?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 367078,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-07T03:00:52.207000",
          "content": "<p>you probably overflow integers.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367148,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-07T07:23:54.670000",
          "content": "<p>Thanks I figured it out but did not delete my post so that it does not look suspicious. It was a silly mistake in my code I could not see with my tired eyes when I posted my question :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 364707,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-08-01T05:01:51.240000",
      "content": "<p>&gt; How many people will suddenly appear in the last week with very good LB scores?</p>\n\n<p>The new participiant (Robert) came into top-11, just 1 submission. How???</p>\n\n<p><em>*</em> Dreaming of the gold medal. Wanna be 11-th. :))))</p>",
      "votes": 0,
      "replies": [
        {
          "id": 364737,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-01T06:21:52.240000",
          "content": "<p>I was serious when I said yuval shared enough to get above 0.7.  Just read carefully what he shared, implement it, et voilà!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 364895,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-01T13:32:23.987000",
          "content": "<p>The \"just 1 submission\" is because I was depending on my local validation for making improvements.  It's safer to do it in this type of competition since there test data would be expected to be similar to train data.  Anyway 11th is not the best position to be in, it reminds me of a saying \"You don't have to be able to run faster than the bear, as long as you are not the slowest\".  I'm the slowest.</p>\n\n<p>@CPMP I think you shouldn't have any problem getting a gold medal.  I think most people like myself who delay submitting were probably just trying to improve their solution and couldn't dedicate the processor time to making a submission from the huge test set.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 364905,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-01T13:46:03.297000",
          "content": "<p>@Robert\nI wonder how long have you solve the problem? Or did you decide to join a few days ago?</p>\n\n<p>What I find surprising is that many new Kaggle contestants appear in the last 1 or 2 weeks (I'm not about this problem, but more generally). Perhaps the reason is that there are already many discussions about a problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 364956,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-01T16:02:11.530000",
          "content": "<p>Don't worry, it wasn't a few days ago :)</p>\n\n<p>I've been at it for a while but I didn't think it was that necessary to make a submission, since I believed the feedback from local validation was good enough.  Plus it takes forever to create a submission.  I think this high cost (processor time) of making submissions is what has caused these late contestants to appear.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365101,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-01T22:43:39.367000",
          "content": "<p>@Robert well done and I agree with your assessment. </p>\n\n<p>I was doing the same thing when I initially joined but read people talking about finding good events, event000001001 being one of the best. So I decided to run my simple and fastest model with a different event randomly then submit to see the score. My simple model's local validation score is not very good but I am able to find good events at the cost of racking up my submission numbers :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365132,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-02T00:16:26.313000",
          "content": "<p>Thanks YaGana.  I don't understand the need to find good events.  You still have to submit solutions for all events.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365134,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-02T00:24:18.540000",
          "content": "<p>@Robert, I was talking about good training events especially with Bay optimization model. Train with one event and predict on test data. Initially it was my best model :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365961,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-03T18:36:24.263000",
          "content": "<p>Robert, don't you want to make a team? \nI'll submit solution in 2 days for 0.66x score (I think). That's near you but still no good and the deadline is near. :(\nMaybe together we can break through 0.7.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365981,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-03T19:22:45.500000",
          "content": "<p>Sergey thanks for the offer but I'm not planning on making a team.  Also it wouldn't be possible for us to combine solutions since my solution is a bit different from the approaches mentioned in the forums.  Most of them do a lot of dbscans then merge the track candidates, I don't have a merge stage.  Thanks again for the offer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365992,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-03T19:45:58.767000",
          "content": "<p>Robert, ok, as you wish. \nBy the way, it is possible to merge final submissions (by the longest track). It gives a boost. I have already merged my own submissions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366031,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-03T21:08:44.640000",
          "content": "<p>Sergey, are you talking about Heng's track extension code?  I've already tried that as is, and it doesn't improve the score for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366037,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-03T21:18:28.260000",
          "content": "<p>No, I mean merging different models like Nicole wrote here: \"merging 4 models of different features (so each model is around 0.62, 0.63, 0.53, 0.54 they meant to find different tracks)\". It is the same as merging thousands of mini-submissions for each pair (z0, R).\nThe simple algorithm assigns a hit to the longest track in all of  (mini-)submissions. Nicole used the more complicated algorithm.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366044,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-03T21:31:56.143000",
          "content": "<p>Ah.. I see.  My algorithm goes a different route.  I don't have multiple tracks to merge.  Each track is finalized as it is created so it's not possible to do any merging since I don't have different versions of each track.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366144,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T05:08:05.953000",
          "content": "<p>Robert, the point Serguey is trying to make is that it may be beneficial to merge your tracks with his.  This is called ensembling.  To your point, ensembling is unusual here as we have to ensemble clusters, but there are ways to do it.  Heng in particular shared existing litterature and packages in the forum.  </p>\n\n<p>It does not mean you have to team of course, but I just wanted to clarify a point you seem to have missed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366152,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-04T05:33:49.077000",
          "content": "<p>Thanks, CPMP! That's what I mean. You described it better than me. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366153,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-04T05:35:27.907000",
          "content": "<p>&gt;  I don't have different versions of each track</p>\n\n<p>Robert, you can also try to ensemble your own models with different params/features/etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366174,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-04T07:45:41.133000",
          "content": "<blockquote>\n  <p>we have to ensemble clusters, but there are ways to do it. Heng in particular shared existing litterature and packages in the forum. </p>\n</blockquote>\n\n<p>Do you mean this topic?\n<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59635\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59635</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366201,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T10:25:30.707000",
          "content": "<p>Heng shared a number of references on cluster ensembling, in another topic I think.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366207,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-04T10:46:11.277000",
          "content": "<p>I've found it! <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58323\">https://www.kaggle.com/c/trackml-particle-identification/discussion/58323</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 366208,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T10:57:07.867000",
          "content": "<p>Right, that's the one!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366283,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-08-04T16:33:12.410000",
          "content": "<p>Sergey and CPMP, thanks for the info.  If there was more time I would try that approach.  Currently my setup does not provide me with different cluster candidates for each point, so I can't do any ensembling at the moment.  I would have to abandon what I have so far and basically start from scratch to go that route, which is much too risky since I am still working on improving my current score.  I'm curious to know if you all are using your own custom ensembling or are you using any packages.   I found this one for python\n<a href=\"https://pypi.org/project/Cluster_Ensembles/\">https://pypi.org/project/Cluster_Ensembles/</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366293,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-04T17:22:25.923000",
          "content": "<p>@Sergey, small correction, I found the log of my short run (30- 60 minutes including post-processing). I randomly chose z-shifts, the reason why I tried this was that @CPMP posted something saying he could get 0.7 within 30 minutes, so I wanted to try if it's possible to get 0.7 within 30 minutes too, so I removed most z-shifts and only chose one for each model. All the features we use were mentioned in this forum, just this combination works well for our code. However, this is our old approach, we're still working on a new one. </p>\n\n<pre><code>0.62101987  (0,0, -3) \n0.57350157  (0,0,-3) \n0.53517226 (0,0,1)\n0.51663778 (0,0,-2)\n 0.54309654 (0,0,-2)\n\nmerged score: 0.66708409\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366312,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-04T18:29:06.567000",
          "content": "<p>@Robert,  you don't get what I want to say, it is absolutely not about you changing anything in what you do so far.  Given I can't find another way to say what I said I'll leave it here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366341,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-04T21:03:20.767000",
          "content": "<p>@Nicole, I took 30 minutes mentioned by @CPMP to mean - 30 mins per event. Is that right if not I am doing something wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366344,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-08-04T21:11:40.660000",
          "content": "<p>@YaGana, it depends on how you cluster your hits and which features you use, there's no fixed amount of time. The more representative features you use, the faster you can get a high score, you are not doing anything wrong, just using a different approach.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 366360,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-04T22:27:26.070000",
          "content": "<p>Thanks @Nicole, I am still at 26 minutes per event but working on tuning the model to improve scores. I will re-evaluate my features on the side to see if I have something that is not very useful there as well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366386,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-05T01:23:36.007000",
          "content": "<blockquote>\n  <p>I took 30 minutes mentioned by @CPMP to mean - 30 mins per event</p>\n</blockquote>\n\n<p>Yes, that's what I meant.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366393,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-08-05T02:03:49.930000",
          "content": "<p>Thanks @CPMP.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366560,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-05T20:17:59.867000",
          "content": "<blockquote>\n  <p>I found this one for python <a href=\"https://pypi.org/project/Cluster_Ensembles/\">https://pypi.org/project/Cluster_Ensembles/</a></p>\n</blockquote>\n\n<p>I've tried it, but without success. I wrote about it here: <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/58323\">https://www.kaggle.com/c/trackml-particle-identification/discussion/58323</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "363371": "How many people will suddenly appear in the last week with very good LB scores?\n\nWill a score above 0.8 (which is my goal) be sufficient to get a gold medal?\n\nThese questions are really open as one can rely on local score to assess progress.  it remind me of the [wikipedia forecasting competition][1] where a number of people (including #1) submitted good solutions without submitting much during the competition.  I was submitting my best there,  and here too.\n\nI really appreciate that demelian is open to show us his current results.  I wish other good contenders were doing the same.  \n\n\n  [1]: https://www.kaggle.com/c/web-traffic-time-series-forecasting",
    "365129": "Now me and @Zidmie are experiencing a feeling of being watched and chased by someone right behind us... :-) @Finnies",
    "364310": "@CPMP, I am still struggling around low score, but I would like to get more higher score of course. So I am not only trying algorithms but also reviewing various kernels and discussions until now (\"standing on the shoulders of giants\").  However, I have a concern that I must not steal the giants' ideas/solutions. I think that I must add my originality on it.\nWhat are \"open\", ideal competitions for you ?\n* I am not good at English, sorry if this is difficult to read.",
    "363425": "You got me. I decide to submit in the last week. But I think there is not much people will above 0.8 since it is hard for a clustering only approach, and DL only is harder.",
    "363924": "Good to see icecuber move in the open.  How many others to come?",
    "385948": "The second \"Throughput\" phase of this competition is online, see [this post][1]\n\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/65525",
    "366794": "Oh, the new participant (bestfitting) has come with the good score. \n",
    "366743": "I added something new to my track extension code and now I see negative values. I wonder if anyone have seen something similar?\n\n",
    "364707": "&gt; How many people will suddenly appear in the last week with very good LB scores?\n\nThe new participiant (Robert) came into top-11, just 1 submission. How???\n\n*** Dreaming of the gold medal. Wanna be 11-th. :))))"
  }
}