{
  "id": 62876,
  "title": "0.8 without supervised learning",
  "url": "/competitions/trackml-particle-identification/discussion/62876",
  "author_name": "",
  "post_date": "2018-08-08T11:57:53.648758400Z",
  "votes": 16,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I just got over 0.8 local score using only DBSCAN and some post processing.  DBSCAN alone gets to approx 0.785.  I thought I'd share it because it contradicts what several of us (including me) wrote before, namely that supervised learning was required to move over 0.8.</p>\n\n<p>I would not be surprised now that some of the people with better score than 0.8 also get there without supervised learning.</p>\n\n<p>The above is for a code that only looks for tracks originating from the z axis.  It finds over 92% of them.  If you do the math, then you'll see it catches some of the other tracks, but not much.</p>",
  "messages": [
    {
      "id": "367726",
      "postDate": "08/08/2018 11:57:53",
      "content": "<p>I just got over 0.8 local score using only DBSCAN and some post processing.  DBSCAN alone gets to approx 0.785.  I thought I'd share it because it contradicts what several of us (including me) wrote before, namely that supervised learning was required to move over 0.8.</p>\n\n<p>I would not be surprised now that some of the people with better score than 0.8 also get there without supervised learning.</p>\n\n<p>The above is for a code that only looks for tracks originating from the z axis.  It finds over 92% of them.  If you do the math, then you'll see it catches some of the other tracks, but not much.</p>",
      "rawMarkdown": "I just got over 0.8 local score using only DBSCAN and some post processing.  DBSCAN alone gets to approx 0.785.  I thought I'd share it because it contradicts what several of us (including me) wrote before, namely that supervised learning was required to move over 0.8.\n\nI would not be surprised now that some of the people with better score than 0.8 also get there without supervised learning.\n\nThe above is for a code that only looks for tracks originating from the z axis.  It finds over 92% of them.  If you do the math, then you'll see it catches some of the other tracks, but not much.",
      "votes": null
    },
    {
      "id": "367872",
      "postDate": "08/08/2018 18:03:59",
      "content": "<p>Thanks for sharing. I've read almost all of your posts and trying to build a proper DBSCAN pipeline, but I can't jump over 0.2 local score (shame!). I believe I've missed something very obvious. </p>\n\n<p>So my pipeline is:\n1) Select features for clustering (2-3 variables)\n2) Do helix unrolling on some angle and shift z somewhere\n3) Merge current clusters to the result ones\n4) Repeat 1-3 \n5) Remove outliers to 0-cluster\n6) Do track extension </p>",
      "rawMarkdown": "Thanks for sharing. I've read almost all of your posts and trying to build a proper DBSCAN pipeline, but I can't jump over 0.2 local score (shame!). I believe I've missed something very obvious. \n\nSo my pipeline is:\n1) Select features for clustering (2-3 variables)\n2) Do helix unrolling on some angle and shift z somewhere\n3) Merge current clusters to the result ones\n4) Repeat 1-3 \n5) Remove outliers to 0-cluster\n6) Do track extension",
      "votes": null
    },
    {
      "id": "367875",
      "postDate": "08/08/2018 18:14:50",
      "content": "<p>Hi, there are some public kernels with DBSCAN that have scores close to 0.5.  You may learn a bit from them.  Note that my 0.785 local score comes without any post processing, no outlier removal nor track extension.  Just to say that you can get some mileage out of clustering itself.</p>",
      "rawMarkdown": "Hi, there are some public kernels with DBSCAN that have scores close to 0.5.  You may learn a bit from them.  Note that my 0.785 local score comes without any post processing, no outlier removal nor track extension.  Just to say that you can get some mileage out of clustering itself.",
      "votes": null
    },
    {
      "id": "367908",
      "postDate": "08/08/2018 19:59:49",
      "content": "<p>Cool! Will you share your code after the competition ends or will you hide it for the second phase?</p>",
      "rawMarkdown": "Cool! Will you share your code after the competition ends or will you hide it for the second phase?",
      "votes": null
    },
    {
      "id": "367918",
      "postDate": "08/08/2018 20:23:52",
      "content": "<p>Hi, probably the latter, but I will share my code  anyway, in few days, or after 2nd phase, for sure.</p>",
      "rawMarkdown": "Hi, probably the latter, but I will share my code  anyway, in few days, or after 2nd phase, for sure.",
      "votes": null
    },
    {
      "id": "367937",
      "postDate": "08/08/2018 21:31:25",
      "content": "<p>How long did it take your DBSCAN to run?</p>",
      "rawMarkdown": "How long did it take your DBSCAN to run?",
      "votes": null
    },
    {
      "id": "368054",
      "postDate": "08/09/2018 06:35:51",
      "content": "<blockquote>\n  <p>How long did it take your DBSCAN to run?</p>\n</blockquote>\n\n<p>About 10 hours per event with an i7 at 4GHz.  Code can probably be made faster, but not much, unless I stop using pandas and undertake a massive rewrite.</p>",
      "rawMarkdown": "&gt;How long did it take your DBSCAN to run?\n\nAbout 10 hours per event with an i7 at 4GHz.  Code can probably be made faster, but not much, unless I stop using pandas and undertake a massive rewrite.",
      "votes": null
    },
    {
      "id": "368072",
      "postDate": "08/09/2018 07:50:24",
      "content": "<p>@Sergey, though our code is not as good, @Liam and myself will share our code esp. the accurate helix unrolling function + accurate 2 features right after this competition  (actually @yuval said very clearly in his post what the 2nd feature is), e.g. 35 seconds to get 0.41 and one minute per event to reach 0.5 using dbscan on my laptop. we've never run an event too long, and we have good merging code, people can find some ideas from it for the 2nd phase <strong>performance</strong> competition.</p>",
      "rawMarkdown": "Sergey, though our code is not as good, @Liam and myself will share our code esp. the accurate helix unrolling function + accurate 2 features right after this competition  (actually @yuval said very clearly in his post what the 2nd feature is), e.g. 35 seconds to get 0.41 and one minute per event to reach 0.5 using dbscan on my laptop. we've never run an event too long, and we have good merging code, people can find some ideas from it for the 2nd phase **performance** competition.",
      "votes": null
    },
    {
      "id": "368073",
      "postDate": "08/09/2018 07:53:53",
      "content": "<p>@Nicole \nGreat! It is interesting to look at your code.</p>",
      "rawMarkdown": "Nicole \nGreat! It is interesting to look at your code.",
      "votes": null
    },
    {
      "id": "368373",
      "postDate": "08/09/2018 19:33:16",
      "content": "<p>Yikes, at 10 hrs/event it would take ~60 days to create a submission!!!\nThat is a lot of compute 😁</p>\n\n<p>Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left. </p>\n\n<p>I think they are leading me to a new submission that should be right around 0.7, maybe a little under, since I need to keep the processing time short to make the Monday deadline.</p>",
      "rawMarkdown": "Yikes, at 10 hrs/event it would take ~60 days to create a submission!!!\nThat is a lot of compute 😁\n\nAnyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left. \n\nI think they are leading me to a new submission that should be right around 0.7, maybe a little under, since I need to keep the processing time short to make the Monday deadline.",
      "votes": null
    },
    {
      "id": "368396",
      "postDate": "08/09/2018 20:16:08",
      "content": "<p>@Nicole, I look forward to seeing your code.</p>",
      "rawMarkdown": "Nicole, I look forward to seeing your code.",
      "votes": null
    },
    {
      "id": "368404",
      "postDate": "08/09/2018 20:33:58",
      "content": "<p>@John, good job!!! :D  Look forward to your score! We feel so good at a silver medal position, that's where we should belong to, kind of like ducks like to play in a muddy pond. You know we suck at geometry... not sure how people can get a hold on it. I'm glad we decided to become software engineers who won't blow up a rocket because we suck at math. :D (though SpaceX would really be cool to work for)</p>",
      "rawMarkdown": "John, good job!!! :D  Look forward to your score! We feel so good at a silver medal position, that's where we should belong to, kind of like ducks like to play in a muddy pond. You know we suck at geometry... not sure how people can get a hold on it. I'm glad we decided to become software engineers who won't blow up a rocket because we suck at math. :D (though SpaceX would really be cool to work for)",
      "votes": null
    },
    {
      "id": "368410",
      "postDate": "08/09/2018 20:43:55",
      "content": "<blockquote>\n  <p>We feel so good at a silver medal position</p>\n</blockquote>\n\n<p>By the way you can swap with Mickey in the private LB.  So your silver medal is in question, maybe gold. :)</p>",
      "rawMarkdown": "&gt; We feel so good at a silver medal position\n\nBy the way you can swap with Mickey in the private LB.  So your silver medal is in question, maybe gold. :)",
      "votes": null
    },
    {
      "id": "368415",
      "postDate": "08/09/2018 21:15:49",
      "content": "<p>@Sergey lol lots of people haven't made submissions for a while including @Mickey and @Heng. I think this position is pretty comfy for us. We're only working on the post processing code right now. I look forward to their scores too!  </p>",
      "rawMarkdown": "Sergey lol lots of people haven't made submissions for a while including @Mickey and @Heng. I think this position is pretty comfy for us. We're only working on the post processing code right now. I look forward to their scores too!",
      "votes": null
    },
    {
      "id": "368423",
      "postDate": "08/09/2018 21:34:56",
      "content": "<p>@Nicole, Thanks!  </p>\n\n<p>Fingers crossed, the cake is currently baking 🍰</p>\n\n<p>With an effective rate of 3 events/hour, my next submission should be sometime on Saturday!</p>",
      "rawMarkdown": "Nicole, Thanks!  \n\nFingers crossed, the cake is currently baking 🍰\n\nWith an effective rate of 3 events/hour, my next submission should be sometime on Saturday!",
      "votes": null
    },
    {
      "id": "368503",
      "postDate": "08/10/2018 03:21:15",
      "content": "<blockquote>\n  <p>Yikes, at 10 hrs/event it would take ~60 days to create a submission!!! That is a lot of compute 😁</p>\n</blockquote>\n\n<p>Outrunner says he is using a lot more time per event, up to 3 days for the largest ones, and it will probably show in the LB ;)</p>\n\n<p>The only way to go is to use parallelism (see the recent discussion about how to do it in Python).  I have access to a 40 core IBM machine, which helps quite a bit ;)</p>",
      "rawMarkdown": "&gt; Yikes, at 10 hrs/event it would take ~60 days to create a submission!!! That is a lot of compute 😁\n\nOutrunner says he is using a lot more time per event, up to 3 days for the largest ones, and it will probably show in the LB ;)\n\nThe only way to go is to use parallelism (see the recent discussion about how to do it in Python).  I have access to a 40 core IBM machine, which helps quite a bit ;)",
      "votes": null
    },
    {
      "id": "369005",
      "postDate": "08/11/2018 15:12:30",
      "content": "<blockquote>\n  <p>Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left.</p>\n  \n  <p>I think they are leading me to a new submission that should be right\n  around 0.7</p>\n</blockquote>\n\n<p><a href=\"/johnhsweeney\">@johnhsweeney</a> , </p>\n\n<p>thanks for the team name!  Seems you used our breadcrumbs effectively!</p>",
      "rawMarkdown": "&gt; Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left.\n&gt; \n&gt; I think they are leading me to a new submission that should be right\n&gt; around 0.7\n\n@johnhsweeney , \n\nthanks for the team name!  Seems you used our breadcrumbs effectively!",
      "votes": null
    },
    {
      "id": "369024",
      "postDate": "08/11/2018 16:40:11",
      "content": "<p>That's not the final results yet...</p>\n\n<p>That is ~2/3 of the updated results ensembled with my previous submission.  If the statistics hold, my final submission should be right around 0.71!</p>\n\n<p>I suspect I could improve further if I knew how to properly handle: </p>\n\n<ol>\n<li>Tracks that extend more than 180 degrees of arc</li>\n<li>Unroll Left and Right handed helices with a single scan (it currently costs me 2 dbscans per 1/R &amp; Z sample</li>\n<li>Probably more effects I haven't even noticed yet!</li>\n</ol>\n\n<p>Thanks again for your generous sharing!</p>",
      "rawMarkdown": "That's not the final results yet...\n\nThat is ~2/3 of the updated results ensembled with my previous submission.  If the statistics hold, my final submission should be right around 0.71!\n\nI suspect I could improve further if I knew how to properly handle: \n\n 1. Tracks that extend more than 180 degrees of arc\n 2. Unroll Left and Right handed helices with a single scan (it currently costs me 2 dbscans per 1/R &amp; Z sample\n 3. Probably more effects I haven't even noticed yet!\n\nThanks again for your generous sharing!",
      "votes": null
    },
    {
      "id": "369608",
      "postDate": "08/13/2018 12:51:28",
      "content": "<p>Seems I overfit a bit as my local score is 0.804 and my LB is a tad below 0.8, currently at 0.7996.</p>\n\n<p>I should have waited a bit before posting this ;)</p>\n\n<p>Nevertheless, I still hope to get over 0.8 in the LB today... even by a tiny margin.</p>",
      "rawMarkdown": "Seems I overfit a bit as my local score is 0.804 and my LB is a tad below 0.8, currently at 0.7996.\n\nI should have waited a bit before posting this ;)\n\nNevertheless, I still hope to get over 0.8 in the LB today... even by a tiny margin.",
      "votes": null
    },
    {
      "id": "369613",
      "postDate": "08/13/2018 12:57:02",
      "content": "<p>@CPMP,  detail, detail, you've proved what dbscan can do ;) pretty impressive. DBSCAN really is a powerful tool for seeking seeds fast and accurately. I never expected DBSCAN-ONLY could reach 0.8 with little post-processing. I'm sure CERN scientists know more exact equations and they could maximize the benefit of it.</p>",
      "rawMarkdown": "CPMP,  detail, detail, you've proved what dbscan can do ;) pretty impressive. DBSCAN really is a powerful tool for seeking seeds fast and accurately. I never expected DBSCAN-ONLY could reach 0.8 with little post-processing. I'm sure CERN scientists know more exact equations and they could maximize the benefit of it.",
      "votes": null
    },
    {
      "id": "369625",
      "postDate": "08/13/2018 13:28:47",
      "content": "<p>Thanks Nicole.</p>",
      "rawMarkdown": "Thanks Nicole.",
      "votes": null
    },
    {
      "id": "369693",
      "postDate": "08/13/2018 16:13:43",
      "content": "<p>Should be interesting to look the code and discuss after the competition ends.</p>",
      "rawMarkdown": "Should be interesting to look the code and discuss after the competition ends.",
      "votes": null
    },
    {
      "id": "369696",
      "postDate": "08/13/2018 16:14:30",
      "content": "<p>Very nice! </p>",
      "rawMarkdown": "Very nice!",
      "votes": null
    },
    {
      "id": "369739",
      "postDate": "08/13/2018 17:17:41",
      "content": "<p>Interesting that my LB score didn't change by even 0.0001 from the results using 2/3 of the new data vs 100%.  There is another thread speculating on which events are included in the public LB.  My experience seems to validate that the last third of the events are not used 😁</p>\n\n<p>On the other hand, my final score of 0.6956 is almost identical to my local validation score... Which for some reason makes me nervous!</p>",
      "rawMarkdown": "Interesting that my LB score didn't change by even 0.0001 from the results using 2/3 of the new data vs 100%.  There is another thread speculating on which events are included in the public LB.  My experience seems to validate that the last third of the events are not used 😁\n\nOn the other hand, my final score of 0.6956 is almost identical to my local validation score... Which for some reason makes me nervous!",
      "votes": null
    },
    {
      "id": "369772",
      "postDate": "08/13/2018 18:05:44",
      "content": "<blockquote>\n  <p>the last third of the events are not used 😁</p>\n</blockquote>\n\n<p>They are most probably used for the private LB score.</p>",
      "rawMarkdown": "&gt; the last third of the events are not used 😁\n\nThey are most probably used for the private LB score.",
      "votes": null
    },
    {
      "id": "369789",
      "postDate": "08/13/2018 18:40:17",
      "content": "<p>Congrats @CPMP. I believe you may get to 0.804 on the private LB. I am saying this because in my setup, I do prediction for each event in the test data and save each to an individual csv file. Then stitch all the 125 files  together afterwards for submission. While waiting for this process, as I am eager to see the progress, so I swap a few event's result in my previous submission with the new ones and submit to the LB. For example at one point substituting results of 7 new events moved my score from 0.5755 to 0.5759, so I am speaking from that experience. However, this swapping of files stops making improvement as from test #60. Obviously these improvements can be events dependent as well.</p>\n\n<p>Good luck everyone and happy Kaggling.</p>",
      "rawMarkdown": "Congrats @CPMP. I believe you may get to 0.804 on the private LB. I am saying this because in my setup, I do prediction for each event in the test data and save each to an individual csv file. Then stitch all the 125 files  together afterwards for submission. While waiting for this process, as I am eager to see the progress, so I swap a few event's result in my previous submission with the new ones and submit to the LB. For example at one point substituting results of 7 new events moved my score from 0.5755 to 0.5759, so I am speaking from that experience. However, this swapping of files stops making improvement as from test #60. Obviously these improvements can be events dependent as well.\n\nGood luck everyone and happy Kaggling.",
      "votes": null
    },
    {
      "id": "369801",
      "postDate": "08/13/2018 19:11:17",
      "content": "<p><a href=\"/johnhsweeney\">@johnhsweeney</a>, Are you sure you submitted? Your last sub is from 2 days ago on the LB.</p>",
      "rawMarkdown": "johnhsweeney, Are you sure you submitted? Your last sub is from 2 days ago on the LB.",
      "votes": null
    },
    {
      "id": "369811",
      "postDate": "08/13/2018 20:00:50",
      "content": "<p>Thanks for asking!  Yes, my final submission was on Saturday afternoon.</p>",
      "rawMarkdown": "Thanks for asking!  Yes, my final submission was on Saturday afternoon.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 367872,
      "author_name": "okhlopkov",
      "author_url": "",
      "post_date": "08/08/2018 18:03:59",
      "content": "<p>Thanks for sharing. I've read almost all of your posts and trying to build a proper DBSCAN pipeline, but I can't jump over 0.2 local score (shame!). I believe I've missed something very obvious. </p>\n\n<p>So my pipeline is:\n1) Select features for clustering (2-3 variables)\n2) Do helix unrolling on some angle and shift z somewhere\n3) Merge current clusters to the result ones\n4) Repeat 1-3 \n5) Remove outliers to 0-cluster\n6) Do track extension </p>",
      "votes": null,
      "replies": [
        {
          "id": 367875,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 18:14:50",
          "content": "<p>Hi, there are some public kernels with DBSCAN that have scores close to 0.5.  You may learn a bit from them.  Note that my 0.785 local score comes without any post processing, no outlier removal nor track extension.  Just to say that you can get some mileage out of clustering itself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 367937,
          "author_name": "posiedon",
          "author_url": "",
          "post_date": "08/08/2018 21:31:25",
          "content": "<p>How long did it take your DBSCAN to run?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368054,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/09/2018 06:35:51",
          "content": "<blockquote>\n  <p>How long did it take your DBSCAN to run?</p>\n</blockquote>\n\n<p>About 10 hours per event with an i7 at 4GHz.  Code can probably be made faster, but not much, unless I stop using pandas and undertake a massive rewrite.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368373,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/09/2018 19:33:16",
          "content": "<p>Yikes, at 10 hrs/event it would take ~60 days to create a submission!!!\nThat is a lot of compute 😁</p>\n\n<p>Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left. </p>\n\n<p>I think they are leading me to a new submission that should be right around 0.7, maybe a little under, since I need to keep the processing time short to make the Monday deadline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368404,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/09/2018 20:33:58",
          "content": "<p>@John, good job!!! :D  Look forward to your score! We feel so good at a silver medal position, that's where we should belong to, kind of like ducks like to play in a muddy pond. You know we suck at geometry... not sure how people can get a hold on it. I'm glad we decided to become software engineers who won't blow up a rocket because we suck at math. :D (though SpaceX would really be cool to work for)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368410,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/09/2018 20:43:55",
          "content": "<blockquote>\n  <p>We feel so good at a silver medal position</p>\n</blockquote>\n\n<p>By the way you can swap with Mickey in the private LB.  So your silver medal is in question, maybe gold. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368415,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/09/2018 21:15:49",
          "content": "<p>@Sergey lol lots of people haven't made submissions for a while including @Mickey and @Heng. I think this position is pretty comfy for us. We're only working on the post processing code right now. I look forward to their scores too!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368423,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/09/2018 21:34:56",
          "content": "<p>@Nicole, Thanks!  </p>\n\n<p>Fingers crossed, the cake is currently baking 🍰</p>\n\n<p>With an effective rate of 3 events/hour, my next submission should be sometime on Saturday!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368503,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2018 03:21:15",
          "content": "<blockquote>\n  <p>Yikes, at 10 hrs/event it would take ~60 days to create a submission!!! That is a lot of compute 😁</p>\n</blockquote>\n\n<p>Outrunner says he is using a lot more time per event, up to 3 days for the largest ones, and it will probably show in the LB ;)</p>\n\n<p>The only way to go is to use parallelism (see the recent discussion about how to do it in Python).  I have access to a 40 core IBM machine, which helps quite a bit ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369005,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/11/2018 15:12:30",
          "content": "<blockquote>\n  <p>Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left.</p>\n  \n  <p>I think they are leading me to a new submission that should be right\n  around 0.7</p>\n</blockquote>\n\n<p><a href=\"/johnhsweeney\">@johnhsweeney</a> , </p>\n\n<p>thanks for the team name!  Seems you used our breadcrumbs effectively!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369024,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/11/2018 16:40:11",
          "content": "<p>That's not the final results yet...</p>\n\n<p>That is ~2/3 of the updated results ensembled with my previous submission.  If the statistics hold, my final submission should be right around 0.71!</p>\n\n<p>I suspect I could improve further if I knew how to properly handle: </p>\n\n<ol>\n<li>Tracks that extend more than 180 degrees of arc</li>\n<li>Unroll Left and Right handed helices with a single scan (it currently costs me 2 dbscans per 1/R &amp; Z sample</li>\n<li>Probably more effects I haven't even noticed yet!</li>\n</ol>\n\n<p>Thanks again for your generous sharing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369739,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/13/2018 17:17:41",
          "content": "<p>Interesting that my LB score didn't change by even 0.0001 from the results using 2/3 of the new data vs 100%.  There is another thread speculating on which events are included in the public LB.  My experience seems to validate that the last third of the events are not used 😁</p>\n\n<p>On the other hand, my final score of 0.6956 is almost identical to my local validation score... Which for some reason makes me nervous!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369772,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/13/2018 18:05:44",
          "content": "<blockquote>\n  <p>the last third of the events are not used 😁</p>\n</blockquote>\n\n<p>They are most probably used for the private LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369801,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/13/2018 19:11:17",
          "content": "<p><a href=\"/johnhsweeney\">@johnhsweeney</a>, Are you sure you submitted? Your last sub is from 2 days ago on the LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369811,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "08/13/2018 20:00:50",
          "content": "<p>Thanks for asking!  Yes, my final submission was on Saturday afternoon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 367908,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/08/2018 19:59:49",
      "content": "<p>Cool! Will you share your code after the competition ends or will you hide it for the second phase?</p>",
      "votes": null,
      "replies": [
        {
          "id": 367918,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2018 20:23:52",
          "content": "<p>Hi, probably the latter, but I will share my code  anyway, in few days, or after 2nd phase, for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368072,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/09/2018 07:50:24",
          "content": "<p>@Sergey, though our code is not as good, @Liam and myself will share our code esp. the accurate helix unrolling function + accurate 2 features right after this competition  (actually @yuval said very clearly in his post what the 2nd feature is), e.g. 35 seconds to get 0.41 and one minute per event to reach 0.5 using dbscan on my laptop. we've never run an event too long, and we have good merging code, people can find some ideas from it for the 2nd phase <strong>performance</strong> competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368073,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/09/2018 07:53:53",
          "content": "<p>@Nicole \nGreat! It is interesting to look at your code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368396,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/09/2018 20:16:08",
          "content": "<p>@Nicole, I look forward to seeing your code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369696,
          "author_name": "muhakabartay",
          "author_url": "",
          "post_date": "08/13/2018 16:14:30",
          "content": "<p>Very nice! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 369608,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/13/2018 12:51:28",
      "content": "<p>Seems I overfit a bit as my local score is 0.804 and my LB is a tad below 0.8, currently at 0.7996.</p>\n\n<p>I should have waited a bit before posting this ;)</p>\n\n<p>Nevertheless, I still hope to get over 0.8 in the LB today... even by a tiny margin.</p>",
      "votes": null,
      "replies": [
        {
          "id": 369613,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/13/2018 12:57:02",
          "content": "<p>@CPMP,  detail, detail, you've proved what dbscan can do ;) pretty impressive. DBSCAN really is a powerful tool for seeking seeds fast and accurately. I never expected DBSCAN-ONLY could reach 0.8 with little post-processing. I'm sure CERN scientists know more exact equations and they could maximize the benefit of it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369625,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/13/2018 13:28:47",
          "content": "<p>Thanks Nicole.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369693,
          "author_name": "muhakabartay",
          "author_url": "",
          "post_date": "08/13/2018 16:13:43",
          "content": "<p>Should be interesting to look the code and discuss after the competition ends.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 369789,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/13/2018 18:40:17",
          "content": "<p>Congrats @CPMP. I believe you may get to 0.804 on the private LB. I am saying this because in my setup, I do prediction for each event in the test data and save each to an individual csv file. Then stitch all the 125 files  together afterwards for submission. While waiting for this process, as I am eager to see the progress, so I swap a few event's result in my previous submission with the new ones and submit to the LB. For example at one point substituting results of 7 new events moved my score from 0.5755 to 0.5759, so I am speaking from that experience. However, this swapping of files stops making improvement as from test #60. Obviously these improvements can be events dependent as well.</p>\n\n<p>Good luck everyone and happy Kaggling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "367726": "I just got over 0.8 local score using only DBSCAN and some post processing.  DBSCAN alone gets to approx 0.785.  I thought I'd share it because it contradicts what several of us (including me) wrote before, namely that supervised learning was required to move over 0.8.\n\nI would not be surprised now that some of the people with better score than 0.8 also get there without supervised learning.\n\nThe above is for a code that only looks for tracks originating from the z axis.  It finds over 92% of them.  If you do the math, then you'll see it catches some of the other tracks, but not much.",
    "367872": "Thanks for sharing. I've read almost all of your posts and trying to build a proper DBSCAN pipeline, but I can't jump over 0.2 local score (shame!). I believe I've missed something very obvious. \n\nSo my pipeline is:\n1) Select features for clustering (2-3 variables)\n2) Do helix unrolling on some angle and shift z somewhere\n3) Merge current clusters to the result ones\n4) Repeat 1-3 \n5) Remove outliers to 0-cluster\n6) Do track extension",
    "367875": "Hi, there are some public kernels with DBSCAN that have scores close to 0.5.  You may learn a bit from them.  Note that my 0.785 local score comes without any post processing, no outlier removal nor track extension.  Just to say that you can get some mileage out of clustering itself.",
    "367908": "Cool! Will you share your code after the competition ends or will you hide it for the second phase?",
    "367918": "Hi, probably the latter, but I will share my code  anyway, in few days, or after 2nd phase, for sure.",
    "367937": "How long did it take your DBSCAN to run?",
    "368054": "&gt;How long did it take your DBSCAN to run?\n\nAbout 10 hours per event with an i7 at 4GHz.  Code can probably be made faster, but not much, unless I stop using pandas and undertake a massive rewrite.",
    "368072": "Sergey, though our code is not as good, @Liam and myself will share our code esp. the accurate helix unrolling function + accurate 2 features right after this competition  (actually @yuval said very clearly in his post what the 2nd feature is), e.g. 35 seconds to get 0.41 and one minute per event to reach 0.5 using dbscan on my laptop. we've never run an event too long, and we have good merging code, people can find some ideas from it for the 2nd phase **performance** competition.",
    "368073": "Nicole \nGreat! It is interesting to look at your code.",
    "368373": "Yikes, at 10 hrs/event it would take ~60 days to create a submission!!!\nThat is a lot of compute 😁\n\nAnyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left. \n\nI think they are leading me to a new submission that should be right around 0.7, maybe a little under, since I need to keep the processing time short to make the Monday deadline.",
    "368396": "Nicole, I look forward to seeing your code.",
    "368404": "John, good job!!! :D  Look forward to your score! We feel so good at a silver medal position, that's where we should belong to, kind of like ducks like to play in a muddy pond. You know we suck at geometry... not sure how people can get a hold on it. I'm glad we decided to become software engineers who won't blow up a rocket because we suck at math. :D (though SpaceX would really be cool to work for)",
    "368410": "&gt; We feel so good at a silver medal position\n\nBy the way you can swap with Mickey in the private LB.  So your silver medal is in question, maybe gold. :)",
    "368415": "Sergey lol lots of people haven't made submissions for a while including @Mickey and @Heng. I think this position is pretty comfy for us. We're only working on the post processing code right now. I look forward to their scores too!",
    "368423": "Nicole, Thanks!  \n\nFingers crossed, the cake is currently baking 🍰\n\nWith an effective rate of 3 events/hour, my next submission should be sometime on Saturday!",
    "368503": "&gt; Yikes, at 10 hrs/event it would take ~60 days to create a submission!!! That is a lot of compute 😁\n\nOutrunner says he is using a lot more time per event, up to 3 days for the largest ones, and it will probably show in the LB ;)\n\nThe only way to go is to use parallelism (see the recent discussion about how to do it in Python).  I have access to a 40 core IBM machine, which helps quite a bit ;)",
    "369005": "&gt; Anyway, I'd like to say thanks for all the breadcrumbs you and Yuval have left.\n&gt; \n&gt; I think they are leading me to a new submission that should be right\n&gt; around 0.7\n\n@johnhsweeney , \n\nthanks for the team name!  Seems you used our breadcrumbs effectively!",
    "369024": "That's not the final results yet...\n\nThat is ~2/3 of the updated results ensembled with my previous submission.  If the statistics hold, my final submission should be right around 0.71!\n\nI suspect I could improve further if I knew how to properly handle: \n\n 1. Tracks that extend more than 180 degrees of arc\n 2. Unroll Left and Right handed helices with a single scan (it currently costs me 2 dbscans per 1/R &amp; Z sample\n 3. Probably more effects I haven't even noticed yet!\n\nThanks again for your generous sharing!",
    "369608": "Seems I overfit a bit as my local score is 0.804 and my LB is a tad below 0.8, currently at 0.7996.\n\nI should have waited a bit before posting this ;)\n\nNevertheless, I still hope to get over 0.8 in the LB today... even by a tiny margin.",
    "369613": "CPMP,  detail, detail, you've proved what dbscan can do ;) pretty impressive. DBSCAN really is a powerful tool for seeking seeds fast and accurately. I never expected DBSCAN-ONLY could reach 0.8 with little post-processing. I'm sure CERN scientists know more exact equations and they could maximize the benefit of it.",
    "369625": "Thanks Nicole.",
    "369693": "Should be interesting to look the code and discuss after the competition ends.",
    "369696": "Very nice!",
    "369739": "Interesting that my LB score didn't change by even 0.0001 from the results using 2/3 of the new data vs 100%.  There is another thread speculating on which events are included in the public LB.  My experience seems to validate that the last third of the events are not used 😁\n\nOn the other hand, my final score of 0.6956 is almost identical to my local validation score... Which for some reason makes me nervous!",
    "369772": "&gt; the last third of the events are not used 😁\n\nThey are most probably used for the private LB score.",
    "369789": "Congrats @CPMP. I believe you may get to 0.804 on the private LB. I am saying this because in my setup, I do prediction for each event in the test data and save each to an individual csv file. Then stitch all the 125 files  together afterwards for submission. While waiting for this process, as I am eager to see the progress, so I swap a few event's result in my previous submission with the new ones and submit to the LB. For example at one point substituting results of 7 new events moved my score from 0.5755 to 0.5759, so I am speaking from that experience. However, this swapping of files stops making improvement as from test #60. Obviously these improvements can be events dependent as well.\n\nGood luck everyone and happy Kaggling.",
    "369801": "johnhsweeney, Are you sure you submitted? Your last sub is from 2 days ago on the LB.",
    "369811": "Thanks for asking!  Yes, my final submission was on Saturday afternoon."
  },
  "source": "meta"
}