{
  "id": 58915,
  "title": "Two races",
  "url": "/competitions/trackml-particle-identification/discussion/58915",
  "author_name": "CPMP",
  "post_date": "2018-06-15T09:55:43.687000",
  "votes": 8,
  "comment_count": 54,
  "views": 0,
  "content": "<p>There are two races in this competition.</p>\n\n<ol>\n<li>Race to the top of LB</li>\n<li>Be the first to submit Grzegorz kernels output.  Seems some folks have scripts that alert them when a good kernel is shared given how fast they react ;)</li>\n</ol>",
  "messages": [
    {
      "id": 343427,
      "postDate": "2018-06-15T09:55:43.687Z",
      "content": "<p>There are two races in this competition.</p>\n\n<ol>\n<li>Race to the top of LB</li>\n<li>Be the first to submit Grzegorz kernels output.  Seems some folks have scripts that alert them when a good kernel is shared given how fast they react ;)</li>\n</ol>",
      "rawMarkdown": "There are two races in this competition.\n\n 1. Race to the top of LB\n 2. Be the first to submit Grzegorz kernels output.  Seems some folks have scripts that alert them when a good kernel is shared given how fast they react ;)",
      "votes": 8
    },
    {
      "id": 345210,
      "postDate": "2018-06-19T12:10:08.663Z",
      "content": "<p>After the shock of the score 0.8, maybe it is time for sharing ideas? I can prepare a kernel which I hope can satisfy majority of the Kagglers active in this competition. I would like to share a more advanced method of ensembling the tracks in the loop of DBSCAN and models with shifted z. I can write it in such a way that it will give no submission output and it will take so much time, that it will be not possible to obtain the full results at Kaggle servers. The score will be 0.54-0.58. As in my first kernel, there will be much place for optimization and your own ideas. What do you think?</p>\n\n<p>Of course not today - today Polish team plays football ;)</p>",
      "rawMarkdown": "After the shock of the score 0.8, maybe it is time for sharing ideas? I can prepare a kernel which I hope can satisfy majority of the Kagglers active in this competition. I would like to share a more advanced method of ensembling the tracks in the loop of DBSCAN and models with shifted z. I can write it in such a way that it will give no submission output and it will take so much time, that it will be not possible to obtain the full results at Kaggle servers. The score will be 0.54-0.58. As in my first kernel, there will be much place for optimization and your own ideas. What do you think?\n\nOf course not today - today Polish team plays football ;)",
      "votes": 5,
      "replies": [
        {
          "id": 345213,
          "postDate": "2018-06-19T12:13:10.963Z",
          "content": "<p>i  second your idea. I believe kagglers can push the score from 0.58 to 0.70</p>",
          "rawMarkdown": "i  second your idea. I believe kagglers can push the score from 0.58 to 0.70",
          "votes": 1
        },
        {
          "id": 345224,
          "postDate": "2018-06-19T12:35:09.477Z",
          "content": "<p>@Grzegorz, Good luck! I hope the Polish team will perform better than the German team. I would only share the idea instead of a kernel so every of us needs to work hard to win a medal. </p>",
          "rawMarkdown": "@Grzegorz, Good luck! I hope the Polish team will perform better than the German team. I would only share the idea instead of a kernel so every of us needs to work hard to win a medal. ",
          "votes": 5
        },
        {
          "id": 345226,
          "postDate": "2018-06-19T12:43:19.037Z",
          "content": "<p><a href=\"/grzegorz\">@grzegorz</a>, I like <a href=\"/nicole\">@nicole</a>'s idea.  Share the idea or even pseudo code and make us work for the benefit!</p>\n\n<p><a href=\"/nicole\">@nicole</a> - I'm guessing that you have created some analytics to gain insight into the \"truth\" ... That would be a welcome kernel too!</p>\n\n<p>PS - I am greatly conflicted by the World Cup.  I grew up in South Texas, so I have a natural affinity for Mexico, my wife is a German citizen, and the primary engineering team I work with is in Gdansk!  Naturally, rooting for team USA is, sadly, not possible!</p>",
          "rawMarkdown": "@grzegorz, I like @nicole's idea.  Share the idea or even pseudo code and make us work for the benefit!\n\n@nicole - I'm guessing that you have created some analytics to gain insight into the \"truth\" ... That would be a welcome kernel too!\n\nPS - I am greatly conflicted by the World Cup.  I grew up in South Texas, so I have a natural affinity for Mexico, my wife is a German citizen, and the primary engineering team I work with is in Gdansk!  Naturally, rooting for team USA is, sadly, not possible!",
          "votes": 3
        },
        {
          "id": 345228,
          "postDate": "2018-06-19T12:46:18.367Z",
          "content": "<p>@Grzegorz, </p>\n\n<p>this probably contradicts what you wrote yesterday:</p>\n\n<blockquote>\n  <p>I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.</p>\n</blockquote>\n\n<p>Note also that the organizers decided to run a Kaggle competition and not a collaborative effort.</p>\n\n<p>My take is that ideas sharing looks more interesting.  For instance, my remark on the origin of tracks and the following discussion this morning probably enlightened many on the need to use various z values at origin.  People who want to share more can team.</p>\n\n<p>This said, I am in no position to dictate anything here.  </p>\n\n<p>Let me conclude by thanking you for asking before sharing.  I'm curious to see what others will say, do they agree with you and Heng, or do they think people should find their own way?  </p>\n\n<p>Edit: i wrote this before seeing Nicole and John's answers.</p>",
          "rawMarkdown": "@Grzegorz, \n\nthis probably contradicts what you wrote yesterday:\n\n&gt; I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.\n\nNote also that the organizers decided to run a Kaggle competition and not a collaborative effort.\n\nMy take is that ideas sharing looks more interesting.  For instance, my remark on the origin of tracks and the following discussion this morning probably enlightened many on the need to use various z values at origin.  People who want to share more can team.\n\nThis said, I am in no position to dictate anything here.  \n\nLet me conclude by thanking you for asking before sharing.  I'm curious to see what others will say, do they agree with you and Heng, or do they think people should find their own way?  \n\nEdit: i wrote this before seeing Nicole and John's answers.",
          "votes": 4
        },
        {
          "id": 345249,
          "postDate": "2018-06-19T13:53:20.310Z",
          "content": "<p>Personally, I am already going to start working on modified origins because of this conversation, however if new ideas are shared, by either concept or kernel, I'd be excited to see them. I definitely felt like I learned more about how the code worked when I recreated HCK's extension code in R, and would be more likely to improve on the code when I had to do the legwork. Alternatively, if a kernel was submitted, more work would be put into other parts of the competition, and the community might push harder towards  .9x. </p>\n\n<p>I don't think submitting a new kernel has much risk of people not having to work for the medals. There is still so much time left in the competition, there will be more entrants able to beat the kernel outputs as we get closer to the deadline.</p>\n\n<p>Really, my biggest worry would be if multiple ideas start getting shared, everyone's ideas will contribute to the computational barrier, ex: user 1 shares an idea that causes good solutions to run track extension twice as long, and user 2 shares an idea that causes good solutions to run track extension twice as often.</p>",
          "rawMarkdown": "Personally, I am already going to start working on modified origins because of this conversation, however if new ideas are shared, by either concept or kernel, I'd be excited to see them. I definitely felt like I learned more about how the code worked when I recreated HCK's extension code in R, and would be more likely to improve on the code when I had to do the legwork. Alternatively, if a kernel was submitted, more work would be put into other parts of the competition, and the community might push harder towards  .9x. \n\nI don't think submitting a new kernel has much risk of people not having to work for the medals. There is still so much time left in the competition, there will be more entrants able to beat the kernel outputs as we get closer to the deadline.\n\nReally, my biggest worry would be if multiple ideas start getting shared, everyone's ideas will contribute to the computational barrier, ex: user 1 shares an idea that causes good solutions to run track extension twice as long, and user 2 shares an idea that causes good solutions to run track extension twice as often.",
          "votes": 1
        },
        {
          "id": 345252,
          "postDate": "2018-06-19T14:04:05.663Z",
          "content": "<p>@John Sweeney...ICELAND!!  Who doesn't like an underdog?</p>",
          "rawMarkdown": "@John Sweeney...ICELAND!!  Who doesn't like an underdog?",
          "votes": 1
        },
        {
          "id": 345382,
          "postDate": "2018-06-19T19:40:51.583Z",
          "content": "<p>The word \"idea\" has at least two meanings. Let me explain it on two examples:</p>\n\n<p>1) \"Maybe it is a good idea to shift our models along z axis or maybe it is not at all, because the most of tracks come from (0,0,0)\"</p>\n\n<p>2) \"It is a good idea to shift models along z axis, because ensembling 7 models gives a boost 0.08 and ensembling 21 models gives a boost 0.105, however the time of calculations is extremely long\".</p>\n\n<p>My thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. Talking about satisfying majority of active Kagglers I thought also about myself. I am not sure if it fully satisfies me to share my checked solutions in the form of few sentences here and there inside someone else's threads. So, mentioned kind of ensembling the tracks will be presented by me  in the form of kernel, when Top100 will exceed 0.58, i.e. when it will be completely useless. You lose nothing - this ensembling of DBSCAN clusters plays similar role and gives the same boost as Heng's extend function, and the effect of ensembling models shifted along z axis was described by me here and there. Best wishes.</p>",
          "rawMarkdown": "The word \"idea\" has at least two meanings. Let me explain it on two examples:\n\n1) \"Maybe it is a good idea to shift our models along z axis or maybe it is not at all, because the most of tracks come from (0,0,0)\"\n\n2) \"It is a good idea to shift models along z axis, because ensembling 7 models gives a boost 0.08 and ensembling 21 models gives a boost 0.105, however the time of calculations is extremely long\".\n\nMy thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. Talking about satisfying majority of active Kagglers I thought also about myself. I am not sure if it fully satisfies me to share my checked solutions in the form of few sentences here and there inside someone else's threads. So, mentioned kind of ensembling the tracks will be presented by me  in the form of kernel, when Top100 will exceed 0.58, i.e. when it will be completely useless. You lose nothing - this ensembling of DBSCAN clusters plays similar role and gives the same boost as Heng's extend function, and the effect of ensembling models shifted along z axis was described by me here and there. Best wishes.",
          "votes": 3
        },
        {
          "id": 345420,
          "postDate": "2018-06-19T21:03:10.867Z",
          "content": "<p>&gt;  I hope the Polish team will perform better than the German team. </p>\n\n<p>Unfortunately the Polish team has lost.\nI also watched my native team (Russia) today and I'm happy we won. (Probably we will lose in the first game at playoff, but anyway).</p>\n\n<p>I like sharing ideas but in the form of algorithms. \"shift models along z axis, because ensembling 7 models gives a boost 0.08\" &lt;--- That's enough for me.</p>",
          "rawMarkdown": "&gt;  I hope the Polish team will perform better than the German team. \n\nUnfortunately the Polish team has lost.\nI also watched my native team (Russia) today and I'm happy we won. (Probably we will lose in the first game at playoff, but anyway).\n\nI like sharing ideas but in the form of algorithms. \"shift models along z axis, because ensembling 7 models gives a boost 0.08\" &lt;--- That's enough for me.",
          "votes": 2
        },
        {
          "id": 345427,
          "postDate": "2018-06-19T21:21:15.450Z",
          "content": "<p>I agree with Sergey.  It is enough of a hint.  With  this level of detail, we have a good direction, but may stumble into something else good we would never find with a directly usable code snippet</p>",
          "rawMarkdown": "I agree with Sergey.  It is enough of a hint.  With  this level of detail, we have a good direction, but may stumble into something else good we would never find with a directly usable code snippet",
          "votes": 1
        },
        {
          "id": 345433,
          "postDate": "2018-06-19T21:48:13.240Z",
          "content": "<p>@Grzegorz, just let us have some fun exploring hidden features ourselves :)  as @John said, we may stumble into something new and better, you've helped us a lot along our journey, highly appreciated! Ensembling 21 models is crazy, but with a 0.1 boost, that's huge. We're reworking on the ensembling code too, since it didn't give us much of a benefit anymore after our single models got much stronger. You're right, I didn't ensemble models with shifted z features, it's enough to know this is an area we can still explore. This competition is very special, not a traditional ML/DL competition but more of a research / brain damaging / sleep depriving competition. :D </p>",
          "rawMarkdown": "@Grzegorz, just let us have some fun exploring hidden features ourselves :)  as @John said, we may stumble into something new and better, you've helped us a lot along our journey, highly appreciated! Ensembling 21 models is crazy, but with a 0.1 boost, that's huge. We're reworking on the ensembling code too, since it didn't give us much of a benefit anymore after our single models got much stronger. You're right, I didn't ensemble models with shifted z features, it's enough to know this is an area we can still explore. This competition is very special, not a traditional ML/DL competition but more of a research / brain damaging / sleep depriving competition. :D ",
          "votes": 2
        },
        {
          "id": 345436,
          "postDate": "2018-06-19T21:58:17.280Z",
          "content": "<p>@Nicole, I don't understand why this competition is so engaging to me, as someone who works in mortgage you'd think I'd be working on the mortgage competition, but I keep coming back here. I've been staying up late working on this so often I've somehow fallen into a biphasic sleep cycle. </p>\n\n<p>@Grzegorz, you do what you want to do, we'll all be interested to see what happens next. </p>",
          "rawMarkdown": "@Nicole, I don't understand why this competition is so engaging to me, as someone who works in mortgage you'd think I'd be working on the mortgage competition, but I keep coming back here. I've been staying up late working on this so often I've somehow fallen into a biphasic sleep cycle. \n\n@Grzegorz, you do what you want to do, we'll all be interested to see what happens next. ",
          "votes": 2
        },
        {
          "id": 345448,
          "postDate": "2018-06-19T22:50:36.387Z",
          "content": "<blockquote>\n  <p>My thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. </p>\n</blockquote>\n\n<p>@Grzegorz, Mine as well, but I rather stick to ideas instead of reusable code.  I fully agree with what @SIlogram <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/58332\">expressed in another ongoing competition</a></p>\n\n<blockquote>\n  <p>First, let me say that I'm not a big fan of the kernels and miss the days when Kagglers would provide general guidelines and hints rather than full-blown solutions</p>\n</blockquote>\n\n<p>And I think you understand now why computing my solution takes time:</p>\n\n<blockquote>\n  <p>however the time of calculations is extremely long</p>\n</blockquote>",
          "rawMarkdown": "&gt; My thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. \n\n@Grzegorz, Mine as well, but I rather stick to ideas instead of reusable code.  I fully agree with what @SIlogram [expressed in another ongoing competition][1]\n \n&gt; First, let me say that I'm not a big fan of the kernels and miss the days when Kagglers would provide general guidelines and hints rather than full-blown solutions\n\nAnd I think you understand now why computing my solution takes time:\n\n&gt; however the time of calculations is extremely long\n\n  [1]: https://www.kaggle.com/c/home-credit-default-risk/discussion/58332",
          "votes": 1
        },
        {
          "id": 345467,
          "postDate": "2018-06-19T23:52:05.107Z",
          "content": "<p>Thank you all for your posts. I would like to be infected with optimism from you.\nI think I'm the world champion in skepticism. Already at the beginning of the competition, I assumed that it is impossible to win by someone who starts from scratch, regardless of what ideas he or she will have. That's why I did not pay too much attention to my ideas. The same ideas invented by you were treated very seriously.</p>\n\n<p>Regarding the winner, I thought that somebody would simply read some wise publication about separation of tracks, implement an algorithm being a version of the Kalman filter or other combinatorial algorithm and that's it. The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.</p>",
          "rawMarkdown": "Thank you all for your posts. I would like to be infected with optimism from you.\nI think I'm the world champion in skepticism. Already at the beginning of the competition, I assumed that it is impossible to win by someone who starts from scratch, regardless of what ideas he or she will have. That's why I did not pay too much attention to my ideas. The same ideas invented by you were treated very seriously.\n\nRegarding the winner, I thought that somebody would simply read some wise publication about separation of tracks, implement an algorithm being a version of the Kalman filter or other combinatorial algorithm and that's it. The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.",
          "votes": 4
        },
        {
          "id": 345642,
          "postDate": "2018-06-20T07:31:54.503Z",
          "content": "<blockquote>\n  <p>The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.</p>\n</blockquote>\n\n<p>You may be right for this, but this is not our problem, is it?  </p>\n\n<p>CERN say in one of their presentation (probably link shared by Heng) that they run this competition with the hope that someone will get good results use a very different technique than the ones they have tried.  </p>",
          "rawMarkdown": "&gt; The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.\n\nYou may be right for this, but this is not our problem, is it?  \n\nCERN say in one of their presentation (probably link shared by Heng) that they run this competition with the hope that someone will get good results use a very different technique than the ones they have tried.  "
        },
        {
          "id": 345741,
          "postDate": "2018-06-20T11:21:55.117Z",
          "content": "<p>It might be that, although in terms of accuracy, the winning result will not be significantly better, in terms of algorithm efficiency, will be better. Well, we don't actually know. Let's see what will happen. </p>",
          "rawMarkdown": "It might be that, although in terms of accuracy, the winning result will not be significantly better, in terms of algorithm efficiency, will be better. Well, we don't actually know. Let's see what will happen. "
        },
        {
          "id": 345767,
          "postDate": "2018-06-20T12:20:17.357Z",
          "content": "<p>To be absolutely clear : the CERN experiments are directly and deeply interested in the results of the challenge, because track reconstruction is a critical problem, see details in <a href=\"https://www.nature.com/articles/d41586-018-05084-2\">https://www.nature.com/articles/d41586-018-05084-2</a>\nIndeed our objective of this kaggle competition is to discover new algorithms. Note that, as indicated in \"Welcome from the organiser\" post, we will launch the second phase of the competition in July (so overlapping with the Kaggle one), with its own set of prizes, where participants will have to submit code, and the CPU time of the evaluation will be evaluated, and participants will be ranked with a score combining the accuracy (== as computed in this first phase) and CPU). More on this soon.</p>",
          "rawMarkdown": "To be absolutely clear : the CERN experiments are directly and deeply interested in the results of the challenge, because track reconstruction is a critical problem, see details in https://www.nature.com/articles/d41586-018-05084-2\nIndeed our objective of this kaggle competition is to discover new algorithms. Note that, as indicated in \"Welcome from the organiser\" post, we will launch the second phase of the competition in July (so overlapping with the Kaggle one), with its own set of prizes, where participants will have to submit code, and the CPU time of the evaluation will be evaluated, and participants will be ranked with a score combining the accuracy (== as computed in this first phase) and CPU). More on this soon.",
          "votes": 3
        },
        {
          "id": 345831,
          "postDate": "2018-06-20T14:49:03.383Z",
          "content": "<p>@David, It is still 2 months to go. Consider extending the set of additional prizes for more symbolic ones. There are rich countries of high salaries and low taxes, e.g. Switzerland, and less rich countries of low salaries and high taxes, e.g. Poland. Even if I create a genius code I will not apply for Tesla V100, because it would be connected with taking credit equal my 5-6 salaries to pay the tax for a material award from USA (2 if from EU).</p>",
          "rawMarkdown": "@David, It is still 2 months to go. Consider extending the set of additional prizes for more symbolic ones. There are rich countries of high salaries and low taxes, e.g. Switzerland, and less rich countries of low salaries and high taxes, e.g. Poland. Even if I create a genius code I will not apply for Tesla V100, because it would be connected with taking credit equal my 5-6 salaries to pay the tax for a material award from USA (2 if from EU).",
          "votes": 2
        },
        {
          "id": 345910,
          "postDate": "2018-06-20T17:45:12.250Z",
          "content": "<p>Grzegorz you deserve (with Heng) a prize for sharing so much.  I hope you understand that nobody has issues with your sharing.  But some, including me, have issues with people merely submitting your kernels output.</p>",
          "rawMarkdown": "Grzegorz you deserve (with Heng) a prize for sharing so much.  I hope you understand that nobody has issues with your sharing.  But some, including me, have issues with people merely submitting your kernels output.",
          "votes": 2
        },
        {
          "id": 346021,
          "postDate": "2018-06-20T23:33:12.027Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 346025,
          "postDate": "2018-06-20T23:42:10.440Z",
          "content": "<p>A week ago I would have guessed that the top 3 scores came from something more advanced than improved clustering.  Now, as CPMP continues to climb towards 0.7, I think maybe everyone is still clustering...</p>\n\n<p>Even 0.8 through clustering?  Why not?</p>\n\n<p>Will a supervised algorithm win or not?</p>",
          "rawMarkdown": "A week ago I would have guessed that the top 3 scores came from something more advanced than improved clustering.  Now, as CPMP continues to climb towards 0.7, I think maybe everyone is still clustering...\n\nEven 0.8 through clustering?  Why not?\n\nWill a supervised algorithm win or not?",
          "votes": 2
        },
        {
          "id": 346207,
          "postDate": "2018-06-21T09:24:19.457Z",
          "content": "<p>@John my 2 cents again, I don't think a pure supervised learning can win this competition, but I guess you mean if a combined solution would. A DL model helps to predict the possibility of hits but at the end of the day you still have to do track fitting (post processing) as what you would do with unsupervised learning. @Heng said his LSTM is better than his backfitting code, worth giving it a shot.  I'm very eager to know if <a href=\"/outrunner\">@outrunner</a> used a supervised learning approach at all, but I doubt she would see my comment. :)</p>",
          "rawMarkdown": "@John my 2 cents again, I don't think a pure supervised learning can win this competition, but I guess you mean if a combined solution would. A DL model helps to predict the possibility of hits but at the end of the day you still have to do track fitting (post processing) as what you would do with unsupervised learning. @Heng said his LSTM is better than his backfitting code, worth giving it a shot.  I'm very eager to know if @outrunner used a supervised learning approach at all, but I doubt she would see my comment. :)",
          "votes": 3
        },
        {
          "id": 346220,
          "postDate": "2018-06-21T10:06:29.020Z",
          "content": "<p>@John, I think I can push my current approach (clustering plus few secret sauce ingredients) above 0.7, but I doubt it can reach 0.8, one reason being is that it focus on tracks starting near the central z axis.  Anyway, I'm pursuing my current track (pun intended ;) ) as it helps me get a better understanding of what we deal with.</p>",
          "rawMarkdown": "@John, I think I can push my current approach (clustering plus few secret sauce ingredients) above 0.7, but I doubt it can reach 0.8, one reason being is that it focus on tracks starting near the central z axis.  Anyway, I'm pursuing my current track (pun intended ;) ) as it helps me get a better understanding of what we deal with.",
          "votes": 1
        },
        {
          "id": 346267,
          "postDate": "2018-06-21T11:46:36.303Z",
          "content": "<p>The score limit if only the particles coming from 0,0,z are under consideration is about 0.82.</p>\n\n<p><a href=\"https://www.kaggle.com/sionek/score-limit-for-0-0-z\">https://www.kaggle.com/sionek/score-limit-for-0-0-z</a></p>",
          "rawMarkdown": "The score limit if only the particles coming from 0,0,z are under consideration is about 0.82.\n\nhttps://www.kaggle.com/sionek/score-limit-for-0-0-z",
          "votes": 6
        },
        {
          "id": 346273,
          "postDate": "2018-06-21T11:55:10.007Z",
          "content": "<p>Right, I have been trying to publish a notebook about it for the last 2 hours without success, Kaggle notebooks are broken now.  Anyway, even if you enlarge the region to 16 mm around the z axis, score is still below 0.85.  That's why I wrote above that identifying tracks that do not originate from origin will be key to winning this competition.</p>",
          "rawMarkdown": "Right, I have been trying to publish a notebook about it for the last 2 hours without success, Kaggle notebooks are broken now.  Anyway, even if you enlarge the region to 16 mm around the z axis, score is still below 0.85.  That's why I wrote above that identifying tracks that do not originate from origin will be key to winning this competition."
        },
        {
          "id": 346276,
          "postDate": "2018-06-21T12:04:10.483Z",
          "content": "<blockquote>\n  <p>identifying tracks that do not originate from origin will be key to winning this competition.</p>\n</blockquote>\n\n<p>Just now we must learn to identify tracks that do originate from origin. To get a score &gt;0.8. :)</p>",
          "rawMarkdown": "&gt; identifying tracks that do not originate from origin will be key to winning this competition.\n\nJust now we must learn to identify tracks that do originate from origin. To get a score &gt;0.8. :)",
          "votes": 1
        },
        {
          "id": 346290,
          "postDate": "2018-06-21T12:18:52.617Z",
          "content": "<p>Interesting! \nI tried a specific method to catch those tracks (something similar to the z shifting, but in x/y directions) : there is a score of 0.18 to catch , and I just got 0.005. And the calculation is very very long... </p>",
          "rawMarkdown": "Interesting! \nI tried a specific method to catch those tracks (something similar to the z shifting, but in x/y directions) : there is a score of 0.18 to catch , and I just got 0.005. And the calculation is very very long... ",
          "votes": 3
        },
        {
          "id": 346294,
          "postDate": "2018-06-21T12:31:10.107Z",
          "content": "<p>@Zidmie, this is the real challenge indeed.</p>",
          "rawMarkdown": "@Zidmie, this is the real challenge indeed."
        },
        {
          "id": 346444,
          "postDate": "2018-06-21T17:50:40.193Z",
          "content": "<p>I've been traveling all week and haven't been able to try any of the excellent \"ideas\" shared here.</p>\n\n<p>I guess I need to do some visualization, but I fundamentally don't understand why particles originating far from x,y = 0,0 are more difficult to find.</p>\n\n<p>Any moving charged particle in the magnetic field will follow a helical path around the axis of the field (x,y = 0,0).</p>\n\n<p>It's path plotted in r,z coordinates will be approximately linear z = mr + z0, (z-z0)/r =  const</p>\n\n<p>I guess searching for appropriate values of z0 can be expensive.</p>\n\n<p>Is my logic flawed?</p>",
          "rawMarkdown": "I've been traveling all week and haven't been able to try any of the excellent \"ideas\" shared here.\n\nI guess I need to do some visualization, but I fundamentally don't understand why particles originating far from x,y = 0,0 are more difficult to find.\n\nAny moving charged particle in the magnetic field will follow a helical path around the axis of the field (x,y = 0,0).\n\nIt's path plotted in r,z coordinates will be approximately linear z = mr + z0, (z-z0)/r =  const\n\nI guess searching for appropriate values of z0 can be expensive.\n\nIs my logic flawed?"
        },
        {
          "id": 346451,
          "postDate": "2018-06-21T18:00:52.107Z",
          "content": "<p>Expanding on the maximum possible score:</p>\n\n<blockquote>\n  <p>using x_window:  500 , y_window:  500 , z_window:  2900 : score\n  0.9973638 , no. of hits: 83.94644 %</p>\n  \n  <p>using x_window:  300 , y_window:  300 , z_window:  1000 : score\n  0.9829006 , no. of hits: 82.78694 %</p>\n  \n  <p>using x_window:  100 , y_window:  100 , z_window:  1000 : score\n  0.9450124 , no. of hits: 80.22416 %</p>\n  \n  <p>using x_window:  100 , y_window:  100 , z_window:  500 : score\n  0.9259655 , no. of hits: 78.79893 %</p>\n  \n  <p>using x_window:  20 , y_window:  20 , z_window:  200 : score 0.8695067\n  , no. of hits: 73.77835 %</p>\n  \n  <p>using x_window:  20 , y_window:  20 , z_window:  20 : score 0.8497927\n  , no. of hits: 71.76945 %</p>\n  \n  <p>using x_window:  0.05 , y_window:  0.05 , z_window:  20 : score\n  0.8198518 , no. of hits: 69.59112 %</p>\n</blockquote>\n\n<p>From the fork I made of Grzegorz Sionkowski's (The Chemist) kernel</p>",
          "rawMarkdown": "Expanding on the maximum possible score:\n\n&gt; using x_window:  500 , y_window:  500 , z_window:  2900 : score\n&gt; 0.9973638 , no. of hits: 83.94644 %\n&gt; \n&gt; \n&gt; using x_window:  300 , y_window:  300 , z_window:  1000 : score\n&gt; 0.9829006 , no. of hits: 82.78694 %\n&gt; \n&gt; \n&gt; using x_window:  100 , y_window:  100 , z_window:  1000 : score\n&gt; 0.9450124 , no. of hits: 80.22416 %\n&gt; \n&gt; \n&gt; using x_window:  100 , y_window:  100 , z_window:  500 : score\n&gt; 0.9259655 , no. of hits: 78.79893 %\n&gt; \n&gt; \n&gt; using x_window:  20 , y_window:  20 , z_window:  200 : score 0.8695067\n&gt; , no. of hits: 73.77835 %\n&gt; \n&gt; \n&gt; using x_window:  20 , y_window:  20 , z_window:  20 : score 0.8497927\n&gt; , no. of hits: 71.76945 %\n&gt; \n&gt; \n&gt; using x_window:  0.05 , y_window:  0.05 , z_window:  20 : score\n&gt; 0.8198518 , no. of hits: 69.59112 %\n\nFrom the fork I made of Grzegorz Sionkowski's (The Chemist) kernel",
          "votes": 2
        },
        {
          "id": 346531,
          "postDate": "2018-06-21T22:31:10.817Z",
          "content": "<p>To John Sweeney:  if a particle is coming from far from x,y=0,0 then its helicoidal trajectory will appear to miss the axis 0.0.\nSee a slightly displaced vertex in CMS  <img src=\"https://cds.cern.ch/record/2280025/files/QCD-f35-53384671-RhoPhi-white-ann.png?subformat=icon-1440\" alt=\"Slightly displaced vertex in CMS\"> (all light green are tracks coming from the origin, the darker one are coming from a displaced vertex).\nand a very displaced vertex in ATLAS <img src=\"https://cds.cern.ch/record/1381516/files/vp1_run165821_evt1605517_rpv.png\" alt=\"Very displaced vertex in ATLAS\"> (missing by a larger distance the origin).</p>",
          "rawMarkdown": "To John Sweeney:  if a particle is coming from far from x,y=0,0 then its helicoidal trajectory will appear to miss the axis 0.0.\nSee a slightly displaced vertex in CMS  ![Slightly displaced vertex in CMS][1] (all light green are tracks coming from the origin, the darker one are coming from a displaced vertex).\nand a very displaced vertex in ATLAS ![Very displaced vertex in ATLAS][2] (missing by a larger distance the origin).\n\n\n  \n\n\n  [1]: https://cds.cern.ch/record/2280025/files/QCD-f35-53384671-RhoPhi-white-ann.png?subformat=icon-1440\n  [2]: https://cds.cern.ch/record/1381516/files/vp1_run165821_evt1605517_rpv.png\n",
          "votes": 2
        },
        {
          "id": 346645,
          "postDate": "2018-06-22T05:20:29.557Z",
          "content": "<p>@John, helix unrolling + clustering code shared in kernels assume that x,y = (0,0) is part of the track.  It does not detect tracks that start far from the origin.</p>",
          "rawMarkdown": "@John, helix unrolling + clustering code shared in kernels assume that x,y = (0,0) is part of the track.  It does not detect tracks that start far from the origin."
        },
        {
          "id": 346672,
          "postDate": "2018-06-22T06:44:23.170Z",
          "content": "<p>@John, Additionally, shared kernels assign only a part of helix to the track - this one while the distance from x,y = (0,0) increases. A part of helix after 0.5 turn (when this distance starts to decrease) is lost or treated as a track of other particle moving along the same helix but in the opposite direction.</p>",
          "rawMarkdown": "@John, Additionally, shared kernels assign only a part of helix to the track - this one while the distance from x,y = (0,0) increases. A part of helix after 0.5 turn (when this distance starts to decrease) is lost or treated as a track of other particle moving along the same helix but in the opposite direction.",
          "votes": 1
        },
        {
          "id": 347078,
          "postDate": "2018-06-23T07:23:05.590Z",
          "content": "<p>@Grzegorz Sionkowski\nThanks for the R kernel, It gave me a big push to code to this competition in R as well. I converted the R code of the 0.456 score script to python and then you made a better one. Now I'll just use your R, with hopefully some improvements :)</p>",
          "rawMarkdown": "@Grzegorz Sionkowski\nThanks for the R kernel, It gave me a big push to code to this competition in R as well. I converted the R code of the 0.456 score script to python and then you made a better one. Now I'll just use your R, with hopefully some improvements :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 344409,
      "postDate": "2018-06-17T23:40:17.567Z",
      "content": "<p>Side note: when your new submission finally scores better that the kernel output, it can be satisfying to pass the big chunk of submissions. </p>",
      "rawMarkdown": "Side note: when your new submission finally scores better that the kernel output, it can be satisfying to pass the big chunk of submissions. ",
      "votes": 4
    },
    {
      "id": 345933,
      "postDate": "2018-06-20T18:45:39.170Z",
      "content": "<p>A few ideas - <a href=\"https://github.com/HEPTrkX\">HEPTrkX</a></p>",
      "rawMarkdown": "A few ideas - [HEPTrkX][1]\n\n\n  [1]: https://github.com/HEPTrkX",
      "votes": 1
    },
    {
      "id": 343911,
      "postDate": "2018-06-16T14:01:44.110Z",
      "content": "<p>My 2 cents: I felt the same way in this competition that I always had to keep up, esp. when we got passed by a mob with exactly the same LB, we figured right away that a new kernel probably got released and we had to abandon/adapt our current approach to the new kernel idea, since ours couldn't get that high score. I felt very tired and wanted to drop from this competition multiple times since we have no math/physics background like @CPMP or @Grzegorz or anyone else I don't know, it's a very difficult competition for us and we have to put lots of time to keep up. And I know lots of people have been putting tons of efforts on working on their own solutions. Running a good kernel can probably get you a silver medal in this competition, I think that's why @CPMP posted this, it wouldn't be fair to other hard workers. However, I also agree any kernel below <code>0.55</code> would be just a benchmark kernel at the end of the race. Not releasing a good kernel in the final weeks will be more sportsmanlike. </p>",
      "rawMarkdown": "My 2 cents: I felt the same way in this competition that I always had to keep up, esp. when we got passed by a mob with exactly the same LB, we figured right away that a new kernel probably got released and we had to abandon/adapt our current approach to the new kernel idea, since ours couldn't get that high score. I felt very tired and wanted to drop from this competition multiple times since we have no math/physics background like @CPMP or @Grzegorz or anyone else I don't know, it's a very difficult competition for us and we have to put lots of time to keep up. And I know lots of people have been putting tons of efforts on working on their own solutions. Running a good kernel can probably get you a silver medal in this competition, I think that's why @CPMP posted this, it wouldn't be fair to other hard workers. However, I also agree any kernel below `0.55` would be just a benchmark kernel at the end of the race. Not releasing a good kernel in the final weeks will be more sportsmanlike. ",
      "votes": 1,
      "replies": [
        {
          "id": 343992,
          "postDate": "2018-06-16T17:49:12.423Z",
          "content": "<p>I doubt simply submitting an output will be worth a silver medal... That'll be the top 50 spots for a competition this size, and there were nearly that many people already ahead of the most recent kernel. </p>\n\n<p>There will probably be a ton of late competition that will be able to beat the benchmarks. </p>",
          "rawMarkdown": "I doubt simply submitting an output will be worth a silver medal... That'll be the top 50 spots for a competition this size, and there were nearly that many people already ahead of the most recent kernel. \n\nThere will probably be a ton of late competition that will be able to beat the benchmarks. "
        },
        {
          "id": 344584,
          "postDate": "2018-06-18T09:25:57.957Z",
          "content": "<blockquote>\n  <p>I doubt simply submitting an output will be worth a silver medal... </p>\n</blockquote>\n\n<p>It depends on how good shared kernels will be.  A kernel at 0.55 may lead to silver IMHO given the number of participants is rather limited in this competition.</p>",
          "rawMarkdown": "&gt; I doubt simply submitting an output will be worth a silver medal... \n\nIt depends on how good shared kernels will be.  A kernel at 0.55 may lead to silver IMHO given the number of participants is rather limited in this competition."
        },
        {
          "id": 344591,
          "postDate": "2018-06-18T09:42:07.183Z",
          "content": "<p>I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.</p>",
          "rawMarkdown": "I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.",
          "votes": 6
        }
      ]
    },
    {
      "id": 346641,
      "postDate": "2018-06-22T04:58:21.427Z",
      "content": "<p>i am surprised that the tracks are pretty straight within radius of 200.</p>\n\n<p>FYI:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9656/Slide1.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9655/Slide2.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "i am surprised that the tracks are pretty straight within radius of 200.\n\nFYI:\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9656/Slide1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9655/Slide2.png",
      "votes": 2
    },
    {
      "id": 343699,
      "postDate": "2018-06-15T21:18:59.873Z",
      "content": "<p>This is a strange issue. On one hand, following his posts are incredibly helpful, but on the other hand it is hard to keep up with the scores. For most of this competition, my results have lagged behind the public solutions. I've finally gotten local validations in the .5x range (Grzegorz's previous work + a few secret features + HKC's track extension code (implemented in R)). I've been working hard to keep up, but with so many skilled contributors it can be challenging to beat the pack.</p>\n\n<p>I worry about another last minute kernel coming out that beats 90% of the public LB... I'm a fan of the two week idea you proposed in another competition... </p>",
      "rawMarkdown": "This is a strange issue. On one hand, following his posts are incredibly helpful, but on the other hand it is hard to keep up with the scores. For most of this competition, my results have lagged behind the public solutions. I've finally gotten local validations in the .5x range (Grzegorz's previous work + a few secret features + HKC's track extension code (implemented in R)). I've been working hard to keep up, but with so many skilled contributors it can be challenging to beat the pack.\n\nI worry about another last minute kernel coming out that beats 90% of the public LB... I'm a fan of the two week idea you proposed in another competition... ",
      "votes": 1,
      "replies": [
        {
          "id": 343727,
          "postDate": "2018-06-15T23:38:25.727Z",
          "content": "<p>@CPMP, <a href=\"/macfarll\">@macfarll</a>, I can feel some negative vibrations in your posts, so let's do some calculations. \nWho remembers the benchmark? What can you say about its quality? If you are talking about the winners reaching or even exceeding the score 0.8, the distance between 0.8 and present public kernel 0.5 is like the distance between present TOP3 (0.68) and the benchmark (0.2). In both cases the winners' kernels are 2.5 times better. So, at the end of the competition, present kernels of the score 0.5 will be as good as the benchmark now. </p>\n\n<p>I do not want to blow up the competition. I am not going to share anything in the last month and anything above 0.55.</p>\n\n<p>And regarding two races, I think, the most important race is the third one - the race of people who do not submit their results and do not share their ideas according to the Russian proverb \"Tishe edesh, dalshe budesh\". I suppose, the results in that race are 0.05 above the results in the race no. 1.</p>",
          "rawMarkdown": "@CPMP, @macfarll, I can feel some negative vibrations in your posts, so let's do some calculations. \nWho remembers the benchmark? What can you say about its quality? If you are talking about the winners reaching or even exceeding the score 0.8, the distance between 0.8 and present public kernel 0.5 is like the distance between present TOP3 (0.68) and the benchmark (0.2). In both cases the winners' kernels are 2.5 times better. So, at the end of the competition, present kernels of the score 0.5 will be as good as the benchmark now. \n\nI do not want to blow up the competition. I am not going to share anything in the last month and anything above 0.55.\n\nAnd regarding two races, I think, the most important race is the third one - the race of people who do not submit their results and do not share their ideas according to the Russian proverb \"Tishe edesh, dalshe budesh\". I suppose, the results in that race are 0.05 above the results in the race no. 1.",
          "votes": 5
        },
        {
          "id": 343730,
          "postDate": "2018-06-16T00:09:00.520Z",
          "content": "<p>I am not upset about your sharing of kernels, I enjoy seeing your work and I think your posts do a good job at inspiring the community. I look forward to any future kernels, as I'm sure I still have plenty to learn from them. I wouldn't be surprised if another kernel comes up with some form of supervised learning scores of .6x by the end of the competition...</p>",
          "rawMarkdown": "I am not upset about your sharing of kernels, I enjoy seeing your work and I think your posts do a good job at inspiring the community. I look forward to any future kernels, as I'm sure I still have plenty to learn from them. I wouldn't be surprised if another kernel comes up with some form of supervised learning scores of .6x by the end of the competition...\n\n"
        },
        {
          "id": 343741,
          "postDate": "2018-06-16T01:11:08.370Z",
          "content": "<blockquote>\n  <p>@CPMP, <a href=\"/macfarll\">@macfarll</a>, I can feel some negative vibrations in your posts</p>\n</blockquote>\n\n<p>@Grzegorz, There is none in mine.  If I had to say something negative then I would say it directly.  I am genuinely impressed by how fast some are to submit your outputs while they are absolutely not active in this competition.  They must have a script that watches the kernel page.  That's what I wrote.</p>\n\n<p>What I am absolutely against is last minute high level score script sharing.  This happened few hours before competition ends in the Talking Data competition, and it ruined the efforts of hundreds of participants.  It ruined it because many did not even see the script, thinking they were done for the competition.  Many of us have asked Kaggle to prevent kernel sharing during the last week of the competition in order to prevent this. Sharing good scripts 2 months before the end is not an issue at all, and your sharing certainly helped fuel interest and hope in this competition.</p>\n\n<p>I consider your race 3 to be part of race 1.  I now see that my post could be interpreted as race 1 being a race to the top of the public LB.  I actually meant a race to the top of the private LB.  That's the one that matters, right?  To your point,  advancing under cover can be a good strategy here.  Lots of us remember the case of Idle Speculation who surprised everyone with a great submission a week before end in Expedia competition, see <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/discussion/21440\">https://www.kaggle.com/c/expedia-hotel-recommendations/discussion/21440</a></p>",
          "rawMarkdown": "&gt; @CPMP, @macfarll, I can feel some negative vibrations in your posts\n\n@Grzegorz, There is none in mine.  If I had to say something negative then I would say it directly.  I am genuinely impressed by how fast some are to submit your outputs while they are absolutely not active in this competition.  They must have a script that watches the kernel page.  That's what I wrote.\n\nWhat I am absolutely against is last minute high level score script sharing.  This happened few hours before competition ends in the Talking Data competition, and it ruined the efforts of hundreds of participants.  It ruined it because many did not even see the script, thinking they were done for the competition.  Many of us have asked Kaggle to prevent kernel sharing during the last week of the competition in order to prevent this. Sharing good scripts 2 months before the end is not an issue at all, and your sharing certainly helped fuel interest and hope in this competition.\n\nI consider your race 3 to be part of race 1.  I now see that my post could be interpreted as race 1 being a race to the top of the public LB.  I actually meant a race to the top of the private LB.  That's the one that matters, right?  To your point,  advancing under cover can be a good strategy here.  Lots of us remember the case of Idle Speculation who surprised everyone with a great submission a week before end in Expedia competition, see https://www.kaggle.com/c/expedia-hotel-recommendations/discussion/21440",
          "votes": 4
        }
      ]
    },
    {
      "id": 345049,
      "postDate": "2018-06-19T05:37:35.493Z",
      "content": "<p>oh, someone got 0.8 on the leaderboard! :D</p>",
      "rawMarkdown": "oh, someone got 0.8 on the leaderboard! :D\n",
      "replies": [
        {
          "id": 345076,
          "postDate": "2018-06-19T06:45:55.770Z",
          "content": "<p>Outrunner from the race no. 3 lost his track. </p>",
          "rawMarkdown": "Outrunner from the race no. 3 lost his track. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 343642,
      "postDate": "2018-06-15T18:32:28.947Z",
      "content": "<p>They hope to get a good place in the leaderboard for Kaggle points. First = higher up.</p>",
      "rawMarkdown": "They hope to get a good place in the leaderboard for Kaggle points. First = higher up.",
      "replies": [
        {
          "id": 344664,
          "postDate": "2018-06-18T13:34:49.387Z",
          "content": "<p>In case of a tie, does the submission time actually count?</p>",
          "rawMarkdown": "In case of a tie, does the submission time actually count?"
        },
        {
          "id": 344678,
          "postDate": "2018-06-18T14:05:02.647Z",
          "content": "<p>The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely...if the code is different.</p>",
          "rawMarkdown": "The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely...if the code is different."
        },
        {
          "id": 344687,
          "postDate": "2018-06-18T14:14:34.453Z",
          "content": "<p>@Diego -</p>\n\n<p>If there is an exact submission score tie, the leaderboard ordering is based on time submitted.</p>\n\n<p>If both the submission scores <em>and</em> submission times are exactly the same, we consult Schrödinger's cat to determine leaderboard ordering. :-)</p>\n\n<p>(In that second case, it's actually deterministic based on db <code>submission_id</code>.)</p>",
          "rawMarkdown": "@Diego -\n\nIf there is an exact submission score tie, the leaderboard ordering is based on time submitted.\n\nIf both the submission scores *and* submission times are exactly the same, we consult Schrödinger's cat to determine leaderboard ordering. :-)\n\n(In that second case, it's actually deterministic based on db `submission_id`.)",
          "votes": 3
        },
        {
          "id": 344691,
          "postDate": "2018-06-18T14:26:54.650Z",
          "content": "<p>&gt; The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely. </p>\n\n<p><a href=\"/droussea\">@droussea</a>  As I write it there are 53 people who submitted the output of a public kernel with a public LB score of 0.4953.  Ties are not only likely, they can involve a very large number of participants.  That's why I wrote the post in the first place.  </p>",
          "rawMarkdown": "&gt; The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely. \n\n@droussea  As I write it there are 53 people who submitted the output of a public kernel with a public LB score of 0.4953.  Ties are not only likely, they can involve a very large number of participants.  That's why I wrote the post in the first place.  ",
          "votes": 1
        },
        {
          "id": 344694,
          "postDate": "2018-06-18T14:29:35.650Z",
          "content": "<p><a href=\"/inversion\">@inversion</a>, thanks for confirming the second race impacts people's rank: the first to submit a shared kernel output gets an advantage over the slower ones.</p>",
          "rawMarkdown": "@inversion, thanks for confirming the second race impacts people's rank: the first to submit a shared kernel output gets an advantage over the slower ones.",
          "votes": 1
        },
        {
          "id": 344705,
          "postDate": "2018-06-18T14:46:54.090Z",
          "content": "<p><a href=\"/inversion\">@inversion</a>:</p>\n\n<blockquote>\n  <p>we consult Schrödinger's cat to determine leaderboard ordering. :-)</p>\n</blockquote>\n\n<p>Incredible, is it still alive? Oops, I forgot - a cat has seven lives.</p>",
          "rawMarkdown": "@inversion:\n&gt; we consult Schrödinger's cat to determine leaderboard ordering. :-)\n\nIncredible, is it still alive? Oops, I forgot - a cat has seven lives.",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 345210,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-06-19T12:10:08.663000",
      "content": "<p>After the shock of the score 0.8, maybe it is time for sharing ideas? I can prepare a kernel which I hope can satisfy majority of the Kagglers active in this competition. I would like to share a more advanced method of ensembling the tracks in the loop of DBSCAN and models with shifted z. I can write it in such a way that it will give no submission output and it will take so much time, that it will be not possible to obtain the full results at Kaggle servers. The score will be 0.54-0.58. As in my first kernel, there will be much place for optimization and your own ideas. What do you think?</p>\n\n<p>Of course not today - today Polish team plays football ;)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 345213,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-19T12:13:10.963000",
          "content": "<p>i  second your idea. I believe kagglers can push the score from 0.58 to 0.70</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345224,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-19T12:35:09.477000",
          "content": "<p>@Grzegorz, Good luck! I hope the Polish team will perform better than the German team. I would only share the idea instead of a kernel so every of us needs to work hard to win a medal. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 345226,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-19T12:43:19.037000",
          "content": "<p><a href=\"/grzegorz\">@grzegorz</a>, I like <a href=\"/nicole\">@nicole</a>'s idea.  Share the idea or even pseudo code and make us work for the benefit!</p>\n\n<p><a href=\"/nicole\">@nicole</a> - I'm guessing that you have created some analytics to gain insight into the \"truth\" ... That would be a welcome kernel too!</p>\n\n<p>PS - I am greatly conflicted by the World Cup.  I grew up in South Texas, so I have a natural affinity for Mexico, my wife is a German citizen, and the primary engineering team I work with is in Gdansk!  Naturally, rooting for team USA is, sadly, not possible!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 345228,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-19T12:46:18.367000",
          "content": "<p>@Grzegorz, </p>\n\n<p>this probably contradicts what you wrote yesterday:</p>\n\n<blockquote>\n  <p>I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.</p>\n</blockquote>\n\n<p>Note also that the organizers decided to run a Kaggle competition and not a collaborative effort.</p>\n\n<p>My take is that ideas sharing looks more interesting.  For instance, my remark on the origin of tracks and the following discussion this morning probably enlightened many on the need to use various z values at origin.  People who want to share more can team.</p>\n\n<p>This said, I am in no position to dictate anything here.  </p>\n\n<p>Let me conclude by thanking you for asking before sharing.  I'm curious to see what others will say, do they agree with you and Heng, or do they think people should find their own way?  </p>\n\n<p>Edit: i wrote this before seeing Nicole and John's answers.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 345249,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-19T13:53:20.310000",
          "content": "<p>Personally, I am already going to start working on modified origins because of this conversation, however if new ideas are shared, by either concept or kernel, I'd be excited to see them. I definitely felt like I learned more about how the code worked when I recreated HCK's extension code in R, and would be more likely to improve on the code when I had to do the legwork. Alternatively, if a kernel was submitted, more work would be put into other parts of the competition, and the community might push harder towards  .9x. </p>\n\n<p>I don't think submitting a new kernel has much risk of people not having to work for the medals. There is still so much time left in the competition, there will be more entrants able to beat the kernel outputs as we get closer to the deadline.</p>\n\n<p>Really, my biggest worry would be if multiple ideas start getting shared, everyone's ideas will contribute to the computational barrier, ex: user 1 shares an idea that causes good solutions to run track extension twice as long, and user 2 shares an idea that causes good solutions to run track extension twice as often.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345252,
          "author_name": "Michael Maguire",
          "author_url": "",
          "post_date": "2018-06-19T14:04:05.663000",
          "content": "<p>@John Sweeney...ICELAND!!  Who doesn't like an underdog?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345382,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-19T19:40:51.583000",
          "content": "<p>The word \"idea\" has at least two meanings. Let me explain it on two examples:</p>\n\n<p>1) \"Maybe it is a good idea to shift our models along z axis or maybe it is not at all, because the most of tracks come from (0,0,0)\"</p>\n\n<p>2) \"It is a good idea to shift models along z axis, because ensembling 7 models gives a boost 0.08 and ensembling 21 models gives a boost 0.105, however the time of calculations is extremely long\".</p>\n\n<p>My thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. Talking about satisfying majority of active Kagglers I thought also about myself. I am not sure if it fully satisfies me to share my checked solutions in the form of few sentences here and there inside someone else's threads. So, mentioned kind of ensembling the tracks will be presented by me  in the form of kernel, when Top100 will exceed 0.58, i.e. when it will be completely useless. You lose nothing - this ensembling of DBSCAN clusters plays similar role and gives the same boost as Heng's extend function, and the effect of ensembling models shifted along z axis was described by me here and there. Best wishes.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 345420,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-06-19T21:03:10.867000",
          "content": "<p>&gt;  I hope the Polish team will perform better than the German team. </p>\n\n<p>Unfortunately the Polish team has lost.\nI also watched my native team (Russia) today and I'm happy we won. (Probably we will lose in the first game at playoff, but anyway).</p>\n\n<p>I like sharing ideas but in the form of algorithms. \"shift models along z axis, because ensembling 7 models gives a boost 0.08\" &lt;--- That's enough for me.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 345427,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-19T21:21:15.450000",
          "content": "<p>I agree with Sergey.  It is enough of a hint.  With  this level of detail, we have a good direction, but may stumble into something else good we would never find with a directly usable code snippet</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345433,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-19T21:48:13.240000",
          "content": "<p>@Grzegorz, just let us have some fun exploring hidden features ourselves :)  as @John said, we may stumble into something new and better, you've helped us a lot along our journey, highly appreciated! Ensembling 21 models is crazy, but with a 0.1 boost, that's huge. We're reworking on the ensembling code too, since it didn't give us much of a benefit anymore after our single models got much stronger. You're right, I didn't ensemble models with shifted z features, it's enough to know this is an area we can still explore. This competition is very special, not a traditional ML/DL competition but more of a research / brain damaging / sleep depriving competition. :D </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 345436,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-19T21:58:17.280000",
          "content": "<p>@Nicole, I don't understand why this competition is so engaging to me, as someone who works in mortgage you'd think I'd be working on the mortgage competition, but I keep coming back here. I've been staying up late working on this so often I've somehow fallen into a biphasic sleep cycle. </p>\n\n<p>@Grzegorz, you do what you want to do, we'll all be interested to see what happens next. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 345448,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-19T22:50:36.387000",
          "content": "<blockquote>\n  <p>My thoughts shared on this forum are not ideas of type 1, they are rather checked solutions, not ideas. </p>\n</blockquote>\n\n<p>@Grzegorz, Mine as well, but I rather stick to ideas instead of reusable code.  I fully agree with what @SIlogram <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/58332\">expressed in another ongoing competition</a></p>\n\n<blockquote>\n  <p>First, let me say that I'm not a big fan of the kernels and miss the days when Kagglers would provide general guidelines and hints rather than full-blown solutions</p>\n</blockquote>\n\n<p>And I think you understand now why computing my solution takes time:</p>\n\n<blockquote>\n  <p>however the time of calculations is extremely long</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 345467,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-19T23:52:05.107000",
          "content": "<p>Thank you all for your posts. I would like to be infected with optimism from you.\nI think I'm the world champion in skepticism. Already at the beginning of the competition, I assumed that it is impossible to win by someone who starts from scratch, regardless of what ideas he or she will have. That's why I did not pay too much attention to my ideas. The same ideas invented by you were treated very seriously.</p>\n\n<p>Regarding the winner, I thought that somebody would simply read some wise publication about separation of tracks, implement an algorithm being a version of the Kalman filter or other combinatorial algorithm and that's it. The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 345642,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-20T07:31:54.503000",
          "content": "<blockquote>\n  <p>The winning code would be of course completely useless for people from CERN, because they have better one and it was not the purpose of the competition. Unfortunately, I still think so.</p>\n</blockquote>\n\n<p>You may be right for this, but this is not our problem, is it?  </p>\n\n<p>CERN say in one of their presentation (probably link shared by Heng) that they run this competition with the hope that someone will get good results use a very different technique than the ones they have tried.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345741,
          "author_name": "Gabriel Preda",
          "author_url": "",
          "post_date": "2018-06-20T11:21:55.117000",
          "content": "<p>It might be that, although in terms of accuracy, the winning result will not be significantly better, in terms of algorithm efficiency, will be better. Well, we don't actually know. Let's see what will happen. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345767,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-06-20T12:20:17.357000",
          "content": "<p>To be absolutely clear : the CERN experiments are directly and deeply interested in the results of the challenge, because track reconstruction is a critical problem, see details in <a href=\"https://www.nature.com/articles/d41586-018-05084-2\">https://www.nature.com/articles/d41586-018-05084-2</a>\nIndeed our objective of this kaggle competition is to discover new algorithms. Note that, as indicated in \"Welcome from the organiser\" post, we will launch the second phase of the competition in July (so overlapping with the Kaggle one), with its own set of prizes, where participants will have to submit code, and the CPU time of the evaluation will be evaluated, and participants will be ranked with a score combining the accuracy (== as computed in this first phase) and CPU). More on this soon.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 345831,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-20T14:49:03.383000",
          "content": "<p>@David, It is still 2 months to go. Consider extending the set of additional prizes for more symbolic ones. There are rich countries of high salaries and low taxes, e.g. Switzerland, and less rich countries of low salaries and high taxes, e.g. Poland. Even if I create a genius code I will not apply for Tesla V100, because it would be connected with taking credit equal my 5-6 salaries to pay the tax for a material award from USA (2 if from EU).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 345910,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-20T17:45:12.250000",
          "content": "<p>Grzegorz you deserve (with Heng) a prize for sharing so much.  I hope you understand that nobody has issues with your sharing.  But some, including me, have issues with people merely submitting your kernels output.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 346021,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-20T23:33:12.027000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346025,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-20T23:42:10.440000",
          "content": "<p>A week ago I would have guessed that the top 3 scores came from something more advanced than improved clustering.  Now, as CPMP continues to climb towards 0.7, I think maybe everyone is still clustering...</p>\n\n<p>Even 0.8 through clustering?  Why not?</p>\n\n<p>Will a supervised algorithm win or not?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 346207,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-21T09:24:19.457000",
          "content": "<p>@John my 2 cents again, I don't think a pure supervised learning can win this competition, but I guess you mean if a combined solution would. A DL model helps to predict the possibility of hits but at the end of the day you still have to do track fitting (post processing) as what you would do with unsupervised learning. @Heng said his LSTM is better than his backfitting code, worth giving it a shot.  I'm very eager to know if <a href=\"/outrunner\">@outrunner</a> used a supervised learning approach at all, but I doubt she would see my comment. :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 346220,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-21T10:06:29.020000",
          "content": "<p>@John, I think I can push my current approach (clustering plus few secret sauce ingredients) above 0.7, but I doubt it can reach 0.8, one reason being is that it focus on tracks starting near the central z axis.  Anyway, I'm pursuing my current track (pun intended ;) ) as it helps me get a better understanding of what we deal with.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 346267,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-21T11:46:36.303000",
          "content": "<p>The score limit if only the particles coming from 0,0,z are under consideration is about 0.82.</p>\n\n<p><a href=\"https://www.kaggle.com/sionek/score-limit-for-0-0-z\">https://www.kaggle.com/sionek/score-limit-for-0-0-z</a></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 346273,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-21T11:55:10.007000",
          "content": "<p>Right, I have been trying to publish a notebook about it for the last 2 hours without success, Kaggle notebooks are broken now.  Anyway, even if you enlarge the region to 16 mm around the z axis, score is still below 0.85.  That's why I wrote above that identifying tracks that do not originate from origin will be key to winning this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346276,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-06-21T12:04:10.483000",
          "content": "<blockquote>\n  <p>identifying tracks that do not originate from origin will be key to winning this competition.</p>\n</blockquote>\n\n<p>Just now we must learn to identify tracks that do originate from origin. To get a score &gt;0.8. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 346290,
          "author_name": "Zidmie",
          "author_url": "",
          "post_date": "2018-06-21T12:18:52.617000",
          "content": "<p>Interesting! \nI tried a specific method to catch those tracks (something similar to the z shifting, but in x/y directions) : there is a score of 0.18 to catch , and I just got 0.005. And the calculation is very very long... </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 346294,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-21T12:31:10.107000",
          "content": "<p>@Zidmie, this is the real challenge indeed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346444,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-21T17:50:40.193000",
          "content": "<p>I've been traveling all week and haven't been able to try any of the excellent \"ideas\" shared here.</p>\n\n<p>I guess I need to do some visualization, but I fundamentally don't understand why particles originating far from x,y = 0,0 are more difficult to find.</p>\n\n<p>Any moving charged particle in the magnetic field will follow a helical path around the axis of the field (x,y = 0,0).</p>\n\n<p>It's path plotted in r,z coordinates will be approximately linear z = mr + z0, (z-z0)/r =  const</p>\n\n<p>I guess searching for appropriate values of z0 can be expensive.</p>\n\n<p>Is my logic flawed?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346451,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-21T18:00:52.107000",
          "content": "<p>Expanding on the maximum possible score:</p>\n\n<blockquote>\n  <p>using x_window:  500 , y_window:  500 , z_window:  2900 : score\n  0.9973638 , no. of hits: 83.94644 %</p>\n  \n  <p>using x_window:  300 , y_window:  300 , z_window:  1000 : score\n  0.9829006 , no. of hits: 82.78694 %</p>\n  \n  <p>using x_window:  100 , y_window:  100 , z_window:  1000 : score\n  0.9450124 , no. of hits: 80.22416 %</p>\n  \n  <p>using x_window:  100 , y_window:  100 , z_window:  500 : score\n  0.9259655 , no. of hits: 78.79893 %</p>\n  \n  <p>using x_window:  20 , y_window:  20 , z_window:  200 : score 0.8695067\n  , no. of hits: 73.77835 %</p>\n  \n  <p>using x_window:  20 , y_window:  20 , z_window:  20 : score 0.8497927\n  , no. of hits: 71.76945 %</p>\n  \n  <p>using x_window:  0.05 , y_window:  0.05 , z_window:  20 : score\n  0.8198518 , no. of hits: 69.59112 %</p>\n</blockquote>\n\n<p>From the fork I made of Grzegorz Sionkowski's (The Chemist) kernel</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 346531,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-06-21T22:31:10.817000",
          "content": "<p>To John Sweeney:  if a particle is coming from far from x,y=0,0 then its helicoidal trajectory will appear to miss the axis 0.0.\nSee a slightly displaced vertex in CMS  <img src=\"https://cds.cern.ch/record/2280025/files/QCD-f35-53384671-RhoPhi-white-ann.png?subformat=icon-1440\" alt=\"Slightly displaced vertex in CMS\"> (all light green are tracks coming from the origin, the darker one are coming from a displaced vertex).\nand a very displaced vertex in ATLAS <img src=\"https://cds.cern.ch/record/1381516/files/vp1_run165821_evt1605517_rpv.png\" alt=\"Very displaced vertex in ATLAS\"> (missing by a larger distance the origin).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 346645,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-22T05:20:29.557000",
          "content": "<p>@John, helix unrolling + clustering code shared in kernels assume that x,y = (0,0) is part of the track.  It does not detect tracks that start far from the origin.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346672,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-22T06:44:23.170000",
          "content": "<p>@John, Additionally, shared kernels assign only a part of helix to the track - this one while the distance from x,y = (0,0) increases. A part of helix after 0.5 turn (when this distance starts to decrease) is lost or treated as a track of other particle moving along the same helix but in the opposite direction.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 347078,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-06-23T07:23:05.590000",
          "content": "<p>@Grzegorz Sionkowski\nThanks for the R kernel, It gave me a big push to code to this competition in R as well. I converted the R code of the 0.456 score script to python and then you made a better one. Now I'll just use your R, with hopefully some improvements :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 344409,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "2018-06-17T23:40:17.567000",
      "content": "<p>Side note: when your new submission finally scores better that the kernel output, it can be satisfying to pass the big chunk of submissions. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 345933,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-06-20T18:45:39.170000",
      "content": "<p>A few ideas - <a href=\"https://github.com/HEPTrkX\">HEPTrkX</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 343911,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-06-16T14:01:44.110000",
      "content": "<p>My 2 cents: I felt the same way in this competition that I always had to keep up, esp. when we got passed by a mob with exactly the same LB, we figured right away that a new kernel probably got released and we had to abandon/adapt our current approach to the new kernel idea, since ours couldn't get that high score. I felt very tired and wanted to drop from this competition multiple times since we have no math/physics background like @CPMP or @Grzegorz or anyone else I don't know, it's a very difficult competition for us and we have to put lots of time to keep up. And I know lots of people have been putting tons of efforts on working on their own solutions. Running a good kernel can probably get you a silver medal in this competition, I think that's why @CPMP posted this, it wouldn't be fair to other hard workers. However, I also agree any kernel below <code>0.55</code> would be just a benchmark kernel at the end of the race. Not releasing a good kernel in the final weeks will be more sportsmanlike. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 343992,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-16T17:49:12.423000",
          "content": "<p>I doubt simply submitting an output will be worth a silver medal... That'll be the top 50 spots for a competition this size, and there were nearly that many people already ahead of the most recent kernel. </p>\n\n<p>There will probably be a ton of late competition that will be able to beat the benchmarks. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344584,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-18T09:25:57.957000",
          "content": "<blockquote>\n  <p>I doubt simply submitting an output will be worth a silver medal... </p>\n</blockquote>\n\n<p>It depends on how good shared kernels will be.  A kernel at 0.55 may lead to silver IMHO given the number of participants is rather limited in this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344591,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-18T09:42:07.183000",
          "content": "<p>I worked hard for each of my medals, so you may be sure I will not publish any kernel which would enable to get a bronze medal for free.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 346641,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-22T04:58:21.427000",
      "content": "<p>i am surprised that the tracks are pretty straight within radius of 200.</p>\n\n<p>FYI:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9656/Slide1.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9655/Slide2.png\" alt=\"enter image description here\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 343699,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "2018-06-15T21:18:59.873000",
      "content": "<p>This is a strange issue. On one hand, following his posts are incredibly helpful, but on the other hand it is hard to keep up with the scores. For most of this competition, my results have lagged behind the public solutions. I've finally gotten local validations in the .5x range (Grzegorz's previous work + a few secret features + HKC's track extension code (implemented in R)). I've been working hard to keep up, but with so many skilled contributors it can be challenging to beat the pack.</p>\n\n<p>I worry about another last minute kernel coming out that beats 90% of the public LB... I'm a fan of the two week idea you proposed in another competition... </p>",
      "votes": 1,
      "replies": [
        {
          "id": 343727,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-15T23:38:25.727000",
          "content": "<p>@CPMP, <a href=\"/macfarll\">@macfarll</a>, I can feel some negative vibrations in your posts, so let's do some calculations. \nWho remembers the benchmark? What can you say about its quality? If you are talking about the winners reaching or even exceeding the score 0.8, the distance between 0.8 and present public kernel 0.5 is like the distance between present TOP3 (0.68) and the benchmark (0.2). In both cases the winners' kernels are 2.5 times better. So, at the end of the competition, present kernels of the score 0.5 will be as good as the benchmark now. </p>\n\n<p>I do not want to blow up the competition. I am not going to share anything in the last month and anything above 0.55.</p>\n\n<p>And regarding two races, I think, the most important race is the third one - the race of people who do not submit their results and do not share their ideas according to the Russian proverb \"Tishe edesh, dalshe budesh\". I suppose, the results in that race are 0.05 above the results in the race no. 1.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 343730,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-16T00:09:00.520000",
          "content": "<p>I am not upset about your sharing of kernels, I enjoy seeing your work and I think your posts do a good job at inspiring the community. I look forward to any future kernels, as I'm sure I still have plenty to learn from them. I wouldn't be surprised if another kernel comes up with some form of supervised learning scores of .6x by the end of the competition...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343741,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-16T01:11:08.370000",
          "content": "<blockquote>\n  <p>@CPMP, <a href=\"/macfarll\">@macfarll</a>, I can feel some negative vibrations in your posts</p>\n</blockquote>\n\n<p>@Grzegorz, There is none in mine.  If I had to say something negative then I would say it directly.  I am genuinely impressed by how fast some are to submit your outputs while they are absolutely not active in this competition.  They must have a script that watches the kernel page.  That's what I wrote.</p>\n\n<p>What I am absolutely against is last minute high level score script sharing.  This happened few hours before competition ends in the Talking Data competition, and it ruined the efforts of hundreds of participants.  It ruined it because many did not even see the script, thinking they were done for the competition.  Many of us have asked Kaggle to prevent kernel sharing during the last week of the competition in order to prevent this. Sharing good scripts 2 months before the end is not an issue at all, and your sharing certainly helped fuel interest and hope in this competition.</p>\n\n<p>I consider your race 3 to be part of race 1.  I now see that my post could be interpreted as race 1 being a race to the top of the public LB.  I actually meant a race to the top of the private LB.  That's the one that matters, right?  To your point,  advancing under cover can be a good strategy here.  Lots of us remember the case of Idle Speculation who surprised everyone with a great submission a week before end in Expedia competition, see <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/discussion/21440\">https://www.kaggle.com/c/expedia-hotel-recommendations/discussion/21440</a></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 345049,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-06-19T05:37:35.493000",
      "content": "<p>oh, someone got 0.8 on the leaderboard! :D</p>",
      "votes": 0,
      "replies": [
        {
          "id": 345076,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-19T06:45:55.770000",
          "content": "<p>Outrunner from the race no. 3 lost his track. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 343642,
      "author_name": "redstr",
      "author_url": "",
      "post_date": "2018-06-15T18:32:28.947000",
      "content": "<p>They hope to get a good place in the leaderboard for Kaggle points. First = higher up.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 344664,
          "author_name": "Diego Vicente",
          "author_url": "",
          "post_date": "2018-06-18T13:34:49.387000",
          "content": "<p>In case of a tie, does the submission time actually count?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344678,
          "author_name": "David Rousseau",
          "author_url": "",
          "post_date": "2018-06-18T14:05:02.647000",
          "content": "<p>The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely...if the code is different.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 344687,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-06-18T14:14:34.453000",
          "content": "<p>@Diego -</p>\n\n<p>If there is an exact submission score tie, the leaderboard ordering is based on time submitted.</p>\n\n<p>If both the submission scores <em>and</em> submission times are exactly the same, we consult Schrödinger's cat to determine leaderboard ordering. :-)</p>\n\n<p>(In that second case, it's actually deterministic based on db <code>submission_id</code>.)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 344691,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-18T14:26:54.650000",
          "content": "<p>&gt; The ranking is determined from the double precision calculation of the score (so beyond the 4 digits appearing on the leaderboard). Tie is extremely unlikely. </p>\n\n<p><a href=\"/droussea\">@droussea</a>  As I write it there are 53 people who submitted the output of a public kernel with a public LB score of 0.4953.  Ties are not only likely, they can involve a very large number of participants.  That's why I wrote the post in the first place.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344694,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-18T14:29:35.650000",
          "content": "<p><a href=\"/inversion\">@inversion</a>, thanks for confirming the second race impacts people's rank: the first to submit a shared kernel output gets an advantage over the slower ones.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 344705,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-18T14:46:54.090000",
          "content": "<p><a href=\"/inversion\">@inversion</a>:</p>\n\n<blockquote>\n  <p>we consult Schrödinger's cat to determine leaderboard ordering. :-)</p>\n</blockquote>\n\n<p>Incredible, is it still alive? Oops, I forgot - a cat has seven lives.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "343427": "There are two races in this competition.\n\n 1. Race to the top of LB\n 2. Be the first to submit Grzegorz kernels output.  Seems some folks have scripts that alert them when a good kernel is shared given how fast they react ;)",
    "345210": "After the shock of the score 0.8, maybe it is time for sharing ideas? I can prepare a kernel which I hope can satisfy majority of the Kagglers active in this competition. I would like to share a more advanced method of ensembling the tracks in the loop of DBSCAN and models with shifted z. I can write it in such a way that it will give no submission output and it will take so much time, that it will be not possible to obtain the full results at Kaggle servers. The score will be 0.54-0.58. As in my first kernel, there will be much place for optimization and your own ideas. What do you think?\n\nOf course not today - today Polish team plays football ;)",
    "344409": "Side note: when your new submission finally scores better that the kernel output, it can be satisfying to pass the big chunk of submissions. ",
    "345933": "A few ideas - [HEPTrkX][1]\n\n\n  [1]: https://github.com/HEPTrkX",
    "343911": "My 2 cents: I felt the same way in this competition that I always had to keep up, esp. when we got passed by a mob with exactly the same LB, we figured right away that a new kernel probably got released and we had to abandon/adapt our current approach to the new kernel idea, since ours couldn't get that high score. I felt very tired and wanted to drop from this competition multiple times since we have no math/physics background like @CPMP or @Grzegorz or anyone else I don't know, it's a very difficult competition for us and we have to put lots of time to keep up. And I know lots of people have been putting tons of efforts on working on their own solutions. Running a good kernel can probably get you a silver medal in this competition, I think that's why @CPMP posted this, it wouldn't be fair to other hard workers. However, I also agree any kernel below `0.55` would be just a benchmark kernel at the end of the race. Not releasing a good kernel in the final weeks will be more sportsmanlike. ",
    "346641": "i am surprised that the tracks are pretty straight within radius of 200.\n\nFYI:\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9656/Slide1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/346641/9655/Slide2.png",
    "343699": "This is a strange issue. On one hand, following his posts are incredibly helpful, but on the other hand it is hard to keep up with the scores. For most of this competition, my results have lagged behind the public solutions. I've finally gotten local validations in the .5x range (Grzegorz's previous work + a few secret features + HKC's track extension code (implemented in R)). I've been working hard to keep up, but with so many skilled contributors it can be challenging to beat the pack.\n\nI worry about another last minute kernel coming out that beats 90% of the public LB... I'm a fan of the two week idea you proposed in another competition... ",
    "345049": "oh, someone got 0.8 on the leaderboard! :D\n",
    "343642": "They hope to get a good place in the leaderboard for Kaggle points. First = higher up."
  }
}