{
  "id": 54400,
  "title": "Kernels Schmernels",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54400",
  "author_name": "Scirpus",
  "post_date": "2018-04-12T20:08:10.445000",
  "votes": 15,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Anyone else going to join competitions just two weeks before the deadline to invest the least possible effort to get the maximum return LOL ;)</p>",
  "messages": [
    {
      "id": 313127,
      "postDate": "2018-04-12T20:08:10.447Z",
      "content": "<p>Anyone else going to join competitions just two weeks before the deadline to invest the least possible effort to get the maximum return LOL ;)</p>",
      "rawMarkdown": "Anyone else going to join competitions just two weeks before the deadline to invest the least possible effort to get the maximum return LOL ;)",
      "votes": 15
    },
    {
      "id": 314856,
      "postDate": "2018-04-16T11:23:44.610Z",
      "content": "<p>I don't totally agree with comments below - there is a big difference between posting kernels with all preprocessing, hyperparams and training iterations to get silver just forking the code and on the other side to post a general idea or approach or hint to feed your brains (great example - <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769</a>). </p>",
      "rawMarkdown": "I don't totally agree with comments below - there is a big difference between posting kernels with all preprocessing, hyperparams and training iterations to get silver just forking the code and on the other side to post a general idea or approach or hint to feed your brains (great example - https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769). ",
      "votes": 7,
      "replies": [
        {
          "id": 314864,
          "postDate": "2018-04-16T11:48:50.980Z",
          "content": "<p>I'll just drop a quick thanks here, it's very motivational to submit something that's well-received :)</p>",
          "rawMarkdown": "I'll just drop a quick thanks here, it's very motivational to submit something that's well-received :)",
          "votes": 5
        },
        {
          "id": 314886,
          "postDate": "2018-04-16T12:37:49.800Z",
          "content": "<p>I agree - great kernel - because it can be used in general, people can learn something that they can take forward.  </p>",
          "rawMarkdown": "I agree - great kernel - because it can be used in general, people can learn something that they can take forward.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 313340,
      "postDate": "2018-04-13T06:02:08.657Z",
      "content": "<p>Yes, ladies and gentleman please stop spawning new kernels, you annihilate our efforts:( Or at least make it less straightforward somehow </p>",
      "rawMarkdown": "Yes, ladies and gentleman please stop spawning new kernels, you annihilate our efforts:( Or at least make it less straightforward somehow ",
      "votes": 7,
      "replies": [
        {
          "id": 313611,
          "postDate": "2018-04-13T15:09:26.380Z",
          "content": "<p>I think the kernels are still just scratching the surface.  It's a very hard problem because of the size of the data.  If the goal is to find the best possible solutions, it makes sense to have some cooperation until we reach the later stages.</p>",
          "rawMarkdown": "I think the kernels are still just scratching the surface.  It's a very hard problem because of the size of the data.  If the goal is to find the best possible solutions, it makes sense to have some cooperation until we reach the later stages.",
          "votes": 8
        },
        {
          "id": 313765,
          "postDate": "2018-04-13T19:58:17.970Z",
          "content": "<p>But this is more like doing random things, finding a good one and saying: see, THIS works! This is not an attempt to understand the problem, so it seems to have no learning value, it only flattens the leaderboard. And I believe magic features usually have some good explanation for their power e.g. the magic ip/os/device/app next_click feature from the latest kernel seems very reasonable, because ip/os/device combination is a pretty good definition of a user, so it is building on a very natural user/app combination.</p>",
          "rawMarkdown": "But this is more like doing random things, finding a good one and saying: see, THIS works! This is not an attempt to understand the problem, so it seems to have no learning value, it only flattens the leaderboard. And I believe magic features usually have some good explanation for their power e.g. the magic ip/os/device/app next_click feature from the latest kernel seems very reasonable, because ip/os/device combination is a pretty good definition of a user, so it is building on a very natural user/app combination.",
          "votes": 4
        },
        {
          "id": 313773,
          "postDate": "2018-04-13T20:19:48.430Z",
          "content": "<p>I still remember that in Mercari Price Suggestion Competition, a kernel could rank 9% on lb was published the day before the last submission day ;), and I dropped from 9% to 15%+. Therefore for me I will keep an eye on kernel section all the time lol.</p>",
          "rawMarkdown": "I still remember that in Mercari Price Suggestion Competition, a kernel could rank 9% on lb was published the day before the last submission day ;), and I dropped from 9% to 15%+. Therefore for me I will keep an eye on kernel section all the time lol.",
          "votes": 2
        },
        {
          "id": 313787,
          "postDate": "2018-04-13T20:42:43.470Z",
          "content": "<p>@Kamil There are a lot of things that seem like reasonable possibilities. Most of the features I've seen in kernels are pretty straightforward, not weird random ideas. But testing which ones work is difficult given the size of the data.  A lot of machine learning involves trying things out to see what works.  Public kernels judged by public LB scores may not be the ideal framework for this, but they're better than nothing (provided one does proper validation).  There is some learning value in learning what works, but the main learning value here is in learning to wrestle with the data.</p>",
          "rawMarkdown": "@Kamil There are a lot of things that seem like reasonable possibilities. Most of the features I've seen in kernels are pretty straightforward, not weird random ideas. But testing which ones work is difficult given the size of the data.  A lot of machine learning involves trying things out to see what works.  Public kernels judged by public LB scores may not be the ideal framework for this, but they're better than nothing (provided one does proper validation).  There is some learning value in learning what works, but the main learning value here is in learning to wrestle with the data.",
          "votes": 4
        },
        {
          "id": 313846,
          "postDate": "2018-04-13T23:51:37.470Z",
          "content": "<p>(You say \"flattens the leaderboard.\" I say \"sets new benchmarks.\")</p>",
          "rawMarkdown": "(You say \"flattens the leaderboard.\" I say \"sets new benchmarks.\")",
          "votes": 7
        },
        {
          "id": 314006,
          "postDate": "2018-04-14T10:53:33.420Z",
          "content": "<blockquote>\n  <p>Therefore for me I will keep an eye on kernel section all the time lol.</p>\n</blockquote>\n\n<p>That's the kind of thing that looks obvious after seeing the actual results of a competition. But when it stills running, it's very hard to know if a kernel mixing 30 other ones is really good or just lucky (we know that ensembles are good, but a small difference by chance can make a particular one looks really good). We end up looking for the winner ones and thinking it'll be easy to identify them in the next competition.</p>",
          "rawMarkdown": "&gt; Therefore for me I will keep an eye on kernel section all the time lol.\n\nThat's the kind of thing that looks obvious after seeing the actual results of a competition. But when it stills running, it's very hard to know if a kernel mixing 30 other ones is really good or just lucky (we know that ensembles are good, but a small difference by chance can make a particular one looks really good). We end up looking for the winner ones and thinking it'll be easy to identify them in the next competition.\n",
          "votes": 1
        },
        {
          "id": 314090,
          "postDate": "2018-04-14T15:27:26.257Z",
          "content": "<p>To be honest, kernels have short life. In any competition, ideas evolve over the period of time and techniques get changed in positive direction if there's a co-operation in terms of feedback, suggestion or a forked version of kernel that shows how it can be improved.  Completely agree that kernel popping out in the last week to destroy leaderboard is really bad.</p>\n\n<p>With regard to this competition, I can still remember the how public LB evolved over the time. (exclude blends):</p>\n\n<pre><code>LB ~ 0.9550 getting started kernels\nLB ~ 0.9631 scale_pos_weight came into the light\nLB ~ 0.9680 Memory optimized kernels \nLB ~ 0.9730 Kernel with user ranking features\nLB ~ 0.9752 Innovative modelling approach i.e FTRL and time delta feats\n</code></pre>\n\n<p>And as on today, we still have 3 more weeks to competition deadline. I am sure that most people don't like kernels section hijacked by blend-bros :). btw, have you noticed that Kaggle also doesn't like blends and have disabled <code>best score</code> feature? </p>",
          "rawMarkdown": "To be honest, kernels have short life. In any competition, ideas evolve over the period of time and techniques get changed in positive direction if there's a co-operation in terms of feedback, suggestion or a forked version of kernel that shows how it can be improved.  Completely agree that kernel popping out in the last week to destroy leaderboard is really bad.\n\nWith regard to this competition, I can still remember the how public LB evolved over the time. (exclude blends):\n\n    LB ~ 0.9550 getting started kernels\n    LB ~ 0.9631 scale_pos_weight came into the light\n    LB ~ 0.9680 Memory optimized kernels \n    LB ~ 0.9730 Kernel with user ranking features\n    LB ~ 0.9752 Innovative modelling approach i.e FTRL and time delta feats\n\nAnd as on today, we still have 3 more weeks to competition deadline. I am sure that most people don't like kernels section hijacked by blend-bros :). btw, have you noticed that Kaggle also doesn't like blends and have disabled `best score` feature? ",
          "votes": 9
        }
      ]
    },
    {
      "id": 315057,
      "postDate": "2018-04-16T17:41:24.913Z",
      "content": "<p>Even being <em>mostly</em> a lurker for the past few years, I tend to agree with a lot that was said here.</p>\n\n<p>Has a \"kernels deadline\" ever been considered? Say 2-4 weeks before deadine?</p>",
      "rawMarkdown": "Even being *mostly* a lurker for the past few years, I tend to agree with a lot that was said here.\n\nHas a \"kernels deadline\" ever been considered? Say 2-4 weeks before deadine?",
      "votes": 3,
      "replies": [
        {
          "id": 315170,
          "postDate": "2018-04-16T20:05:32.030Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 313543,
      "postDate": "2018-04-13T12:44:21.540Z",
      "content": "<p>Oh, please don't give my brain a rational reason to procrastinate until two weeks before the deadline :P</p>",
      "rawMarkdown": "Oh, please don't give my brain a rational reason to procrastinate until two weeks before the deadline :P",
      "votes": 3
    },
    {
      "id": 313392,
      "postDate": "2018-04-13T07:49:24.640Z",
      "content": "<p>Tbf, the new kernels utilizing the time to next click are based off of a kernel that came out quite some time ago (3 weeks ago). It was actually a discussion thread that pointed out that those features were valuable recently :)</p>",
      "rawMarkdown": "Tbf, the new kernels utilizing the time to next click are based off of a kernel that came out quite some time ago (3 weeks ago). It was actually a discussion thread that pointed out that those features were valuable recently :)",
      "votes": 2
    },
    {
      "id": 315013,
      "postDate": "2018-04-16T16:30:44.477Z",
      "content": "<p>For some top kagglers, they have no idea how bad people feel when they spend hours and hours working on original ideas and get screwed by some last minute kernels.  Imagine this, I spend 3 weeks to build a model scoring top 20%, I know I'm not that good. But I still feel I learn a lot and the ranking is fair. Then one night before the deadline, some top players publish a public kernel scoring 10%. If I checked the kernel section that day and decided to use it, I will be among tons of people sharing the same score - probably will rank 10-15% finally. And I will still feel that ranking doesn't mean anything to me - it's not mine! If I didn't then I would be seriously screwed by that kernel - a lot people having &lt;5 submissions will rank higher than me.  While this will have negligible effect on those top players, this type of “sharing\" destroys the fun and fairness for many. </p>\n\n<p>And seriously, are those last minute kernels really for learning/sharing ideas? </p>",
      "rawMarkdown": "For some top kagglers, they have no idea how bad people feel when they spend hours and hours working on original ideas and get screwed by some last minute kernels.  Imagine this, I spend 3 weeks to build a model scoring top 20%, I know I'm not that good. But I still feel I learn a lot and the ranking is fair. Then one night before the deadline, some top players publish a public kernel scoring 10%. If I checked the kernel section that day and decided to use it, I will be among tons of people sharing the same score - probably will rank 10-15% finally. And I will still feel that ranking doesn't mean anything to me - it's not mine! If I didn't then I would be seriously screwed by that kernel - a lot people having &lt;5 submissions will rank higher than me.  While this will have negligible effect on those top players, this type of “sharing\" destroys the fun and fairness for many. \n\nAnd seriously, are those last minute kernels really for learning/sharing ideas? \n",
      "votes": 1,
      "replies": [
        {
          "id": 315053,
          "postDate": "2018-04-16T17:30:57.870Z",
          "content": "<p>Unfortunately human nature is what it is and you will always have people who are lazy and game the system - we all know at least one person we work with that has those characteristics don't we? ;)</p>\n\n<p>I play kaggle to learn and view it the same way I do golf - the only person I am competing with is myself.</p>\n\n<p>Take heart that your hard work will be rewarded in the long run.  </p>\n\n<p>There was a time prior to kernels that encouraged employment agencies to correlate the final leaderboard and medals as an indicator of competence - this is not the case anymore!</p>",
          "rawMarkdown": "Unfortunately human nature is what it is and you will always have people who are lazy and game the system - we all know at least one person we work with that has those characteristics don't we? ;)\n\nI play kaggle to learn and view it the same way I do golf - the only person I am competing with is myself.\n\nTake heart that your hard work will be rewarded in the long run.  \n\nThere was a time prior to kernels that encouraged employment agencies to correlate the final leaderboard and medals as an indicator of competence - this is not the case anymore!",
          "votes": 2
        },
        {
          "id": 315055,
          "postDate": "2018-04-16T17:39:26.807Z",
          "content": "<p>Well, I agree that last-minute high-scoring kernels should be frowned upon, but:</p>\n\n<p>If your ideas are original, chances are you've developed something quite different from that hypothetical last minute leaderboard killer. So you can probably do better than most people by blending it with your own, or by incorporating your own ideas into it.</p>\n\n<p>I remember a few months ago when I was still young in Kaggle years, in the Instacart competition. I was headed for the bronze medal that would give my first expert status. Then a week or two before the end, somebody posted a discussion topic with a set of instructions for how to implement their solution. It wasn't nearly as easy as just running a kernel, but, judging by the leaderboard, a lot of people succeeded in implementing it. And I was one of them. And I made some changes to add my own ideas. It was a lot of work in a short period of time, but there was a lot of learning, and in the end I did manage to get that bronze.</p>",
          "rawMarkdown": "Well, I agree that last-minute high-scoring kernels should be frowned upon, but:\n\nIf your ideas are original, chances are you've developed something quite different from that hypothetical last minute leaderboard killer. So you can probably do better than most people by blending it with your own, or by incorporating your own ideas into it.\n\nI remember a few months ago when I was still young in Kaggle years, in the Instacart competition. I was headed for the bronze medal that would give my first expert status. Then a week or two before the end, somebody posted a discussion topic with a set of instructions for how to implement their solution. It wasn't nearly as easy as just running a kernel, but, judging by the leaderboard, a lot of people succeeded in implementing it. And I was one of them. And I made some changes to add my own ideas. It was a lot of work in a short period of time, but there was a lot of learning, and in the end I did manage to get that bronze.",
          "votes": 2
        },
        {
          "id": 315060,
          "postDate": "2018-04-16T17:47:38.180Z",
          "content": "<p>Also, most public kernels aren't adequately validated. If you've developed a validation framework, and you can fit the approach of a last minute public kernel into your framework, then you can get an idea of how much it's really going to help. You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score.</p>",
          "rawMarkdown": "Also, most public kernels aren't adequately validated. If you've developed a validation framework, and you can fit the approach of a last minute public kernel into your framework, then you can get an idea of how much it's really going to help. You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score."
        },
        {
          "id": 315065,
          "postDate": "2018-04-16T17:53:45.070Z",
          "content": "<p>@Scirpus, I'm not complaining about people taking use of those kernels - that's understandable. I think it's unfair and unnecessary to publish high ranking kernels in the last few days. If they are published earlier,  people have time to learn and merge the idea into their own solutions, which is the rightful purpose of sharing kernels. Or they can be published after the ddl, people will still read and learn from winners' solutions. Personally I think publishing kernels should be banned during the last 2 days or so. So the final ranking represents much more fairness. </p>\n\n<p>Again, I'm only complaining about publishing LAST MINUTE high ranking kernels. I think the only purpose of doing that is showing off instead of sharing knowledge. </p>",
          "rawMarkdown": "@Scirpus, I'm not complaining about people taking use of those kernels - that's understandable. I think it's unfair and unnecessary to publish high ranking kernels in the last few days. If they are published earlier,  people have time to learn and merge the idea into their own solutions, which is the rightful purpose of sharing kernels. Or they can be published after the ddl, people will still read and learn from winners' solutions. Personally I think publishing kernels should be banned during the last 2 days or so. So the final ranking represents much more fairness. \n\nAgain, I'm only complaining about publishing LAST MINUTE high ranking kernels. I think the only purpose of doing that is showing off instead of sharing knowledge. ",
          "votes": 3
        },
        {
          "id": 315072,
          "postDate": "2018-04-16T18:00:39.270Z",
          "content": "<p>@Andy,  I agree -even a week before ddl would be ok, people like you will still have enough time to digest and integrate the ideas. I'm more specifically concerned about those high ranking kernels published in the last 2 days or so.  I've seen this happening and I don't think the publisher intended to communicate good ideas. As you said, it can be poorly validated. But given its high score, many people will still use that solution (we have 2 final subs anyway). And when the competition ends, the private lb is flooded with the same scores and people originally in that range are screwed. </p>",
          "rawMarkdown": "@Andy,  I agree -even a week before ddl would be ok, people like you will still have enough time to digest and integrate the ideas. I'm more specifically concerned about those high ranking kernels published in the last 2 days or so.  I've seen this happening and I don't think the publisher intended to communicate good ideas. As you said, it can be poorly validated. But given its high score, many people will still use that solution (we have 2 final subs anyway). And when the competition ends, the private lb is flooded with the same scores and people originally in that range are screwed. ",
          "votes": 1
        },
        {
          "id": 315081,
          "postDate": "2018-04-16T18:18:26.093Z",
          "content": "<p>I think even with a week to go it can still be very unfair. It's easy to say that public kernels are always a level playing field, but that really isn't true toward the end of the competition because of time investment asymmetry. If someone spends months working on a competition but does not have time to work on it in the final week due to personal/work reasons, they can see most of their effort quickly overshadowed by the 50th place solution that suddenly gets open-sourced. Meanwhile, a second person can hop in during the last week with enough time to make some improvements to the strong baseline, earning a medal with a fraction of the effort and understanding of the problem that the first person had.</p>\n\n<p>In short, the timeline of public kernel release has a dramatic impact on <em>when</em> time investment is most rewarded. This is luck not meritocracy, and really can ruin some of the value of kaggle's gamification in my view. I'm all for public kernels and think they're immensely valuable for learning and knowledge sharing, but I also think it'd be much better if there were a kernels/code share deadline something like 2 weeks before competition deadlines.</p>",
          "rawMarkdown": "I think even with a week to go it can still be very unfair. It's easy to say that public kernels are always a level playing field, but that really isn't true toward the end of the competition because of time investment asymmetry. If someone spends months working on a competition but does not have time to work on it in the final week due to personal/work reasons, they can see most of their effort quickly overshadowed by the 50th place solution that suddenly gets open-sourced. Meanwhile, a second person can hop in during the last week with enough time to make some improvements to the strong baseline, earning a medal with a fraction of the effort and understanding of the problem that the first person had.\n\nIn short, the timeline of public kernel release has a dramatic impact on *when* time investment is most rewarded. This is luck not meritocracy, and really can ruin some of the value of kaggle's gamification in my view. I'm all for public kernels and think they're immensely valuable for learning and knowledge sharing, but I also think it'd be much better if there were a kernels/code share deadline something like 2 weeks before competition deadlines.",
          "votes": 2
        },
        {
          "id": 315094,
          "postDate": "2018-04-16T18:34:57.610Z",
          "content": "<p>@congb - I understand exactly what you mean - I was just giving you some solace for when it happens ;)</p>",
          "rawMarkdown": "@congb - I understand exactly what you mean - I was just giving you some solace for when it happens ;)"
        },
        {
          "id": 315120,
          "postDate": "2018-04-16T19:10:03.313Z",
          "content": "<p>A small change that might help would be to disallow submissions from public kernels after the team-up deadline. That would at least prevent new public kernels with verified public LB scores. And people who wanted to submit new versions of old public kernels could still fork and submit privately.</p>",
          "rawMarkdown": "A small change that might help would be to disallow submissions from public kernels after the team-up deadline. That would at least prevent new public kernels with verified public LB scores. And people who wanted to submit new versions of old public kernels could still fork and submit privately.",
          "votes": 2
        },
        {
          "id": 325092,
          "postDate": "2018-05-08T05:39:20.053Z",
          "content": "<blockquote>\n  <p><strong>Andy Harless wrote</strong></p>\n  \n  <blockquote>\n    <p>Also, most public kernels aren't adequately validated. ...You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score.</p>\n  </blockquote>\n</blockquote>\n\n<p>It would help if Kaggle recorded all of <strong>{training set CV, public + private test-set leaderboard}</strong> scores for kernels, so we could quantify the overfit after-the-fact.</p>",
          "rawMarkdown": "\n&gt; **Andy Harless wrote**\n&gt; \n&gt;&gt; Also, most public kernels aren't adequately validated. ...You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score.\n\nIt would help if Kaggle recorded all of **{training set CV, public + private test-set leaderboard}** scores for kernels, so we could quantify the overfit after-the-fact."
        }
      ]
    },
    {
      "id": 315667,
      "postDate": "2018-04-17T12:56:54.207Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 315172,
      "postDate": "2018-04-16T20:09:21.813Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 314856,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2018-04-16T11:23:44.610000",
      "content": "<p>I don't totally agree with comments below - there is a big difference between posting kernels with all preprocessing, hyperparams and training iterations to get silver just forking the code and on the other side to post a general idea or approach or hint to feed your brains (great example - <a href=\"https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769\">https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769</a>). </p>",
      "votes": 7,
      "replies": [
        {
          "id": 314864,
          "author_name": "NanoMathias",
          "author_url": "",
          "post_date": "2018-04-16T11:48:50.980000",
          "content": "<p>I'll just drop a quick thanks here, it's very motivational to submit something that's well-received :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 314886,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-04-16T12:37:49.800000",
          "content": "<p>I agree - great kernel - because it can be used in general, people can learn something that they can take forward.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 313340,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2018-04-13T06:02:08.657000",
      "content": "<p>Yes, ladies and gentleman please stop spawning new kernels, you annihilate our efforts:( Or at least make it less straightforward somehow </p>",
      "votes": 7,
      "replies": [
        {
          "id": 313611,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-13T15:09:26.380000",
          "content": "<p>I think the kernels are still just scratching the surface.  It's a very hard problem because of the size of the data.  If the goal is to find the best possible solutions, it makes sense to have some cooperation until we reach the later stages.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 313765,
          "author_name": "Kamil Zieliński",
          "author_url": "",
          "post_date": "2018-04-13T19:58:17.970000",
          "content": "<p>But this is more like doing random things, finding a good one and saying: see, THIS works! This is not an attempt to understand the problem, so it seems to have no learning value, it only flattens the leaderboard. And I believe magic features usually have some good explanation for their power e.g. the magic ip/os/device/app next_click feature from the latest kernel seems very reasonable, because ip/os/device combination is a pretty good definition of a user, so it is building on a very natural user/app combination.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 313773,
          "author_name": "bangda",
          "author_url": "",
          "post_date": "2018-04-13T20:19:48.430000",
          "content": "<p>I still remember that in Mercari Price Suggestion Competition, a kernel could rank 9% on lb was published the day before the last submission day ;), and I dropped from 9% to 15%+. Therefore for me I will keep an eye on kernel section all the time lol.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 313787,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-13T20:42:43.470000",
          "content": "<p>@Kamil There are a lot of things that seem like reasonable possibilities. Most of the features I've seen in kernels are pretty straightforward, not weird random ideas. But testing which ones work is difficult given the size of the data.  A lot of machine learning involves trying things out to see what works.  Public kernels judged by public LB scores may not be the ideal framework for this, but they're better than nothing (provided one does proper validation).  There is some learning value in learning what works, but the main learning value here is in learning to wrestle with the data.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 313846,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-13T23:51:37.470000",
          "content": "<p>(You say \"flattens the leaderboard.\" I say \"sets new benchmarks.\")</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 314006,
          "author_name": "Luis Moneda",
          "author_url": "",
          "post_date": "2018-04-14T10:53:33.420000",
          "content": "<blockquote>\n  <p>Therefore for me I will keep an eye on kernel section all the time lol.</p>\n</blockquote>\n\n<p>That's the kind of thing that looks obvious after seeing the actual results of a competition. But when it stills running, it's very hard to know if a kernel mixing 30 other ones is really good or just lucky (we know that ensembles are good, but a small difference by chance can make a particular one looks really good). We end up looking for the winner ones and thinking it'll be easy to identify them in the next competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314090,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-14T15:27:26.257000",
          "content": "<p>To be honest, kernels have short life. In any competition, ideas evolve over the period of time and techniques get changed in positive direction if there's a co-operation in terms of feedback, suggestion or a forked version of kernel that shows how it can be improved.  Completely agree that kernel popping out in the last week to destroy leaderboard is really bad.</p>\n\n<p>With regard to this competition, I can still remember the how public LB evolved over the time. (exclude blends):</p>\n\n<pre><code>LB ~ 0.9550 getting started kernels\nLB ~ 0.9631 scale_pos_weight came into the light\nLB ~ 0.9680 Memory optimized kernels \nLB ~ 0.9730 Kernel with user ranking features\nLB ~ 0.9752 Innovative modelling approach i.e FTRL and time delta feats\n</code></pre>\n\n<p>And as on today, we still have 3 more weeks to competition deadline. I am sure that most people don't like kernels section hijacked by blend-bros :). btw, have you noticed that Kaggle also doesn't like blends and have disabled <code>best score</code> feature? </p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 315057,
      "author_name": "F Bertrand",
      "author_url": "",
      "post_date": "2018-04-16T17:41:24.913000",
      "content": "<p>Even being <em>mostly</em> a lurker for the past few years, I tend to agree with a lot that was said here.</p>\n\n<p>Has a \"kernels deadline\" ever been considered? Say 2-4 weeks before deadine?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 315170,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-16T20:05:32.030000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 313543,
      "author_name": "Luis Moneda",
      "author_url": "",
      "post_date": "2018-04-13T12:44:21.540000",
      "content": "<p>Oh, please don't give my brain a rational reason to procrastinate until two weeks before the deadline :P</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 313392,
      "author_name": "MitchelFung",
      "author_url": "",
      "post_date": "2018-04-13T07:49:24.640000",
      "content": "<p>Tbf, the new kernels utilizing the time to next click are based off of a kernel that came out quite some time ago (3 weeks ago). It was actually a discussion thread that pointed out that those features were valuable recently :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 315013,
      "author_name": "Cong",
      "author_url": "",
      "post_date": "2018-04-16T16:30:44.477000",
      "content": "<p>For some top kagglers, they have no idea how bad people feel when they spend hours and hours working on original ideas and get screwed by some last minute kernels.  Imagine this, I spend 3 weeks to build a model scoring top 20%, I know I'm not that good. But I still feel I learn a lot and the ranking is fair. Then one night before the deadline, some top players publish a public kernel scoring 10%. If I checked the kernel section that day and decided to use it, I will be among tons of people sharing the same score - probably will rank 10-15% finally. And I will still feel that ranking doesn't mean anything to me - it's not mine! If I didn't then I would be seriously screwed by that kernel - a lot people having &lt;5 submissions will rank higher than me.  While this will have negligible effect on those top players, this type of “sharing\" destroys the fun and fairness for many. </p>\n\n<p>And seriously, are those last minute kernels really for learning/sharing ideas? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 315053,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-04-16T17:30:57.870000",
          "content": "<p>Unfortunately human nature is what it is and you will always have people who are lazy and game the system - we all know at least one person we work with that has those characteristics don't we? ;)</p>\n\n<p>I play kaggle to learn and view it the same way I do golf - the only person I am competing with is myself.</p>\n\n<p>Take heart that your hard work will be rewarded in the long run.  </p>\n\n<p>There was a time prior to kernels that encouraged employment agencies to correlate the final leaderboard and medals as an indicator of competence - this is not the case anymore!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 315055,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-16T17:39:26.807000",
          "content": "<p>Well, I agree that last-minute high-scoring kernels should be frowned upon, but:</p>\n\n<p>If your ideas are original, chances are you've developed something quite different from that hypothetical last minute leaderboard killer. So you can probably do better than most people by blending it with your own, or by incorporating your own ideas into it.</p>\n\n<p>I remember a few months ago when I was still young in Kaggle years, in the Instacart competition. I was headed for the bronze medal that would give my first expert status. Then a week or two before the end, somebody posted a discussion topic with a set of instructions for how to implement their solution. It wasn't nearly as easy as just running a kernel, but, judging by the leaderboard, a lot of people succeeded in implementing it. And I was one of them. And I made some changes to add my own ideas. It was a lot of work in a short period of time, but there was a lot of learning, and in the end I did manage to get that bronze.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 315060,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-16T17:47:38.180000",
          "content": "<p>Also, most public kernels aren't adequately validated. If you've developed a validation framework, and you can fit the approach of a last minute public kernel into your framework, then you can get an idea of how much it's really going to help. You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315065,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-16T17:53:45.070000",
          "content": "<p>@Scirpus, I'm not complaining about people taking use of those kernels - that's understandable. I think it's unfair and unnecessary to publish high ranking kernels in the last few days. If they are published earlier,  people have time to learn and merge the idea into their own solutions, which is the rightful purpose of sharing kernels. Or they can be published after the ddl, people will still read and learn from winners' solutions. Personally I think publishing kernels should be banned during the last 2 days or so. So the final ranking represents much more fairness. </p>\n\n<p>Again, I'm only complaining about publishing LAST MINUTE high ranking kernels. I think the only purpose of doing that is showing off instead of sharing knowledge. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 315072,
          "author_name": "Cong",
          "author_url": "",
          "post_date": "2018-04-16T18:00:39.270000",
          "content": "<p>@Andy,  I agree -even a week before ddl would be ok, people like you will still have enough time to digest and integrate the ideas. I'm more specifically concerned about those high ranking kernels published in the last 2 days or so.  I've seen this happening and I don't think the publisher intended to communicate good ideas. As you said, it can be poorly validated. But given its high score, many people will still use that solution (we have 2 final subs anyway). And when the competition ends, the private lb is flooded with the same scores and people originally in that range are screwed. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 315081,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-04-16T18:18:26.093000",
          "content": "<p>I think even with a week to go it can still be very unfair. It's easy to say that public kernels are always a level playing field, but that really isn't true toward the end of the competition because of time investment asymmetry. If someone spends months working on a competition but does not have time to work on it in the final week due to personal/work reasons, they can see most of their effort quickly overshadowed by the 50th place solution that suddenly gets open-sourced. Meanwhile, a second person can hop in during the last week with enough time to make some improvements to the strong baseline, earning a medal with a fraction of the effort and understanding of the problem that the first person had.</p>\n\n<p>In short, the timeline of public kernel release has a dramatic impact on <em>when</em> time investment is most rewarded. This is luck not meritocracy, and really can ruin some of the value of kaggle's gamification in my view. I'm all for public kernels and think they're immensely valuable for learning and knowledge sharing, but I also think it'd be much better if there were a kernels/code share deadline something like 2 weeks before competition deadlines.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 315094,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-04-16T18:34:57.610000",
          "content": "<p>@congb - I understand exactly what you mean - I was just giving you some solace for when it happens ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 315120,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-16T19:10:03.313000",
          "content": "<p>A small change that might help would be to disallow submissions from public kernels after the team-up deadline. That would at least prevent new public kernels with verified public LB scores. And people who wanted to submit new versions of old public kernels could still fork and submit privately.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 325092,
          "author_name": "Stephen McInerney",
          "author_url": "",
          "post_date": "2018-05-08T05:39:20.053000",
          "content": "<blockquote>\n  <p><strong>Andy Harless wrote</strong></p>\n  \n  <blockquote>\n    <p>Also, most public kernels aren't adequately validated. ...You probably would want to blend your results with it anyway, but you don't necessarily want to give it most of the weight. And those who do aren't necessarily helping their private leaderboard score.</p>\n  </blockquote>\n</blockquote>\n\n<p>It would help if Kaggle recorded all of <strong>{training set CV, public + private test-set leaderboard}</strong> scores for kernels, so we could quantify the overfit after-the-fact.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 315667,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-17T12:56:54.207000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 315172,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-16T20:09:21.813000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "313127": "Anyone else going to join competitions just two weeks before the deadline to invest the least possible effort to get the maximum return LOL ;)",
    "314856": "I don't totally agree with comments below - there is a big difference between posting kernels with all preprocessing, hyperparams and training iterations to get silver just forking the code and on the other side to post a general idea or approach or hint to feed your brains (great example - https://www.kaggle.com/nanomathias/bayesian-tuning-of-xgboost-lightgbm-lb-0-9769). ",
    "313340": "Yes, ladies and gentleman please stop spawning new kernels, you annihilate our efforts:( Or at least make it less straightforward somehow ",
    "315057": "Even being *mostly* a lurker for the past few years, I tend to agree with a lot that was said here.\n\nHas a \"kernels deadline\" ever been considered? Say 2-4 weeks before deadine?",
    "313543": "Oh, please don't give my brain a rational reason to procrastinate until two weeks before the deadline :P",
    "313392": "Tbf, the new kernels utilizing the time to next click are based off of a kernel that came out quite some time ago (3 weeks ago). It was actually a discussion thread that pointed out that those features were valuable recently :)",
    "315013": "For some top kagglers, they have no idea how bad people feel when they spend hours and hours working on original ideas and get screwed by some last minute kernels.  Imagine this, I spend 3 weeks to build a model scoring top 20%, I know I'm not that good. But I still feel I learn a lot and the ranking is fair. Then one night before the deadline, some top players publish a public kernel scoring 10%. If I checked the kernel section that day and decided to use it, I will be among tons of people sharing the same score - probably will rank 10-15% finally. And I will still feel that ranking doesn't mean anything to me - it's not mine! If I didn't then I would be seriously screwed by that kernel - a lot people having &lt;5 submissions will rank higher than me.  While this will have negligible effect on those top players, this type of “sharing\" destroys the fun and fairness for many. \n\nAnd seriously, are those last minute kernels really for learning/sharing ideas? \n",
    "315667": "",
    "315172": ""
  }
}