{
  "id": 76792,
  "title": "Kernels' overabuse has started to annoy me",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76792",
  "author_name": "Georgios Sarantitis",
  "post_date": "2019-01-06T20:33:52.427000",
  "votes": 12,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Let's be straight. There are people here who just fork the latest and better (in terms of LB score) kernels and just climb up the ladder. This is just futile and pointless. Its a strategy thats never gonna get you to the top or teach you anything  on the long run. It also has a side-effect for other people: it discourages the rest of us who are trying hard to make a good submission after struggling with data. </p>\n\n<p>This particular competition is illustrative of this point once again: it is very difficult to gain just a few positions and yet there are people in novice level who with 1-2 submissions are currently within medals (not that they' ll stay there but anywayz). Just saying...</p>",
  "messages": [
    {
      "id": 451632,
      "postDate": "2019-01-07T11:59:41.923Z",
      "content": "<p>I'm sorry but this makes no sense to me.</p>\n\n<ul>\n<li>You and the \"ladder climbers\" have equal access to public kernels, the only different, according to you, is that you put actually put effort into competing. If that's true then by the end there's no way they can finish above you, unless you refuse to use public knowledge for some reasons.</li>\n<li>The \"ladder climbers\" discouraging genuine competitors: Why? They fork a kernel, rerun and get the same score, you instead try to improve upon that/ use that to improve your own kernel, why do you care if other people are wasting their time?</li>\n<li>People in novice level being at the top: Ranks are irrelevant, and I don't know what it has to do with your previous points. </li>\n</ul>\n\n<p>There is one thing that I dislike about kernels though, it is people publishing high scoring ones at the very late stage of a competition, but that's less about kernels and more about people's unethical behavior.</p>",
      "rawMarkdown": "I'm sorry but this makes no sense to me.\n\n- You and the \"ladder climbers\" have equal access to public kernels, the only different, according to you, is that you put actually put effort into competing. If that's true then by the end there's no way they can finish above you, unless you refuse to use public knowledge for some reasons.\n- The \"ladder climbers\" discouraging genuine competitors: Why? They fork a kernel, rerun and get the same score, you instead try to improve upon that/ use that to improve your own kernel, why do you care if other people are wasting their time?\n- People in novice level being at the top: Ranks are irrelevant, and I don't know what it has to do with your previous points. \n\nThere is one thing that I dislike about kernels though, it is people publishing high scoring ones at the very late stage of a competition, but that's less about kernels and more about people's unethical behavior.",
      "votes": 29,
      "replies": [
        {
          "id": 451648,
          "postDate": "2019-01-07T12:37:17.853Z",
          "content": "<p>Well, the discouraging part comes when I fall to bed to wake up 300 positions below due to an over-forked last night's kernel with lower score. I don't know if you don't see it that way, for me its heartbreaking... And when that happens really late in the competition we come to your last point where the damage being made is irreversible due to lack of time...</p>\n\n<p>But to set things straight I am not against novices taking high places, even winning competitions, not at all. I have been there and I am not exactly at the top either. But some submissions you can really know simply don't deserve it...</p>",
          "rawMarkdown": "Well, the discouraging part comes when I fall to bed to wake up 300 positions below due to an over-forked last night's kernel with lower score. I don't know if you don't see it that way, for me its heartbreaking... And when that happens really late in the competition we come to your last point where the damage being made is irreversible due to lack of time...\n\nBut to set things straight I am not against novices taking high places, even winning competitions, not at all. I have been there and I am not exactly at the top either. But some submissions you can really know simply don't deserve it...",
          "votes": 1
        },
        {
          "id": 451650,
          "postDate": "2019-01-07T12:41:32.663Z",
          "content": "<p>A few days ago, I saw one public kernel with 0.699...but it only existed one or two hours, haha...</p>",
          "rawMarkdown": "A few days ago, I saw one public kernel with 0.699...but it only existed one or two hours, haha...",
          "votes": 5
        }
      ]
    },
    {
      "id": 451313,
      "postDate": "2019-01-06T20:33:52.427Z",
      "content": "<p>Let's be straight. There are people here who just fork the latest and better (in terms of LB score) kernels and just climb up the ladder. This is just futile and pointless. Its a strategy thats never gonna get you to the top or teach you anything  on the long run. It also has a side-effect for other people: it discourages the rest of us who are trying hard to make a good submission after struggling with data. </p>\n\n<p>This particular competition is illustrative of this point once again: it is very difficult to gain just a few positions and yet there are people in novice level who with 1-2 submissions are currently within medals (not that they' ll stay there but anywayz). Just saying...</p>",
      "rawMarkdown": "Let's be straight. There are people here who just fork the latest and better (in terms of LB score) kernels and just climb up the ladder. This is just futile and pointless. Its a strategy thats never gonna get you to the top or teach you anything  on the long run. It also has a side-effect for other people: it discourages the rest of us who are trying hard to make a good submission after struggling with data. \n\nThis particular competition is illustrative of this point once again: it is very difficult to gain just a few positions and yet there are people in novice level who with 1-2 submissions are currently within medals (not that they' ll stay there but anywayz). Just saying...",
      "votes": 12
    },
    {
      "id": 451723,
      "postDate": "2019-01-07T15:05:37.773Z",
      "content": "<p>Kaggle officially warns about late kernel sharing : </p>\n\n<p><a href=\"https://www.kaggle.com/spoiler-alert\">https://www.kaggle.com/spoiler-alert</a></p>",
      "rawMarkdown": "Kaggle officially warns about late kernel sharing : \n\nhttps://www.kaggle.com/spoiler-alert",
      "votes": 5,
      "replies": [
        {
          "id": 451753,
          "postDate": "2019-01-07T16:00:27.920Z",
          "content": "<p>Nice bro :-)</p>",
          "rawMarkdown": "Nice bro :-)",
          "votes": 2
        }
      ]
    },
    {
      "id": 452529,
      "postDate": "2019-01-08T21:17:06.163Z",
      "content": "<p>Personally, I see the kernels (especially with a good score) as a baseline which everyone has access to. So assuming everyone now has the knowledge from that kernel, I try to learn as many new things from different kernels and combine that knowledge to push my score beyond the kernel score. Its more about - if there is a kernel, fork it, learn the new and good ideas in it and use those along with your own to climb up further. The novices who just fork and submit will never be able to go up in the long run. </p>",
      "rawMarkdown": "Personally, I see the kernels (especially with a good score) as a baseline which everyone has access to. So assuming everyone now has the knowledge from that kernel, I try to learn as many new things from different kernels and combine that knowledge to push my score beyond the kernel score. Its more about - if there is a kernel, fork it, learn the new and good ideas in it and use those along with your own to climb up further. The novices who just fork and submit will never be able to go up in the long run. ",
      "votes": 6
    },
    {
      "id": 451319,
      "postDate": "2019-01-06T20:55:58.143Z",
      "content": "<p>What's the problem? That is open research: \"Err and err and err, but less and less and less\". I learn more if I see some ideas from others, read up on them and try to think about them and incorporate them. You won't get into top ranks by just forking the best kernels.</p>",
      "rawMarkdown": "What's the problem? That is open research: \"Err and err and err, but less and less and less\". I learn more if I see some ideas from others, read up on them and try to think about them and incorporate them. You won't get into top ranks by just forking the best kernels.",
      "votes": 6,
      "replies": [
        {
          "id": 451385,
          "postDate": "2019-01-07T02:02:56.003Z",
          "content": "<p>Exactly, forking other Kernel's is useful to learn. That's why people love kaggle.</p>",
          "rawMarkdown": "Exactly, forking other Kernel's is useful to learn. That's why people love kaggle.",
          "votes": 5
        }
      ]
    },
    {
      "id": 456199,
      "postDate": "2019-01-15T10:34:51.053Z",
      "content": "<p>Maybe Kaggle can open an NLP-type competition to detect very similar kernels and issue a separate ranking for only unique kernels :D that'd weed out the forkers. </p>",
      "rawMarkdown": "Maybe Kaggle can open an NLP-type competition to detect very similar kernels and issue a separate ranking for only unique kernels :D that'd weed out the forkers. ",
      "votes": 3
    },
    {
      "id": 451687,
      "postDate": "2019-01-07T13:45:46.433Z",
      "content": "<p>It would be a shame if people did not share their good kernels for others to learn. At least for me this is perhaps the best part of Kaggle, to learn from others.</p>\n\n<p>But I see your issue too, the competition is for many people about gaming the system, and in that way one has to do the same unless really a top contender. I don't generally fork and submit, but then I generally don't make the medals :). </p>\n\n<p>Hard to balance, I guess. But I like people sharing rather than discouraging that.</p>",
      "rawMarkdown": "It would be a shame if people did not share their good kernels for others to learn. At least for me this is perhaps the best part of Kaggle, to learn from others.\n\nBut I see your issue too, the competition is for many people about gaming the system, and in that way one has to do the same unless really a top contender. I don't generally fork and submit, but then I generally don't make the medals :). \n\nHard to balance, I guess. But I like people sharing rather than discouraging that.",
      "votes": 4
    },
    {
      "id": 456569,
      "postDate": "2019-01-16T04:13:03.117Z",
      "content": "<p>While it is never a great feeling to see your score plunge due to people simply copying and pasting a Kernal, it is certainly better than having a data leak. Additionally, those who may have designed their own Kernal from scratch, but have a lower score should be able to quickly modify a higher performing Kernal to score higher than the copy pasters.  </p>\n\n<p>Generally a company would want to have the best algorithm as well and they would prefer someone who can modify the best performing \"off the shelf\" algorithm than spend time creating one from scratch that does not perform as well.  So these competitions are not much different than life...  I am just glad to see people sharing so I can learn some techniques from others. </p>\n\n<p>However, I would like to see some kind of bonus point structure or otherwise for the submission of code that could run in production. </p>",
      "rawMarkdown": "While it is never a great feeling to see your score plunge due to people simply copying and pasting a Kernal, it is certainly better than having a data leak. Additionally, those who may have designed their own Kernal from scratch, but have a lower score should be able to quickly modify a higher performing Kernal to score higher than the copy pasters.  \n\nGenerally a company would want to have the best algorithm as well and they would prefer someone who can modify the best performing \"off the shelf\" algorithm than spend time creating one from scratch that does not perform as well.  So these competitions are not much different than life...  I am just glad to see people sharing so I can learn some techniques from others. \n\nHowever, I would like to see some kind of bonus point structure or otherwise for the submission of code that could run in production. ",
      "votes": 2
    },
    {
      "id": 457398,
      "postDate": "2019-01-17T10:51:28.223Z",
      "content": "<p>hi  if u see the new kernel with 0.699 score but without any sense you will be angry</p>",
      "rawMarkdown": "hi  if u see the new kernel with 0.699 score but without any sense you will be angry"
    },
    {
      "id": 456368,
      "postDate": "2019-01-15T17:19:04.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 455598,
      "postDate": "2019-01-14T08:24:20.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 451326,
      "postDate": "2019-01-06T21:17:50.510Z",
      "content": "<p>... and it gets even worse:\nI guess there are also people that download submission files from top public kernels, combine by average or voting, copy paste into a notebook and submit. Maybe some of the very high scores with very few submissions were obtained with this method.</p>\n\n<p>Let's be optimistic, winning this competition has a lot to do with getting a stable model (that will output similar probabilities when retrained in phase II) and a robust 'threshold finding algorithm'. </p>\n\n<p>The test.csv is rather small anyway and thus LB scores have high variance. \nMy gut-feeling is that LB f1 variations smaller that 0.005 are irrelevant.\nSo I personalty rely on my own hold-out set.</p>\n\n<p>By the way people novice in Kaggle are not necessarily novice to ML ;-)</p>",
      "rawMarkdown": "... and it gets even worse:\nI guess there are also people that download submission files from top public kernels, combine by average or voting, copy paste into a notebook and submit. Maybe some of the very high scores with very few submissions were obtained with this method.\n\nLet's be optimistic, winning this competition has a lot to do with getting a stable model (that will output similar probabilities when retrained in phase II) and a robust 'threshold finding algorithm'. \n\nThe test.csv is rather small anyway and thus LB scores have high variance. \nMy gut-feeling is that LB f1 variations smaller that 0.005 are irrelevant.\nSo I personalty rely on my own hold-out set.\n\nBy the way people novice in Kaggle are not necessarily novice to ML ;-)",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 451334,
          "postDate": "2019-01-06T21:58:45.347Z",
          "content": "<p>Won't those scores be thrown out in private LB anyways?</p>",
          "rawMarkdown": "Won't those scores be thrown out in private LB anyways?",
          "votes": 1
        },
        {
          "id": 451344,
          "postDate": "2019-01-06T22:36:23.540Z",
          "content": "<p>Private LB will be a much larger sample (376k rows or 320k, I'm still not sure) so I expect the 'lucky values' will be attenuated.</p>",
          "rawMarkdown": "Private LB will be a much larger sample (376k rows or 320k, I'm still not sure) so I expect the 'lucky values' will be attenuated.",
          "isDeleted": true
        },
        {
          "id": 451383,
          "postDate": "2019-01-07T02:00:55.730Z",
          "content": "<p>It is not possible to download and load other submissions, as it is considered as external data.</p>",
          "rawMarkdown": "It is not possible to download and load other submissions, as it is considered as external data.",
          "votes": 1
        },
        {
          "id": 451846,
          "postDate": "2019-01-07T18:56:18.340Z",
          "content": "<p>You can't load the data itself onto the actual kernel, but you could still download the submissions onto your own computer from public kernels. Then combine them all using some kind of average (not on a kaggle kernel, somewhere else). After that, you copy/paste those values into a kaggle kernel and create your submission file using those values. Even though you are using external data, the kaggle kernel won't know this. So what Annabelle suggested is very possible. And since Public LB will remain the same for this competition, it doesn't matter to them that this method is useless for the private test set.</p>",
          "rawMarkdown": "You can't load the data itself onto the actual kernel, but you could still download the submissions onto your own computer from public kernels. Then combine them all using some kind of average (not on a kaggle kernel, somewhere else). After that, you copy/paste those values into a kaggle kernel and create your submission file using those values. Even though you are using external data, the kaggle kernel won't know this. So what Annabelle suggested is very possible. And since Public LB will remain the same for this competition, it doesn't matter to them that this method is useless for the private test set."
        },
        {
          "id": 451859,
          "postDate": "2019-01-07T19:29:37.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/smokerx/test-others\">https://www.kaggle.com/smokerx/test-others</a>\nHere is an example of what you could do. This particular person said that they only did it because the Kaggle kernels were taking too long to submit, but it does illustrate how you can use external data using copy/paste.</p>",
          "rawMarkdown": "https://www.kaggle.com/smokerx/test-others\nHere is an example of what you could do. This particular person said that they only did it because the Kaggle kernels were taking too long to submit, but it does illustrate how you can use external data using copy/paste."
        },
        {
          "id": 451870,
          "postDate": "2019-01-07T19:58:13.677Z",
          "content": "<p>Exactly.\nThat is why I am curious to know if private LB score will be based on 320k totally unseen rows of phase 2 test set or on all 376k rows of which 56k can be downloaded and predictions obtained by combining multiple historical predictions. </p>\n\n<p>I posted a question to <a href=\"/inversion\">@inversion</a> staff but got no reply yet</p>",
          "rawMarkdown": "Exactly.\nThat is why I am curious to know if private LB score will be based on 320k totally unseen rows of phase 2 test set or on all 376k rows of which 56k can be downloaded and predictions obtained by combining multiple historical predictions. \n\nI posted a question to @inversion staff but got no reply yet",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 451902,
          "postDate": "2019-01-07T21:47:56.493Z",
          "content": "<p>Hi @Annablle - With the holidays and some issues getting notifications, I missed your question.</p>\n\n<p>To be very clear about this:\n1. The Public Test is ~56k lines (phase 1)\n2. The phase 2 Test is ~376k lines\n3. The phase 2 Test includes Pubic <em>and</em> Private\n4. The Private rows will be used for the final Leaderboard\n5. The Public rows are used so that when we re-score, the Public Leaderboard isn't lost.</p>",
          "rawMarkdown": "Hi @Annablle - With the holidays and some issues getting notifications, I missed your question.\n\nTo be very clear about this:\n1. The Public Test is ~56k lines (phase 1)\n2. The phase 2 Test is ~376k lines\n3. The phase 2 Test includes Pubic *and* Private\n4. The Private rows will be used for the final Leaderboard\n5. The Public rows are used so that when we re-score, the Public Leaderboard isn't lost.",
          "votes": 3
        },
        {
          "id": 451926,
          "postDate": "2019-01-07T22:55:33.353Z",
          "content": "<p>Thanks for the reply <a href=\"/inversion\">@inversion</a>.\nThat is good news, it means that it is not possible to hand-label the ~56k lines of phase 1 and copy-paste these labels into a notebook to boost score of final private LB.</p>",
          "rawMarkdown": "Thanks for the reply @inversion.\nThat is good news, it means that it is not possible to hand-label the ~56k lines of phase 1 and copy-paste these labels into a notebook to boost score of final private LB.",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 452470,
          "postDate": "2019-01-08T19:09:26.757Z",
          "content": "<p>Hi <a href=\"/inversion\">@inversion</a> </p>\n\n<p>When is the time that phase 2 data will be loaded? Went to the timeline page and couldn't see anything there. </p>\n\n<p>Will we be able to run our kernels on the phase 2 data and be able to select which kernels we want to submit for the final submission?</p>",
          "rawMarkdown": "Hi @inversion \n\nWhen is the time that phase 2 data will be loaded? Went to the timeline page and couldn't see anything there. \n\nWill we be able to run our kernels on the phase 2 data and be able to select which kernels we want to submit for the final submission?"
        },
        {
          "id": 453688,
          "postDate": "2019-01-10T16:00:31.317Z",
          "content": "<p>Hi Rahul,\nthe phase 2 data will not be publicly visible prior to competition end. The kernels will be run on new datasets with no user interaction. We just have to prepare good kernels :-)</p>",
          "rawMarkdown": "Hi Rahul,\nthe phase 2 data will not be publicly visible prior to competition end. The kernels will be run on new datasets with no user interaction. We just have to prepare good kernels :-)"
        },
        {
          "id": 456617,
          "postDate": "2019-01-16T07:44:22.530Z",
          "content": "<p><a href=\"/inversion\">@inversion</a>\nHello! If I have a kernel, it runs 7200s in Public Test(~56k lines). However, it runs 7400s in phase 2 Test (~376k lines). Is that means the kernel will fail in the final standing?\nThanks!</p>",
          "rawMarkdown": "@inversion\nHello! If I have a kernel, it runs 7200s in Public Test(~56k lines). However, it runs 7400s in phase 2 Test (~376k lines). Is that means the kernel will fail in the final standing?\nThanks!\n",
          "votes": 2
        },
        {
          "id": 457438,
          "postDate": "2019-01-17T12:30:53.240Z",
          "content": "<p>Yes, read the guidelines posted.</p>",
          "rawMarkdown": "Yes, read the guidelines posted."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 451632,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2019-01-07T11:59:41.923000",
      "content": "<p>I'm sorry but this makes no sense to me.</p>\n\n<ul>\n<li>You and the \"ladder climbers\" have equal access to public kernels, the only different, according to you, is that you put actually put effort into competing. If that's true then by the end there's no way they can finish above you, unless you refuse to use public knowledge for some reasons.</li>\n<li>The \"ladder climbers\" discouraging genuine competitors: Why? They fork a kernel, rerun and get the same score, you instead try to improve upon that/ use that to improve your own kernel, why do you care if other people are wasting their time?</li>\n<li>People in novice level being at the top: Ranks are irrelevant, and I don't know what it has to do with your previous points. </li>\n</ul>\n\n<p>There is one thing that I dislike about kernels though, it is people publishing high scoring ones at the very late stage of a competition, but that's less about kernels and more about people's unethical behavior.</p>",
      "votes": 29,
      "replies": [
        {
          "id": 451648,
          "author_name": "Georgios Sarantitis",
          "author_url": "",
          "post_date": "2019-01-07T12:37:17.853000",
          "content": "<p>Well, the discouraging part comes when I fall to bed to wake up 300 positions below due to an over-forked last night's kernel with lower score. I don't know if you don't see it that way, for me its heartbreaking... And when that happens really late in the competition we come to your last point where the damage being made is irreversible due to lack of time...</p>\n\n<p>But to set things straight I am not against novices taking high places, even winning competitions, not at all. I have been there and I am not exactly at the top either. But some submissions you can really know simply don't deserve it...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451650,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2019-01-07T12:41:32.663000",
          "content": "<p>A few days ago, I saw one public kernel with 0.699...but it only existed one or two hours, haha...</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 451723,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-01-07T15:05:37.773000",
      "content": "<p>Kaggle officially warns about late kernel sharing : </p>\n\n<p><a href=\"https://www.kaggle.com/spoiler-alert\">https://www.kaggle.com/spoiler-alert</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 451753,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-07T16:00:27.920000",
          "content": "<p>Nice bro :-)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 452529,
      "author_name": "Aditya Vyas",
      "author_url": "",
      "post_date": "2019-01-08T21:17:06.163000",
      "content": "<p>Personally, I see the kernels (especially with a good score) as a baseline which everyone has access to. So assuming everyone now has the knowledge from that kernel, I try to learn as many new things from different kernels and combine that knowledge to push my score beyond the kernel score. Its more about - if there is a kernel, fork it, learn the new and good ideas in it and use those along with your own to climb up further. The novices who just fork and submit will never be able to go up in the long run. </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 451319,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-01-06T20:55:58.143000",
      "content": "<p>What's the problem? That is open research: \"Err and err and err, but less and less and less\". I learn more if I see some ideas from others, read up on them and try to think about them and incorporate them. You won't get into top ranks by just forking the best kernels.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 451385,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-07T02:02:56.003000",
          "content": "<p>Exactly, forking other Kernel's is useful to learn. That's why people love kaggle.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 456199,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2019-01-15T10:34:51.053000",
      "content": "<p>Maybe Kaggle can open an NLP-type competition to detect very similar kernels and issue a separate ranking for only unique kernels :D that'd weed out the forkers. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 451687,
      "author_name": "averagemn",
      "author_url": "",
      "post_date": "2019-01-07T13:45:46.433000",
      "content": "<p>It would be a shame if people did not share their good kernels for others to learn. At least for me this is perhaps the best part of Kaggle, to learn from others.</p>\n\n<p>But I see your issue too, the competition is for many people about gaming the system, and in that way one has to do the same unless really a top contender. I don't generally fork and submit, but then I generally don't make the medals :). </p>\n\n<p>Hard to balance, I guess. But I like people sharing rather than discouraging that.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 456569,
      "author_name": "PaulSave",
      "author_url": "",
      "post_date": "2019-01-16T04:13:03.117000",
      "content": "<p>While it is never a great feeling to see your score plunge due to people simply copying and pasting a Kernal, it is certainly better than having a data leak. Additionally, those who may have designed their own Kernal from scratch, but have a lower score should be able to quickly modify a higher performing Kernal to score higher than the copy pasters.  </p>\n\n<p>Generally a company would want to have the best algorithm as well and they would prefer someone who can modify the best performing \"off the shelf\" algorithm than spend time creating one from scratch that does not perform as well.  So these competitions are not much different than life...  I am just glad to see people sharing so I can learn some techniques from others. </p>\n\n<p>However, I would like to see some kind of bonus point structure or otherwise for the submission of code that could run in production. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 457398,
      "author_name": "young_boy",
      "author_url": "",
      "post_date": "2019-01-17T10:51:28.223000",
      "content": "<p>hi  if u see the new kernel with 0.699 score but without any sense you will be angry</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 456368,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-15T17:19:04.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 455598,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-14T08:24:20.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451326,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-06T21:17:50.510000",
      "content": "<p>... and it gets even worse:\nI guess there are also people that download submission files from top public kernels, combine by average or voting, copy paste into a notebook and submit. Maybe some of the very high scores with very few submissions were obtained with this method.</p>\n\n<p>Let's be optimistic, winning this competition has a lot to do with getting a stable model (that will output similar probabilities when retrained in phase II) and a robust 'threshold finding algorithm'. </p>\n\n<p>The test.csv is rather small anyway and thus LB scores have high variance. \nMy gut-feeling is that LB f1 variations smaller that 0.005 are irrelevant.\nSo I personalty rely on my own hold-out set.</p>\n\n<p>By the way people novice in Kaggle are not necessarily novice to ML ;-)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 451334,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-06T21:58:45.347000",
          "content": "<p>Won't those scores be thrown out in private LB anyways?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451344,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-06T22:36:23.540000",
          "content": "<p>Private LB will be a much larger sample (376k rows or 320k, I'm still not sure) so I expect the 'lucky values' will be attenuated.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451383,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-07T02:00:55.730000",
          "content": "<p>It is not possible to download and load other submissions, as it is considered as external data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451846,
          "author_name": "Julius Hochmuth",
          "author_url": "",
          "post_date": "2019-01-07T18:56:18.340000",
          "content": "<p>You can't load the data itself onto the actual kernel, but you could still download the submissions onto your own computer from public kernels. Then combine them all using some kind of average (not on a kaggle kernel, somewhere else). After that, you copy/paste those values into a kaggle kernel and create your submission file using those values. Even though you are using external data, the kaggle kernel won't know this. So what Annabelle suggested is very possible. And since Public LB will remain the same for this competition, it doesn't matter to them that this method is useless for the private test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451859,
          "author_name": "Julius Hochmuth",
          "author_url": "",
          "post_date": "2019-01-07T19:29:37.350000",
          "content": "<p><a href=\"https://www.kaggle.com/smokerx/test-others\">https://www.kaggle.com/smokerx/test-others</a>\nHere is an example of what you could do. This particular person said that they only did it because the Kaggle kernels were taking too long to submit, but it does illustrate how you can use external data using copy/paste.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451870,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-07T19:58:13.677000",
          "content": "<p>Exactly.\nThat is why I am curious to know if private LB score will be based on 320k totally unseen rows of phase 2 test set or on all 376k rows of which 56k can be downloaded and predictions obtained by combining multiple historical predictions. </p>\n\n<p>I posted a question to <a href=\"/inversion\">@inversion</a> staff but got no reply yet</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451902,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-01-07T21:47:56.493000",
          "content": "<p>Hi @Annablle - With the holidays and some issues getting notifications, I missed your question.</p>\n\n<p>To be very clear about this:\n1. The Public Test is ~56k lines (phase 1)\n2. The phase 2 Test is ~376k lines\n3. The phase 2 Test includes Pubic <em>and</em> Private\n4. The Private rows will be used for the final Leaderboard\n5. The Public rows are used so that when we re-score, the Public Leaderboard isn't lost.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 451926,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-07T22:55:33.353000",
          "content": "<p>Thanks for the reply <a href=\"/inversion\">@inversion</a>.\nThat is good news, it means that it is not possible to hand-label the ~56k lines of phase 1 and copy-paste these labels into a notebook to boost score of final private LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 452470,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2019-01-08T19:09:26.757000",
          "content": "<p>Hi <a href=\"/inversion\">@inversion</a> </p>\n\n<p>When is the time that phase 2 data will be loaded? Went to the timeline page and couldn't see anything there. </p>\n\n<p>Will we be able to run our kernels on the phase 2 data and be able to select which kernels we want to submit for the final submission?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 453688,
          "author_name": "Andrzej Kuro",
          "author_url": "",
          "post_date": "2019-01-10T16:00:31.317000",
          "content": "<p>Hi Rahul,\nthe phase 2 data will not be publicly visible prior to competition end. The kernels will be run on new datasets with no user interaction. We just have to prepare good kernels :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 456617,
          "author_name": "Dilapsky Lee",
          "author_url": "",
          "post_date": "2019-01-16T07:44:22.530000",
          "content": "<p><a href=\"/inversion\">@inversion</a>\nHello! If I have a kernel, it runs 7200s in Public Test(~56k lines). However, it runs 7400s in phase 2 Test (~376k lines). Is that means the kernel will fail in the final standing?\nThanks!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 457438,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-01-17T12:30:53.240000",
          "content": "<p>Yes, read the guidelines posted.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "451632": "I'm sorry but this makes no sense to me.\n\n- You and the \"ladder climbers\" have equal access to public kernels, the only different, according to you, is that you put actually put effort into competing. If that's true then by the end there's no way they can finish above you, unless you refuse to use public knowledge for some reasons.\n- The \"ladder climbers\" discouraging genuine competitors: Why? They fork a kernel, rerun and get the same score, you instead try to improve upon that/ use that to improve your own kernel, why do you care if other people are wasting their time?\n- People in novice level being at the top: Ranks are irrelevant, and I don't know what it has to do with your previous points. \n\nThere is one thing that I dislike about kernels though, it is people publishing high scoring ones at the very late stage of a competition, but that's less about kernels and more about people's unethical behavior.",
    "451313": "Let's be straight. There are people here who just fork the latest and better (in terms of LB score) kernels and just climb up the ladder. This is just futile and pointless. Its a strategy thats never gonna get you to the top or teach you anything  on the long run. It also has a side-effect for other people: it discourages the rest of us who are trying hard to make a good submission after struggling with data. \n\nThis particular competition is illustrative of this point once again: it is very difficult to gain just a few positions and yet there are people in novice level who with 1-2 submissions are currently within medals (not that they' ll stay there but anywayz). Just saying...",
    "451723": "Kaggle officially warns about late kernel sharing : \n\nhttps://www.kaggle.com/spoiler-alert",
    "452529": "Personally, I see the kernels (especially with a good score) as a baseline which everyone has access to. So assuming everyone now has the knowledge from that kernel, I try to learn as many new things from different kernels and combine that knowledge to push my score beyond the kernel score. Its more about - if there is a kernel, fork it, learn the new and good ideas in it and use those along with your own to climb up further. The novices who just fork and submit will never be able to go up in the long run. ",
    "451319": "What's the problem? That is open research: \"Err and err and err, but less and less and less\". I learn more if I see some ideas from others, read up on them and try to think about them and incorporate them. You won't get into top ranks by just forking the best kernels.",
    "456199": "Maybe Kaggle can open an NLP-type competition to detect very similar kernels and issue a separate ranking for only unique kernels :D that'd weed out the forkers. ",
    "451687": "It would be a shame if people did not share their good kernels for others to learn. At least for me this is perhaps the best part of Kaggle, to learn from others.\n\nBut I see your issue too, the competition is for many people about gaming the system, and in that way one has to do the same unless really a top contender. I don't generally fork and submit, but then I generally don't make the medals :). \n\nHard to balance, I guess. But I like people sharing rather than discouraging that.",
    "456569": "While it is never a great feeling to see your score plunge due to people simply copying and pasting a Kernal, it is certainly better than having a data leak. Additionally, those who may have designed their own Kernal from scratch, but have a lower score should be able to quickly modify a higher performing Kernal to score higher than the copy pasters.  \n\nGenerally a company would want to have the best algorithm as well and they would prefer someone who can modify the best performing \"off the shelf\" algorithm than spend time creating one from scratch that does not perform as well.  So these competitions are not much different than life...  I am just glad to see people sharing so I can learn some techniques from others. \n\nHowever, I would like to see some kind of bonus point structure or otherwise for the submission of code that could run in production. ",
    "457398": "hi  if u see the new kernel with 0.699 score but without any sense you will be angry",
    "456368": "",
    "455598": "",
    "451326": "... and it gets even worse:\nI guess there are also people that download submission files from top public kernels, combine by average or voting, copy paste into a notebook and submit. Maybe some of the very high scores with very few submissions were obtained with this method.\n\nLet's be optimistic, winning this competition has a lot to do with getting a stable model (that will output similar probabilities when retrained in phase II) and a robust 'threshold finding algorithm'. \n\nThe test.csv is rather small anyway and thus LB scores have high variance. \nMy gut-feeling is that LB f1 variations smaller that 0.005 are irrelevant.\nSo I personalty rely on my own hold-out set.\n\nBy the way people novice in Kaggle are not necessarily novice to ML ;-)"
  }
}