{
  "id": 74020,
  "title": "When to stop releasing high scoring kernels",
  "url": "/competitions/PLAsTiCC-2018/discussion/74020",
  "author_name": "Subrahmanyam V",
  "post_date": "2018-12-07T16:20:49.954000",
  "votes": 12,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Dear Top Kagglers,\nI would like to see your comments on when it will be fair to stop releasing high scoring kernels.</p>\n\n<p>A new high scoring kernel released when the competetion is closing that can put anybody who runs it into top 100,  sounds like little unfair to all those who are working hard for the tail end of medals.</p>\n\n<p>Every kernel that shares new approach is so useful to the community. But in the sense of competetion, if we can have some kind of agreement on leaving the last week of competetion for just individual effort will be great.</p>\n\n<p>While I type this I feel a bit like a cry baby, and perhaps I should hone up my knowledge so much that this should not matter, but since I am here and it's kind of impacting me, I wanted to raise a general discussion on this.</p>\n\n<p>I also think kaggle admin team should not do any kind of policing, and this should just be some kind of general agreement to leave atleast one week of open competetion.</p>\n\n<p>PS: I built my work on top of 1.080 ideas kernel, and some other nueral net kernel, but those kernels are there atleast for two weeks now.</p>\n\n<p>Please provide your comments, perhaps I am the only one who is feeling like this.</p>",
  "messages": [
    {
      "id": 435178,
      "postDate": "2018-12-07T16:20:49.953Z",
      "content": "<p>Dear Top Kagglers,\nI would like to see your comments on when it will be fair to stop releasing high scoring kernels.</p>\n\n<p>A new high scoring kernel released when the competetion is closing that can put anybody who runs it into top 100,  sounds like little unfair to all those who are working hard for the tail end of medals.</p>\n\n<p>Every kernel that shares new approach is so useful to the community. But in the sense of competetion, if we can have some kind of agreement on leaving the last week of competetion for just individual effort will be great.</p>\n\n<p>While I type this I feel a bit like a cry baby, and perhaps I should hone up my knowledge so much that this should not matter, but since I am here and it's kind of impacting me, I wanted to raise a general discussion on this.</p>\n\n<p>I also think kaggle admin team should not do any kind of policing, and this should just be some kind of general agreement to leave atleast one week of open competetion.</p>\n\n<p>PS: I built my work on top of 1.080 ideas kernel, and some other nueral net kernel, but those kernels are there atleast for two weeks now.</p>\n\n<p>Please provide your comments, perhaps I am the only one who is feeling like this.</p>",
      "rawMarkdown": "Dear Top Kagglers,\nI would like to see your comments on when it will be fair to stop releasing high scoring kernels.\n\nA new high scoring kernel released when the competetion is closing that can put anybody who runs it into top 100,  sounds like little unfair to all those who are working hard for the tail end of medals.\n\nEvery kernel that shares new approach is so useful to the community. But in the sense of competetion, if we can have some kind of agreement on leaving the last week of competetion for just individual effort will be great.\n\nWhile I type this I feel a bit like a cry baby, and perhaps I should hone up my knowledge so much that this should not matter, but since I am here and it's kind of impacting me, I wanted to raise a general discussion on this.\n\nI also think kaggle admin team should not do any kind of policing, and this should just be some kind of general agreement to leave atleast one week of open competetion.\n\nPS: I built my work on top of 1.080 ideas kernel, and some other nueral net kernel, but those kernels are there atleast for two weeks now.\n\nPlease provide your comments, perhaps I am the only one who is feeling like this.",
      "votes": 12
    },
    {
      "id": 435185,
      "postDate": "2018-12-07T16:32:34.587Z",
      "content": "<p>Add them to your ensemble.</p>",
      "rawMarkdown": "Add them to your ensemble.",
      "votes": 7
    },
    {
      "id": 435226,
      "postDate": "2018-12-07T17:55:53.250Z",
      "content": "<p>My ultimate suggestion is: try to not optimize a public kernel, because while you make it better someone will release a better version of it. And you won't be able to get any benefit from ensembling because your model and the public kernel will be originated from the same model, so will have high correlation in the predictions. I try to avoid looking at public kernels to keep my diversity until the point I am stuck with my ideas.</p>",
      "rawMarkdown": "My ultimate suggestion is: try to not optimize a public kernel, because while you make it better someone will release a better version of it. And you won't be able to get any benefit from ensembling because your model and the public kernel will be originated from the same model, so will have high correlation in the predictions. I try to avoid looking at public kernels to keep my diversity until the point I am stuck with my ideas.",
      "votes": 8,
      "replies": [
        {
          "id": 435431,
          "postDate": "2018-12-08T03:02:17.743Z",
          "content": "<p>For beginners, it's still the public kernels. Otherwise I would have run a KNNClassifier on this problem which I actually tried recently. The results are horrible, confusion matrix is blue in only two vertical lines, 90% of the predictions are for class 90 and 42 :)  RandomForestClassifier beats KNN by close margin. </p>\n\n<p>I am learning a lot from Kaggle public kernels in a short time, which I could not do through books and MOOCs.</p>",
          "rawMarkdown": "For beginners, it's still the public kernels. Otherwise I would have run a KNNClassifier on this problem which I actually tried recently. The results are horrible, confusion matrix is blue in only two vertical lines, 90% of the predictions are for class 90 and 42 :)  RandomForestClassifier beats KNN by close margin. \n\nI am learning a lot from Kaggle public kernels in a short time, which I could not do through books and MOOCs.",
          "votes": 2
        }
      ]
    },
    {
      "id": 435290,
      "postDate": "2018-12-07T19:59:33.080Z",
      "content": "<p>This is a recurring topic, see for instance this discussion: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182</a></p>\n\n<p>A proposal that seem to get most votes is to forbid kernel sharing during last week of competition.</p>",
      "rawMarkdown": "This is a recurring topic, see for instance this discussion: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182\n\nA proposal that seem to get most votes is to forbid kernel sharing during last week of competition.",
      "votes": 3,
      "replies": [
        {
          "id": 435389,
          "postDate": "2018-12-08T00:21:11.747Z",
          "content": "<p>I would do two weeks, one week is quite a short time, plus it's a merge deadline. It's nice if people manage to do something on the base of a public kernel before merge deadline. I am a beginner at kaggle though...</p>",
          "rawMarkdown": "I would do two weeks, one week is quite a short time, plus it's a merge deadline. It's nice if people manage to do something on the base of a public kernel before merge deadline. I am a beginner at kaggle though...",
          "votes": 2
        },
        {
          "id": 435970,
          "postDate": "2018-12-09T06:52:12.203Z",
          "content": "<blockquote>\n  <p>I am a beginner at kaggle though…</p>\n</blockquote>\n\n<p>I begin to doubt it.  Beginner maybe, but good performer already. ;)</p>",
          "rawMarkdown": "&gt; I am a beginner at kaggle though…\n\nI begin to doubt it.  Beginner maybe, but good performer already. ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 435191,
      "postDate": "2018-12-07T16:41:26.517Z",
      "content": "<p>No, you are not the only one.</p>\n\n<p>Kernels have generally become a big problem, I think the last week is crucial in a competition and the kernels should not be released during that week, however we still have ten days. </p>\n\n<p>I think people should have more concern about releasing kernels, but today people only thinks in getting votes...</p>",
      "rawMarkdown": "No, you are not the only one.\n\nKernels have generally become a big problem, I think the last week is crucial in a competition and the kernels should not be released during that week, however we still have ten days. \n\nI think people should have more concern about releasing kernels, but today people only thinks in getting votes...\n",
      "votes": 3,
      "replies": [
        {
          "id": 435198,
          "postDate": "2018-12-07T16:45:44.450Z",
          "content": "<p>For sure there are cases of misuse, but I think kernels are the best thing kaggle has to offer. Everyone can participate, no matter their resources. And public kernels can teach a lot </p>",
          "rawMarkdown": "For sure there are cases of misuse, but I think kernels are the best thing kaggle has to offer. Everyone can participate, no matter their resources. And public kernels can teach a lot ",
          "votes": 1
        },
        {
          "id": 435200,
          "postDate": "2018-12-07T16:54:54Z",
          "content": "<p>Yes, I agree with you <a href=\"/iprapas\">@iprapas</a>, my best model is based in the amazing <a href=\"/olivier\">@olivier</a> kernel. I'm talking about of these late kernels with high scores where people are looking only to the votes that can get...</p>",
          "rawMarkdown": "Yes, I agree with you @iprapas, my best model is based in the amazing @olivier kernel. I'm talking about of these late kernels with high scores where people are looking only to the votes that can get...",
          "votes": 1
        },
        {
          "id": 435259,
          "postDate": "2018-12-07T19:05:41.603Z",
          "content": "<p>I learned a lot from downloading a couple of kernels. However, publishing full high score kernels in the last week or ten days does seem disruptive.</p>\n\n<p>Perhaps the same people could achieve the high kernel/discussion scores by just publishing/discussing their best idea? Then we could all learn from them and still have a chance to implement them in our own solutions. After the competition closes the incentive to continue work is very low, at least for me.   </p>",
          "rawMarkdown": "I learned a lot from downloading a couple of kernels. However, publishing full high score kernels in the last week or ten days does seem disruptive.\n\nPerhaps the same people could achieve the high kernel/discussion scores by just publishing/discussing their best idea? Then we could all learn from them and still have a chance to implement them in our own solutions. After the competition closes the incentive to continue work is very low, at least for me.   "
        }
      ]
    },
    {
      "id": 435308,
      "postDate": "2018-12-07T20:39:01.457Z",
      "content": "<p>Subrahmanyam V - unfortunately Kaggle are not going to do anything about this issue as I believe they like lots of high scoring kernels as it artificially inflates the mean scores for their customers (Hint they don't really care about competitors feelings because they could address this if they wanted to!) - luckily Jim hasn't done a true leaderboard killer - as you have used neural nets your model won't be too correlated with his so just do a geometric mean of your output with Jim's and you will zoom up to the leaderboard!  Just have the mindset that all competitions you enter will slowly but surely increase your knowledge.</p>",
      "rawMarkdown": "Subrahmanyam V - unfortunately Kaggle are not going to do anything about this issue as I believe they like lots of high scoring kernels as it artificially inflates the mean scores for their customers (Hint they don't really care about competitors feelings because they could address this if they wanted to!) - luckily Jim hasn't done a true leaderboard killer - as you have used neural nets your model won't be too correlated with his so just do a geometric mean of your output with Jim's and you will zoom up to the leaderboard!  Just have the mindset that all competitions you enter will slowly but surely increase your knowledge.",
      "votes": 4,
      "replies": [
        {
          "id": 435499,
          "postDate": "2018-12-08T05:56:34.300Z",
          "content": "<p>Scripus, I tried that, broke the 1.0 score barrier and I have taken the full benefit on LB :)</p>\n\n<p>I am learning a lot here and I cant thank enough the awesome community and the Kaggle team.</p>",
          "rawMarkdown": "Scripus, I tried that, broke the 1.0 score barrier and I have taken the full benefit on LB :)\n\nI am learning a lot here and I cant thank enough the awesome community and the Kaggle team.",
          "votes": 3
        }
      ]
    },
    {
      "id": 435266,
      "postDate": "2018-12-07T19:13:46.413Z",
      "content": "<p>Thanks for raising this concern.  I am new to the Kaggle community and am guilty of exactly what you describe.  My intention was to share a couple of cool things I did with the community.  The work of others was very helpful to me.  A lot of my learning during this competition was struggling to integrate my code with the code that others had written.  In the process of doing that the code that was difficult for me to understand in the beginning is now clear.</p>\n\n<p>I thought a bit about not posting my work but I didn't consider it any threat to the leaders, who I am nowhere near.  If I don't find more improvements I am unlikely to even finish in the top 100.  I guess time is running short but I am still working.  Do you think I should unshare my kernel?  I felt there was some value in sharing my Smote method and the seven features that I added.  I don't want to be unfair.  I didn't feel I was close enough to the leaders to be an issue.  </p>",
      "rawMarkdown": "Thanks for raising this concern.  I am new to the Kaggle community and am guilty of exactly what you describe.  My intention was to share a couple of cool things I did with the community.  The work of others was very helpful to me.  A lot of my learning during this competition was struggling to integrate my code with the code that others had written.  In the process of doing that the code that was difficult for me to understand in the beginning is now clear.\n\nI thought a bit about not posting my work but I didn't consider it any threat to the leaders, who I am nowhere near.  If I don't find more improvements I am unlikely to even finish in the top 100.  I guess time is running short but I am still working.  Do you think I should unshare my kernel?  I felt there was some value in sharing my Smote method and the seven features that I added.  I don't want to be unfair.  I didn't feel I was close enough to the leaders to be an issue.  ",
      "votes": 4,
      "replies": [
        {
          "id": 435277,
          "postDate": "2018-12-07T19:30:47.847Z",
          "content": "<p>I think there is a lot of value to your kernel and at this point it wouldn't make much sense to make it private. It has close to 400 views and 40 forks</p>",
          "rawMarkdown": "I think there is a lot of value to your kernel and at this point it wouldn't make much sense to make it private. It has close to 400 views and 40 forks"
        },
        {
          "id": 435312,
          "postDate": "2018-12-07T20:45:31.347Z",
          "content": "<ul>\n<li>I'll <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/63940\">echo Walter's advice</a>: if it's going to be disruptive to the current leaderboard, don't share it publicly during the last week. Bear in mind that a kernel that doesn't impact medals can still affect the vast majority of competitors.</li>\n<li>We love to see people sharing their learnings via kernels, but that can be done after the competition has closed.</li>\n</ul>\n\n<p>I'll be making a post reminding everyone of these guidelines once we're a week from close.</p>",
          "rawMarkdown": "- I'll [echo Walter's advice](https://www.kaggle.com/c/home-credit-default-risk/discussion/63940): if it's going to be disruptive to the current leaderboard, don't share it publicly during the last week. Bear in mind that a kernel that doesn't impact medals can still affect the vast majority of competitors.\n- We love to see people sharing their learnings via kernels, but that can be done after the competition has closed.\n\nI'll be making a post reminding everyone of these guidelines once we're a week from close.",
          "votes": 3
        }
      ]
    },
    {
      "id": 435957,
      "postDate": "2018-12-09T06:32:34.310Z",
      "content": "<p>Great question, this question has been asked many times, and the answer usually: \"this is one challenge in kaggle\".</p>",
      "rawMarkdown": "Great question, this question has been asked many times, and the answer usually: \"this is one challenge in kaggle\".",
      "votes": 1
    },
    {
      "id": 435391,
      "postDate": "2018-12-08T00:28:48.833Z",
      "content": "<p>So it seems from some that I should take this down and from others that once shared always shared.  I want to do the right thing.  For now I'm leaving it up but if there's a last week convention to follow I'm OK taking it down.</p>",
      "rawMarkdown": "So it seems from some that I should take this down and from others that once shared always shared.  I want to do the right thing.  For now I'm leaving it up but if there's a last week convention to follow I'm OK taking it down.",
      "votes": 1,
      "replies": [
        {
          "id": 435399,
          "postDate": "2018-12-08T00:58:40.947Z",
          "content": "<p>If you take it down at a time then you impair those who have not seen it before you take it down.</p>",
          "rawMarkdown": "If you take it down at a time then you impair those who have not seen it before you take it down.\n",
          "votes": 3
        },
        {
          "id": 435422,
          "postDate": "2018-12-08T02:41:51.107Z",
          "content": "<p>@Jim Sullivan, it's a good kernel and we will take learnings from it. Please leave it shared. We are not in the last week yet, we have ample time to go through what you have presented and take benefit from it if required.</p>\n\n<p>I only contemplated about what if something bigger (0.95 types of kernel) comes in the last week. That's little scary and hence initiated the discussion for awareness.</p>",
          "rawMarkdown": "@Jim Sullivan, it's a good kernel and we will take learnings from it. Please leave it shared. We are not in the last week yet, we have ample time to go through what you have presented and take benefit from it if required.\n\nI only contemplated about what if something bigger (0.95 types of kernel) comes in the last week. That's little scary and hence initiated the discussion for awareness."
        },
        {
          "id": 435454,
          "postDate": "2018-12-08T04:07:01.597Z",
          "content": "<p>Thanks all for the fruitful discussion.  I will leave the kernel shared.  </p>",
          "rawMarkdown": "Thanks all for the fruitful discussion.  I will leave the kernel shared.  ",
          "votes": 2
        },
        {
          "id": 435564,
          "postDate": "2018-12-08T08:57:54.047Z",
          "content": "<p><a href=\"/jimpsull\">@jimpsull</a>, just wondering, why you share a kernel with a <strong>private</strong> dataset in the last 2 weeks of the competition.</p>\n\n<p>I may have missed something but that means no one can fork and reproduce your results or make a better stack out of it.</p>\n\n<p>So this will only put more pressure on the LB with submissions at around 1.037.</p>\n\n<p>Please correct me if i'm wrong.</p>",
          "rawMarkdown": "@jimpsull, just wondering, why you share a kernel with a **private** dataset in the last 2 weeks of the competition.\n\nI may have missed something but that means no one can fork and reproduce your results or make a better stack out of it.\n\nSo this will only put more pressure on the LB with submissions at around 1.037.\n\nPlease correct me if i'm wrong.",
          "votes": 1
        },
        {
          "id": 435773,
          "postDate": "2018-12-08T18:41:05.073Z",
          "content": "<p>Which dataset is private?  I'm happy to make it public.  There are really only two sources of data which are both public.  I mention those in the kernel.  One is Chai-Ta Tsai's kernel and the other is my Something Different - Test Set Edition.  Those are both public.  </p>",
          "rawMarkdown": "Which dataset is private?  I'm happy to make it public.  There are really only two sources of data which are both public.  I mention those in the kernel.  One is Chai-Ta Tsai's kernel and the other is my Something Different - Test Set Edition.  Those are both public.  "
        },
        {
          "id": 435779,
          "postDate": "2018-12-08T18:52:48.303Z",
          "content": "<p>Thanks for your reply, when I fork the script I don't get one of the folders you use in the kernel.</p>\n\n<p>And BTW in the data tab of the kernel, it says there are 3 datasets and only 2 appear.</p>\n\n<p>The failing folder is : </p>\n\n<p><code>\nprint(os.listdir(\"../input/writefeaturetablefromsmotedartset\"))\n</code></p>",
          "rawMarkdown": "Thanks for your reply, when I fork the script I don't get one of the folders you use in the kernel.\n\nAnd BTW in the data tab of the kernel, it says there are 3 datasets and only 2 appear.\n\nThe failing folder is : \n\n`\nprint(os.listdir(\"../input/writefeaturetablefromsmotedartset\"))\n`"
        },
        {
          "id": 435816,
          "postDate": "2018-12-08T20:49:39.693Z",
          "content": "<p>I just made that kernel public.  Did it help?</p>",
          "rawMarkdown": "I just made that kernel public.  Did it help?"
        },
        {
          "id": 435817,
          "postDate": "2018-12-08T20:51:08.747Z",
          "content": "<p>Btw all it does is write the features instead of the predictions in Chai-Ta Tsais kernel.</p>",
          "rawMarkdown": "Btw all it does is write the features instead of the predictions in Chai-Ta Tsais kernel."
        }
      ]
    },
    {
      "id": 435193,
      "postDate": "2018-12-07T16:42:24.160Z",
      "content": "<p>I believe it is a constant issue in the hunt for medals. Someone who cannot win the competition, believes (s)he can easily get a gold kernel medal by exposing a good solution. Someone who has given thought to the competition and has time near the end, can probably only profit from more ideas coming from a high scoring kernel. Highly improbable that you are going to have exactly the same ideas. My advice would be to try to think more about the journey than the destination, but maybe that's too naive.</p>",
      "rawMarkdown": "I believe it is a constant issue in the hunt for medals. Someone who cannot win the competition, believes (s)he can easily get a gold kernel medal by exposing a good solution. Someone who has given thought to the competition and has time near the end, can probably only profit from more ideas coming from a high scoring kernel. Highly improbable that you are going to have exactly the same ideas. My advice would be to try to think more about the journey than the destination, but maybe that's too naive.",
      "votes": 1
    },
    {
      "id": 436587,
      "postDate": "2018-12-10T15:29:38.360Z",
      "content": "<p>I think it's partially an issue of philosophy; a balance needs to be struck, and that balance between supposed knowledge \"philanthropy\" and potential personal gain is difficult to make sometimes. </p>",
      "rawMarkdown": "I think it's partially an issue of philosophy; a balance needs to be struck, and that balance between supposed knowledge \"philanthropy\" and potential personal gain is difficult to make sometimes. "
    },
    {
      "id": 435297,
      "postDate": "2018-12-07T20:10:38.603Z",
      "content": "<p>It is my understanding that we are not yet in the last week.  True?  I think the 17th is the end.  Would it be considered in good taste for me to make my kernel private on the 10th?</p>",
      "rawMarkdown": "It is my understanding that we are not yet in the last week.  True?  I think the 17th is the end.  Would it be considered in good taste for me to make my kernel private on the 10th?",
      "replies": [
        {
          "id": 435334,
          "postDate": "2018-12-07T21:49:34.007Z",
          "content": "<p>Something that was shared should stay shared IMHO.  </p>",
          "rawMarkdown": "Something that was shared should stay shared IMHO.  ",
          "votes": 6
        }
      ]
    },
    {
      "id": 435209,
      "postDate": "2018-12-07T17:33:09.440Z",
      "content": "<h2>The noblest pleasure is the joy of understanding.</h2>\n\n<p>Leonardo da Vinci</p>",
      "rawMarkdown": "## The noblest pleasure is the joy of understanding. \nLeonardo da Vinci"
    },
    {
      "id": 435862,
      "postDate": "2018-12-08T23:36:08.193Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 435185,
      "author_name": "Sergey Lebedev",
      "author_url": "",
      "post_date": "2018-12-07T16:32:34.587000",
      "content": "<p>Add them to your ensemble.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 435226,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2018-12-07T17:55:53.250000",
      "content": "<p>My ultimate suggestion is: try to not optimize a public kernel, because while you make it better someone will release a better version of it. And you won't be able to get any benefit from ensembling because your model and the public kernel will be originated from the same model, so will have high correlation in the predictions. I try to avoid looking at public kernels to keep my diversity until the point I am stuck with my ideas.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 435431,
          "author_name": "Subrahmanyam V",
          "author_url": "",
          "post_date": "2018-12-08T03:02:17.743000",
          "content": "<p>For beginners, it's still the public kernels. Otherwise I would have run a KNNClassifier on this problem which I actually tried recently. The results are horrible, confusion matrix is blue in only two vertical lines, 90% of the predictions are for class 90 and 42 :)  RandomForestClassifier beats KNN by close margin. </p>\n\n<p>I am learning a lot from Kaggle public kernels in a short time, which I could not do through books and MOOCs.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 435290,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-07T19:59:33.080000",
      "content": "<p>This is a recurring topic, see for instance this discussion: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182</a></p>\n\n<p>A proposal that seem to get most votes is to forbid kernel sharing during last week of competition.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 435389,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-08T00:21:11.747000",
          "content": "<p>I would do two weeks, one week is quite a short time, plus it's a merge deadline. It's nice if people manage to do something on the base of a public kernel before merge deadline. I am a beginner at kaggle though...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 435970,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-09T06:52:12.203000",
          "content": "<blockquote>\n  <p>I am a beginner at kaggle though…</p>\n</blockquote>\n\n<p>I begin to doubt it.  Beginner maybe, but good performer already. ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 435191,
      "author_name": "João Pedro Peinado",
      "author_url": "",
      "post_date": "2018-12-07T16:41:26.517000",
      "content": "<p>No, you are not the only one.</p>\n\n<p>Kernels have generally become a big problem, I think the last week is crucial in a competition and the kernels should not be released during that week, however we still have ten days. </p>\n\n<p>I think people should have more concern about releasing kernels, but today people only thinks in getting votes...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 435198,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "2018-12-07T16:45:44.450000",
          "content": "<p>For sure there are cases of misuse, but I think kernels are the best thing kaggle has to offer. Everyone can participate, no matter their resources. And public kernels can teach a lot </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 435200,
          "author_name": "João Pedro Peinado",
          "author_url": "",
          "post_date": "2018-12-07T16:54:54",
          "content": "<p>Yes, I agree with you <a href=\"/iprapas\">@iprapas</a>, my best model is based in the amazing <a href=\"/olivier\">@olivier</a> kernel. I'm talking about of these late kernels with high scores where people are looking only to the votes that can get...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 435259,
          "author_name": "PeterSorensen",
          "author_url": "",
          "post_date": "2018-12-07T19:05:41.603000",
          "content": "<p>I learned a lot from downloading a couple of kernels. However, publishing full high score kernels in the last week or ten days does seem disruptive.</p>\n\n<p>Perhaps the same people could achieve the high kernel/discussion scores by just publishing/discussing their best idea? Then we could all learn from them and still have a chance to implement them in our own solutions. After the competition closes the incentive to continue work is very low, at least for me.   </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435308,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-12-07T20:39:01.457000",
      "content": "<p>Subrahmanyam V - unfortunately Kaggle are not going to do anything about this issue as I believe they like lots of high scoring kernels as it artificially inflates the mean scores for their customers (Hint they don't really care about competitors feelings because they could address this if they wanted to!) - luckily Jim hasn't done a true leaderboard killer - as you have used neural nets your model won't be too correlated with his so just do a geometric mean of your output with Jim's and you will zoom up to the leaderboard!  Just have the mindset that all competitions you enter will slowly but surely increase your knowledge.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 435499,
          "author_name": "Subrahmanyam V",
          "author_url": "",
          "post_date": "2018-12-08T05:56:34.300000",
          "content": "<p>Scripus, I tried that, broke the 1.0 score barrier and I have taken the full benefit on LB :)</p>\n\n<p>I am learning a lot here and I cant thank enough the awesome community and the Kaggle team.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 435266,
      "author_name": "Jim Sullivan",
      "author_url": "",
      "post_date": "2018-12-07T19:13:46.413000",
      "content": "<p>Thanks for raising this concern.  I am new to the Kaggle community and am guilty of exactly what you describe.  My intention was to share a couple of cool things I did with the community.  The work of others was very helpful to me.  A lot of my learning during this competition was struggling to integrate my code with the code that others had written.  In the process of doing that the code that was difficult for me to understand in the beginning is now clear.</p>\n\n<p>I thought a bit about not posting my work but I didn't consider it any threat to the leaders, who I am nowhere near.  If I don't find more improvements I am unlikely to even finish in the top 100.  I guess time is running short but I am still working.  Do you think I should unshare my kernel?  I felt there was some value in sharing my Smote method and the seven features that I added.  I don't want to be unfair.  I didn't feel I was close enough to the leaders to be an issue.  </p>",
      "votes": 4,
      "replies": [
        {
          "id": 435277,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "2018-12-07T19:30:47.847000",
          "content": "<p>I think there is a lot of value to your kernel and at this point it wouldn't make much sense to make it private. It has close to 400 views and 40 forks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435312,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-12-07T20:45:31.347000",
          "content": "<ul>\n<li>I'll <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/63940\">echo Walter's advice</a>: if it's going to be disruptive to the current leaderboard, don't share it publicly during the last week. Bear in mind that a kernel that doesn't impact medals can still affect the vast majority of competitors.</li>\n<li>We love to see people sharing their learnings via kernels, but that can be done after the competition has closed.</li>\n</ul>\n\n<p>I'll be making a post reminding everyone of these guidelines once we're a week from close.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 435957,
      "author_name": "LongYin/杰少",
      "author_url": "",
      "post_date": "2018-12-09T06:32:34.310000",
      "content": "<p>Great question, this question has been asked many times, and the answer usually: \"this is one challenge in kaggle\".</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 435391,
      "author_name": "Jim Sullivan",
      "author_url": "",
      "post_date": "2018-12-08T00:28:48.833000",
      "content": "<p>So it seems from some that I should take this down and from others that once shared always shared.  I want to do the right thing.  For now I'm leaving it up but if there's a last week convention to follow I'm OK taking it down.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 435399,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-08T00:58:40.947000",
          "content": "<p>If you take it down at a time then you impair those who have not seen it before you take it down.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 435422,
          "author_name": "Subrahmanyam V",
          "author_url": "",
          "post_date": "2018-12-08T02:41:51.107000",
          "content": "<p>@Jim Sullivan, it's a good kernel and we will take learnings from it. Please leave it shared. We are not in the last week yet, we have ample time to go through what you have presented and take benefit from it if required.</p>\n\n<p>I only contemplated about what if something bigger (0.95 types of kernel) comes in the last week. That's little scary and hence initiated the discussion for awareness.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435454,
          "author_name": "Jim Sullivan",
          "author_url": "",
          "post_date": "2018-12-08T04:07:01.597000",
          "content": "<p>Thanks all for the fruitful discussion.  I will leave the kernel shared.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 435564,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-12-08T08:57:54.047000",
          "content": "<p><a href=\"/jimpsull\">@jimpsull</a>, just wondering, why you share a kernel with a <strong>private</strong> dataset in the last 2 weeks of the competition.</p>\n\n<p>I may have missed something but that means no one can fork and reproduce your results or make a better stack out of it.</p>\n\n<p>So this will only put more pressure on the LB with submissions at around 1.037.</p>\n\n<p>Please correct me if i'm wrong.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 435773,
          "author_name": "Jim Sullivan",
          "author_url": "",
          "post_date": "2018-12-08T18:41:05.073000",
          "content": "<p>Which dataset is private?  I'm happy to make it public.  There are really only two sources of data which are both public.  I mention those in the kernel.  One is Chai-Ta Tsai's kernel and the other is my Something Different - Test Set Edition.  Those are both public.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435779,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-12-08T18:52:48.303000",
          "content": "<p>Thanks for your reply, when I fork the script I don't get one of the folders you use in the kernel.</p>\n\n<p>And BTW in the data tab of the kernel, it says there are 3 datasets and only 2 appear.</p>\n\n<p>The failing folder is : </p>\n\n<p><code>\nprint(os.listdir(\"../input/writefeaturetablefromsmotedartset\"))\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435816,
          "author_name": "Jim Sullivan",
          "author_url": "",
          "post_date": "2018-12-08T20:49:39.693000",
          "content": "<p>I just made that kernel public.  Did it help?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435817,
          "author_name": "Jim Sullivan",
          "author_url": "",
          "post_date": "2018-12-08T20:51:08.747000",
          "content": "<p>Btw all it does is write the features instead of the predictions in Chai-Ta Tsais kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435193,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "2018-12-07T16:42:24.160000",
      "content": "<p>I believe it is a constant issue in the hunt for medals. Someone who cannot win the competition, believes (s)he can easily get a gold kernel medal by exposing a good solution. Someone who has given thought to the competition and has time near the end, can probably only profit from more ideas coming from a high scoring kernel. Highly improbable that you are going to have exactly the same ideas. My advice would be to try to think more about the journey than the destination, but maybe that's too naive.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 436587,
      "author_name": "Rustom Ichhaporia",
      "author_url": "",
      "post_date": "2018-12-10T15:29:38.360000",
      "content": "<p>I think it's partially an issue of philosophy; a balance needs to be struck, and that balance between supposed knowledge \"philanthropy\" and potential personal gain is difficult to make sometimes. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435297,
      "author_name": "Jim Sullivan",
      "author_url": "",
      "post_date": "2018-12-07T20:10:38.603000",
      "content": "<p>It is my understanding that we are not yet in the last week.  True?  I think the 17th is the end.  Would it be considered in good taste for me to make my kernel private on the 10th?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435334,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-07T21:49:34.007000",
          "content": "<p>Something that was shared should stay shared IMHO.  </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 435209,
      "author_name": "Eric Vos",
      "author_url": "",
      "post_date": "2018-12-07T17:33:09.440000",
      "content": "<h2>The noblest pleasure is the joy of understanding.</h2>\n\n<p>Leonardo da Vinci</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 435862,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-08T23:36:08.193000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "435178": "Dear Top Kagglers,\nI would like to see your comments on when it will be fair to stop releasing high scoring kernels.\n\nA new high scoring kernel released when the competetion is closing that can put anybody who runs it into top 100,  sounds like little unfair to all those who are working hard for the tail end of medals.\n\nEvery kernel that shares new approach is so useful to the community. But in the sense of competetion, if we can have some kind of agreement on leaving the last week of competetion for just individual effort will be great.\n\nWhile I type this I feel a bit like a cry baby, and perhaps I should hone up my knowledge so much that this should not matter, but since I am here and it's kind of impacting me, I wanted to raise a general discussion on this.\n\nI also think kaggle admin team should not do any kind of policing, and this should just be some kind of general agreement to leave atleast one week of open competetion.\n\nPS: I built my work on top of 1.080 ideas kernel, and some other nueral net kernel, but those kernels are there atleast for two weeks now.\n\nPlease provide your comments, perhaps I am the only one who is feeling like this.",
    "435185": "Add them to your ensemble.",
    "435226": "My ultimate suggestion is: try to not optimize a public kernel, because while you make it better someone will release a better version of it. And you won't be able to get any benefit from ensembling because your model and the public kernel will be originated from the same model, so will have high correlation in the predictions. I try to avoid looking at public kernels to keep my diversity until the point I am stuck with my ideas.",
    "435290": "This is a recurring topic, see for instance this discussion: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182\n\nA proposal that seem to get most votes is to forbid kernel sharing during last week of competition.",
    "435191": "No, you are not the only one.\n\nKernels have generally become a big problem, I think the last week is crucial in a competition and the kernels should not be released during that week, however we still have ten days. \n\nI think people should have more concern about releasing kernels, but today people only thinks in getting votes...\n",
    "435308": "Subrahmanyam V - unfortunately Kaggle are not going to do anything about this issue as I believe they like lots of high scoring kernels as it artificially inflates the mean scores for their customers (Hint they don't really care about competitors feelings because they could address this if they wanted to!) - luckily Jim hasn't done a true leaderboard killer - as you have used neural nets your model won't be too correlated with his so just do a geometric mean of your output with Jim's and you will zoom up to the leaderboard!  Just have the mindset that all competitions you enter will slowly but surely increase your knowledge.",
    "435266": "Thanks for raising this concern.  I am new to the Kaggle community and am guilty of exactly what you describe.  My intention was to share a couple of cool things I did with the community.  The work of others was very helpful to me.  A lot of my learning during this competition was struggling to integrate my code with the code that others had written.  In the process of doing that the code that was difficult for me to understand in the beginning is now clear.\n\nI thought a bit about not posting my work but I didn't consider it any threat to the leaders, who I am nowhere near.  If I don't find more improvements I am unlikely to even finish in the top 100.  I guess time is running short but I am still working.  Do you think I should unshare my kernel?  I felt there was some value in sharing my Smote method and the seven features that I added.  I don't want to be unfair.  I didn't feel I was close enough to the leaders to be an issue.  ",
    "435957": "Great question, this question has been asked many times, and the answer usually: \"this is one challenge in kaggle\".",
    "435391": "So it seems from some that I should take this down and from others that once shared always shared.  I want to do the right thing.  For now I'm leaving it up but if there's a last week convention to follow I'm OK taking it down.",
    "435193": "I believe it is a constant issue in the hunt for medals. Someone who cannot win the competition, believes (s)he can easily get a gold kernel medal by exposing a good solution. Someone who has given thought to the competition and has time near the end, can probably only profit from more ideas coming from a high scoring kernel. Highly improbable that you are going to have exactly the same ideas. My advice would be to try to think more about the journey than the destination, but maybe that's too naive.",
    "436587": "I think it's partially an issue of philosophy; a balance needs to be struck, and that balance between supposed knowledge \"philanthropy\" and potential personal gain is difficult to make sometimes. ",
    "435297": "It is my understanding that we are not yet in the last week.  True?  I think the 17th is the end.  Would it be considered in good taste for me to make my kernel private on the 10th?",
    "435209": "## The noblest pleasure is the joy of understanding. \nLeonardo da Vinci",
    "435862": ""
  }
}