{
  "id": 52779,
  "title": "NO MORE BLENDING IN KERNALS PLEASE",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52779",
  "author_name": "",
  "post_date": "2018-03-23T01:52:44.028800700Z",
  "votes": 80,
  "comment_count": 57,
  "views": 0,
  "content": "<p>These blending kernels is useless to us. Even some insights about how to blend deserve 99 votes. Guys use the kernals' output to generate a blending result just want to Fraud Vote.(骗赞) This damages the ecology of kaggle. </p>\n\n<p>I struggle with this in toxic comment classification. \nDo something interesting and amazing in such a short life.</p>",
  "messages": [
    {
      "id": "301611",
      "postDate": "03/23/2018 01:52:44",
      "content": "<p>These blending kernels is useless to us. Even some insights about how to blend deserve 99 votes. Guys use the kernals' output to generate a blending result just want to Fraud Vote.(骗赞) This damages the ecology of kaggle. </p>\n\n<p>I struggle with this in toxic comment classification. \nDo something interesting and amazing in such a short life.</p>",
      "rawMarkdown": "These blending kernels is useless to us. Even some insights about how to blend deserve 99 votes. Guys use the kernals' output to generate a blending result just want to Fraud Vote.(骗赞) This damages the ecology of kaggle. \n\n\nI struggle with this in toxic comment classification. \nDo something interesting and amazing in such a short life.",
      "votes": null
    },
    {
      "id": "301756",
      "postDate": "03/23/2018 07:44:53",
      "content": "<p>I agree.</p>\n\n<p>These blenders have caused more harm than good in my opinion:</p>\n\n<ul>\n<li><p>people don't want to share their kernels with CREATIVE AND NEW ideas so that they are not put in the blender</p></li>\n<li><p>which leads to no learning from each other and no building upon each other's ideas and knowledge.</p></li>\n</ul>\n\n<p>I personally, would never share my kernel before the competition ends purely because of this. There should be a solution soon. </p>",
      "rawMarkdown": "I agree.\n\nThese blenders have caused more harm than good in my opinion:\n\n- people don't want to share their kernels with CREATIVE AND NEW ideas so that they are not put in the blender\n\n- which leads to no learning from each other and no building upon each other's ideas and knowledge.\n\nI personally, would never share my kernel before the competition ends purely because of this. There should be a solution soon.",
      "votes": null
    },
    {
      "id": "301758",
      "postDate": "03/23/2018 07:47:58",
      "content": "<p>One thing that comes to my mind, which for sure has certain disadvantages as I haven't thought it over is to be able to share publicly the kernel BUT not the output file. In such a way those blender boys will have to at least spend some memory power ( I am talking about computer memory, as we all know that the brain power is limited in such cases)</p>",
      "rawMarkdown": "One thing that comes to my mind, which for sure has certain disadvantages as I haven't thought it over is to be able to share publicly the kernel BUT not the output file. In such a way those blender boys will have to at least spend some memory power ( I am talking about computer memory, as we all know that the brain power is limited in such cases)",
      "votes": null
    },
    {
      "id": "301759",
      "postDate": "03/23/2018 07:48:54",
      "content": "<p>In principle i agree, but - barring manual inspection of kernels by Kaggle - i dont see an automated way to do it.</p>",
      "rawMarkdown": "In principle i agree, but - barring manual inspection of kernels by Kaggle - i dont see an automated way to do it.",
      "votes": null
    },
    {
      "id": "301796",
      "postDate": "03/23/2018 09:05:26",
      "content": "<p>Not sure if possible, but you could post your kernel and just add a line to generate a false output file ? That way you can share your ideas but not your results.</p>",
      "rawMarkdown": "Not sure if possible, but you could post your kernel and just add a line to generate a false output file ? That way you can share your ideas but not your results.",
      "votes": null
    },
    {
      "id": "301799",
      "postDate": "03/23/2018 09:12:17",
      "content": "<p>Absolutely agree.</p>\n\n<p>Only third week of this competition, I think it will be interesting to find some interesting patterns in the data.</p>",
      "rawMarkdown": "Absolutely agree.\n\nOnly third week of this competition, I think it will be interesting to find some interesting patterns in the data.",
      "votes": null
    },
    {
      "id": "301806",
      "postDate": "03/23/2018 09:23:26",
      "content": "<p>Agreed.\nPlease read this discussion and share your opinion.\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802</a></p>",
      "rawMarkdown": "Agreed.\nPlease read this discussion and share your opinion.\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802",
      "votes": null
    },
    {
      "id": "301808",
      "postDate": "03/23/2018 09:24:41",
      "content": "<p>How about making it private if it surpasses 10% of bronze positions. Linking another discussion here -<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802</a> \nPlease do share your opinion Konrad. Thanks!</p>",
      "rawMarkdown": "How about making it private if it surpasses 10% of bronze positions. Linking another discussion here -https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802 \nPlease do share your opinion Konrad. Thanks!",
      "votes": null
    },
    {
      "id": "301815",
      "postDate": "03/23/2018 09:40:13",
      "content": "<p>I understand your line of reasoning here - I think - but that particular solution would give an advantage to the people who grabbed it first. Personally, I don't have that much problem with the blends: they serve as a sort of filter on public kernel ideas, showing me which ones might be useful for implementing myself (i.e. using them in a proper validation setup etc). </p>\n\n<p>There is also a cynical aspect to my take on the problem: there will always be people just gaming the system. I have been riled for quite a while by the fact that in the global ranking there was always this one guy ahead of me, who had a very simple modus operandi: he registered for every competition there was, submitted the benchmark and never looked back. Pure scale effect was enough to put him ahead of me, but I realized that hunting down people like that carried a strong risk of doing more harm than good.</p>",
      "rawMarkdown": "I understand your line of reasoning here - I think - but that particular solution would give an advantage to the people who grabbed it first. Personally, I don't have that much problem with the blends: they serve as a sort of filter on public kernel ideas, showing me which ones might be useful for implementing myself (i.e. using them in a proper validation setup etc). \n\nThere is also a cynical aspect to my take on the problem: there will always be people just gaming the system. I have been riled for quite a while by the fact that in the global ranking there was always this one guy ahead of me, who had a very simple modus operandi: he registered for every competition there was, submitted the benchmark and never looked back. Pure scale effect was enough to put him ahead of me, but I realized that hunting down people like that carried a strong risk of doing more harm than good.",
      "votes": null
    },
    {
      "id": "301817",
      "postDate": "03/23/2018 09:50:24",
      "content": "<p>Very good point!</p>",
      "rawMarkdown": "Very good point!",
      "votes": null
    },
    {
      "id": "301848",
      "postDate": "03/23/2018 11:16:36",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/301848/8854/blend_moar_kernalz.png\" alt=\"enter image description here\"></p>\n\n<h2><a href=\"https://livingthing.danmackinlay.name/deep_learning.html\">adapted from this</a></h2>",
      "rawMarkdown": "![enter image description here][1]\n## [adapted from this][2]\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/301848/8854/blend_moar_kernalz.png\n  [2]: https://livingthing.danmackinlay.name/deep_learning.html",
      "votes": null
    },
    {
      "id": "301859",
      "postDate": "03/23/2018 11:41:36",
      "content": "<p>one solution could be for kernels not to accept user uploads or cross-references to other kernels outputs. each kernel will have to build the output and any blending it might perform starting from the competition datasource.</p>",
      "rawMarkdown": "one solution could be for kernels not to accept user uploads or cross-references to other kernels outputs. each kernel will have to build the output and any blending it might perform starting from the competition datasource.",
      "votes": null
    },
    {
      "id": "301865",
      "postDate": "03/23/2018 11:54:39",
      "content": "<p>Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.</p>",
      "rawMarkdown": "Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.",
      "votes": null
    },
    {
      "id": "301918",
      "postDate": "03/23/2018 13:09:56",
      "content": "<p>There are definitely <a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\">interesting patterns</a> in there :)</p>\n\n<p>(Edit) Preview:\n<a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\"><img src=\"https://s31.postimg.cc/c8e6qsomz/channel_105a.png\" alt=\"Channel 105 clicks\"></a></p>",
      "rawMarkdown": "There are definitely [interesting patterns][1] in there :)\n\n(Edit) Preview:\n[![Channel 105 clicks][2]][1]\n\n  [1]: https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\n  [2]: https://s31.postimg.cc/c8e6qsomz/channel_105a.png",
      "votes": null
    },
    {
      "id": "301929",
      "postDate": "03/23/2018 13:30:07",
      "content": "<p>And <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates/notebook\">here</a> as well ;)</p>",
      "rawMarkdown": "And [here][1] as well ;)\n\n\n  [1]: https://www.kaggle.com/cpmpml/ip-download-rates/notebook",
      "votes": null
    },
    {
      "id": "301990",
      "postDate": "03/23/2018 14:49:06",
      "content": "<p>You are a Soul Painter(灵魂画师)</p>",
      "rawMarkdown": "You are a Soul Painter(灵魂画师)",
      "votes": null
    },
    {
      "id": "302011",
      "postDate": "03/23/2018 15:16:55",
      "content": "<p>Haha, thanks! It's not really my work, the (brilliant) original is <a href=\"https://livingthing.danmackinlay.name/deep_learning.html\">here</a>, I just changed the text in the bottom half, I have made the link in the post a bit clearer...</p>",
      "rawMarkdown": "Haha, thanks! It's not really my work, the (brilliant) original is [here][1], I just changed the text in the bottom half, I have made the link in the post a bit clearer...\n\n\n  [1]: https://livingthing.danmackinlay.name/deep_learning.html",
      "votes": null
    },
    {
      "id": "302028",
      "postDate": "03/23/2018 15:43:39",
      "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <blockquote>\n    <p>Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.</p>\n  </blockquote>\n</blockquote>\n\n<p>Yes, public kernels are usually poison.  Experienced Kagglers know that and benefit from less experienced Kagglers not knowing that.  They are an effective education tool, but more often than not they are effective at teaching precisely the wrong and dangerous lessons.</p>",
      "rawMarkdown": "&gt; **CPMP wrote**\n&gt; \n&gt; &gt; Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.\n\nYes, public kernels are usually poison.  Experienced Kagglers know that and benefit from less experienced Kagglers not knowing that.  They are an effective education tool, but more often than not they are effective at teaching precisely the wrong and dangerous lessons.",
      "votes": null
    },
    {
      "id": "302191",
      "postDate": "03/23/2018 19:27:30",
      "content": "<p>I second @spongebob's concern and would like to share my thoughts. I of course, respect the choice and interest of everyone in this diverse community and consider it okay and tolerable until it's a normal blend (at least with some thought process on blending). </p>\n\n<p>The problem starts when competition becomes the game of random lucky numbers at the very early stage and the idea behind improving base models is forgotten. In my personal opinion, that would be the last thing to try when no time is left to improve any models. </p>\n\n<p>My suggestion would be to <strong>separate/ filter</strong>  blending kernels from real kernels in kernel's tab so that people can choose what they want to explore/learn and when. Alternate option could be to post blending script in discussion section instead of kernels so that true learning material can be accessible easily for majority of Kagglers who are interested to gain and contribute something positive. </p>",
      "rawMarkdown": "I second @spongebob's concern and would like to share my thoughts. I of course, respect the choice and interest of everyone in this diverse community and consider it okay and tolerable until it's a normal blend (at least with some thought process on blending). \n\nThe problem starts when competition becomes the game of random lucky numbers at the very early stage and the idea behind improving base models is forgotten. In my personal opinion, that would be the last thing to try when no time is left to improve any models. \n\nMy suggestion would be to **separate/ filter**  blending kernels from real kernels in kernel's tab so that people can choose what they want to explore/learn and when. Alternate option could be to post blending script in discussion section instead of kernels so that true learning material can be accessible easily for majority of Kagglers who are interested to gain and contribute something positive.",
      "votes": null
    },
    {
      "id": "302247",
      "postDate": "03/23/2018 20:54:12",
      "content": "<p>As a community, we just have to reward actual creative contribution rather than blends. And hey, if the blend is thoughtful and backed by theory, take my upvote!</p>",
      "rawMarkdown": "As a community, we just have to reward actual creative contribution rather than blends. And hey, if the blend is thoughtful and backed by theory, take my upvote!",
      "votes": null
    },
    {
      "id": "302291",
      "postDate": "03/23/2018 21:36:47",
      "content": "<p>I like the idea of having a filter for blends in the Kernels tab. I think it would probably work if they simply require us to declare whether or not it's a blend (maybe when a kernel goes public), because I don't think many people would deliberately lie.  But if necessary they could allow us to filter out kernels with multiple data sources, which is a noisy proxy for blends.</p>",
      "rawMarkdown": "I like the idea of having a filter for blends in the Kernels tab. I think it would probably work if they simply require us to declare whether or not it's a blend (maybe when a kernel goes public), because I don't think many people would deliberately lie.  But if necessary they could allow us to filter out kernels with multiple data sources, which is a noisy proxy for blends.",
      "votes": null
    },
    {
      "id": "302296",
      "postDate": "03/23/2018 21:53:08",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks",
      "votes": null
    },
    {
      "id": "302307",
      "postDate": "03/23/2018 22:19:19",
      "content": "<p>I am a newbie to Kaggle I entirely went overboard with blending in  <a href=\"https://www.kaggle.com/gopisaran/ensemble-of-2-ensembles-lb-0-9696\">this Kernel</a></p>\n\n<p>It was a mere experiment, and I had no intentions to damage the ecology of Kaggle. I am sorry !!\nI now understand that the Kaggle community thrives on improving base models, exploring and learning which is going to be my focus. This discussion was a great learning for a newbie like me. Thanks guys for bringing this point up. Appreciate it!​</p>",
      "rawMarkdown": "I am a newbie to Kaggle I entirely went overboard with blending in  [this Kernel][1]\n\n\n  [1]: https://www.kaggle.com/gopisaran/ensemble-of-2-ensembles-lb-0-9696\n\nIt was a mere experiment, and I had no intentions to damage the ecology of Kaggle. I am sorry !!\nI now understand that the Kaggle community thrives on improving base models, exploring and learning which is going to be my focus. This discussion was a great learning for a newbie like me. Thanks guys for bringing this point up. Appreciate it!​",
      "votes": null
    },
    {
      "id": "302319",
      "postDate": "03/23/2018 22:38:57",
      "content": "<p>These blends are somehow like <a href=\"https://en.wikipedia.org/wiki/Prohibition_in_the_United_States\">these other blends</a>. Prohibition was not the best approach. I'm sure there are better ways to deal with the problem.</p>",
      "rawMarkdown": "These blends are somehow like [these other blends](https://en.wikipedia.org/wiki/Prohibition_in_the_United_States). Prohibition was not the best approach. I'm sure there are better ways to deal with the problem.",
      "votes": null
    },
    {
      "id": "302419",
      "postDate": "03/24/2018 02:53:36",
      "content": "<p>Just release two kernels -- one with your idea and one with your idea included in the Blend of the Day. Presto, now you have two silver medals!</p>",
      "rawMarkdown": "Just release two kernels -- one with your idea and one with your idea included in the Blend of the Day. Presto, now you have two silver medals!",
      "votes": null
    },
    {
      "id": "302495",
      "postDate": "03/24/2018 07:06:52",
      "content": "<p>I Partially agree with @spongebob and I respect every one's views. \n<strong>I agree Blending is also an art and a piece of data science solutions framework. <br>\nBut we should also encourage the Diverse solutions and the effort put in them.</strong> </p>\n\n<p><strong>My suggestion to @Kaggle is</strong> <br>\nto make an <strong>intermediate dead line like ( one/two month after launch of competition)</strong> , <br></p>\n\n<ul>\n<li>Till this date kernels can be made public.</li>\n<li>later this date all new kernels will become private kernels. </li>\n<li>These can be made public post competition. <br>\nBy doing this Blending can still be be done using past public kernels also it encourages Diverse solutions nearing to competition closure. \nThank you</li>\n</ul>",
      "rawMarkdown": "I Partially agree with @spongebob and I respect every one's views. \n**I agree Blending is also an art and a piece of data science solutions framework. <br>\nBut we should also encourage the Diverse solutions and the effort put in them.** \n\n**My suggestion to @Kaggle is** <br>\nto make an **intermediate dead line like ( one/two month after launch of competition)** , <br>\n\n - Till this date kernels can be made public.\n -  later this date all new kernels will become private kernels. \n -  These can be made public post competition. <br>\nBy doing this Blending can still be be done using past public kernels also it encourages Diverse solutions nearing to competition closure. \nThank you",
      "votes": null
    },
    {
      "id": "302535",
      "postDate": "03/24/2018 08:58:56",
      "content": "<p>You don't have to be sorry.  As long as you don't break any rule (or any law) then you're fine.  The comments here are opinions of some kagglers, and they are not defining what is or is not acceptable.</p>",
      "rawMarkdown": "You don't have to be sorry.  As long as you don't break any rule (or any law) then you're fine.  The comments here are opinions of some kagglers, and they are not defining what is or is not acceptable.",
      "votes": null
    },
    {
      "id": "302537",
      "postDate": "03/24/2018 09:01:42",
      "content": "<p>As long as kernels get votes proportionally to how they score on the LB blending will keep going.  And the fact that these kernels receive lots of votes is an indication that the community favor these kernels (as well as kernels with nice graphics).</p>\n\n<p>Just saying ;)</p>",
      "rawMarkdown": "As long as kernels get votes proportionally to how they score on the LB blending will keep going.  And the fact that these kernels receive lots of votes is an indication that the community favor these kernels (as well as kernels with nice graphics).\n\nJust saying ;)",
      "votes": null
    },
    {
      "id": "302586",
      "postDate": "03/24/2018 11:43:09",
      "content": "<p>This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. </p>\n\n<p>Right now the worst a kernel could receive is 0.</p>",
      "rawMarkdown": "This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. \n\nRight now the worst a kernel could receive is 0.",
      "votes": null
    },
    {
      "id": "302651",
      "postDate": "03/24/2018 14:27:57",
      "content": "<p>Oh, I just realized that there is a kernel competition going on, with winners decided by the number of upvotes.  That's just doubling down on the silliness of kernels.</p>",
      "rawMarkdown": "Oh, I just realized that there is a kernel competition going on, with winners decided by the number of upvotes.  That's just doubling down on the silliness of kernels.",
      "votes": null
    },
    {
      "id": "302662",
      "postDate": "03/24/2018 14:45:05",
      "content": "<p>You get what you compensate for.</p>\n\n<p>That's a sentence I hear about sales reps, but it applies here very well ;)</p>",
      "rawMarkdown": "You get what you compensate for.\n\nThat's a sentence I hear about sales reps, but it applies here very well ;)",
      "votes": null
    },
    {
      "id": "302752",
      "postDate": "03/24/2018 17:50:34",
      "content": "<p>You get what you compensate for? So I'm going to get... hmm... never mind.</p>",
      "rawMarkdown": "You get what you compensate for? So I'm going to get... hmm... never mind.",
      "votes": null
    },
    {
      "id": "302761",
      "postDate": "03/24/2018 18:06:20",
      "content": "<p>I look at this from a more root-cause sort-of analysis:</p>\n\n<ul>\n<li>We're incentivized to get more up-doots. </li>\n<li>Newbies or the unwise value quick score increases (without any substantive increase in value).</li>\n<li>What we would like to incentivize is the sharing of creative, instructive, helpful content.</li>\n</ul>\n\n<p>Therefore, we should change the incentives. Perhaps we should have some sort of novelty rating. </p>\n\n<p>I think a good addition of value would be a more descriptive rating system:</p>\n\n<ul>\n<li>Like</li>\n<li>Interesting</li>\n<li>Instructive</li>\n<li>Novel</li>\n<li>etc.</li>\n</ul>\n\n<p>Then Kaggle can change how the Kernel rankings are measured.</p>",
      "rawMarkdown": "I look at this from a more root-cause sort-of analysis:\n\n * We're incentivized to get more up-doots. \n * Newbies or the unwise value quick score increases (without any substantive increase in value).\n * What we would like to incentivize is the sharing of creative, instructive, helpful content.\n\nTherefore, we should change the incentives. Perhaps we should have some sort of novelty rating. \n\nI think a good addition of value would be a more descriptive rating system:\n\n * Like\n * Interesting\n * Instructive\n * Novel\n * etc.\n\nThen Kaggle can change how the Kernel rankings are measured.",
      "votes": null
    },
    {
      "id": "302777",
      "postDate": "03/24/2018 18:15:58",
      "content": "<p>I’d say those form the two main forms of positive feedback loop going on:</p>\n\n<p><em><strong>Nice Graphics</strong></em>\nPost EDA 1 hour after launch → top of Kernels list → views → upvotes → hotness → more views → upvotes etc...</p>\n\n<p><em><strong>Blends</strong></em>\nPost blending kernel incorporating 8 existing models → high public LB position → views → upvotes → hotness → reach top of <em>Kernels</em> list too → views → upvotes ...</p>\n\n<p>The EDA loop is a <strong><em>virtuous circle</em></strong>: those EDAs save the whole crowd having to do the same thing, gives everyone a headstart on the modeling problem. The quality of work is very high, and it <em>leverages the power of the crowd</em>: if all the EDAers mysteriously stopped, Kaggle would pay someone to do that work. The upvote &amp; view count ecology works for EDAs, it seems that the competitive aspect there is to post earliest, be the first, I’d guess at least partly motivated by \"Hotness\". Forks are enabled for EDAs but I've never noticed an EDA fork, let alone see one rank higher than the original work. The EDA sub-community here is great, they have built up valuable portfolios of work, worthy of a prominent place on their CV.</p>\n\n<p>The blending loop is a <strong><em>vicious circle</em></strong>, mundane leaderboard leapfrog, not really worthy of any more comment.</p>\n\n<p>To steal from <a href=\"https://en.wikipedia.org/wiki/Virtuous_circle_and_vicious_circle\">Wikipedia</a> :\n<em>\"These cycles will continue in the direction of their momentum until an external factor intervenes and breaks the cycle.\"</em></p>\n\n<p>I’m wondering when it will. There’s a really interesting (meta) kernel here called <a href=\"https://www.kaggle.com/mlearn/user-engagement-on-kaggle-competitions\">User engagement on Kaggle competitions</a> . The last line of the conclusion is \"... it remains sad that one (initial) competition is enough for many people.\"</p>\n\n<p>The question is: what proportion of new users are impressed by 'blending'? Does it help retention or turn people away sooner?</p>",
      "rawMarkdown": "I’d say those form the two main forms of positive feedback loop going on:\n\n***Nice Graphics***\nPost EDA 1 hour after launch → top of Kernels list → views → upvotes → hotness → more views → upvotes etc...\n\n***Blends***\nPost blending kernel incorporating 8 existing models → high public LB position → views → upvotes → hotness → reach top of *Kernels* list too → views → upvotes ...\n\nThe EDA loop is a ***virtuous circle***: those EDAs save the whole crowd having to do the same thing, gives everyone a headstart on the modeling problem. The quality of work is very high, and it *leverages the power of the crowd*: if all the EDAers mysteriously stopped, Kaggle would pay someone to do that work. The upvote &amp; view count ecology works for EDAs, it seems that the competitive aspect there is to post earliest, be the first, I’d guess at least partly motivated by \"Hotness\". Forks are enabled for EDAs but I've never noticed an EDA fork, let alone see one rank higher than the original work. The EDA sub-community here is great, they have built up valuable portfolios of work, worthy of a prominent place on their CV.\n\nThe blending loop is a ***vicious circle***, mundane leaderboard leapfrog, not really worthy of any more comment.\n\nTo steal from [Wikipedia][1] :\n*\"These cycles will continue in the direction of their momentum until an external factor intervenes and breaks the cycle.\"*\n\nI’m wondering when it will. There’s a really interesting (meta) kernel here called [User engagement on Kaggle competitions][2] . The last line of the conclusion is \"... it remains sad that one (initial) competition is enough for many people.\"\n\nThe question is: what proportion of new users are impressed by 'blending'? Does it help retention or turn people away sooner?\n\n\n  [1]: https://en.wikipedia.org/wiki/Virtuous_circle_and_vicious_circle\n  [2]: https://www.kaggle.com/mlearn/user-engagement-on-kaggle-competitions",
      "votes": null
    },
    {
      "id": "302804",
      "postDate": "03/24/2018 18:44:36",
      "content": "<p>I do blends, and I have very little (aside from amusement) to gain from making them public, since I'm already a kernels master and they have no chance of getting enough votes to win the kernels contest.  But I do think blend kernels are useful as long as they point back to the sources of the original models.  I learned Kaggle mostly by taking high-scoring public kernels and making my own variations on them.  The ability to stack kernels adds another step to this:  you find a high-scoring kernel, look at where it points, and make variations on the inputs.  But since blending is an almost inevitable part of successful final submissions, it makes sense that it should be part of this process too.</p>\n\n<p>I do have a problem with public blends that rely on private inputs.  I think it would be a good idea for Kaggle to put some restrictions on the use of private kernel outputs or outside data uploads as public kernel inputs in the context of a competition.  (And also, as I said in an earlier comment, I think it would be good to allow users to filter out blends in the Kernels tab.)</p>",
      "rawMarkdown": "I do blends, and I have very little (aside from amusement) to gain from making them public, since I'm already a kernels master and they have no chance of getting enough votes to win the kernels contest.  But I do think blend kernels are useful as long as they point back to the sources of the original models.  I learned Kaggle mostly by taking high-scoring public kernels and making my own variations on them.  The ability to stack kernels adds another step to this:  you find a high-scoring kernel, look at where it points, and make variations on the inputs.  But since blending is an almost inevitable part of successful final submissions, it makes sense that it should be part of this process too.\n\nI do have a problem with public blends that rely on private inputs.  I think it would be a good idea for Kaggle to put some restrictions on the use of private kernel outputs or outside data uploads as public kernel inputs in the context of a competition.  (And also, as I said in an earlier comment, I think it would be good to allow users to filter out blends in the Kernels tab.)",
      "votes": null
    },
    {
      "id": "302814",
      "postDate": "03/24/2018 19:04:59",
      "content": "<p>Your post reminds me of VAT (value added tax).  VAT is a tax on added value: you pay that tax on your sales, but you subtract the tax paid by those you buy from.  Translated to kernels, the points you get are based on the votes of your kernel minus the votes of the kernels you forked from, and the points of the kernels that produced the input you are using.  With this metric most blending kernels would have negative value...</p>",
      "rawMarkdown": "Your post reminds me of VAT (value added tax).  VAT is a tax on added value: you pay that tax on your sales, but you subtract the tax paid by those you buy from.  Translated to kernels, the points you get are based on the votes of your kernel minus the votes of the kernels you forked from, and the points of the kernels that produced the input you are using.  With this metric most blending kernels would have negative value...",
      "votes": null
    },
    {
      "id": "302829",
      "postDate": "03/24/2018 19:41:53",
      "content": "<p>I like this idea. VAT should solve most of these problems.</p>",
      "rawMarkdown": "I like this idea. VAT should solve most of these problems.",
      "votes": null
    },
    {
      "id": "302940",
      "postDate": "03/25/2018 03:00:25",
      "content": "<p>With my limited experience on kaggle, I do feel kernels with blend have gone way too overboard over last few competitions. There are people having Kernel expert badge with just Blends as there contribution to this great community. A kernel expert and other contributors holds a lot of respect for me because they have significantly helped in my progressive learning curve. For now it appears to be an easy way of getting cheap up votes. </p>",
      "rawMarkdown": "With my limited experience on kaggle, I do feel kernels with blend have gone way too overboard over last few competitions. There are people having Kernel expert badge with just Blends as there contribution to this great community. A kernel expert and other contributors holds a lot of respect for me because they have significantly helped in my progressive learning curve. For now it appears to be an easy way of getting cheap up votes.",
      "votes": null
    },
    {
      "id": "303198",
      "postDate": "03/25/2018 18:28:19",
      "content": "<p>Votes should count (1 * Performance Tier) times in their respective categories. For example, a Kernel Master's upvote for a kernel should count 4 times toward medals. Medal thresholds could be raised a little and Kagglers with more contribution history would have more sway in the promotion of content. Basically, a \"karma\" system.</p>\n\n<p>Disagree below ; )</p>",
      "rawMarkdown": "Votes should count (1 * Performance Tier) times in their respective categories. For example, a Kernel Master's upvote for a kernel should count 4 times toward medals. Medal thresholds could be raised a little and Kagglers with more contribution history would have more sway in the promotion of content. Basically, a \"karma\" system.\n\nDisagree below ; )",
      "votes": null
    },
    {
      "id": "303405",
      "postDate": "03/26/2018 05:53:18",
      "content": "<p>Good idea, I believe it make sense.</p>",
      "rawMarkdown": "Good idea, I believe it make sense.",
      "votes": null
    },
    {
      "id": "303439",
      "postDate": "03/26/2018 07:25:31",
      "content": "<blockquote>\n  <p>With this metric most blending kernels would have negative value...</p>\n</blockquote>\n\n<p>@CPMP, Bad analogy usually gives nonsense results.  </p>\n\n<p>If we are looking for an analogy, the number of votes is not an analogy of the price of the product, but the number of the buyers of the product in the shop \"All for $1\". Very important thing is, that  many buyers often do not buy new things if they have bought something similar already,  even if the new product is better than the original and the price is only symbolic.</p>",
      "rawMarkdown": "&gt; With this metric most blending kernels would have negative value...\n\n@CPMP, Bad analogy usually gives nonsense results.  \n\nIf we are looking for an analogy, the number of votes is not an analogy of the price of the product, but the number of the buyers of the product in the shop \"All for $1\". Very important thing is, that  many buyers often do not buy new things if they have bought something similar already,  even if the new product is better than the original and the price is only symbolic.",
      "votes": null
    },
    {
      "id": "303465",
      "postDate": "03/26/2018 08:35:38",
      "content": "<p>@Grzegorz,  my analogy is not about the price of products but the sales of a company.  </p>\n\n<p>Isn't the <em>number of number of the buyers of the product in the shop \"All for $1\"</em> exactly the same as the amount of sales in that shop expressed in dollar?  </p>\n\n<p>Seems you just expressed my analogy with your words.  Therefore it must not be that bad. </p>",
      "rawMarkdown": "Grzegorz,  my analogy is not about the price of products but the sales of a company.  \n\nIsn't the *number of number of the buyers of the product in the shop \"All for $1\"* exactly the same as the amount of sales in that shop expressed in dollar?  \n\nSeems you just expressed my analogy with your words.  Therefore it must not be that bad.",
      "votes": null
    },
    {
      "id": "303500",
      "postDate": "03/26/2018 10:09:02",
      "content": "<p>What are these kernAls everybody speak about? ;)</p>",
      "rawMarkdown": "What are these kernAls everybody speak about? ;)",
      "votes": null
    },
    {
      "id": "303533",
      "postDate": "03/26/2018 11:12:01",
      "content": "<p>one could take an example from sites like stackoverflow where votes are not counted unless you are a user who earned your reputation already. </p>",
      "rawMarkdown": "one could take an example from sites like stackoverflow where votes are not counted unless you are a user who earned your reputation already.",
      "votes": null
    },
    {
      "id": "303535",
      "postDate": "03/26/2018 11:15:20",
      "content": "<p>@CPMP, The amount of sales has nothing to do with VAT if the price before and after reselling is the same. And the sales of the companies producing similar products  have nothing to do with the products quality.</p>\n\n<p>But it doesn't matter. I understand your idea of the added value of new (inspired, forked or blended) kernels and generally agree with it.  I  just think,  if the analogies are not good enough, the formulas based on them may be more harmful than useful.</p>",
      "rawMarkdown": "CPMP, The amount of sales has nothing to do with VAT if the price before and after reselling is the same. And the sales of the companies producing similar products  have nothing to do with the products quality.\n\nBut it doesn't matter. I understand your idea of the added value of new (inspired, forked or blended) kernels and generally agree with it.  I  just think,  if the analogies are not good enough, the formulas based on them may be more harmful than useful.",
      "votes": null
    },
    {
      "id": "303548",
      "postDate": "03/26/2018 11:30:36",
      "content": "<p>+1</p>\n\n<p>For discussion here, votes do not count towards medals unless you're above a given level.  I wonder if it is the same for Kernels.</p>",
      "rawMarkdown": "1\n\nFor discussion here, votes do not count towards medals unless you're above a given level.  I wonder if it is the same for Kernels.",
      "votes": null
    },
    {
      "id": "303552",
      "postDate": "03/26/2018 11:33:47",
      "content": "<p>@CPMP - it does seem to be the case for kernels as well, yes.</p>",
      "rawMarkdown": "CPMP - it does seem to be the case for kernels as well, yes.",
      "votes": null
    },
    {
      "id": "303558",
      "postDate": "03/26/2018 11:39:30",
      "content": "<p>@Grzegorz, I'm not sure why you are arguing with.   Glad you agree with the general idea.  I may not have been clear enough, because your idea and mine are exactly the same.  Let me try again.</p>\n\n<p>VAT is the tax on your sales, minus the tax paid by your suppliers.  When there is a single VAT rate then this is equivalent to: VAT is a tax you pay on the difference between your sales, and what you buy from your suppliers, i.e. a VAT is a tax on your gross margin.</p>\n\n<p>It is consistent with your example: if the price before and after reselling is the same (and if you don't buy anything else to sustain your business), then your VAT is 0, as well as your added value.  </p>\n\n<p>I maintain that the number of votes is similar to sales volume (or number of buyers).  When you use that analogy, then your added value is the number of votes you get (your sales), minus the number of votes of the kernels you used (your suppliers sales).</p>\n\n<p>Let me now if this is still not clear enough.</p>",
      "rawMarkdown": "Grzegorz, I'm not sure why you are arguing with.   Glad you agree with the general idea.  I may not have been clear enough, because your idea and mine are exactly the same.  Let me try again.\n\nVAT is the tax on your sales, minus the tax paid by your suppliers.  When there is a single VAT rate then this is equivalent to: VAT is a tax you pay on the difference between your sales, and what you buy from your suppliers, i.e. a VAT is a tax on your gross margin.\n\nIt is consistent with your example: if the price before and after reselling is the same (and if you don't buy anything else to sustain your business), then your VAT is 0, as well as your added value.  \n\nI maintain that the number of votes is similar to sales volume (or number of buyers).  When you use that analogy, then your added value is the number of votes you get (your sales), minus the number of votes of the kernels you used (your suppliers sales).\n\nLet me now if this is still not clear enough.",
      "votes": null
    },
    {
      "id": "303600",
      "postDate": "03/26/2018 12:50:02",
      "content": "<p>Talking about discussion votes, some comments are not bronze in spite of 2 (netto) votes, some of them are bronze in spite of -3 (netto) votes, so there is a dark energy at Kaggle of the strength equal at least 6 votes (peak to peak). \nWe can suppose some notoric haters' down votes are neglected by admins, maybe admins' and masters' votes weight more, etc. , but if there is no transparency in that case, I wouldn't expect any solution clearly separating and evaluating original and blending kernels.  </p>",
      "rawMarkdown": "Talking about discussion votes, some comments are not bronze in spite of 2 (netto) votes, some of them are bronze in spite of -3 (netto) votes, so there is a dark energy at Kaggle of the strength equal at least 6 votes (peak to peak). \nWe can suppose some notoric haters' down votes are neglected by admins, maybe admins' and masters' votes weight more, etc. , but if there is no transparency in that case, I wouldn't expect any solution clearly separating and evaluating original and blending kernels.",
      "votes": null
    },
    {
      "id": "303619",
      "postDate": "03/26/2018 13:09:25",
      "content": "<p>@shivraj: I agree with you here. Blends do not bother me, but it is unfair that they are ranked on the same scale as EDA or generally insightful kernels, and in many cases outperform them in the ranking. I am hoping for a way of weighting their votes, or grading them separately.</p>",
      "rawMarkdown": "shivraj: I agree with you here. Blends do not bother me, but it is unfair that they are ranked on the same scale as EDA or generally insightful kernels, and in many cases outperform them in the ranking. I am hoping for a way of weighting their votes, or grading them separately.",
      "votes": null
    },
    {
      "id": "303630",
      "postDate": "03/26/2018 13:24:32",
      "content": "<p>@CPMP, Let's try at the example:</p>\n\n<p>There is a great LightGBM kernel which obtained 200 votes. I have blended/forked it and obtained 10 votes. Why only 10? Maybe because in the meantime Kagglers have found another ML method much better for the problem, maybe because it is simply the last day of the competition, maybe because of another 99 reasons. Are you 100% sure, my kernel has more or less but negative added value?</p>\n\n<p>If the example above is still not clear, change LightGBM to CNN, me to you, and try again.</p>",
      "rawMarkdown": "CPMP, Let's try at the example:\n\nThere is a great LightGBM kernel which obtained 200 votes. I have blended/forked it and obtained 10 votes. Why only 10? Maybe because in the meantime Kagglers have found another ML method much better for the problem, maybe because it is simply the last day of the competition, maybe because of another 99 reasons. Are you 100% sure, my kernel has more or less but negative added value?\n\nIf the example above is still not clear, change LightGBM to CNN, me to you, and try again.",
      "votes": null
    },
    {
      "id": "303631",
      "postDate": "03/26/2018 13:26:45",
      "content": "<p>I feel for you, being in a similar place.  </p>\n\n<p>I wish down votes would cost points: if you don't have discussion points then you cannot down vote, and each time you down vote you lose points.  I'd put one down vote cost to be a pretty large number of vote points, say 10.</p>",
      "rawMarkdown": "I feel for you, being in a similar place.  \n\nI wish down votes would cost points: if you don't have discussion points then you cannot down vote, and each time you down vote you lose points.  I'd put one down vote cost to be a pretty large number of vote points, say 10.",
      "votes": null
    },
    {
      "id": "303642",
      "postDate": "03/26/2018 13:38:28",
      "content": "<p>@Grzegorz I don't know your kernel hence cannot say if it adds value or not.   I don't know enough to be able to say anything about the quality of your work.  I have no reason to think you are not adding value.</p>\n\n<p>But in general, a kernel that is a weighted average of other kernels output without an explanation of how the weights were selected has negative value.  Weight could have been selected by LB probing for what I know, in which case people reusing it blindly will drop in private LB.   To be specific, in Porto Seguro competition, a blend kernel was very high in public LB, and all those who selected it for final submission dropped by 500 if not more.  That kernel had definitely a negative added value.  No doubt.</p>\n\n<p>Said differently, I am not sure votes reflect the value added by a kernel. Maybe your kernel with 10 votes has very interesting insight that is being overlooked by the community.  And maybe the lgb kernel with lots of votes has no insight compared to other kernels it was derived from.  Without specifics it is hard to say.</p>",
      "rawMarkdown": "Grzegorz I don't know your kernel hence cannot say if it adds value or not.   I don't know enough to be able to say anything about the quality of your work.  I have no reason to think you are not adding value.\n\nBut in general, a kernel that is a weighted average of other kernels output without an explanation of how the weights were selected has negative value.  Weight could have been selected by LB probing for what I know, in which case people reusing it blindly will drop in private LB.   To be specific, in Porto Seguro competition, a blend kernel was very high in public LB, and all those who selected it for final submission dropped by 500 if not more.  That kernel had definitely a negative added value.  No doubt.\n\nSaid differently, I am not sure votes reflect the value added by a kernel. Maybe your kernel with 10 votes has very interesting insight that is being overlooked by the community.  And maybe the lgb kernel with lots of votes has no insight compared to other kernels it was derived from.  Without specifics it is hard to say.",
      "votes": null
    },
    {
      "id": "303654",
      "postDate": "03/26/2018 14:02:29",
      "content": "<p>I have been an admin for 7 years, and I didn't install voting plugin at my forum. IMHO, discussion with arguments is much better. Here, the admins have to fight against haters. @CPMP, Your proposal makes sense. The haters would have less ammo.</p>",
      "rawMarkdown": "I have been an admin for 7 years, and I didn't install voting plugin at my forum. IMHO, discussion with arguments is much better. Here, the admins have to fight against haters. @CPMP, Your proposal makes sense. The haters would have less ammo.",
      "votes": null
    },
    {
      "id": "304005",
      "postDate": "03/26/2018 21:46:57",
      "content": "<p>Very true. I do agree with people's thought of having weighted votes based on one's laurel. </p>\n\n<p>I do see that blends have its primary contribution to the bronze position(obviously lot of sensible blending happens at top positions as well) and one way to overcome them is to up your game and get yourself towards Silver or Gold, but not everybody (fights/can fight) for Gold so it would bother a lot of honest and aspiring learner of Data science who feels Kaggle is a great platform to learn and benchmark their skills.</p>\n\n<p>For now, Most of the Blends signify how lucky was your weights selection.</p>",
      "rawMarkdown": "Very true. I do agree with people's thought of having weighted votes based on one's laurel. \n\nI do see that blends have its primary contribution to the bronze position(obviously lot of sensible blending happens at top positions as well) and one way to overcome them is to up your game and get yourself towards Silver or Gold, but not everybody (fights/can fight) for Gold so it would bother a lot of honest and aspiring learner of Data science who feels Kaggle is a great platform to learn and benchmark their skills.\n\nFor now, Most of the Blends signify how lucky was your weights selection.",
      "votes": null
    },
    {
      "id": "307982",
      "postDate": "04/02/2018 18:50:36",
      "content": "<p>OK, how about <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704\">linear stacking</a> (which is the same thing as blending, really, but more systematic, so maybe you can at least learn the method)?</p>",
      "rawMarkdown": "OK, how about [linear stacking][1] (which is the same thing as blending, really, but more systematic, so maybe you can at least learn the method)?\n\n [1]: https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704",
      "votes": null
    },
    {
      "id": "308110",
      "postDate": "04/03/2018 00:36:08",
      "content": "<p>However you can't allow negative votes because then you will get overrun with robot profiles created to downvote rivals, esp. when there's a kernel competition. Perhaps only allow existing bona-fide members with &gt; threshold longevity, no. of upvoted posts and submissions to have a downvote privilege.</p>\n\n<blockquote>\n  <p><strong>quantumgeek wrote</strong></p>\n  \n  <blockquote>\n    <p>This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. </p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "However you can't allow negative votes because then you will get overrun with robot profiles created to downvote rivals, esp. when there's a kernel competition. Perhaps only allow existing bona-fide members with &gt; threshold longevity, no. of upvoted posts and submissions to have a downvote privilege.\n\n&gt; **quantumgeek wrote**\n&gt; \n&gt; &gt; This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters.",
      "votes": null
    },
    {
      "id": "323817",
      "postDate": "05/06/2018 10:19:50",
      "content": "<p>Personnally, I agree that emphasizing too much blending is an issue as this shows only the last part of the computation and put under the carpet the hard pre blending work. But it is still usefull to be aware that blending can achiever higher score and is in a sense a way to do voting classifyers or any on top ensemble classifyer.  So in short for newbees like me they are usefull but they are not the most important ones.</p>",
      "rawMarkdown": "Personnally, I agree that emphasizing too much blending is an issue as this shows only the last part of the computation and put under the carpet the hard pre blending work. But it is still usefull to be aware that blending can achiever higher score and is in a sense a way to do voting classifyers or any on top ensemble classifyer.  So in short for newbees like me they are usefull but they are not the most important ones.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 301756,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/23/2018 07:44:53",
      "content": "<p>I agree.</p>\n\n<p>These blenders have caused more harm than good in my opinion:</p>\n\n<ul>\n<li><p>people don't want to share their kernels with CREATIVE AND NEW ideas so that they are not put in the blender</p></li>\n<li><p>which leads to no learning from each other and no building upon each other's ideas and knowledge.</p></li>\n</ul>\n\n<p>I personally, would never share my kernel before the competition ends purely because of this. There should be a solution soon. </p>",
      "votes": null,
      "replies": [
        {
          "id": 301759,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/23/2018 07:48:54",
          "content": "<p>In principle i agree, but - barring manual inspection of kernels by Kaggle - i dont see an automated way to do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 301808,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "03/23/2018 09:24:41",
          "content": "<p>How about making it private if it surpasses 10% of bronze positions. Linking another discussion here -<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802</a> \nPlease do share your opinion Konrad. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 301815,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/23/2018 09:40:13",
          "content": "<p>I understand your line of reasoning here - I think - but that particular solution would give an advantage to the people who grabbed it first. Personally, I don't have that much problem with the blends: they serve as a sort of filter on public kernel ideas, showing me which ones might be useful for implementing myself (i.e. using them in a proper validation setup etc). </p>\n\n<p>There is also a cynical aspect to my take on the problem: there will always be people just gaming the system. I have been riled for quite a while by the fact that in the global ranking there was always this one guy ahead of me, who had a very simple modus operandi: he registered for every competition there was, submitted the benchmark and never looked back. Pure scale effect was enough to put him ahead of me, but I realized that hunting down people like that carried a strong risk of doing more harm than good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302419,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "03/24/2018 02:53:36",
          "content": "<p>Just release two kernels -- one with your idea and one with your idea included in the Blend of the Day. Presto, now you have two silver medals!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 301758,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/23/2018 07:47:58",
      "content": "<p>One thing that comes to my mind, which for sure has certain disadvantages as I haven't thought it over is to be able to share publicly the kernel BUT not the output file. In such a way those blender boys will have to at least spend some memory power ( I am talking about computer memory, as we all know that the brain power is limited in such cases)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301796,
      "author_name": "comartel",
      "author_url": "",
      "post_date": "03/23/2018 09:05:26",
      "content": "<p>Not sure if possible, but you could post your kernel and just add a line to generate a false output file ? That way you can share your ideas but not your results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301799,
      "author_name": "nvarganov",
      "author_url": "",
      "post_date": "03/23/2018 09:12:17",
      "content": "<p>Absolutely agree.</p>\n\n<p>Only third week of this competition, I think it will be interesting to find some interesting patterns in the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 301918,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "03/23/2018 13:09:56",
          "content": "<p>There are definitely <a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\">interesting patterns</a> in there :)</p>\n\n<p>(Edit) Preview:\n<a href=\"https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\"><img src=\"https://s31.postimg.cc/c8e6qsomz/channel_105a.png\" alt=\"Channel 105 clicks\"></a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 301929,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/23/2018 13:30:07",
          "content": "<p>And <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates/notebook\">here</a> as well ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 301806,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "03/23/2018 09:23:26",
      "content": "<p>Agreed.\nPlease read this discussion and share your opinion.\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301817,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "03/23/2018 09:50:24",
      "content": "<p>Very good point!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301848,
      "author_name": "jtrotman",
      "author_url": "",
      "post_date": "03/23/2018 11:16:36",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/301848/8854/blend_moar_kernalz.png\" alt=\"enter image description here\"></p>\n\n<h2><a href=\"https://livingthing.danmackinlay.name/deep_learning.html\">adapted from this</a></h2>",
      "votes": null,
      "replies": [
        {
          "id": 301990,
          "author_name": "baomengjiao",
          "author_url": "",
          "post_date": "03/23/2018 14:49:06",
          "content": "<p>You are a Soul Painter(灵魂画师)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302011,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "03/23/2018 15:16:55",
          "content": "<p>Haha, thanks! It's not really my work, the (brilliant) original is <a href=\"https://livingthing.danmackinlay.name/deep_learning.html\">here</a>, I just changed the text in the bottom half, I have made the link in the post a bit clearer...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 301859,
      "author_name": "mzacks",
      "author_url": "",
      "post_date": "03/23/2018 11:41:36",
      "content": "<p>one solution could be for kernels not to accept user uploads or cross-references to other kernels outputs. each kernel will have to build the output and any blending it might perform starting from the competition datasource.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301865,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/23/2018 11:54:39",
      "content": "<p>Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 302028,
          "author_name": "dmitriyguller",
          "author_url": "",
          "post_date": "03/23/2018 15:43:39",
          "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <blockquote>\n    <p>Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.</p>\n  </blockquote>\n</blockquote>\n\n<p>Yes, public kernels are usually poison.  Experienced Kagglers know that and benefit from less experienced Kagglers not knowing that.  They are an effective education tool, but more often than not they are effective at teaching precisely the wrong and dangerous lessons.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302191,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "03/23/2018 19:27:30",
      "content": "<p>I second @spongebob's concern and would like to share my thoughts. I of course, respect the choice and interest of everyone in this diverse community and consider it okay and tolerable until it's a normal blend (at least with some thought process on blending). </p>\n\n<p>The problem starts when competition becomes the game of random lucky numbers at the very early stage and the idea behind improving base models is forgotten. In my personal opinion, that would be the last thing to try when no time is left to improve any models. </p>\n\n<p>My suggestion would be to <strong>separate/ filter</strong>  blending kernels from real kernels in kernel's tab so that people can choose what they want to explore/learn and when. Alternate option could be to post blending script in discussion section instead of kernels so that true learning material can be accessible easily for majority of Kagglers who are interested to gain and contribute something positive. </p>",
      "votes": null,
      "replies": [
        {
          "id": 302291,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/23/2018 21:36:47",
          "content": "<p>I like the idea of having a filter for blends in the Kernels tab. I think it would probably work if they simply require us to declare whether or not it's a blend (maybe when a kernel goes public), because I don't think many people would deliberately lie.  But if necessary they could allow us to filter out kernels with multiple data sources, which is a noisy proxy for blends.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302247,
      "author_name": "nicapotato",
      "author_url": "",
      "post_date": "03/23/2018 20:54:12",
      "content": "<p>As a community, we just have to reward actual creative contribution rather than blends. And hey, if the blend is thoughtful and backed by theory, take my upvote!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302296,
      "author_name": "koseemr",
      "author_url": "",
      "post_date": "03/23/2018 21:53:08",
      "content": "<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302307,
      "author_name": "gopisaran",
      "author_url": "",
      "post_date": "03/23/2018 22:19:19",
      "content": "<p>I am a newbie to Kaggle I entirely went overboard with blending in  <a href=\"https://www.kaggle.com/gopisaran/ensemble-of-2-ensembles-lb-0-9696\">this Kernel</a></p>\n\n<p>It was a mere experiment, and I had no intentions to damage the ecology of Kaggle. I am sorry !!\nI now understand that the Kaggle community thrives on improving base models, exploring and learning which is going to be my focus. This discussion was a great learning for a newbie like me. Thanks guys for bringing this point up. Appreciate it!​</p>",
      "votes": null,
      "replies": [
        {
          "id": 302535,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/24/2018 08:58:56",
          "content": "<p>You don't have to be sorry.  As long as you don't break any rule (or any law) then you're fine.  The comments here are opinions of some kagglers, and they are not defining what is or is not acceptable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302319,
      "author_name": "pvlima",
      "author_url": "",
      "post_date": "03/23/2018 22:38:57",
      "content": "<p>These blends are somehow like <a href=\"https://en.wikipedia.org/wiki/Prohibition_in_the_United_States\">these other blends</a>. Prohibition was not the best approach. I'm sure there are better ways to deal with the problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302495,
      "author_name": "reachkishore",
      "author_url": "",
      "post_date": "03/24/2018 07:06:52",
      "content": "<p>I Partially agree with @spongebob and I respect every one's views. \n<strong>I agree Blending is also an art and a piece of data science solutions framework. <br>\nBut we should also encourage the Diverse solutions and the effort put in them.</strong> </p>\n\n<p><strong>My suggestion to @Kaggle is</strong> <br>\nto make an <strong>intermediate dead line like ( one/two month after launch of competition)</strong> , <br></p>\n\n<ul>\n<li>Till this date kernels can be made public.</li>\n<li>later this date all new kernels will become private kernels. </li>\n<li>These can be made public post competition. <br>\nBy doing this Blending can still be be done using past public kernels also it encourages Diverse solutions nearing to competition closure. \nThank you</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302537,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/24/2018 09:01:42",
      "content": "<p>As long as kernels get votes proportionally to how they score on the LB blending will keep going.  And the fact that these kernels receive lots of votes is an indication that the community favor these kernels (as well as kernels with nice graphics).</p>\n\n<p>Just saying ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 302777,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "03/24/2018 18:15:58",
          "content": "<p>I’d say those form the two main forms of positive feedback loop going on:</p>\n\n<p><em><strong>Nice Graphics</strong></em>\nPost EDA 1 hour after launch → top of Kernels list → views → upvotes → hotness → more views → upvotes etc...</p>\n\n<p><em><strong>Blends</strong></em>\nPost blending kernel incorporating 8 existing models → high public LB position → views → upvotes → hotness → reach top of <em>Kernels</em> list too → views → upvotes ...</p>\n\n<p>The EDA loop is a <strong><em>virtuous circle</em></strong>: those EDAs save the whole crowd having to do the same thing, gives everyone a headstart on the modeling problem. The quality of work is very high, and it <em>leverages the power of the crowd</em>: if all the EDAers mysteriously stopped, Kaggle would pay someone to do that work. The upvote &amp; view count ecology works for EDAs, it seems that the competitive aspect there is to post earliest, be the first, I’d guess at least partly motivated by \"Hotness\". Forks are enabled for EDAs but I've never noticed an EDA fork, let alone see one rank higher than the original work. The EDA sub-community here is great, they have built up valuable portfolios of work, worthy of a prominent place on their CV.</p>\n\n<p>The blending loop is a <strong><em>vicious circle</em></strong>, mundane leaderboard leapfrog, not really worthy of any more comment.</p>\n\n<p>To steal from <a href=\"https://en.wikipedia.org/wiki/Virtuous_circle_and_vicious_circle\">Wikipedia</a> :\n<em>\"These cycles will continue in the direction of their momentum until an external factor intervenes and breaks the cycle.\"</em></p>\n\n<p>I’m wondering when it will. There’s a really interesting (meta) kernel here called <a href=\"https://www.kaggle.com/mlearn/user-engagement-on-kaggle-competitions\">User engagement on Kaggle competitions</a> . The last line of the conclusion is \"... it remains sad that one (initial) competition is enough for many people.\"</p>\n\n<p>The question is: what proportion of new users are impressed by 'blending'? Does it help retention or turn people away sooner?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302586,
      "author_name": "yk1598",
      "author_url": "",
      "post_date": "03/24/2018 11:43:09",
      "content": "<p>This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. </p>\n\n<p>Right now the worst a kernel could receive is 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 308110,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "04/03/2018 00:36:08",
          "content": "<p>However you can't allow negative votes because then you will get overrun with robot profiles created to downvote rivals, esp. when there's a kernel competition. Perhaps only allow existing bona-fide members with &gt; threshold longevity, no. of upvoted posts and submissions to have a downvote privilege.</p>\n\n<blockquote>\n  <p><strong>quantumgeek wrote</strong></p>\n  \n  <blockquote>\n    <p>This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. </p>\n  </blockquote>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302651,
      "author_name": "dmitriyguller",
      "author_url": "",
      "post_date": "03/24/2018 14:27:57",
      "content": "<p>Oh, I just realized that there is a kernel competition going on, with winners decided by the number of upvotes.  That's just doubling down on the silliness of kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 302662,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/24/2018 14:45:05",
          "content": "<p>You get what you compensate for.</p>\n\n<p>That's a sentence I hear about sales reps, but it applies here very well ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302752,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "03/24/2018 17:50:34",
          "content": "<p>You get what you compensate for? So I'm going to get... hmm... never mind.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302761,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "03/24/2018 18:06:20",
      "content": "<p>I look at this from a more root-cause sort-of analysis:</p>\n\n<ul>\n<li>We're incentivized to get more up-doots. </li>\n<li>Newbies or the unwise value quick score increases (without any substantive increase in value).</li>\n<li>What we would like to incentivize is the sharing of creative, instructive, helpful content.</li>\n</ul>\n\n<p>Therefore, we should change the incentives. Perhaps we should have some sort of novelty rating. </p>\n\n<p>I think a good addition of value would be a more descriptive rating system:</p>\n\n<ul>\n<li>Like</li>\n<li>Interesting</li>\n<li>Instructive</li>\n<li>Novel</li>\n<li>etc.</li>\n</ul>\n\n<p>Then Kaggle can change how the Kernel rankings are measured.</p>",
      "votes": null,
      "replies": [
        {
          "id": 302814,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/24/2018 19:04:59",
          "content": "<p>Your post reminds me of VAT (value added tax).  VAT is a tax on added value: you pay that tax on your sales, but you subtract the tax paid by those you buy from.  Translated to kernels, the points you get are based on the votes of your kernel minus the votes of the kernels you forked from, and the points of the kernels that produced the input you are using.  With this metric most blending kernels would have negative value...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302829,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "03/24/2018 19:41:53",
          "content": "<p>I like this idea. VAT should solve most of these problems.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303405,
          "author_name": "lccever",
          "author_url": "",
          "post_date": "03/26/2018 05:53:18",
          "content": "<p>Good idea, I believe it make sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303439,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "03/26/2018 07:25:31",
          "content": "<blockquote>\n  <p>With this metric most blending kernels would have negative value...</p>\n</blockquote>\n\n<p>@CPMP, Bad analogy usually gives nonsense results.  </p>\n\n<p>If we are looking for an analogy, the number of votes is not an analogy of the price of the product, but the number of the buyers of the product in the shop \"All for $1\". Very important thing is, that  many buyers often do not buy new things if they have bought something similar already,  even if the new product is better than the original and the price is only symbolic.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303465,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 08:35:38",
          "content": "<p>@Grzegorz,  my analogy is not about the price of products but the sales of a company.  </p>\n\n<p>Isn't the <em>number of number of the buyers of the product in the shop \"All for $1\"</em> exactly the same as the amount of sales in that shop expressed in dollar?  </p>\n\n<p>Seems you just expressed my analogy with your words.  Therefore it must not be that bad. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303535,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "03/26/2018 11:15:20",
          "content": "<p>@CPMP, The amount of sales has nothing to do with VAT if the price before and after reselling is the same. And the sales of the companies producing similar products  have nothing to do with the products quality.</p>\n\n<p>But it doesn't matter. I understand your idea of the added value of new (inspired, forked or blended) kernels and generally agree with it.  I  just think,  if the analogies are not good enough, the formulas based on them may be more harmful than useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303558,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 11:39:30",
          "content": "<p>@Grzegorz, I'm not sure why you are arguing with.   Glad you agree with the general idea.  I may not have been clear enough, because your idea and mine are exactly the same.  Let me try again.</p>\n\n<p>VAT is the tax on your sales, minus the tax paid by your suppliers.  When there is a single VAT rate then this is equivalent to: VAT is a tax you pay on the difference between your sales, and what you buy from your suppliers, i.e. a VAT is a tax on your gross margin.</p>\n\n<p>It is consistent with your example: if the price before and after reselling is the same (and if you don't buy anything else to sustain your business), then your VAT is 0, as well as your added value.  </p>\n\n<p>I maintain that the number of votes is similar to sales volume (or number of buyers).  When you use that analogy, then your added value is the number of votes you get (your sales), minus the number of votes of the kernels you used (your suppliers sales).</p>\n\n<p>Let me now if this is still not clear enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303630,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "03/26/2018 13:24:32",
          "content": "<p>@CPMP, Let's try at the example:</p>\n\n<p>There is a great LightGBM kernel which obtained 200 votes. I have blended/forked it and obtained 10 votes. Why only 10? Maybe because in the meantime Kagglers have found another ML method much better for the problem, maybe because it is simply the last day of the competition, maybe because of another 99 reasons. Are you 100% sure, my kernel has more or less but negative added value?</p>\n\n<p>If the example above is still not clear, change LightGBM to CNN, me to you, and try again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303642,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 13:38:28",
          "content": "<p>@Grzegorz I don't know your kernel hence cannot say if it adds value or not.   I don't know enough to be able to say anything about the quality of your work.  I have no reason to think you are not adding value.</p>\n\n<p>But in general, a kernel that is a weighted average of other kernels output without an explanation of how the weights were selected has negative value.  Weight could have been selected by LB probing for what I know, in which case people reusing it blindly will drop in private LB.   To be specific, in Porto Seguro competition, a blend kernel was very high in public LB, and all those who selected it for final submission dropped by 500 if not more.  That kernel had definitely a negative added value.  No doubt.</p>\n\n<p>Said differently, I am not sure votes reflect the value added by a kernel. Maybe your kernel with 10 votes has very interesting insight that is being overlooked by the community.  And maybe the lgb kernel with lots of votes has no insight compared to other kernels it was derived from.  Without specifics it is hard to say.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302804,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/24/2018 18:44:36",
      "content": "<p>I do blends, and I have very little (aside from amusement) to gain from making them public, since I'm already a kernels master and they have no chance of getting enough votes to win the kernels contest.  But I do think blend kernels are useful as long as they point back to the sources of the original models.  I learned Kaggle mostly by taking high-scoring public kernels and making my own variations on them.  The ability to stack kernels adds another step to this:  you find a high-scoring kernel, look at where it points, and make variations on the inputs.  But since blending is an almost inevitable part of successful final submissions, it makes sense that it should be part of this process too.</p>\n\n<p>I do have a problem with public blends that rely on private inputs.  I think it would be a good idea for Kaggle to put some restrictions on the use of private kernel outputs or outside data uploads as public kernel inputs in the context of a competition.  (And also, as I said in an earlier comment, I think it would be good to allow users to filter out blends in the Kernels tab.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302940,
      "author_name": "shivrajp",
      "author_url": "",
      "post_date": "03/25/2018 03:00:25",
      "content": "<p>With my limited experience on kaggle, I do feel kernels with blend have gone way too overboard over last few competitions. There are people having Kernel expert badge with just Blends as there contribution to this great community. A kernel expert and other contributors holds a lot of respect for me because they have significantly helped in my progressive learning curve. For now it appears to be an easy way of getting cheap up votes. </p>",
      "votes": null,
      "replies": [
        {
          "id": 303619,
          "author_name": "gimunu",
          "author_url": "",
          "post_date": "03/26/2018 13:09:25",
          "content": "<p>@shivraj: I agree with you here. Blends do not bother me, but it is unfair that they are ranked on the same scale as EDA or generally insightful kernels, and in many cases outperform them in the ranking. I am hoping for a way of weighting their votes, or grading them separately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 304005,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "03/26/2018 21:46:57",
          "content": "<p>Very true. I do agree with people's thought of having weighted votes based on one's laurel. </p>\n\n<p>I do see that blends have its primary contribution to the bronze position(obviously lot of sensible blending happens at top positions as well) and one way to overcome them is to up your game and get yourself towards Silver or Gold, but not everybody (fights/can fight) for Gold so it would bother a lot of honest and aspiring learner of Data science who feels Kaggle is a great platform to learn and benchmark their skills.</p>\n\n<p>For now, Most of the Blends signify how lucky was your weights selection.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303198,
      "author_name": "jmbull",
      "author_url": "",
      "post_date": "03/25/2018 18:28:19",
      "content": "<p>Votes should count (1 * Performance Tier) times in their respective categories. For example, a Kernel Master's upvote for a kernel should count 4 times toward medals. Medal thresholds could be raised a little and Kagglers with more contribution history would have more sway in the promotion of content. Basically, a \"karma\" system.</p>\n\n<p>Disagree below ; )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 303500,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/26/2018 10:09:02",
      "content": "<p>What are these kernAls everybody speak about? ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 303533,
      "author_name": "samehif",
      "author_url": "",
      "post_date": "03/26/2018 11:12:01",
      "content": "<p>one could take an example from sites like stackoverflow where votes are not counted unless you are a user who earned your reputation already. </p>",
      "votes": null,
      "replies": [
        {
          "id": 303548,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 11:30:36",
          "content": "<p>+1</p>\n\n<p>For discussion here, votes do not count towards medals unless you're above a given level.  I wonder if it is the same for Kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303552,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/26/2018 11:33:47",
          "content": "<p>@CPMP - it does seem to be the case for kernels as well, yes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303600,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "03/26/2018 12:50:02",
          "content": "<p>Talking about discussion votes, some comments are not bronze in spite of 2 (netto) votes, some of them are bronze in spite of -3 (netto) votes, so there is a dark energy at Kaggle of the strength equal at least 6 votes (peak to peak). \nWe can suppose some notoric haters' down votes are neglected by admins, maybe admins' and masters' votes weight more, etc. , but if there is no transparency in that case, I wouldn't expect any solution clearly separating and evaluating original and blending kernels.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303631,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/26/2018 13:26:45",
          "content": "<p>I feel for you, being in a similar place.  </p>\n\n<p>I wish down votes would cost points: if you don't have discussion points then you cannot down vote, and each time you down vote you lose points.  I'd put one down vote cost to be a pretty large number of vote points, say 10.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303654,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "03/26/2018 14:02:29",
          "content": "<p>I have been an admin for 7 years, and I didn't install voting plugin at my forum. IMHO, discussion with arguments is much better. Here, the admins have to fight against haters. @CPMP, Your proposal makes sense. The haters would have less ammo.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307982,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/02/2018 18:50:36",
      "content": "<p>OK, how about <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704\">linear stacking</a> (which is the same thing as blending, really, but more systematic, so maybe you can at least learn the method)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323817,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/06/2018 10:19:50",
      "content": "<p>Personnally, I agree that emphasizing too much blending is an issue as this shows only the last part of the computation and put under the carpet the hard pre blending work. But it is still usefull to be aware that blending can achiever higher score and is in a sense a way to do voting classifyers or any on top ensemble classifyer.  So in short for newbees like me they are usefull but they are not the most important ones.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "301611": "These blending kernels is useless to us. Even some insights about how to blend deserve 99 votes. Guys use the kernals' output to generate a blending result just want to Fraud Vote.(骗赞) This damages the ecology of kaggle. \n\n\nI struggle with this in toxic comment classification. \nDo something interesting and amazing in such a short life.",
    "301756": "I agree.\n\nThese blenders have caused more harm than good in my opinion:\n\n- people don't want to share their kernels with CREATIVE AND NEW ideas so that they are not put in the blender\n\n- which leads to no learning from each other and no building upon each other's ideas and knowledge.\n\nI personally, would never share my kernel before the competition ends purely because of this. There should be a solution soon.",
    "301758": "One thing that comes to my mind, which for sure has certain disadvantages as I haven't thought it over is to be able to share publicly the kernel BUT not the output file. In such a way those blender boys will have to at least spend some memory power ( I am talking about computer memory, as we all know that the brain power is limited in such cases)",
    "301759": "In principle i agree, but - barring manual inspection of kernels by Kaggle - i dont see an automated way to do it.",
    "301796": "Not sure if possible, but you could post your kernel and just add a line to generate a false output file ? That way you can share your ideas but not your results.",
    "301799": "Absolutely agree.\n\nOnly third week of this competition, I think it will be interesting to find some interesting patterns in the data.",
    "301806": "Agreed.\nPlease read this discussion and share your opinion.\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802",
    "301808": "How about making it private if it surpasses 10% of bronze positions. Linking another discussion here -https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52802 \nPlease do share your opinion Konrad. Thanks!",
    "301815": "I understand your line of reasoning here - I think - but that particular solution would give an advantage to the people who grabbed it first. Personally, I don't have that much problem with the blends: they serve as a sort of filter on public kernel ideas, showing me which ones might be useful for implementing myself (i.e. using them in a proper validation setup etc). \n\nThere is also a cynical aspect to my take on the problem: there will always be people just gaming the system. I have been riled for quite a while by the fact that in the global ranking there was always this one guy ahead of me, who had a very simple modus operandi: he registered for every competition there was, submitted the benchmark and never looked back. Pure scale effect was enough to put him ahead of me, but I realized that hunting down people like that carried a strong risk of doing more harm than good.",
    "301817": "Very good point!",
    "301848": "![enter image description here][1]\n## [adapted from this][2]\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/301848/8854/blend_moar_kernalz.png\n  [2]: https://livingthing.danmackinlay.name/deep_learning.html",
    "301859": "one solution could be for kernels not to accept user uploads or cross-references to other kernels outputs. each kernel will have to build the output and any blending it might perform starting from the competition datasource.",
    "301865": "Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.",
    "301918": "There are definitely [interesting patterns][1] in there :)\n\n(Edit) Preview:\n[![Channel 105 clicks][2]][1]\n\n  [1]: https://www.kaggle.com/jtrotman/eda-talkingdata-temporal-click-count-plots\n  [2]: https://s31.postimg.cc/c8e6qsomz/channel_105a.png",
    "301929": "And [here][1] as well ;)\n\n\n  [1]: https://www.kaggle.com/cpmpml/ip-download-rates/notebook",
    "301990": "You are a Soul Painter(灵魂画师)",
    "302011": "Haha, thanks! It's not really my work, the (brilliant) original is [here][1], I just changed the text in the bottom half, I have made the link in the post a bit clearer...\n\n\n  [1]: https://livingthing.danmackinlay.name/deep_learning.html",
    "302028": "&gt; **CPMP wrote**\n&gt; \n&gt; &gt; Just add your twist on top of it, and you'll be above it in LB.  In Porto Seguro a blend kernel was massively used.  It was very well ranked on public LB at the end.  However, all those who selected it as final sub where ranked below the 600th rank in the private LB, because it was overfiting a bit compared to models tuned with cross validation.\n\nYes, public kernels are usually poison.  Experienced Kagglers know that and benefit from less experienced Kagglers not knowing that.  They are an effective education tool, but more often than not they are effective at teaching precisely the wrong and dangerous lessons.",
    "302191": "I second @spongebob's concern and would like to share my thoughts. I of course, respect the choice and interest of everyone in this diverse community and consider it okay and tolerable until it's a normal blend (at least with some thought process on blending). \n\nThe problem starts when competition becomes the game of random lucky numbers at the very early stage and the idea behind improving base models is forgotten. In my personal opinion, that would be the last thing to try when no time is left to improve any models. \n\nMy suggestion would be to **separate/ filter**  blending kernels from real kernels in kernel's tab so that people can choose what they want to explore/learn and when. Alternate option could be to post blending script in discussion section instead of kernels so that true learning material can be accessible easily for majority of Kagglers who are interested to gain and contribute something positive.",
    "302247": "As a community, we just have to reward actual creative contribution rather than blends. And hey, if the blend is thoughtful and backed by theory, take my upvote!",
    "302291": "I like the idea of having a filter for blends in the Kernels tab. I think it would probably work if they simply require us to declare whether or not it's a blend (maybe when a kernel goes public), because I don't think many people would deliberately lie.  But if necessary they could allow us to filter out kernels with multiple data sources, which is a noisy proxy for blends.",
    "302296": "thanks",
    "302307": "I am a newbie to Kaggle I entirely went overboard with blending in  [this Kernel][1]\n\n\n  [1]: https://www.kaggle.com/gopisaran/ensemble-of-2-ensembles-lb-0-9696\n\nIt was a mere experiment, and I had no intentions to damage the ecology of Kaggle. I am sorry !!\nI now understand that the Kaggle community thrives on improving base models, exploring and learning which is going to be my focus. This discussion was a great learning for a newbie like me. Thanks guys for bringing this point up. Appreciate it!​",
    "302319": "These blends are somehow like [these other blends](https://en.wikipedia.org/wiki/Prohibition_in_the_United_States). Prohibition was not the best approach. I'm sure there are better ways to deal with the problem.",
    "302419": "Just release two kernels -- one with your idea and one with your idea included in the Blend of the Day. Presto, now you have two silver medals!",
    "302495": "I Partially agree with @spongebob and I respect every one's views. \n**I agree Blending is also an art and a piece of data science solutions framework. <br>\nBut we should also encourage the Diverse solutions and the effort put in them.** \n\n**My suggestion to @Kaggle is** <br>\nto make an **intermediate dead line like ( one/two month after launch of competition)** , <br>\n\n - Till this date kernels can be made public.\n -  later this date all new kernels will become private kernels. \n -  These can be made public post competition. <br>\nBy doing this Blending can still be be done using past public kernels also it encourages Diverse solutions nearing to competition closure. \nThank you",
    "302535": "You don't have to be sorry.  As long as you don't break any rule (or any law) then you're fine.  The comments here are opinions of some kagglers, and they are not defining what is or is not acceptable.",
    "302537": "As long as kernels get votes proportionally to how they score on the LB blending will keep going.  And the fact that these kernels receive lots of votes is an indication that the community favor these kernels (as well as kernels with nice graphics).\n\nJust saying ;)",
    "302586": "This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters. \n\nRight now the worst a kernel could receive is 0.",
    "302651": "Oh, I just realized that there is a kernel competition going on, with winners decided by the number of upvotes.  That's just doubling down on the silliness of kernels.",
    "302662": "You get what you compensate for.\n\nThat's a sentence I hear about sales reps, but it applies here very well ;)",
    "302752": "You get what you compensate for? So I'm going to get... hmm... never mind.",
    "302761": "I look at this from a more root-cause sort-of analysis:\n\n * We're incentivized to get more up-doots. \n * Newbies or the unwise value quick score increases (without any substantive increase in value).\n * What we would like to incentivize is the sharing of creative, instructive, helpful content.\n\nTherefore, we should change the incentives. Perhaps we should have some sort of novelty rating. \n\nI think a good addition of value would be a more descriptive rating system:\n\n * Like\n * Interesting\n * Instructive\n * Novel\n * etc.\n\nThen Kaggle can change how the Kernel rankings are measured.",
    "302777": "I’d say those form the two main forms of positive feedback loop going on:\n\n***Nice Graphics***\nPost EDA 1 hour after launch → top of Kernels list → views → upvotes → hotness → more views → upvotes etc...\n\n***Blends***\nPost blending kernel incorporating 8 existing models → high public LB position → views → upvotes → hotness → reach top of *Kernels* list too → views → upvotes ...\n\nThe EDA loop is a ***virtuous circle***: those EDAs save the whole crowd having to do the same thing, gives everyone a headstart on the modeling problem. The quality of work is very high, and it *leverages the power of the crowd*: if all the EDAers mysteriously stopped, Kaggle would pay someone to do that work. The upvote &amp; view count ecology works for EDAs, it seems that the competitive aspect there is to post earliest, be the first, I’d guess at least partly motivated by \"Hotness\". Forks are enabled for EDAs but I've never noticed an EDA fork, let alone see one rank higher than the original work. The EDA sub-community here is great, they have built up valuable portfolios of work, worthy of a prominent place on their CV.\n\nThe blending loop is a ***vicious circle***, mundane leaderboard leapfrog, not really worthy of any more comment.\n\nTo steal from [Wikipedia][1] :\n*\"These cycles will continue in the direction of their momentum until an external factor intervenes and breaks the cycle.\"*\n\nI’m wondering when it will. There’s a really interesting (meta) kernel here called [User engagement on Kaggle competitions][2] . The last line of the conclusion is \"... it remains sad that one (initial) competition is enough for many people.\"\n\nThe question is: what proportion of new users are impressed by 'blending'? Does it help retention or turn people away sooner?\n\n\n  [1]: https://en.wikipedia.org/wiki/Virtuous_circle_and_vicious_circle\n  [2]: https://www.kaggle.com/mlearn/user-engagement-on-kaggle-competitions",
    "302804": "I do blends, and I have very little (aside from amusement) to gain from making them public, since I'm already a kernels master and they have no chance of getting enough votes to win the kernels contest.  But I do think blend kernels are useful as long as they point back to the sources of the original models.  I learned Kaggle mostly by taking high-scoring public kernels and making my own variations on them.  The ability to stack kernels adds another step to this:  you find a high-scoring kernel, look at where it points, and make variations on the inputs.  But since blending is an almost inevitable part of successful final submissions, it makes sense that it should be part of this process too.\n\nI do have a problem with public blends that rely on private inputs.  I think it would be a good idea for Kaggle to put some restrictions on the use of private kernel outputs or outside data uploads as public kernel inputs in the context of a competition.  (And also, as I said in an earlier comment, I think it would be good to allow users to filter out blends in the Kernels tab.)",
    "302814": "Your post reminds me of VAT (value added tax).  VAT is a tax on added value: you pay that tax on your sales, but you subtract the tax paid by those you buy from.  Translated to kernels, the points you get are based on the votes of your kernel minus the votes of the kernels you forked from, and the points of the kernels that produced the input you are using.  With this metric most blending kernels would have negative value...",
    "302829": "I like this idea. VAT should solve most of these problems.",
    "302940": "With my limited experience on kaggle, I do feel kernels with blend have gone way too overboard over last few competitions. There are people having Kernel expert badge with just Blends as there contribution to this great community. A kernel expert and other contributors holds a lot of respect for me because they have significantly helped in my progressive learning curve. For now it appears to be an easy way of getting cheap up votes.",
    "303198": "Votes should count (1 * Performance Tier) times in their respective categories. For example, a Kernel Master's upvote for a kernel should count 4 times toward medals. Medal thresholds could be raised a little and Kagglers with more contribution history would have more sway in the promotion of content. Basically, a \"karma\" system.\n\nDisagree below ; )",
    "303405": "Good idea, I believe it make sense.",
    "303439": "&gt; With this metric most blending kernels would have negative value...\n\n@CPMP, Bad analogy usually gives nonsense results.  \n\nIf we are looking for an analogy, the number of votes is not an analogy of the price of the product, but the number of the buyers of the product in the shop \"All for $1\". Very important thing is, that  many buyers often do not buy new things if they have bought something similar already,  even if the new product is better than the original and the price is only symbolic.",
    "303465": "Grzegorz,  my analogy is not about the price of products but the sales of a company.  \n\nIsn't the *number of number of the buyers of the product in the shop \"All for $1\"* exactly the same as the amount of sales in that shop expressed in dollar?  \n\nSeems you just expressed my analogy with your words.  Therefore it must not be that bad.",
    "303500": "What are these kernAls everybody speak about? ;)",
    "303533": "one could take an example from sites like stackoverflow where votes are not counted unless you are a user who earned your reputation already.",
    "303535": "CPMP, The amount of sales has nothing to do with VAT if the price before and after reselling is the same. And the sales of the companies producing similar products  have nothing to do with the products quality.\n\nBut it doesn't matter. I understand your idea of the added value of new (inspired, forked or blended) kernels and generally agree with it.  I  just think,  if the analogies are not good enough, the formulas based on them may be more harmful than useful.",
    "303548": "1\n\nFor discussion here, votes do not count towards medals unless you're above a given level.  I wonder if it is the same for Kernels.",
    "303552": "CPMP - it does seem to be the case for kernels as well, yes.",
    "303558": "Grzegorz, I'm not sure why you are arguing with.   Glad you agree with the general idea.  I may not have been clear enough, because your idea and mine are exactly the same.  Let me try again.\n\nVAT is the tax on your sales, minus the tax paid by your suppliers.  When there is a single VAT rate then this is equivalent to: VAT is a tax you pay on the difference between your sales, and what you buy from your suppliers, i.e. a VAT is a tax on your gross margin.\n\nIt is consistent with your example: if the price before and after reselling is the same (and if you don't buy anything else to sustain your business), then your VAT is 0, as well as your added value.  \n\nI maintain that the number of votes is similar to sales volume (or number of buyers).  When you use that analogy, then your added value is the number of votes you get (your sales), minus the number of votes of the kernels you used (your suppliers sales).\n\nLet me now if this is still not clear enough.",
    "303600": "Talking about discussion votes, some comments are not bronze in spite of 2 (netto) votes, some of them are bronze in spite of -3 (netto) votes, so there is a dark energy at Kaggle of the strength equal at least 6 votes (peak to peak). \nWe can suppose some notoric haters' down votes are neglected by admins, maybe admins' and masters' votes weight more, etc. , but if there is no transparency in that case, I wouldn't expect any solution clearly separating and evaluating original and blending kernels.",
    "303619": "shivraj: I agree with you here. Blends do not bother me, but it is unfair that they are ranked on the same scale as EDA or generally insightful kernels, and in many cases outperform them in the ranking. I am hoping for a way of weighting their votes, or grading them separately.",
    "303630": "CPMP, Let's try at the example:\n\nThere is a great LightGBM kernel which obtained 200 votes. I have blended/forked it and obtained 10 votes. Why only 10? Maybe because in the meantime Kagglers have found another ML method much better for the problem, maybe because it is simply the last day of the competition, maybe because of another 99 reasons. Are you 100% sure, my kernel has more or less but negative added value?\n\nIf the example above is still not clear, change LightGBM to CNN, me to you, and try again.",
    "303631": "I feel for you, being in a similar place.  \n\nI wish down votes would cost points: if you don't have discussion points then you cannot down vote, and each time you down vote you lose points.  I'd put one down vote cost to be a pretty large number of vote points, say 10.",
    "303642": "Grzegorz I don't know your kernel hence cannot say if it adds value or not.   I don't know enough to be able to say anything about the quality of your work.  I have no reason to think you are not adding value.\n\nBut in general, a kernel that is a weighted average of other kernels output without an explanation of how the weights were selected has negative value.  Weight could have been selected by LB probing for what I know, in which case people reusing it blindly will drop in private LB.   To be specific, in Porto Seguro competition, a blend kernel was very high in public LB, and all those who selected it for final submission dropped by 500 if not more.  That kernel had definitely a negative added value.  No doubt.\n\nSaid differently, I am not sure votes reflect the value added by a kernel. Maybe your kernel with 10 votes has very interesting insight that is being overlooked by the community.  And maybe the lgb kernel with lots of votes has no insight compared to other kernels it was derived from.  Without specifics it is hard to say.",
    "303654": "I have been an admin for 7 years, and I didn't install voting plugin at my forum. IMHO, discussion with arguments is much better. Here, the admins have to fight against haters. @CPMP, Your proposal makes sense. The haters would have less ammo.",
    "304005": "Very true. I do agree with people's thought of having weighted votes based on one's laurel. \n\nI do see that blends have its primary contribution to the bronze position(obviously lot of sensible blending happens at top positions as well) and one way to overcome them is to up your game and get yourself towards Silver or Gold, but not everybody (fights/can fight) for Gold so it would bother a lot of honest and aspiring learner of Data science who feels Kaggle is a great platform to learn and benchmark their skills.\n\nFor now, Most of the Blends signify how lucky was your weights selection.",
    "307982": "OK, how about [linear stacking][1] (which is the same thing as blending, really, but more systematic, so maybe you can at least learn the method)?\n\n [1]: https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704",
    "308110": "However you can't allow negative votes because then you will get overrun with robot profiles created to downvote rivals, esp. when there's a kernel competition. Perhaps only allow existing bona-fide members with &gt; threshold longevity, no. of upvoted posts and submissions to have a downvote privilege.\n\n&gt; **quantumgeek wrote**\n&gt; \n&gt; &gt; This is definitely a problem. One solution could be to have negative votes in kernels as well. The possibility of getting negative votes might lead to some self-moderation from the posters.",
    "323817": "Personnally, I agree that emphasizing too much blending is an issue as this shows only the last part of the computation and put under the carpet the hard pre blending work. But it is still usefull to be aware that blending can achiever higher score and is in a sense a way to do voting classifyers or any on top ensemble classifyer.  So in short for newbees like me they are usefull but they are not the most important ones."
  },
  "source": "meta"
}