{
  "id": 56631,
  "title": "Is kaggle competition all about blending ?",
  "url": "/competitions/avito-demand-prediction/discussion/56631",
  "author_name": "",
  "post_date": "2018-05-12T09:50:59.926021900Z",
  "votes": 8,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Blending some good kernels seems get good results, without any knowledge or information, even without running a real machine learning model...\nIs it good for data science and our community ?</p>",
  "messages": [
    {
      "id": "327734",
      "postDate": "05/12/2018 09:50:59",
      "content": "<p>Blending some good kernels seems get good results, without any knowledge or information, even without running a real machine learning model...\nIs it good for data science and our community ?</p>",
      "rawMarkdown": "Blending some good kernels seems get good results, without any knowledge or information, even without running a real machine learning model...\nIs it good for data science and our community ?",
      "votes": null
    },
    {
      "id": "327742",
      "postDate": "05/12/2018 10:28:52",
      "content": "<p>No it's not all about blending ..</p>\n\n<p>There are many things to do in this competition.  So you'll need to build very good base models before blending anything if you want reach top 50 AT THE END. ..(Unless someone want to ruin the competiton by publishing a highly scoring script in the last day ^^)</p>",
      "rawMarkdown": "No it's not all about blending ..\n\nThere are many things to do in this competition.  So you'll need to build very good base models before blending anything if you want reach top 50 AT THE END. ..(Unless someone want to ruin the competiton by publishing a highly scoring script in the last day ^^)",
      "votes": null
    },
    {
      "id": "327761",
      "postDate": "05/12/2018 12:00:23",
      "content": "<p>thank you.</p>",
      "rawMarkdown": "thank you.",
      "votes": null
    },
    {
      "id": "327809",
      "postDate": "05/12/2018 14:45:15",
      "content": "<p>As pointed out by <a href=\"/serigne\">@serigne</a>, blending can only get you so far. Some competitions are more \"blendable\" than others, but you still need to have lots of and very <strong>diverse</strong> base models. In this competition, because it involves several different kinds of data, and all of them seem to be important for the total predictive model, it will be very likely that you'll see <strong>many</strong> very diverse models. Furthermore, the image data will not be very easy to work with in kernels, so there will probably be even more \"offline\" good models.</p>\n\n<p>Another thing to point out is that blending is just one form of ensemble methods that are useful in Kaggle competitions. Stacking is another that has been very successful in many competitions, and will likely play a role in this one as well. (I have not tried it yet, but will probably move onto it later.) There are many good expositions to all of these methods in various Kaggle posts and notebooks. </p>\n\n<p>Hope this helps, and good luck with the competition! </p>",
      "rawMarkdown": "As pointed out by @serigne, blending can only get you so far. Some competitions are more \"blendable\" than others, but you still need to have lots of and very **diverse** base models. In this competition, because it involves several different kinds of data, and all of them seem to be important for the total predictive model, it will be very likely that you'll see **many** very diverse models. Furthermore, the image data will not be very easy to work with in kernels, so there will probably be even more \"offline\" good models.\n\nAnother thing to point out is that blending is just one form of ensemble methods that are useful in Kaggle competitions. Stacking is another that has been very successful in many competitions, and will likely play a role in this one as well. (I have not tried it yet, but will probably move onto it later.) There are many good expositions to all of these methods in various Kaggle posts and notebooks. \n\nHope this helps, and good luck with the competition!",
      "votes": null
    },
    {
      "id": "327830",
      "postDate": "05/12/2018 15:23:41",
      "content": "<p>Thank you !</p>",
      "rawMarkdown": "Thank you !",
      "votes": null
    },
    {
      "id": "327838",
      "postDate": "05/12/2018 15:56:37",
      "content": "<p>Look, as the other friends pointed Kaggle has changed the last years with the introduction of kernels, and this is not bad in principle. You can see that there will always be a set of very good kernels in every competition, so what you can do is to:</p>\n\n<ul>\n<li>Improve the existing kernels</li>\n<li>Create you own strong base models</li>\n<li>Blend with the existing public strong models (averaging, stacking, etc)</li>\n<li>Always check the forum and kernels :) </li>\n</ul>\n\n<p>... BUT, you need to be very careful with the public scripts, some of them contain target leaks, not very solid local validation techniques, etc.</p>",
      "rawMarkdown": "Look, as the other friends pointed Kaggle has changed the last years with the introduction of kernels, and this is not bad in principle. You can see that there will always be a set of very good kernels in every competition, so what you can do is to:\n\n - Improve the existing kernels\n - Create you own strong base models\n - Blend with the existing public strong models (averaging, stacking, etc)\n - Always check the forum and kernels :) \n\n... BUT, you need to be very careful with the public scripts, some of them contain target leaks, not very solid local validation techniques, etc.",
      "votes": null
    },
    {
      "id": "327839",
      "postDate": "05/12/2018 16:01:27",
      "content": "<p>I've generally found that if you want to win a gold medal, your single models need to individually win silver medals and if you want to win a silver medal, your single models need to individually win bronze medals.</p>",
      "rawMarkdown": "I've generally found that if you want to win a gold medal, your single models need to individually win silver medals and if you want to win a silver medal, your single models need to individually win bronze medals.",
      "votes": null
    },
    {
      "id": "327842",
      "postDate": "05/12/2018 16:05:16",
      "content": "<p>That’s a good rule of thumb.</p>",
      "rawMarkdown": "That’s a good rule of thumb.",
      "votes": null
    },
    {
      "id": "327843",
      "postDate": "05/12/2018 16:08:26",
      "content": "<p>All kernels are wrong, but some are useful. ;)</p>",
      "rawMarkdown": "All kernels are wrong, but some are useful. ;)",
      "votes": null
    },
    {
      "id": "327848",
      "postDate": "05/12/2018 16:47:24",
      "content": "<p>So seeing that you are right now in the gold territory, how good are your single models?</p>",
      "rawMarkdown": "So seeing that you are right now in the gold territory, how good are your single models?",
      "votes": null
    },
    {
      "id": "327856",
      "postDate": "05/12/2018 17:24:58",
      "content": "<p>I don't just lay out rules and then not follow them. ;)</p>",
      "rawMarkdown": "I don't just lay out rules and then not follow them. ;)",
      "votes": null
    },
    {
      "id": "327868",
      "postDate": "05/12/2018 18:24:33",
      "content": "<p>That's not true. Competitions are mostly about producing good and diverse models. Blending and stacking are usually the easiest part of the job once you have them.</p>",
      "rawMarkdown": "That's not true. Competitions are mostly about producing good and diverse models. Blending and stacking are usually the easiest part of the job once you have them.",
      "votes": null
    },
    {
      "id": "327945",
      "postDate": "05/13/2018 01:46:43",
      "content": "<p>Blend Winner is a bad  trend. Kaggle should try to stop it.</p>",
      "rawMarkdown": "Blend Winner is a bad  trend. Kaggle should try to stop it.",
      "votes": null
    },
    {
      "id": "327971",
      "postDate": "05/13/2018 03:32:42",
      "content": "<p>It sounds good rule to win a medal. Thank you!</p>",
      "rawMarkdown": "It sounds good rule to win a medal. Thank you!",
      "votes": null
    },
    {
      "id": "328041",
      "postDate": "05/13/2018 08:08:13",
      "content": "<p>For me, competitions are about two things:\n - Learning\n - Winning</p>\n\n<p>Neither of those two can be fulfilled by just blending :)</p>\n\n<p>Its like, you go to a restaurant to enjoy your favorite cuisine. Now, blending is more like smell of the cuisine. But, smell will never satisfy you. You need to savour your food to be truly satisfied. Savoring food is more like making strong base models :)</p>\n\n<p>I guess my taste in analogies is not that good :D </p>",
      "rawMarkdown": "For me, competitions are about two things:\n - Learning\n - Winning\n\nNeither of those two can be fulfilled by just blending :)\n\nIts like, you go to a restaurant to enjoy your favorite cuisine. Now, blending is more like smell of the cuisine. But, smell will never satisfy you. You need to savour your food to be truly satisfied. Savoring food is more like making strong base models :)\n\nI guess my taste in analogies is not that good :D",
      "votes": null
    },
    {
      "id": "328083",
      "postDate": "05/13/2018 09:55:56",
      "content": "<p>thanks you for your answer. sounds interesting!</p>",
      "rawMarkdown": "thanks you for your answer. sounds interesting!",
      "votes": null
    },
    {
      "id": "328397",
      "postDate": "05/14/2018 07:45:07",
      "content": "<p>@Jin Ma, the answer is NO. Just like anything else in life, you will find some people looking for shortcuts and quick gratification. Working on producing strong diverse models, creating a good local CV and trusting it is the way to go. Sure you may find a blend of blends kernel winning a silver medal in a contest like the \"Talking data AD tracking challenge\" but do not despair, that kernel could have easily dropped over 1000 points in the private LB. I have seen both happen in my short life on Kaggle.</p>\n\n<p>My advise, shut out the noise, learn from the kernels and discussions. Apply your knowledge to build the best models you can and let the chips fall where they may. I personally focus on the learning, then believe the medals and even a win will come some day :-)</p>",
      "rawMarkdown": "Jin Ma, the answer is NO. Just like anything else in life, you will find some people looking for shortcuts and quick gratification. Working on producing strong diverse models, creating a good local CV and trusting it is the way to go. Sure you may find a blend of blends kernel winning a silver medal in a contest like the \"Talking data AD tracking challenge\" but do not despair, that kernel could have easily dropped over 1000 points in the private LB. I have seen both happen in my short life on Kaggle.\n\nMy advise, shut out the noise, learn from the kernels and discussions. Apply your knowledge to build the best models you can and let the chips fall where they may. I personally focus on the learning, then believe the medals and even a win will come some day :-)",
      "votes": null
    },
    {
      "id": "328452",
      "postDate": "05/14/2018 11:30:07",
      "content": "<p>It hurts me again when I saw  \"Talking data AD tracking challenge\" .</p>",
      "rawMarkdown": "It hurts me again when I saw  \"Talking data AD tracking challenge\" .",
      "votes": null
    },
    {
      "id": "328527",
      "postDate": "05/14/2018 14:41:55",
      "content": "<blockquote>\n  <p>It hurts me again when I saw \"Talking data AD tracking challenge\" .</p>\n</blockquote>\n\n<p>I know that feel bro. <br>\n<img src=\"http://i0.kym-cdn.com/entries/icons/mobile/000/005/393/IKIFEEL.jpg\" alt=\"bro\"></p>",
      "rawMarkdown": "&gt; It hurts me again when I saw \"Talking data AD tracking challenge\" .\n\n\nI know that feel bro.  \n![bro][1]\n\n\n  [1]: http://i0.kym-cdn.com/entries/icons/mobile/000/005/393/IKIFEEL.jpg",
      "votes": null
    },
    {
      "id": "328547",
      "postDate": "05/14/2018 15:40:45",
      "content": "<p>Hi Dimitris,\nWhat do you mean by target leak  ? Does it mean using target variable in some form as a predictor variable in my model ?</p>",
      "rawMarkdown": "Hi Dimitris,\nWhat do you mean by target leak  ? Does it mean using target variable in some form as a predictor variable in my model ?",
      "votes": null
    },
    {
      "id": "328556",
      "postDate": "05/14/2018 15:53:02",
      "content": "<p>There have been a few abuses of the kernels and discussions over the past year or so, but most of them have nothing to do with kernels as such or with blending. They are the byproduct of Kaggle being an open online platform, so it is always possible that someone may release a very good solution on the last day or in the last week, either in kernels or as a link to an outside file/code/repository. It is unfortunate when that happens, and it does tend to frustrate most of us. But even in the \"real world\" there are many instances of similar behavior: your team may be working on a really cool new product, only for your rival to release an open source version of your product a week before yours was scheduled to launch. For better or for worse, those things are part of life, and are valuable life lessons in their own right. </p>\n\n<p>As for ensembling, there are many, many reasons why I like it a <strong>lot</strong>:</p>\n\n<ol>\n<li><p>It lets me explore many different algorithms/solutions/models. Instead of focusing on a single best solution, and stressing over \"dead ends,\" I get to try many things, and feel comfortable that even the suboptimal solutions can play a role in my final model.</p></li>\n<li><p>It makes me appreciate different approaches to solving problems. Both here on Kaggle and in the \"real world\" I have often been part of teams that comprise people from very different backgrounds. We'd have Physicists, Computer Scientists, and Statisticians. We'd have people who emphasize feature engineering, and those who are happy to focus on hyperparameter tuning. I'd work with teams with members living on three different continents. And universally, without exception, all those different approaches would positively impact the final solution. There is a lot of debate these days of value of diversity, but in this concrete, narrow sense, I've seen that it's a very, very valuable asset.</p></li>\n<li><p>It makes me appreciate different platforms/tech stacks. These days there is a lot of debate in ML and DS circles about the relative merit of R vs. Python. I am partial towards Python, but have seen many examples where models developed with both of those languages, <em>even when superficially very similar</em>, blend very well for a better result. The top solution in this competition right now is a blend of R and Python scripts.</p></li>\n<li><p>For models that combine many different forms of data, such as images, text and tabular, you pretty much <strong>have to</strong> use ensemble methods.</p></li>\n<li><p>As has been discussed in several posts, in this particular competition we are probably not trying to estimate some ground truth, but rather predictions of another model. This model, more likely than not, is a very sophisticated ensemble. In order to emulate it in some way, we need to think along the lines of developing strong ensembles of our own. </p></li>\n</ol>",
      "rawMarkdown": "There have been a few abuses of the kernels and discussions over the past year or so, but most of them have nothing to do with kernels as such or with blending. They are the byproduct of Kaggle being an open online platform, so it is always possible that someone may release a very good solution on the last day or in the last week, either in kernels or as a link to an outside file/code/repository. It is unfortunate when that happens, and it does tend to frustrate most of us. But even in the \"real world\" there are many instances of similar behavior: your team may be working on a really cool new product, only for your rival to release an open source version of your product a week before yours was scheduled to launch. For better or for worse, those things are part of life, and are valuable life lessons in their own right. \n\nAs for ensembling, there are many, many reasons why I like it a **lot**:\n\n1. It lets me explore many different algorithms/solutions/models. Instead of focusing on a single best solution, and stressing over \"dead ends,\" I get to try many things, and feel comfortable that even the suboptimal solutions can play a role in my final model.\n\n2. It makes me appreciate different approaches to solving problems. Both here on Kaggle and in the \"real world\" I have often been part of teams that comprise people from very different backgrounds. We'd have Physicists, Computer Scientists, and Statisticians. We'd have people who emphasize feature engineering, and those who are happy to focus on hyperparameter tuning. I'd work with teams with members living on three different continents. And universally, without exception, all those different approaches would positively impact the final solution. There is a lot of debate these days of value of diversity, but in this concrete, narrow sense, I've seen that it's a very, very valuable asset.\n\n3. It makes me appreciate different platforms/tech stacks. These days there is a lot of debate in ML and DS circles about the relative merit of R vs. Python. I am partial towards Python, but have seen many examples where models developed with both of those languages, *even when superficially very similar*, blend very well for a better result. The top solution in this competition right now is a blend of R and Python scripts.\n\n4. For models that combine many different forms of data, such as images, text and tabular, you pretty much **have to** use ensemble methods.\n\n5. As has been discussed in several posts, in this particular competition we are probably not trying to estimate some ground truth, but rather predictions of another model. This model, more likely than not, is a very sophisticated ensemble. In order to emulate it in some way, we need to think along the lines of developing strong ensembles of our own.",
      "votes": null
    },
    {
      "id": "328615",
      "postDate": "05/14/2018 17:57:01",
      "content": "<p>It's like anything in life. There is the easy way, and hard way...</p>",
      "rawMarkdown": "It's like anything in life. There is the easy way, and hard way...",
      "votes": null
    },
    {
      "id": "328655",
      "postDate": "05/14/2018 20:42:51",
      "content": "<p>Ditto @YunFei Duan</p>\n\n<p>Thanks for the hug @Totoro. </p>",
      "rawMarkdown": "Ditto @YunFei Duan\n\nThanks for the hug @Totoro.",
      "votes": null
    },
    {
      "id": "330215",
      "postDate": "05/18/2018 09:40:54",
      "content": "<p>Lies, Damn Lies, and .. Kernels?</p>",
      "rawMarkdown": "Lies, Damn Lies, and .. Kernels?",
      "votes": null
    },
    {
      "id": "330234",
      "postDate": "05/18/2018 10:40:06",
      "content": "<p>Wait for the private leaderboard results before concluding that blending kernels yields good results.</p>",
      "rawMarkdown": "Wait for the private leaderboard results before concluding that blending kernels yields good results.",
      "votes": null
    },
    {
      "id": "330336",
      "postDate": "05/18/2018 15:45:56",
      "content": "<p>Yes Firdaus, this is what target leak usually mean. If you are not careful in your feature selection and engineering it can happen.</p>",
      "rawMarkdown": "Yes Firdaus, this is what target leak usually mean. If you are not careful in your feature selection and engineering it can happen.",
      "votes": null
    },
    {
      "id": "330693",
      "postDate": "05/19/2018 13:11:16",
      "content": "<p>Hi Firdaus, sorry for the late response, yes this is what I meant and it is not always bad.\nYou can read an interesting article by Marios Michaildis (aka Kazanova) here: <a href=\"https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/\">https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/</a></p>",
      "rawMarkdown": "Hi Firdaus, sorry for the late response, yes this is what I meant and it is not always bad.\nYou can read an interesting article by Marios Michaildis (aka Kazanova) here: https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/",
      "votes": null
    },
    {
      "id": "335372",
      "postDate": "05/29/2018 16:50:15",
      "content": "<p>Blending is a critical tool in the arsenal.</p>\n\n<p>From my experience, being a good blender of standard kernels can get you to the top 5%-10% depending on the quality of the solutions provided in the kernels and your ability to do tweaks and tune hyperparameters.  To get beyond the top 5% really takes insight.  </p>\n\n<p>My favorite part of these competitions is the insights shared at the end!  :-)  For example, my current position (while it lasts) is coming from  the insights provided in the Jigsaw Toxic Comment Competition. </p>",
      "rawMarkdown": "Blending is a critical tool in the arsenal.\n\nFrom my experience, being a good blender of standard kernels can get you to the top 5%-10% depending on the quality of the solutions provided in the kernels and your ability to do tweaks and tune hyperparameters.  To get beyond the top 5% really takes insight.  \n\nMy favorite part of these competitions is the insights shared at the end!  :-)  For example, my current position (while it lasts) is coming from  the insights provided in the Jigsaw Toxic Comment Competition.",
      "votes": null
    },
    {
      "id": "335886",
      "postDate": "05/30/2018 15:36:01",
      "content": "<p>Thanks for your reply!!!</p>",
      "rawMarkdown": "Thanks for your reply!!!",
      "votes": null
    },
    {
      "id": "335887",
      "postDate": "05/30/2018 15:40:05",
      "content": "<p>Thanks for your reply!!!</p>",
      "rawMarkdown": "Thanks for your reply!!!",
      "votes": null
    },
    {
      "id": "335888",
      "postDate": "05/30/2018 15:41:22",
      "content": "<p>Thanks for your reply!!!</p>",
      "rawMarkdown": "Thanks for your reply!!!",
      "votes": null
    },
    {
      "id": "335889",
      "postDate": "05/30/2018 15:41:45",
      "content": "<p>Thanks for your reply!!!</p>",
      "rawMarkdown": "Thanks for your reply!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 327742,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/12/2018 10:28:52",
      "content": "<p>No it's not all about blending ..</p>\n\n<p>There are many things to do in this competition.  So you'll need to build very good base models before blending anything if you want reach top 50 AT THE END. ..(Unless someone want to ruin the competiton by publishing a highly scoring script in the last day ^^)</p>",
      "votes": null,
      "replies": [
        {
          "id": 327761,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/12/2018 12:00:23",
          "content": "<p>thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327809,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "05/12/2018 14:45:15",
      "content": "<p>As pointed out by <a href=\"/serigne\">@serigne</a>, blending can only get you so far. Some competitions are more \"blendable\" than others, but you still need to have lots of and very <strong>diverse</strong> base models. In this competition, because it involves several different kinds of data, and all of them seem to be important for the total predictive model, it will be very likely that you'll see <strong>many</strong> very diverse models. Furthermore, the image data will not be very easy to work with in kernels, so there will probably be even more \"offline\" good models.</p>\n\n<p>Another thing to point out is that blending is just one form of ensemble methods that are useful in Kaggle competitions. Stacking is another that has been very successful in many competitions, and will likely play a role in this one as well. (I have not tried it yet, but will probably move onto it later.) There are many good expositions to all of these methods in various Kaggle posts and notebooks. </p>\n\n<p>Hope this helps, and good luck with the competition! </p>",
      "votes": null,
      "replies": [
        {
          "id": 327830,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/12/2018 15:23:41",
          "content": "<p>Thank you !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327838,
      "author_name": "dimitrislev",
      "author_url": "",
      "post_date": "05/12/2018 15:56:37",
      "content": "<p>Look, as the other friends pointed Kaggle has changed the last years with the introduction of kernels, and this is not bad in principle. You can see that there will always be a set of very good kernels in every competition, so what you can do is to:</p>\n\n<ul>\n<li>Improve the existing kernels</li>\n<li>Create you own strong base models</li>\n<li>Blend with the existing public strong models (averaging, stacking, etc)</li>\n<li>Always check the forum and kernels :) </li>\n</ul>\n\n<p>... BUT, you need to be very careful with the public scripts, some of them contain target leaks, not very solid local validation techniques, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 327843,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/12/2018 16:08:26",
          "content": "<p>All kernels are wrong, but some are useful. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328547,
          "author_name": "",
          "author_url": "",
          "post_date": "05/14/2018 15:40:45",
          "content": "<p>Hi Dimitris,\nWhat do you mean by target leak  ? Does it mean using target variable in some form as a predictor variable in my model ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330215,
          "author_name": "nicapotato",
          "author_url": "",
          "post_date": "05/18/2018 09:40:54",
          "content": "<p>Lies, Damn Lies, and .. Kernels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330336,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "05/18/2018 15:45:56",
          "content": "<p>Yes Firdaus, this is what target leak usually mean. If you are not careful in your feature selection and engineering it can happen.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330693,
          "author_name": "dimitrislev",
          "author_url": "",
          "post_date": "05/19/2018 13:11:16",
          "content": "<p>Hi Firdaus, sorry for the late response, yes this is what I meant and it is not always bad.\nYou can read an interesting article by Marios Michaildis (aka Kazanova) here: <a href=\"https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/\">https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335887,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/30/2018 15:40:05",
          "content": "<p>Thanks for your reply!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327839,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "05/12/2018 16:01:27",
      "content": "<p>I've generally found that if you want to win a gold medal, your single models need to individually win silver medals and if you want to win a silver medal, your single models need to individually win bronze medals.</p>",
      "votes": null,
      "replies": [
        {
          "id": 327842,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/12/2018 16:05:16",
          "content": "<p>That’s a good rule of thumb.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327848,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "05/12/2018 16:47:24",
          "content": "<p>So seeing that you are right now in the gold territory, how good are your single models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327856,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/12/2018 17:24:58",
          "content": "<p>I don't just lay out rules and then not follow them. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327971,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/13/2018 03:32:42",
          "content": "<p>It sounds good rule to win a medal. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327868,
      "author_name": "thousandvoices",
      "author_url": "",
      "post_date": "05/12/2018 18:24:33",
      "content": "<p>That's not true. Competitions are mostly about producing good and diverse models. Blending and stacking are usually the easiest part of the job once you have them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 335886,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/30/2018 15:36:01",
          "content": "<p>Thanks for your reply!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327945,
      "author_name": "kownse",
      "author_url": "",
      "post_date": "05/13/2018 01:46:43",
      "content": "<p>Blend Winner is a bad  trend. Kaggle should try to stop it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328041,
      "author_name": "tezdhar",
      "author_url": "",
      "post_date": "05/13/2018 08:08:13",
      "content": "<p>For me, competitions are about two things:\n - Learning\n - Winning</p>\n\n<p>Neither of those two can be fulfilled by just blending :)</p>\n\n<p>Its like, you go to a restaurant to enjoy your favorite cuisine. Now, blending is more like smell of the cuisine. But, smell will never satisfy you. You need to savour your food to be truly satisfied. Savoring food is more like making strong base models :)</p>\n\n<p>I guess my taste in analogies is not that good :D </p>",
      "votes": null,
      "replies": [
        {
          "id": 328083,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/13/2018 09:55:56",
          "content": "<p>thanks you for your answer. sounds interesting!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328397,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/14/2018 07:45:07",
      "content": "<p>@Jin Ma, the answer is NO. Just like anything else in life, you will find some people looking for shortcuts and quick gratification. Working on producing strong diverse models, creating a good local CV and trusting it is the way to go. Sure you may find a blend of blends kernel winning a silver medal in a contest like the \"Talking data AD tracking challenge\" but do not despair, that kernel could have easily dropped over 1000 points in the private LB. I have seen both happen in my short life on Kaggle.</p>\n\n<p>My advise, shut out the noise, learn from the kernels and discussions. Apply your knowledge to build the best models you can and let the chips fall where they may. I personally focus on the learning, then believe the medals and even a win will come some day :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 328452,
          "author_name": "kownse",
          "author_url": "",
          "post_date": "05/14/2018 11:30:07",
          "content": "<p>It hurts me again when I saw  \"Talking data AD tracking challenge\" .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328527,
          "author_name": "ngxbac",
          "author_url": "",
          "post_date": "05/14/2018 14:41:55",
          "content": "<blockquote>\n  <p>It hurts me again when I saw \"Talking data AD tracking challenge\" .</p>\n</blockquote>\n\n<p>I know that feel bro. <br>\n<img src=\"http://i0.kym-cdn.com/entries/icons/mobile/000/005/393/IKIFEEL.jpg\" alt=\"bro\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328655,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "05/14/2018 20:42:51",
          "content": "<p>Ditto @YunFei Duan</p>\n\n<p>Thanks for the hug @Totoro. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328556,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "05/14/2018 15:53:02",
      "content": "<p>There have been a few abuses of the kernels and discussions over the past year or so, but most of them have nothing to do with kernels as such or with blending. They are the byproduct of Kaggle being an open online platform, so it is always possible that someone may release a very good solution on the last day or in the last week, either in kernels or as a link to an outside file/code/repository. It is unfortunate when that happens, and it does tend to frustrate most of us. But even in the \"real world\" there are many instances of similar behavior: your team may be working on a really cool new product, only for your rival to release an open source version of your product a week before yours was scheduled to launch. For better or for worse, those things are part of life, and are valuable life lessons in their own right. </p>\n\n<p>As for ensembling, there are many, many reasons why I like it a <strong>lot</strong>:</p>\n\n<ol>\n<li><p>It lets me explore many different algorithms/solutions/models. Instead of focusing on a single best solution, and stressing over \"dead ends,\" I get to try many things, and feel comfortable that even the suboptimal solutions can play a role in my final model.</p></li>\n<li><p>It makes me appreciate different approaches to solving problems. Both here on Kaggle and in the \"real world\" I have often been part of teams that comprise people from very different backgrounds. We'd have Physicists, Computer Scientists, and Statisticians. We'd have people who emphasize feature engineering, and those who are happy to focus on hyperparameter tuning. I'd work with teams with members living on three different continents. And universally, without exception, all those different approaches would positively impact the final solution. There is a lot of debate these days of value of diversity, but in this concrete, narrow sense, I've seen that it's a very, very valuable asset.</p></li>\n<li><p>It makes me appreciate different platforms/tech stacks. These days there is a lot of debate in ML and DS circles about the relative merit of R vs. Python. I am partial towards Python, but have seen many examples where models developed with both of those languages, <em>even when superficially very similar</em>, blend very well for a better result. The top solution in this competition right now is a blend of R and Python scripts.</p></li>\n<li><p>For models that combine many different forms of data, such as images, text and tabular, you pretty much <strong>have to</strong> use ensemble methods.</p></li>\n<li><p>As has been discussed in several posts, in this particular competition we are probably not trying to estimate some ground truth, but rather predictions of another model. This model, more likely than not, is a very sophisticated ensemble. In order to emulate it in some way, we need to think along the lines of developing strong ensembles of our own. </p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 335889,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/30/2018 15:41:45",
          "content": "<p>Thanks for your reply!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328615,
      "author_name": "lingtop",
      "author_url": "",
      "post_date": "05/14/2018 17:57:01",
      "content": "<p>It's like anything in life. There is the easy way, and hard way...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330234,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/18/2018 10:40:06",
      "content": "<p>Wait for the private leaderboard results before concluding that blending kernels yields good results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 335372,
      "author_name": "larryfreeman",
      "author_url": "",
      "post_date": "05/29/2018 16:50:15",
      "content": "<p>Blending is a critical tool in the arsenal.</p>\n\n<p>From my experience, being a good blender of standard kernels can get you to the top 5%-10% depending on the quality of the solutions provided in the kernels and your ability to do tweaks and tune hyperparameters.  To get beyond the top 5% really takes insight.  </p>\n\n<p>My favorite part of these competitions is the insights shared at the end!  :-)  For example, my current position (while it lasts) is coming from  the insights provided in the Jigsaw Toxic Comment Competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 335888,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "05/30/2018 15:41:22",
          "content": "<p>Thanks for your reply!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "327734": "Blending some good kernels seems get good results, without any knowledge or information, even without running a real machine learning model...\nIs it good for data science and our community ?",
    "327742": "No it's not all about blending ..\n\nThere are many things to do in this competition.  So you'll need to build very good base models before blending anything if you want reach top 50 AT THE END. ..(Unless someone want to ruin the competiton by publishing a highly scoring script in the last day ^^)",
    "327761": "thank you.",
    "327809": "As pointed out by @serigne, blending can only get you so far. Some competitions are more \"blendable\" than others, but you still need to have lots of and very **diverse** base models. In this competition, because it involves several different kinds of data, and all of them seem to be important for the total predictive model, it will be very likely that you'll see **many** very diverse models. Furthermore, the image data will not be very easy to work with in kernels, so there will probably be even more \"offline\" good models.\n\nAnother thing to point out is that blending is just one form of ensemble methods that are useful in Kaggle competitions. Stacking is another that has been very successful in many competitions, and will likely play a role in this one as well. (I have not tried it yet, but will probably move onto it later.) There are many good expositions to all of these methods in various Kaggle posts and notebooks. \n\nHope this helps, and good luck with the competition!",
    "327830": "Thank you !",
    "327838": "Look, as the other friends pointed Kaggle has changed the last years with the introduction of kernels, and this is not bad in principle. You can see that there will always be a set of very good kernels in every competition, so what you can do is to:\n\n - Improve the existing kernels\n - Create you own strong base models\n - Blend with the existing public strong models (averaging, stacking, etc)\n - Always check the forum and kernels :) \n\n... BUT, you need to be very careful with the public scripts, some of them contain target leaks, not very solid local validation techniques, etc.",
    "327839": "I've generally found that if you want to win a gold medal, your single models need to individually win silver medals and if you want to win a silver medal, your single models need to individually win bronze medals.",
    "327842": "That’s a good rule of thumb.",
    "327843": "All kernels are wrong, but some are useful. ;)",
    "327848": "So seeing that you are right now in the gold territory, how good are your single models?",
    "327856": "I don't just lay out rules and then not follow them. ;)",
    "327868": "That's not true. Competitions are mostly about producing good and diverse models. Blending and stacking are usually the easiest part of the job once you have them.",
    "327945": "Blend Winner is a bad  trend. Kaggle should try to stop it.",
    "327971": "It sounds good rule to win a medal. Thank you!",
    "328041": "For me, competitions are about two things:\n - Learning\n - Winning\n\nNeither of those two can be fulfilled by just blending :)\n\nIts like, you go to a restaurant to enjoy your favorite cuisine. Now, blending is more like smell of the cuisine. But, smell will never satisfy you. You need to savour your food to be truly satisfied. Savoring food is more like making strong base models :)\n\nI guess my taste in analogies is not that good :D",
    "328083": "thanks you for your answer. sounds interesting!",
    "328397": "Jin Ma, the answer is NO. Just like anything else in life, you will find some people looking for shortcuts and quick gratification. Working on producing strong diverse models, creating a good local CV and trusting it is the way to go. Sure you may find a blend of blends kernel winning a silver medal in a contest like the \"Talking data AD tracking challenge\" but do not despair, that kernel could have easily dropped over 1000 points in the private LB. I have seen both happen in my short life on Kaggle.\n\nMy advise, shut out the noise, learn from the kernels and discussions. Apply your knowledge to build the best models you can and let the chips fall where they may. I personally focus on the learning, then believe the medals and even a win will come some day :-)",
    "328452": "It hurts me again when I saw  \"Talking data AD tracking challenge\" .",
    "328527": "&gt; It hurts me again when I saw \"Talking data AD tracking challenge\" .\n\n\nI know that feel bro.  \n![bro][1]\n\n\n  [1]: http://i0.kym-cdn.com/entries/icons/mobile/000/005/393/IKIFEEL.jpg",
    "328547": "Hi Dimitris,\nWhat do you mean by target leak  ? Does it mean using target variable in some form as a predictor variable in my model ?",
    "328556": "There have been a few abuses of the kernels and discussions over the past year or so, but most of them have nothing to do with kernels as such or with blending. They are the byproduct of Kaggle being an open online platform, so it is always possible that someone may release a very good solution on the last day or in the last week, either in kernels or as a link to an outside file/code/repository. It is unfortunate when that happens, and it does tend to frustrate most of us. But even in the \"real world\" there are many instances of similar behavior: your team may be working on a really cool new product, only for your rival to release an open source version of your product a week before yours was scheduled to launch. For better or for worse, those things are part of life, and are valuable life lessons in their own right. \n\nAs for ensembling, there are many, many reasons why I like it a **lot**:\n\n1. It lets me explore many different algorithms/solutions/models. Instead of focusing on a single best solution, and stressing over \"dead ends,\" I get to try many things, and feel comfortable that even the suboptimal solutions can play a role in my final model.\n\n2. It makes me appreciate different approaches to solving problems. Both here on Kaggle and in the \"real world\" I have often been part of teams that comprise people from very different backgrounds. We'd have Physicists, Computer Scientists, and Statisticians. We'd have people who emphasize feature engineering, and those who are happy to focus on hyperparameter tuning. I'd work with teams with members living on three different continents. And universally, without exception, all those different approaches would positively impact the final solution. There is a lot of debate these days of value of diversity, but in this concrete, narrow sense, I've seen that it's a very, very valuable asset.\n\n3. It makes me appreciate different platforms/tech stacks. These days there is a lot of debate in ML and DS circles about the relative merit of R vs. Python. I am partial towards Python, but have seen many examples where models developed with both of those languages, *even when superficially very similar*, blend very well for a better result. The top solution in this competition right now is a blend of R and Python scripts.\n\n4. For models that combine many different forms of data, such as images, text and tabular, you pretty much **have to** use ensemble methods.\n\n5. As has been discussed in several posts, in this particular competition we are probably not trying to estimate some ground truth, but rather predictions of another model. This model, more likely than not, is a very sophisticated ensemble. In order to emulate it in some way, we need to think along the lines of developing strong ensembles of our own.",
    "328615": "It's like anything in life. There is the easy way, and hard way...",
    "328655": "Ditto @YunFei Duan\n\nThanks for the hug @Totoro.",
    "330215": "Lies, Damn Lies, and .. Kernels?",
    "330234": "Wait for the private leaderboard results before concluding that blending kernels yields good results.",
    "330336": "Yes Firdaus, this is what target leak usually mean. If you are not careful in your feature selection and engineering it can happen.",
    "330693": "Hi Firdaus, sorry for the late response, yes this is what I meant and it is not always bad.\nYou can read an interesting article by Marios Michaildis (aka Kazanova) here: https://blog.h2o.ai/2018/03/driverless-ai-prevents-overfitting-leakage/",
    "335372": "Blending is a critical tool in the arsenal.\n\nFrom my experience, being a good blender of standard kernels can get you to the top 5%-10% depending on the quality of the solutions provided in the kernels and your ability to do tweaks and tune hyperparameters.  To get beyond the top 5% really takes insight.  \n\nMy favorite part of these competitions is the insights shared at the end!  :-)  For example, my current position (while it lasts) is coming from  the insights provided in the Jigsaw Toxic Comment Competition.",
    "335886": "Thanks for your reply!!!",
    "335887": "Thanks for your reply!!!",
    "335888": "Thanks for your reply!!!",
    "335889": "Thanks for your reply!!!"
  },
  "source": "meta"
}