{
  "id": 53016,
  "title": "Look at 'Blending' from another perspective",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53016",
  "author_name": "",
  "post_date": "2018-03-26T09:43:20.725272600Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I know lots of kagglers think 'Blending' is not a good way to climb LB, and make no sense for modeling.\nBut in ML, we know ensemble is usually a good solution.So I think we should use it like the way below:\nset weights for different submissions ,merge the result and test it in the validation set , then get a optimum weights. And then train different models for relative submission you borrow, and summarize the predictions with the optimize weights as your final submission,but not just Blending Others' Submission directly.  </p>",
  "messages": [
    {
      "id": "303494",
      "postDate": "03/26/2018 09:43:20",
      "content": "<p>I know lots of kagglers think 'Blending' is not a good way to climb LB, and make no sense for modeling.\nBut in ML, we know ensemble is usually a good solution.So I think we should use it like the way below:\nset weights for different submissions ,merge the result and test it in the validation set , then get a optimum weights. And then train different models for relative submission you borrow, and summarize the predictions with the optimize weights as your final submission,but not just Blending Others' Submission directly.  </p>",
      "rawMarkdown": "I know lots of kagglers think 'Blending' is not a good way to climb LB, and make no sense for modeling.\nBut in ML, we know ensemble is usually a good solution.So I think we should use it like the way below:\nset weights for different submissions ,merge the result and test it in the validation set , then get a optimum weights. And then train different models for relative submission you borrow, and summarize the predictions with the optimize weights as your final submission,but not just Blending Others' Submission directly.",
      "votes": null
    },
    {
      "id": "303501",
      "postDate": "03/26/2018 10:14:13",
      "content": "<p>but this method always overfit LB</p>",
      "rawMarkdown": "but this method always overfit LB",
      "votes": null
    },
    {
      "id": "303644",
      "postDate": "03/26/2018 13:40:18",
      "content": "<p>I'm afraid what you describe is precisely how most of these blend are created.  And as said by cherryunix, this most often leads to overfiting the public test data.</p>",
      "rawMarkdown": "I'm afraid what you describe is precisely how most of these blend are created.  And as said by cherryunix, this most often leads to overfiting the public test data.",
      "votes": null
    },
    {
      "id": "303831",
      "postDate": "03/26/2018 17:43:09",
      "content": "<p>I think you've misunderstood the opinion raised about blending. I am sure that nobody denies that blends ( <code>with some logic</code> ) are bad. They are very important (at the end) but <strong>to make any kind of blend, you need to have base models</strong> . The reason why people are frustrated is because of high volume of blends and less base models. Most Kagglers, in my opinion are either students or data science enthusiasts who are here to improve their knowledge or skills in this field. And so far, kernels section was the best learning source. </p>\n\n<p>Imagine the situation where everyone starts blending right from the beginning of competition. That would be the future where anyone with no knowledge in machine learning will participate in competition and will start multiplying random numbers with sample submission to see which number is the luckiest number. </p>",
      "rawMarkdown": "I think you've misunderstood the opinion raised about blending. I am sure that nobody denies that blends ( `with some logic` ) are bad. They are very important (at the end) but **to make any kind of blend, you need to have base models** . The reason why people are frustrated is because of high volume of blends and less base models. Most Kagglers, in my opinion are either students or data science enthusiasts who are here to improve their knowledge or skills in this field. And so far, kernels section was the best learning source. \n\nImagine the situation where everyone starts blending right from the beginning of competition. That would be the future where anyone with no knowledge in machine learning will participate in competition and will start multiplying random numbers with sample submission to see which number is the luckiest number.",
      "votes": null
    },
    {
      "id": "303990",
      "postDate": "03/26/2018 21:21:28",
      "content": "<p>I agree with what's already been said - blend kernels tend toward magic numbers with no validation framework to produce good public LB scores. This voids most educational opportunity that could come from them on how to ensemble. An ensemble kernel in the style of this <a href=\"https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard\">excellent one</a> would be very valuable, but in my experience this is much rarer in active competitions than magic button kernels.</p>\n\n<p>Another point I'd add is that blending (and especially blending of this style) is typically less useful for learning how to build industry-quality ML models. Convoluted mega-ensembles are much harder to interpret and maintain than well-crafted single models, and the marginal predictive power gains don't always justify this cost. Many people who are here (at least partially) for educational purposes would benefit more from learning about good feature engineering and single model construction.</p>\n\n<p>In short, blend kernels tend to be less interesting, less rigorous, and less educational than single model kernels.</p>",
      "rawMarkdown": "I agree with what's already been said - blend kernels tend toward magic numbers with no validation framework to produce good public LB scores. This voids most educational opportunity that could come from them on how to ensemble. An ensemble kernel in the style of this [excellent one][1] would be very valuable, but in my experience this is much rarer in active competitions than magic button kernels.\n\nAnother point I'd add is that blending (and especially blending of this style) is typically less useful for learning how to build industry-quality ML models. Convoluted mega-ensembles are much harder to interpret and maintain than well-crafted single models, and the marginal predictive power gains don't always justify this cost. Many people who are here (at least partially) for educational purposes would benefit more from learning about good feature engineering and single model construction.\n\nIn short, blend kernels tend to be less interesting, less rigorous, and less educational than single model kernels.\n\n\n  [1]: https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard",
      "votes": null
    },
    {
      "id": "304084",
      "postDate": "03/27/2018 02:00:32",
      "content": "<p>maybe you are right, the magic weights could not be random choose</p>",
      "rawMarkdown": "maybe you are right, the magic weights could not be random choose",
      "votes": null
    },
    {
      "id": "304086",
      "postDate": "03/27/2018 02:06:22",
      "content": "<p>Thanks for your explanation. In deed, I was shocked when I found the 'Blending Kernels' at the first time, it's pretty genius but as you said,no knowledge.</p>",
      "rawMarkdown": "Thanks for your explanation. In deed, I was shocked when I found the 'Blending Kernels' at the first time, it's pretty genius but as you said,no knowledge.",
      "votes": null
    },
    {
      "id": "304095",
      "postDate": "03/27/2018 02:30:56",
      "content": "<p>Learn a lot from your subscription, thanks!</p>",
      "rawMarkdown": "Learn a lot from your subscription, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 303501,
      "author_name": "cherryunix",
      "author_url": "",
      "post_date": "03/26/2018 10:14:13",
      "content": "<p>but this method always overfit LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 303644,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/26/2018 13:40:18",
      "content": "<p>I'm afraid what you describe is precisely how most of these blend are created.  And as said by cherryunix, this most often leads to overfiting the public test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 304084,
          "author_name": "lccever",
          "author_url": "",
          "post_date": "03/27/2018 02:00:32",
          "content": "<p>maybe you are right, the magic weights could not be random choose</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303831,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "03/26/2018 17:43:09",
      "content": "<p>I think you've misunderstood the opinion raised about blending. I am sure that nobody denies that blends ( <code>with some logic</code> ) are bad. They are very important (at the end) but <strong>to make any kind of blend, you need to have base models</strong> . The reason why people are frustrated is because of high volume of blends and less base models. Most Kagglers, in my opinion are either students or data science enthusiasts who are here to improve their knowledge or skills in this field. And so far, kernels section was the best learning source. </p>\n\n<p>Imagine the situation where everyone starts blending right from the beginning of competition. That would be the future where anyone with no knowledge in machine learning will participate in competition and will start multiplying random numbers with sample submission to see which number is the luckiest number. </p>",
      "votes": null,
      "replies": [
        {
          "id": 304086,
          "author_name": "lccever",
          "author_url": "",
          "post_date": "03/27/2018 02:06:22",
          "content": "<p>Thanks for your explanation. In deed, I was shocked when I found the 'Blending Kernels' at the first time, it's pretty genius but as you said,no knowledge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303990,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "03/26/2018 21:21:28",
      "content": "<p>I agree with what's already been said - blend kernels tend toward magic numbers with no validation framework to produce good public LB scores. This voids most educational opportunity that could come from them on how to ensemble. An ensemble kernel in the style of this <a href=\"https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard\">excellent one</a> would be very valuable, but in my experience this is much rarer in active competitions than magic button kernels.</p>\n\n<p>Another point I'd add is that blending (and especially blending of this style) is typically less useful for learning how to build industry-quality ML models. Convoluted mega-ensembles are much harder to interpret and maintain than well-crafted single models, and the marginal predictive power gains don't always justify this cost. Many people who are here (at least partially) for educational purposes would benefit more from learning about good feature engineering and single model construction.</p>\n\n<p>In short, blend kernels tend to be less interesting, less rigorous, and less educational than single model kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 304095,
          "author_name": "lccever",
          "author_url": "",
          "post_date": "03/27/2018 02:30:56",
          "content": "<p>Learn a lot from your subscription, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "303494": "I know lots of kagglers think 'Blending' is not a good way to climb LB, and make no sense for modeling.\nBut in ML, we know ensemble is usually a good solution.So I think we should use it like the way below:\nset weights for different submissions ,merge the result and test it in the validation set , then get a optimum weights. And then train different models for relative submission you borrow, and summarize the predictions with the optimize weights as your final submission,but not just Blending Others' Submission directly.",
    "303501": "but this method always overfit LB",
    "303644": "I'm afraid what you describe is precisely how most of these blend are created.  And as said by cherryunix, this most often leads to overfiting the public test data.",
    "303831": "I think you've misunderstood the opinion raised about blending. I am sure that nobody denies that blends ( `with some logic` ) are bad. They are very important (at the end) but **to make any kind of blend, you need to have base models** . The reason why people are frustrated is because of high volume of blends and less base models. Most Kagglers, in my opinion are either students or data science enthusiasts who are here to improve their knowledge or skills in this field. And so far, kernels section was the best learning source. \n\nImagine the situation where everyone starts blending right from the beginning of competition. That would be the future where anyone with no knowledge in machine learning will participate in competition and will start multiplying random numbers with sample submission to see which number is the luckiest number.",
    "303990": "I agree with what's already been said - blend kernels tend toward magic numbers with no validation framework to produce good public LB scores. This voids most educational opportunity that could come from them on how to ensemble. An ensemble kernel in the style of this [excellent one][1] would be very valuable, but in my experience this is much rarer in active competitions than magic button kernels.\n\nAnother point I'd add is that blending (and especially blending of this style) is typically less useful for learning how to build industry-quality ML models. Convoluted mega-ensembles are much harder to interpret and maintain than well-crafted single models, and the marginal predictive power gains don't always justify this cost. Many people who are here (at least partially) for educational purposes would benefit more from learning about good feature engineering and single model construction.\n\nIn short, blend kernels tend to be less interesting, less rigorous, and less educational than single model kernels.\n\n\n  [1]: https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard",
    "304084": "maybe you are right, the magic weights could not be random choose",
    "304086": "Thanks for your explanation. In deed, I was shocked when I found the 'Blending Kernels' at the first time, it's pretty genius but as you said,no knowledge.",
    "304095": "Learn a lot from your subscription, thanks!"
  },
  "source": "meta"
}