{
  "id": 53141,
  "title": "In Defense of Public Blending: Seek Alpha, Not Total Return",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53141",
  "author_name": "Andy Harless",
  "post_date": "2018-03-27T15:19:06.936000",
  "votes": 16,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Unless blending is done in public (and I mean, fully public, with the inputs to the blend being public as well), the public has no way of knowing which single-model kernels are the most promising.  Just as the LB scores of blend kernels are not directly indicative of their value, neither are the LB scores of single-model kernels.</p>\n\n<p>In the investments field, it is common (in the simplest case) to model the total return of a strategy as <code>R = alpha + beta*M</code>, where <code>M</code> is the return on the \"market portfolio,\" the return that you could get just by mixing all available assets without using any special strategy.  The value of a particular strategy is found not in its total return but in its \"alpha.\"</p>\n\n<p>The analogy with Kaggle competitions is far from perfect, but it is still relevant:  <em>the value of a particular model is found not in its individual score but in how it can improve the score of your final blend</em>.  (Often, for example, in competitions that aren't standard deep learning use cases, Kagglers find that their highest single-model scores come from gradient boosted tree models but that neural network models are critical to maximizing their final score.)  For public kernels, the ones that are most valuable are not the ones with the highest single model scores but the ones that make unique contributions to blend scores.</p>\n\n<p>But you can't identify these most valuable kernels unless you have seen the blends.</p>\n\n<p>It's true that blend scores are not indicative of the value of the blend kernel itself, but (assuming the base models are also public kernels that are linked from blend kernel's \"Data\" tab), they are indicative of where you want to look to find the most valuable kernels.</p>",
  "messages": [
    {
      "id": 304416,
      "postDate": "2018-03-27T15:19:06.937Z",
      "content": "<p>Unless blending is done in public (and I mean, fully public, with the inputs to the blend being public as well), the public has no way of knowing which single-model kernels are the most promising.  Just as the LB scores of blend kernels are not directly indicative of their value, neither are the LB scores of single-model kernels.</p>\n\n<p>In the investments field, it is common (in the simplest case) to model the total return of a strategy as <code>R = alpha + beta*M</code>, where <code>M</code> is the return on the \"market portfolio,\" the return that you could get just by mixing all available assets without using any special strategy.  The value of a particular strategy is found not in its total return but in its \"alpha.\"</p>\n\n<p>The analogy with Kaggle competitions is far from perfect, but it is still relevant:  <em>the value of a particular model is found not in its individual score but in how it can improve the score of your final blend</em>.  (Often, for example, in competitions that aren't standard deep learning use cases, Kagglers find that their highest single-model scores come from gradient boosted tree models but that neural network models are critical to maximizing their final score.)  For public kernels, the ones that are most valuable are not the ones with the highest single model scores but the ones that make unique contributions to blend scores.</p>\n\n<p>But you can't identify these most valuable kernels unless you have seen the blends.</p>\n\n<p>It's true that blend scores are not indicative of the value of the blend kernel itself, but (assuming the base models are also public kernels that are linked from blend kernel's \"Data\" tab), they are indicative of where you want to look to find the most valuable kernels.</p>",
      "rawMarkdown": "Unless blending is done in public (and I mean, fully public, with the inputs to the blend being public as well), the public has no way of knowing which single-model kernels are the most promising.  Just as the LB scores of blend kernels are not directly indicative of their value, neither are the LB scores of single-model kernels.\n\nIn the investments field, it is common (in the simplest case) to model the total return of a strategy as `R = alpha + beta*M`, where `M` is the return on the \"market portfolio,\" the return that you could get just by mixing all available assets without using any special strategy.  The value of a particular strategy is found not in its total return but in its \"alpha.\"\n\nThe analogy with Kaggle competitions is far from perfect, but it is still relevant:  *the value of a particular model is found not in its individual score but in how it can improve the score of your final blend*.  (Often, for example, in competitions that aren't standard deep learning use cases, Kagglers find that their highest single-model scores come from gradient boosted tree models but that neural network models are critical to maximizing their final score.)  For public kernels, the ones that are most valuable are not the ones with the highest single model scores but the ones that make unique contributions to blend scores.\n\nBut you can't identify these most valuable kernels unless you have seen the blends.\n\nIt's true that blend scores are not indicative of the value of the blend kernel itself, but (assuming the base models are also public kernels that are linked from blend kernel's \"Data\" tab), they are indicative of where you want to look to find the most valuable kernels.",
      "votes": 16
    },
    {
      "id": 304529,
      "postDate": "2018-03-27T18:01:28.250Z",
      "content": "<p>Issue with blending is that it move people away from the basics: proper EDA, feature engineering, and local validation setting.  </p>\n\n<p>In this competition, it is possible to do way better with a single lgb model than any of the public blends (I have models at 0.971x, and I'm sure top ranked competitors are way higher than that with single models).  </p>",
      "rawMarkdown": "Issue with blending is that it move people away from the basics: proper EDA, feature engineering, and local validation setting.  \n\nIn this competition, it is possible to do way better with a single lgb model than any of the public blends (I have models at 0.971x, and I'm sure top ranked competitors are way higher than that with single models).  \n",
      "votes": 13,
      "replies": [
        {
          "id": 304568,
          "postDate": "2018-03-27T18:51:01.793Z",
          "content": "<p>EDA kernels tend to get the most upvotes anyhow, so I don't think blending is a threat to them.  As for local validation, well, almost all high-scoring public kernels, blends or not, are a distraction: that's something people just have to learn.  It is useful to see the feature engineering in public kernels, and that's one reason I suggest going to the data tab of blend kernels and clicking on the base model links (when they're available).  Aside from public kernels, my experience so far in this competition has convinced me that blending (or stacking) is going to be important: it's not really fair to compare the best non-public single models to the best public blends.  (I have private single models, which won't run on Kaggle because of memory limitations, that do better than the best public blends; but I also have private blends that do considerably better than my best private single models.)</p>",
          "rawMarkdown": "EDA kernels tend to get the most upvotes anyhow, so I don't think blending is a threat to them.  As for local validation, well, almost all high-scoring public kernels, blends or not, are a distraction: that's something people just have to learn.  It is useful to see the feature engineering in public kernels, and that's one reason I suggest going to the data tab of blend kernels and clicking on the base model links (when they're available).  Aside from public kernels, my experience so far in this competition has convinced me that blending (or stacking) is going to be important: it's not really fair to compare the best non-public single models to the best public blends.  (I have private single models, which won't run on Kaggle because of memory limitations, that do better than the best public blends; but I also have private blends that do considerably better than my best private single models.)",
          "votes": 3
        },
        {
          "id": 304629,
          "postDate": "2018-03-27T20:06:07.737Z",
          "content": "<p>I agree with CPMP. Besides, it's still early in the competition... At this point I only focus on feature engineering to improve a simple single model. My current best submission (0.9764) is such a simple model with 4 engineered features. So people should focus on creating more intelligent features because that's where you can improve your models.</p>",
          "rawMarkdown": "I agree with CPMP. Besides, it's still early in the competition... At this point I only focus on feature engineering to improve a simple single model. My current best submission (0.9764) is such a simple model with 4 engineered features. So people should focus on creating more intelligent features because that's where you can improve your models.",
          "votes": 32
        },
        {
          "id": 304650,
          "postDate": "2018-03-27T20:42:20.580Z",
          "content": "<blockquote>\n  <p>My current best submission (0.9764) is such a simple model</p>\n</blockquote>\n\n<p>I knew it! ;)</p>",
          "rawMarkdown": "&gt; My current best submission (0.9764) is such a simple model\n\nI knew it! ;)",
          "votes": 2
        },
        {
          "id": 304675,
          "postDate": "2018-03-27T21:30:01.060Z",
          "content": "<p>@Danijel Kivaranovic, I wish I could upvote your post 100 times. Thank you very much for the hint.</p>",
          "rawMarkdown": "@Danijel Kivaranovic, I wish I could upvote your post 100 times. Thank you very much for the hint.",
          "votes": 3
        },
        {
          "id": 305368,
          "postDate": "2018-03-28T19:27:31.983Z",
          "content": "<p>That is amazing. Can I ask how many data do you use for training?</p>",
          "rawMarkdown": "That is amazing. Can I ask how many data do you use for training?",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 305385,
          "postDate": "2018-03-28T20:00:12.390Z",
          "content": "<p>@Dire sure, I use the full train set for training </p>",
          "rawMarkdown": "@Dire sure, I use the full train set for training ",
          "votes": 1
        },
        {
          "id": 305690,
          "postDate": "2018-03-29T09:28:10.080Z",
          "content": "<p>@Danijel Kivaranovic: wow, a really high score with 4 features only. well done. It is not going that well for me (yet). everytime i think i found a very good feature it tend to overfit like crazy and the model hit early stopping limit. i have been looking at the data a lot and can't figure out what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks. Like a user would click on an app several times, even using the same channel, and not purchase it and then suddenly decides to purchase it. i am missing part of the puzzle it seems. </p>",
          "rawMarkdown": "@Danijel Kivaranovic: wow, a really high score with 4 features only. well done. It is not going that well for me (yet). everytime i think i found a very good feature it tend to overfit like crazy and the model hit early stopping limit. i have been looking at the data a lot and can't figure out what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks. Like a user would click on an app several times, even using the same channel, and not purchase it and then suddenly decides to purchase it. i am missing part of the puzzle it seems. \n"
        },
        {
          "id": 305703,
          "postDate": "2018-03-29T09:43:17.213Z",
          "content": "<p>He didn't say it was a lgb model ;)</p>",
          "rawMarkdown": "He didn't say it was a lgb model ;)",
          "votes": 2
        },
        {
          "id": 305707,
          "postDate": "2018-03-29T09:48:31.317Z",
          "content": "<p>Haha yeah but it is a lgb model. No magic behind there ;)</p>",
          "rawMarkdown": "Haha yeah but it is a lgb model. No magic behind there ;)",
          "votes": 9
        },
        {
          "id": 305738,
          "postDate": "2018-03-29T10:31:50.643Z",
          "content": "<p>so it must be some magical 4 features :) </p>",
          "rawMarkdown": "so it must be some magical 4 features :) "
        },
        {
          "id": 305754,
          "postDate": "2018-03-29T10:51:23.753Z",
          "content": "<p>@Danijel Kivaranovic I was 99.99% sure about that and thanks for confirming. Please pardon me and do one last favor by simply typing  one of these letters: <strong>R</strong>  or <strong>P</strong>  or <strong>C</strong></p>",
          "rawMarkdown": "@Danijel Kivaranovic I was 99.99% sure about that and thanks for confirming. Please pardon me and do one last favor by simply typing  one of these letters: **R**  or **P**  or **C**"
        },
        {
          "id": 305774,
          "postDate": "2018-03-29T11:42:26.063Z",
          "content": "<p>@Sameh... <code>what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks.</code> my thinking is a lot of the time we can't assume this is a single user. Sometimes it may be but there was an earlier post, I think form the admin that many of these ip's are shared and they can also change frequently.... which makes it more challenging. </p>",
          "rawMarkdown": "@Sameh... `what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks.` my thinking is a lot of the time we can't assume this is a single user. Sometimes it may be but there was an earlier post, I think form the admin that many of these ip's are shared and they can also change frequently.... which makes it more challenging. ",
          "votes": 4
        },
        {
          "id": 305961,
          "postDate": "2018-03-29T17:15:02.403Z",
          "content": "<p>@Darragh: yes. what you say make sense. maybe users from the same household or company or any shared space would have the same IP. maybe one way to distinguish them is by calculating the difference in time with previous transaction. if it is in the same second then most likely it is a different user</p>",
          "rawMarkdown": "@Darragh: yes. what you say make sense. maybe users from the same household or company or any shared space would have the same IP. maybe one way to distinguish them is by calculating the difference in time with previous transaction. if it is in the same second then most likely it is a different user"
        },
        {
          "id": 306463,
          "postDate": "2018-03-30T12:59:02.067Z",
          "content": "<p>prediction for the not new IPs are very hard. in my CV, the roc for records with new IPs is very high (close to 99%) while it is close to 95% for records with existing IPs. </p>",
          "rawMarkdown": "prediction for the not new IPs are very hard. in my CV, the roc for records with new IPs is very high (close to 99%) while it is close to 95% for records with existing IPs. "
        }
      ]
    },
    {
      "id": 304425,
      "postDate": "2018-03-27T15:38:44.353Z",
      "content": "<p>I find the concept of blending useful and interesting.  However when I see blends, I'd like to know how the authors decided on different weights.  Is it trial and error, or was there some kind of algorithm than made one choose which model to weight at how much?  A more thorough explanation on why a particular mix of models works and how it was chosen would be helpful.</p>",
      "rawMarkdown": "I find the concept of blending useful and interesting.  However when I see blends, I'd like to know how the authors decided on different weights.  Is it trial and error, or was there some kind of algorithm than made one choose which model to weight at how much?  A more thorough explanation on why a particular mix of models works and how it was chosen would be helpful.",
      "votes": 2,
      "replies": [
        {
          "id": 304433,
          "postDate": "2018-03-27T15:52:54.393Z",
          "content": "<p>A lot of them are based on LB feedback, but most Kaggle Kernels choose weights based on feel and intuition rather than an algorithmic approach.</p>",
          "rawMarkdown": "A lot of them are based on LB feedback, but most Kaggle Kernels choose weights based on feel and intuition rather than an algorithmic approach."
        },
        {
          "id": 304436,
          "postDate": "2018-03-27T16:02:23.787Z",
          "content": "<p>I think there is a lot of trial and error. In the Zillow competition, where the data were small enough to do single-kernel blends, I had some kernels (e.g. <a href=\"https://www.kaggle.com/aharless/xgboost-lightgbm-and-ols-and-nn\">this one</a>, on which see the comments at the bottom of the code) where I adjusted the weights in successive versions by using a quadratic approximation for the relation between weights and LB scores.  (I don't necessarily recommend that strategy, since it will tend to overfit the public LB.)  There is also a lot of intuition, and rules of thumb, like that the highest-scoring individual models should get the highest weight and that less high-scoring models should be included when their predictions have low correlations with the higher-weighted models.</p>\n\n<p>But the respectable way to choose blending weights is to use validation data (either a holdout set or out-of-fold data).  Unfortunately that's difficult to do with public kernels because they aren't typically designed with the same validation procedure in mind, so one would have to rewrite the base model kernels to produce consistent sets of validation predictions.  I may try that (since I should do it anyway, for my own submissions).  I may even try to get into fancier stacking approaches in public.  (In Porto Seguro, there was a <a href=\"https://www.kaggle.com/yekenot/2-level-stacker-silver-solution\">public stacking kernel</a> that got into the silver medal range on the private leaderboard.  As with Zillow, the data were small enough to do the entire stack in one kernel.)</p>",
          "rawMarkdown": "I think there is a lot of trial and error. In the Zillow competition, where the data were small enough to do single-kernel blends, I had some kernels (e.g. [this one][1], on which see the comments at the bottom of the code) where I adjusted the weights in successive versions by using a quadratic approximation for the relation between weights and LB scores.  (I don't necessarily recommend that strategy, since it will tend to overfit the public LB.)  There is also a lot of intuition, and rules of thumb, like that the highest-scoring individual models should get the highest weight and that less high-scoring models should be included when their predictions have low correlations with the higher-weighted models.\n\nBut the respectable way to choose blending weights is to use validation data (either a holdout set or out-of-fold data).  Unfortunately that's difficult to do with public kernels because they aren't typically designed with the same validation procedure in mind, so one would have to rewrite the base model kernels to produce consistent sets of validation predictions.  I may try that (since I should do it anyway, for my own submissions).  I may even try to get into fancier stacking approaches in public.  (In Porto Seguro, there was a [public stacking kernel][2] that got into the silver medal range on the private leaderboard.  As with Zillow, the data were small enough to do the entire stack in one kernel.)\n\n [1]: https://www.kaggle.com/aharless/xgboost-lightgbm-and-ols-and-nn\n [2]: https://www.kaggle.com/yekenot/2-level-stacker-silver-solution"
        },
        {
          "id": 304446,
          "postDate": "2018-03-27T16:20:05.260Z",
          "content": "<p>I think even basic description of rationale/strategy would help, because not everyone knows all the same rules of thumb.  I just learned some from your comment here!  </p>\n\n<p>For this particular competition data is huge.  It literally takes me hours to get any one new adjusted model to re-predict something.  So if writers of kernels could say something \"I based it on intuition\", or \"works well on my local validation\",  or \"I ranked them in order of LB score\", \"order of local val score\", \"i tested 5 permutations\", etc, that would provide more value and a way to access whether this is something others can use/replicate with their model combos.</p>\n\n<p>For instance right now top scoring blender here: <a href=\"https://www.kaggle.com/tunguz/psst-wanna-blend-some-more\">https://www.kaggle.com/tunguz/psst-wanna-blend-some-more</a> has no explanation whatsoever for the many multipliers it uses, or why he chose to break the weights into multiples.  If the author gave at least some reasoning for it, I think people would be both less frustrated and find it more useful.  Same here: <a href=\"https://www.kaggle.com/sergeyzlobin/wanna-blend/code\">https://www.kaggle.com/sergeyzlobin/wanna-blend/code</a> ...  I do wanna blend, but why this way and not other way?  What is the rationale?  </p>\n\n<p>I personally want to learn to blend the models better, cause I get good individual model results, but then get stuck mixing them up.  But other than blindly copy-pasting blend mixes, I don't know how to use these public blends with my own models.</p>",
          "rawMarkdown": "I think even basic description of rationale/strategy would help, because not everyone knows all the same rules of thumb.  I just learned some from your comment here!  \n\nFor this particular competition data is huge.  It literally takes me hours to get any one new adjusted model to re-predict something.  So if writers of kernels could say something \"I based it on intuition\", or \"works well on my local validation\",  or \"I ranked them in order of LB score\", \"order of local val score\", \"i tested 5 permutations\", etc, that would provide more value and a way to access whether this is something others can use/replicate with their model combos.\n\nFor instance right now top scoring blender here: https://www.kaggle.com/tunguz/psst-wanna-blend-some-more has no explanation whatsoever for the many multipliers it uses, or why he chose to break the weights into multiples.  If the author gave at least some reasoning for it, I think people would be both less frustrated and find it more useful.  Same here: https://www.kaggle.com/sergeyzlobin/wanna-blend/code ...  I do wanna blend, but why this way and not other way?  What is the rationale?  \n\nI personally want to learn to blend the models better, cause I get good individual model results, but then get stuck mixing them up.  But other than blindly copy-pasting blend mixes, I don't know how to use these public blends with my own models.",
          "votes": 4
        },
        {
          "id": 304468,
          "postDate": "2018-03-27T16:52:46.610Z",
          "content": "<p>I think they're trying to limit the influence of similar, but high performing models in their blend. Homogeneity reduces the effectiveness of ensemble methods because these isn't as much information gain compared to placing higher weights on different, but objectively worse models.</p>",
          "rawMarkdown": "I think they're trying to limit the influence of similar, but high performing models in their blend. Homogeneity reduces the effectiveness of ensemble methods because these isn't as much information gain compared to placing higher weights on different, but objectively worse models.",
          "votes": 2
        },
        {
          "id": 304484,
          "postDate": "2018-03-27T17:16:00.010Z",
          "content": "<p>Manually tuning blend weight coefficients based on diversity is like trying to choose regularized regression coefficients to account for multicollinearity by hand. Personally I don't understand why people ever do it this way instead of using an actual (regularized) logistic regression as a meta-learner on out of sample predictions -- I'd pretty much always trust that more than my intuition or public LB feedback.</p>",
          "rawMarkdown": "Manually tuning blend weight coefficients based on diversity is like trying to choose regularized regression coefficients to account for multicollinearity by hand. Personally I don't understand why people ever do it this way instead of using an actual (regularized) logistic regression as a meta-learner on out of sample predictions -- I'd pretty much always trust that more than my intuition or public LB feedback.",
          "votes": 4
        },
        {
          "id": 304485,
          "postDate": "2018-03-27T17:16:41.783Z",
          "content": "<p>Also, though it may seem unintuitive, negative coefficients can be extremely useful for linear ensembles.</p>",
          "rawMarkdown": "Also, though it may seem unintuitive, negative coefficients can be extremely useful for linear ensembles.",
          "votes": 3
        },
        {
          "id": 304499,
          "postDate": "2018-03-27T17:27:54.110Z",
          "content": "<p>A lot of optimized Kaggle solutions use neural networks(dnn/cnn/etc) as a meta learner. It's just easier and faster to climb the LB using weights. ;)</p>",
          "rawMarkdown": "A lot of optimized Kaggle solutions use neural networks(dnn/cnn/etc) as a meta learner. It's just easier and faster to climb the LB using weights. ;)",
          "votes": 2
        },
        {
          "id": 304500,
          "postDate": "2018-03-27T17:29:15.993Z",
          "content": "<p>That's fair, it's definitely faster/easier than proper out of sample weight selection</p>",
          "rawMarkdown": "That's fair, it's definitely faster/easier than proper out of sample weight selection",
          "votes": 2
        },
        {
          "id": 304534,
          "postDate": "2018-03-27T18:07:01.823Z",
          "content": "<p>As I said, in a smaller data competition like Porto Seguro, doing legitimate stacking in a public kernel was not that hard.  In this competition, one would almost certainly have to use multiple kernels, and probably they would all involve a non-trivial amount of coding.  We may hope that someone (maybe me) will do that eventually, but for now simple blends based on intuition and trial and error are the best indication available in public for how well different models blend.</p>",
          "rawMarkdown": "As I said, in a smaller data competition like Porto Seguro, doing legitimate stacking in a public kernel was not that hard.  In this competition, one would almost certainly have to use multiple kernels, and probably they would all involve a non-trivial amount of coding.  We may hope that someone (maybe me) will do that eventually, but for now simple blends based on intuition and trial and error are the best indication available in public for how well different models blend.",
          "votes": 2
        },
        {
          "id": 304549,
          "postDate": "2018-03-27T18:23:51.447Z",
          "content": "<p>Yeah, I think it's a fair point and agree now it's totally understandable why it's done. But I do also agree with CPMP that it can be a distraction from methods that people might learn more from and may be more successful like good FE.</p>",
          "rawMarkdown": "Yeah, I think it's a fair point and agree now it's totally understandable why it's done. But I do also agree with CPMP that it can be a distraction from methods that people might learn more from and may be more successful like good FE.",
          "votes": 4
        },
        {
          "id": 304579,
          "postDate": "2018-03-27T19:06:18.893Z",
          "content": "<p>@Joe Eddy What resources would you recommend for </p>\n\n<blockquote>\n  <p>\"negative coefficients can be extremely useful for linear ensembles\"</p>\n</blockquote>\n\n<p>I am trying to wrap my head around stacking with logistic regression. Also how does it differ from @Andy 's <a href=\"https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange\">Logit Stacking</a>?</p>",
          "rawMarkdown": "@Joe Eddy What resources would you recommend for \n&gt; \"negative coefficients can be extremely useful for linear ensembles\"\n\nI am trying to wrap my head around stacking with logistic regression. Also how does it differ from @Andy 's [Logit Stacking](https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange)?"
        },
        {
          "id": 304606,
          "postDate": "2018-03-27T19:39:09.167Z",
          "content": "<p>@Nick Brooks A minor point, but my kernel that you linked isn't really stacking, or even blending in the usual sense (hence the title). That particular kernel is more like bagging (although it's not the usual sort of bagging either, because the samples are disjoint and chosen based on a time cutoff rather than randomness).  The point of that kernel was to get around Kaggle's memory limitation by doing two separate fits to two separate parts of the dataset in two separate kernels that feed into that one.</p>\n\n<p>More generally, though, there are two separate issues with respect to logits.  One issue is how you code the base model predictions. (They are naturally expressed as probabilities in submission files, but logits may be better for stacking/blending.)  The other issue is how you combine the base models. @Joe Eddy is suggesting using a logistic regression, which essentially finds an optimal linear combination of logits (so it would normally be appropriate to express the inputs as logits, since the model will interpret them that way, but it may also get useful information from inputs expressed in other ways).  You could think of my logit blends as just guessing the optimal coefficients (weights) for a logistic regression.  When all our out-of-sample predictions are for the test data, we can't run an actual logistic regression, since we have no targets.  But Joe is suggesting (with good reason) that it would be better to generate out-of-sample predictions for validation data and run an actual logistic regression on those.</p>",
          "rawMarkdown": "@Nick Brooks A minor point, but my kernel that you linked isn't really stacking, or even blending in the usual sense (hence the title). That particular kernel is more like bagging (although it's not the usual sort of bagging either, because the samples are disjoint and chosen based on a time cutoff rather than randomness).  The point of that kernel was to get around Kaggle's memory limitation by doing two separate fits to two separate parts of the dataset in two separate kernels that feed into that one.\n\nMore generally, though, there are two separate issues with respect to logits.  One issue is how you code the base model predictions. (They are naturally expressed as probabilities in submission files, but logits may be better for stacking/blending.)  The other issue is how you combine the base models. @Joe Eddy is suggesting using a logistic regression, which essentially finds an optimal linear combination of logits (so it would normally be appropriate to express the inputs as logits, since the model will interpret them that way, but it may also get useful information from inputs expressed in other ways).  You could think of my logit blends as just guessing the optimal coefficients (weights) for a logistic regression.  When all our out-of-sample predictions are for the test data, we can't run an actual logistic regression, since we have no targets.  But Joe is suggesting (with good reason) that it would be better to generate out-of-sample predictions for validation data and run an actual logistic regression on those.",
          "votes": 4
        },
        {
          "id": 307965,
          "postDate": "2018-04-02T18:20:16.260Z",
          "content": "<p>OK. Now I have a public kernel with <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704\">respectable blending</a> (technically, linear stacking, I think, but the distinction in terminology is not well-established), where you can see where the weights come from.</p>",
          "rawMarkdown": "OK. Now I have a public kernel with [respectable blending][1] (technically, linear stacking, I think, but the distinction in terminology is not well-established), where you can see where the weights come from.\n\n\n  [1]: https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704",
          "votes": 1
        },
        {
          "id": 314058,
          "postDate": "2018-04-14T14:27:23.760Z",
          "content": "<p>But now I'm realizing once again what a nuisance the logistic regularization parameter is.  There's nothing particularly special about the default C=1.  And unless the base model predictions are somehow normalized, the parameter value is nonsense, because the same parameter will give a completely different answer if you change the scale of the inputs.  So we're back to intuition and trial and error.  Do we expect to have more reliable intuitions about meta-learner parameters than about raw weights?</p>",
          "rawMarkdown": "But now I'm realizing once again what a nuisance the logistic regularization parameter is.  There's nothing particularly special about the default C=1.  And unless the base model predictions are somehow normalized, the parameter value is nonsense, because the same parameter will give a completely different answer if you change the scale of the inputs.  So we're back to intuition and trial and error.  Do we expect to have more reliable intuitions about meta-learner parameters than about raw weights?"
        },
        {
          "id": 314063,
          "postDate": "2018-04-14T14:35:32.580Z",
          "content": "<p>What about using cross validation, or at least some train/val split?</p>",
          "rawMarkdown": "What about using cross validation, or at least some train/val split?"
        },
        {
          "id": 314079,
          "postDate": "2018-04-14T14:53:04.720Z",
          "content": "<p>NNs are non-parametric and are strong candidates for meta-learners. I think the big problem with using a FF NN here is that it overfits to the temporal trends of the data.</p>",
          "rawMarkdown": "NNs are non-parametric and are strong candidates for meta-learners. I think the big problem with using a FF NN here is that it overfits to the temporal trends of the data."
        },
        {
          "id": 314091,
          "postDate": "2018-04-14T15:32:20.247Z",
          "content": "<p>As with using cross-validation to choose a logistic regularization parameter, the temporal trends are a pain in the neck.  But I guess, since the data are so large, one may have several opportunities to sub-split the validation data by time to validate stacking parameters.</p>",
          "rawMarkdown": "As with using cross-validation to choose a logistic regularization parameter, the temporal trends are a pain in the neck.  But I guess, since the data are so large, one may have several opportunities to sub-split the validation data by time to validate stacking parameters."
        }
      ]
    },
    {
      "id": 306938,
      "postDate": "2018-03-31T11:23:04.317Z",
      "content": "<p>I may say that high public LB kernels are close to growth stocks. Value stocks are usually better but we lack indicators to get a kernel value (validation strategy, CV score, overfiting...). I don't think there is an equivalent of sales to stock value ratio to select kernels...</p>",
      "rawMarkdown": "I may say that high public LB kernels are close to growth stocks. Value stocks are usually better but we lack indicators to get a kernel value (validation strategy, CV score, overfiting...). I don't think there is an equivalent of sales to stock value ratio to select kernels...",
      "replies": [
        {
          "id": 306978,
          "postDate": "2018-03-31T13:33:17.310Z",
          "content": "<p>We need more cross-value-dation</p>",
          "rawMarkdown": "We need more cross-value-dation",
          "votes": 1
        }
      ]
    },
    {
      "id": 304426,
      "postDate": "2018-03-27T15:40:33.643Z",
      "content": "<p><em>Footnote</em> The analogy, as stated, breaks down because investment strategies are typically scalable, whereas Kaggle strategies typically are not (at least not within the limited context of a particular competition). It might be more precise to use a non-scalable aspect of investment strategy as the analogy:  for example, \"Seek alpha, not a high Sharpe ratio.\"  But that makes it more complicated to explain. And of course most people use multiple-factor models these days, so beta is a vector, and there is more than one relevant market portfolio.  And some of the factors are regarded as being themselves exploitable (\"smart beta\").  And so on.  But all this complication would distract from the main point, which I think still holds.</p>",
      "rawMarkdown": "*Footnote* The analogy, as stated, breaks down because investment strategies are typically scalable, whereas Kaggle strategies typically are not (at least not within the limited context of a particular competition). It might be more precise to use a non-scalable aspect of investment strategy as the analogy:  for example, \"Seek alpha, not a high Sharpe ratio.\"  But that makes it more complicated to explain. And of course most people use multiple-factor models these days, so beta is a vector, and there is more than one relevant market portfolio.  And some of the factors are regarded as being themselves exploitable (\"smart beta\").  And so on.  But all this complication would distract from the main point, which I think still holds."
    }
  ],
  "comments": [
    {
      "id": 304529,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-03-27T18:01:28.250000",
      "content": "<p>Issue with blending is that it move people away from the basics: proper EDA, feature engineering, and local validation setting.  </p>\n\n<p>In this competition, it is possible to do way better with a single lgb model than any of the public blends (I have models at 0.971x, and I'm sure top ranked competitors are way higher than that with single models).  </p>",
      "votes": 13,
      "replies": [
        {
          "id": 304568,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-27T18:51:01.793000",
          "content": "<p>EDA kernels tend to get the most upvotes anyhow, so I don't think blending is a threat to them.  As for local validation, well, almost all high-scoring public kernels, blends or not, are a distraction: that's something people just have to learn.  It is useful to see the feature engineering in public kernels, and that's one reason I suggest going to the data tab of blend kernels and clicking on the base model links (when they're available).  Aside from public kernels, my experience so far in this competition has convinced me that blending (or stacking) is going to be important: it's not really fair to compare the best non-public single models to the best public blends.  (I have private single models, which won't run on Kaggle because of memory limitations, that do better than the best public blends; but I also have private blends that do considerably better than my best private single models.)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 304629,
          "author_name": "Danijel Kivaranovic",
          "author_url": "",
          "post_date": "2018-03-27T20:06:07.737000",
          "content": "<p>I agree with CPMP. Besides, it's still early in the competition... At this point I only focus on feature engineering to improve a simple single model. My current best submission (0.9764) is such a simple model with 4 engineered features. So people should focus on creating more intelligent features because that's where you can improve your models.</p>",
          "votes": 32,
          "replies": []
        },
        {
          "id": 304650,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-03-27T20:42:20.580000",
          "content": "<blockquote>\n  <p>My current best submission (0.9764) is such a simple model</p>\n</blockquote>\n\n<p>I knew it! ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304675,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-03-27T21:30:01.060000",
          "content": "<p>@Danijel Kivaranovic, I wish I could upvote your post 100 times. Thank you very much for the hint.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 305368,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-03-28T19:27:31.983000",
          "content": "<p>That is amazing. Can I ask how many data do you use for training?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 305385,
          "author_name": "Danijel Kivaranovic",
          "author_url": "",
          "post_date": "2018-03-28T20:00:12.390000",
          "content": "<p>@Dire sure, I use the full train set for training </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 305690,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-03-29T09:28:10.080000",
          "content": "<p>@Danijel Kivaranovic: wow, a really high score with 4 features only. well done. It is not going that well for me (yet). everytime i think i found a very good feature it tend to overfit like crazy and the model hit early stopping limit. i have been looking at the data a lot and can't figure out what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks. Like a user would click on an app several times, even using the same channel, and not purchase it and then suddenly decides to purchase it. i am missing part of the puzzle it seems. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 305703,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-03-29T09:43:17.213000",
          "content": "<p>He didn't say it was a lgb model ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 305707,
          "author_name": "Danijel Kivaranovic",
          "author_url": "",
          "post_date": "2018-03-29T09:48:31.317000",
          "content": "<p>Haha yeah but it is a lgb model. No magic behind there ;)</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 305738,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-03-29T10:31:50.643000",
          "content": "<p>so it must be some magical 4 features :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 305754,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-03-29T10:51:23.753000",
          "content": "<p>@Danijel Kivaranovic I was 99.99% sure about that and thanks for confirming. Please pardon me and do one last favor by simply typing  one of these letters: <strong>R</strong>  or <strong>P</strong>  or <strong>C</strong></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 305774,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-03-29T11:42:26.063000",
          "content": "<p>@Sameh... <code>what make a user (ip, device, os) purchase an app after a series of what seems to be random clicks.</code> my thinking is a lot of the time we can't assume this is a single user. Sometimes it may be but there was an earlier post, I think form the admin that many of these ip's are shared and they can also change frequently.... which makes it more challenging. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 305961,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-03-29T17:15:02.403000",
          "content": "<p>@Darragh: yes. what you say make sense. maybe users from the same household or company or any shared space would have the same IP. maybe one way to distinguish them is by calculating the difference in time with previous transaction. if it is in the same second then most likely it is a different user</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 306463,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-03-30T12:59:02.067000",
          "content": "<p>prediction for the not new IPs are very hard. in my CV, the roc for records with new IPs is very high (close to 99%) while it is close to 95% for records with existing IPs. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 304425,
      "author_name": "yulia",
      "author_url": "",
      "post_date": "2018-03-27T15:38:44.353000",
      "content": "<p>I find the concept of blending useful and interesting.  However when I see blends, I'd like to know how the authors decided on different weights.  Is it trial and error, or was there some kind of algorithm than made one choose which model to weight at how much?  A more thorough explanation on why a particular mix of models works and how it was chosen would be helpful.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 304433,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-03-27T15:52:54.393000",
          "content": "<p>A lot of them are based on LB feedback, but most Kaggle Kernels choose weights based on feel and intuition rather than an algorithmic approach.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 304436,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-27T16:02:23.787000",
          "content": "<p>I think there is a lot of trial and error. In the Zillow competition, where the data were small enough to do single-kernel blends, I had some kernels (e.g. <a href=\"https://www.kaggle.com/aharless/xgboost-lightgbm-and-ols-and-nn\">this one</a>, on which see the comments at the bottom of the code) where I adjusted the weights in successive versions by using a quadratic approximation for the relation between weights and LB scores.  (I don't necessarily recommend that strategy, since it will tend to overfit the public LB.)  There is also a lot of intuition, and rules of thumb, like that the highest-scoring individual models should get the highest weight and that less high-scoring models should be included when their predictions have low correlations with the higher-weighted models.</p>\n\n<p>But the respectable way to choose blending weights is to use validation data (either a holdout set or out-of-fold data).  Unfortunately that's difficult to do with public kernels because they aren't typically designed with the same validation procedure in mind, so one would have to rewrite the base model kernels to produce consistent sets of validation predictions.  I may try that (since I should do it anyway, for my own submissions).  I may even try to get into fancier stacking approaches in public.  (In Porto Seguro, there was a <a href=\"https://www.kaggle.com/yekenot/2-level-stacker-silver-solution\">public stacking kernel</a> that got into the silver medal range on the private leaderboard.  As with Zillow, the data were small enough to do the entire stack in one kernel.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 304446,
          "author_name": "yulia",
          "author_url": "",
          "post_date": "2018-03-27T16:20:05.260000",
          "content": "<p>I think even basic description of rationale/strategy would help, because not everyone knows all the same rules of thumb.  I just learned some from your comment here!  </p>\n\n<p>For this particular competition data is huge.  It literally takes me hours to get any one new adjusted model to re-predict something.  So if writers of kernels could say something \"I based it on intuition\", or \"works well on my local validation\",  or \"I ranked them in order of LB score\", \"order of local val score\", \"i tested 5 permutations\", etc, that would provide more value and a way to access whether this is something others can use/replicate with their model combos.</p>\n\n<p>For instance right now top scoring blender here: <a href=\"https://www.kaggle.com/tunguz/psst-wanna-blend-some-more\">https://www.kaggle.com/tunguz/psst-wanna-blend-some-more</a> has no explanation whatsoever for the many multipliers it uses, or why he chose to break the weights into multiples.  If the author gave at least some reasoning for it, I think people would be both less frustrated and find it more useful.  Same here: <a href=\"https://www.kaggle.com/sergeyzlobin/wanna-blend/code\">https://www.kaggle.com/sergeyzlobin/wanna-blend/code</a> ...  I do wanna blend, but why this way and not other way?  What is the rationale?  </p>\n\n<p>I personally want to learn to blend the models better, cause I get good individual model results, but then get stuck mixing them up.  But other than blindly copy-pasting blend mixes, I don't know how to use these public blends with my own models.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 304468,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-03-27T16:52:46.610000",
          "content": "<p>I think they're trying to limit the influence of similar, but high performing models in their blend. Homogeneity reduces the effectiveness of ensemble methods because these isn't as much information gain compared to placing higher weights on different, but objectively worse models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304484,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-27T17:16:00.010000",
          "content": "<p>Manually tuning blend weight coefficients based on diversity is like trying to choose regularized regression coefficients to account for multicollinearity by hand. Personally I don't understand why people ever do it this way instead of using an actual (regularized) logistic regression as a meta-learner on out of sample predictions -- I'd pretty much always trust that more than my intuition or public LB feedback.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 304485,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-27T17:16:41.783000",
          "content": "<p>Also, though it may seem unintuitive, negative coefficients can be extremely useful for linear ensembles.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 304499,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-03-27T17:27:54.110000",
          "content": "<p>A lot of optimized Kaggle solutions use neural networks(dnn/cnn/etc) as a meta learner. It's just easier and faster to climb the LB using weights. ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304500,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-27T17:29:15.993000",
          "content": "<p>That's fair, it's definitely faster/easier than proper out of sample weight selection</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304534,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-27T18:07:01.823000",
          "content": "<p>As I said, in a smaller data competition like Porto Seguro, doing legitimate stacking in a public kernel was not that hard.  In this competition, one would almost certainly have to use multiple kernels, and probably they would all involve a non-trivial amount of coding.  We may hope that someone (maybe me) will do that eventually, but for now simple blends based on intuition and trial and error are the best indication available in public for how well different models blend.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 304549,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-27T18:23:51.447000",
          "content": "<p>Yeah, I think it's a fair point and agree now it's totally understandable why it's done. But I do also agree with CPMP that it can be a distraction from methods that people might learn more from and may be more successful like good FE.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 304579,
          "author_name": "nicapotato",
          "author_url": "",
          "post_date": "2018-03-27T19:06:18.893000",
          "content": "<p>@Joe Eddy What resources would you recommend for </p>\n\n<blockquote>\n  <p>\"negative coefficients can be extremely useful for linear ensembles\"</p>\n</blockquote>\n\n<p>I am trying to wrap my head around stacking with logistic regression. Also how does it differ from @Andy 's <a href=\"https://www.kaggle.com/aharless/ceci-n-est-pas-un-m-lange\">Logit Stacking</a>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 304606,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-27T19:39:09.167000",
          "content": "<p>@Nick Brooks A minor point, but my kernel that you linked isn't really stacking, or even blending in the usual sense (hence the title). That particular kernel is more like bagging (although it's not the usual sort of bagging either, because the samples are disjoint and chosen based on a time cutoff rather than randomness).  The point of that kernel was to get around Kaggle's memory limitation by doing two separate fits to two separate parts of the dataset in two separate kernels that feed into that one.</p>\n\n<p>More generally, though, there are two separate issues with respect to logits.  One issue is how you code the base model predictions. (They are naturally expressed as probabilities in submission files, but logits may be better for stacking/blending.)  The other issue is how you combine the base models. @Joe Eddy is suggesting using a logistic regression, which essentially finds an optimal linear combination of logits (so it would normally be appropriate to express the inputs as logits, since the model will interpret them that way, but it may also get useful information from inputs expressed in other ways).  You could think of my logit blends as just guessing the optimal coefficients (weights) for a logistic regression.  When all our out-of-sample predictions are for the test data, we can't run an actual logistic regression, since we have no targets.  But Joe is suggesting (with good reason) that it would be better to generate out-of-sample predictions for validation data and run an actual logistic regression on those.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 307965,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-02T18:20:16.260000",
          "content": "<p>OK. Now I have a public kernel with <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9704\">respectable blending</a> (technically, linear stacking, I think, but the distinction in terminology is not well-established), where you can see where the weights come from.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 314058,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-14T14:27:23.760000",
          "content": "<p>But now I'm realizing once again what a nuisance the logistic regularization parameter is.  There's nothing particularly special about the default C=1.  And unless the base model predictions are somehow normalized, the parameter value is nonsense, because the same parameter will give a completely different answer if you change the scale of the inputs.  So we're back to intuition and trial and error.  Do we expect to have more reliable intuitions about meta-learner parameters than about raw weights?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314063,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-14T14:35:32.580000",
          "content": "<p>What about using cross validation, or at least some train/val split?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314079,
          "author_name": "Alan Khoa Nguyen",
          "author_url": "",
          "post_date": "2018-04-14T14:53:04.720000",
          "content": "<p>NNs are non-parametric and are strong candidates for meta-learners. I think the big problem with using a FF NN here is that it overfits to the temporal trends of the data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 314091,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-14T15:32:20.247000",
          "content": "<p>As with using cross-validation to choose a logistic regularization parameter, the temporal trends are a pain in the neck.  But I guess, since the data are so large, one may have several opportunities to sub-split the validation data by time to validate stacking parameters.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 306938,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-03-31T11:23:04.317000",
      "content": "<p>I may say that high public LB kernels are close to growth stocks. Value stocks are usually better but we lack indicators to get a kernel value (validation strategy, CV score, overfiting...). I don't think there is an equivalent of sales to stock value ratio to select kernels...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 306978,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-31T13:33:17.310000",
          "content": "<p>We need more cross-value-dation</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 304426,
      "author_name": "Andy Harless",
      "author_url": "",
      "post_date": "2018-03-27T15:40:33.643000",
      "content": "<p><em>Footnote</em> The analogy, as stated, breaks down because investment strategies are typically scalable, whereas Kaggle strategies typically are not (at least not within the limited context of a particular competition). It might be more precise to use a non-scalable aspect of investment strategy as the analogy:  for example, \"Seek alpha, not a high Sharpe ratio.\"  But that makes it more complicated to explain. And of course most people use multiple-factor models these days, so beta is a vector, and there is more than one relevant market portfolio.  And some of the factors are regarded as being themselves exploitable (\"smart beta\").  And so on.  But all this complication would distract from the main point, which I think still holds.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "304416": "Unless blending is done in public (and I mean, fully public, with the inputs to the blend being public as well), the public has no way of knowing which single-model kernels are the most promising.  Just as the LB scores of blend kernels are not directly indicative of their value, neither are the LB scores of single-model kernels.\n\nIn the investments field, it is common (in the simplest case) to model the total return of a strategy as `R = alpha + beta*M`, where `M` is the return on the \"market portfolio,\" the return that you could get just by mixing all available assets without using any special strategy.  The value of a particular strategy is found not in its total return but in its \"alpha.\"\n\nThe analogy with Kaggle competitions is far from perfect, but it is still relevant:  *the value of a particular model is found not in its individual score but in how it can improve the score of your final blend*.  (Often, for example, in competitions that aren't standard deep learning use cases, Kagglers find that their highest single-model scores come from gradient boosted tree models but that neural network models are critical to maximizing their final score.)  For public kernels, the ones that are most valuable are not the ones with the highest single model scores but the ones that make unique contributions to blend scores.\n\nBut you can't identify these most valuable kernels unless you have seen the blends.\n\nIt's true that blend scores are not indicative of the value of the blend kernel itself, but (assuming the base models are also public kernels that are linked from blend kernel's \"Data\" tab), they are indicative of where you want to look to find the most valuable kernels.",
    "304529": "Issue with blending is that it move people away from the basics: proper EDA, feature engineering, and local validation setting.  \n\nIn this competition, it is possible to do way better with a single lgb model than any of the public blends (I have models at 0.971x, and I'm sure top ranked competitors are way higher than that with single models).  \n",
    "304425": "I find the concept of blending useful and interesting.  However when I see blends, I'd like to know how the authors decided on different weights.  Is it trial and error, or was there some kind of algorithm than made one choose which model to weight at how much?  A more thorough explanation on why a particular mix of models works and how it was chosen would be helpful.",
    "306938": "I may say that high public LB kernels are close to growth stocks. Value stocks are usually better but we lack indicators to get a kernel value (validation strategy, CV score, overfiting...). I don't think there is an equivalent of sales to stock value ratio to select kernels...",
    "304426": "*Footnote* The analogy, as stated, breaks down because investment strategies are typically scalable, whereas Kaggle strategies typically are not (at least not within the limited context of a particular competition). It might be more precise to use a non-scalable aspect of investment strategy as the analogy:  for example, \"Seek alpha, not a high Sharpe ratio.\"  But that makes it more complicated to explain. And of course most people use multiple-factor models these days, so beta is a vector, and there is more than one relevant market portfolio.  And some of the factors are regarded as being themselves exploitable (\"smart beta\").  And so on.  But all this complication would distract from the main point, which I think still holds."
  }
}