{
  "id": 54483,
  "title": "How to do bagging/blending properly?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54483",
  "author_name": "",
  "post_date": "2018-04-13T19:04:22.121075900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>For most previous competitions, I use stacking to combine different models and different sets of features. But for this one, running the stacking process for 1 model takes me a day or so (like a 5-fold CV) and I think it's not that practical given the CPU/RAM I have. So I'm looking at more straightforward bagging/blending strategies.</p>\n\n<p>The pred results from different models (py lgbm /frtl/R lgbm/ nn ) are quite different - normally this is a good sign that bagging/blending will give you a big boost. However, I tried simple averaging and averaging after normalization - neither helps as much as I would expect. </p>\n\n<p>So, what do you think is the proper bagging/blending strategy for this competition?</p>",
  "messages": [
    {
      "id": "313728",
      "postDate": "04/13/2018 19:04:22",
      "content": "<p>For most previous competitions, I use stacking to combine different models and different sets of features. But for this one, running the stacking process for 1 model takes me a day or so (like a 5-fold CV) and I think it's not that practical given the CPU/RAM I have. So I'm looking at more straightforward bagging/blending strategies.</p>\n\n<p>The pred results from different models (py lgbm /frtl/R lgbm/ nn ) are quite different - normally this is a good sign that bagging/blending will give you a big boost. However, I tried simple averaging and averaging after normalization - neither helps as much as I would expect. </p>\n\n<p>So, what do you think is the proper bagging/blending strategy for this competition?</p>",
      "rawMarkdown": "For most previous competitions, I use stacking to combine different models and different sets of features. But for this one, running the stacking process for 1 model takes me a day or so (like a 5-fold CV) and I think it's not that practical given the CPU/RAM I have. So I'm looking at more straightforward bagging/blending strategies.\n\nThe pred results from different models (py lgbm /frtl/R lgbm/ nn ) are quite different - normally this is a good sign that bagging/blending will give you a big boost. However, I tried simple averaging and averaging after normalization - neither helps as much as I would expect. \n\nSo, what do you think is the proper bagging/blending strategy for this competition?",
      "votes": null
    },
    {
      "id": "313758",
      "postDate": "04/13/2018 19:45:55",
      "content": "<p>Here is a great write up by @Andy over blending, which tends to blend well <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730\">Andy's Stacker</a>.</p>\n\n<p>Apart from this I am using similar approach but by optimizing my validation set weights with Bayesian optimization and propagating the  weights of validation to the test set.</p>",
      "rawMarkdown": "Here is a great write up by @Andy over blending, which tends to blend well [Andy's Stacker][1].\n\nApart from this I am using similar approach but by optimizing my validation set weights with Bayesian optimization and propagating the  weights of validation to the test set.\n\n\n  [1]: https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730",
      "votes": null
    },
    {
      "id": "313762",
      "postDate": "04/13/2018 19:55:05",
      "content": "<p>Thx! I think Andy's kernel is essentially a stacking process- you need the CV files and the 2nd layer model is a linear/logistic regression. However, computing the CV files is very time-consuming and that is where I'm stuck. Is there a way to blend with only the submission files and get reasonably good boost?</p>",
      "rawMarkdown": "Thx! I think Andy's kernel is essentially a stacking process- you need the CV files and the 2nd layer model is a linear/logistic regression. However, computing the CV files is very time-consuming and that is where I'm stuck. Is there a way to blend with only the submission files and get reasonably good boost?",
      "votes": null
    },
    {
      "id": "313785",
      "postDate": "04/13/2018 20:40:56",
      "content": "<p>giving weights to your blending files according to their validation scores</p>",
      "rawMarkdown": "giving weights to your blending files according to their validation scores",
      "votes": null
    },
    {
      "id": "313788",
      "postDate": "04/13/2018 20:43:20",
      "content": "<p>Its always a best practice to use validation to compute weights, but as you mentioned it might get expensive sometimes. In such cases blending based on correlation of submission files is also a way to go .</p>\n\n<p>You can have a look at this interesting kernel by @Ryan <a href=\"https://www.kaggle.com/reppic/lazy-ensembling-algorithm\">Ryan's Lazy ensembling</a></p>",
      "rawMarkdown": "Its always a best practice to use validation to compute weights, but as you mentioned it might get expensive sometimes. In such cases blending based on correlation of submission files is also a way to go .\n\nYou can have a look at this interesting kernel by @Ryan [Ryan's Lazy ensembling][1]\n\n\n  [1]: https://www.kaggle.com/reppic/lazy-ensembling-algorithm",
      "votes": null
    },
    {
      "id": "313796",
      "postDate": "04/13/2018 20:54:46",
      "content": "<p>@shivraj, this ensembling method looks very interesting to me. I'm trying it as soon as I get my new sub opportunities. thank you a lot!</p>\n\n<p>@Snorlax, I think I tried that without much improvement. When I check the distribution of the predictions by different models, they really look different. And when averaging (even with weights), I feel one or two of them dominate the others. I'm unsure though - am still looking into it. Thx!</p>",
      "rawMarkdown": "shivraj, this ensembling method looks very interesting to me. I'm trying it as soon as I get my new sub opportunities. thank you a lot!\n\n@Snorlax, I think I tried that without much improvement. When I check the distribution of the predictions by different models, they really look different. And when averaging (even with weights), I feel one or two of them dominate the others. I'm unsure though - am still looking into it. Thx!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 313758,
      "author_name": "shivrajp",
      "author_url": "",
      "post_date": "04/13/2018 19:45:55",
      "content": "<p>Here is a great write up by @Andy over blending, which tends to blend well <a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730\">Andy's Stacker</a>.</p>\n\n<p>Apart from this I am using similar approach but by optimizing my validation set weights with Bayesian optimization and propagating the  weights of validation to the test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 313762,
          "author_name": "cizonb",
          "author_url": "",
          "post_date": "04/13/2018 19:55:05",
          "content": "<p>Thx! I think Andy's kernel is essentially a stacking process- you need the CV files and the 2nd layer model is a linear/logistic regression. However, computing the CV files is very time-consuming and that is where I'm stuck. Is there a way to blend with only the submission files and get reasonably good boost?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313785,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/13/2018 20:40:56",
          "content": "<p>giving weights to your blending files according to their validation scores</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313788,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "04/13/2018 20:43:20",
          "content": "<p>Its always a best practice to use validation to compute weights, but as you mentioned it might get expensive sometimes. In such cases blending based on correlation of submission files is also a way to go .</p>\n\n<p>You can have a look at this interesting kernel by @Ryan <a href=\"https://www.kaggle.com/reppic/lazy-ensembling-algorithm\">Ryan's Lazy ensembling</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313796,
          "author_name": "cizonb",
          "author_url": "",
          "post_date": "04/13/2018 20:54:46",
          "content": "<p>@shivraj, this ensembling method looks very interesting to me. I'm trying it as soon as I get my new sub opportunities. thank you a lot!</p>\n\n<p>@Snorlax, I think I tried that without much improvement. When I check the distribution of the predictions by different models, they really look different. And when averaging (even with weights), I feel one or two of them dominate the others. I'm unsure though - am still looking into it. Thx!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "313728": "For most previous competitions, I use stacking to combine different models and different sets of features. But for this one, running the stacking process for 1 model takes me a day or so (like a 5-fold CV) and I think it's not that practical given the CPU/RAM I have. So I'm looking at more straightforward bagging/blending strategies.\n\nThe pred results from different models (py lgbm /frtl/R lgbm/ nn ) are quite different - normally this is a good sign that bagging/blending will give you a big boost. However, I tried simple averaging and averaging after normalization - neither helps as much as I would expect. \n\nSo, what do you think is the proper bagging/blending strategy for this competition?",
    "313758": "Here is a great write up by @Andy over blending, which tends to blend well [Andy's Stacker][1].\n\nApart from this I am using similar approach but by optimizing my validation set weights with Bayesian optimization and propagating the  weights of validation to the test set.\n\n\n  [1]: https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730",
    "313762": "Thx! I think Andy's kernel is essentially a stacking process- you need the CV files and the 2nd layer model is a linear/logistic regression. However, computing the CV files is very time-consuming and that is where I'm stuck. Is there a way to blend with only the submission files and get reasonably good boost?",
    "313785": "giving weights to your blending files according to their validation scores",
    "313788": "Its always a best practice to use validation to compute weights, but as you mentioned it might get expensive sometimes. In such cases blending based on correlation of submission files is also a way to go .\n\nYou can have a look at this interesting kernel by @Ryan [Ryan's Lazy ensembling][1]\n\n\n  [1]: https://www.kaggle.com/reppic/lazy-ensembling-algorithm",
    "313796": "shivraj, this ensembling method looks very interesting to me. I'm trying it as soon as I get my new sub opportunities. thank you a lot!\n\n@Snorlax, I think I tried that without much improvement. When I check the distribution of the predictions by different models, they really look different. And when averaging (even with weights), I feel one or two of them dominate the others. I'm unsure though - am still looking into it. Thx!"
  },
  "source": "meta"
}