{
  "id": 90419,
  "title": "Ensembling best practices for regression problems",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90419",
  "author_name": "",
  "post_date": "2019-04-23T17:56:40.002186400Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I've been looking for strategies in ensembling and found this great summary:\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D\">https://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D</a></p>\n\n<p>However, it seems most if not all the points mentioned are about classification not regression.</p>\n\n<p>What is your advice on ensembling in such competition, where stacking seems to be a bad option (see Andrew's kernel), possibly due to leakage?\nHow would you optimize the weights when blending? <br>\nAre weak models useful at all when blending?</p>\n\n<p>Please share your insights on the best practices for combining regression models.</p>",
  "messages": [
    {
      "id": "521975",
      "postDate": "04/23/2019 17:56:40",
      "content": "<p>I've been looking for strategies in ensembling and found this great summary:\n<a href=\"https://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D\">https://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D</a></p>\n\n<p>However, it seems most if not all the points mentioned are about classification not regression.</p>\n\n<p>What is your advice on ensembling in such competition, where stacking seems to be a bad option (see Andrew's kernel), possibly due to leakage?\nHow would you optimize the weights when blending? <br>\nAre weak models useful at all when blending?</p>\n\n<p>Please share your insights on the best practices for combining regression models.</p>",
      "rawMarkdown": "I've been looking for strategies in ensembling and found this great summary:\nhttps://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D\n\nHowever, it seems most if not all the points mentioned are about classification not regression.\n\nWhat is your advice on ensembling in such competition, where stacking seems to be a bad option (see Andrew's kernel), possibly due to leakage?\nHow would you optimize the weights when blending?  \nAre weak models useful at all when blending?\n\nPlease share your insights on the best practices for combining regression models.",
      "votes": null
    },
    {
      "id": "521986",
      "postDate": "04/23/2019 18:10:30",
      "content": "<p>There is no methodological difference between ensembling for classification and regression. If using neural networks for ensembling, be sure to have a sigmoid output activation at the end for classification so your predictions are bound by 0-1. Leave it linear for regression.</p>\n\n<p>As to the usefulness of weak models, see if <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51058\"><strong>my earlier post</strong></a> answers your question.</p>",
      "rawMarkdown": "There is no methodological difference between ensembling for classification and regression. If using neural networks for ensembling, be sure to have a sigmoid output activation at the end for classification so your predictions are bound by 0-1. Leave it linear for regression.\n\nAs to the usefulness of weak models, see if [__my earlier post__](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51058) answers your question.",
      "votes": null
    },
    {
      "id": "521993",
      "postDate": "04/23/2019 18:29:56",
      "content": "<p>Great post indeed, very nice to read. \nI think when it comes to stacking, regression is no different from classification, but what about blending?\nThis competition, in my opinion, is tricky and dangerous when it comes to stacking, especially depending on the CV strategy</p>",
      "rawMarkdown": "Great post indeed, very nice to read. \nI think when it comes to stacking, regression is no different from classification, but what about blending?\nThis competition, in my opinion, is tricky and dangerous when it comes to stacking, especially depending on the CV strategy",
      "votes": null
    },
    {
      "id": "522008",
      "postDate": "04/23/2019 18:43:56",
      "content": "<p>what is  Andrew's kernel? </p>",
      "rawMarkdown": "what is  Andrew's kernel?",
      "votes": null
    },
    {
      "id": "522010",
      "postDate": "04/23/2019 18:48:14",
      "content": "<p>Sorry, it's this one:\n<a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a></p>\n\n<p>It has a comment: \"It turned out that stacking is much worse than blending on LB.\"</p>",
      "rawMarkdown": "Sorry, it's this one:\nhttps://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\n\nIt has a comment: \"It turned out that stacking is much worse than blending on LB.\"",
      "votes": null
    },
    {
      "id": "522058",
      "postDate": "04/23/2019 20:07:56",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "522113",
      "postDate": "04/23/2019 21:52:16",
      "content": "<p>I'd start with weighted mean of each model predictions.  Just make sure sum of the weights is 1.</p>",
      "rawMarkdown": "I'd start with weighted mean of each model predictions.  Just make sure sum of the weights is 1.",
      "votes": null
    },
    {
      "id": "522133",
      "postDate": "04/23/2019 22:34:02",
      "content": "<p>Would you learn the weights from the CV score or a mixture of CV and LB scores?</p>",
      "rawMarkdown": "Would you learn the weights from the CV score or a mixture of CV and LB scores?",
      "votes": null
    },
    {
      "id": "522138",
      "postDate": "04/23/2019 22:44:41",
      "content": "<blockquote>\n  <p>Would you learn the weights from the CV score or a mixture of CV and LB scores?</p>\n</blockquote>\n\n<p>A proper way to learn weights is from out-of-fold files - see an example <a href=\"https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors\"><strong>here</strong></a> of doing it by brute-force minimization. If you do it from CV and/or LB scores, that would be more in the domain of estimating, eyeballing or guessing.</p>\n\n<p>Since I have been on a self-promotion tour in this thread, here is <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224\"><strong>an argument</strong></a> why everyone should be making out-of-fold predictions along with their test predictions.</p>",
      "rawMarkdown": "&gt; Would you learn the weights from the CV score or a mixture of CV and LB scores?\n\nA proper way to learn weights is from out-of-fold files - see an example [__here__](https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors) of doing it by brute-force minimization. If you do it from CV and/or LB scores, that would be more in the domain of estimating, eyeballing or guessing.\n\nSince I have been on a self-promotion tour in this thread, here is [__an argument__](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224) why everyone should be making out-of-fold predictions along with their test predictions.",
      "votes": null
    },
    {
      "id": "522146",
      "postDate": "04/23/2019 22:56:30",
      "content": "<p>Thanks a lot <a href=\"/tilii7\">@tilii7</a>  that's exactly what I am looking for.  Are you joining this competition at some point?</p>",
      "rawMarkdown": "Thanks a lot @tilii7  that's exactly what I am looking for.  Are you joining this competition at some point?",
      "votes": null
    },
    {
      "id": "522150",
      "postDate": "04/23/2019 23:06:35",
      "content": "<p>You are welcome. The plan is to join the competition if I figure out a reliable validation scheme.</p>",
      "rawMarkdown": "You are welcome. The plan is to join the competition if I figure out a reliable validation scheme.",
      "votes": null
    },
    {
      "id": "522365",
      "postDate": "04/24/2019 10:16:57",
      "content": "<blockquote>\n  <p>if I figure out a reliable validation scheme.</p>\n</blockquote>\n\n<p>You're setting up a very high bar I'm afraid ;)</p>",
      "rawMarkdown": "&gt; if I figure out a reliable validation scheme.\n\nYou're setting up a very high bar I'm afraid ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 521986,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/23/2019 18:10:30",
      "content": "<p>There is no methodological difference between ensembling for classification and regression. If using neural networks for ensembling, be sure to have a sigmoid output activation at the end for classification so your predictions are bound by 0-1. Leave it linear for regression.</p>\n\n<p>As to the usefulness of weak models, see if <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51058\"><strong>my earlier post</strong></a> answers your question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 521993,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/23/2019 18:29:56",
          "content": "<p>Great post indeed, very nice to read. \nI think when it comes to stacking, regression is no different from classification, but what about blending?\nThis competition, in my opinion, is tricky and dangerous when it comes to stacking, especially depending on the CV strategy</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522008,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "04/23/2019 18:43:56",
      "content": "<p>what is  Andrew's kernel? </p>",
      "votes": null,
      "replies": [
        {
          "id": 522010,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/23/2019 18:48:14",
          "content": "<p>Sorry, it's this one:\n<a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a></p>\n\n<p>It has a comment: \"It turned out that stacking is much worse than blending on LB.\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522058,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "04/23/2019 20:07:56",
          "content": "<p>thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522113,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/23/2019 21:52:16",
      "content": "<p>I'd start with weighted mean of each model predictions.  Just make sure sum of the weights is 1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 522133,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/23/2019 22:34:02",
          "content": "<p>Would you learn the weights from the CV score or a mixture of CV and LB scores?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522138,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/23/2019 22:44:41",
          "content": "<blockquote>\n  <p>Would you learn the weights from the CV score or a mixture of CV and LB scores?</p>\n</blockquote>\n\n<p>A proper way to learn weights is from out-of-fold files - see an example <a href=\"https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors\"><strong>here</strong></a> of doing it by brute-force minimization. If you do it from CV and/or LB scores, that would be more in the domain of estimating, eyeballing or guessing.</p>\n\n<p>Since I have been on a self-promotion tour in this thread, here is <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224\"><strong>an argument</strong></a> why everyone should be making out-of-fold predictions along with their test predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522146,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/23/2019 22:56:30",
          "content": "<p>Thanks a lot <a href=\"/tilii7\">@tilii7</a>  that's exactly what I am looking for.  Are you joining this competition at some point?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522150,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/23/2019 23:06:35",
          "content": "<p>You are welcome. The plan is to join the competition if I figure out a reliable validation scheme.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522365,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/24/2019 10:16:57",
          "content": "<blockquote>\n  <p>if I figure out a reliable validation scheme.</p>\n</blockquote>\n\n<p>You're setting up a very high bar I'm afraid ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "521975": "I've been looking for strategies in ensembling and found this great summary:\nhttps://mlwave.com/kaggle-ensembling-guide/?lipi=urn%3Ali%3Apage%3Ad_flagship3_pulse_read%3BPZ4T3JLHTu%2BOWNI0d5kFbg%3D%3D\n\nHowever, it seems most if not all the points mentioned are about classification not regression.\n\nWhat is your advice on ensembling in such competition, where stacking seems to be a bad option (see Andrew's kernel), possibly due to leakage?\nHow would you optimize the weights when blending?  \nAre weak models useful at all when blending?\n\nPlease share your insights on the best practices for combining regression models.",
    "521986": "There is no methodological difference between ensembling for classification and regression. If using neural networks for ensembling, be sure to have a sigmoid output activation at the end for classification so your predictions are bound by 0-1. Leave it linear for regression.\n\nAs to the usefulness of weak models, see if [__my earlier post__](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/51058) answers your question.",
    "521993": "Great post indeed, very nice to read. \nI think when it comes to stacking, regression is no different from classification, but what about blending?\nThis competition, in my opinion, is tricky and dangerous when it comes to stacking, especially depending on the CV strategy",
    "522008": "what is  Andrew's kernel?",
    "522010": "Sorry, it's this one:\nhttps://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\n\nIt has a comment: \"It turned out that stacking is much worse than blending on LB.\"",
    "522058": "thanks!",
    "522113": "I'd start with weighted mean of each model predictions.  Just make sure sum of the weights is 1.",
    "522133": "Would you learn the weights from the CV score or a mixture of CV and LB scores?",
    "522138": "&gt; Would you learn the weights from the CV score or a mixture of CV and LB scores?\n\nA proper way to learn weights is from out-of-fold files - see an example [__here__](https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors) of doing it by brute-force minimization. If you do it from CV and/or LB scores, that would be more in the domain of estimating, eyeballing or guessing.\n\nSince I have been on a self-promotion tour in this thread, here is [__an argument__](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52224) why everyone should be making out-of-fold predictions along with their test predictions.",
    "522146": "Thanks a lot @tilii7  that's exactly what I am looking for.  Are you joining this competition at some point?",
    "522150": "You are welcome. The plan is to join the competition if I figure out a reliable validation scheme.",
    "522365": "&gt; if I figure out a reliable validation scheme.\n\nYou're setting up a very high bar I'm afraid ;)"
  },
  "source": "meta"
}