{
  "id": 185501,
  "title": "Some simple techniques for stacking models",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/185501",
  "author_name": "",
  "post_date": "2020-09-21T05:28:19.372162200Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Now that we're in the \"endgame\" of this competition now, people will start to ensemble/stack their models. Just pure blending might not go a long way, but perhaps these few methods can.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369\" target=\"_blank\">https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369</a> scipy.optimize's lsq_linear</li>\n<li><a href=\"https://github.com/h2oai/pystacknet\" target=\"_blank\">https://github.com/h2oai/pystacknet</a> PyStackNet, a Python implementation of GM Kazanova's library <code>StackNet</code>.</li>\n<li><a href=\"https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw\" target=\"_blank\">https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw</a> Again, this might be unstable but it might work if you're lucky and have thoroughly checked the coefficient.</li>\n</ul>\n<p><strong>Let me know if you have anything else to share pertaining this topic!</strong></p>\n<p>Edit: pertaining to the suggestion of <a href=\"https://www.kaggle.com/ajaychaudhari\" target=\"_blank\">@ajaychaudhari</a> , stacking won't ideally take much time depending on what type of model you use - in the public kernels stacking doesn't take much time AFAIK.</p>\n<p>However if you are using heavier models which take &gt; 2.5 hrs to run, then stacking might not be ideal. </p>",
  "messages": [
    {
      "id": "1020337",
      "postDate": "09/21/2020 05:28:19",
      "content": "<p>Now that we're in the \"endgame\" of this competition now, people will start to ensemble/stack their models. Just pure blending might not go a long way, but perhaps these few methods can.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369\" target=\"_blank\">https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369</a> scipy.optimize's lsq_linear</li>\n<li><a href=\"https://github.com/h2oai/pystacknet\" target=\"_blank\">https://github.com/h2oai/pystacknet</a> PyStackNet, a Python implementation of GM Kazanova's library <code>StackNet</code>.</li>\n<li><a href=\"https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw\" target=\"_blank\">https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw</a> Again, this might be unstable but it might work if you're lucky and have thoroughly checked the coefficient.</li>\n</ul>\n<p><strong>Let me know if you have anything else to share pertaining this topic!</strong></p>\n<p>Edit: pertaining to the suggestion of <a href=\"https://www.kaggle.com/ajaychaudhari\" target=\"_blank\">@ajaychaudhari</a> , stacking won't ideally take much time depending on what type of model you use - in the public kernels stacking doesn't take much time AFAIK.</p>\n<p>However if you are using heavier models which take &gt; 2.5 hrs to run, then stacking might not be ideal. </p>",
      "rawMarkdown": "Now that we're in the \"endgame\" of this competition now, people will start to ensemble/stack their models. Just pure blending might not go a long way, but perhaps these few methods can.\n\n+ https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369 scipy.optimize's lsq_linear\n+ https://github.com/h2oai/pystacknet PyStackNet, a Python implementation of GM Kazanova's library `StackNet`.\n+ https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw Again, this might be unstable but it might work if you're lucky and have thoroughly checked the coefficient.\n\n**Let me know if you have anything else to share pertaining this topic!**\n\nEdit: pertaining to the suggestion of @ajaychaudhari , stacking won't ideally take much time depending on what type of model you use - in the public kernels stacking doesn't take much time AFAIK.\n\nHowever if you are using heavier models which take > 2.5 hrs to run, then stacking might not be ideal.",
      "votes": null
    },
    {
      "id": "1020344",
      "postDate": "09/21/2020 05:38:09",
      "content": "<p>I tried out stacking but notebook ran out of time. Any results you want to share?</p>",
      "rawMarkdown": "I tried out stacking but notebook ran out of time. Any results you want to share?",
      "votes": null
    },
    {
      "id": "1021415",
      "postDate": "09/21/2020 20:16:25",
      "content": "<p>Thank you for sharing.<br>\nIn my previous competition, I used <a href=\"https://www.kaggle.com/mekhdigakhramanian/post-processing-v2\" target=\"_blank\">https://www.kaggle.com/mekhdigakhramanian/post-processing-v2</a> with public submission outputs and 3 my own kernels but it was overfitted,  be careful</p>",
      "rawMarkdown": "Thank you for sharing.\nIn my previous competition, I used https://www.kaggle.com/mekhdigakhramanian/post-processing-v2 with public submission outputs and 3 my own kernels but it was overfitted,  be careful",
      "votes": null
    },
    {
      "id": "1022832",
      "postDate": "09/22/2020 18:48:47",
      "content": "<p>Similar to what you have mentioned in point-1, I tried blending models by finding optimal weights using OOF predictions. Blending gives a slight boost to the CV. </p>\n<p>StackNet might be tricky to implement here given that most of the conventional models are not doing well in this problem and its also going to be a challenge to find a lot of uncorrelated models. Any thoughts on this?</p>",
      "rawMarkdown": "Similar to what you have mentioned in point-1, I tried blending models by finding optimal weights using OOF predictions. Blending gives a slight boost to the CV. \n\nStackNet might be tricky to implement here given that most of the conventional models are not doing well in this problem and its also going to be a challenge to find a lot of uncorrelated models. Any thoughts on this?",
      "votes": null
    },
    {
      "id": "1023180",
      "postDate": "09/23/2020 02:24:04",
      "content": "<blockquote>\n  <p>most of the conventional models are not doing well in this problem</p>\n</blockquote>\n<p>Since most people are using NNs or a variant, I doubt they'd be much interested in using StackNet. LightGBM in a public kernel was pretty decent, but haven't seen much of LGB after that.</p>\n<p>Uncorrelated models too might be a bit of a challenge, so you'll need to run a lot of (computationally expensive) experiments in order to get the best possible models from your experiments.</p>",
      "rawMarkdown": "> most of the conventional models are not doing well in this problem\n\nSince most people are using NNs or a variant, I doubt they'd be much interested in using StackNet. LightGBM in a public kernel was pretty decent, but haven't seen much of LGB after that.\n\nUncorrelated models too might be a bit of a challenge, so you'll need to run a lot of (computationally expensive) experiments in order to get the best possible models from your experiments.",
      "votes": null
    },
    {
      "id": "1025196",
      "postDate": "09/24/2020 11:57:24",
      "content": "<p>informative , Thanks for sharing <a href=\"/nxrprime\">@nxrprime</a> </p>",
      "rawMarkdown": "informative , Thanks for sharing @nxrprime",
      "votes": null
    },
    {
      "id": "1026959",
      "postDate": "09/25/2020 17:28:00",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> </p>",
      "rawMarkdown": "Thanks for sharing @nxrprime",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1020344,
      "author_name": "ajay19",
      "author_url": "",
      "post_date": "09/21/2020 05:38:09",
      "content": "<p>I tried out stacking but notebook ran out of time. Any results you want to share?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1021415,
      "author_name": "",
      "author_url": "",
      "post_date": "09/21/2020 20:16:25",
      "content": "<p>Thank you for sharing.<br>\nIn my previous competition, I used <a href=\"https://www.kaggle.com/mekhdigakhramanian/post-processing-v2\" target=\"_blank\">https://www.kaggle.com/mekhdigakhramanian/post-processing-v2</a> with public submission outputs and 3 my own kernels but it was overfitted,  be careful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1022832,
      "author_name": "abhishekgbhat",
      "author_url": "",
      "post_date": "09/22/2020 18:48:47",
      "content": "<p>Similar to what you have mentioned in point-1, I tried blending models by finding optimal weights using OOF predictions. Blending gives a slight boost to the CV. </p>\n<p>StackNet might be tricky to implement here given that most of the conventional models are not doing well in this problem and its also going to be a challenge to find a lot of uncorrelated models. Any thoughts on this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1023180,
          "author_name": "nxrprime",
          "author_url": "",
          "post_date": "09/23/2020 02:24:04",
          "content": "<blockquote>\n  <p>most of the conventional models are not doing well in this problem</p>\n</blockquote>\n<p>Since most people are using NNs or a variant, I doubt they'd be much interested in using StackNet. LightGBM in a public kernel was pretty decent, but haven't seen much of LGB after that.</p>\n<p>Uncorrelated models too might be a bit of a challenge, so you'll need to run a lot of (computationally expensive) experiments in order to get the best possible models from your experiments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1026959,
      "author_name": "vineeth1999",
      "author_url": "",
      "post_date": "09/25/2020 17:28:00",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1025196,
      "author_name": "pinakimishrads",
      "author_url": "",
      "post_date": "09/24/2020 11:57:24",
      "content": "<p>informative , Thanks for sharing <a href=\"/nxrprime\">@nxrprime</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1020337": "Now that we're in the \"endgame\" of this competition now, people will start to ensemble/stack their models. Just pure blending might not go a long way, but perhaps these few methods can.\n\n+ https://www.kaggle.com/c/ashrae-energy-prediction/discussion/122369 scipy.optimize's lsq_linear\n+ https://github.com/h2oai/pystacknet PyStackNet, a Python implementation of GM Kazanova's library `StackNet`.\n+ https://www.kaggle.com/khoongweihao/post-processing-technique-c-f-1st-place-jigsaw Again, this might be unstable but it might work if you're lucky and have thoroughly checked the coefficient.\n\n**Let me know if you have anything else to share pertaining this topic!**\n\nEdit: pertaining to the suggestion of @ajaychaudhari , stacking won't ideally take much time depending on what type of model you use - in the public kernels stacking doesn't take much time AFAIK.\n\nHowever if you are using heavier models which take > 2.5 hrs to run, then stacking might not be ideal.",
    "1020344": "I tried out stacking but notebook ran out of time. Any results you want to share?",
    "1021415": "Thank you for sharing.\nIn my previous competition, I used https://www.kaggle.com/mekhdigakhramanian/post-processing-v2 with public submission outputs and 3 my own kernels but it was overfitted,  be careful",
    "1022832": "Similar to what you have mentioned in point-1, I tried blending models by finding optimal weights using OOF predictions. Blending gives a slight boost to the CV. \n\nStackNet might be tricky to implement here given that most of the conventional models are not doing well in this problem and its also going to be a challenge to find a lot of uncorrelated models. Any thoughts on this?",
    "1023180": "> most of the conventional models are not doing well in this problem\n\nSince most people are using NNs or a variant, I doubt they'd be much interested in using StackNet. LightGBM in a public kernel was pretty decent, but haven't seen much of LGB after that.\n\nUncorrelated models too might be a bit of a challenge, so you'll need to run a lot of (computationally expensive) experiments in order to get the best possible models from your experiments.",
    "1025196": "informative , Thanks for sharing @nxrprime",
    "1026959": "Thanks for sharing @nxrprime"
  },
  "source": "meta"
}