{
  "id": 55932,
  "title": "Is a Single-Model more reliable to evade shake-up? ",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55932",
  "author_name": "",
  "post_date": "2018-05-03T09:00:40.323999900Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi guys, \nWas wondering if single models are better than ensembles ? Although ensembles are kind of stable if weights are given to different types of models! </p>\n\n<p>18% - Public \n82% - Private</p>\n\n<p>Safer side is to have one submission with your single model and one submission with your ensembles. Please Share you thoughts :) </p>",
  "messages": [
    {
      "id": "322598",
      "postDate": "05/03/2018 09:00:40",
      "content": "<p>Hi guys, \nWas wondering if single models are better than ensembles ? Although ensembles are kind of stable if weights are given to different types of models! </p>\n\n<p>18% - Public \n82% - Private</p>\n\n<p>Safer side is to have one submission with your single model and one submission with your ensembles. Please Share you thoughts :) </p>",
      "rawMarkdown": "Hi guys, \nWas wondering if single models are better than ensembles ? Although ensembles are kind of stable if weights are given to different types of models! \n\n18% - Public \n82% - Private\n\nSafer side is to have one submission with your single model and one submission with your ensembles. Please Share you thoughts :)",
      "votes": null
    },
    {
      "id": "322614",
      "postDate": "05/03/2018 09:38:13",
      "content": "<p>Yesterday I asked a fairly similar question: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55821\">Choosing entries for final submission</a>. I recommend reading the answers that have been given there.</p>",
      "rawMarkdown": "Yesterday I asked a fairly similar question: [Choosing entries for final submission][1]. I recommend reading the answers that have been given there.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55821",
      "votes": null
    },
    {
      "id": "322651",
      "postDate": "05/03/2018 11:31:44",
      "content": "<p>Hey thanks a lot.. will look through :) </p>",
      "rawMarkdown": "Hey thanks a lot.. will look through :)",
      "votes": null
    },
    {
      "id": "322738",
      "postDate": "05/03/2018 14:38:42",
      "content": "<p>No, not necessarily. It very much depends on how you form/validate your ensemble. </p>\n\n<p>As you elude to, a well-constructed ensemble can yield a significant reduction in both bias and variance vs. a single model solution, and would be more likely to generalize well to the test data. It's important to not conflate the conceptual complexity of ensembles with model variance - the two can often be inversely related. So if you construct an ensemble with rigorous methodology (e.g. stacking on out of fold predictions on a representative dataset, and validating the results), you can often expect the ensemble to be \"safer\" than a single model.   </p>\n\n<p>On the other hand, if you blindly ensemble models that score well on the public leaderboard, you may end up with an overfit mess. Remember that the private test hours may exhibit very different behavior from the public hour.</p>",
      "rawMarkdown": "No, not necessarily. It very much depends on how you form/validate your ensemble. \n\nAs you elude to, a well-constructed ensemble can yield a significant reduction in both bias and variance vs. a single model solution, and would be more likely to generalize well to the test data. It's important to not conflate the conceptual complexity of ensembles with model variance - the two can often be inversely related. So if you construct an ensemble with rigorous methodology (e.g. stacking on out of fold predictions on a representative dataset, and validating the results), you can often expect the ensemble to be \"safer\" than a single model.   \n\nOn the other hand, if you blindly ensemble models that score well on the public leaderboard, you may end up with an overfit mess. Remember that the private test hours may exhibit very different behavior from the public hour.",
      "votes": null
    },
    {
      "id": "322754",
      "postDate": "05/03/2018 15:42:27",
      "content": "<p>@Joe Thanks for your thoughts. I'm primarily focusing on improving my single model by adding some features ! </p>",
      "rawMarkdown": "Joe Thanks for your thoughts. I'm primarily focusing on improving my single model by adding some features !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 322614,
      "author_name": "skalskip",
      "author_url": "",
      "post_date": "05/03/2018 09:38:13",
      "content": "<p>Yesterday I asked a fairly similar question: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55821\">Choosing entries for final submission</a>. I recommend reading the answers that have been given there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 322651,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "05/03/2018 11:31:44",
      "content": "<p>Hey thanks a lot.. will look through :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 322738,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/03/2018 14:38:42",
      "content": "<p>No, not necessarily. It very much depends on how you form/validate your ensemble. </p>\n\n<p>As you elude to, a well-constructed ensemble can yield a significant reduction in both bias and variance vs. a single model solution, and would be more likely to generalize well to the test data. It's important to not conflate the conceptual complexity of ensembles with model variance - the two can often be inversely related. So if you construct an ensemble with rigorous methodology (e.g. stacking on out of fold predictions on a representative dataset, and validating the results), you can often expect the ensemble to be \"safer\" than a single model.   </p>\n\n<p>On the other hand, if you blindly ensemble models that score well on the public leaderboard, you may end up with an overfit mess. Remember that the private test hours may exhibit very different behavior from the public hour.</p>",
      "votes": null,
      "replies": [
        {
          "id": 322754,
          "author_name": "prashanththangavel",
          "author_url": "",
          "post_date": "05/03/2018 15:42:27",
          "content": "<p>@Joe Thanks for your thoughts. I'm primarily focusing on improving my single model by adding some features ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "322598": "Hi guys, \nWas wondering if single models are better than ensembles ? Although ensembles are kind of stable if weights are given to different types of models! \n\n18% - Public \n82% - Private\n\nSafer side is to have one submission with your single model and one submission with your ensembles. Please Share you thoughts :)",
    "322614": "Yesterday I asked a fairly similar question: [Choosing entries for final submission][1]. I recommend reading the answers that have been given there.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55821",
    "322651": "Hey thanks a lot.. will look through :)",
    "322738": "No, not necessarily. It very much depends on how you form/validate your ensemble. \n\nAs you elude to, a well-constructed ensemble can yield a significant reduction in both bias and variance vs. a single model solution, and would be more likely to generalize well to the test data. It's important to not conflate the conceptual complexity of ensembles with model variance - the two can often be inversely related. So if you construct an ensemble with rigorous methodology (e.g. stacking on out of fold predictions on a representative dataset, and validating the results), you can often expect the ensemble to be \"safer\" than a single model.   \n\nOn the other hand, if you blindly ensemble models that score well on the public leaderboard, you may end up with an overfit mess. Remember that the private test hours may exhibit very different behavior from the public hour.",
    "322754": "Joe Thanks for your thoughts. I'm primarily focusing on improving my single model by adding some features !"
  },
  "source": "meta"
}