{
  "id": 546385,
  "title": "What's the reason to create exactly 3 submissions/models and use majority voting",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546385",
  "author_name": "",
  "post_date": "2024-11-15T12:38:50.006353Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>I noticed that alot of the public notebooks use 3 submissions/models and then use majority voting on their final submission. <br>\nWhat's the reason for this method and why are the models made the way they are? <br>\nI'm not looking for \"because it gives the best score\". I assume alot is copying what works, ofc. But I would like to learn if that is a normal approach to such a problem.</p>\n<p>Most of the time I see a votingregressor model of lgbm, xgb and catboost, which makes sense since all of them are strong boosting regressors. And another common model is the combination of the previous model and randomforestregressor+gradientboostingregressor. The last one is either another votingregressor or tabnet model with different hyperparameters for lgbm, xgb and catboost.</p>\n<p>Why are the models chosen this way? Only the first one seems obvious to me.</p>\n<p>How do you choose the different hyperparamers for the models? Especially if they use the same method (optimization would lead to the same hyperparameters).</p>\n<p>I would appreciate any insights. Thanks!</p>",
  "messages": [
    {
      "id": "3046379",
      "postDate": "11/15/2024 12:38:50",
      "content": "<p>Hello everyone!</p>\n<p>I noticed that alot of the public notebooks use 3 submissions/models and then use majority voting on their final submission. <br>\nWhat's the reason for this method and why are the models made the way they are? <br>\nI'm not looking for \"because it gives the best score\". I assume alot is copying what works, ofc. But I would like to learn if that is a normal approach to such a problem.</p>\n<p>Most of the time I see a votingregressor model of lgbm, xgb and catboost, which makes sense since all of them are strong boosting regressors. And another common model is the combination of the previous model and randomforestregressor+gradientboostingregressor. The last one is either another votingregressor or tabnet model with different hyperparameters for lgbm, xgb and catboost.</p>\n<p>Why are the models chosen this way? Only the first one seems obvious to me.</p>\n<p>How do you choose the different hyperparamers for the models? Especially if they use the same method (optimization would lead to the same hyperparameters).</p>\n<p>I would appreciate any insights. Thanks!</p>",
      "rawMarkdown": "Hello everyone!\n\nI noticed that alot of the public notebooks use 3 submissions/models and then use majority voting on their final submission. \nWhat's the reason for this method and why are the models made the way they are? \nI'm not looking for \"because it gives the best score\". I assume alot is copying what works, ofc. But I would like to learn if that is a normal approach to such a problem.\n\nMost of the time I see a votingregressor model of lgbm, xgb and catboost, which makes sense since all of them are strong boosting regressors. And another common model is the combination of the previous model and randomforestregressor+gradientboostingregressor. The last one is either another votingregressor or tabnet model with different hyperparameters for lgbm, xgb and catboost.\n\nWhy are the models chosen this way? Only the first one seems obvious to me.\n\nHow do you choose the different hyperparamers for the models? Especially if they use the same method (optimization would lead to the same hyperparameters).\n\nI would appreciate any insights. Thanks!",
      "votes": null
    },
    {
      "id": "3046769",
      "postDate": "11/15/2024 21:14:11",
      "content": "<p>Like you mentioned - copy/paste is the most of the public works and the second part is that this dataset is sensible to random_seed so they are looking to ensemble different models even if it's the same type of it and get a higher score. So they are playing with these seed ensembles and publish them and again a lot of copy / paste so the leaderboard is overflooded with overfitted models. Private LB will shake them up.</p>",
      "rawMarkdown": "Like you mentioned - copy/paste is the most of the public works and the second part is that this dataset is sensible to random_seed so they are looking to ensemble different models even if it's the same type of it and get a higher score. So they are playing with these seed ensembles and publish them and again a lot of copy / paste so the leaderboard is overflooded with overfitted models. Private LB will shake them up.",
      "votes": null
    },
    {
      "id": "3046973",
      "postDate": "11/16/2024 05:29:56",
      "content": "<p>Agreed <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> <br>\nThe data is very small and complex models are unlikely to work here - most public works are using more than necessary complex models and this will surely not sustain </p>",
      "rawMarkdown": "Agreed @eu1234 \nThe data is very small and complex models are unlikely to work here - most public works are using more than necessary complex models and this will surely not sustain",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3046769,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/15/2024 21:14:11",
      "content": "<p>Like you mentioned - copy/paste is the most of the public works and the second part is that this dataset is sensible to random_seed so they are looking to ensemble different models even if it's the same type of it and get a higher score. So they are playing with these seed ensembles and publish them and again a lot of copy / paste so the leaderboard is overflooded with overfitted models. Private LB will shake them up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3046973,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/16/2024 05:29:56",
          "content": "<p>Agreed <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> <br>\nThe data is very small and complex models are unlikely to work here - most public works are using more than necessary complex models and this will surely not sustain </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3046379": "Hello everyone!\n\nI noticed that alot of the public notebooks use 3 submissions/models and then use majority voting on their final submission. \nWhat's the reason for this method and why are the models made the way they are? \nI'm not looking for \"because it gives the best score\". I assume alot is copying what works, ofc. But I would like to learn if that is a normal approach to such a problem.\n\nMost of the time I see a votingregressor model of lgbm, xgb and catboost, which makes sense since all of them are strong boosting regressors. And another common model is the combination of the previous model and randomforestregressor+gradientboostingregressor. The last one is either another votingregressor or tabnet model with different hyperparameters for lgbm, xgb and catboost.\n\nWhy are the models chosen this way? Only the first one seems obvious to me.\n\nHow do you choose the different hyperparamers for the models? Especially if they use the same method (optimization would lead to the same hyperparameters).\n\nI would appreciate any insights. Thanks!",
    "3046769": "Like you mentioned - copy/paste is the most of the public works and the second part is that this dataset is sensible to random_seed so they are looking to ensemble different models even if it's the same type of it and get a higher score. So they are playing with these seed ensembles and publish them and again a lot of copy / paste so the leaderboard is overflooded with overfitted models. Private LB will shake them up.",
    "3046973": "Agreed @eu1234 \nThe data is very small and complex models are unlikely to work here - most public works are using more than necessary complex models and this will surely not sustain"
  },
  "source": "meta"
}