{
  "id": 318004,
  "title": "How to do ensembling ?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/318004",
  "author_name": "",
  "post_date": "2022-04-10T05:50:10.152744700Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can anyone explain good ensembling methods to improve my LB?</p>",
  "messages": [
    {
      "id": "1750799",
      "postDate": "04/10/2022 05:50:10",
      "content": "<p>Can anyone explain good ensembling methods to improve my LB?</p>",
      "rawMarkdown": "Can anyone explain good ensembling methods to improve my LB?",
      "votes": null
    },
    {
      "id": "1750827",
      "postDate": "04/10/2022 06:17:18",
      "content": "<p>You may try the below techniques- </p>\n<ol>\n<li>RandomForestClassifier- This is available in the package sklearn.ensemble. </li>\n<li>XGBClassifier- This is available in the package xgboost</li>\n<li>LGBMClassifier- This is a fast execution based boosted tree method available in the lgbm package</li>\n<li>CatBoostClassifier- This is available in the catboost package</li>\n</ol>\n<p>You may consider parameter tuning using a grid search and cross validation, available in sklearn.model_selection package. GridSearchCV is a commonly used mechanism for this purpose. You may consider a cross-validation strategy like shuffle= True, number of splits, etc. to help you with your tasks. </p>\n<p>You may refer to the sklearn documentation for more details in this regard. The associated link is as below- <br>\n<a href=\"https://scikit-learn.org/stable/\" target=\"_blank\">https://scikit-learn.org/stable/</a></p>",
      "rawMarkdown": "You may try the below techniques- \n1. RandomForestClassifier- This is available in the package sklearn.ensemble. \n2. XGBClassifier- This is available in the package xgboost\n3. LGBMClassifier- This is a fast execution based boosted tree method available in the lgbm package\n4. CatBoostClassifier- This is available in the catboost package\n\nYou may consider parameter tuning using a grid search and cross validation, available in sklearn.model_selection package. GridSearchCV is a commonly used mechanism for this purpose. You may consider a cross-validation strategy like shuffle= True, number of splits, etc. to help you with your tasks. \n\nYou may refer to the sklearn documentation for more details in this regard. The associated link is as below- \nhttps://scikit-learn.org/stable/",
      "votes": null
    },
    {
      "id": "1750880",
      "postDate": "04/10/2022 07:14:52",
      "content": "<p>I want to do ensembling of the embeddings which I get as outcome from my happywhale trained model.<br>\nSorry as I forgot to mention it.</p>",
      "rawMarkdown": "I want to do ensembling of the embeddings which I get as outcome from my happywhale trained model.\nSorry as I forgot to mention it.",
      "votes": null
    },
    {
      "id": "1751945",
      "postDate": "04/11/2022 08:31:32",
      "content": "<p>Check it. Hope it works for you!<br>\n<a href=\"https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\" target=\"_blank\">https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels</a><br>\n<a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458\" target=\"_blank\">https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458</a></p>",
      "rawMarkdown": "Check it. Hope it works for you!\nhttps://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\nhttps://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458",
      "votes": null
    },
    {
      "id": "1752083",
      "postDate": "04/11/2022 11:38:43",
      "content": "<p>You can concatenate the embeddings to each other and then run nearest neighbors on it. </p>",
      "rawMarkdown": "You can concatenate the embeddings to each other and then run nearest neighbors on it.",
      "votes": null
    },
    {
      "id": "1752264",
      "postDate": "04/11/2022 15:12:27",
      "content": "<p>Is it possible to get a good score from ensembling here?<br>\nI got cv of 0.81 on validation dataset but in public submission I got only score of 72.<br>\nHow much 'new_individuals' should I consider. At this time I am assuming it to be 10%</p>",
      "rawMarkdown": "Is it possible to get a good score from ensembling here?\nI got cv of 0.81 on validation dataset but in public submission I got only score of 72.\nHow much 'new_individuals' should I consider. At this time I am assuming it to be 10%",
      "votes": null
    },
    {
      "id": "1752266",
      "postDate": "04/11/2022 15:15:19",
      "content": "<p>Trust your cv.</p>",
      "rawMarkdown": "Trust your cv.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1750827,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/10/2022 06:17:18",
      "content": "<p>You may try the below techniques- </p>\n<ol>\n<li>RandomForestClassifier- This is available in the package sklearn.ensemble. </li>\n<li>XGBClassifier- This is available in the package xgboost</li>\n<li>LGBMClassifier- This is a fast execution based boosted tree method available in the lgbm package</li>\n<li>CatBoostClassifier- This is available in the catboost package</li>\n</ol>\n<p>You may consider parameter tuning using a grid search and cross validation, available in sklearn.model_selection package. GridSearchCV is a commonly used mechanism for this purpose. You may consider a cross-validation strategy like shuffle= True, number of splits, etc. to help you with your tasks. </p>\n<p>You may refer to the sklearn documentation for more details in this regard. The associated link is as below- <br>\n<a href=\"https://scikit-learn.org/stable/\" target=\"_blank\">https://scikit-learn.org/stable/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1750880,
      "author_name": "jainishsavalia",
      "author_url": "",
      "post_date": "04/10/2022 07:14:52",
      "content": "<p>I want to do ensembling of the embeddings which I get as outcome from my happywhale trained model.<br>\nSorry as I forgot to mention it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1752083,
          "author_name": "vexxingbanana",
          "author_url": "",
          "post_date": "04/11/2022 11:38:43",
          "content": "<p>You can concatenate the embeddings to each other and then run nearest neighbors on it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1752264,
          "author_name": "jainishsavalia",
          "author_url": "",
          "post_date": "04/11/2022 15:12:27",
          "content": "<p>Is it possible to get a good score from ensembling here?<br>\nI got cv of 0.81 on validation dataset but in public submission I got only score of 72.<br>\nHow much 'new_individuals' should I consider. At this time I am assuming it to be 10%</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1752266,
          "author_name": "yangranran",
          "author_url": "",
          "post_date": "04/11/2022 15:15:19",
          "content": "<p>Trust your cv.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1751945,
      "author_name": "yangranran",
      "author_url": "",
      "post_date": "04/11/2022 08:31:32",
      "content": "<p>Check it. Hope it works for you!<br>\n<a href=\"https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\" target=\"_blank\">https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels</a><br>\n<a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458\" target=\"_blank\">https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1750799": "Can anyone explain good ensembling methods to improve my LB?",
    "1750827": "You may try the below techniques- \n1. RandomForestClassifier- This is available in the package sklearn.ensemble. \n2. XGBClassifier- This is available in the package xgboost\n3. LGBMClassifier- This is a fast execution based boosted tree method available in the lgbm package\n4. CatBoostClassifier- This is available in the catboost package\n\nYou may consider parameter tuning using a grid search and cross validation, available in sklearn.model_selection package. GridSearchCV is a commonly used mechanism for this purpose. You may consider a cross-validation strategy like shuffle= True, number of splits, etc. to help you with your tasks. \n\nYou may refer to the sklearn documentation for more details in this regard. The associated link is as below- \nhttps://scikit-learn.org/stable/",
    "1750880": "I want to do ensembling of the embeddings which I get as outcome from my happywhale trained model.\nSorry as I forgot to mention it.",
    "1751945": "Check it. Hope it works for you!\nhttps://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\nhttps://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/313458",
    "1752083": "You can concatenate the embeddings to each other and then run nearest neighbors on it.",
    "1752264": "Is it possible to get a good score from ensembling here?\nI got cv of 0.81 on validation dataset but in public submission I got only score of 72.\nHow much 'new_individuals' should I consider. At this time I am assuming it to be 10%",
    "1752266": "Trust your cv."
  },
  "source": "meta"
}