{
  "id": 336229,
  "title": "How to make the model better?",
  "url": "/competitions/amex-default-prediction/discussion/336229",
  "author_name": "",
  "post_date": "2022-07-10T00:46:36.402466600Z",
  "votes": 13,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I can score up to 0.799 when using LGBM alone, but when I use XGBoost + LGBM + CatBoost for model fusion, I find that the scores run out lower than any of the individual runs.<br>\nThis is really puzzling to me. Does it feel like it could be because of the amount of features?</p>",
  "messages": [
    {
      "id": "1849957",
      "postDate": "07/10/2022 00:46:36",
      "content": "<p>I can score up to 0.799 when using LGBM alone, but when I use XGBoost + LGBM + CatBoost for model fusion, I find that the scores run out lower than any of the individual runs.<br>\nThis is really puzzling to me. Does it feel like it could be because of the amount of features?</p>",
      "rawMarkdown": "I can score up to 0.799 when using LGBM alone, but when I use XGBoost + LGBM + CatBoost for model fusion, I find that the scores run out lower than any of the individual runs.\nThis is really puzzling to me. Does it feel like it could be because of the amount of features?",
      "votes": null
    },
    {
      "id": "1850197",
      "postDate": "07/10/2022 07:32:07",
      "content": "<p>My suggestion is to only focus on one model, improve the feature quality (for example remove highly correlated features), and find the best hyper parameters. This should improve your LB score.</p>",
      "rawMarkdown": "My suggestion is to only focus on one model, improve the feature quality (for example remove highly correlated features), and find the best hyper parameters. This should improve your LB score.",
      "votes": null
    },
    {
      "id": "1850438",
      "postDate": "07/10/2022 12:06:30",
      "content": "<p>Thanks, at the moment I am also giving up on the idea of model fusion and going back to improve the quality of a separate model.</p>",
      "rawMarkdown": "Thanks, at the moment I am also giving up on the idea of model fusion and going back to improve the quality of a separate model.",
      "votes": null
    },
    {
      "id": "1850960",
      "postDate": "07/11/2022 00:24:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a>! Could I please ask you why you would remove highly correlated features? Thx a lot for your answer! 🙏🙂</p>",
      "rawMarkdown": "Hey @mohammadrahmati! Could I please ask you why you would remove highly correlated features? Thx a lot for your answer! 🙏🙂",
      "votes": null
    },
    {
      "id": "1850975",
      "postDate": "07/11/2022 00:36:10",
      "content": "<p>It's possible that you're overfitting with some of your models. It's worth playing around with the hyperparameters of each model to see if you can find a better combination. Good luck!</p>",
      "rawMarkdown": "It's possible that you're overfitting with some of your models. It's worth playing around with the hyperparameters of each model to see if you can find a better combination. Good luck!",
      "votes": null
    },
    {
      "id": "1850987",
      "postDate": "07/11/2022 00:54:07",
      "content": "<p>Does the LB or CV get worse? For an extreme experiment, what happens to the OOF CV if you do 98% LGBM + 1% XGB + 1% CatBoost? Does the CV change better or worse compared to LGBM alone?</p>",
      "rawMarkdown": "Does the LB or CV get worse? For an extreme experiment, what happens to the OOF CV if you do 98% LGBM + 1% XGB + 1% CatBoost? Does the CV change better or worse compared to LGBM alone?",
      "votes": null
    },
    {
      "id": "1850997",
      "postDate": "07/11/2022 01:07:04",
      "content": "<p>CV became worse. Later tried nn and found that the single model nn effect is the worst. Haven't considered the effect of 98% LGBM + 1% XGB + 1% CatBoost yet, let me see if my individual models have tuned to a better result first. Thank you very much for the advice！</p>",
      "rawMarkdown": "CV became worse. Later tried nn and found that the single model nn effect is the worst. Haven't considered the effect of 98% LGBM + 1% XGB + 1% CatBoost yet, let me see if my individual models have tuned to a better result first. Thank you very much for the advice！",
      "votes": null
    },
    {
      "id": "1850998",
      "postDate": "07/11/2022 01:07:18",
      "content": "<p>Thanks. Maybe so, I'll try to tune the parameters of each model again and see if I can go further.</p>",
      "rawMarkdown": "Thanks. Maybe so, I'll try to tune the parameters of each model again and see if I can go further.",
      "votes": null
    },
    {
      "id": "1851367",
      "postDate": "07/11/2022 07:32:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <br>\nFeatures with high correlation are more linearly dependent and hence have almost the same effect on the dependent variable. So, when two features have high correlation, we can drop one of the two features. Highly correlated features can be problematic for most of the algorithms.</p>",
      "rawMarkdown": "Hi @radek1, \nFeatures with high correlation are more linearly dependent and hence have almost the same effect on the dependent variable. So, when two features have high correlation, we can drop one of the two features. Highly correlated features can be problematic for most of the algorithms.",
      "votes": null
    },
    {
      "id": "1851638",
      "postDate": "07/11/2022 12:07:32",
      "content": "<p>Blending a new model such as XGB or CatBoost with LGBM won't necessarily boost the CV LB (but it usually does).</p>\n<p>My suggestion above is the trick I use to see if a new model can help my ensemble. When i build a new model, I will blend it with my old best models at 95% weight for old models and 5% weight for new model. If the new ensemble overall CV is better than old models by themselves, then the new model helps.</p>\n<p>Using this technique, I can test lots of new models like XGB, LGBM, GRU, LSTM, Transformer, DAE, TabNet, SVC, LR, etc etc. If a new model helps, then we can give attention on improving the new model. If the new model does not help, then we can avoid wasting time on that new model.</p>\n<p>Later, we can find optimal blend weights for all our models. And the resultant ensemble will be very strong.</p>",
      "rawMarkdown": "Blending a new model such as XGB or CatBoost with LGBM won't necessarily boost the CV LB (but it usually does).\n\nMy suggestion above is the trick I use to see if a new model can help my ensemble. When i build a new model, I will blend it with my old best models at 95% weight for old models and 5% weight for new model. If the new ensemble overall CV is better than old models by themselves, then the new model helps.\n\nUsing this technique, I can test lots of new models like XGB, LGBM, GRU, LSTM, Transformer, DAE, TabNet, SVC, LR, etc etc. If a new model helps, then we can give attention on improving the new model. If the new model does not help, then we can avoid wasting time on that new model.\n\nLater, we can find optimal blend weights for all our models. And the resultant ensemble will be very strong.",
      "votes": null
    },
    {
      "id": "1851715",
      "postDate": "07/11/2022 13:16:54",
      "content": "<p>Wow, your conversation was very helpful to me!</p>\n<p>I'm still dumbfounded that I keep trying and the highest score I ran after model fusion was 0.794. I always thought it was me, but it turns out the synthesized model is rather useless and I don't have to waste time with it.</p>\n<p>Thank you very much for your help, I've learned a lot from you! Thanks again!</p>",
      "rawMarkdown": "Wow, your conversation was very helpful to me!\n\nI'm still dumbfounded that I keep trying and the highest score I ran after model fusion was 0.794. I always thought it was me, but it turns out the synthesized model is rather useless and I don't have to waste time with it.\n\nThank you very much for your help, I've learned a lot from you! Thanks again!",
      "votes": null
    },
    {
      "id": "1852101",
      "postDate": "07/11/2022 19:14:05",
      "content": "<p>You can maybe check the correlation between the predictions. I noticed that even if a model single score is not great, it can boost the blending if it is less correlated. </p>",
      "rawMarkdown": "You can maybe check the correlation between the predictions. I noticed that even if a model single score is not great, it can boost the blending if it is less correlated.",
      "votes": null
    },
    {
      "id": "1852161",
      "postDate": "07/11/2022 20:54:47",
      "content": "<p>Try to blend the seeds.</p>",
      "rawMarkdown": "Try to blend the seeds.",
      "votes": null
    },
    {
      "id": "1852351",
      "postDate": "07/12/2022 02:37:47",
      "content": "<p>Well, thank you for your advice！</p>",
      "rawMarkdown": "Well, thank you for your advice！",
      "votes": null
    },
    {
      "id": "1852352",
      "postDate": "07/12/2022 02:38:46",
      "content": "<p>Thank you buddy！</p>",
      "rawMarkdown": "Thank you buddy！",
      "votes": null
    },
    {
      "id": "1889051",
      "postDate": "08/08/2022 00:30:15",
      "content": "<p>I am also using LGBM model but my score is much more lower than yours.Did u impute the data if yes how did u impute?I am using LGBMClassifier which outputs 0 or 1.However in the Overview it says your submission should contain probabilities.Does your model outputs probabilities?if yes how can LGBMClassifier outputs probabilities rather than categories.Can you please help me</p>",
      "rawMarkdown": "I am also using LGBM model but my score is much more lower than yours.Did u impute the data if yes how did u impute?I am using LGBMClassifier which outputs 0 or 1.However in the Overview it says your submission should contain probabilities.Does your model outputs probabilities?if yes how can LGBMClassifier outputs probabilities rather than categories.Can you please help me",
      "votes": null
    },
    {
      "id": "1889086",
      "postDate": "08/08/2022 01:04:25",
      "content": "<p><strong>This is a contest that requires the output of probabilities and is not a classification contest, you can see an example of the contest results submission in the overview at：</strong><code># For each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format......</code><br>\n<strong>You can also view some descriptions and detailed information about the contest on it.</strong></p>\n<p>1、My score is higher because I used more than one model, LGBM, but multiple models.<br>\n2、In addition to the original features, special engineering was used to add some new features and some feature filtering methods to remove unnecessary features.<br>\n3、You can refer to other people's public code and compare it with yours so you can see where the problem is.<br>\n<a target=\"_blank\">Amex LGBM Dart CV 0.7977</a></p>\n<p>Good luck for you, I hope you get something out of it. Thanks.</p>",
      "rawMarkdown": "**This is a contest that requires the output of probabilities and is not a classification contest, you can see an example of the contest results submission in the overview at：**`# For each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format......`\n**You can also view some descriptions and detailed information about the contest on it.**\n\n1、My score is higher because I used more than one model, LGBM, but multiple models.\n2、In addition to the original features, special engineering was used to add some new features and some feature filtering methods to remove unnecessary features.\n3、You can refer to other people's public code and compare it with yours so you can see where the problem is.\n[Amex LGBM Dart CV 0.7977](url[https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977](url))\n\nGood luck for you, I hope you get something out of it. Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1850197,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/10/2022 07:32:07",
      "content": "<p>My suggestion is to only focus on one model, improve the feature quality (for example remove highly correlated features), and find the best hyper parameters. This should improve your LB score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1850438,
          "author_name": "",
          "author_url": "",
          "post_date": "07/10/2022 12:06:30",
          "content": "<p>Thanks, at the moment I am also giving up on the idea of model fusion and going back to improve the quality of a separate model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1850960,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "07/11/2022 00:24:18",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a>! Could I please ask you why you would remove highly correlated features? Thx a lot for your answer! 🙏🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851367,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/11/2022 07:32:50",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, <br>\nFeatures with high correlation are more linearly dependent and hence have almost the same effect on the dependent variable. So, when two features have high correlation, we can drop one of the two features. Highly correlated features can be problematic for most of the algorithms.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1850975,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/11/2022 00:36:10",
      "content": "<p>It's possible that you're overfitting with some of your models. It's worth playing around with the hyperparameters of each model to see if you can find a better combination. Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1850998,
          "author_name": "",
          "author_url": "",
          "post_date": "07/11/2022 01:07:18",
          "content": "<p>Thanks. Maybe so, I'll try to tune the parameters of each model again and see if I can go further.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1850987,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/11/2022 00:54:07",
      "content": "<p>Does the LB or CV get worse? For an extreme experiment, what happens to the OOF CV if you do 98% LGBM + 1% XGB + 1% CatBoost? Does the CV change better or worse compared to LGBM alone?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1850997,
          "author_name": "",
          "author_url": "",
          "post_date": "07/11/2022 01:07:04",
          "content": "<p>CV became worse. Later tried nn and found that the single model nn effect is the worst. Haven't considered the effect of 98% LGBM + 1% XGB + 1% CatBoost yet, let me see if my individual models have tuned to a better result first. Thank you very much for the advice！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851638,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/11/2022 12:07:32",
          "content": "<p>Blending a new model such as XGB or CatBoost with LGBM won't necessarily boost the CV LB (but it usually does).</p>\n<p>My suggestion above is the trick I use to see if a new model can help my ensemble. When i build a new model, I will blend it with my old best models at 95% weight for old models and 5% weight for new model. If the new ensemble overall CV is better than old models by themselves, then the new model helps.</p>\n<p>Using this technique, I can test lots of new models like XGB, LGBM, GRU, LSTM, Transformer, DAE, TabNet, SVC, LR, etc etc. If a new model helps, then we can give attention on improving the new model. If the new model does not help, then we can avoid wasting time on that new model.</p>\n<p>Later, we can find optimal blend weights for all our models. And the resultant ensemble will be very strong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851715,
          "author_name": "",
          "author_url": "",
          "post_date": "07/11/2022 13:16:54",
          "content": "<p>Wow, your conversation was very helpful to me!</p>\n<p>I'm still dumbfounded that I keep trying and the highest score I ran after model fusion was 0.794. I always thought it was me, but it turns out the synthesized model is rather useless and I don't have to waste time with it.</p>\n<p>Thank you very much for your help, I've learned a lot from you! Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1852101,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "07/11/2022 19:14:05",
      "content": "<p>You can maybe check the correlation between the predictions. I noticed that even if a model single score is not great, it can boost the blending if it is less correlated. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1852351,
          "author_name": "",
          "author_url": "",
          "post_date": "07/12/2022 02:37:47",
          "content": "<p>Well, thank you for your advice！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1852161,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "07/11/2022 20:54:47",
      "content": "<p>Try to blend the seeds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1852352,
          "author_name": "",
          "author_url": "",
          "post_date": "07/12/2022 02:38:46",
          "content": "<p>Thank you buddy！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1889051,
      "author_name": "muraterennal",
      "author_url": "",
      "post_date": "08/08/2022 00:30:15",
      "content": "<p>I am also using LGBM model but my score is much more lower than yours.Did u impute the data if yes how did u impute?I am using LGBMClassifier which outputs 0 or 1.However in the Overview it says your submission should contain probabilities.Does your model outputs probabilities?if yes how can LGBMClassifier outputs probabilities rather than categories.Can you please help me</p>",
      "votes": null,
      "replies": [
        {
          "id": 1889086,
          "author_name": "",
          "author_url": "",
          "post_date": "08/08/2022 01:04:25",
          "content": "<p><strong>This is a contest that requires the output of probabilities and is not a classification contest, you can see an example of the contest results submission in the overview at：</strong><code># For each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format......</code><br>\n<strong>You can also view some descriptions and detailed information about the contest on it.</strong></p>\n<p>1、My score is higher because I used more than one model, LGBM, but multiple models.<br>\n2、In addition to the original features, special engineering was used to add some new features and some feature filtering methods to remove unnecessary features.<br>\n3、You can refer to other people's public code and compare it with yours so you can see where the problem is.<br>\n<a target=\"_blank\">Amex LGBM Dart CV 0.7977</a></p>\n<p>Good luck for you, I hope you get something out of it. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1849957": "I can score up to 0.799 when using LGBM alone, but when I use XGBoost + LGBM + CatBoost for model fusion, I find that the scores run out lower than any of the individual runs.\nThis is really puzzling to me. Does it feel like it could be because of the amount of features?",
    "1850197": "My suggestion is to only focus on one model, improve the feature quality (for example remove highly correlated features), and find the best hyper parameters. This should improve your LB score.",
    "1850438": "Thanks, at the moment I am also giving up on the idea of model fusion and going back to improve the quality of a separate model.",
    "1850960": "Hey @mohammadrahmati! Could I please ask you why you would remove highly correlated features? Thx a lot for your answer! 🙏🙂",
    "1850975": "It's possible that you're overfitting with some of your models. It's worth playing around with the hyperparameters of each model to see if you can find a better combination. Good luck!",
    "1850987": "Does the LB or CV get worse? For an extreme experiment, what happens to the OOF CV if you do 98% LGBM + 1% XGB + 1% CatBoost? Does the CV change better or worse compared to LGBM alone?",
    "1850997": "CV became worse. Later tried nn and found that the single model nn effect is the worst. Haven't considered the effect of 98% LGBM + 1% XGB + 1% CatBoost yet, let me see if my individual models have tuned to a better result first. Thank you very much for the advice！",
    "1850998": "Thanks. Maybe so, I'll try to tune the parameters of each model again and see if I can go further.",
    "1851367": "Hi @radek1, \nFeatures with high correlation are more linearly dependent and hence have almost the same effect on the dependent variable. So, when two features have high correlation, we can drop one of the two features. Highly correlated features can be problematic for most of the algorithms.",
    "1851638": "Blending a new model such as XGB or CatBoost with LGBM won't necessarily boost the CV LB (but it usually does).\n\nMy suggestion above is the trick I use to see if a new model can help my ensemble. When i build a new model, I will blend it with my old best models at 95% weight for old models and 5% weight for new model. If the new ensemble overall CV is better than old models by themselves, then the new model helps.\n\nUsing this technique, I can test lots of new models like XGB, LGBM, GRU, LSTM, Transformer, DAE, TabNet, SVC, LR, etc etc. If a new model helps, then we can give attention on improving the new model. If the new model does not help, then we can avoid wasting time on that new model.\n\nLater, we can find optimal blend weights for all our models. And the resultant ensemble will be very strong.",
    "1851715": "Wow, your conversation was very helpful to me!\n\nI'm still dumbfounded that I keep trying and the highest score I ran after model fusion was 0.794. I always thought it was me, but it turns out the synthesized model is rather useless and I don't have to waste time with it.\n\nThank you very much for your help, I've learned a lot from you! Thanks again!",
    "1852101": "You can maybe check the correlation between the predictions. I noticed that even if a model single score is not great, it can boost the blending if it is less correlated.",
    "1852161": "Try to blend the seeds.",
    "1852351": "Well, thank you for your advice！",
    "1852352": "Thank you buddy！",
    "1889051": "I am also using LGBM model but my score is much more lower than yours.Did u impute the data if yes how did u impute?I am using LGBMClassifier which outputs 0 or 1.However in the Overview it says your submission should contain probabilities.Does your model outputs probabilities?if yes how can LGBMClassifier outputs probabilities rather than categories.Can you please help me",
    "1889086": "**This is a contest that requires the output of probabilities and is not a classification contest, you can see an example of the contest results submission in the overview at：**`# For each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format......`\n**You can also view some descriptions and detailed information about the contest on it.**\n\n1、My score is higher because I used more than one model, LGBM, but multiple models.\n2、In addition to the original features, special engineering was used to add some new features and some feature filtering methods to remove unnecessary features.\n3、You can refer to other people's public code and compare it with yours so you can see where the problem is.\n[Amex LGBM Dart CV 0.7977](url[https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977](url))\n\nGood luck for you, I hope you get something out of it. Thanks."
  },
  "source": "meta"
}