{
  "id": 494915,
  "title": "Finding it hard to improve on the benchmark (LGBM)",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/494915",
  "author_name": "",
  "post_date": "2024-04-18T22:11:36.771186500Z",
  "votes": 13,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Approaching this challenge, I decided to first freeze the feature set and trial various model setups against the starter LGBM set up. The trouble is that everything I have thrown at this problem so far has underperformed against the benchmark (I.e. LGBM with validation dataset). This includes:</p>\n<ul>\n<li>RF, XGB, and all classifiers in py caret</li>\n<li>Hyperparameter tunning with grid search and Bayesian optimisation</li>\n<li>target balancing (manually and with SMOTE)</li>\n<li>dimensionality reduction (PCA / SVD)</li>\n<li>pre-training a model to guess params before full training</li>\n<li>Cross-validation</li>\n</ul>\n<p>Model scores generally look competitive and the validation and testing sets (AUC ~0.65-0.75 and stability up to around 0.5). However, when I submit these the stability drops massively and I struggle to get this over 0.2. Often the more data I use and the more sophisticated a model I use, the lower score I get. This leads me to conclude:</p>\n<ul>\n<li>even moderate attempts to train a model leads to over fitting</li>\n<li>only extremely general solutions will deliver stability </li>\n<li>the model is largely irrelevant (I should really be making more features at this point).</li>\n</ul>\n<p>Interested to hear if anyone has had similar / contrasting experiences</p>",
  "messages": [
    {
      "id": "2759792",
      "postDate": "04/18/2024 22:11:36",
      "content": "<p>Approaching this challenge, I decided to first freeze the feature set and trial various model setups against the starter LGBM set up. The trouble is that everything I have thrown at this problem so far has underperformed against the benchmark (I.e. LGBM with validation dataset). This includes:</p>\n<ul>\n<li>RF, XGB, and all classifiers in py caret</li>\n<li>Hyperparameter tunning with grid search and Bayesian optimisation</li>\n<li>target balancing (manually and with SMOTE)</li>\n<li>dimensionality reduction (PCA / SVD)</li>\n<li>pre-training a model to guess params before full training</li>\n<li>Cross-validation</li>\n</ul>\n<p>Model scores generally look competitive and the validation and testing sets (AUC ~0.65-0.75 and stability up to around 0.5). However, when I submit these the stability drops massively and I struggle to get this over 0.2. Often the more data I use and the more sophisticated a model I use, the lower score I get. This leads me to conclude:</p>\n<ul>\n<li>even moderate attempts to train a model leads to over fitting</li>\n<li>only extremely general solutions will deliver stability </li>\n<li>the model is largely irrelevant (I should really be making more features at this point).</li>\n</ul>\n<p>Interested to hear if anyone has had similar / contrasting experiences</p>",
      "rawMarkdown": "Approaching this challenge, I decided to first freeze the feature set and trial various model setups against the starter LGBM set up. The trouble is that everything I have thrown at this problem so far has underperformed against the benchmark (I.e. LGBM with validation dataset). This includes:\n\n- RF, XGB, and all classifiers in py caret\n- Hyperparameter tunning with grid search and Bayesian optimisation\n- target balancing (manually and with SMOTE)\n- dimensionality reduction (PCA / SVD)\n- pre-training a model to guess params before full training\n- Cross-validation\n\nModel scores generally look competitive and the validation and testing sets (AUC ~0.65-0.75 and stability up to around 0.5). However, when I submit these the stability drops massively and I struggle to get this over 0.2. Often the more data I use and the more sophisticated a model I use, the lower score I get. This leads me to conclude:\n- even moderate attempts to train a model leads to over fitting\n- only extremely general solutions will deliver stability \n- the model is largely irrelevant (I should really be making more features at this point).\n\nInterested to hear if anyone has had similar / contrasting experiences",
      "votes": null
    },
    {
      "id": "2759995",
      "postDate": "04/19/2024 03:53:25",
      "content": "<p>Hi! I get exactly the same results.</p>",
      "rawMarkdown": "Hi! I get exactly the same results.",
      "votes": null
    },
    {
      "id": "2760269",
      "postDate": "04/19/2024 07:27:01",
      "content": "<p>Nice to read. Many people would face the same problem if they started from scratch. I noticed that the models with class weights make the results worse. (0.05~0.1 difference) I think test dataset has more reactive to some features than the train set so the unbalanced model gives you the better score in submission.<br>\nCheck out if you have same experiences as mine.</p>",
      "rawMarkdown": "Nice to read. Many people would face the same problem if they started from scratch. I noticed that the models with class weights make the results worse. (0.05~0.1 difference) I think test dataset has more reactive to some features than the train set so the unbalanced model gives you the better score in submission.\nCheck out if you have same experiences as mine.",
      "votes": null
    },
    {
      "id": "2761625",
      "postDate": "04/20/2024 00:43:27",
      "content": "<p>My AUC on my local training data is .85ish for my validation data. +/- a bit with CV as expected. My stability metric is .6-.7</p>\n<p>When I submit this model it drops considerably to below .5 (our first submission was basically our best and we haven't improved).</p>\n<p>In a general sense, it appears our model is overfit to the training data,</p>\n<p>I have been struggling with this. I am sure there are multiple reasons, but one may be that we did not do strong feature engineering. Our model relies on a few key features as well. These may not be present in the Kaggle test data to the same degree or perhaps they have data drift. Finally, we did not do a lot of categorical data investigation. I have a feeling that some of our aggregations etc. are a bit messy and don't perform well to unseen data. </p>",
      "rawMarkdown": "My AUC on my local training data is .85ish for my validation data. +/- a bit with CV as expected. My stability metric is .6-.7\n\nWhen I submit this model it drops considerably to below .5 (our first submission was basically our best and we haven't improved).\n\nIn a general sense, it appears our model is overfit to the training data,\n\nI have been struggling with this. I am sure there are multiple reasons, but one may be that we did not do strong feature engineering. Our model relies on a few key features as well. These may not be present in the Kaggle test data to the same degree or perhaps they have data drift. Finally, we did not do a lot of categorical data investigation. I have a feeling that some of our aggregations etc. are a bit messy and don't perform well to unseen data.",
      "votes": null
    },
    {
      "id": "2763789",
      "postDate": "04/20/2024 18:30:43",
      "content": "<p>It seems that adding features does not influence much the validation score?</p>",
      "rawMarkdown": "It seems that adding features does not influence much the validation score?",
      "votes": null
    },
    {
      "id": "2766912",
      "postDate": "04/22/2024 02:39:05",
      "content": "<p>Maybe try to start with the high scoring code that people have publicly submitted and see if you can improve upon it.</p>\n<p>In terms of overfitting: There is heaps of data, the drop in stability from the training data to the leaderboard score I would not characterise as a result of overfitting. Rather, the data distribution and dynamics change over time and so you generally expect the model to decrease in performance over time. Indeed, the whole point of the competition is to try to find solutions to this problem.</p>",
      "rawMarkdown": "Maybe try to start with the high scoring code that people have publicly submitted and see if you can improve upon it.\n\nIn terms of overfitting: There is heaps of data, the drop in stability from the training data to the leaderboard score I would not characterise as a result of overfitting. Rather, the data distribution and dynamics change over time and so you generally expect the model to decrease in performance over time. Indeed, the whole point of the competition is to try to find solutions to this problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2759995,
      "author_name": "alexxanderlarko",
      "author_url": "",
      "post_date": "04/19/2024 03:53:25",
      "content": "<p>Hi! I get exactly the same results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2760269,
      "author_name": "gilgarad",
      "author_url": "",
      "post_date": "04/19/2024 07:27:01",
      "content": "<p>Nice to read. Many people would face the same problem if they started from scratch. I noticed that the models with class weights make the results worse. (0.05~0.1 difference) I think test dataset has more reactive to some features than the train set so the unbalanced model gives you the better score in submission.<br>\nCheck out if you have same experiences as mine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2761625,
      "author_name": "nicksalem",
      "author_url": "",
      "post_date": "04/20/2024 00:43:27",
      "content": "<p>My AUC on my local training data is .85ish for my validation data. +/- a bit with CV as expected. My stability metric is .6-.7</p>\n<p>When I submit this model it drops considerably to below .5 (our first submission was basically our best and we haven't improved).</p>\n<p>In a general sense, it appears our model is overfit to the training data,</p>\n<p>I have been struggling with this. I am sure there are multiple reasons, but one may be that we did not do strong feature engineering. Our model relies on a few key features as well. These may not be present in the Kaggle test data to the same degree or perhaps they have data drift. Finally, we did not do a lot of categorical data investigation. I have a feeling that some of our aggregations etc. are a bit messy and don't perform well to unseen data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2763789,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "04/20/2024 18:30:43",
      "content": "<p>It seems that adding features does not influence much the validation score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766912,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "04/22/2024 02:39:05",
      "content": "<p>Maybe try to start with the high scoring code that people have publicly submitted and see if you can improve upon it.</p>\n<p>In terms of overfitting: There is heaps of data, the drop in stability from the training data to the leaderboard score I would not characterise as a result of overfitting. Rather, the data distribution and dynamics change over time and so you generally expect the model to decrease in performance over time. Indeed, the whole point of the competition is to try to find solutions to this problem.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2759792": "Approaching this challenge, I decided to first freeze the feature set and trial various model setups against the starter LGBM set up. The trouble is that everything I have thrown at this problem so far has underperformed against the benchmark (I.e. LGBM with validation dataset). This includes:\n\n- RF, XGB, and all classifiers in py caret\n- Hyperparameter tunning with grid search and Bayesian optimisation\n- target balancing (manually and with SMOTE)\n- dimensionality reduction (PCA / SVD)\n- pre-training a model to guess params before full training\n- Cross-validation\n\nModel scores generally look competitive and the validation and testing sets (AUC ~0.65-0.75 and stability up to around 0.5). However, when I submit these the stability drops massively and I struggle to get this over 0.2. Often the more data I use and the more sophisticated a model I use, the lower score I get. This leads me to conclude:\n- even moderate attempts to train a model leads to over fitting\n- only extremely general solutions will deliver stability \n- the model is largely irrelevant (I should really be making more features at this point).\n\nInterested to hear if anyone has had similar / contrasting experiences",
    "2759995": "Hi! I get exactly the same results.",
    "2760269": "Nice to read. Many people would face the same problem if they started from scratch. I noticed that the models with class weights make the results worse. (0.05~0.1 difference) I think test dataset has more reactive to some features than the train set so the unbalanced model gives you the better score in submission.\nCheck out if you have same experiences as mine.",
    "2761625": "My AUC on my local training data is .85ish for my validation data. +/- a bit with CV as expected. My stability metric is .6-.7\n\nWhen I submit this model it drops considerably to below .5 (our first submission was basically our best and we haven't improved).\n\nIn a general sense, it appears our model is overfit to the training data,\n\nI have been struggling with this. I am sure there are multiple reasons, but one may be that we did not do strong feature engineering. Our model relies on a few key features as well. These may not be present in the Kaggle test data to the same degree or perhaps they have data drift. Finally, we did not do a lot of categorical data investigation. I have a feeling that some of our aggregations etc. are a bit messy and don't perform well to unseen data.",
    "2763789": "It seems that adding features does not influence much the validation score?",
    "2766912": "Maybe try to start with the high scoring code that people have publicly submitted and see if you can improve upon it.\n\nIn terms of overfitting: There is heaps of data, the drop in stability from the training data to the leaderboard score I would not characterise as a result of overfitting. Rather, the data distribution and dynamics change over time and so you generally expect the model to decrease in performance over time. Indeed, the whole point of the competition is to try to find solutions to this problem."
  },
  "source": "meta"
}