{
  "id": 589613,
  "title": "CatBoost Overfitting Issue – Need Help Understanding Poor Submission Score",
  "url": "/competitions/drw-crypto-market-prediction/discussion/589613",
  "author_name": "",
  "post_date": "2025-07-14T06:09:12.870740600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm facing an overfitting issue while training a CatBoost model. Here's what happened:</p>\n<p>🔹 First Notebook (Baseline):<br>\n<a href=\"https://www.kaggle.com/code/qamarmath/catboostregressor-baseline\" target=\"_blank\">https://www.kaggle.com/code/qamarmath/catboostregressor-baseline</a><br>\nTrain Pearson Correlation: 0.3946</p>\n<p>Validation Pearson Correlation: 0.1142</p>\n<p>Submission Score: 0.07</p>\n<p>The model clearly overfit the training data.<br>\n🔹 Second Notebook (After Hyperparameter Tuning):<br>\n<a href=\"https://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost\" target=\"_blank\">https://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost</a><br>\nTrain Pearson Correlation: 0.2250 ✅</p>\n<p>Validation Pearson Correlation: 0.1251 ✅</p>\n<p>Submission Score: 0.04 ❌</p>\n<p>This model shows less overfitting and performs slightly better on validation — but the submission score dropped even more.</p>\n<p>I'm trying to understand why the second model, which looks better on both train and validation, performs worse on the test data.</p>",
  "messages": [
    {
      "id": "3248188",
      "postDate": "07/14/2025 06:09:12",
      "content": "<p>Hi everyone,</p>\n<p>I'm facing an overfitting issue while training a CatBoost model. Here's what happened:</p>\n<p>🔹 First Notebook (Baseline):<br>\n<a href=\"https://www.kaggle.com/code/qamarmath/catboostregressor-baseline\" target=\"_blank\">https://www.kaggle.com/code/qamarmath/catboostregressor-baseline</a><br>\nTrain Pearson Correlation: 0.3946</p>\n<p>Validation Pearson Correlation: 0.1142</p>\n<p>Submission Score: 0.07</p>\n<p>The model clearly overfit the training data.<br>\n🔹 Second Notebook (After Hyperparameter Tuning):<br>\n<a href=\"https://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost\" target=\"_blank\">https://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost</a><br>\nTrain Pearson Correlation: 0.2250 ✅</p>\n<p>Validation Pearson Correlation: 0.1251 ✅</p>\n<p>Submission Score: 0.04 ❌</p>\n<p>This model shows less overfitting and performs slightly better on validation — but the submission score dropped even more.</p>\n<p>I'm trying to understand why the second model, which looks better on both train and validation, performs worse on the test data.</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm facing an overfitting issue while training a CatBoost model. Here's what happened:\n\n🔹 First Notebook (Baseline):\nhttps://www.kaggle.com/code/qamarmath/catboostregressor-baseline\nTrain Pearson Correlation: 0.3946\n\nValidation Pearson Correlation: 0.1142\n\nSubmission Score: 0.07\n\nThe model clearly overfit the training data.\n🔹 Second Notebook (After Hyperparameter Tuning):\nhttps://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost\nTrain Pearson Correlation: 0.2250 ✅\n\nValidation Pearson Correlation: 0.1251 ✅\n\nSubmission Score: 0.04 ❌\n\nThis model shows less overfitting and performs slightly better on validation — but the submission score dropped even more.\n\nI'm trying to understand why the second model, which looks better on both train and validation, performs worse on the test data.",
      "votes": null
    },
    {
      "id": "3248249",
      "postDate": "07/14/2025 07:57:10",
      "content": "<p>The features are non stationnary, meaning that the relation you have modeled in train and validated in val may have completely vanished in the test set. You can try simpler models (capturing real structural patterns and not just temporary statistical behaviors) or ensemble many more models to decrease your variance or even model the different regimes of the market.</p>",
      "rawMarkdown": "The features are non stationnary, meaning that the relation you have modeled in train and validated in val may have completely vanished in the test set. You can try simpler models (capturing real structural patterns and not just temporary statistical behaviors) or ensemble many more models to decrease your variance or even model the different regimes of the market.",
      "votes": null
    },
    {
      "id": "3248299",
      "postDate": "07/14/2025 10:15:32",
      "content": "<p>For a start you need to 1) close the cv gap. There are reasons you are overfitting (feature engineering ?, lack of regularisation ?). and 2) add robust evaluation. For exemple your models can be both 0.05 +/- 0.02 and you current evaluation is just not enough to conclude. </p>",
      "rawMarkdown": "For a start you need to 1) close the cv gap. There are reasons you are overfitting (feature engineering ?, lack of regularisation ?). and 2) add robust evaluation. For exemple your models can be both 0.05 +/- 0.02 and you current evaluation is just not enough to conclude.",
      "votes": null
    },
    {
      "id": "3248365",
      "postDate": "07/14/2025 12:42:06",
      "content": "<p>Agree with the comments from <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> and <a href=\"https://www.kaggle.com/yannfb\" target=\"_blank\">@yannfb</a> .</p>\n<p>The features that show importance in the train dataset may drift/have completely different behavior in the test dataset.</p>\n<p>If you look at the charts from these notebook, you can see that from a stability point of view, there are very few features that have good stability in the test and train dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0</a><br>\n<a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm</a><br>\n<a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb</a></p>\n<p>I have found that in addition to regularization, noise injection, and drop out, using an overfitting discriminator also worked well.</p>",
      "rawMarkdown": "Agree with the comments from @lucasmorin and @yannfb .\n\nThe features that show importance in the train dataset may drift/have completely different behavior in the test dataset.\n\nIf you look at the charts from these notebook, you can see that from a stability point of view, there are very few features that have good stability in the test and train dataset.\n\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb\n\nI have found that in addition to regularization, noise injection, and drop out, using an overfitting discriminator also worked well.",
      "votes": null
    },
    {
      "id": "3248400",
      "postDate": "07/14/2025 13:43:16",
      "content": "<p>To add to this, look into techniques for detecting distributional shift, e.g. adversarial validation. Especially in crypto distributional shifts can be strong.</p>",
      "rawMarkdown": "To add to this, look into techniques for detecting distributional shift, e.g. adversarial validation. Especially in crypto distributional shifts can be strong.",
      "votes": null
    },
    {
      "id": "3248404",
      "postDate": "07/14/2025 13:58:15",
      "content": "<p>Good point, adversarial validation is a great tool.</p>",
      "rawMarkdown": "Good point, adversarial validation is a great tool.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3248249,
      "author_name": "yannfb",
      "author_url": "",
      "post_date": "07/14/2025 07:57:10",
      "content": "<p>The features are non stationnary, meaning that the relation you have modeled in train and validated in val may have completely vanished in the test set. You can try simpler models (capturing real structural patterns and not just temporary statistical behaviors) or ensemble many more models to decrease your variance or even model the different regimes of the market.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3248299,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "07/14/2025 10:15:32",
      "content": "<p>For a start you need to 1) close the cv gap. There are reasons you are overfitting (feature engineering ?, lack of regularisation ?). and 2) add robust evaluation. For exemple your models can be both 0.05 +/- 0.02 and you current evaluation is just not enough to conclude. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3248365,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/14/2025 12:42:06",
      "content": "<p>Agree with the comments from <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> and <a href=\"https://www.kaggle.com/yannfb\" target=\"_blank\">@yannfb</a> .</p>\n<p>The features that show importance in the train dataset may drift/have completely different behavior in the test dataset.</p>\n<p>If you look at the charts from these notebook, you can see that from a stability point of view, there are very few features that have good stability in the test and train dataset.</p>\n<p><a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0</a><br>\n<a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm</a><br>\n<a href=\"https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb\" target=\"_blank\">https://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb</a></p>\n<p>I have found that in addition to regularization, noise injection, and drop out, using an overfitting discriminator also worked well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3248400,
          "author_name": "drshredz",
          "author_url": "",
          "post_date": "07/14/2025 13:43:16",
          "content": "<p>To add to this, look into techniques for detecting distributional shift, e.g. adversarial validation. Especially in crypto distributional shifts can be strong.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3248404,
              "author_name": "taylorsamarel",
              "author_url": "",
              "post_date": "07/14/2025 13:58:15",
              "content": "<p>Good point, adversarial validation is a great tool.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3248188": "Hi everyone,\n\nI'm facing an overfitting issue while training a CatBoost model. Here's what happened:\n\n🔹 First Notebook (Baseline):\nhttps://www.kaggle.com/code/qamarmath/catboostregressor-baseline\nTrain Pearson Correlation: 0.3946\n\nValidation Pearson Correlation: 0.1142\n\nSubmission Score: 0.07\n\nThe model clearly overfit the training data.\n🔹 Second Notebook (After Hyperparameter Tuning):\nhttps://www.kaggle.com/code/qamarmath/hyperparameter-tuning-of-catboost\nTrain Pearson Correlation: 0.2250 ✅\n\nValidation Pearson Correlation: 0.1251 ✅\n\nSubmission Score: 0.04 ❌\n\nThis model shows less overfitting and performs slightly better on validation — but the submission score dropped even more.\n\nI'm trying to understand why the second model, which looks better on both train and validation, performs worse on the test data.",
    "3248249": "The features are non stationnary, meaning that the relation you have modeled in train and validated in val may have completely vanished in the test set. You can try simpler models (capturing real structural patterns and not just temporary statistical behaviors) or ensemble many more models to decrease your variance or even model the different regimes of the market.",
    "3248299": "For a start you need to 1) close the cv gap. There are reasons you are overfitting (feature engineering ?, lack of regularisation ?). and 2) add robust evaluation. For exemple your models can be both 0.05 +/- 0.02 and you current evaluation is just not enough to conclude.",
    "3248365": "Agree with the comments from @lucasmorin and @yannfb .\n\nThe features that show importance in the train dataset may drift/have completely different behavior in the test dataset.\n\nIf you look at the charts from these notebook, you can see that from a stability point of view, there are very few features that have good stability in the test and train dataset.\n\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-using-time-series-pseudo-labels-2-0\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-lightgbm\nhttps://www.kaggle.com/code/taylorsamarel/feature-select-pseudo-labels-x-xgb\n\nI have found that in addition to regularization, noise injection, and drop out, using an overfitting discriminator also worked well.",
    "3248400": "To add to this, look into techniques for detecting distributional shift, e.g. adversarial validation. Especially in crypto distributional shifts can be strong.",
    "3248404": "Good point, adversarial validation is a great tool."
  },
  "source": "meta"
}