{
  "id": 588041,
  "title": "Why does the test set divided on train_df perform well, but the final score is not good.",
  "url": "/competitions/drw-crypto-market-prediction/discussion/588041",
  "author_name": "",
  "post_date": "2025-07-04T04:40:50.318633100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/marcosshen/drw-denoise-cnn\" target=\"_blank\">https://www.kaggle.com/code/marcosshen/drw-denoise-cnn</a></p>",
  "messages": [
    {
      "id": "3240634",
      "postDate": "07/04/2025 04:40:50",
      "content": "<p><a href=\"https://www.kaggle.com/code/marcosshen/drw-denoise-cnn\" target=\"_blank\">https://www.kaggle.com/code/marcosshen/drw-denoise-cnn</a></p>",
      "rawMarkdown": "https://www.kaggle.com/code/marcosshen/drw-denoise-cnn",
      "votes": null
    },
    {
      "id": "3240649",
      "postDate": "07/04/2025 05:05:02",
      "content": "<p>The training data is shuffled, which will lead to overfitting<br>\n<code>X_train, X_temp, y_train, y_temp = train_test_split(\n    X, y, train_size=0.90, random_state=42)</code></p>",
      "rawMarkdown": "The training data is shuffled, which will lead to overfitting\n`X_train, X_temp, y_train, y_temp = train_test_split(\n    X, y, train_size=0.90, random_state=42)`",
      "votes": null
    },
    {
      "id": "3240668",
      "postDate": "07/04/2025 05:31:45",
      "content": "<p>That usually happens when the model learns patterns that don’t hold up on the actual test data — maybe due to overfitting, data leakage, or just a validation split that doesn’t match the test set well. You might want to revisit your cross-validation strategy. Also, trying ensembling can help improve performance by making your model more robust.</p>",
      "rawMarkdown": "That usually happens when the model learns patterns that don’t hold up on the actual test data — maybe due to overfitting, data leakage, or just a validation split that doesn’t match the test set well. You might want to revisit your cross-validation strategy. Also, trying ensembling can help improve performance by making your model more robust.",
      "votes": null
    },
    {
      "id": "3242776",
      "postDate": "07/06/2025 11:12:49",
      "content": "<p><a href=\"https://www.kaggle.com/marcosshen\" target=\"_blank\">@marcosshen</a> I also facing same issue, from training data file, I used like 20% of it as trained data model and the rest 80% for test data validation if model perform well. Pearson score is promising at  0.88.</p>\n<p>But upon generating submitted data with test data, I only got score 0.05. That's too far off :(</p>",
      "rawMarkdown": "marcosshen I also facing same issue, from training data file, I used like 20% of it as trained data model and the rest 80% for test data validation if model perform well. Pearson score is promising at  0.88.\n\nBut upon generating submitted data with test data, I only got score 0.05. That's too far off :(",
      "votes": null
    },
    {
      "id": "3243179",
      "postDate": "07/06/2025 20:53:05",
      "content": "<p>This is a common symptom of overfitting.</p>",
      "rawMarkdown": "This is a common symptom of overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3240649,
      "author_name": "littlecitizen",
      "author_url": "",
      "post_date": "07/04/2025 05:05:02",
      "content": "<p>The training data is shuffled, which will lead to overfitting<br>\n<code>X_train, X_temp, y_train, y_temp = train_test_split(\n    X, y, train_size=0.90, random_state=42)</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3240668,
      "author_name": "puneetsingh010",
      "author_url": "",
      "post_date": "07/04/2025 05:31:45",
      "content": "<p>That usually happens when the model learns patterns that don’t hold up on the actual test data — maybe due to overfitting, data leakage, or just a validation split that doesn’t match the test set well. You might want to revisit your cross-validation strategy. Also, trying ensembling can help improve performance by making your model more robust.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3242776,
      "author_name": "budiarsana",
      "author_url": "",
      "post_date": "07/06/2025 11:12:49",
      "content": "<p><a href=\"https://www.kaggle.com/marcosshen\" target=\"_blank\">@marcosshen</a> I also facing same issue, from training data file, I used like 20% of it as trained data model and the rest 80% for test data validation if model perform well. Pearson score is promising at  0.88.</p>\n<p>But upon generating submitted data with test data, I only got score 0.05. That's too far off :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3243179,
      "author_name": "taylorsamarel",
      "author_url": "",
      "post_date": "07/06/2025 20:53:05",
      "content": "<p>This is a common symptom of overfitting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3240634": "https://www.kaggle.com/code/marcosshen/drw-denoise-cnn",
    "3240649": "The training data is shuffled, which will lead to overfitting\n`X_train, X_temp, y_train, y_temp = train_test_split(\n    X, y, train_size=0.90, random_state=42)`",
    "3240668": "That usually happens when the model learns patterns that don’t hold up on the actual test data — maybe due to overfitting, data leakage, or just a validation split that doesn’t match the test set well. You might want to revisit your cross-validation strategy. Also, trying ensembling can help improve performance by making your model more robust.",
    "3242776": "marcosshen I also facing same issue, from training data file, I used like 20% of it as trained data model and the rest 80% for test data validation if model perform well. Pearson score is promising at  0.88.\n\nBut upon generating submitted data with test data, I only got score 0.05. That's too far off :(",
    "3243179": "This is a common symptom of overfitting."
  },
  "source": "meta"
}