{
  "id": 581075,
  "title": "CV/LB Thread",
  "url": "/competitions/drw-crypto-market-prediction/discussion/581075",
  "author_name": "",
  "post_date": "2025-05-28T08:11:13.078590Z",
  "votes": 15,
  "comment_count": 22,
  "views": 0,
  "content": "<p>What is your current best CV and its corresponding LB?</p>\n<p>Mine is currently a small stack of three models ensembled with Ridge. Below are the results, which I will update as I (hopefully) progress throughout the competition.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Date</strong></th>\n<th><strong>Models</strong></th>\n<th><strong>Ensembling method</strong></th>\n<th><strong>5 Fold CV</strong></th>\n<th><strong>Public LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28. May</td>\n<td>2 LGB, 1 XGB</td>\n<td>Ridge</td>\n<td>0.11716</td>\n<td>0.08865</td>\n</tr>\n<tr>\n<td>6. June</td>\n<td>2 LGB, 1 XGB, 1 AutoGluon</td>\n<td>Ridge</td>\n<td>0.182045</td>\n<td>0.09081</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "3211242",
      "postDate": "05/28/2025 08:11:13",
      "content": "<p>What is your current best CV and its corresponding LB?</p>\n<p>Mine is currently a small stack of three models ensembled with Ridge. Below are the results, which I will update as I (hopefully) progress throughout the competition.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Date</strong></th>\n<th><strong>Models</strong></th>\n<th><strong>Ensembling method</strong></th>\n<th><strong>5 Fold CV</strong></th>\n<th><strong>Public LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28. May</td>\n<td>2 LGB, 1 XGB</td>\n<td>Ridge</td>\n<td>0.11716</td>\n<td>0.08865</td>\n</tr>\n<tr>\n<td>6. June</td>\n<td>2 LGB, 1 XGB, 1 AutoGluon</td>\n<td>Ridge</td>\n<td>0.182045</td>\n<td>0.09081</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "What is your current best CV and its corresponding LB?\n\nMine is currently a small stack of three models ensembled with Ridge. Below are the results, which I will update as I (hopefully) progress throughout the competition.\n\n| **Date** |  **Models**  | **Ensembling method** | **5 Fold CV** | **Public LB** |\n|----------|:------------:|:---------------------:|:-------------:|:-------------:|\n| 28. May  | 2 LGB, 1 XGB |         Ridge         |    0.11716    |    0.08865    |\n| 6. June  | 2 LGB, 1 XGB, 1 AutoGluon |         Ridge         |    0.182045    |    0.09081    |",
      "votes": null
    },
    {
      "id": "3211260",
      "postDate": "05/28/2025 08:28:04",
      "content": "<table>\n<thead>\n<tr>\n<th>Date</th>\n<th>Models</th>\n<th>Ensembling method</th>\n<th>5 Fold CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28. May</td>\n<td>1 XGB</td>\n<td>None</td>\n<td>0.04419</td>\n<td>0.04812</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "| Date     | Models         | Ensembling method | 5 Fold CV | Public LB |\n|----------|----------------|-------------------|-----------|-----------|\n| 28. May  | 1 XGB   | None             | 0.04419   | 0.04812   |",
      "votes": null
    },
    {
      "id": "3211286",
      "postDate": "05/28/2025 09:08:42",
      "content": "<p>CV 0.149, LB 0.08847</p>",
      "rawMarkdown": "CV 0.149, LB 0.08847",
      "votes": null
    },
    {
      "id": "3211301",
      "postDate": "05/28/2025 09:36:28",
      "content": "<p>That's quite a gap <a href=\"https://www.kaggle.com/maxuhl98\" target=\"_blank\">@maxuhl98</a>. Are you using standard KFold?</p>",
      "rawMarkdown": "That's quite a gap @maxuhl98. Are you using standard KFold?",
      "votes": null
    },
    {
      "id": "3211304",
      "postDate": "05/28/2025 09:43:40",
      "content": "<p>shuffle=True in Kfold? </p>",
      "rawMarkdown": "shuffle=True in Kfold?",
      "votes": null
    },
    {
      "id": "3211306",
      "postDate": "05/28/2025 09:44:27",
      "content": "<p>Looks like the major choice is Ridge + 3Trees 😂</p>",
      "rawMarkdown": "Looks like the major choice is Ridge + 3Trees 😂",
      "votes": null
    },
    {
      "id": "3211316",
      "postDate": "05/28/2025 09:50:42",
      "content": "<p>No, I am using TimeSeriesSplit and shuffle the training data after splitting, you can see the function to create the splits in my <a href=\"https://www.kaggle.com/code/maxuhl98/xgboost-fe-timeseriessplit\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "No, I am using TimeSeriesSplit and shuffle the training data after splitting, you can see the function to create the splits in my [notebook](https://www.kaggle.com/code/maxuhl98/xgboost-fe-timeseriessplit)",
      "votes": null
    },
    {
      "id": "3211321",
      "postDate": "05/28/2025 09:54:45",
      "content": "<p>I see, that's a solid strategy, and your CV-LB correlation is very good. You mistyped your CV score in the table. I guess it should be 0.04419 instead of 0.4419.</p>",
      "rawMarkdown": "I see, that's a solid strategy, and your CV-LB correlation is very good. You mistyped your CV score in the table. I guess it should be 0.04419 instead of 0.4419.",
      "votes": null
    },
    {
      "id": "3211325",
      "postDate": "05/28/2025 09:57:59",
      "content": "<p>Thanks for the info, I corrected the typo. May I ask which CV strategy you are using?</p>",
      "rawMarkdown": "Thanks for the info, I corrected the typo. May I ask which CV strategy you are using?",
      "votes": null
    },
    {
      "id": "3211344",
      "postDate": "05/28/2025 10:23:37",
      "content": "<p>I'm using KFold without shuffling, which is somewhat leaky. I've seen time series competitions won with this strategy, so I'm not yet sure whether I want to keep using it or switch to TimeSeriesSplit. The CV-LB correlation is definitely better with the latter, but ensembling becomes a bit trickier.</p>",
      "rawMarkdown": "I'm using KFold without shuffling, which is somewhat leaky. I've seen time series competitions won with this strategy, so I'm not yet sure whether I want to keep using it or switch to TimeSeriesSplit. The CV-LB correlation is definitely better with the latter, but ensembling becomes a bit trickier.",
      "votes": null
    },
    {
      "id": "3211405",
      "postDate": "05/28/2025 11:41:46",
      "content": "<p>CV 0.116, LB 0.094. Used some tricks after CV. How did you tune the models though?</p>",
      "rawMarkdown": "CV 0.116, LB 0.094. Used some tricks after CV. How did you tune the models though?",
      "votes": null
    },
    {
      "id": "3211420",
      "postDate": "05/28/2025 12:07:28",
      "content": "<p>Are you using KFold?</p>\n<p>I tuned them using a combination of Optuna and WandB sweeps. This might not be the best option, as I haven't experimented with other methods yet, but I'm currently optimizing the mean Pearson correlation in my tuning scripts. I have a hunch that optimizing a different metric than the one used in the competition might yield better results, but I haven't gotten that far in testing it yet.</p>",
      "rawMarkdown": "Are you using KFold?\n\nI tuned them using a combination of Optuna and WandB sweeps. This might not be the best option, as I haven't experimented with other methods yet, but I'm currently optimizing the mean Pearson correlation in my tuning scripts. I have a hunch that optimizing a different metric than the one used in the competition might yield better results, but I haven't gotten that far in testing it yet.",
      "votes": null
    },
    {
      "id": "3211433",
      "postDate": "05/28/2025 12:38:15",
      "content": "<ol>\n<li>KFold it is.</li>\n<li>I saw your notebook. For your pipeline that's a sharp observation because Pearson is invariant under shifting and recaling, which is troublesome when you take average.</li>\n<li>I think the label has been reverse engineered as the 1st place got a 0.25+ score.</li>\n</ol>",
      "rawMarkdown": "1. KFold it is.\n2. I saw your notebook. For your pipeline that's a sharp observation because Pearson is invariant under shifting and recaling, which is troublesome when you take average.\n3. I think the label has been reverse engineered as the 1st place got a 0.25+ score.",
      "votes": null
    },
    {
      "id": "3211436",
      "postDate": "05/28/2025 12:39:55",
      "content": "<blockquote>\n  <p>I think the label has been reverse engineered as the 1st place got a 0.25+ score.</p>\n</blockquote>\n<p>Possible also things like using Standard Scaler over Train+Test combined (leakage) which happened a lot in the Solana community competition..</p>",
      "rawMarkdown": ">I think the label has been reverse engineered as the 1st place got a 0.25+ score.\n\nPossible also things like using Standard Scaler over Train+Test combined (leakage) which happened a lot in the Solana community competition..",
      "votes": null
    },
    {
      "id": "3211592",
      "postDate": "05/28/2025 15:30:12",
      "content": "<p>Tuned with early stopping and Optuna-based hyperparameter search; post-CV tricks focused on ensembling and test-time augmentation.</p>",
      "rawMarkdown": "Tuned with early stopping and Optuna-based hyperparameter search; post-CV tricks focused on ensembling and test-time augmentation.",
      "votes": null
    },
    {
      "id": "3215240",
      "postDate": "06/01/2025 21:03:53",
      "content": "<p>Does anyone have tips for boosting model performance?</p>",
      "rawMarkdown": "Does anyone have tips for boosting model performance?",
      "votes": null
    },
    {
      "id": "3215836",
      "postDate": "06/02/2025 17:41:07",
      "content": "<p>I used RNN with residual blocks and My pearson coef was 0.8 on CV.. But on LB its only 0.07</p>",
      "rawMarkdown": "I used RNN with residual blocks and My pearson coef was 0.8 on CV.. But on LB its only 0.07",
      "votes": null
    },
    {
      "id": "3224245",
      "postDate": "06/14/2025 15:05:52",
      "content": "<p>Since we are using TimeSeriesSplit, the 1/6th of the oofs will all be 0s. Won't the affect our CV score?</p>",
      "rawMarkdown": "Since we are using TimeSeriesSplit, the 1/6th of the oofs will all be 0s. Won't the affect our CV score?",
      "votes": null
    },
    {
      "id": "3224284",
      "postDate": "06/14/2025 16:17:03",
      "content": "<p>Usually the Training Data from the First Split ist Not used for oof scoring </p>",
      "rawMarkdown": "Usually the Training Data from the First Split ist Not used for oof scoring",
      "votes": null
    },
    {
      "id": "3224285",
      "postDate": "06/14/2025 16:19:04",
      "content": "<p>Are you sure? I thought most people scored including with the first split?</p>",
      "rawMarkdown": "Are you sure? I thought most people scored including with the first split?",
      "votes": null
    },
    {
      "id": "3224341",
      "postDate": "06/14/2025 17:29:07",
      "content": "<p>Yes they are using the validation data (orange) of the first split but the training data (blue) of the first split is usually not used for scoring (which i assumed to be the zeros you mentioned)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F1d3fab43295c3ce0406c08dbf3fae70b%2Fts-cv.png?generation=1749922285638334&amp;alt=media\" alt=\"Time Series CV Visuallized\"></p>\n<p>As you can see the data from the first blue colored split is not used for scoring and in most scenarios where you use the Time Series CV you do not want to use future data to predict the past leading to most approaches not having any oof predictions for the time period of the training data used in the first split</p>",
      "rawMarkdown": "Yes they are using the validation data (orange) of the first split but the training data (blue) of the first split is usually not used for scoring (which i assumed to be the zeros you mentioned)\n\n![Time Series CV Visuallized](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F1d3fab43295c3ce0406c08dbf3fae70b%2Fts-cv.png?generation=1749922285638334&alt=media)\n\nAs you can see the data from the first blue colored split is not used for scoring and in most scenarios where you use the Time Series CV you do not want to use future data to predict the past leading to most approaches not having any oof predictions for the time period of the training data used in the first split",
      "votes": null
    },
    {
      "id": "3227777",
      "postDate": "06/19/2025 09:06:18",
      "content": "<p>I found that if I simply add all features, the cv will be 0.12+, but LB 0.5-. But if I just add few of features, the cv will be very close to LB, 0.7 to 0.7.</p>",
      "rawMarkdown": "I found that if I simply add all features, the cv will be 0.12+, but LB 0.5-. But if I just add few of features, the cv will be very close to LB, 0.7 to 0.7.",
      "votes": null
    },
    {
      "id": "3236244",
      "postDate": "06/30/2025 06:37:19",
      "content": "<p>CV 0.131 LB 0.110</p>",
      "rawMarkdown": "CV 0.131 LB 0.110",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3211260,
      "author_name": "maxuhl98",
      "author_url": "",
      "post_date": "05/28/2025 08:28:04",
      "content": "<table>\n<thead>\n<tr>\n<th>Date</th>\n<th>Models</th>\n<th>Ensembling method</th>\n<th>5 Fold CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28. May</td>\n<td>1 XGB</td>\n<td>None</td>\n<td>0.04419</td>\n<td>0.04812</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 3211301,
          "author_name": "ravaghi",
          "author_url": "",
          "post_date": "05/28/2025 09:36:28",
          "content": "<p>That's quite a gap <a href=\"https://www.kaggle.com/maxuhl98\" target=\"_blank\">@maxuhl98</a>. Are you using standard KFold?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211316,
              "author_name": "maxuhl98",
              "author_url": "",
              "post_date": "05/28/2025 09:50:42",
              "content": "<p>No, I am using TimeSeriesSplit and shuffle the training data after splitting, you can see the function to create the splits in my <a href=\"https://www.kaggle.com/code/maxuhl98/xgboost-fe-timeseriessplit\" target=\"_blank\">notebook</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3211321,
                  "author_name": "ravaghi",
                  "author_url": "",
                  "post_date": "05/28/2025 09:54:45",
                  "content": "<p>I see, that's a solid strategy, and your CV-LB correlation is very good. You mistyped your CV score in the table. I guess it should be 0.04419 instead of 0.4419.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3211325,
                      "author_name": "maxuhl98",
                      "author_url": "",
                      "post_date": "05/28/2025 09:57:59",
                      "content": "<p>Thanks for the info, I corrected the typo. May I ask which CV strategy you are using?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3211344,
                          "author_name": "ravaghi",
                          "author_url": "",
                          "post_date": "05/28/2025 10:23:37",
                          "content": "<p>I'm using KFold without shuffling, which is somewhat leaky. I've seen time series competitions won with this strategy, so I'm not yet sure whether I want to keep using it or switch to TimeSeriesSplit. The CV-LB correlation is definitely better with the latter, but ensembling becomes a bit trickier.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3211304,
          "author_name": "henrysun",
          "author_url": "",
          "post_date": "05/28/2025 09:43:40",
          "content": "<p>shuffle=True in Kfold? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3211286,
      "author_name": "gtoking",
      "author_url": "",
      "post_date": "05/28/2025 09:08:42",
      "content": "<p>CV 0.149, LB 0.08847</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3211306,
      "author_name": "henrysun",
      "author_url": "",
      "post_date": "05/28/2025 09:44:27",
      "content": "<p>Looks like the major choice is Ridge + 3Trees 😂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3211405,
      "author_name": "edmundyoung",
      "author_url": "",
      "post_date": "05/28/2025 11:41:46",
      "content": "<p>CV 0.116, LB 0.094. Used some tricks after CV. How did you tune the models though?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3211420,
          "author_name": "ravaghi",
          "author_url": "",
          "post_date": "05/28/2025 12:07:28",
          "content": "<p>Are you using KFold?</p>\n<p>I tuned them using a combination of Optuna and WandB sweeps. This might not be the best option, as I haven't experimented with other methods yet, but I'm currently optimizing the mean Pearson correlation in my tuning scripts. I have a hunch that optimizing a different metric than the one used in the competition might yield better results, but I haven't gotten that far in testing it yet.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3211433,
              "author_name": "edmundyoung",
              "author_url": "",
              "post_date": "05/28/2025 12:38:15",
              "content": "<ol>\n<li>KFold it is.</li>\n<li>I saw your notebook. For your pipeline that's a sharp observation because Pearson is invariant under shifting and recaling, which is troublesome when you take average.</li>\n<li>I think the label has been reverse engineered as the 1st place got a 0.25+ score.</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 3211436,
                  "author_name": "julianmukaj",
                  "author_url": "",
                  "post_date": "05/28/2025 12:39:55",
                  "content": "<blockquote>\n  <p>I think the label has been reverse engineered as the 1st place got a 0.25+ score.</p>\n</blockquote>\n<p>Possible also things like using Standard Scaler over Train+Test combined (leakage) which happened a lot in the Solana community competition..</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3211592,
      "author_name": "princeverma1010",
      "author_url": "",
      "post_date": "05/28/2025 15:30:12",
      "content": "<p>Tuned with early stopping and Optuna-based hyperparameter search; post-CV tricks focused on ensembling and test-time augmentation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3215240,
      "author_name": "shiyifengt",
      "author_url": "",
      "post_date": "06/01/2025 21:03:53",
      "content": "<p>Does anyone have tips for boosting model performance?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3215836,
      "author_name": "poitrew",
      "author_url": "",
      "post_date": "06/02/2025 17:41:07",
      "content": "<p>I used RNN with residual blocks and My pearson coef was 0.8 on CV.. But on LB its only 0.07</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3224245,
      "author_name": "paperxd",
      "author_url": "",
      "post_date": "06/14/2025 15:05:52",
      "content": "<p>Since we are using TimeSeriesSplit, the 1/6th of the oofs will all be 0s. Won't the affect our CV score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3224284,
          "author_name": "maxuhl98",
          "author_url": "",
          "post_date": "06/14/2025 16:17:03",
          "content": "<p>Usually the Training Data from the First Split ist Not used for oof scoring </p>",
          "votes": null,
          "replies": [
            {
              "id": 3224285,
              "author_name": "paperxd",
              "author_url": "",
              "post_date": "06/14/2025 16:19:04",
              "content": "<p>Are you sure? I thought most people scored including with the first split?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3224341,
                  "author_name": "maxuhl98",
                  "author_url": "",
                  "post_date": "06/14/2025 17:29:07",
                  "content": "<p>Yes they are using the validation data (orange) of the first split but the training data (blue) of the first split is usually not used for scoring (which i assumed to be the zeros you mentioned)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F1d3fab43295c3ce0406c08dbf3fae70b%2Fts-cv.png?generation=1749922285638334&amp;alt=media\" alt=\"Time Series CV Visuallized\"></p>\n<p>As you can see the data from the first blue colored split is not used for scoring and in most scenarios where you use the Time Series CV you do not want to use future data to predict the past leading to most approaches not having any oof predictions for the time period of the training data used in the first split</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3227777,
      "author_name": "inszva",
      "author_url": "",
      "post_date": "06/19/2025 09:06:18",
      "content": "<p>I found that if I simply add all features, the cv will be 0.12+, but LB 0.5-. But if I just add few of features, the cv will be very close to LB, 0.7 to 0.7.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3236244,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "06/30/2025 06:37:19",
      "content": "<p>CV 0.131 LB 0.110</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3211242": "What is your current best CV and its corresponding LB?\n\nMine is currently a small stack of three models ensembled with Ridge. Below are the results, which I will update as I (hopefully) progress throughout the competition.\n\n| **Date** |  **Models**  | **Ensembling method** | **5 Fold CV** | **Public LB** |\n|----------|:------------:|:---------------------:|:-------------:|:-------------:|\n| 28. May  | 2 LGB, 1 XGB |         Ridge         |    0.11716    |    0.08865    |\n| 6. June  | 2 LGB, 1 XGB, 1 AutoGluon |         Ridge         |    0.182045    |    0.09081    |",
    "3211260": "| Date     | Models         | Ensembling method | 5 Fold CV | Public LB |\n|----------|----------------|-------------------|-----------|-----------|\n| 28. May  | 1 XGB   | None             | 0.04419   | 0.04812   |",
    "3211286": "CV 0.149, LB 0.08847",
    "3211301": "That's quite a gap @maxuhl98. Are you using standard KFold?",
    "3211304": "shuffle=True in Kfold?",
    "3211306": "Looks like the major choice is Ridge + 3Trees 😂",
    "3211316": "No, I am using TimeSeriesSplit and shuffle the training data after splitting, you can see the function to create the splits in my [notebook](https://www.kaggle.com/code/maxuhl98/xgboost-fe-timeseriessplit)",
    "3211321": "I see, that's a solid strategy, and your CV-LB correlation is very good. You mistyped your CV score in the table. I guess it should be 0.04419 instead of 0.4419.",
    "3211325": "Thanks for the info, I corrected the typo. May I ask which CV strategy you are using?",
    "3211344": "I'm using KFold without shuffling, which is somewhat leaky. I've seen time series competitions won with this strategy, so I'm not yet sure whether I want to keep using it or switch to TimeSeriesSplit. The CV-LB correlation is definitely better with the latter, but ensembling becomes a bit trickier.",
    "3211405": "CV 0.116, LB 0.094. Used some tricks after CV. How did you tune the models though?",
    "3211420": "Are you using KFold?\n\nI tuned them using a combination of Optuna and WandB sweeps. This might not be the best option, as I haven't experimented with other methods yet, but I'm currently optimizing the mean Pearson correlation in my tuning scripts. I have a hunch that optimizing a different metric than the one used in the competition might yield better results, but I haven't gotten that far in testing it yet.",
    "3211433": "1. KFold it is.\n2. I saw your notebook. For your pipeline that's a sharp observation because Pearson is invariant under shifting and recaling, which is troublesome when you take average.\n3. I think the label has been reverse engineered as the 1st place got a 0.25+ score.",
    "3211436": ">I think the label has been reverse engineered as the 1st place got a 0.25+ score.\n\nPossible also things like using Standard Scaler over Train+Test combined (leakage) which happened a lot in the Solana community competition..",
    "3211592": "Tuned with early stopping and Optuna-based hyperparameter search; post-CV tricks focused on ensembling and test-time augmentation.",
    "3215240": "Does anyone have tips for boosting model performance?",
    "3215836": "I used RNN with residual blocks and My pearson coef was 0.8 on CV.. But on LB its only 0.07",
    "3224245": "Since we are using TimeSeriesSplit, the 1/6th of the oofs will all be 0s. Won't the affect our CV score?",
    "3224284": "Usually the Training Data from the First Split ist Not used for oof scoring",
    "3224285": "Are you sure? I thought most people scored including with the first split?",
    "3224341": "Yes they are using the validation data (orange) of the first split but the training data (blue) of the first split is usually not used for scoring (which i assumed to be the zeros you mentioned)\n\n![Time Series CV Visuallized](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F1d3fab43295c3ce0406c08dbf3fae70b%2Fts-cv.png?generation=1749922285638334&alt=media)\n\nAs you can see the data from the first blue colored split is not used for scoring and in most scenarios where you use the Time Series CV you do not want to use future data to predict the past leading to most approaches not having any oof predictions for the time period of the training data used in the first split",
    "3227777": "I found that if I simply add all features, the cv will be 0.12+, but LB 0.5-. But if I just add few of features, the cv will be very close to LB, 0.7 to 0.7.",
    "3236244": "CV 0.131 LB 0.110"
  },
  "source": "meta"
}