{
  "id": 543914,
  "title": "big gap between CV and LB",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543914",
  "author_name": "",
  "post_date": "2024-11-02T07:08:45.883007700Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Excuse me, is this normal?<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold\">JS2024 synthetic data with purgedkfold</a></p>",
  "messages": [
    {
      "id": "3034469",
      "postDate": "11/02/2024 07:08:45",
      "content": "<p>Excuse me, is this normal?<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold\">JS2024 synthetic data with purgedkfold</a></p>",
      "rawMarkdown": "Excuse me, is this normal?<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold\">JS2024 synthetic data with purgedkfold</a>",
      "votes": null
    },
    {
      "id": "3034500",
      "postDate": "11/02/2024 07:49:35",
      "content": "<p>Overfitting or data drift, I myself have found the CV to be significantly higher than LB</p>\n<p>I didn't experiment with multiple models because generally speaking financial data should be chosen with a single model to improve accuracy</p>",
      "rawMarkdown": "Overfitting or data drift, I myself have found the CV to be significantly higher than LB\n\nI didn't experiment with multiple models because generally speaking financial data should be chosen with a single model to improve accuracy",
      "votes": null
    },
    {
      "id": "3034511",
      "postDate": "11/02/2024 08:04:03",
      "content": "<p>Other than the gap, do you find that an increase in your CV score aligns with an increase on the LB, or have you noticed that they might be misaligned?</p>",
      "rawMarkdown": "Other than the gap, do you find that an increase in your CV score aligns with an increase on the LB, or have you noticed that they might be misaligned?",
      "votes": null
    },
    {
      "id": "3034617",
      "postDate": "11/02/2024 11:28:50",
      "content": "<p>In light of ISIC2024, I don't put much faith in CV,because CV I can only judge within limited data, small samples may lead to uneven distribution of data across folds, and certain folds may contain more or less samples of a particular category.</p>\n<p>Especially financial data mostly conforms to normal distribution, some extreme quotes tend to result in limited scores per fold</p>",
      "rawMarkdown": "In light of ISIC2024, I don't put much faith in CV,because CV I can only judge within limited data, small samples may lead to uneven distribution of data across folds, and certain folds may contain more or less samples of a particular category.\n\nEspecially financial data mostly conforms to normal distribution, some extreme quotes tend to result in limited scores per fold",
      "votes": null
    },
    {
      "id": "3034713",
      "postDate": "11/02/2024 14:07:19",
      "content": "<p>You mean the lb not the cv here right?<br>\nBecause ISIC is the reason why I don't wish at all to do anything based on that it fits well with lb</p>",
      "rawMarkdown": "You mean the lb not the cv here right?\nBecause ISIC is the reason why I don't wish at all to do anything based on that it fits well with lb",
      "votes": null
    },
    {
      "id": "3035186",
      "postDate": "11/03/2024 05:31:20",
      "content": "<p>Can you explain the ISIC2024 reference? A specific paper presented?<br>\nAlso interesting what is the basis for the previous remark you wrote: \"generally speaking financial data should be chosen with a single model to improve accuracy\"</p>",
      "rawMarkdown": "Can you explain the ISIC2024 reference? A specific paper presented?\nAlso interesting what is the basis for the previous remark you wrote: \"generally speaking financial data should be chosen with a single model to improve accuracy\"",
      "votes": null
    },
    {
      "id": "3037222",
      "postDate": "11/05/2024 14:00:44",
      "content": "<p>I have a feeling that finding a good CV strategy will be of great importance in this competition because the CVs have a gap compared to the LB, so finding a good CV will save a lot of time evaluating the models.</p>",
      "rawMarkdown": "I have a feeling that finding a good CV strategy will be of great importance in this competition because the CVs have a gap compared to the LB, so finding a good CV will save a lot of time evaluating the models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3034500,
      "author_name": "aristotlechen",
      "author_url": "",
      "post_date": "11/02/2024 07:49:35",
      "content": "<p>Overfitting or data drift, I myself have found the CV to be significantly higher than LB</p>\n<p>I didn't experiment with multiple models because generally speaking financial data should be chosen with a single model to improve accuracy</p>",
      "votes": null,
      "replies": [
        {
          "id": 3034511,
          "author_name": "aymanallawi",
          "author_url": "",
          "post_date": "11/02/2024 08:04:03",
          "content": "<p>Other than the gap, do you find that an increase in your CV score aligns with an increase on the LB, or have you noticed that they might be misaligned?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3034617,
              "author_name": "aristotlechen",
              "author_url": "",
              "post_date": "11/02/2024 11:28:50",
              "content": "<p>In light of ISIC2024, I don't put much faith in CV,because CV I can only judge within limited data, small samples may lead to uneven distribution of data across folds, and certain folds may contain more or less samples of a particular category.</p>\n<p>Especially financial data mostly conforms to normal distribution, some extreme quotes tend to result in limited scores per fold</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3034713,
                  "author_name": "aymanallawi",
                  "author_url": "",
                  "post_date": "11/02/2024 14:07:19",
                  "content": "<p>You mean the lb not the cv here right?<br>\nBecause ISIC is the reason why I don't wish at all to do anything based on that it fits well with lb</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3035186,
                  "author_name": "casras",
                  "author_url": "",
                  "post_date": "11/03/2024 05:31:20",
                  "content": "<p>Can you explain the ISIC2024 reference? A specific paper presented?<br>\nAlso interesting what is the basis for the previous remark you wrote: \"generally speaking financial data should be chosen with a single model to improve accuracy\"</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3037222,
      "author_name": "nandodmelo",
      "author_url": "",
      "post_date": "11/05/2024 14:00:44",
      "content": "<p>I have a feeling that finding a good CV strategy will be of great importance in this competition because the CVs have a gap compared to the LB, so finding a good CV will save a lot of time evaluating the models.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3034469": "Excuse me, is this normal?<a href=\"https://www.kaggle.com/code/yunsuxiaozi/js2024-synthetic-data-with-purgedkfold\">JS2024 synthetic data with purgedkfold</a>",
    "3034500": "Overfitting or data drift, I myself have found the CV to be significantly higher than LB\n\nI didn't experiment with multiple models because generally speaking financial data should be chosen with a single model to improve accuracy",
    "3034511": "Other than the gap, do you find that an increase in your CV score aligns with an increase on the LB, or have you noticed that they might be misaligned?",
    "3034617": "In light of ISIC2024, I don't put much faith in CV,because CV I can only judge within limited data, small samples may lead to uneven distribution of data across folds, and certain folds may contain more or less samples of a particular category.\n\nEspecially financial data mostly conforms to normal distribution, some extreme quotes tend to result in limited scores per fold",
    "3034713": "You mean the lb not the cv here right?\nBecause ISIC is the reason why I don't wish at all to do anything based on that it fits well with lb",
    "3035186": "Can you explain the ISIC2024 reference? A specific paper presented?\nAlso interesting what is the basis for the previous remark you wrote: \"generally speaking financial data should be chosen with a single model to improve accuracy\"",
    "3037222": "I have a feeling that finding a good CV strategy will be of great importance in this competition because the CVs have a gap compared to the LB, so finding a good CV will save a lot of time evaluating the models."
  },
  "source": "meta"
}