{
  "id": 451404,
  "title": "about local cv",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/451404",
  "author_name": "",
  "post_date": "2023-10-28T15:41:51.053820300Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm a beginner in a Kaggle competition, and I have a question to ask. What does 'local CV' mean? Does it refer to the mean of MRRMSE in k-fold cross-validation or the last MRRMSE in k-fold cross-validation, or some other data? </p>",
  "messages": [
    {
      "id": "2502947",
      "postDate": "10/28/2023 15:41:51",
      "content": "<p>I'm a beginner in a Kaggle competition, and I have a question to ask. What does 'local CV' mean? Does it refer to the mean of MRRMSE in k-fold cross-validation or the last MRRMSE in k-fold cross-validation, or some other data? </p>",
      "rawMarkdown": "I'm a beginner in a Kaggle competition, and I have a question to ask. What does 'local CV' mean? Does it refer to the mean of MRRMSE in k-fold cross-validation or the last MRRMSE in k-fold cross-validation, or some other data?",
      "votes": null
    },
    {
      "id": "2502983",
      "postDate": "10/28/2023 16:06:31",
      "content": "<p>I think it means what ever metric one uses for local cross validation and it's likely everyone is talking about different ones?</p>",
      "rawMarkdown": "I think it means what ever metric one uses for local cross validation and it's likely everyone is talking about different ones?",
      "votes": null
    },
    {
      "id": "2503002",
      "postDate": "10/28/2023 16:36:10",
      "content": "<p>I thought there was default “local cv”, and it was my lack of understanding of default ”local cv” that made it unclear how everyone was measuring code.  Now I understand, thank you for your response.</p>",
      "rawMarkdown": "I thought there was default “local cv”, and it was my lack of understanding of default ”local cv” that made it unclear how everyone was measuring code.  Now I understand, thank you for your response.",
      "votes": null
    },
    {
      "id": "2503105",
      "postDate": "10/28/2023 19:05:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yanchengwei\" target=\"_blank\">@yanchengwei</a> I suggest not to use a simple k-fold, but a scheme like the one in <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">this notebook</a>. The simple k-fold isn't appropriate for simulating the task of predicting differential expression for an unseen cell type.</p>",
      "rawMarkdown": "Hi @yanchengwei I suggest not to use a simple k-fold, but a scheme like the one in [this notebook](https://www.kaggle.com/code/ambrosm/scp-quickstart). The simple k-fold isn't appropriate for simulating the task of predicting differential expression for an unseen cell type.",
      "votes": null
    },
    {
      "id": "2503370",
      "postDate": "10/29/2023 03:19:57",
      "content": "<p>Thanks for your advice. I'll try. Thanks again</p>",
      "rawMarkdown": "Thanks for your advice. I'll try. Thanks again",
      "votes": null
    },
    {
      "id": "2503825",
      "postDate": "10/29/2023 13:36:35",
      "content": "<p>Hello, when people refer to 'local cv,' they mean the average of the resulting MRRMSEs or any other chosen metric (such as MAE, MSE, Huber, etc.). I use cross-validation (cv) primarily for tuning my models and then train the best-tuned model with the optimal hyperparameters on random data splits. However, I'd like to point out that using cross-validation for the final selection of the best model may not be advisable in this context due to the potential for high uncertainty in the results.</p>",
      "rawMarkdown": "Hello, when people refer to 'local cv,' they mean the average of the resulting MRRMSEs or any other chosen metric (such as MAE, MSE, Huber, etc.). I use cross-validation (cv) primarily for tuning my models and then train the best-tuned model with the optimal hyperparameters on random data splits. However, I'd like to point out that using cross-validation for the final selection of the best model may not be advisable in this context due to the potential for high uncertainty in the results.",
      "votes": null
    },
    {
      "id": "2503859",
      "postDate": "10/29/2023 14:13:59",
      "content": "<p>Thanks for your answer.I get it now</p>",
      "rawMarkdown": "Thanks for your answer.I get it now",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2502983,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "10/28/2023 16:06:31",
      "content": "<p>I think it means what ever metric one uses for local cross validation and it's likely everyone is talking about different ones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503002,
          "author_name": "yanchengwei",
          "author_url": "",
          "post_date": "10/28/2023 16:36:10",
          "content": "<p>I thought there was default “local cv”, and it was my lack of understanding of default ”local cv” that made it unclear how everyone was measuring code.  Now I understand, thank you for your response.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2503105,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "10/28/2023 19:05:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yanchengwei\" target=\"_blank\">@yanchengwei</a> I suggest not to use a simple k-fold, but a scheme like the one in <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">this notebook</a>. The simple k-fold isn't appropriate for simulating the task of predicting differential expression for an unseen cell type.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503370,
          "author_name": "yanchengwei",
          "author_url": "",
          "post_date": "10/29/2023 03:19:57",
          "content": "<p>Thanks for your advice. I'll try. Thanks again</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2503825,
      "author_name": "eliork",
      "author_url": "",
      "post_date": "10/29/2023 13:36:35",
      "content": "<p>Hello, when people refer to 'local cv,' they mean the average of the resulting MRRMSEs or any other chosen metric (such as MAE, MSE, Huber, etc.). I use cross-validation (cv) primarily for tuning my models and then train the best-tuned model with the optimal hyperparameters on random data splits. However, I'd like to point out that using cross-validation for the final selection of the best model may not be advisable in this context due to the potential for high uncertainty in the results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2503859,
          "author_name": "yanchengwei",
          "author_url": "",
          "post_date": "10/29/2023 14:13:59",
          "content": "<p>Thanks for your answer.I get it now</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2502947": "I'm a beginner in a Kaggle competition, and I have a question to ask. What does 'local CV' mean? Does it refer to the mean of MRRMSE in k-fold cross-validation or the last MRRMSE in k-fold cross-validation, or some other data?",
    "2502983": "I think it means what ever metric one uses for local cross validation and it's likely everyone is talking about different ones?",
    "2503002": "I thought there was default “local cv”, and it was my lack of understanding of default ”local cv” that made it unclear how everyone was measuring code.  Now I understand, thank you for your response.",
    "2503105": "Hi @yanchengwei I suggest not to use a simple k-fold, but a scheme like the one in [this notebook](https://www.kaggle.com/code/ambrosm/scp-quickstart). The simple k-fold isn't appropriate for simulating the task of predicting differential expression for an unseen cell type.",
    "2503370": "Thanks for your advice. I'll try. Thanks again",
    "2503825": "Hello, when people refer to 'local cv,' they mean the average of the resulting MRRMSEs or any other chosen metric (such as MAE, MSE, Huber, etc.). I use cross-validation (cv) primarily for tuning my models and then train the best-tuned model with the optimal hyperparameters on random data splits. However, I'd like to point out that using cross-validation for the final selection of the best model may not be advisable in this context due to the potential for high uncertainty in the results.",
    "2503859": "Thanks for your answer.I get it now"
  },
  "source": "meta"
}