{
  "id": 514769,
  "title": "Question about CV",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514769",
  "author_name": "",
  "post_date": "2024-06-25T12:26:37.702439100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys, I was reading some starter notebook today and I notice that they include the training data in their CV calculation. (basically calculate loss on the whole dataset, or using the train folds). Any explaination behind using training data to calculate CV?</p>\n<p>The notebooks I read (really helpful, thanks!):<br>\n<a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV\" target=\"_blank\">https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV</a><br>\n<a href=\"https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV\" target=\"_blank\">https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV</a></p>\n<p>Any help will be appreciated! :)</p>",
  "messages": [
    {
      "id": "2889319",
      "postDate": "06/25/2024 12:26:37",
      "content": "<p>Hi guys, I was reading some starter notebook today and I notice that they include the training data in their CV calculation. (basically calculate loss on the whole dataset, or using the train folds). Any explaination behind using training data to calculate CV?</p>\n<p>The notebooks I read (really helpful, thanks!):<br>\n<a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV\" target=\"_blank\">https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV</a><br>\n<a href=\"https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV\" target=\"_blank\">https://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV</a></p>\n<p>Any help will be appreciated! :)</p>",
      "rawMarkdown": "Hi guys, I was reading some starter notebook today and I notice that they include the training data in their CV calculation. (basically calculate loss on the whole dataset, or using the train folds). Any explaination behind using training data to calculate CV?\n\nThe notebooks I read (really helpful, thanks!):\nhttps://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV\nhttps://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV\n\nAny help will be appreciated! :)",
      "votes": null
    },
    {
      "id": "2889448",
      "postDate": "06/25/2024 13:53:27",
      "content": "<p>Are you asking <strong>why</strong> the notebooks are doing CV calculation on the entire or subset of train data? Then</p>\n<ul>\n<li>The loss that most models are trained on is different from the score function the competition calculated. So, to get an idea of what the approximate submission score would be, CV calculation is done.</li>\n<li>CV scores between different models can be compared to see whether the model performance is better. If this is not reflected in the LB scores, then we can see something is not adding up. E.g., there might be overfitting.</li>\n</ul>",
      "rawMarkdown": "Are you asking **why** the notebooks are doing CV calculation on the entire or subset of train data? Then\n- The loss that most models are trained on is different from the score function the competition calculated. So, to get an idea of what the approximate submission score would be, CV calculation is done.\n- CV scores between different models can be compared to see whether the model performance is better. If this is not reflected in the LB scores, then we can see something is not adding up. E.g., there might be overfitting.",
      "votes": null
    },
    {
      "id": "2889609",
      "postDate": "06/25/2024 15:26:15",
      "content": "<p>That's exactly what I'm looking for. Huge thanks!</p>",
      "rawMarkdown": "That's exactly what I'm looking for. Huge thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2889448,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/25/2024 13:53:27",
      "content": "<p>Are you asking <strong>why</strong> the notebooks are doing CV calculation on the entire or subset of train data? Then</p>\n<ul>\n<li>The loss that most models are trained on is different from the score function the competition calculated. So, to get an idea of what the approximate submission score would be, CV calculation is done.</li>\n<li>CV scores between different models can be compared to see whether the model performance is better. If this is not reflected in the LB scores, then we can see something is not adding up. E.g., there might be overfitting.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2889609,
          "author_name": "sakurayuyuko",
          "author_url": "",
          "post_date": "06/25/2024 15:26:15",
          "content": "<p>That's exactly what I'm looking for. Huge thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2889319": "Hi guys, I was reading some starter notebook today and I notice that they include the training data in their CV calculation. (basically calculate loss on the whole dataset, or using the train folds). Any explaination behind using training data to calculate CV?\n\nThe notebooks I read (really helpful, thanks!):\nhttps://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline/notebook#Calculation-CV\nhttps://www.kaggle.com/code/hugowjd/rsna2024-lsdc-training-densenet#Calculation-CV\n\nAny help will be appreciated! :)",
    "2889448": "Are you asking **why** the notebooks are doing CV calculation on the entire or subset of train data? Then\n- The loss that most models are trained on is different from the score function the competition calculated. So, to get an idea of what the approximate submission score would be, CV calculation is done.\n- CV scores between different models can be compared to see whether the model performance is better. If this is not reflected in the LB scores, then we can see something is not adding up. E.g., there might be overfitting.",
    "2889609": "That's exactly what I'm looking for. Huge thanks!"
  },
  "source": "meta"
}