{
  "id": 535458,
  "title": "Can we trust CV？",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535458",
  "author_name": "",
  "post_date": "2024-09-22T08:40:42.835248600Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I noticed that ：</p>\n<p>Mean Train QWK --&gt; 0.7151  <br>\nMean Validation QWK SCORE: 0.452, <br>\nLB: 0.455  <br>\n——————————————————<br>\nMean Train QWK --&gt; 0.6785, <br>\nMean Validation QWK SCORE: 0.462, <br>\nLB: 0.452  </p>\n<p>This indicates that the scores between online (LB) and offline (cross-validation) are inconsistent. Can we trust the offline CV score? Or is there a better way to approach this?<br>\nAlso, I noticed that this data seems to be quite sensitive to parameter changes. Did you observe this as well?</p>",
  "messages": [
    {
      "id": "2995366",
      "postDate": "09/22/2024 08:40:42",
      "content": "<p>I noticed that ：</p>\n<p>Mean Train QWK --&gt; 0.7151  <br>\nMean Validation QWK SCORE: 0.452, <br>\nLB: 0.455  <br>\n——————————————————<br>\nMean Train QWK --&gt; 0.6785, <br>\nMean Validation QWK SCORE: 0.462, <br>\nLB: 0.452  </p>\n<p>This indicates that the scores between online (LB) and offline (cross-validation) are inconsistent. Can we trust the offline CV score? Or is there a better way to approach this?<br>\nAlso, I noticed that this data seems to be quite sensitive to parameter changes. Did you observe this as well?</p>",
      "rawMarkdown": "I noticed that ：\n\nMean Train QWK --> 0.7151  \nMean Validation QWK SCORE: 0.452, \nLB: 0.455  \n——————————————————\nMean Train QWK --> 0.6785, \nMean Validation QWK SCORE: 0.462, \nLB: 0.452  \n\nThis indicates that the scores between online (LB) and offline (cross-validation) are inconsistent. Can we trust the offline CV score? Or is there a better way to approach this?\nAlso, I noticed that this data seems to be quite sensitive to parameter changes. Did you observe this as well?",
      "votes": null
    },
    {
      "id": "2995373",
      "postDate": "09/22/2024 08:51:46",
      "content": "<p>This is an observation with most competitions involving hard labels. Accuracy score, Kappa score, Matthews correlation, etc. are examples of evaluation metrics that require hard labels and even a couple of correct/ incorrect classifications could result in a sizable penalty <a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> </p>\n<p>Train CV score versus OOF CV is very spread out, (70% - 45%) this is also a concerning observation!</p>",
      "rawMarkdown": "This is an observation with most competitions involving hard labels. Accuracy score, Kappa score, Matthews correlation, etc. are examples of evaluation metrics that require hard labels and even a couple of correct/ incorrect classifications could result in a sizable penalty @wayne127 \n\nTrain CV score versus OOF CV is very spread out, (70% - 45%) this is also a concerning observation!",
      "votes": null
    },
    {
      "id": "2995386",
      "postDate": "09/22/2024 09:15:09",
      "content": "<p>CV: 0.452 LB: 0.455 | CV: 0.481 LB: 0.443 </p>",
      "rawMarkdown": "CV: 0.452 LB: 0.455 | CV: 0.481 LB: 0.443",
      "votes": null
    },
    {
      "id": "2995459",
      "postDate": "09/22/2024 10:51:28",
      "content": "<p>You should trust your CV if you are confident that your validation strategy is good and copies the test environment. Check overfitting by simplifying your model, regularizing, or using cross-validation techniques according to your data. Use the leaderboard for confirmation but don't over-optimize for it. Ensure your model generalizes well across multiple CV folds and isn't just fitting noise.</p>",
      "rawMarkdown": "You should trust your CV if you are confident that your validation strategy is good and copies the test environment. Check overfitting by simplifying your model, regularizing, or using cross-validation techniques according to your data. Use the leaderboard for confirmation but don't over-optimize for it. Ensure your model generalizes well across multiple CV folds and isn't just fitting noise.",
      "votes": null
    },
    {
      "id": "2995528",
      "postDate": "09/22/2024 11:59:06",
      "content": "<p>Okay, thank you very much. I think I will carefully consider this issue after completing data preprocessing and feature engineering to ensure that my model is not overly affected by noisy data. However, I still appreciate it very much!🥳</p>",
      "rawMarkdown": "Okay, thank you very much. I think I will carefully consider this issue after completing data preprocessing and feature engineering to ensure that my model is not overly affected by noisy data. However, I still appreciate it very much!🥳",
      "votes": null
    },
    {
      "id": "2995535",
      "postDate": "09/22/2024 12:01:24",
      "content": "<p>Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible</p>",
      "rawMarkdown": "Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible",
      "votes": null
    },
    {
      "id": "2995678",
      "postDate": "09/22/2024 14:15:47",
      "content": "<p>I believe you are referring to an optimized score with 'Mean Validation QWK Score: 0.452'. In my opinion, we should trust the cross-validation (CV) score more than the optimized score. From my experiments, I've observed cases where a model with a good optimized score, like 0.480, had a lower leaderboard (LB) score of only 0.449. On the other hand, my single LightGBM model had an optimized CV score of 0.462 and an LB score of 0.453.</p>\n<p>In both cases, the key factor to focus on is the normal CV score without optimization. For example, the single LightGBM model had a CV of 0.4011, while the second model had a CV of 0.3869.</p>\n<p>Therefore, if we look closely, the optimized score isn't the main factor we should focus on. This is my opinion.</p>",
      "rawMarkdown": "I believe you are referring to an optimized score with 'Mean Validation QWK Score: 0.452'. In my opinion, we should trust the cross-validation (CV) score more than the optimized score. From my experiments, I've observed cases where a model with a good optimized score, like 0.480, had a lower leaderboard (LB) score of only 0.449. On the other hand, my single LightGBM model had an optimized CV score of 0.462 and an LB score of 0.453.\n\nIn both cases, the key factor to focus on is the normal CV score without optimization. For example, the single LightGBM model had a CV of 0.4011, while the second model had a CV of 0.3869.\n\nTherefore, if we look closely, the optimized score isn't the main factor we should focus on. This is my opinion.",
      "votes": null
    },
    {
      "id": "2995744",
      "postDate": "09/22/2024 14:56:49",
      "content": "<p>A valuable suggestion, I think it's time to change our strategy</p>",
      "rawMarkdown": "A valuable suggestion, I think it's time to change our strategy",
      "votes": null
    },
    {
      "id": "3036536",
      "postDate": "11/04/2024 16:43:50",
      "content": "<p><a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a>  ' Can we trust CV' crucial question?</p>",
      "rawMarkdown": "wayne127  ' Can we trust CV' crucial question?",
      "votes": null
    },
    {
      "id": "3060496",
      "postDate": "12/01/2024 19:27:21",
      "content": "<p>I see such a large variation across folds in almost all the codes I have seen. It means the models are struggling to fit the data, and the Public Leaderboard likely is representing some highly optimized local solutions as QWK is very sensitive to a few outliers.  I also see this by small changes in seeds and hyperparameters. There will probably be a huge shakeup on the Private Leaderboard.   </p>\n<p>There is may also be some bad data, e.g. some parents giving bad results on the survey even though their children are fine. Maybe this is the best that can be done given that the labels are based on a survey. Perhaps if we had some other data set to help calibrate the parents?</p>",
      "rawMarkdown": "I see such a large variation across folds in almost all the codes I have seen. It means the models are struggling to fit the data, and the Public Leaderboard likely is representing some highly optimized local solutions as QWK is very sensitive to a few outliers.  I also see this by small changes in seeds and hyperparameters. There will probably be a huge shakeup on the Private Leaderboard.   \n\nThere is may also be some bad data, e.g. some parents giving bad results on the survey even though their children are fine. Maybe this is the best that can be done given that the labels are based on a survey. Perhaps if we had some other data set to help calibrate the parents?",
      "votes": null
    },
    {
      "id": "3060497",
      "postDate": "12/01/2024 19:29:34",
      "content": "<p><a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a> this is a luck competition, anything can happen here!</p>",
      "rawMarkdown": "saidkoussi this is a luck competition, anything can happen here!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2995373,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "09/22/2024 08:51:46",
      "content": "<p>This is an observation with most competitions involving hard labels. Accuracy score, Kappa score, Matthews correlation, etc. are examples of evaluation metrics that require hard labels and even a couple of correct/ incorrect classifications could result in a sizable penalty <a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> </p>\n<p>Train CV score versus OOF CV is very spread out, (70% - 45%) this is also a concerning observation!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995535,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "09/22/2024 12:01:24",
          "content": "<p>Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2995386,
      "author_name": "ucas0v0zhuoqunli",
      "author_url": "",
      "post_date": "09/22/2024 09:15:09",
      "content": "<p>CV: 0.452 LB: 0.455 | CV: 0.481 LB: 0.443 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2995459,
      "author_name": "manojajj",
      "author_url": "",
      "post_date": "09/22/2024 10:51:28",
      "content": "<p>You should trust your CV if you are confident that your validation strategy is good and copies the test environment. Check overfitting by simplifying your model, regularizing, or using cross-validation techniques according to your data. Use the leaderboard for confirmation but don't over-optimize for it. Ensure your model generalizes well across multiple CV folds and isn't just fitting noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995528,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "09/22/2024 11:59:06",
          "content": "<p>Okay, thank you very much. I think I will carefully consider this issue after completing data preprocessing and feature engineering to ensure that my model is not overly affected by noisy data. However, I still appreciate it very much!🥳</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2995678,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "09/22/2024 14:15:47",
      "content": "<p>I believe you are referring to an optimized score with 'Mean Validation QWK Score: 0.452'. In my opinion, we should trust the cross-validation (CV) score more than the optimized score. From my experiments, I've observed cases where a model with a good optimized score, like 0.480, had a lower leaderboard (LB) score of only 0.449. On the other hand, my single LightGBM model had an optimized CV score of 0.462 and an LB score of 0.453.</p>\n<p>In both cases, the key factor to focus on is the normal CV score without optimization. For example, the single LightGBM model had a CV of 0.4011, while the second model had a CV of 0.3869.</p>\n<p>Therefore, if we look closely, the optimized score isn't the main factor we should focus on. This is my opinion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995744,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "09/22/2024 14:56:49",
          "content": "<p>A valuable suggestion, I think it's time to change our strategy</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3036536,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "11/04/2024 16:43:50",
      "content": "<p><a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a>  ' Can we trust CV' crucial question?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3060497,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/01/2024 19:29:34",
          "content": "<p><a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a> this is a luck competition, anything can happen here!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3060496,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "12/01/2024 19:27:21",
      "content": "<p>I see such a large variation across folds in almost all the codes I have seen. It means the models are struggling to fit the data, and the Public Leaderboard likely is representing some highly optimized local solutions as QWK is very sensitive to a few outliers.  I also see this by small changes in seeds and hyperparameters. There will probably be a huge shakeup on the Private Leaderboard.   </p>\n<p>There is may also be some bad data, e.g. some parents giving bad results on the survey even though their children are fine. Maybe this is the best that can be done given that the labels are based on a survey. Perhaps if we had some other data set to help calibrate the parents?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2995366": "I noticed that ：\n\nMean Train QWK --> 0.7151  \nMean Validation QWK SCORE: 0.452, \nLB: 0.455  \n——————————————————\nMean Train QWK --> 0.6785, \nMean Validation QWK SCORE: 0.462, \nLB: 0.452  \n\nThis indicates that the scores between online (LB) and offline (cross-validation) are inconsistent. Can we trust the offline CV score? Or is there a better way to approach this?\nAlso, I noticed that this data seems to be quite sensitive to parameter changes. Did you observe this as well?",
    "2995373": "This is an observation with most competitions involving hard labels. Accuracy score, Kappa score, Matthews correlation, etc. are examples of evaluation metrics that require hard labels and even a couple of correct/ incorrect classifications could result in a sizable penalty @wayne127 \n\nTrain CV score versus OOF CV is very spread out, (70% - 45%) this is also a concerning observation!",
    "2995386": "CV: 0.452 LB: 0.455 | CV: 0.481 LB: 0.443",
    "2995459": "You should trust your CV if you are confident that your validation strategy is good and copies the test environment. Check overfitting by simplifying your model, regularizing, or using cross-validation techniques according to your data. Use the leaderboard for confirmation but don't over-optimize for it. Ensure your model generalizes well across multiple CV folds and isn't just fitting noise.",
    "2995528": "Okay, thank you very much. I think I will carefully consider this issue after completing data preprocessing and feature engineering to ensure that my model is not overly affected by noisy data. However, I still appreciate it very much!🥳",
    "2995535": "Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible",
    "2995678": "I believe you are referring to an optimized score with 'Mean Validation QWK Score: 0.452'. In my opinion, we should trust the cross-validation (CV) score more than the optimized score. From my experiments, I've observed cases where a model with a good optimized score, like 0.480, had a lower leaderboard (LB) score of only 0.449. On the other hand, my single LightGBM model had an optimized CV score of 0.462 and an LB score of 0.453.\n\nIn both cases, the key factor to focus on is the normal CV score without optimization. For example, the single LightGBM model had a CV of 0.4011, while the second model had a CV of 0.3869.\n\nTherefore, if we look closely, the optimized score isn't the main factor we should focus on. This is my opinion.",
    "2995744": "A valuable suggestion, I think it's time to change our strategy",
    "3036536": "wayne127  ' Can we trust CV' crucial question?",
    "3060496": "I see such a large variation across folds in almost all the codes I have seen. It means the models are struggling to fit the data, and the Public Leaderboard likely is representing some highly optimized local solutions as QWK is very sensitive to a few outliers.  I also see this by small changes in seeds and hyperparameters. There will probably be a huge shakeup on the Private Leaderboard.   \n\nThere is may also be some bad data, e.g. some parents giving bad results on the survey even though their children are fine. Maybe this is the best that can be done given that the labels are based on a survey. Perhaps if we had some other data set to help calibrate the parents?",
    "3060497": "saidkoussi this is a luck competition, anything can happen here!"
  },
  "source": "meta"
}