{
  "id": 547260,
  "title": "Relationship between CV score and LB score",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/547260",
  "author_name": "",
  "post_date": "2024-11-20T15:05:26.315582500Z",
  "votes": 5,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I am performing preprocessing and tuning of hyperparameters to improve model performance based on <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">this code</a>.<br>\nHowever, as the optimized QWK evaluated by CV improves, the LB score decreases. In other words, an inversely proportional phenomenon is observed.We did not perform any preprocessing that could have altered the data distributions of CV and LB.<br>\nIs anyone facing a similar problem?</p>",
  "messages": [
    {
      "id": "3050809",
      "postDate": "11/20/2024 15:05:26",
      "content": "<p>I am performing preprocessing and tuning of hyperparameters to improve model performance based on <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model\" target=\"_blank\">this code</a>.<br>\nHowever, as the optimized QWK evaluated by CV improves, the LB score decreases. In other words, an inversely proportional phenomenon is observed.We did not perform any preprocessing that could have altered the data distributions of CV and LB.<br>\nIs anyone facing a similar problem?</p>",
      "rawMarkdown": "I am performing preprocessing and tuning of hyperparameters to improve model performance based on [this code](https://www.kaggle.com/code/abdmental01/cmi-best-single-model).\nHowever, as the optimized QWK evaluated by CV improves, the LB score decreases. In other words, an inversely proportional phenomenon is observed.We did not perform any preprocessing that could have altered the data distributions of CV and LB.\nIs anyone facing a similar problem?",
      "votes": null
    },
    {
      "id": "3051438",
      "postDate": "11/21/2024 09:44:46",
      "content": "<p>Same here. It seems like the kernel used here yields an optimistically high LB score compared to any other LGBM model. In my opinion, the original LB score might not be reliable.</p>",
      "rawMarkdown": "Same here. It seems like the kernel used here yields an optimistically high LB score compared to any other LGBM model. In my opinion, the original LB score might not be reliable.",
      "votes": null
    },
    {
      "id": "3051442",
      "postDate": "11/21/2024 09:50:41",
      "content": "<p><a href=\"https://www.kaggle.com/theerawitplukmontol\" target=\"_blank\">@theerawitplukmontol</a> <br>\nThank you for your comment. I think that the data currently used to calculate the LB score may have a large distribution deviation from the training data. What do you think that?</p>",
      "rawMarkdown": "theerawitplukmontol \nThank you for your comment. I think that the data currently used to calculate the LB score may have a large distribution deviation from the training data. What do you think that?",
      "votes": null
    },
    {
      "id": "3052659",
      "postDate": "11/22/2024 17:06:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuogawa\" target=\"_blank\">@yuogawa</a>~</p>\n<p>Here are some discussion about the CV-LB relationship: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523\" target=\"_blank\">CV-LB thread</a>, <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535681\" target=\"_blank\">Handle your leaderboard scores with care!</a>, <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/545096\" target=\"_blank\">CV - LB relationship</a>.</p>\n<p>Just for reference. In summary, there seems to be <strong>no</strong> relationship between CV &amp; LB in this competition. Usually, people trust CV more than LB. Good luck~</p>",
      "rawMarkdown": "Hi @yuogawa~\n\nHere are some discussion about the CV-LB relationship: [CV-LB thread](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523), [Handle your leaderboard scores with care!](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535681), [CV - LB relationship](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/545096).\n\nJust for reference. In summary, there seems to be **no** relationship between CV & LB in this competition. Usually, people trust CV more than LB. Good luck~",
      "votes": null
    },
    {
      "id": "3053111",
      "postDate": "11/23/2024 06:50:14",
      "content": "<p>Thank you! I will try my best while taking these discussions into consideration.</p>",
      "rawMarkdown": "Thank you! I will try my best while taking these discussions into consideration.",
      "votes": null
    },
    {
      "id": "3053508",
      "postDate": "11/23/2024 15:09:26",
      "content": "<p>CV 0.484 LB 0.465 - 0.471. It is hard to say what going on. Maybe test data are much different than train one. </p>",
      "rawMarkdown": "CV 0.484 LB 0.465 - 0.471. It is hard to say what going on. Maybe test data are much different than train one.",
      "votes": null
    },
    {
      "id": "3053513",
      "postDate": "11/23/2024 15:17:05",
      "content": "<p>Your cv score seems very good!<br>\nDo you use any algorithms for imputing missing values?</p>",
      "rawMarkdown": "Your cv score seems very good!\nDo you use any algorithms for imputing missing values?",
      "votes": null
    },
    {
      "id": "3053525",
      "postDate": "11/23/2024 15:39:13",
      "content": "<p>I skip them. </p>",
      "rawMarkdown": "I skip them.",
      "votes": null
    },
    {
      "id": "3053873",
      "postDate": "11/24/2024 04:06:04",
      "content": "<p>It looks like the Public LB samples used are all existing in the training set, that means if you find them and match in your submission.csv you would get 1.000 on the LB. Unless kaggle is not using submission.csv for the public LB?</p>",
      "rawMarkdown": "It looks like the Public LB samples used are all existing in the training set, that means if you find them and match in your submission.csv you would get 1.000 on the LB. Unless kaggle is not using submission.csv for the public LB?",
      "votes": null
    },
    {
      "id": "3053876",
      "postDate": "11/24/2024 04:12:54",
      "content": "<p><a href=\"https://www.kaggle.com/lawrencechernin\" target=\"_blank\">@lawrencechernin</a> <br>\nThank you for your valuable comment!</p>",
      "rawMarkdown": "lawrencechernin \nThank you for your valuable comment!",
      "votes": null
    },
    {
      "id": "3053880",
      "postDate": "11/24/2024 04:22:31",
      "content": "<p>Actually, it seems more complicated… There are 9 ids in the test.csv are not found in the training.csv. For those I assign a value of zero and then when I submit this, the score is 0.000 because QWK heavily penalizes wrong guesses.  But one can try find these 9 by guessing iterating. I made my script public here: <a href=\"https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set\" target=\"_blank\">https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set</a>? scriptVersionId=209283327  Better is just to try climb the public leaderboard by exact fitting the 11 as they are in train, and using a model for the other 9. What it means the public LB for this competition is pretty much useless.</p>",
      "rawMarkdown": "Actually, it seems more complicated... There are 9 ids in the test.csv are not found in the training.csv. For those I assign a value of zero and then when I submit this, the score is 0.000 because QWK heavily penalizes wrong guesses.  But one can try find these 9 by guessing iterating. I made my script public here: https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set? scriptVersionId=209283327  Better is just to try climb the public leaderboard by exact fitting the 11 as they are in train, and using a model for the other 9. What it means the public LB for this competition is pretty much useless.",
      "votes": null
    },
    {
      "id": "3054285",
      "postDate": "11/24/2024 14:25:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lawrencechernin\" target=\"_blank\">@lawrencechernin</a> ,<br>\nI believe you are making a misunderstanding : kaggle is not using the 20 test samples for the public LB, or maybe Kaggle is using them but with 38% of 3800 others. </p>\n<p>In <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\" target=\"_blank\"><strong>Data</strong> page</a> :</p>\n<blockquote>\n  <p>Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.</p>\n</blockquote>",
      "rawMarkdown": "Hi @lawrencechernin ,\nI believe you are making a misunderstanding : kaggle is not using the 20 test samples for the public LB, or maybe Kaggle is using them but with 38% of 3800 others. \n\nIn [**Data** page](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data) :\n>Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.",
      "votes": null
    },
    {
      "id": "3054527",
      "postDate": "11/24/2024 19:06:13",
      "content": "<p>Yes, I saw that also, but then what is the purpose of the test set that they provide?</p>",
      "rawMarkdown": "Yes, I saw that also, but then what is the purpose of the test set that they provide?",
      "votes": null
    },
    {
      "id": "3054529",
      "postDate": "11/24/2024 19:10:19",
      "content": "<p>To do a sample prediction. This is only for a coding reason. </p>",
      "rawMarkdown": "To do a sample prediction. This is only for a coding reason.",
      "votes": null
    },
    {
      "id": "3056345",
      "postDate": "11/26/2024 21:18:06",
      "content": "<p>Also my single model Cv = .460 and my ensemble model' Lb Score is .461. I think it can be good for private data. Right?</p>",
      "rawMarkdown": "Also my single model Cv = .460 and my ensemble model' Lb Score is .461. I think it can be good for private data. Right?",
      "votes": null
    },
    {
      "id": "3056371",
      "postDate": "11/26/2024 22:32:57",
      "content": "<p><a href=\"https://www.kaggle.com/trcnveli\" target=\"_blank\">@trcnveli</a> <br>\nThank you for your comment!<br>\nI think it will be good result although i don’t know how score other people are.<br>\nIs your CV reproducible for different seed values? <br>\nIn my case,  varies by up to ±0.05.</p>",
      "rawMarkdown": "trcnveli \nThank you for your comment!\nI think it will be good result although i don’t know how score other people are.\nIs your CV reproducible for different seed values? \nIn my case,  varies by up to ±0.05.",
      "votes": null
    },
    {
      "id": "3057924",
      "postDate": "11/28/2024 20:20:38",
      "content": "<p>Yes. My Cv reproducible for different seed values and I am just wondering it will be so overfit in private set or not. I think, public dataset and train dataset is similar but private is different from them. We will see.</p>",
      "rawMarkdown": "Yes. My Cv reproducible for different seed values and I am just wondering it will be so overfit in private set or not. I think, public dataset and train dataset is similar but private is different from them. We will see.",
      "votes": null
    },
    {
      "id": "3058024",
      "postDate": "11/28/2024 22:22:49",
      "content": "<p><a href=\"https://www.kaggle.com/trcnveli\" target=\"_blank\">@trcnveli</a> <br>\nIn my case, I have a different situation of yours. My CV is about 0.464, but the LB score drops into 0.44. So I doubt if there is a different dispersions between train and public data.<br>\nIs this difference between you and me a merely <strong>luck</strong>? Or are there  any essentional ones? </p>",
      "rawMarkdown": "trcnveli \nIn my case, I have a different situation of yours. My CV is about 0.464, but the LB score drops into 0.44. So I doubt if there is a different dispersions between train and public data.\nIs this difference between you and me a merely **luck**? Or are there  any essentional ones?",
      "votes": null
    },
    {
      "id": "3058793",
      "postDate": "11/30/2024 01:49:10",
      "content": "<p>I think the train and public test sets are quite similar. For now, it seems like neither of us has any issues. We'll see what happens with the private test set. Good luck!</p>",
      "rawMarkdown": "I think the train and public test sets are quite similar. For now, it seems like neither of us has any issues. We'll see what happens with the private test set. Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3051438,
      "author_name": "theerawitplukmontol",
      "author_url": "",
      "post_date": "11/21/2024 09:44:46",
      "content": "<p>Same here. It seems like the kernel used here yields an optimistically high LB score compared to any other LGBM model. In my opinion, the original LB score might not be reliable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3051442,
          "author_name": "yuogawa",
          "author_url": "",
          "post_date": "11/21/2024 09:50:41",
          "content": "<p><a href=\"https://www.kaggle.com/theerawitplukmontol\" target=\"_blank\">@theerawitplukmontol</a> <br>\nThank you for your comment. I think that the data currently used to calculate the LB score may have a large distribution deviation from the training data. What do you think that?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3052659,
      "author_name": "fangzitao",
      "author_url": "",
      "post_date": "11/22/2024 17:06:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuogawa\" target=\"_blank\">@yuogawa</a>~</p>\n<p>Here are some discussion about the CV-LB relationship: <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523\" target=\"_blank\">CV-LB thread</a>, <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535681\" target=\"_blank\">Handle your leaderboard scores with care!</a>, <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/545096\" target=\"_blank\">CV - LB relationship</a>.</p>\n<p>Just for reference. In summary, there seems to be <strong>no</strong> relationship between CV &amp; LB in this competition. Usually, people trust CV more than LB. Good luck~</p>",
      "votes": null,
      "replies": [
        {
          "id": 3053111,
          "author_name": "yuogawa",
          "author_url": "",
          "post_date": "11/23/2024 06:50:14",
          "content": "<p>Thank you! I will try my best while taking these discussions into consideration.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3053508,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "11/23/2024 15:09:26",
      "content": "<p>CV 0.484 LB 0.465 - 0.471. It is hard to say what going on. Maybe test data are much different than train one. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3053513,
          "author_name": "yuogawa",
          "author_url": "",
          "post_date": "11/23/2024 15:17:05",
          "content": "<p>Your cv score seems very good!<br>\nDo you use any algorithms for imputing missing values?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3053525,
              "author_name": "jankowalski2000",
              "author_url": "",
              "post_date": "11/23/2024 15:39:13",
              "content": "<p>I skip them. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3053873,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "11/24/2024 04:06:04",
      "content": "<p>It looks like the Public LB samples used are all existing in the training set, that means if you find them and match in your submission.csv you would get 1.000 on the LB. Unless kaggle is not using submission.csv for the public LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3053876,
          "author_name": "yuogawa",
          "author_url": "",
          "post_date": "11/24/2024 04:12:54",
          "content": "<p><a href=\"https://www.kaggle.com/lawrencechernin\" target=\"_blank\">@lawrencechernin</a> <br>\nThank you for your valuable comment!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3053880,
              "author_name": "lawrencechernin",
              "author_url": "",
              "post_date": "11/24/2024 04:22:31",
              "content": "<p>Actually, it seems more complicated… There are 9 ids in the test.csv are not found in the training.csv. For those I assign a value of zero and then when I submit this, the score is 0.000 because QWK heavily penalizes wrong guesses.  But one can try find these 9 by guessing iterating. I made my script public here: <a href=\"https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set\" target=\"_blank\">https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set</a>? scriptVersionId=209283327  Better is just to try climb the public leaderboard by exact fitting the 11 as they are in train, and using a model for the other 9. What it means the public LB for this competition is pretty much useless.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3054285,
          "author_name": "adaubas",
          "author_url": "",
          "post_date": "11/24/2024 14:25:24",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lawrencechernin\" target=\"_blank\">@lawrencechernin</a> ,<br>\nI believe you are making a misunderstanding : kaggle is not using the 20 test samples for the public LB, or maybe Kaggle is using them but with 38% of 3800 others. </p>\n<p>In <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data\" target=\"_blank\"><strong>Data</strong> page</a> :</p>\n<blockquote>\n  <p>Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.</p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 3054527,
              "author_name": "lawrencechernin",
              "author_url": "",
              "post_date": "11/24/2024 19:06:13",
              "content": "<p>Yes, I saw that also, but then what is the purpose of the test set that they provide?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3054529,
                  "author_name": "jankowalski2000",
                  "author_url": "",
                  "post_date": "11/24/2024 19:10:19",
                  "content": "<p>To do a sample prediction. This is only for a coding reason. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3056345,
      "author_name": "trcnveli",
      "author_url": "",
      "post_date": "11/26/2024 21:18:06",
      "content": "<p>Also my single model Cv = .460 and my ensemble model' Lb Score is .461. I think it can be good for private data. Right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3056371,
          "author_name": "yuogawa",
          "author_url": "",
          "post_date": "11/26/2024 22:32:57",
          "content": "<p><a href=\"https://www.kaggle.com/trcnveli\" target=\"_blank\">@trcnveli</a> <br>\nThank you for your comment!<br>\nI think it will be good result although i don’t know how score other people are.<br>\nIs your CV reproducible for different seed values? <br>\nIn my case,  varies by up to ±0.05.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3057924,
              "author_name": "trcnveli",
              "author_url": "",
              "post_date": "11/28/2024 20:20:38",
              "content": "<p>Yes. My Cv reproducible for different seed values and I am just wondering it will be so overfit in private set or not. I think, public dataset and train dataset is similar but private is different from them. We will see.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3058024,
                  "author_name": "yuogawa",
                  "author_url": "",
                  "post_date": "11/28/2024 22:22:49",
                  "content": "<p><a href=\"https://www.kaggle.com/trcnveli\" target=\"_blank\">@trcnveli</a> <br>\nIn my case, I have a different situation of yours. My CV is about 0.464, but the LB score drops into 0.44. So I doubt if there is a different dispersions between train and public data.<br>\nIs this difference between you and me a merely <strong>luck</strong>? Or are there  any essentional ones? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3058793,
                      "author_name": "trcnveli",
                      "author_url": "",
                      "post_date": "11/30/2024 01:49:10",
                      "content": "<p>I think the train and public test sets are quite similar. For now, it seems like neither of us has any issues. We'll see what happens with the private test set. Good luck!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3050809": "I am performing preprocessing and tuning of hyperparameters to improve model performance based on [this code](https://www.kaggle.com/code/abdmental01/cmi-best-single-model).\nHowever, as the optimized QWK evaluated by CV improves, the LB score decreases. In other words, an inversely proportional phenomenon is observed.We did not perform any preprocessing that could have altered the data distributions of CV and LB.\nIs anyone facing a similar problem?",
    "3051438": "Same here. It seems like the kernel used here yields an optimistically high LB score compared to any other LGBM model. In my opinion, the original LB score might not be reliable.",
    "3051442": "theerawitplukmontol \nThank you for your comment. I think that the data currently used to calculate the LB score may have a large distribution deviation from the training data. What do you think that?",
    "3052659": "Hi @yuogawa~\n\nHere are some discussion about the CV-LB relationship: [CV-LB thread](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535523), [Handle your leaderboard scores with care!](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535681), [CV - LB relationship](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/545096).\n\nJust for reference. In summary, there seems to be **no** relationship between CV & LB in this competition. Usually, people trust CV more than LB. Good luck~",
    "3053111": "Thank you! I will try my best while taking these discussions into consideration.",
    "3053508": "CV 0.484 LB 0.465 - 0.471. It is hard to say what going on. Maybe test data are much different than train one.",
    "3053513": "Your cv score seems very good!\nDo you use any algorithms for imputing missing values?",
    "3053525": "I skip them.",
    "3053873": "It looks like the Public LB samples used are all existing in the training set, that means if you find them and match in your submission.csv you would get 1.000 on the LB. Unless kaggle is not using submission.csv for the public LB?",
    "3053876": "lawrencechernin \nThank you for your valuable comment!",
    "3053880": "Actually, it seems more complicated... There are 9 ids in the test.csv are not found in the training.csv. For those I assign a value of zero and then when I submit this, the score is 0.000 because QWK heavily penalizes wrong guesses.  But one can try find these 9 by guessing iterating. I made my script public here: https://www.kaggle.com/code/lawrencechernin/find-11-test-samples-in-the-train-set? scriptVersionId=209283327  Better is just to try climb the public leaderboard by exact fitting the 11 as they are in train, and using a model for the other 9. What it means the public LB for this competition is pretty much useless.",
    "3054285": "Hi @lawrencechernin ,\nI believe you are making a misunderstanding : kaggle is not using the 20 test samples for the public LB, or maybe Kaggle is using them but with 38% of 3800 others. \n\nIn [**Data** page](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data) :\n>Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.",
    "3054527": "Yes, I saw that also, but then what is the purpose of the test set that they provide?",
    "3054529": "To do a sample prediction. This is only for a coding reason.",
    "3056345": "Also my single model Cv = .460 and my ensemble model' Lb Score is .461. I think it can be good for private data. Right?",
    "3056371": "trcnveli \nThank you for your comment!\nI think it will be good result although i don’t know how score other people are.\nIs your CV reproducible for different seed values? \nIn my case,  varies by up to ±0.05.",
    "3057924": "Yes. My Cv reproducible for different seed values and I am just wondering it will be so overfit in private set or not. I think, public dataset and train dataset is similar but private is different from them. We will see.",
    "3058024": "trcnveli \nIn my case, I have a different situation of yours. My CV is about 0.464, but the LB score drops into 0.44. So I doubt if there is a different dispersions between train and public data.\nIs this difference between you and me a merely **luck**? Or are there  any essentional ones?",
    "3058793": "I think the train and public test sets are quite similar. For now, it seems like neither of us has any issues. We'll see what happens with the private test set. Good luck!"
  },
  "source": "meta"
}