{
  "id": 551850,
  "title": "What will be the final result?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551850",
  "author_name": "",
  "post_date": "2024-12-16T02:37:59.321266200Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>There have been many discussions on the topic of low correlation between LB scores and local CV scores. In my task, I synthesized a total of 4 methods, including different model selection and feature processing, Their current Optimized QWK scores are 0.453, 0.617, 0.613and 0.642, respectively, but this version of the notebook only obtains a LB SCORE of 0.404, which makes me confused and I hope you can help me solve it😭</p>",
  "messages": [
    {
      "id": "3073076",
      "postDate": "12/16/2024 02:37:59",
      "content": "<p>There have been many discussions on the topic of low correlation between LB scores and local CV scores. In my task, I synthesized a total of 4 methods, including different model selection and feature processing, Their current Optimized QWK scores are 0.453, 0.617, 0.613and 0.642, respectively, but this version of the notebook only obtains a LB SCORE of 0.404, which makes me confused and I hope you can help me solve it😭</p>",
      "rawMarkdown": "There have been many discussions on the topic of low correlation between LB scores and local CV scores. In my task, I synthesized a total of 4 methods, including different model selection and feature processing, Their current Optimized QWK scores are 0.453, 0.617, 0.613and 0.642, respectively, but this version of the notebook only obtains a LB SCORE of 0.404, which makes me confused and I hope you can help me solve it😭",
      "votes": null
    },
    {
      "id": "3073092",
      "postDate": "12/16/2024 03:29:00",
      "content": "<p>I also have such confusion. In my tests, although the validation scores of the publicly available notebooks are very high after optimization, the LB is very low. For example, CV: 0.550 LB: 0.344. Could this be a problem with its KNN imputation? Perhaps it's overfitting? Maybe it's an issue with the optimizer? I'm not very clear on this either.</p>",
      "rawMarkdown": "I also have such confusion. In my tests, although the validation scores of the publicly available notebooks are very high after optimization, the LB is very low. For example, CV: 0.550 LB: 0.344. Could this be a problem with its KNN imputation? Perhaps it's overfitting? Maybe it's an issue with the optimizer? I'm not very clear on this either.",
      "votes": null
    },
    {
      "id": "3073100",
      "postDate": "12/16/2024 04:05:23",
      "content": "<p>Computing cv on imputed data is unreasonable in general. The imputed data contains information from the original data, so the cv is obviously higher, especially when you fail to keep them separate in your kfold strategy.</p>",
      "rawMarkdown": "Computing cv on imputed data is unreasonable in general. The imputed data contains information from the original data, so the cv is obviously higher, especially when you fail to keep them separate in your kfold strategy.",
      "votes": null
    },
    {
      "id": "3073143",
      "postDate": "12/16/2024 05:24:28",
      "content": "<p>This will be revealed in the next few days - stay patient for now and enjoy the churn 3 days later <a href=\"https://www.kaggle.com/bloodthristy\" target=\"_blank\">@bloodthristy</a>!</p>",
      "rawMarkdown": "This will be revealed in the next few days - stay patient for now and enjoy the churn 3 days later @bloodthristy!",
      "votes": null
    },
    {
      "id": "3073190",
      "postDate": "12/16/2024 06:39:06",
      "content": "<p>I also had one such CV - LB result 2 months earlier in my baseline work. I think such a CV-LB relation is not stable and worthy. </p>",
      "rawMarkdown": "I also had one such CV - LB result 2 months earlier in my baseline work. I think such a CV-LB relation is not stable and worthy.",
      "votes": null
    },
    {
      "id": "3073501",
      "postDate": "12/16/2024 14:56:13",
      "content": "<p>I think you had so much fun during this competition 😄</p>",
      "rawMarkdown": "I think you had so much fun during this competition 😄",
      "votes": null
    },
    {
      "id": "3073586",
      "postDate": "12/16/2024 16:16:35",
      "content": "<p>In the early EDA, folks figured out that data is very noisy. Even manually filled by luck in some cases, targets as well. I think the most honest approach is using main data only, no imputed target, and no features with many NaNs; the rest of the low-NaN features are filled by interpolation. After that, fit a few very simple, regularized models: 1st submission is a solo model, the best of them; the 2nd is a careful ensemble. Regarding the thresholds, I think it's needed to produce global maximums, and some kind of post-processing may be appropriate.</p>",
      "rawMarkdown": "In the early EDA, folks figured out that data is very noisy. Even manually filled by luck in some cases, targets as well. I think the most honest approach is using main data only, no imputed target, and no features with many NaNs; the rest of the low-NaN features are filled by interpolation. After that, fit a few very simple, regularized models: 1st submission is a solo model, the best of them; the 2nd is a careful ensemble. Regarding the thresholds, I think it's needed to produce global maximums, and some kind of post-processing may be appropriate.",
      "votes": null
    },
    {
      "id": "3076274",
      "postDate": "12/19/2024 20:04:57",
      "content": "<p>4 hours to go now 😰   From past experience, do you have any idea how long it will take for the Private LB results to be posted? I imagine there's a lot of processing to do first…</p>",
      "rawMarkdown": "4 hours to go now 😰   From past experience, do you have any idea how long it will take for the Private LB results to be posted? I imagine there's a lot of processing to do first...",
      "votes": null
    },
    {
      "id": "3076293",
      "postDate": "12/19/2024 20:43:10",
      "content": "<p>Lately, preliminary results appear immediately; the verification, of course, may take longer. But I can’t say anything about today.</p>",
      "rawMarkdown": "Lately, preliminary results appear immediately; the verification, of course, may take longer. But I can’t say anything about today.",
      "votes": null
    },
    {
      "id": "3076336",
      "postDate": "12/19/2024 22:25:33",
      "content": "<p>Every submission we make is included in the private set. Verification might take a bit longer, but our results have already been calculated beforehand. Good luck! <a href=\"https://www.kaggle.com/dan3dewey\" target=\"_blank\">@dan3dewey</a> </p>",
      "rawMarkdown": "Every submission we make is included in the private set. Verification might take a bit longer, but our results have already been calculated beforehand. Good luck! @dan3dewey",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3073092,
      "author_name": "passionfruit216",
      "author_url": "",
      "post_date": "12/16/2024 03:29:00",
      "content": "<p>I also have such confusion. In my tests, although the validation scores of the publicly available notebooks are very high after optimization, the LB is very low. For example, CV: 0.550 LB: 0.344. Could this be a problem with its KNN imputation? Perhaps it's overfitting? Maybe it's an issue with the optimizer? I'm not very clear on this either.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3073190,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/16/2024 06:39:06",
          "content": "<p>I also had one such CV - LB result 2 months earlier in my baseline work. I think such a CV-LB relation is not stable and worthy. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3073100,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "12/16/2024 04:05:23",
      "content": "<p>Computing cv on imputed data is unreasonable in general. The imputed data contains information from the original data, so the cv is obviously higher, especially when you fail to keep them separate in your kfold strategy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073143,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/16/2024 05:24:28",
      "content": "<p>This will be revealed in the next few days - stay patient for now and enjoy the churn 3 days later <a href=\"https://www.kaggle.com/bloodthristy\" target=\"_blank\">@bloodthristy</a>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3073501,
          "author_name": "trcnveli",
          "author_url": "",
          "post_date": "12/16/2024 14:56:13",
          "content": "<p>I think you had so much fun during this competition 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3076274,
          "author_name": "dan3dewey",
          "author_url": "",
          "post_date": "12/19/2024 20:04:57",
          "content": "<p>4 hours to go now 😰   From past experience, do you have any idea how long it will take for the Private LB results to be posted? I imagine there's a lot of processing to do first…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3076293,
              "author_name": "zaakciiru",
              "author_url": "",
              "post_date": "12/19/2024 20:43:10",
              "content": "<p>Lately, preliminary results appear immediately; the verification, of course, may take longer. But I can’t say anything about today.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3076336,
              "author_name": "trcnveli",
              "author_url": "",
              "post_date": "12/19/2024 22:25:33",
              "content": "<p>Every submission we make is included in the private set. Verification might take a bit longer, but our results have already been calculated beforehand. Good luck! <a href=\"https://www.kaggle.com/dan3dewey\" target=\"_blank\">@dan3dewey</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3073586,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "12/16/2024 16:16:35",
      "content": "<p>In the early EDA, folks figured out that data is very noisy. Even manually filled by luck in some cases, targets as well. I think the most honest approach is using main data only, no imputed target, and no features with many NaNs; the rest of the low-NaN features are filled by interpolation. After that, fit a few very simple, regularized models: 1st submission is a solo model, the best of them; the 2nd is a careful ensemble. Regarding the thresholds, I think it's needed to produce global maximums, and some kind of post-processing may be appropriate.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3073076": "There have been many discussions on the topic of low correlation between LB scores and local CV scores. In my task, I synthesized a total of 4 methods, including different model selection and feature processing, Their current Optimized QWK scores are 0.453, 0.617, 0.613and 0.642, respectively, but this version of the notebook only obtains a LB SCORE of 0.404, which makes me confused and I hope you can help me solve it😭",
    "3073092": "I also have such confusion. In my tests, although the validation scores of the publicly available notebooks are very high after optimization, the LB is very low. For example, CV: 0.550 LB: 0.344. Could this be a problem with its KNN imputation? Perhaps it's overfitting? Maybe it's an issue with the optimizer? I'm not very clear on this either.",
    "3073100": "Computing cv on imputed data is unreasonable in general. The imputed data contains information from the original data, so the cv is obviously higher, especially when you fail to keep them separate in your kfold strategy.",
    "3073143": "This will be revealed in the next few days - stay patient for now and enjoy the churn 3 days later @bloodthristy!",
    "3073190": "I also had one such CV - LB result 2 months earlier in my baseline work. I think such a CV-LB relation is not stable and worthy.",
    "3073501": "I think you had so much fun during this competition 😄",
    "3073586": "In the early EDA, folks figured out that data is very noisy. Even manually filled by luck in some cases, targets as well. I think the most honest approach is using main data only, no imputed target, and no features with many NaNs; the rest of the low-NaN features are filled by interpolation. After that, fit a few very simple, regularized models: 1st submission is a solo model, the best of them; the 2nd is a careful ensemble. Regarding the thresholds, I think it's needed to produce global maximums, and some kind of post-processing may be appropriate.",
    "3076274": "4 hours to go now 😰   From past experience, do you have any idea how long it will take for the Private LB results to be posted? I imagine there's a lot of processing to do first...",
    "3076293": "Lately, preliminary results appear immediately; the verification, of course, may take longer. But I can’t say anything about today.",
    "3076336": "Every submission we make is included in the private set. Verification might take a bit longer, but our results have already been calculated beforehand. Good luck! @dan3dewey"
  },
  "source": "meta"
}