{
  "id": 551164,
  "title": "Guess the winner score of this competition?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551164",
  "author_name": "",
  "post_date": "2024-12-11T16:23:17.647361400Z",
  "votes": 12,
  "comment_count": 34,
  "views": 0,
  "content": "<p>I guess the winner (and maybe the gold medal score) is around 0.43-0.45. How about you?</p>",
  "messages": [
    {
      "id": "3069582",
      "postDate": "12/11/2024 16:23:17",
      "content": "<p>I guess the winner (and maybe the gold medal score) is around 0.43-0.45. How about you?</p>",
      "rawMarkdown": "I guess the winner (and maybe the gold medal score) is around 0.43-0.45. How about you?",
      "votes": null
    },
    {
      "id": "3069589",
      "postDate": "12/11/2024 16:26:36",
      "content": "<p>I just started this competition and it feels like ICR all over again. I guess it's easier to overfit to a smaller portion of test set so 0.43-0.45 sounds reasonable. I can get 0.49 oof score with a single model but I have no idea how it will translate to private test set.</p>",
      "rawMarkdown": "I just started this competition and it feels like ICR all over again. I guess it's easier to overfit to a smaller portion of test set so 0.43-0.45 sounds reasonable. I can get 0.49 oof score with a single model but I have no idea how it will translate to private test set.",
      "votes": null
    },
    {
      "id": "3069614",
      "postDate": "12/11/2024 16:47:23",
      "content": "<p>I think it will be around 0.455 - 0.46<br>\nIt is very easy to overfit here and a lot of luck is needed to prevent it. This is ICR part 2 in my opinion <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
      "rawMarkdown": "I think it will be around 0.455 - 0.46\nIt is very easy to overfit here and a lot of luck is needed to prevent it. This is ICR part 2 in my opinion @bibanh @gunesevitan",
      "votes": null
    },
    {
      "id": "3069620",
      "postDate": "12/11/2024 16:54:39",
      "content": "<p>Thanks for sharing feeling. <br>\nHow do you validate your model ?<br>\nIn my case, I use multiple runs of CV as validation and my CV has steadily improved, but my LB hasn't improved at all.</p>\n<p>I’m wondering why public notebooks with data leakage tend to have high LB scores.</p>",
      "rawMarkdown": "Thanks for sharing feeling. \nHow do you validate your model ?\nIn my case, I use multiple runs of CV as validation and my CV has steadily improved, but my LB hasn't improved at all.\n\nI’m wondering why public notebooks with data leakage tend to have high LB scores.",
      "votes": null
    },
    {
      "id": "3069621",
      "postDate": "12/11/2024 16:56:56",
      "content": "<p>I track lots of things like fold score mean, std and oof scores, and some proxy metrics. I haven't looked at public notebooks other than eda notebooks.</p>",
      "rawMarkdown": "I track lots of things like fold score mean, std and oof scores, and some proxy metrics. I haven't looked at public notebooks other than eda notebooks.",
      "votes": null
    },
    {
      "id": "3069624",
      "postDate": "12/11/2024 17:02:01",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> with due respect, I may say that most high scoring public notebooks are unlikely to score well on the private LB. Looking at them is a big risk</p>",
      "rawMarkdown": "gunesevitan with due respect, I may say that most high scoring public notebooks are unlikely to score well on the private LB. Looking at them is a big risk",
      "votes": null
    },
    {
      "id": "3069626",
      "postDate": "12/11/2024 17:04:46",
      "content": "<p>I agree due to fact that it's easier to overfit here compared to ICR because of thresholds. ICR was evaluated on log loss but QWK is a different story.</p>",
      "rawMarkdown": "I agree due to fact that it's easier to overfit here compared to ICR because of thresholds. ICR was evaluated on log loss but QWK is a different story.",
      "votes": null
    },
    {
      "id": "3069649",
      "postDate": "12/11/2024 17:41:45",
      "content": "<p>Is there any case that CV/public does not correlate but CV/private does correlate well ?</p>",
      "rawMarkdown": "Is there any case that CV/public does not correlate but CV/private does correlate well ?",
      "votes": null
    },
    {
      "id": "3069658",
      "postDate": "12/11/2024 18:04:33",
      "content": "<p>Somewhere around 0.476 is my guess.</p>",
      "rawMarkdown": "Somewhere around 0.476 is my guess.",
      "votes": null
    },
    {
      "id": "3069671",
      "postDate": "12/11/2024 18:31:39",
      "content": "<p>Probably not below 0.44.</p>",
      "rawMarkdown": "Probably not below 0.44.",
      "votes": null
    },
    {
      "id": "3069710",
      "postDate": "12/11/2024 19:28:48",
      "content": "<p>Oh it's too high kkkk</p>",
      "rawMarkdown": "Oh it's too high kkkk",
      "votes": null
    },
    {
      "id": "3069908",
      "postDate": "12/12/2024 03:15:48",
      "content": "<p>maybe around .420 will be in gold zone</p>",
      "rawMarkdown": "maybe around .420 will be in gold zone",
      "votes": null
    },
    {
      "id": "3069912",
      "postDate": "12/12/2024 03:27:12",
      "content": "<p>Same here, also just started two days ago. I guess I will pick the submission that I have the most confidence on it performing well on private LB, regardless of CV or public LB. It’s easy to overfit both of them due to how small and noisy the data is 😬</p>\n<p>Luck will be a huge factor I guess…</p>",
      "rawMarkdown": "Same here, also just started two days ago. I guess I will pick the submission that I have the most confidence on it performing well on private LB, regardless of CV or public LB. It’s easy to overfit both of them due to how small and noisy the data is 😬\n\nLuck will be a huge factor I guess…",
      "votes": null
    },
    {
      "id": "3070001",
      "postDate": "12/12/2024 05:51:54",
      "content": "<p>Yeah, most of the shakeup competitions.</p>",
      "rawMarkdown": "Yeah, most of the shakeup competitions.",
      "votes": null
    },
    {
      "id": "3070067",
      "postDate": "12/12/2024 07:42:56",
      "content": "<p>maybe around 0.43</p>",
      "rawMarkdown": "maybe around 0.43",
      "votes": null
    },
    {
      "id": "3070406",
      "postDate": "12/12/2024 16:38:56",
      "content": "<p>I agree with this. Something like finetuning hyperparameters (lightgbm, xgboost, …) and oof threshold searching are just to overfit the public test. It's very hard to get 0.4 score by default.</p>",
      "rawMarkdown": "I agree with this. Something like finetuning hyperparameters (lightgbm, xgboost, ...) and oof threshold searching are just to overfit the public test. It's very hard to get 0.4 score by default.",
      "votes": null
    },
    {
      "id": "3070723",
      "postDate": "12/13/2024 01:38:58",
      "content": "<p>Without tuning thresholds, I could get cv ~0.43 by a single model. With tuning I could get cv 0.488 via ensemble(both only using labelled data). But the LB is only 0.45+.</p>",
      "rawMarkdown": "Without tuning thresholds, I could get cv ~0.43 by a single model. With tuning I could get cv 0.488 via ensemble(both only using labelled data). But the LB is only 0.45+.",
      "votes": null
    },
    {
      "id": "3070740",
      "postDate": "12/13/2024 02:19:19",
      "content": "<p>same as me.<br>\nWith tuning, I got CV0.48-0.49 but around 0.45 in public LB </p>",
      "rawMarkdown": "same as me.\nWith tuning, I got CV0.48-0.49 but around 0.45 in public LB",
      "votes": null
    },
    {
      "id": "3071399",
      "postDate": "12/13/2024 18:11:33",
      "content": "<p>Single NN, only labelled main data, drop columns with NaNs more than 30% -&gt; 21 features. Global-maximum tuned thresholds, CV 0.482 | LB 0.448.</p>",
      "rawMarkdown": "Single NN, only labelled main data, drop columns with NaNs more than 30% -> 21 features. Global-maximum tuned thresholds, CV 0.482 | LB 0.448.",
      "votes": null
    },
    {
      "id": "3071511",
      "postDate": "12/13/2024 20:58:09",
      "content": "<p>0.46 - 0.47 and I give 5% chance the public notebooks will work.</p>",
      "rawMarkdown": "0.46 - 0.47 and I give 5% chance the public notebooks will work.",
      "votes": null
    },
    {
      "id": "3071516",
      "postDate": "12/13/2024 21:02:57",
      "content": "<p>0.476 is possibile. There are almost 7000 submissions in total. </p>",
      "rawMarkdown": "0.476 is possibile. There are almost 7000 submissions in total.",
      "votes": null
    },
    {
      "id": "3071677",
      "postDate": "12/14/2024 05:32:34",
      "content": "<p>Hope so but it's very hard to get this score</p>",
      "rawMarkdown": "Hope so but it's very hard to get this score",
      "votes": null
    },
    {
      "id": "3071714",
      "postDate": "12/14/2024 06:57:27",
      "content": "<p>Thanks for sharing your scores, but what do you mean by without tuning thresholds? Aren't you converting soft predictions to classes at some point using some thresholds?</p>",
      "rawMarkdown": "Thanks for sharing your scores, but what do you mean by without tuning thresholds? Aren't you converting soft predictions to classes at some point using some thresholds?",
      "votes": null
    },
    {
      "id": "3071717",
      "postDate": "12/14/2024 07:07:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> , that means I used the original threshold [30,50,80] or [0.5,1.5,2.5] to transform to 0,1,2,3, and then compute the kappa score.</p>",
      "rawMarkdown": "Hi @gunesevitan , that means I used the original threshold [30,50,80] or [0.5,1.5,2.5] to transform to 0,1,2,3, and then compute the kappa score.",
      "votes": null
    },
    {
      "id": "3071722",
      "postDate": "12/14/2024 07:26:49",
      "content": "<p>I second that……the final LB score should lie between 0.43-0.45</p>",
      "rawMarkdown": "I second that......the final LB score should lie between 0.43-0.45",
      "votes": null
    },
    {
      "id": "3071739",
      "postDate": "12/14/2024 07:55:09",
      "content": "<p>I see… I didn't expect you were using the continuous target. That didn't work for me though. Regression objective on thresholded target worked better for me.</p>",
      "rawMarkdown": "I see... I didn't expect you were using the continuous target. That didn't work for me though. Regression objective on thresholded target worked better for me.",
      "votes": null
    },
    {
      "id": "3071797",
      "postDate": "12/14/2024 10:12:51",
      "content": "<blockquote>\n  <p>I give 5% chance the public notebooks will work.</p>\n</blockquote>\n<p>This is generally true, but at the recent AES 2.0 competition (with QWK metric), teams ranked 33rd through 50th all received exactly the same scores while submitting the same public notebook.</p>\n<p>A small modification, such as changing the seed or adding another notebook with a small weight, could give you gold almost for free.</p>",
      "rawMarkdown": "> I give 5% chance the public notebooks will work.\n\nThis is generally true, but at the recent AES 2.0 competition (with QWK metric), teams ranked 33rd through 50th all received exactly the same scores while submitting the same public notebook.\n\nA small modification, such as changing the seed or adding another notebook with a small weight, could give you gold almost for free.",
      "votes": null
    },
    {
      "id": "3071953",
      "postDate": "12/14/2024 15:01:52",
      "content": "<p>Same here…tuning thresholds is giving massive improvement in the CV but the LB is only marginally better. Maybe it’s just overfitting the CV, I wouldn’t be surprised 😅 ~3k samples is not high enough for the threshold to generalize…</p>\n<p>The fold-based thresholds are also not close to each other 😐</p>",
      "rawMarkdown": "Same here…tuning thresholds is giving massive improvement in the CV but the LB is only marginally better. Maybe it’s just overfitting the CV, I wouldn’t be surprised 😅 ~3k samples is not high enough for the threshold to generalize…\n\nThe fold-based thresholds are also not close to each other 😐",
      "votes": null
    },
    {
      "id": "3072695",
      "postDate": "12/15/2024 15:00:16",
      "content": "<p>I would guess between 0.45 and 0.47</p>",
      "rawMarkdown": "I would guess between 0.45 and 0.47",
      "votes": null
    },
    {
      "id": "3072797",
      "postDate": "12/15/2024 16:33:41",
      "content": "<p>I can’t thank you enough for your help.</p>",
      "rawMarkdown": "I can’t thank you enough for your help.",
      "votes": null
    },
    {
      "id": "3072846",
      "postDate": "12/15/2024 17:18:58",
      "content": "<p>Those public notebooks are becoming hardcore blends while approaching to competition ending. They might actually work since they reduce variance. </p>",
      "rawMarkdown": "Those public notebooks are becoming hardcore blends while approaching to competition ending. They might actually work since they reduce variance.",
      "votes": null
    },
    {
      "id": "3073644",
      "postDate": "12/16/2024 17:23:58",
      "content": "<p>IMO the only real winners will be the competition hosts… The community identified enough data issues that hopefully they will learn how to improve their data collection protocols for next time. 👍</p>",
      "rawMarkdown": "IMO the only real winners will be the competition hosts... The community identified enough data issues that hopefully they will learn how to improve their data collection protocols for next time. 👍",
      "votes": null
    },
    {
      "id": "3074116",
      "postDate": "12/17/2024 10:31:22",
      "content": "<p>I see everyone predicts 0.45 0.46, so what will the results be for those who scored 0.49 0.5 in public test data?</p>",
      "rawMarkdown": "I see everyone predicts 0.45 0.46, so what will the results be for those who scored 0.49 0.5 in public test data?",
      "votes": null
    },
    {
      "id": "3074119",
      "postDate": "12/17/2024 10:37:32",
      "content": "<p>Why do you predict so?</p>",
      "rawMarkdown": "Why do you predict so?",
      "votes": null
    },
    {
      "id": "3074381",
      "postDate": "12/17/2024 15:51:34",
      "content": "<p>I would guess between 0.44 and 0.46</p>",
      "rawMarkdown": "I would guess between 0.44 and 0.46",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3069589,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "12/11/2024 16:26:36",
      "content": "<p>I just started this competition and it feels like ICR all over again. I guess it's easier to overfit to a smaller portion of test set so 0.43-0.45 sounds reasonable. I can get 0.49 oof score with a single model but I have no idea how it will translate to private test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069620,
          "author_name": "clearwaterkzk",
          "author_url": "",
          "post_date": "12/11/2024 16:54:39",
          "content": "<p>Thanks for sharing feeling. <br>\nHow do you validate your model ?<br>\nIn my case, I use multiple runs of CV as validation and my CV has steadily improved, but my LB hasn't improved at all.</p>\n<p>I’m wondering why public notebooks with data leakage tend to have high LB scores.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3069621,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "12/11/2024 16:56:56",
              "content": "<p>I track lots of things like fold score mean, std and oof scores, and some proxy metrics. I haven't looked at public notebooks other than eda notebooks.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3069624,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "12/11/2024 17:02:01",
                  "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> with due respect, I may say that most high scoring public notebooks are unlikely to score well on the private LB. Looking at them is a big risk</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3069626,
                      "author_name": "gunesevitan",
                      "author_url": "",
                      "post_date": "12/11/2024 17:04:46",
                      "content": "<p>I agree due to fact that it's easier to overfit here compared to ICR because of thresholds. ICR was evaluated on log loss but QWK is a different story.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3069649,
                          "author_name": "clearwaterkzk",
                          "author_url": "",
                          "post_date": "12/11/2024 17:41:45",
                          "content": "<p>Is there any case that CV/public does not correlate but CV/private does correlate well ?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3070001,
                              "author_name": "gunesevitan",
                              "author_url": "",
                              "post_date": "12/12/2024 05:51:54",
                              "content": "<p>Yeah, most of the shakeup competitions.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3069912,
          "author_name": "yeoyunsianggeremie",
          "author_url": "",
          "post_date": "12/12/2024 03:27:12",
          "content": "<p>Same here, also just started two days ago. I guess I will pick the submission that I have the most confidence on it performing well on private LB, regardless of CV or public LB. It’s easy to overfit both of them due to how small and noisy the data is 😬</p>\n<p>Luck will be a huge factor I guess…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3069614,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/11/2024 16:47:23",
      "content": "<p>I think it will be around 0.455 - 0.46<br>\nIt is very easy to overfit here and a lot of luck is needed to prevent it. This is ICR part 2 in my opinion <a href=\"https://www.kaggle.com/bibanh\" target=\"_blank\">@bibanh</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3074119,
          "author_name": "damvantai",
          "author_url": "",
          "post_date": "12/17/2024 10:37:32",
          "content": "<p>Why do you predict so?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3069658,
      "author_name": "bhatiji",
      "author_url": "",
      "post_date": "12/11/2024 18:04:33",
      "content": "<p>Somewhere around 0.476 is my guess.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3069710,
          "author_name": "bibanh",
          "author_url": "",
          "post_date": "12/11/2024 19:28:48",
          "content": "<p>Oh it's too high kkkk</p>",
          "votes": null,
          "replies": [
            {
              "id": 3071516,
              "author_name": "jankowalski2000",
              "author_url": "",
              "post_date": "12/13/2024 21:02:57",
              "content": "<p>0.476 is possibile. There are almost 7000 submissions in total. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3069671,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/11/2024 18:31:39",
      "content": "<p>Probably not below 0.44.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3069908,
      "author_name": "trcnveli",
      "author_url": "",
      "post_date": "12/12/2024 03:15:48",
      "content": "<p>maybe around .420 will be in gold zone</p>",
      "votes": null,
      "replies": [
        {
          "id": 3070406,
          "author_name": "bibanh",
          "author_url": "",
          "post_date": "12/12/2024 16:38:56",
          "content": "<p>I agree with this. Something like finetuning hyperparameters (lightgbm, xgboost, …) and oof threshold searching are just to overfit the public test. It's very hard to get 0.4 score by default.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3070723,
              "author_name": "takanashihumbert",
              "author_url": "",
              "post_date": "12/13/2024 01:38:58",
              "content": "<p>Without tuning thresholds, I could get cv ~0.43 by a single model. With tuning I could get cv 0.488 via ensemble(both only using labelled data). But the LB is only 0.45+.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3070740,
                  "author_name": "clearwaterkzk",
                  "author_url": "",
                  "post_date": "12/13/2024 02:19:19",
                  "content": "<p>same as me.<br>\nWith tuning, I got CV0.48-0.49 but around 0.45 in public LB </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3071399,
                      "author_name": "yekenot",
                      "author_url": "",
                      "post_date": "12/13/2024 18:11:33",
                      "content": "<p>Single NN, only labelled main data, drop columns with NaNs more than 30% -&gt; 21 features. Global-maximum tuned thresholds, CV 0.482 | LB 0.448.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 3071714,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "12/14/2024 06:57:27",
                  "content": "<p>Thanks for sharing your scores, but what do you mean by without tuning thresholds? Aren't you converting soft predictions to classes at some point using some thresholds?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3071717,
                      "author_name": "takanashihumbert",
                      "author_url": "",
                      "post_date": "12/14/2024 07:07:36",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> , that means I used the original threshold [30,50,80] or [0.5,1.5,2.5] to transform to 0,1,2,3, and then compute the kappa score.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3071739,
                          "author_name": "gunesevitan",
                          "author_url": "",
                          "post_date": "12/14/2024 07:55:09",
                          "content": "<p>I see… I didn't expect you were using the continuous target. That didn't work for me though. Regression objective on thresholded target worked better for me.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    },
                    {
                      "id": 3071953,
                      "author_name": "yeoyunsianggeremie",
                      "author_url": "",
                      "post_date": "12/14/2024 15:01:52",
                      "content": "<p>Same here…tuning thresholds is giving massive improvement in the CV but the LB is only marginally better. Maybe it’s just overfitting the CV, I wouldn’t be surprised 😅 ~3k samples is not high enough for the threshold to generalize…</p>\n<p>The fold-based thresholds are also not close to each other 😐</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3070067,
      "author_name": "silentsapphirewill",
      "author_url": "",
      "post_date": "12/12/2024 07:42:56",
      "content": "<p>maybe around 0.43</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3071511,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "12/13/2024 20:58:09",
      "content": "<p>0.46 - 0.47 and I give 5% chance the public notebooks will work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3071677,
          "author_name": "bibanh",
          "author_url": "",
          "post_date": "12/14/2024 05:32:34",
          "content": "<p>Hope so but it's very hard to get this score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3071797,
          "author_name": "nazarov",
          "author_url": "",
          "post_date": "12/14/2024 10:12:51",
          "content": "<blockquote>\n  <p>I give 5% chance the public notebooks will work.</p>\n</blockquote>\n<p>This is generally true, but at the recent AES 2.0 competition (with QWK metric), teams ranked 33rd through 50th all received exactly the same scores while submitting the same public notebook.</p>\n<p>A small modification, such as changing the seed or adding another notebook with a small weight, could give you gold almost for free.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3072846,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "12/15/2024 17:18:58",
              "content": "<p>Those public notebooks are becoming hardcore blends while approaching to competition ending. They might actually work since they reduce variance. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3071722,
      "author_name": "dristi0705",
      "author_url": "",
      "post_date": "12/14/2024 07:26:49",
      "content": "<p>I second that……the final LB score should lie between 0.43-0.45</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3072695,
      "author_name": "ssqqzs123",
      "author_url": "",
      "post_date": "12/15/2024 15:00:16",
      "content": "<p>I would guess between 0.45 and 0.47</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3072797,
      "author_name": "rudizgen",
      "author_url": "",
      "post_date": "12/15/2024 16:33:41",
      "content": "<p>I can’t thank you enough for your help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073644,
      "author_name": "josephmarturano",
      "author_url": "",
      "post_date": "12/16/2024 17:23:58",
      "content": "<p>IMO the only real winners will be the competition hosts… The community identified enough data issues that hopefully they will learn how to improve their data collection protocols for next time. 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3074116,
      "author_name": "damvantai",
      "author_url": "",
      "post_date": "12/17/2024 10:31:22",
      "content": "<p>I see everyone predicts 0.45 0.46, so what will the results be for those who scored 0.49 0.5 in public test data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3074381,
      "author_name": "damvantai",
      "author_url": "",
      "post_date": "12/17/2024 15:51:34",
      "content": "<p>I would guess between 0.44 and 0.46</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3069582": "I guess the winner (and maybe the gold medal score) is around 0.43-0.45. How about you?",
    "3069589": "I just started this competition and it feels like ICR all over again. I guess it's easier to overfit to a smaller portion of test set so 0.43-0.45 sounds reasonable. I can get 0.49 oof score with a single model but I have no idea how it will translate to private test set.",
    "3069614": "I think it will be around 0.455 - 0.46\nIt is very easy to overfit here and a lot of luck is needed to prevent it. This is ICR part 2 in my opinion @bibanh @gunesevitan",
    "3069620": "Thanks for sharing feeling. \nHow do you validate your model ?\nIn my case, I use multiple runs of CV as validation and my CV has steadily improved, but my LB hasn't improved at all.\n\nI’m wondering why public notebooks with data leakage tend to have high LB scores.",
    "3069621": "I track lots of things like fold score mean, std and oof scores, and some proxy metrics. I haven't looked at public notebooks other than eda notebooks.",
    "3069624": "gunesevitan with due respect, I may say that most high scoring public notebooks are unlikely to score well on the private LB. Looking at them is a big risk",
    "3069626": "I agree due to fact that it's easier to overfit here compared to ICR because of thresholds. ICR was evaluated on log loss but QWK is a different story.",
    "3069649": "Is there any case that CV/public does not correlate but CV/private does correlate well ?",
    "3069658": "Somewhere around 0.476 is my guess.",
    "3069671": "Probably not below 0.44.",
    "3069710": "Oh it's too high kkkk",
    "3069908": "maybe around .420 will be in gold zone",
    "3069912": "Same here, also just started two days ago. I guess I will pick the submission that I have the most confidence on it performing well on private LB, regardless of CV or public LB. It’s easy to overfit both of them due to how small and noisy the data is 😬\n\nLuck will be a huge factor I guess…",
    "3070001": "Yeah, most of the shakeup competitions.",
    "3070067": "maybe around 0.43",
    "3070406": "I agree with this. Something like finetuning hyperparameters (lightgbm, xgboost, ...) and oof threshold searching are just to overfit the public test. It's very hard to get 0.4 score by default.",
    "3070723": "Without tuning thresholds, I could get cv ~0.43 by a single model. With tuning I could get cv 0.488 via ensemble(both only using labelled data). But the LB is only 0.45+.",
    "3070740": "same as me.\nWith tuning, I got CV0.48-0.49 but around 0.45 in public LB",
    "3071399": "Single NN, only labelled main data, drop columns with NaNs more than 30% -> 21 features. Global-maximum tuned thresholds, CV 0.482 | LB 0.448.",
    "3071511": "0.46 - 0.47 and I give 5% chance the public notebooks will work.",
    "3071516": "0.476 is possibile. There are almost 7000 submissions in total.",
    "3071677": "Hope so but it's very hard to get this score",
    "3071714": "Thanks for sharing your scores, but what do you mean by without tuning thresholds? Aren't you converting soft predictions to classes at some point using some thresholds?",
    "3071717": "Hi @gunesevitan , that means I used the original threshold [30,50,80] or [0.5,1.5,2.5] to transform to 0,1,2,3, and then compute the kappa score.",
    "3071722": "I second that......the final LB score should lie between 0.43-0.45",
    "3071739": "I see... I didn't expect you were using the continuous target. That didn't work for me though. Regression objective on thresholded target worked better for me.",
    "3071797": "> I give 5% chance the public notebooks will work.\n\nThis is generally true, but at the recent AES 2.0 competition (with QWK metric), teams ranked 33rd through 50th all received exactly the same scores while submitting the same public notebook.\n\nA small modification, such as changing the seed or adding another notebook with a small weight, could give you gold almost for free.",
    "3071953": "Same here…tuning thresholds is giving massive improvement in the CV but the LB is only marginally better. Maybe it’s just overfitting the CV, I wouldn’t be surprised 😅 ~3k samples is not high enough for the threshold to generalize…\n\nThe fold-based thresholds are also not close to each other 😐",
    "3072695": "I would guess between 0.45 and 0.47",
    "3072797": "I can’t thank you enough for your help.",
    "3072846": "Those public notebooks are becoming hardcore blends while approaching to competition ending. They might actually work since they reduce variance.",
    "3073644": "IMO the only real winners will be the competition hosts... The community identified enough data issues that hopefully they will learn how to improve their data collection protocols for next time. 👍",
    "3074116": "I see everyone predicts 0.45 0.46, so what will the results be for those who scored 0.49 0.5 in public test data?",
    "3074119": "Why do you predict so?",
    "3074381": "I would guess between 0.44 and 0.46"
  },
  "source": "meta"
}