{
  "id": 535523,
  "title": "CV-LB thread",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535523",
  "author_name": "",
  "post_date": "2024-09-22T14:55:28.391743300Z",
  "votes": 28,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I think it would be good to discuss CV-LB relations as we embark on our respective model pursuits. CV becomes even more important with evaluation metrics like the Kappa score that rely on hard labels rather than soft probabilities. A CV-LB thread with baseline approaches will help one and all evaluate their models effectively and add value to their pursuits.</p>\n<p>I shall commence with my findings insofar, most of which are public. I have used a cv scheme as below-<br>\n<code>cv = StratifiedKFold(n_splits = 5, random_state = 42, shuffle = True)</code></p>\n<table>\n<thead>\n<tr>\n<th>Model type</th>\n<th>Classification/ Regression</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>Regression</td>\n<td>0.464780</td>\n<td>0.453</td>\n</tr>\n<tr>\n<td>LGBM-XGB</td>\n<td>Regression</td>\n<td>0.463967</td>\n<td>0.459</td>\n</tr>\n<tr>\n<td>Catboost</td>\n<td>Regression</td>\n<td>0.457608</td>\n<td>0.442</td>\n</tr>\n<tr>\n<td>Catboost- LGBM-XGB</td>\n<td>Regression</td>\n<td>0.457419</td>\n<td>0.432</td>\n</tr>\n<tr>\n<td>Catboost- LGBM-XGB</td>\n<td>Regression</td>\n<td>0.457849</td>\n<td>0.431</td>\n</tr>\n</tbody>\n</table>\n<p>More details are in the below kernels- </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1</a></li>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2</a></li>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1</a></li>\n</ul>\n<p>Others may add value to the post with their findings as well!<br>\nBest regards!</p>",
  "messages": [
    {
      "id": "2995742",
      "postDate": "09/22/2024 14:55:28",
      "content": "<p>Hello all,</p>\n<p>I think it would be good to discuss CV-LB relations as we embark on our respective model pursuits. CV becomes even more important with evaluation metrics like the Kappa score that rely on hard labels rather than soft probabilities. A CV-LB thread with baseline approaches will help one and all evaluate their models effectively and add value to their pursuits.</p>\n<p>I shall commence with my findings insofar, most of which are public. I have used a cv scheme as below-<br>\n<code>cv = StratifiedKFold(n_splits = 5, random_state = 42, shuffle = True)</code></p>\n<table>\n<thead>\n<tr>\n<th>Model type</th>\n<th>Classification/ Regression</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM</td>\n<td>Regression</td>\n<td>0.464780</td>\n<td>0.453</td>\n</tr>\n<tr>\n<td>LGBM-XGB</td>\n<td>Regression</td>\n<td>0.463967</td>\n<td>0.459</td>\n</tr>\n<tr>\n<td>Catboost</td>\n<td>Regression</td>\n<td>0.457608</td>\n<td>0.442</td>\n</tr>\n<tr>\n<td>Catboost- LGBM-XGB</td>\n<td>Regression</td>\n<td>0.457419</td>\n<td>0.432</td>\n</tr>\n<tr>\n<td>Catboost- LGBM-XGB</td>\n<td>Regression</td>\n<td>0.457849</td>\n<td>0.431</td>\n</tr>\n</tbody>\n</table>\n<p>More details are in the below kernels- </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1</a></li>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2</a></li>\n<li><a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1</a></li>\n</ul>\n<p>Others may add value to the post with their findings as well!<br>\nBest regards!</p>",
      "rawMarkdown": "Hello all,\n\nI think it would be good to discuss CV-LB relations as we embark on our respective model pursuits. CV becomes even more important with evaluation metrics like the Kappa score that rely on hard labels rather than soft probabilities. A CV-LB thread with baseline approaches will help one and all evaluate their models effectively and add value to their pursuits.\n\nI shall commence with my findings insofar, most of which are public. I have used a cv scheme as below-\n`cv = StratifiedKFold(n_splits = 5, random_state = 42, shuffle = True)`\n\n| Model type | Classification/ Regression  | CV | LB |\n| ---               | --- | ------------ | --------------- |\n| LGBM          |  Regression  | 0.464780  |  0.453 |\n| LGBM-XGB |  Regression  |  0.463967 |  0.459 |\n| Catboost     |  Regression  |  0.457608 |  0.442 |\n|  Catboost- LGBM-XGB  |  Regression  |  0.457419  |  0.432 |\n| Catboost- LGBM-XGB   |  Regression  |  0.457849  |  0.431 |\n\n\nMore details are in the below kernels- \n- https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1\n- https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2\n- https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1\n\nOthers may add value to the post with their findings as well!\nBest regards!",
      "votes": null
    },
    {
      "id": "2996015",
      "postDate": "09/22/2024 23:34:25",
      "content": "<h3>Performance Of Single Model LGBM Both Have Same Features + Params Optuna Optimized</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>First Model LGBM</td>\n<td>0.7301</td>\n<td>0.4139</td>\n<td>0.457</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>Second Model LGBM</td>\n<td>0.7151</td>\n<td>0.3942</td>\n<td>0.452</td>\n<td>0.455</td>\n</tr>\n<tr>\n<td>Thrid Model LGBM</td>\n<td>0.7875</td>\n<td>0.4140</td>\n<td>0.440</td>\n<td>0.435</td>\n</tr>\n</tbody>\n</table>\n<h4>In the end again Shakeup. If we see carefully , my First Model has a Good Train , Test and Optimized Score and a Bad LB Score. While The Second Model has Bad Train, Test and Optimized Score Than First one But it Perform Well on LB.</h4>\n<ul>\n<li>CV <code>StratifiedKFold(n_splits=5, shuffle=True, random_state=42)</code></li>\n</ul>\n<p>Edited : </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>New LGBM</td>\n<td>0.5298</td>\n<td>0.3593</td>\n<td>0.469</td>\n<td>0.456</td>\n</tr>\n<tr>\n<td>Second New LGBM</td>\n<td>0.5465</td>\n<td>0.3713</td>\n<td>0.478</td>\n<td>0.453</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>I'm feeling quite confused about what to focus on. Let's see, we still have three months. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </li>\n</ul>",
      "rawMarkdown": "### Performance Of Single Model LGBM Both Have Same Features + Params Optuna Optimized\n\n| Model             | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-----------|----------------|----------------|-------------|\n| First Model LGBM | 0.7301    | 0.4139         | 0.457          | 0.452       |\n| Second Model LGBM | 0.7151    | 0.3942         | 0.452          | 0.455       |\n| Thrid Model LGBM | 0.7875    | 0.4140         | 0.440          | 0.435       |\n\n#### In the end again Shakeup. If we see carefully , my First Model has a Good Train , Test and Optimized Score and a Bad LB Score. While The Second Model has Bad Train, Test and Optimized Score Than First one But it Perform Well on LB. \n\n- CV `StratifiedKFold(n_splits=5, shuffle=True, random_state=42)`\n\nEdited : \n\n| Model                   | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-------------|----------------|----------------|-------------|\n| New LGBM           | 0.5298         | 0.3593             | 0.469                | 0.456\n| Second New LGBM           | 0.5465         | 0.3713             | 0.478                | 0.453\n\n- I'm feeling quite confused about what to focus on. Let's see, we still have three months. @ravi20076",
      "votes": null
    },
    {
      "id": "2996024",
      "postDate": "09/23/2024 00:22:09",
      "content": "<p>I think Second one is Overfitted </p>",
      "rawMarkdown": "I think Second one is Overfitted",
      "votes": null
    },
    {
      "id": "2996364",
      "postDate": "09/23/2024 12:15:13",
      "content": "<h3>Latest Submission | Overfitting LB or What? ! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a></h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>First Model LGBM</td>\n<td>0.7117</td>\n<td>0.4082</td>\n<td>0.457</td>\n<td>0.465</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "### Latest Submission | Overfitting LB or What? ! @ravi20076 \n\n| Model             | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-----------|----------------|----------------|-------------|\n| First Model LGBM | 0.7117    | 0.4082         | 0.457          | 0.465     |",
      "votes": null
    },
    {
      "id": "2996462",
      "postDate": "09/23/2024 14:25:14",
      "content": "<p>This itself elicits the structure of the leaderboard <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
      "rawMarkdown": "This itself elicits the structure of the leaderboard @abdmental01",
      "votes": null
    },
    {
      "id": "2996944",
      "postDate": "09/24/2024 05:25:05",
      "content": "<p>Same with me, some of my better CV model results have LB scores in the range of 0.452- 0.455 <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <br>\nIs this an <a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard\" target=\"_blank\">ICR</a>?</p>",
      "rawMarkdown": "Same with me, some of my better CV model results have LB scores in the range of 0.452- 0.455 @abdmental01 \nIs this an [ICR](https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard)?",
      "votes": null
    },
    {
      "id": "2997210",
      "postDate": "09/24/2024 10:36:04",
      "content": "<p>I didn't participate in ICR, but after seeing the leaderboard, I think the same situation might occur here as well. My Some models that performed well in cross-validation are not doing as well on the leaderboard. It seems like the model is not capturing the data patterns efficiently. Alternatively, it could be that some of the models that are performing well on the leaderboard are better at generalizing and capturing the patterns in the data. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> It's hard to chose which one to trust CV or LB</p>\n<p>Example of my recent Two Single Models : </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV Score</th>\n<th>Optimized Score</th>\n<th>LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>0.4113</td>\n<td>0.457</td>\n<td>0.465</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>0.4094</td>\n<td>0.452</td>\n<td>0.471</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I didn't participate in ICR, but after seeing the leaderboard, I think the same situation might occur here as well. My Some models that performed well in cross-validation are not doing as well on the leaderboard. It seems like the model is not capturing the data patterns efficiently. Alternatively, it could be that some of the models that are performing well on the leaderboard are better at generalizing and capturing the patterns in the data. @ravi20076 It's hard to chose which one to trust CV or LB\n\nExample of my recent Two Single Models : \n\n| Model    | CV Score | Optimized Score | LB Score |\n|----------|----------|-----------------|----------|\n| Model 1  | 0.4113   | 0.457            | 0.465     |\n| Model 2  | 0.4094   | 0.452           | 0.471    |",
      "votes": null
    },
    {
      "id": "2998053",
      "postDate": "09/25/2024 07:27:45",
      "content": "<p>My results </p>\n<ol>\n<li>Model 1 : 10 fold StratifiedKfold same params as public notebooks Cv = 0.4116 , LB = 0.441</li>\n<li>Model 2 : 10 fold StratifiedKfold(Changed Seed) , slight modification from public notebooks , CV = 0.4159 , LB = 0.453</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> seeing the results and  reading the comments here it feels like we would see similar leaderboard as we saw at the end of AES 2.0 scoring competition needless to say there too the scoring metric was QWK I guess😐</p>",
      "rawMarkdown": "My results \n1. Model 1 : 10 fold StratifiedKfold same params as public notebooks Cv = 0.4116 , LB = 0.441\n2. Model 2 : 10 fold StratifiedKfold(Changed Seed) , slight modification from public notebooks , CV = 0.4159 , LB = 0.453\n\n@ravi20076 seeing the results and  reading the comments here it feels like we would see similar leaderboard as we saw at the end of AES 2.0 scoring competition needless to say there too the scoring metric was QWK I guess😐",
      "votes": null
    },
    {
      "id": "2998105",
      "postDate": "09/25/2024 08:39:18",
      "content": "<p>It will be worse in my opinion as this dataset is even smaller than the other one and is more noisy <a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> </p>",
      "rawMarkdown": "It will be worse in my opinion as this dataset is even smaller than the other one and is more noisy @sayedathar11",
      "votes": null
    },
    {
      "id": "2998164",
      "postDate": "09/25/2024 10:07:35",
      "content": "<p>share my exp results.<br>\nI conducted an experiment using Optuna by varying the evaluation metrics to be optimized.<br>\nAs of now, there is no correlation between CV and LB, and the situation is not good.</p>\n<ul>\n<li>Model : LGB Single</li>\n<li>Split : SKF(5Fold)</li>\n<li>FE : 136 Col(AfterFeatureSelection)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>LB↑</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.3894</td>\n<td>0.445</td>\n<td><strong>0.457</strong></td>\n</tr>\n<tr>\n<td>0.398</td>\n<td>0.467</td>\n<td>0.453</td>\n</tr>\n<tr>\n<td>0.4077</td>\n<td>0.438</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>0.3578</td>\n<td><strong>0.484</strong></td>\n<td>0.447</td>\n</tr>\n<tr>\n<td><strong>0.412</strong></td>\n<td>0.467</td>\n<td>0.446</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "share my exp results.\nI conducted an experiment using Optuna by varying the evaluation metrics to be optimized.\nAs of now, there is no correlation between CV and LB, and the situation is not good.\n\n* Model : LGB Single\n* Split : SKF(5Fold)\n* FE : 136 Col(AfterFeatureSelection)\n\n| Validation QWK | Optimized QWK | LB↑ |\n| --- | --- | --- |\n0.3894|0.445|**0.457**|\n0.398|0.467|0.453|\n0.4077|0.438|0.452|\n0.3578|**0.484**|0.447|\n**0.412**|0.467|0.446|",
      "votes": null
    },
    {
      "id": "2998383",
      "postDate": "09/25/2024 15:25:03",
      "content": "<p><a href=\"https://www.kaggle.com/hideyukizushi\" target=\"_blank\">@hideyukizushi</a> <br>\nI also have very similar observations- no CV-LB relation till date<br>\nMy best CV model is around yours - <strong>0.4813567</strong> with LB of <strong>0.447</strong> as well. Most of my good CV scores are providing LB of 0.452 - 0.455</p>",
      "rawMarkdown": "hideyukizushi \nI also have very similar observations- no CV-LB relation till date\nMy best CV model is around yours - **0.4813567** with LB of **0.447** as well. Most of my good CV scores are providing LB of 0.452 - 0.455",
      "votes": null
    },
    {
      "id": "2998758",
      "postDate": "09/26/2024 01:27:33",
      "content": "<p>I just try simple LGBM several times, and CV and LB are almost unrelated, much depending on poor tabular data. </p>",
      "rawMarkdown": "I just try simple LGBM several times, and CV and LB are almost unrelated, much depending on poor tabular data.",
      "votes": null
    },
    {
      "id": "3008329",
      "postDate": "10/06/2024 13:32:05",
      "content": "<p>I didn't have time to work on my models so I used my submissions to see how the LB-Score changes with different random-seeds for two models. </p>\n<p><strong>Model 1:</strong> <a href=\"https://www.kaggle.com/code/lennarthaupts/cmi-detecting-problematic-digital-behavior\" target=\"_blank\">notebook</a></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Optimized</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.443</td>\n<td>0.455</td>\n<td>0.467</td>\n</tr>\n<tr>\n<td>0.449</td>\n<td>0.470</td>\n<td>0.462</td>\n</tr>\n<tr>\n<td>0.441</td>\n<td>0.456</td>\n<td>0.468</td>\n</tr>\n<tr>\n<td>0.440</td>\n<td>0.468</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>0.443</td>\n<td>0.457</td>\n<td>0.455</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Model 2:</strong></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Optimized</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.410</td>\n<td>0.462</td>\n<td>0.446</td>\n</tr>\n<tr>\n<td>0.384</td>\n<td>0.449</td>\n<td>0.442</td>\n</tr>\n<tr>\n<td>0.387</td>\n<td>0.452</td>\n<td><strong>0.472</strong></td>\n</tr>\n<tr>\n<td>0.387</td>\n<td>0.459</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>0.405</td>\n<td>0.473</td>\n<td>0.457</td>\n</tr>\n</tbody>\n</table>\n<p>SKF with 10 folds in both cases. My take away is that overfitting the LB can quickly happen with some random luck.</p>",
      "rawMarkdown": "I didn't have time to work on my models so I used my submissions to see how the LB-Score changes with different random-seeds for two models. \n\n**Model 1:** [notebook](https://www.kaggle.com/code/lennarthaupts/cmi-detecting-problematic-digital-behavior)\n| CV | Optimized |LB|\n| --- | --- |---|\n| 0.443 |0.455|0.467|\n| 0.449|0.470 |0.462|\n|0.441|0.456|0.468|\n|0.440|0.468|0.466|\n| 0.443|0.457|0.455|\n\n**Model 2:**\n| CV | Optimized |LB|\n| --- | --- |---|\n|0.410|0.462|0.446|\n|0.384|0.449|0.442|\n|0.387| 0.452|**0.472**|\n|0.387|0.459|0.466|\n|0.405|0.473|0.457|\n\nSKF with 10 folds in both cases. My take away is that overfitting the LB can quickly happen with some random luck.",
      "votes": null
    },
    {
      "id": "3008559",
      "postDate": "10/06/2024 18:19:51",
      "content": "<p>I guess these 2 scenarios, and some of the result from Lennart shared here, are indicating that test set is less difficult than what we're expecting from train set.</p>",
      "rawMarkdown": "I guess these 2 scenarios, and some of the result from Lennart shared here, are indicating that test set is less difficult than what we're expecting from train set.",
      "votes": null
    },
    {
      "id": "3010141",
      "postDate": "10/08/2024 17:35:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>, hi everyone,</p>\n<p>I've got these results : </p>\n<p>Train QWK optimized : .473<br>\nOOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.<br>\nPublic LB QWK : .435 (of course I'm desappointed, but something tells me to not trust public LB)</p>\n<p>I'd rather to fit optimization thresholds with train samples instead of OOF samples to control overfitting.</p>\n<p>I'm wondering what are best CV QWK today…</p>",
      "rawMarkdown": "Hi @ravi20076, hi everyone,\n\nI've got these results : \n\nTrain QWK optimized : .473\nOOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.\nPublic LB QWK : .435 (of course I'm desappointed, but something tells me to not trust public LB)\n\nI'd rather to fit optimization thresholds with train samples instead of OOF samples to control overfitting.\n\nI'm wondering what are best CV QWK today...",
      "votes": null
    },
    {
      "id": "3012243",
      "postDate": "10/08/2024 19:43:05",
      "content": "<p>The leaderboard is not trustworthy here <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, a couple of incorrectly classified points can bump one down and a lucky seed/ model parameter can bounce one up a lot!</p>\n<p>I changed just the reg_lambda in my lightgbm and my leaderboard increased from 0.452 -&gt; 0.463 earlier in the competition</p>",
      "rawMarkdown": "The leaderboard is not trustworthy here @adaubas, a couple of incorrectly classified points can bump one down and a lucky seed/ model parameter can bounce one up a lot!\n\nI changed just the reg_lambda in my lightgbm and my leaderboard increased from 0.452 -> 0.463 earlier in the competition",
      "votes": null
    },
    {
      "id": "3012265",
      "postDate": "10/08/2024 20:10:43",
      "content": "<p>And did get a better CV score than your previous 0.481 ?</p>\n<p>I'm wondering if we can do really better than 0.48.</p>",
      "rawMarkdown": "And did get a better CV score than your previous 0.481 ?\n\n I'm wondering if we can do really better than 0.48.",
      "votes": null
    },
    {
      "id": "3012401",
      "postDate": "10/09/2024 02:37:14",
      "content": "<p>CV score was nearly the same for both cases (changed by 0.0001) <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>",
      "rawMarkdown": "CV score was nearly the same for both cases (changed by 0.0001) @adaubas",
      "votes": null
    },
    {
      "id": "3012578",
      "postDate": "10/09/2024 07:32:56",
      "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>\n<blockquote>\n  <p>OOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.</p>\n</blockquote>\n<p>Have you tried to fit different random seed into the exact same model and get a few more CV? some of my models are more stable (+/-0.01) but some has large variance (+/-0.05).</p>\n<p>The training data set is too small that Optuna search cannot be fully trusted. I had a few cases the 'best param' has good cv only at that particular random seed.</p>",
      "rawMarkdown": "adaubas \n>OOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.\n\nHave you tried to fit different random seed into the exact same model and get a few more CV? some of my models are more stable (+/-0.01) but some has large variance (+/-0.05).\n\nThe training data set is too small that Optuna search cannot be fully trusted. I had a few cases the 'best param' has good cv only at that particular random seed.",
      "votes": null
    },
    {
      "id": "3014259",
      "postDate": "10/11/2024 05:12:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a>,</p>\n<p>Yes I tried with several random seeds and my OOF QWK optimized after 5 stratified 5 folds is stable : between .004 and .008. </p>",
      "rawMarkdown": "Hi @tomyuen,\n\nYes I tried with several random seeds and my OOF QWK optimized after 5 stratified 5 folds is stable : between .004 and .008.",
      "votes": null
    },
    {
      "id": "3014408",
      "postDate": "10/11/2024 08:11:02",
      "content": "<p>I've had very similar CV/LB scores , I think there will be so much luck involved in this competition . </p>",
      "rawMarkdown": "I've had very similar CV/LB scores , I think there will be so much luck involved in this competition .",
      "votes": null
    },
    {
      "id": "3014552",
      "postDate": "10/11/2024 10:46:00",
      "content": "<p>This in my opinion, is 99% luck and 1% everything else <a href=\"https://www.kaggle.com/mohammedahmedxx12\" target=\"_blank\">@mohammedahmedxx12</a> </p>",
      "rawMarkdown": "This in my opinion, is 99% luck and 1% everything else @mohammedahmedxx12",
      "votes": null
    },
    {
      "id": "3026442",
      "postDate": "10/23/2024 19:46:29",
      "content": "<p>Is this irony sir? </p>",
      "rawMarkdown": "Is this irony sir?",
      "votes": null
    },
    {
      "id": "3026470",
      "postDate": "10/23/2024 20:40:07",
      "content": "<p>Reality Sir <a href=\"https://www.kaggle.com/davids1992\" target=\"_blank\">@davids1992</a> </p>",
      "rawMarkdown": "Reality Sir @davids1992",
      "votes": null
    },
    {
      "id": "3028000",
      "postDate": "10/25/2024 14:14:32",
      "content": "<p>Did you use the actigraphy data in your models?</p>",
      "rawMarkdown": "Did you use the actigraphy data in your models?",
      "votes": null
    },
    {
      "id": "3028644",
      "postDate": "10/26/2024 11:27:31",
      "content": "<p>Yes of course <a href=\"https://www.kaggle.com/deepaksaldanha\" target=\"_blank\">@deepaksaldanha</a> </p>",
      "rawMarkdown": "Yes of course @deepaksaldanha",
      "votes": null
    },
    {
      "id": "3029053",
      "postDate": "10/26/2024 19:38:39",
      "content": "<p>Thanks for your reply, curious to know, since actigraphy data for a major chunk of students in both train and test sets is missing, I'm assuming you've used imputation techniques to fill those gaps, is this why you are seeing so much difference in CV vs LB? could these imputations be creating heavily biased data ?</p>",
      "rawMarkdown": "Thanks for your reply, curious to know, since actigraphy data for a major chunk of students in both train and test sets is missing, I'm assuming you've used imputation techniques to fill those gaps, is this why you are seeing so much difference in CV vs LB? could these imputations be creating heavily biased data ?",
      "votes": null
    },
    {
      "id": "3032412",
      "postDate": "10/30/2024 21:40:06",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> What CV is that? A stratified k-fold, or anything else? 0.4813567 is pretty high…</p>",
      "rawMarkdown": "ravi20076 What CV is that? A stratified k-fold, or anything else? 0.4813567 is pretty high...",
      "votes": null
    },
    {
      "id": "3032444",
      "postDate": "10/30/2024 22:10:51",
      "content": "<p><a href=\"https://www.kaggle.com/davids1992\" target=\"_blank\">@davids1992</a> Stratified 5 fold CV scheme</p>",
      "rawMarkdown": "davids1992 Stratified 5 fold CV scheme",
      "votes": null
    },
    {
      "id": "3035207",
      "postDate": "11/03/2024 06:26:23",
      "content": "<p>Single LGBM regression, the difference between the models is FE and feature selection. Using 10 features atm..</p>\n<p>tried actigraphy data but found it caused increase in train scores but decrease in oof so decided not to use it (yet anyways..).</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>TRAIN</th>\n<th>OOF</th>\n<th>LB</th>\n<th>OOF/LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM 1</td>\n<td>0.493</td>\n<td>0.478</td>\n<td>0.455</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>LGBM 2</td>\n<td>0.490</td>\n<td>0.480</td>\n<td>0.449</td>\n<td>0.464</td>\n</tr>\n<tr>\n<td>LGBM 3</td>\n<td>0.494</td>\n<td>0.481</td>\n<td>0.462</td>\n<td>0.472</td>\n</tr>\n<tr>\n<td>LGBM 4</td>\n<td>0.515</td>\n<td>0.488</td>\n<td>0.463</td>\n<td>0.476</td>\n</tr>\n<tr>\n<td>LGBM 5</td>\n<td>0.504</td>\n<td>0.494</td>\n<td>0.457</td>\n<td>0.476</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Single LGBM regression, the difference between the models is FE and feature selection. Using 10 features atm..\n\ntried actigraphy data but found it caused increase in train scores but decrease in oof so decided not to use it (yet anyways..).\n\n| Model    | TRAIN | OOF | LB | OOF/LB |\n|--------|-------------|-----------|----------|--------------|\n| LGBM 1  | 0.493       | 0.478     | 0.455    | 0.466        |\n| LGBM 2  | 0.490       | 0.480     | 0.449    | 0.464        |\n| LGBM 3  | 0.494       | 0.481     | 0.462    | 0.472        |\n| LGBM 4  | 0.515       | 0.488     | 0.463    | 0.476        |\n| LGBM 5  | 0.504       | 0.494     | 0.457    | 0.476        |",
      "votes": null
    },
    {
      "id": "3037917",
      "postDate": "11/06/2024 11:15:38",
      "content": "<p>looking at the comments, I saw a column Optimized, what does it mean? Could short describe about it?</p>",
      "rawMarkdown": "looking at the comments, I saw a column Optimized, what does it mean? Could short describe about it?",
      "votes": null
    },
    {
      "id": "3038020",
      "postDate": "11/06/2024 13:42:29",
      "content": "<p>most people in this game use a regression+ QWK optimizer to classify cases into 0,1,2,3, instead of a simple classification 0 1 2 3, or turning result into integer based on 0.5 1.5 2.5 threshold.</p>",
      "rawMarkdown": "most people in this game use a regression+ QWK optimizer to classify cases into 0,1,2,3, instead of a simple classification 0 1 2 3, or turning result into integer based on 0.5 1.5 2.5 threshold.",
      "votes": null
    },
    {
      "id": "3038032",
      "postDate": "11/06/2024 13:58:45",
      "content": "<p>and as i can see it gives a good result:)</p>",
      "rawMarkdown": "and as i can see it gives a good result:)",
      "votes": null
    },
    {
      "id": "3077114",
      "postDate": "12/20/2024 15:18:21",
      "content": "<p>Great notebook!</p>",
      "rawMarkdown": "Great notebook!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2996015,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "09/22/2024 23:34:25",
      "content": "<h3>Performance Of Single Model LGBM Both Have Same Features + Params Optuna Optimized</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>First Model LGBM</td>\n<td>0.7301</td>\n<td>0.4139</td>\n<td>0.457</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>Second Model LGBM</td>\n<td>0.7151</td>\n<td>0.3942</td>\n<td>0.452</td>\n<td>0.455</td>\n</tr>\n<tr>\n<td>Thrid Model LGBM</td>\n<td>0.7875</td>\n<td>0.4140</td>\n<td>0.440</td>\n<td>0.435</td>\n</tr>\n</tbody>\n</table>\n<h4>In the end again Shakeup. If we see carefully , my First Model has a Good Train , Test and Optimized Score and a Bad LB Score. While The Second Model has Bad Train, Test and Optimized Score Than First one But it Perform Well on LB.</h4>\n<ul>\n<li>CV <code>StratifiedKFold(n_splits=5, shuffle=True, random_state=42)</code></li>\n</ul>\n<p>Edited : </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>New LGBM</td>\n<td>0.5298</td>\n<td>0.3593</td>\n<td>0.469</td>\n<td>0.456</td>\n</tr>\n<tr>\n<td>Second New LGBM</td>\n<td>0.5465</td>\n<td>0.3713</td>\n<td>0.478</td>\n<td>0.453</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>I'm feeling quite confused about what to focus on. Let's see, we still have three months. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2996024,
          "author_name": "abdmental01",
          "author_url": "",
          "post_date": "09/23/2024 00:22:09",
          "content": "<p>I think Second one is Overfitted </p>",
          "votes": null,
          "replies": [
            {
              "id": 2996944,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "09/24/2024 05:25:05",
              "content": "<p>Same with me, some of my better CV model results have LB scores in the range of 0.452- 0.455 <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <br>\nIs this an <a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard\" target=\"_blank\">ICR</a>?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2997210,
                  "author_name": "abdmental01",
                  "author_url": "",
                  "post_date": "09/24/2024 10:36:04",
                  "content": "<p>I didn't participate in ICR, but after seeing the leaderboard, I think the same situation might occur here as well. My Some models that performed well in cross-validation are not doing as well on the leaderboard. It seems like the model is not capturing the data patterns efficiently. Alternatively, it could be that some of the models that are performing well on the leaderboard are better at generalizing and capturing the patterns in the data. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> It's hard to chose which one to trust CV or LB</p>\n<p>Example of my recent Two Single Models : </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV Score</th>\n<th>Optimized Score</th>\n<th>LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>0.4113</td>\n<td>0.457</td>\n<td>0.465</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>0.4094</td>\n<td>0.452</td>\n<td>0.471</td>\n</tr>\n</tbody>\n</table>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3008559,
                      "author_name": "tomyuen",
                      "author_url": "",
                      "post_date": "10/06/2024 18:19:51",
                      "content": "<p>I guess these 2 scenarios, and some of the result from Lennart shared here, are indicating that test set is less difficult than what we're expecting from train set.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2996364,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "09/23/2024 12:15:13",
      "content": "<h3>Latest Submission | Overfitting LB or What? ! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a></h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train QWK</th>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>First Model LGBM</td>\n<td>0.7117</td>\n<td>0.4082</td>\n<td>0.457</td>\n<td>0.465</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2996462,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/23/2024 14:25:14",
          "content": "<p>This itself elicits the structure of the leaderboard <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2998053,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "09/25/2024 07:27:45",
      "content": "<p>My results </p>\n<ol>\n<li>Model 1 : 10 fold StratifiedKfold same params as public notebooks Cv = 0.4116 , LB = 0.441</li>\n<li>Model 2 : 10 fold StratifiedKfold(Changed Seed) , slight modification from public notebooks , CV = 0.4159 , LB = 0.453</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> seeing the results and  reading the comments here it feels like we would see similar leaderboard as we saw at the end of AES 2.0 scoring competition needless to say there too the scoring metric was QWK I guess😐</p>",
      "votes": null,
      "replies": [
        {
          "id": 2998105,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/25/2024 08:39:18",
          "content": "<p>It will be worse in my opinion as this dataset is even smaller than the other one and is more noisy <a href=\"https://www.kaggle.com/sayedathar11\" target=\"_blank\">@sayedathar11</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2998164,
      "author_name": "hideyukizushi",
      "author_url": "",
      "post_date": "09/25/2024 10:07:35",
      "content": "<p>share my exp results.<br>\nI conducted an experiment using Optuna by varying the evaluation metrics to be optimized.<br>\nAs of now, there is no correlation between CV and LB, and the situation is not good.</p>\n<ul>\n<li>Model : LGB Single</li>\n<li>Split : SKF(5Fold)</li>\n<li>FE : 136 Col(AfterFeatureSelection)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Validation QWK</th>\n<th>Optimized QWK</th>\n<th>LB↑</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.3894</td>\n<td>0.445</td>\n<td><strong>0.457</strong></td>\n</tr>\n<tr>\n<td>0.398</td>\n<td>0.467</td>\n<td>0.453</td>\n</tr>\n<tr>\n<td>0.4077</td>\n<td>0.438</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>0.3578</td>\n<td><strong>0.484</strong></td>\n<td>0.447</td>\n</tr>\n<tr>\n<td><strong>0.412</strong></td>\n<td>0.467</td>\n<td>0.446</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2998383,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/25/2024 15:25:03",
          "content": "<p><a href=\"https://www.kaggle.com/hideyukizushi\" target=\"_blank\">@hideyukizushi</a> <br>\nI also have very similar observations- no CV-LB relation till date<br>\nMy best CV model is around yours - <strong>0.4813567</strong> with LB of <strong>0.447</strong> as well. Most of my good CV scores are providing LB of 0.452 - 0.455</p>",
          "votes": null,
          "replies": [
            {
              "id": 3032412,
              "author_name": "davids1992",
              "author_url": "",
              "post_date": "10/30/2024 21:40:06",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> What CV is that? A stratified k-fold, or anything else? 0.4813567 is pretty high…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3032444,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "10/30/2024 22:10:51",
                  "content": "<p><a href=\"https://www.kaggle.com/davids1992\" target=\"_blank\">@davids1992</a> Stratified 5 fold CV scheme</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2998758,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "09/26/2024 01:27:33",
      "content": "<p>I just try simple LGBM several times, and CV and LB are almost unrelated, much depending on poor tabular data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3008329,
      "author_name": "lennarthaupts",
      "author_url": "",
      "post_date": "10/06/2024 13:32:05",
      "content": "<p>I didn't have time to work on my models so I used my submissions to see how the LB-Score changes with different random-seeds for two models. </p>\n<p><strong>Model 1:</strong> <a href=\"https://www.kaggle.com/code/lennarthaupts/cmi-detecting-problematic-digital-behavior\" target=\"_blank\">notebook</a></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Optimized</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.443</td>\n<td>0.455</td>\n<td>0.467</td>\n</tr>\n<tr>\n<td>0.449</td>\n<td>0.470</td>\n<td>0.462</td>\n</tr>\n<tr>\n<td>0.441</td>\n<td>0.456</td>\n<td>0.468</td>\n</tr>\n<tr>\n<td>0.440</td>\n<td>0.468</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>0.443</td>\n<td>0.457</td>\n<td>0.455</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Model 2:</strong></p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Optimized</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.410</td>\n<td>0.462</td>\n<td>0.446</td>\n</tr>\n<tr>\n<td>0.384</td>\n<td>0.449</td>\n<td>0.442</td>\n</tr>\n<tr>\n<td>0.387</td>\n<td>0.452</td>\n<td><strong>0.472</strong></td>\n</tr>\n<tr>\n<td>0.387</td>\n<td>0.459</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>0.405</td>\n<td>0.473</td>\n<td>0.457</td>\n</tr>\n</tbody>\n</table>\n<p>SKF with 10 folds in both cases. My take away is that overfitting the LB can quickly happen with some random luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3010141,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "10/08/2024 17:35:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>, hi everyone,</p>\n<p>I've got these results : </p>\n<p>Train QWK optimized : .473<br>\nOOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.<br>\nPublic LB QWK : .435 (of course I'm desappointed, but something tells me to not trust public LB)</p>\n<p>I'd rather to fit optimization thresholds with train samples instead of OOF samples to control overfitting.</p>\n<p>I'm wondering what are best CV QWK today…</p>",
      "votes": null,
      "replies": [
        {
          "id": 3012243,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "10/08/2024 19:43:05",
          "content": "<p>The leaderboard is not trustworthy here <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, a couple of incorrectly classified points can bump one down and a lucky seed/ model parameter can bounce one up a lot!</p>\n<p>I changed just the reg_lambda in my lightgbm and my leaderboard increased from 0.452 -&gt; 0.463 earlier in the competition</p>",
          "votes": null,
          "replies": [
            {
              "id": 3012265,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "10/08/2024 20:10:43",
              "content": "<p>And did get a better CV score than your previous 0.481 ?</p>\n<p>I'm wondering if we can do really better than 0.48.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3012401,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "10/09/2024 02:37:14",
                  "content": "<p>CV score was nearly the same for both cases (changed by 0.0001) <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3012578,
                      "author_name": "tomyuen",
                      "author_url": "",
                      "post_date": "10/09/2024 07:32:56",
                      "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> </p>\n<blockquote>\n  <p>OOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.</p>\n</blockquote>\n<p>Have you tried to fit different random seed into the exact same model and get a few more CV? some of my models are more stable (+/-0.01) but some has large variance (+/-0.05).</p>\n<p>The training data set is too small that Optuna search cannot be fully trusted. I had a few cases the 'best param' has good cv only at that particular random seed.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3014259,
                          "author_name": "adaubas",
                          "author_url": "",
                          "post_date": "10/11/2024 05:12:43",
                          "content": "<p>Hi <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a>,</p>\n<p>Yes I tried with several random seeds and my OOF QWK optimized after 5 stratified 5 folds is stable : between .004 and .008. </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3014408,
      "author_name": "mohammedahmedxx12",
      "author_url": "",
      "post_date": "10/11/2024 08:11:02",
      "content": "<p>I've had very similar CV/LB scores , I think there will be so much luck involved in this competition . </p>",
      "votes": null,
      "replies": [
        {
          "id": 3014552,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "10/11/2024 10:46:00",
          "content": "<p>This in my opinion, is 99% luck and 1% everything else <a href=\"https://www.kaggle.com/mohammedahmedxx12\" target=\"_blank\">@mohammedahmedxx12</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 3026442,
              "author_name": "davids1992",
              "author_url": "",
              "post_date": "10/23/2024 19:46:29",
              "content": "<p>Is this irony sir? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3026470,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "10/23/2024 20:40:07",
                  "content": "<p>Reality Sir <a href=\"https://www.kaggle.com/davids1992\" target=\"_blank\">@davids1992</a> </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3028000,
      "author_name": "deepaksaldanha",
      "author_url": "",
      "post_date": "10/25/2024 14:14:32",
      "content": "<p>Did you use the actigraphy data in your models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3028644,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "10/26/2024 11:27:31",
          "content": "<p>Yes of course <a href=\"https://www.kaggle.com/deepaksaldanha\" target=\"_blank\">@deepaksaldanha</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 3029053,
              "author_name": "deepaksaldanha",
              "author_url": "",
              "post_date": "10/26/2024 19:38:39",
              "content": "<p>Thanks for your reply, curious to know, since actigraphy data for a major chunk of students in both train and test sets is missing, I'm assuming you've used imputation techniques to fill those gaps, is this why you are seeing so much difference in CV vs LB? could these imputations be creating heavily biased data ?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3035207,
      "author_name": "bsmelbs",
      "author_url": "",
      "post_date": "11/03/2024 06:26:23",
      "content": "<p>Single LGBM regression, the difference between the models is FE and feature selection. Using 10 features atm..</p>\n<p>tried actigraphy data but found it caused increase in train scores but decrease in oof so decided not to use it (yet anyways..).</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>TRAIN</th>\n<th>OOF</th>\n<th>LB</th>\n<th>OOF/LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGBM 1</td>\n<td>0.493</td>\n<td>0.478</td>\n<td>0.455</td>\n<td>0.466</td>\n</tr>\n<tr>\n<td>LGBM 2</td>\n<td>0.490</td>\n<td>0.480</td>\n<td>0.449</td>\n<td>0.464</td>\n</tr>\n<tr>\n<td>LGBM 3</td>\n<td>0.494</td>\n<td>0.481</td>\n<td>0.462</td>\n<td>0.472</td>\n</tr>\n<tr>\n<td>LGBM 4</td>\n<td>0.515</td>\n<td>0.488</td>\n<td>0.463</td>\n<td>0.476</td>\n</tr>\n<tr>\n<td>LGBM 5</td>\n<td>0.504</td>\n<td>0.494</td>\n<td>0.457</td>\n<td>0.476</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3037917,
      "author_name": "alexeyfil",
      "author_url": "",
      "post_date": "11/06/2024 11:15:38",
      "content": "<p>looking at the comments, I saw a column Optimized, what does it mean? Could short describe about it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3038020,
          "author_name": "tomyuen",
          "author_url": "",
          "post_date": "11/06/2024 13:42:29",
          "content": "<p>most people in this game use a regression+ QWK optimizer to classify cases into 0,1,2,3, instead of a simple classification 0 1 2 3, or turning result into integer based on 0.5 1.5 2.5 threshold.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3038032,
              "author_name": "alexeyfil",
              "author_url": "",
              "post_date": "11/06/2024 13:58:45",
              "content": "<p>and as i can see it gives a good result:)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3077114,
      "author_name": "twangp",
      "author_url": "",
      "post_date": "12/20/2024 15:18:21",
      "content": "<p>Great notebook!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2995742": "Hello all,\n\nI think it would be good to discuss CV-LB relations as we embark on our respective model pursuits. CV becomes even more important with evaluation metrics like the Kappa score that rely on hard labels rather than soft probabilities. A CV-LB thread with baseline approaches will help one and all evaluate their models effectively and add value to their pursuits.\n\nI shall commence with my findings insofar, most of which are public. I have used a cv scheme as below-\n`cv = StratifiedKFold(n_splits = 5, random_state = 42, shuffle = True)`\n\n| Model type | Classification/ Regression  | CV | LB |\n| ---               | --- | ------------ | --------------- |\n| LGBM          |  Regression  | 0.464780  |  0.453 |\n| LGBM-XGB |  Regression  |  0.463967 |  0.459 |\n| Catboost     |  Regression  |  0.457608 |  0.442 |\n|  Catboost- LGBM-XGB  |  Regression  |  0.457419  |  0.432 |\n| Catboost- LGBM-XGB   |  Regression  |  0.457849  |  0.431 |\n\n\nMore details are in the below kernels- \n- https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v1\n- https://www.kaggle.com/code/ravi20076/cmi2024-baseline-v2\n- https://www.kaggle.com/code/ravi20076/cmi2024-packages-v1\n\nOthers may add value to the post with their findings as well!\nBest regards!",
    "2996015": "### Performance Of Single Model LGBM Both Have Same Features + Params Optuna Optimized\n\n| Model             | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-----------|----------------|----------------|-------------|\n| First Model LGBM | 0.7301    | 0.4139         | 0.457          | 0.452       |\n| Second Model LGBM | 0.7151    | 0.3942         | 0.452          | 0.455       |\n| Thrid Model LGBM | 0.7875    | 0.4140         | 0.440          | 0.435       |\n\n#### In the end again Shakeup. If we see carefully , my First Model has a Good Train , Test and Optimized Score and a Bad LB Score. While The Second Model has Bad Train, Test and Optimized Score Than First one But it Perform Well on LB. \n\n- CV `StratifiedKFold(n_splits=5, shuffle=True, random_state=42)`\n\nEdited : \n\n| Model                   | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-------------|----------------|----------------|-------------|\n| New LGBM           | 0.5298         | 0.3593             | 0.469                | 0.456\n| Second New LGBM           | 0.5465         | 0.3713             | 0.478                | 0.453\n\n- I'm feeling quite confused about what to focus on. Let's see, we still have three months. @ravi20076",
    "2996024": "I think Second one is Overfitted",
    "2996364": "### Latest Submission | Overfitting LB or What? ! @ravi20076 \n\n| Model             | Train QWK | Validation QWK | Optimized QWK  | Leaderboard |\n|-------------------|-----------|----------------|----------------|-------------|\n| First Model LGBM | 0.7117    | 0.4082         | 0.457          | 0.465     |",
    "2996462": "This itself elicits the structure of the leaderboard @abdmental01",
    "2996944": "Same with me, some of my better CV model results have LB scores in the range of 0.452- 0.455 @abdmental01 \nIs this an [ICR](https://www.kaggle.com/competitions/icr-identify-age-related-conditions/leaderboard)?",
    "2997210": "I didn't participate in ICR, but after seeing the leaderboard, I think the same situation might occur here as well. My Some models that performed well in cross-validation are not doing as well on the leaderboard. It seems like the model is not capturing the data patterns efficiently. Alternatively, it could be that some of the models that are performing well on the leaderboard are better at generalizing and capturing the patterns in the data. @ravi20076 It's hard to chose which one to trust CV or LB\n\nExample of my recent Two Single Models : \n\n| Model    | CV Score | Optimized Score | LB Score |\n|----------|----------|-----------------|----------|\n| Model 1  | 0.4113   | 0.457            | 0.465     |\n| Model 2  | 0.4094   | 0.452           | 0.471    |",
    "2998053": "My results \n1. Model 1 : 10 fold StratifiedKfold same params as public notebooks Cv = 0.4116 , LB = 0.441\n2. Model 2 : 10 fold StratifiedKfold(Changed Seed) , slight modification from public notebooks , CV = 0.4159 , LB = 0.453\n\n@ravi20076 seeing the results and  reading the comments here it feels like we would see similar leaderboard as we saw at the end of AES 2.0 scoring competition needless to say there too the scoring metric was QWK I guess😐",
    "2998105": "It will be worse in my opinion as this dataset is even smaller than the other one and is more noisy @sayedathar11",
    "2998164": "share my exp results.\nI conducted an experiment using Optuna by varying the evaluation metrics to be optimized.\nAs of now, there is no correlation between CV and LB, and the situation is not good.\n\n* Model : LGB Single\n* Split : SKF(5Fold)\n* FE : 136 Col(AfterFeatureSelection)\n\n| Validation QWK | Optimized QWK | LB↑ |\n| --- | --- | --- |\n0.3894|0.445|**0.457**|\n0.398|0.467|0.453|\n0.4077|0.438|0.452|\n0.3578|**0.484**|0.447|\n**0.412**|0.467|0.446|",
    "2998383": "hideyukizushi \nI also have very similar observations- no CV-LB relation till date\nMy best CV model is around yours - **0.4813567** with LB of **0.447** as well. Most of my good CV scores are providing LB of 0.452 - 0.455",
    "2998758": "I just try simple LGBM several times, and CV and LB are almost unrelated, much depending on poor tabular data.",
    "3008329": "I didn't have time to work on my models so I used my submissions to see how the LB-Score changes with different random-seeds for two models. \n\n**Model 1:** [notebook](https://www.kaggle.com/code/lennarthaupts/cmi-detecting-problematic-digital-behavior)\n| CV | Optimized |LB|\n| --- | --- |---|\n| 0.443 |0.455|0.467|\n| 0.449|0.470 |0.462|\n|0.441|0.456|0.468|\n|0.440|0.468|0.466|\n| 0.443|0.457|0.455|\n\n**Model 2:**\n| CV | Optimized |LB|\n| --- | --- |---|\n|0.410|0.462|0.446|\n|0.384|0.449|0.442|\n|0.387| 0.452|**0.472**|\n|0.387|0.459|0.466|\n|0.405|0.473|0.457|\n\nSKF with 10 folds in both cases. My take away is that overfitting the LB can quickly happen with some random luck.",
    "3008559": "I guess these 2 scenarios, and some of the result from Lennart shared here, are indicating that test set is less difficult than what we're expecting from train set.",
    "3010141": "Hi @ravi20076, hi everyone,\n\nI've got these results : \n\nTrain QWK optimized : .473\nOOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.\nPublic LB QWK : .435 (of course I'm desappointed, but something tells me to not trust public LB)\n\nI'd rather to fit optimization thresholds with train samples instead of OOF samples to control overfitting.\n\nI'm wondering what are best CV QWK today...",
    "3012243": "The leaderboard is not trustworthy here @adaubas, a couple of incorrectly classified points can bump one down and a lucky seed/ model parameter can bounce one up a lot!\n\nI changed just the reg_lambda in my lightgbm and my leaderboard increased from 0.452 -> 0.463 earlier in the competition",
    "3012265": "And did get a better CV score than your previous 0.481 ?\n\n I'm wondering if we can do really better than 0.48.",
    "3012401": "CV score was nearly the same for both cases (changed by 0.0001) @adaubas",
    "3012578": "adaubas \n>OOF QWK optimized : .48 (+/- .006) after 5 stratified trainings over 5 folds.\n\nHave you tried to fit different random seed into the exact same model and get a few more CV? some of my models are more stable (+/-0.01) but some has large variance (+/-0.05).\n\nThe training data set is too small that Optuna search cannot be fully trusted. I had a few cases the 'best param' has good cv only at that particular random seed.",
    "3014259": "Hi @tomyuen,\n\nYes I tried with several random seeds and my OOF QWK optimized after 5 stratified 5 folds is stable : between .004 and .008.",
    "3014408": "I've had very similar CV/LB scores , I think there will be so much luck involved in this competition .",
    "3014552": "This in my opinion, is 99% luck and 1% everything else @mohammedahmedxx12",
    "3026442": "Is this irony sir?",
    "3026470": "Reality Sir @davids1992",
    "3028000": "Did you use the actigraphy data in your models?",
    "3028644": "Yes of course @deepaksaldanha",
    "3029053": "Thanks for your reply, curious to know, since actigraphy data for a major chunk of students in both train and test sets is missing, I'm assuming you've used imputation techniques to fill those gaps, is this why you are seeing so much difference in CV vs LB? could these imputations be creating heavily biased data ?",
    "3032412": "ravi20076 What CV is that? A stratified k-fold, or anything else? 0.4813567 is pretty high...",
    "3032444": "davids1992 Stratified 5 fold CV scheme",
    "3035207": "Single LGBM regression, the difference between the models is FE and feature selection. Using 10 features atm..\n\ntried actigraphy data but found it caused increase in train scores but decrease in oof so decided not to use it (yet anyways..).\n\n| Model    | TRAIN | OOF | LB | OOF/LB |\n|--------|-------------|-----------|----------|--------------|\n| LGBM 1  | 0.493       | 0.478     | 0.455    | 0.466        |\n| LGBM 2  | 0.490       | 0.480     | 0.449    | 0.464        |\n| LGBM 3  | 0.494       | 0.481     | 0.462    | 0.472        |\n| LGBM 4  | 0.515       | 0.488     | 0.463    | 0.476        |\n| LGBM 5  | 0.504       | 0.494     | 0.457    | 0.476        |",
    "3037917": "looking at the comments, I saw a column Optimized, what does it mean? Could short describe about it?",
    "3038020": "most people in this game use a regression+ QWK optimizer to classify cases into 0,1,2,3, instead of a simple classification 0 1 2 3, or turning result into integer based on 0.5 1.5 2.5 threshold.",
    "3038032": "and as i can see it gives a good result:)",
    "3077114": "Great notebook!"
  },
  "source": "meta"
}