{
  "id": 395391,
  "title": "CV and LB Scores",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395391",
  "author_name": "yanqiangmiffy",
  "post_date": "2023-03-17T05:52:51.343000",
  "votes": 40,
  "comment_count": 15,
  "views": 0,
  "content": "<p>The cross-validation strategy I use is <strong>StratifiedGroupKfold</strong>, the labels are <strong>[StartHesitation, Turn, Walking]</strong>, and the groups are <strong>subject</strong>,number of folds is <strong>5</strong></p>\n<p>Below is my score records:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.223</td>\n<td>0.171</td>\n</tr>\n<tr>\n<td>0.199</td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>0.240</td>\n<td>0.213</td>\n</tr>\n<tr>\n<td>0.242</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>0.243</td>\n<td>0.252</td>\n</tr>\n<tr>\n<td>0.26910</td>\n<td>0.256</td>\n</tr>\n<tr>\n<td>0.30</td>\n<td>0.269</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2185515,
      "postDate": "2023-03-17T05:52:51.343Z",
      "content": "<p>The cross-validation strategy I use is <strong>StratifiedGroupKfold</strong>, the labels are <strong>[StartHesitation, Turn, Walking]</strong>, and the groups are <strong>subject</strong>,number of folds is <strong>5</strong></p>\n<p>Below is my score records:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.223</td>\n<td>0.171</td>\n</tr>\n<tr>\n<td>0.199</td>\n<td>0.2</td>\n</tr>\n<tr>\n<td>0.240</td>\n<td>0.213</td>\n</tr>\n<tr>\n<td>0.242</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>0.243</td>\n<td>0.252</td>\n</tr>\n<tr>\n<td>0.26910</td>\n<td>0.256</td>\n</tr>\n<tr>\n<td>0.30</td>\n<td>0.269</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "The cross-validation strategy I use is **StratifiedGroupKfold**, the labels are **[StartHesitation, Turn, Walking]**, and the groups are **subject**,number of folds is **5**\n\nBelow is my score records:\n\n|CV  |  LB|\n| --- | --- |\n| 0.223 |0.171  |\n| 0.199 | 0.2|\n| 0.240 |0.213|\n| 0.242 |  0.231|\n| 0.243 | 0.252 |\n| 0.26910 | 0.256 |\n| 0.30 |0.269  |\n",
      "votes": 39
    },
    {
      "id": 2203030,
      "postDate": "2023-03-30T13:23:14Z",
      "content": "<p>StratifiedGroupKfold y:labels groups:subject<br>\nCV:0.244 LB:0.296</p>\n<p>updated 2023/4/12<br>\nCV0.291 LB 0.325</p>\n<p>GroupKfold　nn single model<br>\nCV:0.384 LB:0.38</p>",
      "rawMarkdown": "StratifiedGroupKfold y:labels groups:subject\nCV:0.244 LB:0.296\n\nupdated 2023/4/12\nCV0.291 LB 0.325\n\nGroupKfold　nn single model\nCV:0.384 LB:0.38",
      "votes": 3,
      "replies": [
        {
          "id": 2213984,
          "postDate": "2023-04-08T03:36:07.477Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2214090,
              "postDate": "2023-04-08T06:15:45.940Z",
              "content": "<p>I think it depends on the distribution of the data, if you think in terms of 5folds, some folds are better than the cv score, some are worse.<br>\nI think it is essentially the same thing.<br>\nWhen considering LB, it is more important to look at the correlation with CV.</p>",
              "rawMarkdown": "I think it depends on the distribution of the data, if you think in terms of 5folds, some folds are better than the cv score, some are worse.\nI think it is essentially the same thing.\nWhen considering LB, it is more important to look at the correlation with CV.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2259672,
          "postDate": "2023-05-15T06:53:45.333Z",
          "content": "<p><a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a> How are you implementing StratifiedGroupKFold? Having a hard time wrapping my head around 3 labels</p>",
          "rawMarkdown": "@hiroakifukuse How are you implementing StratifiedGroupKFold? Having a hard time wrapping my head around 3 labels",
          "replies": [
            {
              "id": 2259974,
              "postDate": "2023-05-15T11:21:07.453Z",
              "content": "<p><a href=\"https://www.kaggle.com/exjustice\" target=\"_blank\">@exjustice</a> For each id, create a matrix of [\"start_hesitate\", \"turn\", \"walk\"] and sum these in the row direction to create a column [\"sum\"].<br>\nNext, subtract [\"sum\"] from 1. This will result in a value of 1 when all values of [\"start_hesitate\", \"turn\", \"walk\"] are 0. The column [\"None\"] can be created.<br>\nThis is connected to create a [\"start_hesitate\", \"turn\", \"walk\", \"None\"] matrix. and perform argmax, a label consisting of 0~3 is completed.<br>\nFinally, do this for all the ids and merge them vertically.<br>\nIf you stratify with this label and group by subject, you can do StratifiedGroupKFold.</p>",
              "rawMarkdown": "@exjustice For each id, create a matrix of [\"start_hesitate\", \"turn\", \"walk\"] and sum these in the row direction to create a column [\"sum\"].\nNext, subtract [\"sum\"] from 1. This will result in a value of 1 when all values of [\"start_hesitate\", \"turn\", \"walk\"] are 0. The column [\"None\"] can be created.\nThis is connected to create a [\"start_hesitate\", \"turn\", \"walk\", \"None\"] matrix. and perform argmax, a label consisting of 0~3 is completed.\nFinally, do this for all the ids and merge them vertically.\nIf you stratify with this label and group by subject, you can do StratifiedGroupKFold.",
              "votes": 2
            },
            {
              "id": 2259977,
              "postDate": "2023-05-15T11:30:37.143Z",
              "content": "<p>Thanks for this and the detailed explanation! Did you end up using the label for training as well? Or simply for stratification?</p>",
              "rawMarkdown": "Thanks for this and the detailed explanation! Did you end up using the label for training as well? Or simply for stratification?\n",
              "votes": 1
            },
            {
              "id": 2259990,
              "postDate": "2023-05-15T11:49:34.537Z",
              "content": "<p>It was used only for stratification.</p>",
              "rawMarkdown": "It was used only for stratification."
            },
            {
              "id": 2261453,
              "postDate": "2023-05-16T10:29:51.837Z",
              "content": "<p>Out of curiosity, are you seeing a consistent CV-&gt; LB Correlation. In other words, improved CV MAP-&gt; improved LB performance? Im not observing any consistency at the moment</p>",
              "rawMarkdown": "Out of curiosity, are you seeing a consistent CV-> LB Correlation. In other words, improved CV MAP-> improved LB performance? Im not observing any consistency at the moment"
            }
          ]
        }
      ]
    },
    {
      "id": 2217598,
      "postDate": "2023-04-11T03:25:29.770Z",
      "content": "<p>StratifiedGroupKfold </p>\n<p>Features and hyperparameters are same for two experiments.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>y:visit,  groups:subject</td>\n<td>all</td>\n<td></td>\n<td>0.294</td>\n</tr>\n<tr>\n<td></td>\n<td>fold0</td>\n<td>0.1699</td>\n<td>0.242</td>\n</tr>\n<tr>\n<td></td>\n<td>fold1</td>\n<td>0.1371</td>\n<td>0.277</td>\n</tr>\n<tr>\n<td></td>\n<td>fold2</td>\n<td>0.2958</td>\n<td>0.275</td>\n</tr>\n<tr>\n<td></td>\n<td>fold3</td>\n<td>0.1131</td>\n<td>0.258</td>\n</tr>\n<tr>\n<td>y:label,  groups:subject</td>\n<td>all</td>\n<td></td>\n<td>0.255</td>\n</tr>\n<tr>\n<td></td>\n<td>fold0</td>\n<td>0.2487</td>\n<td>0.255</td>\n</tr>\n<tr>\n<td></td>\n<td>fold1</td>\n<td>0.2161</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td></td>\n<td>fold2</td>\n<td>0.1222</td>\n<td>0.24</td>\n</tr>\n<tr>\n<td></td>\n<td>fold3</td>\n<td>0.189</td>\n<td>0.258</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "StratifiedGroupKfold \n\nFeatures and hyperparameters are same for two experiments.\n\n| | | CV | LB |\n| --- | --- | --- | --- |\n| y:visit,  groups:subject | all |  | 0.294|\n| | fold0 | 0.1699 | 0.242 |\n| | fold1 | 0.1371 | 0.277 |\n| | fold2 | 0.2958 | 0.275 |\n| | fold3 | 0.1131 | 0.258 |\n| y:label,  groups:subject| all |  | 0.255 |\n| | fold0 | 0.2487 | 0.255 |\n| | fold1 | 0.2161 | 0.231 |\n| | fold2 | 0.1222 | 0.24 |\n| | fold3 | 0.189 | 0.258 |  ",
      "votes": 4,
      "replies": [
        {
          "id": 2218373,
          "postDate": "2023-04-11T15:53:20.653Z",
          "content": "<p>does all mean ensemble of folds or training with the whole data? It is too much variance (0.255 vs 0.294) if it is the latter.</p>",
          "rawMarkdown": "does all mean ensemble of folds or training with the whole data? It is too much variance (0.255 vs 0.294) if it is the latter.",
          "votes": 1,
          "replies": [
            {
              "id": 2219097,
              "postDate": "2023-04-12T09:42:59.527Z",
              "content": "<p>'all' means ensemble of folds.</p>",
              "rawMarkdown": "'all' means ensemble of folds.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2186309,
      "postDate": "2023-03-17T17:44:27.240Z",
      "content": "<p>updated </p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.387</td>\n<td>0.290</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "updated \n|  CV| LB |\n| --- | --- |\n|0.387|0.290 |",
      "votes": 2,
      "replies": [
        {
          "id": 2259671,
          "postDate": "2023-05-15T06:53:01.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/quincyqiang\" target=\"_blank\">@quincyqiang</a> How are you implementing stratifiedGroupKFold? having a hard time wrapping my head around the three labels</p>",
          "rawMarkdown": "@quincyqiang How are you implementing stratifiedGroupKFold? having a hard time wrapping my head around the three labels"
        }
      ]
    },
    {
      "id": 2217703,
      "postDate": "2023-04-11T05:38:39.340Z",
      "content": "<p>you are doing a great job💪</p>",
      "rawMarkdown": "you are doing a great job💪"
    },
    {
      "id": 2230070,
      "postDate": "2023-04-22T01:35:02.033Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2203030,
      "author_name": "Hi F",
      "author_url": "",
      "post_date": "2023-03-30T13:23:14",
      "content": "<p>StratifiedGroupKfold y:labels groups:subject<br>\nCV:0.244 LB:0.296</p>\n<p>updated 2023/4/12<br>\nCV0.291 LB 0.325</p>\n<p>GroupKfold　nn single model<br>\nCV:0.384 LB:0.38</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2213984,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-04-08T03:36:07.477000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2214090,
              "author_name": "Hi F",
              "author_url": "",
              "post_date": "2023-04-08T06:15:45.940000",
              "content": "<p>I think it depends on the distribution of the data, if you think in terms of 5folds, some folds are better than the cv score, some are worse.<br>\nI think it is essentially the same thing.<br>\nWhen considering LB, it is more important to look at the correlation with CV.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2259672,
          "author_name": "Yijie Xu",
          "author_url": "",
          "post_date": "2023-05-15T06:53:45.333000",
          "content": "<p><a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a> How are you implementing StratifiedGroupKFold? Having a hard time wrapping my head around 3 labels</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2259974,
              "author_name": "Hi F",
              "author_url": "",
              "post_date": "2023-05-15T11:21:07.453000",
              "content": "<p><a href=\"https://www.kaggle.com/exjustice\" target=\"_blank\">@exjustice</a> For each id, create a matrix of [\"start_hesitate\", \"turn\", \"walk\"] and sum these in the row direction to create a column [\"sum\"].<br>\nNext, subtract [\"sum\"] from 1. This will result in a value of 1 when all values of [\"start_hesitate\", \"turn\", \"walk\"] are 0. The column [\"None\"] can be created.<br>\nThis is connected to create a [\"start_hesitate\", \"turn\", \"walk\", \"None\"] matrix. and perform argmax, a label consisting of 0~3 is completed.<br>\nFinally, do this for all the ids and merge them vertically.<br>\nIf you stratify with this label and group by subject, you can do StratifiedGroupKFold.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2259977,
              "author_name": "Yijie Xu",
              "author_url": "",
              "post_date": "2023-05-15T11:30:37.143000",
              "content": "<p>Thanks for this and the detailed explanation! Did you end up using the label for training as well? Or simply for stratification?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2259990,
              "author_name": "Hi F",
              "author_url": "",
              "post_date": "2023-05-15T11:49:34.537000",
              "content": "<p>It was used only for stratification.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261453,
              "author_name": "Yijie Xu",
              "author_url": "",
              "post_date": "2023-05-16T10:29:51.837000",
              "content": "<p>Out of curiosity, are you seeing a consistent CV-&gt; LB Correlation. In other words, improved CV MAP-&gt; improved LB performance? Im not observing any consistency at the moment</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2217598,
      "author_name": "mrt💪🥺",
      "author_url": "",
      "post_date": "2023-04-11T03:25:29.770000",
      "content": "<p>StratifiedGroupKfold </p>\n<p>Features and hyperparameters are same for two experiments.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>y:visit,  groups:subject</td>\n<td>all</td>\n<td></td>\n<td>0.294</td>\n</tr>\n<tr>\n<td></td>\n<td>fold0</td>\n<td>0.1699</td>\n<td>0.242</td>\n</tr>\n<tr>\n<td></td>\n<td>fold1</td>\n<td>0.1371</td>\n<td>0.277</td>\n</tr>\n<tr>\n<td></td>\n<td>fold2</td>\n<td>0.2958</td>\n<td>0.275</td>\n</tr>\n<tr>\n<td></td>\n<td>fold3</td>\n<td>0.1131</td>\n<td>0.258</td>\n</tr>\n<tr>\n<td>y:label,  groups:subject</td>\n<td>all</td>\n<td></td>\n<td>0.255</td>\n</tr>\n<tr>\n<td></td>\n<td>fold0</td>\n<td>0.2487</td>\n<td>0.255</td>\n</tr>\n<tr>\n<td></td>\n<td>fold1</td>\n<td>0.2161</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td></td>\n<td>fold2</td>\n<td>0.1222</td>\n<td>0.24</td>\n</tr>\n<tr>\n<td></td>\n<td>fold3</td>\n<td>0.189</td>\n<td>0.258</td>\n</tr>\n</tbody>\n</table>",
      "votes": 4,
      "replies": [
        {
          "id": 2218373,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2023-04-11T15:53:20.653000",
          "content": "<p>does all mean ensemble of folds or training with the whole data? It is too much variance (0.255 vs 0.294) if it is the latter.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2219097,
              "author_name": "mrt💪🥺",
              "author_url": "",
              "post_date": "2023-04-12T09:42:59.527000",
              "content": "<p>'all' means ensemble of folds.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2186309,
      "author_name": "yanqiangmiffy",
      "author_url": "",
      "post_date": "2023-03-17T17:44:27.240000",
      "content": "<p>updated </p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.387</td>\n<td>0.290</td>\n</tr>\n</tbody>\n</table>",
      "votes": 2,
      "replies": [
        {
          "id": 2259671,
          "author_name": "Yijie Xu",
          "author_url": "",
          "post_date": "2023-05-15T06:53:01.507000",
          "content": "<p><a href=\"https://www.kaggle.com/quincyqiang\" target=\"_blank\">@quincyqiang</a> How are you implementing stratifiedGroupKFold? having a hard time wrapping my head around the three labels</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2217703,
      "author_name": "JoeJoeW",
      "author_url": "",
      "post_date": "2023-04-11T05:38:39.340000",
      "content": "<p>you are doing a great job💪</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2230070,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-22T01:35:02.033000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2185515": "The cross-validation strategy I use is **StratifiedGroupKfold**, the labels are **[StartHesitation, Turn, Walking]**, and the groups are **subject**,number of folds is **5**\n\nBelow is my score records:\n\n|CV  |  LB|\n| --- | --- |\n| 0.223 |0.171  |\n| 0.199 | 0.2|\n| 0.240 |0.213|\n| 0.242 |  0.231|\n| 0.243 | 0.252 |\n| 0.26910 | 0.256 |\n| 0.30 |0.269  |\n",
    "2203030": "StratifiedGroupKfold y:labels groups:subject\nCV:0.244 LB:0.296\n\nupdated 2023/4/12\nCV0.291 LB 0.325\n\nGroupKfold　nn single model\nCV:0.384 LB:0.38",
    "2217598": "StratifiedGroupKfold \n\nFeatures and hyperparameters are same for two experiments.\n\n| | | CV | LB |\n| --- | --- | --- | --- |\n| y:visit,  groups:subject | all |  | 0.294|\n| | fold0 | 0.1699 | 0.242 |\n| | fold1 | 0.1371 | 0.277 |\n| | fold2 | 0.2958 | 0.275 |\n| | fold3 | 0.1131 | 0.258 |\n| y:label,  groups:subject| all |  | 0.255 |\n| | fold0 | 0.2487 | 0.255 |\n| | fold1 | 0.2161 | 0.231 |\n| | fold2 | 0.1222 | 0.24 |\n| | fold3 | 0.189 | 0.258 |  ",
    "2186309": "updated \n|  CV| LB |\n| --- | --- |\n|0.387|0.290 |",
    "2217703": "you are doing a great job💪",
    "2230070": ""
  }
}