{
  "id": 394426,
  "title": "Reliable CV strategy?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/394426",
  "author_name": "",
  "post_date": "2023-03-13T13:47:09.958080500Z",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Anyone found some way for a reliable CV strategy that gives similar scores to LB?<br>\nI tried splitting by Subject but then I got a much lower local results than LB.</p>",
  "messages": [
    {
      "id": "2179940",
      "postDate": "03/13/2023 13:47:09",
      "content": "<p>Anyone found some way for a reliable CV strategy that gives similar scores to LB?<br>\nI tried splitting by Subject but then I got a much lower local results than LB.</p>",
      "rawMarkdown": "Anyone found some way for a reliable CV strategy that gives similar scores to LB?\nI tried splitting by Subject but then I got a much lower local results than LB.",
      "votes": null
    },
    {
      "id": "2183968",
      "postDate": "03/16/2023 04:26:44",
      "content": "<p>This doesn't directly answer your question, but I'm currently just breaking all the files into non-overlapping chunks and using a random subset of the chunks for cross validation (so there are situations where part of one file is used for training and some other part of the same file is used for testing). I'm observing local mAP scores that are much higher than the scores I'm getting on the leaderboard, so I figure there is probably <em>something</em> my training and cross-validation splits have in common that I'm overfitting to. Common subjects are one likely explanation.</p>\n<p>Some example scores that illustrate the fact I'm likely overfitting to something that isn't being picked up by my cross-validation are included below. I'm not too terribly worried about this <em>yet</em> because the scores I'm getting locally are positively correlated with the LB scores (with a pearson correlation coefficient of 0.77), but its certainly not ideal.</p>\n<table>\n<thead>\n<tr>\n<th>Local CV score</th>\n<th>Public LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.502</td>\n<td>0.224</td>\n</tr>\n<tr>\n<td>0.44</td>\n<td>0.245</td>\n</tr>\n<tr>\n<td>0.456</td>\n<td>0.251</td>\n</tr>\n<tr>\n<td>0.724</td>\n<td>0.271</td>\n</tr>\n<tr>\n<td>0.717</td>\n<td>0.27</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "This doesn't directly answer your question, but I'm currently just breaking all the files into non-overlapping chunks and using a random subset of the chunks for cross validation (so there are situations where part of one file is used for training and some other part of the same file is used for testing). I'm observing local mAP scores that are much higher than the scores I'm getting on the leaderboard, so I figure there is probably *something* my training and cross-validation splits have in common that I'm overfitting to. Common subjects are one likely explanation.\n\nSome example scores that illustrate the fact I'm likely overfitting to something that isn't being picked up by my cross-validation are included below. I'm not too terribly worried about this *yet* because the scores I'm getting locally are positively correlated with the LB scores (with a pearson correlation coefficient of 0.77), but its certainly not ideal.\n\n| Local CV score | Public LB score |\n| --- | --- |\n| 0.502 | 0.224 |\n| 0.44 | 0.245 |\n| 0.456 | 0.251 |\n| 0.724 | 0.271 |\n| 0.717 | 0.27 |",
      "votes": null
    },
    {
      "id": "2185016",
      "postDate": "03/16/2023 18:52:00",
      "content": "<p>I have the same situation, when I used GroupKFold using subject (this is how the host split the data) the model does not trained at all. If I use StratifiedKfold cv is great, but lb is much lower. Good cv strategy is going to be one of the keys in this competition.</p>",
      "rawMarkdown": "I have the same situation, when I used GroupKFold using subject (this is how the host split the data) the model does not trained at all. If I use StratifiedKfold cv is great, but lb is much lower. Good cv strategy is going to be one of the keys in this competition.",
      "votes": null
    },
    {
      "id": "2185261",
      "postDate": "03/16/2023 22:53:18",
      "content": "<p>It's really nice to see the responses here, because I haven't submitted any results yet, but my baseline results were unreasonably high and I've been trying to figure out how I overtrained so badly on such simple methods.  Thanks for bringing this up.</p>",
      "rawMarkdown": "It's really nice to see the responses here, because I haven't submitted any results yet, but my baseline results were unreasonably high and I've been trying to figure out how I overtrained so badly on such simple methods.  Thanks for bringing this up.",
      "votes": null
    },
    {
      "id": "2186399",
      "postDate": "03/17/2023 18:52:02",
      "content": "<p>Chance level locally, chance level on LB, so everything works fine for me</p>",
      "rawMarkdown": "Chance level locally, chance level on LB, so everything works fine for me",
      "votes": null
    },
    {
      "id": "2190070",
      "postDate": "03/21/2023 02:25:16",
      "content": "<p>I use groupkfold on subject, currently it's reliable</p>",
      "rawMarkdown": "I use groupkfold on subject, currently it's reliable",
      "votes": null
    },
    {
      "id": "2201011",
      "postDate": "03/29/2023 01:35:34",
      "content": "<p>Was anyone able to figure out a reliable cross validation strategy? I used GroupKFold but it is still a good amount higher than the leader board.</p>",
      "rawMarkdown": "Was anyone able to figure out a reliable cross validation strategy? I used GroupKFold but it is still a good amount higher than the leader board.",
      "votes": null
    },
    {
      "id": "2274119",
      "postDate": "05/25/2023 16:25:44",
      "content": "<p>I'm using StratifiedGroupKFold, and I've got local CV 0.337 with LB 0.238.</p>",
      "rawMarkdown": "I'm using StratifiedGroupKFold, and I've got local CV 0.337 with LB 0.238.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2183968,
      "author_name": "jsday96",
      "author_url": "",
      "post_date": "03/16/2023 04:26:44",
      "content": "<p>This doesn't directly answer your question, but I'm currently just breaking all the files into non-overlapping chunks and using a random subset of the chunks for cross validation (so there are situations where part of one file is used for training and some other part of the same file is used for testing). I'm observing local mAP scores that are much higher than the scores I'm getting on the leaderboard, so I figure there is probably <em>something</em> my training and cross-validation splits have in common that I'm overfitting to. Common subjects are one likely explanation.</p>\n<p>Some example scores that illustrate the fact I'm likely overfitting to something that isn't being picked up by my cross-validation are included below. I'm not too terribly worried about this <em>yet</em> because the scores I'm getting locally are positively correlated with the LB scores (with a pearson correlation coefficient of 0.77), but its certainly not ideal.</p>\n<table>\n<thead>\n<tr>\n<th>Local CV score</th>\n<th>Public LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.502</td>\n<td>0.224</td>\n</tr>\n<tr>\n<td>0.44</td>\n<td>0.245</td>\n</tr>\n<tr>\n<td>0.456</td>\n<td>0.251</td>\n</tr>\n<tr>\n<td>0.724</td>\n<td>0.271</td>\n</tr>\n<tr>\n<td>0.717</td>\n<td>0.27</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2185016,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "03/16/2023 18:52:00",
          "content": "<p>I have the same situation, when I used GroupKFold using subject (this is how the host split the data) the model does not trained at all. If I use StratifiedKfold cv is great, but lb is much lower. Good cv strategy is going to be one of the keys in this competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2274119,
              "author_name": "atamazian",
              "author_url": "",
              "post_date": "05/25/2023 16:25:44",
              "content": "<p>I'm using StratifiedGroupKFold, and I've got local CV 0.337 with LB 0.238.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2185261,
      "author_name": "josephreid",
      "author_url": "",
      "post_date": "03/16/2023 22:53:18",
      "content": "<p>It's really nice to see the responses here, because I haven't submitted any results yet, but my baseline results were unreasonably high and I've been trying to figure out how I overtrained so badly on such simple methods.  Thanks for bringing this up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2186399,
      "author_name": "richarddinga",
      "author_url": "",
      "post_date": "03/17/2023 18:52:02",
      "content": "<p>Chance level locally, chance level on LB, so everything works fine for me</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2190070,
      "author_name": "xzj19013742",
      "author_url": "",
      "post_date": "03/21/2023 02:25:16",
      "content": "<p>I use groupkfold on subject, currently it's reliable</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2201011,
      "author_name": "jayticku",
      "author_url": "",
      "post_date": "03/29/2023 01:35:34",
      "content": "<p>Was anyone able to figure out a reliable cross validation strategy? I used GroupKFold but it is still a good amount higher than the leader board.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2179940": "Anyone found some way for a reliable CV strategy that gives similar scores to LB?\nI tried splitting by Subject but then I got a much lower local results than LB.",
    "2183968": "This doesn't directly answer your question, but I'm currently just breaking all the files into non-overlapping chunks and using a random subset of the chunks for cross validation (so there are situations where part of one file is used for training and some other part of the same file is used for testing). I'm observing local mAP scores that are much higher than the scores I'm getting on the leaderboard, so I figure there is probably *something* my training and cross-validation splits have in common that I'm overfitting to. Common subjects are one likely explanation.\n\nSome example scores that illustrate the fact I'm likely overfitting to something that isn't being picked up by my cross-validation are included below. I'm not too terribly worried about this *yet* because the scores I'm getting locally are positively correlated with the LB scores (with a pearson correlation coefficient of 0.77), but its certainly not ideal.\n\n| Local CV score | Public LB score |\n| --- | --- |\n| 0.502 | 0.224 |\n| 0.44 | 0.245 |\n| 0.456 | 0.251 |\n| 0.724 | 0.271 |\n| 0.717 | 0.27 |",
    "2185016": "I have the same situation, when I used GroupKFold using subject (this is how the host split the data) the model does not trained at all. If I use StratifiedKfold cv is great, but lb is much lower. Good cv strategy is going to be one of the keys in this competition.",
    "2185261": "It's really nice to see the responses here, because I haven't submitted any results yet, but my baseline results were unreasonably high and I've been trying to figure out how I overtrained so badly on such simple methods.  Thanks for bringing this up.",
    "2186399": "Chance level locally, chance level on LB, so everything works fine for me",
    "2190070": "I use groupkfold on subject, currently it's reliable",
    "2201011": "Was anyone able to figure out a reliable cross validation strategy? I used GroupKFold but it is still a good amount higher than the leader board.",
    "2274119": "I'm using StratifiedGroupKFold, and I've got local CV 0.337 with LB 0.238."
  },
  "source": "meta"
}