{
  "id": 478277,
  "title": "Whats with the CV in this comp?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/478277",
  "author_name": "Cody_Null",
  "post_date": "2024-02-20T01:30:55.290000",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>As expected, as the competition has gone on LB scores have gotten much better. However, I have noticed that the CV has not massively decreased. I had expected that improving LB to this level would be due to a closing in the gap but it doesn't seem that is the case. To understand this, I thought maybe more CV scores would be helpful. </p>\n<p>CV             LB</p>\n<p>0.592       0.38<br>\n0.578       0.43<br>\n0.603       0.41<br>\n0.61          0.45<br>\n0.67          0.5</p>",
  "messages": [
    {
      "id": 2659570,
      "postDate": "2024-02-20T01:30:55.290Z",
      "content": "<p>As expected, as the competition has gone on LB scores have gotten much better. However, I have noticed that the CV has not massively decreased. I had expected that improving LB to this level would be due to a closing in the gap but it doesn't seem that is the case. To understand this, I thought maybe more CV scores would be helpful. </p>\n<p>CV             LB</p>\n<p>0.592       0.38<br>\n0.578       0.43<br>\n0.603       0.41<br>\n0.61          0.45<br>\n0.67          0.5</p>",
      "rawMarkdown": "As expected, as the competition has gone on LB scores have gotten much better. However, I have noticed that the CV has not massively decreased. I had expected that improving LB to this level would be due to a closing in the gap but it doesn't seem that is the case. To understand this, I thought maybe more CV scores would be helpful. \n\nCV             LB\n\n0.592       0.38\n0.578       0.43\n0.603       0.41\n0.61          0.45\n0.67          0.5\n\n",
      "votes": 7
    },
    {
      "id": 2660919,
      "postDate": "2024-02-20T21:50:37.703Z",
      "content": "<p>I have a hypothesis: as seen in this <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">notebook</a>, I've trained EfficientNetB0 in two stages, where the data is split into two populations based on the total number of votes. The population with less total votes likely has peaked distributions (it is necessarily peaked when total votes = 1), so I train with this data in the first stage. The remaining data is used in the second stage of training. As expected, the final model has a lower CV score on the second population. The public LB score (0.39) is less than the first population CV score (not calculated), but greater than the second population CV score (0.29).</p>\n<p>This seems to indicate the public LB dataset contains less peaked distributions (more votes per sample) than the training dataset. If true, it's possible the private LB dataset might not showcase similar distributions, which could imply a surprising churn when private LB is released.</p>\n<p>What are your thoughts?</p>",
      "rawMarkdown": "I have a hypothesis: as seen in this [notebook](https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39), I've trained EfficientNetB0 in two stages, where the data is split into two populations based on the total number of votes. The population with less total votes likely has peaked distributions (it is necessarily peaked when total votes = 1), so I train with this data in the first stage. The remaining data is used in the second stage of training. As expected, the final model has a lower CV score on the second population. The public LB score (0.39) is less than the first population CV score (not calculated), but greater than the second population CV score (0.29).\n\nThis seems to indicate the public LB dataset contains less peaked distributions (more votes per sample) than the training dataset. If true, it's possible the private LB dataset might not showcase similar distributions, which could imply a surprising churn when private LB is released.\n\nWhat are your thoughts?",
      "votes": 5,
      "replies": [
        {
          "id": 2661085,
          "postDate": "2024-02-21T02:40:49.570Z",
          "content": "<p>Yep I do see that relationship as well, what cv strategy are you using?</p>",
          "rawMarkdown": "Yep I do see that relationship as well, what cv strategy are you using?",
          "votes": 1,
          "replies": [
            {
              "id": 2661616,
              "postDate": "2024-02-21T11:25:27.640Z",
              "content": "<p>I’m getting the best results with GKF on patient_id. I’ve tried SGKF on patient_id and spectrogram_id, but not as good for me. </p>",
              "rawMarkdown": "I’m getting the best results with GKF on patient_id. I’ve tried SGKF on patient_id and spectrogram_id, but not as good for me. ",
              "votes": 1
            },
            {
              "id": 2663928,
              "postDate": "2024-02-22T18:14:14.037Z",
              "content": "<p>Are you worried about overfitting in the 2 stage approach? </p>",
              "rawMarkdown": "Are you worried about overfitting in the 2 stage approach? ",
              "votes": 1
            },
            {
              "id": 2663996,
              "postDate": "2024-02-22T18:44:33.717Z",
              "content": "<p>It is certainly possible. Stage 1 contains a significant amount of seizure examples, so if the private dataset contains many seizures with all votes for seizures (peaked distribution), then my 2 stage approach may not perform well. </p>\n<p>However, given the nature of KL-divergence, we are severely punished for predicting a class with nearly zero probability when the class has nonzero probability. In my opinion, biasing Stage 2 data (more votes in total) should make the model more robust to unseen data.</p>",
              "rawMarkdown": "It is certainly possible. Stage 1 contains a significant amount of seizure examples, so if the private dataset contains many seizures with all votes for seizures (peaked distribution), then my 2 stage approach may not perform well. \n\nHowever, given the nature of KL-divergence, we are severely punished for predicting a class with nearly zero probability when the class has nonzero probability. In my opinion, biasing Stage 2 data (more votes in total) should make the model more robust to unseen data.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2659626,
      "postDate": "2024-02-20T03:37:16.837Z",
      "content": "<ul>\n<li>In my experiment, using EfficientNet for 8spectrograms (Kaggle spectrograms and EEG spectrograms) </li>\n<li>split. GroupKFold on patient_id<ul>\n<li>CV    LB</li>\n<li>0.599441616    0.38</li>\n<li>0.615207684    0.40</li>\n<li>0.622906769    0.41</li></ul></li>\n</ul>",
      "rawMarkdown": "* In my experiment, using EfficientNet for 8spectrograms (Kaggle spectrograms and EEG spectrograms) \n* split. GroupKFold on patient_id\n * CV\tLB\n * 0.599441616\t0.38\n * 0.615207684\t0.40\n * 0.622906769\t0.41",
      "votes": 1
    },
    {
      "id": 2659623,
      "postDate": "2024-02-20T03:31:34.747Z",
      "content": "<p>No ensembles</p>\n<table>\n<thead>\n<tr>\n<th>5-fold GKF</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.577395</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.578273</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.544300</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.536900</td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "No ensembles\n\n|   5-fold GKF  | LB |\n|--------------|----------|\n|   0.577395   |   0.41   |\n|   0.578273   |   0.40   |\n|   0.544300   |   0.40   |\n|   0.536900   |   0.40   |\n",
      "votes": 1
    },
    {
      "id": 2659970,
      "postDate": "2024-02-20T08:52:27.923Z",
      "content": "<p>First of all define what CV strategy do you imply. GKF/SGKF on patient unique ids? Or something else</p>",
      "rawMarkdown": "First of all define what CV strategy do you imply. GKF/SGKF on patient unique ids? Or something else",
      "votes": 2,
      "replies": [
        {
          "id": 2660293,
          "postDate": "2024-02-20T14:12:37.330Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2662553,
      "postDate": "2024-02-22T01:23:03.957Z",
      "content": "<p>(All are single 5-fold models)</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6926</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.6432</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.6332</td>\n<td>0.37</td>\n</tr>\n<tr>\n<td>0.7547</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>0.7306</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.7208</td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>\n<p>It seems my CVs compared to LBs are generally higher than what people got… Any idea what could be the reason? Are these models overfit the public LB?</p>",
      "rawMarkdown": "(All are single 5-fold models)\n \n| CV | LB |\n| ---- | ---- |\n| 0.6926 | 0.41 |\n| 0.6432 | 0.40 |\n| 0.6332 | 0.37 |\n| 0.7547 | 0.48 |\n| 0.7306 | 0.41 |\n| 0.7208 | 0.40 |\n\n\nIt seems my CVs compared to LBs are generally higher than what people got... Any idea what could be the reason? Are these models overfit the public LB?"
    },
    {
      "id": 2660210,
      "postDate": "2024-02-20T12:58:13.083Z",
      "content": "<p>In my experiments, CV is very confusing. Especially when using learning rate schedulers (better LB, worse CV) and data augmentations (worse CV, better LB).</p>",
      "rawMarkdown": "In my experiments, CV is very confusing. Especially when using learning rate schedulers (better LB, worse CV) and data augmentations (worse CV, better LB)."
    },
    {
      "id": 2659950,
      "postDate": "2024-02-20T08:35:07.357Z",
      "content": "<p>I also noticed this when using Data Augmentations.</p>\n<p>CV score increased but my LB score improved.</p>\n<p>CV 0.57 -&gt; 0.60<br>\nLB 0.42 -&gt; 0.40</p>",
      "rawMarkdown": "I also noticed this when using Data Augmentations.\n\nCV score increased but my LB score improved.\n\nCV 0.57 -> 0.60\nLB 0.42 -> 0.40\n"
    }
  ],
  "comments": [
    {
      "id": 2660919,
      "author_name": "Sean R.B. Bearden, Ph.D.",
      "author_url": "",
      "post_date": "2024-02-20T21:50:37.703000",
      "content": "<p>I have a hypothesis: as seen in this <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">notebook</a>, I've trained EfficientNetB0 in two stages, where the data is split into two populations based on the total number of votes. The population with less total votes likely has peaked distributions (it is necessarily peaked when total votes = 1), so I train with this data in the first stage. The remaining data is used in the second stage of training. As expected, the final model has a lower CV score on the second population. The public LB score (0.39) is less than the first population CV score (not calculated), but greater than the second population CV score (0.29).</p>\n<p>This seems to indicate the public LB dataset contains less peaked distributions (more votes per sample) than the training dataset. If true, it's possible the private LB dataset might not showcase similar distributions, which could imply a surprising churn when private LB is released.</p>\n<p>What are your thoughts?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2661085,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2024-02-21T02:40:49.570000",
          "content": "<p>Yep I do see that relationship as well, what cv strategy are you using?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2661616,
              "author_name": "Sean R.B. Bearden, Ph.D.",
              "author_url": "",
              "post_date": "2024-02-21T11:25:27.640000",
              "content": "<p>I’m getting the best results with GKF on patient_id. I’ve tried SGKF on patient_id and spectrogram_id, but not as good for me. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2663928,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2024-02-22T18:14:14.037000",
              "content": "<p>Are you worried about overfitting in the 2 stage approach? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2663996,
              "author_name": "Sean R.B. Bearden, Ph.D.",
              "author_url": "",
              "post_date": "2024-02-22T18:44:33.717000",
              "content": "<p>It is certainly possible. Stage 1 contains a significant amount of seizure examples, so if the private dataset contains many seizures with all votes for seizures (peaked distribution), then my 2 stage approach may not perform well. </p>\n<p>However, given the nature of KL-divergence, we are severely punished for predicting a class with nearly zero probability when the class has nonzero probability. In my opinion, biasing Stage 2 data (more votes in total) should make the model more robust to unseen data.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2659626,
      "author_name": "yukiZ",
      "author_url": "",
      "post_date": "2024-02-20T03:37:16.837000",
      "content": "<ul>\n<li>In my experiment, using EfficientNet for 8spectrograms (Kaggle spectrograms and EEG spectrograms) </li>\n<li>split. GroupKFold on patient_id<ul>\n<li>CV    LB</li>\n<li>0.599441616    0.38</li>\n<li>0.615207684    0.40</li>\n<li>0.622906769    0.41</li></ul></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2659623,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-02-20T03:31:34.747000",
      "content": "<p>No ensembles</p>\n<table>\n<thead>\n<tr>\n<th>5-fold GKF</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.577395</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.578273</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.544300</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.536900</td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2659970,
      "author_name": "Yurnero",
      "author_url": "",
      "post_date": "2024-02-20T08:52:27.923000",
      "content": "<p>First of all define what CV strategy do you imply. GKF/SGKF on patient unique ids? Or something else</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2660293,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-20T14:12:37.330000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2662553,
      "author_name": "Steven_Y",
      "author_url": "",
      "post_date": "2024-02-22T01:23:03.957000",
      "content": "<p>(All are single 5-fold models)</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6926</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.6432</td>\n<td>0.40</td>\n</tr>\n<tr>\n<td>0.6332</td>\n<td>0.37</td>\n</tr>\n<tr>\n<td>0.7547</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>0.7306</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>0.7208</td>\n<td>0.40</td>\n</tr>\n</tbody>\n</table>\n<p>It seems my CVs compared to LBs are generally higher than what people got… Any idea what could be the reason? Are these models overfit the public LB?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2660210,
      "author_name": "Koolo",
      "author_url": "",
      "post_date": "2024-02-20T12:58:13.083000",
      "content": "<p>In my experiments, CV is very confusing. Especially when using learning rate schedulers (better LB, worse CV) and data augmentations (worse CV, better LB).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2659950,
      "author_name": "Feida Wei",
      "author_url": "",
      "post_date": "2024-02-20T08:35:07.357000",
      "content": "<p>I also noticed this when using Data Augmentations.</p>\n<p>CV score increased but my LB score improved.</p>\n<p>CV 0.57 -&gt; 0.60<br>\nLB 0.42 -&gt; 0.40</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2659570": "As expected, as the competition has gone on LB scores have gotten much better. However, I have noticed that the CV has not massively decreased. I had expected that improving LB to this level would be due to a closing in the gap but it doesn't seem that is the case. To understand this, I thought maybe more CV scores would be helpful. \n\nCV             LB\n\n0.592       0.38\n0.578       0.43\n0.603       0.41\n0.61          0.45\n0.67          0.5\n\n",
    "2660919": "I have a hypothesis: as seen in this [notebook](https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39), I've trained EfficientNetB0 in two stages, where the data is split into two populations based on the total number of votes. The population with less total votes likely has peaked distributions (it is necessarily peaked when total votes = 1), so I train with this data in the first stage. The remaining data is used in the second stage of training. As expected, the final model has a lower CV score on the second population. The public LB score (0.39) is less than the first population CV score (not calculated), but greater than the second population CV score (0.29).\n\nThis seems to indicate the public LB dataset contains less peaked distributions (more votes per sample) than the training dataset. If true, it's possible the private LB dataset might not showcase similar distributions, which could imply a surprising churn when private LB is released.\n\nWhat are your thoughts?",
    "2659626": "* In my experiment, using EfficientNet for 8spectrograms (Kaggle spectrograms and EEG spectrograms) \n* split. GroupKFold on patient_id\n * CV\tLB\n * 0.599441616\t0.38\n * 0.615207684\t0.40\n * 0.622906769\t0.41",
    "2659623": "No ensembles\n\n|   5-fold GKF  | LB |\n|--------------|----------|\n|   0.577395   |   0.41   |\n|   0.578273   |   0.40   |\n|   0.544300   |   0.40   |\n|   0.536900   |   0.40   |\n",
    "2659970": "First of all define what CV strategy do you imply. GKF/SGKF on patient unique ids? Or something else",
    "2662553": "(All are single 5-fold models)\n \n| CV | LB |\n| ---- | ---- |\n| 0.6926 | 0.41 |\n| 0.6432 | 0.40 |\n| 0.6332 | 0.37 |\n| 0.7547 | 0.48 |\n| 0.7306 | 0.41 |\n| 0.7208 | 0.40 |\n\n\nIt seems my CVs compared to LBs are generally higher than what people got... Any idea what could be the reason? Are these models overfit the public LB?",
    "2660210": "In my experiments, CV is very confusing. Especially when using learning rate schedulers (better LB, worse CV) and data augmentations (worse CV, better LB).",
    "2659950": "I also noticed this when using Data Augmentations.\n\nCV score increased but my LB score improved.\n\nCV 0.57 -> 0.60\nLB 0.42 -> 0.40\n"
  }
}