{
  "id": 467915,
  "title": "CV vs LB Scores",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467915",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2024-01-14T15:45:45.626000",
  "votes": 60,
  "comment_count": 57,
  "views": 0,
  "content": "<h1>Ensemble Models:</h1>\n<table>\n<thead>\n<tr>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n<th>User</th>\n<th>Comment message</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GKF</td>\n<td><strong>0.500</strong></td>\n<td><strong>0.35</strong></td>\n<td><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></td>\n<td>combine all the ideas from my starter notebooks</td>\n<td><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2609612\" target=\"_blank\">comment here</a></td>\n</tr>\n<tr>\n<td>SGKF</td>\n<td>~</td>\n<td><strong>0.37</strong></td>\n<td><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a></td>\n<td>128,256,4 → 256,256,2 → 512,512,2(resize)</td>\n<td><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2608138\" target=\"_blank\">comment here</a></td>\n</tr>\n</tbody>\n</table>\n<h1>EEG Sequenced based</h1>\n<p>Model : 1D sequence multilabel classifier (Stratified Group KFold)<br>\nSequence:  Eegs pair sequences <a href=\"https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features\" target=\"_blank\">Notebook - Eegs Pairing Analysis &amp; Features</a> for dataset<br>\nCV: ..<br>\nLB: 0.79</p>\n<hr>\n<p>Model : MLP multilabel classifier (Group KFold)<br>\nSequence: Eegs pair sequences Notebook - How To Make Spectrogram from EEG for dataset by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 1.037<br>\nLB: 0.77</p>\n<hr>\n<p>Model : 1D-CNN (Group KFold)<br>\nSequence: Eegs pair sequences Notebook - <a href=\"https://www.kaggle.com/code/anmolgarg1998/hms-cnn1d-inference/notebook\" target=\"_blank\">HMS_cnn1d_inference\n</a> by <a href=\"https://www.kaggle.com/anmolgarg1998\" target=\"_blank\">@anmolgarg1998</a><br>\nCV: -<br>\nLB: 0.69</p>\n<hr>\n<p>Model: Wavenet classifier (GroupKFold)<br>\nSequence: Eeg sequence as wavs <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.91 -&gt; 0.81</p>\n<h2>LB: 0.66 -&gt; 0.52</h2>\n<h1>Spectrogram sequenced based</h1>\n<p>Model: Catboost (GroupKFold)<br>\nSequence: Spectrogram sequences <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">Notebook - CatBoost Starter - [LB 0.67]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.82<br>\nLB: 0.67</p>\n<hr>\n<p>Model: EfficientNetB2 (GroupKFold)<br>\nSequence: Spectrogram sequences <a href=\"b2-starter-lb-0-57/notebook?scriptVersionId=158976580\" target=\"_blank\">Notebook - EfficientNetB2 Starter - [LB 0.57]\n</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.72<br>\nLB: 0.57 </p>\n<hr>\n<p>Model: ResNet34d (Stratified Group KFold )<br>\nSequence: Spectrogram sequences <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-inference\" target=\"_blank\">Notebook - HMS-HBAC: ResNet34d Baseline [Inference]\n</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a><br>\nCV: 0.715</p>\n<h2>LB: 0.49</h2>\n<h1>EEG Sequenced -&gt; Spectrogram &amp; Kaggle Spectrogram sequence based</h1>\n<p>Model: CatBoost starter (GroupKFold)<br>\nSequence: Both <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60\" target=\"_blank\">CatBoost Starter - [LB 0.60]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV 0.74<br>\nLB: 0.60</p>\n<hr>\n<p>Model: EfficientNet starter (GroupKFold)<br>\nSequence: Both <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficientNetB0 Starter - [LB 0.43]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV 0.59</p>\n<h2>LB: 0.43</h2>",
  "messages": [
    {
      "id": 2601663,
      "postDate": "2024-01-14T15:45:45.627Z",
      "content": "<h1>Ensemble Models:</h1>\n<table>\n<thead>\n<tr>\n<th>Folds</th>\n<th>CV</th>\n<th>LB</th>\n<th>User</th>\n<th>Comment message</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GKF</td>\n<td><strong>0.500</strong></td>\n<td><strong>0.35</strong></td>\n<td><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></td>\n<td>combine all the ideas from my starter notebooks</td>\n<td><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2609612\" target=\"_blank\">comment here</a></td>\n</tr>\n<tr>\n<td>SGKF</td>\n<td>~</td>\n<td><strong>0.37</strong></td>\n<td><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a></td>\n<td>128,256,4 → 256,256,2 → 512,512,2(resize)</td>\n<td><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2608138\" target=\"_blank\">comment here</a></td>\n</tr>\n</tbody>\n</table>\n<h1>EEG Sequenced based</h1>\n<p>Model : 1D sequence multilabel classifier (Stratified Group KFold)<br>\nSequence:  Eegs pair sequences <a href=\"https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features\" target=\"_blank\">Notebook - Eegs Pairing Analysis &amp; Features</a> for dataset<br>\nCV: ..<br>\nLB: 0.79</p>\n<hr>\n<p>Model : MLP multilabel classifier (Group KFold)<br>\nSequence: Eegs pair sequences Notebook - How To Make Spectrogram from EEG for dataset by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 1.037<br>\nLB: 0.77</p>\n<hr>\n<p>Model : 1D-CNN (Group KFold)<br>\nSequence: Eegs pair sequences Notebook - <a href=\"https://www.kaggle.com/code/anmolgarg1998/hms-cnn1d-inference/notebook\" target=\"_blank\">HMS_cnn1d_inference\n</a> by <a href=\"https://www.kaggle.com/anmolgarg1998\" target=\"_blank\">@anmolgarg1998</a><br>\nCV: -<br>\nLB: 0.69</p>\n<hr>\n<p>Model: Wavenet classifier (GroupKFold)<br>\nSequence: Eeg sequence as wavs <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.91 -&gt; 0.81</p>\n<h2>LB: 0.66 -&gt; 0.52</h2>\n<h1>Spectrogram sequenced based</h1>\n<p>Model: Catboost (GroupKFold)<br>\nSequence: Spectrogram sequences <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">Notebook - CatBoost Starter - [LB 0.67]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.82<br>\nLB: 0.67</p>\n<hr>\n<p>Model: EfficientNetB2 (GroupKFold)<br>\nSequence: Spectrogram sequences <a href=\"b2-starter-lb-0-57/notebook?scriptVersionId=158976580\" target=\"_blank\">Notebook - EfficientNetB2 Starter - [LB 0.57]\n</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV: 0.72<br>\nLB: 0.57 </p>\n<hr>\n<p>Model: ResNet34d (Stratified Group KFold )<br>\nSequence: Spectrogram sequences <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-inference\" target=\"_blank\">Notebook - HMS-HBAC: ResNet34d Baseline [Inference]\n</a> by <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a><br>\nCV: 0.715</p>\n<h2>LB: 0.49</h2>\n<h1>EEG Sequenced -&gt; Spectrogram &amp; Kaggle Spectrogram sequence based</h1>\n<p>Model: CatBoost starter (GroupKFold)<br>\nSequence: Both <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60\" target=\"_blank\">CatBoost Starter - [LB 0.60]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV 0.74<br>\nLB: 0.60</p>\n<hr>\n<p>Model: EfficientNet starter (GroupKFold)<br>\nSequence: Both <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficientNetB0 Starter - [LB 0.43]</a> by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nCV 0.59</p>\n<h2>LB: 0.43</h2>",
      "rawMarkdown": "# Ensemble Models:\n| Folds | CV | LB | User | Comment message | Comment |\n| --- | --- | --- | --- | --- |\n| GKF| **0.500**| **0.35** | @cdeotte | combine all the ideas from my starter notebooks | [comment here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2609612)\n| SGKF | ~ | **0.37** | @abebe9849 | 128,256,4 → 256,256,2 → 512,512,2(resize) | [comment here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2608138)\n\n# EEG Sequenced based\nModel : 1D sequence multilabel classifier (Stratified Group KFold)\nSequence:  Eegs pair sequences [Notebook - Eegs Pairing Analysis & Features](https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features) for dataset\nCV: ..\nLB: 0.79\n\n------------\nModel : MLP multilabel classifier (Group KFold)\nSequence: Eegs pair sequences Notebook - How To Make Spectrogram from EEG for dataset by @cdeotte\nCV: 1.037\nLB: 0.77\n\n------------\nModel : 1D-CNN (Group KFold)\nSequence: Eegs pair sequences Notebook - [HMS_cnn1d_inference\n](https://www.kaggle.com/code/anmolgarg1998/hms-cnn1d-inference/notebook) by @anmolgarg1998\nCV: -\nLB: 0.69\n\n------------\nModel: Wavenet classifier (GroupKFold)\nSequence: Eeg sequence as wavs [WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52) by @cdeotte\nCV: 0.91 -> 0.81\nLB: 0.66 -> 0.52\n------------\n\n# Spectrogram sequenced based\nModel: Catboost (GroupKFold)\nSequence: Spectrogram sequences [Notebook - CatBoost Starter - [LB 0.67]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67) by @cdeotte\nCV: 0.82\nLB: 0.67\n\n-----------\nModel: EfficientNetB2 (GroupKFold)\nSequence: Spectrogram sequences [Notebook - EfficientNetB2 Starter - [LB 0.57]\n](b2-starter-lb-0-57/notebook?scriptVersionId=158976580) by @cdeotte\nCV: 0.72\nLB: 0.57 \n\n-----------\nModel: ResNet34d (Stratified Group KFold )\nSequence: Spectrogram sequences [Notebook - HMS-HBAC: ResNet34d Baseline [Inference]\n](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-inference) by @ttahara\nCV: 0.715\nLB: 0.49\n-----------\n\n# EEG Sequenced -> Spectrogram & Kaggle Spectrogram sequence based\n\nModel: CatBoost starter (GroupKFold)\nSequence: Both [CatBoost Starter - [LB 0.60]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60) by @cdeotte\nCV 0.74\nLB: 0.60\n\n-----\nModel: EfficientNet starter (GroupKFold)\nSequence: Both [EfficientNetB0 Starter - [LB 0.43]](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43) by @cdeotte\nCV 0.59\nLB: 0.43\n--------\n\n",
      "votes": 59
    },
    {
      "id": 2609612,
      "postDate": "2024-01-19T15:49:09.557Z",
      "content": "<p><strong>5 GKF CV: 0.50, LB: 0.35</strong>. I combine all the ideas from my starter notebooks and I use all my Kaggle datasets:</p>\n<p><strong>Starter Notebooks</strong>:</p>\n<ul>\n<li>EfficientNetB2 starter <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a></li>\n<li>CatBoost starter <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a></li>\n<li>WaveNet starter <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\" target=\"_blank\">here</a></li>\n<li>MLP starter <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Kaggle Datasets</strong></p>\n<ul>\n<li>Brain spectrograms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a></li>\n<li>EEG spectrograms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a></li>\n<li>EEG waveforms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eegs\" target=\"_blank\">here</a></li>\n<li>Kaggle's competition KL Div metric <a href=\"https://www.kaggle.com/datasets/cdeotte/kaggle-kl-div\" target=\"_blank\">here</a></li>\n</ul>",
      "rawMarkdown": "**5 GKF CV: 0.50, LB: 0.35**. I combine all the ideas from my starter notebooks and I use all my Kaggle datasets:\n\n**Starter Notebooks**:\n* EfficientNetB2 starter [here][1]\n* CatBoost starter [here][2]\n* WaveNet starter [here][3]\n* MLP starter [here][4]\n\n**Kaggle Datasets**\n* Brain spectrograms [here][5]\n* EEG spectrograms [here][6]\n* EEG waveforms [here][7]\n* Kaggle's competition KL Div metric [here][8]\n\n[1]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[2]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[3]: https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\n[4]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[5]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[6]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[7]: https://www.kaggle.com/datasets/cdeotte/brain-eegs\n[8]: https://www.kaggle.com/datasets/cdeotte/kaggle-kl-div",
      "votes": 31,
      "replies": [
        {
          "id": 2609707,
          "postDate": "2024-01-19T16:41:03.130Z",
          "content": "<p>Impressive, thanks for sharing. What is GKF here?</p>",
          "rawMarkdown": "Impressive, thanks for sharing. What is GKF here?",
          "votes": 2,
          "replies": [
            {
              "id": 2609712,
              "postDate": "2024-01-19T16:45:26.727Z",
              "content": "<p>Group KFold (on patient id). And my train data (on which CV score is computed) is 17089 samples of all unique <code>eeg_id</code>.</p>",
              "rawMarkdown": "Group KFold (on patient id). And my train data (on which CV score is computed) is 17089 samples of all unique `eeg_id`.",
              "votes": 5
            },
            {
              "id": 2609752,
              "postDate": "2024-01-19T17:12:35.650Z",
              "content": "<p>Is validation also one sample per eeg_id or all of them?</p>",
              "rawMarkdown": "Is validation also one sample per eeg_id or all of them?",
              "votes": 2
            },
            {
              "id": 2609772,
              "postDate": "2024-01-19T17:25:38.640Z",
              "content": "<p>Validation is exactly like my 4 starter notebooks. The shape of OOF is <code>(17089,6)</code>. Each <code>eeg_id</code> gets validated once.</p>",
              "rawMarkdown": "Validation is exactly like my 4 starter notebooks. The shape of OOF is `(17089,6)`. Each `eeg_id` gets validated once.",
              "votes": 5
            },
            {
              "id": 2609787,
              "postDate": "2024-01-19T17:35:50.243Z",
              "content": "<p>That approach looks risky to me. Don't you think validating on all samples of unseen patients would be more robust?</p>",
              "rawMarkdown": "That approach looks risky to me. Don't you think validating on all samples of unseen patients would be more robust?",
              "votes": 3
            },
            {
              "id": 2609788,
              "postDate": "2024-01-19T17:38:07.717Z",
              "content": "<p>Are you saying evaluate on <code>106,800</code> samples?</p>\n<p>The purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.</p>",
              "rawMarkdown": "Are you saying evaluate on `106,800` samples?\n\nThe purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.",
              "votes": 5
            },
            {
              "id": 2609806,
              "postDate": "2024-01-19T17:50:53.323Z",
              "content": "<p>Yes my OOF is 106,800 samples.</p>\n<blockquote>\n  <p>The purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.</p>\n</blockquote>\n<p>This is perfectly valid but I think there is a chance of getting too many easy/hard samples when taking one sample per eeg_id. If I can't figure out a decent sampling logic, I'll probably go on like this.</p>",
              "rawMarkdown": "Yes my OOF is 106,800 samples.\n\n> The purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.\n\nThis is perfectly valid but I think there is a chance of getting too many easy/hard samples when taking one sample per eeg_id. If I can't figure out a decent sampling logic, I'll probably go on like this.",
              "votes": 7
            },
            {
              "id": 2609846,
              "postDate": "2024-01-19T18:15:29.963Z",
              "content": "<p>I currently validate in the same way as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, but I think <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> raises a good point. </p>\n<p><code>~9%</code> of eeg_ids have &gt;1 unique combination of votes. Another approach could be to evaluate on all <code>106,800</code> samples and then taking the mean KLD for each eeg_id?</p>",
              "rawMarkdown": "I currently validate in the same way as @cdeotte, but I think @gunesevitan raises a good point. \n\n`~9%` of eeg_ids have >1 unique combination of votes. Another approach could be to evaluate on all `106,800` samples and then taking the mean KLD for each eeg_id?",
              "votes": 3
            }
          ]
        },
        {
          "id": 2609936,
          "postDate": "2024-01-19T19:24:09.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is it ensemble model by average or weighted?</p>",
          "rawMarkdown": "@cdeotte is it ensemble model by average or weighted?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2622917,
      "postDate": "2024-01-27T19:52:50.767Z",
      "content": "<p>kaggle spectrograms : CV = 0.50, LB = 0.4<br>\nChris's spectrograms : CV = 0.46, LB = 0.4<br>\nEnsemble CV = 0.43, LB = 0.34<br>\nMy CV is a little bit optimistic because I'm using only eegs that have one unique spectrogram to replicate the test set distribution, the other eegs with more than one spec are used for training but not for validation.</p>",
      "rawMarkdown": "kaggle spectrograms : CV = 0.50, LB = 0.4\nChris's spectrograms : CV = 0.46, LB = 0.4\nEnsemble CV = 0.43, LB = 0.34\nMy CV is a little bit optimistic because I'm using only eegs that have one unique spectrogram to replicate the test set distribution, the other eegs with more than one spec are used for training but not for validation.",
      "votes": 13,
      "replies": [
        {
          "id": 2622952,
          "postDate": "2024-01-27T20:28:47.610Z",
          "content": "<p>Great job Ah! Using both Kaggle spectrograms and my spectrograms is how I achieved my LB=0.34 also.</p>\n<p>For reference, my spectrograms (made from EEG) are in Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. (Kaggle spectrograms are in Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a>). How to download these datasets to your local machine is explained <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>. Examples how to use my spectrograms (and Kaggle spectrograms) including how to infer test data is in starter notebooks <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Great job Ah! Using both Kaggle spectrograms and my spectrograms is how I achieved my LB=0.34 also.\n\nFor reference, my spectrograms (made from EEG) are in Kaggle dataset [here][1]. (Kaggle spectrograms are in Kaggle dataset [here][4]). How to download these datasets to your local machine is explained [here][5]. Examples how to use my spectrograms (and Kaggle spectrograms) including how to infer test data is in starter notebooks [here][2] and [here][3]\n\n[1]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[2]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n[3]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60\n[4]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[5]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760",
          "votes": 8,
          "replies": [
            {
              "id": 2622957,
              "postDate": "2024-01-27T20:37:36.887Z",
              "content": "<p>I'm really grateful for the contributions that you've made to this community, thanks a lot!</p>",
              "rawMarkdown": "I'm really grateful for the contributions that you've made to this community, thanks a lot!",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2608138,
      "postDate": "2024-01-18T15:17:09.240Z",
      "content": "<p>5sgkf CV :598 LB .42 kaggle provided spectrogram only</p>",
      "rawMarkdown": " 5sgkf CV :598 LB .42 kaggle provided spectrogram only\n\n\n",
      "votes": 11,
      "replies": [
        {
          "id": 2609417,
          "postDate": "2024-01-19T13:18:46.883Z",
          "content": "<p>Nice job    </p>",
          "rawMarkdown": "Nice job    ",
          "votes": 2
        },
        {
          "id": 2615643,
          "postDate": "2024-01-23T08:15:01.887Z",
          "content": "<p>5sgkf CV :60 LB .43 chirs's sspectrogram ,128,256,4→256,256,2→512,512,2(resize)</p>\n<p>LB0.43+LB0.42 ensemble →LB0.37<br>\n→→The different inputs create diversity in the information being extracted!</p>",
          "rawMarkdown": "5sgkf CV :60 LB .43 chirs's sspectrogram ,128,256,4→256,256,2→512,512,2(resize)\n\nLB0.43+LB0.42 ensemble →LB0.37\n→→The different inputs create diversity in the information being extracted!",
          "votes": 10,
          "replies": [
            {
              "id": 2615762,
              "postDate": "2024-01-23T09:50:47.883Z",
              "content": "<p>Nice job patriot. My new EEG spectrograms in my Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a> are very powerful. They are the reason that I have LB 0.35 now too 😀</p>\n<p>(for reference to others, there is a discussion about them <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>)</p>",
              "rawMarkdown": "Nice job patriot. My new EEG spectrograms in my Kaggle dataset [here][1] are very powerful. They are the reason that I have LB 0.35 now too 😀\n\n(for reference to others, there is a discussion about them [here][2])\n\n[1]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[2]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760",
              "votes": 8
            },
            {
              "id": 2615932,
              "postDate": "2024-01-23T12:05:21.753Z",
              "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> <br>\nCould I ask you about the details of the statement \"128,256,4→256,256,2\" ?<br>\nJust reshape like this?<br>\n[:, :, 0] → [0:128, :, 0]<br>\n[:, :, 1] → [128:256, :, 0]<br>\n[:, :, 2] → [0:128, :, 1]<br>\n[:, :, 3] → [128:256, :, 1]<br>\n(I don't know the order is true.)<br>\nOr other operations?</p>",
              "rawMarkdown": "@abebe9849 \nCould I ask you about the details of the statement \"128,256,4→256,256,2\" ?\nJust reshape like this?\n[:, :, 0] → [0:128, :, 0]\n[:, :, 1] → [128:256, :, 0]\n[:, :, 2] → [0:128, :, 1]\n[:, :, 3] → [128:256, :, 1]\n(I don't know the order is true.)\nOr other operations?",
              "votes": 2
            },
            {
              "id": 2615970,
              "postDate": "2024-01-23T12:36:51.680Z",
              "content": "<p>img = np.load(f\"EEG_Spectrograms/{eeg_id}.npy\")#128,256,4<br>\nimg = np.stack([np.concatenate([img[:,:,0],img[:,:,2]]),np.concatenate([img[:,:,1],img[:,:,3]])],axis=-1)</p>",
              "rawMarkdown": "img = np.load(f\"EEG_Spectrograms/{eeg_id}.npy\")#128,256,4\nimg = np.stack([np.concatenate([img[:,:,0],img[:,:,2]]),np.concatenate([img[:,:,1],img[:,:,3]])],axis=-1)",
              "votes": 6
            },
            {
              "id": 2616303,
              "postDate": "2024-01-23T15:11:02.407Z",
              "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> I appreciate your kindness!</p>",
              "rawMarkdown": "@abebe9849 I appreciate your kindness!",
              "votes": 1
            },
            {
              "id": 2618709,
              "postDate": "2024-01-24T23:38:29.120Z",
              "content": "<p>input:chirs's spectrogram+kaggle provided  CV0.54 LB 0.38 !!</p>",
              "rawMarkdown": "input:chirs's spectrogram+kaggle provided  CV0.54 LB 0.38 !!",
              "votes": 2
            }
          ]
        },
        {
          "id": 2615810,
          "postDate": "2024-01-23T10:35:01.253Z",
          "content": "<p>I agree to this. ..models trained with different spectrograms are doing great in ensemble in public LB and CV . </p>",
          "rawMarkdown": "I agree to this. ..models trained with different spectrograms are doing great in ensemble in public LB and CV . ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2625346,
      "postDate": "2024-01-29T11:02:23.100Z",
      "content": "<p>I use only raw eeg. SGKF . CV - 0.659 / LB - 0.45</p>",
      "rawMarkdown": "I use only raw eeg. SGKF . CV - 0.659 / LB - 0.45",
      "votes": 8
    },
    {
      "id": 2619300,
      "postDate": "2024-01-25T11:52:31.477Z",
      "content": "<table>\n  <tbody>\n    <tr>\n      <td>5-fold<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n      <td>CV<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n      <td>LB<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.605</td>\n      <td>0.42</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.609</td>\n      <td>0.41</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.628</td>\n      <td>0.40</td>\n    </tr>\n  </tbody>\n</table>\n<p>In my experiments, I don't feel the correlation between CV and LB…</p>",
      "rawMarkdown": "<table>\n  <tbody>\n    <tr>\n      <td align=\"center\">5-fold<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n      <td align=\"center\">CV<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n      <td align=\"center\">LB<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.605</td>\n      <td>0.42</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.609</td>\n      <td>0.41</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.628</td>\n      <td>0.40</td>\n    </tr>\n  </tbody>\n</table>\n\nIn my experiments, I don't feel the correlation between CV and LB...",
      "votes": 7
    },
    {
      "id": 2619147,
      "postDate": "2024-01-25T08:56:35.473Z",
      "content": "<p>cv 0.587136049535561 lb 0.57</p>",
      "rawMarkdown": "cv 0.587136049535561 lb 0.57",
      "votes": 7
    },
    {
      "id": 2607180,
      "postDate": "2024-01-18T05:19:09.247Z",
      "content": "<p>Actually I am seeing quite a bit of correlation issue between CV and LB , when the model architecture changes or for ensemble .. the local CV might be good but LB score is bad . <br>\ne.g Local CV .639 single model had only .45 LB<br>\nlocal CV .648 single model has .43 LB<br>\nLocal ensemble CV .580 has .44 LB</p>\n<p>Anyone else same experience ? Only me ?</p>",
      "rawMarkdown": "Actually I am seeing quite a bit of correlation issue between CV and LB , when the model architecture changes or for ensemble .. the local CV might be good but LB score is bad . \ne.g Local CV .639 single model had only .45 LB\nlocal CV .648 single model has .43 LB\nLocal ensemble CV .580 has .44 LB\n\nAnyone else same experience ? Only me ?\n",
      "votes": 7,
      "replies": [
        {
          "id": 2607190,
          "postDate": "2024-01-18T05:25:08.090Z",
          "content": "<p>Yes I have the same issue. My validation folds are fixed and I split on the entire dataset. I'm using stratified group kfold, 5 folds and seed 32.</p>\n<p>OOF: 0.7649 (5 fold mean 0.7639 std 0.0351) -&gt; LB 0.46<br>\nOOF: 0.7400 (5 fold mean 0.7390 std 0.0333) -&gt; LB 0.51<br>\nOOF: 0.7459 (5 fold mean 0.7459 std 0.06) -&gt; LB 0.45<br>\nOOF: 0.7351 (5 fold mean 0.7341 std 0.075) -&gt; LB 0.46<br>\nOOF: 0.7235 (5 fold mean 0.7231 std 0.0453) -&gt; LB 0.48<br>\nOOF: 0.7127 (5 fold mean 0.7126 std 0.0409) -&gt; LB 0.48<br>\nOOF: 0.7047 (5 fold mean 0.7040 std 0.0442) -&gt; LB 0.47<br>\nOOF: 0.6853 (5 fold mean 0.6844 std 0.0444) -&gt; LB 0.46<br>\nOOF: 0.6977 (5 fold mean 0.6970 std 0.0422) -&gt; LB 0.47<br>\nOOF: 0.6884 (5 fold mean 0.6882 std 0.0429) -&gt; LB 0.48<br>\nOOF: 0.6664 (5 fold mean 0.6662 std 0.0410) -&gt; LB 0.48</p>\n<p>I almost made 0.1 improvement over my first submission and my LB score is still same.</p>",
          "rawMarkdown": "Yes I have the same issue. My validation folds are fixed and I split on the entire dataset. I'm using stratified group kfold, 5 folds and seed 32.\n\nOOF: 0.7649 (5 fold mean 0.7639 std 0.0351) -> LB 0.46\nOOF: 0.7400 (5 fold mean 0.7390 std 0.0333) -> LB 0.51\nOOF: 0.7459 (5 fold mean 0.7459 std 0.06) -> LB 0.45\nOOF: 0.7351 (5 fold mean 0.7341 std 0.075) -> LB 0.46\nOOF: 0.7235 (5 fold mean 0.7231 std 0.0453) -> LB 0.48\nOOF: 0.7127 (5 fold mean 0.7126 std 0.0409) -> LB 0.48\nOOF: 0.7047 (5 fold mean 0.7040 std 0.0442) -> LB 0.47\nOOF: 0.6853 (5 fold mean 0.6844 std 0.0444) -> LB 0.46\nOOF: 0.6977 (5 fold mean 0.6970 std 0.0422) -> LB 0.47\nOOF: 0.6884 (5 fold mean 0.6882 std 0.0429) -> LB 0.48\nOOF: 0.6664 (5 fold mean 0.6662 std 0.0410) -> LB 0.48\n\nI almost made 0.1 improvement over my first submission and my LB score is still same.\n",
          "votes": 7,
          "replies": [
            {
              "id": 2608134,
              "postDate": "2024-01-18T15:16:23.260Z",
              "content": "<p>Also facing the same issue. <br>\nCV: 0.6371 -&gt; LB: 0.45<br>\nCV: 0.6279 -&gt; LB: 0.46 </p>",
              "rawMarkdown": "Also facing the same issue. \nCV: 0.6371 -> LB: 0.45\nCV: 0.6279 -> LB: 0.46 ",
              "votes": 4
            }
          ]
        },
        {
          "id": 2619309,
          "postDate": "2024-01-25T11:58:14.887Z",
          "content": "<p>same situation</p>",
          "rawMarkdown": "same situation"
        }
      ]
    },
    {
      "id": 2627387,
      "postDate": "2024-01-30T16:56:17.800Z",
      "content": "<p>My CV-lb results using kaggle provided spectrograms:</p>\n<ol>\n<li>CV =0.635 LB = 0.44</li>\n<li>CV = 0.618 LB = 0.42</li>\n<li>CV = 0.600 LB = 0.42 (best)</li>\n</ol>",
      "rawMarkdown": "My CV-lb results using kaggle provided spectrograms:\n\n1. CV =0.635 LB = 0.44\n2. CV = 0.618 LB = 0.42\n3. CV = 0.600 LB = 0.42 (best)",
      "votes": 5,
      "replies": [
        {
          "id": 2627992,
          "postDate": "2024-01-31T03:04:06.237Z",
          "content": "<p>Pytorch or Tf <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> ?</p>",
          "rawMarkdown": "Pytorch or Tf @chaudharypriyanshu ?",
          "votes": 1,
          "replies": [
            {
              "id": 2628375,
              "postDate": "2024-01-31T09:00:46.803Z",
              "content": "<p>I'm using Pytorch.</p>",
              "rawMarkdown": "I'm using Pytorch.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2642816,
          "postDate": "2024-02-08T12:43:26.983Z",
          "content": "<p>Update:<br>\nI have been working on making my raw eeg based models over a week. Luckily i made them work:<br>\nCV = 0.613, LB = 0.44 (There is still some tuning left to make LB more stable).</p>",
          "rawMarkdown": "Update:\nI have been working on making my raw eeg based models over a week. Luckily i made them work:\nCV = 0.613, LB = 0.44 (There is still some tuning left to make LB more stable).",
          "votes": 2
        }
      ]
    },
    {
      "id": 2611029,
      "postDate": "2024-01-20T14:01:36.550Z",
      "content": "<p>My OOF loss:  0.58. LB: 0.43</p>",
      "rawMarkdown": "My OOF loss:  0.58. LB: 0.43",
      "votes": 5
    },
    {
      "id": 2616788,
      "postDate": "2024-01-23T20:24:46.067Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I updated all my starter notebooks to use more features. Here are the new CV LB scores to update your post</p>\n<ul>\n<li>EfficientNet starter : CV 0.73 LB 0.57 =&gt; CV CV 0.59 LB 0.43</li>\n<li>WaveNet starter : CV 0.91 LB 0.66 =&gt; CV 0.81 LB 0.52</li>\n<li>CatBoost starter : CV 0.82 LB 0.67 =&gt; CV 0.74 LB 0.60</li>\n</ul>\n<p>Also note that the \"EEG Sequenced &amp; Spectrogram sequence based\" model listed in your discussion is by me. (You list the wrong author above). Also my new CV LB is <strong>CV 0.50 LB 0.35</strong>. Thanks!</p>",
      "rawMarkdown": "Hi @seshurajup I updated all my starter notebooks to use more features. Here are the new CV LB scores to update your post\n* EfficientNet starter : CV 0.73 LB 0.57 => CV CV 0.59 LB 0.43\n* WaveNet starter : CV 0.91 LB 0.66 => CV 0.81 LB 0.52\n* CatBoost starter : CV 0.82 LB 0.67 => CV 0.74 LB 0.60\n\nAlso note that the \"EEG Sequenced & Spectrogram sequence based\" model listed in your discussion is by me. (You list the wrong author above). Also my new CV LB is **CV 0.50 LB 0.35**. Thanks!",
      "votes": 3,
      "replies": [
        {
          "id": 2616881,
          "postDate": "2024-01-23T22:47:35.473Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks Chris. Updated</p>",
          "rawMarkdown": "@cdeotte Thanks Chris.~~ I will fix it~~ Updated",
          "votes": 1
        }
      ]
    },
    {
      "id": 2605761,
      "postDate": "2024-01-17T08:17:16.957Z",
      "content": "<p>Single model 5sgkf Cv :648  LB .43 kaggle provided spectrogram only </p>",
      "rawMarkdown": "Single model 5sgkf Cv :648  LB .43 kaggle provided spectrogram only ",
      "votes": 3,
      "replies": [
        {
          "id": 2605913,
          "postDate": "2024-01-17T10:13:55.227Z",
          "content": "<p>Great job Doomsday!</p>",
          "rawMarkdown": "Great job Doomsday!",
          "votes": 3,
          "replies": [
            {
              "id": 2610301,
              "postDate": "2024-01-20T04:07:44.450Z",
              "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> Do you know if sgfk is more computationally intensive than gkf? My notebook keeps crashing when running 5 folds of sgfk, but it handles 10 folds of gkf without any issues. I don't understand what's causing this.</p>",
              "rawMarkdown": "@phoenix9032 Do you know if sgfk is more computationally intensive than gkf? My notebook keeps crashing when running 5 folds of sgfk, but it handles 10 folds of gkf without any issues. I don't understand what's causing this.",
              "votes": 1
            },
            {
              "id": 2610701,
              "postDate": "2024-01-20T09:27:11.733Z",
              "content": "<p><a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a>  and this happens just by changing gkf to sgkf in your kernel ? </p>",
              "rawMarkdown": "@yantxx  and this happens just by changing gkf to sgkf in your kernel ? "
            }
          ]
        }
      ]
    },
    {
      "id": 2611696,
      "postDate": "2024-01-20T22:18:45.930Z",
      "content": "<p>I'm pretty surprised how close deep learning (WaveNet) and machine learning (GBT) techniques are scoring in this competition. I assume some sort of ensemble using these two methods could be the way forward.</p>",
      "rawMarkdown": "I'm pretty surprised how close deep learning (WaveNet) and machine learning (GBT) techniques are scoring in this competition. I assume some sort of ensemble using these two methods could be the way forward.",
      "votes": 4
    },
    {
      "id": 2608632,
      "postDate": "2024-01-19T01:18:54.097Z",
      "content": "<p>CV: 0.6465       LB: 0.43 <br>\nCV: 0.7106        LB: 0.45<br>\nCV: 0.6834       LB: 0.48<br>\nNo strong correlation has been found so far.</p>",
      "rawMarkdown": "CV: 0.6465       LB: 0.43 \nCV: 0.7106        LB: 0.45\nCV: 0.6834       LB: 0.48\nNo strong correlation has been found so far.",
      "votes": 4
    },
    {
      "id": 2601716,
      "postDate": "2024-01-14T16:15:16.773Z",
      "content": "<p>Catboost (slightly different from Chris's version) 10 folds CV: 0.75 LB 0.73 I'm somewhat happy with the CV~LB correlation on this one</p>",
      "rawMarkdown": "Catboost (slightly different from Chris's version) 10 folds CV: 0.75 LB 0.73 I'm somewhat happy with the CV~LB correlation on this one",
      "votes": 4,
      "replies": [
        {
          "id": 2603483,
          "postDate": "2024-01-15T19:46:06.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a> is 10 Folds distribution not well balanced right? is it same for you? - using GroupKFold ? Thanks for sharing</p>",
          "rawMarkdown": "@yantxx is 10 Folds distribution not well balanced right? is it same for you? - using GroupKFold ? Thanks for sharing"
        }
      ]
    },
    {
      "id": 2605006,
      "postDate": "2024-01-16T19:31:14.223Z",
      "content": "<h1>EEG Sequenced based</h1>\n<p>Model : MLP multilabel classifier (Group KFold)<br>\nSequence:  Eegs pair sequences <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">Notebook - How To Make Spectrogram from EEG</a> for dataset<br>\nCV: 1.037<br>\nLB: 0.77</p>\n<blockquote>\n  <p><strong>You've been sharing this link too often. We prevent redundant posts to reduce spam</strong>  - it means do i need to write as comment instead update topic?</p>\n</blockquote>",
      "rawMarkdown": "# EEG Sequenced based\nModel : MLP multilabel classifier (Group KFold)\nSequence:  Eegs pair sequences [Notebook - How To Make Spectrogram from EEG](https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg) for dataset\nCV: 1.037\nLB: 0.77\n\n> **You've been sharing this link too often. We prevent redundant posts to reduce spam**  - it means do i need to write as comment instead update topic?",
      "votes": 1
    },
    {
      "id": 2604502,
      "postDate": "2024-01-16T13:06:21.387Z",
      "content": "<p>I've observed a somewhat significant variance in my Group 5Fold CV:<br>\n<strong>CV: 0.68 LB: 0.49</strong></p>",
      "rawMarkdown": "I've observed a somewhat significant variance in my Group 5Fold CV:\n**CV: 0.68 LB: 0.49**\n",
      "votes": 1,
      "replies": [
        {
          "id": 2604508,
          "postDate": "2024-01-16T13:10:18.593Z",
          "content": "<p><a href=\"https://www.kaggle.com/nightsh4de\" target=\"_blank\">@nightsh4de</a> so GroupKFold got better split.  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> explanation about CV vs LB analysis</p>\n<blockquote>\n  <p>same validation loss.  Gap between train loss and valid loss is smaller and the model is better<br>\n  same validation loss.  Gap between train loss and valid loss is larger and the model is overfitting train</p>\n</blockquote>\n<p>Detailed explanation go through comments <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg/comments#2604160\" target=\"_blank\">Notebook Comments - How To Make Spectrograms from EEG</a></p>",
          "rawMarkdown": "@nightsh4de so GroupKFold got better split.  @cdeotte explanation about CV vs LB analysis\n\n> same validation loss.  Gap between train loss and valid loss is smaller and the model is better\n> same validation loss.  Gap between train loss and valid loss is larger and the model is overfitting train\n\nDetailed explanation go through comments [Notebook Comments - How To Make Spectrograms from EEG](https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg/comments#2604160)",
          "votes": 2,
          "replies": [
            {
              "id": 2604515,
              "postDate": "2024-01-16T13:15:38.410Z",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> yes</p>",
              "rawMarkdown": "@seshurajup yes",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2657781,
      "postDate": "2024-02-18T18:04:04.563Z",
      "content": "<p>I'm still getting started here, but using Kaggle spectrograms only, here are my results for 5-fold cross validation using sgkf on patient_id with a resnet-like architecture. There's still a lot to try (gkf instead of sgkf, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> spectrograms, better augmentations, better preprocessing, better models, etc.), but I thought this comparison was interesting.</p>\n<p>aggregating on eeg_id: Train 0.40 (stdev 0.04), CV 0.79 (stdev 0.03), LB 0.56<br>\naggregating on spectogram_id: Train 0.42 (stdev 0.03), CV 0.72 (stdev 0.02), LB 0.57. </p>\n<p>I'm still trying to determine if aggregating on spectrogram_id is overly optimistic (less data, easier samples, etc), or if aggregating on eeg_id really is overfitting more because some longer spectrograms are over represented.</p>\n<p>Currently, it seems like everyone's doing better on the leaderboard than their cv, which makes me concerned about the private set of the data we'll be evaluated on. I'm paying attention to lower LB scores, but I'm focusing on CV at the moment.</p>",
      "rawMarkdown": "I'm still getting started here, but using Kaggle spectrograms only, here are my results for 5-fold cross validation using sgkf on patient_id with a resnet-like architecture. There's still a lot to try (gkf instead of sgkf, @cdeotte spectrograms, better augmentations, better preprocessing, better models, etc.), but I thought this comparison was interesting.\n\naggregating on eeg_id: Train 0.40 (stdev 0.04), CV 0.79 (stdev 0.03), LB 0.56\naggregating on spectogram_id: Train 0.42 (stdev 0.03), CV 0.72 (stdev 0.02), LB 0.57. \n\nI'm still trying to determine if aggregating on spectrogram_id is overly optimistic (less data, easier samples, etc), or if aggregating on eeg_id really is overfitting more because some longer spectrograms are over represented.\n\nCurrently, it seems like everyone's doing better on the leaderboard than their cv, which makes me concerned about the private set of the data we'll be evaluated on. I'm paying attention to lower LB scores, but I'm focusing on CV at the moment.",
      "votes": 2
    },
    {
      "id": 2623090,
      "postDate": "2024-01-28T00:50:42.290Z",
      "content": "<p>Could anyone share train vs validation scores? My model overfits very quickly after the second epoch. After 4 epochs: train score ~0.35 while validation score ~0.7</p>\n<p>I havent submitted yet to LB because I want to have a model that does not overfit</p>",
      "rawMarkdown": "Could anyone share train vs validation scores? My model overfits very quickly after the second epoch. After 4 epochs: train score ~0.35 while validation score ~0.7\n\nI havent submitted yet to LB because I want to have a model that does not overfit",
      "votes": 2,
      "replies": [
        {
          "id": 2623102,
          "postDate": "2024-01-28T01:08:02.330Z",
          "content": "<p>You can check out version 5 of my EfficientNet starter <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43?scriptVersionId=159901408\" target=\"_blank\">here</a>. Epoch 4 has train loss 0.36 and valid loss 0.61 (this is KL Div loss). If you want to make the train valid gap smaller, then you need to add regularization like data augmentation (in preprocess) or dropout, L1 reg, L2 reg, freeze model parameters, etc (in model architecture). Also the gap is smaller for smaller models like EffNetB0 compared with EffNetB5 </p>",
          "rawMarkdown": "You can check out version 5 of my EfficientNet starter [here][1]. Epoch 4 has train loss 0.36 and valid loss 0.61 (this is KL Div loss). If you want to make the train valid gap smaller, then you need to add regularization like data augmentation (in preprocess) or dropout, L1 reg, L2 reg, freeze model parameters, etc (in model architecture). Also the gap is smaller for smaller models like EffNetB0 compared with EffNetB5 \n\n[1]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43?scriptVersionId=159901408",
          "votes": 3,
          "replies": [
            {
              "id": 2623198,
              "postDate": "2024-01-28T04:05:43.440Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I have actually translated super closely your keras efficientnet notebook into pytorch and I am seeing higher scores :( . I will probably make it public once I can lower the CV a bit</p>",
              "rawMarkdown": "Thanks @cdeotte I have actually translated super closely your keras efficientnet notebook into pytorch and I am seeing higher scores :( . I will probably make it public once I can lower the CV a bit",
              "votes": 2
            },
            {
              "id": 2626820,
              "postDate": "2024-01-30T09:39:18.417Z",
              "content": "<p>I also encountered this problem at first, but I couldn't find the problem. Later I rewrote the code and this phenomenon did not happen again. I still don't know why the code I wrote at the beginning had the same situation as yours. If you find Yes, please tell me the reason😀</p>",
              "rawMarkdown": "I also encountered this problem at first, but I couldn't find the problem. Later I rewrote the code and this phenomenon did not happen again. I still don't know why the code I wrote at the beginning had the same situation as yours. If you find Yes, please tell me the reason😀",
              "votes": 1
            },
            {
              "id": 2627914,
              "postDate": "2024-01-31T00:49:30.390Z",
              "content": "<p>For me i think it was the LR scheduler <a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> </p>",
              "rawMarkdown": "For me i think it was the LR scheduler @gentlezdh "
            }
          ]
        }
      ]
    },
    {
      "id": 2678703,
      "postDate": "2024-03-03T01:59:33.217Z",
      "content": "<p>CV 0.6513 LB 0.66 very closely but very far from Leaderboard … </p>",
      "rawMarkdown": "CV 0.6513 LB 0.66 very closely but very far from Leaderboard ... "
    },
    {
      "id": 2622072,
      "postDate": "2024-01-27T08:26:40.907Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    },
    {
      "id": 2620692,
      "postDate": "2024-01-26T09:49:03.230Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2609612,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-01-19T15:49:09.557000",
      "content": "<p><strong>5 GKF CV: 0.50, LB: 0.35</strong>. I combine all the ideas from my starter notebooks and I use all my Kaggle datasets:</p>\n<p><strong>Starter Notebooks</strong>:</p>\n<ul>\n<li>EfficientNetB2 starter <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">here</a></li>\n<li>CatBoost starter <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\" target=\"_blank\">here</a></li>\n<li>WaveNet starter <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\" target=\"_blank\">here</a></li>\n<li>MLP starter <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Kaggle Datasets</strong></p>\n<ul>\n<li>Brain spectrograms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a></li>\n<li>EEG spectrograms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a></li>\n<li>EEG waveforms <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eegs\" target=\"_blank\">here</a></li>\n<li>Kaggle's competition KL Div metric <a href=\"https://www.kaggle.com/datasets/cdeotte/kaggle-kl-div\" target=\"_blank\">here</a></li>\n</ul>",
      "votes": 31,
      "replies": [
        {
          "id": 2609707,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-19T16:41:03.130000",
          "content": "<p>Impressive, thanks for sharing. What is GKF here?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2609712,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-19T16:45:26.727000",
              "content": "<p>Group KFold (on patient id). And my train data (on which CV score is computed) is 17089 samples of all unique <code>eeg_id</code>.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2609752,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-01-19T17:12:35.650000",
              "content": "<p>Is validation also one sample per eeg_id or all of them?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2609772,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-19T17:25:38.640000",
              "content": "<p>Validation is exactly like my 4 starter notebooks. The shape of OOF is <code>(17089,6)</code>. Each <code>eeg_id</code> gets validated once.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2609787,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-01-19T17:35:50.243000",
              "content": "<p>That approach looks risky to me. Don't you think validating on all samples of unseen patients would be more robust?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2609788,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-19T17:38:07.717000",
              "content": "<p>Are you saying evaluate on <code>106,800</code> samples?</p>\n<p>The purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2609806,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-01-19T17:50:53.323000",
              "content": "<p>Yes my OOF is 106,800 samples.</p>\n<blockquote>\n  <p>The purpose of validation is to evaluate the same way that test data will be evaluated. The test data does not have multiple samples per eeg_id.</p>\n</blockquote>\n<p>This is perfectly valid but I think there is a chance of getting too many easy/hard samples when taking one sample per eeg_id. If I can't figure out a decent sampling logic, I'll probably go on like this.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2609846,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-01-19T18:15:29.963000",
              "content": "<p>I currently validate in the same way as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, but I think <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> raises a good point. </p>\n<p><code>~9%</code> of eeg_ids have &gt;1 unique combination of votes. Another approach could be to evaluate on all <code>106,800</code> samples and then taking the mean KLD for each eeg_id?</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2609936,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-19T19:24:09.100000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is it ensemble model by average or weighted?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2622917,
      "author_name": "Ahmed El Fazouani",
      "author_url": "",
      "post_date": "2024-01-27T19:52:50.767000",
      "content": "<p>kaggle spectrograms : CV = 0.50, LB = 0.4<br>\nChris's spectrograms : CV = 0.46, LB = 0.4<br>\nEnsemble CV = 0.43, LB = 0.34<br>\nMy CV is a little bit optimistic because I'm using only eegs that have one unique spectrogram to replicate the test set distribution, the other eegs with more than one spec are used for training but not for validation.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2622952,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-27T20:28:47.610000",
          "content": "<p>Great job Ah! Using both Kaggle spectrograms and my spectrograms is how I achieved my LB=0.34 also.</p>\n<p>For reference, my spectrograms (made from EEG) are in Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a>. (Kaggle spectrograms are in Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\" target=\"_blank\">here</a>). How to download these datasets to your local machine is explained <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>. Examples how to use my spectrograms (and Kaggle spectrograms) including how to infer test data is in starter notebooks <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60\" target=\"_blank\">here</a></p>",
          "votes": 8,
          "replies": [
            {
              "id": 2622957,
              "author_name": "Ahmed El Fazouani",
              "author_url": "",
              "post_date": "2024-01-27T20:37:36.887000",
              "content": "<p>I'm really grateful for the contributions that you've made to this community, thanks a lot!</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2608138,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2024-01-18T15:17:09.240000",
      "content": "<p>5sgkf CV :598 LB .42 kaggle provided spectrogram only</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2609417,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-19T13:18:46.883000",
          "content": "<p>Nice job    </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2615643,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2024-01-23T08:15:01.887000",
          "content": "<p>5sgkf CV :60 LB .43 chirs's sspectrogram ,128,256,4→256,256,2→512,512,2(resize)</p>\n<p>LB0.43+LB0.42 ensemble →LB0.37<br>\n→→The different inputs create diversity in the information being extracted!</p>",
          "votes": 10,
          "replies": [
            {
              "id": 2615762,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-01-23T09:50:47.883000",
              "content": "<p>Nice job patriot. My new EEG spectrograms in my Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">here</a> are very powerful. They are the reason that I have LB 0.35 now too 😀</p>\n<p>(for reference to others, there is a discussion about them <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>)</p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2615932,
              "author_name": "sy",
              "author_url": "",
              "post_date": "2024-01-23T12:05:21.753000",
              "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> <br>\nCould I ask you about the details of the statement \"128,256,4→256,256,2\" ?<br>\nJust reshape like this?<br>\n[:, :, 0] → [0:128, :, 0]<br>\n[:, :, 1] → [128:256, :, 0]<br>\n[:, :, 2] → [0:128, :, 1]<br>\n[:, :, 3] → [128:256, :, 1]<br>\n(I don't know the order is true.)<br>\nOr other operations?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2615970,
              "author_name": "patriot",
              "author_url": "",
              "post_date": "2024-01-23T12:36:51.680000",
              "content": "<p>img = np.load(f\"EEG_Spectrograms/{eeg_id}.npy\")#128,256,4<br>\nimg = np.stack([np.concatenate([img[:,:,0],img[:,:,2]]),np.concatenate([img[:,:,1],img[:,:,3]])],axis=-1)</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2616303,
              "author_name": "sy",
              "author_url": "",
              "post_date": "2024-01-23T15:11:02.407000",
              "content": "<p><a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> I appreciate your kindness!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2618709,
              "author_name": "patriot",
              "author_url": "",
              "post_date": "2024-01-24T23:38:29.120000",
              "content": "<p>input:chirs's spectrogram+kaggle provided  CV0.54 LB 0.38 !!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2615810,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2024-01-23T10:35:01.253000",
          "content": "<p>I agree to this. ..models trained with different spectrograms are doing great in ensemble in public LB and CV . </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2625346,
      "author_name": "Ayushman Buragohain",
      "author_url": "",
      "post_date": "2024-01-29T11:02:23.100000",
      "content": "<p>I use only raw eeg. SGKF . CV - 0.659 / LB - 0.45</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2619300,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-01-25T11:52:31.477000",
      "content": "<table>\n  <tbody>\n    <tr>\n      <td>5-fold<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n      <td>CV<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n      <td>LB<br>\n        \n        \n        \n        \n        \n        \n        \n        \n      </td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.605</td>\n      <td>0.42</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.609</td>\n      <td>0.41</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.628</td>\n      <td>0.40</td>\n    </tr>\n  </tbody>\n</table>\n<p>In my experiments, I don't feel the correlation between CV and LB…</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2619147,
      "author_name": "HZM",
      "author_url": "",
      "post_date": "2024-01-25T08:56:35.473000",
      "content": "<p>cv 0.587136049535561 lb 0.57</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2607180,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-01-18T05:19:09.247000",
      "content": "<p>Actually I am seeing quite a bit of correlation issue between CV and LB , when the model architecture changes or for ensemble .. the local CV might be good but LB score is bad . <br>\ne.g Local CV .639 single model had only .45 LB<br>\nlocal CV .648 single model has .43 LB<br>\nLocal ensemble CV .580 has .44 LB</p>\n<p>Anyone else same experience ? Only me ?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2607190,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-01-18T05:25:08.090000",
          "content": "<p>Yes I have the same issue. My validation folds are fixed and I split on the entire dataset. I'm using stratified group kfold, 5 folds and seed 32.</p>\n<p>OOF: 0.7649 (5 fold mean 0.7639 std 0.0351) -&gt; LB 0.46<br>\nOOF: 0.7400 (5 fold mean 0.7390 std 0.0333) -&gt; LB 0.51<br>\nOOF: 0.7459 (5 fold mean 0.7459 std 0.06) -&gt; LB 0.45<br>\nOOF: 0.7351 (5 fold mean 0.7341 std 0.075) -&gt; LB 0.46<br>\nOOF: 0.7235 (5 fold mean 0.7231 std 0.0453) -&gt; LB 0.48<br>\nOOF: 0.7127 (5 fold mean 0.7126 std 0.0409) -&gt; LB 0.48<br>\nOOF: 0.7047 (5 fold mean 0.7040 std 0.0442) -&gt; LB 0.47<br>\nOOF: 0.6853 (5 fold mean 0.6844 std 0.0444) -&gt; LB 0.46<br>\nOOF: 0.6977 (5 fold mean 0.6970 std 0.0422) -&gt; LB 0.47<br>\nOOF: 0.6884 (5 fold mean 0.6882 std 0.0429) -&gt; LB 0.48<br>\nOOF: 0.6664 (5 fold mean 0.6662 std 0.0410) -&gt; LB 0.48</p>\n<p>I almost made 0.1 improvement over my first submission and my LB score is still same.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2608134,
              "author_name": "nymfree",
              "author_url": "",
              "post_date": "2024-01-18T15:16:23.260000",
              "content": "<p>Also facing the same issue. <br>\nCV: 0.6371 -&gt; LB: 0.45<br>\nCV: 0.6279 -&gt; LB: 0.46 </p>",
              "votes": 4,
              "replies": []
            }
          ]
        },
        {
          "id": 2619309,
          "author_name": "Donghui Zhang",
          "author_url": "",
          "post_date": "2024-01-25T11:58:14.887000",
          "content": "<p>same situation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2627387,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2024-01-30T16:56:17.800000",
      "content": "<p>My CV-lb results using kaggle provided spectrograms:</p>\n<ol>\n<li>CV =0.635 LB = 0.44</li>\n<li>CV = 0.618 LB = 0.42</li>\n<li>CV = 0.600 LB = 0.42 (best)</li>\n</ol>",
      "votes": 5,
      "replies": [
        {
          "id": 2627992,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2024-01-31T03:04:06.237000",
          "content": "<p>Pytorch or Tf <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2628375,
              "author_name": "Priyanshu Chaudhary",
              "author_url": "",
              "post_date": "2024-01-31T09:00:46.803000",
              "content": "<p>I'm using Pytorch.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2642816,
          "author_name": "Priyanshu Chaudhary",
          "author_url": "",
          "post_date": "2024-02-08T12:43:26.983000",
          "content": "<p>Update:<br>\nI have been working on making my raw eeg based models over a week. Luckily i made them work:<br>\nCV = 0.613, LB = 0.44 (There is still some tuning left to make LB more stable).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2611029,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2024-01-20T14:01:36.550000",
      "content": "<p>My OOF loss:  0.58. LB: 0.43</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2616788,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-01-23T20:24:46.067000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> I updated all my starter notebooks to use more features. Here are the new CV LB scores to update your post</p>\n<ul>\n<li>EfficientNet starter : CV 0.73 LB 0.57 =&gt; CV CV 0.59 LB 0.43</li>\n<li>WaveNet starter : CV 0.91 LB 0.66 =&gt; CV 0.81 LB 0.52</li>\n<li>CatBoost starter : CV 0.82 LB 0.67 =&gt; CV 0.74 LB 0.60</li>\n</ul>\n<p>Also note that the \"EEG Sequenced &amp; Spectrogram sequence based\" model listed in your discussion is by me. (You list the wrong author above). Also my new CV LB is <strong>CV 0.50 LB 0.35</strong>. Thanks!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2616881,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-23T22:47:35.473000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks Chris. Updated</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2605761,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-01-17T08:17:16.957000",
      "content": "<p>Single model 5sgkf Cv :648  LB .43 kaggle provided spectrogram only </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2605913,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-17T10:13:55.227000",
          "content": "<p>Great job Doomsday!</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2610301,
              "author_name": "Yan Teixeira",
              "author_url": "",
              "post_date": "2024-01-20T04:07:44.450000",
              "content": "<p><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a> Do you know if sgfk is more computationally intensive than gkf? My notebook keeps crashing when running 5 folds of sgfk, but it handles 10 folds of gkf without any issues. I don't understand what's causing this.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2610701,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2024-01-20T09:27:11.733000",
              "content": "<p><a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a>  and this happens just by changing gkf to sgkf in your kernel ? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2611696,
      "author_name": "Jacob Sharples",
      "author_url": "",
      "post_date": "2024-01-20T22:18:45.930000",
      "content": "<p>I'm pretty surprised how close deep learning (WaveNet) and machine learning (GBT) techniques are scoring in this competition. I assume some sort of ensemble using these two methods could be the way forward.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2608632,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-19T01:18:54.097000",
      "content": "<p>CV: 0.6465       LB: 0.43 <br>\nCV: 0.7106        LB: 0.45<br>\nCV: 0.6834       LB: 0.48<br>\nNo strong correlation has been found so far.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2601716,
      "author_name": "Yan Teixeira",
      "author_url": "",
      "post_date": "2024-01-14T16:15:16.773000",
      "content": "<p>Catboost (slightly different from Chris's version) 10 folds CV: 0.75 LB 0.73 I'm somewhat happy with the CV~LB correlation on this one</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2603483,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-15T19:46:06.717000",
          "content": "<p><a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a> is 10 Folds distribution not well balanced right? is it same for you? - using GroupKFold ? Thanks for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2605006,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-01-16T19:31:14.223000",
      "content": "<h1>EEG Sequenced based</h1>\n<p>Model : MLP multilabel classifier (Group KFold)<br>\nSequence:  Eegs pair sequences <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">Notebook - How To Make Spectrogram from EEG</a> for dataset<br>\nCV: 1.037<br>\nLB: 0.77</p>\n<blockquote>\n  <p><strong>You've been sharing this link too often. We prevent redundant posts to reduce spam</strong>  - it means do i need to write as comment instead update topic?</p>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2604502,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2024-01-16T13:06:21.387000",
      "content": "<p>I've observed a somewhat significant variance in my Group 5Fold CV:<br>\n<strong>CV: 0.68 LB: 0.49</strong></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2604508,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-16T13:10:18.593000",
          "content": "<p><a href=\"https://www.kaggle.com/nightsh4de\" target=\"_blank\">@nightsh4de</a> so GroupKFold got better split.  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> explanation about CV vs LB analysis</p>\n<blockquote>\n  <p>same validation loss.  Gap between train loss and valid loss is smaller and the model is better<br>\n  same validation loss.  Gap between train loss and valid loss is larger and the model is overfitting train</p>\n</blockquote>\n<p>Detailed explanation go through comments <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrograms-from-eeg/comments#2604160\" target=\"_blank\">Notebook Comments - How To Make Spectrograms from EEG</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2604515,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2024-01-16T13:15:38.410000",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> yes</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2657781,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-02-18T18:04:04.563000",
      "content": "<p>I'm still getting started here, but using Kaggle spectrograms only, here are my results for 5-fold cross validation using sgkf on patient_id with a resnet-like architecture. There's still a lot to try (gkf instead of sgkf, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> spectrograms, better augmentations, better preprocessing, better models, etc.), but I thought this comparison was interesting.</p>\n<p>aggregating on eeg_id: Train 0.40 (stdev 0.04), CV 0.79 (stdev 0.03), LB 0.56<br>\naggregating on spectogram_id: Train 0.42 (stdev 0.03), CV 0.72 (stdev 0.02), LB 0.57. </p>\n<p>I'm still trying to determine if aggregating on spectrogram_id is overly optimistic (less data, easier samples, etc), or if aggregating on eeg_id really is overfitting more because some longer spectrograms are over represented.</p>\n<p>Currently, it seems like everyone's doing better on the leaderboard than their cv, which makes me concerned about the private set of the data we'll be evaluated on. I'm paying attention to lower LB scores, but I'm focusing on CV at the moment.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2623090,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2024-01-28T00:50:42.290000",
      "content": "<p>Could anyone share train vs validation scores? My model overfits very quickly after the second epoch. After 4 epochs: train score ~0.35 while validation score ~0.7</p>\n<p>I havent submitted yet to LB because I want to have a model that does not overfit</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2623102,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-01-28T01:08:02.330000",
          "content": "<p>You can check out version 5 of my EfficientNet starter <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43?scriptVersionId=159901408\" target=\"_blank\">here</a>. Epoch 4 has train loss 0.36 and valid loss 0.61 (this is KL Div loss). If you want to make the train valid gap smaller, then you need to add regularization like data augmentation (in preprocess) or dropout, L1 reg, L2 reg, freeze model parameters, etc (in model architecture). Also the gap is smaller for smaller models like EffNetB0 compared with EffNetB5 </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2623198,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2024-01-28T04:05:43.440000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I have actually translated super closely your keras efficientnet notebook into pytorch and I am seeing higher scores :( . I will probably make it public once I can lower the CV a bit</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2626820,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-01-30T09:39:18.417000",
              "content": "<p>I also encountered this problem at first, but I couldn't find the problem. Later I rewrote the code and this phenomenon did not happen again. I still don't know why the code I wrote at the beginning had the same situation as yours. If you find Yes, please tell me the reason😀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2627914,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2024-01-31T00:49:30.390000",
              "content": "<p>For me i think it was the LR scheduler <a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2678703,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2024-03-03T01:59:33.217000",
      "content": "<p>CV 0.6513 LB 0.66 very closely but very far from Leaderboard … </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2622072,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-27T08:26:40.907000",
      "content": "",
      "votes": -4,
      "replies": []
    },
    {
      "id": 2620692,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-26T09:49:03.230000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2601663": "# Ensemble Models:\n| Folds | CV | LB | User | Comment message | Comment |\n| --- | --- | --- | --- | --- |\n| GKF| **0.500**| **0.35** | @cdeotte | combine all the ideas from my starter notebooks | [comment here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2609612)\n| SGKF | ~ | **0.37** | @abebe9849 | 128,256,4 → 256,256,2 → 512,512,2(resize) | [comment here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467915#2608138)\n\n# EEG Sequenced based\nModel : 1D sequence multilabel classifier (Stratified Group KFold)\nSequence:  Eegs pair sequences [Notebook - Eegs Pairing Analysis & Features](https://www.kaggle.com/code/seshurajup/eegs-pairing-analysis-features) for dataset\nCV: ..\nLB: 0.79\n\n------------\nModel : MLP multilabel classifier (Group KFold)\nSequence: Eegs pair sequences Notebook - How To Make Spectrogram from EEG for dataset by @cdeotte\nCV: 1.037\nLB: 0.77\n\n------------\nModel : 1D-CNN (Group KFold)\nSequence: Eegs pair sequences Notebook - [HMS_cnn1d_inference\n](https://www.kaggle.com/code/anmolgarg1998/hms-cnn1d-inference/notebook) by @anmolgarg1998\nCV: -\nLB: 0.69\n\n------------\nModel: Wavenet classifier (GroupKFold)\nSequence: Eeg sequence as wavs [WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52) by @cdeotte\nCV: 0.91 -> 0.81\nLB: 0.66 -> 0.52\n------------\n\n# Spectrogram sequenced based\nModel: Catboost (GroupKFold)\nSequence: Spectrogram sequences [Notebook - CatBoost Starter - [LB 0.67]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67) by @cdeotte\nCV: 0.82\nLB: 0.67\n\n-----------\nModel: EfficientNetB2 (GroupKFold)\nSequence: Spectrogram sequences [Notebook - EfficientNetB2 Starter - [LB 0.57]\n](b2-starter-lb-0-57/notebook?scriptVersionId=158976580) by @cdeotte\nCV: 0.72\nLB: 0.57 \n\n-----------\nModel: ResNet34d (Stratified Group KFold )\nSequence: Spectrogram sequences [Notebook - HMS-HBAC: ResNet34d Baseline [Inference]\n](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-inference) by @ttahara\nCV: 0.715\nLB: 0.49\n-----------\n\n# EEG Sequenced -> Spectrogram & Kaggle Spectrogram sequence based\n\nModel: CatBoost starter (GroupKFold)\nSequence: Both [CatBoost Starter - [LB 0.60]](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60) by @cdeotte\nCV 0.74\nLB: 0.60\n\n-----\nModel: EfficientNet starter (GroupKFold)\nSequence: Both [EfficientNetB0 Starter - [LB 0.43]](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43) by @cdeotte\nCV 0.59\nLB: 0.43\n--------\n\n",
    "2609612": "**5 GKF CV: 0.50, LB: 0.35**. I combine all the ideas from my starter notebooks and I use all my Kaggle datasets:\n\n**Starter Notebooks**:\n* EfficientNetB2 starter [here][1]\n* CatBoost starter [here][2]\n* WaveNet starter [here][3]\n* MLP starter [here][4]\n\n**Kaggle Datasets**\n* Brain spectrograms [here][5]\n* EEG spectrograms [here][6]\n* EEG waveforms [here][7]\n* Kaggle's competition KL Div metric [here][8]\n\n[1]: https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\n[2]: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-67\n[3]: https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-66\n[4]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[5]: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n[6]: https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\n[7]: https://www.kaggle.com/datasets/cdeotte/brain-eegs\n[8]: https://www.kaggle.com/datasets/cdeotte/kaggle-kl-div",
    "2622917": "kaggle spectrograms : CV = 0.50, LB = 0.4\nChris's spectrograms : CV = 0.46, LB = 0.4\nEnsemble CV = 0.43, LB = 0.34\nMy CV is a little bit optimistic because I'm using only eegs that have one unique spectrogram to replicate the test set distribution, the other eegs with more than one spec are used for training but not for validation.",
    "2608138": " 5sgkf CV :598 LB .42 kaggle provided spectrogram only\n\n\n",
    "2625346": "I use only raw eeg. SGKF . CV - 0.659 / LB - 0.45",
    "2619300": "<table>\n  <tbody>\n    <tr>\n      <td align=\"center\">5-fold<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n      <td align=\"center\">CV<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n      <td align=\"center\">LB<br>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</span>\n        <span>&nbsp;&nbsp;</span>\n      </td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.605</td>\n      <td>0.42</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.609</td>\n      <td>0.41</td>\n    </tr>\n    <tr>\n      <td>StratifiedGroupKFold</td>\n      <td>0.628</td>\n      <td>0.40</td>\n    </tr>\n  </tbody>\n</table>\n\nIn my experiments, I don't feel the correlation between CV and LB...",
    "2619147": "cv 0.587136049535561 lb 0.57",
    "2607180": "Actually I am seeing quite a bit of correlation issue between CV and LB , when the model architecture changes or for ensemble .. the local CV might be good but LB score is bad . \ne.g Local CV .639 single model had only .45 LB\nlocal CV .648 single model has .43 LB\nLocal ensemble CV .580 has .44 LB\n\nAnyone else same experience ? Only me ?\n",
    "2627387": "My CV-lb results using kaggle provided spectrograms:\n\n1. CV =0.635 LB = 0.44\n2. CV = 0.618 LB = 0.42\n3. CV = 0.600 LB = 0.42 (best)",
    "2611029": "My OOF loss:  0.58. LB: 0.43",
    "2616788": "Hi @seshurajup I updated all my starter notebooks to use more features. Here are the new CV LB scores to update your post\n* EfficientNet starter : CV 0.73 LB 0.57 => CV CV 0.59 LB 0.43\n* WaveNet starter : CV 0.91 LB 0.66 => CV 0.81 LB 0.52\n* CatBoost starter : CV 0.82 LB 0.67 => CV 0.74 LB 0.60\n\nAlso note that the \"EEG Sequenced & Spectrogram sequence based\" model listed in your discussion is by me. (You list the wrong author above). Also my new CV LB is **CV 0.50 LB 0.35**. Thanks!",
    "2605761": "Single model 5sgkf Cv :648  LB .43 kaggle provided spectrogram only ",
    "2611696": "I'm pretty surprised how close deep learning (WaveNet) and machine learning (GBT) techniques are scoring in this competition. I assume some sort of ensemble using these two methods could be the way forward.",
    "2608632": "CV: 0.6465       LB: 0.43 \nCV: 0.7106        LB: 0.45\nCV: 0.6834       LB: 0.48\nNo strong correlation has been found so far.",
    "2601716": "Catboost (slightly different from Chris's version) 10 folds CV: 0.75 LB 0.73 I'm somewhat happy with the CV~LB correlation on this one",
    "2605006": "# EEG Sequenced based\nModel : MLP multilabel classifier (Group KFold)\nSequence:  Eegs pair sequences [Notebook - How To Make Spectrogram from EEG](https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg) for dataset\nCV: 1.037\nLB: 0.77\n\n> **You've been sharing this link too often. We prevent redundant posts to reduce spam**  - it means do i need to write as comment instead update topic?",
    "2604502": "I've observed a somewhat significant variance in my Group 5Fold CV:\n**CV: 0.68 LB: 0.49**\n",
    "2657781": "I'm still getting started here, but using Kaggle spectrograms only, here are my results for 5-fold cross validation using sgkf on patient_id with a resnet-like architecture. There's still a lot to try (gkf instead of sgkf, @cdeotte spectrograms, better augmentations, better preprocessing, better models, etc.), but I thought this comparison was interesting.\n\naggregating on eeg_id: Train 0.40 (stdev 0.04), CV 0.79 (stdev 0.03), LB 0.56\naggregating on spectogram_id: Train 0.42 (stdev 0.03), CV 0.72 (stdev 0.02), LB 0.57. \n\nI'm still trying to determine if aggregating on spectrogram_id is overly optimistic (less data, easier samples, etc), or if aggregating on eeg_id really is overfitting more because some longer spectrograms are over represented.\n\nCurrently, it seems like everyone's doing better on the leaderboard than their cv, which makes me concerned about the private set of the data we'll be evaluated on. I'm paying attention to lower LB scores, but I'm focusing on CV at the moment.",
    "2623090": "Could anyone share train vs validation scores? My model overfits very quickly after the second epoch. After 4 epochs: train score ~0.35 while validation score ~0.7\n\nI havent submitted yet to LB because I want to have a model that does not overfit",
    "2678703": "CV 0.6513 LB 0.66 very closely but very far from Leaderboard ... ",
    "2622072": "",
    "2620692": ""
  }
}