{
  "id": 477123,
  "title": "CV VS Public LB score ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477123",
  "author_name": "",
  "post_date": "2024-02-14T19:01:26.092165100Z",
  "votes": 9,
  "comment_count": 18,
  "views": 0,
  "content": "<p>During this competition I have tried and improved multiple models that have been shared. For each of these models I have noticed two things. Each time my cv score increased so did my public LB score increased. Next to this I noticed that the cv scores of different models do not hold much correlation to each other E.g. a better cv score in one model does not mean it will have a better leaderboard score than another completely different architecture. Has anyone else experienced this as well?</p>",
  "messages": [
    {
      "id": "2652483",
      "postDate": "02/14/2024 19:01:26",
      "content": "<p>During this competition I have tried and improved multiple models that have been shared. For each of these models I have noticed two things. Each time my cv score increased so did my public LB score increased. Next to this I noticed that the cv scores of different models do not hold much correlation to each other E.g. a better cv score in one model does not mean it will have a better leaderboard score than another completely different architecture. Has anyone else experienced this as well?</p>",
      "rawMarkdown": "During this competition I have tried and improved multiple models that have been shared. For each of these models I have noticed two things. Each time my cv score increased so did my public LB score increased. Next to this I noticed that the cv scores of different models do not hold much correlation to each other E.g. a better cv score in one model does not mean it will have a better leaderboard score than another completely different architecture. Has anyone else experienced this as well?",
      "votes": null
    },
    {
      "id": "2652632",
      "postDate": "02/14/2024 21:45:53",
      "content": "<p>I have noticed that many models using <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">Chris Deotte's EEG Spectrograms</a> have higher CV KL-Divergence than LB Score. One possibility for this is that the training data is not using the EEG offset seconds, so many EEG spectrograms are not centered on the provided training spectrograms. When making a submission, the raw EEG signal is exactly 50 seconds, so the offset is intrinsically applied.</p>",
      "rawMarkdown": "I have noticed that many models using [Chris Deotte's EEG Spectrograms](https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms) have higher CV KL-Divergence than LB Score. One possibility for this is that the training data is not using the EEG offset seconds, so many EEG spectrograms are not centered on the provided training spectrograms. When making a submission, the raw EEG signal is exactly 50 seconds, so the offset is intrinsically applied.",
      "votes": null
    },
    {
      "id": "2652651",
      "postDate": "02/14/2024 22:49:23",
      "content": "<p>I'll have a look into this, currently I have been using predominantly <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> his public notebooks.</p>",
      "rawMarkdown": "I'll have a look into this, currently I have been using predominantly @cdeotte his public notebooks.",
      "votes": null
    },
    {
      "id": "2653247",
      "postDate": "02/15/2024 10:08:43",
      "content": "<p>it's the same for me, I'm always stuck at 0.34 although there is an improvment for my cv.<br>\nI don't know how the top scorers are breaking this barrier, maybe I should try another model like wavenet or 1dcnn.<br>\nI also think we shouldn't rely so much on th LB score to avoid a shakeup, because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)</p>",
      "rawMarkdown": "it's the same for me, I'm always stuck at 0.34 although there is an improvment for my cv.\nI don't know how the top scorers are breaking this barrier, maybe I should try another model like wavenet or 1dcnn.\nI also think we shouldn't rely so much on th LB score to avoid a shakeup, because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)",
      "votes": null
    },
    {
      "id": "2653254",
      "postDate": "02/15/2024 10:19:55",
      "content": "<p>But is this improvement in your cv marginal? The cv improvement for me is little as is the lb improvement. There is just a large discrepance between the cv score of different architecture models and the public LB score retrieved.</p>",
      "rawMarkdown": "But is this improvement in your cv marginal? The cv improvement for me is little as is the lb improvement. There is just a large discrepance between the cv score of different architecture models and the public LB score retrieved.",
      "votes": null
    },
    {
      "id": "2653278",
      "postDate": "02/15/2024 10:34:43",
      "content": "<p>I improved the model that I'm using for eeg spec from 0.4 LB to 0.36 LB,<br>\nbut ensembling it with kaggle spec is always giving the same LB score 0.34</p>",
      "rawMarkdown": "I improved the model that I'm using for eeg spec from 0.4 LB to 0.36 LB,\nbut ensembling it with kaggle spec is always giving the same LB score 0.34",
      "votes": null
    },
    {
      "id": "2653413",
      "postDate": "02/15/2024 12:40:09",
      "content": "<p>Kinda same situation. Our team improved the CV by 0.04 and we havent pass 0.34 barier</p>",
      "rawMarkdown": "Kinda same situation. Our team improved the CV by 0.04 and we havent pass 0.34 barier",
      "votes": null
    },
    {
      "id": "2653424",
      "postDate": "02/15/2024 12:49:55",
      "content": "<p>Interesting haven't tried ensembling by using different data sources. Another week training ahead it is.</p>",
      "rawMarkdown": "Interesting haven't tried ensembling by using different data sources. Another week training ahead it is.",
      "votes": null
    },
    {
      "id": "2653777",
      "postDate": "02/15/2024 16:36:36",
      "content": "<p>whats your best ensemble CV so far? I am seeing same trend, CV gets better but LB doesn't move. </p>",
      "rawMarkdown": "whats your best ensemble CV so far? I am seeing same trend, CV gets better but LB doesn't move.",
      "votes": null
    },
    {
      "id": "2654015",
      "postDate": "02/15/2024 19:18:32",
      "content": "<p>it's around 0.42-0.43 computed on 9K samples (unique eegs), probably will be close to yours if computed on all samples, and if no one of us is overfitting the lb badly 😅😅</p>",
      "rawMarkdown": "it's around 0.42-0.43 computed on 9K samples (unique eegs), probably will be close to yours if computed on all samples, and if no one of us is overfitting the lb badly 😅😅",
      "votes": null
    },
    {
      "id": "2654493",
      "postDate": "02/16/2024 07:05:01",
      "content": "<p>What do you mean by 'computed on 9k samples (unique eegs)'? </p>",
      "rawMarkdown": "What do you mean by 'computed on 9k samples (unique eegs)'?",
      "votes": null
    },
    {
      "id": "2654529",
      "postDate": "02/16/2024 07:34:59",
      "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> I meant spectrograms that have one unique eeg.<br>\n<a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> was the improvment that you've made (0.34-&gt;0.32) reflected on your cv ?</p>",
      "rawMarkdown": "samson8 I meant spectrograms that have one unique eeg.\n@pheadrus was the improvment that you've made (0.34->0.32) reflected on your cv ?",
      "votes": null
    },
    {
      "id": "2654533",
      "postDate": "02/16/2024 07:39:51",
      "content": "<p>Yes, this time it did reflect on my CV. But sometimes CV improves while LB doesn't. Note this time around I had a couple of very diverse models in the mix. </p>",
      "rawMarkdown": "Yes, this time it did reflect on my CV. But sometimes CV improves while LB doesn't. Note this time around I had a couple of very diverse models in the mix.",
      "votes": null
    },
    {
      "id": "2654543",
      "postDate": "02/16/2024 07:57:07",
      "content": "<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> thanks for explanation. That is interesting validation strategy, but I dont understand why its valid. Do we know that test contains only patients with single eeg record?</p>",
      "rawMarkdown": "ahmedelfazouan thanks for explanation. That is interesting validation strategy, but I dont understand why its valid. Do we know that test contains only patients with single eeg record?",
      "votes": null
    },
    {
      "id": "2654577",
      "postDate": "02/16/2024 08:44:06",
      "content": "<p>Yes, It is known</p>",
      "rawMarkdown": "Yes, It is known",
      "votes": null
    },
    {
      "id": "2655804",
      "postDate": "02/17/2024 07:57:59",
      "content": "<p>Interesting…</p>",
      "rawMarkdown": "Interesting...",
      "votes": null
    },
    {
      "id": "2659231",
      "postDate": "02/19/2024 17:20:11",
      "content": "<p>I have been experimenting using different models and parameters on the same validation data, but I couldn't understand why different architectures have different cv scores but similar lb scores. What could be the reason behind this? </p>",
      "rawMarkdown": "I have been experimenting using different models and parameters on the same validation data, but I couldn't understand why different architectures have different cv scores but similar lb scores. What could be the reason behind this?",
      "votes": null
    },
    {
      "id": "2666693",
      "postDate": "02/24/2024 14:56:54",
      "content": "<blockquote>\n  <p>because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)</p>\n</blockquote>\n<p>that means whole test set has 2640 rows/patients<br>\n<a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> could you please explain how you derived this result about Public LB? is it based on a discussion/confirmation from hosts? </p>",
      "rawMarkdown": ">  because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)\n\nthat means whole test set has 2640 rows/patients\n@ahmedelfazouan could you please explain how you derived this result about Public LB? is it based on a discussion/confirmation from hosts?",
      "votes": null
    },
    {
      "id": "2666716",
      "postDate": "02/24/2024 15:20:57",
      "content": "<p>It was discussed <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "It was discussed [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2652632,
      "author_name": "seanbearden",
      "author_url": "",
      "post_date": "02/14/2024 21:45:53",
      "content": "<p>I have noticed that many models using <a href=\"https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms\" target=\"_blank\">Chris Deotte's EEG Spectrograms</a> have higher CV KL-Divergence than LB Score. One possibility for this is that the training data is not using the EEG offset seconds, so many EEG spectrograms are not centered on the provided training spectrograms. When making a submission, the raw EEG signal is exactly 50 seconds, so the offset is intrinsically applied.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2652651,
          "author_name": "stefanoclss",
          "author_url": "",
          "post_date": "02/14/2024 22:49:23",
          "content": "<p>I'll have a look into this, currently I have been using predominantly <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> his public notebooks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2653247,
      "author_name": "ahmedelfazouan",
      "author_url": "",
      "post_date": "02/15/2024 10:08:43",
      "content": "<p>it's the same for me, I'm always stuck at 0.34 although there is an improvment for my cv.<br>\nI don't know how the top scorers are breaking this barrier, maybe I should try another model like wavenet or 1dcnn.<br>\nI also think we shouldn't rely so much on th LB score to avoid a shakeup, because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2653254,
          "author_name": "stefanoclss",
          "author_url": "",
          "post_date": "02/15/2024 10:19:55",
          "content": "<p>But is this improvement in your cv marginal? The cv improvement for me is little as is the lb improvement. There is just a large discrepance between the cv score of different architecture models and the public LB score retrieved.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2653278,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/15/2024 10:34:43",
              "content": "<p>I improved the model that I'm using for eeg spec from 0.4 LB to 0.36 LB,<br>\nbut ensembling it with kaggle spec is always giving the same LB score 0.34</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2653424,
                  "author_name": "stefanoclss",
                  "author_url": "",
                  "post_date": "02/15/2024 12:49:55",
                  "content": "<p>Interesting haven't tried ensembling by using different data sources. Another week training ahead it is.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2653413,
          "author_name": "samson8",
          "author_url": "",
          "post_date": "02/15/2024 12:40:09",
          "content": "<p>Kinda same situation. Our team improved the CV by 0.04 and we havent pass 0.34 barier</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2653777,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "02/15/2024 16:36:36",
          "content": "<p>whats your best ensemble CV so far? I am seeing same trend, CV gets better but LB doesn't move. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2654015,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/15/2024 19:18:32",
              "content": "<p>it's around 0.42-0.43 computed on 9K samples (unique eegs), probably will be close to yours if computed on all samples, and if no one of us is overfitting the lb badly 😅😅</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2654493,
                  "author_name": "samson8",
                  "author_url": "",
                  "post_date": "02/16/2024 07:05:01",
                  "content": "<p>What do you mean by 'computed on 9k samples (unique eegs)'? </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2654529,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/16/2024 07:34:59",
              "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> I meant spectrograms that have one unique eeg.<br>\n<a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> was the improvment that you've made (0.34-&gt;0.32) reflected on your cv ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2654533,
                  "author_name": "pheadrus",
                  "author_url": "",
                  "post_date": "02/16/2024 07:39:51",
                  "content": "<p>Yes, this time it did reflect on my CV. But sometimes CV improves while LB doesn't. Note this time around I had a couple of very diverse models in the mix. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2654543,
                  "author_name": "samson8",
                  "author_url": "",
                  "post_date": "02/16/2024 07:57:07",
                  "content": "<p><a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> thanks for explanation. That is interesting validation strategy, but I dont understand why its valid. Do we know that test contains only patients with single eeg record?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2654577,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/16/2024 08:44:06",
              "content": "<p>Yes, It is known</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2666693,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "02/24/2024 14:56:54",
          "content": "<blockquote>\n  <p>because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)</p>\n</blockquote>\n<p>that means whole test set has 2640 rows/patients<br>\n<a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> could you please explain how you derived this result about Public LB? is it based on a discussion/confirmation from hosts? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2666716,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/24/2024 15:20:57",
              "content": "<p>It was discussed <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\" target=\"_blank\">here</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2655804,
      "author_name": "aidlre001",
      "author_url": "",
      "post_date": "02/17/2024 07:57:59",
      "content": "<p>Interesting…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2659231,
      "author_name": "turkenm",
      "author_url": "",
      "post_date": "02/19/2024 17:20:11",
      "content": "<p>I have been experimenting using different models and parameters on the same validation data, but I couldn't understand why different architectures have different cv scores but similar lb scores. What could be the reason behind this? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2652483": "During this competition I have tried and improved multiple models that have been shared. For each of these models I have noticed two things. Each time my cv score increased so did my public LB score increased. Next to this I noticed that the cv scores of different models do not hold much correlation to each other E.g. a better cv score in one model does not mean it will have a better leaderboard score than another completely different architecture. Has anyone else experienced this as well?",
    "2652632": "I have noticed that many models using [Chris Deotte's EEG Spectrograms](https://www.kaggle.com/datasets/cdeotte/brain-eeg-spectrograms) have higher CV KL-Divergence than LB Score. One possibility for this is that the training data is not using the EEG offset seconds, so many EEG spectrograms are not centered on the provided training spectrograms. When making a submission, the raw EEG signal is exactly 50 seconds, so the offset is intrinsically applied.",
    "2652651": "I'll have a look into this, currently I have been using predominantly @cdeotte his public notebooks.",
    "2653247": "it's the same for me, I'm always stuck at 0.34 although there is an improvment for my cv.\nI don't know how the top scorers are breaking this barrier, maybe I should try another model like wavenet or 1dcnn.\nI also think we shouldn't rely so much on th LB score to avoid a shakeup, because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)",
    "2653254": "But is this improvement in your cv marginal? The cv improvement for me is little as is the lb improvement. There is just a large discrepance between the cv score of different architecture models and the public LB score retrieved.",
    "2653278": "I improved the model that I'm using for eeg spec from 0.4 LB to 0.36 LB,\nbut ensembling it with kaggle spec is always giving the same LB score 0.34",
    "2653413": "Kinda same situation. Our team improved the CV by 0.04 and we havent pass 0.34 barier",
    "2653424": "Interesting haven't tried ensembling by using different data sources. Another week training ahead it is.",
    "2653777": "whats your best ensemble CV so far? I am seeing same trend, CV gets better but LB doesn't move.",
    "2654015": "it's around 0.42-0.43 computed on 9K samples (unique eegs), probably will be close to yours if computed on all samples, and if no one of us is overfitting the lb badly 😅😅",
    "2654493": "What do you mean by 'computed on 9k samples (unique eegs)'?",
    "2654529": "samson8 I meant spectrograms that have one unique eeg.\n@pheadrus was the improvment that you've made (0.34->0.32) reflected on your cv ?",
    "2654533": "Yes, this time it did reflect on my CV. But sometimes CV improves while LB doesn't. Note this time around I had a couple of very diverse models in the mix.",
    "2654543": "ahmedelfazouan thanks for explanation. That is interesting validation strategy, but I dont understand why its valid. Do we know that test contains only patients with single eeg record?",
    "2654577": "Yes, It is known",
    "2655804": "Interesting...",
    "2659231": "I have been experimenting using different models and parameters on the same validation data, but I couldn't understand why different architectures have different cv scores but similar lb scores. What could be the reason behind this?",
    "2666693": ">  because it contains only 924 samples (2640 * 35 / 100), CV is the king in this case (17 K samples)\n\nthat means whole test set has 2640 rows/patients\n@ahmedelfazouan could you please explain how you derived this result about Public LB? is it based on a discussion/confirmation from hosts?",
    "2666716": "It was discussed [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287)"
  },
  "source": "meta"
}