{
  "id": 477135,
  "title": "EffNetB0 model trained twice, once for each of two training populations - [LB 0.39]",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477135",
  "author_name": "Sean R.B. Bearden, Ph.D.",
  "post_date": "2024-02-14T20:00:12.402000",
  "votes": 120,
  "comment_count": 17,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">EffNetB0 2 Pop Model Train Twice - [LB 0.39]</a></p>\n<p>In the exploration of EEG data, as highlighted by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in his insightful <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>, an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.</p>\n<p>This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.</p>\n<p>In light of these reflections, I adapted my approach in the <a href=\"https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\" target=\"_blank\">EfficientNetB0 Noisy Student notebook</a> to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).</p>\n<p>This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.</p>\n<p>Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of <code>train.csv</code> to 20,183 rows by eliminating duplicates based on a combination of [<code>eeg_id</code>, <code>seizure_vote</code>, <code>lpd_vote</code>, <code>gpd_vote</code>, <code>lrda_vote</code>,<code>grda_vote</code>, <code>other_vote</code>], thus incorporating an additional 3,094 samples compared to the dataset used in <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">original notebook</a>, which was limited to unique <code>eeg_ids</code>.</p>\n<p>Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">cdeotte's work</a>, opting instead to generate distinct spectrograms for each label_id based on <code>eeg_label_offset_seconds</code>. This modification yielded a richer training dataset with nuanced variations.</p>\n<p>The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.</p>\n<p>I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?</p>\n<p>Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!</p>\n<hr>\n<p>Massive thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the inspiration drawn from his <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">notebook</a> and to <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the foundational insights provided in his <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>. If you find the EEG spectrograms useful, please consider upvoting my <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">dataset</a> and associated <a href=\"https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\" target=\"_blank\">notebook</a>.</p>\n<p><em>Note: This EfficientNet model begins with <a href=\"https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\" target=\"_blank\">noisy student weights</a>.</em></p>",
  "messages": [
    {
      "id": 2652549,
      "postDate": "2024-02-14T20:00:12.403Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">EffNetB0 2 Pop Model Train Twice - [LB 0.39]</a></p>\n<p>In the exploration of EEG data, as highlighted by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in his insightful <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>, an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.</p>\n<p>This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.</p>\n<p>In light of these reflections, I adapted my approach in the <a href=\"https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\" target=\"_blank\">EfficientNetB0 Noisy Student notebook</a> to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).</p>\n<p>This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.</p>\n<p>Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of <code>train.csv</code> to 20,183 rows by eliminating duplicates based on a combination of [<code>eeg_id</code>, <code>seizure_vote</code>, <code>lpd_vote</code>, <code>gpd_vote</code>, <code>lrda_vote</code>,<code>grda_vote</code>, <code>other_vote</code>], thus incorporating an additional 3,094 samples compared to the dataset used in <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">original notebook</a>, which was limited to unique <code>eeg_ids</code>.</p>\n<p>Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">cdeotte's work</a>, opting instead to generate distinct spectrograms for each label_id based on <code>eeg_label_offset_seconds</code>. This modification yielded a richer training dataset with nuanced variations.</p>\n<p>The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.</p>\n<p>I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?</p>\n<p>Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!</p>\n<hr>\n<p>Massive thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the inspiration drawn from his <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">notebook</a> and to <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the foundational insights provided in his <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>. If you find the EEG spectrograms useful, please consider upvoting my <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">dataset</a> and associated <a href=\"https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\" target=\"_blank\">notebook</a>.</p>\n<p><em>Note: This EfficientNet model begins with <a href=\"https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\" target=\"_blank\">noisy student weights</a>.</em></p>",
      "rawMarkdown": "[EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]\n\nIn the exploration of EEG data, as highlighted by @pcjimmmy in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.\n\nThis observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.\n\nIn light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).\n\nThis experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.\n\nGiven the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of `train.csv` to 20,183 rows by eliminating duplicates based on a combination of [`eeg_id`, `seizure_vote`, `lpd_vote`, `gpd_vote`, `lrda_vote`,`grda_vote`, `other_vote`], thus incorporating an additional 3,094 samples compared to the dataset used in @cdeotte's [original notebook][1], which was limited to unique `eeg_ids`.\n\nFurther refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on `eeg_label_offset_seconds`. This modification yielded a richer training dataset with nuanced variations.\n\nThe outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.\n\nI am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?\n\nYour feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!\n\n---\nMassive thanks to @cdeotte for the inspiration drawn from his [notebook][1] and to @pcjimmmy for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].\n\n*Note: This EfficientNet model begins with [noisy student weights][5].*\n\n[0]: https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\n[1]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n[2]: https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\n[3]: https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\n[4]: https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\n[5]: https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\n[6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39",
      "votes": 120
    },
    {
      "id": 2653865,
      "postDate": "2024-02-15T17:16:49.993Z",
      "content": "<p>I've read what <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, it was on my list(4th) to try something similiar, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">link here</a><br>\nI would train on a single stage, but change the way I would calculate the target probabilities. The more votes there is, the more I should be confident of that probability.</p>\n<blockquote>\n  <p><strong>Weighted voting</strong> Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.<br>\n  Examples: (votes -&gt; weighted probability)<br>\n  1- [2, 2, 0, 0, 0, 0] -&gt; [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]<br>\n  2- [10, 10, 0, 0, 0, 0] -&gt; [0.5, 0.5, 0, 0, 0, 0]<br>\n  Some math formula is needed.</p>\n</blockquote>\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "I've read what @pcjimmmy, it was on my list(4th) to try something similiar, [link here] (https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666)\nI would train on a single stage, but change the way I would calculate the target probabilities. The more votes there is, the more I should be confident of that probability.\n\n> **Weighted voting** Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.\nExamples: (votes -> weighted probability)\n1- [2, 2, 0, 0, 0, 0] -> [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]\n2- [10, 10, 0, 0, 0, 0] -> [0.5, 0.5, 0, 0, 0, 0]\nSome math formula is needed.\n\nThanks for sharing!",
      "votes": 5,
      "replies": [
        {
          "id": 2653880,
          "postDate": "2024-02-15T17:25:11.287Z",
          "content": "<p>I've tried something similar and found a slight improvement to my CV and LB for the model. Still tinkering with it, but I like the idea!</p>",
          "rawMarkdown": "I've tried something similar and found a slight improvement to my CV and LB for the model. Still tinkering with it, but I like the idea!",
          "votes": 1,
          "replies": [
            {
              "id": 2654387,
              "postDate": "2024-02-16T05:34:13.343Z",
              "content": "<p>How are you getting the votes for test data ? Or just during training you are planning to use vote as a feature to guide the loss ..</p>",
              "rawMarkdown": "How are you getting the votes for test data ? Or just during training you are planning to use vote as a feature to guide the loss ..",
              "votes": 1
            },
            {
              "id": 2654408,
              "postDate": "2024-02-16T05:58:25.630Z",
              "content": "<p>Votes aren’t a feature during training. I’ve only used the votes to break apart the training data. Votes are not inputs into the model, so you don’t need votes in the test data. </p>",
              "rawMarkdown": "Votes aren’t a feature during training. I’ve only used the votes to break apart the training data. Votes are not inputs into the model, so you don’t need votes in the test data. ",
              "votes": 1
            },
            {
              "id": 2655674,
              "postDate": "2024-02-17T06:02:49.027Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2656012,
          "postDate": "2024-02-17T11:45:44.627Z",
          "content": "<p>I tried giving greater loss weight to samples with more votes, I chose samples over 10, but the effect did not increase.</p>",
          "rawMarkdown": "I tried giving greater loss weight to samples with more votes, I chose samples over 10, but the effect did not increase.",
          "votes": 3
        },
        {
          "id": 2661079,
          "postDate": "2024-02-21T02:34:49.980Z",
          "content": "<p>I had the same idea.</p>\n<p>My thinking was: Consider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.</p>\n<p>I used the formula:</p>\n<p>alpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes<br>\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6</p>\n<p>where smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.</p>",
          "rawMarkdown": "I had the same idea.\n\nMy thinking was: Consider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.\n\nI used the formula:\n\nalpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6\n\nwhere smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.",
          "votes": 1,
          "replies": [
            {
              "id": 2661088,
              "postDate": "2024-02-21T02:47:54.420Z",
              "content": "<p>I should also note that with the smoothing, train losses did significantly decrease. I suspect it's because it's a lot easier to predict flat distributions.</p>",
              "rawMarkdown": "I should also note that with the smoothing, train losses did significantly decrease. I suspect it's because it's a lot easier to predict flat distributions.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2652905,
      "postDate": "2024-02-15T04:38:06.313Z",
      "content": "<p>Great idea to take advantage of the number of votes, very insightful, thank you for sharing! </p>",
      "rawMarkdown": "Great idea to take advantage of the number of votes, very insightful, thank you for sharing! ",
      "votes": 1
    },
    {
      "id": 2655870,
      "postDate": "2024-02-17T08:48:47.337Z",
      "content": "<p>Funny, I was just noticing the same thing while looking at Chris Deotte's EfficientNetB2 Starter.  In the very small sample I looked at (like only 5 or 6) there was only ever unanimous consensus where there were 1 or 2 votes.</p>\n<p>My thinking, was it might make sense to weight the training data based on the number of votes each EEG/Spectrogram receives by adding that many copies into the set.  So in that scenario an EEG with 5 votes would affect the training 5 times more than than an EEG with 1 vote.  No idea whether this is an appropriate approach, it may overdo it, but seems like it might achieve some level of regularization against the data receiving very few votes.</p>\n<p>If this overdoes it maybe a log(N) approach or something similar.</p>",
      "rawMarkdown": "Funny, I was just noticing the same thing while looking at Chris Deotte's EfficientNetB2 Starter.  In the very small sample I looked at (like only 5 or 6) there was only ever unanimous consensus where there were 1 or 2 votes.\n\nMy thinking, was it might make sense to weight the training data based on the number of votes each EEG/Spectrogram receives by adding that many copies into the set.  So in that scenario an EEG with 5 votes would affect the training 5 times more than than an EEG with 1 vote.  No idea whether this is an appropriate approach, it may overdo it, but seems like it might achieve some level of regularization against the data receiving very few votes.\n\nIf this overdoes it maybe a log(N) approach or something similar.",
      "votes": 2,
      "replies": [
        {
          "id": 2656040,
          "postDate": "2024-02-17T11:59:13.320Z",
          "content": "<p>Hello, <br>\nI added regularization to the targets labels, which can be controlled by a coefficient.<br>\nYou can find out about this approach in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477498\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Hello, \nI added regularization to the targets labels, which can be controlled by a coefficient.\nYou can find out about this approach in [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477498)",
          "votes": 2
        },
        {
          "id": 2664203,
          "postDate": "2024-02-22T21:19:04.810Z",
          "content": "<p><a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> Great idea! Have you had success with weighting the training data?</p>",
          "rawMarkdown": "@davidlist Great idea! Have you had success with weighting the training data?"
        }
      ]
    },
    {
      "id": 2669194,
      "postDate": "2024-02-26T06:18:49.913Z",
      "content": "<p>thank you for yours</p>\n<blockquote>\n  <p>[EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]</p>\n  <p>In the exploration of EEG data, as highlighted by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.</p>\n  <p>This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.</p>\n  <p>In light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).</p>\n  <p>This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.</p>\n  <p>Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of <code>train.csv</code> to 20,183 rows by eliminating duplicates based on a combination of [<code>eeg_id</code>, <code>seizure_vote</code>, <code>lpd_vote</code>, <code>gpd_vote</code>, <code>lrda_vote</code>,<code>grda_vote</code>, <code>other_vote</code>], thus incorporating an additional 3,094 samples compared to the dataset used in <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s [original notebook][1], which was limited to unique <code>eeg_ids</code>.</p>\n  <p>Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on <code>eeg_label_offset_seconds</code>. This modification yielded a richer training dataset with nuanced variations.</p>\n  <p>The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.</p>\n  <p>I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?</p>\n  <p>Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!</p>\n  <hr>\n  <p>Massive thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the inspiration drawn from his [notebook][1] and to <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].</p>\n  <p><em>Note: This EfficientNet model begins with [noisy student weights][5].</em></p>\n  <p>[0]: <a href=\"https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42</a><br>\n  [1]: <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43</a><br>\n  [2]: <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda</a><br>\n  [3]: <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique</a><br>\n  [4]: <a href=\"https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo</a><br>\n  [5]: <a href=\"https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights</a><br>\n  [6]: <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg</a><br>\n  [7]: <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39</a></p>\n</blockquote>",
      "rawMarkdown": "thank you for yours\n\n\n\n> [EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]\n> \n> In the exploration of EEG data, as highlighted by @pcjimmmy in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.\n> \n> This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.\n> \n> In light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).\n> \n> This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.\n> \n> Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of `train.csv` to 20,183 rows by eliminating duplicates based on a combination of [`eeg_id`, `seizure_vote`, `lpd_vote`, `gpd_vote`, `lrda_vote`,`grda_vote`, `other_vote`], thus incorporating an additional 3,094 samples compared to the dataset used in @cdeotte's [original notebook][1], which was limited to unique `eeg_ids`.\n> \n> Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on `eeg_label_offset_seconds`. This modification yielded a richer training dataset with nuanced variations.\n> \n> The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.\n> \n> I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?\n> \n> Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!\n> \n> ---\n> Massive thanks to @cdeotte for the inspiration drawn from his [notebook][1] and to @pcjimmmy for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].\n> \n> *Note: This EfficientNet model begins with [noisy student weights][5].*\n> \n> [0]: https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\n> [1]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n> [2]: https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\n> [3]: https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\n> [4]: https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\n> [5]: https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\n> [6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n> [7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\n\n"
    },
    {
      "id": 2657080,
      "postDate": "2024-02-18T09:27:53.603Z",
      "content": "<p>Thanks for sharing. <br>\nI want to ask the test.csv without total_evaluators, how about we train a model to  predict each egg/spec less 10 or not as new feautre use on test.csv?</p>",
      "rawMarkdown": "Thanks for sharing. \nI want to ask the test.csv without total_evaluators, how about we train a model to  predict each egg/spec less 10 or not as new feautre use on test.csv?",
      "replies": [
        {
          "id": 2664208,
          "postDate": "2024-02-22T21:28:53.910Z",
          "content": "<p><a href=\"https://www.kaggle.com/kerrysun\" target=\"_blank\">@kerrysun</a> thank you for the suggestion, but not sure I am following. Are you suggesting we create a model to predict how many evaluators were used when voting on the EEGs? My intuition is that the EEG/spec data does not contain info indicating how many evaluators were utilized. Presumably, the number of evaluators is not correlated to the signal itself. </p>\n<p>However, I suppose it is possible that less evaluators were used when the diagnosis was obvious. Seizures tend to have less votes than the other categories. A possible interpretation is that it is easy for experts to agree when a seizure is occurring, so not many experts were consulted. However, the context of the competition seems to invalidate that idea.</p>",
          "rawMarkdown": "@kerrysun thank you for the suggestion, but not sure I am following. Are you suggesting we create a model to predict how many evaluators were used when voting on the EEGs? My intuition is that the EEG/spec data does not contain info indicating how many evaluators were utilized. Presumably, the number of evaluators is not correlated to the signal itself. \n\nHowever, I suppose it is possible that less evaluators were used when the diagnosis was obvious. Seizures tend to have less votes than the other categories. A possible interpretation is that it is easy for experts to agree when a seizure is occurring, so not many experts were consulted. However, the context of the competition seems to invalidate that idea.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2655637,
      "postDate": "2024-02-17T05:32:30.797Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2658772,
      "postDate": "2024-02-19T11:48:50.700Z",
      "content": "<p>Amazing approach thanks for sharing</p>",
      "rawMarkdown": "Amazing approach thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2653865,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-02-15T17:16:49.993000",
      "content": "<p>I've read what <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, it was on my list(4th) to try something similiar, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">link here</a><br>\nI would train on a single stage, but change the way I would calculate the target probabilities. The more votes there is, the more I should be confident of that probability.</p>\n<blockquote>\n  <p><strong>Weighted voting</strong> Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.<br>\n  Examples: (votes -&gt; weighted probability)<br>\n  1- [2, 2, 0, 0, 0, 0] -&gt; [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]<br>\n  2- [10, 10, 0, 0, 0, 0] -&gt; [0.5, 0.5, 0, 0, 0, 0]<br>\n  Some math formula is needed.</p>\n</blockquote>\n<p>Thanks for sharing!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2653880,
          "author_name": "Sean R.B. Bearden, Ph.D.",
          "author_url": "",
          "post_date": "2024-02-15T17:25:11.287000",
          "content": "<p>I've tried something similar and found a slight improvement to my CV and LB for the model. Still tinkering with it, but I like the idea!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2654387,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2024-02-16T05:34:13.343000",
              "content": "<p>How are you getting the votes for test data ? Or just during training you are planning to use vote as a feature to guide the loss ..</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2654408,
              "author_name": "Sean R.B. Bearden, Ph.D.",
              "author_url": "",
              "post_date": "2024-02-16T05:58:25.630000",
              "content": "<p>Votes aren’t a feature during training. I’ve only used the votes to break apart the training data. Votes are not inputs into the model, so you don’t need votes in the test data. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2655674,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-17T06:02:49.027000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2656012,
          "author_name": "MiHu",
          "author_url": "",
          "post_date": "2024-02-17T11:45:44.627000",
          "content": "<p>I tried giving greater loss weight to samples with more votes, I chose samples over 10, but the effect did not increase.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2661079,
          "author_name": "Cael Hasse",
          "author_url": "",
          "post_date": "2024-02-21T02:34:49.980000",
          "content": "<p>I had the same idea.</p>\n<p>My thinking was: Consider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.</p>\n<p>I used the formula:</p>\n<p>alpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes<br>\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6</p>\n<p>where smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2661088,
              "author_name": "Cael Hasse",
              "author_url": "",
              "post_date": "2024-02-21T02:47:54.420000",
              "content": "<p>I should also note that with the smoothing, train losses did significantly decrease. I suspect it's because it's a lot easier to predict flat distributions.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2652905,
      "author_name": "Rob D",
      "author_url": "",
      "post_date": "2024-02-15T04:38:06.313000",
      "content": "<p>Great idea to take advantage of the number of votes, very insightful, thank you for sharing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2655870,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-02-17T08:48:47.337000",
      "content": "<p>Funny, I was just noticing the same thing while looking at Chris Deotte's EfficientNetB2 Starter.  In the very small sample I looked at (like only 5 or 6) there was only ever unanimous consensus where there were 1 or 2 votes.</p>\n<p>My thinking, was it might make sense to weight the training data based on the number of votes each EEG/Spectrogram receives by adding that many copies into the set.  So in that scenario an EEG with 5 votes would affect the training 5 times more than than an EEG with 1 vote.  No idea whether this is an appropriate approach, it may overdo it, but seems like it might achieve some level of regularization against the data receiving very few votes.</p>\n<p>If this overdoes it maybe a log(N) approach or something similar.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2656040,
          "author_name": "Danial Zakaria",
          "author_url": "",
          "post_date": "2024-02-17T11:59:13.320000",
          "content": "<p>Hello, <br>\nI added regularization to the targets labels, which can be controlled by a coefficient.<br>\nYou can find out about this approach in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477498\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2664203,
          "author_name": "Sean R.B. Bearden, Ph.D.",
          "author_url": "",
          "post_date": "2024-02-22T21:19:04.810000",
          "content": "<p><a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> Great idea! Have you had success with weighting the training data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2669194,
      "author_name": "laplace_zfg",
      "author_url": "",
      "post_date": "2024-02-26T06:18:49.913000",
      "content": "<p>thank you for yours</p>\n<blockquote>\n  <p>[EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]</p>\n  <p>In the exploration of EEG data, as highlighted by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.</p>\n  <p>This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.</p>\n  <p>In light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).</p>\n  <p>This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.</p>\n  <p>Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of <code>train.csv</code> to 20,183 rows by eliminating duplicates based on a combination of [<code>eeg_id</code>, <code>seizure_vote</code>, <code>lpd_vote</code>, <code>gpd_vote</code>, <code>lrda_vote</code>,<code>grda_vote</code>, <code>other_vote</code>], thus incorporating an additional 3,094 samples compared to the dataset used in <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s [original notebook][1], which was limited to unique <code>eeg_ids</code>.</p>\n  <p>Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on <code>eeg_label_offset_seconds</code>. This modification yielded a richer training dataset with nuanced variations.</p>\n  <p>The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.</p>\n  <p>I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?</p>\n  <p>Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!</p>\n  <hr>\n  <p>Massive thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the inspiration drawn from his [notebook][1] and to <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].</p>\n  <p><em>Note: This EfficientNet model begins with [noisy student weights][5].</em></p>\n  <p>[0]: <a href=\"https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42</a><br>\n  [1]: <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43</a><br>\n  [2]: <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda</a><br>\n  [3]: <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique</a><br>\n  [4]: <a href=\"https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo</a><br>\n  [5]: <a href=\"https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights</a><br>\n  [6]: <a href=\"https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg</a><br>\n  [7]: <a href=\"https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\" target=\"_blank\">https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39</a></p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2657080,
      "author_name": "kerry sun",
      "author_url": "",
      "post_date": "2024-02-18T09:27:53.603000",
      "content": "<p>Thanks for sharing. <br>\nI want to ask the test.csv without total_evaluators, how about we train a model to  predict each egg/spec less 10 or not as new feautre use on test.csv?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2664208,
          "author_name": "Sean R.B. Bearden, Ph.D.",
          "author_url": "",
          "post_date": "2024-02-22T21:28:53.910000",
          "content": "<p><a href=\"https://www.kaggle.com/kerrysun\" target=\"_blank\">@kerrysun</a> thank you for the suggestion, but not sure I am following. Are you suggesting we create a model to predict how many evaluators were used when voting on the EEGs? My intuition is that the EEG/spec data does not contain info indicating how many evaluators were utilized. Presumably, the number of evaluators is not correlated to the signal itself. </p>\n<p>However, I suppose it is possible that less evaluators were used when the diagnosis was obvious. Seizures tend to have less votes than the other categories. A possible interpretation is that it is easy for experts to agree when a seizure is occurring, so not many experts were consulted. However, the context of the competition seems to invalidate that idea.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2655637,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-17T05:32:30.797000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2658772,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2024-02-19T11:48:50.700000",
      "content": "<p>Amazing approach thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2652549": "[EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]\n\nIn the exploration of EEG data, as highlighted by @pcjimmmy in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.\n\nThis observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.\n\nIn light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).\n\nThis experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.\n\nGiven the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of `train.csv` to 20,183 rows by eliminating duplicates based on a combination of [`eeg_id`, `seizure_vote`, `lpd_vote`, `gpd_vote`, `lrda_vote`,`grda_vote`, `other_vote`], thus incorporating an additional 3,094 samples compared to the dataset used in @cdeotte's [original notebook][1], which was limited to unique `eeg_ids`.\n\nFurther refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on `eeg_label_offset_seconds`. This modification yielded a richer training dataset with nuanced variations.\n\nThe outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.\n\nI am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?\n\nYour feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!\n\n---\nMassive thanks to @cdeotte for the inspiration drawn from his [notebook][1] and to @pcjimmmy for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].\n\n*Note: This EfficientNet model begins with [noisy student weights][5].*\n\n[0]: https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\n[1]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n[2]: https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\n[3]: https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\n[4]: https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\n[5]: https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\n[6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n[7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39",
    "2653865": "I've read what @pcjimmmy, it was on my list(4th) to try something similiar, [link here] (https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666)\nI would train on a single stage, but change the way I would calculate the target probabilities. The more votes there is, the more I should be confident of that probability.\n\n> **Weighted voting** Maybe we could have a weighted voting of some sort, so the more votes there is, the more confident we are in those votes and should give them more weight.\nExamples: (votes -> weighted probability)\n1- [2, 2, 0, 0, 0, 0] -> [0.3, 0.3, 0.1, 0.1, 0.1, 0.1]\n2- [10, 10, 0, 0, 0, 0] -> [0.5, 0.5, 0, 0, 0, 0]\nSome math formula is needed.\n\nThanks for sharing!",
    "2652905": "Great idea to take advantage of the number of votes, very insightful, thank you for sharing! ",
    "2655870": "Funny, I was just noticing the same thing while looking at Chris Deotte's EfficientNetB2 Starter.  In the very small sample I looked at (like only 5 or 6) there was only ever unanimous consensus where there were 1 or 2 votes.\n\nMy thinking, was it might make sense to weight the training data based on the number of votes each EEG/Spectrogram receives by adding that many copies into the set.  So in that scenario an EEG with 5 votes would affect the training 5 times more than than an EEG with 1 vote.  No idea whether this is an appropriate approach, it may overdo it, but seems like it might achieve some level of regularization against the data receiving very few votes.\n\nIf this overdoes it maybe a log(N) approach or something similar.",
    "2669194": "thank you for yours\n\n\n\n> [EffNetB0 2 Pop Model Train Twice - [LB 0.39]][7]\n> \n> In the exploration of EEG data, as highlighted by @pcjimmmy in his insightful [notebook][2], an intriguing pattern emerges concerning the total votes for each label_id. This pattern suggests that the training data may have been aggregated from two distinct sources, characterized by their vote counts: one group ranges from 1-7 votes, while the other spans from 10-28 votes. This division not only hints at the data's dual origin but also underscores a unique class imbalance within each group.\n> \n> This observation led me to ponder the implications of training models on data subsets with minimal votes, especially considering the scenario of a dataset containing entries with a singular vote. Such conditions could predispose models to predict classes with exaggerated confidence, potentially skewing the remaining probabilities towards zero. This scenario is particularly problematic considering the impact on KL-divergence, which escalates significantly when predicted probabilities approach zero for outcomes with a higher actual probability.\n> \n> In light of these reflections, I adapted my approach in the [EfficientNetB0 Noisy Student notebook][0] to explore the effects of artificially adjusting the least probable class prediction, which unfortunately led to a degradation in performance (LB score deteriorating from 0.42 to 0.49).\n> \n> This experience sparked the idea of a two-stage training process, initially focusing on the dataset comprising 1-7 total votes, followed by the dataset with 10-28 votes. The rationale was to harness the learning potential from the peaked distributions in the first stage, subsequently refining the model in the second stage to mitigate the KL-divergence issue.\n> \n> Given the split in data, I sought ways to maximize the training dataset for each stage, leading to the condensation of `train.csv` to 20,183 rows by eliminating duplicates based on a combination of [`eeg_id`, `seizure_vote`, `lpd_vote`, `gpd_vote`, `lrda_vote`,`grda_vote`, `other_vote`], thus incorporating an additional 3,094 samples compared to the dataset used in @cdeotte's [original notebook][1], which was limited to unique `eeg_ids`.\n> \n> Further refining our approach, I diverged from the method of generating EEG spectrograms centered around a 50-second sample, as done in [cdeotte's work][6], opting instead to generate distinct spectrograms for each label_id based on `eeg_label_offset_seconds`. This modification yielded a richer training dataset with nuanced variations.\n> \n> The outcome of this nuanced, two-stage training strategy was promising, achieving a CV score of 0.68 and an LB score of 0.39.\n> \n> I am eager to hear your thoughts on this methodology. How do you approach the challenge of class imbalance and skewed distributions in your models? Have you experimented with similar two-stage training strategies, or do you have alternative techniques to share?\n> \n> Your feedback and insights would be invaluable, not only to refine this approach further but also to foster a richer understanding within our community. Let's discuss!\n> \n> ---\n> Massive thanks to @cdeotte for the inspiration drawn from his [notebook][1] and to @pcjimmmy for the foundational insights provided in his [notebook][2]. If you find the EEG spectrograms useful, please consider upvoting my [dataset][3] and associated [notebook][4].\n> \n> *Note: This EfficientNet model begins with [noisy student weights][5].*\n> \n> [0]: https://www.kaggle.com/code/seanbearden/efficientnetb0-noisy-student-lb-0-42\n> [1]: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n> [2]: https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\n> [3]: https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\n> [4]: https://www.kaggle.com/code/seanbearden/spectrogram-from-unique-eeg-votes-combo\n> [5]: https://www.kaggle.com/datasets/seanbearden/tf-efficientnet-noisy-student-weights\n> [6]: https://www.kaggle.com/code/cdeotte/how-to-make-spectrogram-from-eeg\n> [7]: https://www.kaggle.com/code/seanbearden/effnetb0-2-pop-model-train-twice-lb-0-39\n\n",
    "2657080": "Thanks for sharing. \nI want to ask the test.csv without total_evaluators, how about we train a model to  predict each egg/spec less 10 or not as new feautre use on test.csv?",
    "2655637": "",
    "2658772": "Amazing approach thanks for sharing"
  }
}