{
  "id": 483949,
  "title": "Before it colapse",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/483949",
  "author_name": "SSS",
  "post_date": "2024-03-14T20:44:51.324000",
  "votes": 32,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>\"You take the blue pill - the story ends, you wake up in your bed and believe whatever you want to believe. You take the red pill - you stay in Wonderland and I show you how deep the rabbit hole goes.\"</p>\n<ol>\n<li>Currently, there are many topics/kernels opened which are pro two-stage training, and recently there is the one which warns us about potential harm.</li>\n<li>Few weeks ago I fitted the model with a binary target - (0, 7] voters and [10, 28] voters, I used AUCROC to see if the model can distinguish between two datasets. Well it did (0.62 score, weak though, but still):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef54e7736fcbcd710b6e2c5bf473864c%2Faucroc.png?generation=1710448838631576&amp;alt=media\"></li>\n</ol>\n<h2>Outro</h2>\n<p>Honestly, I am kinda late here but I just want to post it separately, so it won't be lost in the comments elsewhere. The story is that polishing with second stage means heaviliy overfitting to the public dataset. If we are lucky enough, then yes - that will work, though I'd rather trust my <strong>0.4 LB</strong> with global <strong>CV 0.53</strong> than two-stage <strong>0.3 CV</strong> with <strong>LB 0.3</strong> but heavily overfited by using the second stage.</p>\n<p>You make a decision what pill you take.</p>",
  "messages": [
    {
      "id": 2697353,
      "postDate": "2024-03-14T20:44:51.323Z",
      "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>\"You take the blue pill - the story ends, you wake up in your bed and believe whatever you want to believe. You take the red pill - you stay in Wonderland and I show you how deep the rabbit hole goes.\"</p>\n<ol>\n<li>Currently, there are many topics/kernels opened which are pro two-stage training, and recently there is the one which warns us about potential harm.</li>\n<li>Few weeks ago I fitted the model with a binary target - (0, 7] voters and [10, 28] voters, I used AUCROC to see if the model can distinguish between two datasets. Well it did (0.62 score, weak though, but still):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef54e7736fcbcd710b6e2c5bf473864c%2Faucroc.png?generation=1710448838631576&amp;alt=media\"></li>\n</ol>\n<h2>Outro</h2>\n<p>Honestly, I am kinda late here but I just want to post it separately, so it won't be lost in the comments elsewhere. The story is that polishing with second stage means heaviliy overfitting to the public dataset. If we are lucky enough, then yes - that will work, though I'd rather trust my <strong>0.4 LB</strong> with global <strong>CV 0.53</strong> than two-stage <strong>0.3 CV</strong> with <strong>LB 0.3</strong> but heavily overfited by using the second stage.</p>\n<p>You make a decision what pill you take.</p>",
      "rawMarkdown": "Hi,\n\n## Intro\n\"You take the blue pill - the story ends, you wake up in your bed and believe whatever you want to believe. You take the red pill - you stay in Wonderland and I show you how deep the rabbit hole goes.\"\n\n1. Currently, there are many topics/kernels opened which are pro two-stage training, and recently there is the one which warns us about potential harm.\n2. Few weeks ago I fitted the model with a binary target - (0, 7] voters and [10, 28] voters, I used AUCROC to see if the model can distinguish between two datasets. Well it did (0.62 score, weak though, but still):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef54e7736fcbcd710b6e2c5bf473864c%2Faucroc.png?generation=1710448838631576&alt=media)\n\n## Outro\nHonestly, I am kinda late here but I just want to post it separately, so it won't be lost in the comments elsewhere. The story is that polishing with second stage means heaviliy overfitting to the public dataset. If we are lucky enough, then yes - that will work, though I'd rather trust my **0.4 LB** with global **CV 0.53** than two-stage **0.3 CV** with **LB 0.3** but heavily overfited by using the second stage.\n\nYou make a decision what pill you take.",
      "votes": 32
    },
    {
      "id": 2697672,
      "postDate": "2024-03-15T04:08:06.270Z",
      "content": "<p>For dataset classification model, which features do you use?<br>\nDoes your model still work if you selected data with the same diagnoses (e.g., lpd data with voter &lt; 9 vs.  lpd data with voter &gt; 9)?</p>",
      "rawMarkdown": "For dataset classification model, which features do you use?\nDoes your model still work if you selected data with the same diagnoses (e.g., lpd data with voter < 9 vs.  lpd data with voter > 9)?",
      "votes": 3,
      "replies": [
        {
          "id": 2698401,
          "postDate": "2024-03-15T13:34:33.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> well I would leave that to the reader rather. There are things you can do with classification model, split data into different subsets based on votes, specific classes and potentially get even better discriminating power…</p>",
          "rawMarkdown": "@tomooinubushi well I would leave that to the reader rather. There are things you can do with classification model, split data into different subsets based on votes, specific classes and potentially get even better discriminating power..."
        }
      ]
    },
    {
      "id": 2698436,
      "postDate": "2024-03-15T14:11:00.373Z",
      "content": "<p>It would make zero sense if private lb has low quality labels (ie less than 10 evaluators distribution), why would they want a model that generalizes well to the low quality labels? it's counter intuitive of the goal of the competition imo.</p>",
      "rawMarkdown": "It would make zero sense if private lb has low quality labels (ie less than 10 evaluators distribution), why would they want a model that generalizes well to the low quality labels? it's counter intuitive of the goal of the competition imo.",
      "votes": 4,
      "replies": [
        {
          "id": 2698464,
          "postDate": "2024-03-15T14:22:43.780Z",
          "content": "<p>fair enough, though what does make &gt;=10 evaluators better than 3 highly skilled professionals? Am I missing something?</p>",
          "rawMarkdown": "fair enough, though what does make >=10 evaluators better than 3 highly skilled professionals? Am I missing something?",
          "votes": 1,
          "replies": [
            {
              "id": 2698548,
              "postDate": "2024-03-15T15:12:03.890Z",
              "content": "<p>From Overview -</p>\n<blockquote>\n  <p>The EEG segments used in this competition have been annotated, or classified, by a group of experts. In some cases experts completely agree about the correct label. On other cases the experts disagree. We call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.</p>\n</blockquote>\n<p>Thought perhaps more annotators tended to be used for the edge cases or where more disagreements?  Need to check out the different patterns see if that makes a difference as well. </p>",
              "rawMarkdown": "From Overview -\n>The EEG segments used in this competition have been annotated, or classified, by a group of experts. In some cases experts completely agree about the correct label. On other cases the experts disagree. We call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.\n\nThought perhaps more annotators tended to be used for the edge cases or where more disagreements?  Need to check out the different patterns see if that makes a difference as well. ",
              "votes": 2
            },
            {
              "id": 2698551,
              "postDate": "2024-03-15T15:13:45.260Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2698583,
              "postDate": "2024-03-15T15:26:15.630Z",
              "content": "<p>The competition hosts say there will be 3-20 evaluators. There will absolutely be test with low votes. The question is how much! And I think personally hosts should just tell us the distribution to avoid wasting time. Either many top data scientists will fall because they were LB chasing or many good teams who were trying to make good wholistic models will fall. If they tell us the distribution we could prepare for all of us to meet the businesses needs. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "rawMarkdown": "The competition hosts say there will be 3-20 evaluators. There will absolutely be test with low votes. The question is how much! And I think personally hosts should just tell us the distribution to avoid wasting time. Either many top data scientists will fall because they were LB chasing or many good teams who were trying to make good wholistic models will fall. If they tell us the distribution we could prepare for all of us to meet the businesses needs. @sohier ",
              "votes": 2
            },
            {
              "id": 2698937,
              "postDate": "2024-03-15T17:32:03.263Z",
              "content": "<p>If you treat them equally, then why votes &gt;= 10 dataset with only 1.4% seziure dominate instances while for the whole dataset the ratio is about 18%. They are totally different datasets, why your local cv by chance match 35% test dataset very well? And you could assume the other 65% will not? You get 0.3 result on your votes &gt;= 10 dataset both local cv and LB, will you assume the other 65% test dataset will give you score about 0.6? Seems impossible. I think kaggle will randomly split 35% and 65%, I could not think out reasons why they do not.</p>",
              "rawMarkdown": "If you treat them equally, then why votes >= 10 dataset with only 1.4% seziure dominate instances while for the whole dataset the ratio is about 18%. They are totally different datasets, why your local cv by chance match 35% test dataset very well? And you could assume the other 65% will not? You get 0.3 result on your votes >= 10 dataset both local cv and LB, will you assume the other 65% test dataset will give you score about 0.6? Seems impossible. I think kaggle will randomly split 35% and 65%, I could not think out reasons why they do not.",
              "votes": 3
            },
            {
              "id": 2699178,
              "postDate": "2024-03-15T19:48:09.887Z",
              "content": "<p>Not enough data maybe? Maybe that why they evaluate us on low votes + high votes.</p>",
              "rawMarkdown": "Not enough data maybe? Maybe that why they evaluate us on low votes + high votes."
            }
          ]
        }
      ]
    },
    {
      "id": 2710815,
      "postDate": "2024-03-22T15:19:58.150Z",
      "content": "<p>If I were hosting this competition and investing $, I would hope to get a generally applicable mode for actual practice and set the private test to have a similar distribution of the training set as possible. Yet, we may practice in CV, some folds make better scores, and others do not even try to set a similar distribution in training and validation. This might happen to the test used in the public LB. Yet the final test used in private LB is nearly twice as large as that used in public LB, and its distribution may get closer to the one used in the training set we have.</p>",
      "rawMarkdown": "If I were hosting this competition and investing $, I would hope to get a generally applicable mode for actual practice and set the private test to have a similar distribution of the training set as possible. Yet, we may practice in CV, some folds make better scores, and others do not even try to set a similar distribution in training and validation. This might happen to the test used in the public LB. Yet the final test used in private LB is nearly twice as large as that used in public LB, and its distribution may get closer to the one used in the training set we have.",
      "votes": 1,
      "replies": [
        {
          "id": 2711536,
          "postDate": "2024-03-22T22:40:00.030Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2700609,
      "postDate": "2024-03-16T15:54:14.953Z",
      "content": "<p>I think I'm going to take the Blue Pill 🥏, if that refers to the LB 0.3x models, based on the assumption that public and private test datasets came from the same dataset, were shuffled and split randomly, which means similliar data distributions. <br>\nThe competition owners didn't mention or warn of differences between the public and private datasets.<br>\nTwo stage training has some logic behind it, an assumption was that more votes should give us more confidence in the votes, based on that assumption the data were split, and it actually made the model better, which confirms the assumption, this wasn't based on blind luck.</p>",
      "rawMarkdown": "I think I'm going to take the Blue Pill 🥏, if that refers to the LB 0.3x models, based on the assumption that public and private test datasets came from the same dataset, were shuffled and split randomly, which means similliar data distributions. \nThe competition owners didn't mention or warn of differences between the public and private datasets.\nTwo stage training has some logic behind it, an assumption was that more votes should give us more confidence in the votes, based on that assumption the data were split, and it actually made the model better, which confirms the assumption, this wasn't based on blind luck.",
      "votes": 1,
      "replies": [
        {
          "id": 2700621,
          "postDate": "2024-03-16T16:02:18.290Z",
          "content": "<p>You are correct, but in my experience, they don't always warn about it. I have been on kaggle for a year and I can think of only one competition where they have made competitors aware of the drift (not saying there wasn't more though). But I do know in the Benetech Making Graphs Accessible comp, the AMP Parkinsons comp, the ICR comp, the Kaggle LLM Science Exam comp, and the Linking Writing to Process to Quality Writing comp did not warn of this. The only one I know that did warn of it was the SenNet plus HOA comp and it was only because the resolution of images changed, meaning it wouldnt  submit on private data.</p>",
          "rawMarkdown": "You are correct, but in my experience, they don't always warn about it. I have been on kaggle for a year and I can think of only one competition where they have made competitors aware of the drift (not saying there wasn't more though). But I do know in the Benetech Making Graphs Accessible comp, the AMP Parkinsons comp, the ICR comp, the Kaggle LLM Science Exam comp, and the Linking Writing to Process to Quality Writing comp did not warn of this. The only one I know that did warn of it was the SenNet plus HOA comp and it was only because the resolution of images changed, meaning it wouldnt  submit on private data.",
          "votes": 1,
          "replies": [
            {
              "id": 2703899,
              "postDate": "2024-03-18T13:12:27.037Z",
              "content": "<p>Cannot agree more. </p>",
              "rawMarkdown": "Cannot agree more. ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2717239,
          "postDate": "2024-03-26T13:09:31.270Z",
          "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> I wonder how do you do the two-stage training. Correct me if I am wrong -- model is finetuned on data that has smaller KL compared to uniform distribution. If the whole process doesn't involve any change on the valid set (only part of the train set is selected to train model in the second stage), then it shouldn't be able to improve CV from 0.5ish to 0.3ish. I suspect there is leakage, is the train valid split same in the two stages</p>",
          "rawMarkdown": "@nartaa I wonder how do you do the two-stage training. Correct me if I am wrong -- model is finetuned on data that has smaller KL compared to uniform distribution. If the whole process doesn't involve any change on the valid set (only part of the train set is selected to train model in the second stage), then it shouldn't be able to improve CV from 0.5ish to 0.3ish. I suspect there is leakage, is the train valid split same in the two stages",
          "votes": 1,
          "replies": [
            {
              "id": 2717246,
              "postDate": "2024-03-26T13:17:47.663Z",
              "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> You're understanding of how it's done is correct.</p>\n<blockquote>\n  <p>it shouldn't be able to improve CV from 0.5ish to 0.3ish</p>\n</blockquote>\n<p>It doesn't improve CV, It stays similar to the first stage CV.</p>\n<blockquote>\n  <p>I suspect there is leakage, is the train valid split same in the two stages</p>\n</blockquote>\n<p>No leakage in that regard.</p>",
              "rawMarkdown": "@renyiwei You're understanding of how it's done is correct.\n\n> it shouldn't be able to improve CV from 0.5ish to 0.3ish\n\nIt doesn't improve CV, It stays similar to the first stage CV.\n\n> I suspect there is leakage, is the train valid split same in the two stages\n\nNo leakage in that regard.",
              "votes": 1
            },
            {
              "id": 2717255,
              "postDate": "2024-03-26T13:23:41.093Z",
              "content": "<p>In that case, an extra stage training is a pure gamble on the distribution of the private LB (whether is softer or harder compared to local data) -- go big or home, lol. </p>",
              "rawMarkdown": "In that case, an extra stage training is a pure gamble on the distribution of the private LB (whether is softer or harder compared to local data) -- go big or home, lol. "
            }
          ]
        }
      ]
    },
    {
      "id": 2698253,
      "postDate": "2024-03-15T11:24:59.343Z",
      "content": "<p>RE the 2 stage - have wondered about duplication of data when using <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique</a> - </p>\n<blockquote>\n  <p>…the sample size to 20,183. That's 3,094 more samples than the original notebook, each with a slight variation in vote distribution.</p>\n</blockquote>\n<p>so train used for Group KFold CV with 2 stage has 20183 used to split into however many folds whereas using eeg_id without duplicates 17089 is used for splits.</p>\n<p>Not sure if that makes sense, perhaps being done differently by various notebooks, but possibly there is duplication that is affecting training and the predictions from the 2 stage models that overfits public LB?    </p>",
      "rawMarkdown": "RE the 2 stage - have wondered about duplication of data when using https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique - \n> ...the sample size to 20,183. That's 3,094 more samples than the original notebook, each with a slight variation in vote distribution.\n\nso train used for Group KFold CV with 2 stage has 20183 used to split into however many folds whereas using eeg_id without duplicates 17089 is used for splits.\n\nNot sure if that makes sense, perhaps being done differently by various notebooks, but possibly there is duplication that is affecting training and the predictions from the 2 stage models that overfits public LB?    \n",
      "votes": 1
    },
    {
      "id": 2697426,
      "postDate": "2024-03-14T22:19:37.797Z",
      "content": "<p>I think you are right. Really we have no way of knowing. However overfitting has yet to work for me. So I will stick to our more wholistic CV strategy as well.</p>",
      "rawMarkdown": "I think you are right. Really we have no way of knowing. However overfitting has yet to work for me. So I will stick to our more wholistic CV strategy as well.",
      "votes": 1,
      "replies": [
        {
          "id": 2697525,
          "postDate": "2024-03-15T00:22:36.930Z",
          "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> so you guys are not to two-stage thing, are you?  Really, I am looking forward to seeing the winning solutions without two-stage…</p>",
          "rawMarkdown": "@cody11null so you guys are not to two-stage thing, are you?  Really, I am looking forward to seeing the winning solutions without two-stage...",
          "votes": 1,
          "replies": [
            {
              "id": 2697526,
              "postDate": "2024-03-15T00:25:41.423Z",
              "content": "<p>Yeah we are not on the 2 stage train. We are in the same boat!</p>",
              "rawMarkdown": "Yeah we are not on the 2 stage train. We are in the same boat!",
              "votes": 1
            },
            {
              "id": 2697528,
              "postDate": "2024-03-15T00:26:36.020Z",
              "content": "<p>:D I am in the two-stage Titanic… On the stage which sank last in the movie, lol. </p>",
              "rawMarkdown": ":D I am in the two-stage Titanic... On the stage which sank last in the movie, lol. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2720272,
      "postDate": "2024-03-28T07:42:02.657Z",
      "content": "<p>I dont understand why they even want to match the distribution of cases where the evaluators disagree? IMO we should be training on high quality labeled data and potentially using the model as an \"expert\" to handle the edge cases and make votes.</p>",
      "rawMarkdown": "I dont understand why they even want to match the distribution of cases where the evaluators disagree? IMO we should be training on high quality labeled data and potentially using the model as an \"expert\" to handle the edge cases and make votes.",
      "votes": 2
    },
    {
      "id": 2697684,
      "postDate": "2024-03-15T04:32:27.207Z",
      "content": "<p>Why not take both since we can choose two submissions?</p>",
      "rawMarkdown": "Why not take both since we can choose two submissions?",
      "votes": 2,
      "replies": [
        {
          "id": 2698002,
          "postDate": "2024-03-15T08:06:16.430Z",
          "content": "<p>We gonna do that most likely.</p>",
          "rawMarkdown": "We gonna do that most likely.",
          "votes": 1
        },
        {
          "id": 2698949,
          "postDate": "2024-03-15T17:44:52.750Z",
          "content": "<p>Only 2 could be choosen, I would rather  choose 1 with best LB, 1 with best local cv on votes &gt;= 10. </p>",
          "rawMarkdown": "Only 2 could be choosen, I would rather  choose 1 with best LB, 1 with best local cv on votes >= 10. ",
          "votes": 1,
          "replies": [
            {
              "id": 2698960,
              "postDate": "2024-03-15T17:58:24.443Z",
              "content": "<p>Those were always same for us. CV on votes &gt;= 10 is almost perfectly correlated with LB.</p>",
              "rawMarkdown": "Those were always same for us. CV on votes >= 10 is almost perfectly correlated with LB."
            },
            {
              "id": 2699028,
              "postDate": "2024-03-15T18:25:33.847Z",
              "content": "<p>Mostly the same, but not always for me.</p>",
              "rawMarkdown": "Mostly the same, but not always for me."
            },
            {
              "id": 2699062,
              "postDate": "2024-03-15T18:48:56.827Z",
              "content": "<p>I wouldn't use one submission for best LB in that case which is probably better on third decimal or something.</p>",
              "rawMarkdown": "I wouldn't use one submission for best LB in that case which is probably better on third decimal or something.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2697428,
      "postDate": "2024-03-14T22:21:39.530Z",
      "content": "<p>We can always count on you to write something about the risk of following public notebooks and overfitting :) (I mean this as a good thing). </p>\n<p>It will be interesting to see how it shakes out in the end!</p>",
      "rawMarkdown": "We can always count on you to write something about the risk of following public notebooks and overfitting :) (I mean this as a good thing). \n\nIt will be interesting to see how it shakes out in the end!",
      "votes": 2,
      "replies": [
        {
          "id": 2697527,
          "postDate": "2024-03-15T00:25:55.683Z",
          "content": "<p>Hah, I've never been a \"trust-your-cv-will-there-be-a-shake up-I-told-you\" guy. Hope, I am still young and have not turned into grandma waiting for a doom's day to come.</p>",
          "rawMarkdown": "Hah, I've never been a \"trust-your-cv-will-there-be-a-shake up-I-told-you\" guy. Hope, I am still young and have not turned into grandma waiting for a doom's day to come.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2697577,
      "postDate": "2024-03-15T01:43:10.743Z",
      "content": "<p>Also achieves CV 0.53/LB 0.37 without two-stage. But why are my results from the two-stage training different from everyone else's? It hurts CV and LB both.😭<br>\nFor example, one of my two-stage training result:  </p>\n<p>First stage: training with all data. <strong>CV 0.57/LB 0.38</strong>  <br>\nSecond stage: training with samples KL Loss(calculated with GT and 1/6) &lt; 5.5. <strong>CV 0.67/LB 0.44</strong>  </p>\n<p>Second stage v2: training with samples with total_evaluators &gt; 10. <strong>CV 0.59/LB 0.39</strong></p>\n<p>I am using the non-overlapp 17089 eeg_id as training and validation set. Can I ask how you did your second stage of training? I think the second stage will become one of my final two approaches.</p>",
      "rawMarkdown": "Also achieves CV 0.53/LB 0.37 without two-stage. But why are my results from the two-stage training different from everyone else's? It hurts CV and LB both.😭\nFor example, one of my two-stage training result:  \n\nFirst stage: training with all data. **CV 0.57/LB 0.38**  \nSecond stage: training with samples KL Loss(calculated with GT and 1/6) < 5.5. **CV 0.67/LB 0.44**  \n \nSecond stage v2: training with samples with total_evaluators > 10. **CV 0.59/LB 0.39**\n\nI am using the non-overlapp 17089 eeg_id as training and validation set. Can I ask how you did your second stage of training? I think the second stage will become one of my final two approaches.",
      "replies": [
        {
          "id": 2697812,
          "postDate": "2024-03-15T05:39:43.487Z",
          "content": "<p>I used the 20000 data points dataframe with two stages and can get 0.3 single model for lb. I also used a modified version of the three spectrograms data preparation (10m, 50s, 10s). What mine differs from yours is I only used sum_votes &lt; 10 for first stage and sum_votes &gt; 10 for second stage. Trained 5 epochs for each stage. I am a little suspicious about the KL loss method because it's not even about the data source anymore. It's just overfitting the vote pattern of public lb. And I agree with the post that two stage method would require some luck to get a good private score.</p>",
          "rawMarkdown": "I used the 20000 data points dataframe with two stages and can get 0.3 single model for lb. I also used a modified version of the three spectrograms data preparation (10m, 50s, 10s). What mine differs from yours is I only used sum_votes < 10 for first stage and sum_votes > 10 for second stage. Trained 5 epochs for each stage. I am a little suspicious about the KL loss method because it's not even about the data source anymore. It's just overfitting the vote pattern of public lb. And I agree with the post that two stage method would require some luck to get a good private score.",
          "votes": 1,
          "replies": [
            {
              "id": 2697822,
              "postDate": "2024-03-15T05:43:00.003Z",
              "content": "<p>you wanna team up?</p>",
              "rawMarkdown": "you wanna team up?"
            },
            {
              "id": 2711649,
              "postDate": "2024-03-23T01:27:43.090Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2698235,
      "postDate": "2024-03-15T11:06:59.057Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2698395,
          "postDate": "2024-03-15T13:31:08.827Z",
          "content": "<p>I'd rather trust cv on all data because I don't know what is in the hidden part of the test (65%). Thanks on confirming probing test though and that you heavily invested your efforts into subset of n_votes &gt;= 10 believing that it is amazing. I wish there were a solution that generalizes well to never seen data rather than overfit with bunch of voodo kaggle tricks to a specific subset. You might disagree since kaggle is all about gold zone, right?</p>",
          "rawMarkdown": "I'd rather trust cv on all data because I don't know what is in the hidden part of the test (65%). Thanks on confirming probing test though and that you heavily invested your efforts into subset of n_votes >= 10 believing that it is amazing. I wish there were a solution that generalizes well to never seen data rather than overfit with bunch of voodo kaggle tricks to a specific subset. You might disagree since kaggle is all about gold zone, right?",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2697672,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2024-03-15T04:08:06.270000",
      "content": "<p>For dataset classification model, which features do you use?<br>\nDoes your model still work if you selected data with the same diagnoses (e.g., lpd data with voter &lt; 9 vs.  lpd data with voter &gt; 9)?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2698401,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T13:34:33.807000",
          "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> well I would leave that to the reader rather. There are things you can do with classification model, split data into different subsets based on votes, specific classes and potentially get even better discriminating power…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2698436,
      "author_name": "Catharis",
      "author_url": "",
      "post_date": "2024-03-15T14:11:00.373000",
      "content": "<p>It would make zero sense if private lb has low quality labels (ie less than 10 evaluators distribution), why would they want a model that generalizes well to the low quality labels? it's counter intuitive of the goal of the competition imo.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2698464,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T14:22:43.780000",
          "content": "<p>fair enough, though what does make &gt;=10 evaluators better than 3 highly skilled professionals? Am I missing something?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2698548,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2024-03-15T15:12:03.890000",
              "content": "<p>From Overview -</p>\n<blockquote>\n  <p>The EEG segments used in this competition have been annotated, or classified, by a group of experts. In some cases experts completely agree about the correct label. On other cases the experts disagree. We call segments where there are high levels of agreement “idealized” patterns. Cases where ~1/2 of experts give a label as “other” and ~1/2 give one of the remaining five labels, we call “proto patterns”. Cases where experts are approximately split between 2 of the 5 named patterns, we call “edge cases”.</p>\n</blockquote>\n<p>Thought perhaps more annotators tended to be used for the edge cases or where more disagreements?  Need to check out the different patterns see if that makes a difference as well. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2698551,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-15T15:13:45.260000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2698583,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2024-03-15T15:26:15.630000",
              "content": "<p>The competition hosts say there will be 3-20 evaluators. There will absolutely be test with low votes. The question is how much! And I think personally hosts should just tell us the distribution to avoid wasting time. Either many top data scientists will fall because they were LB chasing or many good teams who were trying to make good wholistic models will fall. If they tell us the distribution we could prepare for all of us to meet the businesses needs. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2698937,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-15T17:32:03.263000",
              "content": "<p>If you treat them equally, then why votes &gt;= 10 dataset with only 1.4% seziure dominate instances while for the whole dataset the ratio is about 18%. They are totally different datasets, why your local cv by chance match 35% test dataset very well? And you could assume the other 65% will not? You get 0.3 result on your votes &gt;= 10 dataset both local cv and LB, will you assume the other 65% test dataset will give you score about 0.6? Seems impossible. I think kaggle will randomly split 35% and 65%, I could not think out reasons why they do not.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2699178,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2024-03-15T19:48:09.887000",
              "content": "<p>Not enough data maybe? Maybe that why they evaluate us on low votes + high votes.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2710815,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "2024-03-22T15:19:58.150000",
      "content": "<p>If I were hosting this competition and investing $, I would hope to get a generally applicable mode for actual practice and set the private test to have a similar distribution of the training set as possible. Yet, we may practice in CV, some folds make better scores, and others do not even try to set a similar distribution in training and validation. This might happen to the test used in the public LB. Yet the final test used in private LB is nearly twice as large as that used in public LB, and its distribution may get closer to the one used in the training set we have.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2711536,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-22T22:40:00.030000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2700609,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-03-16T15:54:14.953000",
      "content": "<p>I think I'm going to take the Blue Pill 🥏, if that refers to the LB 0.3x models, based on the assumption that public and private test datasets came from the same dataset, were shuffled and split randomly, which means similliar data distributions. <br>\nThe competition owners didn't mention or warn of differences between the public and private datasets.<br>\nTwo stage training has some logic behind it, an assumption was that more votes should give us more confidence in the votes, based on that assumption the data were split, and it actually made the model better, which confirms the assumption, this wasn't based on blind luck.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2700621,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2024-03-16T16:02:18.290000",
          "content": "<p>You are correct, but in my experience, they don't always warn about it. I have been on kaggle for a year and I can think of only one competition where they have made competitors aware of the drift (not saying there wasn't more though). But I do know in the Benetech Making Graphs Accessible comp, the AMP Parkinsons comp, the ICR comp, the Kaggle LLM Science Exam comp, and the Linking Writing to Process to Quality Writing comp did not warn of this. The only one I know that did warn of it was the SenNet plus HOA comp and it was only because the resolution of images changed, meaning it wouldnt  submit on private data.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2703899,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-03-18T13:12:27.037000",
              "content": "<p>Cannot agree more. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2717239,
          "author_name": "Roy Wei",
          "author_url": "",
          "post_date": "2024-03-26T13:09:31.270000",
          "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> I wonder how do you do the two-stage training. Correct me if I am wrong -- model is finetuned on data that has smaller KL compared to uniform distribution. If the whole process doesn't involve any change on the valid set (only part of the train set is selected to train model in the second stage), then it shouldn't be able to improve CV from 0.5ish to 0.3ish. I suspect there is leakage, is the train valid split same in the two stages</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2717246,
              "author_name": "Danial Zakaria",
              "author_url": "",
              "post_date": "2024-03-26T13:17:47.663000",
              "content": "<p><a href=\"https://www.kaggle.com/renyiwei\" target=\"_blank\">@renyiwei</a> You're understanding of how it's done is correct.</p>\n<blockquote>\n  <p>it shouldn't be able to improve CV from 0.5ish to 0.3ish</p>\n</blockquote>\n<p>It doesn't improve CV, It stays similar to the first stage CV.</p>\n<blockquote>\n  <p>I suspect there is leakage, is the train valid split same in the two stages</p>\n</blockquote>\n<p>No leakage in that regard.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2717255,
              "author_name": "Roy Wei",
              "author_url": "",
              "post_date": "2024-03-26T13:23:41.093000",
              "content": "<p>In that case, an extra stage training is a pure gamble on the distribution of the private LB (whether is softer or harder compared to local data) -- go big or home, lol. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2698253,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2024-03-15T11:24:59.343000",
      "content": "<p>RE the 2 stage - have wondered about duplication of data when using <a href=\"https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique\" target=\"_blank\">https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique</a> - </p>\n<blockquote>\n  <p>…the sample size to 20,183. That's 3,094 more samples than the original notebook, each with a slight variation in vote distribution.</p>\n</blockquote>\n<p>so train used for Group KFold CV with 2 stage has 20183 used to split into however many folds whereas using eeg_id without duplicates 17089 is used for splits.</p>\n<p>Not sure if that makes sense, perhaps being done differently by various notebooks, but possibly there is duplication that is affecting training and the predictions from the 2 stage models that overfits public LB?    </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2697426,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-03-14T22:19:37.797000",
      "content": "<p>I think you are right. Really we have no way of knowing. However overfitting has yet to work for me. So I will stick to our more wholistic CV strategy as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2697525,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T00:22:36.930000",
          "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> so you guys are not to two-stage thing, are you?  Really, I am looking forward to seeing the winning solutions without two-stage…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2697526,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2024-03-15T00:25:41.423000",
              "content": "<p>Yeah we are not on the 2 stage train. We are in the same boat!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2697528,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-03-15T00:26:36.020000",
              "content": "<p>:D I am in the two-stage Titanic… On the stage which sank last in the movie, lol. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2720272,
      "author_name": "calicrypto",
      "author_url": "",
      "post_date": "2024-03-28T07:42:02.657000",
      "content": "<p>I dont understand why they even want to match the distribution of cases where the evaluators disagree? IMO we should be training on high quality labeled data and potentially using the model as an \"expert\" to handle the edge cases and make votes.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2697684,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-03-15T04:32:27.207000",
      "content": "<p>Why not take both since we can choose two submissions?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2698002,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T08:06:16.430000",
          "content": "<p>We gonna do that most likely.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2698949,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2024-03-15T17:44:52.750000",
          "content": "<p>Only 2 could be choosen, I would rather  choose 1 with best LB, 1 with best local cv on votes &gt;= 10. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2698960,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-15T17:58:24.443000",
              "content": "<p>Those were always same for us. CV on votes &gt;= 10 is almost perfectly correlated with LB.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2699028,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2024-03-15T18:25:33.847000",
              "content": "<p>Mostly the same, but not always for me.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2699062,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-15T18:48:56.827000",
              "content": "<p>I wouldn't use one submission for best LB in that case which is probably better on third decimal or something.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2697428,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-03-14T22:21:39.530000",
      "content": "<p>We can always count on you to write something about the risk of following public notebooks and overfitting :) (I mean this as a good thing). </p>\n<p>It will be interesting to see how it shakes out in the end!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2697527,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T00:25:55.683000",
          "content": "<p>Hah, I've never been a \"trust-your-cv-will-there-be-a-shake up-I-told-you\" guy. Hope, I am still young and have not turned into grandma waiting for a doom's day to come.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2697577,
      "author_name": "Wisp Vale",
      "author_url": "",
      "post_date": "2024-03-15T01:43:10.743000",
      "content": "<p>Also achieves CV 0.53/LB 0.37 without two-stage. But why are my results from the two-stage training different from everyone else's? It hurts CV and LB both.😭<br>\nFor example, one of my two-stage training result:  </p>\n<p>First stage: training with all data. <strong>CV 0.57/LB 0.38</strong>  <br>\nSecond stage: training with samples KL Loss(calculated with GT and 1/6) &lt; 5.5. <strong>CV 0.67/LB 0.44</strong>  </p>\n<p>Second stage v2: training with samples with total_evaluators &gt; 10. <strong>CV 0.59/LB 0.39</strong></p>\n<p>I am using the non-overlapp 17089 eeg_id as training and validation set. Can I ask how you did your second stage of training? I think the second stage will become one of my final two approaches.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2697812,
          "author_name": "LLLEEEOOOH",
          "author_url": "",
          "post_date": "2024-03-15T05:39:43.487000",
          "content": "<p>I used the 20000 data points dataframe with two stages and can get 0.3 single model for lb. I also used a modified version of the three spectrograms data preparation (10m, 50s, 10s). What mine differs from yours is I only used sum_votes &lt; 10 for first stage and sum_votes &gt; 10 for second stage. Trained 5 epochs for each stage. I am a little suspicious about the KL loss method because it's not even about the data source anymore. It's just overfitting the vote pattern of public lb. And I agree with the post that two stage method would require some luck to get a good private score.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2697822,
              "author_name": "LLLEEEOOOH",
              "author_url": "",
              "post_date": "2024-03-15T05:43:00.003000",
              "content": "<p>you wanna team up?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2711649,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-23T01:27:43.090000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2698235,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-15T11:06:59.057000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2698395,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-03-15T13:31:08.827000",
          "content": "<p>I'd rather trust cv on all data because I don't know what is in the hidden part of the test (65%). Thanks on confirming probing test though and that you heavily invested your efforts into subset of n_votes &gt;= 10 believing that it is amazing. I wish there were a solution that generalizes well to never seen data rather than overfit with bunch of voodo kaggle tricks to a specific subset. You might disagree since kaggle is all about gold zone, right?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2697353": "Hi,\n\n## Intro\n\"You take the blue pill - the story ends, you wake up in your bed and believe whatever you want to believe. You take the red pill - you stay in Wonderland and I show you how deep the rabbit hole goes.\"\n\n1. Currently, there are many topics/kernels opened which are pro two-stage training, and recently there is the one which warns us about potential harm.\n2. Few weeks ago I fitted the model with a binary target - (0, 7] voters and [10, 28] voters, I used AUCROC to see if the model can distinguish between two datasets. Well it did (0.62 score, weak though, but still):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fef54e7736fcbcd710b6e2c5bf473864c%2Faucroc.png?generation=1710448838631576&alt=media)\n\n## Outro\nHonestly, I am kinda late here but I just want to post it separately, so it won't be lost in the comments elsewhere. The story is that polishing with second stage means heaviliy overfitting to the public dataset. If we are lucky enough, then yes - that will work, though I'd rather trust my **0.4 LB** with global **CV 0.53** than two-stage **0.3 CV** with **LB 0.3** but heavily overfited by using the second stage.\n\nYou make a decision what pill you take.",
    "2697672": "For dataset classification model, which features do you use?\nDoes your model still work if you selected data with the same diagnoses (e.g., lpd data with voter < 9 vs.  lpd data with voter > 9)?",
    "2698436": "It would make zero sense if private lb has low quality labels (ie less than 10 evaluators distribution), why would they want a model that generalizes well to the low quality labels? it's counter intuitive of the goal of the competition imo.",
    "2710815": "If I were hosting this competition and investing $, I would hope to get a generally applicable mode for actual practice and set the private test to have a similar distribution of the training set as possible. Yet, we may practice in CV, some folds make better scores, and others do not even try to set a similar distribution in training and validation. This might happen to the test used in the public LB. Yet the final test used in private LB is nearly twice as large as that used in public LB, and its distribution may get closer to the one used in the training set we have.",
    "2700609": "I think I'm going to take the Blue Pill 🥏, if that refers to the LB 0.3x models, based on the assumption that public and private test datasets came from the same dataset, were shuffled and split randomly, which means similliar data distributions. \nThe competition owners didn't mention or warn of differences between the public and private datasets.\nTwo stage training has some logic behind it, an assumption was that more votes should give us more confidence in the votes, based on that assumption the data were split, and it actually made the model better, which confirms the assumption, this wasn't based on blind luck.",
    "2698253": "RE the 2 stage - have wondered about duplication of data when using https://www.kaggle.com/datasets/seanbearden/eeg-spectrogram-by-lead-id-unique - \n> ...the sample size to 20,183. That's 3,094 more samples than the original notebook, each with a slight variation in vote distribution.\n\nso train used for Group KFold CV with 2 stage has 20183 used to split into however many folds whereas using eeg_id without duplicates 17089 is used for splits.\n\nNot sure if that makes sense, perhaps being done differently by various notebooks, but possibly there is duplication that is affecting training and the predictions from the 2 stage models that overfits public LB?    \n",
    "2697426": "I think you are right. Really we have no way of knowing. However overfitting has yet to work for me. So I will stick to our more wholistic CV strategy as well.",
    "2720272": "I dont understand why they even want to match the distribution of cases where the evaluators disagree? IMO we should be training on high quality labeled data and potentially using the model as an \"expert\" to handle the edge cases and make votes.",
    "2697684": "Why not take both since we can choose two submissions?",
    "2697428": "We can always count on you to write something about the risk of following public notebooks and overfitting :) (I mean this as a good thing). \n\nIt will be interesting to see how it shakes out in the end!",
    "2697577": "Also achieves CV 0.53/LB 0.37 without two-stage. But why are my results from the two-stage training different from everyone else's? It hurts CV and LB both.😭\nFor example, one of my two-stage training result:  \n\nFirst stage: training with all data. **CV 0.57/LB 0.38**  \nSecond stage: training with samples KL Loss(calculated with GT and 1/6) < 5.5. **CV 0.67/LB 0.44**  \n \nSecond stage v2: training with samples with total_evaluators > 10. **CV 0.59/LB 0.39**\n\nI am using the non-overlapp 17089 eeg_id as training and validation set. Can I ask how you did your second stage of training? I think the second stage will become one of my final two approaches.",
    "2698235": ""
  }
}