{
  "id": 492262,
  "title": "Why shake-up didn't happen",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492262",
  "author_name": "Gunes Evitan",
  "post_date": "2024-04-09T03:48:29.582000",
  "votes": 49,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Potential shake-up is discussed in multiple topics here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374</a></li>\n</ul>\n<p>but it didn't happen. I was pretty much sure that it wasn't going to happen. I'm going to explain how we came to that conclusion.</p>\n<p>In early stages, I suffered from CV/LB discrepancy like everyone else. LB score was way better compared to CV score which isn't common in Kaggle competitions. Obviously, LB distribution was different than training set.</p>\n<p>After couple weeks, I teamed up with Jebastin and started this competition officially. We started with reading <a href=\"https://pubmed.ncbi.nlm.nih.gov/36878708/\" target=\"_blank\">SPaRCNet</a> paper. We found that the dataset was very similar, and they split it like this. It can be seen that Test Dataset 4 has only annotations with votes &gt;= 10.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F210d51a546ca638d23b8d12ef5c0e9c1%2F3.png?generation=1712634348452427&amp;alt=media\" alt=\"3\"></p>\n<p>We also found that 1st author of SPaRCNet paper (Jin Jing) and this competition was the same person. We believed that person would use the same split criteria for this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff2bfc62086ed8318b646ff4cf3091321%2F2.png?generation=1712634374839619&amp;alt=media\" alt=\"1\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3cc9b96cbb5069426d0e587ab391da07%2F1.png?generation=1712634383550684&amp;alt=media\" alt=\"2\"></p>\n<p>The reasoning behind this split is it always makes sense to evaluate models on less noisy labels. Annotators are not perfect, and they can make mistakes. That's why there are things like vote count and annotator quality. They kept high number of voters and high quality annotators in test dataset in order to test more accurately. This is very common in medical domain, and I was also testing some of my projects on similar test sets in my previous jobs. This is called the gold standard, and it is the most reliable testing technique. Models should learn the high quality annotations. Otherwise, this project wouldn't be successful.</p>",
  "messages": [
    {
      "id": 2742733,
      "postDate": "2024-04-09T03:48:29.583Z",
      "content": "<p>Potential shake-up is discussed in multiple topics here:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374</a></li>\n</ul>\n<p>but it didn't happen. I was pretty much sure that it wasn't going to happen. I'm going to explain how we came to that conclusion.</p>\n<p>In early stages, I suffered from CV/LB discrepancy like everyone else. LB score was way better compared to CV score which isn't common in Kaggle competitions. Obviously, LB distribution was different than training set.</p>\n<p>After couple weeks, I teamed up with Jebastin and started this competition officially. We started with reading <a href=\"https://pubmed.ncbi.nlm.nih.gov/36878708/\" target=\"_blank\">SPaRCNet</a> paper. We found that the dataset was very similar, and they split it like this. It can be seen that Test Dataset 4 has only annotations with votes &gt;= 10.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F210d51a546ca638d23b8d12ef5c0e9c1%2F3.png?generation=1712634348452427&amp;alt=media\" alt=\"3\"></p>\n<p>We also found that 1st author of SPaRCNet paper (Jin Jing) and this competition was the same person. We believed that person would use the same split criteria for this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff2bfc62086ed8318b646ff4cf3091321%2F2.png?generation=1712634374839619&amp;alt=media\" alt=\"1\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3cc9b96cbb5069426d0e587ab391da07%2F1.png?generation=1712634383550684&amp;alt=media\" alt=\"2\"></p>\n<p>The reasoning behind this split is it always makes sense to evaluate models on less noisy labels. Annotators are not perfect, and they can make mistakes. That's why there are things like vote count and annotator quality. They kept high number of voters and high quality annotators in test dataset in order to test more accurately. This is very common in medical domain, and I was also testing some of my projects on similar test sets in my previous jobs. This is called the gold standard, and it is the most reliable testing technique. Models should learn the high quality annotations. Otherwise, this project wouldn't be successful.</p>",
      "rawMarkdown": "Potential shake-up is discussed in multiple topics here:\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374\n\nbut it didn't happen. I was pretty much sure that it wasn't going to happen. I'm going to explain how we came to that conclusion.\n\nIn early stages, I suffered from CV/LB discrepancy like everyone else. LB score was way better compared to CV score which isn't common in Kaggle competitions. Obviously, LB distribution was different than training set.\n\nAfter couple weeks, I teamed up with Jebastin and started this competition officially. We started with reading [SPaRCNet](https://pubmed.ncbi.nlm.nih.gov/36878708/) paper. We found that the dataset was very similar, and they split it like this. It can be seen that Test Dataset 4 has only annotations with votes >= 10.\n\n ![3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F210d51a546ca638d23b8d12ef5c0e9c1%2F3.png?generation=1712634348452427&alt=media)\n\nWe also found that 1st author of SPaRCNet paper (Jin Jing) and this competition was the same person. We believed that person would use the same split criteria for this competition.\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff2bfc62086ed8318b646ff4cf3091321%2F2.png?generation=1712634374839619&alt=media)\n\n![2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3cc9b96cbb5069426d0e587ab391da07%2F1.png?generation=1712634383550684&alt=media)\n\nThe reasoning behind this split is it always makes sense to evaluate models on less noisy labels. Annotators are not perfect, and they can make mistakes. That's why there are things like vote count and annotator quality. They kept high number of voters and high quality annotators in test dataset in order to test more accurately. This is very common in medical domain, and I was also testing some of my projects on similar test sets in my previous jobs. This is called the gold standard, and it is the most reliable testing technique. Models should learn the high quality annotations. Otherwise, this project wouldn't be successful.",
      "votes": 49
    },
    {
      "id": 2742803,
      "postDate": "2024-04-09T05:17:51.040Z",
      "content": "<p>Performance of the model trained with KL-div loss depends on the given distribution of the training dataset, - therefore we must have probed LB in order to find the \"right\" training distribution, and I didn't like it</p>\n<p>I wonder what competition would look like if orgs used metric that doesn't depend on the absolute values of probabilities (some ranking metric per-class for example)</p>",
      "rawMarkdown": "Performance of the model trained with KL-div loss depends on the given distribution of the training dataset, - therefore we must have probed LB in order to find the \"right\" training distribution, and I didn't like it\n\nI wonder what competition would look like if orgs used metric that doesn't depend on the absolute values of probabilities (some ranking metric per-class for example)",
      "votes": 6,
      "replies": [
        {
          "id": 2743298,
          "postDate": "2024-04-09T11:08:29.480Z",
          "content": "<p>That's correct but it isn't much different than binary classification when you think about it. Annotators classify EEGs and they generate a similar distribution to what models would do if this was binary classification.</p>",
          "rawMarkdown": "That's correct but it isn't much different than binary classification when you think about it. Annotators classify EEGs and they generate a similar distribution to what models would do if this was binary classification."
        }
      ]
    },
    {
      "id": 2743124,
      "postDate": "2024-04-09T08:31:30.623Z",
      "content": "<p>Great spot! High level research makes all the difference! I figured I was missing something with your confidence in previous posts, but I never found it. I guess this is a lesson learned for me and something I can improve on moving forward! Great work on this one!</p>",
      "rawMarkdown": "Great spot! High level research makes all the difference! I figured I was missing something with your confidence in previous posts, but I never found it. I guess this is a lesson learned for me and something I can improve on moving forward! Great work on this one!",
      "votes": 3,
      "replies": [
        {
          "id": 2743303,
          "postDate": "2024-04-09T11:12:37.007Z",
          "content": "<p>Thanks so much! It's hard to find those subtle things. My confidence was coming from my past experience. I worked on medical imaging in my previous job for 2 years and we used same kind of validation on many projects.</p>",
          "rawMarkdown": "Thanks so much! It's hard to find those subtle things. My confidence was coming from my past experience. I worked on medical imaging in my previous job for 2 years and we used same kind of validation on many projects.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2747056,
      "postDate": "2024-04-11T16:31:26.500Z",
      "content": "<p>Nice explanation. I was also confident there would not be a shakeup</p>\n<blockquote>\n  <p>It can be seen that Test Dataset 4 has only annotations with votes &gt;= 10.</p>\n</blockquote>\n<p>Note that I think the Test Dataset 4 (i.e. Kaggle LB test) has annotations with <code>3 &lt;= votes &lt;= 20</code> as Kaggle says in the comp data description. </p>\n<blockquote>\n  <p>Kaggle says: Note that the test samples had between 3 and 20 annotators.</p>\n</blockquote>\n<p>I think Test Dataset 4 is created like this</p>\n<ul>\n<li>Begin with all EEG</li>\n<li>Filter samples with <code>vote&gt;=10</code> (at this point we have both <strong>expert annotations</strong> and <strong>fellowship trained annotations</strong>)</li>\n<li>Now change the targets on these samples by <strong>only keeping 20 experts' annotations</strong> (i.e remove fellowship annotations)</li>\n<li>The labels now have <code>3 &lt;= votes &lt;= 20</code>.</li>\n<li>Split this in half to make Test Dataset 3 and Test Dataset 4</li>\n<li>Use Test Dataset 4 as Kaggle LB </li>\n</ul>\n<p>Hence the final Kaggle leaderboard ground truth labels have <code>3 &lt;= votes &lt;= 20</code> where each vote is from one of the 20 experts</p>",
      "rawMarkdown": "Nice explanation. I was also confident there would not be a shakeup\n>It can be seen that Test Dataset 4 has only annotations with votes >= 10.\n\nNote that I think the Test Dataset 4 (i.e. Kaggle LB test) has annotations with `3 <= votes <= 20` as Kaggle says in the comp data description. \n> Kaggle says: Note that the test samples had between 3 and 20 annotators.\n\nI think Test Dataset 4 is created like this\n* Begin with all EEG\n* Filter samples with `vote>=10` (at this point we have both **expert annotations** and **fellowship trained annotations**)\n* Now change the targets on these samples by **only keeping 20 experts' annotations** (i.e remove fellowship annotations)\n* The labels now have `3 <= votes <= 20`.\n* Split this in half to make Test Dataset 3 and Test Dataset 4\n* Use Test Dataset 4 as Kaggle LB \n\nHence the final Kaggle leaderboard ground truth labels have `3 <= votes <= 20` where each vote is from one of the 20 experts",
      "votes": 1
    },
    {
      "id": 2743313,
      "postDate": "2024-04-09T11:20:24.373Z",
      "content": "<p>It is interesting given the response even dismissal in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/475292#2642052\" target=\"_blank\">this discussion</a> on SPaRCNet</p>\n<blockquote>\n  <p>Otherwise, this project wouldn't be successful.</p>\n</blockquote>\n<p>Unless the objective was to evaluate/replicate some aspects of the paper what is successful could be debated. The gold standard does not exist in the real world (usually it is noisy).</p>\n<p>From the Overview - </p>\n<blockquote>\n  <p>Your work may help rapidly improve electroencephalography pattern classification accuracy, unlocking transformative benefits for neurocritical care, epilepsy, and drug development. Advancement in this area may allow doctors and brain researchers to detect seizures or other brain damage to provide faster and more accurate treatments. </p>\n</blockquote>\n<p>Usually consider that as an objective in competitions.</p>",
      "rawMarkdown": "It is interesting given the response even dismissal in [this discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/475292#2642052) on SPaRCNet\n\n>Otherwise, this project wouldn't be successful.\n\nUnless the objective was to evaluate/replicate some aspects of the paper what is successful could be debated. The gold standard does not exist in the real world (usually it is noisy).\n\nFrom the Overview - \n\n>Your work may help rapidly improve electroencephalography pattern classification accuracy, unlocking transformative benefits for neurocritical care, epilepsy, and drug development. Advancement in this area may allow doctors and brain researchers to detect seizures or other brain damage to provide faster and more accurate treatments. \n\nUsually consider that as an objective in competitions.\n",
      "votes": 1
    },
    {
      "id": 2743289,
      "postDate": "2024-04-09T10:57:49.153Z",
      "content": "<p>Wow, it is indeed a deep and fundamental research of the competition, not a leak, and thank you very much for sharing it!  I stick to the 1-stage model until the end but should listen.</p>",
      "rawMarkdown": "Wow, it is indeed a deep and fundamental research of the competition, not a leak, and thank you very much for sharing it!  I stick to the 1-stage model until the end but should listen.",
      "votes": 1
    },
    {
      "id": 2744515,
      "postDate": "2024-04-10T00:36:46.060Z",
      "content": "<p>Thanks for sharing. I have beginner question.<br>\nIs it common on Kaggle that the info of test data like how to split, data quality, and so on is not be provided by the host?</p>",
      "rawMarkdown": "Thanks for sharing. I have beginner question.\nIs it common on Kaggle that the info of test data like how to split, data quality, and so on is not be provided by the host?\n",
      "votes": 2,
      "replies": [
        {
          "id": 2744536,
          "postDate": "2024-04-10T01:27:06.153Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2748293,
          "postDate": "2024-04-12T11:02:01.163Z",
          "content": "<p>It depends, but most of the time it can inferred easily from the CV/LB scores.</p>",
          "rawMarkdown": "It depends, but most of the time it can inferred easily from the CV/LB scores."
        }
      ]
    },
    {
      "id": 2743230,
      "postDate": "2024-04-09T09:47:19.607Z",
      "content": "<p>I haven't read this paper yet. I choose 2 submission, one for number vote &gt;= 10, the other trained with full data. After hubmap comp, I saw a few data with high quality label performed well on private test. But, when train stage 2 with number vote &gt;= 10, I evaluate model on val data with all vote, kl loss of this increase, while kl loss of 10 vote decrease. Therefore, I choose two distributions num_vote &gt;= 10 and full data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7cf53c9176cd4b70a81230ecdd1d1887%2FScreenshot%20from%202024-04-09%2016-57-33.png?generation=1712656683924177&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2Fdadc4146818c92af594751319a11ade3%2FScreenshot%20from%202024-04-09%2016-58-15.png?generation=1712656717235312&amp;alt=media\"></p>",
      "rawMarkdown": " I haven't read this paper yet. I choose 2 submission, one for number vote >= 10, the other trained with full data. After hubmap comp, I saw a few data with high quality label performed well on private test. But, when train stage 2 with number vote >= 10, I evaluate model on val data with all vote, kl loss of this increase, while kl loss of 10 vote decrease. Therefore, I choose two distributions num_vote >= 10 and full data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7cf53c9176cd4b70a81230ecdd1d1887%2FScreenshot%20from%202024-04-09%2016-57-33.png?generation=1712656683924177&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2Fdadc4146818c92af594751319a11ade3%2FScreenshot%20from%202024-04-09%2016-58-15.png?generation=1712656717235312&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 2743306,
          "postDate": "2024-04-09T11:17:37.633Z",
          "content": "<p>We did the same thing and our results for both submissions are almost same.</p>",
          "rawMarkdown": "We did the same thing and our results for both submissions are almost same.",
          "votes": 1,
          "replies": [
            {
              "id": 2743400,
              "postDate": "2024-04-09T12:39:37.720Z",
              "content": "<p>We did the same as well. Honestly, It was a bit easier for me since I knew I would not fight for a floating point gap between prize winners. So one submission was enough for me.</p>",
              "rawMarkdown": "We did the same as well. Honestly, It was a bit easier for me since I knew I would not fight for a floating point gap between prize winners. So one submission was enough for me.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2756112,
      "postDate": "2024-04-16T21:52:42.063Z",
      "content": "<p>How is it possible that LB score was way better compared to CV score? How can a model learn to predict LB data better than its own train set data?</p>",
      "rawMarkdown": "How is it possible that LB score was way better compared to CV score? How can a model learn to predict LB data better than its own train set data?"
    },
    {
      "id": 2745274,
      "postDate": "2024-04-10T14:15:37.487Z",
      "content": "<p>Thinks for sharing ideas!</p>",
      "rawMarkdown": "Thinks for sharing ideas!"
    }
  ],
  "comments": [
    {
      "id": 2742803,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-09T05:17:51.040000",
      "content": "<p>Performance of the model trained with KL-div loss depends on the given distribution of the training dataset, - therefore we must have probed LB in order to find the \"right\" training distribution, and I didn't like it</p>\n<p>I wonder what competition would look like if orgs used metric that doesn't depend on the absolute values of probabilities (some ranking metric per-class for example)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2743298,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-04-09T11:08:29.480000",
          "content": "<p>That's correct but it isn't much different than binary classification when you think about it. Annotators classify EEGs and they generate a similar distribution to what models would do if this was binary classification.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2743124,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-04-09T08:31:30.623000",
      "content": "<p>Great spot! High level research makes all the difference! I figured I was missing something with your confidence in previous posts, but I never found it. I guess this is a lesson learned for me and something I can improve on moving forward! Great work on this one!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2743303,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-04-09T11:12:37.007000",
          "content": "<p>Thanks so much! It's hard to find those subtle things. My confidence was coming from my past experience. I worked on medical imaging in my previous job for 2 years and we used same kind of validation on many projects.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2747056,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-04-11T16:31:26.500000",
      "content": "<p>Nice explanation. I was also confident there would not be a shakeup</p>\n<blockquote>\n  <p>It can be seen that Test Dataset 4 has only annotations with votes &gt;= 10.</p>\n</blockquote>\n<p>Note that I think the Test Dataset 4 (i.e. Kaggle LB test) has annotations with <code>3 &lt;= votes &lt;= 20</code> as Kaggle says in the comp data description. </p>\n<blockquote>\n  <p>Kaggle says: Note that the test samples had between 3 and 20 annotators.</p>\n</blockquote>\n<p>I think Test Dataset 4 is created like this</p>\n<ul>\n<li>Begin with all EEG</li>\n<li>Filter samples with <code>vote&gt;=10</code> (at this point we have both <strong>expert annotations</strong> and <strong>fellowship trained annotations</strong>)</li>\n<li>Now change the targets on these samples by <strong>only keeping 20 experts' annotations</strong> (i.e remove fellowship annotations)</li>\n<li>The labels now have <code>3 &lt;= votes &lt;= 20</code>.</li>\n<li>Split this in half to make Test Dataset 3 and Test Dataset 4</li>\n<li>Use Test Dataset 4 as Kaggle LB </li>\n</ul>\n<p>Hence the final Kaggle leaderboard ground truth labels have <code>3 &lt;= votes &lt;= 20</code> where each vote is from one of the 20 experts</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2743313,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2024-04-09T11:20:24.373000",
      "content": "<p>It is interesting given the response even dismissal in <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/475292#2642052\" target=\"_blank\">this discussion</a> on SPaRCNet</p>\n<blockquote>\n  <p>Otherwise, this project wouldn't be successful.</p>\n</blockquote>\n<p>Unless the objective was to evaluate/replicate some aspects of the paper what is successful could be debated. The gold standard does not exist in the real world (usually it is noisy).</p>\n<p>From the Overview - </p>\n<blockquote>\n  <p>Your work may help rapidly improve electroencephalography pattern classification accuracy, unlocking transformative benefits for neurocritical care, epilepsy, and drug development. Advancement in this area may allow doctors and brain researchers to detect seizures or other brain damage to provide faster and more accurate treatments. </p>\n</blockquote>\n<p>Usually consider that as an objective in competitions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2743289,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "2024-04-09T10:57:49.153000",
      "content": "<p>Wow, it is indeed a deep and fundamental research of the competition, not a leak, and thank you very much for sharing it!  I stick to the 1-stage model until the end but should listen.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2744515,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-04-10T00:36:46.060000",
      "content": "<p>Thanks for sharing. I have beginner question.<br>\nIs it common on Kaggle that the info of test data like how to split, data quality, and so on is not be provided by the host?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744536,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-10T01:27:06.153000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2748293,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-04-12T11:02:01.163000",
          "content": "<p>It depends, but most of the time it can inferred easily from the CV/LB scores.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2743230,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2024-04-09T09:47:19.607000",
      "content": "<p>I haven't read this paper yet. I choose 2 submission, one for number vote &gt;= 10, the other trained with full data. After hubmap comp, I saw a few data with high quality label performed well on private test. But, when train stage 2 with number vote &gt;= 10, I evaluate model on val data with all vote, kl loss of this increase, while kl loss of 10 vote decrease. Therefore, I choose two distributions num_vote &gt;= 10 and full data.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7cf53c9176cd4b70a81230ecdd1d1887%2FScreenshot%20from%202024-04-09%2016-57-33.png?generation=1712656683924177&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2Fdadc4146818c92af594751319a11ade3%2FScreenshot%20from%202024-04-09%2016-58-15.png?generation=1712656717235312&amp;alt=media\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2743306,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-04-09T11:17:37.633000",
          "content": "<p>We did the same thing and our results for both submissions are almost same.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2743400,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-04-09T12:39:37.720000",
              "content": "<p>We did the same as well. Honestly, It was a bit easier for me since I knew I would not fight for a floating point gap between prize winners. So one submission was enough for me.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2756112,
      "author_name": "Rafal Roszak",
      "author_url": "",
      "post_date": "2024-04-16T21:52:42.063000",
      "content": "<p>How is it possible that LB score was way better compared to CV score? How can a model learn to predict LB data better than its own train set data?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2745274,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-10T14:15:37.487000",
      "content": "<p>Thinks for sharing ideas!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742733": "Potential shake-up is discussed in multiple topics here:\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/487479\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/482548\n* https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471374\n\nbut it didn't happen. I was pretty much sure that it wasn't going to happen. I'm going to explain how we came to that conclusion.\n\nIn early stages, I suffered from CV/LB discrepancy like everyone else. LB score was way better compared to CV score which isn't common in Kaggle competitions. Obviously, LB distribution was different than training set.\n\nAfter couple weeks, I teamed up with Jebastin and started this competition officially. We started with reading [SPaRCNet](https://pubmed.ncbi.nlm.nih.gov/36878708/) paper. We found that the dataset was very similar, and they split it like this. It can be seen that Test Dataset 4 has only annotations with votes >= 10.\n\n ![3](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F210d51a546ca638d23b8d12ef5c0e9c1%2F3.png?generation=1712634348452427&alt=media)\n\nWe also found that 1st author of SPaRCNet paper (Jin Jing) and this competition was the same person. We believed that person would use the same split criteria for this competition.\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2Ff2bfc62086ed8318b646ff4cf3091321%2F2.png?generation=1712634374839619&alt=media)\n\n![2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2706866%2F3cc9b96cbb5069426d0e587ab391da07%2F1.png?generation=1712634383550684&alt=media)\n\nThe reasoning behind this split is it always makes sense to evaluate models on less noisy labels. Annotators are not perfect, and they can make mistakes. That's why there are things like vote count and annotator quality. They kept high number of voters and high quality annotators in test dataset in order to test more accurately. This is very common in medical domain, and I was also testing some of my projects on similar test sets in my previous jobs. This is called the gold standard, and it is the most reliable testing technique. Models should learn the high quality annotations. Otherwise, this project wouldn't be successful.",
    "2742803": "Performance of the model trained with KL-div loss depends on the given distribution of the training dataset, - therefore we must have probed LB in order to find the \"right\" training distribution, and I didn't like it\n\nI wonder what competition would look like if orgs used metric that doesn't depend on the absolute values of probabilities (some ranking metric per-class for example)",
    "2743124": "Great spot! High level research makes all the difference! I figured I was missing something with your confidence in previous posts, but I never found it. I guess this is a lesson learned for me and something I can improve on moving forward! Great work on this one!",
    "2747056": "Nice explanation. I was also confident there would not be a shakeup\n>It can be seen that Test Dataset 4 has only annotations with votes >= 10.\n\nNote that I think the Test Dataset 4 (i.e. Kaggle LB test) has annotations with `3 <= votes <= 20` as Kaggle says in the comp data description. \n> Kaggle says: Note that the test samples had between 3 and 20 annotators.\n\nI think Test Dataset 4 is created like this\n* Begin with all EEG\n* Filter samples with `vote>=10` (at this point we have both **expert annotations** and **fellowship trained annotations**)\n* Now change the targets on these samples by **only keeping 20 experts' annotations** (i.e remove fellowship annotations)\n* The labels now have `3 <= votes <= 20`.\n* Split this in half to make Test Dataset 3 and Test Dataset 4\n* Use Test Dataset 4 as Kaggle LB \n\nHence the final Kaggle leaderboard ground truth labels have `3 <= votes <= 20` where each vote is from one of the 20 experts",
    "2743313": "It is interesting given the response even dismissal in [this discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/475292#2642052) on SPaRCNet\n\n>Otherwise, this project wouldn't be successful.\n\nUnless the objective was to evaluate/replicate some aspects of the paper what is successful could be debated. The gold standard does not exist in the real world (usually it is noisy).\n\nFrom the Overview - \n\n>Your work may help rapidly improve electroencephalography pattern classification accuracy, unlocking transformative benefits for neurocritical care, epilepsy, and drug development. Advancement in this area may allow doctors and brain researchers to detect seizures or other brain damage to provide faster and more accurate treatments. \n\nUsually consider that as an objective in competitions.\n",
    "2743289": "Wow, it is indeed a deep and fundamental research of the competition, not a leak, and thank you very much for sharing it!  I stick to the 1-stage model until the end but should listen.",
    "2744515": "Thanks for sharing. I have beginner question.\nIs it common on Kaggle that the info of test data like how to split, data quality, and so on is not be provided by the host?\n",
    "2743230": " I haven't read this paper yet. I choose 2 submission, one for number vote >= 10, the other trained with full data. After hubmap comp, I saw a few data with high quality label performed well on private test. But, when train stage 2 with number vote >= 10, I evaluate model on val data with all vote, kl loss of this increase, while kl loss of 10 vote decrease. Therefore, I choose two distributions num_vote >= 10 and full data.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2F7cf53c9176cd4b70a81230ecdd1d1887%2FScreenshot%20from%202024-04-09%2016-57-33.png?generation=1712656683924177&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2832287%2Fdadc4146818c92af594751319a11ade3%2FScreenshot%20from%202024-04-09%2016-58-15.png?generation=1712656717235312&alt=media)",
    "2756112": "How is it possible that LB score was way better compared to CV score? How can a model learn to predict LB data better than its own train set data?",
    "2745274": "Thinks for sharing ideas!"
  }
}