{
  "id": 482548,
  "title": "Do you think there will be a shake?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/482548",
  "author_name": "Donghui Zhang",
  "post_date": "2024-03-08T07:02:25.320000",
  "votes": 5,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Currently, the two-stage method is popular in public methods. I think the main reason why this method is effective is because it is found that LB data tends to a certain distribution, but this distribution is far from the distribution of the training set. Private data sets may not and the characteristics presented by LB data, so I think we still need to be wary of this method to avoid overfitting LB data and causing shake.</p>",
  "messages": [
    {
      "id": 2687170,
      "postDate": "2024-03-08T10:41:00.367Z",
      "content": "<p>From the competition data page:</p>\n<blockquote>\n  <p>Note that the test samples had between 3 and 20 annotators.</p>\n</blockquote>\n<p>If we train our models on subsets which are more focused on votes &gt; 10 (skipping almost half of the test data). </p>\n<p>How can we not expect a shakeup?</p>",
      "rawMarkdown": "From the competition data page:\n>Note that the test samples had between 3 and 20 annotators.\n\nIf we train our models on subsets which are more focused on votes > 10 (skipping almost half of the test data). \n\nHow can we not expect a shakeup?",
      "votes": 7,
      "replies": [
        {
          "id": 2687271,
          "postDate": "2024-03-08T12:14:37.843Z",
          "content": "<p>I can't think of any logical reason why public and private would be different, and 35% is a good representative ratio.</p>",
          "rawMarkdown": "I can't think of any logical reason why public and private would be different, and 35% is a good representative ratio.",
          "replies": [
            {
              "id": 2687659,
              "postDate": "2024-03-08T17:14:26.370Z",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> different distributions for public/private lbs is very common in Kaggle competitions </p>",
              "rawMarkdown": "@gunesevitan different distributions for public/private lbs is very common in Kaggle competitions ",
              "votes": 5
            },
            {
              "id": 2687738,
              "postDate": "2024-03-08T18:32:03.823Z",
              "content": "<p>Yes, but most of the time it's part of the problem i.e. using one type of data and predicting another type. Organizers either want you to do domain adaptation (computer vision) or build a stable model (time series). I can't imagine either of those scenarios to be expected here.</p>",
              "rawMarkdown": "Yes, but most of the time it's part of the problem i.e. using one type of data and predicting another type. Organizers either want you to do domain adaptation (computer vision) or build a stable model (time series). I can't imagine either of those scenarios to be expected here."
            },
            {
              "id": 2687945,
              "postDate": "2024-03-08T21:59:55.723Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2688188,
              "postDate": "2024-03-09T04:48:09.860Z",
              "content": "<p>So what this there in public LB 35 % data distribution same as the stage 2 training data so public LB good. If the private 65% contains more 1 - 10 annotators data.  Then there is chance of shake up.. Is my understanding is correct? </p>",
              "rawMarkdown": "So what this there in public LB 35 % data distribution same as the stage 2 training data so public LB good. If the private 65% contains more 1 - 10 annotators data.  Then there is chance of shake up.. Is my understanding is correct? ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2688016,
          "postDate": "2024-03-08T23:42:29.330Z",
          "content": "<p>This is interesting, in train about 50% samples are with exactly 3 annotators.</p>",
          "rawMarkdown": "This is interesting, in train about 50% samples are with exactly 3 annotators.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2686945,
      "postDate": "2024-03-08T07:02:25.320Z",
      "content": "<p>Currently, the two-stage method is popular in public methods. I think the main reason why this method is effective is because it is found that LB data tends to a certain distribution, but this distribution is far from the distribution of the training set. Private data sets may not and the characteristics presented by LB data, so I think we still need to be wary of this method to avoid overfitting LB data and causing shake.</p>",
      "rawMarkdown": "Currently, the two-stage method is popular in public methods. I think the main reason why this method is effective is because it is found that LB data tends to a certain distribution, but this distribution is far from the distribution of the training set. Private data sets may not and the characteristics presented by LB data, so I think we still need to be wary of this method to avoid overfitting LB data and causing shake.",
      "votes": 5
    },
    {
      "id": 2689044,
      "postDate": "2024-03-09T16:05:53.830Z",
      "content": "<p>Here is a Cumulative Distribution of Total Evaluators in <em>eeg_id-unique-train_df</em> for reference.</p>\n<pre><code>sorted_counts = df_train[].value_counts().sort_index()\ncumulative_dist = sorted_counts.cumsum() / sorted_counts.()\nplt.plot(cumulative_dist.index, cumulative_dist, marker=, linestyle=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F5e2f00904e50cfc6e303eea31fddb7a1%2Foutput.png?generation=1710000306691110&amp;alt=media\"></p>",
      "rawMarkdown": "Here is a Cumulative Distribution of Total Evaluators in *eeg_id-unique-train_df* for reference.\n```python\nsorted_counts = df_train[\"total_evaluators\"].value_counts().sort_index()\ncumulative_dist = sorted_counts.cumsum() / sorted_counts.sum()\nplt.plot(cumulative_dist.index, cumulative_dist, marker='o', linestyle='--')\n```\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F5e2f00904e50cfc6e303eea31fddb7a1%2Foutput.png?generation=1710000306691110&alt=media)",
      "votes": 4
    },
    {
      "id": 2687252,
      "postDate": "2024-03-08T11:50:54.420Z",
      "content": "<p>The public to private ratio is ~1:2. <br>\nThe ratio of small and large number of experts on train data is ~1:2. <br>\nThink about it… ))</p>",
      "rawMarkdown": "The public to private ratio is ~1:2. \nThe ratio of small and large number of experts on train data is ~1:2. \nThink about it... ))",
      "votes": 1,
      "replies": [
        {
          "id": 2687394,
          "postDate": "2024-03-08T14:31:20.160Z",
          "content": "<p>If that's true, we need two submissions, one for the public set (10~20 annotators), and one for the private set (3~10 annotators).</p>",
          "rawMarkdown": "If that's true, we need two submissions, one for the public set (10~20 annotators), and one for the private set (3~10 annotators).\n",
          "votes": 2,
          "replies": [
            {
              "id": 2687420,
              "postDate": "2024-03-08T15:17:49.700Z",
              "content": "<p>I think you can make a suitable integration for these two distributions</p>",
              "rawMarkdown": "I think you can make a suitable integration for these two distributions",
              "votes": 1
            },
            {
              "id": 2687555,
              "postDate": "2024-03-08T16:05:15.140Z",
              "content": "<p>so if we ensample stage 1 + stage 2 models combined will avoid the 10~20 annotators and 3~10 annotators different distributions issue.<br>\nBut i am currently submitting only stage 2. so chances of shake is lot form my understanding. </p>",
              "rawMarkdown": "so if we ensample stage 1 + stage 2 models combined will avoid the 10~20 annotators and 3~10 annotators different distributions issue.\nBut i am currently submitting only stage 2. so chances of shake is lot form my understanding. ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2687408,
          "postDate": "2024-03-08T14:56:04.123Z",
          "content": "<p>If the organizers made such a split, it would be extremely unfair towards the competitors (us).<br>\nAnd also make the competition a bit meaningless. <br>\nAs Andrew Ng said in one of his courses, validation should be approximately from the same distribution of the test/final target (not exact words, but close enough, I think)<br>\nOtherwise, what are we even doing other than mindlessly guessing? Lol.</p>",
          "rawMarkdown": "If the organizers made such a split, it would be extremely unfair towards the competitors (us).\nAnd also make the competition a bit meaningless. \nAs Andrew Ng said in one of his courses, validation should be approximately from the same distribution of the test/final target (not exact words, but close enough, I think)\nOtherwise, what are we even doing other than mindlessly guessing? Lol.",
          "votes": 2,
          "replies": [
            {
              "id": 2687499,
              "postDate": "2024-03-08T15:48:38.587Z",
              "content": "<p>Totally agree. It doesn't make sense to do such thing and very uncharacteristic of Kaggle. When public and private test sets are different in competitions, they always tell us and want us to build robust models in competitions like hubmap or ubc ocean. Doing the same thing here doesn't make any sense and fails the objective of this competition.</p>",
              "rawMarkdown": "Totally agree. It doesn't make sense to do such thing and very uncharacteristic of Kaggle. When public and private test sets are different in competitions, they always tell us and want us to build robust models in competitions like hubmap or ubc ocean. Doing the same thing here doesn't make any sense and fails the objective of this competition.",
              "votes": 3
            },
            {
              "id": 2687532,
              "postDate": "2024-03-08T15:58:20.473Z",
              "content": "<p>Eactly, also in the Ribonanza competition there was a case that the private test had a different distribution from public/train, but it was OK since they tolds as it is the case and how the distribution differs, so we could plan our validation scheme accordingly. <br>\nBut if they just put random distribution in the private test, that is different both from the train distribution and the public test distribution, and does not inform us on the matter, it is just…well…idk, data sciencve done wrong?</p>",
              "rawMarkdown": "Eactly, also in the Ribonanza competition there was a case that the private test had a different distribution from public/train, but it was OK since they tolds as it is the case and how the distribution differs, so we could plan our validation scheme accordingly. \nBut if they just put random distribution in the private test, that is different both from the train distribution and the public test distribution, and does not inform us on the matter, it is just...well...idk, data sciencve done wrong?",
              "votes": 1
            },
            {
              "id": 2687585,
              "postDate": "2024-03-08T16:14:28.407Z",
              "content": "<p>and throw away 50k for no value whatsoever</p>",
              "rawMarkdown": "and throw away 50k for no value whatsoever",
              "votes": 2
            },
            {
              "id": 2687948,
              "postDate": "2024-03-08T22:03:02.230Z",
              "content": "<p>I totally agree with this but in reality I see most comps have different distributions and different shakeup. CV normally wins</p>",
              "rawMarkdown": "I totally agree with this but in reality I see most comps have different distributions and different shakeup. CV normally wins",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2687976,
      "postDate": "2024-03-08T22:37:54.163Z",
      "content": "<p>The fact that cv and lb scores are so different make me think there could be some shakeup. </p>",
      "rawMarkdown": "The fact that cv and lb scores are so different make me think there could be some shakeup. ",
      "votes": 2
    },
    {
      "id": 2686953,
      "postDate": "2024-03-08T07:11:55.267Z",
      "content": "<p>I'm not expecting any shake up.</p>",
      "rawMarkdown": "I'm not expecting any shake up.",
      "votes": 1,
      "replies": [
        {
          "id": 2686988,
          "postDate": "2024-03-08T07:30:29.943Z",
          "content": "<p>Hope nothing shakes your position😀</p>",
          "rawMarkdown": "Hope nothing shakes your position😀",
          "votes": 1,
          "replies": [
            {
              "id": 2687038,
              "postDate": "2024-03-08T08:08:56.907Z",
              "content": "<p>We are hoping to reach 1st place 🙏</p>",
              "rawMarkdown": "We are hoping to reach 1st place 🙏",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2688512,
      "postDate": "2024-03-09T09:55:01.607Z",
      "content": "<p>In fact, the reason why I feel this way is because in the two months before the game, most teams were stuck at 0.34. Even though the offline cv was optimized, the LB did not change until the second stage method emerged. The situation has improved, but for the front-row team, this does not affect them:)</p>",
      "rawMarkdown": "In fact, the reason why I feel this way is because in the two months before the game, most teams were stuck at 0.34. Even though the offline cv was optimized, the LB did not change until the second stage method emerged. The situation has improved, but for the front-row team, this does not affect them:)",
      "replies": [
        {
          "id": 2688517,
          "postDate": "2024-03-09T09:58:01.830Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479776#2671166\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479776#2671166</a>  That's why I think what he said makes sense</p>",
          "rawMarkdown": "https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479776#2671166  That's why I think what he said makes sense"
        },
        {
          "id": 2688526,
          "postDate": "2024-03-09T10:06:25.217Z",
          "content": "<p>You mean people in top 20 LB .. Not used stage wise training.. So the shake up in  private LB will be less for the front row people… Is my understanding is correct</p>",
          "rawMarkdown": "You mean people in top 20 LB .. Not used stage wise training.. So the shake up in  private LB will be less for the front row people... Is my understanding is correct",
          "replies": [
            {
              "id": 2688551,
              "postDate": "2024-03-09T10:11:45.287Z",
              "content": "<p>They should probably realize that the occurrence of this situation will make the model more robust. As for whether to use the second stage, I can’t judge.</p>",
              "rawMarkdown": "They should probably realize that the occurrence of this situation will make the model more robust. As for whether to use the second stage, I can’t judge.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2688484,
      "postDate": "2024-03-09T09:14:12.643Z",
      "content": "<p>considering <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard\" target=\"_blank\">SenNet + HOA - Hacking the Human Vasculature in 3D</a> anything is possible.</p>\n<p>It is not just number of annotators, but quality of spectrograms - differences between competition provided and generated from EEG raw data have been noted. For the latter no information has been provided on preprocessing methodology which could be different based on source of original data.<br>\nIt is not specified if source of patients are the same facilities, equipment the same, etc.  And distribution of patterns represented could vary, patients should be different for public test from private test.</p>\n<p>Usually hosts/sponsors are looking for robust solutions or to verify what is already known to work best, so time will tell…<br>\nGood luck to everyone!</p>",
      "rawMarkdown": "considering [SenNet + HOA - Hacking the Human Vasculature in 3D](https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard) anything is possible.\n\nIt is not just number of annotators, but quality of spectrograms - differences between competition provided and generated from EEG raw data have been noted. For the latter no information has been provided on preprocessing methodology which could be different based on source of original data.\nIt is not specified if source of patients are the same facilities, equipment the same, etc.  And distribution of patterns represented could vary, patients should be different for public test from private test.\n\nUsually hosts/sponsors are looking for robust solutions or to verify what is already known to work best, so time will tell...\nGood luck to everyone!"
    },
    {
      "id": 2687343,
      "postDate": "2024-03-08T13:38:21.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2687170,
      "author_name": "Muhammad Ahmed",
      "author_url": "",
      "post_date": "2024-03-08T10:41:00.367000",
      "content": "<p>From the competition data page:</p>\n<blockquote>\n  <p>Note that the test samples had between 3 and 20 annotators.</p>\n</blockquote>\n<p>If we train our models on subsets which are more focused on votes &gt; 10 (skipping almost half of the test data). </p>\n<p>How can we not expect a shakeup?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2687271,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-03-08T12:14:37.843000",
          "content": "<p>I can't think of any logical reason why public and private would be different, and 35% is a good representative ratio.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2687659,
              "author_name": "Muhammad Ahmed",
              "author_url": "",
              "post_date": "2024-03-08T17:14:26.370000",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> different distributions for public/private lbs is very common in Kaggle competitions </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2687738,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-08T18:32:03.823000",
              "content": "<p>Yes, but most of the time it's part of the problem i.e. using one type of data and predicting another type. Organizers either want you to do domain adaptation (computer vision) or build a stable model (time series). I can't imagine either of those scenarios to be expected here.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2687945,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-08T21:59:55.723000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2688188,
              "author_name": "xavier fernando",
              "author_url": "",
              "post_date": "2024-03-09T04:48:09.860000",
              "content": "<p>So what this there in public LB 35 % data distribution same as the stage 2 training data so public LB good. If the private 65% contains more 1 - 10 annotators data.  Then there is chance of shake up.. Is my understanding is correct? </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2688016,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2024-03-08T23:42:29.330000",
          "content": "<p>This is interesting, in train about 50% samples are with exactly 3 annotators.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2689044,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-03-09T16:05:53.830000",
      "content": "<p>Here is a Cumulative Distribution of Total Evaluators in <em>eeg_id-unique-train_df</em> for reference.</p>\n<pre><code>sorted_counts = df_train[].value_counts().sort_index()\ncumulative_dist = sorted_counts.cumsum() / sorted_counts.()\nplt.plot(cumulative_dist.index, cumulative_dist, marker=, linestyle=)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F5e2f00904e50cfc6e303eea31fddb7a1%2Foutput.png?generation=1710000306691110&amp;alt=media\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2687252,
      "author_name": "Kirill Chemrov",
      "author_url": "",
      "post_date": "2024-03-08T11:50:54.420000",
      "content": "<p>The public to private ratio is ~1:2. <br>\nThe ratio of small and large number of experts on train data is ~1:2. <br>\nThink about it… ))</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2687394,
          "author_name": "Yuri Sun",
          "author_url": "",
          "post_date": "2024-03-08T14:31:20.160000",
          "content": "<p>If that's true, we need two submissions, one for the public set (10~20 annotators), and one for the private set (3~10 annotators).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2687420,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-03-08T15:17:49.700000",
              "content": "<p>I think you can make a suitable integration for these two distributions</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2687555,
              "author_name": "xavier fernando",
              "author_url": "",
              "post_date": "2024-03-08T16:05:15.140000",
              "content": "<p>so if we ensample stage 1 + stage 2 models combined will avoid the 10~20 annotators and 3~10 annotators different distributions issue.<br>\nBut i am currently submitting only stage 2. so chances of shake is lot form my understanding. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2687408,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-03-08T14:56:04.123000",
          "content": "<p>If the organizers made such a split, it would be extremely unfair towards the competitors (us).<br>\nAnd also make the competition a bit meaningless. <br>\nAs Andrew Ng said in one of his courses, validation should be approximately from the same distribution of the test/final target (not exact words, but close enough, I think)<br>\nOtherwise, what are we even doing other than mindlessly guessing? Lol.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2687499,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-08T15:48:38.587000",
              "content": "<p>Totally agree. It doesn't make sense to do such thing and very uncharacteristic of Kaggle. When public and private test sets are different in competitions, they always tell us and want us to build robust models in competitions like hubmap or ubc ocean. Doing the same thing here doesn't make any sense and fails the objective of this competition.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2687532,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-03-08T15:58:20.473000",
              "content": "<p>Eactly, also in the Ribonanza competition there was a case that the private test had a different distribution from public/train, but it was OK since they tolds as it is the case and how the distribution differs, so we could plan our validation scheme accordingly. <br>\nBut if they just put random distribution in the private test, that is different both from the train distribution and the public test distribution, and does not inform us on the matter, it is just…well…idk, data sciencve done wrong?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2687585,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-08T16:14:28.407000",
              "content": "<p>and throw away 50k for no value whatsoever</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2687948,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2024-03-08T22:03:02.230000",
              "content": "<p>I totally agree with this but in reality I see most comps have different distributions and different shakeup. CV normally wins</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2687976,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-03-08T22:37:54.163000",
      "content": "<p>The fact that cv and lb scores are so different make me think there could be some shakeup. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2686953,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-03-08T07:11:55.267000",
      "content": "<p>I'm not expecting any shake up.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2686988,
          "author_name": "Donghui Zhang",
          "author_url": "",
          "post_date": "2024-03-08T07:30:29.943000",
          "content": "<p>Hope nothing shakes your position😀</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2687038,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-03-08T08:08:56.907000",
              "content": "<p>We are hoping to reach 1st place 🙏</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2688512,
      "author_name": "Donghui Zhang",
      "author_url": "",
      "post_date": "2024-03-09T09:55:01.607000",
      "content": "<p>In fact, the reason why I feel this way is because in the two months before the game, most teams were stuck at 0.34. Even though the offline cv was optimized, the LB did not change until the second stage method emerged. The situation has improved, but for the front-row team, this does not affect them:)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2688517,
          "author_name": "Donghui Zhang",
          "author_url": "",
          "post_date": "2024-03-09T09:58:01.830000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479776#2671166\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/479776#2671166</a>  That's why I think what he said makes sense</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2688526,
          "author_name": "xavier fernando",
          "author_url": "",
          "post_date": "2024-03-09T10:06:25.217000",
          "content": "<p>You mean people in top 20 LB .. Not used stage wise training.. So the shake up in  private LB will be less for the front row people… Is my understanding is correct</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2688551,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-03-09T10:11:45.287000",
              "content": "<p>They should probably realize that the occurrence of this situation will make the model more robust. As for whether to use the second stage, I can’t judge.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2688484,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2024-03-09T09:14:12.643000",
      "content": "<p>considering <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard\" target=\"_blank\">SenNet + HOA - Hacking the Human Vasculature in 3D</a> anything is possible.</p>\n<p>It is not just number of annotators, but quality of spectrograms - differences between competition provided and generated from EEG raw data have been noted. For the latter no information has been provided on preprocessing methodology which could be different based on source of original data.<br>\nIt is not specified if source of patients are the same facilities, equipment the same, etc.  And distribution of patterns represented could vary, patients should be different for public test from private test.</p>\n<p>Usually hosts/sponsors are looking for robust solutions or to verify what is already known to work best, so time will tell…<br>\nGood luck to everyone!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2687343,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-08T13:38:21.087000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2687170": "From the competition data page:\n>Note that the test samples had between 3 and 20 annotators.\n\nIf we train our models on subsets which are more focused on votes > 10 (skipping almost half of the test data). \n\nHow can we not expect a shakeup?",
    "2686945": "Currently, the two-stage method is popular in public methods. I think the main reason why this method is effective is because it is found that LB data tends to a certain distribution, but this distribution is far from the distribution of the training set. Private data sets may not and the characteristics presented by LB data, so I think we still need to be wary of this method to avoid overfitting LB data and causing shake.",
    "2689044": "Here is a Cumulative Distribution of Total Evaluators in *eeg_id-unique-train_df* for reference.\n```python\nsorted_counts = df_train[\"total_evaluators\"].value_counts().sort_index()\ncumulative_dist = sorted_counts.cumsum() / sorted_counts.sum()\nplt.plot(cumulative_dist.index, cumulative_dist, marker='o', linestyle='--')\n```\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F5e2f00904e50cfc6e303eea31fddb7a1%2Foutput.png?generation=1710000306691110&alt=media)",
    "2687252": "The public to private ratio is ~1:2. \nThe ratio of small and large number of experts on train data is ~1:2. \nThink about it... ))",
    "2687976": "The fact that cv and lb scores are so different make me think there could be some shakeup. ",
    "2686953": "I'm not expecting any shake up.",
    "2688512": "In fact, the reason why I feel this way is because in the two months before the game, most teams were stuck at 0.34. Even though the offline cv was optimized, the LB did not change until the second stage method emerged. The situation has improved, but for the front-row team, this does not affect them:)",
    "2688484": "considering [SenNet + HOA - Hacking the Human Vasculature in 3D](https://www.kaggle.com/competitions/blood-vessel-segmentation/leaderboard) anything is possible.\n\nIt is not just number of annotators, but quality of spectrograms - differences between competition provided and generated from EEG raw data have been noted. For the latter no information has been provided on preprocessing methodology which could be different based on source of original data.\nIt is not specified if source of patients are the same facilities, equipment the same, etc.  And distribution of patterns represented could vary, patients should be different for public test from private test.\n\nUsually hosts/sponsors are looking for robust solutions or to verify what is already known to work best, so time will tell...\nGood luck to everyone!",
    "2687343": ""
  }
}