{
  "id": 492190,
  "title": "High Votes Distribution Wins!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492190",
  "author_name": "",
  "post_date": "2024-04-09T00:15:07.291288200Z",
  "votes": 17,
  "comment_count": 13,
  "views": 0,
  "content": "<p>It seems that in the end the distribution for the final private LB was the same distribution of high votes (or similar) as the public LB. Sadly for my team, we bet against this strategy, assuming there would be a low number of votes for part of the distribution. We had 2 strategies we wanted to select for this causing us not to select our 0.35 private LB solution. Very interested to see why so many seemed so confident this would be the case in this competition and all of the cool strategies many at the top were able to implement! Congrats to all!</p>",
  "messages": [
    {
      "id": "2742495",
      "postDate": "04/09/2024 00:15:07",
      "content": "<p>It seems that in the end the distribution for the final private LB was the same distribution of high votes (or similar) as the public LB. Sadly for my team, we bet against this strategy, assuming there would be a low number of votes for part of the distribution. We had 2 strategies we wanted to select for this causing us not to select our 0.35 private LB solution. Very interested to see why so many seemed so confident this would be the case in this competition and all of the cool strategies many at the top were able to implement! Congrats to all!</p>",
      "rawMarkdown": "It seems that in the end the distribution for the final private LB was the same distribution of high votes (or similar) as the public LB. Sadly for my team, we bet against this strategy, assuming there would be a low number of votes for part of the distribution. We had 2 strategies we wanted to select for this causing us not to select our 0.35 private LB solution. Very interested to see why so many seemed so confident this would be the case in this competition and all of the cool strategies many at the top were able to implement! Congrats to all!",
      "votes": null
    },
    {
      "id": "2742500",
      "postDate": "04/09/2024 00:19:15",
      "content": "<p>35% and 65% should have same distribution…, that's why we choose to belive in cv of votes&gt;=10.</p>",
      "rawMarkdown": "35% and 65% should have same distribution..., that's why we choose to belive in cv of votes>=10.",
      "votes": null
    },
    {
      "id": "2742558",
      "postDate": "04/09/2024 00:55:05",
      "content": "<p>I guess I mean comparing the training data to the Public LB there were 2 different groups high votes and low votes. In the public LB much of the data was high votes data which could be inferred by comparing the CV to the LB.  So we had a lot of discussion about if the private LB would sway back towards the training data with a mix of high and low votes or if it would remain similar to public Lb. Also congrats on your placement! I can’t wait to read more about your efforts on this!</p>",
      "rawMarkdown": "I guess I mean comparing the training data to the Public LB there were 2 different groups high votes and low votes. In the public LB much of the data was high votes data which could be inferred by comparing the CV to the LB.  So we had a lot of discussion about if the private LB would sway back towards the training data with a mix of high and low votes or if it would remain similar to public Lb. Also congrats on your placement! I can’t wait to read more about your efforts on this!",
      "votes": null
    },
    {
      "id": "2742600",
      "postDate": "04/09/2024 01:29:57",
      "content": "<p>Ya we had an inkling when discussing last week that stage 2 is what most are going for and will work in lb ,started quite late  . Though second choice was our STAGE 1 model we just slogged to improve stage 2 and our last sub/selection was like best private and public  . </p>",
      "rawMarkdown": "Ya we had an inkling when discussing last week that stage 2 is what most are going for and will work in lb ,started quite late  . Though second choice was our STAGE 1 model we just slogged to improve stage 2 and our last sub/selection was like best private and public  .",
      "votes": null
    },
    {
      "id": "2742677",
      "postDate": "04/09/2024 02:54:11",
      "content": "<p>I wrote discussion about test data.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243</a></p>\n<p>In the host's paper, they define labels with vote&gt;=10 as high quality labels and prepare test data with high quality labels.</p>",
      "rawMarkdown": "I wrote discussion about test data.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243\n\nIn the host's paper, they define labels with vote>=10 as high quality labels and prepare test data with high quality labels.",
      "votes": null
    },
    {
      "id": "2742685",
      "postDate": "04/09/2024 03:00:41",
      "content": "<p>Wasn't this tested early in the comp? Someone submitted both distributions and high votes one had a lower score. That was a key insight to me, and I thank who did it.</p>",
      "rawMarkdown": "Wasn't this tested early in the comp? Someone submitted both distributions and high votes one had a lower score. That was a key insight to me, and I thank who did it.",
      "votes": null
    },
    {
      "id": "2742686",
      "postDate": "04/09/2024 03:01:23",
      "content": "<p>Yep that makes perfect sense. And conceptually it does too but because the training had other data we fell for the trap it seems </p>",
      "rawMarkdown": "Yep that makes perfect sense. And conceptually it does too but because the training had other data we fell for the trap it seems",
      "votes": null
    },
    {
      "id": "2742706",
      "postDate": "04/09/2024 03:18:38",
      "content": "<p>Intuition for us was that it would make more sense for host to want a model which predicts classes more like ensemble of 10+ experts rather than just 3 experts.</p>",
      "rawMarkdown": "Intuition for us was that it would make more sense for host to want a model which predicts classes more like ensemble of 10+ experts rather than just 3 experts.",
      "votes": null
    },
    {
      "id": "2742743",
      "postDate": "04/09/2024 03:53:17",
      "content": "<p>I also wrote my take here. <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262</a></p>",
      "rawMarkdown": "I also wrote my take here. https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262",
      "votes": null
    },
    {
      "id": "2742868",
      "postDate": "04/09/2024 05:52:28",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> was the author of this idea</p>",
      "rawMarkdown": "pcjimmmy was the author of this idea",
      "votes": null
    },
    {
      "id": "2743312",
      "postDate": "04/09/2024 11:19:16",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> your post was one of the key insights for me. Let me add you in my writeup.</p>\n<p><a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> Thank you for the reference.</p>",
      "rawMarkdown": "pcjimmmy your post was one of the key insights for me. Let me add you in my writeup.\n\n@medali1992 Thank you for the reference.",
      "votes": null
    },
    {
      "id": "2743317",
      "postDate": "04/09/2024 11:31:48",
      "content": "<p>Yes this was tested on the public LB, but often times the public LB and private have different distributions.</p>",
      "rawMarkdown": "Yes this was tested on the public LB, but often times the public LB and private have different distributions.",
      "votes": null
    },
    {
      "id": "2743341",
      "postDate": "04/09/2024 11:49:29",
      "content": "<p>Fair point. I assumed private was similar to public.</p>",
      "rawMarkdown": "Fair point. I assumed private was similar to public.",
      "votes": null
    },
    {
      "id": "2744144",
      "postDate": "04/09/2024 19:19:30",
      "content": "<p>Yes. I think this is to be expected. When we read the host's research paper, they explain how train data is all vote count and test data is vote count &gt;= 10.</p>",
      "rawMarkdown": "Yes. I think this is to be expected. When we read the host's research paper, they explain how train data is all vote count and test data is vote count >= 10.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742500,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "04/09/2024 00:19:15",
      "content": "<p>35% and 65% should have same distribution…, that's why we choose to belive in cv of votes&gt;=10.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742558,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "04/09/2024 00:55:05",
          "content": "<p>I guess I mean comparing the training data to the Public LB there were 2 different groups high votes and low votes. In the public LB much of the data was high votes data which could be inferred by comparing the CV to the LB.  So we had a lot of discussion about if the private LB would sway back towards the training data with a mix of high and low votes or if it would remain similar to public Lb. Also congrats on your placement! I can’t wait to read more about your efforts on this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742600,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "04/09/2024 01:29:57",
      "content": "<p>Ya we had an inkling when discussing last week that stage 2 is what most are going for and will work in lb ,started quite late  . Though second choice was our STAGE 1 model we just slogged to improve stage 2 and our last sub/selection was like best private and public  . </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2742677,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "04/09/2024 02:54:11",
      "content": "<p>I wrote discussion about test data.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243</a></p>\n<p>In the host's paper, they define labels with vote&gt;=10 as high quality labels and prepare test data with high quality labels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742686,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "04/09/2024 03:01:23",
          "content": "<p>Yep that makes perfect sense. And conceptually it does too but because the training had other data we fell for the trap it seems </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2742743,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "04/09/2024 03:53:17",
          "content": "<p>I also wrote my take here. <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742685,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/09/2024 03:00:41",
      "content": "<p>Wasn't this tested early in the comp? Someone submitted both distributions and high votes one had a lower score. That was a key insight to me, and I thank who did it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742868,
          "author_name": "medali1992",
          "author_url": "",
          "post_date": "04/09/2024 05:52:28",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> was the author of this idea</p>",
          "votes": null,
          "replies": [
            {
              "id": 2743312,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "04/09/2024 11:19:16",
              "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> your post was one of the key insights for me. Let me add you in my writeup.</p>\n<p><a href=\"https://www.kaggle.com/medali1992\" target=\"_blank\">@medali1992</a> Thank you for the reference.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2743317,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "04/09/2024 11:31:48",
          "content": "<p>Yes this was tested on the public LB, but often times the public LB and private have different distributions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2743341,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "04/09/2024 11:49:29",
              "content": "<p>Fair point. I assumed private was similar to public.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742706,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "04/09/2024 03:18:38",
      "content": "<p>Intuition for us was that it would make more sense for host to want a model which predicts classes more like ensemble of 10+ experts rather than just 3 experts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2744144,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/09/2024 19:19:30",
      "content": "<p>Yes. I think this is to be expected. When we read the host's research paper, they explain how train data is all vote count and test data is vote count &gt;= 10.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742495": "It seems that in the end the distribution for the final private LB was the same distribution of high votes (or similar) as the public LB. Sadly for my team, we bet against this strategy, assuming there would be a low number of votes for part of the distribution. We had 2 strategies we wanted to select for this causing us not to select our 0.35 private LB solution. Very interested to see why so many seemed so confident this would be the case in this competition and all of the cool strategies many at the top were able to implement! Congrats to all!",
    "2742500": "35% and 65% should have same distribution..., that's why we choose to belive in cv of votes>=10.",
    "2742558": "I guess I mean comparing the training data to the Public LB there were 2 different groups high votes and low votes. In the public LB much of the data was high votes data which could be inferred by comparing the CV to the LB.  So we had a lot of discussion about if the private LB would sway back towards the training data with a mix of high and low votes or if it would remain similar to public Lb. Also congrats on your placement! I can’t wait to read more about your efforts on this!",
    "2742600": "Ya we had an inkling when discussing last week that stage 2 is what most are going for and will work in lb ,started quite late  . Though second choice was our STAGE 1 model we just slogged to improve stage 2 and our last sub/selection was like best private and public  .",
    "2742677": "I wrote discussion about test data.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492243\n\nIn the host's paper, they define labels with vote>=10 as high quality labels and prepare test data with high quality labels.",
    "2742685": "Wasn't this tested early in the comp? Someone submitted both distributions and high votes one had a lower score. That was a key insight to me, and I thank who did it.",
    "2742686": "Yep that makes perfect sense. And conceptually it does too but because the training had other data we fell for the trap it seems",
    "2742706": "Intuition for us was that it would make more sense for host to want a model which predicts classes more like ensemble of 10+ experts rather than just 3 experts.",
    "2742743": "I also wrote my take here. https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492262",
    "2742868": "pcjimmmy was the author of this idea",
    "2743312": "pcjimmmy your post was one of the key insights for me. Let me add you in my writeup.\n\n@medali1992 Thank you for the reference.",
    "2743317": "Yes this was tested on the public LB, but often times the public LB and private have different distributions.",
    "2743341": "Fair point. I assumed private was similar to public.",
    "2744144": "Yes. I think this is to be expected. When we read the host's research paper, they explain how train data is all vote count and test data is vote count >= 10."
  },
  "source": "meta"
}