{
  "id": 467163,
  "title": "Computing Vote Ratio's | Total Number Of Voters In Train Samples",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467163",
  "author_name": "",
  "post_date": "2024-01-11T10:44:28.340556400Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The labels provided in <code>train.csv</code> contain the number of votes for each category.<br>\nHowever, for the submission we need to predict the vote ratio between 0 and 1.<br>\nWhen a certain category has 3 votes, it is unclear if this means 3/3 votes(1.0 ratio) or, for example, 3/6 votes (0.5 ratio).</p>\n<p>How can we obtain the total number of voters in the train samples to allow us to compute the vote ratio?</p>",
  "messages": [
    {
      "id": "2596807",
      "postDate": "01/11/2024 10:44:28",
      "content": "<p>The labels provided in <code>train.csv</code> contain the number of votes for each category.<br>\nHowever, for the submission we need to predict the vote ratio between 0 and 1.<br>\nWhen a certain category has 3 votes, it is unclear if this means 3/3 votes(1.0 ratio) or, for example, 3/6 votes (0.5 ratio).</p>\n<p>How can we obtain the total number of voters in the train samples to allow us to compute the vote ratio?</p>",
      "rawMarkdown": "The labels provided in `train.csv` contain the number of votes for each category.\nHowever, for the submission we need to predict the vote ratio between 0 and 1.\nWhen a certain category has 3 votes, it is unclear if this means 3/3 votes(1.0 ratio) or, for example, 3/6 votes (0.5 ratio).\n\nHow can we obtain the total number of voters in the train samples to allow us to compute the vote ratio?",
      "votes": null
    },
    {
      "id": "2596837",
      "postDate": "01/11/2024 11:03:28",
      "content": "<p>I think you can approach this as a classification problem. You can assign labels with argmax on vote columns and train on those labels. You can convert model outputs to probabilities/vote ratios with a pdf on test time.</p>",
      "rawMarkdown": "I think you can approach this as a classification problem. You can assign labels with argmax on vote columns and train on those labels. You can convert model outputs to probabilities/vote ratios with a pdf on test time.",
      "votes": null
    },
    {
      "id": "2596885",
      "postDate": "01/11/2024 11:45:09",
      "content": "<p>Using the <code>argmax</code> on the labels is an idea!</p>\n<p>What puzzles me is that certain samples have spread out votes, with for example 4 on grda and 8 on other, <code>eeg_id=1347137760</code>.<br>\nA ratio of 0.333 and 0.666 as labels would be a possibility as well.<br>\nRatios as labels would also align more with the validation metric as well.<br>\nThe labelling approach will probably be a main challenge in this competition.</p>",
      "rawMarkdown": "Using the `argmax` on the labels is an idea!\n\nWhat puzzles me is that certain samples have spread out votes, with for example 4 on grda and 8 on other, `eeg_id=1347137760`.\nA ratio of 0.333 and 0.666 as labels would be a possibility as well.\nRatios as labels would also align more with the validation metric as well.\nThe labelling approach will probably be a main challenge in this competition.",
      "votes": null
    },
    {
      "id": "2596888",
      "postDate": "01/11/2024 11:48:35",
      "content": "<p>Annotator disagreements are actually very common in medical domain. One way to tackle that is dropping or giving less weights to samples with high entropy labels.</p>",
      "rawMarkdown": "Annotator disagreements are actually very common in medical domain. One way to tackle that is dropping or giving less weights to samples with high entropy labels.",
      "votes": null
    },
    {
      "id": "2618815",
      "postDate": "01/25/2024 02:38:24",
      "content": "<p>To see the level of disagreement visually, I have a scatter plot of vote entropy vs the number of votes in the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470645\" target=\"_blank\">k-means clusters for classification</a> discussion post. Often an entropy equivalent to spread evenly over 2 or 3 HBAs.</p>",
      "rawMarkdown": "To see the level of disagreement visually, I have a scatter plot of vote entropy vs the number of votes in the [k-means clusters for classification](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470645) discussion post. Often an entropy equivalent to spread evenly over 2 or 3 HBAs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2596837,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/11/2024 11:03:28",
      "content": "<p>I think you can approach this as a classification problem. You can assign labels with argmax on vote columns and train on those labels. You can convert model outputs to probabilities/vote ratios with a pdf on test time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596885,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "01/11/2024 11:45:09",
          "content": "<p>Using the <code>argmax</code> on the labels is an idea!</p>\n<p>What puzzles me is that certain samples have spread out votes, with for example 4 on grda and 8 on other, <code>eeg_id=1347137760</code>.<br>\nA ratio of 0.333 and 0.666 as labels would be a possibility as well.<br>\nRatios as labels would also align more with the validation metric as well.<br>\nThe labelling approach will probably be a main challenge in this competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2596888,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "01/11/2024 11:48:35",
              "content": "<p>Annotator disagreements are actually very common in medical domain. One way to tackle that is dropping or giving less weights to samples with high entropy labels.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2618815,
                  "author_name": "dan3dewey",
                  "author_url": "",
                  "post_date": "01/25/2024 02:38:24",
                  "content": "<p>To see the level of disagreement visually, I have a scatter plot of vote entropy vs the number of votes in the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470645\" target=\"_blank\">k-means clusters for classification</a> discussion post. Often an entropy equivalent to spread evenly over 2 or 3 HBAs.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2596807": "The labels provided in `train.csv` contain the number of votes for each category.\nHowever, for the submission we need to predict the vote ratio between 0 and 1.\nWhen a certain category has 3 votes, it is unclear if this means 3/3 votes(1.0 ratio) or, for example, 3/6 votes (0.5 ratio).\n\nHow can we obtain the total number of voters in the train samples to allow us to compute the vote ratio?",
    "2596837": "I think you can approach this as a classification problem. You can assign labels with argmax on vote columns and train on those labels. You can convert model outputs to probabilities/vote ratios with a pdf on test time.",
    "2596885": "Using the `argmax` on the labels is an idea!\n\nWhat puzzles me is that certain samples have spread out votes, with for example 4 on grda and 8 on other, `eeg_id=1347137760`.\nA ratio of 0.333 and 0.666 as labels would be a possibility as well.\nRatios as labels would also align more with the validation metric as well.\nThe labelling approach will probably be a main challenge in this competition.",
    "2596888": "Annotator disagreements are actually very common in medical domain. One way to tackle that is dropping or giving less weights to samples with high entropy labels.",
    "2618815": "To see the level of disagreement visually, I have a scatter plot of vote entropy vs the number of votes in the [k-means clusters for classification](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470645) discussion post. Often an entropy equivalent to spread evenly over 2 or 3 HBAs."
  },
  "source": "meta"
}