{
  "id": 479786,
  "title": "The Hunt For Weighted Voting Algorithm",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479786",
  "author_name": "",
  "post_date": "2024-02-26T01:07:30.060759500Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So to test this, I am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> notebook: <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43</a></p>\n<p>So basically the intuition is the more votes, the more weight it should have on training. There is no way that 1 voter is as accurate as the some 10000 voters that some of this data has. The problem is to determine how much more weight these greater voting sample size data should have. I don't think I have found anything on this other than here: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666</a></p>\n<p>I encourage people who have found anything on this to comment on their findings here, because I think this may help models, and its kinda fun.</p>\n<p>So anyways things that I have found:</p>\n<p>So the cv score for the original (no weight) is: <strong>0.59015</strong> (std: 0.0296)</p>\n<p>With <strong>log_128(voteCount + 1)</strong> weighting it drops to: <strong>0.57349</strong> (std: 0.0191) (quite stable)</p>\n<p>With <strong>ln(log_128(voteCount))/3+1</strong>: <strong>0.57690</strong> (std: 0.0301)</p>\n<p><a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> idea of:<br>\n<strong>alpha = 1 / (smoothing+np.sqrt(voteCount))</strong><br>\n<strong>(1-alpha) + alpha/6</strong><br>\nworked quite well with a cv score: <strong>0.57803</strong> (std: 0.0234) (smoothing = 1 for this one with going to 2 and 0.5 getting worse)</p>\n<p>Um whatever I did here 1/(1+e^(-log_10(weights/((1.96<em>0.01)/0.1)^2)))</em>1/0.9 (intuition sigmoid curve maybe clean it up): <strong>0.58436</strong> (std: 0.0286)</p>",
  "messages": [
    {
      "id": "2668897",
      "postDate": "02/26/2024 01:07:30",
      "content": "<p>So to test this, I am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> notebook: <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43</a></p>\n<p>So basically the intuition is the more votes, the more weight it should have on training. There is no way that 1 voter is as accurate as the some 10000 voters that some of this data has. The problem is to determine how much more weight these greater voting sample size data should have. I don't think I have found anything on this other than here: <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666</a></p>\n<p>I encourage people who have found anything on this to comment on their findings here, because I think this may help models, and its kinda fun.</p>\n<p>So anyways things that I have found:</p>\n<p>So the cv score for the original (no weight) is: <strong>0.59015</strong> (std: 0.0296)</p>\n<p>With <strong>log_128(voteCount + 1)</strong> weighting it drops to: <strong>0.57349</strong> (std: 0.0191) (quite stable)</p>\n<p>With <strong>ln(log_128(voteCount))/3+1</strong>: <strong>0.57690</strong> (std: 0.0301)</p>\n<p><a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> idea of:<br>\n<strong>alpha = 1 / (smoothing+np.sqrt(voteCount))</strong><br>\n<strong>(1-alpha) + alpha/6</strong><br>\nworked quite well with a cv score: <strong>0.57803</strong> (std: 0.0234) (smoothing = 1 for this one with going to 2 and 0.5 getting worse)</p>\n<p>Um whatever I did here 1/(1+e^(-log_10(weights/((1.96<em>0.01)/0.1)^2)))</em>1/0.9 (intuition sigmoid curve maybe clean it up): <strong>0.58436</strong> (std: 0.0286)</p>",
      "rawMarkdown": "So to test this, I am using @cdeotte notebook: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n\nSo basically the intuition is the more votes, the more weight it should have on training. There is no way that 1 voter is as accurate as the some 10000 voters that some of this data has. The problem is to determine how much more weight these greater voting sample size data should have. I don't think I have found anything on this other than here: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\n\nI encourage people who have found anything on this to comment on their findings here, because I think this may help models, and its kinda fun.\n\nSo anyways things that I have found:\n\nSo the cv score for the original (no weight) is: **0.59015** (std: 0.0296)\n\nWith **log_128(voteCount + 1)** weighting it drops to: **0.57349** (std: 0.0191) (quite stable)\n\nWith **ln(log_128(voteCount))/3+1**: **0.57690** (std: 0.0301)\n\n@caelhasse idea of:\n**alpha = 1 / (smoothing+np.sqrt(voteCount))**\n**(1-alpha) + alpha/6**\nworked quite well with a cv score: **0.57803** (std: 0.0234) (smoothing = 1 for this one with going to 2 and 0.5 getting worse)\n\nUm whatever I did here 1/(1+e^(-log_10(weights/((1.96*0.01)/0.1)^2)))*1/0.9 (intuition sigmoid curve maybe clean it up): **0.58436** (std: 0.0286)",
      "votes": null
    },
    {
      "id": "2668923",
      "postDate": "02/26/2024 02:25:45",
      "content": "<p>The risk about your intuition is \"what did kaggle and the host do to build the test set\"</p>\n<p>My intuition is that the single vote was Dr. Lawrence Hirsch.  (or another expert of similiar skills).   You intuition would fit if a intern was assigned the job of rating some nice eeg's that were available to fill up the data set.</p>\n<p>So - the big question - does the test set have the same distribution of voters as the train?</p>",
      "rawMarkdown": "The risk about your intuition is \"what did kaggle and the host do to build the test set\"\n\nMy intuition is that the single vote was Dr. Lawrence Hirsch.  (or another expert of similiar skills).   You intuition would fit if a intern was assigned the job of rating some nice eeg's that were available to fill up the data set.\n\nSo - the big question - does the test set have the same distribution of voters as the train?",
      "votes": null
    },
    {
      "id": "2668961",
      "postDate": "02/26/2024 03:15:39",
      "content": "<p>Hmmm interesting, maybe that's why dropping cv does not really help the leaderboard. I still do think that more accurate voting should be weighted more, but I'm not sure. I'll probably just pick my best model and submit one with weighted training and another with unweighted (Since 2 can be selected). Still, as they say, always trust that cv score.</p>",
      "rawMarkdown": "Hmmm interesting, maybe that's why dropping cv does not really help the leaderboard. I still do think that more accurate voting should be weighted more, but I'm not sure. I'll probably just pick my best model and submit one with weighted training and another with unweighted (Since 2 can be selected). Still, as they say, always trust that cv score.",
      "votes": null
    },
    {
      "id": "2669560",
      "postDate": "02/26/2024 10:35:21",
      "content": "<p>You're right but that would be a real problem if the model should predict the true label of the patient. What has to predict is the probability distribution of many expert votations. Not?</p>",
      "rawMarkdown": "You're right but that would be a real problem if the model should predict the true label of the patient. What has to predict is the probability distribution of many expert votations. Not?",
      "votes": null
    },
    {
      "id": "2670524",
      "postDate": "02/26/2024 23:29:30",
      "content": "<p>I had an idea related to this that I posted elsewhere:</p>\n<p>Consider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.</p>\n<p>I used the formula:</p>\n<p>alpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes<br>\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6</p>\n<p>where smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.</p>",
      "rawMarkdown": "I had an idea related to this that I posted elsewhere:\n\nConsider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.\n\nI used the formula:\n\nalpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6\n\nwhere smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.",
      "votes": null
    },
    {
      "id": "2671601",
      "postDate": "02/27/2024 15:52:20",
      "content": "<p>Thanks for sharing, keep up the good work, in my opinion, the more votes you have, the higher the quality of the data, that is, the more experts were able to give a verdict on that data</p>",
      "rawMarkdown": "Thanks for sharing, keep up the good work, in my opinion, the more votes you have, the higher the quality of the data, that is, the more experts were able to give a verdict on that data",
      "votes": null
    },
    {
      "id": "2671958",
      "postDate": "02/27/2024 19:40:43",
      "content": "<p>Very interesting, will try it now. I just graphed it in desmos and I definitely like how higher values stop scaling as much in weight.<br>\nEdit: I’m using a smoothing of 2 and just graphed the histogram and it has a very nice distribution (similar to the log one but in my opinion better). Going to try smoothing of 1 and 0 as well.</p>",
      "rawMarkdown": "Very interesting, will try it now. I just graphed it in desmos and I definitely like how higher values stop scaling as much in weight.\nEdit: I’m using a smoothing of 2 and just graphed the histogram and it has a very nice distribution (similar to the log one but in my opinion better). Going to try smoothing of 1 and 0 as well.",
      "votes": null
    },
    {
      "id": "2672145",
      "postDate": "02/28/2024 00:13:14",
      "content": "<p>Your equation worked quite well with a smoothing value of 1, ending with a cv score: 0.57803 (std: 0.0234)</p>",
      "rawMarkdown": "Your equation worked quite well with a smoothing value of 1, ending with a cv score: 0.57803 (std: 0.0234)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2668923,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/26/2024 02:25:45",
      "content": "<p>The risk about your intuition is \"what did kaggle and the host do to build the test set\"</p>\n<p>My intuition is that the single vote was Dr. Lawrence Hirsch.  (or another expert of similiar skills).   You intuition would fit if a intern was assigned the job of rating some nice eeg's that were available to fill up the data set.</p>\n<p>So - the big question - does the test set have the same distribution of voters as the train?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2668961,
          "author_name": "kawaiicoderuwu",
          "author_url": "",
          "post_date": "02/26/2024 03:15:39",
          "content": "<p>Hmmm interesting, maybe that's why dropping cv does not really help the leaderboard. I still do think that more accurate voting should be weighted more, but I'm not sure. I'll probably just pick my best model and submit one with weighted training and another with unweighted (Since 2 can be selected). Still, as they say, always trust that cv score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2669560,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "02/26/2024 10:35:21",
          "content": "<p>You're right but that would be a real problem if the model should predict the true label of the patient. What has to predict is the probability distribution of many expert votations. Not?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2670524,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "02/26/2024 23:29:30",
      "content": "<p>I had an idea related to this that I posted elsewhere:</p>\n<p>Consider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.</p>\n<p>I used the formula:</p>\n<p>alpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes<br>\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6</p>\n<p>where smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2671958,
          "author_name": "kawaiicoderuwu",
          "author_url": "",
          "post_date": "02/27/2024 19:40:43",
          "content": "<p>Very interesting, will try it now. I just graphed it in desmos and I definitely like how higher values stop scaling as much in weight.<br>\nEdit: I’m using a smoothing of 2 and just graphed the histogram and it has a very nice distribution (similar to the log one but in my opinion better). Going to try smoothing of 1 and 0 as well.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2672145,
              "author_name": "kawaiicoderuwu",
              "author_url": "",
              "post_date": "02/28/2024 00:13:14",
              "content": "<p>Your equation worked quite well with a smoothing value of 1, ending with a cv score: 0.57803 (std: 0.0234)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2671601,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "02/27/2024 15:52:20",
      "content": "<p>Thanks for sharing, keep up the good work, in my opinion, the more votes you have, the higher the quality of the data, that is, the more experts were able to give a verdict on that data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2668897": "So to test this, I am using @cdeotte notebook: https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\n\nSo basically the intuition is the more votes, the more weight it should have on training. There is no way that 1 voter is as accurate as the some 10000 voters that some of this data has. The problem is to determine how much more weight these greater voting sample size data should have. I don't think I have found anything on this other than here: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\n\nI encourage people who have found anything on this to comment on their findings here, because I think this may help models, and its kinda fun.\n\nSo anyways things that I have found:\n\nSo the cv score for the original (no weight) is: **0.59015** (std: 0.0296)\n\nWith **log_128(voteCount + 1)** weighting it drops to: **0.57349** (std: 0.0191) (quite stable)\n\nWith **ln(log_128(voteCount))/3+1**: **0.57690** (std: 0.0301)\n\n@caelhasse idea of:\n**alpha = 1 / (smoothing+np.sqrt(voteCount))**\n**(1-alpha) + alpha/6**\nworked quite well with a cv score: **0.57803** (std: 0.0234) (smoothing = 1 for this one with going to 2 and 0.5 getting worse)\n\nUm whatever I did here 1/(1+e^(-log_10(weights/((1.96*0.01)/0.1)^2)))*1/0.9 (intuition sigmoid curve maybe clean it up): **0.58436** (std: 0.0286)",
    "2668923": "The risk about your intuition is \"what did kaggle and the host do to build the test set\"\n\nMy intuition is that the single vote was Dr. Lawrence Hirsch.  (or another expert of similiar skills).   You intuition would fit if a intern was assigned the job of rating some nice eeg's that were available to fill up the data set.\n\nSo - the big question - does the test set have the same distribution of voters as the train?",
    "2668961": "Hmmm interesting, maybe that's why dropping cv does not really help the leaderboard. I still do think that more accurate voting should be weighted more, but I'm not sure. I'll probably just pick my best model and submit one with weighted training and another with unweighted (Since 2 can be selected). Still, as they say, always trust that cv score.",
    "2669560": "You're right but that would be a real problem if the model should predict the true label of the patient. What has to predict is the probability distribution of many expert votations. Not?",
    "2670524": "I had an idea related to this that I posted elsewhere:\n\nConsider the normalised votes as estimates of an underlying probability for a physician to vote for that class. The less votes the more uncertainty in that estimate; even though the estimate may be the most probable value it may not be the median value. When trying to predict out of sample normalised votes, it may be better to make a more conservative prediction.\n\nI used the formula:\n\nalpha = 1/(smoothing + np.sqrt(y_sum)) #y_sum being the number of votes\ntrain[SMOOTH_TARGETS] = (1-alpha)*train[TARGETS] + alpha/6\n\nwhere smoothing is some parameter (higher == less smoothing). In my tests I applied the label smoothing to the training labels and not the validation labels. So far I haven't found any benefits but maybe in the right context or with improvements, someone could get something out of it.",
    "2671601": "Thanks for sharing, keep up the good work, in my opinion, the more votes you have, the higher the quality of the data, that is, the more experts were able to give a verdict on that data",
    "2671958": "Very interesting, will try it now. I just graphed it in desmos and I definitely like how higher values stop scaling as much in weight.\nEdit: I’m using a smoothing of 2 and just graphed the histogram and it has a very nice distribution (similar to the log one but in my opinion better). Going to try smoothing of 1 and 0 as well.",
    "2672145": "Your equation worked quite well with a smoothing value of 1, ending with a cv score: 0.57803 (std: 0.0234)"
  },
  "source": "meta"
}