{
  "id": 477554,
  "title": "A Simple Label Smoothing Scheme Based on the Number of Experts",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477554",
  "author_name": "",
  "post_date": "2024-02-16T16:43:15.401613700Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>An important topic of discussion in this competition has been how to account for the number of experts who have voted on any given example of EEG and Spectrogram data. <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">This helpful post</a> by <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a> introduced a two stage training process based on the apparent separate sources of training data we are given, discovered by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in this <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>. Here I propose a simple method for smoothing class labels based on the number of experts who have voted on a given example.</p>\n<p>Consider a ground truth label representing the probability of some set of classes of the form <code>[1, 0, 0, 0, 0, 0]</code>. Label smoothing penalizes the ground truth class by redistributing a portion the positive label across the other classes, preventing a trained model from becoming overly confident in its predictions and helping to regularize the network. For some small number <code>p</code> the smoothed label looks like <code>[1-p, p/(k-1), p/(k-1), p/(k-1), p/(k-1), p/(k-1) ]</code> where <code>k</code> is the number of classes. Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached. Therefore would like to choose a value of <code>p</code> such that there is a higher degree of smoothing when fewer experts are involved and less smoothing with more experts.</p>\n<p>A simple functional form we can start with is one which increases <code>p</code> when fewer experts have voted and decreases <code>p</code> when there are more experts. As a starting point I propose the following: <code>p = 1/(num_experts + 1)</code> where the <code>+ 1</code> in the denominator is included so that the positive ground truth label is not reduced to 0 when only one expert has voted. Now that our value of <code>p</code> is chosen we then penalize the label which has been decided on in the <code>expert_consensus</code> column of <code>train.csv</code> and distribute to the other classes. As an example consider a 'Seizure' consensus with the label <code>[1, 0, 0, 0, 0, 0]</code> and 3 expert votes. Our value of <code>p</code> is <code>1/(3+1) = 0.25</code> and our smoothed label becomes <code>[0.75, 0.05, 0.05, 0.05, 0.05, 0.05]</code>. </p>\n<p>I have implemented this in my own training and have observed improved regularization and better CV performance in 5-fold training grouped by <code>patient_id</code> in a 1D CNN+RNN model. Of course this is very simple and only serves to act as a starting point for further exploration. Other functional forms for <code>p</code> may improve performance further and more sophisticated methods for choosing how to redistribute the labels could be important (e.g. the case of an even split among experts). Please share your thoughts and ideas below and perhaps where you think this method may fail.</p>",
  "messages": [
    {
      "id": "2655028",
      "postDate": "02/16/2024 16:43:15",
      "content": "<p>An important topic of discussion in this competition has been how to account for the number of experts who have voted on any given example of EEG and Spectrogram data. <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">This helpful post</a> by <a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@seanbearden</a> introduced a two stage training process based on the apparent separate sources of training data we are given, discovered by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> in this <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">notebook</a>. Here I propose a simple method for smoothing class labels based on the number of experts who have voted on a given example.</p>\n<p>Consider a ground truth label representing the probability of some set of classes of the form <code>[1, 0, 0, 0, 0, 0]</code>. Label smoothing penalizes the ground truth class by redistributing a portion the positive label across the other classes, preventing a trained model from becoming overly confident in its predictions and helping to regularize the network. For some small number <code>p</code> the smoothed label looks like <code>[1-p, p/(k-1), p/(k-1), p/(k-1), p/(k-1), p/(k-1) ]</code> where <code>k</code> is the number of classes. Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached. Therefore would like to choose a value of <code>p</code> such that there is a higher degree of smoothing when fewer experts are involved and less smoothing with more experts.</p>\n<p>A simple functional form we can start with is one which increases <code>p</code> when fewer experts have voted and decreases <code>p</code> when there are more experts. As a starting point I propose the following: <code>p = 1/(num_experts + 1)</code> where the <code>+ 1</code> in the denominator is included so that the positive ground truth label is not reduced to 0 when only one expert has voted. Now that our value of <code>p</code> is chosen we then penalize the label which has been decided on in the <code>expert_consensus</code> column of <code>train.csv</code> and distribute to the other classes. As an example consider a 'Seizure' consensus with the label <code>[1, 0, 0, 0, 0, 0]</code> and 3 expert votes. Our value of <code>p</code> is <code>1/(3+1) = 0.25</code> and our smoothed label becomes <code>[0.75, 0.05, 0.05, 0.05, 0.05, 0.05]</code>. </p>\n<p>I have implemented this in my own training and have observed improved regularization and better CV performance in 5-fold training grouped by <code>patient_id</code> in a 1D CNN+RNN model. Of course this is very simple and only serves to act as a starting point for further exploration. Other functional forms for <code>p</code> may improve performance further and more sophisticated methods for choosing how to redistribute the labels could be important (e.g. the case of an even split among experts). Please share your thoughts and ideas below and perhaps where you think this method may fail.</p>",
      "rawMarkdown": "An important topic of discussion in this competition has been how to account for the number of experts who have voted on any given example of EEG and Spectrogram data. [This helpful post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135) by @seanbearden introduced a two stage training process based on the apparent separate sources of training data we are given, discovered by @pcjimmmy in this [notebook](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda). Here I propose a simple method for smoothing class labels based on the number of experts who have voted on a given example.\n\nConsider a ground truth label representing the probability of some set of classes of the form `[1, 0, 0, 0, 0, 0]`. Label smoothing penalizes the ground truth class by redistributing a portion the positive label across the other classes, preventing a trained model from becoming overly confident in its predictions and helping to regularize the network. For some small number `p` the smoothed label looks like `[1-p, p/(k-1), p/(k-1), p/(k-1), p/(k-1), p/(k-1) ]` where `k` is the number of classes. Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached. Therefore would like to choose a value of `p` such that there is a higher degree of smoothing when fewer experts are involved and less smoothing with more experts.\n\nA simple functional form we can start with is one which increases `p` when fewer experts have voted and decreases `p` when there are more experts. As a starting point I propose the following: `p = 1/(num_experts + 1)` where the `+ 1` in the denominator is included so that the positive ground truth label is not reduced to 0 when only one expert has voted. Now that our value of `p` is chosen we then penalize the label which has been decided on in the `expert_consensus` column of `train.csv` and distribute to the other classes. As an example consider a 'Seizure' consensus with the label `[1, 0, 0, 0, 0, 0]` and 3 expert votes. Our value of `p` is `1/(3+1) = 0.25` and our smoothed label becomes `[0.75, 0.05, 0.05, 0.05, 0.05, 0.05]`. \n\nI have implemented this in my own training and have observed improved regularization and better CV performance in 5-fold training grouped by `patient_id` in a 1D CNN+RNN model. Of course this is very simple and only serves to act as a starting point for further exploration. Other functional forms for `p` may improve performance further and more sophisticated methods for choosing how to redistribute the labels could be important (e.g. the case of an even split among experts). Please share your thoughts and ideas below and perhaps where you think this method may fail.",
      "votes": null
    },
    {
      "id": "2655076",
      "postDate": "02/16/2024 17:12:31",
      "content": "<blockquote>\n  <p>Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached.</p>\n</blockquote>\n<p>Unfortunately, this intuition is completelty wrong because different number of experts isn't picked at random across all samples from all patients. Actual labeling algorithm looks pretty cryptic according to <a href=\"https://bdsp.io/content/bdsp-sparcnet/1.1/\" target=\"_blank\">https://bdsp.io/content/bdsp-sparcnet/1.1/</a></p>\n<p>Moreover, in regular clinical practice more uncertain cases tend to involve more experts to diagnosis procedure.</p>",
      "rawMarkdown": ">Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached.\n\nUnfortunately, this intuition is completelty wrong because different number of experts isn't picked at random across all samples from all patients. Actual labeling algorithm looks pretty cryptic according to https://bdsp.io/content/bdsp-sparcnet/1.1/\n\nMoreover, in regular clinical practice more uncertain cases tend to involve more experts to diagnosis procedure.",
      "votes": null
    },
    {
      "id": "2655170",
      "postDate": "02/16/2024 17:49:26",
      "content": "<p>Thank you for the insight. So because experts aren't chosen at random for each patient there will be some biases those experts carry, thereby making the consensus of the label distributions similarly biased? Has it been determined that the Kaggle dataset is a subset of this SPaRCNet data?</p>",
      "rawMarkdown": "Thank you for the insight. So because experts aren't chosen at random for each patient there will be some biases those experts carry, thereby making the consensus of the label distributions similarly biased? Has it been determined that the Kaggle dataset is a subset of this SPaRCNet data?",
      "votes": null
    },
    {
      "id": "2655194",
      "postDate": "02/16/2024 17:59:44",
      "content": "<p>Possible biases in assessment are not clear for me. Good news: experts are independent, so they can't influence each other.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439</a> <code>SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.</code> (c) Kaggle Staff</p>",
      "rawMarkdown": "Possible biases in assessment are not clear for me. Good news: experts are independent, so they can't influence each other.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439 `SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.` (c) Kaggle Staff",
      "votes": null
    },
    {
      "id": "2655227",
      "postDate": "02/16/2024 18:50:05",
      "content": "<p>I have a very basic statistics base so apologize if is a terrible suggestion.<br>\nBut what with  distribute equally votes to each label to match the maximum of expert votes in a single label? Let's say the maximum voted label has 10 expert votes and an example labels has [1, 0, 0, 0, 0, 0], so turn it into [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6].<br>\nJust as your suggestion direction but without parameters.</p>",
      "rawMarkdown": "I have a very basic statistics base so apologize if is a terrible suggestion.\nBut what with  distribute equally votes to each label to match the maximum of expert votes in a single label? Let's say the maximum voted label has 10 expert votes and an example labels has [1, 0, 0, 0, 0, 0], so turn it into [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6].\nJust as your suggestion direction but without parameters.",
      "votes": null
    },
    {
      "id": "2655349",
      "postDate": "02/16/2024 19:53:32",
      "content": "<p>This is certainly something one can try. However, if I understand your suggestion correctly this will apply stronger smoothing to examples where there are more experts who have voted and weaker smoothing when fewer experts are involved. The objective of my proposed scheme is the opposite. In the initial post I had posited that when more experts vote on a particular example then we can have more \"confidence\" in the expert consensus reached, however this may be a bad intuition as pointed out by <a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a>.</p>\n<p>After some thought I believe a more correct interpretation is that when we have more experts voting on a particular example the resulting distribution is more <strong>representative</strong> of how other experts would vote since each expert is independent. However, there may be some detail I am missing in this interpretation. If we assume this interpretation to be the case then we still have the same goal of wanting stronger smoothing when fewer experts are involved because having only one expert vote is not necessarily a good representation of how other experts will vote. </p>\n<p>There are many unknowns in this analysis, such as; Are examples with only 1 vote trivial examples and so more experts aren't needed to correctly diagnose? In this case my interpretation above would be incorrect. Regardless, empirically my smoothing scheme has improved my own results and we may try other schemes to see what is best. I encourage you to try what you have suggested and see if it works, then let us know here.</p>",
      "rawMarkdown": "This is certainly something one can try. However, if I understand your suggestion correctly this will apply stronger smoothing to examples where there are more experts who have voted and weaker smoothing when fewer experts are involved. The objective of my proposed scheme is the opposite. In the initial post I had posited that when more experts vote on a particular example then we can have more \"confidence\" in the expert consensus reached, however this may be a bad intuition as pointed out by @ogurtsov.\n\nAfter some thought I believe a more correct interpretation is that when we have more experts voting on a particular example the resulting distribution is more **representative** of how other experts would vote since each expert is independent. However, there may be some detail I am missing in this interpretation. If we assume this interpretation to be the case then we still have the same goal of wanting stronger smoothing when fewer experts are involved because having only one expert vote is not necessarily a good representation of how other experts will vote. \n\nThere are many unknowns in this analysis, such as; Are examples with only 1 vote trivial examples and so more experts aren't needed to correctly diagnose? In this case my interpretation above would be incorrect. Regardless, empirically my smoothing scheme has improved my own results and we may try other schemes to see what is best. I encourage you to try what you have suggested and see if it works, then let us know here.",
      "votes": null
    },
    {
      "id": "2655355",
      "postDate": "02/16/2024 19:57:00",
      "content": "<p>No. Just contrary. The amount of equally distributed new votes will be the one that turns all labels equally voted. So where there are more expert votations with respect the max, less additional votes will be added. And viceversa.<br>\nif Max expert votes = 10<br>\n[1, 0, 0, 0, 0, 0] -&gt; [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6]<br>\n[2, 0, 0, 0, 0, 0] -&gt; [1 + 8/6, 8/6, 8/6, 8/6, 8/6, 8/6]<br>\n…<br>\n[9, 0, 0, 0, 0, 0] -&gt; [1 + 1/6, 1/6, 1/6, 1/6, 1/6, 1/6]<br>\n[10, 0, 0, 0, 0, 0] -&gt; [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]<br>\nMy english could be better. The idea would be to vote equally with 1/6 the remaining votes of the courrent label to match the maximum observed votes. And finally /= Max expert votes</p>",
      "rawMarkdown": "No. Just contrary. The amount of equally distributed new votes will be the one that turns all labels equally voted. So where there are more expert votations with respect the max, less additional votes will be added. And viceversa.\nif Max expert votes = 10\n[1, 0, 0, 0, 0, 0] -> [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6]\n[2, 0, 0, 0, 0, 0] -> [1 + 8/6, 8/6, 8/6, 8/6, 8/6, 8/6]\n...\n[9, 0, 0, 0, 0, 0] -> [1 + 1/6, 1/6, 1/6, 1/6, 1/6, 1/6]\n[10, 0, 0, 0, 0, 0] -> [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]\nMy english could be better. The idea would be to vote equally with 1/6 the remaining votes of the courrent label to match the maximum observed votes. And finally /= Max expert votes",
      "votes": null
    },
    {
      "id": "2655383",
      "postDate": "02/16/2024 20:21:00",
      "content": "<p>I see, I think I misunderstood your description. So you mean to redistribute votes based on the maximum number of experts for each disorder? For example if 'Seizure' had a max of 10 we would perform the scheme you have provided above and if 'LPD' had a max of 15 we would have something like:</p>\n<p>[1, 0, 0, 0, 0, 0] -&gt; [1 + 14/6, 14/6, 14/6, 14/6, 14/6, 14/6]<br>\n[2, 0, 0, 0, 0, 0] -&gt; [1 + 13/6, 13/6, 13/6, 13/6, 13/6, 13/6]<br>\n…<br>\n[15, 0, 0, 0, 0, 0] -&gt; [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]</p>\n<p>If this is the correct interpretation of your suggestion its very interesting because it encodes some kind of information on how much agreement experts typically reach for specific disorders. </p>",
      "rawMarkdown": "I see, I think I misunderstood your description. So you mean to redistribute votes based on the maximum number of experts for each disorder? For example if 'Seizure' had a max of 10 we would perform the scheme you have provided above and if 'LPD' had a max of 15 we would have something like:\n\n[1, 0, 0, 0, 0, 0] -> [1 + 14/6, 14/6, 14/6, 14/6, 14/6, 14/6]\n[2, 0, 0, 0, 0, 0] -> [1 + 13/6, 13/6, 13/6, 13/6, 13/6, 13/6]\n...\n[15, 0, 0, 0, 0, 0] -> [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]\n\nIf this is the correct interpretation of your suggestion its very interesting because it encodes some kind of information on how much agreement experts typically reach for specific disorders.",
      "votes": null
    },
    {
      "id": "2655390",
      "postDate": "02/16/2024 20:28:11",
      "content": "<p>I actually was thinking into a global maximum, not per seizure. I used first class just to simplify. Just to smooth less confident absolute votations. Not sure if per seizure would be more or less reasonable.<br>\nActually I think is the same result.</p>",
      "rawMarkdown": "I actually was thinking into a global maximum, not per seizure. I used first class just to simplify. Just to smooth less confident absolute votations. Not sure if per seizure would be more or less reasonable.\nActually I think is the same result.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2655076,
      "author_name": "ogurtsov",
      "author_url": "",
      "post_date": "02/16/2024 17:12:31",
      "content": "<blockquote>\n  <p>Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached.</p>\n</blockquote>\n<p>Unfortunately, this intuition is completelty wrong because different number of experts isn't picked at random across all samples from all patients. Actual labeling algorithm looks pretty cryptic according to <a href=\"https://bdsp.io/content/bdsp-sparcnet/1.1/\" target=\"_blank\">https://bdsp.io/content/bdsp-sparcnet/1.1/</a></p>\n<p>Moreover, in regular clinical practice more uncertain cases tend to involve more experts to diagnosis procedure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655170,
          "author_name": "djhuth",
          "author_url": "",
          "post_date": "02/16/2024 17:49:26",
          "content": "<p>Thank you for the insight. So because experts aren't chosen at random for each patient there will be some biases those experts carry, thereby making the consensus of the label distributions similarly biased? Has it been determined that the Kaggle dataset is a subset of this SPaRCNet data?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2655194,
              "author_name": "ogurtsov",
              "author_url": "",
              "post_date": "02/16/2024 17:59:44",
              "content": "<p>Possible biases in assessment are not clear for me. Good news: experts are independent, so they can't influence each other.<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439</a> <code>SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.</code> (c) Kaggle Staff</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2655227,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/16/2024 18:50:05",
      "content": "<p>I have a very basic statistics base so apologize if is a terrible suggestion.<br>\nBut what with  distribute equally votes to each label to match the maximum of expert votes in a single label? Let's say the maximum voted label has 10 expert votes and an example labels has [1, 0, 0, 0, 0, 0], so turn it into [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6].<br>\nJust as your suggestion direction but without parameters.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2655349,
          "author_name": "djhuth",
          "author_url": "",
          "post_date": "02/16/2024 19:53:32",
          "content": "<p>This is certainly something one can try. However, if I understand your suggestion correctly this will apply stronger smoothing to examples where there are more experts who have voted and weaker smoothing when fewer experts are involved. The objective of my proposed scheme is the opposite. In the initial post I had posited that when more experts vote on a particular example then we can have more \"confidence\" in the expert consensus reached, however this may be a bad intuition as pointed out by <a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a>.</p>\n<p>After some thought I believe a more correct interpretation is that when we have more experts voting on a particular example the resulting distribution is more <strong>representative</strong> of how other experts would vote since each expert is independent. However, there may be some detail I am missing in this interpretation. If we assume this interpretation to be the case then we still have the same goal of wanting stronger smoothing when fewer experts are involved because having only one expert vote is not necessarily a good representation of how other experts will vote. </p>\n<p>There are many unknowns in this analysis, such as; Are examples with only 1 vote trivial examples and so more experts aren't needed to correctly diagnose? In this case my interpretation above would be incorrect. Regardless, empirically my smoothing scheme has improved my own results and we may try other schemes to see what is best. I encourage you to try what you have suggested and see if it works, then let us know here.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2655355,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "02/16/2024 19:57:00",
              "content": "<p>No. Just contrary. The amount of equally distributed new votes will be the one that turns all labels equally voted. So where there are more expert votations with respect the max, less additional votes will be added. And viceversa.<br>\nif Max expert votes = 10<br>\n[1, 0, 0, 0, 0, 0] -&gt; [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6]<br>\n[2, 0, 0, 0, 0, 0] -&gt; [1 + 8/6, 8/6, 8/6, 8/6, 8/6, 8/6]<br>\n…<br>\n[9, 0, 0, 0, 0, 0] -&gt; [1 + 1/6, 1/6, 1/6, 1/6, 1/6, 1/6]<br>\n[10, 0, 0, 0, 0, 0] -&gt; [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]<br>\nMy english could be better. The idea would be to vote equally with 1/6 the remaining votes of the courrent label to match the maximum observed votes. And finally /= Max expert votes</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2655383,
                  "author_name": "djhuth",
                  "author_url": "",
                  "post_date": "02/16/2024 20:21:00",
                  "content": "<p>I see, I think I misunderstood your description. So you mean to redistribute votes based on the maximum number of experts for each disorder? For example if 'Seizure' had a max of 10 we would perform the scheme you have provided above and if 'LPD' had a max of 15 we would have something like:</p>\n<p>[1, 0, 0, 0, 0, 0] -&gt; [1 + 14/6, 14/6, 14/6, 14/6, 14/6, 14/6]<br>\n[2, 0, 0, 0, 0, 0] -&gt; [1 + 13/6, 13/6, 13/6, 13/6, 13/6, 13/6]<br>\n…<br>\n[15, 0, 0, 0, 0, 0] -&gt; [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]</p>\n<p>If this is the correct interpretation of your suggestion its very interesting because it encodes some kind of information on how much agreement experts typically reach for specific disorders. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2655390,
                      "author_name": "sacuscreed",
                      "author_url": "",
                      "post_date": "02/16/2024 20:28:11",
                      "content": "<p>I actually was thinking into a global maximum, not per seizure. I used first class just to simplify. Just to smooth less confident absolute votations. Not sure if per seizure would be more or less reasonable.<br>\nActually I think is the same result.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2655028": "An important topic of discussion in this competition has been how to account for the number of experts who have voted on any given example of EEG and Spectrogram data. [This helpful post](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135) by @seanbearden introduced a two stage training process based on the apparent separate sources of training data we are given, discovered by @pcjimmmy in this [notebook](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda). Here I propose a simple method for smoothing class labels based on the number of experts who have voted on a given example.\n\nConsider a ground truth label representing the probability of some set of classes of the form `[1, 0, 0, 0, 0, 0]`. Label smoothing penalizes the ground truth class by redistributing a portion the positive label across the other classes, preventing a trained model from becoming overly confident in its predictions and helping to regularize the network. For some small number `p` the smoothed label looks like `[1-p, p/(k-1), p/(k-1), p/(k-1), p/(k-1), p/(k-1) ]` where `k` is the number of classes. Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached. Therefore would like to choose a value of `p` such that there is a higher degree of smoothing when fewer experts are involved and less smoothing with more experts.\n\nA simple functional form we can start with is one which increases `p` when fewer experts have voted and decreases `p` when there are more experts. As a starting point I propose the following: `p = 1/(num_experts + 1)` where the `+ 1` in the denominator is included so that the positive ground truth label is not reduced to 0 when only one expert has voted. Now that our value of `p` is chosen we then penalize the label which has been decided on in the `expert_consensus` column of `train.csv` and distribute to the other classes. As an example consider a 'Seizure' consensus with the label `[1, 0, 0, 0, 0, 0]` and 3 expert votes. Our value of `p` is `1/(3+1) = 0.25` and our smoothed label becomes `[0.75, 0.05, 0.05, 0.05, 0.05, 0.05]`. \n\nI have implemented this in my own training and have observed improved regularization and better CV performance in 5-fold training grouped by `patient_id` in a 1D CNN+RNN model. Of course this is very simple and only serves to act as a starting point for further exploration. Other functional forms for `p` may improve performance further and more sophisticated methods for choosing how to redistribute the labels could be important (e.g. the case of an even split among experts). Please share your thoughts and ideas below and perhaps where you think this method may fail.",
    "2655076": ">Intuitively, we expect that when a training example has fewer expert votes that there is less confidence in the expert consensus and when the example has more expert votes there is more confidence in the consensus reached.\n\nUnfortunately, this intuition is completelty wrong because different number of experts isn't picked at random across all samples from all patients. Actual labeling algorithm looks pretty cryptic according to https://bdsp.io/content/bdsp-sparcnet/1.1/\n\nMoreover, in regular clinical practice more uncertain cases tend to involve more experts to diagnosis procedure.",
    "2655170": "Thank you for the insight. So because experts aren't chosen at random for each patient there will be some biases those experts carry, thereby making the consensus of the label distributions similarly biased? Has it been determined that the Kaggle dataset is a subset of this SPaRCNet data?",
    "2655194": "Possible biases in assessment are not clear for me. Good news: experts are independent, so they can't influence each other.\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471439 `SPaRCNet is a rough equivalent of the training data. It does not overlap with the test set.` (c) Kaggle Staff",
    "2655227": "I have a very basic statistics base so apologize if is a terrible suggestion.\nBut what with  distribute equally votes to each label to match the maximum of expert votes in a single label? Let's say the maximum voted label has 10 expert votes and an example labels has [1, 0, 0, 0, 0, 0], so turn it into [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6].\nJust as your suggestion direction but without parameters.",
    "2655349": "This is certainly something one can try. However, if I understand your suggestion correctly this will apply stronger smoothing to examples where there are more experts who have voted and weaker smoothing when fewer experts are involved. The objective of my proposed scheme is the opposite. In the initial post I had posited that when more experts vote on a particular example then we can have more \"confidence\" in the expert consensus reached, however this may be a bad intuition as pointed out by @ogurtsov.\n\nAfter some thought I believe a more correct interpretation is that when we have more experts voting on a particular example the resulting distribution is more **representative** of how other experts would vote since each expert is independent. However, there may be some detail I am missing in this interpretation. If we assume this interpretation to be the case then we still have the same goal of wanting stronger smoothing when fewer experts are involved because having only one expert vote is not necessarily a good representation of how other experts will vote. \n\nThere are many unknowns in this analysis, such as; Are examples with only 1 vote trivial examples and so more experts aren't needed to correctly diagnose? In this case my interpretation above would be incorrect. Regardless, empirically my smoothing scheme has improved my own results and we may try other schemes to see what is best. I encourage you to try what you have suggested and see if it works, then let us know here.",
    "2655355": "No. Just contrary. The amount of equally distributed new votes will be the one that turns all labels equally voted. So where there are more expert votations with respect the max, less additional votes will be added. And viceversa.\nif Max expert votes = 10\n[1, 0, 0, 0, 0, 0] -> [1 + 9/6, 9/6, 9/6, 9/6, 9/6, 9/6]\n[2, 0, 0, 0, 0, 0] -> [1 + 8/6, 8/6, 8/6, 8/6, 8/6, 8/6]\n...\n[9, 0, 0, 0, 0, 0] -> [1 + 1/6, 1/6, 1/6, 1/6, 1/6, 1/6]\n[10, 0, 0, 0, 0, 0] -> [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]\nMy english could be better. The idea would be to vote equally with 1/6 the remaining votes of the courrent label to match the maximum observed votes. And finally /= Max expert votes",
    "2655383": "I see, I think I misunderstood your description. So you mean to redistribute votes based on the maximum number of experts for each disorder? For example if 'Seizure' had a max of 10 we would perform the scheme you have provided above and if 'LPD' had a max of 15 we would have something like:\n\n[1, 0, 0, 0, 0, 0] -> [1 + 14/6, 14/6, 14/6, 14/6, 14/6, 14/6]\n[2, 0, 0, 0, 0, 0] -> [1 + 13/6, 13/6, 13/6, 13/6, 13/6, 13/6]\n...\n[15, 0, 0, 0, 0, 0] -> [1 + 0/6, 0/6, 0/6, 0/6, 0/6, 0/6]\n\nIf this is the correct interpretation of your suggestion its very interesting because it encodes some kind of information on how much agreement experts typically reach for specific disorders.",
    "2655390": "I actually was thinking into a global maximum, not per seizure. I used first class just to simplify. Just to smooth less confident absolute votations. Not sure if per seizure would be more or less reasonable.\nActually I think is the same result."
  },
  "source": "meta"
}