{
  "id": 470645,
  "title": "k-means clusters for classification",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470645",
  "author_name": "Daniel Dewey",
  "post_date": "2024-01-24T23:05:34.141000",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In classifiers like <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60?scriptVersionId=159895287\" target=\"_blank\">CatBoost Starter</a> the classes are \"unit vectors\" for each of the HBAs. As the vote-entropy plot below shows, there is not always complete agreement on the particular HBA. In the classifier output this ambiguity would show up in the <code>proba</code> class values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc58d3230020af2496a4f4b96ada44f4c%2Fentropy_of_votes.png?generation=1706137237229795&amp;alt=media\"></p>\n<p>Alternately, some of the ambiguity can be built into the set of classes, for example by assigning classes based on a k-means clustering. Doing simple k-means clustering on the vote probability vectors gives an \"elbow\" at 6 clusters, not surprisingly one for each HBA. However, adding a 7th cluster produces a center away from the HBA peaks and centered at <code>[3.4% 12.1% 5.5% 15.7% 18.8% 44.5%]</code> . Perhaps this captures some of the inherent ambiguity among the classes? Maybe fitting on a target using the cluster-based classes might be more \"orthogonal\" and produce improved&nbsp;results?</p>\n<pre><code># The centers for  clusters:   * ordered by main column *\n# [     ]\n# [         ]\n# [      ]\n# [     ]\n# [     ]\n# [     ]\n# The new th cluster:\n# [     ] \n</code></pre>\n<p>Notebook: <a href=\"https://www.kaggle.com/dan3dewey/hms-2024-brain-clusters\" target=\"_blank\">HMS 2024 - Brain Clusters</a></p>",
  "messages": [
    {
      "id": 2618683,
      "postDate": "2024-01-24T23:05:34.140Z",
      "content": "<p>In classifiers like <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's <a href=\"https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60?scriptVersionId=159895287\" target=\"_blank\">CatBoost Starter</a> the classes are \"unit vectors\" for each of the HBAs. As the vote-entropy plot below shows, there is not always complete agreement on the particular HBA. In the classifier output this ambiguity would show up in the <code>proba</code> class values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc58d3230020af2496a4f4b96ada44f4c%2Fentropy_of_votes.png?generation=1706137237229795&amp;alt=media\"></p>\n<p>Alternately, some of the ambiguity can be built into the set of classes, for example by assigning classes based on a k-means clustering. Doing simple k-means clustering on the vote probability vectors gives an \"elbow\" at 6 clusters, not surprisingly one for each HBA. However, adding a 7th cluster produces a center away from the HBA peaks and centered at <code>[3.4% 12.1% 5.5% 15.7% 18.8% 44.5%]</code> . Perhaps this captures some of the inherent ambiguity among the classes? Maybe fitting on a target using the cluster-based classes might be more \"orthogonal\" and produce improved&nbsp;results?</p>\n<pre><code># The centers for  clusters:   * ordered by main column *\n# [     ]\n# [         ]\n# [      ]\n# [     ]\n# [     ]\n# [     ]\n# The new th cluster:\n# [     ] \n</code></pre>\n<p>Notebook: <a href=\"https://www.kaggle.com/dan3dewey/hms-2024-brain-clusters\" target=\"_blank\">HMS 2024 - Brain Clusters</a></p>",
      "rawMarkdown": "In classifiers like @cdeotte 's [CatBoost Starter](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60?scriptVersionId=159895287) the classes are \"unit vectors\" for each of the HBAs. As the vote-entropy plot below shows, there is not always complete agreement on the particular HBA. In the classifier output this ambiguity would show up in the `proba` class values.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc58d3230020af2496a4f4b96ada44f4c%2Fentropy_of_votes.png?generation=1706137237229795&alt=media)\n\nAlternately, some of the ambiguity can be built into the set of classes, for example by assigning classes based on a k-means clustering. Doing simple k-means clustering on the vote probability vectors gives an \"elbow\" at 6 clusters, not surprisingly one for each HBA. However, adding a 7th cluster produces a center away from the HBA peaks and centered at `[3.4% 12.1% 5.5% 15.7% 18.8% 44.5%]` . Perhaps this captures some of the inherent ambiguity among the classes? Maybe fitting on a target using the cluster-based classes might be more \"orthogonal\" and produce improved results?\n\n```\n# The centers for 7 clusters:   * ordered by main column *\n# [0.97295318 0.00590648 0.00483242 0.00223349 0.00114005 0.01293438]\n# [0.0359414  0.8263946  0.02026326 0.0416359  0.0039719  0.07179293]\n# [0.09039322 0.05736363 0.73959869 0.0036609  0.02577728 0.08320629]\n# [0.01311906 0.05057075 0.00535362 0.76954905 0.05295756 0.10844996]\n# [0.00294371 0.00356274 0.00838875 0.02236721 0.92515794 0.03757965]\n# [0.00794713 0.01162694 0.01349516 0.01482245 0.02306207 0.92904625]\n# The new 7th cluster:\n# [0.03401079 0.12059055 0.05457865 0.15692139 0.18847755 0.44542107] \n```\n\nNotebook: [HMS 2024 - Brain Clusters](https://www.kaggle.com/dan3dewey/hms-2024-brain-clusters)\n",
      "votes": 18
    },
    {
      "id": 2618899,
      "postDate": "2024-01-25T04:52:17.167Z",
      "content": "<p>Good analysis, Daniel! There is one angle that you might not have explored yet and would be worth looking at: for a given patient/eeg_id/spectrogram_id, there are several identical voting distributions (e.g. 3 sz, 0 for all others, for 5 label_ids from a given eeg_id/spectrogram_id combo). One can assume that these identical voting results arise from the same sub-group of experts; one voting result instead of many should suffice. Therefore, the analysis should also be done by disregarding these \"duplicates\". I am curious to see if you would achieve the same results with that streamlined dataset.</p>",
      "rawMarkdown": "Good analysis, Daniel! There is one angle that you might not have explored yet and would be worth looking at: for a given patient/eeg_id/spectrogram_id, there are several identical voting distributions (e.g. 3 sz, 0 for all others, for 5 label_ids from a given eeg_id/spectrogram_id combo). One can assume that these identical voting results arise from the same sub-group of experts; one voting result instead of many should suffice. Therefore, the analysis should also be done by disregarding these \"duplicates\". I am curious to see if you would achieve the same results with that streamlined dataset.",
      "votes": 1
    },
    {
      "id": 2618698,
      "postDate": "2024-01-24T23:25:48.847Z",
      "content": "<p>This is very interesting. Do you mind sharing this notebook? </p>",
      "rawMarkdown": "This is very interesting. Do you mind sharing this notebook? ",
      "votes": 1,
      "replies": [
        {
          "id": 2618774,
          "postDate": "2024-01-25T01:39:46.873Z",
          "content": "<p>I added a link to the notebook and made it public.  Thanks!</p>",
          "rawMarkdown": "I added a link to the notebook and made it public.  Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2621720,
      "postDate": "2024-01-26T23:51:28.363Z",
      "content": "<p><a href=\"https://www.kaggle.com/patrob\" target=\"_blank\">@patrob</a> , your comment about \"identical voting\" within an eeg is very interesting.  I added some code to the end of the notebook to group the train rows in different ways and there is indeed an \"identical vote\" effect. </p>\n<p>First, just grouping with patient_id and adding expert_consensus shows that patients only have on average 2 different HBAs ever assigned to them:</p>\n<pre><code>#     Number    HBAs  a patient: looks  about   .\n#   , grouped : [\"patient_id\"]\n#   , grouped : [\"patient_id\", \"expert_consensus\"]\n#   , grouped : [\"patient_id\", \"rand_id\"] &lt;\n</code></pre>\n<p>Then, grouping on patient_id and eeg_id gives the same number of rows as eeg_id, so as was noted before, each eeg is specific to a single patient. A surprise is that by then grouping as well on expert_consensus, the number of rows grows very little, only 5% more (compared to 54% more when rand_id with 2 choices is included.) This says that the eegs each contain primarily one single HBA type.</p>\n<pre><code>#     Variety  HBAs  votes  patient\n#  , grouped : [\"patient_id\", \"eeg_id\"]\n#  , grouped : [\"patient_id\", \"eeg_id\", \"expert_consensus\"]\n#  , grouped : [\"patient_id\", \"eeg_id\", \"rand_id\"] &lt;\n</code></pre>\n<p>Finally, getting to the voting, I include additionally grouping by total_vote (about 10% more) and by the maximum vote (1.5% more.) This gives 20072 unique patient-eeg-consensus-voting combinations out of the total 106800.</p>\n<pre><code>#  , grouped : [\"patient_id\", . . . _consensus\", \"total_vote\"]\n# 20072 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\",\"max_vote\"]\n</code></pre>\n<p>This suggests that expert assessment is commonly applied to multiple similar HBAs in an eeg, rather than assessing each individually. For example the image below shows the grouping selecting on ones with more than 2 identicals. The right-most column, \"eeg_sub_id\" is actually the counts() in the grouping. For example, the last line shows that patient 65494 in eeg 3758950107 had 12 seizures noted, and they all received 3 seizure votes out of 3 total votes.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc9c7c5edfa184caf02312c363eaf9fc7%2Fmultiple_identical_votes.png?generation=1706313041439152&amp;alt=media\"><br>\nI don't know off hand what the implications are for how these \"identically voted\" samples should be used. Don't do anything different? Choose only one to include? Weight each of them by 1/n_identical? Maybe it's best to wait and see what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has to say  😂</p>",
      "rawMarkdown": "@patrob , your comment about \"identical voting\" within an eeg is very interesting.  I added some code to the end of the notebook to group the train rows in different ways and there is indeed an \"identical vote\" effect. \n\nFirst, just grouping with patient_id and adding expert_consensus shows that patients only have on average 2 different HBAs ever assigned to them:\n\n```\n#     Number of types of HBAs for a patient: looks like about 2 for each.\n#  1950 rows, grouped by: [\"patient_id\"]\n#  3625 rows, grouped by: [\"patient_id\", \"expert_consensus\"]\n#  3799 rows, grouped by: [\"patient_id\", \"rand_id\"] <-- rand_id has 2 values, 0,1\n```\nThen, grouping on patient_id and eeg_id gives the same number of rows as eeg_id, so as was noted before, each eeg is specific to a single patient. A surprise is that by then grouping as well on expert_consensus, the number of rows grows very little, only 5% more (compared to 54% more when rand_id with 2 choices is included.) This says that the eegs each contain primarily one single HBA type.\n```\n#     Variety of HBAs and votes in patient--eeg combinations\n# 17089 rows, grouped by: [\"patient_id\", \"eeg_id\"]\n# 18013 rows, grouped by: [\"patient_id\", \"eeg_id\", \"expert_consensus\"]\n# 26266 rows, grouped by: [\"patient_id\", \"eeg_id\", \"rand_id\"] <-- 2 coices, 0,1\n```\nFinally, getting to the voting, I include additionally grouping by total_vote (about 10% more) and by the maximum vote (1.5% more.) This gives 20072 unique patient-eeg-consensus-voting combinations out of the total 106800.\n```\n# 19783 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\"]\n# 20072 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\",\"max_vote\"]\n```\nThis suggests that expert assessment is commonly applied to multiple similar HBAs in an eeg, rather than assessing each individually. For example the image below shows the grouping selecting on ones with more than 2 identicals. The right-most column, \"eeg_sub_id\" is actually the counts() in the grouping. For example, the last line shows that patient 65494 in eeg 3758950107 had 12 seizures noted, and they all received 3 seizure votes out of 3 total votes.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc9c7c5edfa184caf02312c363eaf9fc7%2Fmultiple_identical_votes.png?generation=1706313041439152&alt=media)\nI don't know off hand what the implications are for how these \"identically voted\" samples should be used. Don't do anything different? Choose only one to include? Weight each of them by 1/n_identical? Maybe it's best to wait and see what @cdeotte has to say  😂"
    },
    {
      "id": 2622135,
      "postDate": "2024-01-27T08:57:41.470Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2618899,
      "author_name": "Patrick Robitaille",
      "author_url": "",
      "post_date": "2024-01-25T04:52:17.167000",
      "content": "<p>Good analysis, Daniel! There is one angle that you might not have explored yet and would be worth looking at: for a given patient/eeg_id/spectrogram_id, there are several identical voting distributions (e.g. 3 sz, 0 for all others, for 5 label_ids from a given eeg_id/spectrogram_id combo). One can assume that these identical voting results arise from the same sub-group of experts; one voting result instead of many should suffice. Therefore, the analysis should also be done by disregarding these \"duplicates\". I am curious to see if you would achieve the same results with that streamlined dataset.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2618698,
      "author_name": "Yan Teixeira",
      "author_url": "",
      "post_date": "2024-01-24T23:25:48.847000",
      "content": "<p>This is very interesting. Do you mind sharing this notebook? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2618774,
          "author_name": "Daniel Dewey",
          "author_url": "",
          "post_date": "2024-01-25T01:39:46.873000",
          "content": "<p>I added a link to the notebook and made it public.  Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2621720,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-01-26T23:51:28.363000",
      "content": "<p><a href=\"https://www.kaggle.com/patrob\" target=\"_blank\">@patrob</a> , your comment about \"identical voting\" within an eeg is very interesting.  I added some code to the end of the notebook to group the train rows in different ways and there is indeed an \"identical vote\" effect. </p>\n<p>First, just grouping with patient_id and adding expert_consensus shows that patients only have on average 2 different HBAs ever assigned to them:</p>\n<pre><code>#     Number    HBAs  a patient: looks  about   .\n#   , grouped : [\"patient_id\"]\n#   , grouped : [\"patient_id\", \"expert_consensus\"]\n#   , grouped : [\"patient_id\", \"rand_id\"] &lt;\n</code></pre>\n<p>Then, grouping on patient_id and eeg_id gives the same number of rows as eeg_id, so as was noted before, each eeg is specific to a single patient. A surprise is that by then grouping as well on expert_consensus, the number of rows grows very little, only 5% more (compared to 54% more when rand_id with 2 choices is included.) This says that the eegs each contain primarily one single HBA type.</p>\n<pre><code>#     Variety  HBAs  votes  patient\n#  , grouped : [\"patient_id\", \"eeg_id\"]\n#  , grouped : [\"patient_id\", \"eeg_id\", \"expert_consensus\"]\n#  , grouped : [\"patient_id\", \"eeg_id\", \"rand_id\"] &lt;\n</code></pre>\n<p>Finally, getting to the voting, I include additionally grouping by total_vote (about 10% more) and by the maximum vote (1.5% more.) This gives 20072 unique patient-eeg-consensus-voting combinations out of the total 106800.</p>\n<pre><code>#  , grouped : [\"patient_id\", . . . _consensus\", \"total_vote\"]\n# 20072 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\",\"max_vote\"]\n</code></pre>\n<p>This suggests that expert assessment is commonly applied to multiple similar HBAs in an eeg, rather than assessing each individually. For example the image below shows the grouping selecting on ones with more than 2 identicals. The right-most column, \"eeg_sub_id\" is actually the counts() in the grouping. For example, the last line shows that patient 65494 in eeg 3758950107 had 12 seizures noted, and they all received 3 seizure votes out of 3 total votes.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc9c7c5edfa184caf02312c363eaf9fc7%2Fmultiple_identical_votes.png?generation=1706313041439152&amp;alt=media\"><br>\nI don't know off hand what the implications are for how these \"identically voted\" samples should be used. Don't do anything different? Choose only one to include? Weight each of them by 1/n_identical? Maybe it's best to wait and see what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> has to say  😂</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2622135,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-27T08:57:41.470000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2618683": "In classifiers like @cdeotte 's [CatBoost Starter](https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-60?scriptVersionId=159895287) the classes are \"unit vectors\" for each of the HBAs. As the vote-entropy plot below shows, there is not always complete agreement on the particular HBA. In the classifier output this ambiguity would show up in the `proba` class values.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc58d3230020af2496a4f4b96ada44f4c%2Fentropy_of_votes.png?generation=1706137237229795&alt=media)\n\nAlternately, some of the ambiguity can be built into the set of classes, for example by assigning classes based on a k-means clustering. Doing simple k-means clustering on the vote probability vectors gives an \"elbow\" at 6 clusters, not surprisingly one for each HBA. However, adding a 7th cluster produces a center away from the HBA peaks and centered at `[3.4% 12.1% 5.5% 15.7% 18.8% 44.5%]` . Perhaps this captures some of the inherent ambiguity among the classes? Maybe fitting on a target using the cluster-based classes might be more \"orthogonal\" and produce improved results?\n\n```\n# The centers for 7 clusters:   * ordered by main column *\n# [0.97295318 0.00590648 0.00483242 0.00223349 0.00114005 0.01293438]\n# [0.0359414  0.8263946  0.02026326 0.0416359  0.0039719  0.07179293]\n# [0.09039322 0.05736363 0.73959869 0.0036609  0.02577728 0.08320629]\n# [0.01311906 0.05057075 0.00535362 0.76954905 0.05295756 0.10844996]\n# [0.00294371 0.00356274 0.00838875 0.02236721 0.92515794 0.03757965]\n# [0.00794713 0.01162694 0.01349516 0.01482245 0.02306207 0.92904625]\n# The new 7th cluster:\n# [0.03401079 0.12059055 0.05457865 0.15692139 0.18847755 0.44542107] \n```\n\nNotebook: [HMS 2024 - Brain Clusters](https://www.kaggle.com/dan3dewey/hms-2024-brain-clusters)\n",
    "2618899": "Good analysis, Daniel! There is one angle that you might not have explored yet and would be worth looking at: for a given patient/eeg_id/spectrogram_id, there are several identical voting distributions (e.g. 3 sz, 0 for all others, for 5 label_ids from a given eeg_id/spectrogram_id combo). One can assume that these identical voting results arise from the same sub-group of experts; one voting result instead of many should suffice. Therefore, the analysis should also be done by disregarding these \"duplicates\". I am curious to see if you would achieve the same results with that streamlined dataset.",
    "2618698": "This is very interesting. Do you mind sharing this notebook? ",
    "2621720": "@patrob , your comment about \"identical voting\" within an eeg is very interesting.  I added some code to the end of the notebook to group the train rows in different ways and there is indeed an \"identical vote\" effect. \n\nFirst, just grouping with patient_id and adding expert_consensus shows that patients only have on average 2 different HBAs ever assigned to them:\n\n```\n#     Number of types of HBAs for a patient: looks like about 2 for each.\n#  1950 rows, grouped by: [\"patient_id\"]\n#  3625 rows, grouped by: [\"patient_id\", \"expert_consensus\"]\n#  3799 rows, grouped by: [\"patient_id\", \"rand_id\"] <-- rand_id has 2 values, 0,1\n```\nThen, grouping on patient_id and eeg_id gives the same number of rows as eeg_id, so as was noted before, each eeg is specific to a single patient. A surprise is that by then grouping as well on expert_consensus, the number of rows grows very little, only 5% more (compared to 54% more when rand_id with 2 choices is included.) This says that the eegs each contain primarily one single HBA type.\n```\n#     Variety of HBAs and votes in patient--eeg combinations\n# 17089 rows, grouped by: [\"patient_id\", \"eeg_id\"]\n# 18013 rows, grouped by: [\"patient_id\", \"eeg_id\", \"expert_consensus\"]\n# 26266 rows, grouped by: [\"patient_id\", \"eeg_id\", \"rand_id\"] <-- 2 coices, 0,1\n```\nFinally, getting to the voting, I include additionally grouping by total_vote (about 10% more) and by the maximum vote (1.5% more.) This gives 20072 unique patient-eeg-consensus-voting combinations out of the total 106800.\n```\n# 19783 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\"]\n# 20072 rows, grouped by: [\"patient_id\", . . . _consensus\", \"total_vote\",\"max_vote\"]\n```\nThis suggests that expert assessment is commonly applied to multiple similar HBAs in an eeg, rather than assessing each individually. For example the image below shows the grouping selecting on ones with more than 2 identicals. The right-most column, \"eeg_sub_id\" is actually the counts() in the grouping. For example, the last line shows that patient 65494 in eeg 3758950107 had 12 seizures noted, and they all received 3 seizure votes out of 3 total votes.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fc9c7c5edfa184caf02312c363eaf9fc7%2Fmultiple_identical_votes.png?generation=1706313041439152&alt=media)\nI don't know off hand what the implications are for how these \"identically voted\" samples should be used. Don't do anything different? Choose only one to include? Weight each of them by 1/n_identical? Maybe it's best to wait and see what @cdeotte has to say  😂",
    "2622135": ""
  }
}