{
  "id": 479153,
  "title": "Regarding the determination of Expert Consensus when there is a tie in the number of votes.",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479153",
  "author_name": "",
  "post_date": "2024-02-23T09:59:53.223978600Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, Kagglers.<br>\nApologies if this has already been discussed in previous discussions.<br>\nAfter conducting some EDA, I discovered data where the votes are tied. In such cases, how was the Expert Consensus(final targets) ultimately determined? Additionally, would it be quite challenging for a model to make these subtle judgments?<br>\nI'm looking for a discussion on this matter.\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8386305%2Fb63bbc64c1403755a3e5f1b08dee6f99%2F2024-02-23%2018.56.56.png?generation=1708682288500579&amp;alt=media\"></p>\n<p>I made this discovery using the following code:</p>\n<pre><code>def filter_rows():\n      [TARGETS]\n    non_zero_values  [v  v   if v  ]\n    if  non_zero_values:\n         \n    max_value  (non_zero_values)\n    unique_values  (non_zero_values)\n     (non_zero_values.()    max_value      unique_values)\n\nfiltered_df  train[train.apply(filter_rows, axis)]\nfiltered_df\n</code></pre>\n<p>Thank you first.</p>",
  "messages": [
    {
      "id": "2664942",
      "postDate": "02/23/2024 09:59:53",
      "content": "<p>Hi, Kagglers.<br>\nApologies if this has already been discussed in previous discussions.<br>\nAfter conducting some EDA, I discovered data where the votes are tied. In such cases, how was the Expert Consensus(final targets) ultimately determined? Additionally, would it be quite challenging for a model to make these subtle judgments?<br>\nI'm looking for a discussion on this matter.\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8386305%2Fb63bbc64c1403755a3e5f1b08dee6f99%2F2024-02-23%2018.56.56.png?generation=1708682288500579&amp;alt=media\"></p>\n<p>I made this discovery using the following code:</p>\n<pre><code>def filter_rows():\n      [TARGETS]\n    non_zero_values  [v  v   if v  ]\n    if  non_zero_values:\n         \n    max_value  (non_zero_values)\n    unique_values  (non_zero_values)\n     (non_zero_values.()    max_value      unique_values)\n\nfiltered_df  train[train.apply(filter_rows, axis)]\nfiltered_df\n</code></pre>\n<p>Thank you first.</p>",
      "rawMarkdown": "Hi, Kagglers.\nApologies if this has already been discussed in previous discussions.\nAfter conducting some EDA, I discovered data where the votes are tied. In such cases, how was the Expert Consensus(final targets) ultimately determined? Additionally, would it be quite challenging for a model to make these subtle judgments?\nI'm looking for a discussion on this matter.\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8386305%2Fb63bbc64c1403755a3e5f1b08dee6f99%2F2024-02-23%2018.56.56.png?generation=1708682288500579&alt=media)\n\n\nI made this discovery using the following code:\n```\ndef filter_rows(row):\n    values = row[TARGETS]\n    non_zero_values = [v for v in values if v != 0]\n    if not non_zero_values:\n        return False\n    max_value = max(non_zero_values)\n    unique_values = set(non_zero_values)\n    return any(non_zero_values.count(value) > 1 and max_value == value for value in unique_values)\n\nfiltered_df = train[train.apply(filter_rows, axis=1)]\nfiltered_df\n```\n\nThank you first.",
      "votes": null
    },
    {
      "id": "2665202",
      "postDate": "02/23/2024 14:11:23",
      "content": "<p>['lpd_vote', 'gpd_vote'] : LPD 165<br>\n['gpd_vote', 'grda_vote'] : GPD 58<br>\n['gpd_vote', 'lrda_vote'] : GPD 3<br>\n['lpd_vote', 'lrda_vote'] : LPD 60<br>\n['lrda_vote', 'grda_vote'] : LRDA 6<br>\n['seizure_vote', 'gpd_vote'] : Seizure    56<br>\n['seizure_vote', 'lpd_vote'] : Seizure    90<br>\n['seizure_vote', 'grda_vote'] : Seizure    3<br>\n['seizure_vote', 'lrda_vote'] : Seizure    6<br>\n['seizure_vote', 'other_vote'] : Seizure    147<br>\n(others always lose)<br>\netc … inside [] is kinds of votes, after \":\" is expert consensus number. <br>\nso if u do compare the strength <strong>SEIZURE&gt;LPD&gt;GPD&gt;LRDA&gt;GRDA&gt;OTHERS</strong><br>\nwill make notebook about this soon</p>",
      "rawMarkdown": "['lpd_vote', 'gpd_vote'] : LPD 165\n['gpd_vote', 'grda_vote'] : GPD 58\n['gpd_vote', 'lrda_vote'] : GPD 3\n['lpd_vote', 'lrda_vote'] : LPD 60\n['lrda_vote', 'grda_vote'] : LRDA 6\n['seizure_vote', 'gpd_vote'] : Seizure    56\n['seizure_vote', 'lpd_vote'] : Seizure    90\n['seizure_vote', 'grda_vote'] : Seizure    3\n['seizure_vote', 'lrda_vote'] : Seizure    6\n['seizure_vote', 'other_vote'] : Seizure    147\n(others always lose)\netc ... inside [] is kinds of votes, after \":\" is expert consensus number. \nso if u do compare the strength **SEIZURE>LPD>GPD>LRDA>GRDA>OTHERS**\nwill make notebook about this soon",
      "votes": null
    },
    {
      "id": "2665379",
      "postDate": "02/23/2024 16:15:58",
      "content": "<p>Expert_consensus is just <strong>argmax</strong>(*_vote).<br>\nargmax in case of equality of votes always returns the <strong>very first vote</strong>, i.e. voice with \"lower index\".<br>\nIt's easy to check:</p>\n<pre><code>df = pd.read_csv()\n\ntars = {: , : , : , : , : , : }\ntargets = [, , , , , ]\n\ndf[] = df.expert_consensus.(tars)\ndf[] = np.argmax(df[targets].values, axis=)\n\n(df[] != df.target).()\n</code></pre>\n<p>0</p>",
      "rawMarkdown": "Expert_consensus is just **argmax**(*_vote).\nargmax in case of equality of votes always returns the **very first vote**, i.e. voice with \"lower index\".\nIt's easy to check:\n\n```python\ndf = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\ntars = {'Seizure': 0, 'LPD': 1, 'GPD': 2, 'LRDA': 3, 'GRDA': 4, 'Other': 5}\ntargets = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\n\ndf['target'] = df.expert_consensus.map(tars)\ndf['argmax_vote'] = np.argmax(df[targets].values, axis=1)\n\n(df['argmax_vote'] != df.target).sum()\n```\n 0",
      "votes": null
    },
    {
      "id": "2665520",
      "postDate": "02/23/2024 17:43:44",
      "content": "<p>\" Additionally, would it be quite challenging for a model to make these subtle judgments?\" Not much. The model has to predict the probabilities, not the consensus.</p>",
      "rawMarkdown": "\" Additionally, would it be quite challenging for a model to make these subtle judgments?\" Not much. The model has to predict the probabilities, not the consensus.",
      "votes": null
    },
    {
      "id": "2665609",
      "postDate": "02/23/2024 18:39:01",
      "content": "<p>Nice to know.  Thanks</p>\n<p>For the data set I am using I found around 4000+ rows of ties.  When chasing another idea last week I had created a column in train for the top 2 vote getters.   With a bit of coding I changed the expert consensus to the second tie getter on a extracted dataframe.  I than merged the 4000+ into my train.   Just getting things running now, but one issue was that I was using eeg_id and eeg_sub_id later for some grouping.  My 4000+ addition were duplicates of course and messed up those results.   So I added 1000 to the eeg_sub_id before I merged the extracted dataframe into train.</p>\n<p>I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the expert_consenus.  Will see if it helps those models.</p>",
      "rawMarkdown": "Nice to know.  Thanks\n\nFor the data set I am using I found around 4000+ rows of ties.  When chasing another idea last week I had created a column in train for the top 2 vote getters.   With a bit of coding I changed the expert consensus to the second tie getter on a extracted dataframe.  I than merged the 4000+ into my train.   Just getting things running now, but one issue was that I was using eeg_id and eeg_sub_id later for some grouping.  My 4000+ addition were duplicates of course and messed up those results.   So I added 1000 to the eeg_sub_id before I merged the extracted dataframe into train.\n\nI am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the expert_consenus.  Will see if it helps those models.",
      "votes": null
    },
    {
      "id": "2665957",
      "postDate": "02/24/2024 02:44:50",
      "content": "<p>Thank you for reply.<br>\nI didn’t know that.<br>\nI was not curious about “expert consensus” because our making model predicts probability.</p>",
      "rawMarkdown": "Thank you for reply.\nI didn’t know that.\nI was not curious about “expert consensus” because our making model predicts probability.",
      "votes": null
    },
    {
      "id": "2666094",
      "postDate": "02/24/2024 06:31:29",
      "content": "<p>Thank you for reply.</p>\n<p>「I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the 　expert_consenus. Will see if it helps those models.」<br>\n→I was thinking same things. Basically, in this competition, we use probabilities that model output, but I was thinking using 'expert_consenus' helps.<br>\nAnd, I was thinking of trying to make a bit change to targets probabilities which has tie votes based on expert_consenus.</p>",
      "rawMarkdown": "Thank you for reply.\n\n「I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the 　expert_consenus. Will see if it helps those models.」\n→I was thinking same things. Basically, in this competition, we use probabilities that model output, but I was thinking using 'expert_consenus' helps.\nAnd, I was thinking of trying to make a bit change to targets probabilities which has tie votes based on expert_consenus.",
      "votes": null
    },
    {
      "id": "2666289",
      "postDate": "02/24/2024 09:41:39",
      "content": "<p>In reality, there are 5 IIIC pattern classes not 6, the 'other' class (non-IIIC patterns) contains segments that were noisy and too poor in quality to classify in one of the 5 classes. you can read more about the process of annotation <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10136013/\" target=\"_blank\">here</a>.<br>\nIt would be interesting If you can find a tie between two of the five classes except others</p>",
      "rawMarkdown": "In reality, there are 5 IIIC pattern classes not 6, the 'other' class (non-IIIC patterns) contains segments that were noisy and too poor in quality to classify in one of the 5 classes. you can read more about the process of annotation [here](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10136013/).\nIt would be interesting If you can find a tie between two of the five classes except others",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2665202,
      "author_name": "beckpro",
      "author_url": "",
      "post_date": "02/23/2024 14:11:23",
      "content": "<p>['lpd_vote', 'gpd_vote'] : LPD 165<br>\n['gpd_vote', 'grda_vote'] : GPD 58<br>\n['gpd_vote', 'lrda_vote'] : GPD 3<br>\n['lpd_vote', 'lrda_vote'] : LPD 60<br>\n['lrda_vote', 'grda_vote'] : LRDA 6<br>\n['seizure_vote', 'gpd_vote'] : Seizure    56<br>\n['seizure_vote', 'lpd_vote'] : Seizure    90<br>\n['seizure_vote', 'grda_vote'] : Seizure    3<br>\n['seizure_vote', 'lrda_vote'] : Seizure    6<br>\n['seizure_vote', 'other_vote'] : Seizure    147<br>\n(others always lose)<br>\netc … inside [] is kinds of votes, after \":\" is expert consensus number. <br>\nso if u do compare the strength <strong>SEIZURE&gt;LPD&gt;GPD&gt;LRDA&gt;GRDA&gt;OTHERS</strong><br>\nwill make notebook about this soon</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2665379,
      "author_name": "maatkara",
      "author_url": "",
      "post_date": "02/23/2024 16:15:58",
      "content": "<p>Expert_consensus is just <strong>argmax</strong>(*_vote).<br>\nargmax in case of equality of votes always returns the <strong>very first vote</strong>, i.e. voice with \"lower index\".<br>\nIt's easy to check:</p>\n<pre><code>df = pd.read_csv()\n\ntars = {: , : , : , : , : , : }\ntargets = [, , , , , ]\n\ndf[] = df.expert_consensus.(tars)\ndf[] = np.argmax(df[targets].values, axis=)\n\n(df[] != df.target).()\n</code></pre>\n<p>0</p>",
      "votes": null,
      "replies": [
        {
          "id": 2665957,
          "author_name": "haruki741",
          "author_url": "",
          "post_date": "02/24/2024 02:44:50",
          "content": "<p>Thank you for reply.<br>\nI didn’t know that.<br>\nI was not curious about “expert consensus” because our making model predicts probability.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2665520,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/23/2024 17:43:44",
      "content": "<p>\" Additionally, would it be quite challenging for a model to make these subtle judgments?\" Not much. The model has to predict the probabilities, not the consensus.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2665609,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/23/2024 18:39:01",
      "content": "<p>Nice to know.  Thanks</p>\n<p>For the data set I am using I found around 4000+ rows of ties.  When chasing another idea last week I had created a column in train for the top 2 vote getters.   With a bit of coding I changed the expert consensus to the second tie getter on a extracted dataframe.  I than merged the 4000+ into my train.   Just getting things running now, but one issue was that I was using eeg_id and eeg_sub_id later for some grouping.  My 4000+ addition were duplicates of course and messed up those results.   So I added 1000 to the eeg_sub_id before I merged the extracted dataframe into train.</p>\n<p>I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the expert_consenus.  Will see if it helps those models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2666094,
          "author_name": "haruki741",
          "author_url": "",
          "post_date": "02/24/2024 06:31:29",
          "content": "<p>Thank you for reply.</p>\n<p>「I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the 　expert_consenus. Will see if it helps those models.」<br>\n→I was thinking same things. Basically, in this competition, we use probabilities that model output, but I was thinking using 'expert_consenus' helps.<br>\nAnd, I was thinking of trying to make a bit change to targets probabilities which has tie votes based on expert_consenus.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2666289,
      "author_name": "ahmedelfazouan",
      "author_url": "",
      "post_date": "02/24/2024 09:41:39",
      "content": "<p>In reality, there are 5 IIIC pattern classes not 6, the 'other' class (non-IIIC patterns) contains segments that were noisy and too poor in quality to classify in one of the 5 classes. you can read more about the process of annotation <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10136013/\" target=\"_blank\">here</a>.<br>\nIt would be interesting If you can find a tie between two of the five classes except others</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2664942": "Hi, Kagglers.\nApologies if this has already been discussed in previous discussions.\nAfter conducting some EDA, I discovered data where the votes are tied. In such cases, how was the Expert Consensus(final targets) ultimately determined? Additionally, would it be quite challenging for a model to make these subtle judgments?\nI'm looking for a discussion on this matter.\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8386305%2Fb63bbc64c1403755a3e5f1b08dee6f99%2F2024-02-23%2018.56.56.png?generation=1708682288500579&alt=media)\n\n\nI made this discovery using the following code:\n```\ndef filter_rows(row):\n    values = row[TARGETS]\n    non_zero_values = [v for v in values if v != 0]\n    if not non_zero_values:\n        return False\n    max_value = max(non_zero_values)\n    unique_values = set(non_zero_values)\n    return any(non_zero_values.count(value) > 1 and max_value == value for value in unique_values)\n\nfiltered_df = train[train.apply(filter_rows, axis=1)]\nfiltered_df\n```\n\nThank you first.",
    "2665202": "['lpd_vote', 'gpd_vote'] : LPD 165\n['gpd_vote', 'grda_vote'] : GPD 58\n['gpd_vote', 'lrda_vote'] : GPD 3\n['lpd_vote', 'lrda_vote'] : LPD 60\n['lrda_vote', 'grda_vote'] : LRDA 6\n['seizure_vote', 'gpd_vote'] : Seizure    56\n['seizure_vote', 'lpd_vote'] : Seizure    90\n['seizure_vote', 'grda_vote'] : Seizure    3\n['seizure_vote', 'lrda_vote'] : Seizure    6\n['seizure_vote', 'other_vote'] : Seizure    147\n(others always lose)\netc ... inside [] is kinds of votes, after \":\" is expert consensus number. \nso if u do compare the strength **SEIZURE>LPD>GPD>LRDA>GRDA>OTHERS**\nwill make notebook about this soon",
    "2665379": "Expert_consensus is just **argmax**(*_vote).\nargmax in case of equality of votes always returns the **very first vote**, i.e. voice with \"lower index\".\nIt's easy to check:\n\n```python\ndf = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\ntars = {'Seizure': 0, 'LPD': 1, 'GPD': 2, 'LRDA': 3, 'GRDA': 4, 'Other': 5}\ntargets = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\n\ndf['target'] = df.expert_consensus.map(tars)\ndf['argmax_vote'] = np.argmax(df[targets].values, axis=1)\n\n(df['argmax_vote'] != df.target).sum()\n```\n 0",
    "2665520": "\" Additionally, would it be quite challenging for a model to make these subtle judgments?\" Not much. The model has to predict the probabilities, not the consensus.",
    "2665609": "Nice to know.  Thanks\n\nFor the data set I am using I found around 4000+ rows of ties.  When chasing another idea last week I had created a column in train for the top 2 vote getters.   With a bit of coding I changed the expert consensus to the second tie getter on a extracted dataframe.  I than merged the 4000+ into my train.   Just getting things running now, but one issue was that I was using eeg_id and eeg_sub_id later for some grouping.  My 4000+ addition were duplicates of course and messed up those results.   So I added 1000 to the eeg_sub_id before I merged the extracted dataframe into train.\n\nI am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the expert_consenus.  Will see if it helps those models.",
    "2665957": "Thank you for reply.\nI didn’t know that.\nI was not curious about “expert consensus” because our making model predicts probability.",
    "2666094": "Thank you for reply.\n\n「I am mostly using the probabilities for my models rather than the expert_consenus but a couple of models do use the 　expert_consenus. Will see if it helps those models.」\n→I was thinking same things. Basically, in this competition, we use probabilities that model output, but I was thinking using 'expert_consenus' helps.\nAnd, I was thinking of trying to make a bit change to targets probabilities which has tie votes based on expert_consenus.",
    "2666289": "In reality, there are 5 IIIC pattern classes not 6, the 'other' class (non-IIIC patterns) contains segments that were noisy and too poor in quality to classify in one of the 5 classes. you can read more about the process of annotation [here](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10136013/).\nIt would be interesting If you can find a tie between two of the five classes except others"
  },
  "source": "meta"
}