{
  "id": 477893,
  "title": "Spectogram Subsections - no vote shifts over 500+ subsections",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477893",
  "author_name": "Raki",
  "post_date": "2024-02-18T10:32:57.990000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>I wanted to evaluate the consistency in spectograms/EEG and came up with the approach of taking the spectograms with most sub_ids and tracking how often the vote-distribution shifts in them. </p>\n<p>I noticed a curious pattern: <br>\nSome don’t have shifts even for 700+ subsections and some shift hundreds of times. </p>\n<p>The corresponding notebook can be found <a href=\"https://www.kaggle.com/raki21/eda-spectogram-subsections\" target=\"_blank\">here</a>.</p>\n<h1>Interpretation</h1>\n<p>At first I thought of 2 interpretations for this lag of shifts in votes: </p>\n<ol>\n<li>some spectograms are just very consistent.</li>\n<li>the expert evaluators don’t change their prediction often and changes in predictions come mainly from changes in evaluators.</li>\n</ol>\n<p>To get a lower bound for changes in evaluators I tracked when the total number of votes changes. It is a lower bound because there could also be a situation that the evaluators change at but the number of experts is the same. We don’t have a way to find these switches. </p>\n<p>What you can see is that almost all shifts appear to be due to evaluator switches.<br>\nOf the 100 spectograms with most subsections, 47 don’t switch the number of evaluators and of those 45 also don’t shift the vote_distribution. =&gt; It appears that experts tend to stick to their first vote across a spectogram. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F714dd6415871f3bc87a79db33517fd2d%2FShifts%20and%20Switches.png?generation=1708252002277438&amp;alt=media\"></p>\n<p>I came up with one additional explanation (3.), probably all of the explanations contribute to the lag of vote shifts to some extent.</p>\n<ol>\n<li>patterns tend to be pretty similar across some spectograms.</li>\n<li>experts have some aversion of changing their vote.</li>\n<li>the organizers tend to give low change spectograms to the same experts, but high change to different experts every few subsections. </li>\n</ol>\n<p>From this we might be able to get an approximation of noise from errors in expert evaluation by seeing how much the evaluation changes across different evaluators, we can for example take a set with 200 switches and get KL-divergence between predictions every 5 switches! I also did this in my EDA but am not confident this captures the noise that well.</p>",
  "messages": [
    {
      "id": 2657140,
      "postDate": "2024-02-18T10:32:57.990Z",
      "content": "<h1>Introduction</h1>\n<p>I wanted to evaluate the consistency in spectograms/EEG and came up with the approach of taking the spectograms with most sub_ids and tracking how often the vote-distribution shifts in them. </p>\n<p>I noticed a curious pattern: <br>\nSome don’t have shifts even for 700+ subsections and some shift hundreds of times. </p>\n<p>The corresponding notebook can be found <a href=\"https://www.kaggle.com/raki21/eda-spectogram-subsections\" target=\"_blank\">here</a>.</p>\n<h1>Interpretation</h1>\n<p>At first I thought of 2 interpretations for this lag of shifts in votes: </p>\n<ol>\n<li>some spectograms are just very consistent.</li>\n<li>the expert evaluators don’t change their prediction often and changes in predictions come mainly from changes in evaluators.</li>\n</ol>\n<p>To get a lower bound for changes in evaluators I tracked when the total number of votes changes. It is a lower bound because there could also be a situation that the evaluators change at but the number of experts is the same. We don’t have a way to find these switches. </p>\n<p>What you can see is that almost all shifts appear to be due to evaluator switches.<br>\nOf the 100 spectograms with most subsections, 47 don’t switch the number of evaluators and of those 45 also don’t shift the vote_distribution. =&gt; It appears that experts tend to stick to their first vote across a spectogram. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F714dd6415871f3bc87a79db33517fd2d%2FShifts%20and%20Switches.png?generation=1708252002277438&amp;alt=media\"></p>\n<p>I came up with one additional explanation (3.), probably all of the explanations contribute to the lag of vote shifts to some extent.</p>\n<ol>\n<li>patterns tend to be pretty similar across some spectograms.</li>\n<li>experts have some aversion of changing their vote.</li>\n<li>the organizers tend to give low change spectograms to the same experts, but high change to different experts every few subsections. </li>\n</ol>\n<p>From this we might be able to get an approximation of noise from errors in expert evaluation by seeing how much the evaluation changes across different evaluators, we can for example take a set with 200 switches and get KL-divergence between predictions every 5 switches! I also did this in my EDA but am not confident this captures the noise that well.</p>",
      "rawMarkdown": "# Introduction \nI wanted to evaluate the consistency in spectograms/EEG and came up with the approach of taking the spectograms with most sub_ids and tracking how often the vote-distribution shifts in them. \n\nI noticed a curious pattern: \nSome don’t have shifts even for 700+ subsections and some shift hundreds of times. \n\nThe corresponding notebook can be found [here](https://www.kaggle.com/raki21/eda-spectogram-subsections).\n\n\n# Interpretation\nAt first I thought of 2 interpretations for this lag of shifts in votes: \n\n1. some spectograms are just very consistent.\n2. the expert evaluators don’t change their prediction often and changes in predictions come mainly from changes in evaluators.\n\nTo get a lower bound for changes in evaluators I tracked when the total number of votes changes. It is a lower bound because there could also be a situation that the evaluators change at but the number of experts is the same. We don’t have a way to find these switches. \n\nWhat you can see is that almost all shifts appear to be due to evaluator switches.\nOf the 100 spectograms with most subsections, 47 don’t switch the number of evaluators and of those 45 also don’t shift the vote_distribution. => It appears that experts tend to stick to their first vote across a spectogram. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F714dd6415871f3bc87a79db33517fd2d%2FShifts%20and%20Switches.png?generation=1708252002277438&alt=media)\n\nI came up with one additional explanation (3.), probably all of the explanations contribute to the lag of vote shifts to some extent.\n\n1. patterns tend to be pretty similar across some spectograms.\n2. experts have some aversion of changing their vote.\n3. the organizers tend to give low change spectograms to the same experts, but high change to different experts every few subsections. \n\nFrom this we might be able to get an approximation of noise from errors in expert evaluation by seeing how much the evaluation changes across different evaluators, we can for example take a set with 200 switches and get KL-divergence between predictions every 5 switches! I also did this in my EDA but am not confident this captures the noise that well.",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2657140": "# Introduction \nI wanted to evaluate the consistency in spectograms/EEG and came up with the approach of taking the spectograms with most sub_ids and tracking how often the vote-distribution shifts in them. \n\nI noticed a curious pattern: \nSome don’t have shifts even for 700+ subsections and some shift hundreds of times. \n\nThe corresponding notebook can be found [here](https://www.kaggle.com/raki21/eda-spectogram-subsections).\n\n\n# Interpretation\nAt first I thought of 2 interpretations for this lag of shifts in votes: \n\n1. some spectograms are just very consistent.\n2. the expert evaluators don’t change their prediction often and changes in predictions come mainly from changes in evaluators.\n\nTo get a lower bound for changes in evaluators I tracked when the total number of votes changes. It is a lower bound because there could also be a situation that the evaluators change at but the number of experts is the same. We don’t have a way to find these switches. \n\nWhat you can see is that almost all shifts appear to be due to evaluator switches.\nOf the 100 spectograms with most subsections, 47 don’t switch the number of evaluators and of those 45 also don’t shift the vote_distribution. => It appears that experts tend to stick to their first vote across a spectogram. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F714dd6415871f3bc87a79db33517fd2d%2FShifts%20and%20Switches.png?generation=1708252002277438&alt=media)\n\nI came up with one additional explanation (3.), probably all of the explanations contribute to the lag of vote shifts to some extent.\n\n1. patterns tend to be pretty similar across some spectograms.\n2. experts have some aversion of changing their vote.\n3. the organizers tend to give low change spectograms to the same experts, but high change to different experts every few subsections. \n\nFrom this we might be able to get an approximation of noise from errors in expert evaluation by seeing how much the evaluation changes across different evaluators, we can for example take a set with 200 switches and get KL-divergence between predictions every 5 switches! I also did this in my EDA but am not confident this captures the noise that well."
  }
}