{
  "id": 480417,
  "title": "Data exploration: to sum up ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/480417",
  "author_name": "TheEventHorizons",
  "post_date": "2024-02-28T12:59:28.782000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In this notebook <a href=\"https://www.kaggle.com/code/theeventhorizons/eda-hms\" target=\"_blank\">https://www.kaggle.com/code/theeventhorizons/eda-hms</a>, I conducted a thorough analysis of the \"train.csv\" dataset and some particular eegs and spectrograms from train_eegs and train_spectrograms files, covering essential aspects such as target variable distribution, identifier significance, and inter-variable relationships. The absence of missing values ensures data reliability. This brief overview sets the stage for a deeper dive into the dataset's complexities.</p>\n<p>Additionally, I incorporated the work of <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a> from their notebook <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook</a>, which I believe significantly contributes to our understanding of the data.</p>\n<p>Below is a summary of the data exploration conducted:</p>\n<h1>1. train.csv:</h1>\n<p><strong>Initial Form Analysis</strong>:</p>\n<ul>\n<li>Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)</li>\n<li>Rows and Columns: (106800, 15)</li>\n<li>Types of Variables: 12 int64, 2 float64, 1 object</li>\n<li>Analysis of Missing Variables: No missing value</li>\n</ul>\n<p><strong>Initial Background Analysis</strong>:</p>\n<ul>\n<li><p>Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000</p></li>\n<li><p>Significance of Variables:</p>\n<ul>\n<li>label_id: Nothing to say since it's a unique number representing each cases in the dataset</li>\n<li>eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.</li>\n<li>spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  </li>\n<li>patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  </li></ul></li>\n<li><p>Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).</p></li>\n<li><p>Relationship Variables/Target: </p>\n<ul>\n<li>patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).</li>\n<li>eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.</li>\n<li>spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds <br>\nbut also the eeg_id we look at.</li></ul></li>\n<li><p>Relationship Variables/variables: </p>\n<ul>\n<li>eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.</li>\n<li>eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.</li></ul></li>\n</ul>\n<p><strong>detailled Analysis (from <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a>)</strong></p>\n<ul>\n<li>It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)</li>\n<li>Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.</li>\n<li>Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.</li>\n<li>The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.<ul>\n<li>Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators</li>\n<li>gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.</li>\n<li>Conclusion: Stuies relying only on the 'first' may miss valuable information. </li></ul></li>\n<li>Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.</li>\n<li>Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.</li>\n<li>The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.</li>\n</ul>\n<h1>2. train_eegs (for a particular eeg):</h1>\n<p>​<br>\n <strong>Form Analysis</strong>:<br>\n​</p>\n<ul>\n<li>Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. </li>\n<li>Types of Variables: 20 float64</li>\n<li>Analysis of Missing Variables: No missing value<br>\n​<br>\n<strong>Background Analysis</strong>:<br>\n​</li>\n<li>Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.<br>\n​</li>\n<li>Relationship Variables/variables: <ul>\n<li>Variables seem to be generally correlated according to the distance in a defined montage</li>\n<li>Correlations between variables change according to the label_id</li></ul></li>\n</ul>\n<h1>3. train_spectrograms (for a particular spectrogram):</h1>\n<p><strong>Form Analysis</strong>:</p>\n<ul>\n<li>Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. </li>\n<li>Types of Variables: 400 float32, 1 int64</li>\n<li>Analysis of Missing Variables: No missing value</li>\n</ul>\n<p><strong>Background Analysis</strong>:</p>\n<ul>\n<li><p>Significance of Variables:</p>\n<ul>\n<li>Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.</li>\n<li>Observe that different spectrograms_sub_id with the same id coincide for a large part </li></ul></li>\n<li><p>Relationship Variables/variables: </p>\n<ul>\n<li>Variables RL and RP with the same suffix number seem to be generally highly correlated (&gt;88%) </li>\n<li>Correlations between variables change according to the label_id</li></ul></li>\n</ul>\n<p>Thank you for reading !  feel free to comment !</p>",
  "messages": [
    {
      "id": 2673035,
      "postDate": "2024-02-28T12:59:28.783Z",
      "content": "<p>In this notebook <a href=\"https://www.kaggle.com/code/theeventhorizons/eda-hms\" target=\"_blank\">https://www.kaggle.com/code/theeventhorizons/eda-hms</a>, I conducted a thorough analysis of the \"train.csv\" dataset and some particular eegs and spectrograms from train_eegs and train_spectrograms files, covering essential aspects such as target variable distribution, identifier significance, and inter-variable relationships. The absence of missing values ensures data reliability. This brief overview sets the stage for a deeper dive into the dataset's complexities.</p>\n<p>Additionally, I incorporated the work of <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a> from their notebook <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook\" target=\"_blank\">https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook</a>, which I believe significantly contributes to our understanding of the data.</p>\n<p>Below is a summary of the data exploration conducted:</p>\n<h1>1. train.csv:</h1>\n<p><strong>Initial Form Analysis</strong>:</p>\n<ul>\n<li>Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)</li>\n<li>Rows and Columns: (106800, 15)</li>\n<li>Types of Variables: 12 int64, 2 float64, 1 object</li>\n<li>Analysis of Missing Variables: No missing value</li>\n</ul>\n<p><strong>Initial Background Analysis</strong>:</p>\n<ul>\n<li><p>Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000</p></li>\n<li><p>Significance of Variables:</p>\n<ul>\n<li>label_id: Nothing to say since it's a unique number representing each cases in the dataset</li>\n<li>eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.</li>\n<li>spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  </li>\n<li>patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  </li></ul></li>\n<li><p>Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).</p></li>\n<li><p>Relationship Variables/Target: </p>\n<ul>\n<li>patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).</li>\n<li>eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.</li>\n<li>spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds <br>\nbut also the eeg_id we look at.</li></ul></li>\n<li><p>Relationship Variables/variables: </p>\n<ul>\n<li>eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.</li>\n<li>eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.</li></ul></li>\n</ul>\n<p><strong>detailled Analysis (from <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a>)</strong></p>\n<ul>\n<li>It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)</li>\n<li>Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.</li>\n<li>Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.</li>\n<li>The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.<ul>\n<li>Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators</li>\n<li>gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.</li>\n<li>Conclusion: Stuies relying only on the 'first' may miss valuable information. </li></ul></li>\n<li>Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.</li>\n<li>Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.</li>\n<li>The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.</li>\n</ul>\n<h1>2. train_eegs (for a particular eeg):</h1>\n<p>​<br>\n <strong>Form Analysis</strong>:<br>\n​</p>\n<ul>\n<li>Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. </li>\n<li>Types of Variables: 20 float64</li>\n<li>Analysis of Missing Variables: No missing value<br>\n​<br>\n<strong>Background Analysis</strong>:<br>\n​</li>\n<li>Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.<br>\n​</li>\n<li>Relationship Variables/variables: <ul>\n<li>Variables seem to be generally correlated according to the distance in a defined montage</li>\n<li>Correlations between variables change according to the label_id</li></ul></li>\n</ul>\n<h1>3. train_spectrograms (for a particular spectrogram):</h1>\n<p><strong>Form Analysis</strong>:</p>\n<ul>\n<li>Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. </li>\n<li>Types of Variables: 400 float32, 1 int64</li>\n<li>Analysis of Missing Variables: No missing value</li>\n</ul>\n<p><strong>Background Analysis</strong>:</p>\n<ul>\n<li><p>Significance of Variables:</p>\n<ul>\n<li>Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.</li>\n<li>Observe that different spectrograms_sub_id with the same id coincide for a large part </li></ul></li>\n<li><p>Relationship Variables/variables: </p>\n<ul>\n<li>Variables RL and RP with the same suffix number seem to be generally highly correlated (&gt;88%) </li>\n<li>Correlations between variables change according to the label_id</li></ul></li>\n</ul>\n<p>Thank you for reading !  feel free to comment !</p>",
      "rawMarkdown": "In this notebook https://www.kaggle.com/code/theeventhorizons/eda-hms, I conducted a thorough analysis of the \"train.csv\" dataset and some particular eegs and spectrograms from train_eegs and train_spectrograms files, covering essential aspects such as target variable distribution, identifier significance, and inter-variable relationships. The absence of missing values ensures data reliability. This brief overview sets the stage for a deeper dive into the dataset's complexities.\n\nAdditionally, I incorporated the work of @pcjimmy from their notebook https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook, which I believe significantly contributes to our understanding of the data.\n\nBelow is a summary of the data exploration conducted:\n\n\n# 1. train.csv:\n\n**Initial Form Analysis**:\n\n- Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)\n- Rows and Columns: (106800, 15)\n- Types of Variables: 12 int64, 2 float64, 1 object\n- Analysis of Missing Variables: No missing value\n\n**Initial Background Analysis**:\n\n- Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000\n\n- Significance of Variables:\n    * label_id: Nothing to say since it's a unique number representing each cases in the dataset\n    * eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.\n    * spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  \n    * patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  \n    \n    \n- Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).\n\n- Relationship Variables/Target: \n    * patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).\n    * eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.\n    * spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds \n      but also the eeg_id we look at.\n      \n- Relationship Variables/variables: \n    * eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.\n    * eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.\n\n\n**detailled Analysis (from @pcjimmy)**\n\n- It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)\n- Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.\n- Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.\n- The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.\n    -  Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators\n    - gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.\n    - Conclusion: Stuies relying only on the 'first' may miss valuable information. \n- Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.\n- Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.\n- The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.\n \n\n# 2. train_eegs (for a particular eeg):\n​\n **Form Analysis**:\n​\n- Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. \n- Types of Variables: 20 float64\n- Analysis of Missing Variables: No missing value\n​\n**Background Analysis**:\n​\n- Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.\n​\n- Relationship Variables/variables: \n    * Variables seem to be generally correlated according to the distance in a defined montage\n    * Correlations between variables change according to the label_id\n\n\n# 3. train_spectrograms (for a particular spectrogram):\n\n**Form Analysis**:\n\n- Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. \n- Types of Variables: 400 float32, 1 int64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables:\n    * Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.\n    * Observe that different spectrograms_sub_id with the same id coincide for a large part \n\n- Relationship Variables/variables: \n    * Variables RL and RP with the same suffix number seem to be generally highly correlated (>88%) \n    * Correlations between variables change according to the label_id\n\n\nThank you for reading !  feel free to comment !",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2673035": "In this notebook https://www.kaggle.com/code/theeventhorizons/eda-hms, I conducted a thorough analysis of the \"train.csv\" dataset and some particular eegs and spectrograms from train_eegs and train_spectrograms files, covering essential aspects such as target variable distribution, identifier significance, and inter-variable relationships. The absence of missing values ensures data reliability. This brief overview sets the stage for a deeper dive into the dataset's complexities.\n\nAdditionally, I incorporated the work of @pcjimmy from their notebook https://www.kaggle.com/code/pcjimmmy/patient-variation-eda/notebook, which I believe significantly contributes to our understanding of the data.\n\nBelow is a summary of the data exploration conducted:\n\n\n# 1. train.csv:\n\n**Initial Form Analysis**:\n\n- Target Variable: (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote)\n- Rows and Columns: (106800, 15)\n- Types of Variables: 12 int64, 2 float64, 1 object\n- Analysis of Missing Variables: No missing value\n\n**Initial Background Analysis**:\n\n- Target Visualization: The target variable is well balanced. Number of expert consensus for each type of activity is between 15000 and 2000\n\n- Significance of Variables:\n    * label_id: Nothing to say since it's a unique number representing each cases in the dataset\n    * eeg_id: There are 17089 different eeg_id. Each eeg_id contains eeg_sub_id which can go from 0 to 742 associated to a eeg_label_offset_seconds which goes from 0.0 to 3372.0. Thus the occurency of a eeg_id goes from 1 to 743. A eeg_id is associated to a unique spectrogram_id.\n    * spectrogram_id: There are 11138 different spectrogram_id Each spectrogram_id contains spectrogram_sub_id which can go from 0 to 1021 associated to a spectrogram_label_offset_seconds which goes from 0.0 to 17632.0?. Thus the occurency of a spectrogram_id goes from 1 to 1022. A spectrogram_id can contain different eeg_id  \n    * patient_id: There are 1950 different patient_id. The occurrence is between 1 and 2215 corresponding to different combination of ('eeg_id','eeg_sub_id', 'eeg_label_offset_seconds', 'spectrogram_id', 'spectrogram_sub_id', 'spectrogram_label_offset_seconds').  \n    \n    \n- Relationship Target/target: it seems that seizure is correlated with grda_vote (24%), other_vote (21%), lrda_vote (17%), gpd_vote (13%), lpd_vote (13%). lrda_vote and gpd_vote (15%). The rest seems to be negligeable (-8%).\n\n- Relationship Variables/Target: \n    * patient_id/expert_consensus: Each patient_id can have more than one expert_consensus corresponding to a particular configuration (seizure_vote, lpd_vote, gpd_vote, lrda_vote, grda_vote, other_vote).\n    * eeg_id, eeg_sub_id, eeg_label_offset_seconds/expert_consensus: Each eeg_id can have different expert consensus based on the eeg_sub_id or eeg_label_offset_seconds we look at.\n    * spectrogram_id, spectrogram_sub_id, spectrogram_label_offset_seconds/expert_consensus: Each spectrogram_id can have different expert consensus based on spectrogram_sub_id, spectrogram_label_offset_seconds \n      but also the eeg_id we look at.\n      \n- Relationship Variables/variables: \n    * eeg_id/spectrogram_id: A eeg_id is associated to a unique spectrogram_id but a spectrogram_id can contain different eeg_id.\n    * eeg_id/patient_id and spectrogram_id/patient_id: Each patient can have several different eeg_id and spectrogram_id.\n\n\n**detailled Analysis (from @pcjimmy)**\n\n- It appears that the training data may have been aggregated from at least two different studies, with varying numbers of evaluators participating the first study has less than 6 voters whereas the second one has more that 9 voters. (Considering the weight of the data, those with more evaluators should carry more significance, particularly if a strong consensus is observed)\n- Rows with a small number of evaluators often receive perfect ratings (1.0), raising questions about their reliability.\n- Rows with very low agreement on consensus (20% or less) present a significant challenge. The decision on how to handle such problematic cases remains uncertain.\n- The rating difficulty varies, with 'seizure_vote' appearing easier to rate compared to 'gpd_vote,' hinting at potential differences in the studies contributing to the training data.\n    -  Seizure_vote - Patients with this condition either exhibit consistent EEG patterns or their symptoms are easily recognizable and agreeable among evaluators\n    - gpd_vote: Patients with this condition either do not display the issue consistently over time (it comes and goes), or it is challenging for evaluators to identify. The plot suggests that there are no EEG recordings for this type where all evaluators unanimously agreed.\n    - Conclusion: Stuies relying only on the 'first' may miss valuable information. \n- Descriptions of eegs where experts agree are termed \"idealized.\" However, questions arise about the confidence in these idealized descriptions when the number of experts is limited. Similarly, when only a single eeg is available for a patient, even with agreement from many experts, the confidence in labeling it as 'ideal' is questioned.\n- Analyzing the distribution of ratings reveals unexpected results. For Seizure, when the number of evaluators is six or fewer, it is the major label. However, with more than nine evaluators, it becomes the least observed. The small vs. large group of evaluators presents a challenge, with the predominance of 'other' in larger groups potentially indicating measurement error or different study characteristics.\n- The issue of outliers, especially in the Seizure category, prompts consideration for removal. Additionally, the expected agreement for 'lateral' is not observed, challenging assumptions about the reliability of right vs. left eeg distributions.\n \n\n# 2. train_eegs (for a particular eeg):\n​\n **Form Analysis**:\n​\n- Rows and Columns: (10000, 20) + (0,1) to facilitate our comprehension, corresponding the 'eeg_label_offset_seconds' column. Each combination ('eeg_id', 'eeg_sub_id') corresponds to a 50 second long subsample starting at time 'eeg_label_offset_seconds' where 200 samples were taken each second. \n- Types of Variables: 20 float64\n- Analysis of Missing Variables: No missing value\n​\n**Background Analysis**:\n​\n- Significance of Variables: Each column represents a measure done by a particular electrode placed on the head of the patient.\n​\n- Relationship Variables/variables: \n    * Variables seem to be generally correlated according to the distance in a defined montage\n    * Correlations between variables change according to the label_id\n\n\n# 3. train_spectrograms (for a particular spectrogram):\n\n**Form Analysis**:\n\n- Rows and Columns: (300, 401). Each combination ('spectrograms_id', 'spectrograms_sub_id') corresponds to a 600 seconds (ten minutes) long subsample starting at time 'spectrograms_label_offset_seconds' where 1 sample was taken each 2 seconds. \n- Types of Variables: 400 float32, 1 int64\n- Analysis of Missing Variables: No missing value\n\n**Background Analysis**:\n\n- Significance of Variables:\n    * Each column signifies a measurement conducted by a specific group of electrodes (LL, LP, RL, RP) positioned at a distinct location on the patient's head, along with a corresponding time column. Each suffix in the column header corresponds to a specific frequency of brain activity.\n    * Observe that different spectrograms_sub_id with the same id coincide for a large part \n\n- Relationship Variables/variables: \n    * Variables RL and RP with the same suffix number seem to be generally highly correlated (>88%) \n    * Correlations between variables change according to the label_id\n\n\nThank you for reading !  feel free to comment !"
  }
}