{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30636,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### In this notebook, I review the degree to which the experts agree on their identification of EEG patterns for the different patterns--i.e., if there are differences in the level of agreement depending on the specific pattern.","metadata":{}},{"cell_type":"code","source":"\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport plotly.graph_objs as go\nimport pandas as pd\nimport plotly.express as px\n","metadata":{"tags":[],"execution":{"iopub.status.busy":"2024-01-18T16:38:47.523236Z","iopub.execute_input":"2024-01-18T16:38:47.523972Z","iopub.status.idle":"2024-01-18T16:38:57.203859Z","shell.execute_reply.started":"2024-01-18T16:38:47.523923Z","shell.execute_reply":"2024-01-18T16:38:57.202157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT = Path.cwd().parent\nINPUT = ROOT / \"input\"\nOUTPUT = ROOT / \"output\"\nSRC = ROOT / \"src\"\n\nDATA = INPUT / \"hms-harmful-brain-activity-classification\"\n\n\nCLASSES = [\"seizure_vote\", \"lpd_vote\", \"gpd_vote\", \"lrda_vote\", \"grda_vote\", \"other_vote\"]\nN_CLASSES = len(CLASSES)\nFOLDS = [0, 1, 2, 3, 4]\nN_FOLDS = len(FOLDS)","metadata":{"tags":[],"execution":{"iopub.status.busy":"2024-01-18T16:39:52.306749Z","iopub.execute_input":"2024-01-18T16:39:52.307531Z","iopub.status.idle":"2024-01-18T16:39:52.316472Z","shell.execute_reply.started":"2024-01-18T16:39:52.307489Z","shell.execute_reply":"2024-01-18T16:39:52.314996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\n\n# convert vote to probability\ntrain[CLASSES] /= train[CLASSES].sum(axis=1).values[:, None]\n\n","metadata":{"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Calculate a confidence score for each eeg, based on the mean probabiltiy for each pattern across the time window.","metadata":{}},{"cell_type":"code","source":"#get expert consensus into this. in order to graph expert consensus vs selected based on confidence\nec=train[['eeg_id', 'expert_consensus']].drop_duplicates()\n\nconfidences=train.groupby('eeg_id', as_index=False)[[\"seizure_vote\", \"lpd_vote\", \"gpd_vote\", \"lrda_vote\", \"grda_vote\", \"other_vote\"]].agg('mean')\n\ncs=[\"seizure_\", \"lpd_\", \"gpd_\", \"lrda_\", \"grda_\", \"other_\"]\nfor c in cs:\n    confidences.rename({f\"{c}vote\": f\"{c}confidence\"}, axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:02.592582Z","iopub.execute_input":"2024-01-18T17:55:02.593034Z","iopub.status.idle":"2024-01-18T17:55:02.640453Z","shell.execute_reply.started":"2024-01-18T17:55:02.593001Z","shell.execute_reply":"2024-01-18T17:55:02.639371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Identify eegs with more than one expert consensus, i.e., those where the expert consensus changes in movement through the time window.","metadata":{}},{"cell_type":"code","source":"\ndef get_expert_consensus_dups(row):\n    confidence_values = row[['seizure_confidence', 'gpd_confidence', 'grda_confidence', 'lpd_confidence', 'lrda_confidence', 'other_confidence']]\n    max_confidence = confidence_values.max()\n    \n    if max_confidence == row['seizure_confidence']:\n        return 'Seizure'\n    elif max_confidence == row['gpd_confidence']:\n        return 'GPD'\n    elif max_confidence == row['grda_confidence']:\n        return 'GRDA'\n    elif max_confidence == row['lpd_confidence']:\n        return 'LPD'\n    elif max_confidence == row['lrda_confidence']:\n        return 'LRDA'\n    elif max_confidence == row['other_confidence']:\n        return 'Other'\n    else:\n        return 'No consensus'\n\n\n\n\n\n#resolve dups\ndups=ec.groupby(\"eeg_id\", as_index=False)['expert_consensus'].agg(\"count\").sort_values(\"expert_consensus\", ascending=False)\ndups=dups.query('expert_consensus>1')\ndups=dups.merge(confidences, on='eeg_id')\ndups['max_confidence']=dups[['seizure_confidence', 'lpd_confidence', 'gpd_confidence',\\\n        'lrda_confidence', 'grda_confidence', 'other_confidence']].max(axis=1)\ndups['expert_consensus'] = dups.apply(get_expert_consensus_dups, axis=1)\ndups['multiple_consensus']=1","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:04.992956Z","iopub.execute_input":"2024-01-18T17:55:04.993445Z","iopub.status.idle":"2024-01-18T17:55:05.541882Z","shell.execute_reply.started":"2024-01-18T17:55:04.993409Z","shell.execute_reply":"2024-01-18T17:55:05.540512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"expert_consensus=train[['eeg_id', 'expert_consensus']].drop_duplicates()\ndup_eeg_ids=dups.eeg_id\nexpert_consensus=expert_consensus.query(\"eeg_id not in @dup_eeg_ids\")","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:09.303938Z","iopub.execute_input":"2024-01-18T17:55:09.304467Z","iopub.status.idle":"2024-01-18T17:55:09.340612Z","shell.execute_reply.started":"2024-01-18T17:55:09.304425Z","shell.execute_reply":"2024-01-18T17:55:09.339373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confidences=confidences.merge(expert_consensus, on ='eeg_id', how='left')\ndups.rename({'expert_consensus':'expert_consensus_dup'},axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:11.276539Z","iopub.execute_input":"2024-01-18T17:55:11.278296Z","iopub.status.idle":"2024-01-18T17:55:11.297443Z","shell.execute_reply.started":"2024-01-18T17:55:11.278209Z","shell.execute_reply":"2024-01-18T17:55:11.295311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confidences=confidences.merge(dups[['eeg_id','expert_consensus_dup', 'multiple_consensus']],\n                  on='eeg_id', how='left')\nconfidences['expert_consensus']=confidences.apply(lambda x: x['expert_consensus_dup'] \\\n                if x['multiple_consensus']==1 else x['expert_consensus'], axis=1 )","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:14.624580Z","iopub.execute_input":"2024-01-18T17:55:14.625161Z","iopub.status.idle":"2024-01-18T17:55:14.939161Z","shell.execute_reply.started":"2024-01-18T17:55:14.625113Z","shell.execute_reply":"2024-01-18T17:55:14.937681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"assert(len(confidences)==len(np.unique(train.eeg_id)))\nassert(len(confidences.query('expert_consensus.isnull()', engine='python'))==0)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T17:55:16.675664Z","iopub.execute_input":"2024-01-18T17:55:16.676129Z","iopub.status.idle":"2024-01-18T17:55:16.692844Z","shell.execute_reply.started":"2024-01-18T17:55:16.676092Z","shell.execute_reply":"2024-01-18T17:55:16.691503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Below shows the distribution of consensus opinions, with cases where multiple consensus labels assigned to the same EEG are included.","metadata":{}},{"cell_type":"code","source":"consensus=train[['eeg_id', 'expert_consensus']].drop_duplicates()\nconsensus= consensus.groupby('expert_consensus', as_index=False)\\\n    .agg('count') \\\n    .rename({'eeg_id':'count'}, axis=1)\n\n\n\ntotal_count = consensus['count'].sum()\n\npercentages = (consensus['count'] / total_count) * 100\n\n\nplt.figure(figsize=(8, 6))\nbars = plt.bar(consensus['expert_consensus'], consensus['count'], color=['red', 'green', 'blue', 'purple', 'orange', 'pink'])\nplt.xlabel(\"Expert Consensus\")\nplt.ylabel(\"Count\")\n\n\nfor bar, percentage in zip(bars, percentages):\n    plt.annotate(f'{percentage:.2f}%', (bar.get_x() + bar.get_width() / 2, bar.get_height()), ha='center', va='bottom')\n\nplt.title(\"EEG Expert Consensus Frequency\")\nplt.xticks(rotation=45, ha=\"right\")  \n# Show the plot\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-01-18T18:37:45.962211Z","iopub.execute_input":"2024-01-18T18:37:45.962697Z","iopub.status.idle":"2024-01-18T18:37:46.375738Z","shell.execute_reply.started":"2024-01-18T18:37:45.962656Z","shell.execute_reply":"2024-01-18T18:37:46.374733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### And below shows the same with de-duping applied.  For EEGs with multiple expert consensus labels, the one with the higher confidence/mean probability is chosen.\n\n#### Since there does not appear to be a significant difference in the frequency, perhaps we can conclude that it is safe to apply some method of de-duplication when preparing the data for training.","metadata":{}},{"cell_type":"code","source":"confidences_agg=confidences[['eeg_id', 'expert_consensus']].drop_duplicates()\nconfidences_agg=confidences_agg[['eeg_id', 'expert_consensus']].drop_duplicates()\nconfidences_agg= confidences_agg.groupby('expert_consensus', as_index=False)\\\n    .agg('count') \\\n    .rename({'eeg_id':'count'}, axis=1)\n\n\n\ntotal_count = confidences_agg['count'].sum()\n\npercentages = (confidences_agg['count'] / total_count) * 100\n\n\nplt.figure(figsize=(8, 6))\nbars = plt.bar(confidences_agg['expert_consensus'], confidences_agg['count'], color=['red', 'green', 'blue', 'purple', 'orange', 'pink'])\nplt.xlabel(\"Expert Consensus\")\nplt.ylabel(\"Count\")\n\n\nfor bar, percentage in zip(bars, percentages):\n    plt.annotate(f'{percentage:.2f}%', (bar.get_x() + bar.get_width() / 2, bar.get_height()), ha='center', va='bottom')\n\nplt.title(\"EEG Expert Consensus Frequency with Dedup\")\nplt.xticks(rotation=45, ha=\"right\") \n\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T18:42:45.319828Z","iopub.execute_input":"2024-01-18T18:42:45.320955Z","iopub.status.idle":"2024-01-18T18:42:45.796733Z","shell.execute_reply.started":"2024-01-18T18:42:45.320908Z","shell.execute_reply":"2024-01-18T18:42:45.795489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_confidence(row):\n    #consensus = row[['expert_consensus']]\n   \n    \n    if row['expert_consensus'] == 'Seizure':\n        return row['seizure_confidence'] \n    elif row['expert_consensus'] == 'GPD':\n        return row['gpd_confidence'] \n    elif row['expert_consensus'] == 'LPD':\n        return row['lpd_confidence'] \n    elif row['expert_consensus'] == 'LRDA':\n        return row['lrda_confidence'] \n    elif row['expert_consensus'] == 'GRDA':\n        return row['grda_confidence']\n    elif row['expert_consensus'] == 'Other':\n        return row['other_confidence']\n    else:\n        return np.NaN\n        \nconfidences['confidence'] = confidences.apply(get_confidence, axis=1)\nassert(len(confidences.query(\"confidence.isnull()\", engine='python'))==0)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-01-18T18:47:28.084769Z","iopub.execute_input":"2024-01-18T18:47:28.085248Z","iopub.status.idle":"2024-01-18T18:47:28.652402Z","shell.execute_reply.started":"2024-01-18T18:47:28.085210Z","shell.execute_reply":"2024-01-18T18:47:28.651370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Below uses the cleaned data to review differences in degree of consensus based on EEG pattern.  The data points represent rounding to the nearest .05; e.g., .90 represents the percentage of confidence values greater or equal to  87.5 and less than 92.5.\n\n#### We see that level of agreeement among experts varies based on the pattern, with better agreement on the labels Seizure, GRDA, and Other, compared to LPD, GPD, and LRDA.","metadata":{}},{"cell_type":"code","source":"\nconfidence_range = [0.3, .35,0.4,.45, 0.5,.55, 0.6,.65, 0.7,.75, 0.8,.85, 0.9,.95, 1.0]\ntick_separation = 0.05\n\n\nplt.figure(figsize=(10, 6))\n\n\nfor consensus_value in confidences['expert_consensus'].unique():\n    subset = confidences[confidences['expert_consensus'] == consensus_value]\n    \n    # Calculate the percentage of confidence values for each confidence range\n    percentages = []\n    for confidence in confidence_range:\n        percentage = (subset['confidence'] >= confidence - tick_separation / 2) & (subset['confidence'] < confidence + tick_separation / 2)\n        \n        percentages.append((percentage.sum() / len(subset)) * 100)\n\n    plt.plot(confidence_range, percentages, marker='o', linestyle='-', label=consensus_value)\n\nplt.xlabel(\"Confidence\")\nplt.ylabel(\"Percentage\")\nplt.title(\"Confidence levels for Each EEG Pattern Label\")\nplt.grid(True)\n\n# Show the plot with legend\nplt.tight_layout()\nplt.xticks(confidence_range)\nplt.legend(loc='best')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-01-18T19:16:00.618136Z","iopub.execute_input":"2024-01-18T19:16:00.619453Z","iopub.status.idle":"2024-01-18T19:16:01.304729Z","shell.execute_reply.started":"2024-01-18T19:16:00.619394Z","shell.execute_reply":"2024-01-18T19:16:01.302752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Below is a different view of the same finding, showing percentage of confidences scores for each label, using confidence score ranges of .10.","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objs as go\nimport pandas as pd\nbar_segment_list=[]\n\n\n\nconfidence_ranges = [(0.2, 0.3), (0.3, 0.4), (0.4, 0.5), (0.5, 0.6), (0.6, 0.7), (0.7, 0.8), (0.8, 0.9), (0.9, 1.0)]\nconfidence_range_labels = ['0.2-0.3', '0.3-0.4', '0.4-0.5', '0.5-0.6', '0.6-0.7', '0.7-0.8', '0.8-0.9', '0.9-1.0']\ncolors = {}\ni=0\nshow_legend={}\nfor l in confidence_ranges:\n    colors[l]=i\n    show_legend[l]=True\n    i+=1\n\nfig = go.Figure()\n\n\nfor confidence_range in confidence_ranges:\n    range_start, range_end = confidence_range\n\n\n    bar_segments = []\n\n    for consensus_value in confidences['expert_consensus'].unique():\n        subset = confidences[(confidences['expert_consensus'] == consensus_value) & (confidences['confidence'] >range_start) & (confidences['confidence'] <= range_end)]\n\n        percentage = (len(subset) / len(confidences[confidences['expert_consensus'] == consensus_value])) * 100\n  \n        bar_segments.append(go.Bar(\n            x=[consensus_value],\n            y=[percentage],\n            name=confidence_range_labels[confidence_ranges.index(confidence_range)],#consensus_value,\n            #marker=dict(color=px.colors.qualitative.Set3\\\n                #[confidences['expert_consensus'].unique().tolist().index(consensus_value)]),\n            marker=dict(color=px.colors.qualitative.Set1[colors[confidence_range]]),\n            hoverinfo='y+name',  # Display percentage and expert consensus label on hover\n            text=f'{percentage:.2f}%',  # Text for hover\n            showlegend=show_legend[confidence_range], \n        ))\n        show_legend[confidence_range] =False\n    \n    fig.add_trace(go.Bar(\n        x=confidences['expert_consensus'].unique(),\n        y=[0], #* len(confidences['expert_consensus'].unique()),  # Y values for the main bar (all zeros)\n        name=confidence_range_labels[confidence_ranges.index(confidence_range)],\n        hoverinfo='skip',  # Skip hover for the main bar\n        showlegend=False,  # Hide the legend for the main bar\n    ))\n\n \n    for bar_segment in bar_segments:\n        fig.add_trace(bar_segment)\n\nfig.update_layout(\n    xaxis_title=\"Expert Consensus\",\n \n    yaxis = go.layout.YAxis(\n        title = 'Confidence Range Frequency', showticklabels=False\n    ),\n    title=\"Consensus Confidence by EEG Pattern Label\",\n    barmode='stack', \n)\n\n\nfig.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-01-18T19:16:21.250674Z","iopub.execute_input":"2024-01-18T19:16:21.251097Z","iopub.status.idle":"2024-01-18T19:16:21.786137Z","shell.execute_reply.started":"2024-01-18T19:16:21.251063Z","shell.execute_reply":"2024-01-18T19:16:21.784865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Next steps:  I have no expertise in this area so I am not sure what to do with the above finding. It also doesn't account for the requirement to predict the event occurring in the middle 10 seconds of thE EEG/spectrogram time windows, see https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010 for details. But perhaps it means that eeg data with higher levels of consensus is of better quality, and eegs with relatively lower consensus can be excluded from training or given less weight.","metadata":{}}]}