{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392733,"sourceType":"datasetVersion","datasetId":4297749},{"sourceId":7392775,"sourceType":"datasetVersion","datasetId":4297782}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"papermill":{"default_parameters":{},"duration":323.289607,"end_time":"2024-01-13T13:38:46.622508","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-01-13T13:33:23.332901","version":"2.4.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Load Libraries","metadata":{"papermill":{"duration":0.007397,"end_time":"2024-01-13T13:33:27.152375","exception":false,"start_time":"2024-01-13T13:33:27.144978","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Further analysis indicates that two different distributions of labels exist dependent on the number of evaluators used.","metadata":{}},{"cell_type":"markdown","source":"The popular shared notebooks are only using the first row of data for each eeg-id.\n\nI believe this is an error and a mis-interpertation of the overview.   \n\nThis notebook shows that valuable information is contained when using all rows per eeg_id.","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd, numpy as np\nimport matplotlib.pyplot as plt\n\n","metadata":{"papermill":{"duration":0.843319,"end_time":"2024-01-13T13:33:28.003192","exception":false,"start_time":"2024-01-13T13:33:27.159873","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T03:19:45.462183Z","iopub.execute_input":"2024-02-12T03:19:45.462932Z","iopub.status.idle":"2024-02-12T03:19:46.697676Z","shell.execute_reply.started":"2024-02-12T03:19:45.462884Z","shell.execute_reply":"2024-02-12T03:19:46.696329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Train Data","metadata":{"papermill":{"duration":0.006939,"end_time":"2024-01-13T13:33:28.017754","exception":false,"start_time":"2024-01-13T13:33:28.010815","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\nTARGETS = df.columns[-6:]\nprint('df shape:', df.shape )\nprint('Targets', list(TARGETS))\ndf = df.sort_values(by=['patient_id', 'eeg_id', 'eeg_sub_id'])\ndf","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":0.323205,"end_time":"2024-01-13T13:33:28.347993","exception":false,"start_time":"2024-01-13T13:33:28.024788","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T03:19:46.700001Z","iopub.execute_input":"2024-02-12T03:19:46.7009Z","iopub.status.idle":"2024-02-12T03:19:47.102002Z","shell.execute_reply.started":"2024-02-12T03:19:46.700864Z","shell.execute_reply":"2024-02-12T03:19:47.100328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Row Accuracy\n\nIn 2003 when I was trained on the 6 sigma process, a standard starting point was to evaluate the measurement system.   Multiple readings on the same 'part' were made.\n\n\nMultiple rows of data are available for many of the eeg_id.  Lets use those looking at how often agreement for the 'expert_consensus' existed for each row.\n\n","metadata":{"papermill":{"duration":0.007773,"end_time":"2024-01-13T13:33:28.363641","exception":false,"start_time":"2024-01-13T13:33:28.355868","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Adding a new column 'total_evaluators' that sums up the six specified columns\ndf['total_evaluators'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].sum(axis=1)\n\ndf.sample(25)  # Display the DataFrame with the new column","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:47.113306Z","iopub.execute_input":"2024-02-12T03:19:47.117666Z","iopub.status.idle":"2024-02-12T03:19:47.24558Z","shell.execute_reply.started":"2024-02-12T03:19:47.117607Z","shell.execute_reply":"2024-02-12T03:19:47.244736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Plotting a histogram for the 'total_evaluators' column in the 'df' DataFrame\n\nplt.figure(figsize=(10, 6))\nplt.hist(df['total_evaluators'], bins=10, color='blue', edgecolor='black')\nplt.title('Histogram of Total Evaluators')\nplt.xlabel('Total Evaluators')\nplt.ylabel('Frequency')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:47.247099Z","iopub.execute_input":"2024-02-12T03:19:47.247726Z","iopub.status.idle":"2024-02-12T03:19:47.55792Z","shell.execute_reply.started":"2024-02-12T03:19:47.247696Z","shell.execute_reply":"2024-02-12T03:19:47.556577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It would seem that the training data might be aggregrated from at least two different studies with different number of persons doing the evaluations.   \n\nWith more evaluators it is likely data that should have more weight if there is good consensus.","metadata":{}},{"cell_type":"code","source":"# Modifying the previous code to add an additional column 'consensus_column' to 'df'\n\n# Finding the column with the largest number for each row and storing the value in 'consensus'\ndf['consensus'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].max(axis=1)\n\n# Identifying the column name that corresponds to the max value for each row\ndf['consensus_column'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].idxmax(axis=1)\n\ndf.head()  # Display the DataFrame with the new columns\n\n\n\ndf.sample(25)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:47.559498Z","iopub.execute_input":"2024-02-12T03:19:47.56037Z","iopub.status.idle":"2024-02-12T03:19:47.66485Z","shell.execute_reply.started":"2024-02-12T03:19:47.560339Z","shell.execute_reply":"2024-02-12T03:19:47.663496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create a new column that shows the percentage agreement\ndf['row_agreement'] = df['consensus']/df['total_evaluators']\ndf.sample(25)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:47.666117Z","iopub.execute_input":"2024-02-12T03:19:47.666477Z","iopub.status.idle":"2024-02-12T03:19:47.710389Z","shell.execute_reply.started":"2024-02-12T03:19:47.666449Z","shell.execute_reply":"2024-02-12T03:19:47.709104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('row_agreement.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:47.711805Z","iopub.execute_input":"2024-02-12T03:19:47.71225Z","iopub.status.idle":"2024-02-12T03:19:48.990665Z","shell.execute_reply.started":"2024-02-12T03:19:47.71222Z","shell.execute_reply":"2024-02-12T03:19:48.989433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting a histogram for the 'row_agreement' column\n\nimport numpy as np\n\n\n# Now, plotting the histogram for 'row_agreement'\nplt.figure(figsize=(10, 6))\nplt.hist(df['row_agreement'], bins=10, color='green', edgecolor='black')\nplt.title('Histogram of Row Agreement')\nplt.xlabel('Row Agreement')\nplt.ylabel('Frequency')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:48.992024Z","iopub.execute_input":"2024-02-12T03:19:48.992459Z","iopub.status.idle":"2024-02-12T03:19:49.273969Z","shell.execute_reply.started":"2024-02-12T03:19:48.992419Z","shell.execute_reply":"2024-02-12T03:19:49.273042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This looks a bit odd.  Again kind of thinking that multiple studies put togeather for for train data, with some of them having poor agreement among evalators.   Many of the 1.0 ratings are for rows with small number of evalators.\n\nRows that have very low agreement on the consensus are a problem to be addressed - not sure what's is best?","metadata":{}},{"cell_type":"code","source":"# Plotting an XY plot for 'row_agreement' vs 'total_evaluators'\n\nplt.figure(figsize=(10, 6))\nplt.scatter(df['row_agreement'], df['total_evaluators'], color='purple', edgecolor='black')\nplt.title('XY Plot of Row Agreement vs Total Evaluators')\nplt.xlabel('Row Agreement')\nplt.ylabel('Total Evaluators')\nplt.grid(True)\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:49.277968Z","iopub.execute_input":"2024-02-12T03:19:49.278364Z","iopub.status.idle":"2024-02-12T03:19:49.877212Z","shell.execute_reply.started":"2024-02-12T03:19:49.278333Z","shell.execute_reply":"2024-02-12T03:19:49.87613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I am tempted to suggest that each of the 6 seizer types has a different degree of difficulty for evaluators.\n","metadata":{}},{"cell_type":"code","source":"# Assuming 'df' has a mechanism to identify which of the 6 columns ('seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote') is the consensus for each row\n# We will generate a plot that shows 'row_agreement' values when each of these columns is the consensus vote\n\n# For demonstration, let's assume 'consensus_column' is a column that indicates which of the 6 columns is the consensus\n# This step is for demonstration purposes and should be replaced with your actual method of determining the consensus column\ndf['consensus_column'] = df[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].idxmax(axis=1)\n\n# Now, let's plot 'row_agreement' for each of the 6 columns when they are the consensus\nplt.figure(figsize=(12, 8))\n\nfor column in ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']:\n    # Filter the DataFrame for rows where this column is the consensus\n    filtered_df = df[df['consensus_column'] == column]\n    # Plotting\n    plt.scatter(filtered_df['row_agreement'], [column] * len(filtered_df), label=column)\n\nplt.title('Row Agreement for Each Column as Consensus')\nplt.xlabel('Row Agreement')\nplt.yticks(['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote'])\nplt.ylabel('Consensus Column')\nplt.legend()\nplt.grid(True)\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:49.878656Z","iopub.execute_input":"2024-02-12T03:19:49.879442Z","iopub.status.idle":"2024-02-12T03:19:55.844893Z","shell.execute_reply.started":"2024-02-12T03:19:49.879409Z","shell.execute_reply":"2024-02-12T03:19:55.844058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Little hard to see the visual information with this plot.\n\nLets revise the plot to distribution curves.","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nplt.figure(figsize=(12, 8))\n\n# Plotting distribution curves for each column\nfor column in ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']:\n    # Filter the DataFrame for rows where this column is the consensus\n    filtered_df = df[df['consensus_column'] == column]\n\n    # Plotting the distribution curve with clipping the x-axis range\n    sns.kdeplot(filtered_df['row_agreement'], label=column, clip=(0, 1.0))\n\nplt.title('Distribution of EEG_ID Agreement for Each Column as Consensus')\nplt.xlabel('Row Agreement')\nplt.ylabel('Density')\nplt.legend()\nplt.grid(True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:55.845799Z","iopub.execute_input":"2024-02-12T03:19:55.846074Z","iopub.status.idle":"2024-02-12T03:19:58.061083Z","shell.execute_reply.started":"2024-02-12T03:19:55.846047Z","shell.execute_reply":"2024-02-12T03:19:58.059882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"OK - this plot a little easier for me to interpert.\n\n1.  Some eeg-id have very poor concensus at 20% agreement or less.   Might not want to use this data ?\n2.  seize_vote seems the easiest to rate \n3.  gpd_vote would seem to be the hardest.  \n4.  Again an appearnce that at least two different studies were aggregated to form our train data.\n","metadata":{}},{"cell_type":"markdown","source":"##  eeg\n\nLets repeat this data analysis by eeg_id.  There are a number of eeg_id that have multiple rows of data.\n\nFirst lets see the distribution of evaluations per egg_id","metadata":{}},{"cell_type":"code","source":"# Assuming 'eeg_id' is a column in the 'df' DataFrame\n# We will generate a histogram that shows the count of rows for each unique 'eeg_id'\n\n# Counting the number of rows for each unique 'eeg_id'\n# Adding the 'eeg_id_counts' to the DataFrame 'df'\n# This will map each 'eeg_id' in 'df' to its count\n\n# First, create a Series with 'eeg_id' as the index and the counts as values\neeg_id_counts = df['eeg_id'].value_counts()\n\n# Mapping each 'eeg_id' in 'df' to its count\ndf['eeg_id_counts'] = df['eeg_id'].map(eeg_id_counts)\n\n\n\n\n# Plotting the histogram\nplt.figure(figsize=(12, 6))\nplt.hist(eeg_id_counts, bins=100, color='orange', edgecolor='black')\nplt.title('Histogram of Row Counts for Each Unique EEG ID')\nplt.xlabel('Number of Rows per EEG ID')\nplt.ylabel('Frequency')\nplt.grid(True)\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:58.062364Z","iopub.execute_input":"2024-02-12T03:19:58.062683Z","iopub.status.idle":"2024-02-12T03:19:58.496625Z","shell.execute_reply.started":"2024-02-12T03:19:58.062657Z","shell.execute_reply":"2024-02-12T03:19:58.495297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like a patient or two hooked up to an eeg on a single session for long time.    \n\nThe current popular shared notebooks are looking at only the first row per unique eeg_id.  That's how they go from 100K to 17K rows of data.\n\nI think using only the first is leaving information on the table, but not sure I want to include 700 rows for a single eeg_id - hmmm   what to do?","metadata":{}},{"cell_type":"code","source":"df.head(25)  # Display the first few rows of the DataFrame to show the new column\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:58.498406Z","iopub.execute_input":"2024-02-12T03:19:58.499182Z","iopub.status.idle":"2024-02-12T03:19:58.538456Z","shell.execute_reply.started":"2024-02-12T03:19:58.499121Z","shell.execute_reply":"2024-02-12T03:19:58.537224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"row_agreement_agg = df.groupby('eeg_id')['consensus'].agg('sum')\n\n# Mapping this aggregated value back to each row in 'df'\ndf['row_consensus_agg'] = df['eeg_id'].map(row_agreement_agg)\n\n\nrow_evaluators_agg = df.groupby('eeg_id')['total_evaluators'].agg('sum')\n\n# Mapping this aggregated value back to each row in 'df'\ndf['row_evaluators_agg'] = df['eeg_id'].map(row_evaluators_agg)\n\ndf['eeg_agreement'] = df['row_consensus_agg']/df['row_evaluators_agg']\n\ndf.head(25)  # Display the first few rows of the DataFrame to show the new column\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:58.540173Z","iopub.execute_input":"2024-02-12T03:19:58.540841Z","iopub.status.idle":"2024-02-12T03:19:58.611363Z","shell.execute_reply.started":"2024-02-12T03:19:58.540784Z","shell.execute_reply":"2024-02-12T03:19:58.610328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('eeg_agreement.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:19:58.612963Z","iopub.execute_input":"2024-02-12T03:19:58.614126Z","iopub.status.idle":"2024-02-12T03:20:00.187839Z","shell.execute_reply.started":"2024-02-12T03:19:58.614065Z","shell.execute_reply":"2024-02-12T03:20:00.18647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nplt.figure(figsize=(12, 8))\n\n# Plotting distribution curves for each column\nfor column in ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']:\n    # Filter the DataFrame for rows where this column is the consensus\n    filtered_df = df[df['consensus_column'] == column]\n\n    # Plotting the distribution curve with clipping the x-axis range\n    sns.kdeplot(filtered_df['eeg_agreement'], label=column, clip=(0, 1.0))\n\nplt.title('Distribution of EEG_ID Agreement for Each Column as Consensus')\nplt.xlabel('Row Agreement')\nplt.ylabel('Density')\nplt.legend()\nplt.grid(True)\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:00.189404Z","iopub.execute_input":"2024-02-12T03:20:00.189783Z","iopub.status.idle":"2024-02-12T03:20:01.687036Z","shell.execute_reply.started":"2024-02-12T03:20:00.189751Z","shell.execute_reply":"2024-02-12T03:20:01.685902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Row and eeg_id appear to have similiar plots.  ","metadata":{}},{"cell_type":"markdown","source":"Expanding on this theme - lets look at patient_id","metadata":{}},{"cell_type":"code","source":"# Assuming 'eeg_id' is a column in the 'df' DataFrame\n# We will generate a histogram that shows the count of rows for each unique 'eeg_id'\n\n# Counting the number of rows for each unique 'eeg_id'\n# Adding the 'eeg_id_counts' to the DataFrame 'df'\n# This will map each 'eeg_id' in 'df' to its count\n\n# First, create a Series with 'eeg_id' as the index and the counts as values\npatient_id_counts = df['patient_id'].value_counts()\n\n# Mapping each 'eeg_id' in 'df' to its count\ndf['patient_id_counts'] = df['patient_id'].map(patient_id_counts)\n\n\n\n\n# Plotting the histogram\nplt.figure(figsize=(12, 6))\nplt.hist(eeg_id_counts, bins=100, color='orange', edgecolor='black')\nplt.title('Histogram of Row Counts for Each Unique Patient ID')\nplt.xlabel('Number of Rows per patient ID')\nplt.ylabel('Frequency')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:01.688606Z","iopub.execute_input":"2024-02-12T03:20:01.689783Z","iopub.status.idle":"2024-02-12T03:20:02.13032Z","shell.execute_reply.started":"2024-02-12T03:20:01.689731Z","shell.execute_reply":"2024-02-12T03:20:02.12911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since a single patient is in a single eeg_id the distributions should be pretty similiar.","metadata":{}},{"cell_type":"code","source":"\n\nrow_agreement_agg = df.groupby('patient_id')['consensus'].agg('sum')\n\n# Mapping this aggregated value back to each row in 'df'\ndf['patient_consensus_agg'] = df['patient_id'].map(row_agreement_agg)\n\n\nrow_evaluators_agg = df.groupby('patient_id')['total_evaluators'].agg('sum')\n\n# Mapping this aggregated value back to each row in 'df'\ndf['patient_evaluators_agg'] = df['patient_id'].map(row_evaluators_agg)\n\ndf['patient_agreement'] = df['patient_consensus_agg']/df['patient_evaluators_agg']\n\ndf.sample(25)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:02.132347Z","iopub.execute_input":"2024-02-12T03:20:02.133235Z","iopub.status.idle":"2024-02-12T03:20:02.194172Z","shell.execute_reply.started":"2024-02-12T03:20:02.13318Z","shell.execute_reply":"2024-02-12T03:20:02.192991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 8))\n\n# Plotting distribution curves for each column\nfor column in ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']:\n    # Filter the DataFrame for rows where this column is the consensus\n    filtered_df = df[df['consensus_column'] == column]\n\n    # Plotting the distribution curve with clipping the x-axis range\n    sns.kdeplot(filtered_df['patient_agreement'], label=column, clip=(0, 1.0))\n\nplt.title('Distribution of Patient Agreement for Each Column as Consensus')\nplt.xlabel('Row Agreement')\nplt.ylabel('Density')\nplt.legend()\nplt.grid(True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:02.195964Z","iopub.execute_input":"2024-02-12T03:20:02.196332Z","iopub.status.idle":"2024-02-12T03:20:03.87434Z","shell.execute_reply.started":"2024-02-12T03:20:02.196302Z","shell.execute_reply":"2024-02-12T03:20:03.873198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Couple of ways to look at this plot.\n\nSeizure_vote - patients with this condition have consistent eeg's\nOR\neasy for evalators to spot and agree.\n\ngpd_vote - patients with this condition don't show the issue over time - it comes and goes\nOR\nvery hard for evalators to spot - the plot suggest we have no eeg for this type where all the evalators agreed \n\n\nOne Conclusion - the shared notebooks that use only the 'first' are missing valuabe information.\n\nAs an initial use of this information I think I will use the patient agreement values as weights in my fork of Chris's catboost model.\n\n\n","metadata":{}},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:03.875884Z","iopub.execute_input":"2024-02-12T03:20:03.87699Z","iopub.status.idle":"2024-02-12T03:20:03.988814Z","shell.execute_reply.started":"2024-02-12T03:20:03.876947Z","shell.execute_reply":"2024-02-12T03:20:03.987506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('train_upgraded.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:03.99052Z","iopub.execute_input":"2024-02-12T03:20:03.991197Z","iopub.status.idle":"2024-02-12T03:20:06.03198Z","shell.execute_reply.started":"2024-02-12T03:20:03.991136Z","shell.execute_reply":"2024-02-12T03:20:06.030444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nplt.hist(df['patient_agreement'], bins=100, color='green', edgecolor='black')\nplt.title('Histogram of Patient Agreement')\nplt.xlabel('Patient Agreement')\nplt.ylabel('Frequency')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:06.033345Z","iopub.execute_input":"2024-02-12T03:20:06.034552Z","iopub.status.idle":"2024-02-12T03:20:06.460953Z","shell.execute_reply.started":"2024-02-12T03:20:06.034516Z","shell.execute_reply":"2024-02-12T03:20:06.459861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The overview describes eeg's were the experts agree as \"idealized\".   If only 3 experts, but they all agree what level of confidence can we have that they are 'ideal'?\n\nIf only a single eeg for a patient, but agreement with a large number of experts, what confidence can we have that they are 'ideal'?\n\nFrom the previous plot we can see that many of the perfect agreement eeg's are 'other' or 'seizure'.  \n","metadata":{}},{"cell_type":"markdown","source":"A couple of the plots above have suggested that our data might be the result of two studies that were combined for this competition.\n\nThe number of evaluators used seems to seperate the two studies.","metadata":{}},{"cell_type":"code","source":"# Create 'large' DataFrame with rows where 'total_evaluators' is greater than 9\nlarge = df[df['total_evaluators'] > 9]\n\n# Create 'small' DataFrame with rows where 'total_evaluators' is less than 10.  actual values are 3 to 6 \nsmall = df[df['total_evaluators'] < 10]\n","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:06.462627Z","iopub.execute_input":"2024-02-12T03:20:06.46331Z","iopub.status.idle":"2024-02-12T03:20:06.493083Z","shell.execute_reply.started":"2024-02-12T03:20:06.463268Z","shell.execute_reply":"2024-02-12T03:20:06.491556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"large[\"expert_consensus\"].value_counts().plot(kind='bar');","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:27.450395Z","iopub.execute_input":"2024-02-12T03:20:27.450772Z","iopub.status.idle":"2024-02-12T03:20:27.672303Z","shell.execute_reply.started":"2024-02-12T03:20:27.450744Z","shell.execute_reply":"2024-02-12T03:20:27.67095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"small[\"expert_consensus\"].value_counts().plot(kind='bar');","metadata":{"execution":{"iopub.status.busy":"2024-02-12T03:20:32.123558Z","iopub.execute_input":"2024-02-12T03:20:32.123969Z","iopub.status.idle":"2024-02-12T03:20:32.353882Z","shell.execute_reply.started":"2024-02-12T03:20:32.123939Z","shell.execute_reply":"2024-02-12T03:20:32.35253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Did not want to see this result.\n\nWhen the number of evaluators is 6 or less, Seizure is the major label, but in the grouping where greater than 9 evaluators were used, this label is the least seen.\n\nNote that all rows are used, a couple of patients have very long eeg records with many evaluations that might change these plots dependent on use of 'first', or 'last', etc.\n\nNo way to know how many evaluators used for the test data.   \n\nThe small vs large group of evaluators presents a real problem - \n\nWith a large group of evaluators the predominat category becomes 'other'.  For many years I tracked defect types for a number of different glass products at number of different producting locations.\nWhen ever 'other' or 'misc' or 'unknown' was the leading type it always suggested a badly trained group of production inspectors.  \n\nMy guess - the large group of data was generated at a ACNS seminar or training session.   Of course, also possible that the large group was highly trained and the 5 available categories were too simplistic for experts.\n\nA key question - was the distribution of eeg's similiar for both groups (so we are seeing measurement error) or was the distribution of eeg's vastly different and we are just seeing the results of completly different studies.\n\nNot sure how to address this question.","metadata":{}}]}