{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71549,"databundleVersionId":8561470,"sourceType":"competition"},{"sourceId":8438871,"sourceType":"datasetVersion","datasetId":5026840}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"#import numpy as np\nimport pandas as pd\nimport os\nimport time\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport plotly.express as px\nimport seaborn as sns\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-18T11:32:17.267196Z","iopub.execute_input":"2024-05-18T11:32:17.267679Z","iopub.status.idle":"2024-05-18T11:32:17.275259Z","shell.execute_reply.started":"2024-05-18T11:32:17.267646Z","shell.execute_reply":"2024-05-18T11:32:17.273961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Importance Lumber Spine**\nThe term \"lumbar spine\" refers to the lower part of the spine consisting of the five vertebrae labeled L1 through L5. T. The lumbar spine is a critical structure that supports much of the upper body's weight and allows for a range of movements, including bending, twisting, and lifting. The lumbar spine bears much of the body's weight, providing structural support and stability. It allows for various movements, including flexion (bending forward), extension (bending backward), lateral flexion (side bending), and rotation. It protects the spinal cord and the nerves that run through it.\n","metadata":{}},{"cell_type":"code","source":"img = mpimg.imread('/kaggle/input/lumber-spine-l1-l5/lumber spine l1-l5.png')  # Replace with your image path\nimgplot = plt.imshow(img)\nplt.axis('off')  # Optional: Turn off the axis\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.277443Z","iopub.execute_input":"2024-05-18T11:32:17.278541Z","iopub.status.idle":"2024-05-18T11:32:17.629713Z","shell.execute_reply.started":"2024-05-18T11:32:17.278454Z","shell.execute_reply":"2024-05-18T11:32:17.627902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Watch this video**\nhttps://www.spine-health.com/video/lumbar-spine-anatomy-video","metadata":{}},{"cell_type":"markdown","source":"**Brief Introdction to Train Data**","metadata":{}},{"cell_type":"markdown","source":"* **train.csv:** This file illustrates study wise severity levels of various lumber spine conditions. For example, study id '124356' has condition as 'spinal_canal_stenosis_l1_l2' with the level of 'Normal/Mild'\n* **train_label_coordinates.csv:** This file contains several features including study_id, series_id, instance_number, condition, level, x, and y. The study_id column in this file is linked to the study_id column in the train.csv file. Additionally, a target column needs to be created by matching the values from the condition and level columns in the train_label_coordinates file with the column names in the train.csv file. To create this target column, combine the values from the condition and level columns in the train_label_coordinates file. The combined values should be in lowercase and have spaces and '/' replaced with underscores (_). This final train file, for model training, is saved in the output directory with the name 'final_train.csv'.\n\n* **test_series_descriptions.csv:** This file is just for further information about the MRI or other techniques used to take images or analyze lumber spine degernation and will not be used in the training process. \n\n**Good luck for further work **","metadata":{}},{"cell_type":"markdown","source":"**Methodolgy for EDA:**\nWe,after loading three train files, have performed EDA as follows\n* train.csv - EDA\n* train_label_coordiantes.csv - EDA\n* Generating final_train_df by combining train_label_coordiantes.csv and train.csv as described above\n* train_series_descriptions.csv - EDA\n","metadata":{}},{"cell_type":"code","source":"# reading all train data files\ntrain_data_df = pd.read_csv('../input/rsna-2024-lumbar-spine-degenerative-classification/train.csv')\ntrain_label_df = pd.read_csv('../input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\ntrain_desc_df = pd.read_csv('../input/rsna-2024-lumbar-spine-degenerative-classification/train_series_descriptions.csv')","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.632507Z","iopub.execute_input":"2024-05-18T11:32:17.633060Z","iopub.status.idle":"2024-05-18T11:32:17.720703Z","shell.execute_reply.started":"2024-05-18T11:32:17.633009Z","shell.execute_reply":"2024-05-18T11:32:17.719810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As described above two datasets are important i.e. train.csv and train_label_coordinates. So first we will explore tain.csv, then train_label_coordinates and at the end the merged dataset i.e. here in this notebook  is termed as final_train_df, will investigated.  ","metadata":{}},{"cell_type":"markdown","source":"**Train.csv - EDA**","metadata":{}},{"cell_type":"code","source":"#Displaying study wise level of each lumber spine conditions\ntrain_data_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.722335Z","iopub.execute_input":"2024-05-18T11:32:17.722949Z","iopub.status.idle":"2024-05-18T11:32:17.746822Z","shell.execute_reply.started":"2024-05-18T11:32:17.722920Z","shell.execute_reply":"2024-05-18T11:32:17.745234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.749796Z","iopub.execute_input":"2024-05-18T11:32:17.750177Z","iopub.status.idle":"2024-05-18T11:32:17.760734Z","shell.execute_reply.started":"2024-05-18T11:32:17.750147Z","shell.execute_reply":"2024-05-18T11:32:17.759036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the lumber spine conditionn in the training data with serial numbers, as you can see there are totoal 25 condistions\nfor i, column in enumerate(train_data_df.columns, start=0):\n    if i==0: #skipping study_id column, here we want to show only Lumbar Spine conditions \n        continue\n    print(f\"{i}. {column}\")\n    if i%5==0:\n        print('-----------------------')\n   ","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.762336Z","iopub.execute_input":"2024-05-18T11:32:17.762786Z","iopub.status.idle":"2024-05-18T11:32:17.774148Z","shell.execute_reply.started":"2024-05-18T11:32:17.762756Z","shell.execute_reply":"2024-05-18T11:32:17.772726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get basic info about the dataset\ntrain_data_df.info()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.775648Z","iopub.execute_input":"2024-05-18T11:32:17.776011Z","iopub.status.idle":"2024-05-18T11:32:17.796535Z","shell.execute_reply.started":"2024-05-18T11:32:17.775982Z","shell.execute_reply":"2024-05-18T11:32:17.794978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary statistics for numerical columns\ntrain_data_df.describe()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.797972Z","iopub.execute_input":"2024-05-18T11:32:17.798325Z","iopub.status.idle":"2024-05-18T11:32:17.815546Z","shell.execute_reply.started":"2024-05-18T11:32:17.798297Z","shell.execute_reply":"2024-05-18T11:32:17.814185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for missing values\ntrain_data_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.819562Z","iopub.execute_input":"2024-05-18T11:32:17.820150Z","iopub.status.idle":"2024-05-18T11:32:17.833259Z","shell.execute_reply.started":"2024-05-18T11:32:17.820103Z","shell.execute_reply":"2024-05-18T11:32:17.831859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Study wise severity levels**","metadata":{}},{"cell_type":"code","source":"from prettytable import PrettyTable\n\n# Initialize a PrettyTable\ntable = PrettyTable()\n\n# Combine all severity columns into a single dataframe for better visualization\nseverity_counts = pd.DataFrame()\ncols = train_data_df.columns[1:] # Exclude the first column (assuming it is an ID column)\nfor column in cols:\n    counts = train_data_df[column].value_counts().reset_index()\n    counts.columns = ['Severity', 'Count']\n    counts['Type'] = column\n    severity_counts = pd.concat([severity_counts, counts])\n\n# Displaying resutls in tabular form\ntable.field_names = [\"Severity\", \"Count\", \"Type\"]\n\n# Add rows to the table\nfor index, row in severity_counts.iterrows():\n    table.add_row([row['Severity'], row['Count'], row['Type']])\n\n# Print the table\nprint(table)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.834807Z","iopub.execute_input":"2024-05-18T11:32:17.835423Z","iopub.status.idle":"2024-05-18T11:32:17.946964Z","shell.execute_reply.started":"2024-05-18T11:32:17.835392Z","shell.execute_reply":"2024-05-18T11:32:17.945505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Correlation Analysis**","metadata":{}},{"cell_type":"code","source":"# Encoding the categorical variables to numerical values for correlation analysis\ndf_temp = train_data_df.copy()\nfor column in cols:\n    df_temp[column] = df_temp[column].astype('category').cat.codes\n\n# Compute the correlation matrix\ncorrelation_matrix = df_temp.corr()\n\n# Plot the heatmap of the correlation matrix\nplt.figure(figsize=(12, 10))\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt='.2f')\nplt.title('Correlation Matrix')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:17.951326Z","iopub.execute_input":"2024-05-18T11:32:17.951770Z","iopub.status.idle":"2024-05-18T11:32:20.603138Z","shell.execute_reply.started":"2024-05-18T11:32:17.951735Z","shell.execute_reply":"2024-05-18T11:32:20.601609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Pairwise comparisons to understand relationships between different stenosis and narrowing types.**","metadata":{}},{"cell_type":"code","source":"# Pairplot for selected columns\nselected_columns = train_data_df.columns[1:10]  # Select a subset for a clearer visualization\nsns.pairplot(train_data_df[selected_columns].astype('category').apply(lambda x: x.cat.codes))\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:20.604661Z","iopub.execute_input":"2024-05-18T11:32:20.605082Z","iopub.status.idle":"2024-05-18T11:32:46.448626Z","shell.execute_reply.started":"2024-05-18T11:32:20.605048Z","shell.execute_reply":"2024-05-18T11:32:46.447316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**train_label_coordinates.csv - EDA**","metadata":{}},{"cell_type":"code","source":"train_label_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:46.450373Z","iopub.execute_input":"2024-05-18T11:32:46.451049Z","iopub.status.idle":"2024-05-18T11:32:46.469376Z","shell.execute_reply.started":"2024-05-18T11:32:46.451014Z","shell.execute_reply":"2024-05-18T11:32:46.467700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_label_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:46.471543Z","iopub.execute_input":"2024-05-18T11:32:46.471971Z","iopub.status.idle":"2024-05-18T11:32:46.482989Z","shell.execute_reply.started":"2024-05-18T11:32:46.471935Z","shell.execute_reply":"2024-05-18T11:32:46.481733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"print('Summary statistics ------------\\n')\nprint()\nprint(train_label_df.describe())\n\n\nprint('\\nDataset information ------------\\n')\nprint()\nprint(train_label_df.info())\n\nprint('\\nDisplaying for missing values ------------\\n')\nprint()\nprint(train_label_df.isnull().sum())\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:46.485126Z","iopub.execute_input":"2024-05-18T11:32:46.485942Z","iopub.status.idle":"2024-05-18T11:32:46.538091Z","shell.execute_reply.started":"2024-05-18T11:32:46.485904Z","shell.execute_reply":"2024-05-18T11:32:46.536851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"**Univariate Analysis - Study ID, Series ID, Instance Number**\nUnivariate analysis is a statistical technique used to analyze and summarize a single variable or attribute.","metadata":{}},{"cell_type":"code","source":"# Unique values and counts\nprint(train_label_df['study_id'].nunique())\nprint(train_label_df['series_id'].nunique())\nprint(train_label_df['instance_number'].nunique())\n\n# Distribution of instance numbers\ntrain_label_df['instance_number'].value_counts().plot(kind='bar')\nplt.title('Distribution of Instance Numbers')\nplt.xlabel('Instance Number')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:46.539737Z","iopub.execute_input":"2024-05-18T11:32:46.540130Z","iopub.status.idle":"2024-05-18T11:32:48.147823Z","shell.execute_reply.started":"2024-05-18T11:32:46.540101Z","shell.execute_reply":"2024-05-18T11:32:48.146441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Bivariate Analysis - Analyzing Condition and Level:**\nBivariate analysis is a statistical technique employed to investigate and understand the relationship between two variables. It focues analyzing the correlation, association, and distribution of two variables to identify patterns, relationships, and trends.\n","metadata":{}},{"cell_type":"code","source":"# Cross-tabulation of condition and level\ncondition_level_ct = pd.crosstab(train_label_df['condition'], train_label_df['level'])\nprint(condition_level_ct)\n\nplt.figure(figsize=(10, 6))\nsns.heatmap(condition_level_ct, annot=True, fmt='d', cmap='viridis')\nplt.title('Condition vs Level')\nplt.xlabel('Level')\nplt.ylabel('Condition')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:48.149784Z","iopub.execute_input":"2024-05-18T11:32:48.150165Z","iopub.status.idle":"2024-05-18T11:32:48.629089Z","shell.execute_reply.started":"2024-05-18T11:32:48.150135Z","shell.execute_reply":"2024-05-18T11:32:48.627497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Study ID vs.Series ID vs Other Variables**","metadata":{}},{"cell_type":"code","source":"# Number of conditions per study\nstudy_condition_counts = train_label_df.groupby('study_id')['condition'].count()\nstudy_condition_counts.hist(bins=20)\nplt.title('Conditions per Study')\nplt.xlabel('Conditions')\nplt.ylabel('Frequency')\nplt.show()\n\n# Number of levels per study\nstudy_level_counts = train_label_df.groupby('study_id')['level'].count()\nstudy_level_counts.hist(bins=20)\nplt.title('Levels per Study')\nplt.xlabel('Levels')\nplt.ylabel('Frequency')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:48.631122Z","iopub.execute_input":"2024-05-18T11:32:48.631544Z","iopub.status.idle":"2024-05-18T11:32:49.284482Z","shell.execute_reply.started":"2024-05-18T11:32:48.631497Z","shell.execute_reply":"2024-05-18T11:32:49.283143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Spatial Analysis - Visualizing X and Y Coordinates**","metadata":{}},{"cell_type":"code","source":"# Plotting X and Y coordinates\nplt.scatter(train_label_df['x'], train_label_df['y'], alpha=0.5)\nplt.title('X vs Y Coordinates')\nplt.xlabel('X')\nplt.ylabel('Y')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:49.286263Z","iopub.execute_input":"2024-05-18T11:32:49.286754Z","iopub.status.idle":"2024-05-18T11:32:49.758676Z","shell.execute_reply.started":"2024-05-18T11:32:49.286712Z","shell.execute_reply":"2024-05-18T11:32:49.757391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" **Plotting distribution of conditions per series**","metadata":{}},{"cell_type":"code","source":"\nseries_condition_counts = train_label_df.groupby('series_id')['condition'].value_counts().unstack()\nseries_condition_counts.plot(kind='bar', stacked=True, figsize=(12, 6))\nplt.title('Distribution of Conditions per Series')\nplt.xlabel('Series ID')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:32:49.760533Z","iopub.execute_input":"2024-05-18T11:32:49.761232Z","iopub.status.idle":"2024-05-18T11:34:50.086404Z","shell.execute_reply.started":"2024-05-18T11:32:49.761192Z","shell.execute_reply":"2024-05-18T11:34:50.084586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Plotting distribution of levels per series**","metadata":{}},{"cell_type":"code","source":"\nseries_level_counts = train_label_df.groupby('series_id')['level'].value_counts().unstack()\nseries_level_counts.plot(kind='bar', stacked=True, figsize=(12, 6))\nplt.title('Distribution of Levels per Series')\nplt.xlabel('Series ID')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:34:50.088451Z","iopub.execute_input":"2024-05-18T11:34:50.088913Z","iopub.status.idle":"2024-05-18T11:36:51.475069Z","shell.execute_reply.started":"2024-05-18T11:34:50.088872Z","shell.execute_reply":"2024-05-18T11:36:51.473329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Plotting correlation between X and Y Coordinates**","metadata":{}},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorrelation_matrix = train_label_df[['x', 'y']].corr()\nprint(correlation_matrix)\n\n# Heatmap of the correlation matrix\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm')\nplt.title('Correlation Matrix of X and Y Coordinates')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:36:51.477218Z","iopub.execute_input":"2024-05-18T11:36:51.477800Z","iopub.status.idle":"2024-05-18T11:36:51.794070Z","shell.execute_reply.started":"2024-05-18T11:36:51.477748Z","shell.execute_reply":"2024-05-18T11:36:51.792704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Investigating distribtion of conditions and levels**","metadata":{}},{"cell_type":"code","source":"# Value counts for condition\nprint(train_label_df['condition'].value_counts())\n\nprint('\\n----------------\\n')\n# Value counts for level\nprint(train_label_df['level'].value_counts())\n\n# Plot the distribution of conditions\ntrain_label_df['condition'].value_counts().plot(kind='bar')\nplt.title('Distribution of Conditions')\nplt.xlabel('Condition')\nplt.ylabel('Count')\nplt.show()\n\n# Plot the distribution of levels\ntrain_label_df['level'].value_counts().plot(kind='bar')\nplt.title('Distribution of Levels')\nplt.xlabel('Level')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:36:51.795897Z","iopub.execute_input":"2024-05-18T11:36:51.796335Z","iopub.status.idle":"2024-05-18T11:36:52.379370Z","shell.execute_reply.started":"2024-05-18T11:36:51.796302Z","shell.execute_reply":"2024-05-18T11:36:52.377862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Creating merged dataframe termed as fina_train_df**","metadata":{}},{"cell_type":"code","source":"# Step 1: Creating the new column by concatenating values from columns entiteld 'conditions\n# and 'level'\ntrain_label_df['new_col'] = train_label_df['condition'].str.lower().str.replace(' ', '_') +  '_' + train_label_df['level'].str.lower().str.replace('/', '_')\n\n# Step 2: Merge the values from train_data_df based on study_id and the newly created column names\ndef get_target_value(row):\n    study_id = row['study_id']\n    new_col = row['new_col']\n    return train_data_df[train_data_df['study_id'] == study_id][new_col].values[0]\n\ntrain_label_df['target'] = train_label_df.apply(get_target_value, axis=1)\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['new_col'], inplace=True)\n\n# Copy the train_label_df to new DataFrame\nfinal_train_df = train_label_df.copy()\n\n# Drop the 'new_col' column \ntrain_label_df.drop(columns=['target'], inplace=True)\n\n\n#final_train_df.to_csv('final_train.csv')\nfinal_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:36:52.381599Z","iopub.execute_input":"2024-05-18T11:36:52.382437Z","iopub.status.idle":"2024-05-18T11:37:12.047150Z","shell.execute_reply.started":"2024-05-18T11:36:52.382390Z","shell.execute_reply":"2024-05-18T11:37:12.045636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**final_train_df - EDA**","metadata":{}},{"cell_type":"markdown","source":"**General EDA**","metadata":{}},{"cell_type":"code","source":"print('Summary statistics ------------\\n')\nprint()\nprint(final_train_df.describe())\n\nprint('\\nDataset information ------------\\n')\nprint()\nprint(final_train_df.info())\n\nprint('\\nDisplaying for missing values ------------\\n')\nprint()\nprint(final_train_df.isnull().sum())\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:12.048849Z","iopub.execute_input":"2024-05-18T11:37:12.050179Z","iopub.status.idle":"2024-05-18T11:37:12.107232Z","shell.execute_reply.started":"2024-05-18T11:37:12.050134Z","shell.execute_reply":"2024-05-18T11:37:12.105708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Plotting the distribution of the target variable**","metadata":{}},{"cell_type":"code","source":"# Check the distribution of the target variable\nprint(final_train_df['target'].value_counts())\n\n# Plot the distribution of the target variable\nsns.countplot(x='target', data=final_train_df)\nplt.title('Distribution of Target Variable')\nplt.xlabel('Target')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:12.108939Z","iopub.execute_input":"2024-05-18T11:37:12.109820Z","iopub.status.idle":"2024-05-18T11:37:12.407582Z","shell.execute_reply.started":"2024-05-18T11:37:12.109772Z","shell.execute_reply":"2024-05-18T11:37:12.406178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Performing Bivariate Analysis**","metadata":{}},{"cell_type":"markdown","source":"**Investigating the distribution of condition and level by target**","metadata":{}},{"cell_type":"code","source":"# Plotting the distribution of conditions\nsns.countplot(x='condition', hue='target', data=final_train_df)\nplt.title('Distribution of Conditions by Target')\nplt.xlabel('Condition')\nplt.ylabel('Count')\nplt.legend(title='Target')\nplt.show()\n\n# Plotting the distribution of levels\nsns.countplot(x='level', hue='target', data=final_train_df)\nplt.title('Distribution of Levels by Target')\nplt.xlabel('Level')\nplt.ylabel('Count')\nplt.legend(title='Target')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:12.409238Z","iopub.execute_input":"2024-05-18T11:37:12.409957Z","iopub.status.idle":"2024-05-18T11:37:13.245432Z","shell.execute_reply.started":"2024-05-18T11:37:12.409918Z","shell.execute_reply":"2024-05-18T11:37:13.244073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"**Analyzing X and Y coordinates by target**","metadata":{}},{"cell_type":"code","source":"# Scatter plot of X and Y coordinates colored by target\nplt.figure(figsize=(10, 6))\nsns.scatterplot(x='x', y='y', hue='target', data=final_train_df, alpha=0.6, palette='viridis')\nplt.title('X vs Y Coordinates by Target')\nplt.xlabel('X')\nplt.ylabel('Y')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:13.247308Z","iopub.execute_input":"2024-05-18T11:37:13.247745Z","iopub.status.idle":"2024-05-18T11:37:15.234069Z","shell.execute_reply.started":"2024-05-18T11:37:13.247708Z","shell.execute_reply":"2024-05-18T11:37:15.232464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Investigating data by pair plot of numerical features","metadata":{}},{"cell_type":"code","source":"# If there are additional numerical features, perform pair plot\nsns.pairplot(final_train_df[['x', 'y', 'instance_number', 'target']], hue='target', palette='viridis')\nplt.suptitle('Pair Plot of Numerical Features by Target', y=1.02)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:15.239293Z","iopub.execute_input":"2024-05-18T11:37:15.239795Z","iopub.status.idle":"2024-05-18T11:37:38.769485Z","shell.execute_reply.started":"2024-05-18T11:37:15.239753Z","shell.execute_reply":"2024-05-18T11:37:38.767883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Visualizing summary of target variable by key features - Study ID and Series ID**","metadata":{}},{"cell_type":"code","source":"# Number of conditions per study_id\nstudy_condition_counts = final_train_df.groupby('study_id')['target'].value_counts().unstack().fillna(0)\nprint(study_condition_counts)\n\n# Plot the number of target occurrences per study_id\nstudy_condition_counts.plot(kind='bar', stacked=True, figsize=(14, 7))\nplt.title('Number of Target Occurrences per Study ID')\nplt.xlabel('Study ID')\nplt.ylabel('Count')\nplt.show()\n\n# Number of conditions per series_id\nseries_condition_counts = final_train_df.groupby('series_id')['target'].value_counts().unstack().fillna(0)\nprint(series_condition_counts)\n\n# Plot the number of target occurrences per series_id\nseries_condition_counts.plot(kind='bar', stacked=True, figsize=(14, 7))\nplt.title('Number of Target Occurrences per Series ID')\nplt.xlabel('Series ID')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:37:38.771646Z","iopub.execute_input":"2024-05-18T11:37:38.773421Z","iopub.status.idle":"2024-05-18T11:39:41.519928Z","shell.execute_reply.started":"2024-05-18T11:37:38.773355Z","shell.execute_reply":"2024-05-18T11:39:41.518386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Visualizing train_series_descriptions.csv**","metadata":{}},{"cell_type":"code","source":"train_desc_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:39:41.521854Z","iopub.execute_input":"2024-05-18T11:39:41.522632Z","iopub.status.idle":"2024-05-18T11:39:41.537593Z","shell.execute_reply.started":"2024-05-18T11:39:41.522586Z","shell.execute_reply":"2024-05-18T11:39:41.536113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_desc_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:39:41.540177Z","iopub.execute_input":"2024-05-18T11:39:41.540773Z","iopub.status.idle":"2024-05-18T11:39:41.550217Z","shell.execute_reply.started":"2024-05-18T11:39:41.540724Z","shell.execute_reply":"2024-05-18T11:39:41.548784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**General EDA**","metadata":{}},{"cell_type":"code","source":"print('Summary statistics ------------\\n')\nprint()\nprint(train_desc_df.describe())\n\nprint('\\nDataset information ------------\\n')\nprint()\nprint(train_desc_df.info())\n\nprint('\\nDisplaying for missing values ------------\\n')\nprint()\nprint(train_desc_df.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:39:41.551828Z","iopub.execute_input":"2024-05-18T11:39:41.552311Z","iopub.status.idle":"2024-05-18T11:39:41.588410Z","shell.execute_reply.started":"2024-05-18T11:39:41.552264Z","shell.execute_reply":"2024-05-18T11:39:41.587128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of unique series descriptions\nseries_descriptions = train_desc_df['series_description'].nunique()\nprint(f\"Number of unique series descriptions: {series_descriptions}\")\n\n# Frequency distribution of series descriptions\nnum_series_description = train_desc_df['series_description'].value_counts()\nprint(num_series_description.head())\n\n# Plot the distribution of the top 20 series descriptions\ntop_series_descriptions = num_series_description.head(20)\nplt.figure(figsize=(12, 6))\nsns.barplot(x=top_series_descriptions.index, y=top_series_descriptions.values, palette='viridis')\nplt.title('Top 20 Series Descriptions by Frequency')\nplt.xlabel('Series Description')\nplt.ylabel('Count')\nplt.xticks(rotation=90)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:39:41.590820Z","iopub.execute_input":"2024-05-18T11:39:41.591345Z","iopub.status.idle":"2024-05-18T11:39:41.866323Z","shell.execute_reply.started":"2024-05-18T11:39:41.591302Z","shell.execute_reply":"2024-05-18T11:39:41.864777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analyzing distribution of imaging techniques (Series Descriptions) - Imaging Techniques Overview:**","metadata":{}},{"cell_type":"code","source":"# Number of unique series descriptions\nseries_descriptions = train_desc_df['series_description'].nunique()\nprint(f\"Number of unique series descriptions: {series_descriptions}\")\n\n# Frequency distribution of series descriptions\nnum_series_description = train_desc_df['series_description'].value_counts()\nprint(num_series_description.head())\n\nplt.figure(figsize=(14, 8))\nsns.countplot(y='series_description', data=train_desc_df, order=train_desc_df['series_description'].value_counts().index, palette='viridis')\nplt.title('Distribution of Series Descriptions')\nplt.xlabel('Count')\nplt.ylabel('Series Description')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T11:39:41.868241Z","iopub.execute_input":"2024-05-18T11:39:41.868816Z","iopub.status.idle":"2024-05-18T11:39:42.142678Z","shell.execute_reply.started":"2024-05-18T11:39:41.868766Z","shell.execute_reply":"2024-05-18T11:39:42.141307Z"},"trusted":true},"execution_count":null,"outputs":[]}]}