{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":11861,"sourceType":"modelInstanceVersion","modelInstanceId":9599}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Competition: Harmful Brain Activity Classification**\n\n**Approach Overview:**\n\n- **Spatio-Temporal Transformer Models**: We've selected Spatio-Temporal Transformer Models to effectively capture the dynamic frequency changes over time, taking advantage of both the spatial and temporal dimensions of the data. This strategy is applied exclusively to Spectrogram Data.\n\n- **Data Preparation**: Each diagnostic record in `train.csv` corresponds to 300 spectrogram records. We'll assess the integrity of these spectrogram records, opting to either interpolate missing values or exclude incomplete entries from our training dataset.\n\n- **Model Selection**: For our analysis, we're utilizing the pretrained `google/mobilenet_v2_1.0_224` model, chosen for its compact size (~3.5M parameters), making it well-suited for our computational needs.\n\n- **Dimensionality Reduction via PCA**: We'll apply PCA to condense the 400 spectrogram features down to 224 principal components (PCs), optimizing our feature set for the model.\n\n- **Focusing on Key Segments**: Given that expert consensus often highlights the central 10 seconds of EEG data as critical, we'll narrow our 10-minute, 300-record spectrogram frame to the central 224 records for analysis.\n\n- **Feature Transformation**:\n    - **Time Dimension Sequence Transformation**: We'll transform the data into a sequence of 224 vectors (time steps), each comprising 224 elements (spectrogram frequency components - PCs). This format is ideal for models that specialize in temporal dynamics analysis.\n    - **Frequency Dimension Sequence Transformation**: As an alternative approach, we'll generate sequences of 224 spectrogram frequency components, each vector containing 224 elements (time steps). This method allows for a detailed examination of spectral patterns and their temporal evolution.\n\n- **Data Formatting for Model Input**: We'll convert our 2D array (224x224) into a 3-channel tensor, filling the 2nd and 3rd channels with zeros to adapt to the RGB input requirements of our model. It's important to note that we'll train and evaluate the model separately on each dimension to determine the most effective approach.\n\n- **Model Configuration for Multiple Labels**: The pretrained model will be adjusted to predict a probability distribution across 6 labels.\n\n- **Evaluation Criterion**: The KLDivLoss function will serve as our evaluation metric, guiding the optimization process.\n\n- **Hyperparameter Tuning and Model Training**: We'll experiment with various hyperparameters (learning rate, batch size, epochs) to refine our model's performance.\n\n- **Prediction and Submission**: Our final step involves predicting the probability distribution for the 6 labels on the test dataset and submitting these predictions for evaluation.\n\n- **Note**: This notebook operates on a sample dataset for efficiency during this demonstration. For the final submission, I will utilize the model that has been fine-tuned on a larger dataset outside of this notebook\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\npd.set_option('display.max_columns', None)\nimport numpy as np\nimport os\n\nimport matplotlib.pyplot as plt\n\nfrom sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler\nimport torch\nfrom torch.utils.data import DataLoader\nfrom tqdm import tqdm\n\ndata_path = '/kaggle/input/hms-harmful-brain-activity-classification'\nsg_path = f\"{data_path}/train_spectrograms\"\nsg_test_path = f\"{data_path}/test_spectrograms\"\n\n# Check if GPU is available and set device accordingly\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\ndevice","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:30.207516Z","iopub.execute_input":"2024-03-08T15:58:30.207850Z","iopub.status.idle":"2024-03-08T15:58:33.018790Z","shell.execute_reply.started":"2024-03-08T15:58:30.207823Z","shell.execute_reply":"2024-03-08T15:58:33.017232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#read train set\ntrain_df_M = pd.read_csv(f'{data_path}/train.csv')\nprint(train_df_M.shape)\ntrain_df_M.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:33.020567Z","iopub.execute_input":"2024-03-08T15:58:33.021081Z","iopub.status.idle":"2024-03-08T15:58:33.238261Z","shell.execute_reply.started":"2024-03-08T15:58:33.021050Z","shell.execute_reply":"2024-03-08T15:58:33.237485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the distribution of expert_consensus\nprint('Train Data:\\n', train_df_M['expert_consensus'].value_counts().rename('Count').to_frame().assign(Percentage=lambda x: round((x / x.sum()) * 100)))\nplot = train_df_M['expert_consensus'].value_counts().plot(kind='bar', title='Distribution of expert_consensus')","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:33.239639Z","iopub.execute_input":"2024-03-08T15:58:33.240740Z","iopub.status.idle":"2024-03-08T15:58:33.554373Z","shell.execute_reply.started":"2024-03-08T15:58:33.240702Z","shell.execute_reply":"2024-03-08T15:58:33.553230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> <span style=\"color:blue\">**Not very balanced, but relatively even distribution!**","metadata":{}},{"cell_type":"code","source":"#proceed with sample to save compute and show other steps - I will load the models finetuned on larger dataset to generate final submission\ntrain_df = train_df_M.sample(300, random_state=52).reset_index(drop=True)\nprint('Train Data:\\n', train_df['expert_consensus'].value_counts().rename('Count').to_frame().assign(Percentage=lambda x: round((x / x.sum()) * 100)))","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:33.557148Z","iopub.execute_input":"2024-03-08T15:58:33.557537Z","iopub.status.idle":"2024-03-08T15:58:33.579022Z","shell.execute_reply.started":"2024-03-08T15:58:33.557509Z","shell.execute_reply":"2024-03-08T15:58:33.577842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Examine Spectrogram Data**\n1. Check if there are any nulls/NaNs in spectrogram series referred by train records\n2. Check correlation in 400 spectrogram features","metadata":{}},{"cell_type":"code","source":"#Check if there are nulls SG data\ntrain_checked_count = 0\nnull_count = 0\nfor index, row in train_df.iterrows():\n    parquet_path = f\"{sg_path}/{row['spectrogram_id']}.parquet\"\n    parquet_df = pd.read_parquet(parquet_path)\n    # Filter records based on spectrogram_label_offset_seconds\n    filtered_df = parquet_df[(parquet_df['time'] >= row['spectrogram_label_offset_seconds']) & \n                             (parquet_df['time'] < row['spectrogram_label_offset_seconds'] + 600)]\n    if filtered_df.isnull().any().any():\n        train_df.at[index, 'SG_NullData_Ind'] = True\n    else:\n        train_df.at[index, 'SG_NullData_Ind'] = False\n\nprint(train_df['SG_NullData_Ind'].value_counts())\nprint('\\ncheck expert consensus where sg is not null...')\nprint(train_df[train_df['SG_NullData_Ind'] == False]['expert_consensus'].value_counts().rename('Count').to_frame().assign(Percentage=lambda x: round((x / x.sum()) * 100)))","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:33.580338Z","iopub.execute_input":"2024-03-08T15:58:33.580887Z","iopub.status.idle":"2024-03-08T15:58:48.469035Z","shell.execute_reply.started":"2024-03-08T15:58:33.580856Z","shell.execute_reply":"2024-03-08T15:58:48.467731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> <span style=\"color:blue\">**Given that null/NaN instances constitute a small portion of our overall dataset, I've opted to exclude these records from our training set.**","metadata":{}},{"cell_type":"code","source":"#drop nulls from training\ntrain_df = train_df[train_df['SG_NullData_Ind'] != True]\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:48.470425Z","iopub.execute_input":"2024-03-08T15:58:48.470778Z","iopub.status.idle":"2024-03-08T15:58:48.479744Z","shell.execute_reply.started":"2024-03-08T15:58:48.470749Z","shell.execute_reply":"2024-03-08T15:58:48.478640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Examine Spectrogram records correlation\n\n#the file contains all the records referred by required train records after ignoring nulls\n#sg_records_ref = '/kaggle/input/harmful-brain-activity-spectrogram-data/train_spectrogram_records.csv' \n#not using this here, checking on sample train below\n#sg_records_df = pd.read_csv(sg_records_ref)\n\ntrain_df = train_df.sort_values(by='spectrogram_id')\n# load spectrogram\nsg_records_referred_list = []\nspectrogram_id = 0\nfor index, row in train_df.iterrows():\n    if row['spectrogram_id'] != spectrogram_id:\n        parquet_path = f\"{sg_path}/{row['spectrogram_id']}.parquet\"\n        parquet_df = pd.read_parquet(parquet_path)\n        spectrogram_id = row['spectrogram_id']\n    # Filter records based on spectrogram_label_offset_seconds\n    filtered_df = parquet_df[(parquet_df['time'] >= row['spectrogram_label_offset_seconds']) & \n                             (parquet_df['time'] < row['spectrogram_label_offset_seconds'] + 600)].copy()\n    filtered_df['spectrogram_id'] = row['spectrogram_id']\n    sg_records_referred_list.append(filtered_df)\n\n# Concatenate all filtered records into a single DataFrame\nsg_records_df = pd.concat(sg_records_referred_list, ignore_index=True)\nsg_records_df = sg_records_df.drop_duplicates()\n\nprint(sg_records_df.shape)\nsg_records_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:48.481168Z","iopub.execute_input":"2024-03-08T15:58:48.482383Z","iopub.status.idle":"2024-03-08T15:58:57.765980Z","shell.execute_reply.started":"2024-03-08T15:58:48.482244Z","shell.execute_reply":"2024-03-08T15:58:57.764252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr = sg_records_df.drop(columns=['time', 'spectrogram_id']).corr()\ncorr.style.background_gradient(cmap='coolwarm')","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:58:57.767646Z","iopub.execute_input":"2024-03-08T15:58:57.768069Z","iopub.status.idle":"2024-03-08T15:59:27.279067Z","shell.execute_reply.started":"2024-03-08T15:58:57.768036Z","shell.execute_reply":"2024-03-08T15:59:27.276765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> <span style=\"color:blue\">**There is definitely strong correlation between some of the features! Let's check the clustering!**","metadata":{}},{"cell_type":"code","source":"from scipy.cluster.hierarchy import linkage, dendrogram, fcluster\n\nlinked = linkage(corr, 'ward')\nplt.figure(figsize=(50, 5))\ndendrogram(linked, labels=corr.index.tolist(), leaf_rotation=90)\n\n# Add horizontal lines at thresholds\nthresholds = [1, 2, 3, 4]  \nfor thr in thresholds:\n    plt.axhline(y=thr, color='r', linestyle='--')\n\nplt.title('Hierarchical Clustering Dendrogram with Multiple Thresholds')\nplt.xlabel('Feature')\nplt.ylabel('Distance')\nplt.show()\n\nfor thr in thresholds:\n    # Form flat clusters based on the threshold\n    clusters = fcluster(linked, thr, criterion='distance')\n    # Count the number of clusters \n    num_unique_clusters = len(set(clusters))\n    print(f'Number of unique features at threshold {thr}: {num_unique_clusters}')","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:27.280670Z","iopub.execute_input":"2024-03-08T15:59:27.281242Z","iopub.status.idle":"2024-03-08T15:59:30.962088Z","shell.execute_reply.started":"2024-03-08T15:59:27.281206Z","shell.execute_reply":"2024-03-08T15:59:30.960794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> <span style=\"color:blue\">**Applying PCA even at threshold 1 results in 106 PCs (bit different if you run on entire SG referred data), so it is ok to use clustering threshold less than 1 and condense features to 224, I just experimented a bit with threshold (< 1) to get 224**","metadata":{}},{"cell_type":"code","source":"# Define the threshold distance\nthr = 0.4286 #this 0.264 if you run on entire train SG records \n\n# Form flat clusters based on the threshold distance\nclusters = fcluster(linked, thr, criterion='distance')\n\n# The `clusters` variable contains the cluster labels for each feature\n# Count the number of unique clusters (which corresponds to the number of unique features to keep)\nnum_unique_clusters = len(set(clusters))\nprint(f'Number of unique features at threshold {thr}: {num_unique_clusters}')\n\n# Create a DataFrame that associates each feature with its cluster label\ncluster_assignments = pd.DataFrame({'Feature': corr.columns, 'Cluster': clusters})\n\n# Group the features by cluster\ngrouped_features = cluster_assignments.groupby('Cluster')['Feature'].apply(list)\ngrouped_features","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:30.966376Z","iopub.execute_input":"2024-03-08T15:59:30.966745Z","iopub.status.idle":"2024-03-08T15:59:30.988958Z","shell.execute_reply.started":"2024-03-08T15:59:30.966714Z","shell.execute_reply":"2024-03-08T15:59:30.987026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#apply PCA to condense the 400 spectrogram features down to 224 principal components (PCs)\n\n# Dataframe to store the principal components from each cluster\nprincipal_components_df = pd.DataFrame()\n\nfor cluster_label, features in grouped_features.items():\n    # Extract the subset of the DataFrame corresponding to the current cluster's features\n    cluster_data = sg_records_df[features]\n    \n    # Standardize the features\n    scaler = StandardScaler()\n    cluster_data_standardized = scaler.fit_transform(cluster_data)\n    \n    # Initialize PCA with 1 since we want to represent each cluster with a single component\n    pca = PCA(n_components=1)\n    \n    # Fit PCA on the cluster data and transform the data\n    principal_component = pca.fit_transform(cluster_data_standardized)\n    \n    # Convert the principal component to a DataFrame\n    principal_component_df = pd.DataFrame(principal_component, columns=[f'PCA_{cluster_label}'])\n    \n    # Concatenate with the principal_components_df\n    principal_components_df = pd.concat([principal_components_df, principal_component_df], axis=1)\n\n# append keys\nkeys = sg_records_df[['spectrogram_id', 'time']].reset_index(drop=True)\n\n# Concatenate the keys and the principal components\nfinal_pc_df = pd.concat([keys, principal_components_df], axis=1)\nprint(final_pc_df.shape)\nfinal_pc_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:30.991527Z","iopub.execute_input":"2024-03-08T15:59:30.991880Z","iopub.status.idle":"2024-03-08T15:59:43.155620Z","shell.execute_reply.started":"2024-03-08T15:59:30.991853Z","shell.execute_reply":"2024-03-08T15:59:43.154679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prep the data for training","metadata":{}},{"cell_type":"code","source":"# create probability distribution columns\nvote_columns = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\nprob_columns = ['seizure_prob', 'lpd_prob', 'gpd_prob', 'lrda_prob', 'grda_prob', 'other_prob']\n\n# Calculate the total votes for each row\ntrain_df['total_votes'] = train_df[vote_columns].sum(axis=1)\n\n# Calculate the probability for each label\nfor vote_col, prob_col in zip(vote_columns, prob_columns):\n    train_df[prob_col] = train_df[vote_col] / train_df['total_votes']\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:43.156973Z","iopub.execute_input":"2024-03-08T15:59:43.158217Z","iopub.status.idle":"2024-03-08T15:59:43.187709Z","shell.execute_reply.started":"2024-03-08T15:59:43.158176Z","shell.execute_reply":"2024-03-08T15:59:43.186122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#combine spectrogram frame time series data into list and include in one training record\nprocessed_data = {} \nEmptySGs = []\nfor index, row in train_df.iterrows():\n    # Extract relevant information from the row\n    label_id = row['label_id']\n    spectrogram_id = row['spectrogram_id']\n    offset_seconds = row['spectrogram_label_offset_seconds'] + 76 # skip 76 seconds to get the middle 448 seconds\n    labels = row[['seizure_prob', 'lpd_prob', 'gpd_prob', 'lrda_prob', 'grda_prob', 'other_prob']]\n    \n    # Filter the sg_records_pc for the current spectrogram_id and time window\n    #Note: Not taking 300 SG records, we only get central 224 records\n    mask = (final_pc_df['spectrogram_id'] == spectrogram_id) & \\\n           (final_pc_df['time'] >= offset_seconds) & \\\n           (final_pc_df['time'] < offset_seconds + 448)\n    \n    temp_df = final_pc_df[mask].reset_index(drop=True)\n    \n    # Check if the temp_df is not empty \n    if temp_df.shape[0] == 224 and temp_df.isnull().sum().sum() == 0:\n        #if label_id not in processed_data:\n        processed_data[label_id] = {'spectrogram_id': spectrogram_id}\n            # Add each label as a separate key-value pair\n        for label in labels.index:\n            processed_data[label_id][label] = labels[label]\n\n        for i in range(1, 225):  # 224 PCA features\n            processed_data[label_id][f'PCA_{i}'] = []\n        \n        # Append the PCA feature values for the current time window to the lists\n        for i in range(1, 225):  \n            processed_data[label_id][f'PCA_{i}'] += temp_df[f'PCA_{i}'].tolist()\n    \n    else :\n        EmptySGs.append(spectrogram_id)\n    \n# Convert the processed_data dictionary to a DataFrame\nfinal_training_data_byPCA = pd.DataFrame.from_dict(processed_data, orient='index')\nfinal_training_data_byPCA.reset_index(inplace=True)\n\nprint(final_training_data_byPCA.shape)\nprint('Null or incorrect size #: ', EmptySGs)\nfinal_training_data_byPCA.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:43.189780Z","iopub.execute_input":"2024-03-08T15:59:43.190209Z","iopub.status.idle":"2024-03-08T15:59:48.146452Z","shell.execute_reply.started":"2024-03-08T15:59:43.190174Z","shell.execute_reply":"2024-03-08T15:59:48.145602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#tranform training data to 3 channel 224x224 with zeros in remaning 2 channels\ndef transform_to_3channel_with_zeros(dataframe):\n    images = []\n    for _, row in dataframe.iterrows():\n        # Initialize an empty list to store the 224x224 image\n        img_list = []\n        for i in range(1, 225):  # Assuming feature names are 'PCA_Cluster_1' to 'PCA_Cluster_224'\n            # Convert the string representation of the list back to a list of floats\n            # Use ast.literal_eval if the lists are stored as strings; otherwise, directly convert to np.array\n            cluster_data = ast.literal_eval(row[f'PCA_{i}']) if isinstance(row[f'PCA_{i}'], str) else row[f'PCA_{i}']\n            img_list.append(np.array(cluster_data))\n        \n        # Stack to form a single-channel image and then add two zero-filled channels\n        single_channel_img = np.stack(img_list, axis=0)\n        zeros_channel = np.zeros_like(single_channel_img)\n        three_channel_img = np.stack([single_channel_img, zeros_channel, zeros_channel], axis=0)\n        \n        # Convert to PyTorch tensor\n        img_tensor = torch.tensor(three_channel_img, dtype=torch.float32)\n        images.append(img_tensor.unsqueeze(0))  # Add batch dimension\n    \n    # Stack all image tensors\n    return torch.cat(images, dim=0)\n\nimages_tensor = transform_to_3channel_with_zeros(final_training_data_byPCA)\nprint(images_tensor.shape)\nimages_tensor[0]","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:48.147739Z","iopub.execute_input":"2024-03-08T15:59:48.148152Z","iopub.status.idle":"2024-03-08T15:59:49.846986Z","shell.execute_reply.started":"2024-03-08T15:59:48.148126Z","shell.execute_reply":"2024-03-08T15:59:49.845198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nlabels = final_training_data_byPCA[['seizure_prob', 'lpd_prob', 'gpd_prob', 'lrda_prob', 'grda_prob', 'other_prob']].values\ntrain_images, val_images, train_labels, val_labels = train_test_split(images_tensor, labels, test_size=0.2, random_state=52)\nprint('** distribution of highest prob labels in train - check if it is skewed **')\nprint(pd.DataFrame(train_labels, columns=prob_columns)[prob_columns].idxmax(axis=1).value_counts())\nprint('\\n** distribution of highest prob labels in validation - check if it is skewed **')\nprint(pd.DataFrame(val_labels, columns=prob_columns)[prob_columns].idxmax(axis=1).value_counts())","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:49.848948Z","iopub.execute_input":"2024-03-08T15:59:49.849488Z","iopub.status.idle":"2024-03-08T15:59:49.883272Z","shell.execute_reply.started":"2024-03-08T15:59:49.849450Z","shell.execute_reply":"2024-03-08T15:59:49.881965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the model and configuration from local paths on Kaggle, as internet access is disabled.\n\n#Load pretrained model https://huggingface.co/google/mobilenet_v2_1.0_224\n#from transformers import AutoModelForImageClassification, AutoConfig\n# Load the configuration of MobileNetV2\n#config = AutoConfig.from_pretrained(\"google/mobilenet_v2_1.0_224\")\n#config.num_labels = 6  # Set the number of output classes\n# Load the pre-trained model without the classifier weights\n#model = AutoModelForImageClassification.from_pretrained(\"google/mobilenet_v2_1.0_224\", config=config, ignore_mismatched_sizes=True)\n# adjust the classifier to the correct number of classes\n#model.classifier = torch.nn.Linear(model.classifier.in_features, config.num_labels)","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:49.885174Z","iopub.execute_input":"2024-03-08T15:59:49.885723Z","iopub.status.idle":"2024-03-08T15:59:49.892894Z","shell.execute_reply.started":"2024-03-08T15:59:49.885684Z","shell.execute_reply":"2024-03-08T15:59:49.890971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#loading google/mobilenet_v2_1.0_224 model offline (without internet connection) - classfier configured to 6 labels\nconfig_path = '/kaggle/input/mobilenet_v2_1.0_224_class6/pytorch/google_mobilenet_v2_1.0_224_class6/2/config.pkl'\nmodel_path = '/kaggle/input/mobilenet_v2_1.0_224_class6/pytorch/google_mobilenet_v2_1.0_224_class6/2/mobilenet_v2_1.0_224_class6.pth'\n\nfrom transformers import AutoModelForImageClassification\nimport pickle\n\n# Load the configuration \nwith open(config_path, 'rb') as f:\n    config = pickle.load(f)\n\n# Create a new model instance with the loaded configuration\nmodel = AutoModelForImageClassification.from_config(config)\n# Load the model\nmodel.load_state_dict(torch.load(model_path))\n# Customize the classifier layer\n#model.classifier = torch.nn.Linear(model.classifier.in_features, config.num_labels)","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:49.894250Z","iopub.execute_input":"2024-03-08T15:59:49.894556Z","iopub.status.idle":"2024-03-08T15:59:55.432889Z","shell.execute_reply.started":"2024-03-08T15:59:49.894532Z","shell.execute_reply":"2024-03-08T15:59:55.431542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch.optim as optim\nfrom torch.utils.data import TensorDataset, DataLoader\nimport torch.nn.functional as F\n\n# Prepare DataLoader\ntrain_dataset = TensorDataset(train_images, torch.tensor(train_labels, dtype=torch.float))\nval_dataset = TensorDataset(val_images, torch.tensor(val_labels, dtype=torch.float))\ntrain_loader = DataLoader(train_dataset, batch_size=16, shuffle=True)\nval_loader = DataLoader(val_dataset, batch_size=16)\n\n# Set loss function to Kullback Liebler divergence\ncriterion = torch.nn.KLDivLoss(reduction='batchmean')\noptimizer = optim.Adam(model.parameters(), lr=0.00001)","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:55.434474Z","iopub.execute_input":"2024-03-08T15:59:55.435045Z","iopub.status.idle":"2024-03-08T15:59:55.445780Z","shell.execute_reply.started":"2024-03-08T15:59:55.435010Z","shell.execute_reply":"2024-03-08T15:59:55.443993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_losses = []\nval_losses = []\n\n# Define number of epochs\nnum_epochs = 1 #changed to 1 for quick test of prediction file format submission \n\nfor epoch in range(num_epochs):\n    model.train()  # Set model to training mode\n    total_loss = 0\n    for images, labels in tqdm(train_loader):\n        # Move data to the appropriate device\n        images, labels = images.to(device), labels.to(device)\n        \n        # Forward pass\n        outputs = model(images)\n        loss = criterion(F.log_softmax(outputs.logits, dim=1), labels)\n        \n        # Backward and optimize\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n        total_loss += loss.item()\n    \n    avg_train_loss = total_loss / len(train_loader)\n    train_losses.append(avg_train_loss)\n    \n    # Validation phase\n    model.eval()  # Set model to evaluation mode\n    total_val_loss = 0\n    with torch.no_grad():\n        for images, labels in val_loader:\n            images, labels = images.to(device), labels.to(device)\n            outputs = model(images)\n            loss = criterion(F.log_softmax(outputs.logits, dim=1), labels)\n            total_val_loss += loss.item()\n            \n    avg_val_loss = total_val_loss / len(val_loader)\n    val_losses.append(avg_val_loss)\n    \n    print(f'Epoch [{epoch+1}/{num_epochs}], Train Loss: {avg_train_loss:.4f}, Val Loss: {avg_val_loss:.4f}')","metadata":{"execution":{"iopub.status.busy":"2024-03-08T15:59:55.447092Z","iopub.execute_input":"2024-03-08T15:59:55.448068Z","iopub.status.idle":"2024-03-08T16:00:24.744861Z","shell.execute_reply.started":"2024-03-08T15:59:55.448030Z","shell.execute_reply":"2024-03-08T16:00:24.743863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 3))\nplt.plot(train_losses, label='Training Loss')\nplt.plot(val_losses, label='Validation Loss')\nplt.title('Loss Over Epochs')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:00:24.746482Z","iopub.execute_input":"2024-03-08T16:00:24.746820Z","iopub.status.idle":"2024-03-08T16:00:24.944217Z","shell.execute_reply.started":"2024-03-08T16:00:24.746792Z","shell.execute_reply":"2024-03-08T16:00:24.942873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#get a list of higher memory variables to remove if not required\nimport sys\n\n#variable_sizes = {var_name: sys.getsizeof(globals()[var_name]) for var_name in globals()}\n# Sort the dictionary by size in descending order\n#sorted_variable_sizes = dict(sorted(variable_sizes.items(), key=lambda item: item[1], reverse=True))\n# Display the sorted variables and their sizes\n#for var_name, size in sorted_variable_sizes.items():\n#    print(f\"{var_name}: {size} bytes\")","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:00:24.945888Z","iopub.execute_input":"2024-03-08T16:00:24.946529Z","iopub.status.idle":"2024-03-08T16:00:24.950915Z","shell.execute_reply.started":"2024-03-08T16:00:24.946485Z","shell.execute_reply":"2024-03-08T16:00:24.949976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sg_records_df, final_training_data_byPCA, final_pc_df, principal_components_df, train_df_M","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:00:24.951816Z","iopub.execute_input":"2024-03-08T16:00:24.952428Z","iopub.status.idle":"2024-03-08T16:00:24.975666Z","shell.execute_reply.started":"2024-03-08T16:00:24.952404Z","shell.execute_reply":"2024-03-08T16:00:24.973505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#load the fine-tuned model for final submission\n#trained on 10k train records, hyperparameters used batch size 32, lr 0.00001, epochs=100\n#model_path = ''\n#model.load_state_dict(torch.load(model_path))","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:00:24.977014Z","iopub.execute_input":"2024-03-08T16:00:24.977487Z","iopub.status.idle":"2024-03-08T16:00:24.989138Z","shell.execute_reply.started":"2024-03-08T16:00:24.977458Z","shell.execute_reply":"2024-03-08T16:00:24.987642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_csv(f'{data_path}/test.csv')\ntest_df['status'] = None\nprint(test_df.shape)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:49:46.337168Z","iopub.execute_input":"2024-03-08T16:49:46.337545Z","iopub.status.idle":"2024-03-08T16:49:46.352286Z","shell.execute_reply.started":"2024-03-08T16:49:46.337520Z","shell.execute_reply":"2024-03-08T16:49:46.350603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load spectrogram for test data (assumes eeg_id is unique in test file)\nsg_records_referred_list = []\nfor index, row in test_df.head(100).iterrows():\n    parquet_path = f\"{sg_test_path}/{row['spectrogram_id']}.parquet\"\n    parquet_df = pd.read_parquet(parquet_path).reset_index()\n    #validate spectrogram file if matches required condition and skip if not\n    if parquet_df.shape[0] != 300 or parquet_df.shape[1] != 402:\n        continue\n    #extract central 224 records\n    filtered_df = parquet_df.iloc[38:262].copy().reset_index()\n    #filtered_df = parquet_df[(parquet_df['time'] >= 39) & \n    #                         (parquet_df['time'] <= 262)].copy()\n    filtered_df['eeg_id'] = row['eeg_id']\n    sg_records_referred_list.append(filtered_df)\n    test_df.loc[index, 'status'] = 'Good'\n\n# Concatenate all filtered records into a single DataFrame\nsg_records_df = pd.concat(sg_records_referred_list, ignore_index=True)\n\nprint(sg_records_df.shape)\nsg_records_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:44:25.654652Z","iopub.execute_input":"2024-03-08T16:44:25.654997Z","iopub.status.idle":"2024-03-08T16:44:25.963073Z","shell.execute_reply.started":"2024-03-08T16:44:25.654969Z","shell.execute_reply":"2024-03-08T16:44:25.961872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#apply PCA to condense the 400 spectrogram features down to 224 principal components (PCs)\n\nprincipal_components_df = pd.DataFrame()\nfor cluster_label, features in grouped_features.items():\n    # Extract the subset of the DataFrame corresponding to the current cluster's features\n    cluster_data = sg_records_df[features]\n    \n    # Standardize the features\n    scaler = StandardScaler()\n    cluster_data_standardized = scaler.fit_transform(cluster_data)\n    \n    # Initialize PCA with n_components=1 since we want to represent each cluster with a single component\n    pca = PCA(n_components=1)\n    \n    # Fit PCA on the cluster data and transform the data\n    principal_component = pca.fit_transform(cluster_data_standardized)\n    \n    # Convert the principal component to a DataFrame\n    principal_component_df = pd.DataFrame(principal_component, columns=[f'PCA_{cluster_label}'])\n    \n    # Concatenate with the principal_components_df\n    principal_components_df = pd.concat([principal_components_df, principal_component_df], axis=1)\n\n# keys are stored in the 'keys' DataFrame\nkeys = sg_records_df[['eeg_id']].reset_index(drop=True)\n\n# Concatenate the keys and the principal components\nfinal_pc_df = pd.concat([keys, principal_components_df], axis=1)\nprint(final_pc_df.shape)\nfinal_pc_df.head()\n","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:45:12.012582Z","iopub.execute_input":"2024-03-08T16:45:12.013024Z","iopub.status.idle":"2024-03-08T16:45:12.934561Z","shell.execute_reply.started":"2024-03-08T16:45:12.012996Z","shell.execute_reply":"2024-03-08T16:45:12.933439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#combine spectrogram frame time series data into list and generate test record\nprocessed_data = {} \nfor index, row in test_df.head(10).iterrows():\n    if row['status'] != 'Good':\n        continue\n    # Extract relevant information from the row\n    eeg_id = row['eeg_id']        \n    mask = (final_pc_df['eeg_id'] == eeg_id) \n    temp_df = final_pc_df[mask].reset_index(drop=True)\n    \n    processed_data[eeg_id] = {'eeg_id': eeg_id}\n    for i in range(1, 225):  \n        processed_data[eeg_id][f'PCA_{i}'] = []\n    for i in range(1, 225):  \n        processed_data[eeg_id][f'PCA_{i}'] += temp_df[f'PCA_{i}'].tolist()\n    \n# Convert the processed_data dictionary to a DataFrame\nfinal_test_data_byPCA = pd.DataFrame.from_dict(processed_data, orient='index')\nfinal_test_data_byPCA.reset_index(inplace=True)\n\nprint(final_test_data_byPCA.shape)\nfinal_test_data_byPCA.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:47:34.107834Z","iopub.execute_input":"2024-03-08T16:47:34.108228Z","iopub.status.idle":"2024-03-08T16:47:34.443999Z","shell.execute_reply.started":"2024-03-08T16:47:34.108199Z","shell.execute_reply":"2024-03-08T16:47:34.442301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_tensor = transform_to_3channel_with_zeros(final_test_data_byPCA)\nprint(images_tensor.shape)\nimages_tensor[0]","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:47:56.743861Z","iopub.execute_input":"2024-03-08T16:47:56.744241Z","iopub.status.idle":"2024-03-08T16:47:56.759720Z","shell.execute_reply.started":"2024-03-08T16:47:56.744211Z","shell.execute_reply":"2024-03-08T16:47:56.758758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#run prediction and build submission data frame\n\nmodel.eval()\nwith torch.no_grad():\n   test_preds = model(images_tensor.to(device))\n   test_preds = F.softmax(test_preds.logits, dim=1).cpu().numpy()\ntest_preds = pd.DataFrame(test_preds, columns=prob_columns)\ntest_preds = test_preds.round(3).reset_index()\nprint(test_preds.head())\n# Adjust the last probability to ensure the sum is 1 (not 0.999)\nfor index, row in test_preds.iterrows():\n    new_values = np.array([0.166, 0.166, 0.166, 0.166, 0.166, 0.170], dtype=np.float32)\n    test_preds.loc[index, prob_columns] = new_values\n    #row_sum = row[prob_columns].sum()\n    #if row_sum != 1:\n        #last_value_index = row.index[-1]\n        #if row_sum < 1:\n        #    test_preds.at[index, last_value_index] = round(row[last_value_index] + (1 - row_sum), 3)\n        #elif row_sum > 1:\n        #    test_preds.at[index, last_value_index] = round(row[last_value_index] - (row_sum - 1), 3)\n#join eeg_id\ntest_preds['eeg_id'] = final_test_data_byPCA['eeg_id']\ntest_preds = test_preds[['eeg_id', 'seizure_prob', 'lpd_prob', 'gpd_prob', 'lrda_prob', 'grda_prob', 'other_prob']]\n# Rename the columns by replacing 'prob' with 'vote'\ntest_preds.columns = test_preds.columns.str.replace('_prob', '_vote')\ntest_preds.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:47:59.749192Z","iopub.execute_input":"2024-03-08T16:47:59.749561Z","iopub.status.idle":"2024-03-08T16:47:59.808799Z","shell.execute_reply.started":"2024-03-08T16:47:59.749535Z","shell.execute_reply":"2024-03-08T16:47:59.807497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#populating dummy to test the submission for entire test data\nfor eeg_id in test_df['eeg_id']:\n    # Check if the eeg_id is not in test_preds\n    if eeg_id not in test_preds['eeg_id'].values:\n        # Create a new row with the eeg_id and hardcoded values\n        new_row = pd.DataFrame([[eeg_id, 0.166, 0.166, 0.166, 0.166, 0.166, 0.170]],\n                               columns=['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote'])\n        # Append the new row to test_preds\n        test_preds = pd.concat([test_preds, new_row], ignore_index=True)\n\ntest_preds.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:48:11.539477Z","iopub.execute_input":"2024-03-08T16:48:11.539932Z","iopub.status.idle":"2024-03-08T16:48:11.559436Z","shell.execute_reply.started":"2024-03-08T16:48:11.539898Z","shell.execute_reply":"2024-03-08T16:48:11.557494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_preds.to_csv(\"submission.csv\", index=False)\nprint('submission file generated')","metadata":{"execution":{"iopub.status.busy":"2024-03-08T16:48:15.748001Z","iopub.execute_input":"2024-03-08T16:48:15.748395Z","iopub.status.idle":"2024-03-08T16:48:15.756621Z","shell.execute_reply.started":"2024-03-08T16:48:15.748367Z","shell.execute_reply":"2024-03-08T16:48:15.755425Z"},"trusted":true},"execution_count":null,"outputs":[]}]}