{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"code","source":"import warnings\nimport timm\nfrom fastai.vision.all import *\nfrom fastcore.parallel import *\n\npath = Path('/kaggle/input/hms-harmful-brain-activity-classification')\n\npath.ls()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-05T17:39:41.084342Z","iopub.execute_input":"2024-03-05T17:39:41.085186Z","iopub.status.idle":"2024-03-05T17:39:53.104706Z","shell.execute_reply.started":"2024-03-05T17:39:41.085144Z","shell.execute_reply":"2024-03-05T17:39:53.103747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background\n\nIn this notebook, I will train (and export) a resnet34 image classification model on stacked and transposed spectrograms instead of the averaged spectrograms that I used in my [previous notebook](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-starter-notebook). Everything else in the training will be the same.\n\nI'll be referencing [Sonu Jha's fastai Starter Notebook](https://www.kaggle.com/code/sonujha090/hms-hbac-fastai-starter/notebook) for the spectrogram processing code.\n\nAs I have done so with my previous notebooks, this is the training notebook and has a [corresponding submission notebook](https://www.kaggle.com/vishalbakshi/hms-hbac-fastai-stacked-images-submit).","metadata":{}},{"cell_type":"markdown","source":"## Converting Training and Test Data to Images","metadata":{}},{"cell_type":"markdown","source":"Here is `process_spec` as shown in the reference notebook. I'll run it on one spectrogram parquet file to show what it looks like:","metadata":{}},{"cell_type":"code","source":"def process_spec(spec_id, split=\"train\"):\n    # read the data\n    data = pd.read_parquet(path/f'{split}_spectrograms'/f'{spec_id}.parquet')\n    \n    # replace NA with 0\n    data = data.fillna(0)\n    \n    # convert DataFrame to array\n    data = data.values[:, 1:]\n    \n    # transpose\n    data = data.T\n    data = data.astype(\"float32\")\n    \n    # convert array to PILImage\n    im = PILImage.create(Image.fromarray((data * 255).astype(np.uint8)))\n    return im","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:44:05.032456Z","iopub.execute_input":"2024-03-05T17:44:05.033414Z","iopub.status.idle":"2024-03-05T17:44:05.040709Z","shell.execute_reply.started":"2024-03-05T17:44:05.033366Z","shell.execute_reply":"2024-03-05T17:44:05.039462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"process_spec('1111500860')","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:44:07.072196Z","iopub.execute_input":"2024-03-05T17:44:07.073389Z","iopub.status.idle":"2024-03-05T17:44:07.417539Z","shell.execute_reply.started":"2024-03-05T17:44:07.073335Z","shell.execute_reply":"2024-03-05T17:44:07.416440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Instead of averaging the four regions, it stacks them on top of each other and transposes them so it's read horizontally instead of vertically.","metadata":{}},{"cell_type":"markdown","source":"As done in the reference code, I'll create temporary folders to store the images:","metadata":{}},{"cell_type":"code","source":"# create temporary folders to hold spectrograms\nSPEC_DIR = \"/tmp/dataset/hms-hbac\"\nos.makedirs(SPEC_DIR+'/train_spectrograms', exist_ok=True)\nos.makedirs(SPEC_DIR+'/test_spectrograms', exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:45:23.279387Z","iopub.execute_input":"2024-03-05T17:45:23.280323Z","iopub.status.idle":"2024-03-05T17:45:23.287907Z","shell.execute_reply.started":"2024-03-05T17:45:23.280282Z","shell.execute_reply":"2024-03-05T17:45:23.286936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And add a line to the `process_spec` function to save the image:","metadata":{}},{"cell_type":"code","source":"def process_spec(spec_id, split=\"train\"):\n    # read the data\n    data = pd.read_parquet(path/f'{split}_spectrograms'/f'{spec_id}.parquet')\n    \n    # replace NA with 0\n    data = data.fillna(0)\n    \n    # convert DataFrame to array\n    data = data.values[:, 1:]\n    \n    # transpose\n    data = data.T\n    data = data.astype(\"float32\")\n    \n    # convert array to PILImage\n    im = PILImage.create(Image.fromarray((data * 255).astype(np.uint8)))\n    im.save(f\"{SPEC_DIR}/{split}_spectrograms/{spec_id}.png\")","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:45:51.753314Z","iopub.execute_input":"2024-03-05T17:45:51.754369Z","iopub.status.idle":"2024-03-05T17:45:51.761429Z","shell.execute_reply.started":"2024-03-05T17:45:51.754330Z","shell.execute_reply":"2024-03-05T17:45:51.760181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll load the training data so I can get all of the `spectrogram_id` values:","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(path/'train.csv')\ndf.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:45:56.159618Z","iopub.execute_input":"2024-03-05T17:45:56.160108Z","iopub.status.idle":"2024-03-05T17:45:56.449952Z","shell.execute_reply.started":"2024-03-05T17:45:56.160069Z","shell.execute_reply":"2024-03-05T17:45:56.448794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec_ids = df[\"spectrogram_id\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:46:00.007854Z","iopub.execute_input":"2024-03-05T17:46:00.008669Z","iopub.status.idle":"2024-03-05T17:46:00.018502Z","shell.execute_reply.started":"2024-03-05T17:46:00.008638Z","shell.execute_reply":"2024-03-05T17:46:00.017328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(spec_ids)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:46:00.958783Z","iopub.execute_input":"2024-03-05T17:46:00.959957Z","iopub.status.idle":"2024-03-05T17:46:00.967211Z","shell.execute_reply.started":"2024-03-05T17:46:00.959920Z","shell.execute_reply":"2024-03-05T17:46:00.965854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And then save them as image files using `fastcore.parallel` which takes about four and a half minutes.","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")\nparallel(process_spec, spec_ids, split='train', n_workers=4)\nwarnings.filterwarnings(\"default\")","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:46:10.207038Z","iopub.execute_input":"2024-03-05T17:46:10.207440Z","iopub.status.idle":"2024-03-05T17:52:26.460958Z","shell.execute_reply.started":"2024-03-05T17:46:10.207412Z","shell.execute_reply":"2024-03-05T17:52:26.459834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll load a training spectrogram image to make sure:","metadata":{}},{"cell_type":"code","source":"Path('/tmp/dataset/hms-hbac/train_spectrograms').ls()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:54:17.297749Z","iopub.execute_input":"2024-03-05T17:54:17.298640Z","iopub.status.idle":"2024-03-05T17:54:17.336108Z","shell.execute_reply.started":"2024-03-05T17:54:17.298587Z","shell.execute_reply":"2024-03-05T17:54:17.335058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"PILImage.create('/tmp/dataset/hms-hbac/train_spectrograms/1829953376.png')","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:54:21.236537Z","iopub.execute_input":"2024-03-05T17:54:21.237520Z","iopub.status.idle":"2024-03-05T17:54:21.297450Z","shell.execute_reply.started":"2024-03-05T17:54:21.237479Z","shell.execute_reply":"2024-03-05T17:54:21.296316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good!","metadata":{}},{"cell_type":"markdown","source":"## Finetuning a Pretrained ResNet34","metadata":{}},{"cell_type":"markdown","source":"I'll wrap all of my code to prep data for training into a single cell.","metadata":{}},{"cell_type":"code","source":"df['img_path'] = '/tmp/dataset/hms-hbac/train_spectrograms/' + df['spectrogram_id'].astype(str) + '.png'\n    \ncols = ['eeg_id', 'spectrogram_id', 'img_path', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\nagg_funcs = {c: 'sum' for c in cols if 'vote' in c}\n\nunique_df = df[cols].groupby(['eeg_id', 'spectrogram_id', 'img_path'], as_index=False).agg(agg_funcs)\nunique_df['target'] = unique_df[[c for c in cols if 'vote' in c]].idxmax(axis=1)\n    \ntrain_bool = [False for _ in range(int(0.8 * len(unique_df)))]\nvalid_bool = [True for _ in range(len(unique_df) - int(0.8 * len(unique_df)))]\nis_valid_bool = pd.Series(train_bool + valid_bool).sample(frac=1).reset_index(drop=True)\n\nunique_df[\"is_valid\"] = is_valid_bool\nunique_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:09.032175Z","iopub.execute_input":"2024-03-05T17:55:09.032550Z","iopub.status.idle":"2024-03-05T17:55:09.239560Z","shell.execute_reply.started":"2024-03-05T17:55:09.032518Z","shell.execute_reply":"2024-03-05T17:55:09.238310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I can create my `DataBlock`. Since I now have paths to images ready to train on, I can use `ImageBlock` for my inputs. I'll also make sure to `Resize` my images so they are 224 x 224 squares.","metadata":{}},{"cell_type":"code","source":"dblock = DataBlock(\n            blocks=(ImageBlock, CategoryBlock),\n            splitter=ColSplitter(),\n            get_x=ColReader('img_path'),\n            get_y=ColReader('target'),\n            item_tfms=Resize(224, method='squish'))\n\ndls = dblock.dataloaders(unique_df)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:14.630193Z","iopub.execute_input":"2024-03-05T17:55:14.631131Z","iopub.status.idle":"2024-03-05T17:55:16.348331Z","shell.execute_reply.started":"2024-03-05T17:55:14.631099Z","shell.execute_reply":"2024-03-05T17:55:16.347240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch(nrows=1, ncols=3)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:19.028802Z","iopub.execute_input":"2024-03-05T17:55:19.029166Z","iopub.status.idle":"2024-03-05T17:55:20.677933Z","shell.execute_reply.started":"2024-03-05T17:55:19.029137Z","shell.execute_reply":"2024-03-05T17:55:20.676920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check dls vocab\ndls.vocab","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:25.735485Z","iopub.execute_input":"2024-03-05T17:55:25.735842Z","iopub.status.idle":"2024-03-05T17:55:25.742349Z","shell.execute_reply.started":"2024-03-05T17:55:25.735814Z","shell.execute_reply":"2024-03-05T17:55:25.741381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"fastai uses `timm` when you specify the architecture as a string:","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, 'resnet34', metrics=accuracy).to_fp16()","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:30.616315Z","iopub.execute_input":"2024-03-05T17:55:30.616735Z","iopub.status.idle":"2024-03-05T17:55:31.911480Z","shell.execute_reply.started":"2024-03-05T17:55:30.616706Z","shell.execute_reply":"2024-03-05T17:55:31.910413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fine_tune(12, 0.01)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T17:55:51.440829Z","iopub.execute_input":"2024-03-05T17:55:51.441343Z","iopub.status.idle":"2024-03-05T18:14:03.593338Z","shell.execute_reply.started":"2024-03-05T17:55:51.441299Z","shell.execute_reply":"2024-03-05T18:14:03.592236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting to note: each epoch took 30-50 seconds longer than the model trained on average spectrograms---not sure if that's just a fluke or if the different image structure changes the training.","metadata":{}},{"cell_type":"code","source":"learn.save('/kaggle/working/hms_hbac_resnet34_stacked', with_opt=False)","metadata":{"execution":{"iopub.status.busy":"2024-03-05T18:16:11.669733Z","iopub.execute_input":"2024-03-05T18:16:11.670704Z","iopub.status.idle":"2024-03-05T18:16:11.888144Z","shell.execute_reply.started":"2024-03-05T18:16:11.670663Z","shell.execute_reply":"2024-03-05T18:16:11.887020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's it for this notebook, next, I'll execute a submission from my [submission notebook](https://www.kaggle.com/vishalbakshi/hms-hbac-fastai-stacked-images-submit) to see how this approach's score compares to the averaged spectrogram approach.\n","metadata":{}}]}