{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7781194,"sourceType":"datasetVersion","datasetId":4553461},{"sourceId":15638,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":13023},{"sourceId":15640,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":13024},{"sourceId":15641,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":13025},{"sourceId":15642,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":13026},{"sourceId":15643,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":13027}],"dockerImageVersionId":30665,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"code","source":"import warnings\nimport timm\nfrom fastai.vision.all import *\nfrom fastcore.parallel import *\n\npath = Path('/kaggle/input/hms-harmful-brain-activity-classification')\n\npath.ls()","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-14T06:43:56.732929Z","iopub.execute_input":"2024-03-14T06:43:56.733258Z","iopub.status.idle":"2024-03-14T06:44:08.541531Z","shell.execute_reply.started":"2024-03-14T06:43:56.733231Z","shell.execute_reply":"2024-03-14T06:44:08.540526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background\n\nThis is the fifth notebook in a series of 5 notebooks where I train different `convnext` image classifiers on training spectrogram images using the fastai library:\n\n- [Part 0](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-planning-small-model-experiments): I plan out my initial small model experiments and create a [Kaggle dataset](https://www.kaggle.com/datasets/vishalbakshi/hms-hbac-training-spectrogram-images) with training spectrogram images.\n- [Part 1 [Train]](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-convnext-small-pt-1-train): I train 48 variants of `convnext_small_in22k` using different `ImageDataLoaders`.\n- [Part 1 [Analysis]](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-convnext-small-pt-1-analysis): I analyze the results from Part 1, run a few more trainings and pick the top convnext models for submission.\n- [Part 2 Train](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-convnext-small-pt-2-train): I train those top convnext models and export them to Kaggle.\n- **Part 2 [Submit] (You are here): I submit those models individually and as ensembles, and document their Kaggle Public Score.**\n\nI'll follow the same approach (experiment -> train and export top models -> submit -> document Kaggle Public Score) for three other families: `vit`, `swin` and `swinv2`. Once I have identified the best `small` models, I'll train their `large` versions, submit them and document the results. I have taken this general approach from Jeremy Howard's [Road to the Top](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1) notebook series (although his notebooks and presentation is much more efficient).","metadata":{}},{"cell_type":"markdown","source":"In this notebook I submit predictions from each of the 5 `convnext_small_in22k` models I've trained, as well as ensembles of all 5 models.","metadata":{}},{"cell_type":"markdown","source":"## Submission Results","metadata":{}},{"cell_type":"markdown","source":"Here are the results for the single-model submissions I made using my `convnext_small_in22k` models, sorted from best to worse Public Score:\n\n|Model #|item method|item img size|batch img size|Public Score|\n|:-:|:-:|:-:|:-:|:-:|\n|3|crop|(400, 311)|None|1.77|\n|2|pad|400|None|1.89|\n|5|pad|(320, 512)|None|1.95|\n|4|squish|(400, 311)|None|1.95|\n|1|squish|(311, 400)|None|2.08|\n\nHere are the results of 5- and 6-model ensemble submissions using the above models. In the 6-model submissions I weighted one of the models twice. \n\n|Model Weighted Twice|Public Score|\n|:-:|:-:|\n|--|1.43|\n|4|1.43|\n|1|1.44|\n|3|1.44|\n|2|1.45|\n|5|1.46|\n\nI want to pick 3 models to move forward with but it's unclear to me which models I should pick. Models 3, 2 and 5 got better individual scores than 1 and 4 (which performed better in ensembles when weighted twice). I tried the 10 combinations of three-model combinations and documented their results:\n\n|Ensemble|Public Score|\n|:-:|:-:|\n|1,2,4|1.45|\n|1,3,4|1.45|\n|1,4,5|1.5|\n|3,4,5|1.51|\n|2,3,4|1.52|\n|1,2,3|1.54|\n|1,3,5|1.54|\n|2,4,5|1.55|\n|1,2,5|1.57|\n|2,3,5|1.58|\n\n\nI'll select models 1, 3 and 4. Models 1 and 4 seem to work well in ensembles scoring the top three 3- and 5-model ensemble scores. Model 3 was the best individual model submission. At least two of the models 1, 3 and 4 are included in each the top five 3-model ensemble submissions.  You could also argue to include model 2, but I'd like to limit my ensemble to 3 models because I'm running out of submissions with the deadline approaching in less than a month. \n","metadata":{}},{"cell_type":"markdown","source":"## Generate Test Data Images","metadata":{}},{"cell_type":"markdown","source":"My models are image classifiers trained on spectrogram images so I need to convert the test parquet data to images. I'm referencing the following notebooks:\n\n- [HMS - HBAC - Fastai Starter](https://www.kaggle.com/code/sonujha090/hms-hbac-fastai-starter)\n- [HMS-HBAC: KerasCV Starter Notebook](https://www.kaggle.com/code/awsaf49/hms-hbac-kerascv-starter-notebook)","metadata":{}},{"cell_type":"code","source":"# create temporary folders to hold spectrograms\nSPEC_DIR = \"/tmp/dataset/hms-hbac\"\nos.makedirs(SPEC_DIR+'/test_spectrograms', exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:11.324987Z","iopub.execute_input":"2024-03-14T06:46:11.325944Z","iopub.status.idle":"2024-03-14T06:46:11.330758Z","shell.execute_reply.started":"2024-03-14T06:46:11.325907Z","shell.execute_reply":"2024-03-14T06:46:11.329709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_spec(spec_id, split=\"train\"):\n    # read the data\n    data = pd.read_parquet(path/f'{split}_spectrograms'/f'{spec_id}.parquet')\n    \n    # replace NA with 0\n    data = data.fillna(0)\n    \n    # convert DataFrame to array\n    data = data.values[:, 1:]\n    \n    # transpose\n    data = data.T\n    data = data.astype(\"float32\")\n    \n    # convert array to PILImage\n    im = PILImage.create(Image.fromarray((data * 255).astype(np.uint8)))\n    im.save(f\"{SPEC_DIR}/{split}_spectrograms/{spec_id}.png\")","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:11.983722Z","iopub.execute_input":"2024-03-14T06:46:11.984062Z","iopub.status.idle":"2024-03-14T06:46:11.989889Z","shell.execute_reply.started":"2024-03-14T06:46:11.984035Z","shell.execute_reply":"2024-03-14T06:46:11.989023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_csv(path/'test.csv')\ntest_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:13.678563Z","iopub.execute_input":"2024-03-14T06:46:13.679235Z","iopub.status.idle":"2024-03-14T06:46:13.701401Z","shell.execute_reply.started":"2024-03-14T06:46:13.679204Z","shell.execute_reply":"2024-03-14T06:46:13.700537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"spec_ids = test_df['spectrogram_id'].unique()\nlen(spec_ids)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:14.058672Z","iopub.execute_input":"2024-03-14T06:46:14.059367Z","iopub.status.idle":"2024-03-14T06:46:14.068927Z","shell.execute_reply.started":"2024-03-14T06:46:14.059338Z","shell.execute_reply":"2024-03-14T06:46:14.067888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")\nparallel(process_spec, spec_ids, split='test', n_workers=4)\nwarnings.filterwarnings(\"default\")","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:14.518490Z","iopub.execute_input":"2024-03-14T06:46:14.519091Z","iopub.status.idle":"2024-03-14T06:46:14.892599Z","shell.execute_reply.started":"2024-03-14T06:46:14.519064Z","shell.execute_reply":"2024-03-14T06:46:14.891414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"PILImage.create(Path('/tmp/dataset/hms-hbac/test_spectrograms').ls()[0])","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:16.412156Z","iopub.execute_input":"2024-03-14T06:46:16.412832Z","iopub.status.idle":"2024-03-14T06:46:16.471628Z","shell.execute_reply.started":"2024-03-14T06:46:16.412797Z","shell.execute_reply":"2024-03-14T06:46:16.470780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating the `DataLoaders` Object","metadata":{}},{"cell_type":"markdown","source":"I already have created a dataset with training images, so I'll load that into my notebook and create a training path `trn_path` to use in my `DataLoaders`.","metadata":{}},{"cell_type":"code","source":"trn_path = Path('/kaggle/input/hms-hbac-training-spectrogram-images/train_spectrograms')","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:19.512487Z","iopub.execute_input":"2024-03-14T06:46:19.513103Z","iopub.status.idle":"2024-03-14T06:46:19.517407Z","shell.execute_reply.started":"2024-03-14T06:46:19.513070Z","shell.execute_reply":"2024-03-14T06:46:19.516439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each model was trained with different `item_tfms` so I'll list them out here.","metadata":{}},{"cell_type":"code","source":"item1 = Resize((311,400), method='squish')\nitem2 = Resize((400), method=ResizeMethod.Pad, pad_mode=PadMode.Zeros)\nitem3 = Resize((400,311))\nitem4 = Resize((400,311), method='squish')\nitem5 = Resize((320,512), method=ResizeMethod.Pad, pad_mode=PadMode.Zeros)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:21.762986Z","iopub.execute_input":"2024-03-14T06:46:21.763352Z","iopub.status.idle":"2024-03-14T06:46:21.769983Z","shell.execute_reply.started":"2024-03-14T06:46:21.763325Z","shell.execute_reply":"2024-03-14T06:46:21.768950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ensemble","metadata":{}},{"cell_type":"markdown","source":"I'll create a list of `DataLoaders`, one for each of the five models that I am using in this ensemble.","metadata":{}},{"cell_type":"code","source":"dls_list = []\nitems = [item1, item2, item3, item4, item5]\n\nfor item in items:\n    dls = ImageDataLoaders.from_folder(\n        trn_path, \n        valid_pct=0.2, \n        item_tfms=item,\n        batch_tfms=None,\n        bs=16)\n    \n    dls_list.append(dls)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:46:25.095116Z","iopub.execute_input":"2024-03-14T06:46:25.095815Z","iopub.status.idle":"2024-03-14T06:46:45.338848Z","shell.execute_reply.started":"2024-03-14T06:46:25.095784Z","shell.execute_reply":"2024-03-14T06:46:45.337806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls_list","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:47:01.711408Z","iopub.execute_input":"2024-03-14T06:47:01.711779Z","iopub.status.idle":"2024-03-14T06:47:01.718718Z","shell.execute_reply.started":"2024-03-14T06:47:01.711751Z","shell.execute_reply":"2024-03-14T06:47:01.717631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, I'll create a list of `Learner`s, one for each model.","metadata":{}},{"cell_type":"code","source":"model_paths = [\n    Path('/kaggle/input/hms_hbac_convnext_small_in22k/pytorch/1/1/hms_hbac_convnext_small_in22k_1'),\n    Path('/kaggle/input/hms_hbac_convnext_small_in22k/pytorch/2/1/hms_hbac_convnext_small_in22k_2'),\n    Path('/kaggle/input/hms_hbac_convnext_small_in22k/pytorch/3/1/hms_hbac_convnext_small_in22k_3'),\n    Path('/kaggle/input/hms_hbac_convnext_small_in22k/pytorch/4/1/hms_hbac_convnext_small_in22k_4'),\n    Path('/kaggle/input/hms_hbac_convnext_small_in22k/pytorch/5/1/hms_hbac_convnext_small_in22k_5')\n]","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:47:03.851780Z","iopub.execute_input":"2024-03-14T06:47:03.852147Z","iopub.status.idle":"2024-03-14T06:47:03.858004Z","shell.execute_reply.started":"2024-03-14T06:47:03.852108Z","shell.execute_reply":"2024-03-14T06:47:03.856838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learners = []\n\nfor idx, model_path in enumerate(model_paths):\n    learn = vision_learner(dls_list[idx], 'convnext_small_in22k', pretrained=False)\n    learn.model_dir = '/kaggle/working/'\n    learn.load(model_path)\n    learners.append(learn)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:47:04.386900Z","iopub.execute_input":"2024-03-14T06:47:04.387246Z","iopub.status.idle":"2024-03-14T06:47:21.313797Z","shell.execute_reply.started":"2024-03-14T06:47:04.387220Z","shell.execute_reply":"2024-03-14T06:47:21.312756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learners","metadata":{"execution":{"iopub.status.busy":"2024-03-14T06:47:24.398402Z","iopub.execute_input":"2024-03-14T06:47:24.399039Z","iopub.status.idle":"2024-03-14T06:47:24.405140Z","shell.execute_reply.started":"2024-03-14T06:47:24.399008Z","shell.execute_reply":"2024-03-14T06:47:24.404162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, I'll get a `DataFrame` of probabilities for each test image from each `Learner` in my 3-model ensemble.","metadata":{}},{"cell_type":"code","source":"probs_df_list = []\n\ntst_files = get_image_files(SPEC_DIR+'/test_spectrograms')\n\nfor idx in range(5):\n    # create test DataLoader\n    tst_dl = dls_list[idx].test_dl(tst_files)\n    \n    # get TTA predictions\n    probs,_= learners[idx].tta(dl=tst_dl)\n    \n    # formatting\n    probs_df = pd.DataFrame(probs, columns=dls_list[idx].vocab)\n    probs_df['eeg_id'] = test_df['eeg_id']\n    probs_df = probs_df[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]\n    \n    probs_df_list.append(probs_df)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T07:00:27.063206Z","iopub.execute_input":"2024-03-14T07:00:27.064259Z","iopub.status.idle":"2024-03-14T07:00:33.151021Z","shell.execute_reply.started":"2024-03-14T07:00:27.064209Z","shell.execute_reply":"2024-03-14T07:00:33.149908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Concatenate all `DataFrames`:","metadata":{}},{"cell_type":"code","source":"all_probs_df = pd.concat([probs_df_list[i] for i in [0, 2, 2, 2, 3]])\nall_probs_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-14T07:16:44.359831Z","iopub.execute_input":"2024-03-14T07:16:44.360278Z","iopub.status.idle":"2024-03-14T07:16:44.370838Z","shell.execute_reply.started":"2024-03-14T07:16:44.360243Z","shell.execute_reply":"2024-03-14T07:16:44.369884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Take the mean value of each vote by `eeg_id`","metadata":{}},{"cell_type":"code","source":"final_probs = all_probs_df.groupby('eeg_id').mean().reset_index()","metadata":{"execution":{"iopub.status.busy":"2024-03-14T07:16:48.398120Z","iopub.execute_input":"2024-03-14T07:16:48.398519Z","iopub.status.idle":"2024-03-14T07:16:48.406640Z","shell.execute_reply.started":"2024-03-14T07:16:48.398490Z","shell.execute_reply":"2024-03-14T07:16:48.405470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_probs.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-14T07:16:49.278738Z","iopub.execute_input":"2024-03-14T07:16:49.279109Z","iopub.status.idle":"2024-03-14T07:16:49.291825Z","shell.execute_reply.started":"2024-03-14T07:16:49.279080Z","shell.execute_reply":"2024-03-14T07:16:49.290725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And export it:","metadata":{}},{"cell_type":"code","source":"final_probs.to_csv('submission.csv', index=False)\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-14T07:16:52.554450Z","iopub.execute_input":"2024-03-14T07:16:52.555441Z","iopub.status.idle":"2024-03-14T07:16:53.567769Z","shell.execute_reply.started":"2024-03-14T07:16:52.555401Z","shell.execute_reply":"2024-03-14T07:16:53.566614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}