{"metadata":{"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.9.13"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install timm==0.6.13 huggingface_hub kaggle pynvml -Uqq\n!jupyter notebook --ServerApp.iopub_data_rate_limit=1.0e10\n!jupyter notebook --ServerApp.iopub_msg_rate_limit=1.0e10\n# !kaggle datasets download -d vishalbakshi/hms-hbac-training-spectrogram-images\n# zipfile.ZipFile('hms-hbac-training-spectrogram-images.zip').extractall('hms-hbac-training-spectrogram-images')","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:33:56.037231Z","iopub.status.busy":"2024-04-01T18:33:56.036631Z","iopub.status.idle":"2024-04-01T18:34:02.133229Z","shell.execute_reply":"2024-04-01T18:34:02.132454Z","shell.execute_reply.started":"2024-04-01T18:33:56.037163Z"},"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nimport timm\nimport gc\n\nfrom fastai.vision.all import *\nfrom fastcore.parallel import *\n\n#path = Path('/kaggle/input/hms-hbac-training-spectrogram-images/train_spectrograms')\npath = Path('/notebooks/hms-hbac-training-spectrogram-images/train_spectrograms')\npath.ls()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:34:02.134603Z","iopub.status.busy":"2024-04-01T18:34:02.134375Z","iopub.status.idle":"2024-04-01T18:34:07.081033Z","shell.execute_reply":"2024-04-01T18:34:07.080594Z","shell.execute_reply.started":"2024-04-01T18:34:02.134581Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background","metadata":{}},{"cell_type":"markdown","source":"In this notebook, I'll train the largest versions of the five models that resulted in the highest score (when weighted twice in a 10-model ensemble) from my previous notebook:\n\n- `convnext_xlarge_in22k`\n- `vit_giant_patch14_224_clip_laion2b` (this is patch14, not the patch16 I was using, but it's the largest 224x224 pretrained model in this timm version)\n\n\nI will not be re-training `swinv2_large_window12_192_22k` model J as this is the largest 192x192 swinv2 model in this timm version. I'll re-use the one I've already trained in the next submission(s).\n","metadata":{}},{"cell_type":"markdown","source":"|Model Name|item method|item img size|batch_tfms|Public Score|\n|:-:|:-:|:-:|:-:|:-:|\n|convnext 1|squish|(311, 400)|None|1.39|\n|convnext 3|crop|(400, 311)|None|1.40|\n|swinv2 J|crop|(400, 311)|aug_transforms(size=192, min_scale=0.75)|1.40|\n|convnext 4|squish|(400, 311)|None|1.41|\n|vit AW|crop|256 --> (400, 311) --> (320, 512)|RandomResizedCropGPU(size=224, min_scale=1.0)|1.41|","metadata":{}},{"cell_type":"code","source":"def train(fn, arch, item, batch, epochs, accum=4):        \n    dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=item,\n        batch_tfms=batch,\n        bs=64//accum)\n    \n    cbs = GradientAccumulation(64) if accum else []\n    learn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()\n    learn.fine_tune(epochs, 0.01)\n    print(error_rate(*learn.tta(dl=dls.valid)))\n    if fn is not None: learn.save(fn, with_opt=False)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:18:54.956688Z","iopub.status.busy":"2024-04-01T04:18:54.956013Z","iopub.status.idle":"2024-04-01T04:18:54.961083Z","shell.execute_reply":"2024-04-01T04:18:54.960629Z","shell.execute_reply.started":"2024-04-01T04:18:54.956669Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def progressive_resizing(fn, sizes, method, batch, arch, epochs, pad_mode=PadMode.Reflection):\n    dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize(sizes[0], method=method, pad_mode=pad_mode),\n        batch_tfms=batch,\n        bs=64//16)\n    \n    cbs = GradientAccumulation(64)\n    learn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()\n    learn.fine_tune(epochs[0], 0.01)\n\n    dls = ImageDataLoaders.from_folder(\n            path, \n            valid_pct=0.2, \n            item_tfms=Resize(sizes[1], method=method, pad_mode=pad_mode),\n            batch_tfms=batch,\n            bs=64//16)\n\n    learn.dls = dls\n    learn.fine_tune(epochs[1], 0.01)\n\n    dls = ImageDataLoaders.from_folder(\n            path, \n            valid_pct=0.2, \n            item_tfms=Resize(sizes[2], method=method, pad_mode=pad_mode),\n            batch_tfms=batch,\n            bs=64//16)\n\n    learn.dls = dls\n    learn.fine_tune(epochs[2], 0.01)\n    print(error_rate(*learn.tta(dl=dls.valid)))\n    if fn is not None: learn.save(fn, with_opt=False)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:34:09.562321Z","iopub.status.busy":"2024-04-01T18:34:09.561303Z","iopub.status.idle":"2024-04-01T18:34:09.569665Z","shell.execute_reply":"2024-04-01T18:34:09.569037Z","shell.execute_reply.started":"2024-04-01T18:34:09.562275Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ndef report_gpu():\n    print(torch.cuda.list_gpu_processes())\n    gc.collect()\n    torch.cuda.empty_cache()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:01:38.342566Z","iopub.status.busy":"2024-04-01T18:01:38.341624Z","iopub.status.idle":"2024-04-01T18:01:38.345606Z","shell.execute_reply":"2024-04-01T18:01:38.345051Z","shell.execute_reply.started":"2024-04-01T18:01:38.342556Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report_gpu()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:01:39.409526Z","iopub.status.busy":"2024-04-01T18:01:39.408743Z","iopub.status.idle":"2024-04-01T18:01:39.508687Z","shell.execute_reply":"2024-04-01T18:01:39.508197Z","shell.execute_reply.started":"2024-04-01T18:01:39.409499Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## convnext_xlarge_in22k","metadata":{}},{"cell_type":"markdown","source":"### Model 1","metadata":{}},{"cell_type":"code","source":"arch = 'convnext_xlarge_in22k'","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:19:19.552187Z","iopub.status.busy":"2024-04-01T04:19:19.551461Z","iopub.status.idle":"2024-04-01T04:19:19.555406Z","shell.execute_reply":"2024-04-01T04:19:19.554791Z","shell.execute_reply.started":"2024-04-01T04:19:19.552155Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arch","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:19:20.665072Z","iopub.status.busy":"2024-04-01T04:19:20.664341Z","iopub.status.idle":"2024-04-01T04:19:20.668640Z","shell.execute_reply":"2024-04-01T04:19:20.668183Z","shell.execute_reply.started":"2024-04-01T04:19:20.665046Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(\n    fn=None, \n    arch=arch, \n    item=Resize((311,400), method='squish'), \n    batch=None, \n    epochs=10, \n    accum=8)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T02:53:54.133065Z","iopub.status.busy":"2024-04-01T02:53:54.132362Z","iopub.status.idle":"2024-04-01T04:00:42.405673Z","shell.execute_reply":"2024-04-01T04:00:42.404994Z","shell.execute_reply.started":"2024-04-01T02:53:54.133037Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The model's error rate is stalling, the validation loss is starting to increase slightly, and the TTA Validation Error Rate is about 16% higher than the smaller models (although it could just be a bad validation set).\n\nI'll run `lr_find` to see if I need to change the learning rate from 0.01 to something else.","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//8)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:03:28.428372Z","iopub.status.busy":"2024-04-01T04:03:28.427764Z","iopub.status.idle":"2024-04-01T04:03:38.049787Z","shell.execute_reply":"2024-04-01T04:03:38.049008Z","shell.execute_reply.started":"2024-04-01T04:03:28.428343Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:05:47.952105Z","iopub.status.busy":"2024-04-01T04:05:47.951483Z","iopub.status.idle":"2024-04-01T04:06:12.465648Z","shell.execute_reply":"2024-04-01T04:06:12.464889Z","shell.execute_reply.started":"2024-04-01T04:05:47.952076Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I have ran `lr_find` 3 times and the plot looks like it does above. I wonder if changing the batch size would improve it:","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=16)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:19:31.296492Z","iopub.status.busy":"2024-04-01T04:19:31.295717Z","iopub.status.idle":"2024-04-01T04:19:36.940010Z","shell.execute_reply":"2024-04-01T04:19:36.938988Z","shell.execute_reply.started":"2024-04-01T04:19:31.296467Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:19:39.909822Z","iopub.status.busy":"2024-04-01T04:19:39.908764Z","iopub.status.idle":"2024-04-01T04:20:28.153253Z","shell.execute_reply":"2024-04-01T04:20:28.152440Z","shell.execute_reply.started":"2024-04-01T04:19:39.909778Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:21:01.724347Z","iopub.status.busy":"2024-04-01T04:21:01.723897Z","iopub.status.idle":"2024-04-01T04:21:41.964062Z","shell.execute_reply":"2024-04-01T04:21:41.963435Z","shell.execute_reply.started":"2024-04-01T04:21:01.724321Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:26:37.769346Z","iopub.status.busy":"2024-04-01T04:26:37.768361Z","iopub.status.idle":"2024-04-01T04:27:19.790106Z","shell.execute_reply":"2024-04-01T04:27:19.789485Z","shell.execute_reply.started":"2024-04-01T04:26:37.769312Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I wonder if the `item_tfms` and `batch_tfms` that worked well for the smaller model no longer apply? I'll try a different set of transforms:","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=RandomResizedCropGPU(size=224, min_scale=1.0),\n        bs=16)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:28:32.830959Z","iopub.status.busy":"2024-04-01T04:28:32.830089Z","iopub.status.idle":"2024-04-01T04:28:41.815196Z","shell.execute_reply":"2024-04-01T04:28:41.809571Z","shell.execute_reply.started":"2024-04-01T04:28:32.830927Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:28:44.235740Z","iopub.status.busy":"2024-04-01T04:28:44.235024Z","iopub.status.idle":"2024-04-01T04:29:05.518072Z","shell.execute_reply":"2024-04-01T04:29:05.517312Z","shell.execute_reply.started":"2024-04-01T04:28:44.235710Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:29:14.825482Z","iopub.status.busy":"2024-04-01T04:29:14.824706Z","iopub.status.idle":"2024-04-01T04:29:36.411211Z","shell.execute_reply":"2024-04-01T04:29:36.410399Z","shell.execute_reply.started":"2024-04-01T04:29:14.825445Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:29:49.200687Z","iopub.status.busy":"2024-04-01T04:29:49.200201Z","iopub.status.idle":"2024-04-01T04:30:10.739004Z","shell.execute_reply":"2024-04-01T04:30:10.738125Z","shell.execute_reply.started":"2024-04-01T04:29:49.200662Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize(400),\n        batch_tfms=RandomResizedCropGPU(size=224, min_scale=1.0),\n        bs=16)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:30:51.144542Z","iopub.status.busy":"2024-04-01T04:30:51.143673Z","iopub.status.idle":"2024-04-01T04:30:58.425564Z","shell.execute_reply":"2024-04-01T04:30:58.424870Z","shell.execute_reply.started":"2024-04-01T04:30:51.144510Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:31:00.025732Z","iopub.status.busy":"2024-04-01T04:31:00.025118Z","iopub.status.idle":"2024-04-01T04:31:21.357224Z","shell.execute_reply":"2024-04-01T04:31:21.356278Z","shell.execute_reply.started":"2024-04-01T04:31:00.025705Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just for a sanity check, I'll compare this with the `lr_find` output for the smaller version of architecture.","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, 'convnext_large_in22k', metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:33:00.425571Z","iopub.status.busy":"2024-04-01T04:33:00.425266Z","iopub.status.idle":"2024-04-01T04:33:54.198402Z","shell.execute_reply":"2024-04-01T04:33:54.197864Z","shell.execute_reply.started":"2024-04-01T04:33:00.425546Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:34:36.508617Z","iopub.status.busy":"2024-04-01T04:34:36.507707Z","iopub.status.idle":"2024-04-01T04:34:51.974944Z","shell.execute_reply":"2024-04-01T04:34:51.974206Z","shell.execute_reply.started":"2024-04-01T04:34:36.508586Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting. I'm getting a similar looking lr_find plot for the smaller architecture. I'll try the transforms that resulted in a low error rate.","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=16)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_large_in22k', metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:37:15.805251Z","iopub.status.busy":"2024-04-01T04:37:15.804831Z","iopub.status.idle":"2024-04-01T04:37:21.889687Z","shell.execute_reply":"2024-04-01T04:37:21.888953Z","shell.execute_reply.started":"2024-04-01T04:37:15.805220Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:37:24.665601Z","iopub.status.busy":"2024-04-01T04:37:24.664573Z","iopub.status.idle":"2024-04-01T04:37:53.534407Z","shell.execute_reply":"2024-04-01T04:37:53.533623Z","shell.execute_reply.started":"2024-04-01T04:37:24.665563Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I noticed that the `lr_find` progress bar doesn't go through the entire dataset. It stops at 95/556 steps, or about 20%. I'll increase the `num_it` parameter and see if that changes anything.","metadata":{}},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:39:19.605249Z","iopub.status.busy":"2024-04-01T04:39:19.604625Z","iopub.status.idle":"2024-04-01T04:41:21.660146Z","shell.execute_reply":"2024-04-01T04:41:21.659208Z","shell.execute_reply.started":"2024-04-01T04:39:19.605219Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:42:15.101298Z","iopub.status.busy":"2024-04-01T04:42:15.100451Z","iopub.status.idle":"2024-04-01T04:44:16.807191Z","shell.execute_reply":"2024-04-01T04:44:16.806216Z","shell.execute_reply.started":"2024-04-01T04:42:15.101242Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:44:16.809294Z","iopub.status.busy":"2024-04-01T04:44:16.808730Z","iopub.status.idle":"2024-04-01T04:46:17.237938Z","shell.execute_reply":"2024-04-01T04:46:17.237150Z","shell.execute_reply.started":"2024-04-01T04:44:16.809269Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I get a slightly better `lr_find` plot. I'll run it on the smallest `convnext` architecture that I have tried, `convnext_small_in22k`.","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=16)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_small_in22k', metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:49:46.884808Z","iopub.status.busy":"2024-04-01T04:49:46.883956Z","iopub.status.idle":"2024-04-01T04:49:59.817935Z","shell.execute_reply":"2024-04-01T04:49:59.817197Z","shell.execute_reply.started":"2024-04-01T04:49:46.884775Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:50:02.229365Z","iopub.status.busy":"2024-04-01T04:50:02.228630Z","iopub.status.idle":"2024-04-01T04:50:59.891461Z","shell.execute_reply":"2024-04-01T04:50:59.890495Z","shell.execute_reply.started":"2024-04-01T04:50:02.229338Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:50:59.893513Z","iopub.status.busy":"2024-04-01T04:50:59.893236Z","iopub.status.idle":"2024-04-01T04:51:56.968516Z","shell.execute_reply":"2024-04-01T04:51:56.967934Z","shell.execute_reply.started":"2024-04-01T04:50:59.893489Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:51:56.969790Z","iopub.status.busy":"2024-04-01T04:51:56.969385Z","iopub.status.idle":"2024-04-01T04:52:54.542684Z","shell.execute_reply":"2024-04-01T04:52:54.541909Z","shell.execute_reply.started":"2024-04-01T04:51:56.969766Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I'm curious---if I set the `seed` to `42` as I did when I trained the original `convnext_small_in22k` model, will I get a similar result as when I first trained, when the TTA Validation Error Rate I got was 0.3543?","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//4)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_small_in22k', metrics=error_rate, cbs=cbs).to_fp16()\nlearn.fine_tune(7, 0.01)\nprint(error_rate(*learn.tta(dl=dls.valid)))","metadata":{"execution":{"iopub.execute_input":"2024-04-01T04:57:33.180461Z","iopub.status.busy":"2024-04-01T04:57:33.179793Z","iopub.status.idle":"2024-04-01T05:12:04.698414Z","shell.execute_reply":"2024-04-01T05:12:04.697686Z","shell.execute_reply.started":"2024-04-01T04:57:33.180435Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm okay with this TTA error rate since it's close enough to the original (0.3543). Now that I've confirmed that there's nothing wonky about how I'm training the model (at least, nothing that I can tell), I'll re-run `lr_find` for the `xlarge` model with `num_it=500`. I'll set the `seed` so it's comparable.","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//4)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_xlarge_in22k', metrics=error_rate, cbs=cbs).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T05:15:27.535649Z","iopub.status.busy":"2024-04-01T05:15:27.535067Z","iopub.status.idle":"2024-04-01T05:15:35.830086Z","shell.execute_reply":"2024-04-01T05:15:35.829503Z","shell.execute_reply.started":"2024-04-01T05:15:27.535622Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T05:15:59.534564Z","iopub.status.busy":"2024-04-01T05:15:59.533722Z","iopub.status.idle":"2024-04-01T05:19:02.606989Z","shell.execute_reply":"2024-04-01T05:19:02.606131Z","shell.execute_reply.started":"2024-04-01T05:15:59.534537Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T05:19:02.609026Z","iopub.status.busy":"2024-04-01T05:19:02.608737Z","iopub.status.idle":"2024-04-01T05:22:00.432756Z","shell.execute_reply":"2024-04-01T05:22:00.432081Z","shell.execute_reply.started":"2024-04-01T05:19:02.608997Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(num_it=500)","metadata":{"execution":{"iopub.execute_input":"2024-04-01T05:22:00.434393Z","iopub.status.busy":"2024-04-01T05:22:00.434183Z","iopub.status.idle":"2024-04-01T05:24:57.796583Z","shell.execute_reply":"2024-04-01T05:24:57.796014Z","shell.execute_reply.started":"2024-04-01T05:22:00.434371Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll train `convnext_xlarge_in22k` three more times---once with the same learning rate (0.01), once with a smaller learning rate (0.001) and once with a larger learning rate (0.03). I'll train for a much larger number of epochs (20), to make sure I'm not training for too few with 10.","metadata":{}},{"cell_type":"markdown","source":"### LR = 0.001","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//8)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_xlarge_in22k', metrics=error_rate, cbs=cbs).to_fp16()\nlearn.fine_tune(20, 0.001)\nprint(error_rate(*learn.tta(dl=dls.valid)))","metadata":{"execution":{"iopub.execute_input":"2024-04-01T05:45:08.660159Z","iopub.status.busy":"2024-04-01T05:45:08.659697Z","iopub.status.idle":"2024-04-01T07:53:06.808209Z","shell.execute_reply":"2024-04-01T07:53:06.807543Z","shell.execute_reply.started":"2024-04-01T05:45:08.660135Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LR = 0.01","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//8)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_xlarge_in22k', metrics=error_rate, cbs=cbs).to_fp16()\nlearn.fine_tune(20, 0.01)\nprint(error_rate(*learn.tta(dl=dls.valid)))","metadata":{"execution":{"iopub.execute_input":"2024-04-01T11:17:42.201231Z","iopub.status.busy":"2024-04-01T11:17:42.200969Z","iopub.status.idle":"2024-04-01T13:24:13.281938Z","shell.execute_reply":"2024-04-01T13:24:13.280944Z","shell.execute_reply.started":"2024-04-01T11:17:42.201212Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report_gpu()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T13:42:47.191684Z","iopub.status.busy":"2024-04-01T13:42:47.191319Z","iopub.status.idle":"2024-04-01T13:42:47.309895Z","shell.execute_reply":"2024-04-01T13:42:47.308681Z","shell.execute_reply.started":"2024-04-01T13:42:47.191651Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LR = 0.03","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=Resize((311,400), method='squish'),\n        batch_tfms=None,\n        bs=64//8)\n    \ncbs = GradientAccumulation(64)\nlearn = vision_learner(dls, 'convnext_xlarge_in22k', metrics=error_rate, cbs=cbs).to_fp16()\nlearn.fine_tune(20, 0.03)\nprint(error_rate(*learn.tta(dl=dls.valid)))","metadata":{"execution":{"iopub.execute_input":"2024-04-01T13:42:50.087428Z","iopub.status.busy":"2024-04-01T13:42:50.087147Z","iopub.status.idle":"2024-04-01T15:49:11.035973Z","shell.execute_reply":"2024-04-01T15:49:11.035189Z","shell.execute_reply.started":"2024-04-01T13:42:50.087407Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report_gpu()","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:33:42.677059Z","iopub.status.busy":"2024-04-01T18:33:42.676541Z","iopub.status.idle":"2024-04-01T18:33:42.807659Z","shell.execute_reply":"2024-04-01T18:33:42.806918Z","shell.execute_reply.started":"2024-04-01T18:33:42.677039Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Learning Rate Summary","metadata":{}},{"cell_type":"markdown","source":"Here is what I observed in the above three training runs:\n\n- The training was the most stable when the learning rate was 0.001. The validation loss started diverging at around 8-10 epochs, with a 40% error rate.\n- Both learning rates of 0.01 and 0.03 resulted in the validation loss diverging at around 4-6 epochs, with a 40% error rate. In both cases, the validation loss kept increasing but the error rate kept decreasing. My guess/interpretation for what this means: the model was overfitting the training data to the point where the validation loss was increasing, but because the error rate kept decreasing, the training and validation set images are similar to each other. That's something I'll ask around about in the fastai community to get a better understanding.","metadata":{}},{"cell_type":"markdown","source":"## vit_giant_patch14_224_clip_laion2b","metadata":{}},{"cell_type":"markdown","source":"### Model AW","metadata":{}},{"cell_type":"code","source":"arch = 'vit_giant_patch14_224_clip_laion2b'","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:53:39.301561Z","iopub.status.busy":"2024-04-01T18:53:39.300914Z","iopub.status.idle":"2024-04-01T18:53:39.304910Z","shell.execute_reply":"2024-04-01T18:53:39.304119Z","shell.execute_reply.started":"2024-04-01T18:53:39.301530Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arch","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:53:40.273586Z","iopub.status.busy":"2024-04-01T18:53:40.273338Z","iopub.status.idle":"2024-04-01T18:53:40.278051Z","shell.execute_reply":"2024-04-01T18:53:40.277322Z","shell.execute_reply.started":"2024-04-01T18:53:40.273569Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"progressive_resizing(\n    fn=None, \n    sizes=[256, (400, 311), (320, 512)], \n    method='crop', \n    batch=RandomResizedCropGPU(size=224, min_scale=1.0), \n    arch=arch, \n    epochs=[4, 7, 10])","metadata":{"execution":{"iopub.execute_input":"2024-04-01T18:34:32.785111Z","iopub.status.busy":"2024-04-01T18:34:32.784105Z","iopub.status.idle":"2024-04-01T18:43:54.555782Z","shell.execute_reply":"2024-04-01T18:43:54.554585Z","shell.execute_reply.started":"2024-04-01T18:34:32.785084Z"},"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even a batch size of 4, which would take over 3 hours to train, is too large for this sized model. Given that I have only a week left in the competition, and that the convnext_xlarge didn't improve the TTA Validation Error Rate significantly, I don't think pursuing a larger model size will greatly benefit my Kaggle score.","metadata":{}},{"cell_type":"markdown","source":"## Next Steps","metadata":{}},{"cell_type":"markdown","source":"Instead of training these much larger models, which are not showing signs of significantly improving my Kaggle score, I'll shift my attention this last week of the competition to exploring the data provided. I have not yet tried to train models on the EEG data and use it in tandem with the spectrogram data. If after exploring that there remains any time, I'll train larger models to see if I can slightly improve my Kaggle public score.","metadata":{}}]}