{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install timm huggingface_hub kaggle -Uqq\n\nimport plotly.express as px\nimport pandas as pd\nimport warnings\nimport timm\nimport gc\n\nfrom fastai.vision.all import *\nfrom fastcore.parallel import *\n\npath = Path('/kaggle/input/hms-hbac-training-spectrogram-images/train_spectrograms')\n# path = Path('/notebooks/hms-hbac-training-spectrogram-images/train_spectrograms')\npath.ls()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.execute_input":"2024-03-10T20:49:16.661914Z","iopub.status.busy":"2024-03-10T20:49:16.661112Z","iopub.status.idle":"2024-03-10T20:49:27.053075Z","shell.execute_reply":"2024-03-10T20:49:27.052501Z","shell.execute_reply.started":"2024-03-10T20:49:16.661853Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background\n\nThis is the third notebook of a series of 5 notebooks where I train different `convnext` image classifiers on training spectrogram images using the fastai library:\n\n- [Part 0](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-planning-small-model-experiments): I plan out my initial small model experiments and create a [Kaggle dataset](https://www.kaggle.com/datasets/vishalbakshi/hms-hbac-training-spectrogram-images) with training spectrogram images.\n- [Part 1 [Train]](https://www.kaggle.com/vishalbakshi/hms-hbac-fastai-image-classifiers-pt-1-train): I train 48 variants of `convnext_small_in22k` using different `ImageDataLoaders`.\n- **Part 1 [Analysis] (You are here): I analyze the results from Part 1, run a few more trainings and pick the top convnext models for submission.**\n- [Part 2 [Train]](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-convnext-small-pt-2-train): I train those top convnext models and export them to Kaggle.\n- [Part 2 [Submit]](https://www.kaggle.com/code/vishalbakshi/hms-hbac-fastai-convnext-small-pt-2-submit): I submit those models individually and as ensembles, and document their Kaggle Public Score.\n\nI'll follow the same approach (experiment -> train and export top models -> submit -> document Kaggle Public Score) for three other families: `vit`, `swin` and `swinv2`. Once I have identified the best `small` models, I'll train their `large` versions, submit them and document the results. I have taken this general approach from Jeremy Howard's [Road to the Top](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1) notebook series (although his notebooks and presentation is much more efficient).","metadata":{}},{"cell_type":"markdown","source":"## Summary","metadata":{}},{"cell_type":"markdown","source":"After training 58 `convnext_small` variants, I will use the following 5 for submission:\n\n\n|`item_tfms` method|`item_tfms` size|`batch_tfms` size|TTA Validation Error Rate|Final Epoch Validation Error Rate|Minutes/Epoch|\n|:-:|:-:|:-:|:-:|:-:|:-:|\n|squish|(311, 400)|None|0.3543|0.354288|1.75|\n|pad|400|None|0.3601|0.359677|2.20|\n|crop|(400, 311)|None|0.3655|0.364616|1.77|\n|squish|(400, 311)|None|0.3646|0.364616|1.75|\n|pad|(320, 512)|None|0.3678|0.373597|2.20|","metadata":{}},{"cell_type":"markdown","source":"## Analyzing Training Results","metadata":{}},{"cell_type":"markdown","source":"I documented my training results from [Part 1 [Train]](https://www.kaggle.com/vishalbakshi/hms-hbac-fastai-image-classifiers-pt-1-train) in a CSV and saved it as a gist. I'll import that gist analyze the results with a goal to pick a few of the top performing models for submission. Note that all images had a fixed height of 400 pixels and a variable width (median 311 pixels).","metadata":{}},{"cell_type":"code","source":"url = 'https://gist.githubusercontent.com/vishalbakshi/d5d4cf1ff73c6daecfd6eb79513f6ada/raw/fa62eb5693b450136be07a73c64cb1d63c5eb52d/hms_hbac_small_model_experiments.csv'\ndf = pd.read_csv(url)\ndf.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-11T21:34:00.455522Z","iopub.execute_input":"2024-03-11T21:34:00.456115Z","iopub.status.idle":"2024-03-11T21:34:00.902501Z","shell.execute_reply.started":"2024-03-11T21:34:00.456076Z","shell.execute_reply":"2024-03-11T21:34:00.901264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Overall, I trained models for:\n\n- three `item_tfms` methods (`squish`, `crop` and `pad`)\n- four `item_tfms` sizes (`400`, `311`, `(311, 400)` and `(400, 311)`)\n- four `batch_tfms` sizes (`None`, `384`, `288`, and `(384, 299)`). \n\nI chose `400` pixels since that is the fixed height of all of the training spectrogram images in my [dataset](https://www.kaggle.com/datasets/vishalbakshi/hms-hbac-training-spectrogram-images), and `311` was the median width. Jeremy recommends using multiples of 32 for `convnext` so I picked `384` and `288` accordingly.","metadata":{}},{"cell_type":"markdown","source":"### Overall Distribution of TTA Validation Error Rate","metadata":{}},{"cell_type":"markdown","source":"|Statistic|Value|\n|:-|:-|\n|Min| 0.3543|\n|Lower Fence| 0.3601|\n|Median| 0.3887|\n|Upper Fence| 0.4091|\n|Max| 0.4091|","metadata":{}},{"cell_type":"code","source":"px.box(df, 'tta_val_error_rate')","metadata":{"execution":{"iopub.status.busy":"2024-03-11T21:34:29.731888Z","iopub.execute_input":"2024-03-11T21:34:29.733308Z","iopub.status.idle":"2024-03-11T21:34:32.080246Z","shell.execute_reply.started":"2024-03-11T21:34:29.733230Z","shell.execute_reply":"2024-03-11T21:34:32.079133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Overall Distribution of Final Epoch Validation Error Rate","metadata":{}},{"cell_type":"markdown","source":"Note that the TTA validation error rate is lower across the board than the final epoch validation error rate, so I'll definitely be using TTA when calculating predictions.","metadata":{}},{"cell_type":"markdown","source":"|Statistic|Value|\n|:-|:-|\n|Min| 0.3543|\n|Lower Fence| 0.3597|\n|Median| 0.4068|\n|Upper Fence| 0.4297|\n|Max| 0.4607|\n","metadata":{}},{"cell_type":"code","source":"px.box(df, 'final_epoch_val_error_rate')","metadata":{"execution":{"iopub.status.busy":"2024-03-11T21:34:38.219312Z","iopub.execute_input":"2024-03-11T21:34:38.219822Z","iopub.status.idle":"2024-03-11T21:34:38.296225Z","shell.execute_reply.started":"2024-03-11T21:34:38.219781Z","shell.execute_reply":"2024-03-11T21:34:38.294810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Top Overall Models by TTA Validation Error Rate","metadata":{}},{"cell_type":"markdown","source":"First, I'll look at my best performing models based on TTA (Test Time Augmentation) validation set error rate.","metadata":{}},{"cell_type":"code","source":"df.sort_values(by='tta_val_error_rate')[:10]","metadata":{"execution":{"iopub.execute_input":"2024-03-10T04:40:12.280911Z","iopub.status.busy":"2024-03-10T04:40:12.280089Z","iopub.status.idle":"2024-03-10T04:40:12.299562Z","shell.execute_reply":"2024-03-10T04:40:12.298552Z","shell.execute_reply.started":"2024-03-10T04:40:12.280876Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- My best performing model was one with `item_tfms=Resize((311,400), method='squish')` and no batch transforms, with a TTA validation error rate of 0.3543. \n- It's significant to note that 9 out of the 10 best performing models did not use batch transformations. In other words, data augmentation during training doesn't seem to improve this model's performance. \n- Six of the top 10 models took less than 2 minutes per epoch. \n- There was a pretty even distribution of item transform methods (3 squish, 4 crop and 3 pad) in the top 10. With the top 3 including each. \n- The top 5 performing `Resize` sizes were `311`, `400` and `(311,400)`. With `400` pixels being included in four of the top five models.","metadata":{}},{"cell_type":"markdown","source":"### Top Overall Models by Final Epoch Validation Error Rate","metadata":{}},{"cell_type":"code","source":"df.sort_values(by='final_epoch_val_error_rate')[:10]","metadata":{"execution":{"iopub.execute_input":"2024-03-10T04:46:53.108763Z","iopub.status.busy":"2024-03-10T04:46:53.108380Z","iopub.status.idle":"2024-03-10T04:46:53.125852Z","shell.execute_reply":"2024-03-10T04:46:53.124620Z","shell.execute_reply.started":"2024-03-10T04:46:53.108734Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The top 2 models are the same as before.\n- The top three `item_tfms` methods are the same as before (squish, pad and crop), although it seems important to note that pad appears only once in the top seven. \n- Seven of the top 10 models took less than 2 minutes per epoch.\n- The dimension of `400` pixels was present in all top 5 models here.\n- Nine models are in both top 10 rankings. \n- The model that enters the top 10 for final epoch validation error rate is the pad method for `item_tfms` with a size of `(400, 311)` and no batch transforms. \n- The model that left the top 10 had a crop `Resize` method sized `(400,311)` and had batch tranforms (and data augmentation) on images sizes `384` pixels square.","metadata":{}},{"cell_type":"code","source":"df.sort_values(by='final_epoch_val_error_rate')[:10].index.difference(df.sort_values(by='tta_val_error_rate')[:10].index)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T04:49:02.468225Z","iopub.status.busy":"2024-03-10T04:49:02.467821Z","iopub.status.idle":"2024-03-10T04:49:02.482444Z","shell.execute_reply":"2024-03-10T04:49:02.481018Z","shell.execute_reply.started":"2024-03-10T04:49:02.468185Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sort_values(by='tta_val_error_rate')[:10].index.difference(df.sort_values(by='final_epoch_val_error_rate')[:10].index)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T04:54:12.754479Z","iopub.status.busy":"2024-03-10T04:54:12.752141Z","iopub.status.idle":"2024-03-10T04:54:12.765454Z","shell.execute_reply":"2024-03-10T04:54:12.764141Z","shell.execute_reply.started":"2024-03-10T04:54:12.754419Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Median Error Rate By Parameter","metadata":{}},{"cell_type":"code","source":"df.groupby('item_method').agg({'tta_val_error_rate': 'median', 'final_epoch_val_error_rate': 'median'})","metadata":{"execution":{"iopub.execute_input":"2024-03-10T05:28:53.506841Z","iopub.status.busy":"2024-03-10T05:28:53.506438Z","iopub.status.idle":"2024-03-10T05:28:53.521137Z","shell.execute_reply":"2024-03-10T05:28:53.519925Z","shell.execute_reply.started":"2024-03-10T05:28:53.506808Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `squish` method performed the best on both types of error rate.","metadata":{}},{"cell_type":"code","source":"df.groupby('item_size').agg({'tta_val_error_rate': 'median', 'final_epoch_val_error_rate': 'median'})","metadata":{"execution":{"iopub.execute_input":"2024-03-10T05:28:06.922014Z","iopub.status.busy":"2024-03-10T05:28:06.921623Z","iopub.status.idle":"2024-03-10T05:28:06.935900Z","shell.execute_reply":"2024-03-10T05:28:06.934978Z","shell.execute_reply.started":"2024-03-10T05:28:06.921984Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"400 x 311 (height x width) images had an advantage in TTA error rate, and 311 x 400 images had an advantage in final epoch error rate.","metadata":{}},{"cell_type":"code","source":"df.batch_size.fillna('None', inplace=True)\ndf.groupby('batch_size').agg({'tta_val_error_rate': 'median', 'final_epoch_val_error_rate': 'median'})","metadata":{"_kg_hide-input":false,"execution":{"iopub.execute_input":"2024-03-10T05:27:51.872018Z","iopub.status.busy":"2024-03-10T05:27:51.870795Z","iopub.status.idle":"2024-03-10T05:27:51.885575Z","shell.execute_reply":"2024-03-10T05:27:51.884487Z","shell.execute_reply.started":"2024-03-10T05:27:51.871977Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As seen in the top 10 rankings, having no batch transforms resulted in the best median error rates.","metadata":{}},{"cell_type":"code","source":"(df.groupby(['item_method', 'item_size'])\n     .agg({'tta_val_error_rate': 'median'})\n     .sort_values(by=['item_method', 'tta_val_error_rate'])\n)","metadata":{"_kg_hide-input":false,"execution":{"iopub.execute_input":"2024-03-10T05:27:33.158573Z","iopub.status.busy":"2024-03-10T05:27:33.158183Z","iopub.status.idle":"2024-03-10T05:27:33.179342Z","shell.execute_reply":"2024-03-10T05:27:33.178175Z","shell.execute_reply.started":"2024-03-10T05:27:33.158542Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When cropping the image, the best performing size (by TTA error rate) was 400 x 311 (height x width). When padding or squishing the image, the best size was 311 x 400 pixels.","metadata":{}},{"cell_type":"code","source":"(df.groupby(['item_method', 'item_size'])\n     .agg({'final_epoch_val_error_rate': 'median'})\n     .sort_values(by=['item_method', 'final_epoch_val_error_rate'])\n)","metadata":{"_kg_hide-input":false,"execution":{"iopub.execute_input":"2024-03-10T06:07:52.216921Z","iopub.status.busy":"2024-03-10T06:07:52.216451Z","iopub.status.idle":"2024-03-10T06:07:52.234574Z","shell.execute_reply":"2024-03-10T06:07:52.233297Z","shell.execute_reply.started":"2024-03-10T06:07:52.216882Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The same method/size combinations have the lowest median final epoch validation error rate.","metadata":{}},{"cell_type":"markdown","source":"## Picking the Top Models Based on Current Results","metadata":{}},{"cell_type":"markdown","source":"Given the current results, I would pick the following five models for submission:","metadata":{}},{"cell_type":"code","source":"df.loc[[8, 36, 28, 12, 0]]","metadata":{"_kg_hide-input":false,"execution":{"iopub.execute_input":"2024-03-10T05:50:21.296408Z","iopub.status.busy":"2024-03-10T05:50:21.295949Z","iopub.status.idle":"2024-03-10T05:50:21.311615Z","shell.execute_reply":"2024-03-10T05:50:21.310517Z","shell.execute_reply.started":"2024-03-10T05:50:21.296373Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"None of these have batch transforms since they decreased performance. The top 3 represent an equal mix of `item_tfms` methods and sizes.  `squish` had the lowest TTA error rate across all `item_tfms` methods so I picked it for the last two. Finally, `(400, 311)` had the lowest median TTA error rate across all sizes so I picked it twice.","metadata":{}},{"cell_type":"markdown","source":"## Running Additional Experiments","metadata":{}},{"cell_type":"markdown","source":"Before I finalize my top five models, I'd like to run a few more experiments. I was curious that the best performing model had an unusual size: 311 pixels in height, 400 pixel width. I wonder if the larger width allowed for more flexibility for those spectrograms that were extremely large (1000+ pixels). I'd like to train a few models that are >400 pixels wide.\n\nI noted that the 2nd and 3rd best TTA error validation rate were for 400 x 400 items (for both pad and crop methods) so I'll try a few models with square images larger than 400 px (using the padding method).\n\nFinally, I'd like to just try a couple of sizes that I haven't thus far: 128 x 128 and 256 x 256. I don't have any reason to believe they will perform well, but I'd like to give small images a shot.","metadata":{}},{"cell_type":"markdown","source":"### Experiment Results","metadata":{}},{"cell_type":"markdown","source":"Here are the results from the additional trainings below:","metadata":{}},{"cell_type":"code","source":"url = 'https://gist.githubusercontent.com/vishalbakshi/d5d4cf1ff73c6daecfd6eb79513f6ada/raw/98f1d48b551b75a61670637361730e23bd4c8a59/hms_hbac_small_model_experiments.csv'\ndf = pd.read_csv(url)\ndf.tail(10)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T17:40:00.004303Z","iopub.status.busy":"2024-03-11T17:40:00.003541Z","iopub.status.idle":"2024-03-11T17:40:00.164382Z","shell.execute_reply":"2024-03-11T17:40:00.163511Z","shell.execute_reply.started":"2024-03-11T17:40:00.004274Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The best performing model out of these last 10 training runs was with a `Resize` size `320` x `512` using the `pad` method. This resulted in a TTA validation error rate of 0.3678, the same as the `400` square `squish`ed images (which I'll replace in my top 5).\n\nHere are my final top 5 models that I will use for submission:","metadata":{}},{"cell_type":"code","source":"df.loc[[8, 36, 28, 12, 50]]","metadata":{"execution":{"iopub.execute_input":"2024-03-11T17:44:14.770740Z","iopub.status.busy":"2024-03-11T17:44:14.770296Z","iopub.status.idle":"2024-03-11T17:44:14.794114Z","shell.execute_reply":"2024-03-11T17:44:14.793104Z","shell.execute_reply.started":"2024-03-11T17:44:14.770706Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm hiding my `train` function so this notebook stays as short as possible. I'll also add the memory/cache clearing lines of code within the function so that I don't have to run them each time.","metadata":{}},{"cell_type":"code","source":"def train(arch, size, item, accum=4, epochs=7):\n    if size is None: \n        batch = None\n    else: \n        batch = aug_transforms(size=size, min_scale=0.75)\n        \n    dls = ImageDataLoaders.from_folder(\n        path, \n        valid_pct=0.2, \n        seed=42,\n        item_tfms=item,\n        batch_tfms=batch,\n        bs=64//accum)\n    \n    cbs = GradientAccumulation(64) if accum else []\n    learn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()\n    learn.fine_tune(epochs, 0.01)\n    print(error_rate(*learn.tta(dl=dls.valid)))\n    gc.collect()\n    torch.cuda.empty_cache()","metadata":{"_kg_hide-input":true,"execution":{"iopub.execute_input":"2024-03-10T06:32:45.593558Z","iopub.status.busy":"2024-03-10T06:32:45.592965Z","iopub.status.idle":"2024-03-10T06:32:45.598852Z","shell.execute_reply":"2024-03-10T06:32:45.597889Z","shell.execute_reply.started":"2024-03-10T06:32:45.593536Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### >400px Wide","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.execute_input":"2024-03-10T20:56:37.557135Z","iopub.status.busy":"2024-03-10T20:56:37.556357Z","iopub.status.idle":"2024-03-10T20:56:37.559889Z","shell.execute_reply":"2024-03-10T20:56:37.559304Z","shell.execute_reply.started":"2024-03-10T20:56:37.557113Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arch = 'convnext_small_in22k'","metadata":{"execution":{"iopub.execute_input":"2024-03-10T20:56:38.654066Z","iopub.status.busy":"2024-03-10T20:56:38.653179Z","iopub.status.idle":"2024-03-10T20:56:38.657004Z","shell.execute_reply":"2024-03-10T20:56:38.656548Z","shell.execute_reply.started":"2024-03-10T20:56:38.654033Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((320, 512), method='squish'))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T06:47:42.419780Z","iopub.status.busy":"2024-03-10T06:47:42.419321Z","iopub.status.idle":"2024-03-10T07:06:01.620556Z","shell.execute_reply":"2024-03-10T07:06:01.619784Z","shell.execute_reply.started":"2024-03-10T06:47:42.419753Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((320, 512)))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T07:06:01.621887Z","iopub.status.busy":"2024-03-10T07:06:01.621704Z","iopub.status.idle":"2024-03-10T07:24:22.252505Z","shell.execute_reply":"2024-03-10T07:24:22.251159Z","shell.execute_reply.started":"2024-03-10T07:06:01.621871Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((320, 512), method=ResizeMethod.Pad, pad_mode=PadMode.Zeros))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T07:24:22.254374Z","iopub.status.busy":"2024-03-10T07:24:22.254082Z","iopub.status.idle":"2024-03-10T07:42:40.931015Z","shell.execute_reply":"2024-03-10T07:42:40.929940Z","shell.execute_reply.started":"2024-03-10T07:24:22.254356Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((384, 768), method='squish'))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T08:30:43.031979Z","iopub.status.busy":"2024-03-10T08:30:43.031661Z","iopub.status.idle":"2024-03-10T09:01:51.388004Z","shell.execute_reply":"2024-03-10T09:01:51.387250Z","shell.execute_reply.started":"2024-03-10T08:30:43.031956Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((384, 768)))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T09:01:51.391331Z","iopub.status.busy":"2024-03-10T09:01:51.391113Z","iopub.status.idle":"2024-03-10T09:32:59.425528Z","shell.execute_reply":"2024-03-10T09:32:59.424954Z","shell.execute_reply.started":"2024-03-10T09:01:51.391314Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize((384, 768), method=ResizeMethod.Pad, pad_mode=PadMode.Zeros))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T09:32:59.427079Z","iopub.status.busy":"2024-03-10T09:32:59.426880Z","iopub.status.idle":"2024-03-10T10:04:07.280982Z","shell.execute_reply":"2024-03-10T10:04:07.280383Z","shell.execute_reply.started":"2024-03-10T09:32:59.427062Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Small Images","metadata":{}},{"cell_type":"code","source":"train(arch, size=None, item=Resize(64))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T10:04:07.282484Z","iopub.status.busy":"2024-03-10T10:04:07.282291Z","iopub.status.idle":"2024-03-10T10:08:43.861056Z","shell.execute_reply":"2024-03-10T10:08:43.860257Z","shell.execute_reply.started":"2024-03-10T10:04:07.282473Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize(128))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T10:08:43.863427Z","iopub.status.busy":"2024-03-10T10:08:43.862778Z","iopub.status.idle":"2024-03-10T10:13:24.341592Z","shell.execute_reply":"2024-03-10T10:13:24.340824Z","shell.execute_reply.started":"2024-03-10T10:08:43.863403Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train(arch, size=None, item=Resize(256))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T10:13:24.343193Z","iopub.status.busy":"2024-03-10T10:13:24.342984Z","iopub.status.idle":"2024-03-10T10:21:52.456409Z","shell.execute_reply":"2024-03-10T10:21:52.455801Z","shell.execute_reply.started":"2024-03-10T10:13:24.343174Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Progressive Resizing\n\nFollowing the approach shown in [chapter 7 of the fastai textbook](https://github.com/fastai/fastbook/blob/master/07_sizing_and_tta.ipynb), I'll train the model on increasing larger sizes and smaller learning rates. However, this does not improve the performance. The model performs much worse (~20% higher error rate).","metadata":{}},{"cell_type":"code","source":"gc.collect()\ntorch.cuda.empty_cache()","metadata":{"execution":{"iopub.execute_input":"2024-03-10T21:49:44.755704Z","iopub.status.busy":"2024-03-10T21:49:44.755197Z","iopub.status.idle":"2024-03-10T21:49:45.209821Z","shell.execute_reply":"2024-03-10T21:49:45.209152Z","shell.execute_reply.started":"2024-03-10T21:49:44.755675Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_dls(size):\n    dls = ImageDataLoaders.from_folder(\n            path, \n            valid_pct=0.2, \n            seed=42,\n            item_tfms=Resize(size=size),\n            batch_tfms=None,\n            bs=16)\n    return dls","metadata":{"execution":{"iopub.execute_input":"2024-03-10T21:49:47.575467Z","iopub.status.busy":"2024-03-10T21:49:47.575163Z","iopub.status.idle":"2024-03-10T21:49:47.579214Z","shell.execute_reply":"2024-03-10T21:49:47.578507Z","shell.execute_reply.started":"2024-03-10T21:49:47.575445Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = get_dls(size=128)\ndls.show_batch(nrows=1, ncols=3)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T21:49:50.619210Z","iopub.status.busy":"2024-03-10T21:49:50.618908Z","iopub.status.idle":"2024-03-10T21:49:51.275855Z","shell.execute_reply":"2024-03-10T21:49:51.275337Z","shell.execute_reply.started":"2024-03-10T21:49:50.619188Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cbs = GradientAccumulation(64)\nlearn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()\nlearn.fine_tune(4, 0.03)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T21:50:12.784832Z","iopub.status.busy":"2024-03-10T21:50:12.784322Z","iopub.status.idle":"2024-03-10T21:52:53.667457Z","shell.execute_reply":"2024-03-10T21:52:53.666732Z","shell.execute_reply.started":"2024-03-10T21:50:12.784810Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.dls = get_dls(size=(311,400)) \nlearn.fine_tune(5, 0.02)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T21:52:53.668862Z","iopub.status.busy":"2024-03-10T21:52:53.668663Z","iopub.status.idle":"2024-03-10T22:02:54.775809Z","shell.execute_reply":"2024-03-10T22:02:54.775087Z","shell.execute_reply.started":"2024-03-10T21:52:53.668842Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.dls = get_dls(size=(320,512)) \nlearn.fine_tune(9, 0.01)","metadata":{"execution":{"iopub.execute_input":"2024-03-10T22:02:54.777194Z","iopub.status.busy":"2024-03-10T22:02:54.777002Z","iopub.status.idle":"2024-03-10T22:24:16.140968Z","shell.execute_reply":"2024-03-10T22:24:16.140218Z","shell.execute_reply.started":"2024-03-10T22:02:54.777173Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(error_rate(*learn.tta(dl=dls.valid)))","metadata":{"execution":{"iopub.execute_input":"2024-03-10T22:24:16.145122Z","iopub.status.busy":"2024-03-10T22:24:16.144420Z","iopub.status.idle":"2024-03-10T22:24:33.320963Z","shell.execute_reply":"2024-03-10T22:24:33.320363Z","shell.execute_reply.started":"2024-03-10T22:24:16.145091Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}