{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hi there! This is the beggining of a series that's basically a copy of [Jeremy's](https://www.kaggle.com/jhoward) [road to the top](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1) series.  \n\nIn that Jeremy makes the argument that:\n>  You might be surprised to discover that the process of doing this was nearly entirely mechanistic and didn't involve any consideration of the actual data or evaluation details at all.\n\nWell, I don't know anything about identifying whales (although I would find it fun to learn a little bit about it), so let's give it a shot!\n\n---\n**Notes:**  \nThis series is being written by [Lucas](https://twitter.com/lucasgvazquez) and [Francesco](https://twitter.com/Fra_Pochetti), we are going to use \"I\" instead of \"we\" just because it's easier, but keep in mind that we say \"I\" we actually mean \"we\".  \n\nThe notebooks are being published in real time before the series is completed, so at this point we still don't know what final results await us, maybe this will fail completely!\n\nBecause the competition is finished, we won't be able to get a placement in the leaderboard :/","metadata":{"papermill":{"duration":0.015941,"end_time":"2022-07-14T11:58:51.592483","exception":false,"start_time":"2022-07-14T11:58:51.576542","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## First steps","metadata":{"papermill":{"duration":0.014348,"end_time":"2022-07-14T11:58:51.622035","exception":false,"start_time":"2022-07-14T11:58:51.607687","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"We'll start by installing and importing required libs (mainly fastai stuff) and getting the data. We'll use [fastkaggle](https://github.com/fastai/fastkaggle), a handy library that will facilitate our life interacting with kaggle, both locally and in their kernels.","metadata":{"papermill":{"duration":0.013394,"end_time":"2022-07-14T11:58:51.649812","exception":false,"start_time":"2022-07-14T11:58:51.636418","status":"completed"},"tags":[]}},{"cell_type":"code","source":"try: import fastkaggle\nexcept ModuleNotFoundError:\n    !pip install -Uq fastkaggle\n\nfrom fastkaggle import *","metadata":{"_kg_hide-output":true,"execution":{"iopub.execute_input":"2022-07-14T11:58:51.680915Z","iopub.status.busy":"2022-07-14T11:58:51.679477Z","iopub.status.idle":"2022-07-14T11:59:06.239481Z","shell.execute_reply":"2022-07-14T11:59:06.237997Z"},"papermill":{"duration":14.579749,"end_time":"2022-07-14T11:59:06.243117","exception":false,"start_time":"2022-07-14T11:58:51.663368","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp = 'noaa-right-whale-recognition'\ndata_dir = setup_comp(comp, install='fastai \"timm>=0.6.2.dev0\"')","metadata":{"execution":{"iopub.execute_input":"2022-07-14T11:59:06.275958Z","iopub.status.busy":"2022-07-14T11:59:06.274366Z","iopub.status.idle":"2022-07-14T11:59:19.631976Z","shell.execute_reply":"2022-07-14T11:59:19.630701Z"},"papermill":{"duration":13.376364,"end_time":"2022-07-14T11:59:19.635056","exception":false,"start_time":"2022-07-14T11:59:06.258692","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.vision.all import *","metadata":{"execution":{"iopub.execute_input":"2022-07-14T11:59:19.666046Z","iopub.status.busy":"2022-07-14T11:59:19.664936Z","iopub.status.idle":"2022-07-14T11:59:24.589749Z","shell.execute_reply":"2022-07-14T11:59:24.588390Z"},"papermill":{"duration":4.943016,"end_time":"2022-07-14T11:59:24.592881","exception":false,"start_time":"2022-07-14T11:59:19.649865","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We're going to be setting a seed to make this notebook easier to reproduce, but don't go crazy about it and try to make all your experiments reproducible, you actually want a bit of variation when experimenting - I just set the seed on the very last moment before I submit to kaggle, when I'm already done with all my experiments.","metadata":{"papermill":{"duration":0.01369,"end_time":"2022-07-14T11:59:24.622038","exception":false,"start_time":"2022-07-14T11:59:24.608348","status":"completed"},"tags":[]}},{"cell_type":"code","source":"set_seed(42)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T11:59:24.653352Z","iopub.status.busy":"2022-07-14T11:59:24.651744Z","iopub.status.idle":"2022-07-14T11:59:24.661855Z","shell.execute_reply":"2022-07-14T11:59:24.660719Z"},"papermill":{"duration":0.028056,"end_time":"2022-07-14T11:59:24.664192","exception":false,"start_time":"2022-07-14T11:59:24.636136","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploring the data","metadata":{"papermill":{"duration":0.01348,"end_time":"2022-07-14T11:59:24.691806","exception":false,"start_time":"2022-07-14T11:59:24.678326","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Let's take an initial quick look at the data, we have all our images (both train and test) in a single zip file (that we will have to extract), and a `csv` file with the labels. ","metadata":{"papermill":{"duration":0.01422,"end_time":"2022-07-14T11:59:24.720059","exception":false,"start_time":"2022-07-14T11:59:24.705839","status":"completed"},"tags":[]}},{"cell_type":"code","source":"imgs_dir = untar_dir(data_dir/'imgs.zip', Path('imgs'))","metadata":{"execution":{"iopub.execute_input":"2022-07-14T11:59:24.749816Z","iopub.status.busy":"2022-07-14T11:59:24.749502Z","iopub.status.idle":"2022-07-14T12:03:19.140396Z","shell.execute_reply":"2022-07-14T12:03:19.135497Z"},"papermill":{"duration":234.421286,"end_time":"2022-07-14T12:03:19.155680","exception":false,"start_time":"2022-07-14T11:59:24.734394","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_files = get_image_files(imgs_dir)\ntargs_df = pd.read_csv(data_dir/'train.csv')","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:19.192046Z","iopub.status.busy":"2022-07-14T12:03:19.191659Z","iopub.status.idle":"2022-07-14T12:03:19.376342Z","shell.execute_reply":"2022-07-14T12:03:19.375002Z"},"papermill":{"duration":0.206591,"end_time":"2022-07-14T12:03:19.378943","exception":false,"start_time":"2022-07-14T12:03:19.172352","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since the task was a bit confusing at first, let me reiterate it and clarify what it is: We're trying to identify which whale appears in each image, categorized by `whaleID` in our dataframe. We are **not** identifying whales by species, but instead whether it's Kayla, Takara, etc.","metadata":{"papermill":{"duration":0.015701,"end_time":"2022-07-14T12:03:20.365711","exception":false,"start_time":"2022-07-14T12:03:20.350010","status":"completed"},"tags":[]}},{"cell_type":"code","source":"targs_df.head()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:20.406165Z","iopub.status.busy":"2022-07-14T12:03:20.405669Z","iopub.status.idle":"2022-07-14T12:03:20.444667Z","shell.execute_reply":"2022-07-14T12:03:20.443454Z"},"papermill":{"duration":0.067039,"end_time":"2022-07-14T12:03:20.447307","exception":false,"start_time":"2022-07-14T12:03:20.380268","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I also always like to take a quick look at some images.","metadata":{"papermill":{"duration":0.014663,"end_time":"2022-07-14T12:03:20.476688","exception":false,"start_time":"2022-07-14T12:03:20.462025","status":"completed"},"tags":[]}},{"cell_type":"code","source":"image_file = random.choice(image_files)\nPILImage.create(image_file)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:20.509046Z","iopub.status.busy":"2022-07-14T12:03:20.506903Z","iopub.status.idle":"2022-07-14T12:03:22.929783Z","shell.execute_reply":"2022-07-14T12:03:22.906628Z"},"papermill":{"duration":2.556057,"end_time":"2022-07-14T12:03:23.047492","exception":false,"start_time":"2022-07-14T12:03:20.491435","status":"completed"},"scrolled":true,"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, while experimenting I found that one of the images specified in the dataframe is not present in the `imgs` folder, so let me already go ahead and remove it.","metadata":{"papermill":{"duration":0.200456,"end_time":"2022-07-14T12:03:23.443705","exception":false,"start_time":"2022-07-14T12:03:23.243249","status":"completed"},"tags":[]}},{"cell_type":"code","source":"targs_df = targs_df[targs_df['Image']!='w_7489.jpg']","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:23.704709Z","iopub.status.busy":"2022-07-14T12:03:23.704144Z","iopub.status.idle":"2022-07-14T12:03:23.724179Z","shell.execute_reply":"2022-07-14T12:03:23.722961Z"},"papermill":{"duration":0.148262,"end_time":"2022-07-14T12:03:23.726633","exception":false,"start_time":"2022-07-14T12:03:23.578371","status":"completed"},"scrolled":true,"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check how many samples and classes we have and how balanced our data is.","metadata":{"papermill":{"duration":0.083793,"end_time":"2022-07-14T12:03:23.890980","exception":false,"start_time":"2022-07-14T12:03:23.807187","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def count_data(df):\n    print(f\"There are {len(df['whaleID'].unique())} classes and {len(targs_df)} samples.\")\n    \n    whale_count = Counter(targs_df['whaleID'])\n    return plt.hist(whale_count.values())","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:24.062143Z","iopub.status.busy":"2022-07-14T12:03:24.061636Z","iopub.status.idle":"2022-07-14T12:03:24.070455Z","shell.execute_reply":"2022-07-14T12:03:24.069023Z"},"papermill":{"duration":0.098178,"end_time":"2022-07-14T12:03:24.073034","exception":false,"start_time":"2022-07-14T12:03:23.974856","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count_data(targs_df)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:24.243081Z","iopub.status.busy":"2022-07-14T12:03:24.242496Z","iopub.status.idle":"2022-07-14T12:03:24.583008Z","shell.execute_reply":"2022-07-14T12:03:24.581909Z"},"papermill":{"duration":0.431196,"end_time":"2022-07-14T12:03:24.585915","exception":false,"start_time":"2022-07-14T12:03:24.154719","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can observe that the vast majority of data consists of whales that showed up less than 20 times, with a good amount showing less than 6. That makes our dataset heavily imbalanced, we can simplify and eliminate this problem for now by creating a subset that only contains whales that showed up more than 20 times.","metadata":{"papermill":{"duration":0.085807,"end_time":"2022-07-14T12:03:27.112793","exception":false,"start_time":"2022-07-14T12:03:27.026986","status":"completed"},"tags":[]}},{"cell_type":"code","source":"count_above = targs_df['whaleID'].value_counts() > 20\ncount_above = list(count_above.index[count_above])\n\nmask = targs_df['whaleID'].apply(lambda o: o in count_above)\ntargs_df = targs_df[mask]","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:27.286651Z","iopub.status.busy":"2022-07-14T12:03:27.286097Z","iopub.status.idle":"2022-07-14T12:03:27.306540Z","shell.execute_reply":"2022-07-14T12:03:27.305333Z"},"papermill":{"duration":0.111904,"end_time":"2022-07-14T12:03:27.309209","exception":false,"start_time":"2022-07-14T12:03:27.197305","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count_data(targs_df)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:27.480274Z","iopub.status.busy":"2022-07-14T12:03:27.477716Z","iopub.status.idle":"2022-07-14T12:03:27.702292Z","shell.execute_reply":"2022-07-14T12:03:27.701102Z"},"papermill":{"duration":0.313871,"end_time":"2022-07-14T12:03:27.705088","exception":false,"start_time":"2022-07-14T12:03:27.391217","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A bit better but still unbalanced. With 745 samples our models should train extremely quick and we'll be able to iterate fast!","metadata":{"papermill":{"duration":0.084375,"end_time":"2022-07-14T12:03:27.872488","exception":false,"start_time":"2022-07-14T12:03:27.788113","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"There's a problem though, the resolution of the images is very high and that will bottleneck our training pipeline! (You will notice the GPU will stay at 0% usage for long periods of time).","metadata":{"papermill":{"duration":0.08445,"end_time":"2022-07-14T12:03:28.042880","exception":false,"start_time":"2022-07-14T12:03:27.958430","status":"completed"},"tags":[]}},{"cell_type":"code","source":"sizes = parallel(image_size, image_files)\nwidths, heights = zip(*sizes)\nplt.hist(widths, label='width')\nplt.hist(heights, label='heights')\nplt.legend()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:28.218563Z","iopub.status.busy":"2022-07-14T12:03:28.217971Z","iopub.status.idle":"2022-07-14T12:03:56.554557Z","shell.execute_reply":"2022-07-14T12:03:56.553247Z"},"papermill":{"duration":28.428907,"end_time":"2022-07-14T12:03:56.557139","exception":false,"start_time":"2022-07-14T12:03:28.128232","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can resize all images to a much smaller size before hand, avoiding overwheling the CPU (which is specially limited in kaggle). Of course `fastai` has a handy function for that; it'll resize all images to a new folder, keeping the original file structure the same.","metadata":{"papermill":{"duration":0.089618,"end_time":"2022-07-14T12:03:56.736004","exception":false,"start_time":"2022-07-14T12:03:56.646386","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if not Path('imgs_288max').exists():\n    resize_images(imgs_dir, dest='imgs_288max', max_size=288, recurse=True, progress=progress_bar)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:03:56.913373Z","iopub.status.busy":"2022-07-14T12:03:56.912754Z","iopub.status.idle":"2022-07-14T12:44:09.964212Z","shell.execute_reply":"2022-07-14T12:44:09.962864Z"},"papermill":{"duration":2413.145486,"end_time":"2022-07-14T12:44:09.967087","exception":false,"start_time":"2022-07-14T12:03:56.821601","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data pipeline","metadata":{"papermill":{"duration":0.131953,"end_time":"2022-07-14T12:44:10.187926","exception":false,"start_time":"2022-07-14T12:44:10.055973","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"We are finally ready to build the data pipeline! Most fastai tutorials use the `DataBlock` API but I really like to use the `Datasets` API. It's a tiny bit more verbose, but in my opinion it's so much easier to understand and use (specially when you start doing out of the box stuff).\n\nThe first thing we need is something to iterate over to get our image-caption pairs. The most natural approach is to use `image_files`, but in this case, both training and testing files are in the same folder so we would have to write some custom logic to separate them. Instead an easier approach is just to iterate over the rows of our dataframe, which only contain training items.","metadata":{"papermill":{"duration":0.088042,"end_time":"2022-07-14T12:44:10.361632","exception":false,"start_time":"2022-07-14T12:44:10.273590","status":"completed"},"tags":[]}},{"cell_type":"code","source":"items = targs_df['Image']\nitem = items[0]\nitem","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:10.535235Z","iopub.status.busy":"2022-07-14T12:44:10.534534Z","iopub.status.idle":"2022-07-14T12:44:10.546606Z","shell.execute_reply":"2022-07-14T12:44:10.545305Z"},"papermill":{"duration":0.102154,"end_time":"2022-07-14T12:44:10.549103","exception":false,"start_time":"2022-07-14T12:44:10.446949","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We need a way of converting each of those items into an image-caption pair.","metadata":{"papermill":{"duration":0.084832,"end_time":"2022-07-14T12:44:10.728666","exception":false,"start_time":"2022-07-14T12:44:10.643834","status":"completed"},"tags":[]}},{"cell_type":"code","source":"image_file = f'imgs_288max/{item}'\nimage = PILImage.create(image_file)\nimage","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:10.898749Z","iopub.status.busy":"2022-07-14T12:44:10.898199Z","iopub.status.idle":"2022-07-14T12:44:10.914718Z","shell.execute_reply":"2022-07-14T12:44:10.913389Z"},"papermill":{"duration":0.105431,"end_time":"2022-07-14T12:44:10.917624","exception":false,"start_time":"2022-07-14T12:44:10.812193","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image2whaleID = {o.Image: o.whaleID for o in targs_df.itertuples()}\nlabel = image2whaleID[item]\nlabel","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:11.093525Z","iopub.status.busy":"2022-07-14T12:44:11.093024Z","iopub.status.idle":"2022-07-14T12:44:11.102882Z","shell.execute_reply":"2022-07-14T12:44:11.101571Z"},"papermill":{"duration":0.103487,"end_time":"2022-07-14T12:44:11.105384","exception":false,"start_time":"2022-07-14T12:44:11.001897","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we just have to put those into their functions and create a `Dataset`. We will also randomly split our data into train and validation.","metadata":{"papermill":{"duration":0.222309,"end_time":"2022-07-14T12:44:11.411662","exception":false,"start_time":"2022-07-14T12:44:11.189353","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def get_image_file(item):\n    return f'imgs_288max/{item}'\n\ndef get_label(item):\n    return image2whaleID[item]","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:11.590244Z","iopub.status.busy":"2022-07-14T12:44:11.589211Z","iopub.status.idle":"2022-07-14T12:44:11.596570Z","shell.execute_reply":"2022-07-14T12:44:11.595420Z"},"papermill":{"duration":0.100759,"end_time":"2022-07-14T12:44:11.599036","exception":false,"start_time":"2022-07-14T12:44:11.498277","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We also have to provide `Categorize` to the label, so it converts the label from a string to a tensor.","metadata":{"papermill":{"duration":0.083304,"end_time":"2022-07-14T12:44:11.765946","exception":false,"start_time":"2022-07-14T12:44:11.682642","status":"completed"},"tags":[]}},{"cell_type":"code","source":"x_pipe = [get_image_file, PILImage.create]\ny_pipe = [get_label, Categorize()]\n\nsplits = RandomSplitter(seed=42)(items)\ndss = Datasets(items, [x_pipe, y_pipe], splits=splits)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:11.957492Z","iopub.status.busy":"2022-07-14T12:44:11.956908Z","iopub.status.idle":"2022-07-14T12:44:12.269638Z","shell.execute_reply":"2022-07-14T12:44:12.268358Z"},"papermill":{"duration":0.422735,"end_time":"2022-07-14T12:44:12.272275","exception":false,"start_time":"2022-07-14T12:44:11.849540","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dss.show(dss[0])","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:12.478327Z","iopub.status.busy":"2022-07-14T12:44:12.477630Z","iopub.status.idle":"2022-07-14T12:44:12.637687Z","shell.execute_reply":"2022-07-14T12:44:12.636296Z"},"papermill":{"duration":0.252738,"end_time":"2022-07-14T12:44:12.641229","exception":false,"start_time":"2022-07-14T12:44:12.388491","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can create our dataloader. Before the dataloader is able to create batches, we need to resize the images to one common size. We can do that by passing transforms to `after_item` which executes on an item by item basis, just before a batch is formed.  \n\nAfter we have our batch, we will apply augmentation transforms. `fastai` provides a bundle of commonly used transforms accessible with `aug_transforms`. The most important parameters to be defined here are:  \n1. `min_scale`: randomly crop the images, maintaining atleast `min_scale` percentage of the original image.  \n2. `size`: resize the crop to the defined `size`.  ","metadata":{"papermill":{"duration":0.127765,"end_time":"2022-07-14T12:44:12.876101","exception":false,"start_time":"2022-07-14T12:44:12.748336","status":"completed"},"tags":[]}},{"cell_type":"code","source":"after_item = [ToTensor(), Resize((192, 288))]\nafter_batch = [IntToFloatTensor(), *aug_transforms(size=(128, 192), min_scale=0.75, flip_vert=True, max_rotate=180)]\ndls = dss.dataloaders(32, after_item=after_item, after_batch=after_batch)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:13.134880Z","iopub.status.busy":"2022-07-14T12:44:13.134280Z","iopub.status.idle":"2022-07-14T12:44:21.927921Z","shell.execute_reply":"2022-07-14T12:44:21.926555Z"},"papermill":{"duration":8.928062,"end_time":"2022-07-14T12:44:21.930946","exception":false,"start_time":"2022-07-14T12:44:13.002884","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch(max_n=3)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:22.110991Z","iopub.status.busy":"2022-07-14T12:44:22.110531Z","iopub.status.idle":"2022-07-14T12:44:22.580250Z","shell.execute_reply":"2022-07-14T12:44:22.578910Z"},"papermill":{"duration":0.569382,"end_time":"2022-07-14T12:44:22.587773","exception":false,"start_time":"2022-07-14T12:44:22.018391","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks about right! Let's define our metric and train our first model `resnet26d`, as it's the fastest one we can find in [the best vision models for fine-tuning](https://www.kaggle.com/code/jhoward/the-best-vision-models-for-fine-tuning).","metadata":{"papermill":{"duration":0.088274,"end_time":"2022-07-14T12:44:22.770913","exception":false,"start_time":"2022-07-14T12:44:22.682639","status":"completed"},"tags":[]}},{"cell_type":"code","source":"metrics = [error_rate]\nlearn = vision_learner(dls, 'resnet26d', metrics=metrics).to_fp16()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:22.982043Z","iopub.status.busy":"2022-07-14T12:44:22.981551Z","iopub.status.idle":"2022-07-14T12:44:27.946493Z","shell.execute_reply":"2022-07-14T12:44:27.944939Z"},"papermill":{"duration":5.092194,"end_time":"2022-07-14T12:44:27.949676","exception":false,"start_time":"2022-07-14T12:44:22.857482","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:28.126587Z","iopub.status.busy":"2022-07-14T12:44:28.126015Z","iopub.status.idle":"2022-07-14T12:44:52.731772Z","shell.execute_reply":"2022-07-14T12:44:52.730350Z"},"papermill":{"duration":24.697547,"end_time":"2022-07-14T12:44:52.734372","exception":false,"start_time":"2022-07-14T12:44:28.036825","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fine_tune(5, 0.002)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:44:52.917952Z","iopub.status.busy":"2022-07-14T12:44:52.917138Z","iopub.status.idle":"2022-07-14T12:45:19.895816Z","shell.execute_reply":"2022-07-14T12:45:19.894427Z"},"papermill":{"duration":27.073979,"end_time":"2022-07-14T12:45:19.898438","exception":false,"start_time":"2022-07-14T12:44:52.824459","status":"completed"},"scrolled":true,"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, that certainly seems horrible! But we don't care right now, we just want to make the first submission as fast as possible!","metadata":{"papermill":{"duration":0.089434,"end_time":"2022-07-14T12:45:20.080036","exception":false,"start_time":"2022-07-14T12:45:19.990602","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Inference","metadata":{"papermill":{"duration":0.089476,"end_time":"2022-07-14T12:45:20.256780","exception":false,"start_time":"2022-07-14T12:45:20.167304","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Remember that all our images, train and test are into the same folder? We need a way to get what images are from the test set first of all! For that we can use the sample submission file that comes with the data.","metadata":{"papermill":{"duration":0.08927,"end_time":"2022-07-14T12:45:20.438161","exception":false,"start_time":"2022-07-14T12:45:20.348891","status":"completed"},"tags":[]}},{"cell_type":"code","source":"test_df = pd.read_csv(data_dir / 'sample_submission.csv')\ntest_df.head()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:20.621212Z","iopub.status.busy":"2022-07-14T12:45:20.619849Z","iopub.status.idle":"2022-07-14T12:45:21.009797Z","shell.execute_reply":"2022-07-14T12:45:21.008530Z"},"papermill":{"duration":0.487158,"end_time":"2022-07-14T12:45:21.012421","exception":false,"start_time":"2022-07-14T12:45:20.525263","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can then easily create a validation dataloader with fastai and predict!","metadata":{"papermill":{"duration":0.091531,"end_time":"2022-07-14T12:45:21.193695","exception":false,"start_time":"2022-07-14T12:45:21.102164","status":"completed"},"tags":[]}},{"cell_type":"code","source":"test_items = test_df['Image']\ntest_dl = learn.dls.test_dl(test_items)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:21.381047Z","iopub.status.busy":"2022-07-14T12:45:21.380525Z","iopub.status.idle":"2022-07-14T12:45:21.389307Z","shell.execute_reply":"2022-07-14T12:45:21.387824Z"},"papermill":{"duration":0.105219,"end_time":"2022-07-14T12:45:21.392132","exception":false,"start_time":"2022-07-14T12:45:21.286913","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds, targs = learn.get_preds(dl=test_dl)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:21.576697Z","iopub.status.busy":"2022-07-14T12:45:21.576151Z","iopub.status.idle":"2022-07-14T12:45:46.826014Z","shell.execute_reply":"2022-07-14T12:45:46.824630Z"},"papermill":{"duration":25.346738,"end_time":"2022-07-14T12:45:46.828566","exception":false,"start_time":"2022-07-14T12:45:21.481828","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.DataFrame(preds.numpy(), columns=learn.dls.vocab)\ndf[\"Image\"] = test_items\ndf.head()","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:47.018569Z","iopub.status.busy":"2022-07-14T12:45:47.017295Z","iopub.status.idle":"2022-07-14T12:45:47.055722Z","shell.execute_reply":"2022-07-14T12:45:47.054362Z"},"papermill":{"duration":0.137378,"end_time":"2022-07-14T12:45:47.058315","exception":false,"start_time":"2022-07-14T12:45:46.920937","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, a problem left to fix is: we removed a bunch of classes from our dataset, which we now have to add back in. Let's give all of the removed classes a probability of 0.\n\nWe're going to use the columns of the sample submission file to know all the whaleIDs, and subtract that from our learner vocab to get all classes we excluded.","metadata":{"papermill":{"duration":0.090005,"end_time":"2022-07-14T12:45:47.235079","exception":false,"start_time":"2022-07-14T12:45:47.145074","status":"completed"},"tags":[]}},{"cell_type":"code","source":"full_vocab = list(test_df.columns[1:])","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:47.415384Z","iopub.status.busy":"2022-07-14T12:45:47.414876Z","iopub.status.idle":"2022-07-14T12:45:47.421229Z","shell.execute_reply":"2022-07-14T12:45:47.419707Z"},"papermill":{"duration":0.099304,"end_time":"2022-07-14T12:45:47.423845","exception":false,"start_time":"2022-07-14T12:45:47.324541","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"diff_vocab = set(full_vocab) - set(learn.dls.vocab)\nfor cat in diff_vocab:\n    df[cat] = 0.0","metadata":{"_kg_hide-output":true,"execution":{"iopub.execute_input":"2022-07-14T12:45:47.605443Z","iopub.status.busy":"2022-07-14T12:45:47.604856Z","iopub.status.idle":"2022-07-14T12:45:47.805425Z","shell.execute_reply":"2022-07-14T12:45:47.804019Z"},"papermill":{"duration":0.297332,"end_time":"2022-07-14T12:45:47.808295","exception":false,"start_time":"2022-07-14T12:45:47.510963","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def submit(df):\n    preds_path = 'first-submission.csv'\n    df.to_csv(preds_path, index=False)\n    if not iskaggle:\n        from kaggle import api\n        api.competition_submit_cli(preds_path, 'resnet26d >20whales', comp)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:47.997740Z","iopub.status.busy":"2022-07-14T12:45:47.996543Z","iopub.status.idle":"2022-07-14T12:45:48.004626Z","shell.execute_reply":"2022-07-14T12:45:48.003439Z"},"papermill":{"duration":0.105688,"end_time":"2022-07-14T12:45:48.007133","exception":false,"start_time":"2022-07-14T12:45:47.901445","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit(df)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:48.190548Z","iopub.status.busy":"2022-07-14T12:45:48.190091Z","iopub.status.idle":"2022-07-14T12:45:50.716857Z","shell.execute_reply":"2022-07-14T12:45:50.715434Z"},"papermill":{"duration":2.624836,"end_time":"2022-07-14T12:45:50.719795","exception":false,"start_time":"2022-07-14T12:45:48.094959","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Whelp, that got us a score of `30.29923`, putting us around at rank 248 out of 364 participants. That's, well, horrible.... Let's find out how horrible this is by trying a uniform guess of 1 for every whale.","metadata":{"papermill":{"duration":0.115727,"end_time":"2022-07-14T12:45:50.924391","exception":false,"start_time":"2022-07-14T12:45:50.808664","status":"completed"},"tags":[]}},{"cell_type":"code","source":"diff_vocab = set(full_vocab) - set(learn.dls.vocab)\nfor cat in full_vocab:\n    df[cat] = 1.0","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:51.108933Z","iopub.status.busy":"2022-07-14T12:45:51.107951Z","iopub.status.idle":"2022-07-14T12:45:51.166455Z","shell.execute_reply":"2022-07-14T12:45:51.165008Z"},"papermill":{"duration":0.155183,"end_time":"2022-07-14T12:45:51.168916","exception":false,"start_time":"2022-07-14T12:45:51.013733","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit(df)","metadata":{"execution":{"iopub.execute_input":"2022-07-14T12:45:51.385695Z","iopub.status.busy":"2022-07-14T12:45:51.385129Z","iopub.status.idle":"2022-07-14T12:45:54.153463Z","shell.execute_reply":"2022-07-14T12:45:54.152041Z"},"papermill":{"duration":2.895623,"end_time":"2022-07-14T12:45:54.156657","exception":false,"start_time":"2022-07-14T12:45:51.261034","status":"completed"},"scrolled":true,"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That improves the score to 6.10255!\n\nAll this work we had, just to be heavily outclassed by a uniform (random) guess!!! (╯°□°）╯︵ ┻━┻","metadata":{"papermill":{"duration":0.089815,"end_time":"2022-07-14T12:45:54.338420","exception":false,"start_time":"2022-07-14T12:45:54.248605","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"We've excluded 418 of classes from our dataset, so very likely that's a reason of this extremely poor performance. Next time we can try training a model with the complete dataset, and see if that is better than random guessing.","metadata":{"papermill":{"duration":0.091393,"end_time":"2022-07-14T12:45:54.518121","exception":false,"start_time":"2022-07-14T12:45:54.426728","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## Lessons learned","metadata":{"papermill":{"duration":0.089216,"end_time":"2022-07-14T12:45:54.707663","exception":false,"start_time":"2022-07-14T12:45:54.618447","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Even though I said in the beggining \"let's just take a quick look at the dataset and get to the modelling step asap\", I think I still spent too much time into over-engineering solutions to simplify the problem, only to get burnt at the end.\n\nNext time, just really try harder to get to training models, doing absolutely nothing more than the necessary to get there.","metadata":{"papermill":{"duration":0.088757,"end_time":"2022-07-14T12:45:54.886669","exception":false,"start_time":"2022-07-14T12:45:54.797912","status":"completed"},"tags":[]}}]}