{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"[Our previous attempt](https://www.kaggle.com/code/lucasgvazquez/road-to-the-top-part-1-a-dive-to-the-bottom) was a complete failure. We spent too much time trying to simplify the task but the result was much worse than a uniform predictor.\n\nLesson learned: Setup a basic model and aim for a first submission ASAP. It might sound a bit against \"best practices\", but when starting a new task, you should not pay much attention to the data. Deep exploratory data analysis (EDA) can give valuable insights, but in the beginning we just need to know the bare minimum to get a model training.\n\nThis is the notebook that should have actually been the first in the series: Let's train a tiny model, with very small images, using the entire dataset.","metadata":{}},{"cell_type":"markdown","source":"If you find this work useful, please remember to click the upvote button at the top. That will help others finding this work, and it will also help us ([Lucas](https://twitter.com/lucasgvazquez) and [Francesco](https://twitter.com/Fra_Pochetti)) to get some recognition and stay motivated to keep producing these :)","metadata":{}},{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"markdown","source":"Downloading libraries and imports same as before. We'll be hiding those cells.","metadata":{}},{"cell_type":"code","source":"try: import fastkaggle\nexcept ModuleNotFoundError:\n    !pip install -Uq fastkaggle\n\nfrom fastkaggle import *","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-13T08:29:19.329112Z","iopub.execute_input":"2022-08-13T08:29:19.329899Z","iopub.status.idle":"2022-08-13T08:29:19.393964Z","shell.execute_reply.started":"2022-08-13T08:29:19.329807Z","shell.execute_reply":"2022-08-13T08:29:19.392867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try: import timm\nexcept ModuleNotFoundError:\n    !pip install \"timm>=0.6.2.dev0\"","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-13T08:29:48.285599Z","iopub.execute_input":"2022-08-13T08:29:48.286269Z","iopub.status.idle":"2022-08-13T08:29:51.410702Z","shell.execute_reply.started":"2022-08-13T08:29:48.286222Z","shell.execute_reply":"2022-08-13T08:29:51.409510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.vision.all import *","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-13T08:30:00.565440Z","iopub.execute_input":"2022-08-13T08:30:00.566486Z","iopub.status.idle":"2022-08-13T08:30:01.216900Z","shell.execute_reply.started":"2022-08-13T08:30:00.566434Z","shell.execute_reply":"2022-08-13T08:30:01.215808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp = 'noaa-right-whale-recognition'\ndata_dir = setup_comp(comp, install='fastai \"timm>=0.6.2.dev0\"')\nimgs_dir = untar_dir(data_dir/'imgs.zip', Path('imgs'))","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-13T08:30:06.369659Z","iopub.execute_input":"2022-08-13T08:30:06.370202Z","iopub.status.idle":"2022-08-13T08:30:20.176118Z","shell.execute_reply.started":"2022-08-13T08:30:06.370154Z","shell.execute_reply":"2022-08-13T08:30:20.174932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once again, let's keep the goal in mind: We want to do the **bare-minimum** to have an initial model training. It's surprisingly hard to define what the bare minimum is. Ideally, we just want to ingest the dataset as it's, randomly split the data, define basic transforms (resize + augmentations), and train the model.\n\nUnfortunately, most tasks will require some extra steps to get there. For sake of clarity, we'll be tagging those as **Extra step**. In the comments, you can give your opinion if these extra steps were really required or if we could have done something differently.","metadata":{}},{"cell_type":"markdown","source":"**Extra step:** What is the average size of images in the dataset? Having images that are too big will bottleneck the CPU (when resizing) and training will be very slow. Let's check the average size of images and resize them if necessary.","metadata":{}},{"cell_type":"code","source":"def plot_sizes(image_sizes, max_n=3000):\n    sizes = parallel(image_size, image_files[:max_n], progress=progress_bar)\n    \n    widths, heights = zip(*sizes)\n    min_x = min(widths + heights)\n    max_x = max(widths + heights)\n    plt.hist(widths, label='width', range=(min_x, max_x))\n    plt.hist(heights, label='height', range=(min_x, max_x))\n    plt.legend();","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:31:21.999110Z","iopub.execute_input":"2022-08-13T08:31:21.999611Z","iopub.status.idle":"2022-08-13T08:31:22.012281Z","shell.execute_reply.started":"2022-08-13T08:31:21.999560Z","shell.execute_reply":"2022-08-13T08:31:22.011020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_files = get_image_files(imgs_dir)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:31:23.607150Z","iopub.execute_input":"2022-08-13T08:31:23.607638Z","iopub.status.idle":"2022-08-13T08:31:23.736519Z","shell.execute_reply.started":"2022-08-13T08:31:23.607596Z","shell.execute_reply":"2022-08-13T08:31:23.735483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_sizes(image_files)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:31:24.869193Z","iopub.execute_input":"2022-08-13T08:31:24.872013Z","iopub.status.idle":"2022-08-13T08:31:33.447318Z","shell.execute_reply.started":"2022-08-13T08:31:24.871968Z","shell.execute_reply":"2022-08-13T08:31:33.446120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the plot above we can see that majority of images are about `3000x2000`. For our first model we would like to try something around `224`. Opening these big images and downsizing by this much will bottleneck the training pipeline (you will notice the GPU will be at 0% usage for long periods).\n\nLet's resize all images to 480 and save them to disk, fastai already provides a handy function for doing this.","metadata":{}},{"cell_type":"code","source":"if not Path('imgs_480max').exists():\n    resize_images(imgs_dir, dest='imgs_480max', max_size=480, progress=progress_bar)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_files = get_image_files('imgs_480max')\nplot_sizes(image_files)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:31:53.078940Z","iopub.execute_input":"2022-08-13T08:31:53.079575Z","iopub.status.idle":"2022-08-13T08:32:01.209748Z","shell.execute_reply.started":"2022-08-13T08:31:53.079528Z","shell.execute_reply":"2022-08-13T08:32:01.208369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Pipeline","metadata":{}},{"cell_type":"markdown","source":"Let's not get into the details here as it was already explained in the [previous](https://www.kaggle.com/code/lucasgvazquez/road-to-the-top-part-1-a-dive-to-the-bottom) kernel.","metadata":{}},{"cell_type":"code","source":"targs_df = pd.read_csv(data_dir/'train.csv')\ntargs_df = targs_df[targs_df['Image']!='w_7489.jpg']\nitems = targs_df['Image']","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:04.625639Z","iopub.execute_input":"2022-08-13T08:32:04.626308Z","iopub.status.idle":"2022-08-13T08:32:04.648010Z","shell.execute_reply.started":"2022-08-13T08:32:04.626260Z","shell.execute_reply":"2022-08-13T08:32:04.647064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_image_path(name):\n    return Path('imgs_480max')/name","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:06.417166Z","iopub.execute_input":"2022-08-13T08:32:06.418025Z","iopub.status.idle":"2022-08-13T08:32:06.423846Z","shell.execute_reply.started":"2022-08-13T08:32:06.417981Z","shell.execute_reply":"2022-08-13T08:32:06.422755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_label(name):\n    return targs_df[targs_df['Image']==name]['whaleID'].item()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:09.001563Z","iopub.execute_input":"2022-08-13T08:32:09.009181Z","iopub.status.idle":"2022-08-13T08:32:09.018631Z","shell.execute_reply.started":"2022-08-13T08:32:09.002034Z","shell.execute_reply":"2022-08-13T08:32:09.017532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_pipe = [get_image_path, PILImage.create]\ny_pipe = [get_label, Categorize()]","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:10.553222Z","iopub.execute_input":"2022-08-13T08:32:10.553686Z","iopub.status.idle":"2022-08-13T08:32:10.562926Z","shell.execute_reply.started":"2022-08-13T08:32:10.553647Z","shell.execute_reply":"2022-08-13T08:32:10.561840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Extra step:** We won't pay much attention to class balance at this point, however, we do need to make sure we have at least 1 sample per class in the training set. In most cases we don't have to worry about that, we would have plenty of examples per class and a random split would almost certainly contain all classes. But this dataset contains a bunch of classes with a single sample, so we have to manually modify the splits to make sure these are on the training set.","metadata":{}},{"cell_type":"code","source":"# make sure at least one of each whale is in training set, then randomly split\nmust_train_whales = targs_df.groupby('whaleID').first()['Image']\n# some magic from [here](https://stackoverflow.com/questions/49823963/get-index-of-one-series-into-another-in-pandas)\nmust_train_ids = pd.Series(targs_df['Image'].index, index=targs_df['Image']).get(must_train_whales)\nmust_train_ids = set(must_train_ids)\n\ntrain_ids, valid_ids = RandomSplitter(seed=42)(items)\nprint(f\"Before: train_ids={len(train_ids)}, valid_ids={len(valid_ids)}\")\ntrain_ids = L(set(train_ids).union(must_train_ids))\nvalid_ids = L(set(valid_ids) - must_train_ids)\nprint(f\"After: train_ids={len(train_ids)}, valid_ids={len(valid_ids)}\")\nsplits = (train_ids, valid_ids)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:12.967703Z","iopub.execute_input":"2022-08-13T08:32:12.968206Z","iopub.status.idle":"2022-08-13T08:32:12.996800Z","shell.execute_reply.started":"2022-08-13T08:32:12.968156Z","shell.execute_reply":"2022-08-13T08:32:12.995803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dss = Datasets(items, [x_pipe, y_pipe], splits=splits)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:13.526623Z","iopub.execute_input":"2022-08-13T08:32:13.529346Z","iopub.status.idle":"2022-08-13T08:32:18.330464Z","shell.execute_reply.started":"2022-08-13T08:32:13.529303Z","shell.execute_reply":"2022-08-13T08:32:18.329371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dss.show(dss[76])","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:18.336015Z","iopub.execute_input":"2022-08-13T08:32:18.338668Z","iopub.status.idle":"2022-08-13T08:32:18.559995Z","shell.execute_reply.started":"2022-08-13T08:32:18.338626Z","shell.execute_reply":"2022-08-13T08:32:18.559030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"after_item = [ToTensor(), Resize((320, 480))]\nafter_batch = [IntToFloatTensor(), *aug_transforms(size=(224, 336))]\n\ndls = dss.dataloaders(32, after_item=after_item, after_batch=after_batch)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:18.564482Z","iopub.execute_input":"2022-08-13T08:32:18.565181Z","iopub.status.idle":"2022-08-13T08:32:23.010598Z","shell.execute_reply.started":"2022-08-13T08:32:18.565141Z","shell.execute_reply":"2022-08-13T08:32:23.009395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:23.018000Z","iopub.execute_input":"2022-08-13T08:32:23.020824Z","iopub.status.idle":"2022-08-13T08:32:24.362494Z","shell.execute_reply.started":"2022-08-13T08:32:23.020779Z","shell.execute_reply":"2022-08-13T08:32:24.361088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training","metadata":{}},{"cell_type":"markdown","source":"No new magic happening here, again refer to the previous blog post for a detailed explanation.","metadata":{}},{"cell_type":"code","source":"metrics = [error_rate]","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:24.364467Z","iopub.execute_input":"2022-08-13T08:32:24.365132Z","iopub.status.idle":"2022-08-13T08:32:24.371359Z","shell.execute_reply.started":"2022-08-13T08:32:24.365089Z","shell.execute_reply":"2022-08-13T08:32:24.370069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn = vision_learner(dls, 'resnet26d', metrics=metrics).to_fp16()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:24.373458Z","iopub.execute_input":"2022-08-13T08:32:24.373839Z","iopub.status.idle":"2022-08-13T08:32:34.173165Z","shell.execute_reply.started":"2022-08-13T08:32:24.373801Z","shell.execute_reply":"2022-08-13T08:32:34.172103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:32:34.175954Z","iopub.execute_input":"2022-08-13T08:32:34.176791Z","iopub.status.idle":"2022-08-13T08:33:12.821541Z","shell.execute_reply.started":"2022-08-13T08:32:34.176746Z","shell.execute_reply":"2022-08-13T08:33:12.820411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fine_tune(10, 0.002)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:37:35.587195Z","iopub.execute_input":"2022-08-13T08:37:35.588017Z","iopub.status.idle":"2022-08-13T08:45:46.724060Z","shell.execute_reply.started":"2022-08-13T08:37:35.587965Z","shell.execute_reply":"2022-08-13T08:45:46.722286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.recorder.plot_loss()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:47:46.962163Z","iopub.execute_input":"2022-08-13T08:47:46.963486Z","iopub.status.idle":"2022-08-13T08:47:47.298901Z","shell.execute_reply.started":"2022-08-13T08:47:46.963438Z","shell.execute_reply":"2022-08-13T08:47:47.297896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Inference","metadata":{}},{"cell_type":"markdown","source":"Same function as before, just copy-pasta.","metadata":{}},{"cell_type":"code","source":"def submit(learn):\n    test_df = pd.read_csv(data_dir / 'sample_submission.csv')\n    test_dl = learn.dls.test_dl(test_df['Image'])\n    \n    preds, targs = learn.get_preds(dl=test_dl)\n    \n    df = pd.DataFrame(preds.numpy(), columns=learn.dls.vocab)\n    df[\"Image\"] = test_df['Image']\n    \n    preds_path = \"submission.csv\"\n    df.to_csv(preds_path, index=False)\n    \n    if not iskaggle:\n        from kaggle import api\n        api.competition_submit_cli(preds_path, \"initial submission\", comp)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:48:01.837934Z","iopub.execute_input":"2022-08-13T08:48:01.838497Z","iopub.status.idle":"2022-08-13T08:48:01.849776Z","shell.execute_reply.started":"2022-08-13T08:48:01.838453Z","shell.execute_reply":"2022-08-13T08:48:01.848465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit(learn)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:48:17.995795Z","iopub.execute_input":"2022-08-13T08:48:17.996312Z","iopub.status.idle":"2022-08-13T08:49:23.237677Z","shell.execute_reply.started":"2022-08-13T08:48:17.996267Z","shell.execute_reply":"2022-08-13T08:49:23.236342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we get a score of 4.54! A massive improvement compared to our last score of `30.29923`, and finally better than random guessing at ` 6.10255`!\n\nIndeed this notebook should have been the first one in the series. It was much easier to produce than the last one and provides a decent baseline to start improving.\n\nFor our next steps, we can dive deeper into what transforms we want to use, scale up the size of images by using a technique known as progressive resizing and only then start thinking of going bigger with models.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}