{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":4521,"databundleVersionId":326986,"sourceType":"competition"}],"dockerImageVersionId":30787,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"markdown","source":"Downloading libraries and imports same as before. We'll be hiding those cells.","metadata":{}},{"cell_type":"code","source":"try: import fastkaggle\nexcept ModuleNotFoundError:\n    !pip install -Uq fastkaggle\n\nfrom fastkaggle import *","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-24T06:57:25.027077Z","iopub.execute_input":"2024-10-24T06:57:25.027417Z","iopub.status.idle":"2024-10-24T06:57:38.341309Z","shell.execute_reply.started":"2024-10-24T06:57:25.027384Z","shell.execute_reply":"2024-10-24T06:57:38.340321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ensure fastkaggle is installed and imported for seamless Kaggle workflows.s available.","metadata":{}},{"cell_type":"code","source":"try: import timm\nexcept ModuleNotFoundError:\n    !pip install \"timm>=0.6.2.dev0\"","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-24T06:57:38.343386Z","iopub.execute_input":"2024-10-24T06:57:38.343991Z","iopub.status.idle":"2024-10-24T06:57:44.231563Z","shell.execute_reply.started":"2024-10-24T06:57:38.343944Z","shell.execute_reply":"2024-10-24T06:57:44.230735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ensure `timm` library is installed for PyTorch image models; install if missing.","metadata":{}},{"cell_type":"code","source":"from fastai.vision.all import *","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-10-24T06:57:44.232747Z","iopub.execute_input":"2024-10-24T06:57:44.233079Z","iopub.status.idle":"2024-10-24T06:57:45.573885Z","shell.execute_reply.started":"2024-10-24T06:57:44.233045Z","shell.execute_reply":"2024-10-24T06:57:45.572914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Import all functions and classes from the `fastai.vision` module for deep learning tasks in computer vision.","metadata":{}},{"cell_type":"code","source":"comp = 'noaa-right-whale-recognition'\ndata_dir = setup_comp(comp, install='fastai \"timm>=0.6.2.dev0\"')\nimgs_dir = untar_dir(data_dir/'imgs.zip', Path('imgs'))","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-24T06:57:45.576525Z","iopub.execute_input":"2024-10-24T06:57:45.576956Z","iopub.status.idle":"2024-10-24T07:01:41.670115Z","shell.execute_reply.started":"2024-10-24T06:57:45.576895Z","shell.execute_reply":"2024-10-24T07:01:41.668756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code prepares the environment by downloading the competition data and extracting image files, while also ensuring that required libraries like fastai and timm are installed for working with deep learning models.\r\n\r\n\r\n\r\n\r\n\r\n\r\n","metadata":{}},{"cell_type":"code","source":"def plot_sizes(image_sizes, max_n=3000):\n    sizes = parallel(image_size, image_files[:max_n], progress=progress_bar)\n    \n    widths, heights = zip(*sizes)\n    min_x = min(widths + heights)\n    max_x = max(widths + heights)\n    plt.hist(widths, label='width', range=(min_x, max_x))\n    plt.hist(heights, label='height', range=(min_x, max_x))\n    plt.legend();","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:01:41.672038Z","iopub.execute_input":"2024-10-24T07:01:41.672765Z","iopub.status.idle":"2024-10-24T07:01:41.684921Z","shell.execute_reply.started":"2024-10-24T07:01:41.672704Z","shell.execute_reply":"2024-10-24T07:01:41.683979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This function computes and visualizes the distribution of image widths and heights to understand their size variations. It uses parallel processing to speed up the dimension extraction and ensures both histograms are plotted with the same range for consistency.","metadata":{}},{"cell_type":"code","source":"image_files = get_image_files(imgs_dir)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:01:41.688214Z","iopub.execute_input":"2024-10-24T07:01:41.690158Z","iopub.status.idle":"2024-10-24T07:01:43.261794Z","shell.execute_reply.started":"2024-10-24T07:01:41.690111Z","shell.execute_reply":"2024-10-24T07:01:43.260760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This line helps gather the image data from the extracted folder (imgs_dir) so it can be further processed, such as for analyzing dimensions, training machine learning models, or visualizing data distributions.","metadata":{}},{"cell_type":"code","source":"plot_sizes(image_files)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:01:43.263037Z","iopub.execute_input":"2024-10-24T07:01:43.263413Z","iopub.status.idle":"2024-10-24T07:01:55.378346Z","shell.execute_reply.started":"2024-10-24T07:01:43.263371Z","shell.execute_reply":"2024-10-24T07:01:55.377388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the plot above we can see that majority of images are about `3000x2000`. For our first model we would like to try something around `224`. Opening these big images and downsizing by this much will bottleneck the training pipeline (you will notice the GPU will be at 0% usage for long periods).\n\nLet's resize all images to 480 and save them to disk, fastai already provides a handy function for doing this.","metadata":{}},{"cell_type":"code","source":"if not Path('imgs_480max').exists():\n    resize_images(imgs_dir, dest='imgs_480max', max_size=480, progress=progress_bar)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:01:55.379813Z","iopub.execute_input":"2024-10-24T07:01:55.380526Z","iopub.status.idle":"2024-10-24T07:12:48.923054Z","shell.execute_reply.started":"2024-10-24T07:01:55.380471Z","shell.execute_reply":"2024-10-24T07:12:48.922127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code ensures that images are resized (if not already done) to make them uniform with a maximum size of 480 pixels, which can help optimize model training and improve memory usage. The resized images are saved to a new folder ('imgs_480max') for further use.","metadata":{}},{"cell_type":"code","source":"image_files = get_image_files('imgs_480max')\nplot_sizes(image_files)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:48.924286Z","iopub.execute_input":"2024-10-24T07:12:48.924610Z","iopub.status.idle":"2024-10-24T07:12:53.281866Z","shell.execute_reply.started":"2024-10-24T07:12:48.924572Z","shell.execute_reply":"2024-10-24T07:12:53.280923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Together, these lines ensure that the image files from the resized directory are processed and visualized, allowing you to assess the dimensions of the resized images, confirming that they meet the desired size criteria and facilitating further analysis or model training.","metadata":{}},{"cell_type":"markdown","source":"## Data Pipeline","metadata":{}},{"cell_type":"code","source":"targs_df = pd.read_csv(data_dir/'train.csv')\ntargs_df = targs_df[targs_df['Image']!='w_7489.jpg']\nitems = targs_df['Image']","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.285463Z","iopub.execute_input":"2024-10-24T07:12:53.285768Z","iopub.status.idle":"2024-10-24T07:12:53.361553Z","shell.execute_reply.started":"2024-10-24T07:12:53.285733Z","shell.execute_reply":"2024-10-24T07:12:53.360760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reads the train.csv file into a DataFrame, excludes the row for the image 'w_7489.jpg', and extracts the remaining image names into the items variable.","metadata":{}},{"cell_type":"code","source":"def get_image_path(name):\n    return Path('imgs_480max')/name","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.363535Z","iopub.execute_input":"2024-10-24T07:12:53.363857Z","iopub.status.idle":"2024-10-24T07:12:53.368399Z","shell.execute_reply.started":"2024-10-24T07:12:53.363822Z","shell.execute_reply":"2024-10-24T07:12:53.367348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\r\nDefines a function `get_image_path(name)` that constructs and returns the file path for an image in the `'imgs_480max'` directory based on the provided image name.","metadata":{}},{"cell_type":"code","source":"def get_label(name):\n    return targs_df[targs_df['Image']==name]['whaleID'].item()","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.369496Z","iopub.execute_input":"2024-10-24T07:12:53.369781Z","iopub.status.idle":"2024-10-24T07:12:53.380359Z","shell.execute_reply.started":"2024-10-24T07:12:53.369750Z","shell.execute_reply":"2024-10-24T07:12:53.379566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Defines a function get_label(name) that retrieves and returns the corresponding whaleID for a given image name from the targs_df DataFrame.","metadata":{}},{"cell_type":"code","source":"x_pipe = [get_image_path, PILImage.create]\ny_pipe = [get_label, Categorize()]","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.381387Z","iopub.execute_input":"2024-10-24T07:12:53.381681Z","iopub.status.idle":"2024-10-24T07:12:53.392110Z","shell.execute_reply.started":"2024-10-24T07:12:53.381645Z","shell.execute_reply":"2024-10-24T07:12:53.391244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sets up processing pipelines `x_pipe` for loading images using their paths and `y_pipe` for retrieving labels and categorizing them.","metadata":{}},{"cell_type":"code","source":"# make sure at least one of each whale is in training set, then randomly split\nmust_train_whales = targs_df.groupby('whaleID').first()['Image']\n# some magic from [here](https://stackoverflow.com/questions/49823963/get-index-of-one-series-into-another-in-pandas)\nmust_train_ids = pd.Series(targs_df['Image'].index, index=targs_df['Image']).get(must_train_whales)\nmust_train_ids = set(must_train_ids)\n\ntrain_ids, valid_ids = RandomSplitter(seed=42)(items)\nprint(f\"Before: train_ids={len(train_ids)}, valid_ids={len(valid_ids)}\")\ntrain_ids = L(set(train_ids).union(must_train_ids))\nvalid_ids = L(set(valid_ids) - must_train_ids)\nprint(f\"After: train_ids={len(train_ids)}, valid_ids={len(valid_ids)}\")\nsplits = (train_ids, valid_ids)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.393151Z","iopub.execute_input":"2024-10-24T07:12:53.393503Z","iopub.status.idle":"2024-10-24T07:12:53.470317Z","shell.execute_reply.started":"2024-10-24T07:12:53.393469Z","shell.execute_reply":"2024-10-24T07:12:53.469442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code ensures that at least one image of each whale species is included in the training set while randomly splitting the remaining images into training and validation sets. It prints the counts of images before and after the adjustments to confirm that the requirements are met.","metadata":{}},{"cell_type":"code","source":"dss = Datasets(items, [x_pipe, y_pipe], splits=splits)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:53.471342Z","iopub.execute_input":"2024-10-24T07:12:53.471639Z","iopub.status.idle":"2024-10-24T07:12:58.181443Z","shell.execute_reply.started":"2024-10-24T07:12:53.471606Z","shell.execute_reply":"2024-10-24T07:12:58.180625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Creates a Datasets object dss using the specified items, processing pipelines (x_pipe for images and y_pipe for labels), and the defined splits for training and validation data.","metadata":{}},{"cell_type":"code","source":"dss.show(dss[76])","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:58.182655Z","iopub.execute_input":"2024-10-24T07:12:58.183016Z","iopub.status.idle":"2024-10-24T07:12:58.439802Z","shell.execute_reply.started":"2024-10-24T07:12:58.182981Z","shell.execute_reply":"2024-10-24T07:12:58.438832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Displays the 76th item in the Datasets object dss, showing the processed image and its corresponding label.","metadata":{}},{"cell_type":"code","source":"after_item = [ToTensor(), Resize((320, 480))]\nafter_batch = [IntToFloatTensor(), *aug_transforms(size=(224, 336))]\n\ndls = dss.dataloaders(32, after_item=after_item, after_batch=after_batch)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:58.440934Z","iopub.execute_input":"2024-10-24T07:12:58.441256Z","iopub.status.idle":"2024-10-24T07:12:59.419209Z","shell.execute_reply.started":"2024-10-24T07:12:58.441216Z","shell.execute_reply":"2024-10-24T07:12:59.418339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prepares data loaders dls from the Datasets object dss, applying specified transformations (after_item for individual items and after_batch for batches) with a batch size of 32.","metadata":{}},{"cell_type":"code","source":"dls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:12:59.420490Z","iopub.execute_input":"2024-10-24T07:12:59.420827Z","iopub.status.idle":"2024-10-24T07:13:00.637108Z","shell.execute_reply.started":"2024-10-24T07:12:59.420786Z","shell.execute_reply":"2024-10-24T07:13:00.636134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training","metadata":{}},{"cell_type":"markdown","source":"No new magic happening here, again refer to the previous blog post for a detailed explanation.","metadata":{}},{"cell_type":"code","source":"metrics = [error_rate]","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:13:00.638623Z","iopub.execute_input":"2024-10-24T07:13:00.639029Z","iopub.status.idle":"2024-10-24T07:13:00.643724Z","shell.execute_reply.started":"2024-10-24T07:13:00.638985Z","shell.execute_reply":"2024-10-24T07:13:00.642829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Defines a list metrics containing the error_rate function to evaluate the model's performance during training or validation.","metadata":{}},{"cell_type":"code","source":"pip install huggingface_hub","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:13:00.645048Z","iopub.execute_input":"2024-10-24T07:13:00.645436Z","iopub.status.idle":"2024-10-24T07:13:12.246359Z","shell.execute_reply.started":"2024-10-24T07:13:00.645393Z","shell.execute_reply":"2024-10-24T07:13:12.245238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Installs the huggingface_hub library, which provides tools for accessing and sharing models and datasets on the Hugging Face Hub.","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, 'resnet26d', metrics=metrics,pretrained = False).to_fp16()","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:13:12.247954Z","iopub.execute_input":"2024-10-24T07:13:12.248307Z","iopub.status.idle":"2024-10-24T07:13:12.542824Z","shell.execute_reply.started":"2024-10-24T07:13:12.248267Z","shell.execute_reply":"2024-10-24T07:13:12.541518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Creates a vision_learner object learn with a ResNet26D architecture using the specified data loaders dls, sets the evaluation metrics, and enables half-precision training with to_fp16().","metadata":{}},{"cell_type":"code","source":"learn.lr_find()","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:13:12.544379Z","iopub.execute_input":"2024-10-24T07:13:12.544993Z","iopub.status.idle":"2024-10-24T07:13:35.984269Z","shell.execute_reply.started":"2024-10-24T07:13:12.544943Z","shell.execute_reply":"2024-10-24T07:13:35.983346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Runs the learning rate finder for the learn object, which helps identify the optimal learning rate for training the model by plotting the loss against different learning rates.","metadata":{}},{"cell_type":"code","source":"learn.fine_tune(10, 0.002)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:13:35.985841Z","iopub.execute_input":"2024-10-24T07:13:35.986226Z","iopub.status.idle":"2024-10-24T07:18:32.858296Z","shell.execute_reply.started":"2024-10-24T07:13:35.986188Z","shell.execute_reply":"2024-10-24T07:18:32.857315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fine-tunes the model using the learn object for 10 epochs with a maximum learning rate of 0.002, adapting the pre-trained weights to the new dataset.","metadata":{}},{"cell_type":"code","source":"learn.recorder.plot_loss()","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:18:32.860047Z","iopub.execute_input":"2024-10-24T07:18:32.860429Z","iopub.status.idle":"2024-10-24T07:18:33.250027Z","shell.execute_reply.started":"2024-10-24T07:18:32.860388Z","shell.execute_reply":"2024-10-24T07:18:33.249021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plots the training and validation loss curves from the learn object’s recorder, visualizing the model's performance over the training epochs.","metadata":{}},{"cell_type":"markdown","source":"## Inference","metadata":{}},{"cell_type":"markdown","source":"Same function as before, just copy-pasta.","metadata":{}},{"cell_type":"code","source":"def submit(learn):\n    test_df = pd.read_csv(data_dir / 'sample_submission.csv')\n    test_dl = learn.dls.test_dl(test_df['Image'])\n    \n    preds, targs = learn.get_preds(dl=test_dl)\n    \n    df = pd.DataFrame(preds.numpy(), columns=learn.dls.vocab)\n    df[\"Image\"] = test_df['Image']\n    \n    preds_path = \"submission.csv\"\n    df.to_csv(preds_path, index=False)\n    \n    if not iskaggle:\n        from kaggle import api\n        api.competition_submit_cli(preds_path, \"initial submission\", comp)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:18:33.251552Z","iopub.execute_input":"2024-10-24T07:18:33.252230Z","iopub.status.idle":"2024-10-24T07:18:33.258653Z","shell.execute_reply.started":"2024-10-24T07:18:33.252194Z","shell.execute_reply":"2024-10-24T07:18:33.257618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit(learn)","metadata":{"execution":{"iopub.status.busy":"2024-10-24T07:18:33.259823Z","iopub.execute_input":"2024-10-24T07:18:33.260161Z","iopub.status.idle":"2024-10-24T07:18:54.881917Z","shell.execute_reply.started":"2024-10-24T07:18:33.260127Z","shell.execute_reply":"2024-10-24T07:18:54.880975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we get a score of 4.54! A massive improvement compared to our last score of `30.29923`, and finally better than random guessing at ` 6.10255`!\n\nIndeed this notebook should have been the first one in the series. It was much easier to produce than the last one and provides a decent baseline to start improving.\n\nFor our next steps, we can dive deeper into what transforms we want to use, scale up the size of images by using a technique known as progressive resizing and only then start thinking of going bigger with models.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}