{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Fast and Agile Resnet34 with fast.ai and fastkaggle","metadata":{}},{"cell_type":"markdown","source":"During the amazing learning experience that was the fastai 2022 course -stay tuned because it will be released for free in a couple of weeks- and Jeremy's live coding sessions, we were guided in an experimenting journey using different models and applying multiple techniques.  \n  \n \nI found resnet34 to be great model for this dataset to experiment with. It is also a very light and fast model to run on my old 8Gb GPU. \n  \nThis Notebook became my baseline model from which I started to test all the techniques that were showed in Jeremy's \"Road to the Top\" series that are linked in the Acknowledgements.\n  \n","metadata":{}},{"cell_type":"markdown","source":"### Acknowledgements\n\n**fastai course:**\n- [Practical Deep Learning for Coders (a UQ collaboration with fast.ai)](https://itee.uq.edu.au/event/2022/practical-deep-learning-coders-uq-fastai)  \n\n**Jeremy's Notebook Series:**\n- [First Steps: Road to the Top, Part 1](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1)\n- [Small models: Road to the Top, Part 2](https://www.kaggle.com/code/jhoward/small-models-road-to-the-top-part-2)\n- [Scaling Up: Road to the Top, Part 3](https://www.kaggle.com/code/jhoward/scaling-up-road-to-the-top-part-3)\n- [Multi-target: Road to the Top, Part 4](https://www.kaggle.com/code/jhoward/multi-target-road-to-the-top-part-4)","metadata":{}},{"cell_type":"markdown","source":"### Installing the libraries\n","metadata":{}},{"cell_type":"code","source":"# fastkaggle allows you to work locally and then submit the results and notebook to Kaggle\n\ntry: import fastkaggle\n\nexcept ModuleNotFoundError:\n    !pip install -Uq fastkaggle\n\nfrom fastkaggle import *","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"competition = 'paddy-disease-classification'\npath = setup_comp(competition, install='fastai \"timm>=0.6.2.dev0\"')\n\nfrom fastai.vision.all import *\nset_seed(42)","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Setting data paths","metadata":{}},{"cell_type":"code","source":"# train images\ntrain_path = path / 'train_images'\ntrain_files = get_image_files(train_path)\n\n# test images\ntest_path = path/'test_images'\ntest_files = get_image_files(test_path).sorted()\n\n# sample submission\nsample_submission = pd.read_csv(path/'sample_submission.csv')\n\n# train labels\ntrain_df = pd.read_csv(path / 'train.csv')\ntrain_df.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Dataloaders for fastai training\nYou can create the dataloader in any of these two ways:\n1. From a `DataBlock`","metadata":{}},{"cell_type":"code","source":"dblock = DataBlock(\n    blocks=(ImageBlock, CategoryBlock),\n    get_items=get_image_files,\n    get_y=parent_label,\n    splitter=RandomSplitter(0.2, seed=42),\n    item_tfms=Resize(480, method='squish'),\n    batch_tfms=aug_transforms(size=224, min_scale=0.75)\n)\n\ndls = dblock.dataloaders(train_path)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2. By using the high level `ImageDataLoaders`","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_folder(\n    train_path, \n    valid_pct=0.2,\n    seed=42,\n    item_tfms=Resize(480, method='squish'),\n    batch_tfms=aug_transforms(size=224, min_scale=0.75)\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create a learner and train","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, resnet34, metrics=error_rate).to_fp16()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And that's it, 16 epochs to get the best baseline for the price ","metadata":{}},{"cell_type":"code","source":"learn.fine_tune(16, 0.005)","metadata":{"scrolled":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Predictions and Test Time Augmentation","metadata":{}},{"cell_type":"markdown","source":"Lets compare the error rate -on the validation set- that are obtained with the normal prediction function and with the predictions we can get applying a technique called Test Time Augmentation (TTA). As you'll see, TTA is easy with fastai.","metadata":{}},{"cell_type":"code","source":"# Get predictions on validation set\nprobs, target = learn.get_preds(dl=dls.valid)\nerror_rate(probs, target)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get TTA predictions on validation set\nprobs, target = learn.tta(dl=dls.valid)\nerror_rate(probs, target)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So you can see a boost with TTA.","metadata":{}},{"cell_type":"markdown","source":"### Predictions on test set","metadata":{}},{"cell_type":"code","source":"# TTA predictions from test images\nprobs, _ = learn.tta(dl=dls.test_dl(test_files))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the index with the greater probability\npreds = probs.argmax(dim=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.vocab[preds]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission","metadata":{}},{"cell_type":"code","source":"sample_submission.label = dls.vocab[preds]\nsample_submission.to_csv('submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusions","metadata":{}},{"cell_type":"markdown","source":"* I found this model being a good baseline, with a good accuracy for its speed and cost.\n* You can try different epochs, learning rates, or even a different seed and see what happens when submitting the results.\n* Then you can apply some of the techniques that Jeremy applied in his series.\n* And then \n\n","metadata":{}},{"cell_type":"code","source":"# Pushing the notebook from my home PC to Kaggle\n\nif not iskaggle:\n    push_notebook(\n        'fmussari', \n        'fast-resnet34-with-fastai',\n        title='Fast Resnet34 with Fastai',\n        file='2022-07. Fast and Agile Resnet34 with Fastai.ipynb',\n        competition=competition, \n        private=True, \n        gpu=True\n    )","metadata":{},"execution_count":null,"outputs":[]}]}