{"cells":[{"metadata":{},"cell_type":"markdown","source":"As the title suggests, we will be using fastai to achieve the best possible score with minimum lines of code.\n\nThis is the training notebook, you can find the [inference notebook here](https://www.kaggle.com/ankursingh12/fastai-plant2021-starter-inference).\n\nLets get started . . . \n\nFirst, we will to import fastai."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from fastai.vision.all import *\n\nseed = 42\nset_seed(seed, reproducible=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We are setting the seed for reproducibility. \n\nNext, we will initialize some path (& other) variables (for use throughout the notebook)\n\nPS: I will be using my version for the dataset. I have resized all the images so that its much faster to load them into the RAM. You can find the dataset [here](https://www.kaggle.com/ankursingh12/resized-plant2021). "},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/plant-pathology-2021-fgvc8')\ndata_path = Path('../input/resized-plant2021')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Data\n\nEnough prep-work! Lets read our data . . ."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv(path/'train.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"hmm, just image names and their label. Looks simple ? Not so soon. The labels are space-delimited strings. Its a multi-label problem. \n\nWe will use Fastai's datablock API to load our data. DataBlock API is simply amazing. Infinitely  flexibility and incredibly powerful. To use the datablock API, you need to define some functions."},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_x(x): return str(data_path/'img_sz_640') + os.path.sep + x['image']\ndef get_y(y): return y['labels']\n\ndatablock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n                   splitter=RandomSplitter(seed=seed),\n                   get_x=get_x, get_y=get_y,\n                   item_tfms = RandomResizedCrop(512),\n                   batch_tfms=[*aug_transforms(mult=2.0,flip_vert=True, size=460), \n                               Normalize.from_stats(*imagenet_stats)])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Don't worry a lot if the above code looks cryptic. You can read the [6th notebook (or chapter)](https://github.com/fastai/fastbook/blob/master/06_multicat.ipynb) in fastbook for details. It explains the topic in the most simplest way possible. And once you master datablock API, you will feel like a Ninja (trust me on this)!\n\nYou are amazing! Now lets create our dataloaders, & then take a look at some images."},{"metadata":{"trusted":true},"cell_type":"code","source":"dls = datablock.dataloaders(df)\ndls.show_batch(max_n=9)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks good to me, what do you think?\n\nWe are done with data, time for some training."},{"metadata":{},"cell_type":"markdown","source":"### Model\n\nFastai has an awesome class which puts everything together, called `cnn_learner`. Here we are using ResNet50. "},{"metadata":{"trusted":true},"cell_type":"code","source":"f1score = F1Score(average='macro')\nlearn = cnn_learner(dls, resnet50, metrics=[accuracy, f1score]).to_fp16()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We are using `accuracy_multi` and `f1score` metrics, because its a multi-label problem and the evaluation metric for the competition is *F1Score*.\n\nFinally, lets train (technically, fine-tune 🤯) our model."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fine_tune(5, 3e-3, wd=0.5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"lets train it some more"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(5, slice(3e-3), wd=0.5) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Okay, done with training! Lets look at some predictions . . ."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.show_results()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Amazing! Lets export the model so that we can deploy it to production 😂. Just kidding, we will (only) use it for inference."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.export(f'resnet50.pkl')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Fastai is extremely flexible and powerful at the same time. This is just the baseline notebook. You can easily build on top of it. Here are some things that you can experiment with:\n\n- Preprocess and Feature Engineering\n- Data Augmentation and External Datasets\n- Different Model Architectures\n- Training Schedule, Optimizer, etc\n- Postprocess\n\nYou can find the **[inference notebook here](https://www.kaggle.com/ankursingh12/fastai-plant2021-starter-inference)**.\n\nHope you had fun reading the notebook. Kindly consider **upvoting**."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}