{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Understanding the dataset"},{"metadata":{},"cell_type":"markdown","source":"****Let us print the dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"import os","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"print(\"There are following directories and files in this dataset\")\nprint(*list(os.listdir(\"../input/cassava-leaf-disease-classification\")),sep = \"\\n\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Choosing the data for model training"},{"metadata":{},"cell_type":"markdown","source":"****We will now count the number of images in the following directories:"},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\ntrain_images_jpg_format = glob.glob('../input/cassava-leaf-disease-classification/train_images/*.jpg')\nlen(train_images_jpg_format)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Importing necessary libraries"},{"metadata":{},"cell_type":"markdown","source":"**Import Pandas** - For data analysis \n**Import Fastai** - For training of deep learning model and predictions.\n\n**Note**: We are using Fastai version 2 (Not previous version of Fastai- which is version 1).\n\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\n\nimport fastai\nfrom fastai.vision.all import *\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Defining the variables and assigning the paths for this notebook"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/cassava-leaf-disease-classification')\nimage_path = Path('../input/cassava-leaf-disease-classification/train_images')\ntraining_data_file = Path('../input/cassava-leaf-disease-classification/train.csv')\nsample_submission_file = Path('../input/cassava-leaf-disease-classification/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analysing the data\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(training_data_file)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Size of Training data \\n\", train_df.shape)\nprint(\"----------------------------------------------------------\")\nprint(\"\\nFirst few samples of data are \\n\",train_df.head())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us print the number of data samples with each output category"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['label'].value_counts()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Selecting a subset of data for training purpose"},{"metadata":{},"cell_type":"markdown","source":"We will select all the training data which has the output category as \"1\",\"2\",\"4\" and 0.196 % of training data which has the output category as \"3\" to have equal number of inputs with the same category of output."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df_output_3 = train_df[train_df['label']==3].sample(frac=0.196,random_state=111)\ntrain_df_output_3.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us join the inputs selected from both the categories and call it a \"new_df\"."},{"metadata":{"trusted":true},"cell_type":"code","source":"new_df = pd.concat([train_df[train_df['label']==0],train_df[train_df['label']==1],train_df[train_df['label']==2],train_df_output_3,train_df[train_df['label']==4]]).reset_index(drop=True)\nnew_df.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Creating the image data loader"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_data_loader = ImageDataLoaders.from_df(new_df, path=image_path,\n                               seed=42, fn_col=0, \n                               label_col=1, \n                               item_tfms=Resize(128), \n                               batch_tfms=aug_transforms(flip_vert=True, max_warp=0.), \n                               bs=128, val_bs=None, shuffle_train=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us check the device type of our \"ImageDataLoader\" to make sure that we are using \"GPU\""},{"metadata":{"trusted":true},"cell_type":"code","source":"image_data_loader.device","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us check few random images from our ImageDataLoader's batch to make sure that images and labels appears correctly in it."},{"metadata":{"trusted":true},"cell_type":"code","source":"image_data_loader.show_batch()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Trainnig the image recognizer model"},{"metadata":{},"cell_type":"markdown","source":"We create a CNN (convolutional neural network) with the following specific details:\n\n* What data we want to train it on? </br> Our data to be used for training is \"image_data_loader\"\n\n* Which architecture to use? </br> We are using Resnet34\n\n* what metric to use for our training evaluation? </br> We have specified it as \"error_rate\""},{"metadata":{"trusted":true},"cell_type":"code","source":"learn = cnn_learner(image_data_loader, resnet34, metrics=error_rate)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us train the model for 2 epochs"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fine_tune(2,freeze_epochs = 4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(4,slice(1e-5,1e-3))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> # Bring it on -  Test data !!"},{"metadata":{},"cell_type":"markdown","source":"Defining the variables and assigning the paths for **test dataset**"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_image_files = Path('../input/cassava-leaf-disease-classification/test_images')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us create a **ImageDataLoader** of our test data set"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_data_loader_test = image_data_loader.test_dl(get_image_files(test_image_files))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Make the predictions using our trained model called \"**learn**\".\n* Ignoring the first two outputs from the model, let us take our final result stored in variable \"**predictions**\""},{"metadata":{"trusted":true},"cell_type":"code","source":"_,_,results = learn.get_preds(dl = image_data_loader_test, with_decoded = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = pd.read_csv(sample_submission_file)\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub['label'] = results","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('my_submission_file.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}