{"cells":[{"metadata":{"_uuid":"52bf7bee263bbde2ef83a9523a89454579443974"},"cell_type":"markdown","source":"In this notebook, I will prepare whale image with Fast AI datablock API. Addressing the train - validation split problem that many might experience using FastAI 1.0 functions. \n\nAlso, I hope this kernel can help starters like me to take advantage of the Fast AI and lay down baseline model. \n\nI am still learning Fast AI and pytorch, please feel free to leave me a comment as I will be learned from my mistakes :) \n\nFor people that are interested, here is the link to the Fast AI documentation \n[fastai doc](https://docs.fast.ai/data_block.html)"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nfrom fastai.vision import *\nfrom fastai.basic_data import *","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"156a4ae1c2bf82d1db833fdafbd83fdbe3f38b4e"},"cell_type":"markdown","source":"# Take a look of the label data"},{"metadata":{"trusted":true,"_uuid":"a239eb9dbcaf8cccc7381d7fdd6b85fc7bdec3af"},"cell_type":"code","source":"df_label = pd.read_csv('../input/train.csv')\ndf_label.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ae584cc6a520948a6b42b5f17d4a5bc9e0559cc"},"cell_type":"markdown","source":"I wont be talking to much about data exploration, as there are already many great kernels talk in depth about the data. \n\nThere are just couple things people might fall in trap later on when creating databunch object\n\n1. There are 5005 unique labels, that means 5005 classes if we load them all into Fast AI. \n2. We will have rare labels that we only have 1 image, as well 2 images...3 images as so on. "},{"metadata":{"trusted":true,"_uuid":"896270ddf88c1c59440c725be1fc8033d63e4b30"},"cell_type":"markdown","source":"# **Using FastAI data block API**\nThe defualt factory method won't work in this case, as it will be hard for you to manually put the data into imageNet style, also, it seems to me very time consuming if you try to create 5005 folders and manually separate each class.\n\nTherefore we can use data block API, following the documation, it will just be 4 steps\n\n1. Provide inputs\n2. Split data\n3. Label data\n4. Create databunch to pass to pytorch"},{"metadata":{"_uuid":"68e49a5af92787d7e68eb644652decd6bffea47f"},"cell_type":"markdown","source":"Here is what we can do\n1. It is vision data, we can create a ImageItemList object\n2. We can do random split, since it is a baseline model. \n3. We have labels in the train.csv files, and label happens to be col 1, so we can call the default function to handle for as "},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"cdd0e918257d7c6562982be75601ef67c0f403c8"},"cell_type":"code","source":"src = (ImageItemList.from_csv('../input/','train.csv',folder='train')\n        .random_split_by_pct()\n        .label_from_df())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"30223f0c4ef43fb377f9f10d08224e07251390e1"},"cell_type":"markdown","source":"***Exception: Your validation data contains a label that isn't present in the training set, please fix your data***\n\nMajor block people might have using Fast AI libarary is during train valid split (it happens to me at least)\n\nIf you call random_split_by_pct(), default will split the data randomly with 80-20. There will be cases that the class with only 1 image splitted into validation set, now your validation set has a class that train set never knows. Therefore you have the FastAI complaining, because the model can't predict that class since the model never sees it during training. \n\nWhale image data set is different as other datasets talked in the class. If you treat it as classfication problem, we are  actually trying to classify an item from 5005 different classes. A good validation set is very important, but we are not addressing it here since we just want to get a baseline model. "},{"metadata":{"_uuid":"46b4a626d43923b6a023aecb61cd6c009c28bcbb"},"cell_type":"markdown","source":"# Novice way to handle \n\nWe already know what problem we have, validation set has classes that train set doesn't. \nWe can fix it by not spliting the data, that we can at least create the databunch and view the data/train the model with default hyper-parameters\n\nBut you lose the ability to tune the model since you don't have a validation set, we will see how we can fix this problem later on. (At least we get the things going)"},{"metadata":{"trusted":true,"_uuid":"9c7e55e6e406c4fa3fc8d1f77cb2e1f316e4a16a"},"cell_type":"code","source":"src = (ImageItemList.from_csv('../input/','train.csv',folder='train')\n        .no_split()\n        .label_from_df())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6114ced4b52ab32f22c7d82d31ff5b46d18c7f93"},"cell_type":"code","source":"data = (src.transform(get_transforms(),size=224)\n       .databunch()\n       .normalize(imagenet_stats))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"95e9976b214110f8ac16e63e53922c1b5886a84f"},"cell_type":"code","source":"data.show_batch(rows=3,figsize=(12,7))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1c6211eb75d73f9e86041c41fa30c3c34e428eb"},"cell_type":"markdown","source":"Something went wrong again, as I researched a bit, this seems to be kaggle kernel / pytorch issue?\nBut you can simply fix this problem by just let 1 single CPU handle the dataloading step"},{"metadata":{"trusted":true,"_uuid":"ff5c6b1f0642dde245c87525a9f17976badb8629"},"cell_type":"code","source":"data = (src.transform(get_transforms(),size=224)\n       .databunch(num_workers=0)\n       .normalize(imagenet_stats))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"32571072d920ea86c943ea403e3174f010e8a315"},"cell_type":"code","source":"data.show_batch(rows=3,figsize=(12,9))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8493221ecb4704d8c62fd7d5446685324e75a469"},"cell_type":"markdown","source":"There we go, now the data is ready, we can start training the model.\nHowever, there is no way you can tune your model, you don't want to use test set to evaluate your model.\n\nWe need somehow give our model a validation set, so we know how well is our model.\n\n**Split the data with 20% to validation set**\n\nWe failed spliting the data before, because we try to split the whole data as 80% train 20% validation.\nBut if the class has only 1 image, we want it to be in the train-set. We will try our best to learn the class, and if we see it in the test set, we want to make right prediction. \n\nTherefore, we can do 20% split on the sub-set of the train set.\n\n1. If the class has more than 2 image, we pick 1 to the validation set\n2. If the class has 2 or less image, we let it stay on the train set\n\nAlso, you can do data agrumentation. To create more images for each small classes to solve the issue (but you still want to make sure you sub-sampling correctly) "},{"metadata":{"_uuid":"285cd3aadbe967d135d8d8fb00e693a59640d719"},"cell_type":"markdown","source":"# **Sub-sampling**\n\n1. Create a copy of the df_label\n2. Create a new column called total to track the number of images in that class\n3. Create a new data frame that sub-sampling from each classes. \n    * If the class has 1 image, we won't take it to validation set\n    * If the class has 2 images, we won't take it to validation set\n    * If the class has 3 images, we take 1 to validation set, and leave 2 in the train set\nand so on. \n\nWe can also adjust the threshold later, for example, if we have 5 images, we take 1 to validation set, and leave the classes with 5 images or less in training. "},{"metadata":{"trusted":true,"_uuid":"88a160afdfceb7b3b10ebde9cb5d211cc2f170f8"},"cell_type":"code","source":"df_test_split = df_label.copy()\ndf_test_split['total'] = df_test_split.groupby('Id')['Id'].transform('count')\ndf_grouped = df_test_split.groupby('Id').apply(lambda x: x.sample(frac=0.2,random_state=47))\ndf_grouped.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"171ba281f6157458528f259879bbf7b9b56b3613"},"cell_type":"markdown","source":"Now we have randomly picked 4389 images from the train set, each of them have at least 3 images originally in the train set. \nWe can take a close look to make sure we are selecting the right ones"},{"metadata":{"trusted":true,"_uuid":"39352ad533d3d6ca53d9da70d82480933d785f47"},"cell_type":"code","source":"df_grouped.tail(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8b76b5c2a96cc6241f55cf4bb0c10015dc656538"},"cell_type":"code","source":"df_merged = pd.merge(left=df_test_split,right=df_grouped,on='Image',how='left',suffixes=('','_y'))\ndf_merged['is_valid'] = df_merged.Id_y.isnull()!=True\ndf_merged.head(20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43cb65cf55879d5827b9073ca0d00033737817ef"},"cell_type":"markdown","source":"Drop the merged colums that we dont need from the final table"},{"metadata":{"trusted":true,"_uuid":"1e99e97906344c7fd29629ef9c9bc152216fd0f1"},"cell_type":"code","source":"df_merged.drop(['Id_y','total_y'],axis=1,inplace=True)\ndf_merged.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ede886f1c8fe86f5fe23718e08eac03ee6b18f9d"},"cell_type":"markdown","source":"Now we are ready, we have marked our original train-set with new columns if it should split to validation set or not. \nWe can start load it to FastAI \n\nSince  .from_csv will need a csv file, we will pack our dataframe to csv"},{"metadata":{"trusted":true,"_uuid":"31af884216f04cfc74862106e632e7ec4dd08fdf"},"cell_type":"code","source":"df_merged.to_csv('validation_random.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e01a68f9780078f7a62d37cf80bafb681a9a7773"},"cell_type":"code","source":"src = (ImageItemList.from_csv('../input/','/kaggle/working/validation_random.csv',folder='train')\n        .split_from_df(col='is_valid')\n        .label_from_df(cols='Id'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43b8c7919673e3315f1a41315c61ae9d0fa91514"},"cell_type":"code","source":"data = (src.transform(get_transforms(max_zoom=1, max_warp=0),resize_method=ResizeMethod.SQUISH,size=224)\n       .databunch(num_workers=0)\n       .normalize(imagenet_stats))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"043d00854f1d4a4642cfb8c1cb5731cc5ecd518d"},"cell_type":"code","source":"data.show_batch(rows=3,figsize=(12,9))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4a60584e72f139547a42bd351935fc05a471a7e8"},"cell_type":"markdown","source":"Now you have databunch object ready, and train-validation set ready. \nYou can start calling learner to build the baseline model.\n\nCouple things note here:\n1. You need to create a MAP5 metric to evaluate your model, FastAI doesn't have default MAP5 metric. \nYou can read more in [here](http:/https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric/) to construct a MAP5 metric to evaluate your model\n\n2. Kaggle kernel doesn't allow you to download model to '../input/, so if you are just calling learner_create you will have some 'READ only' issue. (I havn't figured out a work around, so I trained on other VMs)\n\nThanks for taking your time to read this kernel. Hope this helps FastAI starters. "},{"metadata":{"trusted":true,"_uuid":"2670a84c009e3e986bd02093e899aa01ba9304f3"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}