{"cells":[{"metadata":{"_uuid":"b41b6f01833e0d8f363dc032b77c42c2fcd28448"},"cell_type":"markdown","source":"This is a basic notebook to train a classifier using the fastai library. We will start with a Resnet18 that is light and fast to train."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import fastai\nfrom fastai.vision import *","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d80088e13b66982bf705b669412666173675978f"},"cell_type":"markdown","source":"Let's define some paths."},{"metadata":{"trusted":true,"_uuid":"06091590c14ead57b0371347cbd313aa662bcdfd"},"cell_type":"code","source":"work_dir = Path('/kaggle/working/')\npath = Path('../input')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fca005f17d62131a0ff3adc201b62c334b85a240"},"cell_type":"code","source":"train = 'train_images/train_images'\ntest =  path/'leaderboard_test_data/leaderboard_test_data'\nholdout = path/'leaderboard_holdout_data/leaderboard_holdout_data'\nsample_sub = path/'SampleSubmission.csv'\nlabels = path/'traininglabels.csv'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"894eaccf852e7d0e0c71f1e91239000ff7e555eb"},"cell_type":"code","source":"df = pd.read_csv(labels)\ndf_sample = pd.read_csv(sample_sub)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5ace8fc87710353e5ea2bb773b01ae890745464b"},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"569516214d2a6d83c028fdb4d37169f65ba7ac23"},"cell_type":"code","source":"df.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6e9cda6be24f987f32f2adbbad1c42aa3aa890d"},"cell_type":"markdown","source":"There are some rows with low score, we will look into that later."},{"metadata":{"trusted":true,"_uuid":"a64c8d64715f7d4645af7fbbe45db467b8641cfe"},"cell_type":"code","source":"df[df['score']<0.75]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c89bf4cf4124f1a47506d9b769220bd44f0d5d0"},"cell_type":"markdown","source":"There are 942 observed palmoil plantations, roughly 8% of total images."},{"metadata":{"trusted":true,"_uuid":"7cb4697b19b9f68e39df2c3cc0223b50bab96ca3"},"cell_type":"code","source":"(df.has_oilpalm==1).sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d0f9e5023a33e7f4ad467a9a80d88ae87bf04aad"},"cell_type":"markdown","source":"We have to combine test and holdout for the submission"},{"metadata":{"trusted":true,"_uuid":"9ab6ac3ad7db3869ef3f8d5efd7eb70cef070aee"},"cell_type":"code","source":"test_names = [f for f in test.iterdir()]\nholdout_names = [f for f in holdout.iterdir()]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3312c6231afec80d6610a3baa160137702bac43c"},"cell_type":"markdown","source":"We will use the [datablock](https://docs.fast.ai/data_block.html) API from fastai, it is so elegant! We create an ImageItemList to hold our data, using the `.from_df()`method. \nThen we split the data in train/valid sets, I will use 0.2 (20% of data for validation) and a seed=2019, to be able to reproduce my results.\nFinally we add the test set, nice trick we just sum the lists to get the whole."},{"metadata":{"trusted":true,"_uuid":"f1d53127afb926e75f516ee9ae91c839b7fd4bc1"},"cell_type":"code","source":"src = (ImageItemList.from_df(df, path, folder=train)\n      .random_split_by_pct(0.2, seed=2019)\n      .label_from_df('has_oilpalm')\n      .add_test(test_names+holdout_names))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a48b13a4c257640f82dc3c70e0c126303a20de5"},"cell_type":"markdown","source":"We have to add some data augmentation, `get_transforms()` get us a basic set of data augmentations. And we set size=128 to train faster, you can test with 256, but have to reduce the batch size (bs=16). We have to normalize (substract the mean, divide by std) data for the GPU to work better, as I have not computed the actual value, I will just use ImageNet stats."},{"metadata":{"trusted":true,"_uuid":"1d34ec0a6ae4d57d4448c4bc6b49017600634e2d"},"cell_type":"code","source":"data =  (src.transform(get_transforms(), size=128)\n         .databunch(bs=64)\n         .normalize(imagenet_stats))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10feb57993599bedf4cd6c0cc314b771de197277"},"cell_type":"markdown","source":"Loot at data?"},{"metadata":{"trusted":true,"_uuid":"a228a5ceb9da8da540c022d8897e712ec054517c"},"cell_type":"code","source":"data.show_batch(3, figsize=(10,7))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bdd86f9a9f2154af5ba979ce9e08377266507c1f"},"cell_type":"markdown","source":"Let's impement the competition metric, luckyly it is already implemented in sklearn.  We have to modify it a little bit, \n- First, fastai expects a pair (preds, targets) and sklearn expects (targets, preds)\n- Secondly, sklearn needs to vectors of equal shape. For our case, `preds` has shape (bs, 2), so we take the second column, the one that contains the probabilities of palmoil"},{"metadata":{"trusted":true,"_uuid":"ca5040d788a87ad00fe3d347c41b49c008a0a099"},"cell_type":"code","source":"#This was working perfectly some minutes ago!\nfrom sklearn.metrics import roc_auc_score\ndef auc_score(preds,targets):\n    return torch.tensor(roc_auc_score(targets,preds[:,1]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6fb37cbe537c79666beb745acfc2441e038b31f"},"cell_type":"markdown","source":"For some extrange reason thie metric does not always work as a callback for the learner."},{"metadata":{"trusted":true,"_uuid":"6482e7f5a56cdf1b22b8a5f6c703dc80ca001a28"},"cell_type":"code","source":"learn = create_cnn(data, models.resnet18, \n                   metrics=[accuracy], #<---add aoc metric?\n                   model_dir='/kaggle/working/models')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61f76b9e4d35f460e8aa173b3e258eb915383dd2"},"cell_type":"markdown","source":"## Train\nFirst you have to compute the learning rate and choose the one where it is steeper."},{"metadata":{"trusted":true,"_uuid":"9af7bca13332ef1f816eb9c9296dc321ed83faa8"},"cell_type":"code","source":"learn.lr_find(); learn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0011a30f45cbcb9bce31617064a034987cf2e6fa"},"cell_type":"code","source":"lr = 1e-2","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"053ce352a04f078d3aa593cc066ca552a26d57b0"},"cell_type":"markdown","source":"We will use `fit_one_cycle` to train the model, because it is awesome. First we train only the head of the model."},{"metadata":{"trusted":true,"_uuid":"9d1b226d5ced1baf956cfa404076266666597314"},"cell_type":"code","source":"learn.fit_one_cycle(6, lr)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5bcd4ba876d0b695050fe6b64d1b28b796522fc"},"cell_type":"code","source":"Then we unfreeze and train the whole model, with lower lr.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f6bae4dec981333d440d241c19bc8464859a20b7"},"cell_type":"code","source":"learn.unfreeze()\nlearn.fit_one_cycle(3, slice(1e-4, 1e-3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3d86614ee0f817a5713c0bf0d1bcc190b3cf1513"},"cell_type":"code","source":"learn.save('128')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9452b855a96df4f9bfe1189af29fd68cd7119370"},"cell_type":"markdown","source":"Let's compute our AUC over the validation set."},{"metadata":{"trusted":true,"_uuid":"93d398d2a59d6b8e7e52f08a0514623b37e66235"},"cell_type":"code","source":"p,t = learn.get_preds()\nauc_score(p,t)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7e663da6d363a394074d18e79bc1bca63d2fe043"},"cell_type":"markdown","source":"# View results\n\nWe can reviews our model, to see what it did worng. Probably some of this images even a human has a hard time evaluating."},{"metadata":{"trusted":true,"_uuid":"4093211dbb7f68d80164fc4f67f5e4e192200f4d"},"cell_type":"code","source":"interp = ClassificationInterpretation.from_learner(learn)\nlosses,idxs = interp.top_losses()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab0854cf2c83d9f1f4a98bcdbf5c24919ff0c4df"},"cell_type":"code","source":"interp.plot_top_losses(9, figsize=(15,11))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6508773ab6eab1b8bd4ff04308fefa04719b1cd"},"cell_type":"markdown","source":"## Sub file\nWe have to create our sub file by concatenating both holdout and test names."},{"metadata":{"trusted":true,"_uuid":"8750fbb41ce791d556ede44816d9ec63affbad02"},"cell_type":"code","source":"p,t = learn.get_preds(ds_type=DatasetType.Test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b943557415ea7481b4492f3ac1888644c3a602c9"},"cell_type":"code","source":"p = to_np(p); p.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"92a06ece33cd4cab4f227b10611aed88cab5f517"},"cell_type":"code","source":"ids = np.array([f.name for f in (test_names+holdout_names)]);ids.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f48a774a9311109f87dfd96adad723b25092a1f0"},"cell_type":"code","source":"#We only recover the probs of having palmoil (column 1)\nsub = pd.DataFrame(np.stack([ids, p[:,1]], axis=1), columns=df_sample.columns)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"67ad14edac9bfc39c3bf9b27e24aaf94bbfdf48b"},"cell_type":"code","source":"sub.to_csv(work_dir/'sub.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ee624be0137bb57cd461ba3d1861d2cc1e816ad8"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}