{"cells":[{"metadata":{"_uuid":"c45eead470fae7ee5654fc70609bab7ca7f0fa9d"},"cell_type":"markdown","source":"## TL;DR\n\nThis is a simple classifier based on an Imagenet-trained Resnet18. It's inspired by Iafoss[](https://www.kaggle.com/iafoss)'s kernels in other competititions.\n\n### Update\n\nJust realized Densnet121 is way better, so let's switch to that for now!"},{"metadata":{"_uuid":"88e20188ddffce29fc4c3af7b6fc09bb48cc1fee"},"cell_type":"markdown","source":"## Training "},{"metadata":{"_uuid":"132b11ca41afe41627ed3c0df8b2be39d30f93d2"},"cell_type":"markdown","source":"Let's start by importing our libararies."},{"metadata":{"trusted":true,"_uuid":"9c7c08637a08960677f15998e6e579ab43ce05b9"},"cell_type":"code","source":"!pip install fastai==0.7.0 --no-deps\n!pip install torch==0.4.1 torchvision==0.2.1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from fastai.conv_learner import *\nfrom fastai.dataset import *\n\nimport pandas as pd\nimport numpy as np\nimport os\nfrom sklearn.model_selection import train_test_split\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"MODEL_PATH = 'Dn121_v1'\nTRAIN = '../input/train/'\nTEST = '../input/test/'\nLABELS = '../input/train_labels.csv'\nSAMPLE_SUB = '../input/sample_submission.csv'\nORG_SIZE=96\nBATCH_SIZE = 128","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e0624ab350e370dbff80cac45f33744c48e5633b"},"cell_type":"markdown","source":"The architecture is flexible, I chose Resnet18 since it can fit quite well into a kernel. You may play with this if you want to. "},{"metadata":{"trusted":true,"_uuid":"6ea9033e0200d3d9142b4ee05c45c1dd4f2d8c1d"},"cell_type":"code","source":"arch = dn121 \nnw = 4","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bf2a5c5342e855974efaeb7fe5c2b90f2cf636cf"},"cell_type":"markdown","source":"Next, we prapare out dataset to work with Fastai's pipeline."},{"metadata":{"trusted":true,"_uuid":"d9adfc15b56c7f80f291c66dc6d6f38d4d55e6a2"},"cell_type":"code","source":"train_df = pd.read_csv(LABELS).set_index('id')\ntrain_names = train_df.index.values\ntrain_labels = np.asarray(train_df['label'].values)\nprint(\"Number of positive samples = {:.4f}%\".format(np.count_nonzero(train_labels)*100/len(train_labels)))\ntest_names = [f.replace(\".tif\",\"\") for f in os.listdir(TEST)]\ntr_n, val_n = train_test_split(train_names, test_size=0.15, random_state=42069)\nprint(len(tr_n), len(val_n))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"94af91d70819db979d39a4d77b2e30493498978b"},"cell_type":"code","source":"class HCDDataset(FilesDataset):\n    def __init__(self, fnames, path, transform):\n        self.train_df = train_df\n        super().__init__(fnames, transform, path)\n\n    def get_x(self, i):\n        img = open_image(os.path.join(self.path, self.fnames[i]+\".tif\"))\n        # We crop the center of the original image for faster training time\n        img = img[(ORG_SIZE-self.sz)//2:(ORG_SIZE+self.sz)//2,(ORG_SIZE-self.sz)//2:(ORG_SIZE+self.sz)//2,:]\n        return img\n\n    def get_y(self, i):\n        if (self.path == TEST): return 0\n        return self.train_df.loc[self.fnames[i]]['label']\n\n\n    def get_c(self):\n        return 2\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"140dfae2b41cbe4f770f8d80fcaba0ebc772e983"},"cell_type":"code","source":"def get_data(sz, bs):\n    aug_tfms = [RandomRotate(20, tfm_y=TfmType.NO),\n                RandomDihedral(tfm_y=TfmType.NO)]\n    tfms = tfms_from_model(arch, sz, crop_type=CropType.NO, tfm_y=TfmType.NO,\n                           aug_tfms=aug_tfms)\n    ds = ImageData.get_ds(HCDDataset, (tr_n[:-(len(tr_n) % bs)], TRAIN),\n                          (val_n, TRAIN), tfms, test=(test_names, TEST))\n    md = ImageData(\"./\", ds, bs, num_workers=nw, classes=None)\n    return md\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f8258255beb8fb608abb8a292b07c7161580007e"},"cell_type":"code","source":"md = get_data(96, BATCH_SIZE)\nlearn = ConvLearner.pretrained(arch, md) \nlearn.opt_fn = optim.Adam","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c2a3accea73d13e6b8f81febd988b9e377fb5572"},"cell_type":"markdown","source":"Uncomment these lines to run Fastai's automatic learning rate finder. \n\nThe graphs showed that the loss converged at around learning rate = 1e-2, which means we should set our learning rate a bit higher than that."},{"metadata":{"trusted":true,"_uuid":"04b0332bd91ee3752b8da857d34e566c96a638d4"},"cell_type":"code","source":"# learn.lr_find()\n# learn.sched.plot()\nlr = 2e-2","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"115459f2b3756d4f029cb223a57a80abba5f2992"},"cell_type":"markdown","source":"We start by training only the newly initialized weights, then unfreeze the model and finetune the pretrained weights with reduced learning rate."},{"metadata":{"trusted":true,"_uuid":"f3118d5e2dbe61c8d51d0e33642ea5bb0b516a54"},"cell_type":"code","source":"learn.fit(lr, 1, cycle_len=2)\nlearn.unfreeze()\nlrs = np.array([1e-4, 5e-4, 1.2e-3])\nlearn.fit(lrs, 1, cycle_len=5, use_clr=(20, 16))\nlearn.fit(lrs/4, 1, cycle_len=5, use_clr=(10, 8))\nlearn.fit(lrs/16, 1, cycle_len=5, use_clr=(10, 8))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eabe555ace20e3b36c6266432108d31de1e58282"},"cell_type":"markdown","source":"## Predictions\nLet's add some TTA (Test-Time-Augmentation), this usually increases the accuracy a bit, you may need to do some tests here to compare the result with normal prediction."},{"metadata":{"trusted":true,"_uuid":"6cbfaedbad6bac01b06d87eaf3723dd260b7a51e"},"cell_type":"code","source":"# preds_t,y_t = learn.predict_with_targs(is_test=True) # Predicting without TTA\npreds_t,y_t = learn.TTA(is_test=True, n_aug=8)\npreds_t = np.stack(preds_t, axis=-1)\npreds_t = np.exp(preds_t)\npreds_t = preds_t.mean(axis=-1)[:,1]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"82b4e79d2a05267790200de4860fd25a0a669f7f"},"cell_type":"markdown","source":"Finally, our submission."},{"metadata":{"trusted":true,"_uuid":"f5fbd91e970d375debc3270ebd5b08bb41eeb66e"},"cell_type":"code","source":"sample_df = pd.read_csv(SAMPLE_SUB)\nsample_list = list(sample_df.id)\npred_list = [p for p in preds_t]\npred_dic = dict((key, value) for (key, value) in zip(learn.data.test_ds.fnames,pred_list))\npred_list_cor = [pred_dic[id] for id in sample_list]\ndf = pd.DataFrame({'id':sample_list,'label':pred_list_cor})\ndf.to_csv('submission.csv'.format(MODEL_PATH), header=True, index=False)\n","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}