{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RANZCR 1st Place Solution Training Classification Model\n\nHi all,\n\nWe're very exciting to writing this notebook and the summary of our solution here.\n\nOur final pipeline has 4 training stages but the minimal pipeline I show here has only 2 stages.\n\nThe 5-fold model trained with this minimal pipeline is sufficient to achieve CV 0.968-0.969 and pub/pvt LB 0.972\n\nI published 3 notebooks to demonstrate how our MINIMAL pipeline works.\n\n* Stage1: Segmentation (https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-seg-model-small-ver)\n* Stage2: Classification (This notebook)\n* Inference (https://www.kaggle.com/haqishen/ranzcr-1st-place-soluiton-inference-small-ver)\n\nThis notebook shows how we can train a **classification** model using the original image, and the predicted mask of the segmentation model which we trained in stage 1.\n\nI've trained 5 fold segmentation models (b1 1024, exactly the same setting to the stage1 notebook） locally and uploaded the predicted masks here: https://www.kaggle.com/haqishen/ranzcr-pred-masks-qishen\n\nTo get a similar performance of a single classification model to ours, you only need to do:\n\n* set DEBUG to `False`\n* train for 30 epochs for each fold\n* change the model type to effnet-b3/b4/b5\n\nIn this minimal pipeline we do not use external data, so the classification model input is **5ch**, original image as 3ch, predicted masks as 2ch (w/o trachea bifurcation channel)\n\nIt tooks around 3h for training a single fold classification models by this setting (b1 512)\n\nFinally I got CV score: `0.96788` by B1 1024 segmentationg model + B1 512 classification model\n\nYou should be able to reach CV 0.968-0.969 by simply using B5-B7 backbone to train segmentation model and use B3-B5 backbone to train classification model, without changing any other setting.\n\nOur brief summary of winning solution: https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226633\n\n# Main Ideas\n\n* amp\n* combine original image (3ch) and predicted mask (3ch) as 5ch input\n* optimized augmentation methods\n* use CrossEntropy for ETT tube\n* warmup + cosine scheduler\n\n# Thanks!","metadata":{}},{"cell_type":"code","source":"!pip -q install git+https://github.com/ildoonet/pytorch-gradual-warmup-lr.git","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:35:55.786907Z","iopub.execute_input":"2021-08-18T12:35:55.787289Z","iopub.status.idle":"2021-08-18T12:36:05.310996Z","shell.execute_reply.started":"2021-08-18T12:35:55.787207Z","shell.execute_reply":"2021-08-18T12:36:05.309847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DEBUG = True\n\n# Add smp & timm to kernel w/o Internet\nimport sys\nsys.path = [\n    '../input/smp20210127/segmentation_models.pytorch-master/segmentation_models.pytorch-master/',\n    '../input/smp20210127/EfficientNet-PyTorch-master/EfficientNet-PyTorch-master',\n    '../input/smp20210127/pytorch-image-models-master/pytorch-image-models-master',\n    '../input/smp20210127/pretrained-models.pytorch-master/pretrained-models.pytorch-master',\n] + sys.path","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:36:31.687427Z","iopub.execute_input":"2021-08-18T12:36:31.687753Z","iopub.status.idle":"2021-08-18T12:36:31.692239Z","shell.execute_reply.started":"2021-08-18T12:36:31.687721Z","shell.execute_reply":"2021-08-18T12:36:31.691411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# libraries\nimport os\nimport time\nimport random\nimport numpy as np\nimport pandas as pd\nimport subprocess\nimport cv2\nimport PIL.Image\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nimport torch\nfrom torch.utils.data import DataLoader, Dataset\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torchvision.transforms as transforms\nimport torch.optim as optim\nfrom torch.optim import lr_scheduler\nfrom torch.optim.lr_scheduler import CosineAnnealingLR\nfrom sklearn.metrics import roc_auc_score\nfrom warmup_scheduler import GradualWarmupScheduler\nimport albumentations\nimport timm\nfrom tqdm.notebook import tqdm\nimport torch.cuda.amp as amp\nimport warnings\n\nwarnings.simplefilter('ignore')\nscaler = amp.GradScaler()\ndevice = torch.device('cuda')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-08-18T12:36:35.153753Z","iopub.execute_input":"2021-08-18T12:36:35.154063Z","iopub.status.idle":"2021-08-18T12:36:39.450109Z","shell.execute_reply.started":"2021-08-18T12:36:35.154034Z","shell.execute_reply":"2021-08-18T12:36:39.449128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dual Cutout implementations\nclass CutoutV2(albumentations.DualTransform):\n    def __init__(\n        self,\n        num_holes=8, #Number of patches to cut out of each image.\n        max_h_size=8,#hole最大高度\n        max_w_size=8,#hole最大寬度\n        fill_value=0,#丟棄像素的值。\n        always_apply=False,\n        p=0.5,\n    ):\n        super(CutoutV2, self).__init__(always_apply, p)\n        self.num_holes = num_holes\n        self.max_h_size = max_h_size\n        self.max_w_size = max_w_size\n        self.fill_value = fill_value\n\n    def apply(self, image, fill_value=0, holes=(), **params):\n        #functional.cutout：CoarseDropout of the square regions in the image.\n        #Coarse dropout 是一種數據增強技術，可防止您的模型過度擬合。從訓練圖像中隨機刪除矩形或是刪除許多相似大小的小矩形的技術。\n        return albumentations.functional.cutout(image, holes, fill_value) #Cutout (image, 要歸零的區域數量, 丟棄像素的值)\n\n    def get_params_dependent_on_targets(self, params):\n        img = params[\"image\"]\n        height, width = img.shape[:2]\n\n        holes = []\n        for _n in range(self.num_holes):\n            y = random.randint(0, height)#產生指定範圍內的隨機整數\n            x = random.randint(0, width)\n\n            y1 = np.clip(y - self.max_h_size // 2, 0, height)#clip截取，超出的部分就把它強置為邊界部分\n            y2 = np.clip(y1 + self.max_h_size, 0, height)\n            x1 = np.clip(x - self.max_w_size // 2, 0, width)\n            x2 = np.clip(x1 + self.max_w_size, 0, width)\n            holes.append((x1, y1, x2, y2))\n\n        return {\"holes\": holes}\n\n    @property\n    def targets_as_params(self):\n        return [\"image\"]\n\n    def get_transform_init_args_names(self):\n        return (\"num_holes\", \"max_h_size\", \"max_w_size\")","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:36:47.384099Z","iopub.execute_input":"2021-08-18T12:36:47.384468Z","iopub.status.idle":"2021-08-18T12:36:47.396171Z","shell.execute_reply.started":"2021-08-18T12:36:47.384437Z","shell.execute_reply":"2021-08-18T12:36:47.395278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Config","metadata":{}},{"cell_type":"code","source":"kernel_type = 'enetb1_3ch_512_lr3e4_bs32_30epo'\nenet_type = 'tf_efficientnet_b1_ns'\ndata_dir = '../input/image1000090'\nnum_workers = 2\nnum_classes = 12\nn_ch = 3 #原n_ch = 5\nimage_size = 512\nbatch_size = 32\ninit_lr = 3e-4\nwarmup_epo = 1\n# If DEBUG == True, only run 3 epochs per fold\ncosine_epo = 10 if not DEBUG else 2\nn_epochs = warmup_epo + cosine_epo\nloss_weights = [1., 9.]\nimage_folder = 'train'\n#mask_folder = '../input/ranzcr-pred-masks-qishen/mask_unet++b1_2cbce_1024T15tip_lr1e4_bs4_augv2_30epo/'\n\nlog_dir = './logs'\nmodel_dir = './models'\nos.makedirs(log_dir, exist_ok=True)\nos.makedirs(model_dir, exist_ok=True)\nlog_file = os.path.join(log_dir, f'log_{kernel_type}.txt')","metadata":{"execution":{"iopub.status.busy":"2021-08-18T13:08:42.665279Z","iopub.execute_input":"2021-08-18T13:08:42.665623Z","iopub.status.idle":"2021-08-18T13:08:42.673741Z","shell.execute_reply.started":"2021-08-18T13:08:42.665594Z","shell.execute_reply":"2021-08-18T13:08:42.672915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv('../input/image1000090/train5000.csv')\n\n# If DEBUG == True, use only 100 samples to run.\ndf_train = pd.concat([\n    df_train.query('fold == 0').sample(1000),\n    df_train.query('fold == 1').sample(1000),\n    df_train.query('fold == 2').sample(1000),\n    df_train.query('fold == 3').sample(1000),\n    df_train.query('fold == 4').sample(1000),\n]) if DEBUG else df_train\n\nno_ETT = (df_train[['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal']].values.max(1) == 0).astype(int)\ndf_train.insert(4, column='no_ETT', value=no_ETT)\ndf_train.shape\n","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:47:34.586899Z","iopub.execute_input":"2021-08-18T12:47:34.587262Z","iopub.status.idle":"2021-08-18T12:47:34.662981Z","shell.execute_reply.started":"2021-08-18T12:47:34.587230Z","shell.execute_reply":"2021-08-18T12:47:34.662167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define Dataset","metadata":{}},{"cell_type":"code","source":"class RANZCRDatasetCLS(Dataset):\n\n    def __init__(self, df, mode, transform=None):\n\n        self.df = df.reset_index(drop=True)\n        self.mode = mode\n        self.transform = transform\n\n    def __len__(self):\n        return self.df.shape[0]\n\n    def __getitem__(self, index):\n        row = self.df.iloc[index]\n        image = cv2.imread(os.path.join(data_dir, image_folder, row.StudyInstanceUID + '.jpg'))[:, :, ::-1]\n        #mask = cv2.imread(os.path.join(mask_folder, row.StudyInstanceUID + '.png')).astype(np.float32)[:,:,:2]\n\n        res = self.transform(image=image) #(image=image, mask=mask)\n        image = res['image'].astype(np.float32).transpose(2, 0, 1) / 255.\n       # mask = res['mask'].astype(np.float32).transpose(2, 0, 1) / 255.\n\n        #image = np.concatenate([image, mask], 0)\n\n        if self.mode == 'test':\n            return torch.tensor(image)\n        else:\n            label = row[[\n                'ETT - Abnormal',\n                'ETT - Borderline',\n                'ETT - Normal',\n                'no_ETT',\n                'NGT - Abnormal',\n                'NGT - Borderline',\n                'NGT - Incompletely Imaged',\n                'NGT - Normal',\n                'CVC - Abnormal',\n                'CVC - Borderline',\n                'CVC - Normal',\n                'Swan Ganz Catheter Present'\n            ]].values.astype(float)\n            return torch.tensor(image).float(), torch.tensor(label).float()","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:47:48.514058Z","iopub.execute_input":"2021-08-18T12:47:48.514389Z","iopub.status.idle":"2021-08-18T12:47:48.524261Z","shell.execute_reply.started":"2021-08-18T12:47:48.514360Z","shell.execute_reply":"2021-08-18T12:47:48.523082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Augmentations","metadata":{}},{"cell_type":"code","source":"transforms_train = albumentations.Compose([\n    albumentations.Resize(image_size, image_size),#將輸入調整為給定的高度和寬度。\n])\ntransforms_val = albumentations.Compose([\n    albumentations.Resize(image_size, image_size),\n])","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:47:51.557130Z","iopub.execute_input":"2021-08-18T12:47:51.557489Z","iopub.status.idle":"2021-08-18T12:47:51.562079Z","shell.execute_reply.started":"2021-08-18T12:47:51.557457Z","shell.execute_reply":"2021-08-18T12:47:51.561264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization","metadata":{}},{"cell_type":"code","source":"df_show = df_train.sample(10)\ndataset_show = RANZCRDatasetCLS(df_show, 'train', transform=transforms_train)\nfrom pylab import rcParams\nrcParams['figure.figsize'] = 20,10\n\nf, axarr = plt.subplots(1,5)\nimgs = []\nfor p in range(5):\n    img, label = dataset_show[p]\n    imgs.append(img)\n    axarr[p].imshow(img[:3].transpose(0, 1).transpose(1,2))\n\nf, axarr = plt.subplots(1,5)\nfor p in range(5):\n    axarr[p].imshow(imgs[p][3:].sum(0))","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:47:54.666843Z","iopub.execute_input":"2021-08-18T12:47:54.667179Z","iopub.status.idle":"2021-08-18T12:47:55.118699Z","shell.execute_reply.started":"2021-08-18T12:47:54.667131Z","shell.execute_reply":"2021-08-18T12:47:55.117261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class enetv2(nn.Module):\n    def __init__(self, enet_type, out_dim):\n        super(enetv2, self).__init__()\n        self.enet = timm.create_model(enet_type, True)\n        self.dropout = nn.Dropout(0.5)\n        self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1,n_ch//3+1,1,1)[:, :n_ch])\n        self.myfc = nn.Linear(self.enet.classifier.in_features, out_dim)\n        self.enet.classifier = nn.Identity()\n\n    def extract(self, x):\n        return self.enet(x)\n\n    def forward(self, x):\n        x = self.extract(x)\n        h = self.myfc(self.dropout(x))\n        return h\n\nm = enetv2(enet_type, num_classes)\nm(torch.rand(2,n_ch,image_size, image_size))","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:50:21.925033Z","iopub.execute_input":"2021-08-18T12:50:21.925445Z","iopub.status.idle":"2021-08-18T12:50:25.245118Z","shell.execute_reply.started":"2021-08-18T12:50:21.925410Z","shell.execute_reply":"2021-08-18T12:50:25.244247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loss","metadata":{}},{"cell_type":"code","source":"bce = nn.BCEWithLogitsLoss()\nce = nn.CrossEntropyLoss()\n\n\ndef criterion(logits, targets, lw=loss_weights):\n    loss1 = ce(logits[:, :4], targets[:, :4].argmax(1)) * lw[0]\n    loss2 = bce(logits[:, 4:], targets[:, 4:]) * lw[1]\n    return (loss1 + loss2) / sum(lw)","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:50:29.446939Z","iopub.execute_input":"2021-08-18T12:50:29.447278Z","iopub.status.idle":"2021-08-18T12:50:29.453917Z","shell.execute_reply.started":"2021-08-18T12:50:29.447248Z","shell.execute_reply":"2021-08-18T12:50:29.452976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LR Scheduler","metadata":{}},{"cell_type":"code","source":"class GradualWarmupSchedulerV2(GradualWarmupScheduler):\n    def __init__(self, optimizer, multiplier, total_epoch, after_scheduler=None):\n        super(GradualWarmupSchedulerV2, self).__init__(optimizer, multiplier, total_epoch, after_scheduler)\n    def get_lr(self):\n        if self.last_epoch > self.total_epoch:\n            if self.after_scheduler:\n                if not self.finished:\n                    self.after_scheduler.base_lrs = [base_lr * self.multiplier for base_lr in self.base_lrs]\n                    self.finished = True\n                return self.after_scheduler.get_lr()\n            return [base_lr * self.multiplier for base_lr in self.base_lrs]\n        if self.multiplier == 1.0:\n            return [base_lr * (float(self.last_epoch) / self.total_epoch) for base_lr in self.base_lrs]\n        else:\n            return [base_lr * ((self.multiplier - 1.) * self.last_epoch / self.total_epoch + 1.) for base_lr in self.base_lrs]\n\noptimizer = optim.Adam(m.parameters(), lr=init_lr)\nscheduler_cosine = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, cosine_epo)\nscheduler_warmup = GradualWarmupSchedulerV2(optimizer, multiplier=10, total_epoch=warmup_epo, after_scheduler=scheduler_cosine)\nlrs = []\nfor epoch in range(1, n_epochs+1):\n    scheduler_warmup.step(epoch-1)\n    lrs.append(optimizer.param_groups[0][\"lr\"])\nrcParams['figure.figsize'] = 20,3\nplt.plot(lrs)","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:50:33.244481Z","iopub.execute_input":"2021-08-18T12:50:33.244825Z","iopub.status.idle":"2021-08-18T12:50:33.404148Z","shell.execute_reply.started":"2021-08-18T12:50:33.244797Z","shell.execute_reply":"2021-08-18T12:50:33.403527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train & Validation Function","metadata":{}},{"cell_type":"code","source":"def train_epoch(model, loader, optimizer):\n\n    model.train()\n    train_loss = []\n    bar = tqdm(loader)\n    for (data, targets) in bar:\n\n        optimizer.zero_grad()\n        data, targets = data.to(device), targets.to(device)\n\n        with amp.autocast():\n            logits = model(data)\n            loss = criterion(logits, targets)\n        scaler.scale(loss).backward() \n        scaler.step(optimizer)\n        scaler.update()\n\n        loss_np = loss.item()\n        train_loss.append(loss_np)\n        smooth_loss = sum(train_loss[-50:]) / min(len(train_loss), 50)\n        bar.set_description('loss: %.4f, smth: %.4f' % (loss_np, smooth_loss))\n\n    return np.mean(train_loss)\n\n\ndef valid_epoch(model, loader, get_output=False):\n\n    model.eval()\n    val_loss = []\n    LOGITS = []\n    TARGETS = []\n    with torch.no_grad():\n        for (data, targets) in tqdm(loader):\n            data, targets = data.to(device), targets.to(device)\n            logits = model(data)\n            loss = criterion(logits, targets)\n            val_loss.append(loss.item())\n            LOGITS.append(logits.cpu())\n            TARGETS.append(targets.cpu())\n            \n    val_loss = np.mean(val_loss)\n    LOGITS = torch.cat(LOGITS)\n    LOGITS[:, :4] = LOGITS[:, :4].softmax(1)\n    LOGITS[:, 4:] = LOGITS[:, 4:].sigmoid()\n    TARGETS = torch.cat(TARGETS).numpy()\n\n    if get_output:\n        return LOGITS\n    else:\n        aucs = []\n        for cid in range(num_classes):\n            if cid == 3: continue\n            try:\n                aucs.append( roc_auc_score(TARGETS[:, cid], LOGITS[:, cid]) )\n            except:\n                aucs.append(0.5)\n        return val_loss, aucs","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:50:36.664254Z","iopub.execute_input":"2021-08-18T12:50:36.664581Z","iopub.status.idle":"2021-08-18T12:50:36.807194Z","shell.execute_reply.started":"2021-08-18T12:50:36.664554Z","shell.execute_reply":"2021-08-18T12:50:36.806188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Run!","metadata":{}},{"cell_type":"code","source":"def run(fold):\n    content = 'Fold: ' + str(fold)\n    print(content)\n    with open(log_file, 'a') as appender:\n        appender.write(content + '\\n')\n    train_ = df_train.query(f'fold!={fold}').copy()\n    valid_ = df_train.query(f'fold=={fold}').copy()\n\n    dataset_train = RANZCRDatasetCLS(train_, 'train', transform=transforms_train)\n    dataset_valid = RANZCRDatasetCLS(valid_, 'valid', transform=transforms_val)\n    train_loader = torch.utils.data.DataLoader(dataset_train, batch_size=batch_size, shuffle=True, num_workers=num_workers)\n    valid_loader = torch.utils.data.DataLoader(dataset_valid, batch_size=batch_size, shuffle=False, num_workers=num_workers)\n\n    model = enetv2(enet_type, num_classes)\n    model = model.to(device)\n    aucs_max = 0\n    model_file = os.path.join(model_dir, f'{kernel_type}_best_fold{fold}.pth')\n\n    optimizer = optim.Adam(model.parameters(), lr=init_lr)\n    scheduler_cosine = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, cosine_epo)\n    scheduler_warmup = GradualWarmupSchedulerV2(optimizer, multiplier=10, total_epoch=warmup_epo, after_scheduler=scheduler_cosine)\n    for epoch in range(1, n_epochs+1):\n        print(time.ctime(), 'Epoch:', epoch)\n        scheduler_warmup.step(epoch-1)\n\n        train_loss = train_epoch(model, train_loader, optimizer)\n        val_loss, aucs = valid_epoch(model, valid_loader)\n\n        #optimizer.param_groups[0]：長度為6的字典，包括['amsgrad', 'params', 'lr', 'betas', 'weight_decay', 'eps']這6個參數\n        content = time.ctime() + ' ' + f'Fold {fold} Epoch {epoch}, lr: {optimizer.param_groups[0][\"lr\"]:.7f}, train loss: {train_loss:.4f}, valid loss: {(val_loss):.4f}, aucs: {np.mean(aucs):.4f}.'\n        content += '\\n' + ' '.join([f'{x:.4f}' for x in aucs])\n        print(content)\n        with open(log_file, 'a') as appender:\n            appender.write(content + '\\n')\n\n        if aucs_max < np.mean(aucs):\n            print('aucs increased ({:.6f} --> {:.6f}).  Saving model ...'.format(aucs_max, np.mean(aucs)))\n            torch.save(model.state_dict(), model_file)\n            aucs_max = np.mean(aucs)\n\n    torch.save(model.state_dict(), os.path.join(model_dir, f'{kernel_type}_model_fold{fold}.pth'))\n","metadata":{"execution":{"iopub.status.busy":"2021-08-18T12:50:40.223969Z","iopub.execute_input":"2021-08-18T12:50:40.224304Z","iopub.status.idle":"2021-08-18T12:50:40.237286Z","shell.execute_reply.started":"2021-08-18T12:50:40.224275Z","shell.execute_reply":"2021-08-18T12:50:40.236349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for fold in range(5):\n    run(fold)","metadata":{"execution":{"iopub.status.busy":"2021-08-18T13:08:52.107684Z","iopub.execute_input":"2021-08-18T13:08:52.108239Z","iopub.status.idle":"2021-08-18T13:08:52.669338Z","shell.execute_reply.started":"2021-08-18T13:08:52.108194Z","shell.execute_reply":"2021-08-18T13:08:52.666135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis","metadata":{}},{"cell_type":"code","source":"dfs = []\nLOGITS = []\n\nfor fold in range(5):\n    valid_ = df_train.query(f'fold=={fold}').copy().reset_index(drop=True)\n    dfs.append(valid_)\n    dataset_valid = RANZCRDatasetCLS(valid_, 'valid', transform=transforms_val)\n    valid_loader = torch.utils.data.DataLoader(dataset_valid, batch_size=batch_size, shuffle=False, num_workers=num_workers)\n\n    model = enetv2(enet_type, num_classes)\n    model = model.to(device)\n    model_file = os.path.join(model_dir, f'{kernel_type}_best_fold{fold}.pth')\n    model.load_state_dict(torch.load(model_file), strict=True)\n    model.eval()\n    \n    outputs = valid_epoch(model, valid_loader, get_output=True)\n    # rank prediction\n    outputs = pd.DataFrame(outputs.numpy()).rank(pct=True).values\n    LOGITS.append(outputs)\n\ndfs = pd.concat(dfs).reset_index(drop=True)\nLOGITS = np.concatenate(LOGITS)","metadata":{"execution":{"iopub.status.busy":"2021-07-24T06:32:39.200777Z","iopub.execute_input":"2021-07-24T06:32:39.201166Z","iopub.status.idle":"2021-07-24T06:32:51.957311Z","shell.execute_reply.started":"2021-07-24T06:32:39.201123Z","shell.execute_reply":"2021-07-24T06:32:51.95653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"outputs","metadata":{"execution":{"iopub.status.busy":"2021-07-24T06:32:51.958816Z","iopub.execute_input":"2021-07-24T06:32:51.95921Z","iopub.status.idle":"2021-07-24T06:32:51.972958Z","shell.execute_reply.started":"2021-07-24T06:32:51.959178Z","shell.execute_reply":"2021-07-24T06:32:51.968468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfs","metadata":{"execution":{"iopub.status.busy":"2021-07-24T06:32:51.975639Z","iopub.execute_input":"2021-07-24T06:32:51.976126Z","iopub.status.idle":"2021-07-24T06:32:52.00757Z","shell.execute_reply.started":"2021-07-24T06:32:51.976079Z","shell.execute_reply":"2021-07-24T06:32:52.006674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LOGITS","metadata":{"execution":{"iopub.status.busy":"2021-07-24T06:32:52.009014Z","iopub.execute_input":"2021-07-24T06:32:52.009381Z","iopub.status.idle":"2021-07-24T06:32:52.015584Z","shell.execute_reply.started":"2021-07-24T06:32:52.009333Z","shell.execute_reply":"2021-07-24T06:32:52.014722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aucs = []\nfor cid in range(num_classes):\n    if cid == 3: continue\n    aucs.append( roc_auc_score(dfs.values[:,1:-3].astype(float)[:, cid], LOGITS[:, cid]) )\nprint('oof cv:', np.mean(aucs))","metadata":{"execution":{"iopub.status.busy":"2021-07-24T06:32:52.017735Z","iopub.execute_input":"2021-07-24T06:32:52.01814Z","iopub.status.idle":"2021-07-24T06:32:52.026396Z","shell.execute_reply.started":"2021-07-24T06:32:52.018098Z","shell.execute_reply":"2021-07-24T06:32:52.024909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Appendix: My Local Training Log Fold 0\n\nYou should get similar log to mine only by changing DEBUG = False and train on your local machine with >= 16GB GPU.\n\n```\nThu Mar 18 09:20:58 2021 Epoch 1, lr: 0.0003000, train loss: 0.2412, valid loss: 0.1794, aucs: 0.9323.\n0.9358 0.9338 0.9857 0.9509 0.9435 0.9768 0.9804 0.8689 0.7919 0.8894 0.9977\nThu Mar 18 09:26:52 2021 Epoch 2, lr: 0.0030000, train loss: 0.2339, valid loss: 0.1674, aucs: 0.9339.\n0.9106 0.9435 0.9873 0.9349 0.9305 0.9754 0.9844 0.8759 0.8265 0.9060 0.9974\nThu Mar 18 09:32:48 2021 Epoch 3, lr: 0.0030000, train loss: 0.2071, valid loss: 0.1754, aucs: 0.9421.\n0.9118 0.9458 0.9906 0.9604 0.9606 0.9785 0.9839 0.8981 0.8351 0.9004 0.9976\nThu Mar 18 09:38:35 2021 Epoch 4, lr: 0.0029649, train loss: 0.1994, valid loss: 0.1714, aucs: 0.9415.\n0.9168 0.9501 0.9893 0.9509 0.9549 0.9776 0.9819 0.8849 0.8417 0.9091 0.9991\nThu Mar 18 09:44:21 2021 Epoch 5, lr: 0.0029215, train loss: 0.1945, valid loss: 0.1608, aucs: 0.9445.\n0.9219 0.9558 0.9897 0.9632 0.9650 0.9791 0.9847 0.8886 0.8347 0.9075 0.9994\nThu Mar 18 09:49:59 2021 Epoch 6, lr: 0.0028614, train loss: 0.1902, valid loss: 0.1552, aucs: 0.9499.\n0.9485 0.9469 0.9877 0.9765 0.9642 0.9821 0.9865 0.8922 0.8544 0.9126 0.9975\nThu Mar 18 09:55:53 2021 Epoch 7, lr: 0.0027853, train loss: 0.1856, valid loss: 0.1705, aucs: 0.9474.\n0.9260 0.9523 0.9895 0.9654 0.9659 0.9827 0.9861 0.8923 0.8437 0.9187 0.9990\nThu Mar 18 10:01:41 2021 Epoch 8, lr: 0.0026941, train loss: 0.1819, valid loss: 0.1585, aucs: 0.9476.\n0.9162 0.9528 0.9908 0.9620 0.9660 0.9830 0.9870 0.9027 0.8516 0.9124 0.9993\nThu Mar 18 10:07:21 2021 Epoch 9, lr: 0.0025890, train loss: 0.1802, valid loss: 0.1640, aucs: 0.9456.\n0.9127 0.9581 0.9914 0.9463 0.9445 0.9812 0.9846 0.9077 0.8576 0.9183 0.9990\nThu Mar 18 10:13:02 2021 Epoch 10, lr: 0.0024711, train loss: 0.1775, valid loss: 0.1656, aucs: 0.9481.\n0.9393 0.9594 0.9914 0.9624 0.9631 0.9802 0.9823 0.9023 0.8401 0.9100 0.9990\nThu Mar 18 10:18:39 2021 Epoch 11, lr: 0.0023418, train loss: 0.1742, valid loss: 0.1444, aucs: 0.9550.\n0.9276 0.9625 0.9919 0.9721 0.9718 0.9851 0.9878 0.9105 0.8741 0.9223 0.9993\nThu Mar 18 10:24:15 2021 Epoch 12, lr: 0.0022026, train loss: 0.1721, valid loss: 0.1580, aucs: 0.9554.\n0.9355 0.9607 0.9922 0.9783 0.9654 0.9851 0.9891 0.9086 0.8714 0.9246 0.9986\nThu Mar 18 10:29:58 2021 Epoch 13, lr: 0.0020552, train loss: 0.1699, valid loss: 0.1447, aucs: 0.9575.\n0.9547 0.9634 0.9921 0.9700 0.9711 0.9853 0.9881 0.9151 0.8711 0.9225 0.9991\nThu Mar 18 10:35:36 2021 Epoch 14, lr: 0.0019013, train loss: 0.1673, valid loss: 0.1449, aucs: 0.9544.\n0.9221 0.9605 0.9916 0.9637 0.9736 0.9841 0.9873 0.9104 0.8792 0.9274 0.9988\nThu Mar 18 10:41:14 2021 Epoch 15, lr: 0.0017427, train loss: 0.1641, valid loss: 0.1396, aucs: 0.9613.\n0.9653 0.9650 0.9928 0.9790 0.9738 0.9852 0.9882 0.9180 0.8807 0.9276 0.9986\nThu Mar 18 10:47:00 2021 Epoch 16, lr: 0.0015812, train loss: 0.1610, valid loss: 0.1421, aucs: 0.9632.\n0.9753 0.9684 0.9934 0.9773 0.9756 0.9864 0.9884 0.9183 0.8854 0.9289 0.9982\nThu Mar 18 10:52:53 2021 Epoch 17, lr: 0.0014188, train loss: 0.1590, valid loss: 0.1362, aucs: 0.9612.\n0.9619 0.9637 0.9929 0.9812 0.9681 0.9857 0.9890 0.9184 0.8844 0.9312 0.9966\nThu Mar 18 10:58:42 2021 Epoch 18, lr: 0.0012573, train loss: 0.1571, valid loss: 0.1364, aucs: 0.9627.\n0.9684 0.9652 0.9931 0.9819 0.9760 0.9871 0.9901 0.9184 0.8823 0.9293 0.9983\nThu Mar 18 11:04:34 2021 Epoch 19, lr: 0.0010987, train loss: 0.1545, valid loss: 0.1371, aucs: 0.9630.\n0.9751 0.9650 0.9928 0.9808 0.9750 0.9870 0.9892 0.9165 0.8869 0.9287 0.9963\nThu Mar 18 11:10:28 2021 Epoch 20, lr: 0.0009448, train loss: 0.1524, valid loss: 0.1357, aucs: 0.9626.\n0.9702 0.9667 0.9933 0.9781 0.9764 0.9867 0.9884 0.9176 0.8831 0.9309 0.9973\nThu Mar 18 11:16:17 2021 Epoch 21, lr: 0.0007974, train loss: 0.1514, valid loss: 0.1352, aucs: 0.9631.\n0.9710 0.9645 0.9928 0.9793 0.9746 0.9861 0.9889 0.9224 0.8883 0.9295 0.9968\nThu Mar 18 11:22:06 2021 Epoch 22, lr: 0.0006582, train loss: 0.1482, valid loss: 0.1363, aucs: 0.9630.\n0.9737 0.9654 0.9930 0.9768 0.9717 0.9865 0.9891 0.9221 0.8878 0.9299 0.9971\nThu Mar 18 11:27:59 2021 Epoch 23, lr: 0.0005289, train loss: 0.1463, valid loss: 0.1337, aucs: 0.9643.\n0.9678 0.9675 0.9932 0.9816 0.9769 0.9867 0.9889 0.9232 0.8903 0.9332 0.9977\nThu Mar 18 11:33:51 2021 Epoch 24, lr: 0.0004110, train loss: 0.1444, valid loss: 0.1340, aucs: 0.9648.\n0.9808 0.9644 0.9929 0.9806 0.9751 0.9863 0.9891 0.9245 0.8901 0.9322 0.9970\nThu Mar 18 11:39:41 2021 Epoch 25, lr: 0.0003059, train loss: 0.1429, valid loss: 0.1330, aucs: 0.9650.\n0.9785 0.9655 0.9930 0.9827 0.9766 0.9867 0.9892 0.9238 0.8893 0.9327 0.9971\nThu Mar 18 11:45:35 2021 Epoch 26, lr: 0.0002147, train loss: 0.1417, valid loss: 0.1351, aucs: 0.9648.\n0.9803 0.9662 0.9931 0.9803 0.9754 0.9863 0.9888 0.9232 0.8894 0.9330 0.9970\nThu Mar 18 11:51:26 2021 Epoch 27, lr: 0.0001386, train loss: 0.1403, valid loss: 0.1336, aucs: 0.9653.\n0.9767 0.9672 0.9932 0.9814 0.9773 0.9867 0.9894 0.9251 0.8905 0.9334 0.9970\nThu Mar 18 11:57:16 2021 Epoch 28, lr: 0.0000785, train loss: 0.1393, valid loss: 0.1339, aucs: 0.9656.\n0.9786 0.9671 0.9933 0.9815 0.9776 0.9867 0.9894 0.9258 0.8908 0.9332 0.9973\nThu Mar 18 12:03:07 2021 Epoch 29, lr: 0.0000351, train loss: 0.1380, valid loss: 0.1338, aucs: 0.9655.\n0.9794 0.9670 0.9932 0.9816 0.9771 0.9865 0.9893 0.9255 0.8905 0.9336 0.9972\nThu Mar 18 12:08:59 2021 Epoch 30, lr: 0.0000088, train loss: 0.1387, valid loss: 0.1334, aucs: 0.9656.\n0.9776 0.9670 0.9932 0.9816 0.9775 0.9867 0.9894 0.9264 0.8911 0.9336 0.9975\n```","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}