{"cells":[{"metadata":{"papermill":{"duration":0.046144,"end_time":"2020-12-14T18:40:32.916699","exception":false,"start_time":"2020-12-14T18:40:32.870555","status":"completed"},"tags":[]},"cell_type":"markdown","source":"In this kernel I would use tf.data plus Keras to build a baseline, this type of baseline can be helpful to you in solving similar problems as well."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.execute_input":"2020-12-14T18:40:33.230014Z","iopub.status.busy":"2020-12-14T18:40:33.22739Z","iopub.status.idle":"2020-12-14T18:40:41.696905Z","shell.execute_reply":"2020-12-14T18:40:41.697543Z"},"papermill":{"duration":8.52276,"end_time":"2020-12-14T18:40:41.697722","exception":false,"start_time":"2020-12-14T18:40:33.174962","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential, Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping, ReduceLROnPlateau\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten,GlobalAveragePooling2D,BatchNormalization, Activation\nimport glob\n\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow import keras\n\n\nimport os\n\nfrom tensorflow.compat.v1 import ConfigProto\nfrom tensorflow.compat.v1 import InteractiveSession\nconfig = ConfigProto()\nconfig.gpu_options.allow_growth = True\nsession = InteractiveSession(config=config)\n","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.045087,"end_time":"2020-12-14T18:40:41.787237","exception":false,"start_time":"2020-12-14T18:40:41.74215","status":"completed"},"tags":[]},"cell_type":"markdown","source":"Adding Seed helps to reproduce results. Setting Debug Parameter will run the model on smaller number of epochs to validate the architecture."},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:41.882839Z","iopub.status.busy":"2020-12-14T18:40:41.882059Z","iopub.status.idle":"2020-12-14T18:40:41.886239Z","shell.execute_reply":"2020-12-14T18:40:41.886722Z"},"papermill":{"duration":0.05466,"end_time":"2020-12-14T18:40:41.886862","exception":false,"start_time":"2020-12-14T18:40:41.832202","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"SEED = 42\nDEBUG = False\n\nos.environ['PYTHONHASHSEED'] = str(SEED)\nnp.random.seed(SEED)\ntf.random.set_seed(SEED)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.045947,"end_time":"2020-12-14T18:40:41.977259","exception":false,"start_time":"2020-12-14T18:40:41.931312","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Prepare Data"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.077637Z","iopub.status.busy":"2020-12-14T18:40:42.076915Z","iopub.status.idle":"2020-12-14T18:40:42.123489Z","shell.execute_reply":"2020-12-14T18:40:42.124317Z"},"papermill":{"duration":0.100738,"end_time":"2020-12-14T18:40:42.12454","exception":false,"start_time":"2020-12-14T18:40:42.023802","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(len(df), df['StudyInstanceUID'].nunique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of Records:',len(df), 'Number of Patients:' ,df['PatientID'].nunique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['PatientID'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.044301,"end_time":"2020-12-14T18:40:42.213233","exception":false,"start_time":"2020-12-14T18:40:42.168932","status":"completed"},"tags":[]},"cell_type":"markdown","source":"Class Distribution of dataset:"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.columns","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.30932Z","iopub.status.busy":"2020-12-14T18:40:42.308351Z","iopub.status.idle":"2020-12-14T18:40:42.33966Z","shell.execute_reply":"2020-12-14T18:40:42.340782Z"},"papermill":{"duration":0.08319,"end_time":"2020-12-14T18:40:42.340966","exception":false,"start_time":"2020-12-14T18:40:42.257776","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df['path'] = '../input/ranzcr-clip-catheter-line-classification/train/' + df['StudyInstanceUID']+'.jpg'\nlabels = ['ETT - Abnormal', 'ETT - Borderline',\n       'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline',\n       'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal',\n       'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present']\n\nfor label in labels:\n    print(\"#\"*25)\n    print(label)\n    print(df[label].value_counts(normalize=True) * 100)\n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.438651Z","iopub.status.busy":"2020-12-14T18:40:42.437634Z","iopub.status.idle":"2020-12-14T18:40:42.44072Z","shell.execute_reply":"2020-12-14T18:40:42.439965Z"},"papermill":{"duration":0.053323,"end_time":"2020-12-14T18:40:42.440831","exception":false,"start_time":"2020-12-14T18:40:42.387508","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"if DEBUG:\n    _, df = train_test_split(df, test_size = 0.4, random_state=SEED, shuffle=True)\n","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.045966,"end_time":"2020-12-14T18:40:42.532132","exception":false,"start_time":"2020-12-14T18:40:42.486166","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Data Loader using tf.Data"},{"metadata":{"papermill":{"duration":0.046404,"end_time":"2020-12-14T18:40:42.624603","exception":false,"start_time":"2020-12-14T18:40:42.578199","status":"completed"},"tags":[]},"cell_type":"markdown","source":"### Spliting Dataset"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.723229Z","iopub.status.busy":"2020-12-14T18:40:42.722104Z","iopub.status.idle":"2020-12-14T18:40:42.729924Z","shell.execute_reply":"2020-12-14T18:40:42.729381Z"},"papermill":{"duration":0.059794,"end_time":"2020-12-14T18:40:42.730034","exception":false,"start_time":"2020-12-14T18:40:42.67024","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"X_train, X_valid = train_test_split(df, test_size = 0.1, random_state=SEED, shuffle=True)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.839641Z","iopub.status.busy":"2020-12-14T18:40:42.838193Z","iopub.status.idle":"2020-12-14T18:40:42.868094Z","shell.execute_reply":"2020-12-14T18:40:42.867002Z"},"papermill":{"duration":0.091434,"end_time":"2020-12-14T18:40:42.868199","exception":false,"start_time":"2020-12-14T18:40:42.776765","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"train_ds = tf.data.Dataset.from_tensor_slices((X_train.path.values, X_train[labels].values))\nvalid_ds = tf.data.Dataset.from_tensor_slices((X_valid.path.values, X_valid[labels].values))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:42.968805Z","iopub.status.busy":"2020-12-14T18:40:42.967849Z","iopub.status.idle":"2020-12-14T18:40:43.031876Z","shell.execute_reply":"2020-12-14T18:40:43.030792Z"},"papermill":{"duration":0.11664,"end_time":"2020-12-14T18:40:43.032028","exception":false,"start_time":"2020-12-14T18:40:42.915388","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"for path, label in train_ds.take(5):\n    print ('Path: {}, Label: {}'.format(path, label))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:43.129515Z","iopub.status.busy":"2020-12-14T18:40:43.128585Z","iopub.status.idle":"2020-12-14T18:40:43.139075Z","shell.execute_reply":"2020-12-14T18:40:43.139515Z"},"papermill":{"duration":0.061607,"end_time":"2020-12-14T18:40:43.139636","exception":false,"start_time":"2020-12-14T18:40:43.078029","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"for path, label in valid_ds.take(5):\n    print ('Path: {}, Label: {}'.format(path, label))","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.046077,"end_time":"2020-12-14T18:40:43.232522","exception":false,"start_time":"2020-12-14T18:40:43.186445","status":"completed"},"tags":[]},"cell_type":"markdown","source":"### Data Generator"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:43.332205Z","iopub.status.busy":"2020-12-14T18:40:43.331329Z","iopub.status.idle":"2020-12-14T18:40:43.334837Z","shell.execute_reply":"2020-12-14T18:40:43.334147Z"},"papermill":{"duration":0.054274,"end_time":"2020-12-14T18:40:43.334964","exception":false,"start_time":"2020-12-14T18:40:43.28069","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"AUTOTUNE = tf.data.experimental.AUTOTUNE\n\ntarget_size_dim = 300","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:43.442731Z","iopub.status.busy":"2020-12-14T18:40:43.441816Z","iopub.status.idle":"2020-12-14T18:40:43.444891Z","shell.execute_reply":"2020-12-14T18:40:43.445382Z"},"papermill":{"duration":0.060771,"end_time":"2020-12-14T18:40:43.445509","exception":false,"start_time":"2020-12-14T18:40:43.384738","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def process_data_train(image_path, label):\n    # load the raw data from the file as a string\n    img = tf.io.read_file(image_path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.random_brightness(img, 0.3)\n    img = tf.image.random_flip_left_right(img)\n    img = tf.image.resize(img, [target_size_dim,target_size_dim])\n    return img, label\n\ndef process_data_valid(image_path, label):\n    # load the raw data from the file as a string\n    img = tf.io.read_file(image_path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [target_size_dim,target_size_dim])\n    return img, label\n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:43.545046Z","iopub.status.busy":"2020-12-14T18:40:43.543889Z","iopub.status.idle":"2020-12-14T18:40:43.806319Z","shell.execute_reply":"2020-12-14T18:40:43.805718Z"},"papermill":{"duration":0.31506,"end_time":"2020-12-14T18:40:43.80648","exception":false,"start_time":"2020-12-14T18:40:43.49142","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Set `num_parallel_calls` so multiple images are loaded/processed in parallel.\ntrain_ds = train_ds.map(process_data_train, num_parallel_calls=AUTOTUNE)\nvalid_ds = valid_ds.map(process_data_valid, num_parallel_calls=AUTOTUNE)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:43.911026Z","iopub.status.busy":"2020-12-14T18:40:43.90985Z","iopub.status.idle":"2020-12-14T18:40:44.366299Z","shell.execute_reply":"2020-12-14T18:40:44.367132Z"},"papermill":{"duration":0.511192,"end_time":"2020-12-14T18:40:44.367329","exception":false,"start_time":"2020-12-14T18:40:43.856137","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"for image, label in train_ds.take(1):\n    plt.imshow(image.numpy().astype('uint8'))\n    plt.show()\n    print(\"Image shape: \", image.numpy().shape)\n    print(\"Label: \", labels[np.argmax(label.numpy())])","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.052789,"end_time":"2020-12-14T18:40:44.47564","exception":false,"start_time":"2020-12-14T18:40:44.422851","status":"completed"},"tags":[]},"cell_type":"markdown","source":"### Improving Performance"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:44.587214Z","iopub.status.busy":"2020-12-14T18:40:44.586405Z","iopub.status.idle":"2020-12-14T18:40:44.614834Z","shell.execute_reply":"2020-12-14T18:40:44.614169Z"},"papermill":{"duration":0.087412,"end_time":"2020-12-14T18:40:44.614949","exception":false,"start_time":"2020-12-14T18:40:44.527537","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def configure_for_performance(ds, batch_size = 32):\n    ds = ds.cache('/kaggle/dump.tfcache') ## Due to ram and disk limitation of kaggle you cannot set cache to be true. \n                                    ## Setting this to true can increase performance by 2x\n    \n    ds = ds.repeat()\n    ds = ds.shuffle(buffer_size=1024)\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(buffer_size=AUTOTUNE)\n    return ds\n\nbatch_size = 32\n\ntrain_ds_batch = configure_for_performance(train_ds)\nvalid_ds_batch = valid_ds.batch(batch_size*2)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:44.725061Z","iopub.status.busy":"2020-12-14T18:40:44.723936Z","iopub.status.idle":"2020-12-14T18:40:55.452929Z","shell.execute_reply":"2020-12-14T18:40:55.453758Z"},"papermill":{"duration":10.787192,"end_time":"2020-12-14T18:40:55.454021","exception":false,"start_time":"2020-12-14T18:40:44.666829","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"image_batch, label_batch = next(iter(train_ds_batch))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:55.62901Z","iopub.status.busy":"2020-12-14T18:40:55.627917Z","iopub.status.idle":"2020-12-14T18:40:56.569054Z","shell.execute_reply":"2020-12-14T18:40:56.569574Z"},"papermill":{"duration":1.028802,"end_time":"2020-12-14T18:40:56.569712","exception":false,"start_time":"2020-12-14T18:40:55.54091","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"\nplt.figure(figsize=(10, 10))\nfor i in range(16):\n    ax = plt.subplot(4, 4, i + 1)\n    plt.imshow(image_batch[i].numpy().astype(\"uint8\"))\n    label = labels[np.argmax(label_batch[i].numpy())]\n    plt.title(label)\n    plt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.061062,"end_time":"2020-12-14T18:40:56.695643","exception":false,"start_time":"2020-12-14T18:40:56.634581","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Data Augmentation"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:56.831478Z","iopub.status.busy":"2020-12-14T18:40:56.830781Z","iopub.status.idle":"2020-12-14T18:40:56.851552Z","shell.execute_reply":"2020-12-14T18:40:56.850809Z"},"papermill":{"duration":0.092355,"end_time":"2020-12-14T18:40:56.851679","exception":false,"start_time":"2020-12-14T18:40:56.759324","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"data_augmentation = keras.Sequential(\n    [\n        tf.keras.layers.experimental.preprocessing.RandomRotation(0.05, interpolation='nearest'),\n        #tf.keras.layers.experimental.preprocessing.RandomContrast((0.1))\n    ]\n)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:56.994022Z","iopub.status.busy":"2020-12-14T18:40:56.993268Z","iopub.status.idle":"2020-12-14T18:40:58.344966Z","shell.execute_reply":"2020-12-14T18:40:58.345507Z"},"papermill":{"duration":1.426889,"end_time":"2020-12-14T18:40:58.345643","exception":false,"start_time":"2020-12-14T18:40:56.918754","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"\nplt.figure(figsize=(10, 10))\nfor i in range(16):\n    augmented_images = data_augmentation(image_batch)\n    ax = plt.subplot(4, 4, i + 1)\n    plt.imshow(augmented_images[i].numpy().astype(\"uint8\"))\n    label = labels[np.argmax(label_batch[i].numpy())]\n    plt.title(label)\n    plt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.078356,"end_time":"2020-12-14T18:40:58.505525","exception":false,"start_time":"2020-12-14T18:40:58.427169","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Creating Model"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:58.673695Z","iopub.status.busy":"2020-12-14T18:40:58.672752Z","iopub.status.idle":"2020-12-14T18:40:58.677154Z","shell.execute_reply":"2020-12-14T18:40:58.67654Z"},"papermill":{"duration":0.087677,"end_time":"2020-12-14T18:40:58.677263","exception":false,"start_time":"2020-12-14T18:40:58.589586","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"## Only available in tf2.3+\n\nfrom tensorflow.keras.applications import EfficientNetB3 \n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:58.85636Z","iopub.status.busy":"2020-12-14T18:40:58.850723Z","iopub.status.idle":"2020-12-14T18:40:58.859156Z","shell.execute_reply":"2020-12-14T18:40:58.858642Z"},"papermill":{"duration":0.09709,"end_time":"2020-12-14T18:40:58.85927","exception":false,"start_time":"2020-12-14T18:40:58.76218","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def load_pretrained_model(weights_path, drop_connect, target_size_dim, layers_to_unfreeze=5):\n    model = EfficientNetB3(\n            weights=None, \n            include_top=False, \n            drop_connect_rate=0.4\n        )\n    \n    model.load_weights(weights_path)\n    \n    model.trainable = True\n\n    # for layer in model.layers[-layers_to_unfreeze:]:\n    #     if not isinstance(layer, tf.keras.layers.BatchNormalization): \n    #         layer.trainable = True\n\n    if DEBUG:\n        for layer in model.layers:\n            #print(layer.name, layer.trainable)\n            pass\n\n    return model\n\ndef build_my_model(base_model, optimizer, metrics, loss):\n    \n    inputs = tf.keras.layers.Input(shape=(target_size_dim, target_size_dim, 3))\n    x = data_augmentation(inputs)\n    outputs_eff = base_model(x)\n    global_avg_pooling = GlobalAveragePooling2D()(outputs_eff)\n    dense_1= Dense(256)(global_avg_pooling)\n    bn_1 = BatchNormalization()(dense_1)\n    activation = Activation('relu')(bn_1)\n    dropout = Dropout(0.3)(activation)\n    dense_2 = Dense(len(labels), activation='sigmoid')(dropout)\n\n    my_model = tf.keras.Model(inputs, dense_2)\n    \n    my_model.compile(\n        optimizer=optimizer,\n        loss=loss,\n        metrics=metrics\n    )\n    return my_model\n\n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:59.02237Z","iopub.status.busy":"2020-12-14T18:40:59.021561Z","iopub.status.idle":"2020-12-14T18:40:59.02569Z","shell.execute_reply":"2020-12-14T18:40:59.025193Z"},"papermill":{"duration":0.08822,"end_time":"2020-12-14T18:40:59.025795","exception":false,"start_time":"2020-12-14T18:40:58.937575","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"#!wget https://storage.googleapis.com/keras-applications/efficientnetb3_notop.h5\n## to get model weights","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:59.225573Z","iopub.status.busy":"2020-12-14T18:40:59.224423Z","iopub.status.idle":"2020-12-14T18:40:59.228499Z","shell.execute_reply":"2020-12-14T18:40:59.229132Z"},"papermill":{"duration":0.09106,"end_time":"2020-12-14T18:40:59.229271","exception":false,"start_time":"2020-12-14T18:40:59.138211","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"model_weights_path = '../input/noisystudent/efficientnetb3_notop.h5'\nmodel_weights_path","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:40:59.400894Z","iopub.status.busy":"2020-12-14T18:40:59.398238Z","iopub.status.idle":"2020-12-14T18:41:07.291064Z","shell.execute_reply":"2020-12-14T18:41:07.291659Z"},"papermill":{"duration":7.983467,"end_time":"2020-12-14T18:41:07.291805","exception":false,"start_time":"2020-12-14T18:40:59.308338","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"drop_rate = 0.4 ## value of dropout to be used in loaded network\nbase_model = load_pretrained_model( model_weights_path, drop_rate, target_size_dim )\n\noptimizer = tf.keras.optimizers.Adam(lr = 1e-4)\nmetrics = tf.keras.metrics.AUC(multi_label=True)\nmy_model = build_my_model(base_model, optimizer, metrics = [metrics], loss='binary_crossentropy')\nmy_model.summary()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.079859,"end_time":"2020-12-14T18:41:07.453696","exception":false,"start_time":"2020-12-14T18:41:07.373837","status":"completed"},"tags":[]},"cell_type":"markdown","source":"### Callbacks"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:41:07.619634Z","iopub.status.busy":"2020-12-14T18:41:07.618664Z","iopub.status.idle":"2020-12-14T18:41:07.621969Z","shell.execute_reply":"2020-12-14T18:41:07.621465Z"},"papermill":{"duration":0.090983,"end_time":"2020-12-14T18:41:07.622127","exception":false,"start_time":"2020-12-14T18:41:07.531144","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"weight_path_save = 'best_model.hdf5'\nlast_weight_path = 'last_model.hdf5'\n\ncheckpoint = ModelCheckpoint(weight_path_save, \n                             monitor= 'val_loss', \n                             verbose=1, \n                             save_best_only=True, \n                             mode= 'min', \n                             save_weights_only = False)\ncheckpoint_last = ModelCheckpoint(last_weight_path, \n                             monitor= 'val_loss', \n                             verbose=1, \n                             save_best_only=False, \n                             mode= 'min', \n                             save_weights_only = False)\n\n\nearly = EarlyStopping(monitor= 'val_loss', \n                      mode= 'min', \n                      patience=5)\n\nreduceLROnPlat = ReduceLROnPlateau(monitor='val_loss', factor=0.8, patience=2, verbose=1, mode='auto', epsilon=0.0001, cooldown=5, min_lr=0.00001)\ncallbacks_list = [checkpoint, checkpoint_last, early, reduceLROnPlat]","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.077392,"end_time":"2020-12-14T18:41:07.77705","exception":false,"start_time":"2020-12-14T18:41:07.699658","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Train Model"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:41:07.942061Z","iopub.status.busy":"2020-12-14T18:41:07.941365Z","iopub.status.idle":"2020-12-14T18:41:07.946441Z","shell.execute_reply":"2020-12-14T18:41:07.945748Z"},"papermill":{"duration":0.088692,"end_time":"2020-12-14T18:41:07.946544","exception":false,"start_time":"2020-12-14T18:41:07.857852","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"if DEBUG:\n    epochs = 3\nelse:\n    epochs = 10\n    \nprint(f\"Model will train for {epochs} epochs\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:41:08.111596Z","iopub.status.busy":"2020-12-14T18:41:08.110871Z","iopub.status.idle":"2020-12-14T18:41:08.115605Z","shell.execute_reply":"2020-12-14T18:41:08.11506Z"},"papermill":{"duration":0.091334,"end_time":"2020-12-14T18:41:08.115724","exception":false,"start_time":"2020-12-14T18:41:08.02439","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# from sklearn.utils import class_weight\n\n# classes_to_predict =[0, 1, 2, 3, 4]\n# class_weights = class_weight.compute_class_weight(\"balanced\", classes_to_predict, train_gen.labels)\n# class_weights_dict = {i : class_weights[i] for i,label in enumerate(classes_to_predict)}\n\n# print(class_weights_dict)\n","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T18:41:08.290195Z","iopub.status.busy":"2020-12-14T18:41:08.28952Z","iopub.status.idle":"2020-12-14T20:30:04.900431Z","shell.execute_reply":"2020-12-14T20:30:04.890747Z"},"papermill":{"duration":6536.702218,"end_time":"2020-12-14T20:30:04.90061","exception":false,"start_time":"2020-12-14T18:41:08.198392","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"if DEBUG:\n    history = my_model.fit(train_ds_batch, \n                              validation_data = valid_ds_batch, \n                              epochs = epochs, \n                              callbacks = callbacks_list,\n                               steps_per_epoch = 1,\n                           validation_steps = 1\n                               #class_weight=class_weights_dict\n                              )\nelse:\n    steps_per_epoch = len(X_train) // batch_size\n    history = my_model.fit(train_ds_batch, \n                              validation_data = valid_ds_batch, \n                              epochs = epochs, \n                              callbacks = callbacks_list,\n                                steps_per_epoch = steps_per_epoch\n                               #class_weight=class_weights_dict\n                              )","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:30:11.587211Z","iopub.status.busy":"2020-12-14T20:30:11.586241Z","iopub.status.idle":"2020-12-14T20:30:11.589207Z","shell.execute_reply":"2020-12-14T20:30:11.588689Z"},"papermill":{"duration":3.267097,"end_time":"2020-12-14T20:30:11.589307","exception":false,"start_time":"2020-12-14T20:30:08.32221","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"\ndef plot_hist(hist):\n    plt.figure(figsize=(15,5))\n    local_epochs = len(hist.history[\"auc\"])\n    plt.plot(np.arange(local_epochs), hist.history[\"auc\"], '-o', label='Train AUC',color='#ff7f0e')\n    plt.plot(np.arange(local_epochs), hist.history[\"val_auc\"], '-o',label='Val AUC',color='#1f77b4')\n    plt.xlabel('Epoch',size=14)\n    plt.ylabel('Accuracy',size=14)\n    plt.legend(loc=2)\n    \n    plt2 = plt.gca().twinx()\n    plt2.plot(np.arange(local_epochs) ,history.history['loss'],'-o',label='Train Loss',color='#2ca02c')\n    plt2.plot(np.arange(local_epochs) ,history.history['val_loss'],'-o',label='Val Loss',color='#d62728')\n    plt.legend(loc=3)\n    plt.ylabel('Loss',size=14)\n    plt.title(\"Model Accuracy and loss\")\n    \n    #plt.legend([\"train\", \"validation\"], loc=\"upper left\")\n    \n    plt.savefig('loss.png')\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:30:18.993561Z","iopub.status.busy":"2020-12-14T20:30:18.992563Z","iopub.status.idle":"2020-12-14T20:30:19.556431Z","shell.execute_reply":"2020-12-14T20:30:19.557036Z"},"papermill":{"duration":4.625253,"end_time":"2020-12-14T20:30:19.557177","exception":false,"start_time":"2020-12-14T20:30:14.931924","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot_hist(history)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":3.212068,"end_time":"2020-12-14T20:30:25.981687","exception":false,"start_time":"2020-12-14T20:30:22.769619","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Evaluating Model on Validation Set"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:30:39.282288Z","iopub.status.busy":"2020-12-14T20:30:39.261822Z","iopub.status.idle":"2020-12-14T20:30:39.623581Z","shell.execute_reply":"2020-12-14T20:30:39.622953Z"},"papermill":{"duration":3.815574,"end_time":"2020-12-14T20:30:39.62371","exception":false,"start_time":"2020-12-14T20:30:35.808136","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"my_model.load_weights(weight_path_save) ## load the best model or all your metrics would be on the last run not on the best one","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:30:46.101911Z","iopub.status.busy":"2020-12-14T20:30:46.100758Z","iopub.status.idle":"2020-12-14T20:30:59.05312Z","shell.execute_reply":"2020-12-14T20:30:59.052511Z"},"papermill":{"duration":16.208495,"end_time":"2020-12-14T20:30:59.053262","exception":false,"start_time":"2020-12-14T20:30:42.844767","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"pred_valid_y = my_model.predict(valid_ds_batch,  verbose = True, workers=4)\npred_valid_y_labels = np.argmax(pred_valid_y, axis=-1)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:31:05.871385Z","iopub.status.busy":"2020-12-14T20:31:05.870432Z","iopub.status.idle":"2020-12-14T20:31:12.574234Z","shell.execute_reply":"2020-12-14T20:31:12.573323Z"},"papermill":{"duration":9.983731,"end_time":"2020-12-14T20:31:12.574403","exception":false,"start_time":"2020-12-14T20:31:02.590672","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"valid_labels = np.concatenate([y.numpy() for x, y in valid_ds_batch], axis=0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_labels","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":3.449956,"end_time":"2020-12-14T20:31:33.120029","exception":false,"start_time":"2020-12-14T20:31:29.670073","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Getting Predictions on Test Set"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:31:39.694172Z","iopub.status.busy":"2020-12-14T20:31:39.693241Z","iopub.status.idle":"2020-12-14T20:31:39.696397Z","shell.execute_reply":"2020-12-14T20:31:39.695884Z"},"papermill":{"duration":3.279335,"end_time":"2020-12-14T20:31:39.696514","exception":false,"start_time":"2020-12-14T20:31:36.417179","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"import glob","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:31:46.501799Z","iopub.status.busy":"2020-12-14T20:31:46.501129Z","iopub.status.idle":"2020-12-14T20:31:46.512607Z","shell.execute_reply":"2020-12-14T20:31:46.513103Z"},"papermill":{"duration":3.512657,"end_time":"2020-12-14T20:31:46.513246","exception":false,"start_time":"2020-12-14T20:31:43.000589","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"test_images = glob.glob('../input/ranzcr-clip-catheter-line-classification/test/*.jpg')\n#test_images = test_images * 5\n#print(test_images)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:31:53.421658Z","iopub.status.busy":"2020-12-14T20:31:53.420503Z","iopub.status.idle":"2020-12-14T20:31:53.42885Z","shell.execute_reply":"2020-12-14T20:31:53.428162Z"},"papermill":{"duration":3.693609,"end_time":"2020-12-14T20:31:53.428993","exception":false,"start_time":"2020-12-14T20:31:49.735384","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df_test = pd.DataFrame(np.array(test_images), columns=['Path'])\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:32:01.258394Z","iopub.status.busy":"2020-12-14T20:32:01.257541Z","iopub.status.idle":"2020-12-14T20:32:01.331903Z","shell.execute_reply":"2020-12-14T20:32:01.33129Z"},"papermill":{"duration":3.355682,"end_time":"2020-12-14T20:32:01.332028","exception":false,"start_time":"2020-12-14T20:31:57.976346","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"test_ds = tf.data.Dataset.from_tensor_slices((df_test.Path.values))\n\n\ndef process_test(image_path):\n    # load the raw data from the file as a string\n    img = tf.io.read_file(image_path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [target_size_dim,target_size_dim])\n    return img\n    \ntest_ds = test_ds.map(process_test, num_parallel_calls=AUTOTUNE).batch(batch_size*2)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:32:08.025187Z","iopub.status.busy":"2020-12-14T20:32:08.02422Z","iopub.status.idle":"2020-12-14T20:32:10.618709Z","shell.execute_reply":"2020-12-14T20:32:10.617535Z"},"papermill":{"duration":6.011183,"end_time":"2020-12-14T20:32:10.618838","exception":false,"start_time":"2020-12-14T20:32:04.607655","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"pred_y = my_model.predict(test_ds, workers=4, verbose=1)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_ss = pd.DataFrame(pred_y, columns = labels)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:32:17.326045Z","iopub.status.busy":"2020-12-14T20:32:17.325092Z","iopub.status.idle":"2020-12-14T20:32:17.328673Z","shell.execute_reply":"2020-12-14T20:32:17.329368Z"},"papermill":{"duration":3.537047,"end_time":"2020-12-14T20:32:17.329534","exception":false,"start_time":"2020-12-14T20:32:13.792487","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df_test['image_id'] = df_test.Path.str.split('/').str[-1].str[:-4]\ndf_ss['StudyInstanceUID'] = df_test['image_id']\ndf_ss.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_ss.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cols_reordered = ['StudyInstanceUID', 'ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal',\n       'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal',\n       'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal',\n       'Swan Ganz Catheter Present']\n\ndf_order = df_ss[cols_reordered]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_order.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-14T20:32:23.962388Z","iopub.status.busy":"2020-12-14T20:32:23.961678Z","iopub.status.idle":"2020-12-14T20:32:24.29118Z","shell.execute_reply":"2020-12-14T20:32:24.291917Z"},"papermill":{"duration":3.632097,"end_time":"2020-12-14T20:32:24.292162","exception":false,"start_time":"2020-12-14T20:32:20.660065","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df_order.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.chdir(r'../working')\nfrom IPython.display import FileLink\nFileLink(r'submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":3.248996,"end_time":"2020-12-14T20:32:38.12901","exception":false,"start_time":"2020-12-14T20:32:34.880014","status":"completed"},"tags":[]},"cell_type":"markdown","source":"## Work In Progress. This might not be the final solution\n\n### If you learnt something from this kernel kindly upvote :)**"},{"metadata":{"papermill":{"duration":3.209727,"end_time":"2020-12-14T20:32:44.826042","exception":false,"start_time":"2020-12-14T20:32:41.616315","status":"completed"},"tags":[]},"cell_type":"markdown","source":"Things to take care of:\n\n1. Data leakage : Ideally I would not want the same patient in training as well as validation set, this indicates data leakage\n2. Will try to solve this problem using efficientdet (considering this as a detection problem) as train annotations are given\n3. Use OOF Cross Validation"},{"metadata":{"papermill":{"duration":3.815296,"end_time":"2020-12-14T20:32:59.120356","exception":false,"start_time":"2020-12-14T20:32:55.30506","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}