{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Histopathologic Cancer Detection: CNN (Keras)**\n\nTable of Contents:\n* [About the Data / Problem](#about)\n* [Set Up](#setup)\n* [Exploratory Data Analysis](#eda)\n* [Model Building](#model)\n* [Model Training](#modelt)\n* [Model Validation](#modelv)\n* [Submission](#submission)\n* [Conclusion](#conclusion)\n* [References](#ref)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"about\"></a>\n# **About the Data / Problem**","metadata":{}},{"cell_type":"markdown","source":"This dataset and task can be found on [Kaggle](http://www.kaggle.com/c/histopathologic-cancer-detection/overview). The overarching task here is to 'create an algorithm to identify metastatic cancer in small image patches taken from larger digital pathology scans'. This task translates well to a binary classification problem. More specifically, a convolutional neural network (CNN) will be used to divide the images from the dataset into two classes, i.e. having cancer and not having cancer. \n\nThe dataset consists of microscopic images of lymph node tissue. Each image has a resolution of 96x96 pixels, and the task will be to identify metastatic cancer tissue in a 32x32 pixel center region of the image. According to the Kaggle competition description, the identification of at least 1 pixel of tumor tissue would effectively label the image as positive, i.e. having cancer. The train dataset consists of 220,025 images, while the test dataset contains 57,468 images. \n\nAs this was my first attempt at developing a CNN using Keras, this project was highly influenced by other works, primarily one by [Pablo Gómez](https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93). All other inspirational works have been listed in the reference section. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"setup\"></a>\n# **Set Up**","metadata":{}},{"cell_type":"code","source":"# Loading necessary libraries\nfrom glob import glob \nimport numpy as np\nimport pandas as pd\nimport keras,cv2,os\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten, BatchNormalization, Activation\nfrom keras.layers import Conv2D, MaxPool2D\nfrom tqdm import tqdm_notebook,trange\nimport matplotlib.pyplot as plt\nimport gc ","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2023-07-17T13:10:44.087963Z","iopub.execute_input":"2023-07-17T13:10:44.088274Z","iopub.status.idle":"2023-07-17T13:10:47.140774Z","shell.execute_reply.started":"2023-07-17T13:10:44.088215Z","shell.execute_reply":"2023-07-17T13:10:47.139895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating paths for train and test data\npath = \"../input/\" \ntrain_path = path + 'train/'\ntest_path = path + 'test/'","metadata":{"_uuid":"d72f6d8131277d5a1452d92677bcaaae8117165b","execution":{"iopub.status.busy":"2023-07-17T13:11:31.296453Z","iopub.execute_input":"2023-07-17T13:11:31.296828Z","iopub.status.idle":"2023-07-17T13:11:31.301538Z","shell.execute_reply.started":"2023-07-17T13:11:31.296749Z","shell.execute_reply":"2023-07-17T13:11:31.300629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a dataframe from the train_path filenames and lables\ndf = pd.DataFrame({'path': glob(os.path.join(train_path,'*.tif'))}) \ndf['id'] = df.path.map(lambda x: x.split('/')[3].split(\".\")[0]) \nlabels = pd.read_csv(path+\"train_labels.csv\") \ndf = df.merge(labels, on = \"id\") ","metadata":{"execution":{"iopub.status.busy":"2023-07-17T13:11:32.945518Z","iopub.execute_input":"2023-07-17T13:11:32.945862Z","iopub.status.idle":"2023-07-17T13:11:37.270559Z","shell.execute_reply.started":"2023-07-17T13:11:32.945776Z","shell.execute_reply":"2023-07-17T13:11:37.269785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Viewing the new dataframe\ndf.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-07-17T13:12:51.633769Z","iopub.execute_input":"2023-07-17T13:12:51.634130Z","iopub.status.idle":"2023-07-17T13:12:51.650177Z","shell.execute_reply.started":"2023-07-17T13:12:51.634070Z","shell.execute_reply":"2023-07-17T13:12:51.649388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the info of the new dataframe\ndf.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-17T13:13:09.123272Z","iopub.execute_input":"2023-07-17T13:13:09.123604Z","iopub.status.idle":"2023-07-17T13:13:09.214988Z","shell.execute_reply.started":"2023-07-17T13:13:09.123548Z","shell.execute_reply":"2023-07-17T13:13:09.213921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading N images from df dataframe\ndef load_data(N,df):\n    X = np.zeros([N,96,96,3],dtype=np.uint8) \n    y = np.squeeze(df.as_matrix(columns=['label']))[0:N]\n    for i, row in tqdm_notebook(df.iterrows(), total=N):\n        if i == N:\n            break\n        X[i] = cv2.imread(row['path'])     \n    return X,y","metadata":{"_uuid":"5087ee421fbaf96cbd18df1467cf14c34b4f7289","execution":{"iopub.status.busy":"2023-07-17T13:14:05.922453Z","iopub.execute_input":"2023-07-17T13:14:05.922799Z","iopub.status.idle":"2023-07-17T13:14:05.928769Z","shell.execute_reply.started":"2023-07-17T13:14:05.922726Z","shell.execute_reply":"2023-07-17T13:14:05.927763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading 10000 images\nN=10000\nX,y = load_data(N=N,df=df) ","metadata":{"_uuid":"d2207ea4a34a42c5173b48e4ba90403fe75d6685","execution":{"iopub.status.busy":"2023-07-17T13:14:19.744159Z","iopub.execute_input":"2023-07-17T13:14:19.744482Z","iopub.status.idle":"2023-07-17T13:15:52.960348Z","shell.execute_reply.started":"2023-07-17T13:14:19.744421Z","shell.execute_reply":"2023-07-17T13:15:52.959481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code chunks above, inspired by a great [work](https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93/notebook) on this project, converts all images to the unit8 format. This ensures that all images have a pixel format with integer values from 0 to 225. This will also help reduce the memory footprint of this project. I have also loaded in 10000 images for this project. Sticking with a smaller amount of images will also reduce the memory footprint.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"eda\"></a>\n# **Exploratory Data Analysis**","metadata":{}},{"cell_type":"code","source":"# Visualizing a set of random images\nfig = plt.figure(figsize=(10, 6), dpi=150)\nnp.random.seed(42) \nfor plotNr,idx in enumerate(np.random.randint(0,N,8)):\n    ax = fig.add_subplot(2, 8//2, plotNr+1, xticks=[], yticks=[]) \n    plt.imshow(X[idx]) \n    ax.set_title('Label: ' + str(y[idx])) ","metadata":{"_uuid":"ffed68f7b732222097006f1047744950ad085ed8","execution":{"iopub.status.busy":"2023-07-17T13:21:38.782105Z","iopub.execute_input":"2023-07-17T13:21:38.782429Z","iopub.status.idle":"2023-07-17T13:21:39.614515Z","shell.execute_reply.started":"2023-07-17T13:21:38.782369Z","shell.execute_reply":"2023-07-17T13:21:39.613776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The images above represent random examples of tissue with and without evidence of cancer. Changing the random.seed in the code chunk above will result in different examples being shown. An image labeled as 0 represents tissue WITHOUT cancer and an image labeled as 1 represents tissue WITH cancer. As I have zero knowledge or background in histopathology, it would be unlikely that I could correctly identify a tumor in any given tissue. However, upon visualizing several sets of random images, it seems as though the tissues with large white dots often warrant a positive label. From this, one could deduce that these large white dots are tumors. Although this doesn't necessarily help us here, it is definitely interesting.  ","metadata":{}},{"cell_type":"code","source":"# Visualizing the distribution of positive and negative cases\nfig = plt.figure(figsize=(4, 2),dpi=150)\nplt.bar([1,0], [(y==0).sum(), (y==1).sum()]); \nplt.xticks([1,0],[\"Negative (0) (N={})\".format((y==0).sum()),\"Positive (1) (N={})\".format((y==1).sum())]);\nplt.ylabel(\"Samples\")","metadata":{"_uuid":"9c0351062237909773100c6e87409fe270379652","execution":{"iopub.status.busy":"2023-07-17T13:28:39.701288Z","iopub.execute_input":"2023-07-17T13:28:39.701626Z","iopub.status.idle":"2023-07-17T13:28:39.928541Z","shell.execute_reply.started":"2023-07-17T13:28:39.701568Z","shell.execute_reply":"2023-07-17T13:28:39.927362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above, we can see the value counts of each label in the dataset. Overall, there are 5953 negative images and 4047 positive images. In other words, roughly 60% of the dataset contains images of tissue WITHOUT cancer and 40% WITH cancer. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"model\"></a>\n# **Model Building**","metadata":{"_uuid":"c3bd89f6d9e00472c46b1c9c617e3d090c9fccce"}},{"cell_type":"code","source":"# Loading images from the training set\nN = df[\"path\"].size\nX,y = load_data(N=N,df=df)","metadata":{"_uuid":"30e69a2a034e279da6eeb0fc09f1a58b58fc0401","execution":{"iopub.status.busy":"2023-07-17T13:33:31.192805Z","iopub.execute_input":"2023-07-17T13:33:31.193167Z","iopub.status.idle":"2023-07-17T14:06:15.452221Z","shell.execute_reply.started":"2023-07-17T13:33:31.193108Z","shell.execute_reply":"2023-07-17T14:06:15.451324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting the data into train and validation sets\ntraining_portion = 0.8 \nsplit_idx = int(np.round(training_portion * y.shape[0])) \nnp.random.seed(42) ","metadata":{"_uuid":"60644d18d00811a12f6270ac42f8c745902018a8","execution":{"iopub.status.busy":"2023-07-17T14:06:18.861740Z","iopub.execute_input":"2023-07-17T14:06:18.862094Z","iopub.status.idle":"2023-07-17T14:06:18.867378Z","shell.execute_reply.started":"2023-07-17T14:06:18.862036Z","shell.execute_reply":"2023-07-17T14:06:18.866552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shuffling the data\nidx = np.arange(y.shape[0])\nnp.random.shuffle(idx)\nX = X[idx]\ny = y[idx]","metadata":{"execution":{"iopub.status.busy":"2023-07-17T14:06:22.228099Z","iopub.execute_input":"2023-07-17T14:06:22.228405Z","iopub.status.idle":"2023-07-17T14:06:28.049185Z","shell.execute_reply.started":"2023-07-17T14:06:22.228348Z","shell.execute_reply":"2023-07-17T14:06:28.048267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The model below has been inspired by another great [submission](http://www.kaggle.com/code/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb/notebook) to this competition (also listed in the references). ","metadata":{}},{"cell_type":"code","source":"# Building the model\n# Parameters\nkernel_size = (3,3)\npool_size= (2,2)\nfirst_filters = 32\nsecond_filters = 64\nthird_filters = 128\n\n# Dropout for regularization, 0.3 for conv layers, 0.5 for dense layer\ndropout_conv = 0.3\ndropout_dense = 0.5\n\n# Initializing the model\nmodel = Sequential()\n\n# Building the first conv block\nmodel.add(Conv2D(first_filters, kernel_size, input_shape = (96, 96, 3)))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(Conv2D(first_filters, kernel_size, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(MaxPool2D(pool_size = pool_size)) \nmodel.add(Dropout(dropout_conv))\n\n# Building the second conv block\nmodel.add(Conv2D(second_filters, kernel_size, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(Conv2D(second_filters, kernel_size, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(MaxPool2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\n# Building the third conv block\nmodel.add(Conv2D(third_filters, kernel_size, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(Conv2D(third_filters, kernel_size, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(MaxPool2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\n# Building the dense layer\nmodel.add(Flatten())\nmodel.add(Dense(256, use_bias=False))\nmodel.add(BatchNormalization())\nmodel.add(Activation(\"relu\"))\nmodel.add(Dropout(dropout_dense))\n\n# Converting 0 to 1 using sigmoid activation function\nmodel.add(Dense(1, activation = \"sigmoid\"))","metadata":{"_uuid":"9434b86b47f6132843b328d4fbf326aad01b24c0","execution":{"iopub.status.busy":"2023-07-17T14:06:37.040063Z","iopub.execute_input":"2023-07-17T14:06:37.040395Z","iopub.status.idle":"2023-07-17T14:06:39.626225Z","shell.execute_reply.started":"2023-07-17T14:06:37.040336Z","shell.execute_reply":"2023-07-17T14:06:39.625323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compiling the model\nbatch_size = 50\n\nmodel.compile(loss=keras.losses.binary_crossentropy,\n              optimizer=keras.optimizers.Adam(0.001), \n              metrics=['accuracy'])","metadata":{"_uuid":"43c6aea09f1541f2411ad651ada5a3552304ac0b","execution":{"iopub.status.busy":"2023-07-17T14:06:47.172499Z","iopub.execute_input":"2023-07-17T14:06:47.172959Z","iopub.status.idle":"2023-07-17T14:06:47.226347Z","shell.execute_reply.started":"2023-07-17T14:06:47.172751Z","shell.execute_reply":"2023-07-17T14:06:47.225500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above, we have created our model with three convolutional layers and a dense layer. We will also use a batch size of 50, binary cross entropy and the Adam optimizer with a learning rate of 0.001. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"modelt\"></a>\n# **Model Training**","metadata":{}},{"cell_type":"code","source":"# Training the model for 3 epochs\nepochs = 3 \nfor epoch in range(epochs):\n    iterations = np.floor(split_idx / batch_size).astype(int) \n    loss,acc = 0,0 \n    with trange(iterations) as t: \n        for i in t:\n            start_idx = i * batch_size \n            x_batch = X[start_idx:start_idx+batch_size] \n            y_batch = y[start_idx:start_idx+batch_size] \n            metrics = model.train_on_batch(x_batch, y_batch) \n            loss = loss + metrics[0] \n            acc = acc + metrics[1] \n            t.set_description('Running training epoch ' + str(epoch)) \n            t.set_postfix(loss=\"%.2f\" % round(loss / (i+1),2),acc=\"%.2f\" % round(acc / (i+1),2)) ","metadata":{"_uuid":"594a0db99d65d223695d4dd8d9437d915363e055","execution":{"iopub.status.busy":"2023-07-17T14:08:56.236047Z","iopub.execute_input":"2023-07-17T14:08:56.236410Z","iopub.status.idle":"2023-07-17T14:15:35.930075Z","shell.execute_reply.started":"2023-07-17T14:08:56.236348Z","shell.execute_reply":"2023-07-17T14:15:35.929322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code chunk above trains the model on 3 epochs. As shown above, the first epoch yields an accuracy of 0.85 and a loss of 0.36. The second yields an accuracy of 0.90 and a loss of 0.26. And the third yields an accuracy of 0.91 and a loss of 0.22. From this, we can see that increasing the number of epochs has both increased accuracy and decreased loss. Any future work on this project could test and see how accuracy and loss would be changed if the number of epochs continues to increase. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"modelv\"></a>\n# **Model Validation**","metadata":{}},{"cell_type":"code","source":"# Performing a validation epoch\niterations = np.floor((y.shape[0]-split_idx) / batch_size).astype(int) \nloss,acc = 0,0 \nwith trange(iterations) as t: \n    for i in t:\n        start_idx = i * batch_size \n        x_batch = X[start_idx:start_idx+batch_size] \n        y_batch = y[start_idx:start_idx+batch_size] \n        metrics = model.test_on_batch(x_batch, y_batch) \n        loss = loss + metrics[0] \n        acc = acc + metrics[1] \n        t.set_description('Running training') \n        t.set_description('Running validation')\n        t.set_postfix(loss=\"%.2f\" % round(loss / (i+1),2),acc=\"%.2f\" % round(acc / (i+1),2))\n        \nprint(\"Validation loss:\",loss / iterations)\nprint(\"Validation accuracy:\",acc / iterations)","metadata":{"_uuid":"86b62f51aaedb77b00bfc39a830cd256c7c826d3","execution":{"iopub.status.busy":"2023-07-17T14:19:27.871945Z","iopub.execute_input":"2023-07-17T14:19:27.872432Z","iopub.status.idle":"2023-07-17T14:19:44.946762Z","shell.execute_reply.started":"2023-07-17T14:19:27.872374Z","shell.execute_reply":"2023-07-17T14:19:44.946115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code above tests the performance of our model on the validation set. As seen above, the validation accuracy is 0.915 and the loss is 0.237. Both metrics appear to be quite similar to the performance of the model on the training data. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"submission\"></a>\n# **Submission**","metadata":{"_uuid":"4726acfd86e61d1ad1c65fa0793204c2b128aafc"}},{"cell_type":"markdown","source":"The code below for creating a submission has been adapted from this [project](https://www.kaggle.com/code/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb/notebook).","metadata":{}},{"cell_type":"code","source":"# Clearing up RAM\nX = None\ny = None\ngc.collect();","metadata":{"_uuid":"3217bef913dde489dd23aebceaea82a4e88a1748","execution":{"iopub.status.busy":"2023-07-17T14:22:32.798870Z","iopub.execute_input":"2023-07-17T14:22:32.799229Z","iopub.status.idle":"2023-07-17T14:22:32.904128Z","shell.execute_reply.started":"2023-07-17T14:22:32.799170Z","shell.execute_reply":"2023-07-17T14:22:32.903170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating submission path\nbase_test_dir = path + 'test/' \ntest_files = glob(os.path.join(base_test_dir,'*.tif')) \nsubmission = pd.DataFrame() \nfile_batch = 5000 \nmax_idx = len(test_files) \nfor idx in range(0, max_idx, file_batch): \n    print(\"Indexes: %i - %i\"%(idx, idx+file_batch))\n    test_df = pd.DataFrame({'path': test_files[idx:idx+file_batch]}) \n    test_df['id'] = test_df.path.map(lambda x: x.split('/')[3].split(\".\")[0]) \n    test_df['image'] = test_df['path'].map(cv2.imread) \n    K_test = np.stack(test_df[\"image\"].values) \n    predictions = model.predict(K_test,verbose = 1) \n    test_df['label'] = predictions \n    submission = pd.concat([submission, test_df[[\"id\", \"label\"]]])\nsubmission.head() ","metadata":{"_uuid":"6d5d71c896efb4018434d00d1aab03d3355cb367","execution":{"iopub.status.busy":"2023-07-17T14:23:30.551004Z","iopub.execute_input":"2023-07-17T14:23:30.551364Z","iopub.status.idle":"2023-07-17T14:34:35.194274Z","shell.execute_reply.started":"2023-07-17T14:23:30.551303Z","shell.execute_reply":"2023-07-17T14:34:35.193269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a submission file\nsubmission.to_csv(\"submission.csv\", index = False, header = True) #create the submission file","metadata":{"_uuid":"a6113e9f267e0902a423490b75fa331f524e2aae","execution":{"iopub.status.busy":"2023-07-17T14:34:41.374445Z","iopub.execute_input":"2023-07-17T14:34:41.374761Z","iopub.status.idle":"2023-07-17T14:34:41.894894Z","shell.execute_reply.started":"2023-07-17T14:34:41.374699Z","shell.execute_reply":"2023-07-17T14:34:41.893952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"conclusion\"></a>\n# **Conclusion**","metadata":{}},{"cell_type":"code","source":"# Showing summary of the model used\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-07-17T14:38:31.486774Z","iopub.execute_input":"2023-07-17T14:38:31.487163Z","iopub.status.idle":"2023-07-17T14:38:31.504700Z","shell.execute_reply.started":"2023-07-17T14:38:31.487099Z","shell.execute_reply":"2023-07-17T14:38:31.504003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In conclusion, by using the model above with three epochs, I was able to yield an accuracy of 0.91 and a loss of 0.22 on the training data. On the validation data, the accuracy was roughly similar at 0.915 and the loss was 0.24. This project also yielded a public accuracy score of 0.9484 and a private score of 0.937. \n\nFor future works, it may be interesting to either alter the parameters and structure of the Keras model or to increase the number of epochs that were trained on. Regardless, I am quite happy with the score I have obtained, as ~95% accuracy is a great score. However, if this were applied to real life, I would assume that a higher accuracy would be preferred, so as not to incorrectly diagnose cancer tumors in a tissue sample. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"ref\"></a>\n# **References**","metadata":{}},{"cell_type":"markdown","source":"- https://www.kaggle.com/competitions/histopathologic-cancer-detection/overview\n- https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93\n- https://www.kaggle.com/code/abhinand05/histopathologic-cancer-detection-using-cnns#Loading-Data-and-EDA\n- https://www.kaggle.com/code/akarshu121/cancer-detection-with-cnn-for-beginners#4.-Validation-and-Analysis\n- https://www.kaggle.com/code/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb/notebook#Define-the-model\n- https://www.kaggle.com/code/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb/notebook","metadata":{}}]}