{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Using Deep Learning to detect COVID-19 presence from x-ray scans"},{"metadata":{},"cell_type":"markdown","source":"Disclaimer: I am not a doctor nor a medical researcher. This work is only intended as a source of inspiration for further studies."},{"metadata":{},"cell_type":"markdown","source":"## COVID-19-xray\n"},{"metadata":{},"cell_type":"markdown","source":"The more the pandemic crisis progresses, the more it gets important that countries perform tests to help understand and stop the spread of COVID-19.\nUnfortunately, the capacity for COVID-19 testing is still low in many countries. "},{"metadata":{},"cell_type":"markdown","source":"### How are tests performed?"},{"metadata":{},"cell_type":"markdown","source":"The standard COVID-19 tests are called PCR (Polymerase chain reaction) tests. This family of tests looks for the existence of antibodies of a given infection. Two main issues with this test are:\n\n1. a shortage a tests available worldwide\n2. a patient might be carring the virus without having symptoms. In this case the test fails to identify infected patients"},{"metadata":{},"cell_type":"markdown","source":"[Dr. Joseph Paul Cohen, Postdoctoral Fellow at University of Montreal](https://josephpcohen.com/w/), recently open sourced a [database](https://github.com/ieee8023/covid-chestxray-dataset) containing chest x-ray pictures of patients suffering from the COVID-19 disease. \n"},{"metadata":{},"cell_type":"markdown","source":"The database only contains pictures of patients suffering from COVID-19. In order to build a classifier for xray images we first need to find similar x-ray images of people who are not suffering from the disease.\nIt turns out Kaggle has a database with chest x-ray images of patients suffering of pneumonia and healthy patients. Hence, we are going to use both sources images in our dataset.\n"},{"metadata":{},"cell_type":"markdown","source":"The notebook is organized as follows:\n\n1. Data Preparation  \n\n2. Train Network using Fastai\n\n3. Optimize Network\n\n4. Test\n\n5. What's Next"},{"metadata":{},"cell_type":"markdown","source":"But first let's import necessary libraries"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1. Data Preparation"},{"metadata":{},"cell_type":"markdown","source":"Let's import Fastai, create useful paths and create covid_df"},{"metadata":{"trusted":true},"cell_type":"code","source":"from fastai import *\nfrom fastai.vision import *\n\n# useful paths\ninput_path = Path('/kaggle/input') \ncovid_xray_path = input_path/'xraycovid'\npneumonia_path = input_path/'chest-xray-pneumonia/chest_xray'\n\ncovid_df = pd.read_csv(covid_xray_path/'metadata.csv')\ncovid_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We notice straight away that we have a large number of NaN, let's remove them and see what we are left with"},{"metadata":{"trusted":true},"cell_type":"code","source":"covid_df.dropna(axis=1,inplace=True)\ncovid_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That looks better. We are mainly interested in two columns: ```finding``` and ```filename```. The former tells us wether or not a patient is suffering from the virus whereas the latter tells us the finename. The other interesting column is ```view```. It turns out the view is the angle used when the scan is taken and the most frequently used is PA. PA view stands for Posteroanterior view."},{"metadata":{"trusted":true},"cell_type":"code","source":"covid_df.groupby('view').count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"covid_df.groupby('finding').count()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"PA makes up the majority of the datapoints. Let's keep them and remove the rest."},{"metadata":{"trusted":true},"cell_type":"code","source":"covid_df = covid_df[lambda x: x['view'] == 'PA']\ncovid_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For simplicity, let's also rename the elements in column ```finding``` to be ```positive``` if the patient is suffering from COVID-19 and negative otherwise."},{"metadata":{"trusted":true},"cell_type":"code","source":"covid_df['finding'] = covid_df['finding'].apply(lambda x:'positive' if x == 'COVID-19' else 'negative')\ncovid_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finally, let's replace the ```filename``` column by the entire system path and keep only the two columns we are more interested in"},{"metadata":{"trusted":true},"cell_type":"code","source":"def makeFilename(x = ''):\n    return covid_xray_path/f'images/{x}'\n\ncovid_df['filename'] = covid_df['filename'].apply(makeFilename)\ncovid_df = covid_df[['finding', 'filename']]\ncovid_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We now need to create a dataframe of the same format using the pictures from the other database. Once we have that dataframe, we can use the mighty [ImageDataBunch](https://docs.fast.ai/vision.data.html) methods to create a dataset that we can feed to our convolutional network.  \n\nSince our second database is made up of pictures of both healthy patients and pneumonia suffering patients, we are going to take an equal mix of both. I tried using only images of healthy people from this database but I reflected that since COVID-19 and pneumonia are linked somehow then it might give our network an edge to also contain pneumonia x-rays.\n\nThis is what our ```pneumonia_df``` looks like:"},{"metadata":{"trusted":true},"cell_type":"code","source":"\npneumonia_df = pd.DataFrame([], columns=['finding', 'filename'])\nfolders = ['train/NORMAL', 'val/NORMAL', 'test/NORMAL']\nfor folder in folders:\n    fnames = get_image_files(pneumonia_path/folder)\n    fnames = map(lambda x: ['negative', x], fnames)\n    df = pd.DataFrame(fnames, columns=['finding', 'filename'])\n    pneumonia_df = pneumonia_df.append(df, ignore_index = True)\n\n\n\nfolders = ['train/PNEUMONIA', 'val/PNEUMONIA', 'test/PNEUMONIA']\nfor folder in folders:\n    fnames = get_image_files(pneumonia_path/folder)\n    fnames = map(lambda x: ['negative', x], fnames)\n    df = pd.DataFrame(fnames, columns=['finding', 'filename'])\n    pneumonia_df = pneumonia_df.append(df, ignore_index = True)\n\npneumonia_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As you can see we have 5856 pictures which is about 60 times larger than our covid_df.  \n\nSince we have 92 pictures in our covid_df, I decided to take an equal number of pictures of healthy patients and an equal number of picture of pneumonia patients. In other words, 92 covid_df images, 92 healthy patient images, and 92 pneumonia affected patients. As far as our analysis goes, we are really only interested in covid positive and covid negative. Therefore, both the healthy and pneumonia patients will be labeled as ```negative```"},{"metadata":{},"cell_type":"markdown","source":"NB: Following great suggestions, I received, I am gonna run the Convolutional Net on two dataBunch:\n\n- The first will have covid_df and healthy_df\n- The second one will have covid_df and pneumonia_df\n\nWe will then compare the perfomances and, hopefully, we will get comparable results so that we can have more confidence in our results.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"\nhealthy_df = pd.DataFrame([], columns=['finding', 'filename'])\nfolders = ['train/NORMAL', 'val/NORMAL', 'test/NORMAL']\nfor folder in folders:\n    fnames = get_image_files(pneumonia_path/folder)\n    fnames = map(lambda x: ['negative', x], fnames)\n    df = pd.DataFrame(fnames, columns=['finding', 'filename'])\n    healthy_df = healthy_df.append(df, ignore_index = True)\n    \npneumonia_df = pd.DataFrame([], columns=['finding', 'filename'])\nfolders = ['train/PNEUMONIA', 'val/PNEUMONIA', 'test/PNEUMONIA']\nfor folder in folders:\n    fnames = get_image_files(pneumonia_path/folder)\n    fnames = map(lambda x: ['negative', x], fnames)\n    df = pd.DataFrame(fnames, columns=['finding', 'filename'])\n    pneumonia_df = pneumonia_df.append(df, ignore_index = True)\n\npneumonia_df = pneumonia_df.sample(covid_df.shape[0]).reset_index(drop=True)\n\nhealthy_df = healthy_df.sample(covid_df.shape[0]).reset_index(drop=True)\n\nnegative_df = healthy_df.append(pneumonia_df, ignore_index = True)\n\nnegative_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, we can finally merge our dataframes to get the dataframe needed to build our [ImageDataBunch](https://docs.fast.ai/vision.data.html)."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = covid_df.append(healthy_df, ignore_index = True)\ndf = df.sample(frac=1).reset_index(drop=True)\ndf.sample(20)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. Train Network using Fastai"},{"metadata":{},"cell_type":"markdown","source":"I am going to run the Convolutional Net using two training sets.\nThe first will have  covid_df and healthy_df\nThe second one will have covid_df and pneumonia_df\n\nWe will then compare the perfomances and, hopefully, we will get comparable results so that we can have more confidence in our results."},{"metadata":{},"cell_type":"markdown","source":"## First case: COVID-19 patients and healthy patients"},{"metadata":{},"cell_type":"markdown","source":"We are now ready to create the ImageDataBunch."},{"metadata":{"trusted":true},"cell_type":"code","source":"np.random.seed(42)\ndata = ImageDataBunch.from_df(\n        '/', \n        df, \n        fn_col='filename',\n        label_col='finding',\n        ds_tfms=get_transforms(), ## data augmentation: flip horizozntally\n        size=224, \n        num_workers=4\n    ).normalize(imagenet_stats)\n\ndata\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's take random images from ```data``` to see if they look consistent."},{"metadata":{"trusted":true},"cell_type":"code","source":"data.show_batch(rows=80, figsize=(21,21))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"To my untrained eyes, it looks like the images look consistent. We are going to use a resnet50 and leverage Kaggle free GPU Quota. Let's start training ten cycles."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn = cnn_learner(data, models.resnet50, metrics=error_rate)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's first fit 10 cycles and see how it improves"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like we can do better, let's run ten cycles more."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It looks good. We are going to keep the 5.5% error for now and try the next data frame\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.save('stage-1')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Second case: COVID-19 patients and pneumonia patients"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"df2 = covid_df.append(pneumonia_df, ignore_index = True)\ndf2 = df2.sample(frac=1).reset_index(drop=True)\nnp.random.seed(42)\ndata2 = ImageDataBunch.from_df(\n        '/', \n        df2, \n        fn_col='filename',\n        label_col='finding',\n        ds_tfms=get_transforms(), ## data augmentation: flip horizozntally\n        size=224, \n        num_workers=4\n    ).normalize(imagenet_stats)\n\nlearn2 = cnn_learner(data2, models.resnet50, metrics=error_rate)\nlearn2.fit_one_cycle(5)"},{"metadata":{},"cell_type":"markdown","source":"Let's run some more cycles"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.fit_one_cycle(10)"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.fit_one_cycle(10)"},{"metadata":{},"cell_type":"markdown","source":"That's a nice error rate. Let's save ```learn2``` and starti optimizing"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.save('learn2-stage-1')\n"},{"metadata":{},"cell_type":"markdown","source":"# 3. Optmize"},{"metadata":{},"cell_type":"markdown","source":"Results for the first case were already pretty solid in **Part 2**. We are going to first optimize the results for the first case and then optimize results for the second case. Then, if the two cases accuracy do not differ too much, we will be confident in our result and try to predict random images online."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.load('stage-1')\nlearn.unfreeze()\nlearn.lr_find()\nlearn.recorder.plot()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" The longest downward shape is found in the region around ```1e-4``` let's use that as our starting point"},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(10, max_lr=slice(7e-5,2e-4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I obtained the 0% error rate after a updated my notebook on kaggle and used a balanced dataset. This error rate though, is probably due to the fact that I am still collecting data and would require much more images to have an more stable error rate."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.save('stage-2')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(2, max_lr=slice(7e-5,2e-4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like the error rate is not really moving. With 3.6% error rate we might be satisfied with this first results. We are going to save and plot the confusion matrix."},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.load('stage-2')\ninterp = ClassificationInterpretation.from_learner(learn)\ninterp.plot_confusion_matrix()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Since the error rate is 0, the confusion matrix shows we have no errors."},{"metadata":{},"cell_type":"markdown","source":"## Second case: COVID-19 patients and pneumonia patients"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.load('learn2-stage-1')\n\nlearn2.unfreeze()\n\nlearn2.lr_find()\n\nlearn2.recorder.plot()\n"},{"metadata":{},"cell_type":"markdown","source":" The longest downward shape is found in the region around ???????????```1e-4``` let's use that as our starting point"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.load('learn2-stage-1')\nlearn2.unfreeze()\nlearn2.fit_one_cycle(4, max_lr=slice(7e-5,1e-4))"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"learn2.save('learn2-stage-2')"},{"metadata":{},"cell_type":"markdown","source":"Both cases have an error rate < 3%. Given the scarsity of data, this is a promising first result. Since using both models, covid-19 prediction seems to be consistent, we can be confident enough in its predictions."},{"metadata":{},"cell_type":"markdown","source":"# Test"},{"metadata":{"trusted":true},"cell_type":"code","source":"# learn = cnn_learner(data, models.resnet50, metrics=error_rate)\n# learn.load('stage-3')\n# interp = ClassificationInterpretation.from_learner(learn)\nimg = open_image(input_path/'testimg/df1053d3e8896b53ef140773e10e26_gallery.jpeg')\nlearn.predict(img)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That image is taken from https://radiopaedia.org/images/52197348 and it is an image of a positive patient."},{"metadata":{},"cell_type":"markdown","source":"# What's next?"},{"metadata":{},"cell_type":"markdown","source":"First of all, I would like to incorporate scans from other sources and see if accuracy and generalization might increase.\nToday, while I was about to pusblish this article, I found out that [MIT](https://www.technologyreview.com/s/615399/coronavirus-neural-network-can-help-spot-covid-19-in-chest-x-ray-pneumonia/) has released a database containing xrays images of covid patients. Next, I am going to incorporate MIT's database and see where we get."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}