{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Why this topic?\n* Pneumonia accounts for over 15% of all deaths of children under 5 years old internationally. \n    * In 2015, 920,000 children under the age of 5 died from the disease.\n* In the United States, pneumonia accounts for over 500,000 visits to emergency departments and over 50,000 deaths in 2015, keeping the ailment on the list of top 10 causes of death in the country.\n![image](https://i.ibb.co/kMWtGYc/Lungs.png)","metadata":{}},{"cell_type":"markdown","source":"### Files & Folder descriptions\n* **`labels.csv`** - Contains \n    * `patientIds` \n    * `Target` \n    * `bounding box`  : Top left corner's `x`,`y` co-ordinates, it's `width` & `height`. \n* **`detailed_class_info.csv`** - Classifies each patientID into following category\n    * `Normal` - Health Lungs, \n    * `Lung Opacity` - Pneumonia Detected \n    * `No Lung Opacity / Not Normal` - Disease other than Pneumonia detected\n    \n* **`dataset`** - Contains 26.7K X-ray images in `*.dcm` format","metadata":{}},{"cell_type":"markdown","source":"## Importing dependencies","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom pydicom import dcmread\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport torch\nimport torch.nn as nn\nimport torchvision\nimport torchvision.transforms as transforms\nfrom torch.utils import data\nimport glob, pylab\nimport pydicom\n","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:14.212897Z","iopub.execute_input":"2021-06-14T22:17:14.213379Z","iopub.status.idle":"2021-06-14T22:17:16.706382Z","shell.execute_reply.started":"2021-06-14T22:17:14.213272Z","shell.execute_reply":"2021-06-14T22:17:16.705592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Importing & displaying labelled data from CSV File","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/dataset/labels.csv')\ndf.head(99999)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:16.710766Z","iopub.execute_input":"2021-06-14T22:17:16.711011Z","iopub.status.idle":"2021-06-14T22:17:16.782983Z","shell.execute_reply.started":"2021-06-14T22:17:16.710987Z","shell.execute_reply":"2021-06-14T22:17:16.782165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **`patientId`** : A `patientId`. Each patientId corresponds to a unique image.\n* **`Target`**   : a`Target`, either 0 or 1 for absence or presence of pneumonia, respectively\n* **`x`**        : the upper-left `x` coordinate of the bounding box. \n* **`y`**        : the upper-left `y` coordinate of the bounding box.\n* **`width`**    : the `width` of the bounding box.\n* **`height`**   : the `height` of the bounding box.\n1. Patient 0, the patient does not have pneumonia and so the corresponding bounding box information is set to NaN. \n1. Patient 4 is an example case with pnuemonia\n","metadata":{}},{"cell_type":"code","source":"df_detailed = pd.read_csv('../input/dataset/detailed_class_info.csv')\ndf_detailed.head(99999)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:16.784384Z","iopub.execute_input":"2021-06-14T22:17:16.784728Z","iopub.status.idle":"2021-06-14T22:17:16.834808Z","shell.execute_reply.started":"2021-06-14T22:17:16.784694Z","shell.execute_reply":"2021-06-14T22:17:16.833984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this project, the primary endpoint will be the detection of bounding boxes consisting of a binary classification---e.g. the presence or absence of pneumonia. However, in addition to the binary classification, each bounding box is further categorized into normal or no lung opacity / not normal. This extra third class indicates that while pneumonia was determined not to be present, there was nonetheless some type of abnormality on the image---and oftentimes this finding may mimic the appearance of true pneumonia. This extra class is provided as supplemental information to help improve algorithm accuracy if needed.","metadata":{}},{"cell_type":"code","source":"summary = {}\nfor n, row in df_detailed.iterrows():\n    if row['class'] not in summary:\n        summary[row['class']] = 0\n    summary[row['class']] += 1\n    \nprint(summary)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:16.837877Z","iopub.execute_input":"2021-06-14T22:17:16.838144Z","iopub.status.idle":"2021-06-14T22:17:19.109183Z","shell.execute_reply.started":"2021-06-14T22:17:16.83812Z","shell.execute_reply":"2021-06-14T22:17:19.108307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there is a relatively even split between the three classes, with nearly 1/3rd of the data comprising of pneumonia. Compared to most medical imaging datasets, where the prevalence of disease is quite low, this dataset has been significantly enriched with pathology.","metadata":{}},{"cell_type":"markdown","source":"## Overview of DICOM files and medical images\nMedical images are stored in a special format known as DICOM files (*.dcm). They contain a combination of header metadata as well as underlying raw image arrays for pixel data. In Python, one popular library to access and manipulate DICOM files is the pydicom module. ","metadata":{}},{"cell_type":"code","source":"patientId = df['patientId'][0]\ndcm_file = '../input/dataset/images/%s.dcm' % patientId\ndcm_data = pydicom.read_file(dcm_file)\n\nprint(dcm_data)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:19.112802Z","iopub.execute_input":"2021-06-14T22:17:19.113154Z","iopub.status.idle":"2021-06-14T22:17:19.124872Z","shell.execute_reply.started":"2021-06-14T22:17:19.113115Z","shell.execute_reply":"2021-06-14T22:17:19.123815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the standard headers containing patient identifable information have been anonymized (removed) so we are left with a relatively sparse set of metadata. The primary field we will be accessing is the underlying pixel data as follows:","metadata":{}},{"cell_type":"code","source":"im = dcm_data.pixel_array\nprint(type(im))\nprint(im.dtype)\nprint(im.shape)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:19.127344Z","iopub.execute_input":"2021-06-14T22:17:19.127784Z","iopub.status.idle":"2021-06-14T22:17:19.148826Z","shell.execute_reply.started":"2021-06-14T22:17:19.127742Z","shell.execute_reply":"2021-06-14T22:17:19.148002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Considerations**\n\nAs we can see here, the pixel array data is stored as a Numpy array, a powerful numeric Python library for handling and manipulating matrix data (among other things). In addition, it is apparent here that the original radiographs have been preprocessed for us as follows:\n\nThe relatively high dynamic range, high bit-depth original images have been rescaled to 8-bit encoding (256 grayscales). For the radiologists out there, this means that the images have been windowed and leveled already. In clinical practice, manipulating the image bit-depth is typically done manually by a radiologist to highlight certain disease processes. To visually assess the quality of the automated bit-depth downscaling and for considerations on potentially improving this baseline, consider consultation with a radiologist physician.\n\nThe relativley large original image matrices (typically acquired at >2000 x 2000) have been resized to the data-science friendly shape of 1024 x 1024. For the purposes of this challenge, the diagnosis of most pneumonia cases can typically be made at this resolution. To visually assess the feasibility of diagnosis at this resolution, and to determine the optimal resolution for pneumonia detection (oftentimes can be done at a resolution even smaller than 1024 x 1024), consider consultation with a radiogist physician.","metadata":{}},{"cell_type":"markdown","source":"## Visualizing Examples\n\n","metadata":{}},{"cell_type":"code","source":"label_data = pd.read_csv('../input/dataset/labels.csv')\ncolumns = ['patientId', 'Target']\nlabel_data = label_data.filter(columns)\n\ntrain_labels, test_labels = train_test_split(label_data.values, test_size=0.1)\nimages_f = '../input/dataset/images'\ntrain_paths = [os.path.join(images_f, image[0]) for image in train_labels]\n\ndef imshow(num_to_show=12):\n    \n    plt.figure(figsize=(20,15))\n    \n    for i in range(num_to_show):\n        plt.subplot(3, 4, i+1)\n        plt.grid(True)\n        plt.xticks([])\n        plt.yticks([])\n        \n        img_dcm = dcmread(f'{train_paths[i+20]}.dcm')\n        img_np = img_dcm.pixel_array\n        plt.imshow(img_np, cmap=plt.cm.binary)\n        plt.xlabel(train_labels[i+20][1])\n\nimshow()","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:19.150447Z","iopub.execute_input":"2021-06-14T22:17:19.150814Z","iopub.status.idle":"2021-06-14T22:17:20.744615Z","shell.execute_reply.started":"2021-06-14T22:17:19.150773Z","shell.execute_reply":"2021-06-14T22:17:20.743578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploring the Data and Labels\n\nAny given patient may potentially have many boxes if there are several different suspicious areas of pneumonia. To collapse the current CSV file dataframe into a dictionary with unique entries, we considered the following method:","metadata":{}},{"cell_type":"code","source":"def parse_data(df):\n \n    # --- Define lambda to extract coords in list [y, x, height, width]\n    extract_box = lambda row: [row['y'], row['x'], row['height'], row['width']]\n\n    parsed = {}\n    for n, row in df.iterrows():\n        # --- Initialize patient entry into parsed \n        pid = row['patientId']\n        if pid not in parsed:\n            parsed[pid] = {\n                'dicom': '../input/dataset/images/%s.dcm' % pid,\n                'label': row['Target'],\n                'boxes': []}\n\n        # --- Add box if opacity is present\n        if parsed[pid]['label'] == 1:\n            parsed[pid]['boxes'].append(extract_box(row))\n\n    return parsed\n\nparsed = parse_data(df)\nprint(parsed['00436515-870c-4b36-a041-de91049b9ab4'])","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:20.746301Z","iopub.execute_input":"2021-06-14T22:17:20.746663Z","iopub.status.idle":"2021-06-14T22:17:23.175223Z","shell.execute_reply.started":"2021-06-14T22:17:20.746623Z","shell.execute_reply":"2021-06-14T22:17:23.173684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Boxes\nIn order to overlay color boxes on the original grayscale DICOM files, consider using the following methods (below, the main method `draw()` requires the method `overlay_box()`):","metadata":{}},{"cell_type":"code","source":"def draw(data):\n    \"\"\"\n    Method to draw single patient with bounding box(es) if present \n\n    \"\"\"\n    # --- Open DICOM file\n    d = pydicom.read_file(data['dicom'])\n    im = d.pixel_array\n\n    # --- Convert from single-channel grayscale to 3-channel RGB\n    im = np.stack([im] * 3, axis=2)\n\n    # --- Add boxes with random color if present\n    for box in data['boxes']:\n        rgb = np.floor(np.random.rand(3) * 256).astype('int')\n        im = overlay_box(im=im, box=box, rgb=rgb, stroke=6)\n\n    pylab.imshow(im, cmap=pylab.cm.gist_gray)\n    pylab.axis('off')\n\ndef overlay_box(im, box, rgb, stroke=1):\n    \"\"\"\n    Method to overlay single box on image\n\n    \"\"\"\n    # --- Convert coordinates to integers\n    box = [int(b) for b in box]\n    \n    # --- Extract coordinates\n    y1, x1, height, width = box\n    y2 = y1 + height\n    x2 = x1 + width\n\n    im[y1:y1 + stroke, x1:x2] = rgb\n    im[y2:y2 + stroke, x1:x2] = rgb\n    im[y1:y2, x1:x1 + stroke] = rgb\n    im[y1:y2, x2:x2 + stroke] = rgb\n\n    return im","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.176819Z","iopub.execute_input":"2021-06-14T22:17:23.17718Z","iopub.status.idle":"2021-06-14T22:17:23.186285Z","shell.execute_reply.started":"2021-06-14T22:17:23.177141Z","shell.execute_reply":"2021-06-14T22:17:23.185388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"draw(parsed['00436515-870c-4b36-a041-de91049b9ab4'])","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.187576Z","iopub.execute_input":"2021-06-14T22:17:23.18809Z","iopub.status.idle":"2021-06-14T22:17:23.353041Z","shell.execute_reply.started":"2021-06-14T22:17:23.188049Z","shell.execute_reply":"2021-06-14T22:17:23.352193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring non-Pneumonia disease\nAs we see from the 1st table first patient i.e patient[0] does not have pneumonia however does have another imaging abnormality present. Let's take a closer look:","metadata":{}},{"cell_type":"code","source":"patientId = df_detailed['patientId'][0]\ndraw(parsed[patientId])","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.354247Z","iopub.execute_input":"2021-06-14T22:17:23.354751Z","iopub.status.idle":"2021-06-14T22:17:23.511102Z","shell.execute_reply.started":"2021-06-14T22:17:23.354712Z","shell.execute_reply":"2021-06-14T22:17:23.510274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While the image displayed inline within the notebook is small, as a radiologist it is evident that the patient has several well circumscribed nodular densities in the left lung (right side of image). In addition there is a large chest tube in the right lung (left side of the image) which has been placed to drain fluid accumulation (e.g. pleural effusion) at the right lung base that also demonstrates overlying patchy densities (e.g. possibly atelectasis or partial lung collapse).\n\nAs you can see, there are a number of abnormalities on the image, and the determination that none of these findings correlate to pneumonia is somewhat subjective even among expert physicians. Therefore, as is almost always the case in medical imaging datasets, the provided ground-truth labels are far from 100% objective. We have kept this in mind as we develop our algorithm.","metadata":{}},{"cell_type":"markdown","source":"## Dividing labels for train and testing set","metadata":{}},{"cell_type":"code","source":"label_data = pd.read_csv('../input/dataset/labels.csv')\ncolumns = ['patientId', 'Target']\nlabel_data = label_data.filter(columns)\n\ntrain_labels, test_labels = train_test_split(label_data.values, test_size=0.1)\nprint(\"Training labels: \",train_labels.shape)\nprint(\"Training labels: \",test_labels.shape)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.512629Z","iopub.execute_input":"2021-06-14T22:17:23.512983Z","iopub.status.idle":"2021-06-14T22:17:23.551229Z","shell.execute_reply.started":"2021-06-14T22:17:23.512947Z","shell.execute_reply":"2021-06-14T22:17:23.550425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Composing transformations","metadata":{}},{"cell_type":"code","source":"transform = transforms.Compose([\n    transforms.RandomHorizontalFlip(),\n    transforms.Resize(224),\n    transforms.ToTensor()])","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.5528Z","iopub.execute_input":"2021-06-14T22:17:23.553138Z","iopub.status.idle":"2021-06-14T22:17:23.557721Z","shell.execute_reply.started":"2021-06-14T22:17:23.553102Z","shell.execute_reply":"2021-06-14T22:17:23.556734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Write a custom dataset ","metadata":{}},{"cell_type":"code","source":"class Dataset(data.Dataset):\n    \n    def __init__(self, paths, labels, transform=None):\n        self.paths = paths\n        self.labels = labels\n        self.transform = transform\n    \n    def __getitem__(self, index):\n        image = dcmread(f'{self.paths[index]}.dcm')\n        image = image.pixel_array\n        image = image / 255.0\n\n        image = (255*image).clip(0, 255).astype(np.uint8)\n        image = Image.fromarray(image).convert('RGB')\n\n        label = self.labels[index][1]\n        \n        if self.transform is not None:\n            image = self.transform(image)\n            \n        return image, label\n    \n    def __len__(self):\n        \n        return len(self.paths)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.559413Z","iopub.execute_input":"2021-06-14T22:17:23.559809Z","iopub.status.idle":"2021-06-14T22:17:23.569117Z","shell.execute_reply.started":"2021-06-14T22:17:23.559762Z","shell.execute_reply":"2021-06-14T22:17:23.568177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare training and validation dataloader","metadata":{}},{"cell_type":"code","source":"images_f = '../input/dataset/images'\n\ntrain_paths = [os.path.join(images_f, image[0]) for image in train_labels]\ntest_paths = [os.path.join(images_f, image[0]) for image in test_labels]\n\nprint(len(train_paths))\nprint(len(test_paths))","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.570708Z","iopub.execute_input":"2021-06-14T22:17:23.571136Z","iopub.status.idle":"2021-06-14T22:17:23.644789Z","shell.execute_reply.started":"2021-06-14T22:17:23.571096Z","shell.execute_reply":"2021-06-14T22:17:23.644004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset = Dataset(train_paths, train_labels, transform=transform)\ntest_dataset = Dataset(test_paths, test_labels, transform=transform)\ntrain_loader = data.DataLoader(dataset=train_dataset, batch_size=50, shuffle=True)\ntest_loader = data.DataLoader(dataset=test_dataset, batch_size=50, shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.64622Z","iopub.execute_input":"2021-06-14T22:17:23.646563Z","iopub.status.idle":"2021-06-14T22:17:23.651846Z","shell.execute_reply.started":"2021-06-14T22:17:23.64653Z","shell.execute_reply":"2021-06-14T22:17:23.650857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Specify device","metadata":{}},{"cell_type":"code","source":"device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(device)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.653686Z","iopub.execute_input":"2021-06-14T22:17:23.654096Z","iopub.status.idle":"2021-06-14T22:17:23.717622Z","shell.execute_reply.started":"2021-06-14T22:17:23.65406Z","shell.execute_reply":"2021-06-14T22:17:23.716404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Loading ResNet50 and fine-tune","metadata":{}},{"cell_type":"code","source":"\nmodel = torchvision.models.resnet50(pretrained=True)\nnum_ftrs = model.fc.in_features\n# Here the size of each output sample is set to 2.\n# Alternatively, it can be generalized to nn.Linear(num_ftrs, len(class_names)).\nmodel.fc = nn.Linear(num_ftrs, 2)\n\nmodel.to(device)\n\ncriterion = nn.CrossEntropyLoss()\n\n# Observe that all parameters are being optimized\noptimizer = torch.optim.SGD(model.parameters(), lr=0.001, momentum=0.9)\n\n# Decay LR by a factor of 0.1 every 7 epochs\nexp_lr_scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=7, gamma=0.1)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:23.718877Z","iopub.execute_input":"2021-06-14T22:17:23.719246Z","iopub.status.idle":"2021-06-14T22:17:32.057158Z","shell.execute_reply.started":"2021-06-14T22:17:23.719208Z","shell.execute_reply":"2021-06-14T22:17:32.056394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(model)","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:32.058436Z","iopub.execute_input":"2021-06-14T22:17:32.058785Z","iopub.status.idle":"2021-06-14T22:17:32.065287Z","shell.execute_reply.started":"2021-06-14T22:17:32.058749Z","shell.execute_reply":"2021-06-14T22:17:32.064433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training NN on dataset","metadata":{}},{"cell_type":"code","source":"num_epochs = 1\nbreaking = 0;\n# Train the model\ntotal_step = len(train_loader)\nfor epoch in range(num_epochs):\n    # Training step\n    for i, (images, labels) in tqdm(enumerate(train_loader)):\n        images = images.to(device)\n        labels = labels.to(device)\n        print (labels)\n        # Forward pass\n        outputs = model(images)\n        loss = criterion(outputs, labels)\n        if breaking >= 120:\n            break\n        breaking += 1\n        # Backward and optimize\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n        \n        if (i+1) % 2000 == 0:   \n            print(\"Epoch [{}/{}], Step [{}/{}], Loss: {:.4f}\"\n                   .format(epoch+1, num_epochs, i+1, total_step, loss.item()))\n    ","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:17:32.066574Z","iopub.execute_input":"2021-06-14T22:17:32.067063Z","iopub.status.idle":"2021-06-14T22:18:59.061071Z","shell.execute_reply.started":"2021-06-14T22:17:32.067025Z","shell.execute_reply":"2021-06-14T22:18:59.060042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Testing","metadata":{}},{"cell_type":"code","source":"model.eval()\n\ncorrect = 0\ntotal = 0  \nbreaking = 1;\ntrue_positive = 0;\ntrue_negative = 0;\nfalse_positive = 0;\nfalse_negative = 0;\npredicted_labels = []\ncorrect_labels = []\nfor images, labels in tqdm(test_loader):\n    images = images.to(device)\n    labels = labels.to(device)\n    predictions = model(images)\n    _, predicted = torch.max(predictions, 1)\n    total += labels.size(0)\n    correct += (labels == predicted).sum()\n    for x in enumerate(predicted):\n        predicted_labels.append(x[1].item())\n    for x in enumerate(labels):\n        correct_labels.append(x[1].item())\n    if breaking >= 60:\n        break\n    breaking += 1\n\nfor predicted, actual in zip(predicted_labels, correct_labels):\n    predicted = bool(predicted)\n    actual = bool(actual)\n  \n    if predicted and actual:\n        true_positive += 1\n    if predicted and  not actual:\n        false_positive += 1\n    if not predicted and actual:\n        false_negative += 1\n    if not predicted and not actual:\n        true_negative += 1\n\ntotal = true_positive + true_negative + false_positive + false_negative\n\nprint(f'Recall/Sensitivity: {100 * true_positive/(true_positive + false_negative)}')\nprint(f'Specificity: {100 * true_negative/(true_negative + false_positive)}')\nprint(f'Accu: {100 * (true_positive + true_negative)/total}')\n\n","metadata":{"execution":{"iopub.status.busy":"2021-06-14T22:18:59.062721Z","iopub.execute_input":"2021-06-14T22:18:59.063383Z","iopub.status.idle":"2021-06-14T22:19:33.346564Z","shell.execute_reply.started":"2021-06-14T22:18:59.063328Z","shell.execute_reply":"2021-06-14T22:19:33.345595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*  **Recall**: The higher the numerical value of **sensitivity**, the less likely diagnos tic test returns false-positive results. For example, if sensitivity = 99%, it means: when we conduct a diagnostic test on a patient with certain disease, there is 99% of chance, this patient will be identified as positive. A test with high sensitivity tents to capture all possible positive conditions without missing anyone. **Thus a test with high sensitivity is often used to screen for disease**. \n* The numerical value of **specificity** represents the probability of a test diagnoses a particular disease **without giving false-positive results**. For example, if the specificity of a test is 99%. It means: when we conduct a diagnostic test on a patient without certain disease, there is 99% chance; this patient will be identified as negative.","metadata":{}}]}