{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This is a case of Image data augmentation.\nImage data augmentation is a technique that can be used to artificially expand the size of a training dataset by creating modified versions of images in the dataset.\n\nTraining deep learning neural network models on more data can result in more skillful models, and the augmentation techniques can create variations of the images that can improve the ability of the fit models to generalize what they have learned to new images.\n\nThe Keras deep learning neural network library provides the capability to fit models using image data augmentation via the ImageDataGenerator class.\n how to improve image classification\nby using data augmentation and convolutional neural networks. Model\noverfitting and poor performance are common problems in applying neural network techniques. Approaches to bring intra-class differences down\nand retain sensitivity to the inter-class variations are important to maximize model accuracy and minimize the loss function. The image dataset, the effects of model overfitting were monitored\nwithin different model architectures in combination of data augmentation and hyper-parameter tuning. The model performance was evaluated\nwith train and test accuracy and loss, characteristics derived from the\nconfusion matrices, and visualizations of different model outputs. As a macro-architecture with 500 weighted layers, \nmodel is used for large scale image classification. In the presence of image\ndata augmentation, the overall model train accuracy is 96%, the\ntest accuracy is stabilized at 92%, and both the results of train and test\nlosses are below 0.5. The overall image classification error rate is dropped\nto 8%, while the single class misclassification rates are less than 7.5% in\neight out of ten image classes. Model architecture, hyper-parameter tuning, and data augmentation are essential to reduce model overfitting and\nhelp build a more reliable convolutional neural network model.","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport matplotlib.pyplot as plt\nimport pydicom\nfrom skimage.transform import resize\n","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:53.500045Z","iopub.execute_input":"2021-06-30T12:40:53.500476Z","iopub.status.idle":"2021-06-30T12:40:54.629278Z","shell.execute_reply.started":"2021-06-30T12:40:53.500389Z","shell.execute_reply":"2021-06-30T12:40:54.628483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading the csv to fetch the data and ","metadata":{}},{"cell_type":"code","source":"submission_df = pd.read_csv('/kaggle/input/siim-covid19-detection/sample_submission.csv', index_col=None)\nimage_df = pd.read_csv('/kaggle/input/siim-covid19-detection/train_image_level.csv', index_col=None)\nstudy_df = pd.read_csv('/kaggle/input/siim-covid19-detection/train_study_level.csv', index_col=None)\n\nprint(f\"Train image level csv shape : {image_df.shape}\")\nprint(f\"Train study level csv shape : {study_df.shape}\")","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.631464Z","iopub.execute_input":"2021-06-30T12:40:54.631813Z","iopub.status.idle":"2021-06-30T12:40:54.700780Z","shell.execute_reply.started":"2021-06-30T12:40:54.631777Z","shell.execute_reply":"2021-06-30T12:40:54.700067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Data to plot\nlabels = 'Negative for Pneumonia','Typical Appearance','Indeterminate Appearance','Atypical Appearance'\nsizes = [study_df['Negative for Pneumonia'].sum(), study_df['Typical Appearance'].sum(), study_df['Indeterminate Appearance'].sum(), study_df['Atypical Appearance'].sum()]\ncolors = ['gold', 'yellowgreen', 'lightcoral', 'lightskyblue']\nexplode = (0.1, 0, 0, 0)  # explode 1st slice\n\n# Plot\nplt.pie(sizes, explode=explode, labels=labels, colors=colors,\nautopct='%1.1f%%', shadow=True, startangle=140)\n\nplt.axis('equal')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.703704Z","iopub.execute_input":"2021-06-30T12:40:54.703959Z","iopub.status.idle":"2021-06-30T12:40:54.856630Z","shell.execute_reply.started":"2021-06-30T12:40:54.703934Z","shell.execute_reply":"2021-06-30T12:40:54.855891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above plot, it is clear that the typical appearance has highest number at the same time. Atypical Appearance has the lowest percentage","metadata":{}},{"cell_type":"code","source":"study_df","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.857987Z","iopub.execute_input":"2021-06-30T12:40:54.858474Z","iopub.status.idle":"2021-06-30T12:40:54.877225Z","shell.execute_reply.started":"2021-06-30T12:40:54.858432Z","shell.execute_reply":"2021-06-30T12:40:54.876130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.880857Z","iopub.execute_input":"2021-06-30T12:40:54.881219Z","iopub.status.idle":"2021-06-30T12:40:54.893323Z","shell.execute_reply.started":"2021-06-30T12:40:54.881185Z","shell.execute_reply":"2021-06-30T12:40:54.892120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_df","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.896687Z","iopub.execute_input":"2021-06-30T12:40:54.897080Z","iopub.status.idle":"2021-06-30T12:40:54.910712Z","shell.execute_reply.started":"2021-06-30T12:40:54.897044Z","shell.execute_reply":"2021-06-30T12:40:54.909765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"i=0\nj=0\nk=0\nl=0\nm=0\ntrain_path=[]\nfor dirname, _, filenames in os.walk('/kaggle/input/siim-covid19-detection/train/'):\n    for filename in filenames:\n        str_split =filename.split('.')\n        if str_split[1]==\"dcm\" :\n            df=image_df[image_df['id'].str.contains(str_split[0])]\n            studydf=study_df[study_df.id.isin(df.StudyInstanceUID+'_study')]\n            lbl_folder=studydf.columns[studydf.eq(1).any()] \n            if lbl_folder=='Atypical Appearance':\n                if i<10:\n                    train_path.append(os.path.join(dirname, filename))\n                    i=i+1\n            elif lbl_folder=='Indeterminate Appearance':\n                if j<10:\n                    train_path.append(os.path.join(dirname, filename))\n                    j=j+1\n            elif lbl_folder=='Negative for Pneumonia':\n                if k<10:\n                    train_path.append(os.path.join(dirname, filename))\n                    k=k+1\n            elif lbl_folder=='Typical Appearance':\n                if l<10:\n                    train_path.append(os.path.join(dirname, filename))\n                    l=l+1\n            else :\n                if m<10:\n                    train_path.append(os.path.join(dirname, filename))\n                    m=m+1","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:40:54.912396Z","iopub.execute_input":"2021-06-30T12:40:54.912811Z","iopub.status.idle":"2021-06-30T12:41:52.599153Z","shell.execute_reply.started":"2021-06-30T12:40:54.912775Z","shell.execute_reply":"2021-06-30T12:41:52.598393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for rows in train_path:\n    # specify your image path\n    try:\n        ct_dicom = pydicom.read_file(rows)\n        plt.imshow(ct_dicom.pixel_array, cmap='gray')\n        plt.show()\n    except:\n          pass\n    \n    \n  ","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:41:52.600596Z","iopub.execute_input":"2021-06-30T12:41:52.600930Z","iopub.status.idle":"2021-06-30T12:42:39.918780Z","shell.execute_reply.started":"2021-06-30T12:41:52.600896Z","shell.execute_reply":"2021-06-30T12:42:39.917892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\nimport os\nfrom skimage.transform import resize","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:42:39.920079Z","iopub.execute_input":"2021-06-30T12:42:39.920590Z","iopub.status.idle":"2021-06-30T12:42:40.093667Z","shell.execute_reply.started":"2021-06-30T12:42:39.920552Z","shell.execute_reply":"2021-06-30T12:42:40.092644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.patches as patches\nbox=[]\nfor rows in train_path:\n    #rows=\"/kaggle/input/siim-covid19-detection/train/ea44061da219/bd68195b6b27/a818e98c3f90.dcm\"\n    imgname=rows.rsplit('/', 1)[1]\n    str_split =imgname.split('.')\n    imagename=str_split[0]\n    df=image_df[image_df['id'].str.contains(imagename,regex=True)]\n    data_items = df['boxes'].items()\n    data_list = list(data_items)\n    boxes_df = pd.DataFrame(data_list)\n    if boxes_df.empty:\n        print('DataFrame is empty!')\n    else:\n        boxes_df[1].fillna(0, inplace=True)\n        if boxes_df[1].get(0) != 0:\n            ss=boxes_df[1].get(0).split('},')\n            x1=float(ss[0].split(',')[0].split(':')[1].strip())\n            y1=float(ss[0].split(',')[1].split(':')[1].strip())\n            width1=float(ss[0].split(',')[2].split(':')[1].strip())\n            if '}]' in ss[0].split(',')[3].split(':')[1].strip() :\n              height1=float(ss[0].split(',')[3].split(':')[1].strip().replace(\"}]\", \"\"))\n              box = [\n                       {'x' : x1, 'y' : y1, 'width' : width1, 'height' : height1},\n                    ]\n            else:\n              height1=float(ss[0].split(',')[3].split(':')[1].strip()) \n              x2=float(ss[1].split(',')[0].split(':')[1].strip())\n              y2=float(ss[1].split(',')[1].split(':')[1].strip())\n              width2=float(ss[1].split(',')[2].split(':')[1].strip())\n              height2=float(ss[1].split(',')[3].split(':')[1].replace('}]', '').replace(\"'\", \"\").strip())\n              box = [\n                      {'x' : x1, 'y' : y1, 'width' : width1, 'height' : height1},\n                      {'x' : x2, 'y' : y2, 'width' : width2, 'height' : height2},\n                    ]\n        dfbox= pd.DataFrame(box)\n    fig, a = plt.subplots(1,1)\n    fig.set_size_inches(5,5)\n    dicom_image = pydicom.read_file(rows) \n    try:\n        img = dicom_image.pixel_array\n                #print(path+imagename+'.jpg')\n              #scipy.misc.imsave (path+imagename+'.jpg', img)\n        a.imshow(img, cmap = 'gray')\n        for index, row in dfbox.iterrows(): \n            x, y, width, height  = row['x'], row['y'], row['width'], row['height']\n            rect = patches.Rectangle((x, y),\n                                     width, height,\n                                     linewidth = 2,\n                                     edgecolor = 'r',\n                                     facecolor = 'none')\n\n                # Draw the bounding box on top of the image\n            a.add_patch(rect)\n        plt.show()\n    except:\n        pass\n    ","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:42:40.094954Z","iopub.execute_input":"2021-06-30T12:42:40.095434Z","iopub.status.idle":"2021-06-30T12:43:17.120447Z","shell.execute_reply.started":"2021-06-30T12:42:40.095398Z","shell.execute_reply":"2021-06-30T12:43:17.119738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# define the amount of data that will be used training\nTRAIN_SPLIT = 0.5\n# the amount of validation data will be a percentage of the\n# *training* data\nVAL_SPLIT = 0.5\n# import the necessary packages\nimport random\nimport shutil\nimport os\n# grab the paths to all input images in the original input directory\n# and shuffle them\nrandom.seed(42)\nrandom.shuffle(train_path)\n# compute the training and testing split\ni = int(len(train_path) * TRAIN_SPLIT)\ntrainPaths = train_path[:i]\nvalPaths = train_path[i:]","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:17.123237Z","iopub.execute_input":"2021-06-30T12:43:17.123497Z","iopub.status.idle":"2021-06-30T12:43:17.128553Z","shell.execute_reply.started":"2021-06-30T12:43:17.123470Z","shell.execute_reply":"2021-06-30T12:43:17.127576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = ['Atypical Appearance', 'Indeterminate Appearance', 'Negative for Pneumonia', 'Typical Appearance', 'Not Identified']\nimg_size = 224\nfor colums in labels:\n    class_num = labels.index(colums) \ndata=[]\nfor rows in trainPaths:\n    imgname=rows.rsplit('/', 1)[1]\n    str_split =imgname.split('.')\n    if str_split[1]==\"dcm\" :\n        df=image_df[image_df['id'].str.contains(str_split[0])]\n        studydf=study_df[study_df.id.isin(df.StudyInstanceUID+'_study')]\n        lbl_folder=studydf.columns[studydf.eq(1).any()]\n        \n        try:\n            ds= pydicom.dcmread(rows)\n            img_arr = ds.pixel_array\n            resized_arr = resize(img_arr, (224, 224), anti_aliasing=True)[...,::-1]\n            if lbl_folder[0]==labels[0]:\n                data.append([resized_arr, 0])\n            elif lbl_folder[0]==labels[1]:\n                data.append([resized_arr, 1])\n            elif lbl_folder[0]==labels[2]:\n                data.append([resized_arr, 2])\n            elif lbl_folder[0]==labels[3]:\n                data.append([resized_arr, 3])\n            else:\n                data.append([resized_arr, 4])\n        except : \n            pass\n        \ndstrainimagedetails=np.array(data)  ","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:17.130074Z","iopub.execute_input":"2021-06-30T12:43:17.130426Z","iopub.status.idle":"2021-06-30T12:43:30.669929Z","shell.execute_reply.started":"2021-06-30T12:43:17.130390Z","shell.execute_reply":"2021-06-30T12:43:30.668459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nfrom tensorflow import keras \nfrom tensorflow.keras import Sequential\nfrom tensorflow.keras.layers import Dense, Conv2D , MaxPool2D , Flatten , Dropout\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.optimizers import Adam\nfrom sklearn.metrics import classification_report,confusion_matrix\nimport tensorflow as tf\n","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:30.675564Z","iopub.execute_input":"2021-06-30T12:43:30.675940Z","iopub.status.idle":"2021-06-30T12:43:35.409503Z","shell.execute_reply.started":"2021-06-30T12:43:30.675900Z","shell.execute_reply":"2021-06-30T12:43:35.408412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataval=[]\nfor rows in valPaths:\n    imgname=rows.rsplit('/', 1)[1]\n    str_split =imgname.split('.')\n    if str_split[1]==\"dcm\" :\n        df=image_df[image_df['id'].str.contains(str_split[0])]\n        studydf=study_df[study_df.id.isin(df.StudyInstanceUID+'_study')]\n        lbl_folder=studydf.columns[studydf.eq(1).any()]\n        try:\n            ds= pydicom.dcmread(rows)\n            img_arr = ds.pixel_array\n            resized_arr = resize(img_arr, (224, 224), anti_aliasing=True)[...,::-1]\n            if lbl_folder[0]==labels[0]:\n                dataval.append([resized_arr, 0])\n            elif lbl_folder[0]==labels[1]:\n                dataval.append([resized_arr, 1])\n            elif lbl_folder[0]==labels[2]:\n                dataval.append([resized_arr, 2])\n            elif lbl_folder[0]==labels[3]:\n                dataval.append([resized_arr, 3])\n            else:\n                dataval.append([resized_arr, 4])\n        except Exception as e: print(e)\n               \ndsvalidimagedetails=np.array(dataval)  ","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:35.414405Z","iopub.execute_input":"2021-06-30T12:43:35.414757Z","iopub.status.idle":"2021-06-30T12:43:52.450590Z","shell.execute_reply.started":"2021-06-30T12:43:35.414721Z","shell.execute_reply":"2021-06-30T12:43:52.449652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train = []\ny_train = []\nx_val = []\ny_val = []\n\nfor feature, label in dstrainimagedetails:\n    x_train.append(feature)\n    y_train.append(label)\nfor feature, label in dsvalidimagedetails:\n    x_val.append(feature)\n    y_val.append(label)\n\n# Normalize the data\nx_train = np.array(x_train) / 255\nx_val = np.array(x_val) / 255\n\nx_train =x_train.reshape(-1, 224, 224, 1)\ny_train=y_train = np.array(y_train)\n\nx_val =x_val.reshape(-1, 224, 224, 1)\ny_val = np.array(y_val)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:52.452203Z","iopub.execute_input":"2021-06-30T12:43:52.452565Z","iopub.status.idle":"2021-06-30T12:43:52.468762Z","shell.execute_reply.started":"2021-06-30T12:43:52.452525Z","shell.execute_reply":"2021-06-30T12:43:52.467670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(x_train.shape)\nprint(y_train.shape)\nprint(x_val.shape)\nprint(y_val.shape)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:52.470412Z","iopub.execute_input":"2021-06-30T12:43:52.470832Z","iopub.status.idle":"2021-06-30T12:43:52.480620Z","shell.execute_reply.started":"2021-06-30T12:43:52.470796Z","shell.execute_reply":"2021-06-30T12:43:52.476158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen = ImageDataGenerator(\n        featurewise_center=False,  # set input mean to 0 over the dataset\n        samplewise_center=False,  # set each sample mean to 0\n        featurewise_std_normalization=False,  # divide inputs by std of the dataset\n        samplewise_std_normalization=False,  # divide each input by its std\n        zca_whitening=False,  # apply ZCA whitening\n        rotation_range = 30,  # randomly rotate images in the range (degrees, 0 to 180)\n        zoom_range = 0.2, # Randomly zoom image \n        width_shift_range=0.1,  # randomly shift images horizontally (fraction of total width)\n        height_shift_range=0.1,  # randomly shift images vertically (fraction of total height)\n        horizontal_flip = True,  # randomly flip images\n        vertical_flip=False)  # randomly flip images\n\n\ndatagen.fit(x_train)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:52.482582Z","iopub.execute_input":"2021-06-30T12:43:52.483244Z","iopub.status.idle":"2021-06-30T12:43:52.492929Z","shell.execute_reply.started":"2021-06-30T12:43:52.483198Z","shell.execute_reply":"2021-06-30T12:43:52.491516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = Sequential()\nmodel.add(Conv2D(32,3,padding=\"same\", activation=\"relu\", input_shape=(224,224,1)))\nmodel.add(MaxPool2D())\n\nmodel.add(Conv2D(32, 3, padding=\"same\", activation=\"relu\"))\nmodel.add(MaxPool2D())\n\nmodel.add(Conv2D(64, 3, padding=\"same\", activation=\"relu\"))\nmodel.add(MaxPool2D())\nmodel.add(Dropout(0.4))\n\nmodel.add(Flatten())\nmodel.add(Dense(2048,activation=\"relu\"))\nmodel.add(Dense(4, activation=\"softmax\"))\n\nmodel.build(input_shape=(224,224,3))\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:52.494690Z","iopub.execute_input":"2021-06-30T12:43:52.495357Z","iopub.status.idle":"2021-06-30T12:43:54.566845Z","shell.execute_reply.started":"2021-06-30T12:43:52.495318Z","shell.execute_reply":"2021-06-30T12:43:54.565881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"opt = Adam(lr=0.000001)\nmodel.compile(optimizer = opt , loss = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True) , metrics = ['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:54.568982Z","iopub.execute_input":"2021-06-30T12:43:54.569393Z","iopub.status.idle":"2021-06-30T12:43:54.583942Z","shell.execute_reply.started":"2021-06-30T12:43:54.569351Z","shell.execute_reply":"2021-06-30T12:43:54.583111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(x_val.shape)\nprint(x_train.shape)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:54.586130Z","iopub.execute_input":"2021-06-30T12:43:54.586854Z","iopub.status.idle":"2021-06-30T12:43:54.592523Z","shell.execute_reply.started":"2021-06-30T12:43:54.586817Z","shell.execute_reply":"2021-06-30T12:43:54.591468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x_train,y_train,epochs = 50 , validation_data = (x_val,y_val))\n","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:43:54.594069Z","iopub.execute_input":"2021-06-30T12:43:54.594499Z","iopub.status.idle":"2021-06-30T12:44:02.231458Z","shell.execute_reply.started":"2021-06-30T12:43:54.594464Z","shell.execute_reply":"2021-06-30T12:44:02.230579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"acc = history.history['accuracy']\nval_acc = history.history['val_accuracy']\nloss = history.history['loss']\nval_loss = history.history['val_loss']\n\nepochs_range = range(50)\n\nplt.figure(figsize=(15, 15))\nplt.subplot(2, 2, 1)\nplt.plot(epochs_range, acc, label='Training Accuracy')\nplt.plot(epochs_range, val_acc, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\nplt.subplot(2, 2, 2)\nplt.plot(epochs_range, loss, label='Training Loss')\nplt.plot(epochs_range, val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:02.235197Z","iopub.execute_input":"2021-06-30T12:44:02.235470Z","iopub.status.idle":"2021-06-30T12:44:02.528941Z","shell.execute_reply.started":"2021-06-30T12:44:02.235444Z","shell.execute_reply":"2021-06-30T12:44:02.528115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, show_shapes=True)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:02.530394Z","iopub.execute_input":"2021-06-30T12:44:02.530952Z","iopub.status.idle":"2021-06-30T12:44:03.011999Z","shell.execute_reply.started":"2021-06-30T12:44:02.530910Z","shell.execute_reply":"2021-06-30T12:44:03.011001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict_classes(x_val)\npredictions = predictions.reshape(1,-1)[0]\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:03.013573Z","iopub.execute_input":"2021-06-30T12:44:03.013941Z","iopub.status.idle":"2021-06-30T12:44:03.156536Z","shell.execute_reply.started":"2021-06-30T12:44:03.013903Z","shell.execute_reply":"2021-06-30T12:44:03.153386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_val, predictions, target_names =['Atypical Appearance (Class 0)', 'Indeterminate Appearance (Class 1)', 'Negative for Pneumonia (Class 2)',  'Typical Appearance (Class 4)']))","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:03.157935Z","iopub.execute_input":"2021-06-30T12:44:03.158297Z","iopub.status.idle":"2021-06-30T12:44:03.175094Z","shell.execute_reply.started":"2021-06-30T12:44:03.158259Z","shell.execute_reply":"2021-06-30T12:44:03.173962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Predicting the classes for the test data whose label are not given","metadata":{}},{"cell_type":"code","source":"test_path=[]\nfor dirname, _, filenames in os.walk('/kaggle/input/siim-covid19-detection/test/'):\n    for filename in filenames:\n        str_split =filename.split('.')\n        if str_split[1]==\"dcm\" :\n            test_path.append(os.path.join(dirname, filename))\n","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:03.177006Z","iopub.execute_input":"2021-06-30T12:44:03.177427Z","iopub.status.idle":"2021-06-30T12:44:14.034752Z","shell.execute_reply.started":"2021-06-30T12:44:03.177388Z","shell.execute_reply":"2021-06-30T12:44:14.033980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random.shuffle(test_path)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:14.036045Z","iopub.execute_input":"2021-06-30T12:44:14.036401Z","iopub.status.idle":"2021-06-30T12:44:14.044718Z","shell.execute_reply.started":"2021-06-30T12:44:14.036364Z","shell.execute_reply":"2021-06-30T12:44:14.043959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testpath=[]\ni=0\nfor rows in test_path:\n    if i<51:\n        testpath.append(rows)\n        i=i+1","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:14.046073Z","iopub.execute_input":"2021-06-30T12:44:14.046471Z","iopub.status.idle":"2021-06-30T12:44:14.052433Z","shell.execute_reply.started":"2021-06-30T12:44:14.046432Z","shell.execute_reply":"2021-06-30T12:44:14.051360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"probability_model = tf.keras.Sequential([model, \n                                         tf.keras.layers.Softmax()])","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:14.054033Z","iopub.execute_input":"2021-06-30T12:44:14.054629Z","iopub.status.idle":"2021-06-30T12:44:14.091724Z","shell.execute_reply.started":"2021-06-30T12:44:14.054592Z","shell.execute_reply":"2021-06-30T12:44:14.091040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds_finalsubmission=[]\nfor rows in test_path:        \n    try:\n        img=rows.rsplit('/', 1)[1]\n        img=img.split('.')\n        img=img[0]+'_study'\n        dstest= pydicom.dcmread(rows)\n        img_arrtest = dstest.pixel_array\n        resized_arrtest = resize(img_arrtest, (224, 224), anti_aliasing=True)\n        x_test = np.array(resized_arrtest) / 255\n        x_test =x_test.reshape(-1, 224, 224, 1)\n        predictions = probability_model.predict(x_test)\n        predicted_label = np.argmax(predictions)\n        \n        if predicted_label==0:\n            title='atypical 1 0 0 1 1'\n        elif predicted_label==1:\n            title='intermediate 1 0 0 1 1'\n        elif predicted_label==2:\n            title='negative 1 0 0 1 1'\n        elif predicted_label==3:\n            title='typical 1 0 0 1 1'\n        elif predicted_label==4:\n            img=img[0]+'_image'\n            title='none 1 0 0 1 1'\n        output_data=img+','+title\n        ds_finalsubmission.append(output_data)\n    except: \n        pass\n               \n","metadata":{"execution":{"iopub.status.busy":"2021-06-30T12:44:14.093364Z","iopub.execute_input":"2021-06-30T12:44:14.093709Z","iopub.status.idle":"2021-06-30T13:02:31.054479Z","shell.execute_reply.started":"2021-06-30T12:44:14.093674Z","shell.execute_reply":"2021-06-30T13:02:31.053476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalsubmission=pd.DataFrame({\"Id,PredictionString\":ds_finalsubmission})","metadata":{"execution":{"iopub.status.busy":"2021-06-30T13:02:31.056302Z","iopub.execute_input":"2021-06-30T13:02:31.056679Z","iopub.status.idle":"2021-06-30T13:02:31.063551Z","shell.execute_reply.started":"2021-06-30T13:02:31.056640Z","shell.execute_reply":"2021-06-30T13:02:31.062399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalsubmission","metadata":{"execution":{"iopub.status.busy":"2021-06-30T13:02:31.068442Z","iopub.execute_input":"2021-06-30T13:02:31.068922Z","iopub.status.idle":"2021-06-30T13:02:31.088431Z","shell.execute_reply.started":"2021-06-30T13:02:31.068877Z","shell.execute_reply":"2021-06-30T13:02:31.087292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalsubmission.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2021-06-30T13:02:31.091395Z","iopub.execute_input":"2021-06-30T13:02:31.091978Z","iopub.status.idle":"2021-06-30T13:02:31.318511Z","shell.execute_reply.started":"2021-06-30T13:02:31.091935Z","shell.execute_reply":"2021-06-30T13:02:31.317752Z"},"trusted":true},"execution_count":null,"outputs":[]}]}