{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":26680,"databundleVersionId":2283525,"sourceType":"competition"}],"dockerImageVersionId":30096,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Problem description\n\n****\n\nFive times more deadly than the flu, COVID-19 causes significant morbidity and mortality. Like other pneumonias, pulmonary infection with COVID-19 results in inflammation and fluid in the lungs. COVID-19 looks very similar to other viral and bacterial pneumonias on chest radiographs, which makes it difficult to diagnose. This computer vision model for detection and localization of COVID-19 would help doctors provide a quick and confident diagnosis. As a result, patients could get the right treatment before the most severe effects of the virus take hold.\n\n\nCurrently, COVID-19 can be diagnosed via polymerase chain reaction to detect genetic material from the virus or chest radiograph. However, it can take a few hours and sometimes days before the molecular test results are back. By contrast, chest radiographs can be obtained in minutes. While guidelines exist to help radiologists differentiate COVID-19 from other types of infection, their assessments vary. In addition, non-radiologists could be supported with better localization of the disease, such as with a visual bounding box.\n\n\nIn this competition, the task is to identify and localize COVID-19 abnormalities on chest radiographs. In particular, categorization of the radiographs as negative for pneumonia or typical, indeterminate, or atypical for COVID-19.","metadata":{}},{"cell_type":"markdown","source":"**Categorization of the radiographs:**\n\n* NEGATIVE FOR PNEUMONIA - No lung opacities\n\n* TYPICAL APPEARANCE - Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution\n\n* INDETERMINATE APPEARANCE - Absence of typical findings AND unilateral, central or upper lung predominant distribution\n\n* ATYPICAL APPEARANCE - Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity","metadata":{}},{"cell_type":"markdown","source":"**Input data:**\n\n* train_study_level.csv - the train study-level metadata, with one row for each study, including correct labels.\n* train_image_level.csv - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\n* sample_submission.csv - a sample submission file containing all image- and study-level IDs.\n* train folder - comprises 6334 chest scans in DICOM format, stored in paths with the form study/series/image\n* test folder - The hidden test dataset is of roughly the same scale as the training dataset. Studies in the test set may contain more than one label.","metadata":{}},{"cell_type":"markdown","source":"# Content table\n\n****\n\n1. Importing the libraries\n2. Importing the datasets\n3. Data exploration\n4. Read Dicom files\n5. Feature engineering\n6. Making the model\n7. Compiling the model\n8. References","metadata":{}},{"cell_type":"markdown","source":"# Importing the libraries\n****","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sn\nimport pydicom as dicom # Dicom (Digital Imaging in Medicine) - medical image datasets, storage and transfer\nimport os\nfrom tqdm import tqdm # allows you to output a smart progress bar by wrapping around any iterable\nimport glob # retrieve files/pathnames matching a specified pattern\nimport pprint # pretty-print” arbitrary Python data structures\nimport ast # \nfrom pydicom.pixel_data_handlers.util import apply_voi_lut #\nimport wandb #\nimport keras\nfrom tensorflow.keras import models\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import optimizers\nfrom textwrap import wrap\n\npd.set_option('display.max_columns', 500)","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-11-03T18:51:15.362508Z","iopub.execute_input":"2024-11-03T18:51:15.362853Z","iopub.status.idle":"2024-11-03T18:51:15.372813Z","shell.execute_reply.started":"2024-11-03T18:51:15.362823Z","shell.execute_reply":"2024-11-03T18:51:15.371849Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Importing the datasets\n****","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/siim-covid19-detection/'\ntrain_dir = '/kaggle/input/siim-covid19-detection/train'\ntest_dir = '/kaggle/input/siim-covid19-detection/test'\n\ntrain_image_level = pd.read_csv(path + \"train_image_level.csv\")\ntrain_study_level = pd.read_csv(path + \"train_study_level.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:19.242832Z","iopub.execute_input":"2024-11-03T18:51:19.243188Z","iopub.status.idle":"2024-11-03T18:51:19.27796Z","shell.execute_reply.started":"2024-11-03T18:51:19.243156Z","shell.execute_reply":"2024-11-03T18:51:19.277177Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data exploration\n****","metadata":{}},{"cell_type":"markdown","source":"Let's have a look inside the train_image_level:","metadata":{}},{"cell_type":"code","source":"train_image_level.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:27.291576Z","iopub.execute_input":"2024-11-03T18:51:27.291944Z","iopub.status.idle":"2024-11-03T18:51:27.302838Z","shell.execute_reply.started":"2024-11-03T18:51:27.291907Z","shell.execute_reply":"2024-11-03T18:51:27.301991Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_image_level.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T18:51:32.02389Z","iopub.execute_input":"2024-11-03T18:51:32.024306Z","iopub.status.idle":"2024-11-03T18:51:32.06332Z","shell.execute_reply.started":"2024-11-03T18:51:32.024268Z","shell.execute_reply":"2024-11-03T18:51:32.062511Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are 6334 unique values in the train_image_level dataframe.","metadata":{}},{"cell_type":"markdown","source":"Now let's have a look inside train_study_level dataset:","metadata":{}},{"cell_type":"code","source":"train_study_level.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:36.416216Z","iopub.execute_input":"2024-11-03T18:51:36.416552Z","iopub.status.idle":"2024-11-03T18:51:36.427034Z","shell.execute_reply.started":"2024-11-03T18:51:36.416524Z","shell.execute_reply":"2024-11-03T18:51:36.426076Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_study_level.describe()","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:39.870412Z","iopub.execute_input":"2024-11-03T18:51:39.870753Z","iopub.status.idle":"2024-11-03T18:51:39.897294Z","shell.execute_reply.started":"2024-11-03T18:51:39.870723Z","shell.execute_reply":"2024-11-03T18:51:39.896368Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Distribution of Appearances**","metadata":{}},{"cell_type":"markdown","source":"Note:\n\nWhen you use enumerate(), the function gives you back two loop variables:\n\n1. The count of the current iteration\n2. The value of the item at the current iteration","metadata":{}},{"cell_type":"code","source":"columns = ['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']\nsum = []\n\n# label rotation for clear view\nfig, ax = plt.subplots()\nax.set_xticklabels(labels = columns, rotation = 45)\n\nfor column in columns:\n    plt.bar(column, train_study_level[column].sum())","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:47.1799Z","iopub.execute_input":"2024-11-03T18:51:47.180281Z","iopub.status.idle":"2024-11-03T18:51:47.319337Z","shell.execute_reply.started":"2024-11-03T18:51:47.180249Z","shell.execute_reply":"2024-11-03T18:51:47.31839Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are 6054 rows in the train_study_level dataframe. The number of unique values in study dataframe differs from the unique values in the images dataframe. Let's check how many studies have more than 1 image linked.","metadata":{}},{"cell_type":"code","source":"train_study_level_key = train_study_level.id.str[:-6]\ntraining_set = pd.merge(left = train_study_level, right = train_image_level, how = 'right', left_on = train_study_level_key, right_on = 'StudyInstanceUID')\ntraining_set.drop(['id_x'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:51:53.55639Z","iopub.execute_input":"2024-11-03T18:51:53.556727Z","iopub.status.idle":"2024-11-03T18:51:53.588588Z","shell.execute_reply.started":"2024-11-03T18:51:53.556698Z","shell.execute_reply":"2024-11-03T18:51:53.587909Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's have a look at these studies with multiple images:","metadata":{}},{"cell_type":"code","source":"training_set[training_set.groupby('StudyInstanceUID')['id_y'].transform('size') > 1].sort_values('StudyInstanceUID')","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:00.214429Z","iopub.execute_input":"2024-11-03T18:52:00.214799Z","iopub.status.idle":"2024-11-03T18:52:00.244075Z","shell.execute_reply.started":"2024-11-03T18:52:00.214762Z","shell.execute_reply":"2024-11-03T18:52:00.243271Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Read Dicom files\n****","metadata":{}},{"cell_type":"markdown","source":"Function used to locate image from the path:","metadata":{}},{"cell_type":"code","source":"def extract_image(i):\n    path_train = path + 'train/' + training_set.loc[i, 'StudyInstanceUID']\n    last_folder_in_path = os.listdir(path_train)[0]\n    path_train = path_train + '/{}/'.format(last_folder_in_path)\n    img_id = training_set.loc[i, 'id_y'].replace('_image','.dcm')\n    \n    print(img_id)\n    \n    data_file = dicom.dcmread(path_train + img_id)\n    img = data_file.pixel_array\n    return img","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:06.001907Z","iopub.execute_input":"2024-11-03T18:52:06.002301Z","iopub.status.idle":"2024-11-03T18:52:06.008038Z","shell.execute_reply.started":"2024-11-03T18:52:06.002262Z","shell.execute_reply":"2024-11-03T18:52:06.007076Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Images and rectangles visualization**","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(3,3, figsize=(20,16))\nfig.subplots_adjust(hspace=.1, wspace=.1)\naxes = axes.ravel()\n\nfor row in range(9):\n    img = extract_image(row)\n    if (training_set.loc[row,'boxes'] == training_set.loc[row,'boxes']):\n        boxes = ast.literal_eval(training_set.loc[row,'boxes'])\n        for box in boxes:\n            p = matplotlib.patches.Rectangle((box['x'], box['y']),\n                                              box['width'], box['height'],\n                                              ec = 'r', fc = 'none', lw = 2.\n                                            )\n            axes[row].add_patch(p)\n    axes[row].imshow(img, cmap = 'gray')\n    axes[row].set_title(training_set.loc[row, 'label'].split(' ')[0])\n    axes[row].set_xticklabels([])\n    axes[row].set_yticklabels([])","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:11.394939Z","iopub.execute_input":"2024-11-03T18:52:11.395331Z","iopub.status.idle":"2024-11-03T18:52:19.539925Z","shell.execute_reply.started":"2024-11-03T18:52:11.395296Z","shell.execute_reply":"2024-11-03T18:52:19.539113Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature engineering\n****","metadata":{}},{"cell_type":"markdown","source":"**Opacity_Count** - Count the number of opacities in the image","metadata":{}},{"cell_type":"code","source":"Opacity_Count = training_set['label'].str.count('opacity')\ntraining_set['Opacity_Count'] = Opacity_Count.values","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:25.677362Z","iopub.execute_input":"2024-11-03T18:52:25.677736Z","iopub.status.idle":"2024-11-03T18:52:25.690512Z","shell.execute_reply.started":"2024-11-03T18:52:25.677699Z","shell.execute_reply":"2024-11-03T18:52:25.689481Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Rectange_Area** - Sum of areas of rectangles - assumption : the bigger the rectangle - the bigger the opacity","metadata":{}},{"cell_type":"code","source":"image_rectangles_areas = []\n\nfor row in range(6334):#len(training_set.index)):\n    image_rectangles_area_sum = 0\n    rectangle_area = 0\n    if (training_set.loc[row,'boxes'] == training_set.loc[row,'boxes']):\n        boxes = ast.literal_eval(training_set.loc[row,'boxes'])\n        for box in boxes:\n            rectangle_area = box['width'] * box['height']\n            image_rectangles_area_sum = image_rectangles_area_sum + rectangle_area\n        image_rectangles_areas.append(image_rectangles_area_sum)\n    else: # nan values\n        image_rectangles_area_sum = image_rectangles_area_sum + rectangle_area\n        image_rectangles_areas.append(image_rectangles_area_sum)","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:30.035389Z","iopub.execute_input":"2024-11-03T18:52:30.03574Z","iopub.status.idle":"2024-11-03T18:52:30.417004Z","shell.execute_reply.started":"2024-11-03T18:52:30.035713Z","shell.execute_reply":"2024-11-03T18:52:30.416268Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_set['Rectangle_Area'] = image_rectangles_areas","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:35.417707Z","iopub.execute_input":"2024-11-03T18:52:35.418064Z","iopub.status.idle":"2024-11-03T18:52:35.424773Z","shell.execute_reply.started":"2024-11-03T18:52:35.418008Z","shell.execute_reply":"2024-11-03T18:52:35.423532Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Creating buckets - rectangle areas**","metadata":{}},{"cell_type":"markdown","source":"Distribution of the rectangle areas","metadata":{}},{"cell_type":"code","source":"training_set['Rectangle_Area'] = round(training_set['Rectangle_Area'],2)","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:39.664584Z","iopub.execute_input":"2024-11-03T18:52:39.664929Z","iopub.status.idle":"2024-11-03T18:52:39.669868Z","shell.execute_reply.started":"2024-11-03T18:52:39.664899Z","shell.execute_reply":"2024-11-03T18:52:39.668978Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#pd.qcut(training_set['Rectangle_Area'], q = 4)\n\n#training_set.boxplot(by = \"Negative for Pneumonia\",column = ['Rectangle_Area'],grid = True, layout=(1, 1))\n\ncut_labels_4 = ['0', '<1e6', '<2e6', '<4e6', '<8e6']\ncut_bins = [-1, 0, 1000000, 2000000, 4000000, 8000000]\ntraining_set['Rectangle_Area_Bin'] = pd.cut(training_set['Rectangle_Area'], bins = cut_bins, labels = cut_labels_4)","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:43.669713Z","iopub.execute_input":"2024-11-03T18:52:43.670067Z","iopub.status.idle":"2024-11-03T18:52:43.677598Z","shell.execute_reply.started":"2024-11-03T18:52:43.670029Z","shell.execute_reply":"2024-11-03T18:52:43.676686Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"columns = ['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']\n\nplt.figure(figsize = (16, 14))\nsn.set(font_scale = 1.2)\nsn.set_style('ticks')\n\nfor i, column in enumerate(columns):\n    plt.subplot(3, 3, i + 1)\n    sn.countplot(data = training_set, x = 'Rectangle_Area_Bin', hue = column, palette = ['#d02f52',\"#55a0ee\"])\n    \nsn.despine()","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:52:49.223591Z","iopub.execute_input":"2024-11-03T18:52:49.223924Z","iopub.status.idle":"2024-11-03T18:52:50.07916Z","shell.execute_reply.started":"2024-11-03T18:52:49.223896Z","shell.execute_reply":"2024-11-03T18:52:50.078304Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Ensure 'sum' is not redefined as a variable\ndel sum\n\nopacity = sorted(list(training_set['Rectangle_Area_Bin'].value_counts().index))\n\nfor i in opacity:\n    Count_Series = training_set[training_set['Rectangle_Area_Bin'] == i].iloc[:, [1, 2, 3, 4]].sum()\n    fig = plt.figure(figsize=(12, 3))\n    sn.barplot(x=Count_Series.index, y=Count_Series.values / sum(training_set['Rectangle_Area_Bin'] == i))\n    plt.title('Rectangle_Area_Bin : {}'.format(i))\n    plt.show()\n\n# For the second loop\nopacity = sorted(list(training_set['Opacity_Count'].value_counts().index))\n\nfor i in opacity:\n    Count_Series = training_set[training_set['Opacity_Count'] == i].iloc[:, [1, 2, 3, 4]].sum()\n    fig = plt.figure(figsize=(12, 3))\n    sn.barplot(x=Count_Series.index, y=Count_Series.values / sum(training_set['Opacity_Count'] == i))\n    plt.title('OpacityCount : {}'.format(i))\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:55:47.171384Z","iopub.execute_input":"2024-11-03T18:55:47.171736Z","iopub.status.idle":"2024-11-03T18:55:48.702969Z","shell.execute_reply.started":"2024-11-03T18:55:47.171706Z","shell.execute_reply":"2024-11-03T18:55:48.702111Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Rectangle area and opacity count","metadata":{}},{"cell_type":"markdown","source":"**TBD**: Position of the rectangle by quadrants (4 bins - 4 quadrants)","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Image metadata**","metadata":{}},{"cell_type":"code","source":"training_paths = []\ntrain_directory = \"../input/siim-covid19-detection/train/\"\n\nfor sid in tqdm(training_set['StudyInstanceUID']):\n    training_paths.append(glob.glob(os.path.join(train_directory, sid +\"/*/*\"))[0])\n\ntraining_set['path'] = training_paths","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:56:03.893198Z","iopub.execute_input":"2024-11-03T18:56:03.893547Z","iopub.status.idle":"2024-11-03T18:57:25.043346Z","shell.execute_reply.started":"2024-11-03T18:56:03.893519Z","shell.execute_reply":"2024-11-03T18:57:25.042489Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The pixel values are in the range of 0 to 255. It is easier for us to normalize the data between 0 to 1 and we can do that just by dividing our train and test set by 255.","metadata":{}},{"cell_type":"code","source":"voi_lut=True\nfix_monochrome=True\n\ndef dicom_dataset_to_dict(filename,func):\n    \"\"\"Credit: https://github.com/pydicom/pydicom/issues/319\n               https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\n    \"\"\"\n    \n    dicom_header = dicom.dcmread(filename) \n    \n    #====== DICOM FILE DATA ======\n    dicom_dict = {}\n    repr(dicom_header)\n    for dicom_value in dicom_header.values():\n        if dicom_value.tag == (0x7fe0, 0x0010):\n            #discard pixel data\n            continue\n        if type(dicom_value.value) == dicom.dataset.Dataset:\n            dicom_dict[dicom_value.name] = dicom_dataset_to_dict(dicom_value.value)\n        else:\n            v = _convert_value(dicom_value.value)\n            dicom_dict[dicom_value.name] = v\n      \n    del dicom_dict['Pixel Representation']\n    \n    if func != 'metadata_df':\n        #====== DICOM IMAGE DATA ======\n        # VOI LUT (if available by DICOM device) is used to transform raw DICOM data to \"human-friendly\" view\n        if voi_lut:\n            data = apply_voi_lut(dicom_header.pixel_array, dicom_header)\n        else:\n            data = dicom_header.pixel_array\n        # depending on this value, X-ray may look inverted - fix that:\n        if fix_monochrome and dicom_header.PhotometricInterpretation == \"MONOCHROME1\":\n            data = np.amax(data) - data\n        data = data - np.min(data)\n        data = data / np.max(data)\n        modified_image_data = (data * 255).astype(np.uint8)\n    \n        return dicom_dict, modified_image_data\n    \n    else:\n        return dicom_dict\n\ndef _sanitise_unicode(s):\n    return s.replace(u\"\\u0000\", \"\").strip()\n\ndef _convert_value(v):\n    t = type(v)\n    if t in (list, int, float):\n        cv = v\n    elif t == str:\n        cv = _sanitise_unicode(v)\n    elif t == bytes:\n        s = v.decode('ascii', 'replace')\n        cv = _sanitise_unicode(s)\n    elif t == dicom.valuerep.DSfloat:\n        cv = float(v)\n    elif t == dicom.valuerep.IS:\n        cv = int(v)\n    else:\n        cv = repr(v)\n    return cv\n\n","metadata":{"execution":{"iopub.status.busy":"2024-11-03T18:57:30.323896Z","iopub.execute_input":"2024-11-03T18:57:30.324284Z","iopub.status.idle":"2024-11-03T18:57:30.337481Z","shell.execute_reply.started":"2024-11-03T18:57:30.324251Z","shell.execute_reply":"2024-11-03T18:57:30.336513Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Getting the dictionary data to the dataframe and dropping the columns not needed","metadata":{}},{"cell_type":"code","source":"from multiprocessing import Pool, cpu_count\n\ndef process_file(filename):\n    try:\n        return dicom_dataset_to_dict(filename, 'metadata_df')\n    except:\n        return None\n\nwith Pool(cpu_count()) as pool:\n    metadata = pool.map(process_file, training_set.path)\n\n# Remove None entries from any files that failed to process\nmetadata = [m for m in metadata if m is not None]\ndicom_data_df = pd.DataFrame(metadata)\n","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-11-03T19:10:21.340773Z","iopub.execute_input":"2024-11-03T19:10:21.34115Z","iopub.status.idle":"2024-11-03T19:13:12.631794Z","shell.execute_reply.started":"2024-11-03T19:10:21.341113Z","shell.execute_reply":"2024-11-03T19:13:12.630701Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dicom_data_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:20:12.678148Z","iopub.execute_input":"2024-11-03T19:20:12.678524Z","iopub.status.idle":"2024-11-03T19:20:12.708858Z","shell.execute_reply.started":"2024-11-03T19:20:12.678488Z","shell.execute_reply":"2024-11-03T19:20:12.708089Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dicom_data_df.drop(['Specific Character Set', 'SOP Class UID','SOP Instance UID','Study Date','Study Time','Accession Number','Patient ID','Accession Number','Rows','Columns'], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:20:18.839543Z","iopub.execute_input":"2024-11-03T19:20:18.839896Z","iopub.status.idle":"2024-11-03T19:20:18.874589Z","shell.execute_reply.started":"2024-11-03T19:20:18.839866Z","shell.execute_reply":"2024-11-03T19:20:18.873624Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Get the metadata information as new columns in an existing dataframe","metadata":{}},{"cell_type":"code","source":"training_set_merged = pd.merge(left = training_set, right = dicom_data_df, how = 'left', left_on = 'StudyInstanceUID', right_on = 'Study Instance UID')\ntraining_set_merged.head()\ntraining_set = training_set_merged","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:20:24.184243Z","iopub.execute_input":"2024-11-03T19:20:24.184594Z","iopub.status.idle":"2024-11-03T19:20:24.215809Z","shell.execute_reply.started":"2024-11-03T19:20:24.184565Z","shell.execute_reply":"2024-11-03T19:20:24.215102Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize a new column 'Class' to store pneumonia class\ntraining_set['Class'] = ''\n\nfor row in range(len(training_set)):\n    if training_set.loc[row, 'Negative for Pneumonia'] == 1:\n        pneumonia_class = 'Negative for Pneumonia'\n    elif training_set.loc[row, 'Typical Appearance'] == 1:\n        pneumonia_class = 'Typical Appearance'\n    elif training_set.loc[row, 'Indeterminate Appearance'] == 1:\n        pneumonia_class = 'Indeterminate Appearance'\n    else:\n        pneumonia_class = 'Atypical Appearance'\n    \n    # Assign the class to the row\n    training_set.loc[row, 'Class'] = pneumonia_class\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:21:31.374848Z","iopub.execute_input":"2024-11-03T19:21:31.375233Z","iopub.status.idle":"2024-11-03T19:21:34.135785Z","shell.execute_reply.started":"2024-11-03T19:21:31.375196Z","shell.execute_reply":"2024-11-03T19:21:34.134846Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\n\nconditions = [\n    training_set['Negative for Pneumonia'] == 1,\n    training_set['Typical Appearance'] == 1,\n    training_set['Indeterminate Appearance'] == 1,\n]\n\nchoices = [\n    'Negative for Pneumonia',\n    'Typical Appearance',\n    'Indeterminate Appearance',\n]\n\ntraining_set['Class'] = np.select(conditions, choices, default='Atypical Appearance')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:21:49.239798Z","iopub.execute_input":"2024-11-03T19:21:49.240167Z","iopub.status.idle":"2024-11-03T19:21:49.249568Z","shell.execute_reply.started":"2024-11-03T19:21:49.24013Z","shell.execute_reply":"2024-11-03T19:21:49.248672Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"To be able to use y_col attribute in the ImageDataGenerator, we will create new column 'Class' containing the type of pneumonia.","metadata":{}},{"cell_type":"code","source":"for row in range(6334):\n    if training_set['Negative for Pneumonia'] == 1:\n        pneumonia_class = 'Negative for Pneumonia'\n    elif training_set['Typical Appearance'] == 1:\n        pneumonia_class = 'Typical Appearance'\n    elif training_set['Indeterminate Appearance'] == 1:\n        pneumonia_class = 'Indeterminate Appearance'\n    else:\n        pneumonia_class = 'Atypical Appearance'\n\n    training_set['Class'] = pneumonia_class","metadata":{"execution":{"iopub.status.busy":"2024-11-03T19:21:59.667642Z","iopub.execute_input":"2024-11-03T19:21:59.667979Z","iopub.status.idle":"2024-11-03T19:21:59.696761Z","shell.execute_reply.started":"2024-11-03T19:21:59.667949Z","shell.execute_reply":"2024-11-03T19:21:59.695491Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_set.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:22:05.348756Z","iopub.execute_input":"2024-11-03T19:22:05.349105Z","iopub.status.idle":"2024-11-03T19:22:05.390509Z","shell.execute_reply.started":"2024-11-03T19:22:05.349073Z","shell.execute_reply":"2024-11-03T19:22:05.389533Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**TBD**: Outliers and irregularities in the data","metadata":{}},{"cell_type":"markdown","source":"# Building the model\n****","metadata":{}},{"cell_type":"markdown","source":"EfficientNet is used:\n\nThe EfficientNet scaling method uniformly scales network width, depth, and resolution with a set of fixed scaling coefficients. The base EfficientNet-B0 network is based on the inverted bottleneck residual blocks of MobileNetV2, in addition to squeeze-and-excitation blocks.","metadata":{}},{"cell_type":"markdown","source":"**Understanding how EfficientNets works a little better:**","metadata":{}},{"cell_type":"markdown","source":"The core idea about Efficient Nets is the use of compound scaling - using a weighted scale of three inter-connected hyper parameters of the model - Resolution of the input, Depth of the Network and Width of the Network.\n\n\n![](https://warehouse-camo.ingress.cmh1.psfhosted.org/fe998467d67d4e76b3f0c81fd7d52db053735d7c/68747470733a2f2f6c617465782e636f6465636f67732e636f6d2f706e672e6c617465783f5c696e6c696e652673706163653b5c6470697b3330307d2673706163653b5c62675f77686974652673706163653b5c626567696e7b616c69676e2a7d2673706163653b64657074683a262673706163653b642673706163653b3d2673706163653b5c616c7068612673706163653b5e2673706163653b5c7068692673706163653b5c5c2673706163653b77696474683a262673706163653b772673706163653b3d2673706163653b5c626574612673706163653b5e2673706163653b5c7068692673706163653b5c5c2673706163653b7265736f6c7574696f6e3a262673706163653b722673706163653b3d2673706163653b5c67616d6d612673706163653b5e2673706163653b5c7068692673706163653b5c656e647b616c69676e2a7d//)\n\n\nWhen phi, the compound coefficient, is initially set to 1, we get the base configuration - in this case EfficientNetB0. We then use this configuration in a grid search to find the coefficients alpha, beta and gamma which optimize the following objective under the constraint:\n\n\n![](https://warehouse-camo.ingress.cmh1.psfhosted.org/bc03bbc347eef78c683053ad5e24f5e348c5562b/68747470733a2f2f6c617465782e636f6465636f67732e636f6d2f706e672e6c617465783f5c696e6c696e652673706163653b5c6470697b3330307d2673706163653b5c626567696e7b616c69676e2a7d2673706163653b5c616c7068612673706163653b5c63646f742673706163653b5c626574612673706163653b5e2673706163653b322673706163653b5c63646f742673706163653b5c67616d6d612673706163653b5e2673706163653b322673706163653b265c617070726f782673706163653b322673706163653b5c5c2673706163653b5c616c7068612673706163653b5c67652673706163653b312c2673706163653b5c626574612673706163653b5c67652673706163653b26312c2673706163653b5c67616d6d612673706163653b5c67652673706163653b312673706163653b5c656e647b616c69676e2a7d)\n\n\nOnce these coefficients for alpha, beta and gamma are found, then simply scale phi, the compound coeffieints by different amounts to get a family of models with more capacity and possibly better performance.","metadata":{}},{"cell_type":"markdown","source":"**EfficientNet pros:**\n\nBy using shortcuts directly between the bottlenecks which connects a much fewer number of channels compared to expansion layers, combined with depthwise separable convolution which effectively **reduces computation** by almost a factor of k^2, compared to traditional layers. Where k stands for the kernel size, specifying the height and width of the 2D convolution window.\n\nThe second benefit of EfficientNet, it scales more efficiently by carefully balancing network depth, width, and resolution, which lead to **better performance**.","metadata":{}},{"cell_type":"markdown","source":"****\nNext we set up the infrastructure to run a training job on our dataset. We choose the number of epochs to train for. The more epochs, the better your model is likely to fit your data but training will run for longer.\n\nNext, we set up the network to build the correct number of layers for the number of classes we have in our dataset.","metadata":{}},{"cell_type":"markdown","source":"**Layers:**\n\n\n**The pooling layer** operates upon each feature map separately to create a new set of the same number of pooled feature maps.\nPooling involves selecting a pooling operation, much like a filter to be applied to feature maps.\nAverage Pooling: Calculate the average value for each patch on the feature map.\n\n**Dense** implements the operation: output = activation(dot(input, kernel) + bias) where activation is the element-wise activation function passed as the activation argument, kernel is a weights matrix created by the layer, and bias is a bias vector created by the layer (only applicable if use_bias is True).\nDense is the only actual network layer in the model.\nA Dense layer feeds all outputs from the previous layer to all its neurons, each neuron providing one output to the next layer. A Dense(10) has ten neurons.\n\n**Model** groups layers into an object with training and inference features.\n\n**Adam** is a replacement optimization algorithm for stochastic gradient descent for training deep learning models. Adam combines the best properties of the AdaGrad and RMSProp algorithms to provide an optimization algorithm that can handle sparse gradients on noisy problems.\n\nLoss is a prediction error of Neural Net. And the method to calculate the loss is called Loss Function. In simple words, the Loss is used to calculate the gradients. And gradients are used to update the weights of the Neural Net.\n\n**CategoricalCrossentropy** - crossentropy loss function when there are two or more label classes","metadata":{}},{"cell_type":"code","source":"!pip install efficientnet\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:23:39.172852Z","iopub.execute_input":"2024-11-03T19:23:39.173254Z","iopub.status.idle":"2024-11-03T19:23:48.789423Z","shell.execute_reply.started":"2024-11-03T19:23:39.173213Z","shell.execute_reply":"2024-11-03T19:23:48.788415Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import efficientnet.tfkeras as efn\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:24:03.218179Z","iopub.execute_input":"2024-11-03T19:24:03.218693Z","iopub.status.idle":"2024-11-03T19:24:03.616775Z","shell.execute_reply.started":"2024-11-03T19:24:03.21864Z","shell.execute_reply":"2024-11-03T19:24:03.615926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define IMG_SIZES based on your image dimensions (example: 224x224 for EfficientNetB0)\nIMG_SIZES = [(224, 224), (240, 240), (260, 260), (300, 300), (380, 380), (456, 456), (528, 528), (600, 600)]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:25:26.379721Z","iopub.execute_input":"2024-11-03T19:25:26.380109Z","iopub.status.idle":"2024-11-03T19:25:26.385009Z","shell.execute_reply.started":"2024-11-03T19:25:26.380072Z","shell.execute_reply":"2024-11-03T19:25:26.384Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"****","metadata":{}},{"cell_type":"code","source":"import efficientnet.tfkeras as efn  # Import the EfficientNet models\n\n# Define image sizes for each EfficientNet model\nIMG_SIZES = [(224, 224), (240, 240), (260, 260), (300, 300), (380, 380), (456, 456), (528, 528), (600, 600)]\n\nEFNS = [efn.EfficientNetB0, efn.EfficientNetB1, efn.EfficientNetB2, efn.EfficientNetB3, \n        efn.EfficientNetB4, efn.EfficientNetB5, efn.EfficientNetB6, efn.EfficientNetB7]\n\ndef build_model(dim=IMG_SIZES[0], ef=0):\n    inputs = tf.keras.layers.Input(shape=(*dim, 3))\n    base = EFNS[ef](input_shape=(*dim,3), weights='imagenet', include_top=False)\n    x = base(inputs)\n    \n    # Pooling layer\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    # Dense layers\n    x = tf.keras.layers.Dense(64, activation='relu')(x)\n    x = tf.keras.layers.Dense(4, activation='softmax')(x)\n    \n    model = tf.keras.Model(inputs=inputs, outputs=x)\n    \n    opt = tf.keras.optimizers.Adam(learning_rate=0.001)\n    loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.01)\n    \n    auc = tf.keras.metrics.AUC(curve='ROC', multi_label=True)\n    acc = tf.keras.metrics.CategoricalAccuracy()\n    \n    # Ensure TensorFlow Addons (tfa) is imported for F1Score\n    import tensorflow_addons as tfa\n    f1  = tfa.metrics.F1Score(num_classes=4, average='macro', threshold=None)\n    \n    model.compile(optimizer=opt, loss=loss, metrics=[auc, acc, f1])\n    \n    return model\n","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:25:50.739311Z","iopub.execute_input":"2024-11-03T19:25:50.73971Z","iopub.status.idle":"2024-11-03T19:25:50.751235Z","shell.execute_reply.started":"2024-11-03T19:25:50.739671Z","shell.execute_reply":"2024-11-03T19:25:50.750324Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training the model\n****","metadata":{}},{"cell_type":"markdown","source":"1. Specify where the training and test folders are\n2. Use Keras's **ImageDataGenerator** to **augment** the training data. If you haven't used this library before, or are new to data augmentation, take a look at this link: http://keras.io/preprocessing/image/\n3. Use a pre-trained model called **EfficientNet**\n4. Make predictions on the test images in the test zip file and format the submission.csv file to hold our own submissions!","metadata":{}},{"cell_type":"markdown","source":"Image data augmentation is a technique that can be used to artificially expand the size of a training dataset by creating modified versions of images in the dataset.\n\n**Augmenting the images: ImageDataGenerator:**\n\n* generate **two generators** - one for training, and another for validation. These are stored in train_generator and val_generator. For both, we apply a series of distortions.\n* instead of storing all these new images in a directory, we use the method flow_from_dataframe to dynamically load these images as we train the model\n* **flow_from_dataframe** - takes the dataframe and the path to a directory + generates batches. The generated batches contain augmented/normalized data.\n* all the distortions we made for **train_gen** are not applied to test_gen. This is because we don't want to augment the data in the test directory.","metadata":{}},{"cell_type":"markdown","source":"* **training** dataset - used to fit the model\n* **validation** dataset - used to provide an unbiased evaluation of a model fit on the training dataset while tuning model hyperparameters","metadata":{}},{"cell_type":"code","source":"# Display the first few entries in the 'path' column\nprint(training_set['path'].head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:44:14.639443Z","iopub.execute_input":"2024-11-03T19:44:14.639809Z","iopub.status.idle":"2024-11-03T19:44:14.645571Z","shell.execute_reply.started":"2024-11-03T19:44:14.639779Z","shell.execute_reply":"2024-11-03T19:44:14.644685Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Update the path column in the DataFrame\nbase_directory = '/kaggle/working/'  # Set this to your actual image directory\n\n# Update the path column to reflect the correct directory\ntraining_set['path'] = training_set['path'].str.replace('/path/to/image/folder/', base_directory)\n\n# Display the updated paths to verify\nprint(training_set['path'].head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:44:46.574288Z","iopub.execute_input":"2024-11-03T19:44:46.574633Z","iopub.status.idle":"2024-11-03T19:44:46.588554Z","shell.execute_reply.started":"2024-11-03T19:44:46.574604Z","shell.execute_reply":"2024-11-03T19:44:46.587566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# Check if the images exist\nvalid_paths = training_set['path'].apply(lambda x: os.path.isfile(x))\n\n# Display any invalid paths\ninvalid_paths = training_set[~valid_paths]\n\nif invalid_paths.empty:\n    print(\"All image paths are valid.\")\nelse:\n    print(\"Invalid image paths found:\")\n    print(invalid_paths)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:45:29.533308Z","iopub.execute_input":"2024-11-03T19:45:29.533779Z","iopub.status.idle":"2024-11-03T19:45:29.599665Z","shell.execute_reply.started":"2024-11-03T19:45:29.533741Z","shell.execute_reply":"2024-11-03T19:45:29.598725Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display the invalid paths\nprint(\"Invalid image paths:\")\nprint(invalid_paths[['path']])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:46:20.822617Z","iopub.execute_input":"2024-11-03T19:46:20.822967Z","iopub.status.idle":"2024-11-03T19:46:20.831877Z","shell.execute_reply.started":"2024-11-03T19:46:20.822938Z","shell.execute_reply":"2024-11-03T19:46:20.830921Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!ls /kaggle/input\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:49:17.220762Z","iopub.execute_input":"2024-11-03T19:49:17.221137Z","iopub.status.idle":"2024-11-03T19:49:18.267794Z","shell.execute_reply.started":"2024-11-03T19:49:17.221101Z","shell.execute_reply":"2024-11-03T19:49:18.266595Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\n\n# Verify if the directory exists\nif os.path.exists(dataset_dir):\n    print(\"Dataset directory exists.\")\n    # List all files in the dataset directory\n    files = os.listdir(dataset_dir)\n    print(\"Files in dataset directory:\")\n    for file in files:\n        print(file)\nelse:\n    print(f\"Error: Dataset directory does not exist at {dataset_dir}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:50:20.24623Z","iopub.execute_input":"2024-11-03T19:50:20.246639Z","iopub.status.idle":"2024-11-03T19:50:20.254533Z","shell.execute_reply.started":"2024-11-03T19:50:20.2466Z","shell.execute_reply":"2024-11-03T19:50:20.253466Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Specify the path to the train directory\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    # List all files in the train directory\n    train_files = os.listdir(train_dir)\n    print(\"Files in train directory:\")\n    for file in train_files:\n        print(file)\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:51:24.418244Z","iopub.execute_input":"2024-11-03T19:51:24.418703Z","iopub.status.idle":"2024-11-03T19:51:25.30747Z","shell.execute_reply.started":"2024-11-03T19:51:24.418664Z","shell.execute_reply":"2024-11-03T19:51:25.301425Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Modify this list based on actual image file names or patterns\n# If you have the actual file names from the dataset, replace the below with those names\nexpected_images = ['22353f3ff2d1', 'e6ade0ee672e', 'feffa20fac13']  # Example names\n\nprint(\"\\nChecking for expected image files in train directory...\")\nfor img in expected_images:\n    img_path = os.path.join(train_dir, img)\n    if img in train_files:\n        print(f\"Found: {img}\")\n    else:\n        print(f\"Not found: {img}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:54:33.975745Z","iopub.execute_input":"2024-11-03T19:54:33.976142Z","iopub.status.idle":"2024-11-03T19:54:33.983173Z","shell.execute_reply.started":"2024-11-03T19:54:33.976099Z","shell.execute_reply":"2024-11-03T19:54:33.98226Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from PIL import Image\nimport os\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:55:49.245639Z","iopub.execute_input":"2024-11-03T19:55:49.24603Z","iopub.status.idle":"2024-11-03T19:55:49.249887Z","shell.execute_reply.started":"2024-11-03T19:55:49.245984Z","shell.execute_reply":"2024-11-03T19:55:49.248985Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom PIL import Image\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    train_dirs = os.listdir(train_dir)  # Get list of subdirectories\n    print(\"Directories in train directory (showing first 10):\")\n    \n    # Show only the first 10 directories\n    for dir_name in train_dirs[:10]:  \n        print(dir_name)\n    \n    print(f\"\\nTotal directories found: {len(train_dirs)}\")\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n\n# Attempt to load images from each directory in the train directory\nprint(\"\\nLoading images from the train directories...\")\nfor dir_name in train_dirs[:10]:  # Load images only from the first 10 directories\n    dir_path = os.path.join(train_dir, dir_name)\n    if os.path.isdir(dir_path):  # Check if it's a directory\n        # List image files in the directory\n        image_files = os.listdir(dir_path)\n        print(f\"\\nFound images in directory: {dir_name}\")\n        \n        # Limit to the first 5 images for loading\n        for image_file in image_files[:5]:  \n            img_path = os.path.join(dir_path, image_file)\n            try:\n                # Load the image\n                image = Image.open(img_path)\n                print(f\"Successfully loaded: {image_file}\")\n            except Exception as e:\n                print(f\"Error loading image {image_file}: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:57:44.437952Z","iopub.execute_input":"2024-11-03T19:57:44.438381Z","iopub.status.idle":"2024-11-03T19:57:44.466695Z","shell.execute_reply.started":"2024-11-03T19:57:44.438342Z","shell.execute_reply":"2024-11-03T19:57:44.46585Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom PIL import Image\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    train_dirs = os.listdir(train_dir)  # Get list of subdirectories\n    print(\"Directories in train directory (showing first 10):\")\n    \n    # Show only the first 10 directories\n    for dir_name in train_dirs[:10]:  \n        print(dir_name)\n    \n    print(f\"\\nTotal directories found: {len(train_dirs)}\")\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n\n# Attempt to load images from each directory in the train directory\nprint(\"\\nLoading images from the train directories...\")\nfor dir_name in train_dirs[:10]:  # Load images only from the first 10 directories\n    dir_path = os.path.join(train_dir, dir_name)\n    if os.path.isdir(dir_path):  # Check if it's a directory\n        # List all files in the directory\n        image_files = os.listdir(dir_path)\n        print(f\"\\nFound images in directory: {dir_name}\")\n        \n        # Load only files that are images\n        for image_file in image_files:\n            img_path = os.path.join(dir_path, image_file)\n            # Check if the file is an actual file (not a directory)\n            if os.path.isfile(img_path):\n                try:\n                    # Load the image\n                    image = Image.open(img_path)\n                    print(f\"Successfully loaded: {image_file}\")\n                except Exception as e:\n                    print(f\"Error loading image {image_file}: {e}\")\n            else:\n                print(f\"Skipped (not a file): {image_file}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:58:26.489748Z","iopub.execute_input":"2024-11-03T19:58:26.490125Z","iopub.status.idle":"2024-11-03T19:58:26.517978Z","shell.execute_reply.started":"2024-11-03T19:58:26.490089Z","shell.execute_reply":"2024-11-03T19:58:26.517113Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom PIL import Image\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    train_dirs = os.listdir(train_dir)  # Get list of subdirectories\n    print(\"Directories in train directory (showing first 10):\")\n    \n    # Show only the first 10 directories\n    for dir_name in train_dirs[:10]:  \n        print(dir_name)\n    \n    print(f\"\\nTotal directories found: {len(train_dirs)}\")\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n\n# Function to recursively load images from a directory\ndef load_images_from_directory(dir_path):\n    for root, dirs, files in os.walk(dir_path):\n        for file in files:\n            img_path = os.path.join(root, file)\n            if file.lower().endswith(('.png', '.jpg', '.jpeg', '.tiff', '.bmp', '.gif')):  # Add any other image extensions if needed\n                try:\n                    # Load the image\n                    image = Image.open(img_path)\n                    print(f\"Successfully loaded: {img_path}\")\n                except Exception as e:\n                    print(f\"Error loading image {img_path}: {e}\")\n\n# Attempt to load images from each directory in the train directory\nprint(\"\\nLoading images from the train directories...\")\nfor dir_name in train_dirs[:10]:  # Load images only from the first 10 directories\n    dir_path = os.path.join(train_dir, dir_name)\n    if os.path.isdir(dir_path):  # Check if it's a directory\n        print(f\"\\nSearching in directory: {dir_name}\")\n        load_images_from_directory(dir_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T19:59:14.132067Z","iopub.execute_input":"2024-11-03T19:59:14.132443Z","iopub.status.idle":"2024-11-03T19:59:14.163364Z","shell.execute_reply.started":"2024-11-03T19:59:14.132403Z","shell.execute_reply":"2024-11-03T19:59:14.162463Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom PIL import Image\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    train_dirs = os.listdir(train_dir)  # Get list of subdirectories\n    print(\"Directories in train directory (showing first 10):\")\n    \n    # Show only the first 10 directories\n    for dir_name in train_dirs[:10]:  \n        print(dir_name)\n    \n    print(f\"\\nTotal directories found: {len(train_dirs)}\")\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n\n# Function to recursively load images from a directory\ndef load_images_from_directory(dir_path):\n    total_files = 0\n    image_files = []\n    \n    for root, dirs, files in os.walk(dir_path):\n        print(f\"Checking directory: {root} (Total files: {len(files)})\")  # Print current directory and number of files\n        for file in files:\n            total_files += 1\n            img_path = os.path.join(root, file)\n            if file.lower().endswith(('.png', '.jpg', '.jpeg', '.tiff', '.bmp', '.gif')):  # Add any other image extensions if needed\n                image_files.append(img_path)  # Store valid image paths\n                try:\n                    # Load the image\n                    image = Image.open(img_path)\n                    print(f\"Successfully loaded: {img_path}\")\n                except Exception as e:\n                    print(f\"Error loading image {img_path}: {e}\")\n            else:\n                print(f\"Skipped (not an image file): {img_path}\")\n    \n    print(f\"Total files checked in '{dir_path}': {total_files}\")\n    print(f\"Total image files found: {len(image_files)}\")\n\n# Attempt to load images from each directory in the train directory\nprint(\"\\nLoading images from the train directories...\")\nfor dir_name in train_dirs[:10]:  # Load images only from the first 10 directories\n    dir_path = os.path.join(train_dir, dir_name)\n    if os.path.isdir(dir_path):  # Check if it's a directory\n        print(f\"\\nSearching in directory: {dir_name}\")\n        load_images_from_directory(dir_path)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:00:15.424114Z","iopub.execute_input":"2024-11-03T20:00:15.424503Z","iopub.status.idle":"2024-11-03T20:00:15.4656Z","shell.execute_reply.started":"2024-11-03T20:00:15.424467Z","shell.execute_reply":"2024-11-03T20:00:15.464664Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\n\n# Specify the dataset directory\ndataset_dir = '/kaggle/input/siim-covid19-detection'\ntrain_dir = os.path.join(dataset_dir, 'train')\n\n# Verify if the train directory exists\nif os.path.exists(train_dir):\n    print(\"Train directory exists.\")\n    train_dirs = os.listdir(train_dir)  # Get list of subdirectories\n    print(\"Directories in train directory (showing first 10):\")\n    \n    # Show only the first 10 directories\n    for dir_name in train_dirs[:10]:  \n        print(dir_name)\n    \n    print(f\"\\nTotal directories found: {len(train_dirs)}\")\nelse:\n    print(f\"Error: Train directory does not exist at {train_dir}\")\n\n# Function to recursively load DICOM images from a directory\ndef load_dicom_images_from_directory(dir_path):\n    total_files = 0\n    dicom_files = []\n    \n    for root, dirs, files in os.walk(dir_path):\n        print(f\"Checking directory: {root} (Total files: {len(files)})\")  # Print current directory and number of files\n        for file in files:\n            total_files += 1\n            img_path = os.path.join(root, file)\n            if file.lower().endswith('.dcm'):  # Check if it's a DICOM file\n                dicom_files.append(img_path)  # Store valid DICOM paths\n                try:\n                    # Load the DICOM image\n                    dicom = pydicom.dcmread(img_path)\n                    # Optionally display the image\n                    plt.imshow(dicom.pixel_array, cmap=plt.cm.bone)\n                    plt.axis('off')\n                    plt.show()\n                    print(f\"Successfully loaded DICOM: {img_path}\")\n                except Exception as e:\n                    print(f\"Error loading DICOM {img_path}: {e}\")\n            else:\n                print(f\"Skipped (not a DICOM file): {img_path}\")\n    \n    print(f\"Total files checked in '{dir_path}': {total_files}\")\n    print(f\"Total DICOM files found: {len(dicom_files)}\")\n\n# Attempt to load DICOM images from each directory in the train directory\nprint(\"\\nLoading DICOM images from the train directories...\")\nfor dir_name in train_dirs[:10]:  # Load images only from the first 10 directories\n    dir_path = os.path.join(train_dir, dir_name)\n    if os.path.isdir(dir_path):  # Check if it's a directory\n        print(f\"\\nSearching in directory: {dir_name}\")\n        load_dicom_images_from_directory(dir_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:01:15.485113Z","iopub.execute_input":"2024-11-03T20:01:15.485458Z","iopub.status.idle":"2024-11-03T20:01:26.978571Z","shell.execute_reply.started":"2024-11-03T20:01:15.485427Z","shell.execute_reply":"2024-11-03T20:01:26.977695Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(training_set.columns)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:06:40.712341Z","iopub.execute_input":"2024-11-03T20:06:40.71269Z","iopub.status.idle":"2024-11-03T20:06:40.718135Z","shell.execute_reply.started":"2024-11-03T20:06:40.71266Z","shell.execute_reply":"2024-11-03T20:06:40.71723Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(training_set.columns.tolist())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:08:06.84056Z","iopub.execute_input":"2024-11-03T20:08:06.840945Z","iopub.status.idle":"2024-11-03T20:08:06.845993Z","shell.execute_reply.started":"2024-11-03T20:08:06.840905Z","shell.execute_reply":"2024-11-03T20:08:06.845078Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_set.columns = training_set.columns.str.strip()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:08:25.797841Z","iopub.execute_input":"2024-11-03T20:08:25.798202Z","iopub.status.idle":"2024-11-03T20:08:25.80295Z","shell.execute_reply.started":"2024-11-03T20:08:25.798171Z","shell.execute_reply":"2024-11-03T20:08:25.801912Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Assuming you want to create a single class label from multiple columns\ntraining_set['Class'] = training_set[['Negative for Pneumonia', 'Typical Appearance', \n                                        'Indeterminate Appearance', 'Atypical Appearance']].idxmax(axis=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:08:56.939247Z","iopub.execute_input":"2024-11-03T20:08:56.939641Z","iopub.status.idle":"2024-11-03T20:08:56.945627Z","shell.execute_reply.started":"2024-11-03T20:08:56.939602Z","shell.execute_reply":"2024-11-03T20:08:56.94476Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_dg = data_generator.flow_from_dataframe(\n    dataframe=training_set,\n    directory=None,\n    x_col='path',\n    y_col='Class',  # This should now reference the correct class column\n    target_size=(224, 224),\n    subset='training',\n    batch_size=1024,\n    shuffle=True,\n    class_mode='categorical'\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:09:02.066425Z","iopub.execute_input":"2024-11-03T20:09:02.066753Z","iopub.status.idle":"2024-11-03T20:09:02.159411Z","shell.execute_reply.started":"2024-11-03T20:09:02.066726Z","shell.execute_reply":"2024-11-03T20:09:02.157585Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Strip any whitespace from the column names\ntraining_set.columns = training_set.columns.str.strip()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:09:50.858957Z","iopub.execute_input":"2024-11-03T20:09:50.859344Z","iopub.status.idle":"2024-11-03T20:09:50.863827Z","shell.execute_reply.started":"2024-11-03T20:09:50.8593Z","shell.execute_reply":"2024-11-03T20:09:50.862886Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(training_set.columns.tolist())  # Check the columns again\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:10:02.263696Z","iopub.execute_input":"2024-11-03T20:10:02.264056Z","iopub.status.idle":"2024-11-03T20:10:02.269193Z","shell.execute_reply.started":"2024-11-03T20:10:02.264Z","shell.execute_reply":"2024-11-03T20:10:02.267996Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a 'Class' column based on the other columns\n# Assuming you want to create a single label from multiple columns, we can use idxmax to select the highest class\ntraining_set['Class'] = training_set[['Negative for Pneumonia', 'Typical Appearance', \n                                        'Indeterminate Appearance', 'Atypical Appearance']].idxmax(axis=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:10:15.39945Z","iopub.execute_input":"2024-11-03T20:10:15.399902Z","iopub.status.idle":"2024-11-03T20:10:15.409089Z","shell.execute_reply.started":"2024-11-03T20:10:15.399862Z","shell.execute_reply":"2024-11-03T20:10:15.408117Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(training_set['Class'].unique())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:10:28.963335Z","iopub.execute_input":"2024-11-03T20:10:28.963686Z","iopub.status.idle":"2024-11-03T20:10:28.968515Z","shell.execute_reply.started":"2024-11-03T20:10:28.963654Z","shell.execute_reply":"2024-11-03T20:10:28.967592Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\nfrom keras.preprocessing.image import ImageDataGenerator\n\n# Specify the base directory containing the images\nbase_directory = '/path/to/image/folder'  # Change this to your actual image directory\n\n# Strip any whitespace from the column names\ntraining_set.columns = training_set.columns.str.strip()\n\n# Create full image paths\ntraining_set['path'] = training_set['path'].apply(lambda x: os.path.join(base_directory, x))\n\n# Check values in the relevant columns\nprint(training_set[['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']].head())\n\n# Create a 'Class' column based on existing categories\n# Check for binary indicators (0 or 1) and get the class label based on the maximum\ntraining_set['Class'] = training_set[['Negative for Pneumonia', \n                                        'Typical Appearance', \n                                        'Indeterminate Appearance', \n                                        'Atypical Appearance']].idxmax(axis=1)\n\n# Verify 'Class' column creation\nprint(\"Class column created. Unique classes:\", training_set['Class'].unique())\n\n# If the unique classes are still empty, check for NaN values\nif training_set['Class'].isna().any():\n    print(\"There are NaN values in the Class column. Consider checking the input data.\")\n\n# Create the ImageDataGenerator\ndata_generator = ImageDataGenerator(\n    rescale=1/255,\n    validation_split=0.10,\n    rotation_range=40,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n\n# Use the updated DataFrame to create the data generator\ntry:\n    train_dg = data_generator.flow_from_dataframe(\n        dataframe=training_set,\n        directory=None,                      # Set to None since we created full paths\n        x_col='path',                        # This now contains full paths\n        y_col='Class',                       # This should reference the 'Class' column\n        target_size=(224, 224),              # Set target size according to your model input size\n        subset='training',                   # Use 'training' or 'validation' as needed\n        batch_size=1024,\n        shuffle=True,\n        class_mode='categorical'\n    )\n    print(\"Image paths processed successfully.\")\nexcept KeyError as e:\n    print(f\"KeyError encountered: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:12:27.20617Z","iopub.execute_input":"2024-11-03T20:12:27.206585Z","iopub.status.idle":"2024-11-03T20:12:27.225359Z","shell.execute_reply.started":"2024-11-03T20:12:27.206545Z","shell.execute_reply":"2024-11-03T20:12:27.224275Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\nfrom keras.preprocessing.image import ImageDataGenerator\n\n# Load your dataset (update this with your actual loading mechanism)\n# training_set = pd.read_csv('your_dataset.csv')  # Example placeholder\n\n# Check if the DataFrame is loaded correctly\nprint(\"Initial DataFrame structure:\")\nprint(training_set.head())  # Check the first few rows of the DataFrame\n\n# Check columns and their names\nprint(\"DataFrame columns:\")\nprint(training_set.columns)\n\n# Ensure there are no leading/trailing spaces in the column names\ntraining_set.columns = training_set.columns.str.strip()\n\n# Check if the relevant columns exist\nrequired_columns = ['Negative for Pneumonia', \n                    'Typical Appearance', \n                    'Indeterminate Appearance', \n                    'Atypical Appearance']\nprint(\"Checking for required columns:\")\nfor col in required_columns:\n    print(f\"Column '{col}' exists: {col in training_set.columns}\")\n\n# Create full image paths\nbase_directory = '/path/to/image/folder'  # Update to your actual image directory\ntraining_set['path'] = training_set['path'].apply(lambda x: os.path.join(base_directory, x))\n\n# Print the unique values in the relevant columns before creating the Class column\nif not training_set.empty:\n    print(\"Unique values in relevant columns before class creation:\")\n    for col in required_columns:\n        print(f\"Unique values for '{col}': {training_set[col].unique()}\")\nelse:\n    print(\"Training set is empty. Cannot create classes.\")\n\n# Create the Class column\nif all(col in training_set.columns for col in required_columns):\n    training_set['Class'] = training_set[required_columns].idxmax(axis=1, skipna=True)\nelse:\n    print(\"Not all required columns are present. Cannot create 'Class'.\")\n\n# Verify the Class column creation\nif 'Class' in training_set.columns:\n    print(\"Class column created. Unique classes:\", training_set['Class'].unique())\nelse:\n    print(\"Class column was not created.\")\n\n# Create the ImageDataGenerator\ndata_generator = ImageDataGenerator(\n    rescale=1/255,\n    validation_split=0.10,\n    rotation_range=40,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n\n# Use the updated DataFrame to create the data generator\ntry:\n    train_dg = data_generator.flow_from_dataframe(\n        dataframe=training_set,\n        directory=None,                      # Set to None since we created full paths\n        x_col='path',                        # This now contains full paths\n        y_col='Class',                       # This should reference the 'Class' column\n        target_size=(224, 224),              # Set target size according to your model input size\n        subset='training',                   # Use 'training' or 'validation' as needed\n        batch_size=1024,\n        shuffle=True,\n        class_mode='categorical'\n    )\n    print(\"Image paths processed successfully.\")\nexcept KeyError as e:\n    print(f\"KeyError encountered: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:13:44.606241Z","iopub.execute_input":"2024-11-03T20:13:44.606603Z","iopub.status.idle":"2024-11-03T20:13:44.628368Z","shell.execute_reply.started":"2024-11-03T20:13:44.606572Z","shell.execute_reply":"2024-11-03T20:13:44.627425Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\nfrom keras.preprocessing.image import ImageDataGenerator\n\n# Load your dataset (ensure this path is correct)\ndata_file_path = 'path/to/your/data.csv'  # Update with your actual path\ntry:\n    training_set = pd.read_csv(data_file_path)\n    print(\"Data loaded successfully. Initial DataFrame structure:\")\n    print(training_set.head())\nexcept Exception as e:\n    print(f\"Error loading data: {e}\")\n\n# Check if the DataFrame is loaded correctly\nprint(\"Initial DataFrame structure:\")\nprint(training_set.head())  # Check the first few rows of the DataFrame\n\n# Check columns and their names\nprint(\"DataFrame columns:\")\nprint(training_set.columns)\n\n# Ensure there are no leading/trailing spaces in the column names\ntraining_set.columns = training_set.columns.str.strip()\n\n# Check if the relevant columns exist\nrequired_columns = ['Negative for Pneumonia', \n                    'Typical Appearance', \n                    'Indeterminate Appearance', \n                    'Atypical Appearance']\nprint(\"Checking for required columns:\")\nfor col in required_columns:\n    print(f\"Column '{col}' exists: {col in training_set.columns}\")\n\n# Check if the DataFrame is empty\nif training_set.empty:\n    print(\"Training set is empty. Please check your data source and loading process.\")\nelse:\n    # Create full image paths\n    base_directory = '/path/to/image/folder'  # Update to your actual image directory\n    training_set['path'] = training_set['path'].apply(lambda x: os.path.join(base_directory, x))\n\n    # Print unique values for each relevant column\n    print(\"Unique values in relevant columns before class creation:\")\n    for col in required_columns:\n        print(f\"Unique values for '{col}': {training_set[col].unique()}\")\n\n    # Create the Class column\n    training_set['Class'] = training_set[required_columns].idxmax(axis=1, skipna=True)\n    print(\"Class column created. Unique classes:\", training_set['Class'].unique())\n\n# Create the ImageDataGenerator\ndata_generator = ImageDataGenerator(\n    rescale=1/255,\n    validation_split=0.10,\n    rotation_range=40,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n\n# Use the updated DataFrame to create the data generator\ntry:\n    train_dg = data_generator.flow_from_dataframe(\n        dataframe=training_set,\n        directory=None,                      # Set to None since we created full paths\n        x_col='path',                        # This now contains full paths\n        y_col='Class',                       # This should reference the 'Class' column\n        target_size=(224, 224),              # Set target size according to your model input size\n        subset='training',                   # Use 'training' or 'validation' as needed\n        batch_size=1024,\n        shuffle=True,\n        class_mode='categorical'\n    )\n    print(\"Image paths processed successfully.\")\nexcept KeyError as e:\n    print(f\"KeyError encountered: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:14:30.128389Z","iopub.execute_input":"2024-11-03T20:14:30.128763Z","iopub.status.idle":"2024-11-03T20:14:30.14961Z","shell.execute_reply.started":"2024-11-03T20:14:30.128731Z","shell.execute_reply":"2024-11-03T20:14:30.148472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\n\n# Print the current working directory\nprint(\"Current working directory:\", os.getcwd())\n\n# Specify the path to your actual data file\ndata_file_path = 'your/actual/path/to/data.csv'  # Update to your actual path\n\n# Check if the file exists\nif os.path.exists(data_file_path):\n    print(\"File exists. Proceeding to load data...\")\n    \n    # Load the DataFrame\n    training_set = pd.read_csv(data_file_path)\n    \n    # Print the initial DataFrame structure\n    print(\"Data loaded successfully. Initial DataFrame structure:\")\n    print(training_set.head())  # Check the first few rows\n\n    # Check if the DataFrame is empty\n    if training_set.empty:\n        print(\"Training set is empty. Please check your data source and loading process.\")\n    else:\n        # Proceed with the rest of your code\n        print(\"Initial DataFrame structure:\")\n        print(training_set.info())  # Display DataFrame info\n\n        # Define the required columns for class creation\n        required_columns = ['Negative for Pneumonia', 'Typical Appearance', \n                            'Indeterminate Appearance', 'Atypical Appearance']\n\n        # Check for required columns\n        for col in required_columns:\n            print(f\"Column '{col}' exists: {col in training_set.columns}\")\n\n        # Print the unique values in the relevant columns before creating the Class column\n        print(\"Unique values in relevant columns before class creation:\")\n        print(training_set[required_columns].unique())\n\n        # Create the Class column based on the existing columns\n        def create_class(row):\n            if row['Negative for Pneumonia'] == 1:\n                return 'Negative'\n            elif row['Typical Appearance'] == 1:\n                return 'Typical'\n            elif row['Indeterminate Appearance'] == 1:\n                return 'Indeterminate'\n            elif row['Atypical Appearance'] == 1:\n                return 'Atypical'\n            return 'Unknown'\n\n        training_set['Class'] = training_set.apply(create_class, axis=1)\n\n        # Print the unique classes\n        unique_classes = training_set['Class'].unique()\n        print(f\"Class column created. Unique classes: {unique_classes}\")\n\nelse:\n    print(f\"File does not exist at: {data_file_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:16:06.472255Z","iopub.execute_input":"2024-11-03T20:16:06.472642Z","iopub.status.idle":"2024-11-03T20:16:06.483428Z","shell.execute_reply.started":"2024-11-03T20:16:06.472604Z","shell.execute_reply":"2024-11-03T20:16:06.482466Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\n\n# Print the current working directory\nprint(\"Current working directory:\", os.getcwd())\n\n# List files in the current directory\nprint(\"Files in the current working directory:\")\nprint(os.listdir('/kaggle/working'))\n\n# Specify the path to your actual data file (update the filename as needed)\ndata_file_path = '/kaggle/working/data.csv'  # Update to your actual path after confirming the filename\n\n# Check if the file exists\nif os.path.exists(data_file_path):\n    print(\"File exists. Proceeding to load data...\")\n    \n    # Load the DataFrame\n    training_set = pd.read_csv(data_file_path)\n    \n    # Print the initial DataFrame structure\n    print(\"Data loaded successfully. Initial DataFrame structure:\")\n    print(training_set.head())  # Check the first few rows\n\n    # Check if the DataFrame is empty\n    if training_set.empty:\n        print(\"Training set is empty. Please check your data source and loading process.\")\n    else:\n        # Proceed with the rest of your code\n        print(\"Initial DataFrame structure:\")\n        print(training_set.info())  # Display DataFrame info\n\n        # Define the required columns for class creation\n        required_columns = ['Negative for Pneumonia', 'Typical Appearance', \n                            'Indeterminate Appearance', 'Atypical Appearance']\n\n        # Check for required columns\n        for col in required_columns:\n            print(f\"Column '{col}' exists: {col in training_set.columns}\")\n\n        # Print the unique values in the relevant columns before creating the Class column\n        print(\"Unique values in relevant columns before class creation:\")\n        print(training_set[required_columns].unique())\n\n        # Create the Class column based on the existing columns\n        def create_class(row):\n            if row['Negative for Pneumonia'] == 1:\n                return 'Negative'\n            elif row['Typical Appearance'] == 1:\n                return 'Typical'\n            elif row['Indeterminate Appearance'] == 1:\n                return 'Indeterminate'\n            elif row['Atypical Appearance'] == 1:\n                return 'Atypical'\n            return 'Unknown'\n\n        training_set['Class'] = training_set.apply(create_class, axis=1)\n\n        # Print the unique classes\n        unique_classes = training_set['Class'].unique()\n        print(f\"Class column created. Unique classes: {unique_classes}\")\n\nelse:\n    print(f\"File does not exist at: {data_file_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:17:00.679918Z","iopub.execute_input":"2024-11-03T20:17:00.680285Z","iopub.status.idle":"2024-11-03T20:17:00.692554Z","shell.execute_reply.started":"2024-11-03T20:17:00.680254Z","shell.execute_reply":"2024-11-03T20:17:00.691682Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List files in the /kaggle/input directory\nprint(\"Files in the /kaggle/input directory:\")\nprint(os.listdir('/kaggle/input'))\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:17:37.904845Z","iopub.execute_input":"2024-11-03T20:17:37.905214Z","iopub.status.idle":"2024-11-03T20:17:37.910278Z","shell.execute_reply.started":"2024-11-03T20:17:37.905183Z","shell.execute_reply":"2024-11-03T20:17:37.909445Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# List files in the siim-covid19-detection directory\ninput_directory = '/kaggle/input/siim-covid19-detection'\nprint(\"Files in the siim-covid19-detection directory:\")\nprint(os.listdir(input_directory))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:18:05.886472Z","iopub.execute_input":"2024-11-03T20:18:05.886835Z","iopub.status.idle":"2024-11-03T20:18:05.898663Z","shell.execute_reply.started":"2024-11-03T20:18:05.8868Z","shell.execute_reply":"2024-11-03T20:18:05.897752Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\n\n# Define the path to the train_image_level.csv file\ndata_file_path = '/kaggle/input/siim-covid19-detection/train_image_level.csv'\n\n# Check if the file exists\nif os.path.exists(data_file_path):\n    print(\"File exists. Proceeding to load data...\")\n\n    # Load the DataFrame\n    training_set = pd.read_csv(data_file_path)\n\n    # Print the initial DataFrame structure\n    print(\"Data loaded successfully. Initial DataFrame structure:\")\n    print(training_set.head())  # Check the first few rows\n\n    # Check if the DataFrame is empty\n    if training_set.empty:\n        print(\"Training set is empty. Please check your data source and loading process.\")\n    else:\n        # Proceed with the rest of your code\n        print(\"Initial DataFrame structure:\")\n        print(training_set.info())  # Display DataFrame info\n\n        # Check the unique values in the 'label' column to understand its content\n        unique_labels = training_set['label'].unique()\n        print(\"Unique values in the 'label' column:\")\n        print(unique_labels)\n\n        # Create the Class column based on the 'label'\n        def create_class(label):\n            if 'opacity' in label:\n                return 'Positive'\n            else:\n                return 'Negative'\n\n        training_set['Class'] = training_set['label'].apply(create_class)\n\n        # Print the unique classes\n        unique_classes = training_set['Class'].unique()\n        print(f\"Class column created. Unique classes: {unique_classes}\")\n\nelse:\n    print(f\"File does not exist at: {data_file_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:19:34.949038Z","iopub.execute_input":"2024-11-03T20:19:34.949463Z","iopub.status.idle":"2024-11-03T20:19:34.997687Z","shell.execute_reply.started":"2024-11-03T20:19:34.949428Z","shell.execute_reply":"2024-11-03T20:19:34.996757Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(training_set.head())  # Display the first few rows to inspect the DataFrame\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:20:36.721784Z","iopub.execute_input":"2024-11-03T20:20:36.722149Z","iopub.status.idle":"2024-11-03T20:20:36.730293Z","shell.execute_reply.started":"2024-11-03T20:20:36.722115Z","shell.execute_reply":"2024-11-03T20:20:36.729394Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.preprocessing import LabelEncoder\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Flatten\nfrom keras.optimizers import Adam\nfrom keras.utils import to_categorical\n\n# Load the data\ndata_path = '/kaggle/input/siim-covid19-detection/train_image_level.csv'\ntraining_set = pd.read_csv(data_path)\n\n# Assuming 'label' has the classes, and 'Class' contains your binary labels\nX = training_set['label']  # Change this if needed to your actual feature\ny = training_set['Class']\n\n# Encode labels\nle = LabelEncoder()\ny_encoded = le.fit_transform(y)\n\n# Convert to categorical (if needed)\ny_categorical = to_categorical(y_encoded)\n\n# Set up Stratified K-Folds\nfolds = 5\nskf = StratifiedKFold(n_splits=folds, shuffle=True, random_state=42)\n\n# Initialize model\nmodel = Sequential()\nmodel.add(Flatten(input_shape=(input_shape)))  # Specify your input shape\nmodel.add(Dense(128, activation='relu'))\nmodel.add(Dense(len(le.classes_), activation='softmax'))  # Change output based on classes\n\n# Compile model\nmodel.compile(optimizer=Adam(), loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Train and validate the model using k-folds\nepochs = [12] * folds\nresults = []\n\nfor fold, (train_index, val_index) in enumerate(skf.split(X, y_encoded)):\n    print(f\"Fold {fold + 1}/{folds}\")\n\n    # Split data into training and validation sets\n    X_train, X_val = X.iloc[train_index], X.iloc[val_index]\n    y_train, y_val = y_categorical[train_index], y_categorical[val_index]\n\n    # Train the model\n    history = model.fit(X_train, y_train, epochs=epochs[fold], validation_data=(X_val, y_val))\n    \n    # Store results\n    results.append(history.history)\n\n# You can analyze `results` for performance metrics after training\n","metadata":{"execution":{"iopub.status.busy":"2024-11-03T20:23:38.078692Z","iopub.execute_input":"2024-11-03T20:23:38.079072Z","iopub.status.idle":"2024-11-03T20:23:38.198385Z","shell.execute_reply.started":"2024-11-03T20:23:38.079032Z","shell.execute_reply":"2024-11-03T20:23:38.196947Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**KFold cross validation**\n****\nProvides train/test indices to split data in train/test sets. Split dataset into k consecutive folds (without shuffling by default).\n\nEach fold is then used once as a validation while the k - 1 remaining folds form the training set.\n\nWhen shuffle is True, random_state affects the ordering of the indices, which controls the randomness of each fold. Otherwise, this parameter has no effect. \n\nsplit(X[, y, groups]) - Generate indices to split data into training and test set.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import KFold\n\nskf = KFold(n_splits = folds, shuffle = True, random_state = 0)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:23:47.751945Z","iopub.execute_input":"2024-11-03T20:23:47.752347Z","iopub.status.idle":"2024-11-03T20:23:47.756474Z","shell.execute_reply.started":"2024-11-03T20:23:47.752311Z","shell.execute_reply":"2024-11-03T20:23:47.755481Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"np.arange(num_of_train_files)\nskf.split(np.arange(num_of_train_files))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-03T20:23:54.300443Z","iopub.execute_input":"2024-11-03T20:23:54.300812Z","iopub.status.idle":"2024-11-03T20:23:54.306336Z","shell.execute_reply.started":"2024-11-03T20:23:54.300775Z","shell.execute_reply":"2024-11-03T20:23:54.305421Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# References\n****\n\n* https://github.com/pydicom/pydicom/issues/319\n* https://www.kaggle.com/songseungwon/siim-covid-19-detection-10-step-tutorial-1\n* https://www.kaggle.com/ruchi798/siim-covid-19-detection-eda-data-augmentation#DICOM-data\n* https://www.kaggle.com/awsaf49/siim-covid-19-study-level-train-tpu/comments\n* https://towardsdatascience.com/train-validation-and-test-sets-72cb40cba9e7\n* https://www.kaggle.com/arjunrao2000/beginners-guide-efficientnet-with-keras","metadata":{}}]}