{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nimport seaborn as sns \nimport glob \nimport matplotlib.pyplot as plt \n'''for dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))'''\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the .csv files ","metadata":{}},{"cell_type":"markdown","source":"- **train_study_level.csv** - the train study-level metadata, with one row for each study, including correct labels.\n\n- **train_image_level.csv** - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\n\n- **sample_submission.csv** - a sample submission file containing all image- and study-level IDs.","metadata":{}},{"cell_type":"markdown","source":"# train_image_level.csv\n\n- **id** - unique image identifier\n- **boxes** - bounding boxes in easily-readable dictionary format\n- **label** - the correct prediction label for the provided bounding boxes","metadata":{}},{"cell_type":"code","source":"train_label = pd.read_csv('/kaggle/input/siim-covid19-detection/train_image_level.csv')\ntrain_label.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study = pd.read_csv('/kaggle/input/siim-covid19-detection/train_study_level.csv')\ntrain_study.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(train_study), len(train_label))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = train_study.columns\ncol","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## train_study_level.csv file labels \n\n\n- **id** - unique study identifier\n- **Negative for Pneumonia** - 1 if the study is negative for pneumonia, 0 otherwise\n- **Typical Appearance** - 1 if the study has this appearance, 0 otherwise\n- **Indeterminate Appearance**  - 1 if the study has this appearance, 0 otherwise\n- **Atypical Appearance**  - 1 if the study has this appearance, 0 otherwise","metadata":{}},{"cell_type":"code","source":"plt.figure(1)\nfor i in range(len(col)-1):\n    #     plt.subplot(1,4,i)\n    plt.figure(i)\n    sns.countplot(train_study[col[i+1]])\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(train_study['Typical Appearance'])\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the .dicom images ","metadata":{}},{"cell_type":"code","source":"import pydicom as dcm ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file = glob.glob('/kaggle/input/siim-covid19-detection/train/*/*/*')\nlen(file)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(dcm.dcmread(file[0]).pixel_array, cmap = 'gray')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}