{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Pulmonary fibrosis\n\n**Pulmonary fibrosis** is a lung disease that occurs when lung tissue becomes damaged and scarred. This thickened, stiff tissue makes it more difficult for your lungs to work properly. As pulmonary fibrosis worsens, you become progressively more short of breath.\n\n![Pulmonary fibrosis](https://upload.wikimedia.org/wikipedia/commons/thumb/e/e1/IPF_amiodarone.JPG/300px-IPF_amiodarone.JPG)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Training Data ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* In this notebook we are going to see some Explorative Data Analysis\n\n    1. Finding the correlations of the columns\n    2. Finding the Null Values in the data\n    3. Finding the Different Ages\n    4. Count of the sex \n          * Male or Female\n    5. Class of the data \n    6. Plotting some images ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here we are Just going take the train data for our EDA purpose ","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"data = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Setting the Patients column as the Index\n\ndata = data.set_index(['Patient'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finding the Information of the from here we got \n* total rows is 1549\n* total columns 6\n* All the datatypes\n* size of the dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1>Correlation</h1>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here in this we are having 4 numeric columns and 2 object data type columns   \n* So, here we are working for the correlation, correlation is for only the Numeric data to find the remaining columns we need encode from string to numbers","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data.corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"LabelEncoding is the technique in machine learning which help to encode the string data numeric, now we are going to use from sklearn \n\n* We are encoding Sex column into the numeric\n* We are encoding Smoking Status column into the numeric","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"le = LabelEncoder()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data['Sex'] = le.fit_transform(data['Sex'])\ndata['SmokingStatus'] = le.fit_transform(data['SmokingStatus'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nfig = plt.figure(figsize=(15, 10))\nsns.heatmap(data.corr(), annot=True, \n           cbar_kws={\"orientation\": \"horizontal\"})\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1>Null Values in the Data</h1>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"val = data.isna().sum().values\nlab = data.columns\n\nplt.bar(lab, val)\nplt.xlabel('Labels')\nplt.ylabel(\"Empty or not\")\nplt.title(\"Null values\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1>Ages Counts<h1>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = data['Age'].value_counts().values\nlabs = data['Age'].unique()\n\nfigu = plt.figure(figsize=(15, 5))\nfig = sns.barplot(labs,vals) \nplt.xlabel('Ages')\nplt.ylabel(\"Total Pateints\")\nplt.title(\"Total number of patients with the particular age\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nvals = data['Sex'].value_counts().values\nlabs = ['male', 'Female']\nsns.barplot(vals, labs)\nplt.xlabel('Labels')\nplt.ylabel(\"Number of males or females\")\nplt.title(\"Sex Count\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1> Male and Female SexCount </h1>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = data['Sex'].value_counts().values\nlabs = ['male', 'female']\n\nexplode = []\n\nfor i in range(len(vals)):\n    if max(vals) == vals[i]:\n        explode.append(0.1)\n    else:\n        explode.append(0)\n        \n\nplt.pie(vals, explode, labs, autopct='%1.1f%%',\n        shadow=True, startangle=90)\nplt.title(\"Sex Count\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = data['SmokingStatus'].unique()\nvals","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1>Smoking Status</h1>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = data['SmokingStatus'].value_counts().values\nlabs = ['Ex-smoker', 'Never smoked', 'Currently smokes']\n\nfig = plt.figure(figsize=(7, 7))\n\nexplode = []\n\nfor i in range(len(vals)):\n    if max(vals) == vals[i]:\n        explode.append(0.1)\n    else:\n        explode.append(0)\n        \n\nplt.pie(vals, explode, labs, autopct='%1.1f%%',\n        shadow=True, startangle=90)\nplt.title(\"Sex Count\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = data['SmokingStatus'].value_counts().values\nlabs = ['Ex-smoker', 'Never smoked', 'Currently smokes']\n\n\nfig = plt.figure(figsize=(12, 5))\nsns.barplot(vals, labs)\n\nplt.xlabel('Labels')\nplt.ylabel(\"Number of males or females\")\nplt.title(\"Sex Count\")\nplt.tight_layout()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1> Plotting the Images </h1>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here we are plotting the images with help of matplotlib and pydicom, pydicom is a library to plot medical images which are ending with **.dcm**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430'\n\nimages = [path + '/' + img for img in os.listdir(path) if img.endswith('dcm')]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\nw=10\nh=10\nfig=plt.figure(figsize=(14, 8))\ncolumns = 3\nrows = 3\nfor i in range(1, columns*rows+1):\n    ds = pydicom.dcmread(images[i])\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='hsv') \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# FVC\n\n\n**Forced vital capacity (FVC)** is the total amount of air exhaled during the FEV test. Forced expiratory volume and forced vital capacity are lung function tests that are measured during spirometry. ... Diagnose obstructive lung diseases such as asthma and chronic obstructive pulmonary disease (COPD).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fvc_female = []\nfvc_male = []\n\nfor i in range(len(data)):\n    \n    if data['Sex'][i] == 1 :\n        fvc_female.append(data['FVC'][i])\n    elif data['Sex'][i] == 0:\n        fvc_male.append(data['FVC'][i])\n    else:\n        pass","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(fvc_female, list(range(len(fvc_female))), c='r', label='Male')\nplt.scatter(fvc_male,list(range(len(fvc_male))), c='y', label='Female')\nplt.xlabel(\"FVC\")\nplt.ylabel(\"Range\")\nplt.title(\"Forced vital capacity (FVC)\")\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"The Max week of the patient {max(data['Weeks'])}\")\nprint(f\"The Max week of the patient {min(data['Weeks'])}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Testing data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"data = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Setting the Patients column as the Index\n\ndata = data.set_index(['Patient'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Images of the patients in testset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '../input/osic-pulmonary-fibrosis-progression/test/ID00419637202311204720264'\n\nimages = [path + '/' + img for img in os.listdir(path) if img.endswith('dcm')]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = pydicom.dcmread(images[0])\nval = ds.pixel_array\nimg = np.array(val, dtype='f')\nimg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\nimport cv2\nw=10\nh=10\nfig=plt.figure(figsize=(14, 8))\ncolumns = 3\nrows = 3\nfor i in range(1, columns*rows+1):\n    ds = pydicom.dcmread(images[i])\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='hsv') \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"The Max week of the patient {max(data['Weeks'])}\")\nprint(f\"The Max week of the patient {min(data['Weeks'])}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Working On Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_Data = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = train_Data['Patient'].value_counts()[:25]\nlabs = train_Data['Patient'].unique()[:25]\n\nfig = plt.figure(figsize=(12, 5))\nsns.barplot(vals, labs)\n\nplt.xlabel('Count')\nplt.ylabel(\"Patients IDS\")\nplt.title(\"Unqiue patient Count out of duplicate\")\nplt.tight_layout()\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"nodupData = train_Data.drop_duplicates(subset = 'Patient', keep='first')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"nodupData.set_index('Patient', inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"nodupData","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nUnique_patients = list(train_Data['Patient'].unique())\n\n\ndef getting_group(groupID):\n    \n    return train_Data.groupby('Patient').get_group(groupID)\n\nprint(getting_group(Unique_patients[0]).plot())\nprint(getting_group(Unique_patients[1]).plot())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_Data = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Unique_patients == list(test_Data['Patient'].unique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_Data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"samplt_Data = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#samplt_Data['Patient_Week'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Conclusion\n\nWe have see many visualizationa and we got all about the  data\n\nI belive you have loved the repo \n\nPlease makeUp vote if you like ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}