{"cells":[{"metadata":{},"cell_type":"markdown","source":"<h1 class=\"list-group-item list-group-item-success\" data-toggle=\"list\"  role=\"tab\" aria-controls=\"home\">EDA - OSIC Pulmonary Fibrosis Progression</h1>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn%3AANd9GcSroWe00EY1yelEbhqMg9L8-yjHQWTB5amO8w&usqp=CAU\" alt=\"Meatball Sub\" width=\"800\"/>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Pulmonary fibrosis** : Pulmonary fibrosis is a lung disease that occurs when lung tissue becomes damaged and scarred. This thickened, stiff tissue makes it more difficult for your lungs to work properly. As pulmonary fibrosis worsens, you become progressively more short of breath.\n\nThe scarring associated with pulmonary fibrosis can be caused by a multitude of factors. But in most cases, doctors can't pinpoint what's causing the problem. When a cause can't be found, the condition is termed idiopathic pulmonary fibrosis. [link](https://www.mayoclinic.org/diseases-conditions/pulmonary-fibrosis/symptoms-causes/syc-20353690#:~:text=Pulmonary%20fibrosis%20is%20a%20lung,your%20lungs%20to%20work%20properly.)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from IPython.display import YouTubeVideo\nYouTubeVideo('AfK9LPNj-Zo', width=800, height=300)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd \nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nfrom plotly.offline import init_notebook_mode, iplot\nimport plotly.graph_objects as go\nfrom ipywidgets import widgets\nfrom ipywidgets import *\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set(style=\"darkgrid\")\nfrom scipy.signal import find_peaks\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/osic-pulmonary-fibrosis-progression/train.csv\",delimiter=\",\",encoding=\"latin\", engine='python')\ntest = pd.read_csv(\"../input/osic-pulmonary-fibrosis-progression/test.csv\",delimiter=\",\",encoding=\"latin\", engine='python')\ntrain.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train.info())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.dtypes.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1 - Quik analyse ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"count = train['Patient'].value_counts() \nprint(count) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Number of Patient in the train set {}\".format(len( train['Patient'].unique()))) \nprint(\"Number of Patient in the test set {}\".format(len( test['Patient'].unique()))) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">  \n<b> Observation 1 :</b> Number of Patient in the train set = 176 and the number of patient in the test set = 5. \n</div>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### 2 - FVC (Forced Vital Capacity)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Forced vital capacity (FVC) is the total amount of air exhaled during the FEV test.**\n\nForced expiratory volume and forced vital capacity are lung function tests that are measured during spirometry. Forced expiratory volume is the most important measurement of lung function. It is used to:\n- Diagnose obstructive lung diseases such as asthma and chronic obstructive pulmonary disease (COPD). A person who has asthma or COPD has a lower FEV1 result than a healthy person.\n- See how well medicines used to improve breathing are working.\n- Check if lung disease is getting worse. Decreases in the FEV1 value may mean the lung disease is getting worse.\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"YouTubeVideo('BmYCAp4dRuA', width=800, height=300)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**How are FEV1 and FVC Measured?**\n\nForced expiratory volume in one second (FEV1) and forced vital capacity (FVC) are measured during a pulmonary function test. A diagnostic device called a spirometer measures the amount of air you inhale, exhale and the amount of time it takes for you to exhale completely after a deep breath. For pulmonary function tests, the spirometer attaches to a machine that records your lung function measurements.\n\n**What is FVC?** \n\nThe forced vital capacity (FVC) measurement shows the amount of air a person can forcefully and quickly exhale after taking a deep breath.\nDetermining your FVC helps your doctor diagnose a chronic lung disease, monitor the disease over time and understand the severity of the condition. In general, doctors compare your FVC measurement with the predicted FVC based on your age, height and weight.\n\n**What is FEV1?** \n\nForced expiratory volume is measured during the forced vital capacity test. The forced expiratory volume in one second (FEV1) measurement shows the amount of air a person can forcefully exhale in one second of the FVC test. In addition, doctors can measure forced expiratory volume during the second and third seconds of the FVC test.\nDetermining your FEV1 measurement helps your doctor understand the severity of disease. Typically, lower FEV1 scores show more severe stages of lung disease.\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize = (20, 10))\nax = fig.add_subplot()\ni = 0 \nfor id_patient in train[\"Patient\"].unique()[0:6] : \n    y = train[train[\"Patient\"] == id_patient][\"FVC\"].reset_index(drop=True)\n    df = train[train[\"Patient\"] == id_patient].reset_index(drop=True)\n    max_peaks_index, _ = find_peaks(y, height=0) \n    doublediff2 = np.diff(np.sign(np.diff(-1*y))) \n    min_peaks_index = np.where(doublediff2 == -2)[0] + 1\n    ax.plot(y, color = \"blue\", alpha = .6)\n\n    if i == 0:\n        ax.scatter(x = y[max_peaks_index].index, y = y[max_peaks_index].values, marker = \"^\", s = 150, color = \"green\", alpha = .6, label = \"Peaks\")\n        ax.scatter(x = y[min_peaks_index].index, y = y[min_peaks_index].values, marker = \"v\", s = 150, color = \"red\", alpha = .6, label = \"Troughs\")\n    else :\n        ax.scatter(x = y[max_peaks_index].index, y = y[max_peaks_index].values, marker = \"^\", s = 150, color = \"green\", alpha = .6)\n        ax.scatter(x = y[min_peaks_index].index, y = y[min_peaks_index].values, marker = \"v\", s = 150, color = \"red\", alpha = .6)\n    for max_annot in max_peaks_index[:] :\n        for min_annot in min_peaks_index[:] :\n\n            max_text = df.iloc[max_annot][\"FVC\"]\n            min_text = df.iloc[min_annot][\"FVC\"]\n\n            max_text_w = df.iloc[max_annot][\"Weeks\"]\n            min_text_w = df.iloc[min_annot][\"Weeks\"]\n\n            ax.text(df.index[max_annot], y[max_annot] + 50, s = max_text, fontsize = 12, horizontalalignment = 'center', verticalalignment = 'center')\n            ax.text(df.index[min_annot], y[min_annot] + 50, s = min_text, fontsize = 12, horizontalalignment = 'center', verticalalignment = 'center')\n\n            ax.text(df.index[max_annot], y[max_annot] - 50, s = \"Week : \" + str(max_text_w), fontsize = 12, horizontalalignment = 'center', verticalalignment = 'center')\n            ax.text(df.index[min_annot], y[min_annot] - 50, s = \"Week : \" + str(min_text_w), fontsize = 12, horizontalalignment = 'center', verticalalignment = 'center')\n    ax.text(df.index[0], y[0] + 30, s = id_patient, fontsize = 10, horizontalalignment = 'center', verticalalignment = 'center')\n    i = i + 1\n    ax.legend(loc = \"upper left\", fontsize = 10)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['FVC_mean'] = train['FVC'].groupby(train['Patient']).transform('mean')\ntrain['FVC_max'] = train['FVC'].groupby(train['Patient']).transform('max')\ntrain['FVC_min'] = train['FVC'].groupby(train['Patient']).transform('min')\ntrain['FVC_std'] = train['FVC'].groupby(train['Patient']).transform('std')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize = (12, 6))\nax = fig.add_subplot(111) \n\nfor Smoking in sorted(list(train[\"SmokingStatus\"].unique())):\n    Age = train[train[\"SmokingStatus\"] == Smoking][\"Age\"]\n    FVC_mean = train[train[\"SmokingStatus\"] == Smoking][\"FVC_mean\"]\n    ax.scatter(Age, FVC_mean, label = Smoking, s = 10)\n\nax.spines[\"top\"].set_color(\"None\") \nax.spines[\"right\"].set_color(\"None\")\nax.set_xlabel(\"Age\") \nax.set_ylabel(\"FVC_mean\")\nax.set_title(\"Scatter plot of Age vs FVC_mean.\")\nax.legend(loc = \"upper left\", fontsize = 10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize = (12, 6))\nax = fig.add_subplot(111) \n\nfor Smoking in sorted(list(train[\"SmokingStatus\"].unique())):\n    Percent = train[train[\"SmokingStatus\"] == Smoking][\"Percent\"]\n    FVC_mean = train[train[\"SmokingStatus\"] == Smoking][\"FVC_mean\"]\n    ax.scatter(Percent, FVC_mean, label = Smoking, s = 10)\n\nax.spines[\"top\"].set_color(\"None\") \nax.spines[\"right\"].set_color(\"None\")\n\nax.set_xlabel(\"Percent\") \nax.set_ylabel(\"FVC_mean\")\n\nax.set_title(\"Scatter plot of Percent vs FVC_mean.\")\nax.legend(loc = \"upper left\", fontsize = 10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import squarify\nlabel_value = train[\"SmokingStatus\"].value_counts().to_dict()\nlabels = [\"{} has {} obs\".format(class_, obs) for class_, obs in label_value.items()]\ncolors = [plt.cm.Spectral(i/float(len(labels))) for i in range(len(labels))]\nplt.figure(figsize = (10, 5))\nsquarify.plot(sizes = label_value.values(), label = labels,  color = colors, alpha = 0.8)\nplt.title(\"Smoking Status\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(20, 9))\np = sns.boxplot(x='Sex', y='Age', hue='SmokingStatus', data=train, ax=axes[0])\np.set_title('train')\n\np = sns.boxplot(x='Sex', y='FVC_mean', hue='SmokingStatus', data=train, ax=axes[1])\np.set_title('train')\n\np = sns.boxplot(x='Sex', y='FVC_std', hue='SmokingStatus', data=train, ax=axes[2])\np.set_title('train')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(20, 9))\n\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"Percent\"]\n    g1 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0])\n    g1.set_title('Percent vs SmokingStatus')\ng1.legend()\n\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"Age\"]\n    g2 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1])\n    g2.set_title('Age vs SmokingStatus')\ng2.legend()\n   ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(20, 9))\n\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"Percent\"]\n    g1 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0])\n    g1.set_title('Percent vs Sex')\ng1.legend()\n\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"Age\"]\n    g2 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1])\n    g2.set_title('Age vs Sex')\ng2.legend()\n   ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(20, 11))\n\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"FVC_mean\"]\n    g1 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0,0])\n    g1.set_title('FVC_mean vs SmokingStatus')\ng1.legend()\n\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"FVC_std\"]\n    g2 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0,1])\n    g2.set_title('FVC_std vs SmokingStatus')\ng2.legend()\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"FVC_max\"]\n    g3 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1,0])\n    g3.set_title('FVC_max vs SmokingStatus')\ng3.legend()\n\nfor s in train[\"SmokingStatus\"].unique():\n    x = train[train[\"SmokingStatus\"] == s][\"FVC_min\"]\n    g4 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1,1])\n    g4.set_title('FVC_min vs SmokingStatus')\ng4.legend()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(20, 11))\n\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"FVC_mean\"]\n    g1 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0,0])\n    g1.set_title('FVC_mean vs Sex')\ng1.legend()\n\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"FVC_std\"]\n    g2 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[0,1])\n    g2.set_title('FVC_std vs Sex')\ng2.legend()\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"FVC_max\"]\n    g3 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1,0])\n    g3.set_title('FVC_max vs Sex')\ng3.legend()\n\nfor s in train[\"Sex\"].unique():\n    x = train[train[\"Sex\"] == s][\"FVC_min\"]\n    g4 = sns.distplot(x, kde = True, label = \"{}\".format(s), ax=axes[1,1])\n    g4.set_title('FVC_min vs Sex')\ng4.legend()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### To be continued ...","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}