{"cells":[{"metadata":{},"cell_type":"markdown","source":"<img src='https://www.osicild.org/uploads/1/2/2/7/122798879/editor/kaggle-v01-clipped.png?1569346633'>\n<h1><center>OSIC Pulmonary Fibrosis Progression - EDA</center><h1>\n\n# 0.注釈\n  本記事は https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda を日本語に訳したものになります。<br>\n  内容の誤りや誤訳については予めご容赦ください。\n    \n# 1. <a id='Introduction'>導入 🃏 </a>\n    \n###  1.1 肺線維症とは?\n* [肺線維症とは、肺の組織が傷ついて瘢痕化することで起こる肺の病気です。](https://www.mayoclinic.org/diseases-conditions/pulmonary-fibrosis/symptoms-causes/syc-20353690)  <br>この肥厚して硬くなった組織は、あなたの肺が正常に働くことをより困難にします。<br>あなたがこのタイプの肺の病気についてさらに知りたい場合は、有益なビデオを下にリンクしています。"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from IPython.display import HTML\nHTML('<center><iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/AfK9LPNj-Zo\" frameborder=\"0\" allow=\"accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe></center>')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"###  1.2 OSIC Pulmonary Fibrosis Progression コンペティションとは?\n- このコンペティションでは、肺のCTスキャンに基づいて患者の肺機能の低下の重症度を予測します。吸い込んだ空気と吐いた空気の量を測定するスピロメーターの出力に基づいて肺機能を判断します。課題は、画像、メタデータ、ベースラインFVCを入力として、機械学習技術を用いて予測を行うことです。\n    \n    \n###  1.3 我々は何をすべきか?\n肺のCTスキャンに基づいて、患者さんの肺機能の低下の重症度を予測します。\nつまり、各患者さんの最終的な3回のFVC測定値と、その予測の信頼値を予測します。\n    \n- 本大会のリーダーボードは、テストデータの約1%->15%で算出しています。他の99%->85%で計算されますので、最終的な順位は異なる可能性があります。\n    \n###  1.4 メトリック:ラプラス対数尤度\n![](https://i.imgur.com/tEIZvli.png)\n- 画像クレジット: https://en.wikipedia.org/wiki/Laplace_distribution\n    \n- このコンペティションの評価指標は、ラプラス対数尤度を修正したものです。\n予測値は、ラプラス対数尤度の修正版で評価されます。テストセットの各標本について、`FVC` と `Confidence` の尺度 (標準偏差σ) を予測します。\n\n    70より小さい`Confidence`値は切り取られます。\n\n    $\\large \\sigma_{clipped} = max(\\sigma, 70),$\n\n    また、1000を超えるエラーは、大きなエラーを避けるために切り取られています。\n\n    $\\large \\Delta = min ( |FVC_{true} - FVC_{predicted}|, 1000 ),$\n\n    メトリックは次のように定義されています。\n\n    $\\Large metric = -   \\frac{\\sqrt{2} \\Delta}{\\sigma_{clipped}} - \\ln ( \\sqrt{2} \\sigma_{clipped} ).$\n\n\nそれについての詳細は [Evaluation Page](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/overview/evaluation)にあります。\n\nもしこれが新鮮なもので、あなたにとって価値のあるものだったと感じたら、<font color='orange'>upvoting</font>を検討してみてください。 😄"},{"metadata":{},"cell_type":"markdown","source":"## <font size='5' color='blue'>目次</font> \n\n\n\n* [探索的データ分析の基礎](#1)  \n    * [はじめに - ライブラリのインポート]()\n    * [train.csvの読み込み]()\n    \n \n* [データの探索](#2)   \n     * [訓練・テストデータをチェック]()\n     * [ユニークな患者(ID)]()\n     * ['SmokingStatus' カラムの説明]()\n     * [週の分布]()\n     * [FVC - 強制的な生命維持能力]()\n     * [パーセント列の探索]()\n     * [性別の分布]()\n     * [患者の重なり]()\n     \n \n* [画像の可視化 : DECOM](#3)    \n     * [一つのDECOM画像と情報を可視化]()\n     * [複数のDECOM画像と情報を可視化]()\n     * [gifを用いて可視化]()\n     \n     \n* [DIOCOMファイルの抽出 Info.](#4)\n\n\n* [Pandas プロファイリング](#5)\n     * [Train.csvに対するPandasのプロファイリングレポート]()\n     * [Test.csvに対するPandasのプロファイリングレポート]()\n"},{"metadata":{},"cell_type":"markdown","source":"# 2. <a id='importing'>必要なライブラリのインポート📗</a> "},{"metadata":{"trusted":true,"_kg_hide-input":false,"_kg_hide-output":true},"cell_type":"code","source":"import os\nfrom os import listdir\nimport pandas as pd\nimport numpy as np\nimport glob\nimport tqdm\nfrom typing import Dict\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n# plotly\n!pip install chart_studio\nimport plotly.express as px\nimport chart_studio.plotly as py\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\n# color\nfrom colorama import Fore, Back, Style\n\nimport seaborn as sns\nsns.set(style=\"whitegrid\")\n\n# pydicom\nimport pydicom\n\n# 警告の抑制 \nimport warnings\nwarnings.filterwarnings('ignore')\n\n# プロットの設定\nplt.style.use('fivethirtyeight')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. <a id='reading'>train.csvの読み込み 📚</a>"},{"metadata":{"trusted":true},"cell_type":"code","source":"# 利用可能なファイル一覧\nlist(os.listdir(\"../input/osic-pulmonary-fibrosis-progression\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"IMAGE_PATH = \"../input/osic-pulmonary-fibrosis-progressiont/\"\n\ntrain_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\n\nprint(Fore.YELLOW + 'Training data shape: ',Style.RESET_ALL,train_df.shape)\ntrain_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df.groupby(['SmokingStatus']).count()['Sex'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. <a id='basic'>基本的なデータの探索 🏕️</a> "},{"metadata":{},"cell_type":"markdown","source":"## **一般的な情報**"},{"metadata":{"trusted":true},"cell_type":"code","source":"# NULL値とデータ型\nprint(Fore.YELLOW + 'Train Set !!',Style.RESET_ALL)\nprint(train_df.info())\nprint('-------------')\nprint(Fore.BLUE + 'Test Set !!',Style.RESET_ALL)\nprint(test_df.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Percentカラムの型はfloat64です。"},{"metadata":{},"cell_type":"markdown","source":"### 欠損値"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"train_dfとtest_dfには欠損値はありません。"},{"metadata":{"trusted":true},"cell_type":"code","source":"# データセットに含まれる患者の総数(train+test)\n\nprint(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.BLUE +\"Total Patients in Test set: \",Style.RESET_ALL,test_df['Patient'].count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" `5` : テストセットの患者数"},{"metadata":{},"cell_type":"markdown","source":"## ユニークな患者(ID)"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.YELLOW + \"患者IDの総数は\",Style.RESET_ALL,f\"{train_df['Patient'].count()},\", Fore.BLUE + \"これらの中から一意の ID は\", Style.RESET_ALL, f\"{train_df['Patient'].value_counts().shape[0]}.\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_patient_ids = set(train_df['Patient'].unique())\ntest_patient_ids = set(test_df['Patient'].unique())\n\ntrain_patient_ids.intersection(test_patient_ids)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"テストデータにはすでに `5` の患者が存在しており、訓練データにも同様に見つけることができる。"},{"metadata":{"trusted":true},"cell_type":"code","source":"columns = train_df.keys()\ncolumns = list(columns)\nprint(columns)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### **患者数**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"訓練データでは、1つの「Patient」に対して複数の行があります。なぜならPatientは週数、FVC、Percentが異なるためです。"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['Patient'].value_counts().max()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"テストデータでは、1つのPatientが表示され、Patient idが一意であることを意味しています。"},{"metadata":{"trusted":true},"cell_type":"code","source":"np.quantile(train_df['Patient'].value_counts(), 0.75) - np.quantile(test_df['Patient'].value_counts(), 0.25)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(np.quantile(train_df['Patient'].value_counts(), 0.95))\nprint(np.quantile(test_df['Patient'].value_counts(), 0.95))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## **トレーニング画像フォルダの患者数と画像数**\n* https://www.kaggle.com/yeayates21/osic-simple-image-eda"},{"metadata":{"trusted":true},"cell_type":"code","source":"files = folders = 0\n\npath = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train\"\n\nfor _, dirnames, filenames in os.walk(path):\n  # ^ これの意味は 「この値は使わない」\n    files += len(filenames)\n    folders += len(dirnames)\n#print(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\nprint(Fore.YELLOW +f'{files:,}',Style.RESET_ALL,\"files/images, \" + Fore.BLUE + f'{folders:,}',Style.RESET_ALL ,'folders/patients')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"files = []\nfor _, dirnames, filenames in os.walk(path):\n  # ^ これの意味は 「この値は使わない」\n    files.append(len(filenames))\n\nprint(Fore.YELLOW +f'{round(np.mean(files)):,}',Style.RESET_ALL,'average files/images per patient')\nprint(Fore.BLUE +f'{round(np.max(files)):,}',Style.RESET_ALL, 'max files/images per patient')\nprint(Fore.GREEN +f'{round(np.min(files)):,}',Style.RESET_ALL,'min files/images per patient')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 5. <a id='details'>データ探索の詳細 🎠</a> "},{"metadata":{},"cell_type":"markdown","source":"## 個々の患者データフレームの作成\n\n`175` のユニークな患者に対して、新しいデータフレームを作成します。\n\nThanks [@wjdanalharthi](https://www.kaggle.com/wjdanalharthi)"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df = train_df[['Patient', 'Age', 'Sex', 'SmokingStatus']].drop_duplicates()\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"こちらを利用できます:\n- https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-progression-interactive-eda"},{"metadata":{"trusted":true},"cell_type":"code","source":"# ユニークな患者リストとそのプロパティを作成する\ntrain_dir = '../input/osic-pulmonary-fibrosis-progression/train/'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test/'\n\npatient_ids = os.listdir(train_dir)\npatient_ids = sorted(patient_ids)\n\n#新しい行の作成\nno_of_instances = []\nage = []\nsex = []\nsmoking_status = []\n\nfor patient_id in patient_ids:\n    patient_info = train_df[train_df['Patient'] == patient_id].reset_index()\n    no_of_instances.append(len(os.listdir(train_dir + patient_id)))\n    age.append(patient_info['Age'][0])\n    sex.append(patient_info['Sex'][0])\n    smoking_status.append(patient_info['SmokingStatus'][0])\n\n#患者情報のデータフレームの作成    \npatient_df = pd.DataFrame(list(zip(patient_ids, no_of_instances, age, sex, smoking_status)), \n                                 columns =['Patient', 'no_of_instances', 'Age', 'Sex', 'SmokingStatus'])\nprint(patient_df.info())\npatient_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 'SmokingStatus' カラムの探索"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts().iplot(kind='bar',\n                                              yTitle='Counts', \n                                              linecolor='black', \n                                              opacity=0.7,\n                                              color='blue',\n                                              theme='pearl',\n                                              bargap=0.5,\n                                              gridcolor='white',\n                                              title='Distribution of the SmokingStatus column in the Unique Patient Set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`118` : 元喫煙者\n\n`49` : 一度も吸ったことがない\n\n`9` : 喫煙者"},{"metadata":{},"cell_type":"markdown","source":"## 週の分布"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].value_counts().head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].value_counts().iplot(kind='barh',\n                                      xTitle='Counts(Weeks)', \n                                      linecolor='black', \n                                      opacity=0.7,\n                                      color='#FB8072',\n                                      theme='pearl',\n                                      bargap=0.2,\n                                      gridcolor='white',\n                                      title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Weeks'].iplot(kind='hist',\n                              xTitle='Weeks', \n                              yTitle='Counts',\n                              linecolor='black', \n                              opacity=0.7,\n                              color='#FB8072',\n                              theme='pearl',\n                              bargap=0.2,\n                              gridcolor='white',\n                              title='Distribution of the Weeks in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Weeksにはマイナスの値もあります。\n\nなぜなら、WeeksはベースラインCTの前後の相対的な週数だからです。"},{"metadata":{},"cell_type":"markdown","source":"## 週の年齢分布"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Weeks\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## FVC - 強制的な生命維持能力"},{"metadata":{},"cell_type":"markdown","source":" 強制バイタル容量（FVC）、すなわち吐く空気の量。\n - 記録された肺活量（ml)"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['FVC'].iplot(kind='hist',\n                      xTitle='Lung Capacity(ml)', \n                      linecolor='black', \n                      opacity=0.8,\n                      color='#FB8072',\n                      bargap=0.5,\n                      gridcolor='white',\n                      title='Distribution of the FVC in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### FVCとパーセントの比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Percent\", color='Age')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"FVC は Percent と直線的な関係があるようで、理にかなっています。"},{"metadata":{},"cell_type":"markdown","source":"### FVCと年齢の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Age\", color='Sex')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"男性は年齢に関係なく女性よりもFVCが高いです。"},{"metadata":{},"cell_type":"markdown","source":"### FVCと週の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"FVC\", y=\"Weeks\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### ある患者におけるFVCと週の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = train_df[train_df.Patient == 'ID00228637202259965313869']\nfig = px.line(patient, x=\"Weeks\", y=\"FVC\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"タバコを吸ったことがない人は、喫煙者よりもFVCが低い。元喫煙者の中には非常に高いFVCを持つ人もいます。"},{"metadata":{},"cell_type":"markdown","source":"## パーセント"},{"metadata":{},"cell_type":"markdown","source":"患者のFVCを、類似した特徴を持つ人の典型的なFVCのパーセンテージとして近似した計算フィールド。"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Percent'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['Percent'].iplot(kind='hist',bins=30,color='blue',xTitle='Percent distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 患者のデータフレームにおけるパーセントと喫煙状態の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train_df\nfig = px.violin(df, y='Percent', x='SmokingStatus', box=True, color='Sex', points=\"all\",\n          hover_data=train_df.columns)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = train_df['SmokingStatus'], y = train_df['Percent'], palette = 'Reds')\nax.set_xlabel(xlabel = 'Smoking Habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Percent', fontsize = 15)\nax.set_title(label = 'Distribution of Smoking Status Over Percentage', fontsize = 20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.scatter(train_df, x=\"Age\", y=\"Percent\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### ある患者におけるFVCと週の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = train_df[train_df.Patient == 'ID00228637202259965313869']\nfig = px.line(patient, x=\"Weeks\", y=\"Percent\", color='SmokingStatus')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ユニークな患者の年齢分布"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Age'].iplot(kind='hist',bins=30,color='red',xTitle='Ages of distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 患者のデータフレームにおける年齢と喫煙状態の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Ex-smoker', 'Age'], label = 'Ex-smoker',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Never smoked', 'Age'], label = 'Never smoked',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['SmokingStatus'] == 'Currently smokes', 'Age'], label = 'Currently smokes', shade=True)\n\n# プロットのラベル付け\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nax = sns.violinplot(x = patient_df['SmokingStatus'], y = patient_df['Age'], palette = 'Reds')\nax.set_xlabel(xlabel = 'Smoking habit', fontsize = 15)\nax.set_ylabel(ylabel = 'Age', fontsize = 15)\nax.set_title(label = 'Distribution of Smokers over Age', fontsize = 20)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 患者のデータフレームにおける年齢と性別の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Male', 'Age'], label = 'Male',shade=True)\nsns.kdeplot(patient_df.loc[patient_df['Sex'] == 'Female', 'Age'], label = 'Female',shade=True)\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 性別の分布"},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Sex'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_df['Sex'].value_counts().iplot(kind='bar',\n                                          yTitle='Count', \n                                          linecolor='black', \n                                          opacity=0.7,\n                                          color='blue',\n                                          theme='pearl',\n                                          bargap=0.8,\n                                          gridcolor='white',\n                                          title='Distribution of the Sex column in Patient Dataframe')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`139` : 男性\n\n`37` : 女性"},{"metadata":{},"cell_type":"markdown","source":"### 患者のデータフレームにおける性別と喫煙状態の比較"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 6))\na = sns.countplot(data=patient_df, x='SmokingStatus', hue='Sex')\n\nfor p in a.patches:\n    a.annotate(format(p.get_height(), ','), \n           (p.get_x() + p.get_width() / 2., \n            p.get_height()), ha = 'center', va = 'center', \n           xytext = (0, 4), textcoords = 'offset points')\n\nplt.title('Gender split by SmokingStatus', fontsize=16)\nsns.despine(left=True, bottom=True);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.box(patient_df, x=\"Sex\", y=\"Age\", points=\"all\")\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"## 患者の重なり"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# トレーニングセットの患者IDを抽出\nids_train = train_df.Patient.values\n# バリデーションセットの患者IDを抽出\nids_test = test_df.Patient.values\n#print(Fore.YELLOW +\"Total Patients in Train set: \",Style.RESET_ALL,train_df['Patient'].count())\n# ユニークなIDを識別するために、学習セットIDの「セット」データ構造を作成\nids_train_set = set(ids_train)\nprint(Fore.YELLOW + \"There are\",Style.RESET_ALL,f'{len(ids_train_set)}', Fore.BLUE + 'unique Patient IDs',Style.RESET_ALL,'in the training set')\n# ユニークなIDを識別するために、検証セットIDの「セット」データ構造を作成\nids_test_set = set(ids_test)\nprint(Fore.YELLOW + \"There are\", Style.RESET_ALL, f'{len(ids_test_set)}', Fore.BLUE + 'unique Patient IDs',Style.RESET_ALL,'in the test set')\n\n# セット間の交点を見て、患者の重複を識別\npatient_overlap = list(ids_train_set.intersection(ids_test_set))\nn_overlap = len(patient_overlap)\nprint(Fore.YELLOW + \"There are\", Style.RESET_ALL, f'{n_overlap}', Fore.BLUE + 'Patient IDs',Style.RESET_ALL, 'in both the training and test sets')\nprint('')\nprint(Fore.CYAN + 'These patients are in both the training and test datasets:', Style.RESET_ALL)\nprint(f'{patient_overlap}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`5`の患者はトレーニングデータセットとテストデータセットの両方に存在する。"},{"metadata":{},"cell_type":"markdown","source":"## train.csvのヒートマップ"},{"metadata":{"trusted":true},"cell_type":"code","source":"corrmat = train_df.corr() \nf, ax = plt.subplots(figsize =(9, 8)) \nsns.heatmap(corrmat, ax = ax, cmap = 'RdYlBu_r', linewidths = 0.5) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"以前の可視化情報と比較してみてください。また、以下のPandas Profilingと比較することがあります。"},{"metadata":{},"cell_type":"markdown","source":"# 6. <a id='visual'>画像の可視化 : DECOM 🗺️</a>  "},{"metadata":{},"cell_type":"markdown","source":"`1` の情報を含む画像の種類:\n\n- `.dcm` ファイル: [DICOM files](https://en.wikipedia.org/wiki/DICOM). 医療におけるデジタル画像通信形式で保存されています。超音波やMRIなどのメディカルスキャンの画像＋患者の情報が入っています。"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.YELLOW + 'Train .dcm number of images:',Style.RESET_ALL, len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/train'))), '\\n' +\n      Fore.BLUE + 'Test .dcm number of images:',Style.RESET_ALL, len(list(os.listdir('../input/osic-pulmonary-fibrosis-progression/test'))), '\\n' +\n      '--------------------------------', '\\n' +\n      'There is the same number of images as in train/ test .csv datasets')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"DICOMの画像を見てみましょう。"},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.grid(False)\n    plt.imshow(dataset.pixel_array, cmap='gray') # cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data\ndef show_dcm_info(dataset):\n    print(Fore.YELLOW + \"Filename.........:\",Style.RESET_ALL,file_path)\n    print()\n\n    pat_name = dataset.PatientName\n    display_name = pat_name.family_name + \", \" + pat_name.given_name\n    print(Fore.BLUE + \"Patient's name......:\",Style.RESET_ALL, display_name)\n    print(Fore.BLUE + \"Patient id..........:\",Style.RESET_ALL, dataset.PatientID)\n    print(Fore.BLUE + \"Patient's Sex.......:\",Style.RESET_ALL, dataset.PatientSex)\n    print(Fore.YELLOW + \"Modality............:\",Style.RESET_ALL, dataset.Modality)\n    print(Fore.GREEN + \"Body Part Examined..:\",Style.RESET_ALL, dataset.BodyPartExamined)\n    \n    if 'PixelData' in dataset:\n        rows = int(dataset.Rows)\n        cols = int(dataset.Columns)\n        print(Fore.BLUE + \"Image size.......:\",Style.RESET_ALL,\" {rows:d} x {cols:d}, {size:d} bytes\".format(\n            rows=rows, cols=cols, size=len(dataset.PixelData)))\n        if 'PixelSpacing' in dataset:\n            print(Fore.YELLOW + \"Pixel spacing....:\",Style.RESET_ALL,dataset.PixelSpacing)\n            dataset.PixelSpacing = [1, 1]\n        plt.figure(figsize=(10, 10))\n        plt.imshow(dataset.pixel_array, cmap='gray')\n        plt.show()\nfor file_path in glob.glob('../input/osic-pulmonary-fibrosis-progression/train/*/*.dcm'):\n    dataset = pydicom.dcmread(file_path)\n    show_dcm_info(dataset)\n    break # すべてを見るには、これをコメントアウトしてください。","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# 最初の画像を順に表示\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='gray')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/yeayates21/osic-simple-image-eda\n\nimdir = \"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00123637202217151272140\"\nprint(\"total images for patient ID00123637202217151272140: \", len(os.listdir(imdir)))\n\n# 最初の画像を順に表示\nfig=plt.figure(figsize=(12, 12))\ncolumns = 4\nrows = 5\nimglist = os.listdir(imdir)\nfor i in range(1, columns*rows +1):\n    filename = imdir + \"/\" + str(i) + \".dcm\"\n    ds = pydicom.dcmread(filename)\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap='jet')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# gifを使った可視化\n* https://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"apply_resample = False\n\ndef load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def load_scan(path):\n    slices = [pydicom.read_file(path + '/' + s) for s in os.listdir(path)]\n    slices.sort(key = lambda x: float(x.ImagePositionPatient[2]))\n    try:\n        slice_thickness = np.abs(slices[0].ImagePositionPatient[2] - slices[1].ImagePositionPatient[2])\n    except:\n        slice_thickness = np.abs(slices[0].SliceLocation - slices[1].SliceLocation)\n        \n    for s in slices:\n        s.SliceThickness = slice_thickness\n        \n    return slices","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def get_pixels_hu(slices):\n    image = np.stack([s.pixel_array for s in slices])\n    # int16に変換 \n    # 常に低い値でなければならない (<32k)\n    image = image.astype(np.int16)\n\n    # スキャン外ピクセルを 0 に設定\n    # 切片は通常-1024\n    image[image == -2000] = 0\n    \n    # ハウンズフィールドのユニット（HU）に変換\n    for slice_number in range(len(slices)):\n        \n        intercept = slices[slice_number].RescaleIntercept\n        slope = slices[slice_number].RescaleSlope\n        \n        if slope != 1:\n            image[slice_number] = slope * image[slice_number].astype(np.float64)\n            image[slice_number] = image[slice_number].astype(np.int16)\n            \n        image[slice_number] += np.int16(intercept)\n    \n    return np.array(image, dtype=np.int16)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def set_lungwin(img, hu=[-1200., 600.]):\n    lungwin = np.array(hu)\n    newimg = (img-lungwin[0]) / (lungwin[1]-lungwin[0])\n    newimg[newimg < 0] = 0\n    newimg[newimg > 1] = 1\n    newimg = (newimg * 255).astype('uint8')\n    return newimg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"scans = load_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/')\nscan_array = set_lungwin(get_pixels_hu(scans))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# 1mmへの再サンプル（任意のステップで、z軸のスライス厚が大きいため、本大会には関係ないかもしれません）\n\nfrom scipy.ndimage.interpolation import zoom\n\ndef resample(imgs, spacing, new_spacing):\n    new_shape = np.round(imgs.shape * spacing / new_spacing)\n    true_spacing = spacing * imgs.shape / new_shape\n    resize_factor = new_shape / imgs.shape\n    imgs = zoom(imgs, resize_factor, mode='nearest')\n    return imgs, true_spacing, new_shape\n\nspacing_z = (scans[-1].ImagePositionPatient[2] - scans[0].ImagePositionPatient[2]) / len(scans)\n\nif apply_resample:\n    scan_array_resample = resample(scan_array, np.array(np.array([spacing_z, *scans[0].PixelSpacing])), np.array([1.,1.,1.]))[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import imageio\nfrom IPython.display import Image\n\nimageio.mimsave(\"/tmp/gif.gif\", scan_array, duration=0.0001)\nImage(filename=\"/tmp/gif.gif\", format='png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# アニメーションを使った可視化\n* https://www.kaggle.com/pranavkasela/animating-the-lung-ct-scan"},{"metadata":{},"cell_type":"markdown","source":"アニメーションには、「gifを使った可視化」で用いたscan_arrayを使用しています。"},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"import matplotlib.animation as animation\n\nfig = plt.figure()\n\nims = []\nfor image in scan_array:\n    im = plt.imshow(image, animated=True, cmap=\"Greys\")\n    plt.axis(\"off\")\n    ims.append([im])\n\nani = animation.ArtistAnimation(fig, ims, interval=100, blit=False,\n                                repeat_delay=1000)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HTML(ani.to_jshtml())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HTML(ani.to_html5_video())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 7. <a id='extract'>データフレーム内のDIOCOMファイル情報の抽出 🌊</a>"},{"metadata":{},"cell_type":"markdown","source":"* https://www.kaggle.com/trsekhar123/nb-to-extract-metadata-and-resize-images-train"},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_dicom_meta_data(filename: str) -> Dict:\n    # イメージの読み込み\n    \n    image_data = pydicom.read_file(filename)\n    img=np.array(image_data.pixel_array).flatten()\n    row = {\n        'Patient': image_data.PatientID,\n        'body_part_examined': image_data.BodyPartExamined,\n        'image_position_patient': image_data.ImagePositionPatient,\n        'image_orientation_patient': image_data.ImageOrientationPatient,\n        'photometric_interpretation': image_data.PhotometricInterpretation,\n        'rows': image_data.Rows,\n        'columns': image_data.Columns,\n        'pixel_spacing': image_data.PixelSpacing,\n        'window_center': image_data.WindowCenter,\n        'window_width': image_data.WindowWidth,\n        'modality': image_data.Modality,\n        'StudyInstanceUID': image_data.StudyInstanceUID,\n        'SeriesInstanceUID': image_data.StudyInstanceUID,\n        'StudyID': image_data.StudyInstanceUID, \n        'SamplesPerPixel': image_data.SamplesPerPixel,\n        'BitsAllocated': image_data.BitsAllocated,\n        'BitsStored': image_data.BitsStored,\n        'HighBit': image_data.HighBit,\n        'PixelRepresentation': image_data.PixelRepresentation,\n        'RescaleIntercept': image_data.RescaleIntercept,\n        'RescaleSlope': image_data.RescaleSlope,\n        'img_min': np.min(img),\n        'img_max': np.max(img),\n        'img_mean': np.mean(img),\n        'img_std': np.std(img)}\n\n    return row","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_image_path = '/kaggle/input/osic-pulmonary-fibrosis-progression/train'\ntrain_image_files = glob.glob(os.path.join(train_image_path, '*', '*.dcm'))\n\nmeta_data_df = []\nfor filename in tqdm.tqdm(train_image_files):\n    try:\n        meta_data_df.append(extract_dicom_meta_data(filename))\n    except Exception as e:\n        continue","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"dcomファイルのmeta要素はファイルごとに微妙に違うようです。そのためExceptionを使っています。\n- 気をつけてください！"},{"metadata":{"trusted":true},"cell_type":"code","source":"# dict から pd.DataFrame への変換\nmeta_data_df = pd.DataFrame.from_dict(meta_data_df)\nmeta_data_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"下のコードはまだ意味があるので、そのままにしておきます。"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# ソース: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154658\nfolder='train'\nPATH='../input/osic-pulmonary-fibrosis-progression/'\n\nlast_index = 2\n\ncolumn_names = ['image_name', 'dcm_ImageOrientationPatient', \n                'dcm_ImagePositionPatient', 'dcm_PatientID',\n                'dcm_PatientName', 'dcm_PatientSex'\n                'dcm_rows', 'dcm_columns']\n\ndef extract_DICOM_attributes(folder):\n    patients_folder = list(os.listdir(os.path.join(PATH, folder)))\n    df = pd.DataFrame()\n    \n    i = 0\n    \n    for patient_id in patients_folder:\n   \n        img_path = os.path.join(PATH, folder, patient_id)\n        \n        print(img_path)\n        \n        images = list(os.listdir(img_path))\n        \n        #df = pd.DataFrame()\n\n        for image in images:\n            image_name = image.split(\".\")[0]\n\n            dicom_file_path = os.path.join(img_path,image)\n            dicom_file_dataset = pydicom.read_file(dicom_file_path)\n                \n            '''\n            print(dicom_file_dataset.dir(\"pat\"))\n            print(dicom_file_dataset.data_element(\"ImageOrientationPatient\"))\n            print(dicom_file_dataset.data_element(\"ImagePositionPatient\"))\n            print(dicom_file_dataset.data_element(\"PatientID\"))\n            print(dicom_file_dataset.data_element(\"PatientName\"))\n            print(dicom_file_dataset.data_element(\"PatientSex\"))\n            '''\n            \n            imageOrientationPatient = dicom_file_dataset.ImageOrientationPatient\n            #imagePositionPatient = dicom_file_dataset.ImagePositionPatient\n            patientID = dicom_file_dataset.PatientID\n            patientName = dicom_file_dataset.PatientName\n            patientSex = dicom_file_dataset.PatientSex\n        \n            rows = dicom_file_dataset.Rows\n            cols = dicom_file_dataset.Columns\n            \n            #print(rows)\n            #print(columns)\n            \n            temp_dict = {'image_name': image_name, \n                                    'dcm_ImageOrientationPatient': imageOrientationPatient,\n                                    #'dcm_ImagePositionPatient':imagePositionPatient,\n                                    'dcm_PatientID': patientID, \n                                    'dcm_PatientName': patientName,\n                                    'dcm_PatientSex': patientSex,\n                                    'dcm_rows': rows,\n                                    'dcm_columns': cols}\n\n\n            df = df.append([temp_dict])\n            \n        i += 1\n        \n        if i == last_index:\n            break\n            \n    return df\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"extract_DICOM_attributes('train')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"dcomファイルのmeta要素はファイルごとに微妙に違うようです。\n- 気をつけてください!"},{"metadata":{},"cell_type":"markdown","source":"# <a id='etc'>補論 - Pandas プロファイリング 🌤️</a>"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas_profiling as pdp","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest_df = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"profile_train_df = pdp.ProfileReport(train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_train_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"profile_test_df = pdp.ProfileReport(test_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_test_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## このカーネルが役に立ったら、<font color='orange'>please upvote</font>!\n- また次回お会いしましょう!"},{"metadata":{},"cell_type":"markdown","source":"# 参考\n- https://www.kaggle.com/tarunpaparaju/siim-isic-melanoma-eda-pytorch-baseline\n- https://www.kaggle.com/andradaolteanu/siim-melanoma-competition-eda-augmentations\n- https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter\n- https://www.kaggle.com/yeayates21/osic-simple-image-eda\n- https://www.kaggle.com/aviralpamecha/osic-in-depth-and-advanced-analysis\n- https://www.kaggle.com/danpresil1/dicom-basic-preprocessing-and-visualization"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}