{"cells":[{"metadata":{},"cell_type":"markdown","source":"[Melanoma Classification : EDA starter](https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter/data)の説明を日本語化しました","execution_count":null},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div align='center'><font size=\"6\" color=\"#F39C12\">SIIM-ISIC Melanoma Classification -EDA </font></div>\n<hr>\n\n![](https://impactmelanoma.org/wp-content/uploads/2018/11/Standard-Infographic_0.jpg)\nhttps://impactmelanoma.org/wp-content/uploads/2018/11/Standard-Infographic_0.jpg\n\nメラノーマは、肌の色を決める色素（メラニン）を作るメラノサイトと呼ばれる皮膚細胞から発生する皮膚がんです。メラノーマは、急速に全身に広がるため、皮膚がんの中で最も危険なタイプと考えられていますが、早期に発見されれば、一般的には非常に治療が可能です。\nhttps://www.verywellhealth.com/what-is-melanoma-514215\n\n[Society for Imaging Informatics in Medicine (SIIM)](https://siim.org/page/about_siim)は、医用画像情報学の現在および将来の利用に関心を持つ人々のための主要な医療専門組織です。この学会の使命は、学際的なコミュニティでの教育、研究、イノベーションを通じて、企業全体で医用画像情報学を発展させることにあります。[The International Skin Imaging Collaboration or ISIC](https://siim.org/page/about_siim) Melanoma Projectは、メラノーマの死亡率を減らすために皮膚のデジタル画像処理の応用を促進することを目的とした、学界と産業界のパートナーシップです。\n\nISICメラノーマプロジェクトの包括的な目標は、メラノーマの早期発見の精度と効率を向上させることで、メラノーマに関連した死亡や不必要な生検を減らす努力を支援することです。\n\n## 目的\n\nこのコンテストの目的は、皮膚病変の画像からメラノーマを特定することです。具体的には、同一患者内の画像を用いて、どれがメラノーマを表す可能性が高いかを判断する必要がある。つまり、画像中の病変が悪性か良性かの確率を予測するモデルを作成する必要があります。\n\n## データセット\nデータセットは、.DIOCOM形式の画像で構成されています。\n* DIOCOMフォーマット\n* JPEGディレクトリ内のJPEG形式\n* tfrecords ディレクトリの TFRecord のフォーマット\n\nまた、トレーニング、テスト、提出ファイルからなるメタデータがCSV形式で提供されています。\n\n## 評価指標について\n\nこの問題では、我々の提出物は、**area under the ROC curve**を用いて評価されます。ROC曲線（受信機動作特性曲線）とは、すべての分類しきい値における分類モデルの性能を示すグラフです。この曲線は、2つのパラメータをプロットします。\n\n![](https://imgur.com/yNeAG4M.png)\n\nROC曲線は、異なる分類しきい値でのTPR対FPRをプロットしたものです。分類しきい値を下げると、より多くの項目が陽性として分類され、その結果、偽陽性と真陽性の両方が増加します。次の図は、典型的なROC曲線を示しています。\n\n![](https://imgur.com/N3UOcBF.png)\n\nsource: https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"\n# 1. 必要なライブラリのインポート\n\nノートブックをフォークする場合に備えて、インターネットを `ON` モードにしておいてください。","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"\nfrom os import listdir\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n#plotly\n!pip install chart_studio\nimport plotly.express as px\nimport chart_studio.plotly as py\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True, theme='pearl')\n\nimport seaborn as sns\nsns.set(style=\"whitegrid\")\n\n\n#pydicom\nimport pydicom\n\n# Suppress warnings \nimport warnings\nwarnings.filterwarnings('ignore')\n\n\n# Settings for pretty nice plots\nplt.style.use('fivethirtyeight')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. 画像データセットの読み込み","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# List files available\nprint(os.listdir(\"../input/siim-isic-melanoma-classification\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Defining data path\nIMAGE_PATH = \"../input/siim-isic-melanoma-classification/\"\n\ntrain_df = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest_df = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')\n\n\n#Training data\nprint('Training data shape: ', train_df.shape)\ntrain_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Test data\nprint('Test data shape: ', test_df.shape)\ntest_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.groupby(['benign_malignant']).count()['sex'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. データ探査\n\n## 欠損値","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Null values and Data types\nprint('Train Set')\nprint(train_df.info())\nprint('-------------')\nprint('Test Set')\nprint(test_df.info())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"いくつかの列に欠けている値があります。これらは後ほど処理します。","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## 画像数","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Total number of images in the dataset(train+test)\nprint(\"Total images in Train set: \",train_df['image_name'].count())\nprint(\"Total images in Test set: \",test_df['image_name'].count())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## TrainユニークID","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"The total patient ids are {train_df['patient_id'].count()}, from those the unique ids are {train_df['patient_id'].value_counts().shape[0]} \")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## TestユニークID","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"The total patient ids are {test_df['patient_id'].count()}, from those the unique ids are {test_df['patient_id'].value_counts().shape[0]} \")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"固有の患者数は、総患者数よりも少ない。これは、患者が複数のレコードを持っていることを意味します。","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"columns = train_df.keys()\ncolumns = list(columns)\nprint(columns)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ターゲットカラム探索","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['target'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df['target'].value_counts(normalize=True).iplot(kind='bar',\n                                                      yTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='red',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',\n                                                     \n                                                      title='Distribution of the Target column in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Training男女別分布\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['sex'].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_df['sex'].value_counts(normalize=True).iplot(kind='bar',\n                                                      yTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='green',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',\n                                                     \n                                                      title='Distribution of the Sex column in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 性別 vs ターゲット","execution_count":null},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"z=train_df.groupby(['target','sex'])['benign_malignant'].count().to_frame().reset_index()\nz.style.background_gradient(cmap='Reds')  ","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"sns.catplot(x='target',y='benign_malignant', hue='sex',data=z,kind='bar')\nplt.ylabel('Count')\nplt.xlabel('benign:0 vs malignant:1')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## イメージサイトの場所","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['anatom_site_general_challenge'].value_counts(normalize=True).sort_values()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df['anatom_site_general_challenge'].value_counts(normalize=True).sort_values().iplot(kind='barh',\n                                                      xTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='#FB8072',\n                                                      theme='pearl',\n                                                      bargap=0.2,\n                                                      gridcolor='white',\n                                                      title='Distribution of the imaged site in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 性別によらないイメージサイトの位置","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"z1=train_df.groupby(['sex','anatom_site_general_challenge'])['benign_malignant'].count().to_frame().reset_index()\nz1.style.background_gradient(cmap='Reds')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"sns.catplot(x='anatom_site_general_challenge',y='benign_malignant', hue='sex',data=z1,kind='bar')\nplt.gcf().set_size_inches(10,8)\nplt.xlabel('location of imaged site')\nplt.xticks(rotation=45,fontsize='10', horizontalalignment='right')\nplt.ylabel('count of melanoma cases')\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 患者さんの年齢分布","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['age_approx'].iplot(kind='hist',bins=30,color='orange',xTitle='Age distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 年齢を視覚化するKDEs\n密度プロットでデータを要約し、データの質量がどこにあるかを確認します。[カーネル密度推定プロット](https://chemicalstatistician.wordpress.com/2013/06/09/exploratory-data-analysis-kernel-density-estimation-in-r-on-ozone-pollution-data-in-new-york-and-ozonopolis/)は、単一の変数の分布を示し、平滑化ヒストグラムと考えることができます（これは、各データ点でカーネル（通常はガウス分布）を計算し、すべての個々のカーネルを平均化して、単一の平滑曲線を作成することによって作成されます）。このグラフには、seaborn kdeplotを使用します。\n\n### 年齢の分布（目標値と比較して","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# KDE plot of age that were diagnosed as benign\nsns.kdeplot(train_df.loc[train_df['target'] == 0, 'age_approx'], label = 'Benign',shade=True)\n\n# KDE plot of age that were diagnosed as malignant\nsns.kdeplot(train_df.loc[train_df['target'] == 1, 'age_approx'], label = 'Malignant',shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 男女別の年齢分布","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# KDE plot of age that were diagnosed as benign\nsns.kdeplot(train_df.loc[train_df['sex'] == 'male', 'age_approx'], label = 'Male',shade=True)\n\n# KDE plot of age that were diagnosed as malignant\nsns.kdeplot(train_df.loc[train_df['sex'] == 'female', 'age_approx'], label = 'Female',shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 診断結果の分布","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['diagnosis'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_df['diagnosis'].value_counts(normalize=True).sort_values().iplot(kind='barh',\n                                                      xTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='blue',\n                                                      theme='pearl',\n                                                      bargap=0.2,\n                                                      gridcolor='white',\n                                                      title='Distribution in the training set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test男女別分布","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['sex'].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['sex'].value_counts(normalize=True).iplot(kind='bar',\n                                                      yTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='green',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',\n                                                     \n                                                      title='Distribution of the Sex column in the test set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test画像撮影部位","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['anatom_site_general_challenge'].value_counts(normalize=True).sort_values()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['anatom_site_general_challenge'].value_counts(normalize=True).sort_values().iplot(kind='barh',\n                                                      xTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='#FB8072',\n                                                      theme='pearl',\n                                                      bargap=0.2,\n                                                      gridcolor='white',\n                                                      title='Distribution of the imaged site in the test set')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test 患者さんの年齢分布","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['age_approx'].iplot(kind='hist',bins=30,color='orange',xTitle='Age distribution',yTitle='Count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test 男女別の年齢分布","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# KDE plot of age that were diagnosed as benign\nsns.kdeplot(test_df.loc[test_df['sex'] == 'male', 'age_approx'], label = 'Male',shade=True)\n\n# KDE plot of age that were diagnosed as malignant\nsns.kdeplot(test_df.loc[test_df['sex'] == 'female', 'age_approx'], label = 'Female',shade=True)\n\n# Labeling of plot\nplt.xlabel('Age (years)'); plt.ylabel('Density'); plt.title('Distribution of Ages');\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 患者のオーバーラップ \nトレーニングセットとテストセットの両方に同じ患者の病変画像が現れないことを確認する必要があります。","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Extract patient id's for the training set\nids_train = train_df.patient_id.values\n# Extract patient id's for the validation set\nids_test = test_df.patient_id.values\n\n# Create a \"set\" datastructure of the training set id's to identify unique id's\nids_train_set = set(ids_train)\nprint(f'There are {len(ids_train_set)} unique Patient IDs in the training set')\n# Create a \"set\" datastructure of the validation set id's to identify unique id's\nids_test_set = set(ids_test)\nprint(f'There are {len(ids_test_set)} unique Patient IDs in the training set')\n\n# Identify patient overlap by looking at the intersection between the sets\npatient_overlap = list(ids_train_set.intersection(ids_test_set))\nn_overlap = len(patient_overlap)\nprint(f'There are {n_overlap} Patient IDs in both the training and test sets')\nprint('')\nprint(f'These patients are in both the training and test datasets:')\nprint(f'{patient_overlap}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# 4. 画像の可視化 . JPEG\n\n## ランダムに選択された画像を可視化する","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"images = train_df['image_name'].values\n\n# Extract 9 random images from it\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Display Random Images')\n\n# Adjust the size of your images\nplt.figure(figsize=(10,8))\n\n# Iterate and plot random images\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \n# Adjust subplot parameters to give specified padding\nplt.tight_layout()   ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"JPEG形式の画像はサイズが異なることがわかります。","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"## 良性病変の画像の可視化","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"benign = train_df[train_df['benign_malignant']=='benign']\nmalignant = train_df[train_df['benign_malignant']=='malignant']","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"images = benign['image_name'].values\n\n# Extract 9 random images from it\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Display benign Images')\n\n# Adjust the size of your images\nplt.figure(figsize=(10,8))\n\n# Iterate and plot random images\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \n# Adjust subplot parameters to give specified padding\nplt.tight_layout()   ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 悪性病変の画像の可視化","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"images = malignant['image_name'].values\n\n# Extract 9 random images from it\nrandom_images = [np.random.choice(images+'.jpg') for i in range(9)]\n\n# Location of the image dir\nimg_dir = IMAGE_PATH+'/jpeg/train'\n\nprint('Display malignant Images')\n\n# Adjust the size of your images\nplt.figure(figsize=(10,8))\n\n# Iterate and plot random images\nfor i in range(9):\n    plt.subplot(3, 3, i + 1)\n    img = plt.imread(os.path.join(img_dir, random_images[i]))\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \n# Adjust subplot parameters to give specified padding\nplt.tight_layout()   ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ヒストグラム\n\nヒストグラムは、画像内で様々な色の値がどのくらいの頻度で発生しているか、つまりピクセルの強度値の頻度をグラフィカルに表したものです。RGB色空間では、ピクセルの値は0から255までの範囲で、0は黒、255は白を表します。ヒストグラムを分析することで、画像の明るさ、コントラスト、強度分布を理解することができます。ここで、各カテゴリからランダムに選択したサンプルのヒストグラムを見てみましょう。\n\n### 良性のカテゴリー","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"f = plt.figure(figsize=(16,8))\nf.add_subplot(1,2, 1)\n\nsample_img = benign['image_name'][0]+'.jpg'\nraw_image = plt.imread(os.path.join(img_dir, sample_img))\nplt.imshow(raw_image, cmap='gray')\nplt.colorbar()\nplt.title('Benign Image')\nprint(f\"Image dimensions:  {raw_image.shape[0],raw_image.shape[1]}\")\nprint(f\"Maximum pixel value : {raw_image.max():.1f} ; Minimum pixel value:{raw_image.min():.1f}\")\nprint(f\"Mean value of the pixels : {raw_image.mean():.1f} ; Standard deviation : {raw_image.std():.1f}\")\n\nf.add_subplot(1,2, 2)\n\n#_ = plt.hist(raw_image.ravel(),bins = 256, color = 'orange',)\n_ = plt.hist(raw_image[:, :, 0].ravel(), bins = 256, color = 'red', alpha = 0.5)\n_ = plt.hist(raw_image[:, :, 1].ravel(), bins = 256, color = 'Green', alpha = 0.5)\n_ = plt.hist(raw_image[:, :, 2].ravel(), bins = 256, color = 'Blue', alpha = 0.5)\n_ = plt.xlabel('Intensity Value')\n_ = plt.ylabel('Count')\n_ = plt.legend(['Red_Channel', 'Green_Channel', 'Blue_Channel'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 悪性カテゴリー","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"f = plt.figure(figsize=(16,8))\nf.add_subplot(1,2, 1)\n\nsample_img = malignant['image_name'][235]+'.jpg'\nraw_image = plt.imread(os.path.join(img_dir, sample_img))\nplt.imshow(raw_image, cmap='gray')\nplt.colorbar()\nplt.title('Malignant Image')\nprint(f\"Image dimensions:  {raw_image.shape[0],raw_image.shape[1]}\")\nprint(f\"Maximum pixel value : {raw_image.max():.1f} ; Minimum pixel value:{raw_image.min():.1f}\")\nprint(f\"Mean value of the pixels : {raw_image.mean():.1f} ; Standard deviation : {raw_image.std():.1f}\")\n\nf.add_subplot(1,2, 2)\n\n#_ = plt.hist(raw_image.ravel(),bins = 256, color = 'orange',)\n_ = plt.hist(raw_image[:, :, 0].ravel(), bins = 256, color = 'red', alpha = 0.5)\n_ = plt.hist(raw_image[:, :, 1].ravel(), bins = 256, color = 'Green', alpha = 0.5)\n_ = plt.hist(raw_image[:, :, 2].ravel(), bins = 256, color = 'Blue', alpha = 0.5)\n_ = plt.xlabel('Intensity Value')\n_ = plt.ylabel('Count')\n_ = plt.legend(['Red_Channel', 'Green_Channel', 'Blue_Channel'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# 5 DIOCOMファイルの前処理 \nDICOMは、医用画像を保存・伝送するために最も一般的に使用されており、スキャナー、サーバー、ワークステーション、プリンター、ネットワークハードウェア、複数のメーカーの画像アーカイブ通信システム（PACS）などの医用画像機器を統合することを可能にしています。\n\nDICOM 画像の拡張子は dcm です。DICOMファイルには，ヘッダとデータセットの2つの部分がある。ヘッダは，カプセル化されたデータセットに関する情報を含む。これは、ファイルプリアンブル、DICOMプレフィックス、ファイルメタ要素で構成されています。\n幸いなことに、Pydicomと呼ばれるPythonのライブラリがあり、これを使ってDIOCOMファイルを読むことができます。pydicomを使うと、これらの複雑なファイルを簡単に自然なパイソニック構造に読み込んで簡単に操作することができます。pydicomは、これらの複雑なファイルを簡単に自然なピソ構造に読み込んで、簡単に操作できるようにしてくれます。変更したデータセットは、DICOM形式のファイルに再度書き込むことができます。\n\n数年前のコンテストで発表された、DIOCOM画像ファイルへの素晴らしい入門書となる非常に素晴らしい[kernel](https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data)があります。\nカーネル: https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print (pydicom.__version__)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# https://www.kaggle.com/schlerp/getting-to-know-dicom-and-the-data\ndef show_dcm_info(dataset):\n    print(\"Filename.........:\", file_path)\n    print(\"Storage type.....:\", dataset.SOPClassUID)\n    print()\n\n    pat_name = dataset.PatientName\n    display_name = pat_name.family_name + \", \" + pat_name.given_name\n    print(\"Patient's name......:\", display_name)\n    print(\"Patient id..........:\", dataset.PatientID)\n    print(\"Patient's Age.......:\", dataset.PatientAge)\n    print(\"Patient's Sex.......:\", dataset.PatientSex)\n    print(\"Modality............:\", dataset.Modality)\n    print(\"Body Part Examined..:\", dataset.BodyPartExamined)\n   \n    \n    \n    if 'PixelData' in dataset:\n        rows = int(dataset.Rows)\n        cols = int(dataset.Columns)\n        print(\"Image size.......: {rows:d} x {cols:d}, {size:d} bytes\".format(\n            rows=rows, cols=cols, size=len(dataset.PixelData)))\n        if 'PixelSpacing' in dataset:\n            print(\"Pixel spacing....:\", dataset.PixelSpacing)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.grid(False)\n    plt.imshow(dataset.pixel_array)\n    plt.show()\n    \ni = 1\nnum_to_plot = 5\nfor file_name in os.listdir('../input/siim-isic-melanoma-classification/train/'):\n        file_path = os.path.join('../input/siim-isic-melanoma-classification/train/',file_name)\n        dataset = pydicom.dcmread(file_path)\n        show_dcm_info(dataset)\n        plot_pixel_array(dataset)\n    \n        if i >= num_to_plot:\n            break\n    \n        i += 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## データフレーム内のDIOCOMファイルの情報を抽出する\n\n[Gabriel Preda](https://www.kaggle.com/gpreda) さんが [discussion forum](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154658) で以下のコードを共有してくれました。","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# source: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154658\nfolder='train'\nPATH='../input/siim-isic-melanoma-classification/'\n\ndef extract_DICOM_attributes(folder):\n    images = list(os.listdir(os.path.join(PATH, folder)))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(PATH,folder,image)\n        dicom_file_dataset = pydicom.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n\n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"extract_DICOM_attributes('train')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}