{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🥱😬 TL;DR\n\nTwo major findings:\n\n1. `Atypical Appearance`, `Indeterminate Appearance`, `Negative for Pneumonia` are highly unbalanced classes.\n\n\n2. Pixel Spacing varies a lot between images and train & test data. It should be fixed in preprocessing for better and stable predictions.\n\n\nEverthing else looks normal!","metadata":{"execution":{"iopub.status.busy":"2021-05-26T18:39:59.825870Z","iopub.execute_input":"2021-05-26T18:39:59.826362Z","iopub.status.idle":"2021-05-26T18:39:59.835520Z","shell.execute_reply.started":"2021-05-26T18:39:59.826271Z","shell.execute_reply":"2021-05-26T18:39:59.833356Z"}}},{"cell_type":"markdown","source":"# SIIM-FISABIO-RSNA COVID-19 Detection EDA\n\n👉 [Problem Type] Object Detection and Multiclass Classification Problem","metadata":{}},{"cell_type":"code","source":"!pip install pandas-profiling[notebook] --quiet\n!pip install pydicom --quiet","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:35.555930Z","start_time":"2021-05-26T18:26:35.553673Z"},"scrolled":true,"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:16.908049Z","iopub.execute_input":"2021-05-26T19:27:16.909140Z","iopub.status.idle":"2021-05-26T19:27:31.733850Z","shell.execute_reply.started":"2021-05-26T19:27:16.909066Z","shell.execute_reply":"2021-05-26T19:27:31.732340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.io as pio\nimport pydicom as dicom\nfrom pathlib import Path\nfrom pandas_profiling import ProfileReport\nfrom PIL import Image\nfrom fastprogress import progress_bar\n\nsns.set_style(\"whitegrid\", {'axes.grid' : False})\n%config InlineBackend.figure_format = 'retina'\n\npio.templates.default = \"ggplot2\"","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:38.300377Z","start_time":"2021-05-26T18:26:35.558454Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:31.736757Z","iopub.execute_input":"2021-05-26T19:27:31.737148Z","iopub.status.idle":"2021-05-26T19:27:31.871064Z","shell.execute_reply.started":"2021-05-26T19:27:31.737110Z","shell.execute_reply":"2021-05-26T19:27:31.870096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATADIR = Path(\"../input/siim-covid19-detection\")\nTRAINDIR = DATADIR/\"train\"\nTESTDIR = DATADIR/\"test\"\ntrain_study_filepath = DATADIR/\"train_study_level.csv\"\ntrain_image_filepath = DATADIR/\"train_image_level.csv\"","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:38.436566Z","start_time":"2021-05-26T18:26:38.433806Z"},"execution":{"iopub.status.busy":"2021-05-26T19:27:31.872951Z","iopub.execute_input":"2021-05-26T19:27:31.873449Z","iopub.status.idle":"2021-05-26T19:27:31.878728Z","shell.execute_reply.started":"2021-05-26T19:27:31.873400Z","shell.execute_reply":"2021-05-26T19:27:31.877440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df = pd.read_csv(train_study_filepath)\ntrain_image_df = pd.read_csv(train_image_filepath)","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:38.475885Z","start_time":"2021-05-26T18:26:38.438943Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:31.881254Z","iopub.execute_input":"2021-05-26T19:27:31.881719Z","iopub.status.idle":"2021-05-26T19:27:31.967921Z","shell.execute_reply.started":"2021-05-26T19:27:31.881653Z","shell.execute_reply":"2021-05-26T19:27:31.966722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🔥 Pandas Profiler \nThe pandas `df.describe()` function is great but a little basic for serious exploratory data analysis. *pandas_profiling* extends the pandas DataFrame with `df.profile_report()` for quick data analysis.\n\nThis saves time to write basic EDA code.\n\nThere are two ways to use it:\n1. As a method of dataframe - `df.profile_report()`\n2. `pandas_profiling.ProfileReport()` method. (see below)","metadata":{"tags":[]}},{"cell_type":"code","source":"profile = ProfileReport(train_study_df, title=\"Study Level\")\nprofile","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:42.475025Z","start_time":"2021-05-26T18:26:38.477188Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:31.969694Z","iopub.execute_input":"2021-05-26T19:27:31.970135Z","iopub.status.idle":"2021-05-26T19:27:39.589235Z","shell.execute_reply.started":"2021-05-26T19:27:31.970074Z","shell.execute_reply":"2021-05-26T19:27:39.588111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🤔 Takeaway from Study Level profiler:\n1. Very unbalanced class distribution. \n2. *Typical Appearance* is the most balanced. ~ 52.8% (class 0) and 47.2% (class 1)\n1. *Atypical Appearance* has really low class 1. ~ 92.2% (class 0) and 7.8% (class 1)","metadata":{}},{"cell_type":"code","source":"train_image_df.profile_report(title=\"Image Level\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:44.836008Z","start_time":"2021-05-26T18:26:42.476267Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:39.591047Z","iopub.execute_input":"2021-05-26T19:27:39.591402Z","iopub.status.idle":"2021-05-26T19:27:44.115364Z","shell.execute_reply.started":"2021-05-26T19:27:39.591355Z","shell.execute_reply":"2021-05-26T19:27:44.113893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🤔 Takeaway from Image Level profiler:\n1. boxes has 2040 (32.2%) missing values \n2. label has 2040 (32.2%) `\"none 1 0 0 1 1\"` values.","metadata":{}},{"cell_type":"markdown","source":"### Using plots to understand data\n\nTo make metadata more simpler, let's merge two dataframes (image_level and study_level) into one. This can be done using outer join on`StudyInstanceUID` column from *image level* dataframe and `id` column from *study level* dataframe (just need to remove ..._study from samples).","metadata":{"tags":[]}},{"cell_type":"code","source":"# Remove _study part from id column and save ids in new column named StudyInstanceUID to merge dataframes\ntrain_study_df[\"StudyInstanceUID\"] = train_study_df.id.apply(lambda x: x.split(\"_\")[0])\ntraindf = pd.merge(train_study_df, train_image_df, on=\"StudyInstanceUID\")\ntraindf.rename({\"id_x\": \"id_study\", \"id_y\": \"id_image\"}, axis=1, inplace=True)\ntraindf.head()\n","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:44.857398Z","start_time":"2021-05-26T18:26:44.837205Z"},"execution":{"iopub.status.busy":"2021-05-26T19:27:44.116908Z","iopub.execute_input":"2021-05-26T19:27:44.117222Z","iopub.status.idle":"2021-05-26T19:27:44.151953Z","shell.execute_reply.started":"2021-05-26T19:27:44.117189Z","shell.execute_reply":"2021-05-26T19:27:44.150500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(histfunc=\"count\", x=traindf[\"Typical Appearance\"], name=\"Typical Appearance\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=traindf[\"Atypical Appearance\"], name=\"Atypical Appearance\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=traindf[\"Indeterminate Appearance\"], name=\"Indeterminate Appearance\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=traindf[\"Negative for Pneumonia\"], name=\"Negative for Pneumonia\"))\nfig.update_layout(title_text='Sample count per class') # title of plot\nfig.show()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:45.024611Z","start_time":"2021-05-26T18:26:44.859124Z"},"execution":{"iopub.status.busy":"2021-05-26T19:27:44.155904Z","iopub.execute_input":"2021-05-26T19:27:44.156246Z","iopub.status.idle":"2021-05-26T19:27:44.182833Z","shell.execute_reply.started":"2021-05-26T19:27:44.156214Z","shell.execute_reply":"2021-05-26T19:27:44.181404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DICOM Metadata Analysis\n\nIn this, we will analyse the DCM's metadata. Let's see if we can get any useful information.","metadata":{}},{"cell_type":"code","source":"dcmpaths = TRAINDIR.rglob(\"*.dcm\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:45.028625Z","start_time":"2021-05-26T18:26:45.026537Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:44.186266Z","iopub.execute_input":"2021-05-26T19:27:44.186957Z","iopub.status.idle":"2021-05-26T19:27:44.192572Z","shell.execute_reply.started":"2021-05-26T19:27:44.186911Z","shell.execute_reply":"2021-05-26T19:27:44.191073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_path = next(dcmpaths)\nsample = dicom.dcmread(sample_path)\nprint(\"Image size =\", (sample.Rows, sample.Columns))\nsample","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:45.087025Z","start_time":"2021-05-26T18:26:45.029896Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:44.194290Z","iopub.execute_input":"2021-05-26T19:27:44.194667Z","iopub.status.idle":"2021-05-26T19:27:44.261404Z","shell.execute_reply.started":"2021-05-26T19:27:44.194629Z","shell.execute_reply":"2021-05-26T19:27:44.260068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nplt.imshow(sample.pixel_array)\nplt.axis(\"off\");\nplt.title(\"Sample image\", {'fontsize':20});","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:26:46.632622Z","start_time":"2021-05-26T18:26:45.088212Z"},"tags":[],"execution":{"iopub.status.busy":"2021-05-26T19:27:44.263374Z","iopub.execute_input":"2021-05-26T19:27:44.263858Z","iopub.status.idle":"2021-05-26T19:27:46.018484Z","shell.execute_reply.started":"2021-05-26T19:27:44.263810Z","shell.execute_reply":"2021-05-26T19:27:46.016871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create a dataframe out of metadata for all samples.","metadata":{}},{"cell_type":"code","source":"# TRAINING DATA\ndcmpaths = TRAINDIR.rglob(\"*.dcm\")\nmetadata_traindf = {\n    \"Gender\": [], \"BodyPartExamined\": [], \"ImgHeight\": [], \n    \"ImgWidth\": [], \"ImagerPixelSpacing\": [], \"SOPInstanceUID\": []\n}\nfor path in progress_bar(list(dcmpaths)):\n    sample = dicom.dcmread(path)\n    metadata_traindf[\"Gender\"].append(sample.PatientSex)\n    metadata_traindf[\"BodyPartExamined\"].append(sample.BodyPartExamined)\n    metadata_traindf[\"ImgHeight\"].append(sample.Rows)\n    metadata_traindf[\"ImgWidth\"].append(sample.Columns)\n    metadata_traindf[\"ImagerPixelSpacing\"].append(float(sample.ImagerPixelSpacing[0]))\n    metadata_traindf[\"SOPInstanceUID\"].append(sample.SOPInstanceUID)\n    \nmetadata_traindf = pd.DataFrame.from_dict(metadata_traindf)","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:30:15.195242Z","start_time":"2021-05-26T18:26:46.633814Z"},"execution":{"iopub.status.busy":"2021-05-26T19:27:46.020779Z","iopub.execute_input":"2021-05-26T19:27:46.021228Z","iopub.status.idle":"2021-05-26T19:46:04.502622Z","shell.execute_reply.started":"2021-05-26T19:27:46.021173Z","shell.execute_reply":"2021-05-26T19:46:04.501181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_traindf.profile_report(title=\"DICOM Training Metadata\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:30:19.447528Z","start_time":"2021-05-26T18:30:15.196529Z"},"execution":{"iopub.status.busy":"2021-05-26T19:46:04.504703Z","iopub.execute_input":"2021-05-26T19:46:04.505017Z","iopub.status.idle":"2021-05-26T19:46:14.259226Z","shell.execute_reply.started":"2021-05-26T19:46:04.504987Z","shell.execute_reply":"2021-05-26T19:46:14.258049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TESTING DATA\ndcmpaths = TESTDIR.rglob(\"*.dcm\")\nmetadata_testdf = {\n    \"Gender\": [], \"BodyPartExamined\": [], \"ImgHeight\": [], \n    \"ImgWidth\": [], \"ImagerPixelSpacing\": [], \"SOPInstanceUID\": []\n}\nfor path in progress_bar(list(dcmpaths)):\n    sample = dicom.dcmread(path)\n    metadata_testdf[\"Gender\"].append(sample.PatientSex)\n    metadata_testdf[\"BodyPartExamined\"].append(sample.BodyPartExamined)\n    metadata_testdf[\"ImgHeight\"].append(sample.Rows)\n    metadata_testdf[\"ImgWidth\"].append(sample.Columns)\n    metadata_testdf[\"ImagerPixelSpacing\"].append(float(sample.ImagerPixelSpacing[0]))\n    metadata_testdf[\"SOPInstanceUID\"].append(sample.SOPInstanceUID)\n    \nmetadata_testdf = pd.DataFrame.from_dict(metadata_testdf)","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:30:59.372897Z","start_time":"2021-05-26T18:30:19.448729Z"},"execution":{"iopub.status.busy":"2021-05-26T19:46:14.260908Z","iopub.execute_input":"2021-05-26T19:46:14.261259Z","iopub.status.idle":"2021-05-26T19:49:43.396707Z","shell.execute_reply.started":"2021-05-26T19:46:14.261224Z","shell.execute_reply":"2021-05-26T19:49:43.395058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata_testdf.profile_report(title=\"DICOM Testing Metadata\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.078071Z","start_time":"2021-05-26T18:30:59.374397Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:43.398781Z","iopub.execute_input":"2021-05-26T19:49:43.399294Z","iopub.status.idle":"2021-05-26T19:49:50.183484Z","shell.execute_reply.started":"2021-05-26T19:49:43.399240Z","shell.execute_reply":"2021-05-26T19:49:50.182442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🤔 Takeaway from Metadata profilers:\n1. Both training and test data have almost similiar gender distribution\n2. Both training and test set have ~80% CHEST bodypart scans.\n3. Pixel spacing is not consistent between images as well as training and testing data. We need to fix this in preprocessing because the results might vary.\n\nLet's compare with plots","metadata":{}},{"cell_type":"markdown","source":"### Gender\n\nAs said earlier, distribution across train and test is similar as shown below.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_traindf[\"Gender\"], name=\"Train data Gender\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_testdf[\"Gender\"], name=\"Test dataGender\"))\nfig.update_layout(title_text='Sample count per class') # title of plot\nfig.show()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.129574Z","start_time":"2021-05-26T18:31:05.079156Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.184943Z","iopub.execute_input":"2021-05-26T19:49:50.185258Z","iopub.status.idle":"2021-05-26T19:49:50.239131Z","shell.execute_reply.started":"2021-05-26T19:49:50.185226Z","shell.execute_reply":"2021-05-26T19:49:50.238311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pixel Spacing\n`ImagerPixelSpacing` has similar distribution across train and test data but the pixel spacing is very different (given below). For better predictions, we need to fix the pixel spacing in preprocessing.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_traindf[\"ImagerPixelSpacing\"], histnorm=\"probability\", name=\"Train ImagerPixelSpacing\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_testdf[\"ImagerPixelSpacing\"], histnorm=\"probability\", name=\"Test ImagerPixelSpacing\"))\nfig.update_layout(title_text='Sample count per class (Probability Normalized just to adjust sample scale)') # title of plot\nfig.show()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.157997Z","start_time":"2021-05-26T18:31:05.131266Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.240347Z","iopub.execute_input":"2021-05-26T19:49:50.240817Z","iopub.status.idle":"2021-05-26T19:49:50.264622Z","shell.execute_reply.started":"2021-05-26T19:49:50.240785Z","shell.execute_reply":"2021-05-26T19:49:50.263420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pixel Spacing and Gender\nAnother analysis to understand what is the most common pixel spacing value.","metadata":{}},{"cell_type":"code","source":"px.histogram(metadata_traindf, x='ImagerPixelSpacing', marginal=\"box\", color='Gender', title=\"Train data ImagerPixelSpacing Distribution (based on Gender)\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.322761Z","start_time":"2021-05-26T18:31:05.159336Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.266105Z","iopub.execute_input":"2021-05-26T19:49:50.266401Z","iopub.status.idle":"2021-05-26T19:49:50.573621Z","shell.execute_reply.started":"2021-05-26T19:49:50.266373Z","shell.execute_reply":"2021-05-26T19:49:50.572848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.histogram(metadata_testdf, x='ImagerPixelSpacing', marginal=\"box\", color='Gender', title=\"Test data ImagerPixelSpacing Distribution (based on Gender)\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.411373Z","start_time":"2021-05-26T18:31:05.324489Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.574765Z","iopub.execute_input":"2021-05-26T19:49:50.575198Z","iopub.status.idle":"2021-05-26T19:49:50.670962Z","shell.execute_reply.started":"2021-05-26T19:49:50.575168Z","shell.execute_reply":"2021-05-26T19:49:50.670157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### BodyPart Examined and Gender\nLooks similar distribution Genderwise.","metadata":{}},{"cell_type":"code","source":"metadata_traindf.BodyPartExamined.unique()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.417309Z","start_time":"2021-05-26T18:31:05.412699Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.672135Z","iopub.execute_input":"2021-05-26T19:49:50.672541Z","iopub.status.idle":"2021-05-26T19:49:50.678766Z","shell.execute_reply.started":"2021-05-26T19:49:50.672511Z","shell.execute_reply":"2021-05-26T19:49:50.678019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the same BodyPart has different spellings in metadata. This is maybe due to different scanner. I don't know. If you know then please let me know in the comments.\n\nAnyway, I'm replacing different spellings to one.\nSo, there will be two changes:\n1. Replacing `2- TORAX`, `TORAX`, `TÒRAX`, `T?RAX` with `THORAX`.\n2. Replacing `Pecho` with `PECHO`","metadata":{}},{"cell_type":"code","source":"replace_dict = {\n    '2- TORAX': 'THORAX',\n    'TORAX': 'THORAX',\n    'TÒRAX': 'THORAX',\n    'T?RAX': 'THORAX',\n    'Pecho': 'PECHO'\n}\nmetadata_traindf.BodyPartExamined = metadata_traindf.BodyPartExamined.replace(replace_dict)\nmetadata_testdf.BodyPartExamined = metadata_testdf.BodyPartExamined.replace(replace_dict)","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.430207Z","start_time":"2021-05-26T18:31:05.419337Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.679990Z","iopub.execute_input":"2021-05-26T19:49:50.680401Z","iopub.status.idle":"2021-05-26T19:49:50.699832Z","shell.execute_reply.started":"2021-05-26T19:49:50.680370Z","shell.execute_reply":"2021-05-26T19:49:50.698506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1. Train and Test set distribution comparison","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_traindf[\"BodyPartExamined\"], histnorm=\"probability\", name=\"Train BodyPartExamined\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_testdf[\"BodyPartExamined\"], histnorm=\"probability\", name=\"Test BodyPartExamined\"))\nfig.update_layout(title_text='Sample count per Bodypart (Probability Normalized just to adjust sample scale)') # title of plot\nfig.show()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.476435Z","start_time":"2021-05-26T18:31:05.432241Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.701353Z","iopub.execute_input":"2021-05-26T19:49:50.701688Z","iopub.status.idle":"2021-05-26T19:49:50.759404Z","shell.execute_reply.started":"2021-05-26T19:49:50.701638Z","shell.execute_reply":"2021-05-26T19:49:50.757665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Same thing on Log scale for better visualization. Looks normal to me.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_traindf[\"BodyPartExamined\"], histnorm=\"probability\", name=\"Train BodyPartExamined\"))\nfig.add_trace(go.Histogram(histfunc=\"count\", x=metadata_testdf[\"BodyPartExamined\"], histnorm=\"probability\", name=\"Test BodyPartExamined\"))\nfig.update_layout(title_text='(Log) Sample count per Bodypart (Probability Normalized just to adjust sample scale)') # title of plot\nfig.update_yaxes(type=\"log\")\nfig.show()","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.510719Z","start_time":"2021-05-26T18:31:05.478815Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.764139Z","iopub.execute_input":"2021-05-26T19:49:50.764695Z","iopub.status.idle":"2021-05-26T19:49:50.830431Z","shell.execute_reply.started":"2021-05-26T19:49:50.764620Z","shell.execute_reply":"2021-05-26T19:49:50.829469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 2. BodyPartExamined and Gender","metadata":{}},{"cell_type":"code","source":"px.histogram(metadata_traindf, x='BodyPartExamined', marginal=\"violin\", color='Gender', title=\"Train data BodyPartExamined Distribution (based on Gender)\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.637868Z","start_time":"2021-05-26T18:31:05.511956Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:50.831734Z","iopub.execute_input":"2021-05-26T19:49:50.832060Z","iopub.status.idle":"2021-05-26T19:49:51.036335Z","shell.execute_reply.started":"2021-05-26T19:49:50.832030Z","shell.execute_reply":"2021-05-26T19:49:51.034993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"px.histogram(metadata_testdf, x='BodyPartExamined', marginal=\"violin\", color='Gender', title=\"Test data BodyPartExamined Distribution (based on Gender)\")","metadata":{"ExecuteTime":{"end_time":"2021-05-26T18:31:05.723977Z","start_time":"2021-05-26T18:31:05.639391Z"},"execution":{"iopub.status.busy":"2021-05-26T19:49:51.038118Z","iopub.execute_input":"2021-05-26T19:49:51.038453Z","iopub.status.idle":"2021-05-26T19:49:51.150849Z","shell.execute_reply.started":"2021-05-26T19:49:51.038419Z","shell.execute_reply":"2021-05-26T19:49:51.149767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### That's all! Please upvote if you find this information useful ✌️🙂","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}