{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Objective \nObjective of this kernel is to quickly explore the dataset given and understand the dataset and have an idea about **Diabetic Retinopathy**.\n\nWhat is Diabetic Retinopathy?  \nDR is a damage to the **Retina** caused by complications of diabetes mellitus. This condition can lead to blindness if left untreated.\n\nWhat exactly is happening to the Retina?  \nDR is damage of blood vessels in the retina that happens due to diabetes.\n\nWhat are the symptoms?  \nCommon symptoms are blurred vision, color blindness, floaters and complete loss of vision.\n\nHow does a normal eye and an eye affected with DR looks?  \nThis image from American Optmetric Association shows the difference between a normal eye and a DR eye.\n![Normal Eye vs DR Eye](https://www.aoa.org/Images/public/Diabetic_Retinopathy.jpg)  \n\nWhat other complications occur along with DR?  \n- **Vitreous hemorrhage** - Leak in new blood vessel\n- **Detached Retina** - Scar tissue that pulls the retina away from back of the eye\n- **Galucoma** - Blockage of the flow of fluid in the eye when new blood vessels form\n\nIs this curable?  \nYes to some extent but significant **side-effects** are anticipated in **Proliferative DR** condition."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from IPython.display import YouTubeVideo\nYouTubeVideo('7cEd2ZrItNg')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Prepare Dataset\nLet us prepare the dataset for us to easily navigate across the images.\n- Read the train and test files\n- Create the Diagnostic Label dataframe and merge to train data\n- Enumerate all images and add to the train and test dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/train.csv\")\nprint(\"Shape of train data: {0}\".format(train_df.shape))\ntest_df = pd.read_csv(\"../input/test.csv\")\nprint(\"Shape of test data: {0}\".format(test_df.shape))\n\ndiagnosis_df = pd.DataFrame({\n    'diagnosis': [0, 1, 2, 3, 4],\n    'diagnosis_label': ['No DR', 'Mild', 'Moderate', 'Severe', 'Proliferative DR']\n})\n\ntrain_df = train_df.merge(diagnosis_df, how=\"left\", on=\"diagnosis\")\n\ntrain_image_files = [os.path.join(dp, f) for dp, dn, fn in os.walk(os.path.expanduser(\"../input/train_images\")) for f in fn]\ntrain_images_df = pd.DataFrame({\n    'files': train_image_files,\n    'id_code': [file.split('/')[3].split('.')[0] for file in train_image_files],\n})\ntrain_df = train_df.merge(train_images_df, how=\"left\", on=\"id_code\")\ndel train_images_df\nprint(\"Shape of train data: {0}\".format(train_df.shape))\n\ntest_image_files = [os.path.join(dp, f) for dp, dn, fn in os.walk(os.path.expanduser(\"../input/test_images\")) for f in fn]\ntest_images_df = pd.DataFrame({\n    'files': test_image_files,\n    'id_code': [file.split('/')[3].split('.')[0] for file in test_image_files],\n})\n\n\ntest_df = test_df.merge(test_images_df, how=\"left\", on=\"id_code\")\ndel test_images_df\nprint(\"Shape of test data: {0}\".format(test_df.shape))\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Number of unique diagnosis: {0}\".format(train_df.diagnosis.nunique()))\ndiagnosis_count = train_df.diagnosis.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom PIL import Image\n\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.tools as tls\n\npd.options.mode.chained_assignment = None\npd.options.display.max_columns = 9999\npd.options.display.float_format = '{:20, .2f}'.format","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Sample Distribution by Severity"},{"metadata":{"trusted":true},"cell_type":"code","source":"def render_bar_chart(data_df, column_name, title, filename):\n    series = data_df[column_name].value_counts()\n    count = series.shape[0]\n    \n    trace = go.Bar(x = series.index, y=series.values, marker=dict(\n        color=series.values,\n        showscale=True\n    ))\n    layout = go.Layout(title=title)\n    data = [trace]\n    fig = go.Figure(data=data, layout=layout)\n    py.iplot(fig, filename=filename)\n    \n    \nrender_bar_chart(train_df, 'diagnosis_label', 'Diabetic Retinopathy: Observation Distribution by Severity ', 'members')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## No DR\nExamine the dataset and identify whether we can figure out any significant differences across the various diagnosis of DR images.  \nLet us pick some images randomly and closely watch. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"SAMPLES_TO_EXAMINE = 5\nimport cv2\ndef render_images(files):\n    plt.figure(figsize=(50, 50))\n    row = 1\n    for an_image in files:\n        image = cv2.imread(an_image)[..., [2, 1, 0]]\n        plt.subplot(6, 5, row)\n        plt.imshow(image)\n        row += 1\n    plt.show()\n    \nno_dr_pics = train_df[train_df[\"diagnosis\"] == 0].sample(SAMPLES_TO_EXAMINE)\nrender_images(no_dr_pics.files.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Mild DR\nI couldn't distinguish anything between no DR and early stages of DR images with my naked eye."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"mild_dr_pics = train_df[train_df[\"diagnosis\"] == 1].sample(SAMPLES_TO_EXAMINE)\nrender_images(mild_dr_pics.files.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Moderate DR\nImages where DR was diagnosed as moderate one, I could clearly see some patches and bright spots in almost all images. However the blood vessel appearance remains almost same."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"moderate_dr_pics = train_df[train_df[\"diagnosis\"] == 2].sample(SAMPLES_TO_EXAMINE)\nrender_images(moderate_dr_pics.files.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Severe DR\nIn sever DR diagnostic conditions, blood vesses are quite significant and seen densely. Bright patches are much more evident."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"severe_dr_pics = train_df[train_df[\"diagnosis\"] == 3].sample(SAMPLES_TO_EXAMINE)\nrender_images(severe_dr_pics.files.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Proliferative DR\nIn Proliferative DR images, the images are significantly different from the other 4 classes. There is a haziness and dull. In some of the images, it looks blood is oozing out from the vessels... bit scary."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"preoliferative_dr_pics = train_df[train_df[\"diagnosis\"] == 4].sample(SAMPLES_TO_EXAMINE)\nrender_images(preoliferative_dr_pics.files.values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}