{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\n\nThis kernel is just a quick look at the training dataset image sizes, and a look at some of the images at the lowest and highest diagnosis levels. To see if there is something easily visible to understand what the doctor might be looking at in a classification.\n\nThere is also a [previous competition](https://www.kaggle.com/c/diabetic-retinopathy-detection) on the same topic, with the exact same training labels. It seems to have a much larger training dataset. This set was mentioned multiple times in the [external data thread](). I had trouble adding that competition as a data source (error about loading the data). So I downloaded the data and set it up as a [separate dataset](https://www.kaggle.com/donkeys/retinopathy-train-2015). Had to downscale it quite a bit to max 896x896 pixel sizes, to fit it into the 20GB dataset size limit. But it seems potentially useful.\n\nI am not quite sure how to check exact date of some old competition here on Kaggle, so I just picket the number of years it displays in the past, and went with 2015. So I will call the older set the *2015* set here. Or the *past* set vs the actual current set for the *present* time.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport cv2\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport numpy as np\nimport json\nimport math\nimport PIL\nfrom PIL import ImageOps\nfrom keras.models import Sequential, Model\nfrom keras.layers import Dense, Flatten, Activation, Dropout, GlobalAveragePooling2D\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras import optimizers, applications\nfrom keras.callbacks import ModelCheckpoint, LearningRateScheduler, TensorBoard, EarlyStopping\nfrom keras import backend as K \nfrom keras.utils.np_utils import to_categorical\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nimport keras\n\nfrom tqdm.auto import tqdm\ntqdm.pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -l ../input/","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Number of files in train vs test vs the 2015 training set"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"!ls -l ../input/aptos2019-blindness-detection/train_images | wc -l","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -l ../input/aptos2019-blindness-detection/test_images | wc -l","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -l ../input/retinopathy-train-2015/rescaled_train_896/rescaled_train_896 | wc -l","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Basic metadata"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_path_2015 = \"../input/retinopathy-train-2015/rescaled_train_896/rescaled_train_896/\"\ntrain_path = \"../input/aptos2019-blindness-detection/train_images/\"\ntest_path = \"../input/aptos2019-blindness-detection/test_imges/\"\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(\"../input/aptos2019-blindness-detection/train.csv\")\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_test = pd.read_csv(\"../input/aptos2019-blindness-detection/test.csv\")\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_2015 = pd.read_csv(\"../input/retinopathy-train-2015/rescaled_train_896/trainLabels.csv\")\ndf_train_2015.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First 10 un-ordered files in past and present training sets to see the filenames match the csv columns (\"id_code\" and \"image\"):"},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -lU ../input/aptos2019-blindness-detection/train_images/ | head -10","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls -lU ../input/retinopathy-train-2015/rescaled_train_896/rescaled_train_896 | head -10","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's a match."},{"metadata":{},"cell_type":"markdown","source":"## Collect all metadata to single dataframe(s)"},{"metadata":{"trusted":true},"cell_type":"code","source":"n_rows = df_train.shape[0]\nn_rows","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[\"filename\"] = df_train[\"id_code\"]+\".png\"\ndf_train[\"path\"] = [train_path]*n_rows\n#the year is just to be able to easily separate the past and present datasets later\ndf_train[\"year\"] = [2019]*n_rows\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n_rows_2015 = df_train_2015.shape[0]\nn_rows_2015","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_2015[\"filename\"] = df_train_2015[\"image\"]+\".png\"\ndf_train_2015[\"path\"] = [train_path_2015]*n_rows_2015\ndf_train_2015[\"year\"] = [2015]*n_rows_2015\ndf_train_2015.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_2015.columns = [\"id_code\", \"diagnosis\", \"filename\", \"path\", \"year\"]\ndf_train_2015.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_all = pd.concat([df_train,df_train_2015], axis=0, sort=False).reset_index()\ndf_train_all.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_all.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#replacing df_train with the full set to calculate features and do visualizations all at once, keeping the original (present) just in case\ndf_train_orig = df_train\ndf_train = df_train_all","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Calculate Aspect Ratios etc."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nimg_sizes = []\nwidths = []\nheights = []\naspect_ratios = []\n\nfor index, row in tqdm(df_train.iterrows(), total=df_train.shape[0]):\n    filename = row[\"filename\"]\n    path = row[\"path\"]\n    img_path = os.path.join(path, filename)\n    with open(img_path, 'rb') as f:\n        img = PIL.Image.open(f)\n        img_size = img.size\n        img_sizes.append(img_size)\n        widths.append(img_size[0])\n        heights.append(img_size[1])\n        aspect_ratios.append(img_size[0]/img_size[1])\n\ndf_train[\"width\"] = widths\ndf_train[\"height\"] = heights\ndf_train[\"aspect_ratio\"] = aspect_ratios\ndf_train[\"size\"] = img_sizes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Aspect Ratios\n\nSee that there are no images that are hugely different in size to others:"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted = df_train.sort_values(by=\"aspect_ratio\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted[df_sorted[\"year\"] == 2015].head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted[df_sorted[\"year\"] == 2019].head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The aspect ratios in the past and present seem very close to each other."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted[df_sorted[\"year\"] == 2015].tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted[df_sorted[\"year\"] == 2019].tail()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Look at the Images / Eyes"},{"metadata":{"trusted":true},"cell_type":"code","source":"#This just shows a single image in the notebook\ndef show_img(filename, path):\n        img = PIL.Image.open(f\"{path}/{filename}\")\n        npa = np.array(img)\n        print(npa.shape)\n        #https://stackoverflow.com/questions/35902302/discarding-alpha-channel-from-images-stored-as-numpy-arrays\n#        npa3 = npa[ :, :, :3]\n        print(filename)\n        plt.imshow(npa)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib\n\nfont = {'family' : 'normal',\n        'weight' : 'normal',\n        'size'   : 22}\n\nmatplotlib.rc('font', **font)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## A Random Eye\n\nVisualize the first image in past and present sets to see if they are at all alike:\n"},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"row = df_sorted[df_sorted[\"year\"] == 2019].iloc[0]\nshow_img(row.filename, row.path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"row = df_sorted[df_sorted[\"year\"] == 2015].iloc[0]\nshow_img(row.filename, row.path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 9-Eyes\n\nVisualize 9 images from a set at a time, to learn a bit more about the set at once."},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_first_9(df_to_plot):\n    plt.figure(figsize=[30,30])\n    for x in range(9):\n        path = df_to_plot.iloc[x].path\n        filename = df_to_plot.iloc[x].filename\n        img = PIL.Image.open(f\"{path}/{filename}\")\n        print(filename)\n        plt.subplot(3, 3, x+1)\n        plt.imshow(img)\n        title_str = filename+\", diagnosis: \"+str(df_to_plot.iloc[x].diagnosis)\n        plt.title(title_str)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Smallest Aspect Ratio\n\nThere seem to be no images with aspect ratio < 1, so plotting the smallest aspect ratios (practically the ratio is then 1) should show the most \"square\" images:"},{"metadata":{"trusted":true},"cell_type":"code","source":"del df_sorted\ndf_sorted = df_train.sort_values(by=\"aspect_ratio\", ascending=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2019])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2015])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Generally, the past vs present images seem very similar. Some color differences, although some of the later pics will show both have these more \"orange\" and \"greenish\" ones as well. But a deeper investigation of how the color spaces are distributed in different sets could be interesting."},{"metadata":{},"cell_type":"markdown","source":"## Highest Aspect Ratios\n\nThis should be the ones least \"square\":"},{"metadata":{"trusted":true},"cell_type":"code","source":"del df_sorted\ndf_sorted = df_train.sort_values(by=\"aspect_ratio\", ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2019])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2015])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Diagnosis Values\n\nA look at the highest vs lowest diagnosis values /levels given in the training set. Can we spot some differences? \n\n### Highest / Most Severe Diagnosis:"},{"metadata":{"trusted":true},"cell_type":"code","source":"del df_sorted\ndf_sorted = df_train.sort_values(by=\"diagnosis\", ascending=False)\ndf_sorted.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2019])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2015])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Lowest / Healthiest Diagnosis:"},{"metadata":{"trusted":true},"cell_type":"code","source":"del df_sorted\ndf_sorted = df_train.sort_values(by=\"diagnosis\", ascending=True)\ndf_sorted.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Present"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2019])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Past"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_first_9(df_sorted[df_sorted[\"year\"] == 2015])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I guess the healthier ones look more \"clean\"."},{"metadata":{},"cell_type":"markdown","source":"# Final Size Statistics\n\nOn average, are the files about the same size? Actually might make sense to look at the past and present sets separately since I had to downsize the past significantly. But the idea is there, and it does already show if there are some really small ones.."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Are the smallest files still valid files?"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sorted = df_train.sort_values(by=\"width\", ascending=True)\n\nplot_first_9(df_sorted[df_sorted[\"year\"] == 2019])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nplot_first_9(df_sorted[df_sorted[\"year\"] == 2015])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Conclusions\n\nThe images from both sets seem to be quite similar. Possibly some color differences and other minor differences?"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}