{"cells":[{"metadata":{},"cell_type":"markdown","source":"\n## Purpose of this kernel\n\nHi, I've been interested in image processing world for a long time, and finally joined APTOS competition.\n\nSame as other novice competitors, I begun to read discussions and kernels witten by other competitors, and I especially enjoyed Michael Kazachok's nice kernel (https://www.kaggle.com/miklgr500/auto-encoder), which demonstrate auto encoder technique and apply PCA on image data, because his kernel gave me great insight about how dataset are distributed.\n\n\nIn adittion, I learned elementary techniques to preprocess image data from Kazachok's kernel, so write this kernel which introduce how to read and preprocess images for newcomers like me.\n\nAny comments for clarification and correction are welcome."},{"metadata":{},"cell_type":"markdown","source":"## Import packages"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"import cv2\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Read image\n\nUse cv2(OpenCV) package's imread() to read images."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\ntest_df = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_id = (train_df['id_code'])[1]\npath = f\"../input/train_images/{image_id}.png\"\nimg = cv2.imread(path)\nplt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB)) ## note that cv.imread() gives BGR array, so need convert it to RGB array","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Convert image into gray-scale\n\nConvert a BGR image into a gray-scale image by cv2.cvtColor().\nBecause better performances are achieved with more information in general, I have no confidence that I should always discard information of colors. However, at least there are several reasons which justify to discard them:\n\n1. As mentioned by the organizer, the images were taken with various kinds of cameras, so color information could work as noise. (\"The images were gathered from multiple clinics using a variety of cameras over an extended period of time, which will introduce further variation.\")\n\n2. Images in this dataset are not so colorful, and it's likely that imformation of colors are not so extensively useful for this competitions. (Though I have no perfect confidence about it, because I'm not a doctor!)\n\n3. Discard colors can reduce data size and could save computing time."},{"metadata":{"trusted":true},"cell_type":"code","source":"img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\nplt.gray()\nplt.imshow(img)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Clip and resize\n\nCut off outside of retina image, and resize it into handy size."},{"metadata":{"trusted":true},"cell_type":"code","source":"tol = 5\nmask = img > tol\n\nnz_rows = mask.any(1)  ## filter rows where all values are less than tol out\nnz_cols = mask.any(0)  ## filter cols where all values are less than tol out\n\nimg = img[np.ix_(nz_rows, nz_cols)]\nplt.imshow(img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img = cv2.resize(img, (224, 224))\nplt.imshow(img)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Sharpen image\nSharpen the image by subtracting blured image, and add mean level.\n\nIn fact, I don't understand how magic numbers (used for cv2.GaussianBlur() and cv2.addWeighted()) are determined.\nAnyway, you can find the basic concept of this process (i.e., using Gaussian Blur as low-pass filter) in various articles (the URL below, for example).\n\nhttps://stackoverflow.com/questions/4993082/how-to-sharpen-an-image-in-opencv"},{"metadata":{"trusted":true},"cell_type":"code","source":"kernel_size = (0, 0)\nsigma_XY = 224/10\nimg2 = cv2.GaussianBlur(img, kernel_size, sigma_XY)\nplt.imshow(img2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img = cv2.addWeighted(img, 4, img2, -4, 128)  ## img = 4 * img - 4 * img2 + 128\nplt.imshow(img)  ## sharpened image","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}