{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Inspired by [Applying Ben's Preprocessing](https://www.kaggle.com/banzaibanzer/applying-ben-s-preprocessing)\n\n## My Own Insight: Denoising after Ben's Preprocssing Leads To Better Visual Clarity\n\nApplying Ben's Preprocessing introduces noises(I assume gaussian?) to images at a certain level and thus denoising can help reduce such negative effect. **(Even some characters inked from back of the page becomes more recognizable).**\n\nHowever, the cons:\n1. The denoising takes WAY TOO LONG. On average, each image requires about 21-23 seconds to be denoised.\n\n2. The denoising algorithm in nature cause slight loss of information. However, in my opinion, the loss of information it cause is acceptable compared to its benefits.\n"},{"metadata":{},"cell_type":"markdown","source":"## Prerequisites"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport cv2 # image processing\n\nfrom matplotlib import pyplot as plt # data visualization\n\n# making sure result is reproducible\nSEED = 2019\nnp.random.seed(SEED)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Some Utility Functions"},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_image(image):\n    '''\n        Simply read a single image and convert it RGB in opencv given its filename.\n    '''\n\n    return cv2.cvtColor(cv2.imread(image), cv2.COLOR_BGR2RGB)\n\n\ndef apply_ben_preprocessing(image):\n    '''\n        Apply Ben's preprocessing on a single image in opencv format\n    '''\n    \n    return cv2.addWeighted(image, 4, cv2.GaussianBlur(image, (0,0), 10), -4, 128)\n\n\ndef apply_denoising(image):\n    '''\n        Apply denoising on a single image given it in opencv format.\n        Denoising is done using twice the recommended strength from opencv docs.\n    '''\n    \n    return cv2.fastNlMeansDenoisingColored(image, None, 20, 20, 7, 21)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Let's See the Differences"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\nsamples = train_df.sample(n=10)\n\nfor ID in samples['image_id']:\n    filename = '../input/train_images/{}.jpg'.format(ID)\n    \n    img = read_image(filename)\n    before = apply_ben_preprocessing(img)\n    after = apply_denoising(before)\n\n    fig, ax = plt.subplots(1, 2, figsize=(16, 20))\n    ax[0].imshow(before)\n    ax[1].imshow(after)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## After Thought\n\nIf this kernel helps you in anyway, please give an upvote :)\n\nFeel free to comment and discuss with me below on the subject and I will continue to explore more about this dataset and hopefully I can come up with a model that achieves a decent result. \n\n#### As always, Happy Kaggling! :)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}