{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"## Documents are exposed to harsh conditions once scribed and enter the world. Most of human history is captured in paper form, some going back thousands of years -- yet the knowledge they hold often degrades very quickly. Even in the modern world, as those documents are scanned, faxed, copied, and printed, there are many distortions that get introduced that obscure their original form. Those distortions, or \"noise\", prevent them from being captured in a more immortal, digital form. That is the problem we intend to solve, with your help.\n\nDepart with us on the hero's journey, to chart the uncharted path that restores luster to that which was once lost.","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport glob\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport glob","metadata":{"execution":{"iopub.status.busy":"2022-07-10T15:05:15.093945Z","iopub.execute_input":"2022-07-10T15:05:15.094434Z","iopub.status.idle":"2022-07-10T15:05:15.464415Z","shell.execute_reply.started":"2022-07-10T15:05:15.094348Z","shell.execute_reply":"2022-07-10T15:05:15.463669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualize Data","metadata":{}},{"cell_type":"code","source":"def plot_image_examples(clean, shabby, rows=3, cols=2, title='Cleaned vs Shabby'):\n    fig, axs = plt.subplots(rows, cols, figsize=(30,30))\n    for col in range(cols):\n        idx = np.random.randint(len(clean), size=1)[0]\n\n        clean_img = cv2.imread(clean[idx], 0)\n        shabby_img = cv2.imread(shabby[idx], 0)\n\n        axs[0, col].imshow(shabby_img)\n        axs[0, col].axis('off')\n        \n        axs[1, col].imshow(clean_img)\n        axs[1, col].axis('off')\n        \n        axs[2, col].imshow(cv2.fastNlMeansDenoising(shabby_img, None, 10, 7, 21) )\n        axs[2, col].axis('off')\n            \n    plt.suptitle(title)\n    \nBASE_DIR = \"/kaggle/input/denoising-shabby-pages\"\ncleaned_img_paths = sorted(glob.glob(BASE_DIR+\"/train_cleaned/train_cleaned/*.png\"))\ntrain_shabby_paths = sorted(glob.glob(BASE_DIR+\"/train/train/*.png\"))\nplot_image_examples(cleaned_img_paths,train_shabby_paths)\n\nTEST_DIR = BASE_DIR+\"/test/test/\"\nOUTPUT_DIR = \"/kaggle/working/results\"\n\ntry:\n    os.mkdir(OUTPUT_DIR)\nexcept Exception:\n    pass\nfinal_img_paths = sorted(glob.glob(TEST_DIR+\"*.png\"))\nlen(final_img_paths)","metadata":{"execution":{"iopub.status.busy":"2022-07-03T17:21:05.678019Z","iopub.execute_input":"2022-07-03T17:21:05.678515Z","iopub.status.idle":"2022-07-03T17:21:07.419907Z","shell.execute_reply.started":"2022-07-03T17:21:05.678479Z","shell.execute_reply":"2022-07-03T17:21:07.417458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above shows a very diverse set of examples. But few things wrong are that there are some shabby images and the clean version of those images are colored. It is challenging for a deep learning model to make a colored image from a distorted image (Unless we have a lot of data ofcourse)","metadata":{}},{"cell_type":"code","source":"# Make_submission.py:\nimport random\nimport os\nimport cv2\n\n################################################################################\n# CHANGE THIS\n################################################################################\ncleaned_images = final_img_paths\n\n################################################################################\n# DON'T CHANGE ANYTHING BELOW HERE\n################################################################################\ndef select_pixels(img):\n    y,x = img.shape\n\n    pixels = list()\n\n    for i in range(10000):\n        pixel = (random.randrange(y), random.randrange(x))\n\n        if pixel not in pixels:\n            pixels.append(pixel)\n\n    return pixels\n\nrandom.seed(0)\n#cleaned_images = sorted(os.listdir(cleaned_images_dir))\n\nwith open(\"submission.csv\", \"w\") as submission_file:\n    submission_file.write(\"id,predicted\\n\")\n\n    filenum = 1\n    for image in cleaned_images:\n\n        img = cv2.imread(image)\n        img = cv2.fastNlMeansDenoisingColored(img,None)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n        pixels = select_pixels(img)\n\n        for pixel in pixels:\n            y,x = pixel\n            submission_file.write(\"{}_{}_{},{}\\n\".format(filenum, y, x, img[y][x]/255.0))\n\n        filenum += 1","metadata":{"execution":{"iopub.status.busy":"2022-07-03T17:23:44.361418Z","iopub.execute_input":"2022-07-03T17:23:44.362214Z","iopub.status.idle":"2022-07-03T17:31:47.831725Z","shell.execute_reply.started":"2022-07-03T17:23:44.362178Z","shell.execute_reply":"2022-07-03T17:31:47.8306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}