{"cells":[{"metadata":{},"cell_type":"markdown","source":"### Summary\n\nDuplicates are always harmful for training process: differently labeled duplicates produce noise in the dataset, while equally labeled duplicates lead to data leakage. \n\nIn this short notebook I am looking through image hash of **Plant Pathology 2021** competition dataset with `image_hash` library, finding more than 50 duplicates.\n\n### Update\nDue to recent changes in the `train.csv` file mentioned in **[this discussion](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/228465)**, we have no more the `cider_apple_rust` class. This version (7) is made after the changes."},{"metadata":{},"cell_type":"markdown","source":"### Imports"},{"metadata":{"trusted":true,"_uuid":"4a78d0be9bc398c6f69dffdb9709c7e32361f47d","scrolled":true},"cell_type":"code","source":"from tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport imagehash\nimport PIL\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true},"cell_type":"code","source":"class CFG():\n    \n    threshold = .9\n    img_size = 512\n    seed = 42","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1. Saving downscaled images to boost performance\nComputing hash over original images of very high quality would take nearly 5 hours, thus we downscaling first."},{"metadata":{"trusted":true,"scrolled":true},"cell_type":"code","source":"root = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n\npaths = os.listdir(root)\n\ndf = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv', index_col='image')\n\nfor path in tqdm(paths, total=len(paths)):\n    image = tf.io.read_file(os.path.join(root, path))\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, [CFG.img_size, CFG.img_size])\n    image = tf.cast(image, tf.uint8).numpy()\n    plt.imsave(path, image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 2. Hash computation"},{"metadata":{"trusted":true,"_uuid":"3a94ec9c45f58e7bc62bfeee6c2cdf06d7d92d92","scrolled":true},"cell_type":"code","source":"hash_functions = [\n    imagehash.average_hash,\n    imagehash.phash,\n    imagehash.dhash,\n    imagehash.whash]\n\nimage_ids = []\nhashes = []\n\npaths = tf.io.gfile.glob('./*.jpg')\n\nfor path in tqdm(paths, total=len(paths)):\n\n    image = PIL.Image.open(path)\n\n    hashes.append(np.array([x(image).hash for x in hash_functions]).reshape(-1,))\n    image_ids.append(path.split('/')[-1])\n    \nhashes = np.array(hashes)\nimage_ids = np.array(image_ids)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 3. Run search across hashed images\nWe firstly compare each image hash with all the hashes and then leave only unique pairs of matches"},{"metadata":{"trusted":true,"scrolled":true},"cell_type":"code","source":"duplicate_ids = []\n\nfor i in tqdm(range(len(hashes)), total=len(hashes)):\n    similarity = (hashes[i] == hashes).mean(axis=1)\n    duplicate_ids.append(list(image_ids[similarity > CFG.threshold]))\n    \nduplicates = [frozenset([x] + y) for x, y in zip(image_ids, duplicate_ids)]\nduplicates = set([x for x in duplicates if len(x) > 1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we add some of the duplicates spotted by @kingofarmy in the corresponding **[discussion](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/229851)**:"},{"metadata":{"trusted":true},"cell_type":"code","source":"duplicates_by_kingofarmy = {\n    frozenset(('8dbeda49894d522e.jpg', 'afbe5641896d522a.jpg')),\n    frozenset(('af6292db1b611d98.jpg', 'a56292dadb618d95.jpg')),\n    frozenset(('abf0b5a0df028b17.jpg', 'abf0b5819f028f0f.jpg')),\n    frozenset(('e385830ecacd2d9e.jpg', 'c335971e8acd609e.jpg')),\n    frozenset(('cebdc20f67838631.jpg', 'dfbdc047068b063d.jpg')),\n    frozenset(('f392f11919991cea.jpg', 'f196f11a99d91ce0.jpg'))}\n\nduplicates |= duplicates_by_kingofarmy","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 4. Let's see what is found"},{"metadata":{"trusted":true,"scrolled":true},"cell_type":"code","source":"print(f'Found {len(duplicates)} duplicate pairs:')\nfor row in duplicates:\n    print(', '.join(row))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Writing duplicates to \"duplicates.csv\".')\nwith open('duplicates.csv', 'w') as file:\n    for row in duplicates:\n        file.write(','.join(row) + '\\n')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false},"cell_type":"code","source":"for row in duplicates:\n    \n    figure, axes = plt.subplots(1, len(row), figsize=[5 * len(row), 5])\n\n    for i, image_id in enumerate(row):\n        image = plt.imread(os.path.join('../input/plant-pathology-2021-fgvc8/train_images', image_id))\n        axes[i].imshow(image)\n\n        axes[i].set_title(f'{image_id} - {df.loc[image_id, \"labels\"]}')\n        axes[i].axis('off')\n\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Clear working folder to avoid output pollution"},{"metadata":{"trusted":true,"scrolled":true},"cell_type":"code","source":"for file in tf.io.gfile.glob('./*.jpg'):\n    os.remove(file)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Acknowledgements\n\n* This work is Copy&Edit form @appian **[notebook](https://www.kaggle.com/appian/let-s-find-out-duplicate-images-with-imagehash)** with a lot of changes, but still highly inspired. If you find this notebook useful, please, upvote his work too.\n* Thanks to @kingofarmy for spotting more duplicates in **[his thread](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/229851)**."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}