{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Purpose\n\nThe goal of this notebook is to quickly analyze the dataset for some potential issues that might impact the performance of your model.\n\nFeel free to include or exclude these images in training your model and monitor the scores as you do so. ","metadata":{}},{"cell_type":"markdown","source":"## Installation\n\nWe will use a free tool - [fastdup](https://github.com/visual-layer/fastdup) to quickly analyze the dataset.","metadata":{}},{"cell_type":"code","source":"!pip install fastdup matplotlib -Uq","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:21:57.855769Z","iopub.execute_input":"2023-05-26T04:21:57.856679Z","iopub.status.idle":"2023-05-26T04:22:21.264740Z","shell.execute_reply.started":"2023-05-26T04:21:57.856632Z","shell.execute_reply":"2023-05-26T04:22:21.262603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Run fastdup\n\n`work_dir` - Directory to store the results.\n\n`images_dir` - Directory to your images.","metadata":{}},{"cell_type":"code","source":"import fastdup\n\nwork_dir = \"./fastdup_report\"\nimages_dir = \"/kaggle/input/image-matching-challenge-2023/train\"\n\nfd = fastdup.create(work_dir, images_dir)\nfd.run(ccthreshold=0.9, threshold=0.8, overwrite=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:22:21.269111Z","iopub.execute_input":"2023-05-26T04:22:21.270155Z","iopub.status.idle":"2023-05-26T04:24:45.725316Z","shell.execute_reply.started":"2023-05-26T04:22:21.270062Z","shell.execute_reply":"2023-05-26T04:24:45.723477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's it! Only few lines of code to analyze the dataset. Now let's visualize the results.","metadata":{}},{"cell_type":"markdown","source":"## Duplicate Images\nLet's visualize the duplicates with `duplicates_gallery`. The `num_images` argument specifies the number of duplicates to vizualize.\n\nDistance of `1.0` indicate exact duplicate image. The lower the Distance score, the more dissimilar the image pair is.\n","metadata":{}},{"cell_type":"code","source":"fd.vis.duplicates_gallery(num_images=50)","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:25:16.910296Z","iopub.execute_input":"2023-05-26T04:25:16.910854Z","iopub.status.idle":"2023-05-26T04:25:21.907417Z","shell.execute_reply.started":"2023-05-26T04:25:16.910805Z","shell.execute_reply":"2023-05-26T04:25:21.906494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Outliers\n\nWe can also view a gallery of outliers. Feel free to also specify the `num_images` argument.","metadata":{}},{"cell_type":"code","source":"fd.vis.outliers_gallery()","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:26:29.302052Z","iopub.execute_input":"2023-05-26T04:26:29.302532Z","iopub.status.idle":"2023-05-26T04:26:30.239174Z","shell.execute_reply.started":"2023-05-26T04:26:29.302479Z","shell.execute_reply":"2023-05-26T04:26:30.237157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dark/Bright/Blurry Images\n\nView a gallery of sorted darkest, brightest and blurriest images.\n\n","metadata":{}},{"cell_type":"code","source":"fd.vis.stats_gallery(metric='dark')","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:27:06.932423Z","iopub.execute_input":"2023-05-26T04:27:06.933026Z","iopub.status.idle":"2023-05-26T04:27:07.613957Z","shell.execute_reply.started":"2023-05-26T04:27:06.932975Z","shell.execute_reply":"2023-05-26T04:27:07.612271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fd.vis.stats_gallery(metric='bright')","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:27:31.352726Z","iopub.execute_input":"2023-05-26T04:27:31.353192Z","iopub.status.idle":"2023-05-26T04:27:32.058819Z","shell.execute_reply.started":"2023-05-26T04:27:31.353149Z","shell.execute_reply":"2023-05-26T04:27:32.048692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fd.vis.stats_gallery(metric='blur')","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:27:45.592309Z","iopub.execute_input":"2023-05-26T04:27:45.592801Z","iopub.status.idle":"2023-05-26T04:27:46.033960Z","shell.execute_reply.started":"2023-05-26T04:27:45.592760Z","shell.execute_reply":"2023-05-26T04:27:46.031990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Image Clusters\nWe can visualize similar images together as clusters using the `components_gallery`.\n\nIn the visualization below the keyword `component` refers to cluster. Eg. component 436 means cluster number 436 found in the dataset.\n\nClustering similar images together let's you group pictures of similar looking buildings/scenes together.","metadata":{}},{"cell_type":"code","source":"fd.vis.component_gallery()","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:28:04.270762Z","iopub.execute_input":"2023-05-26T04:28:04.271255Z","iopub.status.idle":"2023-05-26T04:28:43.610800Z","shell.execute_reply.started":"2023-05-26T04:28:04.271214Z","shell.execute_reply":"2023-05-26T04:28:43.609244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you'd like to view a detailed dataframe of what is being visualized in the galleries, run the `.similarity()` function.","metadata":{}},{"cell_type":"code","source":"fd.similarity()","metadata":{"execution":{"iopub.status.busy":"2023-05-26T04:33:16.228210Z","iopub.execute_input":"2023-05-26T04:33:16.229735Z","iopub.status.idle":"2023-05-26T04:33:16.311029Z","shell.execute_reply.started":"2023-05-26T04:33:16.229647Z","shell.execute_reply":"2023-05-26T04:33:16.308605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using the dataframe you can now vizualize the data in whatever format you wish. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}