{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport cv2\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-31T00:05:10.081786Z","iopub.execute_input":"2022-03-31T00:05:10.082136Z","iopub.status.idle":"2022-03-31T00:05:10.483231Z","shell.execute_reply.started":"2022-03-31T00:05:10.082049Z","shell.execute_reply":"2022-03-31T00:05:10.482187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**How clean is the data?**\n\nI ran the ImageHash library against every image in the training_images, trying to find if some pictures of individuals were mis-identified, even though they were taken seconds apart. I found two such instances of this happening.\n\nMy processes basically followed what @appian did here https://www.kaggle.com/code/appian/let-s-find-out-duplicate-images-with-imagehash/notebook. I had to run the hashing locally, as Kaggle kept dying on me. I then had to pay for CoLab+ to run the similarity calculations, as it took about 20GB of RAM to do, which I don't have.\n\nI have supplied the results of my runs as a saved numpy file in the \"whalesimularitydata\". 90.npy contains above 90% similarity, 91.py contains above 91% similarity and so on. Format of each is is a ndarray with each element being [(whale1image,whale1species,whale1ID),(whale2image,whale2species,whale2ID)]\n\nI found that only >0.95 images had a chance of being true positives of mis-identified whales. 0.94 and below were likely false positives.\n\n**How does this impact the competition?**\nIt doesn't as far as I can tell, as these mis-identified images don't show up in the testset. See below for more details. ","metadata":{}},{"cell_type":"code","source":"def show(row1, row2):\n    print('Image: %s / %s' % (row1[2], row2[2]))\n    print('Species: %s / %s' % (row1[1], row2[1]))\n    print('Individual: %s / %s' % (row1[0], row2[0]))\n    \n    image1 = cv2.imread('../input/happy-whale-and-dolphin/train_images/%s' % (row1[2]))\n    image2 = cv2.imread('../input/happy-whale-and-dolphin/train_images/%s' % (row2[2]))\n    image1 = cv2.cvtColor(image1, cv2.COLOR_BGR2RGB)\n    image2 = cv2.cvtColor(image2, cv2.COLOR_BGR2RGB)\n    \n    fig = plt.figure(figsize=(10, 20))\n    fig.add_subplot(1,2,1)\n    plt.imshow(image1)\n    fig.add_subplot(1,2, 2)\n    plt.imshow(image2)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-29T22:21:26.751026Z","iopub.execute_input":"2022-03-29T22:21:26.751811Z","iopub.status.idle":"2022-03-29T22:21:26.760525Z","shell.execute_reply.started":"2022-03-29T22:21:26.75176Z","shell.execute_reply":"2022-03-29T22:21:26.759443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#data is stored in a id, species, image format\npotential_duplicates=np.load(\"../input/whalesimularitydata/95.npy\")\n\nfor whale1, whale2 in potential_duplicates:\n    if whale1[0] != whale2[0]:\n        show(whale1,whale2)","metadata":{"execution":{"iopub.status.busy":"2022-03-29T22:21:28.587308Z","iopub.execute_input":"2022-03-29T22:21:28.58814Z","iopub.status.idle":"2022-03-29T22:21:32.824556Z","shell.execute_reply.started":"2022-03-29T22:21:28.588035Z","shell.execute_reply":"2022-03-29T22:21:32.822879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"../input/happy-whale-and-dolphin/train.csv\")\nhw1 = train[train.individual_id == \"33dfa6052821\"]\nhw2 = train[train.individual_id == \"79fe61c84d86\"]\nprint(\"Number of photos of each of the humpback whales in question\")\nprint(len(hw1), \"and\", len(hw2))\nkw1 = train[train.individual_id == \"1b589cc6179d\"]\nkw2 = train[train.individual_id == \"fc5088954f84\"]\nprint(\"Number of photos of each of the killer whales in question\")\nprint(len(kw1), \"and\", len(kw2))","metadata":{"execution":{"iopub.status.busy":"2022-03-29T22:08:49.26253Z","iopub.execute_input":"2022-03-29T22:08:49.262845Z","iopub.status.idle":"2022-03-29T22:08:49.385071Z","shell.execute_reply.started":"2022-03-29T22:08:49.262807Z","shell.execute_reply":"2022-03-29T22:08:49.384132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# MORE analysis\n\nI went to bed and wanted to find out if these bad actor photos were also in the test set and could possibly result in lower lb scores due to bad data. I couldn't find the photos in the test set, so we're probably safe. I did find many more \"near duplicate\" photos scattered throughout though (81 in total). This is the result of me running the same ImageHash across all the images in the dataset and returning any matches >95% similarity. Obviously some are false positives, but **I invite the competition organizers to manually review the images in question to ensure that mis-identification (like happened in the training set) isn't ALSO happening in the test dataset.**","metadata":{}},{"cell_type":"code","source":"def attempt_to_show(whale1, whale2):\n    species=[[],[]]\n    individual=[[],[]]\n    \n    image1 = cv2.imread('../input/happy-whale-and-dolphin/train_images/%s' % (whale1))\n    image2 = cv2.imread('../input/happy-whale-and-dolphin/train_images/%s' % (whale2))\n    if image1 is None:\n        image1 = cv2.imread('../input/happy-whale-and-dolphin/test_images/%s' % (whale1))\n        species[0]=\"TEST\"\n        individual[0]=\"TEST\"\n    else:\n        species[0] = train[train.image == whale1].iloc[0].species\n        individual[0] = train[train.image == whale1].iloc[0].individual_id\n        \n    if image2 is None:\n        image2 = cv2.imread('../input/happy-whale-and-dolphin/test_images/%s' % (whale2))\n        species[1]=\"TEST\"\n        individual[1]=\"TEST\"\n    else:\n        species[1] = train[train.image == whale2].iloc[0].species\n        individual[1] = train[train.image == whale2].iloc[0].individual_id\n        \n    print('Image: %s / %s' % (whale1, whale2))\n    print('Species: %s / %s' % (species[0],species[1]))\n    print('Individual: %s / %s' % (individual[0], individual[1]))\n    \n    image1 = cv2.cvtColor(image1, cv2.COLOR_BGR2RGB)\n    image2 = cv2.cvtColor(image2, cv2.COLOR_BGR2RGB)\n    \n    fig = plt.figure(figsize=(10, 20))\n    fig.add_subplot(1,2,1)\n    plt.imshow(image1)\n    fig.add_subplot(1,2, 2)\n    plt.imshow(image2)\n    plt.show()\n\n\nall_duplicates=np.load(\"../input/whalesimularitydata/all95.npy\")\ntrain = pd.read_csv(\"../input/happy-whale-and-dolphin/train.csv\")\nfor whale1, whale2 in all_duplicates:\n    attempt_to_show(whale1,whale2)","metadata":{"execution":{"iopub.status.busy":"2022-03-31T00:07:23.241112Z","iopub.execute_input":"2022-03-31T00:07:23.241470Z","iopub.status.idle":"2022-03-31T00:09:38.998667Z","shell.execute_reply.started":"2022-03-31T00:07:23.241433Z","shell.execute_reply":"2022-03-31T00:09:38.997529Z"},"trusted":true},"execution_count":null,"outputs":[]}]}