{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# Introduction\n\nI found out that many image include datetime in EXIF and older images tend to be`new_whale`. \nCheck below example if you are interested.\n\n# Example\n## Import Packages\n"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"from pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom PIL import Image\nfrom PIL.ExifTags import TAGS\n\nfrom tqdm import tqdm","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a7a4c54a38ee0b841c485c46fccb2a2b92bf6c6d"},"cell_type":"markdown","source":"## Make function to get EXIF"},{"metadata":{"trusted":true,"_uuid":"146f8883809937409b8c78c47d2743cc26a5a65b"},"cell_type":"code","source":"def get_exif(file):\n\n    img = Image.open(file)\n\n    # check if img have exif or not\n    try:\n        exif = img._getexif()\n    except AttributeError:\n        return {}\n\n    if exif is None:\n        return {}\n\n    # prepare dict\n    exif_table = {}\n\n    # convert keys\n    for tag_id, value in exif.items():\n        exif_table[TAGS.get(tag_id, tag_id)] = value\n\n    return exif_table","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9e62b2fd6c76ecb8d79f666137cde8beb7694b5f"},"cell_type":"markdown","source":"## Make function to get datetime for each images"},{"metadata":{"trusted":true,"_uuid":"ea22c25ebac01444f58a6574a32bd41d34316a0c"},"cell_type":"code","source":"def get_datetiems(dir_images, target_images):\n\n    datetimes = list()\n\n    for img in tqdm(target_images):\n\n        exif = get_exif(Path(dir_images) / img)\n\n        if 'DateTime' in exif.keys():\n            try:\n                dt = datetime.strptime(exif['DateTime'], '%Y:%m:%d %H:%M:%S')\n            except ValueError:\n                dt = ''\n            datetimes.append(dt)\n        else:\n            datetimes.append('')\n\n    return datetimes","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b71046f7bdf4710baa8b8f6b981ac86b9502a84"},"cell_type":"markdown","source":"## Load train.csv"},{"metadata":{"trusted":true,"_uuid":"6e4605410bf6d6d2a9dc11127884b4734dcd9ea4"},"cell_type":"code","source":"df = pd.read_csv('../input/train.csv')\nprint(df.head())\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"097e5e2bf54201e6b05b0a33657b11abca93604a"},"cell_type":"markdown","source":"## Get datetimes"},{"metadata":{"trusted":true,"_uuid":"e5b5b901a81ebd55b7a8c528322b11619014388c"},"cell_type":"code","source":"df['DateTime'] = get_datetiems('../input/train', df.Image.tolist())\nprint(df.head())\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f22eba457569b8800b67bf8d006a623fe8491545"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"d6c73ffeca18f9e002e73b1bedce536bf90f9cad"},"cell_type":"markdown","source":"## Filter by Id"},{"metadata":{"trusted":true,"_uuid":"87131ad6f9ed9992752a02adb37383e260fa7d14"},"cell_type":"code","source":"df['is_new_whale'] = df['Id'] == 'new_whale'\ndf = df.dropna(axis=0)\nprint('number of images which have datetime: {}'.format(len(df)))\n\ndf['year'] = df['DateTime'].dt.year\n\nnew_whales = df.query('is_new_whale==True').groupby('year').agg(len).Image\nprint('number of new_whales')\nprint(new_whales)\n\nnon_new_whales = df.query('is_new_whale==False').groupby('year').agg(len).Image\nprint('number of non new_whales')\nprint(non_new_whales)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bfb1291a435babb0136525c73aaf9390eaf5f83e"},"cell_type":"markdown","source":"## Make the figure"},{"metadata":{"trusted":true,"_uuid":"b27206072d1bc3b0c33eb870c60c8ec91a577be4"},"cell_type":"code","source":"xticks = list(range(2003, 2019, 1))\nind = np.arange(len(xticks))\nwidth = 0.35\n\nplt.figure(figsize=(12,8))\n\np1 = plt.bar(ind, new_whales.tolist(), width)\np2 = plt.bar(ind, non_new_whales.tolist(), width, bottom=new_whales.tolist())\n\nplt.xticks(ind, xticks)\nplt.xlabel('year')\nplt.ylabel('counts')\nplt.legend((p1[0], p2[0]), ('new_whale', 'non new_whale'))\n\n\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"483f413f8f3c0a5b43eaed2dd083dfd6af8f7645"},"cell_type":"markdown","source":"Most of the image taken in 2005 are new_whale. On the other hand, more than half of images taken in 2016-2018 are not new_whale. \n\n# Conclusion\n\nExif information may be useful to indentify `new_whale`. Although we can't divide them by simple thresholding, we can use this information as one of the features. Does this feature help us ? Please your comment."},{"metadata":{"_uuid":"c12b305daea6eaaa6d6ec49f82d091bd48256591"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"25bdf2919114b55a4fbea4bd6b543a9dcb6bae9f"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"6af457d666152f8d8db35d47a3b08dc21947fb05"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"d4db94002e6532686a04b3bb313963e5bffe3c88"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"580d3d40a20ce05f911e51e03c28b1006722f1f3"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}