{
  "id": 78464,
  "title": "Found EXIF information from images",
  "url": "/competitions/humpback-whale-identification/discussion/78464",
  "author_name": "Yiheng Wang",
  "post_date": "2019-01-24T07:01:59.516000",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I just found that EXIF information for some images were not removed in the dataset. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11096/samples.png\" alt=\"enter image description here\">\nThen, I used the python package \"exifread\" to detect all images, and found that more than 25% images for both train and test set have these data.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11095/exif.png\" alt=\"enter image description here\">\nAre these information useful for our task? I tried to do some basic analysis but failed. \nAlthough didn't find any useful things, I wish to publish this finding here, for everyone, for equity. I planned to upload the code into the kernel, but \"pip install exifread\" seems useless. Thus I just post the simple extract code here. If you are interested, feel free to dig from them : )</p>",
  "messages": [
    {
      "id": 460653,
      "postDate": "2019-01-24T07:01:59.517Z",
      "content": "<p>I just found that EXIF information for some images were not removed in the dataset. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11096/samples.png\" alt=\"enter image description here\">\nThen, I used the python package \"exifread\" to detect all images, and found that more than 25% images for both train and test set have these data.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11095/exif.png\" alt=\"enter image description here\">\nAre these information useful for our task? I tried to do some basic analysis but failed. \nAlthough didn't find any useful things, I wish to publish this finding here, for everyone, for equity. I planned to upload the code into the kernel, but \"pip install exifread\" seems useless. Thus I just post the simple extract code here. If you are interested, feel free to dig from them : )</p>",
      "rawMarkdown": "I just found that EXIF information for some images were not removed in the dataset. \n![enter image description here][1]\nThen, I used the python package \"exifread\" to detect all images, and found that more than 25% images for both train and test set have these data.\n![enter image description here][2]\nAre these information useful for our task? I tried to do some basic analysis but failed. \nAlthough didn't find any useful things, I wish to publish this finding here, for everyone, for equity. I planned to upload the code into the kernel, but \"pip install exifread\" seems useless. Thus I just post the simple extract code here. If you are interested, feel free to dig from them : )\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11096/samples.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11095/exif.png",
      "votes": 11
    },
    {
      "id": 460655,
      "postDate": "2019-01-24T07:04:58.273Z",
      "content": "<p>from pandas import read_csv\nimport warnings\nimport exifread\nwarnings.filterwarnings('ignore')</p>\n\n<p>Path = '../input/'\ntrain_list = [p for <em>,p,</em> in read_csv(Path + 'train.csv').to_records()]\ntest_list = [p for <em>,p,</em> in read_csv(Path + 'sample_submission.csv').to_records()]</p>\n\n<p>exif_dict = {}\ncount = 0\nfor img in train_list:\n    f = open(Path + 'train/' + img, 'rb')\n    try:\n        tags = exifread.process_file(f, details=False)\n        if len(tags) &gt; 0:\n            exif_dict[img] = tags\n    except:\n        count += 1</p>\n\n<p>import pickle\nwith open('exif_train.pickle', 'wb') as f: pickle.dump(exif_dict, f)</p>",
      "rawMarkdown": "from pandas import read_csv\nimport warnings\nimport exifread\nwarnings.filterwarnings('ignore')\n\nPath = '../input/'\ntrain_list = [p for _,p,_ in read_csv(Path + 'train.csv').to_records()]\ntest_list = [p for _,p,_ in read_csv(Path + 'sample_submission.csv').to_records()]\n\nexif_dict = {}\ncount = 0\nfor img in train_list:\n    f = open(Path + 'train/' + img, 'rb')\n    try:\n        tags = exifread.process_file(f, details=False)\n        if len(tags) &gt; 0:\n            exif_dict[img] = tags\n    except:\n        count += 1\n\nimport pickle\nwith open('exif_train.pickle', 'wb') as f: pickle.dump(exif_dict, f)",
      "votes": 4
    },
    {
      "id": 460982,
      "postDate": "2019-01-24T23:29:57.867Z",
      "content": "<p>That's potentially a leak but not a major one since it doesn't concern all the images: you can use the metadata to connect (or at least cluster) the images in the train set and the test set without using the pixels using location, artist, camera type...\nSince the first Avito competition, Kaggle has usually been really careful to remove the metadata unless the sponsor wants the competitors to use them. </p>",
      "rawMarkdown": "That's potentially a leak but not a major one since it doesn't concern all the images: you can use the metadata to connect (or at least cluster) the images in the train set and the test set without using the pixels using location, artist, camera type...\nSince the first Avito competition, Kaggle has usually been really careful to remove the metadata unless the sponsor wants the competitors to use them. ",
      "votes": 1
    },
    {
      "id": 470840,
      "postDate": "2019-02-13T16:40:52.457Z",
      "content": "<p>photographer name and date/time of photo taken seems useful</p>",
      "rawMarkdown": "photographer name and date/time of photo taken seems useful",
      "votes": 2
    },
    {
      "id": 460656,
      "postDate": "2019-01-24T07:06:07.747Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 460655,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2019-01-24T07:04:58.273000",
      "content": "<p>from pandas import read_csv\nimport warnings\nimport exifread\nwarnings.filterwarnings('ignore')</p>\n\n<p>Path = '../input/'\ntrain_list = [p for <em>,p,</em> in read_csv(Path + 'train.csv').to_records()]\ntest_list = [p for <em>,p,</em> in read_csv(Path + 'sample_submission.csv').to_records()]</p>\n\n<p>exif_dict = {}\ncount = 0\nfor img in train_list:\n    f = open(Path + 'train/' + img, 'rb')\n    try:\n        tags = exifread.process_file(f, details=False)\n        if len(tags) &gt; 0:\n            exif_dict[img] = tags\n    except:\n        count += 1</p>\n\n<p>import pickle\nwith open('exif_train.pickle', 'wb') as f: pickle.dump(exif_dict, f)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 460982,
      "author_name": "eagle4",
      "author_url": "",
      "post_date": "2019-01-24T23:29:57.867000",
      "content": "<p>That's potentially a leak but not a major one since it doesn't concern all the images: you can use the metadata to connect (or at least cluster) the images in the train set and the test set without using the pixels using location, artist, camera type...\nSince the first Avito competition, Kaggle has usually been really careful to remove the metadata unless the sponsor wants the competitors to use them. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470840,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-13T16:40:52.457000",
      "content": "<p>photographer name and date/time of photo taken seems useful</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 460656,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-24T07:06:07.747000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "460653": "I just found that EXIF information for some images were not removed in the dataset. \n![enter image description here][1]\nThen, I used the python package \"exifread\" to detect all images, and found that more than 25% images for both train and test set have these data.\n![enter image description here][2]\nAre these information useful for our task? I tried to do some basic analysis but failed. \nAlthough didn't find any useful things, I wish to publish this finding here, for everyone, for equity. I planned to upload the code into the kernel, but \"pip install exifread\" seems useless. Thus I just post the simple extract code here. If you are interested, feel free to dig from them : )\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11096/samples.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/460653/11095/exif.png",
    "460655": "from pandas import read_csv\nimport warnings\nimport exifread\nwarnings.filterwarnings('ignore')\n\nPath = '../input/'\ntrain_list = [p for _,p,_ in read_csv(Path + 'train.csv').to_records()]\ntest_list = [p for _,p,_ in read_csv(Path + 'sample_submission.csv').to_records()]\n\nexif_dict = {}\ncount = 0\nfor img in train_list:\n    f = open(Path + 'train/' + img, 'rb')\n    try:\n        tags = exifread.process_file(f, details=False)\n        if len(tags) &gt; 0:\n            exif_dict[img] = tags\n    except:\n        count += 1\n\nimport pickle\nwith open('exif_train.pickle', 'wb') as f: pickle.dump(exif_dict, f)",
    "460982": "That's potentially a leak but not a major one since it doesn't concern all the images: you can use the metadata to connect (or at least cluster) the images in the train set and the test set without using the pixels using location, artist, camera type...\nSince the first Avito competition, Kaggle has usually been really careful to remove the metadata unless the sponsor wants the competitors to use them. ",
    "470840": "photographer name and date/time of photo taken seems useful",
    "460656": ""
  }
}