{
  "id": 30337,
  "title": "Check Image Size and Missing Data",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30337",
  "author_name": "",
  "post_date": "2017-03-18T18:39:33.134095Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong><a href=\"https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z\">As amaia pointed out</a>,  some of the images have a problem.</strong></p>\n\n<p>I wrote a script to check dataset (height, width, color channels and missing data), <br>\nand ran this script on my local machine and kaggle's kernel script.  </p>\n\n<p>I intend to check 0 byte files and prematured filse (they are partially missing data). <br>\nI attached the output csv file of this script.</p>\n\n<p><strong>In conclusion, I found 6 missing data files.</strong> <br>\n<strong>But there may be more images with other problems.</strong> <br>\n<strong>If you find them, I want you to report them.</strong> </p>\n\n<ul>\n<li><p>0 byte files (Thanks to <a href=\"https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z\">amaia</a>) <br>\nadditional/Type_2/2845.jpg <br>\nadditional/Type_2/5892.jpg <br>\nadditional/Type_2/5893.jpg  </p></li>\n<li><p>Premature end of JPEG file <br>\ntrain/Type_1/1339.jpg (missing about 45% data) <br>\nadditional/Type_1/3068.jpg (missing about 22% data) <br>\nadditional/Type_2/7.jpg (missing about 25% data)  </p></li>\n</ul>\n\n<p>Just to make sure, I wrote kaggle's kernels <br>\n(the output csv files are same between my local machine and kaggle's kernel, of cource). <br>\nIn kaggle's kernel, there is computing-time limitation (1200 sec.). So I divided into 4 kernels. <br>\nEach of outputs can combine easily by the following command in new directory (only includes 4 csv files).  </p>\n\n<p><code>awk 'FNR==1 &amp;&amp; NR!=1{next;}{print}' *.csv &gt; merge.csv</code></p>\n\n<p><a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-1-of-4\">kernel 1</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-2-of-4\">kernel 2</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-3-of-4\">kernel 3</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-4-of-4\">kernel 4</a></p>",
  "messages": [
    {
      "id": "168941",
      "postDate": "03/18/2017 18:39:33",
      "content": "<p><strong><a href=\"https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z\">As amaia pointed out</a>,  some of the images have a problem.</strong></p>\n\n<p>I wrote a script to check dataset (height, width, color channels and missing data), <br>\nand ran this script on my local machine and kaggle's kernel script.  </p>\n\n<p>I intend to check 0 byte files and prematured filse (they are partially missing data). <br>\nI attached the output csv file of this script.</p>\n\n<p><strong>In conclusion, I found 6 missing data files.</strong> <br>\n<strong>But there may be more images with other problems.</strong> <br>\n<strong>If you find them, I want you to report them.</strong> </p>\n\n<ul>\n<li><p>0 byte files (Thanks to <a href=\"https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z\">amaia</a>) <br>\nadditional/Type_2/2845.jpg <br>\nadditional/Type_2/5892.jpg <br>\nadditional/Type_2/5893.jpg  </p></li>\n<li><p>Premature end of JPEG file <br>\ntrain/Type_1/1339.jpg (missing about 45% data) <br>\nadditional/Type_1/3068.jpg (missing about 22% data) <br>\nadditional/Type_2/7.jpg (missing about 25% data)  </p></li>\n</ul>\n\n<p>Just to make sure, I wrote kaggle's kernels <br>\n(the output csv files are same between my local machine and kaggle's kernel, of cource). <br>\nIn kaggle's kernel, there is computing-time limitation (1200 sec.). So I divided into 4 kernels. <br>\nEach of outputs can combine easily by the following command in new directory (only includes 4 csv files).  </p>\n\n<p><code>awk 'FNR==1 &amp;&amp; NR!=1{next;}{print}' *.csv &gt; merge.csv</code></p>\n\n<p><a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-1-of-4\">kernel 1</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-2-of-4\">kernel 2</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-3-of-4\">kernel 3</a>, <a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-4-of-4\">kernel 4</a></p>",
      "rawMarkdown": "**[As amaia pointed out](https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z),  some of the images have a problem.**\n\nI wrote a script to check dataset (height, width, color channels and missing data),  \nand ran this script on my local machine and kaggle's kernel script.  \n\nI intend to check 0 byte files and prematured filse (they are partially missing data).  \nI attached the output csv file of this script.\n\n**In conclusion, I found 6 missing data files.**  \n**But there may be more images with other problems.**  \n**If you find them, I want you to report them.** \n\n* 0 byte files (Thanks to [amaia](https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z))  \nadditional/Type_2/2845.jpg  \nadditional/Type_2/5892.jpg  \nadditional/Type_2/5893.jpg  \n\n* Premature end of JPEG file  \ntrain/Type_1/1339.jpg (missing about 45% data)  \nadditional/Type_1/3068.jpg (missing about 22% data)  \nadditional/Type_2/7.jpg (missing about 25% data)  \n\n\nJust to make sure, I wrote kaggle's kernels  \n(the output csv files are same between my local machine and kaggle's kernel, of cource).  \nIn kaggle's kernel, there is computing-time limitation (1200 sec.). So I divided into 4 kernels.  \nEach of outputs can combine easily by the following command in new directory (only includes 4 csv files).  \n\n```awk 'FNR==1 && NR!=1{next;}{print}' *.csv > merge.csv```\n\n[kernel 1](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-1-of-4), [kernel 2](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-2-of-4), [kernel 3](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-3-of-4), [kernel 4](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-4-of-4)",
      "votes": null
    },
    {
      "id": "168973",
      "postDate": "03/18/2017 21:49:33",
      "content": "<p>Can you do a md5 checksum? These are my checksums:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143</a></p>",
      "rawMarkdown": "Can you do a md5 checksum? These are my checksums:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143",
      "votes": null
    },
    {
      "id": "169006",
      "postDate": "03/19/2017 03:50:38",
      "content": "<blockquote>\n  <p><strong>visoft wrote</strong></p>\n  \n  <blockquote>\n    <p>Can you do a md5 checksum? These are my checksums:\n    <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143</a></p>\n  </blockquote>\n</blockquote>\n\n<p>Thank you for your advice.</p>\n\n<p><a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-hashsum-md5-and-sha1-of-dataset\">I checked hashsum on 2 or 3 independent environment.</a> <br>\nSo, their images are certainly broken.  I'm worried about how to handle their files...</p>",
      "rawMarkdown": "> **visoft wrote**\n> \n> > Can you do a md5 checksum? These are my checksums:\n> https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\n\nThank you for your advice.\n\n[I checked hashsum on 2 or 3 independent environment.](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-hashsum-md5-and-sha1-of-dataset)  \nSo, their images are certainly broken.  I'm worried about how to handle their files...",
      "votes": null
    },
    {
      "id": "169688",
      "postDate": "03/22/2017 05:21:12",
      "content": "<p>I found the same six problem files, including the three that were zero length.</p>\n\n<p>I computed md5sums on all files and found 84 pairs of duplicates, e.g., here are three pairs:</p>\n\n<pre><code>           pathName filenameExt    size nrow ncol                           md5sum\n1              test     230.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n2 additional/Type_1    1825.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n3              test     178.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n4 additional/Type_1    2151.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n5      train/Type_2     165.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44\n6 additional/Type_1    2320.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44      \n</code></pre>\n\n<p>22 of the test images appear to be duplicates with known type assignments!</p>\n\n<p>The variety of image sizes was a bit curious.  Here's a crosstab of the number of rows by the number of columns in the images</p>\n\n<pre><code>       480  640 2448 3088 3096 3264 4128 4160\n480     0   78    0    0    0    0    0    0\n640     1    0    0    0    0    0    0    0\n2322    0    0    0    0    0    0   99    0\n2448    0    0    0    0    0 4181    0    0\n3096    0    0    0    0    0    0 3891    0\n3120    0    0    0    0    0    0    0  446\n3264    0    0  169    0    0    0    0    0\n4128    0    0    0    1   48    0    0    0  \n</code></pre>",
      "rawMarkdown": "I found the same six problem files, including the three that were zero length.\n\nI computed md5sums on all files and found 84 pairs of duplicates, e.g., here are three pairs:\n\n               pathName filenameExt    size nrow ncol                           md5sum\n    1              test     230.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n    2 additional/Type_1    1825.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n    3              test     178.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n    4 additional/Type_1    2151.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n    5      train/Type_2     165.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44\n    6 additional/Type_1    2320.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44      \n\n22 of the test images appear to be duplicates with known type assignments!\n\nThe variety of image sizes was a bit curious.  Here's a crosstab of the number of rows by the number of columns in the images\n\n           480  640 2448 3088 3096 3264 4128 4160\n    480     0   78    0    0    0    0    0    0\n    640     1    0    0    0    0    0    0    0\n    2322    0    0    0    0    0    0   99    0\n    2448    0    0    0    0    0 4181    0    0\n    3096    0    0    0    0    0    0 3891    0\n    3120    0    0    0    0    0    0    0  446\n    3264    0    0  169    0    0    0    0    0\n    4128    0    0    0    1   48    0    0    0",
      "votes": null
    },
    {
      "id": "169689",
      "postDate": "03/22/2017 05:24:59",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "169690",
      "postDate": "03/22/2017 05:25:42",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "169772",
      "postDate": "03/22/2017 14:41:24",
      "content": "<p>Thank you very much for your comment. They are very interesting insights.</p>\n\n<p>The duplication of files are not only train-set duplication, but also wrong-label and test-dataset leakage!! That's a serious problem!</p>\n\n<p>The variety of image may be caused by camera devices like smartphone.\nI did not look in detail of them and meta tags. I feel like I really learned something.</p>",
      "rawMarkdown": "Thank you very much for your comment. They are very interesting insights.\n\nThe duplication of files are not only train-set duplication, but also wrong-label and test-dataset leakage!! That's a serious problem!\n\nThe variety of image may be caused by camera devices like smartphone.\nI did not look in detail of them and meta tags. I feel like I really learned something.",
      "votes": null
    },
    {
      "id": "182673",
      "postDate": "05/15/2017 10:00:38",
      "content": "<p>Thank you for listing out the odd files. \nFor file with 0 bytes 5893.jpg is in Type_1. </p>",
      "rawMarkdown": "Thank you for listing out the odd files. \nFor file with 0 bytes 5893.jpg is in Type_1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 168973,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "03/18/2017 21:49:33",
      "content": "<p>Can you do a md5 checksum? These are my checksums:\n<a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 169006,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "03/19/2017 03:50:38",
          "content": "<blockquote>\n  <p><strong>visoft wrote</strong></p>\n  \n  <blockquote>\n    <p>Can you do a md5 checksum? These are my checksums:\n    <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143</a></p>\n  </blockquote>\n</blockquote>\n\n<p>Thank you for your advice.</p>\n\n<p><a href=\"https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-hashsum-md5-and-sha1-of-dataset\">I checked hashsum on 2 or 3 independent environment.</a> <br>\nSo, their images are certainly broken.  I'm worried about how to handle their files...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169688,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "03/22/2017 05:21:12",
      "content": "<p>I found the same six problem files, including the three that were zero length.</p>\n\n<p>I computed md5sums on all files and found 84 pairs of duplicates, e.g., here are three pairs:</p>\n\n<pre><code>           pathName filenameExt    size nrow ncol                           md5sum\n1              test     230.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n2 additional/Type_1    1825.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n3              test     178.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n4 additional/Type_1    2151.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n5      train/Type_2     165.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44\n6 additional/Type_1    2320.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44      \n</code></pre>\n\n<p>22 of the test images appear to be duplicates with known type assignments!</p>\n\n<p>The variety of image sizes was a bit curious.  Here's a crosstab of the number of rows by the number of columns in the images</p>\n\n<pre><code>       480  640 2448 3088 3096 3264 4128 4160\n480     0   78    0    0    0    0    0    0\n640     1    0    0    0    0    0    0    0\n2322    0    0    0    0    0    0   99    0\n2448    0    0    0    0    0 4181    0    0\n3096    0    0    0    0    0    0 3891    0\n3120    0    0    0    0    0    0    0  446\n3264    0    0  169    0    0    0    0    0\n4128    0    0    0    1   48    0    0    0  \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 169772,
          "author_name": "kambarakun",
          "author_url": "",
          "post_date": "03/22/2017 14:41:24",
          "content": "<p>Thank you very much for your comment. They are very interesting insights.</p>\n\n<p>The duplication of files are not only train-set duplication, but also wrong-label and test-dataset leakage!! That's a serious problem!</p>\n\n<p>The variety of image may be caused by camera devices like smartphone.\nI did not look in detail of them and meta tags. I feel like I really learned something.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 169689,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "03/22/2017 05:24:59",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 169690,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "03/22/2017 05:25:42",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 182673,
      "author_name": "vinurai",
      "author_url": "",
      "post_date": "05/15/2017 10:00:38",
      "content": "<p>Thank you for listing out the odd files. \nFor file with 0 bytes 5893.jpg is in Type_1. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "168941": "**[As amaia pointed out](https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z),  some of the images have a problem.**\n\nI wrote a script to check dataset (height, width, color channels and missing data),  \nand ran this script on my local machine and kaggle's kernel script.  \n\nI intend to check 0 byte files and prematured filse (they are partially missing data).  \nI attached the output csv file of this script.\n\n**In conclusion, I found 6 missing data files.**  \n**But there may be more images with other problems.**  \n**If you find them, I want you to report them.** \n\n* 0 byte files (Thanks to [amaia](https://www.kaggle.com/aamaia/intel-mobileodt-cervical-cancer-screening/three-empty-images-in-additional-7z))  \nadditional/Type_2/2845.jpg  \nadditional/Type_2/5892.jpg  \nadditional/Type_2/5893.jpg  \n\n* Premature end of JPEG file  \ntrain/Type_1/1339.jpg (missing about 45% data)  \nadditional/Type_1/3068.jpg (missing about 22% data)  \nadditional/Type_2/7.jpg (missing about 25% data)  \n\n\nJust to make sure, I wrote kaggle's kernels  \n(the output csv files are same between my local machine and kaggle's kernel, of cource).  \nIn kaggle's kernel, there is computing-time limitation (1200 sec.). So I divided into 4 kernels.  \nEach of outputs can combine easily by the following command in new directory (only includes 4 csv files).  \n\n```awk 'FNR==1 && NR!=1{next;}{print}' *.csv > merge.csv```\n\n[kernel 1](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-1-of-4), [kernel 2](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-2-of-4), [kernel 3](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-3-of-4), [kernel 4](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-all-dataset-size-and-missing-4-of-4)",
    "168973": "Can you do a md5 checksum? These are my checksums:\nhttps://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143",
    "169006": "> **visoft wrote**\n> \n> > Can you do a md5 checksum? These are my checksums:\n> https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30187#168143\n\nThank you for your advice.\n\n[I checked hashsum on 2 or 3 independent environment.](https://www.kaggle.com/kambarakun/intel-mobileodt-cervical-cancer-screening/check-hashsum-md5-and-sha1-of-dataset)  \nSo, their images are certainly broken.  I'm worried about how to handle their files...",
    "169688": "I found the same six problem files, including the three that were zero length.\n\nI computed md5sums on all files and found 84 pairs of duplicates, e.g., here are three pairs:\n\n               pathName filenameExt    size nrow ncol                           md5sum\n    1              test     230.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n    2 additional/Type_1    1825.jpg 2496221 2448 3264 6d35f947b6ed19ab1a8e14780e9cac49\n    3              test     178.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n    4 additional/Type_1    2151.jpg 7883984 3096 4128 7b880e972c354921a1db55316532d4a5\n    5      train/Type_2     165.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44\n    6 additional/Type_1    2320.jpg 5873812 3096 4128 373faf2f38d55b2a827c69e136686c44      \n\n22 of the test images appear to be duplicates with known type assignments!\n\nThe variety of image sizes was a bit curious.  Here's a crosstab of the number of rows by the number of columns in the images\n\n           480  640 2448 3088 3096 3264 4128 4160\n    480     0   78    0    0    0    0    0    0\n    640     1    0    0    0    0    0    0    0\n    2322    0    0    0    0    0    0   99    0\n    2448    0    0    0    0    0 4181    0    0\n    3096    0    0    0    0    0    0 3891    0\n    3120    0    0    0    0    0    0    0  446\n    3264    0    0  169    0    0    0    0    0\n    4128    0    0    0    1   48    0    0    0",
    "169689": "",
    "169690": "",
    "169772": "Thank you very much for your comment. They are very interesting insights.\n\nThe duplication of files are not only train-set duplication, but also wrong-label and test-dataset leakage!! That's a serious problem!\n\nThe variety of image may be caused by camera devices like smartphone.\nI did not look in detail of them and meta tags. I feel like I really learned something.",
    "182673": "Thank you for listing out the odd files. \nFor file with 0 bytes 5893.jpg is in Type_1."
  },
  "source": "meta"
}