{
  "id": 62921,
  "title": "2000 corrupted images in data set",
  "url": "/competitions/airbus-ship-detection/discussion/62921",
  "author_name": "abnerzhang",
  "post_date": "2018-08-09T00:56:55.999000",
  "votes": 25,
  "comment_count": 7,
  "views": 0,
  "content": "<p>There are around 2000 corrupted images in data set, around 1% of total images. \nBe careful with these images.</p>",
  "messages": [
    {
      "id": 367968,
      "postDate": "2018-08-09T00:56:56Z",
      "content": "<p>There are around 2000 corrupted images in data set, around 1% of total images. \nBe careful with these images.</p>",
      "rawMarkdown": "There are around 2000 corrupted images in data set, around 1% of total images. \nBe careful with these images.",
      "votes": 24
    },
    {
      "id": 368675,
      "postDate": "2018-08-10T12:41:08.160Z",
      "content": "<p>I got the list by check whether there is a row or column of original images all the same value, normally all the 0. The model of prediction may not be confused but preprocess like clustering may be affected. </p>",
      "rawMarkdown": "I got the list by check whether there is a row or column of original images all the same value, normally all the 0. The model of prediction may not be confused but preprocess like clustering may be affected. ",
      "votes": 3,
      "replies": [
        {
          "id": 369477,
          "postDate": "2018-08-13T06:18:10.263Z",
          "content": "<p>Indeed, the pre-processing might be impacted. Good thing there are not too many of them, but also normalisation is impacted due to lower variance (but this would be minimal).</p>\n\n<p>I guess possible strategy (both for train &amp; test) is to crop these images and then scale them. However this would cause some \"aspect distortion\".    </p>",
          "rawMarkdown": "Indeed, the pre-processing might be impacted. Good thing there are not too many of them, but also normalisation is impacted due to lower variance (but this would be minimal).\n\nI guess possible strategy (both for train &amp; test) is to crop these images and then scale them. However this would cause some \"aspect distortion\".    "
        },
        {
          "id": 371086,
          "postDate": "2018-08-16T00:40:56.420Z",
          "content": "<p>new data will be provided by <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/63255\">https://www.kaggle.com/c/airbus-ship-detection/discussion/63255</a></p>",
          "rawMarkdown": "new data will be provided by https://www.kaggle.com/c/airbus-ship-detection/discussion/63255"
        }
      ]
    },
    {
      "id": 368218,
      "postDate": "2018-08-09T13:59:14.350Z",
      "content": "<p>Thanks for identifying these images!!! For sure this list can help to create a better balanced training set. Out of curiosity, how did you identify these images?</p>\n\n<p>Not sure I would classify them as corrupt. It looks very much like the artefacts you get as a result of image augmentation. So if for example you scale an image close to the border of the original image you need to decide how to deal with the area that is not present in the original image. </p>\n\n<p>In this case it seems that they selected to use a a solid color (black or blue) to fill the unknown areas. I guess they defined different color values for different augmentations. If they would have always selected black, it would at least looked a bit better.</p>\n\n<p>The good thing is they are present in both the train and test set. So your model should easily start ignoring those large blue and black areas and not mistake them for huge oil tankers ;)</p>",
      "rawMarkdown": "Thanks for identifying these images!!! For sure this list can help to create a better balanced training set. Out of curiosity, how did you identify these images?\n\nNot sure I would classify them as corrupt. It looks very much like the artefacts you get as a result of image augmentation. So if for example you scale an image close to the border of the original image you need to decide how to deal with the area that is not present in the original image. \n\nIn this case it seems that they selected to use a a solid color (black or blue) to fill the unknown areas. I guess they defined different color values for different augmentations. If they would have always selected black, it would at least looked a bit better.\n\nThe good thing is they are present in both the train and test set. So your model should easily start ignoring those large blue and black areas and not mistake them for huge oil tankers ;)\n",
      "votes": 1
    },
    {
      "id": 370590,
      "postDate": "2018-08-15T04:47:42.980Z",
      "content": "<p>As long as the files are valid jpgs, then learning around these rare cases just makes for a more robust model.  However, truly corrupted jpgs like, '6384c3e78.jpg' will break your code on read.  My python code threw,</p>\n\n<pre><code>TypeError: float() argument must be a string or a number, not 'JpegImageFile'\n</code></pre>\n\n<p>when it tried to read this file.  You can check for yourself easily by running,</p>\n\n<pre><code>display 6384c3e78.jpg\n</code></pre>\n\n<p>in your linux terminal.  I got:</p>\n\n<pre><code>display: Premature end of JPEG file `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\ndisplay: Corrupt JPEG data: premature end of data segment `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n</code></pre>\n\n<p>Of course I may just have a bad file.  Can anyone verify?</p>",
      "rawMarkdown": "As long as the files are valid jpgs, then learning around these rare cases just makes for a more robust model.  However, truly corrupted jpgs like, '6384c3e78.jpg' will break your code on read.  My python code threw,\n\n    TypeError: float() argument must be a string or a number, not 'JpegImageFile'\n\nwhen it tried to read this file.  You can check for yourself easily by running,\n\n    display 6384c3e78.jpg\n\nin your linux terminal.  I got:\n\n    display: Premature end of JPEG file `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n    display: Corrupt JPEG data: premature end of data segment `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n\nOf course I may just have a bad file.  Can anyone verify?",
      "replies": [
        {
          "id": 371078,
          "postDate": "2018-08-16T00:12:20.567Z",
          "content": "<p>Yes, 6384c3e78 is un-openable for me too.</p>",
          "rawMarkdown": "Yes, 6384c3e78 is un-openable for me too.",
          "isDeleted": true
        },
        {
          "id": 371085,
          "postDate": "2018-08-16T00:39:29.137Z",
          "content": "<p>this one was discussed in <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/62574\">https://www.kaggle.com/c/airbus-ship-detection/discussion/62574</a>  and the author provided some solutions to deal that one</p>",
          "rawMarkdown": "this one was discussed in https://www.kaggle.com/c/airbus-ship-detection/discussion/62574  and the author provided some solutions to deal that one"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 368675,
      "author_name": "abnerzhang",
      "author_url": "",
      "post_date": "2018-08-10T12:41:08.160000",
      "content": "<p>I got the list by check whether there is a row or column of original images all the same value, normally all the 0. The model of prediction may not be confused but preprocess like clustering may be affected. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 369477,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2018-08-13T06:18:10.263000",
          "content": "<p>Indeed, the pre-processing might be impacted. Good thing there are not too many of them, but also normalisation is impacted due to lower variance (but this would be minimal).</p>\n\n<p>I guess possible strategy (both for train &amp; test) is to crop these images and then scale them. However this would cause some \"aspect distortion\".    </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371086,
          "author_name": "abnerzhang",
          "author_url": "",
          "post_date": "2018-08-16T00:40:56.420000",
          "content": "<p>new data will be provided by <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/63255\">https://www.kaggle.com/c/airbus-ship-detection/discussion/63255</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 368218,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2018-08-09T13:59:14.350000",
      "content": "<p>Thanks for identifying these images!!! For sure this list can help to create a better balanced training set. Out of curiosity, how did you identify these images?</p>\n\n<p>Not sure I would classify them as corrupt. It looks very much like the artefacts you get as a result of image augmentation. So if for example you scale an image close to the border of the original image you need to decide how to deal with the area that is not present in the original image. </p>\n\n<p>In this case it seems that they selected to use a a solid color (black or blue) to fill the unknown areas. I guess they defined different color values for different augmentations. If they would have always selected black, it would at least looked a bit better.</p>\n\n<p>The good thing is they are present in both the train and test set. So your model should easily start ignoring those large blue and black areas and not mistake them for huge oil tankers ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 370590,
      "author_name": "WillieMaddox",
      "author_url": "",
      "post_date": "2018-08-15T04:47:42.980000",
      "content": "<p>As long as the files are valid jpgs, then learning around these rare cases just makes for a more robust model.  However, truly corrupted jpgs like, '6384c3e78.jpg' will break your code on read.  My python code threw,</p>\n\n<pre><code>TypeError: float() argument must be a string or a number, not 'JpegImageFile'\n</code></pre>\n\n<p>when it tried to read this file.  You can check for yourself easily by running,</p>\n\n<pre><code>display 6384c3e78.jpg\n</code></pre>\n\n<p>in your linux terminal.  I got:</p>\n\n<pre><code>display: Premature end of JPEG file `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\ndisplay: Corrupt JPEG data: premature end of data segment `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n</code></pre>\n\n<p>Of course I may just have a bad file.  Can anyone verify?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 371078,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-16T00:12:20.567000",
          "content": "<p>Yes, 6384c3e78 is un-openable for me too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 371085,
          "author_name": "abnerzhang",
          "author_url": "",
          "post_date": "2018-08-16T00:39:29.137000",
          "content": "<p>this one was discussed in <a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/62574\">https://www.kaggle.com/c/airbus-ship-detection/discussion/62574</a>  and the author provided some solutions to deal that one</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "367968": "There are around 2000 corrupted images in data set, around 1% of total images. \nBe careful with these images.",
    "368675": "I got the list by check whether there is a row or column of original images all the same value, normally all the 0. The model of prediction may not be confused but preprocess like clustering may be affected. ",
    "368218": "Thanks for identifying these images!!! For sure this list can help to create a better balanced training set. Out of curiosity, how did you identify these images?\n\nNot sure I would classify them as corrupt. It looks very much like the artefacts you get as a result of image augmentation. So if for example you scale an image close to the border of the original image you need to decide how to deal with the area that is not present in the original image. \n\nIn this case it seems that they selected to use a a solid color (black or blue) to fill the unknown areas. I guess they defined different color values for different augmentations. If they would have always selected black, it would at least looked a bit better.\n\nThe good thing is they are present in both the train and test set. So your model should easily start ignoring those large blue and black areas and not mistake them for huge oil tankers ;)\n",
    "370590": "As long as the files are valid jpgs, then learning around these rare cases just makes for a more robust model.  However, truly corrupted jpgs like, '6384c3e78.jpg' will break your code on read.  My python code threw,\n\n    TypeError: float() argument must be a string or a number, not 'JpegImageFile'\n\nwhen it tried to read this file.  You can check for yourself easily by running,\n\n    display 6384c3e78.jpg\n\nin your linux terminal.  I got:\n\n    display: Premature end of JPEG file `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n    display: Corrupt JPEG data: premature end of data segment `6384c3e78.jpg' @ warning/jpeg.c/JPEGWarningHandler/352.\n\nOf course I may just have a bad file.  Can anyone verify?"
  }
}