{
  "id": 124133,
  "title": "Bad images finding",
  "url": "/competitions/pku-autonomous-driving/discussion/124133",
  "author_name": "",
  "post_date": "2020-01-02T05:32:36.393461700Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>How do you find bad images from dataset? for finding bad images do you visually check every image or is there library or tools or program that can be used to find bad images?\ni want to spend more time finding bad images, a little aid from you will be highly appreciated,thank you in advance!</p>",
  "messages": [
    {
      "id": "708240",
      "postDate": "01/02/2020 05:32:36",
      "content": "<p>How do you find bad images from dataset? for finding bad images do you visually check every image or is there library or tools or program that can be used to find bad images?\ni want to spend more time finding bad images, a little aid from you will be highly appreciated,thank you in advance!</p>",
      "rawMarkdown": "How do you find bad images from dataset? for finding bad images do you visually check every image or is there library or tools or program that can be used to find bad images?\ni want to spend more time finding bad images, a little aid from you will be highly appreciated,thank you in advance!",
      "votes": null
    },
    {
      "id": "708292",
      "postDate": "01/02/2020 06:38:55",
      "content": "<p>Let's define bad images.\n1. Images with bad quality/resolution? So far it seems images are of good quality in this competition.\n2. Bad annotation? They are part of every ML dataset. Usually, neural networks can ignore a few bad samples, for example, 5-10% annotation errors. That's why semi-supervised/weakly-supervised learning works. \n3. Sometimes preprocessing is helpful to improve the quality of images (i.e. noise reduction, contrast enhancement etc.).  It depends on the dataset provided, and one needs to apply his own judgment for that specific problem.\n4. I think it's important to understand the sources of common/easy errors. Sometimes they can be generated from data itself, sometimes can be due to model training style. In either case, they can be time-consuming to find out!</p>",
      "rawMarkdown": "Let's define bad images.\n1. Images with bad quality/resolution? So far it seems images are of good quality in this competition.\n2. Bad annotation? They are part of every ML dataset. Usually, neural networks can ignore a few bad samples, for example, 5-10% annotation errors. That's why semi-supervised/weakly-supervised learning works. \n3. Sometimes preprocessing is helpful to improve the quality of images (i.e. noise reduction, contrast enhancement etc.).  It depends on the dataset provided, and one needs to apply his own judgment for that specific problem.\n4. I think it's important to understand the sources of common/easy errors. Sometimes they can be generated from data itself, sometimes can be due to model training style. In either case, they can be time-consuming to find out!",
      "votes": null
    },
    {
      "id": "708309",
      "postDate": "01/02/2020 06:58:16",
      "content": "<p><a href=\"/sgalib\">@sgalib</a>  thanks for the wisdom mate.\ni was  looking for a way to find more broken images(if exist) like  this post : <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621</a></p>\n\n<p>it is possible that more broken images existing in train directory</p>",
      "rawMarkdown": "sgalib  thanks for the wisdom mate.\ni was  looking for a way to find more broken images(if exist) like  this post : https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621\n\nit is possible that more broken images existing in train directory",
      "votes": null
    },
    {
      "id": "708557",
      "postDate": "01/02/2020 12:31:21",
      "content": "<p>I checked them all manually. This list of \"broken\" images is complete. Fortunately, there are none in the test set, but we have to deal with augmentation there. </p>",
      "rawMarkdown": "I checked them all manually. This list of \"broken\" images is complete. Fortunately, there are none in the test set, but we have to deal with augmentation there.",
      "votes": null
    },
    {
      "id": "708733",
      "postDate": "01/02/2020 15:46:05",
      "content": "<p>Thanks <a href=\"/ilu000\">@ilu000</a> \nThanks \nThanks for the manual checking </p>",
      "rawMarkdown": "Thanks @ilu000 \nThanks \nThanks for the manual checking",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 708292,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/02/2020 06:38:55",
      "content": "<p>Let's define bad images.\n1. Images with bad quality/resolution? So far it seems images are of good quality in this competition.\n2. Bad annotation? They are part of every ML dataset. Usually, neural networks can ignore a few bad samples, for example, 5-10% annotation errors. That's why semi-supervised/weakly-supervised learning works. \n3. Sometimes preprocessing is helpful to improve the quality of images (i.e. noise reduction, contrast enhancement etc.).  It depends on the dataset provided, and one needs to apply his own judgment for that specific problem.\n4. I think it's important to understand the sources of common/easy errors. Sometimes they can be generated from data itself, sometimes can be due to model training style. In either case, they can be time-consuming to find out!</p>",
      "votes": null,
      "replies": [
        {
          "id": 708309,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/02/2020 06:58:16",
          "content": "<p><a href=\"/sgalib\">@sgalib</a>  thanks for the wisdom mate.\ni was  looking for a way to find more broken images(if exist) like  this post : <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621</a></p>\n\n<p>it is possible that more broken images existing in train directory</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708557,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "01/02/2020 12:31:21",
          "content": "<p>I checked them all manually. This list of \"broken\" images is complete. Fortunately, there are none in the test set, but we have to deal with augmentation there. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708733,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/02/2020 15:46:05",
          "content": "<p>Thanks <a href=\"/ilu000\">@ilu000</a> \nThanks \nThanks for the manual checking </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "708240": "How do you find bad images from dataset? for finding bad images do you visually check every image or is there library or tools or program that can be used to find bad images?\ni want to spend more time finding bad images, a little aid from you will be highly appreciated,thank you in advance!",
    "708292": "Let's define bad images.\n1. Images with bad quality/resolution? So far it seems images are of good quality in this competition.\n2. Bad annotation? They are part of every ML dataset. Usually, neural networks can ignore a few bad samples, for example, 5-10% annotation errors. That's why semi-supervised/weakly-supervised learning works. \n3. Sometimes preprocessing is helpful to improve the quality of images (i.e. noise reduction, contrast enhancement etc.).  It depends on the dataset provided, and one needs to apply his own judgment for that specific problem.\n4. I think it's important to understand the sources of common/easy errors. Sometimes they can be generated from data itself, sometimes can be due to model training style. In either case, they can be time-consuming to find out!",
    "708309": "sgalib  thanks for the wisdom mate.\ni was  looking for a way to find more broken images(if exist) like  this post : https://www.kaggle.com/c/pku-autonomous-driving/discussion/117621\n\nit is possible that more broken images existing in train directory",
    "708557": "I checked them all manually. This list of \"broken\" images is complete. Fortunately, there are none in the test set, but we have to deal with augmentation there.",
    "708733": "Thanks @ilu000 \nThanks \nThanks for the manual checking"
  },
  "source": "meta"
}