{
  "id": 313543,
  "title": "One bad image file and CSV has filenames not in images directory",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/313543",
  "author_name": "Gerry",
  "post_date": "2022-03-17T17:39:13.804000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>When I ran my code using the ImageDataGenerator.flow_from_date I found a few issues. For one the file 2017-06-23__16-27-08-730.png appears to be an invalid image file. I tried to read it in using PIL, CV2 and matplot imread. So I adjusted the dataframe to not include this file. Second, there are 5624 file names reference in the CSV file that are not images in the train_images directory. I eliminated those from the dataframe as well. After that things seemed to be OK. </p>",
  "messages": [
    {
      "id": 1726161,
      "postDate": "2022-03-17T17:39:13.803Z",
      "content": "<p>When I ran my code using the ImageDataGenerator.flow_from_date I found a few issues. For one the file 2017-06-23__16-27-08-730.png appears to be an invalid image file. I tried to read it in using PIL, CV2 and matplot imread. So I adjusted the dataframe to not include this file. Second, there are 5624 file names reference in the CSV file that are not images in the train_images directory. I eliminated those from the dataframe as well. After that things seemed to be OK. </p>",
      "rawMarkdown": "When I ran my code using the ImageDataGenerator.flow_from_date I found a few issues. For one the file 2017-06-23__16-27-08-730.png appears to be an invalid image file. I tried to read it in using PIL, CV2 and matplot imread. So I adjusted the dataframe to not include this file. Second, there are 5624 file names reference in the CSV file that are not images in the train_images directory. I eliminated those from the dataframe as well. After that things seemed to be OK. ",
      "votes": 6
    },
    {
      "id": 1727409,
      "postDate": "2022-03-18T00:43:33.293Z",
      "content": "<p>Looks like the dataset has been updated and the bad image files removed. There is still one problem, the train_cultivar_mapping.csv file still has 442 references in it that are NOT image files in train_images</p>",
      "rawMarkdown": "Looks like the dataset has been updated and the bad image files removed. There is still one problem, the train_cultivar_mapping.csv file still has 442 references in it that are NOT image files in train_images",
      "votes": 3,
      "replies": [
        {
          "id": 1730353,
          "postDate": "2022-03-21T07:04:24.773Z",
          "content": "<p>Yes, I also found that there are 442 images missing</p>",
          "rawMarkdown": "Yes, I also found that there are 442 images missing"
        },
        {
          "id": 1733726,
          "postDate": "2022-03-24T14:55:05.453Z",
          "content": "<p>Sorry about this! We did a bit of data cleaning before launch to remove some super bright and super dark images, and it looks like the test set CSV got updated but not the training set one. I've sent the updated file over to the Kaggle team and we'll hopefully get it posted shortly. In the meantime, any of the options that folks have posted for skipping over the missing files should work.</p>",
          "rawMarkdown": "Sorry about this! We did a bit of data cleaning before launch to remove some super bright and super dark images, and it looks like the test set CSV got updated but not the training set one. I've sent the updated file over to the Kaggle team and we'll hopefully get it posted shortly. In the meantime, any of the options that folks have posted for skipping over the missing files should work."
        },
        {
          "id": 1734046,
          "postDate": "2022-03-24T23:06:21.153Z",
          "content": "<p>The corrected train_cultivar_mapping.csv should be available now.</p>",
          "rawMarkdown": "The corrected train_cultivar_mapping.csv should be available now."
        }
      ]
    },
    {
      "id": 1730862,
      "postDate": "2022-03-21T17:26:54.723Z",
      "content": "<p>I did this to make a csv without any missing:</p>\n<pre><code>csv = pd.read_csv('train_cultivar_mapping.csv')\nmissing = []\nfor name in list(csv['image'].unique()):\n  if not os.path.exists('train_images/'+name):\n    missing.append(name)\n\ncsv = csv[csv.image.isin(missing) == False]\ncsv.to_csv('nomissing.csv')\n</code></pre>",
      "rawMarkdown": "I did this to make a csv without any missing:\n```\ncsv = pd.read_csv('train_cultivar_mapping.csv')\nmissing = []\nfor name in list(csv['image'].unique()):\n  if not os.path.exists('train_images/'+name):\n    missing.append(name)\n\ncsv = csv[csv.image.isin(missing) == False]\ncsv.to_csv('nomissing.csv')\n```"
    },
    {
      "id": 1726226,
      "postDate": "2022-03-17T19:19:49.577Z",
      "content": "<p>I also found that in the test images there is a non image file called 'dummy - Shortcut.lnk' </p>",
      "rawMarkdown": "I also found that in the test images there is a non image file called 'dummy - Shortcut.lnk' "
    },
    {
      "id": 1730352,
      "postDate": "2022-03-21T07:03:46.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1727409,
      "author_name": "Gerry",
      "author_url": "",
      "post_date": "2022-03-18T00:43:33.293000",
      "content": "<p>Looks like the dataset has been updated and the bad image files removed. There is still one problem, the train_cultivar_mapping.csv file still has 442 references in it that are NOT image files in train_images</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1730353,
          "author_name": "minghigh",
          "author_url": "",
          "post_date": "2022-03-21T07:04:24.773000",
          "content": "<p>Yes, I also found that there are 442 images missing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1733726,
          "author_name": "Abby Stylianou",
          "author_url": "",
          "post_date": "2022-03-24T14:55:05.453000",
          "content": "<p>Sorry about this! We did a bit of data cleaning before launch to remove some super bright and super dark images, and it looks like the test set CSV got updated but not the training set one. I've sent the updated file over to the Kaggle team and we'll hopefully get it posted shortly. In the meantime, any of the options that folks have posted for skipping over the missing files should work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1734046,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2022-03-24T23:06:21.153000",
          "content": "<p>The corrected train_cultivar_mapping.csv should be available now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1730862,
      "author_name": "Chase Phelps",
      "author_url": "",
      "post_date": "2022-03-21T17:26:54.723000",
      "content": "<p>I did this to make a csv without any missing:</p>\n<pre><code>csv = pd.read_csv('train_cultivar_mapping.csv')\nmissing = []\nfor name in list(csv['image'].unique()):\n  if not os.path.exists('train_images/'+name):\n    missing.append(name)\n\ncsv = csv[csv.image.isin(missing) == False]\ncsv.to_csv('nomissing.csv')\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1726226,
      "author_name": "Gerry",
      "author_url": "",
      "post_date": "2022-03-17T19:19:49.577000",
      "content": "<p>I also found that in the test images there is a non image file called 'dummy - Shortcut.lnk' </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1730352,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-21T07:03:46.543000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1726161": "When I ran my code using the ImageDataGenerator.flow_from_date I found a few issues. For one the file 2017-06-23__16-27-08-730.png appears to be an invalid image file. I tried to read it in using PIL, CV2 and matplot imread. So I adjusted the dataframe to not include this file. Second, there are 5624 file names reference in the CSV file that are not images in the train_images directory. I eliminated those from the dataframe as well. After that things seemed to be OK. ",
    "1727409": "Looks like the dataset has been updated and the bad image files removed. There is still one problem, the train_cultivar_mapping.csv file still has 442 references in it that are NOT image files in train_images",
    "1730862": "I did this to make a csv without any missing:\n```\ncsv = pd.read_csv('train_cultivar_mapping.csv')\nmissing = []\nfor name in list(csv['image'].unique()):\n  if not os.path.exists('train_images/'+name):\n    missing.append(name)\n\ncsv = csv[csv.image.isin(missing) == False]\ncsv.to_csv('nomissing.csv')\n```",
    "1726226": "I also found that in the test images there is a non image file called 'dummy - Shortcut.lnk' ",
    "1730352": ""
  }
}