{
  "id": 207693,
  "title": "Mismatch between csv file columns and number of images",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207693",
  "author_name": "",
  "post_date": "2020-12-30T21:25:40.733459500Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey Fellow Kagglers,</p>\n<p>I am currently trying too load the dataset into TensorFlow using tf.keras.preprocessing.image_dataset_from_directory but there appears to be a mismatch in terms of the dimensions. The dataframe - \"train.csv\" - has 21397 rows while there are about 21402 training images. Does this mean that there are about 5 unlabeled images? If so, what can I do to circumvent my problem? </p>",
  "messages": [
    {
      "id": "1133019",
      "postDate": "12/30/2020 21:25:40",
      "content": "<p>Hey Fellow Kagglers,</p>\n<p>I am currently trying too load the dataset into TensorFlow using tf.keras.preprocessing.image_dataset_from_directory but there appears to be a mismatch in terms of the dimensions. The dataframe - \"train.csv\" - has 21397 rows while there are about 21402 training images. Does this mean that there are about 5 unlabeled images? If so, what can I do to circumvent my problem? </p>",
      "rawMarkdown": "Hey Fellow Kagglers,\n\nI am currently trying too load the dataset into TensorFlow using tf.keras.preprocessing.image_dataset_from_directory but there appears to be a mismatch in terms of the dimensions. The dataframe - \"train.csv\" - has 21397 rows while there are about 21402 training images. Does this mean that there are about 5 unlabeled images? If so, what can I do to circumvent my problem?",
      "votes": null
    },
    {
      "id": "1133039",
      "postDate": "12/30/2020 21:42:59",
      "content": "<p>The last TFRecord is labeled as having fewer images (1327 instead of 1338 for all the others).</p>",
      "rawMarkdown": "The last TFRecord is labeled as having fewer images (1327 instead of 1338 for all the others).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1133039,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "12/30/2020 21:42:59",
      "content": "<p>The last TFRecord is labeled as having fewer images (1327 instead of 1338 for all the others).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1133019": "Hey Fellow Kagglers,\n\nI am currently trying too load the dataset into TensorFlow using tf.keras.preprocessing.image_dataset_from_directory but there appears to be a mismatch in terms of the dimensions. The dataframe - \"train.csv\" - has 21397 rows while there are about 21402 training images. Does this mean that there are about 5 unlabeled images? If so, what can I do to circumvent my problem?",
    "1133039": "The last TFRecord is labeled as having fewer images (1327 instead of 1338 for all the others)."
  },
  "source": "meta"
}