{
  "id": 313438,
  "title": "Test images are nowhere to be found",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/313438",
  "author_name": "",
  "post_date": "2022-03-17T05:27:46.016601300Z",
  "votes": 10,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi there, I am training a CNN model for this competition. While I have problem with submission. I can't find the files in sample_submission.csv either in training image folder or test image folder. As you can test in following code. Is this dataset having mistakes or following some mapping rule? </p>\n<pre><code>import os\nimport pandas as pd\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/test/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\n</code></pre>",
  "messages": [
    {
      "id": "1725445",
      "postDate": "03/17/2022 05:27:46",
      "content": "<p>Hi there, I am training a CNN model for this competition. While I have problem with submission. I can't find the files in sample_submission.csv either in training image folder or test image folder. As you can test in following code. Is this dataset having mistakes or following some mapping rule? </p>\n<pre><code>import os\nimport pandas as pd\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/test/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\n</code></pre>",
      "rawMarkdown": "Hi there, I am training a CNN model for this competition. While I have problem with submission. I can't find the files in sample_submission.csv either in training image folder or test image folder. As you can test in following code. Is this dataset having mistakes or following some mapping rule? \n```\nimport os\nimport pandas as pd\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/test/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\n```",
      "votes": null
    },
    {
      "id": "1725450",
      "postDate": "03/17/2022 05:30:03",
      "content": "<p>Besides a few training images are missing and I have to filter them.</p>",
      "rawMarkdown": "Besides a few training images are missing and I have to filter them.",
      "votes": null
    },
    {
      "id": "1725513",
      "postDate": "03/17/2022 07:04:24",
      "content": "<p>Maybe The Host Forget to rename the files ? I think we could rename them ourselves </p>",
      "rawMarkdown": "Maybe The Host Forget to rename the files ? I think we could rename them ourselves",
      "votes": null
    },
    {
      "id": "1725518",
      "postDate": "03/17/2022 07:07:35",
      "content": "<p>I tried submitting file names in file system, but it's not working.</p>",
      "rawMarkdown": "I tried submitting file names in file system, but it's not working.",
      "votes": null
    },
    {
      "id": "1725550",
      "postDate": "03/17/2022 07:38:46",
      "content": "<p>I'm having the same problem</p>",
      "rawMarkdown": "I'm having the same problem",
      "votes": null
    },
    {
      "id": "1725564",
      "postDate": "03/17/2022 07:52:47",
      "content": "<p>Maybe rename the files in the folder ?</p>",
      "rawMarkdown": "Maybe rename the files in the folder ?",
      "votes": null
    },
    {
      "id": "1725583",
      "postDate": "03/17/2022 08:18:41",
      "content": "<p>I suppose filename field stored in database are encrypted. Competition host forgot to decrypt this field.</p>",
      "rawMarkdown": "I suppose filename field stored in database are encrypted. Competition host forgot to decrypt this field.",
      "votes": null
    },
    {
      "id": "1725618",
      "postDate": "03/17/2022 09:03:22",
      "content": "<p>Can you share the filtered train csv pls ?</p>",
      "rawMarkdown": "Can you share the filtered train csv pls ?",
      "votes": null
    },
    {
      "id": "1725629",
      "postDate": "03/17/2022 09:11:34",
      "content": "<p>You could use following code:</p>\n<pre><code>import pandas as pd\nimport os\ntrain = pd.read_csv(\"../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv\")\ntrain[\"file_path\"] = train[\"image\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\ntrain[\"is_exist\"] = train[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(\"Total Training Samples:\", len(train))\ntrain = train[train.is_exist==True]\nprint(\"Valid Training Samples:\", len(train))\n</code></pre>",
      "rawMarkdown": "You could use following code:\n```python\nimport pandas as pd\nimport os\ntrain = pd.read_csv(\"../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv\")\ntrain[\"file_path\"] = train[\"image\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\ntrain[\"is_exist\"] = train[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(\"Total Training Samples:\", len(train))\ntrain = train[train.is_exist==True]\nprint(\"Valid Training Samples:\", len(train))\n```",
      "votes": null
    },
    {
      "id": "1726107",
      "postDate": "03/17/2022 16:53:45",
      "content": "<p>Thanks for flagging this, and apologies for the error. I will get it fixed this morning. Due to the size of the dataset it will probably be another 4-5 hours before the update is available for download. </p>\n<p>Edit: the corrections are now live. You should only have to download the sample submission file again.</p>",
      "rawMarkdown": "Thanks for flagging this, and apologies for the error. I will get it fixed this morning. Due to the size of the dataset it will probably be another 4-5 hours before the update is available for download. \n\nEdit: the corrections are now live. You should only have to download the sample submission file again.",
      "votes": null
    },
    {
      "id": "1727370",
      "postDate": "03/17/2022 23:18:11",
      "content": "<p>More effective solution:</p>\n<p><code>train_df = pd.read_csv('../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv', index_col = 'image')</code></p>\n<p><code>import os</code></p>\n<p><code>images_present = os.listdir('../input/sorghum-id-fgvc-9/train_images')</code></p>\n<p><code>train_df = train_df.loc[images_present]</code></p>",
      "rawMarkdown": "More effective solution:\n\n`train_df = pd.read_csv('../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv', index_col = 'image')`\n\n`import os`\n\n`images_present = os.listdir('../input/sorghum-id-fgvc-9/train_images')`\n\n`train_df = train_df.loc[images_present]`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1725450,
      "author_name": "lonnieqin",
      "author_url": "",
      "post_date": "03/17/2022 05:30:03",
      "content": "<p>Besides a few training images are missing and I have to filter them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1725618,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/17/2022 09:03:22",
          "content": "<p>Can you share the filtered train csv pls ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725629,
          "author_name": "lonnieqin",
          "author_url": "",
          "post_date": "03/17/2022 09:11:34",
          "content": "<p>You could use following code:</p>\n<pre><code>import pandas as pd\nimport os\ntrain = pd.read_csv(\"../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv\")\ntrain[\"file_path\"] = train[\"image\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\ntrain[\"is_exist\"] = train[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(\"Total Training Samples:\", len(train))\ntrain = train[train.is_exist==True]\nprint(\"Valid Training Samples:\", len(train))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1727370,
          "author_name": "motloch",
          "author_url": "",
          "post_date": "03/17/2022 23:18:11",
          "content": "<p>More effective solution:</p>\n<p><code>train_df = pd.read_csv('../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv', index_col = 'image')</code></p>\n<p><code>import os</code></p>\n<p><code>images_present = os.listdir('../input/sorghum-id-fgvc-9/train_images')</code></p>\n<p><code>train_df = train_df.loc[images_present]</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725513,
      "author_name": "mithilsalunkhe",
      "author_url": "",
      "post_date": "03/17/2022 07:04:24",
      "content": "<p>Maybe The Host Forget to rename the files ? I think we could rename them ourselves </p>",
      "votes": null,
      "replies": [
        {
          "id": 1725518,
          "author_name": "lonnieqin",
          "author_url": "",
          "post_date": "03/17/2022 07:07:35",
          "content": "<p>I tried submitting file names in file system, but it's not working.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725564,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/17/2022 07:52:47",
          "content": "<p>Maybe rename the files in the folder ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725583,
          "author_name": "lonnieqin",
          "author_url": "",
          "post_date": "03/17/2022 08:18:41",
          "content": "<p>I suppose filename field stored in database are encrypted. Competition host forgot to decrypt this field.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1725550,
      "author_name": "tchaye59",
      "author_url": "",
      "post_date": "03/17/2022 07:38:46",
      "content": "<p>I'm having the same problem</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1726107,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "03/17/2022 16:53:45",
      "content": "<p>Thanks for flagging this, and apologies for the error. I will get it fixed this morning. Due to the size of the dataset it will probably be another 4-5 hours before the update is available for download. </p>\n<p>Edit: the corrections are now live. You should only have to download the sample submission file again.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1725445": "Hi there, I am training a CNN model for this competition. While I have problem with submission. I can't find the files in sample_submission.csv either in training image folder or test image folder. As you can test in following code. Is this dataset having mistakes or following some mapping rule? \n```\nimport os\nimport pandas as pd\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/test/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\nsubmission = pd.read_csv(\"../input/sorghum-id-fgvc-9/sample_submission.csv\")\nsubmission[\"file_path\"] = submission[\"filename\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\nsubmission[\"is_exist\"] = submission[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(submission[\"is_exist\"].sum())\n```",
    "1725450": "Besides a few training images are missing and I have to filter them.",
    "1725513": "Maybe The Host Forget to rename the files ? I think we could rename them ourselves",
    "1725518": "I tried submitting file names in file system, but it's not working.",
    "1725550": "I'm having the same problem",
    "1725564": "Maybe rename the files in the folder ?",
    "1725583": "I suppose filename field stored in database are encrypted. Competition host forgot to decrypt this field.",
    "1725618": "Can you share the filtered train csv pls ?",
    "1725629": "You could use following code:\n```python\nimport pandas as pd\nimport os\ntrain = pd.read_csv(\"../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv\")\ntrain[\"file_path\"] = train[\"image\"].apply(lambda image: f\"../input/sorghum-id-fgvc-9/train_images/\" + image)\ntrain[\"is_exist\"] = train[\"file_path\"].apply(lambda file_path: os.path.exists(file_path))\nprint(\"Total Training Samples:\", len(train))\ntrain = train[train.is_exist==True]\nprint(\"Valid Training Samples:\", len(train))\n```",
    "1726107": "Thanks for flagging this, and apologies for the error. I will get it fixed this morning. Due to the size of the dataset it will probably be another 4-5 hours before the update is available for download. \n\nEdit: the corrections are now live. You should only have to download the sample submission file again.",
    "1727370": "More effective solution:\n\n`train_df = pd.read_csv('../input/sorghum-id-fgvc-9/train_cultivar_mapping.csv', index_col = 'image')`\n\n`import os`\n\n`images_present = os.listdir('../input/sorghum-id-fgvc-9/train_images')`\n\n`train_df = train_df.loc[images_present]`"
  },
  "source": "meta"
}