{
  "id": 314661,
  "title": "How to build the submission.csv with secret dataset",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/314661",
  "author_name": "",
  "post_date": "2022-03-23T21:03:06.335336700Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello there,</p>\n<p>I am making an attempt at this kaggle competition. For the moment I hope to achieve around 20% accuracy in 2 hours, but I am not sure how to get the image_id to put in the submission.csv since it is a secret dataset.</p>\n<p>Do I load a test.csv and predict with it ?</p>\n<p>I see some of you already successfully submitted their notebook and I hope you can enlighten me for my eternal gratitude.</p>",
  "messages": [
    {
      "id": "1732930",
      "postDate": "03/23/2022 21:03:06",
      "content": "<p>Hello there,</p>\n<p>I am making an attempt at this kaggle competition. For the moment I hope to achieve around 20% accuracy in 2 hours, but I am not sure how to get the image_id to put in the submission.csv since it is a secret dataset.</p>\n<p>Do I load a test.csv and predict with it ?</p>\n<p>I see some of you already successfully submitted their notebook and I hope you can enlighten me for my eternal gratitude.</p>",
      "rawMarkdown": "Hello there,\n\nI am making an attempt at this kaggle competition. For the moment I hope to achieve around 20% accuracy in 2 hours, but I am not sure how to get the image_id to put in the submission.csv since it is a secret dataset.\n\nDo I load a test.csv and predict with it ?\n\nI see some of you already successfully submitted their notebook and I hope you can enlighten me for my eternal gratitude.",
      "votes": null
    },
    {
      "id": "1732944",
      "postDate": "03/23/2022 21:33:19",
      "content": "<p>I would recommend loading the sample submission and filling in your predictions there.</p>",
      "rawMarkdown": "I would recommend loading the sample submission and filling in your predictions there.",
      "votes": null
    },
    {
      "id": "1732958",
      "postDate": "03/23/2022 21:50:18",
      "content": "<p>Thanks for your answer, I am looking for the data to feed to my model.predict() </p>\n<p>since there is no visible test dataset, how to call it in my notebook ?</p>\n<p>Edit : I'll try to feed all images from test_images to my model for prediction even if only one appears in the dataset. Hope it works because it might take some times …<br>\nMaybe I'll train and save the model in a different notebook and only load and predicts from the competition notebook …</p>",
      "rawMarkdown": "Thanks for your answer, I am looking for the data to feed to my model.predict() \n\nsince there is no visible test dataset, how to call it in my notebook ?\n\nEdit : I'll try to feed all images from test_images to my model for prediction even if only one appears in the dataset. Hope it works because it might take some times ...\nMaybe I'll train and save the model in a different notebook and only load and predicts from the competition notebook ...",
      "votes": null
    },
    {
      "id": "1734493",
      "postDate": "03/25/2022 11:52:08",
      "content": "<p>You can check my <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-classification-inference\" target=\"_blank\">Hotel-ID starter - classification - inference</a> notebook.</p>\n<p>I load all image names in test_images folder, do predictions for each image and save it as submission.csv</p>\n<pre><code>test_df = pd.DataFrame(data={\"image_id\": os.listdir(TEST_DATA_FOLDER), \"hotel_id\": \"\"}).sort_values(by=\"image_id\")\n# do predictions for each image\n# ...\ntest_df.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "rawMarkdown": "You can check my [Hotel-ID starter - classification - inference](https://www.kaggle.com/code/michaln/hotel-id-starter-classification-inference) notebook.\n\nI load all image names in test_images folder, do predictions for each image and save it as submission.csv\n``` python\ntest_df = pd.DataFrame(data={\"image_id\": os.listdir(TEST_DATA_FOLDER), \"hotel_id\": \"\"}).sort_values(by=\"image_id\")\n# do predictions for each image\n# ...\ntest_df.to_csv(\"submission.csv\", index=False)\n```",
      "votes": null
    },
    {
      "id": "1734539",
      "postDate": "03/25/2022 12:32:00",
      "content": "<p>Thank you michaln, I will also check your notebook</p>",
      "rawMarkdown": "Thank you michaln, I will also check your notebook",
      "votes": null
    },
    {
      "id": "1736073",
      "postDate": "03/26/2022 22:54:30",
      "content": "<p>If I could chime in here, the file paths can be obtained using the glob function in the glob package like image_paths = glob(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/*.jpg\"). Then, looping over all image_paths stored as path, os.path.basename(path) returns the filename to be entered into image_id. I see that you have found that the header of the file name column is indeed \"image_id\" instead of \"image.\" The Overview shows the latter which, when used, returns a Submission File Error. Just want to point out that difference as it caused me some confusion.</p>",
      "rawMarkdown": "If I could chime in here, the file paths can be obtained using the glob function in the glob package like image_paths = glob(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/*.jpg\"). Then, looping over all image_paths stored as path, os.path.basename(path) returns the filename to be entered into image_id. I see that you have found that the header of the file name column is indeed \"image_id\" instead of \"image.\" The Overview shows the latter which, when used, returns a Submission File Error. Just want to point out that difference as it caused me some confusion.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1732944,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "03/23/2022 21:33:19",
      "content": "<p>I would recommend loading the sample submission and filling in your predictions there.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1732958,
          "author_name": "laurentpoyet",
          "author_url": "",
          "post_date": "03/23/2022 21:50:18",
          "content": "<p>Thanks for your answer, I am looking for the data to feed to my model.predict() </p>\n<p>since there is no visible test dataset, how to call it in my notebook ?</p>\n<p>Edit : I'll try to feed all images from test_images to my model for prediction even if only one appears in the dataset. Hope it works because it might take some times …<br>\nMaybe I'll train and save the model in a different notebook and only load and predicts from the competition notebook …</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1734493,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "03/25/2022 11:52:08",
      "content": "<p>You can check my <a href=\"https://www.kaggle.com/code/michaln/hotel-id-starter-classification-inference\" target=\"_blank\">Hotel-ID starter - classification - inference</a> notebook.</p>\n<p>I load all image names in test_images folder, do predictions for each image and save it as submission.csv</p>\n<pre><code>test_df = pd.DataFrame(data={\"image_id\": os.listdir(TEST_DATA_FOLDER), \"hotel_id\": \"\"}).sort_values(by=\"image_id\")\n# do predictions for each image\n# ...\ntest_df.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1734539,
          "author_name": "laurentpoyet",
          "author_url": "",
          "post_date": "03/25/2022 12:32:00",
          "content": "<p>Thank you michaln, I will also check your notebook</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1736073,
      "author_name": "solenodon",
      "author_url": "",
      "post_date": "03/26/2022 22:54:30",
      "content": "<p>If I could chime in here, the file paths can be obtained using the glob function in the glob package like image_paths = glob(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/*.jpg\"). Then, looping over all image_paths stored as path, os.path.basename(path) returns the filename to be entered into image_id. I see that you have found that the header of the file name column is indeed \"image_id\" instead of \"image.\" The Overview shows the latter which, when used, returns a Submission File Error. Just want to point out that difference as it caused me some confusion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1732930": "Hello there,\n\nI am making an attempt at this kaggle competition. For the moment I hope to achieve around 20% accuracy in 2 hours, but I am not sure how to get the image_id to put in the submission.csv since it is a secret dataset.\n\nDo I load a test.csv and predict with it ?\n\nI see some of you already successfully submitted their notebook and I hope you can enlighten me for my eternal gratitude.",
    "1732944": "I would recommend loading the sample submission and filling in your predictions there.",
    "1732958": "Thanks for your answer, I am looking for the data to feed to my model.predict() \n\nsince there is no visible test dataset, how to call it in my notebook ?\n\nEdit : I'll try to feed all images from test_images to my model for prediction even if only one appears in the dataset. Hope it works because it might take some times ...\nMaybe I'll train and save the model in a different notebook and only load and predicts from the competition notebook ...",
    "1734493": "You can check my [Hotel-ID starter - classification - inference](https://www.kaggle.com/code/michaln/hotel-id-starter-classification-inference) notebook.\n\nI load all image names in test_images folder, do predictions for each image and save it as submission.csv\n``` python\ntest_df = pd.DataFrame(data={\"image_id\": os.listdir(TEST_DATA_FOLDER), \"hotel_id\": \"\"}).sort_values(by=\"image_id\")\n# do predictions for each image\n# ...\ntest_df.to_csv(\"submission.csv\", index=False)\n```",
    "1734539": "Thank you michaln, I will also check your notebook",
    "1736073": "If I could chime in here, the file paths can be obtained using the glob function in the glob package like image_paths = glob(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/*.jpg\"). Then, looping over all image_paths stored as path, os.path.basename(path) returns the filename to be entered into image_id. I see that you have found that the header of the file name column is indeed \"image_id\" instead of \"image.\" The Overview shows the latter which, when used, returns a Submission File Error. Just want to point out that difference as it caused me some confusion."
  },
  "source": "meta"
}