{
  "id": 171153,
  "title": "Where is the \"hidden\" test data?",
  "url": "/competitions/birdsong-recognition/discussion/171153",
  "author_name": "",
  "post_date": "2020-07-30T15:43:22.579452900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This may be a Kaggle beginner's question, but after some experimentation and documentation reading I have to ask: Where is this \"hidden\" test data located? What does \"hidden\" mean? I think this is not clear from the description of the competition.</p>\n\n<p>The data description says:\n```\nFiles</p>\n\n<p>train_audio The train data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org.</p>\n\n<p>test_audio The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long. They will not all fit in a notebook's memory at the same time. The recordings were taken at three separate remote locations in North America. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n```</p>\n\n<p>This seems to imply that there is a folder <code>test_audio</code> next to <code>train_audio</code>. However, it is not accessible to a notebook this way:</p>\n\n<p>```\nimport os</p>\n\n<p>test_audio_path = '/kaggle/input/birdsong-recognition/test_audio'\nprint(\"test_audio exists: \", os.path.exists(test_audio_path))</p>\n\n<p>for dirname, _, filenames in os.walk(test_audio_path):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))</p>\n\n<p>```</p>\n\n<p>Is the path incorrect? Am I missing something?</p>",
  "messages": [
    {
      "id": "952079",
      "postDate": "07/30/2020 15:43:22",
      "content": "<p>This may be a Kaggle beginner's question, but after some experimentation and documentation reading I have to ask: Where is this \"hidden\" test data located? What does \"hidden\" mean? I think this is not clear from the description of the competition.</p>\n\n<p>The data description says:\n```\nFiles</p>\n\n<p>train_audio The train data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org.</p>\n\n<p>test_audio The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long. They will not all fit in a notebook's memory at the same time. The recordings were taken at three separate remote locations in North America. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n```</p>\n\n<p>This seems to imply that there is a folder <code>test_audio</code> next to <code>train_audio</code>. However, it is not accessible to a notebook this way:</p>\n\n<p>```\nimport os</p>\n\n<p>test_audio_path = '/kaggle/input/birdsong-recognition/test_audio'\nprint(\"test_audio exists: \", os.path.exists(test_audio_path))</p>\n\n<p>for dirname, _, filenames in os.walk(test_audio_path):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))</p>\n\n<p>```</p>\n\n<p>Is the path incorrect? Am I missing something?</p>",
      "rawMarkdown": "This may be a Kaggle beginner's question, but after some experimentation and documentation reading I have to ask: Where is this \"hidden\" test data located? What does \"hidden\" mean? I think this is not clear from the description of the competition.\n\nThe data description says:\n```\nFiles\n\ntrain_audio The train data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org.\n\ntest_audio The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long. They will not all fit in a notebook's memory at the same time. The recordings were taken at three separate remote locations in North America. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n```\n\nThis seems to imply that there is a folder `test_audio` next to `train_audio`. However, it is not accessible to a notebook this way:\n\n```\nimport os\n\ntest_audio_path = '/kaggle/input/birdsong-recognition/test_audio'\nprint(\"test_audio exists: \", os.path.exists(test_audio_path))\n\nfor dirname, _, filenames in os.walk(test_audio_path):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n```\n\nIs the path incorrect? Am I missing something?",
      "votes": null
    },
    {
      "id": "952084",
      "postDate": "07/30/2020 15:53:23",
      "content": "<p>you can read about it at:\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">https://www.kaggle.com/c/birdsong-recognition/discussion/159993</a></p>",
      "rawMarkdown": "you can read about it at:\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/158987\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159993",
      "votes": null
    },
    {
      "id": "952105",
      "postDate": "07/30/2020 16:15:35",
      "content": "<p>Thanks. I think the missing two sentences that cause a lot of the confusion are:</p>\n\n<p>On submission, you do not submit a file with predictions, but a notebook that is then re-run. In this phase, hidden data is visible. </p>\n\n<p>The Kaggle user interface suggests implicitly that a submission is a .csv file. </p>",
      "rawMarkdown": "Thanks. I think the missing two sentences that cause a lot of the confusion are:\n\n  On submission, you do not submit a file with predictions, but a notebook that is then re-run. In this phase, hidden data is visible. \n\nThe Kaggle user interface suggests implicitly that a submission is a .csv file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 952084,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2020 15:53:23",
      "content": "<p>you can read about it at:\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158987\">https://www.kaggle.com/c/birdsong-recognition/discussion/158987</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">https://www.kaggle.com/c/birdsong-recognition/discussion/159993</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 952105,
      "author_name": "christianstaudt",
      "author_url": "",
      "post_date": "07/30/2020 16:15:35",
      "content": "<p>Thanks. I think the missing two sentences that cause a lot of the confusion are:</p>\n\n<p>On submission, you do not submit a file with predictions, but a notebook that is then re-run. In this phase, hidden data is visible. </p>\n\n<p>The Kaggle user interface suggests implicitly that a submission is a .csv file. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "952079": "This may be a Kaggle beginner's question, but after some experimentation and documentation reading I have to ask: Where is this \"hidden\" test data located? What does \"hidden\" mean? I think this is not clear from the description of the competition.\n\nThe data description says:\n```\nFiles\n\ntrain_audio The train data consists of short recordings of individual bird calls generously uploaded by users of xenocanto.org.\n\ntest_audio The hidden test_audio directory contains approximately 150 recordings in mp3 format, each roughly 10 minutes long. They will not all fit in a notebook's memory at the same time. The recordings were taken at three separate remote locations in North America. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n```\n\nThis seems to imply that there is a folder `test_audio` next to `train_audio`. However, it is not accessible to a notebook this way:\n\n```\nimport os\n\ntest_audio_path = '/kaggle/input/birdsong-recognition/test_audio'\nprint(\"test_audio exists: \", os.path.exists(test_audio_path))\n\nfor dirname, _, filenames in os.walk(test_audio_path):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n```\n\nIs the path incorrect? Am I missing something?",
    "952084": "you can read about it at:\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/158987\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159993",
    "952105": "Thanks. I think the missing two sentences that cause a lot of the confusion are:\n\n  On submission, you do not submit a file with predictions, but a notebook that is then re-run. In this phase, hidden data is visible. \n\nThe Kaggle user interface suggests implicitly that a submission is a .csv file."
  },
  "source": "meta"
}