{
  "id": 576529,
  "title": "About test data directory structure",
  "url": "/competitions/image-matching-challenge-2025/discussion/576529",
  "author_name": "",
  "post_date": "2025-05-05T15:20:04.526802200Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,<br>\nI would like to confirm the expected directory structure of the test images.<br>\nIn the <strong>Overview</strong>, it says:</p>\n<blockquote>\n  <p>\"For test data, we combine all images into a single folder.\"</p>\n</blockquote>\n<p>From this, I assumed the test image structure would be <strong>flat</strong>, like:<br>\ntest/image1.png<br>\ntest/image2.png<br>\n…</p>\n<p>However, some baseline-like code seems to assume a <strong>nested directory</strong> structure, such as:<br>\ntest//.png</p>\n<p>This raises the question:</p>\n<blockquote>\n  <p>💬 Which structure should we rely on for submission:  <br>\n  <strong>(A)</strong> the flat structure (<code>test/*.png</code>) mentioned in the overview,  <br>\n  or  <br>\n  <strong>(B)</strong> the nested structure (<code>test/&lt;dataset&gt;/&lt;image&gt;.png</code>) used in the code?</p>\n</blockquote>\n<p>I'm concerned about aligning my code with the actual structure used during evaluation.</p>\n<p>Thanks for clarifying!</p>",
  "messages": [
    {
      "id": "3194252",
      "postDate": "05/05/2025 15:20:04",
      "content": "<p>Hi,<br>\nI would like to confirm the expected directory structure of the test images.<br>\nIn the <strong>Overview</strong>, it says:</p>\n<blockquote>\n  <p>\"For test data, we combine all images into a single folder.\"</p>\n</blockquote>\n<p>From this, I assumed the test image structure would be <strong>flat</strong>, like:<br>\ntest/image1.png<br>\ntest/image2.png<br>\n…</p>\n<p>However, some baseline-like code seems to assume a <strong>nested directory</strong> structure, such as:<br>\ntest//.png</p>\n<p>This raises the question:</p>\n<blockquote>\n  <p>💬 Which structure should we rely on for submission:  <br>\n  <strong>(A)</strong> the flat structure (<code>test/*.png</code>) mentioned in the overview,  <br>\n  or  <br>\n  <strong>(B)</strong> the nested structure (<code>test/&lt;dataset&gt;/&lt;image&gt;.png</code>) used in the code?</p>\n</blockquote>\n<p>I'm concerned about aligning my code with the actual structure used during evaluation.</p>\n<p>Thanks for clarifying!</p>",
      "rawMarkdown": "Hi,\nI would like to confirm the expected directory structure of the test images.\nIn the **Overview**, it says:\n> \"For test data, we combine all images into a single folder.\"\n\nFrom this, I assumed the test image structure would be **flat**, like:\ntest/image1.png\ntest/image2.png\n...\n\nHowever, some baseline-like code seems to assume a **nested directory** structure, such as:\ntest/<dataset>/<image>.png\n\n\nThis raises the question:\n\n> 💬 Which structure should we rely on for submission:  \n> **(A)** the flat structure (`test/*.png`) mentioned in the overview,  \n> or  \n> **(B)** the nested structure (`test/<dataset>/<image>.png`) used in the code?\n\nI'm concerned about aligning my code with the actual structure used during evaluation.\n\nThanks for clarifying!",
      "votes": null
    },
    {
      "id": "3194576",
      "postDate": "05/06/2025 04:17:17",
      "content": "<p>There are multiple datasets, which group different scenes together. So you'd have:</p>\n<pre><code>dataset1/scene1/image1\ndataset1/scene1/image2\ndataset1/scene2/image1\ndataset2/scene1/image1\netc\n</code></pre>\n<p>The scenes are grouped together, and you'll have to separate them. So in the training data, you'll see:</p>\n<pre><code>dataset1/scene1_image1\ndataset1/scene1_image2\ndataset1/scene2_image1\ndataset2/scene1_image1\netc\n</code></pre>\n<p>We concatenate scene and image names just so that it's easier for you to understand the data at a glance. You can't obviously do this on the (hidden) test data, where filenames will be randomized.</p>\n<p>Datasets are completely independent. You should process them one at a time, and we'll average performance over them.</p>",
      "rawMarkdown": "There are multiple datasets, which group different scenes together. So you'd have:\n```\ndataset1/scene1/image1\ndataset1/scene1/image2\ndataset1/scene2/image1\ndataset2/scene1/image1\netc\n```\n\nThe scenes are grouped together, and you'll have to separate them. So in the training data, you'll see:\n```\ndataset1/scene1_image1\ndataset1/scene1_image2\ndataset1/scene2_image1\ndataset2/scene1_image1\netc\n```\n\nWe concatenate scene and image names just so that it's easier for you to understand the data at a glance. You can't obviously do this on the (hidden) test data, where filenames will be randomized.\n\nDatasets are completely independent. You should process them one at a time, and we'll average performance over them.",
      "votes": null
    },
    {
      "id": "3194609",
      "postDate": "05/06/2025 05:40:19",
      "content": "<p>Thank you for your reply.<br>\nI'm new to Kaggle, so please explain a bit more.<br>\nIs the following understanding correct?</p>\n<p>(1) The structure of the test data is assumed to be as follows:<br>\n<code>test/dataset_name/random_image_name.png</code></p>\n<p>(2) The notebook should automatically search for the dataset_name directly under the test folder and automatically obtain the image name under it.</p>\n<p>(3) Image data should not be obtained based on the dataset or image written in <code>sample_submissions.csv</code>.</p>\n<p>Thank you in advance.</p>",
      "rawMarkdown": "Thank you for your reply.\nI'm new to Kaggle, so please explain a bit more.\nIs the following understanding correct?\n\n(1) The structure of the test data is assumed to be as follows:\n`test/dataset_name/random_image_name.png`\n\n(2) The notebook should automatically search for the dataset_name directly under the test folder and automatically obtain the image name under it.\n\n(3) Image data should not be obtained based on the dataset or image written in `sample_submissions.csv`.\n\nThank you in advance.",
      "votes": null
    },
    {
      "id": "3198462",
      "postDate": "05/09/2025 13:03:58",
      "content": "<p>Where is stert public notebook fore this competition?</p>",
      "rawMarkdown": "Where is stert public notebook fore this competition?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3194576,
      "author_name": "eduardtrulls",
      "author_url": "",
      "post_date": "05/06/2025 04:17:17",
      "content": "<p>There are multiple datasets, which group different scenes together. So you'd have:</p>\n<pre><code>dataset1/scene1/image1\ndataset1/scene1/image2\ndataset1/scene2/image1\ndataset2/scene1/image1\netc\n</code></pre>\n<p>The scenes are grouped together, and you'll have to separate them. So in the training data, you'll see:</p>\n<pre><code>dataset1/scene1_image1\ndataset1/scene1_image2\ndataset1/scene2_image1\ndataset2/scene1_image1\netc\n</code></pre>\n<p>We concatenate scene and image names just so that it's easier for you to understand the data at a glance. You can't obviously do this on the (hidden) test data, where filenames will be randomized.</p>\n<p>Datasets are completely independent. You should process them one at a time, and we'll average performance over them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3194609,
          "author_name": "est2mzd",
          "author_url": "",
          "post_date": "05/06/2025 05:40:19",
          "content": "<p>Thank you for your reply.<br>\nI'm new to Kaggle, so please explain a bit more.<br>\nIs the following understanding correct?</p>\n<p>(1) The structure of the test data is assumed to be as follows:<br>\n<code>test/dataset_name/random_image_name.png</code></p>\n<p>(2) The notebook should automatically search for the dataset_name directly under the test folder and automatically obtain the image name under it.</p>\n<p>(3) Image data should not be obtained based on the dataset or image written in <code>sample_submissions.csv</code>.</p>\n<p>Thank you in advance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3198462,
      "author_name": "johndoerry",
      "author_url": "",
      "post_date": "05/09/2025 13:03:58",
      "content": "<p>Where is stert public notebook fore this competition?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3194252": "Hi,\nI would like to confirm the expected directory structure of the test images.\nIn the **Overview**, it says:\n> \"For test data, we combine all images into a single folder.\"\n\nFrom this, I assumed the test image structure would be **flat**, like:\ntest/image1.png\ntest/image2.png\n...\n\nHowever, some baseline-like code seems to assume a **nested directory** structure, such as:\ntest/<dataset>/<image>.png\n\n\nThis raises the question:\n\n> 💬 Which structure should we rely on for submission:  \n> **(A)** the flat structure (`test/*.png`) mentioned in the overview,  \n> or  \n> **(B)** the nested structure (`test/<dataset>/<image>.png`) used in the code?\n\nI'm concerned about aligning my code with the actual structure used during evaluation.\n\nThanks for clarifying!",
    "3194576": "There are multiple datasets, which group different scenes together. So you'd have:\n```\ndataset1/scene1/image1\ndataset1/scene1/image2\ndataset1/scene2/image1\ndataset2/scene1/image1\netc\n```\n\nThe scenes are grouped together, and you'll have to separate them. So in the training data, you'll see:\n```\ndataset1/scene1_image1\ndataset1/scene1_image2\ndataset1/scene2_image1\ndataset2/scene1_image1\netc\n```\n\nWe concatenate scene and image names just so that it's easier for you to understand the data at a glance. You can't obviously do this on the (hidden) test data, where filenames will be randomized.\n\nDatasets are completely independent. You should process them one at a time, and we'll average performance over them.",
    "3194609": "Thank you for your reply.\nI'm new to Kaggle, so please explain a bit more.\nIs the following understanding correct?\n\n(1) The structure of the test data is assumed to be as follows:\n`test/dataset_name/random_image_name.png`\n\n(2) The notebook should automatically search for the dataset_name directly under the test folder and automatically obtain the image name under it.\n\n(3) Image data should not be obtained based on the dataset or image written in `sample_submissions.csv`.\n\nThank you in advance.",
    "3198462": "Where is stert public notebook fore this competition?"
  },
  "source": "meta"
}