{
  "id": 268171,
  "title": "How many images in private test dataset?",
  "url": "/competitions/landmark-retrieval-2021/discussion/268171",
  "author_name": "",
  "post_date": "2021-08-26T09:19:58.800125100Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Does anyone know how many images in private test dataset?</p>",
  "messages": [
    {
      "id": "1491280",
      "postDate": "08/26/2021 09:19:58",
      "content": "<p>Does anyone know how many images in private test dataset?</p>",
      "rawMarkdown": "Does anyone know how many images in private test dataset?",
      "votes": null
    },
    {
      "id": "1492033",
      "postDate": "08/26/2021 20:00:16",
      "content": "<p>Should be two times the public number, around 2400</p>",
      "rawMarkdown": "Should be two times the public number, around 2400",
      "votes": null
    },
    {
      "id": "1492095",
      "postDate": "08/26/2021 21:43:03",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <br>\nIs that estimated by runtime, or did you find any information from the hosts? edit: nevermind, the test samples we see are the full public portion. </p>\n<p>So, we can assume the index to be identical or does it also get larger in inference? I didn't waste a submission to check, yet. </p>",
      "rawMarkdown": "Thanks @wowfattie \nIs that estimated by runtime, or did you find any information from the hosts? edit: nevermind, the test samples we see are the full public portion. \n\nSo, we can assume the index to be identical or does it also get larger in inference? I didn't waste a submission to check, yet.",
      "votes": null
    },
    {
      "id": "1492147",
      "postDate": "08/26/2021 23:59:16",
      "content": "<p>I'm also wondering whether the index set change during submission. If not, this<br>\n could be a 'leaky' kernel competition.</p>",
      "rawMarkdown": "I'm also wondering whether the index set change during submission. If not, this\n could be a 'leaky' kernel competition.",
      "votes": null
    },
    {
      "id": "1492153",
      "postDate": "08/27/2021 00:12:36",
      "content": "<p>Thanks for the answer. <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <br>\nAnd I found some description of last competition.<br>\n<a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/data\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/data</a></p>\n<blockquote>\n  <p>The provided index/ and test/ images in the publicly available dataset are provided to mock the size and structure of the private data, but are otherwise not directly used.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for the answer. @wowfattie \nAnd I found some description of last competition.\nhttps://www.kaggle.com/c/landmark-retrieval-2020/data\n> The provided index/ and test/ images in the publicly available dataset are provided to mock the size and structure of the private data, but are otherwise not directly used.",
      "votes": null
    },
    {
      "id": "1492693",
      "postDate": "08/27/2021 11:28:31",
      "content": "<blockquote>\n  <p>I'm also wondering whether the index set change during submission. If not, this<br>\n   could be a 'leaky' kernel competition.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> apparently, they are exchanged:</p>\n<p>I checked using this little code snippet</p>\n<pre><code>if np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n</code></pre>\n<p>and the submission failed<br>\n<img src=\"https://i.imgur.com/IBHtv12.png\" alt=\"failed\"></p>",
      "rawMarkdown": "> I'm also wondering whether the index set change during submission. If not, this\n>  could be a 'leaky' kernel competition.\n\n@wowfattie apparently, they are exchanged:\n\nI checked using this little code snippet\n```\nif np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n```\nand the submission failed\n![failed](https://i.imgur.com/IBHtv12.png)",
      "votes": null
    },
    {
      "id": "1503109",
      "postDate": "09/05/2021 02:17:02",
      "content": "<p>According to the <a href=\"https://arxiv.org/pdf/2108.08874.pdf\" target=\"_blank\">paper</a>:</p>\n<blockquote>\n  <p>• Public eval dataset (retrieval challenge): 514 images<br>\n  sampled/downloaded from the GLDv2 training dataset<br>\n  and Wikimedia. This only contains the images of the<br>\n  landmarks that are present in the above index dataset.<br>\n  • Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and<br>\n  Wikimedia. This only contains the images of the landmarks that are present in the above index dataset.</p>\n</blockquote>",
      "rawMarkdown": "According to the [paper](https://arxiv.org/pdf/2108.08874.pdf):\n> • Public eval dataset (retrieval challenge): 514 images\nsampled/downloaded from the GLDv2 training dataset\nand Wikimedia. This only contains the images of the\nlandmarks that are present in the above index dataset.\n• Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and\nWikimedia. This only contains the images of the landmarks that are present in the above index dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1492033,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/26/2021 20:00:16",
      "content": "<p>Should be two times the public number, around 2400</p>",
      "votes": null,
      "replies": [
        {
          "id": 1492095,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "08/26/2021 21:43:03",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <br>\nIs that estimated by runtime, or did you find any information from the hosts? edit: nevermind, the test samples we see are the full public portion. </p>\n<p>So, we can assume the index to be identical or does it also get larger in inference? I didn't waste a submission to check, yet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1492147,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "08/26/2021 23:59:16",
          "content": "<p>I'm also wondering whether the index set change during submission. If not, this<br>\n could be a 'leaky' kernel competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1492693,
              "author_name": "ilu000",
              "author_url": "",
              "post_date": "08/27/2021 11:28:31",
              "content": "<blockquote>\n  <p>I'm also wondering whether the index set change during submission. If not, this<br>\n   could be a 'leaky' kernel competition.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> apparently, they are exchanged:</p>\n<p>I checked using this little code snippet</p>\n<pre><code>if np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n</code></pre>\n<p>and the submission failed<br>\n<img src=\"https://i.imgur.com/IBHtv12.png\" alt=\"failed\"></p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1492153,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "08/27/2021 00:12:36",
          "content": "<p>Thanks for the answer. <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> <br>\nAnd I found some description of last competition.<br>\n<a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/data\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/data</a></p>\n<blockquote>\n  <p>The provided index/ and test/ images in the publicly available dataset are provided to mock the size and structure of the private data, but are otherwise not directly used.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1503109,
      "author_name": "skull8888888",
      "author_url": "",
      "post_date": "09/05/2021 02:17:02",
      "content": "<p>According to the <a href=\"https://arxiv.org/pdf/2108.08874.pdf\" target=\"_blank\">paper</a>:</p>\n<blockquote>\n  <p>• Public eval dataset (retrieval challenge): 514 images<br>\n  sampled/downloaded from the GLDv2 training dataset<br>\n  and Wikimedia. This only contains the images of the<br>\n  landmarks that are present in the above index dataset.<br>\n  • Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and<br>\n  Wikimedia. This only contains the images of the landmarks that are present in the above index dataset.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1491280": "Does anyone know how many images in private test dataset?",
    "1492033": "Should be two times the public number, around 2400",
    "1492095": "Thanks @wowfattie \nIs that estimated by runtime, or did you find any information from the hosts? edit: nevermind, the test samples we see are the full public portion. \n\nSo, we can assume the index to be identical or does it also get larger in inference? I didn't waste a submission to check, yet.",
    "1492147": "I'm also wondering whether the index set change during submission. If not, this\n could be a 'leaky' kernel competition.",
    "1492153": "Thanks for the answer. @wowfattie \nAnd I found some description of last competition.\nhttps://www.kaggle.com/c/landmark-retrieval-2020/data\n> The provided index/ and test/ images in the publicly available dataset are provided to mock the size and structure of the private data, but are otherwise not directly used.",
    "1492693": "> I'm also wondering whether the index set change during submission. If not, this\n>  could be a 'leaky' kernel competition.\n\n@wowfattie apparently, they are exchanged:\n\nI checked using this little code snippet\n```\nif np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n```\nand the submission failed\n![failed](https://i.imgur.com/IBHtv12.png)",
    "1503109": "According to the [paper](https://arxiv.org/pdf/2108.08874.pdf):\n> • Public eval dataset (retrieval challenge): 514 images\nsampled/downloaded from the GLDv2 training dataset\nand Wikimedia. This only contains the images of the\nlandmarks that are present in the above index dataset.\n• Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and\nWikimedia. This only contains the images of the landmarks that are present in the above index dataset."
  },
  "source": "meta"
}