{
  "id": 268381,
  "title": "Are \"index\"-images the same at the test-run?",
  "url": "/competitions/landmark-retrieval-2021/discussion/268381",
  "author_name": "",
  "post_date": "2021-08-27T06:55:08.500445800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Does anyone know if the index-images remain the same during the hidden test run?</p>",
  "messages": [
    {
      "id": "1492425",
      "postDate": "08/27/2021 06:55:08",
      "content": "<p>Does anyone know if the index-images remain the same during the hidden test run?</p>",
      "rawMarkdown": "Does anyone know if the index-images remain the same during the hidden test run?",
      "votes": null
    },
    {
      "id": "1492691",
      "postDate": "08/27/2021 11:27:37",
      "content": "<p>Apparently they are different.<br>\nI checked using this little code snippet</p>\n<pre><code>if np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n</code></pre>\n<p>and the submission failed<br>\n<img src=\"https://i.imgur.com/IBHtv12.png\" alt=\"failed\"></p>",
      "rawMarkdown": "Apparently they are different.\nI checked using this little code snippet\n```\nif np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n```\nand the submission failed\n![failed](https://i.imgur.com/IBHtv12.png)",
      "votes": null
    },
    {
      "id": "1496420",
      "postDate": "08/30/2021 11:14:23",
      "content": "<p>Wow… Maybe something is wrong with me, but that's so not obvious. Data section clearly states \"When you submit your notebook, Kaggle will rerun your code on the private dataset\". </p>\n<p>However, I genuinely thought it has to do with private re-run with the same index images. </p>\n<p>Okay, I guess that explains why I have 0 scored submissions when I try to simply upload my .csv (it's possible to do so in other competitions) </p>",
      "rawMarkdown": "Wow... Maybe something is wrong with me, but that's so not obvious. Data section clearly states \"When you submit your notebook, Kaggle will rerun your code on the private dataset\". \n\nHowever, I genuinely thought it has to do with private re-run with the same index images. \n\nOkay, I guess that explains why I have 0 scored submissions when I try to simply upload my .csv (it's possible to do so in other competitions)",
      "votes": null
    },
    {
      "id": "1496425",
      "postDate": "08/30/2021 11:19:46",
      "content": "<p>Indeed it is not obvious, and personally, i'd like maximum information about how the data is structured/exchanged in the submission run.<br>\nA lot can be tested by doing submissions as shown above, but it always feels like unnecessary overhead work. </p>",
      "rawMarkdown": "Indeed it is not obvious, and personally, i'd like maximum information about how the data is structured/exchanged in the submission run.\nA lot can be tested by doing submissions as shown above, but it always feels like unnecessary overhead work.",
      "votes": null
    },
    {
      "id": "1503103",
      "postDate": "09/05/2021 02:03:00",
      "content": "<p>They are indeed different, here is the info from the <a href=\"https://arxiv.org/pdf/2108.08874.pdf\" target=\"_blank\">paper</a>:</p>\n<blockquote>\n  <p>Index dataset (retrieval challenge): 100,000 images<br>\n  sampled from the GLDv2 training dataset.<br>\n  • Public eval dataset (retrieval challenge): 514 images<br>\n  sampled/downloaded from the GLDv2 training dataset<br>\n  and Wikimedia. This only contains the images of the<br>\n  landmarks that are present in the above index dataset.<br>\n  • Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and<br>\n  Wikimedia. This only contains the images of the landmarks that are present in the above index dataset.</p>\n</blockquote>",
      "rawMarkdown": "They are indeed different, here is the info from the [paper](https://arxiv.org/pdf/2108.08874.pdf):\n> Index dataset (retrieval challenge): 100,000 images\nsampled from the GLDv2 training dataset.\n• Public eval dataset (retrieval challenge): 514 images\nsampled/downloaded from the GLDv2 training dataset\nand Wikimedia. This only contains the images of the\nlandmarks that are present in the above index dataset.\n• Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and\nWikimedia. This only contains the images of the landmarks that are present in the above index dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1492691,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "08/27/2021 11:27:37",
      "content": "<p>Apparently they are different.<br>\nI checked using this little code snippet</p>\n<pre><code>if np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n</code></pre>\n<p>and the submission failed<br>\n<img src=\"https://i.imgur.com/IBHtv12.png\" alt=\"failed\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1496420,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "08/30/2021 11:14:23",
          "content": "<p>Wow… Maybe something is wrong with me, but that's so not obvious. Data section clearly states \"When you submit your notebook, Kaggle will rerun your code on the private dataset\". </p>\n<p>However, I genuinely thought it has to do with private re-run with the same index images. </p>\n<p>Okay, I guess that explains why I have 0 scored submissions when I try to simply upload my .csv (it's possible to do so in other competitions) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1496425,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "08/30/2021 11:19:46",
          "content": "<p>Indeed it is not obvious, and personally, i'd like maximum information about how the data is structured/exchanged in the submission run.<br>\nA lot can be tested by doing submissions as shown above, but it always feels like unnecessary overhead work. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1503103,
          "author_name": "skull8888888",
          "author_url": "",
          "post_date": "09/05/2021 02:03:00",
          "content": "<p>They are indeed different, here is the info from the <a href=\"https://arxiv.org/pdf/2108.08874.pdf\" target=\"_blank\">paper</a>:</p>\n<blockquote>\n  <p>Index dataset (retrieval challenge): 100,000 images<br>\n  sampled from the GLDv2 training dataset.<br>\n  • Public eval dataset (retrieval challenge): 514 images<br>\n  sampled/downloaded from the GLDv2 training dataset<br>\n  and Wikimedia. This only contains the images of the<br>\n  landmarks that are present in the above index dataset.<br>\n  • Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and<br>\n  Wikimedia. This only contains the images of the landmarks that are present in the above index dataset.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1492425": "Does anyone know if the index-images remain the same during the hidden test run?",
    "1492691": "Apparently they are different.\nI checked using this little code snippet\n```\nif np.array_equal(index_ids_public, np.array(index_ids)):\n    print(\"match\")\nelse:\n    index_ids = []\n    index_df = pd.DataFrame({\"id\": index_ids})\n    index_df[\"landmark_id\"] = -1\n```\nand the submission failed\n![failed](https://i.imgur.com/IBHtv12.png)",
    "1496420": "Wow... Maybe something is wrong with me, but that's so not obvious. Data section clearly states \"When you submit your notebook, Kaggle will rerun your code on the private dataset\". \n\nHowever, I genuinely thought it has to do with private re-run with the same index images. \n\nOkay, I guess that explains why I have 0 scored submissions when I try to simply upload my .csv (it's possible to do so in other competitions)",
    "1496425": "Indeed it is not obvious, and personally, i'd like maximum information about how the data is structured/exchanged in the submission run.\nA lot can be tested by doing submissions as shown above, but it always feels like unnecessary overhead work.",
    "1503103": "They are indeed different, here is the info from the [paper](https://arxiv.org/pdf/2108.08874.pdf):\n> Index dataset (retrieval challenge): 100,000 images\nsampled from the GLDv2 training dataset.\n• Public eval dataset (retrieval challenge): 514 images\nsampled/downloaded from the GLDv2 training dataset\nand Wikimedia. This only contains the images of the\nlandmarks that are present in the above index dataset.\n• Private eval dataset (retrieval challenge): 1028 sampled/downloaded from the GLDv2 training dataset and\nWikimedia. This only contains the images of the landmarks that are present in the above index dataset."
  },
  "source": "meta"
}