{
  "id": 173518,
  "title": "Any information from metadata?",
  "url": "/competitions/landmark-recognition-2020/discussion/173518",
  "author_name": "",
  "post_date": "2020-08-09T16:10:24.242016100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was wandering if there is any relationship between the train folders organization (the /x/y/z/xyz....jpg structure) and the landmark ids or image size and landmark ids? \nAt a first quick glance, it seems not (seems random). Has anyone explored this further?</p>",
  "messages": [
    {
      "id": "964129",
      "postDate": "08/09/2020 16:10:24",
      "content": "<p>I was wandering if there is any relationship between the train folders organization (the /x/y/z/xyz....jpg structure) and the landmark ids or image size and landmark ids? \nAt a first quick glance, it seems not (seems random). Has anyone explored this further?</p>",
      "rawMarkdown": "I was wandering if there is any relationship between the train folders organization (the /x/y/z/xyz....jpg structure) and the landmark ids or image size and landmark ids? \nAt a first quick glance, it seems not (seems random). Has anyone explored this further?",
      "votes": null
    },
    {
      "id": "965290",
      "postDate": "08/10/2020 14:25:48",
      "content": "<p>Haven't really explored but I guess they should be random</p>",
      "rawMarkdown": "Haven't really explored but I guess they should be random",
      "votes": null
    },
    {
      "id": "965732",
      "postDate": "08/10/2020 19:54:42",
      "content": "<p>It is really random but done on purpose - to prevent I/O bottleneck for old Linux systems and AWS S3. \nLong story short, if you put a lot of files (like 1.2M images in our case) into one folder with ext3 file system you may observe significant I/O degradation due to linear complexity of search - O(n), so it is preferable to divide uniformly all these files into subfolders. In our case, the authors divided all files into 4096 subfolders (16 ^ 3, simply using 3 first chars from names) which contain approximately the same small amount of files that much easier for ext3 to handle. \nIf you interested in more details, please refer: <a href=\"https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w\">https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w</a> </p>\n\n<p>And for AWS S3 case, which authors use to host the GLDv2 dataset, it is simply about how S3 balance load. By design, S3 uses different servers to handle reqest to objects with different prefixes. Roughly, if you have files <code>s3://a/b/c.jpg</code> and <code>s3://a/d/e.jpg</code> then it is much probable that they will be handled by different workers. Hence, by dividing all files into different prefixes you uniformly divide load across all S3 workers and increase I/O to these files.</p>\n\n<p>P.S.: Names also not random, this is simply hex-representation of image order-number in the dataset to prevent name-collision. </p>",
      "rawMarkdown": "It is really random but done on purpose - to prevent I/O bottleneck for old Linux systems and AWS S3. \nLong story short, if you put a lot of files (like 1.2M images in our case) into one folder with ext3 file system you may observe significant I/O degradation due to linear complexity of search - O(n), so it is preferable to divide uniformly all these files into subfolders. In our case, the authors divided all files into 4096 subfolders (16 ^ 3, simply using 3 first chars from names) which contain approximately the same small amount of files that much easier for ext3 to handle. \nIf you interested in more details, please refer: https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w \n\nAnd for AWS S3 case, which authors use to host the GLDv2 dataset, it is simply about how S3 balance load. By design, S3 uses different servers to handle reqest to objects with different prefixes. Roughly, if you have files `s3://a/b/c.jpg` and `s3://a/d/e.jpg` then it is much probable that they will be handled by different workers. Hence, by dividing all files into different prefixes you uniformly divide load across all S3 workers and increase I/O to these files.\n\nP.S.: Names also not random, this is simply hex-representation of image order-number in the dataset to prevent name-collision.",
      "votes": null
    },
    {
      "id": "965737",
      "postDate": "08/10/2020 20:04:06",
      "content": "<p>Couldn't wish for a better explanation. Many thanks!\nIt is funny that this dataset is hosted in S3. Or maybe it is both hosted on S3 and the GCP equivalent (Cloud Storage). Do you have more details on this? I can of course check the github repo which I am going to do now: <a href=\"https://github.com/cvdfoundation/google-landmark\">https://github.com/cvdfoundation/google-landmark</a>. </p>",
      "rawMarkdown": "Couldn't wish for a better explanation. Many thanks!\nIt is funny that this dataset is hosted in S3. Or maybe it is both hosted on S3 and the GCP equivalent (Cloud Storage). Do you have more details on this? I can of course check the github repo which I am going to do now: https://github.com/cvdfoundation/google-landmark.",
      "votes": null
    },
    {
      "id": "965760",
      "postDate": "08/10/2020 20:38:01",
      "content": "<p>I also found this quite interesting, but the authors did not describe why they choose AWS S3 as storage and I do not have any clue why they did so instead of hosting dataset on inhouse Cloud Storage, which obviously would be much cheaper for Google.</p>",
      "rawMarkdown": "I also found this quite interesting, but the authors did not describe why they choose AWS S3 as storage and I do not have any clue why they did so instead of hosting dataset on inhouse Cloud Storage, which obviously would be much cheaper for Google.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 965290,
      "author_name": "chandanverma",
      "author_url": "",
      "post_date": "08/10/2020 14:25:48",
      "content": "<p>Haven't really explored but I guess they should be random</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965732,
      "author_name": "alexkirnas",
      "author_url": "",
      "post_date": "08/10/2020 19:54:42",
      "content": "<p>It is really random but done on purpose - to prevent I/O bottleneck for old Linux systems and AWS S3. \nLong story short, if you put a lot of files (like 1.2M images in our case) into one folder with ext3 file system you may observe significant I/O degradation due to linear complexity of search - O(n), so it is preferable to divide uniformly all these files into subfolders. In our case, the authors divided all files into 4096 subfolders (16 ^ 3, simply using 3 first chars from names) which contain approximately the same small amount of files that much easier for ext3 to handle. \nIf you interested in more details, please refer: <a href=\"https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w\">https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w</a> </p>\n\n<p>And for AWS S3 case, which authors use to host the GLDv2 dataset, it is simply about how S3 balance load. By design, S3 uses different servers to handle reqest to objects with different prefixes. Roughly, if you have files <code>s3://a/b/c.jpg</code> and <code>s3://a/d/e.jpg</code> then it is much probable that they will be handled by different workers. Hence, by dividing all files into different prefixes you uniformly divide load across all S3 workers and increase I/O to these files.</p>\n\n<p>P.S.: Names also not random, this is simply hex-representation of image order-number in the dataset to prevent name-collision. </p>",
      "votes": null,
      "replies": [
        {
          "id": 965737,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/10/2020 20:04:06",
          "content": "<p>Couldn't wish for a better explanation. Many thanks!\nIt is funny that this dataset is hosted in S3. Or maybe it is both hosted on S3 and the GCP equivalent (Cloud Storage). Do you have more details on this? I can of course check the github repo which I am going to do now: <a href=\"https://github.com/cvdfoundation/google-landmark\">https://github.com/cvdfoundation/google-landmark</a>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965760,
          "author_name": "alexkirnas",
          "author_url": "",
          "post_date": "08/10/2020 20:38:01",
          "content": "<p>I also found this quite interesting, but the authors did not describe why they choose AWS S3 as storage and I do not have any clue why they did so instead of hosting dataset on inhouse Cloud Storage, which obviously would be much cheaper for Google.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "964129": "I was wandering if there is any relationship between the train folders organization (the /x/y/z/xyz....jpg structure) and the landmark ids or image size and landmark ids? \nAt a first quick glance, it seems not (seems random). Has anyone explored this further?",
    "965290": "Haven't really explored but I guess they should be random",
    "965732": "It is really random but done on purpose - to prevent I/O bottleneck for old Linux systems and AWS S3. \nLong story short, if you put a lot of files (like 1.2M images in our case) into one folder with ext3 file system you may observe significant I/O degradation due to linear complexity of search - O(n), so it is preferable to divide uniformly all these files into subfolders. In our case, the authors divided all files into 4096 subfolders (16 ^ 3, simply using 3 first chars from names) which contain approximately the same small amount of files that much easier for ext3 to handle. \nIf you interested in more details, please refer: https://serverfault.com/questions/736872/why-there-shouldnt-be-too-many-files-in-one-directory-that-serves-just-static-w \n\nAnd for AWS S3 case, which authors use to host the GLDv2 dataset, it is simply about how S3 balance load. By design, S3 uses different servers to handle reqest to objects with different prefixes. Roughly, if you have files `s3://a/b/c.jpg` and `s3://a/d/e.jpg` then it is much probable that they will be handled by different workers. Hence, by dividing all files into different prefixes you uniformly divide load across all S3 workers and increase I/O to these files.\n\nP.S.: Names also not random, this is simply hex-representation of image order-number in the dataset to prevent name-collision.",
    "965737": "Couldn't wish for a better explanation. Many thanks!\nIt is funny that this dataset is hosted in S3. Or maybe it is both hosted on S3 and the GCP equivalent (Cloud Storage). Do you have more details on this? I can of course check the github repo which I am going to do now: https://github.com/cvdfoundation/google-landmark.",
    "965760": "I also found this quite interesting, but the authors did not describe why they choose AWS S3 as storage and I do not have any clue why they did so instead of hosting dataset on inhouse Cloud Storage, which obviously would be much cheaper for Google."
  },
  "source": "meta"
}