{
  "id": 267056,
  "title": "Using Data",
  "url": "/competitions/landmark-retrieval-2021/discussion/267056",
  "author_name": "",
  "post_date": "2021-08-21T14:36:37.928913800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I had a few questions relating to the data. </p>\n<ol>\n<li><p>I don't exactly understand how it has been organized, and what each folder is for. I read that it's organized according to characteristics, but what does that exactly mean?</p></li>\n<li><p>Is there any way to iterate through the data during training without downloading it? I was asking because it's a very large dataset.</p></li>\n</ol>",
  "messages": [
    {
      "id": "1484722",
      "postDate": "08/21/2021 14:36:37",
      "content": "<p>I had a few questions relating to the data. </p>\n<ol>\n<li><p>I don't exactly understand how it has been organized, and what each folder is for. I read that it's organized according to characteristics, but what does that exactly mean?</p></li>\n<li><p>Is there any way to iterate through the data during training without downloading it? I was asking because it's a very large dataset.</p></li>\n</ol>",
      "rawMarkdown": "I had a few questions relating to the data. \n\n1. I don't exactly understand how it has been organized, and what each folder is for. I read that it's organized according to characteristics, but what does that exactly mean?\n\n2. Is there any way to iterate through the data during training without downloading it? I was asking because it's a very large dataset.",
      "votes": null
    },
    {
      "id": "1484832",
      "postDate": "08/21/2021 15:39:50",
      "content": "<p>maybe <a href=\"https://www.kaggle.com/ronigur/landmark-retrieval-random-basline\" target=\"_blank\">this notebook</a> will help</p>",
      "rawMarkdown": "maybe [this notebook](https://www.kaggle.com/ronigur/landmark-retrieval-random-basline) will help",
      "votes": null
    },
    {
      "id": "1484851",
      "postDate": "08/21/2021 15:53:13",
      "content": "<p>Thank you so much!<br>\nSo, this notebook is run on Kaggle itself, or did you download the data?</p>\n<p>Also, I don't understand the file organization system here. I checked one folder, and the images are really different from each other.</p>\n<p>Lastly, how do I work while making the submission? Where do I get the landmark that I need to match, and what do I match it with?</p>\n<p>I'm really sorry for asking all this, I'm new to Kaggle, and there's a lot of ambiguity for me at least.</p>",
      "rawMarkdown": "Thank you so much!\nSo, this notebook is run on Kaggle itself, or did you download the data?\n\nAlso, I don't understand the file organization system here. I checked one folder, and the images are really different from each other.\n\nLastly, how do I work while making the submission? Where do I get the landmark that I need to match, and what do I match it with?\n\nI'm really sorry for asking all this, I'm new to Kaggle, and there's a lot of ambiguity for me at least.",
      "votes": null
    },
    {
      "id": "1488844",
      "postDate": "08/24/2021 14:31:01",
      "content": "<p>Test Images are your query images, which mean you'll use them to pick similar images in the index folder (match test images to its similar images in the index folder).</p>",
      "rawMarkdown": "Test Images are your query images, which mean you'll use them to pick similar images in the index folder (match test images to its similar images in the index folder).",
      "votes": null
    },
    {
      "id": "1489235",
      "postDate": "08/24/2021 20:04:56",
      "content": "<p>Images are organized by their ID. Images with similar IDs may not be similar.</p>\n<p>For instance, images with IDs '0120085011730a4f' and '01201a307cebe9d7' would be found in folder \"0\\1\\2\\\" where \"0\", \"1\" and \"2\" are the first three characters of their IDs. So basically the path to the images would be \"0\\1\\2\\0120085011730a4f.jpg\" and \"0\\1\\2\\01201a307cebe9d7.jpg\" respectively. </p>\n<p>These are all randomly generated IDs and hence images within the same folder are not necessarily related in any way.</p>\n<p>For iterating through the data, you can load the \"train.csv\" file as a dataframe that has all the IDs for the training images in the dataset. Then you can generate a \"filepath\" column in the dataframe using the logic explained above. Then just iterate over that column and use the values in the \"filepath\" column to load the images as required.</p>",
      "rawMarkdown": "Images are organized by their ID. Images with similar IDs may not be similar.\n\nFor instance, images with IDs '0120085011730a4f' and '01201a307cebe9d7' would be found in folder \"0\\1\\2\\\" where \"0\", \"1\" and \"2\" are the first three characters of their IDs. So basically the path to the images would be \"0\\1\\2\\0120085011730a4f.jpg\" and \"0\\1\\2\\01201a307cebe9d7.jpg\" respectively. \n\nThese are all randomly generated IDs and hence images within the same folder are not necessarily related in any way.\n\nFor iterating through the data, you can load the \"train.csv\" file as a dataframe that has all the IDs for the training images in the dataset. Then you can generate a \"filepath\" column in the dataframe using the logic explained above. Then just iterate over that column and use the values in the \"filepath\" column to load the images as required.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1484832,
      "author_name": "ronigur",
      "author_url": "",
      "post_date": "08/21/2021 15:39:50",
      "content": "<p>maybe <a href=\"https://www.kaggle.com/ronigur/landmark-retrieval-random-basline\" target=\"_blank\">this notebook</a> will help</p>",
      "votes": null,
      "replies": [
        {
          "id": 1484851,
          "author_name": "aschamp",
          "author_url": "",
          "post_date": "08/21/2021 15:53:13",
          "content": "<p>Thank you so much!<br>\nSo, this notebook is run on Kaggle itself, or did you download the data?</p>\n<p>Also, I don't understand the file organization system here. I checked one folder, and the images are really different from each other.</p>\n<p>Lastly, how do I work while making the submission? Where do I get the landmark that I need to match, and what do I match it with?</p>\n<p>I'm really sorry for asking all this, I'm new to Kaggle, and there's a lot of ambiguity for me at least.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1488844,
          "author_name": "tongkhangte",
          "author_url": "",
          "post_date": "08/24/2021 14:31:01",
          "content": "<p>Test Images are your query images, which mean you'll use them to pick similar images in the index folder (match test images to its similar images in the index folder).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1489235,
      "author_name": "mayur7garg",
      "author_url": "",
      "post_date": "08/24/2021 20:04:56",
      "content": "<p>Images are organized by their ID. Images with similar IDs may not be similar.</p>\n<p>For instance, images with IDs '0120085011730a4f' and '01201a307cebe9d7' would be found in folder \"0\\1\\2\\\" where \"0\", \"1\" and \"2\" are the first three characters of their IDs. So basically the path to the images would be \"0\\1\\2\\0120085011730a4f.jpg\" and \"0\\1\\2\\01201a307cebe9d7.jpg\" respectively. </p>\n<p>These are all randomly generated IDs and hence images within the same folder are not necessarily related in any way.</p>\n<p>For iterating through the data, you can load the \"train.csv\" file as a dataframe that has all the IDs for the training images in the dataset. Then you can generate a \"filepath\" column in the dataframe using the logic explained above. Then just iterate over that column and use the values in the \"filepath\" column to load the images as required.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1484722": "I had a few questions relating to the data. \n\n1. I don't exactly understand how it has been organized, and what each folder is for. I read that it's organized according to characteristics, but what does that exactly mean?\n\n2. Is there any way to iterate through the data during training without downloading it? I was asking because it's a very large dataset.",
    "1484832": "maybe [this notebook](https://www.kaggle.com/ronigur/landmark-retrieval-random-basline) will help",
    "1484851": "Thank you so much!\nSo, this notebook is run on Kaggle itself, or did you download the data?\n\nAlso, I don't understand the file organization system here. I checked one folder, and the images are really different from each other.\n\nLastly, how do I work while making the submission? Where do I get the landmark that I need to match, and what do I match it with?\n\nI'm really sorry for asking all this, I'm new to Kaggle, and there's a lot of ambiguity for me at least.",
    "1488844": "Test Images are your query images, which mean you'll use them to pick similar images in the index folder (match test images to its similar images in the index folder).",
    "1489235": "Images are organized by their ID. Images with similar IDs may not be similar.\n\nFor instance, images with IDs '0120085011730a4f' and '01201a307cebe9d7' would be found in folder \"0\\1\\2\\\" where \"0\", \"1\" and \"2\" are the first three characters of their IDs. So basically the path to the images would be \"0\\1\\2\\0120085011730a4f.jpg\" and \"0\\1\\2\\01201a307cebe9d7.jpg\" respectively. \n\nThese are all randomly generated IDs and hence images within the same folder are not necessarily related in any way.\n\nFor iterating through the data, you can load the \"train.csv\" file as a dataframe that has all the IDs for the training images in the dataset. Then you can generate a \"filepath\" column in the dataframe using the logic explained above. Then just iterate over that column and use the values in the \"filepath\" column to load the images as required."
  },
  "source": "meta"
}