{
  "id": 165643,
  "title": "Some Questions about Understanding Data",
  "url": "/competitions/landmark-retrieval-2020/discussion/165643",
  "author_name": "",
  "post_date": "2020-07-10T13:02:16.483677Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am new to retrieval problem.\nWhat do index images represent? Like we have to match test images with index?\nWhat do train images represent? And what do landmark ID represent? </p>",
  "messages": [
    {
      "id": "922966",
      "postDate": "07/10/2020 13:02:16",
      "content": "<p>I am new to retrieval problem.\nWhat do index images represent? Like we have to match test images with index?\nWhat do train images represent? And what do landmark ID represent? </p>",
      "rawMarkdown": "I am new to retrieval problem.\nWhat do index images represent? Like we have to match test images with index?\nWhat do train images represent? And what do landmark ID represent?",
      "votes": null
    },
    {
      "id": "923421",
      "postDate": "07/10/2020 20:30:54",
      "content": "<p>This is part of the large-scale image retrieval &amp; recognition problem. We are tackling the issue of retrieval here! Central to this problem is to use representations to describe the image &amp; its similarities </p>\n\n<p>So if you query the image retrieval system with the following <code>Test</code> image - \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F1114807b86d02ffd37bdb5b7193afdfc%2Fpage07_07.jpg?generation=1594412795038471&amp;alt=media\" alt=\"\"></p>\n\n<p>The system should search its <code>index set</code>and find the following image\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fea3f5694a0c99007535ac455fb9d2818%2Fpage07_08.jpg?generation=1594412860944704&amp;alt=media\" alt=\"\"></p>\n\n<p>It does this by applying a fully convolutional network (FCN) to the image and extracting dense features. These are then compared with the index set. </p>\n\n<p>To answer your questions -</p>\n\n<ul>\n<li><code>Test</code>images are your query image. The system will use dense features to find similar images in the <code>index</code>set </li>\n<li>The Fully connected CNN that's used to extract the dense features is first trained with the images in the <code>training set</code></li>\n</ul>",
      "rawMarkdown": "This is part of the large-scale image retrieval &amp; recognition problem. We are tackling the issue of retrieval here! Central to this problem is to use representations to describe the image &amp; its similarities \n\nSo if you query the image retrieval system with the following `Test` image - \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F1114807b86d02ffd37bdb5b7193afdfc%2Fpage07_07.jpg?generation=1594412795038471&amp;alt=media)\n\nThe system should search its `index set `and find the following image\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fea3f5694a0c99007535ac455fb9d2818%2Fpage07_08.jpg?generation=1594412860944704&amp;alt=media)\n\nIt does this by applying a fully convolutional network (FCN) to the image and extracting dense features. These are then compared with the index set. \n\nTo answer your questions -\n\n- `Test `images are your query image. The system will use dense features to find similar images in the `index `set \n-  The Fully connected CNN that's used to extract the dense features is first trained with the images in the `training set `",
      "votes": null
    },
    {
      "id": "923427",
      "postDate": "07/10/2020 20:45:26",
      "content": "<p>Exactly. But that image at submission time will be passes by dividing to 255.0 or not?</p>",
      "rawMarkdown": "Exactly. But that image at submission time will be passes by dividing to 255.0 or not?",
      "votes": null
    },
    {
      "id": "923431",
      "postDate": "07/10/2020 20:47:55",
      "content": "<p>And how submission script will resize image?</p>",
      "rawMarkdown": "And how submission script will resize image?",
      "votes": null
    },
    {
      "id": "924982",
      "postDate": "07/11/2020 18:46:37",
      "content": "<p>All preprocessing should be done inside your model.\nCheck this example: <a href=\"https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model\">https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model</a>\nCtrl+F : Resize</p>",
      "rawMarkdown": "All preprocessing should be done inside your model.\nCheck this example: https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model\nCtrl+F : Resize",
      "votes": null
    },
    {
      "id": "925875",
      "postDate": "07/12/2020 11:05:35",
      "content": "<p><a href=\"/skylord\">@skylord</a> you mean we will get the features from the <strong>'Test Image'</strong> and extract features from all <strong>'index images</strong>' and find the similarity in some hyperspace right? so each time when e get a test image we will have to loop through all the index images to find the similar one. Is my understanding correct ?? \nif we have to loop through all the index Images every time isn't it gonna be very slow? \nThanks in advance. </p>",
      "rawMarkdown": "skylord you mean we will get the features from the **'Test Image'** and extract features from all **'index images**' and find the similarity in some hyperspace right? so each time when e get a test image we will have to loop through all the index images to find the similar one. Is my understanding correct ?? \nif we have to loop through all the index Images every time isn't it gonna be very slow? \nThanks in advance.",
      "votes": null
    },
    {
      "id": "926433",
      "postDate": "07/12/2020 18:05:40",
      "content": "<p>Large scale image retrievals generally work in 4 stages: \n1. Dense local &amp; global feature extraction from query image \n2. Keypoint selection, so that only a subset of the features are used \n3. Dimensionality reduction &amp; normalisation\n4. Image indexing &amp; retrieval</p>\n\n<p>I have summarised it in this <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166354\">post</a> - </p>\n\n<p>While I am still figuring out how it works in this competition, you can significantly reduce the computation requirements by indexing the images. \nThe first pass is done with the global features and for the k nearest neighbours you can then search using the local features.</p>",
      "rawMarkdown": "Large scale image retrievals generally work in 4 stages: \n1. Dense local &amp; global feature extraction from query image \n2. Keypoint selection, so that only a subset of the features are used \n3. Dimensionality reduction &amp; normalisation\n4. Image indexing &amp; retrieval\n\nI have summarised it in this [post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166354) - \n\nWhile I am still figuring out how it works in this competition, you can significantly reduce the computation requirements by indexing the images. \nThe first pass is done with the global features and for the k nearest neighbours you can then search using the local features.",
      "votes": null
    },
    {
      "id": "928162",
      "postDate": "07/13/2020 19:21:25",
      "content": "<p><a href=\"/skylord\">@skylord</a> Why is landmark ID provided? We are given a query image in the test. We extract its features and find all images from the index set with similar features. Where is landmark ID in this equation? </p>",
      "rawMarkdown": "skylord Why is landmark ID provided? We are given a query image in the test. We extract its features and find all images from the index set with similar features. Where is landmark ID in this equation?",
      "votes": null
    },
    {
      "id": "930930",
      "postDate": "07/15/2020 20:21:52",
      "content": "<p>Sorry for the late response! I don't think its useful for training the model. </p>\n\n<p>There are 5mn images and 200k+ landmarks in the GLDv2 dataset (each landmark has its own landmark id! ) Apparently the top three landmarks are churches, parks &amp; museums ( <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/164769\">Summary of GLDv2 paper</a>)</p>\n\n<p>If you have an image-retrieval system in production (like a landmark recognition app or a tourist place recommend er) you <em>could</em> use the id in your workflow. </p>",
      "rawMarkdown": "Sorry for the late response! I don't think its useful for training the model. \n\nThere are 5mn images and 200k+ landmarks in the GLDv2 dataset (each landmark has its own landmark id! ) Apparently the top three landmarks are churches, parks &amp; museums ( [Summary of GLDv2 paper](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/164769))\n\nIf you have an image-retrieval system in production (like a landmark recognition app or a tourist place recommend er) you *could* use the id in your workflow.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 923421,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "07/10/2020 20:30:54",
      "content": "<p>This is part of the large-scale image retrieval &amp; recognition problem. We are tackling the issue of retrieval here! Central to this problem is to use representations to describe the image &amp; its similarities </p>\n\n<p>So if you query the image retrieval system with the following <code>Test</code> image - \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F1114807b86d02ffd37bdb5b7193afdfc%2Fpage07_07.jpg?generation=1594412795038471&amp;alt=media\" alt=\"\"></p>\n\n<p>The system should search its <code>index set</code>and find the following image\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fea3f5694a0c99007535ac455fb9d2818%2Fpage07_08.jpg?generation=1594412860944704&amp;alt=media\" alt=\"\"></p>\n\n<p>It does this by applying a fully convolutional network (FCN) to the image and extracting dense features. These are then compared with the index set. </p>\n\n<p>To answer your questions -</p>\n\n<ul>\n<li><code>Test</code>images are your query image. The system will use dense features to find similar images in the <code>index</code>set </li>\n<li>The Fully connected CNN that's used to extract the dense features is first trained with the images in the <code>training set</code></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 923427,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "07/10/2020 20:45:26",
          "content": "<p>Exactly. But that image at submission time will be passes by dividing to 255.0 or not?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 923431,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "07/10/2020 20:47:55",
          "content": "<p>And how submission script will resize image?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924982,
          "author_name": "macarrony00",
          "author_url": "",
          "post_date": "07/11/2020 18:46:37",
          "content": "<p>All preprocessing should be done inside your model.\nCheck this example: <a href=\"https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model\">https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model</a>\nCtrl+F : Resize</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 925875,
          "author_name": "aliraza786",
          "author_url": "",
          "post_date": "07/12/2020 11:05:35",
          "content": "<p><a href=\"/skylord\">@skylord</a> you mean we will get the features from the <strong>'Test Image'</strong> and extract features from all <strong>'index images</strong>' and find the similarity in some hyperspace right? so each time when e get a test image we will have to loop through all the index images to find the similar one. Is my understanding correct ?? \nif we have to loop through all the index Images every time isn't it gonna be very slow? \nThanks in advance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926433,
          "author_name": "skylord",
          "author_url": "",
          "post_date": "07/12/2020 18:05:40",
          "content": "<p>Large scale image retrievals generally work in 4 stages: \n1. Dense local &amp; global feature extraction from query image \n2. Keypoint selection, so that only a subset of the features are used \n3. Dimensionality reduction &amp; normalisation\n4. Image indexing &amp; retrieval</p>\n\n<p>I have summarised it in this <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166354\">post</a> - </p>\n\n<p>While I am still figuring out how it works in this competition, you can significantly reduce the computation requirements by indexing the images. \nThe first pass is done with the global features and for the k nearest neighbours you can then search using the local features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928162,
          "author_name": "srinathvj",
          "author_url": "",
          "post_date": "07/13/2020 19:21:25",
          "content": "<p><a href=\"/skylord\">@skylord</a> Why is landmark ID provided? We are given a query image in the test. We extract its features and find all images from the index set with similar features. Where is landmark ID in this equation? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930930,
          "author_name": "skylord",
          "author_url": "",
          "post_date": "07/15/2020 20:21:52",
          "content": "<p>Sorry for the late response! I don't think its useful for training the model. </p>\n\n<p>There are 5mn images and 200k+ landmarks in the GLDv2 dataset (each landmark has its own landmark id! ) Apparently the top three landmarks are churches, parks &amp; museums ( <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/164769\">Summary of GLDv2 paper</a>)</p>\n\n<p>If you have an image-retrieval system in production (like a landmark recognition app or a tourist place recommend er) you <em>could</em> use the id in your workflow. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "922966": "I am new to retrieval problem.\nWhat do index images represent? Like we have to match test images with index?\nWhat do train images represent? And what do landmark ID represent?",
    "923421": "This is part of the large-scale image retrieval &amp; recognition problem. We are tackling the issue of retrieval here! Central to this problem is to use representations to describe the image &amp; its similarities \n\nSo if you query the image retrieval system with the following `Test` image - \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F1114807b86d02ffd37bdb5b7193afdfc%2Fpage07_07.jpg?generation=1594412795038471&amp;alt=media)\n\nThe system should search its `index set `and find the following image\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fea3f5694a0c99007535ac455fb9d2818%2Fpage07_08.jpg?generation=1594412860944704&amp;alt=media)\n\nIt does this by applying a fully convolutional network (FCN) to the image and extracting dense features. These are then compared with the index set. \n\nTo answer your questions -\n\n- `Test `images are your query image. The system will use dense features to find similar images in the `index `set \n-  The Fully connected CNN that's used to extract the dense features is first trained with the images in the `training set `",
    "923427": "Exactly. But that image at submission time will be passes by dividing to 255.0 or not?",
    "923431": "And how submission script will resize image?",
    "924982": "All preprocessing should be done inside your model.\nCheck this example: https://www.kaggle.com/mayukh18/creating-submission-from-your-own-model\nCtrl+F : Resize",
    "925875": "skylord you mean we will get the features from the **'Test Image'** and extract features from all **'index images**' and find the similarity in some hyperspace right? so each time when e get a test image we will have to loop through all the index images to find the similar one. Is my understanding correct ?? \nif we have to loop through all the index Images every time isn't it gonna be very slow? \nThanks in advance.",
    "926433": "Large scale image retrievals generally work in 4 stages: \n1. Dense local &amp; global feature extraction from query image \n2. Keypoint selection, so that only a subset of the features are used \n3. Dimensionality reduction &amp; normalisation\n4. Image indexing &amp; retrieval\n\nI have summarised it in this [post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166354) - \n\nWhile I am still figuring out how it works in this competition, you can significantly reduce the computation requirements by indexing the images. \nThe first pass is done with the global features and for the k nearest neighbours you can then search using the local features.",
    "928162": "skylord Why is landmark ID provided? We are given a query image in the test. We extract its features and find all images from the index set with similar features. Where is landmark ID in this equation?",
    "930930": "Sorry for the late response! I don't think its useful for training the model. \n\nThere are 5mn images and 200k+ landmarks in the GLDv2 dataset (each landmark has its own landmark id! ) Apparently the top three landmarks are churches, parks &amp; museums ( [Summary of GLDv2 paper](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/164769))\n\nIf you have an image-retrieval system in production (like a landmark recognition app or a tourist place recommend er) you *could* use the id in your workflow."
  },
  "source": "meta"
}