{
  "id": 166354,
  "title": "How a large-scale image retrieval system works?",
  "url": "/competitions/landmark-retrieval-2020/discussion/166354",
  "author_name": "",
  "post_date": "2020-07-12T17:05:23.829159100Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>A large-scale retrieval system can be decomposed into the following 4 parts:</p>\n\n<p>1# Dense localized feature extraction\n2# Keypoint selection\n3# Dimensionality Reduction\n4# Indexing &amp; Retrieving</p>\n\n<p><strong>DELF: DEnse Local Feature Extraction</strong></p>\n\n<ul>\n<li>Dense features are extracted from an image by aplying a Fully Convolutional Network (FCN)</li>\n<li><p>To handle scale changes, an image pyramid is constructed and the FCN is applied to each level independently </p></li>\n<li><p>The FCN is constructed by using the feature extraction layers of a CNN trained with a classification loss (eg. triplet loss)</p></li>\n<li>The FCN is taken from a ResNet50 model (pre-trained on ImageNet) </li>\n<li><p>Output of the conv4x_ convolutional block is used</p></li>\n<li><p>An annotated database of landmark images is used to train a network with standard cross-entropy loss for image classification </p></li>\n<li>Input image are centre cropped to 250x250 and then random 224x224 crops are used for training </li>\n<li>Thus local descriptors implicitly learn representations that are more relevant for landmark retrieval problem</li>\n</ul>\n\n<p><strong>Attention based Keypoint Selection</strong></p>\n\n<ul>\n<li>Instead of using the features directly; only a subset is used for image retrievals</li>\n<li>This is important for accuracy &amp; computational efficiency of the process</li>\n<li>Learning with weak supervision</li>\n<li>Training the attention model      </li>\n</ul>\n\n<p><strong>Dimensionality Reduction</strong></p>\n\n<ul>\n<li>Dimensionality of the selected features is reduced to 40 using PCA. </li>\n<li>This gives a good trade-off between compactness &amp; discriminativeness</li>\n<li>Features are again L2 normalized </li>\n</ul>\n\n<p><strong>Image Retrieval System</strong></p>\n\n<ul>\n<li>We extract feature descriptors from query &amp; DB images, where a pre-defined number of local features with the highest attention scores per image are selected </li>\n<li>Image retrieval is based on nearest neighbour search ( acombination of KD-tree &amp; Product Quantization)</li>\n<li>A kd tree means a k-dimensionality of tree space</li>\n<li>When a query is fired, we perform an approximate nearest neighbour search for each local descriptor extracted from the query image</li>\n<li>For the top K nearest local decriptors retrieved from the index, all the matches per db image are aggregated</li>\n<li>Then a geometric verification using RANSAC is employed to find the number of inliers. The number of inliers is then used to score an image</li>\n</ul>\n\n<p>Reference:\n<code>\nLarge Scale Image Retrieval with Attentive Deep Local Features - \nHyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, Bohyung Han\nICCV 2017\n</code></p>",
  "messages": [
    {
      "id": "926370",
      "postDate": "07/12/2020 17:05:23",
      "content": "<p>A large-scale retrieval system can be decomposed into the following 4 parts:</p>\n\n<p>1# Dense localized feature extraction\n2# Keypoint selection\n3# Dimensionality Reduction\n4# Indexing &amp; Retrieving</p>\n\n<p><strong>DELF: DEnse Local Feature Extraction</strong></p>\n\n<ul>\n<li>Dense features are extracted from an image by aplying a Fully Convolutional Network (FCN)</li>\n<li><p>To handle scale changes, an image pyramid is constructed and the FCN is applied to each level independently </p></li>\n<li><p>The FCN is constructed by using the feature extraction layers of a CNN trained with a classification loss (eg. triplet loss)</p></li>\n<li>The FCN is taken from a ResNet50 model (pre-trained on ImageNet) </li>\n<li><p>Output of the conv4x_ convolutional block is used</p></li>\n<li><p>An annotated database of landmark images is used to train a network with standard cross-entropy loss for image classification </p></li>\n<li>Input image are centre cropped to 250x250 and then random 224x224 crops are used for training </li>\n<li>Thus local descriptors implicitly learn representations that are more relevant for landmark retrieval problem</li>\n</ul>\n\n<p><strong>Attention based Keypoint Selection</strong></p>\n\n<ul>\n<li>Instead of using the features directly; only a subset is used for image retrievals</li>\n<li>This is important for accuracy &amp; computational efficiency of the process</li>\n<li>Learning with weak supervision</li>\n<li>Training the attention model      </li>\n</ul>\n\n<p><strong>Dimensionality Reduction</strong></p>\n\n<ul>\n<li>Dimensionality of the selected features is reduced to 40 using PCA. </li>\n<li>This gives a good trade-off between compactness &amp; discriminativeness</li>\n<li>Features are again L2 normalized </li>\n</ul>\n\n<p><strong>Image Retrieval System</strong></p>\n\n<ul>\n<li>We extract feature descriptors from query &amp; DB images, where a pre-defined number of local features with the highest attention scores per image are selected </li>\n<li>Image retrieval is based on nearest neighbour search ( acombination of KD-tree &amp; Product Quantization)</li>\n<li>A kd tree means a k-dimensionality of tree space</li>\n<li>When a query is fired, we perform an approximate nearest neighbour search for each local descriptor extracted from the query image</li>\n<li>For the top K nearest local decriptors retrieved from the index, all the matches per db image are aggregated</li>\n<li>Then a geometric verification using RANSAC is employed to find the number of inliers. The number of inliers is then used to score an image</li>\n</ul>\n\n<p>Reference:\n<code>\nLarge Scale Image Retrieval with Attentive Deep Local Features - \nHyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, Bohyung Han\nICCV 2017\n</code></p>",
      "rawMarkdown": "A large-scale retrieval system can be decomposed into the following 4 parts:\n\n1# Dense localized feature extraction\n2# Keypoint selection\n3# Dimensionality Reduction\n4# Indexing &amp; Retrieving\n\n**DELF: DEnse Local Feature Extraction**\n\n- Dense features are extracted from an image by aplying a Fully Convolutional Network (FCN)\n- To handle scale changes, an image pyramid is constructed and the FCN is applied to each level independently \n\n- The FCN is constructed by using the feature extraction layers of a CNN trained with a classification loss (eg. triplet loss)\n- The FCN is taken from a ResNet50 model (pre-trained on ImageNet) \n- Output of the conv4x_ convolutional block is used\n\n- An annotated database of landmark images is used to train a network with standard cross-entropy loss for image classification \n- Input image are centre cropped to 250x250 and then random 224x224 crops are used for training \n- Thus local descriptors implicitly learn representations that are more relevant for landmark retrieval problem\n\n**Attention based Keypoint Selection**\n\n- Instead of using the features directly; only a subset is used for image retrievals\n- This is important for accuracy &amp; computational efficiency of the process\n- Learning with weak supervision\n- Training the attention model  \t\n\n**Dimensionality Reduction**\n\n- Dimensionality of the selected features is reduced to 40 using PCA. \n- This gives a good trade-off between compactness &amp; discriminativeness\n- Features are again L2 normalized \n\n**Image Retrieval System**\n\n- We extract feature descriptors from query &amp; DB images, where a pre-defined number of local features with the highest attention scores per image are selected \n- Image retrieval is based on nearest neighbour search ( acombination of KD-tree &amp; Product Quantization)\n- A kd tree means a k-dimensionality of tree space\n- When a query is fired, we perform an approximate nearest neighbour search for each local descriptor extracted from the query image\n- For the top K nearest local decriptors retrieved from the index, all the matches per db image are aggregated\n- Then a geometric verification using RANSAC is employed to find the number of inliers. The number of inliers is then used to score an image\n\nReference:\n```\nLarge Scale Image Retrieval with Attentive Deep Local Features - \nHyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, Bohyung Han\nICCV 2017\n```",
      "votes": null
    },
    {
      "id": "965895",
      "postDate": "08/11/2020 00:20:45",
      "content": "<p>Nice one !!!</p>",
      "rawMarkdown": "Nice one !!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 965895,
      "author_name": "sudarshanpatil",
      "author_url": "",
      "post_date": "08/11/2020 00:20:45",
      "content": "<p>Nice one !!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "926370": "A large-scale retrieval system can be decomposed into the following 4 parts:\n\n1# Dense localized feature extraction\n2# Keypoint selection\n3# Dimensionality Reduction\n4# Indexing &amp; Retrieving\n\n**DELF: DEnse Local Feature Extraction**\n\n- Dense features are extracted from an image by aplying a Fully Convolutional Network (FCN)\n- To handle scale changes, an image pyramid is constructed and the FCN is applied to each level independently \n\n- The FCN is constructed by using the feature extraction layers of a CNN trained with a classification loss (eg. triplet loss)\n- The FCN is taken from a ResNet50 model (pre-trained on ImageNet) \n- Output of the conv4x_ convolutional block is used\n\n- An annotated database of landmark images is used to train a network with standard cross-entropy loss for image classification \n- Input image are centre cropped to 250x250 and then random 224x224 crops are used for training \n- Thus local descriptors implicitly learn representations that are more relevant for landmark retrieval problem\n\n**Attention based Keypoint Selection**\n\n- Instead of using the features directly; only a subset is used for image retrievals\n- This is important for accuracy &amp; computational efficiency of the process\n- Learning with weak supervision\n- Training the attention model  \t\n\n**Dimensionality Reduction**\n\n- Dimensionality of the selected features is reduced to 40 using PCA. \n- This gives a good trade-off between compactness &amp; discriminativeness\n- Features are again L2 normalized \n\n**Image Retrieval System**\n\n- We extract feature descriptors from query &amp; DB images, where a pre-defined number of local features with the highest attention scores per image are selected \n- Image retrieval is based on nearest neighbour search ( acombination of KD-tree &amp; Product Quantization)\n- A kd tree means a k-dimensionality of tree space\n- When a query is fired, we perform an approximate nearest neighbour search for each local descriptor extracted from the query image\n- For the top K nearest local decriptors retrieved from the index, all the matches per db image are aggregated\n- Then a geometric verification using RANSAC is employed to find the number of inliers. The number of inliers is then used to score an image\n\nReference:\n```\nLarge Scale Image Retrieval with Attentive Deep Local Features - \nHyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, Bohyung Han\nICCV 2017\n```",
    "965895": "Nice one !!!"
  },
  "source": "meta"
}