{
  "id": 265639,
  "title": "Summarising 2020 1st rank solution - Google Landmark Retrieval ",
  "url": "/competitions/landmark-retrieval-2021/discussion/265639",
  "author_name": "Vishnu Subramanian",
  "post_date": "2021-08-16T11:35:39.492000",
  "votes": 23,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Past kaggle solutions provide great insights into how we can approach the present problem.  Thanks to <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> for his <a href=\"https://arxiv.org/abs/2009.05132\" target=\"_blank\">paper</a> on last year's winning solution. This post is my understanding of the paper and I want to try a similar approach to this year's competition. Some of the key highlights of the solution are</p>\n<ol>\n<li>Data </li>\n<li>Metric Learning</li>\n<li>Transfer learning</li>\n<li>Progressive resizing</li>\n<li>Adjusting weights</li>\n<li>Ensembling </li>\n</ol>\n<h2>Data</h2>\n<p>2 datasets were primarily used. </p>\n<h3>CGLD2</h3>\n<p>This is a cleaned dataset containing 81313 landmark classes. </p>\n<h3>GLD2</h3>\n<p>This dataset contains 203094 classes and contains way more noise. CGLD2 is a subset of this dataset. </p>\n<h2>Metric Learning:</h2>\n<p>Metric learning helps </p>\n<ul>\n<li>when the number of classes is very high </li>\n<li>when the class imbalance is high. </li>\n</ul>\n<p>So my guess is this year's solution may also include metric learning. Some of the key resources I found for understanding metric learning. If you have any good resources, please share them here.</p>\n<ul>\n<li>Keras <a href=\"https://keras.io/examples/vision/metric_learning/\" target=\"_blank\">tutorial</a> </li>\n<li>Using Cross entropy for <a href=\"https://www.youtube.com/watch?v=Jb4Ewl5RzkI&amp;t=184s\" target=\"_blank\">metric learning</a> </li>\n</ul>\n<h2>Transfer learning:</h2>\n<p>The base model (EfficientNet 7) was trained on CGLD2 containing 81313 landmark classes. This model itself was able to rank in the top 5 of the private LB. This base model is then fine-tuned on GLD2 containing 203094. Which helped in improving the score further. It is also mentioned that the model performed worse when trained from scratch on the noisy GLD2 dataset. </p>\n<h2>Progressive resizing:</h2>\n<p>After finetuning the model on the GLD2 dataset with images of size 512<em>512, the model is fine-tuned with larger image sizes of 640</em>640 and 736*736.  </p>\n<h2>Adjusting weights:</h2>\n<p>The model was finetuned with 2 times the loss for the data from CGLD2. </p>\n<h2>Ensemble:</h2>\n<p>Several Efficientnet models were used for ensembling with different weightage based on the score achieved in public LB. </p>",
  "messages": [
    {
      "id": 1474982,
      "postDate": "2021-08-16T11:35:39.493Z",
      "content": "<p>Past kaggle solutions provide great insights into how we can approach the present problem.  Thanks to <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a> for his <a href=\"https://arxiv.org/abs/2009.05132\" target=\"_blank\">paper</a> on last year's winning solution. This post is my understanding of the paper and I want to try a similar approach to this year's competition. Some of the key highlights of the solution are</p>\n<ol>\n<li>Data </li>\n<li>Metric Learning</li>\n<li>Transfer learning</li>\n<li>Progressive resizing</li>\n<li>Adjusting weights</li>\n<li>Ensembling </li>\n</ol>\n<h2>Data</h2>\n<p>2 datasets were primarily used. </p>\n<h3>CGLD2</h3>\n<p>This is a cleaned dataset containing 81313 landmark classes. </p>\n<h3>GLD2</h3>\n<p>This dataset contains 203094 classes and contains way more noise. CGLD2 is a subset of this dataset. </p>\n<h2>Metric Learning:</h2>\n<p>Metric learning helps </p>\n<ul>\n<li>when the number of classes is very high </li>\n<li>when the class imbalance is high. </li>\n</ul>\n<p>So my guess is this year's solution may also include metric learning. Some of the key resources I found for understanding metric learning. If you have any good resources, please share them here.</p>\n<ul>\n<li>Keras <a href=\"https://keras.io/examples/vision/metric_learning/\" target=\"_blank\">tutorial</a> </li>\n<li>Using Cross entropy for <a href=\"https://www.youtube.com/watch?v=Jb4Ewl5RzkI&amp;t=184s\" target=\"_blank\">metric learning</a> </li>\n</ul>\n<h2>Transfer learning:</h2>\n<p>The base model (EfficientNet 7) was trained on CGLD2 containing 81313 landmark classes. This model itself was able to rank in the top 5 of the private LB. This base model is then fine-tuned on GLD2 containing 203094. Which helped in improving the score further. It is also mentioned that the model performed worse when trained from scratch on the noisy GLD2 dataset. </p>\n<h2>Progressive resizing:</h2>\n<p>After finetuning the model on the GLD2 dataset with images of size 512<em>512, the model is fine-tuned with larger image sizes of 640</em>640 and 736*736.  </p>\n<h2>Adjusting weights:</h2>\n<p>The model was finetuned with 2 times the loss for the data from CGLD2. </p>\n<h2>Ensemble:</h2>\n<p>Several Efficientnet models were used for ensembling with different weightage based on the score achieved in public LB. </p>",
      "rawMarkdown": "Past kaggle solutions provide great insights into how we can approach the present problem.  Thanks to @keetar for his [paper](https://arxiv.org/abs/2009.05132) on last year's winning solution. This post is my understanding of the paper and I want to try a similar approach to this year's competition. Some of the key highlights of the solution are\n\n1. Data \n2. Metric Learning\n3. Transfer learning\n4. Progressive resizing\n5. Adjusting weights\n6. Ensembling \n\n\n## Data\n2 datasets were primarily used. \n\n### CGLD2\nThis is a cleaned dataset containing 81313 landmark classes. \n\n### GLD2\nThis dataset contains 203094 classes and contains way more noise. CGLD2 is a subset of this dataset. \n\n## Metric Learning: \nMetric learning helps \n- when the number of classes is very high \n- when the class imbalance is high. \n\nSo my guess is this year's solution may also include metric learning. Some of the key resources I found for understanding metric learning. If you have any good resources, please share them here.\n\n- Keras [tutorial](https://keras.io/examples/vision/metric_learning/) \n- Using Cross entropy for [metric learning](https://www.youtube.com/watch?v=Jb4Ewl5RzkI&t=184s) \n\n## Transfer learning: \nThe base model (EfficientNet 7) was trained on CGLD2 containing 81313 landmark classes. This model itself was able to rank in the top 5 of the private LB. This base model is then fine-tuned on GLD2 containing 203094. Which helped in improving the score further. It is also mentioned that the model performed worse when trained from scratch on the noisy GLD2 dataset. \n\n## Progressive resizing:\nAfter finetuning the model on the GLD2 dataset with images of size 512*512, the model is fine-tuned with larger image sizes of 640*640 and 736*736.  \n\n## Adjusting weights:\nThe model was finetuned with 2 times the loss for the data from CGLD2. \n\n## Ensemble: \nSeveral Efficientnet models were used for ensembling with different weightage based on the score achieved in public LB. \n\n\n\n\n\n\n",
      "votes": 24
    },
    {
      "id": 1494246,
      "postDate": "2021-08-28T14:21:06.407Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/vishnus\" target=\"_blank\">@vishnus</a> ,</p>\n<ol>\n<li>You refer here term: <strong>noisy GLD2 dataset</strong>. Can you please tell me what noisy means here? </li>\n<li>Do you have some notebooks/code on Progressive resizing approach? I Would love to try it. </li>\n<li>Adjusting weights: Sorry, I did not get this,  <strong>2 times the loss</strong> ? What training approach will be? Psuedo code will be fine. </li>\n</ol>",
      "rawMarkdown": "Hey @vishnus ,\n1. You refer here term: **noisy GLD2 dataset**. Can you please tell me what noisy means here? \n2. Do you have some notebooks/code on Progressive resizing approach? I Would love to try it. \n3. Adjusting weights: Sorry, I did not get this,  **2 times the loss** ? What training approach will be? Psuedo code will be fine. ",
      "votes": 1,
      "replies": [
        {
          "id": 1494843,
          "postDate": "2021-08-29T05:56:21.363Z",
          "content": "<ol>\n<li>Noisy means the labels need not be accurate. Let's say an image labeled as TajMahal need not be necessarily the Taj Mahal Monument. </li>\n<li>The idea is simple, you train your model with images of size say 224. After the model gets better, you can train the entire pipeline with images of larger size say 512 but this time use model weights from your older pipeline. </li>\n<li>2 * loss_fn(images from CGLD2) + 1 * loss_fn(images from GLD2)</li>\n</ol>",
          "rawMarkdown": "1. Noisy means the labels need not be accurate. Let's say an image labeled as TajMahal need not be necessarily the Taj Mahal Monument. \n2. The idea is simple, you train your model with images of size say 224. After the model gets better, you can train the entire pipeline with images of larger size say 512 but this time use model weights from your older pipeline. \n3. 2 \\* loss_fn(images from CGLD2) + 1 \\* loss_fn(images from GLD2)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1475042,
      "postDate": "2021-08-16T12:16:49.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1494246,
      "author_name": "Saurav Solanki",
      "author_url": "",
      "post_date": "2021-08-28T14:21:06.407000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/vishnus\" target=\"_blank\">@vishnus</a> ,</p>\n<ol>\n<li>You refer here term: <strong>noisy GLD2 dataset</strong>. Can you please tell me what noisy means here? </li>\n<li>Do you have some notebooks/code on Progressive resizing approach? I Would love to try it. </li>\n<li>Adjusting weights: Sorry, I did not get this,  <strong>2 times the loss</strong> ? What training approach will be? Psuedo code will be fine. </li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1494843,
          "author_name": "Vishnu Subramanian",
          "author_url": "",
          "post_date": "2021-08-29T05:56:21.363000",
          "content": "<ol>\n<li>Noisy means the labels need not be accurate. Let's say an image labeled as TajMahal need not be necessarily the Taj Mahal Monument. </li>\n<li>The idea is simple, you train your model with images of size say 224. After the model gets better, you can train the entire pipeline with images of larger size say 512 but this time use model weights from your older pipeline. </li>\n<li>2 * loss_fn(images from CGLD2) + 1 * loss_fn(images from GLD2)</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1475042,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-16T12:16:49.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1474982": "Past kaggle solutions provide great insights into how we can approach the present problem.  Thanks to @keetar for his [paper](https://arxiv.org/abs/2009.05132) on last year's winning solution. This post is my understanding of the paper and I want to try a similar approach to this year's competition. Some of the key highlights of the solution are\n\n1. Data \n2. Metric Learning\n3. Transfer learning\n4. Progressive resizing\n5. Adjusting weights\n6. Ensembling \n\n\n## Data\n2 datasets were primarily used. \n\n### CGLD2\nThis is a cleaned dataset containing 81313 landmark classes. \n\n### GLD2\nThis dataset contains 203094 classes and contains way more noise. CGLD2 is a subset of this dataset. \n\n## Metric Learning: \nMetric learning helps \n- when the number of classes is very high \n- when the class imbalance is high. \n\nSo my guess is this year's solution may also include metric learning. Some of the key resources I found for understanding metric learning. If you have any good resources, please share them here.\n\n- Keras [tutorial](https://keras.io/examples/vision/metric_learning/) \n- Using Cross entropy for [metric learning](https://www.youtube.com/watch?v=Jb4Ewl5RzkI&t=184s) \n\n## Transfer learning: \nThe base model (EfficientNet 7) was trained on CGLD2 containing 81313 landmark classes. This model itself was able to rank in the top 5 of the private LB. This base model is then fine-tuned on GLD2 containing 203094. Which helped in improving the score further. It is also mentioned that the model performed worse when trained from scratch on the noisy GLD2 dataset. \n\n## Progressive resizing:\nAfter finetuning the model on the GLD2 dataset with images of size 512*512, the model is fine-tuned with larger image sizes of 640*640 and 736*736.  \n\n## Adjusting weights:\nThe model was finetuned with 2 times the loss for the data from CGLD2. \n\n## Ensemble: \nSeveral Efficientnet models were used for ensembling with different weightage based on the score achieved in public LB. \n\n\n\n\n\n\n",
    "1494246": "Hey @vishnus ,\n1. You refer here term: **noisy GLD2 dataset**. Can you please tell me what noisy means here? \n2. Do you have some notebooks/code on Progressive resizing approach? I Would love to try it. \n3. Adjusting weights: Sorry, I did not get this,  **2 times the loss** ? What training approach will be? Psuedo code will be fine. ",
    "1475042": ""
  }
}