{
  "id": 175343,
  "title": "Finetuning the baseline - silver medal solution [0.285 public LB]",
  "url": "/competitions/landmark-retrieval-2020/writeups/qiubit-finetuning-the-baseline-silver-medal-soluti",
  "author_name": "",
  "post_date": "2020-08-18T01:28:38.737Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Since the majority of teams just resubmitted a public model, it turned out that in this competition it was enough just to beat the baseline in order to receive a silver medal, which is what I did.</p>\n<p>My idea for doing that was finetuning a baseline model. It turned out it was much harder than it sounds, due to extensive computational resources required for the task! Here are some pointers on how I was able to train on the huge dataset provided by Google, while using limited computational resources and budget.</p>\n<p>In order to train a model, I decided to use TPU from Colab Pro (9.99$ / month) since it gives performance comparable to multiple GPUs (so allows for fast iteration), and I was unable to get GPU / TPU quota from Google Cloud.</p>\n<p>In order to train on the TPUs, the dataset has to be uploaded to Google Storage, preferably in TFRecord format. I used Google Compute Engine to download and prepare the dataset using scripts from <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training</a> I think it used around 6$ of Google's 300$ free trial.</p>\n<p>To sum up - you could beat the baseline for as little as 9.99$ (or even 0$ if you use Kaggle's TPU, but 30h/week was too stringent for me, although it would be more than enough to train the model), if you used TPUs for training.</p>\n<p>Some tricks I used in order to be able to train on TPUs:</p>\n<ul>\n<li>Reimplement model in Keras to get access to high-level APIs for training on TPUs. This required model rewrite in Keras and using this trick: <a href=\"https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model\" target=\"_blank\">https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model</a></li>\n<li>Only use a subset of 14000 classes (7000 largest and the rest uniformly sampled) out of ~81k available in gldv2_clean dataset, in order to cut the parameters of the ArcFace layer from ~165M to around ~29M. A necessity, since memory on TPU available in Colab is really limited (just 8GB).</li>\n<li>Only retrain a couple of final layers. I was unable to fit into memory of TPU when re-training entire model. Un-freezing consecutive layers gave subsequent public leaderboard improvements.</li>\n<li>Uploading TFRecords to multi-region US Google Storage bucket to avoid bills for Network egress from Google Cloud (all TPUs on Colab are currently located in US).</li>\n</ul>\n<p>Specific methods used during finetuning:</p>\n<ul>\n<li>Data augumentation.</li>\n<li>Train on 800x800 input (TPU input has to have fixed size for entirety of training). It was inspired by last year's solution of team smlyaka, which suggested larger image size for final epochs. The input size of 800x800 was chosen, because I discovered that this is max dimension of any image in gldv2_clean dataset.</li>\n<li>Cyclical learning rate, Learning rate finder: <a href=\"https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/\" target=\"_blank\">https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/</a> <a href=\"https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\" target=\"_blank\">https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/</a></li>\n<li>Ensemble with 0.277 baseline was used to increase score a little bit.</li>\n</ul>",
  "messages": [
    {
      "id": "974528",
      "postDate": "08/18/2020 01:13:49",
      "content": "<p>Since the majority of teams just resubmitted a public model, it turned out that in this competition it was enough just to beat the baseline in order to receive a silver medal, which is what I did.</p>\n<p>My idea for doing that was finetuning a baseline model. It turned out it was much harder than it sounds, due to extensive computational resources required for the task! Here are some pointers on how I was able to train on the huge dataset provided by Google, while using limited computational resources and budget.</p>\n<p>In order to train a model, I decided to use TPU from Colab Pro (9.99$ / month) since it gives performance comparable to multiple GPUs (so allows for fast iteration), and I was unable to get GPU / TPU quota from Google Cloud.</p>\n<p>In order to train on the TPUs, the dataset has to be uploaded to Google Storage, preferably in TFRecord format. I used Google Compute Engine to download and prepare the dataset using scripts from <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training</a> I think it used around 6$ of Google's 300$ free trial.</p>\n<p>To sum up - you could beat the baseline for as little as 9.99$ (or even 0$ if you use Kaggle's TPU, but 30h/week was too stringent for me, although it would be more than enough to train the model), if you used TPUs for training.</p>\n<p>Some tricks I used in order to be able to train on TPUs:</p>\n<ul>\n<li>Reimplement model in Keras to get access to high-level APIs for training on TPUs. This required model rewrite in Keras and using this trick: <a href=\"https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model\" target=\"_blank\">https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model</a></li>\n<li>Only use a subset of 14000 classes (7000 largest and the rest uniformly sampled) out of ~81k available in gldv2_clean dataset, in order to cut the parameters of the ArcFace layer from ~165M to around ~29M. A necessity, since memory on TPU available in Colab is really limited (just 8GB).</li>\n<li>Only retrain a couple of final layers. I was unable to fit into memory of TPU when re-training entire model. Un-freezing consecutive layers gave subsequent public leaderboard improvements.</li>\n<li>Uploading TFRecords to multi-region US Google Storage bucket to avoid bills for Network egress from Google Cloud (all TPUs on Colab are currently located in US).</li>\n</ul>\n<p>Specific methods used during finetuning:</p>\n<ul>\n<li>Data augumentation.</li>\n<li>Train on 800x800 input (TPU input has to have fixed size for entirety of training). It was inspired by last year's solution of team smlyaka, which suggested larger image size for final epochs. The input size of 800x800 was chosen, because I discovered that this is max dimension of any image in gldv2_clean dataset.</li>\n<li>Cyclical learning rate, Learning rate finder: <a href=\"https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/\" target=\"_blank\">https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/</a> <a href=\"https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\" target=\"_blank\">https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/</a></li>\n<li>Ensemble with 0.277 baseline was used to increase score a little bit.</li>\n</ul>",
      "rawMarkdown": "Since the majority of teams just resubmitted a public model, it turned out that in this competition it was enough just to beat the baseline in order to receive a silver medal, which is what I did.\n\nMy idea for doing that was finetuning a baseline model. It turned out it was much harder than it sounds, due to extensive computational resources required for the task! Here are some pointers on how I was able to train on the huge dataset provided by Google, while using limited computational resources and budget.\n\nIn order to train a model, I decided to use TPU from Colab Pro (9.99$ / month) since it gives performance comparable to multiple GPUs (so allows for fast iteration), and I was unable to get GPU / TPU quota from Google Cloud.\n\nIn order to train on the TPUs, the dataset has to be uploaded to Google Storage, preferably in TFRecord format. I used Google Compute Engine to download and prepare the dataset using scripts from https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training I think it used around 6$ of Google's 300$ free trial.\n\nTo sum up - you could beat the baseline for as little as 9.99$ (or even 0$ if you use Kaggle's TPU, but 30h/week was too stringent for me, although it would be more than enough to train the model), if you used TPUs for training.\n\nSome tricks I used in order to be able to train on TPUs:\n\n- Reimplement model in Keras to get access to high-level APIs for training on TPUs. This required model rewrite in Keras and using this trick: https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model\n- Only use a subset of 14000 classes (7000 largest and the rest uniformly sampled) out of ~81k available in gldv2_clean dataset, in order to cut the parameters of the ArcFace layer from ~165M to around ~29M. A necessity, since memory on TPU available in Colab is really limited (just 8GB).\n- Only retrain a couple of final layers. I was unable to fit into memory of TPU when re-training entire model. Un-freezing consecutive layers gave subsequent public leaderboard improvements.\n- Uploading TFRecords to multi-region US Google Storage bucket to avoid bills for Network egress from Google Cloud (all TPUs on Colab are currently located in US).\n\nSpecific methods used during finetuning:\n\n- Data augumentation.\n- Train on 800x800 input (TPU input has to have fixed size for entirety of training). It was inspired by last year's solution of team smlyaka, which suggested larger image size for final epochs. The input size of 800x800 was chosen, because I discovered that this is max dimension of any image in gldv2_clean dataset.\n- Cyclical learning rate, Learning rate finder: https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/ https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\n- Ensemble with 0.277 baseline was used to increase score a little bit.",
      "votes": null
    },
    {
      "id": "974593",
      "postDate": "08/18/2020 01:49:50",
      "content": "<p>Congrats to u</p>",
      "rawMarkdown": "Congrats to u",
      "votes": null
    },
    {
      "id": "974726",
      "postDate": "08/18/2020 03:09:20",
      "content": "<p>Hi, congrats on successfully converting the model. i think re implementing the whole model afresh is the breakthrough. you got it with roughly 14000 classes is amazing.  can you give me some more details of datasets and how did you sample them. how did u split them into query, positive and negative samples?</p>",
      "rawMarkdown": "Hi, congrats on successfully converting the model. i think re implementing the whole model afresh is the breakthrough. you got it with roughly 14000 classes is amazing.  can you give me some more details of datasets and how did you sample them. how did u split them into query, positive and negative samples?",
      "votes": null
    },
    {
      "id": "975875",
      "postDate": "08/18/2020 13:59:49",
      "content": "<p>Wow, Congratulations <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> 😍</p>",
      "rawMarkdown": "Wow, Congratulations @qiubit 😍",
      "votes": null
    },
    {
      "id": "976925",
      "postDate": "08/19/2020 07:18:05",
      "content": "<p>Congratz, well done!<br>\nA question, is it now possible to use Colab Pro outside US? I have been looking forward to it.</p>",
      "rawMarkdown": "Congratz, well done!\nA question, is it now possible to use Colab Pro outside US? I have been looking forward to it.",
      "votes": null
    },
    {
      "id": "1038594",
      "postDate": "10/05/2020 22:52:06",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> - any chance you're able to share code as well please? </p>\n<p>Congrats on a great finish.</p>",
      "rawMarkdown": "Thanks @qiubit - any chance you're able to share code as well please? \n\nCongrats on a great finish.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974593,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "08/18/2020 01:49:50",
      "content": "<p>Congrats to u</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974726,
      "author_name": "udaygurugubelli",
      "author_url": "",
      "post_date": "08/18/2020 03:09:20",
      "content": "<p>Hi, congrats on successfully converting the model. i think re implementing the whole model afresh is the breakthrough. you got it with roughly 14000 classes is amazing.  can you give me some more details of datasets and how did you sample them. how did u split them into query, positive and negative samples?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975875,
      "author_name": "rahim3",
      "author_url": "",
      "post_date": "08/18/2020 13:59:49",
      "content": "<p>Wow, Congratulations <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> 😍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976925,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "08/19/2020 07:18:05",
      "content": "<p>Congratz, well done!<br>\nA question, is it now possible to use Colab Pro outside US? I have been looking forward to it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1038594,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "10/05/2020 22:52:06",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> - any chance you're able to share code as well please? </p>\n<p>Congrats on a great finish.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974528": "Since the majority of teams just resubmitted a public model, it turned out that in this competition it was enough just to beat the baseline in order to receive a silver medal, which is what I did.\n\nMy idea for doing that was finetuning a baseline model. It turned out it was much harder than it sounds, due to extensive computational resources required for the task! Here are some pointers on how I was able to train on the huge dataset provided by Google, while using limited computational resources and budget.\n\nIn order to train a model, I decided to use TPU from Colab Pro (9.99$ / month) since it gives performance comparable to multiple GPUs (so allows for fast iteration), and I was unable to get GPU / TPU quota from Google Cloud.\n\nIn order to train on the TPUs, the dataset has to be uploaded to Google Storage, preferably in TFRecord format. I used Google Compute Engine to download and prepare the dataset using scripts from https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training I think it used around 6$ of Google's 300$ free trial.\n\nTo sum up - you could beat the baseline for as little as 9.99$ (or even 0$ if you use Kaggle's TPU, but 30h/week was too stringent for me, although it would be more than enough to train the model), if you used TPUs for training.\n\nSome tricks I used in order to be able to train on TPUs:\n\n- Reimplement model in Keras to get access to high-level APIs for training on TPUs. This required model rewrite in Keras and using this trick: https://www.kaggle.com/qiubit/fetching-weights-from-baseline-model\n- Only use a subset of 14000 classes (7000 largest and the rest uniformly sampled) out of ~81k available in gldv2_clean dataset, in order to cut the parameters of the ArcFace layer from ~165M to around ~29M. A necessity, since memory on TPU available in Colab is really limited (just 8GB).\n- Only retrain a couple of final layers. I was unable to fit into memory of TPU when re-training entire model. Un-freezing consecutive layers gave subsequent public leaderboard improvements.\n- Uploading TFRecords to multi-region US Google Storage bucket to avoid bills for Network egress from Google Cloud (all TPUs on Colab are currently located in US).\n\nSpecific methods used during finetuning:\n\n- Data augumentation.\n- Train on 800x800 input (TPU input has to have fixed size for entirety of training). It was inspired by last year's solution of team smlyaka, which suggested larger image size for final epochs. The input size of 800x800 was chosen, because I discovered that this is max dimension of any image in gldv2_clean dataset.\n- Cyclical learning rate, Learning rate finder: https://www.pyimagesearch.com/2019/07/29/cyclical-learning-rates-with-keras-and-deep-learning/ https://www.pyimagesearch.com/2019/08/05/keras-learning-rate-finder/\n- Ensemble with 0.277 baseline was used to increase score a little bit.",
    "974593": "Congrats to u",
    "974726": "Hi, congrats on successfully converting the model. i think re implementing the whole model afresh is the breakthrough. you got it with roughly 14000 classes is amazing.  can you give me some more details of datasets and how did you sample them. how did u split them into query, positive and negative samples?",
    "975875": "Wow, Congratulations @qiubit 😍",
    "976925": "Congratz, well done!\nA question, is it now possible to use Colab Pro outside US? I have been looking forward to it.",
    "1038594": "Thanks @qiubit - any chance you're able to share code as well please? \n\nCongrats on a great finish."
  },
  "source": "meta"
}