{
  "id": 175306,
  "title": "I beat baseline.277 with EFN on Keras. Limited time and resource stop me further.",
  "url": "/competitions/landmark-retrieval-2020/discussion/175306",
  "author_name": "wzm@",
  "post_date": "2020-08-17T23:31:04.424000",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>This competition is a research one with really high barrier for beginner.  Only 542 team attended it. Moreover, About 504 teams ( after rank-40) does not beat the official  baseline(.271-.277), while 437 teams copied baseline as their best submit.</p>\n<p>I am a pure beginner who joined the competition 10 days before ending. I first checked up the same competition held last year, read paper and learned from top rank solutions. Then I looked at kernels <a href=\"https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss\" target=\"_blank\">Google Landmark Retrieval - Triplet Loss</a> and <a href=\"https://www.kaggle.com/suruili/arcface-gem-train-on-tpu\" target=\"_blank\">ArcFace + GeM + Train on TPU</a>. I finally built my own kernel of ArcFace + GeM + different efficient-net with pure tf.keras.</p>\n<h3>My LB public score</h3>\n<p>I beat .277 baseline with ensemble B6-4round + B7-2round (4096 embed-size). I try to ensemble B5-2round too  (6144 embed-size) but it takes too much memory and failed.</p>\n<p><strong>Milestone of each model is listed in the following table:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Efficient-Net</th>\n<th>1round</th>\n<th>2round</th>\n<th>4round</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B5</td>\n<td>.196</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>B6</td>\n<td>.243</td>\n<td>.253</td>\n<td>.269</td>\n</tr>\n<tr>\n<td>B7</td>\n<td></td>\n<td>.269</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>The milestones show that:</p>\n<ol>\n<li>Large model is more capable for feature capture.</li>\n<li>These model needs more turns to converge.</li>\n</ol>\n<p><strong>I trained them on Kaggle and Google TPU. Time to train each for 1 round of whole dataset is following:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Efficient-Net</th>\n<th>Batch-size</th>\n<th>TPU version</th>\n<th>time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B5</td>\n<td>8 per core</td>\n<td>TPUv2-8 (free on Colab)</td>\n<td>5h+min</td>\n</tr>\n<tr>\n<td>B6</td>\n<td>32 per core</td>\n<td>TPUv3-8</td>\n<td>2h50+min (&lt; Kaggle 3h limit)</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>16 per core</td>\n<td>TPUv3-8</td>\n<td>4h20+min</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Since I attend late, I really do not have time to go further. Instead, I write a topic here to show my experience.</strong></p>\n<h3>Timeline and experience</h3>\n<p><strong>10d ago - 5d ago</strong>:<br>\n  Check Literature.</p>\n<p><strong>4d ago - 3d ago</strong>:<br>\n  Set up GCS, parsed the 105GB image files to tfrecord and store into  GCS.</p>\n<ul>\n<li>I manually set partitions and used all possible CPU kernels from Kaggle (limit 10) and Colab (limit 5) to parse in parallel.</li>\n<li>The time consumption mainly based on network latency rather than processing. Whether The kernel VM is initialed near to the GCS or not decide the running time.</li>\n<li>Based on the above line, some kernel can finish its job quickly while others may exceeded the 9 hour limit. Hence, I manually scheduled all the kernel, restarted it at stop point when necessary.</li>\n<li>It took 15 CPU to run more than one day to finish the parsing.</li>\n</ul>\n<p><strong>2d ago</strong>:<br>\nExperimented and Built my model kernel. Found the way to transform it to submission format.</p>\n<ul>\n<li>ArcFace layer need labels as input, which will break tf.keras training pipeline with tf.dataset. I modified it with idea from <a href=\"https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py\" target=\"_blank\">https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py</a>.</li>\n<li>To enable Keras model save, I transferred ArcFace layer from above link into separate layer, loss and metric class.</li>\n<li>At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that <strong>the dataset is a subset so the labels is also a subset</strong>. I overcome it with tf.lookup.statichashtable.</li>\n</ul>\n<p><strong>1d ago</strong>:<br>\nRun out of Kaggle and Colab free TPU quota, switch to GCP and continue training.</p>\n<ul>\n<li>I spent some time to follow GCP start tutorial, failed to enable Jupiter notebook from scratch. Then I found the \"Upgrade to Google Cloud AI notebook\" button which directly set all environment.</li>\n<li>I spent some time to find that tpu-name and zone are needed to link VM and TPU.</li>\n<li>I wasted 1.5h seeing that the same model config fits in Kaggle TPUv3-8 memory but exceeds GC-TPUv3-8 memory a lot. I searched and tried different methods, finally found that <strong>Tensorflow version 2.3.xx(dev) has bugs somewhere, switched it back to 2.2.0 worked</strong> . </li>\n<li>Among training, my network broke once and notebook was stopped. I tried to picked up as the exact place in dataset by dataset.skip(). However, after waste 1h quota, I found it does not start training for long time. So I have to retrain it by manually estimate the interrupted place in dataset.</li>\n<li>Among training, my notebook is weirdly stopped. I check console for TPU, it said to be in unknown condition. I spend some time to find it is stoped by scheduler and I had to switch from \"Preemptible\" to \"On-demand\" to continue training.</li>\n<li>I train B6 and B7 for 2 turns on GCP TPU. It cost <strong>$100</strong> from my free 300 credit.</li>\n</ul>\n<p><strong>This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments</strong></p>",
  "messages": [
    {
      "id": 974353,
      "postDate": "2020-08-17T23:31:04.423Z",
      "content": "<p>This competition is a research one with really high barrier for beginner.  Only 542 team attended it. Moreover, About 504 teams ( after rank-40) does not beat the official  baseline(.271-.277), while 437 teams copied baseline as their best submit.</p>\n<p>I am a pure beginner who joined the competition 10 days before ending. I first checked up the same competition held last year, read paper and learned from top rank solutions. Then I looked at kernels <a href=\"https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss\" target=\"_blank\">Google Landmark Retrieval - Triplet Loss</a> and <a href=\"https://www.kaggle.com/suruili/arcface-gem-train-on-tpu\" target=\"_blank\">ArcFace + GeM + Train on TPU</a>. I finally built my own kernel of ArcFace + GeM + different efficient-net with pure tf.keras.</p>\n<h3>My LB public score</h3>\n<p>I beat .277 baseline with ensemble B6-4round + B7-2round (4096 embed-size). I try to ensemble B5-2round too  (6144 embed-size) but it takes too much memory and failed.</p>\n<p><strong>Milestone of each model is listed in the following table:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Efficient-Net</th>\n<th>1round</th>\n<th>2round</th>\n<th>4round</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B5</td>\n<td>.196</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>B6</td>\n<td>.243</td>\n<td>.253</td>\n<td>.269</td>\n</tr>\n<tr>\n<td>B7</td>\n<td></td>\n<td>.269</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>The milestones show that:</p>\n<ol>\n<li>Large model is more capable for feature capture.</li>\n<li>These model needs more turns to converge.</li>\n</ol>\n<p><strong>I trained them on Kaggle and Google TPU. Time to train each for 1 round of whole dataset is following:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Efficient-Net</th>\n<th>Batch-size</th>\n<th>TPU version</th>\n<th>time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B5</td>\n<td>8 per core</td>\n<td>TPUv2-8 (free on Colab)</td>\n<td>5h+min</td>\n</tr>\n<tr>\n<td>B6</td>\n<td>32 per core</td>\n<td>TPUv3-8</td>\n<td>2h50+min (&lt; Kaggle 3h limit)</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>16 per core</td>\n<td>TPUv3-8</td>\n<td>4h20+min</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Since I attend late, I really do not have time to go further. Instead, I write a topic here to show my experience.</strong></p>\n<h3>Timeline and experience</h3>\n<p><strong>10d ago - 5d ago</strong>:<br>\n  Check Literature.</p>\n<p><strong>4d ago - 3d ago</strong>:<br>\n  Set up GCS, parsed the 105GB image files to tfrecord and store into  GCS.</p>\n<ul>\n<li>I manually set partitions and used all possible CPU kernels from Kaggle (limit 10) and Colab (limit 5) to parse in parallel.</li>\n<li>The time consumption mainly based on network latency rather than processing. Whether The kernel VM is initialed near to the GCS or not decide the running time.</li>\n<li>Based on the above line, some kernel can finish its job quickly while others may exceeded the 9 hour limit. Hence, I manually scheduled all the kernel, restarted it at stop point when necessary.</li>\n<li>It took 15 CPU to run more than one day to finish the parsing.</li>\n</ul>\n<p><strong>2d ago</strong>:<br>\nExperimented and Built my model kernel. Found the way to transform it to submission format.</p>\n<ul>\n<li>ArcFace layer need labels as input, which will break tf.keras training pipeline with tf.dataset. I modified it with idea from <a href=\"https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py\" target=\"_blank\">https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py</a>.</li>\n<li>To enable Keras model save, I transferred ArcFace layer from above link into separate layer, loss and metric class.</li>\n<li>At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that <strong>the dataset is a subset so the labels is also a subset</strong>. I overcome it with tf.lookup.statichashtable.</li>\n</ul>\n<p><strong>1d ago</strong>:<br>\nRun out of Kaggle and Colab free TPU quota, switch to GCP and continue training.</p>\n<ul>\n<li>I spent some time to follow GCP start tutorial, failed to enable Jupiter notebook from scratch. Then I found the \"Upgrade to Google Cloud AI notebook\" button which directly set all environment.</li>\n<li>I spent some time to find that tpu-name and zone are needed to link VM and TPU.</li>\n<li>I wasted 1.5h seeing that the same model config fits in Kaggle TPUv3-8 memory but exceeds GC-TPUv3-8 memory a lot. I searched and tried different methods, finally found that <strong>Tensorflow version 2.3.xx(dev) has bugs somewhere, switched it back to 2.2.0 worked</strong> . </li>\n<li>Among training, my network broke once and notebook was stopped. I tried to picked up as the exact place in dataset by dataset.skip(). However, after waste 1h quota, I found it does not start training for long time. So I have to retrain it by manually estimate the interrupted place in dataset.</li>\n<li>Among training, my notebook is weirdly stopped. I check console for TPU, it said to be in unknown condition. I spend some time to find it is stoped by scheduler and I had to switch from \"Preemptible\" to \"On-demand\" to continue training.</li>\n<li>I train B6 and B7 for 2 turns on GCP TPU. It cost <strong>$100</strong> from my free 300 credit.</li>\n</ul>\n<p><strong>This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments</strong></p>",
      "rawMarkdown": "This competition is a research one with really high barrier for beginner.  Only 542 team attended it. Moreover, About 504 teams ( after rank-40) does not beat the official  baseline(.271-.277), while 437 teams copied baseline as their best submit.\n\nI am a pure beginner who joined the competition 10 days before ending. I first checked up the same competition held last year, read paper and learned from top rank solutions. Then I looked at kernels [Google Landmark Retrieval - Triplet Loss](https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss) and [ArcFace + GeM + Train on TPU](https://www.kaggle.com/suruili/arcface-gem-train-on-tpu). I finally built my own kernel of ArcFace + GeM + different efficient-net with pure tf.keras.\n\n### My LB public score\nI beat .277 baseline with ensemble B6-4round + B7-2round (4096 embed-size). I try to ensemble B5-2round too  (6144 embed-size) but it takes too much memory and failed.\n\n**Milestone of each model is listed in the following table:**\n| Efficient-Net | 1round | 2round | 4round |\n| --- | --- | --- | --- |\n| B5  | .196 |  |  |\n| B6  | .243 | .253 | .269 |\n| B7  |  | .269 |  |\n\nThe milestones show that:\n1. Large model is more capable for feature capture.\n2. These model needs more turns to converge.\n\n**I trained them on Kaggle and Google TPU. Time to train each for 1 round of whole dataset is following:**\n|  Efficient-Net| Batch-size | TPU version |time |\n| --- | --- | --- | --- |\n| B5 | 8 per core | TPUv2-8 (free on Colab)  |  5h+min |\n| B6 | 32 per core | TPUv3-8 | 2h50+min (< Kaggle 3h limit)|\n| B7 | 16 per core | TPUv3-8 | 4h20+min |\n\n**Since I attend late, I really do not have time to go further. Instead, I write a topic here to show my experience.**\n\n### Timeline and experience\n\n**10d ago - 5d ago**:\n  Check Literature.\n\n**4d ago - 3d ago**:\n  Set up GCS, parsed the 105GB image files to tfrecord and store into  GCS.\n- I manually set partitions and used all possible CPU kernels from Kaggle (limit 10) and Colab (limit 5) to parse in parallel.\n- The time consumption mainly based on network latency rather than processing. Whether The kernel VM is initialed near to the GCS or not decide the running time.\n- Based on the above line, some kernel can finish its job quickly while others may exceeded the 9 hour limit. Hence, I manually scheduled all the kernel, restarted it at stop point when necessary.\n- It took 15 CPU to run more than one day to finish the parsing.\n\n**2d ago**:\nExperimented and Built my model kernel. Found the way to transform it to submission format.\n- ArcFace layer need labels as input, which will break tf.keras training pipeline with tf.dataset. I modified it with idea from https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py.\n- To enable Keras model save, I transferred ArcFace layer from above link into separate layer, loss and metric class.\n- At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that **the dataset is a subset so the labels is also a subset**. I overcome it with tf.lookup.statichashtable.\n\n**1d ago**:\nRun out of Kaggle and Colab free TPU quota, switch to GCP and continue training.\n- I spent some time to follow GCP start tutorial, failed to enable Jupiter notebook from scratch. Then I found the \"Upgrade to Google Cloud AI notebook\" button which directly set all environment.\n- I spent some time to find that tpu-name and zone are needed to link VM and TPU.\n- I wasted 1.5h seeing that the same model config fits in Kaggle TPUv3-8 memory but exceeds GC-TPUv3-8 memory a lot. I searched and tried different methods, finally found that **Tensorflow version 2.3.xx(dev) has bugs somewhere, switched it back to 2.2.0 worked** . \n- Among training, my network broke once and notebook was stopped. I tried to picked up as the exact place in dataset by dataset.skip(). However, after waste 1h quota, I found it does not start training for long time. So I have to retrain it by manually estimate the interrupted place in dataset.\n- Among training, my notebook is weirdly stopped. I check console for TPU, it said to be in unknown condition. I spend some time to find it is stoped by scheduler and I had to switch from \"Preemptible\" to \"On-demand\" to continue training.\n- I train B6 and B7 for 2 turns on GCP TPU. It cost **$100** from my free 300 credit.\n\n\n\n**This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments**",
      "votes": 15
    },
    {
      "id": 1015282,
      "postDate": "2020-09-18T05:18:29.503Z",
      "content": "<p>Thanks for warning about the Tensorflow 2.3 bug. I think I got stuck in the same situation, will try v2.2 now 😊</p>",
      "rawMarkdown": "Thanks for warning about the Tensorflow 2.3 bug. I think I got stuck in the same situation, will try v2.2 now 😊",
      "votes": 1
    },
    {
      "id": 1038591,
      "postDate": "2020-10-05T22:44:31.090Z",
      "content": "<p>Hey great work! And congrats!</p>\n<blockquote>\n  <p>This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments</p>\n</blockquote>\n<p>I am interested, can you please point me to the kernel? Thanks! :) </p>",
      "rawMarkdown": "Hey great work! And congrats!\n\n> This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments\n\nI am interested, can you please point me to the kernel? Thanks! :) "
    },
    {
      "id": 978858,
      "postDate": "2020-08-20T13:01:33.387Z",
      "content": "<p>You got a medal or not doesn't matter but work is really awesome!! to beat the lb by training your own model was really great!</p>",
      "rawMarkdown": "You got a medal or not doesn't matter but work is really awesome!! to beat the lb by training your own model was really great!"
    },
    {
      "id": 974712,
      "postDate": "2020-08-18T03:00:51.660Z",
      "content": "<p>HI, Thats a nice work. can u give some more details about data set. how did u sample and how did u split them into query, positive and negative samples. I tried with resent and my loss was also NAN but couldn't figured out the reason. train time was too slow takes around 5 days, i used triplet loss and it was not converging well. some times all descriptor values becoming zeros. thanks</p>",
      "rawMarkdown": "HI, Thats a nice work. can u give some more details about data set. how did u sample and how did u split them into query, positive and negative samples. I tried with resent and my loss was also NAN but couldn't figured out the reason. train time was too slow takes around 5 days, i used triplet loss and it was not converging well. some times all descriptor values becoming zeros. thanks",
      "replies": [
        {
          "id": 978615,
          "postDate": "2020-08-20T09:34:39.240Z",
          "content": "<p>Hi, Sorry that I just see your reply. I did not use triplet loss with Siamese training, as the top 2 place in 2019 competition and DELG both use ArcFace with X-entropy loss. </p>\n<ul>\n<li>For triplet loss, you can refer to the input dataset of <a href=\"https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss\" target=\"_blank\">Google Landmark Retrieval - Triplet Loss</a>, it has a preprocess kernel. </li>\n<li>For the NaN problem, maybe you met the same \"subset label\" problem as I mentioned above? </li>\n</ul>",
          "rawMarkdown": "Hi, Sorry that I just see your reply. I did not use triplet loss with Siamese training, as the top 2 place in 2019 competition and DELG both use ArcFace with X-entropy loss. \n- For triplet loss, you can refer to the input dataset of [Google Landmark Retrieval - Triplet Loss](https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss), it has a preprocess kernel. \n- For the NaN problem, maybe you met the same \"subset label\" problem as I mentioned above? "
        }
      ]
    },
    {
      "id": 974385,
      "postDate": "2020-08-18T00:02:39.327Z",
      "content": "<p>Thanks for your sharing. May I know whether you have done any cleaning on the dataset?</p>",
      "rawMarkdown": "Thanks for your sharing. May I know whether you have done any cleaning on the dataset?",
      "replies": [
        {
          "id": 974409,
          "postDate": "2020-08-18T00:13:35.900Z",
          "content": "<p>Thank you for reply. I did not do any dataset cleaning due to limit time. Actually, I did not split it into train-validate set either for the same reason.</p>\n<p>Dataset cleaning is used by 1st place last year but not used by the 2rd place. Plus, the dataset is said to be a cleaned version of the Google Landmarks Dataset v2 (GLDv2).  Hence, I am not sure whether data cleaning is really needed. </p>",
          "rawMarkdown": "Thank you for reply. I did not do any dataset cleaning due to limit time. Actually, I did not split it into train-validate set either for the same reason.\n\nDataset cleaning is used by 1st place last year but not used by the 2rd place. Plus, the dataset is said to be a cleaned version of the Google Landmarks Dataset v2 (GLDv2).  Hence, I am not sure whether data cleaning is really needed. ",
          "votes": 2
        },
        {
          "id": 974449,
          "postDate": "2020-08-18T00:27:15.217Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1015282,
      "author_name": "Chan Kha Vu",
      "author_url": "",
      "post_date": "2020-09-18T05:18:29.503000",
      "content": "<p>Thanks for warning about the Tensorflow 2.3 bug. I think I got stuck in the same situation, will try v2.2 now 😊</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1038591,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-10-05T22:44:31.090000",
      "content": "<p>Hey great work! And congrats!</p>\n<blockquote>\n  <p>This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments</p>\n</blockquote>\n<p>I am interested, can you please point me to the kernel? Thanks! :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 978858,
      "author_name": "Shivam",
      "author_url": "",
      "post_date": "2020-08-20T13:01:33.387000",
      "content": "<p>You got a medal or not doesn't matter but work is really awesome!! to beat the lb by training your own model was really great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974712,
      "author_name": "Uday Kumar Gurugubelli",
      "author_url": "",
      "post_date": "2020-08-18T03:00:51.660000",
      "content": "<p>HI, Thats a nice work. can u give some more details about data set. how did u sample and how did u split them into query, positive and negative samples. I tried with resent and my loss was also NAN but couldn't figured out the reason. train time was too slow takes around 5 days, i used triplet loss and it was not converging well. some times all descriptor values becoming zeros. thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 978615,
          "author_name": "wzm@",
          "author_url": "",
          "post_date": "2020-08-20T09:34:39.240000",
          "content": "<p>Hi, Sorry that I just see your reply. I did not use triplet loss with Siamese training, as the top 2 place in 2019 competition and DELG both use ArcFace with X-entropy loss. </p>\n<ul>\n<li>For triplet loss, you can refer to the input dataset of <a href=\"https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss\" target=\"_blank\">Google Landmark Retrieval - Triplet Loss</a>, it has a preprocess kernel. </li>\n<li>For the NaN problem, maybe you met the same \"subset label\" problem as I mentioned above? </li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974385,
      "author_name": "FP",
      "author_url": "",
      "post_date": "2020-08-18T00:02:39.327000",
      "content": "<p>Thanks for your sharing. May I know whether you have done any cleaning on the dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 974409,
          "author_name": "wzm@",
          "author_url": "",
          "post_date": "2020-08-18T00:13:35.900000",
          "content": "<p>Thank you for reply. I did not do any dataset cleaning due to limit time. Actually, I did not split it into train-validate set either for the same reason.</p>\n<p>Dataset cleaning is used by 1st place last year but not used by the 2rd place. Plus, the dataset is said to be a cleaned version of the Google Landmarks Dataset v2 (GLDv2).  Hence, I am not sure whether data cleaning is really needed. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 974449,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T00:27:15.217000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974353": "This competition is a research one with really high barrier for beginner.  Only 542 team attended it. Moreover, About 504 teams ( after rank-40) does not beat the official  baseline(.271-.277), while 437 teams copied baseline as their best submit.\n\nI am a pure beginner who joined the competition 10 days before ending. I first checked up the same competition held last year, read paper and learned from top rank solutions. Then I looked at kernels [Google Landmark Retrieval - Triplet Loss](https://www.kaggle.com/mattbast/google-landmark-retrieval-triplet-loss) and [ArcFace + GeM + Train on TPU](https://www.kaggle.com/suruili/arcface-gem-train-on-tpu). I finally built my own kernel of ArcFace + GeM + different efficient-net with pure tf.keras.\n\n### My LB public score\nI beat .277 baseline with ensemble B6-4round + B7-2round (4096 embed-size). I try to ensemble B5-2round too  (6144 embed-size) but it takes too much memory and failed.\n\n**Milestone of each model is listed in the following table:**\n| Efficient-Net | 1round | 2round | 4round |\n| --- | --- | --- | --- |\n| B5  | .196 |  |  |\n| B6  | .243 | .253 | .269 |\n| B7  |  | .269 |  |\n\nThe milestones show that:\n1. Large model is more capable for feature capture.\n2. These model needs more turns to converge.\n\n**I trained them on Kaggle and Google TPU. Time to train each for 1 round of whole dataset is following:**\n|  Efficient-Net| Batch-size | TPU version |time |\n| --- | --- | --- | --- |\n| B5 | 8 per core | TPUv2-8 (free on Colab)  |  5h+min |\n| B6 | 32 per core | TPUv3-8 | 2h50+min (< Kaggle 3h limit)|\n| B7 | 16 per core | TPUv3-8 | 4h20+min |\n\n**Since I attend late, I really do not have time to go further. Instead, I write a topic here to show my experience.**\n\n### Timeline and experience\n\n**10d ago - 5d ago**:\n  Check Literature.\n\n**4d ago - 3d ago**:\n  Set up GCS, parsed the 105GB image files to tfrecord and store into  GCS.\n- I manually set partitions and used all possible CPU kernels from Kaggle (limit 10) and Colab (limit 5) to parse in parallel.\n- The time consumption mainly based on network latency rather than processing. Whether The kernel VM is initialed near to the GCS or not decide the running time.\n- Based on the above line, some kernel can finish its job quickly while others may exceeded the 9 hour limit. Hence, I manually scheduled all the kernel, restarted it at stop point when necessary.\n- It took 15 CPU to run more than one day to finish the parsing.\n\n**2d ago**:\nExperimented and Built my model kernel. Found the way to transform it to submission format.\n- ArcFace layer need labels as input, which will break tf.keras training pipeline with tf.dataset. I modified it with idea from https://github.com/ktjonsson/keras-ArcFace/blob/master/src/arcface_layer.py.\n- To enable Keras model save, I transferred ArcFace layer from above link into separate layer, loss and metric class.\n- At first several rounds of training, I found it always stop at half round with NAN. I wasted half a day and about 10h TPU quota to find that **the dataset is a subset so the labels is also a subset**. I overcome it with tf.lookup.statichashtable.\n\n**1d ago**:\nRun out of Kaggle and Colab free TPU quota, switch to GCP and continue training.\n- I spent some time to follow GCP start tutorial, failed to enable Jupiter notebook from scratch. Then I found the \"Upgrade to Google Cloud AI notebook\" button which directly set all environment.\n- I spent some time to find that tpu-name and zone are needed to link VM and TPU.\n- I wasted 1.5h seeing that the same model config fits in Kaggle TPUv3-8 memory but exceeds GC-TPUv3-8 memory a lot. I searched and tried different methods, finally found that **Tensorflow version 2.3.xx(dev) has bugs somewhere, switched it back to 2.2.0 worked** . \n- Among training, my network broke once and notebook was stopped. I tried to picked up as the exact place in dataset by dataset.skip(). However, after waste 1h quota, I found it does not start training for long time. So I have to retrain it by manually estimate the interrupted place in dataset.\n- Among training, my notebook is weirdly stopped. I check console for TPU, it said to be in unknown condition. I spend some time to find it is stoped by scheduler and I had to switch from \"Preemptible\" to \"On-demand\" to continue training.\n- I train B6 and B7 for 2 turns on GCP TPU. It cost **$100** from my free 300 credit.\n\n\n\n**This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments**",
    "1015282": "Thanks for warning about the Tensorflow 2.3 bug. I think I got stuck in the same situation, will try v2.2 now 😊",
    "1038591": "Hey great work! And congrats!\n\n> This is all for my experience in this competition. Hope it is helpful for beginners. If anyone is interested in specific code, I will make my kernel public and put links in comments\n\nI am interested, can you please point me to the kernel? Thanks! :) ",
    "978858": "You got a medal or not doesn't matter but work is really awesome!! to beat the lb by training your own model was really great!",
    "974712": "HI, Thats a nice work. can u give some more details about data set. how did u sample and how did u split them into query, positive and negative samples. I tried with resent and my loss was also NAN but couldn't figured out the reason. train time was too slow takes around 5 days, i used triplet loss and it was not converging well. some times all descriptor values becoming zeros. thanks",
    "974385": "Thanks for your sharing. May I know whether you have done any cleaning on the dataset?"
  }
}