{
  "id": 110941,
  "title": "TPU solution: RetinaNet + NAS-FPN + ResNet/EfficientNet, code shared [24th place]",
  "url": "/competitions/open-images-2019-object-detection/writeups/ods-ai-512-tpus-tpu-solution-retinanet-nas-fpn-res",
  "author_name": "",
  "post_date": "2019-10-04T09:46:22.490Z",
  "votes": 41,
  "comment_count": 13,
  "views": 0,
  "content": "<p>TL;DR: TPUs actually work. They are fast (10-100 times of 1080Ti). There are limitations that make porting difficult. My fixed code and models are at  <a href=\"https://github.com/artyompal/tpu_models\">https://github.com/artyompal/tpu_models</a>.</p>\n\n<hr>\n\n<h1>TPU</h1>\n\n<p>I started this competition quite late. I had 1 month, which is not enough to train a lot of models of GPU. I had some TPU credits, so I decided to give it a try.</p>\n\n<p>I forked TPU models repo and trained RetinaNet + FPN + ResNet50. I trained the model like in this tutorial: <a href=\"https://cloud.google.com/tpu/docs/tutorials/retinanet\">https://cloud.google.com/tpu/docs/tutorials/retinanet</a> and it actually worked. To achieve this, I had to convert my data into TFRecords and also convert my validation data to a COCO JSON file. I realized that our metric (AP@50) is just an ordinary COCO/PASCAL VOC AP@50 metric, as long as there is no hierarchy. So training just worked out of the box.</p>\n\n<p>What's TPU? It's a matrix chip. It's a stand-alone device in Google's local network which connects to your cloud instance as TensorFlow master. It has a lot of HBM memory (I used batch 64 with images up to 1024). It has no HDD and it can't connect to your SSD, so it must use Google Cloud Storage for input (images and labels) and output (weights and TensorBoard logs). I had v2x8 and v3x8 TPUs; 3rd generation is ~50% faster.</p>\n\n<p>RetinaNet+ResNet50, batch size 64, images 640x640 processes 250 samples per second on TPUv2. On TPUv3, it processes 350 samples per second:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F0373ce32c12fbd7aa0b29ba5ee9b288b%2Fperformance.png?generation=1570020130759200&amp;alt=media\" alt=\"\"></p>\n\n<p>One issue about TPUs is, they currently require a completely static TF graph. So you have to resize every image to a fixed resolution. It's a limitation of the current implementation (TensorFlow 1.14). Another issue is, image channels are swapped from NHWC to HWCN. Another thing, it doesn't support some TF operations. And you would need special pretrained weight for TPU, normal TF models will not work.</p>\n\n<p>Oh yeah, the main issue is, one has to use TPUEstimator to train models. tf.Estimators are weird pieces of software which by default erase all of your checkpoints but last five, completely ignoring your validation results. Every Deep Learning framework supports validation score-based ReduceLROnPlateau, but not tf.Estimators. If only we had PyTorch/XLA working... (I didn't try it)</p>\n\n<h1>Training</h1>\n\n<p>Naive training on 443 leaf classes resulted in 0.283 on the public leaderboard. Following last year solutions, I split dataset classes into 6 parts by frequency:\n1. Classes 1-100\n2. Classes 101-200\n3. Classes 201-300\n4. Classes 301-400\n5. Classes 401-432\n6. Human parts (11 classes), as they have a special dataset.</p>\n\n<p>I trained 6 models by parts and got a baseline which resulted in 0.48 on the public leaderboard. We teamed up with Dmitry, who had collected 5 pretrained models from different sources and got a script with Soft-NMS blending. This gave us 0.54 on the public LB. I also trained this model by 5 folds, which resulted in 0.51 LB by itself and 0.55 in the ensemble. I did it because I wanted to do pseudo-labeling for some data cleaning, but ran out of time to finish this.</p>\n\n<h1>Models</h1>\n\n<p>TPU models repo has out-of-the-box support of RetinaNet with NAS-FPN (<a href=\"https://arxiv.org/pdf/1904.07392.pdf\">https://arxiv.org/pdf/1904.07392.pdf</a>) with AutoAugment augmentations (<a href=\"https://arxiv.org/pdf/1805.09501.pdf\">https://arxiv.org/pdf/1805.09501.pdf</a>) and DropBlock regularization for both FPN part and backbone (<a href=\"https://arxiv.org/pdf/1810.12890.pdf\">https://arxiv.org/pdf/1810.12890.pdf</a>). By default, RetinaNet only supports ResNet backbone, but I added support of EfficientNet. I also tried adding SE-ResNeXt backbone, but it worked too slow, maybe because I didn't transpose channels to HWCN. Also, I recall that I had a problem with TPUs not supporting REFLECT padding: <a href=\"https://cloud.google.com/tpu/docs/tensorflow-ops\">https://cloud.google.com/tpu/docs/tensorflow-ops</a>.</p>\n\n<p>Of the ResNet family, I only had pretrained weights for ResNet50. Papers above claim that with some combinations of AutoAugment and DropBlock, RetinaNet trains even better without a pretrained backbone! This wasn't the case for me, though. Anyway, I had a lot of horsepowers,  so I downloaded ImageNet and made my own pretrained models for ResNet101, ResNet152, and ResNet200.</p>\n\n<p>Long story short: NAS-FPN is great, DropBlock really helps, and AutoAugment augmentations do improve model performance.</p>\n\n<h1>Data labeling</h1>\n\n<p>Data labels are really noisy. I manually selected 5 labels per class for the validation set. I made a blacklist of labels which are clearly wrong: i.e. the label is \"cello\", but it's actually a violin.</p>\n\n<h1>Last day inference and final ensemble construction</h1>\n\n<p>We had seriously run out of time. I had to write and fix the distributed inference script in the last two days of the competition. This could not end well, and it didn't. I trained models with ResNet152/200 backbones, but we couldn't make a proper ensemble.</p>\n\n<p>Oh yeah, the last night before the deadline I got an error with inference with EfficientNet models: <a href=\"https://github.com/tensorflow/tensorflow/issues/29574\">https://github.com/tensorflow/tensorflow/issues/29574</a>. So we could not use any of my EfficientNet models in the final ensemble. I probably could work around this issue by hacking TPUEstimator, but... time is money, you know :)</p>\n\n<h1>Conclusions</h1>\n\n<p>Well, it was fun! TPUs are fun. As long as someone else is paying. We spent about $1600 worth of GCP credits, and I paid $75 out of my pocket on top of that.</p>\n\n<p>I'm sharing my forked TPU repo here: <a href=\"https://github.com/artyompal/tpu_models\">https://github.com/artyompal/tpu_models</a>. I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.</p>\n\n<hr>\n\n<p>Thanks to <a href=\"/dsvolkov\">@dsvolkov</a> for being my teammate and for all that help with ensembling, debugging and diagnostics!</p>",
  "messages": [
    {
      "id": "638790",
      "postDate": "10/02/2019 12:30:27",
      "content": "<p>TL;DR: TPUs actually work. They are fast (10-100 times of 1080Ti). There are limitations that make porting difficult. My fixed code and models are at  <a href=\"https://github.com/artyompal/tpu_models\">https://github.com/artyompal/tpu_models</a>.</p>\n\n<hr>\n\n<h1>TPU</h1>\n\n<p>I started this competition quite late. I had 1 month, which is not enough to train a lot of models of GPU. I had some TPU credits, so I decided to give it a try.</p>\n\n<p>I forked TPU models repo and trained RetinaNet + FPN + ResNet50. I trained the model like in this tutorial: <a href=\"https://cloud.google.com/tpu/docs/tutorials/retinanet\">https://cloud.google.com/tpu/docs/tutorials/retinanet</a> and it actually worked. To achieve this, I had to convert my data into TFRecords and also convert my validation data to a COCO JSON file. I realized that our metric (AP@50) is just an ordinary COCO/PASCAL VOC AP@50 metric, as long as there is no hierarchy. So training just worked out of the box.</p>\n\n<p>What's TPU? It's a matrix chip. It's a stand-alone device in Google's local network which connects to your cloud instance as TensorFlow master. It has a lot of HBM memory (I used batch 64 with images up to 1024). It has no HDD and it can't connect to your SSD, so it must use Google Cloud Storage for input (images and labels) and output (weights and TensorBoard logs). I had v2x8 and v3x8 TPUs; 3rd generation is ~50% faster.</p>\n\n<p>RetinaNet+ResNet50, batch size 64, images 640x640 processes 250 samples per second on TPUv2. On TPUv3, it processes 350 samples per second:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F0373ce32c12fbd7aa0b29ba5ee9b288b%2Fperformance.png?generation=1570020130759200&amp;alt=media\" alt=\"\"></p>\n\n<p>One issue about TPUs is, they currently require a completely static TF graph. So you have to resize every image to a fixed resolution. It's a limitation of the current implementation (TensorFlow 1.14). Another issue is, image channels are swapped from NHWC to HWCN. Another thing, it doesn't support some TF operations. And you would need special pretrained weight for TPU, normal TF models will not work.</p>\n\n<p>Oh yeah, the main issue is, one has to use TPUEstimator to train models. tf.Estimators are weird pieces of software which by default erase all of your checkpoints but last five, completely ignoring your validation results. Every Deep Learning framework supports validation score-based ReduceLROnPlateau, but not tf.Estimators. If only we had PyTorch/XLA working... (I didn't try it)</p>\n\n<h1>Training</h1>\n\n<p>Naive training on 443 leaf classes resulted in 0.283 on the public leaderboard. Following last year solutions, I split dataset classes into 6 parts by frequency:\n1. Classes 1-100\n2. Classes 101-200\n3. Classes 201-300\n4. Classes 301-400\n5. Classes 401-432\n6. Human parts (11 classes), as they have a special dataset.</p>\n\n<p>I trained 6 models by parts and got a baseline which resulted in 0.48 on the public leaderboard. We teamed up with Dmitry, who had collected 5 pretrained models from different sources and got a script with Soft-NMS blending. This gave us 0.54 on the public LB. I also trained this model by 5 folds, which resulted in 0.51 LB by itself and 0.55 in the ensemble. I did it because I wanted to do pseudo-labeling for some data cleaning, but ran out of time to finish this.</p>\n\n<h1>Models</h1>\n\n<p>TPU models repo has out-of-the-box support of RetinaNet with NAS-FPN (<a href=\"https://arxiv.org/pdf/1904.07392.pdf\">https://arxiv.org/pdf/1904.07392.pdf</a>) with AutoAugment augmentations (<a href=\"https://arxiv.org/pdf/1805.09501.pdf\">https://arxiv.org/pdf/1805.09501.pdf</a>) and DropBlock regularization for both FPN part and backbone (<a href=\"https://arxiv.org/pdf/1810.12890.pdf\">https://arxiv.org/pdf/1810.12890.pdf</a>). By default, RetinaNet only supports ResNet backbone, but I added support of EfficientNet. I also tried adding SE-ResNeXt backbone, but it worked too slow, maybe because I didn't transpose channels to HWCN. Also, I recall that I had a problem with TPUs not supporting REFLECT padding: <a href=\"https://cloud.google.com/tpu/docs/tensorflow-ops\">https://cloud.google.com/tpu/docs/tensorflow-ops</a>.</p>\n\n<p>Of the ResNet family, I only had pretrained weights for ResNet50. Papers above claim that with some combinations of AutoAugment and DropBlock, RetinaNet trains even better without a pretrained backbone! This wasn't the case for me, though. Anyway, I had a lot of horsepowers,  so I downloaded ImageNet and made my own pretrained models for ResNet101, ResNet152, and ResNet200.</p>\n\n<p>Long story short: NAS-FPN is great, DropBlock really helps, and AutoAugment augmentations do improve model performance.</p>\n\n<h1>Data labeling</h1>\n\n<p>Data labels are really noisy. I manually selected 5 labels per class for the validation set. I made a blacklist of labels which are clearly wrong: i.e. the label is \"cello\", but it's actually a violin.</p>\n\n<h1>Last day inference and final ensemble construction</h1>\n\n<p>We had seriously run out of time. I had to write and fix the distributed inference script in the last two days of the competition. This could not end well, and it didn't. I trained models with ResNet152/200 backbones, but we couldn't make a proper ensemble.</p>\n\n<p>Oh yeah, the last night before the deadline I got an error with inference with EfficientNet models: <a href=\"https://github.com/tensorflow/tensorflow/issues/29574\">https://github.com/tensorflow/tensorflow/issues/29574</a>. So we could not use any of my EfficientNet models in the final ensemble. I probably could work around this issue by hacking TPUEstimator, but... time is money, you know :)</p>\n\n<h1>Conclusions</h1>\n\n<p>Well, it was fun! TPUs are fun. As long as someone else is paying. We spent about $1600 worth of GCP credits, and I paid $75 out of my pocket on top of that.</p>\n\n<p>I'm sharing my forked TPU repo here: <a href=\"https://github.com/artyompal/tpu_models\">https://github.com/artyompal/tpu_models</a>. I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.</p>\n\n<hr>\n\n<p>Thanks to <a href=\"/dsvolkov\">@dsvolkov</a> for being my teammate and for all that help with ensembling, debugging and diagnostics!</p>",
      "rawMarkdown": "TL;DR: TPUs actually work. They are fast (10-100 times of 1080Ti). There are limitations that make porting difficult. My fixed code and models are at  https://github.com/artyompal/tpu_models.\n____________\n# TPU\nI started this competition quite late. I had 1 month, which is not enough to train a lot of models of GPU. I had some TPU credits, so I decided to give it a try.\n\nI forked TPU models repo and trained RetinaNet + FPN + ResNet50. I trained the model like in this tutorial: https://cloud.google.com/tpu/docs/tutorials/retinanet and it actually worked. To achieve this, I had to convert my data into TFRecords and also convert my validation data to a COCO JSON file. I realized that our metric (AP@50) is just an ordinary COCO/PASCAL VOC AP@50 metric, as long as there is no hierarchy. So training just worked out of the box.\n\nWhat's TPU? It's a matrix chip. It's a stand-alone device in Google's local network which connects to your cloud instance as TensorFlow master. It has a lot of HBM memory (I used batch 64 with images up to 1024). It has no HDD and it can't connect to your SSD, so it must use Google Cloud Storage for input (images and labels) and output (weights and TensorBoard logs). I had v2x8 and v3x8 TPUs; 3rd generation is ~50% faster.\n\nRetinaNet+ResNet50, batch size 64, images 640x640 processes 250 samples per second on TPUv2. On TPUv3, it processes 350 samples per second:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F0373ce32c12fbd7aa0b29ba5ee9b288b%2Fperformance.png?generation=1570020130759200&amp;alt=media)\n\nOne issue about TPUs is, they currently require a completely static TF graph. So you have to resize every image to a fixed resolution. It's a limitation of the current implementation (TensorFlow 1.14). Another issue is, image channels are swapped from NHWC to HWCN. Another thing, it doesn't support some TF operations. And you would need special pretrained weight for TPU, normal TF models will not work.\n\nOh yeah, the main issue is, one has to use TPUEstimator to train models. tf.Estimators are weird pieces of software which by default erase all of your checkpoints but last five, completely ignoring your validation results. Every Deep Learning framework supports validation score-based ReduceLROnPlateau, but not tf.Estimators. If only we had PyTorch/XLA working... (I didn't try it)\n\n# Training\nNaive training on 443 leaf classes resulted in 0.283 on the public leaderboard. Following last year solutions, I split dataset classes into 6 parts by frequency:\n1. Classes 1-100\n2. Classes 101-200\n3. Classes 201-300\n4. Classes 301-400\n5. Classes 401-432\n6. Human parts (11 classes), as they have a special dataset.\n\nI trained 6 models by parts and got a baseline which resulted in 0.48 on the public leaderboard. We teamed up with Dmitry, who had collected 5 pretrained models from different sources and got a script with Soft-NMS blending. This gave us 0.54 on the public LB. I also trained this model by 5 folds, which resulted in 0.51 LB by itself and 0.55 in the ensemble. I did it because I wanted to do pseudo-labeling for some data cleaning, but ran out of time to finish this.\n\n# Models\nTPU models repo has out-of-the-box support of RetinaNet with NAS-FPN (https://arxiv.org/pdf/1904.07392.pdf) with AutoAugment augmentations (https://arxiv.org/pdf/1805.09501.pdf) and DropBlock regularization for both FPN part and backbone (https://arxiv.org/pdf/1810.12890.pdf). By default, RetinaNet only supports ResNet backbone, but I added support of EfficientNet. I also tried adding SE-ResNeXt backbone, but it worked too slow, maybe because I didn't transpose channels to HWCN. Also, I recall that I had a problem with TPUs not supporting REFLECT padding: https://cloud.google.com/tpu/docs/tensorflow-ops.\n\nOf the ResNet family, I only had pretrained weights for ResNet50. Papers above claim that with some combinations of AutoAugment and DropBlock, RetinaNet trains even better without a pretrained backbone! This wasn't the case for me, though. Anyway, I had a lot of horsepowers,  so I downloaded ImageNet and made my own pretrained models for ResNet101, ResNet152, and ResNet200.\n\nLong story short: NAS-FPN is great, DropBlock really helps, and AutoAugment augmentations do improve model performance.\n\n# Data labeling\nData labels are really noisy. I manually selected 5 labels per class for the validation set. I made a blacklist of labels which are clearly wrong: i.e. the label is \"cello\", but it's actually a violin.\n\n# Last day inference and final ensemble construction\nWe had seriously run out of time. I had to write and fix the distributed inference script in the last two days of the competition. This could not end well, and it didn't. I trained models with ResNet152/200 backbones, but we couldn't make a proper ensemble.\n\nOh yeah, the last night before the deadline I got an error with inference with EfficientNet models: https://github.com/tensorflow/tensorflow/issues/29574. So we could not use any of my EfficientNet models in the final ensemble. I probably could work around this issue by hacking TPUEstimator, but... time is money, you know :)\n\n# Conclusions\nWell, it was fun! TPUs are fun. As long as someone else is paying. We spent about $1600 worth of GCP credits, and I paid $75 out of my pocket on top of that.\n\nI'm sharing my forked TPU repo here: https://github.com/artyompal/tpu_models. I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.\n\n____________\n\nThanks to @dsvolkov for being my teammate and for all that help with ensembling, debugging and diagnostics!",
      "votes": null
    },
    {
      "id": "638795",
      "postDate": "10/02/2019 12:44:27",
      "content": "<p>Great job! Thanks for sharing the TPU code too, there is, somehow, a serious lack of functioning tpu code out there especially for custom stuff so this is super helpful.</p>",
      "rawMarkdown": "Great job! Thanks for sharing the TPU code too, there is, somehow, a serious lack of functioning tpu code out there especially for custom stuff so this is super helpful.",
      "votes": null
    },
    {
      "id": "638800",
      "postDate": "10/02/2019 12:51:42",
      "content": "<p>Cheers! Google's ready TPU models are quite good, but it's not easy to port something new. For example, I needed something like Cascade R-CNN / HTC for the segmentation task.</p>",
      "rawMarkdown": "Cheers! Google's ready TPU models are quite good, but it's not easy to port something new. For example, I needed something like Cascade R-CNN / HTC for the segmentation task.",
      "votes": null
    },
    {
      "id": "638815",
      "postDate": "10/02/2019 13:08:03",
      "content": "<p>Thanks for sharing :) You are the TRUE WINNER :)</p>",
      "rawMarkdown": "Thanks for sharing :) You are the TRUE WINNER :)",
      "votes": null
    },
    {
      "id": "638884",
      "postDate": "10/02/2019 14:23:58",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/artyomp\">@artyomp</a> </p>",
      "rawMarkdown": "Congrats\nThank you for Sharing your Approach &amp; Insights… @artyomp",
      "votes": null
    },
    {
      "id": "638895",
      "postDate": "10/02/2019 14:30:47",
      "content": "<p>You deserved a special prize from Google for promoting TPU :)</p>",
      "rawMarkdown": "You deserved a special prize from Google for promoting TPU :)",
      "votes": null
    },
    {
      "id": "638912",
      "postDate": "10/02/2019 14:50:59",
      "content": "<p><a href=\"/artyomp\">@artyomp</a> - Thanks a lot for sharing the TPU code. Congrats on winning Silver Medal. It was a challenge for me to use TPU for this competition, now will use your github link for all competitions where there is a scope to use TPU :)</p>",
      "rawMarkdown": "artyomp - Thanks a lot for sharing the TPU code. Congrats on winning Silver Medal. It was a challenge for me to use TPU for this competition, now will use your github link for all competitions where there is a scope to use TPU :)",
      "votes": null
    },
    {
      "id": "638938",
      "postDate": "10/02/2019 15:29:51",
      "content": "<p>Just a job offer would be good enough 😁</p>",
      "rawMarkdown": "Just a job offer would be good enough 😁",
      "votes": null
    },
    {
      "id": "639328",
      "postDate": "10/03/2019 04:43:25",
      "content": "<p>Hi, <a href=\"/artyomp\">@artyomp</a>, Congrats on winning Silver Medal. Thanks for sharing the code</p>",
      "rawMarkdown": "Hi, @artyomp, Congrats on winning Silver Medal. Thanks for sharing the code",
      "votes": null
    },
    {
      "id": "639488",
      "postDate": "10/03/2019 09:20:15",
      "content": "<p>Congrats and thanks for sharing your approach</p>",
      "rawMarkdown": "Congrats and thanks for sharing your approach",
      "votes": null
    },
    {
      "id": "639554",
      "postDate": "10/03/2019 10:43:16",
      "content": "<p><a href=\"/artyomp\">@artyomp</a> thank you so much for posting the code! the solution is crazy, I was waiting for someone to post TPU code :) I'm also trying to move to TPUs and the TFRecords are killing me hahaha\n&gt; I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.</p>\n\n<p>thank you so much.</p>",
      "rawMarkdown": "artyomp thank you so much for posting the code! the solution is crazy, I was waiting for someone to post TPU code :) I'm also trying to move to TPUs and the TFRecords are killing me hahaha\n&gt; I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.\n\nthank you so much.",
      "votes": null
    },
    {
      "id": "640185",
      "postDate": "10/03/2019 23:35:33",
      "content": "<p>Congrats and thanks for sharing the code.</p>",
      "rawMarkdown": "Congrats and thanks for sharing the code.",
      "votes": null
    },
    {
      "id": "640441",
      "postDate": "10/04/2019 03:05:46",
      "content": "<p>Thank you for sharing~</p>",
      "rawMarkdown": "Thank you for sharing~",
      "votes": null
    },
    {
      "id": "2968566",
      "postDate": "08/24/2024 03:24:05",
      "content": "<p>good for tpu.</p>",
      "rawMarkdown": "good for tpu.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2968566,
      "author_name": "bambooclub",
      "author_url": "",
      "post_date": "08/24/2024 03:24:05",
      "content": "<p>good for tpu.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638795,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "10/02/2019 12:44:27",
      "content": "<p>Great job! Thanks for sharing the TPU code too, there is, somehow, a serious lack of functioning tpu code out there especially for custom stuff so this is super helpful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 638800,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "10/02/2019 12:51:42",
          "content": "<p>Cheers! Google's ready TPU models are quite good, but it's not easy to port something new. For example, I needed something like Cascade R-CNN / HTC for the segmentation task.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 638815,
      "author_name": "seriousran",
      "author_url": "",
      "post_date": "10/02/2019 13:08:03",
      "content": "<p>Thanks for sharing :) You are the TRUE WINNER :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638884,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/02/2019 14:23:58",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/artyomp\">@artyomp</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638895,
      "author_name": "dsvolkov",
      "author_url": "",
      "post_date": "10/02/2019 14:30:47",
      "content": "<p>You deserved a special prize from Google for promoting TPU :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 638938,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "10/02/2019 15:29:51",
          "content": "<p>Just a job offer would be good enough 😁</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 638912,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "10/02/2019 14:50:59",
      "content": "<p><a href=\"/artyomp\">@artyomp</a> - Thanks a lot for sharing the TPU code. Congrats on winning Silver Medal. It was a challenge for me to use TPU for this competition, now will use your github link for all competitions where there is a scope to use TPU :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 639328,
      "author_name": "kranthi9",
      "author_url": "",
      "post_date": "10/03/2019 04:43:25",
      "content": "<p>Hi, <a href=\"/artyomp\">@artyomp</a>, Congrats on winning Silver Medal. Thanks for sharing the code</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 639488,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "10/03/2019 09:20:15",
      "content": "<p>Congrats and thanks for sharing your approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 639554,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "10/03/2019 10:43:16",
      "content": "<p><a href=\"/artyomp\">@artyomp</a> thank you so much for posting the code! the solution is crazy, I was waiting for someone to post TPU code :) I'm also trying to move to TPUs and the TFRecords are killing me hahaha\n&gt; I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.</p>\n\n<p>thank you so much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 640185,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "10/03/2019 23:35:33",
      "content": "<p>Congrats and thanks for sharing the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 640441,
      "author_name": "revetice",
      "author_url": "",
      "post_date": "10/04/2019 03:05:46",
      "content": "<p>Thank you for sharing~</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638790": "TL;DR: TPUs actually work. They are fast (10-100 times of 1080Ti). There are limitations that make porting difficult. My fixed code and models are at  https://github.com/artyompal/tpu_models.\n____________\n# TPU\nI started this competition quite late. I had 1 month, which is not enough to train a lot of models of GPU. I had some TPU credits, so I decided to give it a try.\n\nI forked TPU models repo and trained RetinaNet + FPN + ResNet50. I trained the model like in this tutorial: https://cloud.google.com/tpu/docs/tutorials/retinanet and it actually worked. To achieve this, I had to convert my data into TFRecords and also convert my validation data to a COCO JSON file. I realized that our metric (AP@50) is just an ordinary COCO/PASCAL VOC AP@50 metric, as long as there is no hierarchy. So training just worked out of the box.\n\nWhat's TPU? It's a matrix chip. It's a stand-alone device in Google's local network which connects to your cloud instance as TensorFlow master. It has a lot of HBM memory (I used batch 64 with images up to 1024). It has no HDD and it can't connect to your SSD, so it must use Google Cloud Storage for input (images and labels) and output (weights and TensorBoard logs). I had v2x8 and v3x8 TPUs; 3rd generation is ~50% faster.\n\nRetinaNet+ResNet50, batch size 64, images 640x640 processes 250 samples per second on TPUv2. On TPUv3, it processes 350 samples per second:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1859557%2F0373ce32c12fbd7aa0b29ba5ee9b288b%2Fperformance.png?generation=1570020130759200&amp;alt=media)\n\nOne issue about TPUs is, they currently require a completely static TF graph. So you have to resize every image to a fixed resolution. It's a limitation of the current implementation (TensorFlow 1.14). Another issue is, image channels are swapped from NHWC to HWCN. Another thing, it doesn't support some TF operations. And you would need special pretrained weight for TPU, normal TF models will not work.\n\nOh yeah, the main issue is, one has to use TPUEstimator to train models. tf.Estimators are weird pieces of software which by default erase all of your checkpoints but last five, completely ignoring your validation results. Every Deep Learning framework supports validation score-based ReduceLROnPlateau, but not tf.Estimators. If only we had PyTorch/XLA working... (I didn't try it)\n\n# Training\nNaive training on 443 leaf classes resulted in 0.283 on the public leaderboard. Following last year solutions, I split dataset classes into 6 parts by frequency:\n1. Classes 1-100\n2. Classes 101-200\n3. Classes 201-300\n4. Classes 301-400\n5. Classes 401-432\n6. Human parts (11 classes), as they have a special dataset.\n\nI trained 6 models by parts and got a baseline which resulted in 0.48 on the public leaderboard. We teamed up with Dmitry, who had collected 5 pretrained models from different sources and got a script with Soft-NMS blending. This gave us 0.54 on the public LB. I also trained this model by 5 folds, which resulted in 0.51 LB by itself and 0.55 in the ensemble. I did it because I wanted to do pseudo-labeling for some data cleaning, but ran out of time to finish this.\n\n# Models\nTPU models repo has out-of-the-box support of RetinaNet with NAS-FPN (https://arxiv.org/pdf/1904.07392.pdf) with AutoAugment augmentations (https://arxiv.org/pdf/1805.09501.pdf) and DropBlock regularization for both FPN part and backbone (https://arxiv.org/pdf/1810.12890.pdf). By default, RetinaNet only supports ResNet backbone, but I added support of EfficientNet. I also tried adding SE-ResNeXt backbone, but it worked too slow, maybe because I didn't transpose channels to HWCN. Also, I recall that I had a problem with TPUs not supporting REFLECT padding: https://cloud.google.com/tpu/docs/tensorflow-ops.\n\nOf the ResNet family, I only had pretrained weights for ResNet50. Papers above claim that with some combinations of AutoAugment and DropBlock, RetinaNet trains even better without a pretrained backbone! This wasn't the case for me, though. Anyway, I had a lot of horsepowers,  so I downloaded ImageNet and made my own pretrained models for ResNet101, ResNet152, and ResNet200.\n\nLong story short: NAS-FPN is great, DropBlock really helps, and AutoAugment augmentations do improve model performance.\n\n# Data labeling\nData labels are really noisy. I manually selected 5 labels per class for the validation set. I made a blacklist of labels which are clearly wrong: i.e. the label is \"cello\", but it's actually a violin.\n\n# Last day inference and final ensemble construction\nWe had seriously run out of time. I had to write and fix the distributed inference script in the last two days of the competition. This could not end well, and it didn't. I trained models with ResNet152/200 backbones, but we couldn't make a proper ensemble.\n\nOh yeah, the last night before the deadline I got an error with inference with EfficientNet models: https://github.com/tensorflow/tensorflow/issues/29574. So we could not use any of my EfficientNet models in the final ensemble. I probably could work around this issue by hacking TPUEstimator, but... time is money, you know :)\n\n# Conclusions\nWell, it was fun! TPUs are fun. As long as someone else is paying. We spent about $1600 worth of GCP credits, and I paid $75 out of my pocket on top of that.\n\nI'm sharing my forked TPU repo here: https://github.com/artyompal/tpu_models. I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.\n\n____________\n\nThanks to @dsvolkov for being my teammate and for all that help with ensembling, debugging and diagnostics!",
    "638795": "Great job! Thanks for sharing the TPU code too, there is, somehow, a serious lack of functioning tpu code out there especially for custom stuff so this is super helpful.",
    "638800": "Cheers! Google's ready TPU models are quite good, but it's not easy to port something new. For example, I needed something like Cascade R-CNN / HTC for the segmentation task.",
    "638815": "Thanks for sharing :) You are the TRUE WINNER :)",
    "638884": "Congrats\nThank you for Sharing your Approach &amp; Insights… @artyomp",
    "638895": "You deserved a special prize from Google for promoting TPU :)",
    "638912": "artyomp - Thanks a lot for sharing the TPU code. Congrats on winning Silver Medal. It was a challenge for me to use TPU for this competition, now will use your github link for all competitions where there is a scope to use TPU :)",
    "638938": "Just a job offer would be good enough 😁",
    "639328": "Hi, @artyomp, Congrats on winning Silver Medal. Thanks for sharing the code",
    "639488": "Congrats and thanks for sharing your approach",
    "639554": "artyomp thank you so much for posting the code! the solution is crazy, I was waiting for someone to post TPU code :) I'm also trying to move to TPUs and the TFRecords are killing me hahaha\n&gt; I'm also going to upload my pretrained weights for TPU models. I'm editing the repo right now, please let me know if something is missing in the README.md.\n\nthank you so much.",
    "640185": "Congrats and thanks for sharing the code.",
    "640441": "Thank you for sharing~",
    "2968566": "good for tpu."
  },
  "source": "meta"
}