{
  "id": 150454,
  "title": "Current first place solution write-up",
  "url": "/competitions/flower-classification-with-tpus/writeups/kvr777-current-first-place-solution-write-up",
  "author_name": "",
  "post_date": "2020-05-14T10:41:54.267Z",
  "votes": 41,
  "comment_count": 25,
  "views": 0,
  "content": "<p><strong>DISCLAIMER</strong></p>\n\n<p><em>I don't read competition forum very often and as a result I missed  the information that usage of external datasets (I used the ones that were shared by <a href=\"/kirillblinov\">@kirillblinov</a>) was banned several days ago. As a result, there is very high probability that my solution is not eligible for prizes (but I ask confirm it from the organizer's side). However, I've got some interesting findings and maybe it would be interesting for other participants.</em></p>\n\n<p>First of all I would like to thank organizers for the good opportunity to test and evaluate TPU technology in this competition. Also I would like to note some participants who contribute a much: <a href=\"/hengck23\">@hengck23</a>  for sharing fresh ideas, experiments and external datasets, <a href=\"/cdeotte\">@cdeotte</a>  for publishing great notebook and adapting augmentations for TPU kernel and <a href=\"/kirillblinov\">@kirillblinov</a>  for processing and sharing external datasets.</p>\n\n<p>My main goal for this competition was to understand how to use free TPU kernels and what are the limitations when it used for free. I decided to use tensorflow instead of pytorch especially due to the fact that there is a great repository with the models specially adapted for TPU kernels: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official\">https://github.com/tensorflow/tpu/tree/master/models/official</a>. That’s was a plan and below are the results of my experiments.</p>\n\n<p><strong>Technical challenges with free tpu kernel</strong></p>\n\n<ol>\n<li>The most common issue I discovered is a <em>“file system scheme [local] is not implement”</em> error. You can’t save files (checkpoints, logs) on local drive or google drive. Instead, google cloud storage (GCS) is required for input data and output. Mainly due to this fact I didn’t use the models from tensorflow github repo - they built on tf 1.x (efficientnet models) and disk space on GCS is required for temporary checkpoints during training.</li>\n<li>The second issue was that some useful image processing utilites (e.g. augmentations) was removed from tensorflow 2.0 core and placed to tensorflow-addons (tfa). But tfa doesnt work properly with kaggle kernel TPU. This issue was solved by <a href=\"/cdeotte\">@cdeotte</a>  who provided code for main augmentations that is compatible for kaggle TPU kernel.</li>\n<li>The third issue was tensorflow-related: for every kind of model you have to use different tf releases: for example classical resnets-like model was ported to tf 2.x while efficientnet models works only with tf 1.x . When I tried to use these models in tf.compat.v1 regime (with paid GCS), I’ve got multiple depreciation warnings and the training results were bad.</li>\n<li>Last but interesting issue - some useful features (like mixed precision training) were not working with TF 2.1 release on TPU. However it's work well on any TF 2.2 version. But TF 2.2 version was not stable and provided some other bugs (that I could fix)</li>\n</ol>\n\n<p>In the end, I decided to use tf2 keras. For unknown reasons Keras let to save checkpoints on local kernel disk, have pretrained models for tf 2.X and it was widely used by other participants of this competition. </p>\n\n<p><strong>Solution tips and tricks</strong></p>\n\n<ol>\n<li><strong>Data.</strong> I used external datasets prepared by <a href=\"/hengck23\">@hengck23</a> and adapted by <a href=\"/kirillblinov\">@kirillblinov</a> . My experiment showed that openimage and inaturalist decrease accuracy, so I removed them. I used pictures with 512 and 331 sizes. Also with mixed precision training I could use large batches (192) for all model architectures.</li>\n<li><strong>Prepocessing.</strong> I tried autoaugment, randaugment (the basis was official tensorflow implementation for tf 1.X and then adapted code for tf 2.X with TPU) and training without any augmentations. My experiments showed that with big datasets it is better don't use any augmentations.</li>\n<li><strong>Model selection.</strong> I tried all major efficientnet architectures. The best for me were b5 and b6 architecture with noisy-student pre-trained weights (great thanks to <a href=\"/pavel92\">@pavel92</a>) . The classical resnets (50 and 101) as well as se-resnext 50-101 were not as well as efficientnet.</li>\n<li><strong>Model training.</strong> I tried different strategies for cross-validation: 5-fold training, train-validation split, training without validation. As it was noted by many participants, training without validation provide higher score on public leaderboard and my final solution included the models that were trained without validation part. I used exponential decay LR scheduler with warmup. </li>\n<li><strong>Ensembling.</strong> As i decided to use aggressive training strategy (no augmentations, no cross-validation), the ensembling was essential to avoid painful falling on private leaderboard. I trained best architectures (b5, b6, b7) with different picture sizes, try to combine effnet and seresnext models. The winner blend is three models B5-512, B5-331 and B6-512. This solution also provides my  best score on public LB.</li>\n</ol>\n\n<p>As a result, I would say that though TPU kernels have some limitations, little bugs it is great opportunity for the researchers with limited computational capacities. In majority of computer vision competitions I participated with one laptop GPU (8gb) and google colab gpu. The model that usually calculated one day locally can be processed on TPU within one hour or faster. This technology is a great equalizer on kaggle competitions and it could facilitate deep learning researches as well.</p>\n\n<p>Regarding the usage of external dataset, I can confirm and assure that I don't know about the ban for usage the external datasets. Moreover, when I read the forum last time (one or two weeks ago) this topic was discussed and my understanding was that it'is ok to use this external dataset. It would be great to implement some alerting system for all competition participants to spread the news like this. </p>\n\n<p>However, despite the issue like that it was a great time to participate in this competition. Thank you very much for all participants and organizers.</p>\n\n<p><strong>UPD</strong>: </p>\n\n<p>Some participants ask to me describe how did I find right model configuration without having validation part. Below is the high level explanation.</p>\n\n<ol>\n<li>At first step I buit multi-dimensional grid  with the parameters I would like to test:\n a. <strong>Model architectures</strong> (Effnet b4, b5, b6, b7; Resnet 50, 101, Se-Resnext 50, 101, Inception)\n b. <strong>Pictures size</strong> (512, 331, 224)\n c. <strong>Augmentations</strong> (none, autoaugment, randaugment with different params). I selected these \n three  types because it is easy to implement and they work well to beat some benchmarks in image classification tasks\nd. <strong>Losses</strong> (cross-entropy, focal)\ne. <strong>Optimizer and lr</strong> </li>\n<li>Then I made grid-search of best combinations in manual mode. I don't need to test all possible combinations of my params to understand that pic size 224 provide lower score than 512. Validation part was \"valid\" folder and public leaderboard. </li>\n<li>Thanks to TPU and relatively small dataset, I could test a lot of hypothesis and considerably reduce my parameters grid. Then I added external datasets, <strong>but keep the same validation part</strong>. So after some experiments I could confirm that my validation score improved with external data and public score improved as well. At this step I defined best models (best combo of grid params)</li>\n<li>The best models I test with different combination of external datasets and realize that it is possible to eliminate some of them while improving validation and public LB score</li>\n<li>My best models was the models without augmentations (strictly speaking \"no augmentation\" training mode include random left-right flip). And I realized that training loss correlated very well with validation and public LB score. And adding validation part in training process increase Public LB considerably. Due to this I decided to retrain best models with validation part and use the models with the lowest training loss.</li>\n<li>Finally I've got about 10 models with good score. I tried different blend combinations. I  average probabilities because the majority of these models were effnets with 512 and 331 pic sizes. Nice to try here was to test different combination of voting but it was out of my goal to test TPU capabilites and I decided to use the simplest version of blending.</li>\n</ol>",
  "messages": [
    {
      "id": "843900",
      "postDate": "05/12/2020 10:41:24",
      "content": "<p><strong>DISCLAIMER</strong></p>\n\n<p><em>I don't read competition forum very often and as a result I missed  the information that usage of external datasets (I used the ones that were shared by <a href=\"/kirillblinov\">@kirillblinov</a>) was banned several days ago. As a result, there is very high probability that my solution is not eligible for prizes (but I ask confirm it from the organizer's side). However, I've got some interesting findings and maybe it would be interesting for other participants.</em></p>\n\n<p>First of all I would like to thank organizers for the good opportunity to test and evaluate TPU technology in this competition. Also I would like to note some participants who contribute a much: <a href=\"/hengck23\">@hengck23</a>  for sharing fresh ideas, experiments and external datasets, <a href=\"/cdeotte\">@cdeotte</a>  for publishing great notebook and adapting augmentations for TPU kernel and <a href=\"/kirillblinov\">@kirillblinov</a>  for processing and sharing external datasets.</p>\n\n<p>My main goal for this competition was to understand how to use free TPU kernels and what are the limitations when it used for free. I decided to use tensorflow instead of pytorch especially due to the fact that there is a great repository with the models specially adapted for TPU kernels: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official\">https://github.com/tensorflow/tpu/tree/master/models/official</a>. That’s was a plan and below are the results of my experiments.</p>\n\n<p><strong>Technical challenges with free tpu kernel</strong></p>\n\n<ol>\n<li>The most common issue I discovered is a <em>“file system scheme [local] is not implement”</em> error. You can’t save files (checkpoints, logs) on local drive or google drive. Instead, google cloud storage (GCS) is required for input data and output. Mainly due to this fact I didn’t use the models from tensorflow github repo - they built on tf 1.x (efficientnet models) and disk space on GCS is required for temporary checkpoints during training.</li>\n<li>The second issue was that some useful image processing utilites (e.g. augmentations) was removed from tensorflow 2.0 core and placed to tensorflow-addons (tfa). But tfa doesnt work properly with kaggle kernel TPU. This issue was solved by <a href=\"/cdeotte\">@cdeotte</a>  who provided code for main augmentations that is compatible for kaggle TPU kernel.</li>\n<li>The third issue was tensorflow-related: for every kind of model you have to use different tf releases: for example classical resnets-like model was ported to tf 2.x while efficientnet models works only with tf 1.x . When I tried to use these models in tf.compat.v1 regime (with paid GCS), I’ve got multiple depreciation warnings and the training results were bad.</li>\n<li>Last but interesting issue - some useful features (like mixed precision training) were not working with TF 2.1 release on TPU. However it's work well on any TF 2.2 version. But TF 2.2 version was not stable and provided some other bugs (that I could fix)</li>\n</ol>\n\n<p>In the end, I decided to use tf2 keras. For unknown reasons Keras let to save checkpoints on local kernel disk, have pretrained models for tf 2.X and it was widely used by other participants of this competition. </p>\n\n<p><strong>Solution tips and tricks</strong></p>\n\n<ol>\n<li><strong>Data.</strong> I used external datasets prepared by <a href=\"/hengck23\">@hengck23</a> and adapted by <a href=\"/kirillblinov\">@kirillblinov</a> . My experiment showed that openimage and inaturalist decrease accuracy, so I removed them. I used pictures with 512 and 331 sizes. Also with mixed precision training I could use large batches (192) for all model architectures.</li>\n<li><strong>Prepocessing.</strong> I tried autoaugment, randaugment (the basis was official tensorflow implementation for tf 1.X and then adapted code for tf 2.X with TPU) and training without any augmentations. My experiments showed that with big datasets it is better don't use any augmentations.</li>\n<li><strong>Model selection.</strong> I tried all major efficientnet architectures. The best for me were b5 and b6 architecture with noisy-student pre-trained weights (great thanks to <a href=\"/pavel92\">@pavel92</a>) . The classical resnets (50 and 101) as well as se-resnext 50-101 were not as well as efficientnet.</li>\n<li><strong>Model training.</strong> I tried different strategies for cross-validation: 5-fold training, train-validation split, training without validation. As it was noted by many participants, training without validation provide higher score on public leaderboard and my final solution included the models that were trained without validation part. I used exponential decay LR scheduler with warmup. </li>\n<li><strong>Ensembling.</strong> As i decided to use aggressive training strategy (no augmentations, no cross-validation), the ensembling was essential to avoid painful falling on private leaderboard. I trained best architectures (b5, b6, b7) with different picture sizes, try to combine effnet and seresnext models. The winner blend is three models B5-512, B5-331 and B6-512. This solution also provides my  best score on public LB.</li>\n</ol>\n\n<p>As a result, I would say that though TPU kernels have some limitations, little bugs it is great opportunity for the researchers with limited computational capacities. In majority of computer vision competitions I participated with one laptop GPU (8gb) and google colab gpu. The model that usually calculated one day locally can be processed on TPU within one hour or faster. This technology is a great equalizer on kaggle competitions and it could facilitate deep learning researches as well.</p>\n\n<p>Regarding the usage of external dataset, I can confirm and assure that I don't know about the ban for usage the external datasets. Moreover, when I read the forum last time (one or two weeks ago) this topic was discussed and my understanding was that it'is ok to use this external dataset. It would be great to implement some alerting system for all competition participants to spread the news like this. </p>\n\n<p>However, despite the issue like that it was a great time to participate in this competition. Thank you very much for all participants and organizers.</p>\n\n<p><strong>UPD</strong>: </p>\n\n<p>Some participants ask to me describe how did I find right model configuration without having validation part. Below is the high level explanation.</p>\n\n<ol>\n<li>At first step I buit multi-dimensional grid  with the parameters I would like to test:\n a. <strong>Model architectures</strong> (Effnet b4, b5, b6, b7; Resnet 50, 101, Se-Resnext 50, 101, Inception)\n b. <strong>Pictures size</strong> (512, 331, 224)\n c. <strong>Augmentations</strong> (none, autoaugment, randaugment with different params). I selected these \n three  types because it is easy to implement and they work well to beat some benchmarks in image classification tasks\nd. <strong>Losses</strong> (cross-entropy, focal)\ne. <strong>Optimizer and lr</strong> </li>\n<li>Then I made grid-search of best combinations in manual mode. I don't need to test all possible combinations of my params to understand that pic size 224 provide lower score than 512. Validation part was \"valid\" folder and public leaderboard. </li>\n<li>Thanks to TPU and relatively small dataset, I could test a lot of hypothesis and considerably reduce my parameters grid. Then I added external datasets, <strong>but keep the same validation part</strong>. So after some experiments I could confirm that my validation score improved with external data and public score improved as well. At this step I defined best models (best combo of grid params)</li>\n<li>The best models I test with different combination of external datasets and realize that it is possible to eliminate some of them while improving validation and public LB score</li>\n<li>My best models was the models without augmentations (strictly speaking \"no augmentation\" training mode include random left-right flip). And I realized that training loss correlated very well with validation and public LB score. And adding validation part in training process increase Public LB considerably. Due to this I decided to retrain best models with validation part and use the models with the lowest training loss.</li>\n<li>Finally I've got about 10 models with good score. I tried different blend combinations. I  average probabilities because the majority of these models were effnets with 512 and 331 pic sizes. Nice to try here was to test different combination of voting but it was out of my goal to test TPU capabilites and I decided to use the simplest version of blending.</li>\n</ol>",
      "rawMarkdown": "**DISCLAIMER**\n\n*I don't read competition forum very often and as a result I missed  the information that usage of external datasets (I used the ones that were shared by @kirillblinov) was banned several days ago. As a result, there is very high probability that my solution is not eligible for prizes (but I ask confirm it from the organizer's side). However, I've got some interesting findings and maybe it would be interesting for other participants.*\n\nFirst of all I would like to thank organizers for the good opportunity to test and evaluate TPU technology in this competition. Also I would like to note some participants who contribute a much: @hengck23  for sharing fresh ideas, experiments and external datasets, @cdeotte  for publishing great notebook and adapting augmentations for TPU kernel and @kirillblinov  for processing and sharing external datasets.\n\nMy main goal for this competition was to understand how to use free TPU kernels and what are the limitations when it used for free. I decided to use tensorflow instead of pytorch especially due to the fact that there is a great repository with the models specially adapted for TPU kernels: https://github.com/tensorflow/tpu/tree/master/models/official. That’s was a plan and below are the results of my experiments.\n\n**Technical challenges with free tpu kernel**\n\n1. The most common issue I discovered is a *“file system scheme [local] is not implement”* error. You can’t save files (checkpoints, logs) on local drive or google drive. Instead, google cloud storage (GCS) is required for input data and output. Mainly due to this fact I didn’t use the models from tensorflow github repo - they built on tf 1.x (efficientnet models) and disk space on GCS is required for temporary checkpoints during training.\n2. The second issue was that some useful image processing utilites (e.g. augmentations) was removed from tensorflow 2.0 core and placed to tensorflow-addons (tfa). But tfa doesnt work properly with kaggle kernel TPU. This issue was solved by @cdeotte  who provided code for main augmentations that is compatible for kaggle TPU kernel.\n3. The third issue was tensorflow-related: for every kind of model you have to use different tf releases: for example classical resnets-like model was ported to tf 2.x while efficientnet models works only with tf 1.x . When I tried to use these models in tf.compat.v1 regime (with paid GCS), I’ve got multiple depreciation warnings and the training results were bad.\n4. Last but interesting issue - some useful features (like mixed precision training) were not working with TF 2.1 release on TPU. However it's work well on any TF 2.2 version. But TF 2.2 version was not stable and provided some other bugs (that I could fix)\n\nIn the end, I decided to use tf2 keras. For unknown reasons Keras let to save checkpoints on local kernel disk, have pretrained models for tf 2.X and it was widely used by other participants of this competition. \n\n**Solution tips and tricks**\n\n1. **Data.** I used external datasets prepared by @hengck23 and adapted by @kirillblinov . My experiment showed that openimage and inaturalist decrease accuracy, so I removed them. I used pictures with 512 and 331 sizes. Also with mixed precision training I could use large batches (192) for all model architectures.\n2. **Prepocessing.** I tried autoaugment, randaugment (the basis was official tensorflow implementation for tf 1.X and then adapted code for tf 2.X with TPU) and training without any augmentations. My experiments showed that with big datasets it is better don't use any augmentations.\n3. **Model selection.** I tried all major efficientnet architectures. The best for me were b5 and b6 architecture with noisy-student pre-trained weights (great thanks to @pavel92) . The classical resnets (50 and 101) as well as se-resnext 50-101 were not as well as efficientnet.\n4. **Model training.** I tried different strategies for cross-validation: 5-fold training, train-validation split, training without validation. As it was noted by many participants, training without validation provide higher score on public leaderboard and my final solution included the models that were trained without validation part. I used exponential decay LR scheduler with warmup. \n5. **Ensembling.** As i decided to use aggressive training strategy (no augmentations, no cross-validation), the ensembling was essential to avoid painful falling on private leaderboard. I trained best architectures (b5, b6, b7) with different picture sizes, try to combine effnet and seresnext models. The winner blend is three models B5-512, B5-331 and B6-512. This solution also provides my  best score on public LB.\n\nAs a result, I would say that though TPU kernels have some limitations, little bugs it is great opportunity for the researchers with limited computational capacities. In majority of computer vision competitions I participated with one laptop GPU (8gb) and google colab gpu. The model that usually calculated one day locally can be processed on TPU within one hour or faster. This technology is a great equalizer on kaggle competitions and it could facilitate deep learning researches as well.\n\nRegarding the usage of external dataset, I can confirm and assure that I don't know about the ban for usage the external datasets. Moreover, when I read the forum last time (one or two weeks ago) this topic was discussed and my understanding was that it'is ok to use this external dataset. It would be great to implement some alerting system for all competition participants to spread the news like this. \n\nHowever, despite the issue like that it was a great time to participate in this competition. Thank you very much for all participants and organizers.\n\n**UPD**: \n\nSome participants ask to me describe how did I find right model configuration without having validation part. Below is the high level explanation.\n\n1. At first step I buit multi-dimensional grid  with the parameters I would like to test:\n     a. **Model architectures** (Effnet b4, b5, b6, b7; Resnet 50, 101, Se-Resnext 50, 101, Inception)\n     b. **Pictures size** (512, 331, 224)\n     c. **Augmentations** (none, autoaugment, randaugment with different params). I selected these \n     three  types because it is easy to implement and they work well to beat some benchmarks in image classification tasks\n  d. **Losses** (cross-entropy, focal)\n e. **Optimizer and lr** \n2. Then I made grid-search of best combinations in manual mode. I don't need to test all possible combinations of my params to understand that pic size 224 provide lower score than 512. Validation part was \"valid\" folder and public leaderboard. \n3. Thanks to TPU and relatively small dataset, I could test a lot of hypothesis and considerably reduce my parameters grid. Then I added external datasets, **but keep the same validation part**. So after some experiments I could confirm that my validation score improved with external data and public score improved as well. At this step I defined best models (best combo of grid params)\n4.  The best models I test with different combination of external datasets and realize that it is possible to eliminate some of them while improving validation and public LB score\n5.  My best models was the models without augmentations (strictly speaking \"no augmentation\" training mode include random left-right flip). And I realized that training loss correlated very well with validation and public LB score. And adding validation part in training process increase Public LB considerably. Due to this I decided to retrain best models with validation part and use the models with the lowest training loss.\n6. Finally I've got about 10 models with good score. I tried different blend combinations. I  average probabilities because the majority of these models were effnets with 512 and 331 pic sizes. Nice to try here was to test different combination of voting but it was out of my goal to test TPU capabilites and I decided to use the simplest version of blending.",
      "votes": null
    },
    {
      "id": "844118",
      "postDate": "05/12/2020 13:07:50",
      "content": "<p>thx for sharing!</p>\n\n<p>I did not have problems with mixed precision  training.</p>\n\n<p>One has to define the last layer like this:\ntf.keras.layers.Dense(len(CLASSES), activation='softmax',<strong>dtype='float32</strong>')</p>\n\n<p>I think Chris pointed this out in a comment</p>",
      "rawMarkdown": "thx for sharing!\n\nI did not have problems with mixed precision  training.\n\nOne has to define the last layer like this:\ntf.keras.layers.Dense(len(CLASSES), activation='softmax',**dtype='float32**')\n\nI think Chris pointed this out in a comment",
      "votes": null
    },
    {
      "id": "844163",
      "postDate": "05/12/2020 13:39:27",
      "content": "<p>mixed precision training works for me also (tf 2.1)</p>",
      "rawMarkdown": "mixed precision training works for me also (tf 2.1)",
      "votes": null
    },
    {
      "id": "844482",
      "postDate": "05/12/2020 16:24:04",
      "content": "<p>Yes, thank you for the comment. I had the issue described here: <a href=\"https://github.com/tensorflow/tensorflow/issues/34782\">https://github.com/tensorflow/tensorflow/issues/34782</a>\nAnd defining last layer as a float32 it was a possible workaround to solve it for tf 2.1 version. I experimented with different losses and layer configurations, so for me less problematic was to upgrade tf version (it was fixed in tf 2.2 pre-release).</p>",
      "rawMarkdown": "Yes, thank you for the comment. I had the issue described here: https://github.com/tensorflow/tensorflow/issues/34782\nAnd defining last layer as a float32 it was a possible workaround to solve it for tf 2.1 version. I experimented with different losses and layer configurations, so for me less problematic was to upgrade tf version (it was fixed in tf 2.2 pre-release).",
      "votes": null
    },
    {
      "id": "844803",
      "postDate": "05/12/2020 21:32:01",
      "content": "<p>I confirm that saving model and checkpoints from TPU to local disk only works using the Keras HDF5 format. Saving in the newer \"Tensorflow Saved Model\" format from TPU to local disk is currently buggy. Saving in \"TF Saved Model\" format from TPU to GCS works though. You can use that on GCP. Not on Kaggle unfortunately because Kaggle notebooks do not have access to a writable GCS location.</p>",
      "rawMarkdown": "I confirm that saving model and checkpoints from TPU to local disk only works using the Keras HDF5 format. Saving in the newer \"Tensorflow Saved Model\" format from TPU to local disk is currently buggy. Saving in \"TF Saved Model\" format from TPU to GCS works though. You can use that on GCP. Not on Kaggle unfortunately because Kaggle notebooks do not have access to a writable GCS location.",
      "votes": null
    },
    {
      "id": "844823",
      "postDate": "05/12/2020 21:57:01",
      "content": "<p>I just saw this answer after I published a new post. <a href=\"/mgornergoogle\">@mgornergoogle</a> , is there a plan to make saving checkpoints / weights possible even if the model is a subclassed model on Kaggle when using TPU?</p>",
      "rawMarkdown": "I just saw this answer after I published a new post. @mgornergoogle , is there a plan to make saving checkpoints / weights possible even if the model is a subclassed model on Kaggle when using TPU?",
      "votes": null
    },
    {
      "id": "844868",
      "postDate": "05/12/2020 22:36:13",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> hi. I forgot to thank you for the support during this competition. Regarding to saving feature it would be great to implement the possibility save temporary checkpoint produced by tf.compat.v1.estimator.tpu.TPUEstimator  . In this case kagglers can experiment with the variety of pretrained models from this repo: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a> </p>",
      "rawMarkdown": "mgornergoogle hi. I forgot to thank you for the support during this competition. Regarding to saving feature it would be great to implement the possibility save temporary checkpoint produced by tf.compat.v1.estimator.tpu.TPUEstimator  . In this case kagglers can experiment with the variety of pretrained models from this repo: https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet",
      "votes": null
    },
    {
      "id": "844883",
      "postDate": "05/12/2020 22:51:36",
      "content": "<p>EfficientNet is now available in Tensorflow 2 Saved Model format from TF Hub: <a href=\"https://tfhub.dev/google/collections/efficientnet/1\">https://tfhub.dev/google/collections/efficientnet/1</a></p>\n\n<p>You can use it on Kaggle by copying the model to a Kaggle dataset and then using the following code:</p>\n\n<p>```\nfrom kaggle_datasets import KaggleDatasets\nGCS_PATH_SAVEDMODEL = KaggleDatasets().get_gcs_path('your-efficientnet-kaggle-dataset')\nlayer = tf.saved_model.load(GCS_PATH_SAVEDMODEL)</p>\n\n<h1>Cast the loaded model to a TFHub KerasLayer.</h1>\n\n<p>layer = hub.KerasLayer(layer, trainable=True)\n```</p>\n\n<p>(example code <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">here</a>, search for \"hub.KerasLayer\")</p>",
      "rawMarkdown": "EfficientNet is now available in Tensorflow 2 Saved Model format from TF Hub: https://tfhub.dev/google/collections/efficientnet/1\n\nYou can use it on Kaggle by copying the model to a Kaggle dataset and then using the following code:\n\n```\nfrom kaggle_datasets import KaggleDatasets\nGCS_PATH_SAVEDMODEL = KaggleDatasets().get_gcs_path('your-efficientnet-kaggle-dataset')\nlayer = tf.saved_model.load(GCS_PATH_SAVEDMODEL)\n# Cast the loaded model to a TFHub KerasLayer.\nlayer = hub.KerasLayer(layer, trainable=True)\n```\n\n(example code [here](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started), search for \"hub.KerasLayer\")",
      "votes": null
    },
    {
      "id": "844885",
      "postDate": "05/12/2020 22:53:26",
      "content": "<p>What exactly is not working with subclassed models ? Do you have an example ?</p>",
      "rawMarkdown": "What exactly is not working with subclassed models ? Do you have an example ?",
      "votes": null
    },
    {
      "id": "844897",
      "postDate": "05/12/2020 23:10:24",
      "content": "<p>Thank you for the sharing this hub. However this hub is missing some large models that is interesting to train on TPU (B8, L2). Also the github repo contain more checkpoints with different options (noisy students, different augmentations like randaugment). </p>",
      "rawMarkdown": "Thank you for the sharing this hub. However this hub is missing some large models that is interesting to train on TPU (B8, L2). Also the github repo contain more checkpoints with different options (noisy students, different augmentations like randaugment).",
      "votes": null
    },
    {
      "id": "844908",
      "postDate": "05/12/2020 23:30:13",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a>, sorry as I understand subclassed models are Keras models? The repo I mentioned uses pure tensorflow without keras. The error I encountered is  “file system scheme [local] is not implement” because TPUEstimator.train() method try to save temp checkpoints during training. I can create the notebook tomorrow to demonstrate this issue.</p>",
      "rawMarkdown": "mgornergoogle, sorry as I understand subclassed models are Keras models? The repo I mentioned uses pure tensorflow without keras. The error I encountered is  “file system scheme [local] is not implement” because TPUEstimator.train() method try to save temp checkpoints during training. I can create the notebook tomorrow to demonstrate this issue.",
      "votes": null
    },
    {
      "id": "844923",
      "postDate": "05/12/2020 23:45:41",
      "content": "<p>TPUEstimator is legacy TF1.x suff now. It won't get much updates.\nPure Tensorflow without Keras today means using custom training loop APIs like <a href=\"https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy#run\">TPUStrategy.run</a>. Like in this example: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a></p>\n\n<p>I understand your frustration about the pace of TF2 transition of Tensorflow model garden models. I will communicate this feedback back to the TF models team.</p>",
      "rawMarkdown": "TPUEstimator is legacy TF1.x suff now. It won't get much updates.\nPure Tensorflow without Keras today means using custom training loop APIs like [TPUStrategy.run](https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy#run). Like in this example: [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\n\nI understand your frustration about the pace of TF2 transition of Tensorflow model garden models. I will communicate this feedback back to the TF models team.",
      "votes": null
    },
    {
      "id": "844930",
      "postDate": "05/12/2020 23:52:35",
      "content": "<p>This repo seems to be the Keras version of EfficientNet. In the official Tensorflow model Garden:\n<a href=\"https://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/\">https://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/</a></p>\n\n<p>I see this line in the code:\n<code>class EfficientNet(tf.keras.Model):</code></p>",
      "rawMarkdown": "This repo seems to be the Keras version of EfficientNet. In the official Tensorflow model Garden:\nhttps://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/\n\nI see this line in the code:\n`class EfficientNet(tf.keras.Model):`",
      "votes": null
    },
    {
      "id": "844940",
      "postDate": "05/13/2020 00:12:20",
      "content": "<p>Thank you for the kernel with the custom loop. I began my competition from this. I like tf 2 style, it is now looks like pytorch.  But I couldn’t  load pretrained tf 1. models under tf2 regime and was forced to use legacy methods to test official pretrained tf models. I will be waiting for release of pretrained TPU models for TF2. Thank you.</p>",
      "rawMarkdown": "Thank you for the kernel with the custom loop. I began my competition from this. I like tf 2 style, it is now looks like pytorch.  But I couldn’t  load pretrained tf 1. models under tf2 regime and was forced to use legacy methods to test official pretrained tf models. I will be waiting for release of pretrained TPU models for TF2. Thank you.",
      "votes": null
    },
    {
      "id": "845118",
      "postDate": "05/13/2020 04:02:19",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "845847",
      "postDate": "05/13/2020 13:02:13",
      "content": "<p>Congratulations <a href=\"/kromanov\">@kromanov</a> for holding the 1st position and really hope you don't get disqualified. </p>\n\n<p>I have few questions,\n1.  Can you please highlight more about the ensemble strategy you chose. You mention that you used a blend of your top 3 performing models. How exactly does that work?\n2. If you were facing issues in mixed precision training for TF-2.1 how were you able to solve this?</p>\n\n<p>Thanks and Congrats again. </p>",
      "rawMarkdown": "Congratulations @kromanov for holding the 1st position and really hope you don't get disqualified. \n\nI have few questions,\n1.  Can you please highlight more about the ensemble strategy you chose. You mention that you used a blend of your top 3 performing models. How exactly does that work?\n2. If you were facing issues in mixed precision training for TF-2.1 how were you able to solve this?\n\nThanks and Congrats again.",
      "votes": null
    },
    {
      "id": "847005",
      "postDate": "05/14/2020 05:49:03",
      "content": "<p>Thank you <a href=\"/kromanov\">@kromanov</a>  for providing issues and their current workarounds of using TPUs. I hope you get your position what you deserve. I am new to kaggle. Great tips and tricks. 👍 👍 </p>",
      "rawMarkdown": "Thank you @kromanov  for providing issues and their current workarounds of using TPUs. I hope you get your position what you deserve. I am new to kaggle. Great tips and tricks. 👍 👍",
      "votes": null
    },
    {
      "id": "847313",
      "postDate": "05/14/2020 09:47:52",
      "content": "<p>hi <a href=\"/rhtsingh\">@rhtsingh</a> . Thank you.</p>\n\n<ol>\n<li>Strickly speaking it is just blending - I averaged the probabilities that was produced by these three models and then take the class with the max probability for every record</li>\n<li>I just make \"pip tensorflow -U\" command before launching computations. But if your goal is just mixed precision training, it is better to use trick described by <a href=\"/romanweilguny\">@romanweilguny</a> below.</li>\n</ol>",
      "rawMarkdown": "hi @rhtsingh . Thank you.\n\n1. Strickly speaking it is just blending - I averaged the probabilities that was produced by these three models and then take the class with the max probability for every record\n2. I just make \"pip tensorflow -U\" command before launching computations. But if your goal is just mixed precision training, it is better to use trick described by @romanweilguny below.",
      "votes": null
    },
    {
      "id": "847320",
      "postDate": "05/14/2020 09:59:30",
      "content": "<p>Thanks a ton. 😊 </p>",
      "rawMarkdown": "Thanks a ton. 😊",
      "votes": null
    },
    {
      "id": "847358",
      "postDate": "05/14/2020 10:43:01",
      "content": "<p>Hi <a href=\"/redwankarimsony\">@redwankarimsony</a> . Thank you!</p>",
      "rawMarkdown": "Hi @redwankarimsony . Thank you!",
      "votes": null
    },
    {
      "id": "868868",
      "postDate": "05/31/2020 14:43:55",
      "content": "<p>hi <a href=\"/romanweilguny\">@romanweilguny</a> <a href=\"/kromanov\">@kromanov</a> <a href=\"/yihdarshieh\">@yihdarshieh</a>  is adding that snippet is all I need to do in TPU?</p>\n\n<p>```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)</p>\n\n<p>```\nI wont have to add these lines in the beginning?</p>",
      "rawMarkdown": "hi @romanweilguny @kromanov @yihdarshieh  is adding that snippet is all I need to do in TPU?\n\n```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n\n```\nI wont have to add these lines in the beginning?",
      "votes": null
    },
    {
      "id": "868999",
      "postDate": "05/31/2020 16:17:44",
      "content": "<p>yes - but I disabled XLA_ACCELERATE on tpu due to some problems...</p>\n\n<p>if XLA_ACCELERATE:\n    if not tpu:\n        tf.config.optimizer.set_jit(True)\n        print('Accelerated Linear Algebra enabled')</p>",
      "rawMarkdown": "yes - but I disabled XLA_ACCELERATE on tpu due to some problems...\n\nif XLA_ACCELERATE:\n    if not tpu:\n        tf.config.optimizer.set_jit(True)\n        print('Accelerated Linear Algebra enabled')",
      "votes": null
    },
    {
      "id": "869051",
      "postDate": "05/31/2020 17:37:11",
      "content": "<p>I dont see anywhere in the code where you disable it. \nYou enable xla if there's is no TPU i.e. if the runtime is on GPU or CPU. \nCorrect me if Im wrong?</p>",
      "rawMarkdown": "I dont see anywhere in the code where you disable it. \nYou enable xla if there's is no TPU i.e. if the runtime is on GPU or CPU. \nCorrect me if Im wrong?",
      "votes": null
    },
    {
      "id": "869073",
      "postDate": "05/31/2020 18:05:21",
      "content": "<p>I think it is disabled by default. Correct me if I'm wrong?</p>",
      "rawMarkdown": "I think it is disabled by default. Correct me if I'm wrong?",
      "votes": null
    },
    {
      "id": "869150",
      "postDate": "05/31/2020 19:09:06",
      "content": "<p><a href=\"/rhtsingh\">@rhtsingh</a> , hi. In tpu you have to use special type: mixed_bfloat16. See below the snippet:\n```\nif USE_FLOAT16:\n    from tensorflow.keras.mixed_precision import experimental as mixed_precision\n    if tpu: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\n    else: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n    print('Mixed precision enabled')</p>\n\n<p>if XLA_ACCELERATE:\n    tf.config.optimizer.set_jit(True)\n    print('Accelerated Linear Algebra enabled')\n```\nI enabled mixed precision and xla accelerate. With tf 2.2 it doesn't produce any errors</p>",
      "rawMarkdown": "rhtsingh , hi. In tpu you have to use special type: mixed_bfloat16. See below the snippet:\n```\nif USE_FLOAT16:\n    from tensorflow.keras.mixed_precision import experimental as mixed_precision\n    if tpu: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\n    else: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n    print('Mixed precision enabled')\n\nif XLA_ACCELERATE:\n    tf.config.optimizer.set_jit(True)\n    print('Accelerated Linear Algebra enabled')\n```\nI enabled mixed precision and xla accelerate. With tf 2.2 it doesn't produce any errors",
      "votes": null
    },
    {
      "id": "869643",
      "postDate": "06/01/2020 07:10:48",
      "content": "<p>Wow, amazing.! Many many thanks for all the knowledge you have shared here <a href=\"/kromanov\">@kromanov</a> and <a href=\"/romanweilguny\">@romanweilguny</a>  for helping people and your outstanding kernel. </p>",
      "rawMarkdown": "Wow, amazing.! Many many thanks for all the knowledge you have shared here @kromanov and @romanweilguny  for helping people and your outstanding kernel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 844118,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "05/12/2020 13:07:50",
      "content": "<p>thx for sharing!</p>\n\n<p>I did not have problems with mixed precision  training.</p>\n\n<p>One has to define the last layer like this:\ntf.keras.layers.Dense(len(CLASSES), activation='softmax',<strong>dtype='float32</strong>')</p>\n\n<p>I think Chris pointed this out in a comment</p>",
      "votes": null,
      "replies": [
        {
          "id": 844163,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/12/2020 13:39:27",
          "content": "<p>mixed precision training works for me also (tf 2.1)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844482,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/12/2020 16:24:04",
          "content": "<p>Yes, thank you for the comment. I had the issue described here: <a href=\"https://github.com/tensorflow/tensorflow/issues/34782\">https://github.com/tensorflow/tensorflow/issues/34782</a>\nAnd defining last layer as a float32 it was a possible workaround to solve it for tf 2.1 version. I experimented with different losses and layer configurations, so for me less problematic was to upgrade tf version (it was fixed in tf 2.2 pre-release).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868868,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "05/31/2020 14:43:55",
          "content": "<p>hi <a href=\"/romanweilguny\">@romanweilguny</a> <a href=\"/kromanov\">@kromanov</a> <a href=\"/yihdarshieh\">@yihdarshieh</a>  is adding that snippet is all I need to do in TPU?</p>\n\n<p>```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)</p>\n\n<p>```\nI wont have to add these lines in the beginning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868999,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "05/31/2020 16:17:44",
          "content": "<p>yes - but I disabled XLA_ACCELERATE on tpu due to some problems...</p>\n\n<p>if XLA_ACCELERATE:\n    if not tpu:\n        tf.config.optimizer.set_jit(True)\n        print('Accelerated Linear Algebra enabled')</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869051,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "05/31/2020 17:37:11",
          "content": "<p>I dont see anywhere in the code where you disable it. \nYou enable xla if there's is no TPU i.e. if the runtime is on GPU or CPU. \nCorrect me if Im wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869073,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "05/31/2020 18:05:21",
          "content": "<p>I think it is disabled by default. Correct me if I'm wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869150,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/31/2020 19:09:06",
          "content": "<p><a href=\"/rhtsingh\">@rhtsingh</a> , hi. In tpu you have to use special type: mixed_bfloat16. See below the snippet:\n```\nif USE_FLOAT16:\n    from tensorflow.keras.mixed_precision import experimental as mixed_precision\n    if tpu: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\n    else: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n    print('Mixed precision enabled')</p>\n\n<p>if XLA_ACCELERATE:\n    tf.config.optimizer.set_jit(True)\n    print('Accelerated Linear Algebra enabled')\n```\nI enabled mixed precision and xla accelerate. With tf 2.2 it doesn't produce any errors</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869643,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "06/01/2020 07:10:48",
          "content": "<p>Wow, amazing.! Many many thanks for all the knowledge you have shared here <a href=\"/kromanov\">@kromanov</a> and <a href=\"/romanweilguny\">@romanweilguny</a>  for helping people and your outstanding kernel. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 844803,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "05/12/2020 21:32:01",
      "content": "<p>I confirm that saving model and checkpoints from TPU to local disk only works using the Keras HDF5 format. Saving in the newer \"Tensorflow Saved Model\" format from TPU to local disk is currently buggy. Saving in \"TF Saved Model\" format from TPU to GCS works though. You can use that on GCP. Not on Kaggle unfortunately because Kaggle notebooks do not have access to a writable GCS location.</p>",
      "votes": null,
      "replies": [
        {
          "id": 844823,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "05/12/2020 21:57:01",
          "content": "<p>I just saw this answer after I published a new post. <a href=\"/mgornergoogle\">@mgornergoogle</a> , is there a plan to make saving checkpoints / weights possible even if the model is a subclassed model on Kaggle when using TPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844868,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/12/2020 22:36:13",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> hi. I forgot to thank you for the support during this competition. Regarding to saving feature it would be great to implement the possibility save temporary checkpoint produced by tf.compat.v1.estimator.tpu.TPUEstimator  . In this case kagglers can experiment with the variety of pretrained models from this repo: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844883,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "05/12/2020 22:51:36",
          "content": "<p>EfficientNet is now available in Tensorflow 2 Saved Model format from TF Hub: <a href=\"https://tfhub.dev/google/collections/efficientnet/1\">https://tfhub.dev/google/collections/efficientnet/1</a></p>\n\n<p>You can use it on Kaggle by copying the model to a Kaggle dataset and then using the following code:</p>\n\n<p>```\nfrom kaggle_datasets import KaggleDatasets\nGCS_PATH_SAVEDMODEL = KaggleDatasets().get_gcs_path('your-efficientnet-kaggle-dataset')\nlayer = tf.saved_model.load(GCS_PATH_SAVEDMODEL)</p>\n\n<h1>Cast the loaded model to a TFHub KerasLayer.</h1>\n\n<p>layer = hub.KerasLayer(layer, trainable=True)\n```</p>\n\n<p>(example code <a href=\"https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started\">here</a>, search for \"hub.KerasLayer\")</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844885,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "05/12/2020 22:53:26",
          "content": "<p>What exactly is not working with subclassed models ? Do you have an example ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844897,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/12/2020 23:10:24",
          "content": "<p>Thank you for the sharing this hub. However this hub is missing some large models that is interesting to train on TPU (B8, L2). Also the github repo contain more checkpoints with different options (noisy students, different augmentations like randaugment). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844908,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/12/2020 23:30:13",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a>, sorry as I understand subclassed models are Keras models? The repo I mentioned uses pure tensorflow without keras. The error I encountered is  “file system scheme [local] is not implement” because TPUEstimator.train() method try to save temp checkpoints during training. I can create the notebook tomorrow to demonstrate this issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844923,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "05/12/2020 23:45:41",
          "content": "<p>TPUEstimator is legacy TF1.x suff now. It won't get much updates.\nPure Tensorflow without Keras today means using custom training loop APIs like <a href=\"https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy#run\">TPUStrategy.run</a>. Like in this example: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a></p>\n\n<p>I understand your frustration about the pace of TF2 transition of Tensorflow model garden models. I will communicate this feedback back to the TF models team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844930,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "05/12/2020 23:52:35",
          "content": "<p>This repo seems to be the Keras version of EfficientNet. In the official Tensorflow model Garden:\n<a href=\"https://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/\">https://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/</a></p>\n\n<p>I see this line in the code:\n<code>class EfficientNet(tf.keras.Model):</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844940,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/13/2020 00:12:20",
          "content": "<p>Thank you for the kernel with the custom loop. I began my competition from this. I like tf 2 style, it is now looks like pytorch.  But I couldn’t  load pretrained tf 1. models under tf2 regime and was forced to use legacy methods to test official pretrained tf models. I will be waiting for release of pretrained TPU models for TF2. Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 845118,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "05/13/2020 04:02:19",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 845847,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "05/13/2020 13:02:13",
      "content": "<p>Congratulations <a href=\"/kromanov\">@kromanov</a> for holding the 1st position and really hope you don't get disqualified. </p>\n\n<p>I have few questions,\n1.  Can you please highlight more about the ensemble strategy you chose. You mention that you used a blend of your top 3 performing models. How exactly does that work?\n2. If you were facing issues in mixed precision training for TF-2.1 how were you able to solve this?</p>\n\n<p>Thanks and Congrats again. </p>",
      "votes": null,
      "replies": [
        {
          "id": 847313,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/14/2020 09:47:52",
          "content": "<p>hi <a href=\"/rhtsingh\">@rhtsingh</a> . Thank you.</p>\n\n<ol>\n<li>Strickly speaking it is just blending - I averaged the probabilities that was produced by these three models and then take the class with the max probability for every record</li>\n<li>I just make \"pip tensorflow -U\" command before launching computations. But if your goal is just mixed precision training, it is better to use trick described by <a href=\"/romanweilguny\">@romanweilguny</a> below.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 847320,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "05/14/2020 09:59:30",
          "content": "<p>Thanks a ton. 😊 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 847005,
      "author_name": "redwankarimsony",
      "author_url": "",
      "post_date": "05/14/2020 05:49:03",
      "content": "<p>Thank you <a href=\"/kromanov\">@kromanov</a>  for providing issues and their current workarounds of using TPUs. I hope you get your position what you deserve. I am new to kaggle. Great tips and tricks. 👍 👍 </p>",
      "votes": null,
      "replies": [
        {
          "id": 847358,
          "author_name": "kromanov",
          "author_url": "",
          "post_date": "05/14/2020 10:43:01",
          "content": "<p>Hi <a href=\"/redwankarimsony\">@redwankarimsony</a> . Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "843900": "**DISCLAIMER**\n\n*I don't read competition forum very often and as a result I missed  the information that usage of external datasets (I used the ones that were shared by @kirillblinov) was banned several days ago. As a result, there is very high probability that my solution is not eligible for prizes (but I ask confirm it from the organizer's side). However, I've got some interesting findings and maybe it would be interesting for other participants.*\n\nFirst of all I would like to thank organizers for the good opportunity to test and evaluate TPU technology in this competition. Also I would like to note some participants who contribute a much: @hengck23  for sharing fresh ideas, experiments and external datasets, @cdeotte  for publishing great notebook and adapting augmentations for TPU kernel and @kirillblinov  for processing and sharing external datasets.\n\nMy main goal for this competition was to understand how to use free TPU kernels and what are the limitations when it used for free. I decided to use tensorflow instead of pytorch especially due to the fact that there is a great repository with the models specially adapted for TPU kernels: https://github.com/tensorflow/tpu/tree/master/models/official. That’s was a plan and below are the results of my experiments.\n\n**Technical challenges with free tpu kernel**\n\n1. The most common issue I discovered is a *“file system scheme [local] is not implement”* error. You can’t save files (checkpoints, logs) on local drive or google drive. Instead, google cloud storage (GCS) is required for input data and output. Mainly due to this fact I didn’t use the models from tensorflow github repo - they built on tf 1.x (efficientnet models) and disk space on GCS is required for temporary checkpoints during training.\n2. The second issue was that some useful image processing utilites (e.g. augmentations) was removed from tensorflow 2.0 core and placed to tensorflow-addons (tfa). But tfa doesnt work properly with kaggle kernel TPU. This issue was solved by @cdeotte  who provided code for main augmentations that is compatible for kaggle TPU kernel.\n3. The third issue was tensorflow-related: for every kind of model you have to use different tf releases: for example classical resnets-like model was ported to tf 2.x while efficientnet models works only with tf 1.x . When I tried to use these models in tf.compat.v1 regime (with paid GCS), I’ve got multiple depreciation warnings and the training results were bad.\n4. Last but interesting issue - some useful features (like mixed precision training) were not working with TF 2.1 release on TPU. However it's work well on any TF 2.2 version. But TF 2.2 version was not stable and provided some other bugs (that I could fix)\n\nIn the end, I decided to use tf2 keras. For unknown reasons Keras let to save checkpoints on local kernel disk, have pretrained models for tf 2.X and it was widely used by other participants of this competition. \n\n**Solution tips and tricks**\n\n1. **Data.** I used external datasets prepared by @hengck23 and adapted by @kirillblinov . My experiment showed that openimage and inaturalist decrease accuracy, so I removed them. I used pictures with 512 and 331 sizes. Also with mixed precision training I could use large batches (192) for all model architectures.\n2. **Prepocessing.** I tried autoaugment, randaugment (the basis was official tensorflow implementation for tf 1.X and then adapted code for tf 2.X with TPU) and training without any augmentations. My experiments showed that with big datasets it is better don't use any augmentations.\n3. **Model selection.** I tried all major efficientnet architectures. The best for me were b5 and b6 architecture with noisy-student pre-trained weights (great thanks to @pavel92) . The classical resnets (50 and 101) as well as se-resnext 50-101 were not as well as efficientnet.\n4. **Model training.** I tried different strategies for cross-validation: 5-fold training, train-validation split, training without validation. As it was noted by many participants, training without validation provide higher score on public leaderboard and my final solution included the models that were trained without validation part. I used exponential decay LR scheduler with warmup. \n5. **Ensembling.** As i decided to use aggressive training strategy (no augmentations, no cross-validation), the ensembling was essential to avoid painful falling on private leaderboard. I trained best architectures (b5, b6, b7) with different picture sizes, try to combine effnet and seresnext models. The winner blend is three models B5-512, B5-331 and B6-512. This solution also provides my  best score on public LB.\n\nAs a result, I would say that though TPU kernels have some limitations, little bugs it is great opportunity for the researchers with limited computational capacities. In majority of computer vision competitions I participated with one laptop GPU (8gb) and google colab gpu. The model that usually calculated one day locally can be processed on TPU within one hour or faster. This technology is a great equalizer on kaggle competitions and it could facilitate deep learning researches as well.\n\nRegarding the usage of external dataset, I can confirm and assure that I don't know about the ban for usage the external datasets. Moreover, when I read the forum last time (one or two weeks ago) this topic was discussed and my understanding was that it'is ok to use this external dataset. It would be great to implement some alerting system for all competition participants to spread the news like this. \n\nHowever, despite the issue like that it was a great time to participate in this competition. Thank you very much for all participants and organizers.\n\n**UPD**: \n\nSome participants ask to me describe how did I find right model configuration without having validation part. Below is the high level explanation.\n\n1. At first step I buit multi-dimensional grid  with the parameters I would like to test:\n     a. **Model architectures** (Effnet b4, b5, b6, b7; Resnet 50, 101, Se-Resnext 50, 101, Inception)\n     b. **Pictures size** (512, 331, 224)\n     c. **Augmentations** (none, autoaugment, randaugment with different params). I selected these \n     three  types because it is easy to implement and they work well to beat some benchmarks in image classification tasks\n  d. **Losses** (cross-entropy, focal)\n e. **Optimizer and lr** \n2. Then I made grid-search of best combinations in manual mode. I don't need to test all possible combinations of my params to understand that pic size 224 provide lower score than 512. Validation part was \"valid\" folder and public leaderboard. \n3. Thanks to TPU and relatively small dataset, I could test a lot of hypothesis and considerably reduce my parameters grid. Then I added external datasets, **but keep the same validation part**. So after some experiments I could confirm that my validation score improved with external data and public score improved as well. At this step I defined best models (best combo of grid params)\n4.  The best models I test with different combination of external datasets and realize that it is possible to eliminate some of them while improving validation and public LB score\n5.  My best models was the models without augmentations (strictly speaking \"no augmentation\" training mode include random left-right flip). And I realized that training loss correlated very well with validation and public LB score. And adding validation part in training process increase Public LB considerably. Due to this I decided to retrain best models with validation part and use the models with the lowest training loss.\n6. Finally I've got about 10 models with good score. I tried different blend combinations. I  average probabilities because the majority of these models were effnets with 512 and 331 pic sizes. Nice to try here was to test different combination of voting but it was out of my goal to test TPU capabilites and I decided to use the simplest version of blending.",
    "844118": "thx for sharing!\n\nI did not have problems with mixed precision  training.\n\nOne has to define the last layer like this:\ntf.keras.layers.Dense(len(CLASSES), activation='softmax',**dtype='float32**')\n\nI think Chris pointed this out in a comment",
    "844163": "mixed precision training works for me also (tf 2.1)",
    "844482": "Yes, thank you for the comment. I had the issue described here: https://github.com/tensorflow/tensorflow/issues/34782\nAnd defining last layer as a float32 it was a possible workaround to solve it for tf 2.1 version. I experimented with different losses and layer configurations, so for me less problematic was to upgrade tf version (it was fixed in tf 2.2 pre-release).",
    "844803": "I confirm that saving model and checkpoints from TPU to local disk only works using the Keras HDF5 format. Saving in the newer \"Tensorflow Saved Model\" format from TPU to local disk is currently buggy. Saving in \"TF Saved Model\" format from TPU to GCS works though. You can use that on GCP. Not on Kaggle unfortunately because Kaggle notebooks do not have access to a writable GCS location.",
    "844823": "I just saw this answer after I published a new post. @mgornergoogle , is there a plan to make saving checkpoints / weights possible even if the model is a subclassed model on Kaggle when using TPU?",
    "844868": "mgornergoogle hi. I forgot to thank you for the support during this competition. Regarding to saving feature it would be great to implement the possibility save temporary checkpoint produced by tf.compat.v1.estimator.tpu.TPUEstimator  . In this case kagglers can experiment with the variety of pretrained models from this repo: https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet",
    "844883": "EfficientNet is now available in Tensorflow 2 Saved Model format from TF Hub: https://tfhub.dev/google/collections/efficientnet/1\n\nYou can use it on Kaggle by copying the model to a Kaggle dataset and then using the following code:\n\n```\nfrom kaggle_datasets import KaggleDatasets\nGCS_PATH_SAVEDMODEL = KaggleDatasets().get_gcs_path('your-efficientnet-kaggle-dataset')\nlayer = tf.saved_model.load(GCS_PATH_SAVEDMODEL)\n# Cast the loaded model to a TFHub KerasLayer.\nlayer = hub.KerasLayer(layer, trainable=True)\n```\n\n(example code [here](https://www.kaggle.com/kivlichangoogle/jigsaw-multilingual-getting-started), search for \"hub.KerasLayer\")",
    "844885": "What exactly is not working with subclassed models ? Do you have an example ?",
    "844897": "Thank you for the sharing this hub. However this hub is missing some large models that is interesting to train on TPU (B8, L2). Also the github repo contain more checkpoints with different options (noisy students, different augmentations like randaugment).",
    "844908": "mgornergoogle, sorry as I understand subclassed models are Keras models? The repo I mentioned uses pure tensorflow without keras. The error I encountered is  “file system scheme [local] is not implement” because TPUEstimator.train() method try to save temp checkpoints during training. I can create the notebook tomorrow to demonstrate this issue.",
    "844923": "TPUEstimator is legacy TF1.x suff now. It won't get much updates.\nPure Tensorflow without Keras today means using custom training loop APIs like [TPUStrategy.run](https://www.tensorflow.org/api_docs/python/tf/distribute/experimental/TPUStrategy#run). Like in this example: [TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\n\nI understand your frustration about the pace of TF2 transition of Tensorflow model garden models. I will communicate this feedback back to the TF models team.",
    "844930": "This repo seems to be the Keras version of EfficientNet. In the official Tensorflow model Garden:\nhttps://github.com/tensorflow/models/blob/master/official/vision/image_classification/efficientnet/\n\nI see this line in the code:\n`class EfficientNet(tf.keras.Model):`",
    "844940": "Thank you for the kernel with the custom loop. I began my competition from this. I like tf 2 style, it is now looks like pytorch.  But I couldn’t  load pretrained tf 1. models under tf2 regime and was forced to use legacy methods to test official pretrained tf models. I will be waiting for release of pretrained TPU models for TF2. Thank you.",
    "845118": "Thanks for sharing",
    "845847": "Congratulations @kromanov for holding the 1st position and really hope you don't get disqualified. \n\nI have few questions,\n1.  Can you please highlight more about the ensemble strategy you chose. You mention that you used a blend of your top 3 performing models. How exactly does that work?\n2. If you were facing issues in mixed precision training for TF-2.1 how were you able to solve this?\n\nThanks and Congrats again.",
    "847005": "Thank you @kromanov  for providing issues and their current workarounds of using TPUs. I hope you get your position what you deserve. I am new to kaggle. Great tips and tricks. 👍 👍",
    "847313": "hi @rhtsingh . Thank you.\n\n1. Strickly speaking it is just blending - I averaged the probabilities that was produced by these three models and then take the class with the max probability for every record\n2. I just make \"pip tensorflow -U\" command before launching computations. But if your goal is just mixed precision training, it is better to use trick described by @romanweilguny below.",
    "847320": "Thanks a ton. 😊",
    "847358": "Hi @redwankarimsony . Thank you!",
    "868868": "hi @romanweilguny @kromanov @yihdarshieh  is adding that snippet is all I need to do in TPU?\n\n```\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n\n```\nI wont have to add these lines in the beginning?",
    "868999": "yes - but I disabled XLA_ACCELERATE on tpu due to some problems...\n\nif XLA_ACCELERATE:\n    if not tpu:\n        tf.config.optimizer.set_jit(True)\n        print('Accelerated Linear Algebra enabled')",
    "869051": "I dont see anywhere in the code where you disable it. \nYou enable xla if there's is no TPU i.e. if the runtime is on GPU or CPU. \nCorrect me if Im wrong?",
    "869073": "I think it is disabled by default. Correct me if I'm wrong?",
    "869150": "rhtsingh , hi. In tpu you have to use special type: mixed_bfloat16. See below the snippet:\n```\nif USE_FLOAT16:\n    from tensorflow.keras.mixed_precision import experimental as mixed_precision\n    if tpu: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_bfloat16')\n    else: \n        policy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\n    mixed_precision.set_policy(policy)\n    print('Mixed precision enabled')\n\nif XLA_ACCELERATE:\n    tf.config.optimizer.set_jit(True)\n    print('Accelerated Linear Algebra enabled')\n```\nI enabled mixed precision and xla accelerate. With tf 2.2 it doesn't produce any errors",
    "869643": "Wow, amazing.! Many many thanks for all the knowledge you have shared here @kromanov and @romanweilguny  for helping people and your outstanding kernel."
  },
  "source": "meta"
}