{
  "id": 218062,
  "title": "Model fitting very slow using TPU",
  "url": "/competitions/tpu-getting-started/discussion/218062",
  "author_name": "Yann Garcia",
  "post_date": "2021-02-09T06:09:06.407000",
  "votes": 0,
  "comment_count": 4,
  "views": null,
  "content": "<p>Hello,</p>\n<p>First, thanks a lot for this very nice compete.</p>\n<p>My issue is when I execute the fit() method of my model, it take a very long time (I stopped ater 5 hours, only 22 on 32 batches where done).<br>\nI compared my code wit some another participants (thanks to share your notebooks) and I didn't find the issue.<br>\nCan someone indicate me how to proceed to debug with TPU, please? Below, are some details about what I did.</p>\n<p>Thanks a lot</p>\n<p><strong>Notes:</strong><br>\n1) The TPU detection works fine:</p>\n<pre><code>----------------------------- kaggle_tpu_detection -----------------------------\nkaggle_tpu_detection: Running on TPU  grpc://10.0.0.2:8470\nkaggle_tpu_detection: replica=8\nkaggle_tpu_detection: global_path:  gs://kds-c4925d8453be6e0b7b5678d01626bde5eeef16a4885151456ef0d63f\nkaggle_tpu_detection Done\n</code></pre>\n<p>2) My model layout looks correct:</p>\n<pre><code>Downloading data from https://storage.googleapis.com/tensorflow/keras-applications/vgg16/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5\n58892288/58889256 [==============================] - 0s 0us/step\nModel: \"sequential\"\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\nvgg16 (Functional)           (None, 16, 16, 512)       14714688  \n_________________________________________________________________\nglobal_average_pooling2d (Gl (None, 512)               0         \n_________________________________________________________________\ndropout (Dropout)            (None, 512)               0         \n_________________________________________________________________\ndense (Dense)                (None, 104)               53352     \n=================================================================\nTotal params: 14,768,040\nTrainable params: 53,352\nNon-trainable params: 14,714,688\n</code></pre>\n<p>3) My datasets too</p>\n<pre><code>_________________________________________________________________\nDataset: 12753 training images, 3712 validation images, 7382 unlabeled test images\n\n4) I'm using tf.keras.callbacks.LearningRateScheduler with an exponencial curve:\n`Learning rate schedule: 0.0001 to 0.0005 to 0.000101`\n</code></pre>\n<p>5) fit() details:<br>\nBATCH_SIZE ==&gt; 128<br>\nSTEPS_PER_EPOCH ==&gt; 99<br>\nDL_EPOCH_NUM ==&gt; 32</p>",
  "messages": [
    {
      "id": 1253362,
      "postDate": "2021-03-26T16:40:50.497Z",
      "content": "<p>Are you using <code>strategy.scope()</code>?   While using distribution strategies, the variables created within the strategy's scope will be replicated across all the replicas and can be kept in sync using all-reduce algorithms.   </p>\n<p>The following is from <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">my notebook </a>, where I built my model within the strategy scope.  It scores &gt; 0.9.  </p>\n<p>I built my notebook with the basics from <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>'s <a href=\"https://www.kaggle.com/philculliton/a-simple-petals-tf-2-2-notebook\" target=\"_blank\">A Simple Petals TF 2.2 notebook</a>.  Phil's notebook is very helpful to me and is a great place to start, to learn and build.  </p>\n<pre><code>with strategy.scope():    \n    pretrained_model = tf.keras.applications.DenseNet201(\n        weights='imagenet', \n        include_top=False ,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = True # transfer learning\n    model = tf.keras.Sequential([\n        pretrained_model,\n        # the following is commented out: layers applied to the training set instead \n        #data_aug_layers,                    \n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dropout(0.2),  \n        tf.keras.layers.Dense(104, kernel_regularizer=regularizers.L2(0.00011), \n                              activation='softmax')\n    ])\n</code></pre>\n<p>Please take a look at my <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">notebook</a> and drop me some comments.  It shows how to visualize images and implement data augmentation.  I finally added external data, which should have been the first  and easiest step to improve the performance of my model.  I've checked that the rules of this competition allows these external data.  I didn't come across the data until now.    </p>",
      "rawMarkdown": "Are you using `strategy.scope()`?   While using distribution strategies, the variables created within the strategy's scope will be replicated across all the replicas and can be kept in sync using all-reduce algorithms.   \n\nThe following is from [my notebook ](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification), where I built my model within the strategy scope.  It scores > 0.9.  \n\nI built my notebook with the basics from @philculliton's [A Simple Petals TF 2.2 notebook](https://www.kaggle.com/philculliton/a-simple-petals-tf-2-2-notebook).  Phil's notebook is very helpful to me and is a great place to start, to learn and build.  \n\n```\nwith strategy.scope():    \n    pretrained_model = tf.keras.applications.DenseNet201(\n        weights='imagenet', \n        include_top=False ,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = True # transfer learning\n    model = tf.keras.Sequential([\n        pretrained_model,\n        # the following is commented out: layers applied to the training set instead \n        #data_aug_layers,                    \n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dropout(0.2),  \n        tf.keras.layers.Dense(104, kernel_regularizer=regularizers.L2(0.00011), \n                              activation='softmax')\n    ])\n```\nPlease take a look at my [notebook](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification) and drop me some comments.  It shows how to visualize images and implement data augmentation.  I finally added external data, which should have been the first  and easiest step to improve the performance of my model.  I've checked that the rules of this competition allows these external data.  I didn't come across the data until now.    ",
      "replies": [
        {
          "id": 1256971,
          "postDate": "2021-03-30T12:22:37.947Z",
          "content": "<p>Hello Saukha,</p>\n<p>Thanks a lot for your comments.</p>\n<p>Yes I was using the TPU.<br>\nIt took a very long time because of I applied augmentation on both Training and Validation datasets. I think this was my mistake.<br>\nAfter fixing it, I got some more reasonable execution time.</p>\n<p>Thanks a lot,</p>\n<p>Have a nice Easter,</p>\n<p>Yann</p>",
          "rawMarkdown": "Hello Saukha,\n\nThanks a lot for your comments.\n\nYes I was using the TPU.\nIt took a very long time because of I applied augmentation on both Training and Validation datasets. I think this was my mistake.\nAfter fixing it, I got some more reasonable execution time.\n\nThanks a lot,\n\nHave a nice Easter,\n\nYann",
          "votes": 2
        },
        {
          "id": 1283295,
          "postDate": "2021-04-24T19:34:12.307Z",
          "content": "<p>I found that adding a batch norm layer to my model speeds up the training time for my model quite a bit.   This helps a lot since I'm limited to the 30-hour limit for TPU time, especially when training with more external data and very deep model, like EfficientNetB7 that I'm using.   See the following lessons from Andrew Ng's <a href=\"https://www.youtube.com/channel/UCcIXc5mJsHVYTZR1maL5l9w\" target=\"_blank\">DeepLearningAI</a> YouTube channel:  </p>\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=em6dfRxYkYU\" target=\"_blank\">Fitting Batch Norm Into Neural Networks (C2W3L05)</a>  </li>\n<li><a href=\"https://www.youtube.com/watch?v=nUUqwaxLnWs\" target=\"_blank\">Why Does Batch Norm Work? (C2W3L06)</a>  </li>\n</ul>",
          "rawMarkdown": "I found that adding a batch norm layer to my model speeds up the training time for my model quite a bit.   This helps a lot since I'm limited to the 30-hour limit for TPU time, especially when training with more external data and very deep model, like EfficientNetB7 that I'm using.   See the following lessons from Andrew Ng's [DeepLearningAI](https://www.youtube.com/channel/UCcIXc5mJsHVYTZR1maL5l9w) YouTube channel:  \n- [Fitting Batch Norm Into Neural Networks (C2W3L05)](https://www.youtube.com/watch?v=em6dfRxYkYU)  \n- [Why Does Batch Norm Work? (C2W3L06)](https://www.youtube.com/watch?v=nUUqwaxLnWs)  ",
          "votes": 1
        },
        {
          "id": 1285803,
          "postDate": "2021-04-27T09:14:25.913Z",
          "content": "<p>Hello Sau Kha,</p>\n<p>Thanks a lot, I'm going to check in this way ;)</p>\n<p>Yann</p>",
          "rawMarkdown": "Hello Sau Kha,\n\nThanks a lot, I'm going to check in this way ;)\n\nYann"
        }
      ]
    },
    {
      "id": 1192470,
      "postDate": "2021-02-09T06:09:06.407Z",
      "content": "<p>Hello,</p>\n<p>First, thanks a lot for this very nice compete.</p>\n<p>My issue is when I execute the fit() method of my model, it take a very long time (I stopped ater 5 hours, only 22 on 32 batches where done).<br>\nI compared my code wit some another participants (thanks to share your notebooks) and I didn't find the issue.<br>\nCan someone indicate me how to proceed to debug with TPU, please? Below, are some details about what I did.</p>\n<p>Thanks a lot</p>\n<p><strong>Notes:</strong><br>\n1) The TPU detection works fine:</p>\n<pre><code>----------------------------- kaggle_tpu_detection -----------------------------\nkaggle_tpu_detection: Running on TPU  grpc://10.0.0.2:8470\nkaggle_tpu_detection: replica=8\nkaggle_tpu_detection: global_path:  gs://kds-c4925d8453be6e0b7b5678d01626bde5eeef16a4885151456ef0d63f\nkaggle_tpu_detection Done\n</code></pre>\n<p>2) My model layout looks correct:</p>\n<pre><code>Downloading data from https://storage.googleapis.com/tensorflow/keras-applications/vgg16/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5\n58892288/58889256 [==============================] - 0s 0us/step\nModel: \"sequential\"\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\nvgg16 (Functional)           (None, 16, 16, 512)       14714688  \n_________________________________________________________________\nglobal_average_pooling2d (Gl (None, 512)               0         \n_________________________________________________________________\ndropout (Dropout)            (None, 512)               0         \n_________________________________________________________________\ndense (Dense)                (None, 104)               53352     \n=================================================================\nTotal params: 14,768,040\nTrainable params: 53,352\nNon-trainable params: 14,714,688\n</code></pre>\n<p>3) My datasets too</p>\n<pre><code>_________________________________________________________________\nDataset: 12753 training images, 3712 validation images, 7382 unlabeled test images\n\n4) I'm using tf.keras.callbacks.LearningRateScheduler with an exponencial curve:\n`Learning rate schedule: 0.0001 to 0.0005 to 0.000101`\n</code></pre>\n<p>5) fit() details:<br>\nBATCH_SIZE ==&gt; 128<br>\nSTEPS_PER_EPOCH ==&gt; 99<br>\nDL_EPOCH_NUM ==&gt; 32</p>",
      "rawMarkdown": "Hello,\n\nFirst, thanks a lot for this very nice compete.\n\nMy issue is when I execute the fit() method of my model, it take a very long time (I stopped ater 5 hours, only 22 on 32 batches where done).\nI compared my code wit some another participants (thanks to share your notebooks) and I didn't find the issue.\nCan someone indicate me how to proceed to debug with TPU, please? Below, are some details about what I did.\n\nThanks a lot\n\n**Notes:**\n1) The TPU detection works fine:\n```\n----------------------------- kaggle_tpu_detection -----------------------------\nkaggle_tpu_detection: Running on TPU  grpc://10.0.0.2:8470\nkaggle_tpu_detection: replica=8\nkaggle_tpu_detection: global_path:  gs://kds-c4925d8453be6e0b7b5678d01626bde5eeef16a4885151456ef0d63f\nkaggle_tpu_detection Done\n```\n\n2) My model layout looks correct:\n```\nDownloading data from https://storage.googleapis.com/tensorflow/keras-applications/vgg16/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5\n58892288/58889256 [==============================] - 0s 0us/step\nModel: \"sequential\"\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\nvgg16 (Functional)           (None, 16, 16, 512)       14714688  \n_________________________________________________________________\nglobal_average_pooling2d (Gl (None, 512)               0         \n_________________________________________________________________\ndropout (Dropout)            (None, 512)               0         \n_________________________________________________________________\ndense (Dense)                (None, 104)               53352     \n=================================================================\nTotal params: 14,768,040\nTrainable params: 53,352\nNon-trainable params: 14,714,688\n```\n3) My datasets too\n```\n_________________________________________________________________\nDataset: 12753 training images, 3712 validation images, 7382 unlabeled test images\n\n4) I'm using tf.keras.callbacks.LearningRateScheduler with an exponencial curve:\n`Learning rate schedule: 0.0001 to 0.0005 to 0.000101`\n```\n5) fit() details:\nBATCH_SIZE ==> 128\nSTEPS_PER_EPOCH ==> 99\nDL_EPOCH_NUM ==> 32\n"
    }
  ],
  "comments": [
    {
      "id": 1253362,
      "author_name": "Sau Kha",
      "author_url": "",
      "post_date": "2021-03-26T16:40:50.497000",
      "content": "<p>Are you using <code>strategy.scope()</code>?   While using distribution strategies, the variables created within the strategy's scope will be replicated across all the replicas and can be kept in sync using all-reduce algorithms.   </p>\n<p>The following is from <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">my notebook </a>, where I built my model within the strategy scope.  It scores &gt; 0.9.  </p>\n<p>I built my notebook with the basics from <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>'s <a href=\"https://www.kaggle.com/philculliton/a-simple-petals-tf-2-2-notebook\" target=\"_blank\">A Simple Petals TF 2.2 notebook</a>.  Phil's notebook is very helpful to me and is a great place to start, to learn and build.  </p>\n<pre><code>with strategy.scope():    \n    pretrained_model = tf.keras.applications.DenseNet201(\n        weights='imagenet', \n        include_top=False ,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = True # transfer learning\n    model = tf.keras.Sequential([\n        pretrained_model,\n        # the following is commented out: layers applied to the training set instead \n        #data_aug_layers,                    \n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dropout(0.2),  \n        tf.keras.layers.Dense(104, kernel_regularizer=regularizers.L2(0.00011), \n                              activation='softmax')\n    ])\n</code></pre>\n<p>Please take a look at my <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">notebook</a> and drop me some comments.  It shows how to visualize images and implement data augmentation.  I finally added external data, which should have been the first  and easiest step to improve the performance of my model.  I've checked that the rules of this competition allows these external data.  I didn't come across the data until now.    </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1256971,
          "author_name": "Yann Garcia",
          "author_url": "",
          "post_date": "2021-03-30T12:22:37.947000",
          "content": "<p>Hello Saukha,</p>\n<p>Thanks a lot for your comments.</p>\n<p>Yes I was using the TPU.<br>\nIt took a very long time because of I applied augmentation on both Training and Validation datasets. I think this was my mistake.<br>\nAfter fixing it, I got some more reasonable execution time.</p>\n<p>Thanks a lot,</p>\n<p>Have a nice Easter,</p>\n<p>Yann</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1283295,
          "author_name": "Sau Kha",
          "author_url": "",
          "post_date": "2021-04-24T19:34:12.307000",
          "content": "<p>I found that adding a batch norm layer to my model speeds up the training time for my model quite a bit.   This helps a lot since I'm limited to the 30-hour limit for TPU time, especially when training with more external data and very deep model, like EfficientNetB7 that I'm using.   See the following lessons from Andrew Ng's <a href=\"https://www.youtube.com/channel/UCcIXc5mJsHVYTZR1maL5l9w\" target=\"_blank\">DeepLearningAI</a> YouTube channel:  </p>\n<ul>\n<li><a href=\"https://www.youtube.com/watch?v=em6dfRxYkYU\" target=\"_blank\">Fitting Batch Norm Into Neural Networks (C2W3L05)</a>  </li>\n<li><a href=\"https://www.youtube.com/watch?v=nUUqwaxLnWs\" target=\"_blank\">Why Does Batch Norm Work? (C2W3L06)</a>  </li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1285803,
          "author_name": "Yann Garcia",
          "author_url": "",
          "post_date": "2021-04-27T09:14:25.913000",
          "content": "<p>Hello Sau Kha,</p>\n<p>Thanks a lot, I'm going to check in this way ;)</p>\n<p>Yann</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1253362": "Are you using `strategy.scope()`?   While using distribution strategies, the variables created within the strategy's scope will be replicated across all the replicas and can be kept in sync using all-reduce algorithms.   \n\nThe following is from [my notebook ](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification), where I built my model within the strategy scope.  It scores > 0.9.  \n\nI built my notebook with the basics from @philculliton's [A Simple Petals TF 2.2 notebook](https://www.kaggle.com/philculliton/a-simple-petals-tf-2-2-notebook).  Phil's notebook is very helpful to me and is a great place to start, to learn and build.  \n\n```\nwith strategy.scope():    \n    pretrained_model = tf.keras.applications.DenseNet201(\n        weights='imagenet', \n        include_top=False ,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = True # transfer learning\n    model = tf.keras.Sequential([\n        pretrained_model,\n        # the following is commented out: layers applied to the training set instead \n        #data_aug_layers,                    \n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dropout(0.2),  \n        tf.keras.layers.Dense(104, kernel_regularizer=regularizers.L2(0.00011), \n                              activation='softmax')\n    ])\n```\nPlease take a look at my [notebook](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification) and drop me some comments.  It shows how to visualize images and implement data augmentation.  I finally added external data, which should have been the first  and easiest step to improve the performance of my model.  I've checked that the rules of this competition allows these external data.  I didn't come across the data until now.    ",
    "1192470": "Hello,\n\nFirst, thanks a lot for this very nice compete.\n\nMy issue is when I execute the fit() method of my model, it take a very long time (I stopped ater 5 hours, only 22 on 32 batches where done).\nI compared my code wit some another participants (thanks to share your notebooks) and I didn't find the issue.\nCan someone indicate me how to proceed to debug with TPU, please? Below, are some details about what I did.\n\nThanks a lot\n\n**Notes:**\n1) The TPU detection works fine:\n```\n----------------------------- kaggle_tpu_detection -----------------------------\nkaggle_tpu_detection: Running on TPU  grpc://10.0.0.2:8470\nkaggle_tpu_detection: replica=8\nkaggle_tpu_detection: global_path:  gs://kds-c4925d8453be6e0b7b5678d01626bde5eeef16a4885151456ef0d63f\nkaggle_tpu_detection Done\n```\n\n2) My model layout looks correct:\n```\nDownloading data from https://storage.googleapis.com/tensorflow/keras-applications/vgg16/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5\n58892288/58889256 [==============================] - 0s 0us/step\nModel: \"sequential\"\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\nvgg16 (Functional)           (None, 16, 16, 512)       14714688  \n_________________________________________________________________\nglobal_average_pooling2d (Gl (None, 512)               0         \n_________________________________________________________________\ndropout (Dropout)            (None, 512)               0         \n_________________________________________________________________\ndense (Dense)                (None, 104)               53352     \n=================================================================\nTotal params: 14,768,040\nTrainable params: 53,352\nNon-trainable params: 14,714,688\n```\n3) My datasets too\n```\n_________________________________________________________________\nDataset: 12753 training images, 3712 validation images, 7382 unlabeled test images\n\n4) I'm using tf.keras.callbacks.LearningRateScheduler with an exponencial curve:\n`Learning rate schedule: 0.0001 to 0.0005 to 0.000101`\n```\n5) fit() details:\nBATCH_SIZE ==> 128\nSTEPS_PER_EPOCH ==> 99\nDL_EPOCH_NUM ==> 32\n"
  }
}