{
  "id": 130703,
  "title": "Hint (for custom training loop) - Try to use drop_remainder when you change BATCH_SIZE",
  "url": "/competitions/flower-classification-with-tpus/discussion/130703",
  "author_name": "",
  "post_date": "2020-02-15T21:14:26.111488800Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>(Only for custom training loop)</p>\n\n<p>The original batch size is </p>\n\n<pre><code>BATCH_SIZE = 16 * strategy.num_replicas_in_sync\n</code></pre>\n\n<p>I changed it to</p>\n\n<pre><code>BATCH_SIZE = 8 * strategy.num_replicas_in_sync\n</code></pre>\n\n<p>And even in the 1st epoch, the model weights become <code>NaN</code>.</p>\n\n<p>I spent a lot of TPU quota to find the bug, and it turns out that, in the second case, we have</p>\n\n<pre><code>BATCH_SIZE = 64\n</code></pre>\n\n<p>Since the training dataset has <code>12753</code> training images, we have the last batch having size <code>12753 % 64 = 17</code>, which is quite small.</p>\n\n<p>With the default batch size, the last batch has <code>12753 % 128 = 81</code> examples, quite enough.</p>\n\n<p>So I used the following code to ignore the last batch</p>\n\n<pre><code>dataset = dataset.batch(BATCH_SIZE, drop_remainder=True)\n</code></pre>\n\n<p>And now I can continue to train the flowers!</p>",
  "messages": [
    {
      "id": "747012",
      "postDate": "02/15/2020 21:14:26",
      "content": "<p>(Only for custom training loop)</p>\n\n<p>The original batch size is </p>\n\n<pre><code>BATCH_SIZE = 16 * strategy.num_replicas_in_sync\n</code></pre>\n\n<p>I changed it to</p>\n\n<pre><code>BATCH_SIZE = 8 * strategy.num_replicas_in_sync\n</code></pre>\n\n<p>And even in the 1st epoch, the model weights become <code>NaN</code>.</p>\n\n<p>I spent a lot of TPU quota to find the bug, and it turns out that, in the second case, we have</p>\n\n<pre><code>BATCH_SIZE = 64\n</code></pre>\n\n<p>Since the training dataset has <code>12753</code> training images, we have the last batch having size <code>12753 % 64 = 17</code>, which is quite small.</p>\n\n<p>With the default batch size, the last batch has <code>12753 % 128 = 81</code> examples, quite enough.</p>\n\n<p>So I used the following code to ignore the last batch</p>\n\n<pre><code>dataset = dataset.batch(BATCH_SIZE, drop_remainder=True)\n</code></pre>\n\n<p>And now I can continue to train the flowers!</p>",
      "rawMarkdown": "(Only for custom training loop)\n\nThe original batch size is \n    \n    BATCH_SIZE = 16 * strategy.num_replicas_in_sync\n\nI changed it to\n\n    BATCH_SIZE = 8 * strategy.num_replicas_in_sync\n\nAnd even in the 1st epoch, the model weights become `NaN`.\n\nI spent a lot of TPU quota to find the bug, and it turns out that, in the second case, we have\n    \n    BATCH_SIZE = 64\n\nSince the training dataset has `12753` training images, we have the last batch having size `12753 % 64 = 17`, which is quite small.\n\nWith the default batch size, the last batch has `12753 % 128 = 81` examples, quite enough.\n\nSo I used the following code to ignore the last batch\n        \n    dataset = dataset.batch(BATCH_SIZE, drop_remainder=True)\n\nAnd now I can continue to train the flowers!",
      "votes": null
    },
    {
      "id": "747024",
      "postDate": "02/15/2020 21:33:57",
      "content": "<p>Nice catch <a href=\"/yihdarshieh\">@yihdarshieh</a> , I also had some issues while changing batch size.</p>",
      "rawMarkdown": "Nice catch @yihdarshieh , I also had some issues while changing batch size.",
      "votes": null
    },
    {
      "id": "749580",
      "postDate": "02/18/2020 19:43:20",
      "content": "<p>I'm not sure this is accurate. The training dataset is looped and then batched so it is in effect infinite and there is no \"remainder\".</p>",
      "rawMarkdown": "I'm not sure this is accurate. The training dataset is looped and then batched so it is in effect infinite and there is no \"remainder\".",
      "votes": null
    },
    {
      "id": "749667",
      "postDate": "02/18/2020 20:56:16",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , I am using custom training loop, I don't use .repeat(). I should precise it.</p>",
      "rawMarkdown": "mgornergoogle , I am using custom training loop, I don't use .repeat(). I should precise it.",
      "votes": null
    },
    {
      "id": "749670",
      "postDate": "02/18/2020 20:59:17",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "749788",
      "postDate": "02/18/2020 22:39:57",
      "content": "<p>OK, understood. Then an even better solution than drop_remainder (which removes useful data), is to repeat the dataset.</p>",
      "rawMarkdown": "OK, understood. Then an even better solution than drop_remainder (which removes useful data), is to repeat the dataset.",
      "votes": null
    },
    {
      "id": "749790",
      "postDate": "02/18/2020 22:42:39",
      "content": "<p>A bit more info:\nTPUs used to be unable to unable to handle the last partial batch and required drop_remainder=True. This was in TF 1.14. XLA compilation would fail. In TF 2.1, support for the last partial batch was added.</p>\n\n<p>This however does not change the underlying math and it may very well happen that if the batch is tiny, some computation in your model will result in a NaN.</p>",
      "rawMarkdown": "A bit more info:\nTPUs used to be unable to unable to handle the last partial batch and required drop_remainder=True. This was in TF 1.14. XLA compilation would fail. In TF 2.1, support for the last partial batch was added.\n\nThis however does not change the underlying math and it may very well happen that if the batch is tiny, some computation in your model will result in a NaN.",
      "votes": null
    },
    {
      "id": "749806",
      "postDate": "02/18/2020 23:05:10",
      "content": "<p>True. I need to find a good way to do custom training loop with dataset.repeat() -- Look like I have to decide when an epoch is finished by myself.</p>",
      "rawMarkdown": "True. I need to find a good way to do custom training loop with dataset.repeat() -- Look like I have to decide when an epoch is finished by myself.",
      "votes": null
    },
    {
      "id": "749815",
      "postDate": "02/18/2020 23:16:30",
      "content": "<p>Yes, it is a bit unfortunate that the tf.data.Dataset API does not expose the pre-repeat cardinality of the dataset. I guess it is not an easy problem to solve in all cases.</p>\n\n<p>For this dataset, the solution I implemented in the getting started notebook is to write the number of records in the name of the .tfrec file. Counting records can then be done by parsing the file names. The code for this is a regexp in the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">getting started notebook</a>, function <code>count_data_items(filenames)</code>. Yeah, a bit hacky, but very convenient when eyeballing the dataset.</p>",
      "rawMarkdown": "Yes, it is a bit unfortunate that the tf.data.Dataset API does not expose the pre-repeat cardinality of the dataset. I guess it is not an easy problem to solve in all cases.\n\nFor this dataset, the solution I implemented in the getting started notebook is to write the number of records in the name of the .tfrec file. Counting records can then be done by parsing the file names. The code for this is a regexp in the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/), function `count_data_items(filenames)`. Yeah, a bit hacky, but very convenient when eyeballing the dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 747024,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "02/15/2020 21:33:57",
      "content": "<p>Nice catch <a href=\"/yihdarshieh\">@yihdarshieh</a> , I also had some issues while changing batch size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 749580,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/18/2020 19:43:20",
      "content": "<p>I'm not sure this is accurate. The training dataset is looped and then batched so it is in effect infinite and there is no \"remainder\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 749667,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/18/2020 20:56:16",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , I am using custom training loop, I don't use .repeat(). I should precise it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749670,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/18/2020 20:59:17",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 749788,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 22:39:57",
          "content": "<p>OK, understood. Then an even better solution than drop_remainder (which removes useful data), is to repeat the dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749790,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 22:42:39",
          "content": "<p>A bit more info:\nTPUs used to be unable to unable to handle the last partial batch and required drop_remainder=True. This was in TF 1.14. XLA compilation would fail. In TF 2.1, support for the last partial batch was added.</p>\n\n<p>This however does not change the underlying math and it may very well happen that if the batch is tiny, some computation in your model will result in a NaN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749806,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/18/2020 23:05:10",
          "content": "<p>True. I need to find a good way to do custom training loop with dataset.repeat() -- Look like I have to decide when an epoch is finished by myself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749815,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 23:16:30",
          "content": "<p>Yes, it is a bit unfortunate that the tf.data.Dataset API does not expose the pre-repeat cardinality of the dataset. I guess it is not an easy problem to solve in all cases.</p>\n\n<p>For this dataset, the solution I implemented in the getting started notebook is to write the number of records in the name of the .tfrec file. Counting records can then be done by parsing the file names. The code for this is a regexp in the <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/\">getting started notebook</a>, function <code>count_data_items(filenames)</code>. Yeah, a bit hacky, but very convenient when eyeballing the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "747012": "(Only for custom training loop)\n\nThe original batch size is \n    \n    BATCH_SIZE = 16 * strategy.num_replicas_in_sync\n\nI changed it to\n\n    BATCH_SIZE = 8 * strategy.num_replicas_in_sync\n\nAnd even in the 1st epoch, the model weights become `NaN`.\n\nI spent a lot of TPU quota to find the bug, and it turns out that, in the second case, we have\n    \n    BATCH_SIZE = 64\n\nSince the training dataset has `12753` training images, we have the last batch having size `12753 % 64 = 17`, which is quite small.\n\nWith the default batch size, the last batch has `12753 % 128 = 81` examples, quite enough.\n\nSo I used the following code to ignore the last batch\n        \n    dataset = dataset.batch(BATCH_SIZE, drop_remainder=True)\n\nAnd now I can continue to train the flowers!",
    "747024": "Nice catch @yihdarshieh , I also had some issues while changing batch size.",
    "749580": "I'm not sure this is accurate. The training dataset is looped and then batched so it is in effect infinite and there is no \"remainder\".",
    "749667": "mgornergoogle , I am using custom training loop, I don't use .repeat(). I should precise it.",
    "749670": "",
    "749788": "OK, understood. Then an even better solution than drop_remainder (which removes useful data), is to repeat the dataset.",
    "749790": "A bit more info:\nTPUs used to be unable to unable to handle the last partial batch and required drop_remainder=True. This was in TF 1.14. XLA compilation would fail. In TF 2.1, support for the last partial batch was added.\n\nThis however does not change the underlying math and it may very well happen that if the batch is tiny, some computation in your model will result in a NaN.",
    "749806": "True. I need to find a good way to do custom training loop with dataset.repeat() -- Look like I have to decide when an epoch is finished by myself.",
    "749815": "Yes, it is a bit unfortunate that the tf.data.Dataset API does not expose the pre-repeat cardinality of the dataset. I guess it is not an easy problem to solve in all cases.\n\nFor this dataset, the solution I implemented in the getting started notebook is to write the number of records in the name of the .tfrec file. Counting records can then be done by parsing the file names. The code for this is a regexp in the [getting started notebook](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu/), function `count_data_items(filenames)`. Yeah, a bit hacky, but very convenient when eyeballing the dataset."
  },
  "source": "meta"
}