{
  "id": 130422,
  "title": "Custom training with TPU",
  "url": "/competitions/flower-classification-with-tpus/discussion/130422",
  "author_name": "",
  "post_date": "2020-02-14T02:45:41.982161500Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, I just published my kernel for demonstrating how to do custom training with TPU -- <a href=\"https://www.kaggle.com/yihdarshieh/custom-training-with-tpu?scriptVersionId=28624481\">Custom training with TPU</a>.</p>\n\n<p>Hope you find it helpful!</p>",
  "messages": [
    {
      "id": "745630",
      "postDate": "02/14/2020 02:45:41",
      "content": "<p>Hi, I just published my kernel for demonstrating how to do custom training with TPU -- <a href=\"https://www.kaggle.com/yihdarshieh/custom-training-with-tpu?scriptVersionId=28624481\">Custom training with TPU</a>.</p>\n\n<p>Hope you find it helpful!</p>",
      "rawMarkdown": "Hi, I just published my kernel for demonstrating how to do custom training with TPU -- [Custom training with TPU](https://www.kaggle.com/yihdarshieh/custom-training-with-tpu?scriptVersionId=28624481).\n\nHope you find it helpful!",
      "votes": null
    },
    {
      "id": "745633",
      "postDate": "02/14/2020 02:51:14",
      "content": "<p>I personally have question for one part in my kernel</p>\n\n<pre><code>for epoch in range(last_epoch, EPOCHS):\n\n    epoch_start_time = datetime.datetime.now()\n\n    # We need to shuffle the training dataset in every epoch.\n    # I don't know if there is a better way than the following way.\n\n    training_dataset = get_training_dataset()\n    training_dist_dataset = strategy.experimental_distribute_dataset(training_dataset)\n</code></pre>\n\n<p>If we don't call get_training_dataset(), and we just iterate again over the same object <code>training_dataset</code>, do the elements appear in the same order as in the previous iteratings?</p>\n\n<p>If yes, we should avoid this situation. Other than calling <code>get_training_dataset()</code>, is there a better way??</p>",
      "rawMarkdown": "I personally have question for one part in my kernel\n\n    for epoch in range(last_epoch, EPOCHS):\n    \n        epoch_start_time = datetime.datetime.now()\n\n        # We need to shuffle the training dataset in every epoch.\n        # I don't know if there is a better way than the following way.\n        \n        training_dataset = get_training_dataset()\n        training_dist_dataset = strategy.experimental_distribute_dataset(training_dataset)\n\nIf we don't call get_training_dataset(), and we just iterate again over the same object ` training_dataset`, do the elements appear in the same order as in the previous iteratings?\n\nIf yes, we should avoid this situation. Other than calling `get_training_dataset()`, is there a better way??",
      "votes": null
    },
    {
      "id": "746162",
      "postDate": "02/14/2020 17:33:21",
      "content": "<p>Congrats for being the first to experiment with distributed custom training loops !</p>\n\n<p>The training dataset returned by <code>get_training_dataset()</code> is iterable. The training loop should iterate on it. I have a custom training loop sample here:\n<a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb</a></p>",
      "rawMarkdown": "Congrats for being the first to experiment with distributed custom training loops !\n\nThe training dataset returned by `get_training_dataset()` is iterable. The training loop should iterate on it. I have a custom training loop sample here:\nhttps://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb",
      "votes": null
    },
    {
      "id": "749758",
      "postDate": "02/18/2020 22:21:29",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , I observed that, in your notebook above, you use <code>tf.reduce_mean</code> in the loss function and the loss function is called inside <code>train_step</code>.</p>\n\n<p>However, in [https://www.tensorflow.org/tutorials/distribute/custom_training#define_the_loss_function](Custom training with tf.distribute.Strategy), it specifically says the following:</p>\n\n<pre><code>If you're writing a custom training loop, as in this tutorial, you should sum the per example losses and divide the sum by the GLOBAL_BATCH_SIZE: scale_loss = tf.reduce_sum(loss) * (1. / GLOBAL_BATCH_SIZE) or you can use tf.nn.compute_average_loss which takes the per example loss, optional sample weights, and GLOBAL_BATCH_SIZE as arguments and returns the scaled loss.\n</code></pre>\n\n<p>and</p>\n\n<pre><code>Using tf.reduce_mean is not recommended. Doing so divides the loss by actual per replica batch size which may vary step to step.\n</code></pre>\n\n<p>and </p>\n\n<pre><code>Why do this?\n\nThis needs to be done because after the gradients are calculated on each replica, they are synced across the replicas by summing them.\n</code></pre>\n\n<p>I think in your notebook about custom training, the loss is larger than the correct one. (If divide your loss function by the number of replica, which is 8 there, it should be correct).</p>\n\n<p>I didn't run several times, but I tried once to change your loss function and run it. The following is the training results of the last epochs, which has training/val acc ~0.80, which a bit better than the result in the original notebook (0.77/0.75). Be careful, the loss value  in the reported is not the correct value to be reported. We should sum the loss values from all the replica, but I didn't spent my time to do this.</p>\n\n<pre><code>EPOCH:  51\nloss:  0.078226216 , accuracy_:  0.8203125  , val_loss:  0.08729416  , val_acc_:  0.7246094  , lr:  0.012249263566363598\n=======================\nEPOCH:  52\nloss:  0.063389845 , accuracy_:  0.81827444  , val_loss:  0.08148105  , val_acc_:  0.7871094  , lr:  0.01168680038804542\n=======================\nEPOCH:  53\nloss:  0.057514176 , accuracy_:  0.8226902  , val_loss:  0.07985326  , val_acc_:  0.7949219  , lr:  0.011152460368643147\n=======================\nEPOCH:  54\nloss:  0.07713041 , accuracy_:  0.8141984  , val_loss:  0.08316609  , val_acc_:  0.7890625  , lr:  0.01064483735021099\n=======================\nEPOCH:  55\nloss:  0.0724027 , accuracy_:  0.8345788  , val_loss:  0.05968307  , val_acc_:  0.8203125  , lr:  0.010162595482700439\n=======================\nEPOCH:  56\nloss:  0.066552304 , accuracy_:  0.8213315  , val_loss:  0.07661423  , val_acc_:  0.80078125  , lr:  0.009704465708565417\n=======================\nEPOCH:  57\nloss:  0.07287082 , accuracy_:  0.8199728  , val_loss:  0.07448967  , val_acc_:  0.8222656  , lr:  0.009269242423137147\n=======================\nEPOCH:  58\nloss:  0.053799704 , accuracy_:  0.83186144  , val_loss:  0.11246389  , val_acc_:  0.73828125  , lr:  0.008855780301980289\n=======================\nEPOCH:  59\nloss:  0.051523175 , accuracy_:  0.8413723  , val_loss:  0.07519884  , val_acc_:  0.7910156  , lr:  0.008462991286881274\n=======================\nEPOCH:  60\nloss:  0.052226786 , accuracy_:  0.84714675  , val_loss:  0.080301926  , val_acc_:  0.8046875  , lr:  0.008089841722537211\n</code></pre>",
      "rawMarkdown": "mgornergoogle , I observed that, in your notebook above, you use `tf.reduce_mean` in the loss function and the loss function is called inside `train_step`.\n\nHowever, in [https://www.tensorflow.org/tutorials/distribute/custom_training#define_the_loss_function](Custom training with tf.distribute.Strategy), it specifically says the following:\n\n    If you're writing a custom training loop, as in this tutorial, you should sum the per example losses and divide the sum by the GLOBAL_BATCH_SIZE: scale_loss = tf.reduce_sum(loss) * (1. / GLOBAL_BATCH_SIZE) or you can use tf.nn.compute_average_loss which takes the per example loss, optional sample weights, and GLOBAL_BATCH_SIZE as arguments and returns the scaled loss.\n\nand\n\n    Using tf.reduce_mean is not recommended. Doing so divides the loss by actual per replica batch size which may vary step to step.\n\nand \n\n    Why do this?\n\n    This needs to be done because after the gradients are calculated on each replica, they are synced across the replicas by summing them.\n\n\n\nI think in your notebook about custom training, the loss is larger than the correct one. (If divide your loss function by the number of replica, which is 8 there, it should be correct).\n\nI didn't run several times, but I tried once to change your loss function and run it. The following is the training results of the last epochs, which has training/val acc ~0.80, which a bit better than the result in the original notebook (0.77/0.75). Be careful, the loss value  in the reported is not the correct value to be reported. We should sum the loss values from all the replica, but I didn't spent my time to do this.\n\n    EPOCH:  51\n    loss:  0.078226216 , accuracy_:  0.8203125  , val_loss:  0.08729416  , val_acc_:  0.7246094  , lr:  0.012249263566363598\n    =======================\n    EPOCH:  52\n    loss:  0.063389845 , accuracy_:  0.81827444  , val_loss:  0.08148105  , val_acc_:  0.7871094  , lr:  0.01168680038804542\n    =======================\n    EPOCH:  53\n    loss:  0.057514176 , accuracy_:  0.8226902  , val_loss:  0.07985326  , val_acc_:  0.7949219  , lr:  0.011152460368643147\n    =======================\n    EPOCH:  54\n    loss:  0.07713041 , accuracy_:  0.8141984  , val_loss:  0.08316609  , val_acc_:  0.7890625  , lr:  0.01064483735021099\n    =======================\n    EPOCH:  55\n    loss:  0.0724027 , accuracy_:  0.8345788  , val_loss:  0.05968307  , val_acc_:  0.8203125  , lr:  0.010162595482700439\n    =======================\n    EPOCH:  56\n    loss:  0.066552304 , accuracy_:  0.8213315  , val_loss:  0.07661423  , val_acc_:  0.80078125  , lr:  0.009704465708565417\n    =======================\n    EPOCH:  57\n    loss:  0.07287082 , accuracy_:  0.8199728  , val_loss:  0.07448967  , val_acc_:  0.8222656  , lr:  0.009269242423137147\n    =======================\n    EPOCH:  58\n    loss:  0.053799704 , accuracy_:  0.83186144  , val_loss:  0.11246389  , val_acc_:  0.73828125  , lr:  0.008855780301980289\n    =======================\n    EPOCH:  59\n    loss:  0.051523175 , accuracy_:  0.8413723  , val_loss:  0.07519884  , val_acc_:  0.7910156  , lr:  0.008462991286881274\n    =======================\n    EPOCH:  60\n    loss:  0.052226786 , accuracy_:  0.84714675  , val_loss:  0.080301926  , val_acc_:  0.8046875  , lr:  0.008089841722537211",
      "votes": null
    },
    {
      "id": "749808",
      "postDate": "02/18/2020 23:10:17",
      "content": "<p>I have to admit I did not pay much attention to replicating exactly the same loss a with model.fit() since a constant factor in the loss can be compensated for in the learning rate.</p>\n\n<p>You are correct, my loss computation would be wrong if batches were of varying sizes. With a repeated training dataset, batch size does not change throughout training though.</p>\n\n<p>The easiest way to fix it is probably to remove <code>reduce_mean</code> from the <code>loss_fn</code> and let <code>tf.distribute.ReduceOp.MEAN</code> do the work.</p>",
      "rawMarkdown": "I have to admit I did not pay much attention to replicating exactly the same loss a with model.fit() since a constant factor in the loss can be compensated for in the learning rate.\n\nYou are correct, my loss computation would be wrong if batches were of varying sizes. With a repeated training dataset, batch size does not change throughout training though.\n\nThe easiest way to fix it is probably to remove `reduce_mean` from the `loss_fn` and let `tf.distribute.ReduceOp.MEAN` do the work.",
      "votes": null
    },
    {
      "id": "749830",
      "postDate": "02/18/2020 23:33:27",
      "content": "<p>Yes, <code>tf.distribute.ReduceOp</code> should work (I am not familiar with it, currently, I do the ReduceOP in my own way, which is not a good practice ....)</p>",
      "rawMarkdown": "Yes, `tf.distribute.ReduceOp` should work (I am not familiar with it, currently, I do the ReduceOP in my own way, which is not a good practice ....)",
      "votes": null
    },
    {
      "id": "749838",
      "postDate": "02/18/2020 23:48:13",
      "content": "<p>Your own custom reduce is absolutely fine. Get the per-replica loss with <code>strategy.experiemental_local_results(loss)</code> and then you can do anything you want.</p>\n\n<p>Data returned from <code>strategy.experimental_run_v2</code> is in a custom format that has more info for <code>tf.distribute.ReduceOp</code>to work. <code>strategy.experiemental_local_results()</code> turns it into a regular per-replica list.</p>",
      "rawMarkdown": "Your own custom reduce is absolutely fine. Get the per-replica loss with `strategy.experiemental_local_results(loss)` and then you can do anything you want.\n\nData returned from `strategy.experimental_run_v2` is in a custom format that has more info for `tf.distribute.ReduceOp `to work. `strategy.experiemental_local_results()` turns it into a regular per-replica list.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 745633,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "02/14/2020 02:51:14",
      "content": "<p>I personally have question for one part in my kernel</p>\n\n<pre><code>for epoch in range(last_epoch, EPOCHS):\n\n    epoch_start_time = datetime.datetime.now()\n\n    # We need to shuffle the training dataset in every epoch.\n    # I don't know if there is a better way than the following way.\n\n    training_dataset = get_training_dataset()\n    training_dist_dataset = strategy.experimental_distribute_dataset(training_dataset)\n</code></pre>\n\n<p>If we don't call get_training_dataset(), and we just iterate again over the same object <code>training_dataset</code>, do the elements appear in the same order as in the previous iteratings?</p>\n\n<p>If yes, we should avoid this situation. Other than calling <code>get_training_dataset()</code>, is there a better way??</p>",
      "votes": null,
      "replies": [
        {
          "id": 746162,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/14/2020 17:33:21",
          "content": "<p>Congrats for being the first to experiment with distributed custom training loops !</p>\n\n<p>The training dataset returned by <code>get_training_dataset()</code> is iterable. The training loop should iterate on it. I have a custom training loop sample here:\n<a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb\">https://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749758,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/18/2020 22:21:29",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> , I observed that, in your notebook above, you use <code>tf.reduce_mean</code> in the loss function and the loss function is called inside <code>train_step</code>.</p>\n\n<p>However, in [https://www.tensorflow.org/tutorials/distribute/custom_training#define_the_loss_function](Custom training with tf.distribute.Strategy), it specifically says the following:</p>\n\n<pre><code>If you're writing a custom training loop, as in this tutorial, you should sum the per example losses and divide the sum by the GLOBAL_BATCH_SIZE: scale_loss = tf.reduce_sum(loss) * (1. / GLOBAL_BATCH_SIZE) or you can use tf.nn.compute_average_loss which takes the per example loss, optional sample weights, and GLOBAL_BATCH_SIZE as arguments and returns the scaled loss.\n</code></pre>\n\n<p>and</p>\n\n<pre><code>Using tf.reduce_mean is not recommended. Doing so divides the loss by actual per replica batch size which may vary step to step.\n</code></pre>\n\n<p>and </p>\n\n<pre><code>Why do this?\n\nThis needs to be done because after the gradients are calculated on each replica, they are synced across the replicas by summing them.\n</code></pre>\n\n<p>I think in your notebook about custom training, the loss is larger than the correct one. (If divide your loss function by the number of replica, which is 8 there, it should be correct).</p>\n\n<p>I didn't run several times, but I tried once to change your loss function and run it. The following is the training results of the last epochs, which has training/val acc ~0.80, which a bit better than the result in the original notebook (0.77/0.75). Be careful, the loss value  in the reported is not the correct value to be reported. We should sum the loss values from all the replica, but I didn't spent my time to do this.</p>\n\n<pre><code>EPOCH:  51\nloss:  0.078226216 , accuracy_:  0.8203125  , val_loss:  0.08729416  , val_acc_:  0.7246094  , lr:  0.012249263566363598\n=======================\nEPOCH:  52\nloss:  0.063389845 , accuracy_:  0.81827444  , val_loss:  0.08148105  , val_acc_:  0.7871094  , lr:  0.01168680038804542\n=======================\nEPOCH:  53\nloss:  0.057514176 , accuracy_:  0.8226902  , val_loss:  0.07985326  , val_acc_:  0.7949219  , lr:  0.011152460368643147\n=======================\nEPOCH:  54\nloss:  0.07713041 , accuracy_:  0.8141984  , val_loss:  0.08316609  , val_acc_:  0.7890625  , lr:  0.01064483735021099\n=======================\nEPOCH:  55\nloss:  0.0724027 , accuracy_:  0.8345788  , val_loss:  0.05968307  , val_acc_:  0.8203125  , lr:  0.010162595482700439\n=======================\nEPOCH:  56\nloss:  0.066552304 , accuracy_:  0.8213315  , val_loss:  0.07661423  , val_acc_:  0.80078125  , lr:  0.009704465708565417\n=======================\nEPOCH:  57\nloss:  0.07287082 , accuracy_:  0.8199728  , val_loss:  0.07448967  , val_acc_:  0.8222656  , lr:  0.009269242423137147\n=======================\nEPOCH:  58\nloss:  0.053799704 , accuracy_:  0.83186144  , val_loss:  0.11246389  , val_acc_:  0.73828125  , lr:  0.008855780301980289\n=======================\nEPOCH:  59\nloss:  0.051523175 , accuracy_:  0.8413723  , val_loss:  0.07519884  , val_acc_:  0.7910156  , lr:  0.008462991286881274\n=======================\nEPOCH:  60\nloss:  0.052226786 , accuracy_:  0.84714675  , val_loss:  0.080301926  , val_acc_:  0.8046875  , lr:  0.008089841722537211\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749808,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 23:10:17",
          "content": "<p>I have to admit I did not pay much attention to replicating exactly the same loss a with model.fit() since a constant factor in the loss can be compensated for in the learning rate.</p>\n\n<p>You are correct, my loss computation would be wrong if batches were of varying sizes. With a repeated training dataset, batch size does not change throughout training though.</p>\n\n<p>The easiest way to fix it is probably to remove <code>reduce_mean</code> from the <code>loss_fn</code> and let <code>tf.distribute.ReduceOp.MEAN</code> do the work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749830,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "02/18/2020 23:33:27",
          "content": "<p>Yes, <code>tf.distribute.ReduceOp</code> should work (I am not familiar with it, currently, I do the ReduceOP in my own way, which is not a good practice ....)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749838,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/18/2020 23:48:13",
          "content": "<p>Your own custom reduce is absolutely fine. Get the per-replica loss with <code>strategy.experiemental_local_results(loss)</code> and then you can do anything you want.</p>\n\n<p>Data returned from <code>strategy.experimental_run_v2</code> is in a custom format that has more info for <code>tf.distribute.ReduceOp</code>to work. <code>strategy.experiemental_local_results()</code> turns it into a regular per-replica list.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "745630": "Hi, I just published my kernel for demonstrating how to do custom training with TPU -- [Custom training with TPU](https://www.kaggle.com/yihdarshieh/custom-training-with-tpu?scriptVersionId=28624481).\n\nHope you find it helpful!",
    "745633": "I personally have question for one part in my kernel\n\n    for epoch in range(last_epoch, EPOCHS):\n    \n        epoch_start_time = datetime.datetime.now()\n\n        # We need to shuffle the training dataset in every epoch.\n        # I don't know if there is a better way than the following way.\n        \n        training_dataset = get_training_dataset()\n        training_dist_dataset = strategy.experimental_distribute_dataset(training_dataset)\n\nIf we don't call get_training_dataset(), and we just iterate again over the same object ` training_dataset`, do the elements appear in the same order as in the previous iteratings?\n\nIf yes, we should avoid this situation. Other than calling `get_training_dataset()`, is there a better way??",
    "746162": "Congrats for being the first to experiment with distributed custom training loops !\n\nThe training dataset returned by `get_training_dataset()` is iterable. The training loop should iterate on it. I have a custom training loop sample here:\nhttps://colab.research.google.com/github/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_customtrainloop_tf2.1.ipynb",
    "749758": "mgornergoogle , I observed that, in your notebook above, you use `tf.reduce_mean` in the loss function and the loss function is called inside `train_step`.\n\nHowever, in [https://www.tensorflow.org/tutorials/distribute/custom_training#define_the_loss_function](Custom training with tf.distribute.Strategy), it specifically says the following:\n\n    If you're writing a custom training loop, as in this tutorial, you should sum the per example losses and divide the sum by the GLOBAL_BATCH_SIZE: scale_loss = tf.reduce_sum(loss) * (1. / GLOBAL_BATCH_SIZE) or you can use tf.nn.compute_average_loss which takes the per example loss, optional sample weights, and GLOBAL_BATCH_SIZE as arguments and returns the scaled loss.\n\nand\n\n    Using tf.reduce_mean is not recommended. Doing so divides the loss by actual per replica batch size which may vary step to step.\n\nand \n\n    Why do this?\n\n    This needs to be done because after the gradients are calculated on each replica, they are synced across the replicas by summing them.\n\n\n\nI think in your notebook about custom training, the loss is larger than the correct one. (If divide your loss function by the number of replica, which is 8 there, it should be correct).\n\nI didn't run several times, but I tried once to change your loss function and run it. The following is the training results of the last epochs, which has training/val acc ~0.80, which a bit better than the result in the original notebook (0.77/0.75). Be careful, the loss value  in the reported is not the correct value to be reported. We should sum the loss values from all the replica, but I didn't spent my time to do this.\n\n    EPOCH:  51\n    loss:  0.078226216 , accuracy_:  0.8203125  , val_loss:  0.08729416  , val_acc_:  0.7246094  , lr:  0.012249263566363598\n    =======================\n    EPOCH:  52\n    loss:  0.063389845 , accuracy_:  0.81827444  , val_loss:  0.08148105  , val_acc_:  0.7871094  , lr:  0.01168680038804542\n    =======================\n    EPOCH:  53\n    loss:  0.057514176 , accuracy_:  0.8226902  , val_loss:  0.07985326  , val_acc_:  0.7949219  , lr:  0.011152460368643147\n    =======================\n    EPOCH:  54\n    loss:  0.07713041 , accuracy_:  0.8141984  , val_loss:  0.08316609  , val_acc_:  0.7890625  , lr:  0.01064483735021099\n    =======================\n    EPOCH:  55\n    loss:  0.0724027 , accuracy_:  0.8345788  , val_loss:  0.05968307  , val_acc_:  0.8203125  , lr:  0.010162595482700439\n    =======================\n    EPOCH:  56\n    loss:  0.066552304 , accuracy_:  0.8213315  , val_loss:  0.07661423  , val_acc_:  0.80078125  , lr:  0.009704465708565417\n    =======================\n    EPOCH:  57\n    loss:  0.07287082 , accuracy_:  0.8199728  , val_loss:  0.07448967  , val_acc_:  0.8222656  , lr:  0.009269242423137147\n    =======================\n    EPOCH:  58\n    loss:  0.053799704 , accuracy_:  0.83186144  , val_loss:  0.11246389  , val_acc_:  0.73828125  , lr:  0.008855780301980289\n    =======================\n    EPOCH:  59\n    loss:  0.051523175 , accuracy_:  0.8413723  , val_loss:  0.07519884  , val_acc_:  0.7910156  , lr:  0.008462991286881274\n    =======================\n    EPOCH:  60\n    loss:  0.052226786 , accuracy_:  0.84714675  , val_loss:  0.080301926  , val_acc_:  0.8046875  , lr:  0.008089841722537211",
    "749808": "I have to admit I did not pay much attention to replicating exactly the same loss a with model.fit() since a constant factor in the loss can be compensated for in the learning rate.\n\nYou are correct, my loss computation would be wrong if batches were of varying sizes. With a repeated training dataset, batch size does not change throughout training though.\n\nThe easiest way to fix it is probably to remove `reduce_mean` from the `loss_fn` and let `tf.distribute.ReduceOp.MEAN` do the work.",
    "749830": "Yes, `tf.distribute.ReduceOp` should work (I am not familiar with it, currently, I do the ReduceOP in my own way, which is not a good practice ....)",
    "749838": "Your own custom reduce is absolutely fine. Get the per-replica loss with `strategy.experiemental_local_results(loss)` and then you can do anything you want.\n\nData returned from `strategy.experimental_run_v2` is in a custom format that has more info for `tf.distribute.ReduceOp `to work. `strategy.experiemental_local_results()` turns it into a regular per-replica list."
  },
  "source": "meta"
}