{
  "id": 145276,
  "title": "small batch size in custom train loop in tf2?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/145276",
  "author_name": "soomiles",
  "post_date": "2020-04-22T14:21:58.168000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi, is there anyone who codes custom train loop in tensorflow?</p>\n\n<p>I code custom train loop with tf.GradientTape, but it can be trained with four times less batch size rather than model.fit(). My custom train loop raises OOM..</p>\n\n<p>Everything except using custom train loop is same.</p>\n\n<p>I just wonder it's a general problem, or my coding fault.\nIn gpu, batch size can be same both in model.fit() and my custom train loop.\nBut in tpu, batch size should be 4 times smaller..</p>\n\n<p>Is there something I miss in tpu setting? </p>\n\n<p>These are some parts of code.</p>\n\n<p>Builds model:</p>\n\n<p>```\n    with self.strategy.scope():\n        # Network \n        transformer = TFAutoModel.from_pretrained(self.pretrained_model_name)\n        self.model = CoolSystem(transformer, max_len=self.max_len)</p>\n\n<pre><code>    # Loss Function \n    self.loss_fn = tf.keras.losses.BinaryCrossentropy(from_logits=False, reduction=tf.keras.losses.Reduction.NONE)\n</code></pre>\n\n<p>```</p>\n\n<p>Train Loop:</p>\n\n<p>```\n    @tf.function\n    def train_step(self, seq, labels):\n        with tf.GradientTape() as tape:\n            logits = self.model(seq)\n            loss = self.loss_fn(labels, logits)</p>\n\n<pre><code>    gradient = tape.gradient(loss, self.model.trainable_variables)\n    self.optimizer.apply_gradients(zip(gradient, self.model.trainable_variables))\n\n    self.train_loss(loss)\n    self.train_auc(labels, tf.squeeze(logits))\n\n    return loss\n\n@tf.function\ndef distribute_train_step(self, seq, labels):\n    loss = self.strategy.experimental_run_v2(self.train_step, args=(seq, labels))\n\n    loss = self.strategy.reduce(tf.distribute.ReduceOp.SUM, loss, axis=None)  * (1. / self.batch_size)\n\n    return loss\n\ndef train(self):\n    start_time = time.time()\n\n    for idx in range(self.n_train_steps):\n        seq, labels = next(self.train_dataset)\n        # update network\n        loss = self.distribute_train_step(seq, labels)\n\n        print(f\"iter: [{idx+1:6d}/{self.n_train_steps:6d}] time: {time.time() - start_time:4.4f} \\\n        loss: {self.train_loss.result():.6f} auc: {self.train_auc.result():.4f}\", end=\"\\r\")\n    print()\n</code></pre>\n\n<p>```</p>\n\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": 817174,
      "postDate": "2020-04-22T23:01:40.867Z",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> posted a TPU-optimized custom training loop for this competition here:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455</a></p>\n\n<p>Refer to the topic \"<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a>\"  to see what I mean by CTL and TPU-optimized CTL.</p>",
      "rawMarkdown": "@dimitreoliveira posted a TPU-optimized custom training loop for this competition here:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\n\nRefer to the topic \"[TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\"  to see what I mean by CTL and TPU-optimized CTL.",
      "votes": 4,
      "replies": [
        {
          "id": 817440,
          "postDate": "2020-04-23T06:35:40.650Z",
          "content": "<p>Thank you. I'll try and post it if it works!</p>",
          "rawMarkdown": "Thank you. I'll try and post it if it works!"
        },
        {
          "id": 818275,
          "postDate": "2020-04-23T18:33:51.447Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've spent time for fixing this problem, but custom training loop is possible with smaller batch size.\nIn <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130703\">this discussion</a>, he changed batch size smaller. Is it somewhat related to my issue?</p>",
          "rawMarkdown": "@mgornergoogle I've spent time for fixing this problem, but custom training loop is possible with smaller batch size.\nIn [this discussion](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130703), he changed batch size smaller. Is it somewhat related to my issue?"
        },
        {
          "id": 819660,
          "postDate": "2020-04-24T18:55:31.800Z",
          "content": "<p>Sorry, I don't understand what you are trying to achieve and what is not working.</p>",
          "rawMarkdown": "Sorry, I don't understand what you are trying to achieve and what is not working."
        },
        {
          "id": 820300,
          "postDate": "2020-04-25T10:12:45.987Z",
          "content": "<p>My question was 'Does memory leak occur when running with custom train loop in TPU?'\nWhen learning with Keras.Model.fit(), I can train 16 * 8 batch size with XLM_Roberta_large, but when customer train loop, I can train 4 * 8 batch size only.</p>",
          "rawMarkdown": "My question was 'Does memory leak occur when running with custom train loop in TPU?'\nWhen learning with Keras.Model.fit(), I can train 16 * 8 batch size with XLM\\_Roberta\\_large, but when customer train loop, I can train 4 * 8 batch size only."
        }
      ]
    },
    {
      "id": 816928,
      "postDate": "2020-04-22T18:04:58.267Z",
      "content": "<p>Did  you get a chance to check <a href=\"/martingorner\">@martingorner</a>'s Custom Training Loop notebook:\n<a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\">https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu</a>\nAlso this discussion: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a></p>",
      "rawMarkdown": "Did  you get a chance to check @martingorner's Custom Training Loop notebook:\nhttps://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\nAlso this discussion: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\n",
      "votes": 1,
      "replies": [
        {
          "id": 817441,
          "postDate": "2020-04-23T06:35:45.047Z",
          "content": "<p>Thank you. I'll try and post it if it works!</p>",
          "rawMarkdown": "Thank you. I'll try and post it if it works!"
        }
      ]
    },
    {
      "id": 816700,
      "postDate": "2020-04-22T14:21:58.170Z",
      "content": "<p>Hi, is there anyone who codes custom train loop in tensorflow?</p>\n\n<p>I code custom train loop with tf.GradientTape, but it can be trained with four times less batch size rather than model.fit(). My custom train loop raises OOM..</p>\n\n<p>Everything except using custom train loop is same.</p>\n\n<p>I just wonder it's a general problem, or my coding fault.\nIn gpu, batch size can be same both in model.fit() and my custom train loop.\nBut in tpu, batch size should be 4 times smaller..</p>\n\n<p>Is there something I miss in tpu setting? </p>\n\n<p>These are some parts of code.</p>\n\n<p>Builds model:</p>\n\n<p>```\n    with self.strategy.scope():\n        # Network \n        transformer = TFAutoModel.from_pretrained(self.pretrained_model_name)\n        self.model = CoolSystem(transformer, max_len=self.max_len)</p>\n\n<pre><code>    # Loss Function \n    self.loss_fn = tf.keras.losses.BinaryCrossentropy(from_logits=False, reduction=tf.keras.losses.Reduction.NONE)\n</code></pre>\n\n<p>```</p>\n\n<p>Train Loop:</p>\n\n<p>```\n    @tf.function\n    def train_step(self, seq, labels):\n        with tf.GradientTape() as tape:\n            logits = self.model(seq)\n            loss = self.loss_fn(labels, logits)</p>\n\n<pre><code>    gradient = tape.gradient(loss, self.model.trainable_variables)\n    self.optimizer.apply_gradients(zip(gradient, self.model.trainable_variables))\n\n    self.train_loss(loss)\n    self.train_auc(labels, tf.squeeze(logits))\n\n    return loss\n\n@tf.function\ndef distribute_train_step(self, seq, labels):\n    loss = self.strategy.experimental_run_v2(self.train_step, args=(seq, labels))\n\n    loss = self.strategy.reduce(tf.distribute.ReduceOp.SUM, loss, axis=None)  * (1. / self.batch_size)\n\n    return loss\n\ndef train(self):\n    start_time = time.time()\n\n    for idx in range(self.n_train_steps):\n        seq, labels = next(self.train_dataset)\n        # update network\n        loss = self.distribute_train_step(seq, labels)\n\n        print(f\"iter: [{idx+1:6d}/{self.n_train_steps:6d}] time: {time.time() - start_time:4.4f} \\\n        loss: {self.train_loss.result():.6f} auc: {self.train_auc.result():.4f}\", end=\"\\r\")\n    print()\n</code></pre>\n\n<p>```</p>\n\n<p>Thanks for reading!</p>",
      "rawMarkdown": "Hi, is there anyone who codes custom train loop in tensorflow?\n\nI code custom train loop with tf.GradientTape, but it can be trained with four times less batch size rather than model.fit(). My custom train loop raises OOM..\n\nEverything except using custom train loop is same.\n\nI just wonder it's a general problem, or my coding fault.\nIn gpu, batch size can be same both in model.fit() and my custom train loop.\nBut in tpu, batch size should be 4 times smaller..\n\nIs there something I miss in tpu setting? \n\nThese are some parts of code.\n\nBuilds model:\n\n```\n    with self.strategy.scope():\n        # Network \n        transformer = TFAutoModel.from_pretrained(self.pretrained_model_name)\n        self.model = CoolSystem(transformer, max_len=self.max_len)\n        \n        # Loss Function \n        self.loss_fn = tf.keras.losses.BinaryCrossentropy(from_logits=False, reduction=tf.keras.losses.Reduction.NONE)\n\n```\n\nTrain Loop:\n\n```\n    @tf.function\n    def train_step(self, seq, labels):\n        with tf.GradientTape() as tape:\n            logits = self.model(seq)\n            loss = self.loss_fn(labels, logits)\n \n        gradient = tape.gradient(loss, self.model.trainable_variables)\n        self.optimizer.apply_gradients(zip(gradient, self.model.trainable_variables))\n\n        self.train_loss(loss)\n        self.train_auc(labels, tf.squeeze(logits))\n\n        return loss\n    \n    @tf.function\n    def distribute_train_step(self, seq, labels):\n        loss = self.strategy.experimental_run_v2(self.train_step, args=(seq, labels))\n\n        loss = self.strategy.reduce(tf.distribute.ReduceOp.SUM, loss, axis=None)  * (1. / self.batch_size)\n\n        return loss\n\n    def train(self):\n        start_time = time.time()\n        \n        for idx in range(self.n_train_steps):\n            seq, labels = next(self.train_dataset)\n            # update network\n            loss = self.distribute_train_step(seq, labels)\n\n            print(f\"iter: [{idx+1:6d}/{self.n_train_steps:6d}] time: {time.time() - start_time:4.4f} \\\n            loss: {self.train_loss.result():.6f} auc: {self.train_auc.result():.4f}\", end=\"\\r\")\n        print()\n```\n\nThanks for reading!",
      "votes": 1
    },
    {
      "id": 816719,
      "postDate": "2020-04-22T14:32:24.503Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 817174,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-04-22T23:01:40.867000",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> posted a TPU-optimized custom training loop for this competition here:\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455</a></p>\n\n<p>Refer to the topic \"<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">TPU: extreme optimizations</a>\"  to see what I mean by CTL and TPU-optimized CTL.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 817440,
          "author_name": "soomiles",
          "author_url": "",
          "post_date": "2020-04-23T06:35:40.650000",
          "content": "<p>Thank you. I'll try and post it if it works!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 818275,
          "author_name": "soomiles",
          "author_url": "",
          "post_date": "2020-04-23T18:33:51.447000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I've spent time for fixing this problem, but custom training loop is possible with smaller batch size.\nIn <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130703\">this discussion</a>, he changed batch size smaller. Is it somewhat related to my issue?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 819660,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-24T18:55:31.800000",
          "content": "<p>Sorry, I don't understand what you are trying to achieve and what is not working.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 820300,
          "author_name": "soomiles",
          "author_url": "",
          "post_date": "2020-04-25T10:12:45.987000",
          "content": "<p>My question was 'Does memory leak occur when running with custom train loop in TPU?'\nWhen learning with Keras.Model.fit(), I can train 16 * 8 batch size with XLM_Roberta_large, but when customer train loop, I can train 4 * 8 batch size only.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 816928,
      "author_name": "Ilya Figotin",
      "author_url": "",
      "post_date": "2020-04-22T18:04:58.267000",
      "content": "<p>Did  you get a chance to check <a href=\"/martingorner\">@martingorner</a>'s Custom Training Loop notebook:\n<a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\">https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu</a>\nAlso this discussion: <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 817441,
          "author_name": "soomiles",
          "author_url": "",
          "post_date": "2020-04-23T06:35:45.047000",
          "content": "<p>Thank you. I'll try and post it if it works!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 816719,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-22T14:32:24.503000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "817174": "@dimitreoliveira posted a TPU-optimized custom training loop for this competition here:\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\n\nRefer to the topic \"[TPU: extreme optimizations](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443)\"  to see what I mean by CTL and TPU-optimized CTL.",
    "816928": "Did  you get a chance to check @martingorner's Custom Training Loop notebook:\nhttps://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\nAlso this discussion: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\n",
    "816700": "Hi, is there anyone who codes custom train loop in tensorflow?\n\nI code custom train loop with tf.GradientTape, but it can be trained with four times less batch size rather than model.fit(). My custom train loop raises OOM..\n\nEverything except using custom train loop is same.\n\nI just wonder it's a general problem, or my coding fault.\nIn gpu, batch size can be same both in model.fit() and my custom train loop.\nBut in tpu, batch size should be 4 times smaller..\n\nIs there something I miss in tpu setting? \n\nThese are some parts of code.\n\nBuilds model:\n\n```\n    with self.strategy.scope():\n        # Network \n        transformer = TFAutoModel.from_pretrained(self.pretrained_model_name)\n        self.model = CoolSystem(transformer, max_len=self.max_len)\n        \n        # Loss Function \n        self.loss_fn = tf.keras.losses.BinaryCrossentropy(from_logits=False, reduction=tf.keras.losses.Reduction.NONE)\n\n```\n\nTrain Loop:\n\n```\n    @tf.function\n    def train_step(self, seq, labels):\n        with tf.GradientTape() as tape:\n            logits = self.model(seq)\n            loss = self.loss_fn(labels, logits)\n \n        gradient = tape.gradient(loss, self.model.trainable_variables)\n        self.optimizer.apply_gradients(zip(gradient, self.model.trainable_variables))\n\n        self.train_loss(loss)\n        self.train_auc(labels, tf.squeeze(logits))\n\n        return loss\n    \n    @tf.function\n    def distribute_train_step(self, seq, labels):\n        loss = self.strategy.experimental_run_v2(self.train_step, args=(seq, labels))\n\n        loss = self.strategy.reduce(tf.distribute.ReduceOp.SUM, loss, axis=None)  * (1. / self.batch_size)\n\n        return loss\n\n    def train(self):\n        start_time = time.time()\n        \n        for idx in range(self.n_train_steps):\n            seq, labels = next(self.train_dataset)\n            # update network\n            loss = self.distribute_train_step(seq, labels)\n\n            print(f\"iter: [{idx+1:6d}/{self.n_train_steps:6d}] time: {time.time() - start_time:4.4f} \\\n            loss: {self.train_loss.result():.6f} auc: {self.train_auc.result():.4f}\", end=\"\\r\")\n        print()\n```\n\nThanks for reading!",
    "816719": ""
  }
}