{
  "id": 251601,
  "title": "Diff Between Pytorch TPU code and TF TPU Code",
  "url": "/competitions/siim-covid19-detection/discussion/251601",
  "author_name": "",
  "post_date": "2021-07-08T02:19:06.694181200Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello, everyone. I don’t know if you have ever encountered such a situation. In kaggle environment, the tpu code written by tensorflow can run perfectly, but the tpu code written by pytorch xla with the same configuration (batch size, model, learning rate, input, etc.) will raise <code>Out of Memory</code> error or <code>model gradient explosion</code>.</p>\n<p><a href=\"https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments\" target=\"_blank\">My pytorch tpu code</a> has meeted above phenomenon, because keeping the same settings with <a href=\"https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data\" target=\"_blank\">tf tpu implementation</a> will cause <code>Out of Memory</code> error or <code>model gradient explosion</code>, so the configuration of <a href=\"https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments\" target=\"_blank\">my version 3 pytorch implementation</a> is not completely implemented in accordance with the <a href=\"https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data\" target=\"_blank\">tf code</a>.</p>\n<ul>\n<li>If the image size in my version3 implementation becomes 600x600(version3: 400x400), learning rate becomes 1e-3/8(version3: 1e-4/8), and batch size becomes 128(version3: 64), to be consistent with above tf implementation, there will be an oom error which doesn't occur in above tf implementation. (Note: 8 is the num of tpu cores, final learning rate need to multiply with num of tpu cores, so final lr is equal to 1e-3/8*8)</li>\n<li>If I only change the learning rate in my version3 implementation to 1e-3/8(version3: 1e-4/8), there will be a gradient explosion which doesn't occur in above tf implementation.</li>\n</ul>\n<p>I don’t think there is an error in my pytorch xla implementation, so is the above phenomenon a normal phenomenon?</p>",
  "messages": [
    {
      "id": "1380298",
      "postDate": "07/08/2021 02:19:06",
      "content": "<p>Hello, everyone. I don’t know if you have ever encountered such a situation. In kaggle environment, the tpu code written by tensorflow can run perfectly, but the tpu code written by pytorch xla with the same configuration (batch size, model, learning rate, input, etc.) will raise <code>Out of Memory</code> error or <code>model gradient explosion</code>.</p>\n<p><a href=\"https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments\" target=\"_blank\">My pytorch tpu code</a> has meeted above phenomenon, because keeping the same settings with <a href=\"https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data\" target=\"_blank\">tf tpu implementation</a> will cause <code>Out of Memory</code> error or <code>model gradient explosion</code>, so the configuration of <a href=\"https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments\" target=\"_blank\">my version 3 pytorch implementation</a> is not completely implemented in accordance with the <a href=\"https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data\" target=\"_blank\">tf code</a>.</p>\n<ul>\n<li>If the image size in my version3 implementation becomes 600x600(version3: 400x400), learning rate becomes 1e-3/8(version3: 1e-4/8), and batch size becomes 128(version3: 64), to be consistent with above tf implementation, there will be an oom error which doesn't occur in above tf implementation. (Note: 8 is the num of tpu cores, final learning rate need to multiply with num of tpu cores, so final lr is equal to 1e-3/8*8)</li>\n<li>If I only change the learning rate in my version3 implementation to 1e-3/8(version3: 1e-4/8), there will be a gradient explosion which doesn't occur in above tf implementation.</li>\n</ul>\n<p>I don’t think there is an error in my pytorch xla implementation, so is the above phenomenon a normal phenomenon?</p>",
      "rawMarkdown": "Hello, everyone. I don’t know if you have ever encountered such a situation. In kaggle environment, the tpu code written by tensorflow can run perfectly, but the tpu code written by pytorch xla with the same configuration (batch size, model, learning rate, input, etc.) will raise `Out of Memory` error or `model gradient explosion`.\n\n[My pytorch tpu code](https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments) has meeted above phenomenon, because keeping the same settings with [tf tpu implementation](https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data) will cause `Out of Memory` error or `model gradient explosion`, so the configuration of [my version 3 pytorch implementation](https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments) is not completely implemented in accordance with the [tf code](https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data).\n\n- If the image size in my version3 implementation becomes 600x600(version3: 400x400), learning rate becomes 1e-3/8(version3: 1e-4/8), and batch size becomes 128(version3: 64), to be consistent with above tf implementation, there will be an oom error which doesn't occur in above tf implementation. (Note: 8 is the num of tpu cores, final learning rate need to multiply with num of tpu cores, so final lr is equal to 1e-3/8*8)\n- If I only change the learning rate in my version3 implementation to 1e-3/8(version3: 1e-4/8), there will be a gradient explosion which doesn't occur in above tf implementation.\n\nI don’t think there is an error in my pytorch xla implementation, so is the above phenomenon a normal phenomenon?",
      "votes": null
    },
    {
      "id": "1400339",
      "postDate": "07/26/2021 08:12:40",
      "content": "<p>Hi, have you found a solution to this? Is the fix to try different learning rates?</p>",
      "rawMarkdown": "Hi, have you found a solution to this? Is the fix to try different learning rates?",
      "votes": null
    },
    {
      "id": "1400383",
      "postDate": "07/26/2021 08:43:21",
      "content": "<p>I have not yet find a solution to this, may be this is the implementation diff between pytorch and tensorflow.</p>",
      "rawMarkdown": "I have not yet find a solution to this, may be this is the implementation diff between pytorch and tensorflow.",
      "votes": null
    },
    {
      "id": "1400542",
      "postDate": "07/26/2021 11:32:31",
      "content": "<p>Tensorflow TPU setup</p>\n<ol>\n<li><p>gradiant accumulate<br>\n<a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\">https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e</a></p></li>\n<li><p>tpu augment<br>\nmixup: <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\" target=\"_blank\">https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu</a><br>\nrotation: <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96</a></p></li>\n</ol>",
      "rawMarkdown": "Tensorflow TPU setup\n\n1. gradiant accumulate\nhttps://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\n\n2. tpu augment\nmixup: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\nrotation: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
      "votes": null
    },
    {
      "id": "1400792",
      "postDate": "07/26/2021 15:26:24",
      "content": "<p>Thank you for your reply!</p>",
      "rawMarkdown": "Thank you for your reply!",
      "votes": null
    },
    {
      "id": "1402245",
      "postDate": "07/28/2021 03:32:44",
      "content": "<p>I have the same problem. Tensorflow + TPU works much better than PyTorch + TPU. If you want to use TPU with PyTorch, try training with a lower batch size. You can also use PyTorch lightning to write hardware-agnostic code and try comparing gpu and tpu performance</p>",
      "rawMarkdown": "I have the same problem. Tensorflow + TPU works much better than PyTorch + TPU. If you want to use TPU with PyTorch, try training with a lower batch size. You can also use PyTorch lightning to write hardware-agnostic code and try comparing gpu and tpu performance",
      "votes": null
    },
    {
      "id": "1402253",
      "postDate": "07/28/2021 03:41:55",
      "content": "<p>I'll give it a try, thank you for your suggestion~</p>",
      "rawMarkdown": "I'll give it a try, thank you for your suggestion~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1400339,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "07/26/2021 08:12:40",
      "content": "<p>Hi, have you found a solution to this? Is the fix to try different learning rates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1400383,
          "author_name": "shangweichen",
          "author_url": "",
          "post_date": "07/26/2021 08:43:21",
          "content": "<p>I have not yet find a solution to this, may be this is the implementation diff between pytorch and tensorflow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1400542,
      "author_name": "drzhuzhe",
      "author_url": "",
      "post_date": "07/26/2021 11:32:31",
      "content": "<p>Tensorflow TPU setup</p>\n<ol>\n<li><p>gradiant accumulate<br>\n<a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\" target=\"_blank\">https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e</a></p></li>\n<li><p>tpu augment<br>\nmixup: <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\" target=\"_blank\">https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu</a><br>\nrotation: <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96</a></p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1400792,
          "author_name": "shangweichen",
          "author_url": "",
          "post_date": "07/26/2021 15:26:24",
          "content": "<p>Thank you for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1402245,
      "author_name": "readoc",
      "author_url": "",
      "post_date": "07/28/2021 03:32:44",
      "content": "<p>I have the same problem. Tensorflow + TPU works much better than PyTorch + TPU. If you want to use TPU with PyTorch, try training with a lower batch size. You can also use PyTorch lightning to write hardware-agnostic code and try comparing gpu and tpu performance</p>",
      "votes": null,
      "replies": [
        {
          "id": 1402253,
          "author_name": "shangweichen",
          "author_url": "",
          "post_date": "07/28/2021 03:41:55",
          "content": "<p>I'll give it a try, thank you for your suggestion~</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1380298": "Hello, everyone. I don’t know if you have ever encountered such a situation. In kaggle environment, the tpu code written by tensorflow can run perfectly, but the tpu code written by pytorch xla with the same configuration (batch size, model, learning rate, input, etc.) will raise `Out of Memory` error or `model gradient explosion`.\n\n[My pytorch tpu code](https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments) has meeted above phenomenon, because keeping the same settings with [tf tpu implementation](https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data) will cause `Out of Memory` error or `model gradient explosion`, so the configuration of [my version 3 pytorch implementation](https://www.kaggle.com/shangweichen/siim-efnb7-train-pytorch-xla-tpu/comments) is not completely implemented in accordance with the [tf code](https://www.kaggle.com/h053473666/siim-covid19-efnb7-train-study/data).\n\n- If the image size in my version3 implementation becomes 600x600(version3: 400x400), learning rate becomes 1e-3/8(version3: 1e-4/8), and batch size becomes 128(version3: 64), to be consistent with above tf implementation, there will be an oom error which doesn't occur in above tf implementation. (Note: 8 is the num of tpu cores, final learning rate need to multiply with num of tpu cores, so final lr is equal to 1e-3/8*8)\n- If I only change the learning rate in my version3 implementation to 1e-3/8(version3: 1e-4/8), there will be a gradient explosion which doesn't occur in above tf implementation.\n\nI don’t think there is an error in my pytorch xla implementation, so is the above phenomenon a normal phenomenon?",
    "1400339": "Hi, have you found a solution to this? Is the fix to try different learning rates?",
    "1400383": "I have not yet find a solution to this, may be this is the implementation diff between pytorch and tensorflow.",
    "1400542": "Tensorflow TPU setup\n\n1. gradiant accumulate\nhttps://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e\n\n2. tpu augment\nmixup: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\nrotation: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96",
    "1400792": "Thank you for your reply!",
    "1402245": "I have the same problem. Tensorflow + TPU works much better than PyTorch + TPU. If you want to use TPU with PyTorch, try training with a lower batch size. You can also use PyTorch lightning to write hardware-agnostic code and try comparing gpu and tpu performance",
    "1402253": "I'll give it a try, thank you for your suggestion~"
  },
  "source": "meta"
}