{
  "id": 208548,
  "title": "TPU vs GPU score",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/208548",
  "author_name": "",
  "post_date": "2021-01-03T22:43:11.275823Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello here,</p>\n<p>I have started this competition using <a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> notebook using pytorch with TPU (here: <a href=\"https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84)\" target=\"_blank\">https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84)</a>, my lb score is 0.836.<br>\nI then switched the exact same code to GPU, so i can train local.<br>\nAnd the lb score dropped to 0.821.</p>\n<p>Has anyone observed the same behavior ? Any hints ?</p>\n<p>Thanks in advance,</p>\n<p>Nicolas</p>",
  "messages": [
    {
      "id": "1137411",
      "postDate": "01/03/2021 22:43:11",
      "content": "<p>Hello here,</p>\n<p>I have started this competition using <a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> notebook using pytorch with TPU (here: <a href=\"https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84)\" target=\"_blank\">https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84)</a>, my lb score is 0.836.<br>\nI then switched the exact same code to GPU, so i can train local.<br>\nAnd the lb score dropped to 0.821.</p>\n<p>Has anyone observed the same behavior ? Any hints ?</p>\n<p>Thanks in advance,</p>\n<p>Nicolas</p>",
      "rawMarkdown": "Hello here,\n\nI have started this competition using @joshi98kishan notebook using pytorch with TPU (here: https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84), my lb score is 0.836.\nI then switched the exact same code to GPU, so i can train local.\nAnd the lb score dropped to 0.821.\n\nHas anyone observed the same behavior ? Any hints ?\n\nThanks in advance,\n\nNicolas",
      "votes": null
    },
    {
      "id": "1137471",
      "postDate": "01/04/2021 00:57:22",
      "content": "<p>I also compared my own GPU model with TPU one (both are PyTorch), and noticed the same issue.</p>",
      "rawMarkdown": "I also compared my own GPU model with TPU one (both are PyTorch), and noticed the same issue.",
      "votes": null
    },
    {
      "id": "1138302",
      "postDate": "01/04/2021 14:59:41",
      "content": "<p>Interesting. By how much did you change the batch size and the learning rate when moving from TPU to GPU ?</p>",
      "rawMarkdown": "Interesting. By how much did you change the batch size and the learning rate when moving from TPU to GPU ?",
      "votes": null
    },
    {
      "id": "1138345",
      "postDate": "01/04/2021 16:02:38",
      "content": "<p>I did not change anything.</p>",
      "rawMarkdown": "I did not change anything.",
      "votes": null
    },
    {
      "id": "1139019",
      "postDate": "01/05/2021 05:36:16",
      "content": "<p>What people did in some other notebooks  - divide  the batch size  by 8 when moving from TPU to GPU (bs=64 with TPU -&gt; bs=8 with GPU) due to larger TPU memory, and adjust learning rate approximately by the same magnitude. On TPU, the batch data is divided into 8 shards and the each shard, I assume, is treated almost independently (forward and backward passes).<br>\nSee an example for batch size change here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a>   (the learning rate here must be changed manually, but I saw notebooks with automatic learning rate changes, too)</p>\n<p>In any case, there should be a difference in the scores, but not by so much.</p>",
      "rawMarkdown": "What people did in some other notebooks  - divide  the batch size  by 8 when moving from TPU to GPU (bs=64 with TPU -> bs=8 with GPU) due to larger TPU memory, and adjust learning rate approximately by the same magnitude. On TPU, the batch data is divided into 8 shards and the each shard, I assume, is treated almost independently (forward and backward passes).\nSee an example for batch size change here: https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train   (the learning rate here must be changed manually, but I saw notebooks with automatic learning rate changes, too)\n\nIn any case, there should be a difference in the scores, but not by so much.",
      "votes": null
    },
    {
      "id": "1139903",
      "postDate": "01/05/2021 17:31:53",
      "content": "<p>Ok thank you !<br>\nIf it is just a question of batch size and learning rate then it is fine. I ll tune the gpu model and should be able to get a similar score </p>",
      "rawMarkdown": "Ok thank you !\nIf it is just a question of batch size and learning rate then it is fine. I ll tune the gpu model and should be able to get a similar score",
      "votes": null
    },
    {
      "id": "1140762",
      "postDate": "01/06/2021 08:49:06",
      "content": "<p>I acutally divided the learning rate by 8 without changing the batch size, and it gives me similar results</p>",
      "rawMarkdown": "I acutally divided the learning rate by 8 without changing the batch size, and it gives me similar results",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1137471,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "01/04/2021 00:57:22",
      "content": "<p>I also compared my own GPU model with TPU one (both are PyTorch), and noticed the same issue.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1138302,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "01/04/2021 14:59:41",
      "content": "<p>Interesting. By how much did you change the batch size and the learning rate when moving from TPU to GPU ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1138345,
          "author_name": "nyounes",
          "author_url": "",
          "post_date": "01/04/2021 16:02:38",
          "content": "<p>I did not change anything.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139019,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "01/05/2021 05:36:16",
          "content": "<p>What people did in some other notebooks  - divide  the batch size  by 8 when moving from TPU to GPU (bs=64 with TPU -&gt; bs=8 with GPU) due to larger TPU memory, and adjust learning rate approximately by the same magnitude. On TPU, the batch data is divided into 8 shards and the each shard, I assume, is treated almost independently (forward and backward passes).<br>\nSee an example for batch size change here: <a href=\"https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train\" target=\"_blank\">https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train</a>   (the learning rate here must be changed manually, but I saw notebooks with automatic learning rate changes, too)</p>\n<p>In any case, there should be a difference in the scores, but not by so much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139903,
          "author_name": "nyounes",
          "author_url": "",
          "post_date": "01/05/2021 17:31:53",
          "content": "<p>Ok thank you !<br>\nIf it is just a question of batch size and learning rate then it is fine. I ll tune the gpu model and should be able to get a similar score </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140762,
          "author_name": "nyounes",
          "author_url": "",
          "post_date": "01/06/2021 08:49:06",
          "content": "<p>I acutally divided the learning rate by 8 without changing the batch size, and it gives me similar results</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1137411": "Hello here,\n\nI have started this competition using @joshi98kishan notebook using pytorch with TPU (here: https://www.kaggle.com/joshi98kishan/inference-pytorch-tta-sub-0-84), my lb score is 0.836.\nI then switched the exact same code to GPU, so i can train local.\nAnd the lb score dropped to 0.821.\n\nHas anyone observed the same behavior ? Any hints ?\n\nThanks in advance,\n\nNicolas",
    "1137471": "I also compared my own GPU model with TPU one (both are PyTorch), and noticed the same issue.",
    "1138302": "Interesting. By how much did you change the batch size and the learning rate when moving from TPU to GPU ?",
    "1138345": "I did not change anything.",
    "1139019": "What people did in some other notebooks  - divide  the batch size  by 8 when moving from TPU to GPU (bs=64 with TPU -> bs=8 with GPU) due to larger TPU memory, and adjust learning rate approximately by the same magnitude. On TPU, the batch data is divided into 8 shards and the each shard, I assume, is treated almost independently (forward and backward passes).\nSee an example for batch size change here: https://www.kaggle.com/wrrosa/hubmap-tf-with-tpu-efficientunet-512x512-train   (the learning rate here must be changed manually, but I saw notebooks with automatic learning rate changes, too)\n\nIn any case, there should be a difference in the scores, but not by so much.",
    "1139903": "Ok thank you !\nIf it is just a question of batch size and learning rate then it is fine. I ll tune the gpu model and should be able to get a similar score",
    "1140762": "I acutally divided the learning rate by 8 without changing the batch size, and it gives me similar results"
  },
  "source": "meta"
}