{
  "id": 270139,
  "title": "Pytorch Lightning with TPU OOM in Kaggle Kernel",
  "url": "/competitions/landmark-recognition-2021/discussion/270139",
  "author_name": "Jihun Lorenzo Park",
  "post_date": "2021-09-03T18:05:20.641000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Is there anyone who managed the TPU training successful with pytorch lightning in kaggle kernel?<br>\nI am getting out of memory error (OOM) when I set tpu_cores=8,,<br>\nIt only works when I set batch size very small e.g. 2 or 4, but as far as I know, it is wrong usage since ddp backend tries to divide the samples in a batch to multiple device and also it is very slow of course.<br>\nI really appreciate if you give any tips for the successful usage,,,</p>\n<p>I am searching for the github issues and kaggle discussions but I can't find the solution for it</p>",
  "messages": [
    {
      "id": 1502015,
      "postDate": "2021-09-03T18:05:20.640Z",
      "content": "<p>Is there anyone who managed the TPU training successful with pytorch lightning in kaggle kernel?<br>\nI am getting out of memory error (OOM) when I set tpu_cores=8,,<br>\nIt only works when I set batch size very small e.g. 2 or 4, but as far as I know, it is wrong usage since ddp backend tries to divide the samples in a batch to multiple device and also it is very slow of course.<br>\nI really appreciate if you give any tips for the successful usage,,,</p>\n<p>I am searching for the github issues and kaggle discussions but I can't find the solution for it</p>",
      "rawMarkdown": "Is there anyone who managed the TPU training successful with pytorch lightning in kaggle kernel?\nI am getting out of memory error (OOM) when I set tpu_cores=8,,\nIt only works when I set batch size very small e.g. 2 or 4, but as far as I know, it is wrong usage since ddp backend tries to divide the samples in a batch to multiple device and also it is very slow of course.\nI really appreciate if you give any tips for the successful usage,,,\n\nI am searching for the github issues and kaggle discussions but I can't find the solution for it",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1502015": "Is there anyone who managed the TPU training successful with pytorch lightning in kaggle kernel?\nI am getting out of memory error (OOM) when I set tpu_cores=8,,\nIt only works when I set batch size very small e.g. 2 or 4, but as far as I know, it is wrong usage since ddp backend tries to divide the samples in a batch to multiple device and also it is very slow of course.\nI really appreciate if you give any tips for the successful usage,,,\n\nI am searching for the github issues and kaggle discussions but I can't find the solution for it"
  }
}