{
  "id": 214141,
  "title": "Why my kernel restart when I train with TPU?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214141",
  "author_name": "",
  "post_date": "2021-01-25T12:30:24.362991900Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am training my neural network,but I need to train it with 5 fold cross validation,10 epochs per fold.I train efficientnetb4 on GPU,the ram usage is about 2.5G,but train on GPU will spend 2.5 hours per fold.With the limitation of 9 hours GPU,I can't train my neural network with cross validation.So I would like to migrate my code to TPU and use Pytorch xla to accelerate my code,but if I use the same batch size,it will be out of memory immediately,even if I decrease my batch size to 4 images,it still need 10G ram,which is very dangerous,because the kernel will restart if I use 12G ram.So why TPU use 4 times ram compared with GPU?What should I do to decrease the memory and train successfully?</p>",
  "messages": [
    {
      "id": "1169308",
      "postDate": "01/25/2021 12:30:24",
      "content": "<p>I am training my neural network,but I need to train it with 5 fold cross validation,10 epochs per fold.I train efficientnetb4 on GPU,the ram usage is about 2.5G,but train on GPU will spend 2.5 hours per fold.With the limitation of 9 hours GPU,I can't train my neural network with cross validation.So I would like to migrate my code to TPU and use Pytorch xla to accelerate my code,but if I use the same batch size,it will be out of memory immediately,even if I decrease my batch size to 4 images,it still need 10G ram,which is very dangerous,because the kernel will restart if I use 12G ram.So why TPU use 4 times ram compared with GPU?What should I do to decrease the memory and train successfully?</p>",
      "rawMarkdown": "I am training my neural network,but I need to train it with 5 fold cross validation,10 epochs per fold.I train efficientnetb4 on GPU,the ram usage is about 2.5G,but train on GPU will spend 2.5 hours per fold.With the limitation of 9 hours GPU,I can't train my neural network with cross validation.So I would like to migrate my code to TPU and use Pytorch xla to accelerate my code,but if I use the same batch size,it will be out of memory immediately,even if I decrease my batch size to 4 images,it still need 10G ram,which is very dangerous,because the kernel will restart if I use 12G ram.So why TPU use 4 times ram compared with GPU?What should I do to decrease the memory and train successfully?",
      "votes": null
    },
    {
      "id": "1169541",
      "postDate": "01/25/2021 14:50:46",
      "content": "<p>commit your kernel for fold 0 training,once done,train fold 1 in next kernel commit,,then fold 2,,,this way you can train  5 folds and it's safer while using pytorch tpu</p>",
      "rawMarkdown": "commit your kernel for fold 0 training,once done,train fold 1 in next kernel commit,,then fold 2,,,this way you can train  5 folds and it's safer while using pytorch tpu",
      "votes": null
    },
    {
      "id": "1169704",
      "postDate": "01/25/2021 17:31:20",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">Mobassir</a> gave you the best answer - to speed up my learning time on local machines I almost always run just one of the 5 folds - simple <strong>if and pass</strong> added to the code does the trick.  If I really love the result than I run all 5 folds.  </p>\n<p>One thing to check is to insure that your getting 8 replicas when you do your tpu configuration for TPU.</p>\n<p>An issue with use of TPU is that you only get 3 hours vs the 9 hours on GPU so possible you might still run out of time doing all 5 folds.</p>",
      "rawMarkdown": "[Mobassir](https://www.kaggle.com/mobassir) gave you the best answer - to speed up my learning time on local machines I almost always run just one of the 5 folds - simple **if and pass** added to the code does the trick.  If I really love the result than I run all 5 folds.  \n\nOne thing to check is to insure that your getting 8 replicas when you do your tpu configuration for TPU.\n\nAn issue with use of TPU is that you only get 3 hours vs the 9 hours on GPU so possible you might still run out of time doing all 5 folds.",
      "votes": null
    },
    {
      "id": "1169708",
      "postDate": "01/25/2021 17:38:16",
      "content": "<p>correction : <br>\n<strong>you get 9 hour for TPU and 9 Hour for GPU</strong><br>\nkaggle now giving us 9 hour tpu for each and every kernel commit</p>",
      "rawMarkdown": "correction : \n**you get 9 hour for TPU and 9 Hour for GPU**\nkaggle now giving us 9 hour tpu for each and every kernel commit",
      "votes": null
    },
    {
      "id": "1169718",
      "postDate": "01/25/2021 17:48:47",
      "content": "<p>Wow - thanks for the update.</p>",
      "rawMarkdown": "Wow - thanks for the update.",
      "votes": null
    },
    {
      "id": "1170345",
      "postDate": "01/26/2021 06:50:55",
      "content": "<p>Thank you for your help. I once submitted one fold in each kernel. Although this was very effective, it was too troublesome for me because I needed to manage the operation of multiple kernels at the same time instead of one.</p>",
      "rawMarkdown": "Thank you for your help. I once submitted one fold in each kernel. Although this was very effective, it was too troublesome for me because I needed to manage the operation of multiple kernels at the same time instead of one.",
      "votes": null
    },
    {
      "id": "1170351",
      "postDate": "01/26/2021 06:58:10",
      "content": "<p>Thanks for your help.I tried this method before,it's effective,but I still want to know,if high single model accuracy means high 5fold model ensemble accuracy?I hope to train one model each time instead of five model becaure train five model will consume me a lot of GPU time,but if the accuracy of each model is similar?Which means I can use one model accuracy to evaluate my five model ensemble accuracy?</p>",
      "rawMarkdown": "Thanks for your help.I tried this method before,it's effective,but I still want to know,if high single model accuracy means high 5fold model ensemble accuracy?I hope to train one model each time instead of five model becaure train five model will consume me a lot of GPU time,but if the accuracy of each model is similar?Which means I can use one model accuracy to evaluate my five model ensemble accuracy?",
      "votes": null
    },
    {
      "id": "1170420",
      "postDate": "01/26/2021 07:46:02",
      "content": "<p><a href=\"https://www.kaggle.com/jianxinhu\" target=\"_blank\">@jianxinhu</a> follow this for kfold : <a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training</a></p>",
      "rawMarkdown": "jianxinhu follow this for kfold : https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training",
      "votes": null
    },
    {
      "id": "1170628",
      "postDate": "01/26/2021 10:32:50",
      "content": "<p>thanks,it's amazing. </p>",
      "rawMarkdown": "thanks,it's amazing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1169541,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "01/25/2021 14:50:46",
      "content": "<p>commit your kernel for fold 0 training,once done,train fold 1 in next kernel commit,,then fold 2,,,this way you can train  5 folds and it's safer while using pytorch tpu</p>",
      "votes": null,
      "replies": [
        {
          "id": 1170345,
          "author_name": "jianxinhu",
          "author_url": "",
          "post_date": "01/26/2021 06:50:55",
          "content": "<p>Thank you for your help. I once submitted one fold in each kernel. Although this was very effective, it was too troublesome for me because I needed to manage the operation of multiple kernels at the same time instead of one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1170420,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/26/2021 07:46:02",
          "content": "<p><a href=\"https://www.kaggle.com/jianxinhu\" target=\"_blank\">@jianxinhu</a> follow this for kfold : <a href=\"https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training\" target=\"_blank\">https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1170628,
          "author_name": "jianxinhu",
          "author_url": "",
          "post_date": "01/26/2021 10:32:50",
          "content": "<p>thanks,it's amazing. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1169704,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/25/2021 17:31:20",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">Mobassir</a> gave you the best answer - to speed up my learning time on local machines I almost always run just one of the 5 folds - simple <strong>if and pass</strong> added to the code does the trick.  If I really love the result than I run all 5 folds.  </p>\n<p>One thing to check is to insure that your getting 8 replicas when you do your tpu configuration for TPU.</p>\n<p>An issue with use of TPU is that you only get 3 hours vs the 9 hours on GPU so possible you might still run out of time doing all 5 folds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1169708,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/25/2021 17:38:16",
          "content": "<p>correction : <br>\n<strong>you get 9 hour for TPU and 9 Hour for GPU</strong><br>\nkaggle now giving us 9 hour tpu for each and every kernel commit</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169718,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/25/2021 17:48:47",
          "content": "<p>Wow - thanks for the update.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1170351,
          "author_name": "jianxinhu",
          "author_url": "",
          "post_date": "01/26/2021 06:58:10",
          "content": "<p>Thanks for your help.I tried this method before,it's effective,but I still want to know,if high single model accuracy means high 5fold model ensemble accuracy?I hope to train one model each time instead of five model becaure train five model will consume me a lot of GPU time,but if the accuracy of each model is similar?Which means I can use one model accuracy to evaluate my five model ensemble accuracy?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1169308": "I am training my neural network,but I need to train it with 5 fold cross validation,10 epochs per fold.I train efficientnetb4 on GPU,the ram usage is about 2.5G,but train on GPU will spend 2.5 hours per fold.With the limitation of 9 hours GPU,I can't train my neural network with cross validation.So I would like to migrate my code to TPU and use Pytorch xla to accelerate my code,but if I use the same batch size,it will be out of memory immediately,even if I decrease my batch size to 4 images,it still need 10G ram,which is very dangerous,because the kernel will restart if I use 12G ram.So why TPU use 4 times ram compared with GPU?What should I do to decrease the memory and train successfully?",
    "1169541": "commit your kernel for fold 0 training,once done,train fold 1 in next kernel commit,,then fold 2,,,this way you can train  5 folds and it's safer while using pytorch tpu",
    "1169704": "[Mobassir](https://www.kaggle.com/mobassir) gave you the best answer - to speed up my learning time on local machines I almost always run just one of the 5 folds - simple **if and pass** added to the code does the trick.  If I really love the result than I run all 5 folds.  \n\nOne thing to check is to insure that your getting 8 replicas when you do your tpu configuration for TPU.\n\nAn issue with use of TPU is that you only get 3 hours vs the 9 hours on GPU so possible you might still run out of time doing all 5 folds.",
    "1169708": "correction : \n**you get 9 hour for TPU and 9 Hour for GPU**\nkaggle now giving us 9 hour tpu for each and every kernel commit",
    "1169718": "Wow - thanks for the update.",
    "1170345": "Thank you for your help. I once submitted one fold in each kernel. Although this was very effective, it was too troublesome for me because I needed to manage the operation of multiple kernels at the same time instead of one.",
    "1170351": "Thanks for your help.I tried this method before,it's effective,but I still want to know,if high single model accuracy means high 5fold model ensemble accuracy?I hope to train one model each time instead of five model becaure train five model will consume me a lot of GPU time,but if the accuracy of each model is similar?Which means I can use one model accuracy to evaluate my five model ensemble accuracy?",
    "1170420": "jianxinhu follow this for kfold : https://www.kaggle.com/tanlikesmath/cassava-pytorch-xla-tpu-starter-training",
    "1170628": "thanks,it's amazing."
  },
  "source": "meta"
}