{
  "id": 149815,
  "title": "PyTorch TPU Commit Failed [Complete. Exited with code 0.]",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/149815",
  "author_name": "Kerem Turgutlu",
  "post_date": "2020-05-10T02:48:42.425000",
  "votes": 2,
  "comment_count": 19,
  "views": 0,
  "content": "<p>My <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/\">notebook</a> runs fine end to end when I use it in interactive mode with Run All Cells. But when I commit it fails without any error message - it only says failed but <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/log?scriptVersionId=33667547\">logs</a> don't include any error. </p>",
  "messages": [
    {
      "id": 841002,
      "postDate": "2020-05-10T14:38:02.717Z",
      "content": "<p>I strongly suggest you to read <a href=\"https://github.com/pytorch/xla/issues/1870\">this</a> )in case you didn't) thread out a couple of times! There are a lot of tips lying around that can help you with not having OOM as such</p>",
      "rawMarkdown": "I strongly suggest you to read [this](https://github.com/pytorch/xla/issues/1870) )in case you didn't) thread out a couple of times! There are a lot of tips lying around that can help you with not having OOM as such",
      "votes": 3
    },
    {
      "id": 840633,
      "postDate": "2020-05-10T08:21:03.377Z",
      "content": "<p>Last week i wasted all my tpu hours because of this problem of pytorch tpu in this notebook : <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline</a></p>\n\n<p>If you check the changelog section then you will understand.\n2 epoch took 30 minutes but for 8 epoch kernel failed\nAgain same model last night i tried for 7 epoch and it works fine\nIt is possible somewhere in my code the program is having oom problem that i am not able to catch at this moments or there could be pytorch tpu problem.let me know please if you find the solution </p>",
      "rawMarkdown": "Last week i wasted all my tpu hours because of this problem of pytorch tpu in this notebook : https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\n\nIf you check the changelog section then you will understand.\n2 epoch took 30 minutes but for 8 epoch kernel failed\nAgain same model last night i tried for 7 epoch and it works fine\nIt is possible somewhere in my code the program is having oom problem that i am not able to catch at this moments or there could be pytorch tpu problem.let me know please if you find the solution ",
      "votes": 1,
      "replies": [
        {
          "id": 840901,
          "postDate": "2020-05-10T13:29:42.810Z",
          "content": "<p>In my case kernel fails within 3-4 mins after commit, so it either fails at the very beginning of the training or even before that. The funny thing is I am able to run it interactively just fine but when it's committed problem arises.</p>",
          "rawMarkdown": "In my case kernel fails within 3-4 mins after commit, so it either fails at the very beginning of the training or even before that. The funny thing is I am able to run it interactively just fine but when it's committed problem arises.",
          "votes": 1
        },
        {
          "id": 840936,
          "postDate": "2020-05-10T13:58:19.407Z",
          "content": "<p><a href=\"/keremt\">@keremt</a>  i just noticed this issue few minutes ago and i realized it is happening only for this competition today because a lot of people are using tpu in kaggle and kaggle putting us on queue ,so maybe you will like to try again after few hours like i will do?</p>",
          "rawMarkdown": "@keremt  i just noticed this issue few minutes ago and i realized it is happening only for this competition today because a lot of people are using tpu in kaggle and kaggle putting us on queue ,so maybe you will like to try again after few hours like i will do?",
          "votes": 1
        },
        {
          "id": 840967,
          "postDate": "2020-05-10T14:16:08.607Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Yeah, actually now even the interactive version started to fail with OOM error. Don't know if it is related to your observation.</p>",
          "rawMarkdown": "@mobassir Yeah, actually now even the interactive version started to fail with OOM error. Don't know if it is related to your observation.",
          "votes": 1
        },
        {
          "id": 840977,
          "postDate": "2020-05-10T14:24:08.427Z",
          "content": "<p><a href=\"/keremt\">@keremt</a>  i see you are using batch_size = 32\ndidn't try reducing it down to 16 to avoid OOM error?</p>",
          "rawMarkdown": "@keremt  i see you are using batch_size = 32\ndidn't try reducing it down to 16 to avoid OOM error?"
        },
        {
          "id": 840981,
          "postDate": "2020-05-10T14:26:07.420Z",
          "content": "<p>Yeah I did try 16, failed again. But then finally switch back to 32 since it was working in interactive mode, so I thought batch size can't be an issue.</p>",
          "rawMarkdown": "Yeah I did try 16, failed again. But then finally switch back to 32 since it was working in interactive mode, so I thought batch size can't be an issue."
        },
        {
          "id": 840995,
          "postDate": "2020-05-10T14:34:10.570Z",
          "content": "<p>I commited just now and it has been running for 15 mins now and haven't failed yet. Your theory might be true about TPU usage or there is a random error that we have in or code.</p>",
          "rawMarkdown": "I commited just now and it has been running for 15 mins now and haven't failed yet. Your theory might be true about TPU usage or there is a random error that we have in or code.",
          "votes": 1
        },
        {
          "id": 841032,
          "postDate": "2020-05-10T14:48:53.477Z",
          "content": "<p><a href=\"/keremt\">@keremt</a>  please let me know if it works,it will help me to save some tpu hours next time</p>",
          "rawMarkdown": "@keremt  please let me know if it works,it will help me to save some tpu hours next time"
        },
        {
          "id": 841049,
          "postDate": "2020-05-10T14:56:00.400Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> here is the successful <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style\">run</a> (version 30).</p>",
          "rawMarkdown": "@mobassir here is the successful [run](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style) (version 30).",
          "votes": 1
        },
        {
          "id": 841211,
          "postDate": "2020-05-10T17:11:46.457Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 841212,
          "postDate": "2020-05-10T17:12:28.453Z",
          "content": "<p>thats great <a href=\"/keremt\">@keremt</a>  thanks for sharing</p>",
          "rawMarkdown": "thats great @keremt  thanks for sharing"
        },
        {
          "id": 841525,
          "postDate": "2020-05-10T20:42:58.630Z",
          "content": "<p>The problem seems to be inconsistent. I was able to run the kernel with more epochs but it stopped after 3 hrs, I think that is the limit. Then I lowered the number of epochs to finish the kernel in less than 3 hours but this time it again started failing as before.</p>",
          "rawMarkdown": "The problem seems to be inconsistent. I was able to run the kernel with more epochs but it stopped after 3 hrs, I think that is the limit. Then I lowered the number of epochs to finish the kernel in less than 3 hours but this time it again started failing as before.",
          "votes": 1
        },
        {
          "id": 841529,
          "postDate": "2020-05-10T20:49:09.587Z",
          "content": "<p><a href=\"/keremt\">@keremt</a>  as i said you,same thing happened with me in this kernel :  <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline</a></p>\n\n<p>1 epoch took 15 minutes so i tried for 10 epochs and got error after 3 hours,then i tried for  8 epoch then again got error after 3 hours then my tpu quota vanished,so again after 2 days i tried for 7 epoch and now you can see it is working, hard to tell what's wrong\nbut i don't face such issue with keras and also i noticed that keras tpu training is much faster than pytorch tpu</p>",
          "rawMarkdown": "@keremt  as i said you,same thing happened with me in this kernel :  https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\n\n1 epoch took 15 minutes so i tried for 10 epochs and got error after 3 hours,then i tried for  8 epoch then again got error after 3 hours then my tpu quota vanished,so again after 2 days i tried for 7 epoch and now you can see it is working, hard to tell what's wrong\nbut i don't face such issue with keras and also i noticed that keras tpu training is much faster than pytorch tpu",
          "votes": 1
        },
        {
          "id": 841534,
          "postDate": "2020-05-10T20:53:54.677Z",
          "content": "<p>Might as well try Keras then, thanks!</p>",
          "rawMarkdown": "Might as well try Keras then, thanks!",
          "votes": 1
        },
        {
          "id": 844274,
          "postDate": "2020-05-12T14:39:07.013Z",
          "content": "<p>I implemented the same code in Keras, <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-keras\">here</a>  which might not be perfect since I usually use pytorch. It looks like keras version is 6x faster than Pytorch TPU. Pytorch TPU is 20x times faster than the fastai GPU which makes keras TPU 120x faster than what I would usually use. Of course these are very initial observations :)</p>",
          "rawMarkdown": "I implemented the same code in Keras, [here](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-keras)  which might not be perfect since I usually use pytorch. It looks like keras version is 6x faster than Pytorch TPU. Pytorch TPU is 20x times faster than the fastai GPU which makes keras TPU 120x faster than what I would usually use. Of course these are very initial observations :)",
          "votes": 1
        },
        {
          "id": 844322,
          "postDate": "2020-05-12T14:58:11.617Z",
          "content": "<p><a href=\"/keremt\">@keremt</a> my observations are also similar, in this kernel : <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline?scriptVersionId=33790876\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline?scriptVersionId=33790876</a>\ni have been spending a lot of time in this kernel brother but somewhere something is missing and it's hard for me to catch,if you check this kernel,i tried to do almost same thing like xhlulu did in his tf tpu kernel of that same  competition,you can see here i am using same train-val split like xhlulu did with same seed,i am using efficientnetb3 like him and also tried adaptive pooling like him,,the only differences are :</p>\n\n<p>loss function - he used crossentropy and i tried bcewithlogits or mseloss in some versions\n2.scheduler - i am using steplr or 1cycle\nand almost all other things are same,but here are some problems i am facing : </p>\n\n<ol>\n<li>his kernel commit finishes within 2 hours for 10 epoch but my commit takes 3+ hours for 10 epoch\nhe used batch size = 16 and i can't use batch size 16 cause it gives oom error so i am using 10 instead</li>\n</ol>\n\n<p>i am trying to understand where i am making mistakes that is causing such massive gap between 2 same models performance written in 2 different framework</p>\n\n<p>you can see from train.log and  code that i am using 8 core for model training,same data,same images,same model but huge different in training time and model performance, hard to understand what's wrong</p>",
          "rawMarkdown": "@keremt my observations are also similar, in this kernel : https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline?scriptVersionId=33790876\ni have been spending a lot of time in this kernel brother but somewhere something is missing and it's hard for me to catch,if you check this kernel,i tried to do almost same thing like xhlulu did in his tf tpu kernel of that same  competition,you can see here i am using same train-val split like xhlulu did with same seed,i am using efficientnetb3 like him and also tried adaptive pooling like him,,the only differences are :\n\nloss function - he used crossentropy and i tried bcewithlogits or mseloss in some versions\n2.scheduler - i am using steplr or 1cycle\nand almost all other things are same,but here are some problems i am facing : \n\n1. his kernel commit finishes within 2 hours for 10 epoch but my commit takes 3+ hours for 10 epoch\nhe used batch size = 16 and i can't use batch size 16 cause it gives oom error so i am using 10 instead\n\ni am trying to understand where i am making mistakes that is causing such massive gap between 2 same models performance written in 2 different framework\n\nyou can see from train.log and  code that i am using 8 core for model training,same data,same images,same model but huge different in training time and model performance, hard to understand what's wrong",
          "votes": 1
        },
        {
          "id": 844331,
          "postDate": "2020-05-12T15:03:43.783Z",
          "content": "<p>Interesting, I might take a look at this competition if I have some time but efficientnet is large memory intensive model compared to other models like resnets - and my observation is that keras TPU is much better with memory utilization compared to pytorch TPU it may be due to the fact that pytorch copies stuff over CPU and TPU device. Have you found anything useful in this <a href=\"https://github.com/pytorch/xla/issues/1870\">link</a> that <a href=\"/adityaecdrid\">@adityaecdrid</a> shared for optimizing your code? If possible I would stick with Keras for TPUs for now.</p>",
          "rawMarkdown": "Interesting, I might take a look at this competition if I have some time but efficientnet is large memory intensive model compared to other models like resnets - and my observation is that keras TPU is much better with memory utilization compared to pytorch TPU it may be due to the fact that pytorch copies stuff over CPU and TPU device. Have you found anything useful in this [link](https://github.com/pytorch/xla/issues/1870) that @adityaecdrid shared for optimizing your code? If possible I would stick with Keras for TPUs for now.",
          "votes": 1
        },
        {
          "id": 844341,
          "postDate": "2020-05-12T15:08:49.353Z",
          "content": "<p>i couldn't  check that because i am working on several competitions,but 2 days ago i created a github issue ,this one : <a href=\"https://github.com/pytorch/xla/issues/2054#issuecomment-627367729\">https://github.com/pytorch/xla/issues/2054#issuecomment-627367729</a></p>\n\n<p>if you check the  last comment then it answers everything, i just checked that last comment few minutes ago and sharing with you now</p>",
          "rawMarkdown": "i couldn't  check that because i am working on several competitions,but 2 days ago i created a github issue ,this one : https://github.com/pytorch/xla/issues/2054#issuecomment-627367729\n\nif you check the  last comment then it answers everything, i just checked that last comment few minutes ago and sharing with you now",
          "votes": 1
        }
      ]
    },
    {
      "id": 840427,
      "postDate": "2020-05-10T02:48:42.427Z",
      "content": "<p>My <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/\">notebook</a> runs fine end to end when I use it in interactive mode with Run All Cells. But when I commit it fails without any error message - it only says failed but <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/log?scriptVersionId=33667547\">logs</a> don't include any error. </p>",
      "rawMarkdown": "My [notebook](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/) runs fine end to end when I use it in interactive mode with Run All Cells. But when I commit it fails without any error message - it only says failed but [logs](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/log?scriptVersionId=33667547) don't include any error. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 841002,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-05-10T14:38:02.717000",
      "content": "<p>I strongly suggest you to read <a href=\"https://github.com/pytorch/xla/issues/1870\">this</a> )in case you didn't) thread out a couple of times! There are a lot of tips lying around that can help you with not having OOM as such</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 840633,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-05-10T08:21:03.377000",
      "content": "<p>Last week i wasted all my tpu hours because of this problem of pytorch tpu in this notebook : <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline</a></p>\n\n<p>If you check the changelog section then you will understand.\n2 epoch took 30 minutes but for 8 epoch kernel failed\nAgain same model last night i tried for 7 epoch and it works fine\nIt is possible somewhere in my code the program is having oom problem that i am not able to catch at this moments or there could be pytorch tpu problem.let me know please if you find the solution </p>",
      "votes": 1,
      "replies": [
        {
          "id": 840901,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T13:29:42.810000",
          "content": "<p>In my case kernel fails within 3-4 mins after commit, so it either fails at the very beginning of the training or even before that. The funny thing is I am able to run it interactively just fine but when it's committed problem arises.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 840936,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-10T13:58:19.407000",
          "content": "<p><a href=\"/keremt\">@keremt</a>  i just noticed this issue few minutes ago and i realized it is happening only for this competition today because a lot of people are using tpu in kaggle and kaggle putting us on queue ,so maybe you will like to try again after few hours like i will do?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 840967,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T14:16:08.607000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Yeah, actually now even the interactive version started to fail with OOM error. Don't know if it is related to your observation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 840977,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-10T14:24:08.427000",
          "content": "<p><a href=\"/keremt\">@keremt</a>  i see you are using batch_size = 32\ndidn't try reducing it down to 16 to avoid OOM error?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840981,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T14:26:07.420000",
          "content": "<p>Yeah I did try 16, failed again. But then finally switch back to 32 since it was working in interactive mode, so I thought batch size can't be an issue.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 840995,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T14:34:10.570000",
          "content": "<p>I commited just now and it has been running for 15 mins now and haven't failed yet. Your theory might be true about TPU usage or there is a random error that we have in or code.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 841032,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-10T14:48:53.477000",
          "content": "<p><a href=\"/keremt\">@keremt</a>  please let me know if it works,it will help me to save some tpu hours next time</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841049,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T14:56:00.400000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> here is the successful <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style\">run</a> (version 30).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 841211,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-10T17:11:46.457000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841212,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-10T17:12:28.453000",
          "content": "<p>thats great <a href=\"/keremt\">@keremt</a>  thanks for sharing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 841525,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T20:42:58.630000",
          "content": "<p>The problem seems to be inconsistent. I was able to run the kernel with more epochs but it stopped after 3 hrs, I think that is the limit. Then I lowered the number of epochs to finish the kernel in less than 3 hours but this time it again started failing as before.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 841529,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-10T20:49:09.587000",
          "content": "<p><a href=\"/keremt\">@keremt</a>  as i said you,same thing happened with me in this kernel :  <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline</a></p>\n\n<p>1 epoch took 15 minutes so i tried for 10 epochs and got error after 3 hours,then i tried for  8 epoch then again got error after 3 hours then my tpu quota vanished,so again after 2 days i tried for 7 epoch and now you can see it is working, hard to tell what's wrong\nbut i don't face such issue with keras and also i noticed that keras tpu training is much faster than pytorch tpu</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 841534,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-10T20:53:54.677000",
          "content": "<p>Might as well try Keras then, thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 844274,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-12T14:39:07.013000",
          "content": "<p>I implemented the same code in Keras, <a href=\"https://www.kaggle.com/keremt/xlm-roberta-tpu-training-keras\">here</a>  which might not be perfect since I usually use pytorch. It looks like keras version is 6x faster than Pytorch TPU. Pytorch TPU is 20x times faster than the fastai GPU which makes keras TPU 120x faster than what I would usually use. Of course these are very initial observations :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 844322,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-12T14:58:11.617000",
          "content": "<p><a href=\"/keremt\">@keremt</a> my observations are also similar, in this kernel : <a href=\"https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline?scriptVersionId=33790876\">https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline?scriptVersionId=33790876</a>\ni have been spending a lot of time in this kernel brother but somewhere something is missing and it's hard for me to catch,if you check this kernel,i tried to do almost same thing like xhlulu did in his tf tpu kernel of that same  competition,you can see here i am using same train-val split like xhlulu did with same seed,i am using efficientnetb3 like him and also tried adaptive pooling like him,,the only differences are :</p>\n\n<p>loss function - he used crossentropy and i tried bcewithlogits or mseloss in some versions\n2.scheduler - i am using steplr or 1cycle\nand almost all other things are same,but here are some problems i am facing : </p>\n\n<ol>\n<li>his kernel commit finishes within 2 hours for 10 epoch but my commit takes 3+ hours for 10 epoch\nhe used batch size = 16 and i can't use batch size 16 cause it gives oom error so i am using 10 instead</li>\n</ol>\n\n<p>i am trying to understand where i am making mistakes that is causing such massive gap between 2 same models performance written in 2 different framework</p>\n\n<p>you can see from train.log and  code that i am using 8 core for model training,same data,same images,same model but huge different in training time and model performance, hard to understand what's wrong</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 844331,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-05-12T15:03:43.783000",
          "content": "<p>Interesting, I might take a look at this competition if I have some time but efficientnet is large memory intensive model compared to other models like resnets - and my observation is that keras TPU is much better with memory utilization compared to pytorch TPU it may be due to the fact that pytorch copies stuff over CPU and TPU device. Have you found anything useful in this <a href=\"https://github.com/pytorch/xla/issues/1870\">link</a> that <a href=\"/adityaecdrid\">@adityaecdrid</a> shared for optimizing your code? If possible I would stick with Keras for TPUs for now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 844341,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-05-12T15:08:49.353000",
          "content": "<p>i couldn't  check that because i am working on several competitions,but 2 days ago i created a github issue ,this one : <a href=\"https://github.com/pytorch/xla/issues/2054#issuecomment-627367729\">https://github.com/pytorch/xla/issues/2054#issuecomment-627367729</a></p>\n\n<p>if you check the  last comment then it answers everything, i just checked that last comment few minutes ago and sharing with you now</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "841002": "I strongly suggest you to read [this](https://github.com/pytorch/xla/issues/1870) )in case you didn't) thread out a couple of times! There are a lot of tips lying around that can help you with not having OOM as such",
    "840633": "Last week i wasted all my tpu hours because of this problem of pytorch tpu in this notebook : https://www.kaggle.com/mobassir/pytorch-tpu-transfer-learning-baseline\n\nIf you check the changelog section then you will understand.\n2 epoch took 30 minutes but for 8 epoch kernel failed\nAgain same model last night i tried for 7 epoch and it works fine\nIt is possible somewhere in my code the program is having oom problem that i am not able to catch at this moments or there could be pytorch tpu problem.let me know please if you find the solution ",
    "840427": "My [notebook](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/) runs fine end to end when I use it in interactive mode with Run All Cells. But when I commit it fails without any error message - it only says failed but [logs](https://www.kaggle.com/keremt/xlm-roberta-tpu-training-fastai-style/log?scriptVersionId=33667547) don't include any error. "
  }
}