{
  "id": 143005,
  "title": "PyTorch XLA/TPU training",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143005",
  "author_name": "ilovescience",
  "post_date": "2020-04-13T11:59:07.404000",
  "votes": 27,
  "comment_count": 11,
  "views": 0,
  "content": "<p>In <a href=\"/xhlulu\">@xhlulu</a>'s amazing <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">kernel</a>, an XLM-RoBERTA-large model is trained on the TPU using TensorFlow. I wanted to also train this, but with PyTorch. In <a href=\"/abhishek\">@abhishek</a>'s great <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation\">kernel</a>, he uses PyTorch XLA to train a BERT-multilingual-uncased model on the TPU. However, if I switch out the BERT model for the XLM-RoBERTa model (even base model), I got VM OOM errors. Discussing this with other competitors, we had originally concluded it is impossible to train XLM-RoBERTa in Kaggle TPU Kernels. However, I decided to investigate further, and I learned how to better manage my VM memory and eventually get the kernel working.Here are a few things that I did to master PyTorch XLA training :)</p>\n\n<ol>\n<li><p>First I tokenized/encoded the data ahead of time in a separate kernel and saved it. Then I loaded it into the kernel and used it to create a TensorDataset to pass data to the model. This was already done by <a href=\"/xhlulu\">@xhlulu</a>'s kernel as well. This might be better than on-the-fly tokenization/encoding used in <a href=\"/abhishek\">@abhishek</a>'s that might be taking up VM memory during training.</p></li>\n<li><p>I always made sure to delete any unused variables  like the DataFrames (after the creation of a PyTorch dataset) and garbage collect.</p></li>\n<li><p>Even these two changes were not enough. I then discussed this with the PyTorch XLA team on GitHub(thanks @dlilbenzi and <a href=\"/jysohn23\">@jysohn23</a>) and apparently, they <em>just added changes to reduce host memory utilization</em>. So I use the nightly version of PyTorch XLA in my kernels. There are other ways to decrease VM memory utilization as well. I will experiment in the near future with these additional methods in order to improve memory utilization and speed of kernel.</p></li>\n</ol>\n\n<p>However there were a couple of things to watch out for. For example, you need to use bfloat16 when training with xlm-roberta-large, float32 leads to host OOM errors. Also, TPU training is <em>very senstitive</em> to the learning rate. I still am struggling to tune the learning rate.</p>\n\n<h1>Here is the kernel: <a href=\"https://www.kaggle.com/tanlikesmath/xlm-roberta-pytorch-xla-tpu\">training link</a>, <a href=\"https://www.kaggle.com/tanlikesmath/xlm-roberta-inference-pytorch-tpu-xla-1-core\">inference link</a></h1>\n\n<p>The score is terrible, but I will tune the LR soon. Also, I will add validation code soon.</p>\n\n<p>I have it working on xhlulu's subset of data, but I am trying to tune the LR a little further before releasing. Additionally, I have some additional memory optimizations in mind. Also, I plan to post a PyTorch XLA tips thread with many tips I have come across or used when training on the TPU.</p>\n\n<p>I hope this helps! As I just mentioned, there were some additional things I wanted to add, but I wanted to share with the community first. Let me know if you have any questions.</p>",
  "messages": [
    {
      "id": 806048,
      "postDate": "2020-04-13T11:59:07.403Z",
      "content": "<p>In <a href=\"/xhlulu\">@xhlulu</a>'s amazing <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">kernel</a>, an XLM-RoBERTA-large model is trained on the TPU using TensorFlow. I wanted to also train this, but with PyTorch. In <a href=\"/abhishek\">@abhishek</a>'s great <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation\">kernel</a>, he uses PyTorch XLA to train a BERT-multilingual-uncased model on the TPU. However, if I switch out the BERT model for the XLM-RoBERTa model (even base model), I got VM OOM errors. Discussing this with other competitors, we had originally concluded it is impossible to train XLM-RoBERTa in Kaggle TPU Kernels. However, I decided to investigate further, and I learned how to better manage my VM memory and eventually get the kernel working.Here are a few things that I did to master PyTorch XLA training :)</p>\n\n<ol>\n<li><p>First I tokenized/encoded the data ahead of time in a separate kernel and saved it. Then I loaded it into the kernel and used it to create a TensorDataset to pass data to the model. This was already done by <a href=\"/xhlulu\">@xhlulu</a>'s kernel as well. This might be better than on-the-fly tokenization/encoding used in <a href=\"/abhishek\">@abhishek</a>'s that might be taking up VM memory during training.</p></li>\n<li><p>I always made sure to delete any unused variables  like the DataFrames (after the creation of a PyTorch dataset) and garbage collect.</p></li>\n<li><p>Even these two changes were not enough. I then discussed this with the PyTorch XLA team on GitHub(thanks @dlilbenzi and <a href=\"/jysohn23\">@jysohn23</a>) and apparently, they <em>just added changes to reduce host memory utilization</em>. So I use the nightly version of PyTorch XLA in my kernels. There are other ways to decrease VM memory utilization as well. I will experiment in the near future with these additional methods in order to improve memory utilization and speed of kernel.</p></li>\n</ol>\n\n<p>However there were a couple of things to watch out for. For example, you need to use bfloat16 when training with xlm-roberta-large, float32 leads to host OOM errors. Also, TPU training is <em>very senstitive</em> to the learning rate. I still am struggling to tune the learning rate.</p>\n\n<h1>Here is the kernel: <a href=\"https://www.kaggle.com/tanlikesmath/xlm-roberta-pytorch-xla-tpu\">training link</a>, <a href=\"https://www.kaggle.com/tanlikesmath/xlm-roberta-inference-pytorch-tpu-xla-1-core\">inference link</a></h1>\n\n<p>The score is terrible, but I will tune the LR soon. Also, I will add validation code soon.</p>\n\n<p>I have it working on xhlulu's subset of data, but I am trying to tune the LR a little further before releasing. Additionally, I have some additional memory optimizations in mind. Also, I plan to post a PyTorch XLA tips thread with many tips I have come across or used when training on the TPU.</p>\n\n<p>I hope this helps! As I just mentioned, there were some additional things I wanted to add, but I wanted to share with the community first. Let me know if you have any questions.</p>",
      "rawMarkdown": "In @xhlulu's amazing [kernel](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), an XLM-RoBERTA-large model is trained on the TPU using TensorFlow. I wanted to also train this, but with PyTorch. In @abhishek's great [kernel](https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation), he uses PyTorch XLA to train a BERT-multilingual-uncased model on the TPU. However, if I switch out the BERT model for the XLM-RoBERTa model (even base model), I got VM OOM errors. Discussing this with other competitors, we had originally concluded it is impossible to train XLM-RoBERTa in Kaggle TPU Kernels. However, I decided to investigate further, and I learned how to better manage my VM memory and eventually get the kernel working.Here are a few things that I did to master PyTorch XLA training :)\n\n1. First I tokenized/encoded the data ahead of time in a separate kernel and saved it. Then I loaded it into the kernel and used it to create a TensorDataset to pass data to the model. This was already done by @xhlulu's kernel as well. This might be better than on-the-fly tokenization/encoding used in @abhishek's that might be taking up VM memory during training.\n\n2. I always made sure to delete any unused variables  like the DataFrames (after the creation of a PyTorch dataset) and garbage collect.\n\n3. Even these two changes were not enough. I then discussed this with the PyTorch XLA team on GitHub(thanks @dlilbenzi and @jysohn23) and apparently, they _just added changes to reduce host memory utilization_. So I use the nightly version of PyTorch XLA in my kernels. There are other ways to decrease VM memory utilization as well. I will experiment in the near future with these additional methods in order to improve memory utilization and speed of kernel.\n\nHowever there were a couple of things to watch out for. For example, you need to use bfloat16 when training with xlm-roberta-large, float32 leads to host OOM errors. Also, TPU training is _very senstitive_ to the learning rate. I still am struggling to tune the learning rate.\n\n# Here is the kernel: [training link](https://www.kaggle.com/tanlikesmath/xlm-roberta-pytorch-xla-tpu), [inference link](https://www.kaggle.com/tanlikesmath/xlm-roberta-inference-pytorch-tpu-xla-1-core)\n\nThe score is terrible, but I will tune the LR soon. Also, I will add validation code soon.\n\nI have it working on xhlulu's subset of data, but I am trying to tune the LR a little further before releasing. Additionally, I have some additional memory optimizations in mind. Also, I plan to post a PyTorch XLA tips thread with many tips I have come across or used when training on the TPU.\n\nI hope this helps! As I just mentioned, there were some additional things I wanted to add, but I wanted to share with the community first. Let me know if you have any questions.\n",
      "votes": 26
    },
    {
      "id": 806330,
      "postDate": "2020-04-13T16:17:12.920Z",
      "content": "<p>Here you go folks, <a href=\"https://www.kaggle.com/adityaecdrid/sample-working-xlmr-large-8-cores-tpu-pytorch\">XLMR_Large-8-Cores</a>; It was on just <strong>80k samples</strong>,;</p>\n\n<p>Stats on interactive -:</p>\n\n<p><code>\nAUC = 0.935995099948079\nAUC = 0.9306042315680167\n</code></p>\n\n<p>PS: I am running for more rows; Would be nice if people play around with LR as i will soon run out of my free TPU hours :(; Also training is sometimes unstable idk;</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> </p>",
      "rawMarkdown": "Here you go folks, [XLMR_Large-8-Cores](https://www.kaggle.com/adityaecdrid/sample-working-xlmr-large-8-cores-tpu-pytorch); It was on just **80k samples**,;\n\nStats on interactive -:\n\n```\nAUC = 0.935995099948079\nAUC = 0.9306042315680167\n```\n\nPS: I am running for more rows; Would be nice if people play around with LR as i will soon run out of my free TPU hours :(; Also training is sometimes unstable idk;\n\n@philippsinger ",
      "votes": 5
    },
    {
      "id": 810082,
      "postDate": "2020-04-16T17:56:06.257Z",
      "content": "<p>Here you should get all your OOM answers <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-tpu-xlmr-base-with-memmap\">kernel</a>; Just don't put too much pedal on bs 😅</p>\n\n<p>It uses the dataset generated from <a href=\"https://www.kaggle.com/adityaecdrid/memmap-tpu-xlmr-pytorch-pad-on-fly\">here</a>.</p>\n\n<p>In short you can simply use <code>numpy's memmap</code> and we make the dataset which is on disk as if it's on the RAM :) ; (from 10k feet that's what's happening;)</p>\n\n<p>&gt;NB I ran base just because i m out of compute; (large works as well! with no OOM)</p>\n\n<p>Also this can be further improved if you want to go about bucketing though I feel it will be an overhead if we do it;</p>\n\n<p>Happy Kaggling!</p>\n\n<p>cc <a href=\"/mgornergoogle\">@mgornergoogle</a>!</p>",
      "rawMarkdown": "Here you should get all your OOM answers [kernel](https://www.kaggle.com/adityaecdrid/pytorch-tpu-xlmr-base-with-memmap); Just don't put too much pedal on bs 😅\n\nIt uses the dataset generated from [here](https://www.kaggle.com/adityaecdrid/memmap-tpu-xlmr-pytorch-pad-on-fly).\n\nIn short you can simply use `numpy's memmap` and we make the dataset which is on disk as if it's on the RAM :) ; (from 10k feet that's what's happening;)\n\n&gt;NB I ran base just because i m out of compute; (large works as well! with no OOM)\n\nAlso this can be further improved if you want to go about bucketing though I feel it will be an overhead if we do it;\n\nHappy Kaggling!\n\ncc @mgornergoogle!",
      "votes": 1
    },
    {
      "id": 806257,
      "postDate": "2020-04-13T14:54:04.377Z",
      "content": "<p>This is a great kernel thanks. Can you also add evaluation here? I had some serious issues with loss jumping around and evaluation being stuck at 0.6 AUC. Actually your inference kernel also scores very badly... so something is similarly wrong to what I observed.</p>",
      "rawMarkdown": "This is a great kernel thanks. Can you also add evaluation here? I had some serious issues with loss jumping around and evaluation being stuck at 0.6 AUC. Actually your inference kernel also scores very badly... so something is similarly wrong to what I observed.",
      "votes": 1,
      "replies": [
        {
          "id": 806589,
          "postDate": "2020-04-13T21:11:34.200Z",
          "content": "<p>Yes I will add the validation loop soon. I had it originally but got some weird bug. But clearly it can work, as <a href=\"/adityaecdrid\">@adityaecdrid</a> has shown.</p>",
          "rawMarkdown": "Yes I will add the validation loop soon. I had it originally but got some weird bug. But clearly it can work, as @adityaecdrid has shown."
        }
      ]
    },
    {
      "id": 837925,
      "postDate": "2020-05-08T06:44:02.123Z",
      "content": "<p>thanks for setting TPU up fo PyTorch, Great Kernel</p>",
      "rawMarkdown": "thanks for setting TPU up fo PyTorch, Great Kernel"
    },
    {
      "id": 806686,
      "postDate": "2020-04-14T00:34:55.213Z",
      "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p>Great! I'm also struggling with memory issue...</p>\n\n<blockquote>\n  <p>TPU training is very senstitive to the learning rate\n  Interesting, do you think it's TPU property rather than xlm-roberta-large's problem?</p>\n</blockquote>",
      "rawMarkdown": "@tanlikesmath \n\nGreat! I'm also struggling with memory issue...\n\n&gt;TPU training is very senstitive to the learning rate\nInteresting, do you think it's TPU property rather than xlm-roberta-large's problem?\n"
    },
    {
      "id": 806265,
      "postDate": "2020-04-13T15:00:22.693Z",
      "content": "<p>You won't believe it but I was stuck in the same problem and almost called it a day. But then I found this post. Thanks very much :) </p>",
      "rawMarkdown": "You won't believe it but I was stuck in the same problem and almost called it a day. But then I found this post. Thanks very much :) ",
      "replies": [
        {
          "id": 806277,
          "postDate": "2020-04-13T15:10:15.373Z",
          "content": "<p>Don't loose hope so early :)</p>",
          "rawMarkdown": "Don't loose hope so early :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 806185,
      "postDate": "2020-04-13T13:50:55.967Z",
      "content": "<ul>\n<li>Here's the end-to-end <a href=\"https://www.kaggle.com/adityaecdrid/working-xlmr-base-8-cores-tpu-pytorch\">xlmr-base-pytorch-8-cores</a>! </li>\n<li>For xlmr-large-pytorch-8-cores, <a href=\"/tanlikesmath\">@tanlikesmath</a> has already shared end-to-end kernels!\nThanks Again!</li>\n</ul>",
      "rawMarkdown": "- Here's the end-to-end [xlmr-base-pytorch-8-cores](https://www.kaggle.com/adityaecdrid/working-xlmr-base-8-cores-tpu-pytorch)! \n- For xlmr-large-pytorch-8-cores, @tanlikesmath has already shared end-to-end kernels!\nThanks Again!"
    },
    {
      "id": 806053,
      "postDate": "2020-04-13T12:02:04.323Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 806392,
      "postDate": "2020-04-13T17:20:37.010Z",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks."
    }
  ],
  "comments": [
    {
      "id": 806330,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-04-13T16:17:12.920000",
      "content": "<p>Here you go folks, <a href=\"https://www.kaggle.com/adityaecdrid/sample-working-xlmr-large-8-cores-tpu-pytorch\">XLMR_Large-8-Cores</a>; It was on just <strong>80k samples</strong>,;</p>\n\n<p>Stats on interactive -:</p>\n\n<p><code>\nAUC = 0.935995099948079\nAUC = 0.9306042315680167\n</code></p>\n\n<p>PS: I am running for more rows; Would be nice if people play around with LR as i will soon run out of my free TPU hours :(; Also training is sometimes unstable idk;</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 810082,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-04-16T17:56:06.257000",
      "content": "<p>Here you should get all your OOM answers <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-tpu-xlmr-base-with-memmap\">kernel</a>; Just don't put too much pedal on bs 😅</p>\n\n<p>It uses the dataset generated from <a href=\"https://www.kaggle.com/adityaecdrid/memmap-tpu-xlmr-pytorch-pad-on-fly\">here</a>.</p>\n\n<p>In short you can simply use <code>numpy's memmap</code> and we make the dataset which is on disk as if it's on the RAM :) ; (from 10k feet that's what's happening;)</p>\n\n<p>&gt;NB I ran base just because i m out of compute; (large works as well! with no OOM)</p>\n\n<p>Also this can be further improved if you want to go about bucketing though I feel it will be an overhead if we do it;</p>\n\n<p>Happy Kaggling!</p>\n\n<p>cc <a href=\"/mgornergoogle\">@mgornergoogle</a>!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 806257,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-04-13T14:54:04.377000",
      "content": "<p>This is a great kernel thanks. Can you also add evaluation here? I had some serious issues with loss jumping around and evaluation being stuck at 0.6 AUC. Actually your inference kernel also scores very badly... so something is similarly wrong to what I observed.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 806589,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-04-13T21:11:34.200000",
          "content": "<p>Yes I will add the validation loop soon. I had it originally but got some weird bug. But clearly it can work, as <a href=\"/adityaecdrid\">@adityaecdrid</a> has shown.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 837925,
      "author_name": "JayChakalasiya",
      "author_url": "",
      "post_date": "2020-05-08T06:44:02.123000",
      "content": "<p>thanks for setting TPU up fo PyTorch, Great Kernel</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 806686,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-04-14T00:34:55.213000",
      "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p>Great! I'm also struggling with memory issue...</p>\n\n<blockquote>\n  <p>TPU training is very senstitive to the learning rate\n  Interesting, do you think it's TPU property rather than xlm-roberta-large's problem?</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 806265,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2020-04-13T15:00:22.693000",
      "content": "<p>You won't believe it but I was stuck in the same problem and almost called it a day. But then I found this post. Thanks very much :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 806277,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-04-13T15:10:15.373000",
          "content": "<p>Don't loose hope so early :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 806185,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-04-13T13:50:55.967000",
      "content": "<ul>\n<li>Here's the end-to-end <a href=\"https://www.kaggle.com/adityaecdrid/working-xlmr-base-8-cores-tpu-pytorch\">xlmr-base-pytorch-8-cores</a>! </li>\n<li>For xlmr-large-pytorch-8-cores, <a href=\"/tanlikesmath\">@tanlikesmath</a> has already shared end-to-end kernels!\nThanks Again!</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 806053,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-13T12:02:04.323000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 806392,
      "author_name": " Igor Krasovskiy",
      "author_url": "",
      "post_date": "2020-04-13T17:20:37.010000",
      "content": "<p>Thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "806048": "In @xhlulu's amazing [kernel](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta), an XLM-RoBERTA-large model is trained on the TPU using TensorFlow. I wanted to also train this, but with PyTorch. In @abhishek's great [kernel](https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml-w-validation), he uses PyTorch XLA to train a BERT-multilingual-uncased model on the TPU. However, if I switch out the BERT model for the XLM-RoBERTa model (even base model), I got VM OOM errors. Discussing this with other competitors, we had originally concluded it is impossible to train XLM-RoBERTa in Kaggle TPU Kernels. However, I decided to investigate further, and I learned how to better manage my VM memory and eventually get the kernel working.Here are a few things that I did to master PyTorch XLA training :)\n\n1. First I tokenized/encoded the data ahead of time in a separate kernel and saved it. Then I loaded it into the kernel and used it to create a TensorDataset to pass data to the model. This was already done by @xhlulu's kernel as well. This might be better than on-the-fly tokenization/encoding used in @abhishek's that might be taking up VM memory during training.\n\n2. I always made sure to delete any unused variables  like the DataFrames (after the creation of a PyTorch dataset) and garbage collect.\n\n3. Even these two changes were not enough. I then discussed this with the PyTorch XLA team on GitHub(thanks @dlilbenzi and @jysohn23) and apparently, they _just added changes to reduce host memory utilization_. So I use the nightly version of PyTorch XLA in my kernels. There are other ways to decrease VM memory utilization as well. I will experiment in the near future with these additional methods in order to improve memory utilization and speed of kernel.\n\nHowever there were a couple of things to watch out for. For example, you need to use bfloat16 when training with xlm-roberta-large, float32 leads to host OOM errors. Also, TPU training is _very senstitive_ to the learning rate. I still am struggling to tune the learning rate.\n\n# Here is the kernel: [training link](https://www.kaggle.com/tanlikesmath/xlm-roberta-pytorch-xla-tpu), [inference link](https://www.kaggle.com/tanlikesmath/xlm-roberta-inference-pytorch-tpu-xla-1-core)\n\nThe score is terrible, but I will tune the LR soon. Also, I will add validation code soon.\n\nI have it working on xhlulu's subset of data, but I am trying to tune the LR a little further before releasing. Additionally, I have some additional memory optimizations in mind. Also, I plan to post a PyTorch XLA tips thread with many tips I have come across or used when training on the TPU.\n\nI hope this helps! As I just mentioned, there were some additional things I wanted to add, but I wanted to share with the community first. Let me know if you have any questions.\n",
    "806330": "Here you go folks, [XLMR_Large-8-Cores](https://www.kaggle.com/adityaecdrid/sample-working-xlmr-large-8-cores-tpu-pytorch); It was on just **80k samples**,;\n\nStats on interactive -:\n\n```\nAUC = 0.935995099948079\nAUC = 0.9306042315680167\n```\n\nPS: I am running for more rows; Would be nice if people play around with LR as i will soon run out of my free TPU hours :(; Also training is sometimes unstable idk;\n\n@philippsinger ",
    "810082": "Here you should get all your OOM answers [kernel](https://www.kaggle.com/adityaecdrid/pytorch-tpu-xlmr-base-with-memmap); Just don't put too much pedal on bs 😅\n\nIt uses the dataset generated from [here](https://www.kaggle.com/adityaecdrid/memmap-tpu-xlmr-pytorch-pad-on-fly).\n\nIn short you can simply use `numpy's memmap` and we make the dataset which is on disk as if it's on the RAM :) ; (from 10k feet that's what's happening;)\n\n&gt;NB I ran base just because i m out of compute; (large works as well! with no OOM)\n\nAlso this can be further improved if you want to go about bucketing though I feel it will be an overhead if we do it;\n\nHappy Kaggling!\n\ncc @mgornergoogle!",
    "806257": "This is a great kernel thanks. Can you also add evaluation here? I had some serious issues with loss jumping around and evaluation being stuck at 0.6 AUC. Actually your inference kernel also scores very badly... so something is similarly wrong to what I observed.",
    "837925": "thanks for setting TPU up fo PyTorch, Great Kernel",
    "806686": "@tanlikesmath \n\nGreat! I'm also struggling with memory issue...\n\n&gt;TPU training is very senstitive to the learning rate\nInteresting, do you think it's TPU property rather than xlm-roberta-large's problem?\n",
    "806265": "You won't believe it but I was stuck in the same problem and almost called it a day. But then I found this post. Thanks very much :) ",
    "806185": "- Here's the end-to-end [xlmr-base-pytorch-8-cores](https://www.kaggle.com/adityaecdrid/working-xlmr-base-8-cores-tpu-pytorch)! \n- For xlmr-large-pytorch-8-cores, @tanlikesmath has already shared end-to-end kernels!\nThanks Again!",
    "806053": "",
    "806392": "Thanks."
  }
}