{
  "id": 161177,
  "title": "How to run XLM-R large with PyTorch XLA on Kaggle Kernels",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/161177",
  "author_name": "",
  "post_date": "2020-06-24T04:45:59.149324800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I had been working on this competition a while back but could not continue working on it. This past week, I thought about checking it out again. I decided to at least put together some of my learnings regarding PyTorch XLA TPU training. </p>\n\n<p>XLM-R large was quite a big model that is difficult to fit on the host VM for TPU training, but with a few memory optimizations it was possible to run it easily on a Kaggle Kernel. </p>\n\n<p>I wrote a quick tutorial on the basic functionality of PyTorch XLA. I also described what such optimizations are needed. Even though this competition is over, I hope this is helpful for future purposes and you are able to use PyTorch XLA for other applications.</p>\n\n<p><strong>Here is the <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\">kernel</a></strong></p>\n\n<hr>\n\n<p>I tried out Dieter's post-processing trick in the kernel but didn't observe much gain in the public or private LB score. I guess it only works well for some cases.</p>",
  "messages": [
    {
      "id": "899203",
      "postDate": "06/24/2020 04:45:59",
      "content": "<p>I had been working on this competition a while back but could not continue working on it. This past week, I thought about checking it out again. I decided to at least put together some of my learnings regarding PyTorch XLA TPU training. </p>\n\n<p>XLM-R large was quite a big model that is difficult to fit on the host VM for TPU training, but with a few memory optimizations it was possible to run it easily on a Kaggle Kernel. </p>\n\n<p>I wrote a quick tutorial on the basic functionality of PyTorch XLA. I also described what such optimizations are needed. Even though this competition is over, I hope this is helpful for future purposes and you are able to use PyTorch XLA for other applications.</p>\n\n<p><strong>Here is the <a href=\"https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r\">kernel</a></strong></p>\n\n<hr>\n\n<p>I tried out Dieter's post-processing trick in the kernel but didn't observe much gain in the public or private LB score. I guess it only works well for some cases.</p>",
      "rawMarkdown": "I had been working on this competition a while back but could not continue working on it. This past week, I thought about checking it out again. I decided to at least put together some of my learnings regarding PyTorch XLA TPU training. \n\nXLM-R large was quite a big model that is difficult to fit on the host VM for TPU training, but with a few memory optimizations it was possible to run it easily on a Kaggle Kernel. \n\nI wrote a quick tutorial on the basic functionality of PyTorch XLA. I also described what such optimizations are needed. Even though this competition is over, I hope this is helpful for future purposes and you are able to use PyTorch XLA for other applications.\n\n**Here is the [kernel](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r)**\n\n______________\n\nI tried out Dieter's post-processing trick in the kernel but didn't observe much gain in the public or private LB score. I guess it only works well for some cases.",
      "votes": null
    },
    {
      "id": "899558",
      "postDate": "06/24/2020 09:57:49",
      "content": "<p>It is said in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">Dieter's post</a>, \"While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\". So, maybe you only use a single xlm-roberta-large model?</p>",
      "rawMarkdown": "It is said in [Dieter's post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980), \"While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\". So, maybe you only use a single xlm-roberta-large model?",
      "votes": null
    },
    {
      "id": "900290",
      "postDate": "06/24/2020 18:09:42",
      "content": "<p>Ah yes, that's a good point, I missed that statement I guess... Thanks for clarifying!</p>",
      "rawMarkdown": "Ah yes, that's a good point, I missed that statement I guess... Thanks for clarifying!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 899558,
      "author_name": "godelscat",
      "author_url": "",
      "post_date": "06/24/2020 09:57:49",
      "content": "<p>It is said in <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980\">Dieter's post</a>, \"While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\". So, maybe you only use a single xlm-roberta-large model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 900290,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "06/24/2020 18:09:42",
          "content": "<p>Ah yes, that's a good point, I missed that statement I guess... Thanks for clarifying!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "899203": "I had been working on this competition a while back but could not continue working on it. This past week, I thought about checking it out again. I decided to at least put together some of my learnings regarding PyTorch XLA TPU training. \n\nXLM-R large was quite a big model that is difficult to fit on the host VM for TPU training, but with a few memory optimizations it was possible to run it easily on a Kaggle Kernel. \n\nI wrote a quick tutorial on the basic functionality of PyTorch XLA. I also described what such optimizations are needed. Even though this competition is over, I hope this is helpful for future purposes and you are able to use PyTorch XLA for other applications.\n\n**Here is the [kernel](https://www.kaggle.com/tanlikesmath/the-ultimate-pytorch-tpu-tutorial-jigsaw-xlm-r)**\n\n______________\n\nI tried out Dieter's post-processing trick in the kernel but didn't observe much gain in the public or private LB score. I guess it only works well for some cases.",
    "899558": "It is said in [Dieter's post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980), \"While this postprocessing might also help other teams, we think it specifically fixes the divergence of global prediction distributions introduced by having 5 models to predict 6 languages, and might not help much if a team used a single model approach.\". So, maybe you only use a single xlm-roberta-large model?",
    "900290": "Ah yes, that's a good point, I missed that statement I guess... Thanks for clarifying!"
  },
  "source": "meta"
}