{
  "id": 159723,
  "title": "Pytorch TPU improvements",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159723",
  "author_name": "",
  "post_date": "2020-06-18T14:02:43.052908400Z",
  "votes": 89,
  "comment_count": 26,
  "views": 0,
  "content": "<p>One of the goals for me in this competition was to play with TPUs and try to get a stable Pytorch kernel running as it is my framework of choice. I recently was in <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-639485009\">contact with Pytorch devs</a> and they are making continuous improvements regarding XLA integration and are very helpful in addressing user feedback, big thanks to them!</p>\n\n<p>I just wanted to quickly share two things that they implemented after receiving feedback that both help me a lot in my kernels, specifically with respect to memory issues. I will publish a Pytorch TPU kernel after competition ends.</p>\n\n<h3>Wrapping the model</h3>\n\n<p>The devs added a new <a href=\"https://github.com/pytorch/xla/blob/f3353240a303c079f73466acc95cb8a487577a55/torch_xla/distributed/xla_multiprocessing.py#L303\">API</a> that makes model sharing between TPU cores more seamingless. All you need to do is to call the following wrapper outside the training routine:</p>\n\n<p><code>model = xmp.MpModelWrapper(MyModel())</code></p>\n\n<h3>Model saving</h3>\n\n<p>The second big issue I identified is that model saving is one of the main culprits with memory issues using large models as the whole state dict is loaded to CPU memory before dumping the model. Again, a solution was <a href=\"https://github.com/pytorch/xla/pull/2140\">developed</a>:</p>\n\n<p><code>\nimport torch_xla.utils.serialization as xser\nxser.save(model.state_dict(), f\"model.bin\", master_only=True)\nmodel.load_state_dict(xser.load(f\"model.bin\"))\n</code></p>\n\n<h3>Further tips &amp; tricks</h3>\n\n<p>If you are doing data subsampling on each TPU core separately, but always load the full data to the cores, it can quickly lead to memory issues. It is better to manually sample the data beforehand, and only feed the samples to the training functions. Another thing is to delete unused objects as often as possible and run <code>gc.collect()</code> frequently.</p>\n\n<p>EDIT\nKernel can be found here: <a href=\"https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589\">https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589</a></p>",
  "messages": [
    {
      "id": "891885",
      "postDate": "06/18/2020 14:02:43",
      "content": "<p>One of the goals for me in this competition was to play with TPUs and try to get a stable Pytorch kernel running as it is my framework of choice. I recently was in <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-639485009\">contact with Pytorch devs</a> and they are making continuous improvements regarding XLA integration and are very helpful in addressing user feedback, big thanks to them!</p>\n\n<p>I just wanted to quickly share two things that they implemented after receiving feedback that both help me a lot in my kernels, specifically with respect to memory issues. I will publish a Pytorch TPU kernel after competition ends.</p>\n\n<h3>Wrapping the model</h3>\n\n<p>The devs added a new <a href=\"https://github.com/pytorch/xla/blob/f3353240a303c079f73466acc95cb8a487577a55/torch_xla/distributed/xla_multiprocessing.py#L303\">API</a> that makes model sharing between TPU cores more seamingless. All you need to do is to call the following wrapper outside the training routine:</p>\n\n<p><code>model = xmp.MpModelWrapper(MyModel())</code></p>\n\n<h3>Model saving</h3>\n\n<p>The second big issue I identified is that model saving is one of the main culprits with memory issues using large models as the whole state dict is loaded to CPU memory before dumping the model. Again, a solution was <a href=\"https://github.com/pytorch/xla/pull/2140\">developed</a>:</p>\n\n<p><code>\nimport torch_xla.utils.serialization as xser\nxser.save(model.state_dict(), f\"model.bin\", master_only=True)\nmodel.load_state_dict(xser.load(f\"model.bin\"))\n</code></p>\n\n<h3>Further tips &amp; tricks</h3>\n\n<p>If you are doing data subsampling on each TPU core separately, but always load the full data to the cores, it can quickly lead to memory issues. It is better to manually sample the data beforehand, and only feed the samples to the training functions. Another thing is to delete unused objects as often as possible and run <code>gc.collect()</code> frequently.</p>\n\n<p>EDIT\nKernel can be found here: <a href=\"https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589\">https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589</a></p>",
      "rawMarkdown": "One of the goals for me in this competition was to play with TPUs and try to get a stable Pytorch kernel running as it is my framework of choice. I recently was in [contact with Pytorch devs](https://github.com/pytorch/xla/issues/1870#issuecomment-639485009) and they are making continuous improvements regarding XLA integration and are very helpful in addressing user feedback, big thanks to them!\n\nI just wanted to quickly share two things that they implemented after receiving feedback that both help me a lot in my kernels, specifically with respect to memory issues. I will publish a Pytorch TPU kernel after competition ends.\n\n### Wrapping the model\n\nThe devs added a new [API](https://github.com/pytorch/xla/blob/f3353240a303c079f73466acc95cb8a487577a55/torch_xla/distributed/xla_multiprocessing.py#L303) that makes model sharing between TPU cores more seamingless. All you need to do is to call the following wrapper outside the training routine:\n\n`model = xmp.MpModelWrapper(MyModel())`\n\n### Model saving\n\nThe second big issue I identified is that model saving is one of the main culprits with memory issues using large models as the whole state dict is loaded to CPU memory before dumping the model. Again, a solution was [developed](https://github.com/pytorch/xla/pull/2140):\n\n```\nimport torch_xla.utils.serialization as xser\nxser.save(model.state_dict(), f\"model.bin\", master_only=True)\nmodel.load_state_dict(xser.load(f\"model.bin\"))\n```\n\n### Further tips &amp; tricks\n\nIf you are doing data subsampling on each TPU core separately, but always load the full data to the cores, it can quickly lead to memory issues. It is better to manually sample the data beforehand, and only feed the samples to the training functions. Another thing is to delete unused objects as often as possible and run `gc.collect()` frequently.\n\nEDIT\nKernel can be found here: https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589",
      "votes": null
    },
    {
      "id": "891925",
      "postDate": "06/18/2020 14:37:59",
      "content": "<p>looking forward to your kernel!</p>",
      "rawMarkdown": "looking forward to your kernel!",
      "votes": null
    },
    {
      "id": "891968",
      "postDate": "06/18/2020 15:05:51",
      "content": "<p>i had a chat with that xla team too,they said by September hopefully they will be able to increase the performance of pytorch tpu close to or equal to tensorflow tpu performance,so yeah fingers crossed,, looking forward to your kernel maybe after 5 more days we will have that :)</p>",
      "rawMarkdown": "i had a chat with that xla team too,they said by September hopefully they will be able to increase the performance of pytorch tpu close to or equal to tensorflow tpu performance,so yeah fingers crossed,, looking forward to your kernel maybe after 5 more days we will have that :)",
      "votes": null
    },
    {
      "id": "891969",
      "postDate": "06/18/2020 15:06:01",
      "content": "<p>Just joined and I'm trying to find a convenient way to merge computation results from different cores. Wish to find great API / tutorial!</p>",
      "rawMarkdown": "Just joined and I'm trying to find a convenient way to merge computation results from different cores. Wish to find great API / tutorial!",
      "votes": null
    },
    {
      "id": "891979",
      "postDate": "06/18/2020 15:15:15",
      "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> You can do something like this:</p>\n\n<p>```\ndef reduce_fn(vals):\n    return sum(vals) / len(vals)</p>\n\n<p>auc = xm.mesh_reduce('auc_reduce', auc, reduce_fn)\nxm.master_print(f'AUC AVG = {auc}')\n```</p>\n\n<p>If you are referring to e.g., scores from different cores. Otherwise you can also return individual results in your <code>_run()</code> function and collect them in the <code>_mp_fn()</code> function.</p>",
      "rawMarkdown": "yaroshevskiy You can do something like this:\n\n```\ndef reduce_fn(vals):\n    return sum(vals) / len(vals)\n\nauc = xm.mesh_reduce('auc_reduce', auc, reduce_fn)\nxm.master_print(f'AUC AVG = {auc}')\n```\n\nIf you are referring to e.g., scores from different cores. Otherwise you can also return individual results in your `_run()` function and collect them in the `_mp_fn()` function.",
      "votes": null
    },
    {
      "id": "892353",
      "postDate": "06/18/2020 20:22:30",
      "content": "<p>Just note that running <code>gc.collect()</code> is a time consuming process. Don't do it in a loop.</p>",
      "rawMarkdown": "Just note that running `gc.collect()` is a time consuming process. Don't do it in a loop.",
      "votes": null
    },
    {
      "id": "892356",
      "postDate": "06/18/2020 20:26:38",
      "content": "<p><a href=\"/abhishek\">@abhishek</a>  also writing too much log information  slows down performance drastically</p>",
      "rawMarkdown": "abhishek  also writing too much log information  slows down performance drastically",
      "votes": null
    },
    {
      "id": "892443",
      "postDate": "06/18/2020 22:01:47",
      "content": "<p>At the end of each epoch is fine.</p>",
      "rawMarkdown": "At the end of each epoch is fine.",
      "votes": null
    },
    {
      "id": "893733",
      "postDate": "06/19/2020 20:50:00",
      "content": "<p>Thanks for sharing these latest updates and for interacting with the team! Another route that I am taking is trying the PytorchLightning library (one example <a href=\"https://colab.research.google.com/drive/1-_LKx4HwAxl5M6xPJmqAAu444LTDQoa3\">here</a>) and its TPU support. Still experimenting so not yet sure how good it is but it is promising so far (in terms of public API). </p>",
      "rawMarkdown": "Thanks for sharing these latest updates and for interacting with the team! Another route that I am taking is trying the PytorchLightning library (one example [here](https://colab.research.google.com/drive/1-_LKx4HwAxl5M6xPJmqAAu444LTDQoa3)) and its TPU support. Still experimenting so not yet sure how good it is but it is promising so far (in terms of public API).",
      "votes": null
    },
    {
      "id": "894109",
      "postDate": "06/20/2020 07:15:50",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    },
    {
      "id": "896358",
      "postDate": "06/22/2020 05:30:08",
      "content": "<p>Hi <a href=\"/philippsinger\">@philippsinger</a> , Can we go for model pruning to optimize the size ?</p>",
      "rawMarkdown": "Hi @philippsinger , Can we go for model pruning to optimize the size ?",
      "votes": null
    },
    {
      "id": "899964",
      "postDate": "06/24/2020 14:51:47",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "906480",
      "postDate": "06/29/2020 10:31:21",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> still eagerly waiting for your kernel.</p>",
      "rawMarkdown": "philippsinger still eagerly waiting for your kernel.",
      "votes": null
    },
    {
      "id": "909483",
      "postDate": "06/30/2020 16:08:46",
      "content": "<p>Yes, I did not forget, sorry for the delay but got stuck with a lot of other stuff. Will try to publish it this week, it probably won't be pretty, but at least contain the things I mentioned above.</p>",
      "rawMarkdown": "Yes, I did not forget, sorry for the delay but got stuck with a lot of other stuff. Will try to publish it this week, it probably won't be pretty, but at least contain the things I mentioned above.",
      "votes": null
    },
    {
      "id": "910332",
      "postDate": "07/01/2020 05:30:03",
      "content": "<p>good</p>",
      "rawMarkdown": "good",
      "votes": null
    },
    {
      "id": "910452",
      "postDate": "07/01/2020 06:45:50",
      "content": "<p>Doesn't have to be pretty, a working one will be more than enough. Thanks for replying and will be waiting. </p>",
      "rawMarkdown": "Doesn't have to be pretty, a working one will be more than enough. Thanks for replying and will be waiting.",
      "votes": null
    },
    {
      "id": "922920",
      "postDate": "07/10/2020 12:18:58",
      "content": "<p>Hello <a href=\"/philippsinger\">@philippsinger</a> \nThanks for sharing, I am just trying to make multi TPU work\nSo model is shared between cores, this is understable\nBut what about optimizer?\nYou also store model only, how do you store optimizer and which one?\nWhen I look at example kernels model is shared but optimizers are created for each core, why?</p>",
      "rawMarkdown": "Hello @philippsinger \nThanks for sharing, I am just trying to make multi TPU work\nSo model is shared between cores, this is understable\nBut what about optimizer?\nYou also store model only, how do you store optimizer and which one?\nWhen I look at example kernels model is shared but optimizers are created for each core, why?",
      "votes": null
    },
    {
      "id": "926226",
      "postDate": "07/12/2020 15:32:07",
      "content": "<p>Sorry took me a bit, but here is the kernel:\n<a href=\"https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589\">https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589</a></p>",
      "rawMarkdown": "Sorry took me a bit, but here is the kernel:\nhttps://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589",
      "votes": null
    },
    {
      "id": "926229",
      "postDate": "07/12/2020 15:33:47",
      "content": "<p>Many thanks <a href=\"/philippsinger\">@philippsinger</a>  👍 😊 </p>",
      "rawMarkdown": "Many thanks @philippsinger  👍 😊",
      "votes": null
    },
    {
      "id": "927025",
      "postDate": "07/13/2020 06:45:07",
      "content": "<p>Thank you very much for your work.👍 </p>",
      "rawMarkdown": "Thank you very much for your work.👍",
      "votes": null
    },
    {
      "id": "957140",
      "postDate": "08/04/2020 04:59:07",
      "content": "<p>Nice helpful tips</p>",
      "rawMarkdown": "Nice helpful tips",
      "votes": null
    },
    {
      "id": "975134",
      "postDate": "08/18/2020 07:03:01",
      "content": "<p>I finally waited for you, and luckily I didn't give up</p>",
      "rawMarkdown": "I finally waited for you, and luckily I didn't give up",
      "votes": null
    },
    {
      "id": "1130999",
      "postDate": "12/29/2020 13:40:59",
      "content": "<p>That's cool!</p>",
      "rawMarkdown": "That's cool!",
      "votes": null
    },
    {
      "id": "1347480",
      "postDate": "06/13/2021 08:37:30",
      "content": "<p>Thanks for sharing these tricks! </p>",
      "rawMarkdown": "Thanks for sharing these tricks!",
      "votes": null
    },
    {
      "id": "1347482",
      "postDate": "06/13/2021 08:39:13",
      "content": "<p>A year has passed, do you know how this claim is going <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a>? Any blog posts or experiments you have done?</p>",
      "rawMarkdown": "A year has passed, do you know how this claim is going @mobassir? Any blog posts or experiments you have done?",
      "votes": null
    },
    {
      "id": "1347498",
      "postDate": "06/13/2021 09:14:52",
      "content": "<p>i used pytorch xla in both recent past cassava and ranzcr competition,<br>\nif you check the changeLog sections of these 2 kernels : </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\" target=\"_blank\">https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9</a></li>\n<li><a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\" target=\"_blank\">https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease</a><br>\nyou will see some of the problems that i faced while using pytorch xla recently.<br>\nin google colab,pytorch xla is exactly same with v2.8 tpu access as it was before and it is painfully slow<br>\nin kaggle if you use latest version of pytorch xla then you will be able to train models like i did in the attached 2 kernels above.<br>\nit looks good now but OOM issue still exist and still pytorch  xla memory is very small compared to tf tpu</li>\n</ol>\n<p>to me it seems like using pytorch xla you will be able to do experiments similar or better than a rtx 2080ti gpu(within 9 hour kaggle tpu limit)<br>\nyou smartly need to use pytorch xla in your code so that you can avoid OOM as much as you can</p>",
      "rawMarkdown": "i used pytorch xla in both recent past cassava and ranzcr competition,\nif you check the changeLog sections of these 2 kernels : \n1. https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\n2. https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\nyou will see some of the problems that i faced while using pytorch xla recently.\nin google colab,pytorch xla is exactly same with v2.8 tpu access as it was before and it is painfully slow\nin kaggle if you use latest version of pytorch xla then you will be able to train models like i did in the attached 2 kernels above.\nit looks good now but OOM issue still exist and still pytorch  xla memory is very small compared to tf tpu\n\nto me it seems like using pytorch xla you will be able to do experiments similar or better than a rtx 2080ti gpu(within 9 hour kaggle tpu limit)\nyou smartly need to use pytorch xla in your code so that you can avoid OOM as much as you can",
      "votes": null
    },
    {
      "id": "1347759",
      "postDate": "06/13/2021 13:31:16",
      "content": "<p>That's so cool!</p>",
      "rawMarkdown": "That's so cool!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1130999,
      "author_name": "danielecomi",
      "author_url": "",
      "post_date": "12/29/2020 13:40:59",
      "content": "<p>That's cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1347480,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/13/2021 08:37:30",
      "content": "<p>Thanks for sharing these tricks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1347759,
      "author_name": "jhanavibehl",
      "author_url": "",
      "post_date": "06/13/2021 13:31:16",
      "content": "<p>That's so cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891925,
      "author_name": "shangweichen",
      "author_url": "",
      "post_date": "06/18/2020 14:37:59",
      "content": "<p>looking forward to your kernel!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891968,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "06/18/2020 15:05:51",
      "content": "<p>i had a chat with that xla team too,they said by September hopefully they will be able to increase the performance of pytorch tpu close to or equal to tensorflow tpu performance,so yeah fingers crossed,, looking forward to your kernel maybe after 5 more days we will have that :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1347482,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/13/2021 08:39:13",
          "content": "<p>A year has passed, do you know how this claim is going <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a>? Any blog posts or experiments you have done?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1347498,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/13/2021 09:14:52",
          "content": "<p>i used pytorch xla in both recent past cassava and ranzcr competition,<br>\nif you check the changeLog sections of these 2 kernels : </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\" target=\"_blank\">https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9</a></li>\n<li><a href=\"https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\" target=\"_blank\">https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease</a><br>\nyou will see some of the problems that i faced while using pytorch xla recently.<br>\nin google colab,pytorch xla is exactly same with v2.8 tpu access as it was before and it is painfully slow<br>\nin kaggle if you use latest version of pytorch xla then you will be able to train models like i did in the attached 2 kernels above.<br>\nit looks good now but OOM issue still exist and still pytorch  xla memory is very small compared to tf tpu</li>\n</ol>\n<p>to me it seems like using pytorch xla you will be able to do experiments similar or better than a rtx 2080ti gpu(within 9 hour kaggle tpu limit)<br>\nyou smartly need to use pytorch xla in your code so that you can avoid OOM as much as you can</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 891969,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "06/18/2020 15:06:01",
      "content": "<p>Just joined and I'm trying to find a convenient way to merge computation results from different cores. Wish to find great API / tutorial!</p>",
      "votes": null,
      "replies": [
        {
          "id": 891979,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/18/2020 15:15:15",
          "content": "<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> You can do something like this:</p>\n\n<p>```\ndef reduce_fn(vals):\n    return sum(vals) / len(vals)</p>\n\n<p>auc = xm.mesh_reduce('auc_reduce', auc, reduce_fn)\nxm.master_print(f'AUC AVG = {auc}')\n```</p>\n\n<p>If you are referring to e.g., scores from different cores. Otherwise you can also return individual results in your <code>_run()</code> function and collect them in the <code>_mp_fn()</code> function.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 892353,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "06/18/2020 20:22:30",
      "content": "<p>Just note that running <code>gc.collect()</code> is a time consuming process. Don't do it in a loop.</p>",
      "votes": null,
      "replies": [
        {
          "id": 892356,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/18/2020 20:26:38",
          "content": "<p><a href=\"/abhishek\">@abhishek</a>  also writing too much log information  slows down performance drastically</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 892443,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/18/2020 22:01:47",
          "content": "<p>At the end of each epoch is fine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893733,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/19/2020 20:50:00",
      "content": "<p>Thanks for sharing these latest updates and for interacting with the team! Another route that I am taking is trying the PytorchLightning library (one example <a href=\"https://colab.research.google.com/drive/1-_LKx4HwAxl5M6xPJmqAAu444LTDQoa3\">here</a>) and its TPU support. Still experimenting so not yet sure how good it is but it is promising so far (in terms of public API). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 894109,
      "author_name": "muralidhar123",
      "author_url": "",
      "post_date": "06/20/2020 07:15:50",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 896358,
      "author_name": "shishu1421",
      "author_url": "",
      "post_date": "06/22/2020 05:30:08",
      "content": "<p>Hi <a href=\"/philippsinger\">@philippsinger</a> , Can we go for model pruning to optimize the size ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 899964,
      "author_name": "denpa92",
      "author_url": "",
      "post_date": "06/24/2020 14:51:47",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 906480,
      "author_name": "rhtsingh",
      "author_url": "",
      "post_date": "06/29/2020 10:31:21",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> still eagerly waiting for your kernel.</p>",
      "votes": null,
      "replies": [
        {
          "id": 909483,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/30/2020 16:08:46",
          "content": "<p>Yes, I did not forget, sorry for the delay but got stuck with a lot of other stuff. Will try to publish it this week, it probably won't be pretty, but at least contain the things I mentioned above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 910452,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "07/01/2020 06:45:50",
          "content": "<p>Doesn't have to be pretty, a working one will be more than enough. Thanks for replying and will be waiting. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 910332,
      "author_name": "bhavithaparvathaneni",
      "author_url": "",
      "post_date": "07/01/2020 05:30:03",
      "content": "<p>good</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 922920,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/10/2020 12:18:58",
      "content": "<p>Hello <a href=\"/philippsinger\">@philippsinger</a> \nThanks for sharing, I am just trying to make multi TPU work\nSo model is shared between cores, this is understable\nBut what about optimizer?\nYou also store model only, how do you store optimizer and which one?\nWhen I look at example kernels model is shared but optimizers are created for each core, why?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 926226,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "07/12/2020 15:32:07",
      "content": "<p>Sorry took me a bit, but here is the kernel:\n<a href=\"https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589\">https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 926229,
          "author_name": "rhtsingh",
          "author_url": "",
          "post_date": "07/12/2020 15:33:47",
          "content": "<p>Many thanks <a href=\"/philippsinger\">@philippsinger</a>  👍 😊 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927025,
          "author_name": "guozhiyu0914",
          "author_url": "",
          "post_date": "07/13/2020 06:45:07",
          "content": "<p>Thank you very much for your work.👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975134,
          "author_name": "shangweichen",
          "author_url": "",
          "post_date": "08/18/2020 07:03:01",
          "content": "<p>I finally waited for you, and luckily I didn't give up</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 957140,
      "author_name": "varunyadav17",
      "author_url": "",
      "post_date": "08/04/2020 04:59:07",
      "content": "<p>Nice helpful tips</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "891885": "One of the goals for me in this competition was to play with TPUs and try to get a stable Pytorch kernel running as it is my framework of choice. I recently was in [contact with Pytorch devs](https://github.com/pytorch/xla/issues/1870#issuecomment-639485009) and they are making continuous improvements regarding XLA integration and are very helpful in addressing user feedback, big thanks to them!\n\nI just wanted to quickly share two things that they implemented after receiving feedback that both help me a lot in my kernels, specifically with respect to memory issues. I will publish a Pytorch TPU kernel after competition ends.\n\n### Wrapping the model\n\nThe devs added a new [API](https://github.com/pytorch/xla/blob/f3353240a303c079f73466acc95cb8a487577a55/torch_xla/distributed/xla_multiprocessing.py#L303) that makes model sharing between TPU cores more seamingless. All you need to do is to call the following wrapper outside the training routine:\n\n`model = xmp.MpModelWrapper(MyModel())`\n\n### Model saving\n\nThe second big issue I identified is that model saving is one of the main culprits with memory issues using large models as the whole state dict is loaded to CPU memory before dumping the model. Again, a solution was [developed](https://github.com/pytorch/xla/pull/2140):\n\n```\nimport torch_xla.utils.serialization as xser\nxser.save(model.state_dict(), f\"model.bin\", master_only=True)\nmodel.load_state_dict(xser.load(f\"model.bin\"))\n```\n\n### Further tips &amp; tricks\n\nIf you are doing data subsampling on each TPU core separately, but always load the full data to the cores, it can quickly lead to memory issues. It is better to manually sample the data beforehand, and only feed the samples to the training functions. Another thing is to delete unused objects as often as possible and run `gc.collect()` frequently.\n\nEDIT\nKernel can be found here: https://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589",
    "891925": "looking forward to your kernel!",
    "891968": "i had a chat with that xla team too,they said by September hopefully they will be able to increase the performance of pytorch tpu close to or equal to tensorflow tpu performance,so yeah fingers crossed,, looking forward to your kernel maybe after 5 more days we will have that :)",
    "891969": "Just joined and I'm trying to find a convenient way to merge computation results from different cores. Wish to find great API / tutorial!",
    "891979": "yaroshevskiy You can do something like this:\n\n```\ndef reduce_fn(vals):\n    return sum(vals) / len(vals)\n\nauc = xm.mesh_reduce('auc_reduce', auc, reduce_fn)\nxm.master_print(f'AUC AVG = {auc}')\n```\n\nIf you are referring to e.g., scores from different cores. Otherwise you can also return individual results in your `_run()` function and collect them in the `_mp_fn()` function.",
    "892353": "Just note that running `gc.collect()` is a time consuming process. Don't do it in a loop.",
    "892356": "abhishek  also writing too much log information  slows down performance drastically",
    "892443": "At the end of each epoch is fine.",
    "893733": "Thanks for sharing these latest updates and for interacting with the team! Another route that I am taking is trying the PytorchLightning library (one example [here](https://colab.research.google.com/drive/1-_LKx4HwAxl5M6xPJmqAAu444LTDQoa3)) and its TPU support. Still experimenting so not yet sure how good it is but it is promising so far (in terms of public API).",
    "894109": "nice",
    "896358": "Hi @philippsinger , Can we go for model pruning to optimize the size ?",
    "899964": "Thanks for sharing!",
    "906480": "philippsinger still eagerly waiting for your kernel.",
    "909483": "Yes, I did not forget, sorry for the delay but got stuck with a lot of other stuff. Will try to publish it this week, it probably won't be pretty, but at least contain the things I mentioned above.",
    "910332": "good",
    "910452": "Doesn't have to be pretty, a working one will be more than enough. Thanks for replying and will be waiting.",
    "922920": "Hello @philippsinger \nThanks for sharing, I am just trying to make multi TPU work\nSo model is shared between cores, this is understable\nBut what about optimizer?\nYou also store model only, how do you store optimizer and which one?\nWhen I look at example kernels model is shared but optimizers are created for each core, why?",
    "926226": "Sorry took me a bit, but here is the kernel:\nhttps://www.kaggle.com/philippsinger/xlm-roberta-large-pytorch-pytorch-tpu?scriptVersionId=38462589",
    "926229": "Many thanks @philippsinger  👍 😊",
    "927025": "Thank you very much for your work.👍",
    "957140": "Nice helpful tips",
    "975134": "I finally waited for you, and luckily I didn't give up",
    "1130999": "That's cool!",
    "1347480": "Thanks for sharing these tricks!",
    "1347482": "A year has passed, do you know how this claim is going @mobassir? Any blog posts or experiments you have done?",
    "1347498": "i used pytorch xla in both recent past cassava and ranzcr competition,\nif you check the changeLog sections of these 2 kernels : \n1. https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\n2. https://www.kaggle.com/mobassir/vit-pytorch-xla-tpu-for-leaf-disease\nyou will see some of the problems that i faced while using pytorch xla recently.\nin google colab,pytorch xla is exactly same with v2.8 tpu access as it was before and it is painfully slow\nin kaggle if you use latest version of pytorch xla then you will be able to train models like i did in the attached 2 kernels above.\nit looks good now but OOM issue still exist and still pytorch  xla memory is very small compared to tf tpu\n\nto me it seems like using pytorch xla you will be able to do experiments similar or better than a rtx 2080ti gpu(within 9 hour kaggle tpu limit)\nyou smartly need to use pytorch xla in your code so that you can avoid OOM as much as you can",
    "1347759": "That's so cool!"
  },
  "source": "meta"
}