{
  "id": 175066,
  "title": "Issue with PyTorch XLA.",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175066",
  "author_name": "",
  "post_date": "2020-08-17T03:18:13.160483400Z",
  "votes": 5,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I am trying to implement PyTorch with XLA for quite a lot of time now, the value of <code>xm.xrt_world_size()</code> comes out to be 1 instead of 8. I think the XLA is not recognizing all the 8 cores. I tried both the tricks, training 8 models in parallel or distributing data equally on 8 cores.</p>\n<p>I studied, Abhishek's, Alex's and CPMP's notebooks regarding implementation of TPU with Pytorch XLA, but to no help!</p>\n<p>If anyone have any pointer on how to implement Pytorch with XLA, please share anything, any help will be appreciated.</p>\n<p>P.S.: I have now created a <a href=\"https://www.kaggle.com/sarques/issue-with-pytorch-xla-melanoma-classification\" target=\"_blank\">notebook</a> to implement the same, if you find any discrepancy in the implementation of TPU, please tell me, I will be eager to do the needed changes! </p>",
  "messages": [
    {
      "id": "973009",
      "postDate": "08/17/2020 03:18:13",
      "content": "<p>I am trying to implement PyTorch with XLA for quite a lot of time now, the value of <code>xm.xrt_world_size()</code> comes out to be 1 instead of 8. I think the XLA is not recognizing all the 8 cores. I tried both the tricks, training 8 models in parallel or distributing data equally on 8 cores.</p>\n<p>I studied, Abhishek's, Alex's and CPMP's notebooks regarding implementation of TPU with Pytorch XLA, but to no help!</p>\n<p>If anyone have any pointer on how to implement Pytorch with XLA, please share anything, any help will be appreciated.</p>\n<p>P.S.: I have now created a <a href=\"https://www.kaggle.com/sarques/issue-with-pytorch-xla-melanoma-classification\" target=\"_blank\">notebook</a> to implement the same, if you find any discrepancy in the implementation of TPU, please tell me, I will be eager to do the needed changes! </p>",
      "rawMarkdown": "I am trying to implement PyTorch with XLA for quite a lot of time now, the value of `xm.xrt_world_size()` comes out to be 1 instead of 8. I think the XLA is not recognizing all the 8 cores. I tried both the tricks, training 8 models in parallel or distributing data equally on 8 cores.\n\nI studied, Abhishek's, Alex's and CPMP's notebooks regarding implementation of TPU with Pytorch XLA, but to no help!\n\nIf anyone have any pointer on how to implement Pytorch with XLA, please share anything, any help will be appreciated.\n\nP.S.: I have now created a [notebook](https://www.kaggle.com/sarques/issue-with-pytorch-xla-melanoma-classification) to implement the same, if you find any discrepancy in the implementation of TPU, please tell me, I will be eager to do the needed changes!",
      "votes": null
    },
    {
      "id": "973443",
      "postDate": "08/17/2020 09:57:07",
      "content": "<p>If you are using <code>xm.xrt_world_size()</code> outside of <code>xmp.spawn()</code> it will show 1 because thats the amount of cores you are using in your notebook but if you are using it inside your <code>xmp.spawn()</code> loop it will show the correct number of cores. PyTorch is still very experimental and has a long way to go till it can catch up to Tensorflow.</p>\n<p>You can also try to check all cores with <code>xm.get_xla_supported_devices()</code></p>",
      "rawMarkdown": "If you are using `xm.xrt_world_size()` outside of `xmp.spawn()` it will show 1 because thats the amount of cores you are using in your notebook but if you are using it inside your `xmp.spawn()` loop it will show the correct number of cores. PyTorch is still very experimental and has a long way to go till it can catch up to Tensorflow.\n\nYou can also try to check all cores with `xm.get_xla_supported_devices()`",
      "votes": null
    },
    {
      "id": "973448",
      "postDate": "08/17/2020 09:59:59",
      "content": "<p>Is that it? I will surely check it out, thanks for the suggestion, actually when I tried running <code>xmp.spawn()</code> with value of <code>nprocs</code> to be 8, it was giving some weird error, so I thought of sticking to value 1.</p>",
      "rawMarkdown": "Is that it? I will surely check it out, thanks for the suggestion, actually when I tried running `xmp.spawn()` with value of `nprocs` to be 8, it was giving some weird error, so I thought of sticking to value 1.",
      "votes": null
    },
    {
      "id": "973450",
      "postDate": "08/17/2020 10:00:47",
      "content": "<p>Post your error if you like, I might be able to help you out.</p>\n<p>If you want to stick to 1 core, you dont need anything else than:</p>\n<p><code>device=xm.xla_device()</code></p>\n<p>and</p>\n<p><code>xm.optimizer_step(optimizer, barrier=True)</code></p>",
      "rawMarkdown": "Post your error if you like, I might be able to help you out.\n\nIf you want to stick to 1 core, you dont need anything else than:\n\n`device=xm.xla_device()`\n\nand\n\n`xm.optimizer_step(optimizer, barrier=True)`",
      "votes": null
    },
    {
      "id": "973452",
      "postDate": "08/17/2020 10:02:15",
      "content": "<p>I will do it later I guess, I will have to generate it again, hope its not a problem for you! as far as I remember, it was something like <code>device cannot be duplicated.</code></p>\n<p>Yeah, 1 core is no problem, implementing 8 cores is the issue for me.</p>",
      "rawMarkdown": "I will do it later I guess, I will have to generate it again, hope its not a problem for you! as far as I remember, it was something like `device cannot be duplicated.`\n\nYeah, 1 core is no problem, implementing 8 cores is the issue for me.",
      "votes": null
    },
    {
      "id": "973931",
      "postDate": "08/17/2020 16:01:16",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I have pasted the error <a href=\"https://paste.centos.org/view/8718ad42\" target=\"_blank\">here</a>!</p>",
      "rawMarkdown": "Hey @aliabdin1 I have pasted the error [here](https://paste.centos.org/view/8718ad42)!",
      "votes": null
    },
    {
      "id": "973983",
      "postDate": "08/17/2020 16:41:15",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> Thats easy to solve. You should not call any functions like <code>xm.xrt_world_size()</code> or <code>xm.xla_device()</code> outside of the <code>_mp_fn()</code> function. Remove these calls put them only in your <code>_mp_fn()</code> function and restart the kernel. It should run without the error now. Its is a very strange behaviour but calling any xm functions before <code>_mp_fn()</code> causes them to happen.</p>",
      "rawMarkdown": "sarques Thats easy to solve. You should not call any functions like `xm.xrt_world_size()` or `xm.xla_device()` outside of the `_mp_fn()` function. Remove these calls put them only in your `_mp_fn()` function and restart the kernel. It should run without the error now. Its is a very strange behaviour but calling any xm functions before `_mp_fn()` causes them to happen.",
      "votes": null
    },
    {
      "id": "973990",
      "postDate": "08/17/2020 16:46:09",
      "content": "<p>Hey, I am trying it just now, should I remove it from the Fitter class too? I am running Fitter from inside <code>_mp_fn()</code> though?</p>",
      "rawMarkdown": "Hey, I am trying it just now, should I remove it from the Fitter class too? I am running Fitter from inside `_mp_fn()` though?",
      "votes": null
    },
    {
      "id": "973995",
      "postDate": "08/17/2020 16:48:22",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> As long as you call it in your <code>_mp_fn()</code> it should work fine.</p>",
      "rawMarkdown": "sarques As long as you call it in your `_mp_fn()` it should work fine.",
      "votes": null
    },
    {
      "id": "973999",
      "postDate": "08/17/2020 16:51:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> Let me be very clear, you are my hero! 🤠 I want to cry now! 😭😭</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2070357%2Fddbe6893354482d49b61113e2850c375%2FScreenshot%20from%202020-08-17%2022-19-37.png?generation=1597683046257433&amp;alt=media\" alt=\"Image\"></p>",
      "rawMarkdown": "Hey @aliabdin1 Let me be very clear, you are my hero! 🤠 I want to cry now! 😭😭\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2070357%2Fddbe6893354482d49b61113e2850c375%2FScreenshot%20from%202020-08-17%2022-19-37.png?generation=1597683046257433&alt=media)",
      "votes": null
    },
    {
      "id": "974007",
      "postDate": "08/17/2020 16:56:06",
      "content": "<p>I was trying this for 2 months, and I can not believe that this was the error. This little thing made fun of me for 2 months, I can't believe this, this is ridiculous. 🤕  </p>",
      "rawMarkdown": "I was trying this for 2 months, and I can not believe that this was the error. This little thing made fun of me for 2 months, I can't believe this, this is ridiculous. 🤕",
      "votes": null
    },
    {
      "id": "974010",
      "postDate": "08/17/2020 16:58:06",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> Im glad to help ;)</p>\n<p>Now you can use functions like <code>xm.master_print()</code> instead of normal <code>print()</code> so only the master prints and you dont have too many confusing prints at once.</p>",
      "rawMarkdown": "sarques Im glad to help ;)\n\nNow you can use functions like `xm.master_print()` instead of normal `print()` so only the master prints and you dont have too many confusing prints at once.",
      "votes": null
    },
    {
      "id": "974021",
      "postDate": "08/17/2020 17:05:48",
      "content": "<p>Yeah, I will strip them now, thanks again! :)</p>",
      "rawMarkdown": "Yeah, I will strip them now, thanks again! :)",
      "votes": null
    },
    {
      "id": "974298",
      "postDate": "08/17/2020 21:18:06",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> congrats. what sort of speed up you have now with xla on tpu vs pure pytorch? </p>",
      "rawMarkdown": "sarques congrats. what sort of speed up you have now with xla on tpu vs pure pytorch?",
      "votes": null
    },
    {
      "id": "974685",
      "postDate": "08/18/2020 02:34:56",
      "content": "<p>Hey, Actually I will have to see it, it was last night when this all happened, so I just switched off the notebook after the successful run.</p>",
      "rawMarkdown": "Hey, Actually I will have to see it, it was last night when this all happened, so I just switched off the notebook after the successful run.",
      "votes": null
    },
    {
      "id": "976162",
      "postDate": "08/18/2020 17:20:11",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I tried running the program now, but I don't understand why it is getting stuck after 3 epochs, is there any reason behind this too? I saw that Abhishek's notebook also got stuck after some epochs.</p>",
      "rawMarkdown": "Hey @aliabdin1 I tried running the program now, but I don't understand why it is getting stuck after 3 epochs, is there any reason behind this too? I saw that Abhishek's notebook also got stuck after some epochs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 973443,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/17/2020 09:57:07",
      "content": "<p>If you are using <code>xm.xrt_world_size()</code> outside of <code>xmp.spawn()</code> it will show 1 because thats the amount of cores you are using in your notebook but if you are using it inside your <code>xmp.spawn()</code> loop it will show the correct number of cores. PyTorch is still very experimental and has a long way to go till it can catch up to Tensorflow.</p>\n<p>You can also try to check all cores with <code>xm.get_xla_supported_devices()</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 973448,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/17/2020 09:59:59",
          "content": "<p>Is that it? I will surely check it out, thanks for the suggestion, actually when I tried running <code>xmp.spawn()</code> with value of <code>nprocs</code> to be 8, it was giving some weird error, so I thought of sticking to value 1.</p>",
          "votes": null,
          "replies": [
            {
              "id": 973450,
              "author_name": "aliabdin1",
              "author_url": "",
              "post_date": "08/17/2020 10:00:47",
              "content": "<p>Post your error if you like, I might be able to help you out.</p>\n<p>If you want to stick to 1 core, you dont need anything else than:</p>\n<p><code>device=xm.xla_device()</code></p>\n<p>and</p>\n<p><code>xm.optimizer_step(optimizer, barrier=True)</code></p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 973452,
              "author_name": "sarques",
              "author_url": "",
              "post_date": "08/17/2020 10:02:15",
              "content": "<p>I will do it later I guess, I will have to generate it again, hope its not a problem for you! as far as I remember, it was something like <code>device cannot be duplicated.</code></p>\n<p>Yeah, 1 core is no problem, implementing 8 cores is the issue for me.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 973931,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/17/2020 16:01:16",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I have pasted the error <a href=\"https://paste.centos.org/view/8718ad42\" target=\"_blank\">here</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973983,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/17/2020 16:41:15",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> Thats easy to solve. You should not call any functions like <code>xm.xrt_world_size()</code> or <code>xm.xla_device()</code> outside of the <code>_mp_fn()</code> function. Remove these calls put them only in your <code>_mp_fn()</code> function and restart the kernel. It should run without the error now. Its is a very strange behaviour but calling any xm functions before <code>_mp_fn()</code> causes them to happen.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973990,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/17/2020 16:46:09",
      "content": "<p>Hey, I am trying it just now, should I remove it from the Fitter class too? I am running Fitter from inside <code>_mp_fn()</code> though?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973995,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/17/2020 16:48:22",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> As long as you call it in your <code>_mp_fn()</code> it should work fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 973999,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/17/2020 16:51:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> Let me be very clear, you are my hero! 🤠 I want to cry now! 😭😭</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2070357%2Fddbe6893354482d49b61113e2850c375%2FScreenshot%20from%202020-08-17%2022-19-37.png?generation=1597683046257433&amp;alt=media\" alt=\"Image\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974007,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/17/2020 16:56:06",
      "content": "<p>I was trying this for 2 months, and I can not believe that this was the error. This little thing made fun of me for 2 months, I can't believe this, this is ridiculous. 🤕  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974010,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/17/2020 16:58:06",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> Im glad to help ;)</p>\n<p>Now you can use functions like <code>xm.master_print()</code> instead of normal <code>print()</code> so only the master prints and you dont have too many confusing prints at once.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974021,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/17/2020 17:05:48",
      "content": "<p>Yeah, I will strip them now, thanks again! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974298,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "08/17/2020 21:18:06",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> congrats. what sort of speed up you have now with xla on tpu vs pure pytorch? </p>",
      "votes": null,
      "replies": [
        {
          "id": 974685,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/18/2020 02:34:56",
          "content": "<p>Hey, Actually I will have to see it, it was last night when this all happened, so I just switched off the notebook after the successful run.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976162,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/18/2020 17:20:11",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I tried running the program now, but I don't understand why it is getting stuck after 3 epochs, is there any reason behind this too? I saw that Abhishek's notebook also got stuck after some epochs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "973009": "I am trying to implement PyTorch with XLA for quite a lot of time now, the value of `xm.xrt_world_size()` comes out to be 1 instead of 8. I think the XLA is not recognizing all the 8 cores. I tried both the tricks, training 8 models in parallel or distributing data equally on 8 cores.\n\nI studied, Abhishek's, Alex's and CPMP's notebooks regarding implementation of TPU with Pytorch XLA, but to no help!\n\nIf anyone have any pointer on how to implement Pytorch with XLA, please share anything, any help will be appreciated.\n\nP.S.: I have now created a [notebook](https://www.kaggle.com/sarques/issue-with-pytorch-xla-melanoma-classification) to implement the same, if you find any discrepancy in the implementation of TPU, please tell me, I will be eager to do the needed changes!",
    "973443": "If you are using `xm.xrt_world_size()` outside of `xmp.spawn()` it will show 1 because thats the amount of cores you are using in your notebook but if you are using it inside your `xmp.spawn()` loop it will show the correct number of cores. PyTorch is still very experimental and has a long way to go till it can catch up to Tensorflow.\n\nYou can also try to check all cores with `xm.get_xla_supported_devices()`",
    "973448": "Is that it? I will surely check it out, thanks for the suggestion, actually when I tried running `xmp.spawn()` with value of `nprocs` to be 8, it was giving some weird error, so I thought of sticking to value 1.",
    "973450": "Post your error if you like, I might be able to help you out.\n\nIf you want to stick to 1 core, you dont need anything else than:\n\n`device=xm.xla_device()`\n\nand\n\n`xm.optimizer_step(optimizer, barrier=True)`",
    "973452": "I will do it later I guess, I will have to generate it again, hope its not a problem for you! as far as I remember, it was something like `device cannot be duplicated.`\n\nYeah, 1 core is no problem, implementing 8 cores is the issue for me.",
    "973931": "Hey @aliabdin1 I have pasted the error [here](https://paste.centos.org/view/8718ad42)!",
    "973983": "sarques Thats easy to solve. You should not call any functions like `xm.xrt_world_size()` or `xm.xla_device()` outside of the `_mp_fn()` function. Remove these calls put them only in your `_mp_fn()` function and restart the kernel. It should run without the error now. Its is a very strange behaviour but calling any xm functions before `_mp_fn()` causes them to happen.",
    "973990": "Hey, I am trying it just now, should I remove it from the Fitter class too? I am running Fitter from inside `_mp_fn()` though?",
    "973995": "sarques As long as you call it in your `_mp_fn()` it should work fine.",
    "973999": "Hey @aliabdin1 Let me be very clear, you are my hero! 🤠 I want to cry now! 😭😭\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2070357%2Fddbe6893354482d49b61113e2850c375%2FScreenshot%20from%202020-08-17%2022-19-37.png?generation=1597683046257433&alt=media)",
    "974007": "I was trying this for 2 months, and I can not believe that this was the error. This little thing made fun of me for 2 months, I can't believe this, this is ridiculous. 🤕",
    "974010": "sarques Im glad to help ;)\n\nNow you can use functions like `xm.master_print()` instead of normal `print()` so only the master prints and you dont have too many confusing prints at once.",
    "974021": "Yeah, I will strip them now, thanks again! :)",
    "974298": "sarques congrats. what sort of speed up you have now with xla on tpu vs pure pytorch?",
    "974685": "Hey, Actually I will have to see it, it was last night when this all happened, so I just switched off the notebook after the successful run.",
    "976162": "Hey @aliabdin1 I tried running the program now, but I don't understand why it is getting stuck after 3 epochs, is there any reason behind this too? I saw that Abhishek's notebook also got stuck after some epochs."
  },
  "source": "meta"
}