{
  "id": 129820,
  "title": "PyTorch-XLA?",
  "url": "/competitions/flower-classification-with-tpus/discussion/129820",
  "author_name": "Aditya Soni",
  "post_date": "2020-02-10T23:04:20.851000",
  "votes": 25,
  "comment_count": 56,
  "views": 0,
  "content": "<h1>Update</h1>\n\n<ul>\n<li>I am not sure if the below is APT For kernels!</li>\n</ul>\n\n<p>==============================================</p>\n\n<p>I tried setting up PyTorch-XLA using this,</p>\n\n<p>```\nimport collections\nfrom datetime import datetime, timedelta\nimport os\nimport tensorflow as tf\nimport numpy as np</p>\n\n<p>_VersionConfig = collections.namedtuple('_VersionConfig', 'wheels,server')\nVERSION = \"torch_xla==nightly\"  #@param [\"torch_xla==nightly\"]\nCONFIG = {\n    'torch_xla==nightly': _VersionConfig('nightly', 'XRT-dev{}'.format(\n        (datetime.today() - timedelta(1)).strftime('%Y%m%d'))),\n}[VERSION]</p>\n\n<p>DIST_BUCKET = 'gs://tpu-pytorch/wheels'\nTORCH_WHEEL = 'torch-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCH_XLA_WHEEL = 'torch_xla-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCHVISION_WHEEL = 'torchvision-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\n! export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH</p>\n\n<h1>Install TPU compat PyTorch/TPU wheels and dependencies</h1>\n\n<p>!pip uninstall -y torch torchvision\n!gsutil cp \"$DIST_BUCKET/$TORCH_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCH_XLA_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCHVISION_WHEEL\" .\n!pip install \"$TORCH_WHEEL\"\n!pip install \"$TORCH_XLA_WHEEL\"\n!pip install \"$TORCHVISION_WHEEL\"\n!apt-get install libomp5 -y\n!apt-get install libopenblas-dev -y</p>\n\n<p>import torch_xla.core.xla_model as xm\nimport torch_xla.distributed.data_parallel as dp # <a href=\"http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading\">http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading</a>\nimport torch_xla.distributed.xla_multiprocessing as xmp # <a href=\"http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing\">http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing</a>\nimport torch_xla.distributed.parallel_loader as pl</p>\n\n<p>print(xm.xrt_world_size()) # don't do this though, just to illustrate it installed things\n```\nIt installs correctly!</p>",
  "messages": [
    {
      "id": 741730,
      "postDate": "2020-02-10T23:04:20.850Z",
      "content": "<h1>Update</h1>\n\n<ul>\n<li>I am not sure if the below is APT For kernels!</li>\n</ul>\n\n<p>==============================================</p>\n\n<p>I tried setting up PyTorch-XLA using this,</p>\n\n<p>```\nimport collections\nfrom datetime import datetime, timedelta\nimport os\nimport tensorflow as tf\nimport numpy as np</p>\n\n<p>_VersionConfig = collections.namedtuple('_VersionConfig', 'wheels,server')\nVERSION = \"torch_xla==nightly\"  #@param [\"torch_xla==nightly\"]\nCONFIG = {\n    'torch_xla==nightly': _VersionConfig('nightly', 'XRT-dev{}'.format(\n        (datetime.today() - timedelta(1)).strftime('%Y%m%d'))),\n}[VERSION]</p>\n\n<p>DIST_BUCKET = 'gs://tpu-pytorch/wheels'\nTORCH_WHEEL = 'torch-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCH_XLA_WHEEL = 'torch_xla-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCHVISION_WHEEL = 'torchvision-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\n! export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH</p>\n\n<h1>Install TPU compat PyTorch/TPU wheels and dependencies</h1>\n\n<p>!pip uninstall -y torch torchvision\n!gsutil cp \"$DIST_BUCKET/$TORCH_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCH_XLA_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCHVISION_WHEEL\" .\n!pip install \"$TORCH_WHEEL\"\n!pip install \"$TORCH_XLA_WHEEL\"\n!pip install \"$TORCHVISION_WHEEL\"\n!apt-get install libomp5 -y\n!apt-get install libopenblas-dev -y</p>\n\n<p>import torch_xla.core.xla_model as xm\nimport torch_xla.distributed.data_parallel as dp # <a href=\"http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading\">http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading</a>\nimport torch_xla.distributed.xla_multiprocessing as xmp # <a href=\"http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing\">http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing</a>\nimport torch_xla.distributed.parallel_loader as pl</p>\n\n<p>print(xm.xrt_world_size()) # don't do this though, just to illustrate it installed things\n```\nIt installs correctly!</p>",
      "rawMarkdown": "# Update\n- I am not sure if the below is APT For kernels!\n\n\n==============================================\n\nI tried setting up PyTorch-XLA using this,\n\n```\nimport collections\nfrom datetime import datetime, timedelta\nimport os\nimport tensorflow as tf\nimport numpy as np\n\n_VersionConfig = collections.namedtuple('_VersionConfig', 'wheels,server')\nVERSION = \"torch_xla==nightly\"  #@param [\"torch_xla==nightly\"]\nCONFIG = {\n    'torch_xla==nightly': _VersionConfig('nightly', 'XRT-dev{}'.format(\n        (datetime.today() - timedelta(1)).strftime('%Y%m%d'))),\n}[VERSION]\n\nDIST_BUCKET = 'gs://tpu-pytorch/wheels'\nTORCH_WHEEL = 'torch-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCH_XLA_WHEEL = 'torch_xla-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCHVISION_WHEEL = 'torchvision-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\n! export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH\n\n# Install TPU compat PyTorch/TPU wheels and dependencies\n!pip uninstall -y torch torchvision\n!gsutil cp \"$DIST_BUCKET/$TORCH_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCH_XLA_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCHVISION_WHEEL\" .\n!pip install \"$TORCH_WHEEL\"\n!pip install \"$TORCH_XLA_WHEEL\"\n!pip install \"$TORCHVISION_WHEEL\"\n!apt-get install libomp5 -y\n!apt-get install libopenblas-dev -y\n\nimport torch_xla.core.xla_model as xm\nimport torch_xla.distributed.data_parallel as dp # http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading\nimport torch_xla.distributed.xla_multiprocessing as xmp # http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing\nimport torch_xla.distributed.parallel_loader as pl\n\nprint(xm.xrt_world_size()) # don't do this though, just to illustrate it installed things\n```\nIt installs correctly!",
      "votes": 25
    },
    {
      "id": 750959,
      "postDate": "2020-02-19T21:48:13.160Z",
      "content": "<p>Dear PyTorch and TPU fans.</p>\n\n<p>At this point, we have been able to confirm that there is additional work to be done on the Kaggle end to make the PyTorch/TPU integration work. It is unfortunately not possible to make it work in user space. Thank you for your hacking attempts though. They gave us good insights into what was need. We are in touch with the PyTorch TPU team at Google and the next step for us is to finish the technical analysis of the situation with them.</p>",
      "rawMarkdown": "Dear PyTorch and TPU fans.\n\nAt this point, we have been able to confirm that there is additional work to be done on the Kaggle end to make the PyTorch/TPU integration work. It is unfortunately not possible to make it work in user space. Thank you for your hacking attempts though. They gave us good insights into what was need. We are in touch with the PyTorch TPU team at Google and the next step for us is to finish the technical analysis of the situation with them.",
      "votes": 20,
      "replies": [
        {
          "id": 750969,
          "postDate": "2020-02-19T21:59:08.300Z",
          "content": "<p>Oh no!!! :D</p>",
          "rawMarkdown": "Oh no!!! :D",
          "votes": 3
        },
        {
          "id": 750981,
          "postDate": "2020-02-19T22:12:09.857Z",
          "content": "<p>This is unfortunate but I am glad you are working with the PyTorch XLA team to solve this issue. Do you have an expected time frame for the integration of PyTorch XLA with the Kaggle TPUs? Do you expect it to happen before the end of this competition?</p>",
          "rawMarkdown": "This is unfortunate but I am glad you are working with the PyTorch XLA team to solve this issue. Do you have an expected time frame for the integration of PyTorch XLA with the Kaggle TPUs? Do you expect it to happen before the end of this competition?"
        },
        {
          "id": 751010,
          "postDate": "2020-02-19T23:20:23.970Z",
          "content": "<p>Also, <a href=\"/mgornergoogle\">@mgornergoogle</a> could you please elaborate what is the current issue that needs to be fixed? Because it was possible for PyTorch XLA to at least recognize the device. Is the error related to the <a href=\"https://github.com/pytorch/xla/issues/1619\">GitHub issue</a>?</p>",
          "rawMarkdown": "Also, @mgornergoogle could you please elaborate what is the current issue that needs to be fixed? Because it was possible for PyTorch XLA to at least recognize the device. Is the error related to the [GitHub issue](https://github.com/pytorch/xla/issues/1619)?"
        },
        {
          "id": 751098,
          "postDate": "2020-02-20T01:54:23.430Z",
          "content": "<p>No, this is independent from the potential TFRecord reader issue. The issue is PyTorch supporting code running on the Cloud TPU itself.</p>",
          "rawMarkdown": "No, this is independent from the potential TFRecord reader issue. The issue is PyTorch supporting code running on the Cloud TPU itself.",
          "votes": 3
        },
        {
          "id": 751261,
          "postDate": "2020-02-20T04:35:56.050Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the clarification. Is it expected for this separate issue to be fixed before the end of the competition?</p>",
          "rawMarkdown": "@mgornergoogle Thanks for the clarification. Is it expected for this separate issue to be fixed before the end of the competition?"
        },
        {
          "id": 751323,
          "postDate": "2020-02-20T05:50:41.477Z",
          "content": "<p>I don't have enough info to answer at this point. We are meeting with the PyTorch TPU team next week to sync our roadmaps.</p>",
          "rawMarkdown": "I don't have enough info to answer at this point. We are meeting with the PyTorch TPU team next week to sync our roadmaps.",
          "votes": 1
        },
        {
          "id": 751345,
          "postDate": "2020-02-20T06:23:59.957Z",
          "content": "<p>Thank you for your response. Please let us know once the time frame is determined.</p>",
          "rawMarkdown": "Thank you for your response. Please let us know once the time frame is determined."
        },
        {
          "id": 757571,
          "postDate": "2020-02-26T23:25:12.647Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's been a week since the last response. Is there an update on the timeframe?</p>",
          "rawMarkdown": "@mgornergoogle It's been a week since the last response. Is there an update on the timeframe?"
        },
        {
          "id": 758577,
          "postDate": "2020-02-28T00:06:56.723Z",
          "content": "<p>Yes, the timeframe will be &gt; 1 month. That's all I can say for now. Still scoping the work.</p>",
          "rawMarkdown": "Yes, the timeframe will be &gt; 1 month. That's all I can say for now. Still scoping the work."
        },
        {
          "id": 758694,
          "postDate": "2020-02-28T03:41:42.850Z",
          "content": "<p>Ok thanks for the update. I hope there will be enough time to participate in this competition with PyTorch.</p>",
          "rawMarkdown": "Ok thanks for the update. I hope there will be enough time to participate in this competition with PyTorch."
        },
        {
          "id": 758695,
          "postDate": "2020-02-28T03:42:38.287Z",
          "content": "<p>Please let us know of any additional updates on the time frame and progress. Greatly appreciate it! </p>",
          "rawMarkdown": "Please let us know of any additional updates on the time frame and progress. Greatly appreciate it! "
        }
      ]
    },
    {
      "id": 749549,
      "postDate": "2020-02-18T19:18:04.450Z",
      "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> you could try adding <code>XRT_TPU_CONFIG</code> environment variable before your code as follows:\n<code>os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470\"</code>.</p>\n\n<p>This should unblock you one step further. However, we cannot guarantee that training will work as we have not yet worked on TPU support in PyTorch.</p>",
      "rawMarkdown": "@adityaecdrid you could try adding `XRT_TPU_CONFIG` environment variable before your code as follows:\n`os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470\"`.\n\nThis should unblock you one step further. However, we cannot guarantee that training will work as we have not yet worked on TPU support in PyTorch.",
      "votes": 3,
      "replies": [
        {
          "id": 749665,
          "postDate": "2020-02-18T20:50:20.120Z",
          "content": "<p>Yes, this works, but PyTorch XLA is still unusable for this competition unless the <a href=\"https://github.com/pytorch/xla/issues/1619\">TFrecords issue</a> is fixed.</p>\n\n<p>I see you are part of the Kaggle Team. Is the Kaggle Team working on integrating PyTorch XLA with the Kaggle TPUs?</p>",
          "rawMarkdown": "Yes, this works, but PyTorch XLA is still unusable for this competition unless the [TFrecords issue](https://github.com/pytorch/xla/issues/1619) is fixed.\n\nI see you are part of the Kaggle Team. Is the Kaggle Team working on integrating PyTorch XLA with the Kaggle TPUs?"
        },
        {
          "id": 749849,
          "postDate": "2020-02-19T00:00:17.113Z",
          "content": "<p>Thanks for getting back kaggle team!</p>",
          "rawMarkdown": "Thanks for getting back kaggle team!"
        },
        {
          "id": 749858,
          "postDate": "2020-02-19T00:08:05.153Z",
          "content": "<p>I responded in the <a href=\"https://github.com/pytorch/xla/issues/1619\">TFRecords issue</a>. I do not see anything wrong with the data returned. The only thing that seems wrong is that you are trying to access an \"id\" field that does not exist in the TFRecords and instead of getting an error, you are getting bogus data. Not sure why. There is a workaround: do not try to access fields that do not exist in the dataset.</p>",
          "rawMarkdown": "I responded in the [TFRecords issue](https://github.com/pytorch/xla/issues/1619). I do not see anything wrong with the data returned. The only thing that seems wrong is that you are trying to access an \"id\" field that does not exist in the TFRecords and instead of getting an error, you are getting bogus data. Not sure why. There is a workaround: do not try to access fields that do not exist in the dataset."
        },
        {
          "id": 749866,
          "postDate": "2020-02-19T00:28:06.643Z",
          "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> can you share the Colab code sample you are referring to? To my knowledge there shouldn't be anything relevant on that port. </p>",
          "rawMarkdown": "@adityaecdrid can you share the Colab code sample you are referring to? To my knowledge there shouldn't be anything relevant on that port. "
        },
        {
          "id": 749878,
          "postDate": "2020-02-19T00:51:03.470Z",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a>  Hey <a href=\"https://gist.github.com/AdityaSoni19031997/f877ebb73dd1b10c1758505eac08abae\">here</a> it's (gist link), me not sure whether i am heading towards correct direction though 😅  (because on GCP Vm's it isn't there in the shell scripts they use to set up everything for you);</p>",
          "rawMarkdown": "@ifigotin  Hey [here](https://gist.github.com/AdityaSoni19031997/f877ebb73dd1b10c1758505eac08abae) it's (gist link), me not sure whether i am heading towards correct direction though 😅  (because on GCP Vm's it isn't there in the shell scripts they use to set up everything for you);",
          "votes": 2
        }
      ]
    },
    {
      "id": 741740,
      "postDate": "2020-02-10T23:16:32.713Z",
      "content": "<p>PyTorch is not yet officially supported on Kaggle TPUs. This launch focuses on tpus + Keras and Tensorflow 2.1. We did not test with PyTorch.\nHowever, if you manage to make it work with PyTorch, please let us know.</p>",
      "rawMarkdown": "PyTorch is not yet officially supported on Kaggle TPUs. This launch focuses on tpus + Keras and Tensorflow 2.1. We did not test with PyTorch.\nHowever, if you manage to make it work with PyTorch, please let us know.",
      "votes": 3,
      "replies": [
        {
          "id": 741748,
          "postDate": "2020-02-10T23:29:16.767Z",
          "content": "<p>Cool Martin Will Update! </p>\n\n<ul>\n<li>It seems it's gonna be a bump ride because datasets is also being saved/read as per tf format? (as you showed in your Great Starter kernel!)</li>\n</ul>\n\n<p>Thanks for your time!</p>\n\n<p>When i executed,</p>\n\n<p><code>\nimport torch_xla.core.xla_model as xm\nxm.xrt_world_size() ## gives 1\ndevices = xm.get_xla_supported_devices() ## This never completes execution and nbs freezes i guess\n</code></p>\n\n<p>Quick Experiments, So I can be wrong!</p>",
          "rawMarkdown": "Cool Martin Will Update! \n\n- It seems it's gonna be a bump ride because datasets is also being saved/read as per tf format? (as you showed in your Great Starter kernel!)\n\nThanks for your time!\n\nWhen i executed,\n\n```\nimport torch_xla.core.xla_model as xm\nxm.xrt_world_size() ## gives 1\ndevices = xm.get_xla_supported_devices() ## This never completes execution and nbs freezes i guess\n```\n\nQuick Experiments, So I can be wrong!"
        },
        {
          "id": 741871,
          "postDate": "2020-02-11T01:23:43.223Z",
          "content": "<p>I have the same issue ☹️ \nCan't execute <code>torch_xla._XLAC._xla_get_devices()</code></p>",
          "rawMarkdown": "I have the same issue ☹️ \nCan't execute `torch_xla._XLAC._xla_get_devices()`",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 741882,
          "postDate": "2020-02-11T01:40:16.650Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 741884,
          "postDate": "2020-02-11T01:41:19.363Z",
          "content": "<p>Datasets are only read. and yes, the dataset for this competition is provided in TFRecord format for convenience. I would not expect the Dataset format to be an issue for PyTorch. Is it ?</p>",
          "rawMarkdown": "Datasets are only read. and yes, the dataset for this competition is provided in TFRecord format for convenience. I would not expect the Dataset format to be an issue for PyTorch. Is it ?"
        },
        {
          "id": 741885,
          "postDate": "2020-02-11T01:42:54.453Z",
          "content": "<p>PyTorch XLA has a TfRecord Reader over <a href=\"https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py\">here</a></p>",
          "rawMarkdown": "PyTorch XLA has a TfRecord Reader over [here](https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py)",
          "votes": 3
        },
        {
          "id": 741888,
          "postDate": "2020-02-11T01:47:42.343Z",
          "content": "<p>Preliminary tests seem to indicate it works fine:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F674553%2F17cd2302ee6ca4b61f2b8a94a54401e9%2Fpytorch_xla_tfrecord.PNG?generation=1581385710518353&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Preliminary tests seem to indicate it works fine:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F674553%2F17cd2302ee6ca4b61f2b8a94a54401e9%2Fpytorch_xla_tfrecord.PNG?generation=1581385710518353&amp;alt=media)\n\n",
          "votes": 1
        },
        {
          "id": 741945,
          "postDate": "2020-02-11T02:33:19.433Z",
          "content": "<p>Great!</p>",
          "rawMarkdown": "Great!"
        },
        {
          "id": 741959,
          "postDate": "2020-02-11T02:40:46.527Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Actually, upon further inspection this does not look correct. The image tensor is not divisible by 512. Also, there are multiple id's for a single image it seems. What do you think? </p>",
          "rawMarkdown": "@mgornergoogle Actually, upon further inspection this does not look correct. The image tensor is not divisible by 512. Also, there are multiple id's for a single image it seems. What do you think? "
        },
        {
          "id": 743028,
          "postDate": "2020-02-11T18:18:11.413Z",
          "content": "<p>The images in the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers\">five flowers</a> or the the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/data\">104 flowers</a> datasets are JPEG encoded. You need to decode them first. Then you will get 512x512 pixel images.</p>\n\n<p>The multiple id's for a single image are not normal. Maybe you are looking at a batch and not a single data element?</p>",
          "rawMarkdown": "The images in the [five flowers](https://www.kaggle.com/mgornergoogle/five-flowers) or the the [104 flowers](https://www.kaggle.com/c/flower-classification-with-tpus/data) datasets are JPEG encoded. You need to decode them first. Then you will get 512x512 pixel images.\n\nThe multiple id's for a single image are not normal. Maybe you are looking at a batch and not a single data element?",
          "votes": 1
        },
        {
          "id": 743189,
          "postDate": "2020-02-11T21:48:28.577Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> \nDo I have to decode directly from the bytes? Or do I decode from the image tensor? PyTorch doesn't seem to have a built-in function for this as opposed to TensorFlow.</p>\n\n<p>Regarding the multiple ids, technically this function is supposed to read only a single example. If I was looking at a batch, shouldn't I be seeing multiple classes though? Maybe there is a bug in the way the example is formatted? This might be expected as the function does not let me define the expected format of the TFRecord like TF does.</p>\n\n<p>If I have the raw bytes, is there a way to convert this into a dictionary? Because Pytorch XLA also has a <a href=\"https://github.com/pytorch/xla/blob/e68039df347d0023a71800e0ebfea5c08168e5f6/torch_xla/utils/tf_record_reader.py#L33\">function</a> that will read the entire dataset to bytes, and I wonder if I can convert that to a proper PyTorch dataset? </p>",
          "rawMarkdown": "@mgornergoogle \nDo I have to decode directly from the bytes? Or do I decode from the image tensor? PyTorch doesn't seem to have a built-in function for this as opposed to TensorFlow.\n\nRegarding the multiple ids, technically this function is supposed to read only a single example. If I was looking at a batch, shouldn't I be seeing multiple classes though? Maybe there is a bug in the way the example is formatted? This might be expected as the function does not let me define the expected format of the TFRecord like TF does.\n\nIf I have the raw bytes, is there a way to convert this into a dictionary? Because Pytorch XLA also has a [function](https://github.com/pytorch/xla/blob/e68039df347d0023a71800e0ebfea5c08168e5f6/torch_xla/utils/tf_record_reader.py#L33) that will read the entire dataset to bytes, and I wonder if I can convert that to a proper PyTorch dataset? ",
          "votes": 1
        },
        {
          "id": 743201,
          "postDate": "2020-02-11T22:00:13.110Z",
          "content": "<p>I'm afraid my experience with Pytorch is minimal... I can try to help on TFRecords though.\nIn a TFRecord (or \"example\") as it is sometimes called in the documentation, you can have only three types of \"things\":\n- a list of bytestrings\n- a list of int64\n- a list of float32\nIn the 100 flowers dataset, the format of each TFRecord of labeled data is:\n\"image\": list of bytestrings containing 1 bytestring (the JPEG-ecoded image bytes)\n\"label\": list of int64 containing 1 int64</p>\n\n<p>I hope this helps some.</p>",
          "rawMarkdown": "I'm afraid my experience with Pytorch is minimal... I can try to help on TFRecords though.\nIn a TFRecord (or \"example\") as it is sometimes called in the documentation, you can have only three types of \"things\":\n- a list of bytestrings\n- a list of int64\n- a list of float32\nIn the 100 flowers dataset, the format of each TFRecord of labeled data is:\n\"image\": list of bytestrings containing 1 bytestring (the JPEG-ecoded image bytes)\n\"label\": list of int64 containing 1 int64\n\nI hope this helps some.",
          "votes": 3
        },
        {
          "id": 744285,
          "postDate": "2020-02-12T17:42:05.937Z",
          "content": "<p>Here's how one can atleast decode those images back <a href=\"https://gist.github.com/dlibenzi/0075e27fca67ce31f7a6d701d77de48a\">gist</a></p>\n\n<p>PS I am not the author but verified it works like a charm but speed is somewhat we need to worry about; Next step might be to wrap it around a PyTorch DataSet (Iterator) (maybe?)</p>",
          "rawMarkdown": "Here's how one can atleast decode those images back [gist](https://gist.github.com/dlibenzi/0075e27fca67ce31f7a6d701d77de48a)\n\nPS I am not the author but verified it works like a charm but speed is somewhat we need to worry about; Next step might be to wrap it around a PyTorch DataSet (Iterator) (maybe?)",
          "votes": 1
        },
        {
          "id": 744440,
          "postDate": "2020-02-12T20:41:15.977Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> As Aditya said, there is a way to decode the images in PyTorch. However, the format is still wrong. Please see <a href=\"https://github.com/pytorch/xla/issues/1619\">this GitHub issue</a> for more information.</p>",
          "rawMarkdown": "@mgornergoogle As Aditya said, there is a way to decode the images in PyTorch. However, the format is still wrong. Please see [this GitHub issue](https://github.com/pytorch/xla/issues/1619) for more information.",
          "votes": 1
        },
        {
          "id": 746240,
          "postDate": "2020-02-14T19:20:48.403Z",
          "content": "<p>I had a look at the GitHub issue and rewsponded there:</p>\n\n<p>In the decoded TFRecords posted by @tmabraham, for training data, \"class\" and \"image\" seem to be correct, \"id\" is not a label that exists in this dataset so I do not know what is being returned there. I would venture that it's a case of bad error handling in the API. </p>",
          "rawMarkdown": "I had a look at the GitHub issue and rewsponded there:\n\nIn the decoded TFRecords posted by @tmabraham, for training data, \"class\" and \"image\" seem to be correct, \"id\" is not a label that exists in this dataset so I do not know what is being returned there. I would venture that it's a case of bad error handling in the API. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 743485,
      "postDate": "2020-02-12T05:27:03.667Z",
      "content": "<p>I'm struggling to install Pytorch-XLA <a href=\"/adityaecdrid\">@adityaecdrid</a>, the code snippet you have given runs properly, but when I try to import torch or torch-xla, it gives some weird errors. Are you able to import torch-xla?</p>",
      "rawMarkdown": "I'm struggling to install Pytorch-XLA @adityaecdrid, the code snippet you have given runs properly, but when I try to import torch or torch-xla, it gives some weird errors. Are you able to import torch-xla?",
      "votes": 1,
      "replies": [
        {
          "id": 743534,
          "postDate": "2020-02-12T06:12:55.400Z",
          "content": "<p><a href=\"/tarunpaparaju\">@tarunpaparaju</a>  Well i have released it as a <a href=\"https://www.kaggle.com/adityaecdrid/quest-to-use-use-pytorch-xla\">nbs</a>, check it once...</p>",
          "rawMarkdown": "@tarunpaparaju  Well i have released it as a [nbs](https://www.kaggle.com/adityaecdrid/quest-to-use-use-pytorch-xla), check it once...",
          "votes": 2
        },
        {
          "id": 743607,
          "postDate": "2020-02-12T06:20:59.487Z",
          "content": "<p>Thanks a lot <a href=\"/adityaecdrid\">@adityaecdrid</a> 😄 </p>",
          "rawMarkdown": "Thanks a lot @adityaecdrid 😄 "
        },
        {
          "id": 744147,
          "postDate": "2020-02-12T16:02:31.600Z",
          "content": "<p>I was able to import it with your notebook, but when I execute <code>dev = xm.xla_device()</code> it runs indefinitely.</p>",
          "rawMarkdown": "I was able to import it with your notebook, but when I execute `dev = xm.xla_device()` it runs indefinitely."
        },
        {
          "id": 744198,
          "postDate": "2020-02-12T16:42:27.740Z",
          "content": "<p>Yep precisely the issue i also faced hence dropped my experiments for now; Something is kinda wrong here because TF says it has 8 but PyTorch only detects one..\nNB Take my words with a pinch of salt, i am a beginner wrt TPU's and XLA's (and can be completely wrong)</p>",
          "rawMarkdown": "Yep precisely the issue i also faced hence dropped my experiments for now; Something is kinda wrong here because TF says it has 8 but PyTorch only detects one..\nNB Take my words with a pinch of salt, i am a beginner wrt TPU's and XLA's (and can be completely wrong)"
        },
        {
          "id": 938602,
          "postDate": "2020-07-21T16:15:40.870Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I am trying to implement multiprocessing on TPU with PyTorch and I am getting this weird error saying it can't replicate devices to 8 because there is only one. Also, when I tried <code>xm.xrt_world_size()</code> I got 1. So, it seems I have the same problem as you, Can you tell me any more about  this and if possible, how to deal with it? Any help will be greatly appreciated! :)</p>",
          "rawMarkdown": "Hey @adityaecdrid I am trying to implement multiprocessing on TPU with PyTorch and I am getting this weird error saying it can't replicate devices to 8 because there is only one. Also, when I tried `xm.xrt_world_size()` I got 1. So, it seems I have the same problem as you, Can you tell me any more about  this and if possible, how to deal with it? Any help will be greatly appreciated! :)"
        },
        {
          "id": 941304,
          "postDate": "2020-07-23T06:56:13.443Z",
          "content": "<p>Don't do xm.xrt_world_size() in a global scope; You should refer the examples from the xla repo; They are very good IMO.</p>\n\n<p><a href=\"https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/multi-core-alexnet-fashion-mnist.ipynb\">https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/multi-core-alexnet-fashion-mnist.ipynb</a></p>",
          "rawMarkdown": "Don't do xm.xrt_world_size() in a global scope; You should refer the examples from the xla repo; They are very good IMO.\n\nhttps://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/multi-core-alexnet-fashion-mnist.ipynb"
        },
        {
          "id": 941389,
          "postDate": "2020-07-23T07:44:50.557Z",
          "content": "<p>Sure, will see it, also, someone told me that there is a problem with Pytorch XLA on Kaggle due to some recwnt update and I should rather search for another option, like pytorch lightening or tensorflow, is it true?</p>",
          "rawMarkdown": "Sure, will see it, also, someone told me that there is a problem with Pytorch XLA on Kaggle due to some recwnt update and I should rather search for another option, like pytorch lightening or tensorflow, is it true?"
        },
        {
          "id": 941594,
          "postDate": "2020-07-23T10:10:39.527Z",
          "content": "<p>I am not quite sure about that claim that XLA doesn't work, maybe you can try forking some xla kernel from Abhishek's kernels and just run it to confirm.. Yesterday, i tested it on Colab, XLA worked great! fine!</p>",
          "rawMarkdown": "I am not quite sure about that claim that XLA doesn't work, maybe you can try forking some xla kernel from Abhishek's kernels and just run it to confirm.. Yesterday, i tested it on Colab, XLA worked great! fine!",
          "votes": -1
        },
        {
          "id": 941620,
          "postDate": "2020-07-23T10:26:54.900Z",
          "content": "<p>Cool, thanks, will do it for sure! :)</p>",
          "rawMarkdown": "Cool, thanks, will do it for sure! :)"
        }
      ]
    },
    {
      "id": 744243,
      "postDate": "2020-02-12T17:17:28.643Z",
      "content": "<p>Sp guys the above won't work as we think because that's how you do it for Colab nbs, but this is a pod instance we are getting access to, so Kaggle team itself has to pull a docker image or something like that needs to be done as per the instructions in the PyTorch XLA Official Github..</p>\n\n<p>Please Avoid the installation steps. Sorry for the mess i created;</p>",
      "rawMarkdown": "Sp guys the above won't work as we think because that's how you do it for Colab nbs, but this is a pod instance we are getting access to, so Kaggle team itself has to pull a docker image or something like that needs to be done as per the instructions in the PyTorch XLA Official Github..\n\nPlease Avoid the installation steps. Sorry for the mess i created;",
      "votes": 2,
      "replies": [
        {
          "id": 744303,
          "postDate": "2020-02-12T18:02:26.430Z",
          "content": "<p>The TPU you are getting is a single board (4 chips, 8 cores) not a pod (TPU pod = 2 to 256 TPU boards). But I do not think this is the issue.\nThere is probably some work needed on our end on the docker image and its config. I can confirm that this work has not been done yet and TPU support in PyTorch was not part of this launch.</p>",
          "rawMarkdown": "The TPU you are getting is a single board (4 chips, 8 cores) not a pod (TPU pod = 2 to 256 TPU boards). But I do not think this is the issue.\nThere is probably some work needed on our end on the docker image and its config. I can confirm that this work has not been done yet and TPU support in PyTorch was not part of this launch.",
          "votes": 2
        },
        {
          "id": 744304,
          "postDate": "2020-02-12T18:06:09.863Z",
          "content": "<p>Just in case this helps, there's a <a href=\"https://github.com/pytorch/xla/issues/1447#issuecomment-585324126\">shell script on GCP VM's</a> that can install the same on the docker image maybe, link points to Github issues... \n- PS i am very new to TPU, Just Exploring it! Thanks for the correction Martin!</p>\n\n<blockquote>\n  <p>TPU support in PyTorch was not part of this launch.</p>\n</blockquote>\n\n<p>Would be really great if this can be figured out in upcoming weeks or so!</p>",
          "rawMarkdown": "Just in case this helps, there's a [shell script on GCP VM's](https://github.com/pytorch/xla/issues/1447#issuecomment-585324126) that can install the same on the docker image maybe, link points to Github issues... \n- PS i am very new to TPU, Just Exploring it! Thanks for the correction Martin!\n\n&gt;TPU support in PyTorch was not part of this launch.\n\nWould be really great if this can be figured out in upcoming weeks or so!"
        },
        {
          "id": 744438,
          "postDate": "2020-02-12T20:39:58.607Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I actually don't think there is any error with the TPU setup, and if I were to follow similar instructions to how PyTorch XLA is set up on GCP, I think it should work. However, right now the major issue is correctly opening up the data.</p>",
          "rawMarkdown": "@mgornergoogle I actually don't think there is any error with the TPU setup, and if I were to follow similar instructions to how PyTorch XLA is set up on GCP, I think it should work. However, right now the major issue is correctly opening up the data."
        }
      ]
    },
    {
      "id": 886057,
      "postDate": "2020-06-14T17:18:11.497Z",
      "content": "<p>Hello,\nFirst, thank you, everyone, for the help with codes and resources to solve this issue. \nAs most of the comments are dated 4 months ago, I was curious if this issue has been solved yet? I am browsing the net and trying to train a model using Pytorch with Kaggle TPU and it is not working. \nIf there is any trick or alternative to make it work, I will really appreciate the help.</p>",
      "rawMarkdown": "Hello,\nFirst, thank you, everyone, for the help with codes and resources to solve this issue. \nAs most of the comments are dated 4 months ago, I was curious if this issue has been solved yet? I am browsing the net and trying to train a model using Pytorch with Kaggle TPU and it is not working. \nIf there is any trick or alternative to make it work, I will really appreciate the help.",
      "replies": [
        {
          "id": 941298,
          "postDate": "2020-07-23T06:54:32.387Z",
          "content": "<p>It's has been resolved already; Apologies for a delayed response, I was not checking in kaggle lately)</p>",
          "rawMarkdown": "It's has been resolved already; Apologies for a delayed response, I was not checking in kaggle lately)"
        },
        {
          "id": 949069,
          "postDate": "2020-07-28T12:23:27.927Z",
          "content": "<p>Thank you! Great news :)</p>",
          "rawMarkdown": "Thank you! Great news :)"
        }
      ]
    },
    {
      "id": 806379,
      "postDate": "2020-04-13T17:12:02.280Z",
      "content": "<p>It's not easy:)</p>",
      "rawMarkdown": "It's not easy:)",
      "replies": [
        {
          "id": 806384,
          "postDate": "2020-04-13T17:17:02.057Z",
          "content": "<p>It's indeed easy now <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">refer</a> :)</p>",
          "rawMarkdown": "It's indeed easy now [refer](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005) :)"
        }
      ]
    },
    {
      "id": 748544,
      "postDate": "2020-02-17T16:39:49.107Z",
      "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> have you able to make it working?</p>",
      "rawMarkdown": "@adityaecdrid have you able to make it working?",
      "replies": [
        {
          "id": 749556,
          "postDate": "2020-02-18T19:20:49.940Z",
          "content": "<p>Try adding <code>os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470</code> as in the comment above. This may help, but again, no guarantees that things will work further.</p>",
          "rawMarkdown": "Try adding `os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470` as in the comment above. This may help, but again, no guarantees that things will work further.",
          "votes": 2
        }
      ]
    },
    {
      "id": 746086,
      "postDate": "2020-02-14T15:51:56.133Z",
      "content": "<p>Just a Minor Update, i guess PyTorch users need to wait for start using TF for the comp and TPU's, because i tried running a Bert script which runs without any issues on colab here (for pytorch),But it's literally stuck <strong>indefinitely</strong>.... :(((</p>",
      "rawMarkdown": "Just a Minor Update, i guess PyTorch users need to wait for start using TF for the comp and TPU's, because i tried running a Bert script which runs without any issues on colab here (for pytorch),But it's literally stuck **indefinitely**.... :((("
    },
    {
      "id": 750751,
      "postDate": "2020-02-19T16:59:15.323Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 750959,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-02-19T21:48:13.160000",
      "content": "<p>Dear PyTorch and TPU fans.</p>\n\n<p>At this point, we have been able to confirm that there is additional work to be done on the Kaggle end to make the PyTorch/TPU integration work. It is unfortunately not possible to make it work in user space. Thank you for your hacking attempts though. They gave us good insights into what was need. We are in touch with the PyTorch TPU team at Google and the next step for us is to finish the technical analysis of the situation with them.</p>",
      "votes": 20,
      "replies": [
        {
          "id": 750969,
          "author_name": "Abhishek Thakur",
          "author_url": "",
          "post_date": "2020-02-19T21:59:08.300000",
          "content": "<p>Oh no!!! :D</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 750981,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-19T22:12:09.857000",
          "content": "<p>This is unfortunate but I am glad you are working with the PyTorch XLA team to solve this issue. Do you have an expected time frame for the integration of PyTorch XLA with the Kaggle TPUs? Do you expect it to happen before the end of this competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751010,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-19T23:20:23.970000",
          "content": "<p>Also, <a href=\"/mgornergoogle\">@mgornergoogle</a> could you please elaborate what is the current issue that needs to be fixed? Because it was possible for PyTorch XLA to at least recognize the device. Is the error related to the <a href=\"https://github.com/pytorch/xla/issues/1619\">GitHub issue</a>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751098,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-20T01:54:23.430000",
          "content": "<p>No, this is independent from the potential TFRecord reader issue. The issue is PyTorch supporting code running on the Cloud TPU itself.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 751261,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-20T04:35:56.050000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for the clarification. Is it expected for this separate issue to be fixed before the end of the competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751323,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-20T05:50:41.477000",
          "content": "<p>I don't have enough info to answer at this point. We are meeting with the PyTorch TPU team next week to sync our roadmaps.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 751345,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-20T06:23:59.957000",
          "content": "<p>Thank you for your response. Please let us know once the time frame is determined.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757571,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-26T23:25:12.647000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> It's been a week since the last response. Is there an update on the timeframe?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758577,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-28T00:06:56.723000",
          "content": "<p>Yes, the timeframe will be &gt; 1 month. That's all I can say for now. Still scoping the work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758694,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-28T03:41:42.850000",
          "content": "<p>Ok thanks for the update. I hope there will be enough time to participate in this competition with PyTorch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758695,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-28T03:42:38.287000",
          "content": "<p>Please let us know of any additional updates on the time frame and progress. Greatly appreciate it! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 749549,
      "author_name": "Ilya Figotin",
      "author_url": "",
      "post_date": "2020-02-18T19:18:04.450000",
      "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> you could try adding <code>XRT_TPU_CONFIG</code> environment variable before your code as follows:\n<code>os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470\"</code>.</p>\n\n<p>This should unblock you one step further. However, we cannot guarantee that training will work as we have not yet worked on TPU support in PyTorch.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 749665,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-18T20:50:20.120000",
          "content": "<p>Yes, this works, but PyTorch XLA is still unusable for this competition unless the <a href=\"https://github.com/pytorch/xla/issues/1619\">TFrecords issue</a> is fixed.</p>\n\n<p>I see you are part of the Kaggle Team. Is the Kaggle Team working on integrating PyTorch XLA with the Kaggle TPUs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 749849,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-19T00:00:17.113000",
          "content": "<p>Thanks for getting back kaggle team!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 749858,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-19T00:08:05.153000",
          "content": "<p>I responded in the <a href=\"https://github.com/pytorch/xla/issues/1619\">TFRecords issue</a>. I do not see anything wrong with the data returned. The only thing that seems wrong is that you are trying to access an \"id\" field that does not exist in the TFRecords and instead of getting an error, you are getting bogus data. Not sure why. There is a workaround: do not try to access fields that do not exist in the dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 749866,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-02-19T00:28:06.643000",
          "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> can you share the Colab code sample you are referring to? To my knowledge there shouldn't be anything relevant on that port. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 749878,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-19T00:51:03.470000",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a>  Hey <a href=\"https://gist.github.com/AdityaSoni19031997/f877ebb73dd1b10c1758505eac08abae\">here</a> it's (gist link), me not sure whether i am heading towards correct direction though 😅  (because on GCP Vm's it isn't there in the shell scripts they use to set up everything for you);</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 741740,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-02-10T23:16:32.713000",
      "content": "<p>PyTorch is not yet officially supported on Kaggle TPUs. This launch focuses on tpus + Keras and Tensorflow 2.1. We did not test with PyTorch.\nHowever, if you manage to make it work with PyTorch, please let us know.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 741748,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-10T23:29:16.767000",
          "content": "<p>Cool Martin Will Update! </p>\n\n<ul>\n<li>It seems it's gonna be a bump ride because datasets is also being saved/read as per tf format? (as you showed in your Great Starter kernel!)</li>\n</ul>\n\n<p>Thanks for your time!</p>\n\n<p>When i executed,</p>\n\n<p><code>\nimport torch_xla.core.xla_model as xm\nxm.xrt_world_size() ## gives 1\ndevices = xm.get_xla_supported_devices() ## This never completes execution and nbs freezes i guess\n</code></p>\n\n<p>Quick Experiments, So I can be wrong!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741871,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-11T01:23:43.223000",
          "content": "<p>I have the same issue ☹️ \nCan't execute <code>torch_xla._XLAC._xla_get_devices()</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741882,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-11T01:40:16.650000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741884,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T01:41:19.363000",
          "content": "<p>Datasets are only read. and yes, the dataset for this competition is provided in TFRecord format for convenience. I would not expect the Dataset format to be an issue for PyTorch. Is it ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741885,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-11T01:42:54.453000",
          "content": "<p>PyTorch XLA has a TfRecord Reader over <a href=\"https://github.com/pytorch/xla/blob/master/torch_xla/utils/tf_record_reader.py\">here</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 741888,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-11T01:47:42.343000",
          "content": "<p>Preliminary tests seem to indicate it works fine:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F674553%2F17cd2302ee6ca4b61f2b8a94a54401e9%2Fpytorch_xla_tfrecord.PNG?generation=1581385710518353&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741945,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T02:33:19.433000",
          "content": "<p>Great!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 741959,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-11T02:40:46.527000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> Actually, upon further inspection this does not look correct. The image tensor is not divisible by 512. Also, there are multiple id's for a single image it seems. What do you think? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743028,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T18:18:11.413000",
          "content": "<p>The images in the <a href=\"https://www.kaggle.com/mgornergoogle/five-flowers\">five flowers</a> or the the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/data\">104 flowers</a> datasets are JPEG encoded. You need to decode them first. Then you will get 512x512 pixel images.</p>\n\n<p>The multiple id's for a single image are not normal. Maybe you are looking at a batch and not a single data element?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 743189,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-11T21:48:28.577000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> \nDo I have to decode directly from the bytes? Or do I decode from the image tensor? PyTorch doesn't seem to have a built-in function for this as opposed to TensorFlow.</p>\n\n<p>Regarding the multiple ids, technically this function is supposed to read only a single example. If I was looking at a batch, shouldn't I be seeing multiple classes though? Maybe there is a bug in the way the example is formatted? This might be expected as the function does not let me define the expected format of the TFRecord like TF does.</p>\n\n<p>If I have the raw bytes, is there a way to convert this into a dictionary? Because Pytorch XLA also has a <a href=\"https://github.com/pytorch/xla/blob/e68039df347d0023a71800e0ebfea5c08168e5f6/torch_xla/utils/tf_record_reader.py#L33\">function</a> that will read the entire dataset to bytes, and I wonder if I can convert that to a proper PyTorch dataset? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 743201,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-11T22:00:13.110000",
          "content": "<p>I'm afraid my experience with Pytorch is minimal... I can try to help on TFRecords though.\nIn a TFRecord (or \"example\") as it is sometimes called in the documentation, you can have only three types of \"things\":\n- a list of bytestrings\n- a list of int64\n- a list of float32\nIn the 100 flowers dataset, the format of each TFRecord of labeled data is:\n\"image\": list of bytestrings containing 1 bytestring (the JPEG-ecoded image bytes)\n\"label\": list of int64 containing 1 int64</p>\n\n<p>I hope this helps some.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 744285,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-12T17:42:05.937000",
          "content": "<p>Here's how one can atleast decode those images back <a href=\"https://gist.github.com/dlibenzi/0075e27fca67ce31f7a6d701d77de48a\">gist</a></p>\n\n<p>PS I am not the author but verified it works like a charm but speed is somewhat we need to worry about; Next step might be to wrap it around a PyTorch DataSet (Iterator) (maybe?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744440,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-12T20:41:15.977000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> As Aditya said, there is a way to decode the images in PyTorch. However, the format is still wrong. Please see <a href=\"https://github.com/pytorch/xla/issues/1619\">this GitHub issue</a> for more information.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 746240,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-14T19:20:48.403000",
          "content": "<p>I had a look at the GitHub issue and rewsponded there:</p>\n\n<p>In the decoded TFRecords posted by @tmabraham, for training data, \"class\" and \"image\" seem to be correct, \"id\" is not a label that exists in this dataset so I do not know what is being returned there. I would venture that it's a case of bad error handling in the API. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 743485,
      "author_name": "Tarun Paparaju",
      "author_url": "",
      "post_date": "2020-02-12T05:27:03.667000",
      "content": "<p>I'm struggling to install Pytorch-XLA <a href=\"/adityaecdrid\">@adityaecdrid</a>, the code snippet you have given runs properly, but when I try to import torch or torch-xla, it gives some weird errors. Are you able to import torch-xla?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 743534,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-12T06:12:55.400000",
          "content": "<p><a href=\"/tarunpaparaju\">@tarunpaparaju</a>  Well i have released it as a <a href=\"https://www.kaggle.com/adityaecdrid/quest-to-use-use-pytorch-xla\">nbs</a>, check it once...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 743607,
          "author_name": "Tarun Paparaju",
          "author_url": "",
          "post_date": "2020-02-12T06:20:59.487000",
          "content": "<p>Thanks a lot <a href=\"/adityaecdrid\">@adityaecdrid</a> 😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744147,
          "author_name": "Seph Pace",
          "author_url": "",
          "post_date": "2020-02-12T16:02:31.600000",
          "content": "<p>I was able to import it with your notebook, but when I execute <code>dev = xm.xla_device()</code> it runs indefinitely.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744198,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-12T16:42:27.740000",
          "content": "<p>Yep precisely the issue i also faced hence dropped my experiments for now; Something is kinda wrong here because TF says it has 8 but PyTorch only detects one..\nNB Take my words with a pinch of salt, i am a beginner wrt TPU's and XLA's (and can be completely wrong)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938602,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-21T16:15:40.870000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I am trying to implement multiprocessing on TPU with PyTorch and I am getting this weird error saying it can't replicate devices to 8 because there is only one. Also, when I tried <code>xm.xrt_world_size()</code> I got 1. So, it seems I have the same problem as you, Can you tell me any more about  this and if possible, how to deal with it? Any help will be greatly appreciated! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941304,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-07-23T06:56:13.443000",
          "content": "<p>Don't do xm.xrt_world_size() in a global scope; You should refer the examples from the xla repo; They are very good IMO.</p>\n\n<p><a href=\"https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/multi-core-alexnet-fashion-mnist.ipynb\">https://colab.research.google.com/github/pytorch/xla/blob/master/contrib/colab/multi-core-alexnet-fashion-mnist.ipynb</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941389,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-23T07:44:50.557000",
          "content": "<p>Sure, will see it, also, someone told me that there is a problem with Pytorch XLA on Kaggle due to some recwnt update and I should rather search for another option, like pytorch lightening or tensorflow, is it true?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941594,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-07-23T10:10:39.527000",
          "content": "<p>I am not quite sure about that claim that XLA doesn't work, maybe you can try forking some xla kernel from Abhishek's kernels and just run it to confirm.. Yesterday, i tested it on Colab, XLA worked great! fine!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 941620,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-07-23T10:26:54.900000",
          "content": "<p>Cool, thanks, will do it for sure! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 744243,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-02-12T17:17:28.643000",
      "content": "<p>Sp guys the above won't work as we think because that's how you do it for Colab nbs, but this is a pod instance we are getting access to, so Kaggle team itself has to pull a docker image or something like that needs to be done as per the instructions in the PyTorch XLA Official Github..</p>\n\n<p>Please Avoid the installation steps. Sorry for the mess i created;</p>",
      "votes": 2,
      "replies": [
        {
          "id": 744303,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-02-12T18:02:26.430000",
          "content": "<p>The TPU you are getting is a single board (4 chips, 8 cores) not a pod (TPU pod = 2 to 256 TPU boards). But I do not think this is the issue.\nThere is probably some work needed on our end on the docker image and its config. I can confirm that this work has not been done yet and TPU support in PyTorch was not part of this launch.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 744304,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-02-12T18:06:09.863000",
          "content": "<p>Just in case this helps, there's a <a href=\"https://github.com/pytorch/xla/issues/1447#issuecomment-585324126\">shell script on GCP VM's</a> that can install the same on the docker image maybe, link points to Github issues... \n- PS i am very new to TPU, Just Exploring it! Thanks for the correction Martin!</p>\n\n<blockquote>\n  <p>TPU support in PyTorch was not part of this launch.</p>\n</blockquote>\n\n<p>Would be really great if this can be figured out in upcoming weeks or so!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744438,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2020-02-12T20:39:58.607000",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I actually don't think there is any error with the TPU setup, and if I were to follow similar instructions to how PyTorch XLA is set up on GCP, I think it should work. However, right now the major issue is correctly opening up the data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 886057,
      "author_name": "Ghada",
      "author_url": "",
      "post_date": "2020-06-14T17:18:11.497000",
      "content": "<p>Hello,\nFirst, thank you, everyone, for the help with codes and resources to solve this issue. \nAs most of the comments are dated 4 months ago, I was curious if this issue has been solved yet? I am browsing the net and trying to train a model using Pytorch with Kaggle TPU and it is not working. \nIf there is any trick or alternative to make it work, I will really appreciate the help.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 941298,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-07-23T06:54:32.387000",
          "content": "<p>It's has been resolved already; Apologies for a delayed response, I was not checking in kaggle lately)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949069,
          "author_name": "Ghada",
          "author_url": "",
          "post_date": "2020-07-28T12:23:27.927000",
          "content": "<p>Thank you! Great news :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 806379,
      "author_name": " Igor Krasovskiy",
      "author_url": "",
      "post_date": "2020-04-13T17:12:02.280000",
      "content": "<p>It's not easy:)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 806384,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-04-13T17:17:02.057000",
          "content": "<p>It's indeed easy now <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">refer</a> :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 748544,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-02-17T16:39:49.107000",
      "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> have you able to make it working?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 749556,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-02-18T19:20:49.940000",
          "content": "<p>Try adding <code>os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470</code> as in the comment above. This may help, but again, no guarantees that things will work further.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 746086,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-02-14T15:51:56.133000",
      "content": "<p>Just a Minor Update, i guess PyTorch users need to wait for start using TF for the comp and TPU's, because i tried running a Bert script which runs without any issues on colab here (for pytorch),But it's literally stuck <strong>indefinitely</strong>.... :(((</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 750751,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-19T16:59:15.323000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "741730": "# Update\n- I am not sure if the below is APT For kernels!\n\n\n==============================================\n\nI tried setting up PyTorch-XLA using this,\n\n```\nimport collections\nfrom datetime import datetime, timedelta\nimport os\nimport tensorflow as tf\nimport numpy as np\n\n_VersionConfig = collections.namedtuple('_VersionConfig', 'wheels,server')\nVERSION = \"torch_xla==nightly\"  #@param [\"torch_xla==nightly\"]\nCONFIG = {\n    'torch_xla==nightly': _VersionConfig('nightly', 'XRT-dev{}'.format(\n        (datetime.today() - timedelta(1)).strftime('%Y%m%d'))),\n}[VERSION]\n\nDIST_BUCKET = 'gs://tpu-pytorch/wheels'\nTORCH_WHEEL = 'torch-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCH_XLA_WHEEL = 'torch_xla-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\nTORCHVISION_WHEEL = 'torchvision-{}-cp36-cp36m-linux_x86_64.whl'.format(CONFIG.wheels)\n! export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH\n\n# Install TPU compat PyTorch/TPU wheels and dependencies\n!pip uninstall -y torch torchvision\n!gsutil cp \"$DIST_BUCKET/$TORCH_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCH_XLA_WHEEL\" .\n!gsutil cp \"$DIST_BUCKET/$TORCHVISION_WHEEL\" .\n!pip install \"$TORCH_WHEEL\"\n!pip install \"$TORCH_XLA_WHEEL\"\n!pip install \"$TORCHVISION_WHEEL\"\n!apt-get install libomp5 -y\n!apt-get install libopenblas-dev -y\n\nimport torch_xla.core.xla_model as xm\nimport torch_xla.distributed.data_parallel as dp # http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multithreading\nimport torch_xla.distributed.xla_multiprocessing as xmp # http://pytorch.org/xla/index.html#running-on-multiple-xla-devices-with-multiprocessing\nimport torch_xla.distributed.parallel_loader as pl\n\nprint(xm.xrt_world_size()) # don't do this though, just to illustrate it installed things\n```\nIt installs correctly!",
    "750959": "Dear PyTorch and TPU fans.\n\nAt this point, we have been able to confirm that there is additional work to be done on the Kaggle end to make the PyTorch/TPU integration work. It is unfortunately not possible to make it work in user space. Thank you for your hacking attempts though. They gave us good insights into what was need. We are in touch with the PyTorch TPU team at Google and the next step for us is to finish the technical analysis of the situation with them.",
    "749549": "@adityaecdrid you could try adding `XRT_TPU_CONFIG` environment variable before your code as follows:\n`os.environ['XRT_TPU_CONFIG']=\"tpu_worker;0;10.0.0.2:8470\"`.\n\nThis should unblock you one step further. However, we cannot guarantee that training will work as we have not yet worked on TPU support in PyTorch.",
    "741740": "PyTorch is not yet officially supported on Kaggle TPUs. This launch focuses on tpus + Keras and Tensorflow 2.1. We did not test with PyTorch.\nHowever, if you manage to make it work with PyTorch, please let us know.",
    "743485": "I'm struggling to install Pytorch-XLA @adityaecdrid, the code snippet you have given runs properly, but when I try to import torch or torch-xla, it gives some weird errors. Are you able to import torch-xla?",
    "744243": "Sp guys the above won't work as we think because that's how you do it for Colab nbs, but this is a pod instance we are getting access to, so Kaggle team itself has to pull a docker image or something like that needs to be done as per the instructions in the PyTorch XLA Official Github..\n\nPlease Avoid the installation steps. Sorry for the mess i created;",
    "886057": "Hello,\nFirst, thank you, everyone, for the help with codes and resources to solve this issue. \nAs most of the comments are dated 4 months ago, I was curious if this issue has been solved yet? I am browsing the net and trying to train a model using Pytorch with Kaggle TPU and it is not working. \nIf there is any trick or alternative to make it work, I will really appreciate the help.",
    "806379": "It's not easy:)",
    "748544": "@adityaecdrid have you able to make it working?",
    "746086": "Just a Minor Update, i guess PyTorch users need to wait for start using TF for the comp and TPU's, because i tried running a Bert script which runs without any issues on colab here (for pytorch),But it's literally stuck **indefinitely**.... :(((",
    "750751": ""
  }
}