{
  "id": 149991,
  "title": "Inference using XLM-Roberta on PyTorch TPU",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/149991",
  "author_name": "",
  "post_date": "2020-05-10T17:59:56.691608500Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n\n<p>I'm trying PyTorch using TPU. I saw this notebook <a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a> which I try to replicate. When I set n_procs to 1 in xmp.spawn function() it works but it is super slow. When I set it to 8, it only predicts 1/8 of the test set. </p>\n\n<p>Anyone having the same issue as me?</p>",
  "messages": [
    {
      "id": "841300",
      "postDate": "05/10/2020 17:59:56",
      "content": "<p>Hello everyone, </p>\n\n<p>I'm trying PyTorch using TPU. I saw this notebook <a href=\"https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta</a> which I try to replicate. When I set n_procs to 1 in xmp.spawn function() it works but it is super slow. When I set it to 8, it only predicts 1/8 of the test set. </p>\n\n<p>Anyone having the same issue as me?</p>",
      "rawMarkdown": "Hello everyone, \n\nI'm trying PyTorch using TPU. I saw this notebook https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta which I try to replicate. When I set n_procs to 1 in xmp.spawn function() it works but it is super slow. When I set it to 8, it only predicts 1/8 of the test set. \n\nAnyone having the same issue as me?",
      "votes": null
    },
    {
      "id": "841354",
      "postDate": "05/10/2020 18:44:12",
      "content": "<p>Hello,\nI guess that 8 files are produced using n_procs=8. you need then to merge those files in one.\nYou need also to be sure about the order of the records in each file to match the correct line id for submission, otherwise, you will get a poor submission ;). \nOn my side, i simply keep the id in the dataloader and predict without activating bfloat16, otherwise, it's a mess as for exemple 470=471 with bfloat16.</p>\n\n<p>There's an option too with xm.rendezvous(...) that allows cores to synchronize and pass data as bytes... I did not try this as it needs to deal with bytes 🙈 </p>",
      "rawMarkdown": "Hello,\nI guess that 8 files are produced using n_procs=8. you need then to merge those files in one.\nYou need also to be sure about the order of the records in each file to match the correct line id for submission, otherwise, you will get a poor submission ;). \nOn my side, i simply keep the id in the dataloader and predict without activating bfloat16, otherwise, it's a mess as for exemple 470=471 with bfloat16.\n\nThere's an option too with xm.rendezvous(...) that allows cores to synchronize and pass data as bytes... I did not try this as it needs to deal with bytes 🙈",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 841354,
      "author_name": "seb6084",
      "author_url": "",
      "post_date": "05/10/2020 18:44:12",
      "content": "<p>Hello,\nI guess that 8 files are produced using n_procs=8. you need then to merge those files in one.\nYou need also to be sure about the order of the records in each file to match the correct line id for submission, otherwise, you will get a poor submission ;). \nOn my side, i simply keep the id in the dataloader and predict without activating bfloat16, otherwise, it's a mess as for exemple 470=471 with bfloat16.</p>\n\n<p>There's an option too with xm.rendezvous(...) that allows cores to synchronize and pass data as bytes... I did not try this as it needs to deal with bytes 🙈 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "841300": "Hello everyone, \n\nI'm trying PyTorch using TPU. I saw this notebook https://www.kaggle.com/shonenkov/tpu-inference-super-fast-xlmroberta which I try to replicate. When I set n_procs to 1 in xmp.spawn function() it works but it is super slow. When I set it to 8, it only predicts 1/8 of the test set. \n\nAnyone having the same issue as me?",
    "841354": "Hello,\nI guess that 8 files are produced using n_procs=8. you need then to merge those files in one.\nYou need also to be sure about the order of the records in each file to match the correct line id for submission, otherwise, you will get a poor submission ;). \nOn my side, i simply keep the id in the dataloader and predict without activating bfloat16, otherwise, it's a mess as for exemple 470=471 with bfloat16.\n\nThere's an option too with xm.rendezvous(...) that allows cores to synchronize and pass data as bytes... I did not try this as it needs to deal with bytes 🙈"
  },
  "source": "meta"
}