{
  "id": 142624,
  "title": "Doing Things on Fly Should Help With TPU's OOM?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/142624",
  "author_name": "",
  "post_date": "2020-04-11T12:25:16.313663500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Well we all know why OOM happens; </p>\n\n<p>So to circumvent that, we can use the idea inspired from from the fact when we use GPU's, we can do things on fly aka at batch-level;</p>\n\n<p>&gt; NB When i tested below, it kinda runs perfectly for single core; for multiple cores, i am trying to see why it's taking little longer to run..</p>\n\n<p>Sample Code Attached; (the below can be combined as a custom collate_fn as well)</p>\n\n<p>As a <a href=\"https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly/\">kernel</a></p>\n\n<p>Will be glad if you can share your thoughts?</p>\n\n<p>Thanks,\nAditya.</p>\n\n<hr>\n\n<p>In your TPU usage exp, does something like this can help Martin? cc <a href=\"/mgornergoogle\">@mgornergoogle</a></p>",
  "messages": [
    {
      "id": "804276",
      "postDate": "04/11/2020 12:25:16",
      "content": "<p>Well we all know why OOM happens; </p>\n\n<p>So to circumvent that, we can use the idea inspired from from the fact when we use GPU's, we can do things on fly aka at batch-level;</p>\n\n<p>&gt; NB When i tested below, it kinda runs perfectly for single core; for multiple cores, i am trying to see why it's taking little longer to run..</p>\n\n<p>Sample Code Attached; (the below can be combined as a custom collate_fn as well)</p>\n\n<p>As a <a href=\"https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly/\">kernel</a></p>\n\n<p>Will be glad if you can share your thoughts?</p>\n\n<p>Thanks,\nAditya.</p>\n\n<hr>\n\n<p>In your TPU usage exp, does something like this can help Martin? cc <a href=\"/mgornergoogle\">@mgornergoogle</a></p>",
      "rawMarkdown": "Well we all know why OOM happens; \n\nSo to circumvent that, we can use the idea inspired from from the fact when we use GPU's, we can do things on fly aka at batch-level;\n\n&gt; NB When i tested below, it kinda runs perfectly for single core; for multiple cores, i am trying to see why it's taking little longer to run..\n\nSample Code Attached; (the below can be combined as a custom collate_fn as well)\n\nAs a [kernel](https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly/)\n\nWill be glad if you can share your thoughts?\n\nThanks,\nAditya.\n\n-------------------\nIn your TPU usage exp, does something like this can help Martin? cc @mgornergoogle",
      "votes": null
    },
    {
      "id": "806345",
      "postDate": "04/13/2020 16:25:53",
      "content": "<p>The PyTorchXLA team has made some memory optimizations (<a href=\"https://github.com/pytorch/xla/issues/1870\">GitHub thread here</a>). Did you try that before going to more exotic solutions ?</p>",
      "rawMarkdown": "The PyTorchXLA team has made some memory optimizations ([GitHub thread here](https://github.com/pytorch/xla/issues/1870)). Did you try that before going to more exotic solutions ?",
      "votes": null
    },
    {
      "id": "806347",
      "postDate": "04/13/2020 16:28:23",
      "content": "<p>Yep! After that optimization only, We have got it working but Martin; Check this <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">thread</a></p>",
      "rawMarkdown": "Yep! After that optimization only, We have got it working but Martin; Check this [thread](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 806345,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "04/13/2020 16:25:53",
      "content": "<p>The PyTorchXLA team has made some memory optimizations (<a href=\"https://github.com/pytorch/xla/issues/1870\">GitHub thread here</a>). Did you try that before going to more exotic solutions ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 806347,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/13/2020 16:28:23",
          "content": "<p>Yep! After that optimization only, We have got it working but Martin; Check this <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005\">thread</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "804276": "Well we all know why OOM happens; \n\nSo to circumvent that, we can use the idea inspired from from the fact when we use GPU's, we can do things on fly aka at batch-level;\n\n&gt; NB When i tested below, it kinda runs perfectly for single core; for multiple cores, i am trying to see why it's taking little longer to run..\n\nSample Code Attached; (the below can be combined as a custom collate_fn as well)\n\nAs a [kernel](https://www.kaggle.com/adityaecdrid/sample-tpu-xlmr-pytorch-pad-on-fly/)\n\nWill be glad if you can share your thoughts?\n\nThanks,\nAditya.\n\n-------------------\nIn your TPU usage exp, does something like this can help Martin? cc @mgornergoogle",
    "806345": "The PyTorchXLA team has made some memory optimizations ([GitHub thread here](https://github.com/pytorch/xla/issues/1870)). Did you try that before going to more exotic solutions ?",
    "806347": "Yep! After that optimization only, We have got it working but Martin; Check this [thread](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/143005)"
  },
  "source": "meta"
}