{
  "id": 138265,
  "title": "I’m going live for this competition",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138265",
  "author_name": "",
  "post_date": "2020-03-24T10:29:13.044783800Z",
  "votes": 69,
  "comment_count": 26,
  "views": 0,
  "content": "<p>It’s only fair if I share it here: I will be showing how to start with this competition live today! \nFind more information here: <a href=\"https://twitter.com/abhi1thakur/status/1242395976138186752\">https://twitter.com/abhi1thakur/status/1242395976138186752</a></p>\n\n<p>P.S. I am sharing it in the forums because I was asked to do so by some competitors when I shared a similar video for bengali.ai competition.</p>",
  "messages": [
    {
      "id": "784568",
      "postDate": "03/24/2020 10:29:13",
      "content": "<p>It’s only fair if I share it here: I will be showing how to start with this competition live today! \nFind more information here: <a href=\"https://twitter.com/abhi1thakur/status/1242395976138186752\">https://twitter.com/abhi1thakur/status/1242395976138186752</a></p>\n\n<p>P.S. I am sharing it in the forums because I was asked to do so by some competitors when I shared a similar video for bengali.ai competition.</p>",
      "rawMarkdown": "It’s only fair if I share it here: I will be showing how to start with this competition live today! \nFind more information here: https://twitter.com/abhi1thakur/status/1242395976138186752\n\nP.S. I am sharing it in the forums because I was asked to do so by some competitors when I shared a similar video for bengali.ai competition.",
      "votes": null
    },
    {
      "id": "785010",
      "postDate": "03/24/2020 17:44:26",
      "content": "<p>If you missed it, you can watch it here: <a href=\"https://youtu.be/vvr_f-X_LaI\">https://youtu.be/vvr_f-X_LaI</a></p>\n\n<p>TPU training kernel: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training</a>\nInference kernel: <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml\">https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml</a></p>",
      "rawMarkdown": "If you missed it, you can watch it here: https://youtu.be/vvr_f-X_LaI\n\n\nTPU training kernel: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\nInference kernel: https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml",
      "votes": null
    },
    {
      "id": "785198",
      "postDate": "03/24/2020 21:51:59",
      "content": "<p>Thank you for the PyTorch + BERT + TPU  getting started materials!</p>",
      "rawMarkdown": "Thank you for the PyTorch + BERT + TPU  getting started materials!",
      "votes": null
    },
    {
      "id": "785281",
      "postDate": "03/25/2020 00:09:12",
      "content": "<p>Tell us if you find a way to make the 8 TPU cores work. I also asked the PyTorch-XLA team. Unfortunately, I suspect this has to do with the limited RAM available on the Kaggle VM...</p>",
      "rawMarkdown": "Tell us if you find a way to make the 8 TPU cores work. I also asked the PyTorch-XLA team. Unfortunately, I suspect this has to do with the limited RAM available on the Kaggle VM...",
      "votes": null
    },
    {
      "id": "785289",
      "postDate": "03/25/2020 00:21:08",
      "content": "<p> 8 cores do work for bert-uncased but not for bert-multilingual</p>",
      "rawMarkdown": "8 cores do work for bert-uncased but not for bert-multilingual",
      "votes": null
    },
    {
      "id": "785608",
      "postDate": "03/25/2020 07:44:35",
      "content": "<p><a href=\"/abhishek\">@abhishek</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/adityaecdrid\">@adityaecdrid</a> with small modification it is indeed possible to train the bert-multlingual using all 8 cores. check out my <a href=\"https://www.kaggle.com/dhananjay3/bert-multi-lingual-tpu-training\">fork here</a>.</p>",
      "rawMarkdown": "abhishek @mgornergoogle @adityaecdrid with small modification it is indeed possible to train the bert-multlingual using all 8 cores. check out my [fork here](https://www.kaggle.com/dhananjay3/bert-multi-lingual-tpu-training).",
      "votes": null
    },
    {
      "id": "785612",
      "postDate": "03/25/2020 07:49:44",
      "content": "<p>Thanks for the fork <a href=\"/dhananjay3\">@dhananjay3</a> . Is it a wise idea to move model outside? Does it not effect anything? Im asking because all examples and tutorials I saw had the model inside.</p>",
      "rawMarkdown": "Thanks for the fork @dhananjay3 . Is it a wise idea to move model outside? Does it not effect anything? Im asking because all examples and tutorials I saw had the model inside.",
      "votes": null
    },
    {
      "id": "785618",
      "postDate": "03/25/2020 07:57:05",
      "content": "<p>Also, the reported AUC seems around 0.5. So its not working as it should?</p>",
      "rawMarkdown": "Also, the reported AUC seems around 0.5. So its not working as it should?",
      "votes": null
    },
    {
      "id": "785622",
      "postDate": "03/25/2020 08:01:23",
      "content": "<p>del</p>",
      "rawMarkdown": "del",
      "votes": null
    },
    {
      "id": "785637",
      "postDate": "03/25/2020 08:18:49",
      "content": "<p><a href=\"/abhishek\">@abhishek</a>  I don't think it's incorrect to do as it will be copied to each device separately.  Plus this will be the only way to get training correctly if there is some randomness while creating model. i.e. if starting weights differ in each instance then the distributed training will not work as far as I understand. the low score is because of high effective learning rate (8x) as compared to your notebook.  </p>",
      "rawMarkdown": "abhishek  I don't think it's incorrect to do as it will be copied to each device separately.  Plus this will be the only way to get training correctly if there is some randomness while creating model. i.e. if starting weights differ in each instance then the distributed training will not work as far as I understand. the low score is because of high effective learning rate (8x) as compared to your notebook.",
      "votes": null
    },
    {
      "id": "785647",
      "postDate": "03/25/2020 08:25:23",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I was looking into this a little bit and here's what I found.</p>\n\n<p>So in TF, you have <code>TPUStrategy</code> which is said to be similar to <code>MirroredStrategy</code> (for multi-GPU) but for TPUs. <code>MirroredStrategy</code> makes a replica of the model on each GPU. It would be fair to assume that <code>TPUStrategy</code> is doing the same for the TPU cores. Isn't this the same method PyTorch XLA multiprocessing using? So would TF also have problems with BERT multilingual models?</p>\n\n<p>The only way I can think of right now for TF to work is if TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?</p>",
      "rawMarkdown": "mgornergoogle I was looking into this a little bit and here's what I found.\n\nSo in TF, you have `TPUStrategy` which is said to be similar to `MirroredStrategy` (for multi-GPU) but for TPUs. `MirroredStrategy` makes a replica of the model on each GPU. It would be fair to assume that `TPUStrategy` is doing the same for the TPU cores. Isn't this the same method PyTorch XLA multiprocessing using? So would TF also have problems with BERT multilingual models?\n\nThe only way I can think of right now for TF to work is if TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?",
      "votes": null
    },
    {
      "id": "785652",
      "postDate": "03/25/2020 08:30:02",
      "content": "<blockquote>\n  <p>the low score is because of high effective learning rate (8x) as compared to your notebook.</p>\n</blockquote>\n\n<p>Easy way to verify, just set it correctly then because IMHO it's not correct as per XLA docs etc :)</p>",
      "rawMarkdown": "&gt; the low score is because of high effective learning rate (8x) as compared to your notebook.\n\nEasy way to verify, just set it correctly then because IMHO it's not correct as per XLA docs etc :)",
      "votes": null
    },
    {
      "id": "785662",
      "postDate": "03/25/2020 08:42:51",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> based on your suggested modification, i wrote a 8 core version of my original kernel. you can find it here: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores</a></p>\n\n<p>It gives an AUC of 0.84 after first epoch on the validation set.</p>\n\n<p>However, the kernel throws an error (OOM) when it reaches second epoch.</p>",
      "rawMarkdown": "dhananjay3 based on your suggested modification, i wrote a 8 core version of my original kernel. you can find it here: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\n\nIt gives an AUC of 0.84 after first epoch on the validation set.\n\nHowever, the kernel throws an error (OOM) when it reaches second epoch.",
      "votes": null
    },
    {
      "id": "785679",
      "postDate": "03/25/2020 09:08:48",
      "content": "<p><a href=\"/abhishek\">@abhishek</a> try reducing num_workers.\n<a href=\"/adityaecdrid\">@adityaecdrid</a> the next version of my notebook fixes the learning rate and the result is as expected so no there is nothing inherently wrong in this method.</p>",
      "rawMarkdown": "abhishek try reducing num_workers.\n@adityaecdrid the next version of my notebook fixes the learning rate and the result is as expected so no there is nothing inherently wrong in this method.",
      "votes": null
    },
    {
      "id": "785680",
      "postDate": "03/25/2020 09:11:42",
      "content": "<p>For whatever reason it might work <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902</a></p>",
      "rawMarkdown": "For whatever reason it might work https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902",
      "votes": null
    },
    {
      "id": "785681",
      "postDate": "03/25/2020 09:12:02",
      "content": "<p>instead i went for reduction of batch size and that seemed to work</p>",
      "rawMarkdown": "instead i went for reduction of batch size and that seemed to work",
      "votes": null
    },
    {
      "id": "785688",
      "postDate": "03/25/2020 09:18:35",
      "content": "<p><a href=\"/abhishek\">@abhishek</a> just to let you know you are currently using 64 as per core batch_size which gives effective batch size of 512. I don't see any point of having that big batch as this reduces num of steps per epoch which eventually increases total training time.</p>",
      "rawMarkdown": "abhishek just to let you know you are currently using 64 as per core batch_size which gives effective batch size of 512. I don't see any point of having that big batch as this reduces num of steps per epoch which eventually increases total training time.",
      "votes": null
    },
    {
      "id": "785882",
      "postDate": "03/25/2020 13:40:31",
      "content": "<p>Thanks for this to all involved. I have one question:\nIs the validation score calculated on the full valid data (8000 records)?</p>\n\n<p>When I print the length of the valid vector it has only length 1000 (divided by 8 forks). Doesn't this need to be collected somewhere and then combined?</p>\n\n<p>Edit: OK I think I understand, it just prints for master node. Is there a way to collect across processes?</p>",
      "rawMarkdown": "Thanks for this to all involved. I have one question:\nIs the validation score calculated on the full valid data (8000 records)?\n\nWhen I print the length of the valid vector it has only length 1000 (divided by 8 forks). Doesn't this need to be collected somewhere and then combined?\n\nEdit: OK I think I understand, it just prints for master node. Is there a way to collect across processes?",
      "votes": null
    },
    {
      "id": "785933",
      "postDate": "03/25/2020 14:20:28",
      "content": "<p>Hey folks, just attaching the reply from XLA Devs regarding what Dhananjay did,</p>\n\n<blockquote>\n  <p>It is not advised to create data structures outside the MP function target call stack.\n  That does not mean you have to create them directly inside the function, as long as they are created from any function which has the MP target function as parent in the call stack.\n  Or simply, avoid creating source code where models are global (that is, they get created when the module is loaded, without any function to be called).\n  When using the pytorch multiprocessing spawn start_method, any global data will have to be pickled, so that puts further restrictions on what you can create on global scope.\n  While for models that might work, you cannot call any XLA function besides xmp.spawn() (the set is a bit wider, but it is better to stay on the safe side) outside of the MP target function call stack.\n  Same issue with pickling applies to datasets.\n  On top of that, for multicore training you will need distributed samplers, which need the core ordinal, which is only available once you are within the MP target function call stack.</p>\n</blockquote>",
      "rawMarkdown": "Hey folks, just attaching the reply from XLA Devs regarding what Dhananjay did,\n\n&gt;It is not advised to create data structures outside the MP function target call stack.\nThat does not mean you have to create them directly inside the function, as long as they are created from any function which has the MP target function as parent in the call stack.\nOr simply, avoid creating source code where models are global (that is, they get created when the module is loaded, without any function to be called).\nWhen using the pytorch multiprocessing spawn start_method, any global data will have to be pickled, so that puts further restrictions on what you can create on global scope.\n&gt; While for models that might work, you cannot call any XLA function besides xmp.spawn() (the set is a bit wider, but it is better to stay on the safe side) outside of the MP target function call stack.\n&gt; Same issue with pickling applies to datasets.\nOn top of that, for multicore training you will need distributed samplers, which need the core ordinal, which is only available once you are within the MP target function call stack.",
      "votes": null
    },
    {
      "id": "786116",
      "postDate": "03/25/2020 16:34:16",
      "content": "<blockquote>\n  <p>TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?</p>\n</blockquote>\n\n<p>I'm trying to get a definitive answer from the PyTorch-XLA devs but I think that is what is happening.</p>",
      "rawMarkdown": "&gt; TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?\n\nI'm trying to get a definitive answer from the PyTorch-XLA devs but I think that is what is happening.",
      "votes": null
    },
    {
      "id": "786232",
      "postDate": "03/25/2020 18:06:59",
      "content": "<p>Yes. Thats the case afaik. :)</p>",
      "rawMarkdown": "Yes. Thats the case afaik. :)",
      "votes": null
    },
    {
      "id": "787219",
      "postDate": "03/26/2020 16:10:51",
      "content": "<p>Hi Abhishek,</p>\n\n<p>This is NLPclassification problem, Target variable is either 0 or 1 ( Comments is not toxic or toxic), But many many cases target variable is negative,  How we can interpret the results in terms of toxic or not toxic</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Abhishek,\n\nThis is NLPclassification problem, Target variable is either 0 or 1 ( Comments is not toxic or toxic), But many many cases target variable is negative,  How we can interpret the results in terms of toxic or not toxic\n\nThanks",
      "votes": null
    },
    {
      "id": "787677",
      "postDate": "03/27/2020 01:39:47",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    },
    {
      "id": "788742",
      "postDate": "03/28/2020 00:32:03",
      "content": "<p>I got some answers from the PyTorch/XLA team. Here are the two possible approaches (in pseudocode)</p>\n\n<p>1) With the model instantiated in the mp_fn:</p>\n\n<p>```\ndef _mp_fn():\n   model = instantiate_my_model()\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()</p>\n\n<p>def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...</p>\n\n<p>device = xm.xla_device()\n   model.to(device)</p>\n\n<p>for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)</p>\n\n<p>xmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```</p>\n\n<p>2) With the model instantiated outside of the mp_fn:\n```</p>\n\n<h1>model instantiated in global scope</h1>\n\n<p>model = instantiate_my_model()</p>\n\n<p>def _mp_fn():\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()</p>\n\n<p>def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...</p>\n\n<p>device = xm.xla_device()\n   model.to(device)</p>\n\n<p>for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)</p>\n\n<p>xmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```</p>\n\n<p>Approach #1 instantiates the model 8 times as the process is forked. With a large model this will run out of memory.\nApproach #2 instantiates the model only once, in global scope. When the process is forked, memory pages corresponding to the model will be marked COW (Copy on Write) and since the model is copied to TPU by <code>model.to(device)</code>, the version on the VM is constant and will not be copied.</p>\n\n<p>The only limitation to remember is that all XLA functions must be called from the mp_fn, which is the case here.</p>",
      "rawMarkdown": "I got some answers from the PyTorch/XLA team. Here are the two possible approaches (in pseudocode)\n\n1) With the model instantiated in the mp_fn:\n\n```\ndef _mp_fn():\n   model = instantiate_my_model()\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()\n\n   def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...\n\n   device = xm.xla_device()\n   model.to(device)\n\n   for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)\n\nxmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```\n\n2) With the model instantiated outside of the mp_fn:\n```\n# model instantiated in global scope\nmodel = instantiate_my_model()\n\ndef _mp_fn():\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()\n\n   def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...\n\n   device = xm.xla_device()\n   model.to(device)\n\n   for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)\n\nxmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```\n\nApproach #1 instantiates the model 8 times as the process is forked. With a large model this will run out of memory.\nApproach #2 instantiates the model only once, in global scope. When the process is forked, memory pages corresponding to the model will be marked COW (Copy on Write) and since the model is copied to TPU by `model.to(device)`, the version on the VM is constant and will not be copied.\n\nThe only limitation to remember is that all XLA functions must be called from the mp_fn, which is the case here.",
      "votes": null
    },
    {
      "id": "788747",
      "postDate": "03/28/2020 00:44:28",
      "content": "<p>Thanks for the clarification! I assume the dataloader cannot be defined outside of the training loop function because it is dependent on the xla functions to define the appropriate distributed sampler?</p>",
      "rawMarkdown": "Thanks for the clarification! I assume the dataloader cannot be defined outside of the training loop function because it is dependent on the xla functions to define the appropriate distributed sampler?",
      "votes": null
    },
    {
      "id": "788845",
      "postDate": "03/28/2020 04:33:26",
      "content": "<p>The way <a href=\"/abhishek\">@abhishek</a> did it <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">here</a> works. The data he loads in the global context is just pandas dataframes. The data loader objects acting on them are defined in the mp_fn. I believe this also prevents the data from being replicated in memory 8 times.</p>",
      "rawMarkdown": "The way @abhishek did it [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores) works. The data he loads in the global context is just pandas dataframes. The data loader objects acting on them are defined in the mp_fn. I believe this also prevents the data from being replicated in memory 8 times.",
      "votes": null
    },
    {
      "id": "828876",
      "postDate": "05/01/2020 10:51:13",
      "content": "<p>Why is it that the TPU idle is around 50%. I was trying to run <a href=\"/abhishek\">@abhishek</a> 's notebook. Will there be a difference if I go with TF rather than using pytorch xla?</p>",
      "rawMarkdown": "Why is it that the TPU idle is around 50%. I was trying to run @abhishek 's notebook. Will there be a difference if I go with TF rather than using pytorch xla?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 785010,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "03/24/2020 17:44:26",
      "content": "<p>If you missed it, you can watch it here: <a href=\"https://youtu.be/vvr_f-X_LaI\">https://youtu.be/vvr_f-X_LaI</a></p>\n\n<p>TPU training kernel: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training</a>\nInference kernel: <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml\">https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 785198,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/24/2020 21:51:59",
          "content": "<p>Thank you for the PyTorch + BERT + TPU  getting started materials!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785281,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/25/2020 00:09:12",
          "content": "<p>Tell us if you find a way to make the 8 TPU cores work. I also asked the PyTorch-XLA team. Unfortunately, I suspect this has to do with the limited RAM available on the Kaggle VM...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785289,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 00:21:08",
          "content": "<p> 8 cores do work for bert-uncased but not for bert-multilingual</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785608,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "03/25/2020 07:44:35",
          "content": "<p><a href=\"/abhishek\">@abhishek</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"/adityaecdrid\">@adityaecdrid</a> with small modification it is indeed possible to train the bert-multlingual using all 8 cores. check out my <a href=\"https://www.kaggle.com/dhananjay3/bert-multi-lingual-tpu-training\">fork here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785612,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "03/25/2020 07:49:44",
          "content": "<p>Thanks for the fork <a href=\"/dhananjay3\">@dhananjay3</a> . Is it a wise idea to move model outside? Does it not effect anything? Im asking because all examples and tutorials I saw had the model inside.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785618,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "03/25/2020 07:57:05",
          "content": "<p>Also, the reported AUC seems around 0.5. So its not working as it should?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785622,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 08:01:23",
          "content": "<p>del</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785637,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "03/25/2020 08:18:49",
          "content": "<p><a href=\"/abhishek\">@abhishek</a>  I don't think it's incorrect to do as it will be copied to each device separately.  Plus this will be the only way to get training correctly if there is some randomness while creating model. i.e. if starting weights differ in each instance then the distributed training will not work as far as I understand. the low score is because of high effective learning rate (8x) as compared to your notebook.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785647,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/25/2020 08:25:23",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I was looking into this a little bit and here's what I found.</p>\n\n<p>So in TF, you have <code>TPUStrategy</code> which is said to be similar to <code>MirroredStrategy</code> (for multi-GPU) but for TPUs. <code>MirroredStrategy</code> makes a replica of the model on each GPU. It would be fair to assume that <code>TPUStrategy</code> is doing the same for the TPU cores. Isn't this the same method PyTorch XLA multiprocessing using? So would TF also have problems with BERT multilingual models?</p>\n\n<p>The only way I can think of right now for TF to work is if TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785652,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 08:30:02",
          "content": "<blockquote>\n  <p>the low score is because of high effective learning rate (8x) as compared to your notebook.</p>\n</blockquote>\n\n<p>Easy way to verify, just set it correctly then because IMHO it's not correct as per XLA docs etc :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785662,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "03/25/2020 08:42:51",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> based on your suggested modification, i wrote a 8 core version of my original kernel. you can find it here: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores</a></p>\n\n<p>It gives an AUC of 0.84 after first epoch on the validation set.</p>\n\n<p>However, the kernel throws an error (OOM) when it reaches second epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785679,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "03/25/2020 09:08:48",
          "content": "<p><a href=\"/abhishek\">@abhishek</a> try reducing num_workers.\n<a href=\"/adityaecdrid\">@adityaecdrid</a> the next version of my notebook fixes the learning rate and the result is as expected so no there is nothing inherently wrong in this method.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785680,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 09:11:42",
          "content": "<p>For whatever reason it might work <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785681,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "03/25/2020 09:12:02",
          "content": "<p>instead i went for reduction of batch size and that seemed to work</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785688,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "03/25/2020 09:18:35",
          "content": "<p><a href=\"/abhishek\">@abhishek</a> just to let you know you are currently using 64 as per core batch_size which gives effective batch size of 512. I don't see any point of having that big batch as this reduces num of steps per epoch which eventually increases total training time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785882,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/25/2020 13:40:31",
          "content": "<p>Thanks for this to all involved. I have one question:\nIs the validation score calculated on the full valid data (8000 records)?</p>\n\n<p>When I print the length of the valid vector it has only length 1000 (divided by 8 forks). Doesn't this need to be collected somewhere and then combined?</p>\n\n<p>Edit: OK I think I understand, it just prints for master node. Is there a way to collect across processes?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 785933,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 14:20:28",
          "content": "<p>Hey folks, just attaching the reply from XLA Devs regarding what Dhananjay did,</p>\n\n<blockquote>\n  <p>It is not advised to create data structures outside the MP function target call stack.\n  That does not mean you have to create them directly inside the function, as long as they are created from any function which has the MP target function as parent in the call stack.\n  Or simply, avoid creating source code where models are global (that is, they get created when the module is loaded, without any function to be called).\n  When using the pytorch multiprocessing spawn start_method, any global data will have to be pickled, so that puts further restrictions on what you can create on global scope.\n  While for models that might work, you cannot call any XLA function besides xmp.spawn() (the set is a bit wider, but it is better to stay on the safe side) outside of the MP target function call stack.\n  Same issue with pickling applies to datasets.\n  On top of that, for multicore training you will need distributed samplers, which need the core ordinal, which is only available once you are within the MP target function call stack.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 786116,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/25/2020 16:34:16",
          "content": "<blockquote>\n  <p>TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?</p>\n</blockquote>\n\n<p>I'm trying to get a definitive answer from the PyTorch-XLA devs but I think that is what is happening.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 786232,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "03/25/2020 18:06:59",
          "content": "<p>Yes. Thats the case afaik. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 787219,
      "author_name": "praveengovi",
      "author_url": "",
      "post_date": "03/26/2020 16:10:51",
      "content": "<p>Hi Abhishek,</p>\n\n<p>This is NLPclassification problem, Target variable is either 0 or 1 ( Comments is not toxic or toxic), But many many cases target variable is negative,  How we can interpret the results in terms of toxic or not toxic</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 787677,
      "author_name": "shamrat",
      "author_url": "",
      "post_date": "03/27/2020 01:39:47",
      "content": "<p>👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 788742,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/28/2020 00:32:03",
      "content": "<p>I got some answers from the PyTorch/XLA team. Here are the two possible approaches (in pseudocode)</p>\n\n<p>1) With the model instantiated in the mp_fn:</p>\n\n<p>```\ndef _mp_fn():\n   model = instantiate_my_model()\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()</p>\n\n<p>def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...</p>\n\n<p>device = xm.xla_device()\n   model.to(device)</p>\n\n<p>for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)</p>\n\n<p>xmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```</p>\n\n<p>2) With the model instantiated outside of the mp_fn:\n```</p>\n\n<h1>model instantiated in global scope</h1>\n\n<p>model = instantiate_my_model()</p>\n\n<p>def _mp_fn():\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()</p>\n\n<p>def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...</p>\n\n<p>device = xm.xla_device()\n   model.to(device)</p>\n\n<p>for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)</p>\n\n<p>xmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```</p>\n\n<p>Approach #1 instantiates the model 8 times as the process is forked. With a large model this will run out of memory.\nApproach #2 instantiates the model only once, in global scope. When the process is forked, memory pages corresponding to the model will be marked COW (Copy on Write) and since the model is copied to TPU by <code>model.to(device)</code>, the version on the VM is constant and will not be copied.</p>\n\n<p>The only limitation to remember is that all XLA functions must be called from the mp_fn, which is the case here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 788747,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "03/28/2020 00:44:28",
          "content": "<p>Thanks for the clarification! I assume the dataloader cannot be defined outside of the training loop function because it is dependent on the xla functions to define the appropriate distributed sampler?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788845,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/28/2020 04:33:26",
          "content": "<p>The way <a href=\"/abhishek\">@abhishek</a> did it <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">here</a> works. The data he loads in the global context is just pandas dataframes. The data loader objects acting on them are defined in the mp_fn. I believe this also prevents the data from being replicated in memory 8 times.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 828876,
      "author_name": "noelmat",
      "author_url": "",
      "post_date": "05/01/2020 10:51:13",
      "content": "<p>Why is it that the TPU idle is around 50%. I was trying to run <a href=\"/abhishek\">@abhishek</a> 's notebook. Will there be a difference if I go with TF rather than using pytorch xla?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "784568": "It’s only fair if I share it here: I will be showing how to start with this competition live today! \nFind more information here: https://twitter.com/abhi1thakur/status/1242395976138186752\n\nP.S. I am sharing it in the forums because I was asked to do so by some competitors when I shared a similar video for bengali.ai competition.",
    "785010": "If you missed it, you can watch it here: https://youtu.be/vvr_f-X_LaI\n\n\nTPU training kernel: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\nInference kernel: https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml",
    "785198": "Thank you for the PyTorch + BERT + TPU  getting started materials!",
    "785281": "Tell us if you find a way to make the 8 TPU cores work. I also asked the PyTorch-XLA team. Unfortunately, I suspect this has to do with the limited RAM available on the Kaggle VM...",
    "785289": "8 cores do work for bert-uncased but not for bert-multilingual",
    "785608": "abhishek @mgornergoogle @adityaecdrid with small modification it is indeed possible to train the bert-multlingual using all 8 cores. check out my [fork here](https://www.kaggle.com/dhananjay3/bert-multi-lingual-tpu-training).",
    "785612": "Thanks for the fork @dhananjay3 . Is it a wise idea to move model outside? Does it not effect anything? Im asking because all examples and tutorials I saw had the model inside.",
    "785618": "Also, the reported AUC seems around 0.5. So its not working as it should?",
    "785622": "del",
    "785637": "abhishek  I don't think it's incorrect to do as it will be copied to each device separately.  Plus this will be the only way to get training correctly if there is some randomness while creating model. i.e. if starting weights differ in each instance then the distributed training will not work as far as I understand. the low score is because of high effective learning rate (8x) as compared to your notebook.",
    "785647": "mgornergoogle I was looking into this a little bit and here's what I found.\n\nSo in TF, you have `TPUStrategy` which is said to be similar to `MirroredStrategy` (for multi-GPU) but for TPUs. `MirroredStrategy` makes a replica of the model on each GPU. It would be fair to assume that `TPUStrategy` is doing the same for the TPU cores. Isn't this the same method PyTorch XLA multiprocessing using? So would TF also have problems with BERT multilingual models?\n\nThe only way I can think of right now for TF to work is if TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?",
    "785652": "&gt; the low score is because of high effective learning rate (8x) as compared to your notebook.\n\nEasy way to verify, just set it correctly then because IMHO it's not correct as per XLA docs etc :)",
    "785662": "dhananjay3 based on your suggested modification, i wrote a 8 core version of my original kernel. you can find it here: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\n\nIt gives an AUC of 0.84 after first epoch on the validation set.\n\nHowever, the kernel throws an error (OOM) when it reaches second epoch.",
    "785679": "abhishek try reducing num_workers.\n@adityaecdrid the next version of my notebook fixes the learning rate and the result is as expected so no there is nothing inherently wrong in this method.",
    "785680": "For whatever reason it might work https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores?scriptVersionId=30798902",
    "785681": "instead i went for reduction of batch size and that seemed to work",
    "785688": "abhishek just to let you know you are currently using 64 as per core batch_size which gives effective batch size of 512. I don't see any point of having that big batch as this reduces num of steps per epoch which eventually increases total training time.",
    "785882": "Thanks for this to all involved. I have one question:\nIs the validation score calculated on the full valid data (8000 records)?\n\nWhen I print the length of the valid vector it has only length 1000 (divided by 8 forks). Doesn't this need to be collected somewhere and then combined?\n\nEdit: OK I think I understand, it just prints for master node. Is there a way to collect across processes?",
    "785933": "Hey folks, just attaching the reply from XLA Devs regarding what Dhananjay did,\n\n&gt;It is not advised to create data structures outside the MP function target call stack.\nThat does not mean you have to create them directly inside the function, as long as they are created from any function which has the MP target function as parent in the call stack.\nOr simply, avoid creating source code where models are global (that is, they get created when the module is loaded, without any function to be called).\nWhen using the pytorch multiprocessing spawn start_method, any global data will have to be pickled, so that puts further restrictions on what you can create on global scope.\n&gt; While for models that might work, you cannot call any XLA function besides xmp.spawn() (the set is a bit wider, but it is better to stay on the safe side) outside of the MP target function call stack.\n&gt; Same issue with pickling applies to datasets.\nOn top of that, for multicore training you will need distributed samplers, which need the core ordinal, which is only available once you are within the MP target function call stack.",
    "786116": "&gt; TF is directly creating the replicas on the TPU, while PyTorch XLA makes the replicas on the VM, then distributes them to the TPU. Is this what is happening?\n\nI'm trying to get a definitive answer from the PyTorch-XLA devs but I think that is what is happening.",
    "786232": "Yes. Thats the case afaik. :)",
    "787219": "Hi Abhishek,\n\nThis is NLPclassification problem, Target variable is either 0 or 1 ( Comments is not toxic or toxic), But many many cases target variable is negative,  How we can interpret the results in terms of toxic or not toxic\n\nThanks",
    "787677": "👍",
    "788742": "I got some answers from the PyTorch/XLA team. Here are the two possible approaches (in pseudocode)\n\n1) With the model instantiated in the mp_fn:\n\n```\ndef _mp_fn():\n   model = instantiate_my_model()\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()\n\n   def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...\n\n   device = xm.xla_device()\n   model.to(device)\n\n   for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)\n\nxmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```\n\n2) With the model instantiated outside of the mp_fn:\n```\n# model instantiated in global scope\nmodel = instantiate_my_model()\n\ndef _mp_fn():\n   data_loader = instantiate_my_dataset_loader()\n   optimizer = instantiate_my_optimizer()\n\n   def train_step(model, optimizer, data)\n      ...\n      xm.optimizer_step(optimizer)\n      ...\n\n   device = xm.xla_device()\n   model.to(device)\n\n   for epoch in range(EPOCHS):\n      device_data = pl.ParallelLoader(data_loader, [device])\n      train_step(model, optimizer, device_data)\n\nxmp.spawn(_mp_fn, nprocs=8, start_method='fork')\n```\n\nApproach #1 instantiates the model 8 times as the process is forked. With a large model this will run out of memory.\nApproach #2 instantiates the model only once, in global scope. When the process is forked, memory pages corresponding to the model will be marked COW (Copy on Write) and since the model is copied to TPU by `model.to(device)`, the version on the VM is constant and will not be copied.\n\nThe only limitation to remember is that all XLA functions must be called from the mp_fn, which is the case here.",
    "788747": "Thanks for the clarification! I assume the dataloader cannot be defined outside of the training loop function because it is dependent on the xla functions to define the appropriate distributed sampler?",
    "788845": "The way @abhishek did it [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores) works. The data he loads in the global context is just pandas dataframes. The data loader objects acting on them are defined in the mp_fn. I believe this also prevents the data from being replicated in memory 8 times.",
    "828876": "Why is it that the TPU idle is around 50%. I was trying to run @abhishek 's notebook. Will there be a difference if I go with TF rather than using pytorch xla?"
  },
  "source": "meta"
}