{
  "id": 165910,
  "title": "bad result when training with 8 TPU cores - OK on single core",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165910",
  "author_name": "",
  "post_date": "2020-07-11T09:36:40.588325800Z",
  "votes": 4,
  "comment_count": 26,
  "views": 0,
  "content": "<p>In the last days I was working to make PyTorch work with TPU on Google Colab then on Kaggle. I have code which works on GPU and on single core TPU with similar results (loss and score). However, when I run it with 8 cores the results is much worse.</p>\n\n<p>Example - training 64x64 images for 3 epochs (I tried multiple times on both Colab and Kaggle and tried also more epochs):</p>\n\n<ul>\n<li>GPU - valid losses: 0.008, 0.005, 0.006</li>\n<li>TPU single core - valid losses: 0.007, 0.006, 0.005</li>\n<li>TPU 8 cores - valid losses: 0.047, 0.006, 0.010</li>\n</ul>\n\n<p>Train loss is similar in each case, so looks like TPU 8 cores is overfitting very quickly. \nI was thinking the problem is score calculation on data subset, that's why I presented here losses instead scores.</p>\n\n<p>for 8 cores I start this way (models is 1-element list in this case):</p>\n\n<p><code>\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=8,\n          start_method='fork')\n</code></p>\n\n<p>for 1 core I start this way:</p>\n\n<p><code>\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=1,\n          start_method='fork')\n</code></p>\n\n<p>I use following sampler:</p>\n\n<p><code>\n        print(\"create train sampler xrt_world_size {} oridinal {} \".format(xm.xrt_world_size(), xm.get_ordinal()))\n        train_sampler = torch.utils.data.distributed.DistributedSampler(\n            train_generator,\n            num_replicas=xm.xrt_world_size(),\n            rank=xm.get_ordinal(),\n            shuffle=True\n        )\n</code></p>\n\n<p>and I see xrt_world_size is 8 and ordinals are changing\n(I was thinking maybe each core uses same subset of data - it would explain overfitting)</p>\n\n<p>I use following loss:</p>\n\n<p>```\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, use_gpu, use_tpu, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        if (use_gpu or use_tpu):\n            self.alpha = torch.tensor([alpha, 1-alpha]).to(device)\n        else:\n            self.alpha = torch.tensor([alpha, 1-alpha])\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>\n\n<p>following loader:</p>\n\n<p><code>\n        para_loader = pl.ParallelLoader(data_loader, [device])\n</code></p>\n\n<p>and that's my train/valid loop:</p>\n\n<p>```\n            if (train_mode):\n                optimizer.zero_grad()</p>\n\n<pre><code>        output = model(x, meta)\n        loss = criterrion(output, target.unsqueeze(1).type_as(output))\n\n        if (train_mode):                \n            if (use_tpu):\n                loss.backward()                \n                xm.optimizer_step(optimizer)\n                pass\n            else:\n                loss.backward()                \n                optimizer.step()                          \n\n        y_preds[i].append(output.cpu().detach().numpy())\n\n        if (use_tpu):\n            reduced_loss = xm.mesh_reduce('loss_reduce', loss, reduce_fn)       \n            train_losses[i].update(reduced_loss.item(), data_loader.batch_size)\n        else:\n            train_losses[i].update(loss.item(), data_loader.batch_size)\n</code></pre>\n\n<p>```</p>\n\n<p>I tried to analyse available kernels but can't find many examples of PyTorch + TPU for 8 cores except this one:\n<a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu</a>\nby <a href=\"/abhishek\">@abhishek</a>\nand this one:\n<a href=\"https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch\">https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch</a>\nby <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> </p>\n\n<p>and I can't spot what I am doing incorrectly</p>\n\n<p>I suspect the hint is here:</p>\n\n<p>```\nprint(\"ordinal {} train_loss {} eval loss{}\".format(xm.get_ordinal(), train_log[\"losses\"][i], eval_log[\"losses\"][i]))</p>\n\n<p>```</p>\n\n<p>```\nordinal 4 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 7 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 6 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 3 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 5 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 1 train_loss 0.013886405224911868 eval loss0.0475122332572937</p>\n\n<p>```</p>\n\n<p>as you can see for each ordinal I have exactly same train_loss and eval_loss which should not happen, right? But sampler should create different subsets for each core so losses should be different?</p>\n\n<p>if you have any tips please share,</p>",
  "messages": [
    {
      "id": "924163",
      "postDate": "07/11/2020 09:36:40",
      "content": "<p>In the last days I was working to make PyTorch work with TPU on Google Colab then on Kaggle. I have code which works on GPU and on single core TPU with similar results (loss and score). However, when I run it with 8 cores the results is much worse.</p>\n\n<p>Example - training 64x64 images for 3 epochs (I tried multiple times on both Colab and Kaggle and tried also more epochs):</p>\n\n<ul>\n<li>GPU - valid losses: 0.008, 0.005, 0.006</li>\n<li>TPU single core - valid losses: 0.007, 0.006, 0.005</li>\n<li>TPU 8 cores - valid losses: 0.047, 0.006, 0.010</li>\n</ul>\n\n<p>Train loss is similar in each case, so looks like TPU 8 cores is overfitting very quickly. \nI was thinking the problem is score calculation on data subset, that's why I presented here losses instead scores.</p>\n\n<p>for 8 cores I start this way (models is 1-element list in this case):</p>\n\n<p><code>\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=8,\n          start_method='fork')\n</code></p>\n\n<p>for 1 core I start this way:</p>\n\n<p><code>\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=1,\n          start_method='fork')\n</code></p>\n\n<p>I use following sampler:</p>\n\n<p><code>\n        print(\"create train sampler xrt_world_size {} oridinal {} \".format(xm.xrt_world_size(), xm.get_ordinal()))\n        train_sampler = torch.utils.data.distributed.DistributedSampler(\n            train_generator,\n            num_replicas=xm.xrt_world_size(),\n            rank=xm.get_ordinal(),\n            shuffle=True\n        )\n</code></p>\n\n<p>and I see xrt_world_size is 8 and ordinals are changing\n(I was thinking maybe each core uses same subset of data - it would explain overfitting)</p>\n\n<p>I use following loss:</p>\n\n<p>```\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, use_gpu, use_tpu, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        if (use_gpu or use_tpu):\n            self.alpha = torch.tensor([alpha, 1-alpha]).to(device)\n        else:\n            self.alpha = torch.tensor([alpha, 1-alpha])\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>\n\n<p>following loader:</p>\n\n<p><code>\n        para_loader = pl.ParallelLoader(data_loader, [device])\n</code></p>\n\n<p>and that's my train/valid loop:</p>\n\n<p>```\n            if (train_mode):\n                optimizer.zero_grad()</p>\n\n<pre><code>        output = model(x, meta)\n        loss = criterrion(output, target.unsqueeze(1).type_as(output))\n\n        if (train_mode):                \n            if (use_tpu):\n                loss.backward()                \n                xm.optimizer_step(optimizer)\n                pass\n            else:\n                loss.backward()                \n                optimizer.step()                          \n\n        y_preds[i].append(output.cpu().detach().numpy())\n\n        if (use_tpu):\n            reduced_loss = xm.mesh_reduce('loss_reduce', loss, reduce_fn)       \n            train_losses[i].update(reduced_loss.item(), data_loader.batch_size)\n        else:\n            train_losses[i].update(loss.item(), data_loader.batch_size)\n</code></pre>\n\n<p>```</p>\n\n<p>I tried to analyse available kernels but can't find many examples of PyTorch + TPU for 8 cores except this one:\n<a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu</a>\nby <a href=\"/abhishek\">@abhishek</a>\nand this one:\n<a href=\"https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch\">https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch</a>\nby <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> </p>\n\n<p>and I can't spot what I am doing incorrectly</p>\n\n<p>I suspect the hint is here:</p>\n\n<p>```\nprint(\"ordinal {} train_loss {} eval loss{}\".format(xm.get_ordinal(), train_log[\"losses\"][i], eval_log[\"losses\"][i]))</p>\n\n<p>```</p>\n\n<p>```\nordinal 4 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 7 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 6 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 3 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 5 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 1 train_loss 0.013886405224911868 eval loss0.0475122332572937</p>\n\n<p>```</p>\n\n<p>as you can see for each ordinal I have exactly same train_loss and eval_loss which should not happen, right? But sampler should create different subsets for each core so losses should be different?</p>\n\n<p>if you have any tips please share,</p>",
      "rawMarkdown": "In the last days I was working to make PyTorch work with TPU on Google Colab then on Kaggle. I have code which works on GPU and on single core TPU with similar results (loss and score). However, when I run it with 8 cores the results is much worse.\n\nExample - training 64x64 images for 3 epochs (I tried multiple times on both Colab and Kaggle and tried also more epochs):\n\n- GPU - valid losses: 0.008, 0.005, 0.006\n- TPU single core - valid losses: 0.007, 0.006, 0.005\n- TPU 8 cores - valid losses: 0.047, 0.006, 0.010\n\nTrain loss is similar in each case, so looks like TPU 8 cores is overfitting very quickly. \nI was thinking the problem is score calculation on data subset, that's why I presented here losses instead scores.\n\n\nfor 8 cores I start this way (models is 1-element list in this case):\n\n```\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=8,\n          start_method='fork')\n```\n\nfor 1 core I start this way:\n\n```\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=1,\n          start_method='fork')\n```\n\nI use following sampler:\n\n```\n        print(\"create train sampler xrt_world_size {} oridinal {} \".format(xm.xrt_world_size(), xm.get_ordinal()))\n        train_sampler = torch.utils.data.distributed.DistributedSampler(\n            train_generator,\n            num_replicas=xm.xrt_world_size(),\n            rank=xm.get_ordinal(),\n            shuffle=True\n        )\n```\n\nand I see xrt_world_size is 8 and ordinals are changing\n(I was thinking maybe each core uses same subset of data - it would explain overfitting)\n\nI use following loss:\n\n```\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, use_gpu, use_tpu, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        if (use_gpu or use_tpu):\n            self.alpha = torch.tensor([alpha, 1-alpha]).to(device)\n        else:\n            self.alpha = torch.tensor([alpha, 1-alpha])\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```\n\nfollowing loader:\n\n```\n        para_loader = pl.ParallelLoader(data_loader, [device])\n```\n\nand that's my train/valid loop:\n\n```\n            if (train_mode):\n                optimizer.zero_grad()\n                \n            output = model(x, meta)\n            loss = criterrion(output, target.unsqueeze(1).type_as(output))\n\n            if (train_mode):                \n                if (use_tpu):\n                    loss.backward()                \n                    xm.optimizer_step(optimizer)\n                    pass\n                else:\n                    loss.backward()                \n                    optimizer.step()                          \n            \n            y_preds[i].append(output.cpu().detach().numpy())\n            \n            if (use_tpu):\n                reduced_loss = xm.mesh_reduce('loss_reduce', loss, reduce_fn)       \n                train_losses[i].update(reduced_loss.item(), data_loader.batch_size)\n            else:\n                train_losses[i].update(loss.item(), data_loader.batch_size)\n\n```\n\n\nI tried to analyse available kernels but can't find many examples of PyTorch + TPU for 8 cores except this one:\nhttps://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\nby @abhishek\nand this one:\nhttps://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch\nby @gopidurgaprasad \n\nand I can't spot what I am doing incorrectly\n\nI suspect the hint is here:\n\n```\nprint(\"ordinal {} train_loss {} eval loss{}\".format(xm.get_ordinal(), train_log[\"losses\"][i], eval_log[\"losses\"][i]))\n\n```\n\n```\nordinal 4 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 7 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 6 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 3 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 5 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 1 train_loss 0.013886405224911868 eval loss0.0475122332572937\n\n```\n\n\nas you can see for each ordinal I have exactly same train_loss and eval_loss which should not happen, right? But sampler should create different subsets for each core so losses should be different?\n\nif you have any tips please share,",
      "votes": null
    },
    {
      "id": "924227",
      "postDate": "07/11/2020 10:42:21",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Hi,</p>\n\n<p>for 8 core 8 models, you don't need sampler and also give the same learning rate, train_dataset and train_loader must be inside the run() function.</p>",
      "rawMarkdown": "jacekpoplawski Hi,\n\nfor 8 core 8 models, you don't need sampler and also give the same learning rate, train_dataset and train_loader must be inside the run() function.",
      "votes": null
    },
    {
      "id": "924233",
      "postDate": "07/11/2020 10:45:12",
      "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I have one model, dataset and loader are inside run, model is created before </p>",
      "rawMarkdown": "gopidurgaprasad I have one model, dataset and loader are inside run, model is created before",
      "votes": null
    },
    {
      "id": "924244",
      "postDate": "07/11/2020 10:48:46",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> I am not getting clearly </p>",
      "rawMarkdown": "jacekpoplawski I am not getting clearly",
      "votes": null
    },
    {
      "id": "924256",
      "postDate": "07/11/2020 10:53:48",
      "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I want to train one model with 8 TPU cores, when I train on single core I have same result as with GPU - validation loss is OK, but when I use 8 cores I see validation loss is big and I also see each core has exactly same loss, which suggests something is wrong</p>",
      "rawMarkdown": "gopidurgaprasad I want to train one model with 8 TPU cores, when I train on single core I have same result as with GPU - validation loss is OK, but when I use 8 cores I see validation loss is big and I also see each core has exactly same loss, which suggests something is wrong",
      "votes": null
    },
    {
      "id": "924260",
      "postDate": "07/11/2020 10:58:51",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> probably some suggestions,\n- use <code>xm.optimizer_step(optimizer, barrier=True)</code> make sure <code>barrier=True</code>\n- dont use <code>xm.master_print()</code></p>",
      "rawMarkdown": "jacekpoplawski probably some suggestions,\n- use `xm.optimizer_step(optimizer, barrier=True)` make sure `barrier=True`\n- dont use `xm.master_print()`",
      "votes": null
    },
    {
      "id": "924272",
      "postDate": "07/11/2020 11:02:46",
      "content": "<p>hm, I am not using barrier=True at all\naccording to documentation\n\"xm.optimizer_step(optimizer) no longer needs a barrier. ParallelLoader automatically creates an XLA barrier that evalutes the graph.\"\n<a href=\"https://pytorch.org/xla/release/1.5/index.html\">https://pytorch.org/xla/release/1.5/index.html</a></p>",
      "rawMarkdown": "hm, I am not using barrier=True at all\naccording to documentation\n\"xm.optimizer_step(optimizer) no longer needs a barrier. ParallelLoader automatically creates an XLA barrier that evalutes the graph.\"\nhttps://pytorch.org/xla/release/1.5/index.html",
      "votes": null
    },
    {
      "id": "924288",
      "postDate": "07/11/2020 11:07:05",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  yes, but we are training different models in different cores, barrier makes it separate.\nI think so.</p>",
      "rawMarkdown": "jacekpoplawski  yes, but we are training different models in different cores, barrier makes it separate.\nI think so.",
      "votes": null
    },
    {
      "id": "924291",
      "postDate": "07/11/2020 11:08:11",
      "content": "<p>in my case I want to train single (same) model on each core</p>",
      "rawMarkdown": "in my case I want to train single (same) model on each core",
      "votes": null
    },
    {
      "id": "924299",
      "postDate": "07/11/2020 11:10:40",
      "content": "<p>for example 1 model on 8 different fold right.\nin that case 1 model thing like 8 different models.</p>",
      "rawMarkdown": "for example 1 model on 8 different fold right.\nin that case 1 model thing like 8 different models.",
      "votes": null
    },
    {
      "id": "924307",
      "postDate": "07/11/2020 11:15:12",
      "content": "<p>no, I use only one model:</p>\n\n<p>first I create it (outside the _mp_fn function):</p>\n\n<p><code>model = xmp.MpModelWrapper(Net(config[\"net_{}\".format(net)], nfeatures, log))</code></p>\n\n<p>then inside _mp_fun function I do following:</p>\n\n<p>```\nif (use_gpu or use_tpu):\n    model = model.to(device)</p>\n\n<p>optimizer = optim.Adam(model.parameters(), lr = lr)</p>\n\n<p>scheduler = ReduceLROnPlateau(optimizer, mode=\"min\", patience=scheduler_patience, factor=scheduler_factor, eps=scheduler_eps, verbose=True)                                  </p>\n\n<p>```</p>",
      "rawMarkdown": "no, I use only one model:\n\nfirst I create it (outside the _mp_fn function):\n\n`model = xmp.MpModelWrapper(Net(config[\"net_{}\".format(net)], nfeatures, log))`\n\nthen inside _mp_fun function I do following:\n\n```\nif (use_gpu or use_tpu):\n    model = model.to(device)\n            \noptimizer = optim.Adam(model.parameters(), lr = lr)\n\nscheduler = ReduceLROnPlateau(optimizer, mode=\"min\", patience=scheduler_patience, factor=scheduler_factor, eps=scheduler_eps, verbose=True)                                  \n\n```",
      "votes": null
    },
    {
      "id": "924314",
      "postDate": "07/11/2020 11:17:31",
      "content": "<p>at the end, you want to train one model on 8 folds in using 8 cores right.</p>",
      "rawMarkdown": "at the end, you want to train one model on 8 folds in using 8 cores right.",
      "votes": null
    },
    {
      "id": "924315",
      "postDate": "07/11/2020 11:18:36",
      "content": "<p>No, I use single fold right now and I want only one model with multicore TPU.</p>",
      "rawMarkdown": "No, I use single fold right now and I want only one model with multicore TPU.",
      "votes": null
    },
    {
      "id": "924323",
      "postDate": "07/11/2020 11:23:26",
      "content": "<p>got it now, check my notebook <a href=\"https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch?scriptVersionId=35516168\">version 2 </a></p>",
      "rawMarkdown": "got it now, check my notebook [version 2 ](https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch?scriptVersionId=35516168)",
      "votes": null
    },
    {
      "id": "924329",
      "postDate": "07/11/2020 11:28:44",
      "content": "<p>Thank you, I will analyse it</p>\n\n<p>the first thing I noted is \n<code>lr=args.learning_rate * xm.xrt_world_size()</code>\nI have seen it also in another kernel but when I tried to multiply my lr by 8 the training was even worse</p>",
      "rawMarkdown": "Thank you, I will analyse it\n\nthe first thing I noted is \n`lr=args.learning_rate * xm.xrt_world_size()`\nI have seen it also in another kernel but when I tried to multiply my lr by 8 the training was even worse",
      "votes": null
    },
    {
      "id": "924344",
      "postDate": "07/11/2020 11:38:44",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> ,\nif your using 1 TPU core, then xm.xrt_world_size() = 1\nif your using 8 TPU core, then xm.xrt_world_size() = 8</p>\n\n<p>another trick is just to multiply with 0.5\n<code>lr = args.learning_rate * 0.5</code></p>",
      "rawMarkdown": "jacekpoplawski ,\nif your using 1 TPU core, then xm.xrt_world_size() = 1\nif your using 8 TPU core, then xm.xrt_world_size() = 8\n\nanother trick is just to multiply with 0.5\n`lr = args.learning_rate * 0.5`",
      "votes": null
    },
    {
      "id": "924345",
      "postDate": "07/11/2020 11:40:21",
      "content": "<p>I did following experiment - I print sum of all targets in the batch for each core, </p>\n\n<p><code>print(\"device {} ordinal {} all_true sum: {}\".format(device, xm.get_ordinal(),sum(all_true)))</code></p>\n\n<p>the result is:\n<code>\ndevice xla:0 ordinal 4 all_true sum: [59.]\ndevice xla:0 ordinal 5 all_true sum: [62.]\ndevice xla:0 ordinal 3 all_true sum: [66.]\ndevice xla:0 ordinal 6 all_true sum: [54.]\ndevice xla:0 ordinal 2 all_true sum: [48.]\ndevice xla:0 ordinal 1 all_true sum: [60.]\ndevice xla:0 ordinal 7 all_true sum: [59.]\n</code></p>\n\n<p>so sampler works correctly, each core has different data to train</p>",
      "rawMarkdown": "I did following experiment - I print sum of all targets in the batch for each core, \n\n`    print(\"device {} ordinal {} all_true sum: {}\".format(device, xm.get_ordinal(),sum(all_true)))`\n\nthe result is:\n```\ndevice xla:0 ordinal 4 all_true sum: [59.]\ndevice xla:0 ordinal 5 all_true sum: [62.]\ndevice xla:0 ordinal 3 all_true sum: [66.]\ndevice xla:0 ordinal 6 all_true sum: [54.]\ndevice xla:0 ordinal 2 all_true sum: [48.]\ndevice xla:0 ordinal 1 all_true sum: [60.]\ndevice xla:0 ordinal 7 all_true sum: [59.]\n```\n\nso sampler works correctly, each core has different data to train",
      "votes": null
    },
    {
      "id": "924349",
      "postDate": "07/11/2020 11:42:51",
      "content": "<p>yes, thats cool</p>",
      "rawMarkdown": "yes, thats cool",
      "votes": null
    },
    {
      "id": "924353",
      "postDate": "07/11/2020 11:47:59",
      "content": "<p>wait I am lost, why *8 can be replaced by *0.5?</p>",
      "rawMarkdown": "wait I am lost, why *8 can be replaced by *0.5?",
      "votes": null
    },
    {
      "id": "924376",
      "postDate": "07/11/2020 12:00:14",
      "content": "<p>Yes, instead of multiply with 8 cores, just multiply with 0.5, one of the trick used in previous NLP competitions </p>",
      "rawMarkdown": "Yes, instead of multiply with 8 cores, just multiply with 0.5, one of the trick used in previous NLP competitions",
      "votes": null
    },
    {
      "id": "924381",
      "postDate": "07/11/2020 12:02:28",
      "content": "<p>I tested lr*8 and the result is:</p>\n\n<p>```\nepoch 0\nmodel 0 train loss: 0.16237319139763712, eval loss: 421542832.0\nmodel 0 train score: 0.5492101409667254, eval score: 0.6900212314225054</p>\n\n<p>epoch 1\nmodel 0 train loss: 0.02815218740142882, eval loss: 371038.859375\nmodel 0 train score: 0.5612284146810536, eval score: 0.4214861995753715</p>\n\n<p>epoch 2\nmodel 0 train loss: 0.015614275098778307, eval loss: 1599.6704711914062\nmodel 0 train score: 0.5566262458583855, eval score: 0.549723991507431\n```</p>",
      "rawMarkdown": "I tested lr*8 and the result is:\n\n```\nepoch 0\nmodel 0 train loss: 0.16237319139763712, eval loss: 421542832.0\nmodel 0 train score: 0.5492101409667254, eval score: 0.6900212314225054\n\nepoch 1\nmodel 0 train loss: 0.02815218740142882, eval loss: 371038.859375\nmodel 0 train score: 0.5612284146810536, eval score: 0.4214861995753715\n\nepoch 2\nmodel 0 train loss: 0.015614275098778307, eval loss: 1599.6704711914062\nmodel 0 train score: 0.5566262458583855, eval score: 0.549723991507431\n```",
      "votes": null
    },
    {
      "id": "924478",
      "postDate": "07/11/2020 13:01:00",
      "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> looks like decreasing lr helps, but this is counterintuitive, maybe the formula to multiplicate per number of cores is just wrong for current implementation</p>",
      "rawMarkdown": "gopidurgaprasad looks like decreasing lr helps, but this is counterintuitive, maybe the formula to multiplicate per number of cores is just wrong for current implementation",
      "votes": null
    },
    {
      "id": "932271",
      "postDate": "07/16/2020 22:24:00",
      "content": "<p>hello again <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> </p>\n\n<p>could you explain \"dont use xm.master_print()\"?\njust found that my kaggle kernel crashing can be related to my usage of print</p>",
      "rawMarkdown": "hello again @gopidurgaprasad \n\ncould you explain \"dont use xm.master_print()\"?\njust found that my kaggle kernel crashing can be related to my usage of print",
      "votes": null
    },
    {
      "id": "932330",
      "postDate": "07/17/2020 01:12:19",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> xm.master_print() dose average the 8 core's results ruffle,\nIf it string printed one time, \nThe problem with xm.master_print() is if some core's don't respond in runtime other core's also waiting for that some fixed time then it's terminate, this is my understanding. I am not sure </p>",
      "rawMarkdown": "jacekpoplawski xm.master_print() dose average the 8 core's results ruffle,\nIf it string printed one time, \nThe problem with xm.master_print() is if some core's don't respond in runtime other core's also waiting for that some fixed time then it's terminate, this is my understanding. I am not sure",
      "votes": null
    },
    {
      "id": "933534",
      "postDate": "07/17/2020 18:55:36",
      "content": "<p>When you train on multiple cores, you usually want to increase the batch size to take advantage of the matmul hardware in each TPU core. But with bigger batches, there is a speedup only if each batch also does \"more training\". That is why you also want to increase the learning rate. Scaling the learning rate by the nb of cores is just an initial guidance though. The ideal LR schedule is usually somewhere close to that but not exactly. Since you are using a new batch size, you need to re-tune the learning rate.</p>",
      "rawMarkdown": "When you train on multiple cores, you usually want to increase the batch size to take advantage of the matmul hardware in each TPU core. But with bigger batches, there is a speedup only if each batch also does \"more training\". That is why you also want to increase the learning rate. Scaling the learning rate by the nb of cores is just an initial guidance though. The ideal LR schedule is usually somewhere close to that but not exactly. Since you are using a new batch size, you need to re-tune the learning rate.",
      "votes": null
    },
    {
      "id": "933553",
      "postDate": "07/17/2020 19:15:32",
      "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I have great reply from pytorch/xla <a href=\"https://github.com/pytorch/xla/issues/2358\">https://github.com/pytorch/xla/issues/2358</a></p>",
      "rawMarkdown": "gopidurgaprasad I have great reply from pytorch/xla https://github.com/pytorch/xla/issues/2358",
      "votes": null
    },
    {
      "id": "934067",
      "postDate": "07/18/2020 08:02:31",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> \nThanks, that's really cool</p>",
      "rawMarkdown": "jacekpoplawski \nThanks, that's really cool",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 924227,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "07/11/2020 10:42:21",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Hi,</p>\n\n<p>for 8 core 8 models, you don't need sampler and also give the same learning rate, train_dataset and train_loader must be inside the run() function.</p>",
      "votes": null,
      "replies": [
        {
          "id": 924233,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 10:45:12",
          "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I have one model, dataset and loader are inside run, model is created before </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924244,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 10:48:46",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> I am not getting clearly </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924256,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 10:53:48",
          "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I want to train one model with 8 TPU cores, when I train on single core I have same result as with GPU - validation loss is OK, but when I use 8 cores I see validation loss is big and I also see each core has exactly same loss, which suggests something is wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924260,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 10:58:51",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> probably some suggestions,\n- use <code>xm.optimizer_step(optimizer, barrier=True)</code> make sure <code>barrier=True</code>\n- dont use <code>xm.master_print()</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924272,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:02:46",
          "content": "<p>hm, I am not using barrier=True at all\naccording to documentation\n\"xm.optimizer_step(optimizer) no longer needs a barrier. ParallelLoader automatically creates an XLA barrier that evalutes the graph.\"\n<a href=\"https://pytorch.org/xla/release/1.5/index.html\">https://pytorch.org/xla/release/1.5/index.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924288,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:07:05",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a>  yes, but we are training different models in different cores, barrier makes it separate.\nI think so.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924291,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:08:11",
          "content": "<p>in my case I want to train single (same) model on each core</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924299,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:10:40",
          "content": "<p>for example 1 model on 8 different fold right.\nin that case 1 model thing like 8 different models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924307,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:15:12",
          "content": "<p>no, I use only one model:</p>\n\n<p>first I create it (outside the _mp_fn function):</p>\n\n<p><code>model = xmp.MpModelWrapper(Net(config[\"net_{}\".format(net)], nfeatures, log))</code></p>\n\n<p>then inside _mp_fun function I do following:</p>\n\n<p>```\nif (use_gpu or use_tpu):\n    model = model.to(device)</p>\n\n<p>optimizer = optim.Adam(model.parameters(), lr = lr)</p>\n\n<p>scheduler = ReduceLROnPlateau(optimizer, mode=\"min\", patience=scheduler_patience, factor=scheduler_factor, eps=scheduler_eps, verbose=True)                                  </p>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924314,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:17:31",
          "content": "<p>at the end, you want to train one model on 8 folds in using 8 cores right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924315,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:18:36",
          "content": "<p>No, I use single fold right now and I want only one model with multicore TPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924323,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:23:26",
          "content": "<p>got it now, check my notebook <a href=\"https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch?scriptVersionId=35516168\">version 2 </a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924329,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:28:44",
          "content": "<p>Thank you, I will analyse it</p>\n\n<p>the first thing I noted is \n<code>lr=args.learning_rate * xm.xrt_world_size()</code>\nI have seen it also in another kernel but when I tried to multiply my lr by 8 the training was even worse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924344,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:38:44",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> ,\nif your using 1 TPU core, then xm.xrt_world_size() = 1\nif your using 8 TPU core, then xm.xrt_world_size() = 8</p>\n\n<p>another trick is just to multiply with 0.5\n<code>lr = args.learning_rate * 0.5</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924353,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 11:47:59",
          "content": "<p>wait I am lost, why *8 can be replaced by *0.5?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924376,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 12:00:14",
          "content": "<p>Yes, instead of multiply with 8 cores, just multiply with 0.5, one of the trick used in previous NLP competitions </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924381,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 12:02:28",
          "content": "<p>I tested lr*8 and the result is:</p>\n\n<p>```\nepoch 0\nmodel 0 train loss: 0.16237319139763712, eval loss: 421542832.0\nmodel 0 train score: 0.5492101409667254, eval score: 0.6900212314225054</p>\n\n<p>epoch 1\nmodel 0 train loss: 0.02815218740142882, eval loss: 371038.859375\nmodel 0 train score: 0.5612284146810536, eval score: 0.4214861995753715</p>\n\n<p>epoch 2\nmodel 0 train loss: 0.015614275098778307, eval loss: 1599.6704711914062\nmodel 0 train score: 0.5566262458583855, eval score: 0.549723991507431\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924478,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/11/2020 13:01:00",
          "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> looks like decreasing lr helps, but this is counterintuitive, maybe the formula to multiplicate per number of cores is just wrong for current implementation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932271,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/16/2020 22:24:00",
          "content": "<p>hello again <a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> </p>\n\n<p>could you explain \"dont use xm.master_print()\"?\njust found that my kaggle kernel crashing can be related to my usage of print</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932330,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/17/2020 01:12:19",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> xm.master_print() dose average the 8 core's results ruffle,\nIf it string printed one time, \nThe problem with xm.master_print() is if some core's don't respond in runtime other core's also waiting for that some fixed time then it's terminate, this is my understanding. I am not sure </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933534,
          "author_name": "martingorner",
          "author_url": "",
          "post_date": "07/17/2020 18:55:36",
          "content": "<p>When you train on multiple cores, you usually want to increase the batch size to take advantage of the matmul hardware in each TPU core. But with bigger batches, there is a speedup only if each batch also does \"more training\". That is why you also want to increase the learning rate. Scaling the learning rate by the nb of cores is just an initial guidance though. The ideal LR schedule is usually somewhere close to that but not exactly. Since you are using a new batch size, you need to re-tune the learning rate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933553,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/17/2020 19:15:32",
          "content": "<p><a href=\"/gopidurgaprasad\">@gopidurgaprasad</a> I have great reply from pytorch/xla <a href=\"https://github.com/pytorch/xla/issues/2358\">https://github.com/pytorch/xla/issues/2358</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934067,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/18/2020 08:02:31",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> \nThanks, that's really cool</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 924345,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/11/2020 11:40:21",
      "content": "<p>I did following experiment - I print sum of all targets in the batch for each core, </p>\n\n<p><code>print(\"device {} ordinal {} all_true sum: {}\".format(device, xm.get_ordinal(),sum(all_true)))</code></p>\n\n<p>the result is:\n<code>\ndevice xla:0 ordinal 4 all_true sum: [59.]\ndevice xla:0 ordinal 5 all_true sum: [62.]\ndevice xla:0 ordinal 3 all_true sum: [66.]\ndevice xla:0 ordinal 6 all_true sum: [54.]\ndevice xla:0 ordinal 2 all_true sum: [48.]\ndevice xla:0 ordinal 1 all_true sum: [60.]\ndevice xla:0 ordinal 7 all_true sum: [59.]\n</code></p>\n\n<p>so sampler works correctly, each core has different data to train</p>",
      "votes": null,
      "replies": [
        {
          "id": 924349,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "07/11/2020 11:42:51",
          "content": "<p>yes, thats cool</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "924163": "In the last days I was working to make PyTorch work with TPU on Google Colab then on Kaggle. I have code which works on GPU and on single core TPU with similar results (loss and score). However, when I run it with 8 cores the results is much worse.\n\nExample - training 64x64 images for 3 epochs (I tried multiple times on both Colab and Kaggle and tried also more epochs):\n\n- GPU - valid losses: 0.008, 0.005, 0.006\n- TPU single core - valid losses: 0.007, 0.006, 0.005\n- TPU 8 cores - valid losses: 0.047, 0.006, 0.010\n\nTrain loss is similar in each case, so looks like TPU 8 cores is overfitting very quickly. \nI was thinking the problem is score calculation on data subset, that's why I presented here losses instead scores.\n\n\nfor 8 cores I start this way (models is 1-element list in this case):\n\n```\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=8,\n          start_method='fork')\n```\n\nfor 1 core I start this way:\n\n```\ndef _mp_fn(rank, config):\n  log = {}\n  run_system.run_system(config, models, log)\nxmp.spawn(_mp_fn, args=(config,), nprocs=1,\n          start_method='fork')\n```\n\nI use following sampler:\n\n```\n        print(\"create train sampler xrt_world_size {} oridinal {} \".format(xm.xrt_world_size(), xm.get_ordinal()))\n        train_sampler = torch.utils.data.distributed.DistributedSampler(\n            train_generator,\n            num_replicas=xm.xrt_world_size(),\n            rank=xm.get_ordinal(),\n            shuffle=True\n        )\n```\n\nand I see xrt_world_size is 8 and ordinals are changing\n(I was thinking maybe each core uses same subset of data - it would explain overfitting)\n\nI use following loss:\n\n```\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, use_gpu, use_tpu, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        if (use_gpu or use_tpu):\n            self.alpha = torch.tensor([alpha, 1-alpha]).to(device)\n        else:\n            self.alpha = torch.tensor([alpha, 1-alpha])\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```\n\nfollowing loader:\n\n```\n        para_loader = pl.ParallelLoader(data_loader, [device])\n```\n\nand that's my train/valid loop:\n\n```\n            if (train_mode):\n                optimizer.zero_grad()\n                \n            output = model(x, meta)\n            loss = criterrion(output, target.unsqueeze(1).type_as(output))\n\n            if (train_mode):                \n                if (use_tpu):\n                    loss.backward()                \n                    xm.optimizer_step(optimizer)\n                    pass\n                else:\n                    loss.backward()                \n                    optimizer.step()                          \n            \n            y_preds[i].append(output.cpu().detach().numpy())\n            \n            if (use_tpu):\n                reduced_loss = xm.mesh_reduce('loss_reduce', loss, reduce_fn)       \n                train_losses[i].update(reduced_loss.item(), data_loader.batch_size)\n            else:\n                train_losses[i].update(loss.item(), data_loader.batch_size)\n\n```\n\n\nI tried to analyse available kernels but can't find many examples of PyTorch + TPU for 8 cores except this one:\nhttps://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\nby @abhishek\nand this one:\nhttps://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch\nby @gopidurgaprasad \n\nand I can't spot what I am doing incorrectly\n\nI suspect the hint is here:\n\n```\nprint(\"ordinal {} train_loss {} eval loss{}\".format(xm.get_ordinal(), train_log[\"losses\"][i], eval_log[\"losses\"][i]))\n\n```\n\n```\nordinal 4 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 7 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 6 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 3 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 5 train_loss 0.013886405224911868 eval loss0.0475122332572937\nordinal 1 train_loss 0.013886405224911868 eval loss0.0475122332572937\n\n```\n\n\nas you can see for each ordinal I have exactly same train_loss and eval_loss which should not happen, right? But sampler should create different subsets for each core so losses should be different?\n\nif you have any tips please share,",
    "924227": "jacekpoplawski Hi,\n\nfor 8 core 8 models, you don't need sampler and also give the same learning rate, train_dataset and train_loader must be inside the run() function.",
    "924233": "gopidurgaprasad I have one model, dataset and loader are inside run, model is created before",
    "924244": "jacekpoplawski I am not getting clearly",
    "924256": "gopidurgaprasad I want to train one model with 8 TPU cores, when I train on single core I have same result as with GPU - validation loss is OK, but when I use 8 cores I see validation loss is big and I also see each core has exactly same loss, which suggests something is wrong",
    "924260": "jacekpoplawski probably some suggestions,\n- use `xm.optimizer_step(optimizer, barrier=True)` make sure `barrier=True`\n- dont use `xm.master_print()`",
    "924272": "hm, I am not using barrier=True at all\naccording to documentation\n\"xm.optimizer_step(optimizer) no longer needs a barrier. ParallelLoader automatically creates an XLA barrier that evalutes the graph.\"\nhttps://pytorch.org/xla/release/1.5/index.html",
    "924288": "jacekpoplawski  yes, but we are training different models in different cores, barrier makes it separate.\nI think so.",
    "924291": "in my case I want to train single (same) model on each core",
    "924299": "for example 1 model on 8 different fold right.\nin that case 1 model thing like 8 different models.",
    "924307": "no, I use only one model:\n\nfirst I create it (outside the _mp_fn function):\n\n`model = xmp.MpModelWrapper(Net(config[\"net_{}\".format(net)], nfeatures, log))`\n\nthen inside _mp_fun function I do following:\n\n```\nif (use_gpu or use_tpu):\n    model = model.to(device)\n            \noptimizer = optim.Adam(model.parameters(), lr = lr)\n\nscheduler = ReduceLROnPlateau(optimizer, mode=\"min\", patience=scheduler_patience, factor=scheduler_factor, eps=scheduler_eps, verbose=True)                                  \n\n```",
    "924314": "at the end, you want to train one model on 8 folds in using 8 cores right.",
    "924315": "No, I use single fold right now and I want only one model with multicore TPU.",
    "924323": "got it now, check my notebook [version 2 ](https://www.kaggle.com/gopidurgaprasad/siim-8-folds-with-8-tpu-cores-pytorch?scriptVersionId=35516168)",
    "924329": "Thank you, I will analyse it\n\nthe first thing I noted is \n`lr=args.learning_rate * xm.xrt_world_size()`\nI have seen it also in another kernel but when I tried to multiply my lr by 8 the training was even worse",
    "924344": "jacekpoplawski ,\nif your using 1 TPU core, then xm.xrt_world_size() = 1\nif your using 8 TPU core, then xm.xrt_world_size() = 8\n\nanother trick is just to multiply with 0.5\n`lr = args.learning_rate * 0.5`",
    "924345": "I did following experiment - I print sum of all targets in the batch for each core, \n\n`    print(\"device {} ordinal {} all_true sum: {}\".format(device, xm.get_ordinal(),sum(all_true)))`\n\nthe result is:\n```\ndevice xla:0 ordinal 4 all_true sum: [59.]\ndevice xla:0 ordinal 5 all_true sum: [62.]\ndevice xla:0 ordinal 3 all_true sum: [66.]\ndevice xla:0 ordinal 6 all_true sum: [54.]\ndevice xla:0 ordinal 2 all_true sum: [48.]\ndevice xla:0 ordinal 1 all_true sum: [60.]\ndevice xla:0 ordinal 7 all_true sum: [59.]\n```\n\nso sampler works correctly, each core has different data to train",
    "924349": "yes, thats cool",
    "924353": "wait I am lost, why *8 can be replaced by *0.5?",
    "924376": "Yes, instead of multiply with 8 cores, just multiply with 0.5, one of the trick used in previous NLP competitions",
    "924381": "I tested lr*8 and the result is:\n\n```\nepoch 0\nmodel 0 train loss: 0.16237319139763712, eval loss: 421542832.0\nmodel 0 train score: 0.5492101409667254, eval score: 0.6900212314225054\n\nepoch 1\nmodel 0 train loss: 0.02815218740142882, eval loss: 371038.859375\nmodel 0 train score: 0.5612284146810536, eval score: 0.4214861995753715\n\nepoch 2\nmodel 0 train loss: 0.015614275098778307, eval loss: 1599.6704711914062\nmodel 0 train score: 0.5566262458583855, eval score: 0.549723991507431\n```",
    "924478": "gopidurgaprasad looks like decreasing lr helps, but this is counterintuitive, maybe the formula to multiplicate per number of cores is just wrong for current implementation",
    "932271": "hello again @gopidurgaprasad \n\ncould you explain \"dont use xm.master_print()\"?\njust found that my kaggle kernel crashing can be related to my usage of print",
    "932330": "jacekpoplawski xm.master_print() dose average the 8 core's results ruffle,\nIf it string printed one time, \nThe problem with xm.master_print() is if some core's don't respond in runtime other core's also waiting for that some fixed time then it's terminate, this is my understanding. I am not sure",
    "933534": "When you train on multiple cores, you usually want to increase the batch size to take advantage of the matmul hardware in each TPU core. But with bigger batches, there is a speedup only if each batch also does \"more training\". That is why you also want to increase the learning rate. Scaling the learning rate by the nb of cores is just an initial guidance though. The ideal LR schedule is usually somewhere close to that but not exactly. Since you are using a new batch size, you need to re-tune the learning rate.",
    "933553": "gopidurgaprasad I have great reply from pytorch/xla https://github.com/pytorch/xla/issues/2358",
    "934067": "jacekpoplawski \nThanks, that's really cool"
  },
  "source": "meta"
}