{
  "id": 141636,
  "title": "[SOLVED] Completely Trapped with BUGS!",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141636",
  "author_name": "",
  "post_date": "2020-04-06T23:49:21.505572Z",
  "votes": 3,
  "comment_count": 27,
  "views": 0,
  "content": "<p>First I tried custom tensorflow to control loss and metrics and other layers of the model, which resulted in bugs, I did make the pytorch version of the same and now again this bug where\n<a href=\"https://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch\">https://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch</a></p>\n\n<p>Exception: process 0 terminated with signal SIGKILL</p>\n\n<p>I have no way to compete without the bug getting a fix either in tensorflow or torch. Guess its most likely a dead end for me atleast for now.</p>\n\n<p>I ask the kaggle community if they have encountered this error or have found any fix.</p>\n\n<p>Thank You!</p>",
  "messages": [
    {
      "id": "799952",
      "postDate": "04/06/2020 23:49:21",
      "content": "<p>First I tried custom tensorflow to control loss and metrics and other layers of the model, which resulted in bugs, I did make the pytorch version of the same and now again this bug where\n<a href=\"https://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch\">https://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch</a></p>\n\n<p>Exception: process 0 terminated with signal SIGKILL</p>\n\n<p>I have no way to compete without the bug getting a fix either in tensorflow or torch. Guess its most likely a dead end for me atleast for now.</p>\n\n<p>I ask the kaggle community if they have encountered this error or have found any fix.</p>\n\n<p>Thank You!</p>",
      "rawMarkdown": "First I tried custom tensorflow to control loss and metrics and other layers of the model, which resulted in bugs, I did make the pytorch version of the same and now again this bug where\nhttps://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch\n\nException: process 0 terminated with signal SIGKILL\n\nI have no way to compete without the bug getting a fix either in tensorflow or torch. Guess its most likely a dead end for me atleast for now.\n\nI ask the kaggle community if they have encountered this error or have found any fix.\n\nThank You!",
      "votes": null
    },
    {
      "id": "800129",
      "postDate": "04/07/2020 05:59:30",
      "content": "<p>Mine worked well with TPU on tf.keras.  I used pre-tokenized though.</p>",
      "rawMarkdown": "Mine worked well with TPU on tf.keras.  I used pre-tokenized though.",
      "votes": null
    },
    {
      "id": "800168",
      "postDate": "04/07/2020 06:31:12",
      "content": "<p>I am talking specifically about tf.keras custom training see this webpage for example what I mean by custom training pipeline. <a href=\"https://www.tensorflow.org/tutorials/distribute/custom_training\">https://www.tensorflow.org/tutorials/distribute/custom_training</a></p>",
      "rawMarkdown": "I am talking specifically about tf.keras custom training see this webpage for example what I mean by custom training pipeline. https://www.tensorflow.org/tutorials/distribute/custom_training",
      "votes": null
    },
    {
      "id": "800904",
      "postDate": "04/07/2020 21:09:48",
      "content": "<p><code>nprocs=1</code> - And it fixed this problem for me. But training is very slow (~approximately 2xP100). It is not full solution...</p>",
      "rawMarkdown": "`nprocs=1` - And it fixed this problem for me. But training is very slow (~approximately 2xP100). It is not full solution...",
      "votes": null
    },
    {
      "id": "800908",
      "postDate": "04/07/2020 21:19:06",
      "content": "<p>The problem did not get the fix for me unfortunately, btw, a question out of context, are you getting your current score with single model or what is your best with single model, you can skip the question if you want to, but I would like to hear to get an idea where I stand. Thank You!</p>",
      "rawMarkdown": "The problem did not get the fix for me unfortunately, btw, a question out of context, are you getting your current score with single model or what is your best with single model, you can skip the question if you want to, but I would like to hear to get an idea where I stand. Thank You!",
      "votes": null
    },
    {
      "id": "803015",
      "postDate": "04/10/2020 04:12:20",
      "content": "<p>I have been running into the same bug when working with PyTorch.</p>\n\n<p>I believe this has to do with XLA using up RAM. I constantly use up all my RAM, which causes the SIGKILL error. If you take a look at this: <a href=\"https://github.com/pytorch/xla/issues/1280\">https://github.com/pytorch/xla/issues/1280</a></p>\n\n<p>They talk about how each of the 8 TPU processes replicates the model, which takes up a considerable amount of RAM. In terms of large transformer models (I was using XLM-RoBERTa), there simply isn't enough RAM to support the model being replicated 8 times. I've seen smaller models such as BERT being successfully loaded as seen in this kernel: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores</a></p>\n\n<p>On the other hand, TensorFlow appears to support this better, as XLM-RoBERTa was successfully used in this kernel: <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>I'm going to switch over to TensorFlow, as Torch-XLA is just unusable with these memory issues.</p>",
      "rawMarkdown": "I have been running into the same bug when working with PyTorch.\n\nI believe this has to do with XLA using up RAM. I constantly use up all my RAM, which causes the SIGKILL error. If you take a look at this: https://github.com/pytorch/xla/issues/1280\n\nThey talk about how each of the 8 TPU processes replicates the model, which takes up a considerable amount of RAM. In terms of large transformer models (I was using XLM-RoBERTa), there simply isn't enough RAM to support the model being replicated 8 times. I've seen smaller models such as BERT being successfully loaded as seen in this kernel: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\n\nOn the other hand, TensorFlow appears to support this better, as XLM-RoBERTa was successfully used in this kernel: https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\nI'm going to switch over to TensorFlow, as Torch-XLA is just unusable with these memory issues.",
      "votes": null
    },
    {
      "id": "803025",
      "postDate": "04/10/2020 04:47:01",
      "content": "<p>Same Problem, I was using roberta, If I change that to bert, it works..., still has to use tensorflow only.</p>",
      "rawMarkdown": "Same Problem, I was using roberta, If I change that to bert, it works..., still has to use tensorflow only.",
      "votes": null
    },
    {
      "id": "803556",
      "postDate": "04/10/2020 16:39:52",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> posted a TPU-optimized custom training loop notebook for this competition here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455</a></p>\n\n<p>I have posted the explanations about CTL and TPU-optimized CTL here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a></p>",
      "rawMarkdown": "dimitreoliveira posted a TPU-optimized custom training loop notebook for this competition here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\n\nI have posted the explanations about CTL and TPU-optimized CTL here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443",
      "votes": null
    },
    {
      "id": "803718",
      "postDate": "04/10/2020 19:29:09",
      "content": "<p>PyTorch team seemed to make some updates to the runtime which should alleviate OOM issues (see the latest comments here: <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-612043779\">https://github.com/pytorch/xla/issues/1870#issuecomment-612043779</a>.\nThe changes should be in the latest nightly PyTorch/XLA.</p>",
      "rawMarkdown": "PyTorch team seemed to make some updates to the runtime which should alleviate OOM issues (see the latest comments here: [https://github.com/pytorch/xla/issues/1870#issuecomment-612043779](https://github.com/pytorch/xla/issues/1870#issuecomment-612043779).\nThe changes should be in the latest nightly PyTorch/XLA.",
      "votes": null
    },
    {
      "id": "803774",
      "postDate": "04/10/2020 20:27:45",
      "content": "<p>Tested it, and it doesn't seem to work. The collaborator there was using nprocs=1 instead of nprocs=8.</p>",
      "rawMarkdown": "Tested it, and it doesn't seem to work. The collaborator there was using nprocs=1 instead of nprocs=8.",
      "votes": null
    },
    {
      "id": "803799",
      "postDate": "04/10/2020 21:05:59",
      "content": "<p>As Davide replied on github issue, you are supposed to use <strong>nightly</strong> PyTorch/XLA (to get his latest changes):\n<code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --version nightly --apt-packages libomp5 libopenblas-dev\n</code>\nI.e. with this argument to the script: --version <strong>nightly</strong> .\nBy default their script now uses other version.\nWith nightly, he was able to use nprocs=8 based on the comments.</p>",
      "rawMarkdown": "As Davide replied on github issue, you are supposed to use **nightly** PyTorch/XLA (to get his latest changes):\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --version nightly --apt-packages libomp5 libopenblas-dev\n```\nI.e. with this argument to the script: --version **nightly** .\nBy default their script now uses other version.\nWith nightly, he was able to use nprocs=8 based on the comments.",
      "votes": null
    },
    {
      "id": "804157",
      "postDate": "04/11/2020 09:28:28",
      "content": "<p>I am always running out of kernel memory after spawning the 8 processes? Any idea what can be the reason here? Is it duplicating the training data also within the kernel for the processes (which are only 4)?</p>",
      "rawMarkdown": "I am always running out of kernel memory after spawning the 8 processes? Any idea what can be the reason here? Is it duplicating the training data also within the kernel for the processes (which are only 4)?",
      "votes": null
    },
    {
      "id": "804163",
      "postDate": "04/11/2020 09:42:02",
      "content": "<p>Don't duplicate your model over all 8 process Psi; You can define the dataset outside as well;  Plus we need to think where all we can save memory as well;</p>",
      "rawMarkdown": "Don't duplicate your model over all 8 process Psi; You can define the dataset outside as well;  Plus we need to think where all we can save memory as well;",
      "votes": null
    },
    {
      "id": "804256",
      "postDate": "04/11/2020 12:09:16",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, What do you think about doing things on fly in case of TPUs as well ?</p>",
      "rawMarkdown": "philippsinger, What do you think about doing things on fly in case of TPUs as well ?",
      "votes": null
    },
    {
      "id": "804377",
      "postDate": "04/11/2020 14:16:52",
      "content": "<p>I am going back to trying pytorch on TPU and now I get this error:</p>\n\n<p><code>tensorflow/compiler/xla/xla_client/tf_logging.cc:11] Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)</code></p>\n\n<p>Any idea?</p>",
      "rawMarkdown": "I am going back to trying pytorch on TPU and now I get this error:\n\n`tensorflow/compiler/xla/xla_client/tf_logging.cc:11] Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)`\n\nAny idea?",
      "votes": null
    },
    {
      "id": "804387",
      "postDate": "04/11/2020 14:37:05",
      "content": "<p>Seeing it for the first time! Can you share the whole tb with the code block line that triggered it?</p>",
      "rawMarkdown": "Seeing it for the first time! Can you share the whole tb with the code block line that triggered it?",
      "votes": null
    },
    {
      "id": "804396",
      "postDate": "04/11/2020 14:51:18",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "804398",
      "postDate": "04/11/2020 14:53:19",
      "content": "<blockquote>\n  <p>model = net.to(device)\n  This is making 8 copies, no? That's not the best way to do it.</p>\n</blockquote>\n\n<p>Refer this Alex, <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-612217012\">https://github.com/pytorch/xla/issues/1870#issuecomment-612217012</a></p>",
      "rawMarkdown": "&gt;model = net.to(device)\nThis is making 8 copies, no? That's not the best way to do it.\n\nRefer this Alex, https://github.com/pytorch/xla/issues/1870#issuecomment-612217012",
      "votes": null
    },
    {
      "id": "804401",
      "postDate": "04/11/2020 14:57:39",
      "content": "<p>Just calling <code>xma.save</code></p>",
      "rawMarkdown": "Just calling `xma.save`",
      "votes": null
    },
    {
      "id": "804442",
      "postDate": "04/11/2020 15:26:22",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "804447",
      "postDate": "04/11/2020 15:34:48",
      "content": "<p>del</p>",
      "rawMarkdown": "del",
      "votes": null
    },
    {
      "id": "804459",
      "postDate": "04/11/2020 15:50:47",
      "content": "<p>You used <code>os.environ['XLA_USE_BF16'] = \"1\"</code> - thank you. Now it works. </p>",
      "rawMarkdown": "You used `os.environ['XLA_USE_BF16'] = \"1\"` - thank you. Now it works.",
      "votes": null
    },
    {
      "id": "804467",
      "postDate": "04/11/2020 15:54:00",
      "content": "<p>Pleasure; i was helpful; That's half the problem resolved; In this colab nbs of mine, <a href=\"https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw\">https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw</a> ; the RAM is increasing over time and SIGKILL happens abruptly; \nNB RAM is ~13 GB on the colab;</p>",
      "rawMarkdown": "Pleasure; i was helpful; That's half the problem resolved; In this colab nbs of mine, https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw ; the RAM is increasing over time and SIGKILL happens abruptly; \nNB RAM is ~13 GB on the colab;",
      "votes": null
    },
    {
      "id": "805908",
      "postDate": "04/13/2020 08:09:40",
      "content": "<p>So i also get this error as well but at the very end, there's a SIGKILL as well for me; Can you confirm Psi? And if you fixed it, can you share how? (On Colab)</p>",
      "rawMarkdown": "So i also get this error as well but at the very end, there's a SIGKILL as well for me; Can you confirm Psi? And if you fixed it, can you share how? (On Colab)",
      "votes": null
    },
    {
      "id": "874595",
      "postDate": "06/05/2020 05:51:18",
      "content": "<p>Hi, all!\nNew problem \nException: process 0 terminated with signal SIGSEGV</p>\n\n<p>What's happened? And how to fix it?</p>",
      "rawMarkdown": "Hi, all!\nNew problem \nException: process 0 terminated with signal SIGSEGV\n\nWhat's happened? And how to fix it?",
      "votes": null
    },
    {
      "id": "876922",
      "postDate": "06/07/2020 06:53:04",
      "content": "<p>hey have you found the solution <a href=\"/andrilko\">@andrilko</a> </p>",
      "rawMarkdown": "hey have you found the solution @andrilko",
      "votes": null
    },
    {
      "id": "877373",
      "postDate": "06/07/2020 14:47:34",
      "content": "<p>Nope, sorry. =\\</p>",
      "rawMarkdown": "Nope, sorry. =\\",
      "votes": null
    },
    {
      "id": "918135",
      "postDate": "07/07/2020 02:47:06",
      "content": "<p>Hi, did you solve this problem? <code>Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)</code></p>",
      "rawMarkdown": "Hi, did you solve this problem? `Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 800129,
      "author_name": "parmarsuraj99",
      "author_url": "",
      "post_date": "04/07/2020 05:59:30",
      "content": "<p>Mine worked well with TPU on tf.keras.  I used pre-tokenized though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 800168,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/07/2020 06:31:12",
          "content": "<p>I am talking specifically about tf.keras custom training see this webpage for example what I mean by custom training pipeline. <a href=\"https://www.tensorflow.org/tutorials/distribute/custom_training\">https://www.tensorflow.org/tutorials/distribute/custom_training</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 803556,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/10/2020 16:39:52",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> posted a TPU-optimized custom training loop notebook for this competition here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455</a></p>\n\n<p>I have posted the explanations about CTL and TPU-optimized CTL here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 800904,
      "author_name": "shonenkov",
      "author_url": "",
      "post_date": "04/07/2020 21:09:48",
      "content": "<p><code>nprocs=1</code> - And it fixed this problem for me. But training is very slow (~approximately 2xP100). It is not full solution...</p>",
      "votes": null,
      "replies": [
        {
          "id": 800908,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/07/2020 21:19:06",
          "content": "<p>The problem did not get the fix for me unfortunately, btw, a question out of context, are you getting your current score with single model or what is your best with single model, you can skip the question if you want to, but I would like to hear to get an idea where I stand. Thank You!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 803015,
      "author_name": "jerryqu",
      "author_url": "",
      "post_date": "04/10/2020 04:12:20",
      "content": "<p>I have been running into the same bug when working with PyTorch.</p>\n\n<p>I believe this has to do with XLA using up RAM. I constantly use up all my RAM, which causes the SIGKILL error. If you take a look at this: <a href=\"https://github.com/pytorch/xla/issues/1280\">https://github.com/pytorch/xla/issues/1280</a></p>\n\n<p>They talk about how each of the 8 TPU processes replicates the model, which takes up a considerable amount of RAM. In terms of large transformer models (I was using XLM-RoBERTa), there simply isn't enough RAM to support the model being replicated 8 times. I've seen smaller models such as BERT being successfully loaded as seen in this kernel: <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores</a></p>\n\n<p>On the other hand, TensorFlow appears to support this better, as XLM-RoBERTa was successfully used in this kernel: <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>I'm going to switch over to TensorFlow, as Torch-XLA is just unusable with these memory issues.</p>",
      "votes": null,
      "replies": [
        {
          "id": 803025,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/10/2020 04:47:01",
          "content": "<p>Same Problem, I was using roberta, If I change that to bert, it works..., still has to use tensorflow only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 803718,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/10/2020 19:29:09",
          "content": "<p>PyTorch team seemed to make some updates to the runtime which should alleviate OOM issues (see the latest comments here: <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-612043779\">https://github.com/pytorch/xla/issues/1870#issuecomment-612043779</a>.\nThe changes should be in the latest nightly PyTorch/XLA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 803774,
          "author_name": "jerryqu",
          "author_url": "",
          "post_date": "04/10/2020 20:27:45",
          "content": "<p>Tested it, and it doesn't seem to work. The collaborator there was using nprocs=1 instead of nprocs=8.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 803799,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/10/2020 21:05:59",
          "content": "<p>As Davide replied on github issue, you are supposed to use <strong>nightly</strong> PyTorch/XLA (to get his latest changes):\n<code>\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --version nightly --apt-packages libomp5 libopenblas-dev\n</code>\nI.e. with this argument to the script: --version <strong>nightly</strong> .\nBy default their script now uses other version.\nWith nightly, he was able to use nprocs=8 based on the comments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804157,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/11/2020 09:28:28",
          "content": "<p>I am always running out of kernel memory after spawning the 8 processes? Any idea what can be the reason here? Is it duplicating the training data also within the kernel for the processes (which are only 4)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804163,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 09:42:02",
          "content": "<p>Don't duplicate your model over all 8 process Psi; You can define the dataset outside as well;  Plus we need to think where all we can save memory as well;</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804256,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 12:09:16",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, What do you think about doing things on fly in case of TPUs as well ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804396,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "04/11/2020 14:51:18",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 804398,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 14:53:19",
          "content": "<blockquote>\n  <p>model = net.to(device)\n  This is making 8 copies, no? That's not the best way to do it.</p>\n</blockquote>\n\n<p>Refer this Alex, <a href=\"https://github.com/pytorch/xla/issues/1870#issuecomment-612217012\">https://github.com/pytorch/xla/issues/1870#issuecomment-612217012</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804442,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "04/11/2020 15:26:22",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 804447,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 15:34:48",
          "content": "<p>del</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804459,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "04/11/2020 15:50:47",
          "content": "<p>You used <code>os.environ['XLA_USE_BF16'] = \"1\"</code> - thank you. Now it works. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804467,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 15:54:00",
          "content": "<p>Pleasure; i was helpful; That's half the problem resolved; In this colab nbs of mine, <a href=\"https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw\">https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw</a> ; the RAM is increasing over time and SIGKILL happens abruptly; \nNB RAM is ~13 GB on the colab;</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 804377,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/11/2020 14:16:52",
      "content": "<p>I am going back to trying pytorch on TPU and now I get this error:</p>\n\n<p><code>tensorflow/compiler/xla/xla_client/tf_logging.cc:11] Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)</code></p>\n\n<p>Any idea?</p>",
      "votes": null,
      "replies": [
        {
          "id": 804387,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/11/2020 14:37:05",
          "content": "<p>Seeing it for the first time! Can you share the whole tb with the code block line that triggered it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804401,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/11/2020 14:57:39",
          "content": "<p>Just calling <code>xma.save</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 805908,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/13/2020 08:09:40",
          "content": "<p>So i also get this error as well but at the very end, there's a SIGKILL as well for me; Can you confirm Psi? And if you fixed it, can you share how? (On Colab)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918135,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "07/07/2020 02:47:06",
          "content": "<p>Hi, did you solve this problem? <code>Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 874595,
      "author_name": "andrilko",
      "author_url": "",
      "post_date": "06/05/2020 05:51:18",
      "content": "<p>Hi, all!\nNew problem \nException: process 0 terminated with signal SIGSEGV</p>\n\n<p>What's happened? And how to fix it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 876922,
          "author_name": "pranshu29",
          "author_url": "",
          "post_date": "06/07/2020 06:53:04",
          "content": "<p>hey have you found the solution <a href=\"/andrilko\">@andrilko</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 877373,
          "author_name": "andrilko",
          "author_url": "",
          "post_date": "06/07/2020 14:47:34",
          "content": "<p>Nope, sorry. =\\</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "799952": "First I tried custom tensorflow to control loss and metrics and other layers of the model, which resulted in bugs, I did make the pytorch version of the same and now again this bug where\nhttps://www.kaggle.com/tanlikesmath/simple-xlmr-tpu-pytorch\n\nException: process 0 terminated with signal SIGKILL\n\nI have no way to compete without the bug getting a fix either in tensorflow or torch. Guess its most likely a dead end for me atleast for now.\n\nI ask the kaggle community if they have encountered this error or have found any fix.\n\nThank You!",
    "800129": "Mine worked well with TPU on tf.keras.  I used pre-tokenized though.",
    "800168": "I am talking specifically about tf.keras custom training see this webpage for example what I mean by custom training pipeline. https://www.tensorflow.org/tutorials/distribute/custom_training",
    "800904": "`nprocs=1` - And it fixed this problem for me. But training is very slow (~approximately 2xP100). It is not full solution...",
    "800908": "The problem did not get the fix for me unfortunately, btw, a question out of context, are you getting your current score with single model or what is your best with single model, you can skip the question if you want to, but I would like to hear to get an idea where I stand. Thank You!",
    "803015": "I have been running into the same bug when working with PyTorch.\n\nI believe this has to do with XLA using up RAM. I constantly use up all my RAM, which causes the SIGKILL error. If you take a look at this: https://github.com/pytorch/xla/issues/1280\n\nThey talk about how each of the 8 TPU processes replicates the model, which takes up a considerable amount of RAM. In terms of large transformer models (I was using XLM-RoBERTa), there simply isn't enough RAM to support the model being replicated 8 times. I've seen smaller models such as BERT being successfully loaded as seen in this kernel: https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\n\nOn the other hand, TensorFlow appears to support this better, as XLM-RoBERTa was successfully used in this kernel: https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\nI'm going to switch over to TensorFlow, as Torch-XLA is just unusable with these memory issues.",
    "803025": "Same Problem, I was using roberta, If I change that to bert, it works..., still has to use tensorflow only.",
    "803556": "dimitreoliveira posted a TPU-optimized custom training loop notebook for this competition here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/139455\n\nI have posted the explanations about CTL and TPU-optimized CTL here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443",
    "803718": "PyTorch team seemed to make some updates to the runtime which should alleviate OOM issues (see the latest comments here: [https://github.com/pytorch/xla/issues/1870#issuecomment-612043779](https://github.com/pytorch/xla/issues/1870#issuecomment-612043779).\nThe changes should be in the latest nightly PyTorch/XLA.",
    "803774": "Tested it, and it doesn't seem to work. The collaborator there was using nprocs=1 instead of nprocs=8.",
    "803799": "As Davide replied on github issue, you are supposed to use **nightly** PyTorch/XLA (to get his latest changes):\n```\n!curl https://raw.githubusercontent.com/pytorch/xla/master/contrib/scripts/env-setup.py -o pytorch-xla-env-setup.py\n!python pytorch-xla-env-setup.py --version nightly --apt-packages libomp5 libopenblas-dev\n```\nI.e. with this argument to the script: --version **nightly** .\nBy default their script now uses other version.\nWith nightly, he was able to use nprocs=8 based on the comments.",
    "804157": "I am always running out of kernel memory after spawning the 8 processes? Any idea what can be the reason here? Is it duplicating the training data also within the kernel for the processes (which are only 4)?",
    "804163": "Don't duplicate your model over all 8 process Psi; You can define the dataset outside as well;  Plus we need to think where all we can save memory as well;",
    "804256": "philippsinger, What do you think about doing things on fly in case of TPUs as well ?",
    "804377": "I am going back to trying pytorch on TPU and now I get this error:\n\n`tensorflow/compiler/xla/xla_client/tf_logging.cc:11] Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)`\n\nAny idea?",
    "804387": "Seeing it for the first time! Can you share the whole tb with the code block line that triggered it?",
    "804396": "",
    "804398": "&gt;model = net.to(device)\nThis is making 8 copies, no? That's not the best way to do it.\n\nRefer this Alex, https://github.com/pytorch/xla/issues/1870#issuecomment-612217012",
    "804401": "Just calling `xma.save`",
    "804442": "",
    "804447": "del",
    "804459": "You used `os.environ['XLA_USE_BF16'] = \"1\"` - thank you. Now it works.",
    "804467": "Pleasure; i was helpful; That's half the problem resolved; In this colab nbs of mine, https://colab.research.google.com/drive/1wKU8El2C_hF60460EYKOXhIeibO4Pxiw#scrollTo=exkRegYELsgw ; the RAM is increasing over time and SIGKILL happens abruptly; \nNB RAM is ~13 GB on the colab;",
    "805908": "So i also get this error as well but at the very end, there's a SIGKILL as well for me; Can you confirm Psi? And if you fixed it, can you share how? (On Colab)",
    "874595": "Hi, all!\nNew problem \nException: process 0 terminated with signal SIGSEGV\n\nWhat's happened? And how to fix it?",
    "876922": "hey have you found the solution @andrilko",
    "877373": "Nope, sorry. =\\",
    "918135": "Hi, did you solve this problem? `Failed to meet rendezvous 'torch_xla.core.xla_model.save': Socket closed (14)`"
  },
  "source": "meta"
}