{
  "id": 140007,
  "title": "InvalidArgumentError error while using TPU",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/140007",
  "author_name": "",
  "post_date": "2020-03-30T23:08:51.513752800Z",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F95c4e8da80b617bbc06713f6cc1752ba%2FScreenshot%20from%202020-03-30%2020-05-44.png?generation=1585609582606935&amp;alt=media\" alt=\"\">\nI would like to report this bug, to see if anyone else got this, or have any idea how to solve it.</p>\n\n<p>I have created a dataset output from a kernel, then I load this data and try to train a model with TPU and get this error, but the curious thing is, I can train a model with this same data if I use CPU or GPU, so I guess that the problem is loading that data on the TPU core.</p>",
  "messages": [
    {
      "id": "792127",
      "postDate": "03/30/2020 23:08:51",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F95c4e8da80b617bbc06713f6cc1752ba%2FScreenshot%20from%202020-03-30%2020-05-44.png?generation=1585609582606935&amp;alt=media\" alt=\"\">\nI would like to report this bug, to see if anyone else got this, or have any idea how to solve it.</p>\n\n<p>I have created a dataset output from a kernel, then I load this data and try to train a model with TPU and get this error, but the curious thing is, I can train a model with this same data if I use CPU or GPU, so I guess that the problem is loading that data on the TPU core.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F95c4e8da80b617bbc06713f6cc1752ba%2FScreenshot%20from%202020-03-30%2020-05-44.png?generation=1585609582606935&amp;alt=media)\nI would like to report this bug, to see if anyone else got this, or have any idea how to solve it.\n\nI have created a dataset output from a kernel, then I load this data and try to train a model with TPU and get this error, but the curious thing is, I can train a model with this same data if I use CPU or GPU, so I guess that the problem is loading that data on the TPU core.",
      "votes": null
    },
    {
      "id": "792152",
      "postDate": "03/30/2020 23:48:19",
      "content": "<p>Yes. I had the same problem . It happend for tensorflow version. There are no problem in pytorch xla. </p>",
      "rawMarkdown": "Yes. I had the same problem . It happend for tensorflow version. There are no problem in pytorch xla.",
      "votes": null
    },
    {
      "id": "792169",
      "postDate": "03/31/2020 00:19:24",
      "content": "<p>This is usually a sign that the TPU has crashed. The next operation you try to run on the TPU after the crash produces this error.</p>",
      "rawMarkdown": "This is usually a sign that the TPU has crashed. The next operation you try to run on the TPU after the crash produces this error.",
      "votes": null
    },
    {
      "id": "792188",
      "postDate": "03/31/2020 00:46:26",
      "content": "<p>Here is the interesting part, I have 2 datasets, output from different kernels, but they are generated the same way (one from the 1st Jigsaw competition and other from the 2nd), but I can use normally the first dataset, the second is the one that gives me problems.</p>",
      "rawMarkdown": "Here is the interesting part, I have 2 datasets, output from different kernels, but they are generated the same way (one from the 1st Jigsaw competition and other from the 2nd), but I can use normally the first dataset, the second is the one that gives me problems.",
      "votes": null
    },
    {
      "id": "792194",
      "postDate": "03/31/2020 00:52:36",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> , you used the same dataset and just changed to Pytorch, and the error disappeared?</p>",
      "rawMarkdown": "qinhui1999 , you used the same dataset and just changed to Pytorch, and the error disappeared?",
      "votes": null
    },
    {
      "id": "792209",
      "postDate": "03/31/2020 01:32:04",
      "content": "<p>In pytorch, I used the max_seq_length=320, bs=64. But in tensorflow,  I used the max_seq_length=512, bs =128. Maybe that is the key factor.   Too much memories required,  it made the TPU crash.  I will test again.</p>",
      "rawMarkdown": "In pytorch, I used the max_seq_length=320, bs=64. But in tensorflow,  I used the max_seq_length=512, bs =128. Maybe that is the key factor.   Too much memories required,  it made the TPU crash.  I will test again.",
      "votes": null
    },
    {
      "id": "792211",
      "postDate": "03/31/2020 01:39:34",
      "content": "<p>I think this is not the case, I've tried, with smaller dataset and batch size, but got the same error...</p>",
      "rawMarkdown": "I think this is not the case, I've tried, with smaller dataset and batch size, but got the same error...",
      "votes": null
    },
    {
      "id": "792227",
      "postDate": "03/31/2020 01:51:10",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for your help. I changed the batch_size  into 64, now it works in tensorflow.   To my surprise, tf tpu version only need about 5g vm memory. But my pytorch  xla tpu version need over 15g vm memory.  I often meet memory crash  in pytorch xla tpu kernel.  But in tf tpu version , it is fine. </p>",
      "rawMarkdown": "dimitreoliveira @mgornergoogle Thanks for your help. I changed the batch_size  into 64, now it works in tensorflow.   To my surprise, tf tpu version only need about 5g vm memory. But my pytorch  xla tpu version need over 15g vm memory.  I often meet memory crash  in pytorch xla tpu kernel.  But in tf tpu version , it is fine.",
      "votes": null
    },
    {
      "id": "792325",
      "postDate": "03/31/2020 05:07:56",
      "content": "<p>Hi <a href=\"/dimitreoliveira\">@dimitreoliveira</a> : Have you been able to solve this issue? I experienced the same issue and wasted all my TPU hours trying to fix it, unsuccessfully:(</p>",
      "rawMarkdown": "Hi @dimitreoliveira : Have you been able to solve this issue? I experienced the same issue and wasted all my TPU hours trying to fix it, unsuccessfully:(",
      "votes": null
    },
    {
      "id": "792414",
      "postDate": "03/31/2020 07:19:07",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> The bug is that the toxic targets in 2018's data are int . But the toxic targets in 2019's data are float.  Hence TPU can not work for mixing int with float.  I cast the float into int32 for the toxic targets, then TPU can run my custom datas. </p>",
      "rawMarkdown": "dimitreoliveira The bug is that the toxic targets in 2018's data are int . But the toxic targets in 2019's data are float.  Hence TPU can not work for mixing int with float.  I cast the float into int32 for the toxic targets, then TPU can run my custom datas.",
      "votes": null
    },
    {
      "id": "792702",
      "postDate": "03/31/2020 13:43:12",
      "content": "<p>Hey <a href=\"/qinhui1999\">@qinhui1999</a> you got it, this was a tricky issue to get, mainly because for me it runs on CPU and GPU fine, but on TPU I get this weird error, but if I cast the labels of both datasets to <code>int</code> or <code>float</code> then I get to train fine, thanks!</p>",
      "rawMarkdown": "Hey @qinhui1999 you got it, this was a tricky issue to get, mainly because for me it runs on CPU and GPU fine, but on TPU I get this weird error, but if I cast the labels of both datasets to `int` or `float` then I get to train fine, thanks!",
      "votes": null
    },
    {
      "id": "792704",
      "postDate": "03/31/2020 13:44:15",
      "content": "<p>Yes as pointed by <a href=\"/qinhui1999\">@qinhui1999</a> , for TPU you need to have the labels on the same type. I've wasted a lot of TPU hours as well ☹️ .</p>",
      "rawMarkdown": "Yes as pointed by @qinhui1999 , for TPU you need to have the labels on the same type. I've wasted a lot of TPU hours as well ☹️ .",
      "votes": null
    },
    {
      "id": "793101",
      "postDate": "03/31/2020 19:55:56",
      "content": "<p>Good catch. Down the line, remember that TPUs can only handle ints and floats. Even int8 will need to be promoted to int. And also bfloat16 if you use them explicitly.</p>",
      "rawMarkdown": "Good catch. Down the line, remember that TPUs can only handle ints and floats. Even int8 will need to be promoted to int. And also bfloat16 if you use them explicitly.",
      "votes": null
    },
    {
      "id": "793102",
      "postDate": "03/31/2020 19:57:22",
      "content": "<p>Yes, that's expected, in Tensorflow, there is practically nothing running on the Kaggle VM when you train. PyTorch uses a different design.</p>",
      "rawMarkdown": "Yes, that's expected, in Tensorflow, there is practically nothing running on the Kaggle VM when you train. PyTorch uses a different design.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 792152,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/30/2020 23:48:19",
      "content": "<p>Yes. I had the same problem . It happend for tensorflow version. There are no problem in pytorch xla. </p>",
      "votes": null,
      "replies": [
        {
          "id": 792194,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 00:52:36",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> , you used the same dataset and just changed to Pytorch, and the error disappeared?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792209,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "03/31/2020 01:32:04",
          "content": "<p>In pytorch, I used the max_seq_length=320, bs=64. But in tensorflow,  I used the max_seq_length=512, bs =128. Maybe that is the key factor.   Too much memories required,  it made the TPU crash.  I will test again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792211,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 01:39:34",
          "content": "<p>I think this is not the case, I've tried, with smaller dataset and batch size, but got the same error...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 792169,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/31/2020 00:19:24",
      "content": "<p>This is usually a sign that the TPU has crashed. The next operation you try to run on the TPU after the crash produces this error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 792188,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 00:46:26",
          "content": "<p>Here is the interesting part, I have 2 datasets, output from different kernels, but they are generated the same way (one from the 1st Jigsaw competition and other from the 2nd), but I can use normally the first dataset, the second is the one that gives me problems.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 792227,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/31/2020 01:51:10",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> <a href=\"/mgornergoogle\">@mgornergoogle</a> Thanks for your help. I changed the batch_size  into 64, now it works in tensorflow.   To my surprise, tf tpu version only need about 5g vm memory. But my pytorch  xla tpu version need over 15g vm memory.  I often meet memory crash  in pytorch xla tpu kernel.  But in tf tpu version , it is fine. </p>",
      "votes": null,
      "replies": [
        {
          "id": 793102,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/31/2020 19:57:22",
          "content": "<p>Yes, that's expected, in Tensorflow, there is practically nothing running on the Kaggle VM when you train. PyTorch uses a different design.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 792325,
      "author_name": "kvsnoufal",
      "author_url": "",
      "post_date": "03/31/2020 05:07:56",
      "content": "<p>Hi <a href=\"/dimitreoliveira\">@dimitreoliveira</a> : Have you been able to solve this issue? I experienced the same issue and wasted all my TPU hours trying to fix it, unsuccessfully:(</p>",
      "votes": null,
      "replies": [
        {
          "id": 792704,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 13:44:15",
          "content": "<p>Yes as pointed by <a href=\"/qinhui1999\">@qinhui1999</a> , for TPU you need to have the labels on the same type. I've wasted a lot of TPU hours as well ☹️ .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 792414,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/31/2020 07:19:07",
      "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> The bug is that the toxic targets in 2018's data are int . But the toxic targets in 2019's data are float.  Hence TPU can not work for mixing int with float.  I cast the float into int32 for the toxic targets, then TPU can run my custom datas. </p>",
      "votes": null,
      "replies": [
        {
          "id": 792702,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 13:43:12",
          "content": "<p>Hey <a href=\"/qinhui1999\">@qinhui1999</a> you got it, this was a tricky issue to get, mainly because for me it runs on CPU and GPU fine, but on TPU I get this weird error, but if I cast the labels of both datasets to <code>int</code> or <code>float</code> then I get to train fine, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793101,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/31/2020 19:55:56",
          "content": "<p>Good catch. Down the line, remember that TPUs can only handle ints and floats. Even int8 will need to be promoted to int. And also bfloat16 if you use them explicitly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "792127": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F95c4e8da80b617bbc06713f6cc1752ba%2FScreenshot%20from%202020-03-30%2020-05-44.png?generation=1585609582606935&amp;alt=media)\nI would like to report this bug, to see if anyone else got this, or have any idea how to solve it.\n\nI have created a dataset output from a kernel, then I load this data and try to train a model with TPU and get this error, but the curious thing is, I can train a model with this same data if I use CPU or GPU, so I guess that the problem is loading that data on the TPU core.",
    "792152": "Yes. I had the same problem . It happend for tensorflow version. There are no problem in pytorch xla.",
    "792169": "This is usually a sign that the TPU has crashed. The next operation you try to run on the TPU after the crash produces this error.",
    "792188": "Here is the interesting part, I have 2 datasets, output from different kernels, but they are generated the same way (one from the 1st Jigsaw competition and other from the 2nd), but I can use normally the first dataset, the second is the one that gives me problems.",
    "792194": "qinhui1999 , you used the same dataset and just changed to Pytorch, and the error disappeared?",
    "792209": "In pytorch, I used the max_seq_length=320, bs=64. But in tensorflow,  I used the max_seq_length=512, bs =128. Maybe that is the key factor.   Too much memories required,  it made the TPU crash.  I will test again.",
    "792211": "I think this is not the case, I've tried, with smaller dataset and batch size, but got the same error...",
    "792227": "dimitreoliveira @mgornergoogle Thanks for your help. I changed the batch_size  into 64, now it works in tensorflow.   To my surprise, tf tpu version only need about 5g vm memory. But my pytorch  xla tpu version need over 15g vm memory.  I often meet memory crash  in pytorch xla tpu kernel.  But in tf tpu version , it is fine.",
    "792325": "Hi @dimitreoliveira : Have you been able to solve this issue? I experienced the same issue and wasted all my TPU hours trying to fix it, unsuccessfully:(",
    "792414": "dimitreoliveira The bug is that the toxic targets in 2018's data are int . But the toxic targets in 2019's data are float.  Hence TPU can not work for mixing int with float.  I cast the float into int32 for the toxic targets, then TPU can run my custom datas.",
    "792702": "Hey @qinhui1999 you got it, this was a tricky issue to get, mainly because for me it runs on CPU and GPU fine, but on TPU I get this weird error, but if I cast the labels of both datasets to `int` or `float` then I get to train fine, thanks!",
    "792704": "Yes as pointed by @qinhui1999 , for TPU you need to have the labels on the same type. I've wasted a lot of TPU hours as well ☹️ .",
    "793101": "Good catch. Down the line, remember that TPUs can only handle ints and floats. Even int8 will need to be promoted to int. And also bfloat16 if you use them explicitly.",
    "793102": "Yes, that's expected, in Tensorflow, there is practically nothing running on the Kaggle VM when you train. PyTorch uses a different design."
  },
  "source": "meta"
}