{
  "id": 138511,
  "title": "PyTorch XLA questions and resources thread",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138511",
  "author_name": "",
  "post_date": "2020-03-25T08:34:23.303726200Z",
  "votes": 10,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I see <a href=\"/philippsinger\">@philippsinger</a> made a thread here for TPU questions. I think it is worth it to create a separate thread regarding PyTorch XLA for TPU training.</p>\n\n<p>Here is also a list of useful PyTorch XLA resources:\n1. <a href=\"https://github.com/pytorch/xla\">PyTorch XLA repository</a>\n2. <a href=\"https://github.com/pytorch/xla/blob/master/API_GUIDE.md\">API guide</a>\n3. <a href=\"https://pytorch.org/xla/\">Official documentation</a>\n4. <a href=\"https://github.com/pytorch/xla/tree/master/contrib/colab\">Colab examples</a>\n5. <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138137\">Kaggle discussion on PyTorch XLA support</a>\n6. <a href=\"/abhishek\">@abhishek</a>'s <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265\">discussion post</a>: <a href=\"https://youtu.be/vvr_f-X_LaI\">YouTube video</a>, <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">training kernel</a>, <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml\">inference kernel</a> \nEDIT: 8 cores working <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">here</a></p>\n\n<p>Feel free to comment with your questions and other resources available.</p>",
  "messages": [
    {
      "id": "785655",
      "postDate": "03/25/2020 08:34:23",
      "content": "<p>I see <a href=\"/philippsinger\">@philippsinger</a> made a thread here for TPU questions. I think it is worth it to create a separate thread regarding PyTorch XLA for TPU training.</p>\n\n<p>Here is also a list of useful PyTorch XLA resources:\n1. <a href=\"https://github.com/pytorch/xla\">PyTorch XLA repository</a>\n2. <a href=\"https://github.com/pytorch/xla/blob/master/API_GUIDE.md\">API guide</a>\n3. <a href=\"https://pytorch.org/xla/\">Official documentation</a>\n4. <a href=\"https://github.com/pytorch/xla/tree/master/contrib/colab\">Colab examples</a>\n5. <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138137\">Kaggle discussion on PyTorch XLA support</a>\n6. <a href=\"/abhishek\">@abhishek</a>'s <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265\">discussion post</a>: <a href=\"https://youtu.be/vvr_f-X_LaI\">YouTube video</a>, <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training\">training kernel</a>, <a href=\"https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml\">inference kernel</a> \nEDIT: 8 cores working <a href=\"https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores\">here</a></p>\n\n<p>Feel free to comment with your questions and other resources available.</p>",
      "rawMarkdown": "I see @philippsinger made a thread [here]() for TPU questions. I think it is worth it to create a separate thread regarding PyTorch XLA for TPU training.\n\nHere is also a list of useful PyTorch XLA resources:\n1. [PyTorch XLA repository](https://github.com/pytorch/xla)\n2. [API guide](https://github.com/pytorch/xla/blob/master/API_GUIDE.md)\n3. [Official documentation](https://pytorch.org/xla/)\n4. [Colab examples](https://github.com/pytorch/xla/tree/master/contrib/colab)\n5. [Kaggle discussion on PyTorch XLA support](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138137)\n6. @abhishek's [discussion post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265): [YouTube video](https://youtu.be/vvr_f-X_LaI), [training kernel](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training), [inference kernel](https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml) \nEDIT: 8 cores working [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores)\n\nFeel free to comment with your questions and other resources available.",
      "votes": null
    },
    {
      "id": "785674",
      "postDate": "03/25/2020 08:59:28",
      "content": "<p><a href=\"https://github.com/pytorch/xla/issues/1819\">Here</a>, I have added an issue as well on XLA Repo for some clarifications regarding the comment and the strategy as mentioned <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265#785608\">here</a></p>",
      "rawMarkdown": "[Here](https://github.com/pytorch/xla/issues/1819), I have added an issue as well on XLA Repo for some clarifications regarding the comment and the strategy as mentioned [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265#785608)",
      "votes": null
    },
    {
      "id": "785813",
      "postDate": "03/25/2020 12:18:40",
      "content": "<p>Did someone encounter kernels being stuck in commit mode?</p>",
      "rawMarkdown": "Did someone encounter kernels being stuck in commit mode?",
      "votes": null
    },
    {
      "id": "785816",
      "postDate": "03/25/2020 12:22:52",
      "content": "<p>Not recently but in the morning yes; &lt;2 mins kernel never got committed (non-TPU)</p>",
      "rawMarkdown": "Not recently but in the morning yes; &lt;2 mins kernel never got committed (non-TPU)",
      "votes": null
    },
    {
      "id": "786209",
      "postDate": "03/25/2020 17:45:13",
      "content": "<p>When I run the notebook , in the TPU details ,MXU efficiency shows 9.00%. Is that too low , and normal? Running 5 epochs , with a batch size of 64 and max_len = 192 had about train steps equal to about 1900, and the notebook timed out. Is there any way to optimize the MXU efficiency on TPU , or is there some mistake in my code? (I'm using Abhishek's notebook with 8 cores, which I wrote on my own so there is a possibility of having some mistakes there.) </p>",
      "rawMarkdown": "When I run the notebook , in the TPU details ,MXU efficiency shows 9.00%. Is that too low , and normal? Running 5 epochs , with a batch size of 64 and max_len = 192 had about train steps equal to about 1900, and the notebook timed out. Is there any way to optimize the MXU efficiency on TPU , or is there some mistake in my code? (I'm using Abhishek's notebook with 8 cores, which I wrote on my own so there is a possibility of having some mistakes there.)",
      "votes": null
    },
    {
      "id": "787542",
      "postDate": "03/26/2020 21:42:24",
      "content": "<p>I do have encountered the same thing twice which is quite weird. I had to restart the session and then again committed, and that worked.</p>",
      "rawMarkdown": "I do have encountered the same thing twice which is quite weird. I had to restart the session and then again committed, and that worked.",
      "votes": null
    },
    {
      "id": "787726",
      "postDate": "03/27/2020 02:26:49",
      "content": "<p>Can you explain what it means to be stuck in commit? Do you mean it's running and never finishing? Feel free to send links to those commits to me (the viewer url is what I need).</p>\n\n<p>I can investigate what's up with them.</p>",
      "rawMarkdown": "Can you explain what it means to be stuck in commit? Do you mean it's running and never finishing? Feel free to send links to those commits to me (the viewer url is what I need).\n\nI can investigate what's up with them.",
      "votes": null
    },
    {
      "id": "787945",
      "postDate": "03/27/2020 08:42:16",
      "content": "<p>Hi, What happened earlier was that when I tried to execute all shells from the starting of the notebook till training shell, everything worked normally and I got first loss update from the TPU server, but after the first loss message, I have waited for 20-25 minutes and it was kind of stucked, I mean the training part. But when I restarted the session and re-run the same thing, It worked flawlessly. With the same code, training was properly working over TPU session and I was able to get loss updates after every 10 sec. Not sure if it was related to PyTorch XLA or Kaggle kernel. For the link, As there are multiple versions of that same notebook, So I do not remember, which version was it. :( </p>",
      "rawMarkdown": "Hi, What happened earlier was that when I tried to execute all shells from the starting of the notebook till training shell, everything worked normally and I got first loss update from the TPU server, but after the first loss message, I have waited for 20-25 minutes and it was kind of stucked, I mean the training part. But when I restarted the session and re-run the same thing, It worked flawlessly. With the same code, training was properly working over TPU session and I was able to get loss updates after every 10 sec. Not sure if it was related to PyTorch XLA or Kaggle kernel. For the link, As there are multiple versions of that same notebook, So I do not remember, which version was it. :(",
      "votes": null
    },
    {
      "id": "788207",
      "postDate": "03/27/2020 13:33:27",
      "content": "<p>Okay thanks, glad it's working now though. If you run into it again feel free to send me the link and I'll take a look.</p>",
      "rawMarkdown": "Okay thanks, glad it's working now though. If you run into it again feel free to send me the link and I'll take a look.",
      "votes": null
    },
    {
      "id": "788380",
      "postDate": "03/27/2020 16:09:39",
      "content": "<p>It's weird that with the latest kernel I have submitted now, It is again getting stucked when I'm using no of epoch just 1 more than what I'm currently using. I'm confident that it should take max 1.5 Hrs to complete, as the current one with 1 less epoch is taking 2281.1s to complete, but instead, it's getting stucked. Due to this, commit keeps expiring after running for 3 Hrs limit and I had to decrease the epoch number by 1. Now unfortunately, my TPU quota also burned out for this week as my 2 commits wasted due to this thing. It would be much helpful if you can point me in the right direction about what could be wrong.</p>",
      "rawMarkdown": "It's weird that with the latest kernel I have submitted now, It is again getting stucked when I'm using no of epoch just 1 more than what I'm currently using. I'm confident that it should take max 1.5 Hrs to complete, as the current one with 1 less epoch is taking 2281.1s to complete, but instead, it's getting stucked. Due to this, commit keeps expiring after running for 3 Hrs limit and I had to decrease the epoch number by 1. Now unfortunately, my TPU quota also burned out for this week as my 2 commits wasted due to this thing. It would be much helpful if you can point me in the right direction about what could be wrong.",
      "votes": null
    },
    {
      "id": "788386",
      "postDate": "03/27/2020 16:17:08",
      "content": "<p>To clarify, is it actually 'getting stuck' or does the log indicate it's getting to the final epoch but that's not finishing within the 3h?</p>\n\n<p>I'm less of an expert on data science code, my expertise if on the code execution system for notebooks. I'd still need the viewer URL to be able to look deeper into your notebook.</p>",
      "rawMarkdown": "To clarify, is it actually 'getting stuck' or does the log indicate it's getting to the final epoch but that's not finishing within the 3h?\n\nI'm less of an expert on data science code, my expertise if on the code execution system for notebooks. I'd still need the viewer URL to be able to look deeper into your notebook.",
      "votes": null
    },
    {
      "id": "788396",
      "postDate": "03/27/2020 16:26:42",
      "content": "<p>Its getting stucked. I ran it during the interactive session to confirm after my 2 commits timed out and certainly that was the reason for the timeout. I have sent you an email for my kernel link. Please find it.</p>",
      "rawMarkdown": "Its getting stucked. I ran it during the interactive session to confirm after my 2 commits timed out and certainly that was the reason for the timeout. I have sent you an email for my kernel link. Please find it.",
      "votes": null
    },
    {
      "id": "788420",
      "postDate": "03/27/2020 16:51:23",
      "content": "<p>It's not clear to me that anything in our systems is actually getting stuck, I think the code is actually taking &gt; 3hrs and just not finishing in time. How many epochs do you run when it runs under 3 hours vs. over?</p>\n\n<p>If possible you might want to try using \"script\" mode instead of notebook mode and dump logs at critical points to get an idea of the timing.</p>\n\n<p>The logs for notebook mode are not that great at this point in time and can make it difficult to track progress. I'll see if any of the TPU experts on our team are aware of any other possible issues. </p>",
      "rawMarkdown": "It's not clear to me that anything in our systems is actually getting stuck, I think the code is actually taking &gt; 3hrs and just not finishing in time. How many epochs do you run when it runs under 3 hours vs. over?\n\nIf possible you might want to try using \"script\" mode instead of notebook mode and dump logs at critical points to get an idea of the timing.\n\nThe logs for notebook mode are not that great at this point in time and can make it difficult to track progress. I'll see if any of the TPU experts on our team are aware of any other possible issues.",
      "votes": null
    },
    {
      "id": "788436",
      "postDate": "03/27/2020 17:06:53",
      "content": "<p>Well, there is just 1 epoch difference between the code which completed in just 2281.1 seconds Vs the code which didn't even completed in 10800 seconds which is definitely not adding up right. BTW, once I will get my TPU limit back in next week, I will try to run this through scripts as you suggested and will see if it persists or what. Thanks.</p>",
      "rawMarkdown": "Well, there is just 1 epoch difference between the code which completed in just 2281.1 seconds Vs the code which didn't even completed in 10800 seconds which is definitely not adding up right. BTW, once I will get my TPU limit back in next week, I will try to run this through scripts as you suggested and will see if it persists or what. Thanks.",
      "votes": null
    },
    {
      "id": "1710094",
      "postDate": "03/02/2022 18:37:40",
      "content": "<p>I am trying to commit a tensorflow notebook with TPU, but it is finishing very quickly within 5 seconds, and is running non-TPU, it is not at all training the model and not giving any output. I tried this 4-5 times in a row, same thing is happening every time.</p>\n<p>I also tried restarting the notebook also, still getting the same issue.</p>\n<p>Can someone please suggest how to solve this issue? thanks in advance!</p>",
      "rawMarkdown": "I am trying to commit a tensorflow notebook with TPU, but it is finishing very quickly within 5 seconds, and is running non-TPU, it is not at all training the model and not giving any output. I tried this 4-5 times in a row, same thing is happening every time.\n\nI also tried restarting the notebook also, still getting the same issue.\n\nCan someone please suggest how to solve this issue? thanks in advance!",
      "votes": null
    },
    {
      "id": "1710099",
      "postDate": "03/02/2022 18:43:43",
      "content": "<p><a href=\"https://www.kaggle.com/nanditab35\" target=\"_blank\">@nanditab35</a> When you use \"Commit\" are you selecting \"Save &amp; Run All\" or \"Quick Version\"?</p>\n<p>Quick version just copies your notebook as it looks in the editor, \"Save &amp; Run All\" runs the notebook from start to finish with the proper settings (TPU etc.)</p>",
      "rawMarkdown": "nanditab35 When you use \"Commit\" are you selecting \"Save & Run All\" or \"Quick Version\"?\n\nQuick version just copies your notebook as it looks in the editor, \"Save & Run All\" runs the notebook from start to finish with the proper settings (TPU etc.)",
      "votes": null
    },
    {
      "id": "1710460",
      "postDate": "03/03/2022 04:54:05",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> for the suggestion. This actually solved my issue. Thanks again!</p>",
      "rawMarkdown": "Thanks @herbison for the suggestion. This actually solved my issue. Thanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1710094,
      "author_name": "nanditab35",
      "author_url": "",
      "post_date": "03/02/2022 18:37:40",
      "content": "<p>I am trying to commit a tensorflow notebook with TPU, but it is finishing very quickly within 5 seconds, and is running non-TPU, it is not at all training the model and not giving any output. I tried this 4-5 times in a row, same thing is happening every time.</p>\n<p>I also tried restarting the notebook also, still getting the same issue.</p>\n<p>Can someone please suggest how to solve this issue? thanks in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1710099,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/02/2022 18:43:43",
          "content": "<p><a href=\"https://www.kaggle.com/nanditab35\" target=\"_blank\">@nanditab35</a> When you use \"Commit\" are you selecting \"Save &amp; Run All\" or \"Quick Version\"?</p>\n<p>Quick version just copies your notebook as it looks in the editor, \"Save &amp; Run All\" runs the notebook from start to finish with the proper settings (TPU etc.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1710460,
          "author_name": "nanditab35",
          "author_url": "",
          "post_date": "03/03/2022 04:54:05",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/herbison\" target=\"_blank\">@herbison</a> for the suggestion. This actually solved my issue. Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 785674,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "03/25/2020 08:59:28",
      "content": "<p><a href=\"https://github.com/pytorch/xla/issues/1819\">Here</a>, I have added an issue as well on XLA Repo for some clarifications regarding the comment and the strategy as mentioned <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265#785608\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 785813,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "03/25/2020 12:18:40",
      "content": "<p>Did someone encounter kernels being stuck in commit mode?</p>",
      "votes": null,
      "replies": [
        {
          "id": 785816,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/25/2020 12:22:52",
          "content": "<p>Not recently but in the morning yes; &lt;2 mins kernel never got committed (non-TPU)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787542,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/26/2020 21:42:24",
          "content": "<p>I do have encountered the same thing twice which is quite weird. I had to restart the session and then again committed, and that worked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787726,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/27/2020 02:26:49",
          "content": "<p>Can you explain what it means to be stuck in commit? Do you mean it's running and never finishing? Feel free to send links to those commits to me (the viewer url is what I need).</p>\n\n<p>I can investigate what's up with them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787945,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/27/2020 08:42:16",
          "content": "<p>Hi, What happened earlier was that when I tried to execute all shells from the starting of the notebook till training shell, everything worked normally and I got first loss update from the TPU server, but after the first loss message, I have waited for 20-25 minutes and it was kind of stucked, I mean the training part. But when I restarted the session and re-run the same thing, It worked flawlessly. With the same code, training was properly working over TPU session and I was able to get loss updates after every 10 sec. Not sure if it was related to PyTorch XLA or Kaggle kernel. For the link, As there are multiple versions of that same notebook, So I do not remember, which version was it. :( </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788207,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/27/2020 13:33:27",
          "content": "<p>Okay thanks, glad it's working now though. If you run into it again feel free to send me the link and I'll take a look.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788380,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/27/2020 16:09:39",
          "content": "<p>It's weird that with the latest kernel I have submitted now, It is again getting stucked when I'm using no of epoch just 1 more than what I'm currently using. I'm confident that it should take max 1.5 Hrs to complete, as the current one with 1 less epoch is taking 2281.1s to complete, but instead, it's getting stucked. Due to this, commit keeps expiring after running for 3 Hrs limit and I had to decrease the epoch number by 1. Now unfortunately, my TPU quota also burned out for this week as my 2 commits wasted due to this thing. It would be much helpful if you can point me in the right direction about what could be wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788386,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/27/2020 16:17:08",
          "content": "<p>To clarify, is it actually 'getting stuck' or does the log indicate it's getting to the final epoch but that's not finishing within the 3h?</p>\n\n<p>I'm less of an expert on data science code, my expertise if on the code execution system for notebooks. I'd still need the viewer URL to be able to look deeper into your notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788396,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/27/2020 16:26:42",
          "content": "<p>Its getting stucked. I ran it during the interactive session to confirm after my 2 commits timed out and certainly that was the reason for the timeout. I have sent you an email for my kernel link. Please find it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788420,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "03/27/2020 16:51:23",
          "content": "<p>It's not clear to me that anything in our systems is actually getting stuck, I think the code is actually taking &gt; 3hrs and just not finishing in time. How many epochs do you run when it runs under 3 hours vs. over?</p>\n\n<p>If possible you might want to try using \"script\" mode instead of notebook mode and dump logs at critical points to get an idea of the timing.</p>\n\n<p>The logs for notebook mode are not that great at this point in time and can make it difficult to track progress. I'll see if any of the TPU experts on our team are aware of any other possible issues. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 788436,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/27/2020 17:06:53",
          "content": "<p>Well, there is just 1 epoch difference between the code which completed in just 2281.1 seconds Vs the code which didn't even completed in 10800 seconds which is definitely not adding up right. BTW, once I will get my TPU limit back in next week, I will try to run this through scripts as you suggested and will see if it persists or what. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 786209,
      "author_name": "p4rallax",
      "author_url": "",
      "post_date": "03/25/2020 17:45:13",
      "content": "<p>When I run the notebook , in the TPU details ,MXU efficiency shows 9.00%. Is that too low , and normal? Running 5 epochs , with a batch size of 64 and max_len = 192 had about train steps equal to about 1900, and the notebook timed out. Is there any way to optimize the MXU efficiency on TPU , or is there some mistake in my code? (I'm using Abhishek's notebook with 8 cores, which I wrote on my own so there is a possibility of having some mistakes there.) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "785655": "I see @philippsinger made a thread [here]() for TPU questions. I think it is worth it to create a separate thread regarding PyTorch XLA for TPU training.\n\nHere is also a list of useful PyTorch XLA resources:\n1. [PyTorch XLA repository](https://github.com/pytorch/xla)\n2. [API guide](https://github.com/pytorch/xla/blob/master/API_GUIDE.md)\n3. [Official documentation](https://pytorch.org/xla/)\n4. [Colab examples](https://github.com/pytorch/xla/tree/master/contrib/colab)\n5. [Kaggle discussion on PyTorch XLA support](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138137)\n6. @abhishek's [discussion post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265): [YouTube video](https://youtu.be/vvr_f-X_LaI), [training kernel](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training), [inference kernel](https://www.kaggle.com/abhishek/inference-of-bert-tpu-model-ml) \nEDIT: 8 cores working [here](https://www.kaggle.com/abhishek/bert-multi-lingual-tpu-training-8-cores)\n\nFeel free to comment with your questions and other resources available.",
    "785674": "[Here](https://github.com/pytorch/xla/issues/1819), I have added an issue as well on XLA Repo for some clarifications regarding the comment and the strategy as mentioned [here](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138265#785608)",
    "785813": "Did someone encounter kernels being stuck in commit mode?",
    "785816": "Not recently but in the morning yes; &lt;2 mins kernel never got committed (non-TPU)",
    "786209": "When I run the notebook , in the TPU details ,MXU efficiency shows 9.00%. Is that too low , and normal? Running 5 epochs , with a batch size of 64 and max_len = 192 had about train steps equal to about 1900, and the notebook timed out. Is there any way to optimize the MXU efficiency on TPU , or is there some mistake in my code? (I'm using Abhishek's notebook with 8 cores, which I wrote on my own so there is a possibility of having some mistakes there.)",
    "787542": "I do have encountered the same thing twice which is quite weird. I had to restart the session and then again committed, and that worked.",
    "787726": "Can you explain what it means to be stuck in commit? Do you mean it's running and never finishing? Feel free to send links to those commits to me (the viewer url is what I need).\n\nI can investigate what's up with them.",
    "787945": "Hi, What happened earlier was that when I tried to execute all shells from the starting of the notebook till training shell, everything worked normally and I got first loss update from the TPU server, but after the first loss message, I have waited for 20-25 minutes and it was kind of stucked, I mean the training part. But when I restarted the session and re-run the same thing, It worked flawlessly. With the same code, training was properly working over TPU session and I was able to get loss updates after every 10 sec. Not sure if it was related to PyTorch XLA or Kaggle kernel. For the link, As there are multiple versions of that same notebook, So I do not remember, which version was it. :(",
    "788207": "Okay thanks, glad it's working now though. If you run into it again feel free to send me the link and I'll take a look.",
    "788380": "It's weird that with the latest kernel I have submitted now, It is again getting stucked when I'm using no of epoch just 1 more than what I'm currently using. I'm confident that it should take max 1.5 Hrs to complete, as the current one with 1 less epoch is taking 2281.1s to complete, but instead, it's getting stucked. Due to this, commit keeps expiring after running for 3 Hrs limit and I had to decrease the epoch number by 1. Now unfortunately, my TPU quota also burned out for this week as my 2 commits wasted due to this thing. It would be much helpful if you can point me in the right direction about what could be wrong.",
    "788386": "To clarify, is it actually 'getting stuck' or does the log indicate it's getting to the final epoch but that's not finishing within the 3h?\n\nI'm less of an expert on data science code, my expertise if on the code execution system for notebooks. I'd still need the viewer URL to be able to look deeper into your notebook.",
    "788396": "Its getting stucked. I ran it during the interactive session to confirm after my 2 commits timed out and certainly that was the reason for the timeout. I have sent you an email for my kernel link. Please find it.",
    "788420": "It's not clear to me that anything in our systems is actually getting stuck, I think the code is actually taking &gt; 3hrs and just not finishing in time. How many epochs do you run when it runs under 3 hours vs. over?\n\nIf possible you might want to try using \"script\" mode instead of notebook mode and dump logs at critical points to get an idea of the timing.\n\nThe logs for notebook mode are not that great at this point in time and can make it difficult to track progress. I'll see if any of the TPU experts on our team are aware of any other possible issues.",
    "788436": "Well, there is just 1 epoch difference between the code which completed in just 2281.1 seconds Vs the code which didn't even completed in 10800 seconds which is definitely not adding up right. BTW, once I will get my TPU limit back in next week, I will try to run this through scripts as you suggested and will see if it persists or what. Thanks.",
    "1710094": "I am trying to commit a tensorflow notebook with TPU, but it is finishing very quickly within 5 seconds, and is running non-TPU, it is not at all training the model and not giving any output. I tried this 4-5 times in a row, same thing is happening every time.\n\nI also tried restarting the notebook also, still getting the same issue.\n\nCan someone please suggest how to solve this issue? thanks in advance!",
    "1710099": "nanditab35 When you use \"Commit\" are you selecting \"Save & Run All\" or \"Quick Version\"?\n\nQuick version just copies your notebook as it looks in the editor, \"Save & Run All\" runs the notebook from start to finish with the proper settings (TPU etc.)",
    "1710460": "Thanks @herbison for the suggestion. This actually solved my issue. Thanks again!"
  },
  "source": "meta"
}