{
  "id": 213439,
  "title": "TPU training broken?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/213439",
  "author_name": "",
  "post_date": "2021-01-22T21:11:44.090483Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is it just for me or is TPU training this last few days broken? There seems to be an issue with paralleldataloader/tqdm i couldn't exactly figure out what is wrong with my code which was working at the start of the week, and i tried forking a couple notebook which use TPU's and they don't seem to run as well. Anyone ran into this issue? Any workaround?</p>\n<p>I did a quicktest on google colab as well and the same code for pytorch doesn't seem to be working there either.</p>",
  "messages": [
    {
      "id": "1165293",
      "postDate": "01/22/2021 21:11:44",
      "content": "<p>Is it just for me or is TPU training this last few days broken? There seems to be an issue with paralleldataloader/tqdm i couldn't exactly figure out what is wrong with my code which was working at the start of the week, and i tried forking a couple notebook which use TPU's and they don't seem to run as well. Anyone ran into this issue? Any workaround?</p>\n<p>I did a quicktest on google colab as well and the same code for pytorch doesn't seem to be working there either.</p>",
      "rawMarkdown": "Is it just for me or is TPU training this last few days broken? There seems to be an issue with paralleldataloader/tqdm i couldn't exactly figure out what is wrong with my code which was working at the start of the week, and i tried forking a couple notebook which use TPU's and they don't seem to run as well. Anyone ran into this issue? Any workaround?\n\nI did a quicktest on google colab as well and the same code for pytorch doesn't seem to be working there either.",
      "votes": null
    },
    {
      "id": "1165297",
      "postDate": "01/22/2021 21:15:42",
      "content": "<p>what do you mean by <strong>doesn't seem to be working?</strong></p>\n<p>please be specific,what error you are getting exactly?<br>\nplease share the full error log </p>",
      "rawMarkdown": "what do you mean by **doesn't seem to be working?**\n\n\nplease be specific,what error you are getting exactly?\nplease share the full error log",
      "votes": null
    },
    {
      "id": "1165299",
      "postDate": "01/22/2021 21:19:28",
      "content": "<p>It doesn't explicitly spits an error. When the function reaches for instance:<code>for i,images,labels in tqdm(enumerate(parallel_train_loader)):</code> it doesn't go further, it simply stays there indefinitely</p>",
      "rawMarkdown": "It doesn't explicitly spits an error. When the function reaches for instance:`for i,images,labels in tqdm(enumerate(parallel_train_loader)):` it doesn't go further, it simply stays there indefinitely",
      "votes": null
    },
    {
      "id": "1165300",
      "postDate": "01/22/2021 21:22:15",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> i tested on a fork of your faster pytorch notebook and it happened there as well. 15 min plus without the code starting the training loop.</p>",
      "rawMarkdown": "mobassir i tested on a fork of your faster pytorch notebook and it happened there as well. 15 min plus without the code starting the training loop.",
      "votes": null
    },
    {
      "id": "1165301",
      "postDate": "01/22/2021 21:25:43",
      "content": "<p>try one things : <br>\n<strong>set num_workers = 0 and train,wait for 30 minutes at least</strong> and then let me know if it is still getting stuck or not, i will get tpu quota back tomorrow,if it doesn't solve your issue,i will try and see how to solve this issue,thank you</p>",
      "rawMarkdown": "try one things : \n**set num_workers = 0 and train,wait for 30 minutes at least** and then let me know if it is still getting stuck or not, i will get tpu quota back tomorrow,if it doesn't solve your issue,i will try and see how to solve this issue,thank you",
      "votes": null
    },
    {
      "id": "1165502",
      "postDate": "01/23/2021 03:06:04",
      "content": "<p>Thanks for answering me, setting num_workers to 0 seemed to work for me. Happy kaggling</p>",
      "rawMarkdown": "Thanks for answering me, setting num_workers to 0 seemed to work for me. Happy kaggling",
      "votes": null
    },
    {
      "id": "1165742",
      "postDate": "01/23/2021 07:51:10",
      "content": "<p><a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> i just checked now and even with num_workers = 4 my notebook faster pytorch is working(just checked now),please check again if it is working for you now or not</p>",
      "rawMarkdown": "capiru i just checked now and even with num_workers = 4 my notebook faster pytorch is working(just checked now),please check again if it is working for you now or not",
      "votes": null
    },
    {
      "id": "1174017",
      "postDate": "01/28/2021 08:19:50",
      "content": "<p>i have the same issue too, and i change num_workers = 0 does not work. What’s  more, in interactive session it works, but when save&amp;commit, it run indefinitely until time up…..</p>",
      "rawMarkdown": "i have the same issue too, and i change num_workers = 0 does not work. What’s  more, in interactive session it works, but when save&commit, it run indefinitely until time up.....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1165297,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "01/22/2021 21:15:42",
      "content": "<p>what do you mean by <strong>doesn't seem to be working?</strong></p>\n<p>please be specific,what error you are getting exactly?<br>\nplease share the full error log </p>",
      "votes": null,
      "replies": [
        {
          "id": 1165299,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/22/2021 21:19:28",
          "content": "<p>It doesn't explicitly spits an error. When the function reaches for instance:<code>for i,images,labels in tqdm(enumerate(parallel_train_loader)):</code> it doesn't go further, it simply stays there indefinitely</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1165300,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/22/2021 21:22:15",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> i tested on a fork of your faster pytorch notebook and it happened there as well. 15 min plus without the code starting the training loop.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1165301,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/22/2021 21:25:43",
          "content": "<p>try one things : <br>\n<strong>set num_workers = 0 and train,wait for 30 minutes at least</strong> and then let me know if it is still getting stuck or not, i will get tpu quota back tomorrow,if it doesn't solve your issue,i will try and see how to solve this issue,thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1165502,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "01/23/2021 03:06:04",
          "content": "<p>Thanks for answering me, setting num_workers to 0 seemed to work for me. Happy kaggling</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1165742,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/23/2021 07:51:10",
          "content": "<p><a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> i just checked now and even with num_workers = 4 my notebook faster pytorch is working(just checked now),please check again if it is working for you now or not</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1174017,
      "author_name": "yifannir",
      "author_url": "",
      "post_date": "01/28/2021 08:19:50",
      "content": "<p>i have the same issue too, and i change num_workers = 0 does not work. What’s  more, in interactive session it works, but when save&amp;commit, it run indefinitely until time up…..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1165293": "Is it just for me or is TPU training this last few days broken? There seems to be an issue with paralleldataloader/tqdm i couldn't exactly figure out what is wrong with my code which was working at the start of the week, and i tried forking a couple notebook which use TPU's and they don't seem to run as well. Anyone ran into this issue? Any workaround?\n\nI did a quicktest on google colab as well and the same code for pytorch doesn't seem to be working there either.",
    "1165297": "what do you mean by **doesn't seem to be working?**\n\n\nplease be specific,what error you are getting exactly?\nplease share the full error log",
    "1165299": "It doesn't explicitly spits an error. When the function reaches for instance:`for i,images,labels in tqdm(enumerate(parallel_train_loader)):` it doesn't go further, it simply stays there indefinitely",
    "1165300": "mobassir i tested on a fork of your faster pytorch notebook and it happened there as well. 15 min plus without the code starting the training loop.",
    "1165301": "try one things : \n**set num_workers = 0 and train,wait for 30 minutes at least** and then let me know if it is still getting stuck or not, i will get tpu quota back tomorrow,if it doesn't solve your issue,i will try and see how to solve this issue,thank you",
    "1165502": "Thanks for answering me, setting num_workers to 0 seemed to work for me. Happy kaggling",
    "1165742": "capiru i just checked now and even with num_workers = 4 my notebook faster pytorch is working(just checked now),please check again if it is working for you now or not",
    "1174017": "i have the same issue too, and i change num_workers = 0 does not work. What’s  more, in interactive session it works, but when save&commit, it run indefinitely until time up....."
  },
  "source": "meta"
}