{
  "id": 155216,
  "title": "GPU Problems",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155216",
  "author_name": "Sauman Das",
  "post_date": "2020-05-31T19:55:06.352000",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n\n<p>I have been experiencing some issues with my GPU specifically with this competition. The GPU worked when I was training (like I could see the bar go up), but on the last step, when it is supposed to also do the validation prediction, the gpu bar does not go up at all. The picture below should show what I am talking about. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fc3a82f02cbec57837e415260fd21533e%2Ftroubleshooting.PNG?generation=1590954835257960&amp;alt=media\" alt=\"\">\nIt just stays on that last step for a long time. Is anybody else experiencing something similar? Any advice on how to fix this would be appreciated. </p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 869185,
      "postDate": "2020-05-31T19:55:06.353Z",
      "content": "<p>Hello everyone,</p>\n\n<p>I have been experiencing some issues with my GPU specifically with this competition. The GPU worked when I was training (like I could see the bar go up), but on the last step, when it is supposed to also do the validation prediction, the gpu bar does not go up at all. The picture below should show what I am talking about. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fc3a82f02cbec57837e415260fd21533e%2Ftroubleshooting.PNG?generation=1590954835257960&amp;alt=media\" alt=\"\">\nIt just stays on that last step for a long time. Is anybody else experiencing something similar? Any advice on how to fix this would be appreciated. </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hello everyone,\n\nI have been experiencing some issues with my GPU specifically with this competition. The GPU worked when I was training (like I could see the bar go up), but on the last step, when it is supposed to also do the validation prediction, the gpu bar does not go up at all. The picture below should show what I am talking about. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fc3a82f02cbec57837e415260fd21533e%2Ftroubleshooting.PNG?generation=1590954835257960&amp;alt=media)\nIt just stays on that last step for a long time. Is anybody else experiencing something similar? Any advice on how to fix this would be appreciated. \n\nThanks!",
      "votes": 3
    },
    {
      "id": 893888,
      "postDate": "2020-06-20T04:09:09.703Z",
      "content": "<p>Thank you for responding <a href=\"/fireheart7\">@fireheart7</a> .\nI've been using PyTorch and the device name has been mentioned in the code already.\nmodel.to(device)\nBatch is also moved to the device for every iteration.\nThere is some other issue with the GPU.</p>",
      "rawMarkdown": "Thank you for responding @fireheart7 .\nI've been using PyTorch and the device name has been mentioned in the code already.\nmodel.to(device)\nBatch is also moved to the device for every iteration.\nThere is some other issue with the GPU.",
      "replies": [
        {
          "id": 893902,
          "postDate": "2020-06-20T04:30:09.490Z",
          "content": "<p>Did you convert the tensors to cuda tensors and compiled the model using CUDA before fitting it ? </p>",
          "rawMarkdown": "Did you convert the tensors to cuda tensors and compiled the model using CUDA before fitting it ? "
        },
        {
          "id": 893951,
          "postDate": "2020-06-20T05:00:00.153Z",
          "content": "<p>Yes, <a href=\"/fireheart7\">@fireheart7</a> I have done it.</p>",
          "rawMarkdown": "Yes, @fireheart7 I have done it."
        }
      ]
    },
    {
      "id": 893551,
      "postDate": "2020-06-19T17:58:11.480Z",
      "content": "<p>Hi!</p>\n\n<p>This is happening because you aren't mentioning the device to run on. Simply turning on the GPU button won't burn up the GPU. For the action to start, you have to mention the device name.</p>\n\n<p>with tf.device(\"/device:GPU:0\"):\n  history = model.fit(x_train, y_train,epochs = EPOCHS, verbose = 1,\n                     batch_size = BATCH_SIZE, validation_data = (x_val, y_val))</p>\n\n<p>Here GPU:0 is the device name, which the kaggle GPU we aim to burn for our training.</p>\n\n<p>Hope this helps,\nAll the best mate!</p>",
      "rawMarkdown": "Hi!\n\nThis is happening because you aren't mentioning the device to run on. Simply turning on the GPU button won't burn up the GPU. For the action to start, you have to mention the device name.\n\nwith tf.device(\"/device:GPU:0\"):\n  history = model.fit(x_train, y_train,epochs = EPOCHS, verbose = 1,\n                     batch_size = BATCH_SIZE, validation_data = (x_val, y_val))\n\nHere GPU:0 is the device name, which the kaggle GPU we aim to burn for our training.\n\nHope this helps,\nAll the best mate!",
      "replies": [
        {
          "id": 896046,
          "postDate": "2020-06-21T19:22:47.323Z",
          "content": "<p>Hi <a href=\"/fireheart7\">@fireheart7</a> ,</p>\n\n<p>I tried writing the program again using the explicit call to the GPU. However, for some reason it still doesn't seem to be utilizing the GPU (according to the bar and ETA). I'm not exactly sure what the problem is but here is another picture. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fcd8774b89c98ff00a32c78f2ecfb53e2%2Fimage.PNG?generation=1592770596401215&amp;alt=media\" alt=\"\"></p>\n\n<p>Thanks,\nSauman</p>",
          "rawMarkdown": "Hi @fireheart7 ,\n\nI tried writing the program again using the explicit call to the GPU. However, for some reason it still doesn't seem to be utilizing the GPU (according to the bar and ETA). I'm not exactly sure what the problem is but here is another picture. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fcd8774b89c98ff00a32c78f2ecfb53e2%2Fimage.PNG?generation=1592770596401215&amp;alt=media)\n\nThanks,\nSauman"
        }
      ]
    },
    {
      "id": 890418,
      "postDate": "2020-06-17T13:57:16.577Z",
      "content": "<p>I have also been facing the same problem, even worse, GPU doesn't work for the training process as well. \nI am using resnet-34 architechture.</p>",
      "rawMarkdown": "I have also been facing the same problem, even worse, GPU doesn't work for the training process as well. \nI am using resnet-34 architechture.\n"
    },
    {
      "id": 885787,
      "postDate": "2020-06-14T13:38:17.107Z",
      "content": "<p>exactly same thing has happend to me while training model using pytorch. please help if you find solution</p>",
      "rawMarkdown": "exactly same thing has happend to me while training model using pytorch. please help if you find solution\n"
    },
    {
      "id": 869331,
      "postDate": "2020-06-01T00:14:37.140Z",
      "content": "<p>The first 489 steps are training. That last 490th step is when your model predicts the entire validation set. So that last step takes longer and afterward it should display your validation AUC. If you wait long enough does that last step finish? Or does is wait forever?</p>",
      "rawMarkdown": "The first 489 steps are training. That last 490th step is when your model predicts the entire validation set. So that last step takes longer and afterward it should display your validation AUC. If you wait long enough does that last step finish? Or does is wait forever?",
      "replies": [
        {
          "id": 870087,
          "postDate": "2020-06-01T13:35:36.830Z",
          "content": "<p>Yesterday, for that last step, it took about 2.5 hours. I'm pretty sure the GPU was not being utilized for some reason (according to the bar at least). \nThanks!</p>",
          "rawMarkdown": "Yesterday, for that last step, it took about 2.5 hours. I'm pretty sure the GPU was not being utilized for some reason (according to the bar at least). \nThanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 935229,
      "postDate": "2020-07-19T07:16:34.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 893888,
      "author_name": "Jay_128",
      "author_url": "",
      "post_date": "2020-06-20T04:09:09.703000",
      "content": "<p>Thank you for responding <a href=\"/fireheart7\">@fireheart7</a> .\nI've been using PyTorch and the device name has been mentioned in the code already.\nmodel.to(device)\nBatch is also moved to the device for every iteration.\nThere is some other issue with the GPU.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 893902,
          "author_name": "Aditya Baurai",
          "author_url": "",
          "post_date": "2020-06-20T04:30:09.490000",
          "content": "<p>Did you convert the tensors to cuda tensors and compiled the model using CUDA before fitting it ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 893951,
          "author_name": "Jay_128",
          "author_url": "",
          "post_date": "2020-06-20T05:00:00.153000",
          "content": "<p>Yes, <a href=\"/fireheart7\">@fireheart7</a> I have done it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 893551,
      "author_name": "Aditya Baurai",
      "author_url": "",
      "post_date": "2020-06-19T17:58:11.480000",
      "content": "<p>Hi!</p>\n\n<p>This is happening because you aren't mentioning the device to run on. Simply turning on the GPU button won't burn up the GPU. For the action to start, you have to mention the device name.</p>\n\n<p>with tf.device(\"/device:GPU:0\"):\n  history = model.fit(x_train, y_train,epochs = EPOCHS, verbose = 1,\n                     batch_size = BATCH_SIZE, validation_data = (x_val, y_val))</p>\n\n<p>Here GPU:0 is the device name, which the kaggle GPU we aim to burn for our training.</p>\n\n<p>Hope this helps,\nAll the best mate!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 896046,
          "author_name": "Sauman Das",
          "author_url": "",
          "post_date": "2020-06-21T19:22:47.323000",
          "content": "<p>Hi <a href=\"/fireheart7\">@fireheart7</a> ,</p>\n\n<p>I tried writing the program again using the explicit call to the GPU. However, for some reason it still doesn't seem to be utilizing the GPU (according to the bar and ETA). I'm not exactly sure what the problem is but here is another picture. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fcd8774b89c98ff00a32c78f2ecfb53e2%2Fimage.PNG?generation=1592770596401215&amp;alt=media\" alt=\"\"></p>\n\n<p>Thanks,\nSauman</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 890418,
      "author_name": "Jay_128",
      "author_url": "",
      "post_date": "2020-06-17T13:57:16.577000",
      "content": "<p>I have also been facing the same problem, even worse, GPU doesn't work for the training process as well. \nI am using resnet-34 architechture.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 885787,
      "author_name": "manpreet singh",
      "author_url": "",
      "post_date": "2020-06-14T13:38:17.107000",
      "content": "<p>exactly same thing has happend to me while training model using pytorch. please help if you find solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 869331,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-01T00:14:37.140000",
      "content": "<p>The first 489 steps are training. That last 490th step is when your model predicts the entire validation set. So that last step takes longer and afterward it should display your validation AUC. If you wait long enough does that last step finish? Or does is wait forever?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 870087,
          "author_name": "Sauman Das",
          "author_url": "",
          "post_date": "2020-06-01T13:35:36.830000",
          "content": "<p>Yesterday, for that last step, it took about 2.5 hours. I'm pretty sure the GPU was not being utilized for some reason (according to the bar at least). \nThanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 935229,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-19T07:16:34.710000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "869185": "Hello everyone,\n\nI have been experiencing some issues with my GPU specifically with this competition. The GPU worked when I was training (like I could see the bar go up), but on the last step, when it is supposed to also do the validation prediction, the gpu bar does not go up at all. The picture below should show what I am talking about. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4670843%2Fc3a82f02cbec57837e415260fd21533e%2Ftroubleshooting.PNG?generation=1590954835257960&amp;alt=media)\nIt just stays on that last step for a long time. Is anybody else experiencing something similar? Any advice on how to fix this would be appreciated. \n\nThanks!",
    "893888": "Thank you for responding @fireheart7 .\nI've been using PyTorch and the device name has been mentioned in the code already.\nmodel.to(device)\nBatch is also moved to the device for every iteration.\nThere is some other issue with the GPU.",
    "893551": "Hi!\n\nThis is happening because you aren't mentioning the device to run on. Simply turning on the GPU button won't burn up the GPU. For the action to start, you have to mention the device name.\n\nwith tf.device(\"/device:GPU:0\"):\n  history = model.fit(x_train, y_train,epochs = EPOCHS, verbose = 1,\n                     batch_size = BATCH_SIZE, validation_data = (x_val, y_val))\n\nHere GPU:0 is the device name, which the kaggle GPU we aim to burn for our training.\n\nHope this helps,\nAll the best mate!",
    "890418": "I have also been facing the same problem, even worse, GPU doesn't work for the training process as well. \nI am using resnet-34 architechture.\n",
    "885787": "exactly same thing has happend to me while training model using pytorch. please help if you find solution\n",
    "869331": "The first 489 steps are training. That last 490th step is when your model predicts the entire validation set. So that last step takes longer and afterward it should display your validation AUC. If you wait long enough does that last step finish? Or does is wait forever?",
    "935229": ""
  }
}