{
  "id": 201505,
  "title": "Different GPU is allocated every week?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201505",
  "author_name": "",
  "post_date": "2020-12-05T10:53:19.840006700Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>--Update--<br>\nTraining speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away…</p>\n<hr>\n<p>I run a kaggle notebook which I made before, and found xgboost inference with GPU becomes very slow and result got worse despite setting seed, comparing to previous run (a week ago). I'm sure that condition is totally same because I just fork the notebook and run it.</p>\n<p>My guess about inference speed is different spec gpu might be allocated every week, similar to google colab GPU. But this is just my guess. Anyone knows why?</p>",
  "messages": [
    {
      "id": "1102818",
      "postDate": "12/05/2020 10:53:19",
      "content": "<p>--Update--<br>\nTraining speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away…</p>\n<hr>\n<p>I run a kaggle notebook which I made before, and found xgboost inference with GPU becomes very slow and result got worse despite setting seed, comparing to previous run (a week ago). I'm sure that condition is totally same because I just fork the notebook and run it.</p>\n<p>My guess about inference speed is different spec gpu might be allocated every week, similar to google colab GPU. But this is just my guess. Anyone knows why?</p>",
      "rawMarkdown": "Update--\nTraining speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away...\n\n--------\nI run a kaggle notebook which I made before, and found xgboost inference with GPU becomes very slow and result got worse despite setting seed, comparing to previous run (a week ago). I'm sure that condition is totally same because I just fork the notebook and run it.\n\nMy guess about inference speed is different spec gpu might be allocated every week, similar to google colab GPU. But this is just my guess. Anyone knows why?",
      "votes": null
    },
    {
      "id": "1102858",
      "postDate": "12/05/2020 12:02:39",
      "content": "<p>Kaggle introduced a dynamic weekly GPU quota allocation since some time.</p>",
      "rawMarkdown": "Kaggle introduced a dynamic weekly GPU quota allocation since some time.",
      "votes": null
    },
    {
      "id": "1102875",
      "postDate": "12/05/2020 12:33:43",
      "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> <br>\nYour \"quota allocation\" means different spec GPU? (I know hours of quota is different every week) <br>\nIf you know related page, I'd like to know it</p>",
      "rawMarkdown": "stecasasso \nYour \"quota allocation\" means different spec GPU? (I know hours of quota is different every week) \nIf you know related page, I'd like to know it",
      "votes": null
    },
    {
      "id": "1103200",
      "postDate": "12/05/2020 17:58:44",
      "content": "<p>I posted a similiar query sometime back when I noticed my exact same code having different run times (and as a result I was have timeouts on my notebooks)</p>\n<p>I was told that have only a single spec GPU. Not sure where the variability comes from, but it is probably a temporary thing. What %age variation are you observing ?</p>",
      "rawMarkdown": "I posted a similiar query sometime back when I noticed my exact same code having different run times (and as a result I was have timeouts on my notebooks)\n\nI was told that have only a single spec GPU. Not sure where the variability comes from, but it is probably a temporary thing. What %age variation are you observing ?",
      "votes": null
    },
    {
      "id": "1103525",
      "postDate": "12/06/2020 02:21:46",
      "content": "<p><a href=\"https://www.kaggle.com/watzisname\" target=\"_blank\">@watzisname</a> <br>\nIt took only about 1min to train in previous time, while this time took about 2h. This difference becomes heavy burden for me. It is questionable for GPU to work, but GPU time is properly decreasing and cudf works so GPU seems in use. Despite temporary, it is fatal to wait next week..</p>",
      "rawMarkdown": "watzisname \nIt took only about 1min to train in previous time, while this time took about 2h. This difference becomes heavy burden for me. It is questionable for GPU to work, but GPU time is properly decreasing and cudf works so GPU seems in use. Despite temporary, it is fatal to wait next week..",
      "votes": null
    },
    {
      "id": "1103732",
      "postDate": "12/06/2020 07:37:37",
      "content": "<p>Sorry, I think I misunderstood your question. No, I am not aware of different type of GPU. I guess one can print the device name at runtime to check</p>",
      "rawMarkdown": "Sorry, I think I misunderstood your question. No, I am not aware of different type of GPU. I guess one can print the device name at runtime to check",
      "votes": null
    },
    {
      "id": "1103848",
      "postDate": "12/06/2020 11:12:34",
      "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> <br>\nOK. I haven't checked device name, so I'll try to do <code>$ nvidia-smi</code>. Thank you!</p>",
      "rawMarkdown": "stecasasso \nOK. I haven't checked device name, so I'll try to do `$ nvidia-smi`. Thank you!",
      "votes": null
    },
    {
      "id": "1103898",
      "postDate": "12/06/2020 12:14:15",
      "content": "<p>you can also do:</p>\n<p><code>\nimport torch; torch.cuda.get_device_name()\n</code></p>",
      "rawMarkdown": "you can also do:\n\n`\nimport torch; torch.cuda.get_device_name()\n`",
      "votes": null
    },
    {
      "id": "1103921",
      "postDate": "12/06/2020 12:27:02",
      "content": "<p>Thanks, I'll also try it </p>\n<p>===update===<br>\nI'll leave it</p>\n<p>Tesla P100-PCIE-16GB<br>\nSun Dec  6 14:17:21 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  Tesla P100-PCIE…  Off  | 00000000:00:04.0 Off |                    0 |<br>\n| N/A   36C    P0    26W / 250W |      2MiB / 16280MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|  No running processes found                                                 |<br>\n+-----------------------------------------------------------------------------+</p>",
      "rawMarkdown": "Thanks, I'll also try it \n\n===update===\nI'll leave it\n\nTesla P100-PCIE-16GB\nSun Dec  6 14:17:21 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  Tesla P100-PCIE...  Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   36C    P0    26W / 250W |      2MiB / 16280MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|  No running processes found                                                 |\n+-----------------------------------------------------------------------------+",
      "votes": null
    },
    {
      "id": "1109757",
      "postDate": "12/12/2020 03:25:33",
      "content": "<p>Training speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away.</p>\n<p>Tesla P100-PCIE-16GB<br>\nSat Dec 12 03:23:07 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  Tesla P100-PCIE…  Off  | 00000000:00:04.0 Off |                    0 |<br>\n| N/A   38C    P0    28W / 250W |      2MiB / 16280MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|  No running processes found                                                 |<br>\n+-----------------------------------------------------------------------------+</p>",
      "rawMarkdown": "Training speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away.\n\nTesla P100-PCIE-16GB\nSat Dec 12 03:23:07 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  Tesla P100-PCIE...  Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   38C    P0    28W / 250W |      2MiB / 16280MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|  No running processes found                                                 |\n+-----------------------------------------------------------------------------+",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102858,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "12/05/2020 12:02:39",
      "content": "<p>Kaggle introduced a dynamic weekly GPU quota allocation since some time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102875,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/05/2020 12:33:43",
          "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> <br>\nYour \"quota allocation\" means different spec GPU? (I know hours of quota is different every week) <br>\nIf you know related page, I'd like to know it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103732,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/06/2020 07:37:37",
          "content": "<p>Sorry, I think I misunderstood your question. No, I am not aware of different type of GPU. I guess one can print the device name at runtime to check</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103848,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/06/2020 11:12:34",
          "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> <br>\nOK. I haven't checked device name, so I'll try to do <code>$ nvidia-smi</code>. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103898,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/06/2020 12:14:15",
          "content": "<p>you can also do:</p>\n<p><code>\nimport torch; torch.cuda.get_device_name()\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103921,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/06/2020 12:27:02",
          "content": "<p>Thanks, I'll also try it </p>\n<p>===update===<br>\nI'll leave it</p>\n<p>Tesla P100-PCIE-16GB<br>\nSun Dec  6 14:17:21 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  Tesla P100-PCIE…  Off  | 00000000:00:04.0 Off |                    0 |<br>\n| N/A   36C    P0    26W / 250W |      2MiB / 16280MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|  No running processes found                                                 |<br>\n+-----------------------------------------------------------------------------+</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109757,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/12/2020 03:25:33",
          "content": "<p>Training speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away.</p>\n<p>Tesla P100-PCIE-16GB<br>\nSat Dec 12 03:23:07 2020       <br>\n+-----------------------------------------------------------------------------+<br>\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |<br>\n|-------------------------------+----------------------+----------------------+<br>\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |<br>\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |<br>\n|                               |                      |               MIG M. |<br>\n|===============================+======================+======================|<br>\n|   0  Tesla P100-PCIE…  Off  | 00000000:00:04.0 Off |                    0 |<br>\n| N/A   38C    P0    28W / 250W |      2MiB / 16280MiB |      0%      Default |<br>\n|                               |                      |                  N/A |<br>\n+-------------------------------+----------------------+----------------------+</p>\n<p>+-----------------------------------------------------------------------------+<br>\n| Processes:                                                                  |<br>\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |<br>\n|        ID   ID                                                   Usage      |<br>\n|=============================================================================|<br>\n|  No running processes found                                                 |<br>\n+-----------------------------------------------------------------------------+</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103200,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "12/05/2020 17:58:44",
      "content": "<p>I posted a similiar query sometime back when I noticed my exact same code having different run times (and as a result I was have timeouts on my notebooks)</p>\n<p>I was told that have only a single spec GPU. Not sure where the variability comes from, but it is probably a temporary thing. What %age variation are you observing ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103525,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/06/2020 02:21:46",
          "content": "<p><a href=\"https://www.kaggle.com/watzisname\" target=\"_blank\">@watzisname</a> <br>\nIt took only about 1min to train in previous time, while this time took about 2h. This difference becomes heavy burden for me. It is questionable for GPU to work, but GPU time is properly decreasing and cudf works so GPU seems in use. Despite temporary, it is fatal to wait next week..</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102818": "Update--\nTraining speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away...\n\n--------\nI run a kaggle notebook which I made before, and found xgboost inference with GPU becomes very slow and result got worse despite setting seed, comparing to previous run (a week ago). I'm sure that condition is totally same because I just fork the notebook and run it.\n\nMy guess about inference speed is different spec gpu might be allocated every week, similar to google colab GPU. But this is just my guess. Anyone knows why?",
    "1102858": "Kaggle introduced a dynamic weekly GPU quota allocation since some time.",
    "1102875": "stecasasso \nYour \"quota allocation\" means different spec GPU? (I know hours of quota is different every week) \nIf you know related page, I'd like to know it",
    "1103200": "I posted a similiar query sometime back when I noticed my exact same code having different run times (and as a result I was have timeouts on my notebooks)\n\nI was told that have only a single spec GPU. Not sure where the variability comes from, but it is probably a temporary thing. What %age variation are you observing ?",
    "1103525": "watzisname \nIt took only about 1min to train in previous time, while this time took about 2h. This difference becomes heavy burden for me. It is questionable for GPU to work, but GPU time is properly decreasing and cudf works so GPU seems in use. Despite temporary, it is fatal to wait next week..",
    "1103732": "Sorry, I think I misunderstood your question. No, I am not aware of different type of GPU. I guess one can print the device name at runtime to check",
    "1103848": "stecasasso \nOK. I haven't checked device name, so I'll try to do `$ nvidia-smi`. Thank you!",
    "1103898": "you can also do:\n\n`\nimport torch; torch.cuda.get_device_name()\n`",
    "1103921": "Thanks, I'll also try it \n\n===update===\nI'll leave it\n\nTesla P100-PCIE-16GB\nSun Dec  6 14:17:21 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  Tesla P100-PCIE...  Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   36C    P0    26W / 250W |      2MiB / 16280MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|  No running processes found                                                 |\n+-----------------------------------------------------------------------------+",
    "1109757": "Training speed comes back in this week and GPU spec is almost no different. I don't know why, but problem goes away.\n\nTesla P100-PCIE-16GB\nSat Dec 12 03:23:07 2020       \n+-----------------------------------------------------------------------------+\n| NVIDIA-SMI 450.51.06    Driver Version: 450.51.06    CUDA Version: 11.0     |\n|-------------------------------+----------------------+----------------------+\n| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |\n|                               |                      |               MIG M. |\n|===============================+======================+======================|\n|   0  Tesla P100-PCIE...  Off  | 00000000:00:04.0 Off |                    0 |\n| N/A   38C    P0    28W / 250W |      2MiB / 16280MiB |      0%      Default |\n|                               |                      |                  N/A |\n+-------------------------------+----------------------+----------------------+\n                                                                               \n+-----------------------------------------------------------------------------+\n| Processes:                                                                  |\n|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |\n|        ID   ID                                                   Usage      |\n|=============================================================================|\n|  No running processes found                                                 |\n+-----------------------------------------------------------------------------+"
  },
  "source": "meta"
}