{
  "id": 398275,
  "title": "I dont know if GPU is working",
  "url": "/competitions/asl-signs/discussion/398275",
  "author_name": "",
  "post_date": "2023-03-29T09:46:09.788479600Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>The GPU usage bar on the Kaggle kernel is being shown as 6%, while the CPU usage is completely filled ,even 181%(red).</p>\n<p>Hence I don't think  if gpu is working and how to deal with it.</p>\n<p>Can you please look into this and help me resolve this?</p>\n<p>Thanks,</p>",
  "messages": [
    {
      "id": "2201425",
      "postDate": "03/29/2023 09:46:09",
      "content": "<p>Hi,</p>\n<p>The GPU usage bar on the Kaggle kernel is being shown as 6%, while the CPU usage is completely filled ,even 181%(red).</p>\n<p>Hence I don't think  if gpu is working and how to deal with it.</p>\n<p>Can you please look into this and help me resolve this?</p>\n<p>Thanks,</p>",
      "rawMarkdown": "Hi,\n\nThe GPU usage bar on the Kaggle kernel is being shown as 6%, while the CPU usage is completely filled ,even 181%(red).\n\nHence I don't think  if gpu is working and how to deal with it.\n\nCan you please look into this and help me resolve this?\n\nThanks,",
      "votes": null
    },
    {
      "id": "2201493",
      "postDate": "03/29/2023 11:06:32",
      "content": "<p>I assume you have a data loading bottleneck. It is a very common case - your GPU process your data much faster than you are able to load it. That's why you have a high CPU load and low GPU util. Do you parse each parquet file during training?</p>",
      "rawMarkdown": "I assume you have a data loading bottleneck. It is a very common case - your GPU process your data much faster than you are able to load it. That's why you have a high CPU load and low GPU util. Do you parse each parquet file during training?",
      "votes": null
    },
    {
      "id": "2204819",
      "postDate": "04/01/2023 02:03:00",
      "content": "<p>I had the same problem. And I pass a batch of parquet files through dataset API. I am not quite sure what causes the error.</p>",
      "rawMarkdown": "I had the same problem. And I pass a batch of parquet files through dataset API. I am not quite sure what causes the error.",
      "votes": null
    },
    {
      "id": "2205248",
      "postDate": "04/01/2023 11:18:20",
      "content": "<p>Iterating over all the parquet files takes approximately 35 minutes. It is really slow and your model is unlikely to take half an hour per epoch. I would recommend to </p>\n<ol>\n<li>Migrate to .tfrecords. This allows me to reduce dataset iteration from 35m to 1m30s, which is significant. </li>\n<li>Consider using tf.data.Dataset.cache(), which cache read the dataset and uses that cache for the rest of epochs. I achieve up to 6s per epoch without any preprocessed dataset on the disc.</li>\n</ol>",
      "rawMarkdown": "Iterating over all the parquet files takes approximately 35 minutes. It is really slow and your model is unlikely to take half an hour per epoch. I would recommend to \n\n1. Migrate to .tfrecords. This allows me to reduce dataset iteration from 35m to 1m30s, which is significant. \n2. Consider using tf.data.Dataset.cache(), which cache read the dataset and uses that cache for the rest of epochs. I achieve up to 6s per epoch without any preprocessed dataset on the disc.",
      "votes": null
    },
    {
      "id": "2208283",
      "postDate": "04/04/2023 00:16:20",
      "content": "<p>GPU is ~only~ used for matrix multiplication. When loading and copying data that will be using the CPU and RAM. Kaggle's free notebooks are rather constrained in this regard. If possible download the data and do the preprocessing locally - creating .npy files or .tfrecords as <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> suggested. Then you can upload these files for the training and make full use of Kaggle's free GPU computation </p>",
      "rawMarkdown": "GPU is ~only~ used for matrix multiplication. When loading and copying data that will be using the CPU and RAM. Kaggle's free notebooks are rather constrained in this regard. If possible download the data and do the preprocessing locally - creating .npy files or .tfrecords as @meowmeowmeowmeowmeow suggested. Then you can upload these files for the training and make full use of Kaggle's free GPU computation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2201493,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "03/29/2023 11:06:32",
      "content": "<p>I assume you have a data loading bottleneck. It is a very common case - your GPU process your data much faster than you are able to load it. That's why you have a high CPU load and low GPU util. Do you parse each parquet file during training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2204819,
          "author_name": "gssdatamanager",
          "author_url": "",
          "post_date": "04/01/2023 02:03:00",
          "content": "<p>I had the same problem. And I pass a batch of parquet files through dataset API. I am not quite sure what causes the error.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2205248,
              "author_name": "meowmeowmeowmeowmeow",
              "author_url": "",
              "post_date": "04/01/2023 11:18:20",
              "content": "<p>Iterating over all the parquet files takes approximately 35 minutes. It is really slow and your model is unlikely to take half an hour per epoch. I would recommend to </p>\n<ol>\n<li>Migrate to .tfrecords. This allows me to reduce dataset iteration from 35m to 1m30s, which is significant. </li>\n<li>Consider using tf.data.Dataset.cache(), which cache read the dataset and uses that cache for the rest of epochs. I achieve up to 6s per epoch without any preprocessed dataset on the disc.</li>\n</ol>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2208283,
      "author_name": "blankfish",
      "author_url": "",
      "post_date": "04/04/2023 00:16:20",
      "content": "<p>GPU is ~only~ used for matrix multiplication. When loading and copying data that will be using the CPU and RAM. Kaggle's free notebooks are rather constrained in this regard. If possible download the data and do the preprocessing locally - creating .npy files or .tfrecords as <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> suggested. Then you can upload these files for the training and make full use of Kaggle's free GPU computation </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2201425": "Hi,\n\nThe GPU usage bar on the Kaggle kernel is being shown as 6%, while the CPU usage is completely filled ,even 181%(red).\n\nHence I don't think  if gpu is working and how to deal with it.\n\nCan you please look into this and help me resolve this?\n\nThanks,",
    "2201493": "I assume you have a data loading bottleneck. It is a very common case - your GPU process your data much faster than you are able to load it. That's why you have a high CPU load and low GPU util. Do you parse each parquet file during training?",
    "2204819": "I had the same problem. And I pass a batch of parquet files through dataset API. I am not quite sure what causes the error.",
    "2205248": "Iterating over all the parquet files takes approximately 35 minutes. It is really slow and your model is unlikely to take half an hour per epoch. I would recommend to \n\n1. Migrate to .tfrecords. This allows me to reduce dataset iteration from 35m to 1m30s, which is significant. \n2. Consider using tf.data.Dataset.cache(), which cache read the dataset and uses that cache for the rest of epochs. I achieve up to 6s per epoch without any preprocessed dataset on the disc.",
    "2208283": "GPU is ~only~ used for matrix multiplication. When loading and copying data that will be using the CPU and RAM. Kaggle's free notebooks are rather constrained in this regard. If possible download the data and do the preprocessing locally - creating .npy files or .tfrecords as @meowmeowmeowmeowmeow suggested. Then you can upload these files for the training and make full use of Kaggle's free GPU computation"
  },
  "source": "meta"
}