{
  "id": 207977,
  "title": "Very slow training even with GPU",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207977",
  "author_name": "",
  "post_date": "2021-01-01T09:26:40.196020Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm new to Deep Learning, and it is my first time using kaggle kernel (actually first time ever training online), and even if I used GPU acceleration, training speed is still slow.</p>\n<p>With (150,150,3) * 64 batch_size input, tensorflow pretrained ResNet50V2(excluding top, freezing all layers) + 2 untrained fully connected layers (256 and 5) + some other layers (batchnorm, activation and dropout), it takes 300-400s to train each epoch on kaggle kernel. However, on my <strong>laptop</strong> with 1660ti + i5-9300h, it takes only 230-235s per epoch to train with exact same model and data (and same code except the paths to files). </p>\n<p>Tesla p100 is no doubt way better than a laptop 1660ti, but how can the fastest epoch on kaggle kernel be slower than the slowest epoch on my pc? Can someone point out possible problems?<br>\nThanks!</p>\n<p>ps. it is not a temporary issue since I tried several times last three days, but it never gets fast, I even tried explicitly call with tf.device('/GPU:0'): nothing changed. And the gpu meter at the top is  0% at all time (maybe it is not even using gpu at all?)<br>\n and using tpu only gives me \"No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node\"</p>\n<p>update: I noticed that kernel cpu is always over 80 percent, maybe IO is making it slow? I am using flow_from_directory, which seems like how many people load data, if this is the problem, what is a better work around for this?</p>",
  "messages": [
    {
      "id": "1134411",
      "postDate": "01/01/2021 09:26:40",
      "content": "<p>I'm new to Deep Learning, and it is my first time using kaggle kernel (actually first time ever training online), and even if I used GPU acceleration, training speed is still slow.</p>\n<p>With (150,150,3) * 64 batch_size input, tensorflow pretrained ResNet50V2(excluding top, freezing all layers) + 2 untrained fully connected layers (256 and 5) + some other layers (batchnorm, activation and dropout), it takes 300-400s to train each epoch on kaggle kernel. However, on my <strong>laptop</strong> with 1660ti + i5-9300h, it takes only 230-235s per epoch to train with exact same model and data (and same code except the paths to files). </p>\n<p>Tesla p100 is no doubt way better than a laptop 1660ti, but how can the fastest epoch on kaggle kernel be slower than the slowest epoch on my pc? Can someone point out possible problems?<br>\nThanks!</p>\n<p>ps. it is not a temporary issue since I tried several times last three days, but it never gets fast, I even tried explicitly call with tf.device('/GPU:0'): nothing changed. And the gpu meter at the top is  0% at all time (maybe it is not even using gpu at all?)<br>\n and using tpu only gives me \"No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node\"</p>\n<p>update: I noticed that kernel cpu is always over 80 percent, maybe IO is making it slow? I am using flow_from_directory, which seems like how many people load data, if this is the problem, what is a better work around for this?</p>",
      "rawMarkdown": "I'm new to Deep Learning, and it is my first time using kaggle kernel (actually first time ever training online), and even if I used GPU acceleration, training speed is still slow.\n\nWith (150,150,3) * 64 batch_size input, tensorflow pretrained ResNet50V2(excluding top, freezing all layers) + 2 untrained fully connected layers (256 and 5) + some other layers (batchnorm, activation and dropout), it takes 300-400s to train each epoch on kaggle kernel. However, on my **laptop** with 1660ti + i5-9300h, it takes only 230-235s per epoch to train with exact same model and data (and same code except the paths to files). \n\nTesla p100 is no doubt way better than a laptop 1660ti, but how can the fastest epoch on kaggle kernel be slower than the slowest epoch on my pc? Can someone point out possible problems?\nThanks!\n\nps. it is not a temporary issue since I tried several times last three days, but it never gets fast, I even tried explicitly call with tf.device('/GPU:0'): nothing changed. And the gpu meter at the top is  0% at all time (maybe it is not even using gpu at all?)\n and using tpu only gives me \"No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node\"\n\nupdate: I noticed that kernel cpu is always over 80 percent, maybe IO is making it slow? I am using flow_from_directory, which seems like how many people load data, if this is the problem, what is a better work around for this?",
      "votes": null
    },
    {
      "id": "1134510",
      "postDate": "01/01/2021 11:10:00",
      "content": "<p><a href=\"https://www.kaggle.com/mervynyang\" target=\"_blank\">@mervynyang</a> The Kaggle GPU kernels have a performance bottleneck due to the poor CPU performance. However, I feel like your situation is caused because you are not utilizing the GPU efficiently. For reference, I use PyTorch with 448x448 image size and a batch size of 16 with EfficientnetB4 and it takes around 800s per epoch on GPU. As shown in the image below, CPU usage is always above 80% and above 90% for GPU on my Kaggle kernel. </p>\n<p>The same setup with TPU and TFRecords using Tensorflow takes about 38s per epoch. Read about using TFRecords and <code>TFRecordDataset</code> for your data pipeline even for GPU <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\">here</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554871%2Fec719bdd48de66b00473e394a5156812%2Fkaggle_monitoring.png?generation=1609499291656543&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "mervynyang The Kaggle GPU kernels have a performance bottleneck due to the poor CPU performance. However, I feel like your situation is caused because you are not utilizing the GPU efficiently. For reference, I use PyTorch with 448x448 image size and a batch size of 16 with EfficientnetB4 and it takes around 800s per epoch on GPU. As shown in the image below, CPU usage is always above 80% and above 90% for GPU on my Kaggle kernel. \n\nThe same setup with TPU and TFRecords using Tensorflow takes about 38s per epoch. Read about using TFRecords and `TFRecordDataset` for your data pipeline even for GPU [here](https://www.tensorflow.org/tutorials/load_data/tfrecord).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554871%2Fec719bdd48de66b00473e394a5156812%2Fkaggle_monitoring.png?generation=1609499291656543&alt=media)",
      "votes": null
    },
    {
      "id": "1135502",
      "postDate": "01/02/2021 09:56:19",
      "content": "<p>Thank you so much! It seems like something's wrong with GPU, and no matter how I tune, it never works, I had to rewrite everything to tfrec + tpu, and now it runs fast!</p>",
      "rawMarkdown": "Thank you so much! It seems like something's wrong with GPU, and no matter how I tune, it never works, I had to rewrite everything to tfrec + tpu, and now it runs fast!",
      "votes": null
    },
    {
      "id": "1136295",
      "postDate": "01/02/2021 23:36:50",
      "content": "<p>Yeah, the IO can be huge bottleneck. Kaggle don't have SSD and the CPU are slightly slower, so when loading from the Hard Disk with lot of augmentation, it can be very slow if you are not prefetching properly.</p>",
      "rawMarkdown": "Yeah, the IO can be huge bottleneck. Kaggle don't have SSD and the CPU are slightly slower, so when loading from the Hard Disk with lot of augmentation, it can be very slow if you are not prefetching properly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1134510,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "01/01/2021 11:10:00",
      "content": "<p><a href=\"https://www.kaggle.com/mervynyang\" target=\"_blank\">@mervynyang</a> The Kaggle GPU kernels have a performance bottleneck due to the poor CPU performance. However, I feel like your situation is caused because you are not utilizing the GPU efficiently. For reference, I use PyTorch with 448x448 image size and a batch size of 16 with EfficientnetB4 and it takes around 800s per epoch on GPU. As shown in the image below, CPU usage is always above 80% and above 90% for GPU on my Kaggle kernel. </p>\n<p>The same setup with TPU and TFRecords using Tensorflow takes about 38s per epoch. Read about using TFRecords and <code>TFRecordDataset</code> for your data pipeline even for GPU <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\" target=\"_blank\">here</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554871%2Fec719bdd48de66b00473e394a5156812%2Fkaggle_monitoring.png?generation=1609499291656543&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1135502,
          "author_name": "mervynyang",
          "author_url": "",
          "post_date": "01/02/2021 09:56:19",
          "content": "<p>Thank you so much! It seems like something's wrong with GPU, and no matter how I tune, it never works, I had to rewrite everything to tfrec + tpu, and now it runs fast!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136295,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "01/02/2021 23:36:50",
      "content": "<p>Yeah, the IO can be huge bottleneck. Kaggle don't have SSD and the CPU are slightly slower, so when loading from the Hard Disk with lot of augmentation, it can be very slow if you are not prefetching properly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1134411": "I'm new to Deep Learning, and it is my first time using kaggle kernel (actually first time ever training online), and even if I used GPU acceleration, training speed is still slow.\n\nWith (150,150,3) * 64 batch_size input, tensorflow pretrained ResNet50V2(excluding top, freezing all layers) + 2 untrained fully connected layers (256 and 5) + some other layers (batchnorm, activation and dropout), it takes 300-400s to train each epoch on kaggle kernel. However, on my **laptop** with 1660ti + i5-9300h, it takes only 230-235s per epoch to train with exact same model and data (and same code except the paths to files). \n\nTesla p100 is no doubt way better than a laptop 1660ti, but how can the fastest epoch on kaggle kernel be slower than the slowest epoch on my pc? Can someone point out possible problems?\nThanks!\n\nps. it is not a temporary issue since I tried several times last three days, but it never gets fast, I even tried explicitly call with tf.device('/GPU:0'): nothing changed. And the gpu meter at the top is  0% at all time (maybe it is not even using gpu at all?)\n and using tpu only gives me \"No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node\"\n\nupdate: I noticed that kernel cpu is always over 80 percent, maybe IO is making it slow? I am using flow_from_directory, which seems like how many people load data, if this is the problem, what is a better work around for this?",
    "1134510": "mervynyang The Kaggle GPU kernels have a performance bottleneck due to the poor CPU performance. However, I feel like your situation is caused because you are not utilizing the GPU efficiently. For reference, I use PyTorch with 448x448 image size and a batch size of 16 with EfficientnetB4 and it takes around 800s per epoch on GPU. As shown in the image below, CPU usage is always above 80% and above 90% for GPU on my Kaggle kernel. \n\nThe same setup with TPU and TFRecords using Tensorflow takes about 38s per epoch. Read about using TFRecords and `TFRecordDataset` for your data pipeline even for GPU [here](https://www.tensorflow.org/tutorials/load_data/tfrecord).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1554871%2Fec719bdd48de66b00473e394a5156812%2Fkaggle_monitoring.png?generation=1609499291656543&alt=media)",
    "1135502": "Thank you so much! It seems like something's wrong with GPU, and no matter how I tune, it never works, I had to rewrite everything to tfrec + tpu, and now it runs fast!",
    "1136295": "Yeah, the IO can be huge bottleneck. Kaggle don't have SSD and the CPU are slightly slower, so when loading from the Hard Disk with lot of augmentation, it can be very slow if you are not prefetching properly."
  },
  "source": "meta"
}