{
  "id": 268756,
  "title": "Transfer learning on EfficientNet taking too long",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/268756",
  "author_name": "",
  "post_date": "2021-08-28T18:33:06.767662700Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Like many other people I am using transfer learning with EfficientNetB7 from tensorflow to classify a signal. I use the image spectogram images kindly provided here by another Kaggle user <a href=\"https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb\" target=\"_blank\">https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb</a></p>\n<p>I am using 140,000 images i batches of 32<br>\nJust to report what I am doing to already speed things up.<br>\n1) I import the images using image_dataset_from_directory in a (256 x 256) image_size<br>\n2) I use tf.data.experimental.AUTOTUNE to prefetch the images so that the gpu is always active<br>\n3) I use EfficientNetB7 from tensorflow. I remove the top layers and freeze the remaining weights. I then add layers of my own for classification. I also use preprocess_inputs. I only have 2561 trainable parameters</p>\n<p>I am using a M1 Mac. Its still taking &gt;2 hrs to go through the entire dataset. Any tips or any guesses on what I might be doing wrong?</p>\n<p>Thanks,<br>\nAkshay</p>",
  "messages": [
    {
      "id": "1494519",
      "postDate": "08/28/2021 18:33:06",
      "content": "<p>Hi,</p>\n<p>Like many other people I am using transfer learning with EfficientNetB7 from tensorflow to classify a signal. I use the image spectogram images kindly provided here by another Kaggle user <a href=\"https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb\" target=\"_blank\">https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb</a></p>\n<p>I am using 140,000 images i batches of 32<br>\nJust to report what I am doing to already speed things up.<br>\n1) I import the images using image_dataset_from_directory in a (256 x 256) image_size<br>\n2) I use tf.data.experimental.AUTOTUNE to prefetch the images so that the gpu is always active<br>\n3) I use EfficientNetB7 from tensorflow. I remove the top layers and freeze the remaining weights. I then add layers of my own for classification. I also use preprocess_inputs. I only have 2561 trainable parameters</p>\n<p>I am using a M1 Mac. Its still taking &gt;2 hrs to go through the entire dataset. Any tips or any guesses on what I might be doing wrong?</p>\n<p>Thanks,<br>\nAkshay</p>",
      "rawMarkdown": "Hi,\n\nLike many other people I am using transfer learning with EfficientNetB7 from tensorflow to classify a signal. I use the image spectogram images kindly provided here by another Kaggle user https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb\n\nI am using 140,000 images i batches of 32\nJust to report what I am doing to already speed things up.\n1) I import the images using image_dataset_from_directory in a (256 x 256) image_size\n2) I use tf.data.experimental.AUTOTUNE to prefetch the images so that the gpu is always active\n3) I use EfficientNetB7 from tensorflow. I remove the top layers and freeze the remaining weights. I then add layers of my own for classification. I also use preprocess_inputs. I only have 2561 trainable parameters\n\nI am using a M1 Mac. Its still taking >2 hrs to go through the entire dataset. Any tips or any guesses on what I might be doing wrong?\n\nThanks,\nAkshay",
      "votes": null
    },
    {
      "id": "1494534",
      "postDate": "08/28/2021 18:44:12",
      "content": "<blockquote>\n  <p>Any tips or any guesses on what I might be doing wrong?</p>\n</blockquote>\n<p>Using B7 on local non-server PC. Even with 10-20x speedup in inference mode with frozen layers, it is still a big net that needs to pass your images through it, and it takes time.</p>",
      "rawMarkdown": "> Any tips or any guesses on what I might be doing wrong?\n\nUsing B7 on local non-server PC. Even with 10-20x speedup in inference mode with frozen layers, it is still a big net that needs to pass your images through it, and it takes time.",
      "votes": null
    },
    {
      "id": "1494537",
      "postDate": "08/28/2021 18:45:00",
      "content": "<p><a href=\"https://www.kaggle.com/aghalsa\" target=\"_blank\">@aghalsa</a> M1 Mac with 16GB, I tried too. It is not good yet to train large models. Kaggle/Colab is much better. See this article from w&amp;b - <a href=\"https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg\" target=\"_blank\">https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg</a>. Maybe M1X for training :D</p>",
      "rawMarkdown": "aghalsa M1 Mac with 16GB, I tried too. It is not good yet to train large models. Kaggle/Colab is much better. See this article from w&b - https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg. Maybe M1X for training :D",
      "votes": null
    },
    {
      "id": "1494547",
      "postDate": "08/28/2021 18:54:45",
      "content": "<p>Thanks. Guess Il will start using collab.  Is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?</p>",
      "rawMarkdown": "Thanks. Guess Il will start using collab.  Is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?",
      "votes": null
    },
    {
      "id": "1494554",
      "postDate": "08/28/2021 19:00:02",
      "content": "<p>Thanks. Is there a way to guesstimate how long it should take. Also is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?</p>",
      "rawMarkdown": "Thanks. Is there a way to guesstimate how long it should take. Also is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?",
      "votes": null
    },
    {
      "id": "1494562",
      "postDate": "08/28/2021 19:14:40",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training</a> - 1 Epoch with B7 - 256, this notebook able to do in ~ 5mins @ Kaggle TPU.</p>",
      "rawMarkdown": "https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training - 1 Epoch with B7 - 256, this notebook able to do in ~ 5mins @ Kaggle TPU.",
      "votes": null
    },
    {
      "id": "1494575",
      "postDate": "08/28/2021 19:34:56",
      "content": "<p>If you would not change earlier layers/input images, then you can pass all your images through those layers once, save results and use those results as \"images\" for upper layer training.</p>\n<p>Doing it this way might be worse than just using a smaller net, though.</p>",
      "rawMarkdown": "If you would not change earlier layers/input images, then you can pass all your images through those layers once, save results and use those results as \"images\" for upper layer training.\n\nDoing it this way might be worse than just using a smaller net, though.",
      "votes": null
    },
    {
      "id": "1498431",
      "postDate": "09/01/2021 01:37:27",
      "content": "<p><a href=\"https://www.kaggle.com/aghalsa\" target=\"_blank\">@aghalsa</a> to give you a comparison I'm using Keras/TF 2.0 with a TPU all run in Kaggle notebooks:</p>\n<ul>\n<li>I use TFRecords for read performance</li>\n<li>An optimised CWT (<a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a>) instead of nnAudio/CQT. The performance of these is similar but much better than librosa</li>\n<li>EFN</li>\n<li>Split my data 95%/5%.</li>\n</ul>\n<p>For experimentation I set n_scales to 128 and stride to 32 to create a 128x128 image with three channels that is passed to B0. A full epoch takes under 3min. </p>\n<p>If I switch to 256x256 CWT and B7 it is about 15 min per epoch.</p>\n<p>In my experience the TPU is 8x faster than a GPU which is 100x faster than a CPU (using Kaggle resources). The bottom line is to get high performance you need to use a GPU or TPU with an optimised CWT and probably with TFRecords.</p>",
      "rawMarkdown": "aghalsa to give you a comparison I'm using Keras/TF 2.0 with a TPU all run in Kaggle notebooks:\n* I use TFRecords for read performance\n* An optimised CWT (https://github.com/Kevin-McIsaac/cmorlet-tensorflow) instead of nnAudio/CQT. The performance of these is similar but much better than librosa\n* EFN\n* Split my data 95%/5%.\n\nFor experimentation I set n_scales to 128 and stride to 32 to create a 128x128 image with three channels that is passed to B0. A full epoch takes under 3min. \n\nIf I switch to 256x256 CWT and B7 it is about 15 min per epoch.\n\nIn my experience the TPU is 8x faster than a GPU which is 100x faster than a CPU (using Kaggle resources). The bottom line is to get high performance you need to use a GPU or TPU with an optimised CWT and probably with TFRecords.",
      "votes": null
    },
    {
      "id": "1561009",
      "postDate": "10/27/2021 09:09:17",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1494534,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "08/28/2021 18:44:12",
      "content": "<blockquote>\n  <p>Any tips or any guesses on what I might be doing wrong?</p>\n</blockquote>\n<p>Using B7 on local non-server PC. Even with 10-20x speedup in inference mode with frozen layers, it is still a big net that needs to pass your images through it, and it takes time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1494554,
          "author_name": "aghalsa",
          "author_url": "",
          "post_date": "08/28/2021 19:00:02",
          "content": "<p>Thanks. Is there a way to guesstimate how long it should take. Also is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494575,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "08/28/2021 19:34:56",
          "content": "<p>If you would not change earlier layers/input images, then you can pass all your images through those layers once, save results and use those results as \"images\" for upper layer training.</p>\n<p>Doing it this way might be worse than just using a smaller net, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1494537,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "08/28/2021 18:45:00",
      "content": "<p><a href=\"https://www.kaggle.com/aghalsa\" target=\"_blank\">@aghalsa</a> M1 Mac with 16GB, I tried too. It is not good yet to train large models. Kaggle/Colab is much better. See this article from w&amp;b - <a href=\"https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg\" target=\"_blank\">https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg</a>. Maybe M1X for training :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 1494547,
          "author_name": "aghalsa",
          "author_url": "",
          "post_date": "08/28/2021 18:54:45",
          "content": "<p>Thanks. Guess Il will start using collab.  Is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1494562,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "08/28/2021 19:14:40",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training</a> - 1 Epoch with B7 - 256, this notebook able to do in ~ 5mins @ Kaggle TPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1498431,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "09/01/2021 01:37:27",
      "content": "<p><a href=\"https://www.kaggle.com/aghalsa\" target=\"_blank\">@aghalsa</a> to give you a comparison I'm using Keras/TF 2.0 with a TPU all run in Kaggle notebooks:</p>\n<ul>\n<li>I use TFRecords for read performance</li>\n<li>An optimised CWT (<a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">https://github.com/Kevin-McIsaac/cmorlet-tensorflow</a>) instead of nnAudio/CQT. The performance of these is similar but much better than librosa</li>\n<li>EFN</li>\n<li>Split my data 95%/5%.</li>\n</ul>\n<p>For experimentation I set n_scales to 128 and stride to 32 to create a 128x128 image with three channels that is passed to B0. A full epoch takes under 3min. </p>\n<p>If I switch to 256x256 CWT and B7 it is about 15 min per epoch.</p>\n<p>In my experience the TPU is 8x faster than a GPU which is 100x faster than a CPU (using Kaggle resources). The bottom line is to get high performance you need to use a GPU or TPU with an optimised CWT and probably with TFRecords.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1561009,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 09:09:17",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1494519": "Hi,\n\nLike many other people I am using transfer learning with EfficientNetB7 from tensorflow to classify a signal. I use the image spectogram images kindly provided here by another Kaggle user https://www.kaggle.com/coldfir3/g2net-cqt-dataset-pt1-jpgrgb\n\nI am using 140,000 images i batches of 32\nJust to report what I am doing to already speed things up.\n1) I import the images using image_dataset_from_directory in a (256 x 256) image_size\n2) I use tf.data.experimental.AUTOTUNE to prefetch the images so that the gpu is always active\n3) I use EfficientNetB7 from tensorflow. I remove the top layers and freeze the remaining weights. I then add layers of my own for classification. I also use preprocess_inputs. I only have 2561 trainable parameters\n\nI am using a M1 Mac. Its still taking >2 hrs to go through the entire dataset. Any tips or any guesses on what I might be doing wrong?\n\nThanks,\nAkshay",
    "1494534": "> Any tips or any guesses on what I might be doing wrong?\n\nUsing B7 on local non-server PC. Even with 10-20x speedup in inference mode with frozen layers, it is still a big net that needs to pass your images through it, and it takes time.",
    "1494537": "aghalsa M1 Mac with 16GB, I tried too. It is not good yet to train large models. Kaggle/Colab is much better. See this article from w&b - https://wandb.ai/vanpelt/m1-benchmark/reports/Can-Apple-s-M1-help-you-train-models-faster-cheaper-than-NVIDIA-s-V100---VmlldzozNTkyMzg. Maybe M1X for training :D",
    "1494547": "Thanks. Guess Il will start using collab.  Is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?",
    "1494554": "Thanks. Is there a way to guesstimate how long it should take. Also is there a way to preprocess all the inputs once beforehand so I can play with meta parameters without having through wait this long again?",
    "1494562": "https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-training - 1 Epoch with B7 - 256, this notebook able to do in ~ 5mins @ Kaggle TPU.",
    "1494575": "If you would not change earlier layers/input images, then you can pass all your images through those layers once, save results and use those results as \"images\" for upper layer training.\n\nDoing it this way might be worse than just using a smaller net, though.",
    "1498431": "aghalsa to give you a comparison I'm using Keras/TF 2.0 with a TPU all run in Kaggle notebooks:\n* I use TFRecords for read performance\n* An optimised CWT (https://github.com/Kevin-McIsaac/cmorlet-tensorflow) instead of nnAudio/CQT. The performance of these is similar but much better than librosa\n* EFN\n* Split my data 95%/5%.\n\nFor experimentation I set n_scales to 128 and stride to 32 to create a 128x128 image with three channels that is passed to B0. A full epoch takes under 3min. \n\nIf I switch to 256x256 CWT and B7 it is about 15 min per epoch.\n\nIn my experience the TPU is 8x faster than a GPU which is 100x faster than a CPU (using Kaggle resources). The bottom line is to get high performance you need to use a GPU or TPU with an optimised CWT and probably with TFRecords.",
    "1561009": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}