{
  "id": 312977,
  "title": "(HELP)CNNs taking more time than transformers",
  "url": "/competitions/happy-whale-and-dolphin/discussion/312977",
  "author_name": "",
  "post_date": "2022-03-15T03:52:03.280176200Z",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>We have been experimenting with different architecture types, images sizes and many thing and observed that CNNs take always more time to train than swin and other transformers.</p>\n<p>For image_size = 224, swin_base224 was taking 6 mins per epoch, whereas effnet_b4 was taking 15 mins. We are training on 3090 with AMP. Except model architecture everything is same.(using timm for pretrained weights)</p>\n<p>Can someone help telling us why it is happening.<br>\nThanks</p>",
  "messages": [
    {
      "id": "1723015",
      "postDate": "03/15/2022 03:52:03",
      "content": "<p>We have been experimenting with different architecture types, images sizes and many thing and observed that CNNs take always more time to train than swin and other transformers.</p>\n<p>For image_size = 224, swin_base224 was taking 6 mins per epoch, whereas effnet_b4 was taking 15 mins. We are training on 3090 with AMP. Except model architecture everything is same.(using timm for pretrained weights)</p>\n<p>Can someone help telling us why it is happening.<br>\nThanks</p>",
      "rawMarkdown": "We have been experimenting with different architecture types, images sizes and many thing and observed that CNNs take always more time to train than swin and other transformers.\n\nFor image_size = 224, swin_base224 was taking 6 mins per epoch, whereas effnet_b4 was taking 15 mins. We are training on 3090 with AMP. Except model architecture everything is same.(using timm for pretrained weights)\n\nCan someone help telling us why it is happening.\nThanks",
      "votes": null
    },
    {
      "id": "1723042",
      "postDate": "03/15/2022 04:34:32",
      "content": "<p>For 3090, you probably want to use cuda 11.1 to best CNN speeds, I have 4x 3090, my speed is 3 minutes per epoch for b5-384</p>",
      "rawMarkdown": "For 3090, you probably want to use cuda 11.1 to best CNN speeds, I have 4x 3090, my speed is 3 minutes per epoch for b5-384",
      "votes": null
    },
    {
      "id": "1723052",
      "postDate": "03/15/2022 04:49:06",
      "content": "<p>Thanks for the suggestion but yes its cuda 11.1.<br>\nAlso do you know whats the training speed for one 3090 gpu of yours.</p>",
      "rawMarkdown": "Thanks for the suggestion but yes its cuda 11.1.\nAlso do you know whats the training speed for one 3090 gpu of yours.",
      "votes": null
    },
    {
      "id": "1723251",
      "postDate": "03/15/2022 09:01:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> </p>\n<p>Have you seen this <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111292\" target=\"_blank\">thread</a> I think they have fixed it now but you could try checking if <code>model.set_swish(memory_efficient=True)</code> this argument helps. </p>\n<p>If you are okay with sharing a minimal script here, I can test on my 3090 too-if it helps. </p>",
      "rawMarkdown": "Hi @mrinath \n\nHave you seen this [thread](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111292) I think they have fixed it now but you could try checking if `model.set_swish(memory_efficient=True)` this argument helps. \n\nIf you are okay with sharing a minimal script here, I can test on my 3090 too-if it helps.",
      "votes": null
    },
    {
      "id": "1723267",
      "postDate": "03/15/2022 09:15:02",
      "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> May I ask, are your 4x3090s in a single box? </p>\n<p>If you are okay to share-I'm curious how are you making it work, what power supply and power limiting are you using? </p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "harshitsheoran May I ask, are your 4x3090s in a single box? \n\nIf you are okay to share-I'm curious how are you making it work, what power supply and power limiting are you using? \n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1723284",
      "postDate": "03/15/2022 09:39:45",
      "content": "<p>Yes, its 4x3090 in a single box, and yes, it was one heck of a build to plan and most of parts were needed to imported, its all water cooled too, I can go on and on about the build but you can watch this, <a href=\"url\" target=\"_blank\">https://www.youtube.com/watch?v=dxgxlg8g5wo</a> they helped me build it</p>",
      "rawMarkdown": "Yes, its 4x3090 in a single box, and yes, it was one heck of a build to plan and most of parts were needed to imported, its all water cooled too, I can go on and on about the build but you can watch this, [https://www.youtube.com/watch?v=dxgxlg8g5wo](url) they helped me build it",
      "votes": null
    },
    {
      "id": "1723286",
      "postDate": "03/15/2022 09:41:22",
      "content": "<p>My setup gives about 3.5-3.9x boost against a single 3090, so your 3090 should be like 2-3 times faster, are you using the cudnn?, btw the bottleneck might not even be 3090, a 3090 is so powerful that to opitimize the other setup is really crucial, it my case, a training costs me about 32 threads and 200 gigs of ram (I most likely dont do ram method unless I use small models, for big models I can make work with my 10GBPS raid 0 I/O) to make it like really really fast and to unlock the limits of my all 3090 setup without I/O bottleneck.</p>",
      "rawMarkdown": "My setup gives about 3.5-3.9x boost against a single 3090, so your 3090 should be like 2-3 times faster, are you using the cudnn?, btw the bottleneck might not even be 3090, a 3090 is so powerful that to opitimize the other setup is really crucial, it my case, a training costs me about 32 threads and 200 gigs of ram (I most likely dont do ram method unless I use small models, for big models I can make work with my 10GBPS raid 0 I/O) to make it like really really fast and to unlock the limits of my all 3090 setup without I/O bottleneck.",
      "votes": null
    },
    {
      "id": "1723469",
      "postDate": "03/15/2022 12:43:27",
      "content": "<p>Thanks for the offer,but the trraining code is quite messy now, and If i share even a snippet it would take tremendous time to figure that out.<br>\nIts almost same as this notebook <a href=\"https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\" target=\"_blank\">https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter</a></p>\n<p>I have not used that set swish thing, I will try and report</p>",
      "rawMarkdown": "Thanks for the offer,but the trraining code is quite messy now, and If i share even a snippet it would take tremendous time to figure that out.\nIts almost same as this notebook https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\n\nI have not used that set swish thing, I will try and report",
      "votes": null
    },
    {
      "id": "1723470",
      "postDate": "03/15/2022 12:45:22",
      "content": "<p>I see, I too think 3090 is not the main problem here, just curious how a transformer model with almost 3x number of parameters is training as 2x fast compared to effnet,.</p>",
      "rawMarkdown": "I see, I too think 3090 is not the main problem here, just curious how a transformer model with almost 3x number of parameters is training as 2x fast compared to effnet,.",
      "votes": null
    },
    {
      "id": "1725595",
      "postDate": "03/17/2022 08:41:52",
      "content": "<p>one query <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> , how do you set no. of workers in your dataloader?</p>",
      "rawMarkdown": "one query @harshitsheoran , how do you set no. of workers in your dataloader?",
      "votes": null
    },
    {
      "id": "1725848",
      "postDate": "03/17/2022 13:29:51",
      "content": "<p>In pytorch:</p>\n<pre><code>train_loader = DataLoader(train_dataset, batch_size=CFG.TRN_BS, shuffle=True, num_workers=8, pin_memory=False)\nvalid_loader = DataLoader(valid_dataset, batch_size=CFG.VAL_BS, shuffle=False, num_workers=8, pin_memory=False)\n</code></pre>\n<p>pin memory False if you want to save some gpu ram, if you want a very little bit of extra speed, especially if your I/O is bottleneck pin_memory=True</p>\n<p>In tensorflow, there are many public kernels with something like<br>\n<code>train_dataset.map(train_augment, num_parallel_calls=8)</code></p>",
      "rawMarkdown": "In pytorch:\n\n```\ntrain_loader = DataLoader(train_dataset, batch_size=CFG.TRN_BS, shuffle=True, num_workers=8, pin_memory=False)\nvalid_loader = DataLoader(valid_dataset, batch_size=CFG.VAL_BS, shuffle=False, num_workers=8, pin_memory=False)\n```\n\npin memory False if you want to save some gpu ram, if you want a very little bit of extra speed, especially if your I/O is bottleneck pin_memory=True\n\n\nIn tensorflow, there are many public kernels with something like\n`train_dataset.map(train_augment, num_parallel_calls=8)`",
      "votes": null
    },
    {
      "id": "1726193",
      "postDate": "03/17/2022 18:16:55",
      "content": "<p>Are you using the tf_* version effnets from timm? These are generally slower. </p>",
      "rawMarkdown": "Are you using the tf_* version effnets from timm? These are generally slower.",
      "votes": null
    },
    {
      "id": "1726238",
      "postDate": "03/17/2022 19:32:45",
      "content": "<p>Thanks, but I tried for both tf and non tf effnets, but almost same time for both</p>",
      "rawMarkdown": "Thanks, but I tried for both tf and non tf effnets, but almost same time for both",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1723042,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "03/15/2022 04:34:32",
      "content": "<p>For 3090, you probably want to use cuda 11.1 to best CNN speeds, I have 4x 3090, my speed is 3 minutes per epoch for b5-384</p>",
      "votes": null,
      "replies": [
        {
          "id": 1723052,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/15/2022 04:49:06",
          "content": "<p>Thanks for the suggestion but yes its cuda 11.1.<br>\nAlso do you know whats the training speed for one 3090 gpu of yours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723267,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/15/2022 09:15:02",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> May I ask, are your 4x3090s in a single box? </p>\n<p>If you are okay to share-I'm curious how are you making it work, what power supply and power limiting are you using? </p>\n<p>Thanks in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723284,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/15/2022 09:39:45",
          "content": "<p>Yes, its 4x3090 in a single box, and yes, it was one heck of a build to plan and most of parts were needed to imported, its all water cooled too, I can go on and on about the build but you can watch this, <a href=\"url\" target=\"_blank\">https://www.youtube.com/watch?v=dxgxlg8g5wo</a> they helped me build it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723286,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/15/2022 09:41:22",
          "content": "<p>My setup gives about 3.5-3.9x boost against a single 3090, so your 3090 should be like 2-3 times faster, are you using the cudnn?, btw the bottleneck might not even be 3090, a 3090 is so powerful that to opitimize the other setup is really crucial, it my case, a training costs me about 32 threads and 200 gigs of ram (I most likely dont do ram method unless I use small models, for big models I can make work with my 10GBPS raid 0 I/O) to make it like really really fast and to unlock the limits of my all 3090 setup without I/O bottleneck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723470,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/15/2022 12:45:22",
          "content": "<p>I see, I too think 3090 is not the main problem here, just curious how a transformer model with almost 3x number of parameters is training as 2x fast compared to effnet,.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725595,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/17/2022 08:41:52",
          "content": "<p>one query <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> , how do you set no. of workers in your dataloader?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725848,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/17/2022 13:29:51",
          "content": "<p>In pytorch:</p>\n<pre><code>train_loader = DataLoader(train_dataset, batch_size=CFG.TRN_BS, shuffle=True, num_workers=8, pin_memory=False)\nvalid_loader = DataLoader(valid_dataset, batch_size=CFG.VAL_BS, shuffle=False, num_workers=8, pin_memory=False)\n</code></pre>\n<p>pin memory False if you want to save some gpu ram, if you want a very little bit of extra speed, especially if your I/O is bottleneck pin_memory=True</p>\n<p>In tensorflow, there are many public kernels with something like<br>\n<code>train_dataset.map(train_augment, num_parallel_calls=8)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1723251,
      "author_name": "init27",
      "author_url": "",
      "post_date": "03/15/2022 09:01:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> </p>\n<p>Have you seen this <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111292\" target=\"_blank\">thread</a> I think they have fixed it now but you could try checking if <code>model.set_swish(memory_efficient=True)</code> this argument helps. </p>\n<p>If you are okay with sharing a minimal script here, I can test on my 3090 too-if it helps. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1723469,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/15/2022 12:43:27",
          "content": "<p>Thanks for the offer,but the trraining code is quite messy now, and If i share even a snippet it would take tremendous time to figure that out.<br>\nIts almost same as this notebook <a href=\"https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\" target=\"_blank\">https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter</a></p>\n<p>I have not used that set swish thing, I will try and report</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1726193,
      "author_name": "benihime91",
      "author_url": "",
      "post_date": "03/17/2022 18:16:55",
      "content": "<p>Are you using the tf_* version effnets from timm? These are generally slower. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1726238,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/17/2022 19:32:45",
          "content": "<p>Thanks, but I tried for both tf and non tf effnets, but almost same time for both</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1723015": "We have been experimenting with different architecture types, images sizes and many thing and observed that CNNs take always more time to train than swin and other transformers.\n\nFor image_size = 224, swin_base224 was taking 6 mins per epoch, whereas effnet_b4 was taking 15 mins. We are training on 3090 with AMP. Except model architecture everything is same.(using timm for pretrained weights)\n\nCan someone help telling us why it is happening.\nThanks",
    "1723042": "For 3090, you probably want to use cuda 11.1 to best CNN speeds, I have 4x 3090, my speed is 3 minutes per epoch for b5-384",
    "1723052": "Thanks for the suggestion but yes its cuda 11.1.\nAlso do you know whats the training speed for one 3090 gpu of yours.",
    "1723251": "Hi @mrinath \n\nHave you seen this [thread](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/111292) I think they have fixed it now but you could try checking if `model.set_swish(memory_efficient=True)` this argument helps. \n\nIf you are okay with sharing a minimal script here, I can test on my 3090 too-if it helps.",
    "1723267": "harshitsheoran May I ask, are your 4x3090s in a single box? \n\nIf you are okay to share-I'm curious how are you making it work, what power supply and power limiting are you using? \n\nThanks in advance!",
    "1723284": "Yes, its 4x3090 in a single box, and yes, it was one heck of a build to plan and most of parts were needed to imported, its all water cooled too, I can go on and on about the build but you can watch this, [https://www.youtube.com/watch?v=dxgxlg8g5wo](url) they helped me build it",
    "1723286": "My setup gives about 3.5-3.9x boost against a single 3090, so your 3090 should be like 2-3 times faster, are you using the cudnn?, btw the bottleneck might not even be 3090, a 3090 is so powerful that to opitimize the other setup is really crucial, it my case, a training costs me about 32 threads and 200 gigs of ram (I most likely dont do ram method unless I use small models, for big models I can make work with my 10GBPS raid 0 I/O) to make it like really really fast and to unlock the limits of my all 3090 setup without I/O bottleneck.",
    "1723469": "Thanks for the offer,but the trraining code is quite messy now, and If i share even a snippet it would take tremendous time to figure that out.\nIts almost same as this notebook https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\n\nI have not used that set swish thing, I will try and report",
    "1723470": "I see, I too think 3090 is not the main problem here, just curious how a transformer model with almost 3x number of parameters is training as 2x fast compared to effnet,.",
    "1725595": "one query @harshitsheoran , how do you set no. of workers in your dataloader?",
    "1725848": "In pytorch:\n\n```\ntrain_loader = DataLoader(train_dataset, batch_size=CFG.TRN_BS, shuffle=True, num_workers=8, pin_memory=False)\nvalid_loader = DataLoader(valid_dataset, batch_size=CFG.VAL_BS, shuffle=False, num_workers=8, pin_memory=False)\n```\n\npin memory False if you want to save some gpu ram, if you want a very little bit of extra speed, especially if your I/O is bottleneck pin_memory=True\n\n\nIn tensorflow, there are many public kernels with something like\n`train_dataset.map(train_augment, num_parallel_calls=8)`",
    "1726193": "Are you using the tf_* version effnets from timm? These are generally slower.",
    "1726238": "Thanks, but I tried for both tf and non tf effnets, but almost same time for both"
  },
  "source": "meta"
}