{
  "id": 215544,
  "title": "My vision transformer submission raises a timeout",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215544",
  "author_name": "",
  "post_date": "2021-01-30T10:31:04.647686600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>After successful training of efficientNet b-5<br>\nI have been trying to submit my vision transformer(B-16, 512x512 size),<br>\nBut Timeout occurs.<br>\nWhat should i do to fix the problem?</p>",
  "messages": [
    {
      "id": "1177515",
      "postDate": "01/30/2021 10:31:04",
      "content": "<p>After successful training of efficientNet b-5<br>\nI have been trying to submit my vision transformer(B-16, 512x512 size),<br>\nBut Timeout occurs.<br>\nWhat should i do to fix the problem?</p>",
      "rawMarkdown": "After successful training of efficientNet b-5\nI have been trying to submit my vision transformer(B-16, 512x512 size),\nBut Timeout occurs.\nWhat should i do to fix the problem?",
      "votes": null
    },
    {
      "id": "1177808",
      "postDate": "01/30/2021 14:18:24",
      "content": "<p>Many people are reporting long inference time with ViT. I don't know if it's a bug or it's just too slow.<br>\nI personally do inference with just 1 fold:</p>\n<ul>\n<li>1 Fold inference: 2.5 hours for ViT-l-16 and 1.5 hour for ViT-b-16. <em>(with batch_size=8)</em></li>\n<li>if you use TTA then expect the inference time to be multiplied by the <code>N_TTA</code>. Setting <code>TTA</code> to <code>False</code> and doing inference with just 1 fold will take less than 2 hours with vit-b16.</li>\n<li>Increase the batch_size if the memory allows it, in my case I use <code>batch_size=32</code>, it's faster and I don't get out of memory issues.</li>\n</ul>\n<p>EDIT: I just tried ViT with pytorch and its inference time is normal, <em>5 folds vit-b16 in 1.5 hours!!!</em> Apparently there is a problem with the vit-keras, maybe inference uses CPU instead of GPU!</p>",
      "rawMarkdown": "Many people are reporting long inference time with ViT. I don't know if it's a bug or it's just too slow.\nI personally do inference with just 1 fold:\n* 1 Fold inference: 2.5 hours for ViT-l-16 and 1.5 hour for ViT-b-16. *(with batch_size=8)*\n* if you use TTA then expect the inference time to be multiplied by the `N_TTA`. Setting `TTA` to `False` and doing inference with just 1 fold will take less than 2 hours with vit-b16.\n* Increase the batch_size if the memory allows it, in my case I use `batch_size=32`, it's faster and I don't get out of memory issues.\n\nEDIT: I just tried ViT with pytorch and its inference time is normal, *5 folds vit-b16 in 1.5 hours!!!* Apparently there is a problem with the vit-keras, maybe inference uses CPU instead of GPU!",
      "votes": null
    },
    {
      "id": "1178629",
      "postDate": "01/31/2021 03:19:42",
      "content": "<p>I barely submitted my vit-b16 written in tensorflow.<br>\nMy vit-b16 took around 8 hours<br>\nI decided to give up vit_keras and transfer to pytorch!<br>\nThank you so much</p>",
      "rawMarkdown": "I barely submitted my vit-b16 written in tensorflow.\nMy vit-b16 took around 8 hours\nI decided to give up vit_keras and transfer to pytorch!\nThank you so much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177808,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "01/30/2021 14:18:24",
      "content": "<p>Many people are reporting long inference time with ViT. I don't know if it's a bug or it's just too slow.<br>\nI personally do inference with just 1 fold:</p>\n<ul>\n<li>1 Fold inference: 2.5 hours for ViT-l-16 and 1.5 hour for ViT-b-16. <em>(with batch_size=8)</em></li>\n<li>if you use TTA then expect the inference time to be multiplied by the <code>N_TTA</code>. Setting <code>TTA</code> to <code>False</code> and doing inference with just 1 fold will take less than 2 hours with vit-b16.</li>\n<li>Increase the batch_size if the memory allows it, in my case I use <code>batch_size=32</code>, it's faster and I don't get out of memory issues.</li>\n</ul>\n<p>EDIT: I just tried ViT with pytorch and its inference time is normal, <em>5 folds vit-b16 in 1.5 hours!!!</em> Apparently there is a problem with the vit-keras, maybe inference uses CPU instead of GPU!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1178629,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "01/31/2021 03:19:42",
          "content": "<p>I barely submitted my vit-b16 written in tensorflow.<br>\nMy vit-b16 took around 8 hours<br>\nI decided to give up vit_keras and transfer to pytorch!<br>\nThank you so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1177515": "After successful training of efficientNet b-5\nI have been trying to submit my vision transformer(B-16, 512x512 size),\nBut Timeout occurs.\nWhat should i do to fix the problem?",
    "1177808": "Many people are reporting long inference time with ViT. I don't know if it's a bug or it's just too slow.\nI personally do inference with just 1 fold:\n* 1 Fold inference: 2.5 hours for ViT-l-16 and 1.5 hour for ViT-b-16. *(with batch_size=8)*\n* if you use TTA then expect the inference time to be multiplied by the `N_TTA`. Setting `TTA` to `False` and doing inference with just 1 fold will take less than 2 hours with vit-b16.\n* Increase the batch_size if the memory allows it, in my case I use `batch_size=32`, it's faster and I don't get out of memory issues.\n\nEDIT: I just tried ViT with pytorch and its inference time is normal, *5 folds vit-b16 in 1.5 hours!!!* Apparently there is a problem with the vit-keras, maybe inference uses CPU instead of GPU!",
    "1178629": "I barely submitted my vit-b16 written in tensorflow.\nMy vit-b16 took around 8 hours\nI decided to give up vit_keras and transfer to pytorch!\nThank you so much"
  },
  "source": "meta"
}