{
  "id": 199856,
  "title": "Predicting Inference Time in EfficientNets",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199856",
  "author_name": "ITK8191",
  "post_date": "2020-11-27T16:52:59.644000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I investigated how much time EfficientNets require inference with <a href=\"https://www.kaggle.com/itsuki9180/efficientnet-and-cutmixup-with-tpu-predict-phase\" target=\"_blank\">my notebook</a>.<br>\nAs a method, I had the NN infer 21397 training images and used that time to calculate the inference time for about 15,000 test images.<br>\nThose results are summarized in the following table.<br>\nThe measurements were taken only once, so there is a large margin of error. <br>\nThe unit is seconds.</p>\n<p>512x512px</p>\n<table>\n<thead>\n<tr>\n<th>NN_type</th>\n<th>inference time(21397 imgs)</th>\n<th>inference time * 15k / 21397</th>\n<th># params of NN</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>372.650</td>\n<td>261.239</td>\n<td>5.3M</td>\n</tr>\n<tr>\n<td>EfficientNet-B1</td>\n<td>438.743</td>\n<td>307.573</td>\n<td>7.8M</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>457.931</td>\n<td>321.024</td>\n<td>9.2M</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>497.378</td>\n<td>348.678</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>630.423</td>\n<td>441.947</td>\n<td>19M</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>824.770</td>\n<td>578.190</td>\n<td>30M</td>\n</tr>\n<tr>\n<td>EfficientNet-B6</td>\n<td>1050.680</td>\n<td>736.561</td>\n<td>43M</td>\n</tr>\n<tr>\n<td>EfficientNet-B7</td>\n<td>1397.224</td>\n<td>979.499</td>\n<td>66M</td>\n</tr>\n</tbody>\n</table>\n<p>384x384px</p>\n<table>\n<thead>\n<tr>\n<th>NN_type</th>\n<th>inference time(21397 imgs)</th>\n<th>inference time * 15k / 21397</th>\n<th># params of NN</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>350.203</td>\n<td>245.503</td>\n<td>5.3M</td>\n</tr>\n<tr>\n<td>EfficientNet-B1</td>\n<td>380.065</td>\n<td>266.438</td>\n<td>7.8M</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>395.923</td>\n<td>277.555</td>\n<td>9.2M</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>425.116</td>\n<td>298.020</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>515.450</td>\n<td>361.347</td>\n<td>19M</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>677.162</td>\n<td>474.712</td>\n<td>30M</td>\n</tr>\n<tr>\n<td>EfficientNet-B6</td>\n<td>826.581</td>\n<td>579.460</td>\n<td>43M</td>\n</tr>\n<tr>\n<td>EfficientNet-B7</td>\n<td>1107.916</td>\n<td>776.685</td>\n<td>66M</td>\n</tr>\n</tbody>\n</table>\n<p>256x256 is during measurement…</p>\n<p>examples:<br>\nusing 512px EffNetB0 1Fold noTTa, this leads 261.239 * 1 * 1 = 261.239sec<br>\nusing 512px EffNetB3 5Folds 5TTA, this leads 348.678 * 5 * 5 = 8716.95sec</p>\n<p>Ignoring the time it takes for Augmentation, pre-processing and post-processing, so inference actually requires a little more time.</p>\n<p>In those that are not my notebooks the time may be different.<br>\nDifferent environments such as Pytorch may also differ in time.<br>\nAnyway, let's create a notebook that makes the most of 32400 seconds!</p>\n<p>references:<br>\n<a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks</a></p>",
  "messages": [
    {
      "id": 1093376,
      "postDate": "2020-11-27T16:52:59.643Z",
      "content": "<p>I investigated how much time EfficientNets require inference with <a href=\"https://www.kaggle.com/itsuki9180/efficientnet-and-cutmixup-with-tpu-predict-phase\" target=\"_blank\">my notebook</a>.<br>\nAs a method, I had the NN infer 21397 training images and used that time to calculate the inference time for about 15,000 test images.<br>\nThose results are summarized in the following table.<br>\nThe measurements were taken only once, so there is a large margin of error. <br>\nThe unit is seconds.</p>\n<p>512x512px</p>\n<table>\n<thead>\n<tr>\n<th>NN_type</th>\n<th>inference time(21397 imgs)</th>\n<th>inference time * 15k / 21397</th>\n<th># params of NN</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>372.650</td>\n<td>261.239</td>\n<td>5.3M</td>\n</tr>\n<tr>\n<td>EfficientNet-B1</td>\n<td>438.743</td>\n<td>307.573</td>\n<td>7.8M</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>457.931</td>\n<td>321.024</td>\n<td>9.2M</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>497.378</td>\n<td>348.678</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>630.423</td>\n<td>441.947</td>\n<td>19M</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>824.770</td>\n<td>578.190</td>\n<td>30M</td>\n</tr>\n<tr>\n<td>EfficientNet-B6</td>\n<td>1050.680</td>\n<td>736.561</td>\n<td>43M</td>\n</tr>\n<tr>\n<td>EfficientNet-B7</td>\n<td>1397.224</td>\n<td>979.499</td>\n<td>66M</td>\n</tr>\n</tbody>\n</table>\n<p>384x384px</p>\n<table>\n<thead>\n<tr>\n<th>NN_type</th>\n<th>inference time(21397 imgs)</th>\n<th>inference time * 15k / 21397</th>\n<th># params of NN</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B0</td>\n<td>350.203</td>\n<td>245.503</td>\n<td>5.3M</td>\n</tr>\n<tr>\n<td>EfficientNet-B1</td>\n<td>380.065</td>\n<td>266.438</td>\n<td>7.8M</td>\n</tr>\n<tr>\n<td>EfficientNet-B2</td>\n<td>395.923</td>\n<td>277.555</td>\n<td>9.2M</td>\n</tr>\n<tr>\n<td>EfficientNet-B3</td>\n<td>425.116</td>\n<td>298.020</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>EfficientNet-B4</td>\n<td>515.450</td>\n<td>361.347</td>\n<td>19M</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>677.162</td>\n<td>474.712</td>\n<td>30M</td>\n</tr>\n<tr>\n<td>EfficientNet-B6</td>\n<td>826.581</td>\n<td>579.460</td>\n<td>43M</td>\n</tr>\n<tr>\n<td>EfficientNet-B7</td>\n<td>1107.916</td>\n<td>776.685</td>\n<td>66M</td>\n</tr>\n</tbody>\n</table>\n<p>256x256 is during measurement…</p>\n<p>examples:<br>\nusing 512px EffNetB0 1Fold noTTa, this leads 261.239 * 1 * 1 = 261.239sec<br>\nusing 512px EffNetB3 5Folds 5TTA, this leads 348.678 * 5 * 5 = 8716.95sec</p>\n<p>Ignoring the time it takes for Augmentation, pre-processing and post-processing, so inference actually requires a little more time.</p>\n<p>In those that are not my notebooks the time may be different.<br>\nDifferent environments such as Pytorch may also differ in time.<br>\nAnyway, let's create a notebook that makes the most of 32400 seconds!</p>\n<p>references:<br>\n<a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks</a></p>",
      "rawMarkdown": "I investigated how much time EfficientNets require inference with [my notebook](https://www.kaggle.com/itsuki9180/efficientnet-and-cutmixup-with-tpu-predict-phase).\nAs a method, I had the NN infer 21397 training images and used that time to calculate the inference time for about 15,000 test images.\nThose results are summarized in the following table.\nThe measurements were taken only once, so there is a large margin of error. \nThe unit is seconds.\n\n512x512px\n| NN_type | inference time(21397 imgs) | inference time * 15k / 21397| # params of NN|\n| --- | --- |\n| EfficientNet-B0 | 372.650 |261.239|5.3M|\n| EfficientNet-B1 | 438.743|307.573|7.8M|\n| EfficientNet-B2 | 457.931|321.024|9.2M|\n| EfficientNet-B3 | 497.378 |348.678|12M|\n| EfficientNet-B4 | 630.423|441.947|19M|\n| EfficientNet-B5 | 824.770 |578.190|30M|\n| EfficientNet-B6 | 1050.680|736.561|43M|\n| EfficientNet-B7 | 1397.224|979.499|66M|\n\n384x384px\n| NN_type | inference time(21397 imgs) | inference time * 15k / 21397| # params of NN|\n| --- | --- |\n| EfficientNet-B0 | 350.203|245.503|5.3M|\n| EfficientNet-B1 | 380.065|266.438|7.8M|\n| EfficientNet-B2 | 395.923|277.555|9.2M|\n| EfficientNet-B3 | 425.116|298.020|12M|\n| EfficientNet-B4 | 515.450|361.347|19M|\n| EfficientNet-B5 | 677.162|474.712|30M|\n| EfficientNet-B6 | 826.581|579.460|43M|\n| EfficientNet-B7 | 1107.916|776.685|66M|\n\n 256x256 is during measurement...\n\nexamples:\nusing 512px EffNetB0 1Fold noTTa, this leads 261.239 * 1 * 1 = 261.239sec\nusing 512px EffNetB3 5Folds 5TTA, this leads 348.678 * 5 * 5 = 8716.95sec\n\nIgnoring the time it takes for Augmentation, pre-processing and post-processing, so inference actually requires a little more time.\n\nIn those that are not my notebooks the time may be different.\nDifferent environments such as Pytorch may also differ in time.\nAnyway, let's create a notebook that makes the most of 32400 seconds!\n\nreferences:\n[EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks](https://arxiv.org/abs/1905.11946)\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1093376": "I investigated how much time EfficientNets require inference with [my notebook](https://www.kaggle.com/itsuki9180/efficientnet-and-cutmixup-with-tpu-predict-phase).\nAs a method, I had the NN infer 21397 training images and used that time to calculate the inference time for about 15,000 test images.\nThose results are summarized in the following table.\nThe measurements were taken only once, so there is a large margin of error. \nThe unit is seconds.\n\n512x512px\n| NN_type | inference time(21397 imgs) | inference time * 15k / 21397| # params of NN|\n| --- | --- |\n| EfficientNet-B0 | 372.650 |261.239|5.3M|\n| EfficientNet-B1 | 438.743|307.573|7.8M|\n| EfficientNet-B2 | 457.931|321.024|9.2M|\n| EfficientNet-B3 | 497.378 |348.678|12M|\n| EfficientNet-B4 | 630.423|441.947|19M|\n| EfficientNet-B5 | 824.770 |578.190|30M|\n| EfficientNet-B6 | 1050.680|736.561|43M|\n| EfficientNet-B7 | 1397.224|979.499|66M|\n\n384x384px\n| NN_type | inference time(21397 imgs) | inference time * 15k / 21397| # params of NN|\n| --- | --- |\n| EfficientNet-B0 | 350.203|245.503|5.3M|\n| EfficientNet-B1 | 380.065|266.438|7.8M|\n| EfficientNet-B2 | 395.923|277.555|9.2M|\n| EfficientNet-B3 | 425.116|298.020|12M|\n| EfficientNet-B4 | 515.450|361.347|19M|\n| EfficientNet-B5 | 677.162|474.712|30M|\n| EfficientNet-B6 | 826.581|579.460|43M|\n| EfficientNet-B7 | 1107.916|776.685|66M|\n\n 256x256 is during measurement...\n\nexamples:\nusing 512px EffNetB0 1Fold noTTa, this leads 261.239 * 1 * 1 = 261.239sec\nusing 512px EffNetB3 5Folds 5TTA, this leads 348.678 * 5 * 5 = 8716.95sec\n\nIgnoring the time it takes for Augmentation, pre-processing and post-processing, so inference actually requires a little more time.\n\nIn those that are not my notebooks the time may be different.\nDifferent environments such as Pytorch may also differ in time.\nAnyway, let's create a notebook that makes the most of 32400 seconds!\n\nreferences:\n[EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks](https://arxiv.org/abs/1905.11946)\n"
  }
}