{
  "id": 204950,
  "title": "EfficientNet Keras Baselines (LB 0.957 - single model)",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/204950",
  "author_name": "xhlulu",
  "post_date": "2020-12-17T16:43:38.855000",
  "votes": 148,
  "comment_count": 63,
  "views": 0,
  "content": "<p>I've published a few notebooks, so I'd like to aggregate them in this post and give more details. This is a work in progress.</p>\n<h2>Scores by model</h2>\n<p>All the models I trained were EfficientNets using TensorFlow Keras. Depending on the size they were either trained on GPU or TPUs. <strong>The scores were all calculated after the rescoring started so they should be up-to-date</strong>.</p>\n<p>You can find the respective notebooks here:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\" target=\"_blank\">Full workflow on GPU (train and submit)</a></li>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">Train on TPU</a></li>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-submission\" target=\"_blank\">Submit on GPU</a></li>\n</ol>\n<p>Note that [1] is for the smaller models (B2 and B3), and [2-3] are for the larger models (B6-B7). Although [1] is simpler, it take a lot of time to submit so I recommend taking a look at [3] if you want to have faster submissions.</p>\n<p>Here's a summary table of the scores:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Img Size</th>\n<th>Valid AUC</th>\n<th>LB</th>\n<th>Accelerator</th>\n<th>Weights</th>\n<th>Version</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>224</td>\n<td>0.9454</td>\n<td>0.883</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>15</td>\n</tr>\n<tr>\n<td>B1</td>\n<td>240</td>\n<td>0.9618</td>\n<td>0.888</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>14</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9278</td>\n<td>0.918</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>10</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9587</td>\n<td>0.912</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>11</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9713</td>\n<td>0.908</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>12</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>300</td>\n<td>0.9202</td>\n<td>0.909</td>\n<td>GPU</td>\n<td>ImageNet</td>\n<td>9</td>\n</tr>\n<tr>\n<td>B5</td>\n<td>456</td>\n<td>0.9382</td>\n<td>0.944</td>\n<td>TPU</td>\n<td>ImageNet</td>\n<td>9</td>\n</tr>\n<tr>\n<td>B6</td>\n<td>528</td>\n<td>0.9415</td>\n<td>0.949</td>\n<td>TPU</td>\n<td>Noisy Student</td>\n<td>6</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>600</td>\n<td>0.9455</td>\n<td>0.953</td>\n<td>TPU</td>\n<td>Noisy Student</td>\n<td>7</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>600</td>\n<td>0.9431</td>\n<td>0.957</td>\n<td>TPU</td>\n<td>ImageNet</td>\n<td>8</td>\n</tr>\n</tbody>\n</table>\n<h2>Hyperparameters and training details</h2>\n<p>Here are some details about how I trained the models:</p>\n<ul>\n<li><strong>Training augmentation</strong>: Simply random left-right and top-bottom flipping</li>\n<li><strong>Optimizer</strong>: Adam with an initial learning rate of 0.001 and no other tuning</li>\n<li><strong>Metrics</strong>: Multi-label AUROC (as opposed to flattened), corresponding to the competition metric</li>\n<li><strong>Scheduling</strong>: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve</li>\n<li><strong>Number of epochs</strong>: 10-15 for early version of GPU, 20 for later version and for TPU</li>\n<li><strong>Batch Size</strong>: 16 for GPU and 16*8=128 for the TPU</li>\n<li><strong>Model saving</strong> Save model with the best validation AUC after every epoch</li>\n</ul>\n<p>I did not try any of the following:</p>\n<ul>\n<li><strong>Cross-validation</strong></li>\n<li><strong>TTA</strong></li>\n<li><strong>Ensembling/Stacking</strong></li>\n</ul>\n<h2>Other notes</h2>\n<ul>\n<li>In more recent versions, I've changed the implementation from <code>qubvel/callidor</code> to the one in <code>tensorflow.keras.applications</code>. Since it is not officially released for v2.2.0, I made <a href=\"https://www.kaggle.com/xhlulu/tf-keras-efficientnet\" target=\"_blank\">a utility script</a> for the TPU notebook. For the other notebooks, I'm loading directly from <code>tensorflow</code> since they are on v2.3.1</li>\n</ul>",
  "messages": [
    {
      "id": 1117019,
      "postDate": "2020-12-17T16:43:38.857Z",
      "content": "<p>I've published a few notebooks, so I'd like to aggregate them in this post and give more details. This is a work in progress.</p>\n<h2>Scores by model</h2>\n<p>All the models I trained were EfficientNets using TensorFlow Keras. Depending on the size they were either trained on GPU or TPUs. <strong>The scores were all calculated after the rescoring started so they should be up-to-date</strong>.</p>\n<p>You can find the respective notebooks here:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\" target=\"_blank\">Full workflow on GPU (train and submit)</a></li>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">Train on TPU</a></li>\n<li><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-submission\" target=\"_blank\">Submit on GPU</a></li>\n</ol>\n<p>Note that [1] is for the smaller models (B2 and B3), and [2-3] are for the larger models (B6-B7). Although [1] is simpler, it take a lot of time to submit so I recommend taking a look at [3] if you want to have faster submissions.</p>\n<p>Here's a summary table of the scores:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Img Size</th>\n<th>Valid AUC</th>\n<th>LB</th>\n<th>Accelerator</th>\n<th>Weights</th>\n<th>Version</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>224</td>\n<td>0.9454</td>\n<td>0.883</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>15</td>\n</tr>\n<tr>\n<td>B1</td>\n<td>240</td>\n<td>0.9618</td>\n<td>0.888</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>14</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9278</td>\n<td>0.918</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>10</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9587</td>\n<td>0.912</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>11</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>260</td>\n<td>0.9713</td>\n<td>0.908</td>\n<td>GPU</td>\n<td>Noisy Student</td>\n<td>12</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>300</td>\n<td>0.9202</td>\n<td>0.909</td>\n<td>GPU</td>\n<td>ImageNet</td>\n<td>9</td>\n</tr>\n<tr>\n<td>B5</td>\n<td>456</td>\n<td>0.9382</td>\n<td>0.944</td>\n<td>TPU</td>\n<td>ImageNet</td>\n<td>9</td>\n</tr>\n<tr>\n<td>B6</td>\n<td>528</td>\n<td>0.9415</td>\n<td>0.949</td>\n<td>TPU</td>\n<td>Noisy Student</td>\n<td>6</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>600</td>\n<td>0.9455</td>\n<td>0.953</td>\n<td>TPU</td>\n<td>Noisy Student</td>\n<td>7</td>\n</tr>\n<tr>\n<td>B7</td>\n<td>600</td>\n<td>0.9431</td>\n<td>0.957</td>\n<td>TPU</td>\n<td>ImageNet</td>\n<td>8</td>\n</tr>\n</tbody>\n</table>\n<h2>Hyperparameters and training details</h2>\n<p>Here are some details about how I trained the models:</p>\n<ul>\n<li><strong>Training augmentation</strong>: Simply random left-right and top-bottom flipping</li>\n<li><strong>Optimizer</strong>: Adam with an initial learning rate of 0.001 and no other tuning</li>\n<li><strong>Metrics</strong>: Multi-label AUROC (as opposed to flattened), corresponding to the competition metric</li>\n<li><strong>Scheduling</strong>: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve</li>\n<li><strong>Number of epochs</strong>: 10-15 for early version of GPU, 20 for later version and for TPU</li>\n<li><strong>Batch Size</strong>: 16 for GPU and 16*8=128 for the TPU</li>\n<li><strong>Model saving</strong> Save model with the best validation AUC after every epoch</li>\n</ul>\n<p>I did not try any of the following:</p>\n<ul>\n<li><strong>Cross-validation</strong></li>\n<li><strong>TTA</strong></li>\n<li><strong>Ensembling/Stacking</strong></li>\n</ul>\n<h2>Other notes</h2>\n<ul>\n<li>In more recent versions, I've changed the implementation from <code>qubvel/callidor</code> to the one in <code>tensorflow.keras.applications</code>. Since it is not officially released for v2.2.0, I made <a href=\"https://www.kaggle.com/xhlulu/tf-keras-efficientnet\" target=\"_blank\">a utility script</a> for the TPU notebook. For the other notebooks, I'm loading directly from <code>tensorflow</code> since they are on v2.3.1</li>\n</ul>",
      "rawMarkdown": "I've published a few notebooks, so I'd like to aggregate them in this post and give more details. This is a work in progress.\n\n## Scores by model\n\nAll the models I trained were EfficientNets using TensorFlow Keras. Depending on the size they were either trained on GPU or TPUs. **The scores were all calculated after the rescoring started so they should be up-to-date**.\n\nYou can find the respective notebooks here:\n1. [Full workflow on GPU (train and submit)](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit)\n2. [Train on TPU](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training)\n3. [Submit on GPU](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-submission)\n\nNote that [1] is for the smaller models (B2 and B3), and [2-3] are for the larger models (B6-B7). Although [1] is simpler, it take a lot of time to submit so I recommend taking a look at [3] if you want to have faster submissions.\n\nHere's a summary table of the scores:\n\n\n| Model | Img Size | Valid AUC | LB    | Accelerator | Weights       | Version |\n|-------|----------|-----------|-------|-------------|---------------|---------|\n| B0    | 224      | 0.9454    | 0.883 | GPU         | Noisy Student | 15      |\n| B1    | 240      | 0.9618    | 0.888 | GPU         | Noisy Student | 14      |\n| B2    | 260      | 0.9278    | 0.918 | GPU         | Noisy Student | 10      |\n| B2    | 260      | 0.9587    | 0.912 | GPU         | Noisy Student | 11      |\n| B2    | 260      | 0.9713    | 0.908 | GPU         | Noisy Student | 12      |\n| B3    | 300      | 0.9202    | 0.909 | GPU         | ImageNet      | 9       |\n| B5    | 456      | 0.9382    | 0.944 | TPU         | ImageNet      | 9       |\n| B6    | 528      | 0.9415    | 0.949 | TPU         | Noisy Student | 6       |\n| B7    | 600      | 0.9455    | 0.953 | TPU         | Noisy Student | 7       |\n| B7    | 600      | 0.9431    | 0.957 | TPU         | ImageNet      | 8       |\n\n## Hyperparameters and training details\n\nHere are some details about how I trained the models:\n\n* **Training augmentation**: Simply random left-right and top-bottom flipping\n* **Optimizer**: Adam with an initial learning rate of 0.001 and no other tuning\n* **Metrics**: Multi-label AUROC (as opposed to flattened), corresponding to the competition metric\n* **Scheduling**: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve\n* **Number of epochs**: 10-15 for early version of GPU, 20 for later version and for TPU\n* **Batch Size**: 16 for GPU and 16*8=128 for the TPU\n* **Model saving** Save model with the best validation AUC after every epoch\n\nI did not try any of the following:\n* **Cross-validation**\n* **TTA**\n* **Ensembling/Stacking**\n\n## Other notes\n\n* In more recent versions, I've changed the implementation from `qubvel/callidor` to the one in `tensorflow.keras.applications`. Since it is not officially released for v2.2.0, I made [a utility script](https://www.kaggle.com/xhlulu/tf-keras-efficientnet) for the TPU notebook. For the other notebooks, I'm loading directly from `tensorflow` since they are on v2.3.1",
      "votes": 148
    },
    {
      "id": 1119390,
      "postDate": "2020-12-20T02:23:35.403Z",
      "content": "<p>latest <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> pytorch mode: ResNet-200D (pretain at imagenet 320x320)<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\n<a href=\"https://twitter.com/wightmanr/status/1340105026786656256?s=20\" target=\"_blank\">https://twitter.com/wightmanr/status/1340105026786656256?s=20</a></p>\n<p>quote \"Compared to an ImageNet-1k only EfficientNet-B5 w/ RandAugment, the 200D is 30% faster on a GPU and better top-1.\"</p>\n<p>single fold, 2xTTA 640x640 resnet200: LB 0.959</p>\n<p>local cv (2xTTA)</p>\n<pre><code>probability : (6017, 11)\nlabel : (6017, 11)\nsubmit_dir : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/valid/local-00009000_model\ninitial_checkpoint : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/checkpoint/00009000_model.pth\n\nfold1:  resnet200 640x640 : LB 0.959\nloss : 0.145014\nauc      : [0.9869, 0.9564, 0.9888, 0.9621, 0.9571, 0.9837, 0.9866, 0.9186, 0.8671, 0.9086, 0.9996]\nauc_flat : 0.978218\nauc_mean : 0.955946\n</code></pre>\n<hr>\n<p>comparison to efficient net</p>\n<pre><code>fold2:  effb2 600x600 : LB 0.951\nloss : 0.143558\nauc      : [0.9564, 0.9547, 0.9899, 0.9479, 0.9264, 0.9757, 0.9817, 0.8984, 0.8354, 0.9020, 0.9996]\nauc_flat : 0.976190\nauc_mean : 0.942560\n\n-----\n\nfold1:  effb5 600x600 : LB 0.944\nloss : 0.153872\nauc      : [0.9907, 0.9513, 0.9890, 0.9624, 0.9551, 0.9796, 0.9847, 0.9298, 0.8604, 0.9086, 0.9989]\nauc_flat : 0.977150\nauc_mean : 0.955487\n\n-----\n\nfold1:  effb7 600x600 : LB 0.954\nloss : 0.165448\nauc      : [0.9892, 0.9479, 0.9870, 0.9543, 0.9447, 0.9807, 0.9844, 0.9139, 0.8565, 0.9044, 0.9995]\nauc_flat : 0.975849\nauc_mean : 0.951142\n</code></pre>",
      "rawMarkdown": "latest @rwightman pytorch mode: ResNet-200D (pretain at imagenet 320x320)\nhttps://github.com/rwightman/pytorch-image-models\nhttps://twitter.com/wightmanr/status/1340105026786656256?s=20\n\nquote \"Compared to an ImageNet-1k only EfficientNet-B5 w/ RandAugment, the 200D is 30% faster on a GPU and better top-1.\"\n\nsingle fold, 2xTTA 640x640 resnet200: LB 0.959\n\nlocal cv (2xTTA)\n```\nprobability : (6017, 11)\nlabel : (6017, 11)\nsubmit_dir : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/valid/local-00009000_model\ninitial_checkpoint : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/checkpoint/00009000_model.pth\n\nfold1:  resnet200 640x640 : LB 0.959\nloss : 0.145014\nauc      : [0.9869, 0.9564, 0.9888, 0.9621, 0.9571, 0.9837, 0.9866, 0.9186, 0.8671, 0.9086, 0.9996]\nauc_flat : 0.978218\nauc_mean : 0.955946\n```\n\n---\n\ncomparison to efficient net\n\n```\nfold2:  effb2 600x600 : LB 0.951\nloss : 0.143558\nauc      : [0.9564, 0.9547, 0.9899, 0.9479, 0.9264, 0.9757, 0.9817, 0.8984, 0.8354, 0.9020, 0.9996]\nauc_flat : 0.976190\nauc_mean : 0.942560\n\n-----\n\nfold1:  effb5 600x600 : LB 0.944\nloss : 0.153872\nauc      : [0.9907, 0.9513, 0.9890, 0.9624, 0.9551, 0.9796, 0.9847, 0.9298, 0.8604, 0.9086, 0.9989]\nauc_flat : 0.977150\nauc_mean : 0.955487\n\n-----\n\nfold1:  effb7 600x600 : LB 0.954\nloss : 0.165448\nauc      : [0.9892, 0.9479, 0.9870, 0.9543, 0.9447, 0.9807, 0.9844, 0.9139, 0.8565, 0.9044, 0.9995]\nauc_flat : 0.975849\nauc_mean : 0.951142\n\n```",
      "votes": 12,
      "replies": [
        {
          "id": 1119407,
          "postDate": "2020-12-20T02:56:59.037Z",
          "content": "<p>What are 2xTTA if I may ask? Hflip and vflip? Thanks.</p>",
          "rawMarkdown": "What are 2xTTA if I may ask? Hflip and vflip? Thanks."
        },
        {
          "id": 1119409,
          "postDate": "2020-12-20T03:01:54.023Z",
          "content": "<p>orginal + hflip</p>",
          "rawMarkdown": "orginal + hflip",
          "votes": 3
        },
        {
          "id": 1119437,
          "postDate": "2020-12-20T04:25:07.617Z",
          "content": "<p>How do you use 'resnet200d' .I use timm to load 'resnet200d' but I get some error like this:'Pretrained model URL is invalid, using random initialization.'Thanks!</p>",
          "rawMarkdown": "How do you use 'resnet200d' .I use timm to load 'resnet200d' but I get some error like this:'Pretrained model URL is invalid, using random initialization.'Thanks!"
        },
        {
          "id": 1119440,
          "postDate": "2020-12-20T04:28:31.217Z",
          "content": "<p>you can download \"resnet200d_ra2-bdba9bf9.pth\" manually and use torch to load it:</p>\n<pre><code>    net = ResNet(\n        Bottleneck,\n        layers=[3, 24, 36, 3],\n        num_classes=1000,\n        in_chans=3,\n        cardinality=1,\n        base_width=64,\n        stem_width=32,\n        stem_type='deep',\n        output_stride=32,\n        block_reduce_first=1,\n        down_kernel_size=1,\n        avg_down=True,\n        act_layer=nn.ReLU,\n        norm_layer=nn.BatchNorm2d,\n        aa_layer=None,\n        drop_rate=0.0,\n        drop_path_rate=0.,\n        drop_block_rate=0.,\n        global_pool='avg',\n        zero_init_last_bn=True,\n        block_args=None\n    )\n    print(net)\n    pretrain_state_dict = torch.load('resnet200d_ra2-bdba9bf9.pth', map_location=lambda storage, loc: storage)\n\n    s = net.load_state_dict(pretrain_state_dict, strict=True)\n    print(s)\n</code></pre>",
          "rawMarkdown": "you can download \"resnet200d_ra2-bdba9bf9.pth\" manually and use torch to load it:\n\n```\n    net = ResNet(\n        Bottleneck,\n        layers=[3, 24, 36, 3],\n        num_classes=1000,\n        in_chans=3,\n        cardinality=1,\n        base_width=64,\n        stem_width=32,\n        stem_type='deep',\n        output_stride=32,\n        block_reduce_first=1,\n        down_kernel_size=1,\n        avg_down=True,\n        act_layer=nn.ReLU,\n        norm_layer=nn.BatchNorm2d,\n        aa_layer=None,\n        drop_rate=0.0,\n        drop_path_rate=0.,\n        drop_block_rate=0.,\n        global_pool='avg',\n        zero_init_last_bn=True,\n        block_args=None\n    )\n    print(net)\n    pretrain_state_dict = torch.load('resnet200d_ra2-bdba9bf9.pth', map_location=lambda storage, loc: storage)\n\n    s = net.load_state_dict(pretrain_state_dict, strict=True)\n    print(s)\n\n```",
          "votes": 3
        },
        {
          "id": 1119444,
          "postDate": "2020-12-20T04:35:47.743Z",
          "content": "<p>Thanks!By the way,what batch size do you use?Thanks!</p>",
          "rawMarkdown": "Thanks!By the way,what batch size do you use?Thanks!"
        },
        {
          "id": 1119455,
          "postDate": "2020-12-20T04:51:28.933Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1215649,
      "postDate": "2021-02-23T20:32:39.870Z",
      "content": "<p>Great Work , thanks a lot for sharing , </p>\n<p>How much time it takes for training with EfficientNet B7 ? </p>",
      "rawMarkdown": "Great Work , thanks a lot for sharing , \n\nHow much time it takes for training with EfficientNet B7 ? ",
      "votes": 7
    },
    {
      "id": 1214584,
      "postDate": "2021-02-23T01:35:27.873Z",
      "content": "<p>Why is the valid AUC basically the same for all, but the LB increases?</p>",
      "rawMarkdown": "Why is the valid AUC basically the same for all, but the LB increases?",
      "votes": 5
    },
    {
      "id": 1149264,
      "postDate": "2021-01-11T18:17:28.413Z",
      "content": "<p>When you train these models from pretrained weights, do you first only train the new fully connected top layer for a few epochs and then train everything together afterwards, or do you just train everything right from the start?</p>",
      "rawMarkdown": "When you train these models from pretrained weights, do you first only train the new fully connected top layer for a few epochs and then train everything together afterwards, or do you just train everything right from the start?",
      "votes": 3
    },
    {
      "id": 1126176,
      "postDate": "2020-12-25T12:08:51.270Z",
      "content": "<p>Isn't vertical flip wrong given the fact that x-rays will come only in upright position?</p>",
      "rawMarkdown": "Isn't vertical flip wrong given the fact that x-rays will come only in upright position?",
      "votes": 3,
      "replies": [
        {
          "id": 1128499,
          "postDate": "2020-12-27T13:39:50.097Z",
          "content": "<p>I've debated that for myself and did not use it for that reason in <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">the notebook I made public</a> and in what I otherwise do (and also limited the extent of rotations). I guess a human radiologist could kind of look at the picture the wrong way around and cope (but most likely would of course flip it around…)? In the end, if it helps the model (as indicated by CV), I would still use it, I just put it very low down on my priority list of things to try.</p>",
          "rawMarkdown": "I've debated that for myself and did not use it for that reason in [the notebook I made public](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) and in what I otherwise do (and also limited the extent of rotations). I guess a human radiologist could kind of look at the picture the wrong way around and cope (but most likely would of course flip it around...)? In the end, if it helps the model (as indicated by CV), I would still use it, I just put it very low down on my priority list of things to try."
        },
        {
          "id": 1128543,
          "postDate": "2020-12-27T14:32:07.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> what you mentioned it's not trivial. In theory ,we do augmentations after see what images the model missclasified and choose the augmentations that benefit that images. By example, if the model missclasified  a image with zoom, we use a zoom augmentation. This time i don't see any pattern in the missclassified images, however it's still a good idea to use augmentations as a <strong>regularization</strong> technique. </p>",
          "rawMarkdown": "@harveenchadha what you mentioned it's not trivial. In theory ,we do augmentations after see what images the model missclasified and choose the augmentations that benefit that images. By example, if the model missclasified  a image with zoom, we use a zoom augmentation. This time i don't see any pattern in the missclassified images, however it's still a good idea to use augmentations as a **regularization** technique. "
        },
        {
          "id": 1128551,
          "postDate": "2020-12-27T14:42:08.303Z",
          "content": "<p>you just need \"more data\". here are some insights</p>\n<ol>\n<li><p>last time in the early days, people create \"virtual support sample\" to change the boundary of SVM. it is found that these \"virtual sample\" does not resemble the actual samples.</p></li>\n<li><p>in some papers, people use extreme augmentation to create more data and treat them as unlabelled. (the augmentation is too extreme to destroy the original label). then weak/self supervised learning method is applied.</p></li>\n<li><p>it is found that by adding noise, deep learning actually improved and become more robust</p></li>\n<li><p>while a sample that is adversarial modified looks very much like the original sample, the network can be easily fooled (aka adversarial attack). the way human looks at the sample is not the same as the machine would</p></li>\n</ol>\n<p>rubbish in, rubbish out …<br>\ngold in, rubbish out …<br>\nrubbish in, gold out … !!!</p>",
          "rawMarkdown": "you just need \"more data\". here are some insights\n\n1. last time in the early days, people create \"virtual support sample\" to change the boundary of SVM. it is found that these \"virtual sample\" does not resemble the actual samples.\n\n2. in some papers, people use extreme augmentation to create more data and treat them as unlabelled. (the augmentation is too extreme to destroy the original label). then weak/self supervised learning method is applied.\n\n3. it is found that by adding noise, deep learning actually improved and become more robust\n\n4. while a sample that is adversarial modified looks very much like the original sample, the network can be easily fooled (aka adversarial attack). the way human looks at the sample is not the same as the machine would\n\nrubbish in, rubbish out ...\ngold in, rubbish out ...\nrubbish in, gold out ... !!!",
          "votes": 4
        },
        {
          "id": 1135906,
          "postDate": "2021-01-02T15:34:54.803Z",
          "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> The effect of specific augmentation methods can be easily tested. Here is the recipe:</p>\n<ol>\n<li>Train the model without the vertical flip, record the validation auc.</li>\n<li>Train the model with the vertical flip, record the validation auc.</li>\n</ol>\n<p>If the validation auc in 1 is greater than the validation auc in 2, vertical flip is useless.</p>",
          "rawMarkdown": "@harveenchadha The effect of specific augmentation methods can be easily tested. Here is the recipe:\n1. Train the model without the vertical flip, record the validation auc.\n2. Train the model with the vertical flip, record the validation auc.\n\nIf the validation auc in 1 is greater than the validation auc in 2, vertical flip is useless.",
          "votes": 1
        },
        {
          "id": 1136099,
          "postDate": "2021-01-02T18:14:35.857Z",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> Yup that is right unless you don't have problem of unlimited compute power :) </p>",
          "rawMarkdown": "@tolgadincer Yup that is right unless you don't have problem of unlimited compute power :) "
        }
      ]
    },
    {
      "id": 1117025,
      "postDate": "2020-12-17T16:50:21.650Z",
      "content": "<p>thanks for the results! <br>\ni wonder if the improvement is due to the image size?<br>\ne.g. i have efficientnetb2 on 512x512 and get valid AUC of 0.936+</p>",
      "rawMarkdown": "thanks for the results! \ni wonder if the improvement is due to the image size?\ne.g. i have efficientnetb2 on 512x512 and get valid AUC of 0.936+",
      "votes": 3,
      "replies": [
        {
          "id": 1117028,
          "postDate": "2020-12-17T16:53:17.547Z",
          "content": "<p>I'm suspecting that my model has been underfitting, since the training/validation losses are very close to each other. Will try to run it for longer (15 epochs) and see if it improves.</p>",
          "rawMarkdown": "I'm suspecting that my model has been underfitting, since the training/validation losses are very close to each other. Will try to run it for longer (15 epochs) and see if it improves.",
          "votes": 1
        },
        {
          "id": 1117346,
          "postDate": "2020-12-18T00:54:00.920Z",
          "content": "<p>Hi,I use the efficientnetb4 on 512x512,and 5 fold.But I get local cv 0.9078 and lb 0.935.Do you know this is why?why you use efficientnetb2 can get 0.936?Thanks!</p>",
          "rawMarkdown": "Hi,I use the efficientnetb4 on 512x512,and 5 fold.But I get local cv 0.9078 and lb 0.935.Do you know this is why?why you use efficientnetb2 can get 0.936?Thanks!"
        },
        {
          "id": 1118536,
          "postDate": "2020-12-19T06:54:18.557Z",
          "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> </p>\n<p>i confirm the effectiveness of using deeper network like B5,B7. I publish some results above. Here are some quick note:</p>\n<ul>\n<li>size alone does not explain the improvement. e.g. 600-efficientnetb2 is slight better than 512-efficientnetb2. But 600-efficientnetb5 can make a huge difference.<br>\n(I think it is the larger receptive field and this can be verified by CAM map. I haven't investigated the shortcut-drop-rate parameter, which can also be a reason. efficientnetb5 has higher shortcut drop rate)</li>\n</ul>",
          "rawMarkdown": "@xhlulu \n\ni confirm the effectiveness of using deeper network like B5,B7. I publish some results above. Here are some quick note:\n\n- size alone does not explain the improvement. e.g. 600-efficientnetb2 is slight better than 512-efficientnetb2. But 600-efficientnetb5 can make a huge difference.\n(I think it is the larger receptive field and this can be verified by CAM map. I haven't investigated the shortcut-drop-rate parameter, which can also be a reason. efficientnetb5 has higher shortcut drop rate)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1147758,
      "postDate": "2021-01-10T17:15:24.283Z",
      "content": "<p>Thank you so much for your works, it has really motivated me to explore this competition more now</p>",
      "rawMarkdown": "Thank you so much for your works, it has really motivated me to explore this competition more now\n",
      "votes": 1
    },
    {
      "id": 1118246,
      "postDate": "2020-12-18T21:56:50.957Z",
      "content": "<p>With my efficientB0 model, I got training AUC of 0.91, validation of 0.90, and LB of 0.90. So, my understanding is I have an underfitting problem. What do you suggest (other than increasing the image size) on GPU? </p>",
      "rawMarkdown": "With my efficientB0 model, I got training AUC of 0.91, validation of 0.90, and LB of 0.90. So, my understanding is I have an underfitting problem. What do you suggest (other than increasing the image size) on GPU? ",
      "votes": 1,
      "replies": [
        {
          "id": 1121268,
          "postDate": "2020-12-21T14:02:41.243Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/sinamhd9\" target=\"_blank\">@sinamhd9</a>, at beginning I refer your notebook for my first submission, but with modified parameters which really work well. i.e</p>\n<ol>\n<li>Using <strong>EfficientB4</strong> with lower batchsize(10~14)</li>\n<li>From my observation, <strong>vertical flipping not contributes significantly in model generalization</strong>.</li>\n<li>Using K-fold methods is better.</li>\n</ol>\n<p>Hope it helps <br>\nThanks <br>\n~ AKhilesh</p>",
          "rawMarkdown": "Hey @sinamhd9, at beginning I refer your notebook for my first submission, but with modified parameters which really work well. i.e\n1. Using **EfficientB4** with lower batchsize(10~14)\n2. From my observation, **vertical flipping not contributes significantly in model generalization**.\n3. Using K-fold methods is better.\n\nHope it helps \nThanks \n~ AKhilesh",
          "votes": 1
        }
      ]
    },
    {
      "id": 1118241,
      "postDate": "2020-12-18T21:43:13.600Z",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> Thanks for sharing! I am using more augmentation methods such as (brightness change, shear, rotation, etc..). <a href=\"https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1\" target=\"_blank\">https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1</a></p>\n<p>Since this is a medical application I am not sure if flipping horizontally or vertically will be problematic or not. What do you think?</p>",
      "rawMarkdown": "@xhlulu Thanks for sharing! I am using more augmentation methods such as (brightness change, shear, rotation, etc..). https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1\n\nSince this is a medical application I am not sure if flipping horizontally or vertically will be problematic or not. What do you think?",
      "votes": 1,
      "replies": [
        {
          "id": 1118257,
          "postDate": "2020-12-18T22:13:25.560Z",
          "content": "<p>From a theoretical perspective I don't think it poses a problem per se, but a rotation might be more effective since it would more closely represent what would happen in real life (scans in different positions).</p>\n<p>Moreover I'm not an expert in medical imaging, I think Dr. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> would likely have a better idea about this than me.</p>",
          "rawMarkdown": "From a theoretical perspective I don't think it poses a problem per se, but a rotation might be more effective since it would more closely represent what would happen in real life (scans in different positions).\n\nMoreover I'm not an expert in medical imaging, I think Dr. @vaillant would likely have a better idea about this than me.",
          "votes": 1
        },
        {
          "id": 1118259,
          "postDate": "2020-12-18T22:15:17.440Z",
          "content": "<p>I'm also not sure about the brightness change and shear, since those don't seem to be common attributes of CT scans, but again I might be wrong.</p>",
          "rawMarkdown": "I'm also not sure about the brightness change and shear, since those don't seem to be common attributes of CT scans, but again I might be wrong."
        },
        {
          "id": 1118269,
          "postDate": "2020-12-18T22:29:25.067Z",
          "content": "<p>Thanks for your reply! Me neither, I was curious to see if they have an effect on better generalization or not.</p>",
          "rawMarkdown": "Thanks for your reply! Me neither, I was curious to see if they have an effect on better generalization or not.",
          "votes": 1
        },
        {
          "id": 1118286,
          "postDate": "2020-12-18T22:57:06.737Z",
          "content": "<p>Clinical implications aside, I think they are worth trying in case they do result in better generalization.</p>",
          "rawMarkdown": "Clinical implications aside, I think they are worth trying in case they do result in better generalization."
        },
        {
          "id": 1118335,
          "postDate": "2020-12-19T00:50:18.040Z",
          "content": "<p>Technically, horizontal/vertical flipping would result in images that we would not see regularly in clinical practice. </p>\n<p>However, I use these augmentations during training anyways to reduce overfitting. </p>",
          "rawMarkdown": "Technically, horizontal/vertical flipping would result in images that we would not see regularly in clinical practice. \n\nHowever, I use these augmentations during training anyways to reduce overfitting. ",
          "votes": 5
        },
        {
          "id": 1118338,
          "postDate": "2020-12-19T01:07:30.180Z",
          "content": "<p>That totally makes sense! Thank you for sharing!</p>",
          "rawMarkdown": "That totally makes sense! Thank you for sharing!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1117360,
      "postDate": "2020-12-18T01:19:26.103Z",
      "content": "<p>Interesting. Is there a reason you skipped B4 and B5 (other than you can't do it all?)? I guess, you might also argue that B0 to B2 might be good to experimenting quickly, while B6/B7 are final model material, while the ones inbetween are neither here nor there? Was that the thinking?</p>\n<p>Mostly as a compromise between time and performance, my first try was this stuff that I shared with an EfficientNet-B4 (vanilla B4, not noisy student) that gets close to the B6 results you posted (LB 0.945) (<a href=\"https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4\" target=\"_blank\">inference</a> and <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">training</a> notebooks). I'd guess that's because you did not use much augmentation &amp; a simple LR schedule (reduce LR on plateau vs. some augmentation &amp; cosine-annealing in my case)? </p>\n<p>And, wow, B7 trained fast on TPU!</p>",
      "rawMarkdown": "Interesting. Is there a reason you skipped B4 and B5 (other than you can't do it all?)? I guess, you might also argue that B0 to B2 might be good to experimenting quickly, while B6/B7 are final model material, while the ones inbetween are neither here nor there? Was that the thinking?\n\nMostly as a compromise between time and performance, my first try was this stuff that I shared with an EfficientNet-B4 (vanilla B4, not noisy student) that gets close to the B6 results you posted (LB 0.945) ([inference](https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4) and [training](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) notebooks). I'd guess that's because you did not use much augmentation & a simple LR schedule (reduce LR on plateau vs. some augmentation & cosine-annealing in my case)? \n\nAnd, wow, B7 trained fast on TPU!",
      "votes": 2,
      "replies": [
        {
          "id": 1117364,
          "postDate": "2020-12-18T01:29:59.870Z",
          "content": "<p>I haven't had the bandwidth to try everything, but I will come to those eventually!</p>",
          "rawMarkdown": "I haven't had the bandwidth to try everything, but I will come to those eventually!"
        }
      ]
    },
    {
      "id": 1242003,
      "postDate": "2021-03-17T10:39:47.827Z",
      "content": "<p>Thank you for your work. Unfortunatelly I was struggling to achieve decent performance with a copy of your notebook (auc below 0.7). Anyone else had similar problem?</p>",
      "rawMarkdown": "Thank you for your work. Unfortunatelly I was struggling to achieve decent performance with a copy of your notebook (auc below 0.7). Anyone else had similar problem?"
    },
    {
      "id": 1146190,
      "postDate": "2021-01-09T15:35:43.917Z",
      "content": "<p><a href=\"https://www.kaggle.com/xhulu\" target=\"_blank\">@xhulu</a> thank you very much for the awesome notebooks - you're the best!!!! I have tried running your notebooks and I have always ended up (after 20 epochs) with around 0.85 val_auc, whereas your table is showing 0.95 and above. Are the results you have presented from the first run of the models that you list or are you using another method that we are not aware of? Thank you in advance.</p>",
      "rawMarkdown": "@xhulu thank you very much for the awesome notebooks - you're the best!!!! I have tried running your notebooks and I have always ended up (after 20 epochs) with around 0.85 val_auc, whereas your table is showing 0.95 and above. Are the results you have presented from the first run of the models that you list or are you using another method that we are not aware of? Thank you in advance.",
      "replies": [
        {
          "id": 1147660,
          "postDate": "2021-01-10T16:25:18.480Z",
          "content": "<p><a href=\"https://www.kaggle.com/samuelliebana\" target=\"_blank\">@samuelliebana</a> , ran <a href=\"https://www.kaggle.com/xhulu\" target=\"_blank\">@xhulu</a>  's notebook successfully. The lower value may attribute to mis-matching image_size you might be passing to <code>efficientNetB0</code>. Can you check? In addition, check the prediction for test image, did you forget to update the image_size? </p>",
          "rawMarkdown": "@samuelliebana , ran @xhulu  's notebook successfully. The lower value may attribute to mis-matching image_size you might be passing to `efficientNetB0`. Can you check? In addition, check the prediction for test image, did you forget to update the image_size? ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1144682,
      "postDate": "2021-01-08T15:54:19.627Z",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>, Thank you sharing simply GOLD notebooks! Get to learn a lot. </p>",
      "rawMarkdown": "@xhlulu, Thank you sharing simply GOLD notebooks! Get to learn a lot. "
    },
    {
      "id": 1135315,
      "postDate": "2021-01-02T05:58:36.050Z",
      "content": "<p>Great jobs</p>",
      "rawMarkdown": "Great jobs"
    },
    {
      "id": 1128467,
      "postDate": "2020-12-27T13:12:02.330Z",
      "content": "<p>It takes an hour and a half to train one <strong>B7</strong> model, on TPU, how do you plan to do cross-validation if TPU stops after 3 hours?</p>",
      "rawMarkdown": "It takes an hour and a half to train one **B7** model, on TPU, how do you plan to do cross-validation if TPU stops after 3 hours?",
      "replies": [
        {
          "id": 1128501,
          "postDate": "2020-12-27T13:41:31.013Z",
          "content": "<p>Not sure what the problem is - you have 30 hours per week and there's several weeks to go. As long as inference is fast enough, training time would not seem to be a problem (if you train in separate notebooks and use the trained models for inference)?</p>",
          "rawMarkdown": "Not sure what the problem is - you have 30 hours per week and there's several weeks to go. As long as inference is fast enough, training time would not seem to be a problem (if you train in separate notebooks and use the trained models for inference)?"
        },
        {
          "id": 1128534,
          "postDate": "2020-12-27T14:25:25.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> my bad, the comment has been corrected.</p>\n<p>We have 30 hours per week, but the TPU stops after 3 hours, if a model takes 1 hour and a half to train, and we want to do, by example 3 fold cross-validation, the 3 hours of TPU wouldn't be enough.</p>",
          "rawMarkdown": "@bjoernholzhauer my bad, the comment has been corrected.\n\nWe have 30 hours per week, but the TPU stops after 3 hours, if a model takes 1 hour and a half to train, and we want to do, by example 3 fold cross-validation, the 3 hours of TPU wouldn't be enough."
        },
        {
          "id": 1128556,
          "postDate": "2020-12-27T14:44:28.627Z",
          "content": "<p>Just run multiple notebooks, one per fold. It's not like you could not have multiple notebooks as inputs to a submission kernel. Sure, it's not very elegant, but totally possible.</p>",
          "rawMarkdown": "Just run multiple notebooks, one per fold. It's not like you could not have multiple notebooks as inputs to a submission kernel. Sure, it's not very elegant, but totally possible.",
          "votes": 1
        },
        {
          "id": 1128558,
          "postDate": "2020-12-27T14:45:53.967Z",
          "content": "<p>Whoa, that's something new, i'm going to try it, thanks for the idea.</p>",
          "rawMarkdown": "Whoa, that's something new, i'm going to try it, thanks for the idea.",
          "votes": 1
        },
        {
          "id": 1128969,
          "postDate": "2020-12-28T00:11:26.783Z",
          "content": "<p>I don't recommend playing with larger models and larger resolution at the beginning of a competition.😄</p>",
          "rawMarkdown": "I don't recommend playing with larger models and larger resolution at the beginning of a competition.😄",
          "votes": 3
        },
        {
          "id": 1128977,
          "postDate": "2020-12-28T00:35:42.967Z",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>, recently i did read the book \"Deep Learning for Coders with Fastai and PyTorch: AI Applications Without a PhD\", where the author wrote that we must try first complex models and finetune if it doesn't work well, after that we can use less complex models. Now that's part of my way to approach any problem with DL.</p>",
          "rawMarkdown": "@underwearfitting, recently i did read the book \"Deep Learning for Coders with Fastai and PyTorch: AI Applications Without a PhD\", where the author wrote that we must try first complex models and finetune if it doesn't work well, after that we can use less complex models. Now that's part of my way to approach any problem with DL."
        },
        {
          "id": 1129085,
          "postDate": "2020-12-28T04:28:04.843Z",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> you mean you won your first medal in MOA? That competition is not related to computer vision at all. The time it takes to run an experiment in MOA is nearly negligible. Maybe you wanna watch this video: <a href=\"https://www.youtube.com/watch?v=L1QKTPb6V_I\" target=\"_blank\">https://www.youtube.com/watch?v=L1QKTPb6V_I</a>.</p>",
          "rawMarkdown": "@hiramcho you mean you won your first medal in MOA? That competition is not related to computer vision at all. The time it takes to run an experiment in MOA is nearly negligible. Maybe you wanna watch this video: https://www.youtube.com/watch?v=L1QKTPb6V_I.",
          "votes": 7
        },
        {
          "id": 1129561,
          "postDate": "2020-12-28T12:40:41.373Z",
          "content": "<p>I understand your concern but if I have desired results I don't want to Change my approach. If at the end of this competition doesn't works, i've learned that you was right. </p>",
          "rawMarkdown": "I understand your concern but if I have desired results I don't want to Change my approach. If at the end of this competition doesn't works, i've learned that you was right. "
        },
        {
          "id": 1129858,
          "postDate": "2020-12-28T15:43:02.087Z",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Thanks For the link I really enjoyed watching it )) </p>",
          "rawMarkdown": "@underwearfitting Thanks For the link I really enjoyed watching it )) ",
          "votes": 3
        },
        {
          "id": 1162375,
          "postDate": "2021-01-21T05:46:03.240Z",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> - in case you or others have not seen this, TPU has same time as GPU now.</p>\n<p><a href=\"https://www.kaggle.com/product-feedback/202409\" target=\"_blank\">https://www.kaggle.com/product-feedback/202409</a><br>\nTPU execution time increased to 9 hours!</p>",
          "rawMarkdown": "@hiramcho - in case you or others have not seen this, TPU has same time as GPU now.\n\nhttps://www.kaggle.com/product-feedback/202409\nTPU execution time increased to 9 hours!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1118906,
      "postDate": "2020-12-19T14:13:07.357Z",
      "content": "<p>Hello, what GPU did you use?</p>",
      "rawMarkdown": "Hello, what GPU did you use?",
      "replies": [
        {
          "id": 1119016,
          "postDate": "2020-12-19T16:23:45.743Z",
          "content": "<p>I think kaggle uses p100</p>",
          "rawMarkdown": "I think kaggle uses p100"
        }
      ]
    },
    {
      "id": 1118393,
      "postDate": "2020-12-19T03:07:52.740Z",
      "content": "<p><code>0.9713</code> with new metrics?</p>",
      "rawMarkdown": "`0.9713` with new metrics?"
    },
    {
      "id": 1118201,
      "postDate": "2020-12-18T20:33:58.470Z",
      "content": "<p>LB score compares between GPU and TPU is really noticeable. I like to see if someone adds some PyTorch observation GPU. </p>",
      "rawMarkdown": "LB score compares between GPU and TPU is really noticeable. I like to see if someone adds some PyTorch observation GPU. ",
      "replies": [
        {
          "id": 1118206,
          "postDate": "2020-12-18T20:40:12.800Z",
          "content": "<p>I don't think the accelerator makes a difference per se, but the model size (B3 vs B7) and the batch size is an important reason. I'm using a batch size of 16 for GPU and 16*8=128 for the TPU. </p>",
          "rawMarkdown": "I don't think the accelerator makes a difference per se, but the model size (B3 vs B7) and the batch size is an important reason. I'm using a batch size of 16 for GPU and 16*8=128 for the TPU. ",
          "votes": 3
        },
        {
          "id": 1118221,
          "postDate": "2020-12-18T21:02:59.533Z",
          "content": "<p>Of course, I didn't mean accelerator but the advantages of using TPU. </p>",
          "rawMarkdown": "Of course, I didn't mean accelerator but the advantages of using TPU. ",
          "votes": 1
        },
        {
          "id": 1118563,
          "postDate": "2020-12-19T07:32:38.187Z",
          "content": "<p>Note also that some of the pre-trained models are <a href=\"https://arxiv.org/abs/1911.04252\" target=\"_blank\">noisy student </a> ones (do better on ImageNet, which you might suspect could improve performance here, too) and some are not. That makes the picture a little less clear.</p>",
          "rawMarkdown": "Note also that some of the pre-trained models are [noisy student ](https://arxiv.org/abs/1911.04252) ones (do better on ImageNet, which you might suspect could improve performance here, too) and some are not. That makes the picture a little less clear.",
          "votes": 2
        },
        {
          "id": 1119021,
          "postDate": "2020-12-19T16:27:09.277Z",
          "content": "<p>I'm super curious about whether noisy students makes a difference on downstream task, or they are over fitting on imagenet. I wish there's a uniform image recognition benchmark for this.</p>",
          "rawMarkdown": "I'm super curious about whether noisy students makes a difference on downstream task, or they are over fitting on imagenet. I wish there's a uniform image recognition benchmark for this.",
          "votes": 1
        },
        {
          "id": 1119146,
          "postDate": "2020-12-19T18:42:29.147Z",
          "content": "<p>Noisy student didn't make the cut-off for <a href=\"https://arxiv.org/abs/1902.10811#:~:text=Our%20results%20suggest%20that%20the,in%20the%20original%20test%20sets.\" target=\"_blank\">Do ImageNet Classifiers Generalize to ImageNet</a>, but their dataset would be one logical thing to try. I'd assume the noisy student models would do well there (but that's the exact same task and mostly assesses overfitting to the test set.).</p>\n<p>For down-stream tasks, I do wonder whether Kaggle competitions would not be one of the more obvious choices, or perhaps Open Images or the COCO Dataset?</p>",
          "rawMarkdown": "Noisy student didn't make the cut-off for [Do ImageNet Classifiers Generalize to ImageNet](https://arxiv.org/abs/1902.10811#:~:text=Our%20results%20suggest%20that%20the,in%20the%20original%20test%20sets.), but their dataset would be one logical thing to try. I'd assume the noisy student models would do well there (but that's the exact same task and mostly assesses overfitting to the test set.).\n\nFor down-stream tasks, I do wonder whether Kaggle competitions would not be one of the more obvious choices, or perhaps Open Images or the COCO Dataset?"
        }
      ]
    },
    {
      "id": 1117115,
      "postDate": "2020-12-17T18:30:48.803Z",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> <br>\nThanks for sharing. :)</p>\n<p>I've gone through the GPU implementation of yours. I saw you didn't do Group-KFold, but just plain train-test-split. Is there any particular reason or you've planned to add it later? </p>",
      "rawMarkdown": "@xhlulu \nThanks for sharing. :)\n\nI've gone through the GPU implementation of yours. I saw you didn't do Group-KFold, but just plain train-test-split. Is there any particular reason or you've planned to add it later? \n\n\n",
      "replies": [
        {
          "id": 1117284,
          "postDate": "2020-12-17T21:50:08.197Z",
          "content": "<p>It usually takes a lot of time to train a GPU model, so adding K-Fold will increase the time by K with only minimal gains (+/- 1%). As a consequence I quickly run out of GPU quota, which I could've used to experiment with crucial hyperparameters (e.g. learning rate, optimizer, model size, image size).</p>\n<p>I feel that publishing very simple notebooks like these ones give a chance to others to try out their favourite techniques without adding too much computation overhead and cognitive loads, and also share their solutions as notebooks (although it will eventually surpass the baseline score). Finally single models are more useful clinically since they are faster to run and can be interpreted using traditional methods (e.g. GradCAM).</p>",
          "rawMarkdown": "It usually takes a lot of time to train a GPU model, so adding K-Fold will increase the time by K with only minimal gains (+/- 1%). As a consequence I quickly run out of GPU quota, which I could've used to experiment with crucial hyperparameters (e.g. learning rate, optimizer, model size, image size).\n\nI feel that publishing very simple notebooks like these ones give a chance to others to try out their favourite techniques without adding too much computation overhead and cognitive loads, and also share their solutions as notebooks (although it will eventually surpass the baseline score). Finally single models are more useful clinically since they are faster to run and can be interpreted using traditional methods (e.g. GradCAM).",
          "votes": 10
        },
        {
          "id": 1117343,
          "postDate": "2020-12-18T00:52:15.823Z",
          "content": "<p>I agree. I was just wondering why you didn't use that 😄</p>\n<p>However, with group k-fold training, setting with <code>multi_label=True</code> gave me pretty constant val AUC whereas without it (False), the val AUC was pretty good (~0.9). I'm not quite sure, maybe computing with column-wise AUC was giving constant output due to the imbalance data point or something else. </p>",
          "rawMarkdown": "I agree. I was just wondering why you didn't use that 😄\n\nHowever, with group k-fold training, setting with `multi_label=True` gave me pretty constant val AUC whereas without it (False), the val AUC was pretty good (~0.9). I'm not quite sure, maybe computing with column-wise AUC was giving constant output due to the imbalance data point or something else. "
        }
      ]
    },
    {
      "id": 1117037,
      "postDate": "2020-12-17T17:06:27.073Z",
      "content": "<p>Thanks for sharing!<br>\np.s. It looks like evaluation metric has been updated.</p>",
      "rawMarkdown": "Thanks for sharing!\np.s. It looks like evaluation metric has been updated.",
      "replies": [
        {
          "id": 1117060,
          "postDate": "2020-12-17T17:30:13.530Z",
          "content": "<p>It seems like the top 10 have not been updated yet based on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> observations</p>",
          "rawMarkdown": "It seems like the top 10 have not been updated yet based on @cdeotte observations"
        }
      ]
    },
    {
      "id": 1214573,
      "postDate": "2021-02-23T01:25:21.047Z",
      "content": "<p>Great experiments. Thanks for sharing!</p>",
      "rawMarkdown": "Great experiments. Thanks for sharing!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1119390,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-20T02:23:35.403000",
      "content": "<p>latest <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> pytorch mode: ResNet-200D (pretain at imagenet 320x320)<br>\n<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\n<a href=\"https://twitter.com/wightmanr/status/1340105026786656256?s=20\" target=\"_blank\">https://twitter.com/wightmanr/status/1340105026786656256?s=20</a></p>\n<p>quote \"Compared to an ImageNet-1k only EfficientNet-B5 w/ RandAugment, the 200D is 30% faster on a GPU and better top-1.\"</p>\n<p>single fold, 2xTTA 640x640 resnet200: LB 0.959</p>\n<p>local cv (2xTTA)</p>\n<pre><code>probability : (6017, 11)\nlabel : (6017, 11)\nsubmit_dir : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/valid/local-00009000_model\ninitial_checkpoint : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/checkpoint/00009000_model.pth\n\nfold1:  resnet200 640x640 : LB 0.959\nloss : 0.145014\nauc      : [0.9869, 0.9564, 0.9888, 0.9621, 0.9571, 0.9837, 0.9866, 0.9186, 0.8671, 0.9086, 0.9996]\nauc_flat : 0.978218\nauc_mean : 0.955946\n</code></pre>\n<hr>\n<p>comparison to efficient net</p>\n<pre><code>fold2:  effb2 600x600 : LB 0.951\nloss : 0.143558\nauc      : [0.9564, 0.9547, 0.9899, 0.9479, 0.9264, 0.9757, 0.9817, 0.8984, 0.8354, 0.9020, 0.9996]\nauc_flat : 0.976190\nauc_mean : 0.942560\n\n-----\n\nfold1:  effb5 600x600 : LB 0.944\nloss : 0.153872\nauc      : [0.9907, 0.9513, 0.9890, 0.9624, 0.9551, 0.9796, 0.9847, 0.9298, 0.8604, 0.9086, 0.9989]\nauc_flat : 0.977150\nauc_mean : 0.955487\n\n-----\n\nfold1:  effb7 600x600 : LB 0.954\nloss : 0.165448\nauc      : [0.9892, 0.9479, 0.9870, 0.9543, 0.9447, 0.9807, 0.9844, 0.9139, 0.8565, 0.9044, 0.9995]\nauc_flat : 0.975849\nauc_mean : 0.951142\n</code></pre>",
      "votes": 12,
      "replies": [
        {
          "id": 1119407,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-12-20T02:56:59.037000",
          "content": "<p>What are 2xTTA if I may ask? Hflip and vflip? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119409,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-20T03:01:54.023000",
          "content": "<p>orginal + hflip</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1119437,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2020-12-20T04:25:07.617000",
          "content": "<p>How do you use 'resnet200d' .I use timm to load 'resnet200d' but I get some error like this:'Pretrained model URL is invalid, using random initialization.'Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119440,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-20T04:28:31.217000",
          "content": "<p>you can download \"resnet200d_ra2-bdba9bf9.pth\" manually and use torch to load it:</p>\n<pre><code>    net = ResNet(\n        Bottleneck,\n        layers=[3, 24, 36, 3],\n        num_classes=1000,\n        in_chans=3,\n        cardinality=1,\n        base_width=64,\n        stem_width=32,\n        stem_type='deep',\n        output_stride=32,\n        block_reduce_first=1,\n        down_kernel_size=1,\n        avg_down=True,\n        act_layer=nn.ReLU,\n        norm_layer=nn.BatchNorm2d,\n        aa_layer=None,\n        drop_rate=0.0,\n        drop_path_rate=0.,\n        drop_block_rate=0.,\n        global_pool='avg',\n        zero_init_last_bn=True,\n        block_args=None\n    )\n    print(net)\n    pretrain_state_dict = torch.load('resnet200d_ra2-bdba9bf9.pth', map_location=lambda storage, loc: storage)\n\n    s = net.load_state_dict(pretrain_state_dict, strict=True)\n    print(s)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1119444,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2020-12-20T04:35:47.743000",
          "content": "<p>Thanks!By the way,what batch size do you use?Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119455,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-20T04:51:28.933000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1215649,
      "author_name": "Salim Khazem",
      "author_url": "",
      "post_date": "2021-02-23T20:32:39.870000",
      "content": "<p>Great Work , thanks a lot for sharing , </p>\n<p>How much time it takes for training with EfficientNet B7 ? </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1214584,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-02-23T01:35:27.873000",
      "content": "<p>Why is the valid AUC basically the same for all, but the LB increases?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1149264,
      "author_name": "Raivo Koot",
      "author_url": "",
      "post_date": "2021-01-11T18:17:28.413000",
      "content": "<p>When you train these models from pretrained weights, do you first only train the new fully connected top layer for a few epochs and then train everything together afterwards, or do you just train everything right from the start?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1126176,
      "author_name": "Harveen Singh Chadha",
      "author_url": "",
      "post_date": "2020-12-25T12:08:51.270000",
      "content": "<p>Isn't vertical flip wrong given the fact that x-rays will come only in upright position?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1128499,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-27T13:39:50.097000",
          "content": "<p>I've debated that for myself and did not use it for that reason in <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">the notebook I made public</a> and in what I otherwise do (and also limited the extent of rotations). I guess a human radiologist could kind of look at the picture the wrong way around and cope (but most likely would of course flip it around…)? In the end, if it helps the model (as indicated by CV), I would still use it, I just put it very low down on my priority list of things to try.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128543,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-27T14:32:07.400000",
          "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> what you mentioned it's not trivial. In theory ,we do augmentations after see what images the model missclasified and choose the augmentations that benefit that images. By example, if the model missclasified  a image with zoom, we use a zoom augmentation. This time i don't see any pattern in the missclassified images, however it's still a good idea to use augmentations as a <strong>regularization</strong> technique. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128551,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-27T14:42:08.303000",
          "content": "<p>you just need \"more data\". here are some insights</p>\n<ol>\n<li><p>last time in the early days, people create \"virtual support sample\" to change the boundary of SVM. it is found that these \"virtual sample\" does not resemble the actual samples.</p></li>\n<li><p>in some papers, people use extreme augmentation to create more data and treat them as unlabelled. (the augmentation is too extreme to destroy the original label). then weak/self supervised learning method is applied.</p></li>\n<li><p>it is found that by adding noise, deep learning actually improved and become more robust</p></li>\n<li><p>while a sample that is adversarial modified looks very much like the original sample, the network can be easily fooled (aka adversarial attack). the way human looks at the sample is not the same as the machine would</p></li>\n</ol>\n<p>rubbish in, rubbish out …<br>\ngold in, rubbish out …<br>\nrubbish in, gold out … !!!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1135906,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2021-01-02T15:34:54.803000",
          "content": "<p><a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> The effect of specific augmentation methods can be easily tested. Here is the recipe:</p>\n<ol>\n<li>Train the model without the vertical flip, record the validation auc.</li>\n<li>Train the model with the vertical flip, record the validation auc.</li>\n</ol>\n<p>If the validation auc in 1 is greater than the validation auc in 2, vertical flip is useless.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1136099,
          "author_name": "Harveen Singh Chadha",
          "author_url": "",
          "post_date": "2021-01-02T18:14:35.857000",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a> Yup that is right unless you don't have problem of unlimited compute power :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117025,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-17T16:50:21.650000",
      "content": "<p>thanks for the results! <br>\ni wonder if the improvement is due to the image size?<br>\ne.g. i have efficientnetb2 on 512x512 and get valid AUC of 0.936+</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1117028,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-17T16:53:17.547000",
          "content": "<p>I'm suspecting that my model has been underfitting, since the training/validation losses are very close to each other. Will try to run it for longer (15 epochs) and see if it improves.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1117346,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2020-12-18T00:54:00.920000",
          "content": "<p>Hi,I use the efficientnetb4 on 512x512,and 5 fold.But I get local cv 0.9078 and lb 0.935.Do you know this is why?why you use efficientnetb2 can get 0.936?Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118536,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-19T06:54:18.557000",
          "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> </p>\n<p>i confirm the effectiveness of using deeper network like B5,B7. I publish some results above. Here are some quick note:</p>\n<ul>\n<li>size alone does not explain the improvement. e.g. 600-efficientnetb2 is slight better than 512-efficientnetb2. But 600-efficientnetb5 can make a huge difference.<br>\n(I think it is the larger receptive field and this can be verified by CAM map. I haven't investigated the shortcut-drop-rate parameter, which can also be a reason. efficientnetb5 has higher shortcut drop rate)</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1147758,
      "author_name": "Digvijay Yadav",
      "author_url": "",
      "post_date": "2021-01-10T17:15:24.283000",
      "content": "<p>Thank you so much for your works, it has really motivated me to explore this competition more now</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1118246,
      "author_name": "Sina Mehdinia",
      "author_url": "",
      "post_date": "2020-12-18T21:56:50.957000",
      "content": "<p>With my efficientB0 model, I got training AUC of 0.91, validation of 0.90, and LB of 0.90. So, my understanding is I have an underfitting problem. What do you suggest (other than increasing the image size) on GPU? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1121268,
          "author_name": "Akhilesh D. Kapse",
          "author_url": "",
          "post_date": "2020-12-21T14:02:41.243000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/sinamhd9\" target=\"_blank\">@sinamhd9</a>, at beginning I refer your notebook for my first submission, but with modified parameters which really work well. i.e</p>\n<ol>\n<li>Using <strong>EfficientB4</strong> with lower batchsize(10~14)</li>\n<li>From my observation, <strong>vertical flipping not contributes significantly in model generalization</strong>.</li>\n<li>Using K-fold methods is better.</li>\n</ol>\n<p>Hope it helps <br>\nThanks <br>\n~ AKhilesh</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1118241,
      "author_name": "Sina Mehdinia",
      "author_url": "",
      "post_date": "2020-12-18T21:43:13.600000",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> Thanks for sharing! I am using more augmentation methods such as (brightness change, shear, rotation, etc..). <a href=\"https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1\" target=\"_blank\">https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1</a></p>\n<p>Since this is a medical application I am not sure if flipping horizontally or vertically will be problematic or not. What do you think?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1118257,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-18T22:13:25.560000",
          "content": "<p>From a theoretical perspective I don't think it poses a problem per se, but a rotation might be more effective since it would more closely represent what would happen in real life (scans in different positions).</p>\n<p>Moreover I'm not an expert in medical imaging, I think Dr. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> would likely have a better idea about this than me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1118259,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-18T22:15:17.440000",
          "content": "<p>I'm also not sure about the brightness change and shear, since those don't seem to be common attributes of CT scans, but again I might be wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118269,
          "author_name": "Sina Mehdinia",
          "author_url": "",
          "post_date": "2020-12-18T22:29:25.067000",
          "content": "<p>Thanks for your reply! Me neither, I was curious to see if they have an effect on better generalization or not.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1118286,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-18T22:57:06.737000",
          "content": "<p>Clinical implications aside, I think they are worth trying in case they do result in better generalization.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118335,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-12-19T00:50:18.040000",
          "content": "<p>Technically, horizontal/vertical flipping would result in images that we would not see regularly in clinical practice. </p>\n<p>However, I use these augmentations during training anyways to reduce overfitting. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1118338,
          "author_name": "Sina Mehdinia",
          "author_url": "",
          "post_date": "2020-12-19T01:07:30.180000",
          "content": "<p>That totally makes sense! Thank you for sharing!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1117360,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2020-12-18T01:19:26.103000",
      "content": "<p>Interesting. Is there a reason you skipped B4 and B5 (other than you can't do it all?)? I guess, you might also argue that B0 to B2 might be good to experimenting quickly, while B6/B7 are final model material, while the ones inbetween are neither here nor there? Was that the thinking?</p>\n<p>Mostly as a compromise between time and performance, my first try was this stuff that I shared with an EfficientNet-B4 (vanilla B4, not noisy student) that gets close to the B6 results you posted (LB 0.945) (<a href=\"https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4\" target=\"_blank\">inference</a> and <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">training</a> notebooks). I'd guess that's because you did not use much augmentation &amp; a simple LR schedule (reduce LR on plateau vs. some augmentation &amp; cosine-annealing in my case)? </p>\n<p>And, wow, B7 trained fast on TPU!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1117364,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-18T01:29:59.870000",
          "content": "<p>I haven't had the bandwidth to try everything, but I will come to those eventually!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1242003,
      "author_name": "Jaroslav Tuma",
      "author_url": "",
      "post_date": "2021-03-17T10:39:47.827000",
      "content": "<p>Thank you for your work. Unfortunatelly I was struggling to achieve decent performance with a copy of your notebook (auc below 0.7). Anyone else had similar problem?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1146190,
      "author_name": "Samuel Liebana",
      "author_url": "",
      "post_date": "2021-01-09T15:35:43.917000",
      "content": "<p><a href=\"https://www.kaggle.com/xhulu\" target=\"_blank\">@xhulu</a> thank you very much for the awesome notebooks - you're the best!!!! I have tried running your notebooks and I have always ended up (after 20 epochs) with around 0.85 val_auc, whereas your table is showing 0.95 and above. Are the results you have presented from the first run of the models that you list or are you using another method that we are not aware of? Thank you in advance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1147660,
          "author_name": "CBR",
          "author_url": "",
          "post_date": "2021-01-10T16:25:18.480000",
          "content": "<p><a href=\"https://www.kaggle.com/samuelliebana\" target=\"_blank\">@samuelliebana</a> , ran <a href=\"https://www.kaggle.com/xhulu\" target=\"_blank\">@xhulu</a>  's notebook successfully. The lower value may attribute to mis-matching image_size you might be passing to <code>efficientNetB0</code>. Can you check? In addition, check the prediction for test image, did you forget to update the image_size? </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1144682,
      "author_name": "CBR",
      "author_url": "",
      "post_date": "2021-01-08T15:54:19.627000",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>, Thank you sharing simply GOLD notebooks! Get to learn a lot. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1135315,
      "author_name": "qingning",
      "author_url": "",
      "post_date": "2021-01-02T05:58:36.050000",
      "content": "<p>Great jobs</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1128467,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-12-27T13:12:02.330000",
      "content": "<p>It takes an hour and a half to train one <strong>B7</strong> model, on TPU, how do you plan to do cross-validation if TPU stops after 3 hours?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1128501,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-27T13:41:31.013000",
          "content": "<p>Not sure what the problem is - you have 30 hours per week and there's several weeks to go. As long as inference is fast enough, training time would not seem to be a problem (if you train in separate notebooks and use the trained models for inference)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128534,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-27T14:25:25.080000",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> my bad, the comment has been corrected.</p>\n<p>We have 30 hours per week, but the TPU stops after 3 hours, if a model takes 1 hour and a half to train, and we want to do, by example 3 fold cross-validation, the 3 hours of TPU wouldn't be enough.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128556,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-27T14:44:28.627000",
          "content": "<p>Just run multiple notebooks, one per fold. It's not like you could not have multiple notebooks as inputs to a submission kernel. Sure, it's not very elegant, but totally possible.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128558,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-27T14:45:53.967000",
          "content": "<p>Whoa, that's something new, i'm going to try it, thanks for the idea.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128969,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-12-28T00:11:26.783000",
          "content": "<p>I don't recommend playing with larger models and larger resolution at the beginning of a competition.😄</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128977,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-28T00:35:42.967000",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a>, recently i did read the book \"Deep Learning for Coders with Fastai and PyTorch: AI Applications Without a PhD\", where the author wrote that we must try first complex models and finetune if it doesn't work well, after that we can use less complex models. Now that's part of my way to approach any problem with DL.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129085,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-12-28T04:28:04.843000",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> you mean you won your first medal in MOA? That competition is not related to computer vision at all. The time it takes to run an experiment in MOA is nearly negligible. Maybe you wanna watch this video: <a href=\"https://www.youtube.com/watch?v=L1QKTPb6V_I\" target=\"_blank\">https://www.youtube.com/watch?v=L1QKTPb6V_I</a>.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1129561,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-28T12:40:41.373000",
          "content": "<p>I understand your concern but if I have desired results I don't want to Change my approach. If at the end of this competition doesn't works, i've learned that you was right. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129858,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "2020-12-28T15:43:02.087000",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Thanks For the link I really enjoyed watching it )) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1162375,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2021-01-21T05:46:03.240000",
          "content": "<p><a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> - in case you or others have not seen this, TPU has same time as GPU now.</p>\n<p><a href=\"https://www.kaggle.com/product-feedback/202409\" target=\"_blank\">https://www.kaggle.com/product-feedback/202409</a><br>\nTPU execution time increased to 9 hours!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1118906,
      "author_name": "Stanislau Korkuts",
      "author_url": "",
      "post_date": "2020-12-19T14:13:07.357000",
      "content": "<p>Hello, what GPU did you use?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1119016,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-19T16:23:45.743000",
          "content": "<p>I think kaggle uses p100</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1118393,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-12-19T03:07:52.740000",
      "content": "<p><code>0.9713</code> with new metrics?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1118201,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-12-18T20:33:58.470000",
      "content": "<p>LB score compares between GPU and TPU is really noticeable. I like to see if someone adds some PyTorch observation GPU. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1118206,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-18T20:40:12.800000",
          "content": "<p>I don't think the accelerator makes a difference per se, but the model size (B3 vs B7) and the batch size is an important reason. I'm using a batch size of 16 for GPU and 16*8=128 for the TPU. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1118221,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-12-18T21:02:59.533000",
          "content": "<p>Of course, I didn't mean accelerator but the advantages of using TPU. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1118563,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-19T07:32:38.187000",
          "content": "<p>Note also that some of the pre-trained models are <a href=\"https://arxiv.org/abs/1911.04252\" target=\"_blank\">noisy student </a> ones (do better on ImageNet, which you might suspect could improve performance here, too) and some are not. That makes the picture a little less clear.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1119021,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-19T16:27:09.277000",
          "content": "<p>I'm super curious about whether noisy students makes a difference on downstream task, or they are over fitting on imagenet. I wish there's a uniform image recognition benchmark for this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1119146,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2020-12-19T18:42:29.147000",
          "content": "<p>Noisy student didn't make the cut-off for <a href=\"https://arxiv.org/abs/1902.10811#:~:text=Our%20results%20suggest%20that%20the,in%20the%20original%20test%20sets.\" target=\"_blank\">Do ImageNet Classifiers Generalize to ImageNet</a>, but their dataset would be one logical thing to try. I'd assume the noisy student models would do well there (but that's the exact same task and mostly assesses overfitting to the test set.).</p>\n<p>For down-stream tasks, I do wonder whether Kaggle competitions would not be one of the more obvious choices, or perhaps Open Images or the COCO Dataset?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117115,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-12-17T18:30:48.803000",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> <br>\nThanks for sharing. :)</p>\n<p>I've gone through the GPU implementation of yours. I saw you didn't do Group-KFold, but just plain train-test-split. Is there any particular reason or you've planned to add it later? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1117284,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-17T21:50:08.197000",
          "content": "<p>It usually takes a lot of time to train a GPU model, so adding K-Fold will increase the time by K with only minimal gains (+/- 1%). As a consequence I quickly run out of GPU quota, which I could've used to experiment with crucial hyperparameters (e.g. learning rate, optimizer, model size, image size).</p>\n<p>I feel that publishing very simple notebooks like these ones give a chance to others to try out their favourite techniques without adding too much computation overhead and cognitive loads, and also share their solutions as notebooks (although it will eventually surpass the baseline score). Finally single models are more useful clinically since they are faster to run and can be interpreted using traditional methods (e.g. GradCAM).</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1117343,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-12-18T00:52:15.823000",
          "content": "<p>I agree. I was just wondering why you didn't use that 😄</p>\n<p>However, with group k-fold training, setting with <code>multi_label=True</code> gave me pretty constant val AUC whereas without it (False), the val AUC was pretty good (~0.9). I'm not quite sure, maybe computing with column-wise AUC was giving constant output due to the imbalance data point or something else. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117037,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-12-17T17:06:27.073000",
      "content": "<p>Thanks for sharing!<br>\np.s. It looks like evaluation metric has been updated.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1117060,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-12-17T17:30:13.530000",
          "content": "<p>It seems like the top 10 have not been updated yet based on <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> observations</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1214573,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-02-23T01:25:21.047000",
      "content": "<p>Great experiments. Thanks for sharing!</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1117019": "I've published a few notebooks, so I'd like to aggregate them in this post and give more details. This is a work in progress.\n\n## Scores by model\n\nAll the models I trained were EfficientNets using TensorFlow Keras. Depending on the size they were either trained on GPU or TPUs. **The scores were all calculated after the rescoring started so they should be up-to-date**.\n\nYou can find the respective notebooks here:\n1. [Full workflow on GPU (train and submit)](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit)\n2. [Train on TPU](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training)\n3. [Submit on GPU](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-submission)\n\nNote that [1] is for the smaller models (B2 and B3), and [2-3] are for the larger models (B6-B7). Although [1] is simpler, it take a lot of time to submit so I recommend taking a look at [3] if you want to have faster submissions.\n\nHere's a summary table of the scores:\n\n\n| Model | Img Size | Valid AUC | LB    | Accelerator | Weights       | Version |\n|-------|----------|-----------|-------|-------------|---------------|---------|\n| B0    | 224      | 0.9454    | 0.883 | GPU         | Noisy Student | 15      |\n| B1    | 240      | 0.9618    | 0.888 | GPU         | Noisy Student | 14      |\n| B2    | 260      | 0.9278    | 0.918 | GPU         | Noisy Student | 10      |\n| B2    | 260      | 0.9587    | 0.912 | GPU         | Noisy Student | 11      |\n| B2    | 260      | 0.9713    | 0.908 | GPU         | Noisy Student | 12      |\n| B3    | 300      | 0.9202    | 0.909 | GPU         | ImageNet      | 9       |\n| B5    | 456      | 0.9382    | 0.944 | TPU         | ImageNet      | 9       |\n| B6    | 528      | 0.9415    | 0.949 | TPU         | Noisy Student | 6       |\n| B7    | 600      | 0.9455    | 0.953 | TPU         | Noisy Student | 7       |\n| B7    | 600      | 0.9431    | 0.957 | TPU         | ImageNet      | 8       |\n\n## Hyperparameters and training details\n\nHere are some details about how I trained the models:\n\n* **Training augmentation**: Simply random left-right and top-bottom flipping\n* **Optimizer**: Adam with an initial learning rate of 0.001 and no other tuning\n* **Metrics**: Multi-label AUROC (as opposed to flattened), corresponding to the competition metric\n* **Scheduling**: Reduce learning rate by 10x every 3 epochs where the valid AUC did not improve\n* **Number of epochs**: 10-15 for early version of GPU, 20 for later version and for TPU\n* **Batch Size**: 16 for GPU and 16*8=128 for the TPU\n* **Model saving** Save model with the best validation AUC after every epoch\n\nI did not try any of the following:\n* **Cross-validation**\n* **TTA**\n* **Ensembling/Stacking**\n\n## Other notes\n\n* In more recent versions, I've changed the implementation from `qubvel/callidor` to the one in `tensorflow.keras.applications`. Since it is not officially released for v2.2.0, I made [a utility script](https://www.kaggle.com/xhlulu/tf-keras-efficientnet) for the TPU notebook. For the other notebooks, I'm loading directly from `tensorflow` since they are on v2.3.1",
    "1119390": "latest @rwightman pytorch mode: ResNet-200D (pretain at imagenet 320x320)\nhttps://github.com/rwightman/pytorch-image-models\nhttps://twitter.com/wightmanr/status/1340105026786656256?s=20\n\nquote \"Compared to an ImageNet-1k only EfficientNet-B5 w/ RandAugment, the 200D is 30% faster on a GPU and better top-1.\"\n\nsingle fold, 2xTTA 640x640 resnet200: LB 0.959\n\nlocal cv (2xTTA)\n```\nprobability : (6017, 11)\nlabel : (6017, 11)\nsubmit_dir : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/valid/local-00009000_model\ninitial_checkpoint : /root/share1/kaggle/2020/ranzcr/result/resnet200/640-aug-2/fold1/checkpoint/00009000_model.pth\n\nfold1:  resnet200 640x640 : LB 0.959\nloss : 0.145014\nauc      : [0.9869, 0.9564, 0.9888, 0.9621, 0.9571, 0.9837, 0.9866, 0.9186, 0.8671, 0.9086, 0.9996]\nauc_flat : 0.978218\nauc_mean : 0.955946\n```\n\n---\n\ncomparison to efficient net\n\n```\nfold2:  effb2 600x600 : LB 0.951\nloss : 0.143558\nauc      : [0.9564, 0.9547, 0.9899, 0.9479, 0.9264, 0.9757, 0.9817, 0.8984, 0.8354, 0.9020, 0.9996]\nauc_flat : 0.976190\nauc_mean : 0.942560\n\n-----\n\nfold1:  effb5 600x600 : LB 0.944\nloss : 0.153872\nauc      : [0.9907, 0.9513, 0.9890, 0.9624, 0.9551, 0.9796, 0.9847, 0.9298, 0.8604, 0.9086, 0.9989]\nauc_flat : 0.977150\nauc_mean : 0.955487\n\n-----\n\nfold1:  effb7 600x600 : LB 0.954\nloss : 0.165448\nauc      : [0.9892, 0.9479, 0.9870, 0.9543, 0.9447, 0.9807, 0.9844, 0.9139, 0.8565, 0.9044, 0.9995]\nauc_flat : 0.975849\nauc_mean : 0.951142\n\n```",
    "1215649": "Great Work , thanks a lot for sharing , \n\nHow much time it takes for training with EfficientNet B7 ? ",
    "1214584": "Why is the valid AUC basically the same for all, but the LB increases?",
    "1149264": "When you train these models from pretrained weights, do you first only train the new fully connected top layer for a few epochs and then train everything together afterwards, or do you just train everything right from the start?",
    "1126176": "Isn't vertical flip wrong given the fact that x-rays will come only in upright position?",
    "1117025": "thanks for the results! \ni wonder if the improvement is due to the image size?\ne.g. i have efficientnetb2 on 512x512 and get valid AUC of 0.936+",
    "1147758": "Thank you so much for your works, it has really motivated me to explore this competition more now\n",
    "1118246": "With my efficientB0 model, I got training AUC of 0.91, validation of 0.90, and LB of 0.90. So, my understanding is I have an underfitting problem. What do you suggest (other than increasing the image size) on GPU? ",
    "1118241": "@xhlulu Thanks for sharing! I am using more augmentation methods such as (brightness change, shear, rotation, etc..). https://www.kaggle.com/sinamhd9/keras-models-image-data-generator-part1\n\nSince this is a medical application I am not sure if flipping horizontally or vertically will be problematic or not. What do you think?",
    "1117360": "Interesting. Is there a reason you skipped B4 and B5 (other than you can't do it all?)? I guess, you might also argue that B0 to B2 might be good to experimenting quickly, while B6/B7 are final model material, while the ones inbetween are neither here nor there? Was that the thinking?\n\nMostly as a compromise between time and performance, my first try was this stuff that I shared with an EfficientNet-B4 (vanilla B4, not noisy student) that gets close to the B6 results you posted (LB 0.945) ([inference](https://www.kaggle.com/bjoernholzhauer/inference-for-trained-fastai-efficientnet-b4) and [training](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) notebooks). I'd guess that's because you did not use much augmentation & a simple LR schedule (reduce LR on plateau vs. some augmentation & cosine-annealing in my case)? \n\nAnd, wow, B7 trained fast on TPU!",
    "1242003": "Thank you for your work. Unfortunatelly I was struggling to achieve decent performance with a copy of your notebook (auc below 0.7). Anyone else had similar problem?",
    "1146190": "@xhulu thank you very much for the awesome notebooks - you're the best!!!! I have tried running your notebooks and I have always ended up (after 20 epochs) with around 0.85 val_auc, whereas your table is showing 0.95 and above. Are the results you have presented from the first run of the models that you list or are you using another method that we are not aware of? Thank you in advance.",
    "1144682": "@xhlulu, Thank you sharing simply GOLD notebooks! Get to learn a lot. ",
    "1135315": "Great jobs",
    "1128467": "It takes an hour and a half to train one **B7** model, on TPU, how do you plan to do cross-validation if TPU stops after 3 hours?",
    "1118906": "Hello, what GPU did you use?",
    "1118393": "`0.9713` with new metrics?",
    "1118201": "LB score compares between GPU and TPU is really noticeable. I like to see if someone adds some PyTorch observation GPU. ",
    "1117115": "@xhlulu \nThanks for sharing. :)\n\nI've gone through the GPU implementation of yours. I saw you didn't do Group-KFold, but just plain train-test-split. Is there any particular reason or you've planned to add it later? \n\n\n",
    "1117037": "Thanks for sharing!\np.s. It looks like evaluation metric has been updated.",
    "1214573": "Great experiments. Thanks for sharing!"
  }
}