{
  "id": 155668,
  "title": "Reaching 0.9 with ResNet18 and 256x256 images",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155668",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2020-06-02T14:49:12.194000",
  "votes": 39,
  "comment_count": 46,
  "views": 0,
  "content": "<p>My current setup:</p>\n\n<ul>\n<li>ResNet18</li>\n<li>Images resized to 256x256 </li>\n<li>Adam</li>\n<li>0.0001 LR for ResNet and 0.02 for randomly initialized classification layer</li>\n<li>4 rotations + vertical flips (8 options total) on training</li>\n<li>10 epochs</li>\n<li>batch size is 64</li>\n<li>3 folds (group + stratify)</li>\n<li>Average of the last 4 epochs -&gt; gives slightly more stable results than just last/best epoch</li>\n<li>Training time is less than an hour on a 1080ti</li>\n</ul>\n\n<p>0.899 CV, 0.903 LB</p>",
  "messages": [
    {
      "id": 871656,
      "postDate": "2020-06-02T14:49:12.193Z",
      "content": "<p>My current setup:</p>\n\n<ul>\n<li>ResNet18</li>\n<li>Images resized to 256x256 </li>\n<li>Adam</li>\n<li>0.0001 LR for ResNet and 0.02 for randomly initialized classification layer</li>\n<li>4 rotations + vertical flips (8 options total) on training</li>\n<li>10 epochs</li>\n<li>batch size is 64</li>\n<li>3 folds (group + stratify)</li>\n<li>Average of the last 4 epochs -&gt; gives slightly more stable results than just last/best epoch</li>\n<li>Training time is less than an hour on a 1080ti</li>\n</ul>\n\n<p>0.899 CV, 0.903 LB</p>",
      "rawMarkdown": "My current setup:\n\n* ResNet18\n* Images resized to 256x256 \n* Adam\n* 0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\n* 4 rotations + vertical flips (8 options total) on training\n* 10 epochs\n* batch size is 64\n* 3 folds (group + stratify)\n* Average of the last 4 epochs -&gt; gives slightly more stable results than just last/best epoch\n* Training time is less than an hour on a 1080ti\n\n0.899 CV, 0.903 LB",
      "votes": 39
    },
    {
      "id": 871707,
      "postDate": "2020-06-02T15:29:11.737Z",
      "content": "<p>Great job. If people want to recreate this in Kaggle notebooks using GPU or TPU, I uploaded TFRecords containing 256x256 <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>. Then after you write you code, you can switch the 256x256 dataset with my 512x512 or 768x768 for a higher CV/LB. Enjoy!</p>\n\n<p>Rotation augmentation for TFRecords is explained <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a>. And CutMix and MixUp <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">here</a></p>",
      "rawMarkdown": "Great job. If people want to recreate this in Kaggle notebooks using GPU or TPU, I uploaded TFRecords containing 256x256 [here][1]. Then after you write you code, you can switch the 256x256 dataset with my 512x512 or 768x768 for a higher CV/LB. Enjoy!\n\nRotation augmentation for TFRecords is explained [here][2]. And CutMix and MixUp [here][3]\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[2]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\n[3]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935",
      "votes": 7
    },
    {
      "id": 871745,
      "postDate": "2020-06-02T15:55:19.697Z",
      "content": "<p>Interestingly, so far switching to 512x512 resolutions makes the result <strong>worse</strong>. I suspect that's because ResNet18 has quite a small receptive field and couldn't fully utilize higher resolution.</p>",
      "rawMarkdown": "Interestingly, so far switching to 512x512 resolutions makes the result **worse**. I suspect that's because ResNet18 has quite a small receptive field and couldn't fully utilize higher resolution.",
      "votes": 3,
      "replies": [
        {
          "id": 871749,
          "postDate": "2020-06-02T15:58:58.120Z",
          "content": "<p>yes, that is interesting</p>",
          "rawMarkdown": "yes, that is interesting"
        },
        {
          "id": 876217,
          "postDate": "2020-06-06T14:41:14.853Z",
          "content": "<p>Better try that resolution on which it was pretrained.This will boost the model's efficiency...</p>",
          "rawMarkdown": "Better try that resolution on which it was pretrained.This will boost the model's efficiency...",
          "votes": 1
        }
      ]
    },
    {
      "id": 884438,
      "postDate": "2020-06-13T11:28:55.857Z",
      "content": "<p>&gt; 4 rotations + vertical flips (8 options total)</p>\n\n<p>Is this a good way of implementing it?\n<code>\nimport albumentations as A\naugs = A.Compose([\n    A.OneOf([\n        A.Rotate(limit=(0, 0), p=1),\n        A.Rotate(limit=(90, 90), p=1),\n        A.Rotate(limit=(180, 180), p=1),\n        A.Rotate(limit=(270, 270), p=1)]),\n    A.VerticalFlip(p=0.5)\n])\n</code>\nOr is there a more concise code?</p>",
      "rawMarkdown": "&gt; 4 rotations + vertical flips (8 options total)\n\nIs this a good way of implementing it?\n```\nimport albumentations as A\naugs = A.Compose([\n    A.OneOf([\n        A.Rotate(limit=(0, 0), p=1),\n        A.Rotate(limit=(90, 90), p=1),\n        A.Rotate(limit=(180, 180), p=1),\n        A.Rotate(limit=(270, 270), p=1)]),\n    A.VerticalFlip(p=0.5)\n])\n```\nOr is there a more concise code?",
      "votes": -1,
      "replies": [
        {
          "id": 884611,
          "postDate": "2020-06-13T13:44:41.153Z",
          "content": "<p>I guess I used <code>RandomRotate90</code> :)</p>",
          "rawMarkdown": "I guess I used `RandomRotate90` :)",
          "votes": 1
        },
        {
          "id": 884640,
          "postDate": "2020-06-13T14:11:33.870Z",
          "content": "<p>Well, it seems this is much more simple :)</p>",
          "rawMarkdown": "Well, it seems this is much more simple :)"
        }
      ]
    },
    {
      "id": 876010,
      "postDate": "2020-06-06T11:08:26.473Z",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a> have you tried more epochs?</p>",
      "rawMarkdown": "@ddanevskyi have you tried more epochs?",
      "votes": -1,
      "replies": [
        {
          "id": 876028,
          "postDate": "2020-06-06T11:18:35.860Z",
          "content": "<p>Yes. For this setup, there was no difference.</p>",
          "rawMarkdown": "Yes. For this setup, there was no difference."
        },
        {
          "id": 876144,
          "postDate": "2020-06-06T13:27:52.363Z",
          "content": "<p>are these results are on the provided dataset only or you used external data as well?</p>",
          "rawMarkdown": "are these results are on the provided dataset only or you used external data as well?",
          "votes": -1
        },
        {
          "id": 876171,
          "postDate": "2020-06-06T13:53:26.873Z",
          "content": "<p>I'm using only the competition data.</p>",
          "rawMarkdown": "I'm using only the competition data."
        }
      ]
    },
    {
      "id": 871790,
      "postDate": "2020-06-02T16:28:18.627Z",
      "content": "<p>Really interesting results. Could you share a notebook for the same? Also , what exactly do you mean by average of the last 4 epochs? Also,  do you think it is possible to run notebooks for this comp on a 6GB 1660Ti?</p>",
      "rawMarkdown": "Really interesting results. Could you share a notebook for the same? Also , what exactly do you mean by average of the last 4 epochs? Also,  do you think it is possible to run notebooks for this comp on a 6GB 1660Ti?",
      "votes": -2,
      "replies": [
        {
          "id": 871860,
          "postDate": "2020-06-02T17:19:24.323Z",
          "content": "<p>I just averaged predictions on the test set from the last 4 epochs (I predict both validation and test on each epoch). </p>\n\n<p>As for hardware, 1660Ti might be fine for shallow models (such as ResNet18), but deeper models (which are a must for a high score) are better be trained on TPUs.</p>",
          "rawMarkdown": "I just averaged predictions on the test set from the last 4 epochs (I predict both validation and test on each epoch). \n\nAs for hardware, 1660Ti might be fine for shallow models (such as ResNet18), but deeper models (which are a must for a high score) are better be trained on TPUs.",
          "votes": 1
        },
        {
          "id": 871896,
          "postDate": "2020-06-02T18:00:37.773Z",
          "content": "<p>Oh , I see. Thanks for the reply. Guess I'll just have to work with training with checkpoints on Kaggle TPU's :(</p>",
          "rawMarkdown": "Oh , I see. Thanks for the reply. Guess I'll just have to work with training with checkpoints on Kaggle TPU's :("
        }
      ]
    },
    {
      "id": 888415,
      "postDate": "2020-06-16T10:34:50.153Z",
      "content": "<p>hi！I tried to train the dataset with resnet50, the number of pictures with true label is 584, the number of false is 32000+, that is imbalanced, I trained resnet50 using your strategy, when I predict, the output label for the test images are all 0, can you tell me why?  Thank you!  </p>",
      "rawMarkdown": "hi！I tried to train the dataset with resnet50, the number of pictures with true label is 584, the number of false is 32000+, that is imbalanced, I trained resnet50 using your strategy, when I predict, the output label for the test images are all 0, can you tell me why?  Thank you!  ",
      "replies": [
        {
          "id": 888437,
          "postDate": "2020-06-16T11:01:40.483Z",
          "content": "<p>I can't. What is your AUC score on the training set?</p>",
          "rawMarkdown": "I can't. What is your AUC score on the training set?"
        },
        {
          "id": 888458,
          "postDate": "2020-06-16T11:21:55.710Z",
          "content": "<p>I am changing the code to add this function, is there a threshold to tell the bad from the good, I set it 0.5 before, need I change it to a bigger one ?</p>",
          "rawMarkdown": "I am changing the code to add this function, is there a threshold to tell the bad from the good, I set it 0.5 before, need I change it to a bigger one ?"
        },
        {
          "id": 888472,
          "postDate": "2020-06-16T11:34:12.873Z",
          "content": "<p>You can keep it 0 or <code>None</code> too, just share the final AUC score on the training set here! :)</p>",
          "rawMarkdown": "You can keep it 0 or `None` too, just share the final AUC score on the training set here! :)"
        }
      ]
    },
    {
      "id": 882245,
      "postDate": "2020-06-11T16:55:47.547Z",
      "content": "<p>Hey, I have a issue with my setup, I am using,</p>\n\n<ul>\n<li>Efficientnet-b0</li>\n<li>224x224</li>\n<li>Adam</li>\n<li>0.0001 LR </li>\n<li>Random Horizontal and Vertical flips for training data</li>\n<li>10 epochs</li>\n<li>batch_size = 64</li>\n<li>5 folds (Group)</li>\n</ul>\n\n<p>I am getting a CV of about 0.72 and that too is fluctuating a lot, it sometimes drops to even 0.3. :/ </p>\n\n<p>I am still a newbie, any suggestions on how to improve this would be highly appreciated. </p>",
      "rawMarkdown": "Hey, I have a issue with my setup, I am using,\n\n- Efficientnet-b0\n- 224x224\n- Adam\n- 0.0001 LR \n- Random Horizontal and Vertical flips for training data\n- 10 epochs\n- batch_size = 64\n- 5 folds (Group)\n\nI am getting a CV of about 0.72 and that too is fluctuating a lot, it sometimes drops to even 0.3. :/ \n\nI am still a newbie, any suggestions on how to improve this would be highly appreciated. ",
      "replies": [
        {
          "id": 882271,
          "postDate": "2020-06-11T17:15:27.297Z",
          "content": "<p>Sounds like a bug. Do you normalize images? Are you sure your training/validation/testing pipelines are the same?</p>",
          "rawMarkdown": "Sounds like a bug. Do you normalize images? Are you sure your training/validation/testing pipelines are the same?"
        },
        {
          "id": 882278,
          "postDate": "2020-06-11T17:19:32.947Z",
          "content": "<p>Yes, I am normalizing them. Yeah, pipeline is same too.</p>\n\n<p>I am thinking of running it for more epochs now, maybe 20, to see if it starts to converge then.</p>",
          "rawMarkdown": "Yes, I am normalizing them. Yeah, pipeline is same too.\n\nI am thinking of running it for more epochs now, maybe 20, to see if it starts to converge then."
        },
        {
          "id": 882357,
          "postDate": "2020-06-11T18:30:51.577Z",
          "content": "<p>It should converge in a couple of epochs, are you reducing your lr during training with a scheduler? </p>",
          "rawMarkdown": "It should converge in a couple of epochs, are you reducing your lr during training with a scheduler? "
        },
        {
          "id": 882627,
          "postDate": "2020-06-12T02:22:52.707Z",
          "content": "<p>Yes, I am using ReduceLROnPlateau as a Scheduler with a patience of 3.</p>\n\n<p>Also, I tried reducing the number of groups to 3 and auc is now going upto 0.8 now, originally it was only reaching 0.72 and it seems much stable than before.</p>\n\n<p>Although I know I can get at least 0.9 CV with this setup, it still is a long way to go.</p>",
          "rawMarkdown": "Yes, I am using ReduceLROnPlateau as a Scheduler with a patience of 3.\n\nAlso, I tried reducing the number of groups to 3 and auc is now going upto 0.8 now, originally it was only reaching 0.72 and it seems much stable than before.\n\nAlthough I know I can get at least 0.9 CV with this setup, it still is a long way to go."
        },
        {
          "id": 883925,
          "postDate": "2020-06-13T04:57:42.793Z",
          "content": "<p>Also I want to add one more thing, I am predicting on 2 classes with cross entropy loss, can it be causing this much fluctuation and less auc?</p>",
          "rawMarkdown": "Also I want to add one more thing, I am predicting on 2 classes with cross entropy loss, can it be causing this much fluctuation and less auc?"
        },
        {
          "id": 883957,
          "postDate": "2020-06-13T06:31:20.873Z",
          "content": "<p>I presume almost everyone is using cross-entropy.</p>",
          "rawMarkdown": "I presume almost everyone is using cross-entropy."
        },
        {
          "id": 883973,
          "postDate": "2020-06-13T06:34:48.163Z",
          "content": "<p>Okay, then there must be something else which is off, I will be looking for that then. Thanks! :)</p>",
          "rawMarkdown": "Okay, then there must be something else which is off, I will be looking for that then. Thanks! :)"
        },
        {
          "id": 884489,
          "postDate": "2020-06-13T12:13:17.547Z",
          "content": "<p>Can I share my notebook with anyone here? it may be a good deal for me to just get it quick checked with someone else, maybe I can't see what's going wrong and I have already dedicated 7 days to this particular problem. :/</p>",
          "rawMarkdown": "Can I share my notebook with anyone here? it may be a good deal for me to just get it quick checked with someone else, maybe I can't see what's going wrong and I have already dedicated 7 days to this particular problem. :/"
        },
        {
          "id": 884617,
          "postDate": "2020-06-13T13:46:48.153Z",
          "content": "<p>You could just make your notebook public. Sharing with someone privately without being in a team is prohibited.</p>",
          "rawMarkdown": "You could just make your notebook public. Sharing with someone privately without being in a team is prohibited."
        },
        {
          "id": 884624,
          "postDate": "2020-06-13T13:53:47.127Z",
          "content": "<p>Sure, I will do that, will comment here after making it public. Thanks! :)</p>",
          "rawMarkdown": "Sure, I will do that, will comment here after making it public. Thanks! :)"
        },
        {
          "id": 884729,
          "postDate": "2020-06-13T15:26:26.410Z",
          "content": "<p>I made the notebook <a href=\"https://www.kaggle.com/sarques/melanoma-classification\">public</a>.</p>\n\n<p>It performed a lot better than the last time as that time the score used to drop to 0.4 sometimes but I think it will still be good if you can take a look. </p>\n\n<p>I took some part of it from other notebooks too as I was just learning how to work with PyTorch. :)</p>",
          "rawMarkdown": "I made the notebook [public](https://www.kaggle.com/sarques/melanoma-classification).\n\nIt performed a lot better than the last time as that time the score used to drop to 0.4 sometimes but I think it will still be good if you can take a look. \n\nI took some part of it from other notebooks too as I was just learning how to work with PyTorch. :)"
        }
      ]
    },
    {
      "id": 873057,
      "postDate": "2020-06-03T18:25:52.160Z",
      "content": "<p>\"0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\"- How LR can be different for ResNet end classification layer ? How to write it in tf.keras ?</p>",
      "rawMarkdown": "\"0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\"- How LR can be different for ResNet end classification layer ? How to write it in tf.keras ?",
      "replies": [
        {
          "id": 873070,
          "postDate": "2020-06-03T18:43:50.940Z",
          "content": "<p>you would do do something like this where model.base is your features and fc is your classifier:</p>\n\n<p><code>\noptimizer = optim.AdamW([{'params': model.base.parameters(), 'lr ':1e-4 },\n                         {'params': model.fc.parameters(), 'lr ':2e-2 },\n                        ], lr=lr)\n</code></p>\n\n<p>This is usefull if you are using a pretrained architechture</p>",
          "rawMarkdown": "you would do do something like this where model.base is your features and fc is your classifier:\n\n```\noptimizer = optim.AdamW([{'params': model.base.parameters(), 'lr ':1e-4 },\n                         {'params': model.fc.parameters(), 'lr ':2e-2 },\n                        ], lr=lr)\n```\n\nThis is usefull if you are using a pretrained architechture",
          "votes": 6
        },
        {
          "id": 873081,
          "postDate": "2020-06-03T19:08:10.077Z",
          "content": "<p>how to initialize model.base and model.fc ? Is it example in tf.keras?</p>",
          "rawMarkdown": "how to initialize model.base and model.fc ? Is it example in tf.keras?"
        },
        {
          "id": 873094,
          "postDate": "2020-06-03T19:31:07.093Z",
          "content": "<p>Sorry i dont really know how to do it in keras</p>",
          "rawMarkdown": "Sorry i dont really know how to do it in keras",
          "votes": 1
        },
        {
          "id": 873176,
          "postDate": "2020-06-03T21:54:56.533Z",
          "content": "<p>hmm - did not hear this before - sounds plausible - can one say this gives more often better results than one lr for head and backbone? What is your experience?</p>",
          "rawMarkdown": "hmm - did not hear this before - sounds plausible - can one say this gives more often better results than one lr for head and backbone? What is your experience?"
        },
        {
          "id": 873206,
          "postDate": "2020-06-03T22:40:48.910Z",
          "content": "<p>For me it doesnt really change much, but the idea is since the features are pretrained they just need finetuning and the classifier parameters are set randomly so they need more trainning. However i think this idea is more valuable when training with images similar to imagenet, not sure though..</p>",
          "rawMarkdown": "For me it doesnt really change much, but the idea is since the features are pretrained they just need finetuning and the classifier parameters are set randomly so they need more trainning. However i think this idea is more valuable when training with images similar to imagenet, not sure though.."
        },
        {
          "id": 873462,
          "postDate": "2020-06-04T06:59:38.260Z",
          "content": "<p>You can do it with <code>tf.keras.backend.stop_gradient()</code></p>\n\n<pre><code>    inp = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    net = efn.EfficientNetB3(weights='imagenet', include_top=False)\n    x = net(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = 0.005*x + 0.995*tf.keras.backend.stop_gradient(x)\n    x = tf.keras.layers.Dense(512, activation='relu')(x)\n    x = tf.keras.layers.Dense(128, activation='relu')(x)\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n</code></pre>\n\n<p>Then if <code>LR = 0.02</code>, you stop 99.5% of that, so the EfficientNet (backbone) only gets <code>LR = 0.0001</code></p>",
          "rawMarkdown": "You can do it with `tf.keras.backend.stop_gradient()`\n\n        inp = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n        net = efn.EfficientNetB3(weights='imagenet', include_top=False)\n        x = net(inp)\n        x = tf.keras.layers.GlobalAveragePooling2D()(x)\n        x = 0.005*x + 0.995*tf.keras.backend.stop_gradient(x)\n        x = tf.keras.layers.Dense(512, activation='relu')(x)\n        x = tf.keras.layers.Dense(128, activation='relu')(x)\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\nThen if `LR = 0.02`, you stop 99.5% of that, so the EfficientNet (backbone) only gets `LR = 0.0001`",
          "votes": 10
        }
      ]
    },
    {
      "id": 871987,
      "postDate": "2020-06-02T19:33:53.350Z",
      "content": "<p>Are you using TPUs for training the model? I have similar model architecture but while training, I encountered following error:</p>\n\n<p>NotFoundError: {{function_node __inference_distributed_function_431538}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}} . </p>\n\n<p>Registered:  \n[[PyFunc]] \n[[MultiDeviceIteratorGetNextFromShard]] \n[[RemoteCall]] \n[[IteratorGetNextAsOptional]]</p>\n\n<p>Please help.</p>\n\n<p>Thank You!</p>",
      "rawMarkdown": "Are you using TPUs for training the model? I have similar model architecture but while training, I encountered following error:\n\nNotFoundError: {{function_node __inference_distributed_function_431538}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}} . \n\nRegistered: ",
      "replies": [
        {
          "id": 872006,
          "postDate": "2020-06-02T20:01:19.147Z",
          "content": "<p>I'm currently using GPU but planning to switch to TPU later. Will let you know how it went.</p>",
          "rawMarkdown": "I'm currently using GPU but planning to switch to TPU later. Will let you know how it went."
        }
      ]
    },
    {
      "id": 871760,
      "postDate": "2020-06-02T16:07:08.110Z",
      "content": "<p>Hey very nice results, ive also experimented with resnet18 and was able to get similar results! However what loss are you using? for me FocalLoss gave me slightly better results!</p>",
      "rawMarkdown": "Hey very nice results, ive also experimented with resnet18 and was able to get similar results! However what loss are you using? for me FocalLoss gave me slightly better results!",
      "replies": [
        {
          "id": 871842,
          "postDate": "2020-06-02T17:11:53.190Z",
          "content": "<p>I'm on BCE. </p>",
          "rawMarkdown": "I'm on BCE. "
        }
      ]
    },
    {
      "id": 873355,
      "postDate": "2020-06-04T04:33:53.353Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 873448,
          "postDate": "2020-06-04T06:45:20.100Z",
          "content": "<p>4 types of rotations and 2 types of vertical flips (regular image + flipped) makes 8 total different augmentation options for training</p>",
          "rawMarkdown": "4 types of rotations and 2 types of vertical flips (regular image + flipped) makes 8 total different augmentation options for training",
          "votes": 1
        },
        {
          "id": 873636,
          "postDate": "2020-06-04T10:14:11.777Z",
          "content": "<p>Exactly.</p>",
          "rawMarkdown": "Exactly."
        },
        {
          "id": 873761,
          "postDate": "2020-06-04T12:09:40.343Z",
          "content": "<p>Why wouldn't you just randomrotate in a range? Is it better to do 4 different rotation choices?</p>",
          "rawMarkdown": "Why wouldn't you just randomrotate in a range? Is it better to do 4 different rotation choices?"
        },
        {
          "id": 873821,
          "postDate": "2020-06-04T13:04:01.747Z",
          "content": "<p>Good question. One advantage of 90 deg rotations is that they do not require interpolation, which results in slightly better image quality. However, since I resize the images anyway, this is unlikely to be a big deal. I guess you could just try different options and compare them.</p>",
          "rawMarkdown": "Good question. One advantage of 90 deg rotations is that they do not require interpolation, which results in slightly better image quality. However, since I resize the images anyway, this is unlikely to be a big deal. I guess you could just try different options and compare them."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 871707,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-02T15:29:11.737000",
      "content": "<p>Great job. If people want to recreate this in Kaggle notebooks using GPU or TPU, I uploaded TFRecords containing 256x256 <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>. Then after you write you code, you can switch the 256x256 dataset with my 512x512 or 768x768 for a higher CV/LB. Enjoy!</p>\n\n<p>Rotation augmentation for TFRecords is explained <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\">here</a>. And CutMix and MixUp <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\">here</a></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 871745,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2020-06-02T15:55:19.697000",
      "content": "<p>Interestingly, so far switching to 512x512 resolutions makes the result <strong>worse</strong>. I suspect that's because ResNet18 has quite a small receptive field and couldn't fully utilize higher resolution.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 871749,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T15:58:58.120000",
          "content": "<p>yes, that is interesting</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876217,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-06T14:41:14.853000",
          "content": "<p>Better try that resolution on which it was pretrained.This will boost the model's efficiency...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 884438,
      "author_name": "Andrey Lukyanenko",
      "author_url": "",
      "post_date": "2020-06-13T11:28:55.857000",
      "content": "<p>&gt; 4 rotations + vertical flips (8 options total)</p>\n\n<p>Is this a good way of implementing it?\n<code>\nimport albumentations as A\naugs = A.Compose([\n    A.OneOf([\n        A.Rotate(limit=(0, 0), p=1),\n        A.Rotate(limit=(90, 90), p=1),\n        A.Rotate(limit=(180, 180), p=1),\n        A.Rotate(limit=(270, 270), p=1)]),\n    A.VerticalFlip(p=0.5)\n])\n</code>\nOr is there a more concise code?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 884611,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-13T13:44:41.153000",
          "content": "<p>I guess I used <code>RandomRotate90</code> :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 884640,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2020-06-13T14:11:33.870000",
          "content": "<p>Well, it seems this is much more simple :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 876010,
      "author_name": "Gaurav Sharma",
      "author_url": "",
      "post_date": "2020-06-06T11:08:26.473000",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a> have you tried more epochs?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 876028,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-06T11:18:35.860000",
          "content": "<p>Yes. For this setup, there was no difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876144,
          "author_name": "Gaurav Sharma",
          "author_url": "",
          "post_date": "2020-06-06T13:27:52.363000",
          "content": "<p>are these results are on the provided dataset only or you used external data as well?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 876171,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-06T13:53:26.873000",
          "content": "<p>I'm using only the competition data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 871790,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2020-06-02T16:28:18.627000",
      "content": "<p>Really interesting results. Could you share a notebook for the same? Also , what exactly do you mean by average of the last 4 epochs? Also,  do you think it is possible to run notebooks for this comp on a 6GB 1660Ti?</p>",
      "votes": -2,
      "replies": [
        {
          "id": 871860,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-02T17:19:24.323000",
          "content": "<p>I just averaged predictions on the test set from the last 4 epochs (I predict both validation and test on each epoch). </p>\n\n<p>As for hardware, 1660Ti might be fine for shallow models (such as ResNet18), but deeper models (which are a must for a high score) are better be trained on TPUs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 871896,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2020-06-02T18:00:37.773000",
          "content": "<p>Oh , I see. Thanks for the reply. Guess I'll just have to work with training with checkpoints on Kaggle TPU's :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 888415,
      "author_name": "wqz111",
      "author_url": "",
      "post_date": "2020-06-16T10:34:50.153000",
      "content": "<p>hi！I tried to train the dataset with resnet50, the number of pictures with true label is 584, the number of false is 32000+, that is imbalanced, I trained resnet50 using your strategy, when I predict, the output label for the test images are all 0, can you tell me why?  Thank you!  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 888437,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-16T11:01:40.483000",
          "content": "<p>I can't. What is your AUC score on the training set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888458,
          "author_name": "wqz111",
          "author_url": "",
          "post_date": "2020-06-16T11:21:55.710000",
          "content": "<p>I am changing the code to add this function, is there a threshold to tell the bad from the good, I set it 0.5 before, need I change it to a bigger one ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 888472,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-16T11:34:12.873000",
          "content": "<p>You can keep it 0 or <code>None</code> too, just share the final AUC score on the training set here! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 882245,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-06-11T16:55:47.547000",
      "content": "<p>Hey, I have a issue with my setup, I am using,</p>\n\n<ul>\n<li>Efficientnet-b0</li>\n<li>224x224</li>\n<li>Adam</li>\n<li>0.0001 LR </li>\n<li>Random Horizontal and Vertical flips for training data</li>\n<li>10 epochs</li>\n<li>batch_size = 64</li>\n<li>5 folds (Group)</li>\n</ul>\n\n<p>I am getting a CV of about 0.72 and that too is fluctuating a lot, it sometimes drops to even 0.3. :/ </p>\n\n<p>I am still a newbie, any suggestions on how to improve this would be highly appreciated. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 882271,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-11T17:15:27.297000",
          "content": "<p>Sounds like a bug. Do you normalize images? Are you sure your training/validation/testing pipelines are the same?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 882278,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-11T17:19:32.947000",
          "content": "<p>Yes, I am normalizing them. Yeah, pipeline is same too.</p>\n\n<p>I am thinking of running it for more epochs now, maybe 20, to see if it starts to converge then.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 882357,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-11T18:30:51.577000",
          "content": "<p>It should converge in a couple of epochs, are you reducing your lr during training with a scheduler? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 882627,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-12T02:22:52.707000",
          "content": "<p>Yes, I am using ReduceLROnPlateau as a Scheduler with a patience of 3.</p>\n\n<p>Also, I tried reducing the number of groups to 3 and auc is now going upto 0.8 now, originally it was only reaching 0.72 and it seems much stable than before.</p>\n\n<p>Although I know I can get at least 0.9 CV with this setup, it still is a long way to go.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883925,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-13T04:57:42.793000",
          "content": "<p>Also I want to add one more thing, I am predicting on 2 classes with cross entropy loss, can it be causing this much fluctuation and less auc?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883957,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-13T06:31:20.873000",
          "content": "<p>I presume almost everyone is using cross-entropy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883973,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-13T06:34:48.163000",
          "content": "<p>Okay, then there must be something else which is off, I will be looking for that then. Thanks! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884489,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-13T12:13:17.547000",
          "content": "<p>Can I share my notebook with anyone here? it may be a good deal for me to just get it quick checked with someone else, maybe I can't see what's going wrong and I have already dedicated 7 days to this particular problem. :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884617,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-13T13:46:48.153000",
          "content": "<p>You could just make your notebook public. Sharing with someone privately without being in a team is prohibited.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884624,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-13T13:53:47.127000",
          "content": "<p>Sure, I will do that, will comment here after making it public. Thanks! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884729,
          "author_name": "Gajendra Saraswat",
          "author_url": "",
          "post_date": "2020-06-13T15:26:26.410000",
          "content": "<p>I made the notebook <a href=\"https://www.kaggle.com/sarques/melanoma-classification\">public</a>.</p>\n\n<p>It performed a lot better than the last time as that time the score used to drop to 0.4 sometimes but I think it will still be good if you can take a look. </p>\n\n<p>I took some part of it from other notebooks too as I was just learning how to work with PyTorch. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 873057,
      "author_name": "Jonas Matuzas",
      "author_url": "",
      "post_date": "2020-06-03T18:25:52.160000",
      "content": "<p>\"0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\"- How LR can be different for ResNet end classification layer ? How to write it in tf.keras ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 873070,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-03T18:43:50.940000",
          "content": "<p>you would do do something like this where model.base is your features and fc is your classifier:</p>\n\n<p><code>\noptimizer = optim.AdamW([{'params': model.base.parameters(), 'lr ':1e-4 },\n                         {'params': model.fc.parameters(), 'lr ':2e-2 },\n                        ], lr=lr)\n</code></p>\n\n<p>This is usefull if you are using a pretrained architechture</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 873081,
          "author_name": "Jonas Matuzas",
          "author_url": "",
          "post_date": "2020-06-03T19:08:10.077000",
          "content": "<p>how to initialize model.base and model.fc ? Is it example in tf.keras?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873094,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-03T19:31:07.093000",
          "content": "<p>Sorry i dont really know how to do it in keras</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 873176,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-06-03T21:54:56.533000",
          "content": "<p>hmm - did not hear this before - sounds plausible - can one say this gives more often better results than one lr for head and backbone? What is your experience?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873206,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-03T22:40:48.910000",
          "content": "<p>For me it doesnt really change much, but the idea is since the features are pretrained they just need finetuning and the classifier parameters are set randomly so they need more trainning. However i think this idea is more valuable when training with images similar to imagenet, not sure though..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873462,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-04T06:59:38.260000",
          "content": "<p>You can do it with <code>tf.keras.backend.stop_gradient()</code></p>\n\n<pre><code>    inp = tf.keras.layers.Input(shape=(*IMAGE_SIZE,3))\n    net = efn.EfficientNetB3(weights='imagenet', include_top=False)\n    x = net(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = 0.005*x + 0.995*tf.keras.backend.stop_gradient(x)\n    x = tf.keras.layers.Dense(512, activation='relu')(x)\n    x = tf.keras.layers.Dense(128, activation='relu')(x)\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n</code></pre>\n\n<p>Then if <code>LR = 0.02</code>, you stop 99.5% of that, so the EfficientNet (backbone) only gets <code>LR = 0.0001</code></p>",
          "votes": 10,
          "replies": []
        }
      ]
    },
    {
      "id": 871987,
      "author_name": "SudipPadhye",
      "author_url": "",
      "post_date": "2020-06-02T19:33:53.350000",
      "content": "<p>Are you using TPUs for training the model? I have similar model architecture but while training, I encountered following error:</p>\n\n<p>NotFoundError: {{function_node __inference_distributed_function_431538}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}} . </p>\n\n<p>Registered:  \n[[PyFunc]] \n[[MultiDeviceIteratorGetNextFromShard]] \n[[RemoteCall]] \n[[IteratorGetNextAsOptional]]</p>\n\n<p>Please help.</p>\n\n<p>Thank You!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 872006,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-02T20:01:19.147000",
          "content": "<p>I'm currently using GPU but planning to switch to TPU later. Will let you know how it went.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 871760,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-06-02T16:07:08.110000",
      "content": "<p>Hey very nice results, ive also experimented with resnet18 and was able to get similar results! However what loss are you using? for me FocalLoss gave me slightly better results!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 871842,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-02T17:11:53.190000",
          "content": "<p>I'm on BCE. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 873355,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T04:33:53.353000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 873448,
          "author_name": "Hunter Mitchell",
          "author_url": "",
          "post_date": "2020-06-04T06:45:20.100000",
          "content": "<p>4 types of rotations and 2 types of vertical flips (regular image + flipped) makes 8 total different augmentation options for training</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 873636,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-04T10:14:11.777000",
          "content": "<p>Exactly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873761,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-04T12:09:40.343000",
          "content": "<p>Why wouldn't you just randomrotate in a range? Is it better to do 4 different rotation choices?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873821,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2020-06-04T13:04:01.747000",
          "content": "<p>Good question. One advantage of 90 deg rotations is that they do not require interpolation, which results in slightly better image quality. However, since I resize the images anyway, this is unlikely to be a big deal. I guess you could just try different options and compare them.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "871656": "My current setup:\n\n* ResNet18\n* Images resized to 256x256 \n* Adam\n* 0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\n* 4 rotations + vertical flips (8 options total) on training\n* 10 epochs\n* batch size is 64\n* 3 folds (group + stratify)\n* Average of the last 4 epochs -&gt; gives slightly more stable results than just last/best epoch\n* Training time is less than an hour on a 1080ti\n\n0.899 CV, 0.903 LB",
    "871707": "Great job. If people want to recreate this in Kaggle notebooks using GPU or TPU, I uploaded TFRecords containing 256x256 [here][1]. Then after you write you code, you can switch the 256x256 dataset with my 512x512 or 768x768 for a higher CV/LB. Enjoy!\n\nRotation augmentation for TFRecords is explained [here][2]. And CutMix and MixUp [here][3]\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[2]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132191\n[3]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935",
    "871745": "Interestingly, so far switching to 512x512 resolutions makes the result **worse**. I suspect that's because ResNet18 has quite a small receptive field and couldn't fully utilize higher resolution.",
    "884438": "&gt; 4 rotations + vertical flips (8 options total)\n\nIs this a good way of implementing it?\n```\nimport albumentations as A\naugs = A.Compose([\n    A.OneOf([\n        A.Rotate(limit=(0, 0), p=1),\n        A.Rotate(limit=(90, 90), p=1),\n        A.Rotate(limit=(180, 180), p=1),\n        A.Rotate(limit=(270, 270), p=1)]),\n    A.VerticalFlip(p=0.5)\n])\n```\nOr is there a more concise code?",
    "876010": "@ddanevskyi have you tried more epochs?",
    "871790": "Really interesting results. Could you share a notebook for the same? Also , what exactly do you mean by average of the last 4 epochs? Also,  do you think it is possible to run notebooks for this comp on a 6GB 1660Ti?",
    "888415": "hi！I tried to train the dataset with resnet50, the number of pictures with true label is 584, the number of false is 32000+, that is imbalanced, I trained resnet50 using your strategy, when I predict, the output label for the test images are all 0, can you tell me why?  Thank you!  ",
    "882245": "Hey, I have a issue with my setup, I am using,\n\n- Efficientnet-b0\n- 224x224\n- Adam\n- 0.0001 LR \n- Random Horizontal and Vertical flips for training data\n- 10 epochs\n- batch_size = 64\n- 5 folds (Group)\n\nI am getting a CV of about 0.72 and that too is fluctuating a lot, it sometimes drops to even 0.3. :/ \n\nI am still a newbie, any suggestions on how to improve this would be highly appreciated. ",
    "873057": "\"0.0001 LR for ResNet and 0.02 for randomly initialized classification layer\"- How LR can be different for ResNet end classification layer ? How to write it in tf.keras ?",
    "871987": "Are you using TPUs for training the model? I have similar model architecture but while training, I encountered following error:\n\nNotFoundError: {{function_node __inference_distributed_function_431538}} No registered 'PyFunc' OpKernel for 'CPU' devices compatible with node {{node PyFunc}} . \n\nRegistered: ",
    "871760": "Hey very nice results, ive also experimented with resnet18 and was able to get similar results! However what loss are you using? for me FocalLoss gave me slightly better results!",
    "873355": ""
  }
}