{
  "id": 155201,
  "title": "dealing with fluctuating validation AUC in a single fold ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155201",
  "author_name": "hengck23",
  "post_date": "2020-05-31T17:50:20.860000",
  "votes": 38,
  "comment_count": 23,
  "views": 0,
  "content": "<p>this is what you may see if you do AUC validation at each training iterations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F640e8e8a3b40c849496aa4551e93a507%2FSelection_037.png?generation=1590947354140255&amp;alt=media\" alt=\"\"></p>\n\n<p>it is not stable. you would not know where to stop. if you make submission to kaggler server, public LB seems not stable too. This is because of the very low number of +ve test samples in validation set.</p>",
  "messages": [
    {
      "id": 869060,
      "postDate": "2020-05-31T17:50:20.860Z",
      "content": "<p>this is what you may see if you do AUC validation at each training iterations:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F640e8e8a3b40c849496aa4551e93a507%2FSelection_037.png?generation=1590947354140255&amp;alt=media\" alt=\"\"></p>\n\n<p>it is not stable. you would not know where to stop. if you make submission to kaggler server, public LB seems not stable too. This is because of the very low number of +ve test samples in validation set.</p>",
      "rawMarkdown": "this is what you may see if you do AUC validation at each training iterations:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F640e8e8a3b40c849496aa4551e93a507%2FSelection_037.png?generation=1590947354140255&amp;alt=media)\n\nit is not stable. you would not know where to stop. if you make submission to kaggler server, public LB seems not stable too. This is because of the very low number of +ve test samples in validation set.",
      "votes": 38
    },
    {
      "id": 869062,
      "postDate": "2020-05-31T17:52:17.893Z",
      "content": "<p>instead of selecting one best model, you can average several models from different training iterations from a single training fold:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcb45fa08e186d0ee74f53a29cded2911%2FSelection_038.png?generation=1590947488907445&amp;alt=media\" alt=\"\"></p>\n\n<p>it should lead to better \"average validation AUC\" (and hopefully better public LB as well)</p>",
      "rawMarkdown": "instead of selecting one best model, you can average several models from different training iterations from a single training fold:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcb45fa08e186d0ee74f53a29cded2911%2FSelection_038.png?generation=1590947488907445&amp;alt=media)\n\nit should lead to better \"average validation AUC\" (and hopefully better public LB as well)",
      "votes": 6,
      "replies": [
        {
          "id": 869064,
          "postDate": "2020-05-31T17:53:43.337Z",
          "content": "<p>as a comparison, here are the distributions of models at different iterations and different validation AUC\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F48a521f89e6c2fd66cc82d5016562bb6%2FSelection_039.png?generation=1590947614481252&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "as a comparison, here are the distributions of models at different iterations and different validation AUC\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F48a521f89e6c2fd66cc82d5016562bb6%2FSelection_039.png?generation=1590947614481252&amp;alt=media)\n"
        },
        {
          "id": 869068,
          "postDate": "2020-05-31T17:58:26.087Z",
          "content": "<p>you can read this paper:\n<a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/ISIC/Perez_Solo_or_Ensemble_Choosing_a_CNN_Architecture_for_Melanoma_Classification_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/ISIC/Perez_Solo_or_Ensemble_Choosing_a_CNN_Architecture_for_Melanoma_Classification_CVPRW_2019_paper.pdf</a></p>\n\n<p><a href=\"http://repositorio.unicamp.br/bitstream/REPOSIP/335713/1/Perez_FabioViniciusMoreira_M.pdf\">http://repositorio.unicamp.br/bitstream/REPOSIP/335713/1/Perez_FabioViniciusMoreira_M.pdf</a></p>\n\n<p>\"Solo or Ensemble? Choosing a CNN Architecture for Melanoma Classification\"\n- Fabio Perez, cvpr worshop 2019\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe8e8ab968922e1e531c5dfe4117b3eeb%2Fmedia_users_user_228264_project_358048_images_x3.png?generation=1590947903714924&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "you can read this paper:\nhttp://openaccess.thecvf.com/content_CVPRW_2019/papers/ISIC/Perez_Solo_or_Ensemble_Choosing_a_CNN_Architecture_for_Melanoma_Classification_CVPRW_2019_paper.pdf\n\nhttp://repositorio.unicamp.br/bitstream/REPOSIP/335713/1/Perez_FabioViniciusMoreira_M.pdf\n\n\"Solo or Ensemble? Choosing a CNN Architecture for Melanoma Classification\"\n- Fabio Perez, cvpr worshop 2019\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe8e8ab968922e1e531c5dfe4117b3eeb%2Fmedia_users_user_228264_project_358048_images_x3.png?generation=1590947903714924&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 880305,
      "postDate": "2020-06-10T07:03:50.967Z",
      "content": "<p>more comprehension results comparing efficientnetb1,b2,b3 on 384x288,512x384</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F96398050142d2ff63d00fff3416bc62f%2FSelection_119.png?generation=1591772628190468&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "more comprehension results comparing efficientnetb1,b2,b3 on 384x288,512x384\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F96398050142d2ff63d00fff3416bc62f%2FSelection_119.png?generation=1591772628190468&amp;alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 880322,
          "postDate": "2020-06-10T07:14:36.260Z",
          "content": "<ul>\n<li><p>it seems that one reason for the fluctuating validation AUC is the use of \"look ahead optimizer\".</p></li>\n<li><p>replace look ahead with sgd at fine-tuning reduces the fluctuation. the resulting SWA ensemble improves validation AUC, but not public LB AUC </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd74148c1aae9e13d65026ddd7019c182%2FSelection_120.png?generation=1591773266517358&amp;alt=media\" alt=\"\"></p></li>\n</ul>",
          "rawMarkdown": "- it seems that one reason for the fluctuating validation AUC is the use of \"look ahead optimizer\".\n\n- replace look ahead with sgd at fine-tuning reduces the fluctuation. the resulting SWA ensemble improves validation AUC, but not public LB AUC \n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd74148c1aae9e13d65026ddd7019c182%2FSelection_120.png?generation=1591773266517358&amp;alt=media)\n\n",
          "votes": 1
        },
        {
          "id": 880435,
          "postDate": "2020-06-10T09:50:46.663Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> sorry for bothering you with beginner questions, I just try to understand...</p>\n\n<p>1) batch_sizes on tpu can be way higher than 50, wouldn't it help to work with eg 256\n2) looking at the validation loss image, it seems you are breaking the weights - so wouldn't it be better to use a less agressive learning rate?\n3) am I right, that you are pre-learning a base model first (for eg 20 epochs) and then do cv finetuning, to save time?</p>\n\n<p>I really like to follow your work!</p>",
          "rawMarkdown": "@hengck23 sorry for bothering you with beginner questions, I just try to understand...\n\n1) batch_sizes on tpu can be way higher than 50, wouldn't it help to work with eg 256\n2) looking at the validation loss image, it seems you are breaking the weights - so wouldn't it be better to use a less agressive learning rate?\n3) am I right, that you are pre-learning a base model first (for eg 20 epochs) and then do cv finetuning, to save time?\n\nI really like to follow your work!"
        },
        {
          "id": 880456,
          "postDate": "2020-06-10T10:12:12.280Z",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> </p>\n\n<p>1) i am not using tpu, so there is a limit to the batch size i can use. in practice, there is a optimal batch size for different problem. good values could be between 32 to 512.  you may also want to read about:</p>\n\n<p>\"FOUR THINGS EVERYONE SHOULD KNOW TO IMPROVE BATCH NORMALIZATION\"\n<a href=\"https://openreview.net/forum?id=HJx8HANFDH\">https://openreview.net/forum?id=HJx8HANFDH</a></p>\n\n<p>2) I am using radam and it is supposed to automatically adjusting learning rate.</p>\n\n<p>3) not true. the performance of 512x384 is not up to my expectation. so i am debugging the training pipleline. From past experience, i reckon that there is something wrong with batch_size, learning rate or other hyperparameters, etc. Hence i start to do fine-tuning, etc</p>\n\n<p>If the training pipeline is perfect,  512x384 should not perform worse than 384x288 (assuming that you can modify everything from model, hyperparameters, etc). Performance should only be limited by the amount of information in the data (assuming that you can perfectly extract information from the data). </p>\n\n<p>512x384 has all information of 384x288. In fact 512x384 may contain more information as it is of higher resolution</p>",
          "rawMarkdown": "@romanweilguny \n\n1) i am not using tpu, so there is a limit to the batch size i can use. in practice, there is a optimal batch size for different problem. good values could be between 32 to 512.  you may also want to read about:\n\n\"FOUR THINGS EVERYONE SHOULD KNOW TO IMPROVE BATCH NORMALIZATION\"\nhttps://openreview.net/forum?id=HJx8HANFDH\n\n2) I am using radam and it is supposed to automatically adjusting learning rate.\n\n3) not true. the performance of 512x384 is not up to my expectation. so i am debugging the training pipleline. From past experience, i reckon that there is something wrong with batch_size, learning rate or other hyperparameters, etc. Hence i start to do fine-tuning, etc\n\nIf the training pipeline is perfect,  512x384 should not perform worse than 384x288 (assuming that you can modify everything from model, hyperparameters, etc). Performance should only be limited by the amount of information in the data (assuming that you can perfectly extract information from the data). \n\n512x384 has all information of 384x288. In fact 512x384 may contain more information as it is of higher resolution",
          "votes": 2
        },
        {
          "id": 880866,
          "postDate": "2020-06-10T15:56:43.553Z",
          "content": "<p>updated results for efficientnet b0 256x256</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd403fe67ab3f1648a35078c42de13f7f%2FSelection_121.png?generation=1591804601557271&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "updated results for efficientnet b0 256x256\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd403fe67ab3f1648a35078c42de13f7f%2FSelection_121.png?generation=1591804601557271&amp;alt=media)\n"
        },
        {
          "id": 880870,
          "postDate": "2020-06-10T15:59:39.527Z",
          "content": "<p>the loss used for all experiments is \"soft margin focal loss\"</p>\n\n<p>```</p>\n\n<h1>ICME.2019Skeleton-BasedActionRecognitionwithSynchronousLocalandNon-LocalSpatio-TemporalLearningandFrequencyAttention.pdf</h1>\n\n<h1>Soft-margin focal loss</h1>\n\n<p>def criterion_margin_focal_binary_cross_entropy(logit, truth):\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)</p>\n\n<pre><code>logit = logit.view(-1)\ntruth = truth.view(-1)\nlog_pos = -F.logsigmoid( logit)\nlog_neg = -F.logsigmoid(-logit)\n\nlog_prob = truth*log_pos + (1-truth)*log_neg\nprob = torch.exp(-log_prob)\nmargin = torch.log(em +(1-em)*prob)\n\nweight = truth*weight_pos + (1-truth)*weight_neg\nloss = margin + weight*(1 - prob) ** gamma * log_prob\n\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "the loss used for all experiments is \"soft margin focal loss\"\n\n```\n\n\n# ICME.2019Skeleton-BasedActionRecognitionwithSynchronousLocalandNon-LocalSpatio-TemporalLearningandFrequencyAttention.pdf\n# Soft-margin focal loss\ndef criterion_margin_focal_binary_cross_entropy(logit, truth):\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)\n\n    logit = logit.view(-1)\n    truth = truth.view(-1)\n    log_pos = -F.logsigmoid( logit)\n    log_neg = -F.logsigmoid(-logit)\n\n    log_prob = truth*log_pos + (1-truth)*log_neg\n    prob = torch.exp(-log_prob)\n    margin = torch.log(em +(1-em)*prob)\n\n    weight = truth*weight_pos + (1-truth)*weight_neg\n    loss = margin + weight*(1 - prob) ** gamma * log_prob\n\n    loss = loss.mean()\n    return loss\n\n```",
          "votes": 8
        },
        {
          "id": 880902,
          "postDate": "2020-06-10T16:18:42.163Z",
          "content": "<p>Thanks so much for posting this - super helpful to see the approach and specific implementation details</p>",
          "rawMarkdown": "Thanks so much for posting this - super helpful to see the approach and specific implementation details"
        },
        {
          "id": 880993,
          "postDate": "2020-06-10T17:26:51.727Z",
          "content": "<p>experiment results on bce, focal loss, focal loss + margin <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2f33293718047725ce12acf2c3726ede%2FSelection_123.png?generation=1591810171550457&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "experiment results on bce, focal loss, focal loss + margin  \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2f33293718047725ce12acf2c3726ede%2FSelection_123.png?generation=1591810171550457&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 881034,
          "postDate": "2020-06-10T18:00:54.547Z",
          "content": "<p>typical training log file of an \"easy fold\" with look-ahead radam optimser on soft margin focal loss</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F9d4e4606bf8f96ca93e5de0cf3f615cd%2FSelection_128.png?generation=1591812052118507&amp;alt=media\" alt=\"\"></p>\n\n<p>true positive rate (tpr) and false positive rate (fpr) are computed at threshold=0.10. </p>\n\n<p>dataset detail\n```\n** dataset setting **\nbatch_size = 32\ntrain_dataset : \n    len   = 66479\n    mode  = train\n    image_size = (512, 384)\n    external = ['isic_2019', 'isic_2016', 'isic_archive']\n    split = split/by_patient_id_200/fold2_train_30265.pickle\n    target =\n        0 60940  (0.917)\n        1  5539  (0.083)</p>\n\n<p>valid_dataset : \n    len   = 2861\n    mode  = train\n    image_size = (512, 384)\n    external = []\n    split = split/by_patient_id_200/fold2_valid_2861.pickle\n    target =\n        0  2818  (0.985)\n        1    43  (0.015)</p>\n\n<p>```</p>",
          "rawMarkdown": "typical training log file of an \"easy fold\" with look-ahead radam optimser on soft margin focal loss\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F9d4e4606bf8f96ca93e5de0cf3f615cd%2FSelection_128.png?generation=1591812052118507&amp;alt=media)\n\ntrue positive rate (tpr) and false positive rate (fpr) are computed at threshold=0.10. \n\ndataset detail\n```\n** dataset setting **\nbatch_size = 32\ntrain_dataset : \n\tlen   = 66479\n\tmode  = train\n\timage_size = (512, 384)\n\texternal = ['isic_2019', 'isic_2016', 'isic_archive']\n\tsplit = split/by_patient_id_200/fold2_train_30265.pickle\n\ttarget =\n\t\t0 60940  (0.917)\n\t\t1  5539  (0.083)\n\nvalid_dataset : \n\tlen   = 2861\n\tmode  = train\n\timage_size = (512, 384)\n\texternal = []\n\tsplit = split/by_patient_id_200/fold2_valid_2861.pickle\n\ttarget =\n\t\t0  2818  (0.985)\n\t\t1    43  (0.015)\n\n```\n\n"
        },
        {
          "id": 881296,
          "postDate": "2020-06-10T22:24:01.017Z",
          "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a> for all this info! Just a question, did you use a class balance sampler with focal loss? To my understanding focal loss is to handle imbalanced data but if we have a balanced sampling does that work well?</p>",
          "rawMarkdown": "Thanks @hengck23 for all this info! Just a question, did you use a class balance sampler with focal loss? To my understanding focal loss is to handle imbalanced data but if we have a balanced sampling does that work well?"
        },
        {
          "id": 883871,
          "postDate": "2020-06-13T02:54:27.960Z",
          "content": "<p>It can also be weighting the easy and hard samples.</p>",
          "rawMarkdown": "It can also be weighting the easy and hard samples."
        },
        {
          "id": 917598,
          "postDate": "2020-07-06T16:18:50.377Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>think you will like this:</p>\n\n<p><a href=\"https://pechyonkin.me/stochastic-weight-averaging/\">https://pechyonkin.me/stochastic-weight-averaging/</a></p>",
          "rawMarkdown": "@hengck23 \n\nthink you will like this:\n\nhttps://pechyonkin.me/stochastic-weight-averaging/"
        }
      ]
    },
    {
      "id": 881015,
      "postDate": "2020-06-10T17:44:36.667Z",
      "content": "<p>how much does each model agree with each other?\nthe graph is for ensemble at public LB =0.946 \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffe08e9692e405a8c6a7f7e0a3b293188%2FSelection_126.png?generation=1591811055524824&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff608762c7cae73be5a08496d13a4505f%2FSelection_127.png?generation=1591811074376245&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how much does each model agree with each other?\nthe graph is for ensemble at public LB =0.946 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffe08e9692e405a8c6a7f7e0a3b293188%2FSelection_126.png?generation=1591811055524824&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff608762c7cae73be5a08496d13a4505f%2FSelection_127.png?generation=1591811074376245&amp;alt=media)\n",
      "votes": 1
    },
    {
      "id": 885860,
      "postDate": "2020-06-14T14:25:14.290Z",
      "content": "<p>Which Loss function gives better result?? Thank you</p>",
      "rawMarkdown": "Which Loss function gives better result?? Thank you\n"
    },
    {
      "id": 878785,
      "postDate": "2020-06-08T20:52:06.837Z",
      "content": "<p>What do you think about this: <a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\">https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/</a>\nI tried it and it always gives me a boost in CV. However lb is verrrry unstable so im not sure what to beleive</p>",
      "rawMarkdown": "What do you think about this: https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\nI tried it and it always gives me a boost in CV. However lb is verrrry unstable so im not sure what to beleive"
    },
    {
      "id": 872224,
      "postDate": "2020-06-03T02:25:33.703Z",
      "content": "<p>I've seen this type of fluctuating AUC during training before when the target is unbalanced (in IEEE Fraud Comp). It seems that maximizing cross entropy doesn't maximize AUC. So for the same cross entropy, AUC can take on different values. It would be nice if we could directly maximize AUC.</p>",
      "rawMarkdown": "I've seen this type of fluctuating AUC during training before when the target is unbalanced (in IEEE Fraud Comp). It seems that maximizing cross entropy doesn't maximize AUC. So for the same cross entropy, AUC can take on different values. It would be nice if we could directly maximize AUC.",
      "replies": [
        {
          "id": 872557,
          "postDate": "2020-06-03T09:58:50.043Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nyou can \"directly\" maximize AUC. any ranking loss or large margin loss should help. as an example, this is sigmoid based ranking loss i have tried. There seems to be some \"optimistic\" improvement. but nothing is conclusive now as i am still doing experiments.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a922021a904ffd4d3a5be3b12ff31c5%2FSelection_086.png?generation=1591178323412302&amp;alt=media\" alt=\"\"></p>\n\n<p>also, there is the famous google paper and code:\n\"Scalable Learning of Non-Decomposable Objectives\"\n<a href=\"https://arxiv.org/pdf/1608.04802.pdf\">https://arxiv.org/pdf/1608.04802.pdf</a>\n<a href=\"https://github.com/facebookresearch/pytext/blob/master/pytext/loss/loss.py\">https://github.com/facebookresearch/pytext/blob/master/pytext/loss/loss.py</a></p>\n\n<p>the main problem of ranking based loss is that your batch-size and pair-sampling has to be correct. This may be non trivial.</p>",
          "rawMarkdown": "@cdeotte \nyou can \"directly\" maximize AUC. any ranking loss or large margin loss should help. as an example, this is sigmoid based ranking loss i have tried. There seems to be some \"optimistic\" improvement. but nothing is conclusive now as i am still doing experiments.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a922021a904ffd4d3a5be3b12ff31c5%2FSelection_086.png?generation=1591178323412302&amp;alt=media)\n\n\n\nalso, there is the famous google paper and code:\n\"Scalable Learning of Non-Decomposable Objectives\"\nhttps://arxiv.org/pdf/1608.04802.pdf\nhttps://github.com/facebookresearch/pytext/blob/master/pytext/loss/loss.py\n\nthe main problem of ranking based loss is that your batch-size and pair-sampling has to be correct. This may be non trivial.\n",
          "votes": 7
        }
      ]
    },
    {
      "id": 872001,
      "postDate": "2020-06-02T19:53:45.263Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  hmm- am I misunderstanding something?\nyour image size geometry is 384x512 while the images have width&gt;height geometry?</p>",
      "rawMarkdown": "@hengck23  hmm- am I misunderstanding something?\nyour image size geometry is 384x512 while the images have width&gt;height geometry?",
      "replies": [
        {
          "id": 872016,
          "postDate": "2020-06-02T20:23:39.880Z",
          "content": "<p>height=384, width=512 in the experiments here</p>",
          "rawMarkdown": "height=384, width=512 in the experiments here"
        }
      ]
    },
    {
      "id": 883872,
      "postDate": "2020-06-13T02:54:54.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 869062,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-05-31T17:52:17.893000",
      "content": "<p>instead of selecting one best model, you can average several models from different training iterations from a single training fold:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcb45fa08e186d0ee74f53a29cded2911%2FSelection_038.png?generation=1590947488907445&amp;alt=media\" alt=\"\"></p>\n\n<p>it should lead to better \"average validation AUC\" (and hopefully better public LB as well)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 869064,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-31T17:53:43.337000",
          "content": "<p>as a comparison, here are the distributions of models at different iterations and different validation AUC\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F48a521f89e6c2fd66cc82d5016562bb6%2FSelection_039.png?generation=1590947614481252&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869068,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-31T17:58:26.087000",
          "content": "<p>you can read this paper:\n<a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/ISIC/Perez_Solo_or_Ensemble_Choosing_a_CNN_Architecture_for_Melanoma_Classification_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/ISIC/Perez_Solo_or_Ensemble_Choosing_a_CNN_Architecture_for_Melanoma_Classification_CVPRW_2019_paper.pdf</a></p>\n\n<p><a href=\"http://repositorio.unicamp.br/bitstream/REPOSIP/335713/1/Perez_FabioViniciusMoreira_M.pdf\">http://repositorio.unicamp.br/bitstream/REPOSIP/335713/1/Perez_FabioViniciusMoreira_M.pdf</a></p>\n\n<p>\"Solo or Ensemble? Choosing a CNN Architecture for Melanoma Classification\"\n- Fabio Perez, cvpr worshop 2019\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe8e8ab968922e1e531c5dfe4117b3eeb%2Fmedia_users_user_228264_project_358048_images_x3.png?generation=1590947903714924&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 880305,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-10T07:03:50.967000",
      "content": "<p>more comprehension results comparing efficientnetb1,b2,b3 on 384x288,512x384</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F96398050142d2ff63d00fff3416bc62f%2FSelection_119.png?generation=1591772628190468&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 880322,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T07:14:36.260000",
          "content": "<ul>\n<li><p>it seems that one reason for the fluctuating validation AUC is the use of \"look ahead optimizer\".</p></li>\n<li><p>replace look ahead with sgd at fine-tuning reduces the fluctuation. the resulting SWA ensemble improves validation AUC, but not public LB AUC </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd74148c1aae9e13d65026ddd7019c182%2FSelection_120.png?generation=1591773266517358&amp;alt=media\" alt=\"\"></p></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 880435,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-06-10T09:50:46.663000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> sorry for bothering you with beginner questions, I just try to understand...</p>\n\n<p>1) batch_sizes on tpu can be way higher than 50, wouldn't it help to work with eg 256\n2) looking at the validation loss image, it seems you are breaking the weights - so wouldn't it be better to use a less agressive learning rate?\n3) am I right, that you are pre-learning a base model first (for eg 20 epochs) and then do cv finetuning, to save time?</p>\n\n<p>I really like to follow your work!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880456,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T10:12:12.280000",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> </p>\n\n<p>1) i am not using tpu, so there is a limit to the batch size i can use. in practice, there is a optimal batch size for different problem. good values could be between 32 to 512.  you may also want to read about:</p>\n\n<p>\"FOUR THINGS EVERYONE SHOULD KNOW TO IMPROVE BATCH NORMALIZATION\"\n<a href=\"https://openreview.net/forum?id=HJx8HANFDH\">https://openreview.net/forum?id=HJx8HANFDH</a></p>\n\n<p>2) I am using radam and it is supposed to automatically adjusting learning rate.</p>\n\n<p>3) not true. the performance of 512x384 is not up to my expectation. so i am debugging the training pipleline. From past experience, i reckon that there is something wrong with batch_size, learning rate or other hyperparameters, etc. Hence i start to do fine-tuning, etc</p>\n\n<p>If the training pipeline is perfect,  512x384 should not perform worse than 384x288 (assuming that you can modify everything from model, hyperparameters, etc). Performance should only be limited by the amount of information in the data (assuming that you can perfectly extract information from the data). </p>\n\n<p>512x384 has all information of 384x288. In fact 512x384 may contain more information as it is of higher resolution</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 880866,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T15:56:43.553000",
          "content": "<p>updated results for efficientnet b0 256x256</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd403fe67ab3f1648a35078c42de13f7f%2FSelection_121.png?generation=1591804601557271&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880870,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T15:59:39.527000",
          "content": "<p>the loss used for all experiments is \"soft margin focal loss\"</p>\n\n<p>```</p>\n\n<h1>ICME.2019Skeleton-BasedActionRecognitionwithSynchronousLocalandNon-LocalSpatio-TemporalLearningandFrequencyAttention.pdf</h1>\n\n<h1>Soft-margin focal loss</h1>\n\n<p>def criterion_margin_focal_binary_cross_entropy(logit, truth):\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)</p>\n\n<pre><code>logit = logit.view(-1)\ntruth = truth.view(-1)\nlog_pos = -F.logsigmoid( logit)\nlog_neg = -F.logsigmoid(-logit)\n\nlog_prob = truth*log_pos + (1-truth)*log_neg\nprob = torch.exp(-log_prob)\nmargin = torch.log(em +(1-em)*prob)\n\nweight = truth*weight_pos + (1-truth)*weight_neg\nloss = margin + weight*(1 - prob) ** gamma * log_prob\n\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 880902,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-06-10T16:18:42.163000",
          "content": "<p>Thanks so much for posting this - super helpful to see the approach and specific implementation details</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880993,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T17:26:51.727000",
          "content": "<p>experiment results on bce, focal loss, focal loss + margin <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2f33293718047725ce12acf2c3726ede%2FSelection_123.png?generation=1591810171550457&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 881034,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-10T18:00:54.547000",
          "content": "<p>typical training log file of an \"easy fold\" with look-ahead radam optimser on soft margin focal loss</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F9d4e4606bf8f96ca93e5de0cf3f615cd%2FSelection_128.png?generation=1591812052118507&amp;alt=media\" alt=\"\"></p>\n\n<p>true positive rate (tpr) and false positive rate (fpr) are computed at threshold=0.10. </p>\n\n<p>dataset detail\n```\n** dataset setting **\nbatch_size = 32\ntrain_dataset : \n    len   = 66479\n    mode  = train\n    image_size = (512, 384)\n    external = ['isic_2019', 'isic_2016', 'isic_archive']\n    split = split/by_patient_id_200/fold2_train_30265.pickle\n    target =\n        0 60940  (0.917)\n        1  5539  (0.083)</p>\n\n<p>valid_dataset : \n    len   = 2861\n    mode  = train\n    image_size = (512, 384)\n    external = []\n    split = split/by_patient_id_200/fold2_valid_2861.pickle\n    target =\n        0  2818  (0.985)\n        1    43  (0.015)</p>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 881296,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-06-10T22:24:01.017000",
          "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a> for all this info! Just a question, did you use a class balance sampler with focal loss? To my understanding focal loss is to handle imbalanced data but if we have a balanced sampling does that work well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883871,
          "author_name": "Jun Liu",
          "author_url": "",
          "post_date": "2020-06-13T02:54:27.960000",
          "content": "<p>It can also be weighting the easy and hard samples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917598,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-06T16:18:50.377000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>think you will like this:</p>\n\n<p><a href=\"https://pechyonkin.me/stochastic-weight-averaging/\">https://pechyonkin.me/stochastic-weight-averaging/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 881015,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-10T17:44:36.667000",
      "content": "<p>how much does each model agree with each other?\nthe graph is for ensemble at public LB =0.946 \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffe08e9692e405a8c6a7f7e0a3b293188%2FSelection_126.png?generation=1591811055524824&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff608762c7cae73be5a08496d13a4505f%2FSelection_127.png?generation=1591811074376245&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 885860,
      "author_name": "PIkachu",
      "author_url": "",
      "post_date": "2020-06-14T14:25:14.290000",
      "content": "<p>Which Loss function gives better result?? Thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 878785,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-06-08T20:52:06.837000",
      "content": "<p>What do you think about this: <a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\">https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/</a>\nI tried it and it always gives me a boost in CV. However lb is verrrry unstable so im not sure what to beleive</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 872224,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-03T02:25:33.703000",
      "content": "<p>I've seen this type of fluctuating AUC during training before when the target is unbalanced (in IEEE Fraud Comp). It seems that maximizing cross entropy doesn't maximize AUC. So for the same cross entropy, AUC can take on different values. It would be nice if we could directly maximize AUC.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 872557,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-03T09:58:50.043000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> \nyou can \"directly\" maximize AUC. any ranking loss or large margin loss should help. as an example, this is sigmoid based ranking loss i have tried. There seems to be some \"optimistic\" improvement. but nothing is conclusive now as i am still doing experiments.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F6a922021a904ffd4d3a5be3b12ff31c5%2FSelection_086.png?generation=1591178323412302&amp;alt=media\" alt=\"\"></p>\n\n<p>also, there is the famous google paper and code:\n\"Scalable Learning of Non-Decomposable Objectives\"\n<a href=\"https://arxiv.org/pdf/1608.04802.pdf\">https://arxiv.org/pdf/1608.04802.pdf</a>\n<a href=\"https://github.com/facebookresearch/pytext/blob/master/pytext/loss/loss.py\">https://github.com/facebookresearch/pytext/blob/master/pytext/loss/loss.py</a></p>\n\n<p>the main problem of ranking based loss is that your batch-size and pair-sampling has to be correct. This may be non trivial.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 872001,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-06-02T19:53:45.263000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  hmm- am I misunderstanding something?\nyour image size geometry is 384x512 while the images have width&gt;height geometry?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 872016,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-02T20:23:39.880000",
          "content": "<p>height=384, width=512 in the experiments here</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 883872,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-13T02:54:54.833000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "869060": "this is what you may see if you do AUC validation at each training iterations:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F640e8e8a3b40c849496aa4551e93a507%2FSelection_037.png?generation=1590947354140255&amp;alt=media)\n\nit is not stable. you would not know where to stop. if you make submission to kaggler server, public LB seems not stable too. This is because of the very low number of +ve test samples in validation set.",
    "869062": "instead of selecting one best model, you can average several models from different training iterations from a single training fold:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcb45fa08e186d0ee74f53a29cded2911%2FSelection_038.png?generation=1590947488907445&amp;alt=media)\n\nit should lead to better \"average validation AUC\" (and hopefully better public LB as well)",
    "880305": "more comprehension results comparing efficientnetb1,b2,b3 on 384x288,512x384\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F96398050142d2ff63d00fff3416bc62f%2FSelection_119.png?generation=1591772628190468&amp;alt=media)\n",
    "881015": "how much does each model agree with each other?\nthe graph is for ensemble at public LB =0.946 \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffe08e9692e405a8c6a7f7e0a3b293188%2FSelection_126.png?generation=1591811055524824&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff608762c7cae73be5a08496d13a4505f%2FSelection_127.png?generation=1591811074376245&amp;alt=media)\n",
    "885860": "Which Loss function gives better result?? Thank you\n",
    "878785": "What do you think about this: https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\nI tried it and it always gives me a boost in CV. However lb is verrrry unstable so im not sure what to beleive",
    "872224": "I've seen this type of fluctuating AUC during training before when the target is unbalanced (in IEEE Fraud Comp). It seems that maximizing cross entropy doesn't maximize AUC. So for the same cross entropy, AUC can take on different values. It would be nice if we could directly maximize AUC.",
    "872001": "@hengck23  hmm- am I misunderstanding something?\nyour image size geometry is 384x512 while the images have width&gt;height geometry?",
    "883872": ""
  }
}