{
  "id": 79086,
  "title": "Climbing the mountain (ideas for boosting the score)",
  "url": "/competitions/humpback-whale-identification/discussion/79086",
  "author_name": "",
  "post_date": "2019-01-30T18:30:22.238653100Z",
  "votes": 54,
  "comment_count": 60,
  "views": 0,
  "content": "<p>Greetings,\nAfter many experiments and substantial refactoring the model, <a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">my kernel</a> could reach <strong>0.805 LB</strong> score for <strong>training within the kernel time limit</strong> (no ensembling and external weights apart from ones pretrained on ImageNet) <strong>with using a similarity distance based approach</strong>. I just wanted to summarize milestones of the kernel score improvement (you can check details in the kernel), and I hope you will find them useful and could apply to your models. Also you can share your ides on the improvement of the score. The provided values of score show the result obtained after running the model within the kernel time limit, each new line indicates the modifications made to the model and the corresponding value of public LB.</p>\n\n<p>1) Triplet network with triplet loss (random selection of images): ~0.3 LB \n2) Batch all contrastive loss (compute loss based on all pairs in a batch, for 96 images per batch this method gives 9120 comparisons instead of 32 triplets): 0.5 LB\n3) Rectangular crops based on bounding boxes (without image distortion) instead of random ones: 0.55 LB\n4) Increase the dimensionality of the embedding space from 64 to 256 + ResNeXt50 instead of ResNet34: 0.606 LB (V1 and 2 of <a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">this kernel</a>)\n5) Metric learning <code>d^2 = (x1-x2).T*A*(x1-x2)</code> instead of Euclidean distance: 0.655 LB\n6) Optimized form of metric (check kernel for details) + hard negative example mining + lr optimization + going to smaller images of size 128x384 instead of 192x576 + averaging only nonzero contributions when loss is computed: 0.699 LB  (V6)\n7) Switching to bounding box crops rescaled to square images of size 224x224: 0.740 (0.748 in a preliminary test). (V7)\n8) Adding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.771 (0.780 in a preliminary test). (V9)\n9) DenseNet169 backbone: 0.800 (V11)\n10) DenseNet121 backbone: 0.805 (V13)</p>",
  "messages": [
    {
      "id": "463855",
      "postDate": "01/30/2019 18:30:22",
      "content": "<p>Greetings,\nAfter many experiments and substantial refactoring the model, <a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">my kernel</a> could reach <strong>0.805 LB</strong> score for <strong>training within the kernel time limit</strong> (no ensembling and external weights apart from ones pretrained on ImageNet) <strong>with using a similarity distance based approach</strong>. I just wanted to summarize milestones of the kernel score improvement (you can check details in the kernel), and I hope you will find them useful and could apply to your models. Also you can share your ides on the improvement of the score. The provided values of score show the result obtained after running the model within the kernel time limit, each new line indicates the modifications made to the model and the corresponding value of public LB.</p>\n\n<p>1) Triplet network with triplet loss (random selection of images): ~0.3 LB \n2) Batch all contrastive loss (compute loss based on all pairs in a batch, for 96 images per batch this method gives 9120 comparisons instead of 32 triplets): 0.5 LB\n3) Rectangular crops based on bounding boxes (without image distortion) instead of random ones: 0.55 LB\n4) Increase the dimensionality of the embedding space from 64 to 256 + ResNeXt50 instead of ResNet34: 0.606 LB (V1 and 2 of <a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">this kernel</a>)\n5) Metric learning <code>d^2 = (x1-x2).T*A*(x1-x2)</code> instead of Euclidean distance: 0.655 LB\n6) Optimized form of metric (check kernel for details) + hard negative example mining + lr optimization + going to smaller images of size 128x384 instead of 192x576 + averaging only nonzero contributions when loss is computed: 0.699 LB  (V6)\n7) Switching to bounding box crops rescaled to square images of size 224x224: 0.740 (0.748 in a preliminary test). (V7)\n8) Adding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.771 (0.780 in a preliminary test). (V9)\n9) DenseNet169 backbone: 0.800 (V11)\n10) DenseNet121 backbone: 0.805 (V13)</p>",
      "rawMarkdown": "Greetings,\nAfter many experiments and substantial refactoring the model, [my kernel][1] could reach **0.805 LB** score for **training within the kernel time limit** (no ensembling and external weights apart from ones pretrained on ImageNet) **with using a similarity distance based approach**. I just wanted to summarize milestones of the kernel score improvement (you can check details in the kernel), and I hope you will find them useful and could apply to your models. Also you can share your ides on the improvement of the score. The provided values of score show the result obtained after running the model within the kernel time limit, each new line indicates the modifications made to the model and the corresponding value of public LB.\n\n1) Triplet network with triplet loss (random selection of images): ~0.3 LB \n2) Batch all contrastive loss (compute loss based on all pairs in a batch, for 96 images per batch this method gives 9120 comparisons instead of 32 triplets): 0.5 LB\n3) Rectangular crops based on bounding boxes (without image distortion) instead of random ones: 0.55 LB\n4) Increase the dimensionality of the embedding space from 64 to 256 + ResNeXt50 instead of ResNet34: 0.606 LB (V1 and 2 of [this kernel][2])\n5) Metric learning `d^2 = (x1-x2).T*A*(x1-x2)` instead of Euclidean distance: 0.655 LB\n6) Optimized form of metric (check kernel for details) + hard negative example mining + lr optimization + going to smaller images of size 128x384 instead of 192x576 + averaging only nonzero contributions when loss is computed: 0.699 LB  (V6)\n7) Switching to bounding box crops rescaled to square images of size 224x224: 0.740 (0.748 in a preliminary test). (V7)\n8) Adding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.771 (0.780 in a preliminary test). (V9)\n9) DenseNet169 backbone: 0.800 (V11)\n10) DenseNet121 backbone: 0.805 (V13)\n\n  [1]: https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\n  [2]: https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit",
      "votes": null
    },
    {
      "id": "463873",
      "postDate": "01/30/2019 19:36:57",
      "content": "<p>just by curiosity, are you just doing a classification or a Siamese kind of arch?  With so few samples of each class, I wouldn't expect a classification to work, or does it?\nWhich is the simplest model that can score high doing a classification?</p>",
      "rawMarkdown": "just by curiosity, are you just doing a classification or a Siamese kind of arch?  With so few samples of each class, I wouldn't expect a classification to work, or does it?\nWhich is the simplest model that can score high doing a classification?",
      "votes": null
    },
    {
      "id": "463880",
      "postDate": "01/30/2019 20:23:44",
      "content": "<p>I use similarity based approach. Due to numerous classes presented by just several examples I didn't run classification approach for this competition. Though, if one just drop all classes that appear in the train less, let's say than 20 times. the classification could work to some extend, I expect. The thing with classification is that it is very simple to reach ~0.6 LB, and there are many public kernels on that. However, I don't know if image classification can give higher score. \nMeanwhile, similarity based approach requires much more elaborated work done on the code and development of the model. In particular, just a simplistic Siamese or Triplet network with random selection of pairs or triplets gives only ~0.3 LB. <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">Martin in his kernel</a> did quite a great job in utilizing LAP and Metric learning. Fine-tuning his model trained for 400 epochs could give <a href=\"https://www.kaggle.com/seesee/siamese-pretrained-0-822/notebook\">0.822 LB</a>.\nIn this post I tried to summarize what I did to be able to reach ~0.75 LB with a similarity based approach for training from scratch (ImageNet pretrained weights) within the kernel time limit.</p>",
      "rawMarkdown": "I use similarity based approach. Due to numerous classes presented by just several examples I didn't run classification approach for this competition. Though, if one just drop all classes that appear in the train less, let's say than 20 times. the classification could work to some extend, I expect. The thing with classification is that it is very simple to reach ~0.6 LB, and there are many public kernels on that. However, I don't know if image classification can give higher score. \nMeanwhile, similarity based approach requires much more elaborated work done on the code and development of the model. In particular, just a simplistic Siamese or Triplet network with random selection of pairs or triplets gives only ~0.3 LB. [Martin in his kernel][1] did quite a great job in utilizing LAP and Metric learning. Fine-tuning his model trained for 400 epochs could give [0.822 LB][2].\nIn this post I tried to summarize what I did to be able to reach ~0.75 LB with a similarity based approach for training from scratch (ImageNet pretrained weights) within the kernel time limit.\n\n\n  [1]: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n  [2]: https://www.kaggle.com/seesee/siamese-pretrained-0-822/notebook",
      "votes": null
    },
    {
      "id": "463977",
      "postDate": "01/31/2019 01:52:35",
      "content": "<p>Hi, how do you get bounding box info? Do you use martinpiotte's pretrained model from <a href=\"https://www.kaggle.com/martinpiotte/bounding-box-model\">https://www.kaggle.com/martinpiotte/bounding-box-model</a> ?</p>",
      "rawMarkdown": "Hi, how do you get bounding box info? Do you use martinpiotte's pretrained model from https://www.kaggle.com/martinpiotte/bounding-box-model ?",
      "votes": null
    },
    {
      "id": "463988",
      "postDate": "01/31/2019 02:32:27",
      "content": "<p>Yes, I used a fork based on this model <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a></p>",
      "rawMarkdown": "Yes, I used a fork based on this model https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes",
      "votes": null
    },
    {
      "id": "464101",
      "postDate": "01/31/2019 07:42:01",
      "content": "<p>Great work <a href=\"/iafoss\">@iafoss</a>! (as usual)</p>",
      "rawMarkdown": "Great work @iafoss! (as usual)",
      "votes": null
    },
    {
      "id": "464109",
      "postDate": "01/31/2019 08:09:59",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "464183",
      "postDate": "01/31/2019 10:12:07",
      "content": "<p>Great Work and thank you for your insight !</p>\n\n<p>Did you try some alternatives to the contrastive loss ?</p>\n\n<p>I think that you may reach a better score using some related work, for instance:\n<a href=\"http://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf\">http://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf</a></p>\n\n<p>There is also an interesting paper beeing discussed here:\n<a href=\"https://openreview.net/forum?id=HkxLXnAcFQ\">https://openreview.net/forum?id=HkxLXnAcFQ</a></p>\n\n<p>The changes required to test these ideas are minimal. I just started the competition, I did not submitt anything up to now. Using the ideas presented above, I could train a resnet-18 on 256x256 images that reach a local map@5 of 0.9+ for known whales. I do not know yet if this is good or not. The convergence of Multi-class N-pair Loss is very fast and very easy to implement. Basically all are variants of a softmax loss based on a distance measures between a subset of classes presented in a minibatch.</p>",
      "rawMarkdown": "Great Work and thank you for your insight !\n\nDid you try some alternatives to the contrastive loss ?\n\nI think that you may reach a better score using some related work, for instance:\nhttp://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf\n\nThere is also an interesting paper beeing discussed here:\nhttps://openreview.net/forum?id=HkxLXnAcFQ\n\nThe changes required to test these ideas are minimal. I just started the competition, I did not submitt anything up to now. Using the ideas presented above, I could train a resnet-18 on 256x256 images that reach a local map@5 of 0.9+ for known whales. I do not know yet if this is good or not. The convergence of Multi-class N-pair Loss is very fast and very easy to implement. Basically all are variants of a softmax loss based on a distance measures between a subset of classes presented in a minibatch.",
      "votes": null
    },
    {
      "id": "464198",
      "postDate": "01/31/2019 11:01:40",
      "content": "<p>Thanks for your reply, emm, I was also using martinpiotte's pretrained model, I got same results with  <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a>, but training with bbox gives worse results, there might be sth wrong in my code ....</p>",
      "rawMarkdown": "Thanks for your reply, emm, I was also using martinpiotte's pretrained model, I got same results with  https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes, but training with bbox gives worse results, there might be sth wrong in my code ....",
      "votes": null
    },
    {
      "id": "464206",
      "postDate": "01/31/2019 11:28:40",
      "content": "<p>I was able to get 0.774 public LB using the first paper mentioned above using 224x224 images (0.90 top 1 accuracy on known whales). It's pretty variable though - most of my models are bouncing around the 0.70-0.77 range at the moment. I'm going to try increasing the image size and see how that goes.</p>",
      "rawMarkdown": "I was able to get 0.774 public LB using the first paper mentioned above using 224x224 images (0.90 top 1 accuracy on known whales). It's pretty variable though - most of my models are bouncing around the 0.70-0.77 range at the moment. I'm going to try increasing the image size and see how that goes.",
      "votes": null
    },
    {
      "id": "464211",
      "postDate": "01/31/2019 11:35:12",
      "content": "<p>Thank you for sharing. I will continue to play around these ideas as well. </p>",
      "rawMarkdown": "Thank you for sharing. I will continue to play around these ideas as well.",
      "votes": null
    },
    {
      "id": "464383",
      "postDate": "01/31/2019 18:40:24",
      "content": "<p><a href=\"/jeandebleau\">@jeandebleau</a>, thank you for sharing the papers, I'll try to look more carefully into them. The things I also tried are based on batch hard triplet loss mentioned by <a href=\"/mnpinto\">@mnpinto</a>: <a href=\"https://arxiv.org/pdf/1901.03662.pdf\">https://arxiv.org/pdf/1901.03662.pdf</a> and <a href=\"https://arxiv.org/pdf/1703.07737.pdf\">https://arxiv.org/pdf/1703.07737.pdf</a> . Though, I couldn't get better results when I used the approach from these paper. Probably, such method may need more time to converge. Also the problem with triplet loss may be that despite the loss does amazing job in keeping images of the same kind together, it doesn't care how close from each other are the images of the same class. In other words it doesn't try to bring predictions for the same labels into one point in the embedding space. If there was no class <code>new_whale</code>, such method would work really the best. However, I assign <code>new_whale</code> label based on a fixed distance in the embedding space. For example, if there are no neighbors at a distance d0 for a particular image, <code>new_whale</code> is the first prediction for this image. The problem is that if the density of points with the same label in the embedding space is different, one my assign new_whale to points having low density. However, again, probably I just didn't train it for long enough.</p>",
      "rawMarkdown": "jeandebleau, thank you for sharing the papers, I'll try to look more carefully into them. The things I also tried are based on batch hard triplet loss mentioned by @mnpinto: https://arxiv.org/pdf/1901.03662.pdf and https://arxiv.org/pdf/1703.07737.pdf . Though, I couldn't get better results when I used the approach from these paper. Probably, such method may need more time to converge. Also the problem with triplet loss may be that despite the loss does amazing job in keeping images of the same kind together, it doesn't care how close from each other are the images of the same class. In other words it doesn't try to bring predictions for the same labels into one point in the embedding space. If there was no class `new_whale`, such method would work really the best. However, I assign `new_whale` label based on a fixed distance in the embedding space. For example, if there are no neighbors at a distance d0 for a particular image, `new_whale` is the first prediction for this image. The problem is that if the density of points with the same label in the embedding space is different, one my assign new_whale to points having low density. However, again, probably I just didn't train it for long enough.",
      "votes": null
    },
    {
      "id": "464401",
      "postDate": "01/31/2019 20:05:39",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> thank you for your kernel. Did you try to use FaceNet, VGG face models (they seem to work for triplet loss before)?</p>",
      "rawMarkdown": "iafoss thank you for your kernel. Did you try to use FaceNet, VGG face models (they seem to work for triplet loss before)?",
      "votes": null
    },
    {
      "id": "464418",
      "postDate": "01/31/2019 20:47:25",
      "content": "<p>@Blonde, I tried to use only models pretrained on ImageNet, though if there are models pretrained exclusively on face recognition of similar comparison based tasks it definitely worth a try. \nAlso, I would expect that larger models work better for this competition since when I looked to some hard triplets I couldn't do better than just random guessing... Though too large models are difficult to train given the GPU RAM limitation(((</p>",
      "rawMarkdown": "Blonde, I tried to use only models pretrained on ImageNet, though if there are models pretrained exclusively on face recognition of similar comparison based tasks it definitely worth a try. \nAlso, I would expect that larger models work better for this competition since when I looked to some hard triplets I couldn't do better than just random guessing... Though too large models are difficult to train given the GPU RAM limitation(((",
      "votes": null
    },
    {
      "id": "464451",
      "postDate": "01/31/2019 22:40:40",
      "content": "<p>yes, batch size it a limitation for triplet loss... did you try to start with even smaller images, 112x112 or 96x96 and large batch, could be better before moving to 224x224, although it depends on the network you choose</p>",
      "rawMarkdown": "yes, batch size it a limitation for triplet loss... did you try to start with even smaller images, 112x112 or 96x96 and large batch, could be better before moving to 224x224, although it depends on the network you choose",
      "votes": null
    },
    {
      "id": "464486",
      "postDate": "02/01/2019 00:48:35",
      "content": "<p>I didn't. For largest model I tried so far, ResNeXt50, I can run 96 224x224 images per batch with using half precision that gives ~10k compared pairs of images per batch with all batch approach. When I ran the code on 2 GPUs I can get ~40k comparisons that I expect to be enough at the initial stage of training. However for bigger models one may start from smaller images to have large enough batches.</p>",
      "rawMarkdown": "I didn't. For largest model I tried so far, ResNeXt50, I can run 96 224x224 images per batch with using half precision that gives ~10k compared pairs of images per batch with all batch approach. When I ran the code on 2 GPUs I can get ~40k comparisons that I expect to be enough at the initial stage of training. However for bigger models one may start from smaller images to have large enough batches.",
      "votes": null
    },
    {
      "id": "464741",
      "postDate": "02/01/2019 12:15:36",
      "content": "<p>Hey <a href=\"/iafoss\">@iafoss</a>, I'm implementing the triplet network like the two papers you mentioned and so far it seems to be working. I've got a 0.7 LB using a ResNet18, batch-hard mining without any augmentations on 256x256 pixel images. It seems it still have plenty of room for improvements, however I don't know if it could the 0.9x. </p>",
      "rawMarkdown": "Hey @iafoss, I'm implementing the triplet network like the two papers you mentioned and so far it seems to be working. I've got a 0.7 LB using a ResNet18, batch-hard mining without any augmentations on 256x256 pixel images. It seems it still have plenty of room for improvements, however I don't know if it could the 0.9x.",
      "votes": null
    },
    {
      "id": "465129",
      "postDate": "02/02/2019 11:57:39",
      "content": "<p>Thank you, I always learn a lot from you.</p>",
      "rawMarkdown": "Thank you, I always learn a lot from you.",
      "votes": null
    },
    {
      "id": "466452",
      "postDate": "02/05/2019 12:03:55",
      "content": "<p>Thank  <a href=\"/iafoss\">@iafoss</a> and <a href=\"/jeandebleau\">@jeandebleau</a> for suggested papers. I implemented all of them. But none of them is better than <code>Online Triplet Loss</code>. With Resnet18 backbone, no augmentations, 224x224 images, I could achieve 0.755 LB after 100 epochs (1 hour for training). I refer triplet loss, architecture in here: \n<a href=\"https://github.com/adambielski/siamese-triplet#online-triplet-selection\">https://github.com/adambielski/siamese-triplet#online-triplet-selection</a>. \nHope it is useful.</p>",
      "rawMarkdown": "Thank  @iafoss and @jeandebleau for suggested papers. I implemented all of them. But none of them is better than `Online Triplet Loss`. With Resnet18 backbone, no augmentations, 224x224 images, I could achieve 0.755 LB after 100 epochs (1 hour for training). I refer triplet loss, architecture in here: \nhttps://github.com/adambielski/siamese-triplet#online-triplet-selection. \nHope it is useful.",
      "votes": null
    },
    {
      "id": "466478",
      "postDate": "02/05/2019 13:09:58",
      "content": "<p><a href=\"/backaggle\">@backaggle</a> I assume you meant 1 hour per epoch</p>",
      "rawMarkdown": "backaggle I assume you meant 1 hour per epoch",
      "votes": null
    },
    {
      "id": "466483",
      "postDate": "02/05/2019 13:19:47",
      "content": "<p><a href=\"/valanm\">@valanm</a> I mean 1 hour for 100 epochs.</p>",
      "rawMarkdown": "valanm I mean 1 hour for 100 epochs.",
      "votes": null
    },
    {
      "id": "466498",
      "postDate": "02/05/2019 13:45:26",
      "content": "<p>thx, that's great. i will look into it. \nbtw, i reached 0.78 with <a href=\"/iafoss\">@iafoss</a> kernel (training within kernel time) - changed bs, augmentation and head</p>",
      "rawMarkdown": "thx, that's great. i will look into it. \nbtw, i reached 0.78 with @iafoss kernel (training within kernel time) - changed bs, augmentation and head",
      "votes": null
    },
    {
      "id": "466562",
      "postDate": "02/05/2019 16:19:54",
      "content": "<p><a href=\"/valanm\">@valanm</a> , Thank you for sharing your insights. It looks that shear augmentation  is quite helpful <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79384\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79384</a> . I'm curious, if you made the head more complicated or simplified to to one similar to Martin's kernel, where the head is just a pooling layer.</p>\n\n<p>Just an update on one more test I have performed. I tried a network predicting score based on x1 - x2 and x1*x2 features and trained with Focal loss calculated in batch all manner, but the result, ~0.72, is worse than my public kernel.</p>",
      "rawMarkdown": "valanm , Thank you for sharing your insights. It looks that shear augmentation  is quite helpful https://www.kaggle.com/c/humpback-whale-identification/discussion/79384 . I'm curious, if you made the head more complicated or simplified to to one similar to Martin's kernel, where the head is just a pooling layer.\n\nJust an update on one more test I have performed. I tried a network predicting score based on x1 - x2 and x1*x2 features and trained with Focal loss calculated in batch all manner, but the result, ~0.72, is worse than my public kernel.",
      "votes": null
    },
    {
      "id": "467195",
      "postDate": "02/06/2019 16:04:37",
      "content": "<p>One more thing that boosted the score:\nAdding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.740 -&gt; 0.771 (0.780 in a preliminary test).</p>",
      "rawMarkdown": "One more thing that boosted the score:\nAdding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.740 -&gt; 0.771 (0.780 in a preliminary test).",
      "votes": null
    },
    {
      "id": "467324",
      "postDate": "02/06/2019 21:55:22",
      "content": "<p>you may want to consider this little trick i commented here:</p>\n\n<p><a href=\"https://www.kaggle.com/seesee/siamese-pretrained-0-822\">https://www.kaggle.com/seesee/siamese-pretrained-0-822</a></p>\n\n<hr>\n\n<p>\"... i visually inspect your results. Many of the new whales (i.e. top1=new__whale) are not really really new whales, and the true id is actually top2 or 3 prediction.</p>\n\n<p>Since we know that there are about 27% new-whale you can do a reassignment by choosing only the top most likely new-whale (i.e threshold base on rank instead of numeric threshold, etc.)</p>\n\n<p>A rough calculation gives improvement of +0.04 if the new__whale are replaced correctly by rank2 or rank3 prediction.\"</p>",
      "rawMarkdown": "you may want to consider this little trick i commented here:\n\nhttps://www.kaggle.com/seesee/siamese-pretrained-0-822\n\n---\n\n\"... i visually inspect your results. Many of the new whales (i.e. top1=new__whale) are not really really new whales, and the true id is actually top2 or 3 prediction.\n\nSince we know that there are about 27% new-whale you can do a reassignment by choosing only the top most likely new-whale (i.e threshold base on rank instead of numeric threshold, etc.)\n\nA rough calculation gives improvement of +0.04 if the new__whale are replaced correctly by rank2 or rank3 prediction.\"",
      "votes": null
    },
    {
      "id": "467373",
      "postDate": "02/07/2019 01:11:14",
      "content": "<p>Hi, Heng. So how do we judge that top1=new_whale is wrong?</p>",
      "rawMarkdown": "Hi, Heng. So how do we judge that top1=new_whale is wrong?",
      "votes": null
    },
    {
      "id": "467671",
      "postDate": "02/07/2019 14:43:13",
      "content": "<p>head is similar to your kernel (tiny changes - slightly more complicated) :)</p>",
      "rawMarkdown": "head is similar to your kernel (tiny changes - slightly more complicated) :)",
      "votes": null
    },
    {
      "id": "468088",
      "postDate": "02/08/2019 09:19:28",
      "content": "<p>I found a useful repo here: <a href=\"https://github.com/bnulihaixia/Deep_metric\">https://github.com/bnulihaixia/Deep_metric</a> . \nThe <code>WeightLoss</code> helped me to reach <code>0.785</code> LB by resnet18, 224x224, no aug which is an improvement compared to <code>Online TripletLoss</code> (0.755)</p>",
      "rawMarkdown": "I found a useful repo here: https://github.com/bnulihaixia/Deep_metric . \nThe `WeightLoss` helped me to reach `0.785` LB by resnet18, 224x224, no aug which is an improvement compared to `Online TripletLoss` (0.755)",
      "votes": null
    },
    {
      "id": "468169",
      "postDate": "02/08/2019 12:24:15",
      "content": "<p>@Iafoss Great work. Thank you. This work help me a lot.</p>",
      "rawMarkdown": "Iafoss Great work. Thank you. This work help me a lot.",
      "votes": null
    },
    {
      "id": "470836",
      "postDate": "02/13/2019 16:32:37",
      "content": "<p>After starting using DenseNet169 backbone I got an improvement from 0.771 to 0.800.</p>",
      "rawMarkdown": "After starting using DenseNet169 backbone I got an improvement from 0.771 to 0.800.",
      "votes": null
    },
    {
      "id": "471865",
      "postDate": "02/15/2019 02:16:45",
      "content": "<p>Great work! After changing image size from 244 to 512, I could get around 0.834 LB. I'll switch to SGD with proper learning rate schedule and train it for a longer time to see where it can reach.</p>",
      "rawMarkdown": "Great work! After changing image size from 244 to 512, I could get around 0.834 LB. I'll switch to SGD with proper learning rate schedule and train it for a longer time to see where it can reach.",
      "votes": null
    },
    {
      "id": "471896",
      "postDate": "02/15/2019 03:54:20",
      "content": "<p>Thanks for checking it, looking forward to hearing from you.\nAnother thing that helped to get 0.814 within the kernel time limit is adding batch norm after flattening and increasing the size of the intermediate head layer and embedding dimensionality to 512.</p>",
      "rawMarkdown": "Thanks for checking it, looking forward to hearing from you.\nAnother thing that helped to get 0.814 within the kernel time limit is adding batch norm after flattening and increasing the size of the intermediate head layer and embedding dimensionality to 512.",
      "votes": null
    },
    {
      "id": "472576",
      "postDate": "02/16/2019 07:40:46",
      "content": "<p>I just change sz=512, but get LB about 0.5. Are there any other value that I should change? Thanks!</p>",
      "rawMarkdown": "I just change sz=512, but get LB about 0.5. Are there any other value that I should change? Thanks!",
      "votes": null
    },
    {
      "id": "472739",
      "postDate": "02/16/2019 15:11:47",
      "content": "<p><a href=\"/dilapsky\">@dilapsky</a> yes you have to adjust the learning rate with the new batch size and new image size \nlearner.lr_find()</p>",
      "rawMarkdown": "dilapsky yes you have to adjust the learning rate with the new batch size and new image size \nlearner.lr_find()",
      "votes": null
    },
    {
      "id": "472740",
      "postDate": "02/16/2019 15:12:33",
      "content": "<p>check <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/80624\">https://www.kaggle.com/c/humpback-whale-identification/discussion/80624</a></p>",
      "rawMarkdown": "check https://www.kaggle.com/c/humpback-whale-identification/discussion/80624",
      "votes": null
    },
    {
      "id": "472758",
      "postDate": "02/16/2019 16:09:09",
      "content": "<p><a href=\"/dilapsky\">@dilapsky</a> , It also may be a problem with just too small batch itself, resulted for example by batch normalization. In protein competition <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a> some people used gradient accumulation to handle it. Also one can freeze bn, but if I remember correctly, it doesn't work in fast.ai 0.7 with half precision. Not sure if adjusting parameters of bn, such as momentum, can help. Another thing, training for some certain number of epochs is needed, in the very beginning the the result may drop a little bit.</p>",
      "rawMarkdown": "dilapsky , It also may be a problem with just too small batch itself, resulted for example by batch normalization. In protein competition https://www.kaggle.com/c/human-protein-atlas-image-classification some people used gradient accumulation to handle it. Also one can freeze bn, but if I remember correctly, it doesn't work in fast.ai 0.7 with half precision. Not sure if adjusting parameters of bn, such as momentum, can help. Another thing, training for some certain number of epochs is needed, in the very beginning the the result may drop a little bit.",
      "votes": null
    },
    {
      "id": "472913",
      "postDate": "02/16/2019 22:05:01",
      "content": "<p>I have also observed variability in the results from run to run. I had to re-run the same notebook and got ~0.02 less in LB, then to be sure that I haven't change anything, I ran the notebook again, and got the same like before (better by ~0.02). </p>\n\n<p>Perhaps this is due to the random choice of  training images that lead to slightly better/worse training depending on pure chance of image pickup, which is telling how much important is the image selection  in similarity based NN.</p>",
      "rawMarkdown": "I have also observed variability in the results from run to run. I had to re-run the same notebook and got ~0.02 less in LB, then to be sure that I haven't change anything, I ran the notebook again, and got the same like before (better by ~0.02). \n\nPerhaps this is due to the random choice of  training images that lead to slightly better/worse training depending on pure chance of image pickup, which is telling how much important is the image selection  in similarity based NN.",
      "votes": null
    },
    {
      "id": "472922",
      "postDate": "02/16/2019 22:32:59",
      "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a>\nThere's a paper from 2018 that  supports your approach to tweaking LRs when decreasing/increasing batchsize:</p>\n\n<p><a href=\"https://arxiv.org/abs/1706.02677\">https://arxiv.org/abs/1706.02677</a></p>",
      "rawMarkdown": "msmelguizo\nThere's a paper from 2018 that  supports your approach to tweaking LRs when decreasing/increasing batchsize:\n\nhttps://arxiv.org/abs/1706.02677",
      "votes": null
    },
    {
      "id": "472971",
      "postDate": "02/17/2019 02:34:00",
      "content": "<p>Hey <a href=\"/dilapsky\">@dilapsky</a>, if you change nothing just set img size to 512 with batch size 16, the score would be around 0.83. And @Iafoss, I've tried larger embedding dimensionality 512 with img size 512 and noticed that the val loss becomes quite unstable. After training for a longer time, val loss becomes inf. Still trying to figure out why. And congratulations on your huge leap in LB score! Did you get this score based on your kernel?</p>",
      "rawMarkdown": "Hey @dilapsky, if you change nothing just set img size to 512 with batch size 16, the score would be around 0.83. And @Iafoss, I've tried larger embedding dimensionality 512 with img size 512 and noticed that the val loss becomes quite unstable. After training for a longer time, val loss becomes inf. Still trying to figure out why. And congratulations on your huge leap in LB score! Did you get this score based on your kernel?",
      "votes": null
    },
    {
      "id": "472992",
      "postDate": "02/17/2019 03:40:30",
      "content": "<p>Thanks, it is based on ensembling and cumulative efforts of all teammates. Regarding unstable val, I had something like u (val loss was fluctuation and increasing while competition score got better) before I added compactification term to the loss. Probably, large  embedding dimensionality exaggerates this effect. </p>",
      "rawMarkdown": "Thanks, it is based on ensembling and cumulative efforts of all teammates. Regarding unstable val, I had something like u (val loss was fluctuation and increasing while competition score got better) before I added compactification term to the loss. Probably, large  embedding dimensionality exaggerates this effect.",
      "votes": null
    },
    {
      "id": "473154",
      "postDate": "02/17/2019 12:36:17",
      "content": "<p><a href=\"/iafoss\">@iafoss</a>, Can you please provide more details on how to handle the batch size using gradient accumulation? I  don't mind if I have to train full precision. I have both a 1080Ti and a 2080 card (two different computers) and it is painful to watch how bad the results of the 2080 card are. Thanks again for your great explanations.</p>",
      "rawMarkdown": "iafoss, Can you please provide more details on how to handle the batch size using gradient accumulation? I  don't mind if I have to train full precision. I have both a 1080Ti and a 2080 card (two different computers) and it is painful to watch how bad the results of the 2080 card are. Thanks again for your great explanations.",
      "votes": null
    },
    {
      "id": "473312",
      "postDate": "02/17/2019 19:02:53",
      "content": "<p>I didn't try it by my own yet, but gradient accumulation was quite effective for protein competition. Our team didn't go images larger than 512x512 in that competition mainly because after decreasing bs below 16 model performance was dropping. Meanwhile, many people who used larger image resolution reported using gradient accumulation.\nThe idea of this method is updating weights only at each n-th step with using accumulated gradients for all n steps. This feature is missing in both fast.ai 0.7 and 1.0. For v1.0, which I'm less familiar with, you can check the following discussion <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219\">https://forums.fast.ai/t/accumulating-gradients/33219</a> . In fast.ai 0.7 you just need to slightly modify Stepper class making sure that you call self.m.zero_grad() at the beginning of n-th step and self.opt.step() at the end of n-1-th step. I'm not 100% sure that some other adjustment are not needed since I didn't do it yet. Another thing one may try is freezing bn, but in this case half prescription doesn't work because of a bug in fast.ai library, if I remember correctly. Another thing people used to handle small batches is replacement of bn by instance normalization. I'm not sure how effective is that since I never used it. </p>",
      "rawMarkdown": "I didn't try it by my own yet, but gradient accumulation was quite effective for protein competition. Our team didn't go images larger than 512x512 in that competition mainly because after decreasing bs below 16 model performance was dropping. Meanwhile, many people who used larger image resolution reported using gradient accumulation.\nThe idea of this method is updating weights only at each n-th step with using accumulated gradients for all n steps. This feature is missing in both fast.ai 0.7 and 1.0. For v1.0, which I'm less familiar with, you can check the following discussion https://forums.fast.ai/t/accumulating-gradients/33219 . In fast.ai 0.7 you just need to slightly modify Stepper class making sure that you call self.m.zero_grad() at the beginning of n-th step and self.opt.step() at the end of n-1-th step. I'm not 100% sure that some other adjustment are not needed since I didn't do it yet. Another thing one may try is freezing bn, but in this case half prescription doesn't work because of a bug in fast.ai library, if I remember correctly. Another thing people used to handle small batches is replacement of bn by instance normalization. I'm not sure how effective is that since I never used it.",
      "votes": null
    },
    {
      "id": "473360",
      "postDate": "02/17/2019 21:08:26",
      "content": "<p>I tried to use accumulation in v1 for the protein challenge following that thread. I managed to get something that did not crash before training, but it seemed to behave weirdly, maybe because I was also using fp16. I plan to look deeper into it when I get a chance, a solid method for gradient accumulation in fastai v1 would be good for everyone. Code is below (if you look closely you may notice some it is Iafoss's code)</p>\n\n<pre><code>from fastai.callbacks import *\npath='.'\n\nclass myOptimWrapper(OptimWrapper):\nn = 8\nistep, izero_grad = 1, 1\ncnt = 0\n\ndef step(self):  \n    if self.istep == self.n :\n        super().step()\n        self.cnt += 1\n        self.istep = 1\n    else :\n        self.istep += 1\n\ndef zero_grad(self):      \n    if self.izero_grad == self.n :\n        super().zero_grad()\n        self.izero_grad = 1\n    else :\n        self.izero_grad += 1\n\n<a href=\"/dataclass\">@dataclass</a>\nclass StepEpochEnd(Callback):\n    learn:Learner\n    def on_epoch_end(self, **kwargs):\n        print(\"real step and zero grad\")\n        self.learn.opt.real_step()\n        self.learn.opt.real_zero_grad()\n\ndef my_create_opt(self, lr:Floats, wd:Floats=0.)-&gt;None:\n    \"Create optimizer with `lr` learning rate and `wd` weight decay.\"\n    self.opt = myOptimWrapper.create(self.opt_func, lr, self.layer_groups,\n                                     wd=wd, true_wd=self.true_wd, bn_wd=self.bn_wd)\n\nLearner.create_opt = my_create_opt\n\nlearn = create_cnn(\ndata,\nresnet50,\ncut=-2,\nsplit_on= _resnet_split,\nloss_func=F.binary_cross_entropy_with_logits, #FocalLoss(logits=True,alpha=alpha_log), \npath=path,    \nmetrics=[f1_score, f1_callback.f1],\ncallback_fns=[partial(GradientClipping, clip=1),\n              partial(EarlyStoppingCallback, monitor='val_loss', min_delta=0.01, patience=6),\n              #partial(StepEpochEnd),\n              ],\n)\n</code></pre>",
      "rawMarkdown": "I tried to use accumulation in v1 for the protein challenge following that thread. I managed to get something that did not crash before training, but it seemed to behave weirdly, maybe because I was also using fp16. I plan to look deeper into it when I get a chance, a solid method for gradient accumulation in fastai v1 would be good for everyone. Code is below (if you look closely you may notice some it is Iafoss's code)\n\n    from fastai.callbacks import *\n    path='.'\n\n    class myOptimWrapper(OptimWrapper):\n    n = 8\n    istep, izero_grad = 1, 1\n    cnt = 0\n\n    def step(self):  \n        if self.istep == self.n :\n            super().step()\n            self.cnt += 1\n            self.istep = 1\n        else :\n            self.istep += 1\n\n    def zero_grad(self):      \n        if self.izero_grad == self.n :\n            super().zero_grad()\n            self.izero_grad = 1\n        else :\n            self.izero_grad += 1\n\n    @dataclass\n    class StepEpochEnd(Callback):\n        learn:Learner\n        def on_epoch_end(self, **kwargs):\n            print(\"real step and zero grad\")\n            self.learn.opt.real_step()\n            self.learn.opt.real_zero_grad()\n\n    def my_create_opt(self, lr:Floats, wd:Floats=0.)-&gt;None:\n        \"Create optimizer with `lr` learning rate and `wd` weight decay.\"\n        self.opt = myOptimWrapper.create(self.opt_func, lr, self.layer_groups,\n                                         wd=wd, true_wd=self.true_wd, bn_wd=self.bn_wd)\n\n    Learner.create_opt = my_create_opt\n\n    learn = create_cnn(\n    data,\n    resnet50,\n    cut=-2,\n    split_on= _resnet_split,\n    loss_func=F.binary_cross_entropy_with_logits, #FocalLoss(logits=True,alpha=alpha_log), \n    path=path,    \n    metrics=[f1_score, f1_callback.f1],\n    callback_fns=[partial(GradientClipping, clip=1),\n                  partial(EarlyStoppingCallback, monitor='val_loss', min_delta=0.01, patience=6),\n                  #partial(StepEpochEnd),\n                  ],\n    )",
      "votes": null
    },
    {
      "id": "473412",
      "postDate": "02/18/2019 00:08:45",
      "content": "<p>I'm not 100% sure how it is done in fast.ai v1, but I think zero_grad is called before step. In fast.ai 0.7 Stepper class, if I remember correctly, first gradients are set to zero, then forward and backward passes are calculated, followed by weight update. So what may be going on in the above code, at n-th step gradient is set to zero followed by calculation of gradient for n-th step and using this value (only computed for one step) for weight update, others are lost. Correct me, if it is not the case. Also, for FP16 in the code there may be an additional condition, I do not remember right now. </p>",
      "rawMarkdown": "I'm not 100% sure how it is done in fast.ai v1, but I think zero_grad is called before step. In fast.ai 0.7 Stepper class, if I remember correctly, first gradients are set to zero, then forward and backward passes are calculated, followed by weight update. So what may be going on in the above code, at n-th step gradient is set to zero followed by calculation of gradient for n-th step and using this value (only computed for one step) for weight update, others are lost. Correct me, if it is not the case. Also, for FP16 in the code there may be an additional condition, I do not remember right now.",
      "votes": null
    },
    {
      "id": "473680",
      "postDate": "02/18/2019 11:08:29",
      "content": "<p>I am very new to pytorch and fastai, Thanks for all of you! BTW, I want to fine-tunning  from 224 to 360, to 512 if I can.So I need to reload the previous model. However, I edit the code just by adding learner.load(‘model’) before training, but nothing happened. I mean, loading model can execute but no improvement in result. Are there any errors I trapped into? I only add that one line, and fastai doc indicate that this is just loading weight. Thanks!</p>",
      "rawMarkdown": "I am very new to pytorch and fastai, Thanks for all of you! BTW, I want to fine-tunning  from 224 to 360, to 512 if I can.So I need to reload the previous model. However, I edit the code just by adding learner.load(‘model’) before training, but nothing happened. I mean, loading model can execute but no improvement in result. Are there any errors I trapped into? I only add that one line, and fastai doc indicate that this is just loading weight. Thanks!",
      "votes": null
    },
    {
      "id": "473728",
      "postDate": "02/18/2019 12:39:12",
      "content": "<p>@Iafoss why did you choose DenseNet in particular and not other architecture?</p>\n\n<p>There are a few other arch that can get better accuracies on imagenet :\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model</a></p>",
      "rawMarkdown": "Iafoss why did you choose DenseNet in particular and not other architecture?\n\nThere are a few other arch that can get better accuracies on imagenet :\nhttps://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model",
      "votes": null
    },
    {
      "id": "473873",
      "postDate": "02/18/2019 16:05:26",
      "content": "<p><a href=\"/dilapsky\">@dilapsky</a> , At the later stage training is quite slow and it is likely that you faced with, but there can be several other things. Usually image resolution is chosen to be multiple to 32, i.e. 224, 384, and 512. If you do it at kaggle you need to set an appropriate path first.</p>",
      "rawMarkdown": "dilapsky , At the later stage training is quite slow and it is likely that you faced with, but there can be several other things. Usually image resolution is chosen to be multiple to 32, i.e. 224, 384, and 512. If you do it at kaggle you need to set an appropriate path first.",
      "votes": null
    },
    {
      "id": "473889",
      "postDate": "02/18/2019 16:36:57",
      "content": "<p><a href=\"/hwasiti\">@hwasiti</a> , Actually, you are quite limited with the chose of the model given limitations of GPU RAM and large image resolution. If you do not have several GPU with 16+ GB of RAM you are likely should not looking for using big models. The rest are ResNet18, ResNet34, ResNeXt50(SE), DenseNet121, DenseNet169. You also may consider ResNet101(SE), DenseNet201, Inception, and NasNet if you have several GPUs or willing to invest time to fight with small batches. I tried ResNet34 and ResNeXt50SE, but there performance was worse than ResNeXt50. SE I expect is more difficult to retrain, and I never used it for the production models so far. The same with NasNet, I tried it several times but the results were poor. I guess it may be also quite difficult to retrain or just architecture itself is overfitted to ImageNet, and doesn't perform well on other image sets. And I also usually ignore ResNet50+ since they have about the same memory requirement as corresponding ResNeXt, but worse performance.</p>",
      "rawMarkdown": "hwasiti , Actually, you are quite limited with the chose of the model given limitations of GPU RAM and large image resolution. If you do not have several GPU with 16+ GB of RAM you are likely should not looking for using big models. The rest are ResNet18, ResNet34, ResNeXt50(SE), DenseNet121, DenseNet169. You also may consider ResNet101(SE), DenseNet201, Inception, and NasNet if you have several GPUs or willing to invest time to fight with small batches. I tried ResNet34 and ResNeXt50SE, but there performance was worse than ResNeXt50. SE I expect is more difficult to retrain, and I never used it for the production models so far. The same with NasNet, I tried it several times but the results were poor. I guess it may be also quite difficult to retrain or just architecture itself is overfitted to ImageNet, and doesn't perform well on other image sets. And I also usually ignore ResNet50+ since they have about the same memory requirement as corresponding ResNeXt, but worse performance.",
      "votes": null
    },
    {
      "id": "474108",
      "postDate": "02/19/2019 00:20:45",
      "content": "<p>That's a nice observation @lafoss. \nI've always went straight for SEResNeXt50 instead of ResNeXt50. I'll try the non-SE counterpart to see how it goes. As for the more \"designed\" architectures like NasNet I also think ti is made to overfit to ImageNet.</p>",
      "rawMarkdown": "That's a nice observation @lafoss. \nI've always went straight for SEResNeXt50 instead of ResNeXt50. I'll try the non-SE counterpart to see how it goes. As for the more \"designed\" architectures like NasNet I also think ti is made to overfit to ImageNet.",
      "votes": null
    },
    {
      "id": "474124",
      "postDate": "02/19/2019 01:02:28",
      "content": "<p>Given the uncertainty on how to do gradient accumulation properly, I asked <a href=\"/sgugger\">@sgugger</a> on the fastai forum whether they can include it in fastai v1. I hope @kcturgutlu method works well and all fastai user will be able to use it if he can do a PR. \nPlease support this request on that <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/24?u=hwasiti\">forum thread</a> if you are interested on integrating such feature into fastai.</p>\n\n<p>I think this is something that almost always has a utility with any image recognition problem. \nWho doesn't want to try checking whether increasing image size will give better accuracy? Our new limit will be the max image size that a GPU memory can handle with only 1 batch size.</p>",
      "rawMarkdown": "Given the uncertainty on how to do gradient accumulation properly, I asked @sgugger on the fastai forum whether they can include it in fastai v1. I hope @kcturgutlu method works well and all fastai user will be able to use it if he can do a PR. \nPlease support this request on that [forum thread][1] if you are interested on integrating such feature into fastai.\n\nI think this is something that almost always has a utility with any image recognition problem. \nWho doesn't want to try checking whether increasing image size will give better accuracy? Our new limit will be the max image size that a GPU memory can handle with only 1 batch size.\n\n\n  [1]: https://forums.fast.ai/t/accumulating-gradients/33219/24?u=hwasiti",
      "votes": null
    },
    {
      "id": "474323",
      "postDate": "02/19/2019 08:16:31",
      "content": "<p>Sounds crazy but as a matter of fact one can also use &lt;1 image per batch</p>",
      "rawMarkdown": "Sounds crazy but as a matter of fact one can also use &lt;1 image per batch",
      "votes": null
    },
    {
      "id": "474355",
      "postDate": "02/19/2019 08:58:58",
      "content": "<p>Seems Batch Normalization shouldn't be the same, when using gradient accumulation. That's why people are getting lower accuracy with grad. accum. \nSee:\n<a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti\">https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti</a></p>\n\n<p><a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti\">https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti</a></p>\n\n<p>and maybe that means we have to fiddle with the arch too..</p>",
      "rawMarkdown": "Seems Batch Normalization shouldn't be the same, when using gradient accumulation. That's why people are getting lower accuracy with grad. accum. \nSee:\n[https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti][1]\n\n[https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti][2]\n\nand maybe that means we have to fiddle with the arch too..\n\n  [1]: https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti\n  [2]: https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti",
      "votes": null
    },
    {
      "id": "474368",
      "postDate": "02/19/2019 09:17:28",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a>, were you able to run Lafoss's script with image size 512 and batch size 16 on Kaggle kernel? I keep getting out of memory error even with img sz 320 ans bsz 16. I forked version-11 of the kernel.</p>\n\n<p><a href=\"/iafoss\">@iafoss</a> thanks so much for all the sharing. For the lates version-13 of your kernel, I keep getting memory error when I run it wthout any changes. Do you know that happens? Any advise?</p>",
      "rawMarkdown": "syoya1997, were you able to run Lafoss's script with image size 512 and batch size 16 on Kaggle kernel? I keep getting out of memory error even with img sz 320 ans bsz 16. I forked version-11 of the kernel.\n\n@iafoss thanks so much for all the sharing. For the lates version-13 of your kernel, I keep getting memory error when I run it wthout any changes. Do you know that happens? Any advise?",
      "votes": null
    },
    {
      "id": "474400",
      "postDate": "02/19/2019 09:50:09",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, I'm not sure whether it would work on kaggle kernl but it's ok on my local GTX 1080Tis with img size 512 and batch size 16. </p>",
      "rawMarkdown": "sheriytm, I'm not sure whether it would work on kaggle kernl but it's ok on my local GTX 1080Tis with img size 512 and batch size 16.",
      "votes": null
    },
    {
      "id": "474503",
      "postDate": "02/19/2019 13:13:11",
      "content": "<p>Perhaps you have 2 GPUs. I have 2 GPUs 1080Ti and that is the only way to run bs 16 with SZ 512. \nand btw, the kernel will run on all your gpus by default.</p>",
      "rawMarkdown": "Perhaps you have 2 GPUs. I have 2 GPUs 1080Ti and that is the only way to run bs 16 with SZ 512. \nand btw, the kernel will run on all your gpus by default.",
      "votes": null
    },
    {
      "id": "474505",
      "postDate": "02/19/2019 13:21:52",
      "content": "<p>Yes. I use 2 GPUs 1080Ti to run with img sz 512 and bs 16. And may I ask what's your best score with different img size. I can just reach 0.847 with 512 and 0.820~ with 224.</p>",
      "rawMarkdown": "Yes. I use 2 GPUs 1080Ti to run with img sz 512 and bs 16. And may I ask what's your best score with different img size. I can just reach 0.847 with 512 and 0.820~ with 224.",
      "votes": null
    },
    {
      "id": "474520",
      "postDate": "02/19/2019 13:52:58",
      "content": "<p>The 2 GPU run haven't finished yet. The 512 bs 8 =&gt; LB 0.821. This is my best score so far with fastai.  I have better scores but those from the keras model.</p>",
      "rawMarkdown": "The 2 GPU run haven't finished yet. The 512 bs 8 =&gt; LB 0.821. This is my best score so far with fastai.  I have better scores but those from the keras model.",
      "votes": null
    },
    {
      "id": "474625",
      "postDate": "02/19/2019 16:34:41",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a> , it's strange, I ran several similar kernels at kaggle without memory errors. It should fit in 12 GB. If you are running it on your own computer, may it be something else occupying part of GPU memory.</p>",
      "rawMarkdown": "sheriytm , it's strange, I ran several similar kernels at kaggle without memory errors. It should fit in 12 GB. If you are running it on your own computer, may it be something else occupying part of GPU memory.",
      "votes": null
    },
    {
      "id": "474772",
      "postDate": "02/19/2019 19:26:58",
      "content": "<p><a href=\"/iafoss\">@iafoss</a>, I am running it on Kaggle but kept getting the memory error. I wanted to try it out on the Kaggle kernel first. I really don't know why I cannot run it successfully with no change.</p>",
      "rawMarkdown": "iafoss, I am running it on Kaggle but kept getting the memory error. I wanted to try it out on the Kaggle kernel first. I really don't know why I cannot run it successfully with no change.",
      "votes": null
    },
    {
      "id": "474980",
      "postDate": "02/20/2019 04:50:20",
      "content": "<p>Could you share the link of the public keras kernel you're using? I want to see whether there's something more I can do to improve the score. Thanks :)</p>",
      "rawMarkdown": "Could you share the link of the public keras kernel you're using? I want to see whether there's something more I can do to improve the score. Thanks :)",
      "votes": null
    },
    {
      "id": "475073",
      "postDate": "02/20/2019 07:54:21",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a>, I am using <a href=\"/iafoss\">@iafoss</a>'s kernel <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit\">here</a> ver-11 and ver-13 as seperate kernels. I am able to run ver-11 with different parameters and got lower result (0.785) but ver-13 is the one giving memory error when run without any change.</p>",
      "rawMarkdown": "syoya1997, I am using @iafoss's kernel [here][1] ver-11 and ver-13 as seperate kernels. I am able to run ver-11 with different parameters and got lower result (0.785) but ver-13 is the one giving memory error when run without any change.\n\n\n  [1]: https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 463873,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "01/30/2019 19:36:57",
      "content": "<p>just by curiosity, are you just doing a classification or a Siamese kind of arch?  With so few samples of each class, I wouldn't expect a classification to work, or does it?\nWhich is the simplest model that can score high doing a classification?</p>",
      "votes": null,
      "replies": [
        {
          "id": 463880,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/30/2019 20:23:44",
          "content": "<p>I use similarity based approach. Due to numerous classes presented by just several examples I didn't run classification approach for this competition. Though, if one just drop all classes that appear in the train less, let's say than 20 times. the classification could work to some extend, I expect. The thing with classification is that it is very simple to reach ~0.6 LB, and there are many public kernels on that. However, I don't know if image classification can give higher score. \nMeanwhile, similarity based approach requires much more elaborated work done on the code and development of the model. In particular, just a simplistic Siamese or Triplet network with random selection of pairs or triplets gives only ~0.3 LB. <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">Martin in his kernel</a> did quite a great job in utilizing LAP and Metric learning. Fine-tuning his model trained for 400 epochs could give <a href=\"https://www.kaggle.com/seesee/siamese-pretrained-0-822/notebook\">0.822 LB</a>.\nIn this post I tried to summarize what I did to be able to reach ~0.75 LB with a similarity based approach for training from scratch (ImageNet pretrained weights) within the kernel time limit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 463977,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "01/31/2019 01:52:35",
      "content": "<p>Hi, how do you get bounding box info? Do you use martinpiotte's pretrained model from <a href=\"https://www.kaggle.com/martinpiotte/bounding-box-model\">https://www.kaggle.com/martinpiotte/bounding-box-model</a> ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 463988,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/31/2019 02:32:27",
          "content": "<p>Yes, I used a fork based on this model <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464198,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "01/31/2019 11:01:40",
          "content": "<p>Thanks for your reply, emm, I was also using martinpiotte's pretrained model, I got same results with  <a href=\"https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes\">https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes</a>, but training with bbox gives worse results, there might be sth wrong in my code ....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 464101,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "01/31/2019 07:42:01",
      "content": "<p>Great work <a href=\"/iafoss\">@iafoss</a>! (as usual)</p>",
      "votes": null,
      "replies": [
        {
          "id": 464109,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/31/2019 08:09:59",
          "content": "<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 464183,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "01/31/2019 10:12:07",
      "content": "<p>Great Work and thank you for your insight !</p>\n\n<p>Did you try some alternatives to the contrastive loss ?</p>\n\n<p>I think that you may reach a better score using some related work, for instance:\n<a href=\"http://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf\">http://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf</a></p>\n\n<p>There is also an interesting paper beeing discussed here:\n<a href=\"https://openreview.net/forum?id=HkxLXnAcFQ\">https://openreview.net/forum?id=HkxLXnAcFQ</a></p>\n\n<p>The changes required to test these ideas are minimal. I just started the competition, I did not submitt anything up to now. Using the ideas presented above, I could train a resnet-18 on 256x256 images that reach a local map@5 of 0.9+ for known whales. I do not know yet if this is good or not. The convergence of Multi-class N-pair Loss is very fast and very easy to implement. Basically all are variants of a softmax loss based on a distance measures between a subset of classes presented in a minibatch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 464206,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "01/31/2019 11:28:40",
          "content": "<p>I was able to get 0.774 public LB using the first paper mentioned above using 224x224 images (0.90 top 1 accuracy on known whales). It's pretty variable though - most of my models are bouncing around the 0.70-0.77 range at the moment. I'm going to try increasing the image size and see how that goes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464211,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "01/31/2019 11:35:12",
          "content": "<p>Thank you for sharing. I will continue to play around these ideas as well. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464383,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/31/2019 18:40:24",
          "content": "<p><a href=\"/jeandebleau\">@jeandebleau</a>, thank you for sharing the papers, I'll try to look more carefully into them. The things I also tried are based on batch hard triplet loss mentioned by <a href=\"/mnpinto\">@mnpinto</a>: <a href=\"https://arxiv.org/pdf/1901.03662.pdf\">https://arxiv.org/pdf/1901.03662.pdf</a> and <a href=\"https://arxiv.org/pdf/1703.07737.pdf\">https://arxiv.org/pdf/1703.07737.pdf</a> . Though, I couldn't get better results when I used the approach from these paper. Probably, such method may need more time to converge. Also the problem with triplet loss may be that despite the loss does amazing job in keeping images of the same kind together, it doesn't care how close from each other are the images of the same class. In other words it doesn't try to bring predictions for the same labels into one point in the embedding space. If there was no class <code>new_whale</code>, such method would work really the best. However, I assign <code>new_whale</code> label based on a fixed distance in the embedding space. For example, if there are no neighbors at a distance d0 for a particular image, <code>new_whale</code> is the first prediction for this image. The problem is that if the density of points with the same label in the embedding space is different, one my assign new_whale to points having low density. However, again, probably I just didn't train it for long enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464401,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "01/31/2019 20:05:39",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> thank you for your kernel. Did you try to use FaceNet, VGG face models (they seem to work for triplet loss before)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464418,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/31/2019 20:47:25",
          "content": "<p>@Blonde, I tried to use only models pretrained on ImageNet, though if there are models pretrained exclusively on face recognition of similar comparison based tasks it definitely worth a try. \nAlso, I would expect that larger models work better for this competition since when I looked to some hard triplets I couldn't do better than just random guessing... Though too large models are difficult to train given the GPU RAM limitation(((</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464451,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "01/31/2019 22:40:40",
          "content": "<p>yes, batch size it a limitation for triplet loss... did you try to start with even smaller images, 112x112 or 96x96 and large batch, could be better before moving to 224x224, although it depends on the network you choose</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464486,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/01/2019 00:48:35",
          "content": "<p>I didn't. For largest model I tried so far, ResNeXt50, I can run 96 224x224 images per batch with using half precision that gives ~10k compared pairs of images per batch with all batch approach. When I ran the code on 2 GPUs I can get ~40k comparisons that I expect to be enough at the initial stage of training. However for bigger models one may start from smaller images to have large enough batches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464741,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/01/2019 12:15:36",
          "content": "<p>Hey <a href=\"/iafoss\">@iafoss</a>, I'm implementing the triplet network like the two papers you mentioned and so far it seems to be working. I've got a 0.7 LB using a ResNet18, batch-hard mining without any augmentations on 256x256 pixel images. It seems it still have plenty of room for improvements, however I don't know if it could the 0.9x. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466452,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "02/05/2019 12:03:55",
          "content": "<p>Thank  <a href=\"/iafoss\">@iafoss</a> and <a href=\"/jeandebleau\">@jeandebleau</a> for suggested papers. I implemented all of them. But none of them is better than <code>Online Triplet Loss</code>. With Resnet18 backbone, no augmentations, 224x224 images, I could achieve 0.755 LB after 100 epochs (1 hour for training). I refer triplet loss, architecture in here: \n<a href=\"https://github.com/adambielski/siamese-triplet#online-triplet-selection\">https://github.com/adambielski/siamese-triplet#online-triplet-selection</a>. \nHope it is useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466478,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "02/05/2019 13:09:58",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> I assume you meant 1 hour per epoch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466483,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "02/05/2019 13:19:47",
          "content": "<p><a href=\"/valanm\">@valanm</a> I mean 1 hour for 100 epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466498,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "02/05/2019 13:45:26",
          "content": "<p>thx, that's great. i will look into it. \nbtw, i reached 0.78 with <a href=\"/iafoss\">@iafoss</a> kernel (training within kernel time) - changed bs, augmentation and head</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466562,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/05/2019 16:19:54",
          "content": "<p><a href=\"/valanm\">@valanm</a> , Thank you for sharing your insights. It looks that shear augmentation  is quite helpful <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79384\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79384</a> . I'm curious, if you made the head more complicated or simplified to to one similar to Martin's kernel, where the head is just a pooling layer.</p>\n\n<p>Just an update on one more test I have performed. I tried a network predicting score based on x1 - x2 and x1*x2 features and trained with Focal loss calculated in batch all manner, but the result, ~0.72, is worse than my public kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 467671,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "02/07/2019 14:43:13",
          "content": "<p>head is similar to your kernel (tiny changes - slightly more complicated) :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465129,
      "author_name": "shisususu",
      "author_url": "",
      "post_date": "02/02/2019 11:57:39",
      "content": "<p>Thank you, I always learn a lot from you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 467195,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "02/06/2019 16:04:37",
      "content": "<p>One more thing that boosted the score:\nAdding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.740 -&gt; 0.771 (0.780 in a preliminary test).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 467324,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/06/2019 21:55:22",
      "content": "<p>you may want to consider this little trick i commented here:</p>\n\n<p><a href=\"https://www.kaggle.com/seesee/siamese-pretrained-0-822\">https://www.kaggle.com/seesee/siamese-pretrained-0-822</a></p>\n\n<hr>\n\n<p>\"... i visually inspect your results. Many of the new whales (i.e. top1=new__whale) are not really really new whales, and the true id is actually top2 or 3 prediction.</p>\n\n<p>Since we know that there are about 27% new-whale you can do a reassignment by choosing only the top most likely new-whale (i.e threshold base on rank instead of numeric threshold, etc.)</p>\n\n<p>A rough calculation gives improvement of +0.04 if the new__whale are replaced correctly by rank2 or rank3 prediction.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 467373,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "02/07/2019 01:11:14",
          "content": "<p>Hi, Heng. So how do we judge that top1=new_whale is wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472740,
          "author_name": "kerukun",
          "author_url": "",
          "post_date": "02/16/2019 15:12:33",
          "content": "<p>check <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/80624\">https://www.kaggle.com/c/humpback-whale-identification/discussion/80624</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 468088,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "02/08/2019 09:19:28",
      "content": "<p>I found a useful repo here: <a href=\"https://github.com/bnulihaixia/Deep_metric\">https://github.com/bnulihaixia/Deep_metric</a> . \nThe <code>WeightLoss</code> helped me to reach <code>0.785</code> LB by resnet18, 224x224, no aug which is an improvement compared to <code>Online TripletLoss</code> (0.755)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 468169,
      "author_name": "swahengb",
      "author_url": "",
      "post_date": "02/08/2019 12:24:15",
      "content": "<p>@Iafoss Great work. Thank you. This work help me a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470836,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "02/13/2019 16:32:37",
      "content": "<p>After starting using DenseNet169 backbone I got an improvement from 0.771 to 0.800.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471865,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "02/15/2019 02:16:45",
      "content": "<p>Great work! After changing image size from 244 to 512, I could get around 0.834 LB. I'll switch to SGD with proper learning rate schedule and train it for a longer time to see where it can reach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 471896,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/15/2019 03:54:20",
          "content": "<p>Thanks for checking it, looking forward to hearing from you.\nAnother thing that helped to get 0.814 within the kernel time limit is adding batch norm after flattening and increasing the size of the intermediate head layer and embedding dimensionality to 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472576,
          "author_name": "dilapsky",
          "author_url": "",
          "post_date": "02/16/2019 07:40:46",
          "content": "<p>I just change sz=512, but get LB about 0.5. Are there any other value that I should change? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472739,
          "author_name": "msmelguizo",
          "author_url": "",
          "post_date": "02/16/2019 15:11:47",
          "content": "<p><a href=\"/dilapsky\">@dilapsky</a> yes you have to adjust the learning rate with the new batch size and new image size \nlearner.lr_find()</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472758,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/16/2019 16:09:09",
          "content": "<p><a href=\"/dilapsky\">@dilapsky</a> , It also may be a problem with just too small batch itself, resulted for example by batch normalization. In protein competition <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a> some people used gradient accumulation to handle it. Also one can freeze bn, but if I remember correctly, it doesn't work in fast.ai 0.7 with half precision. Not sure if adjusting parameters of bn, such as momentum, can help. Another thing, training for some certain number of epochs is needed, in the very beginning the the result may drop a little bit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472913,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/16/2019 22:05:01",
          "content": "<p>I have also observed variability in the results from run to run. I had to re-run the same notebook and got ~0.02 less in LB, then to be sure that I haven't change anything, I ran the notebook again, and got the same like before (better by ~0.02). </p>\n\n<p>Perhaps this is due to the random choice of  training images that lead to slightly better/worse training depending on pure chance of image pickup, which is telling how much important is the image selection  in similarity based NN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472922,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/16/2019 22:32:59",
          "content": "<p><a href=\"/msmelguizo\">@msmelguizo</a>\nThere's a paper from 2018 that  supports your approach to tweaking LRs when decreasing/increasing batchsize:</p>\n\n<p><a href=\"https://arxiv.org/abs/1706.02677\">https://arxiv.org/abs/1706.02677</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472971,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/17/2019 02:34:00",
          "content": "<p>Hey <a href=\"/dilapsky\">@dilapsky</a>, if you change nothing just set img size to 512 with batch size 16, the score would be around 0.83. And @Iafoss, I've tried larger embedding dimensionality 512 with img size 512 and noticed that the val loss becomes quite unstable. After training for a longer time, val loss becomes inf. Still trying to figure out why. And congratulations on your huge leap in LB score! Did you get this score based on your kernel?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472992,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/17/2019 03:40:30",
          "content": "<p>Thanks, it is based on ensembling and cumulative efforts of all teammates. Regarding unstable val, I had something like u (val loss was fluctuation and increasing while competition score got better) before I added compactification term to the loss. Probably, large  embedding dimensionality exaggerates this effect. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473154,
          "author_name": "msmelguizo",
          "author_url": "",
          "post_date": "02/17/2019 12:36:17",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, Can you please provide more details on how to handle the batch size using gradient accumulation? I  don't mind if I have to train full precision. I have both a 1080Ti and a 2080 card (two different computers) and it is painful to watch how bad the results of the 2080 card are. Thanks again for your great explanations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473312,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/17/2019 19:02:53",
          "content": "<p>I didn't try it by my own yet, but gradient accumulation was quite effective for protein competition. Our team didn't go images larger than 512x512 in that competition mainly because after decreasing bs below 16 model performance was dropping. Meanwhile, many people who used larger image resolution reported using gradient accumulation.\nThe idea of this method is updating weights only at each n-th step with using accumulated gradients for all n steps. This feature is missing in both fast.ai 0.7 and 1.0. For v1.0, which I'm less familiar with, you can check the following discussion <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219\">https://forums.fast.ai/t/accumulating-gradients/33219</a> . In fast.ai 0.7 you just need to slightly modify Stepper class making sure that you call self.m.zero_grad() at the beginning of n-th step and self.opt.step() at the end of n-1-th step. I'm not 100% sure that some other adjustment are not needed since I didn't do it yet. Another thing one may try is freezing bn, but in this case half prescription doesn't work because of a bug in fast.ai library, if I remember correctly. Another thing people used to handle small batches is replacement of bn by instance normalization. I'm not sure how effective is that since I never used it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473360,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "02/17/2019 21:08:26",
          "content": "<p>I tried to use accumulation in v1 for the protein challenge following that thread. I managed to get something that did not crash before training, but it seemed to behave weirdly, maybe because I was also using fp16. I plan to look deeper into it when I get a chance, a solid method for gradient accumulation in fastai v1 would be good for everyone. Code is below (if you look closely you may notice some it is Iafoss's code)</p>\n\n<pre><code>from fastai.callbacks import *\npath='.'\n\nclass myOptimWrapper(OptimWrapper):\nn = 8\nistep, izero_grad = 1, 1\ncnt = 0\n\ndef step(self):  \n    if self.istep == self.n :\n        super().step()\n        self.cnt += 1\n        self.istep = 1\n    else :\n        self.istep += 1\n\ndef zero_grad(self):      \n    if self.izero_grad == self.n :\n        super().zero_grad()\n        self.izero_grad = 1\n    else :\n        self.izero_grad += 1\n\n<a href=\"/dataclass\">@dataclass</a>\nclass StepEpochEnd(Callback):\n    learn:Learner\n    def on_epoch_end(self, **kwargs):\n        print(\"real step and zero grad\")\n        self.learn.opt.real_step()\n        self.learn.opt.real_zero_grad()\n\ndef my_create_opt(self, lr:Floats, wd:Floats=0.)-&gt;None:\n    \"Create optimizer with `lr` learning rate and `wd` weight decay.\"\n    self.opt = myOptimWrapper.create(self.opt_func, lr, self.layer_groups,\n                                     wd=wd, true_wd=self.true_wd, bn_wd=self.bn_wd)\n\nLearner.create_opt = my_create_opt\n\nlearn = create_cnn(\ndata,\nresnet50,\ncut=-2,\nsplit_on= _resnet_split,\nloss_func=F.binary_cross_entropy_with_logits, #FocalLoss(logits=True,alpha=alpha_log), \npath=path,    \nmetrics=[f1_score, f1_callback.f1],\ncallback_fns=[partial(GradientClipping, clip=1),\n              partial(EarlyStoppingCallback, monitor='val_loss', min_delta=0.01, patience=6),\n              #partial(StepEpochEnd),\n              ],\n)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473412,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/18/2019 00:08:45",
          "content": "<p>I'm not 100% sure how it is done in fast.ai v1, but I think zero_grad is called before step. In fast.ai 0.7 Stepper class, if I remember correctly, first gradients are set to zero, then forward and backward passes are calculated, followed by weight update. So what may be going on in the above code, at n-th step gradient is set to zero followed by calculation of gradient for n-th step and using this value (only computed for one step) for weight update, others are lost. Correct me, if it is not the case. Also, for FP16 in the code there may be an additional condition, I do not remember right now. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473680,
          "author_name": "dilapsky",
          "author_url": "",
          "post_date": "02/18/2019 11:08:29",
          "content": "<p>I am very new to pytorch and fastai, Thanks for all of you! BTW, I want to fine-tunning  from 224 to 360, to 512 if I can.So I need to reload the previous model. However, I edit the code just by adding learner.load(‘model’) before training, but nothing happened. I mean, loading model can execute but no improvement in result. Are there any errors I trapped into? I only add that one line, and fastai doc indicate that this is just loading weight. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 473873,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/18/2019 16:05:26",
          "content": "<p><a href=\"/dilapsky\">@dilapsky</a> , At the later stage training is quite slow and it is likely that you faced with, but there can be several other things. Usually image resolution is chosen to be multiple to 32, i.e. 224, 384, and 512. If you do it at kaggle you need to set an appropriate path first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474124,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/19/2019 01:02:28",
          "content": "<p>Given the uncertainty on how to do gradient accumulation properly, I asked <a href=\"/sgugger\">@sgugger</a> on the fastai forum whether they can include it in fastai v1. I hope @kcturgutlu method works well and all fastai user will be able to use it if he can do a PR. \nPlease support this request on that <a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/24?u=hwasiti\">forum thread</a> if you are interested on integrating such feature into fastai.</p>\n\n<p>I think this is something that almost always has a utility with any image recognition problem. \nWho doesn't want to try checking whether increasing image size will give better accuracy? Our new limit will be the max image size that a GPU memory can handle with only 1 batch size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474323,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "02/19/2019 08:16:31",
          "content": "<p>Sounds crazy but as a matter of fact one can also use &lt;1 image per batch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474355,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/19/2019 08:58:58",
          "content": "<p>Seems Batch Normalization shouldn't be the same, when using gradient accumulation. That's why people are getting lower accuracy with grad. accum. \nSee:\n<a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti\">https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti</a></p>\n\n<p><a href=\"https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti\">https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti</a></p>\n\n<p>and maybe that means we have to fiddle with the arch too..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474368,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/19/2019 09:17:28",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a>, were you able to run Lafoss's script with image size 512 and batch size 16 on Kaggle kernel? I keep getting out of memory error even with img sz 320 ans bsz 16. I forked version-11 of the kernel.</p>\n\n<p><a href=\"/iafoss\">@iafoss</a> thanks so much for all the sharing. For the lates version-13 of your kernel, I keep getting memory error when I run it wthout any changes. Do you know that happens? Any advise?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474400,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/19/2019 09:50:09",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, I'm not sure whether it would work on kaggle kernl but it's ok on my local GTX 1080Tis with img size 512 and batch size 16. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474503,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/19/2019 13:13:11",
          "content": "<p>Perhaps you have 2 GPUs. I have 2 GPUs 1080Ti and that is the only way to run bs 16 with SZ 512. \nand btw, the kernel will run on all your gpus by default.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474505,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/19/2019 13:21:52",
          "content": "<p>Yes. I use 2 GPUs 1080Ti to run with img sz 512 and bs 16. And may I ask what's your best score with different img size. I can just reach 0.847 with 512 and 0.820~ with 224.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474520,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/19/2019 13:52:58",
          "content": "<p>The 2 GPU run haven't finished yet. The 512 bs 8 =&gt; LB 0.821. This is my best score so far with fastai.  I have better scores but those from the keras model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474625,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/19/2019 16:34:41",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> , it's strange, I ran several similar kernels at kaggle without memory errors. It should fit in 12 GB. If you are running it on your own computer, may it be something else occupying part of GPU memory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474772,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/19/2019 19:26:58",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, I am running it on Kaggle but kept getting the memory error. I wanted to try it out on the Kaggle kernel first. I really don't know why I cannot run it successfully with no change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474980,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/20/2019 04:50:20",
          "content": "<p>Could you share the link of the public keras kernel you're using? I want to see whether there's something more I can do to improve the score. Thanks :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 475073,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "02/20/2019 07:54:21",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a>, I am using <a href=\"/iafoss\">@iafoss</a>'s kernel <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit\">here</a> ver-11 and ver-13 as seperate kernels. I am able to run ver-11 with different parameters and got lower result (0.785) but ver-13 is the one giving memory error when run without any change.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 473728,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "02/18/2019 12:39:12",
      "content": "<p>@Iafoss why did you choose DenseNet in particular and not other architecture?</p>\n\n<p>There are a few other arch that can get better accuracies on imagenet :\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 473889,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/18/2019 16:36:57",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a> , Actually, you are quite limited with the chose of the model given limitations of GPU RAM and large image resolution. If you do not have several GPU with 16+ GB of RAM you are likely should not looking for using big models. The rest are ResNet18, ResNet34, ResNeXt50(SE), DenseNet121, DenseNet169. You also may consider ResNet101(SE), DenseNet201, Inception, and NasNet if you have several GPUs or willing to invest time to fight with small batches. I tried ResNet34 and ResNeXt50SE, but there performance was worse than ResNeXt50. SE I expect is more difficult to retrain, and I never used it for the production models so far. The same with NasNet, I tried it several times but the results were poor. I guess it may be also quite difficult to retrain or just architecture itself is overfitted to ImageNet, and doesn't perform well on other image sets. And I also usually ignore ResNet50+ since they have about the same memory requirement as corresponding ResNeXt, but worse performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 474108,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/19/2019 00:20:45",
          "content": "<p>That's a nice observation @lafoss. \nI've always went straight for SEResNeXt50 instead of ResNeXt50. I'll try the non-SE counterpart to see how it goes. As for the more \"designed\" architectures like NasNet I also think ti is made to overfit to ImageNet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "463855": "Greetings,\nAfter many experiments and substantial refactoring the model, [my kernel][1] could reach **0.805 LB** score for **training within the kernel time limit** (no ensembling and external weights apart from ones pretrained on ImageNet) **with using a similarity distance based approach**. I just wanted to summarize milestones of the kernel score improvement (you can check details in the kernel), and I hope you will find them useful and could apply to your models. Also you can share your ides on the improvement of the score. The provided values of score show the result obtained after running the model within the kernel time limit, each new line indicates the modifications made to the model and the corresponding value of public LB.\n\n1) Triplet network with triplet loss (random selection of images): ~0.3 LB \n2) Batch all contrastive loss (compute loss based on all pairs in a batch, for 96 images per batch this method gives 9120 comparisons instead of 32 triplets): 0.5 LB\n3) Rectangular crops based on bounding boxes (without image distortion) instead of random ones: 0.55 LB\n4) Increase the dimensionality of the embedding space from 64 to 256 + ResNeXt50 instead of ResNet34: 0.606 LB (V1 and 2 of [this kernel][2])\n5) Metric learning `d^2 = (x1-x2).T*A*(x1-x2)` instead of Euclidean distance: 0.655 LB\n6) Optimized form of metric (check kernel for details) + hard negative example mining + lr optimization + going to smaller images of size 128x384 instead of 192x576 + averaging only nonzero contributions when loss is computed: 0.699 LB  (V6)\n7) Switching to bounding box crops rescaled to square images of size 224x224: 0.740 (0.748 in a preliminary test). (V7)\n8) Adding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.771 (0.780 in a preliminary test). (V9)\n9) DenseNet169 backbone: 0.800 (V11)\n10) DenseNet121 backbone: 0.805 (V13)\n\n  [1]: https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\n  [2]: https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit",
    "463873": "just by curiosity, are you just doing a classification or a Siamese kind of arch?  With so few samples of each class, I wouldn't expect a classification to work, or does it?\nWhich is the simplest model that can score high doing a classification?",
    "463880": "I use similarity based approach. Due to numerous classes presented by just several examples I didn't run classification approach for this competition. Though, if one just drop all classes that appear in the train less, let's say than 20 times. the classification could work to some extend, I expect. The thing with classification is that it is very simple to reach ~0.6 LB, and there are many public kernels on that. However, I don't know if image classification can give higher score. \nMeanwhile, similarity based approach requires much more elaborated work done on the code and development of the model. In particular, just a simplistic Siamese or Triplet network with random selection of pairs or triplets gives only ~0.3 LB. [Martin in his kernel][1] did quite a great job in utilizing LAP and Metric learning. Fine-tuning his model trained for 400 epochs could give [0.822 LB][2].\nIn this post I tried to summarize what I did to be able to reach ~0.75 LB with a similarity based approach for training from scratch (ImageNet pretrained weights) within the kernel time limit.\n\n\n  [1]: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n  [2]: https://www.kaggle.com/seesee/siamese-pretrained-0-822/notebook",
    "463977": "Hi, how do you get bounding box info? Do you use martinpiotte's pretrained model from https://www.kaggle.com/martinpiotte/bounding-box-model ?",
    "463988": "Yes, I used a fork based on this model https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes",
    "464101": "Great work @iafoss! (as usual)",
    "464109": "Thanks",
    "464183": "Great Work and thank you for your insight !\n\nDid you try some alternatives to the contrastive loss ?\n\nI think that you may reach a better score using some related work, for instance:\nhttp://www.nec-labs.com/uploads/images/Department-Images/MediaAnalytics/papers/nips16_npairmetriclearning.pdf\n\nThere is also an interesting paper beeing discussed here:\nhttps://openreview.net/forum?id=HkxLXnAcFQ\n\nThe changes required to test these ideas are minimal. I just started the competition, I did not submitt anything up to now. Using the ideas presented above, I could train a resnet-18 on 256x256 images that reach a local map@5 of 0.9+ for known whales. I do not know yet if this is good or not. The convergence of Multi-class N-pair Loss is very fast and very easy to implement. Basically all are variants of a softmax loss based on a distance measures between a subset of classes presented in a minibatch.",
    "464198": "Thanks for your reply, emm, I was also using martinpiotte's pretrained model, I got same results with  https://www.kaggle.com/suicaokhoailang/generating-whale-bounding-boxes, but training with bbox gives worse results, there might be sth wrong in my code ....",
    "464206": "I was able to get 0.774 public LB using the first paper mentioned above using 224x224 images (0.90 top 1 accuracy on known whales). It's pretty variable though - most of my models are bouncing around the 0.70-0.77 range at the moment. I'm going to try increasing the image size and see how that goes.",
    "464211": "Thank you for sharing. I will continue to play around these ideas as well.",
    "464383": "jeandebleau, thank you for sharing the papers, I'll try to look more carefully into them. The things I also tried are based on batch hard triplet loss mentioned by @mnpinto: https://arxiv.org/pdf/1901.03662.pdf and https://arxiv.org/pdf/1703.07737.pdf . Though, I couldn't get better results when I used the approach from these paper. Probably, such method may need more time to converge. Also the problem with triplet loss may be that despite the loss does amazing job in keeping images of the same kind together, it doesn't care how close from each other are the images of the same class. In other words it doesn't try to bring predictions for the same labels into one point in the embedding space. If there was no class `new_whale`, such method would work really the best. However, I assign `new_whale` label based on a fixed distance in the embedding space. For example, if there are no neighbors at a distance d0 for a particular image, `new_whale` is the first prediction for this image. The problem is that if the density of points with the same label in the embedding space is different, one my assign new_whale to points having low density. However, again, probably I just didn't train it for long enough.",
    "464401": "iafoss thank you for your kernel. Did you try to use FaceNet, VGG face models (they seem to work for triplet loss before)?",
    "464418": "Blonde, I tried to use only models pretrained on ImageNet, though if there are models pretrained exclusively on face recognition of similar comparison based tasks it definitely worth a try. \nAlso, I would expect that larger models work better for this competition since when I looked to some hard triplets I couldn't do better than just random guessing... Though too large models are difficult to train given the GPU RAM limitation(((",
    "464451": "yes, batch size it a limitation for triplet loss... did you try to start with even smaller images, 112x112 or 96x96 and large batch, could be better before moving to 224x224, although it depends on the network you choose",
    "464486": "I didn't. For largest model I tried so far, ResNeXt50, I can run 96 224x224 images per batch with using half precision that gives ~10k compared pairs of images per batch with all batch approach. When I ran the code on 2 GPUs I can get ~40k comparisons that I expect to be enough at the initial stage of training. However for bigger models one may start from smaller images to have large enough batches.",
    "464741": "Hey @iafoss, I'm implementing the triplet network like the two papers you mentioned and so far it seems to be working. I've got a 0.7 LB using a ResNet18, batch-hard mining without any augmentations on 256x256 pixel images. It seems it still have plenty of room for improvements, however I don't know if it could the 0.9x.",
    "465129": "Thank you, I always learn a lot from you.",
    "466452": "Thank  @iafoss and @jeandebleau for suggested papers. I implemented all of them. But none of them is better than `Online Triplet Loss`. With Resnet18 backbone, no augmentations, 224x224 images, I could achieve 0.755 LB after 100 epochs (1 hour for training). I refer triplet loss, architecture in here: \nhttps://github.com/adambielski/siamese-triplet#online-triplet-selection. \nHope it is useful.",
    "466478": "backaggle I assume you meant 1 hour per epoch",
    "466483": "valanm I mean 1 hour for 100 epochs.",
    "466498": "thx, that's great. i will look into it. \nbtw, i reached 0.78 with @iafoss kernel (training within kernel time) - changed bs, augmentation and head",
    "466562": "valanm , Thank you for sharing your insights. It looks that shear augmentation  is quite helpful https://www.kaggle.com/c/humpback-whale-identification/discussion/79384 . I'm curious, if you made the head more complicated or simplified to to one similar to Martin's kernel, where the head is just a pooling layer.\n\nJust an update on one more test I have performed. I tried a network predicting score based on x1 - x2 and x1*x2 features and trained with Focal loss calculated in batch all manner, but the result, ~0.72, is worse than my public kernel.",
    "467195": "One more thing that boosted the score:\nAdding a compactification term to the loss (L2 regularization of all distances within a batch). This change encourages the network to group the predictions in the embedding space instead of increasing the distance between them: 0.740 -&gt; 0.771 (0.780 in a preliminary test).",
    "467324": "you may want to consider this little trick i commented here:\n\nhttps://www.kaggle.com/seesee/siamese-pretrained-0-822\n\n---\n\n\"... i visually inspect your results. Many of the new whales (i.e. top1=new__whale) are not really really new whales, and the true id is actually top2 or 3 prediction.\n\nSince we know that there are about 27% new-whale you can do a reassignment by choosing only the top most likely new-whale (i.e threshold base on rank instead of numeric threshold, etc.)\n\nA rough calculation gives improvement of +0.04 if the new__whale are replaced correctly by rank2 or rank3 prediction.\"",
    "467373": "Hi, Heng. So how do we judge that top1=new_whale is wrong?",
    "467671": "head is similar to your kernel (tiny changes - slightly more complicated) :)",
    "468088": "I found a useful repo here: https://github.com/bnulihaixia/Deep_metric . \nThe `WeightLoss` helped me to reach `0.785` LB by resnet18, 224x224, no aug which is an improvement compared to `Online TripletLoss` (0.755)",
    "468169": "Iafoss Great work. Thank you. This work help me a lot.",
    "470836": "After starting using DenseNet169 backbone I got an improvement from 0.771 to 0.800.",
    "471865": "Great work! After changing image size from 244 to 512, I could get around 0.834 LB. I'll switch to SGD with proper learning rate schedule and train it for a longer time to see where it can reach.",
    "471896": "Thanks for checking it, looking forward to hearing from you.\nAnother thing that helped to get 0.814 within the kernel time limit is adding batch norm after flattening and increasing the size of the intermediate head layer and embedding dimensionality to 512.",
    "472576": "I just change sz=512, but get LB about 0.5. Are there any other value that I should change? Thanks!",
    "472739": "dilapsky yes you have to adjust the learning rate with the new batch size and new image size \nlearner.lr_find()",
    "472740": "check https://www.kaggle.com/c/humpback-whale-identification/discussion/80624",
    "472758": "dilapsky , It also may be a problem with just too small batch itself, resulted for example by batch normalization. In protein competition https://www.kaggle.com/c/human-protein-atlas-image-classification some people used gradient accumulation to handle it. Also one can freeze bn, but if I remember correctly, it doesn't work in fast.ai 0.7 with half precision. Not sure if adjusting parameters of bn, such as momentum, can help. Another thing, training for some certain number of epochs is needed, in the very beginning the the result may drop a little bit.",
    "472913": "I have also observed variability in the results from run to run. I had to re-run the same notebook and got ~0.02 less in LB, then to be sure that I haven't change anything, I ran the notebook again, and got the same like before (better by ~0.02). \n\nPerhaps this is due to the random choice of  training images that lead to slightly better/worse training depending on pure chance of image pickup, which is telling how much important is the image selection  in similarity based NN.",
    "472922": "msmelguizo\nThere's a paper from 2018 that  supports your approach to tweaking LRs when decreasing/increasing batchsize:\n\nhttps://arxiv.org/abs/1706.02677",
    "472971": "Hey @dilapsky, if you change nothing just set img size to 512 with batch size 16, the score would be around 0.83. And @Iafoss, I've tried larger embedding dimensionality 512 with img size 512 and noticed that the val loss becomes quite unstable. After training for a longer time, val loss becomes inf. Still trying to figure out why. And congratulations on your huge leap in LB score! Did you get this score based on your kernel?",
    "472992": "Thanks, it is based on ensembling and cumulative efforts of all teammates. Regarding unstable val, I had something like u (val loss was fluctuation and increasing while competition score got better) before I added compactification term to the loss. Probably, large  embedding dimensionality exaggerates this effect.",
    "473154": "iafoss, Can you please provide more details on how to handle the batch size using gradient accumulation? I  don't mind if I have to train full precision. I have both a 1080Ti and a 2080 card (two different computers) and it is painful to watch how bad the results of the 2080 card are. Thanks again for your great explanations.",
    "473312": "I didn't try it by my own yet, but gradient accumulation was quite effective for protein competition. Our team didn't go images larger than 512x512 in that competition mainly because after decreasing bs below 16 model performance was dropping. Meanwhile, many people who used larger image resolution reported using gradient accumulation.\nThe idea of this method is updating weights only at each n-th step with using accumulated gradients for all n steps. This feature is missing in both fast.ai 0.7 and 1.0. For v1.0, which I'm less familiar with, you can check the following discussion https://forums.fast.ai/t/accumulating-gradients/33219 . In fast.ai 0.7 you just need to slightly modify Stepper class making sure that you call self.m.zero_grad() at the beginning of n-th step and self.opt.step() at the end of n-1-th step. I'm not 100% sure that some other adjustment are not needed since I didn't do it yet. Another thing one may try is freezing bn, but in this case half prescription doesn't work because of a bug in fast.ai library, if I remember correctly. Another thing people used to handle small batches is replacement of bn by instance normalization. I'm not sure how effective is that since I never used it.",
    "473360": "I tried to use accumulation in v1 for the protein challenge following that thread. I managed to get something that did not crash before training, but it seemed to behave weirdly, maybe because I was also using fp16. I plan to look deeper into it when I get a chance, a solid method for gradient accumulation in fastai v1 would be good for everyone. Code is below (if you look closely you may notice some it is Iafoss's code)\n\n    from fastai.callbacks import *\n    path='.'\n\n    class myOptimWrapper(OptimWrapper):\n    n = 8\n    istep, izero_grad = 1, 1\n    cnt = 0\n\n    def step(self):  \n        if self.istep == self.n :\n            super().step()\n            self.cnt += 1\n            self.istep = 1\n        else :\n            self.istep += 1\n\n    def zero_grad(self):      \n        if self.izero_grad == self.n :\n            super().zero_grad()\n            self.izero_grad = 1\n        else :\n            self.izero_grad += 1\n\n    @dataclass\n    class StepEpochEnd(Callback):\n        learn:Learner\n        def on_epoch_end(self, **kwargs):\n            print(\"real step and zero grad\")\n            self.learn.opt.real_step()\n            self.learn.opt.real_zero_grad()\n\n    def my_create_opt(self, lr:Floats, wd:Floats=0.)-&gt;None:\n        \"Create optimizer with `lr` learning rate and `wd` weight decay.\"\n        self.opt = myOptimWrapper.create(self.opt_func, lr, self.layer_groups,\n                                         wd=wd, true_wd=self.true_wd, bn_wd=self.bn_wd)\n\n    Learner.create_opt = my_create_opt\n\n    learn = create_cnn(\n    data,\n    resnet50,\n    cut=-2,\n    split_on= _resnet_split,\n    loss_func=F.binary_cross_entropy_with_logits, #FocalLoss(logits=True,alpha=alpha_log), \n    path=path,    \n    metrics=[f1_score, f1_callback.f1],\n    callback_fns=[partial(GradientClipping, clip=1),\n                  partial(EarlyStoppingCallback, monitor='val_loss', min_delta=0.01, patience=6),\n                  #partial(StepEpochEnd),\n                  ],\n    )",
    "473412": "I'm not 100% sure how it is done in fast.ai v1, but I think zero_grad is called before step. In fast.ai 0.7 Stepper class, if I remember correctly, first gradients are set to zero, then forward and backward passes are calculated, followed by weight update. So what may be going on in the above code, at n-th step gradient is set to zero followed by calculation of gradient for n-th step and using this value (only computed for one step) for weight update, others are lost. Correct me, if it is not the case. Also, for FP16 in the code there may be an additional condition, I do not remember right now.",
    "473680": "I am very new to pytorch and fastai, Thanks for all of you! BTW, I want to fine-tunning  from 224 to 360, to 512 if I can.So I need to reload the previous model. However, I edit the code just by adding learner.load(‘model’) before training, but nothing happened. I mean, loading model can execute but no improvement in result. Are there any errors I trapped into? I only add that one line, and fastai doc indicate that this is just loading weight. Thanks!",
    "473728": "Iafoss why did you choose DenseNet in particular and not other architecture?\n\nThere are a few other arch that can get better accuracies on imagenet :\nhttps://github.com/Cadene/pretrained-models.pytorch/blob/master/README.md#accuracy-on-validation-set-single-model",
    "473873": "dilapsky , At the later stage training is quite slow and it is likely that you faced with, but there can be several other things. Usually image resolution is chosen to be multiple to 32, i.e. 224, 384, and 512. If you do it at kaggle you need to set an appropriate path first.",
    "473889": "hwasiti , Actually, you are quite limited with the chose of the model given limitations of GPU RAM and large image resolution. If you do not have several GPU with 16+ GB of RAM you are likely should not looking for using big models. The rest are ResNet18, ResNet34, ResNeXt50(SE), DenseNet121, DenseNet169. You also may consider ResNet101(SE), DenseNet201, Inception, and NasNet if you have several GPUs or willing to invest time to fight with small batches. I tried ResNet34 and ResNeXt50SE, but there performance was worse than ResNeXt50. SE I expect is more difficult to retrain, and I never used it for the production models so far. The same with NasNet, I tried it several times but the results were poor. I guess it may be also quite difficult to retrain or just architecture itself is overfitted to ImageNet, and doesn't perform well on other image sets. And I also usually ignore ResNet50+ since they have about the same memory requirement as corresponding ResNeXt, but worse performance.",
    "474108": "That's a nice observation @lafoss. \nI've always went straight for SEResNeXt50 instead of ResNeXt50. I'll try the non-SE counterpart to see how it goes. As for the more \"designed\" architectures like NasNet I also think ti is made to overfit to ImageNet.",
    "474124": "Given the uncertainty on how to do gradient accumulation properly, I asked @sgugger on the fastai forum whether they can include it in fastai v1. I hope @kcturgutlu method works well and all fastai user will be able to use it if he can do a PR. \nPlease support this request on that [forum thread][1] if you are interested on integrating such feature into fastai.\n\nI think this is something that almost always has a utility with any image recognition problem. \nWho doesn't want to try checking whether increasing image size will give better accuracy? Our new limit will be the max image size that a GPU memory can handle with only 1 batch size.\n\n\n  [1]: https://forums.fast.ai/t/accumulating-gradients/33219/24?u=hwasiti",
    "474323": "Sounds crazy but as a matter of fact one can also use &lt;1 image per batch",
    "474355": "Seems Batch Normalization shouldn't be the same, when using gradient accumulation. That's why people are getting lower accuracy with grad. accum. \nSee:\n[https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti][1]\n\n[https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti][2]\n\nand maybe that means we have to fiddle with the arch too..\n\n  [1]: https://forums.fast.ai/t/accumulating-gradients/33219/35?u=hwasiti\n  [2]: https://forums.fast.ai/t/accumulating-gradients/33219/42?u=hwasiti",
    "474368": "syoya1997, were you able to run Lafoss's script with image size 512 and batch size 16 on Kaggle kernel? I keep getting out of memory error even with img sz 320 ans bsz 16. I forked version-11 of the kernel.\n\n@iafoss thanks so much for all the sharing. For the lates version-13 of your kernel, I keep getting memory error when I run it wthout any changes. Do you know that happens? Any advise?",
    "474400": "sheriytm, I'm not sure whether it would work on kaggle kernl but it's ok on my local GTX 1080Tis with img size 512 and batch size 16.",
    "474503": "Perhaps you have 2 GPUs. I have 2 GPUs 1080Ti and that is the only way to run bs 16 with SZ 512. \nand btw, the kernel will run on all your gpus by default.",
    "474505": "Yes. I use 2 GPUs 1080Ti to run with img sz 512 and bs 16. And may I ask what's your best score with different img size. I can just reach 0.847 with 512 and 0.820~ with 224.",
    "474520": "The 2 GPU run haven't finished yet. The 512 bs 8 =&gt; LB 0.821. This is my best score so far with fastai.  I have better scores but those from the keras model.",
    "474625": "sheriytm , it's strange, I ran several similar kernels at kaggle without memory errors. It should fit in 12 GB. If you are running it on your own computer, may it be something else occupying part of GPU memory.",
    "474772": "iafoss, I am running it on Kaggle but kept getting the memory error. I wanted to try it out on the Kaggle kernel first. I really don't know why I cannot run it successfully with no change.",
    "474980": "Could you share the link of the public keras kernel you're using? I want to see whether there's something more I can do to improve the score. Thanks :)",
    "475073": "syoya1997, I am using @iafoss's kernel [here][1] ver-11 and ver-13 as seperate kernels. I am able to run ver-11 with different parameters and got lower result (0.785) but ver-13 is the one giving memory error when run without any change.\n\n\n  [1]: https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit"
  },
  "source": "meta"
}