{
  "id": 183629,
  "title": "EfficientNet Imagenet vs Noisy Student",
  "url": "/competitions/landmark-recognition-2020/discussion/183629",
  "author_name": "",
  "post_date": "2020-09-17T14:21:13.110491300Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We have all seen this picture<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa8a2466bc9ad2f64a21032fc7f589647%2FSkrmklipp.jpg?generation=1600337323798529&amp;alt=media\" alt=\"\"></p>\n<p>If feels obvious to just use Noisy Student in every case, right. But do we see the same results and gap between them when we use EfficientNet Noisy Student instead of EfficientNet Imagenet in various of problems, that was the reason for reading more about the background to the results.</p>\n<p>We can see in the paper is that following steps are used to get the Noisy Student results:<br>\n\"On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled andpseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment\"<br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\n<a href=\"https://github.com/google-research/noisystudent\" target=\"_blank\">https://github.com/google-research/noisystudent</a></p>\n<p>So Noisy Student is a training framework.<br>\nAnd a framework to use with any architectures and models, not only EfficientNet.<br>\n\"Noisy Student Training leads to an improvement of 1.3% on the baseline(ResNet-50) model,<br>\nwhich shows that Noisy Student Training is effective for architectures other than EfficientNet.\"</p>\n<p>What we get though within the model's architecture is noise methods such as stochastic depth.</p>\n<p>When it comes to EfficientNet AdvProp it have the whole solution implemented in the architectures.<br>\n\"Key to our method is the usage of a separate auxiliary batch norm for adversarial examples, as they have different underlying distributions to normal examples. We show that AdvProp improves a wide range of models on various image recognition tasks and performs better when the models are bigger\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F6b30f473af6276c580dc8474482d8abd%2FSkrmklipp.jpg?generation=1600345942457686&amp;alt=media\" alt=\"\"><br>\nEfficientNet AdvProp<br>\nAdversarial Examples Improve Image Recognition<br>\n<a href=\"https://arxiv.org/pdf/1911.09665.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.09665.pdf</a></p>\n<p><strong>Conclusion</strong>:<br>\nFirst and foremost, what an amazing work from the Google Research Team, after reading 10+ research papers, what an achievement to move the area to a new level. A lot of work done in a short time, impressive. 🙏</p>\n<p>If I have not understood the whole thing completely wrong ( disclaimer), if one wants the same results and gap as in the picture in other problems one have to use the noisy student training framework not just the pretrained weights, even though the extra trained images that comes with Noisy Student weights and the noise methods will certainly will help to get the best results in some tasks. Otherwise maybe it's equivalent to use the Imagenet-weights or maybe even better use the EfficientNet AdvProp?</p>\n<p>If one takes time and read the training steps, papers, architectures, pre/post processing etc one can find many great ideas, methods and modules to use in new tasks and problems, methods like augmix, autoaugment, randaugment, hardswish just to name a few.</p>\n<p>AugMix<br>\n<a href=\"https://arxiv.org/pdf/1912.02781.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.02781.pdf</a><br>\nAutoAugment<br>\n<a href=\"https://arxiv.org/pdf/1805.09501.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.09501.pdf</a><br>\nRandAugment<br>\n<a href=\"https://arxiv.org/pdf/1909.13719.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.13719.pdf</a></p>\n<p>Thanks again to the Google Research Team, and will be exciting to see what comes next 🙂</p>\n<p><strong>Update</strong>:<br>\nHere is a comparision for this task, B1ns and B1imgnet after 24 epochs<br>\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.</p>\n<p>B1ns<br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578<br>\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699<br>\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034<br>\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092<br>\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691<br>\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507<br>\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408<br>\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272</p>\n<p>B1imgnet<br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422<br>\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105<br>\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652<br>\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018<br>\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082<br>\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385<br>\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561<br>\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989</p>\n<p>Better ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy with Imagenet.<br>\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.</p>",
  "messages": [
    {
      "id": "1014535",
      "postDate": "09/17/2020 14:21:13",
      "content": "<p>We have all seen this picture<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa8a2466bc9ad2f64a21032fc7f589647%2FSkrmklipp.jpg?generation=1600337323798529&amp;alt=media\" alt=\"\"></p>\n<p>If feels obvious to just use Noisy Student in every case, right. But do we see the same results and gap between them when we use EfficientNet Noisy Student instead of EfficientNet Imagenet in various of problems, that was the reason for reading more about the background to the results.</p>\n<p>We can see in the paper is that following steps are used to get the Noisy Student results:<br>\n\"On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled andpseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment\"<br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\n<a href=\"https://github.com/google-research/noisystudent\" target=\"_blank\">https://github.com/google-research/noisystudent</a></p>\n<p>So Noisy Student is a training framework.<br>\nAnd a framework to use with any architectures and models, not only EfficientNet.<br>\n\"Noisy Student Training leads to an improvement of 1.3% on the baseline(ResNet-50) model,<br>\nwhich shows that Noisy Student Training is effective for architectures other than EfficientNet.\"</p>\n<p>What we get though within the model's architecture is noise methods such as stochastic depth.</p>\n<p>When it comes to EfficientNet AdvProp it have the whole solution implemented in the architectures.<br>\n\"Key to our method is the usage of a separate auxiliary batch norm for adversarial examples, as they have different underlying distributions to normal examples. We show that AdvProp improves a wide range of models on various image recognition tasks and performs better when the models are bigger\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F6b30f473af6276c580dc8474482d8abd%2FSkrmklipp.jpg?generation=1600345942457686&amp;alt=media\" alt=\"\"><br>\nEfficientNet AdvProp<br>\nAdversarial Examples Improve Image Recognition<br>\n<a href=\"https://arxiv.org/pdf/1911.09665.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.09665.pdf</a></p>\n<p><strong>Conclusion</strong>:<br>\nFirst and foremost, what an amazing work from the Google Research Team, after reading 10+ research papers, what an achievement to move the area to a new level. A lot of work done in a short time, impressive. 🙏</p>\n<p>If I have not understood the whole thing completely wrong ( disclaimer), if one wants the same results and gap as in the picture in other problems one have to use the noisy student training framework not just the pretrained weights, even though the extra trained images that comes with Noisy Student weights and the noise methods will certainly will help to get the best results in some tasks. Otherwise maybe it's equivalent to use the Imagenet-weights or maybe even better use the EfficientNet AdvProp?</p>\n<p>If one takes time and read the training steps, papers, architectures, pre/post processing etc one can find many great ideas, methods and modules to use in new tasks and problems, methods like augmix, autoaugment, randaugment, hardswish just to name a few.</p>\n<p>AugMix<br>\n<a href=\"https://arxiv.org/pdf/1912.02781.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.02781.pdf</a><br>\nAutoAugment<br>\n<a href=\"https://arxiv.org/pdf/1805.09501.pdf\" target=\"_blank\">https://arxiv.org/pdf/1805.09501.pdf</a><br>\nRandAugment<br>\n<a href=\"https://arxiv.org/pdf/1909.13719.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.13719.pdf</a></p>\n<p>Thanks again to the Google Research Team, and will be exciting to see what comes next 🙂</p>\n<p><strong>Update</strong>:<br>\nHere is a comparision for this task, B1ns and B1imgnet after 24 epochs<br>\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.</p>\n<p>B1ns<br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578<br>\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699<br>\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034<br>\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092<br>\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691<br>\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507<br>\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408<br>\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272</p>\n<p>B1imgnet<br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422<br>\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105<br>\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652<br>\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018<br>\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082<br>\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385<br>\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561<br>\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989</p>\n<p>Better ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy with Imagenet.<br>\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.</p>",
      "rawMarkdown": "We have all seen this picture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa8a2466bc9ad2f64a21032fc7f589647%2FSkrmklipp.jpg?generation=1600337323798529&alt=media)\n\nIf feels obvious to just use Noisy Student in every case, right. But do we see the same results and gap between them when we use EfficientNet Noisy Student instead of EfficientNet Imagenet in various of problems, that was the reason for reading more about the background to the results.\n\nWe can see in the paper is that following steps are used to get the Noisy Student results:\n\"On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled andpseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment\"\nhttps://arxiv.org/pdf/1911.04252.pdf\nhttps://github.com/google-research/noisystudent\n\nSo Noisy Student is a training framework.\nAnd a framework to use with any architectures and models, not only EfficientNet.\n\"Noisy Student Training leads to an improvement of 1.3% on the baseline(ResNet-50) model,\nwhich shows that Noisy Student Training is effective for architectures other than EfficientNet.\"\n\nWhat we get though within the model's architecture is noise methods such as stochastic depth.\n\nWhen it comes to EfficientNet AdvProp it have the whole solution implemented in the architectures.\n\"Key to our method is the usage of a separate auxiliary batch norm for adversarial examples, as they have different underlying distributions to normal examples. We show that AdvProp improves a wide range of models on various image recognition tasks and performs better when the models are bigger\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F6b30f473af6276c580dc8474482d8abd%2FSkrmklipp.jpg?generation=1600345942457686&alt=media)\nEfficientNet AdvProp\nAdversarial Examples Improve Image Recognition\nhttps://arxiv.org/pdf/1911.09665.pdf\n\n**Conclusion**:\nFirst and foremost, what an amazing work from the Google Research Team, after reading 10+ research papers, what an achievement to move the area to a new level. A lot of work done in a short time, impressive. 🙏\n\nIf I have not understood the whole thing completely wrong ( disclaimer), if one wants the same results and gap as in the picture in other problems one have to use the noisy student training framework not just the pretrained weights, even though the extra trained images that comes with Noisy Student weights and the noise methods will certainly will help to get the best results in some tasks. Otherwise maybe it's equivalent to use the Imagenet-weights or maybe even better use the EfficientNet AdvProp?\n\nIf one takes time and read the training steps, papers, architectures, pre/post processing etc one can find many great ideas, methods and modules to use in new tasks and problems, methods like augmix, autoaugment, randaugment, hardswish just to name a few.\n\nAugMix\nhttps://arxiv.org/pdf/1912.02781.pdf\nAutoAugment\nhttps://arxiv.org/pdf/1805.09501.pdf\nRandAugment\nhttps://arxiv.org/pdf/1909.13719.pdf\n\nThanks again to the Google Research Team, and will be exciting to see what comes next 🙂\n\n**Update**:\nHere is a comparision for this task, B1ns and B1imgnet after 24 epochs\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.\n\nB1ns\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272\n\nB1imgnet\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989\n\nBetter ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy with Imagenet.\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.",
      "votes": null
    },
    {
      "id": "1015456",
      "postDate": "09/18/2020 07:45:34",
      "content": "<p>Here is a comparision for this task, B1ns and B1imgnet after 24 epochs<br>\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.</p>\n<p>B1ns <br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578<br>\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699<br>\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034<br>\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092<br>\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691<br>\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507<br>\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408<br>\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272</p>\n<p>B1imgnet <br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422<br>\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105<br>\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652<br>\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018<br>\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082<br>\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385<br>\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561<br>\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989</p>\n<p>Better ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy  with Imagenet.<br>\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.</p>",
      "rawMarkdown": "Here is a comparision for this task, B1ns and B1imgnet after 24 epochs\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.\n\nB1ns \nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272\n\nB1imgnet \nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989\n\nBetter ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy  with Imagenet.\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.",
      "votes": null
    },
    {
      "id": "1016647",
      "postDate": "09/19/2020 06:09:05",
      "content": "<p>Do they converge to the same values in the end? Also are you training on a local machine or Colab/Kaggle TPUs?</p>",
      "rawMarkdown": "Do they converge to the same values in the end? Also are you training on a local machine or Colab/Kaggle TPUs?",
      "votes": null
    },
    {
      "id": "1017831",
      "postDate": "09/19/2020 09:37:14",
      "content": "<p>Not trained them to the final yet, out of quota, but now we have a new week I'll continue. I'm using Kaggle, it's a better version of TPU and faster.</p>",
      "rawMarkdown": "Not trained them to the final yet, out of quota, but now we have a new week I'll continue. I'm using Kaggle, it's a better version of TPU and faster.",
      "votes": null
    },
    {
      "id": "1019626",
      "postDate": "09/20/2020 14:47:54",
      "content": "<p>In previous competitions, I have seen that noisy student performance is always better than imagenet.</p>",
      "rawMarkdown": "In previous competitions, I have seen that noisy student performance is always better than imagenet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1015456,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/18/2020 07:45:34",
      "content": "<p>Here is a comparision for this task, B1ns and B1imgnet after 24 epochs<br>\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.</p>\n<p>B1ns <br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578<br>\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699<br>\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034<br>\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092<br>\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691<br>\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507<br>\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408<br>\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272</p>\n<p>B1imgnet <br>\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy<br>\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422<br>\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105<br>\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652<br>\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018<br>\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082<br>\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385<br>\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561<br>\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989</p>\n<p>Better ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy  with Imagenet.<br>\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1016647,
      "author_name": "josealways123",
      "author_url": "",
      "post_date": "09/19/2020 06:09:05",
      "content": "<p>Do they converge to the same values in the end? Also are you training on a local machine or Colab/Kaggle TPUs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1017831,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/19/2020 09:37:14",
          "content": "<p>Not trained them to the final yet, out of quota, but now we have a new week I'll continue. I'm using Kaggle, it's a better version of TPU and faster.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1019626,
      "author_name": "vikrant06",
      "author_url": "",
      "post_date": "09/20/2020 14:47:54",
      "content": "<p>In previous competitions, I have seen that noisy student performance is always better than imagenet.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1014535": "We have all seen this picture\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa8a2466bc9ad2f64a21032fc7f589647%2FSkrmklipp.jpg?generation=1600337323798529&alt=media)\n\nIf feels obvious to just use Noisy Student in every case, right. But do we see the same results and gap between them when we use EfficientNet Noisy Student instead of EfficientNet Imagenet in various of problems, that was the reason for reading more about the background to the results.\n\nWe can see in the paper is that following steps are used to get the Noisy Student results:\n\"On ImageNet, we first train an EfficientNet model on labeled images and use it as a teacher to generate pseudo labels for 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled andpseudo labeled images. We iterate this process by putting back the student as the teacher. During the learning of the student, we inject noise such as dropout, stochastic depth, and data augmentation via RandAugment\"\nhttps://arxiv.org/pdf/1911.04252.pdf\nhttps://github.com/google-research/noisystudent\n\nSo Noisy Student is a training framework.\nAnd a framework to use with any architectures and models, not only EfficientNet.\n\"Noisy Student Training leads to an improvement of 1.3% on the baseline(ResNet-50) model,\nwhich shows that Noisy Student Training is effective for architectures other than EfficientNet.\"\n\nWhat we get though within the model's architecture is noise methods such as stochastic depth.\n\nWhen it comes to EfficientNet AdvProp it have the whole solution implemented in the architectures.\n\"Key to our method is the usage of a separate auxiliary batch norm for adversarial examples, as they have different underlying distributions to normal examples. We show that AdvProp improves a wide range of models on various image recognition tasks and performs better when the models are bigger\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F6b30f473af6276c580dc8474482d8abd%2FSkrmklipp.jpg?generation=1600345942457686&alt=media)\nEfficientNet AdvProp\nAdversarial Examples Improve Image Recognition\nhttps://arxiv.org/pdf/1911.09665.pdf\n\n**Conclusion**:\nFirst and foremost, what an amazing work from the Google Research Team, after reading 10+ research papers, what an achievement to move the area to a new level. A lot of work done in a short time, impressive. 🙏\n\nIf I have not understood the whole thing completely wrong ( disclaimer), if one wants the same results and gap as in the picture in other problems one have to use the noisy student training framework not just the pretrained weights, even though the extra trained images that comes with Noisy Student weights and the noise methods will certainly will help to get the best results in some tasks. Otherwise maybe it's equivalent to use the Imagenet-weights or maybe even better use the EfficientNet AdvProp?\n\nIf one takes time and read the training steps, papers, architectures, pre/post processing etc one can find many great ideas, methods and modules to use in new tasks and problems, methods like augmix, autoaugment, randaugment, hardswish just to name a few.\n\nAugMix\nhttps://arxiv.org/pdf/1912.02781.pdf\nAutoAugment\nhttps://arxiv.org/pdf/1805.09501.pdf\nRandAugment\nhttps://arxiv.org/pdf/1909.13719.pdf\n\nThanks again to the Google Research Team, and will be exciting to see what comes next 🙂\n\n**Update**:\nHere is a comparision for this task, B1ns and B1imgnet after 24 epochs\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.\n\nB1ns\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272\n\nB1imgnet\nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989\n\nBetter ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy with Imagenet.\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.",
    "1015456": "Here is a comparision for this task, B1ns and B1imgnet after 24 epochs\nBoth are identical and are trained with SGD, GEM, SyncBatchNormalization, Arc and with Augmentation.\n\nB1ns \nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,5.365260124206543,0.27216899394989014,4.826018333435059,0.3817462921142578\n1,4.775322437286377,0.3190956711769104,4.431529521942139,0.4186016917228699\n2,4.29673433303833,0.36153286695480347,4.118555068969727,0.4463777244091034\n3,3.913407564163208,0.3987823724746704,3.8740103244781494,0.4709269106388092\n4,3.604736328125,0.43042701482772827,3.6706597805023193,0.4920278489589691\n5,3.34704327583313,0.45897015929222107,3.5063326358795166,0.5086681246757507\n6,3.1319875717163086,0.4835450351238251,3.368483781814575,0.523030698299408\n7,2.9487040042877197,0.5052289962768555,3.2536683082580566,0.534640908241272\n\nB1imgnet \nepoch,loss,sparse_categorical_accuracy,val_loss,val_sparse_categorical_accuracy\n0,4.273810863494873,0.36669015884399414,4.216897010803223,0.4438152313232422\n1,3.8364150524139404,0.40967440605163574,3.932481527328491,0.4705789387226105\n2,3.4913442134857178,0.4462076723575592,3.7162606716156006,0.4907624125480652\n3,3.2166695594787598,0.4768737852573395,3.5417873859405518,0.5079088807106018\n4,2.9916982650756836,0.5034031271934509,3.3971595764160156,0.522587776184082\n5,2.804619550704956,0.5265983939170837,3.280860424041748,0.5353052616119385\n6,2.650902509689331,0.545634925365448,3.17187762260437,0.5454602837562561\n7,2.5167949199676514,0.5631154775619507,3.088815927505493,0.5544131398200989\n\nBetter ratio between train-loss/val-loss with Noisy Student but better val-loss/val-accuracy  with Imagenet.\nIn this case I should have continued with B1ns due to slower converge and train/val ratio but to early to tell, small differences between them, better train a full process and see the results.",
    "1016647": "Do they converge to the same values in the end? Also are you training on a local machine or Colab/Kaggle TPUs?",
    "1017831": "Not trained them to the final yet, out of quota, but now we have a new week I'll continue. I'm using Kaggle, it's a better version of TPU and faster.",
    "1019626": "In previous competitions, I have seen that noisy student performance is always better than imagenet."
  },
  "source": "meta"
}