{
  "id": 165256,
  "title": "Advantage of noisy-student",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165256",
  "author_name": "Vishnu R",
  "post_date": "2020-07-09T03:21:56.160000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>I came to know about a weight called <strong>noisy-student</strong> while going through <a href=\"https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/comments\">a kernel</a> by @sgladysh </p>\n\n<p>What is the advantage of using <strong>noisy-student</strong> over <strong>imagenet</strong>(default weight) of efficientnet?</p>\n\n<p>Ref: <a href=\"https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook\">https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook</a></p>",
  "messages": [
    {
      "id": 921048,
      "postDate": "2020-07-09T03:46:03.897Z",
      "content": "<p>Best place to start really is <a href=\"https://arxiv.org/abs/1911.04252\">the paper</a> itself. In short, noisy student style training extends the idea of <a href=\"https://medium.com/@nainaakash012/rethinking-pre-training-and-self-training-53d489b53cbc\">self-training</a> and <a href=\"https://arxiv.org/abs/1503.02531\">model distillation</a> with the use of equal-or-larger student models and noise added to the student during learning.</p>\n\n<p>All the official tensorflow efficientnet <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">noisy student tpu chechpoints</a> additionally come with <a href=\"https://arxiv.org/abs/1909.13719\">randaugment</a> nowadays. On the other hand, if you're a pytorch user, I have a PR open on <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/pull/201\">EfficientNet-PyTorch</a>, which has code you can use to convert the tf checkpoints to .pth for use with your experiments.</p>\n\n<p>Regarding this competition's dataset, compared to vanilla efficientnet (from my tests) you get a very small gain by using noisy student. Take it with a grain of salt, but I think the gain is smaller than with regular CV competitions because a lot of the convolutional weights have to be trained 30-40 epochs to be able to ID the stenography changes, so it's harder to keep all the benefit from the pretrained model. Might be worth-while experimenting with freezing some layers rather than train the whole model.</p>",
      "rawMarkdown": "Best place to start really is [the paper](https://arxiv.org/abs/1911.04252) itself. In short, noisy student style training extends the idea of [self-training](https://medium.com/@nainaakash012/rethinking-pre-training-and-self-training-53d489b53cbc) and [model distillation](https://arxiv.org/abs/1503.02531) with the use of equal-or-larger student models and noise added to the student during learning.\n\nAll the official tensorflow efficientnet [noisy student tpu chechpoints](https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet) additionally come with [randaugment](https://arxiv.org/abs/1909.13719) nowadays. On the other hand, if you're a pytorch user, I have a PR open on [EfficientNet-PyTorch](https://github.com/lukemelas/EfficientNet-PyTorch/pull/201), which has code you can use to convert the tf checkpoints to .pth for use with your experiments.\n\nRegarding this competition's dataset, compared to vanilla efficientnet (from my tests) you get a very small gain by using noisy student. Take it with a grain of salt, but I think the gain is smaller than with regular CV competitions because a lot of the convolutional weights have to be trained 30-40 epochs to be able to ID the stenography changes, so it's harder to keep all the benefit from the pretrained model. Might be worth-while experimenting with freezing some layers rather than train the whole model.",
      "votes": 7,
      "replies": [
        {
          "id": 921136,
          "postDate": "2020-07-09T05:07:22.050Z",
          "content": "<p>Thank you <a href=\"/authman\">@authman</a> for the detailed explanation. I will try to freeze some layers and try.</p>",
          "rawMarkdown": "Thank you @authman for the detailed explanation. I will try to freeze some layers and try.",
          "votes": 1
        }
      ]
    },
    {
      "id": 923021,
      "postDate": "2020-07-10T13:39:21.620Z",
      "content": "<p><a href=\"/vishnurapps\">@vishnurapps</a> \n<a href=\"https://github.com/qubvel/efficientnet#models\">From</a>:</p>\n\n<p>| Architecture   | top1* Imagenet| top1* Noisy-Student| \n| -------------- | :----: |:---:|\n| EfficientNetB0 | 0.772  |0.788|\n| EfficientNetB1 | 0.791  |0.815|\n| EfficientNetB2 | 0.802  |0.824|\n| EfficientNetB3 | 0.816  |0.841|\n| EfficientNetB4 | 0.830  |0.853|\n| EfficientNetB5 | 0.837  |0.861|\n| EfficientNetB6 | 0.841  |0.864|\n| EfficientNetB7 | 0.844  |0.869|</p>",
      "rawMarkdown": "@vishnurapps \n[From](https://github.com/qubvel/efficientnet#models):\n\n| Architecture   | top1* Imagenet| top1* Noisy-Student| \n| -------------- | :----: |:---:|\n| EfficientNetB0 | 0.772  |0.788|\n| EfficientNetB1 | 0.791  |0.815|\n| EfficientNetB2 | 0.802  |0.824|\n| EfficientNetB3 | 0.816  |0.841|\n| EfficientNetB4 | 0.830  |0.853|\n| EfficientNetB5 | 0.837  |0.861|\n| EfficientNetB6 | 0.841  |0.864|\n| EfficientNetB7 | 0.844  |0.869|",
      "votes": 1
    },
    {
      "id": 921033,
      "postDate": "2020-07-09T03:21:56.160Z",
      "content": "<p>Hi All,</p>\n\n<p>I came to know about a weight called <strong>noisy-student</strong> while going through <a href=\"https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/comments\">a kernel</a> by @sgladysh </p>\n\n<p>What is the advantage of using <strong>noisy-student</strong> over <strong>imagenet</strong>(default weight) of efficientnet?</p>\n\n<p>Ref: <a href=\"https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook\">https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook</a></p>",
      "rawMarkdown": "Hi All,\n\nI came to know about a weight called **noisy-student** while going through [a kernel](https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/comments) by @sgladysh \n\nWhat is the advantage of using **noisy-student** over **imagenet**(default weight) of efficientnet?\n\nRef: https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook",
      "votes": 2
    },
    {
      "id": 923982,
      "postDate": "2020-07-11T07:58:23.297Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 921048,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-07-09T03:46:03.897000",
      "content": "<p>Best place to start really is <a href=\"https://arxiv.org/abs/1911.04252\">the paper</a> itself. In short, noisy student style training extends the idea of <a href=\"https://medium.com/@nainaakash012/rethinking-pre-training-and-self-training-53d489b53cbc\">self-training</a> and <a href=\"https://arxiv.org/abs/1503.02531\">model distillation</a> with the use of equal-or-larger student models and noise added to the student during learning.</p>\n\n<p>All the official tensorflow efficientnet <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">noisy student tpu chechpoints</a> additionally come with <a href=\"https://arxiv.org/abs/1909.13719\">randaugment</a> nowadays. On the other hand, if you're a pytorch user, I have a PR open on <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/pull/201\">EfficientNet-PyTorch</a>, which has code you can use to convert the tf checkpoints to .pth for use with your experiments.</p>\n\n<p>Regarding this competition's dataset, compared to vanilla efficientnet (from my tests) you get a very small gain by using noisy student. Take it with a grain of salt, but I think the gain is smaller than with regular CV competitions because a lot of the convolutional weights have to be trained 30-40 epochs to be able to ID the stenography changes, so it's harder to keep all the benefit from the pretrained model. Might be worth-while experimenting with freezing some layers rather than train the whole model.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 921136,
          "author_name": "Vishnu R",
          "author_url": "",
          "post_date": "2020-07-09T05:07:22.050000",
          "content": "<p>Thank you <a href=\"/authman\">@authman</a> for the detailed explanation. I will try to freeze some layers and try.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 923021,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-10T13:39:21.620000",
      "content": "<p><a href=\"/vishnurapps\">@vishnurapps</a> \n<a href=\"https://github.com/qubvel/efficientnet#models\">From</a>:</p>\n\n<p>| Architecture   | top1* Imagenet| top1* Noisy-Student| \n| -------------- | :----: |:---:|\n| EfficientNetB0 | 0.772  |0.788|\n| EfficientNetB1 | 0.791  |0.815|\n| EfficientNetB2 | 0.802  |0.824|\n| EfficientNetB3 | 0.816  |0.841|\n| EfficientNetB4 | 0.830  |0.853|\n| EfficientNetB5 | 0.837  |0.861|\n| EfficientNetB6 | 0.841  |0.864|\n| EfficientNetB7 | 0.844  |0.869|</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923982,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T07:58:23.297000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "921048": "Best place to start really is [the paper](https://arxiv.org/abs/1911.04252) itself. In short, noisy student style training extends the idea of [self-training](https://medium.com/@nainaakash012/rethinking-pre-training-and-self-training-53d489b53cbc) and [model distillation](https://arxiv.org/abs/1503.02531) with the use of equal-or-larger student models and noise added to the student during learning.\n\nAll the official tensorflow efficientnet [noisy student tpu chechpoints](https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet) additionally come with [randaugment](https://arxiv.org/abs/1909.13719) nowadays. On the other hand, if you're a pytorch user, I have a PR open on [EfficientNet-PyTorch](https://github.com/lukemelas/EfficientNet-PyTorch/pull/201), which has code you can use to convert the tf checkpoints to .pth for use with your experiments.\n\nRegarding this competition's dataset, compared to vanilla efficientnet (from my tests) you get a very small gain by using noisy student. Take it with a grain of salt, but I think the gain is smaller than with regular CV competitions because a lot of the convolutional weights have to be trained 30-40 epochs to be able to ID the stenography changes, so it's harder to keep all the benefit from the pretrained model. Might be worth-while experimenting with freezing some layers rather than train the whole model.",
    "923021": "@vishnurapps \n[From](https://github.com/qubvel/efficientnet#models):\n\n| Architecture   | top1* Imagenet| top1* Noisy-Student| \n| -------------- | :----: |:---:|\n| EfficientNetB0 | 0.772  |0.788|\n| EfficientNetB1 | 0.791  |0.815|\n| EfficientNetB2 | 0.802  |0.824|\n| EfficientNetB3 | 0.816  |0.841|\n| EfficientNetB4 | 0.830  |0.853|\n| EfficientNetB5 | 0.837  |0.861|\n| EfficientNetB6 | 0.841  |0.864|\n| EfficientNetB7 | 0.844  |0.869|",
    "921033": "Hi All,\n\nI came to know about a weight called **noisy-student** while going through [a kernel](https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/comments) by @sgladysh \n\nWhat is the advantage of using **noisy-student** over **imagenet**(default weight) of efficientnet?\n\nRef: https://www.kaggle.com/sgladysh/alaska2-steganalysis-tpu-efficientnet-b7-b6-b5-b4/notebook",
    "923982": ""
  }
}