{
  "id": 249175,
  "title": "【rerun】[CV: 0.983 -> 0.836, LB: 0.974 -> 0.727] New Vision Transformer: VOLO",
  "url": "/competitions/seti-breakthrough-listen/discussion/249175",
  "author_name": "Tawara",
  "post_date": "2021-06-27T02:01:26.825000",
  "votes": 49,
  "comment_count": 18,
  "views": 0,
  "content": "<p>A new paper on Vision Transfomer was submitted to arXiv three days ago. It's called <em>VOLO: Vision Outlooker</em>.</p>\n<h5>VOLO: Vision Outlooker for Visual Recognition [<a href=\"https://arxiv.org/abs/2106.13112\" target=\"_blank\">arXiv:2106.13112</a>]</h5>\n<p>　Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan</p>\n<p>VOLO-D5, which is the largest in the VOLO models, achieved 87.1% top-1 accuracy on ImageNet-1K classification <strong>without using any extra training data</strong>.</p>\n<p><img src=\"https://raw.githubusercontent.com/sail-sg/volo/main/figures/compare.png\" alt=\"Comparison with CaiT and NFNet\"></p>\n<p>To our delight, the implementation and pretrained weights has already been published by the authors [<a href=\"https://github.com/sail-sg/volo\" target=\"_blank\">GitHub</a>] .<br>\nI've uploaded them <strong>on Kaggle Datasets</strong>. Now you can use them from Kaggle Notebooks:<br>\n<a href=\"https://www.kaggle.com/ttahara/volo-package\" target=\"_blank\">https://www.kaggle.com/ttahara/volo-package</a></p>\n<p>I've published a new baseline using <em>VOLO-D1</em> as an example of use it. The experimental setup is roughly the same as the <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\" target=\"_blank\">my previous resnet18d baseline</a>.</p>\n<ul>\n<li>Training(<a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68249988\" target=\"_blank\">fold0</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68250008\" target=\"_blank\">fold1</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281727\" target=\"_blank\">fold2</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281737\" target=\"_blank\">fold3</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68335611\" target=\"_blank\">fold4</a>)</li>\n<li><a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-inference?scriptVersionId=68361639\" target=\"_blank\">Inference Notebook</a></li>\n</ul>\n<p>This baseline got  0.727 <strong>with image_size=256x256</strong>.<br>\nI'm thinking that it might achieve better score and could be useful for ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>fold ＼ model</th>\n<th>ResNet18D(<strong>320x320</strong>)</th>\n<th><em>VOLO-D1</em>(<strong>256x256</strong>)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td> 0.8358</td>\n<td> 0.8501</td>\n</tr>\n<tr>\n<td>1</td>\n<td> 0.8147</td>\n<td> 0.8266</td>\n</tr>\n<tr>\n<td>2</td>\n<td> 0.8242</td>\n<td> 0.8328</td>\n</tr>\n<tr>\n<td>3</td>\n<td> 0.8250</td>\n<td> 0.8354</td>\n</tr>\n<tr>\n<td>4</td>\n<td> 0.8276</td>\n<td> 0.8368</td>\n</tr>\n<tr>\n<td>OOF</td>\n<td> 0.8248</td>\n<td> 0.8362</td>\n</tr>\n<tr>\n<td>Public LB</td>\n<td> 0.726</td>\n<td> 0.727</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nI published the notebooks before updating competition data for rapid reporting.  <br>\nI will re-run these notebooks after update ;)</p>",
  "messages": [
    {
      "id": 1366604,
      "postDate": "2021-06-27T02:01:26.827Z",
      "content": "<p>A new paper on Vision Transfomer was submitted to arXiv three days ago. It's called <em>VOLO: Vision Outlooker</em>.</p>\n<h5>VOLO: Vision Outlooker for Visual Recognition [<a href=\"https://arxiv.org/abs/2106.13112\" target=\"_blank\">arXiv:2106.13112</a>]</h5>\n<p>　Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan</p>\n<p>VOLO-D5, which is the largest in the VOLO models, achieved 87.1% top-1 accuracy on ImageNet-1K classification <strong>without using any extra training data</strong>.</p>\n<p><img src=\"https://raw.githubusercontent.com/sail-sg/volo/main/figures/compare.png\" alt=\"Comparison with CaiT and NFNet\"></p>\n<p>To our delight, the implementation and pretrained weights has already been published by the authors [<a href=\"https://github.com/sail-sg/volo\" target=\"_blank\">GitHub</a>] .<br>\nI've uploaded them <strong>on Kaggle Datasets</strong>. Now you can use them from Kaggle Notebooks:<br>\n<a href=\"https://www.kaggle.com/ttahara/volo-package\" target=\"_blank\">https://www.kaggle.com/ttahara/volo-package</a></p>\n<p>I've published a new baseline using <em>VOLO-D1</em> as an example of use it. The experimental setup is roughly the same as the <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\" target=\"_blank\">my previous resnet18d baseline</a>.</p>\n<ul>\n<li>Training(<a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68249988\" target=\"_blank\">fold0</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68250008\" target=\"_blank\">fold1</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281727\" target=\"_blank\">fold2</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281737\" target=\"_blank\">fold3</a>, <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68335611\" target=\"_blank\">fold4</a>)</li>\n<li><a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-inference?scriptVersionId=68361639\" target=\"_blank\">Inference Notebook</a></li>\n</ul>\n<p>This baseline got  0.727 <strong>with image_size=256x256</strong>.<br>\nI'm thinking that it might achieve better score and could be useful for ensemble.</p>\n<table>\n<thead>\n<tr>\n<th>fold ＼ model</th>\n<th>ResNet18D(<strong>320x320</strong>)</th>\n<th><em>VOLO-D1</em>(<strong>256x256</strong>)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td> 0.8358</td>\n<td> 0.8501</td>\n</tr>\n<tr>\n<td>1</td>\n<td> 0.8147</td>\n<td> 0.8266</td>\n</tr>\n<tr>\n<td>2</td>\n<td> 0.8242</td>\n<td> 0.8328</td>\n</tr>\n<tr>\n<td>3</td>\n<td> 0.8250</td>\n<td> 0.8354</td>\n</tr>\n<tr>\n<td>4</td>\n<td> 0.8276</td>\n<td> 0.8368</td>\n</tr>\n<tr>\n<td>OOF</td>\n<td> 0.8248</td>\n<td> 0.8362</td>\n</tr>\n<tr>\n<td>Public LB</td>\n<td> 0.726</td>\n<td> 0.727</td>\n</tr>\n</tbody>\n</table>\n<p><br><br>\nI published the notebooks before updating competition data for rapid reporting.  <br>\nI will re-run these notebooks after update ;)</p>",
      "rawMarkdown": "A new paper on Vision Transfomer was submitted to arXiv three days ago. It's called _VOLO: Vision Outlooker_.\n\n##### VOLO: Vision Outlooker for Visual Recognition [[arXiv:2106.13112](https://arxiv.org/abs/2106.13112)]   \n　Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan\n\nVOLO-D5, which is the largest in the VOLO models, achieved 87.1% top-1 accuracy on ImageNet-1K classification **without using any extra training data**.\n\n![Comparison with CaiT and NFNet](https://raw.githubusercontent.com/sail-sg/volo/main/figures/compare.png)\n\nTo our delight, the implementation and pretrained weights has already been published by the authors [[GitHub](https://github.com/sail-sg/volo)] .\nI've uploaded them **on Kaggle Datasets**. Now you can use them from Kaggle Notebooks:\nhttps://www.kaggle.com/ttahara/volo-package\n\nI've published a new baseline using _VOLO-D1_ as an example of use it. The experimental setup is roughly the same as the [my previous resnet18d baseline](https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline).\n\n* Training([fold0](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68249988), [fold1](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68250008), [fold2](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281727), [fold3](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281737), [fold4](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68335611))\n* [Inference Notebook](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-inference?scriptVersionId=68361639)\n\nThis baseline got ~~0.974~~ 0.727 **with image_size=256x256**.\nI'm thinking that it might achieve better score and could be useful for ensemble.\n\n| fold ＼ model | ResNet18D(**320x320**) | _VOLO-D1_(**256x256**)|\n|:-------------:|:-----------------------:|:----------------------:|\n| 0 | ~~0.9827~~ 0.8358 | ~~0.9817~~ 0.8501 |\n| 1 | ~~0.9817~~ 0.8147 | ~~0.9833~~ 0.8266 |\n| 2 | ~~0.9843~~ 0.8242 | ~~0.9835~~ 0.8328 |\n| 3 | ~~0.9828~~ 0.8250 | ~~0.9843~~ 0.8354 |\n| 4 | ~~0.9842~~ 0.8276 | ~~0.9835~~ 0.8368 |\n| OOF | ~~0.9830~~ 0.8248 | ~~0.9831~~ 0.8362 |\n| Public LB | ~~0.974~~ 0.726 | ~~0.974~~ 0.727 |\n\n<br>\nI published the notebooks before updating competition data for rapid reporting.  \nI will re-run these notebooks after update ;)",
      "votes": 49
    },
    {
      "id": 1371431,
      "postDate": "2021-07-01T02:17:09.180Z",
      "content": "<p>model:volo_d1-384<br>\nimage_size:640<br>\n1fold:cv 99.01 lb 97.7</p>",
      "rawMarkdown": "model:volo_d1-384\nimage_size:640\n1fold:cv 99.01 lb 97.7",
      "votes": 4,
      "replies": [
        {
          "id": 1372650,
          "postDate": "2021-07-01T22:50:58.747Z",
          "content": "<p>Thank you for sharing.  <br>\nThe result is worse than I expected 🤔</p>\n<p>If you don't mind, can you tell me about your experimental setup?</p>",
          "rawMarkdown": "Thank you for sharing.  \nThe result is worse than I expected 🤔\n\nIf you don't mind, can you tell me about your experimental setup?",
          "votes": 1
        },
        {
          "id": 1372747,
          "postDate": "2021-07-02T01:48:52.380Z",
          "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <br>\nlr=1e-4<br>\nepochs=25<br>\ncutmix+mixup 1:1<br>\ndouble rtx3090 1fold 50000 + s<br>\nObviously, GPU is a little hard to run. It needs TPU!</p>",
          "rawMarkdown": "@ttahara \nlr=1e-4\nepochs=25\ncutmix+mixup 1:1\ndouble rtx3090 1fold 50000 + s\nObviously, GPU is a little hard to run. It needs TPU!",
          "votes": 1
        },
        {
          "id": 1375182,
          "postDate": "2021-07-04T01:12:41.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> <br>\nThanks. Sorry for late reply.  <br>\nYes, ViTs including this model require more and more GPUs or TPUs 😓</p>\n<p>Here are a few my thoughts for training VOLO.</p>\n<h5>Image Size</h5>\n<p>640x640 is so large … I think it is good to start from smaller size (384x384 or 512x512).</p>\n<h5>learning rate and weight decay rate</h5>\n<p>In paper, authors recommend us to set learning rate 1.6e-3 ×<code>batch_size</code> / 1024 for VOLO-D1 and weight decay rate 5.0e-2 for all VOLO models. (see Table 3)</p>\n<h5>regularization techniques</h5>\n<p>In addition to normal data augmentations, authors used <em>Stochastic Depth</em> and  <em>Token Labeling objective with MixToken</em>.    </p>\n<p>I don't know much about the latter, and it seems a bit difficult to use this because we need to prepare token labels.  <br>\nOn the other hand, the former can be easily used by setting a hyperparameter <strong><em>drop_path_rate</em></strong>.</p>\n<pre><code>import timm\nfrom volo.models import volo_d1  # register model to timm\nfrom volo.utils import load_pretrained_weights\n\npretrained_model_path =  \"../input/volo-package/d1_384_85.2.pth.tar\"\n\nmodel = timm.create_model(\n    \"volo_d1\", img_size=384,\n    mix_token=False, return_dense=False, drop_path_rate=0.1)\n\nload_pretrained_weights(model, pretrained_model_path, strict=False)\n</code></pre>",
          "rawMarkdown": "@zhangeng \nThanks. Sorry for late reply.  \nYes, ViTs including this model require more and more GPUs or TPUs 😓\n\nHere are a few my thoughts for training VOLO.\n\n##### Image Size\n640x640 is so large ... I think it is good to start from smaller size (384x384 or 512x512).\n\n##### learning rate and weight decay rate\nIn paper, authors recommend us to set learning rate 1.6e-3 ×`batch_size` / 1024 for VOLO-D1 and weight decay rate 5.0e-2 for all VOLO models. (see Table 3)\n\n##### regularization techniques\nIn addition to normal data augmentations, authors used _Stochastic Depth_ and  _Token Labeling objective with MixToken_.    \n\nI don't know much about the latter, and it seems a bit difficult to use this because we need to prepare token labels.  \nOn the other hand, the former can be easily used by setting a hyperparameter **_drop_path_rate_**.\n\n```\nimport timm\nfrom volo.models import volo_d1  # register model to timm\nfrom volo.utils import load_pretrained_weights\n\npretrained_model_path =  \"../input/volo-package/d1_384_85.2.pth.tar\"\n\nmodel = timm.create_model(\n    \"volo_d1\", img_size=384,\n    mix_token=False, return_dense=False, drop_path_rate=0.1)\n\nload_pretrained_weights(model, pretrained_model_path, strict=False)\n```",
          "votes": 4
        },
        {
          "id": 1376791,
          "postDate": "2021-07-05T10:35:54.837Z",
          "content": "<p>I learned two things<br>\n1.lr=&gt; learning rate 1.6e-3 ×batch_size / 1024<br>\n2.setting a hyperparameter drop_path_rate.<br>\nThank you for your help. After the data is updated, I must try again to change the hyperparameter in these two places</p>",
          "rawMarkdown": "\nI learned two things\n1.lr=> learning rate 1.6e-3 ×batch_size / 1024\n2.setting a hyperparameter drop_path_rate.\n\nThank you for your help. After the data is updated, I must try again to change the hyperparameter in these two places\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1395720,
      "postDate": "2021-07-21T13:47:19.067Z",
      "content": "<p><strong>Update</strong>: I ran training &amp; inference notebooks again for the updated competition dataset.</p>",
      "rawMarkdown": "**Update**: I ran training & inference notebooks again for the updated competition dataset.",
      "votes": 2
    },
    {
      "id": 1370815,
      "postDate": "2021-06-30T12:25:12.037Z",
      "content": "<p>I am trying to implement VOLO in tensorflow. Any Help would be appreciated. You can Dm if you want to help (sorry if this is Off topic).</p>",
      "rawMarkdown": "I am trying to implement VOLO in tensorflow. Any Help would be appreciated. You can Dm if you want to help (sorry if this is Off topic).",
      "votes": 2,
      "replies": [
        {
          "id": 1371433,
          "postDate": "2021-07-01T02:19:04.107Z",
          "content": "<p>hello!Even if the tensorflow Volo model is implemented, how to transfer the weight is still a problem!</p>",
          "rawMarkdown": "hello!Even if the tensorflow Volo model is implemented, how to transfer the weight is still a problem!"
        },
        {
          "id": 1371453,
          "postDate": "2021-07-01T02:49:43.260Z",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> that surely is a problem but we can retrain the model on ImageNet maybe</p>",
          "rawMarkdown": "@zhangeng that surely is a problem but we can retrain the model on ImageNet maybe",
          "votes": 1
        }
      ]
    },
    {
      "id": 1408284,
      "postDate": "2021-08-02T12:00:52.420Z",
      "content": "<p>When I apply your code:</p>\n<pre><code>model_name = base_name.split(\"-\")[0]\nassert timm.is_model(model_name),  \"you can use only models in timm.\"\n</code></pre>\n<p>I found that there is no volo_d1 in my timm library. How can I add volo to the timm library?</p>",
      "rawMarkdown": "When I apply your code:\n```python\nmodel_name = base_name.split(\"-\")[0]\nassert timm.is_model(model_name),  \"you can use only models in timm.\"\n```\n I found that there is no volo_d1 in my timm library. How can I add volo to the timm library?"
    },
    {
      "id": 1371700,
      "postDate": "2021-07-01T07:24:58.717Z",
      "content": "<p>interesting</p>",
      "rawMarkdown": "interesting"
    },
    {
      "id": 1367594,
      "postDate": "2021-06-27T22:35:47.410Z",
      "content": "<p>thx - VERY interesting!!</p>\n<p>Did you do any testing with droprate? 0.5 seems rather high….</p>",
      "rawMarkdown": "thx - VERY interesting!!\n\nDid you do any testing with droprate? 0.5 seems rather high....",
      "replies": [
        {
          "id": 1369810,
          "postDate": "2021-06-29T15:52:03.503Z",
          "content": "<p>Thanks.</p>\n<p>Which drop_rate are you referring to?</p>",
          "rawMarkdown": "Thanks.\n\nWhich drop_rate are you referring to?",
          "votes": 2
        },
        {
          "id": 1370521,
          "postDate": "2021-06-30T08:08:10.603Z",
          "content": "<p>sorry for being not exact. I meant <br>\n nn.ReLU(), nn.Dropout(0.5),]) in the head of the volo model.<br>\nEfficientNets are trained with 0.2-0.4 dropout in the head as proposed by the paper. So 0.5 seems quite a heavy regularisation.</p>",
          "rawMarkdown": "sorry for being not exact. I meant \n nn.ReLU(), nn.Dropout(0.5),]) in the head of the volo model.\nEfficientNets are trained with 0.2-0.4 dropout in the head as proposed by the paper. So 0.5 seems quite a heavy regularisation.",
          "votes": 1
        },
        {
          "id": 1372648,
          "postDate": "2021-07-01T22:48:10.737Z",
          "content": "<p>OK, I see.</p>\n<pre><code># # prepare head classifier\nif dims_head[0] is None:\n    dims_head[0] = in_features\n\nlayers_list = []\nfor i in range(len(dims_head) - 2):\n    in_dim, out_dim = dims_head[i: i + 2]\n    layers_list.extend([\n        nn.Linear(in_dim, out_dim),\n        nn.ReLU(), nn.Dropout(0.5),])\nlayers_list.append(\n    nn.Linear(dims_head[-2], dims_head[-1]))\nself.head = nn.Sequential(*layers_list)\n</code></pre>\n<p>I write the code for customizing head classifier.  <br>\nBut I used <strong>a linear layer as head</strong> in this notebook (by setting <code>dims_head = [384, 1]</code> )</p>",
          "rawMarkdown": "OK, I see.\n\n```\n# # prepare head classifier\nif dims_head[0] is None:\n    dims_head[0] = in_features\n\nlayers_list = []\nfor i in range(len(dims_head) - 2):\n    in_dim, out_dim = dims_head[i: i + 2]\n    layers_list.extend([\n        nn.Linear(in_dim, out_dim),\n        nn.ReLU(), nn.Dropout(0.5),])\nlayers_list.append(\n    nn.Linear(dims_head[-2], dims_head[-1]))\nself.head = nn.Sequential(*layers_list)\n```\n\nI write the code for customizing head classifier.  \nBut I used **a linear layer as head** in this notebook (by setting `dims_head = [384, 1]` )",
          "votes": 1
        }
      ]
    },
    {
      "id": 1367128,
      "postDate": "2021-06-27T12:54:50.677Z",
      "content": "<p>Is there any tensorflow implementation out There ?</p>",
      "rawMarkdown": "Is there any tensorflow implementation out There ?",
      "replies": [
        {
          "id": 1367138,
          "postDate": "2021-06-27T13:01:07.797Z",
          "content": "<p>As far as I know, any tensorflow implementation doesn't exist. It's only been three days since the paper was published.</p>\n<p>I used the official pytorch implementation by authors.</p>",
          "rawMarkdown": "As far as I know, any tensorflow implementation doesn't exist. It's only been three days since the paper was published.\n\nI used the official pytorch implementation by authors.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1369049,
      "postDate": "2021-06-29T05:35:28.740Z",
      "content": "<p>cool! thanks for sharing!</p>",
      "rawMarkdown": "cool! thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1371431,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-01T02:17:09.180000",
      "content": "<p>model:volo_d1-384<br>\nimage_size:640<br>\n1fold:cv 99.01 lb 97.7</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1372650,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-07-01T22:50:58.747000",
          "content": "<p>Thank you for sharing.  <br>\nThe result is worse than I expected 🤔</p>\n<p>If you don't mind, can you tell me about your experimental setup?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1372747,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-07-02T01:48:52.380000",
          "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <br>\nlr=1e-4<br>\nepochs=25<br>\ncutmix+mixup 1:1<br>\ndouble rtx3090 1fold 50000 + s<br>\nObviously, GPU is a little hard to run. It needs TPU!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1375182,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-07-04T01:12:41.993000",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> <br>\nThanks. Sorry for late reply.  <br>\nYes, ViTs including this model require more and more GPUs or TPUs 😓</p>\n<p>Here are a few my thoughts for training VOLO.</p>\n<h5>Image Size</h5>\n<p>640x640 is so large … I think it is good to start from smaller size (384x384 or 512x512).</p>\n<h5>learning rate and weight decay rate</h5>\n<p>In paper, authors recommend us to set learning rate 1.6e-3 ×<code>batch_size</code> / 1024 for VOLO-D1 and weight decay rate 5.0e-2 for all VOLO models. (see Table 3)</p>\n<h5>regularization techniques</h5>\n<p>In addition to normal data augmentations, authors used <em>Stochastic Depth</em> and  <em>Token Labeling objective with MixToken</em>.    </p>\n<p>I don't know much about the latter, and it seems a bit difficult to use this because we need to prepare token labels.  <br>\nOn the other hand, the former can be easily used by setting a hyperparameter <strong><em>drop_path_rate</em></strong>.</p>\n<pre><code>import timm\nfrom volo.models import volo_d1  # register model to timm\nfrom volo.utils import load_pretrained_weights\n\npretrained_model_path =  \"../input/volo-package/d1_384_85.2.pth.tar\"\n\nmodel = timm.create_model(\n    \"volo_d1\", img_size=384,\n    mix_token=False, return_dense=False, drop_path_rate=0.1)\n\nload_pretrained_weights(model, pretrained_model_path, strict=False)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1376791,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-07-05T10:35:54.837000",
          "content": "<p>I learned two things<br>\n1.lr=&gt; learning rate 1.6e-3 ×batch_size / 1024<br>\n2.setting a hyperparameter drop_path_rate.<br>\nThank you for your help. After the data is updated, I must try again to change the hyperparameter in these two places</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1395720,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-07-21T13:47:19.067000",
      "content": "<p><strong>Update</strong>: I ran training &amp; inference notebooks again for the updated competition dataset.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1370815,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-30T12:25:12.037000",
      "content": "<p>I am trying to implement VOLO in tensorflow. Any Help would be appreciated. You can Dm if you want to help (sorry if this is Off topic).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1371433,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-07-01T02:19:04.107000",
          "content": "<p>hello!Even if the tensorflow Volo model is implemented, how to transfer the weight is still a problem!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1371453,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-01T02:49:43.260000",
          "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> that surely is a problem but we can retrain the model on ImageNet maybe</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1408284,
      "author_name": "Gainover",
      "author_url": "",
      "post_date": "2021-08-02T12:00:52.420000",
      "content": "<p>When I apply your code:</p>\n<pre><code>model_name = base_name.split(\"-\")[0]\nassert timm.is_model(model_name),  \"you can use only models in timm.\"\n</code></pre>\n<p>I found that there is no volo_d1 in my timm library. How can I add volo to the timm library?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1371700,
      "author_name": "karthikeyan",
      "author_url": "",
      "post_date": "2021-07-01T07:24:58.717000",
      "content": "<p>interesting</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1367594,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2021-06-27T22:35:47.410000",
      "content": "<p>thx - VERY interesting!!</p>\n<p>Did you do any testing with droprate? 0.5 seems rather high….</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1369810,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-29T15:52:03.503000",
          "content": "<p>Thanks.</p>\n<p>Which drop_rate are you referring to?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1370521,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2021-06-30T08:08:10.603000",
          "content": "<p>sorry for being not exact. I meant <br>\n nn.ReLU(), nn.Dropout(0.5),]) in the head of the volo model.<br>\nEfficientNets are trained with 0.2-0.4 dropout in the head as proposed by the paper. So 0.5 seems quite a heavy regularisation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1372648,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-07-01T22:48:10.737000",
          "content": "<p>OK, I see.</p>\n<pre><code># # prepare head classifier\nif dims_head[0] is None:\n    dims_head[0] = in_features\n\nlayers_list = []\nfor i in range(len(dims_head) - 2):\n    in_dim, out_dim = dims_head[i: i + 2]\n    layers_list.extend([\n        nn.Linear(in_dim, out_dim),\n        nn.ReLU(), nn.Dropout(0.5),])\nlayers_list.append(\n    nn.Linear(dims_head[-2], dims_head[-1]))\nself.head = nn.Sequential(*layers_list)\n</code></pre>\n<p>I write the code for customizing head classifier.  <br>\nBut I used <strong>a linear layer as head</strong> in this notebook (by setting <code>dims_head = [384, 1]</code> )</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1367128,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-27T12:54:50.677000",
      "content": "<p>Is there any tensorflow implementation out There ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1367138,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-27T13:01:07.797000",
          "content": "<p>As far as I know, any tensorflow implementation doesn't exist. It's only been three days since the paper was published.</p>\n<p>I used the official pytorch implementation by authors.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1369049,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2021-06-29T05:35:28.740000",
      "content": "<p>cool! thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1366604": "A new paper on Vision Transfomer was submitted to arXiv three days ago. It's called _VOLO: Vision Outlooker_.\n\n##### VOLO: Vision Outlooker for Visual Recognition [[arXiv:2106.13112](https://arxiv.org/abs/2106.13112)]   \n　Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan\n\nVOLO-D5, which is the largest in the VOLO models, achieved 87.1% top-1 accuracy on ImageNet-1K classification **without using any extra training data**.\n\n![Comparison with CaiT and NFNet](https://raw.githubusercontent.com/sail-sg/volo/main/figures/compare.png)\n\nTo our delight, the implementation and pretrained weights has already been published by the authors [[GitHub](https://github.com/sail-sg/volo)] .\nI've uploaded them **on Kaggle Datasets**. Now you can use them from Kaggle Notebooks:\nhttps://www.kaggle.com/ttahara/volo-package\n\nI've published a new baseline using _VOLO-D1_ as an example of use it. The experimental setup is roughly the same as the [my previous resnet18d baseline](https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline).\n\n* Training([fold0](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68249988), [fold1](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68250008), [fold2](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281727), [fold3](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68281737), [fold4](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-training?scriptVersionId=68335611))\n* [Inference Notebook](https://www.kaggle.com/ttahara/rerun-seti-e-t-volo-d1-baseline-inference?scriptVersionId=68361639)\n\nThis baseline got ~~0.974~~ 0.727 **with image_size=256x256**.\nI'm thinking that it might achieve better score and could be useful for ensemble.\n\n| fold ＼ model | ResNet18D(**320x320**) | _VOLO-D1_(**256x256**)|\n|:-------------:|:-----------------------:|:----------------------:|\n| 0 | ~~0.9827~~ 0.8358 | ~~0.9817~~ 0.8501 |\n| 1 | ~~0.9817~~ 0.8147 | ~~0.9833~~ 0.8266 |\n| 2 | ~~0.9843~~ 0.8242 | ~~0.9835~~ 0.8328 |\n| 3 | ~~0.9828~~ 0.8250 | ~~0.9843~~ 0.8354 |\n| 4 | ~~0.9842~~ 0.8276 | ~~0.9835~~ 0.8368 |\n| OOF | ~~0.9830~~ 0.8248 | ~~0.9831~~ 0.8362 |\n| Public LB | ~~0.974~~ 0.726 | ~~0.974~~ 0.727 |\n\n<br>\nI published the notebooks before updating competition data for rapid reporting.  \nI will re-run these notebooks after update ;)",
    "1371431": "model:volo_d1-384\nimage_size:640\n1fold:cv 99.01 lb 97.7",
    "1395720": "**Update**: I ran training & inference notebooks again for the updated competition dataset.",
    "1370815": "I am trying to implement VOLO in tensorflow. Any Help would be appreciated. You can Dm if you want to help (sorry if this is Off topic).",
    "1408284": "When I apply your code:\n```python\nmodel_name = base_name.split(\"-\")[0]\nassert timm.is_model(model_name),  \"you can use only models in timm.\"\n```\n I found that there is no volo_d1 in my timm library. How can I add volo to the timm library?",
    "1371700": "interesting",
    "1367594": "thx - VERY interesting!!\n\nDid you do any testing with droprate? 0.5 seems rather high....",
    "1367128": "Is there any tensorflow implementation out There ?",
    "1369049": "cool! thanks for sharing!"
  }
}