{
  "id": 213183,
  "title": "VGG is back. RepVGG ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/213183",
  "author_name": "",
  "post_date": "2021-01-21T19:48:07.323948100Z",
  "votes": 17,
  "comment_count": 5,
  "views": 0,
  "content": "<blockquote>\n  <p>We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3x3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet.</p>\n</blockquote>\n<p><a href=\"https://arxiv.org/pdf/2101.03697v1.pdf\" target=\"_blank\">paper</a><br>\n<a href=\"https://github.com/DingXiaoH/RepVGG\" target=\"_blank\">pytorch code</a></p>",
  "messages": [
    {
      "id": "1163616",
      "postDate": "01/21/2021 19:48:07",
      "content": "<blockquote>\n  <p>We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3x3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet.</p>\n</blockquote>\n<p><a href=\"https://arxiv.org/pdf/2101.03697v1.pdf\" target=\"_blank\">paper</a><br>\n<a href=\"https://github.com/DingXiaoH/RepVGG\" target=\"_blank\">pytorch code</a></p>",
      "rawMarkdown": "> We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3x3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet.\n\n[paper](https://arxiv.org/pdf/2101.03697v1.pdf)\n[pytorch code](https://github.com/DingXiaoH/RepVGG)",
      "votes": null
    },
    {
      "id": "1164051",
      "postDate": "01/22/2021 05:34:27",
      "content": "<p>noticed that too.<br>\nbut have not tried it yet.</p>",
      "rawMarkdown": "noticed that too.\nbut have not tried it yet.",
      "votes": null
    },
    {
      "id": "1164210",
      "postDate": "01/22/2021 08:02:31",
      "content": "<p>i have uploaded all the weight files here : <a href=\"https://www.kaggle.com/mobassir/repvgg\" target=\"_blank\">https://www.kaggle.com/mobassir/repvgg</a><br>\nsoon planning to give it a try, thanks for sharing</p>",
      "rawMarkdown": "i have uploaded all the weight files here : https://www.kaggle.com/mobassir/repvgg\nsoon planning to give it a try, thanks for sharing",
      "votes": null
    },
    {
      "id": "1166092",
      "postDate": "01/23/2021 11:36:34",
      "content": "<p>It would be great, if someone could provide a simple example with some explanation in notebook.</p>",
      "rawMarkdown": "It would be great, if someone could provide a simple example with some explanation in notebook.",
      "votes": null
    },
    {
      "id": "1169323",
      "postDate": "01/25/2021 12:40:19",
      "content": "<p>I tried it, seems not really good in this exact case. </p>",
      "rawMarkdown": "I tried it, seems not really good in this exact case.",
      "votes": null
    },
    {
      "id": "1170079",
      "postDate": "01/26/2021 00:38:27",
      "content": "<p>Same, i tested the model on the validation data and it gives me ~90% accuracy. But when i deploy the model for inference it gives me 51% accuracy, when the B0 version is used. Maybe i am doing something wrong when i load the pretrained model. Btw, i added some extra layers to achieve that performance.</p>\n<p>Can u please tell me how did you train and load the model? I did it this way:</p>\n<p>class RepVggModel(nn.Module):</p>\n<pre><code>def __init__(self, model, n_classes):\n    super().__init__()\n    self.n_classes = n_classes\n    self.model = model\n    self.model.linear = nn.Identity()\n    self.fc1 = nn.Linear(1280, 512) # 1280\n    self.bn = nn.BatchNorm1d(num_features=512) # 4096\n    self.relu = nn.ReLU(True)\n    self.dropout = nn.Dropout(0.5) # 0.3\n    self.linear = nn.Linear(512, self.n_classes) # 2*512\n\n\ndef forward(self, x):\n    features = self.model(x)\n    features = self.fc1(features)\n    features = self.bn(features)\n    x = self.relu(features)\n    x = self.dropout(x)\n    x = self.linear(x)\n    #x = torch.softmax(x, dim = 1)\n\n    return x\n</code></pre>\n<h1>for training</h1>\n<p>n_classes = 5<br>\nmodel = create_RepVGG_B0(deploy=False)<br>\nmodel = RepVggModel(model, n_classes)</p>\n<h1>for inference</h1>\n<p>checkpoint = torch.load(model_name)<br>\nmodel_deploy = create_RepVGG_B0(deploy = False)<br>\nmodel_deploy = RepVggModel(model_deploy, n_classes)</p>\n<p>model_deploy.load_state_dict(checkpoint)<br>\nfor parameter in model_deploy.parameters():<br>\nparameter.requires_grad = False<br>\nmodel_deploy.eval()</p>\n<p>I also did some fine-tuning:</p>\n<p>def freeze_model_parametrs(model, n_layers):<br>\n    print('=====\\nFrozen\\n=====')</p>\n<pre><code>for name, param in list(model.named_parameters())[:n_layers]:\n    param.requires_grad = False\n\nprint('=====\\nTrainable\\n=====')\nfor name, param in list(model.named_parameters())[n_layers:]:\n    param.requires_grad = True\n\npytorch_total_params = sum(p.numel() for p in model.parameters())\npytorch_total_params_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\npytorch_total_params_frozen = sum(p.numel() for p in model.parameters() if p.requires_grad == False)\n\nprint('\\nTotal parameters: ', pytorch_total_params)\nprint('\\nTotal trainable parameters: ', pytorch_total_params_trainable)\nprint('\\nTotal frozen parameters: ', pytorch_total_params_frozen)\n\nparams = filter(lambda p: p.requires_grad, model.parameters())\n\nreturn model, params\n</code></pre>\n<p>For example, if I set deploy = True for inference, the model doesn't load properly, since I added some intermediate layers and the classification output is different.</p>\n<p>Am I doing something wrong in here during inference?</p>\n<p>Cheers!</p>",
      "rawMarkdown": "Same, i tested the model on the validation data and it gives me ~90% accuracy. But when i deploy the model for inference it gives me 51% accuracy, when the B0 version is used. Maybe i am doing something wrong when i load the pretrained model. Btw, i added some extra layers to achieve that performance.\n\nCan u please tell me how did you train and load the model? I did it this way:\n\nclass RepVggModel(nn.Module):\n    \n    def __init__(self, model, n_classes):\n        super().__init__()\n        self.n_classes = n_classes\n        self.model = model\n        self.model.linear = nn.Identity()\n        self.fc1 = nn.Linear(1280, 512) # 1280\n        self.bn = nn.BatchNorm1d(num_features=512) # 4096\n        self.relu = nn.ReLU(True)\n        self.dropout = nn.Dropout(0.5) # 0.3\n        self.linear = nn.Linear(512, self.n_classes) # 2*512\n\n    \n    def forward(self, x):\n        features = self.model(x)\n        features = self.fc1(features)\n        features = self.bn(features)\n        x = self.relu(features)\n        x = self.dropout(x)\n        x = self.linear(x)\n        #x = torch.softmax(x, dim = 1)\n        \n        return x\n\n#for training\nn_classes = 5\nmodel = create_RepVGG_B0(deploy=False)\nmodel = RepVggModel(model, n_classes)\n\n#for inference\ncheckpoint = torch.load(model_name)\nmodel_deploy = create_RepVGG_B0(deploy = False)\nmodel_deploy = RepVggModel(model_deploy, n_classes)\n\nmodel_deploy.load_state_dict(checkpoint)\nfor parameter in model_deploy.parameters():\nparameter.requires_grad = False\nmodel_deploy.eval()\n\nI also did some fine-tuning:\n\ndef freeze_model_parametrs(model, n_layers):\n    print('=====\\nFrozen\\n=====')\n\n    for name, param in list(model.named_parameters())[:n_layers]:\n        param.requires_grad = False\n\n    print('=====\\nTrainable\\n=====')\n    for name, param in list(model.named_parameters())[n_layers:]:\n        param.requires_grad = True\n\n    pytorch_total_params = sum(p.numel() for p in model.parameters())\n    pytorch_total_params_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\n    pytorch_total_params_frozen = sum(p.numel() for p in model.parameters() if p.requires_grad == False)\n                                          \n    print('\\nTotal parameters: ', pytorch_total_params)\n    print('\\nTotal trainable parameters: ', pytorch_total_params_trainable)\n    print('\\nTotal frozen parameters: ', pytorch_total_params_frozen)\n    \n    params = filter(lambda p: p.requires_grad, model.parameters())\n    \n    return model, params\n\nFor example, if I set deploy = True for inference, the model doesn't load properly, since I added some intermediate layers and the classification output is different.\n\nAm I doing something wrong in here during inference?\n\nCheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1164051,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "01/22/2021 05:34:27",
      "content": "<p>noticed that too.<br>\nbut have not tried it yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1164210,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "01/22/2021 08:02:31",
      "content": "<p>i have uploaded all the weight files here : <a href=\"https://www.kaggle.com/mobassir/repvgg\" target=\"_blank\">https://www.kaggle.com/mobassir/repvgg</a><br>\nsoon planning to give it a try, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1166092,
      "author_name": "luisgomes21",
      "author_url": "",
      "post_date": "01/23/2021 11:36:34",
      "content": "<p>It would be great, if someone could provide a simple example with some explanation in notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1169323,
      "author_name": "arli2016",
      "author_url": "",
      "post_date": "01/25/2021 12:40:19",
      "content": "<p>I tried it, seems not really good in this exact case. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1170079,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "01/26/2021 00:38:27",
      "content": "<p>Same, i tested the model on the validation data and it gives me ~90% accuracy. But when i deploy the model for inference it gives me 51% accuracy, when the B0 version is used. Maybe i am doing something wrong when i load the pretrained model. Btw, i added some extra layers to achieve that performance.</p>\n<p>Can u please tell me how did you train and load the model? I did it this way:</p>\n<p>class RepVggModel(nn.Module):</p>\n<pre><code>def __init__(self, model, n_classes):\n    super().__init__()\n    self.n_classes = n_classes\n    self.model = model\n    self.model.linear = nn.Identity()\n    self.fc1 = nn.Linear(1280, 512) # 1280\n    self.bn = nn.BatchNorm1d(num_features=512) # 4096\n    self.relu = nn.ReLU(True)\n    self.dropout = nn.Dropout(0.5) # 0.3\n    self.linear = nn.Linear(512, self.n_classes) # 2*512\n\n\ndef forward(self, x):\n    features = self.model(x)\n    features = self.fc1(features)\n    features = self.bn(features)\n    x = self.relu(features)\n    x = self.dropout(x)\n    x = self.linear(x)\n    #x = torch.softmax(x, dim = 1)\n\n    return x\n</code></pre>\n<h1>for training</h1>\n<p>n_classes = 5<br>\nmodel = create_RepVGG_B0(deploy=False)<br>\nmodel = RepVggModel(model, n_classes)</p>\n<h1>for inference</h1>\n<p>checkpoint = torch.load(model_name)<br>\nmodel_deploy = create_RepVGG_B0(deploy = False)<br>\nmodel_deploy = RepVggModel(model_deploy, n_classes)</p>\n<p>model_deploy.load_state_dict(checkpoint)<br>\nfor parameter in model_deploy.parameters():<br>\nparameter.requires_grad = False<br>\nmodel_deploy.eval()</p>\n<p>I also did some fine-tuning:</p>\n<p>def freeze_model_parametrs(model, n_layers):<br>\n    print('=====\\nFrozen\\n=====')</p>\n<pre><code>for name, param in list(model.named_parameters())[:n_layers]:\n    param.requires_grad = False\n\nprint('=====\\nTrainable\\n=====')\nfor name, param in list(model.named_parameters())[n_layers:]:\n    param.requires_grad = True\n\npytorch_total_params = sum(p.numel() for p in model.parameters())\npytorch_total_params_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\npytorch_total_params_frozen = sum(p.numel() for p in model.parameters() if p.requires_grad == False)\n\nprint('\\nTotal parameters: ', pytorch_total_params)\nprint('\\nTotal trainable parameters: ', pytorch_total_params_trainable)\nprint('\\nTotal frozen parameters: ', pytorch_total_params_frozen)\n\nparams = filter(lambda p: p.requires_grad, model.parameters())\n\nreturn model, params\n</code></pre>\n<p>For example, if I set deploy = True for inference, the model doesn't load properly, since I added some intermediate layers and the classification output is different.</p>\n<p>Am I doing something wrong in here during inference?</p>\n<p>Cheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1163616": "> We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3x3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet.\n\n[paper](https://arxiv.org/pdf/2101.03697v1.pdf)\n[pytorch code](https://github.com/DingXiaoH/RepVGG)",
    "1164051": "noticed that too.\nbut have not tried it yet.",
    "1164210": "i have uploaded all the weight files here : https://www.kaggle.com/mobassir/repvgg\nsoon planning to give it a try, thanks for sharing",
    "1166092": "It would be great, if someone could provide a simple example with some explanation in notebook.",
    "1169323": "I tried it, seems not really good in this exact case.",
    "1170079": "Same, i tested the model on the validation data and it gives me ~90% accuracy. But when i deploy the model for inference it gives me 51% accuracy, when the B0 version is used. Maybe i am doing something wrong when i load the pretrained model. Btw, i added some extra layers to achieve that performance.\n\nCan u please tell me how did you train and load the model? I did it this way:\n\nclass RepVggModel(nn.Module):\n    \n    def __init__(self, model, n_classes):\n        super().__init__()\n        self.n_classes = n_classes\n        self.model = model\n        self.model.linear = nn.Identity()\n        self.fc1 = nn.Linear(1280, 512) # 1280\n        self.bn = nn.BatchNorm1d(num_features=512) # 4096\n        self.relu = nn.ReLU(True)\n        self.dropout = nn.Dropout(0.5) # 0.3\n        self.linear = nn.Linear(512, self.n_classes) # 2*512\n\n    \n    def forward(self, x):\n        features = self.model(x)\n        features = self.fc1(features)\n        features = self.bn(features)\n        x = self.relu(features)\n        x = self.dropout(x)\n        x = self.linear(x)\n        #x = torch.softmax(x, dim = 1)\n        \n        return x\n\n#for training\nn_classes = 5\nmodel = create_RepVGG_B0(deploy=False)\nmodel = RepVggModel(model, n_classes)\n\n#for inference\ncheckpoint = torch.load(model_name)\nmodel_deploy = create_RepVGG_B0(deploy = False)\nmodel_deploy = RepVggModel(model_deploy, n_classes)\n\nmodel_deploy.load_state_dict(checkpoint)\nfor parameter in model_deploy.parameters():\nparameter.requires_grad = False\nmodel_deploy.eval()\n\nI also did some fine-tuning:\n\ndef freeze_model_parametrs(model, n_layers):\n    print('=====\\nFrozen\\n=====')\n\n    for name, param in list(model.named_parameters())[:n_layers]:\n        param.requires_grad = False\n\n    print('=====\\nTrainable\\n=====')\n    for name, param in list(model.named_parameters())[n_layers:]:\n        param.requires_grad = True\n\n    pytorch_total_params = sum(p.numel() for p in model.parameters())\n    pytorch_total_params_trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\n    pytorch_total_params_frozen = sum(p.numel() for p in model.parameters() if p.requires_grad == False)\n                                          \n    print('\\nTotal parameters: ', pytorch_total_params)\n    print('\\nTotal trainable parameters: ', pytorch_total_params_trainable)\n    print('\\nTotal frozen parameters: ', pytorch_total_params_frozen)\n    \n    params = filter(lambda p: p.requires_grad, model.parameters())\n    \n    return model, params\n\nFor example, if I set deploy = True for inference, the model doesn't load properly, since I added some intermediate layers and the classification output is different.\n\nAm I doing something wrong in here during inference?\n\nCheers!"
  },
  "source": "meta"
}