{
  "id": 168542,
  "title": "Single Effnet-B0 Private LB 0.921: How to Modify Effnet Architecture",
  "url": "/competitions/alaska2-image-steganalysis/discussion/168542",
  "author_name": "Qishen Ha",
  "post_date": "2020-07-21T03:40:41.597000",
  "votes": 74,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi everyone. Congrats to all winners and thanks to the organizers for this interesting (and harsh) competition!</p>\n<p>Here I want to show you how we modified the EfficientNet-B0 architecture to get Private LB 0.921 (Public LB 0.930~)</p>\n<p>Firstly, add 3 Conv2d Layers before EfficientNet like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd81542263f329ce3547f7206279f088f%2Fp1.png?generation=1595302919316310&amp;alt=media\" alt=\"\"></p>\n<p>Secondly, remove block #5 and #6 in the original  EfficientNet architecture like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd75494dcecbe1e77b30bf41e93916f66%2F2.png?generation=1595302943135345&amp;alt=media\" alt=\"\"></p>\n<p>Simply doing this 2 tricks we were able to boost our local AUC of 30epo experiment from 0.916 -&gt; 0.921<br>\nThen further train for another 30epo (60epo in total), we reach local AUC 0.926. With TTA around 0.930</p>\n<h2>Code</h2>\n<pre><code>import timm\n\nclass enetv2(nn.Module):\n    def __init__(self, backbone, out_dim):\n        super(enetv2, self).__init__()\n        self.conv1 = nn.Conv2d(3, 6, 3, stride=1, padding=1, bias=False)\n        self.conv2 = nn.Conv2d(6, 12, 3, stride=1, padding=1, bias=False)\n        self.conv3 = nn.Conv2d(12, 36, 3, stride=1, padding=1, bias=False)\n        self.mybn1 = nn.BatchNorm2d(6)\n        self.mybn2 = nn.BatchNorm2d(12)\n        self.mybn3 = nn.BatchNorm2d(36)\n\n        self.enet = timm.create_model('efficientnet_b0', pretrained=True)\n        self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1, 12, 1, 1))\n\n        self.dropout = nn.Dropout(0.5)\n        self.enet.blocks[5] = nn.Identity()\n        self.enet.blocks[6] = nn.Sequential(\n            nn.Conv2d(self.enet.blocks[4][2].conv_pwl.out_channels, self.enet.conv_head.in_channels, 1),\n            nn.BatchNorm2d(self.enet.conv_head.in_channels),\n            nn.ReLU6(),\n        )\n        self.myfc = nn.Linear(self.enet.classifier.in_features, out_dim)\n        self.enet.classifier = nn.Identity()\n\n    def extract(self, x):\n        x = F.relu6(self.mybn1(self.conv1(x)))\n        x = F.relu6(self.mybn2(self.conv2(x)))\n        x = F.relu6(self.mybn3(self.conv3(x)))\n        x = self.enet(x)\n        return x\n\n    def forward(self, x):\n        x = self.extract(x)\n        x = self.myfc(self.dropout(x))\n        return x\n</code></pre>",
  "messages": [
    {
      "id": 937474,
      "postDate": "2020-07-21T03:40:41.597Z",
      "content": "<p>Hi everyone. Congrats to all winners and thanks to the organizers for this interesting (and harsh) competition!</p>\n<p>Here I want to show you how we modified the EfficientNet-B0 architecture to get Private LB 0.921 (Public LB 0.930~)</p>\n<p>Firstly, add 3 Conv2d Layers before EfficientNet like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd81542263f329ce3547f7206279f088f%2Fp1.png?generation=1595302919316310&amp;alt=media\" alt=\"\"></p>\n<p>Secondly, remove block #5 and #6 in the original  EfficientNet architecture like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd75494dcecbe1e77b30bf41e93916f66%2F2.png?generation=1595302943135345&amp;alt=media\" alt=\"\"></p>\n<p>Simply doing this 2 tricks we were able to boost our local AUC of 30epo experiment from 0.916 -&gt; 0.921<br>\nThen further train for another 30epo (60epo in total), we reach local AUC 0.926. With TTA around 0.930</p>\n<h2>Code</h2>\n<pre><code>import timm\n\nclass enetv2(nn.Module):\n    def __init__(self, backbone, out_dim):\n        super(enetv2, self).__init__()\n        self.conv1 = nn.Conv2d(3, 6, 3, stride=1, padding=1, bias=False)\n        self.conv2 = nn.Conv2d(6, 12, 3, stride=1, padding=1, bias=False)\n        self.conv3 = nn.Conv2d(12, 36, 3, stride=1, padding=1, bias=False)\n        self.mybn1 = nn.BatchNorm2d(6)\n        self.mybn2 = nn.BatchNorm2d(12)\n        self.mybn3 = nn.BatchNorm2d(36)\n\n        self.enet = timm.create_model('efficientnet_b0', pretrained=True)\n        self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1, 12, 1, 1))\n\n        self.dropout = nn.Dropout(0.5)\n        self.enet.blocks[5] = nn.Identity()\n        self.enet.blocks[6] = nn.Sequential(\n            nn.Conv2d(self.enet.blocks[4][2].conv_pwl.out_channels, self.enet.conv_head.in_channels, 1),\n            nn.BatchNorm2d(self.enet.conv_head.in_channels),\n            nn.ReLU6(),\n        )\n        self.myfc = nn.Linear(self.enet.classifier.in_features, out_dim)\n        self.enet.classifier = nn.Identity()\n\n    def extract(self, x):\n        x = F.relu6(self.mybn1(self.conv1(x)))\n        x = F.relu6(self.mybn2(self.conv2(x)))\n        x = F.relu6(self.mybn3(self.conv3(x)))\n        x = self.enet(x)\n        return x\n\n    def forward(self, x):\n        x = self.extract(x)\n        x = self.myfc(self.dropout(x))\n        return x\n</code></pre>",
      "rawMarkdown": "Hi everyone. Congrats to all winners and thanks to the organizers for this interesting (and harsh) competition!\n\nHere I want to show you how we modified the EfficientNet-B0 architecture to get Private LB 0.921 (Public LB 0.930~)\n\nFirstly, add 3 Conv2d Layers before EfficientNet like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd81542263f329ce3547f7206279f088f%2Fp1.png?generation=1595302919316310&amp;alt=media)\n\n\nSecondly, remove block #5 and #6 in the original  EfficientNet architecture like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd75494dcecbe1e77b30bf41e93916f66%2F2.png?generation=1595302943135345&amp;alt=media)\n\n\nSimply doing this 2 tricks we were able to boost our local AUC of 30epo experiment from 0.916 -&gt; 0.921\nThen further train for another 30epo (60epo in total), we reach local AUC 0.926. With TTA around 0.930\n\n## Code\n\n```\nimport timm\n\nclass enetv2(nn.Module):\n    def __init__(self, backbone, out_dim):\n        super(enetv2, self).__init__()\n        self.conv1 = nn.Conv2d(3, 6, 3, stride=1, padding=1, bias=False)\n        self.conv2 = nn.Conv2d(6, 12, 3, stride=1, padding=1, bias=False)\n        self.conv3 = nn.Conv2d(12, 36, 3, stride=1, padding=1, bias=False)\n        self.mybn1 = nn.BatchNorm2d(6)\n        self.mybn2 = nn.BatchNorm2d(12)\n        self.mybn3 = nn.BatchNorm2d(36)\n\n        self.enet = timm.create_model('efficientnet_b0', pretrained=True)\n        self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1, 12, 1, 1))\n\n        self.dropout = nn.Dropout(0.5)\n        self.enet.blocks[5] = nn.Identity()\n        self.enet.blocks[6] = nn.Sequential(\n            nn.Conv2d(self.enet.blocks[4][2].conv_pwl.out_channels, self.enet.conv_head.in_channels, 1),\n            nn.BatchNorm2d(self.enet.conv_head.in_channels),\n            nn.ReLU6(),\n        )\n        self.myfc = nn.Linear(self.enet.classifier.in_features, out_dim)\n        self.enet.classifier = nn.Identity()\n\n    def extract(self, x):\n        x = F.relu6(self.mybn1(self.conv1(x)))\n        x = F.relu6(self.mybn2(self.conv2(x)))\n        x = F.relu6(self.mybn3(self.conv3(x)))\n        x = self.enet(x)\n        return x\n\n    def forward(self, x):\n        x = self.extract(x)\n        x = self.myfc(self.dropout(x))\n        return x\n```",
      "votes": 73
    },
    {
      "id": 944765,
      "postDate": "2020-07-25T10:29:58.933Z",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> </p>\n\n<p>thanks for the writeup. I implement your modified modified B0 and can get private 0.925 and public 0.936.\ni train 2 cyclic round each for about 40 epoch. (The first cycle gives  0.918/0.927). I was swish() activation for all modifications.</p>\n\n<p>It seems to me that if trained  properly, the even simple model 's score can be quite high. But it takes long days to train and hence it is difficult to do experiments. I suspect a 3rd cycle would gave better results, maybe around +2 improvement.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8f4e608918cf1f7bee7115e68657458%2FSelection_038.png?generation=1595672958616927&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@haqishen \n\nthanks for the writeup. I implement your modified modified B0 and can get private 0.925 and public 0.936.\ni train 2 cyclic round each for about 40 epoch. (The first cycle gives  0.918/0.927). I was swish() activation for all modifications.\n\nIt seems to me that if trained  properly, the even simple model 's score can be quite high. But it takes long days to train and hence it is difficult to do experiments. I suspect a 3rd cycle would gave better results, maybe around +2 improvement.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8f4e608918cf1f7bee7115e68657458%2FSelection_038.png?generation=1595672958616927&amp;alt=media)\n",
      "votes": 7,
      "replies": [
        {
          "id": 945171,
          "postDate": "2020-07-25T16:04:00.040Z",
          "content": "<p>Yep. \nWe don't have enough machine and time to do proper experiments so just train 60epo for 1 cyclic for final models.</p>",
          "rawMarkdown": "Yep. \nWe don't have enough machine and time to do proper experiments so just train 60epo for 1 cyclic for final models.",
          "votes": 1
        }
      ]
    },
    {
      "id": 953729,
      "postDate": "2020-08-01T03:52:04.363Z",
      "content": "<p>Appreciate the sharing QiShen. It is interesting in current AI world that experimentation based on prior model performance and experiences hunch make the improvement nowadays ahead of theory and math.  Such analysis of improvements often come after experimentations today.    Good work. Dr.</p>",
      "rawMarkdown": "Appreciate the sharing QiShen. It is interesting in current AI world that experimentation based on prior model performance and experiences hunch make the improvement nowadays ahead of theory and math.  Such analysis of improvements often come after experimentations today.    Good work. Dr.",
      "votes": 1
    },
    {
      "id": 937724,
      "postDate": "2020-07-21T06:20:26.300Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<p>Really liked the way you experimented. Thank you for sharing. </p>",
      "rawMarkdown": "Congrats @haqishen \n\nReally liked the way you experimented. Thank you for sharing. ",
      "votes": 1
    },
    {
      "id": 937496,
      "postDate": "2020-07-21T03:55:04.993Z",
      "content": "<p>That's awesome to get such a score with a single model, I wonder what was the motivation behind this? </p>",
      "rawMarkdown": "That's awesome to get such a score with a single model, I wonder what was the motivation behind this? ",
      "votes": 1,
      "replies": [
        {
          "id": 937535,
          "postDate": "2020-07-21T04:28:05.047Z",
          "content": "<p>Low level feature in this dataset is quite important so we added some conv layer before backbone.<br>\nThen tried to feed 64x64 feature map directly into global pooling and got better score.<br>\nAt last we found that remove those blocks (#5#6) which run on feature map 32x32 can got even better score…</p>",
          "rawMarkdown": "Low level feature in this dataset is quite important so we added some conv layer before backbone.\nThen tried to feed 64x64 feature map directly into global pooling and got better score.\nAt last we found that remove those blocks (#5#6) which run on feature map 32x32 can got even better score…",
          "votes": 5
        },
        {
          "id": 937963,
          "postDate": "2020-07-21T09:14:24.513Z",
          "content": "<p>Awesome and congratulations</p>",
          "rawMarkdown": "Awesome and congratulations",
          "votes": 1
        },
        {
          "id": 945046,
          "postDate": "2020-07-25T14:32:53.267Z",
          "content": "<p>Congratulations mate, I had the same question as <a href=\"/tanulsingh077\">@tanulsingh077</a>.\nIt's all about trial and error I guess? \nAlso, I think that if you develop enough, you get some good instincts, I strongly believe.</p>",
          "rawMarkdown": "Congratulations mate, I had the same question as @tanulsingh077.\nIt's all about trial and error I guess? \nAlso, I think that if you develop enough, you get some good instincts, I strongly believe.",
          "votes": 1
        },
        {
          "id": 945249,
          "postDate": "2020-07-25T17:18:30.383Z",
          "content": "<p>To be honest yes, we don't have 100% confidence on our tricks, before the experiment result came up.</p>",
          "rawMarkdown": "To be honest yes, we don't have 100% confidence on our tricks, before the experiment result came up."
        }
      ]
    },
    {
      "id": 937484,
      "postDate": "2020-07-21T03:45:16.923Z",
      "content": "<p>Great work. I wonder why did you remove block #5 and #6?</p>",
      "rawMarkdown": "Great work. I wonder why did you remove block #5 and #6?",
      "votes": 1,
      "replies": [
        {
          "id": 937532,
          "postDate": "2020-07-21T04:25:22.043Z",
          "content": "<p>We wanted to try 64x64 feature map feed into global pooling at first so we modified the <code>stride=2</code> to 1 in block #5.<br>\nThen we tried to remove those blocks (#5#6) which run on feature map 32x32 originally and got better score…</p>",
          "rawMarkdown": "We wanted to try 64x64 feature map feed into global pooling at first so we modified the `stride=2` to 1 in block #5.\nThen we tried to remove those blocks (#5#6) which run on feature map 32x32 originally and got better score...",
          "votes": 1
        },
        {
          "id": 937718,
          "postDate": "2020-07-21T06:17:34.870Z",
          "content": "<p>oh, interesting results! i will try your method. thanks.</p>",
          "rawMarkdown": "oh, interesting results! i will try your method. thanks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 937699,
      "postDate": "2020-07-21T06:08:08.117Z",
      "content": "<p>Any ways on how to remove it from pretrained effnet model in \"pip install effiecient-net\"? <a href=\"/haqishen\">@haqishen</a> </p>",
      "rawMarkdown": "Any ways on how to remove it from pretrained effnet model in \"pip install effiecient-net\"? @haqishen ",
      "replies": [
        {
          "id": 937813,
          "postDate": "2020-07-21T07:11:32.450Z",
          "content": "<p>Of course, it's basically the same as this one I shown.</p>",
          "rawMarkdown": "Of course, it's basically the same as this one I shown."
        }
      ]
    },
    {
      "id": 940385,
      "postDate": "2020-07-22T22:40:58.283Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 939240,
      "postDate": "2020-07-22T05:45:50.687Z",
      "content": "<p>Thanks, great work!</p>",
      "rawMarkdown": "Thanks, great work!",
      "votes": 1
    },
    {
      "id": 937613,
      "postDate": "2020-07-21T05:25:50.193Z",
      "content": "<p>Awesome!\nThanks for sharing 😃 </p>",
      "rawMarkdown": "Awesome!\nThanks for sharing 😃 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 944765,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-25T10:29:58.933000",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> </p>\n\n<p>thanks for the writeup. I implement your modified modified B0 and can get private 0.925 and public 0.936.\ni train 2 cyclic round each for about 40 epoch. (The first cycle gives  0.918/0.927). I was swish() activation for all modifications.</p>\n\n<p>It seems to me that if trained  properly, the even simple model 's score can be quite high. But it takes long days to train and hence it is difficult to do experiments. I suspect a 3rd cycle would gave better results, maybe around +2 improvement.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8f4e608918cf1f7bee7115e68657458%2FSelection_038.png?generation=1595672958616927&amp;alt=media\" alt=\"\"></p>",
      "votes": 7,
      "replies": [
        {
          "id": 945171,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-07-25T16:04:00.040000",
          "content": "<p>Yep. \nWe don't have enough machine and time to do proper experiments so just train 60epo for 1 cyclic for final models.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 953729,
      "author_name": "Patrick Chan",
      "author_url": "",
      "post_date": "2020-08-01T03:52:04.363000",
      "content": "<p>Appreciate the sharing QiShen. It is interesting in current AI world that experimentation based on prior model performance and experiences hunch make the improvement nowadays ahead of theory and math.  Such analysis of improvements often come after experimentations today.    Good work. Dr.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937724,
      "author_name": "Vishnu R",
      "author_url": "",
      "post_date": "2020-07-21T06:20:26.300000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<p>Really liked the way you experimented. Thank you for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937496,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2020-07-21T03:55:04.993000",
      "content": "<p>That's awesome to get such a score with a single model, I wonder what was the motivation behind this? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 937535,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-07-21T04:28:05.047000",
          "content": "<p>Low level feature in this dataset is quite important so we added some conv layer before backbone.<br>\nThen tried to feed 64x64 feature map directly into global pooling and got better score.<br>\nAt last we found that remove those blocks (#5#6) which run on feature map 32x32 can got even better score…</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 937963,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-07-21T09:14:24.513000",
          "content": "<p>Awesome and congratulations</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945046,
          "author_name": "Alin Cijov",
          "author_url": "",
          "post_date": "2020-07-25T14:32:53.267000",
          "content": "<p>Congratulations mate, I had the same question as <a href=\"/tanulsingh077\">@tanulsingh077</a>.\nIt's all about trial and error I guess? \nAlso, I think that if you develop enough, you get some good instincts, I strongly believe.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945249,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-07-25T17:18:30.383000",
          "content": "<p>To be honest yes, we don't have 100% confidence on our tricks, before the experiment result came up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 937484,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2020-07-21T03:45:16.923000",
      "content": "<p>Great work. I wonder why did you remove block #5 and #6?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 937532,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-07-21T04:25:22.043000",
          "content": "<p>We wanted to try 64x64 feature map feed into global pooling at first so we modified the <code>stride=2</code> to 1 in block #5.<br>\nThen we tried to remove those blocks (#5#6) which run on feature map 32x32 originally and got better score…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 937718,
          "author_name": "Wonho Song",
          "author_url": "",
          "post_date": "2020-07-21T06:17:34.870000",
          "content": "<p>oh, interesting results! i will try your method. thanks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 937699,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-07-21T06:08:08.117000",
      "content": "<p>Any ways on how to remove it from pretrained effnet model in \"pip install effiecient-net\"? <a href=\"/haqishen\">@haqishen</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 937813,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-07-21T07:11:32.450000",
          "content": "<p>Of course, it's basically the same as this one I shown.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 940385,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-22T22:40:58.283000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 939240,
      "author_name": "zheng lv",
      "author_url": "",
      "post_date": "2020-07-22T05:45:50.687000",
      "content": "<p>Thanks, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 937613,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-07-21T05:25:50.193000",
      "content": "<p>Awesome!\nThanks for sharing 😃 </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937474": "Hi everyone. Congrats to all winners and thanks to the organizers for this interesting (and harsh) competition!\n\nHere I want to show you how we modified the EfficientNet-B0 architecture to get Private LB 0.921 (Public LB 0.930~)\n\nFirstly, add 3 Conv2d Layers before EfficientNet like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd81542263f329ce3547f7206279f088f%2Fp1.png?generation=1595302919316310&amp;alt=media)\n\n\nSecondly, remove block #5 and #6 in the original  EfficientNet architecture like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F448347%2Fd75494dcecbe1e77b30bf41e93916f66%2F2.png?generation=1595302943135345&amp;alt=media)\n\n\nSimply doing this 2 tricks we were able to boost our local AUC of 30epo experiment from 0.916 -&gt; 0.921\nThen further train for another 30epo (60epo in total), we reach local AUC 0.926. With TTA around 0.930\n\n## Code\n\n```\nimport timm\n\nclass enetv2(nn.Module):\n    def __init__(self, backbone, out_dim):\n        super(enetv2, self).__init__()\n        self.conv1 = nn.Conv2d(3, 6, 3, stride=1, padding=1, bias=False)\n        self.conv2 = nn.Conv2d(6, 12, 3, stride=1, padding=1, bias=False)\n        self.conv3 = nn.Conv2d(12, 36, 3, stride=1, padding=1, bias=False)\n        self.mybn1 = nn.BatchNorm2d(6)\n        self.mybn2 = nn.BatchNorm2d(12)\n        self.mybn3 = nn.BatchNorm2d(36)\n\n        self.enet = timm.create_model('efficientnet_b0', pretrained=True)\n        self.enet.conv_stem.weight = nn.Parameter(self.enet.conv_stem.weight.repeat(1, 12, 1, 1))\n\n        self.dropout = nn.Dropout(0.5)\n        self.enet.blocks[5] = nn.Identity()\n        self.enet.blocks[6] = nn.Sequential(\n            nn.Conv2d(self.enet.blocks[4][2].conv_pwl.out_channels, self.enet.conv_head.in_channels, 1),\n            nn.BatchNorm2d(self.enet.conv_head.in_channels),\n            nn.ReLU6(),\n        )\n        self.myfc = nn.Linear(self.enet.classifier.in_features, out_dim)\n        self.enet.classifier = nn.Identity()\n\n    def extract(self, x):\n        x = F.relu6(self.mybn1(self.conv1(x)))\n        x = F.relu6(self.mybn2(self.conv2(x)))\n        x = F.relu6(self.mybn3(self.conv3(x)))\n        x = self.enet(x)\n        return x\n\n    def forward(self, x):\n        x = self.extract(x)\n        x = self.myfc(self.dropout(x))\n        return x\n```",
    "944765": "@haqishen \n\nthanks for the writeup. I implement your modified modified B0 and can get private 0.925 and public 0.936.\ni train 2 cyclic round each for about 40 epoch. (The first cycle gives  0.918/0.927). I was swish() activation for all modifications.\n\nIt seems to me that if trained  properly, the even simple model 's score can be quite high. But it takes long days to train and hence it is difficult to do experiments. I suspect a 3rd cycle would gave better results, maybe around +2 improvement.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8f4e608918cf1f7bee7115e68657458%2FSelection_038.png?generation=1595672958616927&amp;alt=media)\n",
    "953729": "Appreciate the sharing QiShen. It is interesting in current AI world that experimentation based on prior model performance and experiences hunch make the improvement nowadays ahead of theory and math.  Such analysis of improvements often come after experimentations today.    Good work. Dr.",
    "937724": "Congrats @haqishen \n\nReally liked the way you experimented. Thank you for sharing. ",
    "937496": "That's awesome to get such a score with a single model, I wonder what was the motivation behind this? ",
    "937484": "Great work. I wonder why did you remove block #5 and #6?",
    "937699": "Any ways on how to remove it from pretrained effnet model in \"pip install effiecient-net\"? @haqishen ",
    "940385": "",
    "939240": "Thanks, great work!",
    "937613": "Awesome!\nThanks for sharing 😃 "
  }
}