{
  "id": 585194,
  "title": "Anyteam can make Swin Transformer work?",
  "url": "/competitions/waveform-inversion/discussion/585194",
  "author_name": "",
  "post_date": "2025-06-18T15:56:40.769070500Z",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I tried Swin Transformer today and found a significant performance gap compared to ConvNeXt. In other classification and regression tasks, Swin Transformer usually performs quite well. </p>\n<pre><code> torchvision.models.swin_transformer  swin_b, Swin_B_Weights, swin_v2_b, Swin_V2_B_Weights\n torch\n torch.nn  nn\n\n (nn.Module):\n     ():\n        ().__init__()\n        .model = swin_b(weights=weights)\n        \n         .model.head \n        .out_indices = out_indices\n        ._update_stem()\n\n     ():\n        \n        module = .model.features[][]\n        new_conv = nn.Conv2d(\n            in_channels=in_channels,\n            out_channels=module.out_channels,\n            kernel_size=(, ), \n            stride=(,),\n            padding=(, ), \n            dilation=module.dilation,\n            groups=module.groups,\n            bias=module.bias   ,\n            padding_mode=module.padding_mode\n        )\n        \n         torch.no_grad():\n            old_weight = module.weight.data \n            \n            new_weight = old_weight[:,:,:,:].repeat(, , , )[:,:in_channels] \n            new_conv.weight.copy_(new_weight)\n             module.bias   :\n                new_conv.bias.copy_(module.bias)\n        .model.features[][] = new_conv\n\n\n     ():\n        features = []\n        \n         i, module  (.model.features):\n            x = module(x)\n            \n            \n             i  [, , , ]:  \n                \n                \n                \n                features.append(x.permute(, , , ))\n         features\n</code></pre>",
  "messages": [
    {
      "id": "3227167",
      "postDate": "06/18/2025 15:56:40",
      "content": "<p>I tried Swin Transformer today and found a significant performance gap compared to ConvNeXt. In other classification and regression tasks, Swin Transformer usually performs quite well. </p>\n<pre><code> torchvision.models.swin_transformer  swin_b, Swin_B_Weights, swin_v2_b, Swin_V2_B_Weights\n torch\n torch.nn  nn\n\n (nn.Module):\n     ():\n        ().__init__()\n        .model = swin_b(weights=weights)\n        \n         .model.head \n        .out_indices = out_indices\n        ._update_stem()\n\n     ():\n        \n        module = .model.features[][]\n        new_conv = nn.Conv2d(\n            in_channels=in_channels,\n            out_channels=module.out_channels,\n            kernel_size=(, ), \n            stride=(,),\n            padding=(, ), \n            dilation=module.dilation,\n            groups=module.groups,\n            bias=module.bias   ,\n            padding_mode=module.padding_mode\n        )\n        \n         torch.no_grad():\n            old_weight = module.weight.data \n            \n            new_weight = old_weight[:,:,:,:].repeat(, , , )[:,:in_channels] \n            new_conv.weight.copy_(new_weight)\n             module.bias   :\n                new_conv.bias.copy_(module.bias)\n        .model.features[][] = new_conv\n\n\n     ():\n        features = []\n        \n         i, module  (.model.features):\n            x = module(x)\n            \n            \n             i  [, , , ]:  \n                \n                \n                \n                features.append(x.permute(, , , ))\n         features\n</code></pre>",
      "rawMarkdown": "I tried Swin Transformer today and found a significant performance gap compared to ConvNeXt. In other classification and regression tasks, Swin Transformer usually performs quite well. \n\n```\nfrom torchvision.models.swin_transformer import swin_b, Swin_B_Weights, swin_v2_b, Swin_V2_B_Weights\nimport torch\nimport torch.nn as nn\n\nclass SwinBFeatures(nn.Module):\n    def __init__(self, weights=Swin_B_Weights.DEFAULT, out_indices=(0, 1, 2, 3)):\n        super().__init__()\n        self.model = swin_b(weights=weights)\n        # print(self.model)\n        del self.model.head \n        self.out_indices = out_indices\n        self._update_stem()\n        \n    def _update_stem(self, in_channels=5):\n        # create Conv2d layer\n        module = self.model.features[0][0]\n        new_conv = nn.Conv2d(\n            in_channels=in_channels,\n            out_channels=module.out_channels,\n            kernel_size=(4, 4), \n            stride=(4,1),\n            padding=(0, 4), \n            dilation=module.dilation,\n            groups=module.groups,\n            bias=module.bias is not None,\n            padding_mode=module.padding_mode\n        )\n        # adjust weight\n        with torch.no_grad():\n            old_weight = module.weight.data \n            # print(old_weight.shape)\n            new_weight = old_weight[:,:,:,:].repeat(1, 2, 1, 1)[:,:in_channels] \n            new_conv.weight.copy_(new_weight)\n            if module.bias is not None:\n                new_conv.bias.copy_(module.bias)\n        self.model.features[0][0] = new_conv\n\n        \n    def forward(self, x):\n        features = []\n        # print(f\"Input shape: {x.shape}\")\n        for i, module in enumerate(self.model.features):\n            x = module(x)\n            # print(f\"Layer {i} output: {x.shape}\")\n            # if i in [0, 2, 4, 6]:\n            if i in [1, 3, 5, 7]:  \n                # stage_idx = (i - 1) // 2\n                # if stage_idx in self.out_indices:\n                # [B, H, W, C] -> [B, C, H, W]\n                features.append(x.permute(0, 3, 1, 2))\n        return features\n```",
      "votes": null
    },
    {
      "id": "3227177",
      "postDate": "06/18/2025 16:16:10",
      "content": "<p>Don't know about the public baselines, but when I had a score of about 23-25, I did test that on my baseline, swin was converging similarly to my primary model at the time, but it was a lot slower to train so that was the last of what I tested.</p>",
      "rawMarkdown": "Don't know about the public baselines, but when I had a score of about 23-25, I did test that on my baseline, swin was converging similarly to my primary model at the time, but it was a lot slower to train so that was the last of what I tested.",
      "votes": null
    },
    {
      "id": "3227193",
      "postDate": "06/18/2025 16:40:43",
      "content": "<p>I only checked the loss and validation score for the first 5 epochs. Since it converged significantly slower than ConvNeXt and CAFormer, I killed the training.</p>",
      "rawMarkdown": "I only checked the loss and validation score for the first 5 epochs. Since it converged significantly slower than ConvNeXt and CAFormer, I killed the training.",
      "votes": null
    },
    {
      "id": "3227208",
      "postDate": "06/18/2025 16:55:29",
      "content": "<p>Same, I only checked for first like 10 epochs (took like 3 hours), the convergence in my case was actually better than my primary model, then killed because the improvement was not worth almost double the training time</p>",
      "rawMarkdown": "Same, I only checked for first like 10 epochs (took like 3 hours), the convergence in my case was actually better than my primary model, then killed because the improvement was not worth almost double the training time",
      "votes": null
    },
    {
      "id": "3227210",
      "postDate": "06/18/2025 16:56:20",
      "content": "<p>Transformer needs smaller learning rate and the input value range mast be similar to the pretrain samples. Now we modify the stem, care has to be taken.</p>\n<p>A way to check is convert into gray scale image and use 3 channel first annd just resize to 72*72, and compare against different backbone like vit and cnn. They should all give same results</p>",
      "rawMarkdown": "Transformer needs smaller learning rate and the input value range mast be similar to the pretrain samples. Now we modify the stem, care has to be taken.\n\nA way to check is convert into gray scale image and use 3 channel first annd just resize to 72*72, and compare against different backbone like vit and cnn. They should all give same results",
      "votes": null
    },
    {
      "id": "3227236",
      "postDate": "06/18/2025 17:14:52",
      "content": "<p>I tried learning rates of 2e-4 and 2e-5, but the results weren't great. I'm not very experienced with CNN models, so you might want to give it a try as well.</p>",
      "rawMarkdown": "I tried learning rates of 2e-4 and 2e-5, but the results weren't great. I'm not very experienced with CNN models, so you might want to give it a try as well.",
      "votes": null
    },
    {
      "id": "3227259",
      "postDate": "06/18/2025 18:11:50",
      "content": "<p>one way to check proper working of neural nets (either CNN,VIT, or other type) is to check the distribution of their activation values.</p>\n<p>so I can imagine first I thrown in a few imagenet images, and get plot of the histogram of activations. then I thrown in the semisc data (and new stem) and I should same activation range.</p>\n<p>quite a lot of effort to do this.</p>\n<hr>\n<p>but since we already have one working transformer (caformer) and CNN(covnext) I will skip this and focus on other methods (like using physics, generate ore data etc)…</p>\n<p>but in previous competitions, my experience is that VIT  often fails to work if the values after stem is not correct.</p>\n<p>finally, I did briefly try maxvit and nextvit. they still can converge but at slower rate. note that i experiment for few epoches only.</p>",
      "rawMarkdown": "one way to check proper working of neural nets (either CNN,VIT, or other type) is to check the distribution of their activation values.\n\nso I can imagine first I thrown in a few imagenet images, and get plot of the histogram of activations. then I thrown in the semisc data (and new stem) and I should same activation range.\n\nquite a lot of effort to do this.\n\n---\n\nbut since we already have one working transformer (caformer) and CNN(covnext) I will skip this and focus on other methods (like using physics, generate ore data etc)...\n\nbut in previous competitions, my experience is that VIT  often fails to work if the values after stem is not correct.\n\nfinally, I did briefly try maxvit and nextvit. they still can converge but at slower rate. note that i experiment for few epoches only.",
      "votes": null
    },
    {
      "id": "3227475",
      "postDate": "06/19/2025 03:11:09",
      "content": "<p>Thanks, I will also give up tuning Swin Transformer and focus on finding other useful methods since we are very closing to the end.</p>",
      "rawMarkdown": "Thanks, I will also give up tuning Swin Transformer and focus on finding other useful methods since we are very closing to the end.",
      "votes": null
    },
    {
      "id": "3229683",
      "postDate": "06/22/2025 00:50:00",
      "content": "<p>I tried a hybrid approach that combined convolutional and transformer-based ideas, but it didn’t perform well</p>",
      "rawMarkdown": "I tried a hybrid approach that combined convolutional and transformer-based ideas, but it didn’t perform well",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3227177,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "06/18/2025 16:16:10",
      "content": "<p>Don't know about the public baselines, but when I had a score of about 23-25, I did test that on my baseline, swin was converging similarly to my primary model at the time, but it was a lot slower to train so that was the last of what I tested.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3227193,
          "author_name": "hydantess",
          "author_url": "",
          "post_date": "06/18/2025 16:40:43",
          "content": "<p>I only checked the loss and validation score for the first 5 epochs. Since it converged significantly slower than ConvNeXt and CAFormer, I killed the training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3227208,
              "author_name": "harshitsheoran",
              "author_url": "",
              "post_date": "06/18/2025 16:55:29",
              "content": "<p>Same, I only checked for first like 10 epochs (took like 3 hours), the convergence in my case was actually better than my primary model, then killed because the improvement was not worth almost double the training time</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3227210,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/18/2025 16:56:20",
      "content": "<p>Transformer needs smaller learning rate and the input value range mast be similar to the pretrain samples. Now we modify the stem, care has to be taken.</p>\n<p>A way to check is convert into gray scale image and use 3 channel first annd just resize to 72*72, and compare against different backbone like vit and cnn. They should all give same results</p>",
      "votes": null,
      "replies": [
        {
          "id": 3227236,
          "author_name": "hydantess",
          "author_url": "",
          "post_date": "06/18/2025 17:14:52",
          "content": "<p>I tried learning rates of 2e-4 and 2e-5, but the results weren't great. I'm not very experienced with CNN models, so you might want to give it a try as well.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3227259,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/18/2025 18:11:50",
              "content": "<p>one way to check proper working of neural nets (either CNN,VIT, or other type) is to check the distribution of their activation values.</p>\n<p>so I can imagine first I thrown in a few imagenet images, and get plot of the histogram of activations. then I thrown in the semisc data (and new stem) and I should same activation range.</p>\n<p>quite a lot of effort to do this.</p>\n<hr>\n<p>but since we already have one working transformer (caformer) and CNN(covnext) I will skip this and focus on other methods (like using physics, generate ore data etc)…</p>\n<p>but in previous competitions, my experience is that VIT  often fails to work if the values after stem is not correct.</p>\n<p>finally, I did briefly try maxvit and nextvit. they still can converge but at slower rate. note that i experiment for few epoches only.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3227475,
                  "author_name": "hydantess",
                  "author_url": "",
                  "post_date": "06/19/2025 03:11:09",
                  "content": "<p>Thanks, I will also give up tuning Swin Transformer and focus on finding other useful methods since we are very closing to the end.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3229683,
      "author_name": "luyeloveyalin",
      "author_url": "",
      "post_date": "06/22/2025 00:50:00",
      "content": "<p>I tried a hybrid approach that combined convolutional and transformer-based ideas, but it didn’t perform well</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3227167": "I tried Swin Transformer today and found a significant performance gap compared to ConvNeXt. In other classification and regression tasks, Swin Transformer usually performs quite well. \n\n```\nfrom torchvision.models.swin_transformer import swin_b, Swin_B_Weights, swin_v2_b, Swin_V2_B_Weights\nimport torch\nimport torch.nn as nn\n\nclass SwinBFeatures(nn.Module):\n    def __init__(self, weights=Swin_B_Weights.DEFAULT, out_indices=(0, 1, 2, 3)):\n        super().__init__()\n        self.model = swin_b(weights=weights)\n        # print(self.model)\n        del self.model.head \n        self.out_indices = out_indices\n        self._update_stem()\n        \n    def _update_stem(self, in_channels=5):\n        # create Conv2d layer\n        module = self.model.features[0][0]\n        new_conv = nn.Conv2d(\n            in_channels=in_channels,\n            out_channels=module.out_channels,\n            kernel_size=(4, 4), \n            stride=(4,1),\n            padding=(0, 4), \n            dilation=module.dilation,\n            groups=module.groups,\n            bias=module.bias is not None,\n            padding_mode=module.padding_mode\n        )\n        # adjust weight\n        with torch.no_grad():\n            old_weight = module.weight.data \n            # print(old_weight.shape)\n            new_weight = old_weight[:,:,:,:].repeat(1, 2, 1, 1)[:,:in_channels] \n            new_conv.weight.copy_(new_weight)\n            if module.bias is not None:\n                new_conv.bias.copy_(module.bias)\n        self.model.features[0][0] = new_conv\n\n        \n    def forward(self, x):\n        features = []\n        # print(f\"Input shape: {x.shape}\")\n        for i, module in enumerate(self.model.features):\n            x = module(x)\n            # print(f\"Layer {i} output: {x.shape}\")\n            # if i in [0, 2, 4, 6]:\n            if i in [1, 3, 5, 7]:  \n                # stage_idx = (i - 1) // 2\n                # if stage_idx in self.out_indices:\n                # [B, H, W, C] -> [B, C, H, W]\n                features.append(x.permute(0, 3, 1, 2))\n        return features\n```",
    "3227177": "Don't know about the public baselines, but when I had a score of about 23-25, I did test that on my baseline, swin was converging similarly to my primary model at the time, but it was a lot slower to train so that was the last of what I tested.",
    "3227193": "I only checked the loss and validation score for the first 5 epochs. Since it converged significantly slower than ConvNeXt and CAFormer, I killed the training.",
    "3227208": "Same, I only checked for first like 10 epochs (took like 3 hours), the convergence in my case was actually better than my primary model, then killed because the improvement was not worth almost double the training time",
    "3227210": "Transformer needs smaller learning rate and the input value range mast be similar to the pretrain samples. Now we modify the stem, care has to be taken.\n\nA way to check is convert into gray scale image and use 3 channel first annd just resize to 72*72, and compare against different backbone like vit and cnn. They should all give same results",
    "3227236": "I tried learning rates of 2e-4 and 2e-5, but the results weren't great. I'm not very experienced with CNN models, so you might want to give it a try as well.",
    "3227259": "one way to check proper working of neural nets (either CNN,VIT, or other type) is to check the distribution of their activation values.\n\nso I can imagine first I thrown in a few imagenet images, and get plot of the histogram of activations. then I thrown in the semisc data (and new stem) and I should same activation range.\n\nquite a lot of effort to do this.\n\n---\n\nbut since we already have one working transformer (caformer) and CNN(covnext) I will skip this and focus on other methods (like using physics, generate ore data etc)...\n\nbut in previous competitions, my experience is that VIT  often fails to work if the values after stem is not correct.\n\nfinally, I did briefly try maxvit and nextvit. they still can converge but at slower rate. note that i experiment for few epoches only.",
    "3227475": "Thanks, I will also give up tuning Swin Transformer and focus on finding other useful methods since we are very closing to the end.",
    "3229683": "I tried a hybrid approach that combined convolutional and transformer-based ideas, but it didn’t perform well"
  },
  "source": "meta"
}