{
  "id": 246558,
  "title": "Learned Image Resizing",
  "url": "/competitions/seti-breakthrough-listen/discussion/246558",
  "author_name": "Tucker Arrants",
  "post_date": "2021-06-15T22:05:46.184000",
  "votes": 38,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I recently released <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">a kernel</a> in which I try to implement 'learned image resizing' as proposed in <a href=\"https://arxiv.org/abs/2103.09950\" target=\"_blank\">this paper</a>. Instead of resizing images via an interpolation algorithm, we train a model to resize images jointly with the main training task. A very simple way of doing this would be to take <code>512 x 512</code> input images and convolve them to <code>256 x 256</code> before passing them to the main CNN backbone, which was done successfully <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118255\" target=\"_blank\">here</a>. </p>\n<p>The code for the learned image resizer in PyTorch is:</p>\n<pre><code>class CNNWithResizer(nn.Module):\n    def __init__(self, cfg, pretrained=False):\n        super().__init__()\n        self.cfg = cfg\n        self.n = 16\n        self.slope = .1\n        self.r = 1\n        self.cnn = timm.create_model(self.cfg.model_name, pretrained=pretrained, in_chans=1)\n\n        if hasattr(self.cnn, \"fc\"):\n            nb_ft = self.cnn.fc.in_features\n            self.cnn.fc = nn.Identity()\n        elif hasattr(self.cnn, \"_fc\"):\n            nb_ft = self.cnn._fc.in_features\n            self.cnn._fc = nn.Identity()\n        elif hasattr(self.cnn, \"classifier\"):\n            nb_ft = self.cnn.classifier.in_features\n            self.cnn.classifier = nn.Identity()\n        elif hasattr(self.cnn, \"last_linear\"):\n            nb_ft = self.cnn.last_linear.in_features\n            self.cnn.last_linear = nn.Identity()\n        elif hasattr(self.cnn, \"head\"):\n            nb_ft = self.cnn.head.in_features\n            self.cnn.head = nn.Identity()\n\n        self.block1 = nn.Sequential(\n                nn.Conv2d(1, self.n, kernel_size=(7, 7), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(1, 1), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.BatchNorm2d(self.n))\n        self.block2 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block3 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block4 = nn.Sequential(\n                nn.Conv2d(self.n, 1, kernel_size=(7, 7), stride=(1, 1), padding=(3, 3), bias=False))\n        self.fc = nn.Linear(nb_ft, self.cfg.target_size)\n\n    def forward(self, x):\n        res1 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block1(x)\n        res2 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block2(res2)\n        x += res2\n        if self.r &gt; 1:\n            for _ in range(self.r):\n                res2 = x\n                x = self.block2(x)\n                x += res2\n        x = self.block3(x)\n        x += res2\n        x = self.block4(x)\n        x += res1\n        x = self.cnn(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n<p>I am running a set of experiments to compare performance and will have the results by the end of the week. Below are my results so far:</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Resizer input size</th>\n<th>Resizer output size</th>\n<th>Batch size</th>\n<th>Folds</th>\n<th>CV</th>\n<th>Time (epoch)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>Standard resizing</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9869</td>\n<td>125 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>384 x 384</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9871</td>\n<td>338 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>256 x  819</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9884</td>\n<td>380 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>512 x 512</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td><strong>0.9904</strong></td>\n<td>402 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>768 x 768</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9902</td>\n<td>570 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>256 x  819</td>\n<td>320 x 320</td>\n<td>64</td>\n<td>5</td>\n<td>0.9895</td>\n<td>525 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>512 x  512</td>\n<td>320 x 320</td>\n<td>64</td>\n<td>5</td>\n<td>0.9896</td>\n<td>547 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>640 x  640</td>\n<td>320 x 320</td>\n<td>40</td>\n<td>5</td>\n<td>0.9896</td>\n<td>634 secs</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, all models were trained on a Quadro RTX 8000.</p>",
  "messages": [
    {
      "id": 1350894,
      "postDate": "2021-06-15T22:05:46.183Z",
      "content": "<p>I recently released <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">a kernel</a> in which I try to implement 'learned image resizing' as proposed in <a href=\"https://arxiv.org/abs/2103.09950\" target=\"_blank\">this paper</a>. Instead of resizing images via an interpolation algorithm, we train a model to resize images jointly with the main training task. A very simple way of doing this would be to take <code>512 x 512</code> input images and convolve them to <code>256 x 256</code> before passing them to the main CNN backbone, which was done successfully <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118255\" target=\"_blank\">here</a>. </p>\n<p>The code for the learned image resizer in PyTorch is:</p>\n<pre><code>class CNNWithResizer(nn.Module):\n    def __init__(self, cfg, pretrained=False):\n        super().__init__()\n        self.cfg = cfg\n        self.n = 16\n        self.slope = .1\n        self.r = 1\n        self.cnn = timm.create_model(self.cfg.model_name, pretrained=pretrained, in_chans=1)\n\n        if hasattr(self.cnn, \"fc\"):\n            nb_ft = self.cnn.fc.in_features\n            self.cnn.fc = nn.Identity()\n        elif hasattr(self.cnn, \"_fc\"):\n            nb_ft = self.cnn._fc.in_features\n            self.cnn._fc = nn.Identity()\n        elif hasattr(self.cnn, \"classifier\"):\n            nb_ft = self.cnn.classifier.in_features\n            self.cnn.classifier = nn.Identity()\n        elif hasattr(self.cnn, \"last_linear\"):\n            nb_ft = self.cnn.last_linear.in_features\n            self.cnn.last_linear = nn.Identity()\n        elif hasattr(self.cnn, \"head\"):\n            nb_ft = self.cnn.head.in_features\n            self.cnn.head = nn.Identity()\n\n        self.block1 = nn.Sequential(\n                nn.Conv2d(1, self.n, kernel_size=(7, 7), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(1, 1), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.BatchNorm2d(self.n))\n        self.block2 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block3 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block4 = nn.Sequential(\n                nn.Conv2d(self.n, 1, kernel_size=(7, 7), stride=(1, 1), padding=(3, 3), bias=False))\n        self.fc = nn.Linear(nb_ft, self.cfg.target_size)\n\n    def forward(self, x):\n        res1 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block1(x)\n        res2 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block2(res2)\n        x += res2\n        if self.r &gt; 1:\n            for _ in range(self.r):\n                res2 = x\n                x = self.block2(x)\n                x += res2\n        x = self.block3(x)\n        x += res2\n        x = self.block4(x)\n        x += res1\n        x = self.cnn(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n<p>I am running a set of experiments to compare performance and will have the results by the end of the week. Below are my results so far:</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Resizer input size</th>\n<th>Resizer output size</th>\n<th>Batch size</th>\n<th>Folds</th>\n<th>CV</th>\n<th>Time (epoch)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B0</td>\n<td>Standard resizing</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9869</td>\n<td>125 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>384 x 384</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9871</td>\n<td>338 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>256 x  819</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9884</td>\n<td>380 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>512 x 512</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td><strong>0.9904</strong></td>\n<td>402 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>768 x 768</td>\n<td>256 x 256</td>\n<td>64</td>\n<td>5</td>\n<td>0.9902</td>\n<td>570 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>256 x  819</td>\n<td>320 x 320</td>\n<td>64</td>\n<td>5</td>\n<td>0.9895</td>\n<td>525 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>512 x  512</td>\n<td>320 x 320</td>\n<td>64</td>\n<td>5</td>\n<td>0.9896</td>\n<td>547 secs</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>640 x  640</td>\n<td>320 x 320</td>\n<td>40</td>\n<td>5</td>\n<td>0.9896</td>\n<td>634 secs</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, all models were trained on a Quadro RTX 8000.</p>",
      "rawMarkdown": "I recently released [a kernel](https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing) in which I try to implement 'learned image resizing' as proposed in [this paper](https://arxiv.org/abs/2103.09950). Instead of resizing images via an interpolation algorithm, we train a model to resize images jointly with the main training task. A very simple way of doing this would be to take `512 x 512` input images and convolve them to `256 x 256` before passing them to the main CNN backbone, which was done successfully [here](https://www.kaggle.com/c/understanding_cloud_organization/discussion/118255). \n\nThe code for the learned image resizer in PyTorch is:\n\n<pre><code>class CNNWithResizer(nn.Module):\n    def __init__(self, cfg, pretrained=False):\n        super().__init__()\n        self.cfg = cfg\n        self.n = 16\n        self.slope = .1\n        self.r = 1\n        self.cnn = timm.create_model(self.cfg.model_name, pretrained=pretrained, in_chans=1)\n\n        if hasattr(self.cnn, \"fc\"):\n            nb_ft = self.cnn.fc.in_features\n            self.cnn.fc = nn.Identity()\n        elif hasattr(self.cnn, \"_fc\"):\n            nb_ft = self.cnn._fc.in_features\n            self.cnn._fc = nn.Identity()\n        elif hasattr(self.cnn, \"classifier\"):\n            nb_ft = self.cnn.classifier.in_features\n            self.cnn.classifier = nn.Identity()\n        elif hasattr(self.cnn, \"last_linear\"):\n            nb_ft = self.cnn.last_linear.in_features\n            self.cnn.last_linear = nn.Identity()\n        elif hasattr(self.cnn, \"head\"):\n            nb_ft = self.cnn.head.in_features\n            self.cnn.head = nn.Identity()\n        \n        self.block1 = nn.Sequential(\n                nn.Conv2d(1, self.n, kernel_size=(7, 7), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(1, 1), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.BatchNorm2d(self.n))\n        self.block2 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block3 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block4 = nn.Sequential(\n                nn.Conv2d(self.n, 1, kernel_size=(7, 7), stride=(1, 1), padding=(3, 3), bias=False))\n        self.fc = nn.Linear(nb_ft, self.cfg.target_size)\n\n    def forward(self, x):\n        res1 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block1(x)\n        res2 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block2(res2)\n        x += res2\n        if self.r > 1:\n            for _ in range(self.r):\n                res2 = x\n                x = self.block2(x)\n                x += res2\n        x = self.block3(x)\n        x += res2\n        x = self.block4(x)\n        x += res1\n        x = self.cnn(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n\nI am running a set of experiments to compare performance and will have the results by the end of the week. Below are my results so far:\n\n| Encoder | Resizer input size | Resizer output size | Batch size | Folds| CV | Time (epoch)|\n| -------- | ---------- | ------------ | ----------- |  ---- | --- | ------ |\n| B0 | Standard resizing | 256 x 256| 64 | 5 | 0.9869 |125 secs|\n| B0 | 384 x 384| 256 x 256| 64 | 5 | 0.9871|338 secs|\n| B0 | 256 x  819 | 256 x 256| 64 | 5 | 0.9884 |380 secs|\n| B0 | 512 x 512| 256 x 256| 64 | 5 | **0.9904**|402 secs|\n| B0 | 768 x 768| 256 x 256| 64 | 5 | 0.9902 |570 secs|\n| B0 | 256 x  819| 320 x 320| 64 | 5 | 0.9895 |525 secs|\n| B0 | 512 x  512| 320 x 320| 64 | 5 | 0.9896|547 secs|\n| B0 | 640 x  640| 320 x 320| 40 | 5 | 0.9896|634 secs|\n\n\nFor reference, all models were trained on a Quadro RTX 8000.",
      "votes": 37
    },
    {
      "id": 1451991,
      "postDate": "2021-08-05T13:46:09.787Z",
      "content": "<p>it's a good story, but I cannot find any difference between the \"head for learning image resize\" and original \"3<em>3 head with stride=2\", if I remove the first layer in ResNet18 for 112</em>112 images, and then recover it for 224, that means I propose \"learned image resize?\" it's a totally funny idea. </p>",
      "rawMarkdown": "it's a good story, but I cannot find any difference between the \"head for learning image resize\" and original \"3*3 head with stride=2\", if I remove the first layer in ResNet18 for 112*112 images, and then recover it for 224, that means I propose \"learned image resize?\" it's a totally funny idea. ",
      "votes": 3
    },
    {
      "id": 1358347,
      "postDate": "2021-06-20T11:40:44.980Z",
      "content": "<p>I think this is one of the most important topic.<br>\nHow small can we reduce the image size (maintaining needle size)?</p>",
      "rawMarkdown": "I think this is one of the most important topic.\nHow small can we reduce the image size (maintaining needle size)?",
      "votes": 1
    },
    {
      "id": 1355610,
      "postDate": "2021-06-18T12:11:26.187Z",
      "content": "<p>very interesting -thx.<br>\nit would be interesting comparing this to native 512x512, and orig-&gt;512x512 and 640x640-&gt;512x512. Probably would need some hp tuning as the paper states…</p>",
      "rawMarkdown": "very interesting -thx.\nit would be interesting comparing this to native 512x512, and orig->512x512 and 640x640->512x512. Probably would need some hp tuning as the paper states...",
      "votes": 1,
      "replies": [
        {
          "id": 1355647,
          "postDate": "2021-06-18T12:50:40.683Z",
          "content": "<p>I have compared regular <code>512x512</code> to <code>1024x1024-&gt;512x512</code> and observed a <code>.002</code> CV increase. I will test <code>768x768-&gt;512x512</code> and <code>640x640-&gt;512x512</code> but probably won't have the results for a week or so. </p>\n<p><code>Original-&gt;512x512</code> is interesting…technically we would be upscaling one axis and downscaling the other, but total # of pixels is higher for <code>512x512</code> than original resolution. In this case, the bilinear feature resizer would act as an inverse bottleneck. </p>",
          "rawMarkdown": "I have compared regular `512x512` to `1024x1024->512x512` and observed a `.002` CV increase. I will test `768x768->512x512` and `640x640->512x512` but probably won't have the results for a week or so. \n\n`Original->512x512` is interesting...technically we would be upscaling one axis and downscaling the other, but total # of pixels is higher for `512x512` than original resolution. In this case, the bilinear feature resizer would act as an inverse bottleneck. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1354153,
      "postDate": "2021-06-17T12:09:48.533Z",
      "content": "<p>An interesting extension to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's discussion before :<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195</a></p>\n<p>Just curious if you are willing to share the approximate Epoch value in which your model starts to converge while testing?</p>\n<p>Thanks for Sharing 😊</p>",
      "rawMarkdown": "An interesting extension to @ttahara 's discussion before :\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195\n\nJust curious if you are willing to share the approximate Epoch value in which your model starts to converge while testing?\n\nThanks for Sharing 😊",
      "votes": 1,
      "replies": [
        {
          "id": 1354178,
          "postDate": "2021-06-17T12:28:28.853Z",
          "content": "<p>Convergence is around 20-25 epochs. All models above were trained for 25 epochs.</p>",
          "rawMarkdown": "Convergence is around 20-25 epochs. All models above were trained for 25 epochs.",
          "votes": 2
        },
        {
          "id": 1354180,
          "postDate": "2021-06-17T12:30:06.653Z",
          "content": "<p>Thanks! May try this on larger models (If I ever get enough GPU hours D;).</p>",
          "rawMarkdown": "Thanks! May try this on larger models (If I ever get enough GPU hours D;).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1354081,
      "postDate": "2021-06-17T11:26:20.703Z",
      "content": "<p>it's very useful thanks for sharing this idea👍 <a href=\"https://www.kaggle.com/Tucker\" target=\"_blank\">@Tucker</a></p>",
      "rawMarkdown": "it's very useful thanks for sharing this idea👍 @Tucker",
      "votes": 1
    },
    {
      "id": 1354018,
      "postDate": "2021-06-17T10:40:05.520Z",
      "content": "<p>Thanks for sharing this idea, could you share the results of the traditional opencv resize process to make a clearer comparison between traditional resizing and learned resizing?</p>",
      "rawMarkdown": "Thanks for sharing this idea, could you share the results of the traditional opencv resize process to make a clearer comparison between traditional resizing and learned resizing?",
      "votes": 1,
      "replies": [
        {
          "id": 1354138,
          "postDate": "2021-06-17T12:05:01.840Z",
          "content": "<p>Of course, I will have results comparing standard resizing to learned resizing by tomorrow morning. I will also re-run all experiments once the host releases new data.</p>",
          "rawMarkdown": "Of course, I will have results comparing standard resizing to learned resizing by tomorrow morning. I will also re-run all experiments once the host releases new data.",
          "votes": 3
        },
        {
          "id": 1354184,
          "postDate": "2021-06-17T12:32:58.353Z",
          "content": "<p>Thanks for your reply</p>",
          "rawMarkdown": "Thanks for your reply",
          "votes": 1
        }
      ]
    },
    {
      "id": 1350937,
      "postDate": "2021-06-16T00:17:21.643Z",
      "content": "<p>Wow. Thanks that’s something I want to try now haha</p>",
      "rawMarkdown": "Wow. Thanks that’s something I want to try now haha",
      "votes": 1
    },
    {
      "id": 1397024,
      "postDate": "2021-07-22T17:30:58.800Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> , looks you still use A.Resize() to convert original (256, 6x273) image to (512, 512) before you feed to CNNWithResizer model, but this A.Resize() is still interpolation algo, I`m wondering if this A.Resize is necessary here, or directly feed original (256, 6*273) to CNNWithResizer model is better, could you please help to explain a bit? Thanks.</p>",
      "rawMarkdown": "Thanks for sharing @tuckerarrants , looks you still use A.Resize() to convert original (256, 6x273) image to (512, 512) before you feed to CNNWithResizer model, but this A.Resize() is still interpolation algo, I`m wondering if this A.Resize is necessary here, or directly feed original (256, 6*273) to CNNWithResizer model is better, could you please help to explain a bit? Thanks.",
      "replies": [
        {
          "id": 1397078,
          "postDate": "2021-07-22T18:38:23.353Z",
          "content": "<p>You can experiment with both. I got slightly better results resizing to <code>512 x 512</code> first (with standard interpolation) and then downsizing to <code>256 x 256</code> with the learned resizer. In this case, I downsized it to save training time. </p>\n<p>If you want to feed the CNN backbone original resolution features, then you would not resize them first. Best thing to do is try both and compare their CV results. </p>",
          "rawMarkdown": "You can experiment with both. I got slightly better results resizing to `512 x 512` first (with standard interpolation) and then downsizing to `256 x 256` with the learned resizer. In this case, I downsized it to save training time. \n\nIf you want to feed the CNN backbone original resolution features, then you would not resize them first. Best thing to do is try both and compare their CV results. "
        },
        {
          "id": 1398998,
          "postDate": "2021-07-24T18:39:53.277Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> . When I experiment Learnable resizer I found a strange issue: numerical instable, every time I rerun the same code I got different loss and CV value. My implementation is follow this github <a href=\"https://github.com/KushajveerSingh/resize_network_cv\" target=\"_blank\">link</a>, I checked the code and didn`t find anything wrong. May I know whether you also met numerical stability issue with your Learnable resizer code? Thanks.</p>",
          "rawMarkdown": "Thanks @tuckerarrants . When I experiment Learnable resizer I found a strange issue: numerical instable, every time I rerun the same code I got different loss and CV value. My implementation is follow this github [link](https://github.com/KushajveerSingh/resize_network_cv), I checked the code and didn`t find anything wrong. May I know whether you also met numerical stability issue with your Learnable resizer code? Thanks."
        },
        {
          "id": 1400836,
          "postDate": "2021-07-26T16:16:08.007Z",
          "content": "<p>I do not have any issues with numerical instability. I have run the above linked kernel several times and gotten identical results. </p>",
          "rawMarkdown": "I do not have any issues with numerical instability. I have run the above linked kernel several times and gotten identical results. "
        },
        {
          "id": 1401317,
          "postDate": "2021-07-27T07:56:47.750Z",
          "content": "<p>I have addressed that my numerical instability issue is due to F.interpolate have nondeterministic behaviour when back-prop with up-sampling. You code never need up-sample in your CNNWithResizer as your input to resizer is always bigger than output. Thanks for your help.</p>\n<p>And another interesting found is that if I disable grad(with torch.no_grad) for F.interpolate, then this learnable resizer is useless at all, maybe the key of the learnable reisze is to back-prob on F.interpolate?</p>",
          "rawMarkdown": "I have addressed that my numerical instability issue is due to F.interpolate have nondeterministic behaviour when back-prop with up-sampling. You code never need up-sample in your CNNWithResizer as your input to resizer is always bigger than output. Thanks for your help.\n\nAnd another interesting found is that if I disable grad(with torch.no_grad) for F.interpolate, then this learnable resizer is useless at all, maybe the key of the learnable reisze is to back-prob on F.interpolate?",
          "votes": 2
        }
      ]
    },
    {
      "id": 1353037,
      "postDate": "2021-06-16T19:34:48.610Z",
      "content": "<p>Nice <a href=\"https://www.kaggle.com/Tucker\" target=\"_blank\">@Tucker</a> Arrants Thanks for sharing</p>",
      "rawMarkdown": "Nice @Tucker Arrants Thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1451991,
      "author_name": "SleepMaster",
      "author_url": "",
      "post_date": "2021-08-05T13:46:09.787000",
      "content": "<p>it's a good story, but I cannot find any difference between the \"head for learning image resize\" and original \"3<em>3 head with stride=2\", if I remove the first layer in ResNet18 for 112</em>112 images, and then recover it for 224, that means I propose \"learned image resize?\" it's a totally funny idea. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1358347,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-20T11:40:44.980000",
      "content": "<p>I think this is one of the most important topic.<br>\nHow small can we reduce the image size (maintaining needle size)?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1355610,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2021-06-18T12:11:26.187000",
      "content": "<p>very interesting -thx.<br>\nit would be interesting comparing this to native 512x512, and orig-&gt;512x512 and 640x640-&gt;512x512. Probably would need some hp tuning as the paper states…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1355647,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-06-18T12:50:40.683000",
          "content": "<p>I have compared regular <code>512x512</code> to <code>1024x1024-&gt;512x512</code> and observed a <code>.002</code> CV increase. I will test <code>768x768-&gt;512x512</code> and <code>640x640-&gt;512x512</code> but probably won't have the results for a week or so. </p>\n<p><code>Original-&gt;512x512</code> is interesting…technically we would be upscaling one axis and downscaling the other, but total # of pixels is higher for <code>512x512</code> than original resolution. In this case, the bilinear feature resizer would act as an inverse bottleneck. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1354153,
      "author_name": "maze508",
      "author_url": "",
      "post_date": "2021-06-17T12:09:48.533000",
      "content": "<p>An interesting extension to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's discussion before :<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195</a></p>\n<p>Just curious if you are willing to share the approximate Epoch value in which your model starts to converge while testing?</p>\n<p>Thanks for Sharing 😊</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1354178,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-06-17T12:28:28.853000",
          "content": "<p>Convergence is around 20-25 epochs. All models above were trained for 25 epochs.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1354180,
          "author_name": "maze508",
          "author_url": "",
          "post_date": "2021-06-17T12:30:06.653000",
          "content": "<p>Thanks! May try this on larger models (If I ever get enough GPU hours D;).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1354081,
      "author_name": "M.Fauzan Alfariz",
      "author_url": "",
      "post_date": "2021-06-17T11:26:20.703000",
      "content": "<p>it's very useful thanks for sharing this idea👍 <a href=\"https://www.kaggle.com/Tucker\" target=\"_blank\">@Tucker</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1354018,
      "author_name": "YTEP (Jiazhi Yang)",
      "author_url": "",
      "post_date": "2021-06-17T10:40:05.520000",
      "content": "<p>Thanks for sharing this idea, could you share the results of the traditional opencv resize process to make a clearer comparison between traditional resizing and learned resizing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1354138,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-06-17T12:05:01.840000",
          "content": "<p>Of course, I will have results comparing standard resizing to learned resizing by tomorrow morning. I will also re-run all experiments once the host releases new data.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1354184,
          "author_name": "YTEP (Jiazhi Yang)",
          "author_url": "",
          "post_date": "2021-06-17T12:32:58.353000",
          "content": "<p>Thanks for your reply</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1350937,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-06-16T00:17:21.643000",
      "content": "<p>Wow. Thanks that’s something I want to try now haha</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1397024,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-07-22T17:30:58.800000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> , looks you still use A.Resize() to convert original (256, 6x273) image to (512, 512) before you feed to CNNWithResizer model, but this A.Resize() is still interpolation algo, I`m wondering if this A.Resize is necessary here, or directly feed original (256, 6*273) to CNNWithResizer model is better, could you please help to explain a bit? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1397078,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-07-22T18:38:23.353000",
          "content": "<p>You can experiment with both. I got slightly better results resizing to <code>512 x 512</code> first (with standard interpolation) and then downsizing to <code>256 x 256</code> with the learned resizer. In this case, I downsized it to save training time. </p>\n<p>If you want to feed the CNN backbone original resolution features, then you would not resize them first. Best thing to do is try both and compare their CV results. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1398998,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-07-24T18:39:53.277000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> . When I experiment Learnable resizer I found a strange issue: numerical instable, every time I rerun the same code I got different loss and CV value. My implementation is follow this github <a href=\"https://github.com/KushajveerSingh/resize_network_cv\" target=\"_blank\">link</a>, I checked the code and didn`t find anything wrong. May I know whether you also met numerical stability issue with your Learnable resizer code? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1400836,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-07-26T16:16:08.007000",
          "content": "<p>I do not have any issues with numerical instability. I have run the above linked kernel several times and gotten identical results. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1401317,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-07-27T07:56:47.750000",
          "content": "<p>I have addressed that my numerical instability issue is due to F.interpolate have nondeterministic behaviour when back-prop with up-sampling. You code never need up-sample in your CNNWithResizer as your input to resizer is always bigger than output. Thanks for your help.</p>\n<p>And another interesting found is that if I disable grad(with torch.no_grad) for F.interpolate, then this learnable resizer is useless at all, maybe the key of the learnable reisze is to back-prob on F.interpolate?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1353037,
      "author_name": "Faisal Qureshi",
      "author_url": "",
      "post_date": "2021-06-16T19:34:48.610000",
      "content": "<p>Nice <a href=\"https://www.kaggle.com/Tucker\" target=\"_blank\">@Tucker</a> Arrants Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1350894": "I recently released [a kernel](https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing) in which I try to implement 'learned image resizing' as proposed in [this paper](https://arxiv.org/abs/2103.09950). Instead of resizing images via an interpolation algorithm, we train a model to resize images jointly with the main training task. A very simple way of doing this would be to take `512 x 512` input images and convolve them to `256 x 256` before passing them to the main CNN backbone, which was done successfully [here](https://www.kaggle.com/c/understanding_cloud_organization/discussion/118255). \n\nThe code for the learned image resizer in PyTorch is:\n\n<pre><code>class CNNWithResizer(nn.Module):\n    def __init__(self, cfg, pretrained=False):\n        super().__init__()\n        self.cfg = cfg\n        self.n = 16\n        self.slope = .1\n        self.r = 1\n        self.cnn = timm.create_model(self.cfg.model_name, pretrained=pretrained, in_chans=1)\n\n        if hasattr(self.cnn, \"fc\"):\n            nb_ft = self.cnn.fc.in_features\n            self.cnn.fc = nn.Identity()\n        elif hasattr(self.cnn, \"_fc\"):\n            nb_ft = self.cnn._fc.in_features\n            self.cnn._fc = nn.Identity()\n        elif hasattr(self.cnn, \"classifier\"):\n            nb_ft = self.cnn.classifier.in_features\n            self.cnn.classifier = nn.Identity()\n        elif hasattr(self.cnn, \"last_linear\"):\n            nb_ft = self.cnn.last_linear.in_features\n            self.cnn.last_linear = nn.Identity()\n        elif hasattr(self.cnn, \"head\"):\n            nb_ft = self.cnn.head.in_features\n            self.cnn.head = nn.Identity()\n        \n        self.block1 = nn.Sequential(\n                nn.Conv2d(1, self.n, kernel_size=(7, 7), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(1, 1), stride=(1,1), padding=(1, 1), bias=False),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.BatchNorm2d(self.n))\n        self.block2 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n),\n                nn.LeakyReLU(negative_slope=self.slope),\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block3 = nn.Sequential(\n                nn.Conv2d(self.n, self.n, kernel_size=(3, 3), stride=(1,1), padding=(1, 1), bias=False),\n                nn.BatchNorm2d(self.n))\n        self.block4 = nn.Sequential(\n                nn.Conv2d(self.n, 1, kernel_size=(7, 7), stride=(1, 1), padding=(3, 3), bias=False))\n        self.fc = nn.Linear(nb_ft, self.cfg.target_size)\n\n    def forward(self, x):\n        res1 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block1(x)\n        res2 = F.interpolate(x, size=(256, 256), mode='bilinear')\n        x = self.block2(res2)\n        x += res2\n        if self.r > 1:\n            for _ in range(self.r):\n                res2 = x\n                x = self.block2(x)\n                x += res2\n        x = self.block3(x)\n        x += res2\n        x = self.block4(x)\n        x += res1\n        x = self.cnn(x)\n        x = self.fc(x)\n        return x\n</code></pre>\n\nI am running a set of experiments to compare performance and will have the results by the end of the week. Below are my results so far:\n\n| Encoder | Resizer input size | Resizer output size | Batch size | Folds| CV | Time (epoch)|\n| -------- | ---------- | ------------ | ----------- |  ---- | --- | ------ |\n| B0 | Standard resizing | 256 x 256| 64 | 5 | 0.9869 |125 secs|\n| B0 | 384 x 384| 256 x 256| 64 | 5 | 0.9871|338 secs|\n| B0 | 256 x  819 | 256 x 256| 64 | 5 | 0.9884 |380 secs|\n| B0 | 512 x 512| 256 x 256| 64 | 5 | **0.9904**|402 secs|\n| B0 | 768 x 768| 256 x 256| 64 | 5 | 0.9902 |570 secs|\n| B0 | 256 x  819| 320 x 320| 64 | 5 | 0.9895 |525 secs|\n| B0 | 512 x  512| 320 x 320| 64 | 5 | 0.9896|547 secs|\n| B0 | 640 x  640| 320 x 320| 40 | 5 | 0.9896|634 secs|\n\n\nFor reference, all models were trained on a Quadro RTX 8000.",
    "1451991": "it's a good story, but I cannot find any difference between the \"head for learning image resize\" and original \"3*3 head with stride=2\", if I remove the first layer in ResNet18 for 112*112 images, and then recover it for 224, that means I propose \"learned image resize?\" it's a totally funny idea. ",
    "1358347": "I think this is one of the most important topic.\nHow small can we reduce the image size (maintaining needle size)?",
    "1355610": "very interesting -thx.\nit would be interesting comparing this to native 512x512, and orig->512x512 and 640x640->512x512. Probably would need some hp tuning as the paper states...",
    "1354153": "An interesting extension to @ttahara 's discussion before :\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195\n\nJust curious if you are willing to share the approximate Epoch value in which your model starts to converge while testing?\n\nThanks for Sharing 😊",
    "1354081": "it's very useful thanks for sharing this idea👍 @Tucker",
    "1354018": "Thanks for sharing this idea, could you share the results of the traditional opencv resize process to make a clearer comparison between traditional resizing and learned resizing?",
    "1350937": "Wow. Thanks that’s something I want to try now haha",
    "1397024": "Thanks for sharing @tuckerarrants , looks you still use A.Resize() to convert original (256, 6x273) image to (512, 512) before you feed to CNNWithResizer model, but this A.Resize() is still interpolation algo, I`m wondering if this A.Resize is necessary here, or directly feed original (256, 6*273) to CNNWithResizer model is better, could you please help to explain a bit? Thanks.",
    "1353037": "Nice @Tucker Arrants Thanks for sharing"
  }
}