{
  "id": 468437,
  "title": "On the nature of training with different size non-square images, and why you want to stick to the square ones",
  "url": "/competitions/blood-vessel-segmentation/discussion/468437",
  "author_name": "SSS",
  "post_date": "2024-01-16T16:43:59.321000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi, <br>\n(<strong>TL:DR square images perform better</strong>)<br>\nI want to share some thoughts after running some experiments with square and non-square images. <br>\nIt might be intuitive for some of you to use square images during the training. Though I tried to understand what really happens.<br>\nModern ML frameworks and architectures let you easily fit various size/shape images into it. Lets look at VGG11 first layer Conv2d(), pass some input and see what happens.</p>\n<pre><code>()\nconv2d = torch.nn.Conv2d(, , kernel_size=(, ), stride=(, ), padding=(, ))\nc0 = conv2d(batch0)\n()\n(</code></pre>\n<p>Well, the model learns weights of 64 filters for each image, nothing really interesting. Like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ff1d512c5a9a3845de350398da4a242c4%2Fconv.png?generation=1705422498580232&amp;alt=media\"><br>\nWhat happens next is the MaxPooling2d:</p>\n<pre><code>pool2d = torch.nn.MaxPool2d(kernel_size=, stride=, padding=, dilation=, ceil_mode=)\n(pool2d(c0).shape)\ntorch.Size([, , , ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4da799bb21c482323f673d0643463a27%2Fpool.png?generation=1705422725550285&amp;alt=media\"><br>\n<strong>Now imagine you are passing non square image of different sizes</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F3738d3eeaf5bf3d730b82adf690ba042%2Fnon%20square..png?generation=1705423494882827&amp;alt=media\"><br>\nThe quadrants will partially share the local pixel stats of the other quadrants which won`t happen in the square images. The same pooling layer filter leaks stats from other quadrants and adding more images of the different shape will increase the variance even more. It is fine when there is a max pooling applied, the model at least learn something. But once I change it to avg pooling, the dice score becomes 0. This also sheds some light on why training along only z-axis yields better results for me. <br>\nI hope that make sense, I might edit the post, since I still trying to frame these findings the right way.</p>",
  "messages": [
    {
      "id": 2604789,
      "postDate": "2024-01-16T16:43:59.320Z",
      "content": "<p>Hi, <br>\n(<strong>TL:DR square images perform better</strong>)<br>\nI want to share some thoughts after running some experiments with square and non-square images. <br>\nIt might be intuitive for some of you to use square images during the training. Though I tried to understand what really happens.<br>\nModern ML frameworks and architectures let you easily fit various size/shape images into it. Lets look at VGG11 first layer Conv2d(), pass some input and see what happens.</p>\n<pre><code>()\nconv2d = torch.nn.Conv2d(, , kernel_size=(, ), stride=(, ), padding=(, ))\nc0 = conv2d(batch0)\n()\n(</code></pre>\n<p>Well, the model learns weights of 64 filters for each image, nothing really interesting. Like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ff1d512c5a9a3845de350398da4a242c4%2Fconv.png?generation=1705422498580232&amp;alt=media\"><br>\nWhat happens next is the MaxPooling2d:</p>\n<pre><code>pool2d = torch.nn.MaxPool2d(kernel_size=, stride=, padding=, dilation=, ceil_mode=)\n(pool2d(c0).shape)\ntorch.Size([, , , ])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4da799bb21c482323f673d0643463a27%2Fpool.png?generation=1705422725550285&amp;alt=media\"><br>\n<strong>Now imagine you are passing non square image of different sizes</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F3738d3eeaf5bf3d730b82adf690ba042%2Fnon%20square..png?generation=1705423494882827&amp;alt=media\"><br>\nThe quadrants will partially share the local pixel stats of the other quadrants which won`t happen in the square images. The same pooling layer filter leaks stats from other quadrants and adding more images of the different shape will increase the variance even more. It is fine when there is a max pooling applied, the model at least learn something. But once I change it to avg pooling, the dice score becomes 0. This also sheds some light on why training along only z-axis yields better results for me. <br>\nI hope that make sense, I might edit the post, since I still trying to frame these findings the right way.</p>",
      "rawMarkdown": "\nHi, \n\n(**TL:DR square images perform better**)\nI want to share some thoughts after running some experiments with square and non-square images. \n\nIt might be intuitive for some of you to use square images during the training. Though I tried to understand what really happens.\nModern ML frameworks and architectures let you easily fit various size/shape images into it. Lets look at VGG11 first layer Conv2d(), pass some input and see what happens.\n\n```python\nprint(f'->VGG11 first conv = {model.model.encoder.features[0]}')\nconv2d = torch.nn.Conv2d(1, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))\nc0 = conv2d(batch0)\nprint(f'->Shapes {batch0[0].shape=}')\nprint(f'->Shapes after conv {c0.shape=}\n\n\n>> ->VGG11 first conv = Conv2d(1, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))\n>> ->Shapes  batch0[0].shape=torch.Size([1, 1312, 928])\n>> ->Shapes after conv c0.shape=torch.Size([16, 64, 1312, 928])\n\n```\nWell, the model learns weights of 64 filters for each image, nothing really interesting. Like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ff1d512c5a9a3845de350398da4a242c4%2Fconv.png?generation=1705422498580232&alt=media)\n\nWhat happens next is the MaxPooling2d:\n```python\n>>> pool2d = torch.nn.MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False)\n... print(pool2d(c0).shape)\ntorch.Size([16, 64, 656, 464])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4da799bb21c482323f673d0643463a27%2Fpool.png?generation=1705422725550285&alt=media)\n\n**Now imagine you are passing non square image of different sizes**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F3738d3eeaf5bf3d730b82adf690ba042%2Fnon%20square..png?generation=1705423494882827&alt=media)\n\nThe quadrants will partially share the local pixel stats of the other quadrants which won`t happen in the square images. The same pooling layer filter leaks stats from other quadrants and adding more images of the different shape will increase the variance even more. It is fine when there is a max pooling applied, the model at least learn something. But once I change it to avg pooling, the dice score becomes 0. This also sheds some light on why training along only z-axis yields better results for me. \n\nI hope that make sense, I might edit the post, since I still trying to frame these findings the right way.",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2604789": "\nHi, \n\n(**TL:DR square images perform better**)\nI want to share some thoughts after running some experiments with square and non-square images. \n\nIt might be intuitive for some of you to use square images during the training. Though I tried to understand what really happens.\nModern ML frameworks and architectures let you easily fit various size/shape images into it. Lets look at VGG11 first layer Conv2d(), pass some input and see what happens.\n\n```python\nprint(f'->VGG11 first conv = {model.model.encoder.features[0]}')\nconv2d = torch.nn.Conv2d(1, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))\nc0 = conv2d(batch0)\nprint(f'->Shapes {batch0[0].shape=}')\nprint(f'->Shapes after conv {c0.shape=}\n\n\n>> ->VGG11 first conv = Conv2d(1, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1))\n>> ->Shapes  batch0[0].shape=torch.Size([1, 1312, 928])\n>> ->Shapes after conv c0.shape=torch.Size([16, 64, 1312, 928])\n\n```\nWell, the model learns weights of 64 filters for each image, nothing really interesting. Like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Ff1d512c5a9a3845de350398da4a242c4%2Fconv.png?generation=1705422498580232&alt=media)\n\nWhat happens next is the MaxPooling2d:\n```python\n>>> pool2d = torch.nn.MaxPool2d(kernel_size=2, stride=2, padding=0, dilation=1, ceil_mode=False)\n... print(pool2d(c0).shape)\ntorch.Size([16, 64, 656, 464])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F4da799bb21c482323f673d0643463a27%2Fpool.png?generation=1705422725550285&alt=media)\n\n**Now imagine you are passing non square image of different sizes**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F3738d3eeaf5bf3d730b82adf690ba042%2Fnon%20square..png?generation=1705423494882827&alt=media)\n\nThe quadrants will partially share the local pixel stats of the other quadrants which won`t happen in the square images. The same pooling layer filter leaks stats from other quadrants and adding more images of the different shape will increase the variance even more. It is fine when there is a max pooling applied, the model at least learn something. But once I change it to avg pooling, the dice score becomes 0. This also sheds some light on why training along only z-axis yields better results for me. \n\nI hope that make sense, I might edit the post, since I still trying to frame these findings the right way."
  }
}