{
  "id": 234099,
  "title": "What Aspect Ratio Are You Using As Model Input And Why?",
  "url": "/competitions/bms-molecular-translation/discussion/234099",
  "author_name": "",
  "post_date": "2021-04-22T16:23:18.719760800Z",
  "votes": 14,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Hi there. As the title states, I'm curious to see what aspect ratios others are using.</p>\n<p>I know it has been (tentatively) <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233834\" target=\"_blank\">shown that a 2:1 aspect ratio (appropriately cropped) may be the most efficient method of capturing relevant information</a>.</p>\n<p>However, others have suggested that the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233927\" target=\"_blank\">margins might contain useable information to potentially recover ghost methyl groups</a>.</p>\n<p>Personally, I'm still just using a padded-to-square input image (<strong><code>384x384x3</code></strong>). This is probably not ideal, but I wanted to query other Kagglers to see their thought process. I plan on running a minor benchmarking comparison in the near future as well.</p>",
  "messages": [
    {
      "id": "1281104",
      "postDate": "04/22/2021 16:23:18",
      "content": "<p>Hi there. As the title states, I'm curious to see what aspect ratios others are using.</p>\n<p>I know it has been (tentatively) <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233834\" target=\"_blank\">shown that a 2:1 aspect ratio (appropriately cropped) may be the most efficient method of capturing relevant information</a>.</p>\n<p>However, others have suggested that the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/233927\" target=\"_blank\">margins might contain useable information to potentially recover ghost methyl groups</a>.</p>\n<p>Personally, I'm still just using a padded-to-square input image (<strong><code>384x384x3</code></strong>). This is probably not ideal, but I wanted to query other Kagglers to see their thought process. I plan on running a minor benchmarking comparison in the near future as well.</p>",
      "rawMarkdown": "Hi there. As the title states, I'm curious to see what aspect ratios others are using.\n\nI know it has been (tentatively) [shown that a 2:1 aspect ratio (appropriately cropped) may be the most efficient method of capturing relevant information](https://www.kaggle.com/c/bms-molecular-translation/discussion/233834).\n\nHowever, others have suggested that the [margins might contain useable information to potentially recover ghost methyl groups](https://www.kaggle.com/c/bms-molecular-translation/discussion/233927).\n\nPersonally, I'm still just using a padded-to-square input image (**`384x384x3`**). This is probably not ideal, but I wanted to query other Kagglers to see their thought process. I plan on running a minor benchmarking comparison in the near future as well.",
      "votes": null
    },
    {
      "id": "1281243",
      "postDate": "04/22/2021 18:32:01",
      "content": "<p>In my first experiments, I was using the 1.72 ratio as stated in <a href=\"https://www.kaggle.com/markwijkhuizen/advanced-image-cleaning-and-tfrecord-generation\" target=\"_blank\"> Mark's TF records notebook</a>. Then I noticed that this ratio is calculated before the cropping. Then I started to use the cropped ratio (2.17) based on the <a href=\"https://www.kaggle.com/michaelwolff/bms-inchi-cropped-img-sizes-for-best-resolution\" target=\"_blank\">resolution analysis notebook</a>.</p>\n<p>I am not sure about the effect of optimizing the ratio on the model performance, but it decreases the training time for sure.</p>",
      "rawMarkdown": "In my first experiments, I was using the 1.72 ratio as stated in [ Mark's TF records notebook](https://www.kaggle.com/markwijkhuizen/advanced-image-cleaning-and-tfrecord-generation). Then I noticed that this ratio is calculated before the cropping. Then I started to use the cropped ratio (2.17) based on the [resolution analysis notebook](https://www.kaggle.com/michaelwolff/bms-inchi-cropped-img-sizes-for-best-resolution).\n\nI am not sure about the effect of optimizing the ratio on the model performance, but it decreases the training time for sure.",
      "votes": null
    },
    {
      "id": "1281257",
      "postDate": "04/22/2021 18:44:16",
      "content": "<p>Awesome. Thanks for the info. I was a bit uncertain if rectangular input images would be ok to use… but from my research, and the feedback from everyone on Kaggle, it seems that square is just a choice dictated by simplicity and that I can change over my network to using rectangular inputs.</p>\n<p>Thanks again!</p>",
      "rawMarkdown": "Awesome. Thanks for the info. I was a bit uncertain if rectangular input images would be ok to use... but from my research, and the feedback from everyone on Kaggle, it seems that square is just a choice dictated by simplicity and that I can change over my network to using rectangular inputs.\n\nThanks again!",
      "votes": null
    },
    {
      "id": "1282373",
      "postDate": "04/23/2021 21:15:38",
      "content": "<p>Hi, <br>\nI also switched from Mark's ratio of 1.75 to 2.0 (512x256) and got a slight improvement on the LB of 0.4. That was with the 'old' attention model (version 3) in his notebook with LB scores of around 14.</p>\n<p>I also read that the input res should be tuned to the encoder model you want to use. For example EfficientNet B0 has a 'natural' resolution of 224. I.e. that is the resolution the network would expect if you used the whole pretrained model for image classification.</p>\n<p>I got that from these two sites, <a href=\"https://patrick-llgc.github.io/Learning-Deep-Learning/paper_notes/efficientnet.html\" target=\"_blank\">here</a> and <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">here</a>, also with the resolutions of the higher models. I'm not quite clear on what the possible side effects are when you use it with a higher resolution (without the final classification layer).</p>\n<p>An advantage of squares is that you can try 90 degree rotation augmentation.</p>",
      "rawMarkdown": "Hi, \nI also switched from Mark's ratio of 1.75 to 2.0 (512x256) and got a slight improvement on the LB of 0.4. That was with the 'old' attention model (version 3) in his notebook with LB scores of around 14.\n\nI also read that the input res should be tuned to the encoder model you want to use. For example EfficientNet B0 has a 'natural' resolution of 224. I.e. that is the resolution the network would expect if you used the whole pretrained model for image classification.\n\nI got that from these two sites, [here](https://patrick-llgc.github.io/Learning-Deep-Learning/paper_notes/efficientnet.html) and [here](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/), also with the resolutions of the higher models. I'm not quite clear on what the possible side effects are when you use it with a higher resolution (without the final classification layer).\n\nAn advantage of squares is that you can try 90 degree rotation augmentation.",
      "votes": null
    },
    {
      "id": "1282403",
      "postDate": "04/23/2021 22:20:09",
      "content": "<p>From what I see, Efficientnet rebuilds the architecture based on the given image resolution. Do you know if the same is possible for other pre-trained models like ResNet? My other question is what happens to the pre-trained weights when the architecture is rebuilt by the new resolution?</p>",
      "rawMarkdown": "From what I see, Efficientnet rebuilds the architecture based on the given image resolution. Do you know if the same is possible for other pre-trained models like ResNet? My other question is what happens to the pre-trained weights when the architecture is rebuilt by the new resolution?",
      "votes": null
    },
    {
      "id": "1282424",
      "postDate": "04/23/2021 23:11:36",
      "content": "<p>To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size: </p>\n<p>224x224 =&gt; 3.15 LB<br>\n320x320 =&gt; 2.43 LB<br>\n384x384 =&gt; 2.13 LB</p>",
      "rawMarkdown": "To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size: \n\n224x224 => 3.15 LB\n320x320 => 2.43 LB\n384x384 => 2.13 LB",
      "votes": null
    },
    {
      "id": "1282492",
      "postDate": "04/24/2021 02:08:27",
      "content": "<p>the above results are based on transformer or attention-lstm?</p>",
      "rawMarkdown": "the above results are based on transformer or attention-lstm?",
      "votes": null
    },
    {
      "id": "1283063",
      "postDate": "04/24/2021 14:45:18",
      "content": "<p>I am ignoring right now aspect ratio and resize to a quadratic image size. So no padding involved. Weeks ago I have tried to respect the aspect ratio, but the result was worse. Intuitively, if you keep the aspect ratio and pad, your effective resolution will be smaller. Here it seems really bigger resolution -&gt; better score.</p>",
      "rawMarkdown": "I am ignoring right now aspect ratio and resize to a quadratic image size. So no padding involved. Weeks ago I have tried to respect the aspect ratio, but the result was worse. Intuitively, if you keep the aspect ratio and pad, your effective resolution will be smaller. Here it seems really bigger resolution -> better score.",
      "votes": null
    },
    {
      "id": "1283099",
      "postDate": "04/24/2021 15:24:01",
      "content": "<p>Even if original aspect ratio is not preserved, image shape other than quadratic probably is beneficial.</p>\n<p>If you have an image that is twice as wide as tall, and you reshape it to a quadratic image, then you loose more information (~×2 as much?) horizontally. But if you reshape it to an image with the same area as the quadratic one but that is a bit wider and less tall, then you preserve more information in the end with the same amount of computation.</p>",
      "rawMarkdown": "Even if original aspect ratio is not preserved, image shape other than quadratic probably is beneficial.\n\nIf you have an image that is twice as wide as tall, and you reshape it to a quadratic image, then you loose more information (~×2 as much?) horizontally. But if you reshape it to an image with the same area as the quadratic one but that is a bit wider and less tall, then you preserve more information in the end with the same amount of computation.",
      "votes": null
    },
    {
      "id": "1283111",
      "postDate": "04/24/2021 15:36:47",
      "content": "<p>If you want rotatation as data augmentation, then you will have also to pad.</p>",
      "rawMarkdown": "If you want rotatation as data augmentation, then you will have also to pad.",
      "votes": null
    },
    {
      "id": "1283197",
      "postDate": "04/24/2021 16:57:41",
      "content": "<p>I think it uses the same pretrained weights independent of the resolution. Also, the number of parameters in the encoder is independent of the resolution (I just checked). The parameters of a CNN are mostly in the convolutional kernels, which can be applied to different input sizes without changing anything. Just the output dimension scales with the resolution, e.g. EffNet B0: 224x224 produces 7x7 features, 512x256 produces 16x8 features. </p>",
      "rawMarkdown": "I think it uses the same pretrained weights independent of the resolution. Also, the number of parameters in the encoder is independent of the resolution (I just checked). The parameters of a CNN are mostly in the convolutional kernels, which can be applied to different input sizes without changing anything. Just the output dimension scales with the resolution, e.g. EffNet B0: 224x224 produces 7x7 features, 512x256 produces 16x8 features.",
      "votes": null
    },
    {
      "id": "1283261",
      "postDate": "04/24/2021 18:39:13",
      "content": "<p>Thanks, I got it now! It scans the image with the same weights, but the only change is the number of operations based on the given size.</p>",
      "rawMarkdown": "Thanks, I got it now! It scans the image with the same weights, but the only change is the number of operations based on the given size.",
      "votes": null
    },
    {
      "id": "1283278",
      "postDate": "04/24/2021 18:59:53",
      "content": "<p>Sorry, that I don't understand. Why do you need to pad to rotate? If you resize to quadratic shape, like you described, you could just rotate then, no?</p>\n<p>I wouldn't have expected that distorting the non quadratic images to a quadratic shape would work out better in the end. For example angles are not conserved in this transformation. Thanks for this hint, very interesting…</p>",
      "rawMarkdown": "Sorry, that I don't understand. Why do you need to pad to rotate? If you resize to quadratic shape, like you described, you could just rotate then, no?\n\nI wouldn't have expected that distorting the non quadratic images to a quadratic shape would work out better in the end. For example angles are not conserved in this transformation. Thanks for this hint, very interesting...",
      "votes": null
    },
    {
      "id": "1283349",
      "postDate": "04/24/2021 21:07:38",
      "content": "<p>Transformer</p>",
      "rawMarkdown": "Transformer",
      "votes": null
    },
    {
      "id": "1283467",
      "postDate": "04/25/2021 00:44:12",
      "content": "<p>the giant vision transformer from timm models is:</p>\n<p>vit_huge_patch14_224_in21k<br>\n                            patch_size=14, embed_dim=1280, depth=32, num_heads=16, representation_size=1280</p>\n<p>i wonder if anyone has the resource to try this </p>",
      "rawMarkdown": "the giant vision transformer from timm models is:\n\n vit_huge_patch14_224_in21k\n                            patch_size=14, embed_dim=1280, depth=32, num_heads=16, representation_size=1280\n\ni wonder if anyone has the resource to try this ~~(effective input size is 1280x1280)~~",
      "votes": null
    },
    {
      "id": "1283471",
      "postDate": "04/25/2021 00:47:10",
      "content": "<p>if you analyze the error by size, you will find that for the longest seq (largest image), results improve as input size increase</p>\n<p>for the shortest seq (smallest images), the results peak at the best size of about 320. it decreases  after this</p>\n<p>there is no best size that fits all. i am talking about just resizing the input image without considering the aspect ratio.</p>",
      "rawMarkdown": "if you analyze the error by size, you will find that for the longest seq (largest image), results improve as input size increase\n\nfor the shortest seq (smallest images), the results peak at the best size of about 320. it decreases  after this\n\nthere is no best size that fits all. i am talking about just resizing the input image without considering the aspect ratio.",
      "votes": null
    },
    {
      "id": "1283478",
      "postDate": "04/25/2021 01:00:04",
      "content": "<p>How did you make your experiments in order to get these results?</p>",
      "rawMarkdown": "How did you make your experiments in order to get these results?",
      "votes": null
    },
    {
      "id": "1283480",
      "postDate": "04/25/2021 01:03:52",
      "content": "<p>\"Weeks ago I have tried to respect the aspect ratio, but the result was worse.\"</p>\n<p>i updated some results with preserved const aspect ratio when resizing<br>\n(see excel table in the post)</p>\n<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>CV definitely worsens. But LB could be similar or better </p>",
      "rawMarkdown": "\"Weeks ago I have tried to respect the aspect ratio, but the result was worse.\"\n\n\ni updated some results with preserved const aspect ratio when resizing\n(see excel table in the post)\n\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n\nCV definitely worsens. But LB could be similar or better",
      "votes": null
    },
    {
      "id": "1283499",
      "postDate": "04/25/2021 01:40:29",
      "content": "<p>I thought <code>_224_</code> means that your input size is 224x224.</p>",
      "rawMarkdown": "I thought `_224_` means that your input size is 224x224.",
      "votes": null
    },
    {
      "id": "1283505",
      "postDate": "04/25/2021 01:49:27",
      "content": "<p>you can correct. i misread the meaning of representative size.</p>\n<p>in the paper \"For ImageNet results in Table 2, we fine-tuned at higher resolution:<br>\n512 for ViT-L/16 and 518 for ViT-H/14,\" i suppose higher resolution models are not released.</p>",
      "rawMarkdown": "you can correct. i misread the meaning of representative size.\n\nin the paper \"For ImageNet results in Table 2, we fine-tuned at higher resolution:\n512 for ViT-L/16 and 518 for ViT-H/14,\" i suppose higher resolution models are not released.",
      "votes": null
    },
    {
      "id": "1288650",
      "postDate": "04/30/2021 07:51:03",
      "content": "<p>Update: increasing even more makes performance worse. </p>\n<p>512x512 =&gt; 2.36 LB <br>\n(all other configuration the same as above experiments)</p>",
      "rawMarkdown": "Update: increasing even more makes performance worse. \n\n512x512 => 2.36 LB \n(all other configuration the same as above experiments)",
      "votes": null
    },
    {
      "id": "1289312",
      "postDate": "04/30/2021 20:45:59",
      "content": "<p>breakdown your LB score by seq length. each length has its optimum score for different image size,</p>",
      "rawMarkdown": "breakdown your LB score by seq length. each length has its optimum score for different image size,",
      "votes": null
    },
    {
      "id": "1304427",
      "postDate": "05/12/2021 15:58:33",
      "content": "<p>If you are still thinking about the aspect ratio: with ignoring aspect ratio and using a quadratic image size, single model is able to get LB0.75.</p>",
      "rawMarkdown": "If you are still thinking about the aspect ratio: with ignoring aspect ratio and using a quadratic image size, single model is able to get LB0.75.",
      "votes": null
    },
    {
      "id": "1304833",
      "postDate": "05/12/2021 23:30:05",
      "content": "<p>Single model with beam search + normalization ?</p>",
      "rawMarkdown": "Single model with beam search + normalization ?",
      "votes": null
    },
    {
      "id": "1304861",
      "postDate": "05/13/2021 00:21:17",
      "content": "<p>this is something that puzzles me.<br>\ni would think that in theory, keeping the aspect ratio would be better. But in experiments, just resizing it to square gives the best results.</p>\n<p>if you keep aspect ratio:</p>\n<ul>\n<li>the drawing of the function group is consistent. bond length, etc is consistent. The model show learn the graph structure better?</li>\n</ul>\n<p>if you ignore aspect ratio:</p>\n<ul>\n<li>some drawings will be enlarged, some will be down-sized. Just by looking at part of the image, you know what is the \"molecule size\".</li>\n<li>for cases where the image is enlarged, it may be better for convolution CNN (because of striding in conv kernel). But the transformer doesn't use striding, so i am not sure where a bigger image helps the transformer? enlarge image +patch embed = original image + smaller patch embed?</li>\n</ul>",
      "rawMarkdown": "this is something that puzzles me.\ni would think that in theory, keeping the aspect ratio would be better. But in experiments, just resizing it to square gives the best results.\n\nif you keep aspect ratio:\n- the drawing of the function group is consistent. bond length, etc is consistent. The model show learn the graph structure better?\n\nif you ignore aspect ratio:\n- some drawings will be enlarged, some will be down-sized. Just by looking at part of the image, you know what is the \"molecule size\".\n- for cases where the image is enlarged, it may be better for convolution CNN (because of striding in conv kernel). But the transformer doesn't use striding, so i am not sure where a bigger image helps the transformer? enlarge image +patch embed = original image + smaller patch embed?",
      "votes": null
    },
    {
      "id": "1304863",
      "postDate": "05/13/2021 00:23:19",
      "content": "<p>on a side note, i note that there is only one type of rotation in train images (e.g. only +90 degree. -90 is absent)</p>\n<p>Is that the same for test? I look through several test images, it seems that  it follows the same rule as train (but there are too many test images … and i didn't check all)</p>",
      "rawMarkdown": "on a side note, i note that there is only one type of rotation in train images (e.g. only +90 degree. -90 is absent)\n\nIs that the same for test? I look through several test images, it seems that  it follows the same rule as train (but there are too many test images ... and i didn't check all)",
      "votes": null
    },
    {
      "id": "1305502",
      "postDate": "05/13/2021 10:21:24",
      "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> no beam search.</p>",
      "rawMarkdown": "nofreewill no beam search.",
      "votes": null
    },
    {
      "id": "1305509",
      "postDate": "05/13/2021 10:29:36",
      "content": "<p>I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.<br>\nCould it be that quadratic is better for transformers, but not LSTMs?</p>",
      "rawMarkdown": "I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?",
      "votes": null
    },
    {
      "id": "1305555",
      "postDate": "05/13/2021 11:10:36",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> wow, that's impressive. Thanks!</p>",
      "rawMarkdown": "tugstugi wow, that's impressive. Thanks!",
      "votes": null
    },
    {
      "id": "1305573",
      "postDate": "05/13/2021 11:22:16",
      "content": "<pre><code>I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?\n</code></pre>\n<p>I run some small experiments early on with <code>LSTM</code> and<code>resnet34</code><br>\n<code>CV: 6.01</code> with preserving aspect ratio<br>\n<code>CV:604</code> with not preserving aspect ratio</p>\n<p>I observed only small difference but as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  pointed out for me it felt more logical to preserve aspect ratio.</p>\n<p>my best <code>LSTM</code> has cv of <code>1.30</code> and LB of <code>1.30</code> using image resizing with preserving aspect ratio.  and transformer based model with preserving aspect ratio has <code>CV</code> of <code>0.88</code>, and LB of <code>0.87</code>. I believe Its possible for my model to reach CV of <code>0.78-0.80</code> with slightly modified architecture and preserving aspect ratio.</p>\n<p>I know <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> will ask these questions, so I am answering in advance, <code>no beam search</code>, <code>single_model</code> with <code>normalization</code></p>",
      "rawMarkdown": "```\nI tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?\n```\n\nI run some small experiments early on with `LSTM` and` resnet34 `\n`CV: 6.01` with preserving aspect ratio\n`CV:604` with not preserving aspect ratio\n\nI observed only small difference but as @hengck23  pointed out for me it felt more logical to preserve aspect ratio.\n\nmy best `LSTM` has cv of `1.30` and LB of `1.30` using image resizing with preserving aspect ratio.  and transformer based model with preserving aspect ratio has `CV` of `0.88`, and LB of `0.87`. I believe Its possible for my model to reach CV of `0.78-0.80` with slightly modified architecture and preserving aspect ratio.\n\nI know @nofreewill will ask these questions, so I am answering in advance, `no beam search`, `single_model` with `normalization`",
      "votes": null
    },
    {
      "id": "1305655",
      "postDate": "05/13/2021 12:10:06",
      "content": "<p>Thanks for the answer! :D</p>",
      "rawMarkdown": "Thanks for the answer! :D",
      "votes": null
    },
    {
      "id": "1306508",
      "postDate": "05/13/2021 20:28:45",
      "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Have you ever compared:</p>\n<ol>\n<li>Images of the same effective image size? Eg. 1:2 vs 1.41:1.41 - Same number of pixels.</li>\n<li>Images where the height remains the same, but the width changes. E.g. 256:512 vs256:256</li>\n</ol>\n<p>I think squared works better because your image is large in the height direction which is better for all the not so 'short in height' images. And increasing 384<em>384 to 384</em>768 won't help because it is already large enough in the width direction for most cases. And the few exceptions don't make a difference considering that some images might be stretched too much.<br>\nBUT I don't think squared is the key - it's more about finding the best expected height &amp; width (for your model) given the train images.</p>",
      "rawMarkdown": "drhabib Have you ever compared:\n1. Images of the same effective image size? Eg. 1:2 vs 1.41:1.41 - Same number of pixels.\n2. Images where the height remains the same, but the width changes. E.g. 256:512 vs256:256\n\nI think squared works better because your image is large in the height direction which is better for all the not so 'short in height' images. And increasing 384*384 to 384*768 won't help because it is already large enough in the width direction for most cases. And the few exceptions don't make a difference considering that some images might be stretched too much.\nBUT I don't think squared is the key - it's more about finding the best expected height & width (for your model) given the train images.",
      "votes": null
    },
    {
      "id": "1309721",
      "postDate": "05/16/2021 08:27:48",
      "content": "<p><code>To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size:</code></p>\n<p>Can you pls tell how you came up with those numbers then ? Did you trained on subset of data or is it something else ?</p>",
      "rawMarkdown": "`To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size:`\n\nCan you pls tell how you came up with those numbers then ? Did you trained on subset of data or is it something else ?",
      "votes": null
    },
    {
      "id": "1309733",
      "postDate": "05/16/2021 08:41:26",
      "content": "<p>He didn't experiment with different aspect ratios (height/width).<br>\nHe experimented with different image sizes that all had aspect ratio of 1.</p>",
      "rawMarkdown": "He didn't experiment with different aspect ratios (height/width).\nHe experimented with different image sizes that all had aspect ratio of 1.",
      "votes": null
    },
    {
      "id": "1309736",
      "postDate": "05/16/2021 08:45:24",
      "content": "<p>Oh I see, didn't read the comment carefully. Thanks for the clarification.</p>",
      "rawMarkdown": "Oh I see, didn't read the comment carefully. Thanks for the clarification.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1281243,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "04/22/2021 18:32:01",
      "content": "<p>In my first experiments, I was using the 1.72 ratio as stated in <a href=\"https://www.kaggle.com/markwijkhuizen/advanced-image-cleaning-and-tfrecord-generation\" target=\"_blank\"> Mark's TF records notebook</a>. Then I noticed that this ratio is calculated before the cropping. Then I started to use the cropped ratio (2.17) based on the <a href=\"https://www.kaggle.com/michaelwolff/bms-inchi-cropped-img-sizes-for-best-resolution\" target=\"_blank\">resolution analysis notebook</a>.</p>\n<p>I am not sure about the effect of optimizing the ratio on the model performance, but it decreases the training time for sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1281257,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "04/22/2021 18:44:16",
          "content": "<p>Awesome. Thanks for the info. I was a bit uncertain if rectangular input images would be ok to use… but from my research, and the feedback from everyone on Kaggle, it seems that square is just a choice dictated by simplicity and that I can change over my network to using rectangular inputs.</p>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1282373,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "04/23/2021 21:15:38",
          "content": "<p>Hi, <br>\nI also switched from Mark's ratio of 1.75 to 2.0 (512x256) and got a slight improvement on the LB of 0.4. That was with the 'old' attention model (version 3) in his notebook with LB scores of around 14.</p>\n<p>I also read that the input res should be tuned to the encoder model you want to use. For example EfficientNet B0 has a 'natural' resolution of 224. I.e. that is the resolution the network would expect if you used the whole pretrained model for image classification.</p>\n<p>I got that from these two sites, <a href=\"https://patrick-llgc.github.io/Learning-Deep-Learning/paper_notes/efficientnet.html\" target=\"_blank\">here</a> and <a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">here</a>, also with the resolutions of the higher models. I'm not quite clear on what the possible side effects are when you use it with a higher resolution (without the final classification layer).</p>\n<p>An advantage of squares is that you can try 90 degree rotation augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1282403,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "04/23/2021 22:20:09",
          "content": "<p>From what I see, Efficientnet rebuilds the architecture based on the given image resolution. Do you know if the same is possible for other pre-trained models like ResNet? My other question is what happens to the pre-trained weights when the architecture is rebuilt by the new resolution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283197,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "04/24/2021 16:57:41",
          "content": "<p>I think it uses the same pretrained weights independent of the resolution. Also, the number of parameters in the encoder is independent of the resolution (I just checked). The parameters of a CNN are mostly in the convolutional kernels, which can be applied to different input sizes without changing anything. Just the output dimension scales with the resolution, e.g. EffNet B0: 224x224 produces 7x7 features, 512x256 produces 16x8 features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283261,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "04/24/2021 18:39:13",
          "content": "<p>Thanks, I got it now! It scans the image with the same weights, but the only change is the number of operations based on the given size.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1282424,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "04/23/2021 23:11:36",
      "content": "<p>To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size: </p>\n<p>224x224 =&gt; 3.15 LB<br>\n320x320 =&gt; 2.43 LB<br>\n384x384 =&gt; 2.13 LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 1282492,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/24/2021 02:08:27",
          "content": "<p>the above results are based on transformer or attention-lstm?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1283349,
              "author_name": "talktocharles",
              "author_url": "",
              "post_date": "04/24/2021 21:07:38",
              "content": "<p>Transformer</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1283478,
          "author_name": "sorkun",
          "author_url": "",
          "post_date": "04/25/2021 01:00:04",
          "content": "<p>How did you make your experiments in order to get these results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1288650,
          "author_name": "talktocharles",
          "author_url": "",
          "post_date": "04/30/2021 07:51:03",
          "content": "<p>Update: increasing even more makes performance worse. </p>\n<p>512x512 =&gt; 2.36 LB <br>\n(all other configuration the same as above experiments)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1289312,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/30/2021 20:45:59",
          "content": "<p>breakdown your LB score by seq length. each length has its optimum score for different image size,</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1309721,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "05/16/2021 08:27:48",
          "content": "<p><code>To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size:</code></p>\n<p>Can you pls tell how you came up with those numbers then ? Did you trained on subset of data or is it something else ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1309733,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/16/2021 08:41:26",
          "content": "<p>He didn't experiment with different aspect ratios (height/width).<br>\nHe experimented with different image sizes that all had aspect ratio of 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1309736,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "05/16/2021 08:45:24",
          "content": "<p>Oh I see, didn't read the comment carefully. Thanks for the clarification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283063,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "04/24/2021 14:45:18",
      "content": "<p>I am ignoring right now aspect ratio and resize to a quadratic image size. So no padding involved. Weeks ago I have tried to respect the aspect ratio, but the result was worse. Intuitively, if you keep the aspect ratio and pad, your effective resolution will be smaller. Here it seems really bigger resolution -&gt; better score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1283099,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "04/24/2021 15:24:01",
          "content": "<p>Even if original aspect ratio is not preserved, image shape other than quadratic probably is beneficial.</p>\n<p>If you have an image that is twice as wide as tall, and you reshape it to a quadratic image, then you loose more information (~×2 as much?) horizontally. But if you reshape it to an image with the same area as the quadratic one but that is a bit wider and less tall, then you preserve more information in the end with the same amount of computation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283111,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/24/2021 15:36:47",
          "content": "<p>If you want rotatation as data augmentation, then you will have also to pad.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283278,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "04/24/2021 18:59:53",
          "content": "<p>Sorry, that I don't understand. Why do you need to pad to rotate? If you resize to quadratic shape, like you described, you could just rotate then, no?</p>\n<p>I wouldn't have expected that distorting the non quadratic images to a quadratic shape would work out better in the end. For example angles are not conserved in this transformation. Thanks for this hint, very interesting…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283471,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/25/2021 00:47:10",
          "content": "<p>if you analyze the error by size, you will find that for the longest seq (largest image), results improve as input size increase</p>\n<p>for the shortest seq (smallest images), the results peak at the best size of about 320. it decreases  after this</p>\n<p>there is no best size that fits all. i am talking about just resizing the input image without considering the aspect ratio.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283480,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/25/2021 01:03:52",
          "content": "<p>\"Weeks ago I have tried to respect the aspect ratio, but the result was worse.\"</p>\n<p>i updated some results with preserved const aspect ratio when resizing<br>\n(see excel table in the post)</p>\n<p><a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>CV definitely worsens. But LB could be similar or better </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283467,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/25/2021 00:44:12",
      "content": "<p>the giant vision transformer from timm models is:</p>\n<p>vit_huge_patch14_224_in21k<br>\n                            patch_size=14, embed_dim=1280, depth=32, num_heads=16, representation_size=1280</p>\n<p>i wonder if anyone has the resource to try this </p>",
      "votes": null,
      "replies": [
        {
          "id": 1283499,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/25/2021 01:40:29",
          "content": "<p>I thought <code>_224_</code> means that your input size is 224x224.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283505,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/25/2021 01:49:27",
          "content": "<p>you can correct. i misread the meaning of representative size.</p>\n<p>in the paper \"For ImageNet results in Table 2, we fine-tuned at higher resolution:<br>\n512 for ViT-L/16 and 518 for ViT-H/14,\" i suppose higher resolution models are not released.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1304427,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "05/12/2021 15:58:33",
      "content": "<p>If you are still thinking about the aspect ratio: with ignoring aspect ratio and using a quadratic image size, single model is able to get LB0.75.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1304833,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/12/2021 23:30:05",
          "content": "<p>Single model with beam search + normalization ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1304861,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/13/2021 00:21:17",
          "content": "<p>this is something that puzzles me.<br>\ni would think that in theory, keeping the aspect ratio would be better. But in experiments, just resizing it to square gives the best results.</p>\n<p>if you keep aspect ratio:</p>\n<ul>\n<li>the drawing of the function group is consistent. bond length, etc is consistent. The model show learn the graph structure better?</li>\n</ul>\n<p>if you ignore aspect ratio:</p>\n<ul>\n<li>some drawings will be enlarged, some will be down-sized. Just by looking at part of the image, you know what is the \"molecule size\".</li>\n<li>for cases where the image is enlarged, it may be better for convolution CNN (because of striding in conv kernel). But the transformer doesn't use striding, so i am not sure where a bigger image helps the transformer? enlarge image +patch embed = original image + smaller patch embed?</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1304863,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/13/2021 00:23:19",
          "content": "<p>on a side note, i note that there is only one type of rotation in train images (e.g. only +90 degree. -90 is absent)</p>\n<p>Is that the same for test? I look through several test images, it seems that  it follows the same rule as train (but there are too many test images … and i didn't check all)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305502,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "05/13/2021 10:21:24",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> no beam search.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305509,
          "author_name": "michaelwolff",
          "author_url": "",
          "post_date": "05/13/2021 10:29:36",
          "content": "<p>I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.<br>\nCould it be that quadratic is better for transformers, but not LSTMs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305555,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/13/2021 11:10:36",
          "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> wow, that's impressive. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305573,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "05/13/2021 11:22:16",
          "content": "<pre><code>I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?\n</code></pre>\n<p>I run some small experiments early on with <code>LSTM</code> and<code>resnet34</code><br>\n<code>CV: 6.01</code> with preserving aspect ratio<br>\n<code>CV:604</code> with not preserving aspect ratio</p>\n<p>I observed only small difference but as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  pointed out for me it felt more logical to preserve aspect ratio.</p>\n<p>my best <code>LSTM</code> has cv of <code>1.30</code> and LB of <code>1.30</code> using image resizing with preserving aspect ratio.  and transformer based model with preserving aspect ratio has <code>CV</code> of <code>0.88</code>, and LB of <code>0.87</code>. I believe Its possible for my model to reach CV of <code>0.78-0.80</code> with slightly modified architecture and preserving aspect ratio.</p>\n<p>I know <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> will ask these questions, so I am answering in advance, <code>no beam search</code>, <code>single_model</code> with <code>normalization</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305655,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/13/2021 12:10:06",
          "content": "<p>Thanks for the answer! :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306508,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "05/13/2021 20:28:45",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Have you ever compared:</p>\n<ol>\n<li>Images of the same effective image size? Eg. 1:2 vs 1.41:1.41 - Same number of pixels.</li>\n<li>Images where the height remains the same, but the width changes. E.g. 256:512 vs256:256</li>\n</ol>\n<p>I think squared works better because your image is large in the height direction which is better for all the not so 'short in height' images. And increasing 384<em>384 to 384</em>768 won't help because it is already large enough in the width direction for most cases. And the few exceptions don't make a difference considering that some images might be stretched too much.<br>\nBUT I don't think squared is the key - it's more about finding the best expected height &amp; width (for your model) given the train images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1281104": "Hi there. As the title states, I'm curious to see what aspect ratios others are using.\n\nI know it has been (tentatively) [shown that a 2:1 aspect ratio (appropriately cropped) may be the most efficient method of capturing relevant information](https://www.kaggle.com/c/bms-molecular-translation/discussion/233834).\n\nHowever, others have suggested that the [margins might contain useable information to potentially recover ghost methyl groups](https://www.kaggle.com/c/bms-molecular-translation/discussion/233927).\n\nPersonally, I'm still just using a padded-to-square input image (**`384x384x3`**). This is probably not ideal, but I wanted to query other Kagglers to see their thought process. I plan on running a minor benchmarking comparison in the near future as well.",
    "1281243": "In my first experiments, I was using the 1.72 ratio as stated in [ Mark's TF records notebook](https://www.kaggle.com/markwijkhuizen/advanced-image-cleaning-and-tfrecord-generation). Then I noticed that this ratio is calculated before the cropping. Then I started to use the cropped ratio (2.17) based on the [resolution analysis notebook](https://www.kaggle.com/michaelwolff/bms-inchi-cropped-img-sizes-for-best-resolution).\n\nI am not sure about the effect of optimizing the ratio on the model performance, but it decreases the training time for sure.",
    "1281257": "Awesome. Thanks for the info. I was a bit uncertain if rectangular input images would be ok to use... but from my research, and the feedback from everyone on Kaggle, it seems that square is just a choice dictated by simplicity and that I can change over my network to using rectangular inputs.\n\nThanks again!",
    "1282373": "Hi, \nI also switched from Mark's ratio of 1.75 to 2.0 (512x256) and got a slight improvement on the LB of 0.4. That was with the 'old' attention model (version 3) in his notebook with LB scores of around 14.\n\nI also read that the input res should be tuned to the encoder model you want to use. For example EfficientNet B0 has a 'natural' resolution of 224. I.e. that is the resolution the network would expect if you used the whole pretrained model for image classification.\n\nI got that from these two sites, [here](https://patrick-llgc.github.io/Learning-Deep-Learning/paper_notes/efficientnet.html) and [here](https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/), also with the resolutions of the higher models. I'm not quite clear on what the possible side effects are when you use it with a higher resolution (without the final classification layer).\n\nAn advantage of squares is that you can try 90 degree rotation augmentation.",
    "1282403": "From what I see, Efficientnet rebuilds the architecture based on the given image resolution. Do you know if the same is possible for other pre-trained models like ResNet? My other question is what happens to the pre-trained weights when the architecture is rebuilt by the new resolution?",
    "1282424": "To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size: \n\n224x224 => 3.15 LB\n320x320 => 2.43 LB\n384x384 => 2.13 LB",
    "1282492": "the above results are based on transformer or attention-lstm?",
    "1283063": "I am ignoring right now aspect ratio and resize to a quadratic image size. So no padding involved. Weeks ago I have tried to respect the aspect ratio, but the result was worse. Intuitively, if you keep the aspect ratio and pad, your effective resolution will be smaller. Here it seems really bigger resolution -> better score.",
    "1283099": "Even if original aspect ratio is not preserved, image shape other than quadratic probably is beneficial.\n\nIf you have an image that is twice as wide as tall, and you reshape it to a quadratic image, then you loose more information (~×2 as much?) horizontally. But if you reshape it to an image with the same area as the quadratic one but that is a bit wider and less tall, then you preserve more information in the end with the same amount of computation.",
    "1283111": "If you want rotatation as data augmentation, then you will have also to pad.",
    "1283197": "I think it uses the same pretrained weights independent of the resolution. Also, the number of parameters in the encoder is independent of the resolution (I just checked). The parameters of a CNN are mostly in the convolutional kernels, which can be applied to different input sizes without changing anything. Just the output dimension scales with the resolution, e.g. EffNet B0: 224x224 produces 7x7 features, 512x256 produces 16x8 features.",
    "1283261": "Thanks, I got it now! It scans the image with the same weights, but the only change is the number of operations based on the given size.",
    "1283278": "Sorry, that I don't understand. Why do you need to pad to rotate? If you resize to quadratic shape, like you described, you could just rotate then, no?\n\nI wouldn't have expected that distorting the non quadratic images to a quadratic shape would work out better in the end. For example angles are not conserved in this transformation. Thanks for this hint, very interesting...",
    "1283349": "Transformer",
    "1283467": "the giant vision transformer from timm models is:\n\n vit_huge_patch14_224_in21k\n                            patch_size=14, embed_dim=1280, depth=32, num_heads=16, representation_size=1280\n\ni wonder if anyone has the resource to try this ~~(effective input size is 1280x1280)~~",
    "1283471": "if you analyze the error by size, you will find that for the longest seq (largest image), results improve as input size increase\n\nfor the shortest seq (smallest images), the results peak at the best size of about 320. it decreases  after this\n\nthere is no best size that fits all. i am talking about just resizing the input image without considering the aspect ratio.",
    "1283478": "How did you make your experiments in order to get these results?",
    "1283480": "\"Weeks ago I have tried to respect the aspect ratio, but the result was worse.\"\n\n\ni updated some results with preserved const aspect ratio when resizing\n(see excel table in the post)\n\nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n\nCV definitely worsens. But LB could be similar or better",
    "1283499": "I thought `_224_` means that your input size is 224x224.",
    "1283505": "you can correct. i misread the meaning of representative size.\n\nin the paper \"For ImageNet results in Table 2, we fine-tuned at higher resolution:\n512 for ViT-L/16 and 518 for ViT-H/14,\" i suppose higher resolution models are not released.",
    "1288650": "Update: increasing even more makes performance worse. \n\n512x512 => 2.36 LB \n(all other configuration the same as above experiments)",
    "1289312": "breakdown your LB score by seq length. each length has its optimum score for different image size,",
    "1304427": "If you are still thinking about the aspect ratio: with ignoring aspect ratio and using a quadratic image size, single model is able to get LB0.75.",
    "1304833": "Single model with beam search + normalization ?",
    "1304861": "this is something that puzzles me.\ni would think that in theory, keeping the aspect ratio would be better. But in experiments, just resizing it to square gives the best results.\n\nif you keep aspect ratio:\n- the drawing of the function group is consistent. bond length, etc is consistent. The model show learn the graph structure better?\n\nif you ignore aspect ratio:\n- some drawings will be enlarged, some will be down-sized. Just by looking at part of the image, you know what is the \"molecule size\".\n- for cases where the image is enlarged, it may be better for convolution CNN (because of striding in conv kernel). But the transformer doesn't use striding, so i am not sure where a bigger image helps the transformer? enlarge image +patch embed = original image + smaller patch embed?",
    "1304863": "on a side note, i note that there is only one type of rotation in train images (e.g. only +90 degree. -90 is absent)\n\nIs that the same for test? I look through several test images, it seems that  it follows the same rule as train (but there are too many test images ... and i didn't check all)",
    "1305502": "nofreewill no beam search.",
    "1305509": "I tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?",
    "1305555": "tugstugi wow, that's impressive. Thanks!",
    "1305573": "```\nI tested quadratic vs ratio 2 image resolutions with the same number of pixels with LSTMs and Efficientnet B0 encoder. There ratio 2 seems to work best, then quadratic with original aspect ratio and worst, quadratic with resize ignoring aspect ratio.\nCould it be that quadratic is better for transformers, but not LSTMs?\n```\n\nI run some small experiments early on with `LSTM` and` resnet34 `\n`CV: 6.01` with preserving aspect ratio\n`CV:604` with not preserving aspect ratio\n\nI observed only small difference but as @hengck23  pointed out for me it felt more logical to preserve aspect ratio.\n\nmy best `LSTM` has cv of `1.30` and LB of `1.30` using image resizing with preserving aspect ratio.  and transformer based model with preserving aspect ratio has `CV` of `0.88`, and LB of `0.87`. I believe Its possible for my model to reach CV of `0.78-0.80` with slightly modified architecture and preserving aspect ratio.\n\nI know @nofreewill will ask these questions, so I am answering in advance, `no beam search`, `single_model` with `normalization`",
    "1305655": "Thanks for the answer! :D",
    "1306508": "drhabib Have you ever compared:\n1. Images of the same effective image size? Eg. 1:2 vs 1.41:1.41 - Same number of pixels.\n2. Images where the height remains the same, but the width changes. E.g. 256:512 vs256:256\n\nI think squared works better because your image is large in the height direction which is better for all the not so 'short in height' images. And increasing 384*384 to 384*768 won't help because it is already large enough in the width direction for most cases. And the few exceptions don't make a difference considering that some images might be stretched too much.\nBUT I don't think squared is the key - it's more about finding the best expected height & width (for your model) given the train images.",
    "1309721": "`To be honest, I have not had the time or compute available to experiment with different aspect ratios. However I have found a benefit from increasing image size:`\n\nCan you pls tell how you came up with those numbers then ? Did you trained on subset of data or is it something else ?",
    "1309733": "He didn't experiment with different aspect ratios (height/width).\nHe experimented with different image sizes that all had aspect ratio of 1.",
    "1309736": "Oh I see, didn't read the comment carefully. Thanks for the clarification."
  },
  "source": "meta"
}