{
  "id": 311760,
  "title": "Object Detection: How to handle image of different sizes & aspect ratios?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/311760",
  "author_name": "",
  "post_date": "2022-03-08T17:09:15.702725300Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am trying to train a <a href=\"https://github.com/amdegroot/ssd.pytorch\" target=\"_blank\">pytorch SSD</a> using the crowd sourced annotations, but I found that the training images have different image sizes and aspect ratios.</p>\n<ol>\n<li>What is the best practice to batch those images of different sizes?</li>\n<li>Is there any hyper-parameters that should be calibrated to handle the difference in aspect ratio?</li>\n<li>if someone has experience in the repo, do you know what <code>min_dim</code> (in config) mean? I didn't find an explanation in the repo.</li>\n</ol>",
  "messages": [
    {
      "id": "1716133",
      "postDate": "03/08/2022 17:09:15",
      "content": "<p>I am trying to train a <a href=\"https://github.com/amdegroot/ssd.pytorch\" target=\"_blank\">pytorch SSD</a> using the crowd sourced annotations, but I found that the training images have different image sizes and aspect ratios.</p>\n<ol>\n<li>What is the best practice to batch those images of different sizes?</li>\n<li>Is there any hyper-parameters that should be calibrated to handle the difference in aspect ratio?</li>\n<li>if someone has experience in the repo, do you know what <code>min_dim</code> (in config) mean? I didn't find an explanation in the repo.</li>\n</ol>",
      "rawMarkdown": "I am trying to train a [pytorch SSD](https://github.com/amdegroot/ssd.pytorch) using the crowd sourced annotations, but I found that the training images have different image sizes and aspect ratios.\n1. What is the best practice to batch those images of different sizes?\n2. Is there any hyper-parameters that should be calibrated to handle the difference in aspect ratio?\n3. if someone has experience in the repo, do you know what `min_dim` (in config) mean? I didn't find an explanation in the repo.",
      "votes": null
    },
    {
      "id": "1716876",
      "postDate": "03/09/2022 13:33:20",
      "content": "<p>I don't know about how to use this SSD repo, but use Yolo, it is much easier :)</p>",
      "rawMarkdown": "I don't know about how to use this SSD repo, but use Yolo, it is much easier :)",
      "votes": null
    },
    {
      "id": "1716897",
      "postDate": "03/09/2022 13:53:01",
      "content": "<p>Thanks for the suggestion! but I wanna use this repo for a reason.<br>\ncoz I wanna try <a href=\"https://github.com/NVlabs/AL-MDN\" target=\"_blank\">an active learning repo</a> on annotated data, and its built upon this SSD repo</p>",
      "rawMarkdown": "Thanks for the suggestion! but I wanna use this repo for a reason.\ncoz I wanna try [an active learning repo](https://github.com/NVlabs/AL-MDN) on annotated data, and its built upon this SSD repo",
      "votes": null
    },
    {
      "id": "1716990",
      "postDate": "03/09/2022 15:03:09",
      "content": "<p>I have got some insights from <a href=\"https://github.com/ultralytics/yolov5/blob/master/utils/datasets.py#L467\" target=\"_blank\">Yolo v5 source code</a> on point (1):</p>\n<ol>\n<li>sort training images by aspect ratio</li>\n<li>arrange batch based on the sorted orders</li>\n<li>determine the shape of each batch by the min./ max. aspect ratio of its samples</li>\n<li>pad the borders for samples to match the target shape</li>\n</ol>\n<p>One downside of this approach is that it will disable shuffling during training, which sacrifice the batch-wise variations.</p>",
      "rawMarkdown": "I have got some insights from [Yolo v5 source code](https://github.com/ultralytics/yolov5/blob/master/utils/datasets.py#L467) on point (1):\n1. sort training images by aspect ratio\n2. arrange batch based on the sorted orders\n3. determine the shape of each batch by the min./ max. aspect ratio of its samples\n4. pad the borders for samples to match the target shape\n\nOne downside of this approach is that it will disable shuffling during training, which sacrifice the batch-wise variations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1716876,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "03/09/2022 13:33:20",
      "content": "<p>I don't know about how to use this SSD repo, but use Yolo, it is much easier :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1716897,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "03/09/2022 13:53:01",
          "content": "<p>Thanks for the suggestion! but I wanna use this repo for a reason.<br>\ncoz I wanna try <a href=\"https://github.com/NVlabs/AL-MDN\" target=\"_blank\">an active learning repo</a> on annotated data, and its built upon this SSD repo</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1716990,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "03/09/2022 15:03:09",
      "content": "<p>I have got some insights from <a href=\"https://github.com/ultralytics/yolov5/blob/master/utils/datasets.py#L467\" target=\"_blank\">Yolo v5 source code</a> on point (1):</p>\n<ol>\n<li>sort training images by aspect ratio</li>\n<li>arrange batch based on the sorted orders</li>\n<li>determine the shape of each batch by the min./ max. aspect ratio of its samples</li>\n<li>pad the borders for samples to match the target shape</li>\n</ol>\n<p>One downside of this approach is that it will disable shuffling during training, which sacrifice the batch-wise variations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1716133": "I am trying to train a [pytorch SSD](https://github.com/amdegroot/ssd.pytorch) using the crowd sourced annotations, but I found that the training images have different image sizes and aspect ratios.\n1. What is the best practice to batch those images of different sizes?\n2. Is there any hyper-parameters that should be calibrated to handle the difference in aspect ratio?\n3. if someone has experience in the repo, do you know what `min_dim` (in config) mean? I didn't find an explanation in the repo.",
    "1716876": "I don't know about how to use this SSD repo, but use Yolo, it is much easier :)",
    "1716897": "Thanks for the suggestion! but I wanna use this repo for a reason.\ncoz I wanna try [an active learning repo](https://github.com/NVlabs/AL-MDN) on annotated data, and its built upon this SSD repo",
    "1716990": "I have got some insights from [Yolo v5 source code](https://github.com/ultralytics/yolov5/blob/master/utils/datasets.py#L467) on point (1):\n1. sort training images by aspect ratio\n2. arrange batch based on the sorted orders\n3. determine the shape of each batch by the min./ max. aspect ratio of its samples\n4. pad the borders for samples to match the target shape\n\nOne downside of this approach is that it will disable shuffling during training, which sacrifice the batch-wise variations."
  },
  "source": "meta"
}