{
  "id": 122731,
  "title": "Images cropped to character & resized to 100x100 (Dataset)",
  "url": "/competitions/bengaliai-cv19/discussion/122731",
  "author_name": "Maxime Lenormand",
  "post_date": "2019-12-22T14:12:21.377000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>I saw multiple people working on changing the dataset to a faster .feather type in order to shorten load time.</p>\n\n<p>As I'm working on low end hardware (i5 with mx150 gpu &amp; 8Gb of RAM on a laptop) trying to reduce the dataset size is something I try to put some effort in so this was an interesting start.</p>\n\n<p>So I've created a dataset of 100x100 images, but rather than simply reducing the image size, I cropped them to only keep the character rather than only the white space around them too.\nThis is a bit more tricky than simply using OpenCV2 and create a bounding box, because some characters contain different independent symbols. So this requires making a exterior bounding box to capture the entire character.</p>\n\n<p>Here is a quick visual of an original image and a cropped &amp; resized one:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F8c5c94d68becaed426b4325a52ec3359%2FCropped_bengaliAI.png?generation=1577023738570962&amp;alt=media\" alt=\"\"></p>\n\n<p>For more details, I made a <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">notebook</a>. Feel free to use it if you want!</p>\n\n<p>And the dataset can be found <a href=\"https://www.kaggle.com/maxlenormand/cropped-resized-bengaliai-images#train_data_3.feather\">here</a>.</p>",
  "messages": [
    {
      "id": 700734,
      "postDate": "2019-12-22T14:12:21.377Z",
      "content": "<p>Hi!</p>\n\n<p>I saw multiple people working on changing the dataset to a faster .feather type in order to shorten load time.</p>\n\n<p>As I'm working on low end hardware (i5 with mx150 gpu &amp; 8Gb of RAM on a laptop) trying to reduce the dataset size is something I try to put some effort in so this was an interesting start.</p>\n\n<p>So I've created a dataset of 100x100 images, but rather than simply reducing the image size, I cropped them to only keep the character rather than only the white space around them too.\nThis is a bit more tricky than simply using OpenCV2 and create a bounding box, because some characters contain different independent symbols. So this requires making a exterior bounding box to capture the entire character.</p>\n\n<p>Here is a quick visual of an original image and a cropped &amp; resized one:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F8c5c94d68becaed426b4325a52ec3359%2FCropped_bengaliAI.png?generation=1577023738570962&amp;alt=media\" alt=\"\"></p>\n\n<p>For more details, I made a <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">notebook</a>. Feel free to use it if you want!</p>\n\n<p>And the dataset can be found <a href=\"https://www.kaggle.com/maxlenormand/cropped-resized-bengaliai-images#train_data_3.feather\">here</a>.</p>",
      "rawMarkdown": "Hi!\n\nI saw multiple people working on changing the dataset to a faster .feather type in order to shorten load time.\n\nAs I'm working on low end hardware (i5 with mx150 gpu &amp; 8Gb of RAM on a laptop) trying to reduce the dataset size is something I try to put some effort in so this was an interesting start.\n\nSo I've created a dataset of 100x100 images, but rather than simply reducing the image size, I cropped them to only keep the character rather than only the white space around them too.\nThis is a bit more tricky than simply using OpenCV2 and create a bounding box, because some characters contain different independent symbols. So this requires making a exterior bounding box to capture the entire character.\n\nHere is a quick visual of an original image and a cropped &amp; resized one:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F8c5c94d68becaed426b4325a52ec3359%2FCropped_bengaliAI.png?generation=1577023738570962&amp;alt=media)\n\nFor more details, I made a [notebook](https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images). Feel free to use it if you want!\n\nAnd the dataset can be found [here](https://www.kaggle.com/maxlenormand/cropped-resized-bengaliai-images#train_data_3.feather).",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "700734": "Hi!\n\nI saw multiple people working on changing the dataset to a faster .feather type in order to shorten load time.\n\nAs I'm working on low end hardware (i5 with mx150 gpu &amp; 8Gb of RAM on a laptop) trying to reduce the dataset size is something I try to put some effort in so this was an interesting start.\n\nSo I've created a dataset of 100x100 images, but rather than simply reducing the image size, I cropped them to only keep the character rather than only the white space around them too.\nThis is a bit more tricky than simply using OpenCV2 and create a bounding box, because some characters contain different independent symbols. So this requires making a exterior bounding box to capture the entire character.\n\nHere is a quick visual of an original image and a cropped &amp; resized one:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F8c5c94d68becaed426b4325a52ec3359%2FCropped_bengaliAI.png?generation=1577023738570962&amp;alt=media)\n\nFor more details, I made a [notebook](https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images). Feel free to use it if you want!\n\nAnd the dataset can be found [here](https://www.kaggle.com/maxlenormand/cropped-resized-bengaliai-images#train_data_3.feather)."
  }
}