{
  "id": 134651,
  "title": "Why cropping black pixels does not work",
  "url": "/competitions/bengaliai-cv19/discussion/134651",
  "author_name": "Quan",
  "post_date": "2020-03-09T14:23:34.275000",
  "votes": 23,
  "comment_count": 19,
  "views": 0,
  "content": "<p>So after being confused on why training on cropped images sucks, I finally come up with a good explaination.\nThe code I used for cropping image\n<code>\nshapes = {}\ndef crop_image(image):\n    mid = (image.max() - image.min()) / 4\n    mask = image &amp;gt; mid\n    image = image[np.ix_(mask.any(1),mask.any(0))]\n    (h, w) = image.shape\n    shapes_key = \"{}x{}\".format(h, w)\n    if shapes_key not in shapes.keys():\n        shapes[shapes_key] = 0\n    shapes[shapes_key] += 1\n    return image\n</code>\n\"shapes\" is a dictionary where key is the image size in HxW. \nAfter running on all images, the number of different shapes (len(shapes.keys())) is over 14k! \nAnd more than quarter of them have weird aspect ratio, or small shapes like 31x21. \nSo after resizing these images (say 128x128), it just add noise, distortion and garbage information (some images have thin stripes and some has really thick stripes). \nSo it defeats the purpose of preprocessing (preprocessing supposes to enhance the quality of the dataset).\nI've finally found peace.</p>",
  "messages": [
    {
      "id": 767360,
      "postDate": "2020-03-09T14:23:34.277Z",
      "content": "<p>So after being confused on why training on cropped images sucks, I finally come up with a good explaination.\nThe code I used for cropping image\n<code>\nshapes = {}\ndef crop_image(image):\n    mid = (image.max() - image.min()) / 4\n    mask = image &amp;gt; mid\n    image = image[np.ix_(mask.any(1),mask.any(0))]\n    (h, w) = image.shape\n    shapes_key = \"{}x{}\".format(h, w)\n    if shapes_key not in shapes.keys():\n        shapes[shapes_key] = 0\n    shapes[shapes_key] += 1\n    return image\n</code>\n\"shapes\" is a dictionary where key is the image size in HxW. \nAfter running on all images, the number of different shapes (len(shapes.keys())) is over 14k! \nAnd more than quarter of them have weird aspect ratio, or small shapes like 31x21. \nSo after resizing these images (say 128x128), it just add noise, distortion and garbage information (some images have thin stripes and some has really thick stripes). \nSo it defeats the purpose of preprocessing (preprocessing supposes to enhance the quality of the dataset).\nI've finally found peace.</p>",
      "rawMarkdown": "So after being confused on why training on cropped images sucks, I finally come up with a good explaination.\nThe code I used for cropping image\n```\nshapes = {}\ndef crop_image(image):\n    mid = (image.max() - image.min()) / 4\n    mask = image &gt; mid\n    image = image[np.ix_(mask.any(1),mask.any(0))]\n    (h, w) = image.shape\n    shapes_key = \"{}x{}\".format(h, w)\n    if shapes_key not in shapes.keys():\n        shapes[shapes_key] = 0\n    shapes[shapes_key] += 1\n    return image\n```\n\"shapes\" is a dictionary where key is the image size in HxW. \nAfter running on all images, the number of different shapes (len(shapes.keys())) is over 14k! \nAnd more than quarter of them have weird aspect ratio, or small shapes like 31x21. \nSo after resizing these images (say 128x128), it just add noise, distortion and garbage information (some images have thin stripes and some has really thick stripes). \nSo it defeats the purpose of preprocessing (preprocessing supposes to enhance the quality of the dataset).\nI've finally found peace.",
      "votes": 23
    },
    {
      "id": 770047,
      "postDate": "2020-03-12T14:24:20.967Z",
      "content": "<p>What about padding the cropped images to preserve the aspect ratio?</p>\n<pre><code>extra_pad = 4 \nh , w = cropped_img.shape\nl = max(h , w ) + extra_pad                 # the size of image after padding\npadded_img = np.pad(cropped_img, (((l-h)//2,(l-h)//2),((l-w)//2,(l-w)//2) ), mode='constant')\n</code></pre>\n<p>And finally resizing to required size</p>",
      "rawMarkdown": "What about padding the cropped images to preserve the aspect ratio?\n\n    extra_pad = 4 \n    h , w = cropped_img.shape\n    l = max(h , w ) + extra_pad                 # the size of image after padding\n    padded_img = np.pad(cropped_img, (((l-h)//2,(l-h)//2),((l-w)//2,(l-w)//2) ), mode='constant')\n\n\nAnd finally resizing to required size",
      "votes": 3
    },
    {
      "id": 772466,
      "postDate": "2020-03-15T14:32:52.780Z",
      "content": "<p>Here is my notebook exploring the topic\n- <a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\">https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/</a></p>\n\n<p>```python</p>\n\n<h1>Source: <a href=\"https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy\">https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy</a></h1>\n\n<h1>This is the fast method that simply remove all empty rows/columns</h1>\n\n<h1>NOTE: assumes inverted</h1>\n\n<p>def crop_image(img, tol=0):\n    mask = img &gt; tol\n    return img[np.ix_(mask.any(1),mask.any(0))]</p>\n\n<h1>DOCS: <a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html</a></h1>\n\n<h1>NOTE: assumes inverted</h1>\n\n<p>def crop_center_image(img, tol=0):\n    org_shape   = img.shape\n    img_cropped = crop_image(img)\n    pad_x       = (org_shape[0] - img_cropped.shape[0])/2\n    pad_y       = (org_shape[1] - img_cropped.shape[1])/2\n    padding     = (\n        (math.floor(pad_x), math.ceil(pad_x)),\n        (math.floor(pad_y), math.ceil(pad_y))\n    )\n    img_center = np.pad(img_cropped, padding, 'constant', constant_values=0)\n    return img_center</p>\n\n<h1>Source: <a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\">https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/</a></h1>\n\n<h1>noinspection PyArgumentList</h1>\n\n<p>def transform_X(train: DataFrame, denoise=True, normalize=True, center=True, invert=True, resize=2, resize_fn=None) -&gt; np.ndarray:\n    train = (train.drop(columns='image_id', errors='ignore')\n             .values.astype('uint8')                   # unit8 for initial data processing\n             .reshape(-1, 137, 236)                    # 2D arrays for inline image processing\n    )\n    gc.collect(); sleep(1)</p>\n\n<pre><code># Invert for processing\n# Colors   |   0 = black      | 255 = white\n# invert   |   0 = background | 255 = line\n# original | 255 = background |   0 = line\ntrain = (255-train)\n\nif denoise:                     \n    # Rescale lines to maximum brightness, and set background values (less than 2x mean()) to 0\n    train = np.array([ train[i] + (255-train[i].max())              for i in range(train.shape[0]) ])        \n    train = np.array([ train[i] * (train[i] &gt;= np.mean(train[i])*2) for i in range(train.shape[0]) ])                                  \n\nif isinstance(resize, bool) and resize == True:\n    resize = 2    # Reduce image size by 2x\nif resize and resize != 1:                  \n    # NOTEBOOK: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n    # Out of the different resize functions:\n    # - np.mean(dtype=uint8) produces produces fragmented images (needs float16 to work properly - but RAM intensive)\n    # - np.median() produces the most accurate downsampling\n    # - np.max() produces an enhanced image with thicker lines (maybe slightly easier to read)\n    # - np.min() produces a  dehanced image with thiner lines (harder to read)\n    resize_fn = resize_fn or (np.max if invert else np.min)\n    cval      = 0 if invert else 255\n    train = skimage.measure.block_reduce(train, (1, resize,resize), cval=cval, func=resize_fn)\n\nif center:\n    # NOTE: crop_center_image() assumes inverted\n    train = np.array([\n        crop_center_image(train[i,:,:])\n        for i in range(train.shape[0])\n    ])\n\n# Un-invert if invert==False\nif not invert: train = (255-train)\n\nif normalize:\n    train = train.astype('float16') / 255.0   # prevent division cast: int -&gt; float64\n\ntrain = train.reshape(*train.shape, 1)        # 4D ndarray for tensorflow CNN\n\ngc.collect(); sleep(1)\nreturn train\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Here is my notebook exploring the topic\n- https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n\n```python\n# Source: https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy\n\n# This is the fast method that simply remove all empty rows/columns\n# NOTE: assumes inverted\ndef crop_image(img, tol=0):\n    mask = img &gt; tol\n    return img[np.ix_(mask.any(1),mask.any(0))]\n\n\n# DOCS: https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html\n# NOTE: assumes inverted\ndef crop_center_image(img, tol=0):\n    org_shape   = img.shape\n    img_cropped = crop_image(img)\n    pad_x       = (org_shape[0] - img_cropped.shape[0])/2\n    pad_y       = (org_shape[1] - img_cropped.shape[1])/2\n    padding     = (\n        (math.floor(pad_x), math.ceil(pad_x)),\n        (math.floor(pad_y), math.ceil(pad_y))\n    )\n    img_center = np.pad(img_cropped, padding, 'constant', constant_values=0)\n    return img_center\n\n\n# Source: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n# noinspection PyArgumentList\ndef transform_X(train: DataFrame, denoise=True, normalize=True, center=True, invert=True, resize=2, resize_fn=None) -&gt; np.ndarray:\n    train = (train.drop(columns='image_id', errors='ignore')\n             .values.astype('uint8')                   # unit8 for initial data processing\n             .reshape(-1, 137, 236)                    # 2D arrays for inline image processing\n    )\n    gc.collect(); sleep(1)\n\n    # Invert for processing\n    # Colors   |   0 = black      | 255 = white\n    # invert   |   0 = background | 255 = line\n    # original | 255 = background |   0 = line\n    train = (255-train)\n    \n    if denoise:                     \n        # Rescale lines to maximum brightness, and set background values (less than 2x mean()) to 0\n        train = np.array([ train[i] + (255-train[i].max())              for i in range(train.shape[0]) ])        \n        train = np.array([ train[i] * (train[i] &gt;= np.mean(train[i])*2) for i in range(train.shape[0]) ])                                  \n            \n    if isinstance(resize, bool) and resize == True:\n        resize = 2    # Reduce image size by 2x\n    if resize and resize != 1:                  \n        # NOTEBOOK: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n        # Out of the different resize functions:\n        # - np.mean(dtype=uint8) produces produces fragmented images (needs float16 to work properly - but RAM intensive)\n        # - np.median() produces the most accurate downsampling\n        # - np.max() produces an enhanced image with thicker lines (maybe slightly easier to read)\n        # - np.min() produces a  dehanced image with thiner lines (harder to read)\n        resize_fn = resize_fn or (np.max if invert else np.min)\n        cval      = 0 if invert else 255\n        train = skimage.measure.block_reduce(train, (1, resize,resize), cval=cval, func=resize_fn)\n            \n    if center:\n        # NOTE: crop_center_image() assumes inverted\n        train = np.array([\n            crop_center_image(train[i,:,:])\n            for i in range(train.shape[0])\n        ])\n        \n    # Un-invert if invert==False\n    if not invert: train = (255-train)\n\n    if normalize:\n        train = train.astype('float16') / 255.0   # prevent division cast: int -&gt; float64\n\n    train = train.reshape(*train.shape, 1)        # 4D ndarray for tensorflow CNN\n\n    gc.collect(); sleep(1)\n    return train\n```"
    },
    {
      "id": 769836,
      "postDate": "2020-03-12T10:17:47.303Z",
      "content": "<p>So what you did for size of images instead of cropping? just resize? or none of them. feeding original images size to model?</p>",
      "rawMarkdown": "So what you did for size of images instead of cropping? just resize? or none of them. feeding original images size to model?",
      "replies": [
        {
          "id": 770746,
          "postDate": "2020-03-13T10:51:57.643Z",
          "content": "<p>Not using crop,just resize is a good choice .Larger sizes maybe lead to higher scores</p>",
          "rawMarkdown": "Not using crop,just resize is a good choice .Larger sizes maybe lead to higher scores"
        }
      ]
    },
    {
      "id": 769638,
      "postDate": "2020-03-12T05:29:31.960Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice\n"
    },
    {
      "id": 767782,
      "postDate": "2020-03-10T04:57:46.700Z",
      "content": "<p>For me even subtracting from 255 and normalizing by max gave poor results, still don't know why. I'm worried there is some noise in the dataset that the model is learning and even the public dataset has the same noise, would be bad if its true.</p>",
      "rawMarkdown": "For me even subtracting from 255 and normalizing by max gave poor results, still don't know why. I'm worried there is some noise in the dataset that the model is learning and even the public dataset has the same noise, would be bad if its true.",
      "replies": [
        {
          "id": 767801,
          "postDate": "2020-03-10T05:23:56.893Z",
          "content": "<p>Can i ask when u normalize by max do you mean maximum of each individual picture or just plain divide 255</p>",
          "rawMarkdown": "Can i ask when u normalize by max do you mean maximum of each individual picture or just plain divide 255"
        },
        {
          "id": 767816,
          "postDate": "2020-03-10T05:52:33.817Z",
          "content": "<p>Yeah here's the pseudocode</p>\n\n<h1>uint8 image with white background</h1>\n\n<p>image = 255 - image \nimage = image/image.max() # Because some images the lines are darker and some are too light\nimage = augment(image) \nimage = float(image)/255. </p>\n\n<p>Even this gives poor performance on both CV and LB for me.</p>",
          "rawMarkdown": "Yeah here's the pseudocode\n\n# uint8 image with white background\nimage = 255 - image \nimage = image/image.max() # Because some images the lines are darker and some are too light\nimage = augment(image) \nimage = float(image)/255. \n\nEven this gives poor performance on both CV and LB for me.\n",
          "votes": 1
        },
        {
          "id": 767821,
          "postDate": "2020-03-10T06:00:54.557Z",
          "content": "<p>hmm this seems odd because you are dividing twice losing some information due to precision issues? May be the last line shouldnt /255 as the range of the image  is already 0-1.</p>",
          "rawMarkdown": "hmm this seems odd because you are dividing twice losing some information due to precision issues? May be the last line shouldnt /255 as the range of the image  is already 0-1."
        },
        {
          "id": 767914,
          "postDate": "2020-03-10T08:36:54.927Z",
          "content": "<p>Ok sorry my bad I typed the pseudo code wrong -</p>\n\n<p>image = 255 - image\nimage = unit8(float(image)/image.max()*255) # Because some images the lines are darker and some are too light\nimage = augment(image) # Primarily because many of the augmentations are in uint8 domain\nimage = float(image)/255.</p>\n\n<p>This is what I tried .. there are many people who were using the same during first few months.</p>",
          "rawMarkdown": "Ok sorry my bad I typed the pseudo code wrong -\n\nimage = 255 - image\nimage = unit8(float(image)/image.max()*255) # Because some images the lines are darker and some are too light\nimage = augment(image) # Primarily because many of the augmentations are in uint8 domain\nimage = float(image)/255.\n\nThis is what I tried .. there are many people who were using the same during first few months.",
          "votes": 1
        },
        {
          "id": 770782,
          "postDate": "2020-03-13T11:56:48.543Z",
          "content": "<p>My thought is that the diversity of max intensity between all images can somehow prevent training from overfitting (like adjusting brightness or else).</p>",
          "rawMarkdown": "My thought is that the diversity of max intensity between all images can somehow prevent training from overfitting (like adjusting brightness or else)."
        }
      ]
    },
    {
      "id": 767758,
      "postDate": "2020-03-10T04:10:47.823Z",
      "content": "<p>Nice explanation. This can also explain why augmentations like shiftscale or randomresizedcrop doesn't work much for me. Maybe we should choose those methods which don't change the character shape.</p>",
      "rawMarkdown": "Nice explanation. This can also explain why augmentations like shiftscale or randomresizedcrop doesn't work much for me. Maybe we should choose those methods which don't change the character shape."
    },
    {
      "id": 767674,
      "postDate": "2020-03-10T01:38:27.050Z",
      "content": "<p>great!</p>",
      "rawMarkdown": "great!"
    },
    {
      "id": 767593,
      "postDate": "2020-03-09T21:20:36.483Z",
      "content": "<p>Thanks for your test! It solved the confusion I had from the beginning. 👍 </p>",
      "rawMarkdown": "Thanks for your test! It solved the confusion I had from the beginning. 👍 "
    },
    {
      "id": 767473,
      "postDate": "2020-03-09T17:20:59.803Z",
      "content": "<p>or it has some data leak?)</p>",
      "rawMarkdown": "or it has some data leak?)",
      "replies": [
        {
          "id": 767693,
          "postDate": "2020-03-10T02:33:26.437Z",
          "content": "<p>I don't think there's any data leak</p>",
          "rawMarkdown": "I don't think there's any data leak"
        }
      ]
    },
    {
      "id": 767416,
      "postDate": "2020-03-09T15:49:30.753Z",
      "content": "<p>thanks for sharing <a href=\"/quandapro\">@quandapro</a> </p>",
      "rawMarkdown": "thanks for sharing @quandapro ",
      "replies": [
        {
          "id": 767694,
          "postDate": "2020-03-10T02:33:46.167Z",
          "content": "<p>You're welcome bro</p>",
          "rawMarkdown": "You're welcome bro",
          "votes": 1
        }
      ]
    },
    {
      "id": 767631,
      "postDate": "2020-03-09T23:11:37.950Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 770047,
      "author_name": "M K Prasanth",
      "author_url": "",
      "post_date": "2020-03-12T14:24:20.967000",
      "content": "<p>What about padding the cropped images to preserve the aspect ratio?</p>\n<pre><code>extra_pad = 4 \nh , w = cropped_img.shape\nl = max(h , w ) + extra_pad                 # the size of image after padding\npadded_img = np.pad(cropped_img, (((l-h)//2,(l-h)//2),((l-w)//2,(l-w)//2) ), mode='constant')\n</code></pre>\n<p>And finally resizing to required size</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 772466,
      "author_name": "James McGuigan",
      "author_url": "",
      "post_date": "2020-03-15T14:32:52.780000",
      "content": "<p>Here is my notebook exploring the topic\n- <a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\">https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/</a></p>\n\n<p>```python</p>\n\n<h1>Source: <a href=\"https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy\">https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy</a></h1>\n\n<h1>This is the fast method that simply remove all empty rows/columns</h1>\n\n<h1>NOTE: assumes inverted</h1>\n\n<p>def crop_image(img, tol=0):\n    mask = img &gt; tol\n    return img[np.ix_(mask.any(1),mask.any(0))]</p>\n\n<h1>DOCS: <a href=\"https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html\">https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html</a></h1>\n\n<h1>NOTE: assumes inverted</h1>\n\n<p>def crop_center_image(img, tol=0):\n    org_shape   = img.shape\n    img_cropped = crop_image(img)\n    pad_x       = (org_shape[0] - img_cropped.shape[0])/2\n    pad_y       = (org_shape[1] - img_cropped.shape[1])/2\n    padding     = (\n        (math.floor(pad_x), math.ceil(pad_x)),\n        (math.floor(pad_y), math.ceil(pad_y))\n    )\n    img_center = np.pad(img_cropped, padding, 'constant', constant_values=0)\n    return img_center</p>\n\n<h1>Source: <a href=\"https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\">https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/</a></h1>\n\n<h1>noinspection PyArgumentList</h1>\n\n<p>def transform_X(train: DataFrame, denoise=True, normalize=True, center=True, invert=True, resize=2, resize_fn=None) -&gt; np.ndarray:\n    train = (train.drop(columns='image_id', errors='ignore')\n             .values.astype('uint8')                   # unit8 for initial data processing\n             .reshape(-1, 137, 236)                    # 2D arrays for inline image processing\n    )\n    gc.collect(); sleep(1)</p>\n\n<pre><code># Invert for processing\n# Colors   |   0 = black      | 255 = white\n# invert   |   0 = background | 255 = line\n# original | 255 = background |   0 = line\ntrain = (255-train)\n\nif denoise:                     \n    # Rescale lines to maximum brightness, and set background values (less than 2x mean()) to 0\n    train = np.array([ train[i] + (255-train[i].max())              for i in range(train.shape[0]) ])        \n    train = np.array([ train[i] * (train[i] &gt;= np.mean(train[i])*2) for i in range(train.shape[0]) ])                                  \n\nif isinstance(resize, bool) and resize == True:\n    resize = 2    # Reduce image size by 2x\nif resize and resize != 1:                  \n    # NOTEBOOK: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n    # Out of the different resize functions:\n    # - np.mean(dtype=uint8) produces produces fragmented images (needs float16 to work properly - but RAM intensive)\n    # - np.median() produces the most accurate downsampling\n    # - np.max() produces an enhanced image with thicker lines (maybe slightly easier to read)\n    # - np.min() produces a  dehanced image with thiner lines (harder to read)\n    resize_fn = resize_fn or (np.max if invert else np.min)\n    cval      = 0 if invert else 255\n    train = skimage.measure.block_reduce(train, (1, resize,resize), cval=cval, func=resize_fn)\n\nif center:\n    # NOTE: crop_center_image() assumes inverted\n    train = np.array([\n        crop_center_image(train[i,:,:])\n        for i in range(train.shape[0])\n    ])\n\n# Un-invert if invert==False\nif not invert: train = (255-train)\n\nif normalize:\n    train = train.astype('float16') / 255.0   # prevent division cast: int -&gt; float64\n\ntrain = train.reshape(*train.shape, 1)        # 4D ndarray for tensorflow CNN\n\ngc.collect(); sleep(1)\nreturn train\n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 769836,
      "author_name": "masoud parpanchi",
      "author_url": "",
      "post_date": "2020-03-12T10:17:47.303000",
      "content": "<p>So what you did for size of images instead of cropping? just resize? or none of them. feeding original images size to model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 770746,
          "author_name": "ecnu_ybx",
          "author_url": "",
          "post_date": "2020-03-13T10:51:57.643000",
          "content": "<p>Not using crop,just resize is a good choice .Larger sizes maybe lead to higher scores</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 769638,
      "author_name": "Niket Vohra",
      "author_url": "",
      "post_date": "2020-03-12T05:29:31.960000",
      "content": "<p>nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767782,
      "author_name": "Dipam Chakraborty",
      "author_url": "",
      "post_date": "2020-03-10T04:57:46.700000",
      "content": "<p>For me even subtracting from 255 and normalizing by max gave poor results, still don't know why. I'm worried there is some noise in the dataset that the model is learning and even the public dataset has the same noise, would be bad if its true.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 767801,
          "author_name": "garyong",
          "author_url": "",
          "post_date": "2020-03-10T05:23:56.893000",
          "content": "<p>Can i ask when u normalize by max do you mean maximum of each individual picture or just plain divide 255</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767816,
          "author_name": "Dipam Chakraborty",
          "author_url": "",
          "post_date": "2020-03-10T05:52:33.817000",
          "content": "<p>Yeah here's the pseudocode</p>\n\n<h1>uint8 image with white background</h1>\n\n<p>image = 255 - image \nimage = image/image.max() # Because some images the lines are darker and some are too light\nimage = augment(image) \nimage = float(image)/255. </p>\n\n<p>Even this gives poor performance on both CV and LB for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767821,
          "author_name": "garyong",
          "author_url": "",
          "post_date": "2020-03-10T06:00:54.557000",
          "content": "<p>hmm this seems odd because you are dividing twice losing some information due to precision issues? May be the last line shouldnt /255 as the range of the image  is already 0-1.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767914,
          "author_name": "Dipam Chakraborty",
          "author_url": "",
          "post_date": "2020-03-10T08:36:54.927000",
          "content": "<p>Ok sorry my bad I typed the pseudo code wrong -</p>\n\n<p>image = 255 - image\nimage = unit8(float(image)/image.max()*255) # Because some images the lines are darker and some are too light\nimage = augment(image) # Primarily because many of the augmentations are in uint8 domain\nimage = float(image)/255.</p>\n\n<p>This is what I tried .. there are many people who were using the same during first few months.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770782,
          "author_name": "AllenChangTW",
          "author_url": "",
          "post_date": "2020-03-13T11:56:48.543000",
          "content": "<p>My thought is that the diversity of max intensity between all images can somehow prevent training from overfitting (like adjusting brightness or else).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 767758,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-03-10T04:10:47.823000",
      "content": "<p>Nice explanation. This can also explain why augmentations like shiftscale or randomresizedcrop doesn't work much for me. Maybe we should choose those methods which don't change the character shape.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767674,
      "author_name": "0chen",
      "author_url": "",
      "post_date": "2020-03-10T01:38:27.050000",
      "content": "<p>great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767593,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-09T21:20:36.483000",
      "content": "<p>Thanks for your test! It solved the confusion I had from the beginning. 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767473,
      "author_name": "Alex Shonenkov",
      "author_url": "",
      "post_date": "2020-03-09T17:20:59.803000",
      "content": "<p>or it has some data leak?)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 767693,
          "author_name": "Quan",
          "author_url": "",
          "post_date": "2020-03-10T02:33:26.437000",
          "content": "<p>I don't think there's any data leak</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 767416,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-03-09T15:49:30.753000",
      "content": "<p>thanks for sharing <a href=\"/quandapro\">@quandapro</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 767694,
          "author_name": "Quan",
          "author_url": "",
          "post_date": "2020-03-10T02:33:46.167000",
          "content": "<p>You're welcome bro</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 767631,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-09T23:11:37.950000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "767360": "So after being confused on why training on cropped images sucks, I finally come up with a good explaination.\nThe code I used for cropping image\n```\nshapes = {}\ndef crop_image(image):\n    mid = (image.max() - image.min()) / 4\n    mask = image &gt; mid\n    image = image[np.ix_(mask.any(1),mask.any(0))]\n    (h, w) = image.shape\n    shapes_key = \"{}x{}\".format(h, w)\n    if shapes_key not in shapes.keys():\n        shapes[shapes_key] = 0\n    shapes[shapes_key] += 1\n    return image\n```\n\"shapes\" is a dictionary where key is the image size in HxW. \nAfter running on all images, the number of different shapes (len(shapes.keys())) is over 14k! \nAnd more than quarter of them have weird aspect ratio, or small shapes like 31x21. \nSo after resizing these images (say 128x128), it just add noise, distortion and garbage information (some images have thin stripes and some has really thick stripes). \nSo it defeats the purpose of preprocessing (preprocessing supposes to enhance the quality of the dataset).\nI've finally found peace.",
    "770047": "What about padding the cropped images to preserve the aspect ratio?\n\n    extra_pad = 4 \n    h , w = cropped_img.shape\n    l = max(h , w ) + extra_pad                 # the size of image after padding\n    padded_img = np.pad(cropped_img, (((l-h)//2,(l-h)//2),((l-w)//2,(l-w)//2) ), mode='constant')\n\n\nAnd finally resizing to required size",
    "772466": "Here is my notebook exploring the topic\n- https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n\n```python\n# Source: https://codereview.stackexchange.com/questions/132914/crop-black-border-of-image-using-numpy\n\n# This is the fast method that simply remove all empty rows/columns\n# NOTE: assumes inverted\ndef crop_image(img, tol=0):\n    mask = img &gt; tol\n    return img[np.ix_(mask.any(1),mask.any(0))]\n\n\n# DOCS: https://docs.scipy.org/doc/numpy/reference/generated/numpy.pad.html\n# NOTE: assumes inverted\ndef crop_center_image(img, tol=0):\n    org_shape   = img.shape\n    img_cropped = crop_image(img)\n    pad_x       = (org_shape[0] - img_cropped.shape[0])/2\n    pad_y       = (org_shape[1] - img_cropped.shape[1])/2\n    padding     = (\n        (math.floor(pad_x), math.ceil(pad_x)),\n        (math.floor(pad_y), math.ceil(pad_y))\n    )\n    img_center = np.pad(img_cropped, padding, 'constant', constant_values=0)\n    return img_center\n\n\n# Source: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n# noinspection PyArgumentList\ndef transform_X(train: DataFrame, denoise=True, normalize=True, center=True, invert=True, resize=2, resize_fn=None) -&gt; np.ndarray:\n    train = (train.drop(columns='image_id', errors='ignore')\n             .values.astype('uint8')                   # unit8 for initial data processing\n             .reshape(-1, 137, 236)                    # 2D arrays for inline image processing\n    )\n    gc.collect(); sleep(1)\n\n    # Invert for processing\n    # Colors   |   0 = black      | 255 = white\n    # invert   |   0 = background | 255 = line\n    # original | 255 = background |   0 = line\n    train = (255-train)\n    \n    if denoise:                     \n        # Rescale lines to maximum brightness, and set background values (less than 2x mean()) to 0\n        train = np.array([ train[i] + (255-train[i].max())              for i in range(train.shape[0]) ])        \n        train = np.array([ train[i] * (train[i] &gt;= np.mean(train[i])*2) for i in range(train.shape[0]) ])                                  \n            \n    if isinstance(resize, bool) and resize == True:\n        resize = 2    # Reduce image size by 2x\n    if resize and resize != 1:                  \n        # NOTEBOOK: https://www.kaggle.com/jamesmcguigan/bengali-ai-image-processing/\n        # Out of the different resize functions:\n        # - np.mean(dtype=uint8) produces produces fragmented images (needs float16 to work properly - but RAM intensive)\n        # - np.median() produces the most accurate downsampling\n        # - np.max() produces an enhanced image with thicker lines (maybe slightly easier to read)\n        # - np.min() produces a  dehanced image with thiner lines (harder to read)\n        resize_fn = resize_fn or (np.max if invert else np.min)\n        cval      = 0 if invert else 255\n        train = skimage.measure.block_reduce(train, (1, resize,resize), cval=cval, func=resize_fn)\n            \n    if center:\n        # NOTE: crop_center_image() assumes inverted\n        train = np.array([\n            crop_center_image(train[i,:,:])\n            for i in range(train.shape[0])\n        ])\n        \n    # Un-invert if invert==False\n    if not invert: train = (255-train)\n\n    if normalize:\n        train = train.astype('float16') / 255.0   # prevent division cast: int -&gt; float64\n\n    train = train.reshape(*train.shape, 1)        # 4D ndarray for tensorflow CNN\n\n    gc.collect(); sleep(1)\n    return train\n```",
    "769836": "So what you did for size of images instead of cropping? just resize? or none of them. feeding original images size to model?",
    "769638": "nice\n",
    "767782": "For me even subtracting from 255 and normalizing by max gave poor results, still don't know why. I'm worried there is some noise in the dataset that the model is learning and even the public dataset has the same noise, would be bad if its true.",
    "767758": "Nice explanation. This can also explain why augmentations like shiftscale or randomresizedcrop doesn't work much for me. Maybe we should choose those methods which don't change the character shape.",
    "767674": "great!",
    "767593": "Thanks for your test! It solved the confusion I had from the beginning. 👍 ",
    "767473": "or it has some data leak?)",
    "767416": "thanks for sharing @quandapro ",
    "767631": ""
  }
}