{"metadata": {"language_info": {"file_extension": ".py", "name": "python", "version": "3.6.1", "mimetype": "text/x-python", "nbconvert_exporter": "python", "codemirror_mode": {"name": "ipython", "version": 3}, "pygments_lexer": "ipython3"}, "kernelspec": {"display_name": "Python 3", "name": "python3", "language": "python"}}, "nbformat_minor": 1, "cells": [{"metadata": {"_cell_guid": "07257c1c-e3bd-4885-8c79-e52f45f647eb", "_uuid": "e52591953a08da9f7bfe43ef4691cb93f0183328"}, "cell_type": "markdown", "source": ["# Acknowledgments\n", "\n", "I've tried some of the kernels and my implementation for self-learning.  \n", "Reference of the kernels are bellow. Thanks everyone for great posts!\n", "\n", "* [Fast, tested RLE](https://www.kaggle.com/stainsby/fast-tested-rle) by [Sam Stains](https://www.kaggle.com/stainsby)\n", "* [Fast Run Length Encode](https://www.kaggle.com/paulorzp/fast-run-length-encode) by [Paulo Pinto](https://www.kaggle.com/paulorzp)\n", "* [Even Faster Run Length Encoder](https://www.kaggle.com/hackerpoet/even-faster-run-length-encoder) by [Kevin H](https://www.kaggle.com/hackerpoet) and [jeffeverett](https://www.kaggle.com/jeffeverett)"]}, {"metadata": {"_cell_guid": "d327e048-ba9c-4903-977b-637fa6bbecbf", "_uuid": "57517aead30384d07228ff8c87cd7968a77257ce"}, "cell_type": "markdown", "source": ["## Summary\n", "* about 6-7 ms / image: [Fast, tested RLE](https://www.kaggle.com/stainsby/fast-tested-rle) by [Sam Stains](https://www.kaggle.com/stainsby)\n", "* about 180 ms / image: [Fast Run Length Encode](https://www.kaggle.com/paulorzp/fast-run-length-encode) by [Paulo Pinto](https://www.kaggle.com/paulorzp)\n", "* about 18- ms / image: [Even Faster Run Length Encoder](https://www.kaggle.com/hackerpoet/even-faster-run-length-encoder) by [Kevin H](https://www.kaggle.com/hackerpoet) and [jeffeverett](https://www.kaggle.com/jeffeverett)\n", "* about 7-8 ms / image: my code\n", "\n", "I tried 1 - 100 image.  \n", "In this kernel's speed-test method, time-increases with the number of images.  \n", "\n", "**I want to share if there is a faster way:)**"]}, {"metadata": {"_cell_guid": "8a1bcb61-f722-47f3-be16-22a2f387bb97", "_uuid": "1906c53afd52d368fa10d4bdf7531220f4ef9a51"}, "cell_type": "markdown", "source": ["## Initalization"]}, {"metadata": {"_cell_guid": "f0695cc1-aa0d-4351-9604-b819f903c65d", "_uuid": "18d3ec9b3ac47b1f20b0d6baa4c195986116bafc", "collapsed": true}, "cell_type": "code", "source": ["%matplotlib inline\n", "\n", "import os\n", "import re\n", "import scipy.misc\n", "import numpy as np\n", "import matplotlib.pyplot as plt\n", "from PIL import Image"], "outputs": [], "execution_count": 1}, {"metadata": {"_cell_guid": "a1d555bb-5553-404b-99bd-62fee7fd9320", "_uuid": "22506aaccd51d6618ca53fa48f4dd8f3001a6102"}, "cell_type": "code", "source": ["# Get filenames list of image files\n", "re_masks = re.compile('(^(.+?)_[0-9]+?)_mask\\.gif$')\n", "list_masks = [name for name in os.listdir('../input/train_masks/') if re_masks.search(name)]\n", "list_masks.sort()\n", "list_masks[0:5]"], "outputs": [], "execution_count": 2}, {"metadata": {"_cell_guid": "fa5e3ea7-b315-41f3-9469-1f086d5962ee", "scrolled": true, "_uuid": "2deecb82ae0c6f6882de1477509b63815a3f2441"}, "cell_type": "code", "source": ["# Check image loading\n", "np_img_tmp = np.uint8(Image.open('../input/train_masks/' + list_masks[0]))\n", "plt.imshow(np_img_tmp)\n", "print('min:%s, max: %s' % (np_img_tmp.min(), np_img_tmp.max()))\n", "np_img_tmp"], "outputs": [], "execution_count": 3}, {"metadata": {"_cell_guid": "2f843f0d-6bf6-4164-ab90-4517f80a386e", "_uuid": "d91972353157f7e46534ad5791c438b139c46b30"}, "cell_type": "code", "source": ["# Make speed-test dataset, shape = (N, 1280, 1918), N=100\n", "list_np_dataset_nx1280x1918 = []\n", "for img_mask in list_masks[0:100]:\n", "    list_np_dataset_nx1280x1918.append(np.uint8(Image.open('../input/train_masks/' + img_mask)))\n", "np_dataset_nx1280x1918 = np.array(list_np_dataset_nx1280x1918)\n", "np_dataset_nx1280x1918.shape"], "outputs": [], "execution_count": 4}, {"metadata": {"_cell_guid": "89ed2927-af00-4b92-bc5f-396a6b411efd", "_uuid": "688c52339f8c28e4bc156e111111e72366156fb5"}, "cell_type": "markdown", "source": ["## [Fast, tested RLE](https://www.kaggle.com/stainsby/fast-tested-rle) by [Sam Stains](https://www.kaggle.com/stainsby)"]}, {"metadata": {"_cell_guid": "f19007f1-8ad4-4df2-935a-332b7bc88dba", "_uuid": "19ce12f43f3bfa410a566d12515b7535f4f18eda", "collapsed": true}, "cell_type": "code", "source": ["# REF: https://www.kaggle.com/stainsby/fast-tested-rle\n", "\n", "def rle_encode(mask_image):\n", "    pixels = mask_image.flatten()\n", "    # We avoid issues with '1' at the start or end (at the corners of \n", "    # the original image) by setting those pixels to '0' explicitly.\n", "    # We do not expect these to be non-zero for an accurate mask, \n", "    # so this should not harm the score.\n", "    pixels[0] = 0\n", "    pixels[-1] = 0\n", "    runs = np.where(pixels[1:] != pixels[:-1])[0] + 2\n", "    runs[1::2] = runs[1::2] - runs[:-1:2]\n", "    return runs\n", "\n", "\n", "def rle_to_string(runs):\n", "    return ' '.join(str(x) for x in runs)"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "2a9854cd-6f59-437a-b8bd-35b9cfa77e83", "_uuid": "b29ecdfbddd97e4282f28e70047960c496db3408", "collapsed": true}, "cell_type": "code", "source": ["# For run like map-function\n", "def mapfunc_img_to_rle_stainsby(np_img_mask_nx1280x1918):\n", "    # Get RLE string from 1280 * 1918 image by N * 1280 * 1918 dataset\n", "    return [rle_to_string(rle_encode(np_img_mask_1280x1918)) for np_img_mask_1280x1918 in np_img_mask_nx1280x1918]"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "baa7887a-7abf-4f03-9af1-5592919910f3", "_uuid": "2bf990c21a81a30d9467d0d1f17f519c60440bde", "collapsed": true}, "cell_type": "code", "source": ["# convert 1 images\n", "mapfunc_img_to_rle_stainsby(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "ed96402a-1f68-43e9-aeb4-02743daf06ab", "_uuid": "42f6273f0a27ba30d7674e654874f2e3595139ef", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 1 images\n", "mapfunc_img_to_rle_stainsby(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "a6ed748e-eb84-4b6b-b3f9-2ab0e3d99b8d", "_uuid": "92571ec099bfa74e839d224432e49322b9da6177", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 10 images\n", "mapfunc_img_to_rle_stainsby(np_dataset_nx1280x1918[0:10])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "bdccf30a-3086-429f-b1ce-79e983370e43", "_uuid": "6ea736029aa678f6ca396a44d467248f31143cba"}, "cell_type": "markdown", "source": ["## [Fast Run Length Encode](https://www.kaggle.com/paulorzp/fast-run-length-encode) by [Paulo Pinto](https://www.kaggle.com/paulorzp)"]}, {"metadata": {"_cell_guid": "0cce208f-e9b0-445b-9f46-1f86b712e370", "_uuid": "530feb460ec17950f6ba0923efef0af61fd4b191", "collapsed": true}, "cell_type": "code", "source": ["# REF: https://www.kaggle.com/paulorzp/fast-run-length-encode\n", "\n", "def rle (img):\n", "    '''\n", "    img: numpy array, 1 - mask, 0 - background\n", "    Returns run length as string formated\n", "    '''\n", "    bytes = np.where(img.flatten()==1)[0]\n", "    runs = []\n", "    prev = -2\n", "    for b in bytes:\n", "        if (b>prev+1): runs.extend((b+1, 0))\n", "        runs[-1] += 1\n", "        prev = b\n", "    \n", "    return ' '.join([str(i) for i in runs])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "38ccdbc5-b36e-4100-8c19-5f85c73eecab", "_uuid": "decfb5ae161b00f94ed5e298a6b4f1de85dcad63", "collapsed": true}, "cell_type": "code", "source": ["# For run like map-function\n", "def mapfunc_img_to_rle_paulorzp(np_img_mask_nx1280x1918):\n", "    # Get RLE string from 1280 * 1918 image by N * 1280 * 1918 dataset\n", "    return [rle(np_img_mask_1280x1918) for np_img_mask_1280x1918 in np_img_mask_nx1280x1918]"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "9abdc189-4898-4d5b-9b51-c2e480d60a1e", "_uuid": "60e253f6aabf9f33c93e780ca566d46b7938fdae", "collapsed": true}, "cell_type": "code", "source": ["# convert 1 images\n", "mapfunc_img_to_rle_paulorzp(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "47294a4b-0cb6-4406-88e3-160225a88337", "_uuid": "7b06d0d06f7fcc9a1c2ba7fc7e7cd729a5102f54", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 1 images\n", "mapfunc_img_to_rle_paulorzp(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "c712f3bb-ab52-4ba4-816b-b46d219ff795", "_uuid": "594742013f723e992601a8cf703d8443fb597246", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 10 images\n", "mapfunc_img_to_rle_paulorzp(np_dataset_nx1280x1918[0:10])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "dd2a1eb0-6849-47cc-9e04-74ea2a8e1df5", "_uuid": "6d5217a00e9324e768a4b40d958b8981d62e7b8e"}, "cell_type": "markdown", "source": ["## [Even Faster Run Length Encoder](https://www.kaggle.com/hackerpoet/even-faster-run-length-encoder) by [Kevin H](https://www.kaggle.com/hackerpoet) and [jeffeverett](https://www.kaggle.com/jeffeverett)\n", "\n"]}, {"metadata": {"_cell_guid": "9dcafdf1-22ca-48d6-928b-1f7e23253c95", "_uuid": "d36e8262c8d8c7cb687fd8a230186d155f9cfc0d", "collapsed": true}, "cell_type": "code", "source": ["# REF: https://www.kaggle.com/hackerpoet/even-faster-run-length-encoder\n", "\n", "def rle_kevin_and_jeffeverett(img):\n", "    flat_img = img.flatten()\n", "    flat_img = np.where(flat_img > 0.5, 1, 0).astype(np.uint8)\n", "    flat_img = np.insert(flat_img, [0, len(flat_img)], [0, 0])\n", "\n", "    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n", "    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n", "    starts_ix = np.where(starts)[0] + 1\n", "    ends_ix = np.where(ends)[0] + 1\n", "    lengths = ends_ix - starts_ix\n", "\n", "    encoding = ''\n", "    for idx in range(len(starts_ix)):\n", "        encoding += '%d %d ' % (starts_ix[idx], lengths[idx])\n", "    return encoding"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "cf6c3ef5-727d-4d35-995e-32bff934a2b2", "_uuid": "e9ce03bc644caa6c6ca99e9a843603b7267b4a88", "collapsed": true}, "cell_type": "code", "source": ["# For run like map-function\n", "def mapfunc_img_to_rle_kevin_and_jeffeverett(np_img_mask_nx1280x1918):\n", "    # Get RLE string from 1280 * 1918 image by N * 1280 * 1918 dataset\n", "    return [rle_kevin_and_jeffeverett(np_img_mask_1280x1918) for np_img_mask_1280x1918 in np_img_mask_nx1280x1918]"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "60d3bc81-f02f-4e5c-8cb0-e964806df915", "scrolled": false, "_uuid": "a439ca68659bfd275ca077287e7db1d8a449b940", "collapsed": true}, "cell_type": "code", "source": ["# convert 1 images\n", "mapfunc_img_to_rle_kevin_and_jeffeverett(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "03a57958-2c49-483e-95ab-cf48892230bc", "_uuid": "ac1c84e142ca8d72e52a0390f14eb19f62e016a2", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 1 images\n", "mapfunc_img_to_rle_kevin_and_jeffeverett(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "36374c02-f57c-4ea5-8717-5714c5bcf30a", "_uuid": "df1175ae23689bb893f5894d6353078a5acaa3b3", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 10 images\n", "mapfunc_img_to_rle_kevin_and_jeffeverett(np_dataset_nx1280x1918[0:10])"], "outputs": [], "execution_count": null}, {"metadata": {"_cell_guid": "4aba33a0-e58d-4b0a-bb95-6b32259a12f5", "_uuid": "4f43573b977bc0afe9e168d0a4ec63e918cac3e8"}, "cell_type": "markdown", "source": ["## My code\n", "**numpy.diff() and numpy.where() is very powerful.**  \n", "TODO: buffer-overflow?"]}, {"metadata": {"_cell_guid": "480b197a-3984-41a6-9bb6-49ddd4ed56b8", "_uuid": "f6abd7b6b1aa86743356d2186e8e80791e1b62c3", "collapsed": true}, "cell_type": "code", "source": ["def np_2d_img_to_str_rle(np_mask_img):\n", "    np_mask_img_vec = np_mask_img.reshape(np_mask_img.shape[0] * np_mask_img.shape[1] )\n", "    np_diff = np.diff(np_mask_img_vec)\n", "    np_where_start = np.where(np_diff == 1)[0]\n", "    np_where_end = np.where(np_diff == 255)[0] # -1 -> 255 in np.uint8, buffer-overflow\n", "    start = np_where_start + 2\n", "    end = np_where_end - start + 2\n", "    list_output = ['%s %s' % (start[i], end[i]) for i in range(len(np_where_start))]\n", "    str_output = ' '.join(list_output)\n", "    return str_output\n", "\n", "# For run like map-function\n", "def mapfunc_np_2d_img_to_str_rle(np_img_mask_nx1280x1918):\n", "    # Get RLE string from 1280 * 1918 image by N * 1280 * 1918 dataset\n", "    return [np_2d_img_to_str_rle(np_img_mask_1280x1918) for np_img_mask_1280x1918 in np_img_mask_nx1280x1918]"], "outputs": [], "execution_count": 7}, {"metadata": {"_cell_guid": "ae77bd4f-acab-4d41-9ea4-1981778fecf5", "_uuid": "f0e8f1026c8804e91aa3e336e26a61448863b89f"}, "cell_type": "code", "source": ["# convert 1 images\n", "mapfunc_np_2d_img_to_str_rle(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": 8}, {"metadata": {"_cell_guid": "08d1fc57-884b-444d-9c56-1eb0b20cb2ab", "_uuid": "42d5d3999f234d8c2e870aaa4d811a83539751e7"}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 1 images\n", "mapfunc_np_2d_img_to_str_rle(np_dataset_nx1280x1918[0:1])"], "outputs": [], "execution_count": 9}, {"metadata": {"_cell_guid": "e2936b19-467f-4b1e-9eee-4fe644037928", "_uuid": "8c37871364952705d328a70d47deb2c5a6c2a6db"}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 10 images\n", "mapfunc_np_2d_img_to_str_rle(np_dataset_nx1280x1918[0:10])"], "outputs": [], "execution_count": 10}, {"metadata": {"_cell_guid": "76d40a98-a7ce-46f4-a101-f4472c8c62eb", "_uuid": "ac330b55a323bb73526819481b7226613a9e2f91", "collapsed": true}, "cell_type": "code", "source": ["%%timeit -n20 -r5 -p5\n", "# convert 10 images\n", "mapfunc_np_2d_img_to_str_rle(np_dataset_nx1280x1918)"], "outputs": [], "execution_count": null}], "nbformat": 4}