{
  "id": 197823,
  "title": "Run Length Encoding - HELP!!",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/197823",
  "author_name": "Aditya Baurai",
  "post_date": "2020-11-18T08:05:21.547000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So, I came across this popular notebook for generating mask using Run Length Encoding(and I saw many notebooks using it too) : </p>\n<p><strong><a href=\"https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode\" target=\"_blank\">https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode</a></strong></p>\n<h2>The code :</h2>\n<p>def rle2mask(mask_rle, shape=(1600,256)):<br>\n    '''<br>\n    mask_rle: run-length as string formated (start length)<br>\n    shape: (width,height) of array to return <br>\n    Returns numpy array, 1 - mask, 0 - background</p>\n<pre><code>'''\ns = mask_rle.split()\nstarts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\nstarts -= 1\nends = starts + lengths\nimg = np.zeros(shape[0]*shape[1], dtype=np.uint8)\nfor lo, hi in zip(starts, ends):\n    img[lo:hi] = 1\nreturn img.reshape(shape).T\n</code></pre>\n<hr>\n<p>Following are my questions : <br>\nQ1. What is the purpose of <strong>starts = starts - 1</strong>?<br>\nQ2. Why are we adding starts to lengths?<br>\nQ3. Upon googling RLE was defined as a sequence - <em>number of times a character appeared | value of the character</em>. Eg: AAAAA ………so it's RLE will be 5A…Now, following the similar logic for the statement <strong>starts, lengths = [np.array(x) for x in (s[0:][::2], s[1:][::2])]</strong> ……..shouldn't it be <strong>lengths</strong> first as in that 5A example we place the frequency(length) first then the character.</p>\n<p>It will be a huge help if someone please clarifies these!</p>\n<p><em>Thank you and have a great day</em><br>\n~<strong>Aditya</strong></p>",
  "messages": [
    {
      "id": 1082771,
      "postDate": "2020-11-18T08:05:21.547Z",
      "content": "<p>So, I came across this popular notebook for generating mask using Run Length Encoding(and I saw many notebooks using it too) : </p>\n<p><strong><a href=\"https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode\" target=\"_blank\">https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode</a></strong></p>\n<h2>The code :</h2>\n<p>def rle2mask(mask_rle, shape=(1600,256)):<br>\n    '''<br>\n    mask_rle: run-length as string formated (start length)<br>\n    shape: (width,height) of array to return <br>\n    Returns numpy array, 1 - mask, 0 - background</p>\n<pre><code>'''\ns = mask_rle.split()\nstarts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\nstarts -= 1\nends = starts + lengths\nimg = np.zeros(shape[0]*shape[1], dtype=np.uint8)\nfor lo, hi in zip(starts, ends):\n    img[lo:hi] = 1\nreturn img.reshape(shape).T\n</code></pre>\n<hr>\n<p>Following are my questions : <br>\nQ1. What is the purpose of <strong>starts = starts - 1</strong>?<br>\nQ2. Why are we adding starts to lengths?<br>\nQ3. Upon googling RLE was defined as a sequence - <em>number of times a character appeared | value of the character</em>. Eg: AAAAA ………so it's RLE will be 5A…Now, following the similar logic for the statement <strong>starts, lengths = [np.array(x) for x in (s[0:][::2], s[1:][::2])]</strong> ……..shouldn't it be <strong>lengths</strong> first as in that 5A example we place the frequency(length) first then the character.</p>\n<p>It will be a huge help if someone please clarifies these!</p>\n<p><em>Thank you and have a great day</em><br>\n~<strong>Aditya</strong></p>",
      "rawMarkdown": "So, I came across this popular notebook for generating mask using Run Length Encoding(and I saw many notebooks using it too) : \n\n**https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode**\n\nThe code :\n-------------------------------------------------------------\ndef rle2mask(mask_rle, shape=(1600,256)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n\n-------------------------------------------------------------\n\nFollowing are my questions : \nQ1. What is the purpose of **starts = starts - 1**?\nQ2. Why are we adding starts to lengths?\nQ3. Upon googling RLE was defined as a sequence - *number of times a character appeared | value of the character*. Eg: AAAAA .........so it's RLE will be 5A...Now, following the similar logic for the statement **starts, lengths = [np.array(x) for x in (s[0:][::2], s[1:][::2])]** ........shouldn't it be **lengths** first as in that 5A example we place the frequency(length) first then the character.\n\nIt will be a huge help if someone please clarifies these!\n\n*Thank you and have a great day*\n~**Aditya**",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1082771": "So, I came across this popular notebook for generating mask using Run Length Encoding(and I saw many notebooks using it too) : \n\n**https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode**\n\nThe code :\n-------------------------------------------------------------\ndef rle2mask(mask_rle, shape=(1600,256)):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n\n-------------------------------------------------------------\n\nFollowing are my questions : \nQ1. What is the purpose of **starts = starts - 1**?\nQ2. Why are we adding starts to lengths?\nQ3. Upon googling RLE was defined as a sequence - *number of times a character appeared | value of the character*. Eg: AAAAA .........so it's RLE will be 5A...Now, following the similar logic for the statement **starts, lengths = [np.array(x) for x in (s[0:][::2], s[1:][::2])]** ........shouldn't it be **lengths** first as in that 5A example we place the frequency(length) first then the character.\n\nIt will be a huge help if someone please clarifies these!\n\n*Thank you and have a great day*\n~**Aditya**"
  }
}