{
  "id": 279923,
  "title": "Understanding Run Length Encoding",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/279923",
  "author_name": "",
  "post_date": "2021-10-19T18:58:07.138671300Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<h3>About</h3>\n<ul>\n<li><p><strong>Run-length encoding (RLE)</strong> is a form of <strong>lossless data compression</strong> in which runs of data (sequences in which the same data value occurs in many consecutive data elements) are stored as a single data value and count, rather than as the original run. </p></li>\n<li><p>This is most efficient on data that contains many such runs.</p></li>\n<li><p>For example simple graphic images such as icons, line drawings, Conway's Game of Life, and animations. </p></li>\n<li><p>For files that do not have many runs, RLE could increase the file size.</p></li>\n<li><p>RLE is suited for compressing any type of data regardless of its information content, but the content of the data will affect the compression ratio achieved by RLE. </p></li>\n<li><p>Although most RLE algorithms cannot achieve the high compression ratios of the more advanced compression methods, RLE is both <em>easy to implement and quick to execute</em>, making it a good alternative to either using a complex compression algorithm or leaving your image data uncompressed.</p></li>\n<li><p>RLE works by reducing the physical size of a repeating string of characters. </p></li>\n<li><p>This repeating string, called a run, is typically encoded into two bytes. </p></li>\n<li><p>The first byte represents the number of characters in the run and is called the run count. In practice, an encoded run may contain 1 to 128 or 256 characters; the run count usually contains as the number of characters minus one (a value in the range of 0 to 127 or 255). </p></li>\n<li><p>The second byte is the value of the character in the run, which is in the range of 0 to 255, and is called the run value.  </p></li>\n</ul>\n<hr>\n<h3>Example - 1</h3>\n<p><strong>Uncompressed</strong>, a character run of 15 A characters would normally require 15 bytes to store:  <br>\n<code>AAAAAAAAAAAAAAA</code></p>\n<p>The same string after <strong>RLE encoding</strong> would require only two bytes:  <br>\n<code>15A</code></p>\n<p>The <code>15A</code> code generated to represent the character string is called an <code>RLE packet</code>. Here, the first byte, <code>15</code>, is the run count and contains the number of repetitions. The second byte, <code>A</code>, is the run value and contains the actual repeated value in the run.</p>\n<hr>\n<h3>Example - 2</h3>\n<p>Consider a screen containing plain black text on a solid white background, over hypothetical scan line, it can be rendered as follows:  </p>\n<p><code>12W1B12W3B24W1B14W2W</code>  </p>\n<p>This can be interpreted as a sequence of twelve Ws, one B, twelve Ws, three Bs, etc., and represents the original 67 characters in only 18.   </p>\n<p>While the actual format used for the storage of images is generally binary rather than ASCII characters like this, the principle remains the same.</p>\n<hr>\n<h3>Implementation:</h3>\n<blockquote>\n  <p>Perform Run–length encoding (RLE) data compression algorithm on string <code>str</code></p>\n</blockquote>\n<pre><code>def encode(s):\n\n    encoding = \"\" # stores output string\n\n    i = 0\n    while i &lt; len(s):\n        # count occurrences of character at index `i`\n        count = 1\n\n        while i + 1 &lt; len(s) and s[i] == s[i + 1]:\n            count = count + 1\n            i = i + 1\n\n        # append current character and its count to the result\n        encoding += str(count) + s[i]\n        i = i + 1\n\n    return encoding\n\n\nif __name__ == '__main__':\n\n    s = 'ABBCCCD'\n    print(encode(s))\n</code></pre>\n<p><strong>Output:</strong><br>\n<code>1A2B3C1D</code></p>\n<hr>\n<h3>Additional Resources</h3>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Run-length_encoding\" target=\"_blank\">Run-length encoding</a></li>\n<li><a href=\"https://www.khanacademy.org/computing/computers-and-internet/xcae6f4a7ff015e7d:digital-information/xcae6f4a7ff015e7d:data-compression/a/simple-image-compression\" target=\"_blank\">Lossless image compression</a></li>\n<li><a href=\"https://www.csfieldguide.org.nz/en/chapters/coding-compression/run-length-encoding/\" target=\"_blank\">Coding - Compression Run length encoding</a></li>\n<li><a href=\"https://www.fileformat.info/mirror/egff/ch09_03.htm\" target=\"_blank\">Run-Length Encoding (RLE)</a></li>\n</ul>\n<hr>",
  "messages": [
    {
      "id": "1550476",
      "postDate": "10/19/2021 18:58:07",
      "content": "<h3>About</h3>\n<ul>\n<li><p><strong>Run-length encoding (RLE)</strong> is a form of <strong>lossless data compression</strong> in which runs of data (sequences in which the same data value occurs in many consecutive data elements) are stored as a single data value and count, rather than as the original run. </p></li>\n<li><p>This is most efficient on data that contains many such runs.</p></li>\n<li><p>For example simple graphic images such as icons, line drawings, Conway's Game of Life, and animations. </p></li>\n<li><p>For files that do not have many runs, RLE could increase the file size.</p></li>\n<li><p>RLE is suited for compressing any type of data regardless of its information content, but the content of the data will affect the compression ratio achieved by RLE. </p></li>\n<li><p>Although most RLE algorithms cannot achieve the high compression ratios of the more advanced compression methods, RLE is both <em>easy to implement and quick to execute</em>, making it a good alternative to either using a complex compression algorithm or leaving your image data uncompressed.</p></li>\n<li><p>RLE works by reducing the physical size of a repeating string of characters. </p></li>\n<li><p>This repeating string, called a run, is typically encoded into two bytes. </p></li>\n<li><p>The first byte represents the number of characters in the run and is called the run count. In practice, an encoded run may contain 1 to 128 or 256 characters; the run count usually contains as the number of characters minus one (a value in the range of 0 to 127 or 255). </p></li>\n<li><p>The second byte is the value of the character in the run, which is in the range of 0 to 255, and is called the run value.  </p></li>\n</ul>\n<hr>\n<h3>Example - 1</h3>\n<p><strong>Uncompressed</strong>, a character run of 15 A characters would normally require 15 bytes to store:  <br>\n<code>AAAAAAAAAAAAAAA</code></p>\n<p>The same string after <strong>RLE encoding</strong> would require only two bytes:  <br>\n<code>15A</code></p>\n<p>The <code>15A</code> code generated to represent the character string is called an <code>RLE packet</code>. Here, the first byte, <code>15</code>, is the run count and contains the number of repetitions. The second byte, <code>A</code>, is the run value and contains the actual repeated value in the run.</p>\n<hr>\n<h3>Example - 2</h3>\n<p>Consider a screen containing plain black text on a solid white background, over hypothetical scan line, it can be rendered as follows:  </p>\n<p><code>12W1B12W3B24W1B14W2W</code>  </p>\n<p>This can be interpreted as a sequence of twelve Ws, one B, twelve Ws, three Bs, etc., and represents the original 67 characters in only 18.   </p>\n<p>While the actual format used for the storage of images is generally binary rather than ASCII characters like this, the principle remains the same.</p>\n<hr>\n<h3>Implementation:</h3>\n<blockquote>\n  <p>Perform Run–length encoding (RLE) data compression algorithm on string <code>str</code></p>\n</blockquote>\n<pre><code>def encode(s):\n\n    encoding = \"\" # stores output string\n\n    i = 0\n    while i &lt; len(s):\n        # count occurrences of character at index `i`\n        count = 1\n\n        while i + 1 &lt; len(s) and s[i] == s[i + 1]:\n            count = count + 1\n            i = i + 1\n\n        # append current character and its count to the result\n        encoding += str(count) + s[i]\n        i = i + 1\n\n    return encoding\n\n\nif __name__ == '__main__':\n\n    s = 'ABBCCCD'\n    print(encode(s))\n</code></pre>\n<p><strong>Output:</strong><br>\n<code>1A2B3C1D</code></p>\n<hr>\n<h3>Additional Resources</h3>\n<ul>\n<li><a href=\"https://en.wikipedia.org/wiki/Run-length_encoding\" target=\"_blank\">Run-length encoding</a></li>\n<li><a href=\"https://www.khanacademy.org/computing/computers-and-internet/xcae6f4a7ff015e7d:digital-information/xcae6f4a7ff015e7d:data-compression/a/simple-image-compression\" target=\"_blank\">Lossless image compression</a></li>\n<li><a href=\"https://www.csfieldguide.org.nz/en/chapters/coding-compression/run-length-encoding/\" target=\"_blank\">Coding - Compression Run length encoding</a></li>\n<li><a href=\"https://www.fileformat.info/mirror/egff/ch09_03.htm\" target=\"_blank\">Run-Length Encoding (RLE)</a></li>\n</ul>\n<hr>",
      "rawMarkdown": "### About\n\n- **Run-length encoding (RLE)** is a form of **lossless data compression** in which runs of data (sequences in which the same data value occurs in many consecutive data elements) are stored as a single data value and count, rather than as the original run. \n- This is most efficient on data that contains many such runs.\n- For example simple graphic images such as icons, line drawings, Conway's Game of Life, and animations. \n- For files that do not have many runs, RLE could increase the file size.\n- RLE is suited for compressing any type of data regardless of its information content, but the content of the data will affect the compression ratio achieved by RLE. \n- Although most RLE algorithms cannot achieve the high compression ratios of the more advanced compression methods, RLE is both *easy to implement and quick to execute*, making it a good alternative to either using a complex compression algorithm or leaving your image data uncompressed.\n  \n- RLE works by reducing the physical size of a repeating string of characters. \n- This repeating string, called a run, is typically encoded into two bytes. \n- The first byte represents the number of characters in the run and is called the run count. In practice, an encoded run may contain 1 to 128 or 256 characters; the run count usually contains as the number of characters minus one (a value in the range of 0 to 127 or 255). \n- The second byte is the value of the character in the run, which is in the range of 0 to 255, and is called the run value.  \n  \n---\n  \n### Example - 1 \n  \n**Uncompressed**, a character run of 15 A characters would normally require 15 bytes to store:  \n`AAAAAAAAAAAAAAA`\n\nThe same string after **RLE encoding** would require only two bytes:  \n`15A`\n\nThe `15A` code generated to represent the character string is called an `RLE packet`. Here, the first byte, `15`, is the run count and contains the number of repetitions. The second byte, `A`, is the run value and contains the actual repeated value in the run.\n   \n---\n  \n### Example - 2\n  \nConsider a screen containing plain black text on a solid white background, over hypothetical scan line, it can be rendered as follows:  \n  \n`12W1B12W3B24W1B14W2W`  \n  \nThis can be interpreted as a sequence of twelve Ws, one B, twelve Ws, three Bs, etc., and represents the original 67 characters in only 18.   \n  \nWhile the actual format used for the storage of images is generally binary rather than ASCII characters like this, the principle remains the same.\n  \n---\n  \n### Implementation:\n> Perform Run–length encoding (RLE) data compression algorithm on string `str`\n```\ndef encode(s):\n \n    encoding = \"\" # stores output string\n \n    i = 0\n    while i < len(s):\n        # count occurrences of character at index `i`\n        count = 1\n \n        while i + 1 < len(s) and s[i] == s[i + 1]:\n            count = count + 1\n            i = i + 1\n \n        # append current character and its count to the result\n        encoding += str(count) + s[i]\n        i = i + 1\n \n    return encoding\n \n \nif __name__ == '__main__':\n \n    s = 'ABBCCCD'\n    print(encode(s))\n```\n\n**Output:**\n`1A2B3C1D`\n  \n---\n  \n### Additional Resources\n- [Run-length encoding](https://en.wikipedia.org/wiki/Run-length_encoding)\n- [Lossless image compression](https://www.khanacademy.org/computing/computers-and-internet/xcae6f4a7ff015e7d:digital-information/xcae6f4a7ff015e7d:data-compression/a/simple-image-compression)\n- [Coding - Compression Run length encoding](https://www.csfieldguide.org.nz/en/chapters/coding-compression/run-length-encoding/)\n- [Run-Length Encoding (RLE)](https://www.fileformat.info/mirror/egff/ch09_03.htm)\n\n---",
      "votes": null
    },
    {
      "id": "1550957",
      "postDate": "10/20/2021 07:18:42",
      "content": "<p>Nice article ! now i understand RLE</p>",
      "rawMarkdown": "Nice article ! now i understand RLE",
      "votes": null
    },
    {
      "id": "1551180",
      "postDate": "10/20/2021 11:40:12",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/vatsalsavaliya\" target=\"_blank\">@vatsalsavaliya</a> </p>",
      "rawMarkdown": "Thank you @vatsalsavaliya",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1550957,
      "author_name": "vatsalsavaliya",
      "author_url": "",
      "post_date": "10/20/2021 07:18:42",
      "content": "<p>Nice article ! now i understand RLE</p>",
      "votes": null,
      "replies": [
        {
          "id": 1551180,
          "author_name": "ishandutta",
          "author_url": "",
          "post_date": "10/20/2021 11:40:12",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/vatsalsavaliya\" target=\"_blank\">@vatsalsavaliya</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1550476": "### About\n\n- **Run-length encoding (RLE)** is a form of **lossless data compression** in which runs of data (sequences in which the same data value occurs in many consecutive data elements) are stored as a single data value and count, rather than as the original run. \n- This is most efficient on data that contains many such runs.\n- For example simple graphic images such as icons, line drawings, Conway's Game of Life, and animations. \n- For files that do not have many runs, RLE could increase the file size.\n- RLE is suited for compressing any type of data regardless of its information content, but the content of the data will affect the compression ratio achieved by RLE. \n- Although most RLE algorithms cannot achieve the high compression ratios of the more advanced compression methods, RLE is both *easy to implement and quick to execute*, making it a good alternative to either using a complex compression algorithm or leaving your image data uncompressed.\n  \n- RLE works by reducing the physical size of a repeating string of characters. \n- This repeating string, called a run, is typically encoded into two bytes. \n- The first byte represents the number of characters in the run and is called the run count. In practice, an encoded run may contain 1 to 128 or 256 characters; the run count usually contains as the number of characters minus one (a value in the range of 0 to 127 or 255). \n- The second byte is the value of the character in the run, which is in the range of 0 to 255, and is called the run value.  \n  \n---\n  \n### Example - 1 \n  \n**Uncompressed**, a character run of 15 A characters would normally require 15 bytes to store:  \n`AAAAAAAAAAAAAAA`\n\nThe same string after **RLE encoding** would require only two bytes:  \n`15A`\n\nThe `15A` code generated to represent the character string is called an `RLE packet`. Here, the first byte, `15`, is the run count and contains the number of repetitions. The second byte, `A`, is the run value and contains the actual repeated value in the run.\n   \n---\n  \n### Example - 2\n  \nConsider a screen containing plain black text on a solid white background, over hypothetical scan line, it can be rendered as follows:  \n  \n`12W1B12W3B24W1B14W2W`  \n  \nThis can be interpreted as a sequence of twelve Ws, one B, twelve Ws, three Bs, etc., and represents the original 67 characters in only 18.   \n  \nWhile the actual format used for the storage of images is generally binary rather than ASCII characters like this, the principle remains the same.\n  \n---\n  \n### Implementation:\n> Perform Run–length encoding (RLE) data compression algorithm on string `str`\n```\ndef encode(s):\n \n    encoding = \"\" # stores output string\n \n    i = 0\n    while i < len(s):\n        # count occurrences of character at index `i`\n        count = 1\n \n        while i + 1 < len(s) and s[i] == s[i + 1]:\n            count = count + 1\n            i = i + 1\n \n        # append current character and its count to the result\n        encoding += str(count) + s[i]\n        i = i + 1\n \n    return encoding\n \n \nif __name__ == '__main__':\n \n    s = 'ABBCCCD'\n    print(encode(s))\n```\n\n**Output:**\n`1A2B3C1D`\n  \n---\n  \n### Additional Resources\n- [Run-length encoding](https://en.wikipedia.org/wiki/Run-length_encoding)\n- [Lossless image compression](https://www.khanacademy.org/computing/computers-and-internet/xcae6f4a7ff015e7d:digital-information/xcae6f4a7ff015e7d:data-compression/a/simple-image-compression)\n- [Coding - Compression Run length encoding](https://www.csfieldguide.org.nz/en/chapters/coding-compression/run-length-encoding/)\n- [Run-Length Encoding (RLE)](https://www.fileformat.info/mirror/egff/ch09_03.htm)\n\n---",
    "1550957": "Nice article ! now i understand RLE",
    "1551180": "Thank you @vatsalsavaliya"
  },
  "source": "meta"
}