{
  "id": 411007,
  "title": "inklabels_rle.csv differs from the generated csv using provided code snippet",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/411007",
  "author_name": "",
  "post_date": "2023-05-17T10:27:20.203178200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello! 👋</p>\n<p>I just compared the rle script output from the provided file and discovered, that they are very similar, but not equal. How can that be?  </p>\n<p>Used code:</p>\n<pre><code>\n numpy  np\n PIL  Image\n filecmp\n\n\n  (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img &gt; , , ).astype(np.uint8)\n\n    starts = np.array((flat_img[:-] == ) &amp; (flat_img[:] == ))\n    ends = np.array((flat_img[:-] == ) &amp; (flat_img[:] == ))\n    starts_ix = np.where(starts)[] + \n    ends_ix = np.where(ends)[] + \n    lengths = ends_ix - starts_ix\n\n     starts_ix, lengths\n\ninklabels = np.array(Image.(), dtype=np.uint8)\nstarts_ix, lengths = rle(inklabels)\ninklabels_rle = .join((, ((starts_ix, lengths), ())))\n( + inklabels_rle, file=(, ))\n\nexisting_path = \n\n filecmp.cmp(, existing_path):\n    ()\n:\n    ()\n</code></pre>\n<p>Examplary difference between the two:</p>\n<p>Provided file<br>\n196 4947266 205 4947745 273 <strong>4951651 2</strong> 4951892 195 4952516 204 4952994 274</p>\n<p>Generated file<br>\n196 4947266 205 4947745 273 4951892 195 4952516 204 4952994 274</p>\n<p>Cheers ✌️</p>",
  "messages": [
    {
      "id": "2263091",
      "postDate": "05/17/2023 10:27:20",
      "content": "<p>Hello! 👋</p>\n<p>I just compared the rle script output from the provided file and discovered, that they are very similar, but not equal. How can that be?  </p>\n<p>Used code:</p>\n<pre><code>\n numpy  np\n PIL  Image\n filecmp\n\n\n  (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img &gt; , , ).astype(np.uint8)\n\n    starts = np.array((flat_img[:-] == ) &amp; (flat_img[:] == ))\n    ends = np.array((flat_img[:-] == ) &amp; (flat_img[:] == ))\n    starts_ix = np.where(starts)[] + \n    ends_ix = np.where(ends)[] + \n    lengths = ends_ix - starts_ix\n\n     starts_ix, lengths\n\ninklabels = np.array(Image.(), dtype=np.uint8)\nstarts_ix, lengths = rle(inklabels)\ninklabels_rle = .join((, ((starts_ix, lengths), ())))\n( + inklabels_rle, file=(, ))\n\nexisting_path = \n\n filecmp.cmp(, existing_path):\n    ()\n:\n    ()\n</code></pre>\n<p>Examplary difference between the two:</p>\n<p>Provided file<br>\n196 4947266 205 4947745 273 <strong>4951651 2</strong> 4951892 195 4952516 204 4952994 274</p>\n<p>Generated file<br>\n196 4947266 205 4947745 273 4951892 195 4952516 204 4952994 274</p>\n<p>Cheers ✌️</p>",
      "rawMarkdown": "Hello! 👋\n\nI just compared the rle script output from the provided file and discovered, that they are very similar, but not equal. How can that be?  \n\nUsed code:\n\n```python\n#!/usr/bin/python3\nimport numpy as np\nfrom PIL import Image\nimport filecmp\n\n# Fast run length encoding, from https://www.kaggle.com/code/hackerpoet/even-faster-run-length-encoder/script\ndef rle (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img > 0.5, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n\n    return starts_ix, lengths\n\ninklabels = np.array(Image.open('/vesuvius-challenge-ink-detection/train/3/inklabels.png'), dtype=np.uint8)\nstarts_ix, lengths = rle(inklabels)\ninklabels_rle = \" \".join(map(str, sum(zip(starts_ix, lengths), ())))\nprint(\"Id,Predicted\\n3,\" + inklabels_rle, file=open('inklabels_rle_3.c.csv', 'w'))\n\nexisting_path = \"/vesuvius-challenge-ink-detection/train/3/inklabels_rle.csv\"\n\nif filecmp.cmp('inklabels_rle_3.csv', existing_path):\n    print(\"The files are identical.\")\nelse:\n    print(\"The files are different.\")\n```\n\nExamplary difference between the two:\n\nProvided file\n196 4947266 205 4947745 273 **4951651 2** 4951892 195 4952516 204 4952994 274\n\nGenerated file\n196 4947266 205 4947745 273 4951892 195 4952516 204 4952994 274\n\nCheers ✌️",
      "votes": null
    },
    {
      "id": "2263470",
      "postDate": "05/17/2023 16:07:13",
      "content": "<p>I <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/405000#2245003\" target=\"_blank\">pointed out the discrepancies</a> after the ink labels were updated—in the PNGs only, not the RLE files, which is why there are discrepancies.</p>\n<p>Last I checked, it hadn't been updated yet, but I haven't checked for a bit.</p>",
      "rawMarkdown": "I [pointed out the discrepancies](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/405000#2245003) after the ink labels were updated—in the PNGs only, not the RLE files, which is why there are discrepancies.\n\nLast I checked, it hadn't been updated yet, but I haven't checked for a bit.",
      "votes": null
    },
    {
      "id": "2265946",
      "postDate": "05/19/2023 15:58:03",
      "content": "<p>The dataset is updated with matching RLE files now. Thanks!</p>",
      "rawMarkdown": "The dataset is updated with matching RLE files now. Thanks!",
      "votes": null
    },
    {
      "id": "2281491",
      "postDate": "05/31/2023 00:11:38",
      "content": "<p>There seems to be difference that I can see as well<br>\nfor set 2</p>\n<p>def rle (img):<br>\n    flat_img = img.flatten()<br>\n    flat_img = np.where(flat_img &gt; 0.5, 1, 0).astype(np.uint8)</p>\n<pre><code>starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))\nends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))\nstarts_ix = np.where(starts)[0] + 2\nends_ix = np.where(ends)[0] + 2\nlengths = ends_ix - starts_ix\n\nreturn starts_ix, lengths\n</code></pre>\n<p>labelArray = np.array(Image.open(\"inklabels.png\"), dtype=np.uint8)<br>\nlabel_starts_ix, label_lengths = rle(labelArray)<br>\nlabelRle = inklabels_rle = \" \".join(map(str, sum(zip(label_starts_ix, label_lengths), ())))</p>\n<p>Will not be the same as in the csv file<br>\nfilename = \"inklabels_rle.csv\"<br>\ndf = pd.read_csv(filename)<br>\ncompetitionRle = df[\"Predicted\"]<br>\ncompetitionRle = competitionRle[0]</p>\n<p>labelRle == competitionRle</p>\n<p>returns False</p>",
      "rawMarkdown": "There seems to be difference that I can see as well\nfor set 2\n\ndef rle (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img > 0.5, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n\n    return starts_ix, lengths\nlabelArray = np.array(Image.open(\"inklabels.png\"), dtype=np.uint8)\nlabel_starts_ix, label_lengths = rle(labelArray)\nlabelRle = inklabels_rle = \" \".join(map(str, sum(zip(label_starts_ix, label_lengths), ())))\n\nWill not be the same as in the csv file\nfilename = \"inklabels_rle.csv\"\ndf = pd.read_csv(filename)\ncompetitionRle = df[\"Predicted\"]\ncompetitionRle = competitionRle[0]\n\nlabelRle == competitionRle\n\nreturns False",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2263470,
      "author_name": "costella",
      "author_url": "",
      "post_date": "05/17/2023 16:07:13",
      "content": "<p>I <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/405000#2245003\" target=\"_blank\">pointed out the discrepancies</a> after the ink labels were updated—in the PNGs only, not the RLE files, which is why there are discrepancies.</p>\n<p>Last I checked, it hadn't been updated yet, but I haven't checked for a bit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2265946,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "05/19/2023 15:58:03",
      "content": "<p>The dataset is updated with matching RLE files now. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2281491,
      "author_name": "abualabed",
      "author_url": "",
      "post_date": "05/31/2023 00:11:38",
      "content": "<p>There seems to be difference that I can see as well<br>\nfor set 2</p>\n<p>def rle (img):<br>\n    flat_img = img.flatten()<br>\n    flat_img = np.where(flat_img &gt; 0.5, 1, 0).astype(np.uint8)</p>\n<pre><code>starts = np.array((flat_img[:-1] == 0) &amp; (flat_img[1:] == 1))\nends = np.array((flat_img[:-1] == 1) &amp; (flat_img[1:] == 0))\nstarts_ix = np.where(starts)[0] + 2\nends_ix = np.where(ends)[0] + 2\nlengths = ends_ix - starts_ix\n\nreturn starts_ix, lengths\n</code></pre>\n<p>labelArray = np.array(Image.open(\"inklabels.png\"), dtype=np.uint8)<br>\nlabel_starts_ix, label_lengths = rle(labelArray)<br>\nlabelRle = inklabels_rle = \" \".join(map(str, sum(zip(label_starts_ix, label_lengths), ())))</p>\n<p>Will not be the same as in the csv file<br>\nfilename = \"inklabels_rle.csv\"<br>\ndf = pd.read_csv(filename)<br>\ncompetitionRle = df[\"Predicted\"]<br>\ncompetitionRle = competitionRle[0]</p>\n<p>labelRle == competitionRle</p>\n<p>returns False</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2263091": "Hello! 👋\n\nI just compared the rle script output from the provided file and discovered, that they are very similar, but not equal. How can that be?  \n\nUsed code:\n\n```python\n#!/usr/bin/python3\nimport numpy as np\nfrom PIL import Image\nimport filecmp\n\n# Fast run length encoding, from https://www.kaggle.com/code/hackerpoet/even-faster-run-length-encoder/script\ndef rle (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img > 0.5, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n\n    return starts_ix, lengths\n\ninklabels = np.array(Image.open('/vesuvius-challenge-ink-detection/train/3/inklabels.png'), dtype=np.uint8)\nstarts_ix, lengths = rle(inklabels)\ninklabels_rle = \" \".join(map(str, sum(zip(starts_ix, lengths), ())))\nprint(\"Id,Predicted\\n3,\" + inklabels_rle, file=open('inklabels_rle_3.c.csv', 'w'))\n\nexisting_path = \"/vesuvius-challenge-ink-detection/train/3/inklabels_rle.csv\"\n\nif filecmp.cmp('inklabels_rle_3.csv', existing_path):\n    print(\"The files are identical.\")\nelse:\n    print(\"The files are different.\")\n```\n\nExamplary difference between the two:\n\nProvided file\n196 4947266 205 4947745 273 **4951651 2** 4951892 195 4952516 204 4952994 274\n\nGenerated file\n196 4947266 205 4947745 273 4951892 195 4952516 204 4952994 274\n\nCheers ✌️",
    "2263470": "I [pointed out the discrepancies](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/405000#2245003) after the ink labels were updated—in the PNGs only, not the RLE files, which is why there are discrepancies.\n\nLast I checked, it hadn't been updated yet, but I haven't checked for a bit.",
    "2265946": "The dataset is updated with matching RLE files now. Thanks!",
    "2281491": "There seems to be difference that I can see as well\nfor set 2\n\ndef rle (img):\n    flat_img = img.flatten()\n    flat_img = np.where(flat_img > 0.5, 1, 0).astype(np.uint8)\n\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n\n    return starts_ix, lengths\nlabelArray = np.array(Image.open(\"inklabels.png\"), dtype=np.uint8)\nlabel_starts_ix, label_lengths = rle(labelArray)\nlabelRle = inklabels_rle = \" \".join(map(str, sum(zip(label_starts_ix, label_lengths), ())))\n\nWill not be the same as in the csv file\nfilename = \"inklabels_rle.csv\"\ndf = pd.read_csv(filename)\ncompetitionRle = df[\"Predicted\"]\ncompetitionRle = competitionRle[0]\n\nlabelRle == competitionRle\n\nreturns False"
  },
  "source": "meta"
}