{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-10-18T05:42:59.994384Z","iopub.execute_input":"2021-10-18T05:42:59.994856Z","iopub.status.idle":"2021-10-18T05:42:59.99861Z","shell.execute_reply.started":"2021-10-18T05:42:59.994792Z","shell.execute_reply":"2021-10-18T05:42:59.997864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Notebook purpose:\n- The purpose of this notebook is to ask: when projecting 1D coordinates to 2D one, how np.reshape should be use\n- In my idea:\n    - **np.reshape()** (default) chunk the 1D array row-by-row, which means left-> right then top -> bottom.\n    - However, it should be instead, **np.reshape(order='F')**, column-by-column, which means from top -> bottom then left to right","metadata":{}},{"cell_type":"markdown","source":"### According to evaluation page:\nhttps://www.kaggle.com/c/sartorius-cell-instance-segmentation/overview/evaluation\n- The competition format requires a space delimited list of pairs. For example, '1 3 10 5' implies pixels 1,2,3,10,11,12,13,14 are to be included in the mask. The pixels are one-indexed\nand numbered from **top to bottom, then left to right**: 1 is pixel 1,1; 2 is pixel 2,1, etc.","metadata":{}},{"cell_type":"markdown","source":"### Also see the explanation here:\nhttps://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/278936\n- The  first comment demonstrates how RLE works","metadata":{}},{"cell_type":"markdown","source":"### Compare the 2 types of reshaping using the above example","metadata":{}},{"cell_type":"markdown","source":"![](https://user-images.githubusercontent.com/17668390/137576003-a887b201-7cc0-4e22-975f-d7727094d1d0.png)","metadata":{}},{"cell_type":"markdown","source":"There are 48 pixels. The top left is numbered 1 then you go down that column until you hit number 8. The top of the next column starts number 9, then down to 16, etc.\n\nIf walk the pixels from 1 to 48, a line of yellow begins at pixel 11 for length 5, then another begins at 19 for length 5, then another begins at 27 for length 5 and the last begins at 37 for length 3.\n\nSo the RLE is \"11 5 19 5 27 5 37 3\".","metadata":{}},{"cell_type":"code","source":"arr2d = np.array([[0,0,0,0,0,0],\n                  [0,0,0,0,0,0],\n                  [0,1,1,1,0,0],\n                  [0,1,1,1,0,0],\n                  [0,1,1,1,1,0],\n                  [0,1,1,1,1,0],\n                  [0,1,1,1,1,0],\n                  [0,0,0,0,0,0]])","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:04.768927Z","iopub.execute_input":"2021-10-18T05:43:04.769392Z","iopub.status.idle":"2021-10-18T05:43:04.776539Z","shell.execute_reply.started":"2021-10-18T05:43:04.769344Z","shell.execute_reply":"2021-10-18T05:43:04.775614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"encoded_str = \"11 5 19 5 27 5 37 3\"","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:06.158297Z","iopub.execute_input":"2021-10-18T05:43:06.158872Z","iopub.status.idle":"2021-10-18T05:43:06.162558Z","shell.execute_reply.started":"2021-10-18T05:43:06.158836Z","shell.execute_reply":"2021-10-18T05:43:06.16186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Default reshape (row-by-row)\n- From notebook: https://www.kaggle.com/dschettler8845/sartorius-segmentation-eda-efficientdet-tf/notebook#helper_functions","metadata":{}},{"cell_type":"code","source":"def rle_decode(mask_rle, shape):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (height,width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape)  ","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:07.10388Z","iopub.execute_input":"2021-10-18T05:43:07.104477Z","iopub.status.idle":"2021-10-18T05:43:07.111206Z","shell.execute_reply.started":"2021-10-18T05:43:07.104436Z","shell.execute_reply":"2021-10-18T05:43:07.11045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = rle_decode(encoded_str, shape=(8,6))\nmask","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:08.01451Z","iopub.execute_input":"2021-10-18T05:43:08.015532Z","iopub.status.idle":"2021-10-18T05:43:08.02746Z","shell.execute_reply.started":"2021-10-18T05:43:08.015477Z","shell.execute_reply":"2021-10-18T05:43:08.026543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# chec if mask == arr2d\n(mask == arr2d).all()","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:20.550041Z","iopub.execute_input":"2021-10-18T05:43:20.550336Z","iopub.status.idle":"2021-10-18T05:43:20.555166Z","shell.execute_reply.started":"2021-10-18T05:43:20.550302Z","shell.execute_reply":"2021-10-18T05:43:20.554565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Proposed reshape (column-by-column):\n- I think it should be","metadata":{}},{"cell_type":"code","source":"def rle_decode_top_to_bot_first(mask_rle, shape):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (height,width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape, order='F')  # Reshape from top -> bottom first","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:31.502284Z","iopub.execute_input":"2021-10-18T05:43:31.502699Z","iopub.status.idle":"2021-10-18T05:43:31.510693Z","shell.execute_reply.started":"2021-10-18T05:43:31.502667Z","shell.execute_reply":"2021-10-18T05:43:31.510082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = rle_decode_top_to_bot_first(encoded_str, shape=(8,6))\nmask","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:31.693441Z","iopub.execute_input":"2021-10-18T05:43:31.693929Z","iopub.status.idle":"2021-10-18T05:43:31.699725Z","shell.execute_reply.started":"2021-10-18T05:43:31.693893Z","shell.execute_reply":"2021-10-18T05:43:31.698848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# chec if mask == arr2d\n(mask == arr2d).all()","metadata":{"execution":{"iopub.status.busy":"2021-10-18T05:43:31.832434Z","iopub.execute_input":"2021-10-18T05:43:31.833285Z","iopub.status.idle":"2021-10-18T05:43:31.839055Z","shell.execute_reply.started":"2021-10-18T05:43:31.833245Z","shell.execute_reply":"2021-10-18T05:43:31.838181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}