{
  "id": 222132,
  "title": "Even Faster Cell Segmentation",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/222132",
  "author_name": "Raman",
  "post_date": "2021-02-25T13:01:40.028000",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p><a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a> has shared an <a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">awesome notebook showing how to speed up the cell segmentation</a>. </p>\n<p>On top of it, I've noticed additional room for improvement and reduced the <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>'s faster segmentation execution time by 40-50%.</p>\n<p>The implementation expects images scaled to 0-1 range so that I can reuse the output of my data loaders directly.</p>\n<p>The details on code modifications can be found in <a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation/commits/master\" target=\"_blank\">the GitHub fork</a>.</p>\n<p>I've also created a <a href=\"https://www.kaggle.com/samusram/hpacellsegmentatorraman\" target=\"_blank\">dataset with the optimized implementation</a> so that it can be leveraged by anyone interested during submissions (building on top of the <a href=\"https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet\" target=\"_blank\">amazing work</a> by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> ).</p>\n<p>The optimized fork does not support padding, as it wasn't used anyways, and it expects the image size instead of the scaling factor (512 by default, as it seems to work the best). The CellSegmentator can be instantiated as follows</p>\n<pre><code>segmentator_even_faster = cellsegmentator.CellSegmentator(\n    NUC_MODEL,\n    CELL_MODEL,\n    device=\"cuda\",\n    multi_channel_model=True,\n)\n</code></pre>\n<p>.</p>\n<p>The comparison with <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>'s execution time as well as the example of usage can be found in <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">this notebook</a>.</p>\n<p>Hopefully, the shared faster segmentation would speed up someone's data preprocessing, i.e., public data labeling, or it hopefully might buy some time for someone's awesome larger models during submission.</p>",
  "messages": [
    {
      "id": 1217965,
      "postDate": "2021-02-25T13:01:40.030Z",
      "content": "<p>Hi everyone,</p>\n<p><a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a> has shared an <a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">awesome notebook showing how to speed up the cell segmentation</a>. </p>\n<p>On top of it, I've noticed additional room for improvement and reduced the <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>'s faster segmentation execution time by 40-50%.</p>\n<p>The implementation expects images scaled to 0-1 range so that I can reuse the output of my data loaders directly.</p>\n<p>The details on code modifications can be found in <a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation/commits/master\" target=\"_blank\">the GitHub fork</a>.</p>\n<p>I've also created a <a href=\"https://www.kaggle.com/samusram/hpacellsegmentatorraman\" target=\"_blank\">dataset with the optimized implementation</a> so that it can be leveraged by anyone interested during submissions (building on top of the <a href=\"https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet\" target=\"_blank\">amazing work</a> by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> ).</p>\n<p>The optimized fork does not support padding, as it wasn't used anyways, and it expects the image size instead of the scaling factor (512 by default, as it seems to work the best). The CellSegmentator can be instantiated as follows</p>\n<pre><code>segmentator_even_faster = cellsegmentator.CellSegmentator(\n    NUC_MODEL,\n    CELL_MODEL,\n    device=\"cuda\",\n    multi_channel_model=True,\n)\n</code></pre>\n<p>.</p>\n<p>The comparison with <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>'s execution time as well as the example of usage can be found in <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">this notebook</a>.</p>\n<p>Hopefully, the shared faster segmentation would speed up someone's data preprocessing, i.e., public data labeling, or it hopefully might buy some time for someone's awesome larger models during submission.</p>",
      "rawMarkdown": "Hi everyone,\n\n@linshokaku has shared an [awesome notebook showing how to speed up the cell segmentation](https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation). \n\nOn top of it, I've noticed additional room for improvement and reduced the @linshokaku's faster segmentation execution time by 40-50%.\n\nThe implementation expects images scaled to 0-1 range so that I can reuse the output of my data loaders directly.\n\nThe details on code modifications can be found in [the GitHub fork](https://github.com/SamusRam/HPA-Cell-Segmentation/commits/master).\n\nI've also created a [dataset with the optimized implementation](https://www.kaggle.com/samusram/hpacellsegmentatorraman) so that it can be leveraged by anyone interested during submissions (building on top of the [amazing work](https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet) by @rdizzl3 ).\n\nThe optimized fork does not support padding, as it wasn't used anyways, and it expects the image size instead of the scaling factor (512 by default, as it seems to work the best). The CellSegmentator can be instantiated as follows\n```\nsegmentator_even_faster = cellsegmentator.CellSegmentator(\n    NUC_MODEL,\n    CELL_MODEL,\n    device=\"cuda\",\n    multi_channel_model=True,\n)\n```.\n\nThe comparison with @linshokaku's execution time as well as the example of usage can be found in [this notebook](https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation).\n\nHopefully, the shared faster segmentation would speed up someone's data preprocessing, i.e., public data labeling, or it hopefully might buy some time for someone's awesome larger models during submission.",
      "votes": 14
    },
    {
      "id": 1284557,
      "postDate": "2021-04-26T04:51:54.160Z",
      "content": "<p>I tried and checked that it is much faster than original. However, I found that it can't be used to generate encode string for cell segmentation. <br>\nHave you found some other solution to address encoding cell segmentation problem if padding = False?</p>",
      "rawMarkdown": "I tried and checked that it is much faster than original. However, I found that it can't be used to generate encode string for cell segmentation. \nHave you found some other solution to address encoding cell segmentation problem if padding = False?",
      "replies": [
        {
          "id": 1284745,
          "postDate": "2021-04-26T08:33:55.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/mainguyenanhvu\" target=\"_blank\">@mainguyenanhvu</a> , could you please clarify what issue you are facing? Perhaps, I'd be able to help once I understand what you struggle with.</p>\n<p>Otherwise, the segmentation itself provides masks just as the initial segmentation code does. Please, see <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\"><strong>this notebook</strong></a> for example usage.</p>\n<p>In <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training\" target=\"_blank\">his awesome training notebook</a>, Darien ( <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> )  <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training#1225433\" target=\"_blank\"><strong>comments</strong></a> that he has successfully substituted the original segmentation with my slightly modified version, and I've noticed others using the modified version successfully as well (including myself :)). So I believe it's not the case that</p>\n<blockquote>\n  <p>it can't be used to generate encode string for cell segmentation.</p>\n</blockquote>\n<p>Once you obtain <code>mask</code> array according to <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">the example notebook</a>, you basically should do something like</p>\n<pre><code>def encode_binary_mask(mask: np.ndarray) -&gt; t.Text:\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n\n    # check input mask --\n    if mask.dtype != np.bool:\n        raise ValueError(\n            \"encode_binary_mask expects a binary mask, received dtype == %s\" %\n            mask.dtype)\n\n    mask = np.squeeze(mask)\n    if len(mask.shape) != 2:\n        raise ValueError(\n            \"encode_binary_mask expects a 2d mask, received shape == %s\" %\n            mask.shape)\n\n    # convert input mask to expected COCO API input --\n    mask_to_encode = mask.reshape(mask.shape[0], mask.shape[1], 1)\n    mask_to_encode = mask_to_encode.astype(np.uint8)\n    mask_to_encode = np.asfortranarray(mask_to_encode)\n\n    # RLE encode mask --\n    encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n\n    # compress and base64 encoding --\n    binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str.decode()\n\n\ncell_mask_bool = mask == cell_i # cell_i starts from 1 \nmask_rle = encode_binary_mask(cell_mask_bool)\n</code></pre>",
          "rawMarkdown": "@mainguyenanhvu , could you please clarify what issue you are facing? Perhaps, I'd be able to help once I understand what you struggle with.\n\nOtherwise, the segmentation itself provides masks just as the initial segmentation code does. Please, see [**this notebook**](https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation) for example usage.\n\nIn [his awesome training notebook](https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training), Darien ( @dschettler8845 )  [**comments**](https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training#1225433) that he has successfully substituted the original segmentation with my slightly modified version, and I've noticed others using the modified version successfully as well (including myself :)). So I believe it's not the case that\n> it can't be used to generate encode string for cell segmentation.\n\nOnce you obtain `mask` array according to [the example notebook](https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation), you basically should do something like\n\n```\ndef encode_binary_mask(mask: np.ndarray) -> t.Text:\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n\n    # check input mask --\n    if mask.dtype != np.bool:\n        raise ValueError(\n            \"encode_binary_mask expects a binary mask, received dtype == %s\" %\n            mask.dtype)\n\n    mask = np.squeeze(mask)\n    if len(mask.shape) != 2:\n        raise ValueError(\n            \"encode_binary_mask expects a 2d mask, received shape == %s\" %\n            mask.shape)\n\n    # convert input mask to expected COCO API input --\n    mask_to_encode = mask.reshape(mask.shape[0], mask.shape[1], 1)\n    mask_to_encode = mask_to_encode.astype(np.uint8)\n    mask_to_encode = np.asfortranarray(mask_to_encode)\n\n    # RLE encode mask --\n    encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n\n    # compress and base64 encoding --\n    binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str.decode()\n\n\ncell_mask_bool = mask == cell_i # cell_i starts from 1 \nmask_rle = encode_binary_mask(cell_mask_bool)\n```",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1284557,
      "author_name": "Anh-Vu Mai-Nguyen",
      "author_url": "",
      "post_date": "2021-04-26T04:51:54.160000",
      "content": "<p>I tried and checked that it is much faster than original. However, I found that it can't be used to generate encode string for cell segmentation. <br>\nHave you found some other solution to address encoding cell segmentation problem if padding = False?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1284745,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-04-26T08:33:55.637000",
          "content": "<p><a href=\"https://www.kaggle.com/mainguyenanhvu\" target=\"_blank\">@mainguyenanhvu</a> , could you please clarify what issue you are facing? Perhaps, I'd be able to help once I understand what you struggle with.</p>\n<p>Otherwise, the segmentation itself provides masks just as the initial segmentation code does. Please, see <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\"><strong>this notebook</strong></a> for example usage.</p>\n<p>In <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training\" target=\"_blank\">his awesome training notebook</a>, Darien ( <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> )  <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training#1225433\" target=\"_blank\"><strong>comments</strong></a> that he has successfully substituted the original segmentation with my slightly modified version, and I've noticed others using the modified version successfully as well (including myself :)). So I believe it's not the case that</p>\n<blockquote>\n  <p>it can't be used to generate encode string for cell segmentation.</p>\n</blockquote>\n<p>Once you obtain <code>mask</code> array according to <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">the example notebook</a>, you basically should do something like</p>\n<pre><code>def encode_binary_mask(mask: np.ndarray) -&gt; t.Text:\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n\n    # check input mask --\n    if mask.dtype != np.bool:\n        raise ValueError(\n            \"encode_binary_mask expects a binary mask, received dtype == %s\" %\n            mask.dtype)\n\n    mask = np.squeeze(mask)\n    if len(mask.shape) != 2:\n        raise ValueError(\n            \"encode_binary_mask expects a 2d mask, received shape == %s\" %\n            mask.shape)\n\n    # convert input mask to expected COCO API input --\n    mask_to_encode = mask.reshape(mask.shape[0], mask.shape[1], 1)\n    mask_to_encode = mask_to_encode.astype(np.uint8)\n    mask_to_encode = np.asfortranarray(mask_to_encode)\n\n    # RLE encode mask --\n    encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n\n    # compress and base64 encoding --\n    binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str.decode()\n\n\ncell_mask_bool = mask == cell_i # cell_i starts from 1 \nmask_rle = encode_binary_mask(cell_mask_bool)\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1217965": "Hi everyone,\n\n@linshokaku has shared an [awesome notebook showing how to speed up the cell segmentation](https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation). \n\nOn top of it, I've noticed additional room for improvement and reduced the @linshokaku's faster segmentation execution time by 40-50%.\n\nThe implementation expects images scaled to 0-1 range so that I can reuse the output of my data loaders directly.\n\nThe details on code modifications can be found in [the GitHub fork](https://github.com/SamusRam/HPA-Cell-Segmentation/commits/master).\n\nI've also created a [dataset with the optimized implementation](https://www.kaggle.com/samusram/hpacellsegmentatorraman) so that it can be leveraged by anyone interested during submissions (building on top of the [amazing work](https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet) by @rdizzl3 ).\n\nThe optimized fork does not support padding, as it wasn't used anyways, and it expects the image size instead of the scaling factor (512 by default, as it seems to work the best). The CellSegmentator can be instantiated as follows\n```\nsegmentator_even_faster = cellsegmentator.CellSegmentator(\n    NUC_MODEL,\n    CELL_MODEL,\n    device=\"cuda\",\n    multi_channel_model=True,\n)\n```.\n\nThe comparison with @linshokaku's execution time as well as the example of usage can be found in [this notebook](https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation).\n\nHopefully, the shared faster segmentation would speed up someone's data preprocessing, i.e., public data labeling, or it hopefully might buy some time for someone's awesome larger models during submission.",
    "1284557": "I tried and checked that it is much faster than original. However, I found that it can't be used to generate encode string for cell segmentation. \nHave you found some other solution to address encoding cell segmentation problem if padding = False?"
  }
}