{
  "id": 215351,
  "title": "How to Speed Up HPA Cell Segmentation?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/215351",
  "author_name": "Darien Schettler",
  "post_date": "2021-01-29T14:06:10.232000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all, I want to do cell segmentation using the <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\"><strong>HPACellSegmentor Tool</strong></a>.</p>\n<p>It is running incredibly slowly though (1.5 minutes for 8 images - 60+ hours required to do the training data if I let it run). I'm using a GPU and a batch size of 8.  </p>\n<p>I have included my code below. Can anyone offer any insights into how to make this code faster? Has anyone used this code to generate the masks already and created a dataset using this tool? </p>\n<pre><code>def create_segmentation_maps(img_dir, batch_size=8):\n    all_mask_rles = {}\n\n    # [[micro-tubules(red)], [endoplasmic-reticulum(yellow)], [nucleus(blue)]]\n    images = [sorted(glob(img_dir + '/' + f'*_{c}.png')) for c in [\"red\", \"yellow\", \"blue\"]]\n\n    # Create segmentor model\n    segmentator = cellsegmentator.CellSegmentator()\n\n    for i in tqdm(range(0, len(images[0]), batch_size), total=len(images[0])//batch_size):\n\n        # Get batch of images\n        sub_images = [img_channel_list[i:i+batch_size] for img_channel_list in images]\n        # Do segmentation\n        cell_segmentations = segmentator.pred_cells(sub_images)\n        nuc_segmentations = segmentator.pred_nuclei(sub_images[2])\n\n        # post-processing\n        for i, path in enumerate(sub_images[0]):\n            img_id = path.replace(\"_red.png\", \"\").rsplit(\"/\", 1)[1]\n            nuc_mask, cell_mask = label_cell(nuc_segmentations[i], cell_segmentations[i])\n            new_name = os.path.basename(path).replace('red','mask')\n            all_mask_rles[img_id] = [rle_encoding(cell_mask, mask_val=i) for i in range(1, np.max(cell_mask)+1)]\n    return all_mask_rles\n\ndef rle_encoding(img, mask_val=1):\n    \"\"\"\n    Turns our masks into RLE encoding to easily store them\n    and feed them into models later on\n    https://en.wikipedia.org/wiki/Run-length_encoding\n    \"\"\"\n    dots = np.where(img.T.flatten() == mask_val)[0]\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if (b&gt;prev+1): run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n\n    return ' '.join([str(x) for x in run_lengths])\n\ndef rle_to_mask(rle_string,height,width):\n    rows,cols = height,width\n    rle_numbers = [int(num_string) for num_string in rle_string.split(' ')]\n    rle_pairs = np.array(rle_numbers).reshape(-1,2)\n    img = np.zeros(rows*cols,dtype=np.uint8)\n    for index,length in rle_pairs:\n        index -= 1\n        img[index:index+length] = 255\n    img = img.reshape(cols,rows)\n    img = img.T\n    return img\n\n\ntrain_masks = create_segmentation_maps(img_dir=TRAIN_IMG_DIR)\ntest_masks = create_segmentation_maps(img_dir=TEST_IMG_DIR)\n</code></pre>",
  "messages": [
    {
      "id": 1176109,
      "postDate": "2021-01-29T14:06:10.233Z",
      "content": "<p>Hi all, I want to do cell segmentation using the <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\"><strong>HPACellSegmentor Tool</strong></a>.</p>\n<p>It is running incredibly slowly though (1.5 minutes for 8 images - 60+ hours required to do the training data if I let it run). I'm using a GPU and a batch size of 8.  </p>\n<p>I have included my code below. Can anyone offer any insights into how to make this code faster? Has anyone used this code to generate the masks already and created a dataset using this tool? </p>\n<pre><code>def create_segmentation_maps(img_dir, batch_size=8):\n    all_mask_rles = {}\n\n    # [[micro-tubules(red)], [endoplasmic-reticulum(yellow)], [nucleus(blue)]]\n    images = [sorted(glob(img_dir + '/' + f'*_{c}.png')) for c in [\"red\", \"yellow\", \"blue\"]]\n\n    # Create segmentor model\n    segmentator = cellsegmentator.CellSegmentator()\n\n    for i in tqdm(range(0, len(images[0]), batch_size), total=len(images[0])//batch_size):\n\n        # Get batch of images\n        sub_images = [img_channel_list[i:i+batch_size] for img_channel_list in images]\n        # Do segmentation\n        cell_segmentations = segmentator.pred_cells(sub_images)\n        nuc_segmentations = segmentator.pred_nuclei(sub_images[2])\n\n        # post-processing\n        for i, path in enumerate(sub_images[0]):\n            img_id = path.replace(\"_red.png\", \"\").rsplit(\"/\", 1)[1]\n            nuc_mask, cell_mask = label_cell(nuc_segmentations[i], cell_segmentations[i])\n            new_name = os.path.basename(path).replace('red','mask')\n            all_mask_rles[img_id] = [rle_encoding(cell_mask, mask_val=i) for i in range(1, np.max(cell_mask)+1)]\n    return all_mask_rles\n\ndef rle_encoding(img, mask_val=1):\n    \"\"\"\n    Turns our masks into RLE encoding to easily store them\n    and feed them into models later on\n    https://en.wikipedia.org/wiki/Run-length_encoding\n    \"\"\"\n    dots = np.where(img.T.flatten() == mask_val)[0]\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if (b&gt;prev+1): run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n\n    return ' '.join([str(x) for x in run_lengths])\n\ndef rle_to_mask(rle_string,height,width):\n    rows,cols = height,width\n    rle_numbers = [int(num_string) for num_string in rle_string.split(' ')]\n    rle_pairs = np.array(rle_numbers).reshape(-1,2)\n    img = np.zeros(rows*cols,dtype=np.uint8)\n    for index,length in rle_pairs:\n        index -= 1\n        img[index:index+length] = 255\n    img = img.reshape(cols,rows)\n    img = img.T\n    return img\n\n\ntrain_masks = create_segmentation_maps(img_dir=TRAIN_IMG_DIR)\ntest_masks = create_segmentation_maps(img_dir=TEST_IMG_DIR)\n</code></pre>",
      "rawMarkdown": "Hi all, I want to do cell segmentation using the [**HPACellSegmentor Tool**](https://github.com/CellProfiling/HPA-Cell-Segmentation).\n\nIt is running incredibly slowly though (1.5 minutes for 8 images - 60+ hours required to do the training data if I let it run). I'm using a GPU and a batch size of 8.  \n\nI have included my code below. Can anyone offer any insights into how to make this code faster? Has anyone used this code to generate the masks already and created a dataset using this tool? \n\n```python\n\ndef create_segmentation_maps(img_dir, batch_size=8):\n    all_mask_rles = {}\n    \n    # [[micro-tubules(red)], [endoplasmic-reticulum(yellow)], [nucleus(blue)]]\n    images = [sorted(glob(img_dir + '/' + f'*_{c}.png')) for c in [\"red\", \"yellow\", \"blue\"]]\n    \n    # Create segmentor model\n    segmentator = cellsegmentator.CellSegmentator()\n    \n    for i in tqdm(range(0, len(images[0]), batch_size), total=len(images[0])//batch_size):\n        \n        # Get batch of images\n        sub_images = [img_channel_list[i:i+batch_size] for img_channel_list in images]\n        # Do segmentation\n        cell_segmentations = segmentator.pred_cells(sub_images)\n        nuc_segmentations = segmentator.pred_nuclei(sub_images[2])\n\n        # post-processing\n        for i, path in enumerate(sub_images[0]):\n            img_id = path.replace(\"_red.png\", \"\").rsplit(\"/\", 1)[1]\n            nuc_mask, cell_mask = label_cell(nuc_segmentations[i], cell_segmentations[i])\n            new_name = os.path.basename(path).replace('red','mask')\n            all_mask_rles[img_id] = [rle_encoding(cell_mask, mask_val=i) for i in range(1, np.max(cell_mask)+1)]\n    return all_mask_rles\n    \ndef rle_encoding(img, mask_val=1):\n    \"\"\"\n    Turns our masks into RLE encoding to easily store them\n    and feed them into models later on\n    https://en.wikipedia.org/wiki/Run-length_encoding\n    \"\"\"\n    dots = np.where(img.T.flatten() == mask_val)[0]\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if (b>prev+1): run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n        \n    return ' '.join([str(x) for x in run_lengths])\n\ndef rle_to_mask(rle_string,height,width):\n    rows,cols = height,width\n    rle_numbers = [int(num_string) for num_string in rle_string.split(' ')]\n    rle_pairs = np.array(rle_numbers).reshape(-1,2)\n    img = np.zeros(rows*cols,dtype=np.uint8)\n    for index,length in rle_pairs:\n        index -= 1\n        img[index:index+length] = 255\n    img = img.reshape(cols,rows)\n    img = img.T\n    return img\n\n    \ntrain_masks = create_segmentation_maps(img_dir=TRAIN_IMG_DIR)\ntest_masks = create_segmentation_maps(img_dir=TEST_IMG_DIR)\n```",
      "votes": 7
    },
    {
      "id": 1195964,
      "postDate": "2021-02-11T07:07:22.470Z",
      "content": "<p>I'm also facing this issue  HPACellSegmentor Tool is extremely slow 🤒🤒. When I try to use multiprocessing, it locks</p>",
      "rawMarkdown": "I'm also facing this issue  HPACellSegmentor Tool is extremely slow 🤒🤒. When I try to use multiprocessing, it locks",
      "votes": 1,
      "replies": [
        {
          "id": 1195997,
          "postDate": "2021-02-11T07:40:38.837Z",
          "content": "<p>I see that another kind competitor <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> has created a segmentation mask dataset here <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773</a></p>\n<p>I haven't checked it personally, but based on the number of upvotes, it would be useful. Good luck!</p>",
          "rawMarkdown": "I see that another kind competitor @its7171 has created a segmentation mask dataset here https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773\n\nI haven't checked it personally, but based on the number of upvotes, it would be useful. Good luck!"
        },
        {
          "id": 1196027,
          "postDate": "2021-02-11T07:57:52.560Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>  for the quick support. For training, it is ok to use this dataset. My problem is at inference when I submit the notebook. I have to build the masks for the LB dataset before generating a prediction which is very slow and makes my notebook timeout.</p>",
          "rawMarkdown": "Thanks, @lnhtrang  for the quick support. For training, it is ok to use this dataset. My problem is at inference when I submit the notebook. I have to build the masks for the LB dataset before generating a prediction which is very slow and makes my notebook timeout.",
          "votes": 1
        },
        {
          "id": 1284521,
          "postDate": "2021-04-26T03:45:05.893Z",
          "content": "<p>I faced the same issue as you did <a href=\"https://www.kaggle.com/tchaye59\" target=\"_blank\">@tchaye59</a>. Do you address it? Can you help me? Thank you very much.</p>",
          "rawMarkdown": "I faced the same issue as you did @tchaye59. Do you address it? Can you help me? Thank you very much.",
          "votes": 1
        },
        {
          "id": 1284567,
          "postDate": "2021-04-26T05:08:00.680Z",
          "content": "<p>I use the faster segmentation tool <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation</a></p>",
          "rawMarkdown": "I use the faster segmentation tool https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation"
        }
      ]
    },
    {
      "id": 1178239,
      "postDate": "2021-01-30T18:42:30.827Z",
      "content": "<p>I added multiprocessing to this and it's working a bit faster. I will attempt to do this over the weekend and upload the RLEs as a dataset on Monday.</p>",
      "rawMarkdown": "I added multiprocessing to this and it's working a bit faster. I will attempt to do this over the weekend and upload the RLEs as a dataset on Monday.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1195964,
      "author_name": "Jude TCHAYE",
      "author_url": "",
      "post_date": "2021-02-11T07:07:22.470000",
      "content": "<p>I'm also facing this issue  HPACellSegmentor Tool is extremely slow 🤒🤒. When I try to use multiprocessing, it locks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1195997,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-02-11T07:40:38.837000",
          "content": "<p>I see that another kind competitor <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> has created a segmentation mask dataset here <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773</a></p>\n<p>I haven't checked it personally, but based on the number of upvotes, it would be useful. Good luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1196027,
          "author_name": "Jude TCHAYE",
          "author_url": "",
          "post_date": "2021-02-11T07:57:52.560000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>  for the quick support. For training, it is ok to use this dataset. My problem is at inference when I submit the notebook. I have to build the masks for the LB dataset before generating a prediction which is very slow and makes my notebook timeout.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1284521,
          "author_name": "Anh-Vu Mai-Nguyen",
          "author_url": "",
          "post_date": "2021-04-26T03:45:05.893000",
          "content": "<p>I faced the same issue as you did <a href=\"https://www.kaggle.com/tchaye59\" target=\"_blank\">@tchaye59</a>. Do you address it? Can you help me? Thank you very much.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1284567,
          "author_name": "Jude TCHAYE",
          "author_url": "",
          "post_date": "2021-04-26T05:08:00.680000",
          "content": "<p>I use the faster segmentation tool <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\">https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1178239,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-01-30T18:42:30.827000",
      "content": "<p>I added multiprocessing to this and it's working a bit faster. I will attempt to do this over the weekend and upload the RLEs as a dataset on Monday.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1176109": "Hi all, I want to do cell segmentation using the [**HPACellSegmentor Tool**](https://github.com/CellProfiling/HPA-Cell-Segmentation).\n\nIt is running incredibly slowly though (1.5 minutes for 8 images - 60+ hours required to do the training data if I let it run). I'm using a GPU and a batch size of 8.  \n\nI have included my code below. Can anyone offer any insights into how to make this code faster? Has anyone used this code to generate the masks already and created a dataset using this tool? \n\n```python\n\ndef create_segmentation_maps(img_dir, batch_size=8):\n    all_mask_rles = {}\n    \n    # [[micro-tubules(red)], [endoplasmic-reticulum(yellow)], [nucleus(blue)]]\n    images = [sorted(glob(img_dir + '/' + f'*_{c}.png')) for c in [\"red\", \"yellow\", \"blue\"]]\n    \n    # Create segmentor model\n    segmentator = cellsegmentator.CellSegmentator()\n    \n    for i in tqdm(range(0, len(images[0]), batch_size), total=len(images[0])//batch_size):\n        \n        # Get batch of images\n        sub_images = [img_channel_list[i:i+batch_size] for img_channel_list in images]\n        # Do segmentation\n        cell_segmentations = segmentator.pred_cells(sub_images)\n        nuc_segmentations = segmentator.pred_nuclei(sub_images[2])\n\n        # post-processing\n        for i, path in enumerate(sub_images[0]):\n            img_id = path.replace(\"_red.png\", \"\").rsplit(\"/\", 1)[1]\n            nuc_mask, cell_mask = label_cell(nuc_segmentations[i], cell_segmentations[i])\n            new_name = os.path.basename(path).replace('red','mask')\n            all_mask_rles[img_id] = [rle_encoding(cell_mask, mask_val=i) for i in range(1, np.max(cell_mask)+1)]\n    return all_mask_rles\n    \ndef rle_encoding(img, mask_val=1):\n    \"\"\"\n    Turns our masks into RLE encoding to easily store them\n    and feed them into models later on\n    https://en.wikipedia.org/wiki/Run-length_encoding\n    \"\"\"\n    dots = np.where(img.T.flatten() == mask_val)[0]\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if (b>prev+1): run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n        \n    return ' '.join([str(x) for x in run_lengths])\n\ndef rle_to_mask(rle_string,height,width):\n    rows,cols = height,width\n    rle_numbers = [int(num_string) for num_string in rle_string.split(' ')]\n    rle_pairs = np.array(rle_numbers).reshape(-1,2)\n    img = np.zeros(rows*cols,dtype=np.uint8)\n    for index,length in rle_pairs:\n        index -= 1\n        img[index:index+length] = 255\n    img = img.reshape(cols,rows)\n    img = img.T\n    return img\n\n    \ntrain_masks = create_segmentation_maps(img_dir=TRAIN_IMG_DIR)\ntest_masks = create_segmentation_maps(img_dir=TEST_IMG_DIR)\n```",
    "1195964": "I'm also facing this issue  HPACellSegmentor Tool is extremely slow 🤒🤒. When I try to use multiprocessing, it locks",
    "1178239": "I added multiprocessing to this and it's working a bit faster. I will attempt to do this over the weekend and upload the RLEs as a dataset on Monday."
  }
}