{
  "id": 288376,
  "title": "Can someone tell me how to generate the right format of submission?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/288376",
  "author_name": "Hey Guan",
  "post_date": "2021-11-17T11:07:11.604000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I  am confusing about how to save our predicted masks into csv files as the competition required. I knew how to encode rle and I got my predicted masks for a single image with shape of 1<em>520</em>704. If I directly encode a mask to RLE, there would be only one row containing all information for this mask. But I see competition required us that each row in our submission represents a single predicted nucleus segmentation for the given&nbsp;Image_Id. How can I split my masks to be many nucleus segmentation? Many thanks!</p>",
  "messages": [
    {
      "id": 1585871,
      "postDate": "2021-11-17T16:10:06.960Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/heyguan\" target=\"_blank\">@heyguan</a> please check out this block of code, it uses <code>cv2.connectedComponents</code> for separating the cell bodies.</p>\n<pre><code>def post_process(probability, threshold=0.5, min_size=300):\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = []\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            a_prediction = np.zeros((520, 704), np.float32)\n            a_prediction[p] = 1\n            predictions.append(a_prediction)\n    return predictions\n</code></pre>\n<p>It returns a list of images which is consists of every predicted cell body of that image, and after that, you can loop through each cell from the list and encode them into RLE.  For reference you can see the TTA part of this NB <a href=\"https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA\" target=\"_blank\">https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA</a></p>",
      "rawMarkdown": "Hey @heyguan please check out this block of code, it uses `cv2.connectedComponents` for separating the cell bodies.\n\n```python\ndef post_process(probability, threshold=0.5, min_size=300):\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = []\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() > min_size:\n            a_prediction = np.zeros((520, 704), np.float32)\n            a_prediction[p] = 1\n            predictions.append(a_prediction)\n    return predictions\n```\n\nIt returns a list of images which is consists of every predicted cell body of that image, and after that, you can loop through each cell from the list and encode them into RLE.  For reference you can see the TTA part of this NB https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA",
      "votes": 5,
      "replies": [
        {
          "id": 1585929,
          "postDate": "2021-11-17T16:51:08.780Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1585583,
      "postDate": "2021-11-17T11:07:11.603Z",
      "content": "<p>Hi, I  am confusing about how to save our predicted masks into csv files as the competition required. I knew how to encode rle and I got my predicted masks for a single image with shape of 1<em>520</em>704. If I directly encode a mask to RLE, there would be only one row containing all information for this mask. But I see competition required us that each row in our submission represents a single predicted nucleus segmentation for the given&nbsp;Image_Id. How can I split my masks to be many nucleus segmentation? Many thanks!</p>",
      "rawMarkdown": "Hi, I  am confusing about how to save our predicted masks into csv files as the competition required. I knew how to encode rle and I got my predicted masks for a single image with shape of 1*520*704. If I directly encode a mask to RLE, there would be only one row containing all information for this mask. But I see competition required us that each row in our submission represents a single predicted nucleus segmentation for the given Image_Id. How can I split my masks to be many nucleus segmentation? Many thanks!",
      "votes": 1
    },
    {
      "id": 1600368,
      "postDate": "2021-11-30T10:39:39.720Z",
      "content": "<p><a href=\"https://www.kaggle.com/Hey\" target=\"_blank\">@Hey</a> Guan, Could you help me with this issue by sharing the code of this part?</p>",
      "rawMarkdown": "@Hey Guan, Could you help me with this issue by sharing the code of this part?\n",
      "replies": [
        {
          "id": 1602089,
          "postDate": "2021-12-01T18:58:11.623Z",
          "content": "<p>Sorry I still have no idea</p>",
          "rawMarkdown": "Sorry I still have no idea"
        }
      ]
    },
    {
      "id": 1585928,
      "postDate": "2021-11-17T16:50:28.520Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1585871,
      "author_name": "somuSan",
      "author_url": "",
      "post_date": "2021-11-17T16:10:06.960000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/heyguan\" target=\"_blank\">@heyguan</a> please check out this block of code, it uses <code>cv2.connectedComponents</code> for separating the cell bodies.</p>\n<pre><code>def post_process(probability, threshold=0.5, min_size=300):\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = []\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() &gt; min_size:\n            a_prediction = np.zeros((520, 704), np.float32)\n            a_prediction[p] = 1\n            predictions.append(a_prediction)\n    return predictions\n</code></pre>\n<p>It returns a list of images which is consists of every predicted cell body of that image, and after that, you can loop through each cell from the list and encode them into RLE.  For reference you can see the TTA part of this NB <a href=\"https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA\" target=\"_blank\">https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1585929,
          "author_name": "Hey Guan",
          "author_url": "",
          "post_date": "2021-11-17T16:51:08.780000",
          "content": "<p>Thanks a lot!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1600368,
      "author_name": "Fatma Mazen",
      "author_url": "",
      "post_date": "2021-11-30T10:39:39.720000",
      "content": "<p><a href=\"https://www.kaggle.com/Hey\" target=\"_blank\">@Hey</a> Guan, Could you help me with this issue by sharing the code of this part?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1602089,
          "author_name": "Hey Guan",
          "author_url": "",
          "post_date": "2021-12-01T18:58:11.623000",
          "content": "<p>Sorry I still have no idea</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1585928,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-17T16:50:28.520000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1585871": "Hey @heyguan please check out this block of code, it uses `cv2.connectedComponents` for separating the cell bodies.\n\n```python\ndef post_process(probability, threshold=0.5, min_size=300):\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = []\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() > min_size:\n            a_prediction = np.zeros((520, 704), np.float32)\n            a_prediction[p] = 1\n            predictions.append(a_prediction)\n    return predictions\n```\n\nIt returns a list of images which is consists of every predicted cell body of that image, and after that, you can loop through each cell from the list and encode them into RLE.  For reference you can see the TTA part of this NB https://www.kaggle.com/soumya9977/residual-unet-with-attention-eda-tta-tf-data/notebook#Basic-TTA",
    "1585583": "Hi, I  am confusing about how to save our predicted masks into csv files as the competition required. I knew how to encode rle and I got my predicted masks for a single image with shape of 1*520*704. If I directly encode a mask to RLE, there would be only one row containing all information for this mask. But I see competition required us that each row in our submission represents a single predicted nucleus segmentation for the given Image_Id. How can I split my masks to be many nucleus segmentation? Many thanks!",
    "1600368": "@Hey Guan, Could you help me with this issue by sharing the code of this part?\n",
    "1585928": ""
  }
}