{
  "id": 298090,
  "title": "Who else has tried U-Net in this competion? ",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/298090",
  "author_name": "",
  "post_date": "2021-12-31T16:04:52.093257100Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey,</p>\n<p>first of all congrats to all winners, and to those who had fun, gained knowledge and just had a good time with this competition.</p>\n<p>i just want to share my insights in this competition. This was my first project in instance segmentation - because I had a good access to the idea of U-Net i did the Unet-Watershed approach. Because I couldn't find any high scoring Unets (correct me please if I am wrong) I want to explain my try. Therefore I used Keras/ TF.</p>\n<p>All in all I got LB scores of about 0.11 - not that much I know:)</p>\n<p>What I tried:</p>\n<ul>\n<li><p>resizing to (256, 256) - for all cell types and also one modell for every cell type</p></li>\n<li><p>resizing to (512, 512) - for all cell types and also one modell for every cell type</p></li>\n<li><p>tiling images (resized original images to (520,705) and then generated patches of (260, 235) in <br>\norder to use the full image/mask. I resized those tiles then to (256, 256) and used them for training.  <br>\nWhile submission, I resized the mask also to (520, 705), and did the whole process again, and then <br>\nrebuild the single tiles back to the original mask - for all cell types and also one modell for every cell <br>\ntype.</p></li>\n</ul>\n<p>I trained first with binary crossentropy about 15 epochs and then changed to 30 - 40 epochs of dice loss - both with Adam and its default settings, which got me to a dice coef meitrc of about 0.75-0.8. I just these tow stages of training just because it lead to the best results.</p>\n<p>I also used different watershed settings for the three cell types - which I derived manually with some examples for each cell type.</p>\n<p>I roughly had pretty much the same augmentation settings (ImageDataGenerator):</p>\n<ul>\n<li>rotation_range = 10 (0 - 10)</li>\n<li>width_shift_range = 0.25 (0. - 0.25)</li>\n<li>height_shift_range = 0.25 (0. - 0.25)</li>\n<li>shear_range = 0.1 (0. - 0.1)</li>\n<li>zoom_range = 0.1 (0. - 0.1)</li>\n<li>fill_mode = 'reflect'</li>\n<li>horizontal_flip = True</li>\n<li>vertical_flip = True</li>\n<li>drop_rate = 0.1 (0.05-0.-125)<br>\nI also tried to evaluate different thresholds, but couln't see any differences.</li>\n</ul>\n<p>What were your experiences? Was it the wrong approch?</p>\n<p>I wish you all a good start for 2022 with hopefully lots of exiting competition?</p>\n<p>Cheers, Andreas</p>",
  "messages": [
    {
      "id": "1634341",
      "postDate": "12/31/2021 16:04:52",
      "content": "<p>Hey,</p>\n<p>first of all congrats to all winners, and to those who had fun, gained knowledge and just had a good time with this competition.</p>\n<p>i just want to share my insights in this competition. This was my first project in instance segmentation - because I had a good access to the idea of U-Net i did the Unet-Watershed approach. Because I couldn't find any high scoring Unets (correct me please if I am wrong) I want to explain my try. Therefore I used Keras/ TF.</p>\n<p>All in all I got LB scores of about 0.11 - not that much I know:)</p>\n<p>What I tried:</p>\n<ul>\n<li><p>resizing to (256, 256) - for all cell types and also one modell for every cell type</p></li>\n<li><p>resizing to (512, 512) - for all cell types and also one modell for every cell type</p></li>\n<li><p>tiling images (resized original images to (520,705) and then generated patches of (260, 235) in <br>\norder to use the full image/mask. I resized those tiles then to (256, 256) and used them for training.  <br>\nWhile submission, I resized the mask also to (520, 705), and did the whole process again, and then <br>\nrebuild the single tiles back to the original mask - for all cell types and also one modell for every cell <br>\ntype.</p></li>\n</ul>\n<p>I trained first with binary crossentropy about 15 epochs and then changed to 30 - 40 epochs of dice loss - both with Adam and its default settings, which got me to a dice coef meitrc of about 0.75-0.8. I just these tow stages of training just because it lead to the best results.</p>\n<p>I also used different watershed settings for the three cell types - which I derived manually with some examples for each cell type.</p>\n<p>I roughly had pretty much the same augmentation settings (ImageDataGenerator):</p>\n<ul>\n<li>rotation_range = 10 (0 - 10)</li>\n<li>width_shift_range = 0.25 (0. - 0.25)</li>\n<li>height_shift_range = 0.25 (0. - 0.25)</li>\n<li>shear_range = 0.1 (0. - 0.1)</li>\n<li>zoom_range = 0.1 (0. - 0.1)</li>\n<li>fill_mode = 'reflect'</li>\n<li>horizontal_flip = True</li>\n<li>vertical_flip = True</li>\n<li>drop_rate = 0.1 (0.05-0.-125)<br>\nI also tried to evaluate different thresholds, but couln't see any differences.</li>\n</ul>\n<p>What were your experiences? Was it the wrong approch?</p>\n<p>I wish you all a good start for 2022 with hopefully lots of exiting competition?</p>\n<p>Cheers, Andreas</p>",
      "rawMarkdown": "Hey,\n\nfirst of all congrats to all winners, and to those who had fun, gained knowledge and just had a good time with this competition.\n\ni just want to share my insights in this competition. This was my first project in instance segmentation - because I had a good access to the idea of U-Net i did the Unet-Watershed approach. Because I couldn't find any high scoring Unets (correct me please if I am wrong) I want to explain my try. Therefore I used Keras/ TF.\n\nAll in all I got LB scores of about 0.11 - not that much I know:)\n\nWhat I tried:\n\n- resizing to (256, 256) - for all cell types and also one modell for every cell type\n\n- resizing to (512, 512) - for all cell types and also one modell for every cell type\n\n- tiling images (resized original images to (520,705) and then generated patches of (260, 235) in \n order to use the full image/mask. I resized those tiles then to (256, 256) and used them for training.  \n While submission, I resized the mask also to (520, 705), and did the whole process again, and then \n rebuild the single tiles back to the original mask - for all cell types and also one modell for every cell \n type.\n\nI trained first with binary crossentropy about 15 epochs and then changed to 30 - 40 epochs of dice loss - both with Adam and its default settings, which got me to a dice coef meitrc of about 0.75-0.8. I just these tow stages of training just because it lead to the best results.\n\nI also used different watershed settings for the three cell types - which I derived manually with some examples for each cell type.\n\nI roughly had pretty much the same augmentation settings (ImageDataGenerator):\n\n- rotation_range = 10 (0 - 10)\n- width_shift_range = 0.25 (0. - 0.25)\n- height_shift_range = 0.25 (0. - 0.25)\n- shear_range = 0.1 (0. - 0.1)\n- zoom_range = 0.1 (0. - 0.1)\n- fill_mode = 'reflect'\n- horizontal_flip = True\n- vertical_flip = True\n- drop_rate = 0.1 (0.05-0.-125)\nI also tried to evaluate different thresholds, but couln't see any differences.\n\nWhat were your experiences? Was it the wrong approch?\n\nI wish you all a good start for 2022 with hopefully lots of exiting competition?\n\nCheers, Andreas",
      "votes": null
    },
    {
      "id": "1635510",
      "postDate": "01/01/2022 20:12:24",
      "content": "<p>Cellpose is a unet model. Therefore many top teams have been using unet.</p>\n<p>I am not sure this is the answer you wanted, but it is nevertheless true.  You probably ask about unet with only segementation mask as target, instead of the spatial gradients used by cellpose.</p>",
      "rawMarkdown": "Cellpose is a unet model. Therefore many top teams have been using unet.\n\nI am not sure this is the answer you wanted, but it is nevertheless true.  You probably ask about unet with only segementation mask as target, instead of the spatial gradients used by cellpose.",
      "votes": null
    },
    {
      "id": "1638433",
      "postDate": "01/04/2022 18:07:18",
      "content": "<p>Hi!</p>\n<p>I have tried some submissions with U-net: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/carlosgut/sartorius-complete-from-eda-to-submit-unet</a>. At the end, I got similar results as you mention.</p>\n<p>I have mainly worked with image size 256x256 and I focused on TF/keras without using any external source or repository. The model was a custom U-net, trained for around 70 epochs, with adam optimizer, cross-binary entropy and dice metric.</p>\n<p>Things that I have experimented: color/b&amp;w, size, data augmentation, epochs, train-val split, threshold, data power enhacements…</p>\n<p>After all submissions, here it goes my conclusions (may or may not be accurate, but it is what I have seen):</p>\n<ul>\n<li>Having the subsmision file properly made is more important than any other approach: U-net returns the whole image, not the isolated points. Breaking them, remove the lonely pixels, correct minimun size of nucleus, boosted the score much more than any experimentation (perhaps this is only on lower scores). </li>\n<li>Color does not add any useful value</li>\n<li>Validation split of 12% improved the results rather than a bigger split (15 or 20%).</li>\n<li>With my training, slightly larger epochs improved the results. Going up to 70 epochs (with earlystopper and learning rate decrease), improved around 0.1 the LB score.</li>\n<li>Threshold was not helpful in general, but it does matter if you take into consideration cell type (astro type has wider area). I think many high score solutions use this.</li>\n</ul>\n<p>Those are some of things that I have experience. I am not an expert, just recently landed here :).</p>\n<p>I am also curious on any feedback, experiences or comments of this topic.<br>\nCheers!</p>",
      "rawMarkdown": "Hi!\n\nI have tried some submissions with U-net: [https://www.kaggle.com/carlosgut/sartorius-complete-from-eda-to-submit-unet](url). At the end, I got similar results as you mention.\n\nI have mainly worked with image size 256x256 and I focused on TF/keras without using any external source or repository. The model was a custom U-net, trained for around 70 epochs, with adam optimizer, cross-binary entropy and dice metric.\n\nThings that I have experimented: color/b&w, size, data augmentation, epochs, train-val split, threshold, data power enhacements...\n\nAfter all submissions, here it goes my conclusions (may or may not be accurate, but it is what I have seen):\n\n- Having the subsmision file properly made is more important than any other approach: U-net returns the whole image, not the isolated points. Breaking them, remove the lonely pixels, correct minimun size of nucleus, boosted the score much more than any experimentation (perhaps this is only on lower scores). \n- Color does not add any useful value\n- Validation split of 12% improved the results rather than a bigger split (15 or 20%).\n- With my training, slightly larger epochs improved the results. Going up to 70 epochs (with earlystopper and learning rate decrease), improved around 0.1 the LB score.\n- Threshold was not helpful in general, but it does matter if you take into consideration cell type (astro type has wider area). I think many high score solutions use this.\n\n\nThose are some of things that I have experience. I am not an expert, just recently landed here :).\n\n\nI am also curious on any feedback, experiences or comments of this topic.\nCheers!",
      "votes": null
    },
    {
      "id": "1639727",
      "postDate": "01/05/2022 21:17:35",
      "content": "<p>Thanks a lot for your answer :) </p>",
      "rawMarkdown": "Thanks a lot for your answer :)",
      "votes": null
    },
    {
      "id": "1639728",
      "postDate": "01/05/2022 21:18:00",
      "content": "<p>Okay, even this is new to me, thx that helps! </p>",
      "rawMarkdown": "Okay, even this is new to me, thx that helps!",
      "votes": null
    },
    {
      "id": "1641295",
      "postDate": "01/07/2022 09:20:10",
      "content": "<p>First, congratulations to the winners and thanks to the organizers for a challenging and interesting competition. For what it's worth, here is my experience with U-nets in this contest:</p>\n<p>I managed 0.211 on Private LB using a fairly simple U-net enhanced with a few home-grown tricks. It would have been 0.213 but my last submission finished scoring a few minutes after the deadline. For the 0.211 submission I ensembled 10 similar models trained on my own computers in R, Keras, and Tensorflow. Hardware included x86_64 workstations running Ubuntu Linux, and a single Nvidia Quadro GP100 GPU. Ensembling was done by averaging raw numeric model predictions over all the models before any post-processing. My Public and Private LB scores were somewhat higher than my validation scores but correlated well with them, and my best two submissions on Public LB were also best on Private LB.</p>\n<p>Although my U-net originally predicted only a single mask over all the cells, I added 4 more prediction outputs to the model to try to help with segmentation during post-processing: a mask over \"boundaries between cell perimeters\", a total count of cells in the image, the total fraction of the image contained in those cells, and probabilities of the 3 cell types. I didn't have any special method of dealing with the cell overlaps in the training set; each pixel was essentially randomly assigned to a single cell. To ensure its independence from the cell mask prediction, my boundary mask prediction came from a separate set of \"Up\" layers starting from the center of the U-net. Altogether, my net included 42 convolutional layers, spread among 4 \"down\", 1 center, and 8 \"up\" sections (4 for the mask, and 4 for the perimeter boundaries).</p>\n<p>The perimeter boundary mask for each training image was the union, over all individual cell masks for that image, of the results of dilating each mask slightly using a kernel size of 3 and computing the intersection (overlap) of that dilated mask with the union of all the other cell masks. The resulting binary mask consisted of very thin lines separating adjacent cell masks.</p>\n<p>For training, all images and masks were resized to 512x512. Rather than dividing training into epochs, instead I called train_on_batch for thousands of batches each containing 8 randomly-selected images, and decayed the learning rate when progress in reducing the loss stalled. Augmentations (training only) included flip, flop, transpose, and a type of non-linear warping that transformed images (and their masks) as if they were reflected by a curved mirror.</p>\n<p>Post-processing consisted of some optional \"temperature sharpening\" of the predicted mask pixel intensities, followed by searching for the \"best (lowest) mask score\" over several mask rounding thresholds starting at 0.002 and incremented by 0.002 until a trend toward worse scores was detected. For each candidate threshold, pixel intensities in the predicted mask that equaled or exceeded the threshold were rounded to 1 and the remainder set to 0. Then an optional \"opening\" (erosion followed by dilation) was performed on the binary mask. Then the predicted perimeter boundary mask was rounded to binary using a specified threshold value, and was used like a pair of scissors to cut thin lines through the predicted cell mask to help EBImage::bwlabel, which labels connected sets of pixels, to identify individual cells. The mask score was computed as a function of discrepancies: 1 - between the cell count predicted directly by the model and the count of objects produced from the predicted mask by bwlabel, and 2 - between the fraction of the image's pixels contained in cells predicted directly by the model and the fraction computed from the predicted mask. Each of those differences was raised to a specified power and the result multiplied by a specified factor, then the two difference terms were added together to produce the total mask score.</p>\n<p>After the best mask was determined, those predicted cell masks whose area (pixel count) was less than a specified minimum were deleted.</p>\n<p>A bunch of hyperparameters, most of which were 3-element vectors indexed by the highest probability cell type prediction, controlled the post-processing, and I did some automatic tuning to try to find good values. With this dataset that sounds like a recipe for overfitting, but at least to the limited extent I did it (it was very time intensive) it seemed to improve performance both in my validation runs and on the LB.</p>\n<p>I only hit on the \"cell perimeter boundary\" idea late in the competition, and my results suggest that it may be possible to improve its performance considerably, although probably not nearly enough to be competitive with Detectron2, Mask R-cnn, etc.</p>\n<p>So, what do readers think: does the above description of my program suggest any useful ideas, or is it just a cautionary tale of what methods to avoid when approaching an object detection problem like this contest?</p>",
      "rawMarkdown": "First, congratulations to the winners and thanks to the organizers for a challenging and interesting competition. For what it's worth, here is my experience with U-nets in this contest:\n\nI managed 0.211 on Private LB using a fairly simple U-net enhanced with a few home-grown tricks. It would have been 0.213 but my last submission finished scoring a few minutes after the deadline. For the 0.211 submission I ensembled 10 similar models trained on my own computers in R, Keras, and Tensorflow. Hardware included x86_64 workstations running Ubuntu Linux, and a single Nvidia Quadro GP100 GPU. Ensembling was done by averaging raw numeric model predictions over all the models before any post-processing. My Public and Private LB scores were somewhat higher than my validation scores but correlated well with them, and my best two submissions on Public LB were also best on Private LB.\n\nAlthough my U-net originally predicted only a single mask over all the cells, I added 4 more prediction outputs to the model to try to help with segmentation during post-processing: a mask over \"boundaries between cell perimeters\", a total count of cells in the image, the total fraction of the image contained in those cells, and probabilities of the 3 cell types. I didn't have any special method of dealing with the cell overlaps in the training set; each pixel was essentially randomly assigned to a single cell. To ensure its independence from the cell mask prediction, my boundary mask prediction came from a separate set of \"Up\" layers starting from the center of the U-net. Altogether, my net included 42 convolutional layers, spread among 4 \"down\", 1 center, and 8 \"up\" sections (4 for the mask, and 4 for the perimeter boundaries).\n\nThe perimeter boundary mask for each training image was the union, over all individual cell masks for that image, of the results of dilating each mask slightly using a kernel size of 3 and computing the intersection (overlap) of that dilated mask with the union of all the other cell masks. The resulting binary mask consisted of very thin lines separating adjacent cell masks.\n\nFor training, all images and masks were resized to 512x512. Rather than dividing training into epochs, instead I called train_on_batch for thousands of batches each containing 8 randomly-selected images, and decayed the learning rate when progress in reducing the loss stalled. Augmentations (training only) included flip, flop, transpose, and a type of non-linear warping that transformed images (and their masks) as if they were reflected by a curved mirror.\n\nPost-processing consisted of some optional \"temperature sharpening\" of the predicted mask pixel intensities, followed by searching for the \"best (lowest) mask score\" over several mask rounding thresholds starting at 0.002 and incremented by 0.002 until a trend toward worse scores was detected. For each candidate threshold, pixel intensities in the predicted mask that equaled or exceeded the threshold were rounded to 1 and the remainder set to 0. Then an optional \"opening\" (erosion followed by dilation) was performed on the binary mask. Then the predicted perimeter boundary mask was rounded to binary using a specified threshold value, and was used like a pair of scissors to cut thin lines through the predicted cell mask to help EBImage::bwlabel, which labels connected sets of pixels, to identify individual cells. The mask score was computed as a function of discrepancies: 1 - between the cell count predicted directly by the model and the count of objects produced from the predicted mask by bwlabel, and 2 - between the fraction of the image's pixels contained in cells predicted directly by the model and the fraction computed from the predicted mask. Each of those differences was raised to a specified power and the result multiplied by a specified factor, then the two difference terms were added together to produce the total mask score.\n\nAfter the best mask was determined, those predicted cell masks whose area (pixel count) was less than a specified minimum were deleted.\n\nA bunch of hyperparameters, most of which were 3-element vectors indexed by the highest probability cell type prediction, controlled the post-processing, and I did some automatic tuning to try to find good values. With this dataset that sounds like a recipe for overfitting, but at least to the limited extent I did it (it was very time intensive) it seemed to improve performance both in my validation runs and on the LB.\n\nI only hit on the \"cell perimeter boundary\" idea late in the competition, and my results suggest that it may be possible to improve its performance considerably, although probably not nearly enough to be competitive with Detectron2, Mask R-cnn, etc.\n\nSo, what do readers think: does the above description of my program suggest any useful ideas, or is it just a cautionary tale of what methods to avoid when approaching an object detection problem like this contest?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1635510,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/01/2022 20:12:24",
      "content": "<p>Cellpose is a unet model. Therefore many top teams have been using unet.</p>\n<p>I am not sure this is the answer you wanted, but it is nevertheless true.  You probably ask about unet with only segementation mask as target, instead of the spatial gradients used by cellpose.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1639728,
          "author_name": "andreashorlbeck",
          "author_url": "",
          "post_date": "01/05/2022 21:18:00",
          "content": "<p>Okay, even this is new to me, thx that helps! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1638433,
      "author_name": "carlosgut",
      "author_url": "",
      "post_date": "01/04/2022 18:07:18",
      "content": "<p>Hi!</p>\n<p>I have tried some submissions with U-net: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/carlosgut/sartorius-complete-from-eda-to-submit-unet</a>. At the end, I got similar results as you mention.</p>\n<p>I have mainly worked with image size 256x256 and I focused on TF/keras without using any external source or repository. The model was a custom U-net, trained for around 70 epochs, with adam optimizer, cross-binary entropy and dice metric.</p>\n<p>Things that I have experimented: color/b&amp;w, size, data augmentation, epochs, train-val split, threshold, data power enhacements…</p>\n<p>After all submissions, here it goes my conclusions (may or may not be accurate, but it is what I have seen):</p>\n<ul>\n<li>Having the subsmision file properly made is more important than any other approach: U-net returns the whole image, not the isolated points. Breaking them, remove the lonely pixels, correct minimun size of nucleus, boosted the score much more than any experimentation (perhaps this is only on lower scores). </li>\n<li>Color does not add any useful value</li>\n<li>Validation split of 12% improved the results rather than a bigger split (15 or 20%).</li>\n<li>With my training, slightly larger epochs improved the results. Going up to 70 epochs (with earlystopper and learning rate decrease), improved around 0.1 the LB score.</li>\n<li>Threshold was not helpful in general, but it does matter if you take into consideration cell type (astro type has wider area). I think many high score solutions use this.</li>\n</ul>\n<p>Those are some of things that I have experience. I am not an expert, just recently landed here :).</p>\n<p>I am also curious on any feedback, experiences or comments of this topic.<br>\nCheers!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1639727,
          "author_name": "andreashorlbeck",
          "author_url": "",
          "post_date": "01/05/2022 21:17:35",
          "content": "<p>Thanks a lot for your answer :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1641295,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "01/07/2022 09:20:10",
      "content": "<p>First, congratulations to the winners and thanks to the organizers for a challenging and interesting competition. For what it's worth, here is my experience with U-nets in this contest:</p>\n<p>I managed 0.211 on Private LB using a fairly simple U-net enhanced with a few home-grown tricks. It would have been 0.213 but my last submission finished scoring a few minutes after the deadline. For the 0.211 submission I ensembled 10 similar models trained on my own computers in R, Keras, and Tensorflow. Hardware included x86_64 workstations running Ubuntu Linux, and a single Nvidia Quadro GP100 GPU. Ensembling was done by averaging raw numeric model predictions over all the models before any post-processing. My Public and Private LB scores were somewhat higher than my validation scores but correlated well with them, and my best two submissions on Public LB were also best on Private LB.</p>\n<p>Although my U-net originally predicted only a single mask over all the cells, I added 4 more prediction outputs to the model to try to help with segmentation during post-processing: a mask over \"boundaries between cell perimeters\", a total count of cells in the image, the total fraction of the image contained in those cells, and probabilities of the 3 cell types. I didn't have any special method of dealing with the cell overlaps in the training set; each pixel was essentially randomly assigned to a single cell. To ensure its independence from the cell mask prediction, my boundary mask prediction came from a separate set of \"Up\" layers starting from the center of the U-net. Altogether, my net included 42 convolutional layers, spread among 4 \"down\", 1 center, and 8 \"up\" sections (4 for the mask, and 4 for the perimeter boundaries).</p>\n<p>The perimeter boundary mask for each training image was the union, over all individual cell masks for that image, of the results of dilating each mask slightly using a kernel size of 3 and computing the intersection (overlap) of that dilated mask with the union of all the other cell masks. The resulting binary mask consisted of very thin lines separating adjacent cell masks.</p>\n<p>For training, all images and masks were resized to 512x512. Rather than dividing training into epochs, instead I called train_on_batch for thousands of batches each containing 8 randomly-selected images, and decayed the learning rate when progress in reducing the loss stalled. Augmentations (training only) included flip, flop, transpose, and a type of non-linear warping that transformed images (and their masks) as if they were reflected by a curved mirror.</p>\n<p>Post-processing consisted of some optional \"temperature sharpening\" of the predicted mask pixel intensities, followed by searching for the \"best (lowest) mask score\" over several mask rounding thresholds starting at 0.002 and incremented by 0.002 until a trend toward worse scores was detected. For each candidate threshold, pixel intensities in the predicted mask that equaled or exceeded the threshold were rounded to 1 and the remainder set to 0. Then an optional \"opening\" (erosion followed by dilation) was performed on the binary mask. Then the predicted perimeter boundary mask was rounded to binary using a specified threshold value, and was used like a pair of scissors to cut thin lines through the predicted cell mask to help EBImage::bwlabel, which labels connected sets of pixels, to identify individual cells. The mask score was computed as a function of discrepancies: 1 - between the cell count predicted directly by the model and the count of objects produced from the predicted mask by bwlabel, and 2 - between the fraction of the image's pixels contained in cells predicted directly by the model and the fraction computed from the predicted mask. Each of those differences was raised to a specified power and the result multiplied by a specified factor, then the two difference terms were added together to produce the total mask score.</p>\n<p>After the best mask was determined, those predicted cell masks whose area (pixel count) was less than a specified minimum were deleted.</p>\n<p>A bunch of hyperparameters, most of which were 3-element vectors indexed by the highest probability cell type prediction, controlled the post-processing, and I did some automatic tuning to try to find good values. With this dataset that sounds like a recipe for overfitting, but at least to the limited extent I did it (it was very time intensive) it seemed to improve performance both in my validation runs and on the LB.</p>\n<p>I only hit on the \"cell perimeter boundary\" idea late in the competition, and my results suggest that it may be possible to improve its performance considerably, although probably not nearly enough to be competitive with Detectron2, Mask R-cnn, etc.</p>\n<p>So, what do readers think: does the above description of my program suggest any useful ideas, or is it just a cautionary tale of what methods to avoid when approaching an object detection problem like this contest?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1634341": "Hey,\n\nfirst of all congrats to all winners, and to those who had fun, gained knowledge and just had a good time with this competition.\n\ni just want to share my insights in this competition. This was my first project in instance segmentation - because I had a good access to the idea of U-Net i did the Unet-Watershed approach. Because I couldn't find any high scoring Unets (correct me please if I am wrong) I want to explain my try. Therefore I used Keras/ TF.\n\nAll in all I got LB scores of about 0.11 - not that much I know:)\n\nWhat I tried:\n\n- resizing to (256, 256) - for all cell types and also one modell for every cell type\n\n- resizing to (512, 512) - for all cell types and also one modell for every cell type\n\n- tiling images (resized original images to (520,705) and then generated patches of (260, 235) in \n order to use the full image/mask. I resized those tiles then to (256, 256) and used them for training.  \n While submission, I resized the mask also to (520, 705), and did the whole process again, and then \n rebuild the single tiles back to the original mask - for all cell types and also one modell for every cell \n type.\n\nI trained first with binary crossentropy about 15 epochs and then changed to 30 - 40 epochs of dice loss - both with Adam and its default settings, which got me to a dice coef meitrc of about 0.75-0.8. I just these tow stages of training just because it lead to the best results.\n\nI also used different watershed settings for the three cell types - which I derived manually with some examples for each cell type.\n\nI roughly had pretty much the same augmentation settings (ImageDataGenerator):\n\n- rotation_range = 10 (0 - 10)\n- width_shift_range = 0.25 (0. - 0.25)\n- height_shift_range = 0.25 (0. - 0.25)\n- shear_range = 0.1 (0. - 0.1)\n- zoom_range = 0.1 (0. - 0.1)\n- fill_mode = 'reflect'\n- horizontal_flip = True\n- vertical_flip = True\n- drop_rate = 0.1 (0.05-0.-125)\nI also tried to evaluate different thresholds, but couln't see any differences.\n\nWhat were your experiences? Was it the wrong approch?\n\nI wish you all a good start for 2022 with hopefully lots of exiting competition?\n\nCheers, Andreas",
    "1635510": "Cellpose is a unet model. Therefore many top teams have been using unet.\n\nI am not sure this is the answer you wanted, but it is nevertheless true.  You probably ask about unet with only segementation mask as target, instead of the spatial gradients used by cellpose.",
    "1638433": "Hi!\n\nI have tried some submissions with U-net: [https://www.kaggle.com/carlosgut/sartorius-complete-from-eda-to-submit-unet](url). At the end, I got similar results as you mention.\n\nI have mainly worked with image size 256x256 and I focused on TF/keras without using any external source or repository. The model was a custom U-net, trained for around 70 epochs, with adam optimizer, cross-binary entropy and dice metric.\n\nThings that I have experimented: color/b&w, size, data augmentation, epochs, train-val split, threshold, data power enhacements...\n\nAfter all submissions, here it goes my conclusions (may or may not be accurate, but it is what I have seen):\n\n- Having the subsmision file properly made is more important than any other approach: U-net returns the whole image, not the isolated points. Breaking them, remove the lonely pixels, correct minimun size of nucleus, boosted the score much more than any experimentation (perhaps this is only on lower scores). \n- Color does not add any useful value\n- Validation split of 12% improved the results rather than a bigger split (15 or 20%).\n- With my training, slightly larger epochs improved the results. Going up to 70 epochs (with earlystopper and learning rate decrease), improved around 0.1 the LB score.\n- Threshold was not helpful in general, but it does matter if you take into consideration cell type (astro type has wider area). I think many high score solutions use this.\n\n\nThose are some of things that I have experience. I am not an expert, just recently landed here :).\n\n\nI am also curious on any feedback, experiences or comments of this topic.\nCheers!",
    "1639727": "Thanks a lot for your answer :)",
    "1639728": "Okay, even this is new to me, thx that helps!",
    "1641295": "First, congratulations to the winners and thanks to the organizers for a challenging and interesting competition. For what it's worth, here is my experience with U-nets in this contest:\n\nI managed 0.211 on Private LB using a fairly simple U-net enhanced with a few home-grown tricks. It would have been 0.213 but my last submission finished scoring a few minutes after the deadline. For the 0.211 submission I ensembled 10 similar models trained on my own computers in R, Keras, and Tensorflow. Hardware included x86_64 workstations running Ubuntu Linux, and a single Nvidia Quadro GP100 GPU. Ensembling was done by averaging raw numeric model predictions over all the models before any post-processing. My Public and Private LB scores were somewhat higher than my validation scores but correlated well with them, and my best two submissions on Public LB were also best on Private LB.\n\nAlthough my U-net originally predicted only a single mask over all the cells, I added 4 more prediction outputs to the model to try to help with segmentation during post-processing: a mask over \"boundaries between cell perimeters\", a total count of cells in the image, the total fraction of the image contained in those cells, and probabilities of the 3 cell types. I didn't have any special method of dealing with the cell overlaps in the training set; each pixel was essentially randomly assigned to a single cell. To ensure its independence from the cell mask prediction, my boundary mask prediction came from a separate set of \"Up\" layers starting from the center of the U-net. Altogether, my net included 42 convolutional layers, spread among 4 \"down\", 1 center, and 8 \"up\" sections (4 for the mask, and 4 for the perimeter boundaries).\n\nThe perimeter boundary mask for each training image was the union, over all individual cell masks for that image, of the results of dilating each mask slightly using a kernel size of 3 and computing the intersection (overlap) of that dilated mask with the union of all the other cell masks. The resulting binary mask consisted of very thin lines separating adjacent cell masks.\n\nFor training, all images and masks were resized to 512x512. Rather than dividing training into epochs, instead I called train_on_batch for thousands of batches each containing 8 randomly-selected images, and decayed the learning rate when progress in reducing the loss stalled. Augmentations (training only) included flip, flop, transpose, and a type of non-linear warping that transformed images (and their masks) as if they were reflected by a curved mirror.\n\nPost-processing consisted of some optional \"temperature sharpening\" of the predicted mask pixel intensities, followed by searching for the \"best (lowest) mask score\" over several mask rounding thresholds starting at 0.002 and incremented by 0.002 until a trend toward worse scores was detected. For each candidate threshold, pixel intensities in the predicted mask that equaled or exceeded the threshold were rounded to 1 and the remainder set to 0. Then an optional \"opening\" (erosion followed by dilation) was performed on the binary mask. Then the predicted perimeter boundary mask was rounded to binary using a specified threshold value, and was used like a pair of scissors to cut thin lines through the predicted cell mask to help EBImage::bwlabel, which labels connected sets of pixels, to identify individual cells. The mask score was computed as a function of discrepancies: 1 - between the cell count predicted directly by the model and the count of objects produced from the predicted mask by bwlabel, and 2 - between the fraction of the image's pixels contained in cells predicted directly by the model and the fraction computed from the predicted mask. Each of those differences was raised to a specified power and the result multiplied by a specified factor, then the two difference terms were added together to produce the total mask score.\n\nAfter the best mask was determined, those predicted cell masks whose area (pixel count) was less than a specified minimum were deleted.\n\nA bunch of hyperparameters, most of which were 3-element vectors indexed by the highest probability cell type prediction, controlled the post-processing, and I did some automatic tuning to try to find good values. With this dataset that sounds like a recipe for overfitting, but at least to the limited extent I did it (it was very time intensive) it seemed to improve performance both in my validation runs and on the LB.\n\nI only hit on the \"cell perimeter boundary\" idea late in the competition, and my results suggest that it may be possible to improve its performance considerably, although probably not nearly enough to be competitive with Detectron2, Mask R-cnn, etc.\n\nSo, what do readers think: does the above description of my program suggest any useful ideas, or is it just a cautionary tale of what methods to avoid when approaching an object detection problem like this contest?"
  },
  "source": "meta"
}