{
  "id": 285384,
  "title": "LIVECell paper summarized",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/285384",
  "author_name": "",
  "post_date": "2021-11-04T14:16:37.004200500Z",
  "votes": 63,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've spent some time digging into the LIVECell dataset created by the contest hosts, here is the summary of things I've found interesting in the context of this competition.</p>\n<h2>What is it?</h2>\n<blockquote>\n  <p>LIVECell (Label-free In Vitro image Examples of Cells) is a new dataset of manually annotated, label-free, phase-contrast images of 2D cell culture.</p>\n</blockquote>\n<p>\"Label free\" here is used in the medical sense - no fluorescent labels added to the samples. It is labeled in the machine learning sense as there are manual annotations.</p>\n<h2>What is it for?</h2>\n<ul>\n<li>Introduces the largest high-quality resource for label-free cell segmentation.</li>\n<li>Presents trained models developed to segment individual cells, for application in new research to enable label-free single-cell studies.</li>\n<li>Proposes a suite of benchmarks, which will readily facilitate continued development and performance comparison of future models.</li>\n</ul>\n<h2>What's in images?</h2>\n<ul>\n<li>5239 manually annotated, expert-validated, Incucyte HD phase-contrast microscopy images.</li>\n<li>There are eight types of cells chosen to maximize diversity of morphologies. </li>\n<li>All the images were captured by the same equipment, same magnification, resolution and cropped to the same size.</li>\n</ul>\n<h2>What are the annotations?</h2>\n<ul>\n<li>More than a total of 1.6 million annotated cells.</li>\n<li>Manually annotated by a managed team of professional annotators, followed by two stages of quality assurance. The first level was performed by the annotation managers and second round by an experienced cell biologist.</li>\n</ul>\n<h2>What were the experiments?</h2>\n<ul>\n<li>Two architectures were tested using inherently different object detection mechanisms: anchor-based and anchor-free. The anchor-based model was an adapted version of Cascade Mask RCNN using a ResNest-200 backbone. The anchor-free model was based on CenterMask.</li>\n<li>Nine models of each architecture were trained on LIVECell, one model on the whole dataset (LIVECell-wide train and evaluate benchmark) and one for each of the eight cell types</li>\n<li>For normalization, pixel intensity values were centered around zero by subtracting by the global average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). To reduce the risk of overfitting, all training used multi-scale data augmentation meaning that image sizes were randomly changed from the original 520 × 704 pixels to a size with the same ratios, but shortest side set to one of (440, 480, 520, 580, 620) pixels.</li>\n<li>All models were trained on a NVIDIA DGX-1 server hosting eight NVIDIA Tesla V100 GPUs with 32 Gb of GPU RAM each, dual 20-Core 2.2 GHz Intel Xeon CPUs and 512 Gb system RAM.</li>\n<li>Microsoft COCO evaluation protocol was used but slightly modified to better reflect cell sizes.</li>\n</ul>\n<h2>What were the findings?</h2>\n<ul>\n<li>When trained and evaluated on all of LIVECell, the two models achieve very similar accuracy. Though anchor based model pulls ahead on smaller subsets of data.</li>\n<li>Certain cell types in LIVECell, predominately BV-2, BT-474, Huh7 and SH-SY5Y, form densely packed clusters where it is not possible even for an experienced cell biologist to detect boundaries between individual cells</li>\n<li>There is a large performance difference across cell types. In particular SH-SY5Y is very sensitive to IoU threshold.</li>\n<li>Models perform better on each individual cell-type test set when trained on all cell types compared to training on that single cell type - indicating that a cell-type universal model is preferable to a specific one.</li>\n</ul>\n<h2>References</h2>\n<ul>\n<li>The article: <a href=\"https://www.nature.com/articles/s41592-021-01249-6\" target=\"_blank\">https://www.nature.com/articles/s41592-021-01249-6</a></li>\n<li>The repository: <a href=\"https://github.com/sartorius-research/LIVECell\" target=\"_blank\">https://github.com/sartorius-research/LIVECell</a></li>\n</ul>",
  "messages": [
    {
      "id": "1570970",
      "postDate": "11/04/2021 14:16:37",
      "content": "<p>I've spent some time digging into the LIVECell dataset created by the contest hosts, here is the summary of things I've found interesting in the context of this competition.</p>\n<h2>What is it?</h2>\n<blockquote>\n  <p>LIVECell (Label-free In Vitro image Examples of Cells) is a new dataset of manually annotated, label-free, phase-contrast images of 2D cell culture.</p>\n</blockquote>\n<p>\"Label free\" here is used in the medical sense - no fluorescent labels added to the samples. It is labeled in the machine learning sense as there are manual annotations.</p>\n<h2>What is it for?</h2>\n<ul>\n<li>Introduces the largest high-quality resource for label-free cell segmentation.</li>\n<li>Presents trained models developed to segment individual cells, for application in new research to enable label-free single-cell studies.</li>\n<li>Proposes a suite of benchmarks, which will readily facilitate continued development and performance comparison of future models.</li>\n</ul>\n<h2>What's in images?</h2>\n<ul>\n<li>5239 manually annotated, expert-validated, Incucyte HD phase-contrast microscopy images.</li>\n<li>There are eight types of cells chosen to maximize diversity of morphologies. </li>\n<li>All the images were captured by the same equipment, same magnification, resolution and cropped to the same size.</li>\n</ul>\n<h2>What are the annotations?</h2>\n<ul>\n<li>More than a total of 1.6 million annotated cells.</li>\n<li>Manually annotated by a managed team of professional annotators, followed by two stages of quality assurance. The first level was performed by the annotation managers and second round by an experienced cell biologist.</li>\n</ul>\n<h2>What were the experiments?</h2>\n<ul>\n<li>Two architectures were tested using inherently different object detection mechanisms: anchor-based and anchor-free. The anchor-based model was an adapted version of Cascade Mask RCNN using a ResNest-200 backbone. The anchor-free model was based on CenterMask.</li>\n<li>Nine models of each architecture were trained on LIVECell, one model on the whole dataset (LIVECell-wide train and evaluate benchmark) and one for each of the eight cell types</li>\n<li>For normalization, pixel intensity values were centered around zero by subtracting by the global average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). To reduce the risk of overfitting, all training used multi-scale data augmentation meaning that image sizes were randomly changed from the original 520 × 704 pixels to a size with the same ratios, but shortest side set to one of (440, 480, 520, 580, 620) pixels.</li>\n<li>All models were trained on a NVIDIA DGX-1 server hosting eight NVIDIA Tesla V100 GPUs with 32 Gb of GPU RAM each, dual 20-Core 2.2 GHz Intel Xeon CPUs and 512 Gb system RAM.</li>\n<li>Microsoft COCO evaluation protocol was used but slightly modified to better reflect cell sizes.</li>\n</ul>\n<h2>What were the findings?</h2>\n<ul>\n<li>When trained and evaluated on all of LIVECell, the two models achieve very similar accuracy. Though anchor based model pulls ahead on smaller subsets of data.</li>\n<li>Certain cell types in LIVECell, predominately BV-2, BT-474, Huh7 and SH-SY5Y, form densely packed clusters where it is not possible even for an experienced cell biologist to detect boundaries between individual cells</li>\n<li>There is a large performance difference across cell types. In particular SH-SY5Y is very sensitive to IoU threshold.</li>\n<li>Models perform better on each individual cell-type test set when trained on all cell types compared to training on that single cell type - indicating that a cell-type universal model is preferable to a specific one.</li>\n</ul>\n<h2>References</h2>\n<ul>\n<li>The article: <a href=\"https://www.nature.com/articles/s41592-021-01249-6\" target=\"_blank\">https://www.nature.com/articles/s41592-021-01249-6</a></li>\n<li>The repository: <a href=\"https://github.com/sartorius-research/LIVECell\" target=\"_blank\">https://github.com/sartorius-research/LIVECell</a></li>\n</ul>",
      "rawMarkdown": "I've spent some time digging into the LIVECell dataset created by the contest hosts, here is the summary of things I've found interesting in the context of this competition.\n\n## What is it?\n> LIVECell (Label-free In Vitro image Examples of Cells) is a new dataset of manually annotated, label-free, phase-contrast images of 2D cell culture.\n\n\"Label free\" here is used in the medical sense - no fluorescent labels added to the samples. It is labeled in the machine learning sense as there are manual annotations.\n\n## What is it for?\n* Introduces the largest high-quality resource for label-free cell segmentation.\n* Presents trained models developed to segment individual cells, for application in new research to enable label-free single-cell studies.\n* Proposes a suite of benchmarks, which will readily facilitate continued development and performance comparison of future models.\n\n## What's in images?\n* 5239 manually annotated, expert-validated, Incucyte HD phase-contrast microscopy images.\n* There are eight types of cells chosen to maximize diversity of morphologies. \n* All the images were captured by the same equipment, same magnification, resolution and cropped to the same size.\n\n## What are the annotations?\n* More than a total of 1.6 million annotated cells.\n* Manually annotated by a managed team of professional annotators, followed by two stages of quality assurance. The first level was performed by the annotation managers and second round by an experienced cell biologist.\n\n## What were the experiments?\n* Two architectures were tested using inherently different object detection mechanisms: anchor-based and anchor-free. The anchor-based model was an adapted version of Cascade Mask RCNN using a ResNest-200 backbone. The anchor-free model was based on CenterMask.\n* Nine models of each architecture were trained on LIVECell, one model on the whole dataset (LIVECell-wide train and evaluate benchmark) and one for each of the eight cell types\n* For normalization, pixel intensity values were centered around zero by subtracting by the global average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). To reduce the risk of overfitting, all training used multi-scale data augmentation meaning that image sizes were randomly changed from the original 520 × 704 pixels to a size with the same ratios, but shortest side set to one of (440, 480, 520, 580, 620) pixels.\n* All models were trained on a NVIDIA DGX-1 server hosting eight NVIDIA Tesla V100 GPUs with 32 Gb of GPU RAM each, dual 20-Core 2.2 GHz Intel Xeon CPUs and 512 Gb system RAM.\n* Microsoft COCO evaluation protocol was used but slightly modified to better reflect cell sizes.\n\n## What were the findings?\n* When trained and evaluated on all of LIVECell, the two models achieve very similar accuracy. Though anchor based model pulls ahead on smaller subsets of data.\n* Certain cell types in LIVECell, predominately BV-2, BT-474, Huh7 and SH-SY5Y, form densely packed clusters where it is not possible even for an experienced cell biologist to detect boundaries between individual cells\n* There is a large performance difference across cell types. In particular SH-SY5Y is very sensitive to IoU threshold.\n* Models perform better on each individual cell-type test set when trained on all cell types compared to training on that single cell type - indicating that a cell-type universal model is preferable to a specific one.\n\n## References\n* The article: https://www.nature.com/articles/s41592-021-01249-6\n* The repository: https://github.com/sartorius-research/LIVECell",
      "votes": null
    },
    {
      "id": "1571013",
      "postDate": "11/04/2021 14:40:03",
      "content": "<p>nice write-up</p>",
      "rawMarkdown": "nice write-up",
      "votes": null
    },
    {
      "id": "1571018",
      "postDate": "11/04/2021 14:42:42",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>! I thought it could be useful and save others some time.</p>",
      "rawMarkdown": "Thanks, @namgalielei! I thought it could be useful and save others some time.",
      "votes": null
    },
    {
      "id": "1571058",
      "postDate": "11/04/2021 15:04:42",
      "content": "<p>Great write-up Slawek, I hope that there is things from our paper and released resources that can help you out during the competition. </p>",
      "rawMarkdown": "Great write-up Slawek, I hope that there is things from our paper and released resources that can help you out during the competition.",
      "votes": null
    },
    {
      "id": "1571065",
      "postDate": "11/04/2021 15:12:14",
      "content": "<p>Thanks Christoffer! Yes, it was definitely helpful to read about previous experiments. </p>",
      "rawMarkdown": "Thanks Christoffer! Yes, it was definitely helpful to read about previous experiments.",
      "votes": null
    },
    {
      "id": "1571999",
      "postDate": "11/05/2021 10:43:46",
      "content": "<p>Good summary~ Thank you~</p>",
      "rawMarkdown": "Good summary~ Thank you~",
      "votes": null
    },
    {
      "id": "1572139",
      "postDate": "11/05/2021 12:54:54",
      "content": "<p>thx for sharing</p>",
      "rawMarkdown": "thx for sharing",
      "votes": null
    },
    {
      "id": "1578770",
      "postDate": "11/11/2021 10:22:43",
      "content": "<p>I'm a final term Software Engineering Student and I had the opportunity to be selected to do my final term project in a laboratory that works on Scientific computing for image-based system biology. the project will take about 6 months. I'm still fairly new to the domain. I passed the Tensorflow certificate I worked a bit with Deep learning and Vision Transformers. But my journey only started less than a year ago. I'm facing a problem choosing the subject of my project It's a research project that also needs to include coding. any ideas?</p>",
      "rawMarkdown": "I'm a final term Software Engineering Student and I had the opportunity to be selected to do my final term project in a laboratory that works on Scientific computing for image-based system biology. the project will take about 6 months. I'm still fairly new to the domain. I passed the Tensorflow certificate I worked a bit with Deep learning and Vision Transformers. But my journey only started less than a year ago. I'm facing a problem choosing the subject of my project It's a research project that also needs to include coding. any ideas?",
      "votes": null
    },
    {
      "id": "1602714",
      "postDate": "12/02/2021 04:51:16",
      "content": "<p>Has anyone tried this normalization:</p>\n<blockquote>\n  <p>average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). </p>\n</blockquote>",
      "rawMarkdown": "Has anyone tried this normalization:\n> average pixel value for the dataset (128), and then divided by the global standard deviation (11.58).",
      "votes": null
    },
    {
      "id": "1617492",
      "postDate": "12/14/2021 06:06:05",
      "content": "<p>cool. thanks for sharing</p>",
      "rawMarkdown": "cool. thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1571013,
      "author_name": "namgalielei",
      "author_url": "",
      "post_date": "11/04/2021 14:40:03",
      "content": "<p>nice write-up</p>",
      "votes": null,
      "replies": [
        {
          "id": 1571018,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/04/2021 14:42:42",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a>! I thought it could be useful and save others some time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1571058,
      "author_name": "christoffersartorius",
      "author_url": "",
      "post_date": "11/04/2021 15:04:42",
      "content": "<p>Great write-up Slawek, I hope that there is things from our paper and released resources that can help you out during the competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1571065,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/04/2021 15:12:14",
          "content": "<p>Thanks Christoffer! Yes, it was definitely helpful to read about previous experiments. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1571999,
      "author_name": "mingookkim",
      "author_url": "",
      "post_date": "11/05/2021 10:43:46",
      "content": "<p>Good summary~ Thank you~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1572139,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "11/05/2021 12:54:54",
      "content": "<p>thx for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1578770,
      "author_name": "rouarouatbi",
      "author_url": "",
      "post_date": "11/11/2021 10:22:43",
      "content": "<p>I'm a final term Software Engineering Student and I had the opportunity to be selected to do my final term project in a laboratory that works on Scientific computing for image-based system biology. the project will take about 6 months. I'm still fairly new to the domain. I passed the Tensorflow certificate I worked a bit with Deep learning and Vision Transformers. But my journey only started less than a year ago. I'm facing a problem choosing the subject of my project It's a research project that also needs to include coding. any ideas?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1602714,
      "author_name": "mlneo07",
      "author_url": "",
      "post_date": "12/02/2021 04:51:16",
      "content": "<p>Has anyone tried this normalization:</p>\n<blockquote>\n  <p>average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). </p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1617492,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "12/14/2021 06:06:05",
      "content": "<p>cool. thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1570970": "I've spent some time digging into the LIVECell dataset created by the contest hosts, here is the summary of things I've found interesting in the context of this competition.\n\n## What is it?\n> LIVECell (Label-free In Vitro image Examples of Cells) is a new dataset of manually annotated, label-free, phase-contrast images of 2D cell culture.\n\n\"Label free\" here is used in the medical sense - no fluorescent labels added to the samples. It is labeled in the machine learning sense as there are manual annotations.\n\n## What is it for?\n* Introduces the largest high-quality resource for label-free cell segmentation.\n* Presents trained models developed to segment individual cells, for application in new research to enable label-free single-cell studies.\n* Proposes a suite of benchmarks, which will readily facilitate continued development and performance comparison of future models.\n\n## What's in images?\n* 5239 manually annotated, expert-validated, Incucyte HD phase-contrast microscopy images.\n* There are eight types of cells chosen to maximize diversity of morphologies. \n* All the images were captured by the same equipment, same magnification, resolution and cropped to the same size.\n\n## What are the annotations?\n* More than a total of 1.6 million annotated cells.\n* Manually annotated by a managed team of professional annotators, followed by two stages of quality assurance. The first level was performed by the annotation managers and second round by an experienced cell biologist.\n\n## What were the experiments?\n* Two architectures were tested using inherently different object detection mechanisms: anchor-based and anchor-free. The anchor-based model was an adapted version of Cascade Mask RCNN using a ResNest-200 backbone. The anchor-free model was based on CenterMask.\n* Nine models of each architecture were trained on LIVECell, one model on the whole dataset (LIVECell-wide train and evaluate benchmark) and one for each of the eight cell types\n* For normalization, pixel intensity values were centered around zero by subtracting by the global average pixel value for the dataset (128), and then divided by the global standard deviation (11.58). To reduce the risk of overfitting, all training used multi-scale data augmentation meaning that image sizes were randomly changed from the original 520 × 704 pixels to a size with the same ratios, but shortest side set to one of (440, 480, 520, 580, 620) pixels.\n* All models were trained on a NVIDIA DGX-1 server hosting eight NVIDIA Tesla V100 GPUs with 32 Gb of GPU RAM each, dual 20-Core 2.2 GHz Intel Xeon CPUs and 512 Gb system RAM.\n* Microsoft COCO evaluation protocol was used but slightly modified to better reflect cell sizes.\n\n## What were the findings?\n* When trained and evaluated on all of LIVECell, the two models achieve very similar accuracy. Though anchor based model pulls ahead on smaller subsets of data.\n* Certain cell types in LIVECell, predominately BV-2, BT-474, Huh7 and SH-SY5Y, form densely packed clusters where it is not possible even for an experienced cell biologist to detect boundaries between individual cells\n* There is a large performance difference across cell types. In particular SH-SY5Y is very sensitive to IoU threshold.\n* Models perform better on each individual cell-type test set when trained on all cell types compared to training on that single cell type - indicating that a cell-type universal model is preferable to a specific one.\n\n## References\n* The article: https://www.nature.com/articles/s41592-021-01249-6\n* The repository: https://github.com/sartorius-research/LIVECell",
    "1571013": "nice write-up",
    "1571018": "Thanks, @namgalielei! I thought it could be useful and save others some time.",
    "1571058": "Great write-up Slawek, I hope that there is things from our paper and released resources that can help you out during the competition.",
    "1571065": "Thanks Christoffer! Yes, it was definitely helpful to read about previous experiments.",
    "1571999": "Good summary~ Thank you~",
    "1572139": "thx for sharing",
    "1578770": "I'm a final term Software Engineering Student and I had the opportunity to be selected to do my final term project in a laboratory that works on Scientific computing for image-based system biology. the project will take about 6 months. I'm still fairly new to the domain. I passed the Tensorflow certificate I worked a bit with Deep learning and Vision Transformers. But my journey only started less than a year ago. I'm facing a problem choosing the subject of my project It's a research project that also needs to include coding. any ideas?",
    "1602714": "Has anyone tried this normalization:\n> average pixel value for the dataset (128), and then divided by the global standard deviation (11.58).",
    "1617492": "cool. thanks for sharing"
  },
  "source": "meta"
}