{
  "id": 216984,
  "title": "Cell Crop Tile Datasets (256x256)",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/216984",
  "author_name": "Darien Schettler",
  "post_date": "2021-02-04T19:18:46.601000",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi there.</p>\n<p>I have made <strong>4 datasets</strong> and would like to share them. All datasets contain <strong>square, 256x256 images centred around an identified cell</strong>.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-red-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Red Channel (Microtubules)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-green-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Green Channel (Protein of Interest)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-blue-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Blue Channel (Nucleus)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-yellow-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Yellow Channel (Endoplasmic Reticulum)</strong></a></li>\n</ul>\n<p><br></p>\n<p><strong>The process to make the datasets is as follows:</strong></p>\n<hr>\n<p><strong>Step 1:</strong> Run CellSegmentator tool <em>(or use <a href=\"https://www.kaggle.com/its7171/hpa-mask\" target=\"_blank\">Tito's Public dataset</a> of mask images, or use <a href=\"https://www.kaggle.com/dschettler8845/hpa-processed-train-dataframe-with-cellwise-rle\" target=\"_blank\">my public dataset containing updated train.csv with RLEs</a>)</em> to identify masks of all cells in all images.<br>\n<strong>Step 2:</strong> Identify the bounding box for each cell mask in all images<br>\n<strong>Step 3:</strong> Pad each bounding box to be square (use black space, pad content towards the centre)<br>\n<strong>Step 4:</strong> Resize each bounding box to be 256x256 (this is big… bigger than in the original image often) <br>\n<strong>Step 5:</strong> Open and stack RGBY channels into a single image and cut out the relevant squares surrounding all identified cells.<br>\n<strong>Step 6:</strong> Save these cellwise crops to a new folder and label accordingly.</p>\n<hr>\n<p><em>Apologies in advance for the messed up folder structure of the datasets. It, at least, is consistent in its weirdness. If you need help parsing it please let me know and I will post a code snippet below.</em></p>\n<hr>\n<p><br></p>\n<p><strong><em>Credit to <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> for his public mask dataset.</em></strong></p>",
  "messages": [
    {
      "id": 1186356,
      "postDate": "2021-02-04T19:18:46.600Z",
      "content": "<p>Hi there.</p>\n<p>I have made <strong>4 datasets</strong> and would like to share them. All datasets contain <strong>square, 256x256 images centred around an identified cell</strong>.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-red-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Red Channel (Microtubules)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-green-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Green Channel (Protein of Interest)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-blue-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Blue Channel (Nucleus)</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/human-protein-atlas-yellow-cell-tile-dataset?rvi=1\" target=\"_blank\"><strong>Yellow Channel (Endoplasmic Reticulum)</strong></a></li>\n</ul>\n<p><br></p>\n<p><strong>The process to make the datasets is as follows:</strong></p>\n<hr>\n<p><strong>Step 1:</strong> Run CellSegmentator tool <em>(or use <a href=\"https://www.kaggle.com/its7171/hpa-mask\" target=\"_blank\">Tito's Public dataset</a> of mask images, or use <a href=\"https://www.kaggle.com/dschettler8845/hpa-processed-train-dataframe-with-cellwise-rle\" target=\"_blank\">my public dataset containing updated train.csv with RLEs</a>)</em> to identify masks of all cells in all images.<br>\n<strong>Step 2:</strong> Identify the bounding box for each cell mask in all images<br>\n<strong>Step 3:</strong> Pad each bounding box to be square (use black space, pad content towards the centre)<br>\n<strong>Step 4:</strong> Resize each bounding box to be 256x256 (this is big… bigger than in the original image often) <br>\n<strong>Step 5:</strong> Open and stack RGBY channels into a single image and cut out the relevant squares surrounding all identified cells.<br>\n<strong>Step 6:</strong> Save these cellwise crops to a new folder and label accordingly.</p>\n<hr>\n<p><em>Apologies in advance for the messed up folder structure of the datasets. It, at least, is consistent in its weirdness. If you need help parsing it please let me know and I will post a code snippet below.</em></p>\n<hr>\n<p><br></p>\n<p><strong><em>Credit to <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> for his public mask dataset.</em></strong></p>",
      "rawMarkdown": "Hi there.\n\nI have made **4 datasets** and would like to share them. All datasets contain **square, 256x256 images centred around an identified cell**.\n- [**Red Channel (Microtubules)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-red-cell-tile-dataset?rvi=1)\n- [**Green Channel (Protein of Interest)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-green-cell-tile-dataset?rvi=1)\n- [**Blue Channel (Nucleus)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-blue-cell-tile-dataset?rvi=1)\n- [**Yellow Channel (Endoplasmic Reticulum)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-yellow-cell-tile-dataset?rvi=1)\n\n<br>\n\n**The process to make the datasets is as follows:**\n\n---\n\n**Step 1:** Run CellSegmentator tool *(or use [Tito's Public dataset](https://www.kaggle.com/its7171/hpa-mask) of mask images, or use [my public dataset containing updated train.csv with RLEs](https://www.kaggle.com/dschettler8845/hpa-processed-train-dataframe-with-cellwise-rle))* to identify masks of all cells in all images.\n**Step 2:** Identify the bounding box for each cell mask in all images\n**Step 3:** Pad each bounding box to be square (use black space, pad content towards the centre)\n**Step 4:** Resize each bounding box to be 256x256 (this is big... bigger than in the original image often) \n**Step 5:** Open and stack RGBY channels into a single image and cut out the relevant squares surrounding all identified cells.\n**Step 6:** Save these cellwise crops to a new folder and label accordingly.\n\n---\n\n*Apologies in advance for the messed up folder structure of the datasets. It, at least, is consistent in its weirdness. If you need help parsing it please let me know and I will post a code snippet below.*\n\n---\n\n<br>\n\n***Credit to @its7171 for his public mask dataset.***\n",
      "votes": 13
    },
    {
      "id": 1212688,
      "postDate": "2021-02-21T13:36:56.590Z",
      "content": "<p>Did you only use the images with one label associated?<br>\nIf yes, you might have also some negatives wrongly categorized as a certain category?</p>",
      "rawMarkdown": "Did you only use the images with one label associated?\nIf yes, you might have also some negatives wrongly categorized as a certain category?",
      "votes": 1,
      "replies": [
        {
          "id": 1212754,
          "postDate": "2021-02-21T14:52:20.550Z",
          "content": "<p>You are 100% correct. This dataset is a first iteration of what will likely require many steps.</p>\n<p>I recently took this dataset and applied manual heuristics to filter out the misclassified negative classes. It looks like between 30-50% of the cell-wise images are actually negative (assuming my manual heuristics are not garbage).</p>\n<p>I'll be posting that dataset once I've had a chance to conduct more experimentation.</p>",
          "rawMarkdown": "You are 100% correct. This dataset is a first iteration of what will likely require many steps.\n\nI recently took this dataset and applied manual heuristics to filter out the misclassified negative classes. It looks like between 30-50% of the cell-wise images are actually negative (assuming my manual heuristics are not garbage).\n\nI'll be posting that dataset once I've had a chance to conduct more experimentation.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1227386,
      "postDate": "2021-03-05T13:33:53.020Z",
      "content": "<p>Hi. So for segmentation if you had used the HPA tool it would have given a binary output image right with 0 as the background and 1 as the foreground. Have you uploaded the output of the segmentation separately?</p>",
      "rawMarkdown": "Hi. So for segmentation if you had used the HPA tool it would have given a binary output image right with 0 as the background and 1 as the foreground. Have you uploaded the output of the segmentation separately?",
      "replies": [
        {
          "id": 1227545,
          "postDate": "2021-03-05T16:23:39.803Z",
          "content": "<p>No. The mask dataset that correlates to these datasets can be found <a href=\"https://www.kaggle.com/its7171/hpa-mask\" target=\"_blank\"><strong>here</strong></a> and was created by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\"><strong>Tito</strong></a>.</p>\n<hr>\n<p>If you'd like to see an updated dataset (and corresponding mask information), please find the two links below. These datasets were generated using the <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\"><strong>Even Faster Cell Segmentation</strong></a> version of the CellSegmentator tool.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/hpa-rgb-tilewise-dataset-v20\" target=\"_blank\"><strong>Link to RGB Dataset</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/hpa-train-data-with-additional-metadata\" target=\"_blank\"><strong>Link to updated CSV dataset containing mask and bounding box information</strong></a></li>\n</ul>",
          "rawMarkdown": "No. The mask dataset that correlates to these datasets can be found [**here**](https://www.kaggle.com/its7171/hpa-mask) and was created by [**Tito**](https://www.kaggle.com/its7171).\n\n---\n\nIf you'd like to see an updated dataset (and corresponding mask information), please find the two links below. These datasets were generated using the [**Even Faster Cell Segmentation**](https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation) version of the CellSegmentator tool.\n- [**Link to RGB Dataset**](https://www.kaggle.com/dschettler8845/hpa-rgb-tilewise-dataset-v20)\n- [**Link to updated CSV dataset containing mask and bounding box information**](https://www.kaggle.com/dschettler8845/hpa-train-data-with-additional-metadata)",
          "votes": 1
        },
        {
          "id": 1227547,
          "postDate": "2021-03-05T16:28:27.813Z",
          "content": "<p>Also, I believe that the HPA tool does not output a binary mask. It outputs a mask where each cell is masked by a single value. i.e. Cell 1 is masked by 1s, Cell 2 is masked by 2s, etc. This allows us to identify overlapping cells with more ease.</p>\n<p>To go from this to individual binary cell maks, you simply use something like this…</p>\n<p><b></b></p>\n<pre><code># `all_masks` is the output I described above\n# `individual_masks` will be a list containing a binary mask for each cell\n\nindividual_masks = [\n  np.where(all_masks==i, 1, 0) \\\n  for i in range(1, all_masks.max()+1)\n]\n</code></pre>\n<p></p>",
          "rawMarkdown": "Also, I believe that the HPA tool does not output a binary mask. It outputs a mask where each cell is masked by a single value. i.e. Cell 1 is masked by 1s, Cell 2 is masked by 2s, etc. This allows us to identify overlapping cells with more ease.\n\nTo go from this to individual binary cell maks, you simply use something like this...\n\n<b>\n\n```python\n\n# `all_masks` is the output I described above\n# `individual_masks` will be a list containing a binary mask for each cell\n\nindividual_masks = [\n  np.where(all_masks==i, 1, 0) \\\n  for i in range(1, all_masks.max()+1)\n]\n\n```\n\n</b>",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1212688,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-02-21T13:36:56.590000",
      "content": "<p>Did you only use the images with one label associated?<br>\nIf yes, you might have also some negatives wrongly categorized as a certain category?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1212754,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-21T14:52:20.550000",
          "content": "<p>You are 100% correct. This dataset is a first iteration of what will likely require many steps.</p>\n<p>I recently took this dataset and applied manual heuristics to filter out the misclassified negative classes. It looks like between 30-50% of the cell-wise images are actually negative (assuming my manual heuristics are not garbage).</p>\n<p>I'll be posting that dataset once I've had a chance to conduct more experimentation.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1227386,
      "author_name": "RSASHWIN",
      "author_url": "",
      "post_date": "2021-03-05T13:33:53.020000",
      "content": "<p>Hi. So for segmentation if you had used the HPA tool it would have given a binary output image right with 0 as the background and 1 as the foreground. Have you uploaded the output of the segmentation separately?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1227545,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-05T16:23:39.803000",
          "content": "<p>No. The mask dataset that correlates to these datasets can be found <a href=\"https://www.kaggle.com/its7171/hpa-mask\" target=\"_blank\"><strong>here</strong></a> and was created by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\"><strong>Tito</strong></a>.</p>\n<hr>\n<p>If you'd like to see an updated dataset (and corresponding mask information), please find the two links below. These datasets were generated using the <a href=\"https://www.kaggle.com/samusram/even-faster-hpa-cell-segmentation\" target=\"_blank\"><strong>Even Faster Cell Segmentation</strong></a> version of the CellSegmentator tool.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/dschettler8845/hpa-rgb-tilewise-dataset-v20\" target=\"_blank\"><strong>Link to RGB Dataset</strong></a></li>\n<li><a href=\"https://www.kaggle.com/dschettler8845/hpa-train-data-with-additional-metadata\" target=\"_blank\"><strong>Link to updated CSV dataset containing mask and bounding box information</strong></a></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1227547,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-05T16:28:27.813000",
          "content": "<p>Also, I believe that the HPA tool does not output a binary mask. It outputs a mask where each cell is masked by a single value. i.e. Cell 1 is masked by 1s, Cell 2 is masked by 2s, etc. This allows us to identify overlapping cells with more ease.</p>\n<p>To go from this to individual binary cell maks, you simply use something like this…</p>\n<p><b></b></p>\n<pre><code># `all_masks` is the output I described above\n# `individual_masks` will be a list containing a binary mask for each cell\n\nindividual_masks = [\n  np.where(all_masks==i, 1, 0) \\\n  for i in range(1, all_masks.max()+1)\n]\n</code></pre>\n<p></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1186356": "Hi there.\n\nI have made **4 datasets** and would like to share them. All datasets contain **square, 256x256 images centred around an identified cell**.\n- [**Red Channel (Microtubules)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-red-cell-tile-dataset?rvi=1)\n- [**Green Channel (Protein of Interest)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-green-cell-tile-dataset?rvi=1)\n- [**Blue Channel (Nucleus)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-blue-cell-tile-dataset?rvi=1)\n- [**Yellow Channel (Endoplasmic Reticulum)**](https://www.kaggle.com/dschettler8845/human-protein-atlas-yellow-cell-tile-dataset?rvi=1)\n\n<br>\n\n**The process to make the datasets is as follows:**\n\n---\n\n**Step 1:** Run CellSegmentator tool *(or use [Tito's Public dataset](https://www.kaggle.com/its7171/hpa-mask) of mask images, or use [my public dataset containing updated train.csv with RLEs](https://www.kaggle.com/dschettler8845/hpa-processed-train-dataframe-with-cellwise-rle))* to identify masks of all cells in all images.\n**Step 2:** Identify the bounding box for each cell mask in all images\n**Step 3:** Pad each bounding box to be square (use black space, pad content towards the centre)\n**Step 4:** Resize each bounding box to be 256x256 (this is big... bigger than in the original image often) \n**Step 5:** Open and stack RGBY channels into a single image and cut out the relevant squares surrounding all identified cells.\n**Step 6:** Save these cellwise crops to a new folder and label accordingly.\n\n---\n\n*Apologies in advance for the messed up folder structure of the datasets. It, at least, is consistent in its weirdness. If you need help parsing it please let me know and I will post a code snippet below.*\n\n---\n\n<br>\n\n***Credit to @its7171 for his public mask dataset.***\n",
    "1212688": "Did you only use the images with one label associated?\nIf yes, you might have also some negatives wrongly categorized as a certain category?",
    "1227386": "Hi. So for segmentation if you had used the HPA tool it would have given a binary output image right with 0 as the background and 1 as the foreground. Have you uploaded the output of the segmentation separately?"
  }
}