{
  "id": 229586,
  "title": "Finally Completed Single Label Cell Level Dataset",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/229586",
  "author_name": "Ayush Thakur",
  "post_date": "2021-03-30T22:38:56.727000",
  "votes": 26,
  "comment_count": 13,
  "views": 0,
  "content": "<p>For someday now I was working on preparing a single label (images with only one label associated with it) cell level (individual cells are labeled) dataset to train a better image classifier. Here's a quick summary.</p>\n<h1>Steps</h1>\n<h3>Cell Level Image Dataset Creation</h3>\n<ul>\n<li>Start with the HPA competition image and using the HPA Segmentation tool segment it to get <code>cell_mask</code> and <code>nuclei_mask</code>. </li>\n<li>Using <code>cell_mask</code> get individual cell patches.</li>\n<li>Using `nuclei_mask' get individual nucleus patches. The contour of the nucleus patch is computed. Using this it is determined if the nucleus is bordering the image. I am discarding the cells whose nucleus is bordering the image. <strong>I am assuming that such a cell would not have \"complete\" protein information</strong>. </li>\n<li><strong>Note</strong> that I am creating two separate RGB and protein image datasets. The DSViz table shown in the image can give you an idea of my dataset. (<a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_split_0/34b460a9c6ea265a4e95/files/dataset.table.json\" target=\"_blank\">W&amp;B DSViz table link</a>)<br>\n<img src=\"https://i.imgur.com/fU8hjlb.png\" alt=\"img\"></li>\n</ul>\n<p>Kernels used:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/dataset-versioning-for-single-label-dataset\" target=\"_blank\">Dataset Versioning for Single Label Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation\" target=\"_blank\">Using W&amp;B DSViz for Cell Level Dataset Creation</a></li>\n</ul>\n<p>Kaggle Dataset:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\" target=\"_blank\">https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset</a></li>\n</ul>\n<h3>Data Distribution</h3>\n<p>The image below is the label distribution of the single label cell level dataset.<br>\n<img src=\"https://i.imgur.com/S9BKxvd.png\" alt=\"img\"></p>\n<h3>Extra HPA Public Data</h3>\n<p>Certainly, there is an insane class imbalance. To mitigate the imbalance to some extent I used HPA Pubic Data. I decided to create separate single label cell level datasets for these classes - Nuclear Membrane, Actin Filaments, Aggresome, Mitotic Spindle, and Negative. </p>\n<p>Kernels used:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/cell-level-extra-data-download\" target=\"_blank\">Cell Level Extra Data Download</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation\" target=\"_blank\">Single Label Cell Label Dataset Creation w\\ W&amp;B</a><br>\n(Apologies for not documenting them well.)</li>\n</ul>\n<p>Kaggle Dataset:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicmitoticspindle\" target=\"_blank\">HPA: Single Label Public Mitotic Spindle</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicactinfilaments\" target=\"_blank\">HPA: Single Label Public ACTIN FILAMENTS</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicnegative\" target=\"_blank\">HPA: Single Label Public Negative</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicnuclearmembrane\" target=\"_blank\">HPA: Single Label Public Nuclear Membrane</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicaggresome\" target=\"_blank\">HPA: Single Label Public aggresome</a></li>\n</ul>\n<h3>Dataset Version Control and Sanity In-check</h3>\n<p>I used Weights and Biases artifacts to keep my sanity intact. There was a lot of back and forth with the dataset creation process. I used artifacts to keep a track of the <code>.csv</code> files produced. The interactive graph view allowed me to not repeat any step or mess up anything. Here's the graph view of the resulting artifact. (<a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/slcl_full_dataset/fadaa879407c3c7397b0/graph\" target=\"_blank\">Link to my artifacts page</a>)</p>\n<p><img src=\"https://i.imgur.com/yCD7kJm.png\" alt=\"img\"></p>\n<h3>Final Data Distribution</h3>\n<p>Here's the current distribution of labels. Mitotic Spindle is really scarce and would be hard to model against. <br>\n<img src=\"https://i.imgur.com/JmBc0vf.png\" alt=\"img\"></p>\n<h1>Next Steps</h1>\n<ul>\n<li>I need to create clever splits of the dataset to train with. I will share the splits and my strategy for the same once I have them ready.</li>\n<li>Train classifiers and evaluate. </li>\n</ul>",
  "messages": [
    {
      "id": 1257484,
      "postDate": "2021-03-30T22:38:56.727Z",
      "content": "<p>For someday now I was working on preparing a single label (images with only one label associated with it) cell level (individual cells are labeled) dataset to train a better image classifier. Here's a quick summary.</p>\n<h1>Steps</h1>\n<h3>Cell Level Image Dataset Creation</h3>\n<ul>\n<li>Start with the HPA competition image and using the HPA Segmentation tool segment it to get <code>cell_mask</code> and <code>nuclei_mask</code>. </li>\n<li>Using <code>cell_mask</code> get individual cell patches.</li>\n<li>Using `nuclei_mask' get individual nucleus patches. The contour of the nucleus patch is computed. Using this it is determined if the nucleus is bordering the image. I am discarding the cells whose nucleus is bordering the image. <strong>I am assuming that such a cell would not have \"complete\" protein information</strong>. </li>\n<li><strong>Note</strong> that I am creating two separate RGB and protein image datasets. The DSViz table shown in the image can give you an idea of my dataset. (<a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_split_0/34b460a9c6ea265a4e95/files/dataset.table.json\" target=\"_blank\">W&amp;B DSViz table link</a>)<br>\n<img src=\"https://i.imgur.com/fU8hjlb.png\" alt=\"img\"></li>\n</ul>\n<p>Kernels used:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/dataset-versioning-for-single-label-dataset\" target=\"_blank\">Dataset Versioning for Single Label Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation\" target=\"_blank\">Using W&amp;B DSViz for Cell Level Dataset Creation</a></li>\n</ul>\n<p>Kaggle Dataset:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\" target=\"_blank\">https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset</a></li>\n</ul>\n<h3>Data Distribution</h3>\n<p>The image below is the label distribution of the single label cell level dataset.<br>\n<img src=\"https://i.imgur.com/S9BKxvd.png\" alt=\"img\"></p>\n<h3>Extra HPA Public Data</h3>\n<p>Certainly, there is an insane class imbalance. To mitigate the imbalance to some extent I used HPA Pubic Data. I decided to create separate single label cell level datasets for these classes - Nuclear Membrane, Actin Filaments, Aggresome, Mitotic Spindle, and Negative. </p>\n<p>Kernels used:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/cell-level-extra-data-download\" target=\"_blank\">Cell Level Extra Data Download</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation\" target=\"_blank\">Single Label Cell Label Dataset Creation w\\ W&amp;B</a><br>\n(Apologies for not documenting them well.)</li>\n</ul>\n<p>Kaggle Dataset:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicmitoticspindle\" target=\"_blank\">HPA: Single Label Public Mitotic Spindle</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicactinfilaments\" target=\"_blank\">HPA: Single Label Public ACTIN FILAMENTS</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicnegative\" target=\"_blank\">HPA: Single Label Public Negative</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicnuclearmembrane\" target=\"_blank\">HPA: Single Label Public Nuclear Membrane</a></li>\n<li><a href=\"https://www.kaggle.com/ayuraj/singlelabelpublicaggresome\" target=\"_blank\">HPA: Single Label Public aggresome</a></li>\n</ul>\n<h3>Dataset Version Control and Sanity In-check</h3>\n<p>I used Weights and Biases artifacts to keep my sanity intact. There was a lot of back and forth with the dataset creation process. I used artifacts to keep a track of the <code>.csv</code> files produced. The interactive graph view allowed me to not repeat any step or mess up anything. Here's the graph view of the resulting artifact. (<a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/slcl_full_dataset/fadaa879407c3c7397b0/graph\" target=\"_blank\">Link to my artifacts page</a>)</p>\n<p><img src=\"https://i.imgur.com/yCD7kJm.png\" alt=\"img\"></p>\n<h3>Final Data Distribution</h3>\n<p>Here's the current distribution of labels. Mitotic Spindle is really scarce and would be hard to model against. <br>\n<img src=\"https://i.imgur.com/JmBc0vf.png\" alt=\"img\"></p>\n<h1>Next Steps</h1>\n<ul>\n<li>I need to create clever splits of the dataset to train with. I will share the splits and my strategy for the same once I have them ready.</li>\n<li>Train classifiers and evaluate. </li>\n</ul>",
      "rawMarkdown": "For someday now I was working on preparing a single label (images with only one label associated with it) cell level (individual cells are labeled) dataset to train a better image classifier. Here's a quick summary.\n\n# Steps\n\n### Cell Level Image Dataset Creation\n*  Start with the HPA competition image and using the HPA Segmentation tool segment it to get `cell_mask` and `nuclei_mask`. \n* Using `cell_mask` get individual cell patches.\n* Using `nuclei_mask' get individual nucleus patches. The contour of the nucleus patch is computed. Using this it is determined if the nucleus is bordering the image. I am discarding the cells whose nucleus is bordering the image. **I am assuming that such a cell would not have \"complete\" protein information**. \n* **Note** that I am creating two separate RGB and protein image datasets. The DSViz table shown in the image can give you an idea of my dataset. ([W&B DSViz table link](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_split_0/34b460a9c6ea265a4e95/files/dataset.table.json))\n![img](https://i.imgur.com/fU8hjlb.png)\n\nKernels used:\n* [Dataset Versioning for Single Label Dataset](https://www.kaggle.com/ayuraj/dataset-versioning-for-single-label-dataset)\n* [Using W&B DSViz for Cell Level Dataset Creation](https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation)\n\nKaggle Dataset:\n* https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\n\n### Data Distribution\n\nThe image below is the label distribution of the single label cell level dataset.\n![img](https://i.imgur.com/S9BKxvd.png)\n\n### Extra HPA Public Data\nCertainly, there is an insane class imbalance. To mitigate the imbalance to some extent I used HPA Pubic Data. I decided to create separate single label cell level datasets for these classes - Nuclear Membrane, Actin Filaments, Aggresome, Mitotic Spindle, and Negative. \n\nKernels used:\n* [Cell Level Extra Data Download](https://www.kaggle.com/ayuraj/cell-level-extra-data-download)\n* [Single Label Cell Label Dataset Creation w\\ W&B](https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation)\n(Apologies for not documenting them well.)\n\nKaggle Dataset:\n* [HPA: Single Label Public Mitotic Spindle](https://www.kaggle.com/ayuraj/singlelabelpublicmitoticspindle)\n* [HPA: Single Label Public ACTIN FILAMENTS](https://www.kaggle.com/ayuraj/singlelabelpublicactinfilaments)\n* [HPA: Single Label Public Negative](https://www.kaggle.com/ayuraj/singlelabelpublicnegative)\n* [HPA: Single Label Public Nuclear Membrane](https://www.kaggle.com/ayuraj/singlelabelpublicnuclearmembrane)\n* [HPA: Single Label Public aggresome](https://www.kaggle.com/ayuraj/singlelabelpublicaggresome)\n\n### Dataset Version Control and Sanity In-check\nI used Weights and Biases artifacts to keep my sanity intact. There was a lot of back and forth with the dataset creation process. I used artifacts to keep a track of the `.csv` files produced. The interactive graph view allowed me to not repeat any step or mess up anything. Here's the graph view of the resulting artifact. ([Link to my artifacts page](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/slcl_full_dataset/fadaa879407c3c7397b0/graph))\n\n![img](https://i.imgur.com/yCD7kJm.png)\n\n### Final Data Distribution\nHere's the current distribution of labels. Mitotic Spindle is really scarce and would be hard to model against. \n![img](https://i.imgur.com/JmBc0vf.png)\n\n# Next Steps\n* I need to create clever splits of the dataset to train with. I will share the splits and my strategy for the same once I have them ready.\n* Train classifiers and evaluate. ",
      "votes": 26
    },
    {
      "id": 1257998,
      "postDate": "2021-03-31T09:12:21.160Z",
      "content": "<p>Great work on this! And thank you for sharing your hard work with the rest of the community.</p>\n<p>I would like to note a couple of things, just to keep in mind when using this:</p>\n<ul>\n<li><p>HPACellSeg can in some instances accidentally group some cells together so you might want to double check that the single cell image image only contains a single nuclei. It should mostly be fine, especially since you excluded cells that touched the borders of the image.</p></li>\n<li><p>In images with a single label, there can be cells that should be considered negative. It's most common of course to see that the pattern is omnipresent but could be worth double checking.</p></li>\n</ul>\n<p>Maybe this was already sanity checked, but figured it was worth mentioning.</p>",
      "rawMarkdown": "Great work on this! And thank you for sharing your hard work with the rest of the community.\n\nI would like to note a couple of things, just to keep in mind when using this:\n\n- HPACellSeg can in some instances accidentally group some cells together so you might want to double check that the single cell image image only contains a single nuclei. It should mostly be fine, especially since you excluded cells that touched the borders of the image.\n\n- In images with a single label, there can be cells that should be considered negative. It's most common of course to see that the pattern is omnipresent but could be worth double checking.\n\nMaybe this was already sanity checked, but figured it was worth mentioning.",
      "votes": 5,
      "replies": [
        {
          "id": 1258037,
          "postDate": "2021-03-31T09:40:00.527Z",
          "content": "<p>Thank you, Casper. These are great pointers that you mentioned. </p>\n<ul>\n<li>There might be just traces of images with more than one cell. I made sure that each image got only one nucleus. </li>\n<li>For single label image-level images, it was a better assumption to assign each cell to an image-level label. The issue as you pointed out is that some of these cells might be negative class. I will double-check. Thank you for this pointer.</li>\n</ul>",
          "rawMarkdown": "Thank you, Casper. These are great pointers that you mentioned. \n\n* There might be just traces of images with more than one cell. I made sure that each image got only one nucleus. \n* For single label image-level images, it was a better assumption to assign each cell to an image-level label. The issue as you pointed out is that some of these cells might be negative class. I will double-check. Thank you for this pointer.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1267119,
      "postDate": "2021-04-08T10:28:26.317Z",
      "content": "<p>Hi, I'm using the single-label-cell-level artifact csv file from this <a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_cell_level/1552ea5a8576a5b97179/files\" target=\"_blank\">link</a></p>\n<p>Are all the files present in the Kaggle dataset? </p>\n<p>I'm getting empty tensors for quite a few files, and getting proper images for some. It maybe an issue with something else though , I am using GCS path on Kaggle to access it. When I'm loading an image  , it loads as an empty tensor , however when I open it in the dataset , it appears to be fine. Probably an issue in my code. Any idea what could be wrong?</p>",
      "rawMarkdown": "Hi, I'm using the single-label-cell-level artifact csv file from this [link](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_cell_level/1552ea5a8576a5b97179/files)\n\nAre all the files present in the Kaggle dataset? \n\nI'm getting empty tensors for quite a few files, and getting proper images for some. It maybe an issue with something else though , I am using GCS path on Kaggle to access it. When I'm loading an image  , it loads as an empty tensor , however when I open it in the dataset , it appears to be fine. Probably an issue in my code. Any idea what could be wrong?",
      "votes": 1,
      "replies": [
        {
          "id": 1267340,
          "postDate": "2021-04-08T13:29:04.560Z",
          "content": "<p>Update - It is working now ,had some leakage in my code. However , training is taking about 1.5 hrs per epoch on TPU. Maybe because it has ~220k small images?</p>",
          "rawMarkdown": "Update - It is working now ,had some leakage in my code. However , training is taking about 1.5 hrs per epoch on TPU. Maybe because it has ~220k small images?"
        },
        {
          "id": 1267417,
          "postDate": "2021-04-08T14:01:24.967Z",
          "content": "<p>Yeah, it's gonna take a lot of time and thus it would be better if you undersample classes with a lot of labels and train smaller models on different folds of the dataset. </p>\n<p>This is the reason I made available the entire dataset as it is. Since the images have kind of noisy labels it would not hurt to discard images from the majority class. Up to you to come up with strategies. Luck. :)</p>",
          "rawMarkdown": "Yeah, it's gonna take a lot of time and thus it would be better if you undersample classes with a lot of labels and train smaller models on different folds of the dataset. \n\nThis is the reason I made available the entire dataset as it is. Since the images have kind of noisy labels it would not hurt to discard images from the majority class. Up to you to come up with strategies. Luck. :)\n\n",
          "votes": 2
        },
        {
          "id": 1267448,
          "postDate": "2021-04-08T14:18:33.097Z",
          "content": "<p>Right. I was thinking of making TFRecords of all the 220k files and use them, but I think I’ll try undersampling first. Thanks!</p>",
          "rawMarkdown": "Right. I was thinking of making TFRecords of all the 220k files and use them, but I think I’ll try undersampling first. Thanks!"
        }
      ]
    },
    {
      "id": 1257650,
      "postDate": "2021-03-31T02:34:27.263Z",
      "content": "<p>Wonderful work. Thank you for doing all of this and making it available to everyone. Very appreciated!!</p>",
      "rawMarkdown": "Wonderful work. Thank you for doing all of this and making it available to everyone. Very appreciated!!",
      "votes": 2,
      "replies": [
        {
          "id": 1258030,
          "postDate": "2021-03-31T09:35:18.023Z",
          "content": "<p>Thank you, Darien. </p>",
          "rawMarkdown": "Thank you, Darien. "
        }
      ]
    },
    {
      "id": 1267323,
      "postDate": "2021-04-08T13:18:56.687Z",
      "content": "<p>Great work! Thanks a lot! Do you provide any matching between cell and label? May be csv file?</p>",
      "rawMarkdown": "Great work! Thanks a lot! Do you provide any matching between cell and label? May be csv file?",
      "replies": [
        {
          "id": 1267339,
          "postDate": "2021-04-08T13:28:16.270Z",
          "content": "<p>I have made a <a href=\"https://www.kaggle.com/p4rallax/hpasinglelabelcellcsv\" target=\"_blank\">dataset</a> using the csv file for it. Use the file with 'only' in the name :)</p>",
          "rawMarkdown": "I have made a [dataset](https://www.kaggle.com/p4rallax/hpasinglelabelcellcsv) using the csv file for it. Use the file with 'only' in the name :)"
        }
      ]
    },
    {
      "id": 1258390,
      "postDate": "2021-03-31T15:12:11.663Z",
      "content": "<p>Nice work.  Thanks for sharing! <br>\nHowever , the link   :  <a href=\"https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\" target=\"_blank\">https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset</a><br>\ndoesn't seem to work for me. could you please check that?</p>",
      "rawMarkdown": "Nice work.  Thanks for sharing! \nHowever , the link   :  https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\ndoesn't seem to work for me. could you please check that?",
      "replies": [
        {
          "id": 1258505,
          "postDate": "2021-03-31T17:12:56.773Z",
          "content": "<p>Can you check it again? My bad it was private. </p>",
          "rawMarkdown": "Can you check it again? My bad it was private. "
        },
        {
          "id": 1258543,
          "postDate": "2021-03-31T17:43:51.860Z",
          "content": "<p>Working now , thanks !</p>",
          "rawMarkdown": "Working now , thanks !"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1257998,
      "author_name": "Casper Winsnes",
      "author_url": "",
      "post_date": "2021-03-31T09:12:21.160000",
      "content": "<p>Great work on this! And thank you for sharing your hard work with the rest of the community.</p>\n<p>I would like to note a couple of things, just to keep in mind when using this:</p>\n<ul>\n<li><p>HPACellSeg can in some instances accidentally group some cells together so you might want to double check that the single cell image image only contains a single nuclei. It should mostly be fine, especially since you excluded cells that touched the borders of the image.</p></li>\n<li><p>In images with a single label, there can be cells that should be considered negative. It's most common of course to see that the pattern is omnipresent but could be worth double checking.</p></li>\n</ul>\n<p>Maybe this was already sanity checked, but figured it was worth mentioning.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1258037,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-03-31T09:40:00.527000",
          "content": "<p>Thank you, Casper. These are great pointers that you mentioned. </p>\n<ul>\n<li>There might be just traces of images with more than one cell. I made sure that each image got only one nucleus. </li>\n<li>For single label image-level images, it was a better assumption to assign each cell to an image-level label. The issue as you pointed out is that some of these cells might be negative class. I will double-check. Thank you for this pointer.</li>\n</ul>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1267119,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2021-04-08T10:28:26.317000",
      "content": "<p>Hi, I'm using the single-label-cell-level artifact csv file from this <a href=\"https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_cell_level/1552ea5a8576a5b97179/files\" target=\"_blank\">link</a></p>\n<p>Are all the files present in the Kaggle dataset? </p>\n<p>I'm getting empty tensors for quite a few files, and getting proper images for some. It maybe an issue with something else though , I am using GCS path on Kaggle to access it. When I'm loading an image  , it loads as an empty tensor , however when I open it in the dataset , it appears to be fine. Probably an issue in my code. Any idea what could be wrong?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1267340,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-04-08T13:29:04.560000",
          "content": "<p>Update - It is working now ,had some leakage in my code. However , training is taking about 1.5 hrs per epoch on TPU. Maybe because it has ~220k small images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1267417,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-04-08T14:01:24.967000",
          "content": "<p>Yeah, it's gonna take a lot of time and thus it would be better if you undersample classes with a lot of labels and train smaller models on different folds of the dataset. </p>\n<p>This is the reason I made available the entire dataset as it is. Since the images have kind of noisy labels it would not hurt to discard images from the majority class. Up to you to come up with strategies. Luck. :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1267448,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-04-08T14:18:33.097000",
          "content": "<p>Right. I was thinking of making TFRecords of all the 220k files and use them, but I think I’ll try undersampling first. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1257650,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-03-31T02:34:27.263000",
      "content": "<p>Wonderful work. Thank you for doing all of this and making it available to everyone. Very appreciated!!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1258030,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-03-31T09:35:18.023000",
          "content": "<p>Thank you, Darien. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1267323,
      "author_name": "NaStroy",
      "author_url": "",
      "post_date": "2021-04-08T13:18:56.687000",
      "content": "<p>Great work! Thanks a lot! Do you provide any matching between cell and label? May be csv file?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1267339,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-04-08T13:28:16.270000",
          "content": "<p>I have made a <a href=\"https://www.kaggle.com/p4rallax/hpasinglelabelcellcsv\" target=\"_blank\">dataset</a> using the csv file for it. Use the file with 'only' in the name :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258390,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2021-03-31T15:12:11.663000",
      "content": "<p>Nice work.  Thanks for sharing! <br>\nHowever , the link   :  <a href=\"https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\" target=\"_blank\">https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset</a><br>\ndoesn't seem to work for me. could you please check that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1258505,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-03-31T17:12:56.773000",
          "content": "<p>Can you check it again? My bad it was private. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258543,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-03-31T17:43:51.860000",
          "content": "<p>Working now , thanks !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1257484": "For someday now I was working on preparing a single label (images with only one label associated with it) cell level (individual cells are labeled) dataset to train a better image classifier. Here's a quick summary.\n\n# Steps\n\n### Cell Level Image Dataset Creation\n*  Start with the HPA competition image and using the HPA Segmentation tool segment it to get `cell_mask` and `nuclei_mask`. \n* Using `cell_mask` get individual cell patches.\n* Using `nuclei_mask' get individual nucleus patches. The contour of the nucleus patch is computed. Using this it is determined if the nucleus is bordering the image. I am discarding the cells whose nucleus is bordering the image. **I am assuming that such a cell would not have \"complete\" protein information**. \n* **Note** that I am creating two separate RGB and protein image datasets. The DSViz table shown in the image can give you an idea of my dataset. ([W&B DSViz table link](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_split_0/34b460a9c6ea265a4e95/files/dataset.table.json))\n![img](https://i.imgur.com/fU8hjlb.png)\n\nKernels used:\n* [Dataset Versioning for Single Label Dataset](https://www.kaggle.com/ayuraj/dataset-versioning-for-single-label-dataset)\n* [Using W&B DSViz for Cell Level Dataset Creation](https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation)\n\nKaggle Dataset:\n* https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\n\n### Data Distribution\n\nThe image below is the label distribution of the single label cell level dataset.\n![img](https://i.imgur.com/S9BKxvd.png)\n\n### Extra HPA Public Data\nCertainly, there is an insane class imbalance. To mitigate the imbalance to some extent I used HPA Pubic Data. I decided to create separate single label cell level datasets for these classes - Nuclear Membrane, Actin Filaments, Aggresome, Mitotic Spindle, and Negative. \n\nKernels used:\n* [Cell Level Extra Data Download](https://www.kaggle.com/ayuraj/cell-level-extra-data-download)\n* [Single Label Cell Label Dataset Creation w\\ W&B](https://www.kaggle.com/ayuraj/using-w-b-dsviz-for-cell-level-dataset-creation)\n(Apologies for not documenting them well.)\n\nKaggle Dataset:\n* [HPA: Single Label Public Mitotic Spindle](https://www.kaggle.com/ayuraj/singlelabelpublicmitoticspindle)\n* [HPA: Single Label Public ACTIN FILAMENTS](https://www.kaggle.com/ayuraj/singlelabelpublicactinfilaments)\n* [HPA: Single Label Public Negative](https://www.kaggle.com/ayuraj/singlelabelpublicnegative)\n* [HPA: Single Label Public Nuclear Membrane](https://www.kaggle.com/ayuraj/singlelabelpublicnuclearmembrane)\n* [HPA: Single Label Public aggresome](https://www.kaggle.com/ayuraj/singlelabelpublicaggresome)\n\n### Dataset Version Control and Sanity In-check\nI used Weights and Biases artifacts to keep my sanity intact. There was a lot of back and forth with the dataset creation process. I used artifacts to keep a track of the `.csv` files produced. The interactive graph view allowed me to not repeat any step or mess up anything. Here's the graph view of the resulting artifact. ([Link to my artifacts page](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/slcl_full_dataset/fadaa879407c3c7397b0/graph))\n\n![img](https://i.imgur.com/yCD7kJm.png)\n\n### Final Data Distribution\nHere's the current distribution of labels. Mitotic Spindle is really scarce and would be hard to model against. \n![img](https://i.imgur.com/JmBc0vf.png)\n\n# Next Steps\n* I need to create clever splits of the dataset to train with. I will share the splits and my strategy for the same once I have them ready.\n* Train classifiers and evaluate. ",
    "1257998": "Great work on this! And thank you for sharing your hard work with the rest of the community.\n\nI would like to note a couple of things, just to keep in mind when using this:\n\n- HPACellSeg can in some instances accidentally group some cells together so you might want to double check that the single cell image image only contains a single nuclei. It should mostly be fine, especially since you excluded cells that touched the borders of the image.\n\n- In images with a single label, there can be cells that should be considered negative. It's most common of course to see that the pattern is omnipresent but could be worth double checking.\n\nMaybe this was already sanity checked, but figured it was worth mentioning.",
    "1267119": "Hi, I'm using the single-label-cell-level artifact csv file from this [link](https://wandb.ai/ayush-thakur/hpa/artifacts/dataset/single_label_cell_level/1552ea5a8576a5b97179/files)\n\nAre all the files present in the Kaggle dataset? \n\nI'm getting empty tensors for quite a few files, and getting proper images for some. It maybe an issue with something else though , I am using GCS path on Kaggle to access it. When I'm loading an image  , it loads as an empty tensor , however when I open it in the dataset , it appears to be fine. Probably an issue in my code. Any idea what could be wrong?",
    "1257650": "Wonderful work. Thank you for doing all of this and making it available to everyone. Very appreciated!!",
    "1267323": "Great work! Thanks a lot! Do you provide any matching between cell and label? May be csv file?",
    "1258390": "Nice work.  Thanks for sharing! \nHowever , the link   :  https://www.kaggle.com/ayuraj/hpa-single-label-cell-level-dataset\ndoesn't seem to work for me. could you please check that?"
  }
}