{"cells":[{"metadata":{},"cell_type":"markdown","source":"[](http://)# Human Protein Atlas - Single Cell Classification EDA \n\n单细胞分类"},{"metadata":{},"cell_type":"markdown","source":"Great thanks to the original kernel https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda\n\n本文在其基础上增添了部分中文翻译"},{"metadata":{},"cell_type":"markdown","source":"> Cells are the basic structural units of all living organisms. \n> The bustling activity of the human cell is maintained as proteins perform specific tasks in designated compartments, called organelles. \n\n细胞是所有生物体的基本结构单位。人体细胞的活动源于蛋白质在特定的细胞器中执行特定任务。\n\n[\"The Human Cell\"](https://youtu.be/P4gz6DrZOOI) by The Human Protein Atlas\n\nI am using this EDA notebook to share insights about the Human Protein Atlas - Single Cell Classification competition. I have also prepared a separate baseline submission notebook available [here](https://www.kaggle.com/thedrcat/hpa-baseline-cell-segmentation).\n\n### Nice people that enjoy this notebook also press the upvote button :) "},{"metadata":{},"cell_type":"markdown","source":"各类别名称与描述：\n\n| 标签 | 英文名称 | 中文名称 | 描述 |\n|-|-|-|-|\n| 0. | Nucleoplasm | 细胞核（核质） | 细胞核在细胞的中心，并且可以借助于蓝色通道中的信号来识别。核质的染色可以包括整个核或不含核仁(Class 2)的区域. |\n| 1. | Nuclear membrane | 核膜 | 核膜表现为围绕着细胞核的一个薄圆。它并不是完全光滑的，有时在核内也可以看到膜的褶皱为小圆圈或小点。 |\n| 2. | Nucleoli | 核仁 | 核仁可以被看作是核质中略微拉长的圆形区域，通常在蓝色DAPI通道中显示出更弱的染色。核仁的数量和大小因细胞类型而异。 |\n| 3. | Nucleoli fibrillar center | 核仁纤维中心 | 核仁纤维中心可根据细胞类型的不同，在核仁中出现点状簇或单个较大的斑点。 |\n| 4. | Nuclear speckles | 核散斑 | 核散斑是亚核结构。在荧光显微镜水平,他们出现不规则的点状结构 |\n| 5. | Nuclear bodies | 核体 | 核体是核质中可见的明显斑点。它们的形状、大小和数量因体型以及细胞类型而异，但与核散斑相比，通常比较圆润。 |\n| 6. | Endoplasmic reticulum | 内质网 | 内质网(ER)是通过胞浆中的网络状染色来识别的，通常靠近细胞核的染色较强，靠近细胞边缘的染色较弱。ER可以借助黄色ER通道中的染色来识别。 |\n| 7. | Golgi apparatus | 高尔基体 | 高尔基体是一个相当大的细胞器，位于细胞核旁边，靠近中心体，红通道中的微管就来源于此。高尔基体的外观呈折叠带状，但其形状和大小可因细胞类型和细胞的各种过程而变化。 |\n| 8. | Intermediate filaments | 中间丝 | 中间丝经常表现出一种略微纠缠的结构，每隔一段时间就会有股交叉。它们可能与微管相似，但与红色微管通道中的染色不相吻合。中间丝可以延伸到整个胞浆中，或集中在靠近细胞核的区域。 |\n| 9. | Actin filaments | 肌动蛋白丝 | 肌动蛋白丝可以被看作是长而直的丝束，也可以看作是较细的丝的分支网络。它们通常位于细胞边缘附近。 |\n| 10. | Microtubules | 微管 | 微管被看作是延伸到整个细胞中的细线。几乎总是可以检测到它们全部起源的中心（中心体）。与红色通道中的染色重叠。 |\n| 11. | Mitotic spindle | 有丝分裂纺锤体 | 有丝分裂的纺锤体可以被看作是一个错综复杂的微管结构，从分裂细胞（有丝分裂）两端的每个中心体辐射出来。在这个阶段，细胞的染色质被浓缩，通过强烈的DAPI染色可见。在有丝分裂进展过程中，有丝分裂纺锤体的大小和确切形状发生变化，清楚地反映了有丝分裂的不同阶段。 |\n| 12. | Centrosome | 中心体 | 这一类包括中心体和中心粒卫星。它们可以被看作是微管起源处靠近细胞核的一小块区域，具有或多或少的明显染色。当细胞分裂时，两个中心体移动到细胞的两端，形成有丝分裂主轴的两极。 |\n| 13. | Plasma membrane | 细胞质膜 | 这一类包括质膜和细胞连接。两者都在细胞的外缘。浆膜有时在细胞周围出现或多或少的明显边缘，偶尔有特征性的突起或褶皱。在一些细胞系中，染色可在整个细胞内均匀分布。在相邻细胞之间的接触部位可观察到细胞连接。 |\n| 14. | Mitochondria | 线粒体 | 线粒体是细胞质中的小棒状单位，常沿微管呈线状分布。 |\n| 15. | Aggresome | 聚集体/异常蛋白包涵体/聚集小体/侵袭体 | 侵袭体可以被看作是一种致密的胞质包涵体，它通常被发现靠近细胞核，在微管网络被破坏的区域。 |\n| 16. | Cytosol | 细胞质基质/胞浆 | 胞浆从浆膜延伸至核膜。它可以是光滑的或颗粒状的，靠近核的地方染色通常较强。 |\n| 17. | Vesicles and punctate cytosolic patterns | 囊泡和点状细胞形态 | 这一类包括细胞质中的小圆室。囊泡、过氧体(脂质代谢)、内膜体(分拣室)、溶酶体(分子降解或吞噬死亡分子)、脂滴(脂肪储存)、胞质体(胞浆中的独特颗粒)。它们是高度动态的，在数量和大小上随环境和细胞线索而变化。它们可以是圆形的，也可以是更加细长的。 |\n| 18. | Negative | 其他 | 这一类包括阴性染色和非特异性图案。这意味着细胞没有绿色染色（阴性），或有染色但不能从染色中解读出图案（非特异性）。 |"},{"metadata":{},"cell_type":"markdown","source":"# Contents\n\n1. [Data](#data)\n    * Training and submission files 训练与提交\n    * Labels distribution 标签分布\n    * Combinations of classes 类别间关系\n2. [Images](#images)\n    * Channels 图像通道\n    * Combining Channels 通道组合\n    * Visualize Single Classes 可视化\n3. [Metric](#metric) 评价标准\n4. [Competition Notes](#notes) \n5. [Solution Approaches](#solutions)\n6. [Resources](#resources)\n    * Datasets 数据集"},{"metadata":{},"cell_type":"markdown","source":"# 1. Data <a class=\"anchor\" id=\"data\"></a>\n\nI will start by exploring the training data, including images and the label distribution. More to come as I start studying the domain. "},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"!pip install ipyplot -q\n\nimport pandas as pd\nimport numpy as np\nfrom fastai.vision.all import *\nimport ipyplot\nimport imageio\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\npath = Path('../input/hpa-single-cell-image-classification')\ndf_train = pd.read_csv(path/'train.csv')\ndf_sub = pd.read_csv(path/'sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1.1. Training and submission files"},{"metadata":{},"cell_type":"markdown","source":"Let's start by looking into the training csv. We get image id and label column, which contains the classes corresponding to each image. Looking at this data, it appears to be a multi-label classification challenge..."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now looking at the sample submission, we realize this is actually not a multi-label classification, but instead an instance segmentation task. For each image, we are being asked to segment each single cell contained in the image, identify the class of that cell, and submit a string containing:\n1. The class of each cell\n2. Confidence of the class prediction\n3. Segmentation mask for each cell, encoded in a fancy scheme\n\nThe fact that our training data is image-level classes, and our task is to predict cell masks and their corresponding classes, explains why the organizers call it a *weakly supervised multi-label classification problem*.\nAlso, the following comments from the organizers are important: \n- some cells can be associated with multiple classes - in that case, we need to predict separate detections for each class using the same mask\n\nBased on an early comment by the host, I had initially assumed that confidence is not relevant and can always be set to 1. It was clarified in [this discussion](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885) that this is not the case. You can check out [this excellent kernel](https://www.kaggle.com/its7171/map-understanding-with-code-and-its-tips) to understand how the competition metric works and how the confidence score is relevant. \n\nFor a demonstration how to create the prediction string encoding, see my [baseline kernel](https://www.kaggle.com/thedrcat/hpa-baseline-cell-segmentation)..\n\n所谓弱监督学习，是指训练数据中，只提供了图像级别的标注（每一个图像可能对应多个类别，其中包括多个细胞，每个细胞也可能对应多个类别），而推理时，需要实现语义分割的功能。分割可以通过HPACellSegmentor这个包来实现，所以我们至少需要实现图像中，每个细胞的分类，所谓 Single Cell Classification。"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_path = path/'train'\ntrain_images = train_path.ls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(df_train), len(train_images), len(train_images)/len(df_train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can also see there are 4 images in the train folder for each image ID, corresponding to red, blue, green and yellow channels. Here's the official description:\n> Each sample consists of four files. Each file represents a different filter on the subcellular protein patterns represented by the sample. The format should be [filename]_[filter color].png for the PNG files. Colors are red for microtubule channels, blue for nuclei channels, yellow for Endoplasmic Reticulum (ER) channels, and green for protein of interest.\n\n每个图像有**4**个通道，对应4张png图像，文件名为[filename]_[filter color].png。R表示微管，B表示细胞核，Y表示内质网，G表示关注的蛋白质"},{"metadata":{},"cell_type":"markdown","source":"## 1.2. Labels distribution\n\nLet's first check how many unique labels we have in the dataset."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"all_labels = df_train.Label.unique().tolist()\nall_labels = '|'.join(all_labels)\nall_labels = all_labels.split('|')\nall_labels = list(set(all_labels))\nnum_unique_labels = len(all_labels)\nall_labels = sorted(all_labels, key=int)\nall_labels = ' '.join(all_labels)\nprint(f'{num_unique_labels} unique labels, values: {all_labels}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Ok, so we have 19 unique labels from 0 to 18. What is the distribution of these labels? I wonder especially how many images have only one label (unique counts) and what is the overall number of images for each label? Looking at the chart below, we can see it's an imbalanced dataset and labels 11 and 18 are going to be especially problematic. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"labels = [str(i) for i in range(19)]\n\nunique_counts = {}\nfor lbl in labels:\n    unique_counts[lbl] = len(df_train[df_train.Label == lbl])\n\nfull_counts = {}\nfor lbl in labels:\n    count = 0\n    for row_label in df_train['Label']:\n        if lbl in row_label.split('|'): count += 1\n    full_counts[lbl] = count\n    \ncounts = list(zip(full_counts.keys(), full_counts.values(), unique_counts.values()))\ncounts = np.array(sorted(counts, key=lambda x:-x[1]))\ncounts = pd.DataFrame(counts, columns=['label', 'full_count', 'unique_count'])\n\nsns.set(style=\"whitegrid\")\nf, ax = plt.subplots(figsize=(16, 12))\n\nsns.set_color_codes(\"pastel\")\nsns.barplot(x=\"full_count\", y=\"label\", data=counts, order=counts.label.values,\n            label=\"full count\", color=\"b\", orient = 'h')\n\n# Plot the crashes where alcohol was involved\nsns.set_color_codes(\"muted\")\nsns.barplot(x=\"unique_count\", y=\"label\", data=counts, order=counts.label.values,\n            label=\"unique count\", color=\"b\", orient = 'h')\n\n# Add a legend and informative axis label\nax.legend(ncol=2, loc=\"lower right\", frameon=True)\nax.set(ylabel=\"\",\n       xlabel=\"Counts\")\nsns.despine(left=True, bottom=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What do these labels represent? The descriptions below are taken from [this notebook shared by the organizers](https://www.kaggle.com/lnhtrang/single-cell-patterns).\n\n| Label | Name | Description |\n|-|-|-|\n| 0. | Nucleoplasm | The nucleus is found in the center of cell and can be identified with the help of the signal in the blue nucleus channel. A staining of the nucleoplasm may include the whole nucleus or of the nucleus without the regions known as nucleoli (Class 2). |\n| 1. | Nuclear membrane | The nuclear membrane appears as a thin circle around the nucleus. It is not perfectly smooth and sometimes it is also possible to see the folds of the membrane as small circles or dots inside the nucleus. |\n| 2. | Nucleoli | Nucleoli can be seen as slightly elongated circular areas in the nucleoplasm, which usually display a much weaker staining in the blue DAPI channel. The number and size of nucleoli varies between cell types. |\n| 3. | Nucleoli fibrillar center | Nucleoli fibrillary center can appear as a spotty cluster or as a single bigger spot in the nucleolus, depending on the cell type. |\n| 4. | Nuclear speckles | Nuclear speckles can be seen as irregular and mottled spots inside the nucleoplasm. |\n| 5. | Nuclear bodies | Nuclear bodies are visible as distinct spots in the nucleoplasm. They vary in shape, size and numbers depending on the type of bodies as well as cell type, but are usually more rounded compared to nuclear speckles. |\n| 6. | Endoplasmic reticulum | The endoplasmic reticulum (ER) is recognized by a network-like staining in the cytosol, which is usually stronger close to the nucleus and weaker close to the edges of the cell. The ER can be identified with the help of the staining in the yellow ER channel. |\n| 7. | Golgi apparatus | The Golgi apparatus is a rather large organelle that is located next to the nucleus, close to the centrosome, from which the microtubules in the red channel originate. It has a folded ribbon-like appearance, but the shape and size can vary between cell types, and in response to cellular various processes. |\n| 8. | Intermediate filaments | Intermediate filaments often exhibit a slightly tangled structure with strands crossing every so often. They can appear similar to microtubules, but do not match well with the staining in the red microtubule channel. Intermediate filaments may extend through the whole cytosol, or be concentrated in an area close to the nucleus. |\n| 9. | Actin filaments | Actin filaments can be seen as long and rather straight bundles of filaments or as branched networks of thinner filaments. They are usually located close to the edges of the cells. |\n| 10. | Microtubules | Microtubules are seen as thin strands that stretch throughout the whole cell. It is almost always possible to detect the center from which they all originate (the centrosome). And yes, as you might have guessed, this overlaps the staining in the red channel. |\n| 11. | Mitotic spindle | The mitotic spindle can be seen as an intricate structure of microtubules radiating from each of the centrosomes at opposite ends of a dividing cell (mitosis). At this stage, the chromatin of the cell is condensed, as visible by intense DAPI staining. The size and exact shape of the mitotic spindle changes during mitotic progression, clearly reflecting the different stages of mitosis. |\n| 12. | Centrosome | This class includes centrosomes and centriolar satellites. They can be seen as a more or less distinct staining of a small area at the origin of the microtubules, close to the nucleus. When a cell is dividing, the two centrosomes move to opposite ends of the cell and form the poles of the mitotic spindle. |\n| 13. | Plasma membrane | This class includes plasma membrane and cell junctions. Both are at the outer edge of the cell. Plasma membrane sometimes appears as a more or less distinct edge around the cell, occasionally with characteristic protrusions or ruffles. In some cell lines, the staining can be uniform across the entire cell. Cell junctions can be observed at contact sites between neighboring cells. |\n| 14. | Mitochondria | Mitochondria are small rod-like units in the cytosol, which are often distributed in a thread-like pattern along microtubules. |\n| 15. | Aggresome | An aggresome can be seen as a dense cytoplasmic inclusion, which is usually found close to the nucleus, in a region where the microtubule network is disrupted. |\n| 16. | Cytosol | The cytosol extends from the plasma membrane to the nuclear membrane. It can appear smooth or granular, and the staining is often stronger close to the nucleus. |\n| 17. | Vesicles and punctate cytosolic patterns | This class includes small circular compartments in the cytosol: Vesicles, Peroxisomes (lipid metabolism), Endosomes (sorting compartments), Lysosomes (degradation of molecules or eating up dead molecules), Lipid droplets (fat storage), Cytoplasmic bodies (distinct granules in the cytosol). They are highly dynamic, varying in numbers and size in response to environmental and cellular cues. They can be round or more elongated. |\n| 18. | Negative | This class include negative stainings and unspecific patterns. This means that the cells have no green staining (negative), or have staining but no pattern can be deciphered from the staining (unspecific). |"},{"metadata":{},"cell_type":"markdown","source":"| 标签 | 英文名称 | 中文名称 | 描述 |\n|-|-|-|-|\n| 0. | Nucleoplasm | 细胞核（核质） | 细胞核在细胞的中心，并且可以借助于蓝色通道中的信号来识别。核质的染色可以包括整个核或不含核仁(Class 2)的区域. |\n| 1. | Nuclear membrane | 核膜 | 核膜表现为围绕着细胞核的一个薄圆。它并不是完全光滑的，有时在核内也可以看到膜的褶皱为小圆圈或小点。 |\n| 2. | Nucleoli | 核仁 | 核仁可以被看作是核质中略微拉长的圆形区域，通常在蓝色DAPI通道中显示出更弱的染色。核仁的数量和大小因细胞类型而异。 |\n| 3. | Nucleoli fibrillar center | 核仁纤维中心 | 核仁纤维中心可根据细胞类型的不同，在核仁中出现点状簇或单个较大的斑点。 |\n| 4. | Nuclear speckles | 核散斑 | 核散斑是亚核结构。在荧光显微镜水平,他们出现不规则的点状结构 |\n| 5. | Nuclear bodies | 核体 | 核体是核质中可见的明显斑点。它们的形状、大小和数量因体型以及细胞类型而异，但与核散斑相比，通常比较圆润。 |\n| 6. | Endoplasmic reticulum | 内质网 | 内质网(ER)是通过胞浆中的网络状染色来识别的，通常靠近细胞核的染色较强，靠近细胞边缘的染色较弱。ER可以借助黄色ER通道中的染色来识别。 |\n| 7. | Golgi apparatus | 高尔基体 | 高尔基体是一个相当大的细胞器，位于细胞核旁边，靠近中心体，红通道中的微管就来源于此。高尔基体的外观呈折叠带状，但其形状和大小可因细胞类型和细胞的各种过程而变化。 |\n| 8. | Intermediate filaments | 中间丝 | 中间丝经常表现出一种略微纠缠的结构，每隔一段时间就会有股交叉。它们可能与微管相似，但与红色微管通道中的染色不相吻合。中间丝可以延伸到整个胞浆中，或集中在靠近细胞核的区域。 |\n| 9. | Actin filaments | 肌动蛋白丝 | 肌动蛋白丝可以被看作是长而直的丝束，也可以看作是较细的丝的分支网络。它们通常位于细胞边缘附近。 |\n| 10. | Microtubules | 微管 | 微管被看作是延伸到整个细胞中的细线。几乎总是可以检测到它们全部起源的中心（中心体）。与红色通道中的染色重叠。 |\n| 11. | Mitotic spindle | 有丝分裂纺锤体 | 有丝分裂的纺锤体可以被看作是一个错综复杂的微管结构，从分裂细胞（有丝分裂）两端的每个中心体辐射出来。在这个阶段，细胞的染色质被浓缩，通过强烈的DAPI染色可见。在有丝分裂进展过程中，有丝分裂纺锤体的大小和确切形状发生变化，清楚地反映了有丝分裂的不同阶段。 |\n| 12. | Centrosome | 中心体 | 这一类包括中心体和中心粒卫星。它们可以被看作是微管起源处靠近细胞核的一小块区域，具有或多或少的明显染色。当细胞分裂时，两个中心体移动到细胞的两端，形成有丝分裂主轴的两极。 |\n| 13. | Plasma membrane | 细胞质膜 | 这一类包括质膜和细胞连接。两者都在细胞的外缘。浆膜有时在细胞周围出现或多或少的明显边缘，偶尔有特征性的突起或褶皱。在一些细胞系中，染色可在整个细胞内均匀分布。在相邻细胞之间的接触部位可观察到细胞连接。 |\n| 14. | Mitochondria | 线粒体 | 线粒体是细胞质中的小棒状单位，常沿微管呈线状分布。 |\n| 15. | Aggresome | 聚集体/异常蛋白包涵体/聚集小体/侵袭体 | 侵袭体可以被看作是一种致密的胞质包涵体，它通常被发现靠近细胞核，在微管网络被破坏的区域。 |\n| 16. | Cytosol | 细胞质基质/胞浆 | 胞浆从浆膜延伸至核膜。它可以是光滑的或颗粒状的，靠近核的地方染色通常较强。 |\n| 17. | Vesicles and punctate cytosolic patterns | 囊泡和点状细胞形态 | 这一类包括细胞质中的小圆室。囊泡、过氧体(脂质代谢)、内膜体(分拣室)、溶酶体(分子降解或吞噬死亡分子)、脂滴(脂肪储存)、胞质体(胞浆中的独特颗粒)。它们是高度动态的，在数量和大小上随环境和细胞线索而变化。它们可以是圆形的，也可以是更加细长的。 |\n| 18. | Negative | 其他 | 这一类包括阴性染色和非特异性图案。这意味着细胞没有绿色染色（阴性），或有染色但不能从染色中解读出图案（非特异性）。 |"},{"metadata":{},"cell_type":"markdown","source":"Ok, so the class 18 which we found with very few examples is actually the \"everything else\" category - this can influence our labeling strategy. \n\n## 1.3. Combinations of classes\n\nWe already have a feeling for the volume of images with a single class. What do the combinations of these classes look like? "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f'Number of unique combinations of classes in the train set: {len(df_train.Label.value_counts())}')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train['num_classes'] = df_train['Label'].apply(lambda r: len(r.split('|')))\ndf_train['num_classes'].value_counts().plot.bar(title='Examples with multiple labels', xlabel='number of labels per example', ylabel='# train examples')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"for x in labels: df_train[x] = df_train['Label'].apply(lambda r: int(x in r.split('|')))\n\ndf_labels = df_train[labels]\ncoocc = df_labels.T.dot(df_labels)\nfor i in range(19): coocc.iloc[i,i] = int(counts.unique_count[counts.label == str(i)].values[0])\n\nfig, ax = plt.subplots(figsize=(16, 13))\nsns.heatmap(coocc, cmap=\"Blues\", linewidth=0.3, cbar_kws={\"shrink\": .8})\ntitle = 'How often do individual classes cooccur in the train set?\\n(Single class occurrences are on the diagonal)\\n'\nplt.title(title, loc='center', fontsize=18)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see from the coocurrence plot that the classes 0 (nucleoplasm) and 16 (cytosol) tend to cooccur most frequently with other classes, and especially with each other. These are also the most frequent classes."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# # I'll try another approach where we normalize the frequencies and compare them to random distribution based on class weights. \n# label_probs = counts.full_count.values.astype(int)/counts.full_count.values.astype(int).sum()\n# label_probs = np.expand_dims(label_probs, axis=0)\n# lprobs = label_probs.T @ label_probs\n# df_probs = pd.DataFrame(lprobs, columns=labels)\n\n# coocc_new = coocc.copy().astype(float)\n# coocc_new = coocc_new.reset_index(drop=True)\n# for i, row in coocc_new.iterrows():\n# #     total = row.values.sum()\n#     for j in range(len(row)):\n#         row[j] /= len(df_train)\n\n# norm_cooc = coocc_new - df_probs\n\n# fig, ax = plt.subplots(figsize=(16, 13))\n# sns.heatmap(norm_cooc, cmap=\"Blues\", linewidth=0.3, cbar_kws={\"shrink\": .8})\n# title = 'Actual vs expected class coocurrence\\n'\n# plt.title(title, loc='center', fontsize=18)\n# plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. Images <a class=\"anchor\" id=\"images\"></a>\n\n## 2.1. Channels"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"mt = [path/'train'/(id+'_red.png') for id in df_train.ID.values]\ner = [path/'train'/(id+'_yellow.png') for id in df_train.ID.values]\nnu = [path/'train'/(id+'_blue.png') for id in df_train.ID.values]\npr = [path/'train'/(id+'_green.png') for id in df_train.ID.values]\nimages = [mt, er, nu, pr]\ntitles = ['microtubules', 'endoplasmic reticulum', 'nucleus', 'protein of interest']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's again try to understand the image data provided, and see if we can visualize it. All images have 4 channels: \n- Red (Microtubules)\n- Green (Protein of interest)\n- Blue (Nucleus)\n- Yellow (Endoplasmic reticulum)\n\nWhat does [Wikipedia](https://en.wikipedia.org/wiki/Cytoskeleton) tell us these channels represent?\n- **Microtubules** are polymers of tubulin that form part of the cytoskeleton and provide structure and shape to eukaryotic cells. Microtubules can grow as long as 50 micrometres and are highly dynamic. The microtubule cytoskeleton is involved in the transport of material within cells, carried out by motor proteins that move on the surface of the microtubule. They are involved in maintaining the structure of the cell and, together with microfilaments and intermediate filaments, they form the cytoskeleton...\n- **Nucleus** is a membrane-bound organelle found in eukaryotic cells. Eukaryotes usually have a single nucleus, but a few cell types, such as mammalian red blood cells, have no nuclei, and a few others including osteoclasts have many. The cell nucleus contains all of the cell's genome, except for the small amount of mitochondrial DNA. The nucleus maintains the integrity of genes and controls the activities of the cell by regulating gene expression. The nucleus is, therefore, the control center of the cell.\n- The **endoplasmic reticulum** (ER) is, in essence, the transportation system of the eukaryotic cell, and has many other important functions such as protein folding. It is a type of organelle made up of two subunits – rough endoplasmic reticulum (RER), and smooth endoplasmic reticulum (SER). The endoplasmic reticulum is found in most eukaryotic cells and forms an interconnected network of flattened, membrane-enclosed sacs known as cisternae (in the RER), and tubular structures in the SER.\n- **Protein of interest** - I still need to understand this better, I assume the researchers are interested in certain types of proteins, which are not disclosed, and these proteins are marked in this channel.（其实结合类别名称，可以比较清楚的知道关注的蛋白质是什么意思）\n\nLet's see an example of each channel below. \n\n\nR表示微管，B表示细胞核，Y表示内质网，G表示关注的蛋白质"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(3,4, figsize=(20,12))\nfor j in range(3):\n    for i in range(4):\n        img = plt.imread(images[i][j])\n        if j == 0: ax[j,i].set_title(titles[i])\n        ax[j,i].imshow(img)\n        ax[j,i].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 2.2. Combining Channels\n\nThese are 4 channels, how do we display them in a single image? According to a comment from competition host, [we could blend them](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214603). Looking at the notebook shared by the hosts, they visualize using 3 channels: red, yellow, blue. Let's see how it looks. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(20,50))\nfor i in range(3):\n    microtubule = plt.imread(mt[i])    \n    endoplasmicrec = plt.imread(er[i])    \n    nuclei = plt.imread(nu[i])\n    img = np.dstack((microtubule, endoplasmicrec, nuclei))\n    ax[i].imshow(img)\n    ax[i].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What if we visualize the red, green and blue channels instead? "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize=(20,50))\nfor i in range(3):\n    microtubule = plt.imread(mt[i])    \n    protein = plt.imread(pr[i])    \n    nuclei = plt.imread(nu[i])\n    img = np.dstack((microtubule, protein, nuclei))\n    ax[i].imshow(img)\n    ax[i].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Ok, these look similar. I'll explore the blending later on, for now let's stick with the organizers' approach.\n\n## 2.3. Visualize single classes\n\nIn the following section I'd like to visualize an image for each class (only images representing single class)."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"class_images = []\nfor label in labels:\n    r_img = df_train[df_train.Label == label].reset_index(drop=True).ID.loc[0] + '_red.png'\n    y_img = df_train[df_train.Label == label].reset_index(drop=True).ID.loc[0] + '_yellow.png'\n    b_img = df_train[df_train.Label == label].reset_index(drop=True).ID.loc[0] + '_blue.png'\n    r = imageio.imread(path/'train'/r_img)\n    y = imageio.imread(path/'train'/y_img)\n    b = imageio.imread(path/'train'/b_img)\n    rgb = np.dstack((r,y,b))\n    class_images.append(PILImage.create(rgb))\n\ncodes = [\n'0. Nucleoplasm',\n'1. Nuclear membrane',\n'2. Nucleoli',\n'3. Nucleoli fibrillar center',\n'4. Nuclear speckles',\n'5. Nuclear bodies',\n'6. Endoplasmic reticulum',\n'7. Golgi apparatus',\n'8. Intermediate filaments',\n'9. Actin filaments',\n'10. Microtubules',\n'11. Mitotic spindle',\n'12. Centrosome',\n'13. Plasma membrane',\n'14. Mitochondria',\n'15. Aggresome',\n'16. Cytosol',\n'17. Vesicles and punctate cytosolic patterns',\n'18. Negative'\n]\n\nipyplot.plot_images(images=class_images, labels=codes, max_images=19, img_width=300)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. Metric <a class=\"anchor\" id=\"metric\"></a>\n\nPlaceholder given the info on metric change.\n\n> The metrics works like this: For each image, all predicted masks are matched with ground-truth masks. Then mAP is calculated. If the model predict a class 0 wrongly for cell 1 in image A, this is counted as 1 False Positive for class 0. If the model predicts a class on a region of the image that is not a cell, this is also counted as 1 False Positive for that class. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214659) 感觉是通过IOU判断是否命中，然后判断每个细胞的类别"},{"metadata":{},"cell_type":"markdown","source":"Here's a link to [some tips on mAP from @its7171](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217158):\n>* The lower the confidence threshold, the higher the mAP we will get. 置信度阈值越低，mAP越高？\n>* Class dependence is linear. \n>* small sample class is equaly important as bit class. 小样本的类别和大样本的类别***同样重要***"},{"metadata":{},"cell_type":"markdown","source":"# 4. Competition Notes <a class=\"anchor\" id=\"notes\"></a>\n\nHere is a collection of notes from various notebooks and forum posts\n> The same type of data was used in a previous [Human Protein Atlas Classification challenge](https://www.kaggle.com/c/human-protein-atlas-image-classification) that aimed to classify image level labels ([paper](https://www.nature.com/articles/s41592-019-0658-6.epdf?no_publisher_access=1&r3_referer=nature)). The results are published and you can read the discussion on the strategies of the winning teams. In the hidden test set of this challenge, we purposely chose images with high single cell variation (SCV). Therefore I believe you won’t gain significant advantages with a metric learning approach (like the winning solution last challenge). [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 上次的比赛是图像级的分类，而这次是细胞级的分类，而且这次的图像中，每张图片中细胞的多样性更大，尤其是在测试集中，所以图像级的分类可能会有较大的问题，Metric learning方法可能不会有太好的效果？\n\n> HPACellSegmentation was used as a basis followed my manual correction of the masks for the ground truth. Classes per single cell were subsequently assigned by annotators. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214616) 测试集的ground truth来自于HPACellSegmentation模型的分割，然后人工做了修正，每个细胞的类别由人工标注，\n\n> We cannot tell you exactly what the proteins (of interest) are now. All we can say is that in this competition a large part of all human proteins are included in either the train or test set. These proteins are involved in a variety of biological processes. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214616) proteins (of interest)具体是啥，目前还不能透露\n\n> The dataset is acquired in a highly standardized way using one imaging modality (confocal microscopy). However, the dataset comprises 17 different cell types of highly different morphology, which affect the protein patterns of the different organelles. All image samples are represented by four filters (stored as individual files), the protein of interest (green) plus three cellular landmarks: nucleus (blue), microtubules (red), endoplasmic reticulum (yellow). The green filter should hence be used to predict the label, and the other filters are used as references. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/data) 图像由激光扫描共聚焦显微镜生成，proteins (of interest)用于预测类别，而其他三个通道用于描述细胞\n\n> The labels you will get for training are image level labels while the task is to predict cell level labels. That is to say, each training image contains a number of cells that have collectively been labeled as described above and the prediction task is to look at images of the same type and predict the labels of each individual cell within those images. As the training labels are a collective label for all the cells in an image, it means that each labeled pattern can be seen in the image but not necessarily that each cell within the image expresses the pattern. This imprecise labeling is what we refer to as weak. During the challenge you will both need to segment the cells in the images and predict the labels of those segmented cells. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/data) 所谓弱监督学习，单张训练图像中，可能存在某种标注的pattern，但是其中所有的细胞不一定都表达了这种pattern，而推理时，我们要预测出每个细胞的pattern\n\n> Note that a mask MAY have more than one class. If that is the case, predict separate detections for each class using the same mask. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation) 每个细胞可能有不止一个类别\n\n> we get the absolute X and Y coordinates for the proteins from the green channel but that information is useless without the context provided by the other channels. The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. For example, if there is overlap between the red channel (Microtubules) and the green (protein) we can understand that the protein localizes to the microtubules. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215981) 通过绿色通道去判断细胞的类别，如果绿色和红色重叠，那就是微管。\n\n> Note that Negative cells will be present in images labeled as other classes as well. This means that, for example, if image A has image-level label 0 1, not all single cells inside this image will have a pattern and some cells can have label 18 (Negative). [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) “其他”这个类别可能在别的图像中也存在，比方说一个图像标注是0和1，那某个细胞可能是“其他”\n\n> The labels are the organelles/structures in which the proteins are located. So you are correct in assuming that the labels only apply to the cells where green is present (as the green is the protein). If there are 4 cells and 3 have green staining and the image-level labels are Mitochondria, and Nucleoplasm.\nCell 1: Green looks to be in the Mitochondria. Therefore, the cell level is Mitochondria.\nCell 2: Green looks to be in the Nucleoplasm. Therefore, the cell level is Nucleoplasm.\nCell 3: Green looks to be in the Nucleoplasm and Mitochondria. Therefore, the cell level label is Nucleoplasm and Mitochondria\nCell 4: No green or green is not present in any organelle. Therefore, the cell level label is Negative. [comment by @dschettler8845](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215736) 类别是绿色通道所在的位置，如果绿色在线粒体，那就是线粒体。\n\n> Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it! [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 如果图像边缘的细胞被遮挡超过一半，就不预测。\n\n> For some images, all cells will inherit the labels from image-level labels. For example, the first image in this section has image-level label Nucleoplasm & Mitochrondria, and all the single cells in this image display both patterns. For other images (especially in the test set), each cell will have only a part of the image-level label, or completely new label (18 Negative). For example, the middle image in this section has image-level label ER & Golgi, but the single cells in this image have either only ER or only Golgi (part of image-level label). And the first image in this section has image-level label Nucleoplasm, Cytokinetic bridge and Mitotic spindle. However, many individual cells only have Nucleoplasm as label, and many don’t have green staining, hence they are Negative. [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 单张训练图像中，可能存在某种标注的pattern，但是其中所有的细胞不一定都表达了这种pattern\n\n> Just like people, all cells are different in terms of their morphology (shape, size) and they can also differ in the spatial distribution of their proteins. This is called single cell variability (SCV), and has been observed across images in HPA and been described in numerous scientific publications. The phenomenon of how genetically identical cells can show functional heterogeneity (for example in response to drug treatment), is poorly understood. We want to study this observed variation to better understand the functional heterogeneity of cells. Therefore a single cell classifier is more helpful than a whole image classifier (as developed in the previous challenge). [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 虽然基因相同，但不同细胞之间依然可以表现出各种形态上的差异。这称为SCV。\n\n> A cell line is, like your explanation says, a population of cells that can be grown for a long time. We grow cells in our lab and take some of them for each imaging experiment. This means that we will have multiple images from the same cell line. It does not mean that we have the same cells between images, as different cells will be sampled for experiments. \nAs for if characteristics of the cell line would improve the classification accuracy, that is certainly an interesting theory. In one of our previous papers (https://www.nature.com/articles/nbt.4225), we did show that including different cell lines in our training data helped accuracy across the board (Figure 5C of that paper). This could mean cell line information may be helpful for a neural network, but we have not tested, much less proved, that theory. If you think it will be helpful, feel free to try it. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215981) 细胞系是指一个培育的细胞种群，之前的实验证明在训练集包含多个细胞系可以提升效果。或许在训练本问题中也有一定的帮助。\n\n> In total, there are 19 classes (0-18) that will be scored. These patterns sometimes vary among 17 cell lines in this competition: A-431, A549, EFO-21, HAP1, HEK 293, HUVEC TERT2, HaCaT, HeLa, PC-3, RH-30, RPTEC TERT1, SH-SY5Y, SK-MEL-30, SiHa, U-2 OS, U-251 MG, hTCEpi. [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 本次数据涉及以上17个细胞系。\n\n> Single or muliple-patterns can be seen together in a cell. Any combination of patterns can be possible for a single cell, except 18 (if the cell's staining is Negative, there isn't any other pattern). [link](https://www.kaggle.com/lnhtrang/single-cell-patterns) 每个细胞可能不止一个pattern，但是18是例外，因为18表示识别到任何pattern。\n\n> the red channel will help you to identify a microtubule pattern (label 10) in the green channel, and the yellow channel will help you identify an ER pattern (label 6) in the green channel. The reference channels (red/green/blue) can also help you to get an overview of the cell - certain patterns will always be distributed in a certain way in correlation to these. For example, labels 2 to 5 are all within the nucleus (i.e. the region of each cell with a signal in the blue channel). [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215736) R通道可以于识别微管，Y可以用于识别ER，结合RGB可以获得关于整个细胞的全局信息。\n\n> All images have all the four channels, and signals from the markers (blue, yellow, red) are present in all cells in the image, independent of the green channel that you are classifying, in order to help you identify where the cells are, as well as where certain structures and regions within the cells are. This can in turn help you to segment the cells and to classify each cell to one or more label(s) according to the signal in the green channel. Note that each of the image-level labels can be present in all or in just a fraction of the individual cells in the image, and that some individual cells may also have additional labels, that are not mentioned among the image-level labels. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215736) byr通道在所有细胞中都存在，且独立于绿色通道。\n\n> there may be images where a single cell have a label that was neglected at the image level (for example a weak cytosolic pattern). Managing such weak-labels is part of this challenge, and it's impact is something we will analyze at the end of the challenge. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215736) 有些图像中，可能有些类别在细胞中存在，但是没有标注在整个图像中。\n\n> the images in the hidden test set are 16-bit. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215428) 测试集的图像是16位的，而训练集是8位的\n\n> Every cell in an image will have one or multiple labels. If there is no specific green pattern (noisy or no signal) it should be labeled as negative. Note that border cells are only included as ground truth when there is enough information to decide on the label (as described in the data tab, and the single cell pattern notebook), i.e. when the majority of the cell area is represented in the image.\nA cell can have multiple labels, but will never have Negative AND some-other-label. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/216924) \n\n> I agree the Negative label is a bit unclear, and I too hope for answers of the above. Also, in the training data, can we consider this category as a \"garbage term\" or a \"none of the above\". If so, wouldn't it be equivalent to a category vector consisting only of zeros (or False or Null etc). Or is there some subtle distinction I don't see? >>>> Yes you are correct, it is a \"none of the above\" label, and can be treated as such. For practical scoring reasons on the Kaggle platform we're using the format where 'Negative' has a label of its own (18). [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/216924)\n\n> In this competition, you should submit the whole cell segmentation. From a research perspective, individual organelle segmentation is definitely an interesting task. But since they are morphologically so different as you see in the Single Cell Pattern notebook, and so are their functions, it is more efficient to develop and analyze a model for each organelle of interest. In fact this is what happens in the biological image analysis space right now with individual deep learning models for Mitochondria, Golgi etc. There are more considerations that we took into account when designing this challenge. We want to study the integrated cell, with organelles in relation to it. Cell segmentation task can be considered an auxiliary task (depending on your approach to modelling of course) on a generalized deep learning model of the cell. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309)\n\n> It appears that 146 of the images contain duplicate classes (just one label duplicated) in their labels. This duplication only happens with the 9, 13, & 17 labels (Actin Filaments, Plasma Membrane, & Vesicles) >>>> In this competition, we merged some spatially similar (and functionally related) patterns into the same class. This was mentioned on the Single Cell Pattern notebook. The duplicates simply results from conversion code. For example, if an image has image-level label as Vesicles and Peroxisomes, then the code results in 17|17. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806) 有些图像有重复的label，这是因为转换时出了问题，忽略就好？\n\n> Image-level label uncertainty The image-level labels are what we refer to as weak or noisy. During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that the labels present in the majority of the images will be annotated. For negative samples it is not uncommon that lets say 4 images show no staining, while the remaining 2 show some unspecific staining or some granular pattern. If you compare the image-level label with the precise pattern observed in any given cell from this group of images, the label will be correct for the vast majority of cells, but perhaps not for all of them (as in your example 1 and 3 above). [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748)\n\n> Single cell label accuracy in test set The test set consists of images where each single cell has been annotated independently. Hence the accuracy of these labels is much better, and will be correct for each cell in every image. The statement below made by Trang Le, @lnhtrang, in our notebook explaining the patterns is correct for how the test set was annotated. *This class (Negative) includes negative stainings and unspecific patterns. This means that the cells have no green staining (negative) or have staining but no pattern can be deciphered from the staining (unspecific).* [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748)\n\n> More negative images can be found in the HPA public image data (external data), where they are annotated as \"No Staining\". [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748)\n\n> How to recognize unspecific patterns Unspecific patterns arise from the use of antibody reagents that do not function well, and do not bind to their intended target protein, but rather to all cellular structures (sometimes excluding the nucleolus). The second example from the top, and the cells of weaker intensity in the top image, are text-book examples of this. [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748)"},{"metadata":{},"cell_type":"markdown","source":"# 5. Solution Approaches <a class=\"anchor\" id=\"solutions\"></a>\n\nHere is a collection of links to various solution approaches shared in competition notebooks and discussions. \n\n* Puzzle-CAM by @phalanx - weakly supervised semantic segmentation + hpasegmentor solution\n    * https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217395\n    * https://github.com/OFRIN/PuzzleCAM\n* mmdetection for segmentation by @its7171 - convert masks to bboxes and train an object detection model\n    * https://www.kaggle.com/its7171/mmdetection-for-segmentation-training\n    * https://www.kaggle.com/its7171/mmdetection-for-segmentation-inference\n* HPA - classifier + explainability + segmentation by @samusram - build a classifier on top of image-level labels (multi-label classification), then use an explainability technique Guided-GRADCAM to extract regions responsible for particular class prediction, and then assign the segmented cells to particular classes based on the overlap with Grad-CAM outputs.\n    * https://www.kaggle.com/samusram/hpa-classifier-explainability-segmentation\n* Classification of segmented and cropped cells by @dschettler8845\n    * https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-training\n    * https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-inference\n    "},{"metadata":{},"cell_type":"markdown","source":"# 6. Resources <a class=\"anchor\" id=\"resources\"></a>\n\nThese resources may be helpful to learn more about the domain and/or approach these challenge. \n\n* Pretrained segmentation models\n    * [Cellpose](https://www.cellpose.org/)\n    * [HPACellSegmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)\n    * [EmbedSeg - Embedding-based Instance Segmentation](http://github.com/juglab/EmbedSeg) via @tanlikesmath\n* [Hosts' notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) demonstrating how you can make use of public HPA data and HPACellSegmentation model\n* [Hosts' guide explaining the patterns and cells in the microscope images](https://www.kaggle.com/lnhtrang/single-cell-patterns)\n* [Paper](https://www.nature.com/articles/s41592-019-0658-6.epdf?no_publisher_access=1&r3_referer=nature) describing results and approaches from the first HPA competition\n* An assumption that the labels were created via online players in an [online game](https://wiki.eveuniversity.org/Project_Discovery:_Human_Protein_Atlas) - via @ryches, [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215122). This may be valid for train set, competition hosts explained that the test set was labelled by experts at HAP [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141)\n* [Comparative Study on Automated Cell Nuclei Segmentation Methods for Cytology Pleural Effusion Images](https://www.hindawi.com/journals/jhe/2018/9240389/) - via @weka511, [link](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/216132)\n* [Organelles by HPA](https://www.proteinatlas.org/humanproteome/cell/organelle)\n* [Dictionary by HPA - Cell Structure part](https://www.proteinatlas.org/learn/dictionary/cell)\n* [A few tips on mAP by tito](https://www.kaggle.com/its7171/a-few-tips-on-map?scriptVersionId=53590041)\n\n\n## 6.1. Datasets\n\n* [Segmentation mask dataset](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215773) by @its7171\n* [Cell crop tile datasets](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/216984) by @dschettler8845\n* [16bit training set](https://www.kaggle.com/philculliton/hpa-2020-16bit-training-set) - the original train set is 8bit, while test set is 16bit. It might be useful for the final models to get trained on 16bit images as well. "},{"metadata":{},"cell_type":"markdown","source":"# To be continued...\n\nPlease upvote if you found this helpful, you will make my day :-)"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}