{
  "id": 72673,
  "title": "Questions/Concerns About the Quality of Competition Data and Pre-processing",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/72673",
  "author_name": "",
  "post_date": "2018-11-26T04:41:46.651639500Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>Hello! I'm new to this challenge and had some burning questions about the \"out-of-the-box\" quality of the competition data. I would greatly appreciate any answers to any of the questions I have. Additionally, if there are certain threads that have already answered the question or would be a good reference, I'd appreciate that as well. Thank you very much!</strong></p>\n\n<p><strong>The biggest question I would have to ask would be:</strong>\n“How much cleaning of the data needs to be done before any machine learning is done? How much pre-processing would be needed to turn the images into valuable inputs?”</p>\n\n<p><strong>And some other specifics about cleaning data would be:</strong>\nDoes each image contain a single cell only? Are the edges of the image the boundaries of a single cell?</p>\n\n<p>Does the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).</p>\n\n<p>Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clusteredwhen attempting to determine the subcellular location?</p>\n\n<p>In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training? </p>\n\n<p>How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?</p>\n\n<p>How has the data been collected, and how has the protein itself been isolated and identified?</p>\n\n<p>Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?</p>\n\n<p>Is each image at the same magnification? How do the cell sizes vary throughout the many images?</p>\n\n<p><strong>Thank you once again!</strong></p>",
  "messages": [
    {
      "id": "427753",
      "postDate": "11/26/2018 04:41:46",
      "content": "<p><strong>Hello! I'm new to this challenge and had some burning questions about the \"out-of-the-box\" quality of the competition data. I would greatly appreciate any answers to any of the questions I have. Additionally, if there are certain threads that have already answered the question or would be a good reference, I'd appreciate that as well. Thank you very much!</strong></p>\n\n<p><strong>The biggest question I would have to ask would be:</strong>\n“How much cleaning of the data needs to be done before any machine learning is done? How much pre-processing would be needed to turn the images into valuable inputs?”</p>\n\n<p><strong>And some other specifics about cleaning data would be:</strong>\nDoes each image contain a single cell only? Are the edges of the image the boundaries of a single cell?</p>\n\n<p>Does the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).</p>\n\n<p>Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clusteredwhen attempting to determine the subcellular location?</p>\n\n<p>In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training? </p>\n\n<p>How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?</p>\n\n<p>How has the data been collected, and how has the protein itself been isolated and identified?</p>\n\n<p>Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?</p>\n\n<p>Is each image at the same magnification? How do the cell sizes vary throughout the many images?</p>\n\n<p><strong>Thank you once again!</strong></p>",
      "rawMarkdown": "**Hello! I'm new to this challenge and had some burning questions about the \"out-of-the-box\" quality of the competition data. I would greatly appreciate any answers to any of the questions I have. Additionally, if there are certain threads that have already answered the question or would be a good reference, I'd appreciate that as well. Thank you very much!**\n\n**The biggest question I would have to ask would be:**\n“How much cleaning of the data needs to be done before any machine learning is done? How much pre-processing would be needed to turn the images into valuable inputs?”\n\n**And some other specifics about cleaning data would be:**\nDoes each image contain a single cell only? Are the edges of the image the boundaries of a single cell?\n\nDoes the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).\n\nIs each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clusteredwhen attempting to determine the subcellular location?\n\nIn the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training? \n\nHow are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?\n\nHow has the data been collected, and how has the protein itself been isolated and identified?\n\nCan we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?\n\nIs each image at the same magnification? How do the cell sizes vary throughout the many images?\n\n**Thank you once again!**",
      "votes": null
    },
    {
      "id": "427805",
      "postDate": "11/26/2018 07:26:09",
      "content": "<p>cleaning: not sure what you mean. The images are high quality photos and can be used directly for training.</p>\n\n<p>single cell? In general there are several cells per image. Several on the kernels show pictures so definitely check them out.</p>\n\n<p>3D: not sure - the images don't appear to show much interference from 3D though I am not sure what that would look like. </p>\n\n<p>specks: specks are not proteins - I believe that the specks are too big to be individual protein molecules. They are cell structures.</p>\n\n<p>by filters, do you mean red/blue/yellow? If so, they are just samples stained to show three reference structures in the cell. They are a big help to learning.I recommend to read Emma's writeup and the paper she references.</p>\n\n<p>magnification: I believe the images have the same magnification - they all come from the same (type of) instrument.</p>",
      "rawMarkdown": "cleaning: not sure what you mean. The images are high quality photos and can be used directly for training.\n\nsingle cell? In general there are several cells per image. Several on the kernels show pictures so definitely check them out.\n\n3D: not sure - the images don't appear to show much interference from 3D though I am not sure what that would look like. \n\nspecks: specks are not proteins - I believe that the specks are too big to be individual protein molecules. They are cell structures.\n\nby filters, do you mean red/blue/yellow? If so, they are just samples stained to show three reference structures in the cell. They are a big help to learning.I recommend to read Emma's writeup and the paper she references.\n\nmagnification: I believe the images have the same magnification - they all come from the same (type of) instrument.",
      "votes": null
    },
    {
      "id": "427940",
      "postDate": "11/26/2018 12:30:00",
      "content": "<p>pete is quite correct but I will elaborate a bit.</p>\n\n<p>Q: Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clustered when attempting to determine the subcellular location?</p>\n\n<p>A: Not sure what you mean by speck, some images have large green spots that are very discrete, others have much smaller and diffuse green speckles, one should assume that every green pixel is in fact a protein (though some images have obvious artifacts, they are rare).  Focusing only on large clusters probably means focusing on some labels while ignoring others.</p>\n\n<p>Q: In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training?</p>\n\n<p>A: If by filters you mean the color channels, blue is fluorescence from DAPI which binds strongly to A-T rich regions of dsDNA. When bound to double stranded DNA and excited by ultraviolet light DAPI has emission peak at 460nm, so its very blue. Something to keep in mind is that while DAPI has a much lower affinity for RNA it will bind to it and RNA is everywhere in the cell, mostly at bound ribosomes on the rough ER but also some will be cytosolic and mitochondria contain both RNA and their own DNA. DAPI bound to RNA has an emission peak close to 500nm, which is green, so there can be some overlap between DAPI and the common green fluorescent tags like GFP or flourescein. Remember GFP is most excited by blue light around 470nm which is quite a bit lower than the wavelength that used to excite DAPI, but there will always be non-zero spillover. I doubt trying to take that into account will make much difference here, but you never know. The 'nucleus' label should mean some protein (green channel) was found in or on the nucleus.</p>\n\n<p>Q: How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?</p>\n\n<p>A: Whats interesting (and a little maddening) about this challenge is that we are not really localizing proteins per se, rather we are indeed attempting to identify the sub-cellular regions themselves -- class25: cytosol, for example, is the entire interior of the cell, it is the fluid filling the cell and exists everywhere, though the fluid-filled interior of membrane-bound structures, like the interior of the nuclear membrane, lumen of the ER, mitochondrial cristae, etc. don't count as cytoplasm. To your specific question, the difference between microtubules and microtubule ends is not arbitrary at all as they will be associated with quite different proteins. For example, microtubules are primarily composed of two proteins, alpha- and beta-tubulin and are major components of the cytoskeleton. Cells are motile or at least can change their shapes more or less, which requires microtubles to be assembled and disassembled as needed, this occurs at the ends and is facilitated by many other proteins and enzymes including numerous GTPases and ATPases like katanin. Microtubules also serve as highways for motor proteins that move cargo around the cell like dyneins and kinesins, of which there are dozens of known subtypes in humans. </p>\n\n<p>Q: How has the data been collected, and how has the protein itself been isolated and identified?</p>\n\n<p>A: From Sullivan et al. (2018) the proteins of interest were identified using a standard protocol, briefly, cells are incubated with a primary antibody against whatever protein and a secondary antibody attached to whatever fluorescent tag is used to visualize the primary antibody. The actual fluorescent tag is in reality two antibody lengths away from the protein itself, so ~10-20nm, but that is beyond the resolution limit for light microscopy and variability from that will not affect anything here. What could have an effect is the efficacy and specificity of the antibodies, the paper refers to the precise antibodies used and their specs can be found from the vendor, Sigma Aldrich.\nThe sizes of the cells will vary by sample and cell-type, this is quite variable. Though only 17 cell lines were used in the paper, it describes an older version of the atlas (v14, while these images most likely come from v18), there may be additional cell types, though if so I doubt its many. If there is any information available linking the image Ids to cell types, I have not found it.\nAs to the acquisition I will quote the paper:</p>\n\n<p>\" Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 µm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). \"</p>\n\n<p>So yes, all images were at the same magnification. See the supplemental methods for complete details. </p>\n\n<p>Q: Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?</p>\n\n<p>A: From the paper:</p>\n\n<p>\"RESULTS\nSubcellular distribution of proteins in microscopy images\nEach sample in the HPA Cell Atlas consists of human cells that are\nimmunofluorescently labeled for one protein of interest and three\nreference markers: DAPI for the nucleus and antibody-based labeling\nof microtubules and the endoplasmic reticulum. High-resolution\nimages were acquired using confocal microscopy (Fig. 1a). The resulting\nimages were annotated to determine the localization(s) of the\nprotein of interest with the help of the three cellular reference markers.\"</p>\n\n<p>So single proteins only per image was the intention and design, the results depend to an extent on antibody specificity and experimenter skill. Information on the former will be in the Sigma catalog, and I see no reason to doubt the later. </p>\n\n<p>I hope this is helpful, good luck!</p>",
      "rawMarkdown": "pete is quite correct but I will elaborate a bit.\n\nQ: Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clustered when attempting to determine the subcellular location?\n\nA: Not sure what you mean by speck, some images have large green spots that are very discrete, others have much smaller and diffuse green speckles, one should assume that every green pixel is in fact a protein (though some images have obvious artifacts, they are rare).  Focusing only on large clusters probably means focusing on some labels while ignoring others.\n\nQ: In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training?\n\nA: If by filters you mean the color channels, blue is fluorescence from DAPI which binds strongly to A-T rich regions of dsDNA. When bound to double stranded DNA and excited by ultraviolet light DAPI has emission peak at 460nm, so its very blue. Something to keep in mind is that while DAPI has a much lower affinity for RNA it will bind to it and RNA is everywhere in the cell, mostly at bound ribosomes on the rough ER but also some will be cytosolic and mitochondria contain both RNA and their own DNA. DAPI bound to RNA has an emission peak close to 500nm, which is green, so there can be some overlap between DAPI and the common green fluorescent tags like GFP or flourescein. Remember GFP is most excited by blue light around 470nm which is quite a bit lower than the wavelength that used to excite DAPI, but there will always be non-zero spillover. I doubt trying to take that into account will make much difference here, but you never know. The 'nucleus' label should mean some protein (green channel) was found in or on the nucleus.\n\nQ: How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?\n\nA: Whats interesting (and a little maddening) about this challenge is that we are not really localizing proteins per se, rather we are indeed attempting to identify the sub-cellular regions themselves -- class25: cytosol, for example, is the entire interior of the cell, it is the fluid filling the cell and exists everywhere, though the fluid-filled interior of membrane-bound structures, like the interior of the nuclear membrane, lumen of the ER, mitochondrial cristae, etc. don't count as cytoplasm. To your specific question, the difference between microtubules and microtubule ends is not arbitrary at all as they will be associated with quite different proteins. For example, microtubules are primarily composed of two proteins, alpha- and beta-tubulin and are major components of the cytoskeleton. Cells are motile or at least can change their shapes more or less, which requires microtubles to be assembled and disassembled as needed, this occurs at the ends and is facilitated by many other proteins and enzymes including numerous GTPases and ATPases like katanin. Microtubules also serve as highways for motor proteins that move cargo around the cell like dyneins and kinesins, of which there are dozens of known subtypes in humans. \n\nQ: How has the data been collected, and how has the protein itself been isolated and identified?\n\nA: From Sullivan et al. (2018) the proteins of interest were identified using a standard protocol, briefly, cells are incubated with a primary antibody against whatever protein and a secondary antibody attached to whatever fluorescent tag is used to visualize the primary antibody. The actual fluorescent tag is in reality two antibody lengths away from the protein itself, so ~10-20nm, but that is beyond the resolution limit for light microscopy and variability from that will not affect anything here. What could have an effect is the efficacy and specificity of the antibodies, the paper refers to the precise antibodies used and their specs can be found from the vendor, Sigma Aldrich.\nThe sizes of the cells will vary by sample and cell-type, this is quite variable. Though only 17 cell lines were used in the paper, it describes an older version of the atlas (v14, while these images most likely come from v18), there may be additional cell types, though if so I doubt its many. If there is any information available linking the image Ids to cell types, I have not found it.\nAs to the acquisition I will quote the paper:\n\n\" Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 µm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). \"\n\nSo yes, all images were at the same magnification. See the supplemental methods for complete details. \n\nQ: Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?\n\nA: From the paper:\n\n\"RESULTS\nSubcellular distribution of proteins in microscopy images\nEach sample in the HPA Cell Atlas consists of human cells that are\nimmunofluorescently labeled for one protein of interest and three\nreference markers: DAPI for the nucleus and antibody-based labeling\nof microtubules and the endoplasmic reticulum. High-resolution\nimages were acquired using confocal microscopy (Fig. 1a). The resulting\nimages were annotated to determine the localization(s) of the\nprotein of interest with the help of the three cellular reference markers.\"\n\nSo single proteins only per image was the intention and design, the results depend to an extent on antibody specificity and experimenter skill. Information on the former will be in the Sigma catalog, and I see no reason to doubt the later. \n\nI hope this is helpful, good luck!",
      "votes": null
    },
    {
      "id": "428206",
      "postDate": "11/26/2018 23:14:57",
      "content": "<p>Certainly useful for me (layman) - thanks.</p>\n\n<p>One of the  questions was:</p>\n\n<p>Does the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).</p>\n\n<p>Thinking more on this, I would venture that the samples are very thin in depth and so they do not contain cells that overlap in depth. So there would be no 3D problem.</p>",
      "rawMarkdown": "Certainly useful for me (layman) - thanks.\n\nOne of the  questions was:\n\nDoes the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).\n\nThinking more on this, I would venture that the samples are very thin in depth and so they do not contain cells that overlap in depth. So there would be no 3D problem.",
      "votes": null
    },
    {
      "id": "428293",
      "postDate": "11/27/2018 03:19:39",
      "content": "<p>Glad to be of service! I'm new to ML but not at all to cell and molec bio research and I'm happy to answer any questions I'm able to. </p>\n\n<p>The question about depth is a good one. Yes, the cells are very thin - they are mounted on a glass microscope slide and covered with a silicate cover-slip that both holds them still for photography and flattens them a bit to make imaging easier. One typically transfers the cells to the slide using a pipette, after incubation with the secondary antibody the cells are mechanically and/or chemically detached from the culture plates, suspended in PBS and glycerol and a drop of this suspension is placed on the slide.</p>\n\n<p>While the cells are indeed thin they are not so thin that very useful depth information cannot be acquired, this is one of the main benefits of confocal microscopy. The depth information that is useful is not about cells being on top of each other, which from the images seems to not be much of a problem, but that protein distribution within cells is indeed 3D. However, the images provided here are projections of a z-stack so we do not actually have access to the depth information and I don't believe there is enough here to attempt to reconstruct it -- without knowing which proteins we are looking at makes inferring depth from what we have probably impossible or at least prohibitively impractical. The HPA dudes do have the z-stacks, they are a little vague in the paper on that detail, only mentioning that z-stacks were acquired from 6 FOVs - that does not mean each image is a projection of 6 z-positions though, just that of the 6 FOVs there must have been at least two depths or there would be no z-stack. \nExpanding from the supplemental methods a bit:</p>\n\n<p>\"Proteins are cataloged serially using in-house generated antibodies and\nimmunostaining in a gene-centric manner as described in detail previously7.\nBriefly, the spatial distribution of each protein is studied in three cell lines\nout of a panel of 17; U-2 OS and two additional selected to have the highest\nRNA expression level of the corresponding gene. Each antibody-cell line\n‘sample’ is then imaged to produce a minimum of two images per sample\n(average 2.93 images per sample). Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 μm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). Each channel is stored as a\nseparate 2,048 × 2,048 16-bit ome-tiff.\"</p>\n\n<p>There is a lot of information in the possession of the HPA people that would make this challenge much more useful, imo. Perhaps they will provide some of it for the next stage. Neither they, nor anyone else, would try to do what this challenge is asking with only the information given as part of basic research - but I suppose that keeps things interesting. </p>",
      "rawMarkdown": "Glad to be of service! I'm new to ML but not at all to cell and molec bio research and I'm happy to answer any questions I'm able to. \n\nThe question about depth is a good one. Yes, the cells are very thin - they are mounted on a glass microscope slide and covered with a silicate cover-slip that both holds them still for photography and flattens them a bit to make imaging easier. One typically transfers the cells to the slide using a pipette, after incubation with the secondary antibody the cells are mechanically and/or chemically detached from the culture plates, suspended in PBS and glycerol and a drop of this suspension is placed on the slide.\n\nWhile the cells are indeed thin they are not so thin that very useful depth information cannot be acquired, this is one of the main benefits of confocal microscopy. The depth information that is useful is not about cells being on top of each other, which from the images seems to not be much of a problem, but that protein distribution within cells is indeed 3D. However, the images provided here are projections of a z-stack so we do not actually have access to the depth information and I don't believe there is enough here to attempt to reconstruct it -- without knowing which proteins we are looking at makes inferring depth from what we have probably impossible or at least prohibitively impractical. The HPA dudes do have the z-stacks, they are a little vague in the paper on that detail, only mentioning that z-stacks were acquired from 6 FOVs - that does not mean each image is a projection of 6 z-positions though, just that of the 6 FOVs there must have been at least two depths or there would be no z-stack. \nExpanding from the supplemental methods a bit:\n\n\"Proteins are cataloged serially using in-house generated antibodies and\nimmunostaining in a gene-centric manner as described in detail previously7.\nBriefly, the spatial distribution of each protein is studied in three cell lines\nout of a panel of 17; U-2 OS and two additional selected to have the highest\nRNA expression level of the corresponding gene. Each antibody-cell line\n‘sample’ is then imaged to produce a minimum of two images per sample\n(average 2.93 images per sample). Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 μm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). Each channel is stored as a\nseparate 2,048 × 2,048 16-bit ome-tiff.\"\n\nThere is a lot of information in the possession of the HPA people that would make this challenge much more useful, imo. Perhaps they will provide some of it for the next stage. Neither they, nor anyone else, would try to do what this challenge is asking with only the information given as part of basic research - but I suppose that keeps things interesting.",
      "votes": null
    },
    {
      "id": "428658",
      "postDate": "11/27/2018 17:07:51",
      "content": "<p>OK while we're at it, here is another question. I have spent a little time looking at the big tif images and they appear to me to be spectacular in their improved resolution. The question is: for humans trying to work with these images for protein classification, how useful is the extra resolution? </p>",
      "rawMarkdown": "OK while we're at it, here is another question. I have spent a little time looking at the big tif images and they appear to me to be spectacular in their improved resolution. The question is: for humans trying to work with these images for protein classification, how useful is the extra resolution?",
      "votes": null
    },
    {
      "id": "428799",
      "postDate": "11/27/2018 22:49:19",
      "content": "<p>Resolution is everything, how can we know something is there if we can't see it? Of course there are many ways to infer the existence of something without direct observation, but biology is an empirical science so no amount of data is too much, ideally. 2080x2080 TIFs are very modest really, compared to what microscopes like the one used to acquire these data are capable of producing. </p>\n\n<p>We mentioned depth information before, the labs that have and use these microscopes are acquiring raw image data at far higher resolutions than 2080. When you bring in the z-stacks you will get up to maybe a dozen or more images of the same field at different focal planes (depending on the equipment and skill of the operator), add the color filters and you end up with some high-dimensional volumetric data, which is great, but it becomes a lot to make sense of. They just gave us those 4 false-colored channels, but there can be quite a bit more than 4 for the same field--we know the absorbance and emission spectra of the fluorescent tags we use and there are filters capable of sufficiently separating wavelengths of ~10nm difference. </p>\n\n<p>All this is to say that cells are very small, and the 10's of thousands of unique proteins are much smaller and while the cell is alive, they are constantly in motion, either being trafficked around the cell, trafficking other things about, or doing their job as scaffolding or as enzymes, often simultaneously. We need all the resolution because we have no idea what could be there until we see it, or see something its causing. </p>\n\n<p>I do routinely use fluorescence microscopy of various types in my research, but my focus now is on electron microscopy as there are hard limits on the resolution of light microscopy and what I am studying cannot be quantitatively measured any other way. Large confocal stacks can be quite large, but I'm used to working with hundreds of ~500MB serial single images around 30k x 30k. Different tools for different problems, but I really enjoy all the pretty pictures of cells they gave us, now if only I could make some sense of them.\ncheers</p>",
      "rawMarkdown": "Resolution is everything, how can we know something is there if we can't see it? Of course there are many ways to infer the existence of something without direct observation, but biology is an empirical science so no amount of data is too much, ideally. 2080x2080 TIFs are very modest really, compared to what microscopes like the one used to acquire these data are capable of producing. \n\nWe mentioned depth information before, the labs that have and use these microscopes are acquiring raw image data at far higher resolutions than 2080. When you bring in the z-stacks you will get up to maybe a dozen or more images of the same field at different focal planes (depending on the equipment and skill of the operator), add the color filters and you end up with some high-dimensional volumetric data, which is great, but it becomes a lot to make sense of. They just gave us those 4 false-colored channels, but there can be quite a bit more than 4 for the same field--we know the absorbance and emission spectra of the fluorescent tags we use and there are filters capable of sufficiently separating wavelengths of ~10nm difference. \n\nAll this is to say that cells are very small, and the 10's of thousands of unique proteins are much smaller and while the cell is alive, they are constantly in motion, either being trafficked around the cell, trafficking other things about, or doing their job as scaffolding or as enzymes, often simultaneously. We need all the resolution because we have no idea what could be there until we see it, or see something its causing. \n\nI do routinely use fluorescence microscopy of various types in my research, but my focus now is on electron microscopy as there are hard limits on the resolution of light microscopy and what I am studying cannot be quantitatively measured any other way. Large confocal stacks can be quite large, but I'm used to working with hundreds of ~500MB serial single images around 30k x 30k. Different tools for different problems, but I really enjoy all the pretty pictures of cells they gave us, now if only I could make some sense of them.\ncheers",
      "votes": null
    },
    {
      "id": "428876",
      "postDate": "11/28/2018 02:38:57",
      "content": "<p>Thank you! Your answers cleared up quite a bit. I understand that the filters will help greatly in identifying the proteins, however, there are some 28 subcellular locations and 3 filters - for the 25 remaining locations, how are they supposed to be identified and classified? Additionally, does each protein image specify what protein is to be identified and the cell type? I am having trouble identifying what really needs to be done here -  are we looking to map every single instance of a certain protein in every location of every type of cell, and do that for many proteins? In any case, without knowing the specific protein to identify, it seems like that would be impossible.</p>",
      "rawMarkdown": "Thank you! Your answers cleared up quite a bit. I understand that the filters will help greatly in identifying the proteins, however, there are some 28 subcellular locations and 3 filters - for the 25 remaining locations, how are they supposed to be identified and classified? Additionally, does each protein image specify what protein is to be identified and the cell type? I am having trouble identifying what really needs to be done here -  are we looking to map every single instance of a certain protein in every location of every type of cell, and do that for many proteins? In any case, without knowing the specific protein to identify, it seems like that would be impossible.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 427805,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "11/26/2018 07:26:09",
      "content": "<p>cleaning: not sure what you mean. The images are high quality photos and can be used directly for training.</p>\n\n<p>single cell? In general there are several cells per image. Several on the kernels show pictures so definitely check them out.</p>\n\n<p>3D: not sure - the images don't appear to show much interference from 3D though I am not sure what that would look like. </p>\n\n<p>specks: specks are not proteins - I believe that the specks are too big to be individual protein molecules. They are cell structures.</p>\n\n<p>by filters, do you mean red/blue/yellow? If so, they are just samples stained to show three reference structures in the cell. They are a big help to learning.I recommend to read Emma's writeup and the paper she references.</p>\n\n<p>magnification: I believe the images have the same magnification - they all come from the same (type of) instrument.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427940,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "11/26/2018 12:30:00",
      "content": "<p>pete is quite correct but I will elaborate a bit.</p>\n\n<p>Q: Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clustered when attempting to determine the subcellular location?</p>\n\n<p>A: Not sure what you mean by speck, some images have large green spots that are very discrete, others have much smaller and diffuse green speckles, one should assume that every green pixel is in fact a protein (though some images have obvious artifacts, they are rare).  Focusing only on large clusters probably means focusing on some labels while ignoring others.</p>\n\n<p>Q: In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training?</p>\n\n<p>A: If by filters you mean the color channels, blue is fluorescence from DAPI which binds strongly to A-T rich regions of dsDNA. When bound to double stranded DNA and excited by ultraviolet light DAPI has emission peak at 460nm, so its very blue. Something to keep in mind is that while DAPI has a much lower affinity for RNA it will bind to it and RNA is everywhere in the cell, mostly at bound ribosomes on the rough ER but also some will be cytosolic and mitochondria contain both RNA and their own DNA. DAPI bound to RNA has an emission peak close to 500nm, which is green, so there can be some overlap between DAPI and the common green fluorescent tags like GFP or flourescein. Remember GFP is most excited by blue light around 470nm which is quite a bit lower than the wavelength that used to excite DAPI, but there will always be non-zero spillover. I doubt trying to take that into account will make much difference here, but you never know. The 'nucleus' label should mean some protein (green channel) was found in or on the nucleus.</p>\n\n<p>Q: How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?</p>\n\n<p>A: Whats interesting (and a little maddening) about this challenge is that we are not really localizing proteins per se, rather we are indeed attempting to identify the sub-cellular regions themselves -- class25: cytosol, for example, is the entire interior of the cell, it is the fluid filling the cell and exists everywhere, though the fluid-filled interior of membrane-bound structures, like the interior of the nuclear membrane, lumen of the ER, mitochondrial cristae, etc. don't count as cytoplasm. To your specific question, the difference between microtubules and microtubule ends is not arbitrary at all as they will be associated with quite different proteins. For example, microtubules are primarily composed of two proteins, alpha- and beta-tubulin and are major components of the cytoskeleton. Cells are motile or at least can change their shapes more or less, which requires microtubles to be assembled and disassembled as needed, this occurs at the ends and is facilitated by many other proteins and enzymes including numerous GTPases and ATPases like katanin. Microtubules also serve as highways for motor proteins that move cargo around the cell like dyneins and kinesins, of which there are dozens of known subtypes in humans. </p>\n\n<p>Q: How has the data been collected, and how has the protein itself been isolated and identified?</p>\n\n<p>A: From Sullivan et al. (2018) the proteins of interest were identified using a standard protocol, briefly, cells are incubated with a primary antibody against whatever protein and a secondary antibody attached to whatever fluorescent tag is used to visualize the primary antibody. The actual fluorescent tag is in reality two antibody lengths away from the protein itself, so ~10-20nm, but that is beyond the resolution limit for light microscopy and variability from that will not affect anything here. What could have an effect is the efficacy and specificity of the antibodies, the paper refers to the precise antibodies used and their specs can be found from the vendor, Sigma Aldrich.\nThe sizes of the cells will vary by sample and cell-type, this is quite variable. Though only 17 cell lines were used in the paper, it describes an older version of the atlas (v14, while these images most likely come from v18), there may be additional cell types, though if so I doubt its many. If there is any information available linking the image Ids to cell types, I have not found it.\nAs to the acquisition I will quote the paper:</p>\n\n<p>\" Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 µm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). \"</p>\n\n<p>So yes, all images were at the same magnification. See the supplemental methods for complete details. </p>\n\n<p>Q: Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?</p>\n\n<p>A: From the paper:</p>\n\n<p>\"RESULTS\nSubcellular distribution of proteins in microscopy images\nEach sample in the HPA Cell Atlas consists of human cells that are\nimmunofluorescently labeled for one protein of interest and three\nreference markers: DAPI for the nucleus and antibody-based labeling\nof microtubules and the endoplasmic reticulum. High-resolution\nimages were acquired using confocal microscopy (Fig. 1a). The resulting\nimages were annotated to determine the localization(s) of the\nprotein of interest with the help of the three cellular reference markers.\"</p>\n\n<p>So single proteins only per image was the intention and design, the results depend to an extent on antibody specificity and experimenter skill. Information on the former will be in the Sigma catalog, and I see no reason to doubt the later. </p>\n\n<p>I hope this is helpful, good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 428206,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "11/26/2018 23:14:57",
          "content": "<p>Certainly useful for me (layman) - thanks.</p>\n\n<p>One of the  questions was:</p>\n\n<p>Does the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).</p>\n\n<p>Thinking more on this, I would venture that the samples are very thin in depth and so they do not contain cells that overlap in depth. So there would be no 3D problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428293,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "11/27/2018 03:19:39",
          "content": "<p>Glad to be of service! I'm new to ML but not at all to cell and molec bio research and I'm happy to answer any questions I'm able to. </p>\n\n<p>The question about depth is a good one. Yes, the cells are very thin - they are mounted on a glass microscope slide and covered with a silicate cover-slip that both holds them still for photography and flattens them a bit to make imaging easier. One typically transfers the cells to the slide using a pipette, after incubation with the secondary antibody the cells are mechanically and/or chemically detached from the culture plates, suspended in PBS and glycerol and a drop of this suspension is placed on the slide.</p>\n\n<p>While the cells are indeed thin they are not so thin that very useful depth information cannot be acquired, this is one of the main benefits of confocal microscopy. The depth information that is useful is not about cells being on top of each other, which from the images seems to not be much of a problem, but that protein distribution within cells is indeed 3D. However, the images provided here are projections of a z-stack so we do not actually have access to the depth information and I don't believe there is enough here to attempt to reconstruct it -- without knowing which proteins we are looking at makes inferring depth from what we have probably impossible or at least prohibitively impractical. The HPA dudes do have the z-stacks, they are a little vague in the paper on that detail, only mentioning that z-stacks were acquired from 6 FOVs - that does not mean each image is a projection of 6 z-positions though, just that of the 6 FOVs there must have been at least two depths or there would be no z-stack. \nExpanding from the supplemental methods a bit:</p>\n\n<p>\"Proteins are cataloged serially using in-house generated antibodies and\nimmunostaining in a gene-centric manner as described in detail previously7.\nBriefly, the spatial distribution of each protein is studied in three cell lines\nout of a panel of 17; U-2 OS and two additional selected to have the highest\nRNA expression level of the corresponding gene. Each antibody-cell line\n‘sample’ is then imaged to produce a minimum of two images per sample\n(average 2.93 images per sample). Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 μm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). Each channel is stored as a\nseparate 2,048 × 2,048 16-bit ome-tiff.\"</p>\n\n<p>There is a lot of information in the possession of the HPA people that would make this challenge much more useful, imo. Perhaps they will provide some of it for the next stage. Neither they, nor anyone else, would try to do what this challenge is asking with only the information given as part of basic research - but I suppose that keeps things interesting. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428658,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "11/27/2018 17:07:51",
          "content": "<p>OK while we're at it, here is another question. I have spent a little time looking at the big tif images and they appear to me to be spectacular in their improved resolution. The question is: for humans trying to work with these images for protein classification, how useful is the extra resolution? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428799,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "11/27/2018 22:49:19",
          "content": "<p>Resolution is everything, how can we know something is there if we can't see it? Of course there are many ways to infer the existence of something without direct observation, but biology is an empirical science so no amount of data is too much, ideally. 2080x2080 TIFs are very modest really, compared to what microscopes like the one used to acquire these data are capable of producing. </p>\n\n<p>We mentioned depth information before, the labs that have and use these microscopes are acquiring raw image data at far higher resolutions than 2080. When you bring in the z-stacks you will get up to maybe a dozen or more images of the same field at different focal planes (depending on the equipment and skill of the operator), add the color filters and you end up with some high-dimensional volumetric data, which is great, but it becomes a lot to make sense of. They just gave us those 4 false-colored channels, but there can be quite a bit more than 4 for the same field--we know the absorbance and emission spectra of the fluorescent tags we use and there are filters capable of sufficiently separating wavelengths of ~10nm difference. </p>\n\n<p>All this is to say that cells are very small, and the 10's of thousands of unique proteins are much smaller and while the cell is alive, they are constantly in motion, either being trafficked around the cell, trafficking other things about, or doing their job as scaffolding or as enzymes, often simultaneously. We need all the resolution because we have no idea what could be there until we see it, or see something its causing. </p>\n\n<p>I do routinely use fluorescence microscopy of various types in my research, but my focus now is on electron microscopy as there are hard limits on the resolution of light microscopy and what I am studying cannot be quantitatively measured any other way. Large confocal stacks can be quite large, but I'm used to working with hundreds of ~500MB serial single images around 30k x 30k. Different tools for different problems, but I really enjoy all the pretty pictures of cells they gave us, now if only I could make some sense of them.\ncheers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 428876,
          "author_name": "andstun",
          "author_url": "",
          "post_date": "11/28/2018 02:38:57",
          "content": "<p>Thank you! Your answers cleared up quite a bit. I understand that the filters will help greatly in identifying the proteins, however, there are some 28 subcellular locations and 3 filters - for the 25 remaining locations, how are they supposed to be identified and classified? Additionally, does each protein image specify what protein is to be identified and the cell type? I am having trouble identifying what really needs to be done here -  are we looking to map every single instance of a certain protein in every location of every type of cell, and do that for many proteins? In any case, without knowing the specific protein to identify, it seems like that would be impossible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "427753": "**Hello! I'm new to this challenge and had some burning questions about the \"out-of-the-box\" quality of the competition data. I would greatly appreciate any answers to any of the questions I have. Additionally, if there are certain threads that have already answered the question or would be a good reference, I'd appreciate that as well. Thank you very much!**\n\n**The biggest question I would have to ask would be:**\n“How much cleaning of the data needs to be done before any machine learning is done? How much pre-processing would be needed to turn the images into valuable inputs?”\n\n**And some other specifics about cleaning data would be:**\nDoes each image contain a single cell only? Are the edges of the image the boundaries of a single cell?\n\nDoes the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).\n\nIs each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clusteredwhen attempting to determine the subcellular location?\n\nIn the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training? \n\nHow are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?\n\nHow has the data been collected, and how has the protein itself been isolated and identified?\n\nCan we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?\n\nIs each image at the same magnification? How do the cell sizes vary throughout the many images?\n\n**Thank you once again!**",
    "427805": "cleaning: not sure what you mean. The images are high quality photos and can be used directly for training.\n\nsingle cell? In general there are several cells per image. Several on the kernels show pictures so definitely check them out.\n\n3D: not sure - the images don't appear to show much interference from 3D though I am not sure what that would look like. \n\nspecks: specks are not proteins - I believe that the specks are too big to be individual protein molecules. They are cell structures.\n\nby filters, do you mean red/blue/yellow? If so, they are just samples stained to show three reference structures in the cell. They are a big help to learning.I recommend to read Emma's writeup and the paper she references.\n\nmagnification: I believe the images have the same magnification - they all come from the same (type of) instrument.",
    "427940": "pete is quite correct but I will elaborate a bit.\n\nQ: Is each speck under the green filter a protein? If so, how significant are the spots that are not highly clustered? Should we focus on points where more proteins are clustered when attempting to determine the subcellular location?\n\nA: Not sure what you mean by speck, some images have large green spots that are very discrete, others have much smaller and diffuse green speckles, one should assume that every green pixel is in fact a protein (though some images have obvious artifacts, they are rare).  Focusing only on large clusters probably means focusing on some labels while ignoring others.\n\nQ: In the case of the filters, would the “nucleus” region depict the actual cell nucleus or does it just highlight more proteins found in that nucleus area? Are the filters meant to be used as part of training?\n\nA: If by filters you mean the color channels, blue is fluorescence from DAPI which binds strongly to A-T rich regions of dsDNA. When bound to double stranded DNA and excited by ultraviolet light DAPI has emission peak at 460nm, so its very blue. Something to keep in mind is that while DAPI has a much lower affinity for RNA it will bind to it and RNA is everywhere in the cell, mostly at bound ribosomes on the rough ER but also some will be cytosolic and mitochondria contain both RNA and their own DNA. DAPI bound to RNA has an emission peak close to 500nm, which is green, so there can be some overlap between DAPI and the common green fluorescent tags like GFP or flourescein. Remember GFP is most excited by blue light around 470nm which is quite a bit lower than the wavelength that used to excite DAPI, but there will always be non-zero spillover. I doubt trying to take that into account will make much difference here, but you never know. The 'nucleus' label should mean some protein (green channel) was found in or on the nucleus.\n\nQ: How are subcellular locations other than those highlighted by the filters supposed be determined? The difference between \"microtubules\" and \"microtubule ends\" seems quite arbitrary; why are the even necessary?\n\nA: Whats interesting (and a little maddening) about this challenge is that we are not really localizing proteins per se, rather we are indeed attempting to identify the sub-cellular regions themselves -- class25: cytosol, for example, is the entire interior of the cell, it is the fluid filling the cell and exists everywhere, though the fluid-filled interior of membrane-bound structures, like the interior of the nuclear membrane, lumen of the ER, mitochondrial cristae, etc. don't count as cytoplasm. To your specific question, the difference between microtubules and microtubule ends is not arbitrary at all as they will be associated with quite different proteins. For example, microtubules are primarily composed of two proteins, alpha- and beta-tubulin and are major components of the cytoskeleton. Cells are motile or at least can change their shapes more or less, which requires microtubles to be assembled and disassembled as needed, this occurs at the ends and is facilitated by many other proteins and enzymes including numerous GTPases and ATPases like katanin. Microtubules also serve as highways for motor proteins that move cargo around the cell like dyneins and kinesins, of which there are dozens of known subtypes in humans. \n\nQ: How has the data been collected, and how has the protein itself been isolated and identified?\n\nA: From Sullivan et al. (2018) the proteins of interest were identified using a standard protocol, briefly, cells are incubated with a primary antibody against whatever protein and a secondary antibody attached to whatever fluorescent tag is used to visualize the primary antibody. The actual fluorescent tag is in reality two antibody lengths away from the protein itself, so ~10-20nm, but that is beyond the resolution limit for light microscopy and variability from that will not affect anything here. What could have an effect is the efficacy and specificity of the antibodies, the paper refers to the precise antibodies used and their specs can be found from the vendor, Sigma Aldrich.\nThe sizes of the cells will vary by sample and cell-type, this is quite variable. Though only 17 cell lines were used in the paper, it describes an older version of the atlas (v14, while these images most likely come from v18), there may be additional cell types, though if so I doubt its many. If there is any information available linking the image Ids to cell types, I have not found it.\nAs to the acquisition I will quote the paper:\n\n\" Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 µm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). \"\n\nSo yes, all images were at the same magnification. See the supplemental methods for complete details. \n\nQ: Can we be sure that there is only one protein identified? Is it possible for a single cell image to depict multiple proteins?\n\nA: From the paper:\n\n\"RESULTS\nSubcellular distribution of proteins in microscopy images\nEach sample in the HPA Cell Atlas consists of human cells that are\nimmunofluorescently labeled for one protein of interest and three\nreference markers: DAPI for the nucleus and antibody-based labeling\nof microtubules and the endoplasmic reticulum. High-resolution\nimages were acquired using confocal microscopy (Fig. 1a). The resulting\nimages were annotated to determine the localization(s) of the\nprotein of interest with the help of the three cellular reference markers.\"\n\nSo single proteins only per image was the intention and design, the results depend to an extent on antibody specificity and experimenter skill. Information on the former will be in the Sigma catalog, and I see no reason to doubt the later. \n\nI hope this is helpful, good luck!",
    "428206": "Certainly useful for me (layman) - thanks.\n\nOne of the  questions was:\n\nDoes the protein image take cell depth into account? In the case of one protein locale being in front of another, how would the location be identified? (This is more of a 3D geometric problem).\n\nThinking more on this, I would venture that the samples are very thin in depth and so they do not contain cells that overlap in depth. So there would be no 3D problem.",
    "428293": "Glad to be of service! I'm new to ML but not at all to cell and molec bio research and I'm happy to answer any questions I'm able to. \n\nThe question about depth is a good one. Yes, the cells are very thin - they are mounted on a glass microscope slide and covered with a silicate cover-slip that both holds them still for photography and flattens them a bit to make imaging easier. One typically transfers the cells to the slide using a pipette, after incubation with the secondary antibody the cells are mechanically and/or chemically detached from the culture plates, suspended in PBS and glycerol and a drop of this suspension is placed on the slide.\n\nWhile the cells are indeed thin they are not so thin that very useful depth information cannot be acquired, this is one of the main benefits of confocal microscopy. The depth information that is useful is not about cells being on top of each other, which from the images seems to not be much of a problem, but that protein distribution within cells is indeed 3D. However, the images provided here are projections of a z-stack so we do not actually have access to the depth information and I don't believe there is enough here to attempt to reconstruct it -- without knowing which proteins we are looking at makes inferring depth from what we have probably impossible or at least prohibitively impractical. The HPA dudes do have the z-stacks, they are a little vague in the paper on that detail, only mentioning that z-stacks were acquired from 6 FOVs - that does not mean each image is a projection of 6 z-positions though, just that of the 6 FOVs there must have been at least two depths or there would be no z-stack. \nExpanding from the supplemental methods a bit:\n\n\"Proteins are cataloged serially using in-house generated antibodies and\nimmunostaining in a gene-centric manner as described in detail previously7.\nBriefly, the spatial distribution of each protein is studied in three cell lines\nout of a panel of 17; U-2 OS and two additional selected to have the highest\nRNA expression level of the corresponding gene. Each antibody-cell line\n‘sample’ is then imaged to produce a minimum of two images per sample\n(average 2.93 images per sample). Each ‘image’ in the HPA Cell Atlas consists\nof four channels acquired sequentially with a Leica SP5 confocal microscope\n(DM6000CS) equipped with a 63× HCX PL APO 1.40 oil CS objective (Leica\nMicrosystems). The settings for each image were as followed: Pinhole 1 Airy\nunit, 16bit acquisition and a pixel size of 0.08 μm. The detector gain measuring\nthe signal of each antibody was adjusted to a maximum of 800 V to avoid\nstrong background noise. A small part of the plates was imaged automatically\nusing the MatrixScreener M3 in LAS AF software (Leica Microsystem). Here,\nz-stacks at six FOVs were acquired. False-colored channels represent the protein\nof interest (green), DAPI labeling of the nucleus (blue), microtubules\n(red), and the endoplasmic reticulum (yellow). Each channel is stored as a\nseparate 2,048 × 2,048 16-bit ome-tiff.\"\n\nThere is a lot of information in the possession of the HPA people that would make this challenge much more useful, imo. Perhaps they will provide some of it for the next stage. Neither they, nor anyone else, would try to do what this challenge is asking with only the information given as part of basic research - but I suppose that keeps things interesting.",
    "428658": "OK while we're at it, here is another question. I have spent a little time looking at the big tif images and they appear to me to be spectacular in their improved resolution. The question is: for humans trying to work with these images for protein classification, how useful is the extra resolution?",
    "428799": "Resolution is everything, how can we know something is there if we can't see it? Of course there are many ways to infer the existence of something without direct observation, but biology is an empirical science so no amount of data is too much, ideally. 2080x2080 TIFs are very modest really, compared to what microscopes like the one used to acquire these data are capable of producing. \n\nWe mentioned depth information before, the labs that have and use these microscopes are acquiring raw image data at far higher resolutions than 2080. When you bring in the z-stacks you will get up to maybe a dozen or more images of the same field at different focal planes (depending on the equipment and skill of the operator), add the color filters and you end up with some high-dimensional volumetric data, which is great, but it becomes a lot to make sense of. They just gave us those 4 false-colored channels, but there can be quite a bit more than 4 for the same field--we know the absorbance and emission spectra of the fluorescent tags we use and there are filters capable of sufficiently separating wavelengths of ~10nm difference. \n\nAll this is to say that cells are very small, and the 10's of thousands of unique proteins are much smaller and while the cell is alive, they are constantly in motion, either being trafficked around the cell, trafficking other things about, or doing their job as scaffolding or as enzymes, often simultaneously. We need all the resolution because we have no idea what could be there until we see it, or see something its causing. \n\nI do routinely use fluorescence microscopy of various types in my research, but my focus now is on electron microscopy as there are hard limits on the resolution of light microscopy and what I am studying cannot be quantitatively measured any other way. Large confocal stacks can be quite large, but I'm used to working with hundreds of ~500MB serial single images around 30k x 30k. Different tools for different problems, but I really enjoy all the pretty pictures of cells they gave us, now if only I could make some sense of them.\ncheers",
    "428876": "Thank you! Your answers cleared up quite a bit. I understand that the filters will help greatly in identifying the proteins, however, there are some 28 subcellular locations and 3 filters - for the 25 remaining locations, how are they supposed to be identified and classified? Additionally, does each protein image specify what protein is to be identified and the cell type? I am having trouble identifying what really needs to be done here -  are we looking to map every single instance of a certain protein in every location of every type of cell, and do that for many proteins? In any case, without knowing the specific protein to identify, it seems like that would be impossible."
  },
  "source": "meta"
}