{
  "id": 73410,
  "title": "Why this is difficult - with eggsamples",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73410",
  "author_name": "",
  "post_date": "2018-12-02T21:08:35.048341200Z",
  "votes": 54,
  "comment_count": 22,
  "views": 0,
  "content": "<p>If I were to pick a real-life example of what cells look like the most, I’d say eggs. Aside from relative uniformity in shapes and sizes among eggs that doesn’t necessarily translate to cells, other features are clearly correlated. An oval middle part that is clearly separated from the rest and holds the goodies? Check. A proteinaceous and semi-liquid substance that surrounds it? Check. A barrier that separates the contents from the outside? Check. Even if you don’t agree with all aspects of my analogy, I’ll give you one that should bolster my case.</p>\n\n<p><img src=\"https://i.postimg.cc/NM6xbStS/sunny-up.jpg\" alt=\"enter image description here\"></p>\n\n<p>That’s a sunny side up fried egg. Or, if you want to be more technical, an egg that was taken out of its shell and laid down on the heated flat surface. Cells we are looking at are exactly like that: they have been laid down on a flat microscopy slide, and they have been fixed (which is like frying but without heat). The egg yolk looks mostly oval on the surface and usually holds its shape well unless agitated, just like nucleus. The egg white has a general shape but much more variability than yolk, just like the cytoplasm surrounding the nucleus. Any classification of cells starts with these two components, and it isn’t by accident that nucleus/nucleoplasm (class 0) and cytosol/cytoplasm (class 25) are most abundant. So, how hard is it to tell these two apart in cell images?</p>\n\n<p>To answer the question, I did a simple transfer learning exercise with VGG16 as a starting model. Took the top layers off and froze the rest, and pre-trained with my own top layers trained for two categories; that was followed by training with all layers unfrozen. There were ~6000 images in the train set and ~1500 images were used for validation.</p>\n\n<p>Here is an image from the validation group, with cytoplasmic protein signal:</p>\n\n<p><img src=\"https://i.postimg.cc/qqh6jD9y/test3.png\" alt=\"enter image description here\"></p>\n\n<p>The network correctly thinks that this signal is in cytosol (0.99999 probability). Here is what the network actually learned:</p>\n\n<p><img src=\"https://i.postimg.cc/ZndxsxgB/verdict3.png\" alt=\"enter image description here\"></p>\n\n<p>The left side shows top 5 feature areas that lead to conclusion that the signal is not in nucleus. The right side shows top 10 feature areas leading to cytoplasm conclusion (green) or those leading to not-nucleus conclusion (pink). It seems like the pattern is learned really well, as long as we are dealing with a nice field of cells.</p>\n\n<p>But here is the thing: this classifier is only ~97% accurate. Take a 5-year old and show her 100 fried eggs sunny side up, and I’ll bet you that she will tell yolks and whites apart with 100% accuracy. Why isn’t our classifier able to better tell apart the two most prominent cellular features?</p>\n\n<p>I took activations from the penultimate network layer and embedded validation data points using t-SNE. Cytosol images are red, nucleus is blue. This particular representation is meant to prove <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72795\"><strong>my earlier point</strong></a> that faces are everywhere.</p>\n\n<p><img src=\"https://i.postimg.cc/7LqcYmb3/cells-smiley-02.gif\" alt=\"enter image description here\"></p>\n\n<p>For an actual embedding, see the attached gif file. It should be clear that there are two nicely defined clusters that have only few stray points – this is expected based on 97% accuracy. It is those strays that interest me, so I laid down the points into a grid that’s easier to look at:</p>\n\n<p><img src=\"https://i.postimg.cc/D0DhcnGd/cells-square-grid-01.png\" alt=\"enter image description here\"></p>\n\n<p>Now we put together cell images to match the above grid:</p>\n\n<p><img src=\"https://i.postimg.cc/hvFWJpRp/scaled-custom-binary-rgb-tsne.png\" alt=\"enter image description here\"></p>\n\n<p>This image is too small to see the details, so I will attach a higher-resolution image if you want to take a look. What I focus on is those stray data points, or those that are on the boundaries between the two colors. I’ll let you make your own conclusions after you inspect the larger image, but I will make two points here. First, at least some of our misclassified images are due to technical problems. Here is an image that you can find by starting from the upper left corner, and going 5 images in X direction (to the right) and 24 in Y direction (down). This image is on the boundary, has a nuclear signal, but is classified as cytoplasmic (0.83 probability).</p>\n\n<p><img src=\"https://i.postimg.cc/NfdtD86f/6f062840-bba9-11e8-b2ba-ac1f6b6435d0-rgb.png\" alt=\"enter image description here\"></p>\n\n<p>There is a color shift in this image and the channels do not line up. It has been discussed before <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70608\"><strong>here</strong></a> and I thought it would not be very prevalent. However, I have since found at least 5 shifted images, and that’s without looking very hard for them. Anyway, when this happens, it totally confuses the network.</p>\n\n<p><img src=\"https://i.postimg.cc/L5bnrccN/verdict5.png\" alt=\"enter image description here\"></p>\n\n<p>I still think this shouldn’t happen often enough to have a major effect.</p>\n\n<p>My second point has to do with natural heterogeneity of these images. I will blow up two sections from the grid above: one from the upper right corner (nuclear images) and from lower right (cytoplasmic images). These are predicted with high confidence and well separated from each other, so one would expect a clear and uniform pattern in them.</p>\n\n<p>Nuclear images first:</p>\n\n<p><img src=\"https://i.postimg.cc/NGz8hS73/cutout-custom-binary-rgb-tsne.png\" alt=\"enter image description here\"></p>\n\n<p>These images have several things in common: medium- to small-sized cells, lots of blue and red and much less green color. That makes sense, because nuclei are smaller than cytoplasm in terms of surface, so there won’t be much outright green signal. Here is where the natural differences in protein expression kick in: if protein of interest is hugely overexpressed, its green signal would simply overpower the blue coming from DNA, and we’d see green nuclei. When its signal is comparable to DAPI (blue), we’d see cyan/turquoise color from equal blue/green mixing. Finally, poorly expressed proteins would give very little green signal, so our nuclei would be blue. Note that “poorly” is meant in relative terms here, as even well-expressed proteins can’t always match the DAPI signal. In reality, we see all these scenarios in the image above, with straight-up blue being most prominent. It is actually very impressive that a network picks up these disparate levels of blue and green mixing mean the same thing.</p>\n\n<p>Now cytoplasmic images:</p>\n\n<p><img src=\"https://i.postimg.cc/yxMSJqLW/cutout-custom-binary-rgb-tsne-2.png\" alt=\"enter image description here\"></p>\n\n<p>These are mostly larger cells (by that I mean larger images of cells), with lots of blue and green and little pure red color. Different levels of green and red mixing give us yellow and orange shades in cytoplasm. In relative terms, the closer we have the cytoplasm to pure red color, the less our protein of interest is expressed relative to microtubules. Conversely, a cytoplasmic signal that is closer to green means that our protein is over-expressed relative to microtubules.</p>\n\n<p>I’ll pre-emptively answer the question: there is no defined set of rules governing protein expression, even when they are expressed in the same system (same cells, same promoters). Needs of the cell for proteins vary greatly as well, and their numbers vary over few orders of magnitude. Beyond that, small proteins express better than large ones, mostly because they are easier to make. I dealt with proteins that give huge signal 8 hours after transfection, while others take 24 hours. Add to that morphological differences between cell types, different growth rates, different expression efficiencies, different antibodies used for detection, different technicians performing experiments, different microscopes, different magnifications, and we are classifying something that is not as simple as eggs that are sunny side up:</p>\n\n<p><img src=\"https://i.postimg.cc/8cysW7Ny/aid1693528-v4-728px-Make-Deep-Fried-Eggs-Step-7-Version-2.jpg\" alt=\"enter image description here\"></p>\n\n<p>Now add 26 classes on top of nucleus and cytoplasm, and suddenly our challenge becomes:</p>\n\n<p><img src=\"https://i.postimg.cc/NMcywhdx/Fried-Egg-Cover.png\" alt=\"enter image description here\"></p>\n\n<p>Turns out telling yolks and whites apart is tricky business. Maybe I should stop whinnying about 97% accuracy.</p>",
  "messages": [
    {
      "id": "431745",
      "postDate": "12/02/2018 21:08:35",
      "content": "<p>If I were to pick a real-life example of what cells look like the most, I’d say eggs. Aside from relative uniformity in shapes and sizes among eggs that doesn’t necessarily translate to cells, other features are clearly correlated. An oval middle part that is clearly separated from the rest and holds the goodies? Check. A proteinaceous and semi-liquid substance that surrounds it? Check. A barrier that separates the contents from the outside? Check. Even if you don’t agree with all aspects of my analogy, I’ll give you one that should bolster my case.</p>\n\n<p><img src=\"https://i.postimg.cc/NM6xbStS/sunny-up.jpg\" alt=\"enter image description here\"></p>\n\n<p>That’s a sunny side up fried egg. Or, if you want to be more technical, an egg that was taken out of its shell and laid down on the heated flat surface. Cells we are looking at are exactly like that: they have been laid down on a flat microscopy slide, and they have been fixed (which is like frying but without heat). The egg yolk looks mostly oval on the surface and usually holds its shape well unless agitated, just like nucleus. The egg white has a general shape but much more variability than yolk, just like the cytoplasm surrounding the nucleus. Any classification of cells starts with these two components, and it isn’t by accident that nucleus/nucleoplasm (class 0) and cytosol/cytoplasm (class 25) are most abundant. So, how hard is it to tell these two apart in cell images?</p>\n\n<p>To answer the question, I did a simple transfer learning exercise with VGG16 as a starting model. Took the top layers off and froze the rest, and pre-trained with my own top layers trained for two categories; that was followed by training with all layers unfrozen. There were ~6000 images in the train set and ~1500 images were used for validation.</p>\n\n<p>Here is an image from the validation group, with cytoplasmic protein signal:</p>\n\n<p><img src=\"https://i.postimg.cc/qqh6jD9y/test3.png\" alt=\"enter image description here\"></p>\n\n<p>The network correctly thinks that this signal is in cytosol (0.99999 probability). Here is what the network actually learned:</p>\n\n<p><img src=\"https://i.postimg.cc/ZndxsxgB/verdict3.png\" alt=\"enter image description here\"></p>\n\n<p>The left side shows top 5 feature areas that lead to conclusion that the signal is not in nucleus. The right side shows top 10 feature areas leading to cytoplasm conclusion (green) or those leading to not-nucleus conclusion (pink). It seems like the pattern is learned really well, as long as we are dealing with a nice field of cells.</p>\n\n<p>But here is the thing: this classifier is only ~97% accurate. Take a 5-year old and show her 100 fried eggs sunny side up, and I’ll bet you that she will tell yolks and whites apart with 100% accuracy. Why isn’t our classifier able to better tell apart the two most prominent cellular features?</p>\n\n<p>I took activations from the penultimate network layer and embedded validation data points using t-SNE. Cytosol images are red, nucleus is blue. This particular representation is meant to prove <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72795\"><strong>my earlier point</strong></a> that faces are everywhere.</p>\n\n<p><img src=\"https://i.postimg.cc/7LqcYmb3/cells-smiley-02.gif\" alt=\"enter image description here\"></p>\n\n<p>For an actual embedding, see the attached gif file. It should be clear that there are two nicely defined clusters that have only few stray points – this is expected based on 97% accuracy. It is those strays that interest me, so I laid down the points into a grid that’s easier to look at:</p>\n\n<p><img src=\"https://i.postimg.cc/D0DhcnGd/cells-square-grid-01.png\" alt=\"enter image description here\"></p>\n\n<p>Now we put together cell images to match the above grid:</p>\n\n<p><img src=\"https://i.postimg.cc/hvFWJpRp/scaled-custom-binary-rgb-tsne.png\" alt=\"enter image description here\"></p>\n\n<p>This image is too small to see the details, so I will attach a higher-resolution image if you want to take a look. What I focus on is those stray data points, or those that are on the boundaries between the two colors. I’ll let you make your own conclusions after you inspect the larger image, but I will make two points here. First, at least some of our misclassified images are due to technical problems. Here is an image that you can find by starting from the upper left corner, and going 5 images in X direction (to the right) and 24 in Y direction (down). This image is on the boundary, has a nuclear signal, but is classified as cytoplasmic (0.83 probability).</p>\n\n<p><img src=\"https://i.postimg.cc/NfdtD86f/6f062840-bba9-11e8-b2ba-ac1f6b6435d0-rgb.png\" alt=\"enter image description here\"></p>\n\n<p>There is a color shift in this image and the channels do not line up. It has been discussed before <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70608\"><strong>here</strong></a> and I thought it would not be very prevalent. However, I have since found at least 5 shifted images, and that’s without looking very hard for them. Anyway, when this happens, it totally confuses the network.</p>\n\n<p><img src=\"https://i.postimg.cc/L5bnrccN/verdict5.png\" alt=\"enter image description here\"></p>\n\n<p>I still think this shouldn’t happen often enough to have a major effect.</p>\n\n<p>My second point has to do with natural heterogeneity of these images. I will blow up two sections from the grid above: one from the upper right corner (nuclear images) and from lower right (cytoplasmic images). These are predicted with high confidence and well separated from each other, so one would expect a clear and uniform pattern in them.</p>\n\n<p>Nuclear images first:</p>\n\n<p><img src=\"https://i.postimg.cc/NGz8hS73/cutout-custom-binary-rgb-tsne.png\" alt=\"enter image description here\"></p>\n\n<p>These images have several things in common: medium- to small-sized cells, lots of blue and red and much less green color. That makes sense, because nuclei are smaller than cytoplasm in terms of surface, so there won’t be much outright green signal. Here is where the natural differences in protein expression kick in: if protein of interest is hugely overexpressed, its green signal would simply overpower the blue coming from DNA, and we’d see green nuclei. When its signal is comparable to DAPI (blue), we’d see cyan/turquoise color from equal blue/green mixing. Finally, poorly expressed proteins would give very little green signal, so our nuclei would be blue. Note that “poorly” is meant in relative terms here, as even well-expressed proteins can’t always match the DAPI signal. In reality, we see all these scenarios in the image above, with straight-up blue being most prominent. It is actually very impressive that a network picks up these disparate levels of blue and green mixing mean the same thing.</p>\n\n<p>Now cytoplasmic images:</p>\n\n<p><img src=\"https://i.postimg.cc/yxMSJqLW/cutout-custom-binary-rgb-tsne-2.png\" alt=\"enter image description here\"></p>\n\n<p>These are mostly larger cells (by that I mean larger images of cells), with lots of blue and green and little pure red color. Different levels of green and red mixing give us yellow and orange shades in cytoplasm. In relative terms, the closer we have the cytoplasm to pure red color, the less our protein of interest is expressed relative to microtubules. Conversely, a cytoplasmic signal that is closer to green means that our protein is over-expressed relative to microtubules.</p>\n\n<p>I’ll pre-emptively answer the question: there is no defined set of rules governing protein expression, even when they are expressed in the same system (same cells, same promoters). Needs of the cell for proteins vary greatly as well, and their numbers vary over few orders of magnitude. Beyond that, small proteins express better than large ones, mostly because they are easier to make. I dealt with proteins that give huge signal 8 hours after transfection, while others take 24 hours. Add to that morphological differences between cell types, different growth rates, different expression efficiencies, different antibodies used for detection, different technicians performing experiments, different microscopes, different magnifications, and we are classifying something that is not as simple as eggs that are sunny side up:</p>\n\n<p><img src=\"https://i.postimg.cc/8cysW7Ny/aid1693528-v4-728px-Make-Deep-Fried-Eggs-Step-7-Version-2.jpg\" alt=\"enter image description here\"></p>\n\n<p>Now add 26 classes on top of nucleus and cytoplasm, and suddenly our challenge becomes:</p>\n\n<p><img src=\"https://i.postimg.cc/NMcywhdx/Fried-Egg-Cover.png\" alt=\"enter image description here\"></p>\n\n<p>Turns out telling yolks and whites apart is tricky business. Maybe I should stop whinnying about 97% accuracy.</p>",
      "rawMarkdown": "If I were to pick a real-life example of what cells look like the most, I’d say eggs. Aside from relative uniformity in shapes and sizes among eggs that doesn’t necessarily translate to cells, other features are clearly correlated. An oval middle part that is clearly separated from the rest and holds the goodies? Check. A proteinaceous and semi-liquid substance that surrounds it? Check. A barrier that separates the contents from the outside? Check. Even if you don’t agree with all aspects of my analogy, I’ll give you one that should bolster my case.\n\n![enter image description here][1]\n\nThat’s a sunny side up fried egg. Or, if you want to be more technical, an egg that was taken out of its shell and laid down on the heated flat surface. Cells we are looking at are exactly like that: they have been laid down on a flat microscopy slide, and they have been fixed (which is like frying but without heat). The egg yolk looks mostly oval on the surface and usually holds its shape well unless agitated, just like nucleus. The egg white has a general shape but much more variability than yolk, just like the cytoplasm surrounding the nucleus. Any classification of cells starts with these two components, and it isn’t by accident that nucleus/nucleoplasm (class 0) and cytosol/cytoplasm (class 25) are most abundant. So, how hard is it to tell these two apart in cell images?\n\nTo answer the question, I did a simple transfer learning exercise with VGG16 as a starting model. Took the top layers off and froze the rest, and pre-trained with my own top layers trained for two categories; that was followed by training with all layers unfrozen. There were ~6000 images in the train set and ~1500 images were used for validation.\n\nHere is an image from the validation group, with cytoplasmic protein signal:\n\n![enter image description here][2]\n\nThe network correctly thinks that this signal is in cytosol (0.99999 probability). Here is what the network actually learned:\n\n![enter image description here][3]\n\nThe left side shows top 5 feature areas that lead to conclusion that the signal is not in nucleus. The right side shows top 10 feature areas leading to cytoplasm conclusion (green) or those leading to not-nucleus conclusion (pink). It seems like the pattern is learned really well, as long as we are dealing with a nice field of cells.\n\nBut here is the thing: this classifier is only ~97% accurate. Take a 5-year old and show her 100 fried eggs sunny side up, and I’ll bet you that she will tell yolks and whites apart with 100% accuracy. Why isn’t our classifier able to better tell apart the two most prominent cellular features?\n\nI took activations from the penultimate network layer and embedded validation data points using t-SNE. Cytosol images are red, nucleus is blue. This particular representation is meant to prove [__my earlier point__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72795) that faces are everywhere.\n\n![enter image description here][4]\n\nFor an actual embedding, see the attached gif file. It should be clear that there are two nicely defined clusters that have only few stray points – this is expected based on 97% accuracy. It is those strays that interest me, so I laid down the points into a grid that’s easier to look at:\n\n![enter image description here][5]\n\nNow we put together cell images to match the above grid:\n\n![enter image description here][6]\n\nThis image is too small to see the details, so I will attach a higher-resolution image if you want to take a look. What I focus on is those stray data points, or those that are on the boundaries between the two colors. I’ll let you make your own conclusions after you inspect the larger image, but I will make two points here. First, at least some of our misclassified images are due to technical problems. Here is an image that you can find by starting from the upper left corner, and going 5 images in X direction (to the right) and 24 in Y direction (down). This image is on the boundary, has a nuclear signal, but is classified as cytoplasmic (0.83 probability).\n\n![enter image description here][7]\n\nThere is a color shift in this image and the channels do not line up. It has been discussed before [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70608) and I thought it would not be very prevalent. However, I have since found at least 5 shifted images, and that’s without looking very hard for them. Anyway, when this happens, it totally confuses the network.\n\n![enter image description here][8]\n\nI still think this shouldn’t happen often enough to have a major effect.\n\nMy second point has to do with natural heterogeneity of these images. I will blow up two sections from the grid above: one from the upper right corner (nuclear images) and from lower right (cytoplasmic images). These are predicted with high confidence and well separated from each other, so one would expect a clear and uniform pattern in them.\n\nNuclear images first:\n\n![enter image description here][9]\n\nThese images have several things in common: medium- to small-sized cells, lots of blue and red and much less green color. That makes sense, because nuclei are smaller than cytoplasm in terms of surface, so there won’t be much outright green signal. Here is where the natural differences in protein expression kick in: if protein of interest is hugely overexpressed, its green signal would simply overpower the blue coming from DNA, and we’d see green nuclei. When its signal is comparable to DAPI (blue), we’d see cyan/turquoise color from equal blue/green mixing. Finally, poorly expressed proteins would give very little green signal, so our nuclei would be blue. Note that “poorly” is meant in relative terms here, as even well-expressed proteins can’t always match the DAPI signal. In reality, we see all these scenarios in the image above, with straight-up blue being most prominent. It is actually very impressive that a network picks up these disparate levels of blue and green mixing mean the same thing.\n\nNow cytoplasmic images:\n\n![enter image description here][10]\n\nThese are mostly larger cells (by that I mean larger images of cells), with lots of blue and green and little pure red color. Different levels of green and red mixing give us yellow and orange shades in cytoplasm. In relative terms, the closer we have the cytoplasm to pure red color, the less our protein of interest is expressed relative to microtubules. Conversely, a cytoplasmic signal that is closer to green means that our protein is over-expressed relative to microtubules.\n\nI’ll pre-emptively answer the question: there is no defined set of rules governing protein expression, even when they are expressed in the same system (same cells, same promoters). Needs of the cell for proteins vary greatly as well, and their numbers vary over few orders of magnitude. Beyond that, small proteins express better than large ones, mostly because they are easier to make. I dealt with proteins that give huge signal 8 hours after transfection, while others take 24 hours. Add to that morphological differences between cell types, different growth rates, different expression efficiencies, different antibodies used for detection, different technicians performing experiments, different microscopes, different magnifications, and we are classifying something that is not as simple as eggs that are sunny side up:\n\n![enter image description here][11]\n\nNow add 26 classes on top of nucleus and cytoplasm, and suddenly our challenge becomes:\n\n![enter image description here][12]\n\nTurns out telling yolks and whites apart is tricky business. Maybe I should stop whinnying about 97% accuracy.\n\n  [1]: https://i.postimg.cc/NM6xbStS/sunny-up.jpg\n  [2]: https://i.postimg.cc/qqh6jD9y/test3.png\n  [3]: https://i.postimg.cc/ZndxsxgB/verdict3.png\n  [4]: https://i.postimg.cc/7LqcYmb3/cells-smiley-02.gif\n  [5]: https://i.postimg.cc/D0DhcnGd/cells-square-grid-01.png\n  [6]: https://i.postimg.cc/hvFWJpRp/scaled-custom-binary-rgb-tsne.png\n  [7]: https://i.postimg.cc/NfdtD86f/6f062840-bba9-11e8-b2ba-ac1f6b6435d0-rgb.png\n  [8]: https://i.postimg.cc/L5bnrccN/verdict5.png\n  [9]: https://i.postimg.cc/NGz8hS73/cutout-custom-binary-rgb-tsne.png\n  [10]: https://i.postimg.cc/yxMSJqLW/cutout-custom-binary-rgb-tsne-2.png\n  [11]: https://i.postimg.cc/8cysW7Ny/aid1693528-v4-728px-Make-Deep-Fried-Eggs-Step-7-Version-2.jpg\n  [12]: https://i.postimg.cc/NMcywhdx/Fried-Egg-Cover.png",
      "votes": null
    },
    {
      "id": "431746",
      "postDate": "12/02/2018 21:10:20",
      "content": "<p>Full-size grid is attached here.</p>",
      "rawMarkdown": "Full-size grid is attached here.",
      "votes": null
    },
    {
      "id": "431747",
      "postDate": "12/02/2018 21:11:02",
      "content": "<p>t-SNE animation is attached here.</p>",
      "rawMarkdown": "t-SNE animation is attached here.",
      "votes": null
    },
    {
      "id": "432367",
      "postDate": "12/03/2018 19:19:22",
      "content": "<p>huh?</p>",
      "rawMarkdown": "huh?",
      "votes": null
    },
    {
      "id": "432379",
      "postDate": "12/03/2018 19:35:40",
      "content": "<p><a href=\"/longhorns2102\">@longhorns2102</a> Anything in particular that's confusing you?</p>",
      "rawMarkdown": "longhorns2102 Anything in particular that's confusing you?",
      "votes": null
    },
    {
      "id": "434136",
      "postDate": "12/06/2018 00:19:36",
      "content": "<p>Nice post!\nFirst question - what is the exact definition of \"expression\"? Does it just mean general contrast, or contrast between green and red, or something else?\nAnd you mention \"These are mostly larger cells (by that I mean larger images of cells)\". So can absolute size be an indicator of type, or is it just an artifact of something else (microscope zoom level or something)? I guess I'm asking can I augment with random zoom/cropping to avoid over-fitting - or will that make it lose some crucial information?</p>",
      "rawMarkdown": "Nice post!\nFirst question - what is the exact definition of \"expression\"? Does it just mean general contrast, or contrast between green and red, or something else?\nAnd you mention \"These are mostly larger cells (by that I mean larger images of cells)\". So can absolute size be an indicator of type, or is it just an artifact of something else (microscope zoom level or something)? I guess I'm asking can I augment with random zoom/cropping to avoid over-fitting - or will that make it lose some crucial information?",
      "votes": null
    },
    {
      "id": "434151",
      "postDate": "12/06/2018 01:01:26",
      "content": "<p>How are you creating the images of feature areas and the tSNE representations? Is there a guide or something that describes how to extract this?</p>",
      "rawMarkdown": "How are you creating the images of feature areas and the tSNE representations? Is there a guide or something that describes how to extract this?",
      "votes": null
    },
    {
      "id": "434157",
      "postDate": "12/06/2018 01:24:16",
      "content": "<p><a href=\"/hshmhashemi\">@hshmhashemi</a> Expression indicates the amount of protein made in cells. The fluorescent signal is proportional to the amount of target proteins. I was making a point that different proteins are not made in cell in equal amounts even under otherwise identical conditions, so that always creates problems with signal normalization.</p>\n\n<p>We have been told that uniform magnification was used for most images. My eyes tell me otherwise. Cells do differ in sizes, but it should not be as much as smallest cells in upper right corner vs. largest cells in lower right corner. There is easily 4-5x size difference, and in my experience that's not the case. Anyway, I suggest asymmetric zoom [0.9, 1.3] (maybe up to 1.5).</p>",
      "rawMarkdown": "hshmhashemi Expression indicates the amount of protein made in cells. The fluorescent signal is proportional to the amount of target proteins. I was making a point that different proteins are not made in cell in equal amounts even under otherwise identical conditions, so that always creates problems with signal normalization.\n\nWe have been told that uniform magnification was used for most images. My eyes tell me otherwise. Cells do differ in sizes, but it should not be as much as smallest cells in upper right corner vs. largest cells in lower right corner. There is easily 4-5x size difference, and in my experience that's not the case. Anyway, I suggest asymmetric zoom [0.9, 1.3] (maybe up to 1.5).",
      "votes": null
    },
    {
      "id": "434162",
      "postDate": "12/06/2018 01:29:54",
      "content": "<p><a href=\"/ldm314\">@ldm314</a> <a href=\"https://github.com/marcotcr/lime\"><strong>Lime</strong></a> was used to explain the images. t-SNE grid creation is described <a href=\"https://github.com/prabodhhere/tsne-grid\"><strong>here</strong></a> and <a href=\"https://github.com/bmcfee/RasterFairy\"><strong>RasterFairy</strong></a> was used for animations.</p>",
      "rawMarkdown": "ldm314 [__Lime__](https://github.com/marcotcr/lime) was used to explain the images. t-SNE grid creation is described [__here__](https://github.com/prabodhhere/tsne-grid) and [__RasterFairy__](https://github.com/bmcfee/RasterFairy) was used for animations.",
      "votes": null
    },
    {
      "id": "434164",
      "postDate": "12/06/2018 01:37:16",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "434208",
      "postDate": "12/06/2018 03:45:22",
      "content": "<p>Nice post！Thanks！</p>",
      "rawMarkdown": "Nice post！Thanks！",
      "votes": null
    },
    {
      "id": "434328",
      "postDate": "12/06/2018 07:50:18",
      "content": "<p>Excellent post, Tilli, thanks! For hshmhashemi, tilii is correct, proteins are made by ribosomes translating mRNA so increased and decreased expression refers to the observable quantity of some protein in this case. What causes changes in expression levels can range from environmental factors to the previous activity states on the cell.</p>\n\n<p>Some of the cells to look significantly different in size, but for any given image I think it is highly unlikely the magnification would be different in different parts of the same image. But this is a fairly big dataset and errors and artifacts always happen. Some of the label imbalance is due to the transient nature of several of the labels, anything to do with cell division in particular. These are cultured human cells and they do divide and grow, but the chances of catching one mid mitosis is just much more unlikely than not.</p>\n\n<p>I don't actually believe knowing the cell types would help very much in this challenge since we dont know what the antibodies are actually targeting, they just want us to identify zones and structures in the cell. Actin is the only actual protein in the entire list of labels.</p>\n\n<p>Also it is certainly true that the level of fluorescence is proportional to expression, its also a function of the specificity and affinity of the antibody. All the cells in a given image will have been incubated in the same set of antibodies, but from that alone it is very tough to infer much that is useful. For example, each label is under no obligation to imply one protein, in fact any label could actually be more than one individual protein, ie, cytosol is most of the cell an there are a lot of labels with only cytosol, but that does not mean only one protein was labeled, it could be the case but we have no way be sure. Even images with multiple labels could actually be showing one or more unique proteins, its quite possible for one protein to be highly expressed in multiple compartments of the cell at once. I wish I could say two or more labels means at least two proteins, but that conclusion is not warranted. </p>\n\n<p>What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.</p>",
      "rawMarkdown": "Excellent post, Tilli, thanks! For hshmhashemi, tilii is correct, proteins are made by ribosomes translating mRNA so increased and decreased expression refers to the observable quantity of some protein in this case. What causes changes in expression levels can range from environmental factors to the previous activity states on the cell.\n\nSome of the cells to look significantly different in size, but for any given image I think it is highly unlikely the magnification would be different in different parts of the same image. But this is a fairly big dataset and errors and artifacts always happen. Some of the label imbalance is due to the transient nature of several of the labels, anything to do with cell division in particular. These are cultured human cells and they do divide and grow, but the chances of catching one mid mitosis is just much more unlikely than not.\n\nI don't actually believe knowing the cell types would help very much in this challenge since we dont know what the antibodies are actually targeting, they just want us to identify zones and structures in the cell. Actin is the only actual protein in the entire list of labels.\n\nAlso it is certainly true that the level of fluorescence is proportional to expression, its also a function of the specificity and affinity of the antibody. All the cells in a given image will have been incubated in the same set of antibodies, but from that alone it is very tough to infer much that is useful. For example, each label is under no obligation to imply one protein, in fact any label could actually be more than one individual protein, ie, cytosol is most of the cell an there are a lot of labels with only cytosol, but that does not mean only one protein was labeled, it could be the case but we have no way be sure. Even images with multiple labels could actually be showing one or more unique proteins, its quite possible for one protein to be highly expressed in multiple compartments of the cell at once. I wish I could say two or more labels means at least two proteins, but that conclusion is not warranted. \n\nWhat I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.",
      "votes": null
    },
    {
      "id": "434715",
      "postDate": "12/06/2018 20:51:08",
      "content": "<blockquote>\n  <p>What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.</p>\n</blockquote>\n\n<p>It may be a good idea to extract individual cells, but I would not try to define a subset that meets any standard. First, what would that standard be? Second, this is meant to be a general classifier, and what happens when you have to classify an image where no cell meets that standard?</p>\n\n<p>Don't think that you need to tell your models to ignore positional variance. There is useful information in that feature. If anything, I think one should be using image shifts and rotations to increase positional variance and improve the generalization.</p>",
      "rawMarkdown": "&gt; What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.\n\nIt may be a good idea to extract individual cells, but I would not try to define a subset that meets any standard. First, what would that standard be? Second, this is meant to be a general classifier, and what happens when you have to classify an image where no cell meets that standard?\n\nDon't think that you need to tell your models to ignore positional variance. There is useful information in that feature. If anything, I think one should be using image shifts and rotations to increase positional variance and improve the generalization.",
      "votes": null
    },
    {
      "id": "434726",
      "postDate": "12/06/2018 21:09:35",
      "content": "<p>I made a case before that differences in cell sizes in images are unlikely to reflect natural variability. Here is a boundary group of cells with nuclear and cytoplasmic signals (around 5 o'clock in the square grid image above):</p>\n\n<p><img src=\"https://i.ibb.co/4YYFB0S/cell-sizes.png\" alt=\"enter image description here\"></p>\n\n<p>Cells in top two rows and the right side of the third row have nuclear signals, while the rest are cytoplasmic. I estimated earlier that size differences between cells were on the order of 4-5x, but this image alone shows that the size range is closer to 10x. To the best of my knowledge, there is no such size difference among human cells commonly used for imaging. I still think that our images were not collected at uniform magnification.</p>",
      "rawMarkdown": "I made a case before that differences in cell sizes in images are unlikely to reflect natural variability. Here is a boundary group of cells with nuclear and cytoplasmic signals (around 5 o'clock in the square grid image above):\n\n![enter image description here][1]\n\nCells in top two rows and the right side of the third row have nuclear signals, while the rest are cytoplasmic. I estimated earlier that size differences between cells were on the order of 4-5x, but this image alone shows that the size range is closer to 10x. To the best of my knowledge, there is no such size difference among human cells commonly used for imaging. I still think that our images were not collected at uniform magnification.\n\n  [1]: https://i.ibb.co/4YYFB0S/cell-sizes.png",
      "votes": null
    },
    {
      "id": "434747",
      "postDate": "12/06/2018 21:58:12",
      "content": "<p>The only domain knowledge I have is that mitochondria is the powerhouse of the cell.</p>\n\n<p>I guess I won't get far.</p>",
      "rawMarkdown": "The only domain knowledge I have is that mitochondria is the powerhouse of the cell.\n\nI guess I won't get far.",
      "votes": null
    },
    {
      "id": "434757",
      "postDate": "12/06/2018 22:25:10",
      "content": "<blockquote>\n  <p>I guess I won't get far.</p>\n</blockquote>\n\n<p>Your present LB spot argues otherwise ;-) I would guess - and hope there is nobody offended by it - that plenty of our top 50 competitors don't have full domain knowledge. My own domain knowledge is of limited use until I actually decide to compete.</p>",
      "rawMarkdown": "&gt; I guess I won't get far.\n\nYour present LB spot argues otherwise ;-) I would guess - and hope there is nobody offended by it - that plenty of our top 50 competitors don't have full domain knowledge. My own domain knowledge is of limited use until I actually decide to compete.",
      "votes": null
    },
    {
      "id": "434763",
      "postDate": "12/06/2018 22:45:11",
      "content": "<p>Ever since the beginning of this competition i felt that this challenge is actually more suitable for object detection. One of the problems of the multilabel approach is that the network must put \"extra effort\" into learning which is which. Putting it into eggs, suppised we have lots of photos of fried eggs of all kinds: duck, pigeon, ostrich and also has browns and scrabled eggs. Each photo shows several of these on the hot plate. A 5 yo child will have lots of trouble finding out which is which if all he had was a title like \"duck and pigeon egg\". However, i can understand the amount of work that involve markup for object detection.... </p>",
      "rawMarkdown": "Ever since the beginning of this competition i felt that this challenge is actually more suitable for object detection. One of the problems of the multilabel approach is that the network must put \"extra effort\" into learning which is which. Putting it into eggs, suppised we have lots of photos of fried eggs of all kinds: duck, pigeon, ostrich and also has browns and scrabled eggs. Each photo shows several of these on the hot plate. A 5 yo child will have lots of trouble finding out which is which if all he had was a title like \"duck and pigeon egg\". However, i can understand the amount of work that involve markup for object detection....",
      "votes": null
    },
    {
      "id": "434850",
      "postDate": "12/07/2018 03:24:08",
      "content": "<p>Yes, those are good points. The images we are given contain multiple cells, so training on only single cell examples of the label might completely fail to generalize to the arbitrary cell numbers/image they want us to predict on. I don't have a great intuition for this sort of thing yet, though - I'm inclined to believe you.</p>",
      "rawMarkdown": "Yes, those are good points. The images we are given contain multiple cells, so training on only single cell examples of the label might completely fail to generalize to the arbitrary cell numbers/image they want us to predict on. I don't have a great intuition for this sort of thing yet, though - I'm inclined to believe you.",
      "votes": null
    },
    {
      "id": "435187",
      "postDate": "12/07/2018 16:34:07",
      "content": "<p>great @Tilii!</p>\n\n<p>thanks for sharing!</p>",
      "rawMarkdown": "great @Tilii!\n\nthanks for sharing!",
      "votes": null
    },
    {
      "id": "435203",
      "postDate": "12/07/2018 16:59:25",
      "content": "<p>I see you too, are a man of culture</p>",
      "rawMarkdown": "I see you too, are a man of culture",
      "votes": null
    },
    {
      "id": "435815",
      "postDate": "12/08/2018 20:39:05",
      "content": "<p>Quite insightful! Thanks for sharing :)</p>",
      "rawMarkdown": "Quite insightful! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "437280",
      "postDate": "12/11/2018 17:07:23",
      "content": "<p>a chicken egg doesn't just look like a single cell; it is one.  hence the similarity</p>",
      "rawMarkdown": "a chicken egg doesn't just look like a single cell; it is one.  hence the similarity",
      "votes": null
    },
    {
      "id": "443028",
      "postDate": "12/20/2018 22:52:41",
      "content": "<p>Just wanted to say this is a brilliant post - thank you for sharing.</p>",
      "rawMarkdown": "Just wanted to say this is a brilliant post - thank you for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 431746,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "12/02/2018 21:10:20",
      "content": "<p>Full-size grid is attached here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 431747,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "12/02/2018 21:11:02",
      "content": "<p>t-SNE animation is attached here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 432367,
      "author_name": "longhorns2102",
      "author_url": "",
      "post_date": "12/03/2018 19:19:22",
      "content": "<p>huh?</p>",
      "votes": null,
      "replies": [
        {
          "id": 432379,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/03/2018 19:35:40",
          "content": "<p><a href=\"/longhorns2102\">@longhorns2102</a> Anything in particular that's confusing you?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434136,
      "author_name": "hshmhashemi",
      "author_url": "",
      "post_date": "12/06/2018 00:19:36",
      "content": "<p>Nice post!\nFirst question - what is the exact definition of \"expression\"? Does it just mean general contrast, or contrast between green and red, or something else?\nAnd you mention \"These are mostly larger cells (by that I mean larger images of cells)\". So can absolute size be an indicator of type, or is it just an artifact of something else (microscope zoom level or something)? I guess I'm asking can I augment with random zoom/cropping to avoid over-fitting - or will that make it lose some crucial information?</p>",
      "votes": null,
      "replies": [
        {
          "id": 434157,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/06/2018 01:24:16",
          "content": "<p><a href=\"/hshmhashemi\">@hshmhashemi</a> Expression indicates the amount of protein made in cells. The fluorescent signal is proportional to the amount of target proteins. I was making a point that different proteins are not made in cell in equal amounts even under otherwise identical conditions, so that always creates problems with signal normalization.</p>\n\n<p>We have been told that uniform magnification was used for most images. My eyes tell me otherwise. Cells do differ in sizes, but it should not be as much as smallest cells in upper right corner vs. largest cells in lower right corner. There is easily 4-5x size difference, and in my experience that's not the case. Anyway, I suggest asymmetric zoom [0.9, 1.3] (maybe up to 1.5).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434328,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "12/06/2018 07:50:18",
          "content": "<p>Excellent post, Tilli, thanks! For hshmhashemi, tilii is correct, proteins are made by ribosomes translating mRNA so increased and decreased expression refers to the observable quantity of some protein in this case. What causes changes in expression levels can range from environmental factors to the previous activity states on the cell.</p>\n\n<p>Some of the cells to look significantly different in size, but for any given image I think it is highly unlikely the magnification would be different in different parts of the same image. But this is a fairly big dataset and errors and artifacts always happen. Some of the label imbalance is due to the transient nature of several of the labels, anything to do with cell division in particular. These are cultured human cells and they do divide and grow, but the chances of catching one mid mitosis is just much more unlikely than not.</p>\n\n<p>I don't actually believe knowing the cell types would help very much in this challenge since we dont know what the antibodies are actually targeting, they just want us to identify zones and structures in the cell. Actin is the only actual protein in the entire list of labels.</p>\n\n<p>Also it is certainly true that the level of fluorescence is proportional to expression, its also a function of the specificity and affinity of the antibody. All the cells in a given image will have been incubated in the same set of antibodies, but from that alone it is very tough to infer much that is useful. For example, each label is under no obligation to imply one protein, in fact any label could actually be more than one individual protein, ie, cytosol is most of the cell an there are a lot of labels with only cytosol, but that does not mean only one protein was labeled, it could be the case but we have no way be sure. Even images with multiple labels could actually be showing one or more unique proteins, its quite possible for one protein to be highly expressed in multiple compartments of the cell at once. I wish I could say two or more labels means at least two proteins, but that conclusion is not warranted. </p>\n\n<p>What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434715,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/06/2018 20:51:08",
          "content": "<blockquote>\n  <p>What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.</p>\n</blockquote>\n\n<p>It may be a good idea to extract individual cells, but I would not try to define a subset that meets any standard. First, what would that standard be? Second, this is meant to be a general classifier, and what happens when you have to classify an image where no cell meets that standard?</p>\n\n<p>Don't think that you need to tell your models to ignore positional variance. There is useful information in that feature. If anything, I think one should be using image shifts and rotations to increase positional variance and improve the generalization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434850,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "12/07/2018 03:24:08",
          "content": "<p>Yes, those are good points. The images we are given contain multiple cells, so training on only single cell examples of the label might completely fail to generalize to the arbitrary cell numbers/image they want us to predict on. I don't have a great intuition for this sort of thing yet, though - I'm inclined to believe you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434151,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "12/06/2018 01:01:26",
      "content": "<p>How are you creating the images of feature areas and the tSNE representations? Is there a guide or something that describes how to extract this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 434162,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/06/2018 01:29:54",
          "content": "<p><a href=\"/ldm314\">@ldm314</a> <a href=\"https://github.com/marcotcr/lime\"><strong>Lime</strong></a> was used to explain the images. t-SNE grid creation is described <a href=\"https://github.com/prabodhhere/tsne-grid\"><strong>here</strong></a> and <a href=\"https://github.com/bmcfee/RasterFairy\"><strong>RasterFairy</strong></a> was used for animations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434164,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/06/2018 01:37:16",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434208,
      "author_name": "zedali",
      "author_url": "",
      "post_date": "12/06/2018 03:45:22",
      "content": "<p>Nice post！Thanks！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434726,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "12/06/2018 21:09:35",
      "content": "<p>I made a case before that differences in cell sizes in images are unlikely to reflect natural variability. Here is a boundary group of cells with nuclear and cytoplasmic signals (around 5 o'clock in the square grid image above):</p>\n\n<p><img src=\"https://i.ibb.co/4YYFB0S/cell-sizes.png\" alt=\"enter image description here\"></p>\n\n<p>Cells in top two rows and the right side of the third row have nuclear signals, while the rest are cytoplasmic. I estimated earlier that size differences between cells were on the order of 4-5x, but this image alone shows that the size range is closer to 10x. To the best of my knowledge, there is no such size difference among human cells commonly used for imaging. I still think that our images were not collected at uniform magnification.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434747,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "12/06/2018 21:58:12",
      "content": "<p>The only domain knowledge I have is that mitochondria is the powerhouse of the cell.</p>\n\n<p>I guess I won't get far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 434757,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/06/2018 22:25:10",
          "content": "<blockquote>\n  <p>I guess I won't get far.</p>\n</blockquote>\n\n<p>Your present LB spot argues otherwise ;-) I would guess - and hope there is nobody offended by it - that plenty of our top 50 competitors don't have full domain knowledge. My own domain knowledge is of limited use until I actually decide to compete.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435203,
          "author_name": "michaeltam",
          "author_url": "",
          "post_date": "12/07/2018 16:59:25",
          "content": "<p>I see you too, are a man of culture</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434763,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "12/06/2018 22:45:11",
      "content": "<p>Ever since the beginning of this competition i felt that this challenge is actually more suitable for object detection. One of the problems of the multilabel approach is that the network must put \"extra effort\" into learning which is which. Putting it into eggs, suppised we have lots of photos of fried eggs of all kinds: duck, pigeon, ostrich and also has browns and scrabled eggs. Each photo shows several of these on the hot plate. A 5 yo child will have lots of trouble finding out which is which if all he had was a title like \"duck and pigeon egg\". However, i can understand the amount of work that involve markup for object detection.... </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435187,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "12/07/2018 16:34:07",
      "content": "<p>great @Tilii!</p>\n\n<p>thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435815,
      "author_name": "cavriends",
      "author_url": "",
      "post_date": "12/08/2018 20:39:05",
      "content": "<p>Quite insightful! Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437280,
      "author_name": "darylrevock",
      "author_url": "",
      "post_date": "12/11/2018 17:07:23",
      "content": "<p>a chicken egg doesn't just look like a single cell; it is one.  hence the similarity</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443028,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "12/20/2018 22:52:41",
      "content": "<p>Just wanted to say this is a brilliant post - thank you for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "431745": "If I were to pick a real-life example of what cells look like the most, I’d say eggs. Aside from relative uniformity in shapes and sizes among eggs that doesn’t necessarily translate to cells, other features are clearly correlated. An oval middle part that is clearly separated from the rest and holds the goodies? Check. A proteinaceous and semi-liquid substance that surrounds it? Check. A barrier that separates the contents from the outside? Check. Even if you don’t agree with all aspects of my analogy, I’ll give you one that should bolster my case.\n\n![enter image description here][1]\n\nThat’s a sunny side up fried egg. Or, if you want to be more technical, an egg that was taken out of its shell and laid down on the heated flat surface. Cells we are looking at are exactly like that: they have been laid down on a flat microscopy slide, and they have been fixed (which is like frying but without heat). The egg yolk looks mostly oval on the surface and usually holds its shape well unless agitated, just like nucleus. The egg white has a general shape but much more variability than yolk, just like the cytoplasm surrounding the nucleus. Any classification of cells starts with these two components, and it isn’t by accident that nucleus/nucleoplasm (class 0) and cytosol/cytoplasm (class 25) are most abundant. So, how hard is it to tell these two apart in cell images?\n\nTo answer the question, I did a simple transfer learning exercise with VGG16 as a starting model. Took the top layers off and froze the rest, and pre-trained with my own top layers trained for two categories; that was followed by training with all layers unfrozen. There were ~6000 images in the train set and ~1500 images were used for validation.\n\nHere is an image from the validation group, with cytoplasmic protein signal:\n\n![enter image description here][2]\n\nThe network correctly thinks that this signal is in cytosol (0.99999 probability). Here is what the network actually learned:\n\n![enter image description here][3]\n\nThe left side shows top 5 feature areas that lead to conclusion that the signal is not in nucleus. The right side shows top 10 feature areas leading to cytoplasm conclusion (green) or those leading to not-nucleus conclusion (pink). It seems like the pattern is learned really well, as long as we are dealing with a nice field of cells.\n\nBut here is the thing: this classifier is only ~97% accurate. Take a 5-year old and show her 100 fried eggs sunny side up, and I’ll bet you that she will tell yolks and whites apart with 100% accuracy. Why isn’t our classifier able to better tell apart the two most prominent cellular features?\n\nI took activations from the penultimate network layer and embedded validation data points using t-SNE. Cytosol images are red, nucleus is blue. This particular representation is meant to prove [__my earlier point__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72795) that faces are everywhere.\n\n![enter image description here][4]\n\nFor an actual embedding, see the attached gif file. It should be clear that there are two nicely defined clusters that have only few stray points – this is expected based on 97% accuracy. It is those strays that interest me, so I laid down the points into a grid that’s easier to look at:\n\n![enter image description here][5]\n\nNow we put together cell images to match the above grid:\n\n![enter image description here][6]\n\nThis image is too small to see the details, so I will attach a higher-resolution image if you want to take a look. What I focus on is those stray data points, or those that are on the boundaries between the two colors. I’ll let you make your own conclusions after you inspect the larger image, but I will make two points here. First, at least some of our misclassified images are due to technical problems. Here is an image that you can find by starting from the upper left corner, and going 5 images in X direction (to the right) and 24 in Y direction (down). This image is on the boundary, has a nuclear signal, but is classified as cytoplasmic (0.83 probability).\n\n![enter image description here][7]\n\nThere is a color shift in this image and the channels do not line up. It has been discussed before [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70608) and I thought it would not be very prevalent. However, I have since found at least 5 shifted images, and that’s without looking very hard for them. Anyway, when this happens, it totally confuses the network.\n\n![enter image description here][8]\n\nI still think this shouldn’t happen often enough to have a major effect.\n\nMy second point has to do with natural heterogeneity of these images. I will blow up two sections from the grid above: one from the upper right corner (nuclear images) and from lower right (cytoplasmic images). These are predicted with high confidence and well separated from each other, so one would expect a clear and uniform pattern in them.\n\nNuclear images first:\n\n![enter image description here][9]\n\nThese images have several things in common: medium- to small-sized cells, lots of blue and red and much less green color. That makes sense, because nuclei are smaller than cytoplasm in terms of surface, so there won’t be much outright green signal. Here is where the natural differences in protein expression kick in: if protein of interest is hugely overexpressed, its green signal would simply overpower the blue coming from DNA, and we’d see green nuclei. When its signal is comparable to DAPI (blue), we’d see cyan/turquoise color from equal blue/green mixing. Finally, poorly expressed proteins would give very little green signal, so our nuclei would be blue. Note that “poorly” is meant in relative terms here, as even well-expressed proteins can’t always match the DAPI signal. In reality, we see all these scenarios in the image above, with straight-up blue being most prominent. It is actually very impressive that a network picks up these disparate levels of blue and green mixing mean the same thing.\n\nNow cytoplasmic images:\n\n![enter image description here][10]\n\nThese are mostly larger cells (by that I mean larger images of cells), with lots of blue and green and little pure red color. Different levels of green and red mixing give us yellow and orange shades in cytoplasm. In relative terms, the closer we have the cytoplasm to pure red color, the less our protein of interest is expressed relative to microtubules. Conversely, a cytoplasmic signal that is closer to green means that our protein is over-expressed relative to microtubules.\n\nI’ll pre-emptively answer the question: there is no defined set of rules governing protein expression, even when they are expressed in the same system (same cells, same promoters). Needs of the cell for proteins vary greatly as well, and their numbers vary over few orders of magnitude. Beyond that, small proteins express better than large ones, mostly because they are easier to make. I dealt with proteins that give huge signal 8 hours after transfection, while others take 24 hours. Add to that morphological differences between cell types, different growth rates, different expression efficiencies, different antibodies used for detection, different technicians performing experiments, different microscopes, different magnifications, and we are classifying something that is not as simple as eggs that are sunny side up:\n\n![enter image description here][11]\n\nNow add 26 classes on top of nucleus and cytoplasm, and suddenly our challenge becomes:\n\n![enter image description here][12]\n\nTurns out telling yolks and whites apart is tricky business. Maybe I should stop whinnying about 97% accuracy.\n\n  [1]: https://i.postimg.cc/NM6xbStS/sunny-up.jpg\n  [2]: https://i.postimg.cc/qqh6jD9y/test3.png\n  [3]: https://i.postimg.cc/ZndxsxgB/verdict3.png\n  [4]: https://i.postimg.cc/7LqcYmb3/cells-smiley-02.gif\n  [5]: https://i.postimg.cc/D0DhcnGd/cells-square-grid-01.png\n  [6]: https://i.postimg.cc/hvFWJpRp/scaled-custom-binary-rgb-tsne.png\n  [7]: https://i.postimg.cc/NfdtD86f/6f062840-bba9-11e8-b2ba-ac1f6b6435d0-rgb.png\n  [8]: https://i.postimg.cc/L5bnrccN/verdict5.png\n  [9]: https://i.postimg.cc/NGz8hS73/cutout-custom-binary-rgb-tsne.png\n  [10]: https://i.postimg.cc/yxMSJqLW/cutout-custom-binary-rgb-tsne-2.png\n  [11]: https://i.postimg.cc/8cysW7Ny/aid1693528-v4-728px-Make-Deep-Fried-Eggs-Step-7-Version-2.jpg\n  [12]: https://i.postimg.cc/NMcywhdx/Fried-Egg-Cover.png",
    "431746": "Full-size grid is attached here.",
    "431747": "t-SNE animation is attached here.",
    "432367": "huh?",
    "432379": "longhorns2102 Anything in particular that's confusing you?",
    "434136": "Nice post!\nFirst question - what is the exact definition of \"expression\"? Does it just mean general contrast, or contrast between green and red, or something else?\nAnd you mention \"These are mostly larger cells (by that I mean larger images of cells)\". So can absolute size be an indicator of type, or is it just an artifact of something else (microscope zoom level or something)? I guess I'm asking can I augment with random zoom/cropping to avoid over-fitting - or will that make it lose some crucial information?",
    "434151": "How are you creating the images of feature areas and the tSNE representations? Is there a guide or something that describes how to extract this?",
    "434157": "hshmhashemi Expression indicates the amount of protein made in cells. The fluorescent signal is proportional to the amount of target proteins. I was making a point that different proteins are not made in cell in equal amounts even under otherwise identical conditions, so that always creates problems with signal normalization.\n\nWe have been told that uniform magnification was used for most images. My eyes tell me otherwise. Cells do differ in sizes, but it should not be as much as smallest cells in upper right corner vs. largest cells in lower right corner. There is easily 4-5x size difference, and in my experience that's not the case. Anyway, I suggest asymmetric zoom [0.9, 1.3] (maybe up to 1.5).",
    "434162": "ldm314 [__Lime__](https://github.com/marcotcr/lime) was used to explain the images. t-SNE grid creation is described [__here__](https://github.com/prabodhhere/tsne-grid) and [__RasterFairy__](https://github.com/bmcfee/RasterFairy) was used for animations.",
    "434164": "Thanks!",
    "434208": "Nice post！Thanks！",
    "434328": "Excellent post, Tilli, thanks! For hshmhashemi, tilii is correct, proteins are made by ribosomes translating mRNA so increased and decreased expression refers to the observable quantity of some protein in this case. What causes changes in expression levels can range from environmental factors to the previous activity states on the cell.\n\nSome of the cells to look significantly different in size, but for any given image I think it is highly unlikely the magnification would be different in different parts of the same image. But this is a fairly big dataset and errors and artifacts always happen. Some of the label imbalance is due to the transient nature of several of the labels, anything to do with cell division in particular. These are cultured human cells and they do divide and grow, but the chances of catching one mid mitosis is just much more unlikely than not.\n\nI don't actually believe knowing the cell types would help very much in this challenge since we dont know what the antibodies are actually targeting, they just want us to identify zones and structures in the cell. Actin is the only actual protein in the entire list of labels.\n\nAlso it is certainly true that the level of fluorescence is proportional to expression, its also a function of the specificity and affinity of the antibody. All the cells in a given image will have been incubated in the same set of antibodies, but from that alone it is very tough to infer much that is useful. For example, each label is under no obligation to imply one protein, in fact any label could actually be more than one individual protein, ie, cytosol is most of the cell an there are a lot of labels with only cytosol, but that does not mean only one protein was labeled, it could be the case but we have no way be sure. Even images with multiple labels could actually be showing one or more unique proteins, its quite possible for one protein to be highly expressed in multiple compartments of the cell at once. I wish I could say two or more labels means at least two proteins, but that conclusion is not warranted. \n\nWhat I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.",
    "434715": "&gt; What I'd like to be able to do, and have been trying despite my n00bness, is to extract each individual cell from the images and use each one, or at least some subset that meets some quality standard, as an example of the label. All cells in an image supposedly did get the same antibodies, so any cell from a given image should count as an example of the multilabel. The positions of the cells in the image is random, but the distribution of protein in the cells is not. I'm unsure how to tell my models how to ignore the variance from the cell positions.\n\nIt may be a good idea to extract individual cells, but I would not try to define a subset that meets any standard. First, what would that standard be? Second, this is meant to be a general classifier, and what happens when you have to classify an image where no cell meets that standard?\n\nDon't think that you need to tell your models to ignore positional variance. There is useful information in that feature. If anything, I think one should be using image shifts and rotations to increase positional variance and improve the generalization.",
    "434726": "I made a case before that differences in cell sizes in images are unlikely to reflect natural variability. Here is a boundary group of cells with nuclear and cytoplasmic signals (around 5 o'clock in the square grid image above):\n\n![enter image description here][1]\n\nCells in top two rows and the right side of the third row have nuclear signals, while the rest are cytoplasmic. I estimated earlier that size differences between cells were on the order of 4-5x, but this image alone shows that the size range is closer to 10x. To the best of my knowledge, there is no such size difference among human cells commonly used for imaging. I still think that our images were not collected at uniform magnification.\n\n  [1]: https://i.ibb.co/4YYFB0S/cell-sizes.png",
    "434747": "The only domain knowledge I have is that mitochondria is the powerhouse of the cell.\n\nI guess I won't get far.",
    "434757": "&gt; I guess I won't get far.\n\nYour present LB spot argues otherwise ;-) I would guess - and hope there is nobody offended by it - that plenty of our top 50 competitors don't have full domain knowledge. My own domain knowledge is of limited use until I actually decide to compete.",
    "434763": "Ever since the beginning of this competition i felt that this challenge is actually more suitable for object detection. One of the problems of the multilabel approach is that the network must put \"extra effort\" into learning which is which. Putting it into eggs, suppised we have lots of photos of fried eggs of all kinds: duck, pigeon, ostrich and also has browns and scrabled eggs. Each photo shows several of these on the hot plate. A 5 yo child will have lots of trouble finding out which is which if all he had was a title like \"duck and pigeon egg\". However, i can understand the amount of work that involve markup for object detection....",
    "434850": "Yes, those are good points. The images we are given contain multiple cells, so training on only single cell examples of the label might completely fail to generalize to the arbitrary cell numbers/image they want us to predict on. I don't have a great intuition for this sort of thing yet, though - I'm inclined to believe you.",
    "435187": "great @Tilii!\n\nthanks for sharing!",
    "435203": "I see you too, are a man of culture",
    "435815": "Quite insightful! Thanks for sharing :)",
    "437280": "a chicken egg doesn't just look like a single cell; it is one.  hence the similarity",
    "443028": "Just wanted to say this is a brilliant post - thank you for sharing."
  },
  "source": "meta"
}