{
  "id": 215981,
  "title": "Checking my understanding",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/215981",
  "author_name": "Mensch",
  "post_date": "2021-02-01T05:47:53.775000",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The objective is to locate the position of the protein - which appears in green - in each cell.</p>\n<p>Each cell has all of the 18 regions associated with it (associated with labels) or not stained (label 18).</p>\n<p>The submission corresponding to a test image will require: determining the cell mask (for each cell marked with protein), classifying the region where the protein is situated into one of the 18 labels. So, if there are 3 proteins (green) marked in a cell, then for the same mask there will be 3 labels. This is done for each cell in the image, which consists of protein(s).</p>",
  "messages": [
    {
      "id": 1180136,
      "postDate": "2021-02-01T05:47:53.777Z",
      "content": "<p>The objective is to locate the position of the protein - which appears in green - in each cell.</p>\n<p>Each cell has all of the 18 regions associated with it (associated with labels) or not stained (label 18).</p>\n<p>The submission corresponding to a test image will require: determining the cell mask (for each cell marked with protein), classifying the region where the protein is situated into one of the 18 labels. So, if there are 3 proteins (green) marked in a cell, then for the same mask there will be 3 labels. This is done for each cell in the image, which consists of protein(s).</p>",
      "rawMarkdown": "The objective is to locate the position of the protein - which appears in green - in each cell.\n\nEach cell has all of the 18 regions associated with it (associated with labels) or not stained (label 18).\n\nThe submission corresponding to a test image will require: determining the cell mask (for each cell marked with protein), classifying the region where the protein is situated into one of the 18 labels. So, if there are 3 proteins (green) marked in a cell, then for the same mask there will be 3 labels. This is done for each cell in the image, which consists of protein(s).",
      "votes": 14
    },
    {
      "id": 1180179,
      "postDate": "2021-02-01T06:24:34.190Z",
      "content": "<p>The objective is to detect if each cell has which label(s) or classify each cell to class(es). <br>\nYou don't need to find and submit the position of the protein in each cell, which is a considerably harder task. What you do for each image is to segment cell mask and classify each cell into one or more of 19 labels (18 organelle labels + Negative).<br>\nPlease check this out to understand the patterns <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns</a></p>\n<p>Please also check out this discussion <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141</a></p>",
      "rawMarkdown": "The objective is to detect if each cell has which label(s) or classify each cell to class(es). \nYou don't need to find and submit the position of the protein in each cell, which is a considerably harder task. What you do for each image is to segment cell mask and classify each cell into one or more of 19 labels (18 organelle labels + Negative).\nPlease check this out to understand the patterns https://www.kaggle.com/lnhtrang/single-cell-patterns\n\nPlease also check out this discussion https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141",
      "votes": 2,
      "replies": [
        {
          "id": 1180419,
          "postDate": "2021-02-01T08:39:37.407Z",
          "content": "<p>Thank you for the clarification.  </p>\n<p>A. But it actually raises a few more questions for me. My understanding was that the green image already provides the coordinate location of the protein, ie. the position of the protein within the cell. The challenge is to identify if the protein is in the nucleoli or plasma membrane or any of these 18 parts (organelles?) of the cell. After all, every cell has these 18 organelles. So, the challenge is to identify as to in which organelle the protein lies, which defines the nature of the cell, making it different from its ancestor or sister.</p>\n<p>Kindly correct me where I am wrong.</p>\n<p>B. Furthermore, I was wondering what is a cell line (in <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns)\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns)</a>. Found this explanation:<br>\n\"Cell line is a general term that applies to a defined population of cells that can be maintained in culture for an extended period of time, retaining the stability of certain phenotypes and functions. Cell lines are usually clonal, meaning that the entire population originated from a single common ancestor cell.\"</p>\n<p>So, do you capture multiple images from each cell line? Does it mean that the same cells are repeated across training images? Also, would it not improve the classification accuracy, if the characteristics of a cell line are also incorporated in a model?</p>",
          "rawMarkdown": "Thank you for the clarification.  \n\nA. But it actually raises a few more questions for me. My understanding was that the green image already provides the coordinate location of the protein, ie. the position of the protein within the cell. The challenge is to identify if the protein is in the nucleoli or plasma membrane or any of these 18 parts (organelles?) of the cell. After all, every cell has these 18 organelles. So, the challenge is to identify as to in which organelle the protein lies, which defines the nature of the cell, making it different from its ancestor or sister.\n\nKindly correct me where I am wrong.\n\n\nB. Furthermore, I was wondering what is a cell line (in https://www.kaggle.com/lnhtrang/single-cell-patterns). Found this explanation:\n\"Cell line is a general term that applies to a defined population of cells that can be maintained in culture for an extended period of time, retaining the stability of certain phenotypes and functions. Cell lines are usually clonal, meaning that the entire population originated from a single common ancestor cell.\"\n\nSo, do you capture multiple images from each cell line? Does it mean that the same cells are repeated across training images? Also, would it not improve the classification accuracy, if the characteristics of a cell line are also incorporated in a model?"
        },
        {
          "id": 1180789,
          "postDate": "2021-02-01T13:24:04.940Z",
          "content": "<p>For A:  <br>\nYes, we get the absolute X and Y coordinates for the proteins from the green channel but that information is useless without the context provided by the other channels. The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. For example, if there is overlap between the red channel (Microtubules) and the green (protein) we can understand that the protein localizes to the microtubules.  <br>\nDo note that proteins can localize to multiple organelles.</p>\n<p>I'm not sure I understand what you are referring to with the ancestor/sister cells part of the question? Each cell will indeed contain all the organelles and proteins can localize to any of these within each cell. This localization may be different between cells in the same image or the same between them, depending on the protein and it's function.</p>\n<p>For B:  <br>\nA cell line is, like your explanation says, a population of cells that can be grown for a long time. We grow cells in our lab and take some of them for each imaging experiment. This means that we will have multiple images from the same cell line. It does not mean that we have the same cells between images, as different cells will be sampled for experiments.</p>\n<p>As for if characteristics of the cell line would improve the classification accuracy, that is certainly an interesting theory. In one of our previous papers (<a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">https://www.nature.com/articles/nbt.4225</a>), we did show that including different cell lines in our training data helped accuracy across the board (Figure 5C of that paper). This could mean cell line information may be helpful for a neural network, but we have not tested, much less proved, that theory. If you think it will be helpful, feel free to try it.</p>",
          "rawMarkdown": "For A:  \nYes, we get the absolute X and Y coordinates for the proteins from the green channel but that information is useless without the context provided by the other channels. The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. For example, if there is overlap between the red channel (Microtubules) and the green (protein) we can understand that the protein localizes to the microtubules.  \nDo note that proteins can localize to multiple organelles.\n\nI'm not sure I understand what you are referring to with the ancestor/sister cells part of the question? Each cell will indeed contain all the organelles and proteins can localize to any of these within each cell. This localization may be different between cells in the same image or the same between them, depending on the protein and it's function.\n\nFor B:  \nA cell line is, like your explanation says, a population of cells that can be grown for a long time. We grow cells in our lab and take some of them for each imaging experiment. This means that we will have multiple images from the same cell line. It does not mean that we have the same cells between images, as different cells will be sampled for experiments.\n\nAs for if characteristics of the cell line would improve the classification accuracy, that is certainly an interesting theory. In one of our previous papers ([https://www.nature.com/articles/nbt.4225](https://www.nature.com/articles/nbt.4225)), we did show that including different cell lines in our training data helped accuracy across the board (Figure 5C of that paper). This could mean cell line information may be helpful for a neural network, but we have not tested, much less proved, that theory. If you think it will be helpful, feel free to try it.",
          "votes": 3
        },
        {
          "id": 1180955,
          "postDate": "2021-02-01T14:55:25.020Z",
          "content": "<p>I get it now. Thanks for clarifying, patiently. :)</p>",
          "rawMarkdown": "I get it now. Thanks for clarifying, patiently. :)",
          "votes": 1
        },
        {
          "id": 1183793,
          "postDate": "2021-02-03T08:14:31.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> Very nicely explained👍. I had a similar query, and the following sentence was an eye-opener:-</p>\n<blockquote>\n  <p>The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. </p>\n</blockquote>\n<p>This clearly summarizes the aim of the project in a single line.</p>",
          "rawMarkdown": "@cwinsnes Very nicely explained👍. I had a similar query, and the following sentence was an eye-opener:-\n\n> The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. \n\nThis clearly summarizes the aim of the project in a single line.",
          "votes": 1
        },
        {
          "id": 1207271,
          "postDate": "2021-02-17T19:18:21.603Z",
          "content": "<p>Thus to summarize the competition task in my understanding:</p>\n<ul>\n<li>We know where the protein is situated from the green channel. We can even get the absolute X and Y coordinate for the same.</li>\n<li>The task is to predict which organelle(s) does the protein belongs to in the cell. </li>\n<li>We do have image-level labels about which organelle the protein is localized but that may not be the case for every cell in the image.</li>\n<li>Thus using weak supervision from image-level labels(or other priors) we need to first segment the cells and then predict the name of the organelle. </li>\n</ul>\n<p>Correct me if I am wrong <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a>.</p>",
          "rawMarkdown": "Thus to summarize the competition task in my understanding:\n\n- We know where the protein is situated from the green channel. We can even get the absolute X and Y coordinate for the same.\n- The task is to predict which organelle(s) does the protein belongs to in the cell. \n- We do have image-level labels about which organelle the protein is localized but that may not be the case for every cell in the image.\n- Thus using weak supervision from image-level labels(or other priors) we need to first segment the cells and then predict the name of the organelle. \n\nCorrect me if I am wrong @cwinsnes.",
          "votes": 1
        },
        {
          "id": 1208532,
          "postDate": "2021-02-18T10:19:18.067Z",
          "content": "<p>Yes, that seems about correct to me! </p>\n<p>A couple of notes:</p>\n<ul>\n<li><p>We don't usually use the words \"belongs to\" but instead use the words \"localizes to\".</p></li>\n<li><p>We don't predict the name of the organelle (they are known), but we predict to which organelle each protein localizes to!</p></li>\n<li><p>The absolute X, Y coordinates that we can get for a protein are of course in relation to the image, not to the cells, by simply checking \"where is the image green\". This in itself is not very useful, which is why we look at the green channel in relation to the other channels.</p></li>\n</ul>\n<p>I'm sure you already understood this, but I want to be sure that I don't accidentally lead you down the wrong path if I misunderstood your wording.</p>",
          "rawMarkdown": "Yes, that seems about correct to me! \n\nA couple of notes:\n\n* We don't usually use the words \"belongs to\" but instead use the words \"localizes to\".\n\n* We don't predict the name of the organelle (they are known), but we predict to which organelle each protein localizes to!\n\n* The absolute X, Y coordinates that we can get for a protein are of course in relation to the image, not to the cells, by simply checking \"where is the image green\". This in itself is not very useful, which is why we look at the green channel in relation to the other channels.\n\nI'm sure you already understood this, but I want to be sure that I don't accidentally lead you down the wrong path if I misunderstood your wording.",
          "votes": 3
        },
        {
          "id": 1208871,
          "postDate": "2021-02-18T14:25:10.593Z",
          "content": "<p>Thank you for the clarification. </p>",
          "rawMarkdown": "Thank you for the clarification. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1180179,
      "author_name": "Trang Le",
      "author_url": "",
      "post_date": "2021-02-01T06:24:34.190000",
      "content": "<p>The objective is to detect if each cell has which label(s) or classify each cell to class(es). <br>\nYou don't need to find and submit the position of the protein in each cell, which is a considerably harder task. What you do for each image is to segment cell mask and classify each cell into one or more of 19 labels (18 organelle labels + Negative).<br>\nPlease check this out to understand the patterns <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns</a></p>\n<p>Please also check out this discussion <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1180419,
          "author_name": "Mensch",
          "author_url": "",
          "post_date": "2021-02-01T08:39:37.407000",
          "content": "<p>Thank you for the clarification.  </p>\n<p>A. But it actually raises a few more questions for me. My understanding was that the green image already provides the coordinate location of the protein, ie. the position of the protein within the cell. The challenge is to identify if the protein is in the nucleoli or plasma membrane or any of these 18 parts (organelles?) of the cell. After all, every cell has these 18 organelles. So, the challenge is to identify as to in which organelle the protein lies, which defines the nature of the cell, making it different from its ancestor or sister.</p>\n<p>Kindly correct me where I am wrong.</p>\n<p>B. Furthermore, I was wondering what is a cell line (in <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns)\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns)</a>. Found this explanation:<br>\n\"Cell line is a general term that applies to a defined population of cells that can be maintained in culture for an extended period of time, retaining the stability of certain phenotypes and functions. Cell lines are usually clonal, meaning that the entire population originated from a single common ancestor cell.\"</p>\n<p>So, do you capture multiple images from each cell line? Does it mean that the same cells are repeated across training images? Also, would it not improve the classification accuracy, if the characteristics of a cell line are also incorporated in a model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1180789,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-02-01T13:24:04.940000",
          "content": "<p>For A:  <br>\nYes, we get the absolute X and Y coordinates for the proteins from the green channel but that information is useless without the context provided by the other channels. The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. For example, if there is overlap between the red channel (Microtubules) and the green (protein) we can understand that the protein localizes to the microtubules.  <br>\nDo note that proteins can localize to multiple organelles.</p>\n<p>I'm not sure I understand what you are referring to with the ancestor/sister cells part of the question? Each cell will indeed contain all the organelles and proteins can localize to any of these within each cell. This localization may be different between cells in the same image or the same between them, depending on the protein and it's function.</p>\n<p>For B:  <br>\nA cell line is, like your explanation says, a population of cells that can be grown for a long time. We grow cells in our lab and take some of them for each imaging experiment. This means that we will have multiple images from the same cell line. It does not mean that we have the same cells between images, as different cells will be sampled for experiments.</p>\n<p>As for if characteristics of the cell line would improve the classification accuracy, that is certainly an interesting theory. In one of our previous papers (<a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">https://www.nature.com/articles/nbt.4225</a>), we did show that including different cell lines in our training data helped accuracy across the board (Figure 5C of that paper). This could mean cell line information may be helpful for a neural network, but we have not tested, much less proved, that theory. If you think it will be helpful, feel free to try it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1180955,
          "author_name": "Mensch",
          "author_url": "",
          "post_date": "2021-02-01T14:55:25.020000",
          "content": "<p>I get it now. Thanks for clarifying, patiently. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1183793,
          "author_name": "Manish Kumar Mohanty",
          "author_url": "",
          "post_date": "2021-02-03T08:14:31.243000",
          "content": "<p><a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> Very nicely explained👍. I had a similar query, and the following sentence was an eye-opener:-</p>\n<blockquote>\n  <p>The challenge lies in understanding which patterns can be seen in each cell, i.e. understanding which organelle(s) (cell part) the protein localizes to in the specific cell. </p>\n</blockquote>\n<p>This clearly summarizes the aim of the project in a single line.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1207271,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-17T19:18:21.603000",
          "content": "<p>Thus to summarize the competition task in my understanding:</p>\n<ul>\n<li>We know where the protein is situated from the green channel. We can even get the absolute X and Y coordinate for the same.</li>\n<li>The task is to predict which organelle(s) does the protein belongs to in the cell. </li>\n<li>We do have image-level labels about which organelle the protein is localized but that may not be the case for every cell in the image.</li>\n<li>Thus using weak supervision from image-level labels(or other priors) we need to first segment the cells and then predict the name of the organelle. </li>\n</ul>\n<p>Correct me if I am wrong <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1208532,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-02-18T10:19:18.067000",
          "content": "<p>Yes, that seems about correct to me! </p>\n<p>A couple of notes:</p>\n<ul>\n<li><p>We don't usually use the words \"belongs to\" but instead use the words \"localizes to\".</p></li>\n<li><p>We don't predict the name of the organelle (they are known), but we predict to which organelle each protein localizes to!</p></li>\n<li><p>The absolute X, Y coordinates that we can get for a protein are of course in relation to the image, not to the cells, by simply checking \"where is the image green\". This in itself is not very useful, which is why we look at the green channel in relation to the other channels.</p></li>\n</ul>\n<p>I'm sure you already understood this, but I want to be sure that I don't accidentally lead you down the wrong path if I misunderstood your wording.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1208871,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-18T14:25:10.593000",
          "content": "<p>Thank you for the clarification. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1180136": "The objective is to locate the position of the protein - which appears in green - in each cell.\n\nEach cell has all of the 18 regions associated with it (associated with labels) or not stained (label 18).\n\nThe submission corresponding to a test image will require: determining the cell mask (for each cell marked with protein), classifying the region where the protein is situated into one of the 18 labels. So, if there are 3 proteins (green) marked in a cell, then for the same mask there will be 3 labels. This is done for each cell in the image, which consists of protein(s).",
    "1180179": "The objective is to detect if each cell has which label(s) or classify each cell to class(es). \nYou don't need to find and submit the position of the protein in each cell, which is a considerably harder task. What you do for each image is to segment cell mask and classify each cell into one or more of 19 labels (18 organelle labels + Negative).\nPlease check this out to understand the patterns https://www.kaggle.com/lnhtrang/single-cell-patterns\n\nPlease also check out this discussion https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141"
  }
}