{
  "id": 227893,
  "title": "What do we predict? ",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/227893",
  "author_name": "MatveySafroshkin",
  "post_date": "2021-03-22T16:49:27.479000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Good day everyone! <br>\nI spent quite a lot of time reading forum and looking at the images. <br>\nI don't really understand if we are predicting a label for cell or for cell components.</p>\n<p>Following this topic : <a href=\"https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda</a>  and this one: <a href=\"https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic\" target=\"_blank\">https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic</a> </p>\n<p>I have several concerns. <br>\nCould you please correct me : </p>\n<ul>\n<li>NUCLEUS <br>\n1) If we have a blue nucleus, we have a nucleoplasm, <br>\n2)if we have a nucleoplasm, there should be a membrane. If there is no membrane, blue nucleus can not stay in the center. (we can somehow find a mask of membrane using dilatation/erosion on nucleus)<br>\n3) Nucleoli, Nucleoli fibrillar center and Nuclear speckles, Nuclear bodies should be found in nucleuses. </li>\n</ul>\n<p>*OUTERCELL: <br>\nGoldi apparatus :  makes a cell look \"furry\"<br>\nMicrotubes: looks like a red web on the cell  <br>\nMitochondria : a red hairy cell <br>\ncytosol: looks like a green cloud with  dark center </p>\n<p>What does single label mitochondria mean?  Is it a correct labeling? <br>\nHow can a cell be without nucleus?  <br>\nThank you in advance.<br>\nBest regards, <br>\nMat</p>",
  "messages": [
    {
      "id": 1249846,
      "postDate": "2021-03-23T15:07:31.337Z",
      "content": "<p>Mat</p>\n<p>All cells have many of the classes we are attempting to predict.  We are not attempting a simple body part identification.  The red, blue and yellow images can be combined into a single image that provides the structure of each cell.  </p>\n<p>Step 1 - the images provided contain a varying number of cells.  We need to segment that large image into the correct number of single cells.  The private test set was segmented by method(s) shared in the overview and used in many of the public kernels.  After automated segmentation humans did a further correction of the segments.  Since we are provided no ground truth for segmentation this is an important step but not the one that should be our major focus.  </p>\n<p>Step 2 - the cells have been stained with an antibody.  Our model is attempting to predict which parts of each single cell has the antibody attached to.  Antibodies will attach to one or more different parts of a cell.  The methods used for this staining result in the \"green\" image being the indicator of the intensity of the attachment.  Each row in the training data represents a different cell line and a different antibody.  The label for that row tells us which part or parts of the cell the antibody has attached to. </p>\n<p>The task is complicated because the label provided for the large image MAY NOT be correct for every single cell contained within that large image.  </p>\n<p>So in a sense we are throwing an unknown antibody on a cell.  It's sticking to only specific parts of the cell.  Our job is to identify (for each row) what part of the cell has been stained (is green).  The <strong>green image is the key</strong> - a model could be partly successful using only the green image.   But including the red, blue and yellow allows a model to see the green in relationship to the cell structure.</p>\n<p>I have been doing lots of these competitions over that past several years - IMO this is one of the most complex ones I have attempted.  We need to match the segmentation performed by the hosts but with out the assistance of any ground truth, we need to label each single cell but the training ground truth does not apply to every single cell within an image, the metric is both difficult to understand and hard to incorporate into model training, the segmentation is a slow process and kaggle time constraints end up causing many of your early attempts at prediction to have out of time error.</p>",
      "rawMarkdown": "Mat\n\nAll cells have many of the classes we are attempting to predict.  We are not attempting a simple body part identification.  The red, blue and yellow images can be combined into a single image that provides the structure of each cell.  \n\nStep 1 - the images provided contain a varying number of cells.  We need to segment that large image into the correct number of single cells.  The private test set was segmented by method(s) shared in the overview and used in many of the public kernels.  After automated segmentation humans did a further correction of the segments.  Since we are provided no ground truth for segmentation this is an important step but not the one that should be our major focus.  \n\nStep 2 - the cells have been stained with an antibody.  Our model is attempting to predict which parts of each single cell has the antibody attached to.  Antibodies will attach to one or more different parts of a cell.  The methods used for this staining result in the \"green\" image being the indicator of the intensity of the attachment.  Each row in the training data represents a different cell line and a different antibody.  The label for that row tells us which part or parts of the cell the antibody has attached to. \n\nThe task is complicated because the label provided for the large image MAY NOT be correct for every single cell contained within that large image.  \n\nSo in a sense we are throwing an unknown antibody on a cell.  It's sticking to only specific parts of the cell.  Our job is to identify (for each row) what part of the cell has been stained (is green).  The **green image is the key** - a model could be partly successful using only the green image.   But including the red, blue and yellow allows a model to see the green in relationship to the cell structure.\n\nI have been doing lots of these competitions over that past several years - IMO this is one of the most complex ones I have attempted.  We need to match the segmentation performed by the hosts but with out the assistance of any ground truth, we need to label each single cell but the training ground truth does not apply to every single cell within an image, the metric is both difficult to understand and hard to incorporate into model training, the segmentation is a slow process and kaggle time constraints end up causing many of your early attempts at prediction to have out of time error.\n\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1250182,
          "postDate": "2021-03-23T20:22:58.680Z",
          "content": "<p>Hello, are you sure that hosts will use shared segmentator for the test data? I think that I read somewhere that test data was manually segmentated and then each cell was manually labeled.<br>\nIf you are right it means that we don't need to try to improve segmentation at all.</p>",
          "rawMarkdown": "Hello, are you sure that hosts will use shared segmentator for the test data? I think that I read somewhere that test data was manually segmentated and then each cell was manually labeled.\nIf you are right it means that we don't need to try to improve segmentation at all.",
          "votes": 1
        },
        {
          "id": 1250456,
          "postDate": "2021-03-24T04:03:45.183Z",
          "content": "<p>I read that they used the hpa segmentation and than annotated the results for the final.  I believe a host indicated it was 90% good if we just used the hpa segmentation.</p>\n<p>We do need to improve segmentation - it can be too slow if not used in well written script.  When I can get a LB score higher than .50 than I plan to try for some improvement in both speed and quality. </p>",
          "rawMarkdown": "I read that they used the hpa segmentation and than annotated the results for the final.  I believe a host indicated it was 90% good if we just used the hpa segmentation.\n\nWe do need to improve segmentation - it can be too slow if not used in well written script.  When I can get a LB score higher than .50 than I plan to try for some improvement in both speed and quality. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1248526,
      "postDate": "2021-03-22T16:49:27.480Z",
      "content": "<p>Good day everyone! <br>\nI spent quite a lot of time reading forum and looking at the images. <br>\nI don't really understand if we are predicting a label for cell or for cell components.</p>\n<p>Following this topic : <a href=\"https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda</a>  and this one: <a href=\"https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic\" target=\"_blank\">https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic</a> </p>\n<p>I have several concerns. <br>\nCould you please correct me : </p>\n<ul>\n<li>NUCLEUS <br>\n1) If we have a blue nucleus, we have a nucleoplasm, <br>\n2)if we have a nucleoplasm, there should be a membrane. If there is no membrane, blue nucleus can not stay in the center. (we can somehow find a mask of membrane using dilatation/erosion on nucleus)<br>\n3) Nucleoli, Nucleoli fibrillar center and Nuclear speckles, Nuclear bodies should be found in nucleuses. </li>\n</ul>\n<p>*OUTERCELL: <br>\nGoldi apparatus :  makes a cell look \"furry\"<br>\nMicrotubes: looks like a red web on the cell  <br>\nMitochondria : a red hairy cell <br>\ncytosol: looks like a green cloud with  dark center </p>\n<p>What does single label mitochondria mean?  Is it a correct labeling? <br>\nHow can a cell be without nucleus?  <br>\nThank you in advance.<br>\nBest regards, <br>\nMat</p>",
      "rawMarkdown": "Good day everyone! \nI spent quite a lot of time reading forum and looking at the images. \nI don't really understand if we are predicting a label for cell or for cell components.\n \nFollowing this topic : https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda  and this one: https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic \n\nI have several concerns. \nCould you please correct me : \n\n* NUCLEUS \n1) If we have a blue nucleus, we have a nucleoplasm, \n2)if we have a nucleoplasm, there should be a membrane. If there is no membrane, blue nucleus can not stay in the center. (we can somehow find a mask of membrane using dilatation/erosion on nucleus)\n3) Nucleoli, Nucleoli fibrillar center and Nuclear speckles, Nuclear bodies should be found in nucleuses. \n\n*OUTERCELL: \nGoldi apparatus :  makes a cell look \"furry\"\nMicrotubes: looks like a red web on the cell  \nMitochondria : a red hairy cell \ncytosol: looks like a green cloud with  dark center \n\n\nWhat does single label mitochondria mean?  Is it a correct labeling? \nHow can a cell be without nucleus?  \nThank you in advance.\nBest regards, \nMat\n\n",
      "votes": 2
    },
    {
      "id": 1250927,
      "postDate": "2021-03-24T11:13:09.120Z",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, Thank you a lot for your answer. It makes things a bit clearer. Why no one tried to make to label each cell individually?  Is it very hard to do? <br>\nAs you mentioned, your model it trying to predict, where antibodies are attached to, but why do we have mitochondria class?  Or does it mean, that antibodies are attached to Mitochondria ?</p>",
      "rawMarkdown": "@pcjimmmy, Thank you a lot for your answer. It makes things a bit clearer. Why no one tried to make to label each cell individually?  Is it very hard to do? \nAs you mentioned, your model it trying to predict, where antibodies are attached to, but why do we have mitochondria class?  Or does it mean, that antibodies are attached to Mitochondria ?",
      "replies": [
        {
          "id": 1251296,
          "postDate": "2021-03-24T16:49:06.877Z",
          "content": "<p>Glad I could make things a little clearer.   I tried to think of a good analogy or two to help make it more clear.  No great examples, but here's a couple of the weak ones.</p>\n<ol>\n<li><p>Certain fabrics and certain laundry detergents can make your clothes very bright under a UV lamp - If I had a bunch of photos of folks in a night club taken under UV I might challenge you to segment out each individual and than tell me which piece of clothing had been washed with laundry detergent X.  All the men would be wearing shirts - but we would only want to label those shirts that were washed with detergent X.  Now the task would be complicated by the issue that someone washed his shirt three weeks ago and spent some time outside in the rain.  Or complicated by the need to identify men vs women shirts.  Or not include florescent pants, etc.  </p></li>\n<li><p>I take a bunch of images of car engines with thermal camera.  All engines have parts that can be put in a bunch of different classes.  I might weakly label engines for defects and challenge you to identify the part class that is over/under heated and responsible for the poor performance.  I don't want you to label each spark plug, but I do want you to label plugs that are too hot or too cold.  </p></li>\n</ol>\n<p>When I look at the single cell images it does not appear to be that difficult to  label each cell if you know what your doing.  On my todo list is to manually label single cell images for those classes with a small number of samples in the class.  The Mitotic spindle is very distinct in appearance and in some of the larger images it appears that only a single cell in the large image has this attachment point.  My segmentation ended up generating 491K of single cell images.  (will need to read the rules close to insure that hand labeling of training data is acceptable).  That's a lot of labeling to do !</p>\n<p>But I think the hosts want to try and advance the science by challenging kagglers to develop code that can take weakly labeled images and build good models.  In this area of science and many others it does seem like most of the expense of getting a good model is the hand label process.</p>\n<p>The hosts did provide a notebook showing examples of all the classes.<br>\n<a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#14.-Mitochondria\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns#14.-Mitochondria</a></p>",
          "rawMarkdown": "Glad I could make things a little clearer.   I tried to think of a good analogy or two to help make it more clear.  No great examples, but here's a couple of the weak ones.\n\n1.  Certain fabrics and certain laundry detergents can make your clothes very bright under a UV lamp - If I had a bunch of photos of folks in a night club taken under UV I might challenge you to segment out each individual and than tell me which piece of clothing had been washed with laundry detergent X.  All the men would be wearing shirts - but we would only want to label those shirts that were washed with detergent X.  Now the task would be complicated by the issue that someone washed his shirt three weeks ago and spent some time outside in the rain.  Or complicated by the need to identify men vs women shirts.  Or not include florescent pants, etc.  \n\n2.  I take a bunch of images of car engines with thermal camera.  All engines have parts that can be put in a bunch of different classes.  I might weakly label engines for defects and challenge you to identify the part class that is over/under heated and responsible for the poor performance.  I don't want you to label each spark plug, but I do want you to label plugs that are too hot or too cold.  \n\n\n\nWhen I look at the single cell images it does not appear to be that difficult to  label each cell if you know what your doing.  On my todo list is to manually label single cell images for those classes with a small number of samples in the class.  The Mitotic spindle is very distinct in appearance and in some of the larger images it appears that only a single cell in the large image has this attachment point.  My segmentation ended up generating 491K of single cell images.  (will need to read the rules close to insure that hand labeling of training data is acceptable).  That's a lot of labeling to do !\n\nBut I think the hosts want to try and advance the science by challenging kagglers to develop code that can take weakly labeled images and build good models.  In this area of science and many others it does seem like most of the expense of getting a good model is the hand label process.\n\nThe hosts did provide a notebook showing examples of all the classes.\nhttps://www.kaggle.com/lnhtrang/single-cell-patterns#14.-Mitochondria\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1249846,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-03-23T15:07:31.337000",
      "content": "<p>Mat</p>\n<p>All cells have many of the classes we are attempting to predict.  We are not attempting a simple body part identification.  The red, blue and yellow images can be combined into a single image that provides the structure of each cell.  </p>\n<p>Step 1 - the images provided contain a varying number of cells.  We need to segment that large image into the correct number of single cells.  The private test set was segmented by method(s) shared in the overview and used in many of the public kernels.  After automated segmentation humans did a further correction of the segments.  Since we are provided no ground truth for segmentation this is an important step but not the one that should be our major focus.  </p>\n<p>Step 2 - the cells have been stained with an antibody.  Our model is attempting to predict which parts of each single cell has the antibody attached to.  Antibodies will attach to one or more different parts of a cell.  The methods used for this staining result in the \"green\" image being the indicator of the intensity of the attachment.  Each row in the training data represents a different cell line and a different antibody.  The label for that row tells us which part or parts of the cell the antibody has attached to. </p>\n<p>The task is complicated because the label provided for the large image MAY NOT be correct for every single cell contained within that large image.  </p>\n<p>So in a sense we are throwing an unknown antibody on a cell.  It's sticking to only specific parts of the cell.  Our job is to identify (for each row) what part of the cell has been stained (is green).  The <strong>green image is the key</strong> - a model could be partly successful using only the green image.   But including the red, blue and yellow allows a model to see the green in relationship to the cell structure.</p>\n<p>I have been doing lots of these competitions over that past several years - IMO this is one of the most complex ones I have attempted.  We need to match the segmentation performed by the hosts but with out the assistance of any ground truth, we need to label each single cell but the training ground truth does not apply to every single cell within an image, the metric is both difficult to understand and hard to incorporate into model training, the segmentation is a slow process and kaggle time constraints end up causing many of your early attempts at prediction to have out of time error.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1250182,
          "author_name": "Mikhail Gurevich",
          "author_url": "",
          "post_date": "2021-03-23T20:22:58.680000",
          "content": "<p>Hello, are you sure that hosts will use shared segmentator for the test data? I think that I read somewhere that test data was manually segmentated and then each cell was manually labeled.<br>\nIf you are right it means that we don't need to try to improve segmentation at all.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1250456,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-03-24T04:03:45.183000",
          "content": "<p>I read that they used the hpa segmentation and than annotated the results for the final.  I believe a host indicated it was 90% good if we just used the hpa segmentation.</p>\n<p>We do need to improve segmentation - it can be too slow if not used in well written script.  When I can get a LB score higher than .50 than I plan to try for some improvement in both speed and quality. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1250927,
      "author_name": "MatveySafroshkin",
      "author_url": "",
      "post_date": "2021-03-24T11:13:09.120000",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, Thank you a lot for your answer. It makes things a bit clearer. Why no one tried to make to label each cell individually?  Is it very hard to do? <br>\nAs you mentioned, your model it trying to predict, where antibodies are attached to, but why do we have mitochondria class?  Or does it mean, that antibodies are attached to Mitochondria ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1251296,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-03-24T16:49:06.877000",
          "content": "<p>Glad I could make things a little clearer.   I tried to think of a good analogy or two to help make it more clear.  No great examples, but here's a couple of the weak ones.</p>\n<ol>\n<li><p>Certain fabrics and certain laundry detergents can make your clothes very bright under a UV lamp - If I had a bunch of photos of folks in a night club taken under UV I might challenge you to segment out each individual and than tell me which piece of clothing had been washed with laundry detergent X.  All the men would be wearing shirts - but we would only want to label those shirts that were washed with detergent X.  Now the task would be complicated by the issue that someone washed his shirt three weeks ago and spent some time outside in the rain.  Or complicated by the need to identify men vs women shirts.  Or not include florescent pants, etc.  </p></li>\n<li><p>I take a bunch of images of car engines with thermal camera.  All engines have parts that can be put in a bunch of different classes.  I might weakly label engines for defects and challenge you to identify the part class that is over/under heated and responsible for the poor performance.  I don't want you to label each spark plug, but I do want you to label plugs that are too hot or too cold.  </p></li>\n</ol>\n<p>When I look at the single cell images it does not appear to be that difficult to  label each cell if you know what your doing.  On my todo list is to manually label single cell images for those classes with a small number of samples in the class.  The Mitotic spindle is very distinct in appearance and in some of the larger images it appears that only a single cell in the large image has this attachment point.  My segmentation ended up generating 491K of single cell images.  (will need to read the rules close to insure that hand labeling of training data is acceptable).  That's a lot of labeling to do !</p>\n<p>But I think the hosts want to try and advance the science by challenging kagglers to develop code that can take weakly labeled images and build good models.  In this area of science and many others it does seem like most of the expense of getting a good model is the hand label process.</p>\n<p>The hosts did provide a notebook showing examples of all the classes.<br>\n<a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#14.-Mitochondria\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns#14.-Mitochondria</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1249846": "Mat\n\nAll cells have many of the classes we are attempting to predict.  We are not attempting a simple body part identification.  The red, blue and yellow images can be combined into a single image that provides the structure of each cell.  \n\nStep 1 - the images provided contain a varying number of cells.  We need to segment that large image into the correct number of single cells.  The private test set was segmented by method(s) shared in the overview and used in many of the public kernels.  After automated segmentation humans did a further correction of the segments.  Since we are provided no ground truth for segmentation this is an important step but not the one that should be our major focus.  \n\nStep 2 - the cells have been stained with an antibody.  Our model is attempting to predict which parts of each single cell has the antibody attached to.  Antibodies will attach to one or more different parts of a cell.  The methods used for this staining result in the \"green\" image being the indicator of the intensity of the attachment.  Each row in the training data represents a different cell line and a different antibody.  The label for that row tells us which part or parts of the cell the antibody has attached to. \n\nThe task is complicated because the label provided for the large image MAY NOT be correct for every single cell contained within that large image.  \n\nSo in a sense we are throwing an unknown antibody on a cell.  It's sticking to only specific parts of the cell.  Our job is to identify (for each row) what part of the cell has been stained (is green).  The **green image is the key** - a model could be partly successful using only the green image.   But including the red, blue and yellow allows a model to see the green in relationship to the cell structure.\n\nI have been doing lots of these competitions over that past several years - IMO this is one of the most complex ones I have attempted.  We need to match the segmentation performed by the hosts but with out the assistance of any ground truth, we need to label each single cell but the training ground truth does not apply to every single cell within an image, the metric is both difficult to understand and hard to incorporate into model training, the segmentation is a slow process and kaggle time constraints end up causing many of your early attempts at prediction to have out of time error.\n\n\n\n",
    "1248526": "Good day everyone! \nI spent quite a lot of time reading forum and looking at the images. \nI don't really understand if we are predicting a label for cell or for cell components.\n \nFollowing this topic : https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda  and this one: https://www.kaggle.com/leoprovorov/hpa-part-1-understanding-the-topic \n\nI have several concerns. \nCould you please correct me : \n\n* NUCLEUS \n1) If we have a blue nucleus, we have a nucleoplasm, \n2)if we have a nucleoplasm, there should be a membrane. If there is no membrane, blue nucleus can not stay in the center. (we can somehow find a mask of membrane using dilatation/erosion on nucleus)\n3) Nucleoli, Nucleoli fibrillar center and Nuclear speckles, Nuclear bodies should be found in nucleuses. \n\n*OUTERCELL: \nGoldi apparatus :  makes a cell look \"furry\"\nMicrotubes: looks like a red web on the cell  \nMitochondria : a red hairy cell \ncytosol: looks like a green cloud with  dark center \n\n\nWhat does single label mitochondria mean?  Is it a correct labeling? \nHow can a cell be without nucleus?  \nThank you in advance.\nBest regards, \nMat\n\n",
    "1250927": "@pcjimmmy, Thank you a lot for your answer. It makes things a bit clearer. Why no one tried to make to label each cell individually?  Is it very hard to do? \nAs you mentioned, your model it trying to predict, where antibodies are attached to, but why do we have mitochondria class?  Or does it mean, that antibodies are attached to Mitochondria ?"
  }
}