{
  "id": 69223,
  "title": "Previous publication: data cleaning step",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69223",
  "author_name": "",
  "post_date": "2018-10-21T17:16:15.861088500Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi guys, </p>\n\n<p>I was wondering whether you read the previous publication (<a href=\"https://www.nature.com/articles/nbt.4225\">https://www.nature.com/articles/nbt.4225</a>) on the classification procedure and what you think about the steps the authors implemented there. In particular, the step regarding data cleaning is interesting. The authors decide to throw out all cells whose nucleus is at the image border, though the cytoplasm of the cell can contact the edge of the image. Has someone tried both approaches? I could imagine some cells that fit these criteria are adding more noise than data since potentially large part of the cell is missing. Maybe it's better to cut out cells that touch the edge altogether?</p>\n\n<p>Best,\nLukas</p>",
  "messages": [
    {
      "id": "407736",
      "postDate": "10/21/2018 17:16:15",
      "content": "<p>Hi guys, </p>\n\n<p>I was wondering whether you read the previous publication (<a href=\"https://www.nature.com/articles/nbt.4225\">https://www.nature.com/articles/nbt.4225</a>) on the classification procedure and what you think about the steps the authors implemented there. In particular, the step regarding data cleaning is interesting. The authors decide to throw out all cells whose nucleus is at the image border, though the cytoplasm of the cell can contact the edge of the image. Has someone tried both approaches? I could imagine some cells that fit these criteria are adding more noise than data since potentially large part of the cell is missing. Maybe it's better to cut out cells that touch the edge altogether?</p>\n\n<p>Best,\nLukas</p>",
      "rawMarkdown": "Hi guys, \n\nI was wondering whether you read the previous publication (https://www.nature.com/articles/nbt.4225) on the classification procedure and what you think about the steps the authors implemented there. In particular, the step regarding data cleaning is interesting. The authors decide to throw out all cells whose nucleus is at the image border, though the cytoplasm of the cell can contact the edge of the image. Has someone tried both approaches? I could imagine some cells that fit these criteria are adding more noise than data since potentially large part of the cell is missing. Maybe it's better to cut out cells that touch the edge altogether?\n\nBest,\nLukas",
      "votes": null
    },
    {
      "id": "407784",
      "postDate": "10/21/2018 18:46:23",
      "content": "<p>One tuning approach I have on my list is to make a network that will give me bounding boxes of cells to process.  Then from the large scale images I can get smaller images of cells to train a network on.</p>",
      "rawMarkdown": "One tuning approach I have on my list is to make a network that will give me bounding boxes of cells to process.  Then from the large scale images I can get smaller images of cells to train a network on.",
      "votes": null
    },
    {
      "id": "407794",
      "postDate": "10/21/2018 19:04:43",
      "content": "<p>Ok, so the idea would be to break down one image into several images, each containing one cell and feeding them into the network. One could potentially maintain the original image size and center the cell based on the center of mass of the nucleus in order to circumvent issues with different bounding box sizes and additionally make the pictures more comparable. Ultimately, you would still have to implement a decision rule for the orginal picture where the single cells came from. But that is something I will probably try.  </p>",
      "rawMarkdown": "Ok, so the idea would be to break down one image into several images, each containing one cell and feeding them into the network. One could potentially maintain the original image size and center the cell based on the center of mass of the nucleus in order to circumvent issues with different bounding box sizes and additionally make the pictures more comparable. Ultimately, you would still have to implement a decision rule for the orginal picture where the single cells came from. But that is something I will probably try.",
      "votes": null
    },
    {
      "id": "407847",
      "postDate": "10/21/2018 22:04:19",
      "content": "<p>Exactly my thought, some manual annotation will be needed to identify the cells. My custom model is doing ok so far, but I think this would improve things. </p>",
      "rawMarkdown": "Exactly my thought, some manual annotation will be needed to identify the cells. My custom model is doing ok so far, but I think this would improve things.",
      "votes": null
    },
    {
      "id": "407906",
      "postDate": "10/22/2018 01:20:26",
      "content": "<p>It seems a great idea that break down one image ito several images.</p>",
      "rawMarkdown": "It seems a great idea that break down one image ito several images.",
      "votes": null
    },
    {
      "id": "408405",
      "postDate": "10/22/2018 20:25:49",
      "content": "<p>Breaking down the images to the cell level only works if the proteins labeled are present in all the cells. `Does someone with domain knowledge know if this is the case? I feel like some proteins can be found in only one or two cells on some images.</p>",
      "rawMarkdown": "Breaking down the images to the cell level only works if the proteins labeled are present in all the cells. `Does someone with domain knowledge know if this is the case? I feel like some proteins can be found in only one or two cells on some images.",
      "votes": null
    },
    {
      "id": "408411",
      "postDate": "10/22/2018 20:43:00",
      "content": "<blockquote>\n  <p>I feel like some proteins can be found in only one or two cells on some images.</p>\n</blockquote>\n\n<p>There is a stochastic aspect of protein expression, just like in any other process that happens because of random molecular collisions and (dis)associations. Still, the differences shouldn't be so large that in a field of neighboring cells there are some with good expression and others without any protein. Even if that's happening - any examples? - I can't imagine that it would be more than a small fraction of images.</p>",
      "rawMarkdown": "&gt; I feel like some proteins can be found in only one or two cells on some images.\n\nThere is a stochastic aspect of protein expression, just like in any other process that happens because of random molecular collisions and (dis)associations. Still, the differences shouldn't be so large that in a field of neighboring cells there are some with good expression and others without any protein. Even if that's happening - any examples? - I can't imagine that it would be more than a small fraction of images.",
      "votes": null
    },
    {
      "id": "408655",
      "postDate": "10/23/2018 08:39:17",
      "content": "<p>We do indeed see expression of certain patterns in a fraction of the cells. The fraction of images in which we see a large amount of heterogeneity (including intensity changes, not only trans-locations/missing locations) this is the case is &gt;10%, so it is significant. That said, as mentioned, we chose to split up the image on a per-cell level and then deal with noise introduced by this afterwards.</p>",
      "rawMarkdown": "We do indeed see expression of certain patterns in a fraction of the cells. The fraction of images in which we see a large amount of heterogeneity (including intensity changes, not only trans-locations/missing locations) this is the case is &gt;10%, so it is significant. That said, as mentioned, we chose to split up the image on a per-cell level and then deal with noise introduced by this afterwards.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 407784,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/21/2018 18:46:23",
      "content": "<p>One tuning approach I have on my list is to make a network that will give me bounding boxes of cells to process.  Then from the large scale images I can get smaller images of cells to train a network on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 407794,
      "author_name": "lukano",
      "author_url": "",
      "post_date": "10/21/2018 19:04:43",
      "content": "<p>Ok, so the idea would be to break down one image into several images, each containing one cell and feeding them into the network. One could potentially maintain the original image size and center the cell based on the center of mass of the nucleus in order to circumvent issues with different bounding box sizes and additionally make the pictures more comparable. Ultimately, you would still have to implement a decision rule for the orginal picture where the single cells came from. But that is something I will probably try.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 407847,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/21/2018 22:04:19",
          "content": "<p>Exactly my thought, some manual annotation will be needed to identify the cells. My custom model is doing ok so far, but I think this would improve things. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 407906,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "10/22/2018 01:20:26",
      "content": "<p>It seems a great idea that break down one image ito several images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 408405,
      "author_name": "pforet95",
      "author_url": "",
      "post_date": "10/22/2018 20:25:49",
      "content": "<p>Breaking down the images to the cell level only works if the proteins labeled are present in all the cells. `Does someone with domain knowledge know if this is the case? I feel like some proteins can be found in only one or two cells on some images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 408411,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "10/22/2018 20:43:00",
          "content": "<blockquote>\n  <p>I feel like some proteins can be found in only one or two cells on some images.</p>\n</blockquote>\n\n<p>There is a stochastic aspect of protein expression, just like in any other process that happens because of random molecular collisions and (dis)associations. Still, the differences shouldn't be so large that in a field of neighboring cells there are some with good expression and others without any protein. Even if that's happening - any examples? - I can't imagine that it would be more than a small fraction of images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 408655,
          "author_name": "devins",
          "author_url": "",
          "post_date": "10/23/2018 08:39:17",
          "content": "<p>We do indeed see expression of certain patterns in a fraction of the cells. The fraction of images in which we see a large amount of heterogeneity (including intensity changes, not only trans-locations/missing locations) this is the case is &gt;10%, so it is significant. That said, as mentioned, we chose to split up the image on a per-cell level and then deal with noise introduced by this afterwards.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "407736": "Hi guys, \n\nI was wondering whether you read the previous publication (https://www.nature.com/articles/nbt.4225) on the classification procedure and what you think about the steps the authors implemented there. In particular, the step regarding data cleaning is interesting. The authors decide to throw out all cells whose nucleus is at the image border, though the cytoplasm of the cell can contact the edge of the image. Has someone tried both approaches? I could imagine some cells that fit these criteria are adding more noise than data since potentially large part of the cell is missing. Maybe it's better to cut out cells that touch the edge altogether?\n\nBest,\nLukas",
    "407784": "One tuning approach I have on my list is to make a network that will give me bounding boxes of cells to process.  Then from the large scale images I can get smaller images of cells to train a network on.",
    "407794": "Ok, so the idea would be to break down one image into several images, each containing one cell and feeding them into the network. One could potentially maintain the original image size and center the cell based on the center of mass of the nucleus in order to circumvent issues with different bounding box sizes and additionally make the pictures more comparable. Ultimately, you would still have to implement a decision rule for the orginal picture where the single cells came from. But that is something I will probably try.",
    "407847": "Exactly my thought, some manual annotation will be needed to identify the cells. My custom model is doing ok so far, but I think this would improve things.",
    "407906": "It seems a great idea that break down one image ito several images.",
    "408405": "Breaking down the images to the cell level only works if the proteins labeled are present in all the cells. `Does someone with domain knowledge know if this is the case? I feel like some proteins can be found in only one or two cells on some images.",
    "408411": "&gt; I feel like some proteins can be found in only one or two cells on some images.\n\nThere is a stochastic aspect of protein expression, just like in any other process that happens because of random molecular collisions and (dis)associations. Still, the differences shouldn't be so large that in a field of neighboring cells there are some with good expression and others without any protein. Even if that's happening - any examples? - I can't imagine that it would be more than a small fraction of images.",
    "408655": "We do indeed see expression of certain patterns in a fraction of the cells. The fraction of images in which we see a large amount of heterogeneity (including intensity changes, not only trans-locations/missing locations) this is the case is &gt;10%, so it is significant. That said, as mentioned, we chose to split up the image on a per-cell level and then deal with noise introduced by this afterwards."
  },
  "source": "meta"
}