{
  "id": 239827,
  "title": "Hybrid approach - ML and GOP (good old programming)",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239827",
  "author_name": "Philip Hucklesby",
  "post_date": "2021-05-17T18:59:13.166000",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Also from me, many thanks to the organisers for this awesome competition !<br>\nAnd doubly awesome watching the solutions: thanks to the participants for this learning experience.</p>\n<p>Myself, new to all this and watching from the sidelines, I didn't really get past a baseline attempt.<br>\nNevertheless, I share here what I have done - maybe something in it is of interest to someone.</p>\n<p>Inspired by the discussion with Luke: <a href=\"https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia\" target=\"_blank\">https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia</a>, I decided first to see what I could do without machine learning:</p>\n<ol>\n<li><p>What seems to work very well is recognising the nuclei with a simple threshold/findContours using standard openCV routines.</p></li>\n<li><p>Encouraged by this and since I then had contours of the nuclei as lists of line segments, I tried expanding the contours in a sort of poor-man's flood-fill to find the cell boundaries.  This was crude, but also worked surprisingly well, though with some characteristic spikes that I didn't get around to improving upon.</p></li>\n<li><p>I packaged this segmentation into a class CellSegmentatorOpenCV, with the idea that it might provide a substitutable component to the usually used CellSegmentator.  In the end, the only part of the interface that I used was the routine pred_cells, but at least regarding the method of consuming the images it should prove reasonably compatible for anyone who cares to try it.</p></li>\n<li><p>Working with the nucleus boundary, I created masks for identifiable parts of the cells: a ring just inside the nucleus (nuclear membrane) and a ring just outside the nucleus (microtubules, aggrosome …)</p></li>\n<li><p>Since findContours did a good job on the nucleus, I used it inside the nucleus to try to identify the nucleoli, nucleoli fibrillar center, nuclear speckles, nuclear bodies.</p></li>\n<li><p>Working with only a very few images from the training set (up to 20 of each label, choosing images without multiple labels), I collected some simple statistics (count, sum, mean, median, variance, correlation coefficients) in the regions identified. Storing these per Image, per cell and per \"Nucleolus\", I got a wide Excel table that I could eyeball for patterns that seemed strong enough to point towards some sort of identification.</p></li>\n<li><p>Built a graph data structure using networkx, modifying the cell segmentation to add the cells as nodes and the edges signifying neighbouring vicinity in the image - this was an easy extention of the poor-man's floodfill. (A next step which I didn't get around to would have been to identify neighbourhood to black background areas of the image and therefore identify groups of cells as disconnected subgraphs)</p></li>\n<li><p>I noticed that the difference between some labels seems to be mainly in texture: microtubules, endoplasmic reticulum, mitochondria and cytosol seem to occupy the same space but look different to the eye.  This seemed really to be a case for machine learning. So I generated some boxes in the space between the nucleus and the cell boundary with a view to testing them for texture.</p></li>\n<li><p>Added a test on the image segments in the boxes, (method, though maybe not appropriate, directly out of the tutorial here: <a href=\"https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/\" target=\"_blank\">https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/</a>)</p></li>\n</ol>\n<p>Competition time ran out on me at that point. This non-ML baseline scored 0.053 (that's right, with an extra zero after the decimal point !) This was using the standard CellSegmentator to generate the RLE masks although using the CellSegmentatorOpenCV for everything else (which made a small score difference). The texture test in Point 9 took the score to 0.069, but as a late submission.</p>\n<p>I dubbed the notebook Dataministic - sorry for the pun - here it is: <a href=\"https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg\" target=\"_blank\">https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg</a></p>",
  "messages": [
    {
      "id": 1312023,
      "postDate": "2021-05-17T18:59:13.167Z",
      "content": "<p>Hi all,</p>\n<p>Also from me, many thanks to the organisers for this awesome competition !<br>\nAnd doubly awesome watching the solutions: thanks to the participants for this learning experience.</p>\n<p>Myself, new to all this and watching from the sidelines, I didn't really get past a baseline attempt.<br>\nNevertheless, I share here what I have done - maybe something in it is of interest to someone.</p>\n<p>Inspired by the discussion with Luke: <a href=\"https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia\" target=\"_blank\">https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia</a>, I decided first to see what I could do without machine learning:</p>\n<ol>\n<li><p>What seems to work very well is recognising the nuclei with a simple threshold/findContours using standard openCV routines.</p></li>\n<li><p>Encouraged by this and since I then had contours of the nuclei as lists of line segments, I tried expanding the contours in a sort of poor-man's flood-fill to find the cell boundaries.  This was crude, but also worked surprisingly well, though with some characteristic spikes that I didn't get around to improving upon.</p></li>\n<li><p>I packaged this segmentation into a class CellSegmentatorOpenCV, with the idea that it might provide a substitutable component to the usually used CellSegmentator.  In the end, the only part of the interface that I used was the routine pred_cells, but at least regarding the method of consuming the images it should prove reasonably compatible for anyone who cares to try it.</p></li>\n<li><p>Working with the nucleus boundary, I created masks for identifiable parts of the cells: a ring just inside the nucleus (nuclear membrane) and a ring just outside the nucleus (microtubules, aggrosome …)</p></li>\n<li><p>Since findContours did a good job on the nucleus, I used it inside the nucleus to try to identify the nucleoli, nucleoli fibrillar center, nuclear speckles, nuclear bodies.</p></li>\n<li><p>Working with only a very few images from the training set (up to 20 of each label, choosing images without multiple labels), I collected some simple statistics (count, sum, mean, median, variance, correlation coefficients) in the regions identified. Storing these per Image, per cell and per \"Nucleolus\", I got a wide Excel table that I could eyeball for patterns that seemed strong enough to point towards some sort of identification.</p></li>\n<li><p>Built a graph data structure using networkx, modifying the cell segmentation to add the cells as nodes and the edges signifying neighbouring vicinity in the image - this was an easy extention of the poor-man's floodfill. (A next step which I didn't get around to would have been to identify neighbourhood to black background areas of the image and therefore identify groups of cells as disconnected subgraphs)</p></li>\n<li><p>I noticed that the difference between some labels seems to be mainly in texture: microtubules, endoplasmic reticulum, mitochondria and cytosol seem to occupy the same space but look different to the eye.  This seemed really to be a case for machine learning. So I generated some boxes in the space between the nucleus and the cell boundary with a view to testing them for texture.</p></li>\n<li><p>Added a test on the image segments in the boxes, (method, though maybe not appropriate, directly out of the tutorial here: <a href=\"https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/\" target=\"_blank\">https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/</a>)</p></li>\n</ol>\n<p>Competition time ran out on me at that point. This non-ML baseline scored 0.053 (that's right, with an extra zero after the decimal point !) This was using the standard CellSegmentator to generate the RLE masks although using the CellSegmentatorOpenCV for everything else (which made a small score difference). The texture test in Point 9 took the score to 0.069, but as a late submission.</p>\n<p>I dubbed the notebook Dataministic - sorry for the pun - here it is: <a href=\"https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg\" target=\"_blank\">https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg</a></p>",
      "rawMarkdown": "Hi all,\n\nAlso from me, many thanks to the organisers for this awesome competition !\nAnd doubly awesome watching the solutions: thanks to the participants for this learning experience.\n\nMyself, new to all this and watching from the sidelines, I didn't really get past a baseline attempt.\nNevertheless, I share here what I have done - maybe something in it is of interest to someone.\n\nInspired by the discussion with Luke: https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia, I decided first to see what I could do without machine learning:\n\n1. What seems to work very well is recognising the nuclei with a simple threshold/findContours using standard openCV routines.\n\n2. Encouraged by this and since I then had contours of the nuclei as lists of line segments, I tried expanding the contours in a sort of poor-man's flood-fill to find the cell boundaries.  This was crude, but also worked surprisingly well, though with some characteristic spikes that I didn't get around to improving upon.\n\n3. I packaged this segmentation into a class CellSegmentatorOpenCV, with the idea that it might provide a substitutable component to the usually used CellSegmentator.  In the end, the only part of the interface that I used was the routine pred_cells, but at least regarding the method of consuming the images it should prove reasonably compatible for anyone who cares to try it.\n\n4. Working with the nucleus boundary, I created masks for identifiable parts of the cells: a ring just inside the nucleus (nuclear membrane) and a ring just outside the nucleus (microtubules, aggrosome ...)\n\n5. Since findContours did a good job on the nucleus, I used it inside the nucleus to try to identify the nucleoli, nucleoli fibrillar center, nuclear speckles, nuclear bodies.\n\n6. Working with only a very few images from the training set (up to 20 of each label, choosing images without multiple labels), I collected some simple statistics (count, sum, mean, median, variance, correlation coefficients) in the regions identified. Storing these per Image, per cell and per \"Nucleolus\", I got a wide Excel table that I could eyeball for patterns that seemed strong enough to point towards some sort of identification.\n\n7. Built a graph data structure using networkx, modifying the cell segmentation to add the cells as nodes and the edges signifying neighbouring vicinity in the image - this was an easy extention of the poor-man's floodfill. (A next step which I didn't get around to would have been to identify neighbourhood to black background areas of the image and therefore identify groups of cells as disconnected subgraphs)\n\n8. I noticed that the difference between some labels seems to be mainly in texture: microtubules, endoplasmic reticulum, mitochondria and cytosol seem to occupy the same space but look different to the eye.  This seemed really to be a case for machine learning. So I generated some boxes in the space between the nucleus and the cell boundary with a view to testing them for texture.\n\n9. Added a test on the image segments in the boxes, (method, though maybe not appropriate, directly out of the tutorial here: https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/)\n\nCompetition time ran out on me at that point. This non-ML baseline scored 0.053 (that's right, with an extra zero after the decimal point !) This was using the standard CellSegmentator to generate the RLE masks although using the CellSegmentatorOpenCV for everything else (which made a small score difference). The texture test in Point 9 took the score to 0.069, but as a late submission.\n\nI dubbed the notebook Dataministic - sorry for the pun - here it is: https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg",
      "votes": 10
    },
    {
      "id": 1312162,
      "postDate": "2021-05-17T21:39:37.650Z",
      "content": "<p>Thanks for sharing. You probably looked at the images more closely, gave more thought than many people scored better than you! 💯</p>",
      "rawMarkdown": "Thanks for sharing. You probably looked at the images more closely, gave more thought than many people scored better than you! 💯",
      "replies": [
        {
          "id": 1313901,
          "postDate": "2021-05-18T19:28:04.150Z",
          "content": "<p>Thanks for the kind feedback, Shai - delighted to see it was worth writing about</p>",
          "rawMarkdown": "Thanks for the kind feedback, Shai - delighted to see it was worth writing about",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1312162,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2021-05-17T21:39:37.650000",
      "content": "<p>Thanks for sharing. You probably looked at the images more closely, gave more thought than many people scored better than you! 💯</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1313901,
          "author_name": "Philip Hucklesby",
          "author_url": "",
          "post_date": "2021-05-18T19:28:04.150000",
          "content": "<p>Thanks for the kind feedback, Shai - delighted to see it was worth writing about</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1312023": "Hi all,\n\nAlso from me, many thanks to the organisers for this awesome competition !\nAnd doubly awesome watching the solutions: thanks to the participants for this learning experience.\n\nMyself, new to all this and watching from the sidelines, I didn't really get past a baseline attempt.\nNevertheless, I share here what I have done - maybe something in it is of interest to someone.\n\nInspired by the discussion with Luke: https://www.kaggle.com/lukemerrick/who-needs-a-gpu-shallow-learning-in-julia, I decided first to see what I could do without machine learning:\n\n1. What seems to work very well is recognising the nuclei with a simple threshold/findContours using standard openCV routines.\n\n2. Encouraged by this and since I then had contours of the nuclei as lists of line segments, I tried expanding the contours in a sort of poor-man's flood-fill to find the cell boundaries.  This was crude, but also worked surprisingly well, though with some characteristic spikes that I didn't get around to improving upon.\n\n3. I packaged this segmentation into a class CellSegmentatorOpenCV, with the idea that it might provide a substitutable component to the usually used CellSegmentator.  In the end, the only part of the interface that I used was the routine pred_cells, but at least regarding the method of consuming the images it should prove reasonably compatible for anyone who cares to try it.\n\n4. Working with the nucleus boundary, I created masks for identifiable parts of the cells: a ring just inside the nucleus (nuclear membrane) and a ring just outside the nucleus (microtubules, aggrosome ...)\n\n5. Since findContours did a good job on the nucleus, I used it inside the nucleus to try to identify the nucleoli, nucleoli fibrillar center, nuclear speckles, nuclear bodies.\n\n6. Working with only a very few images from the training set (up to 20 of each label, choosing images without multiple labels), I collected some simple statistics (count, sum, mean, median, variance, correlation coefficients) in the regions identified. Storing these per Image, per cell and per \"Nucleolus\", I got a wide Excel table that I could eyeball for patterns that seemed strong enough to point towards some sort of identification.\n\n7. Built a graph data structure using networkx, modifying the cell segmentation to add the cells as nodes and the edges signifying neighbouring vicinity in the image - this was an easy extention of the poor-man's floodfill. (A next step which I didn't get around to would have been to identify neighbourhood to black background areas of the image and therefore identify groups of cells as disconnected subgraphs)\n\n8. I noticed that the difference between some labels seems to be mainly in texture: microtubules, endoplasmic reticulum, mitochondria and cytosol seem to occupy the same space but look different to the eye.  This seemed really to be a case for machine learning. So I generated some boxes in the space between the nucleus and the cell boundary with a view to testing them for texture.\n\n9. Added a test on the image segments in the boxes, (method, though maybe not appropriate, directly out of the tutorial here: https://www.pyimagesearch.com/2015/12/07/local-binary-patterns-with-python-opencv/)\n\nCompetition time ran out on me at that point. This non-ML baseline scored 0.053 (that's right, with an extra zero after the decimal point !) This was using the standard CellSegmentator to generate the RLE masks although using the CellSegmentatorOpenCV for everything else (which made a small score difference). The texture test in Point 9 took the score to 0.069, but as a late submission.\n\nI dubbed the notebook Dataministic - sorry for the pun - here it is: https://www.kaggle.com/philiphucklesby/dataministic-with-cellseg",
    "1312162": "Thanks for sharing. You probably looked at the images more closely, gave more thought than many people scored better than you! 💯"
  }
}