{
  "id": 103683,
  "title": "CellProfiler software and the Cell Painting assay",
  "url": "/competitions/recursion-cellular-image-classification/discussion/103683",
  "author_name": "Nicholas Schafer",
  "post_date": "2019-08-11T01:36:13.618000",
  "votes": 13,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello, fellow cell classifiers - </p>\n\n<p>I am currently participating in my first round of Kaggle competitions. I was interested in this one because of my background in biophysics. I actually interviewed at Recursion a few years ago and have been following their progress with interest ever since. I can see how the subject of this competition is an important problem that they need to solve, and I'm hoping that all of you will be able to provide them with a solution that works well. </p>\n\n<p>What I found most striking while reading through the public discussion boards and kernels was how unfamiliar most of the terminology and techniques being used would be to a classically trained biologist, or to a classically trained scientist of almost any type, for that matter. This is clearly a computer vision problem, but even among computer vision problems it does not square with the typical characterization of computer vision problems as being amongst those problems that are easy for human's to solve but difficult for computers. Just imagine trying to tell apart more than 1,000 different phenotypes simply by looking under the microscope! Even desperate graduate students couldn't be pushed that far.</p>\n\n<p>And that brings me to the next most striking aspect of the competition: the top of the leaderboard. In a multiclass classification problem with more than 1,000 classes, the top of the leaderboard currently shows a 99% accuracy. Having read the discussion board, I realize that this is not, strictly speaking, a 1,000+ class classification problem. Because of a data leak, the problem can be broken down into 4 subproblems of 277 classes (be sure to check this out on the discussion boards if you haven't seen it already). Nonetheless, 99% accuracy seems quite high to me. Perhaps the private leaderboard will eventually show that some of those solutions are overfitting to the public leaderboard data, but with only 15 submissions to date, it seems possible that SharksWithLasers really does have a good solution to Recurion's problem. So, how did they do it?</p>\n\n<p>Sorry, but I don't know how they did it. However, could they have been helped by using... Recursion's own software? A little bit of digging turned up the following potentially helpful resources: \n<a href=\"https://github.com/recursionpharma/CellProfiler\">https://github.com/recursionpharma/CellProfiler</a>\n<a href=\"https://github.com/CellProfiler/CellProfiler-Analyst\">https://github.com/CellProfiler/CellProfiler-Analyst</a>\n<a href=\"https://cellprofiler.org/\">https://cellprofiler.org/</a>\n<a href=\"https://www.nature.com/articles/nprot.2016.105\">https://www.nature.com/articles/nprot.2016.105</a></p>\n\n<p>Unsurprisingly, Recursion has already been working on this problem, namely being able to distinguish biological differences from batch differences. The quotes below are from the Nature Methods paper linked above. Properly configured, the Cell Profiler software can be used to create \"~1,500 morphological features ... to produce a rich profile that is suitable for the detection of subtle phenotypes\". They also note in their description of the protocol (experimental and computational) that \"feature extraction and data analysis take an additional 1-2 weeks.\" The Cell Painting assay described in that paper is essentially the same as was used to produce the data in this competition. The features produced by Cell Profiler include \"staining intensities, textural patterns, size, and shape of the labeled cellular structures, as well as correlations between stains across channels, and adjacency relationships between cells and among intracellular structures.\" In that paper they also make direct reference to the types of problems that this competition is focused on. \"A perennial concern with assay development is that any technical sources of variation can have an impact on all the wells and/or plates such that any biological signals are overwhelmed by systematic noise introduced by sample preparation.\" \"The image feature extraction workflow for Cell Painting is divided into three asks,, each of which is performed by a CellProfiler pipeline: (i) illumination, (ii) quality control (QC), and (iii) morphological feature extraction.\" \"Last, data analysis across separately performed experiments is likely to be complicated, requiring proper control over potentially substantial effects of differences in cell seeding, growth, and other <em>batch-related</em> or other systematic artifacts. Protocols for such cases have not yet been developed.\"</p>\n\n<p>Well, what are you waiting for? Go forth and develop :-) A potentially useful CellProfiler pipeline can be found here:\n<a href=\"https://github.com/gigascience/paper-bray2017\">https://github.com/gigascience/paper-bray2017</a></p>\n\n<p>And others here:\n<a href=\"https://cellprofiler.org/examples/published_pipelines\">https://cellprofiler.org/examples/published_pipelines</a></p>\n\n<p>Best,\nNick</p>",
  "messages": [
    {
      "id": 596613,
      "postDate": "2019-08-11T01:36:13.617Z",
      "content": "<p>Hello, fellow cell classifiers - </p>\n\n<p>I am currently participating in my first round of Kaggle competitions. I was interested in this one because of my background in biophysics. I actually interviewed at Recursion a few years ago and have been following their progress with interest ever since. I can see how the subject of this competition is an important problem that they need to solve, and I'm hoping that all of you will be able to provide them with a solution that works well. </p>\n\n<p>What I found most striking while reading through the public discussion boards and kernels was how unfamiliar most of the terminology and techniques being used would be to a classically trained biologist, or to a classically trained scientist of almost any type, for that matter. This is clearly a computer vision problem, but even among computer vision problems it does not square with the typical characterization of computer vision problems as being amongst those problems that are easy for human's to solve but difficult for computers. Just imagine trying to tell apart more than 1,000 different phenotypes simply by looking under the microscope! Even desperate graduate students couldn't be pushed that far.</p>\n\n<p>And that brings me to the next most striking aspect of the competition: the top of the leaderboard. In a multiclass classification problem with more than 1,000 classes, the top of the leaderboard currently shows a 99% accuracy. Having read the discussion board, I realize that this is not, strictly speaking, a 1,000+ class classification problem. Because of a data leak, the problem can be broken down into 4 subproblems of 277 classes (be sure to check this out on the discussion boards if you haven't seen it already). Nonetheless, 99% accuracy seems quite high to me. Perhaps the private leaderboard will eventually show that some of those solutions are overfitting to the public leaderboard data, but with only 15 submissions to date, it seems possible that SharksWithLasers really does have a good solution to Recurion's problem. So, how did they do it?</p>\n\n<p>Sorry, but I don't know how they did it. However, could they have been helped by using... Recursion's own software? A little bit of digging turned up the following potentially helpful resources: \n<a href=\"https://github.com/recursionpharma/CellProfiler\">https://github.com/recursionpharma/CellProfiler</a>\n<a href=\"https://github.com/CellProfiler/CellProfiler-Analyst\">https://github.com/CellProfiler/CellProfiler-Analyst</a>\n<a href=\"https://cellprofiler.org/\">https://cellprofiler.org/</a>\n<a href=\"https://www.nature.com/articles/nprot.2016.105\">https://www.nature.com/articles/nprot.2016.105</a></p>\n\n<p>Unsurprisingly, Recursion has already been working on this problem, namely being able to distinguish biological differences from batch differences. The quotes below are from the Nature Methods paper linked above. Properly configured, the Cell Profiler software can be used to create \"~1,500 morphological features ... to produce a rich profile that is suitable for the detection of subtle phenotypes\". They also note in their description of the protocol (experimental and computational) that \"feature extraction and data analysis take an additional 1-2 weeks.\" The Cell Painting assay described in that paper is essentially the same as was used to produce the data in this competition. The features produced by Cell Profiler include \"staining intensities, textural patterns, size, and shape of the labeled cellular structures, as well as correlations between stains across channels, and adjacency relationships between cells and among intracellular structures.\" In that paper they also make direct reference to the types of problems that this competition is focused on. \"A perennial concern with assay development is that any technical sources of variation can have an impact on all the wells and/or plates such that any biological signals are overwhelmed by systematic noise introduced by sample preparation.\" \"The image feature extraction workflow for Cell Painting is divided into three asks,, each of which is performed by a CellProfiler pipeline: (i) illumination, (ii) quality control (QC), and (iii) morphological feature extraction.\" \"Last, data analysis across separately performed experiments is likely to be complicated, requiring proper control over potentially substantial effects of differences in cell seeding, growth, and other <em>batch-related</em> or other systematic artifacts. Protocols for such cases have not yet been developed.\"</p>\n\n<p>Well, what are you waiting for? Go forth and develop :-) A potentially useful CellProfiler pipeline can be found here:\n<a href=\"https://github.com/gigascience/paper-bray2017\">https://github.com/gigascience/paper-bray2017</a></p>\n\n<p>And others here:\n<a href=\"https://cellprofiler.org/examples/published_pipelines\">https://cellprofiler.org/examples/published_pipelines</a></p>\n\n<p>Best,\nNick</p>",
      "rawMarkdown": "Hello, fellow cell classifiers - \n\nI am currently participating in my first round of Kaggle competitions. I was interested in this one because of my background in biophysics. I actually interviewed at Recursion a few years ago and have been following their progress with interest ever since. I can see how the subject of this competition is an important problem that they need to solve, and I'm hoping that all of you will be able to provide them with a solution that works well. \n\nWhat I found most striking while reading through the public discussion boards and kernels was how unfamiliar most of the terminology and techniques being used would be to a classically trained biologist, or to a classically trained scientist of almost any type, for that matter. This is clearly a computer vision problem, but even among computer vision problems it does not square with the typical characterization of computer vision problems as being amongst those problems that are easy for human's to solve but difficult for computers. Just imagine trying to tell apart more than 1,000 different phenotypes simply by looking under the microscope! Even desperate graduate students couldn't be pushed that far.\n\nAnd that brings me to the next most striking aspect of the competition: the top of the leaderboard. In a multiclass classification problem with more than 1,000 classes, the top of the leaderboard currently shows a 99% accuracy. Having read the discussion board, I realize that this is not, strictly speaking, a 1,000+ class classification problem. Because of a data leak, the problem can be broken down into 4 subproblems of 277 classes (be sure to check this out on the discussion boards if you haven't seen it already). Nonetheless, 99% accuracy seems quite high to me. Perhaps the private leaderboard will eventually show that some of those solutions are overfitting to the public leaderboard data, but with only 15 submissions to date, it seems possible that SharksWithLasers really does have a good solution to Recurion's problem. So, how did they do it?\n\nSorry, but I don't know how they did it. However, could they have been helped by using... Recursion's own software? A little bit of digging turned up the following potentially helpful resources: \nhttps://github.com/recursionpharma/CellProfiler\nhttps://github.com/CellProfiler/CellProfiler-Analyst\nhttps://cellprofiler.org/\nhttps://www.nature.com/articles/nprot.2016.105\n\nUnsurprisingly, Recursion has already been working on this problem, namely being able to distinguish biological differences from batch differences. The quotes below are from the Nature Methods paper linked above. Properly configured, the Cell Profiler software can be used to create \"~1,500 morphological features ... to produce a rich profile that is suitable for the detection of subtle phenotypes\". They also note in their description of the protocol (experimental and computational) that \"feature extraction and data analysis take an additional 1-2 weeks.\" The Cell Painting assay described in that paper is essentially the same as was used to produce the data in this competition. The features produced by Cell Profiler include \"staining intensities, textural patterns, size, and shape of the labeled cellular structures, as well as correlations between stains across channels, and adjacency relationships between cells and among intracellular structures.\" In that paper they also make direct reference to the types of problems that this competition is focused on. \"A perennial concern with assay development is that any technical sources of variation can have an impact on all the wells and/or plates such that any biological signals are overwhelmed by systematic noise introduced by sample preparation.\" \"The image feature extraction workflow for Cell Painting is divided into three asks,, each of which is performed by a CellProfiler pipeline: (i) illumination, (ii) quality control (QC), and (iii) morphological feature extraction.\" \"Last, data analysis across separately performed experiments is likely to be complicated, requiring proper control over potentially substantial effects of differences in cell seeding, growth, and other *batch-related* or other systematic artifacts. Protocols for such cases have not yet been developed.\"\n\nWell, what are you waiting for? Go forth and develop :-) A potentially useful CellProfiler pipeline can be found here:\nhttps://github.com/gigascience/paper-bray2017\n\nAnd others here:\nhttps://cellprofiler.org/examples/published_pipelines\n\nBest,\nNick",
      "votes": 13
    },
    {
      "id": 596817,
      "postDate": "2019-08-11T10:54:34.633Z",
      "content": "<p><a href=\"/npschafer\">@npschafer</a> Thanks for the info. \n1. We are at 0.925 and we still have a lot to do in the \"conventional ML\" area, i.e. without using any Recursion software (I don't know about the 0.99 but I estimate it's possible to get &gt;0.95 without it).\n2. I estimate that by the end of the competition a few teams will get to a Public LB of 1.0.</p>",
      "rawMarkdown": "@npschafer Thanks for the info. \n1. We are at 0.925 and we still have a lot to do in the \"conventional ML\" area, i.e. without using any Recursion software (I don't know about the 0.99 but I estimate it's possible to get &gt;0.95 without it).\n2. I estimate that by the end of the competition a few teams will get to a Public LB of 1.0.",
      "votes": 11,
      "replies": [
        {
          "id": 596967,
          "postDate": "2019-08-11T15:38:11.113Z",
          "content": "<p>Thanks for sharing a bit about your top-performing solution. I had been wondering whether anyone near the top had gotten there with \"conventional ML\" and without using bespoke cell profiling techniques. Evidently, it is possible! This probably shows my ignorance re: the power of computer vision methods more than anything else, but I think that it also speaks to the amount and quality of the data provided. Nice work!</p>",
          "rawMarkdown": "Thanks for sharing a bit about your top-performing solution. I had been wondering whether anyone near the top had gotten there with \"conventional ML\" and without using bespoke cell profiling techniques. Evidently, it is possible! This probably shows my ignorance re: the power of computer vision methods more than anything else, but I think that it also speaks to the amount and quality of the data provided. Nice work!"
        }
      ]
    },
    {
      "id": 596902,
      "postDate": "2019-08-11T13:31:05.410Z",
      "content": "<p>Thanks for this invaluable information! I immediately tried CellProfiler but unfortunately found that the software will take a huge amount of time to process all the images we have. (Doing illumination correction step only for the images from U2OS-01 experiment showed an estimated time of 14 hours with 4 cores... maybe it'll take a month to process all the images) But I'm still thinking that this image-classification competition can be translated into tabular data-based classification competition with conventional cell profiling approach and extremely curious about the performance of the classifier exploiting single-cell based features. Again, thank you for such an interesting resource!</p>",
      "rawMarkdown": "Thanks for this invaluable information! I immediately tried CellProfiler but unfortunately found that the software will take a huge amount of time to process all the images we have. (Doing illumination correction step only for the images from U2OS-01 experiment showed an estimated time of 14 hours with 4 cores... maybe it'll take a month to process all the images) But I'm still thinking that this image-classification competition can be translated into tabular data-based classification competition with conventional cell profiling approach and extremely curious about the performance of the classifier exploiting single-cell based features. Again, thank you for such an interesting resource!",
      "votes": 1,
      "replies": [
        {
          "id": 596963,
          "postDate": "2019-08-11T15:34:39.320Z",
          "content": "<p>I was thinking the same thing re: converting this from a computer vision problem into a tabular data problem. They claim that the CellProfiler software can be run efficiently in distributed computing environments, which, as you point out, will seemingly be necessary when you are working with data sets of this size (or larger). With all of the information that Recursion provided, I'm wondering what their rationale was for not mentioning the CellProfiler software. Perhaps they were hoping that people would come up with good solutions that are largely orthogonal to traditional cell profiling techniques. And based on <a href=\"/yuval6967\">@yuval6967</a>'s comments, it seems that people have been able to do so!</p>",
          "rawMarkdown": "I was thinking the same thing re: converting this from a computer vision problem into a tabular data problem. They claim that the CellProfiler software can be run efficiently in distributed computing environments, which, as you point out, will seemingly be necessary when you are working with data sets of this size (or larger). With all of the information that Recursion provided, I'm wondering what their rationale was for not mentioning the CellProfiler software. Perhaps they were hoping that people would come up with good solutions that are largely orthogonal to traditional cell profiling techniques. And based on @yuval6967's comments, it seems that people have been able to do so!"
        }
      ]
    },
    {
      "id": 596651,
      "postDate": "2019-08-11T03:43:49.227Z",
      "content": "<p>Yes, it certainly puzzles me how teams who surpassed the 0.9 score keep improving while the rests are fairly stagnant in our positions. </p>\n\n<p>I wonder would it then be appropriate to perform 4 classification tasks instead of 1, given that we only account for the <code>sirna</code> when we do predictions for our test set now, as shown by nosound in his kernel.</p>\n\n<p>Another thing is that most people (including me) seem to not have found the given controls useful even though chimael mentioned that this is a \"relative\" image classification task.</p>",
      "rawMarkdown": "Yes, it certainly puzzles me how teams who surpassed the 0.9 score keep improving while the rests are fairly stagnant in our positions. \n\nI wonder would it then be appropriate to perform 4 classification tasks instead of 1, given that we only account for the `sirna` when we do predictions for our test set now, as shown by nosound in his kernel.\n\nAnother thing is that most people (including me) seem to not have found the given controls useful even though chimael mentioned that this is a \"relative\" image classification task.",
      "votes": 1,
      "replies": [
        {
          "id": 596761,
          "postDate": "2019-08-11T08:29:22.013Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 596969,
          "postDate": "2019-08-11T15:41:23.867Z",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> see <a href=\"/yuval6967\">@yuval6967</a>'s comment above. Apparently, it is not necessary to apply traditional cell profiling techniques in order to get a good solution. That being said, I don't think that CellProfiler's usefulness has yet been ruled out. And yes, I think that there is potential gain to be had in leveraging all of the information available to us including the fact that the siRNAs always come in groups of 277. Good luck!</p>",
          "rawMarkdown": "@wjshenggggg see @yuval6967's comment above. Apparently, it is not necessary to apply traditional cell profiling techniques in order to get a good solution. That being said, I don't think that CellProfiler's usefulness has yet been ruled out. And yes, I think that there is potential gain to be had in leveraging all of the information available to us including the fact that the siRNAs always come in groups of 277. Good luck!"
        },
        {
          "id": 597659,
          "postDate": "2019-08-12T16:13:31.023Z",
          "content": "<p>It is really amazing to see a LB at 0.99. However, I noticed that the test data are from the same cell lines as the training data. Thus a classsifier specific for each cell line may be the way to go. </p>\n\n<p>If it is really the winner's solution, then the models will be useless on other cell lines. I wonder if this competition design was made on purpose.</p>",
          "rawMarkdown": "It is really amazing to see a LB at 0.99. However, I noticed that the test data are from the same cell lines as the training data. Thus a classsifier specific for each cell line may be the way to go. \n\nIf it is really the winner's solution, then the models will be useless on other cell lines. I wonder if this competition design was made on purpose.",
          "votes": 1
        },
        {
          "id": 597672,
          "postDate": "2019-08-12T16:34:04.607Z",
          "content": "<p>hi <a href=\"/chimael\">@chimael</a>, thanks for the input. That being said, is it correct that if we build cell line specific classifiers, \"controls\" would not be as helpful as you suggested earlier. Since it's cell line specific, we do not need to frame this as a \"relative\" image classification task.</p>",
          "rawMarkdown": "hi @chimael, thanks for the input. That being said, is it correct that if we build cell line specific classifiers, \"controls\" would not be as helpful as you suggested earlier. Since it's cell line specific, we do not need to frame this as a \"relative\" image classification task.",
          "votes": 1
        },
        {
          "id": 598121,
          "postDate": "2019-08-13T06:51:28.930Z",
          "content": "<p>I think your are correct. A cell specific classifier don't need controls to interpret the pictures of treatments. \nI was first hoping that this competition would show that we can produce a model that can recognize the effects of an siRNA across cell lines. Now I think that this competition will produce models that can recognize different phenotypes of a cell line. I may have misunderstood the intention of this competition. The function of specific siRNA may not be the main focus of the competition host.</p>",
          "rawMarkdown": "I think your are correct. A cell specific classifier don't need controls to interpret the pictures of treatments. \nI was first hoping that this competition would show that we can produce a model that can recognize the effects of an siRNA across cell lines. Now I think that this competition will produce models that can recognize different phenotypes of a cell line. I may have misunderstood the intention of this competition. The function of specific siRNA may not be the main focus of the competition host.",
          "votes": 1
        },
        {
          "id": 598435,
          "postDate": "2019-08-13T15:15:25.547Z",
          "content": "<p>Thank you, chimael. I agree with the way you understand the host's motivation in hosting the competition. It might just be unfortunate that the setup is not as perfect as expected - pointed out by Berton, \"Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment by plate.\"</p>",
          "rawMarkdown": "Thank you, chimael. I agree with the way you understand the host's motivation in hosting the competition. It might just be unfortunate that the setup is not as perfect as expected - pointed out by Berton, \"Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment by plate.\""
        }
      ]
    },
    {
      "id": 597038,
      "postDate": "2019-08-11T17:40:15.887Z",
      "content": "<p>Just so proper people are attributed correctly, CellProfiler was/is developed by Anne Carpenter's group at the Broad Institute rather than by Recursion.</p>\n\n<p>To add to the CellProfiler vs. CNN discussion, believe it or not morphological cell profiling is a pretty active field and I encourage you to have a look at some papers, but currently it's found that CNN's are out-performing more classical segmentation approaches (CellProfiler) in classification tasks of small molecules. <a href=\"https://www.biorxiv.org/content/10.1101/161422v1\">Here's</a> one such example.</p>",
      "rawMarkdown": "Just so proper people are attributed correctly, CellProfiler was/is developed by Anne Carpenter's group at the Broad Institute rather than by Recursion.\n\nTo add to the CellProfiler vs. CNN discussion, believe it or not morphological cell profiling is a pretty active field and I encourage you to have a look at some papers, but currently it's found that CNN's are out-performing more classical segmentation approaches (CellProfiler) in classification tasks of small molecules. [Here's](https://www.biorxiv.org/content/10.1101/161422v1) one such example.",
      "votes": 2,
      "replies": [
        {
          "id": 597095,
          "postDate": "2019-08-11T19:55:16.167Z",
          "content": "<p>Good catch, <a href=\"/swarchal\">@swarchal</a>. Thanks for the correction. I incorrectly attributed the software because the first Github repository that I ran into was actually Recursion's fork of the original repository and, also, there are Recursion co-authors on the paper that I read. Now it makes more sense why Recursion would not have felt the need to mention CellProfiler in the background information for this competition.</p>\n\n<p>Are you under the impression that a pipeline that takes the features that CellProfiler produces and combines them with other machine learning techniques on the resulting tabulated data would not be competitive? If so, then the basic premise of my post seems to be incorrect, and it is probably not worth people's time to try to get CellProfiler to run on all of the data provided by the competition.</p>",
          "rawMarkdown": "Good catch, @swarchal. Thanks for the correction. I incorrectly attributed the software because the first Github repository that I ran into was actually Recursion's fork of the original repository and, also, there are Recursion co-authors on the paper that I read. Now it makes more sense why Recursion would not have felt the need to mention CellProfiler in the background information for this competition.\n\nAre you under the impression that a pipeline that takes the features that CellProfiler produces and combines them with other machine learning techniques on the resulting tabulated data would not be competitive? If so, then the basic premise of my post seems to be incorrect, and it is probably not worth people's time to try to get CellProfiler to run on all of the data provided by the competition."
        },
        {
          "id": 597107,
          "postDate": "2019-08-11T20:43:09.693Z",
          "content": "<p>I don't want to say it wouldn't be competitive, most of the published research comparing these methods at classifying cell morphology has focused on a single dataset (BBBC021). It would be great if a CellProfiler or similar software approach did well, as the features extracted are much more interpretable for biologists, e.g which features are changing in response to a particular siRNA?</p>\n\n<p>If anyone decides to go down the CellProfiler route, there's a nice review <a href=\"https://www.nature.com/articles/nmeth.4397\">here</a>, though I am a bit biased ;).</p>",
          "rawMarkdown": "I don't want to say it wouldn't be competitive, most of the published research comparing these methods at classifying cell morphology has focused on a single dataset (BBBC021). It would be great if a CellProfiler or similar software approach did well, as the features extracted are much more interpretable for biologists, e.g which features are changing in response to a particular siRNA?\n\nIf anyone decides to go down the CellProfiler route, there's a nice review [here](https://www.nature.com/articles/nmeth.4397), though I am a bit biased ;).",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 596817,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-08-11T10:54:34.633000",
      "content": "<p><a href=\"/npschafer\">@npschafer</a> Thanks for the info. \n1. We are at 0.925 and we still have a lot to do in the \"conventional ML\" area, i.e. without using any Recursion software (I don't know about the 0.99 but I estimate it's possible to get &gt;0.95 without it).\n2. I estimate that by the end of the competition a few teams will get to a Public LB of 1.0.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 596967,
          "author_name": "Nicholas Schafer",
          "author_url": "",
          "post_date": "2019-08-11T15:38:11.113000",
          "content": "<p>Thanks for sharing a bit about your top-performing solution. I had been wondering whether anyone near the top had gotten there with \"conventional ML\" and without using bespoke cell profiling techniques. Evidently, it is possible! This probably shows my ignorance re: the power of computer vision methods more than anything else, but I think that it also speaks to the amount and quality of the data provided. Nice work!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 596902,
      "author_name": "dohlee",
      "author_url": "",
      "post_date": "2019-08-11T13:31:05.410000",
      "content": "<p>Thanks for this invaluable information! I immediately tried CellProfiler but unfortunately found that the software will take a huge amount of time to process all the images we have. (Doing illumination correction step only for the images from U2OS-01 experiment showed an estimated time of 14 hours with 4 cores... maybe it'll take a month to process all the images) But I'm still thinking that this image-classification competition can be translated into tabular data-based classification competition with conventional cell profiling approach and extremely curious about the performance of the classifier exploiting single-cell based features. Again, thank you for such an interesting resource!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 596963,
          "author_name": "Nicholas Schafer",
          "author_url": "",
          "post_date": "2019-08-11T15:34:39.320000",
          "content": "<p>I was thinking the same thing re: converting this from a computer vision problem into a tabular data problem. They claim that the CellProfiler software can be run efficiently in distributed computing environments, which, as you point out, will seemingly be necessary when you are working with data sets of this size (or larger). With all of the information that Recursion provided, I'm wondering what their rationale was for not mentioning the CellProfiler software. Perhaps they were hoping that people would come up with good solutions that are largely orthogonal to traditional cell profiling techniques. And based on <a href=\"/yuval6967\">@yuval6967</a>'s comments, it seems that people have been able to do so!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 596651,
      "author_name": "wjsheng",
      "author_url": "",
      "post_date": "2019-08-11T03:43:49.227000",
      "content": "<p>Yes, it certainly puzzles me how teams who surpassed the 0.9 score keep improving while the rests are fairly stagnant in our positions. </p>\n\n<p>I wonder would it then be appropriate to perform 4 classification tasks instead of 1, given that we only account for the <code>sirna</code> when we do predictions for our test set now, as shown by nosound in his kernel.</p>\n\n<p>Another thing is that most people (including me) seem to not have found the given controls useful even though chimael mentioned that this is a \"relative\" image classification task.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 596761,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-11T08:29:22.013000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 596969,
          "author_name": "Nicholas Schafer",
          "author_url": "",
          "post_date": "2019-08-11T15:41:23.867000",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> see <a href=\"/yuval6967\">@yuval6967</a>'s comment above. Apparently, it is not necessary to apply traditional cell profiling techniques in order to get a good solution. That being said, I don't think that CellProfiler's usefulness has yet been ruled out. And yes, I think that there is potential gain to be had in leveraging all of the information available to us including the fact that the siRNAs always come in groups of 277. Good luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 597659,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "2019-08-12T16:13:31.023000",
          "content": "<p>It is really amazing to see a LB at 0.99. However, I noticed that the test data are from the same cell lines as the training data. Thus a classsifier specific for each cell line may be the way to go. </p>\n\n<p>If it is really the winner's solution, then the models will be useless on other cell lines. I wonder if this competition design was made on purpose.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 597672,
          "author_name": "wjsheng",
          "author_url": "",
          "post_date": "2019-08-12T16:34:04.607000",
          "content": "<p>hi <a href=\"/chimael\">@chimael</a>, thanks for the input. That being said, is it correct that if we build cell line specific classifiers, \"controls\" would not be as helpful as you suggested earlier. Since it's cell line specific, we do not need to frame this as a \"relative\" image classification task.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 598121,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "2019-08-13T06:51:28.930000",
          "content": "<p>I think your are correct. A cell specific classifier don't need controls to interpret the pictures of treatments. \nI was first hoping that this competition would show that we can produce a model that can recognize the effects of an siRNA across cell lines. Now I think that this competition will produce models that can recognize different phenotypes of a cell line. I may have misunderstood the intention of this competition. The function of specific siRNA may not be the main focus of the competition host.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 598435,
          "author_name": "wjsheng",
          "author_url": "",
          "post_date": "2019-08-13T15:15:25.547000",
          "content": "<p>Thank you, chimael. I agree with the way you understand the host's motivation in hosting the competition. It might just be unfortunate that the setup is not as perfect as expected - pointed out by Berton, \"Recursion stores the 1,108 treatment siRNAs in source plates containing 277 siRNA each, and transfers siRNA from these source plates to the destination plates used in an experiment by plate.\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 597038,
      "author_name": "Scott Warchal",
      "author_url": "",
      "post_date": "2019-08-11T17:40:15.887000",
      "content": "<p>Just so proper people are attributed correctly, CellProfiler was/is developed by Anne Carpenter's group at the Broad Institute rather than by Recursion.</p>\n\n<p>To add to the CellProfiler vs. CNN discussion, believe it or not morphological cell profiling is a pretty active field and I encourage you to have a look at some papers, but currently it's found that CNN's are out-performing more classical segmentation approaches (CellProfiler) in classification tasks of small molecules. <a href=\"https://www.biorxiv.org/content/10.1101/161422v1\">Here's</a> one such example.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 597095,
          "author_name": "Nicholas Schafer",
          "author_url": "",
          "post_date": "2019-08-11T19:55:16.167000",
          "content": "<p>Good catch, <a href=\"/swarchal\">@swarchal</a>. Thanks for the correction. I incorrectly attributed the software because the first Github repository that I ran into was actually Recursion's fork of the original repository and, also, there are Recursion co-authors on the paper that I read. Now it makes more sense why Recursion would not have felt the need to mention CellProfiler in the background information for this competition.</p>\n\n<p>Are you under the impression that a pipeline that takes the features that CellProfiler produces and combines them with other machine learning techniques on the resulting tabulated data would not be competitive? If so, then the basic premise of my post seems to be incorrect, and it is probably not worth people's time to try to get CellProfiler to run on all of the data provided by the competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 597107,
          "author_name": "Scott Warchal",
          "author_url": "",
          "post_date": "2019-08-11T20:43:09.693000",
          "content": "<p>I don't want to say it wouldn't be competitive, most of the published research comparing these methods at classifying cell morphology has focused on a single dataset (BBBC021). It would be great if a CellProfiler or similar software approach did well, as the features extracted are much more interpretable for biologists, e.g which features are changing in response to a particular siRNA?</p>\n\n<p>If anyone decides to go down the CellProfiler route, there's a nice review <a href=\"https://www.nature.com/articles/nmeth.4397\">here</a>, though I am a bit biased ;).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "596613": "Hello, fellow cell classifiers - \n\nI am currently participating in my first round of Kaggle competitions. I was interested in this one because of my background in biophysics. I actually interviewed at Recursion a few years ago and have been following their progress with interest ever since. I can see how the subject of this competition is an important problem that they need to solve, and I'm hoping that all of you will be able to provide them with a solution that works well. \n\nWhat I found most striking while reading through the public discussion boards and kernels was how unfamiliar most of the terminology and techniques being used would be to a classically trained biologist, or to a classically trained scientist of almost any type, for that matter. This is clearly a computer vision problem, but even among computer vision problems it does not square with the typical characterization of computer vision problems as being amongst those problems that are easy for human's to solve but difficult for computers. Just imagine trying to tell apart more than 1,000 different phenotypes simply by looking under the microscope! Even desperate graduate students couldn't be pushed that far.\n\nAnd that brings me to the next most striking aspect of the competition: the top of the leaderboard. In a multiclass classification problem with more than 1,000 classes, the top of the leaderboard currently shows a 99% accuracy. Having read the discussion board, I realize that this is not, strictly speaking, a 1,000+ class classification problem. Because of a data leak, the problem can be broken down into 4 subproblems of 277 classes (be sure to check this out on the discussion boards if you haven't seen it already). Nonetheless, 99% accuracy seems quite high to me. Perhaps the private leaderboard will eventually show that some of those solutions are overfitting to the public leaderboard data, but with only 15 submissions to date, it seems possible that SharksWithLasers really does have a good solution to Recurion's problem. So, how did they do it?\n\nSorry, but I don't know how they did it. However, could they have been helped by using... Recursion's own software? A little bit of digging turned up the following potentially helpful resources: \nhttps://github.com/recursionpharma/CellProfiler\nhttps://github.com/CellProfiler/CellProfiler-Analyst\nhttps://cellprofiler.org/\nhttps://www.nature.com/articles/nprot.2016.105\n\nUnsurprisingly, Recursion has already been working on this problem, namely being able to distinguish biological differences from batch differences. The quotes below are from the Nature Methods paper linked above. Properly configured, the Cell Profiler software can be used to create \"~1,500 morphological features ... to produce a rich profile that is suitable for the detection of subtle phenotypes\". They also note in their description of the protocol (experimental and computational) that \"feature extraction and data analysis take an additional 1-2 weeks.\" The Cell Painting assay described in that paper is essentially the same as was used to produce the data in this competition. The features produced by Cell Profiler include \"staining intensities, textural patterns, size, and shape of the labeled cellular structures, as well as correlations between stains across channels, and adjacency relationships between cells and among intracellular structures.\" In that paper they also make direct reference to the types of problems that this competition is focused on. \"A perennial concern with assay development is that any technical sources of variation can have an impact on all the wells and/or plates such that any biological signals are overwhelmed by systematic noise introduced by sample preparation.\" \"The image feature extraction workflow for Cell Painting is divided into three asks,, each of which is performed by a CellProfiler pipeline: (i) illumination, (ii) quality control (QC), and (iii) morphological feature extraction.\" \"Last, data analysis across separately performed experiments is likely to be complicated, requiring proper control over potentially substantial effects of differences in cell seeding, growth, and other *batch-related* or other systematic artifacts. Protocols for such cases have not yet been developed.\"\n\nWell, what are you waiting for? Go forth and develop :-) A potentially useful CellProfiler pipeline can be found here:\nhttps://github.com/gigascience/paper-bray2017\n\nAnd others here:\nhttps://cellprofiler.org/examples/published_pipelines\n\nBest,\nNick",
    "596817": "@npschafer Thanks for the info. \n1. We are at 0.925 and we still have a lot to do in the \"conventional ML\" area, i.e. without using any Recursion software (I don't know about the 0.99 but I estimate it's possible to get &gt;0.95 without it).\n2. I estimate that by the end of the competition a few teams will get to a Public LB of 1.0.",
    "596902": "Thanks for this invaluable information! I immediately tried CellProfiler but unfortunately found that the software will take a huge amount of time to process all the images we have. (Doing illumination correction step only for the images from U2OS-01 experiment showed an estimated time of 14 hours with 4 cores... maybe it'll take a month to process all the images) But I'm still thinking that this image-classification competition can be translated into tabular data-based classification competition with conventional cell profiling approach and extremely curious about the performance of the classifier exploiting single-cell based features. Again, thank you for such an interesting resource!",
    "596651": "Yes, it certainly puzzles me how teams who surpassed the 0.9 score keep improving while the rests are fairly stagnant in our positions. \n\nI wonder would it then be appropriate to perform 4 classification tasks instead of 1, given that we only account for the `sirna` when we do predictions for our test set now, as shown by nosound in his kernel.\n\nAnother thing is that most people (including me) seem to not have found the given controls useful even though chimael mentioned that this is a \"relative\" image classification task.",
    "597038": "Just so proper people are attributed correctly, CellProfiler was/is developed by Anne Carpenter's group at the Broad Institute rather than by Recursion.\n\nTo add to the CellProfiler vs. CNN discussion, believe it or not morphological cell profiling is a pretty active field and I encourage you to have a look at some papers, but currently it's found that CNN's are out-performing more classical segmentation approaches (CellProfiler) in classification tasks of small molecules. [Here's](https://www.biorxiv.org/content/10.1101/161422v1) one such example."
  }
}