{
  "id": 99171,
  "title": "Experimental Biology Tips",
  "url": "/competitions/recursion-cellular-image-classification/discussion/99171",
  "author_name": "",
  "post_date": "2019-07-09T09:15:47.761255800Z",
  "votes": 116,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I have a relevant background in experimental molecular biology (Ph.D. with 15y as a researcher). I am very interested in this challenge, but since I know that I will not have the time to participate, I thought I could help you all with some knowledge on small interference RNA (siRNA).\nsiRNA are small regulatory RNAs that are used as regulators of protein production from genes. As you may guess, the regulation of protein levels in cells is critical to cell biology and altered regulation is frequently identified in diseases such as cancer. </p>\n\n<h3>Biology</h3>\n\n<p>siRNAs are artificial small RNAs that are experimentally introduced in cells to alter some specific protein levels. Interestingly, since the discovery of this phenomenon, some endogenous small RNA that assume a natural regulatory role in the cell was discovered. These small RNAs are called micro RNA (miRNA). One difference is that miRNA are regulating many protein levels at once contrary to the siRNA that is designed to be very specific to a protein of interest. It is important to know this because siRNA use the same molecular pathways as miRNA and we can frequently detect some 'off-target' effects of siRNA that was designed to one specific protein. This can sometimes complicate the interpretation of knock-down experiments since the observed phenotype may be due to the down-regulation of another protein than the targeted one.</p>\n\n<h3>Knock-down efficiency</h3>\n\n<p>Another consideration to keep in mind is that siRNA can downregulate protein expression but rarely abrogate its expression. Some other experimental procedures are used to completely knock-out a gene product (CRISPR/Cas9 or traditional transgene recombination). Usually, efficient siRNA can down-regulate protein production by 80%, but this may vary depending on the cell type and physiological context. Thus in this challenge, please note that positive controls are siRNA with proven phenotypical effect on cells however treatment siRNA may not all have a distinguishable phenotype on cell pictures.</p>\n\n<h3>Experimental bias</h3>\n\n<p>siRNA can be introduced chemically (transfection) into cells or biologically (transduction via engineered retrovirus). The efficiency can be quite different for the two methods. The transduction is better due to the selection of infected cells. The transfection is quicker but may result in poor transfection rates depending on the cell type. We currently don't know the type of siRNA delivery chosen here. But if the transfection was chosen, we might expect a variability between replicates.</p>\n\n<h3>Analysis</h3>\n\n<p>Usually, after siRNA knock-down of a protein of interest, researchers may asses qualitatively or quantitatively the phenotypic effect on cells. In this setting, we don't know the targeted proteins, thus it is unlikely that we can come up with a possible effect on the cell based on the function of the protein. Instead, we are tasked with the evaluation of all relevant features we may derive from the pictures and the positive controls. \nPhenotypic alteration is observed relative to negative control. Unfortunately, the authors have chosen untreated cells as a negative control, which is not a true negative control. A better alternative is a siRNA that does not have a target so that we can exclude the siRNA delivery method from the phenotypic analysis. \nThe method of analysis by an experimented cell biologist would be to observe a handful of individual cells in a treated well versus selected cells in the negative control. Obviously, you would like to select cells representative of the whole well but it is very subjective and error-prone. Then you would compare the intensity and shapes of the different markers (6 channels) in the two conditions. Quantitative evaluation is the best for statistical analysis, however, qualitative analysis is most common.</p>\n\n<h3>Features</h3>\n\n<p>siRNA may affect many biological processes. Consequences of some effects may be observed on the selected markers of the cells. The chosen markers are:\n* nuclei (blue), \n* endoplasmic reticulum (green), \n* actin (red), \n* nucleoli (cyan), \n* mitochondria (magenta), \n* and Golgi apparatus (yellow).\nThese are reported on the legend of figure 6 on <a href=\"https://www.rxrx.ai/\">https://www.rxrx.ai/</a>.\nNuclei are the compartment where genomic DNA is stored. Nucleoli are sites in the nucleus where ribosomal (the molecular structure for protein assembly) RNA are produced. Endoplasmic Reticulum is the compartment around the nucleus were proteins start their synthesis. Protein synthesis continues (folding) in the Golgi compartment. Next is mitochondria, the site of energy production. And finally, the Actin structure under the cell membrane.</p>\n\n<p>Here are some of the features I could think of:\n* The relative number of cells in the field to monitor the effect on proliferation.\n* The number of \"kissing\" nucleus that could be indicative of the number of cells in one particular phase of the mitosis (a division of cell).\n* The size of the nucleus versus the size of the cytoplasm.\n* The number of Nucleoli per nucleus\n* The shapes of nucleus and cytoplasm (roundness)\n* The number and size of colonies (a cluster of cells)\n* The number of mitochondria per cytoplasm\nextraction \nHowever, rather than designing features, one may prefer automatic extraction via ConvNet.</p>\n\n<h3>Relative Truth</h3>\n\n<p>Regular image classification tasks use supervised training with labels. These labels are true in the absolute. Meaning that the value of the label is not changed by its context. In this challenge, however, the truthness of a label is relative. Indeed, the quantification of treatment is only valid compared to a control state. Thus the picture of a siRNA treated cell can only be interpreted relative to a control picture. The task is to interpret pairs of pictures for the classification of the treated cells. Interestingly, we have 31 control conditions for each siRNA treated replicates. These control conditions can be interpreted as the extreme phenotypic pictures for the positive controls (upper bounds of the range) and the base phenotype for the negative control (lower bound of the range). Each treatment picture can be defined in this 31-dimensional space. From there a regular classifier may use this 31-dimensional vector for training and prediction.</p>",
  "messages": [
    {
      "id": "571186",
      "postDate": "07/09/2019 09:15:47",
      "content": "<p>I have a relevant background in experimental molecular biology (Ph.D. with 15y as a researcher). I am very interested in this challenge, but since I know that I will not have the time to participate, I thought I could help you all with some knowledge on small interference RNA (siRNA).\nsiRNA are small regulatory RNAs that are used as regulators of protein production from genes. As you may guess, the regulation of protein levels in cells is critical to cell biology and altered regulation is frequently identified in diseases such as cancer. </p>\n\n<h3>Biology</h3>\n\n<p>siRNAs are artificial small RNAs that are experimentally introduced in cells to alter some specific protein levels. Interestingly, since the discovery of this phenomenon, some endogenous small RNA that assume a natural regulatory role in the cell was discovered. These small RNAs are called micro RNA (miRNA). One difference is that miRNA are regulating many protein levels at once contrary to the siRNA that is designed to be very specific to a protein of interest. It is important to know this because siRNA use the same molecular pathways as miRNA and we can frequently detect some 'off-target' effects of siRNA that was designed to one specific protein. This can sometimes complicate the interpretation of knock-down experiments since the observed phenotype may be due to the down-regulation of another protein than the targeted one.</p>\n\n<h3>Knock-down efficiency</h3>\n\n<p>Another consideration to keep in mind is that siRNA can downregulate protein expression but rarely abrogate its expression. Some other experimental procedures are used to completely knock-out a gene product (CRISPR/Cas9 or traditional transgene recombination). Usually, efficient siRNA can down-regulate protein production by 80%, but this may vary depending on the cell type and physiological context. Thus in this challenge, please note that positive controls are siRNA with proven phenotypical effect on cells however treatment siRNA may not all have a distinguishable phenotype on cell pictures.</p>\n\n<h3>Experimental bias</h3>\n\n<p>siRNA can be introduced chemically (transfection) into cells or biologically (transduction via engineered retrovirus). The efficiency can be quite different for the two methods. The transduction is better due to the selection of infected cells. The transfection is quicker but may result in poor transfection rates depending on the cell type. We currently don't know the type of siRNA delivery chosen here. But if the transfection was chosen, we might expect a variability between replicates.</p>\n\n<h3>Analysis</h3>\n\n<p>Usually, after siRNA knock-down of a protein of interest, researchers may asses qualitatively or quantitatively the phenotypic effect on cells. In this setting, we don't know the targeted proteins, thus it is unlikely that we can come up with a possible effect on the cell based on the function of the protein. Instead, we are tasked with the evaluation of all relevant features we may derive from the pictures and the positive controls. \nPhenotypic alteration is observed relative to negative control. Unfortunately, the authors have chosen untreated cells as a negative control, which is not a true negative control. A better alternative is a siRNA that does not have a target so that we can exclude the siRNA delivery method from the phenotypic analysis. \nThe method of analysis by an experimented cell biologist would be to observe a handful of individual cells in a treated well versus selected cells in the negative control. Obviously, you would like to select cells representative of the whole well but it is very subjective and error-prone. Then you would compare the intensity and shapes of the different markers (6 channels) in the two conditions. Quantitative evaluation is the best for statistical analysis, however, qualitative analysis is most common.</p>\n\n<h3>Features</h3>\n\n<p>siRNA may affect many biological processes. Consequences of some effects may be observed on the selected markers of the cells. The chosen markers are:\n* nuclei (blue), \n* endoplasmic reticulum (green), \n* actin (red), \n* nucleoli (cyan), \n* mitochondria (magenta), \n* and Golgi apparatus (yellow).\nThese are reported on the legend of figure 6 on <a href=\"https://www.rxrx.ai/\">https://www.rxrx.ai/</a>.\nNuclei are the compartment where genomic DNA is stored. Nucleoli are sites in the nucleus where ribosomal (the molecular structure for protein assembly) RNA are produced. Endoplasmic Reticulum is the compartment around the nucleus were proteins start their synthesis. Protein synthesis continues (folding) in the Golgi compartment. Next is mitochondria, the site of energy production. And finally, the Actin structure under the cell membrane.</p>\n\n<p>Here are some of the features I could think of:\n* The relative number of cells in the field to monitor the effect on proliferation.\n* The number of \"kissing\" nucleus that could be indicative of the number of cells in one particular phase of the mitosis (a division of cell).\n* The size of the nucleus versus the size of the cytoplasm.\n* The number of Nucleoli per nucleus\n* The shapes of nucleus and cytoplasm (roundness)\n* The number and size of colonies (a cluster of cells)\n* The number of mitochondria per cytoplasm\nextraction \nHowever, rather than designing features, one may prefer automatic extraction via ConvNet.</p>\n\n<h3>Relative Truth</h3>\n\n<p>Regular image classification tasks use supervised training with labels. These labels are true in the absolute. Meaning that the value of the label is not changed by its context. In this challenge, however, the truthness of a label is relative. Indeed, the quantification of treatment is only valid compared to a control state. Thus the picture of a siRNA treated cell can only be interpreted relative to a control picture. The task is to interpret pairs of pictures for the classification of the treated cells. Interestingly, we have 31 control conditions for each siRNA treated replicates. These control conditions can be interpreted as the extreme phenotypic pictures for the positive controls (upper bounds of the range) and the base phenotype for the negative control (lower bound of the range). Each treatment picture can be defined in this 31-dimensional space. From there a regular classifier may use this 31-dimensional vector for training and prediction.</p>",
      "rawMarkdown": "I have a relevant background in experimental molecular biology (Ph.D. with 15y as a researcher). I am very interested in this challenge, but since I know that I will not have the time to participate, I thought I could help you all with some knowledge on small interference RNA (siRNA).\nsiRNA are small regulatory RNAs that are used as regulators of protein production from genes. As you may guess, the regulation of protein levels in cells is critical to cell biology and altered regulation is frequently identified in diseases such as cancer. \n\n###Biology\nsiRNAs are artificial small RNAs that are experimentally introduced in cells to alter some specific protein levels. Interestingly, since the discovery of this phenomenon, some endogenous small RNA that assume a natural regulatory role in the cell was discovered. These small RNAs are called micro RNA (miRNA). One difference is that miRNA are regulating many protein levels at once contrary to the siRNA that is designed to be very specific to a protein of interest. It is important to know this because siRNA use the same molecular pathways as miRNA and we can frequently detect some 'off-target' effects of siRNA that was designed to one specific protein. This can sometimes complicate the interpretation of knock-down experiments since the observed phenotype may be due to the down-regulation of another protein than the targeted one.\n\n### Knock-down efficiency\nAnother consideration to keep in mind is that siRNA can downregulate protein expression but rarely abrogate its expression. Some other experimental procedures are used to completely knock-out a gene product (CRISPR/Cas9 or traditional transgene recombination). Usually, efficient siRNA can down-regulate protein production by 80%, but this may vary depending on the cell type and physiological context. Thus in this challenge, please note that positive controls are siRNA with proven phenotypical effect on cells however treatment siRNA may not all have a distinguishable phenotype on cell pictures.\n\n### Experimental bias\nsiRNA can be introduced chemically (transfection) into cells or biologically (transduction via engineered retrovirus). The efficiency can be quite different for the two methods. The transduction is better due to the selection of infected cells. The transfection is quicker but may result in poor transfection rates depending on the cell type. We currently don't know the type of siRNA delivery chosen here. But if the transfection was chosen, we might expect a variability between replicates.\n\n### Analysis\nUsually, after siRNA knock-down of a protein of interest, researchers may asses qualitatively or quantitatively the phenotypic effect on cells. In this setting, we don't know the targeted proteins, thus it is unlikely that we can come up with a possible effect on the cell based on the function of the protein. Instead, we are tasked with the evaluation of all relevant features we may derive from the pictures and the positive controls. \nPhenotypic alteration is observed relative to negative control. Unfortunately, the authors have chosen untreated cells as a negative control, which is not a true negative control. A better alternative is a siRNA that does not have a target so that we can exclude the siRNA delivery method from the phenotypic analysis. \nThe method of analysis by an experimented cell biologist would be to observe a handful of individual cells in a treated well versus selected cells in the negative control. Obviously, you would like to select cells representative of the whole well but it is very subjective and error-prone. Then you would compare the intensity and shapes of the different markers (6 channels) in the two conditions. Quantitative evaluation is the best for statistical analysis, however, qualitative analysis is most common.\n\n### Features\nsiRNA may affect many biological processes. Consequences of some effects may be observed on the selected markers of the cells. The chosen markers are:\n* nuclei (blue), \n* endoplasmic reticulum (green), \n* actin (red), \n* nucleoli (cyan), \n* mitochondria (magenta), \n* and Golgi apparatus (yellow).\nThese are reported on the legend of figure 6 on https://www.rxrx.ai/.\nNuclei are the compartment where genomic DNA is stored. Nucleoli are sites in the nucleus where ribosomal (the molecular structure for protein assembly) RNA are produced. Endoplasmic Reticulum is the compartment around the nucleus were proteins start their synthesis. Protein synthesis continues (folding) in the Golgi compartment. Next is mitochondria, the site of energy production. And finally, the Actin structure under the cell membrane.\n\nHere are some of the features I could think of:\n* The relative number of cells in the field to monitor the effect on proliferation.\n* The number of \"kissing\" nucleus that could be indicative of the number of cells in one particular phase of the mitosis (a division of cell).\n* The size of the nucleus versus the size of the cytoplasm.\n* The number of Nucleoli per nucleus\n* The shapes of nucleus and cytoplasm (roundness)\n* The number and size of colonies (a cluster of cells)\n* The number of mitochondria per cytoplasm\nextraction \nHowever, rather than designing features, one may prefer automatic extraction via ConvNet.\n\n### Relative Truth\nRegular image classification tasks use supervised training with labels. These labels are true in the absolute. Meaning that the value of the label is not changed by its context. In this challenge, however, the truthness of a label is relative. Indeed, the quantification of treatment is only valid compared to a control state. Thus the picture of a siRNA treated cell can only be interpreted relative to a control picture. The task is to interpret pairs of pictures for the classification of the treated cells. Interestingly, we have 31 control conditions for each siRNA treated replicates. These control conditions can be interpreted as the extreme phenotypic pictures for the positive controls (upper bounds of the range) and the base phenotype for the negative control (lower bound of the range). Each treatment picture can be defined in this 31-dimensional space. From there a regular classifier may use this 31-dimensional vector for training and prediction.",
      "votes": null
    },
    {
      "id": "571310",
      "postDate": "07/09/2019 13:37:06",
      "content": "<p>Wow, many thanks to you for this terrific explanation, this is priceless. However the meaning of \"control\" is still not very clear to me. I've read rxrx.ai and what I understand currently is that \"positive control\" is some kind of reference \"state\" to which we should compare \"negative controls\" to make decision about applied siRNA? Also we can treat untreated well as another \"positive control\", I am right? Could you please correct me if I am wrong.</p>",
      "rawMarkdown": "Wow, many thanks to you for this terrific explanation, this is priceless. However the meaning of \"control\" is still not very clear to me. I've read rxrx.ai and what I understand currently is that \"positive control\" is some kind of reference \"state\" to which we should compare \"negative controls\" to make decision about applied siRNA? Also we can treat untreated well as another \"positive control\", I am right? Could you please correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "571481",
      "postDate": "07/09/2019 16:26:10",
      "content": "<p>Thank you. I am looking forward to read the solution you come up with.</p>\n\n<p>The controls of an experiment is a very important concept. The positive controls can be seen as the most extreme phenotype that an experimental siRNA might achieve. Thus indeed it is a reference, but not to compare the negative control, but rather to compare the experimental siRNA.</p>\n\n<p>The untreated cell is negative control. You definitely right when you say that the negative control should be used as another positive control. We make the distinction between negative and positive because they are supposed to be at both extremes of the phenotypic specter [0-1]. </p>",
      "rawMarkdown": "Thank you. I am looking forward to read the solution you come up with.\n\nThe controls of an experiment is a very important concept. The positive controls can be seen as the most extreme phenotype that an experimental siRNA might achieve. Thus indeed it is a reference, but not to compare the negative control, but rather to compare the experimental siRNA.\n\nThe untreated cell is negative control. You definitely right when you say that the negative control should be used as another positive control. We make the distinction between negative and positive because they are supposed to be at both extremes of the phenotypic specter [0-1].",
      "votes": null
    },
    {
      "id": "571619",
      "postDate": "07/09/2019 20:47:38",
      "content": "<p>Thank you very much for clarification. I've checked <code>train_controls.csv</code> and <code>test_controls.csv</code> and looks like there are only 18 unique control siRNA which are present in all experiments and plates, does it mean that i can't define each treatment picture in 31-dimensional space as you mentioned?</p>",
      "rawMarkdown": "Thank you very much for clarification. I've checked `train_controls.csv` and `test_controls.csv` and looks like there are only 18 unique control siRNA which are present in all experiments and plates, does it mean that i can't define each treatment picture in 31-dimensional space as you mentioned?",
      "votes": null
    },
    {
      "id": "571703",
      "postDate": "07/10/2019 02:08:32",
      "content": "<p>Most of the plate do have 30 pos. control and 1 neg. control. It is more reasonable to do so when designing experiment. However, some plates have only 28 or 29 pos. controls which might be due to failed siRNA treatment...\nThe experimental design grid is like this:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F247447%2Fba620572f2a41f2eb92deb8eb2ceb054%2FUntitled.png?generation=1562724360665563&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Most of the plate do have 30 pos. control and 1 neg. control. It is more reasonable to do so when designing experiment. However, some plates have only 28 or 29 pos. controls which might be due to failed siRNA treatment...\nThe experimental design grid is like this:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F247447%2Fba620572f2a41f2eb92deb8eb2ceb054%2FUntitled.png?generation=1562724360665563&amp;alt=media)",
      "votes": null
    },
    {
      "id": "571730",
      "postDate": "07/10/2019 03:31:26",
      "content": "<p>There is indeed 31 controls per experiment. However, I did not find experiments with 28 or 29 positive controls. Please see this kernel: <a href=\"https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp\">https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp</a></p>",
      "rawMarkdown": "There is indeed 31 controls per experiment. However, I did not find experiments with 28 or 29 positive controls. Please see this kernel: https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp",
      "votes": null
    },
    {
      "id": "571752",
      "postDate": "07/10/2019 04:08:58",
      "content": "<p>31 controls per experiment, Yes.\n31 controls per plate, No.</p>\n\n<p><code>\ndf = pd.read_csv('train_controls.csv')\ndf.groupby('experiment')['sirna'].nunique()\ndf.groupby(['experiment', 'plate'])['sirna'].nunique()\ndf.groupby(['experiment', 'plate', 'well_type'])['sirna'].nunique()\n</code></p>",
      "rawMarkdown": "31 controls per experiment, Yes.\n31 controls per plate, No.\n\n```\ndf = pd.read_csv('train_controls.csv')\ndf.groupby('experiment')['sirna'].nunique()\ndf.groupby(['experiment', 'plate'])['sirna'].nunique()\ndf.groupby(['experiment', 'plate', 'well_type'])['sirna'].nunique()\n```",
      "votes": null
    },
    {
      "id": "572088",
      "postDate": "07/10/2019 12:47:51",
      "content": "<p>Wow really thanks for your sharing.</p>",
      "rawMarkdown": "Wow really thanks for your sharing.",
      "votes": null
    },
    {
      "id": "574481",
      "postDate": "07/14/2019 00:02:47",
      "content": "<p>Thank you for writing this. This is exactly the kind information that I was looking for.</p>\n\n<p>So, the input to the ConvNet can be the image (6 channels) from the site for which we are querying and images from all 31 controls (positive and negative) for <em>that</em> experiment (sorted). That's 6 * (1 + 31) = 126 channels</p>\n\n<p>Does this make sense?</p>",
      "rawMarkdown": "Thank you for writing this. This is exactly the kind information that I was looking for.\n\nSo, the input to the ConvNet can be the image (6 channels) from the site for which we are querying and images from all 31 controls (positive and negative) for *that* experiment (sorted). That's 6 * (1 + 31) = 126 channels\n\nDoes this make sense?",
      "votes": null
    },
    {
      "id": "574579",
      "postDate": "07/14/2019 06:05:17",
      "content": "<p>I don't know if it's even possible to have 126 channel input. However, at some point your network will need to evaluate your current image compared to each control or maybe a selection of it.\nI was thinking of some embedding for picture. I didn't came across this kind of solution for pictures yet. Do you have any idea how to deal with this problem?</p>",
      "rawMarkdown": "I don't know if it's even possible to have 126 channel input. However, at some point your network will need to evaluate your current image compared to each control or maybe a selection of it.\nI was thinking of some embedding for picture. I didn't came across this kind of solution for pictures yet. Do you have any idea how to deal with this problem?",
      "votes": null
    },
    {
      "id": "574681",
      "postDate": "07/14/2019 10:06:32",
      "content": "<p>Really thank you for sharing this. I have a question. How does this cell classification task relate to drug discovery? </p>",
      "rawMarkdown": "Really thank you for sharing this. I have a question. How does this cell classification task relate to drug discovery?",
      "votes": null
    },
    {
      "id": "574699",
      "postDate": "07/14/2019 10:45:14",
      "content": "<p>That's a good question. Current method for drug discovery is to screen large number of chemicals on particular cell models. If a drug can revert a phenotype (such as cell proliferation for cancer cells) you can observe this on cell culture. To do this by high throughput you need to invent a quantitative test for each disease you want to screen. If you could come up with a cell imaging technique that's both quantitative and high throughput, you will greatly accelerate drug discovery.</p>",
      "rawMarkdown": "That's a good question. Current method for drug discovery is to screen large number of chemicals on particular cell models. If a drug can revert a phenotype (such as cell proliferation for cancer cells) you can observe this on cell culture. To do this by high throughput you need to invent a quantitative test for each disease you want to screen. If you could come up with a cell imaging technique that's both quantitative and high throughput, you will greatly accelerate drug discovery.",
      "votes": null
    },
    {
      "id": "574768",
      "postDate": "07/14/2019 13:13:35",
      "content": "<p>Thank you! It seems quite meaningful. So for now is this cell observation job done manually? </p>",
      "rawMarkdown": "Thank you! It seems quite meaningful. So for now is this cell observation job done manually?",
      "votes": null
    },
    {
      "id": "574823",
      "postDate": "07/14/2019 14:27:49",
      "content": "<p>Yes, to my knowledge, cell imaging analysis is most of the time qualitative and done manually by a trained cell biologist.</p>",
      "rawMarkdown": "Yes, to my knowledge, cell imaging analysis is most of the time qualitative and done manually by a trained cell biologist.",
      "votes": null
    },
    {
      "id": "577646",
      "postDate": "07/16/2019 22:10:24",
      "content": "<p>Huge thanks for this!</p>",
      "rawMarkdown": "Huge thanks for this!",
      "votes": null
    },
    {
      "id": "577652",
      "postDate": "07/16/2019 22:24:42",
      "content": "<p>126 channel input is possible, but it makes training very slow (it puts pressure on CPU for creating the batches). \nRe \"embedding for images\":\nThe closest thing to embedding that I can think of is extracting features from the images and using the features. But that's what a convNet does, anyway.</p>",
      "rawMarkdown": "126 channel input is possible, but it makes training very slow (it puts pressure on CPU for creating the batches). \nRe \"embedding for images\":\nThe closest thing to embedding that I can think of is extracting features from the images and using the features. But that's what a convNet does, anyway.",
      "votes": null
    },
    {
      "id": "584202",
      "postDate": "07/25/2019 14:48:19",
      "content": "<p>Does anyone know what types of noise typically come from Fluorescent microscopy? \nPossibly, the type of instrumentation that may be used so that I can get an idea of how to deal how to move further?</p>",
      "rawMarkdown": "Does anyone know what types of noise typically come from Fluorescent microscopy? \nPossibly, the type of instrumentation that may be used so that I can get an idea of how to deal how to move further?",
      "votes": null
    },
    {
      "id": "585571",
      "postDate": "07/27/2019 16:55:02",
      "content": "<p>According to the <a href=\"https://www.rxrx.ai/\">RxRx1 site</a>, each experiment has the same 30 control siRNAs on every plate. Does that mean that only 30 of the 1108 siRNAs are used as positive controls, or would a different batch number use a different 30?</p>",
      "rawMarkdown": "According to the [RxRx1 site](https://www.rxrx.ai/), each experiment has the same 30 control siRNAs on every plate. Does that mean that only 30 of the 1108 siRNAs are used as positive controls, or would a different batch number use a different 30?",
      "votes": null
    },
    {
      "id": "586426",
      "postDate": "07/29/2019 06:47:12",
      "content": "<p>For each 1108 siRNA you have 30 positive controls and one or more negative controls. Each time you run a replicate experiment (another batch) you need to include all the controls to account for changes in experimental conditions. With this design you are sure that the differences you observe when using an experimental siRNA, compared to a positive siRNA or negative siRNA, is only due to the biological response and not a variation during the experiment.</p>",
      "rawMarkdown": "For each 1108 siRNA you have 30 positive controls and one or more negative controls. Each time you run a replicate experiment (another batch) you need to include all the controls to account for changes in experimental conditions. With this design you are sure that the differences you observe when using an experimental siRNA, compared to a positive siRNA or negative siRNA, is only due to the biological response and not a variation during the experiment.",
      "votes": null
    },
    {
      "id": "586564",
      "postDate": "07/29/2019 10:21:48",
      "content": "<p>Thanks <a href=\"/chimael\">@chimael</a>.</p>",
      "rawMarkdown": "Thanks @chimael.",
      "votes": null
    },
    {
      "id": "592714",
      "postDate": "08/05/2019 18:03:59",
      "content": "<p><a href=\"/vshmyhlo\">@vshmyhlo</a> this thread should help you understand the goal of controls <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591</a> .</p>",
      "rawMarkdown": "vshmyhlo this thread should help you understand the goal of controls https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591 .",
      "votes": null
    },
    {
      "id": "823198",
      "postDate": "04/27/2020 13:39:09",
      "content": "<p>Great tips, thanks for sharing!</p>",
      "rawMarkdown": "Great tips, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1547793",
      "postDate": "10/17/2021 15:38:14",
      "content": "<p>Nice read. Thanks for sharing <a href=\"https://www.kaggle.com/chimael\" target=\"_blank\">@chimael</a> </p>",
      "rawMarkdown": "Nice read. Thanks for sharing @chimael",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1547793,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "10/17/2021 15:38:14",
      "content": "<p>Nice read. Thanks for sharing <a href=\"https://www.kaggle.com/chimael\" target=\"_blank\">@chimael</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 571310,
      "author_name": "vshmyhlo",
      "author_url": "",
      "post_date": "07/09/2019 13:37:06",
      "content": "<p>Wow, many thanks to you for this terrific explanation, this is priceless. However the meaning of \"control\" is still not very clear to me. I've read rxrx.ai and what I understand currently is that \"positive control\" is some kind of reference \"state\" to which we should compare \"negative controls\" to make decision about applied siRNA? Also we can treat untreated well as another \"positive control\", I am right? Could you please correct me if I am wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 571481,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/09/2019 16:26:10",
          "content": "<p>Thank you. I am looking forward to read the solution you come up with.</p>\n\n<p>The controls of an experiment is a very important concept. The positive controls can be seen as the most extreme phenotype that an experimental siRNA might achieve. Thus indeed it is a reference, but not to compare the negative control, but rather to compare the experimental siRNA.</p>\n\n<p>The untreated cell is negative control. You definitely right when you say that the negative control should be used as another positive control. We make the distinction between negative and positive because they are supposed to be at both extremes of the phenotypic specter [0-1]. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571619,
          "author_name": "vshmyhlo",
          "author_url": "",
          "post_date": "07/09/2019 20:47:38",
          "content": "<p>Thank you very much for clarification. I've checked <code>train_controls.csv</code> and <code>test_controls.csv</code> and looks like there are only 18 unique control siRNA which are present in all experiments and plates, does it mean that i can't define each treatment picture in 31-dimensional space as you mentioned?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571703,
          "author_name": "zongtseng",
          "author_url": "",
          "post_date": "07/10/2019 02:08:32",
          "content": "<p>Most of the plate do have 30 pos. control and 1 neg. control. It is more reasonable to do so when designing experiment. However, some plates have only 28 or 29 pos. controls which might be due to failed siRNA treatment...\nThe experimental design grid is like this:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F247447%2Fba620572f2a41f2eb92deb8eb2ceb054%2FUntitled.png?generation=1562724360665563&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571730,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/10/2019 03:31:26",
          "content": "<p>There is indeed 31 controls per experiment. However, I did not find experiments with 28 or 29 positive controls. Please see this kernel: <a href=\"https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp\">https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 571752,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "07/10/2019 04:08:58",
          "content": "<p>31 controls per experiment, Yes.\n31 controls per plate, No.</p>\n\n<p><code>\ndf = pd.read_csv('train_controls.csv')\ndf.groupby('experiment')['sirna'].nunique()\ndf.groupby(['experiment', 'plate'])['sirna'].nunique()\ndf.groupby(['experiment', 'plate', 'well_type'])['sirna'].nunique()\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592714,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/05/2019 18:03:59",
          "content": "<p><a href=\"/vshmyhlo\">@vshmyhlo</a> this thread should help you understand the goal of controls <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591</a> .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 572088,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "07/10/2019 12:47:51",
      "content": "<p>Wow really thanks for your sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 574481,
      "author_name": "rrezaii",
      "author_url": "",
      "post_date": "07/14/2019 00:02:47",
      "content": "<p>Thank you for writing this. This is exactly the kind information that I was looking for.</p>\n\n<p>So, the input to the ConvNet can be the image (6 channels) from the site for which we are querying and images from all 31 controls (positive and negative) for <em>that</em> experiment (sorted). That's 6 * (1 + 31) = 126 channels</p>\n\n<p>Does this make sense?</p>",
      "votes": null,
      "replies": [
        {
          "id": 574579,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/14/2019 06:05:17",
          "content": "<p>I don't know if it's even possible to have 126 channel input. However, at some point your network will need to evaluate your current image compared to each control or maybe a selection of it.\nI was thinking of some embedding for picture. I didn't came across this kind of solution for pictures yet. Do you have any idea how to deal with this problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577652,
          "author_name": "rrezaii",
          "author_url": "",
          "post_date": "07/16/2019 22:24:42",
          "content": "<p>126 channel input is possible, but it makes training very slow (it puts pressure on CPU for creating the batches). \nRe \"embedding for images\":\nThe closest thing to embedding that I can think of is extracting features from the images and using the features. But that's what a convNet does, anyway.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 574681,
      "author_name": "shentao",
      "author_url": "",
      "post_date": "07/14/2019 10:06:32",
      "content": "<p>Really thank you for sharing this. I have a question. How does this cell classification task relate to drug discovery? </p>",
      "votes": null,
      "replies": [
        {
          "id": 574699,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/14/2019 10:45:14",
          "content": "<p>That's a good question. Current method for drug discovery is to screen large number of chemicals on particular cell models. If a drug can revert a phenotype (such as cell proliferation for cancer cells) you can observe this on cell culture. To do this by high throughput you need to invent a quantitative test for each disease you want to screen. If you could come up with a cell imaging technique that's both quantitative and high throughput, you will greatly accelerate drug discovery.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 574768,
          "author_name": "shentao",
          "author_url": "",
          "post_date": "07/14/2019 13:13:35",
          "content": "<p>Thank you! It seems quite meaningful. So for now is this cell observation job done manually? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 574823,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/14/2019 14:27:49",
          "content": "<p>Yes, to my knowledge, cell imaging analysis is most of the time qualitative and done manually by a trained cell biologist.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 577646,
      "author_name": "michelml",
      "author_url": "",
      "post_date": "07/16/2019 22:10:24",
      "content": "<p>Huge thanks for this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 584202,
      "author_name": "mrfugu69",
      "author_url": "",
      "post_date": "07/25/2019 14:48:19",
      "content": "<p>Does anyone know what types of noise typically come from Fluorescent microscopy? \nPossibly, the type of instrumentation that may be used so that I can get an idea of how to deal how to move further?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 585571,
      "author_name": "kenkrige",
      "author_url": "",
      "post_date": "07/27/2019 16:55:02",
      "content": "<p>According to the <a href=\"https://www.rxrx.ai/\">RxRx1 site</a>, each experiment has the same 30 control siRNAs on every plate. Does that mean that only 30 of the 1108 siRNAs are used as positive controls, or would a different batch number use a different 30?</p>",
      "votes": null,
      "replies": [
        {
          "id": 586426,
          "author_name": "chimael",
          "author_url": "",
          "post_date": "07/29/2019 06:47:12",
          "content": "<p>For each 1108 siRNA you have 30 positive controls and one or more negative controls. Each time you run a replicate experiment (another batch) you need to include all the controls to account for changes in experimental conditions. With this design you are sure that the differences you observe when using an experimental siRNA, compared to a positive siRNA or negative siRNA, is only due to the biological response and not a variation during the experiment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 586564,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "07/29/2019 10:21:48",
          "content": "<p>Thanks <a href=\"/chimael\">@chimael</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823198,
      "author_name": "duboviy",
      "author_url": "",
      "post_date": "04/27/2020 13:39:09",
      "content": "<p>Great tips, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "571186": "I have a relevant background in experimental molecular biology (Ph.D. with 15y as a researcher). I am very interested in this challenge, but since I know that I will not have the time to participate, I thought I could help you all with some knowledge on small interference RNA (siRNA).\nsiRNA are small regulatory RNAs that are used as regulators of protein production from genes. As you may guess, the regulation of protein levels in cells is critical to cell biology and altered regulation is frequently identified in diseases such as cancer. \n\n###Biology\nsiRNAs are artificial small RNAs that are experimentally introduced in cells to alter some specific protein levels. Interestingly, since the discovery of this phenomenon, some endogenous small RNA that assume a natural regulatory role in the cell was discovered. These small RNAs are called micro RNA (miRNA). One difference is that miRNA are regulating many protein levels at once contrary to the siRNA that is designed to be very specific to a protein of interest. It is important to know this because siRNA use the same molecular pathways as miRNA and we can frequently detect some 'off-target' effects of siRNA that was designed to one specific protein. This can sometimes complicate the interpretation of knock-down experiments since the observed phenotype may be due to the down-regulation of another protein than the targeted one.\n\n### Knock-down efficiency\nAnother consideration to keep in mind is that siRNA can downregulate protein expression but rarely abrogate its expression. Some other experimental procedures are used to completely knock-out a gene product (CRISPR/Cas9 or traditional transgene recombination). Usually, efficient siRNA can down-regulate protein production by 80%, but this may vary depending on the cell type and physiological context. Thus in this challenge, please note that positive controls are siRNA with proven phenotypical effect on cells however treatment siRNA may not all have a distinguishable phenotype on cell pictures.\n\n### Experimental bias\nsiRNA can be introduced chemically (transfection) into cells or biologically (transduction via engineered retrovirus). The efficiency can be quite different for the two methods. The transduction is better due to the selection of infected cells. The transfection is quicker but may result in poor transfection rates depending on the cell type. We currently don't know the type of siRNA delivery chosen here. But if the transfection was chosen, we might expect a variability between replicates.\n\n### Analysis\nUsually, after siRNA knock-down of a protein of interest, researchers may asses qualitatively or quantitatively the phenotypic effect on cells. In this setting, we don't know the targeted proteins, thus it is unlikely that we can come up with a possible effect on the cell based on the function of the protein. Instead, we are tasked with the evaluation of all relevant features we may derive from the pictures and the positive controls. \nPhenotypic alteration is observed relative to negative control. Unfortunately, the authors have chosen untreated cells as a negative control, which is not a true negative control. A better alternative is a siRNA that does not have a target so that we can exclude the siRNA delivery method from the phenotypic analysis. \nThe method of analysis by an experimented cell biologist would be to observe a handful of individual cells in a treated well versus selected cells in the negative control. Obviously, you would like to select cells representative of the whole well but it is very subjective and error-prone. Then you would compare the intensity and shapes of the different markers (6 channels) in the two conditions. Quantitative evaluation is the best for statistical analysis, however, qualitative analysis is most common.\n\n### Features\nsiRNA may affect many biological processes. Consequences of some effects may be observed on the selected markers of the cells. The chosen markers are:\n* nuclei (blue), \n* endoplasmic reticulum (green), \n* actin (red), \n* nucleoli (cyan), \n* mitochondria (magenta), \n* and Golgi apparatus (yellow).\nThese are reported on the legend of figure 6 on https://www.rxrx.ai/.\nNuclei are the compartment where genomic DNA is stored. Nucleoli are sites in the nucleus where ribosomal (the molecular structure for protein assembly) RNA are produced. Endoplasmic Reticulum is the compartment around the nucleus were proteins start their synthesis. Protein synthesis continues (folding) in the Golgi compartment. Next is mitochondria, the site of energy production. And finally, the Actin structure under the cell membrane.\n\nHere are some of the features I could think of:\n* The relative number of cells in the field to monitor the effect on proliferation.\n* The number of \"kissing\" nucleus that could be indicative of the number of cells in one particular phase of the mitosis (a division of cell).\n* The size of the nucleus versus the size of the cytoplasm.\n* The number of Nucleoli per nucleus\n* The shapes of nucleus and cytoplasm (roundness)\n* The number and size of colonies (a cluster of cells)\n* The number of mitochondria per cytoplasm\nextraction \nHowever, rather than designing features, one may prefer automatic extraction via ConvNet.\n\n### Relative Truth\nRegular image classification tasks use supervised training with labels. These labels are true in the absolute. Meaning that the value of the label is not changed by its context. In this challenge, however, the truthness of a label is relative. Indeed, the quantification of treatment is only valid compared to a control state. Thus the picture of a siRNA treated cell can only be interpreted relative to a control picture. The task is to interpret pairs of pictures for the classification of the treated cells. Interestingly, we have 31 control conditions for each siRNA treated replicates. These control conditions can be interpreted as the extreme phenotypic pictures for the positive controls (upper bounds of the range) and the base phenotype for the negative control (lower bound of the range). Each treatment picture can be defined in this 31-dimensional space. From there a regular classifier may use this 31-dimensional vector for training and prediction.",
    "571310": "Wow, many thanks to you for this terrific explanation, this is priceless. However the meaning of \"control\" is still not very clear to me. I've read rxrx.ai and what I understand currently is that \"positive control\" is some kind of reference \"state\" to which we should compare \"negative controls\" to make decision about applied siRNA? Also we can treat untreated well as another \"positive control\", I am right? Could you please correct me if I am wrong.",
    "571481": "Thank you. I am looking forward to read the solution you come up with.\n\nThe controls of an experiment is a very important concept. The positive controls can be seen as the most extreme phenotype that an experimental siRNA might achieve. Thus indeed it is a reference, but not to compare the negative control, but rather to compare the experimental siRNA.\n\nThe untreated cell is negative control. You definitely right when you say that the negative control should be used as another positive control. We make the distinction between negative and positive because they are supposed to be at both extremes of the phenotypic specter [0-1].",
    "571619": "Thank you very much for clarification. I've checked `train_controls.csv` and `test_controls.csv` and looks like there are only 18 unique control siRNA which are present in all experiments and plates, does it mean that i can't define each treatment picture in 31-dimensional space as you mentioned?",
    "571703": "Most of the plate do have 30 pos. control and 1 neg. control. It is more reasonable to do so when designing experiment. However, some plates have only 28 or 29 pos. controls which might be due to failed siRNA treatment...\nThe experimental design grid is like this:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F247447%2Fba620572f2a41f2eb92deb8eb2ceb054%2FUntitled.png?generation=1562724360665563&amp;alt=media)",
    "571730": "There is indeed 31 controls per experiment. However, I did not find experiments with 28 or 29 positive controls. Please see this kernel: https://www.kaggle.com/chimael/08-recursion-positive-controls-per-exp",
    "571752": "31 controls per experiment, Yes.\n31 controls per plate, No.\n\n```\ndf = pd.read_csv('train_controls.csv')\ndf.groupby('experiment')['sirna'].nunique()\ndf.groupby(['experiment', 'plate'])['sirna'].nunique()\ndf.groupby(['experiment', 'plate', 'well_type'])['sirna'].nunique()\n```",
    "572088": "Wow really thanks for your sharing.",
    "574481": "Thank you for writing this. This is exactly the kind information that I was looking for.\n\nSo, the input to the ConvNet can be the image (6 channels) from the site for which we are querying and images from all 31 controls (positive and negative) for *that* experiment (sorted). That's 6 * (1 + 31) = 126 channels\n\nDoes this make sense?",
    "574579": "I don't know if it's even possible to have 126 channel input. However, at some point your network will need to evaluate your current image compared to each control or maybe a selection of it.\nI was thinking of some embedding for picture. I didn't came across this kind of solution for pictures yet. Do you have any idea how to deal with this problem?",
    "574681": "Really thank you for sharing this. I have a question. How does this cell classification task relate to drug discovery?",
    "574699": "That's a good question. Current method for drug discovery is to screen large number of chemicals on particular cell models. If a drug can revert a phenotype (such as cell proliferation for cancer cells) you can observe this on cell culture. To do this by high throughput you need to invent a quantitative test for each disease you want to screen. If you could come up with a cell imaging technique that's both quantitative and high throughput, you will greatly accelerate drug discovery.",
    "574768": "Thank you! It seems quite meaningful. So for now is this cell observation job done manually?",
    "574823": "Yes, to my knowledge, cell imaging analysis is most of the time qualitative and done manually by a trained cell biologist.",
    "577646": "Huge thanks for this!",
    "577652": "126 channel input is possible, but it makes training very slow (it puts pressure on CPU for creating the batches). \nRe \"embedding for images\":\nThe closest thing to embedding that I can think of is extracting features from the images and using the features. But that's what a convNet does, anyway.",
    "584202": "Does anyone know what types of noise typically come from Fluorescent microscopy? \nPossibly, the type of instrumentation that may be used so that I can get an idea of how to deal how to move further?",
    "585571": "According to the [RxRx1 site](https://www.rxrx.ai/), each experiment has the same 30 control siRNAs on every plate. Does that mean that only 30 of the 1108 siRNAs are used as positive controls, or would a different batch number use a different 30?",
    "586426": "For each 1108 siRNA you have 30 positive controls and one or more negative controls. Each time you run a replicate experiment (another batch) you need to include all the controls to account for changes in experimental conditions. With this design you are sure that the differences you observe when using an experimental siRNA, compared to a positive siRNA or negative siRNA, is only due to the biological response and not a variation during the experiment.",
    "586564": "Thanks @chimael.",
    "592714": "vshmyhlo this thread should help you understand the goal of controls https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/101826#latest-591591 .",
    "823198": "Great tips, thanks for sharing!",
    "1547793": "Nice read. Thanks for sharing @chimael"
  },
  "source": "meta"
}