{
  "id": 108716,
  "title": "Why are controls important?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/108716",
  "author_name": "",
  "post_date": "2019-09-13T12:52:44.423919100Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>All kinds of hints show that effectively using control data seems to be essential to good LB . Still don't understand why are controls important: from what I understand, controls across experiments are different just by the batch effects, but don't two samples with same siRNAs have the same relationship and serve the same purpose?</p>",
  "messages": [
    {
      "id": "625780",
      "postDate": "09/13/2019 12:52:44",
      "content": "<p>All kinds of hints show that effectively using control data seems to be essential to good LB . Still don't understand why are controls important: from what I understand, controls across experiments are different just by the batch effects, but don't two samples with same siRNAs have the same relationship and serve the same purpose?</p>",
      "rawMarkdown": "All kinds of hints show that effectively using control data seems to be essential to good LB . Still don't understand why are controls important: from what I understand, controls across experiments are different just by the batch effects, but don't two samples with same siRNAs have the same relationship and serve the same purpose?",
      "votes": null
    },
    {
      "id": "625784",
      "postDate": "09/13/2019 12:57:10",
      "content": "<p>The only idea I have about using controls currently is to use L2 between extracted features to regularize the model. Have not tried this out yet; wonder if anyone will be kind enough to point at how to use the controls</p>",
      "rawMarkdown": "The only idea I have about using controls currently is to use L2 between extracted features to regularize the model. Have not tried this out yet; wonder if anyone will be kind enough to point at how to use the controls",
      "votes": null
    },
    {
      "id": "626219",
      "postDate": "09/14/2019 03:15:45",
      "content": "<p>I am a beginner in machine learning and do not know why 'control' is important in this case.\nBut in biological experiments, the control is important to compare the effects of specific siRNA.</p>\n\n<p>RNA interference is a phenomenon in which the expression of a gene is suppressed by targeting a specific gene sequence.</p>\n\n<p>A positive control is a siRNA that greatly reduces the amount of transcription of a specific gene (for example, 90%).\nThe siRNA not targeting a known gene is treated as a negative control.</p>\n\n<p>I hope it helps something for you.</p>",
      "rawMarkdown": "I am a beginner in machine learning and do not know why 'control' is important in this case.\nBut in biological experiments, the control is important to compare the effects of specific siRNA.\n\nRNA interference is a phenomenon in which the expression of a gene is suppressed by targeting a specific gene sequence.\n\nA positive control is a siRNA that greatly reduces the amount of transcription of a specific gene (for example, 90%).\nThe siRNA not targeting a known gene is treated as a negative control.\n\nI hope it helps something for you.",
      "votes": null
    },
    {
      "id": "626946",
      "postDate": "09/15/2019 06:32:03",
      "content": "<p>Did your L2 experiment provide good results?</p>",
      "rawMarkdown": "Did your L2 experiment provide good results?",
      "votes": null
    },
    {
      "id": "627382",
      "postDate": "09/15/2019 21:19:36",
      "content": "<p>Hey there, </p>\n\n<p>I am new to Kaggle and deep learning and do not know how controls are important for training in this context. However, a common sense would be to use them for normalization of the data. <strong>I might be wrong here but this is my reasoning:</strong></p>\n\n<p>Negative control represents state 0 and positive control 1, they define [lower, upper] boundary for each plate. siRNA effect should be somewhere in the middle. Normalization should be done separately for each channel and each plate (to account for technical noise). Unfortunately, I am not really sure how to implement this. Otherwise, I would try it. </p>\n\n<p>I am also sharing some ideas that might be useful for someone.</p>\n\n<p>Why to use - and + controls? \nIf you don't have controls you cannot be sure that the effect you observe is due to the siRNA. If you observe no effect in siRNA treatment, it could be due to the siRNA itself or some external factor. This is why you need + control. It should always produce the effect. If it did not then there was smth wrong with the whole experiment. - controls are used for a similar reason. If you see the effect of siRNA then you need to check the - control, as it should have no effect (untreated). If it does, then again some external factors influenced the results.  </p>\n\n<p>Why to use different cell lines? \nDifferent siRNA will have different effects on various cell types. For drug development, both the safety and effectiveness is important. Lets say we have two cell lines: one is cancer line (test the effectiveness of the siRNA against cancer) and another one is healthy (test the safety of the siRNA). The simplest measure of the drug effectiveness is a cell density. A good siRNA lead would kill cancer cells and have no effect on healthy cells, and so on. </p>\n\n<p>ps google cell lines</p>",
      "rawMarkdown": "Hey there, \n\nI am new to Kaggle and deep learning and do not know how controls are important for training in this context. However, a common sense would be to use them for normalization of the data. **I might be wrong here but this is my reasoning:**\n\nNegative control represents state 0 and positive control 1, they define [lower, upper] boundary for each plate. siRNA effect should be somewhere in the middle. Normalization should be done separately for each channel and each plate (to account for technical noise). Unfortunately, I am not really sure how to implement this. Otherwise, I would try it. \n\nI am also sharing some ideas that might be useful for someone.\n \nWhy to use - and + controls? \nIf you don't have controls you cannot be sure that the effect you observe is due to the siRNA. If you observe no effect in siRNA treatment, it could be due to the siRNA itself or some external factor. This is why you need + control. It should always produce the effect. If it did not then there was smth wrong with the whole experiment. - controls are used for a similar reason. If you see the effect of siRNA then you need to check the - control, as it should have no effect (untreated). If it does, then again some external factors influenced the results.  \n\nWhy to use different cell lines? \nDifferent siRNA will have different effects on various cell types. For drug development, both the safety and effectiveness is important. Lets say we have two cell lines: one is cancer line (test the effectiveness of the siRNA against cancer) and another one is healthy (test the safety of the siRNA). The simplest measure of the drug effectiveness is a cell density. A good siRNA lead would kill cancer cells and have no effect on healthy cells, and so on. \n\nps google cell lines",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 625784,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "09/13/2019 12:57:10",
      "content": "<p>The only idea I have about using controls currently is to use L2 between extracted features to regularize the model. Have not tried this out yet; wonder if anyone will be kind enough to point at how to use the controls</p>",
      "votes": null,
      "replies": [
        {
          "id": 626946,
          "author_name": "abyaadrafid",
          "author_url": "",
          "post_date": "09/15/2019 06:32:03",
          "content": "<p>Did your L2 experiment provide good results?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 626219,
      "author_name": "kfurudate",
      "author_url": "",
      "post_date": "09/14/2019 03:15:45",
      "content": "<p>I am a beginner in machine learning and do not know why 'control' is important in this case.\nBut in biological experiments, the control is important to compare the effects of specific siRNA.</p>\n\n<p>RNA interference is a phenomenon in which the expression of a gene is suppressed by targeting a specific gene sequence.</p>\n\n<p>A positive control is a siRNA that greatly reduces the amount of transcription of a specific gene (for example, 90%).\nThe siRNA not targeting a known gene is treated as a negative control.</p>\n\n<p>I hope it helps something for you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 627382,
      "author_name": "aspak1",
      "author_url": "",
      "post_date": "09/15/2019 21:19:36",
      "content": "<p>Hey there, </p>\n\n<p>I am new to Kaggle and deep learning and do not know how controls are important for training in this context. However, a common sense would be to use them for normalization of the data. <strong>I might be wrong here but this is my reasoning:</strong></p>\n\n<p>Negative control represents state 0 and positive control 1, they define [lower, upper] boundary for each plate. siRNA effect should be somewhere in the middle. Normalization should be done separately for each channel and each plate (to account for technical noise). Unfortunately, I am not really sure how to implement this. Otherwise, I would try it. </p>\n\n<p>I am also sharing some ideas that might be useful for someone.</p>\n\n<p>Why to use - and + controls? \nIf you don't have controls you cannot be sure that the effect you observe is due to the siRNA. If you observe no effect in siRNA treatment, it could be due to the siRNA itself or some external factor. This is why you need + control. It should always produce the effect. If it did not then there was smth wrong with the whole experiment. - controls are used for a similar reason. If you see the effect of siRNA then you need to check the - control, as it should have no effect (untreated). If it does, then again some external factors influenced the results.  </p>\n\n<p>Why to use different cell lines? \nDifferent siRNA will have different effects on various cell types. For drug development, both the safety and effectiveness is important. Lets say we have two cell lines: one is cancer line (test the effectiveness of the siRNA against cancer) and another one is healthy (test the safety of the siRNA). The simplest measure of the drug effectiveness is a cell density. A good siRNA lead would kill cancer cells and have no effect on healthy cells, and so on. </p>\n\n<p>ps google cell lines</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "625780": "All kinds of hints show that effectively using control data seems to be essential to good LB . Still don't understand why are controls important: from what I understand, controls across experiments are different just by the batch effects, but don't two samples with same siRNAs have the same relationship and serve the same purpose?",
    "625784": "The only idea I have about using controls currently is to use L2 between extracted features to regularize the model. Have not tried this out yet; wonder if anyone will be kind enough to point at how to use the controls",
    "626219": "I am a beginner in machine learning and do not know why 'control' is important in this case.\nBut in biological experiments, the control is important to compare the effects of specific siRNA.\n\nRNA interference is a phenomenon in which the expression of a gene is suppressed by targeting a specific gene sequence.\n\nA positive control is a siRNA that greatly reduces the amount of transcription of a specific gene (for example, 90%).\nThe siRNA not targeting a known gene is treated as a negative control.\n\nI hope it helps something for you.",
    "626946": "Did your L2 experiment provide good results?",
    "627382": "Hey there, \n\nI am new to Kaggle and deep learning and do not know how controls are important for training in this context. However, a common sense would be to use them for normalization of the data. **I might be wrong here but this is my reasoning:**\n\nNegative control represents state 0 and positive control 1, they define [lower, upper] boundary for each plate. siRNA effect should be somewhere in the middle. Normalization should be done separately for each channel and each plate (to account for technical noise). Unfortunately, I am not really sure how to implement this. Otherwise, I would try it. \n\nI am also sharing some ideas that might be useful for someone.\n \nWhy to use - and + controls? \nIf you don't have controls you cannot be sure that the effect you observe is due to the siRNA. If you observe no effect in siRNA treatment, it could be due to the siRNA itself or some external factor. This is why you need + control. It should always produce the effect. If it did not then there was smth wrong with the whole experiment. - controls are used for a similar reason. If you see the effect of siRNA then you need to check the - control, as it should have no effect (untreated). If it does, then again some external factors influenced the results.  \n\nWhy to use different cell lines? \nDifferent siRNA will have different effects on various cell types. For drug development, both the safety and effectiveness is important. Lets say we have two cell lines: one is cancer line (test the effectiveness of the siRNA against cancer) and another one is healthy (test the safety of the siRNA). The simplest measure of the drug effectiveness is a cell density. A good siRNA lead would kill cancer cells and have no effect on healthy cells, and so on. \n\nps google cell lines"
  },
  "source": "meta"
}