{
  "id": 441845,
  "title": "Clarifying the aim of the challenge",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/441845",
  "author_name": "",
  "post_date": "2023-09-20T11:19:23.310064Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Dear Organising Team,</p>\n<p>Thanks for putting this challenge together. We strongly believe that it could make for an interesting comparison of approaches to quantifying perturbation effects in single cells. We would like to gain a better understanding of the precise question this challenge aims to answer.</p>\n<p>In our eyes, the scientific question the experiment tries to answer is to estimate the effect of a drug on a cell type, more precisely with unobserved combinations of (drug, cell type). Mathematically, this is defined as the conditional average treatment effect (CATE):</p>\n<p><code>CATE(drug, cell type) = E[Gene Expression | do(Drug=drug), Cell Type=cell type] - E[Gene Expression | do(Drug=control), Cell Type=cell type]</code></p>\n<p>Note that if you’re not familiar with causal inference notation that this definition uses do-calculus, described for example here: <a href=\"http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf\" target=\"_blank\">http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf</a>, chapter 3.1 and 3.2. In brief, this operator indicates an idealized intervention.</p>\n<p>Moving from scientific question to challenge, the organisers have modelled gene expression as a linear function of cell type, drug and covariates (library, plate) and estimated this CATE as the coefficient of a drug in this linear model. In doing so, they have provided an answer to the scientific question. Now, the challenge’s task for the participants is to predict the -log10(p-value) of this estimate. Therefore, the goal of the challenge as it is currently set up is not the estimation of the CATE.</p>\n<p>The Overview states: “We’re interested in understanding the problem of generalizing perturbation responses across cell lines.” We would be very interested in helping to answer that question (CATE). However, we do not currently understand the purpose of predicting p-values output by a statistical model. We would like to acknowledge that we are not experts in single-cell analysis and may well misunderstand portions of the challenge, any clarifications would therefore be greatly appreciated.</p>",
  "messages": [
    {
      "id": "2448000",
      "postDate": "09/20/2023 11:19:23",
      "content": "<p>Dear Organising Team,</p>\n<p>Thanks for putting this challenge together. We strongly believe that it could make for an interesting comparison of approaches to quantifying perturbation effects in single cells. We would like to gain a better understanding of the precise question this challenge aims to answer.</p>\n<p>In our eyes, the scientific question the experiment tries to answer is to estimate the effect of a drug on a cell type, more precisely with unobserved combinations of (drug, cell type). Mathematically, this is defined as the conditional average treatment effect (CATE):</p>\n<p><code>CATE(drug, cell type) = E[Gene Expression | do(Drug=drug), Cell Type=cell type] - E[Gene Expression | do(Drug=control), Cell Type=cell type]</code></p>\n<p>Note that if you’re not familiar with causal inference notation that this definition uses do-calculus, described for example here: <a href=\"http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf\" target=\"_blank\">http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf</a>, chapter 3.1 and 3.2. In brief, this operator indicates an idealized intervention.</p>\n<p>Moving from scientific question to challenge, the organisers have modelled gene expression as a linear function of cell type, drug and covariates (library, plate) and estimated this CATE as the coefficient of a drug in this linear model. In doing so, they have provided an answer to the scientific question. Now, the challenge’s task for the participants is to predict the -log10(p-value) of this estimate. Therefore, the goal of the challenge as it is currently set up is not the estimation of the CATE.</p>\n<p>The Overview states: “We’re interested in understanding the problem of generalizing perturbation responses across cell lines.” We would be very interested in helping to answer that question (CATE). However, we do not currently understand the purpose of predicting p-values output by a statistical model. We would like to acknowledge that we are not experts in single-cell analysis and may well misunderstand portions of the challenge, any clarifications would therefore be greatly appreciated.</p>",
      "rawMarkdown": "Dear Organising Team,\n\nThanks for putting this challenge together. We strongly believe that it could make for an interesting comparison of approaches to quantifying perturbation effects in single cells. We would like to gain a better understanding of the precise question this challenge aims to answer.\n\nIn our eyes, the scientific question the experiment tries to answer is to estimate the effect of a drug on a cell type, more precisely with unobserved combinations of (drug, cell type). Mathematically, this is defined as the conditional average treatment effect (CATE):\n\n`CATE(drug, cell type) = E[Gene Expression | do(Drug=drug), Cell Type=cell type] - E[Gene Expression | do(Drug=control), Cell Type=cell type]`\n\nNote that if you’re not familiar with causal inference notation that this definition uses do-calculus, described for example here: http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf, chapter 3.1 and 3.2. In brief, this operator indicates an idealized intervention.\n\nMoving from scientific question to challenge, the organisers have modelled gene expression as a linear function of cell type, drug and covariates (library, plate) and estimated this CATE as the coefficient of a drug in this linear model. In doing so, they have provided an answer to the scientific question. Now, the challenge’s task for the participants is to predict the -log10(p-value) of this estimate. Therefore, the goal of the challenge as it is currently set up is not the estimation of the CATE.\n\nThe Overview states: “We’re interested in understanding the problem of generalizing perturbation responses across cell lines.” We would be very interested in helping to answer that question (CATE). However, we do not currently understand the purpose of predicting p-values output by a statistical model. We would like to acknowledge that we are not experts in single-cell analysis and may well misunderstand portions of the challenge, any clarifications would therefore be greatly appreciated.",
      "votes": null
    },
    {
      "id": "2448372",
      "postDate": "09/20/2023 14:50:28",
      "content": "<p>This seems to be a duplicate of <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706</a></p>\n<p>I'm starting to think it would be helpful to run an info session. Please upvote if you'd be interested.</p>",
      "rawMarkdown": "This seems to be a duplicate of https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\n\nI'm starting to think it would be helpful to run an info session. Please upvote if you'd be interested.",
      "votes": null
    },
    {
      "id": "2448393",
      "postDate": "09/20/2023 14:59:04",
      "content": "<p>An info session would be great, thank you.</p>",
      "rawMarkdown": "An info session would be great, thank you.",
      "votes": null
    },
    {
      "id": "2456611",
      "postDate": "09/26/2023 10:14:08",
      "content": "<p>Any update on this info session?</p>",
      "rawMarkdown": "Any update on this info session?",
      "votes": null
    },
    {
      "id": "2456698",
      "postDate": "09/26/2023 11:16:14",
      "content": "<p>I think the aim of the competition is clear (predict the transformed p-value produced by the limma model). This can be taken entirely out of the biological context and experimental setup: \"there is this random matrix whose elements are presumably dependent. use this dependence to predict the missing values\".</p>\n<p>However, I agree that the aim of the competition does not align with the goal of the study. Everything has been overly simplified, to the point, I believe, of making the experimental or even biological aspects irrelevant. If the only information used for prediction is the name of the compound, we lose all benefits of modelling more complex interactions.</p>\n<p>It's mostly reduced to a missing value imputation problem. Still, we can leverage some structure in the data to, presumably, help with the imputation.</p>\n<p>That being said, that's why they have the \"judges prize\": to motivate us to create models using the biological data. There, the discrepancy between study goals and competition goals will make things a bit awkward in my opinion: \"if the goal were to predict CATE (say), we would do X, but since it is not, we cannot justify doing X\".</p>",
      "rawMarkdown": "I think the aim of the competition is clear (predict the transformed p-value produced by the limma model). This can be taken entirely out of the biological context and experimental setup: \"there is this random matrix whose elements are presumably dependent. use this dependence to predict the missing values\".\n\nHowever, I agree that the aim of the competition does not align with the goal of the study. Everything has been overly simplified, to the point, I believe, of making the experimental or even biological aspects irrelevant. If the only information used for prediction is the name of the compound, we lose all benefits of modelling more complex interactions.\n\nIt's mostly reduced to a missing value imputation problem. Still, we can leverage some structure in the data to, presumably, help with the imputation.\n\nThat being said, that's why they have the \"judges prize\": to motivate us to create models using the biological data. There, the discrepancy between study goals and competition goals will make things a bit awkward in my opinion: \"if the goal were to predict CATE (say), we would do X, but since it is not, we cannot justify doing X\".",
      "votes": null
    },
    {
      "id": "2458309",
      "postDate": "09/27/2023 13:50:43",
      "content": "<p>Hi all, please see details here: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493</a></p>",
      "rawMarkdown": "Hi all, please see details here: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2448372,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "09/20/2023 14:50:28",
      "content": "<p>This seems to be a duplicate of <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706</a></p>\n<p>I'm starting to think it would be helpful to run an info session. Please upvote if you'd be interested.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2448393,
          "author_name": "jeskowagner",
          "author_url": "",
          "post_date": "09/20/2023 14:59:04",
          "content": "<p>An info session would be great, thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2456611,
          "author_name": "olabayle",
          "author_url": "",
          "post_date": "09/26/2023 10:14:08",
          "content": "<p>Any update on this info session?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2458309,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "09/27/2023 13:50:43",
          "content": "<p>Hi all, please see details here: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2456698,
      "author_name": "wiwh404",
      "author_url": "",
      "post_date": "09/26/2023 11:16:14",
      "content": "<p>I think the aim of the competition is clear (predict the transformed p-value produced by the limma model). This can be taken entirely out of the biological context and experimental setup: \"there is this random matrix whose elements are presumably dependent. use this dependence to predict the missing values\".</p>\n<p>However, I agree that the aim of the competition does not align with the goal of the study. Everything has been overly simplified, to the point, I believe, of making the experimental or even biological aspects irrelevant. If the only information used for prediction is the name of the compound, we lose all benefits of modelling more complex interactions.</p>\n<p>It's mostly reduced to a missing value imputation problem. Still, we can leverage some structure in the data to, presumably, help with the imputation.</p>\n<p>That being said, that's why they have the \"judges prize\": to motivate us to create models using the biological data. There, the discrepancy between study goals and competition goals will make things a bit awkward in my opinion: \"if the goal were to predict CATE (say), we would do X, but since it is not, we cannot justify doing X\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2448000": "Dear Organising Team,\n\nThanks for putting this challenge together. We strongly believe that it could make for an interesting comparison of approaches to quantifying perturbation effects in single cells. We would like to gain a better understanding of the precise question this challenge aims to answer.\n\nIn our eyes, the scientific question the experiment tries to answer is to estimate the effect of a drug on a cell type, more precisely with unobserved combinations of (drug, cell type). Mathematically, this is defined as the conditional average treatment effect (CATE):\n\n`CATE(drug, cell type) = E[Gene Expression | do(Drug=drug), Cell Type=cell type] - E[Gene Expression | do(Drug=control), Cell Type=cell type]`\n\nNote that if you’re not familiar with causal inference notation that this definition uses do-calculus, described for example here: http://bayes.cs.ucla.edu/PRIMER/primer-ch3.pdf, chapter 3.1 and 3.2. In brief, this operator indicates an idealized intervention.\n\nMoving from scientific question to challenge, the organisers have modelled gene expression as a linear function of cell type, drug and covariates (library, plate) and estimated this CATE as the coefficient of a drug in this linear model. In doing so, they have provided an answer to the scientific question. Now, the challenge’s task for the participants is to predict the -log10(p-value) of this estimate. Therefore, the goal of the challenge as it is currently set up is not the estimation of the CATE.\n\nThe Overview states: “We’re interested in understanding the problem of generalizing perturbation responses across cell lines.” We would be very interested in helping to answer that question (CATE). However, we do not currently understand the purpose of predicting p-values output by a statistical model. We would like to acknowledge that we are not experts in single-cell analysis and may well misunderstand portions of the challenge, any clarifications would therefore be greatly appreciated.",
    "2448372": "This seems to be a duplicate of https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440706\n\nI'm starting to think it would be helpful to run an info session. Please upvote if you'd be interested.",
    "2448393": "An info session would be great, thank you.",
    "2456611": "Any update on this info session?",
    "2456698": "I think the aim of the competition is clear (predict the transformed p-value produced by the limma model). This can be taken entirely out of the biological context and experimental setup: \"there is this random matrix whose elements are presumably dependent. use this dependence to predict the missing values\".\n\nHowever, I agree that the aim of the competition does not align with the goal of the study. Everything has been overly simplified, to the point, I believe, of making the experimental or even biological aspects irrelevant. If the only information used for prediction is the name of the compound, we lose all benefits of modelling more complex interactions.\n\nIt's mostly reduced to a missing value imputation problem. Still, we can leverage some structure in the data to, presumably, help with the imputation.\n\nThat being said, that's why they have the \"judges prize\": to motivate us to create models using the biological data. There, the discrepancy between study goals and competition goals will make things a bit awkward in my opinion: \"if the goal were to predict CATE (say), we would do X, but since it is not, we cannot justify doing X\".",
    "2458309": "Hi all, please see details here: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493"
  },
  "source": "meta"
}