{
  "id": 440706,
  "title": "The values to predict",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/440706",
  "author_name": "",
  "post_date": "2023-09-16T00:00:16.349708500Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I probably missed something obvious. It is not super clear what we need to predict. I saw a post saying that the target is to predict -log10(p-value) from limma. Is that correct? Or should we predict log fold change or something more quantitative about the effect size (instead of significance)?</p>",
  "messages": [
    {
      "id": "2441011",
      "postDate": "09/16/2023 00:00:16",
      "content": "<p>I probably missed something obvious. It is not super clear what we need to predict. I saw a post saying that the target is to predict -log10(p-value) from limma. Is that correct? Or should we predict log fold change or something more quantitative about the effect size (instead of significance)?</p>",
      "rawMarkdown": "I probably missed something obvious. It is not super clear what we need to predict. I saw a post saying that the target is to predict -log10(p-value) from limma. Is that correct? Or should we predict log fold change or something more quantitative about the effect size (instead of significance)?",
      "votes": null
    },
    {
      "id": "2442350",
      "postDate": "09/16/2023 23:42:14",
      "content": "<p>Per the Data page:</p>\n<p>\"The input to your model will be a tuple of cell_type and sm_name and the output of your model will be predicted signed -log10(p-values) for all 18211 genes.\"</p>",
      "rawMarkdown": "Per the Data page:\n\n\"The input to your model will be a tuple of cell_type and sm_name and the output of your model will be predicted signed -log10(p-values) for all 18211 genes.\"",
      "votes": null
    },
    {
      "id": "2443122",
      "postDate": "09/17/2023 14:35:27",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2445322",
      "postDate": "09/18/2023 18:27:22",
      "content": "<p>Hi Scott, thanks for your clarification. Before finding this thread I was confused about the task as well.</p>\n<p>The data pages also states that</p>\n<blockquote>\n  <p>Your task is to predict differential expression values for Myeloid and B cells for a majority of compounds.</p>\n</blockquote>\n<p>which I find to be misleading, if we are indeed to predict the p-value of the limma model.</p>",
      "rawMarkdown": "Hi Scott, thanks for your clarification. Before finding this thread I was confused about the task as well.\n\nThe data pages also states that\n> Your task is to predict differential expression values for Myeloid and B cells for a majority of compounds.\n\nwhich I find to be misleading, if we are indeed to predict the p-value of the limma model.",
      "votes": null
    },
    {
      "id": "2446765",
      "postDate": "09/19/2023 15:45:40",
      "content": "<p>Yes this is still highly confusing. Even if I understand the task, I don't understand what is the purpose of predicting the inference output of a linear model. What else than a linear model should we do?</p>",
      "rawMarkdown": "Yes this is still highly confusing. Even if I understand the task, I don't understand what is the purpose of predicting the inference output of a linear model. What else than a linear model should we do?",
      "votes": null
    },
    {
      "id": "2446834",
      "postDate": "09/19/2023 16:37:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/olabayle\" target=\"_blank\">@olabayle</a>, I think there's some confusion about the role of the differential expression analysis. </p>\n<p>Let's take a simple case. Say you have 6 raw gene expression samples, 3 treatment samples and 3 control samples. Let's also say there's only one gene. If you wanted to calculate the effect size of the perturbation treatment, you could calculate a T-score between the 3 treatment and 3 control samples. Another approach would be to fit a linear model, where the explanatory variable is the treatment condition and the independent variable is the gene expression. You can calculate a p-value associated with this coefficient as an alternative to a T score. The p-value is a single score for the perturbation effect output from a linear model (note we went from 6 raw measurements at the beginning to one p-value)</p>\n<p>In the simple case, using a linear model might seem overkill, but our case is much, much more complicated. We have tens of thousands of genes, each of which has a different baseline expression. We also have additional experimental covariates we want to regress out of our signal. We also have some Bayesian priors about the nature of differential expression. All of these we want to capture in a model that output a value that estimates the effect of a perturbation in a cell type.</p>\n<p>For the competition, we hide some of the data. You don't get to see either the raw measurements or the estimate of the effect size we learn from that raw data for some (cell type, compound) combinations. Your job is a tensor-completion task to fill in the blanks.</p>\n<p>I hope this makes sense. If you want to learn more, I suggest reading:</p>\n<ul>\n<li><a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8</a></li>\n<li><a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">https://www.nature.com/articles/s41467-021-25960-2</a></li>\n</ul>",
      "rawMarkdown": "Hi @olabayle, I think there's some confusion about the role of the differential expression analysis. \n\nLet's take a simple case. Say you have 6 raw gene expression samples, 3 treatment samples and 3 control samples. Let's also say there's only one gene. If you wanted to calculate the effect size of the perturbation treatment, you could calculate a T-score between the 3 treatment and 3 control samples. Another approach would be to fit a linear model, where the explanatory variable is the treatment condition and the independent variable is the gene expression. You can calculate a p-value associated with this coefficient as an alternative to a T score. The p-value is a single score for the perturbation effect output from a linear model (note we went from 6 raw measurements at the beginning to one p-value)\n\nIn the simple case, using a linear model might seem overkill, but our case is much, much more complicated. We have tens of thousands of genes, each of which has a different baseline expression. We also have additional experimental covariates we want to regress out of our signal. We also have some Bayesian priors about the nature of differential expression. All of these we want to capture in a model that output a value that estimates the effect of a perturbation in a cell type.\n\nFor the competition, we hide some of the data. You don't get to see either the raw measurements or the estimate of the effect size we learn from that raw data for some (cell type, compound) combinations. Your job is a tensor-completion task to fill in the blanks.\n\nI hope this makes sense. If you want to learn more, I suggest reading:\n\n* https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8\n* https://www.nature.com/articles/s41467-021-25960-2",
      "votes": null
    },
    {
      "id": "2446859",
      "postDate": "09/19/2023 17:02:59",
      "content": "<p>I think the problem is that I thought the challenge was actually about estimating this effect size whereas you have already done it using a linear model. I don't want to leak to much information but I understood the goal of the challenge as the estimation of the conditional treatment effect on gene expression (GE), i.e : CATE(drug, cell type) = E[GE|do(drug=d), cell_type] - E[GE|do(drug=control), cell_type]. This does not refer to any modeling assumption and leaves participants free to build their own molecular/cell-type specific representations. The only thing that needs to be provided is the aggregate gene expression per cell type (at the well level). And technical bias can be adjusted for using more advanced methods than linear regression. Because we are tasked to predict some p-values output by a linear model it seems the goal is to retro engineer the inputs to the limma model.</p>",
      "rawMarkdown": "I think the problem is that I thought the challenge was actually about estimating this effect size whereas you have already done it using a linear model. I don't want to leak to much information but I understood the goal of the challenge as the estimation of the conditional treatment effect on gene expression (GE), i.e : CATE(drug, cell type) = E[GE|do(drug=d), cell_type] - E[GE|do(drug=control), cell_type]. This does not refer to any modeling assumption and leaves participants free to build their own molecular/cell-type specific representations. The only thing that needs to be provided is the aggregate gene expression per cell type (at the well level). And technical bias can be adjusted for using more advanced methods than linear regression. Because we are tasked to predict some p-values output by a linear model it seems the goal is to retro engineer the inputs to the limma model.",
      "votes": null
    },
    {
      "id": "2448368",
      "postDate": "09/20/2023 14:48:10",
      "content": "<p>I still think there's confusion here about the input and output of the competition.</p>\n<p>In inference, all you get is two strings for the <code>cell_type</code> and <code>perturbation</code> conditions. Your model then outputs predicts DE values for 18,000 genes. </p>\n<p>You don't get the raw expression values for this combination of <code>cell_type</code> and <code>perturbation</code> for compounds in the public / private LB. There are multiple valid ways to predict the compound effect here. It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE. Another is just to train a matrix completion model on the training data and directly output DE values. The choice is yours.</p>",
      "rawMarkdown": "I still think there's confusion here about the input and output of the competition.\n\nIn inference, all you get is two strings for the `cell_type` and `perturbation` conditions. Your model then outputs predicts DE values for 18,000 genes. \n\nYou don't get the raw expression values for this combination of `cell_type` and `perturbation` for compounds in the public / private LB. There are multiple valid ways to predict the compound effect here. It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE. Another is just to train a matrix completion model on the training data and directly output DE values. The choice is yours.",
      "votes": null
    },
    {
      "id": "2448460",
      "postDate": "09/20/2023 15:46:39",
      "content": "<p>See also: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372</a></p>",
      "rawMarkdown": "See also: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372",
      "votes": null
    },
    {
      "id": "2457095",
      "postDate": "09/26/2023 16:02:15",
      "content": "<p>If the targets are the -log10(p-value), then when I try to reverse it by doing -10^x, then how come the number ranges are not between 0-1? Am I misunderstanding something?</p>\n<p>Edit: Nvm, just saw that there's some sign(LFC) multiplied to it.</p>",
      "rawMarkdown": "If the targets are the -log10(p-value), then when I try to reverse it by doing -10^x, then how come the number ranges are not between 0-1? Am I misunderstanding something?\n\nEdit: Nvm, just saw that there's some sign(LFC) multiplied to it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2442350,
      "author_name": "scottgigante",
      "author_url": "",
      "post_date": "09/16/2023 23:42:14",
      "content": "<p>Per the Data page:</p>\n<p>\"The input to your model will be a tuple of cell_type and sm_name and the output of your model will be predicted signed -log10(p-values) for all 18211 genes.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 2443122,
          "author_name": "yupenghe",
          "author_url": "",
          "post_date": "09/17/2023 14:35:27",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2445322,
          "author_name": "stefanstark",
          "author_url": "",
          "post_date": "09/18/2023 18:27:22",
          "content": "<p>Hi Scott, thanks for your clarification. Before finding this thread I was confused about the task as well.</p>\n<p>The data pages also states that</p>\n<blockquote>\n  <p>Your task is to predict differential expression values for Myeloid and B cells for a majority of compounds.</p>\n</blockquote>\n<p>which I find to be misleading, if we are indeed to predict the p-value of the limma model.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2446765,
              "author_name": "olabayle",
              "author_url": "",
              "post_date": "09/19/2023 15:45:40",
              "content": "<p>Yes this is still highly confusing. Even if I understand the task, I don't understand what is the purpose of predicting the inference output of a linear model. What else than a linear model should we do?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2446834,
                  "author_name": "danielburkhardt",
                  "author_url": "",
                  "post_date": "09/19/2023 16:37:35",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/olabayle\" target=\"_blank\">@olabayle</a>, I think there's some confusion about the role of the differential expression analysis. </p>\n<p>Let's take a simple case. Say you have 6 raw gene expression samples, 3 treatment samples and 3 control samples. Let's also say there's only one gene. If you wanted to calculate the effect size of the perturbation treatment, you could calculate a T-score between the 3 treatment and 3 control samples. Another approach would be to fit a linear model, where the explanatory variable is the treatment condition and the independent variable is the gene expression. You can calculate a p-value associated with this coefficient as an alternative to a T score. The p-value is a single score for the perturbation effect output from a linear model (note we went from 6 raw measurements at the beginning to one p-value)</p>\n<p>In the simple case, using a linear model might seem overkill, but our case is much, much more complicated. We have tens of thousands of genes, each of which has a different baseline expression. We also have additional experimental covariates we want to regress out of our signal. We also have some Bayesian priors about the nature of differential expression. All of these we want to capture in a model that output a value that estimates the effect of a perturbation in a cell type.</p>\n<p>For the competition, we hide some of the data. You don't get to see either the raw measurements or the estimate of the effect size we learn from that raw data for some (cell type, compound) combinations. Your job is a tensor-completion task to fill in the blanks.</p>\n<p>I hope this makes sense. If you want to learn more, I suggest reading:</p>\n<ul>\n<li><a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8</a></li>\n<li><a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">https://www.nature.com/articles/s41467-021-25960-2</a></li>\n</ul>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2446859,
                      "author_name": "olabayle",
                      "author_url": "",
                      "post_date": "09/19/2023 17:02:59",
                      "content": "<p>I think the problem is that I thought the challenge was actually about estimating this effect size whereas you have already done it using a linear model. I don't want to leak to much information but I understood the goal of the challenge as the estimation of the conditional treatment effect on gene expression (GE), i.e : CATE(drug, cell type) = E[GE|do(drug=d), cell_type] - E[GE|do(drug=control), cell_type]. This does not refer to any modeling assumption and leaves participants free to build their own molecular/cell-type specific representations. The only thing that needs to be provided is the aggregate gene expression per cell type (at the well level). And technical bias can be adjusted for using more advanced methods than linear regression. Because we are tasked to predict some p-values output by a linear model it seems the goal is to retro engineer the inputs to the limma model.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2448368,
                          "author_name": "danielburkhardt",
                          "author_url": "",
                          "post_date": "09/20/2023 14:48:10",
                          "content": "<p>I still think there's confusion here about the input and output of the competition.</p>\n<p>In inference, all you get is two strings for the <code>cell_type</code> and <code>perturbation</code> conditions. Your model then outputs predicts DE values for 18,000 genes. </p>\n<p>You don't get the raw expression values for this combination of <code>cell_type</code> and <code>perturbation</code> for compounds in the public / private LB. There are multiple valid ways to predict the compound effect here. It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE. Another is just to train a matrix completion model on the training data and directly output DE values. The choice is yours.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    },
                    {
                      "id": 2457095,
                      "author_name": "yuqizheng",
                      "author_url": "",
                      "post_date": "09/26/2023 16:02:15",
                      "content": "<p>If the targets are the -log10(p-value), then when I try to reverse it by doing -10^x, then how come the number ranges are not between 0-1? Am I misunderstanding something?</p>\n<p>Edit: Nvm, just saw that there's some sign(LFC) multiplied to it.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 2448460,
              "author_name": "jeskowagner",
              "author_url": "",
              "post_date": "09/20/2023 15:46:39",
              "content": "<p>See also: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2441011": "I probably missed something obvious. It is not super clear what we need to predict. I saw a post saying that the target is to predict -log10(p-value) from limma. Is that correct? Or should we predict log fold change or something more quantitative about the effect size (instead of significance)?",
    "2442350": "Per the Data page:\n\n\"The input to your model will be a tuple of cell_type and sm_name and the output of your model will be predicted signed -log10(p-values) for all 18211 genes.\"",
    "2443122": "Thank you!",
    "2445322": "Hi Scott, thanks for your clarification. Before finding this thread I was confused about the task as well.\n\nThe data pages also states that\n> Your task is to predict differential expression values for Myeloid and B cells for a majority of compounds.\n\nwhich I find to be misleading, if we are indeed to predict the p-value of the limma model.",
    "2446765": "Yes this is still highly confusing. Even if I understand the task, I don't understand what is the purpose of predicting the inference output of a linear model. What else than a linear model should we do?",
    "2446834": "Hi @olabayle, I think there's some confusion about the role of the differential expression analysis. \n\nLet's take a simple case. Say you have 6 raw gene expression samples, 3 treatment samples and 3 control samples. Let's also say there's only one gene. If you wanted to calculate the effect size of the perturbation treatment, you could calculate a T-score between the 3 treatment and 3 control samples. Another approach would be to fit a linear model, where the explanatory variable is the treatment condition and the independent variable is the gene expression. You can calculate a p-value associated with this coefficient as an alternative to a T score. The p-value is a single score for the perturbation effect output from a linear model (note we went from 6 raw measurements at the beginning to one p-value)\n\nIn the simple case, using a linear model might seem overkill, but our case is much, much more complicated. We have tens of thousands of genes, each of which has a different baseline expression. We also have additional experimental covariates we want to regress out of our signal. We also have some Bayesian priors about the nature of differential expression. All of these we want to capture in a model that output a value that estimates the effect of a perturbation in a cell type.\n\nFor the competition, we hide some of the data. You don't get to see either the raw measurements or the estimate of the effect size we learn from that raw data for some (cell type, compound) combinations. Your job is a tensor-completion task to fill in the blanks.\n\nI hope this makes sense. If you want to learn more, I suggest reading:\n\n* https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8\n* https://www.nature.com/articles/s41467-021-25960-2",
    "2446859": "I think the problem is that I thought the challenge was actually about estimating this effect size whereas you have already done it using a linear model. I don't want to leak to much information but I understood the goal of the challenge as the estimation of the conditional treatment effect on gene expression (GE), i.e : CATE(drug, cell type) = E[GE|do(drug=d), cell_type] - E[GE|do(drug=control), cell_type]. This does not refer to any modeling assumption and leaves participants free to build their own molecular/cell-type specific representations. The only thing that needs to be provided is the aggregate gene expression per cell type (at the well level). And technical bias can be adjusted for using more advanced methods than linear regression. Because we are tasked to predict some p-values output by a linear model it seems the goal is to retro engineer the inputs to the limma model.",
    "2448368": "I still think there's confusion here about the input and output of the competition.\n\nIn inference, all you get is two strings for the `cell_type` and `perturbation` conditions. Your model then outputs predicts DE values for 18,000 genes. \n\nYou don't get the raw expression values for this combination of `cell_type` and `perturbation` for compounds in the public / private LB. There are multiple valid ways to predict the compound effect here. It's true that one valid approach would be to train a generative model to simulate inputs to the limma model and then run DE. Another is just to train a matrix completion model on the training data and directly output DE values. The choice is yours.",
    "2448460": "See also: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441845#2448372",
    "2457095": "If the targets are the -log10(p-value), then when I try to reverse it by doing -10^x, then how come the number ranges are not between 0-1? Am I misunderstanding something?\n\nEdit: Nvm, just saw that there's some sign(LFC) multiplied to it."
  },
  "source": "meta"
}