{
  "id": 440625,
  "title": "Clarification on DE analysis and de_train.parquet",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/440625",
  "author_name": "",
  "post_date": "2023-09-15T17:28:30.065699600Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all, <br>\nCould someone explain why DE analysis and modeling was conducted/shared in the first place? As I understand it, the goal is to predict DE at the cell type/small-molecule/gene level. I don't understand how the <code>de_train.parquet</code> data is useful in anyway as a \"ground truth dataset\" as it seems to come from a linear model (Limma) prediction according to the challenge description.</p>\n<p>Edit:<br>\nIn other words I am making sure the goal of the challenge is <strong>NOT</strong> to run a Limma model.</p>",
  "messages": [
    {
      "id": "2440723",
      "postDate": "09/15/2023 17:28:30",
      "content": "<p>Hi all, <br>\nCould someone explain why DE analysis and modeling was conducted/shared in the first place? As I understand it, the goal is to predict DE at the cell type/small-molecule/gene level. I don't understand how the <code>de_train.parquet</code> data is useful in anyway as a \"ground truth dataset\" as it seems to come from a linear model (Limma) prediction according to the challenge description.</p>\n<p>Edit:<br>\nIn other words I am making sure the goal of the challenge is <strong>NOT</strong> to run a Limma model.</p>",
      "rawMarkdown": "Hi all, \nCould someone explain why DE analysis and modeling was conducted/shared in the first place? As I understand it, the goal is to predict DE at the cell type/small-molecule/gene level. I don't understand how the `de_train.parquet` data is useful in anyway as a \"ground truth dataset\" as it seems to come from a linear model (Limma) prediction according to the challenge description.\n\nEdit:\nIn other words I am making sure the goal of the challenge is **NOT** to run a Limma model.",
      "votes": null
    },
    {
      "id": "2440928",
      "postDate": "09/15/2023 20:22:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/olabayle\" target=\"_blank\">@olabayle</a>, I'm not sure I fully understand your question here. The goal is to predict differential expression when you haven't seen a perturbation for a particular cell type.</p>\n<p>The reason we use Limma is to identify which parts of the gene expression profile in the perturbation condition is caused by the perturbation, and not caused by technical noise. This paper might be helpful: <a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">https://www.nature.com/articles/s41467-021-25960-2</a></p>\n<p>Let me know if you have any further questions.</p>",
      "rawMarkdown": "Hi @olabayle, I'm not sure I fully understand your question here. The goal is to predict differential expression when you haven't seen a perturbation for a particular cell type.\n\nThe reason we use Limma is to identify which parts of the gene expression profile in the perturbation condition is caused by the perturbation, and not caused by technical noise. This paper might be helpful: https://www.nature.com/articles/s41467-021-25960-2\n\nLet me know if you have any further questions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2440928,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "09/15/2023 20:22:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/olabayle\" target=\"_blank\">@olabayle</a>, I'm not sure I fully understand your question here. The goal is to predict differential expression when you haven't seen a perturbation for a particular cell type.</p>\n<p>The reason we use Limma is to identify which parts of the gene expression profile in the perturbation condition is caused by the perturbation, and not caused by technical noise. This paper might be helpful: <a href=\"https://www.nature.com/articles/s41467-021-25960-2\" target=\"_blank\">https://www.nature.com/articles/s41467-021-25960-2</a></p>\n<p>Let me know if you have any further questions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2440723": "Hi all, \nCould someone explain why DE analysis and modeling was conducted/shared in the first place? As I understand it, the goal is to predict DE at the cell type/small-molecule/gene level. I don't understand how the `de_train.parquet` data is useful in anyway as a \"ground truth dataset\" as it seems to come from a linear model (Limma) prediction according to the challenge description.\n\nEdit:\nIn other words I am making sure the goal of the challenge is **NOT** to run a Limma model.",
    "2440928": "Hi @olabayle, I'm not sure I fully understand your question here. The goal is to predict differential expression when you haven't seen a perturbation for a particular cell type.\n\nThe reason we use Limma is to identify which parts of the gene expression profile in the perturbation condition is caused by the perturbation, and not caused by technical noise. This paper might be helpful: https://www.nature.com/articles/s41467-021-25960-2\n\nLet me know if you have any further questions."
  },
  "source": "meta"
}