{
  "id": 443792,
  "title": "Clarification on DE analysis",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/443792",
  "author_name": "Phi Bya",
  "post_date": "2023-09-28T20:05:47.816000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The following text is in the data tab:</p>\n<p>\"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature. We then fit a linear model to the pseudobulked counts data using Limma and include the library (row), plate, and donor as technical covariates and compound as the experimental covariate. Here, pseudobulked means we summed the raw counts for all cells of a given type for each well in the experiment.\"</p>\n<p>This is confusing to me, by \"averaging the raw gene expression counts in each cell of a specific type in each sample\", do you mean averaging the raw gene expression counts of <em>all cells</em> belonging to a specific type in each sample?  But after that: \"pseudobulked means we <strong>summed</strong> the raw counts for all cells of a given type for each well in the experiment\". Here I assume each sample is a well in a plate, is it <em>sum</em> or <em>average</em>?</p>\n<p>Another question I have is: how were p-values and fold change for the combination of (compound and cell type) obtained when the model was fitted only once using all samples but there is no interaction term between compound and cell type in the model? Sorry if I am being ignorant on this.</p>",
  "messages": [
    {
      "id": 2460489,
      "postDate": "2023-09-28T20:05:47.817Z",
      "content": "<p>The following text is in the data tab:</p>\n<p>\"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature. We then fit a linear model to the pseudobulked counts data using Limma and include the library (row), plate, and donor as technical covariates and compound as the experimental covariate. Here, pseudobulked means we summed the raw counts for all cells of a given type for each well in the experiment.\"</p>\n<p>This is confusing to me, by \"averaging the raw gene expression counts in each cell of a specific type in each sample\", do you mean averaging the raw gene expression counts of <em>all cells</em> belonging to a specific type in each sample?  But after that: \"pseudobulked means we <strong>summed</strong> the raw counts for all cells of a given type for each well in the experiment\". Here I assume each sample is a well in a plate, is it <em>sum</em> or <em>average</em>?</p>\n<p>Another question I have is: how were p-values and fold change for the combination of (compound and cell type) obtained when the model was fitted only once using all samples but there is no interaction term between compound and cell type in the model? Sorry if I am being ignorant on this.</p>",
      "rawMarkdown": "The following text is in the data tab:\n\n\"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature. We then fit a linear model to the pseudobulked counts data using Limma and include the library (row), plate, and donor as technical covariates and compound as the experimental covariate. Here, pseudobulked means we summed the raw counts for all cells of a given type for each well in the experiment.\"\n\nThis is confusing to me, by \"averaging the raw gene expression counts in each cell of a specific type in each sample\", do you mean averaging the raw gene expression counts of *all cells* belonging to a specific type in each sample?  But after that: \"pseudobulked means we **summed** the raw counts for all cells of a given type for each well in the experiment\". Here I assume each sample is a well in a plate, is it *sum* or *average*?\n\nAnother question I have is: how were p-values and fold change for the combination of (compound and cell type) obtained when the model was fitted only once using all samples but there is no interaction term between compound and cell type in the model? Sorry if I am being ignorant on this.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2460489": "The following text is in the data tab:\n\n\"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature. We then fit a linear model to the pseudobulked counts data using Limma and include the library (row), plate, and donor as technical covariates and compound as the experimental covariate. Here, pseudobulked means we summed the raw counts for all cells of a given type for each well in the experiment.\"\n\nThis is confusing to me, by \"averaging the raw gene expression counts in each cell of a specific type in each sample\", do you mean averaging the raw gene expression counts of *all cells* belonging to a specific type in each sample?  But after that: \"pseudobulked means we **summed** the raw counts for all cells of a given type for each well in the experiment\". Here I assume each sample is a well in a plate, is it *sum* or *average*?\n\nAnother question I have is: how were p-values and fold change for the combination of (compound and cell type) obtained when the model was fitted only once using all samples but there is no interaction term between compound and cell type in the model? Sorry if I am being ignorant on this."
  }
}