{
  "id": 448033,
  "title": "pseudobulking",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/448033",
  "author_name": "",
  "post_date": "2023-10-18T06:50:52.528374900Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the data description: \"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature\". However, in code \"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\", pseudobulking was implemented by summing the raw gene expression (through function \"sum_by\").  Has anyone noticed this problem? </p>",
  "messages": [
    {
      "id": "2486764",
      "postDate": "10/18/2023 06:50:52",
      "content": "<p>In the data description: \"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature\". However, in code \"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\", pseudobulking was implemented by summing the raw gene expression (through function \"sum_by\").  Has anyone noticed this problem? </p>",
      "rawMarkdown": "In the data description: \"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature\". However, in code \"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\", pseudobulking was implemented by summing the raw gene expression (through function \"sum_by\").  Has anyone noticed this problem?",
      "votes": null
    },
    {
      "id": "2486904",
      "postDate": "10/18/2023 08:29:57",
      "content": "<p>Yes! But I <strong>guess</strong> this not change the result, because the <code>limma_fit.r</code> code will normalize it using the <code>calcNormFactors</code> in <code>edgeR</code>.</p>\n<p>The bigger problem is that the <code>pseudobulk</code> dose not match the <code>ground truth pseudobulk</code>, neither the shape nor the value.<br>\nI'm still trying to figure it out, vary frustrated.</p>",
      "rawMarkdown": "Yes! But I **guess** this not change the result, because the `limma_fit.r` code will normalize it using the `calcNormFactors` in `edgeR`.\n\nThe bigger problem is that the `pseudobulk` dose not match the `ground truth pseudobulk`, neither the shape nor the value.\nI'm still trying to figure it out, vary frustrated.",
      "votes": null
    },
    {
      "id": "2487252",
      "postDate": "10/18/2023 13:27:35",
      "content": "<p>Yes, we also find that the pseudobulk dose not match the ground truth pseudobulk! <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
      "rawMarkdown": "Yes, we also find that the pseudobulk dose not match the ground truth pseudobulk! @danielburkhardt",
      "votes": null
    },
    {
      "id": "2488443",
      "postDate": "10/19/2023 09:06:10",
      "content": "<p><strong>Update</strong>: In my notebook <a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3\" target=\"_blank\">op2-02-pseudobulk-v3</a>, comparing the <code>pseudobulk</code> vs <code>ground truth pseudobulk</code>.</p>\n<ul>\n<li>I found the <strong>donor_id</strong> and <strong>plate_name</strong> are different in the metadata.</li>\n<li>After fixed the metadata, I also found the <strong>value</strong> (count) of some genes were <strong>zero</strong> in <code>pseudobulk</code> vs <code>ground truth pseudobulk</code>.</li>\n</ul>\n<hr>\n<p>So I run Limma using <a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a> and compared the DE results of <code>pseudobulk</code> and<code>ground truth pseudobulk</code> with <code>de_train</code>,</p>\n<p>The results show that: </p>\n<ul>\n<li>MRRMSE(<code>ground truth pseudobulk</code> DE, <code>de_train</code>) = 0.033</li>\n<li>MRRMSE(<code>pseudobulk</code> DE, <code>de_train</code>) = <strong>0.185</strong> <strong>!!!</strong></li>\n<li>MRRMSE(<code>ground truth pseudobulk</code> DE, <code>pseudobulk</code> DE) = <strong>0.197</strong> <strong>!!!</strong></li>\n</ul>\n<p>The results and plot are in <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a>.</p>\n<hr>\n<p>These <strong>inconsistency</strong> of the data really annoyed me. <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Any comments? </p>",
      "rawMarkdown": "**Update**: In my notebook [op2-02-pseudobulk-v3](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3), comparing the `pseudobulk` vs `ground truth pseudobulk`.\n- I found the **donor_id** and **plate_name** are different in the metadata.\n- After fixed the metadata, I also found the **value** (count) of some genes were **zero** in `pseudobulk` vs `ground truth pseudobulk`.\n\n---\n\nSo I run Limma using [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma) and compared the DE results of `pseudobulk` and`ground truth pseudobulk` with `de_train`,\n\n\nThe results show that: \n- MRRMSE(`ground truth pseudobulk` DE, `de_train`) = 0.033\n- MRRMSE(`pseudobulk` DE, `de_train`) = **0.185** **!!!**\n- MRRMSE(`ground truth pseudobulk` DE, `pseudobulk` DE) = **0.197** **!!!**\n\nThe results and plot are in [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results).\n\n---\n\nThese **inconsistency** of the data really annoyed me. @danielburkhardt Any comments?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2486904,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "10/18/2023 08:29:57",
      "content": "<p>Yes! But I <strong>guess</strong> this not change the result, because the <code>limma_fit.r</code> code will normalize it using the <code>calcNormFactors</code> in <code>edgeR</code>.</p>\n<p>The bigger problem is that the <code>pseudobulk</code> dose not match the <code>ground truth pseudobulk</code>, neither the shape nor the value.<br>\nI'm still trying to figure it out, vary frustrated.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2487252,
          "author_name": "stevenxr",
          "author_url": "",
          "post_date": "10/18/2023 13:27:35",
          "content": "<p>Yes, we also find that the pseudobulk dose not match the ground truth pseudobulk! <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2488443,
          "author_name": "awater1223",
          "author_url": "",
          "post_date": "10/19/2023 09:06:10",
          "content": "<p><strong>Update</strong>: In my notebook <a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3\" target=\"_blank\">op2-02-pseudobulk-v3</a>, comparing the <code>pseudobulk</code> vs <code>ground truth pseudobulk</code>.</p>\n<ul>\n<li>I found the <strong>donor_id</strong> and <strong>plate_name</strong> are different in the metadata.</li>\n<li>After fixed the metadata, I also found the <strong>value</strong> (count) of some genes were <strong>zero</strong> in <code>pseudobulk</code> vs <code>ground truth pseudobulk</code>.</li>\n</ul>\n<hr>\n<p>So I run Limma using <a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a> and compared the DE results of <code>pseudobulk</code> and<code>ground truth pseudobulk</code> with <code>de_train</code>,</p>\n<p>The results show that: </p>\n<ul>\n<li>MRRMSE(<code>ground truth pseudobulk</code> DE, <code>de_train</code>) = 0.033</li>\n<li>MRRMSE(<code>pseudobulk</code> DE, <code>de_train</code>) = <strong>0.185</strong> <strong>!!!</strong></li>\n<li>MRRMSE(<code>ground truth pseudobulk</code> DE, <code>pseudobulk</code> DE) = <strong>0.197</strong> <strong>!!!</strong></li>\n</ul>\n<p>The results and plot are in <a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a>.</p>\n<hr>\n<p>These <strong>inconsistency</strong> of the data really annoyed me. <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Any comments? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2486764": "In the data description: \"We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature\". However, in code \"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\", pseudobulking was implemented by summing the raw gene expression (through function \"sum_by\").  Has anyone noticed this problem?",
    "2486904": "Yes! But I **guess** this not change the result, because the `limma_fit.r` code will normalize it using the `calcNormFactors` in `edgeR`.\n\nThe bigger problem is that the `pseudobulk` dose not match the `ground truth pseudobulk`, neither the shape nor the value.\nI'm still trying to figure it out, vary frustrated.",
    "2487252": "Yes, we also find that the pseudobulk dose not match the ground truth pseudobulk! @danielburkhardt",
    "2488443": "**Update**: In my notebook [op2-02-pseudobulk-v3](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v3), comparing the `pseudobulk` vs `ground truth pseudobulk`.\n- I found the **donor_id** and **plate_name** are different in the metadata.\n- After fixed the metadata, I also found the **value** (count) of some genes were **zero** in `pseudobulk` vs `ground truth pseudobulk`.\n\n---\n\nSo I run Limma using [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma) and compared the DE results of `pseudobulk` and`ground truth pseudobulk` with `de_train`,\n\n\nThe results show that: \n- MRRMSE(`ground truth pseudobulk` DE, `de_train`) = 0.033\n- MRRMSE(`pseudobulk` DE, `de_train`) = **0.185** **!!!**\n- MRRMSE(`ground truth pseudobulk` DE, `pseudobulk` DE) = **0.197** **!!!**\n\nThe results and plot are in [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results).\n\n---\n\nThese **inconsistency** of the data really annoyed me. @danielburkhardt Any comments?"
  },
  "source": "meta"
}