{
  "id": 443580,
  "title": "Limma model clarification",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/443580",
  "author_name": "",
  "post_date": "2023-09-27T23:31:37.443913500Z",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi there! So based on the input, it seems like our goal is to compute cell type specific p-values for each compound. However, the limma model shown doesn't seem to have an interaction term between cell type and compound, is there another way to compute cell type specific p-values or should there be an interaction term? </p>\n<p>Also the image says that each cell is an observation where before the description said that the data was pseudobulked, could you clarify which one is correct?</p>",
  "messages": [
    {
      "id": "2458949",
      "postDate": "09/27/2023 23:31:37",
      "content": "<p>Hi there! So based on the input, it seems like our goal is to compute cell type specific p-values for each compound. However, the limma model shown doesn't seem to have an interaction term between cell type and compound, is there another way to compute cell type specific p-values or should there be an interaction term? </p>\n<p>Also the image says that each cell is an observation where before the description said that the data was pseudobulked, could you clarify which one is correct?</p>",
      "rawMarkdown": "Hi there! So based on the input, it seems like our goal is to compute cell type specific p-values for each compound. However, the limma model shown doesn't seem to have an interaction term between cell type and compound, is there another way to compute cell type specific p-values or should there be an interaction term? \n\nAlso the image says that each cell is an observation where before the description said that the data was pseudobulked, could you clarify which one is correct?",
      "votes": null
    },
    {
      "id": "2460117",
      "postDate": "09/28/2023 15:26:00",
      "content": "<p>Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.</p>",
      "rawMarkdown": "Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.",
      "votes": null
    },
    {
      "id": "2460521",
      "postDate": "09/28/2023 20:42:49",
      "content": "<p>Hi, Daniel<br>\nThanks for your update. I am still confused about how to calculate the p-value for the combination of cell type and compounds. Just as <a href=\"https://www.kaggle.com/yauwning\" target=\"_blank\">@yauwning</a> said, there is no interaction term. Could you please provide the code to compute the p-value? (A mixture of R and python code is also great!)</p>\n<blockquote>\n  <p>Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.</p>\n</blockquote>",
      "rawMarkdown": "Hi, Daniel\nThanks for your update. I am still confused about how to calculate the p-value for the combination of cell type and compounds. Just as @yauwning said, there is no interaction term. Could you please provide the code to compute the p-value? (A mixture of R and python code is also great!)\n\n> Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.",
      "votes": null
    },
    {
      "id": "2466478",
      "postDate": "10/03/2023 21:00:19",
      "content": "<p>Thank you for all your work on the competition. Has the limma model description been updated yet with the cell type/perturbation interaction term? On the data page, it still looks to me as if the model does not have the interaction. See screen shot. Thanks for your help!</p>",
      "rawMarkdown": "Thank you for all your work on the competition. Has the limma model description been updated yet with the cell type/perturbation interaction term? On the data page, it still looks to me as if the model does not have the interaction. See screen shot. Thanks for your help!",
      "votes": null
    },
    {
      "id": "2466513",
      "postDate": "10/03/2023 22:21:18",
      "content": "<p>FYI looks like the code for running DE has been posted, they run limma separately for each cell type: <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\" target=\"_blank\">https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb</a></p>",
      "rawMarkdown": "FYI looks like the code for running DE has been posted, they run limma separately for each cell type: https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb",
      "votes": null
    },
    {
      "id": "2467911",
      "postDate": "10/05/2023 01:51:54",
      "content": "<p>Tried to run this local.  Took a while to remember how to get files from aws :)</p>\n<p>But cannot complete as I can't get past error as it seems to be hard coded to run on saturncloud.  Would appreciate any help on getting this cell to run (having more fun getting kaggle to not do silly stuff with the formatting of the stuff below)</p>\n<p>FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'</p>\n<blockquote>\n  <p>cell_types = bulk_adata.obs['cell_type'].unique()<br>\n  de_dfs = []</p>\n</blockquote>\n<p>for cell_type in cell_types:<br>\n    cell_type_selection = bulk_adata.obs['cell_type'].eq(cell_type)<br>\n    cell_type_bulk_adata = bulk_adata[cell_type_selection].copy()</p>\n<pre><code> = run_limma_for_cell_(cell_type_bulk_adata)\n\n.append(de_df)\n</code></pre>\n<p>de_dfs = c.compute(de_dfs, sync=True)<br>\nde_df = pd.concat(de_dfs)&gt;</p>",
      "rawMarkdown": "Tried to run this local.  Took a while to remember how to get files from aws :)\n\nBut cannot complete as I can't get past error as it seems to be hard coded to run on saturncloud.  Would appreciate any help on getting this cell to run (having more fun getting kaggle to not do silly stuff with the formatting of the stuff below)\n\nFileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'\n\n >cell_types = bulk_adata.obs['cell_type'].unique()\nde_dfs = []\n\nfor cell_type in cell_types:\n    cell_type_selection = bulk_adata.obs['cell_type'].eq(cell_type)\n    cell_type_bulk_adata = bulk_adata[cell_type_selection].copy()\n    \n    de_df = run_limma_for_cell_type(cell_type_bulk_adata)\n    \n    de_dfs.append(de_df)\n\nde_dfs = c.compute(de_dfs, sync=True)\nde_df = pd.concat(de_dfs)>",
      "votes": null
    },
    {
      "id": "2469602",
      "postDate": "10/06/2023 13:19:47",
      "content": "<p>Don't have R on my local machine.  Will install and see if I can get past this cell.</p>\n<p>There are a few flavors of R available - can the host shed some light on the version installed on SaturnCloud?</p>",
      "rawMarkdown": "Don't have R on my local machine.  Will install and see if I can get past this cell.\n\nThere are a few flavors of R available - can the host shed some light on the version installed on SaturnCloud?",
      "votes": null
    },
    {
      "id": "2469693",
      "postDate": "10/06/2023 14:25:20",
      "content": "<p>I'm going to be honest, I think you're way better off running on Saturn Cloud if possible. You can access instances there via ssh if that's more convenient. It took us a few days to get the packages set up. I'll work with <a href=\"https://www.kaggle.com/hhhuuugggooo\" target=\"_blank\">@hhhuuugggooo</a> to help get more details about the environment.</p>\n<p>Is there a reason you don't want to use Saturn? </p>",
      "rawMarkdown": "I'm going to be honest, I think you're way better off running on Saturn Cloud if possible. You can access instances there via ssh if that's more convenient. It took us a few days to get the packages set up. I'll work with @hhhuuugggooo to help get more details about the environment.\n\nIs there a reason you don't want to use Saturn?",
      "votes": null
    },
    {
      "id": "2469706",
      "postDate": "10/06/2023 14:39:34",
      "content": "<p>Thanks - no good reason beyond fact I spent a ton of money over the last few years and ended up with 4 local machines that I have trouble keeping busy :)</p>\n<p>I have found that keeping Python running successfully on local machines almost as much fun as coding in Python.  So running R scripts from Python would be a new learning experience.</p>",
      "rawMarkdown": "Thanks - no good reason beyond fact I spent a ton of money over the last few years and ended up with 4 local machines that I have trouble keeping busy :)\n\nI have found that keeping Python running successfully on local machines almost as much fun as coding in Python.  So running R scripts from Python would be a new learning experience.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2460117,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "09/28/2023 15:26:00",
      "content": "<p>Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2460521,
          "author_name": "xinmingtu",
          "author_url": "",
          "post_date": "09/28/2023 20:42:49",
          "content": "<p>Hi, Daniel<br>\nThanks for your update. I am still confused about how to calculate the p-value for the combination of cell type and compounds. Just as <a href=\"https://www.kaggle.com/yauwning\" target=\"_blank\">@yauwning</a> said, there is no interaction term. Could you please provide the code to compute the p-value? (A mixture of R and python code is also great!)</p>\n<blockquote>\n  <p>Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2466478,
      "author_name": "jameslong",
      "author_url": "",
      "post_date": "10/03/2023 21:00:19",
      "content": "<p>Thank you for all your work on the competition. Has the limma model description been updated yet with the cell type/perturbation interaction term? On the data page, it still looks to me as if the model does not have the interaction. See screen shot. Thanks for your help!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2466513,
      "author_name": "yauwning",
      "author_url": "",
      "post_date": "10/03/2023 22:21:18",
      "content": "<p>FYI looks like the code for running DE has been posted, they run limma separately for each cell type: <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb\" target=\"_blank\">https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2467911,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "10/05/2023 01:51:54",
          "content": "<p>Tried to run this local.  Took a while to remember how to get files from aws :)</p>\n<p>But cannot complete as I can't get past error as it seems to be hard coded to run on saturncloud.  Would appreciate any help on getting this cell to run (having more fun getting kaggle to not do silly stuff with the formatting of the stuff below)</p>\n<p>FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'</p>\n<blockquote>\n  <p>cell_types = bulk_adata.obs['cell_type'].unique()<br>\n  de_dfs = []</p>\n</blockquote>\n<p>for cell_type in cell_types:<br>\n    cell_type_selection = bulk_adata.obs['cell_type'].eq(cell_type)<br>\n    cell_type_bulk_adata = bulk_adata[cell_type_selection].copy()</p>\n<pre><code> = run_limma_for_cell_(cell_type_bulk_adata)\n\n.append(de_df)\n</code></pre>\n<p>de_dfs = c.compute(de_dfs, sync=True)<br>\nde_df = pd.concat(de_dfs)&gt;</p>",
          "votes": null,
          "replies": [
            {
              "id": 2469602,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "10/06/2023 13:19:47",
              "content": "<p>Don't have R on my local machine.  Will install and see if I can get past this cell.</p>\n<p>There are a few flavors of R available - can the host shed some light on the version installed on SaturnCloud?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2469693,
                  "author_name": "danielburkhardt",
                  "author_url": "",
                  "post_date": "10/06/2023 14:25:20",
                  "content": "<p>I'm going to be honest, I think you're way better off running on Saturn Cloud if possible. You can access instances there via ssh if that's more convenient. It took us a few days to get the packages set up. I'll work with <a href=\"https://www.kaggle.com/hhhuuugggooo\" target=\"_blank\">@hhhuuugggooo</a> to help get more details about the environment.</p>\n<p>Is there a reason you don't want to use Saturn? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2469706,
                      "author_name": "pcjimmmy",
                      "author_url": "",
                      "post_date": "10/06/2023 14:39:34",
                      "content": "<p>Thanks - no good reason beyond fact I spent a ton of money over the last few years and ended up with 4 local machines that I have trouble keeping busy :)</p>\n<p>I have found that keeping Python running successfully on local machines almost as much fun as coding in Python.  So running R scripts from Python would be a new learning experience.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2458949": "Hi there! So based on the input, it seems like our goal is to compute cell type specific p-values for each compound. However, the limma model shown doesn't seem to have an interaction term between cell type and compound, is there another way to compute cell type specific p-values or should there be an interaction term? \n\nAlso the image says that each cell is an observation where before the description said that the data was pseudobulked, could you clarify which one is correct?",
    "2460117": "Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.",
    "2460521": "Hi, Daniel\nThanks for your update. I am still confused about how to calculate the p-value for the combination of cell type and compounds. Just as @yauwning said, there is no interaction term. Could you please provide the code to compute the p-value? (A mixture of R and python code is also great!)\n\n> Good catch! I've updated the figure and added some text. We fit the limma model to the pseudobulked raw counts.",
    "2466478": "Thank you for all your work on the competition. Has the limma model description been updated yet with the cell type/perturbation interaction term? On the data page, it still looks to me as if the model does not have the interaction. See screen shot. Thanks for your help!",
    "2466513": "FYI looks like the code for running DE has been posted, they run limma separately for each cell type: https://github.com/openproblems-bio/neurips-2023-scripts/blob/main/compute_de.ipynb",
    "2467911": "Tried to run this local.  Took a while to remember how to get files from aws :)\n\nBut cannot complete as I can't get past error as it seems to be hard coded to run on saturncloud.  Would appreciate any help on getting this cell to run (having more fun getting kaggle to not do silly stuff with the formatting of the stuff below)\n\nFileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'FileNotFoundError: [Errno 2] No such file or directory: '/opt/saturncloud/envs/rscript/bin/Rscript'\n\n >cell_types = bulk_adata.obs['cell_type'].unique()\nde_dfs = []\n\nfor cell_type in cell_types:\n    cell_type_selection = bulk_adata.obs['cell_type'].eq(cell_type)\n    cell_type_bulk_adata = bulk_adata[cell_type_selection].copy()\n    \n    de_df = run_limma_for_cell_type(cell_type_bulk_adata)\n    \n    de_dfs.append(de_df)\n\nde_dfs = c.compute(de_dfs, sync=True)\nde_df = pd.concat(de_dfs)>",
    "2469602": "Don't have R on my local machine.  Will install and see if I can get past this cell.\n\nThere are a few flavors of R available - can the host shed some light on the version installed on SaturnCloud?",
    "2469693": "I'm going to be honest, I think you're way better off running on Saturn Cloud if possible. You can access instances there via ssh if that's more convenient. It took us a few days to get the packages set up. I'll work with @hhhuuugggooo to help get more details about the environment.\n\nIs there a reason you don't want to use Saturn?",
    "2469706": "Thanks - no good reason beyond fact I spent a ton of money over the last few years and ended up with 4 local machines that I have trouble keeping busy :)\n\nI have found that keeping Python running successfully on local machines almost as much fun as coding in Python.  So running R scripts from Python would be a new learning experience."
  },
  "source": "meta"
}