{
  "id": 446034,
  "title": "Run the differential expression code using Kaggle kernels",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/446034",
  "author_name": "",
  "post_date": "2023-10-10T03:35:40.990888300Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm glad to share that I have successfully run the differential expression code (<code>pseudobulk</code> and <code>Limma</code>) using <strong>three Kaggle kernels</strong>.</p>\n<hr>\n<p>The openproblems-bio recently (Oct. 3, 2023) released the code for differential expression: <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts\" target=\"_blank\">repo</a>,  discussion in <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444986\" target=\"_blank\">here</a>.</p>\n<p>There are three steps in the main code <strong><code>compute_de.ipynb</code></strong>:</p>\n<ol>\n<li>pseudobulk</li>\n<li>run limma (using R and dask)</li>\n<li>check the results (compared with <code>de_train.parquet</code>)</li>\n</ol>\n<hr>\n<p>So I split the code into three Kaggle kernels (notebook):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2\" target=\"_blank\">op2-02-pseudobulk-v2</a><ul>\n<li><strong>generate pseudobulk <code>bulk_adata.h5ad</code></strong></li>\n<li>it contains the full code from openproblems-bio, but the limma part is not running in this notebook</li>\n<li>I wrote some glue code to generate the intermediate files for limma (generate <strong><code>input.h5ad</code></strong> files for each cell type and a bash code <strong><code>r_code.sh</code></strong> for running R)</li></ul></li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a><ul>\n<li><strong>setup R environment for limma using a Python kernel</strong><ul>\n<li>it costs me a lot of time</li>\n<li>the Python kernel works better than R kernel even for a R environment (@_@)</li></ul></li>\n<li><strong>running limma</strong></li>\n<li><strong>view the voom plot</strong></li>\n<li>using the output from <a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2\" target=\"_blank\">op2-02-pseudobulk-v2</a></li></ul></li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a><ul>\n<li><strong>merge the limma results and check</strong> (compared with <code>de_train.parquet</code>)</li>\n<li>using the output from <a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a></li></ul></li>\n</ul>\n<hr>\n<p>BTW (ads for my other notebooks):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\" target=\"_blank\">op2-00-basic-metadata-eda</a>: basic view the data, the <strong>data split</strong> part contains the compounds <strong><code>sm_lincs_id</code></strong> in <strong>public and private LB</strong>.</li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-01-singlecell-adata-eda\" target=\"_blank\">op2-01-singlecell-adata-eda</a>: <strong>UMAP</strong> view for the single cell data.</li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v1\" target=\"_blank\">op2-02-pseudobulk-v1</a>: before openproblems-bio release the code, I tried to reproduce the differential expression using <a href=\"https://github.com/owkin/PyDESeq2\" target=\"_blank\"><strong>PyDESeq2</strong></a>, but not working yet.</li>\n</ul>",
  "messages": [
    {
      "id": "2475610",
      "postDate": "10/10/2023 03:35:40",
      "content": "<p>I'm glad to share that I have successfully run the differential expression code (<code>pseudobulk</code> and <code>Limma</code>) using <strong>three Kaggle kernels</strong>.</p>\n<hr>\n<p>The openproblems-bio recently (Oct. 3, 2023) released the code for differential expression: <a href=\"https://github.com/openproblems-bio/neurips-2023-scripts\" target=\"_blank\">repo</a>,  discussion in <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444986\" target=\"_blank\">here</a>.</p>\n<p>There are three steps in the main code <strong><code>compute_de.ipynb</code></strong>:</p>\n<ol>\n<li>pseudobulk</li>\n<li>run limma (using R and dask)</li>\n<li>check the results (compared with <code>de_train.parquet</code>)</li>\n</ol>\n<hr>\n<p>So I split the code into three Kaggle kernels (notebook):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2\" target=\"_blank\">op2-02-pseudobulk-v2</a><ul>\n<li><strong>generate pseudobulk <code>bulk_adata.h5ad</code></strong></li>\n<li>it contains the full code from openproblems-bio, but the limma part is not running in this notebook</li>\n<li>I wrote some glue code to generate the intermediate files for limma (generate <strong><code>input.h5ad</code></strong> files for each cell type and a bash code <strong><code>r_code.sh</code></strong> for running R)</li></ul></li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a><ul>\n<li><strong>setup R environment for limma using a Python kernel</strong><ul>\n<li>it costs me a lot of time</li>\n<li>the Python kernel works better than R kernel even for a R environment (@_@)</li></ul></li>\n<li><strong>running limma</strong></li>\n<li><strong>view the voom plot</strong></li>\n<li>using the output from <a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2\" target=\"_blank\">op2-02-pseudobulk-v2</a></li></ul></li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results\" target=\"_blank\">op2-04-merge-limma-results</a><ul>\n<li><strong>merge the limma results and check</strong> (compared with <code>de_train.parquet</code>)</li>\n<li>using the output from <a href=\"https://www.kaggle.com/code/awater1223/op2-03-limma\" target=\"_blank\">op2-03-limma</a></li></ul></li>\n</ul>\n<hr>\n<p>BTW (ads for my other notebooks):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\" target=\"_blank\">op2-00-basic-metadata-eda</a>: basic view the data, the <strong>data split</strong> part contains the compounds <strong><code>sm_lincs_id</code></strong> in <strong>public and private LB</strong>.</li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-01-singlecell-adata-eda\" target=\"_blank\">op2-01-singlecell-adata-eda</a>: <strong>UMAP</strong> view for the single cell data.</li>\n<li><a href=\"https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v1\" target=\"_blank\">op2-02-pseudobulk-v1</a>: before openproblems-bio release the code, I tried to reproduce the differential expression using <a href=\"https://github.com/owkin/PyDESeq2\" target=\"_blank\"><strong>PyDESeq2</strong></a>, but not working yet.</li>\n</ul>",
      "rawMarkdown": "I'm glad to share that I have successfully run the differential expression code (`pseudobulk` and `Limma`) using **three Kaggle kernels**.\n\n---\n\nThe openproblems-bio recently (Oct. 3, 2023) released the code for differential expression: [repo](https://github.com/openproblems-bio/neurips-2023-scripts),  discussion in [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493) and [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444986).\n\nThere are three steps in the main code **`compute_de.ipynb`**:\n1. pseudobulk\n2. run limma (using R and dask)\n3. check the results (compared with `de_train.parquet`)\n\n---\n\nSo I split the code into three Kaggle kernels (notebook):\n- [op2-02-pseudobulk-v2](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2)\n  - **generate pseudobulk `bulk_adata.h5ad`**\n  - it contains the full code from openproblems-bio, but the limma part is not running in this notebook\n  - I wrote some glue code to generate the intermediate files for limma (generate **`input.h5ad`** files for each cell type and a bash code **`r_code.sh`** for running R)\n- [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma)\n  - **setup R environment for limma using a Python kernel**\n      - it costs me a lot of time\n      - the Python kernel works better than R kernel even for a R environment (@_@)\n  - **running limma**\n  - **view the voom plot**\n  - using the output from [op2-02-pseudobulk-v2](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2)\n- [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results)\n  - **merge the limma results and check** (compared with `de_train.parquet`)\n  - using the output from [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma)\n\n---\n\nBTW (ads for my other notebooks):\n- [op2-00-basic-metadata-eda](https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda): basic view the data, the **data split** part contains the compounds **`sm_lincs_id`** in **public and private LB**.\n- [op2-01-singlecell-adata-eda](https://www.kaggle.com/code/awater1223/op2-01-singlecell-adata-eda): **UMAP** view for the single cell data.\n- [op2-02-pseudobulk-v1](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v1): before openproblems-bio release the code, I tried to reproduce the differential expression using [**PyDESeq2**](https://github.com/owkin/PyDESeq2), but not working yet.",
      "votes": null
    },
    {
      "id": "2497264",
      "postDate": "10/24/2023 14:50:25",
      "content": "<p>Thank you so much for your contribution! This is really useful. Your code works perfectly on kaggle. I am wondering if you have tried to use your local machine to run the code? On my local machine, when I run the R code, there is always an error:  Loading H5AD\\nError: AnnDataReadError: Above error raised while reading key '/layers' of type  from /. <br>\nDo you have any idea about the issue? Thank you a lot! </p>",
      "rawMarkdown": "Thank you so much for your contribution! This is really useful. Your code works perfectly on kaggle. I am wondering if you have tried to use your local machine to run the code? On my local machine, when I run the R code, there is always an error:  Loading H5AD\\nError: AnnDataReadError: Above error raised while reading key '/layers' of type <class 'h5py._hl.group.Group'> from /. \nDo you have any idea about the issue? Thank you a lot!",
      "votes": null
    },
    {
      "id": "2497935",
      "postDate": "10/25/2023 02:47:03",
      "content": "<p>Maybe try another version of anndata?</p>",
      "rawMarkdown": "Maybe try another version of anndata?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2497264,
      "author_name": "shuhuiwang1",
      "author_url": "",
      "post_date": "10/24/2023 14:50:25",
      "content": "<p>Thank you so much for your contribution! This is really useful. Your code works perfectly on kaggle. I am wondering if you have tried to use your local machine to run the code? On my local machine, when I run the R code, there is always an error:  Loading H5AD\\nError: AnnDataReadError: Above error raised while reading key '/layers' of type  from /. <br>\nDo you have any idea about the issue? Thank you a lot! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2497935,
          "author_name": "awater1223",
          "author_url": "",
          "post_date": "10/25/2023 02:47:03",
          "content": "<p>Maybe try another version of anndata?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2475610": "I'm glad to share that I have successfully run the differential expression code (`pseudobulk` and `Limma`) using **three Kaggle kernels**.\n\n---\n\nThe openproblems-bio recently (Oct. 3, 2023) released the code for differential expression: [repo](https://github.com/openproblems-bio/neurips-2023-scripts),  discussion in [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443493) and [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444986).\n\nThere are three steps in the main code **`compute_de.ipynb`**:\n1. pseudobulk\n2. run limma (using R and dask)\n3. check the results (compared with `de_train.parquet`)\n\n---\n\nSo I split the code into three Kaggle kernels (notebook):\n- [op2-02-pseudobulk-v2](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2)\n  - **generate pseudobulk `bulk_adata.h5ad`**\n  - it contains the full code from openproblems-bio, but the limma part is not running in this notebook\n  - I wrote some glue code to generate the intermediate files for limma (generate **`input.h5ad`** files for each cell type and a bash code **`r_code.sh`** for running R)\n- [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma)\n  - **setup R environment for limma using a Python kernel**\n      - it costs me a lot of time\n      - the Python kernel works better than R kernel even for a R environment (@_@)\n  - **running limma**\n  - **view the voom plot**\n  - using the output from [op2-02-pseudobulk-v2](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v2)\n- [op2-04-merge-limma-results](https://www.kaggle.com/code/awater1223/op2-04-merge-limma-results)\n  - **merge the limma results and check** (compared with `de_train.parquet`)\n  - using the output from [op2-03-limma](https://www.kaggle.com/code/awater1223/op2-03-limma)\n\n---\n\nBTW (ads for my other notebooks):\n- [op2-00-basic-metadata-eda](https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda): basic view the data, the **data split** part contains the compounds **`sm_lincs_id`** in **public and private LB**.\n- [op2-01-singlecell-adata-eda](https://www.kaggle.com/code/awater1223/op2-01-singlecell-adata-eda): **UMAP** view for the single cell data.\n- [op2-02-pseudobulk-v1](https://www.kaggle.com/code/awater1223/op2-02-pseudobulk-v1): before openproblems-bio release the code, I tried to reproduce the differential expression using [**PyDESeq2**](https://github.com/owkin/PyDESeq2), but not working yet.",
    "2497264": "Thank you so much for your contribution! This is really useful. Your code works perfectly on kaggle. I am wondering if you have tried to use your local machine to run the code? On my local machine, when I run the R code, there is always an error:  Loading H5AD\\nError: AnnDataReadError: Above error raised while reading key '/layers' of type <class 'h5py._hl.group.Group'> from /. \nDo you have any idea about the issue? Thank you a lot!",
    "2497935": "Maybe try another version of anndata?"
  },
  "source": "meta"
}