{
  "id": 438879,
  "title": "Structuring the Dataset",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/438879",
  "author_name": "",
  "post_date": "2023-09-12T21:33:40.217892300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Single cell data can be easy to work with in Seurat (R) and Scanpy (Python). Here, the data is available in .parquet format. Has anyone worked with this format before for single cell data and know how to set up Seurat objects for example from .parquet? </p>",
  "messages": [
    {
      "id": "2435336",
      "postDate": "09/12/2023 21:33:40",
      "content": "<p>Single cell data can be easy to work with in Seurat (R) and Scanpy (Python). Here, the data is available in .parquet format. Has anyone worked with this format before for single cell data and know how to set up Seurat objects for example from .parquet? </p>",
      "rawMarkdown": "Single cell data can be easy to work with in Seurat (R) and Scanpy (Python). Here, the data is available in .parquet format. Has anyone worked with this format before for single cell data and know how to set up Seurat objects for example from .parquet?",
      "votes": null
    },
    {
      "id": "2441515",
      "postDate": "09/16/2023 08:21:41",
      "content": "<p>Read in pandas first,then convert the data into h5ad</p>",
      "rawMarkdown": "Read in pandas first,then convert the data into h5ad",
      "votes": null
    },
    {
      "id": "2459935",
      "postDate": "09/28/2023 13:27:14",
      "content": "<p>Hello, I got struggle in create h5ad file. Can you help me to figure it? <br>\nThank you!</p>",
      "rawMarkdown": "Hello, I got struggle in create h5ad file. Can you help me to figure it? \nThank you!",
      "votes": null
    },
    {
      "id": "2464117",
      "postDate": "10/02/2023 00:44:42",
      "content": "<p>Hello, with R you can read .parquet with arrow package, like so:</p>\n<p><code>library(arrow)</code>  <br>\n<code>de_train &lt;- read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')</code></p>\n<p>for analysis with Seurat, it may be useful to refer to the <a href=\"https://satijalab.org/seurat/archive/v3.2/pbmc3k_tutorial.html\" target=\"_blank\">tutorial</a>, or my example notebook <a href=\"https://www.kaggle.com/code/antoninadolgorukova/mmscel-basic-analysis-with-seurat\" target=\"_blank\">here</a>.</p>\n<p>You need a matrix-like object with unnormalized data, cells by columns, and genes by rows. I am not sure, actually, do we have this data? The multiome_train.parquet contains baseline gene expression, but only for a part of genes in de_train.</p>",
      "rawMarkdown": "Hello, with R you can read .parquet with arrow package, like so:\n\n`library(arrow)`  \n`de_train <- read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')`\n\nfor analysis with Seurat, it may be useful to refer to the [tutorial](https://satijalab.org/seurat/archive/v3.2/pbmc3k_tutorial.html), or my example notebook [here](https://www.kaggle.com/code/antoninadolgorukova/mmscel-basic-analysis-with-seurat).\n\nYou need a matrix-like object with unnormalized data, cells by columns, and genes by rows. I am not sure, actually, do we have this data? The multiome_train.parquet contains baseline gene expression, but only for a part of genes in de_train.",
      "votes": null
    },
    {
      "id": "2464333",
      "postDate": "10/02/2023 07:03:41",
      "content": "<p>Python users may use that way:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle</a><br>\nsee also discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871</a></p>",
      "rawMarkdown": "Python users may use that way:\nhttps://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle\nsee also discussion:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2441515,
      "author_name": "yanchengwei",
      "author_url": "",
      "post_date": "09/16/2023 08:21:41",
      "content": "<p>Read in pandas first,then convert the data into h5ad</p>",
      "votes": null,
      "replies": [
        {
          "id": 2459935,
          "author_name": "ohonganh",
          "author_url": "",
          "post_date": "09/28/2023 13:27:14",
          "content": "<p>Hello, I got struggle in create h5ad file. Can you help me to figure it? <br>\nThank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2464117,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "10/02/2023 00:44:42",
      "content": "<p>Hello, with R you can read .parquet with arrow package, like so:</p>\n<p><code>library(arrow)</code>  <br>\n<code>de_train &lt;- read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')</code></p>\n<p>for analysis with Seurat, it may be useful to refer to the <a href=\"https://satijalab.org/seurat/archive/v3.2/pbmc3k_tutorial.html\" target=\"_blank\">tutorial</a>, or my example notebook <a href=\"https://www.kaggle.com/code/antoninadolgorukova/mmscel-basic-analysis-with-seurat\" target=\"_blank\">here</a>.</p>\n<p>You need a matrix-like object with unnormalized data, cells by columns, and genes by rows. I am not sure, actually, do we have this data? The multiome_train.parquet contains baseline gene expression, but only for a part of genes in de_train.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2464333,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/02/2023 07:03:41",
      "content": "<p>Python users may use that way:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle</a><br>\nsee also discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2435336": "Single cell data can be easy to work with in Seurat (R) and Scanpy (Python). Here, the data is available in .parquet format. Has anyone worked with this format before for single cell data and know how to set up Seurat objects for example from .parquet?",
    "2441515": "Read in pandas first,then convert the data into h5ad",
    "2459935": "Hello, I got struggle in create h5ad file. Can you help me to figure it? \nThank you!",
    "2464117": "Hello, with R you can read .parquet with arrow package, like so:\n\n`library(arrow)`  \n`de_train <- read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')`\n\nfor analysis with Seurat, it may be useful to refer to the [tutorial](https://satijalab.org/seurat/archive/v3.2/pbmc3k_tutorial.html), or my example notebook [here](https://www.kaggle.com/code/antoninadolgorukova/mmscel-basic-analysis-with-seurat).\n\nYou need a matrix-like object with unnormalized data, cells by columns, and genes by rows. I am not sure, actually, do we have this data? The multiome_train.parquet contains baseline gene expression, but only for a part of genes in de_train.",
    "2464333": "Python users may use that way:\nhttps://www.kaggle.com/code/alexandervc/op2-rna-seq-data-scanpy-adata-cell-cycle\nsee also discussion:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/443871"
  },
  "source": "meta"
}