{
  "id": 454107,
  "title": "New adata_excluded_ids File",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/454107",
  "author_name": "",
  "post_date": "2023-11-08T23:57:14.255605700Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>We just added a file <code>adata_excluded_ids.csv</code> to the competition dataset. This file contains <code>obs_id</code> and <code>gene</code> pairs that can be excluded from <code>adata_train.parquet</code> to obtain a filtered version of this dataset.</p>\n<p>Using Polars, you can recover the filtered version like so:</p>\n<pre><code> polars  pl\n pathlib  \n\ndata_dir = Path(\"/kaggle/input/open-problems-single-cell-perturbations\")\nadata = pl.read_parquet(data_dir / )\nexcluded_ids = pl.read_csv(data_dir / )\n\nadata_filtered = adata.(excluded_ids, =[, ], how=)\n</code></pre>\n<p>The competition host will be posting a description of this filtered data shortly.</p>",
  "messages": [
    {
      "id": "2517984",
      "postDate": "11/08/2023 23:57:14",
      "content": "<p>Hi everyone,</p>\n<p>We just added a file <code>adata_excluded_ids.csv</code> to the competition dataset. This file contains <code>obs_id</code> and <code>gene</code> pairs that can be excluded from <code>adata_train.parquet</code> to obtain a filtered version of this dataset.</p>\n<p>Using Polars, you can recover the filtered version like so:</p>\n<pre><code> polars  pl\n pathlib  \n\ndata_dir = Path(\"/kaggle/input/open-problems-single-cell-perturbations\")\nadata = pl.read_parquet(data_dir / )\nexcluded_ids = pl.read_csv(data_dir / )\n\nadata_filtered = adata.(excluded_ids, =[, ], how=)\n</code></pre>\n<p>The competition host will be posting a description of this filtered data shortly.</p>",
      "rawMarkdown": "Hi everyone,\n\nWe just added a file `adata_excluded_ids.csv` to the competition dataset. This file contains `obs_id` and `gene` pairs that can be excluded from `adata_train.parquet` to obtain a filtered version of this dataset.\n\nUsing Polars, you can recover the filtered version like so:\n\n```\nimport polars as pl\nfrom pathlib import Path\n\ndata_dir = Path(\"/kaggle/input/open-problems-single-cell-perturbations\")\nadata = pl.read_parquet(data_dir / 'adata_train.parquet')\nexcluded_ids = pl.read_csv(data_dir / 'excluded_ids.csv')\n\nadata_filtered = adata.join(excluded_ids, on=['obs_id', 'gene'], how='anti')\n```\n\nThe competition host will be posting a description of this filtered data shortly.",
      "votes": null
    },
    {
      "id": "2543099",
      "postDate": "11/29/2023 19:04:06",
      "content": "<p>To clarify Ryan's comment above: this filtered version of <code>adata_train.parquet</code> reflects the filtering that was done prior to computing differential expression. If you compute differential expression without this filtering, you will obtain results that are very correlated with but not numerically equivalent to <code>de_train.parquet</code>. </p>\n<p>The values that were filtered out from <code>adata_train.parquet</code> are mostly negligible; over 97% of the values that were filtered out were for a single transcript count, and the 99.9th percentile of filtered values was for only 4 counts. We do not expect this filtering to have a significant impact on model training/development. We only decided to include this extra filtering in the competition data for the sake of completeness.</p>",
      "rawMarkdown": "To clarify Ryan's comment above: this filtered version of `adata_train.parquet` reflects the filtering that was done prior to computing differential expression. If you compute differential expression without this filtering, you will obtain results that are very correlated with but not numerically equivalent to `de_train.parquet`. \n\nThe values that were filtered out from `adata_train.parquet` are mostly negligible; over 97% of the values that were filtered out were for a single transcript count, and the 99.9th percentile of filtered values was for only 4 counts. We do not expect this filtering to have a significant impact on model training/development. We only decided to include this extra filtering in the competition data for the sake of completeness.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2543099,
      "author_name": "andrewbenz",
      "author_url": "",
      "post_date": "11/29/2023 19:04:06",
      "content": "<p>To clarify Ryan's comment above: this filtered version of <code>adata_train.parquet</code> reflects the filtering that was done prior to computing differential expression. If you compute differential expression without this filtering, you will obtain results that are very correlated with but not numerically equivalent to <code>de_train.parquet</code>. </p>\n<p>The values that were filtered out from <code>adata_train.parquet</code> are mostly negligible; over 97% of the values that were filtered out were for a single transcript count, and the 99.9th percentile of filtered values was for only 4 counts. We do not expect this filtering to have a significant impact on model training/development. We only decided to include this extra filtering in the competition data for the sake of completeness.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2517984": "Hi everyone,\n\nWe just added a file `adata_excluded_ids.csv` to the competition dataset. This file contains `obs_id` and `gene` pairs that can be excluded from `adata_train.parquet` to obtain a filtered version of this dataset.\n\nUsing Polars, you can recover the filtered version like so:\n\n```\nimport polars as pl\nfrom pathlib import Path\n\ndata_dir = Path(\"/kaggle/input/open-problems-single-cell-perturbations\")\nadata = pl.read_parquet(data_dir / 'adata_train.parquet')\nexcluded_ids = pl.read_csv(data_dir / 'excluded_ids.csv')\n\nadata_filtered = adata.join(excluded_ids, on=['obs_id', 'gene'], how='anti')\n```\n\nThe competition host will be posting a description of this filtered data shortly.",
    "2543099": "To clarify Ryan's comment above: this filtered version of `adata_train.parquet` reflects the filtering that was done prior to computing differential expression. If you compute differential expression without this filtering, you will obtain results that are very correlated with but not numerically equivalent to `de_train.parquet`. \n\nThe values that were filtered out from `adata_train.parquet` are mostly negligible; over 97% of the values that were filtered out were for a single transcript count, and the 99.9th percentile of filtered values was for only 4 counts. We do not expect this filtering to have a significant impact on model training/development. We only decided to include this extra filtering in the competition data for the sake of completeness."
  },
  "source": "meta"
}