{
  "id": 448892,
  "title": "Is there the mapping of gene symbol to Ensembl gene id?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/448892",
  "author_name": "",
  "post_date": "2023-10-22T04:40:33.188447200Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Is there any good way to get the mapping of gene symbols in <code>de_train.parquet</code> to Ensembl gene id?</p>\n<p>In 10x cellranger output, a file like \"features.tsv.gz\" maps gene symbols to Ensembl gene ids. Is there a mapping like \"features.tsv.gz\"? </p>\n<p>I tried to download \"2020-A\" from <a href=\"https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads;\" target=\"_blank\">https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads;</a> refer to <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440959\" target=\"_blank\">this discussion</a>. I parsed <code>genes.gtf</code>, but about 2000 gene symbols cannot convert to Emsenbl gene ids. </p>\n<p>I used this very simple script to check gene symbols:</p>\n<pre><code> polars  pl\nde_df = pl.read_parquet()\ngenes = (de_df.columns[:])\n\n () -&gt; [, ]:\n    genes_map = {}\n     (path)  f:\n         i, line  (f):\n             line.startswith():\n                \n\n            line = line.strip()\n            elems = line.split()\n            attrs = elems[]\n            d = {}\n             attr  attrs.split():\n                attr = attr.strip()\n                 (attr) == :\n                    \n                k, v = attr.split()\n                k = k.strip()\n                v = v.strip()\n                d[k] = v\n\n            genes_map[d[]] = d[]\n\n     genes_map\n\ngenes_map = read_gtf()\n\n(((genes) - (genes_map.values())))\n\n</code></pre>",
  "messages": [
    {
      "id": "2491816",
      "postDate": "10/22/2023 04:40:33",
      "content": "<p>Is there any good way to get the mapping of gene symbols in <code>de_train.parquet</code> to Ensembl gene id?</p>\n<p>In 10x cellranger output, a file like \"features.tsv.gz\" maps gene symbols to Ensembl gene ids. Is there a mapping like \"features.tsv.gz\"? </p>\n<p>I tried to download \"2020-A\" from <a href=\"https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads;\" target=\"_blank\">https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads;</a> refer to <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440959\" target=\"_blank\">this discussion</a>. I parsed <code>genes.gtf</code>, but about 2000 gene symbols cannot convert to Emsenbl gene ids. </p>\n<p>I used this very simple script to check gene symbols:</p>\n<pre><code> polars  pl\nde_df = pl.read_parquet()\ngenes = (de_df.columns[:])\n\n () -&gt; [, ]:\n    genes_map = {}\n     (path)  f:\n         i, line  (f):\n             line.startswith():\n                \n\n            line = line.strip()\n            elems = line.split()\n            attrs = elems[]\n            d = {}\n             attr  attrs.split():\n                attr = attr.strip()\n                 (attr) == :\n                    \n                k, v = attr.split()\n                k = k.strip()\n                v = v.strip()\n                d[k] = v\n\n            genes_map[d[]] = d[]\n\n     genes_map\n\ngenes_map = read_gtf()\n\n(((genes) - (genes_map.values())))\n\n</code></pre>",
      "rawMarkdown": "Is there any good way to get the mapping of gene symbols in `de_train.parquet` to Ensembl gene id?\n\nIn 10x cellranger output, a file like \"features.tsv.gz\" maps gene symbols to Ensembl gene ids. Is there a mapping like \"features.tsv.gz\"? \n\nI tried to download \"2020-A\" from https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads; refer to [this discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440959). I parsed `genes.gtf`, but about 2000 gene symbols cannot convert to Emsenbl gene ids. \n\nI used this very simple script to check gene symbols:\n\n```python\nimport polars as pl\nde_df = pl.read_parquet(\"./data/de_train.parquet\")\ngenes = set(de_df.columns[5:])\n\ndef read_gtf(path: str) -> dict[str, str]:\n    genes_map = {}\n    with open(path) as f:\n        for i, line in enumerate(f):\n            if line.startswith(\"#\"):\n                continue\n\n            line = line.strip()\n            elems = line.split(\"\\t\")\n            attrs = elems[8]\n            d = {}\n            for attr in attrs.split(\";\"):\n                attr = attr.strip()\n                if len(attr) == 0:\n                    continue\n                k, v = attr.split(\" \")\n                k = k.strip(\"\\\"\")\n                v = v.strip(\"\\\"\")\n                d[k] = v\n\n            genes_map[d[\"gene_id\"]] = d[\"gene_name\"]\n\n    return genes_map\n\ngenes_map = read_gtf(\"./refdata-gex-GRCh38-2020-A/genes/genes.gtf\")\n\nprint(len(set(genes) - set(genes_map.values())))\n# 2399\n```",
      "votes": null
    },
    {
      "id": "2496097",
      "postDate": "10/23/2023 18:13:51",
      "content": "<p>Have you checked: <a href=\"https://docs.mygene.info/projects/mygene-py/en/latest/\" target=\"_blank\">https://docs.mygene.info/projects/mygene-py/en/latest/</a><br>\nI want to map gene symbols to NCBI GenID. However, I find not easy way to do that in python. For example this one <code>CH17-340M24.3</code>, though can easily found on NCBI website, but I haven't find a strightforward way to map it to GenID. </p>",
      "rawMarkdown": "Have you checked: https://docs.mygene.info/projects/mygene-py/en/latest/\nI want to map gene symbols to NCBI GenID. However, I find not easy way to do that in python. For example this one `CH17-340M24.3`, though can easily found on NCBI website, but I haven't find a strightforward way to map it to GenID.",
      "votes": null
    },
    {
      "id": "2504539",
      "postDate": "10/30/2023 01:15:56",
      "content": "<p>Thanks for telling us about the useful tools. But unfortunately, I cannot find some gene symbols in the mygene database.</p>\n<p>Like below:</p>\n<pre><code>mg.query(, species=)\n\n</code></pre>\n<p>Some of them are not an exact match, so it is difficult to determine what gene symbol is</p>\n<pre><code>mg.query(, species=)\n</code></pre>\n<pre><code>'took' \n 'total' \n 'max_score' \n 'hits' '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'ZNF503 antisense RNA '\n   'symbol' 'ZNF503-AS1'\n   'taxid' \n  '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'ZNF503 antisense RNA '\n   'symbol' 'ZNF503-AS2'\n   'taxid' \n  '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'zinc finger protein '\n   'symbol' 'ZNF503'\n   'taxid' \n</code></pre>\n<p>To create this dataset, there must be reference data like gtf or gff. I would like to know what reference data was used to create this dataset.</p>",
      "rawMarkdown": "Thanks for telling us about the useful tools. But unfortunately, I cannot find some gene symbols in the mygene database.\n\nLike below:\n```python\nmg.query(\"AL031281.3\", species=\"human\")\n# {'took': 26, 'total': 0, 'max_score': None, 'hits': []}\n```\n\nSome of them are not an exact match, so it is difficult to determine what gene symbol is\n\n```python\nmg.query(\"ZNF503-1\", species=\"human\")\n```\n\n```json\n{'took': 48,\n 'total': 3,\n 'max_score': 20.526493,\n 'hits': [{'_id': '253264',\n   '_score': 20.526493,\n   'entrezgene': '253264',\n   'name': 'ZNF503 antisense RNA 1',\n   'symbol': 'ZNF503-AS1',\n   'taxid': 9606},\n  {'_id': '100131213',\n   '_score': 19.651655,\n   'entrezgene': '100131213',\n   'name': 'ZNF503 antisense RNA 2',\n   'symbol': 'ZNF503-AS2',\n   'taxid': 9606},\n  {'_id': '84858',\n   '_score': 6.26353,\n   'entrezgene': '84858',\n   'name': 'zinc finger protein 503',\n   'symbol': 'ZNF503',\n   'taxid': 9606}]}\n```\n\nTo create this dataset, there must be reference data like gtf or gff. I would like to know what reference data was used to create this dataset.",
      "votes": null
    },
    {
      "id": "2510628",
      "postDate": "11/03/2023 06:41:51",
      "content": "<p><a href=\"https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv\" target=\"_blank\">https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv</a></p>\n<p>here is a gene info table containing 1-1 mapping of ensembl_id and gene symbol.</p>",
      "rawMarkdown": "https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv\n\nhere is a gene info table containing 1-1 mapping of ensembl_id and gene symbol.",
      "votes": null
    },
    {
      "id": "2525950",
      "postDate": "11/15/2023 13:45:52",
      "content": "<p>I also struggled with this - it turns out that the target (gene) IDs in the de_train table are not HGNC_symbols but external gene names. After doing a sweeping, there was a good match with GRCh38 release94 - only 80 column IDs could not be found. They all contained a -[0-9]* pattern at the end of strings, otherwise they looked like gene_symbols. After removing this bit after the hyphen, this remaining 80 could be also mapped to GRCh38v94. </p>\n<p>Please find the modified target_ids and the corresponding ENSEMBL_ID to target_id tables here <br>\n(the 'external_gene_name_cap' column to look for the target_id - there are 32 cases where multiple EMSEMBL_ID map to the target_id name):</p>\n<p><a href=\"https://github.com/Xaft/kaggle/tree/main/sc%20perturbation\" target=\"_blank\">https://github.com/Xaft/kaggle/tree/main/sc%20perturbation</a> </p>",
      "rawMarkdown": "I also struggled with this - it turns out that the target (gene) IDs in the de_train table are not HGNC_symbols but external gene names. After doing a sweeping, there was a good match with GRCh38 release94 - only 80 column IDs could not be found. They all contained a -[0-9]* pattern at the end of strings, otherwise they looked like gene_symbols. After removing this bit after the hyphen, this remaining 80 could be also mapped to GRCh38v94. \n\nPlease find the modified target_ids and the corresponding ENSEMBL_ID to target_id tables here \n(the 'external_gene_name_cap' column to look for the target_id - there are 32 cases where multiple EMSEMBL_ID map to the target_id name):\n\nhttps://github.com/Xaft/kaggle/tree/main/sc%20perturbation",
      "votes": null
    },
    {
      "id": "2526089",
      "postDate": "11/15/2023 15:16:20",
      "content": "<p>Just in case here is some example to work with Python \"mygene\" on Kaggle:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mygene-package-work-example-1\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mygene-package-work-example-1</a></p>",
      "rawMarkdown": "Just in case here is some example to work with Python \"mygene\" on Kaggle:\nhttps://www.kaggle.com/code/alexandervc/mygene-package-work-example-1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2496097,
      "author_name": "entron",
      "author_url": "",
      "post_date": "10/23/2023 18:13:51",
      "content": "<p>Have you checked: <a href=\"https://docs.mygene.info/projects/mygene-py/en/latest/\" target=\"_blank\">https://docs.mygene.info/projects/mygene-py/en/latest/</a><br>\nI want to map gene symbols to NCBI GenID. However, I find not easy way to do that in python. For example this one <code>CH17-340M24.3</code>, though can easily found on NCBI website, but I haven't find a strightforward way to map it to GenID. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2504539,
          "author_name": "illuminationk27",
          "author_url": "",
          "post_date": "10/30/2023 01:15:56",
          "content": "<p>Thanks for telling us about the useful tools. But unfortunately, I cannot find some gene symbols in the mygene database.</p>\n<p>Like below:</p>\n<pre><code>mg.query(, species=)\n\n</code></pre>\n<p>Some of them are not an exact match, so it is difficult to determine what gene symbol is</p>\n<pre><code>mg.query(, species=)\n</code></pre>\n<pre><code>'took' \n 'total' \n 'max_score' \n 'hits' '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'ZNF503 antisense RNA '\n   'symbol' 'ZNF503-AS1'\n   'taxid' \n  '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'ZNF503 antisense RNA '\n   'symbol' 'ZNF503-AS2'\n   'taxid' \n  '_id' ''\n   '_score' \n   'entrezgene' ''\n   'name' 'zinc finger protein '\n   'symbol' 'ZNF503'\n   'taxid' \n</code></pre>\n<p>To create this dataset, there must be reference data like gtf or gff. I would like to know what reference data was used to create this dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2510628,
      "author_name": "superdanielshao",
      "author_url": "",
      "post_date": "11/03/2023 06:41:51",
      "content": "<p><a href=\"https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv\" target=\"_blank\">https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv</a></p>\n<p>here is a gene info table containing 1-1 mapping of ensembl_id and gene symbol.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2525950,
      "author_name": "imregaspar",
      "author_url": "",
      "post_date": "11/15/2023 13:45:52",
      "content": "<p>I also struggled with this - it turns out that the target (gene) IDs in the de_train table are not HGNC_symbols but external gene names. After doing a sweeping, there was a good match with GRCh38 release94 - only 80 column IDs could not be found. They all contained a -[0-9]* pattern at the end of strings, otherwise they looked like gene_symbols. After removing this bit after the hyphen, this remaining 80 could be also mapped to GRCh38v94. </p>\n<p>Please find the modified target_ids and the corresponding ENSEMBL_ID to target_id tables here <br>\n(the 'external_gene_name_cap' column to look for the target_id - there are 32 cases where multiple EMSEMBL_ID map to the target_id name):</p>\n<p><a href=\"https://github.com/Xaft/kaggle/tree/main/sc%20perturbation\" target=\"_blank\">https://github.com/Xaft/kaggle/tree/main/sc%20perturbation</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2526089,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/15/2023 15:16:20",
      "content": "<p>Just in case here is some example to work with Python \"mygene\" on Kaggle:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mygene-package-work-example-1\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mygene-package-work-example-1</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2491816": "Is there any good way to get the mapping of gene symbols in `de_train.parquet` to Ensembl gene id?\n\nIn 10x cellranger output, a file like \"features.tsv.gz\" maps gene symbols to Ensembl gene ids. Is there a mapping like \"features.tsv.gz\"? \n\nI tried to download \"2020-A\" from https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads; refer to [this discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/440959). I parsed `genes.gtf`, but about 2000 gene symbols cannot convert to Emsenbl gene ids. \n\nI used this very simple script to check gene symbols:\n\n```python\nimport polars as pl\nde_df = pl.read_parquet(\"./data/de_train.parquet\")\ngenes = set(de_df.columns[5:])\n\ndef read_gtf(path: str) -> dict[str, str]:\n    genes_map = {}\n    with open(path) as f:\n        for i, line in enumerate(f):\n            if line.startswith(\"#\"):\n                continue\n\n            line = line.strip()\n            elems = line.split(\"\\t\")\n            attrs = elems[8]\n            d = {}\n            for attr in attrs.split(\";\"):\n                attr = attr.strip()\n                if len(attr) == 0:\n                    continue\n                k, v = attr.split(\" \")\n                k = k.strip(\"\\\"\")\n                v = v.strip(\"\\\"\")\n                d[k] = v\n\n            genes_map[d[\"gene_id\"]] = d[\"gene_name\"]\n\n    return genes_map\n\ngenes_map = read_gtf(\"./refdata-gex-GRCh38-2020-A/genes/genes.gtf\")\n\nprint(len(set(genes) - set(genes_map.values())))\n# 2399\n```",
    "2496097": "Have you checked: https://docs.mygene.info/projects/mygene-py/en/latest/\nI want to map gene symbols to NCBI GenID. However, I find not easy way to do that in python. For example this one `CH17-340M24.3`, though can easily found on NCBI website, but I haven't find a strightforward way to map it to GenID.",
    "2504539": "Thanks for telling us about the useful tools. But unfortunately, I cannot find some gene symbols in the mygene database.\n\nLike below:\n```python\nmg.query(\"AL031281.3\", species=\"human\")\n# {'took': 26, 'total': 0, 'max_score': None, 'hits': []}\n```\n\nSome of them are not an exact match, so it is difficult to determine what gene symbol is\n\n```python\nmg.query(\"ZNF503-1\", species=\"human\")\n```\n\n```json\n{'took': 48,\n 'total': 3,\n 'max_score': 20.526493,\n 'hits': [{'_id': '253264',\n   '_score': 20.526493,\n   'entrezgene': '253264',\n   'name': 'ZNF503 antisense RNA 1',\n   'symbol': 'ZNF503-AS1',\n   'taxid': 9606},\n  {'_id': '100131213',\n   '_score': 19.651655,\n   'entrezgene': '100131213',\n   'name': 'ZNF503 antisense RNA 2',\n   'symbol': 'ZNF503-AS2',\n   'taxid': 9606},\n  {'_id': '84858',\n   '_score': 6.26353,\n   'entrezgene': '84858',\n   'name': 'zinc finger protein 503',\n   'symbol': 'ZNF503',\n   'taxid': 9606}]}\n```\n\nTo create this dataset, there must be reference data like gtf or gff. I would like to know what reference data was used to create this dataset.",
    "2510628": "https://huggingface.co/datasets/ctheodoris/Genecorpus-30M/blob/main/example_input_files/gene_info_table.csv\n\nhere is a gene info table containing 1-1 mapping of ensembl_id and gene symbol.",
    "2525950": "I also struggled with this - it turns out that the target (gene) IDs in the de_train table are not HGNC_symbols but external gene names. After doing a sweeping, there was a good match with GRCh38 release94 - only 80 column IDs could not be found. They all contained a -[0-9]* pattern at the end of strings, otherwise they looked like gene_symbols. After removing this bit after the hyphen, this remaining 80 could be also mapped to GRCh38v94. \n\nPlease find the modified target_ids and the corresponding ENSEMBL_ID to target_id tables here \n(the 'external_gene_name_cap' column to look for the target_id - there are 32 cases where multiple EMSEMBL_ID map to the target_id name):\n\nhttps://github.com/Xaft/kaggle/tree/main/sc%20perturbation",
    "2526089": "Just in case here is some example to work with Python \"mygene\" on Kaggle:\nhttps://www.kaggle.com/code/alexandervc/mygene-package-work-example-1"
  },
  "source": "meta"
}