{
  "id": 349559,
  "title": "Associating ATAC peaks to Genes",
  "url": "/competitions/open-problems-multimodal/discussion/349559",
  "author_name": "",
  "post_date": "2022-09-01T18:50:10.572546Z",
  "votes": 42,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It was great seeing this discussion about linking CITE variables and gene names. I also want to throw out that it's possible to map the ATAC data to genes.</p>\n<p>You can download the genome reference data used for the Multiome data from 10x:</p>\n<pre><code>curl -O https://cf.10xgenomics.com/supp/cell-exp/refdata-gex-GRCh38-2020-A.tar.gz\n</code></pre>\n<p>When you unpack that tarball, you'll find a GTF file called <code>genes.gtf</code> that has information about the genes associated with each genomic region. You could then use the following code to get close features.</p>\n<pre><code>## Load GTF file\n\nimport pandas as pd\nimport numpy as np\n\ntenx_gtf = pd.read_table(\"/home/jovyan/data/refdata-gex-GRCh38-2020-A/genes/genes.gtf\", skiprows=5, header=None)\n\nGTF_HEADER  = ['seqname', 'source', 'feature', 'start', 'end', 'score',\n               'strand', 'frame', 'info']\n\ntenx_gtf.columns = GTF_HEADER\n\n## Helper function to find close genes\ndef find_close_features(peak_coords, gtf, feature_type=\"gene\", max_distance=10000):\n    '''Takes a ATAC peak_coords like chr1:100029425-100029883 and a GTF file and returns genes within max_distance from the peak'''\n\n    chrom = peak_coords.split(\":\")[0]\n    start = int(peak_coords.split(\":\")[1].split(\"-\")[0])\n    end   = int(peak_coords.split(\":\")[1].split(\"-\")[1])\n\n    # Slice the GTF to the relevant chromosome\n    chrom_mask = (gtf[\"seqname\"] == chrom) &amp; (gtf[\"feature\"] == feature_type)\n    genes_on_chrom = gtf[chrom_mask]\n\n    # Calculate distances to the same chromosome\n    dists = np.min(\n    (\n        np.abs(start - genes_on_chrom[\"start\"]),\n        np.abs(start - genes_on_chrom[\"end\"]) , \n        np.abs(end - genes_on_chrom[\"start\"]), \n        np.abs(end - genes_on_chrom[\"end\"])\n    ), axis=0\n    )\n\n    # Filter through features less than max_distance\n    feature_names = []\n    for gene in genes_on_chrom[dists &lt; max_distance][\"info\"]:\n        gene_info = {}\n        for keyval in  gene.split(\"; \"):\n            key, val = keyval.split(\" \")\n            if isinstance(val, str):\n                val = val.replace('\"',\"\")\n                val = val.rstrip(\";\")\n            gene_info[key] = val\n        feature_names.append(gene_info[\"gene_name\"])\n\n    return feature_names\n\n## Example Usage\nfind_close_features( \"chr1:100029425-100029883\", tenx_gtf)\n</code></pre>\n<p>Which will return <code>['SLC35A3', 'MFSD14A']</code> for this example.</p>\n<p>Let me know if this is helpful, I could throw this in a notebook.</p>",
  "messages": [
    {
      "id": "1922855",
      "postDate": "09/01/2022 18:50:10",
      "content": "<p>It was great seeing this discussion about linking CITE variables and gene names. I also want to throw out that it's possible to map the ATAC data to genes.</p>\n<p>You can download the genome reference data used for the Multiome data from 10x:</p>\n<pre><code>curl -O https://cf.10xgenomics.com/supp/cell-exp/refdata-gex-GRCh38-2020-A.tar.gz\n</code></pre>\n<p>When you unpack that tarball, you'll find a GTF file called <code>genes.gtf</code> that has information about the genes associated with each genomic region. You could then use the following code to get close features.</p>\n<pre><code>## Load GTF file\n\nimport pandas as pd\nimport numpy as np\n\ntenx_gtf = pd.read_table(\"/home/jovyan/data/refdata-gex-GRCh38-2020-A/genes/genes.gtf\", skiprows=5, header=None)\n\nGTF_HEADER  = ['seqname', 'source', 'feature', 'start', 'end', 'score',\n               'strand', 'frame', 'info']\n\ntenx_gtf.columns = GTF_HEADER\n\n## Helper function to find close genes\ndef find_close_features(peak_coords, gtf, feature_type=\"gene\", max_distance=10000):\n    '''Takes a ATAC peak_coords like chr1:100029425-100029883 and a GTF file and returns genes within max_distance from the peak'''\n\n    chrom = peak_coords.split(\":\")[0]\n    start = int(peak_coords.split(\":\")[1].split(\"-\")[0])\n    end   = int(peak_coords.split(\":\")[1].split(\"-\")[1])\n\n    # Slice the GTF to the relevant chromosome\n    chrom_mask = (gtf[\"seqname\"] == chrom) &amp; (gtf[\"feature\"] == feature_type)\n    genes_on_chrom = gtf[chrom_mask]\n\n    # Calculate distances to the same chromosome\n    dists = np.min(\n    (\n        np.abs(start - genes_on_chrom[\"start\"]),\n        np.abs(start - genes_on_chrom[\"end\"]) , \n        np.abs(end - genes_on_chrom[\"start\"]), \n        np.abs(end - genes_on_chrom[\"end\"])\n    ), axis=0\n    )\n\n    # Filter through features less than max_distance\n    feature_names = []\n    for gene in genes_on_chrom[dists &lt; max_distance][\"info\"]:\n        gene_info = {}\n        for keyval in  gene.split(\"; \"):\n            key, val = keyval.split(\" \")\n            if isinstance(val, str):\n                val = val.replace('\"',\"\")\n                val = val.rstrip(\";\")\n            gene_info[key] = val\n        feature_names.append(gene_info[\"gene_name\"])\n\n    return feature_names\n\n## Example Usage\nfind_close_features( \"chr1:100029425-100029883\", tenx_gtf)\n</code></pre>\n<p>Which will return <code>['SLC35A3', 'MFSD14A']</code> for this example.</p>\n<p>Let me know if this is helpful, I could throw this in a notebook.</p>",
      "rawMarkdown": "It was great seeing this discussion about linking CITE variables and gene names. I also want to throw out that it's possible to map the ATAC data to genes.\n\nYou can download the genome reference data used for the Multiome data from 10x:\n\n```\ncurl -O https://cf.10xgenomics.com/supp/cell-exp/refdata-gex-GRCh38-2020-A.tar.gz\n```\n\nWhen you unpack that tarball, you'll find a GTF file called `genes.gtf` that has information about the genes associated with each genomic region. You could then use the following code to get close features.\n\n```python\n## Load GTF file\n\nimport pandas as pd\nimport numpy as np\n\ntenx_gtf = pd.read_table(\"/home/jovyan/data/refdata-gex-GRCh38-2020-A/genes/genes.gtf\", skiprows=5, header=None)\n\nGTF_HEADER  = ['seqname', 'source', 'feature', 'start', 'end', 'score',\n               'strand', 'frame', 'info']\n\ntenx_gtf.columns = GTF_HEADER\n\n## Helper function to find close genes\ndef find_close_features(peak_coords, gtf, feature_type=\"gene\", max_distance=10000):\n    '''Takes a ATAC peak_coords like chr1:100029425-100029883 and a GTF file and returns genes within max_distance from the peak'''\n    \n    chrom = peak_coords.split(\":\")[0]\n    start = int(peak_coords.split(\":\")[1].split(\"-\")[0])\n    end   = int(peak_coords.split(\":\")[1].split(\"-\")[1])\n    \n    # Slice the GTF to the relevant chromosome\n    chrom_mask = (gtf[\"seqname\"] == chrom) & (gtf[\"feature\"] == feature_type)\n    genes_on_chrom = gtf[chrom_mask]\n    \n    # Calculate distances to the same chromosome\n    dists = np.min(\n    (\n        np.abs(start - genes_on_chrom[\"start\"]),\n        np.abs(start - genes_on_chrom[\"end\"]) , \n        np.abs(end - genes_on_chrom[\"start\"]), \n        np.abs(end - genes_on_chrom[\"end\"])\n    ), axis=0\n    )\n    \n    # Filter through features less than max_distance\n    feature_names = []\n    for gene in genes_on_chrom[dists < max_distance][\"info\"]:\n        gene_info = {}\n        for keyval in  gene.split(\"; \"):\n            key, val = keyval.split(\" \")\n            if isinstance(val, str):\n                val = val.replace('\"',\"\")\n                val = val.rstrip(\";\")\n            gene_info[key] = val\n        feature_names.append(gene_info[\"gene_name\"])\n    \n    return feature_names\n\n## Example Usage\nfind_close_features( \"chr1:100029425-100029883\", tenx_gtf)\n```\nWhich will return `['SLC35A3', 'MFSD14A']` for this example.\n\nLet me know if this is helpful, I could throw this in a notebook.",
      "votes": null
    },
    {
      "id": "1922891",
      "postDate": "09/01/2022 19:51:49",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Thank you ! Yes, please put it in a notebook !</p>\n<p>PS<br>\nThe discussion you probably mean that discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617</a></p>",
      "rawMarkdown": "danielburkhardt Thank you ! Yes, please put it in a notebook !\n\nPS\nThe discussion you probably mean that discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617",
      "votes": null
    },
    {
      "id": "2011848",
      "postDate": "10/31/2022 21:06:04",
      "content": "<p>I have seen that the suffix (A of A_B) of the 22k columns inputs correspond to 18k names of targets multiome (rna too).</p>",
      "rawMarkdown": "I have seen that the suffix (A of A_B) of the 22k columns inputs correspond to 18k names of targets multiome (rna too).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1922891,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/01/2022 19:51:49",
      "content": "<p><a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> Thank you ! Yes, please put it in a notebook !</p>\n<p>PS<br>\nThe discussion you probably mean that discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2011848,
      "author_name": "pierretisseur",
      "author_url": "",
      "post_date": "10/31/2022 21:06:04",
      "content": "<p>I have seen that the suffix (A of A_B) of the 22k columns inputs correspond to 18k names of targets multiome (rna too).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1922855": "It was great seeing this discussion about linking CITE variables and gene names. I also want to throw out that it's possible to map the ATAC data to genes.\n\nYou can download the genome reference data used for the Multiome data from 10x:\n\n```\ncurl -O https://cf.10xgenomics.com/supp/cell-exp/refdata-gex-GRCh38-2020-A.tar.gz\n```\n\nWhen you unpack that tarball, you'll find a GTF file called `genes.gtf` that has information about the genes associated with each genomic region. You could then use the following code to get close features.\n\n```python\n## Load GTF file\n\nimport pandas as pd\nimport numpy as np\n\ntenx_gtf = pd.read_table(\"/home/jovyan/data/refdata-gex-GRCh38-2020-A/genes/genes.gtf\", skiprows=5, header=None)\n\nGTF_HEADER  = ['seqname', 'source', 'feature', 'start', 'end', 'score',\n               'strand', 'frame', 'info']\n\ntenx_gtf.columns = GTF_HEADER\n\n## Helper function to find close genes\ndef find_close_features(peak_coords, gtf, feature_type=\"gene\", max_distance=10000):\n    '''Takes a ATAC peak_coords like chr1:100029425-100029883 and a GTF file and returns genes within max_distance from the peak'''\n    \n    chrom = peak_coords.split(\":\")[0]\n    start = int(peak_coords.split(\":\")[1].split(\"-\")[0])\n    end   = int(peak_coords.split(\":\")[1].split(\"-\")[1])\n    \n    # Slice the GTF to the relevant chromosome\n    chrom_mask = (gtf[\"seqname\"] == chrom) & (gtf[\"feature\"] == feature_type)\n    genes_on_chrom = gtf[chrom_mask]\n    \n    # Calculate distances to the same chromosome\n    dists = np.min(\n    (\n        np.abs(start - genes_on_chrom[\"start\"]),\n        np.abs(start - genes_on_chrom[\"end\"]) , \n        np.abs(end - genes_on_chrom[\"start\"]), \n        np.abs(end - genes_on_chrom[\"end\"])\n    ), axis=0\n    )\n    \n    # Filter through features less than max_distance\n    feature_names = []\n    for gene in genes_on_chrom[dists < max_distance][\"info\"]:\n        gene_info = {}\n        for keyval in  gene.split(\"; \"):\n            key, val = keyval.split(\" \")\n            if isinstance(val, str):\n                val = val.replace('\"',\"\")\n                val = val.rstrip(\";\")\n            gene_info[key] = val\n        feature_names.append(gene_info[\"gene_name\"])\n    \n    return feature_names\n\n## Example Usage\nfind_close_features( \"chr1:100029425-100029883\", tenx_gtf)\n```\nWhich will return `['SLC35A3', 'MFSD14A']` for this example.\n\nLet me know if this is helpful, I could throw this in a notebook.",
    "1922891": "danielburkhardt Thank you ! Yes, please put it in a notebook !\n\nPS\nThe discussion you probably mean that discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/347617",
    "2011848": "I have seen that the suffix (A of A_B) of the 22k columns inputs correspond to 18k names of targets multiome (rna too)."
  },
  "source": "meta"
}