{
  "id": 347617,
  "title": "What are the coordinates of the ATAC-seq peaks?",
  "url": "/competitions/open-problems-multimodal/discussion/347617",
  "author_name": "",
  "post_date": "2022-08-24T20:42:18.230742700Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Is there any more detailed description about how the <code>train/test_multi_inputs.h5</code> is generated? I expect to see some genomic coordinates associated these peaks so that we know how the peak looks like in gene promoter or 3UTR? </p>\n<p>I've never processed single-cell ATAC-seq data before, but that's something I expect?</p>\n<p>Thanks,<br>\nYichao</p>",
  "messages": [
    {
      "id": "1912580",
      "postDate": "08/24/2022 20:42:18",
      "content": "<p>Hi,</p>\n<p>Is there any more detailed description about how the <code>train/test_multi_inputs.h5</code> is generated? I expect to see some genomic coordinates associated these peaks so that we know how the peak looks like in gene promoter or 3UTR? </p>\n<p>I've never processed single-cell ATAC-seq data before, but that's something I expect?</p>\n<p>Thanks,<br>\nYichao</p>",
      "rawMarkdown": "Hi,\n\nIs there any more detailed description about how the `train/test_multi_inputs.h5` is generated? I expect to see some genomic coordinates associated these peaks so that we know how the peak looks like in gene promoter or 3UTR? \n\nI've never processed single-cell ATAC-seq data before, but that's something I expect?\n\nThanks,\nYichao",
      "votes": null
    },
    {
      "id": "1913691",
      "postDate": "08/25/2022 12:56:29",
      "content": "<p>Yup, you can find the feature information in the associated <code>.h5</code> columns for the multiome. The description of those columns can be found in the Data Description tab under File and Field descriptions:</p>\n<blockquote>\n  <p>train/test_multi_inputs.h5 - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and <strong>columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).</strong></p>\n</blockquote>",
      "rawMarkdown": "Yup, you can find the feature information in the associated `.h5` columns for the multiome. The description of those columns can be found in the Data Description tab under File and Field descriptions:\n\n> train/test_multi_inputs.h5 - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and **columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).**",
      "votes": null
    },
    {
      "id": "1913797",
      "postDate": "08/25/2022 14:08:15",
      "content": "<p>Thanks for the quick reply. Yes, I saw that. But the columns are gene names, e.g., ENSG00000121410. It's very likely that in one gene, there are multiple peaks. And what about intergenic regions?</p>\n<p>So the ATAC-seq peak counts, if without TF-IDF, are kind of averaged peaks per gene?</p>",
      "rawMarkdown": "Thanks for the quick reply. Yes, I saw that. But the columns are gene names, e.g., ENSG00000121410. It's very likely that in one gene, there are multiple peaks. And what about intergenic regions?\n\nSo the ATAC-seq peak counts, if without TF-IDF, are kind of averaged peaks per gene?",
      "votes": null
    },
    {
      "id": "1914142",
      "postDate": "08/25/2022 19:05:11",
      "content": "<p>Hi,</p>\n<p>actually, those starting with \"ENSG\" are columns in gene expression data, not DNA accessibility (ATAC) data.</p>\n<p>Among the 228942 columns in ATAC data, 228878 have names like \"chr[NUMBER]:START:END\". A small sample:</p>\n<p>['chr20:59765313-59766220', 'chr11:19250190-19251086', 'chr6:112001912-112002831', 'chr3:11258639-11259424', 'chr17:27457566-27458329'].</p>",
      "rawMarkdown": "Hi,\n\nactually, those starting with \"ENSG\" are columns in gene expression data, not DNA accessibility (ATAC) data.\n\nAmong the 228942 columns in ATAC data, 228878 have names like \"chr[NUMBER]:START:END\". A small sample:\n\n['chr20:59765313-59766220', 'chr11:19250190-19251086', 'chr6:112001912-112002831', 'chr3:11258639-11259424', 'chr17:27457566-27458329'].",
      "votes": null
    },
    {
      "id": "1928922",
      "postDate": "09/06/2022 17:43:32",
      "content": "<p>Thanks! How did I miss that …</p>",
      "rawMarkdown": "Thanks! How did I miss that ...",
      "votes": null
    },
    {
      "id": "2011835",
      "postDate": "10/31/2022 20:34:53",
      "content": "<p>I have seen that the suffixe A of A_B name of citeSeq inputs (rna) (22k) correspond to 18k of multiome columns (rna target ). I wonder why the name in citeseq is longer. </p>",
      "rawMarkdown": "I have seen that the suffixe A of A_B name of citeSeq inputs (rna) (22k) correspond to 18k of multiome columns (rna target ). I wonder why the name in citeseq is longer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913691,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "08/25/2022 12:56:29",
      "content": "<p>Yup, you can find the feature information in the associated <code>.h5</code> columns for the multiome. The description of those columns can be found in the Data Description tab under File and Field descriptions:</p>\n<blockquote>\n  <p>train/test_multi_inputs.h5 - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and <strong>columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).</strong></p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1913797,
          "author_name": "unfashionable",
          "author_url": "",
          "post_date": "08/25/2022 14:08:15",
          "content": "<p>Thanks for the quick reply. Yes, I saw that. But the columns are gene names, e.g., ENSG00000121410. It's very likely that in one gene, there are multiple peaks. And what about intergenic regions?</p>\n<p>So the ATAC-seq peak counts, if without TF-IDF, are kind of averaged peaks per gene?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914142,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/25/2022 19:05:11",
          "content": "<p>Hi,</p>\n<p>actually, those starting with \"ENSG\" are columns in gene expression data, not DNA accessibility (ATAC) data.</p>\n<p>Among the 228942 columns in ATAC data, 228878 have names like \"chr[NUMBER]:START:END\". A small sample:</p>\n<p>['chr20:59765313-59766220', 'chr11:19250190-19251086', 'chr6:112001912-112002831', 'chr3:11258639-11259424', 'chr17:27457566-27458329'].</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1928922,
          "author_name": "unfashionable",
          "author_url": "",
          "post_date": "09/06/2022 17:43:32",
          "content": "<p>Thanks! How did I miss that …</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2011835,
      "author_name": "pierretisseur",
      "author_url": "",
      "post_date": "10/31/2022 20:34:53",
      "content": "<p>I have seen that the suffixe A of A_B name of citeSeq inputs (rna) (22k) correspond to 18k of multiome columns (rna target ). I wonder why the name in citeseq is longer. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912580": "Hi,\n\nIs there any more detailed description about how the `train/test_multi_inputs.h5` is generated? I expect to see some genomic coordinates associated these peaks so that we know how the peak looks like in gene promoter or 3UTR? \n\nI've never processed single-cell ATAC-seq data before, but that's something I expect?\n\nThanks,\nYichao",
    "1913691": "Yup, you can find the feature information in the associated `.h5` columns for the multiome. The description of those columns can be found in the Data Description tab under File and Field descriptions:\n\n> train/test_multi_inputs.h5 - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and **columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).**",
    "1913797": "Thanks for the quick reply. Yes, I saw that. But the columns are gene names, e.g., ENSG00000121410. It's very likely that in one gene, there are multiple peaks. And what about intergenic regions?\n\nSo the ATAC-seq peak counts, if without TF-IDF, are kind of averaged peaks per gene?",
    "1914142": "Hi,\n\nactually, those starting with \"ENSG\" are columns in gene expression data, not DNA accessibility (ATAC) data.\n\nAmong the 228942 columns in ATAC data, 228878 have names like \"chr[NUMBER]:START:END\". A small sample:\n\n['chr20:59765313-59766220', 'chr11:19250190-19251086', 'chr6:112001912-112002831', 'chr3:11258639-11259424', 'chr17:27457566-27458329'].",
    "1928922": "Thanks! How did I miss that ...",
    "2011835": "I have seen that the suffixe A of A_B name of citeSeq inputs (rna) (22k) correspond to 18k of multiome columns (rna target ). I wonder why the name in citeseq is longer."
  },
  "source": "meta"
}