{
  "id": 450439,
  "title": "ATAC-seq analysis",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/450439",
  "author_name": "",
  "post_date": "2023-10-24T12:39:42.295021400Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all,<br>\nAfter exploring with little success SMILES embedding and gene ontology analysis, I am now focusing on adding the ATAC-seq dataset to the features for each cell type. I hope to use the baseline ATAC-seq data as the embedding of the cell types' epigenetic state. <br>\nIt would be nice to have a topic dedicated to it if people want to brainstorm. <br>\nI am currently exploring the notebooks from the last competition.</p>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453</a></p>\n<p><a href=\"https://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried\" target=\"_blank\">https://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried</a></p>",
  "messages": [
    {
      "id": "2497074",
      "postDate": "10/24/2023 12:39:42",
      "content": "<p>Hi all,<br>\nAfter exploring with little success SMILES embedding and gene ontology analysis, I am now focusing on adding the ATAC-seq dataset to the features for each cell type. I hope to use the baseline ATAC-seq data as the embedding of the cell types' epigenetic state. <br>\nIt would be nice to have a topic dedicated to it if people want to brainstorm. <br>\nI am currently exploring the notebooks from the last competition.</p>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453</a></p>\n<p><a href=\"https://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried\" target=\"_blank\">https://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried</a></p>",
      "rawMarkdown": "Hi all,\nAfter exploring with little success SMILES embedding and gene ontology analysis, I am now focusing on adding the ATAC-seq dataset to the features for each cell type. I hope to use the baseline ATAC-seq data as the embedding of the cell types' epigenetic state. \nIt would be nice to have a topic dedicated to it if people want to brainstorm. \nI am currently exploring the notebooks from the last competition.\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453\n\nhttps://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried",
      "votes": null
    },
    {
      "id": "2497147",
      "postDate": "10/24/2023 13:34:51",
      "content": "<p>Thanks for posting ! <br>\nIn my mind that is one of the pain-points how to incorporate the ATAC-seq and other data.</p>\n<p>The only idea comes to mind so far - try to created aggregates over cell type, and separeately over compounds of any data related to ATAC-seq - that can be used as features for the present challenge.</p>\n<p>For example we can take PCA or some signatures from ATAC-seq data, then groupby(cell_type/compound).aggergete( ) - and we get some features. </p>\n<p>But have not tried that yet. </p>",
      "rawMarkdown": "Thanks for posting ! \nIn my mind that is one of the pain-points how to incorporate the ATAC-seq and other data.\n\nThe only idea comes to mind so far - try to created aggregates over cell type, and separeately over compounds of any data related to ATAC-seq - that can be used as features for the present challenge.\n\nFor example we can take PCA or some signatures from ATAC-seq data, then groupby(cell_type/compound).aggergete( ) - and we get some features. \n\nBut have not tried that yet.",
      "votes": null
    },
    {
      "id": "2498135",
      "postDate": "10/25/2023 06:52:26",
      "content": "<p>I used a MultiVI model to encode the GEX+ATAC expression into one latent space. The whole step can be found here.<br>\n<a href=\"https://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings\" target=\"_blank\">https://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F207a2bb08fc3f2a8398ae8b5ece11cda%2F41592_2023_1909_Fig1_HTML.webp?generation=1698216669301009&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2Fdcda296b97f6e41ec5840bd1a7e80a01%2Foutput.png?generation=1698216731946545&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I used a MultiVI model to encode the GEX+ATAC expression into one latent space. The whole step can be found here.\nhttps://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F207a2bb08fc3f2a8398ae8b5ece11cda%2F41592_2023_1909_Fig1_HTML.webp?generation=1698216669301009&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2Fdcda296b97f6e41ec5840bd1a7e80a01%2Foutput.png?generation=1698216731946545&alt=media)",
      "votes": null
    },
    {
      "id": "2498353",
      "postDate": "10/25/2023 08:56:54",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2502660",
      "postDate": "10/28/2023 11:18:10",
      "content": "<p>I've recently processed the ATAC-seq data and obtained results that were unexpected. Being relatively new to this, I'd appreciate insights from those with more experience. Notably, the FRIP score is quite low, and the TSS enrichment profile doesn't center as anticipated. Pls see the attachment.</p>",
      "rawMarkdown": "I've recently processed the ATAC-seq data and obtained results that were unexpected. Being relatively new to this, I'd appreciate insights from those with more experience. Notably, the FRIP score is quite low, and the TSS enrichment profile doesn't center as anticipated. Pls see the attachment.",
      "votes": null
    },
    {
      "id": "2524791",
      "postDate": "11/14/2023 14:33:03",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jalilnourisa\" target=\"_blank\">@jalilnourisa</a>,</p>\n<p>I find this video is a nice introduction to ATAC-seq. <br>\n<a href=\"https://www.youtube.com/watch?v=e2396GKFMRY\" target=\"_blank\">https://www.youtube.com/watch?v=e2396GKFMRY</a></p>\n<p>I am sure there are nice tools in Python too, like this one:<br>\n<a href=\"https://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html\" target=\"_blank\">https://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html</a></p>\n<p>Here is the notebook explaining how to convert toanndata:<br>\n<a href=\"https://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook\" target=\"_blank\">https://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook</a></p>\n<p>From what I have been reading, a low FRIP score could indicate either technical issues with the library preparation or issues during the peak calling process.</p>\n<p>Useful links I found:<br>\n<a href=\"https://www.biostars.org/p/9489225/\" target=\"_blank\">https://www.biostars.org/p/9489225/</a><br>\n<a href=\"https://www.encodeproject.org/atac-seq/\" target=\"_blank\">https://www.encodeproject.org/atac-seq/</a></p>\n<p>So, scores under 0.2 are concerning as they indicate either an issue with the bioinformatic pipeline or a technical issue.</p>\n<p>Here, the ATAC-seq peak counts were transformed with TF-IDF, which might explain why you get strange values. That would likely be an issue for the R and Python pipelines above.<br>\nSee a good link about TF-IDF transformation of peaks data: <a href=\"https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/\" target=\"_blank\">https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/</a>. </p>",
      "rawMarkdown": "Hello @jalilnourisa,\n\nI find this video is a nice introduction to ATAC-seq. \nhttps://www.youtube.com/watch?v=e2396GKFMRY\n\nI am sure there are nice tools in Python too, like this one:\nhttps://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html\n\nHere is the notebook explaining how to convert toanndata:\nhttps://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook\n\nFrom what I have been reading, a low FRIP score could indicate either technical issues with the library preparation or issues during the peak calling process.\n\nUseful links I found:\nhttps://www.biostars.org/p/9489225/\nhttps://www.encodeproject.org/atac-seq/\n\nSo, scores under 0.2 are concerning as they indicate either an issue with the bioinformatic pipeline or a technical issue.\n\nHere, the ATAC-seq peak counts were transformed with TF-IDF, which might explain why you get strange values. That would likely be an issue for the R and Python pipelines above.\nSee a good link about TF-IDF transformation of peaks data: https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/.",
      "votes": null
    },
    {
      "id": "2528181",
      "postDate": "11/17/2023 06:50:50",
      "content": "<p>Thank you for starting this discussion and the resources shared.  I would like to confirm that there is only baseline data, correct?  In other words, I have not seen ATAC-seq data for the cells after exposure to the small molecules.</p>",
      "rawMarkdown": "Thank you for starting this discussion and the resources shared.  I would like to confirm that there is only baseline data, correct?  In other words, I have not seen ATAC-seq data for the cells after exposure to the small molecules.",
      "votes": null
    },
    {
      "id": "2528894",
      "postDate": "11/17/2023 17:47:11",
      "content": "<p>Yes, you are correct. You only have the baseline. If you know which part of the genome is accessible (in a relaxed state chromatin region) for each cell type, you can potentially predict which genes are more likely to be expressed in a given cell type. </p>",
      "rawMarkdown": "Yes, you are correct. You only have the baseline. If you know which part of the genome is accessible (in a relaxed state chromatin region) for each cell type, you can potentially predict which genes are more likely to be expressed in a given cell type.",
      "votes": null
    },
    {
      "id": "2528927",
      "postDate": "11/17/2023 18:22:19",
      "content": "<p>Ok great, thanks for confirming. <br>\nI haven't figured out how to relate the Peaks data to a Cell type and Gene region. See file pasted here.  I'd appreciate input.</p>",
      "rawMarkdown": "Ok great, thanks for confirming. \nI haven't figured out how to relate the Peaks data to a Cell type and Gene region. See file pasted here.  I'd appreciate input.",
      "votes": null
    },
    {
      "id": "2528939",
      "postDate": "11/17/2023 18:36:55",
      "content": "<p>I am exploring several things. I have not yet found a straightforward method. The data here was pre-processed (peaks to TF-IDF). I am trying to see what methods will be compatible with that. </p>",
      "rawMarkdown": "I am exploring several things. I have not yet found a straightforward method. The data here was pre-processed (peaks to TF-IDF). I am trying to see what methods will be compatible with that.",
      "votes": null
    },
    {
      "id": "2530223",
      "postDate": "11/18/2023 22:59:41",
      "content": "<p>I am also trying to hook the multiome data to the training data. <br>\nI figured how to associate the peaks to individual genes. <br>\nI would like to associate the peaks to the cell types, but I don't know if that is possible. <br>\nI assume that the peaks were measured for each cell_type?<br>\nIf that is the case then how can you tell what cell type does a peak belongs to? </p>",
      "rawMarkdown": "I am also trying to hook the multiome data to the training data. \nI figured how to associate the peaks to individual genes. \nI would like to associate the peaks to the cell types, but I don't know if that is possible. \nI assume that the peaks were measured for each cell_type?\nIf that is the case then how can you tell what cell type does a peak belongs to?",
      "votes": null
    },
    {
      "id": "2530681",
      "postDate": "11/19/2023 12:11:41",
      "content": "<p>Thanks great idea</p>",
      "rawMarkdown": "Thanks great idea",
      "votes": null
    },
    {
      "id": "2531025",
      "postDate": "11/19/2023 19:21:01",
      "content": "<p>You can combine multiome_obs_meta with multiome_train using the obs_id as a key for an inner join. <br>\nThis will link intervals with cell type and donors. <br>\nHow did you associate the peaks to individual genes?</p>",
      "rawMarkdown": "You can combine multiome_obs_meta with multiome_train using the obs_id as a key for an inner join. \nThis will link intervals with cell type and donors. \nHow did you associate the peaks to individual genes?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2497147,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/24/2023 13:34:51",
      "content": "<p>Thanks for posting ! <br>\nIn my mind that is one of the pain-points how to incorporate the ATAC-seq and other data.</p>\n<p>The only idea comes to mind so far - try to created aggregates over cell type, and separeately over compounds of any data related to ATAC-seq - that can be used as features for the present challenge.</p>\n<p>For example we can take PCA or some signatures from ATAC-seq data, then groupby(cell_type/compound).aggergete( ) - and we get some features. </p>\n<p>But have not tried that yet. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2498135,
      "author_name": "superdanielshao",
      "author_url": "",
      "post_date": "10/25/2023 06:52:26",
      "content": "<p>I used a MultiVI model to encode the GEX+ATAC expression into one latent space. The whole step can be found here.<br>\n<a href=\"https://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings\" target=\"_blank\">https://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F207a2bb08fc3f2a8398ae8b5ece11cda%2F41592_2023_1909_Fig1_HTML.webp?generation=1698216669301009&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2Fdcda296b97f6e41ec5840bd1a7e80a01%2Foutput.png?generation=1698216731946545&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2498353,
          "author_name": "wguesdon",
          "author_url": "",
          "post_date": "10/25/2023 08:56:54",
          "content": "<p>Thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2502660,
      "author_name": "jalilnourisa",
      "author_url": "",
      "post_date": "10/28/2023 11:18:10",
      "content": "<p>I've recently processed the ATAC-seq data and obtained results that were unexpected. Being relatively new to this, I'd appreciate insights from those with more experience. Notably, the FRIP score is quite low, and the TSS enrichment profile doesn't center as anticipated. Pls see the attachment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2524791,
          "author_name": "wguesdon",
          "author_url": "",
          "post_date": "11/14/2023 14:33:03",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/jalilnourisa\" target=\"_blank\">@jalilnourisa</a>,</p>\n<p>I find this video is a nice introduction to ATAC-seq. <br>\n<a href=\"https://www.youtube.com/watch?v=e2396GKFMRY\" target=\"_blank\">https://www.youtube.com/watch?v=e2396GKFMRY</a></p>\n<p>I am sure there are nice tools in Python too, like this one:<br>\n<a href=\"https://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html\" target=\"_blank\">https://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html</a></p>\n<p>Here is the notebook explaining how to convert toanndata:<br>\n<a href=\"https://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook\" target=\"_blank\">https://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook</a></p>\n<p>From what I have been reading, a low FRIP score could indicate either technical issues with the library preparation or issues during the peak calling process.</p>\n<p>Useful links I found:<br>\n<a href=\"https://www.biostars.org/p/9489225/\" target=\"_blank\">https://www.biostars.org/p/9489225/</a><br>\n<a href=\"https://www.encodeproject.org/atac-seq/\" target=\"_blank\">https://www.encodeproject.org/atac-seq/</a></p>\n<p>So, scores under 0.2 are concerning as they indicate either an issue with the bioinformatic pipeline or a technical issue.</p>\n<p>Here, the ATAC-seq peak counts were transformed with TF-IDF, which might explain why you get strange values. That would likely be an issue for the R and Python pipelines above.<br>\nSee a good link about TF-IDF transformation of peaks data: <a href=\"https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/\" target=\"_blank\">https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/</a>. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2528181,
              "author_name": "rosariomollo",
              "author_url": "",
              "post_date": "11/17/2023 06:50:50",
              "content": "<p>Thank you for starting this discussion and the resources shared.  I would like to confirm that there is only baseline data, correct?  In other words, I have not seen ATAC-seq data for the cells after exposure to the small molecules.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2528894,
                  "author_name": "wguesdon",
                  "author_url": "",
                  "post_date": "11/17/2023 17:47:11",
                  "content": "<p>Yes, you are correct. You only have the baseline. If you know which part of the genome is accessible (in a relaxed state chromatin region) for each cell type, you can potentially predict which genes are more likely to be expressed in a given cell type. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2528927,
                      "author_name": "rosariomollo",
                      "author_url": "",
                      "post_date": "11/17/2023 18:22:19",
                      "content": "<p>Ok great, thanks for confirming. <br>\nI haven't figured out how to relate the Peaks data to a Cell type and Gene region. See file pasted here.  I'd appreciate input.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2528939,
                          "author_name": "wguesdon",
                          "author_url": "",
                          "post_date": "11/17/2023 18:36:55",
                          "content": "<p>I am exploring several things. I have not yet found a straightforward method. The data here was pre-processed (peaks to TF-IDF). I am trying to see what methods will be compatible with that. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2530223,
                              "author_name": "adsar0",
                              "author_url": "",
                              "post_date": "11/18/2023 22:59:41",
                              "content": "<p>I am also trying to hook the multiome data to the training data. <br>\nI figured how to associate the peaks to individual genes. <br>\nI would like to associate the peaks to the cell types, but I don't know if that is possible. <br>\nI assume that the peaks were measured for each cell_type?<br>\nIf that is the case then how can you tell what cell type does a peak belongs to? </p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2531025,
                                  "author_name": "wguesdon",
                                  "author_url": "",
                                  "post_date": "11/19/2023 19:21:01",
                                  "content": "<p>You can combine multiome_obs_meta with multiome_train using the obs_id as a key for an inner join. <br>\nThis will link intervals with cell type and donors. <br>\nHow did you associate the peaks to individual genes?</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2530681,
      "author_name": "edwin49",
      "author_url": "",
      "post_date": "11/19/2023 12:11:41",
      "content": "<p>Thanks great idea</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2497074": "Hi all,\nAfter exploring with little success SMILES embedding and gene ontology analysis, I am now focusing on adding the ATAC-seq dataset to the features for each cell type. I hope to use the baseline ATAC-seq data as the embedding of the cell types' epigenetic state. \nIt would be nice to have a topic dedicated to it if people want to brainstorm. \nI am currently exploring the notebooks from the last competition.\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/366453\n\nhttps://www.kaggle.com/code/llttyy/open-problem-biological-ideas/notebook#Other-ideas-I-have-tried",
    "2497147": "Thanks for posting ! \nIn my mind that is one of the pain-points how to incorporate the ATAC-seq and other data.\n\nThe only idea comes to mind so far - try to created aggregates over cell type, and separeately over compounds of any data related to ATAC-seq - that can be used as features for the present challenge.\n\nFor example we can take PCA or some signatures from ATAC-seq data, then groupby(cell_type/compound).aggergete( ) - and we get some features. \n\nBut have not tried that yet.",
    "2498135": "I used a MultiVI model to encode the GEX+ATAC expression into one latent space. The whole step can be found here.\nhttps://www.kaggle.com/code/superdanielshao/op2-multivi-for-additional-cell-embeddings\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F207a2bb08fc3f2a8398ae8b5ece11cda%2F41592_2023_1909_Fig1_HTML.webp?generation=1698216669301009&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2Fdcda296b97f6e41ec5840bd1a7e80a01%2Foutput.png?generation=1698216731946545&alt=media)",
    "2498353": "Thanks for sharing!",
    "2502660": "I've recently processed the ATAC-seq data and obtained results that were unexpected. Being relatively new to this, I'd appreciate insights from those with more experience. Notably, the FRIP score is quite low, and the TSS enrichment profile doesn't center as anticipated. Pls see the attachment.",
    "2524791": "Hello @jalilnourisa,\n\nI find this video is a nice introduction to ATAC-seq. \nhttps://www.youtube.com/watch?v=e2396GKFMRY\n\nI am sure there are nice tools in Python too, like this one:\nhttps://muon-tutorials.readthedocs.io/en/latest/single-cell-rna-atac/pbmc10k/2-Chromatin-Accessibility-Processing.html\n\nHere is the notebook explaining how to convert toanndata:\nhttps://www.kaggle.com/code/jeskowagner/converting-input-files-to-anndata/notebook\n\nFrom what I have been reading, a low FRIP score could indicate either technical issues with the library preparation or issues during the peak calling process.\n\nUseful links I found:\nhttps://www.biostars.org/p/9489225/\nhttps://www.encodeproject.org/atac-seq/\n\nSo, scores under 0.2 are concerning as they indicate either an issue with the bioinformatic pipeline or a technical issue.\n\nHere, the ATAC-seq peak counts were transformed with TF-IDF, which might explain why you get strange values. That would likely be an issue for the R and Python pipelines above.\nSee a good link about TF-IDF transformation of peaks data: https://divingintogeneticsandgenomics.com/post/clustering-scatacseq-data-the-tf-idf-way/.",
    "2528181": "Thank you for starting this discussion and the resources shared.  I would like to confirm that there is only baseline data, correct?  In other words, I have not seen ATAC-seq data for the cells after exposure to the small molecules.",
    "2528894": "Yes, you are correct. You only have the baseline. If you know which part of the genome is accessible (in a relaxed state chromatin region) for each cell type, you can potentially predict which genes are more likely to be expressed in a given cell type.",
    "2528927": "Ok great, thanks for confirming. \nI haven't figured out how to relate the Peaks data to a Cell type and Gene region. See file pasted here.  I'd appreciate input.",
    "2528939": "I am exploring several things. I have not yet found a straightforward method. The data here was pre-processed (peaks to TF-IDF). I am trying to see what methods will be compatible with that.",
    "2530223": "I am also trying to hook the multiome data to the training data. \nI figured how to associate the peaks to individual genes. \nI would like to associate the peaks to the cell types, but I don't know if that is possible. \nI assume that the peaks were measured for each cell_type?\nIf that is the case then how can you tell what cell type does a peak belongs to?",
    "2530681": "Thanks great idea",
    "2531025": "You can combine multiome_obs_meta with multiome_train using the obs_id as a key for an inner join. \nThis will link intervals with cell type and donors. \nHow did you associate the peaks to individual genes?"
  },
  "source": "meta"
}