{
  "id": 350863,
  "title": "BioQuestion 02: Feauture Importance for Multiome - how far from the target (gene) its important features are placed ?  DNA-distance",
  "url": "/competitions/open-problems-multimodal/discussion/350863",
  "author_name": "",
  "post_date": "2022-09-07T11:48:35.314118Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p><strong>Disclaimer.</strong> That might improve score, might not, but any outcome would be of interest for research community. <br>\nSo everyone is welcome to collaborate  - hopefully produce a paper - see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion1</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion2</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661\" target=\"_blank\">Discussion3\n</a></p>\n<p>For Multiome - each target (gene) has certain position on the DNA (two numbers - chromosome number, and position on the chromosome - start and end). These data are \"standard\" and can be found \"everywhere\" (okay, there are some details, but less clarify it later).   </p>\n<p>The same is for features of Multiome - each has chromosome number and start+end  (see names of the columns). </p>\n<p>The biological intuition says the concrete target (gene) can be affected ONLY BY NEARBY features. Here \"nearby\" - means in the sense of the position on DNA. Very rough estimates seems to be from 100 000 to 2 000 000. </p>\n<p><strong>Research question:</strong> Analyze <strong>how far</strong> important features are placed from each target (gene) . (\"Far\" means - distance along chromosome - how many nucleotides are in between). </p>\n<p>So for example,  expectation - features from say chromosome \"A\" - should not affect on genes (targets) on chromosome \"B\". Features from \"patches\" e.g. GL<strong>*</strong> should not affect at all (?) (may be??).</p>\n<p>The code from orgs related to the question: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559</a><br>\nThat notebook might be useful: <a href=\"https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity\" target=\"_blank\">https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity</a> </p>\n<p>There are some practical estimates from biology and some theoretical ideas like -  \"TAD\"s (<a href=\"https://en.wikipedia.org/wiki/Topologically_associating_domain\" target=\"_blank\">https://en.wikipedia.org/wiki/Topologically_associating_domain</a>)  are related to boundaries of \"influence\".<br>\nBut it would be interesting to understand and interpret what DS methods can given. <br>\nSome info on TADs: <a href=\"http://dna.cs.miami.edu/TADKB/\" target=\"_blank\">http://dna.cs.miami.edu/TADKB/</a></p>\n<p>PS <br>\nPrevious question: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>",
  "messages": [
    {
      "id": "1929832",
      "postDate": "09/07/2022 11:48:35",
      "content": "<p><strong>Disclaimer.</strong> That might improve score, might not, but any outcome would be of interest for research community. <br>\nSo everyone is welcome to collaborate  - hopefully produce a paper - see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion1</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293\" target=\"_blank\">Discussion2</a>, <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661\" target=\"_blank\">Discussion3\n</a></p>\n<p>For Multiome - each target (gene) has certain position on the DNA (two numbers - chromosome number, and position on the chromosome - start and end). These data are \"standard\" and can be found \"everywhere\" (okay, there are some details, but less clarify it later).   </p>\n<p>The same is for features of Multiome - each has chromosome number and start+end  (see names of the columns). </p>\n<p>The biological intuition says the concrete target (gene) can be affected ONLY BY NEARBY features. Here \"nearby\" - means in the sense of the position on DNA. Very rough estimates seems to be from 100 000 to 2 000 000. </p>\n<p><strong>Research question:</strong> Analyze <strong>how far</strong> important features are placed from each target (gene) . (\"Far\" means - distance along chromosome - how many nucleotides are in between). </p>\n<p>So for example,  expectation - features from say chromosome \"A\" - should not affect on genes (targets) on chromosome \"B\". Features from \"patches\" e.g. GL<strong>*</strong> should not affect at all (?) (may be??).</p>\n<p>The code from orgs related to the question: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559</a><br>\nThat notebook might be useful: <a href=\"https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity\" target=\"_blank\">https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity</a> </p>\n<p>There are some practical estimates from biology and some theoretical ideas like -  \"TAD\"s (<a href=\"https://en.wikipedia.org/wiki/Topologically_associating_domain\" target=\"_blank\">https://en.wikipedia.org/wiki/Topologically_associating_domain</a>)  are related to boundaries of \"influence\".<br>\nBut it would be interesting to understand and interpret what DS methods can given. <br>\nSome info on TADs: <a href=\"http://dna.cs.miami.edu/TADKB/\" target=\"_blank\">http://dna.cs.miami.edu/TADKB/</a></p>\n<p>PS <br>\nPrevious question: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>",
      "rawMarkdown": "**Disclaimer.** That might improve score, might not, but any outcome would be of interest for research community. \nSo everyone is welcome to collaborate  - hopefully produce a paper - see [Discussion1](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion2](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion3\n](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661)\n\nFor Multiome - each target (gene) has certain position on the DNA (two numbers - chromosome number, and position on the chromosome - start and end). These data are \"standard\" and can be found \"everywhere\" (okay, there are some details, but less clarify it later).   \n\nThe same is for features of Multiome - each has chromosome number and start+end  (see names of the columns). \n\nThe biological intuition says the concrete target (gene) can be affected ONLY BY NEARBY features. Here \"nearby\" - means in the sense of the position on DNA. Very rough estimates seems to be from 100 000 to 2 000 000. \n\n**Research question:** Analyze **how far** important features are placed from each target (gene) . (\"Far\" means - distance along chromosome - how many nucleotides are in between). \n\nSo for example,  expectation - features from say chromosome \"A\" - should not affect on genes (targets) on chromosome \"B\". Features from \"patches\" e.g. GL***** should not affect at all (?) (may be??).\n\nThe code from orgs related to the question: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\nThat notebook might be useful: https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity \n\nThere are some practical estimates from biology and some theoretical ideas like -  \"TAD\"s (https://en.wikipedia.org/wiki/Topologically_associating_domain)  are related to boundaries of \"influence\".\nBut it would be interesting to understand and interpret what DS methods can given. \nSome info on TADs: http://dna.cs.miami.edu/TADKB/\n\n\nPS \nPrevious question: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856",
      "votes": null
    },
    {
      "id": "1931385",
      "postDate": "09/08/2022 16:28:55",
      "content": "<p>That picture explains why somewhat far regions of DNA may influence each other - it is because they are close in 3d space. <br>\nMoreover there are quite extensive studies on 3D structure - and the protein CTCF plays the key role, a kind of binding to parts of the DNA together. The type of data on 3D structure is called Hi-C - <a href=\"https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique\" target=\"_blank\">https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique</a>)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F966ec603e7af3bcc93c1528c7df5cdc3%2Fphoto_2022-09-08_18-16-58.jpg?generation=1662653855604670&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "That picture explains why somewhat far regions of DNA may influence each other - it is because they are close in 3d space. \nMoreover there are quite extensive studies on 3D structure - and the protein CTCF plays the key role, a kind of binding to parts of the DNA together. The type of data on 3D structure is called Hi-C - https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F966ec603e7af3bcc93c1528c7df5cdc3%2Fphoto_2022-09-08_18-16-58.jpg?generation=1662653855604670&alt=media)",
      "votes": null
    },
    {
      "id": "1933699",
      "postDate": "09/10/2022 17:28:03",
      "content": "<p>For enhancers and promoters, there are proteins that bind to these regions and lead to increases or decreases in transcription.  In other words, increased or decreased levels of those proteins and their binding to enhancers or promoters will lead to increases or decreases in mRNA production.  Keep in mind that there are proteins that are repressors and their binding to certain DNA regions will lead to a decrease in mRNA for that particular gene.</p>\n<p>The correlation between mRNA concentration and protein concentration is not as strong as we might think, but certainly a lack of mRNA for a particular protein means none of that protein is produced.  More specifically, no new copies of that protein are produced, but existing copies are still present.  Just to complicate it a bit further, those existing copies may not be stable and may degrade over time.</p>\n<p>All of this can be described as cis-acting factors, meaning the proteins that bind to enhancers/promoters have an effect directly on transcription from that gene location.  There are also trans-acting factors which are proteins that bind somewhere far away from the gene of interest - not on its promoter or enhancer and maybe on another chromosome - but still have an impact.  Imagine a protein that increases the concentration of another protein which is a cis-acting factor for a given gene.  That original protein, a trans-acting factor in this example, has an impact on transcription from that gene, but indirectly though a cis-acting factor.</p>\n<p>The process from gene to mRNA to protein is incredibly complex and difficult to disentangle.  You are right, there are likely relationships that can be exploited based on the data we have, however, it certainly won't be obvious, in my opinion.</p>",
      "rawMarkdown": "For enhancers and promoters, there are proteins that bind to these regions and lead to increases or decreases in transcription.  In other words, increased or decreased levels of those proteins and their binding to enhancers or promoters will lead to increases or decreases in mRNA production.  Keep in mind that there are proteins that are repressors and their binding to certain DNA regions will lead to a decrease in mRNA for that particular gene.\n\nThe correlation between mRNA concentration and protein concentration is not as strong as we might think, but certainly a lack of mRNA for a particular protein means none of that protein is produced.  More specifically, no new copies of that protein are produced, but existing copies are still present.  Just to complicate it a bit further, those existing copies may not be stable and may degrade over time.\n\nAll of this can be described as cis-acting factors, meaning the proteins that bind to enhancers/promoters have an effect directly on transcription from that gene location.  There are also trans-acting factors which are proteins that bind somewhere far away from the gene of interest - not on its promoter or enhancer and maybe on another chromosome - but still have an impact.  Imagine a protein that increases the concentration of another protein which is a cis-acting factor for a given gene.  That original protein, a trans-acting factor in this example, has an impact on transcription from that gene, but indirectly though a cis-acting factor.\n\nThe process from gene to mRNA to protein is incredibly complex and difficult to disentangle.  You are right, there are likely relationships that can be exploited based on the data we have, however, it certainly won't be obvious, in my opinion.",
      "votes": null
    },
    {
      "id": "1934896",
      "postDate": "09/11/2022 16:44:01",
      "content": "<p>Interesting. I just published some work I had been doing this week on inputs/targets correlations (see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725)\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725)</a>. </p>\n<p>I had noticed some intriguing patterns of targets being very similarly affected by several inputs in Multiome. Maybe this locality aspect is the biological explanation. I'll have to check that. </p>\n<p>And btw, since you seem to have some in-domain knowledge, I'd be very interested in some biology feedback on my purely data-driven observations.</p>",
      "rawMarkdown": "Interesting. I just published some work I had been doing this week on inputs/targets correlations (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725). \n\nI had noticed some intriguing patterns of targets being very similarly affected by several inputs in Multiome. Maybe this locality aspect is the biological explanation. I'll have to check that. \n\nAnd btw, since you seem to have some in-domain knowledge, I'd be very interested in some biology feedback on my purely data-driven observations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1931385,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/08/2022 16:28:55",
      "content": "<p>That picture explains why somewhat far regions of DNA may influence each other - it is because they are close in 3d space. <br>\nMoreover there are quite extensive studies on 3D structure - and the protein CTCF plays the key role, a kind of binding to parts of the DNA together. The type of data on 3D structure is called Hi-C - <a href=\"https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique\" target=\"_blank\">https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique</a>)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F966ec603e7af3bcc93c1528c7df5cdc3%2Fphoto_2022-09-08_18-16-58.jpg?generation=1662653855604670&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1933699,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "09/10/2022 17:28:03",
          "content": "<p>For enhancers and promoters, there are proteins that bind to these regions and lead to increases or decreases in transcription.  In other words, increased or decreased levels of those proteins and their binding to enhancers or promoters will lead to increases or decreases in mRNA production.  Keep in mind that there are proteins that are repressors and their binding to certain DNA regions will lead to a decrease in mRNA for that particular gene.</p>\n<p>The correlation between mRNA concentration and protein concentration is not as strong as we might think, but certainly a lack of mRNA for a particular protein means none of that protein is produced.  More specifically, no new copies of that protein are produced, but existing copies are still present.  Just to complicate it a bit further, those existing copies may not be stable and may degrade over time.</p>\n<p>All of this can be described as cis-acting factors, meaning the proteins that bind to enhancers/promoters have an effect directly on transcription from that gene location.  There are also trans-acting factors which are proteins that bind somewhere far away from the gene of interest - not on its promoter or enhancer and maybe on another chromosome - but still have an impact.  Imagine a protein that increases the concentration of another protein which is a cis-acting factor for a given gene.  That original protein, a trans-acting factor in this example, has an impact on transcription from that gene, but indirectly though a cis-acting factor.</p>\n<p>The process from gene to mRNA to protein is incredibly complex and difficult to disentangle.  You are right, there are likely relationships that can be exploited based on the data we have, however, it certainly won't be obvious, in my opinion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1934896,
      "author_name": "fabiencrom",
      "author_url": "",
      "post_date": "09/11/2022 16:44:01",
      "content": "<p>Interesting. I just published some work I had been doing this week on inputs/targets correlations (see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725)\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725)</a>. </p>\n<p>I had noticed some intriguing patterns of targets being very similarly affected by several inputs in Multiome. Maybe this locality aspect is the biological explanation. I'll have to check that. </p>\n<p>And btw, since you seem to have some in-domain knowledge, I'd be very interested in some biology feedback on my purely data-driven observations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1929832": "**Disclaimer.** That might improve score, might not, but any outcome would be of interest for research community. \nSo everyone is welcome to collaborate  - hopefully produce a paper - see [Discussion1](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion2](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348293), [Discussion3\n](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348661)\n\nFor Multiome - each target (gene) has certain position on the DNA (two numbers - chromosome number, and position on the chromosome - start and end). These data are \"standard\" and can be found \"everywhere\" (okay, there are some details, but less clarify it later).   \n\nThe same is for features of Multiome - each has chromosome number and start+end  (see names of the columns). \n\nThe biological intuition says the concrete target (gene) can be affected ONLY BY NEARBY features. Here \"nearby\" - means in the sense of the position on DNA. Very rough estimates seems to be from 100 000 to 2 000 000. \n\n**Research question:** Analyze **how far** important features are placed from each target (gene) . (\"Far\" means - distance along chromosome - how many nucleotides are in between). \n\nSo for example,  expectation - features from say chromosome \"A\" - should not affect on genes (targets) on chromosome \"B\". Features from \"patches\" e.g. GL***** should not affect at all (?) (may be??).\n\nThe code from orgs related to the question: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349559\nThat notebook might be useful: https://www.kaggle.com/code/masato114/msci-multiome-using-geneactivity \n\nThere are some practical estimates from biology and some theoretical ideas like -  \"TAD\"s (https://en.wikipedia.org/wiki/Topologically_associating_domain)  are related to boundaries of \"influence\".\nBut it would be interesting to understand and interpret what DS methods can given. \nSome info on TADs: http://dna.cs.miami.edu/TADKB/\n\n\nPS \nPrevious question: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856",
    "1931385": "That picture explains why somewhat far regions of DNA may influence each other - it is because they are close in 3d space. \nMoreover there are quite extensive studies on 3D structure - and the protein CTCF plays the key role, a kind of binding to parts of the DNA together. The type of data on 3D structure is called Hi-C - https://en.wikipedia.org/wiki/Hi-C_(genomic_analysis_technique)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F966ec603e7af3bcc93c1528c7df5cdc3%2Fphoto_2022-09-08_18-16-58.jpg?generation=1662653855604670&alt=media)",
    "1933699": "For enhancers and promoters, there are proteins that bind to these regions and lead to increases or decreases in transcription.  In other words, increased or decreased levels of those proteins and their binding to enhancers or promoters will lead to increases or decreases in mRNA production.  Keep in mind that there are proteins that are repressors and their binding to certain DNA regions will lead to a decrease in mRNA for that particular gene.\n\nThe correlation between mRNA concentration and protein concentration is not as strong as we might think, but certainly a lack of mRNA for a particular protein means none of that protein is produced.  More specifically, no new copies of that protein are produced, but existing copies are still present.  Just to complicate it a bit further, those existing copies may not be stable and may degrade over time.\n\nAll of this can be described as cis-acting factors, meaning the proteins that bind to enhancers/promoters have an effect directly on transcription from that gene location.  There are also trans-acting factors which are proteins that bind somewhere far away from the gene of interest - not on its promoter or enhancer and maybe on another chromosome - but still have an impact.  Imagine a protein that increases the concentration of another protein which is a cis-acting factor for a given gene.  That original protein, a trans-acting factor in this example, has an impact on transcription from that gene, but indirectly though a cis-acting factor.\n\nThe process from gene to mRNA to protein is incredibly complex and difficult to disentangle.  You are right, there are likely relationships that can be exploited based on the data we have, however, it certainly won't be obvious, in my opinion.",
    "1934896": "Interesting. I just published some work I had been doing this week on inputs/targets correlations (see https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725). \n\nI had noticed some intriguing patterns of targets being very similarly affected by several inputs in Multiome. Maybe this locality aspect is the biological explanation. I'll have to check that. \n\nAnd btw, since you seem to have some in-domain knowledge, I'd be very interested in some biology feedback on my purely data-driven observations."
  },
  "source": "meta"
}