{
  "id": 349766,
  "title": "Question about boundaries of DNA data",
  "url": "/competitions/open-problems-multimodal/discussion/349766",
  "author_name": "",
  "post_date": "2022-09-02T15:52:37.534960500Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>One of the datasets we have in this competition is the one about chromatin accessibility data. This dataset holds accessibility values for different regions of the DNA called peeks. Multiple peeks can constitute a single gene.</p>\n<p>To get a better understanding of the data, I did some analysis of the dataset. My attempt can be found here: <a href=\"https://www.kaggle.com/code/leohash/cell-feature-generation\" target=\"_blank\">https://www.kaggle.com/code/leohash/cell-feature-generation</a></p>\n<p>One thing that really surprised me is that most accessible chromatin regions don't have an accessible neighbor. Since multiple peeks can constitute a single gene and I assume access to DNA happens on a \"gene level\" I was expecting many neighboring accessible parts. Can someone explain to me what the biological reason behind that is?</p>\n<p>Kind regards!</p>",
  "messages": [
    {
      "id": "1923919",
      "postDate": "09/02/2022 15:52:37",
      "content": "<p>One of the datasets we have in this competition is the one about chromatin accessibility data. This dataset holds accessibility values for different regions of the DNA called peeks. Multiple peeks can constitute a single gene.</p>\n<p>To get a better understanding of the data, I did some analysis of the dataset. My attempt can be found here: <a href=\"https://www.kaggle.com/code/leohash/cell-feature-generation\" target=\"_blank\">https://www.kaggle.com/code/leohash/cell-feature-generation</a></p>\n<p>One thing that really surprised me is that most accessible chromatin regions don't have an accessible neighbor. Since multiple peeks can constitute a single gene and I assume access to DNA happens on a \"gene level\" I was expecting many neighboring accessible parts. Can someone explain to me what the biological reason behind that is?</p>\n<p>Kind regards!</p>",
      "rawMarkdown": "One of the datasets we have in this competition is the one about chromatin accessibility data. This dataset holds accessibility values for different regions of the DNA called peeks. Multiple peeks can constitute a single gene.\n\nTo get a better understanding of the data, I did some analysis of the dataset. My attempt can be found here: https://www.kaggle.com/code/leohash/cell-feature-generation\n\nOne thing that really surprised me is that most accessible chromatin regions don't have an accessible neighbor. Since multiple peeks can constitute a single gene and I assume access to DNA happens on a \"gene level\" I was expecting many neighboring accessible parts. Can someone explain to me what the biological reason behind that is?\n\nKind regards!",
      "votes": null
    },
    {
      "id": "1923971",
      "postDate": "09/02/2022 16:33:19",
      "content": "<p>Sounds very interesting ! Thanks for sharing !<br>\nBut is the link to the notebook correct ? It seems there is not output in it … </p>",
      "rawMarkdown": "Sounds very interesting ! Thanks for sharing !\nBut is the link to the notebook correct ? It seems there is not output in it ...",
      "votes": null
    },
    {
      "id": "1924506",
      "postDate": "09/03/2022 05:46:59",
      "content": "<p>Can you check the Data Tab of the notebook. The relevant information is written to a file metadata.csv. I choose this way to make it easy to use the data in other notebooks as input.</p>",
      "rawMarkdown": "Can you check the Data Tab of the notebook. The relevant information is written to a file metadata.csv. I choose this way to make it easy to use the data in other notebooks as input.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1923971,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/02/2022 16:33:19",
      "content": "<p>Sounds very interesting ! Thanks for sharing !<br>\nBut is the link to the notebook correct ? It seems there is not output in it … </p>",
      "votes": null,
      "replies": [
        {
          "id": 1924506,
          "author_name": "leohash",
          "author_url": "",
          "post_date": "09/03/2022 05:46:59",
          "content": "<p>Can you check the Data Tab of the notebook. The relevant information is written to a file metadata.csv. I choose this way to make it easy to use the data in other notebooks as input.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1923919": "One of the datasets we have in this competition is the one about chromatin accessibility data. This dataset holds accessibility values for different regions of the DNA called peeks. Multiple peeks can constitute a single gene.\n\nTo get a better understanding of the data, I did some analysis of the dataset. My attempt can be found here: https://www.kaggle.com/code/leohash/cell-feature-generation\n\nOne thing that really surprised me is that most accessible chromatin regions don't have an accessible neighbor. Since multiple peeks can constitute a single gene and I assume access to DNA happens on a \"gene level\" I was expecting many neighboring accessible parts. Can someone explain to me what the biological reason behind that is?\n\nKind regards!",
    "1923971": "Sounds very interesting ! Thanks for sharing !\nBut is the link to the notebook correct ? It seems there is not output in it ...",
    "1924506": "Can you check the Data Tab of the notebook. The relevant information is written to a file metadata.csv. I choose this way to make it easy to use the data in other notebooks as input."
  },
  "source": "meta"
}