{
  "id": 346798,
  "title": "How to overcome lack of knowledge?",
  "url": "/competitions/open-problems-multimodal/discussion/346798",
  "author_name": "",
  "post_date": "2022-08-21T11:41:13.311721600Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This competition is very special and needs extremly specialized knowledge.</p>\n<p>I find it hard reading about this topic (even the data tab) because I would have to start on a very basic level to understand what we are supposed to predict here and how the data helps us to do so.</p>\n<p>Does anyone face the same issue and if yes, what are you doing about it?</p>",
  "messages": [
    {
      "id": "1908126",
      "postDate": "08/21/2022 11:41:13",
      "content": "<p>This competition is very special and needs extremly specialized knowledge.</p>\n<p>I find it hard reading about this topic (even the data tab) because I would have to start on a very basic level to understand what we are supposed to predict here and how the data helps us to do so.</p>\n<p>Does anyone face the same issue and if yes, what are you doing about it?</p>",
      "rawMarkdown": "This competition is very special and needs extremly specialized knowledge.\n\nI find it hard reading about this topic (even the data tab) because I would have to start on a very basic level to understand what we are supposed to predict here and how the data helps us to do so.\n\nDoes anyone face the same issue and if yes, what are you doing about it?",
      "votes": null
    },
    {
      "id": "1908499",
      "postDate": "08/21/2022 18:15:04",
      "content": "<p>Honestly, I find that I tend to have a poor understanding of a dataset just by reading about it - particularly for a competition like this where the test set is split in an unusual fashion and is outside my field anyway. I think I read the Data Description about ten times and still didn't really have a feel for it. In order to get a better understanding of what the challenge is and how the data fits together, there is no substitute for just diving in and exploring the data itself manually. Remember that an in-depth understanding of biochemistry is in no way essential for this competition (although a bit of domain knowledge never hurts) and ultimately, this is an ML problem. </p>",
      "rawMarkdown": "Honestly, I find that I tend to have a poor understanding of a dataset just by reading about it - particularly for a competition like this where the test set is split in an unusual fashion and is outside my field anyway. I think I read the Data Description about ten times and still didn't really have a feel for it. In order to get a better understanding of what the challenge is and how the data fits together, there is no substitute for just diving in and exploring the data itself manually. Remember that an in-depth understanding of biochemistry is in no way essential for this competition (although a bit of domain knowledge never hurts) and ultimately, this is an ML problem.",
      "votes": null
    },
    {
      "id": "1908527",
      "postDate": "08/21/2022 18:38:14",
      "content": "<p>We gave a talk about last year’s competition that may be helpful context:<br>\n<a href=\"https://www.youtube.com/watch?v=ZXDILOyiy7A\" target=\"_blank\">https://www.youtube.com/watch?v=ZXDILOyiy7A</a></p>\n<p>Abstractly, there are three layers to the central dogma of biology: DNA, RNA, protein. Here the challenge is to predict features of one layer based on features of another layer, where benchmarking is possible because technologies now exist to measure both layers in the same cell for hundreds of thousands of cells.</p>\n<p>So one can start with any model that takes a vector and predicts a vector. We’re of course also interested to see whether biological knowledge will inspire higher-performing and/or more interpretable model architectures.</p>\n<p>For example, the knowledge that the DNA features are ordered by location along the genome (indicating that location’s physical accessibility), each RNA feature corresponds to a DNA location, and each protein feature corresponds to an RNA feature.</p>",
      "rawMarkdown": "We gave a talk about last year’s competition that may be helpful context:\nhttps://www.youtube.com/watch?v=ZXDILOyiy7A\n\nAbstractly, there are three layers to the central dogma of biology: DNA, RNA, protein. Here the challenge is to predict features of one layer based on features of another layer, where benchmarking is possible because technologies now exist to measure both layers in the same cell for hundreds of thousands of cells.\n\nSo one can start with any model that takes a vector and predicts a vector. We’re of course also interested to see whether biological knowledge will inspire higher-performing and/or more interpretable model architectures.\n\nFor example, the knowledge that the DNA features are ordered by location along the genome (indicating that location’s physical accessibility), each RNA feature corresponds to a DNA location, and each protein feature corresponds to an RNA feature.",
      "votes": null
    },
    {
      "id": "1908802",
      "postDate": "08/22/2022 03:38:57",
      "content": "<p>Maybe you can learn by reading the code posted by others</p>",
      "rawMarkdown": "Maybe you can learn by reading the code posted by others",
      "votes": null
    },
    {
      "id": "1909298",
      "postDate": "08/22/2022 13:50:43",
      "content": "<p>In addition, we created a short info page for a similar competition we hosted last year. It contains some descriptions about the underlying data type and plenty of links to learning resources throughout.</p>\n<p>👉 <a href=\"https://openproblems.bio/neurips_docs/data/about_multimodal/\" target=\"_blank\">Open Problems - About multimodal single-cell data</a></p>",
      "rawMarkdown": "In addition, we created a short info page for a similar competition we hosted last year. It contains some descriptions about the underlying data type and plenty of links to learning resources throughout.\n\n👉 [Open Problems - About multimodal single-cell data](https://openproblems.bio/neurips_docs/data/about_multimodal/)",
      "votes": null
    },
    {
      "id": "1914371",
      "postDate": "08/26/2022 03:40:46",
      "content": "<p>same here… had a very difficult time reading the data description 😂</p>",
      "rawMarkdown": "same here... had a very difficult time reading the data description 😂",
      "votes": null
    },
    {
      "id": "1914995",
      "postDate": "08/26/2022 15:42:22",
      "content": "<p>I am considering jumping into this competition.  My approach will be simple, read all the informative and popular topics to understand (a little) and spark ideas. Then start ONLY from cloning top public notebook(s) with high score + good explanations both. </p>\n<p>My last competition used \"anonymized\" data. I will treat this data similarly, since I have no domain knowledge. Others will post good ideas based on domain knowledge and I will use my general problem solving skills to estimate which ideas are the most likely to help and pursue those ideas first. </p>",
      "rawMarkdown": "I am considering jumping into this competition.  My approach will be simple, read all the informative and popular topics to understand (a little) and spark ideas. Then start ONLY from cloning top public notebook(s) with high score + good explanations both. \n\nMy last competition used \"anonymized\" data. I will treat this data similarly, since I have no domain knowledge. Others will post good ideas based on domain knowledge and I will use my general problem solving skills to estimate which ideas are the most likely to help and pursue those ideas first.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1908499,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "08/21/2022 18:15:04",
      "content": "<p>Honestly, I find that I tend to have a poor understanding of a dataset just by reading about it - particularly for a competition like this where the test set is split in an unusual fashion and is outside my field anyway. I think I read the Data Description about ten times and still didn't really have a feel for it. In order to get a better understanding of what the challenge is and how the data fits together, there is no substitute for just diving in and exploring the data itself manually. Remember that an in-depth understanding of biochemistry is in no way essential for this competition (although a bit of domain knowledge never hurts) and ultimately, this is an ML problem. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1908527,
      "author_name": "jonathanbloom",
      "author_url": "",
      "post_date": "08/21/2022 18:38:14",
      "content": "<p>We gave a talk about last year’s competition that may be helpful context:<br>\n<a href=\"https://www.youtube.com/watch?v=ZXDILOyiy7A\" target=\"_blank\">https://www.youtube.com/watch?v=ZXDILOyiy7A</a></p>\n<p>Abstractly, there are three layers to the central dogma of biology: DNA, RNA, protein. Here the challenge is to predict features of one layer based on features of another layer, where benchmarking is possible because technologies now exist to measure both layers in the same cell for hundreds of thousands of cells.</p>\n<p>So one can start with any model that takes a vector and predicts a vector. We’re of course also interested to see whether biological knowledge will inspire higher-performing and/or more interpretable model architectures.</p>\n<p>For example, the knowledge that the DNA features are ordered by location along the genome (indicating that location’s physical accessibility), each RNA feature corresponds to a DNA location, and each protein feature corresponds to an RNA feature.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1909298,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "08/22/2022 13:50:43",
          "content": "<p>In addition, we created a short info page for a similar competition we hosted last year. It contains some descriptions about the underlying data type and plenty of links to learning resources throughout.</p>\n<p>👉 <a href=\"https://openproblems.bio/neurips_docs/data/about_multimodal/\" target=\"_blank\">Open Problems - About multimodal single-cell data</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1908802,
      "author_name": "shawnasan",
      "author_url": "",
      "post_date": "08/22/2022 03:38:57",
      "content": "<p>Maybe you can learn by reading the code posted by others</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914371,
      "author_name": "snaker",
      "author_url": "",
      "post_date": "08/26/2022 03:40:46",
      "content": "<p>same here… had a very difficult time reading the data description 😂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914995,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/26/2022 15:42:22",
      "content": "<p>I am considering jumping into this competition.  My approach will be simple, read all the informative and popular topics to understand (a little) and spark ideas. Then start ONLY from cloning top public notebook(s) with high score + good explanations both. </p>\n<p>My last competition used \"anonymized\" data. I will treat this data similarly, since I have no domain knowledge. Others will post good ideas based on domain knowledge and I will use my general problem solving skills to estimate which ideas are the most likely to help and pursue those ideas first. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1908126": "This competition is very special and needs extremly specialized knowledge.\n\nI find it hard reading about this topic (even the data tab) because I would have to start on a very basic level to understand what we are supposed to predict here and how the data helps us to do so.\n\nDoes anyone face the same issue and if yes, what are you doing about it?",
    "1908499": "Honestly, I find that I tend to have a poor understanding of a dataset just by reading about it - particularly for a competition like this where the test set is split in an unusual fashion and is outside my field anyway. I think I read the Data Description about ten times and still didn't really have a feel for it. In order to get a better understanding of what the challenge is and how the data fits together, there is no substitute for just diving in and exploring the data itself manually. Remember that an in-depth understanding of biochemistry is in no way essential for this competition (although a bit of domain knowledge never hurts) and ultimately, this is an ML problem.",
    "1908527": "We gave a talk about last year’s competition that may be helpful context:\nhttps://www.youtube.com/watch?v=ZXDILOyiy7A\n\nAbstractly, there are three layers to the central dogma of biology: DNA, RNA, protein. Here the challenge is to predict features of one layer based on features of another layer, where benchmarking is possible because technologies now exist to measure both layers in the same cell for hundreds of thousands of cells.\n\nSo one can start with any model that takes a vector and predicts a vector. We’re of course also interested to see whether biological knowledge will inspire higher-performing and/or more interpretable model architectures.\n\nFor example, the knowledge that the DNA features are ordered by location along the genome (indicating that location’s physical accessibility), each RNA feature corresponds to a DNA location, and each protein feature corresponds to an RNA feature.",
    "1908802": "Maybe you can learn by reading the code posted by others",
    "1909298": "In addition, we created a short info page for a similar competition we hosted last year. It contains some descriptions about the underlying data type and plenty of links to learning resources throughout.\n\n👉 [Open Problems - About multimodal single-cell data](https://openproblems.bio/neurips_docs/data/about_multimodal/)",
    "1914371": "same here... had a very difficult time reading the data description 😂",
    "1914995": "I am considering jumping into this competition.  My approach will be simple, read all the informative and popular topics to understand (a little) and spark ideas. Then start ONLY from cloning top public notebook(s) with high score + good explanations both. \n\nMy last competition used \"anonymized\" data. I will treat this data similarly, since I have no domain knowledge. Others will post good ideas based on domain knowledge and I will use my general problem solving skills to estimate which ideas are the most likely to help and pursue those ideas first."
  },
  "source": "meta"
}