{
  "id": 494685,
  "title": "Does strong background in Medicinal Chemistry is required?",
  "url": "/competitions/leash-BELKA/discussion/494685",
  "author_name": "",
  "post_date": "2024-04-18T04:17:33.782788700Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I think yes, In order to predict the best fit of a ligand and you have to have strong knowledge of medicinal chemistry because at the end of the day you are predicting the drug if you don't know about the chemistry of drug how will you develop the algorithm.</p>",
  "messages": [
    {
      "id": "2758266",
      "postDate": "04/18/2024 04:17:33",
      "content": "<p>I think yes, In order to predict the best fit of a ligand and you have to have strong knowledge of medicinal chemistry because at the end of the day you are predicting the drug if you don't know about the chemistry of drug how will you develop the algorithm.</p>",
      "rawMarkdown": "I think yes, In order to predict the best fit of a ligand and you have to have strong knowledge of medicinal chemistry because at the end of the day you are predicting the drug if you don't know about the chemistry of drug how will you develop the algorithm.",
      "votes": null
    },
    {
      "id": "2758372",
      "postDate": "04/18/2024 06:20:22",
      "content": "<p>the answer is both yes and no, at current state of technology.</p>\n<p>if the answer yes ( background knowledge is required), it means that a datascience degree ALONE is pretty useless. you need additional degree (e.g. business, bio informatics) on your application for domain knowledge. </p>\n<p>if the answer is no ( background knowledge is NOT required), then automated AI system (like drug design, essay scoring system) are working like human. Then many jobs will be completely repalced (i.e. with human intervention) by AI system.</p>\n<hr>\n<p>The correct answer is really depends on how how data you have. if you have ALL (or sufficiently large) data, then you can solve the problem from pure data point of view (i.e. just treat data such vectors, \"pure numbers without the meaning behind it\") </p>\n<hr>\n<p>an analogy is openAI sora. i can generate video without knowing anything about the world (physics, etc …) It \"seems correct\", but not really \"correct\".</p>",
      "rawMarkdown": "the answer is both yes and no, at current state of technology.\n\nif the answer yes ( background knowledge is required), it means that a datascience degree ALONE is pretty useless. you need additional degree (e.g. business, bio informatics) on your application for domain knowledge. \n\nif the answer is no ( background knowledge is NOT required), then automated AI system (like drug design, essay scoring system) are working like human. Then many jobs will be completely repalced (i.e. with human intervention) by AI system.\n\n---\n\nThe correct answer is really depends on how how data you have. if you have ALL (or sufficiently large) data, then you can solve the problem from pure data point of view (i.e. just treat data such vectors, \"pure numbers without the meaning behind it\") \n\n---\n\nan analogy is openAI sora. i can generate video without knowing anything about the world (physics, etc ...) It \"seems correct\", but not really \"correct\".",
      "votes": null
    },
    {
      "id": "2758714",
      "postDate": "04/18/2024 10:02:34",
      "content": "<p>Was sora created just by feeding it tons of data? I think they probably took a more thoughtful approach, since both data and training capacity is limited</p>",
      "rawMarkdown": "Was sora created just by feeding it tons of data? I think they probably took a more thoughtful approach, since both data and training capacity is limited",
      "votes": null
    },
    {
      "id": "2759213",
      "postDate": "04/18/2024 14:42:55",
      "content": "<p>This is a pretty complicated task - lots of smart domain understanding people in the world working to understand it.  If you have a strong background in the domain I have found that it helps remove some of the wrong paths you might take in creating features that work, or eliminating duplication of features, etc.</p>\n<p>But you still have to go down a lot of paths - one of my best competition results was with a team that had 3 graduate chemists and two really good Python coders on a similar task.  Another good result was a team that had only no graduates in the domain but 4 really good coders.</p>\n<p>IMO this competition needs some heavy duty compute power, lots of data science and a touch of biochemistry.</p>",
      "rawMarkdown": "This is a pretty complicated task - lots of smart domain understanding people in the world working to understand it.  If you have a strong background in the domain I have found that it helps remove some of the wrong paths you might take in creating features that work, or eliminating duplication of features, etc.\n\nBut you still have to go down a lot of paths - one of my best competition results was with a team that had 3 graduate chemists and two really good Python coders on a similar task.  Another good result was a team that had only no graduates in the domain but 4 really good coders.\n\nIMO this competition needs some heavy duty compute power, lots of data science and a touch of biochemistry.",
      "votes": null
    },
    {
      "id": "2759327",
      "postDate": "04/18/2024 15:55:57",
      "content": "<p>While we think there is a lot of domain knowledge that is important in this problem, we set off to build these enormous datasets exactly because we think this problem can be solved with enough data.</p>\n<p>I highly recommend everyone read a short essay by the great Richard Sutton laying out this argument, known as The Bitter Lesson <a href=\"http://www.incompleteideas.net/IncIdeas/BitterLesson.html\" target=\"_blank\">http://www.incompleteideas.net/IncIdeas/BitterLesson.html</a></p>",
      "rawMarkdown": "While we think there is a lot of domain knowledge that is important in this problem, we set off to build these enormous datasets exactly because we think this problem can be solved with enough data.\n\nI highly recommend everyone read a short essay by the great Richard Sutton laying out this argument, known as The Bitter Lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html",
      "votes": null
    },
    {
      "id": "2759641",
      "postDate": "04/18/2024 19:52:40",
      "content": "<p>Kaggle seems like the perfect place to answer that question! But you'll have to wait until the end of the competition to get the answer…</p>",
      "rawMarkdown": "Kaggle seems like the perfect place to answer that question! But you'll have to wait until the end of the competition to get the answer...",
      "votes": null
    },
    {
      "id": "2773920",
      "postDate": "04/25/2024 01:02:00",
      "content": "<p>how about each team put the number of chemist in their team name? Then we can tell from the leader board now.<br>\ne.g. xxx-name[1]</p>",
      "rawMarkdown": "how about each team put the number of chemist in their team name? Then we can tell from the leader board now.\ne.g. xxx-name[1]",
      "votes": null
    },
    {
      "id": "2775342",
      "postDate": "04/25/2024 16:19:38",
      "content": "<p>A lot of data science is about removing human intervention required. </p>\n<p>We know that bio chem expertise is helpful. We also know that a good data scientist can use ML to effectively analyze 300 million training samples and give helpful predictions, with no domain expertise required. </p>\n<p>But winning a competition is really it's own little world. Any one trick or a hundred clever ideas or nothing clever just good comprehensive data science, any or all or none of the above could happen to do better than everyone else. </p>\n<p>At the core, though, I guess I'll take a stand and say of course domain expertise isn't required. Why?</p>\n<p>Libraries like rdkit and tools like diffdock. Written by domain experts, usable by anyone, or at least by any data scientist! And helpful kaggle experts sharing publicly and answering questions like <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> </p>",
      "rawMarkdown": "A lot of data science is about removing human intervention required. \n\nWe know that bio chem expertise is helpful. We also know that a good data scientist can use ML to effectively analyze 300 million training samples and give helpful predictions, with no domain expertise required. \n\nBut winning a competition is really it's own little world. Any one trick or a hundred clever ideas or nothing clever just good comprehensive data science, any or all or none of the above could happen to do better than everyone else. \n\nAt the core, though, I guess I'll take a stand and say of course domain expertise isn't required. Why?\n\nLibraries like rdkit and tools like diffdock. Written by domain experts, usable by anyone, or at least by any data scientist! And helpful kaggle experts sharing publicly and answering questions like @chemdatafarmer",
      "votes": null
    },
    {
      "id": "2775870",
      "postDate": "04/25/2024 21:10:30",
      "content": "<p>Thanks for the shout-out <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> :) happy to help.</p>",
      "rawMarkdown": "Thanks for the shout-out @roberthatch :) happy to help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2758372,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/18/2024 06:20:22",
      "content": "<p>the answer is both yes and no, at current state of technology.</p>\n<p>if the answer yes ( background knowledge is required), it means that a datascience degree ALONE is pretty useless. you need additional degree (e.g. business, bio informatics) on your application for domain knowledge. </p>\n<p>if the answer is no ( background knowledge is NOT required), then automated AI system (like drug design, essay scoring system) are working like human. Then many jobs will be completely repalced (i.e. with human intervention) by AI system.</p>\n<hr>\n<p>The correct answer is really depends on how how data you have. if you have ALL (or sufficiently large) data, then you can solve the problem from pure data point of view (i.e. just treat data such vectors, \"pure numbers without the meaning behind it\") </p>\n<hr>\n<p>an analogy is openAI sora. i can generate video without knowing anything about the world (physics, etc …) It \"seems correct\", but not really \"correct\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 2758714,
          "author_name": "gyulamaloveczky4",
          "author_url": "",
          "post_date": "04/18/2024 10:02:34",
          "content": "<p>Was sora created just by feeding it tons of data? I think they probably took a more thoughtful approach, since both data and training capacity is limited</p>",
          "votes": null,
          "replies": [
            {
              "id": 2759327,
              "author_name": "andrewdblevins",
              "author_url": "",
              "post_date": "04/18/2024 15:55:57",
              "content": "<p>While we think there is a lot of domain knowledge that is important in this problem, we set off to build these enormous datasets exactly because we think this problem can be solved with enough data.</p>\n<p>I highly recommend everyone read a short essay by the great Richard Sutton laying out this argument, known as The Bitter Lesson <a href=\"http://www.incompleteideas.net/IncIdeas/BitterLesson.html\" target=\"_blank\">http://www.incompleteideas.net/IncIdeas/BitterLesson.html</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2759213,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "04/18/2024 14:42:55",
      "content": "<p>This is a pretty complicated task - lots of smart domain understanding people in the world working to understand it.  If you have a strong background in the domain I have found that it helps remove some of the wrong paths you might take in creating features that work, or eliminating duplication of features, etc.</p>\n<p>But you still have to go down a lot of paths - one of my best competition results was with a team that had 3 graduate chemists and two really good Python coders on a similar task.  Another good result was a team that had only no graduates in the domain but 4 really good coders.</p>\n<p>IMO this competition needs some heavy duty compute power, lots of data science and a touch of biochemistry.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2759641,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "04/18/2024 19:52:40",
      "content": "<p>Kaggle seems like the perfect place to answer that question! But you'll have to wait until the end of the competition to get the answer…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2773920,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/25/2024 01:02:00",
          "content": "<p>how about each team put the number of chemist in their team name? Then we can tell from the leader board now.<br>\ne.g. xxx-name[1]</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2775342,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/25/2024 16:19:38",
      "content": "<p>A lot of data science is about removing human intervention required. </p>\n<p>We know that bio chem expertise is helpful. We also know that a good data scientist can use ML to effectively analyze 300 million training samples and give helpful predictions, with no domain expertise required. </p>\n<p>But winning a competition is really it's own little world. Any one trick or a hundred clever ideas or nothing clever just good comprehensive data science, any or all or none of the above could happen to do better than everyone else. </p>\n<p>At the core, though, I guess I'll take a stand and say of course domain expertise isn't required. Why?</p>\n<p>Libraries like rdkit and tools like diffdock. Written by domain experts, usable by anyone, or at least by any data scientist! And helpful kaggle experts sharing publicly and answering questions like <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2775870,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/25/2024 21:10:30",
          "content": "<p>Thanks for the shout-out <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> :) happy to help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2758266": "I think yes, In order to predict the best fit of a ligand and you have to have strong knowledge of medicinal chemistry because at the end of the day you are predicting the drug if you don't know about the chemistry of drug how will you develop the algorithm.",
    "2758372": "the answer is both yes and no, at current state of technology.\n\nif the answer yes ( background knowledge is required), it means that a datascience degree ALONE is pretty useless. you need additional degree (e.g. business, bio informatics) on your application for domain knowledge. \n\nif the answer is no ( background knowledge is NOT required), then automated AI system (like drug design, essay scoring system) are working like human. Then many jobs will be completely repalced (i.e. with human intervention) by AI system.\n\n---\n\nThe correct answer is really depends on how how data you have. if you have ALL (or sufficiently large) data, then you can solve the problem from pure data point of view (i.e. just treat data such vectors, \"pure numbers without the meaning behind it\") \n\n---\n\nan analogy is openAI sora. i can generate video without knowing anything about the world (physics, etc ...) It \"seems correct\", but not really \"correct\".",
    "2758714": "Was sora created just by feeding it tons of data? I think they probably took a more thoughtful approach, since both data and training capacity is limited",
    "2759213": "This is a pretty complicated task - lots of smart domain understanding people in the world working to understand it.  If you have a strong background in the domain I have found that it helps remove some of the wrong paths you might take in creating features that work, or eliminating duplication of features, etc.\n\nBut you still have to go down a lot of paths - one of my best competition results was with a team that had 3 graduate chemists and two really good Python coders on a similar task.  Another good result was a team that had only no graduates in the domain but 4 really good coders.\n\nIMO this competition needs some heavy duty compute power, lots of data science and a touch of biochemistry.",
    "2759327": "While we think there is a lot of domain knowledge that is important in this problem, we set off to build these enormous datasets exactly because we think this problem can be solved with enough data.\n\nI highly recommend everyone read a short essay by the great Richard Sutton laying out this argument, known as The Bitter Lesson http://www.incompleteideas.net/IncIdeas/BitterLesson.html",
    "2759641": "Kaggle seems like the perfect place to answer that question! But you'll have to wait until the end of the competition to get the answer...",
    "2773920": "how about each team put the number of chemist in their team name? Then we can tell from the leader board now.\ne.g. xxx-name[1]",
    "2775342": "A lot of data science is about removing human intervention required. \n\nWe know that bio chem expertise is helpful. We also know that a good data scientist can use ML to effectively analyze 300 million training samples and give helpful predictions, with no domain expertise required. \n\nBut winning a competition is really it's own little world. Any one trick or a hundred clever ideas or nothing clever just good comprehensive data science, any or all or none of the above could happen to do better than everyone else. \n\nAt the core, though, I guess I'll take a stand and say of course domain expertise isn't required. Why?\n\nLibraries like rdkit and tools like diffdock. Written by domain experts, usable by anyone, or at least by any data scientist! And helpful kaggle experts sharing publicly and answering questions like @chemdatafarmer",
    "2775870": "Thanks for the shout-out @roberthatch :) happy to help."
  },
  "source": "meta"
}