{
  "id": 502970,
  "title": "What Might the Winning Score be?",
  "url": "/competitions/leash-BELKA/discussion/502970",
  "author_name": "John Mitchell",
  "post_date": "2024-05-15T15:01:06.199000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>[Written in terms of \"old\" score - hope to update soon]</p>\n<p>\"Just for fun\", I'm speculating about what the ultimate winning LB score (both private and corresponding public are interesting) might be at the end of this competition.</p>\n<p>The technology behind BELKA is really clever and effective, nonetheless it is a high-throughput combinatorial chemistry kind of method and the DNA tags will affect binding to some extent. If I were to imagine a best possible measured set of  gold standard binding affinities from the best conceivable experiments, I'd expect them to differ significantly from the BELKA ones.</p>\n<p>Of the methods tried in this competition, many of the codes have concentrated on a 2D-descriptor-based cheminformatics approach. As of now, this is maxing out in my hands just below 0.600 LB, though some have reported slightly higher scores.</p>\n<p>We may also see more innovative AI approaches based on tokenizing SMILES etc., the limited available information suggests to me that these may also end up scoring around 0.600.</p>\n<p>There's clearly scope for some involvement of docking in this competition, though it's perhaps likely to involve also a degree of AI-based selection of key compounds to dock and prediction of other molecules' scores rather than the expense of docking each and every compound.</p>\n<p>We may also see some MD or similar simulation approaches, but again I'd anticipate this being combined with clever models to select the key compounds to simulate and to predict affinities of others, due to the expense of using an approach like that on a dataset of this size. There may also be other kinds of model that I haven't considered.</p>\n<p>I speculate that the best solutions will be ensembles incorporating \"cheminformatics truth\", \"docking truth\", \"MD truth\", \"generative AI truth\" etc. The LB score necessarily compares those not with the \"best possible gold standard binding assay truth\", but with \"BELKA truth\". I expect this to put a limit on the best scores achieved.</p>\n<p>To be honest, I don't have a great feel for the micro-averaged precision metric. I thought about comparing it with ROC-AUC.  If I understand correctly, a random model gets 0.008 AP, 0.500 AUC. A perfect mocel gets 1.000 AP, 1.000 AUC. </p>\n<p>The nearest thing to another point on the calibration curve that I have comes from fitting two otherwise equivalent cheminformatics models, one using AP as the metric, the other using AUC. What I found is that the cv for the two models gave: 0.6965 AP, 0.9559 AUC (possibly overfitted). There are a lot of caveats to accepting that as a straight conversion, but if the metric were AUC I certainly wouldn't expect the winning private LB score to be as high as 0.9559 AUC.</p>\n<p>Thus, my speculation is that the ultimate winning private LB score won't be much above 0.700. Do people agree with this, or have different views? </p>",
  "messages": [
    {
      "id": 2814813,
      "postDate": "2024-05-15T15:01:06.200Z",
      "content": "<p>[Written in terms of \"old\" score - hope to update soon]</p>\n<p>\"Just for fun\", I'm speculating about what the ultimate winning LB score (both private and corresponding public are interesting) might be at the end of this competition.</p>\n<p>The technology behind BELKA is really clever and effective, nonetheless it is a high-throughput combinatorial chemistry kind of method and the DNA tags will affect binding to some extent. If I were to imagine a best possible measured set of  gold standard binding affinities from the best conceivable experiments, I'd expect them to differ significantly from the BELKA ones.</p>\n<p>Of the methods tried in this competition, many of the codes have concentrated on a 2D-descriptor-based cheminformatics approach. As of now, this is maxing out in my hands just below 0.600 LB, though some have reported slightly higher scores.</p>\n<p>We may also see more innovative AI approaches based on tokenizing SMILES etc., the limited available information suggests to me that these may also end up scoring around 0.600.</p>\n<p>There's clearly scope for some involvement of docking in this competition, though it's perhaps likely to involve also a degree of AI-based selection of key compounds to dock and prediction of other molecules' scores rather than the expense of docking each and every compound.</p>\n<p>We may also see some MD or similar simulation approaches, but again I'd anticipate this being combined with clever models to select the key compounds to simulate and to predict affinities of others, due to the expense of using an approach like that on a dataset of this size. There may also be other kinds of model that I haven't considered.</p>\n<p>I speculate that the best solutions will be ensembles incorporating \"cheminformatics truth\", \"docking truth\", \"MD truth\", \"generative AI truth\" etc. The LB score necessarily compares those not with the \"best possible gold standard binding assay truth\", but with \"BELKA truth\". I expect this to put a limit on the best scores achieved.</p>\n<p>To be honest, I don't have a great feel for the micro-averaged precision metric. I thought about comparing it with ROC-AUC.  If I understand correctly, a random model gets 0.008 AP, 0.500 AUC. A perfect mocel gets 1.000 AP, 1.000 AUC. </p>\n<p>The nearest thing to another point on the calibration curve that I have comes from fitting two otherwise equivalent cheminformatics models, one using AP as the metric, the other using AUC. What I found is that the cv for the two models gave: 0.6965 AP, 0.9559 AUC (possibly overfitted). There are a lot of caveats to accepting that as a straight conversion, but if the metric were AUC I certainly wouldn't expect the winning private LB score to be as high as 0.9559 AUC.</p>\n<p>Thus, my speculation is that the ultimate winning private LB score won't be much above 0.700. Do people agree with this, or have different views? </p>",
      "rawMarkdown": "[Written in terms of \"old\" score - hope to update soon]\n\n\"Just for fun\", I'm speculating about what the ultimate winning LB score (both private and corresponding public are interesting) might be at the end of this competition.\n\nThe technology behind BELKA is really clever and effective, nonetheless it is a high-throughput combinatorial chemistry kind of method and the DNA tags will affect binding to some extent. If I were to imagine a best possible measured set of  gold standard binding affinities from the best conceivable experiments, I'd expect them to differ significantly from the BELKA ones.\n\nOf the methods tried in this competition, many of the codes have concentrated on a 2D-descriptor-based cheminformatics approach. As of now, this is maxing out in my hands just below 0.600 LB, though some have reported slightly higher scores.\n\nWe may also see more innovative AI approaches based on tokenizing SMILES etc., the limited available information suggests to me that these may also end up scoring around 0.600.\n\nThere's clearly scope for some involvement of docking in this competition, though it's perhaps likely to involve also a degree of AI-based selection of key compounds to dock and prediction of other molecules' scores rather than the expense of docking each and every compound.\n\nWe may also see some MD or similar simulation approaches, but again I'd anticipate this being combined with clever models to select the key compounds to simulate and to predict affinities of others, due to the expense of using an approach like that on a dataset of this size. There may also be other kinds of model that I haven't considered.\n\nI speculate that the best solutions will be ensembles incorporating \"cheminformatics truth\", \"docking truth\", \"MD truth\", \"generative AI truth\" etc. The LB score necessarily compares those not with the \"best possible gold standard binding assay truth\", but with \"BELKA truth\". I expect this to put a limit on the best scores achieved.\n\nTo be honest, I don't have a great feel for the micro-averaged precision metric. I thought about comparing it with ROC-AUC.  If I understand correctly, a random model gets 0.008 AP, 0.500 AUC. A perfect mocel gets 1.000 AP, 1.000 AUC. \n\nThe nearest thing to another point on the calibration curve that I have comes from fitting two otherwise equivalent cheminformatics models, one using AP as the metric, the other using AUC. What I found is that the cv for the two models gave: 0.6965 AP, 0.9559 AUC (possibly overfitted). There are a lot of caveats to accepting that as a straight conversion, but if the metric were AUC I certainly wouldn't expect the winning private LB score to be as high as 0.9559 AUC.\n\nThus, my speculation is that the ultimate winning private LB score won't be much above 0.700. Do people agree with this, or have different views? "
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2814813": "[Written in terms of \"old\" score - hope to update soon]\n\n\"Just for fun\", I'm speculating about what the ultimate winning LB score (both private and corresponding public are interesting) might be at the end of this competition.\n\nThe technology behind BELKA is really clever and effective, nonetheless it is a high-throughput combinatorial chemistry kind of method and the DNA tags will affect binding to some extent. If I were to imagine a best possible measured set of  gold standard binding affinities from the best conceivable experiments, I'd expect them to differ significantly from the BELKA ones.\n\nOf the methods tried in this competition, many of the codes have concentrated on a 2D-descriptor-based cheminformatics approach. As of now, this is maxing out in my hands just below 0.600 LB, though some have reported slightly higher scores.\n\nWe may also see more innovative AI approaches based on tokenizing SMILES etc., the limited available information suggests to me that these may also end up scoring around 0.600.\n\nThere's clearly scope for some involvement of docking in this competition, though it's perhaps likely to involve also a degree of AI-based selection of key compounds to dock and prediction of other molecules' scores rather than the expense of docking each and every compound.\n\nWe may also see some MD or similar simulation approaches, but again I'd anticipate this being combined with clever models to select the key compounds to simulate and to predict affinities of others, due to the expense of using an approach like that on a dataset of this size. There may also be other kinds of model that I haven't considered.\n\nI speculate that the best solutions will be ensembles incorporating \"cheminformatics truth\", \"docking truth\", \"MD truth\", \"generative AI truth\" etc. The LB score necessarily compares those not with the \"best possible gold standard binding assay truth\", but with \"BELKA truth\". I expect this to put a limit on the best scores achieved.\n\nTo be honest, I don't have a great feel for the micro-averaged precision metric. I thought about comparing it with ROC-AUC.  If I understand correctly, a random model gets 0.008 AP, 0.500 AUC. A perfect mocel gets 1.000 AP, 1.000 AUC. \n\nThe nearest thing to another point on the calibration curve that I have comes from fitting two otherwise equivalent cheminformatics models, one using AP as the metric, the other using AUC. What I found is that the cv for the two models gave: 0.6965 AP, 0.9559 AUC (possibly overfitted). There are a lot of caveats to accepting that as a straight conversion, but if the metric were AUC I certainly wouldn't expect the winning private LB score to be as high as 0.9559 AUC.\n\nThus, my speculation is that the ultimate winning private LB score won't be much above 0.700. Do people agree with this, or have different views? "
  }
}