{
  "id": 492846,
  "title": "[LB0.613, low GPU resources!] my experimental results ",
  "url": "/competitions/leash-BELKA/discussion/492846",
  "author_name": "hengck23",
  "post_date": "2024-04-11T06:50:03.937000",
  "votes": 103,
  "comment_count": 126,
  "views": 0,
  "content": "<h2>[Acknowledgement]</h2>\n<p>\"We extend our thanks to HP for providing the HP Z8 Fury Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"<br>\n<a href=\"https://www.hp.com/us-en/workstations/z8-fury.html\" target=\"_blank\">https://www.hp.com/us-en/workstations/z8-fury.html</a></p>\n<p>specifications:</p>\n<ul>\n<li>2x RTX 6000 Ada GPU (48 GB ram each) :  xgboost training of 10M molecule ecfps at 45 minutes!</li>\n<li>Intel® Xeon® W9 Processor / 56cores : fast extraction with rdkit, etc</li>\n<li>256 GB ram</li>\n</ul>\n<hr>\n<h2>[experimental results]</h2>\n<p>useful:</p>\n<p>ecfp is binary vector. so you can decrease memory by using np.packbits, that is about 98m x 256 bytes</p>\n<p>pytorch version of np.unpackbits</p>\n<ul>\n<li><a href=\"https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a\" target=\"_blank\">https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a</a></li>\n<li><a href=\"https://github.com/pytorch/pytorch/issues/32867\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/32867</a></li>\n</ul>\n<p>processed dataset for discussion:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd971971a363e7bf8de359ffdaf72ecf2%2FSelection_077.png?generation=1715120816304983&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2746313,
      "postDate": "2024-04-11T06:50:03.937Z",
      "content": "<h2>[Acknowledgement]</h2>\n<p>\"We extend our thanks to HP for providing the HP Z8 Fury Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"<br>\n<a href=\"https://www.hp.com/us-en/workstations/z8-fury.html\" target=\"_blank\">https://www.hp.com/us-en/workstations/z8-fury.html</a></p>\n<p>specifications:</p>\n<ul>\n<li>2x RTX 6000 Ada GPU (48 GB ram each) :  xgboost training of 10M molecule ecfps at 45 minutes!</li>\n<li>Intel® Xeon® W9 Processor / 56cores : fast extraction with rdkit, etc</li>\n<li>256 GB ram</li>\n</ul>\n<hr>\n<h2>[experimental results]</h2>\n<p>useful:</p>\n<p>ecfp is binary vector. so you can decrease memory by using np.packbits, that is about 98m x 256 bytes</p>\n<p>pytorch version of np.unpackbits</p>\n<ul>\n<li><a href=\"https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a\" target=\"_blank\">https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a</a></li>\n<li><a href=\"https://github.com/pytorch/pytorch/issues/32867\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/32867</a></li>\n</ul>\n<p>processed dataset for discussion:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd971971a363e7bf8de359ffdaf72ecf2%2FSelection_077.png?generation=1715120816304983&amp;alt=media\"></p>",
      "rawMarkdown": "## [Acknowledgement]\n\n\"We extend our thanks to HP for providing the HP Z8 Fury Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"\nhttps://www.hp.com/us-en/workstations/z8-fury.html\n\nspecifications:\n- 2x RTX 6000 Ada GPU (48 GB ram each) :  xgboost training of 10M molecule ecfps at 45 minutes!\n- Intel® Xeon® W9 Processor / 56cores : fast extraction with rdkit, etc\n- 256 GB ram\n\n---\n\n##[experimental results]\nuseful:\n\necfp is binary vector. so you can decrease memory by using np.packbits, that is about 98m x 256 bytes\n\npytorch version of np.unpackbits\n- https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a\n- https://github.com/pytorch/pytorch/issues/32867\n\nprocessed dataset for discussion:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd971971a363e7bf8de359ffdaf72ecf2%2FSelection_077.png?generation=1715120816304983&alt=media)",
      "votes": 100
    },
    {
      "id": 2748792,
      "postDate": "2024-04-12T16:16:15.880Z",
      "content": "<p>dataset for the above discussion have been shared!!!!<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p>In summary:</p>\n<ul>\n<li>train.reduced.parquet : 98_415_610 training SMILES and their information</li>\n<li>train.ecfp4.packed.npz : Features extracted using rdkit  <ul>\n<li>AllChem.GetMorganFingerprintAsBitVect(mol, 2, 2048)                                                    </li>\n<li>repack with np.packbits() to give 98_415_610 x 256 feature matrix</li></ul></li>\n<li>train.bind.npz : 98_415_610 x 3 target matrix</li>\n<li>test.reduced.parquet/ test.ecfp4.packed.npz : similarly processed for the test SMILES</li>\n<li>all_buildingblock.csv: building blocks id used in train.reduced.parquet/test.reduced.parquet</li>\n<li>fold0.parquet: train_share,valid_share,valid_nonshare splits for the experiments in the discussion</li>\n</ul>",
      "rawMarkdown": "dataset for the above discussion have been shared!!!!\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\nIn summary:\n- train.reduced.parquet : 98_415_610 training SMILES and their information\n- train.ecfp4.packed.npz : Features extracted using rdkit  \n   - AllChem.GetMorganFingerprintAsBitVect(mol, 2, 2048)\t\t\t\t\t\t\t\t\t\t\t\t\t\n   - repack with np.packbits() to give 98_415_610 x 256 feature matrix\n- train.bind.npz : 98_415_610 x 3 target matrix\n- test.reduced.parquet/ test.ecfp4.packed.npz : similarly processed for the test SMILES\n- all_buildingblock.csv: building blocks id used in train.reduced.parquet/test.reduced.parquet\n- fold0.parquet: train_share,valid_share,valid_nonshare splits for the experiments in the discussion",
      "votes": 13,
      "replies": [
        {
          "id": 2748795,
          "postDate": "2024-04-12T16:17:43.113Z",
          "content": "<p>i think with proper xgboost tuning and ensemble of good data splits/folds, you can get LB of about 0.575. If you discover good hyper-parameters, please share here! Thanks!</p>",
          "rawMarkdown": "i think with proper xgboost tuning and ensemble of good data splits/folds, you can get LB of about 0.575. If you discover good hyper-parameters, please share here! Thanks!",
          "votes": 3
        }
      ]
    },
    {
      "id": 2746508,
      "postDate": "2024-04-11T09:15:43.460Z",
      "content": "<p>Thanks for sharing ! </p>\n<p>My colleagues and me sometimes organize introductory webinars on Kaggle competitions <br>\nSome video records here: <a href=\"https://www.youtube.com/watch?v=aqUOz3nFYm4&amp;list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw\" target=\"_blank\">https://www.youtube.com/watch?v=aqUOz3nFYm4&amp;list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw</a></p>\n<p>Is there any chance that you can make an introduction to that competition ? <br>\nJust intro survey - data, metrics, public solutions, whatsever.<br>\nIt is not a private sharing - we make  anounces here on Kaggle forum - everybody can join via zoom<br>\nand ask questions/comments/whatsever. Video will be later on Youtube.<br>\nTiming about 40 minutes </p>\n<p>Sometimes organizers join the zoom and make their comments also:<br>\n<a href=\"https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170\" target=\"_blank\">https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170</a></p>\n<p>In general it seems quite helpful for community. <br>\nAnother examples here:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825</a><br>\n<a href=\"https://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975\" target=\"_blank\">https://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975</a></p>",
      "rawMarkdown": "Thanks for sharing ! \n\nMy colleagues and me sometimes organize introductory webinars on Kaggle competitions \nSome video records here: https://www.youtube.com/watch?v=aqUOz3nFYm4&list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw\n\nIs there any chance that you can make an introduction to that competition ? \nJust intro survey - data, metrics, public solutions, whatsever.\nIt is not a private sharing - we make  anounces here on Kaggle forum - everybody can join via zoom\nand ask questions/comments/whatsever. Video will be later on Youtube.\nTiming about 40 minutes \n\nSometimes organizers join the zoom and make their comments also:\nhttps://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170\n\nIn general it seems quite helpful for community. \nAnother examples here:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825\nhttps://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975",
      "votes": 11
    },
    {
      "id": 2765385,
      "postDate": "2024-04-21T05:52:32.497Z",
      "content": "<p>transformer (pure SMILES string) results out !!!!<br>\n<a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">https://huggingface.co/ibm/MoLFormer-XL-both-10pct</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa4b79498cbc6dd566f0dcb15a31d6272%2FSelection_032.png?generation=1713678750507300&amp;alt=media\"></p>",
      "rawMarkdown": "transformer (pure SMILES string) results out !!!!\nhttps://huggingface.co/ibm/MoLFormer-XL-both-10pct\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa4b79498cbc6dd566f0dcb15a31d6272%2FSelection_032.png?generation=1713678750507300&alt=media)",
      "votes": 7,
      "replies": [
        {
          "id": 2804090,
          "postDate": "2024-05-09T19:31:07.377Z",
          "content": "<p>Care to share how much training data this network saw? Or how many FF layers after the transformer? I used 3 layers post-transformer and trained on about half the training data to get a score of ~0.5.</p>",
          "rawMarkdown": "Care to share how much training data this network saw? Or how many FF layers after the transformer? I used 3 layers post-transformer and trained on about half the training data to get a score of ~0.5.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2776878,
      "postDate": "2024-04-26T11:42:08.483Z",
      "content": "<p>one of my idea<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F642a737a82d59e4b1ed1fba2b312cc8e%2FSelection_051.png?generation=1714131726333596&amp;alt=media\"></p>",
      "rawMarkdown": "one of my idea\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F642a737a82d59e4b1ed1fba2b312cc8e%2FSelection_051.png?generation=1714131726333596&alt=media)",
      "votes": 6,
      "replies": [
        {
          "id": 2776899,
          "postDate": "2024-04-26T11:49:45.487Z",
          "content": "<p>A similar idea: if you are interested in multitask learning to help embed the test and train in a similar embedding space, you might want to consider predicting some calculated properties like cLogP (for many enzymes, binding loosely tracks with cLogP) or TPSA (which is a reference to the polar surface area of the molecule). </p>\n<p>RDKit has some tools to calculate descriptors like these, as a start.</p>",
          "rawMarkdown": "A similar idea: if you are interested in multitask learning to help embed the test and train in a similar embedding space, you might want to consider predicting some calculated properties like cLogP (for many enzymes, binding loosely tracks with cLogP) or TPSA (which is a reference to the polar surface area of the molecule). \n\nRDKit has some tools to calculate descriptors like these, as a start.",
          "votes": 1,
          "replies": [
            {
              "id": 2776922,
              "postDate": "2024-04-26T12:02:15.167Z",
              "content": "<p>ref<br>\n<a href=\"https://dl.acm.org/doi/10.1145/3233547.3233548\" target=\"_blank\">https://dl.acm.org/doi/10.1145/3233547.3233548</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3532116953f23315366443db19af4f0d%2FSelection_052.png?generation=1714132907090436&amp;alt=media\"></p>",
              "rawMarkdown": "ref\nhttps://dl.acm.org/doi/10.1145/3233547.3233548\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3532116953f23315366443db19af4f0d%2FSelection_052.png?generation=1714132907090436&alt=media)",
              "votes": 1
            },
            {
              "id": 2776927,
              "postDate": "2024-04-26T12:05:51.517Z",
              "content": "<p>\" calculated properties like cLogP  …\"</p>\n<p>There are a couple of possible self-supervised task</p>\n<ul>\n<li>MLM : masked language modeling</li>\n<li>MTR: mult task regression (e.g. cLoP, …. i think anything related to shape and surface, electricty, energy)</li>\n<li>chemical reaction</li>\n</ul>\n<p>thanks!</p>",
              "rawMarkdown": "\" calculated properties like cLogP  ...\"\n\nThere are a couple of possible self-supervised task\n- MLM : masked language modeling\n- MTR: mult task regression (e.g. cLoP, .... i think anything related to shape and surface, electricty, energy)\n- chemical reaction\n\nthanks!",
              "votes": 1
            },
            {
              "id": 2780098,
              "postDate": "2024-04-28T03:47:15.373Z",
              "content": "<p>yet another way to pretrain with test data …. basically it is \"smart clustering\" via VAE<br>\nlikewise, denoised diffusion also work (imagine train --&gt; noise --&gt; test)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbff4b909984bfcc298489abd9665ca16%2FSelection_055.png?generation=1714275971114985&amp;alt=media\"></p>",
              "rawMarkdown": "yet another way to pretrain with test data .... basically it is \"smart clustering\" via VAE\nlikewise, denoised diffusion also work (imagine train --> noise --> test)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbff4b909984bfcc298489abd9665ca16%2FSelection_055.png?generation=1714275971114985&alt=media)\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2746350,
      "postDate": "2024-04-11T07:07:40.860Z",
      "content": "<p>With you around this comp' would be extra fun </p>",
      "rawMarkdown": "With you around this comp' would be extra fun ",
      "votes": 5
    },
    {
      "id": 2747013,
      "postDate": "2024-04-11T15:41:38.363Z",
      "content": "<p>useful software tools. implementations, etc:</p>\n<ol>\n<li>multi-process df apply:<br>\n<a href=\"https://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores\" target=\"_blank\">https://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores</a><br>\n<a href=\"https://github.com/ddelange/mapply\" target=\"_blank\">https://github.com/ddelange/mapply</a></li>\n</ol>\n<pre><code>import mapply\nmapply.init(\n    =50,#-1,\n    =,\n)\ndf[] = df[].mapply(to_ecfp4_fun)\n</code></pre>",
      "rawMarkdown": "useful software tools. implementations, etc:\n1. multi-process df apply:\nhttps://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores\nhttps://github.com/ddelange/mapply\n\n```\nimport mapply\nmapply.init(\n    n_workers=50,#-1,\n    progressbar=True,\n)\ndf['ecfp'] = df['molecule_smiles'].mapply(to_ecfp4_fun)\n```",
      "votes": 6,
      "replies": [
        {
          "id": 2747046,
          "postDate": "2024-04-11T16:16:54.933Z",
          "content": "<p>how to write fast submission code … run in minutes</p>\n<pre><code>probability = xgb(x_valid)\n = (, ) \n\nid = reduced_df]()\n\n\nid = id(-)\nprobability = probability(-)\n = id!=-\n==)\n\nsubmit_df = pd({\n    :id,\n    :probability,\n})\n submit_df(f,index=False)\n</code></pre>\n<p>reduced_df:<br>\n(id is -1 if protein is not evaluated)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b3628a98221b1f366382dc9e56e9ea2%2FSelection_022.png?generation=1712852204921481&amp;alt=media\"></p>",
          "rawMarkdown": "how to write fast submission code ... run in minutes\n\n```\nprobability = xgb.predict_proba(x_valid)\n#probability.shape = (878022, 3) \n\nid = reduced_df[['id_BRD4', 'id_HSA', 'id_sEH']].to_numpy()\nassert(probability.shape==id.shape)\n\nid = id.reshape(-1)\nprobability = probability.reshape(-1)\nmask = id!=-1\nassert(mask.sum()==1674896)\n\nsubmit_df = pd.DataFrame({\n\t'id':id[mask],\n\t'binds':probability[mask],\n})\n submit_df.to_csv(f'submission.csv',index=False)\n```\n\nreduced\\_df:\n(id is -1 if protein is not evaluated)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b3628a98221b1f366382dc9e56e9ea2%2FSelection_022.png?generation=1712852204921481&alt=media)",
          "votes": 4,
          "replies": [
            {
              "id": 2759822,
              "postDate": "2024-04-18T22:49:40.710Z",
              "content": "<p>This is brilliant. Thank you for sharing!</p>",
              "rawMarkdown": "This is brilliant. Thank you for sharing!"
            }
          ]
        },
        {
          "id": 2747524,
          "postDate": "2024-04-11T23:58:38.013Z",
          "content": "<p>This is pretty neat! Having the progress bar is nice too.</p>",
          "rawMarkdown": "This is pretty neat! Having the progress bar is nice too.",
          "votes": 1
        },
        {
          "id": 2748892,
          "postDate": "2024-04-12T17:59:28.237Z",
          "content": "<ol>\n<li>Easy Custom Losses for Tree Boosters using PyTorch<br>\n<a href=\"https://towardsdatascience.com/easy-custom-losses-for-tree-boosters-using-pytorch-57ffaa0b2eb3\" target=\"_blank\">https://towardsdatascience.com/easy-custom-losses-for-tree-boosters-using-pytorch-57ffaa0b2eb3</a><br>\n<a href=\"https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\" target=\"_blank\">https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec</a></li>\n</ol>",
          "rawMarkdown": "2. Easy Custom Losses for Tree Boosters using PyTorch\nhttps://towardsdatascience.com/easy-custom-losses-for-tree-boosters-using-pytorch-57ffaa0b2eb3\nhttps://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec",
          "votes": 2
        }
      ]
    },
    {
      "id": 2799761,
      "postDate": "2024-05-07T23:01:47.243Z",
      "content": "<p>i just realize that you may have more training data</p>\n<ol>\n<li>molecule + bind label</li>\n<li>building block + soft bind label (average over all molecule using that block). you can treat it as molecule=block+wildcard here</li>\n</ol>",
      "rawMarkdown": "i just realize that you may have more training data\n1. molecule + bind label\n2. building block + soft bind label (average over all molecule using that block). you can treat it as molecule=block+wildcard here",
      "votes": 3
    },
    {
      "id": 2766963,
      "postDate": "2024-04-22T03:34:06.023Z",
      "content": "<p>a pretty smart method that uses both labelled and unlabelled data</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa103fefd904bcddb90413a393dee5344%2FSelection_033.png?generation=1713756844349343&amp;alt=media\"></p>",
      "rawMarkdown": "a pretty smart method that uses both labelled and unlabelled data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa103fefd904bcddb90413a393dee5344%2FSelection_033.png?generation=1713756844349343&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2767009,
          "postDate": "2024-04-22T04:25:42.573Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F57be4534a687d8f516f898bc3aa77eff%2FScreenshot%202024-04-22%20at%2012-23-50%20ICAN%20Interpretable%20cross-attention%20network%20for%20identifying%20drug%20and%20target%20protein%20interactions%20-%20journal.pone.0276609.pdf.png?generation=1713759870859276&amp;alt=media\"></p>\n<p>Another paper describing a very similar approach:<br>\nICAN: Interpretable cross-attention network for identifying drug and target protein interactions</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F57be4534a687d8f516f898bc3aa77eff%2FScreenshot%202024-04-22%20at%2012-23-50%20ICAN%20Interpretable%20cross-attention%20network%20for%20identifying%20drug%20and%20target%20protein%20interactions%20-%20journal.pone.0276609.pdf.png?generation=1713759870859276&alt=media)\n\nAnother paper describing a very similar approach:\nICAN: Interpretable cross-attention network for identifying drug and target protein interactions",
          "votes": 2,
          "replies": [
            {
              "id": 2767112,
              "postDate": "2024-04-22T05:51:08.660Z",
              "content": "<p>there are a couple of similar papers. best to get those with github codes and experiment results from well known dataset. I foresee most deep net solution will be slow. how to sample training data could be an issue</p>",
              "rawMarkdown": "there are a couple of similar papers. best to get those with github codes and experiment results from well known dataset. I foresee most deep net solution will be slow. how to sample training data could be an issue"
            },
            {
              "id": 2767685,
              "postDate": "2024-04-22T13:36:13.940Z",
              "content": "<p>Agreed, efficient use of compute will be key.</p>",
              "rawMarkdown": "Agreed, efficient use of compute will be key."
            },
            {
              "id": 2817098,
              "postDate": "2024-05-16T17:41:49.310Z",
              "content": "<p>With only three types of proteins, is there too little protein data? If so, how can we tackle this problem?</p>",
              "rawMarkdown": "With only three types of proteins, is there too little protein data? If so, how can we tackle this problem?"
            }
          ]
        },
        {
          "id": 2768571,
          "postDate": "2024-04-23T00:09:17.210Z",
          "content": "<p>a few good github repo that i am using:<br>\n<a href=\"https://github.com/larngroup/DTITR\" target=\"_blank\">https://github.com/larngroup/DTITR</a><br>\n<a href=\"https://github.com/ZXT0212/CAT-DTI\" target=\"_blank\">https://github.com/ZXT0212/CAT-DTI</a><br>\n<a href=\"https://github.com/peizhenbai/DrugBAN\" target=\"_blank\">https://github.com/peizhenbai/DrugBAN</a></p>",
          "rawMarkdown": "a few good github repo that i am using:\nhttps://github.com/larngroup/DTITR\nhttps://github.com/ZXT0212/CAT-DTI\nhttps://github.com/peizhenbai/DrugBAN\n\n",
          "votes": 2,
          "replies": [
            {
              "id": 2770751,
              "postDate": "2024-04-24T02:37:49.427Z",
              "content": "<p>Great sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!! Did you get to try any of above methods yet for this competition?</p>",
              "rawMarkdown": "Great sharing @hengck23!! Did you get to try any of above methods yet for this competition?"
            }
          ]
        }
      ]
    },
    {
      "id": 2765165,
      "postDate": "2024-04-21T01:43:43.427Z",
      "content": "<p>yet another external data!!!!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4137ec834e1c8fd91b077bc2745775f%2FSelection_030.png?generation=1713663808992163&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4190bd04b74691820e74f3824afc4fa%2FSelection_031.png?generation=1713663821080691&amp;alt=media\"></p>",
      "rawMarkdown": "yet another external data!!!!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4137ec834e1c8fd91b077bc2745775f%2FSelection_030.png?generation=1713663808992163&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4190bd04b74691820e74f3824afc4fa%2FSelection_031.png?generation=1713663821080691&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2770487,
          "postDate": "2024-04-23T21:14:40.463Z",
          "content": "<p>I'll take a look at generating a dataset from this close to our competition data when I have time (likely this weekend). Thanks for finding!</p>",
          "rawMarkdown": "I'll take a look at generating a dataset from this close to our competition data when I have time (likely this weekend). Thanks for finding!",
          "replies": [
            {
              "id": 2770565,
              "postDate": "2024-04-23T22:34:17.323Z",
              "content": "<p>i have been looking at other papers and dataset. i think it actually make more sense to convert kaggle dataset to other dataset, i.e. replacing [Dy]</p>",
              "rawMarkdown": "i have been looking at other papers and dataset. i think it actually make more sense to convert kaggle dataset to other dataset, i.e. replacing [Dy]"
            },
            {
              "id": 2770636,
              "postDate": "2024-04-24T00:28:09.887Z",
              "content": "<p>Have you found consistency in what they use to model the DNA attachment point? I've done a little experimenting on my end switching Dy out for polyethyleneglycol like linkers and from a modeling perspective, I've seen similar results so far (which surprised me a little).</p>\n<p>One thing I have not tried is modeling the DNA attachment point as something small like a methyl group. Intuitively I'd think that providing the model some information on where the DNA attached would be valuable, but perhaps it's not.</p>",
              "rawMarkdown": "Have you found consistency in what they use to model the DNA attachment point? I've done a little experimenting on my end switching Dy out for polyethyleneglycol like linkers and from a modeling perspective, I've seen similar results so far (which surprised me a little).\n\nOne thing I have not tried is modeling the DNA attachment point as something small like a methyl group. Intuitively I'd think that providing the model some information on where the DNA attached would be valuable, but perhaps it's not."
            },
            {
              "id": 2770758,
              "postDate": "2024-04-24T02:46:36.887Z",
              "content": "<p>I've been using your C/methyl conversion in my data pipeline (rest in progress).<br>\nIs switching methyl/linker to PEG as simple as swapping the C/Dy for OCCOCCO (PEG smile)?</p>\n<p>Also thanks for all the domain knowledge you've shared!</p>",
              "rawMarkdown": "I've been using your C/methyl conversion in my data pipeline (rest in progress).\nIs switching methyl/linker to PEG as simple as swapping the C/Dy for OCCOCCO (PEG smile)?\n\nAlso thanks for all the domain knowledge you've shared!"
            },
            {
              "id": 2771702,
              "postDate": "2024-04-24T11:23:48.150Z",
              "content": "<p>It's a little bit more complicated than that, but not much. I plan to post a notebook about it this weekend.</p>\n<p>Also, you are very welcome! Happy to help and share some knowledge of chemistry.</p>",
              "rawMarkdown": "It's a little bit more complicated than that, but not much. I plan to post a notebook about it this weekend.\n\nAlso, you are very welcome! Happy to help and share some knowledge of chemistry.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2748338,
      "postDate": "2024-04-12T11:49:45.287Z",
      "content": "<p>anyone one to make baseline for \"DiffDock\"?</p>",
      "rawMarkdown": "anyone one to make baseline for \"DiffDock\"?",
      "votes": 3,
      "replies": [
        {
          "id": 2748932,
          "postDate": "2024-04-12T18:14:32.807Z",
          "content": "<p>There is a <a href=\"https://github.com/suneelbvs/DiffDock\" target=\"_blank\">Colab diffdock notebook</a> for anyone that is up for the challenge. With some tinkering, it can probably also be made to run on Kaggle.</p>",
          "rawMarkdown": "There is a [Colab diffdock notebook](https://github.com/suneelbvs/DiffDock) for anyone that is up for the challenge. With some tinkering, it can probably also be made to run on Kaggle.",
          "votes": 1,
          "replies": [
            {
              "id": 2749654,
              "postDate": "2024-04-13T07:05:43.363Z",
              "content": "<p>DiffDock has a new version very recently, so I wouldn't try a colab notebook from 2 years ago except maybe as reference</p>",
              "rawMarkdown": "DiffDock has a new version very recently, so I wouldn't try a colab notebook from 2 years ago except maybe as reference",
              "votes": 1
            },
            {
              "id": 2749756,
              "postDate": "2024-04-13T08:18:36.173Z",
              "content": "<p>Good point, so maybe a lot of tinkering hehe</p>",
              "rawMarkdown": "Good point, so maybe a lot of tinkering hehe"
            }
          ]
        },
        {
          "id": 2748942,
          "postDate": "2024-04-12T18:20:17.240Z",
          "content": "<p>I am working on building structural models. Not with DiffDock, but more traditional physics-based docking engines.</p>",
          "rawMarkdown": "I am working on building structural models. Not with DiffDock, but more traditional physics-based docking engines.",
          "votes": 1
        },
        {
          "id": 2850582,
          "postDate": "2024-06-02T08:25:00.930Z",
          "content": "<p>I recently made it to run, the thing is, it is very slow! Impossible to run so many compounds</p>",
          "rawMarkdown": "I recently made it to run, the thing is, it is very slow! Impossible to run so many compounds"
        }
      ]
    },
    {
      "id": 2768605,
      "postDate": "2024-04-23T00:41:20.493Z",
      "content": "<p>yet another results with Topological Torsion Fingerprint.<br>\nIt appears to me that just pure SMILE string model performance limit is 0.67 (for same block).</p>\n<p>there aren't huge different for transformer based, fingerprint based, etc ….<br>\nthe data has very very long tails …. (i.e. the positive cases don't cluster together in the chemical space to gives high peak density, instead they are spreading around)<br>\nThis make it difficult to generalized to unseen blocks, especially for fingerprint methods …</p>\n<p>need to think of new strategy</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e01cbdf17b3bb57233617c6cb3f5135%2FSelection_035.png?generation=1713832518151318&amp;alt=media\"> </p>",
      "rawMarkdown": "yet another results with Topological Torsion Fingerprint.\nIt appears to me that just pure SMILE string model performance limit is 0.67 (for same block).\n\nthere aren't huge different for transformer based, fingerprint based, etc ....\nthe data has very very long tails .... (i.e. the positive cases don't cluster together in the chemical space to gives high peak density, instead they are spreading around)\nThis make it difficult to generalized to unseen blocks, especially for fingerprint methods ...\n\nneed to think of new strategy\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e01cbdf17b3bb57233617c6cb3f5135%2FSelection_035.png?generation=1713832518151318&alt=media) ",
      "votes": 4
    },
    {
      "id": 2767460,
      "postDate": "2024-04-22T10:21:56.747Z",
      "content": "<p>average precision trick?<br>\n<a href=\"https://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596</a></p>\n<p>\"This works because Average Precision metric is always improved when adding more predictions if the new predictions have a lower score than all previous predictions.\"</p>\n<p>hmm</p>\n<p>actually there is another way to look at our problem:<br>\nstage1:  predict no binding <br>\nstage2: predict binding score/rank (only less than 1% of data from stage1. enable use of more time consuming methods?)</p>",
      "rawMarkdown": "average precision trick?\nhttps://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596\n\n\"This works because Average Precision metric is always improved when adding more predictions if the new predictions have a lower score than all previous predictions.\"\n\nhmm\n\nactually there is another way to look at our problem:\nstage1:  predict no binding \nstage2: predict binding score/rank (only less than 1% of data from stage1. enable use of more time consuming methods?)",
      "votes": 4
    },
    {
      "id": 2758038,
      "postDate": "2024-04-17T22:18:20.453Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, what are your positive-to-negative ratios in your val sets? As was discussed <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/492126#2742181\" target=\"_blank\">here</a>, the metric strongly depends on the ratio. I suggest enforcing a known, constant ratio to facilitate easy comparison between results. The obvious one is the 1/128 ratio since this is the test set ratio…with this, I currently have validation scores of 0.53 for the all-BBs_shared val set, 0.29 for at least one BB is not shared and at least one BB is shared val set, and 0.018 for truly none of BBs is shared…yea the last one is tough lol. Probably different core would be even harder.</p>\n<p>EDIT: I got confused, the ratio should be 1/125. </p>",
      "rawMarkdown": "@hengck23, what are your positive-to-negative ratios in your val sets? As was discussed [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/492126#2742181), the metric strongly depends on the ratio. I suggest enforcing a known, constant ratio to facilitate easy comparison between results. The obvious one is the 1/128 ratio since this is the test set ratio...with this, I currently have validation scores of 0.53 for the all-BBs_shared val set, 0.29 for at least one BB is not shared and at least one BB is shared val set, and 0.018 for truly none of BBs is shared...yea the last one is tough lol. Probably different core would be even harder.\n\nEDIT: I got confused, the ratio should be 1/125. ",
      "votes": 4,
      "replies": [
        {
          "id": 2758217,
          "postDate": "2024-04-18T03:10:38.957Z",
          "content": "<p>be careful!  1/125 is only for public.</p>\n<p>always keep in mind: you have full access of the test data. you can measure some magic statistics on it. this is the trick to winning</p>",
          "rawMarkdown": "be careful!  1/125 is only for public.\n\nalways keep in mind: you have full access of the test data. you can measure some magic statistics on it. this is the trick to winning",
          "votes": 2,
          "replies": [
            {
              "id": 2758728,
              "postDate": "2024-04-18T10:11:41.747Z",
              "content": "<p>Will test.csv be the actual public leaderboard test? Or what do you mean?</p>",
              "rawMarkdown": "Will test.csv be the actual public leaderboard test? Or what do you mean?",
              "votes": 1
            },
            {
              "id": 2765402,
              "postDate": "2024-04-21T06:00:26.990Z",
              "content": "<p>Yes the test data is the full test data, including public and private leaderboard components <a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> </p>",
              "rawMarkdown": "Yes the test data is the full test data, including public and private leaderboard components @gyulamaloveczky4 ",
              "votes": 1
            },
            {
              "id": 2768461,
              "postDate": "2024-04-22T21:09:27.973Z",
              "content": "<p>a submission of zero (or other constant) will revealed the percentage of all positive cases.</p>\n<p>what happen if you submit a couple of different constants, e.g.</p>\n<ul>\n<li>different values for different protein, different block,  etc</li>\n</ul>\n<p>hint :google for probing the kaggle leaderboard for average precision </p>",
              "rawMarkdown": "a submission of zero (or other constant) will revealed the percentage of all positive cases.\n\nwhat happen if you submit a couple of different constants, e.g.\n- different values for different protein, different block,  etc\n\nhint :google for probing the kaggle leaderboard for average precision ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2758737,
          "postDate": "2024-04-18T10:16:50.893Z",
          "content": "<p>AP is a great metric because you can measure ranking performance of your model regardless of the pos/neg ratio, AP CV should correlate with LB even if pos/neg ratio is a bit shifted (we don't really care about absolute values of probabilities here)</p>",
          "rawMarkdown": "AP is a great metric because you can measure ranking performance of your model regardless of the pos/neg ratio, AP CV should correlate with LB even if pos/neg ratio is a bit shifted (we don't really care about absolute values of probabilities here)",
          "replies": [
            {
              "id": 2758939,
              "postDate": "2024-04-18T11:58:56.593Z",
              "content": "<p>True, as long as you use the same validation set. But if you want to compare the validation scores of two different sets (or people) you need to use the same ratio</p>",
              "rawMarkdown": "True, as long as you use the same validation set. But if you want to compare the validation scores of two different sets (or people) you need to use the same ratio",
              "votes": 1
            },
            {
              "id": 2769342,
              "postDate": "2024-04-23T09:43:06.153Z",
              "content": "<p>ratio squad?</p>",
              "rawMarkdown": "ratio squad?"
            }
          ]
        }
      ]
    },
    {
      "id": 2754473,
      "postDate": "2024-04-16T04:28:05.317Z",
      "content": "<p>i just realise that we need to probe the num of non train blocks in public test first. this can be done by two submission. a normal one and another with non train block masked off with magic value and observe the difference. </p>\n<p>we need to verify private and public share the same distribution of train and non train block.</p>",
      "rawMarkdown": "i just realise that we need to probe the num of non train blocks in public test first. this can be done by two submission. a normal one and another with non train block masked off with magic value and observe the difference. \n\nwe need to verify private and public share the same distribution of train and non train block.",
      "votes": 4,
      "replies": [
        {
          "id": 2754528,
          "postDate": "2024-04-16T05:01:48.073Z",
          "content": "<p>There's like 34 unique buildingblock1s, with triazine core, not in train. And there's like 36 unique buildingblock1s, with different core, not in train. </p>\n<p>Depends how many subs you want to spend, but it would be useful to know, for each of those 70 unique BBs, and for each protein target, so 70x3, are some in public test? Are all in public test? Maybe half are in public test, half in private? If something like half and half, but no overlap (so public LB doesn't help \"learn\" the unseen private LB) then, maybe you can find out the groupings just by making a graph of which unique BBs are included with which other unique BBs. </p>\n<p>Example:<br>\nBB1: 1,2,3,4<br>\nBB2: 5-12<br>\nBB3: 13-20</p>\n<p>Test1: 1,5,13<br>\nTest2: 1,5,14<br>\n…<br>\nTestN: 2,5,13<br>\n1 and 2 both \"connect\" to 13, so they're in the same group. But maybe 5,13,14 etc never pair with 3 or 4, that would imply that one group (1,2,5,13,14,…) is either public or private, and the other group (3,4,…) is the other leaderboard group</p>",
          "rawMarkdown": "There's like 34 unique buildingblock1s, with triazine core, not in train. And there's like 36 unique buildingblock1s, with different core, not in train. \n\nDepends how many subs you want to spend, but it would be useful to know, for each of those 70 unique BBs, and for each protein target, so 70x3, are some in public test? Are all in public test? Maybe half are in public test, half in private? If something like half and half, but no overlap (so public LB doesn't help \"learn\" the unseen private LB) then, maybe you can find out the groupings just by making a graph of which unique BBs are included with which other unique BBs. \n\nExample:\nBB1: 1,2,3,4\nBB2: 5-12\nBB3: 13-20\n\nTest1: 1,5,13\nTest2: 1,5,14\n...\nTestN: 2,5,13\n1 and 2 both \"connect\" to 13, so they're in the same group. But maybe 5,13,14 etc never pair with 3 or 4, that would imply that one group (1,2,5,13,14,...) is either public or private, and the other group (3,4,...) is the other leaderboard group\n\n",
          "replies": [
            {
              "id": 2755986,
              "postDate": "2024-04-16T19:11:01.533Z",
              "content": "<p>For triazines in test with BBs not in train, there's two separate blocks:<br>\n17 BB_a, 36 BB_b/c. (presumably public LB)<br>\n17 BB_a, 36 BB_b/c. (presumably private LB)</p>\n<p>And every possible combination is:<br>\n2 groups, times 3 proteins, times 17 BB_a, times 36 BB_b/c 2 combinations, the same BB is allowed to be taken twice, so it's:<br>\n2x3x17x(36+1)x36/2 = 67,932. Which is exactly the number of rows I observed. So that checks out.</p>\n<p>However, for the other group of 500,000 rows, it all appears to be one grouping. So I think this critical group we will probably find is NOT in public test at all! 500,000 rows is 167k per protein. Since the organizers said:</p>\n<blockquote>\n  <p>We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<p>and 360K - 200K = 160K, so I suspect the public and private are equivalent EXCEPT for this big curveball of the non-triazine molecules with unseen BBs will be only in private LB most likely.</p>\n<p>If someone proves this, please post back here!</p>",
              "rawMarkdown": "For triazines in test with BBs not in train, there's two separate blocks:\n17 BB_a, 36 BB_b/c. (presumably public LB)\n17 BB_a, 36 BB_b/c. (presumably private LB)\n\nAnd every possible combination is:\n2 groups, times 3 proteins, times 17 BB_a, times 36 BB_b/c 2 combinations, the same BB is allowed to be taken twice, so it's:\n2x3x17x(36+1)x36/2 = 67,932. Which is exactly the number of rows I observed. So that checks out.\n\nHowever, for the other group of 500,000 rows, it all appears to be one grouping. So I think this critical group we will probably find is NOT in public test at all! 500,000 rows is 167k per protein. Since the organizers said:\n\n> We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein\n\nand 360K - 200K = 160K, so I suspect the public and private are equivalent EXCEPT for this big curveball of the non-triazine molecules with unseen BBs will be only in private LB most likely.\n\nIf someone proves this, please post back here!\n\n",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2780103,
      "postDate": "2024-04-28T03:51:08.867Z",
      "content": "<p>this can be used for augmentation!<br>\n<a href=\"https://chao1224.github.io/MoleculeSTM\" target=\"_blank\">https://chao1224.github.io/MoleculeSTM</a></p>",
      "rawMarkdown": "this can be used for augmentation!\nhttps://chao1224.github.io/MoleculeSTM",
      "votes": 1
    },
    {
      "id": 2774731,
      "postDate": "2024-04-25T10:00:15.547Z",
      "content": "<p>In the specifications of your HP Z8 Fury Data Science Workstation you have \"256 MB random\".  Is this a typo that should be \"256 GB random\"?</p>",
      "rawMarkdown": "In the specifications of your HP Z8 Fury Data Science Workstation you have \"256 MB random\".  Is this a typo that should be \"256 GB random\"?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2774754,
          "postDate": "2024-04-25T10:10:46.637Z",
          "content": "<p>oops! in <br>\nGB. thanks!</p>",
          "rawMarkdown": "oops! in \nGB. thanks!"
        }
      ]
    },
    {
      "id": 2760019,
      "postDate": "2024-04-19T04:11:48.083Z",
      "content": "<p>for the past few days i have been training neural nets (e.g. string based transformer). Initial results seems to suggest to me that </p>\n<ul>\n<li>the choice of building blocks are not random.  </li>\n<li>there are too much data … some active learning or sampling is required (if we know how the building blocks are generated, it would be useful. e.g. rule based or the host uses some ML/AI-assisted software)</li>\n</ul>",
      "rawMarkdown": "for the past few days i have been training neural nets (e.g. string based transformer). Initial results seems to suggest to me that \n- the choice of building blocks are not random.  \n- there are too much data ... some active learning or sampling is required (if we know how the building blocks are generated, it would be useful. e.g. rule based or the host uses some ML/AI-assisted software)",
      "votes": 1,
      "replies": [
        {
          "id": 2770641,
          "postDate": "2024-04-24T00:32:13.700Z",
          "content": "<p>I have not tried to figure out how the authors chose their dataset yet, but I do know that it's quite common to choose building blocks based on their calculated physical properties. This way, the dataset is biased towards final molecules with properties that more closely resemble drugs. For readings on this, perhaps Lipinski's rules would be a good start for some basic properties you'd like to see in a final molecule (these are common guidelines, not really hard and fast rules).</p>",
          "rawMarkdown": "I have not tried to figure out how the authors chose their dataset yet, but I do know that it's quite common to choose building blocks based on their calculated physical properties. This way, the dataset is biased towards final molecules with properties that more closely resemble drugs. For readings on this, perhaps Lipinski's rules would be a good start for some basic properties you'd like to see in a final molecule (these are common guidelines, not really hard and fast rules).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2754293,
      "postDate": "2024-04-15T23:52:51.580Z",
      "content": "<p>DTI plug and play :<br>\n<a href=\"https://github.com/yazdanimehdi/DeepDrugDomain\" target=\"_blank\">https://github.com/yazdanimehdi/DeepDrugDomain</a></p>",
      "rawMarkdown": "DTI plug and play :\nhttps://github.com/yazdanimehdi/DeepDrugDomain",
      "votes": 1
    },
    {
      "id": 2753393,
      "postDate": "2024-04-15T13:16:53.537Z",
      "content": "<p>How do you construct the label for multi-label classification? <br>\nSome of the molecules bind to the multiple proteins, and XGBoost doesn't support this setup</p>\n<pre><code>pd.).value\n    \n     \n       \n          \nName: count, dtype: \n</code></pre>",
      "rawMarkdown": "How do you construct the label for multi-label classification? \nSome of the molecules bind to the multiple proteins, and XGBoost doesn't support this setup\n```\npd.Series(train_bind.sum(axis=1)).value_counts()\n0    96905831\n1     1429714\n2       80003\n3          62\nName: count, dtype: int64\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2753491,
          "postDate": "2024-04-15T14:13:20.047Z",
          "content": "<p>for cpu:  <br>\n<a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html\" target=\"_blank\">https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html</a><br>\nStarting from version 1.6, XGBoost has experimental support for multi-output regression and multi-label classification with Python package.</p>\n<p>for gpu:  <br>\nyou can write your custom loss. (e.g. kl divergence loss)<br>\nbut interesting, i find that the following 3 results are almost the same</p>\n<ul>\n<li>one model,  input = [feature + protein feature] where  protein feature =[1,0,0],[0,1,0],[0,0,1]</li>\n<li>three model, each is binary classification for each protein</li>\n<li>one model, just set target= Lx3 array and objective=\"binary:logistic\"<br>\nXGBoost uses one-vs-rest method<br>\n(<a href=\"https://discuss.xgboost.ai/t/understanding-leaf-values/1840\" target=\"_blank\">https://discuss.xgboost.ai/t/understanding-leaf-values/1840</a>)</li>\n</ul>\n<p>you can try a toy example, e.g. y =np.random.choice(2,(100,3)) and overfit a xgbost model. you should see it can gives multi-label output using the above methods.</p>",
          "rawMarkdown": "for cpu:  \nhttps://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html\nStarting from version 1.6, XGBoost has experimental support for multi-output regression and multi-label classification with Python package.\n\nfor gpu:  \nyou can write your custom loss. (e.g. kl divergence loss)\nbut interesting, i find that the following 3 results are almost the same\n- one model,  input = [feature + protein feature] where  protein feature =[1,0,0],[0,1,0],[0,0,1]\n- three model, each is binary classification for each protein\n- one model, just set target= Lx3 array and objective=\"binary:logistic\"\nXGBoost uses one-vs-rest method\n(https://discuss.xgboost.ai/t/understanding-leaf-values/1840)\n\nyou can try a toy example, e.g. y =np.random.choice(2,(100,3)) and overfit a xgbost model. you should see it can gives multi-label output using the above methods.",
          "votes": 3,
          "replies": [
            {
              "id": 2753538,
              "postDate": "2024-04-15T14:51:13.407Z",
              "content": "<pre><code>    x_train = np.random.choice(256,(20,256)).astype(np.uint8)\n    y_train = np.random.choice(2,(20,3)).astype(np.uint8)\n\n    clf = XGBClassifier(\n        =6,\n        =100,\n        # =,\n        =,\n        =,\n        objective  =, #, \n        learning_rate= 0.3,\n        device = ,\n        =42,\n        =1,\n    )\n\n    clf.fit(\n        x_train, y_train,\n        =\n    )\n\n    probability = clf.predict_proba(x_train)\n    predict = (probability&gt;0.5).astype(np.uint8)\n    (probability)\n    (predict)\n    (y_train)\n    ((y_train!=predict).sum())\n</code></pre>\n<p>y train</p>\n<pre><code>\n\n\n\n\n\n\n\n</code></pre>\n<p>probability</p>\n<pre><code>[[   ]\n [     ]\n [  ]\n [     ]\n [  ]\n [  ]\n [      ]\n [   ]\ny_train!=predict).sum() gives zero\n</code></pre>",
              "rawMarkdown": "```\n\tx_train = np.random.choice(256,(20,256)).astype(np.uint8)\n\ty_train = np.random.choice(2,(20,3)).astype(np.uint8)\n\n\tclf = XGBClassifier(\n\t\tmax_depth=6,\n\t\tn_estimators=100,\n\t\t# multi_strategy=\"multi_output_tree\",\n\t\tbooster='gbtree',\n\t\ttree_method='hist',\n\t\tobjective  ='binary:logistic', #'reg:logistic', \n\t\tlearning_rate= 0.3,\n\t\tdevice = 'cuda:1',\n\t\trandom_state=42,\n\t\tverbosity=1,\n\t)\n\n\tclf.fit(\n\t\tx_train, y_train,\n\t\tverbose=True\n\t)\n\n\tprobability = clf.predict_proba(x_train)\n\tpredict = (probability>0.5).astype(np.uint8)\n\tprint(probability)\n\tprint(predict)\n\tprint(y_train)\n\tprint((y_train!=predict).sum())\n\n```\n\ny train\n```\n[[0 0 1]\n [0 0 0]\n [0 0 0]\n [1 0 1]\n [1 0 0]\n [0 0 0]\n [1 1 1]\n [1 1 1]\n\n```\nprobability\n```\n[[0.09288329 0.10505337 0.9170576 ]\n [0.1229272  0.114103   0.05714793]\n [0.11309299 0.08702122 0.08370278]\n [0.8736928  0.086165   0.74388474]\n [0.80896735 0.03677367 0.12687069]\n [0.18387493 0.14618503 0.04674026]\n [0.9360719  0.7926505  0.861303  ]\n [0.92951655 0.7828865  0.94130194]\ny_train!=predict).sum() gives zero\n```",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2747361,
      "postDate": "2024-04-11T20:19:02.817Z",
      "content": "<p>since it is micro average precision, one has to be careful in calibration, class-balancing or multi-class/single-class setup:</p>\n<pre><code>probability=np.([\n,\n,\n,\n,\n,\n,\n,\n,\n,\n])\ny=np.([\n,\n,\n,\n,\n,\n,\n,\n,\n,\n])\nscore (all micro) \nscore (, binds_BRD4) \nscore (, binds_HSA) \nscore (, binds_sEH) \n</code></pre>",
      "rawMarkdown": "since it is micro average precision, one has to be careful in calibration, class-balancing or multi-class/single-class setup:\n```\n\nprobability=np.array([\n\t[0.9,0.2,0.2],\n\t[0.9,0.2,0.2],\n\t[0.9,0.2,0.2],\n\t[0.6,0.3,0.2],\n\t[0.5,0.3,0.2],\n\t[0.1,0.3,0.2],\n\t[0.1,0.2,0.5],\n\t[0.1,0.2,0.5],\n\t[0.1,0.2,0.5],\n])\ny=np.array([\n\t[1,0,0],\n\t[1,0,0],\n\t[1,0,0],\n\t[0,1,0],\n\t[0,1,0],\n\t[0,1,0],\n\t[0,0,1],\n\t[0,0,1],\n\t[0,0,1],\n])\nscore (all micro) 0.856060606060606\nscore (0, binds_BRD4) 1.0\nscore (1, binds_HSA) 1.0\nscore (2, binds_sEH) 1.0\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2747952,
          "postDate": "2024-04-12T07:00:57.613Z",
          "content": "<p>Thanks for sharing! Has any post-processing helped to achieve higher cv/lb score?</p>",
          "rawMarkdown": "Thanks for sharing! Has any post-processing helped to achieve higher cv/lb score?",
          "votes": 1,
          "replies": [
            {
              "id": 2748308,
              "postDate": "2024-04-12T11:26:12.063Z",
              "content": "<p>i leave this for future plan. improving accuracy for test building blocks not found in training is my top prority now.</p>\n<p>but you can google for differentiable micro average precision loss</p>",
              "rawMarkdown": "i leave this for future plan. improving accuracy for test building blocks not found in training is my top prority now.\n\nbut you can google for differentiable micro average precision loss",
              "votes": 1
            },
            {
              "id": 2748332,
              "postDate": "2024-04-12T11:43:26.317Z",
              "content": "<p>Your numbers for valid_noshare are very promising! It seems we can really generalize to new BBs. The question now is only how much.</p>",
              "rawMarkdown": "Your numbers for valid_noshare are very promising! It seems we can really generalize to new BBs. The question now is only how much."
            },
            {
              "id": 2748342,
              "postDate": "2024-04-12T11:51:21.550Z",
              "content": "<p>\"Your numbers for valid_noshare are very promising!\"</p>\n<p>any model that uses train size larger 68_152_325 (see subsample column) will not have validation metrics. (in this case, the training set include the validation set).</p>\n<p>i would say that current model for valid nonshare has precision half of that valid share.  <br>\nbut according to papers, SOTA can probably increase that to 70%</p>",
              "rawMarkdown": "\"Your numbers for valid_noshare are very promising!\"\n\nany model that uses train size larger 68\\_152\\_325 (see subsample column) will not have validation metrics. (in this case, the training set include the validation set).\n\ni would say that current model for valid nonshare has precision half of that valid share.  \nbut according to papers, SOTA can probably increase that to 70%\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2776107,
      "postDate": "2024-04-26T02:12:11.753Z",
      "content": "<p>i found my ideal algorithm<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0c62172bd70306926304d48f42f18fe6%2FSelection_047.png?generation=1714097530102992&amp;alt=media\"></p>\n<p><a href=\"https://github.com/luwei0917/TankBind\" target=\"_blank\">https://github.com/luwei0917/TankBind</a></p>\n<p>\"TankBind also support virtual screening. In our example here, for the WDR domain of LRRK2 protein, we can screen 10,000 drug candidates in 2 minutes (or 1M in around 3 hours) with a single GPU. Check out\"</p>",
      "rawMarkdown": "i found my ideal algorithm\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0c62172bd70306926304d48f42f18fe6%2FSelection_047.png?generation=1714097530102992&alt=media)\n\nhttps://github.com/luwei0917/TankBind\n\n\"TankBind also support virtual screening. In our example here, for the WDR domain of LRRK2 protein, we can screen 10,000 drug candidates in 2 minutes (or 1M in around 3 hours) with a single GPU. Check out\"",
      "votes": 2,
      "replies": [
        {
          "id": 2778802,
          "postDate": "2024-04-27T11:17:32.807Z",
          "content": "<p>Hi, if you trust that DTI prediction may be the optimal solution for this task, may be you can try some of the latest models, such as: </p>\n<ol>\n<li><a href=\"https://github.com/Zhang-Runze/PackDock\" target=\"_blank\">https://github.com/Zhang-Runze/PackDock</a></li>\n<li><a href=\"https://github.com/HBioquant/DiffBindFR\" target=\"_blank\">https://github.com/HBioquant/DiffBindFR</a></li>\n</ol>\n<p>The rationale behind this is that some early DTI prediction methods have been criticized in recent papers for their unreasonable docking poses and poor generalization ability, e. g. <a href=\"https://arxiv.org/abs/2308.05777\" target=\"_blank\">https://arxiv.org/abs/2308.05777</a>. </p>",
          "rawMarkdown": "Hi, if you trust that DTI prediction may be the optimal solution for this task, may be you can try some of the latest models, such as: \n1. https://github.com/Zhang-Runze/PackDock\n2. https://github.com/HBioquant/DiffBindFR\n\nThe rationale behind this is that some early DTI prediction methods have been criticized in recent papers for their unreasonable docking poses and poor generalization ability, e. g. https://arxiv.org/abs/2308.05777. ",
          "votes": 2,
          "replies": [
            {
              "id": 2778817,
              "postDate": "2024-04-27T11:31:10.350Z",
              "content": "<p>\" unreasonable docking poses and poor generalization ability\"<br>\ni agree with these.</p>\n<p>from competition point of view:</p>\n<ul>\n<li>for the non sharing blocks, most competitors are going to perform poorly.</li>\n<li>ML does not work if the train data significantly differs from these nonshare test.<br>\n(you may to refer to <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity</a>)</li>\n<li>if you can do better in these test subset (even just mere 20% to 40% of these), maybe you can have better private score. </li>\n</ul>\n<p>so my strategy is simple:</p>\n<ul>\n<li>QSAR for majority of the results</li>\n<li>self-supervsied, external data/block-based modeling to improve nonshare blocks</li>\n<li>DTI to improve/anaylse to results</li>\n<li>docking for a small \"selected\" subset</li>\n</ul>",
              "rawMarkdown": "\" unreasonable docking poses and poor generalization ability\"\ni agree with these.\n\nfrom competition point of view:\n- for the non sharing blocks, most competitors are going to perform poorly.\n- ML does not work if the train data significantly differs from these nonshare test.\n(you may to refer to https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity)\n- if you can do better in these test subset (even just mere 20% to 40% of these), maybe you can have better private score. \n\nso my strategy is simple:\n- QSAR for majority of the results\n- self-supervsied, external data/block-based modeling to improve nonshare blocks\n- DTI to improve/anaylse to results\n- docking for a small \"selected\" subset",
              "votes": 1
            },
            {
              "id": 2780192,
              "postDate": "2024-04-28T05:03:40.963Z",
              "content": "<blockquote>\n  <p>QSAR for majority of the results<br>\n  self-supervsied, external data/block-based modeling to improve nonshare blocks<br>\n  DTI to improve/anaylse to results<br>\n  docking for a small \"selected\" subset</p>\n</blockquote>\n<p>Great! Indeed, for real-world projects, we have indeed employed approachs like this. </p>",
              "rawMarkdown": ">QSAR for majority of the results\nself-supervsied, external data/block-based modeling to improve nonshare blocks\nDTI to improve/anaylse to results\ndocking for a small \"selected\" subset\n\n\nGreat! Indeed, for real-world projects, we have indeed employed approachs like this. "
            }
          ]
        }
      ]
    },
    {
      "id": 2774889,
      "postDate": "2024-04-25T11:14:52.020Z",
      "content": "<p>i wonder if something like this is possible?<br>\nwe have the blocks and the reaction results.(for both train and test)</p>\n<p>Contextual Molecule Representation Learning from Chemical Reaction Knowledge <br>\n<a href=\"https://arxiv.org/html/2402.13779v1\" target=\"_blank\">https://arxiv.org/html/2402.13779v1</a></p>\n<ul>\n<li>transformer based SSL method</li>\n<li>Molecular Representation Learning (MRL)</li>\n</ul>\n<p>\"REMO, a self-supervised learning framework that takes advantage of well-defined atom-combination rules in common chemistry. Specifically, REMO pre-trains graph/Transformer encoders on 1.7 million known chemical reactions in the literature. We propose two pre-training objectives: Masked Reaction Centre Reconstruction (MRCR) and Reaction Centre Identification (RCI). REMO offers a novel solution to MRL by exploiting the underlying shared patterns in chemical reactions as context for pre-training, which effectively infers meaningful representations of common chemistry knowledge\"</p>",
      "rawMarkdown": "i wonder if something like this is possible?\nwe have the blocks and the reaction results.(for both train and test)\n\nContextual Molecule Representation Learning from Chemical Reaction Knowledge \nhttps://arxiv.org/html/2402.13779v1\n- transformer based SSL method\n- Molecular Representation Learning (MRL)\n\n\"REMO, a self-supervised learning framework that takes advantage of well-defined atom-combination rules in common chemistry. Specifically, REMO pre-trains graph/Transformer encoders on 1.7 million known chemical reactions in the literature. We propose two pre-training objectives: Masked Reaction Centre Reconstruction (MRCR) and Reaction Centre Identification (RCI). REMO offers a novel solution to MRL by exploiting the underlying shared patterns in chemical reactions as context for pre-training, which effectively infers meaningful representations of common chemistry knowledge\"",
      "votes": 2
    },
    {
      "id": 2773938,
      "postDate": "2024-04-25T01:21:45.663Z",
      "content": "<p>Wherever you are, the competition will be very exciting. Inspired by you, I have also started a similar discussion in BirdCLEF2024, and will continue to update it.</p>",
      "rawMarkdown": "Wherever you are, the competition will be very exciting. Inspired by you, I have also started a similar discussion in BirdCLEF2024, and will continue to update it.",
      "votes": 2
    },
    {
      "id": 2771701,
      "postDate": "2024-04-24T11:23:44.463Z",
      "content": "<p>Stochastic Optimization of Areas Under Precision-Recall Curves with Provable Convergence<br>\n<a href=\"https://arxiv.org/abs/2104.08736\" target=\"_blank\">https://arxiv.org/abs/2104.08736</a></p>\n<p>example usage<br>\n<a href=\"https://docs.libauc.org/examples/auprc.html\" target=\"_blank\">https://docs.libauc.org/examples/auprc.html</a></p>\n<pre><code>model = ResNet18(=, =None, =1)\nmodel = model.cuda()\n\nloss_fn = APLoss(=len(trainSet), =margin, =gamma)\noptimizer = SOAP(model.parameters(), =lr, =, =weight_decay)\n</code></pre>\n<p>see also<br>\n<a href=\"https://github.com/divelab/MoleculeX\" target=\"_blank\">https://github.com/divelab/MoleculeX</a></p>\n<p>In addition, AdvProp is able to deal with tasks in which samples from different classes are highly imbalanced. In these cases, we employ advanced loss functions that optimize various areas under curves (AUC), such as areas under the receiver operating characteristic (AUROC) and the precision recall curve (AUPRC). AdvProp has been used to participate in the AI Cures open challenge for COVID-19 and is now ranked #1 in terms of both AUROC and AUPRC on the leaderboard. </p>",
      "rawMarkdown": "Stochastic Optimization of Areas Under Precision-Recall Curves with Provable Convergence\nhttps://arxiv.org/abs/2104.08736\n\nexample usage\nhttps://docs.libauc.org/examples/auprc.html\n\n```\nmodel = ResNet18(pretrained=False, last_activation=None, num_classes=1)\nmodel = model.cuda()\n\nloss_fn = APLoss(data_len=len(trainSet), margin=margin, gamma=gamma)\noptimizer = SOAP(model.parameters(), lr=lr, mode='adam', weight_decay=weight_decay)\n```\n\nsee also\nhttps://github.com/divelab/MoleculeX\n\nIn addition, AdvProp is able to deal with tasks in which samples from different classes are highly imbalanced. In these cases, we employ advanced loss functions that optimize various areas under curves (AUC), such as areas under the receiver operating characteristic (AUROC) and the precision recall curve (AUPRC). AdvProp has been used to participate in the AI Cures open challenge for COVID-19 and is now ranked #1 in terms of both AUROC and AUPRC on the leaderboard. ",
      "votes": 2,
      "replies": [
        {
          "id": 2850573,
          "postDate": "2024-06-02T08:18:31.437Z",
          "content": "<p>Interesting to see you find my friend’s work on advprop, I’m trying to optimize it for this dataset</p>",
          "rawMarkdown": "Interesting to see you find my friend’s work on advprop, I’m trying to optimize it for this dataset"
        }
      ]
    },
    {
      "id": 2775243,
      "postDate": "2024-04-25T15:04:45.750Z",
      "content": "<p>i suddenly have an idea.<br>\nfor the test smiles, run some tools or simulations (e.g. energy,  force field, conformers?) to get some measurements related to binding. learn model to predict such quality on test.</p>\n<p>can check correlation of measurements with target bind on train, tsNE, etc</p>",
      "rawMarkdown": "i suddenly have an idea.\nfor the test smiles, run some tools or simulations (e.g. energy,  force field, conformers?) to get some measurements related to binding. learn model to predict such quality on test.\n\ncan check correlation of measurements with target bind on train, tsNE, etc",
      "replies": [
        {
          "id": 2775778,
          "postDate": "2024-04-25T19:38:01.787Z",
          "content": "<p>do you mean to say \"for the <strong>train</strong> smiles…\" ? if not, could you explain a little more what you mean?</p>",
          "rawMarkdown": "do you mean to say \"for the **train** smiles...\" ? if not, could you explain a little more what you mean?",
          "replies": [
            {
              "id": 2775825,
              "postDate": "2024-04-25T20:27:55.740Z",
              "content": "<p>the test smiles</p>",
              "rawMarkdown": "the test smiles"
            }
          ]
        }
      ]
    },
    {
      "id": 2754267,
      "postDate": "2024-04-15T23:38:16.950Z",
      "content": "<p>what? joint protein-molecue fingerprint? interesting ….<br>\n<a href=\"https://academic.oup.com/bioinformatics/article/35/8/1334/5092926\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/35/8/1334/5092926</a></p>\n<p>Development of a protein–ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions</p>",
      "rawMarkdown": "what? joint protein-molecue fingerprint? interesting ....\nhttps://academic.oup.com/bioinformatics/article/35/8/1334/5092926\n\nDevelopment of a protein–ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions",
      "votes": 2
    },
    {
      "id": 2748402,
      "postDate": "2024-04-12T12:26:35.590Z",
      "content": "<p>Training on 98M samples with 2048 bits each gives almost 200GB of training data. Do I understand that correctly? Just wondering what hardware you use for it.</p>",
      "rawMarkdown": "Training on 98M samples with 2048 bits each gives almost 200GB of training data. Do I understand that correctly? Just wondering what hardware you use for it.",
      "votes": 2,
      "replies": [
        {
          "id": 2748476,
          "postDate": "2024-04-12T13:18:06.317Z",
          "content": "<p>use one byte to represent 8 bit.<br>\nso 2048 bit=256 byte.</p>\n<p>see np.packbits.</p>",
          "rawMarkdown": "use one byte to represent 8 bit.\nso 2048 bit=256 byte.\n\nsee np.packbits.",
          "votes": 2,
          "replies": [
            {
              "id": 2748546,
              "postDate": "2024-04-12T13:38:58.040Z",
              "content": "<p>My bad. Thanks!</p>",
              "rawMarkdown": "My bad. Thanks!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2753694,
      "postDate": "2024-04-15T16:05:54.067Z",
      "content": "<p>With you around this comp' would be extra fun</p>",
      "rawMarkdown": "With you around this comp' would be extra fun",
      "votes": -5
    },
    {
      "id": 2875171,
      "postDate": "2024-06-16T23:39:34.410Z",
      "content": "<p>Curious have you tried the 77M version of chemberta or molformer?</p>",
      "rawMarkdown": "Curious have you tried the 77M version of chemberta or molformer?"
    },
    {
      "id": 2816398,
      "postDate": "2024-05-16T09:44:15.263Z",
      "content": "<p>BRD4 dataset <br>\n<a href=\"https://github.com/shiwentao00/Graphsite-classifier/tree/master\" target=\"_blank\">https://github.com/shiwentao00/Graphsite-classifier/tree/master</a><br>\nGraphSite: Ligand Binding Site Classification with Deep Graph Learning<br>\n<a href=\"https://www.mdpi.com/2218-273X/12/8/1053\" target=\"_blank\">https://www.mdpi.com/2218-273X/12/8/1053</a></p>",
      "rawMarkdown": "BRD4 dataset \nhttps://github.com/shiwentao00/Graphsite-classifier/tree/master\nGraphSite: Ligand Binding Site Classification with Deep Graph Learning\nhttps://www.mdpi.com/2218-273X/12/8/1053"
    },
    {
      "id": 2804674,
      "postDate": "2024-05-10T06:44:37.237Z",
      "content": "<p>some probing tutorial:<br>\nfor more information refer to <a href=\"https://stats.stackexchange.com/questions/394494/calculating-sklearns-average-precision-by-hand\" target=\"_blank\">https://stats.stackexchange.com/questions/394494/calculating-sklearns-average-precision-by-hand</a></p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/how-to-probe\" target=\"_blank\">https://www.kaggle.com/code/hengck23/how-to-probe</a></p>",
      "rawMarkdown": "some probing tutorial:\nfor more information refer to https://stats.stackexchange.com/questions/394494/calculating-sklearns-average-precision-by-hand\n\nhttps://www.kaggle.com/code/hengck23/how-to-probe"
    },
    {
      "id": 2794938,
      "postDate": "2024-05-05T15:26:51.117Z",
      "content": "<p>Really nice work on the FP+SMILES model <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>",
      "rawMarkdown": "Really nice work on the FP+SMILES model @hengck23 !"
    },
    {
      "id": 2794282,
      "postDate": "2024-05-05T07:41:50.030Z",
      "content": "<p>old results for archive:</p>\n<p>WARNING: </p>\n<ul>\n<li>results may not be optimized/tuned (since i am still experimenting … <br>\nit will be updated throughout the competition till last 2 weeks)  </li>\n<li><strong>the split are not correct, valid nonshare actually used share blocks by mistake !!! (see discussion below)</strong></li>\n</ul>\n<hr>\n<p>i think i can conclude transformer beats xgboost+fingerprint!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a7872cf8060014be810dd706bab1133%2FSelection_036.png?generation=1713929270479439&amp;alt=media\"></p>\n<p>new results with gpu training<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca8428acaa3ffa2501850471183ca09b%2FSelection_026.png?generation=1713066709109703&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdaa7c3665fe31df78e61ffac106f4d7e%2FSelection_024.png?generation=1712911082033973&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2e854b52473673388a32ab6c76d93b6%2FSelection_017.png?generation=1712818198030642&amp;alt=media\"></p>",
      "rawMarkdown": "old results for archive:\n\n\nWARNING: \n- results may not be optimized/tuned (since i am still experimenting ... \n   it will be updated throughout the competition till last 2 weeks)  \n- **the split are not correct, valid nonshare actually used share blocks by mistake !!! (see discussion below)**\n\n---\ni think i can conclude transformer beats xgboost+fingerprint!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a7872cf8060014be810dd706bab1133%2FSelection_036.png?generation=1713929270479439&alt=media)\n\nnew results with gpu training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca8428acaa3ffa2501850471183ca09b%2FSelection_026.png?generation=1713066709109703&alt=media)\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdaa7c3665fe31df78e61ffac106f4d7e%2FSelection_024.png?generation=1712911082033973&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2e854b52473673388a32ab6c76d93b6%2FSelection_017.png?generation=1712818198030642&alt=media)"
    },
    {
      "id": 2776056,
      "postDate": "2024-04-26T00:56:49.790Z",
      "content": "<p><a href=\"https://www.thesgc.org/news/structural-genomics-consortium-and-x-chem-enter-collaboration-unlock-human-proteome-and-promote\" target=\"_blank\">https://www.thesgc.org/news/structural-genomics-consortium-and-x-chem-enter-collaboration-unlock-human-proteome-and-promote</a><br>\nThese datasets, curated in an ML-ready format, will be posted to a public portal to be used for model building. …  but X-Chem and the SGC expect the DEL-ML data portal will be ready for public access starting in early 2024.</p>",
      "rawMarkdown": "https://www.thesgc.org/news/structural-genomics-consortium-and-x-chem-enter-collaboration-unlock-human-proteome-and-promote\nThese datasets, curated in an ML-ready format, will be posted to a public portal to be used for model building. ...  but X-Chem and the SGC expect the DEL-ML data portal will be ready for public access starting in early 2024."
    },
    {
      "id": 2775830,
      "postDate": "2024-04-25T20:36:06.070Z",
      "content": "<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F188a7ef0c66e995e004ceb9a9dd276da%2FSelection_046.png?generation=1714077718909224&amp;alt=media\"><p></p>\n<p>bb considerations</p>",
      "rawMarkdown": "![https://www.youtube.com/watch?v=AuNXa0nNkAc](https://www.youtube.com/watch?v=AuNXa0nNkAc)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F188a7ef0c66e995e004ceb9a9dd276da%2FSelection_046.png?generation=1714077718909224&alt=media)\n\nbb considerations"
    },
    {
      "id": 2775703,
      "postDate": "2024-04-25T18:44:39.913Z",
      "content": "<p>there is one trick.<br>\none can \"google\" for compound that is close to the test molecule …. write an LLM agent for that</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fda40f675564d703d03c4f25b0cffa68b%2FSelection_044.png?generation=1714070663174093&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7cfa93726f6544be9304b8fb5f6929a8%2FSelection_043.png?generation=1714070675606349&amp;alt=media\"></p>",
      "rawMarkdown": "there is one trick.\none can \"google\" for compound that is close to the test molecule .... write an LLM agent for that\n \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fda40f675564d703d03c4f25b0cffa68b%2FSelection_044.png?generation=1714070663174093&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7cfa93726f6544be9304b8fb5f6929a8%2FSelection_043.png?generation=1714070675606349&alt=media)\n"
    },
    {
      "id": 2775219,
      "postDate": "2024-04-25T14:48:35.107Z",
      "content": "<p>yet another external data<br>\n<a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/60c741e2567dfedeb7ec3e52\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/60c741e2567dfedeb7ec3e52</a><br>\n(see Supplementary materials)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F688846386de6863b1ddc4e84f7d3598a%2FSelection_041.png?generation=1714056462488116&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F035dffab1de71e8666b952f275e21f1f%2FSelection_042.png?generation=1714056474616688&amp;alt=media\"></p>",
      "rawMarkdown": "yet another external data\nhttps://chemrxiv.org/engage/chemrxiv/article-details/60c741e2567dfedeb7ec3e52\n(see Supplementary materials)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F688846386de6863b1ddc4e84f7d3598a%2FSelection_041.png?generation=1714056462488116&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F035dffab1de71e8666b952f275e21f1f%2FSelection_042.png?generation=1714056474616688&alt=media)"
    },
    {
      "id": 2774925,
      "postDate": "2024-04-25T11:47:07.613Z",
      "content": "<p>i have completed </p>\n<ol>\n<li>fingerprint + xgboost</li>\n<li>SMILE stsing transmformer</li>\n</ol>\n<p>Now i am trying (in process of training …) 2d/3d GNN (i.e. input = molecular graph).<br>\nIf anyone has results in GNN, or good papers/github repo to recommend, or tips in handling huge data of grpah, it is most welcomed!</p>",
      "rawMarkdown": "i have completed \n1. fingerprint + xgboost\n2. SMILE stsing transmformer\n\nNow i am trying (in process of training ...) 2d/3d GNN (i.e. input = molecular graph).\nIf anyone has results in GNN, or good papers/github repo to recommend, or tips in handling huge data of grpah, it is most welcomed!",
      "replies": [
        {
          "id": 2775040,
          "postDate": "2024-04-25T12:57:50.753Z",
          "content": "<p>The edges data is sparse. So keep it as sparse matrix and turn to dense just before feeding to the model. Max number of nodes is around 70 (much less than smile length) so total size of data is manageable. Less than *2 of smile data size. I have not tried gnn, but tried nodes transformer+feeding edges data to attention (similar to top solutions in Ribonanza) and could not beat smile transformer yet- smile transformer ~0.67 val vs. ~0.605 val for nodes transformer+injected edges. I'm waiting to see your results :)</p>",
          "rawMarkdown": "The edges data is sparse. So keep it as sparse matrix and turn to dense just before feeding to the model. Max number of nodes is around 70 (much less than smile length) so total size of data is manageable. Less than *2 of smile data size. I have not tried gnn, but tried nodes transformer+feeding edges data to attention (similar to top solutions in Ribonanza) and could not beat smile transformer yet- smile transformer ~0.67 val vs. ~0.605 val for nodes transformer+injected edges. I'm waiting to see your results :)"
        },
        {
          "id": 2775078,
          "postDate": "2024-04-25T13:17:25.933Z",
          "content": "<p><a href=\"https://github.com/wengong-jin/hgraph2graph/tree/master\" target=\"_blank\">https://github.com/wengong-jin/hgraph2graph/tree/master</a><br>\n(paper in github)<br>\nI was looking at this earlier as an encoder but I'm trying molformer first since it can be used pretrained.<br>\nI have however generated the hgraph vocab for train + test with a C/Methane Dy substitute.</p>\n<p>Edit: <a href=\"https://www.kaggle.com/datasets/sroger/belka-hgraph-vocab/\" target=\"_blank\">https://www.kaggle.com/datasets/sroger/belka-hgraph-vocab/</a></p>",
          "rawMarkdown": "https://github.com/wengong-jin/hgraph2graph/tree/master\n(paper in github)\nI was looking at this earlier as an encoder but I'm trying molformer first since it can be used pretrained.\nI have however generated the hgraph vocab for train + test with a C/Methane Dy substitute.\n\nEdit: https://www.kaggle.com/datasets/sroger/belka-hgraph-vocab/",
          "replies": [
            {
              "id": 2775189,
              "postDate": "2024-04-25T14:25:50.677Z",
              "content": "<p>thanks all.</p>\n<hr>\n<p>\"~0.67 val vs. ~0.605 val \"<br>\ni have a feeling that there may not be much help from GNN … but let's wait for a few days.</p>\n<p>I am wondering how others have achieved LB better than 0.600. Is it they are have better nonshare blocks score or better share block score (which means local cv should reach 0.700)?</p>\n<p>or are they \"dealing with blocks\"?</p>\n<p>fingerprint is much like 2d graph, so i am hoping 3d information would work.<br>\nhowever, generating 3d information itself is a problem (not accurate).</p>\n<p>if it doesn't my next bet is SSL(self supervised learning) on test molecules.</p>\n<hr>\n<p>\" train + test with a C/Methane Dy substitute.\"<br>\nyou can show examples of pair (before,after conversion), i can do experiments at my side.<br>\n(e.g. molformer or chembeta)<br>\nwould be better if you some code.</p>",
              "rawMarkdown": "thanks all.\n\n----\n\"~0.67 val vs. ~0.605 val \"\ni have a feeling that there may not be much help from GNN ... but let's wait for a few days.\n\nI am wondering how others have achieved LB better than 0.600. Is it they are have better nonshare blocks score or better share block score (which means local cv should reach 0.700)?\n\nor are they \"dealing with blocks\"?\n\nfingerprint is much like 2d graph, so i am hoping 3d information would work.\nhowever, generating 3d information itself is a problem (not accurate).\n\nif it doesn't my next bet is SSL(self supervised learning) on test molecules.\n\n---\n\n\" train + test with a C/Methane Dy substitute.\"\nyou can show examples of pair (before,after conversion), i can do experiments at my side.\n(e.g. molformer or chembeta)\nwould be better if you some code."
            },
            {
              "id": 2775192,
              "postDate": "2024-04-25T14:32:55.160Z",
              "content": "<p><a href=\"https://arxiv.org/pdf/2310.13769.pdf\" target=\"_blank\">https://arxiv.org/pdf/2310.13769.pdf</a><br>\nCompositional Deep Probabilistic Models of DNA-Encoded Libraries</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e085abf23adc893f9aa0d416679c260%2FSelection_040.png?generation=1714055546823055&amp;alt=media\"></p>",
              "rawMarkdown": "https://arxiv.org/pdf/2310.13769.pdf\nCompositional Deep Probabilistic Models of DNA-Encoded Libraries\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3e085abf23adc893f9aa0d416679c260%2FSelection_040.png?generation=1714055546823055&alt=media)"
            },
            {
              "id": 2775224,
              "postDate": "2024-04-25T14:49:43.023Z",
              "content": "<p>\"I am wondering how others have achieved LB better than 0.600\"<br>\nI have two guesses currently:</p>\n<ol>\n<li>ensembling- have you tried ensembling several transformer seeds and/or XGB?</li>\n<li>Smile augmentation- smiles are not unique. While we probably do have enough negative smiles, I think positive smiles augmentation would help.</li>\n</ol>",
              "rawMarkdown": "\"I am wondering how others have achieved LB better than 0.600\"\nI have two guesses currently:\n1. ensembling- have you tried ensembling several transformer seeds and/or XGB?\n2. Smile augmentation- smiles are not unique. While we probably do have enough negative smiles, I think positive smiles augmentation would help.",
              "votes": 1
            },
            {
              "id": 2775294,
              "postDate": "2024-04-25T15:39:05.960Z",
              "content": "<p>For myself, I'm getting right about .700 on share block CV, and haven't started doing any enhancements or non-share block work yet. </p>",
              "rawMarkdown": "For myself, I'm getting right about .700 on share block CV, and haven't started doing any enhancements or non-share block work yet. ",
              "votes": 1
            },
            {
              "id": 2775440,
              "postDate": "2024-04-25T16:58:22.993Z",
              "content": "<p>\"For myself, I'm getting right about .700 \"</p>\n<p>that is very good!!</p>",
              "rawMarkdown": "\"For myself, I'm getting right about .700 \"\n\nthat is very good!!",
              "votes": 1
            },
            {
              "id": 2775571,
              "postDate": "2024-04-25T17:37:59.577Z",
              "content": "<p>Thanks!</p>\n<p></p>\n<p>Edit: just a test dataset pipeline bug, fixed now and correlating fine. My best trusted baseline was 0.695 CV, 0.597 LB, and that went to 0.700 CV, 0.600 LB today, so that's pretty close to exact correlation so far.</p>",
              "rawMarkdown": "Thanks!\n\n~~...But it's not correlating with LB 🤦🏻‍♂️~~\n\nEdit: just a test dataset pipeline bug, fixed now and correlating fine. My best trusted baseline was 0.695 CV, 0.597 LB, and that went to 0.700 CV, 0.600 LB today, so that's pretty close to exact correlation so far."
            },
            {
              "id": 2775674,
              "postDate": "2024-04-25T18:35:21.840Z",
              "content": "<p>my estimate is quite correct:<br>\nCV = LB + m</p>\n<p>m is the performance of nonshare block, which is close to zero (or test distribution of (ratio of nonshare )*(zero submission)). zero submission is assume model cannot detect nonshare block at all. </p>",
              "rawMarkdown": "my estimate is quite correct:\nCV = LB + m\n\nm is the performance of nonshare block, which is close to zero (or test distribution of (ratio of nonshare )*(zero submission)). zero submission is assume model cannot detect nonshare block at all. "
            },
            {
              "id": 2775683,
              "postDate": "2024-04-25T18:38:05.603Z",
              "content": "<p>wanted to share some info about potential of the transformers, - </p>\n<p>data: all 1's (about 1.4kk), randomly sample 20kk of zero samples<br>\ntransformer model: 0.589 LB</p>",
              "rawMarkdown": "wanted to share some info about potential of the transformers, - \n\ndata: all 1's (about 1.4kk), randomly sample 20kk of zero samples\ntransformer model: 0.589 LB",
              "votes": 5
            },
            {
              "id": 2776086,
              "postDate": "2024-04-26T01:32:48.620Z",
              "content": "<p>thanks a lot!</p>",
              "rawMarkdown": "thanks a lot!"
            },
            {
              "id": 2776451,
              "postDate": "2024-04-26T07:19:00.877Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> for sharing. Is it basic transformer or specific transformer for chemistry ? Will apreciate if you can point to a paper or something.</p>",
              "rawMarkdown": "Thanks @martynoveduard for sharing. Is it basic transformer or specific transformer for chemistry ? Will apreciate if you can point to a paper or something.",
              "votes": 1
            },
            {
              "id": 2776489,
              "postDate": "2024-04-26T07:48:53.673Z",
              "content": "<p>I just trained mine from scratch, haven't tried existing pre-trained models</p>",
              "rawMarkdown": "I just trained mine from scratch, haven't tried existing pre-trained models",
              "votes": 1
            },
            {
              "id": 2776523,
              "postDate": "2024-04-26T08:08:28.753Z",
              "content": "<p>Thought I'd add to this now that I've made a sub.<br>\nNaive folds/sampling + frozen molformer encoder and some other layers: LB:0.539 CV:0.607</p>",
              "rawMarkdown": "Thought I'd add to this now that I've made a sub.\nNaive folds/sampling + frozen molformer encoder and some other layers: LB:0.539 CV:0.607"
            },
            {
              "id": 2776596,
              "postDate": "2024-04-26T08:43:36.453Z",
              "content": "<p>if anyone want to use pretrained model (frozen), i think one has to replace the [Dy]</p>",
              "rawMarkdown": "if anyone want to use pretrained model (frozen), i think one has to replace the [Dy]"
            },
            {
              "id": 2776605,
              "postDate": "2024-04-26T08:48:14.073Z",
              "content": "<pre><code>from rdkit import Chem\nfrom rdkit.Chem import AllChem\n\nm_Dy, m_C = Chem., Chem.\n\ndef proc -&gt; str:\n    m = Chem.\n    m = AllChem.\n    Chem.\n    return Chem.\n</code></pre>\n<p>This is the relevant snippet in my dataset, credits to <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a></p>",
              "rawMarkdown": "```\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\n\nm_Dy, m_C = Chem.MolFromSmiles(\"[Dy]\"), Chem.MolFromSmiles(\"C\")\n\ndef proc_smile(smile: str) -> str:\n    m = Chem.MolFromSmiles(smile)\n    m = AllChem.ReplaceSubstructs(m, m_Dy, m_C)[0]\n    Chem.SanitizeMol(m)\n    return Chem.MolToSmiles(m)\n```\n\nThis is the relevant snippet in my dataset, credits to @chemdatafarmer",
              "votes": 1
            },
            {
              "id": 2776891,
              "postDate": "2024-04-26T11:45:22.617Z",
              "content": "<p>\"Thought I'd add to this now that I've made a sub.\"<br>\nmake separate submission for share and nonshare blocks. then you can see if pretrain helps in nonsharing blocks</p>",
              "rawMarkdown": "\"Thought I'd add to this now that I've made a sub.\"\nmake separate submission for share and nonshare blocks. then you can see if pretrain helps in nonsharing blocks"
            },
            {
              "id": 2778251,
              "postDate": "2024-04-27T04:02:34.360Z",
              "content": "<p>I'll look into that, haven't done anything with the blocks so far.</p>",
              "rawMarkdown": "I'll look into that, haven't done anything with the blocks so far."
            },
            {
              "id": 2778409,
              "postDate": "2024-04-27T06:37:01.243Z",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> <br>\n\"I just trained mine from scratch, haven't tried existing pre-trained models\"<br>\ndo you use your own tokenizer?</p>",
              "rawMarkdown": "@martynoveduard \n\"I just trained mine from scratch, haven't tried existing pre-trained models\"\ndo you use your own tokenizer?"
            }
          ]
        }
      ]
    },
    {
      "id": 2774264,
      "postDate": "2024-04-25T05:42:14.977Z",
      "content": "<p>this is what i done when i do not know when to stop training. just average the just few checkpoints<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19b0eb03e19d34bae44873186d47055c%2FSelection_039.png?generation=1714023733193529&amp;alt=media\"></p>",
      "rawMarkdown": "this is what i done when i do not know when to stop training. just average the just few checkpoints\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19b0eb03e19d34bae44873186d47055c%2FSelection_039.png?generation=1714023733193529&alt=media)"
    },
    {
      "id": 2773983,
      "postDate": "2024-04-25T02:30:26.197Z",
      "content": "<p>cascade model, active learning and submission</p>\n<p>our final model is a cascade of models, each model has high recall of say 99%, but moderate precision say 50%.<br>\nassume each model has recall =0.99, FPrate =0.10, then a cascade of of 5 such model will be<br>\nrecall=pow(0.99,5)=still high,  FPrate =pow(0.1,5)=very low = 1e-6</p>\n<p>how it work:</p>\n<pre><code>.  data \n.   discarded  model0()&lt; threshold0,     model\n.   discarded  model1()&lt; threshold1,     model\n.   discarded  model2()&lt; threshold2,     model\n....\n few data remains\n(here we can use complicated methods like docking, MD simulation, complex fingerprint, large deep net, etc)\n</code></pre>\n<p>at submission</p>\n<p>````</p>\n<ol>\n<li>submit all zero rank (e.g. lb score =s0)</li>\n<li>for rejected sample, submit as rank=0, for remaining sample rank=1 <br>\n(if new lb score s1==s0, then you have 100 recall in public test)</li>\n<li>likewise, new accepted sample will have rank2 (i.e. submission has only 3 values 0,1,2)<br>\n3, etc ….<br>\n….</li>\n</ol>\n<p>the trick is that try to keeps0=s1=s2=s3 … then you have zero miss and left with very little data for next model</p>\n<p>```</p>\n<p>reference: <br>\ncascade adaboost for face detection. i think xgboost can be modified to do this.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb2302b19ce617b9d488c9a303ad1aeb%2FSelection_038.png?generation=1714012410058499&amp;alt=media\"></p>",
      "rawMarkdown": "cascade model, active learning and submission\n\nour final model is a cascade of models, each model has high recall of say 99%, but moderate precision say 50%.\nassume each model has recall =0.99, FPrate =0.10, then a cascade of of 5 such model will be\nrecall=pow(0.99,5)=still high,  FPrate =pow(0.1,5)=very low = 1e-6\n\nhow it work:\n```\n0. all data x\n1. x is discarded if model0(x)< threshold0, else go to next model\n2. x is discarded if model1(x)< threshold1, else go to next model\n3. x is discarded if model2(x)< threshold2, else go to next model\n....\nonly few data remains\n(here we can use complicated methods like docking, MD simulation, complex fingerprint, large deep net, etc)\n\n```\n\nat submission\n\n````\n0. submit all zero rank (e.g. lb score =s0)\n1. for rejected sample, submit as rank=0, for remaining sample rank=1 \n(if new lb score s1==s0, then you have 100 recall in public test)\n2. likewise, new accepted sample will have rank2 (i.e. submission has only 3 values 0,1,2)\n3, etc ....\n....\n\nthe trick is that try to keeps0=s1=s2=s3 ... then you have zero miss and left with very little data for next model\n\n```\n\nreference: \ncascade adaboost for face detection. i think xgboost can be modified to do this.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb2302b19ce617b9d488c9a303ad1aeb%2FSelection_038.png?generation=1714012410058499&alt=media)\n\n"
    },
    {
      "id": 2756352,
      "postDate": "2024-04-17T02:19:55.753Z",
      "content": "<p>diffdock: try nvidia online demo<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18ff3e4e995521d5794c1f5080d1cc8b%2FSelection_027.png?generation=1713320365421742&amp;alt=media\"></p>\n<p><a href=\"https://developer.nvidia.com/blog/new-models-molmim-and-diffdock-power-molecule-generation-and-molecular-docking-in-bionemo/\" target=\"_blank\">https://developer.nvidia.com/blog/new-models-molmim-and-diffdock-power-molecule-generation-and-molecular-docking-in-bionemo/</a></p>\n<p>maybe can apply for cloud beta services for benchmarking? if not, you can setup on your own gpu.</p>",
      "rawMarkdown": "diffdock: try nvidia online demo\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18ff3e4e995521d5794c1f5080d1cc8b%2FSelection_027.png?generation=1713320365421742&alt=media)\n\nhttps://developer.nvidia.com/blog/new-models-molmim-and-diffdock-power-molecule-generation-and-molecular-docking-in-bionemo/\n\nmaybe can apply for cloud beta services for benchmarking? if not, you can setup on your own gpu."
    },
    {
      "id": 2753345,
      "postDate": "2024-04-15T12:47:53.483Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for giving us a great insight on this competition! I have a small question though. What strategy did you use to sub-sample the dataset into 25m?</p>",
      "rawMarkdown": "@hengck23 Thank you for giving us a great insight on this competition! I have a small question though. What strategy did you use to sub-sample the dataset into 25m?"
    },
    {
      "id": 2748860,
      "postDate": "2024-04-12T17:30:33.590Z",
      "content": "<p>Did you split by maximizing nonshare groups via clustering? Because no way a naive splitting would give 13% nonshare. When I split each BB group to 5F it results in something on the order of 1% nonshare (~1M molecules out of ~100M). And only ~50% train. When I do 10F it drops to ~0.1% nonshare. I fyou did clustering, I suggest you check #positive in each group because it may not be random.</p>",
      "rawMarkdown": "Did you split by maximizing nonshare groups via clustering? Because no way a naive splitting would give 13% nonshare. When I split each BB group to 5F it results in something on the order of 1% nonshare (~1M molecules out of ~100M). And only ~50% train. When I do 10F it drops to ~0.1% nonshare. I fyou did clustering, I suggest you check #positive in each group because it may not be random.",
      "replies": [
        {
          "id": 2748863,
          "postDate": "2024-04-12T17:36:39.897Z",
          "content": "<p>check the split in my public dataset <a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a>.</p>\n<p>you first just select e.g. 5% nonshare building blocks for each bb1,bb2bb3. this will give you about 15% \"nonshare SMILE\", i.e. \"valid nonshare\". Then for the remaining \"share SMILE\", random split to \"train share\" and \"valid share\".</p>\n<hr>\n<p>a better way is to select nonshare train building blocks that is similar to nonshare test building blocks.<br>\n(e.g. Tanimoto Similarity)</p>",
          "rawMarkdown": "check the split in my public dataset https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset.\n\nyou first just select e.g. 5% nonshare building blocks for each bb1,bb2bb3. this will give you about 15% \"nonshare SMILE\", i.e. \"valid nonshare\". Then for the remaining \"share SMILE\", random split to \"train share\" and \"valid share\".\n\n---\n\na better way is to select nonshare train building blocks that is similar to nonshare test building blocks.\n(e.g. Tanimoto Similarity)",
          "votes": 2,
          "replies": [
            {
              "id": 2748905,
              "postDate": "2024-04-12T18:05:55.867Z",
              "content": "<p>Ahhh I see. I thought that nonshare='none of the BBs is shared', but it is actually 'at least one of the BBs is not shared' which is a bit different haha.<br>\nI wonder what the scores are for truly nonshared. This would also probably explain at least partly the valid/test gap.</p>",
              "rawMarkdown": "Ahhh I see. I thought that nonshare='none of the BBs is shared', but it is actually 'at least one of the BBs is not shared' which is a bit different haha.\nI wonder what the scores are for truly nonshared. This would also probably explain at least partly the valid/test gap.",
              "votes": 1
            },
            {
              "id": 2748912,
              "postDate": "2024-04-12T18:08:17.393Z",
              "content": "<p>none of the bb is shared is correct</p>",
              "rawMarkdown": "none of the bb is shared is correct\n"
            },
            {
              "id": 2748940,
              "postDate": "2024-04-12T18:19:35.737Z",
              "content": "<p>By your split? No, since you choose several non-share for each BBs, resulting in molecules with e.g. nonshare for BB1 but shared for BB2. Which are inside the nonshared split. So not truly 'none is shared'. Truly nonshared split is much smaller.</p>",
              "rawMarkdown": "By your split? No, since you choose several non-share for each BBs, resulting in molecules with e.g. nonshare for BB1 but shared for BB2. Which are inside the nonshared split. So not truly 'none is shared'. Truly nonshared split is much smaller.",
              "votes": 1
            },
            {
              "id": 2748957,
              "postDate": "2024-04-12T18:31:41.197Z",
              "content": "<p>it is my mistake. i thought bb1,bb2,bb3 are non overlapping.</p>",
              "rawMarkdown": "it is my mistake. i thought bb1,bb2,bb3 are non overlapping."
            },
            {
              "id": 2748986,
              "postDate": "2024-04-12T18:50:53.827Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2748993,
              "postDate": "2024-04-12T18:57:42.090Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2749035,
              "postDate": "2024-04-12T19:44:25.307Z",
              "content": "<p>If you want to do by similarity, you could do a sphere exclusion clustering in rdkit and hold certain clusters out. See here for one way to implement a sphere exclusion clustering: <a href=\"https://rdkit.blogspot.com/2020/11/sphere-exclusion-clustering-with-rdkit.html\" target=\"_blank\">https://rdkit.blogspot.com/2020/11/sphere-exclusion-clustering-with-rdkit.html</a>.</p>",
              "rawMarkdown": "If you want to do by similarity, you could do a sphere exclusion clustering in rdkit and hold certain clusters out. See here for one way to implement a sphere exclusion clustering: https://rdkit.blogspot.com/2020/11/sphere-exclusion-clustering-with-rdkit.html.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2747606,
      "postDate": "2024-04-12T02:05:21.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> You may find it interesting that the LB score is very similar with LGBM instead of XGBoost, shows similar improvement when scaling up to more estimators, and trains much faster. Could be helpful for experiments. I don't have exact time measurements, unfortunately.</p>",
      "rawMarkdown": "@hengck23 You may find it interesting that the LB score is very similar with LGBM instead of XGBoost, shows similar improvement when scaling up to more estimators, and trains much faster. Could be helpful for experiments. I don't have exact time measurements, unfortunately.",
      "replies": [
        {
          "id": 2748482,
          "postDate": "2024-04-12T13:20:51.477Z",
          "content": "<p>i haven't try xgboost gpu yet (tree=hist(gpu)). maybe that is fast enough</p>",
          "rawMarkdown": "i haven't try xgboost gpu yet (tree=hist(gpu)). maybe that is fast enough"
        }
      ]
    },
    {
      "id": 2747954,
      "postDate": "2024-04-12T07:01:42.257Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2757671,
      "postDate": "2024-04-17T16:47:19.327Z",
      "content": "<p>Thanks for sharing new ideas!!</p>",
      "rawMarkdown": "Thanks for sharing new ideas!!"
    },
    {
      "id": 2748731,
      "postDate": "2024-04-12T15:08:47.207Z",
      "content": "<p>thanks for sharing idea!</p>",
      "rawMarkdown": "thanks for sharing idea!"
    }
  ],
  "comments": [
    {
      "id": 2748792,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-12T16:16:15.880000",
      "content": "<p>dataset for the above discussion have been shared!!!!<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p>In summary:</p>\n<ul>\n<li>train.reduced.parquet : 98_415_610 training SMILES and their information</li>\n<li>train.ecfp4.packed.npz : Features extracted using rdkit  <ul>\n<li>AllChem.GetMorganFingerprintAsBitVect(mol, 2, 2048)                                                    </li>\n<li>repack with np.packbits() to give 98_415_610 x 256 feature matrix</li></ul></li>\n<li>train.bind.npz : 98_415_610 x 3 target matrix</li>\n<li>test.reduced.parquet/ test.ecfp4.packed.npz : similarly processed for the test SMILES</li>\n<li>all_buildingblock.csv: building blocks id used in train.reduced.parquet/test.reduced.parquet</li>\n<li>fold0.parquet: train_share,valid_share,valid_nonshare splits for the experiments in the discussion</li>\n</ul>",
      "votes": 13,
      "replies": [
        {
          "id": 2748795,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-12T16:17:43.113000",
          "content": "<p>i think with proper xgboost tuning and ensemble of good data splits/folds, you can get LB of about 0.575. If you discover good hyper-parameters, please share here! Thanks!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2746508,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2024-04-11T09:15:43.460000",
      "content": "<p>Thanks for sharing ! </p>\n<p>My colleagues and me sometimes organize introductory webinars on Kaggle competitions <br>\nSome video records here: <a href=\"https://www.youtube.com/watch?v=aqUOz3nFYm4&amp;list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw\" target=\"_blank\">https://www.youtube.com/watch?v=aqUOz3nFYm4&amp;list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw</a></p>\n<p>Is there any chance that you can make an introduction to that competition ? <br>\nJust intro survey - data, metrics, public solutions, whatsever.<br>\nIt is not a private sharing - we make  anounces here on Kaggle forum - everybody can join via zoom<br>\nand ask questions/comments/whatsever. Video will be later on Youtube.<br>\nTiming about 40 minutes </p>\n<p>Sometimes organizers join the zoom and make their comments also:<br>\n<a href=\"https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170\" target=\"_blank\">https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170</a></p>\n<p>In general it seems quite helpful for community. <br>\nAnother examples here:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825</a><br>\n<a href=\"https://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975\" target=\"_blank\">https://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975</a></p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 2765385,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-21T05:52:32.497000",
      "content": "<p>transformer (pure SMILES string) results out !!!!<br>\n<a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">https://huggingface.co/ibm/MoLFormer-XL-both-10pct</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa4b79498cbc6dd566f0dcb15a31d6272%2FSelection_032.png?generation=1713678750507300&amp;alt=media\"></p>",
      "votes": 7,
      "replies": [
        {
          "id": 2804090,
          "author_name": "Scott K",
          "author_url": "",
          "post_date": "2024-05-09T19:31:07.377000",
          "content": "<p>Care to share how much training data this network saw? Or how many FF layers after the transformer? I used 3 layers post-transformer and trained on about half the training data to get a score of ~0.5.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2776878,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-26T11:42:08.483000",
      "content": "<p>one of my idea<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F642a737a82d59e4b1ed1fba2b312cc8e%2FSelection_051.png?generation=1714131726333596&amp;alt=media\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2776899,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-04-26T11:49:45.487000",
          "content": "<p>A similar idea: if you are interested in multitask learning to help embed the test and train in a similar embedding space, you might want to consider predicting some calculated properties like cLogP (for many enzymes, binding loosely tracks with cLogP) or TPSA (which is a reference to the polar surface area of the molecule). </p>\n<p>RDKit has some tools to calculate descriptors like these, as a start.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2776922,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-26T12:02:15.167000",
              "content": "<p>ref<br>\n<a href=\"https://dl.acm.org/doi/10.1145/3233547.3233548\" target=\"_blank\">https://dl.acm.org/doi/10.1145/3233547.3233548</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3532116953f23315366443db19af4f0d%2FSelection_052.png?generation=1714132907090436&amp;alt=media\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2776927,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-26T12:05:51.517000",
              "content": "<p>\" calculated properties like cLogP  …\"</p>\n<p>There are a couple of possible self-supervised task</p>\n<ul>\n<li>MLM : masked language modeling</li>\n<li>MTR: mult task regression (e.g. cLoP, …. i think anything related to shape and surface, electricty, energy)</li>\n<li>chemical reaction</li>\n</ul>\n<p>thanks!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2780098,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-28T03:47:15.373000",
              "content": "<p>yet another way to pretrain with test data …. basically it is \"smart clustering\" via VAE<br>\nlikewise, denoised diffusion also work (imagine train --&gt; noise --&gt; test)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbff4b909984bfcc298489abd9665ca16%2FSelection_055.png?generation=1714275971114985&amp;alt=media\"></p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2746350,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-11T07:07:40.860000",
      "content": "<p>With you around this comp' would be extra fun </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2747013,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-11T15:41:38.363000",
      "content": "<p>useful software tools. implementations, etc:</p>\n<ol>\n<li>multi-process df apply:<br>\n<a href=\"https://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores\" target=\"_blank\">https://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores</a><br>\n<a href=\"https://github.com/ddelange/mapply\" target=\"_blank\">https://github.com/ddelange/mapply</a></li>\n</ol>\n<pre><code>import mapply\nmapply.init(\n    =50,#-1,\n    =,\n)\ndf[] = df[].mapply(to_ecfp4_fun)\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 2747046,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-11T16:16:54.933000",
          "content": "<p>how to write fast submission code … run in minutes</p>\n<pre><code>probability = xgb(x_valid)\n = (, ) \n\nid = reduced_df]()\n\n\nid = id(-)\nprobability = probability(-)\n = id!=-\n==)\n\nsubmit_df = pd({\n    :id,\n    :probability,\n})\n submit_df(f,index=False)\n</code></pre>\n<p>reduced_df:<br>\n(id is -1 if protein is not evaluated)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b3628a98221b1f366382dc9e56e9ea2%2FSelection_022.png?generation=1712852204921481&amp;alt=media\"></p>",
          "votes": 4,
          "replies": [
            {
              "id": 2759822,
              "author_name": "Steven Hewitt",
              "author_url": "",
              "post_date": "2024-04-18T22:49:40.710000",
              "content": "<p>This is brilliant. Thank you for sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2747524,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-04-11T23:58:38.013000",
          "content": "<p>This is pretty neat! Having the progress bar is nice too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2748892,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-12T17:59:28.237000",
          "content": "<ol>\n<li>Easy Custom Losses for Tree Boosters using PyTorch<br>\n<a href=\"https://towardsdatascience.com/easy-custom-losses-for-tree-boosters-using-pytorch-57ffaa0b2eb3\" target=\"_blank\">https://towardsdatascience.com/easy-custom-losses-for-tree-boosters-using-pytorch-57ffaa0b2eb3</a><br>\n<a href=\"https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\" target=\"_blank\">https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec</a></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2799761,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-07T23:01:47.243000",
      "content": "<p>i just realize that you may have more training data</p>\n<ol>\n<li>molecule + bind label</li>\n<li>building block + soft bind label (average over all molecule using that block). you can treat it as molecule=block+wildcard here</li>\n</ol>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2766963,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-22T03:34:06.023000",
      "content": "<p>a pretty smart method that uses both labelled and unlabelled data</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa103fefd904bcddb90413a393dee5344%2FSelection_033.png?generation=1713756844349343&amp;alt=media\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2767009,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2024-04-22T04:25:42.573000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3586013%2F57be4534a687d8f516f898bc3aa77eff%2FScreenshot%202024-04-22%20at%2012-23-50%20ICAN%20Interpretable%20cross-attention%20network%20for%20identifying%20drug%20and%20target%20protein%20interactions%20-%20journal.pone.0276609.pdf.png?generation=1713759870859276&amp;alt=media\"></p>\n<p>Another paper describing a very similar approach:<br>\nICAN: Interpretable cross-attention network for identifying drug and target protein interactions</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2767112,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-22T05:51:08.660000",
              "content": "<p>there are a couple of similar papers. best to get those with github codes and experiment results from well known dataset. I foresee most deep net solution will be slow. how to sample training data could be an issue</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2767685,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-04-22T13:36:13.940000",
              "content": "<p>Agreed, efficient use of compute will be key.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2817098,
              "author_name": "genki0520",
              "author_url": "",
              "post_date": "2024-05-16T17:41:49.310000",
              "content": "<p>With only three types of proteins, is there too little protein data? If so, how can we tackle this problem?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2768571,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-23T00:09:17.210000",
          "content": "<p>a few good github repo that i am using:<br>\n<a href=\"https://github.com/larngroup/DTITR\" target=\"_blank\">https://github.com/larngroup/DTITR</a><br>\n<a href=\"https://github.com/ZXT0212/CAT-DTI\" target=\"_blank\">https://github.com/ZXT0212/CAT-DTI</a><br>\n<a href=\"https://github.com/peizhenbai/DrugBAN\" target=\"_blank\">https://github.com/peizhenbai/DrugBAN</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2770751,
              "author_name": "Patrick Chan",
              "author_url": "",
              "post_date": "2024-04-24T02:37:49.427000",
              "content": "<p>Great sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!! Did you get to try any of above methods yet for this competition?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2765165,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-21T01:43:43.427000",
      "content": "<p>yet another external data!!!!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4137ec834e1c8fd91b077bc2745775f%2FSelection_030.png?generation=1713663808992163&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4190bd04b74691820e74f3824afc4fa%2FSelection_031.png?generation=1713663821080691&amp;alt=media\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2770487,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-04-23T21:14:40.463000",
          "content": "<p>I'll take a look at generating a dataset from this close to our competition data when I have time (likely this weekend). Thanks for finding!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2770565,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-23T22:34:17.323000",
              "content": "<p>i have been looking at other papers and dataset. i think it actually make more sense to convert kaggle dataset to other dataset, i.e. replacing [Dy]</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2770636,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-04-24T00:28:09.887000",
              "content": "<p>Have you found consistency in what they use to model the DNA attachment point? I've done a little experimenting on my end switching Dy out for polyethyleneglycol like linkers and from a modeling perspective, I've seen similar results so far (which surprised me a little).</p>\n<p>One thing I have not tried is modeling the DNA attachment point as something small like a methyl group. Intuitively I'd think that providing the model some information on where the DNA attached would be valuable, but perhaps it's not.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2770758,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-04-24T02:46:36.887000",
              "content": "<p>I've been using your C/methyl conversion in my data pipeline (rest in progress).<br>\nIs switching methyl/linker to PEG as simple as swapping the C/Dy for OCCOCCO (PEG smile)?</p>\n<p>Also thanks for all the domain knowledge you've shared!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2771702,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-04-24T11:23:48.150000",
              "content": "<p>It's a little bit more complicated than that, but not much. I plan to post a notebook about it this weekend.</p>\n<p>Also, you are very welcome! Happy to help and share some knowledge of chemistry.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2748338,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-12T11:49:45.287000",
      "content": "<p>anyone one to make baseline for \"DiffDock\"?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2748932,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-04-12T18:14:32.807000",
          "content": "<p>There is a <a href=\"https://github.com/suneelbvs/DiffDock\" target=\"_blank\">Colab diffdock notebook</a> for anyone that is up for the challenge. With some tinkering, it can probably also be made to run on Kaggle.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2749654,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-04-13T07:05:43.363000",
              "content": "<p>DiffDock has a new version very recently, so I wouldn't try a colab notebook from 2 years ago except maybe as reference</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2749756,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-04-13T08:18:36.173000",
              "content": "<p>Good point, so maybe a lot of tinkering hehe</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2748942,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2024-04-12T18:20:17.240000",
          "content": "<p>I am working on building structural models. Not with DiffDock, but more traditional physics-based docking engines.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2850582,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-06-02T08:25:00.930000",
          "content": "<p>I recently made it to run, the thing is, it is very slow! Impossible to run so many compounds</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2768605,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-23T00:41:20.493000",
      "content": "<p>yet another results with Topological Torsion Fingerprint.<br>\nIt appears to me that just pure SMILE string model performance limit is 0.67 (for same block).</p>\n<p>there aren't huge different for transformer based, fingerprint based, etc ….<br>\nthe data has very very long tails …. (i.e. the positive cases don't cluster together in the chemical space to gives high peak density, instead they are spreading around)<br>\nThis make it difficult to generalized to unseen blocks, especially for fingerprint methods …</p>\n<p>need to think of new strategy</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e01cbdf17b3bb57233617c6cb3f5135%2FSelection_035.png?generation=1713832518151318&amp;alt=media\"> </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2767460,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-22T10:21:56.747000",
      "content": "<p>average precision trick?<br>\n<a href=\"https://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596</a></p>\n<p>\"This works because Average Precision metric is always improved when adding more predictions if the new predictions have a lower score than all previous predictions.\"</p>\n<p>hmm</p>\n<p>actually there is another way to look at our problem:<br>\nstage1:  predict no binding <br>\nstage2: predict binding score/rank (only less than 1% of data from stage1. enable use of more time consuming methods?)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2758038,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-17T22:18:20.453000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, what are your positive-to-negative ratios in your val sets? As was discussed <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/492126#2742181\" target=\"_blank\">here</a>, the metric strongly depends on the ratio. I suggest enforcing a known, constant ratio to facilitate easy comparison between results. The obvious one is the 1/128 ratio since this is the test set ratio…with this, I currently have validation scores of 0.53 for the all-BBs_shared val set, 0.29 for at least one BB is not shared and at least one BB is shared val set, and 0.018 for truly none of BBs is shared…yea the last one is tough lol. Probably different core would be even harder.</p>\n<p>EDIT: I got confused, the ratio should be 1/125. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2758217,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-18T03:10:38.957000",
          "content": "<p>be careful!  1/125 is only for public.</p>\n<p>always keep in mind: you have full access of the test data. you can measure some magic statistics on it. this is the trick to winning</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2758728,
              "author_name": "Gyula Maloveczky4",
              "author_url": "",
              "post_date": "2024-04-18T10:11:41.747000",
              "content": "<p>Will test.csv be the actual public leaderboard test? Or what do you mean?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2765402,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-04-21T06:00:26.990000",
              "content": "<p>Yes the test data is the full test data, including public and private leaderboard components <a href=\"https://www.kaggle.com/gyulamaloveczky4\" target=\"_blank\">@gyulamaloveczky4</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2768461,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-22T21:09:27.973000",
              "content": "<p>a submission of zero (or other constant) will revealed the percentage of all positive cases.</p>\n<p>what happen if you submit a couple of different constants, e.g.</p>\n<ul>\n<li>different values for different protein, different block,  etc</li>\n</ul>\n<p>hint :google for probing the kaggle leaderboard for average precision </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2758737,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-04-18T10:16:50.893000",
          "content": "<p>AP is a great metric because you can measure ranking performance of your model regardless of the pos/neg ratio, AP CV should correlate with LB even if pos/neg ratio is a bit shifted (we don't really care about absolute values of probabilities here)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2758939,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-04-18T11:58:56.593000",
              "content": "<p>True, as long as you use the same validation set. But if you want to compare the validation scores of two different sets (or people) you need to use the same ratio</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2769342,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-04-23T09:43:06.153000",
              "content": "<p>ratio squad?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2754473,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-16T04:28:05.317000",
      "content": "<p>i just realise that we need to probe the num of non train blocks in public test first. this can be done by two submission. a normal one and another with non train block masked off with magic value and observe the difference. </p>\n<p>we need to verify private and public share the same distribution of train and non train block.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2754528,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-04-16T05:01:48.073000",
          "content": "<p>There's like 34 unique buildingblock1s, with triazine core, not in train. And there's like 36 unique buildingblock1s, with different core, not in train. </p>\n<p>Depends how many subs you want to spend, but it would be useful to know, for each of those 70 unique BBs, and for each protein target, so 70x3, are some in public test? Are all in public test? Maybe half are in public test, half in private? If something like half and half, but no overlap (so public LB doesn't help \"learn\" the unseen private LB) then, maybe you can find out the groupings just by making a graph of which unique BBs are included with which other unique BBs. </p>\n<p>Example:<br>\nBB1: 1,2,3,4<br>\nBB2: 5-12<br>\nBB3: 13-20</p>\n<p>Test1: 1,5,13<br>\nTest2: 1,5,14<br>\n…<br>\nTestN: 2,5,13<br>\n1 and 2 both \"connect\" to 13, so they're in the same group. But maybe 5,13,14 etc never pair with 3 or 4, that would imply that one group (1,2,5,13,14,…) is either public or private, and the other group (3,4,…) is the other leaderboard group</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2755986,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-04-16T19:11:01.533000",
              "content": "<p>For triazines in test with BBs not in train, there's two separate blocks:<br>\n17 BB_a, 36 BB_b/c. (presumably public LB)<br>\n17 BB_a, 36 BB_b/c. (presumably private LB)</p>\n<p>And every possible combination is:<br>\n2 groups, times 3 proteins, times 17 BB_a, times 36 BB_b/c 2 combinations, the same BB is allowed to be taken twice, so it's:<br>\n2x3x17x(36+1)x36/2 = 67,932. Which is exactly the number of rows I observed. So that checks out.</p>\n<p>However, for the other group of 500,000 rows, it all appears to be one grouping. So I think this critical group we will probably find is NOT in public test at all! 500,000 rows is 167k per protein. Since the organizers said:</p>\n<blockquote>\n  <p>We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein</p>\n</blockquote>\n<p>and 360K - 200K = 160K, so I suspect the public and private are equivalent EXCEPT for this big curveball of the non-triazine molecules with unseen BBs will be only in private LB most likely.</p>\n<p>If someone proves this, please post back here!</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2780103,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-28T03:51:08.867000",
      "content": "<p>this can be used for augmentation!<br>\n<a href=\"https://chao1224.github.io/MoleculeSTM\" target=\"_blank\">https://chao1224.github.io/MoleculeSTM</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2774731,
      "author_name": "David J. Slate",
      "author_url": "",
      "post_date": "2024-04-25T10:00:15.547000",
      "content": "<p>In the specifications of your HP Z8 Fury Data Science Workstation you have \"256 MB random\".  Is this a typo that should be \"256 GB random\"?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2774754,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-25T10:10:46.637000",
          "content": "<p>oops! in <br>\nGB. thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2760019,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-19T04:11:48.083000",
      "content": "<p>for the past few days i have been training neural nets (e.g. string based transformer). Initial results seems to suggest to me that </p>\n<ul>\n<li>the choice of building blocks are not random.  </li>\n<li>there are too much data … some active learning or sampling is required (if we know how the building blocks are generated, it would be useful. e.g. rule based or the host uses some ML/AI-assisted software)</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2770641,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-04-24T00:32:13.700000",
          "content": "<p>I have not tried to figure out how the authors chose their dataset yet, but I do know that it's quite common to choose building blocks based on their calculated physical properties. This way, the dataset is biased towards final molecules with properties that more closely resemble drugs. For readings on this, perhaps Lipinski's rules would be a good start for some basic properties you'd like to see in a final molecule (these are common guidelines, not really hard and fast rules).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2754293,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-15T23:52:51.580000",
      "content": "<p>DTI plug and play :<br>\n<a href=\"https://github.com/yazdanimehdi/DeepDrugDomain\" target=\"_blank\">https://github.com/yazdanimehdi/DeepDrugDomain</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2753393,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-15T13:16:53.537000",
      "content": "<p>How do you construct the label for multi-label classification? <br>\nSome of the molecules bind to the multiple proteins, and XGBoost doesn't support this setup</p>\n<pre><code>pd.).value\n    \n     \n       \n          \nName: count, dtype: \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2753491,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-15T14:13:20.047000",
          "content": "<p>for cpu:  <br>\n<a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html\" target=\"_blank\">https://xgboost.readthedocs.io/en/stable/tutorials/multioutput.html</a><br>\nStarting from version 1.6, XGBoost has experimental support for multi-output regression and multi-label classification with Python package.</p>\n<p>for gpu:  <br>\nyou can write your custom loss. (e.g. kl divergence loss)<br>\nbut interesting, i find that the following 3 results are almost the same</p>\n<ul>\n<li>one model,  input = [feature + protein feature] where  protein feature =[1,0,0],[0,1,0],[0,0,1]</li>\n<li>three model, each is binary classification for each protein</li>\n<li>one model, just set target= Lx3 array and objective=\"binary:logistic\"<br>\nXGBoost uses one-vs-rest method<br>\n(<a href=\"https://discuss.xgboost.ai/t/understanding-leaf-values/1840\" target=\"_blank\">https://discuss.xgboost.ai/t/understanding-leaf-values/1840</a>)</li>\n</ul>\n<p>you can try a toy example, e.g. y =np.random.choice(2,(100,3)) and overfit a xgbost model. you should see it can gives multi-label output using the above methods.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2753538,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-15T14:51:13.407000",
              "content": "<pre><code>    x_train = np.random.choice(256,(20,256)).astype(np.uint8)\n    y_train = np.random.choice(2,(20,3)).astype(np.uint8)\n\n    clf = XGBClassifier(\n        =6,\n        =100,\n        # =,\n        =,\n        =,\n        objective  =, #, \n        learning_rate= 0.3,\n        device = ,\n        =42,\n        =1,\n    )\n\n    clf.fit(\n        x_train, y_train,\n        =\n    )\n\n    probability = clf.predict_proba(x_train)\n    predict = (probability&gt;0.5).astype(np.uint8)\n    (probability)\n    (predict)\n    (y_train)\n    ((y_train!=predict).sum())\n</code></pre>\n<p>y train</p>\n<pre><code>\n\n\n\n\n\n\n\n</code></pre>\n<p>probability</p>\n<pre><code>[[   ]\n [     ]\n [  ]\n [     ]\n [  ]\n [  ]\n [      ]\n [   ]\ny_train!=predict).sum() gives zero\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2747361,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-11T20:19:02.817000",
      "content": "<p>since it is micro average precision, one has to be careful in calibration, class-balancing or multi-class/single-class setup:</p>\n<pre><code>probability=np.([\n,\n,\n,\n,\n,\n,\n,\n,\n,\n])\ny=np.([\n,\n,\n,\n,\n,\n,\n,\n,\n,\n])\nscore (all micro) \nscore (, binds_BRD4) \nscore (, binds_HSA) \nscore (, binds_sEH) \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2747952,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-04-12T07:00:57.613000",
          "content": "<p>Thanks for sharing! Has any post-processing helped to achieve higher cv/lb score?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2748308,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-12T11:26:12.063000",
              "content": "<p>i leave this for future plan. improving accuracy for test building blocks not found in training is my top prority now.</p>\n<p>but you can google for differentiable micro average precision loss</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2748332,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-04-12T11:43:26.317000",
              "content": "<p>Your numbers for valid_noshare are very promising! It seems we can really generalize to new BBs. The question now is only how much.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2748342,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-12T11:51:21.550000",
              "content": "<p>\"Your numbers for valid_noshare are very promising!\"</p>\n<p>any model that uses train size larger 68_152_325 (see subsample column) will not have validation metrics. (in this case, the training set include the validation set).</p>\n<p>i would say that current model for valid nonshare has precision half of that valid share.  <br>\nbut according to papers, SOTA can probably increase that to 70%</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2776107,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-26T02:12:11.753000",
      "content": "<p>i found my ideal algorithm<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0c62172bd70306926304d48f42f18fe6%2FSelection_047.png?generation=1714097530102992&amp;alt=media\"></p>\n<p><a href=\"https://github.com/luwei0917/TankBind\" target=\"_blank\">https://github.com/luwei0917/TankBind</a></p>\n<p>\"TankBind also support virtual screening. In our example here, for the WDR domain of LRRK2 protein, we can screen 10,000 drug candidates in 2 minutes (or 1M in around 3 hours) with a single GPU. Check out\"</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2778802,
          "author_name": "Wang-Lin-Boop",
          "author_url": "",
          "post_date": "2024-04-27T11:17:32.807000",
          "content": "<p>Hi, if you trust that DTI prediction may be the optimal solution for this task, may be you can try some of the latest models, such as: </p>\n<ol>\n<li><a href=\"https://github.com/Zhang-Runze/PackDock\" target=\"_blank\">https://github.com/Zhang-Runze/PackDock</a></li>\n<li><a href=\"https://github.com/HBioquant/DiffBindFR\" target=\"_blank\">https://github.com/HBioquant/DiffBindFR</a></li>\n</ol>\n<p>The rationale behind this is that some early DTI prediction methods have been criticized in recent papers for their unreasonable docking poses and poor generalization ability, e. g. <a href=\"https://arxiv.org/abs/2308.05777\" target=\"_blank\">https://arxiv.org/abs/2308.05777</a>. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2778817,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-27T11:31:10.350000",
              "content": "<p>\" unreasonable docking poses and poor generalization ability\"<br>\ni agree with these.</p>\n<p>from competition point of view:</p>\n<ul>\n<li>for the non sharing blocks, most competitors are going to perform poorly.</li>\n<li>ML does not work if the train data significantly differs from these nonshare test.<br>\n(you may to refer to <a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/belka-generalizability-similarity</a>)</li>\n<li>if you can do better in these test subset (even just mere 20% to 40% of these), maybe you can have better private score. </li>\n</ul>\n<p>so my strategy is simple:</p>\n<ul>\n<li>QSAR for majority of the results</li>\n<li>self-supervsied, external data/block-based modeling to improve nonshare blocks</li>\n<li>DTI to improve/anaylse to results</li>\n<li>docking for a small \"selected\" subset</li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2780192,
              "author_name": "Wang-Lin-Boop",
              "author_url": "",
              "post_date": "2024-04-28T05:03:40.963000",
              "content": "<blockquote>\n  <p>QSAR for majority of the results<br>\n  self-supervsied, external data/block-based modeling to improve nonshare blocks<br>\n  DTI to improve/anaylse to results<br>\n  docking for a small \"selected\" subset</p>\n</blockquote>\n<p>Great! Indeed, for real-world projects, we have indeed employed approachs like this. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2774889,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-25T11:14:52.020000",
      "content": "<p>i wonder if something like this is possible?<br>\nwe have the blocks and the reaction results.(for both train and test)</p>\n<p>Contextual Molecule Representation Learning from Chemical Reaction Knowledge <br>\n<a href=\"https://arxiv.org/html/2402.13779v1\" target=\"_blank\">https://arxiv.org/html/2402.13779v1</a></p>\n<ul>\n<li>transformer based SSL method</li>\n<li>Molecular Representation Learning (MRL)</li>\n</ul>\n<p>\"REMO, a self-supervised learning framework that takes advantage of well-defined atom-combination rules in common chemistry. Specifically, REMO pre-trains graph/Transformer encoders on 1.7 million known chemical reactions in the literature. We propose two pre-training objectives: Masked Reaction Centre Reconstruction (MRCR) and Reaction Centre Identification (RCI). REMO offers a novel solution to MRL by exploiting the underlying shared patterns in chemical reactions as context for pre-training, which effectively infers meaningful representations of common chemistry knowledge\"</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2773938,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-04-25T01:21:45.663000",
      "content": "<p>Wherever you are, the competition will be very exciting. Inspired by you, I have also started a similar discussion in BirdCLEF2024, and will continue to update it.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2771701,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-24T11:23:44.463000",
      "content": "<p>Stochastic Optimization of Areas Under Precision-Recall Curves with Provable Convergence<br>\n<a href=\"https://arxiv.org/abs/2104.08736\" target=\"_blank\">https://arxiv.org/abs/2104.08736</a></p>\n<p>example usage<br>\n<a href=\"https://docs.libauc.org/examples/auprc.html\" target=\"_blank\">https://docs.libauc.org/examples/auprc.html</a></p>\n<pre><code>model = ResNet18(=, =None, =1)\nmodel = model.cuda()\n\nloss_fn = APLoss(=len(trainSet), =margin, =gamma)\noptimizer = SOAP(model.parameters(), =lr, =, =weight_decay)\n</code></pre>\n<p>see also<br>\n<a href=\"https://github.com/divelab/MoleculeX\" target=\"_blank\">https://github.com/divelab/MoleculeX</a></p>\n<p>In addition, AdvProp is able to deal with tasks in which samples from different classes are highly imbalanced. In these cases, we employ advanced loss functions that optimize various areas under curves (AUC), such as areas under the receiver operating characteristic (AUROC) and the precision recall curve (AUPRC). AdvProp has been used to participate in the AI Cures open challenge for COVID-19 and is now ranked #1 in terms of both AUROC and AUPRC on the leaderboard. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2850573,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-06-02T08:18:31.437000",
          "content": "<p>Interesting to see you find my friend’s work on advprop, I’m trying to optimize it for this dataset</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2775243,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-25T15:04:45.750000",
      "content": "<p>i suddenly have an idea.<br>\nfor the test smiles, run some tools or simulations (e.g. energy,  force field, conformers?) to get some measurements related to binding. learn model to predict such quality on test.</p>\n<p>can check correlation of measurements with target bind on train, tsNE, etc</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2775778,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2024-04-25T19:38:01.787000",
          "content": "<p>do you mean to say \"for the <strong>train</strong> smiles…\" ? if not, could you explain a little more what you mean?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2775825,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-25T20:27:55.740000",
              "content": "<p>the test smiles</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2754267,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-15T23:38:16.950000",
      "content": "<p>what? joint protein-molecue fingerprint? interesting ….<br>\n<a href=\"https://academic.oup.com/bioinformatics/article/35/8/1334/5092926\" target=\"_blank\">https://academic.oup.com/bioinformatics/article/35/8/1334/5092926</a></p>\n<p>Development of a protein–ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2748402,
      "author_name": "Oleg Kokorin",
      "author_url": "",
      "post_date": "2024-04-12T12:26:35.590000",
      "content": "<p>Training on 98M samples with 2048 bits each gives almost 200GB of training data. Do I understand that correctly? Just wondering what hardware you use for it.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2748476,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-12T13:18:06.317000",
          "content": "<p>use one byte to represent 8 bit.<br>\nso 2048 bit=256 byte.</p>\n<p>see np.packbits.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2748546,
              "author_name": "Oleg Kokorin",
              "author_url": "",
              "post_date": "2024-04-12T13:38:58.040000",
              "content": "<p>My bad. Thanks!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2753694,
      "author_name": "Harsh Upadhyay",
      "author_url": "",
      "post_date": "2024-04-15T16:05:54.067000",
      "content": "<p>With you around this comp' would be extra fun</p>",
      "votes": -5,
      "replies": []
    },
    {
      "id": 2875171,
      "author_name": "Leon Zhou",
      "author_url": "",
      "post_date": "2024-06-16T23:39:34.410000",
      "content": "<p>Curious have you tried the 77M version of chemberta or molformer?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2816398,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-16T09:44:15.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2804674,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-10T06:44:37.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2794938,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-05T15:26:51.117000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2794282,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-05T07:41:50.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2776056,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-26T00:56:49.790000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2775830,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T20:36:06.070000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2775703,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T18:44:39.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2775219,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T14:48:35.107000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2774925,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T11:47:07.613000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2775040,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-25T12:57:50.753000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2775078,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-25T13:17:25.933000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2775189,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T14:25:50.677000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2775192,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T14:32:55.160000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2775224,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T14:49:43.023000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2775294,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T15:39:05.960000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2775440,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T16:58:22.993000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2775571,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T17:37:59.577000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2775674,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T18:35:21.840000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2775683,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-25T18:38:05.603000",
              "content": "",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2776086,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T01:32:48.620000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2776451,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T07:19:00.877000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2776489,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T07:48:53.673000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2776523,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T08:08:28.753000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2776596,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T08:43:36.453000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2776605,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T08:48:14.073000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2776891,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-26T11:45:22.617000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2778251,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-27T04:02:34.360000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2778409,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-27T06:37:01.243000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2774264,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T05:42:14.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2773983,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-25T02:30:26.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2756352,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-17T02:19:55.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2753345,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-15T12:47:53.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2748860,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-12T17:30:33.590000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2748863,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-12T17:36:39.897000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2748905,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:05:55.867000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2748912,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:08:17.393000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2748940,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:19:35.737000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2748957,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:31:41.197000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2748986,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:50:53.827000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2748993,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T18:57:42.090000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2749035,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-12T19:44:25.307000",
              "content": "",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2747606,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-12T02:05:21.030000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2748482,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-12T13:20:51.477000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2747954,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-12T07:01:42.257000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2757671,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-17T16:47:19.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2748731,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-12T15:08:47.207000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2746313": "## [Acknowledgement]\n\n\"We extend our thanks to HP for providing the HP Z8 Fury Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"\nhttps://www.hp.com/us-en/workstations/z8-fury.html\n\nspecifications:\n- 2x RTX 6000 Ada GPU (48 GB ram each) :  xgboost training of 10M molecule ecfps at 45 minutes!\n- Intel® Xeon® W9 Processor / 56cores : fast extraction with rdkit, etc\n- 256 GB ram\n\n---\n\n##[experimental results]\nuseful:\n\necfp is binary vector. so you can decrease memory by using np.packbits, that is about 98m x 256 bytes\n\npytorch version of np.unpackbits\n- https://gist.github.com/vadimkantorov/30ea6d278bc492abf6ad328c6965613a\n- https://github.com/pytorch/pytorch/issues/32867\n\nprocessed dataset for discussion:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd971971a363e7bf8de359ffdaf72ecf2%2FSelection_077.png?generation=1715120816304983&alt=media)",
    "2748792": "dataset for the above discussion have been shared!!!!\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\nIn summary:\n- train.reduced.parquet : 98_415_610 training SMILES and their information\n- train.ecfp4.packed.npz : Features extracted using rdkit  \n   - AllChem.GetMorganFingerprintAsBitVect(mol, 2, 2048)\t\t\t\t\t\t\t\t\t\t\t\t\t\n   - repack with np.packbits() to give 98_415_610 x 256 feature matrix\n- train.bind.npz : 98_415_610 x 3 target matrix\n- test.reduced.parquet/ test.ecfp4.packed.npz : similarly processed for the test SMILES\n- all_buildingblock.csv: building blocks id used in train.reduced.parquet/test.reduced.parquet\n- fold0.parquet: train_share,valid_share,valid_nonshare splits for the experiments in the discussion",
    "2746508": "Thanks for sharing ! \n\nMy colleagues and me sometimes organize introductory webinars on Kaggle competitions \nSome video records here: https://www.youtube.com/watch?v=aqUOz3nFYm4&list=PL1GnA8S_asBANHU95kOPOPEh3moPm5uhw\n\nIs there any chance that you can make an introduction to that competition ? \nJust intro survey - data, metrics, public solutions, whatsever.\nIt is not a private sharing - we make  anounces here on Kaggle forum - everybody can join via zoom\nand ask questions/comments/whatsever. Video will be later on Youtube.\nTiming about 40 minutes \n\nSometimes organizers join the zoom and make their comments also:\nhttps://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/421170\n\nIn general it seems quite helpful for community. \nAnother examples here:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444825\nhttps://www.kaggle.com/competitions/commonlit-evaluate-student-summaries/discussion/429975",
    "2765385": "transformer (pure SMILES string) results out !!!!\nhttps://huggingface.co/ibm/MoLFormer-XL-both-10pct\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa4b79498cbc6dd566f0dcb15a31d6272%2FSelection_032.png?generation=1713678750507300&alt=media)",
    "2776878": "one of my idea\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F642a737a82d59e4b1ed1fba2b312cc8e%2FSelection_051.png?generation=1714131726333596&alt=media)",
    "2746350": "With you around this comp' would be extra fun ",
    "2747013": "useful software tools. implementations, etc:\n1. multi-process df apply:\nhttps://stackoverflow.com/questions/45545110/make-pandas-dataframe-apply-use-all-cores\nhttps://github.com/ddelange/mapply\n\n```\nimport mapply\nmapply.init(\n    n_workers=50,#-1,\n    progressbar=True,\n)\ndf['ecfp'] = df['molecule_smiles'].mapply(to_ecfp4_fun)\n```",
    "2799761": "i just realize that you may have more training data\n1. molecule + bind label\n2. building block + soft bind label (average over all molecule using that block). you can treat it as molecule=block+wildcard here",
    "2766963": "a pretty smart method that uses both labelled and unlabelled data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa103fefd904bcddb90413a393dee5344%2FSelection_033.png?generation=1713756844349343&alt=media)",
    "2765165": "yet another external data!!!!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4137ec834e1c8fd91b077bc2745775f%2FSelection_030.png?generation=1713663808992163&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4190bd04b74691820e74f3824afc4fa%2FSelection_031.png?generation=1713663821080691&alt=media)",
    "2748338": "anyone one to make baseline for \"DiffDock\"?",
    "2768605": "yet another results with Topological Torsion Fingerprint.\nIt appears to me that just pure SMILE string model performance limit is 0.67 (for same block).\n\nthere aren't huge different for transformer based, fingerprint based, etc ....\nthe data has very very long tails .... (i.e. the positive cases don't cluster together in the chemical space to gives high peak density, instead they are spreading around)\nThis make it difficult to generalized to unseen blocks, especially for fingerprint methods ...\n\nneed to think of new strategy\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5e01cbdf17b3bb57233617c6cb3f5135%2FSelection_035.png?generation=1713832518151318&alt=media) ",
    "2767460": "average precision trick?\nhttps://www.kaggle.com/competitions/child-mind-institute-detect-sleep-states/discussion/459596\n\n\"This works because Average Precision metric is always improved when adding more predictions if the new predictions have a lower score than all previous predictions.\"\n\nhmm\n\nactually there is another way to look at our problem:\nstage1:  predict no binding \nstage2: predict binding score/rank (only less than 1% of data from stage1. enable use of more time consuming methods?)",
    "2758038": "@hengck23, what are your positive-to-negative ratios in your val sets? As was discussed [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/492126#2742181), the metric strongly depends on the ratio. I suggest enforcing a known, constant ratio to facilitate easy comparison between results. The obvious one is the 1/128 ratio since this is the test set ratio...with this, I currently have validation scores of 0.53 for the all-BBs_shared val set, 0.29 for at least one BB is not shared and at least one BB is shared val set, and 0.018 for truly none of BBs is shared...yea the last one is tough lol. Probably different core would be even harder.\n\nEDIT: I got confused, the ratio should be 1/125. ",
    "2754473": "i just realise that we need to probe the num of non train blocks in public test first. this can be done by two submission. a normal one and another with non train block masked off with magic value and observe the difference. \n\nwe need to verify private and public share the same distribution of train and non train block.",
    "2780103": "this can be used for augmentation!\nhttps://chao1224.github.io/MoleculeSTM",
    "2774731": "In the specifications of your HP Z8 Fury Data Science Workstation you have \"256 MB random\".  Is this a typo that should be \"256 GB random\"?\n",
    "2760019": "for the past few days i have been training neural nets (e.g. string based transformer). Initial results seems to suggest to me that \n- the choice of building blocks are not random.  \n- there are too much data ... some active learning or sampling is required (if we know how the building blocks are generated, it would be useful. e.g. rule based or the host uses some ML/AI-assisted software)",
    "2754293": "DTI plug and play :\nhttps://github.com/yazdanimehdi/DeepDrugDomain",
    "2753393": "How do you construct the label for multi-label classification? \nSome of the molecules bind to the multiple proteins, and XGBoost doesn't support this setup\n```\npd.Series(train_bind.sum(axis=1)).value_counts()\n0    96905831\n1     1429714\n2       80003\n3          62\nName: count, dtype: int64\n```",
    "2747361": "since it is micro average precision, one has to be careful in calibration, class-balancing or multi-class/single-class setup:\n```\n\nprobability=np.array([\n\t[0.9,0.2,0.2],\n\t[0.9,0.2,0.2],\n\t[0.9,0.2,0.2],\n\t[0.6,0.3,0.2],\n\t[0.5,0.3,0.2],\n\t[0.1,0.3,0.2],\n\t[0.1,0.2,0.5],\n\t[0.1,0.2,0.5],\n\t[0.1,0.2,0.5],\n])\ny=np.array([\n\t[1,0,0],\n\t[1,0,0],\n\t[1,0,0],\n\t[0,1,0],\n\t[0,1,0],\n\t[0,1,0],\n\t[0,0,1],\n\t[0,0,1],\n\t[0,0,1],\n])\nscore (all micro) 0.856060606060606\nscore (0, binds_BRD4) 1.0\nscore (1, binds_HSA) 1.0\nscore (2, binds_sEH) 1.0\n```",
    "2776107": "i found my ideal algorithm\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0c62172bd70306926304d48f42f18fe6%2FSelection_047.png?generation=1714097530102992&alt=media)\n\nhttps://github.com/luwei0917/TankBind\n\n\"TankBind also support virtual screening. In our example here, for the WDR domain of LRRK2 protein, we can screen 10,000 drug candidates in 2 minutes (or 1M in around 3 hours) with a single GPU. Check out\"",
    "2774889": "i wonder if something like this is possible?\nwe have the blocks and the reaction results.(for both train and test)\n\nContextual Molecule Representation Learning from Chemical Reaction Knowledge \nhttps://arxiv.org/html/2402.13779v1\n- transformer based SSL method\n- Molecular Representation Learning (MRL)\n\n\"REMO, a self-supervised learning framework that takes advantage of well-defined atom-combination rules in common chemistry. Specifically, REMO pre-trains graph/Transformer encoders on 1.7 million known chemical reactions in the literature. We propose two pre-training objectives: Masked Reaction Centre Reconstruction (MRCR) and Reaction Centre Identification (RCI). REMO offers a novel solution to MRL by exploiting the underlying shared patterns in chemical reactions as context for pre-training, which effectively infers meaningful representations of common chemistry knowledge\"",
    "2773938": "Wherever you are, the competition will be very exciting. Inspired by you, I have also started a similar discussion in BirdCLEF2024, and will continue to update it.",
    "2771701": "Stochastic Optimization of Areas Under Precision-Recall Curves with Provable Convergence\nhttps://arxiv.org/abs/2104.08736\n\nexample usage\nhttps://docs.libauc.org/examples/auprc.html\n\n```\nmodel = ResNet18(pretrained=False, last_activation=None, num_classes=1)\nmodel = model.cuda()\n\nloss_fn = APLoss(data_len=len(trainSet), margin=margin, gamma=gamma)\noptimizer = SOAP(model.parameters(), lr=lr, mode='adam', weight_decay=weight_decay)\n```\n\nsee also\nhttps://github.com/divelab/MoleculeX\n\nIn addition, AdvProp is able to deal with tasks in which samples from different classes are highly imbalanced. In these cases, we employ advanced loss functions that optimize various areas under curves (AUC), such as areas under the receiver operating characteristic (AUROC) and the precision recall curve (AUPRC). AdvProp has been used to participate in the AI Cures open challenge for COVID-19 and is now ranked #1 in terms of both AUROC and AUPRC on the leaderboard. ",
    "2775243": "i suddenly have an idea.\nfor the test smiles, run some tools or simulations (e.g. energy,  force field, conformers?) to get some measurements related to binding. learn model to predict such quality on test.\n\ncan check correlation of measurements with target bind on train, tsNE, etc",
    "2754267": "what? joint protein-molecue fingerprint? interesting ....\nhttps://academic.oup.com/bioinformatics/article/35/8/1334/5092926\n\nDevelopment of a protein–ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions",
    "2748402": "Training on 98M samples with 2048 bits each gives almost 200GB of training data. Do I understand that correctly? Just wondering what hardware you use for it.",
    "2753694": "With you around this comp' would be extra fun",
    "2875171": "Curious have you tried the 77M version of chemberta or molformer?",
    "2816398": "BRD4 dataset \nhttps://github.com/shiwentao00/Graphsite-classifier/tree/master\nGraphSite: Ligand Binding Site Classification with Deep Graph Learning\nhttps://www.mdpi.com/2218-273X/12/8/1053",
    "2804674": "some probing tutorial:\nfor more information refer to https://stats.stackexchange.com/questions/394494/calculating-sklearns-average-precision-by-hand\n\nhttps://www.kaggle.com/code/hengck23/how-to-probe",
    "2794938": "Really nice work on the FP+SMILES model @hengck23 !",
    "2794282": "old results for archive:\n\n\nWARNING: \n- results may not be optimized/tuned (since i am still experimenting ... \n   it will be updated throughout the competition till last 2 weeks)  \n- **the split are not correct, valid nonshare actually used share blocks by mistake !!! (see discussion below)**\n\n---\ni think i can conclude transformer beats xgboost+fingerprint!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a7872cf8060014be810dd706bab1133%2FSelection_036.png?generation=1713929270479439&alt=media)\n\nnew results with gpu training\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca8428acaa3ffa2501850471183ca09b%2FSelection_026.png?generation=1713066709109703&alt=media)\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdaa7c3665fe31df78e61ffac106f4d7e%2FSelection_024.png?generation=1712911082033973&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2e854b52473673388a32ab6c76d93b6%2FSelection_017.png?generation=1712818198030642&alt=media)",
    "2776056": "https://www.thesgc.org/news/structural-genomics-consortium-and-x-chem-enter-collaboration-unlock-human-proteome-and-promote\nThese datasets, curated in an ML-ready format, will be posted to a public portal to be used for model building. ...  but X-Chem and the SGC expect the DEL-ML data portal will be ready for public access starting in early 2024.",
    "2775830": "![https://www.youtube.com/watch?v=AuNXa0nNkAc](https://www.youtube.com/watch?v=AuNXa0nNkAc)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F188a7ef0c66e995e004ceb9a9dd276da%2FSelection_046.png?generation=1714077718909224&alt=media)\n\nbb considerations",
    "2775703": "there is one trick.\none can \"google\" for compound that is close to the test molecule .... write an LLM agent for that\n \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fda40f675564d703d03c4f25b0cffa68b%2FSelection_044.png?generation=1714070663174093&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7cfa93726f6544be9304b8fb5f6929a8%2FSelection_043.png?generation=1714070675606349&alt=media)\n",
    "2775219": "yet another external data\nhttps://chemrxiv.org/engage/chemrxiv/article-details/60c741e2567dfedeb7ec3e52\n(see Supplementary materials)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F688846386de6863b1ddc4e84f7d3598a%2FSelection_041.png?generation=1714056462488116&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F035dffab1de71e8666b952f275e21f1f%2FSelection_042.png?generation=1714056474616688&alt=media)",
    "2774925": "i have completed \n1. fingerprint + xgboost\n2. SMILE stsing transmformer\n\nNow i am trying (in process of training ...) 2d/3d GNN (i.e. input = molecular graph).\nIf anyone has results in GNN, or good papers/github repo to recommend, or tips in handling huge data of grpah, it is most welcomed!",
    "2774264": "this is what i done when i do not know when to stop training. just average the just few checkpoints\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F19b0eb03e19d34bae44873186d47055c%2FSelection_039.png?generation=1714023733193529&alt=media)",
    "2773983": "cascade model, active learning and submission\n\nour final model is a cascade of models, each model has high recall of say 99%, but moderate precision say 50%.\nassume each model has recall =0.99, FPrate =0.10, then a cascade of of 5 such model will be\nrecall=pow(0.99,5)=still high,  FPrate =pow(0.1,5)=very low = 1e-6\n\nhow it work:\n```\n0. all data x\n1. x is discarded if model0(x)< threshold0, else go to next model\n2. x is discarded if model1(x)< threshold1, else go to next model\n3. x is discarded if model2(x)< threshold2, else go to next model\n....\nonly few data remains\n(here we can use complicated methods like docking, MD simulation, complex fingerprint, large deep net, etc)\n\n```\n\nat submission\n\n````\n0. submit all zero rank (e.g. lb score =s0)\n1. for rejected sample, submit as rank=0, for remaining sample rank=1 \n(if new lb score s1==s0, then you have 100 recall in public test)\n2. likewise, new accepted sample will have rank2 (i.e. submission has only 3 values 0,1,2)\n3, etc ....\n....\n\nthe trick is that try to keeps0=s1=s2=s3 ... then you have zero miss and left with very little data for next model\n\n```\n\nreference: \ncascade adaboost for face detection. i think xgboost can be modified to do this.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb2302b19ce617b9d488c9a303ad1aeb%2FSelection_038.png?generation=1714012410058499&alt=media)\n\n",
    "2756352": "diffdock: try nvidia online demo\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F18ff3e4e995521d5794c1f5080d1cc8b%2FSelection_027.png?generation=1713320365421742&alt=media)\n\nhttps://developer.nvidia.com/blog/new-models-molmim-and-diffdock-power-molecule-generation-and-molecular-docking-in-bionemo/\n\nmaybe can apply for cloud beta services for benchmarking? if not, you can setup on your own gpu.",
    "2753345": "@hengck23 Thank you for giving us a great insight on this competition! I have a small question though. What strategy did you use to sub-sample the dataset into 25m?",
    "2748860": "Did you split by maximizing nonshare groups via clustering? Because no way a naive splitting would give 13% nonshare. When I split each BB group to 5F it results in something on the order of 1% nonshare (~1M molecules out of ~100M). And only ~50% train. When I do 10F it drops to ~0.1% nonshare. I fyou did clustering, I suggest you check #positive in each group because it may not be random.",
    "2747606": "@hengck23 You may find it interesting that the LB score is very similar with LGBM instead of XGBoost, shows similar improvement when scaling up to more estimators, and trains much faster. Could be helpful for experiments. I don't have exact time measurements, unfortunately.",
    "2747954": "",
    "2757671": "Thanks for sharing new ideas!!",
    "2748731": "thanks for sharing idea!"
  }
}