{
  "id": 505985,
  "title": "[lb0.420 at train=40m] sphere-net: fast 3d-Graph-NN is here !",
  "url": "/competitions/leash-BELKA/discussion/505985",
  "author_name": "",
  "post_date": "2024-05-20T03:42:11.357952900Z",
  "votes": 28,
  "comment_count": 18,
  "views": 0,
  "content": "<p>This is work in progress.<br>\nMy first target is SMP (spherical message passing) sphere-net.<br>\nif you have other good 3d graph-NN, please advised here! Thanks.</p>\n<p>[paper] Spherical Message Passing for 3D Graph Networks - Yi Liu, ICLR 2022 <br>\n<a href=\"https://paperswithcode.com/paper/spherical-message-passing-for-3d-graph\" target=\"_blank\">https://paperswithcode.com/paper/spherical-message-passing-for-3d-graph</a></p>\n<p>first:</p>\n<ul>\n<li>fast conformer generation, with sphere-net example<br>\n<a href=\"https://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator\" target=\"_blank\">https://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator</a></li>\n</ul>\n<p>some generated conformers (<strong>10 million</strong>) found at:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p>submission:<br>\ntrain=10million conformer: lb=0.383</p>\n<p>for the below: did not retrain just add new data and finetune. retrain should give slightly better results …<br>\ntrain=20million : lb=0.387<br>\ntrain=40million : lb=0.420<br>\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F983685dd4eb3f2a9bc2cbc1fcc9a14a8%2FSelection_132.png?generation=1716176444335864&amp;alt=media\"></p>\n<p>results look good:<br>\n[paper] 'High-Quality Conformer Generation with CONFORGE: Algorithm and Performance Assessment' - T Seidel , acs 2023<br>\n<a href=\"https://pubs.acs.org/doi/10.1021/acs.jcim.3c00563\" target=\"_blank\">https://pubs.acs.org/doi/10.1021/acs.jcim.3c00563</a><br>\n<a href=\"https://github.com/molinfo-vienna/CDPKit\" target=\"_blank\">https://github.com/molinfo-vienna/CDPKit</a></p>\n<p>with parallel processing, I was able to generate 10 million conformers with 4 hr</p>",
  "messages": [
    {
      "id": "2824788",
      "postDate": "05/20/2024 03:42:11",
      "content": "<p>This is work in progress.<br>\nMy first target is SMP (spherical message passing) sphere-net.<br>\nif you have other good 3d graph-NN, please advised here! Thanks.</p>\n<p>[paper] Spherical Message Passing for 3D Graph Networks - Yi Liu, ICLR 2022 <br>\n<a href=\"https://paperswithcode.com/paper/spherical-message-passing-for-3d-graph\" target=\"_blank\">https://paperswithcode.com/paper/spherical-message-passing-for-3d-graph</a></p>\n<p>first:</p>\n<ul>\n<li>fast conformer generation, with sphere-net example<br>\n<a href=\"https://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator\" target=\"_blank\">https://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator</a></li>\n</ul>\n<p>some generated conformers (<strong>10 million</strong>) found at:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p>submission:<br>\ntrain=10million conformer: lb=0.383</p>\n<p>for the below: did not retrain just add new data and finetune. retrain should give slightly better results …<br>\ntrain=20million : lb=0.387<br>\ntrain=40million : lb=0.420<br>\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F983685dd4eb3f2a9bc2cbc1fcc9a14a8%2FSelection_132.png?generation=1716176444335864&amp;alt=media\"></p>\n<p>results look good:<br>\n[paper] 'High-Quality Conformer Generation with CONFORGE: Algorithm and Performance Assessment' - T Seidel , acs 2023<br>\n<a href=\"https://pubs.acs.org/doi/10.1021/acs.jcim.3c00563\" target=\"_blank\">https://pubs.acs.org/doi/10.1021/acs.jcim.3c00563</a><br>\n<a href=\"https://github.com/molinfo-vienna/CDPKit\" target=\"_blank\">https://github.com/molinfo-vienna/CDPKit</a></p>\n<p>with parallel processing, I was able to generate 10 million conformers with 4 hr</p>",
      "rawMarkdown": "This is work in progress.\nMy first target is SMP (spherical message passing) sphere-net.\nif you have other good 3d graph-NN, please advised here! Thanks.\n\n[paper] Spherical Message Passing for 3D Graph Networks - Yi Liu, ICLR 2022 \nhttps://paperswithcode.com/paper/spherical-message-passing-for-3d-graph\n\n\nfirst:\n- fast conformer generation, with sphere-net example\nhttps://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator\n\nsome generated conformers (**10 million**) found at:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\nsubmission:\ntrain=10million conformer: lb=0.383\n\nfor the below: did not retrain just add new data and finetune. retrain should give slightly better results ...\ntrain=20million : lb=0.387\ntrain=40million : lb=0.420\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F983685dd4eb3f2a9bc2cbc1fcc9a14a8%2FSelection_132.png?generation=1716176444335864&alt=media)\n\nresults look good:\n[paper] 'High-Quality Conformer Generation with CONFORGE: Algorithm and Performance Assessment' - T Seidel , acs 2023\nhttps://pubs.acs.org/doi/10.1021/acs.jcim.3c00563\nhttps://github.com/molinfo-vienna/CDPKit\n\nwith parallel processing, I was able to generate 10 million conformers with 4 hr",
      "votes": null
    },
    {
      "id": "2827142",
      "postDate": "05/21/2024 09:07:38",
      "content": "<p>spherical CNN has been introduced by the host. Here is a quick visual to explain what it is<br>\n<a href=\"https://research.google/blog/scalable-spherical-cnns-for-scientific-applications/\" target=\"_blank\">https://research.google/blog/scalable-spherical-cnns-for-scientific-applications/</a><br>\ncode: <a href=\"https://github.com/google-research/spherical-cnn/\" target=\"_blank\">https://github.com/google-research/spherical-cnn/</a><br>\n(maybe someone can try this jax code on kaggle tpu)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F45020dd36f6f830bddf5f59c12677a39%2Fimage6.gif?generation=1716282431762462&amp;alt=media\"></p>\n<p>Each atom is represented by a set of spherical signals accumulating physical interactions with other atoms of each type (shown in the three panels on the right). For example, the oxygen atom (O; top panel) has a channel for oxygen (indicated by the sphere labeled “O” on the left) and hydrogen (“H”, right). The accumulated Coulomb forces on the oxygen atom with respect to the two hydrogen atoms is indicated by the red shaded regions on the bottom of the sphere labeled “H”. Because the oxygen atom contributes no forces to itself, the “O” sphere is uniform. We include extra channels for the Van der Waals forces.</p>",
      "rawMarkdown": "spherical CNN has been introduced by the host. Here is a quick visual to explain what it is\nhttps://research.google/blog/scalable-spherical-cnns-for-scientific-applications/\ncode: https://github.com/google-research/spherical-cnn/\n(maybe someone can try this jax code on kaggle tpu)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F45020dd36f6f830bddf5f59c12677a39%2Fimage6.gif?generation=1716282431762462&alt=media)\n\nEach atom is represented by a set of spherical signals accumulating physical interactions with other atoms of each type (shown in the three panels on the right). For example, the oxygen atom (O; top panel) has a channel for oxygen (indicated by the sphere labeled “O” on the left) and hydrogen (“H”, right). The accumulated Coulomb forces on the oxygen atom with respect to the two hydrogen atoms is indicated by the red shaded regions on the bottom of the sphere labeled “H”. Because the oxygen atom contributes no forces to itself, the “O” sphere is uniform. We include extra channels for the Van der Waals forces.",
      "votes": null
    },
    {
      "id": "2827150",
      "postDate": "05/21/2024 09:10:20",
      "content": "<p>this is somehow smiliar to sphere-net (SMP)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea38e45ffbcd1a5dff71c224c5fb2eb5%2FSelection_133.png?generation=1716282593210660&amp;alt=media\"></p>\n<p>it wold be interesting to see comparison of the two</p>",
      "rawMarkdown": "this is somehow smiliar to sphere-net (SMP)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea38e45ffbcd1a5dff71c224c5fb2eb5%2FSelection_133.png?generation=1716282593210660&alt=media)\n\nit wold be interesting to see comparison of the two",
      "votes": null
    },
    {
      "id": "2828291",
      "postDate": "05/22/2024 02:49:49",
      "content": "<p>first 3d gnn sphere net result!!!</p>\n<p><strong>UPDATE: there is a bug in my code. I am only training with my validation set of 400k.</strong><br>\n(i will update the results later)</p>\n<p>i hope there is no bugs in my code. results are too good (though not unreasonable).<br>\ni make making submission, let's wait a little longer for conformer generation of test molecules to complete…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b4f81639e062ecaad8a646eed5c67a%2FSelection_137.png?generation=1716346013257540&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b89ac916f89788115e16c2b82104d90%2FSelection_138.png?generation=1716346035358504&amp;alt=media\"></p>\n<hr>\n<p>plan:</p>\n<ul>\n<li>release conformer dataset (part of train + test)</li>\n<li>release illustrative training code (let kaggler check for bugs if any) </li>\n</ul>",
      "rawMarkdown": "first 3d gnn sphere net result!!!\n\n**UPDATE: there is a bug in my code. I am only training with my validation set of 400k.**\n(i will update the results later)\n\ni hope there is no bugs in my code. results are too good (though not unreasonable).\ni make making submission, let's wait a little longer for conformer generation of test molecules to complete...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b4f81639e062ecaad8a646eed5c67a%2FSelection_137.png?generation=1716346013257540&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b89ac916f89788115e16c2b82104d90%2FSelection_138.png?generation=1716346035358504&alt=media)\n\n---\nplan:\n- release conformer dataset (part of train + test)\n- release illustrative training code (let kaggler check for bugs if any)",
      "votes": null
    },
    {
      "id": "2828658",
      "postDate": "05/22/2024 07:48:27",
      "content": "<p>wow! thanks so much for sharing. </p>",
      "rawMarkdown": "wow! thanks so much for sharing.",
      "votes": null
    },
    {
      "id": "2828716",
      "postDate": "05/22/2024 08:27:21",
      "content": "<p>made a submission</p>\n<ul>\n<li>lb0.287 for 1 M training sample<br>\n(for invalid conformers (2640 out of 1674896) and nonshare blocks, i use values from another submission).</li>\n</ul>\n<p>The values is much lower than expected. Then I discover something interesting:</p>\n<pre><code> .\n\n . \n\n . (fraction of invalid conformers)\n</code></pre>\n<ol>\n<li>my fraction of invalid conformers in the report is wrong.  </li>\n<li>invalid conformers have unusual high fraction of BRD4 binds. This could be why my validation is so good. </li>\n</ol>\n<p>this is only one error type in formation generation:<br>\nConfGen.ReturnCode.FORCEFIELD_SETUP_FAILED: 'force field setup failed',</p>\n<p>smiles of invalid conformers:</p>\n<pre><code>b\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\n</code></pre>",
      "rawMarkdown": "made a submission\n- lb0.287 for 1 M training sample\n(for invalid conformers (2640 out of 1674896) and nonshare blocks, i use values from another submission).\n\nThe values is much lower than expected. Then I discover something interesting:\n\n```\nnonshare_status_df 0.0\n[nan nan nan]\n\nvalid_status_df 0.0026325 \n[0.01899335 0.00189934 0.00094967]\n\ntrain_status_df 0.0027133216695826045 (fraction of invalid conformers)\n[0.01473839 0.00073692 0.00036846] (fraction of positive in invalid conformers)\n['BRD4', 'HSA', 'sEH']\n\n```\n1. my fraction of invalid conformers in the report is wrong.  \n2. invalid conformers have unusual high fraction of BRD4 binds. This could be why my validation is so good. \n\nthis is only one error type in formation generation:\nConfGen.ReturnCode.FORCEFIELD_SETUP_FAILED: 'force field setup failed',\n\nsmiles of invalid conformers:\n```\n\n939 b'C#CCOc1cccc(CNc2nc(NCC3CCC(F)(F)CC3)nc(N[C@@H](CC#C)CC(=O)NC)n2)c1'\n942 b'C#CCOc1cccc(CNc2nc(NCC3CCCn4ccnc43)nc(N[C@@H](CC#C)CC(=O)NC)n2)c1'\n1320 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2(OC)CCC2)nc(Nc2ccc(C#C)cc2)n1'\n1673 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2(O)CC2)nc(Nc2ccc(C#C)cc2)n1'\n2596 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCNC(=O)C(=C)C)nc(Nc2cc(-c3ccc(Cl)cc3)sc2C(=O)OC)n1'\n2929 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCOCC(=C)C)nc(Nc2cc(F)cc(C(=O)OC)c2)n1'\n4060 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(=C)Cl)nc(Nc2c(O)ncnc2O)n1'\n4513 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2CCC(=C)CC2)nc(NCC2CCOCC23CCCC3)n1'\n4927 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(=O)NCC=C)nc(NCc2ccc(Oc3ccc(Cl)cc3Cl)c(C)c2)n1'\n5385 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(C)OCC=C)nc(NCC(c2cccs2)N2CCOCC2)n1'\n5744 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(Nc2ccc(C#N)c(C(F)(F)F)c2)n1'\n5759 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(NCC2(c3ccc(Cl)cc3Cl)CCCC2)n1'\n5820 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(Nc2cc(Cl)c([N+](=O)[O-])cn2)n1'\n6037 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCOCC=C)nc(NCc2cscc2C)n1'\n6318 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSCC=C)nc(Nc2c(C(=O)OC)c[nH]c2C(=O)OC)n1'\n6408 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSCC=C)nc(Nc2cc(C)n(C)n2)n1'\n7156 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCc2cn(C)c(=O)[nH]c2=O)nc(Nc2cccc(NC(C)=O)n2)n1'\n7853 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSC(C)=O)nc(NCC2CCCOC2)n1'\n```",
      "votes": null
    },
    {
      "id": "2829046",
      "postDate": "05/22/2024 11:37:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7ae809ea1080db7dc95429e3072c09d%2FSelection_139.png?generation=1716377800939366&amp;alt=media\"></p>\n<p>conformer data has been updated:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7ae809ea1080db7dc95429e3072c09d%2FSelection_139.png?generation=1716377800939366&alt=media)\n\nconformer data has been updated:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset",
      "votes": null
    },
    {
      "id": "2829783",
      "postDate": "05/22/2024 19:34:37",
      "content": "<p>I am very interested in your research. Thank you for sharing the dataset and code!</p>",
      "rawMarkdown": "I am very interested in your research. Thank you for sharing the dataset and code!",
      "votes": null
    },
    {
      "id": "2830729",
      "postDate": "05/23/2024 10:38:27",
      "content": "<p><a href=\"https://arxiv.org/abs/2206.08515\" target=\"_blank\">https://arxiv.org/abs/2206.08515</a><br>\n[paper] ComENet: Towards Complete and Efficient Message Passing for 3D Molecular Graphs</p>\n<p>This is a much faster version of SphereNet.<br>\nBoth ComENet and SphereNet implemented in the same DIG package that.<br>\nThey use the same input. Hence you can use back your SphereNet training code.</p>\n<p>performance is slight worse than SphereNet but you get 2x speed improvement and 2/3 memory reduction.<br>\nI was able to use batchsize=3k for ComENet.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F413807324f8844968b401d7b00ec0458%2FSelection_141.png?generation=1716460694236012&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F08c77382b1ae9ae5edbd1cd95851dfee%2FSelection_140.png?generation=1716460705895330&amp;alt=media\"></p>",
      "rawMarkdown": "https://arxiv.org/abs/2206.08515\n[paper] ComENet: Towards Complete and Efficient Message Passing for 3D Molecular Graphs\n\nThis is a much faster version of SphereNet.\nBoth ComENet and SphereNet implemented in the same DIG package that.\nThey use the same input. Hence you can use back your SphereNet training code.\n\nperformance is slight worse than SphereNet but you get 2x speed improvement and 2/3 memory reduction.\nI was able to use batchsize=3k for ComENet.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F413807324f8844968b401d7b00ec0458%2FSelection_141.png?generation=1716460694236012&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F08c77382b1ae9ae5edbd1cd95851dfee%2FSelection_140.png?generation=1716460705895330&alt=media)",
      "votes": null
    },
    {
      "id": "2830734",
      "postDate": "05/23/2024 10:44:27",
      "content": "<p>if find the BUG:<br>\n<strong>UPDATE: there is a bug in my code. I am only training with my validation set of 400k.</strong><br>\n(i will update the results later)</p>",
      "rawMarkdown": "if find the BUG:\n**UPDATE: there is a bug in my code. I am only training with my validation set of 400k.**\n(i will update the results later)",
      "votes": null
    },
    {
      "id": "2833510",
      "postDate": "05/24/2024 09:11:49",
      "content": "<p>final corrected results:<br>\nsubmission:<br>\ntrain=10million conformer: lb=0.383<br>\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb4f841f4e4ff39368aacaf4180848ba%2FSelection_143.png?generation=1716541869990366&amp;alt=media\"></p>\n<p>note: this is using comenet (a faster but slightly less accurate version of sphere net)</p>",
      "rawMarkdown": "final corrected results:\nsubmission:\ntrain=10million conformer: lb=0.383\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb4f841f4e4ff39368aacaf4180848ba%2FSelection_143.png?generation=1716541869990366&alt=media)\n\nnote: this is using comenet (a faster but slightly less accurate version of sphere net)",
      "votes": null
    },
    {
      "id": "2835481",
      "postDate": "05/25/2024 10:40:58",
      "content": "<p>i have put up about 4.5 million conforge conformers (lowest energy) at my dataset link.<br>\nmaybe some kagglers can compare results of other conformer generators ( e.g. rdkit, or faster sckit-fingerprints)</p>\n<p>my results with spherenet and comenet show that the quality of conforge is good.<br>\nother kagglers can do experiments on 3d fingerfprint, or molecule transformer/graph that uses xyz information.<br>\nif you do, please post results here.</p>\n<p>if there are good results, i can put more conformers at the database. i current have about 30million.</p>",
      "rawMarkdown": "i have put up about 4.5 million conforge conformers (lowest energy) at my dataset link.\nmaybe some kagglers can compare results of other conformer generators ( e.g. rdkit, or faster sckit-fingerprints)\n\nmy results with spherenet and comenet show that the quality of conforge is good.\nother kagglers can do experiments on 3d fingerfprint, or molecule transformer/graph that uses xyz information.\nif you do, please post results here.\n\nif there are good results, i can put more conformers at the database. i current have about 30million.",
      "votes": null
    },
    {
      "id": "2836505",
      "postDate": "05/26/2024 01:15:31",
      "content": "<p>a good lib to extract 3d atom/bond feature using rdkit:<br>\n<a href=\"https://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py\" target=\"_blank\">https://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py</a> </p>\n<p>see Compound3DKit and CompoundKit<br>\nrelated paper:ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction<br>\n<a href=\"https://arxiv.org/pdf/2106.06130\" target=\"_blank\">https://arxiv.org/pdf/2106.06130</a></p>",
      "rawMarkdown": "a good lib to extract 3d atom/bond feature using rdkit:\nhttps://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py \n\nsee Compound3DKit and CompoundKit\nrelated paper:ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction\nhttps://arxiv.org/pdf/2106.06130",
      "votes": null
    },
    {
      "id": "2840993",
      "postDate": "05/28/2024 10:53:24",
      "content": "<p>Is there public dataset for the 3d information? </p>",
      "rawMarkdown": "Is there public dataset for the 3d information?",
      "votes": null
    },
    {
      "id": "2841720",
      "postDate": "05/28/2024 17:07:53",
      "content": "<p>Maybe you should extract 3d feature by yourself</p>",
      "rawMarkdown": "Maybe you should extract 3d feature by yourself",
      "votes": null
    },
    {
      "id": "2842377",
      "postDate": "05/29/2024 03:55:45",
      "content": "<p>I do not have extra machine for it 🤔 … I had a simple but very fast structure for these 3d molecular task. For better performance, 3d information is necessary. </p>\n<p>But I do not think we can have accurate 3d information.</p>",
      "rawMarkdown": "I do not have extra machine for it 🤔 ... I had a simple but very fast structure for these 3d molecular task. For better performance, 3d information is necessary. \n\nBut I do not think we can have accurate 3d information.",
      "votes": null
    },
    {
      "id": "2842508",
      "postDate": "05/29/2024 05:52:16",
      "content": "<p>i have put 10 millions conformers in my dataset link for toy experiments if anyone is interested in. it should be enough for comparsion experiments.</p>",
      "rawMarkdown": "i have put 10 millions conformers in my dataset link for toy experiments if anyone is interested in. it should be enough for comparsion experiments.",
      "votes": null
    },
    {
      "id": "2842512",
      "postDate": "05/29/2024 05:56:47",
      "content": "<p>\"But I do not think we can have accurate 3d information.\"</p>\n<pre><code>       :\n             = Chem.AddHs(mol)\n            res = AllChem.EmbedMultipleConfs(, numConfs=numConfs)\n            \n            res = AllChem.MMFFOptimizeMoleculeConfs()\n             = Chem.RemoveHs()\n            index = np.argmin([x[]  x  res])\n            energy = res[index][]\n            conf = .GetConformer(id=int(index))\n        except:\n             = mol\n            AllChem.Compute2DCoords()\n            energy = \n            conf = .GetConformer()\n</code></pre>\n<p>creating ensemble of coformers should help.</p>",
      "rawMarkdown": "\"But I do not think we can have accurate 3d information.\"\n\n```\n\n       try:\n            new_mol = Chem.AddHs(mol)\n            res = AllChem.EmbedMultipleConfs(new_mol, numConfs=numConfs)\n            ### MMFF generates multiple conformations\n            res = AllChem.MMFFOptimizeMoleculeConfs(new_mol)\n            new_mol = Chem.RemoveHs(new_mol)\n            index = np.argmin([x[1] for x in res])\n            energy = res[index][1]\n            conf = new_mol.GetConformer(id=int(index))\n        except:\n            new_mol = mol\n            AllChem.Compute2DCoords(new_mol)\n            energy = 0\n            conf = new_mol.GetConformer()\n```\n\ncreating ensemble of coformers should help.",
      "votes": null
    },
    {
      "id": "2846825",
      "postDate": "05/31/2024 09:22:11",
      "content": "<p>How did you train dealing with the unbalanced classes, with stratified K fold cross validation? It might be beneficial to generate many conformers only for the binders to make predictions for these more robust</p>",
      "rawMarkdown": "How did you train dealing with the unbalanced classes, with stratified K fold cross validation? It might be beneficial to generate many conformers only for the binders to make predictions for these more robust",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2827142,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/21/2024 09:07:38",
      "content": "<p>spherical CNN has been introduced by the host. Here is a quick visual to explain what it is<br>\n<a href=\"https://research.google/blog/scalable-spherical-cnns-for-scientific-applications/\" target=\"_blank\">https://research.google/blog/scalable-spherical-cnns-for-scientific-applications/</a><br>\ncode: <a href=\"https://github.com/google-research/spherical-cnn/\" target=\"_blank\">https://github.com/google-research/spherical-cnn/</a><br>\n(maybe someone can try this jax code on kaggle tpu)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F45020dd36f6f830bddf5f59c12677a39%2Fimage6.gif?generation=1716282431762462&amp;alt=media\"></p>\n<p>Each atom is represented by a set of spherical signals accumulating physical interactions with other atoms of each type (shown in the three panels on the right). For example, the oxygen atom (O; top panel) has a channel for oxygen (indicated by the sphere labeled “O” on the left) and hydrogen (“H”, right). The accumulated Coulomb forces on the oxygen atom with respect to the two hydrogen atoms is indicated by the red shaded regions on the bottom of the sphere labeled “H”. Because the oxygen atom contributes no forces to itself, the “O” sphere is uniform. We include extra channels for the Van der Waals forces.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2827150,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/21/2024 09:10:20",
          "content": "<p>this is somehow smiliar to sphere-net (SMP)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea38e45ffbcd1a5dff71c224c5fb2eb5%2FSelection_133.png?generation=1716282593210660&amp;alt=media\"></p>\n<p>it wold be interesting to see comparison of the two</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2828291,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/22/2024 02:49:49",
      "content": "<p>first 3d gnn sphere net result!!!</p>\n<p><strong>UPDATE: there is a bug in my code. I am only training with my validation set of 400k.</strong><br>\n(i will update the results later)</p>\n<p>i hope there is no bugs in my code. results are too good (though not unreasonable).<br>\ni make making submission, let's wait a little longer for conformer generation of test molecules to complete…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b4f81639e062ecaad8a646eed5c67a%2FSelection_137.png?generation=1716346013257540&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b89ac916f89788115e16c2b82104d90%2FSelection_138.png?generation=1716346035358504&amp;alt=media\"></p>\n<hr>\n<p>plan:</p>\n<ul>\n<li>release conformer dataset (part of train + test)</li>\n<li>release illustrative training code (let kaggler check for bugs if any) </li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2828658,
          "author_name": "thedrcat",
          "author_url": "",
          "post_date": "05/22/2024 07:48:27",
          "content": "<p>wow! thanks so much for sharing. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2828716,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/22/2024 08:27:21",
              "content": "<p>made a submission</p>\n<ul>\n<li>lb0.287 for 1 M training sample<br>\n(for invalid conformers (2640 out of 1674896) and nonshare blocks, i use values from another submission).</li>\n</ul>\n<p>The values is much lower than expected. Then I discover something interesting:</p>\n<pre><code> .\n\n . \n\n . (fraction of invalid conformers)\n</code></pre>\n<ol>\n<li>my fraction of invalid conformers in the report is wrong.  </li>\n<li>invalid conformers have unusual high fraction of BRD4 binds. This could be why my validation is so good. </li>\n</ol>\n<p>this is only one error type in formation generation:<br>\nConfGen.ReturnCode.FORCEFIELD_SETUP_FAILED: 'force field setup failed',</p>\n<p>smiles of invalid conformers:</p>\n<pre><code>b\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\nb\n</code></pre>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2830734,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/23/2024 10:44:27",
              "content": "<p>if find the BUG:<br>\n<strong>UPDATE: there is a bug in my code. I am only training with my validation set of 400k.</strong><br>\n(i will update the results later)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2829046,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/22/2024 11:37:09",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7ae809ea1080db7dc95429e3072c09d%2FSelection_139.png?generation=1716377800939366&amp;alt=media\"></p>\n<p>conformer data has been updated:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2829783,
      "author_name": "shinaggle",
      "author_url": "",
      "post_date": "05/22/2024 19:34:37",
      "content": "<p>I am very interested in your research. Thank you for sharing the dataset and code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2830729,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/23/2024 10:38:27",
      "content": "<p><a href=\"https://arxiv.org/abs/2206.08515\" target=\"_blank\">https://arxiv.org/abs/2206.08515</a><br>\n[paper] ComENet: Towards Complete and Efficient Message Passing for 3D Molecular Graphs</p>\n<p>This is a much faster version of SphereNet.<br>\nBoth ComENet and SphereNet implemented in the same DIG package that.<br>\nThey use the same input. Hence you can use back your SphereNet training code.</p>\n<p>performance is slight worse than SphereNet but you get 2x speed improvement and 2/3 memory reduction.<br>\nI was able to use batchsize=3k for ComENet.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F413807324f8844968b401d7b00ec0458%2FSelection_141.png?generation=1716460694236012&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F08c77382b1ae9ae5edbd1cd95851dfee%2FSelection_140.png?generation=1716460705895330&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2833510,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/24/2024 09:11:49",
      "content": "<p>final corrected results:<br>\nsubmission:<br>\ntrain=10million conformer: lb=0.383<br>\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb4f841f4e4ff39368aacaf4180848ba%2FSelection_143.png?generation=1716541869990366&amp;alt=media\"></p>\n<p>note: this is using comenet (a faster but slightly less accurate version of sphere net)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2835481,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/25/2024 10:40:58",
      "content": "<p>i have put up about 4.5 million conforge conformers (lowest energy) at my dataset link.<br>\nmaybe some kagglers can compare results of other conformer generators ( e.g. rdkit, or faster sckit-fingerprints)</p>\n<p>my results with spherenet and comenet show that the quality of conforge is good.<br>\nother kagglers can do experiments on 3d fingerfprint, or molecule transformer/graph that uses xyz information.<br>\nif you do, please post results here.</p>\n<p>if there are good results, i can put more conformers at the database. i current have about 30million.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2836505,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/26/2024 01:15:31",
      "content": "<p>a good lib to extract 3d atom/bond feature using rdkit:<br>\n<a href=\"https://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py\" target=\"_blank\">https://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py</a> </p>\n<p>see Compound3DKit and CompoundKit<br>\nrelated paper:ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction<br>\n<a href=\"https://arxiv.org/pdf/2106.06130\" target=\"_blank\">https://arxiv.org/pdf/2106.06130</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2840993,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "05/28/2024 10:53:24",
      "content": "<p>Is there public dataset for the 3d information? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2841720,
          "author_name": "lblhandsome",
          "author_url": "",
          "post_date": "05/28/2024 17:07:53",
          "content": "<p>Maybe you should extract 3d feature by yourself</p>",
          "votes": null,
          "replies": [
            {
              "id": 2842377,
              "author_name": "yuanzhezhou",
              "author_url": "",
              "post_date": "05/29/2024 03:55:45",
              "content": "<p>I do not have extra machine for it 🤔 … I had a simple but very fast structure for these 3d molecular task. For better performance, 3d information is necessary. </p>\n<p>But I do not think we can have accurate 3d information.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2842508,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "05/29/2024 05:52:16",
                  "content": "<p>i have put 10 millions conformers in my dataset link for toy experiments if anyone is interested in. it should be enough for comparsion experiments.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2842512,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "05/29/2024 05:56:47",
                      "content": "<p>\"But I do not think we can have accurate 3d information.\"</p>\n<pre><code>       :\n             = Chem.AddHs(mol)\n            res = AllChem.EmbedMultipleConfs(, numConfs=numConfs)\n            \n            res = AllChem.MMFFOptimizeMoleculeConfs()\n             = Chem.RemoveHs()\n            index = np.argmin([x[]  x  res])\n            energy = res[index][]\n            conf = .GetConformer(id=int(index))\n        except:\n             = mol\n            AllChem.Compute2DCoords()\n            energy = \n            conf = .GetConformer()\n</code></pre>\n<p>creating ensemble of coformers should help.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2846825,
      "author_name": "rubenkaggle",
      "author_url": "",
      "post_date": "05/31/2024 09:22:11",
      "content": "<p>How did you train dealing with the unbalanced classes, with stratified K fold cross validation? It might be beneficial to generate many conformers only for the binders to make predictions for these more robust</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2824788": "This is work in progress.\nMy first target is SMP (spherical message passing) sphere-net.\nif you have other good 3d graph-NN, please advised here! Thanks.\n\n[paper] Spherical Message Passing for 3D Graph Networks - Yi Liu, ICLR 2022 \nhttps://paperswithcode.com/paper/spherical-message-passing-for-3d-graph\n\n\nfirst:\n- fast conformer generation, with sphere-net example\nhttps://www.kaggle.com/code/hengck23/conforge-open-source-conformer-generator\n\nsome generated conformers (**10 million**) found at:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\nsubmission:\ntrain=10million conformer: lb=0.383\n\nfor the below: did not retrain just add new data and finetune. retrain should give slightly better results ...\ntrain=20million : lb=0.387\ntrain=40million : lb=0.420\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F983685dd4eb3f2a9bc2cbc1fcc9a14a8%2FSelection_132.png?generation=1716176444335864&alt=media)\n\nresults look good:\n[paper] 'High-Quality Conformer Generation with CONFORGE: Algorithm and Performance Assessment' - T Seidel , acs 2023\nhttps://pubs.acs.org/doi/10.1021/acs.jcim.3c00563\nhttps://github.com/molinfo-vienna/CDPKit\n\nwith parallel processing, I was able to generate 10 million conformers with 4 hr",
    "2827142": "spherical CNN has been introduced by the host. Here is a quick visual to explain what it is\nhttps://research.google/blog/scalable-spherical-cnns-for-scientific-applications/\ncode: https://github.com/google-research/spherical-cnn/\n(maybe someone can try this jax code on kaggle tpu)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F45020dd36f6f830bddf5f59c12677a39%2Fimage6.gif?generation=1716282431762462&alt=media)\n\nEach atom is represented by a set of spherical signals accumulating physical interactions with other atoms of each type (shown in the three panels on the right). For example, the oxygen atom (O; top panel) has a channel for oxygen (indicated by the sphere labeled “O” on the left) and hydrogen (“H”, right). The accumulated Coulomb forces on the oxygen atom with respect to the two hydrogen atoms is indicated by the red shaded regions on the bottom of the sphere labeled “H”. Because the oxygen atom contributes no forces to itself, the “O” sphere is uniform. We include extra channels for the Van der Waals forces.",
    "2827150": "this is somehow smiliar to sphere-net (SMP)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fea38e45ffbcd1a5dff71c224c5fb2eb5%2FSelection_133.png?generation=1716282593210660&alt=media)\n\nit wold be interesting to see comparison of the two",
    "2828291": "first 3d gnn sphere net result!!!\n\n**UPDATE: there is a bug in my code. I am only training with my validation set of 400k.**\n(i will update the results later)\n\ni hope there is no bugs in my code. results are too good (though not unreasonable).\ni make making submission, let's wait a little longer for conformer generation of test molecules to complete...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe7b4f81639e062ecaad8a646eed5c67a%2FSelection_137.png?generation=1716346013257540&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5b89ac916f89788115e16c2b82104d90%2FSelection_138.png?generation=1716346035358504&alt=media)\n\n---\nplan:\n- release conformer dataset (part of train + test)\n- release illustrative training code (let kaggler check for bugs if any)",
    "2828658": "wow! thanks so much for sharing.",
    "2828716": "made a submission\n- lb0.287 for 1 M training sample\n(for invalid conformers (2640 out of 1674896) and nonshare blocks, i use values from another submission).\n\nThe values is much lower than expected. Then I discover something interesting:\n\n```\nnonshare_status_df 0.0\n[nan nan nan]\n\nvalid_status_df 0.0026325 \n[0.01899335 0.00189934 0.00094967]\n\ntrain_status_df 0.0027133216695826045 (fraction of invalid conformers)\n[0.01473839 0.00073692 0.00036846] (fraction of positive in invalid conformers)\n['BRD4', 'HSA', 'sEH']\n\n```\n1. my fraction of invalid conformers in the report is wrong.  \n2. invalid conformers have unusual high fraction of BRD4 binds. This could be why my validation is so good. \n\nthis is only one error type in formation generation:\nConfGen.ReturnCode.FORCEFIELD_SETUP_FAILED: 'force field setup failed',\n\nsmiles of invalid conformers:\n```\n\n939 b'C#CCOc1cccc(CNc2nc(NCC3CCC(F)(F)CC3)nc(N[C@@H](CC#C)CC(=O)NC)n2)c1'\n942 b'C#CCOc1cccc(CNc2nc(NCC3CCCn4ccnc43)nc(N[C@@H](CC#C)CC(=O)NC)n2)c1'\n1320 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2(OC)CCC2)nc(Nc2ccc(C#C)cc2)n1'\n1673 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2(O)CC2)nc(Nc2ccc(C#C)cc2)n1'\n2596 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCNC(=O)C(=C)C)nc(Nc2cc(-c3ccc(Cl)cc3)sc2C(=O)OC)n1'\n2929 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCOCC(=C)C)nc(Nc2cc(F)cc(C(=O)OC)c2)n1'\n4060 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(=C)Cl)nc(Nc2c(O)ncnc2O)n1'\n4513 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC2CCC(=C)CC2)nc(NCC2CCOCC23CCCC3)n1'\n4927 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(=O)NCC=C)nc(NCc2ccc(Oc3ccc(Cl)cc3Cl)c(C)c2)n1'\n5385 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCC(C)OCC=C)nc(NCC(c2cccs2)N2CCOCC2)n1'\n5744 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(Nc2ccc(C#N)c(C(F)(F)F)c2)n1'\n5759 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(NCC2(c3ccc(Cl)cc3Cl)CCCC2)n1'\n5820 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCCOCC=C)nc(Nc2cc(Cl)c([N+](=O)[O-])cn2)n1'\n6037 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCOCC=C)nc(NCc2cscc2C)n1'\n6318 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSCC=C)nc(Nc2c(C(=O)OC)c[nH]c2C(=O)OC)n1'\n6408 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSCC=C)nc(Nc2cc(C)n(C)n2)n1'\n7156 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCc2cn(C)c(=O)[nH]c2=O)nc(Nc2cccc(NC(C)=O)n2)n1'\n7853 b'C#CC[C@@H](CC(=O)NC)Nc1nc(NCCSC(C)=O)nc(NCC2CCCOC2)n1'\n```",
    "2829046": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7ae809ea1080db7dc95429e3072c09d%2FSelection_139.png?generation=1716377800939366&alt=media)\n\nconformer data has been updated:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset",
    "2829783": "I am very interested in your research. Thank you for sharing the dataset and code!",
    "2830729": "https://arxiv.org/abs/2206.08515\n[paper] ComENet: Towards Complete and Efficient Message Passing for 3D Molecular Graphs\n\nThis is a much faster version of SphereNet.\nBoth ComENet and SphereNet implemented in the same DIG package that.\nThey use the same input. Hence you can use back your SphereNet training code.\n\nperformance is slight worse than SphereNet but you get 2x speed improvement and 2/3 memory reduction.\nI was able to use batchsize=3k for ComENet.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F413807324f8844968b401d7b00ec0458%2FSelection_141.png?generation=1716460694236012&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F08c77382b1ae9ae5edbd1cd95851dfee%2FSelection_140.png?generation=1716460705895330&alt=media)",
    "2830734": "if find the BUG:\n**UPDATE: there is a bug in my code. I am only training with my validation set of 400k.**\n(i will update the results later)",
    "2833510": "final corrected results:\nsubmission:\ntrain=10million conformer: lb=0.383\ncode and logfile: refer to folder \"conformer-3d-gnn-example-code\" at the dataset link above\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcb4f841f4e4ff39368aacaf4180848ba%2FSelection_143.png?generation=1716541869990366&alt=media)\n\nnote: this is using comenet (a faster but slightly less accurate version of sphere net)",
    "2835481": "i have put up about 4.5 million conforge conformers (lowest energy) at my dataset link.\nmaybe some kagglers can compare results of other conformer generators ( e.g. rdkit, or faster sckit-fingerprints)\n\nmy results with spherenet and comenet show that the quality of conforge is good.\nother kagglers can do experiments on 3d fingerfprint, or molecule transformer/graph that uses xyz information.\nif you do, please post results here.\n\nif there are good results, i can put more conformers at the database. i current have about 30million.",
    "2836505": "a good lib to extract 3d atom/bond feature using rdkit:\nhttps://github.com/PaddlePaddle/PaddleHelix/blob/dev/pahelix/utils/compound_tools.py \n\nsee Compound3DKit and CompoundKit\nrelated paper:ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction\nhttps://arxiv.org/pdf/2106.06130",
    "2840993": "Is there public dataset for the 3d information?",
    "2841720": "Maybe you should extract 3d feature by yourself",
    "2842377": "I do not have extra machine for it 🤔 ... I had a simple but very fast structure for these 3d molecular task. For better performance, 3d information is necessary. \n\nBut I do not think we can have accurate 3d information.",
    "2842508": "i have put 10 millions conformers in my dataset link for toy experiments if anyone is interested in. it should be enough for comparsion experiments.",
    "2842512": "\"But I do not think we can have accurate 3d information.\"\n\n```\n\n       try:\n            new_mol = Chem.AddHs(mol)\n            res = AllChem.EmbedMultipleConfs(new_mol, numConfs=numConfs)\n            ### MMFF generates multiple conformations\n            res = AllChem.MMFFOptimizeMoleculeConfs(new_mol)\n            new_mol = Chem.RemoveHs(new_mol)\n            index = np.argmin([x[1] for x in res])\n            energy = res[index][1]\n            conf = new_mol.GetConformer(id=int(index))\n        except:\n            new_mol = mol\n            AllChem.Compute2DCoords(new_mol)\n            energy = 0\n            conf = new_mol.GetConformer()\n```\n\ncreating ensemble of coformers should help.",
    "2846825": "How did you train dealing with the unbalanced classes, with stratified K fold cross validation? It might be beneficial to generate many conformers only for the binders to make predictions for these more robust"
  },
  "source": "meta"
}