{
  "id": 518956,
  "title": "0.480 public LB solutions! (ensemble of 7 fingerprint/SMILES/atom-level features models)",
  "url": "/competitions/leash-BELKA/writeups/np-p-0-480-public-lb-solutions-ensemble-of-7-finge",
  "author_name": "",
  "post_date": "2024-07-10T06:23:23.287Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>All codes are released in <a href=\"https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB\" target=\"_blank\">https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB</a>.</p>\n<h2>USAGE</h2>\n<pre><code>python train_BELKA.py --modelType FprMLP --EMA             \npython train_BELKA.py --modelType DeepFM/DeepFM2 --EMA     \npython train_BELKA.py --modelType PseLabAttn --EMA         \npython train_BELKA.py --modelType GraphMLP/GraphMLP2 --EMA \npython train_BELKA_lgb.py                                      \n</code></pre>\n<p>FprMLP/DeepFM/DeepFM2 are all based on the molecular fingerprint features only, and achieve <strong>0.620~0.645</strong> in validation(15-fold) and <strong>0.432</strong> in public LB;</p>\n<p>PseLabAttn/GraphMLP are feature-mixture model, achieving <strong>0.650</strong> in validation(15-fold) and <strong>0.458/0.398</strong> in public LB; (this GNN didn't consider the bond type)</p>\n<p>lgb is fingerprint-based lightgbm model, achieving <strong>0.615</strong> in validation(15-fold) and <strong>0.377</strong> in public LB;</p>\n<p>ensemble all of them can lead to about <strong>0.480</strong> in public LB. </p>\n<h2>CONCLUSION</h2>\n<p>For public LB, the <strong>fingerprint-based model or atom/bond feature-based GNN</strong> will achieve great performance. (bond features seem very important in public LB, my GNN didn't consider the bond features so it's only 0.398 in public LB)</p>\n<p>For private LB, the <strong>SMILES-basd model</strong> will achieve great performance. </p>",
  "messages": [
    {
      "id": "2912710",
      "postDate": "07/09/2024 03:39:03",
      "content": "<p>All codes are released in <a href=\"https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB\" target=\"_blank\">https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB</a>.</p>\n<h2>USAGE</h2>\n<pre><code>python train_BELKA.py --modelType FprMLP --EMA             \npython train_BELKA.py --modelType DeepFM/DeepFM2 --EMA     \npython train_BELKA.py --modelType PseLabAttn --EMA         \npython train_BELKA.py --modelType GraphMLP/GraphMLP2 --EMA \npython train_BELKA_lgb.py                                      \n</code></pre>\n<p>FprMLP/DeepFM/DeepFM2 are all based on the molecular fingerprint features only, and achieve <strong>0.620~0.645</strong> in validation(15-fold) and <strong>0.432</strong> in public LB;</p>\n<p>PseLabAttn/GraphMLP are feature-mixture model, achieving <strong>0.650</strong> in validation(15-fold) and <strong>0.458/0.398</strong> in public LB; (this GNN didn't consider the bond type)</p>\n<p>lgb is fingerprint-based lightgbm model, achieving <strong>0.615</strong> in validation(15-fold) and <strong>0.377</strong> in public LB;</p>\n<p>ensemble all of them can lead to about <strong>0.480</strong> in public LB. </p>\n<h2>CONCLUSION</h2>\n<p>For public LB, the <strong>fingerprint-based model or atom/bond feature-based GNN</strong> will achieve great performance. (bond features seem very important in public LB, my GNN didn't consider the bond features so it's only 0.398 in public LB)</p>\n<p>For private LB, the <strong>SMILES-basd model</strong> will achieve great performance. </p>",
      "rawMarkdown": "All codes are released in https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB.\n\n## USAGE\n\n```python\npython train_BELKA.py --modelType FprMLP --EMA True            # for 7 fingerprint-based MLP model\npython train_BELKA.py --modelType DeepFM/DeepFM2 --EMA True    # for 4 fingerprint-based DeepFM model\npython train_BELKA.py --modelType PseLabAttn --EMA True        # for SMILES/ECFP/atom features-based RNN-Transformer model\npython train_BELKA.py --modelType GraphMLP/GraphMLP2 --EMA True# for SMILES/FCFP/atom features-based GNN model\npython train_BELKA_lgb.py                                      # for 7 fingerprint-based lgb model\n```\n\nFprMLP/DeepFM/DeepFM2 are all based on the molecular fingerprint features only, and achieve **0.620~0.645** in validation(15-fold) and **0.432** in public LB;\n\nPseLabAttn/GraphMLP are feature-mixture model, achieving **0.650** in validation(15-fold) and **0.458/0.398** in public LB; (this GNN didn't consider the bond type)\n\nlgb is fingerprint-based lightgbm model, achieving **0.615** in validation(15-fold) and **0.377** in public LB;\n\nensemble all of them can lead to about **0.480** in public LB. \n\n## CONCLUSION\n\nFor public LB, the **fingerprint-based model or atom/bond feature-based GNN** will achieve great performance. (bond features seem very important in public LB, my GNN didn't consider the bond features so it's only 0.398 in public LB)\n\nFor private LB, the **SMILES-basd model** will achieve great performance.",
      "votes": null
    },
    {
      "id": "2914262",
      "postDate": "07/09/2024 19:49:55",
      "content": "<p>Amazing solutions! Thanks for sharing!</p>",
      "rawMarkdown": "Amazing solutions! Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2914262,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "07/09/2024 19:49:55",
      "content": "<p>Amazing solutions! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2912710": "All codes are released in https://github.com/wudejian789/NeurIPS2024-BELKA-0.480-in-public-LB.\n\n## USAGE\n\n```python\npython train_BELKA.py --modelType FprMLP --EMA True            # for 7 fingerprint-based MLP model\npython train_BELKA.py --modelType DeepFM/DeepFM2 --EMA True    # for 4 fingerprint-based DeepFM model\npython train_BELKA.py --modelType PseLabAttn --EMA True        # for SMILES/ECFP/atom features-based RNN-Transformer model\npython train_BELKA.py --modelType GraphMLP/GraphMLP2 --EMA True# for SMILES/FCFP/atom features-based GNN model\npython train_BELKA_lgb.py                                      # for 7 fingerprint-based lgb model\n```\n\nFprMLP/DeepFM/DeepFM2 are all based on the molecular fingerprint features only, and achieve **0.620~0.645** in validation(15-fold) and **0.432** in public LB;\n\nPseLabAttn/GraphMLP are feature-mixture model, achieving **0.650** in validation(15-fold) and **0.458/0.398** in public LB; (this GNN didn't consider the bond type)\n\nlgb is fingerprint-based lightgbm model, achieving **0.615** in validation(15-fold) and **0.377** in public LB;\n\nensemble all of them can lead to about **0.480** in public LB. \n\n## CONCLUSION\n\nFor public LB, the **fingerprint-based model or atom/bond feature-based GNN** will achieve great performance. (bond features seem very important in public LB, my GNN didn't consider the bond features so it's only 0.398 in public LB)\n\nFor private LB, the **SMILES-basd model** will achieve great performance.",
    "2914262": "Amazing solutions! Thanks for sharing!"
  },
  "source": "meta"
}