{
  "id": 518936,
  "title": "Closing Thoughts and Future Directions",
  "url": "/competitions/leash-BELKA/discussion/518936",
  "author_name": "Andrew D. Blevins",
  "post_date": "2024-07-09T00:09:56.932000",
  "votes": 42,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Dear BELKA competitors,<br>\nThank you for your participation in this challenge! We hope it has been an interesting and rewarding experience. Your efforts have provided valuable insights into the complexities of ML-driven drug discovery.</p>\n<h3>Competition Overview</h3>\n<p>The BELKA competition aimed to advance ML in drug discovery by providing a dataset of ~300M protein/molecule interactions as training material for modeling. Our goal was to encourage the development of models that could generalize across chemical space (which is quite vast!) and help discover new life-saving medicines. To evaluate generalization, the private test set contained a large number of molecules from a different corner of chemical space provided in the training and public test sets. </p>\n<h3>Leaderboard shakeup</h3>\n<p>One of the most striking aspects of this competition was the significant shakeup between the public and private leaderboards. This highlights the challenge of creating models that generalize well to new chemical space. <br>\nNotably, many of the prize winners seem to be using the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">1DCNN tutorial notebook</a> provided by user <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">Ah</a>. We extend special thanks to Ah for introducing this method, which proved to be highly effective.<br>\nWe'd also like to give special congratulations to Hengck23, looseRs, and GORNA for their submission robustness, showing less variance between public and private scores.</p>\n<h3>Key takeaways</h3>\n<ol>\n<li><strong>Generalization is hard:</strong> The leaderboard shakeup underscores the difficulty of generalizing in this domain.</li>\n<li><strong>Diversity of approaches:</strong> We saw a wide array of molecule representation techniques, from ECFP features, 1DCNNs, GNNs, transformers, and some docking. We look forward to reading these reports on what worked and, especially, what did not!</li>\n<li><strong>Unexpected outcomes:</strong> The success of the 1DCNN approach, derived from a tutorial notebook, offers valuable insights into the high variance of this problem. We suspected that 300M data points would not be enough to solve this problem generally, but now we think that we did not include enough building blocks or proteins in the validation sets. As big as this challenge was, we think many more examples, and much more diversity, is likely to be required to solve this problem.</li>\n<li><strong>The challenge ahead:</strong> Results suggest this problem is more complex than is commonly understood, and that we have a long way to go before we can trust ML to design drugs that are not extremely similar to the training set.</li>\n</ol>\n<h3>Looking forward</h3>\n<p>What kind of dataset would be needed to actually solve this problem? We'd love to hear your thoughts on this.</p>\n<h3>Call to action</h3>\n<ol>\n<li><strong>Share your experiences:</strong> We encourage all participants to share their approaches, challenges faced, and lessons learned.</li>\n<li><strong>Discuss future directions:</strong> What do you think are the next steps in advancing ML for drug discovery?</li>\n<li><strong>Stay engaged:</strong> While the competition is ending, the problem of using ML to cure diseases remains. We hope this encourages many ML experts to continue working on this problem.</li>\n</ol>\n<h3>We’re not done with BELKA yet</h3>\n<p>We appreciate that binary labels may not be satisfying to everyone; moreover, there are many questions to be asked on the behaviors underlying the data we made here. To facilitate deeper exploration, we’ll be releasing all of the replicates of sequencing read counts for all molecule/protein pairs surfaced in this competition - some 3.6B physical measurements - on a public portal in the upcoming months: <a href=\"https://polarishub.io/\" target=\"_blank\">https://polarishub.io/</a>. <br>\nStay tuned!</p>\n<h3>Conclusion</h3>\n<p>The BELKA competition has highlighted both the progress made and the challenges ahead in ML-driven drug discovery. Your contributions have advanced the field and pointed out areas for future focus.<br>\nThank you again for your participation. We're eager to see how the insights from this competition will shape the future of drug discovery and ML.<br>\nBest regards, The BELKA competition team</p>",
  "messages": [
    {
      "id": 2912498,
      "postDate": "2024-07-09T00:09:56.933Z",
      "content": "<p>Dear BELKA competitors,<br>\nThank you for your participation in this challenge! We hope it has been an interesting and rewarding experience. Your efforts have provided valuable insights into the complexities of ML-driven drug discovery.</p>\n<h3>Competition Overview</h3>\n<p>The BELKA competition aimed to advance ML in drug discovery by providing a dataset of ~300M protein/molecule interactions as training material for modeling. Our goal was to encourage the development of models that could generalize across chemical space (which is quite vast!) and help discover new life-saving medicines. To evaluate generalization, the private test set contained a large number of molecules from a different corner of chemical space provided in the training and public test sets. </p>\n<h3>Leaderboard shakeup</h3>\n<p>One of the most striking aspects of this competition was the significant shakeup between the public and private leaderboards. This highlights the challenge of creating models that generalize well to new chemical space. <br>\nNotably, many of the prize winners seem to be using the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">1DCNN tutorial notebook</a> provided by user <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">Ah</a>. We extend special thanks to Ah for introducing this method, which proved to be highly effective.<br>\nWe'd also like to give special congratulations to Hengck23, looseRs, and GORNA for their submission robustness, showing less variance between public and private scores.</p>\n<h3>Key takeaways</h3>\n<ol>\n<li><strong>Generalization is hard:</strong> The leaderboard shakeup underscores the difficulty of generalizing in this domain.</li>\n<li><strong>Diversity of approaches:</strong> We saw a wide array of molecule representation techniques, from ECFP features, 1DCNNs, GNNs, transformers, and some docking. We look forward to reading these reports on what worked and, especially, what did not!</li>\n<li><strong>Unexpected outcomes:</strong> The success of the 1DCNN approach, derived from a tutorial notebook, offers valuable insights into the high variance of this problem. We suspected that 300M data points would not be enough to solve this problem generally, but now we think that we did not include enough building blocks or proteins in the validation sets. As big as this challenge was, we think many more examples, and much more diversity, is likely to be required to solve this problem.</li>\n<li><strong>The challenge ahead:</strong> Results suggest this problem is more complex than is commonly understood, and that we have a long way to go before we can trust ML to design drugs that are not extremely similar to the training set.</li>\n</ol>\n<h3>Looking forward</h3>\n<p>What kind of dataset would be needed to actually solve this problem? We'd love to hear your thoughts on this.</p>\n<h3>Call to action</h3>\n<ol>\n<li><strong>Share your experiences:</strong> We encourage all participants to share their approaches, challenges faced, and lessons learned.</li>\n<li><strong>Discuss future directions:</strong> What do you think are the next steps in advancing ML for drug discovery?</li>\n<li><strong>Stay engaged:</strong> While the competition is ending, the problem of using ML to cure diseases remains. We hope this encourages many ML experts to continue working on this problem.</li>\n</ol>\n<h3>We’re not done with BELKA yet</h3>\n<p>We appreciate that binary labels may not be satisfying to everyone; moreover, there are many questions to be asked on the behaviors underlying the data we made here. To facilitate deeper exploration, we’ll be releasing all of the replicates of sequencing read counts for all molecule/protein pairs surfaced in this competition - some 3.6B physical measurements - on a public portal in the upcoming months: <a href=\"https://polarishub.io/\" target=\"_blank\">https://polarishub.io/</a>. <br>\nStay tuned!</p>\n<h3>Conclusion</h3>\n<p>The BELKA competition has highlighted both the progress made and the challenges ahead in ML-driven drug discovery. Your contributions have advanced the field and pointed out areas for future focus.<br>\nThank you again for your participation. We're eager to see how the insights from this competition will shape the future of drug discovery and ML.<br>\nBest regards, The BELKA competition team</p>",
      "rawMarkdown": "\nDear BELKA competitors,\nThank you for your participation in this challenge! We hope it has been an interesting and rewarding experience. Your efforts have provided valuable insights into the complexities of ML-driven drug discovery.\n### Competition Overview\nThe BELKA competition aimed to advance ML in drug discovery by providing a dataset of ~300M protein/molecule interactions as training material for modeling. Our goal was to encourage the development of models that could generalize across chemical space (which is quite vast!) and help discover new life-saving medicines. To evaluate generalization, the private test set contained a large number of molecules from a different corner of chemical space provided in the training and public test sets. \n### Leaderboard shakeup\nOne of the most striking aspects of this competition was the significant shakeup between the public and private leaderboards. This highlights the challenge of creating models that generalize well to new chemical space. \nNotably, many of the prize winners seem to be using the [1DCNN tutorial notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) provided by user [Ah](https://www.kaggle.com/ahmedelfazouan). We extend special thanks to Ah for introducing this method, which proved to be highly effective.\nWe'd also like to give special congratulations to Hengck23, looseRs, and GORNA for their submission robustness, showing less variance between public and private scores.\n### Key takeaways\n1. **Generalization is hard:** The leaderboard shakeup underscores the difficulty of generalizing in this domain.\n2. **Diversity of approaches:** We saw a wide array of molecule representation techniques, from ECFP features, 1DCNNs, GNNs, transformers, and some docking. We look forward to reading these reports on what worked and, especially, what did not!\n3. **Unexpected outcomes:** The success of the 1DCNN approach, derived from a tutorial notebook, offers valuable insights into the high variance of this problem. We suspected that 300M data points would not be enough to solve this problem generally, but now we think that we did not include enough building blocks or proteins in the validation sets. As big as this challenge was, we think many more examples, and much more diversity, is likely to be required to solve this problem.\n4. **The challenge ahead:** Results suggest this problem is more complex than is commonly understood, and that we have a long way to go before we can trust ML to design drugs that are not extremely similar to the training set.\n### Looking forward\nWhat kind of dataset would be needed to actually solve this problem? We'd love to hear your thoughts on this.\n### Call to action\n1. **Share your experiences:** We encourage all participants to share their approaches, challenges faced, and lessons learned.\n2. **Discuss future directions:** What do you think are the next steps in advancing ML for drug discovery?\n3. **Stay engaged:** While the competition is ending, the problem of using ML to cure diseases remains. We hope this encourages many ML experts to continue working on this problem.\n### We’re not done with BELKA yet\nWe appreciate that binary labels may not be satisfying to everyone; moreover, there are many questions to be asked on the behaviors underlying the data we made here. To facilitate deeper exploration, we’ll be releasing all of the replicates of sequencing read counts for all molecule/protein pairs surfaced in this competition - some 3.6B physical measurements - on a public portal in the upcoming months: [https://polarishub.io/](https://polarishub.io/). \nStay tuned!\n### Conclusion\nThe BELKA competition has highlighted both the progress made and the challenges ahead in ML-driven drug discovery. Your contributions have advanced the field and pointed out areas for future focus.\nThank you again for your participation. We're eager to see how the insights from this competition will shape the future of drug discovery and ML.\n\nBest regards, The BELKA competition team\n",
      "votes": 42
    },
    {
      "id": 2912721,
      "postDate": "2024-07-09T03:45:46.387Z",
      "content": "<p>I thank the host, <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, for making this great competition with the long-term goal to advanced the science &amp; engineer of drug discovery. Special thank to <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for shared fruitful discussions since the begining. Also <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> (Greysnow) who simplified the dataset for us. <a href=\"https://www.kaggle.com/w5833946\" target=\"_blank\">@w5833946</a> (Ruby) to finally convinced the host to change the metric into the more meaningful direction.</p>\n<hr>\n<h3>Brief personal summary :</h3>\n<p>Falling from 26th place to 128th place due to weak CV design.</p>\n<p>In this competition, I decide to focus on writing my own pipeline and get less distraction. With 1-2 months, I have built 4 pipelines : </p>\n<p><strong>Chemberta, Molformer, <a href=\"https://docs.dgl.ai/en/0.8.x/generated/dgl.nn.pytorch.conv.EGATConv.html\" target=\"_blank\">2D Edge-GAT GNN</a>, <a href=\"https://arxiv.org/pdf/2102.09844.pdf\" target=\"_blank\">3D-Equivariance GNN</a></strong></p>\n<p>I have to learn a lot on RDKits, DGLLife to make EGAT and EGNN works. EGNN are totally difficult to applied as 3D structure is so unreliable with RDKits construction and requires much space and computation resources.</p>\n<p>(Toward the end, I prefered 2D-GNN over Transformer because it doesn't need SMILES syntax augmentation.)</p>\n<p>I tried to fight overfitting with pseudo labeling, external BindDB and sEH data, but it does not help due to too vast space of unknown. I investigated 2d-UMAP of my best model (EGAT) and found that my model mostly have no clue on non-triazine and non-shared BB data.</p>\n<p>Similar to many participants, my <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518943#2912665\" target=\"_blank\">1-fold Molformer alone would get gold</a>, but all other models did not generalize well. So, averaging on all models overfits the public LB.  </p>\n<p><strong>What I learned for this short-term competition</strong> still a rookie mistake, though I already have some experience. I should have built 2 separated CVs to measure both public and private LB.</p>\n<p><strong>Total cost</strong> I rent an RTX-4090 machine so that I can trained 98M data in the last 2 weeks after all the pipelines are successfully built on Kaggle. Total cost was around 110 USD.</p>\n<hr>\n<h2>Looking forward on the long-term goal of Belka</h2>\n<ul>\n<li>Would love to understand more on \"non-triazine\" and \"non-shared BB\" performance. <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/503232\" target=\"_blank\">As the host mentioned</a>, the private LB consists of 3 parts \"non-triazine\", \"shared BB\" and \"non-shared BB\". </li>\n</ul>\n<p>Therefore, <strong>what if we eliminate the shared-BB contribution and focus solely on the 2 unknowns??</strong> <br>\nWhat are the best model scores on this competition regarding this? Is the best model even reach 0.1 on the total unknown? <br>\n<a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, I would be much appreciate if you can provide insights on this aspect.</p>\n<p><strong>Possible future direction of dataset:</strong> As we known, physics of binding are based on surface interaction of the target protein and ligands. Is it possible/practical to design dataset / goal to predict this 3D information?</p>\n<p>If we are still predict targets with classification/regression, it would needs billions of training data to generalize over different ranges of BBs and molecules-cores, similar to the scale of training data used by ChatGPT and Diffusion models.</p>",
      "rawMarkdown": "I thank the host, @andrewdblevins, for making this great competition with the long-term goal to advanced the science & engineer of drug discovery. Special thank to @chemdatafarmer @roberthatch and @hengck23 for shared fruitful discussions since the begining. Also @shlomoron (Greysnow) who simplified the dataset for us. @w5833946 (Ruby) to finally convinced the host to change the metric into the more meaningful direction.\n\n---\n\n### Brief personal summary : \nFalling from 26th place to 128th place due to weak CV design.\n \nIn this competition, I decide to focus on writing my own pipeline and get less distraction. With 1-2 months, I have built 4 pipelines : \n\n**Chemberta, Molformer, [2D Edge-GAT GNN](https://docs.dgl.ai/en/0.8.x/generated/dgl.nn.pytorch.conv.EGATConv.html), [3D-Equivariance GNN](https://arxiv.org/pdf/2102.09844.pdf)**\n\nI have to learn a lot on RDKits, DGLLife to make EGAT and EGNN works. EGNN are totally difficult to applied as 3D structure is so unreliable with RDKits construction and requires much space and computation resources.\n\n(Toward the end, I prefered 2D-GNN over Transformer because it doesn't need SMILES syntax augmentation.)\n\nI tried to fight overfitting with pseudo labeling, external BindDB and sEH data, but it does not help due to too vast space of unknown. I investigated 2d-UMAP of my best model (EGAT) and found that my model mostly have no clue on non-triazine and non-shared BB data.\n\nSimilar to many participants, my [1-fold Molformer alone would get gold](https://www.kaggle.com/competitions/leash-BELKA/discussion/518943#2912665), but all other models did not generalize well. So, averaging on all models overfits the public LB.  \n\n**What I learned for this short-term competition** still a rookie mistake, though I already have some experience. I should have built 2 separated CVs to measure both public and private LB.\n\n**Total cost** I rent an RTX-4090 machine so that I can trained 98M data in the last 2 weeks after all the pipelines are successfully built on Kaggle. Total cost was around 110 USD.\n\n---\n\n## Looking forward on the long-term goal of Belka\n\n- Would love to understand more on \"non-triazine\" and \"non-shared BB\" performance. [As the host mentioned](https://www.kaggle.com/competitions/leash-BELKA/discussion/503232), the private LB consists of 3 parts \"non-triazine\", \"shared BB\" and \"non-shared BB\". \n\t\nTherefore, **what if we eliminate the shared-BB contribution and focus solely on the 2 unknowns??** \nWhat are the best model scores on this competition regarding this? Is the best model even reach 0.1 on the total unknown? \n@andrewdblevins, I would be much appreciate if you can provide insights on this aspect.\n\n**Possible future direction of dataset:** As we known, physics of binding are based on surface interaction of the target protein and ligands. Is it possible/practical to design dataset / goal to predict this 3D information?\n\nIf we are still predict targets with classification/regression, it would needs billions of training data to generalize over different ranges of BBs and molecules-cores, similar to the scale of training data used by ChatGPT and Diffusion models.",
      "votes": 10,
      "replies": [
        {
          "id": 2917595,
          "postDate": "2024-07-11T17:35:02.187Z",
          "content": "<p>Thanks for sharing! May I ask where you rent the RTX-4090 machine?</p>",
          "rawMarkdown": "Thanks for sharing! May I ask where you rent the RTX-4090 machine?",
          "votes": 1,
          "replies": [
            {
              "id": 2919509,
              "postDate": "2024-07-13T01:04:39.347Z",
              "content": "<p><a href=\"https://www.kaggle.com/lililycai\" target=\"_blank\">@lililycai</a> In this competition I rent at vast.ai , in some other times I also considered jarvislabs.ai and runpod.ai :)</p>",
              "rawMarkdown": "@lililycai In this competition I rent at vast.ai , in some other times I also considered jarvislabs.ai and runpod.ai :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2912529,
      "postDate": "2024-07-09T00:49:41.613Z",
      "content": "<p>Thanks for hosting this competition and providing the community with this dataset! It was a fun comp and the first time I've had my hands on this much chemical data. I'm looking forward to seeing folks write-ups.</p>",
      "rawMarkdown": "Thanks for hosting this competition and providing the community with this dataset! It was a fun comp and the first time I've had my hands on this much chemical data. I'm looking forward to seeing folks write-ups.",
      "votes": 6,
      "replies": [
        {
          "id": 2914597,
          "postDate": "2024-07-10T04:17:11.223Z",
          "content": "<p>Thank you for providing great notebooks and dataset!</p>",
          "rawMarkdown": "Thank you for providing great notebooks and dataset!",
          "votes": 1,
          "replies": [
            {
              "id": 2976953,
              "postDate": "2024-09-02T13:01:42.473Z",
              "content": "<p>You are very welcome :) it was fun!</p>",
              "rawMarkdown": "You are very welcome :) it was fun!"
            }
          ]
        }
      ]
    },
    {
      "id": 2976609,
      "postDate": "2024-09-02T06:17:14.257Z",
      "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n<p>nvidia has just make a prediction of small molecule binding affinity perdiction.<br>\nThis may be of interested to you!</p>\n<p>NVIDIA<br>\nGenerative Virtual Screening for Drug Discovery: Search and optimize a library of small molecules to identify chemical structures that bind to a target protein.<br>\n<a href=\"https://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery\" target=\"_blank\">https://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F90c1b11078002f119b2e6a5aa38b3f44%2FSelection_999(5990).png?generation=1725258055052233&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@andrewdblevins \n\nnvidia has just make a prediction of small molecule binding affinity perdiction.\nThis may be of interested to you!\n\nNVIDIA\nGenerative Virtual Screening for Drug Discovery: Search and optimize a library of small molecules to identify chemical structures that bind to a target protein.\nhttps://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F90c1b11078002f119b2e6a5aa38b3f44%2FSelection_999(5990).png?generation=1725258055052233&alt=media)",
      "votes": 4
    },
    {
      "id": 2914509,
      "postDate": "2024-07-10T02:04:00.507Z",
      "content": "<p>Are any of the models and weights useful for you as the competition organizer?</p>",
      "rawMarkdown": "Are any of the models and weights useful for you as the competition organizer?",
      "votes": 3
    },
    {
      "id": 2918888,
      "postDate": "2024-07-12T15:24:27.577Z",
      "content": "<p>From Kaggle we would also like to thank everyone for participating!</p>\n<p>This competition ended with 9143 registrations and 2357 participants on 1946 teams. We had 37022 submissions from 91 countries. For 500 users 21 in the top 100!), this was their first competition. Thank you all for your hard work in this competition and congratulations to our winners and to those who gained a new ranking!  </p>\n<p>The top potential winning teams will be contacted via email for the next steps. We look forward to learning more about their winning solutions.</p>\n<p>We've cleaned the leaderboard and disqualified some teams that have violated the rules. If you think you were removed by mistake, or believe you have evidence that suggests another team cheated, please contact <a href=\"https://www.kaggle.com/compliance\" target=\"_blank\">compliance</a>. Please fill in all the fields honestly.</p>\n<p>We highly encourage you to post a solution write-up about your approach and solution in the forums (see <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">instructions</a>). You may also refer to <a href=\"https://www.kaggle.com/solution-write-up-documentation\" target=\"_blank\">Kaggle Solution Write-Up Documentation</a> for guidance. You are  also encouraged to <a href=\"https://www.kaggle.com/docs/models#publishing-a-model\" target=\"_blank\">publish your models on Kaggle Models</a>!</p>\n<p>Thanks for continuing to make Kaggle a great place to learn, practice, and test data science techniques!</p>\n<p>Happy Modeling!</p>",
      "rawMarkdown": "From Kaggle we would also like to thank everyone for participating!\n\nThis competition ended with 9143 registrations and 2357 participants on 1946 teams. We had 37022 submissions from 91 countries. For 500 users 21 in the top 100!), this was their first competition. Thank you all for your hard work in this competition and congratulations to our winners and to those who gained a new ranking!  \n\nThe top potential winning teams will be contacted via email for the next steps. We look forward to learning more about their winning solutions.\n\nWe've cleaned the leaderboard and disqualified some teams that have violated the rules. If you think you were removed by mistake, or believe you have evidence that suggests another team cheated, please contact [compliance](https://www.kaggle.com/compliance). Please fill in all the fields honestly.\n\nWe highly encourage you to post a solution write-up about your approach and solution in the forums (see [instructions](https://www.kaggle.com/discussions/product-feedback/373153)). You may also refer to [Kaggle Solution Write-Up Documentation](https://www.kaggle.com/solution-write-up-documentation) for guidance. You are  also encouraged to [publish your models on Kaggle Models](https://www.kaggle.com/docs/models#publishing-a-model)!\n\nThanks for continuing to make Kaggle a great place to learn, practice, and test data science techniques!\n\nHappy Modeling!\n",
      "votes": 4
    },
    {
      "id": 2912801,
      "postDate": "2024-07-09T04:37:56.563Z",
      "content": "<p>Greetings! I wonder if you can share per-target group scores in format of .csv for both public and private? Would be fun to analyze it. </p>",
      "rawMarkdown": "Greetings! I wonder if you can share per-target group scores in format of .csv for both public and private? Would be fun to analyze it. ",
      "votes": 4,
      "replies": [
        {
          "id": 2912825,
          "postDate": "2024-07-09T05:10:04.850Z",
          "content": "<p>We are also curious about this, CV/LB scores behaved differently across various targets and bb split status. We would like to see the 3×3 (3×2 for LB) mAP score for each target and each split.</p>",
          "rawMarkdown": "We are also curious about this, CV/LB scores behaved differently across various targets and bb split status. We would like to see the 3×3 (3×2 for LB) mAP score for each target and each split.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2912986,
      "postDate": "2024-07-09T06:52:26.633Z",
      "content": "<p>What of the top student prize?</p>",
      "rawMarkdown": "What of the top student prize?",
      "votes": 2,
      "replies": [
        {
          "id": 2915790,
          "postDate": "2024-07-10T16:38:37.270Z",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> <a href=\"https://www.kaggle.com/ianquigley\" target=\"_blank\">@ianquigley</a> please bring some light to the question. I think we can be a student prize contenders but it's impossible for me to determine status of all 6th…12th place participants.</p>",
          "rawMarkdown": "@andrewdblevins @ianquigley please bring some light to the question. I think we can be a student prize contenders but it's impossible for me to determine status of all 6th...12th place participants.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2912582,
      "postDate": "2024-07-09T01:26:15.883Z",
      "content": "<p>Thanks for the debrief.</p>\n<p>This was my first Kaggle competition and I am hooked.</p>\n<p>It's such a thrill to navigate all the possible paths to victory, and then engineer your way through the mechanics.</p>\n<p>Having NO domain expertise made this a double thrill. I learned a lot in the 20 or so days I worked on this.</p>\n<p>Citizen science is rad.</p>",
      "rawMarkdown": "Thanks for the debrief.\n\nThis was my first Kaggle competition and I am hooked.\n\nIt's such a thrill to navigate all the possible paths to victory, and then engineer your way through the mechanics.\n\nHaving NO domain expertise made this a double thrill. I learned a lot in the 20 or so days I worked on this.\n\nCitizen science is rad.",
      "votes": 2
    },
    {
      "id": 2912531,
      "postDate": "2024-07-09T00:51:49.457Z",
      "content": "<p>With how far we have to go, I wonder if a future competition aiming predict three (not as hard) buckets would, ironically, do better at advancing generality?</p>\n<ul>\n<li>Shared BBs (1, 2, 3) -&gt; (1, 2, 3)</li>\n<li>2/3rds shared (1, 2, 3) -&gt; (4, 2, 3)</li>\n<li>1/3rd shared (1, 2, 3) -&gt; (1, 5, 6)</li>\n</ul>\n<p>This would force competitors to incrementally figure out how to predict unknown substitutions.</p>\n<p>I also mentioned in passing at some point, but a metric focusing on 'top 1%' predictions or something similar might also be better. We know from this competition we can't effectively predict most binding events on unseen data, but with this metric we don't really see whether or not we can predict a fairly small pool of novel candidates from which we'll get a high percentage of binding events.</p>\n<p>Another interesting prediction task would be: given train data with these 1000 building blocks, and test data comprising of 500 or 1000 other building blocks, rank each building block in test by how many binding combinations it will hit with the other bbs in the test dataset. Or instead of rank, pick the 500 out of the 1000 that will give the most binding hits or something like that. Maybe that all depends on the business goals, I guess.</p>",
      "rawMarkdown": "With how far we have to go, I wonder if a future competition aiming predict three (not as hard) buckets would, ironically, do better at advancing generality?\n\n* Shared BBs (1, 2, 3) -> (1, 2, 3)\n* 2/3rds shared (1, 2, 3) -> (4, 2, 3)\n* 1/3rd shared (1, 2, 3) -> (1, 5, 6)\n\nThis would force competitors to incrementally figure out how to predict unknown substitutions.\n\nI also mentioned in passing at some point, but a metric focusing on 'top 1%' predictions or something similar might also be better. We know from this competition we can't effectively predict most binding events on unseen data, but with this metric we don't really see whether or not we can predict a fairly small pool of novel candidates from which we'll get a high percentage of binding events.\n\nAnother interesting prediction task would be: given train data with these 1000 building blocks, and test data comprising of 500 or 1000 other building blocks, rank each building block in test by how many binding combinations it will hit with the other bbs in the test dataset. Or instead of rank, pick the 500 out of the 1000 that will give the most binding hits or something like that. Maybe that all depends on the business goals, I guess.",
      "votes": 2
    },
    {
      "id": 2956957,
      "postDate": "2024-08-12T16:54:58.893Z",
      "content": "<p>Thank you for hosting this competition! It would be really great to continue improving our models with the ground truth data for test set and read counts would be amazing as well. Do you have a timeline for when these might become available?</p>",
      "rawMarkdown": "Thank you for hosting this competition! It would be really great to continue improving our models with the ground truth data for test set and read counts would be amazing as well. Do you have a timeline for when these might become available?",
      "replies": [
        {
          "id": 2961726,
          "postDate": "2024-08-16T20:53:06.637Z",
          "content": "<p>The Polaris team tells us by the end of September.</p>\n<p><a href=\"https://polarishub.io/\" target=\"_blank\">https://polarishub.io/</a></p>",
          "rawMarkdown": "The Polaris team tells us by the end of September.\n\nhttps://polarishub.io/"
        }
      ]
    },
    {
      "id": 2916344,
      "postDate": "2024-07-10T23:37:24.270Z",
      "content": "<p>Our team has tried sincerely to solve the problem, but we were removed from the leaderboard. Organizers, could you please tell us the reason? The models used to create the submission files included Xgboost, 1DCNN, Chemberta, and MoLFormer.</p>\n<p>PS: I have sent an email to the Competitions Compliance Team regarding this.</p>",
      "rawMarkdown": "Our team has tried sincerely to solve the problem, but we were removed from the leaderboard. Organizers, could you please tell us the reason? The models used to create the submission files included Xgboost, 1DCNN, Chemberta, and MoLFormer.\n\nPS: I have sent an email to the Competitions Compliance Team regarding this.",
      "replies": [
        {
          "id": 2916348,
          "postDate": "2024-07-10T23:46:57.440Z",
          "content": "<p>I posted an outline of the solution I created before the competition ended.</p>\n<hr>\n<p>First of all, I would like to thank the organizers and participants of the competition. I learned a lot from the public notes and discussions. I would also like to thank the three people who formed my team.</p>\n<p>The simple solution is as follows.</p>\n<p>・Split into shared and non-shared<br>\nThe training data given this time is strong for shared, but not suitable for non-shared.</p>\n<p>Therefore, it was necessary to split it into shared and non-shared and use a model suitable for each.</p>\n<p>・Use a pre-trained model for non-shared<br>\nUsing a pre-trained model can strengthen predictions for unknown bb.</p>\n<p>・Perform an ensemble with many models<br>\nSince the final submission is CSV, there is no actual limit on inference time.</p>\n<p>Also, this time the target was about 0.5%, which was imbalanced data, so we extended the public LB by ensuring diversity.</p>",
          "rawMarkdown": "I posted an outline of the solution I created before the competition ended.\n\n-----------------\n\nFirst of all, I would like to thank the organizers and participants of the competition. I learned a lot from the public notes and discussions. I would also like to thank the three people who formed my team.\n\nThe simple solution is as follows.\n\n・Split into shared and non-shared\nThe training data given this time is strong for shared, but not suitable for non-shared.\n\nTherefore, it was necessary to split it into shared and non-shared and use a model suitable for each.\n\n・Use a pre-trained model for non-shared\nUsing a pre-trained model can strengthen predictions for unknown bb.\n\n・Perform an ensemble with many models\nSince the final submission is CSV, there is no actual limit on inference time.\n\nAlso, this time the target was about 0.5%, which was imbalanced data, so we extended the public LB by ensuring diversity."
        },
        {
          "id": 2916368,
          "postDate": "2024-07-11T00:23:42.463Z",
          "content": "<p>Since I have the opportunity, I will also share the information about the best private LB I got, 0.297. I hope it will be helpful in the future.</p>\n<p>The baseline of the public note is as follows.</p>\n<p>Base Line: <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data</a></p>\n<h4>Hyperparameters</h4>\n<ul>\n<li>EPOCHS = 25</li>\n<li>BATCH_SIZE = 10000</li>\n<li>LR = 8e-4</li>\n<li>WD = 0.01</li>\n<li>NBR_FOLDS = 15</li>\n<li>SELECTED_FOLDS = [0]</li>\n<li>SEED = 2027</li>\n</ul>\n<p>I don't know if the hyperparameter tuning was just right or if it fit the random numbers well. But there is no doubt that 1DCNN was effective for private LB.</p>",
          "rawMarkdown": "Since I have the opportunity, I will also share the information about the best private LB I got, 0.297. I hope it will be helpful in the future.\n\nThe baseline of the public note is as follows.\n\nBase Line: https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\n\n#### Hyperparameters\n- EPOCHS = 25\n- BATCH_SIZE = 10000\n- LR = 8e-4\n- WD = 0.01\n- NBR_FOLDS = 15\n- SELECTED_FOLDS = [0]\n- SEED = 2027\n\nI don't know if the hyperparameter tuning was just right or if it fit the random numbers well. But there is no doubt that 1DCNN was effective for private LB."
        }
      ]
    },
    {
      "id": 2914916,
      "postDate": "2024-07-10T08:35:33.467Z",
      "content": "<p>It has been mentioned in different ways, but it would be very interesting and useful to have not only the target values for the public and private sets, but also the various model predictions, if that's possible.  A .csv of the test compound ids with columns of the actual classification and as many model predictions the is practical could be very illuminating regarding generalization and target-specific effects.</p>",
      "rawMarkdown": "It has been mentioned in different ways, but it would be very interesting and useful to have not only the target values for the public and private sets, but also the various model predictions, if that's possible.  A .csv of the test compound ids with columns of the actual classification and as many model predictions the is practical could be very illuminating regarding generalization and target-specific effects.",
      "replies": [
        {
          "id": 2915844,
          "postDate": "2024-07-10T17:06:14.060Z",
          "content": "<p>give me a bit before releasing the answer key, but here are the split groups used<br>\n<a href=\"https://www.kaggle.com/datasets/andrewdblevins/leash-belka-split-groups\" target=\"_blank\">https://www.kaggle.com/datasets/andrewdblevins/leash-belka-split-groups</a></p>",
          "rawMarkdown": "give me a bit before releasing the answer key, but here are the split groups used\nhttps://www.kaggle.com/datasets/andrewdblevins/leash-belka-split-groups",
          "votes": 2
        }
      ]
    },
    {
      "id": 2913104,
      "postDate": "2024-07-09T08:56:39.737Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> for this really challenging dataset.<br>\nI look forward to the real read counts and possibly also 'no-target selection' results for control. This should be a more realistic dataset and might even allow out-of-domain prediction (albeit restricted to the training diversity) from just the competition dataset w/o also using external activity reports from, say, the ZINC database.</p>",
      "rawMarkdown": "Thanks @andrewdblevins for this really challenging dataset.\nI look forward to the real read counts and possibly also 'no-target selection' results for control. This should be a more realistic dataset and might even allow out-of-domain prediction (albeit restricted to the training diversity) from just the competition dataset w/o also using external activity reports from, say, the ZINC database.",
      "replies": [
        {
          "id": 2913331,
          "postDate": "2024-07-09T12:24:22.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/bernhardrohde\" target=\"_blank\">@bernhardrohde</a> \"no-target selection\" will be included in the data release, we appreciate that the community may want to try different hit-calling strategies.</p>",
          "rawMarkdown": "@bernhardrohde \"no-target selection\" will be included in the data release, we appreciate that the community may want to try different hit-calling strategies.",
          "votes": 1,
          "replies": [
            {
              "id": 2923527,
              "postDate": "2024-07-16T00:50:00.870Z",
              "content": "<p>Yeah, maybe having the dataset w/ the actual raw read counts can be useful. If some model can predict this, maybe it would allow to rank-order molecules more precisely.</p>",
              "rawMarkdown": "Yeah, maybe having the dataset w/ the actual raw read counts can be useful. If some model can predict this, maybe it would allow to rank-order molecules more precisely.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2916616,
      "postDate": "2024-07-11T06:16:05.247Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2914125,
      "postDate": "2024-07-09T18:25:41.817Z",
      "content": "<p>Very nice! Thanks for the brief!</p>",
      "rawMarkdown": "Very nice! Thanks for the brief!"
    }
  ],
  "comments": [
    {
      "id": 2912721,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2024-07-09T03:45:46.387000",
      "content": "<p>I thank the host, <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, for making this great competition with the long-term goal to advanced the science &amp; engineer of drug discovery. Special thank to <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for shared fruitful discussions since the begining. Also <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> (Greysnow) who simplified the dataset for us. <a href=\"https://www.kaggle.com/w5833946\" target=\"_blank\">@w5833946</a> (Ruby) to finally convinced the host to change the metric into the more meaningful direction.</p>\n<hr>\n<h3>Brief personal summary :</h3>\n<p>Falling from 26th place to 128th place due to weak CV design.</p>\n<p>In this competition, I decide to focus on writing my own pipeline and get less distraction. With 1-2 months, I have built 4 pipelines : </p>\n<p><strong>Chemberta, Molformer, <a href=\"https://docs.dgl.ai/en/0.8.x/generated/dgl.nn.pytorch.conv.EGATConv.html\" target=\"_blank\">2D Edge-GAT GNN</a>, <a href=\"https://arxiv.org/pdf/2102.09844.pdf\" target=\"_blank\">3D-Equivariance GNN</a></strong></p>\n<p>I have to learn a lot on RDKits, DGLLife to make EGAT and EGNN works. EGNN are totally difficult to applied as 3D structure is so unreliable with RDKits construction and requires much space and computation resources.</p>\n<p>(Toward the end, I prefered 2D-GNN over Transformer because it doesn't need SMILES syntax augmentation.)</p>\n<p>I tried to fight overfitting with pseudo labeling, external BindDB and sEH data, but it does not help due to too vast space of unknown. I investigated 2d-UMAP of my best model (EGAT) and found that my model mostly have no clue on non-triazine and non-shared BB data.</p>\n<p>Similar to many participants, my <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518943#2912665\" target=\"_blank\">1-fold Molformer alone would get gold</a>, but all other models did not generalize well. So, averaging on all models overfits the public LB.  </p>\n<p><strong>What I learned for this short-term competition</strong> still a rookie mistake, though I already have some experience. I should have built 2 separated CVs to measure both public and private LB.</p>\n<p><strong>Total cost</strong> I rent an RTX-4090 machine so that I can trained 98M data in the last 2 weeks after all the pipelines are successfully built on Kaggle. Total cost was around 110 USD.</p>\n<hr>\n<h2>Looking forward on the long-term goal of Belka</h2>\n<ul>\n<li>Would love to understand more on \"non-triazine\" and \"non-shared BB\" performance. <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/503232\" target=\"_blank\">As the host mentioned</a>, the private LB consists of 3 parts \"non-triazine\", \"shared BB\" and \"non-shared BB\". </li>\n</ul>\n<p>Therefore, <strong>what if we eliminate the shared-BB contribution and focus solely on the 2 unknowns??</strong> <br>\nWhat are the best model scores on this competition regarding this? Is the best model even reach 0.1 on the total unknown? <br>\n<a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a>, I would be much appreciate if you can provide insights on this aspect.</p>\n<p><strong>Possible future direction of dataset:</strong> As we known, physics of binding are based on surface interaction of the target protein and ligands. Is it possible/practical to design dataset / goal to predict this 3D information?</p>\n<p>If we are still predict targets with classification/regression, it would needs billions of training data to generalize over different ranges of BBs and molecules-cores, similar to the scale of training data used by ChatGPT and Diffusion models.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2917595,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-07-11T17:35:02.187000",
          "content": "<p>Thanks for sharing! May I ask where you rent the RTX-4090 machine?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2919509,
              "author_name": "Neuron Engineer",
              "author_url": "",
              "post_date": "2024-07-13T01:04:39.347000",
              "content": "<p><a href=\"https://www.kaggle.com/lililycai\" target=\"_blank\">@lililycai</a> In this competition I rent at vast.ai , in some other times I also considered jarvislabs.ai and runpod.ai :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2912529,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-07-09T00:49:41.613000",
      "content": "<p>Thanks for hosting this competition and providing the community with this dataset! It was a fun comp and the first time I've had my hands on this much chemical data. I'm looking forward to seeing folks write-ups.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2914597,
          "author_name": "Swikwislkdjc",
          "author_url": "",
          "post_date": "2024-07-10T04:17:11.223000",
          "content": "<p>Thank you for providing great notebooks and dataset!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2976953,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-09-02T13:01:42.473000",
              "content": "<p>You are very welcome :) it was fun!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2976609,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-09-02T06:17:14.257000",
      "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> </p>\n<p>nvidia has just make a prediction of small molecule binding affinity perdiction.<br>\nThis may be of interested to you!</p>\n<p>NVIDIA<br>\nGenerative Virtual Screening for Drug Discovery: Search and optimize a library of small molecules to identify chemical structures that bind to a target protein.<br>\n<a href=\"https://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery\" target=\"_blank\">https://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F90c1b11078002f119b2e6a5aa38b3f44%2FSelection_999(5990).png?generation=1725258055052233&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2914509,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2024-07-10T02:04:00.507000",
      "content": "<p>Are any of the models and weights useful for you as the competition organizer?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2918888,
      "author_name": "Maggie",
      "author_url": "",
      "post_date": "2024-07-12T15:24:27.577000",
      "content": "<p>From Kaggle we would also like to thank everyone for participating!</p>\n<p>This competition ended with 9143 registrations and 2357 participants on 1946 teams. We had 37022 submissions from 91 countries. For 500 users 21 in the top 100!), this was their first competition. Thank you all for your hard work in this competition and congratulations to our winners and to those who gained a new ranking!  </p>\n<p>The top potential winning teams will be contacted via email for the next steps. We look forward to learning more about their winning solutions.</p>\n<p>We've cleaned the leaderboard and disqualified some teams that have violated the rules. If you think you were removed by mistake, or believe you have evidence that suggests another team cheated, please contact <a href=\"https://www.kaggle.com/compliance\" target=\"_blank\">compliance</a>. Please fill in all the fields honestly.</p>\n<p>We highly encourage you to post a solution write-up about your approach and solution in the forums (see <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">instructions</a>). You may also refer to <a href=\"https://www.kaggle.com/solution-write-up-documentation\" target=\"_blank\">Kaggle Solution Write-Up Documentation</a> for guidance. You are  also encouraged to <a href=\"https://www.kaggle.com/docs/models#publishing-a-model\" target=\"_blank\">publish your models on Kaggle Models</a>!</p>\n<p>Thanks for continuing to make Kaggle a great place to learn, practice, and test data science techniques!</p>\n<p>Happy Modeling!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2912801,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-07-09T04:37:56.563000",
      "content": "<p>Greetings! I wonder if you can share per-target group scores in format of .csv for both public and private? Would be fun to analyze it. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2912825,
          "author_name": "S_Kikuchi",
          "author_url": "",
          "post_date": "2024-07-09T05:10:04.850000",
          "content": "<p>We are also curious about this, CV/LB scores behaved differently across various targets and bb split status. We would like to see the 3×3 (3×2 for LB) mAP score for each target and each split.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2912986,
      "author_name": "Vivek Joshy",
      "author_url": "",
      "post_date": "2024-07-09T06:52:26.633000",
      "content": "<p>What of the top student prize?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2915790,
          "author_name": "Ogurtsov",
          "author_url": "",
          "post_date": "2024-07-10T16:38:37.270000",
          "content": "<p><a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> <a href=\"https://www.kaggle.com/ianquigley\" target=\"_blank\">@ianquigley</a> please bring some light to the question. I think we can be a student prize contenders but it's impossible for me to determine status of all 6th…12th place participants.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2912582,
      "author_name": "Matt McDonagh",
      "author_url": "",
      "post_date": "2024-07-09T01:26:15.883000",
      "content": "<p>Thanks for the debrief.</p>\n<p>This was my first Kaggle competition and I am hooked.</p>\n<p>It's such a thrill to navigate all the possible paths to victory, and then engineer your way through the mechanics.</p>\n<p>Having NO domain expertise made this a double thrill. I learned a lot in the 20 or so days I worked on this.</p>\n<p>Citizen science is rad.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2912531,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-07-09T00:51:49.457000",
      "content": "<p>With how far we have to go, I wonder if a future competition aiming predict three (not as hard) buckets would, ironically, do better at advancing generality?</p>\n<ul>\n<li>Shared BBs (1, 2, 3) -&gt; (1, 2, 3)</li>\n<li>2/3rds shared (1, 2, 3) -&gt; (4, 2, 3)</li>\n<li>1/3rd shared (1, 2, 3) -&gt; (1, 5, 6)</li>\n</ul>\n<p>This would force competitors to incrementally figure out how to predict unknown substitutions.</p>\n<p>I also mentioned in passing at some point, but a metric focusing on 'top 1%' predictions or something similar might also be better. We know from this competition we can't effectively predict most binding events on unseen data, but with this metric we don't really see whether or not we can predict a fairly small pool of novel candidates from which we'll get a high percentage of binding events.</p>\n<p>Another interesting prediction task would be: given train data with these 1000 building blocks, and test data comprising of 500 or 1000 other building blocks, rank each building block in test by how many binding combinations it will hit with the other bbs in the test dataset. Or instead of rank, pick the 500 out of the 1000 that will give the most binding hits or something like that. Maybe that all depends on the business goals, I guess.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2956957,
      "author_name": "Terra Sztain",
      "author_url": "",
      "post_date": "2024-08-12T16:54:58.893000",
      "content": "<p>Thank you for hosting this competition! It would be really great to continue improving our models with the ground truth data for test set and read counts would be amazing as well. Do you have a timeline for when these might become available?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2961726,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-08-16T20:53:06.637000",
          "content": "<p>The Polaris team tells us by the end of September.</p>\n<p><a href=\"https://polarishub.io/\" target=\"_blank\">https://polarishub.io/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2916344,
      "author_name": "HayatoFujihara",
      "author_url": "",
      "post_date": "2024-07-10T23:37:24.270000",
      "content": "<p>Our team has tried sincerely to solve the problem, but we were removed from the leaderboard. Organizers, could you please tell us the reason? The models used to create the submission files included Xgboost, 1DCNN, Chemberta, and MoLFormer.</p>\n<p>PS: I have sent an email to the Competitions Compliance Team regarding this.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2916348,
          "author_name": "HayatoFujihara",
          "author_url": "",
          "post_date": "2024-07-10T23:46:57.440000",
          "content": "<p>I posted an outline of the solution I created before the competition ended.</p>\n<hr>\n<p>First of all, I would like to thank the organizers and participants of the competition. I learned a lot from the public notes and discussions. I would also like to thank the three people who formed my team.</p>\n<p>The simple solution is as follows.</p>\n<p>・Split into shared and non-shared<br>\nThe training data given this time is strong for shared, but not suitable for non-shared.</p>\n<p>Therefore, it was necessary to split it into shared and non-shared and use a model suitable for each.</p>\n<p>・Use a pre-trained model for non-shared<br>\nUsing a pre-trained model can strengthen predictions for unknown bb.</p>\n<p>・Perform an ensemble with many models<br>\nSince the final submission is CSV, there is no actual limit on inference time.</p>\n<p>Also, this time the target was about 0.5%, which was imbalanced data, so we extended the public LB by ensuring diversity.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2916368,
          "author_name": "HayatoFujihara",
          "author_url": "",
          "post_date": "2024-07-11T00:23:42.463000",
          "content": "<p>Since I have the opportunity, I will also share the information about the best private LB I got, 0.297. I hope it will be helpful in the future.</p>\n<p>The baseline of the public note is as follows.</p>\n<p>Base Line: <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data</a></p>\n<h4>Hyperparameters</h4>\n<ul>\n<li>EPOCHS = 25</li>\n<li>BATCH_SIZE = 10000</li>\n<li>LR = 8e-4</li>\n<li>WD = 0.01</li>\n<li>NBR_FOLDS = 15</li>\n<li>SELECTED_FOLDS = [0]</li>\n<li>SEED = 2027</li>\n</ul>\n<p>I don't know if the hyperparameter tuning was just right or if it fit the random numbers well. But there is no doubt that 1DCNN was effective for private LB.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2914916,
      "author_name": "KirkDCO",
      "author_url": "",
      "post_date": "2024-07-10T08:35:33.467000",
      "content": "<p>It has been mentioned in different ways, but it would be very interesting and useful to have not only the target values for the public and private sets, but also the various model predictions, if that's possible.  A .csv of the test compound ids with columns of the actual classification and as many model predictions the is practical could be very illuminating regarding generalization and target-specific effects.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2915844,
          "author_name": "Andrew D. Blevins",
          "author_url": "",
          "post_date": "2024-07-10T17:06:14.060000",
          "content": "<p>give me a bit before releasing the answer key, but here are the split groups used<br>\n<a href=\"https://www.kaggle.com/datasets/andrewdblevins/leash-belka-split-groups\" target=\"_blank\">https://www.kaggle.com/datasets/andrewdblevins/leash-belka-split-groups</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2913104,
      "author_name": "Bernhard Rohde",
      "author_url": "",
      "post_date": "2024-07-09T08:56:39.737000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/andrewdblevins\" target=\"_blank\">@andrewdblevins</a> for this really challenging dataset.<br>\nI look forward to the real read counts and possibly also 'no-target selection' results for control. This should be a more realistic dataset and might even allow out-of-domain prediction (albeit restricted to the training diversity) from just the competition dataset w/o also using external activity reports from, say, the ZINC database.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2913331,
          "author_name": "Ian Quigley",
          "author_url": "",
          "post_date": "2024-07-09T12:24:22.350000",
          "content": "<p><a href=\"https://www.kaggle.com/bernhardrohde\" target=\"_blank\">@bernhardrohde</a> \"no-target selection\" will be included in the data release, we appreciate that the community may want to try different hit-calling strategies.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2923527,
              "author_name": "Francois Berenger",
              "author_url": "",
              "post_date": "2024-07-16T00:50:00.870000",
              "content": "<p>Yeah, maybe having the dataset w/ the actual raw read counts can be useful. If some model can predict this, maybe it would allow to rank-order molecules more precisely.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2916616,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-11T06:16:05.247000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2914125,
      "author_name": "Gcdelgado",
      "author_url": "",
      "post_date": "2024-07-09T18:25:41.817000",
      "content": "<p>Very nice! Thanks for the brief!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2912498": "\nDear BELKA competitors,\nThank you for your participation in this challenge! We hope it has been an interesting and rewarding experience. Your efforts have provided valuable insights into the complexities of ML-driven drug discovery.\n### Competition Overview\nThe BELKA competition aimed to advance ML in drug discovery by providing a dataset of ~300M protein/molecule interactions as training material for modeling. Our goal was to encourage the development of models that could generalize across chemical space (which is quite vast!) and help discover new life-saving medicines. To evaluate generalization, the private test set contained a large number of molecules from a different corner of chemical space provided in the training and public test sets. \n### Leaderboard shakeup\nOne of the most striking aspects of this competition was the significant shakeup between the public and private leaderboards. This highlights the challenge of creating models that generalize well to new chemical space. \nNotably, many of the prize winners seem to be using the [1DCNN tutorial notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) provided by user [Ah](https://www.kaggle.com/ahmedelfazouan). We extend special thanks to Ah for introducing this method, which proved to be highly effective.\nWe'd also like to give special congratulations to Hengck23, looseRs, and GORNA for their submission robustness, showing less variance between public and private scores.\n### Key takeaways\n1. **Generalization is hard:** The leaderboard shakeup underscores the difficulty of generalizing in this domain.\n2. **Diversity of approaches:** We saw a wide array of molecule representation techniques, from ECFP features, 1DCNNs, GNNs, transformers, and some docking. We look forward to reading these reports on what worked and, especially, what did not!\n3. **Unexpected outcomes:** The success of the 1DCNN approach, derived from a tutorial notebook, offers valuable insights into the high variance of this problem. We suspected that 300M data points would not be enough to solve this problem generally, but now we think that we did not include enough building blocks or proteins in the validation sets. As big as this challenge was, we think many more examples, and much more diversity, is likely to be required to solve this problem.\n4. **The challenge ahead:** Results suggest this problem is more complex than is commonly understood, and that we have a long way to go before we can trust ML to design drugs that are not extremely similar to the training set.\n### Looking forward\nWhat kind of dataset would be needed to actually solve this problem? We'd love to hear your thoughts on this.\n### Call to action\n1. **Share your experiences:** We encourage all participants to share their approaches, challenges faced, and lessons learned.\n2. **Discuss future directions:** What do you think are the next steps in advancing ML for drug discovery?\n3. **Stay engaged:** While the competition is ending, the problem of using ML to cure diseases remains. We hope this encourages many ML experts to continue working on this problem.\n### We’re not done with BELKA yet\nWe appreciate that binary labels may not be satisfying to everyone; moreover, there are many questions to be asked on the behaviors underlying the data we made here. To facilitate deeper exploration, we’ll be releasing all of the replicates of sequencing read counts for all molecule/protein pairs surfaced in this competition - some 3.6B physical measurements - on a public portal in the upcoming months: [https://polarishub.io/](https://polarishub.io/). \nStay tuned!\n### Conclusion\nThe BELKA competition has highlighted both the progress made and the challenges ahead in ML-driven drug discovery. Your contributions have advanced the field and pointed out areas for future focus.\nThank you again for your participation. We're eager to see how the insights from this competition will shape the future of drug discovery and ML.\n\nBest regards, The BELKA competition team\n",
    "2912721": "I thank the host, @andrewdblevins, for making this great competition with the long-term goal to advanced the science & engineer of drug discovery. Special thank to @chemdatafarmer @roberthatch and @hengck23 for shared fruitful discussions since the begining. Also @shlomoron (Greysnow) who simplified the dataset for us. @w5833946 (Ruby) to finally convinced the host to change the metric into the more meaningful direction.\n\n---\n\n### Brief personal summary : \nFalling from 26th place to 128th place due to weak CV design.\n \nIn this competition, I decide to focus on writing my own pipeline and get less distraction. With 1-2 months, I have built 4 pipelines : \n\n**Chemberta, Molformer, [2D Edge-GAT GNN](https://docs.dgl.ai/en/0.8.x/generated/dgl.nn.pytorch.conv.EGATConv.html), [3D-Equivariance GNN](https://arxiv.org/pdf/2102.09844.pdf)**\n\nI have to learn a lot on RDKits, DGLLife to make EGAT and EGNN works. EGNN are totally difficult to applied as 3D structure is so unreliable with RDKits construction and requires much space and computation resources.\n\n(Toward the end, I prefered 2D-GNN over Transformer because it doesn't need SMILES syntax augmentation.)\n\nI tried to fight overfitting with pseudo labeling, external BindDB and sEH data, but it does not help due to too vast space of unknown. I investigated 2d-UMAP of my best model (EGAT) and found that my model mostly have no clue on non-triazine and non-shared BB data.\n\nSimilar to many participants, my [1-fold Molformer alone would get gold](https://www.kaggle.com/competitions/leash-BELKA/discussion/518943#2912665), but all other models did not generalize well. So, averaging on all models overfits the public LB.  \n\n**What I learned for this short-term competition** still a rookie mistake, though I already have some experience. I should have built 2 separated CVs to measure both public and private LB.\n\n**Total cost** I rent an RTX-4090 machine so that I can trained 98M data in the last 2 weeks after all the pipelines are successfully built on Kaggle. Total cost was around 110 USD.\n\n---\n\n## Looking forward on the long-term goal of Belka\n\n- Would love to understand more on \"non-triazine\" and \"non-shared BB\" performance. [As the host mentioned](https://www.kaggle.com/competitions/leash-BELKA/discussion/503232), the private LB consists of 3 parts \"non-triazine\", \"shared BB\" and \"non-shared BB\". \n\t\nTherefore, **what if we eliminate the shared-BB contribution and focus solely on the 2 unknowns??** \nWhat are the best model scores on this competition regarding this? Is the best model even reach 0.1 on the total unknown? \n@andrewdblevins, I would be much appreciate if you can provide insights on this aspect.\n\n**Possible future direction of dataset:** As we known, physics of binding are based on surface interaction of the target protein and ligands. Is it possible/practical to design dataset / goal to predict this 3D information?\n\nIf we are still predict targets with classification/regression, it would needs billions of training data to generalize over different ranges of BBs and molecules-cores, similar to the scale of training data used by ChatGPT and Diffusion models.",
    "2912529": "Thanks for hosting this competition and providing the community with this dataset! It was a fun comp and the first time I've had my hands on this much chemical data. I'm looking forward to seeing folks write-ups.",
    "2976609": "@andrewdblevins \n\nnvidia has just make a prediction of small molecule binding affinity perdiction.\nThis may be of interested to you!\n\nNVIDIA\nGenerative Virtual Screening for Drug Discovery: Search and optimize a library of small molecules to identify chemical structures that bind to a target protein.\nhttps://build.nvidia.com/nvidia/generative-virtual-screening-for-drug-discovery\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F90c1b11078002f119b2e6a5aa38b3f44%2FSelection_999(5990).png?generation=1725258055052233&alt=media)",
    "2914509": "Are any of the models and weights useful for you as the competition organizer?",
    "2918888": "From Kaggle we would also like to thank everyone for participating!\n\nThis competition ended with 9143 registrations and 2357 participants on 1946 teams. We had 37022 submissions from 91 countries. For 500 users 21 in the top 100!), this was their first competition. Thank you all for your hard work in this competition and congratulations to our winners and to those who gained a new ranking!  \n\nThe top potential winning teams will be contacted via email for the next steps. We look forward to learning more about their winning solutions.\n\nWe've cleaned the leaderboard and disqualified some teams that have violated the rules. If you think you were removed by mistake, or believe you have evidence that suggests another team cheated, please contact [compliance](https://www.kaggle.com/compliance). Please fill in all the fields honestly.\n\nWe highly encourage you to post a solution write-up about your approach and solution in the forums (see [instructions](https://www.kaggle.com/discussions/product-feedback/373153)). You may also refer to [Kaggle Solution Write-Up Documentation](https://www.kaggle.com/solution-write-up-documentation) for guidance. You are  also encouraged to [publish your models on Kaggle Models](https://www.kaggle.com/docs/models#publishing-a-model)!\n\nThanks for continuing to make Kaggle a great place to learn, practice, and test data science techniques!\n\nHappy Modeling!\n",
    "2912801": "Greetings! I wonder if you can share per-target group scores in format of .csv for both public and private? Would be fun to analyze it. ",
    "2912986": "What of the top student prize?",
    "2912582": "Thanks for the debrief.\n\nThis was my first Kaggle competition and I am hooked.\n\nIt's such a thrill to navigate all the possible paths to victory, and then engineer your way through the mechanics.\n\nHaving NO domain expertise made this a double thrill. I learned a lot in the 20 or so days I worked on this.\n\nCitizen science is rad.",
    "2912531": "With how far we have to go, I wonder if a future competition aiming predict three (not as hard) buckets would, ironically, do better at advancing generality?\n\n* Shared BBs (1, 2, 3) -> (1, 2, 3)\n* 2/3rds shared (1, 2, 3) -> (4, 2, 3)\n* 1/3rd shared (1, 2, 3) -> (1, 5, 6)\n\nThis would force competitors to incrementally figure out how to predict unknown substitutions.\n\nI also mentioned in passing at some point, but a metric focusing on 'top 1%' predictions or something similar might also be better. We know from this competition we can't effectively predict most binding events on unseen data, but with this metric we don't really see whether or not we can predict a fairly small pool of novel candidates from which we'll get a high percentage of binding events.\n\nAnother interesting prediction task would be: given train data with these 1000 building blocks, and test data comprising of 500 or 1000 other building blocks, rank each building block in test by how many binding combinations it will hit with the other bbs in the test dataset. Or instead of rank, pick the 500 out of the 1000 that will give the most binding hits or something like that. Maybe that all depends on the business goals, I guess.",
    "2956957": "Thank you for hosting this competition! It would be really great to continue improving our models with the ground truth data for test set and read counts would be amazing as well. Do you have a timeline for when these might become available?",
    "2916344": "Our team has tried sincerely to solve the problem, but we were removed from the leaderboard. Organizers, could you please tell us the reason? The models used to create the submission files included Xgboost, 1DCNN, Chemberta, and MoLFormer.\n\nPS: I have sent an email to the Competitions Compliance Team regarding this.",
    "2914916": "It has been mentioned in different ways, but it would be very interesting and useful to have not only the target values for the public and private sets, but also the various model predictions, if that's possible.  A .csv of the test compound ids with columns of the actual classification and as many model predictions the is practical could be very illuminating regarding generalization and target-specific effects.",
    "2913104": "Thanks @andrewdblevins for this really challenging dataset.\nI look forward to the real read counts and possibly also 'no-target selection' results for control. This should be a more realistic dataset and might even allow out-of-domain prediction (albeit restricted to the training diversity) from just the competition dataset w/o also using external activity reports from, say, the ZINC database.",
    "2916616": "",
    "2914125": "Very nice! Thanks for the brief!"
  }
}