{
  "id": 519815,
  "title": "8th place private solution - a single SMILES-based transformer model is all we need. And a proper benchmarking",
  "url": "/competitions/leash-BELKA/writeups/vlad-vinogradov-8th-place-private-solution-a-singl",
  "author_name": "",
  "post_date": "2024-07-16T19:51:02.803Z",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/overview</a>  <br>\nData context: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/data\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/data</a></p>\n<h1>Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2Fac5a96eb60939a87e40bfdc3a0bfd59c%2Fleash-bio-solution.png?generation=1720838439586161&amp;alt=media\"><br>\nA summary of the solution: ligand-based, SMILES-only, multi-target transformer model trained independently on shared and non-shared building blocks splits. A special hit identification data splitting technique, and k-fold training and validation are used for non-shared blocks. The models are independently averaged for the shared and non-shared parts of the test set.</p>\n<h1>Data splits</h1>\n<p>This competition hinges on the fact that similar molecules tend to have similar properties - this is the main assumption in lead optimization of initially identified hits. However, the correct evaluation strategy is rarely applied, which is why <a href=\"https://polarishub.io/\" target=\"_blank\">Polaris</a>, a collective effort to establish unbiased benchmarks for drug discovery, was launched recently. Moreover, it's hard to say what is correct for a specific task. Yet, it was shown recently by my colleague Simon Steshin that a splitting strategy by 0.4 Tanimoto similarity should be considered for novel hits identification (<a href=\"https://arxiv.org/abs/2310.06399\" target=\"_blank\">Lo-Hi: Practical ML Drug Discovery Benchmark</a>). Perhaps we can apply a similar approach to build a more generalizable model in other tasks, particularly in this competition where 66% of the private score is dedicated to the out-of-distribution molecules.</p>\n<h2>Non-shared BBs</h2>\n<p>The molecules provided by the organizers consist of three building blocks, so we may consider splitting by the similarity of building blocks. Let’s examine the distribution of test building blocks to training building blocks:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F4348e00ecff143d7bcc71fa29175851a%2Ftest2train_bb_tanimoto_dists.png?generation=1720837870293702&amp;alt=media\" alt=\"Distribution of Maximum Tanimoto Similarities Between Test and Training Building Blocks\"></p>\n<p>There are many building blocks shared between train and test sets (Max Tanimoto is 1.0), and we obviously need to avoid the same occurrence in our non-shared train/valid split. We see that BB1 doesn't intersect with BB2, and BB2 significantly intersects with BB3. Besides the clear duplicates, there are many similar building blocks, which, as per the above assumption, will introduce a strong bias into our evaluation.</p>\n<h3>Hit identification split by building blocks (HIBB)</h3>\n<p>So, let's eliminate this bias. We will use the Hi splitting algorithm from the Lo-Hi benchmark which solves a Balanced Vertex Minimum k-Cut problem to construct such training and validation sets that the closest molecules between the two sets will be at least the cutoff Tanimoto distance apart (<a href=\"https://github.com/SteshinSS/lohi_splitter\" target=\"_blank\">code</a>). To decide what threshold to use, let's look at the training set's Tanimoto similarity distributions per block position, considering only the same positions (we will deal with BB2-BB3 cross-positions later):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F47e417d61abf31653b9885b3c3044426%2Ftrain2train_bb_tanimoto_dists.png?generation=1720837978207046&amp;alt=media\" alt=\"Distribution of Maximum Tanimoto Similarities Between Training Building Blocks To Each Other (excl. the same)\"></p>\n<p>It's an open question of what threshold to set, and most likely, using multiple ones is the best approach. I chose the thresholds close to the mean of the per-position similarities (don't ask me why): 0.7 for BB1&lt;-&gt;BB2 and 0.4 for (BB2+BB3&lt;-&gt;BB2+BB3). Yes, I merged BB2 with BB3, as they are known to intersect each other. Ok, this split is the hardest one. Training models on it will quickly show that nothing is working (in other words, structure similarity drives a lot of correlation with activity).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F938d225e60ef54b163b8be07b2a5b688%2Fleash-bio-solution-HIBB-split.png?generation=1720838534349396&amp;alt=media\" alt=\"HIBB split: 0.4 Tanimoto dissimilarity split by building blocks\"></p>\n<p>Pay attention to the exclusivity of building blocks - if a block appears in one set, it cannot appear in another set in any possible combination.</p>\n<h3>Weighted building blocks splits (BB)</h3>\n<p>Although the Hi split is useful, it might be over-pessimistic for the competition's problem, and it removes a lot of data due to the exclusive assignment of building blocks to the train/valid sets. So, I decided to create a simpler split close to what other participants did when replicating the host's split but different in a few aspects:</p>\n<ul>\n<li>The blocks are again exclusive to the train/valid sets</li>\n<li>There's an increased amount of blocks at each position by 3 times compared to the BB split from the great <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\" target=\"_blank\">notebook</a> of <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> and work of both <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">here</a></li>\n<li>There are rolling 5 folds where, for each fold, different blocks are selected at each position</li>\n<li>Training data consists of a mix of <code>any positive</code> molecules (the ones that bind to at least one target) and the same amount of <code>all negative</code> molecules (no binding to all targets) randomly sampled from the corresponding fold's building blocks</li>\n<li>Validation data consists of a mix of what's left from the <code>any positive</code> set and an amount of <code>all negative</code> compounds to approximately match the imbalance of the original training dataset, i.e., <code>any positive</code> rate is ~1.5%</li>\n</ul>\n<p>Overall, the folds look like this (randomization not taken into account for simplicity):</p>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>BB1</th>\n<th>BB2&amp;BB3</th>\n<th>BB3\\BB2</th>\n<th>train size</th>\n<th>train pos rate</th>\n<th>valid size</th>\n<th>valid pos rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0-50</td>\n<td>0-101</td>\n<td>0-5</td>\n<td>1 875 494</td>\n<td>50%</td>\n<td>356 620</td>\n<td>1.67%</td>\n</tr>\n<tr>\n<td>2</td>\n<td>51-101</td>\n<td>102-203</td>\n<td>6-11</td>\n<td>1 777 354</td>\n<td>50%</td>\n<td>406 629</td>\n<td>1.62%</td>\n</tr>\n<tr>\n<td>3</td>\n<td>102-152</td>\n<td>204-305</td>\n<td>12-17</td>\n<td>1 430 522</td>\n<td>50%</td>\n<td>363 503</td>\n<td>2.96%</td>\n</tr>\n<tr>\n<td>4</td>\n<td>153-203</td>\n<td>306-407</td>\n<td>18-23</td>\n<td>1 924 372</td>\n<td>50%</td>\n<td>311 943</td>\n<td>1.47%</td>\n</tr>\n<tr>\n<td>5</td>\n<td>204-254</td>\n<td>408-509</td>\n<td>24-29</td>\n<td>1 903 992</td>\n<td>50%</td>\n<td>313 838</td>\n<td>1.47%</td>\n</tr>\n</tbody>\n</table>\n<p>As with the HIBB split, the validation set is of similar size per fold, but we now have much more training data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F3d2b650b86bb046c919b1cf20ec2418e%2Fleash-bio-solution-BB-split.png?generation=1720838175494292&amp;alt=media\" alt=\"BB split: weighted building blocks splits\"></p>\n<h2>Shared BBs</h2>\n<p>The strategy for the building blocks shared between train and test sets is clear—we just need to overfit to the known building blocks while covering as much data as possible. So, here, I used all 98M molecules, of which I randomly selected 1% for validation and 99% for training.</p>\n<h1>Models and training</h1>\n<p>To put it simply, the training strategies I used are the following:</p>\n<ul>\n<li><strong>shared BBs</strong>: overfitting while tracking performance on validation, taking the last checkpoint</li>\n<li><strong>non-shared BBs</strong>: taking the best checkpoint on the validation set (the same strategy for both HIBB and BB splits)</li>\n</ul>\n<p>I tried A LOT of models and performed extensive hyperparameters tuning (especially on the non-shared BBs) for the following models:</p>\n<ul>\n<li><a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">MolFormer</a>, <a href=\"https://huggingface.co/entropy/roberta_zinc_480m\" target=\"_blank\">RoBERTa ZINC</a>, <a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">ChemBERTa</a>: all showed similar quality with MolFormer and a custom RoBERTa performing better in an ensemble</li>\n<li>GPS++: SOTA full-attention GNN with atoms and bonds features (impl. adopted from <a href=\"https://github.com/datamol-io/graphium\" target=\"_blank\">Graphium</a>). For atoms, the following features are used: atomic-number, group, period, total-valence, degree, formal-charge, radical-electron, hybridization, chirality, implicit-valence, num_h_atoms, aromatic, in-ring, electronegativity. For bonds: bond-type-onehot, stereo, in-ring</li>\n<li>MPNN++: GNN with the same atoms and bonds features (impl. adopted from <a href=\"https://github.com/datamol-io/graphium\" target=\"_blank\">Graphium</a>), see performance in the <a href=\"https://arxiv.org/abs/2404.11568\" target=\"_blank\">scaling GNNs paper</a>. In my case, MPNN++ also performed substantially better than GPS++, potentially due to not attending everything to everything, thus having a stronger inductive bias</li>\n<li>Tanimoto similarity scores of building blocks to the top-50 closest molecules put together in a multi-head attention model where queries/keys are the similarities and values are the ground truth values. It was supposed that this kind of a model (I call it the <code>SimAttn</code> model) could learn from the closest ranked list of molecules, and it did, but I wasn't able to reach high enough AP with it</li>\n<li>XGBoost on RDKit 210 descriptors normalized with standard scaler and ECFP4</li>\n<li>MLP on RDKit + ECFP4 + similarity features</li>\n<li>Meta MLP model on RDKit + ECFP4 + similarity feats + best performing MolFormer: slightly improved performance on BB split, but not enough to convince me to submit</li>\n</ul>\n<p>From the above, the best performing models are the fine-tuned MolFormer (45M parameters) and a custom RoBERTa (~8M parameters) as described in <a href=\"https://arxiv.org/abs/2406.14572\" target=\"_blank\">my company's recent paper</a>. The customization is mainly related to the 500-sized BPE SMILES tokenizer, 15% of masked tokens in 30% of cases, and SMILES re-enumeration in 50% of cases.</p>\n<h1>Conclusion</h1>\n<p>My take on generalizability to unseen chemical subspaces:</p>\n<ol>\n<li>This isn't a solved problem, not only for this competition but everywhere in the public domain</li>\n<li>We can address the issue with extensive benchmarking</li>\n<li>SMILES-based transformer models perform slightly better than the SOTA GNN models</li>\n<li>Scaling number of parameters <strong>improves</strong> the quality on out-of-distribution molecules</li>\n<li>Scaling SMILES-based transformers is way easier than GNNs due to the absence of pre-processing, so scaling to billion-size models and datasets to improve generalizability should be considered in the follow-up research</li>\n</ol>\n<h1>Acknowledgements</h1>\n<p>I thank Copilot and Continue.dev for being my coding teammates all the time. I also acknowledge my colleague Simon for his public work on molecule benchmarking. I thank other participants for their solutions and discussions; I didn't find time to contribute to the discussions, but I was aligned with many of them. Huge thanks to Leash Bio and Kaggle for organizing the competition. And I'm sorry for the teams shuffled on the private leaderboard. I had a solution for 0.299 ten days before the end of the competition, so I expect the same to happen to many other teams. I believe they all tried hard to crack this unsolved but intriguing small molecules generalizability problem.</p>",
  "messages": [
    {
      "id": "2919585",
      "postDate": "07/13/2024 02:46:10",
      "content": "<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/overview</a>  <br>\nData context: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/data\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/data</a></p>\n<h1>Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2Fac5a96eb60939a87e40bfdc3a0bfd59c%2Fleash-bio-solution.png?generation=1720838439586161&amp;alt=media\"><br>\nA summary of the solution: ligand-based, SMILES-only, multi-target transformer model trained independently on shared and non-shared building blocks splits. A special hit identification data splitting technique, and k-fold training and validation are used for non-shared blocks. The models are independently averaged for the shared and non-shared parts of the test set.</p>\n<h1>Data splits</h1>\n<p>This competition hinges on the fact that similar molecules tend to have similar properties - this is the main assumption in lead optimization of initially identified hits. However, the correct evaluation strategy is rarely applied, which is why <a href=\"https://polarishub.io/\" target=\"_blank\">Polaris</a>, a collective effort to establish unbiased benchmarks for drug discovery, was launched recently. Moreover, it's hard to say what is correct for a specific task. Yet, it was shown recently by my colleague Simon Steshin that a splitting strategy by 0.4 Tanimoto similarity should be considered for novel hits identification (<a href=\"https://arxiv.org/abs/2310.06399\" target=\"_blank\">Lo-Hi: Practical ML Drug Discovery Benchmark</a>). Perhaps we can apply a similar approach to build a more generalizable model in other tasks, particularly in this competition where 66% of the private score is dedicated to the out-of-distribution molecules.</p>\n<h2>Non-shared BBs</h2>\n<p>The molecules provided by the organizers consist of three building blocks, so we may consider splitting by the similarity of building blocks. Let’s examine the distribution of test building blocks to training building blocks:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F4348e00ecff143d7bcc71fa29175851a%2Ftest2train_bb_tanimoto_dists.png?generation=1720837870293702&amp;alt=media\" alt=\"Distribution of Maximum Tanimoto Similarities Between Test and Training Building Blocks\"></p>\n<p>There are many building blocks shared between train and test sets (Max Tanimoto is 1.0), and we obviously need to avoid the same occurrence in our non-shared train/valid split. We see that BB1 doesn't intersect with BB2, and BB2 significantly intersects with BB3. Besides the clear duplicates, there are many similar building blocks, which, as per the above assumption, will introduce a strong bias into our evaluation.</p>\n<h3>Hit identification split by building blocks (HIBB)</h3>\n<p>So, let's eliminate this bias. We will use the Hi splitting algorithm from the Lo-Hi benchmark which solves a Balanced Vertex Minimum k-Cut problem to construct such training and validation sets that the closest molecules between the two sets will be at least the cutoff Tanimoto distance apart (<a href=\"https://github.com/SteshinSS/lohi_splitter\" target=\"_blank\">code</a>). To decide what threshold to use, let's look at the training set's Tanimoto similarity distributions per block position, considering only the same positions (we will deal with BB2-BB3 cross-positions later):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F47e417d61abf31653b9885b3c3044426%2Ftrain2train_bb_tanimoto_dists.png?generation=1720837978207046&amp;alt=media\" alt=\"Distribution of Maximum Tanimoto Similarities Between Training Building Blocks To Each Other (excl. the same)\"></p>\n<p>It's an open question of what threshold to set, and most likely, using multiple ones is the best approach. I chose the thresholds close to the mean of the per-position similarities (don't ask me why): 0.7 for BB1&lt;-&gt;BB2 and 0.4 for (BB2+BB3&lt;-&gt;BB2+BB3). Yes, I merged BB2 with BB3, as they are known to intersect each other. Ok, this split is the hardest one. Training models on it will quickly show that nothing is working (in other words, structure similarity drives a lot of correlation with activity).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F938d225e60ef54b163b8be07b2a5b688%2Fleash-bio-solution-HIBB-split.png?generation=1720838534349396&amp;alt=media\" alt=\"HIBB split: 0.4 Tanimoto dissimilarity split by building blocks\"></p>\n<p>Pay attention to the exclusivity of building blocks - if a block appears in one set, it cannot appear in another set in any possible combination.</p>\n<h3>Weighted building blocks splits (BB)</h3>\n<p>Although the Hi split is useful, it might be over-pessimistic for the competition's problem, and it removes a lot of data due to the exclusive assignment of building blocks to the train/valid sets. So, I decided to create a simpler split close to what other participants did when replicating the host's split but different in a few aspects:</p>\n<ul>\n<li>The blocks are again exclusive to the train/valid sets</li>\n<li>There's an increased amount of blocks at each position by 3 times compared to the BB split from the great <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/\" target=\"_blank\">notebook</a> of <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> and work of both <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496576\" target=\"_blank\">here</a></li>\n<li>There are rolling 5 folds where, for each fold, different blocks are selected at each position</li>\n<li>Training data consists of a mix of <code>any positive</code> molecules (the ones that bind to at least one target) and the same amount of <code>all negative</code> molecules (no binding to all targets) randomly sampled from the corresponding fold's building blocks</li>\n<li>Validation data consists of a mix of what's left from the <code>any positive</code> set and an amount of <code>all negative</code> compounds to approximately match the imbalance of the original training dataset, i.e., <code>any positive</code> rate is ~1.5%</li>\n</ul>\n<p>Overall, the folds look like this (randomization not taken into account for simplicity):</p>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>BB1</th>\n<th>BB2&amp;BB3</th>\n<th>BB3\\BB2</th>\n<th>train size</th>\n<th>train pos rate</th>\n<th>valid size</th>\n<th>valid pos rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0-50</td>\n<td>0-101</td>\n<td>0-5</td>\n<td>1 875 494</td>\n<td>50%</td>\n<td>356 620</td>\n<td>1.67%</td>\n</tr>\n<tr>\n<td>2</td>\n<td>51-101</td>\n<td>102-203</td>\n<td>6-11</td>\n<td>1 777 354</td>\n<td>50%</td>\n<td>406 629</td>\n<td>1.62%</td>\n</tr>\n<tr>\n<td>3</td>\n<td>102-152</td>\n<td>204-305</td>\n<td>12-17</td>\n<td>1 430 522</td>\n<td>50%</td>\n<td>363 503</td>\n<td>2.96%</td>\n</tr>\n<tr>\n<td>4</td>\n<td>153-203</td>\n<td>306-407</td>\n<td>18-23</td>\n<td>1 924 372</td>\n<td>50%</td>\n<td>311 943</td>\n<td>1.47%</td>\n</tr>\n<tr>\n<td>5</td>\n<td>204-254</td>\n<td>408-509</td>\n<td>24-29</td>\n<td>1 903 992</td>\n<td>50%</td>\n<td>313 838</td>\n<td>1.47%</td>\n</tr>\n</tbody>\n</table>\n<p>As with the HIBB split, the validation set is of similar size per fold, but we now have much more training data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F3d2b650b86bb046c919b1cf20ec2418e%2Fleash-bio-solution-BB-split.png?generation=1720838175494292&amp;alt=media\" alt=\"BB split: weighted building blocks splits\"></p>\n<h2>Shared BBs</h2>\n<p>The strategy for the building blocks shared between train and test sets is clear—we just need to overfit to the known building blocks while covering as much data as possible. So, here, I used all 98M molecules, of which I randomly selected 1% for validation and 99% for training.</p>\n<h1>Models and training</h1>\n<p>To put it simply, the training strategies I used are the following:</p>\n<ul>\n<li><strong>shared BBs</strong>: overfitting while tracking performance on validation, taking the last checkpoint</li>\n<li><strong>non-shared BBs</strong>: taking the best checkpoint on the validation set (the same strategy for both HIBB and BB splits)</li>\n</ul>\n<p>I tried A LOT of models and performed extensive hyperparameters tuning (especially on the non-shared BBs) for the following models:</p>\n<ul>\n<li><a href=\"https://huggingface.co/ibm/MoLFormer-XL-both-10pct\" target=\"_blank\">MolFormer</a>, <a href=\"https://huggingface.co/entropy/roberta_zinc_480m\" target=\"_blank\">RoBERTa ZINC</a>, <a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">ChemBERTa</a>: all showed similar quality with MolFormer and a custom RoBERTa performing better in an ensemble</li>\n<li>GPS++: SOTA full-attention GNN with atoms and bonds features (impl. adopted from <a href=\"https://github.com/datamol-io/graphium\" target=\"_blank\">Graphium</a>). For atoms, the following features are used: atomic-number, group, period, total-valence, degree, formal-charge, radical-electron, hybridization, chirality, implicit-valence, num_h_atoms, aromatic, in-ring, electronegativity. For bonds: bond-type-onehot, stereo, in-ring</li>\n<li>MPNN++: GNN with the same atoms and bonds features (impl. adopted from <a href=\"https://github.com/datamol-io/graphium\" target=\"_blank\">Graphium</a>), see performance in the <a href=\"https://arxiv.org/abs/2404.11568\" target=\"_blank\">scaling GNNs paper</a>. In my case, MPNN++ also performed substantially better than GPS++, potentially due to not attending everything to everything, thus having a stronger inductive bias</li>\n<li>Tanimoto similarity scores of building blocks to the top-50 closest molecules put together in a multi-head attention model where queries/keys are the similarities and values are the ground truth values. It was supposed that this kind of a model (I call it the <code>SimAttn</code> model) could learn from the closest ranked list of molecules, and it did, but I wasn't able to reach high enough AP with it</li>\n<li>XGBoost on RDKit 210 descriptors normalized with standard scaler and ECFP4</li>\n<li>MLP on RDKit + ECFP4 + similarity features</li>\n<li>Meta MLP model on RDKit + ECFP4 + similarity feats + best performing MolFormer: slightly improved performance on BB split, but not enough to convince me to submit</li>\n</ul>\n<p>From the above, the best performing models are the fine-tuned MolFormer (45M parameters) and a custom RoBERTa (~8M parameters) as described in <a href=\"https://arxiv.org/abs/2406.14572\" target=\"_blank\">my company's recent paper</a>. The customization is mainly related to the 500-sized BPE SMILES tokenizer, 15% of masked tokens in 30% of cases, and SMILES re-enumeration in 50% of cases.</p>\n<h1>Conclusion</h1>\n<p>My take on generalizability to unseen chemical subspaces:</p>\n<ol>\n<li>This isn't a solved problem, not only for this competition but everywhere in the public domain</li>\n<li>We can address the issue with extensive benchmarking</li>\n<li>SMILES-based transformer models perform slightly better than the SOTA GNN models</li>\n<li>Scaling number of parameters <strong>improves</strong> the quality on out-of-distribution molecules</li>\n<li>Scaling SMILES-based transformers is way easier than GNNs due to the absence of pre-processing, so scaling to billion-size models and datasets to improve generalizability should be considered in the follow-up research</li>\n</ol>\n<h1>Acknowledgements</h1>\n<p>I thank Copilot and Continue.dev for being my coding teammates all the time. I also acknowledge my colleague Simon for his public work on molecule benchmarking. I thank other participants for their solutions and discussions; I didn't find time to contribute to the discussions, but I was aligned with many of them. Huge thanks to Leash Bio and Kaggle for organizing the competition. And I'm sorry for the teams shuffled on the private leaderboard. I had a solution for 0.299 ten days before the end of the competition, so I expect the same to happen to many other teams. I believe they all tried hard to crack this unsolved but intriguing small molecules generalizability problem.</p>",
      "rawMarkdown": "# Context\nBusiness context: https://www.kaggle.com/competitions/leash-BELKA/overview  \nData context: https://www.kaggle.com/competitions/leash-BELKA/data\n\n# Summary\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2Fac5a96eb60939a87e40bfdc3a0bfd59c%2Fleash-bio-solution.png?generation=1720838439586161&alt=media)\nA summary of the solution: ligand-based, SMILES-only, multi-target transformer model trained independently on shared and non-shared building blocks splits. A special hit identification data splitting technique, and k-fold training and validation are used for non-shared blocks. The models are independently averaged for the shared and non-shared parts of the test set.\n\n# Data splits\nThis competition hinges on the fact that similar molecules tend to have similar properties - this is the main assumption in lead optimization of initially identified hits. However, the correct evaluation strategy is rarely applied, which is why [Polaris](https://polarishub.io/), a collective effort to establish unbiased benchmarks for drug discovery, was launched recently. Moreover, it's hard to say what is correct for a specific task. Yet, it was shown recently by my colleague Simon Steshin that a splitting strategy by 0.4 Tanimoto similarity should be considered for novel hits identification ([Lo-Hi: Practical ML Drug Discovery Benchmark](https://arxiv.org/abs/2310.06399)). Perhaps we can apply a similar approach to build a more generalizable model in other tasks, particularly in this competition where 66% of the private score is dedicated to the out-of-distribution molecules.\n\n## Non-shared BBs\nThe molecules provided by the organizers consist of three building blocks, so we may consider splitting by the similarity of building blocks. Let’s examine the distribution of test building blocks to training building blocks:\n\n![Distribution of Maximum Tanimoto Similarities Between Test and Training Building Blocks](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F4348e00ecff143d7bcc71fa29175851a%2Ftest2train_bb_tanimoto_dists.png?generation=1720837870293702&alt=media)\n\nThere are many building blocks shared between train and test sets (Max Tanimoto is 1.0), and we obviously need to avoid the same occurrence in our non-shared train/valid split. We see that BB1 doesn't intersect with BB2, and BB2 significantly intersects with BB3. Besides the clear duplicates, there are many similar building blocks, which, as per the above assumption, will introduce a strong bias into our evaluation.\n\n### Hit identification split by building blocks (HIBB)\nSo, let's eliminate this bias. We will use the Hi splitting algorithm from the Lo-Hi benchmark which solves a Balanced Vertex Minimum k-Cut problem to construct such training and validation sets that the closest molecules between the two sets will be at least the cutoff Tanimoto distance apart ([code](https://github.com/SteshinSS/lohi_splitter)). To decide what threshold to use, let's look at the training set's Tanimoto similarity distributions per block position, considering only the same positions (we will deal with BB2-BB3 cross-positions later):\n\n![Distribution of Maximum Tanimoto Similarities Between Training Building Blocks To Each Other (excl. the same)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F47e417d61abf31653b9885b3c3044426%2Ftrain2train_bb_tanimoto_dists.png?generation=1720837978207046&alt=media)\n\nIt's an open question of what threshold to set, and most likely, using multiple ones is the best approach. I chose the thresholds close to the mean of the per-position similarities (don't ask me why): 0.7 for BB1<->BB2 and 0.4 for (BB2+BB3<->BB2+BB3). Yes, I merged BB2 with BB3, as they are known to intersect each other. Ok, this split is the hardest one. Training models on it will quickly show that nothing is working (in other words, structure similarity drives a lot of correlation with activity).\n\n![HIBB split: 0.4 Tanimoto dissimilarity split by building blocks](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F938d225e60ef54b163b8be07b2a5b688%2Fleash-bio-solution-HIBB-split.png?generation=1720838534349396&alt=media)\n\nPay attention to the exclusivity of building blocks - if a block appears in one set, it cannot appear in another set in any possible combination.\n\n### Weighted building blocks splits (BB)\nAlthough the Hi split is useful, it might be over-pessimistic for the competition's problem, and it removes a lot of data due to the exclusive assignment of building blocks to the train/valid sets. So, I decided to create a simpler split close to what other participants did when replicating the host's split but different in a few aspects:\n- The blocks are again exclusive to the train/valid sets\n- There's an increased amount of blocks at each position by 3 times compared to the BB split from the great [notebook](https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/) of @thedrcat and work of both @roberthatch and @hengck23 [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)\n- There are rolling 5 folds where, for each fold, different blocks are selected at each position\n- Training data consists of a mix of `any positive` molecules (the ones that bind to at least one target) and the same amount of `all negative` molecules (no binding to all targets) randomly sampled from the corresponding fold's building blocks\n- Validation data consists of a mix of what's left from the `any positive` set and an amount of `all negative` compounds to approximately match the imbalance of the original training dataset, i.e., `any positive` rate is ~1.5%\n\nOverall, the folds look like this (randomization not taken into account for simplicity):\n|  Fold | BB1 | BB2&BB3 | BB3\\BB2 | train size | train pos rate | valid size | valid pos rate |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 | 0-50 | 0-101 | 0-5 | 1 875 494 | 50% | 356 620 | 1.67%\n| 2 | 51-101 | 102-203 | 6-11 | 1 777 354 | 50% | 406 629 | 1.62%\n| 3 | 102-152 | 204-305 | 12-17 | 1 430 522 | 50% | 363 503 | 2.96%\n| 4 | 153-203 | 306-407 | 18-23 | 1 924 372 | 50% | 311 943 | 1.47%\n| 5 | 204-254 | 408-509 | 24-29 | 1 903 992 | 50% | 313 838 | 1.47%\n\nAs with the HIBB split, the validation set is of similar size per fold, but we now have much more training data:\n\n![BB split: weighted building blocks splits](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F3d2b650b86bb046c919b1cf20ec2418e%2Fleash-bio-solution-BB-split.png?generation=1720838175494292&alt=media)\n\n## Shared BBs\nThe strategy for the building blocks shared between train and test sets is clear—we just need to overfit to the known building blocks while covering as much data as possible. So, here, I used all 98M molecules, of which I randomly selected 1% for validation and 99% for training.\n\n# Models and training\nTo put it simply, the training strategies I used are the following:\n- **shared BBs**: overfitting while tracking performance on validation, taking the last checkpoint\n- **non-shared BBs**: taking the best checkpoint on the validation set (the same strategy for both HIBB and BB splits)\n\nI tried A LOT of models and performed extensive hyperparameters tuning (especially on the non-shared BBs) for the following models:\n- [MolFormer](https://huggingface.co/ibm/MoLFormer-XL-both-10pct), [RoBERTa ZINC](https://huggingface.co/entropy/roberta_zinc_480m), [ChemBERTa](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR): all showed similar quality with MolFormer and a custom RoBERTa performing better in an ensemble\n- GPS++: SOTA full-attention GNN with atoms and bonds features (impl. adopted from [Graphium](https://github.com/datamol-io/graphium)). For atoms, the following features are used: atomic-number, group, period, total-valence, degree, formal-charge, radical-electron, hybridization, chirality, implicit-valence, num_h_atoms, aromatic, in-ring, electronegativity. For bonds: bond-type-onehot, stereo, in-ring\n- MPNN++: GNN with the same atoms and bonds features (impl. adopted from [Graphium](https://github.com/datamol-io/graphium)), see performance in the [scaling GNNs paper](https://arxiv.org/abs/2404.11568). In my case, MPNN++ also performed substantially better than GPS++, potentially due to not attending everything to everything, thus having a stronger inductive bias\n- Tanimoto similarity scores of building blocks to the top-50 closest molecules put together in a multi-head attention model where queries/keys are the similarities and values are the ground truth values. It was supposed that this kind of a model (I call it the `SimAttn` model) could learn from the closest ranked list of molecules, and it did, but I wasn't able to reach high enough AP with it\n- XGBoost on RDKit 210 descriptors normalized with standard scaler and ECFP4\n- MLP on RDKit + ECFP4 + similarity features\n- Meta MLP model on RDKit + ECFP4 + similarity feats + best performing MolFormer: slightly improved performance on BB split, but not enough to convince me to submit\n\nFrom the above, the best performing models are the fine-tuned MolFormer (45M parameters) and a custom RoBERTa (~8M parameters) as described in [my company's recent paper](https://arxiv.org/abs/2406.14572). The customization is mainly related to the 500-sized BPE SMILES tokenizer, 15% of masked tokens in 30% of cases, and SMILES re-enumeration in 50% of cases.\n\n# Conclusion\nMy take on generalizability to unseen chemical subspaces:\n1. This isn't a solved problem, not only for this competition but everywhere in the public domain\n2. We can address the issue with extensive benchmarking\n3. SMILES-based transformer models perform slightly better than the SOTA GNN models\n4. Scaling number of parameters **improves** the quality on out-of-distribution molecules\n5. Scaling SMILES-based transformers is way easier than GNNs due to the absence of pre-processing, so scaling to billion-size models and datasets to improve generalizability should be considered in the follow-up research\n\n# Acknowledgements\nI thank Copilot and Continue.dev for being my coding teammates all the time. I also acknowledge my colleague Simon for his public work on molecule benchmarking. I thank other participants for their solutions and discussions; I didn't find time to contribute to the discussions, but I was aligned with many of them. Huge thanks to Leash Bio and Kaggle for organizing the competition. And I'm sorry for the teams shuffled on the private leaderboard. I had a solution for 0.299 ten days before the end of the competition, so I expect the same to happen to many other teams. I believe they all tried hard to crack this unsolved but intriguing small molecules generalizability problem.",
      "votes": null
    },
    {
      "id": "2919777",
      "postDate": "07/13/2024 07:39:46",
      "content": "<p>Congratulations on the gold, it's a very interesting splitting approach, thanks for sharing!</p>\n<p>And what was the difference in the results with random splitting by BBs (e.g. holding a subset of BBs for validation) vs. splitting by similarity?? </p>\n<p>How much did the splitting you used improve the LB score for the non-shared part (private and public)?</p>",
      "rawMarkdown": "Congratulations on the gold, it's a very interesting splitting approach, thanks for sharing!\n\nAnd what was the difference in the results with random splitting by BBs (e.g. holding a subset of BBs for validation) vs. splitting by similarity?? \n\nHow much did the splitting you used improve the LB score for the non-shared part (private and public)?",
      "votes": null
    },
    {
      "id": "2920275",
      "postDate": "07/13/2024 14:50:33",
      "content": "<p>This solution highlights an innovative approach using a SMILES-based transformer model for the competition.. The strategy of training independently on shared and non-shared building block splits, alongside the use of a special hit identification data splitting technique (HIBB), appears robust for handling out-of-distribution molecules effectively…The extensive model exploration and hyperparameter tuning efforts are commendable, showcasing a thorough investigation into optimizing model performance across different molecular subspaces.. Overall.. this solution sets a high standard for future research in small molecule drug discovery…!!!!</p>",
      "rawMarkdown": "This solution highlights an innovative approach using a SMILES-based transformer model for the competition.. The strategy of training independently on shared and non-shared building block splits, alongside the use of a special hit identification data splitting technique (HIBB), appears robust for handling out-of-distribution molecules effectively...The extensive model exploration and hyperparameter tuning efforts are commendable, showcasing a thorough investigation into optimizing model performance across different molecular subspaces.. Overall.. this solution sets a high standard for future research in small molecule drug discovery...!!!!",
      "votes": null
    },
    {
      "id": "2923450",
      "postDate": "07/15/2024 22:53:00",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> ! Thanks for the great question.<br>\nI had to run some additional submissions to have a clear answer. Look at these results - they perfectly show that the HIBB split dominates for the non-shared BBs while it's the worst split for the shared BBs (due to underfitting and early stopping). And vice versa, random split with its overfitting on all data performs the worst on non-shared BBs but is best on shared BBs. And the BB split is somewhere between both scenarios.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Split type</th>\n<th>Train data</th>\n<th>Private LB (non-shared)</th>\n<th>Public LB (non-shared)</th>\n<th>Private LB (shared)</th>\n<th>Public LB (shared)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MolFormer (10 epochs, class weights = neg frac)</td>\n<td>random</td>\n<td>99% of all 98M dataset (random)</td>\n<td>0.04688</td>\n<td>0.05891</td>\n<td><strong>0.19826</strong></td>\n<td><strong>0.30954</strong></td>\n</tr>\n<tr>\n<td>MolFormer x 5 fold ensemble: avg (each best on valid)</td>\n<td>BB split</td>\n<td>~1.8M pos/neg balance (see solution)</td>\n<td>0.07169</td>\n<td>0.09743</td>\n<td>0.13066</td>\n<td>0.20767</td>\n</tr>\n<tr>\n<td>MolFormer (best on valid per target)</td>\n<td>HIBB split</td>\n<td>~362k (pos rate ~35%)</td>\n<td><strong>0.08100</strong></td>\n<td><strong>0.10530</strong></td>\n<td>0.05856</td>\n<td>0.08926</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hey @antoninadolgorukova ! Thanks for the great question.\nI had to run some additional submissions to have a clear answer. Look at these results - they perfectly show that the HIBB split dominates for the non-shared BBs while it's the worst split for the shared BBs (due to underfitting and early stopping). And vice versa, random split with its overfitting on all data performs the worst on non-shared BBs but is best on shared BBs. And the BB split is somewhere between both scenarios.\n\n| Model | Split type | Train data | Private LB (non-shared) | Public LB (non-shared) | Private LB (shared) | Public LB (shared) |\n| --- | --- | --- | --- | --- |\n| MolFormer (10 epochs, class weights = neg frac) | random | 99% of all 98M dataset (random) | 0.04688 | 0.05891 | **0.19826** | **0.30954** |\n| MolFormer x 5 fold ensemble: avg (each best on valid) | BB split | ~1.8M pos/neg balance (see solution) | 0.07169 | 0.09743 | 0.13066 | 0.20767 |\n| MolFormer (best on valid per target) | HIBB split | ~362k (pos rate ~35%) | **0.08100** | **0.10530** | 0.05856 | 0.08926 |",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2919777,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "07/13/2024 07:39:46",
      "content": "<p>Congratulations on the gold, it's a very interesting splitting approach, thanks for sharing!</p>\n<p>And what was the difference in the results with random splitting by BBs (e.g. holding a subset of BBs for validation) vs. splitting by similarity?? </p>\n<p>How much did the splitting you used improve the LB score for the non-shared part (private and public)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2923450,
          "author_name": "vladvin",
          "author_url": "",
          "post_date": "07/15/2024 22:53:00",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> ! Thanks for the great question.<br>\nI had to run some additional submissions to have a clear answer. Look at these results - they perfectly show that the HIBB split dominates for the non-shared BBs while it's the worst split for the shared BBs (due to underfitting and early stopping). And vice versa, random split with its overfitting on all data performs the worst on non-shared BBs but is best on shared BBs. And the BB split is somewhere between both scenarios.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Split type</th>\n<th>Train data</th>\n<th>Private LB (non-shared)</th>\n<th>Public LB (non-shared)</th>\n<th>Private LB (shared)</th>\n<th>Public LB (shared)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MolFormer (10 epochs, class weights = neg frac)</td>\n<td>random</td>\n<td>99% of all 98M dataset (random)</td>\n<td>0.04688</td>\n<td>0.05891</td>\n<td><strong>0.19826</strong></td>\n<td><strong>0.30954</strong></td>\n</tr>\n<tr>\n<td>MolFormer x 5 fold ensemble: avg (each best on valid)</td>\n<td>BB split</td>\n<td>~1.8M pos/neg balance (see solution)</td>\n<td>0.07169</td>\n<td>0.09743</td>\n<td>0.13066</td>\n<td>0.20767</td>\n</tr>\n<tr>\n<td>MolFormer (best on valid per target)</td>\n<td>HIBB split</td>\n<td>~362k (pos rate ~35%)</td>\n<td><strong>0.08100</strong></td>\n<td><strong>0.10530</strong></td>\n<td>0.05856</td>\n<td>0.08926</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2920275,
      "author_name": "nivestats",
      "author_url": "",
      "post_date": "07/13/2024 14:50:33",
      "content": "<p>This solution highlights an innovative approach using a SMILES-based transformer model for the competition.. The strategy of training independently on shared and non-shared building block splits, alongside the use of a special hit identification data splitting technique (HIBB), appears robust for handling out-of-distribution molecules effectively…The extensive model exploration and hyperparameter tuning efforts are commendable, showcasing a thorough investigation into optimizing model performance across different molecular subspaces.. Overall.. this solution sets a high standard for future research in small molecule drug discovery…!!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2919585": "# Context\nBusiness context: https://www.kaggle.com/competitions/leash-BELKA/overview  \nData context: https://www.kaggle.com/competitions/leash-BELKA/data\n\n# Summary\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2Fac5a96eb60939a87e40bfdc3a0bfd59c%2Fleash-bio-solution.png?generation=1720838439586161&alt=media)\nA summary of the solution: ligand-based, SMILES-only, multi-target transformer model trained independently on shared and non-shared building blocks splits. A special hit identification data splitting technique, and k-fold training and validation are used for non-shared blocks. The models are independently averaged for the shared and non-shared parts of the test set.\n\n# Data splits\nThis competition hinges on the fact that similar molecules tend to have similar properties - this is the main assumption in lead optimization of initially identified hits. However, the correct evaluation strategy is rarely applied, which is why [Polaris](https://polarishub.io/), a collective effort to establish unbiased benchmarks for drug discovery, was launched recently. Moreover, it's hard to say what is correct for a specific task. Yet, it was shown recently by my colleague Simon Steshin that a splitting strategy by 0.4 Tanimoto similarity should be considered for novel hits identification ([Lo-Hi: Practical ML Drug Discovery Benchmark](https://arxiv.org/abs/2310.06399)). Perhaps we can apply a similar approach to build a more generalizable model in other tasks, particularly in this competition where 66% of the private score is dedicated to the out-of-distribution molecules.\n\n## Non-shared BBs\nThe molecules provided by the organizers consist of three building blocks, so we may consider splitting by the similarity of building blocks. Let’s examine the distribution of test building blocks to training building blocks:\n\n![Distribution of Maximum Tanimoto Similarities Between Test and Training Building Blocks](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F4348e00ecff143d7bcc71fa29175851a%2Ftest2train_bb_tanimoto_dists.png?generation=1720837870293702&alt=media)\n\nThere are many building blocks shared between train and test sets (Max Tanimoto is 1.0), and we obviously need to avoid the same occurrence in our non-shared train/valid split. We see that BB1 doesn't intersect with BB2, and BB2 significantly intersects with BB3. Besides the clear duplicates, there are many similar building blocks, which, as per the above assumption, will introduce a strong bias into our evaluation.\n\n### Hit identification split by building blocks (HIBB)\nSo, let's eliminate this bias. We will use the Hi splitting algorithm from the Lo-Hi benchmark which solves a Balanced Vertex Minimum k-Cut problem to construct such training and validation sets that the closest molecules between the two sets will be at least the cutoff Tanimoto distance apart ([code](https://github.com/SteshinSS/lohi_splitter)). To decide what threshold to use, let's look at the training set's Tanimoto similarity distributions per block position, considering only the same positions (we will deal with BB2-BB3 cross-positions later):\n\n![Distribution of Maximum Tanimoto Similarities Between Training Building Blocks To Each Other (excl. the same)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F47e417d61abf31653b9885b3c3044426%2Ftrain2train_bb_tanimoto_dists.png?generation=1720837978207046&alt=media)\n\nIt's an open question of what threshold to set, and most likely, using multiple ones is the best approach. I chose the thresholds close to the mean of the per-position similarities (don't ask me why): 0.7 for BB1<->BB2 and 0.4 for (BB2+BB3<->BB2+BB3). Yes, I merged BB2 with BB3, as they are known to intersect each other. Ok, this split is the hardest one. Training models on it will quickly show that nothing is working (in other words, structure similarity drives a lot of correlation with activity).\n\n![HIBB split: 0.4 Tanimoto dissimilarity split by building blocks](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F938d225e60ef54b163b8be07b2a5b688%2Fleash-bio-solution-HIBB-split.png?generation=1720838534349396&alt=media)\n\nPay attention to the exclusivity of building blocks - if a block appears in one set, it cannot appear in another set in any possible combination.\n\n### Weighted building blocks splits (BB)\nAlthough the Hi split is useful, it might be over-pessimistic for the competition's problem, and it removes a lot of data due to the exclusive assignment of building blocks to the train/valid sets. So, I decided to create a simpler split close to what other participants did when replicating the host's split but different in a few aspects:\n- The blocks are again exclusive to the train/valid sets\n- There's an increased amount of blocks at each position by 3 times compared to the BB split from the great [notebook](https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host/) of @thedrcat and work of both @roberthatch and @hengck23 [here](https://www.kaggle.com/competitions/leash-BELKA/discussion/496576)\n- There are rolling 5 folds where, for each fold, different blocks are selected at each position\n- Training data consists of a mix of `any positive` molecules (the ones that bind to at least one target) and the same amount of `all negative` molecules (no binding to all targets) randomly sampled from the corresponding fold's building blocks\n- Validation data consists of a mix of what's left from the `any positive` set and an amount of `all negative` compounds to approximately match the imbalance of the original training dataset, i.e., `any positive` rate is ~1.5%\n\nOverall, the folds look like this (randomization not taken into account for simplicity):\n|  Fold | BB1 | BB2&BB3 | BB3\\BB2 | train size | train pos rate | valid size | valid pos rate |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 | 0-50 | 0-101 | 0-5 | 1 875 494 | 50% | 356 620 | 1.67%\n| 2 | 51-101 | 102-203 | 6-11 | 1 777 354 | 50% | 406 629 | 1.62%\n| 3 | 102-152 | 204-305 | 12-17 | 1 430 522 | 50% | 363 503 | 2.96%\n| 4 | 153-203 | 306-407 | 18-23 | 1 924 372 | 50% | 311 943 | 1.47%\n| 5 | 204-254 | 408-509 | 24-29 | 1 903 992 | 50% | 313 838 | 1.47%\n\nAs with the HIBB split, the validation set is of similar size per fold, but we now have much more training data:\n\n![BB split: weighted building blocks splits](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F312334%2F3d2b650b86bb046c919b1cf20ec2418e%2Fleash-bio-solution-BB-split.png?generation=1720838175494292&alt=media)\n\n## Shared BBs\nThe strategy for the building blocks shared between train and test sets is clear—we just need to overfit to the known building blocks while covering as much data as possible. So, here, I used all 98M molecules, of which I randomly selected 1% for validation and 99% for training.\n\n# Models and training\nTo put it simply, the training strategies I used are the following:\n- **shared BBs**: overfitting while tracking performance on validation, taking the last checkpoint\n- **non-shared BBs**: taking the best checkpoint on the validation set (the same strategy for both HIBB and BB splits)\n\nI tried A LOT of models and performed extensive hyperparameters tuning (especially on the non-shared BBs) for the following models:\n- [MolFormer](https://huggingface.co/ibm/MoLFormer-XL-both-10pct), [RoBERTa ZINC](https://huggingface.co/entropy/roberta_zinc_480m), [ChemBERTa](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR): all showed similar quality with MolFormer and a custom RoBERTa performing better in an ensemble\n- GPS++: SOTA full-attention GNN with atoms and bonds features (impl. adopted from [Graphium](https://github.com/datamol-io/graphium)). For atoms, the following features are used: atomic-number, group, period, total-valence, degree, formal-charge, radical-electron, hybridization, chirality, implicit-valence, num_h_atoms, aromatic, in-ring, electronegativity. For bonds: bond-type-onehot, stereo, in-ring\n- MPNN++: GNN with the same atoms and bonds features (impl. adopted from [Graphium](https://github.com/datamol-io/graphium)), see performance in the [scaling GNNs paper](https://arxiv.org/abs/2404.11568). In my case, MPNN++ also performed substantially better than GPS++, potentially due to not attending everything to everything, thus having a stronger inductive bias\n- Tanimoto similarity scores of building blocks to the top-50 closest molecules put together in a multi-head attention model where queries/keys are the similarities and values are the ground truth values. It was supposed that this kind of a model (I call it the `SimAttn` model) could learn from the closest ranked list of molecules, and it did, but I wasn't able to reach high enough AP with it\n- XGBoost on RDKit 210 descriptors normalized with standard scaler and ECFP4\n- MLP on RDKit + ECFP4 + similarity features\n- Meta MLP model on RDKit + ECFP4 + similarity feats + best performing MolFormer: slightly improved performance on BB split, but not enough to convince me to submit\n\nFrom the above, the best performing models are the fine-tuned MolFormer (45M parameters) and a custom RoBERTa (~8M parameters) as described in [my company's recent paper](https://arxiv.org/abs/2406.14572). The customization is mainly related to the 500-sized BPE SMILES tokenizer, 15% of masked tokens in 30% of cases, and SMILES re-enumeration in 50% of cases.\n\n# Conclusion\nMy take on generalizability to unseen chemical subspaces:\n1. This isn't a solved problem, not only for this competition but everywhere in the public domain\n2. We can address the issue with extensive benchmarking\n3. SMILES-based transformer models perform slightly better than the SOTA GNN models\n4. Scaling number of parameters **improves** the quality on out-of-distribution molecules\n5. Scaling SMILES-based transformers is way easier than GNNs due to the absence of pre-processing, so scaling to billion-size models and datasets to improve generalizability should be considered in the follow-up research\n\n# Acknowledgements\nI thank Copilot and Continue.dev for being my coding teammates all the time. I also acknowledge my colleague Simon for his public work on molecule benchmarking. I thank other participants for their solutions and discussions; I didn't find time to contribute to the discussions, but I was aligned with many of them. Huge thanks to Leash Bio and Kaggle for organizing the competition. And I'm sorry for the teams shuffled on the private leaderboard. I had a solution for 0.299 ten days before the end of the competition, so I expect the same to happen to many other teams. I believe they all tried hard to crack this unsolved but intriguing small molecules generalizability problem.",
    "2919777": "Congratulations on the gold, it's a very interesting splitting approach, thanks for sharing!\n\nAnd what was the difference in the results with random splitting by BBs (e.g. holding a subset of BBs for validation) vs. splitting by similarity?? \n\nHow much did the splitting you used improve the LB score for the non-shared part (private and public)?",
    "2920275": "This solution highlights an innovative approach using a SMILES-based transformer model for the competition.. The strategy of training independently on shared and non-shared building block splits, alongside the use of a special hit identification data splitting technique (HIBB), appears robust for handling out-of-distribution molecules effectively...The extensive model exploration and hyperparameter tuning efforts are commendable, showcasing a thorough investigation into optimizing model performance across different molecular subspaces.. Overall.. this solution sets a high standard for future research in small molecule drug discovery...!!!!",
    "2923450": "Hey @antoninadolgorukova ! Thanks for the great question.\nI had to run some additional submissions to have a clear answer. Look at these results - they perfectly show that the HIBB split dominates for the non-shared BBs while it's the worst split for the shared BBs (due to underfitting and early stopping). And vice versa, random split with its overfitting on all data performs the worst on non-shared BBs but is best on shared BBs. And the BB split is somewhere between both scenarios.\n\n| Model | Split type | Train data | Private LB (non-shared) | Public LB (non-shared) | Private LB (shared) | Public LB (shared) |\n| --- | --- | --- | --- | --- |\n| MolFormer (10 epochs, class weights = neg frac) | random | 99% of all 98M dataset (random) | 0.04688 | 0.05891 | **0.19826** | **0.30954** |\n| MolFormer x 5 fold ensemble: avg (each best on valid) | BB split | ~1.8M pos/neg balance (see solution) | 0.07169 | 0.09743 | 0.13066 | 0.20767 |\n| MolFormer (best on valid per target) | HIBB split | ~362k (pos rate ~35%) | **0.08100** | **0.10530** | 0.05856 | 0.08926 |"
  },
  "source": "meta"
}