{
  "id": 518993,
  "title": "11th place solution - SSL Pretraining, Multi-models, Multi-representations and Luck",
  "url": "/competitions/leash-BELKA/writeups/ng-11th-place-solution-ssl-pretraining-multi-model",
  "author_name": "",
  "post_date": "2024-07-10T06:59:09.573Z",
  "votes": 33,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First, I would like to thank Kaggle and competition's host for another interesting challenge, which bring me headaches for a long time :D\nThank you to all participants, especially who actively discuss and provide many good insights/code in the forum. As usual on Kaggle, I have a great time learning new things.</p>\n<h2>First impression</h2>\n<ul>\n<li>I trained an Embedding + MLP models on building block id only, e.g [BB1, BB2, BB3] = [100, 200, 300]. Yes, it just see the ID. This model scores shared_AP=65.5 on StratifiedKFold(n_splits=20), which indeed surpass some of my SMILES+CNN1d without looking for what each building block looks like. This strongly indicate that model could memorize target binding and overfiting is nearby.</li>\n<li>Many attempts but actually I failed to setup a trustworthy CV scheme. In such situation, I decided to be blind at all and try to focus the fancy word \"diversity\": Multiple models, multiple input representations and SSL pretraining. I think that strategy indeed survived me in this shuffle competition.</li>\n</ul>\n<h2>How to Cross-Validate?</h2>\n<p>As mentioned above, I don't know.</p>\n<p>My split strategy for each fold:</p>\n<pre><code>1. [11k samples] Hold out a number of building blocks for validation: 17 BB1 + 36 BB2 with positive-fraction-aware balance on each fold.\n2. [78k samples] Leave 20% molecules with most regular scaffolds (&gt; 6116 mols/scaffold) for training, then do a Scaffold Split on the remaining 80% molecules\n3. [103k samples] Stratified Random split on the remaining\n</code></pre>\n<p>Total: 11k non-shared + 181k share, simulate the LB\nThis strategy create a little \"harder\" than pure random split on the shared part.</p>\n<p>The hard part lie in 11k non-shared. Some observation on non-shared CV:</p>\n<ul>\n<li><p>CV scores vary largely between different folds. I find it hard to identify a trend/correlation, or which is work/not work</p></li>\n<li><p>Early stopping should help, but it also hard to answer WHEN? Usually best score is found in first 2 epochs, but also could be much longer\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fa266b63c8b17261dc0768365c16b09ab%2Fnonshare-cv.png?generation=1720515545119485&amp;alt=media\" alt=\"\">\n<em>Each column in form: <code>score (best_epoch)</code>. All results belong to a single split, with nonshare positive rate <code>[0.571, 0.638, 0.714] %</code></em></p></li>\n<li><p>Suitable hyperparams also vary largely between folds. It looks like best setting with best score in fold A could give worst score on fold B :(</p></li>\n<li><p>For one particular fold: score also change significantly due to just small change in hyperparams like random seed, validate interval, .etc. (13.5 -&gt; 32.4, what?)</p></li>\n<li><p>Overfit a model on LB-nonshared-similar building blocks improve nonshared LB score. <code>Building block ECFP6 -&gt; Tanimoto similarity -&gt; Linear Assignment Matching</code> to select most similar/disimilar 170 BB1 + 360 BB2/3, then overfit a 1D-CNN on this subset. LB score for similar=8.8, disimilar=3.9. This indicate that we can improve LB by focusing more on these similar samples. I later select most similar 17 BB1 + 36 BB2/3 as validation set and try to improve CV on this split, but improved CV led to decrease in non-shared LB, which confused me</p></li>\n</ul>\n<p>In conclusion, more and more experiments messed up my mind. The variation is too large and seem to be random for me. I gave up and focus on <strong>SSL pretraining</strong>, improve the diversity by using <strong>multiple models</strong> and <strong>multiple input representations</strong>, and allmost all models were <strong>blindly train on all competition data with no cross-validation</strong>, since I did not trust a single CV split alone and train on multiple CV splits seem to be too time-intensive</p>\n<h2>SSL Pretraining</h2>\n<p>2 pretraining tasks:</p>\n<ul>\n<li>MLM (Masked Language Modeling) with standard setting: 15% masked tokens, 80%-10%-10% replaced with <code>[MASK]</code>/random tokens/keep unchanged.</li>\n<li>MTR (Multi-Task Regression): regress 189 pre-calculated RDKIT Descriptors. These target values are min-max normalized in contrast to CDF Transform in literature.</li>\n</ul>\n<p><code>Joint MTR + MLM pretraining</code> + <code>SMILES Enumeration</code> has success as mentioned in <a href=\"https://github.com/BenevolentAI/MolBERT\" target=\"_blank\">MolBERT</a>, which inspired me to give it a try</p>\n<p>Pretraining using all data (both train + test), with larger sampling weights to test dataset to equalize the frequency of each building blocks and be \"more familiar\" with new domain test dataset.</p>\n<p>I trained 3 models to be finetuned on competition task:</p>\n<ul>\n<li>Squeezeformer + <code>MTR</code></li>\n<li>Squeezeformer + <code>Joint MTR + MLM</code></li>\n<li>Roberta + <code>Joint MTR + MLM</code></li>\n</ul>\n<h2>Modeling</h2>\n<p>All models simultaneously predict 3 targets. Some models are not converged in the last day, and training must be terminated to finish in time:</p>\n<ul>\n<li><p><strong>Molecule Fingerprint + MLP</strong></p>\n<ul>\n<li>Fingerprints: ECFP6, Topological Torsion, MHFP</li>\n<li>MLP with hidden channels [1024, 1024], dropout = 0.3</li>\n<li>Tried KAN but did not outperform MLP</li>\n<li><a href=\"https://openreview.net/forum?id=NLFqlDeuzt\" target=\"_blank\">Independent Feature Matching (IFM)</a> gave no boost</li>\n<li>Feature Selection based on SHAP also gave no boost. For a long time I had believe Feature Selection is the key to reduce share-nonshare gap and overfitting caused by noisy signal, but had no success with it.</li></ul></li>\n<li><p><strong>String-based 1D-CNN</strong></p>\n<ul>\n<li>Input representations: SMILES, Atom-In-Smiles (AIS), SELFIES, DeepSMILES</li>\n<li>Model: Improved from the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">awesome public notebook</a>, just add MaskedBatchNorm, MaskedAttentionPooling and scale to larger size (depth = 6, dim=128)</li>\n<li>All input representations give equally good result on shared part. SELFIES give slightly worse results on non-shared and converge slower than other input representations</li>\n<li>Deeper -&gt; better generalisation. <code>(depth=8, dim=64)</code> is <code>&gt;</code> <code>(depth=3, dim=512)</code> on non-shared, but <code>&lt;</code> on shared. <code>(depth=3, dim=512)</code> also better remember the training set (higher train AP) while archive similar share AP. This need more experiments to confirm.</li></ul></li>\n<li><p><strong>String-based Hybrid CNN-Transformer</strong></p>\n<ul>\n<li>Input representation: SMILES, Atom-In-Smiles</li>\n<li>Model: Squeezeformer, depth = 6, dim=64/96/128</li>\n<li>No positional encoding, CNN do the job. ROPE decrease CV score</li>\n<li>Pretrain: <code>No pretrain</code> or <code>MTR</code> or <code>MLM+MTR</code></li></ul></li>\n<li><p><strong>String-based Transformer</strong></p>\n<ul>\n<li>Input representation: SMILES</li>\n<li>Model: Roberta, depth=6, dim=256, Absolute PE</li>\n<li>Pretrain: <code>MLM+MTR</code></li></ul></li>\n<li><p><strong>String-based MAMBA</strong></p>\n<ul>\n<li>Input representation: SMILES</li>\n<li>Model: Mamba, depth=8, dim=128</li></ul></li>\n<li><p><strong>GNNs</strong></p>\n<ul>\n<li>GIN with depth=5, dim=300, finetune from public pretrained <a href=\"https://github.com/junxia97/Mole-BERT\" target=\"_blank\">Mole-BERT</a> (pretrained on ZINC)</li>\n<li>GCN/GraphSAGE with depth=5, dim=128</li></ul></li>\n<li><p><strong>Catboost</strong></p>\n<ul>\n<li>Ensemble of 12 folds, covering all train dataset. Each fold contain all positive samples and 5.0 times of negative samples.</li>\n<li>Max_depth=10, lr=0.2, iterations=4000, bootstrap_type='No'</li></ul></li>\n</ul>\n<p>Some modeling results used in final ensemble. I think analyze share/nonshared/new library scores is needed for further analysis, and I will update that results later.</p>\n<table>\n<thead>\n<tr>\n<th>File Name</th>\n<th>Private Leaderboard</th>\n<th>Public Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808_all</code></td>\n<td><strong>0.301</strong></td>\n<td>0.399</td>\n</tr>\n<tr>\n<td><code>final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412_all</code></td>\n<td>0.290</td>\n<td>0.373</td>\n</tr>\n<tr>\n<td><code>ensemble_v2</code></td>\n<td><strong><em>0.278</em></strong></td>\n<td><strong><em>0.448</em></strong></td>\n</tr>\n<tr>\n<td><code>ensemble_v1</code></td>\n<td>0.277</td>\n<td>0.435</td>\n</tr>\n<tr>\n<td><code>ensemble_v0</code></td>\n<td>0.274</td>\n<td>0.438</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-8-64_all</code></td>\n<td>0.270</td>\n<td>0.389</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-3-512_all</code></td>\n<td>0.257</td>\n<td>0.386</td>\n</tr>\n<tr>\n<td><code>final_selfies-pregen_cnn1d-3-512_all</code></td>\n<td>0.253</td>\n<td>0.367</td>\n</tr>\n<tr>\n<td><code>final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898_all</code></td>\n<td>0.251</td>\n<td>0.407</td>\n</tr>\n<tr>\n<td><code>final_selfies-pregen_cnn1d-6-128_all</code></td>\n<td>0.244</td>\n<td>0.365</td>\n</tr>\n<tr>\n<td><code>final_ais-pregen_cnn1d-6-128_all</code></td>\n<td>0.244</td>\n<td>0.395</td>\n</tr>\n<tr>\n<td><code>final_sage_all</code></td>\n<td>0.242</td>\n<td><strong>0.433</strong></td>\n</tr>\n<tr>\n<td><code>final_deepsmiles-pregen_cnn1d-6-128_all</code></td>\n<td>0.242</td>\n<td>0.388</td>\n</tr>\n<tr>\n<td><code>final_pretrain-gin_all</code></td>\n<td>0.240</td>\n<td>0.385</td>\n</tr>\n<tr>\n<td><code>topo_mlp_all</code></td>\n<td>0.236</td>\n<td>0.377</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-6-128_all</code></td>\n<td>0.235</td>\n<td>0.392</td>\n</tr>\n<tr>\n<td><code>final_ais_squeeze_all</code></td>\n<td>0.235</td>\n<td>0.388</td>\n</tr>\n<tr>\n<td><code>final_gcn_all</code></td>\n<td>0.233</td>\n<td>0.396</td>\n</tr>\n<tr>\n<td><code>final_mamba_smileschar_all</code></td>\n<td>0.233</td>\n<td>0.394</td>\n</tr>\n<tr>\n<td><code>squeezeformer_0.1033_0.6376_all</code></td>\n<td>0.230</td>\n<td>0.410</td>\n</tr>\n<tr>\n<td><code>catboost_12chunks_all</code></td>\n<td>0.224</td>\n<td>0.385</td>\n</tr>\n<tr>\n<td><code>ecfp6_mlp_all</code></td>\n<td>0.218</td>\n<td>0.359</td>\n</tr>\n<tr>\n<td><code>mhfp_mlp_all</code></td>\n<td>0.193</td>\n<td>0.307</td>\n</tr>\n</tbody>\n</table>\n<p>Best single model was a <strong>Roberta with join MLM + MTR pretraining</strong>, scored <strong>0.301</strong> on Private Leaderboard and 0.399 on Public Leaderboard.</p>\n<p>Simple weighted ensemble with heuristic weights scores <strong>PB=27.8 and LB=44.8</strong>\nIndeed there are some single model with higher PB scores, but ensemble reduce variance and give more stable/trustworthy results. This result is luckily enough for a gold medal.</p>\n<h2>Code</h2>\n<ul>\n<li>Training code: <a href=\"https://github.com/dangnh0611/kaggle_leash_belka\" target=\"_blank\">https://github.com/dangnh0611/kaggle_leash_belka</a></li>\n</ul>\n<p>Thanks for your attention !</p>",
  "messages": [
    {
      "id": "2913113",
      "postDate": "07/09/2024 09:02:53",
      "content": "<p>First, I would like to thank Kaggle and competition's host for another interesting challenge, which bring me headaches for a long time :D\nThank you to all participants, especially who actively discuss and provide many good insights/code in the forum. As usual on Kaggle, I have a great time learning new things.</p>\n<h2>First impression</h2>\n<ul>\n<li>I trained an Embedding + MLP models on building block id only, e.g [BB1, BB2, BB3] = [100, 200, 300]. Yes, it just see the ID. This model scores shared_AP=65.5 on StratifiedKFold(n_splits=20), which indeed surpass some of my SMILES+CNN1d without looking for what each building block looks like. This strongly indicate that model could memorize target binding and overfiting is nearby.</li>\n<li>Many attempts but actually I failed to setup a trustworthy CV scheme. In such situation, I decided to be blind at all and try to focus the fancy word \"diversity\": Multiple models, multiple input representations and SSL pretraining. I think that strategy indeed survived me in this shuffle competition.</li>\n</ul>\n<h2>How to Cross-Validate?</h2>\n<p>As mentioned above, I don't know.</p>\n<p>My split strategy for each fold:</p>\n<pre><code>1. [11k samples] Hold out a number of building blocks for validation: 17 BB1 + 36 BB2 with positive-fraction-aware balance on each fold.\n2. [78k samples] Leave 20% molecules with most regular scaffolds (&gt; 6116 mols/scaffold) for training, then do a Scaffold Split on the remaining 80% molecules\n3. [103k samples] Stratified Random split on the remaining\n</code></pre>\n<p>Total: 11k non-shared + 181k share, simulate the LB\nThis strategy create a little \"harder\" than pure random split on the shared part.</p>\n<p>The hard part lie in 11k non-shared. Some observation on non-shared CV:</p>\n<ul>\n<li><p>CV scores vary largely between different folds. I find it hard to identify a trend/correlation, or which is work/not work</p></li>\n<li><p>Early stopping should help, but it also hard to answer WHEN? Usually best score is found in first 2 epochs, but also could be much longer\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fa266b63c8b17261dc0768365c16b09ab%2Fnonshare-cv.png?generation=1720515545119485&amp;alt=media\" alt=\"\">\n<em>Each column in form: <code>score (best_epoch)</code>. All results belong to a single split, with nonshare positive rate <code>[0.571, 0.638, 0.714] %</code></em></p></li>\n<li><p>Suitable hyperparams also vary largely between folds. It looks like best setting with best score in fold A could give worst score on fold B :(</p></li>\n<li><p>For one particular fold: score also change significantly due to just small change in hyperparams like random seed, validate interval, .etc. (13.5 -&gt; 32.4, what?)</p></li>\n<li><p>Overfit a model on LB-nonshared-similar building blocks improve nonshared LB score. <code>Building block ECFP6 -&gt; Tanimoto similarity -&gt; Linear Assignment Matching</code> to select most similar/disimilar 170 BB1 + 360 BB2/3, then overfit a 1D-CNN on this subset. LB score for similar=8.8, disimilar=3.9. This indicate that we can improve LB by focusing more on these similar samples. I later select most similar 17 BB1 + 36 BB2/3 as validation set and try to improve CV on this split, but improved CV led to decrease in non-shared LB, which confused me</p></li>\n</ul>\n<p>In conclusion, more and more experiments messed up my mind. The variation is too large and seem to be random for me. I gave up and focus on <strong>SSL pretraining</strong>, improve the diversity by using <strong>multiple models</strong> and <strong>multiple input representations</strong>, and allmost all models were <strong>blindly train on all competition data with no cross-validation</strong>, since I did not trust a single CV split alone and train on multiple CV splits seem to be too time-intensive</p>\n<h2>SSL Pretraining</h2>\n<p>2 pretraining tasks:</p>\n<ul>\n<li>MLM (Masked Language Modeling) with standard setting: 15% masked tokens, 80%-10%-10% replaced with <code>[MASK]</code>/random tokens/keep unchanged.</li>\n<li>MTR (Multi-Task Regression): regress 189 pre-calculated RDKIT Descriptors. These target values are min-max normalized in contrast to CDF Transform in literature.</li>\n</ul>\n<p><code>Joint MTR + MLM pretraining</code> + <code>SMILES Enumeration</code> has success as mentioned in <a href=\"https://github.com/BenevolentAI/MolBERT\" target=\"_blank\">MolBERT</a>, which inspired me to give it a try</p>\n<p>Pretraining using all data (both train + test), with larger sampling weights to test dataset to equalize the frequency of each building blocks and be \"more familiar\" with new domain test dataset.</p>\n<p>I trained 3 models to be finetuned on competition task:</p>\n<ul>\n<li>Squeezeformer + <code>MTR</code></li>\n<li>Squeezeformer + <code>Joint MTR + MLM</code></li>\n<li>Roberta + <code>Joint MTR + MLM</code></li>\n</ul>\n<h2>Modeling</h2>\n<p>All models simultaneously predict 3 targets. Some models are not converged in the last day, and training must be terminated to finish in time:</p>\n<ul>\n<li><p><strong>Molecule Fingerprint + MLP</strong></p>\n<ul>\n<li>Fingerprints: ECFP6, Topological Torsion, MHFP</li>\n<li>MLP with hidden channels [1024, 1024], dropout = 0.3</li>\n<li>Tried KAN but did not outperform MLP</li>\n<li><a href=\"https://openreview.net/forum?id=NLFqlDeuzt\" target=\"_blank\">Independent Feature Matching (IFM)</a> gave no boost</li>\n<li>Feature Selection based on SHAP also gave no boost. For a long time I had believe Feature Selection is the key to reduce share-nonshare gap and overfitting caused by noisy signal, but had no success with it.</li></ul></li>\n<li><p><strong>String-based 1D-CNN</strong></p>\n<ul>\n<li>Input representations: SMILES, Atom-In-Smiles (AIS), SELFIES, DeepSMILES</li>\n<li>Model: Improved from the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">awesome public notebook</a>, just add MaskedBatchNorm, MaskedAttentionPooling and scale to larger size (depth = 6, dim=128)</li>\n<li>All input representations give equally good result on shared part. SELFIES give slightly worse results on non-shared and converge slower than other input representations</li>\n<li>Deeper -&gt; better generalisation. <code>(depth=8, dim=64)</code> is <code>&gt;</code> <code>(depth=3, dim=512)</code> on non-shared, but <code>&lt;</code> on shared. <code>(depth=3, dim=512)</code> also better remember the training set (higher train AP) while archive similar share AP. This need more experiments to confirm.</li></ul></li>\n<li><p><strong>String-based Hybrid CNN-Transformer</strong></p>\n<ul>\n<li>Input representation: SMILES, Atom-In-Smiles</li>\n<li>Model: Squeezeformer, depth = 6, dim=64/96/128</li>\n<li>No positional encoding, CNN do the job. ROPE decrease CV score</li>\n<li>Pretrain: <code>No pretrain</code> or <code>MTR</code> or <code>MLM+MTR</code></li></ul></li>\n<li><p><strong>String-based Transformer</strong></p>\n<ul>\n<li>Input representation: SMILES</li>\n<li>Model: Roberta, depth=6, dim=256, Absolute PE</li>\n<li>Pretrain: <code>MLM+MTR</code></li></ul></li>\n<li><p><strong>String-based MAMBA</strong></p>\n<ul>\n<li>Input representation: SMILES</li>\n<li>Model: Mamba, depth=8, dim=128</li></ul></li>\n<li><p><strong>GNNs</strong></p>\n<ul>\n<li>GIN with depth=5, dim=300, finetune from public pretrained <a href=\"https://github.com/junxia97/Mole-BERT\" target=\"_blank\">Mole-BERT</a> (pretrained on ZINC)</li>\n<li>GCN/GraphSAGE with depth=5, dim=128</li></ul></li>\n<li><p><strong>Catboost</strong></p>\n<ul>\n<li>Ensemble of 12 folds, covering all train dataset. Each fold contain all positive samples and 5.0 times of negative samples.</li>\n<li>Max_depth=10, lr=0.2, iterations=4000, bootstrap_type='No'</li></ul></li>\n</ul>\n<p>Some modeling results used in final ensemble. I think analyze share/nonshared/new library scores is needed for further analysis, and I will update that results later.</p>\n<table>\n<thead>\n<tr>\n<th>File Name</th>\n<th>Private Leaderboard</th>\n<th>Public Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808_all</code></td>\n<td><strong>0.301</strong></td>\n<td>0.399</td>\n</tr>\n<tr>\n<td><code>final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412_all</code></td>\n<td>0.290</td>\n<td>0.373</td>\n</tr>\n<tr>\n<td><code>ensemble_v2</code></td>\n<td><strong><em>0.278</em></strong></td>\n<td><strong><em>0.448</em></strong></td>\n</tr>\n<tr>\n<td><code>ensemble_v1</code></td>\n<td>0.277</td>\n<td>0.435</td>\n</tr>\n<tr>\n<td><code>ensemble_v0</code></td>\n<td>0.274</td>\n<td>0.438</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-8-64_all</code></td>\n<td>0.270</td>\n<td>0.389</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-3-512_all</code></td>\n<td>0.257</td>\n<td>0.386</td>\n</tr>\n<tr>\n<td><code>final_selfies-pregen_cnn1d-3-512_all</code></td>\n<td>0.253</td>\n<td>0.367</td>\n</tr>\n<tr>\n<td><code>final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898_all</code></td>\n<td>0.251</td>\n<td>0.407</td>\n</tr>\n<tr>\n<td><code>final_selfies-pregen_cnn1d-6-128_all</code></td>\n<td>0.244</td>\n<td>0.365</td>\n</tr>\n<tr>\n<td><code>final_ais-pregen_cnn1d-6-128_all</code></td>\n<td>0.244</td>\n<td>0.395</td>\n</tr>\n<tr>\n<td><code>final_sage_all</code></td>\n<td>0.242</td>\n<td><strong>0.433</strong></td>\n</tr>\n<tr>\n<td><code>final_deepsmiles-pregen_cnn1d-6-128_all</code></td>\n<td>0.242</td>\n<td>0.388</td>\n</tr>\n<tr>\n<td><code>final_pretrain-gin_all</code></td>\n<td>0.240</td>\n<td>0.385</td>\n</tr>\n<tr>\n<td><code>topo_mlp_all</code></td>\n<td>0.236</td>\n<td>0.377</td>\n</tr>\n<tr>\n<td><code>final_smileschar-pregen_cnn1d-6-128_all</code></td>\n<td>0.235</td>\n<td>0.392</td>\n</tr>\n<tr>\n<td><code>final_ais_squeeze_all</code></td>\n<td>0.235</td>\n<td>0.388</td>\n</tr>\n<tr>\n<td><code>final_gcn_all</code></td>\n<td>0.233</td>\n<td>0.396</td>\n</tr>\n<tr>\n<td><code>final_mamba_smileschar_all</code></td>\n<td>0.233</td>\n<td>0.394</td>\n</tr>\n<tr>\n<td><code>squeezeformer_0.1033_0.6376_all</code></td>\n<td>0.230</td>\n<td>0.410</td>\n</tr>\n<tr>\n<td><code>catboost_12chunks_all</code></td>\n<td>0.224</td>\n<td>0.385</td>\n</tr>\n<tr>\n<td><code>ecfp6_mlp_all</code></td>\n<td>0.218</td>\n<td>0.359</td>\n</tr>\n<tr>\n<td><code>mhfp_mlp_all</code></td>\n<td>0.193</td>\n<td>0.307</td>\n</tr>\n</tbody>\n</table>\n<p>Best single model was a <strong>Roberta with join MLM + MTR pretraining</strong>, scored <strong>0.301</strong> on Private Leaderboard and 0.399 on Public Leaderboard.</p>\n<p>Simple weighted ensemble with heuristic weights scores <strong>PB=27.8 and LB=44.8</strong>\nIndeed there are some single model with higher PB scores, but ensemble reduce variance and give more stable/trustworthy results. This result is luckily enough for a gold medal.</p>\n<h2>Code</h2>\n<ul>\n<li>Training code: <a href=\"https://github.com/dangnh0611/kaggle_leash_belka\" target=\"_blank\">https://github.com/dangnh0611/kaggle_leash_belka</a></li>\n</ul>\n<p>Thanks for your attention !</p>",
      "rawMarkdown": "First, I would like to thank Kaggle and competition's host for another interesting challenge, which bring me headaches for a long time :D\nThank you to all participants, especially who actively discuss and provide many good insights/code in the forum. As usual on Kaggle, I have a great time learning new things.\n\n## First impression\n- I trained an Embedding + MLP models on building block id only, e.g [BB1, BB2, BB3] = [100, 200, 300]. Yes, it just see the ID. This model scores shared_AP=65.5 on StratifiedKFold(n_splits=20), which indeed surpass some of my SMILES+CNN1d without looking for what each building block looks like. This strongly indicate that model could memorize target binding and overfiting is nearby.\n- Many attempts but actually I failed to setup a trustworthy CV scheme. In such situation, I decided to be blind at all and try to focus the fancy word \"diversity\": Multiple models, multiple input representations and SSL pretraining. I think that strategy indeed survived me in this shuffle competition.\n\n\n## How to Cross-Validate?\nAs mentioned above, I don't know.\n\nMy split strategy for each fold:\n\n    1. [11k samples] Hold out a number of building blocks for validation: 17 BB1 + 36 BB2 with positive-fraction-aware balance on each fold.\n    2. [78k samples] Leave 20% molecules with most regular scaffolds (> 6116 mols/scaffold) for training, then do a Scaffold Split on the remaining 80% molecules\n    3. [103k samples] Stratified Random split on the remaining\nTotal: 11k non-shared + 181k share, simulate the LB\nThis strategy create a little \"harder\" than pure random split on the shared part.\n\nThe hard part lie in 11k non-shared. Some observation on non-shared CV:\n- CV scores vary largely between different folds. I find it hard to identify a trend/correlation, or which is work/not work\n- Early stopping should help, but it also hard to answer WHEN? Usually best score is found in first 2 epochs, but also could be much longer\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fa266b63c8b17261dc0768365c16b09ab%2Fnonshare-cv.png?generation=1720515545119485&alt=media)\n    *Each column in form: `score (best_epoch)`. All results belong to a single split, with nonshare positive rate `[0.571, 0.638, 0.714] %`*\n\n- Suitable hyperparams also vary largely between folds. It looks like best setting with best score in fold A could give worst score on fold B :(\n- For one particular fold: score also change significantly due to just small change in hyperparams like random seed, validate interval, .etc. (13.5 -> 32.4, what?)\n- Overfit a model on LB-nonshared-similar building blocks improve nonshared LB score. `Building block ECFP6 -> Tanimoto similarity -> Linear Assignment Matching` to select most similar/disimilar 170 BB1 + 360 BB2/3, then overfit a 1D-CNN on this subset. LB score for similar=8.8, disimilar=3.9. This indicate that we can improve LB by focusing more on these similar samples. I later select most similar 17 BB1 + 36 BB2/3 as validation set and try to improve CV on this split, but improved CV led to decrease in non-shared LB, which confused me\n\nIn conclusion, more and more experiments messed up my mind. The variation is too large and seem to be random for me. I gave up and focus on **SSL pretraining**, improve the diversity by using **multiple models** and **multiple input representations**, and allmost all models were **blindly train on all competition data with no cross-validation**, since I did not trust a single CV split alone and train on multiple CV splits seem to be too time-intensive\n\n\n## SSL Pretraining\n2 pretraining tasks:\n- MLM (Masked Language Modeling) with standard setting: 15% masked tokens, 80%-10%-10% replaced with `[MASK]`/random tokens/keep unchanged.\n- MTR (Multi-Task Regression): regress 189 pre-calculated RDKIT Descriptors. These target values are min-max normalized in contrast to CDF Transform in literature.\n\n`Joint MTR + MLM pretraining` + `SMILES Enumeration` has success as mentioned in [MolBERT](https://github.com/BenevolentAI/MolBERT), which inspired me to give it a try\n\nPretraining using all data (both train + test), with larger sampling weights to test dataset to equalize the frequency of each building blocks and be \"more familiar\" with new domain test dataset.\n\nI trained 3 models to be finetuned on competition task:\n- Squeezeformer + `MTR`\n- Squeezeformer + `Joint MTR + MLM`\n- Roberta + `Joint MTR + MLM`\n \n\n## Modeling\n\nAll models simultaneously predict 3 targets. Some models are not converged in the last day, and training must be terminated to finish in time:\n\n- **Molecule Fingerprint + MLP**\n    + Fingerprints: ECFP6, Topological Torsion, MHFP\n    + MLP with hidden channels [1024, 1024], dropout = 0.3\n    + Tried KAN but did not outperform MLP\n    + [Independent Feature Matching (IFM)](https://openreview.net/forum?id=NLFqlDeuzt) gave no boost\n    + Feature Selection based on SHAP also gave no boost. For a long time I had believe Feature Selection is the key to reduce share-nonshare gap and overfitting caused by noisy signal, but had no success with it.\n\n- **String-based 1D-CNN**\n    + Input representations: SMILES, Atom-In-Smiles (AIS), SELFIES, DeepSMILES\n    + Model: Improved from the [awesome public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data), just add MaskedBatchNorm, MaskedAttentionPooling and scale to larger size (depth = 6, dim=128)\n    + All input representations give equally good result on shared part. SELFIES give slightly worse results on non-shared and converge slower than other input representations\n    + Deeper -> better generalisation. `(depth=8, dim=64)` is `>` `(depth=3, dim=512)` on non-shared, but `<` on shared. `(depth=3, dim=512)` also better remember the training set (higher train AP) while archive similar share AP. This need more experiments to confirm.\n\n- **String-based Hybrid CNN-Transformer**\n    + Input representation: SMILES, Atom-In-Smiles\n    + Model: Squeezeformer, depth = 6, dim=64/96/128\n    + No positional encoding, CNN do the job. ROPE decrease CV score\n    + Pretrain: `No pretrain` or `MTR` or `MLM+MTR`\n\n- **String-based Transformer**\n    + Input representation: SMILES\n    + Model: Roberta, depth=6, dim=256, Absolute PE\n    + Pretrain: `MLM+MTR`\n\n- **String-based MAMBA**\n    + Input representation: SMILES\n    + Model: Mamba, depth=8, dim=128\n- **GNNs**\n    + GIN with depth=5, dim=300, finetune from public pretrained [Mole-BERT](https://github.com/junxia97/Mole-BERT) (pretrained on ZINC)\n    + GCN/GraphSAGE with depth=5, dim=128\n- **Catboost**\n    + Ensemble of 12 folds, covering all train dataset. Each fold contain all positive samples and 5.0 times of negative samples.\n    + Max_depth=10, lr=0.2, iterations=4000, bootstrap_type='No'\n\nSome modeling results used in final ensemble. I think analyze share/nonshared/new library scores is needed for further analysis, and I will update that results later.\n\n\n| File Name | Private Leaderboard | Public Leaderboard |\n| --- | --- | --- |\n| `final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808_all` | **0.301** | 0.399 |\n| `final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412_all` | 0.290 | 0.373 |\n| `ensemble_v2` | ***0.278*** | ***0.448*** |\n| `ensemble_v1` | 0.277 | 0.435 |\n| `ensemble_v0` | 0.274 | 0.438 |\n| `final_smileschar-pregen_cnn1d-8-64_all` | 0.270 | 0.389 |\n| `final_smileschar-pregen_cnn1d-3-512_all` | 0.257 | 0.386 |\n| `final_selfies-pregen_cnn1d-3-512_all` | 0.253 | 0.367 |\n| `final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898_all` | 0.251 | 0.407 |\n| `final_selfies-pregen_cnn1d-6-128_all` | 0.244 | 0.365 |\n| `final_ais-pregen_cnn1d-6-128_all` | 0.244 | 0.395 |\n| `final_sage_all` | 0.242 | **0.433** |\n| `final_deepsmiles-pregen_cnn1d-6-128_all` | 0.242 | 0.388 |\n| `final_pretrain-gin_all` | 0.240 | 0.385 |\n| `topo_mlp_all` | 0.236 | 0.377 |\n| `final_smileschar-pregen_cnn1d-6-128_all` | 0.235 | 0.392 |\n| `final_ais_squeeze_all` | 0.235 | 0.388 |\n| `final_gcn_all` | 0.233 | 0.396 |\n| `final_mamba_smileschar_all` | 0.233 | 0.394 |\n| `squeezeformer_0.1033_0.6376_all` | 0.230 | 0.410 |\n| `catboost_12chunks_all` | 0.224 | 0.385 |\n| `ecfp6_mlp_all` | 0.218 | 0.359 |\n| `mhfp_mlp_all` | 0.193 | 0.307 |\n\n\n\nBest single model was a **Roberta with join MLM + MTR pretraining**, scored **0.301** on Private Leaderboard and 0.399 on Public Leaderboard.\n\nSimple weighted ensemble with heuristic weights scores **PB=27.8 and LB=44.8**\nIndeed there are some single model with higher PB scores, but ensemble reduce variance and give more stable/trustworthy results. This result is luckily enough for a gold medal.\n\n## Code\n- Training code: https://github.com/dangnh0611/kaggle_leash_belka\n\nThanks for your attention !",
      "votes": null
    },
    {
      "id": "2913666",
      "postDate": "07/09/2024 15:42:32",
      "content": "<p>Great to see so many models! Congrats!</p>",
      "rawMarkdown": "Great to see so many models! Congrats!",
      "votes": null
    },
    {
      "id": "2914572",
      "postDate": "07/10/2024 03:42:26",
      "content": "<p>Thank you for sharing! I'm really impressed by the SSL pretraining approach! I'm looking forward to seeing your code :)</p>\n<p>Also, if it's okay with you, can I ask which models you used for the ensemble? Did you only use the finetuned three models, or were there more?</p>",
      "rawMarkdown": "Thank you for sharing! I'm really impressed by the SSL pretraining approach! I'm looking forward to seeing your code :)\n\nAlso, if it's okay with you, can I ask which models you used for the ensemble? Did you only use the finetuned three models, or were there more?",
      "votes": null
    },
    {
      "id": "2914788",
      "postDate": "07/10/2024 06:54:05",
      "content": "<p>Thank you! The code is <a href=\"https://github.com/dangnh0611/kaggle_leash_belka\" target=\"_blank\">released</a>. There are more than three models in PB=27.8 ensemble, with different weights for share/non-share groups:</p>\n<pre><code>SHARE = [\n    \n    \n    \n    \n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    \n    \n    \n    (, ),\n    \n    \n    (, ),\n    (, ),\n    \n    \n    \n]\n\nNONSHARE = [\n    \n    \n    \n    \n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, )\n]\n</code></pre>",
      "rawMarkdown": "Thank you! The code is [released](https://github.com/dangnh0611/kaggle_leash_belka). There are more than three models in PB=27.8 ensemble, with different weights for share/non-share groups:\n \n```python\nSHARE = [\n    # ('catboost_12chunks', 0.0),\n    # ('ecfp6_mlp', 0.0),\n    # ('topo_mlp', 0.0),\n    # ('mhfp_mlp', 0.0),\n    ('final_smileschar-pregen_cnn1d-6-128', 1),\n    ('final_ais-pregen_cnn1d-6-128', 1),\n    ('final_selfies-pregen_cnn1d-6-128', 1),\n    ('final_deepsmiles-pregen_cnn1d-6-128', 1),\n    # ('final_smileschar-pregen_cnn1d-8-64', 0.0),\n    # ('final_smileschar-pregen_cnn1d-3-512', 0.0),\n    # ('final_selfies-pregen_cnn1d-3-512', 0.0),\n    ('final_ais_squeeze', 1.0),\n    # ('final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412', 0.0),\n    # ('final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898', 0.0),\n    ('final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808', 1.0),\n    ('final_mamba_smileschar', 1.0),\n    # ('final_pretrain-gin', 0.0),\n    # ('final_sage', 0.0),\n    # ('final_gcn', 0.0)\n]\n\nNONSHARE = [\n    # ('catboost_12chunks', 0.0),\n    # ('ecfp6_mlp', 0.0),\n    # ('topo_mlp', 0.0),\n    # ('mhfp_mlp', 0.0),\n    ('final_smileschar-pregen_cnn1d-6-128', 0.5),\n    ('final_ais-pregen_cnn1d-6-128', 0.5),\n    ('final_selfies-pregen_cnn1d-6-128', 0.5),\n    ('final_deepsmiles-pregen_cnn1d-6-128', 0.5),\n    ('final_smileschar-pregen_cnn1d-8-64', 0.2),\n    ('final_smileschar-pregen_cnn1d-3-512', 0.2),\n    ('final_selfies-pregen_cnn1d-3-512', 0.2),\n    ('squeeze_first_version', 3.0),\n    ('final_ais_squeeze', 1.0),\n    ('final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412', 0.8),\n    ('final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898', 0.8),\n    ('final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808', 1.0),\n    ('final_mamba_smileschar', 1.0),\n    ('final_pretrain-gin', 0.5),\n    ('final_sage', 0.1),\n    ('final_gcn', 0.1)\n]\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2913666,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "07/09/2024 15:42:32",
      "content": "<p>Great to see so many models! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914572,
      "author_name": "kyu999",
      "author_url": "",
      "post_date": "07/10/2024 03:42:26",
      "content": "<p>Thank you for sharing! I'm really impressed by the SSL pretraining approach! I'm looking forward to seeing your code :)</p>\n<p>Also, if it's okay with you, can I ask which models you used for the ensemble? Did you only use the finetuned three models, or were there more?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2914788,
          "author_name": "dangnh0611",
          "author_url": "",
          "post_date": "07/10/2024 06:54:05",
          "content": "<p>Thank you! The code is <a href=\"https://github.com/dangnh0611/kaggle_leash_belka\" target=\"_blank\">released</a>. There are more than three models in PB=27.8 ensemble, with different weights for share/non-share groups:</p>\n<pre><code>SHARE = [\n    \n    \n    \n    \n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    \n    \n    \n    (, ),\n    \n    \n    (, ),\n    (, ),\n    \n    \n    \n]\n\nNONSHARE = [\n    \n    \n    \n    \n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, ),\n    (, )\n]\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2913113": "First, I would like to thank Kaggle and competition's host for another interesting challenge, which bring me headaches for a long time :D\nThank you to all participants, especially who actively discuss and provide many good insights/code in the forum. As usual on Kaggle, I have a great time learning new things.\n\n## First impression\n- I trained an Embedding + MLP models on building block id only, e.g [BB1, BB2, BB3] = [100, 200, 300]. Yes, it just see the ID. This model scores shared_AP=65.5 on StratifiedKFold(n_splits=20), which indeed surpass some of my SMILES+CNN1d without looking for what each building block looks like. This strongly indicate that model could memorize target binding and overfiting is nearby.\n- Many attempts but actually I failed to setup a trustworthy CV scheme. In such situation, I decided to be blind at all and try to focus the fancy word \"diversity\": Multiple models, multiple input representations and SSL pretraining. I think that strategy indeed survived me in this shuffle competition.\n\n\n## How to Cross-Validate?\nAs mentioned above, I don't know.\n\nMy split strategy for each fold:\n\n    1. [11k samples] Hold out a number of building blocks for validation: 17 BB1 + 36 BB2 with positive-fraction-aware balance on each fold.\n    2. [78k samples] Leave 20% molecules with most regular scaffolds (> 6116 mols/scaffold) for training, then do a Scaffold Split on the remaining 80% molecules\n    3. [103k samples] Stratified Random split on the remaining\nTotal: 11k non-shared + 181k share, simulate the LB\nThis strategy create a little \"harder\" than pure random split on the shared part.\n\nThe hard part lie in 11k non-shared. Some observation on non-shared CV:\n- CV scores vary largely between different folds. I find it hard to identify a trend/correlation, or which is work/not work\n- Early stopping should help, but it also hard to answer WHEN? Usually best score is found in first 2 epochs, but also could be much longer\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fa266b63c8b17261dc0768365c16b09ab%2Fnonshare-cv.png?generation=1720515545119485&alt=media)\n    *Each column in form: `score (best_epoch)`. All results belong to a single split, with nonshare positive rate `[0.571, 0.638, 0.714] %`*\n\n- Suitable hyperparams also vary largely between folds. It looks like best setting with best score in fold A could give worst score on fold B :(\n- For one particular fold: score also change significantly due to just small change in hyperparams like random seed, validate interval, .etc. (13.5 -> 32.4, what?)\n- Overfit a model on LB-nonshared-similar building blocks improve nonshared LB score. `Building block ECFP6 -> Tanimoto similarity -> Linear Assignment Matching` to select most similar/disimilar 170 BB1 + 360 BB2/3, then overfit a 1D-CNN on this subset. LB score for similar=8.8, disimilar=3.9. This indicate that we can improve LB by focusing more on these similar samples. I later select most similar 17 BB1 + 36 BB2/3 as validation set and try to improve CV on this split, but improved CV led to decrease in non-shared LB, which confused me\n\nIn conclusion, more and more experiments messed up my mind. The variation is too large and seem to be random for me. I gave up and focus on **SSL pretraining**, improve the diversity by using **multiple models** and **multiple input representations**, and allmost all models were **blindly train on all competition data with no cross-validation**, since I did not trust a single CV split alone and train on multiple CV splits seem to be too time-intensive\n\n\n## SSL Pretraining\n2 pretraining tasks:\n- MLM (Masked Language Modeling) with standard setting: 15% masked tokens, 80%-10%-10% replaced with `[MASK]`/random tokens/keep unchanged.\n- MTR (Multi-Task Regression): regress 189 pre-calculated RDKIT Descriptors. These target values are min-max normalized in contrast to CDF Transform in literature.\n\n`Joint MTR + MLM pretraining` + `SMILES Enumeration` has success as mentioned in [MolBERT](https://github.com/BenevolentAI/MolBERT), which inspired me to give it a try\n\nPretraining using all data (both train + test), with larger sampling weights to test dataset to equalize the frequency of each building blocks and be \"more familiar\" with new domain test dataset.\n\nI trained 3 models to be finetuned on competition task:\n- Squeezeformer + `MTR`\n- Squeezeformer + `Joint MTR + MLM`\n- Roberta + `Joint MTR + MLM`\n \n\n## Modeling\n\nAll models simultaneously predict 3 targets. Some models are not converged in the last day, and training must be terminated to finish in time:\n\n- **Molecule Fingerprint + MLP**\n    + Fingerprints: ECFP6, Topological Torsion, MHFP\n    + MLP with hidden channels [1024, 1024], dropout = 0.3\n    + Tried KAN but did not outperform MLP\n    + [Independent Feature Matching (IFM)](https://openreview.net/forum?id=NLFqlDeuzt) gave no boost\n    + Feature Selection based on SHAP also gave no boost. For a long time I had believe Feature Selection is the key to reduce share-nonshare gap and overfitting caused by noisy signal, but had no success with it.\n\n- **String-based 1D-CNN**\n    + Input representations: SMILES, Atom-In-Smiles (AIS), SELFIES, DeepSMILES\n    + Model: Improved from the [awesome public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data), just add MaskedBatchNorm, MaskedAttentionPooling and scale to larger size (depth = 6, dim=128)\n    + All input representations give equally good result on shared part. SELFIES give slightly worse results on non-shared and converge slower than other input representations\n    + Deeper -> better generalisation. `(depth=8, dim=64)` is `>` `(depth=3, dim=512)` on non-shared, but `<` on shared. `(depth=3, dim=512)` also better remember the training set (higher train AP) while archive similar share AP. This need more experiments to confirm.\n\n- **String-based Hybrid CNN-Transformer**\n    + Input representation: SMILES, Atom-In-Smiles\n    + Model: Squeezeformer, depth = 6, dim=64/96/128\n    + No positional encoding, CNN do the job. ROPE decrease CV score\n    + Pretrain: `No pretrain` or `MTR` or `MLM+MTR`\n\n- **String-based Transformer**\n    + Input representation: SMILES\n    + Model: Roberta, depth=6, dim=256, Absolute PE\n    + Pretrain: `MLM+MTR`\n\n- **String-based MAMBA**\n    + Input representation: SMILES\n    + Model: Mamba, depth=8, dim=128\n- **GNNs**\n    + GIN with depth=5, dim=300, finetune from public pretrained [Mole-BERT](https://github.com/junxia97/Mole-BERT) (pretrained on ZINC)\n    + GCN/GraphSAGE with depth=5, dim=128\n- **Catboost**\n    + Ensemble of 12 folds, covering all train dataset. Each fold contain all positive samples and 5.0 times of negative samples.\n    + Max_depth=10, lr=0.2, iterations=4000, bootstrap_type='No'\n\nSome modeling results used in final ensemble. I think analyze share/nonshared/new library scores is needed for further analysis, and I will update that results later.\n\n\n| File Name | Private Leaderboard | Public Leaderboard |\n| --- | --- | --- |\n| `final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808_all` | **0.301** | 0.399 |\n| `final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412_all` | 0.290 | 0.373 |\n| `ensemble_v2` | ***0.278*** | ***0.448*** |\n| `ensemble_v1` | 0.277 | 0.435 |\n| `ensemble_v0` | 0.274 | 0.438 |\n| `final_smileschar-pregen_cnn1d-8-64_all` | 0.270 | 0.389 |\n| `final_smileschar-pregen_cnn1d-3-512_all` | 0.257 | 0.386 |\n| `final_selfies-pregen_cnn1d-3-512_all` | 0.253 | 0.367 |\n| `final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898_all` | 0.251 | 0.407 |\n| `final_selfies-pregen_cnn1d-6-128_all` | 0.244 | 0.365 |\n| `final_ais-pregen_cnn1d-6-128_all` | 0.244 | 0.395 |\n| `final_sage_all` | 0.242 | **0.433** |\n| `final_deepsmiles-pregen_cnn1d-6-128_all` | 0.242 | 0.388 |\n| `final_pretrain-gin_all` | 0.240 | 0.385 |\n| `topo_mlp_all` | 0.236 | 0.377 |\n| `final_smileschar-pregen_cnn1d-6-128_all` | 0.235 | 0.392 |\n| `final_ais_squeeze_all` | 0.235 | 0.388 |\n| `final_gcn_all` | 0.233 | 0.396 |\n| `final_mamba_smileschar_all` | 0.233 | 0.394 |\n| `squeezeformer_0.1033_0.6376_all` | 0.230 | 0.410 |\n| `catboost_12chunks_all` | 0.224 | 0.385 |\n| `ecfp6_mlp_all` | 0.218 | 0.359 |\n| `mhfp_mlp_all` | 0.193 | 0.307 |\n\n\n\nBest single model was a **Roberta with join MLM + MTR pretraining**, scored **0.301** on Private Leaderboard and 0.399 on Public Leaderboard.\n\nSimple weighted ensemble with heuristic weights scores **PB=27.8 and LB=44.8**\nIndeed there are some single model with higher PB scores, but ensemble reduce variance and give more stable/trustworthy results. This result is luckily enough for a gold medal.\n\n## Code\n- Training code: https://github.com/dangnh0611/kaggle_leash_belka\n\nThanks for your attention !",
    "2913666": "Great to see so many models! Congrats!",
    "2914572": "Thank you for sharing! I'm really impressed by the SSL pretraining approach! I'm looking forward to seeing your code :)\n\nAlso, if it's okay with you, can I ask which models you used for the ensemble? Did you only use the finetuned three models, or were there more?",
    "2914788": "Thank you! The code is [released](https://github.com/dangnh0611/kaggle_leash_belka). There are more than three models in PB=27.8 ensemble, with different weights for share/non-share groups:\n \n```python\nSHARE = [\n    # ('catboost_12chunks', 0.0),\n    # ('ecfp6_mlp', 0.0),\n    # ('topo_mlp', 0.0),\n    # ('mhfp_mlp', 0.0),\n    ('final_smileschar-pregen_cnn1d-6-128', 1),\n    ('final_ais-pregen_cnn1d-6-128', 1),\n    ('final_selfies-pregen_cnn1d-6-128', 1),\n    ('final_deepsmiles-pregen_cnn1d-6-128', 1),\n    # ('final_smileschar-pregen_cnn1d-8-64', 0.0),\n    # ('final_smileschar-pregen_cnn1d-3-512', 0.0),\n    # ('final_selfies-pregen_cnn1d-3-512', 0.0),\n    ('final_ais_squeeze', 1.0),\n    # ('final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412', 0.0),\n    # ('final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898', 0.0),\n    ('final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808', 1.0),\n    ('final_mamba_smileschar', 1.0),\n    # ('final_pretrain-gin', 0.0),\n    # ('final_sage', 0.0),\n    # ('final_gcn', 0.0)\n]\n\nNONSHARE = [\n    # ('catboost_12chunks', 0.0),\n    # ('ecfp6_mlp', 0.0),\n    # ('topo_mlp', 0.0),\n    # ('mhfp_mlp', 0.0),\n    ('final_smileschar-pregen_cnn1d-6-128', 0.5),\n    ('final_ais-pregen_cnn1d-6-128', 0.5),\n    ('final_selfies-pregen_cnn1d-6-128', 0.5),\n    ('final_deepsmiles-pregen_cnn1d-6-128', 0.5),\n    ('final_smileschar-pregen_cnn1d-8-64', 0.2),\n    ('final_smileschar-pregen_cnn1d-3-512', 0.2),\n    ('final_selfies-pregen_cnn1d-3-512', 0.2),\n    ('squeeze_first_version', 3.0),\n    ('final_ais_squeeze', 1.0),\n    ('final_squeeze-mtr_chartokenize-02_ep11_0.033802_0.651412', 0.8),\n    ('final_squeeze-mtr-mlm_chartokenize-02_ep11_0.11199_0.648898', 0.8),\n    ('final_roberta-mtr-mlm_chartokenize-all_ep11.2_na_0.668808', 1.0),\n    ('final_mamba_smileschar', 1.0),\n    ('final_pretrain-gin', 0.5),\n    ('final_sage', 0.1),\n    ('final_gcn', 0.1)\n]\n```"
  },
  "source": "meta"
}