{
  "id": 519020,
  "title": "1st place solution [UPDATED]",
  "url": "/competitions/leash-BELKA/discussion/519020",
  "author_name": "Victor Shlepov",
  "post_date": "2024-07-09T11:04:51.178000",
  "votes": 85,
  "comment_count": 39,
  "views": 0,
  "content": "<p>[UPDATE]</p>\n<p>a) Here is the <a href=\"https://www.kaggle.com/datasets/victorshlepov/belka2024-finalsolution\" target=\"_blank\">dataset</a> with the (i) code  - model and data processing utils (ii) SMILES encoder vocabulary. I will upload processed training data too (just to save ones time). Once I recover the the trained weights I will add them as well - I have only the light version of the model left (tf.keras.export), so I will likely retrain it from scratch. For those of you who prefer Kaggle notebooks - here is <a href=\"https://www.kaggle.com/code/victorshlepov/belka2024-1stplacesolution/edit\" target=\"_blank\">one</a> (but I would not really process data here - it should take close to infinity). </p>\n<p>b) The architecture is fairly simple and model is very flat - just 4 encoder layers with 8 heads. With a vocabulary size of just 43 tokens I end-up with a fixed dimensionality of 32. I've tried 64 and 16 too - they do not perform.</p>\n<p>c) I guess I used <a href=\"https://pypi.org/project/atomInSmiles/\" target=\"_blank\">atomInSmiles</a> in a somewhat incorrect way and end up with a schema where separate tokens are either, atom (C, H, S, etc)  or digits, or anything in square brackets, like [C@@] are distinct tokens. I leave it to chemistry practitioners to decide what this mess really means :)</p>\n<p>d) Pre-training. I pre-trained model from scratch in two stages:</p>\n<ul>\n<li><p>MLM - the standard prediction of masked tokens (15% of which 80% are masked, 10% are replaced with random token and 10% are retained - classics from \"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding\"). I used dynamic mini-batch alpha weights for CategoricalFocalCrossEntropy, but it was just for fun. Not sure it contributed much. II've trained for about 100 epochs 10K steps each with 2028 samples per batch - the model processed dataset for about 20 times. Note that I've combined all the data here - train, test and external data (reference is in the original text below).</p></li>\n<li><p>SMILES-to-ECFP (size=2048, include_chirality=True). Same model, just a different head (Dense layer with sigmoid activation) and locked embeddings. Some 20-50 epochs, if I recall correctly. The model did not performed great (there's many papers saying that SMILES encoders generally have a difficult time to predict topological fingerprints, and it was exactly the case) with MAP around 0.4, however, I guess that's exactly where it learned some useful representations.</p></li>\n<li><p>My motivation here was to train model on some general task without taking a major overfitting risk. I picked ECFP for 2 reasons (a) performance - they were fast enough to compute, especially with <a href=\"https://scikit-fingerprints.github.io/scikit-fingerprints/index.html\" target=\"_blank\">scikit-fingerprints</a> library, and (b) predicting fingerprints with no predefined meanings for each bit position (unlike MACCS or PubChem) is a challenging task for SMILES transformer, which is good - \"no pain - no gain\", as they say…</p>\n<p>And, yes, I still wonder what \"chirality\" is :)</p></li>\n</ul>\n<p>e) Training. Combined BELKA train set and external data. Masked loss and metrics since external data has labels just for sEH protein.</p>\n<p>f) Validation. I put aside 3% of blocks from train set so my validation set included molecules with one ore more non-shared blocks - some 9 million of samples.</p>\n<p>g) Tech - A100 on google collab.</p>\n<p>That's it. As being said - no magic, just a pure luck and randomness…</p>\n<hr>\n<p>Frankly speaking, the final LB results came as a bit of a surprise to me. The winning model is a very basic encoder: Self-Attention -&gt; FeedForward with 4 layers and 8 heads per layer and  key/value dimension of 32. The classics from the Transformers chapter of Tensorflow tutorials :)</p>\n<p>I used the <a href=\"https://pypi.org/project/atomInSmiles/\" target=\"_blank\">atomInSmiles</a> tokenizer, but I did it incorrectly, so my tokenization scheme was almost character-based. I have not used any pre-trained models like ChemBERTa or similar.</p>\n<p>The difference might have come from a two-stage pre-training schedule: (a) MLM - with 15% masking rate (b) SMILES to ECFP prediction. I'm not good at chemistry, to tell the truth, but I guess the second stage is where the encoder \"learned\" to extract some meaningful results from the SMILES.</p>\n<p>Oh, last but not least: I used the data provided by the competition host and the dataset from <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">\"Building Block-Based Binding Predictions for DNA-Encoded Libraries\"</a>, sited by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and preprocessed by <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> early in the competition.</p>\n<p>Now, the list of things that didn't work out as expected:<br>\n1) Complex tokenization schemes: bi- and tri-grams, atomInSmiles<br>\n2) Any model with a depth above 32 and more than 6 encoder layers<br>\n3) Multi-input models (SMILES + fingerprints)<br>\n4) Pre-training on a larger dataset—I spent about a month experimenting with ZINC…<br>\n5) Custom loss functions—BinaryFocusLoss was just fine<br>\n6) Gated fusion of building blocks<br>\n7) And many more—I will update the list.</p>",
  "messages": [
    {
      "id": 2913246,
      "postDate": "2024-07-09T11:04:51.180Z",
      "content": "<p>[UPDATE]</p>\n<p>a) Here is the <a href=\"https://www.kaggle.com/datasets/victorshlepov/belka2024-finalsolution\" target=\"_blank\">dataset</a> with the (i) code  - model and data processing utils (ii) SMILES encoder vocabulary. I will upload processed training data too (just to save ones time). Once I recover the the trained weights I will add them as well - I have only the light version of the model left (tf.keras.export), so I will likely retrain it from scratch. For those of you who prefer Kaggle notebooks - here is <a href=\"https://www.kaggle.com/code/victorshlepov/belka2024-1stplacesolution/edit\" target=\"_blank\">one</a> (but I would not really process data here - it should take close to infinity). </p>\n<p>b) The architecture is fairly simple and model is very flat - just 4 encoder layers with 8 heads. With a vocabulary size of just 43 tokens I end-up with a fixed dimensionality of 32. I've tried 64 and 16 too - they do not perform.</p>\n<p>c) I guess I used <a href=\"https://pypi.org/project/atomInSmiles/\" target=\"_blank\">atomInSmiles</a> in a somewhat incorrect way and end up with a schema where separate tokens are either, atom (C, H, S, etc)  or digits, or anything in square brackets, like [C@@] are distinct tokens. I leave it to chemistry practitioners to decide what this mess really means :)</p>\n<p>d) Pre-training. I pre-trained model from scratch in two stages:</p>\n<ul>\n<li><p>MLM - the standard prediction of masked tokens (15% of which 80% are masked, 10% are replaced with random token and 10% are retained - classics from \"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding\"). I used dynamic mini-batch alpha weights for CategoricalFocalCrossEntropy, but it was just for fun. Not sure it contributed much. II've trained for about 100 epochs 10K steps each with 2028 samples per batch - the model processed dataset for about 20 times. Note that I've combined all the data here - train, test and external data (reference is in the original text below).</p></li>\n<li><p>SMILES-to-ECFP (size=2048, include_chirality=True). Same model, just a different head (Dense layer with sigmoid activation) and locked embeddings. Some 20-50 epochs, if I recall correctly. The model did not performed great (there's many papers saying that SMILES encoders generally have a difficult time to predict topological fingerprints, and it was exactly the case) with MAP around 0.4, however, I guess that's exactly where it learned some useful representations.</p></li>\n<li><p>My motivation here was to train model on some general task without taking a major overfitting risk. I picked ECFP for 2 reasons (a) performance - they were fast enough to compute, especially with <a href=\"https://scikit-fingerprints.github.io/scikit-fingerprints/index.html\" target=\"_blank\">scikit-fingerprints</a> library, and (b) predicting fingerprints with no predefined meanings for each bit position (unlike MACCS or PubChem) is a challenging task for SMILES transformer, which is good - \"no pain - no gain\", as they say…</p>\n<p>And, yes, I still wonder what \"chirality\" is :)</p></li>\n</ul>\n<p>e) Training. Combined BELKA train set and external data. Masked loss and metrics since external data has labels just for sEH protein.</p>\n<p>f) Validation. I put aside 3% of blocks from train set so my validation set included molecules with one ore more non-shared blocks - some 9 million of samples.</p>\n<p>g) Tech - A100 on google collab.</p>\n<p>That's it. As being said - no magic, just a pure luck and randomness…</p>\n<hr>\n<p>Frankly speaking, the final LB results came as a bit of a surprise to me. The winning model is a very basic encoder: Self-Attention -&gt; FeedForward with 4 layers and 8 heads per layer and  key/value dimension of 32. The classics from the Transformers chapter of Tensorflow tutorials :)</p>\n<p>I used the <a href=\"https://pypi.org/project/atomInSmiles/\" target=\"_blank\">atomInSmiles</a> tokenizer, but I did it incorrectly, so my tokenization scheme was almost character-based. I have not used any pre-trained models like ChemBERTa or similar.</p>\n<p>The difference might have come from a two-stage pre-training schedule: (a) MLM - with 15% masking rate (b) SMILES to ECFP prediction. I'm not good at chemistry, to tell the truth, but I guess the second stage is where the encoder \"learned\" to extract some meaningful results from the SMILES.</p>\n<p>Oh, last but not least: I used the data provided by the competition host and the dataset from <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">\"Building Block-Based Binding Predictions for DNA-Encoded Libraries\"</a>, sited by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and preprocessed by <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> early in the competition.</p>\n<p>Now, the list of things that didn't work out as expected:<br>\n1) Complex tokenization schemes: bi- and tri-grams, atomInSmiles<br>\n2) Any model with a depth above 32 and more than 6 encoder layers<br>\n3) Multi-input models (SMILES + fingerprints)<br>\n4) Pre-training on a larger dataset—I spent about a month experimenting with ZINC…<br>\n5) Custom loss functions—BinaryFocusLoss was just fine<br>\n6) Gated fusion of building blocks<br>\n7) And many more—I will update the list.</p>",
      "rawMarkdown": "[UPDATE]\n\na) Here is the [dataset](https://www.kaggle.com/datasets/victorshlepov/belka2024-finalsolution) with the (i) code  - model and data processing utils (ii) SMILES encoder vocabulary. I will upload processed training data too (just to save ones time). Once I recover the the trained weights I will add them as well - I have only the light version of the model left (tf.keras.export), so I will likely retrain it from scratch. For those of you who prefer Kaggle notebooks - here is [one](https://www.kaggle.com/code/victorshlepov/belka2024-1stplacesolution/edit) (but I would not really process data here - it should take close to infinity). \n\nb) The architecture is fairly simple and model is very flat - just 4 encoder layers with 8 heads. With a vocabulary size of just 43 tokens I end-up with a fixed dimensionality of 32. I've tried 64 and 16 too - they do not perform.\n\nc) I guess I used [atomInSmiles](https://pypi.org/project/atomInSmiles/) in a somewhat incorrect way and end up with a schema where separate tokens are either, atom (C, H, S, etc)  or digits, or anything in square brackets, like [C@@] are distinct tokens. I leave it to chemistry practitioners to decide what this mess really means :)\n\nd) Pre-training. I pre-trained model from scratch in two stages:\n\n- MLM - the standard prediction of masked tokens (15% of which 80% are masked, 10% are replaced with random token and 10% are retained - classics from \"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding\"). I used dynamic mini-batch alpha weights for CategoricalFocalCrossEntropy, but it was just for fun. Not sure it contributed much. II've trained for about 100 epochs 10K steps each with 2028 samples per batch - the model processed dataset for about 20 times. Note that I've combined all the data here - train, test and external data (reference is in the original text below).\n\n- SMILES-to-ECFP (size=2048, include_chirality=True). Same model, just a different head (Dense layer with sigmoid activation) and locked embeddings. Some 20-50 epochs, if I recall correctly. The model did not performed great (there's many papers saying that SMILES encoders generally have a difficult time to predict topological fingerprints, and it was exactly the case) with MAP around 0.4, however, I guess that's exactly where it learned some useful representations.\n\n- My motivation here was to train model on some general task without taking a major overfitting risk. I picked ECFP for 2 reasons (a) performance - they were fast enough to compute, especially with [scikit-fingerprints](https://scikit-fingerprints.github.io/scikit-fingerprints/index.html) library, and (b) predicting fingerprints with no predefined meanings for each bit position (unlike MACCS or PubChem) is a challenging task for SMILES transformer, which is good - \"no pain - no gain\", as they say...\n\n And, yes, I still wonder what \"chirality\" is :)\n\ne) Training. Combined BELKA train set and external data. Masked loss and metrics since external data has labels just for sEH protein.\n\nf) Validation. I put aside 3% of blocks from train set so my validation set included molecules with one ore more non-shared blocks - some 9 million of samples.\n\ng) Tech - A100 on google collab.\n\nThat's it. As being said - no magic, just a pure luck and randomness...\n\n----\n\nFrankly speaking, the final LB results came as a bit of a surprise to me. The winning model is a very basic encoder: Self-Attention -> FeedForward with 4 layers and 8 heads per layer and  key/value dimension of 32. The classics from the Transformers chapter of Tensorflow tutorials :)\n\nI used the [atomInSmiles](https://pypi.org/project/atomInSmiles/) tokenizer, but I did it incorrectly, so my tokenization scheme was almost character-based. I have not used any pre-trained models like ChemBERTa or similar.\n\nThe difference might have come from a two-stage pre-training schedule: (a) MLM - with 15% masking rate (b) SMILES to ECFP prediction. I'm not good at chemistry, to tell the truth, but I guess the second stage is where the encoder \"learned\" to extract some meaningful results from the SMILES.\n\nOh, last but not least: I used the data provided by the competition host and the dataset from [\"Building Block-Based Binding Predictions for DNA-Encoded Libraries\"](https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57), sited by @hengck23 and preprocessed by @chemdatafarmer early in the competition.\n\nNow, the list of things that didn't work out as expected:\n1) Complex tokenization schemes: bi- and tri-grams, atomInSmiles\n2) Any model with a depth above 32 and more than 6 encoder layers\n3) Multi-input models (SMILES + fingerprints)\n4) Pre-training on a larger dataset—I spent about a month experimenting with ZINC...\n5) Custom loss functions—BinaryFocusLoss was just fine\n6) Gated fusion of building blocks\n7) And many more—I will update the list.",
      "votes": 85
    },
    {
      "id": 2931387,
      "postDate": "2024-07-22T01:57:14.293Z",
      "content": "<p>I think the extracted tokens (and so your use of atomInSMILES) are correct. I know SMILES pretty well and I looked at your tokens dictionary.</p>",
      "rawMarkdown": "I think the extracted tokens (and so your use of atomInSMILES) are correct. I know SMILES pretty well and I looked at your tokens dictionary.",
      "votes": 1
    },
    {
      "id": 2930971,
      "postDate": "2024-07-21T14:56:06.050Z",
      "content": "<p>Quite informative post, congrats on the win!</p>",
      "rawMarkdown": "Quite informative post, congrats on the win!",
      "votes": 1
    },
    {
      "id": 2928252,
      "postDate": "2024-07-19T05:00:26.220Z",
      "content": "<p>Congratulations!</p>\n<p>And thank you, the post sure is insightful </p>",
      "rawMarkdown": "Congratulations!\n\nAnd thank you, the post sure is insightful ",
      "votes": 1
    },
    {
      "id": 2926996,
      "postDate": "2024-07-18T07:35:21.537Z",
      "content": "<p>Woww! Simple approach and yet powerful….can you share your notebook if possible?</p>",
      "rawMarkdown": "Woww! Simple approach and yet powerful....can you share your notebook if possible?",
      "votes": 1,
      "replies": [
        {
          "id": 2931010,
          "postDate": "2024-07-21T15:25:53.887Z",
          "content": "<p>Sure, just published it</p>",
          "rawMarkdown": "Sure, just published it"
        }
      ]
    },
    {
      "id": 2919764,
      "postDate": "2024-07-13T07:27:30.923Z",
      "content": "<p>Congratulations!!! Great work!!!</p>",
      "rawMarkdown": "Congratulations!!! Great work!!!",
      "votes": 1
    },
    {
      "id": 2919433,
      "postDate": "2024-07-12T21:16:46.620Z",
      "content": "<p>Congratulations on the win! Thank you for sharing the approach as well as the things that didn't work, it's truly informative.</p>",
      "rawMarkdown": "Congratulations on the win! Thank you for sharing the approach as well as the things that didn't work, it's truly informative.",
      "votes": 1
    },
    {
      "id": 2919397,
      "postDate": "2024-07-12T20:37:53.230Z",
      "content": "<p>Congratulations on the winning solution! Please, link it as your team's solution using this instruction: <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/373153</a></p>",
      "rawMarkdown": "Congratulations on the winning solution! Please, link it as your team's solution using this instruction: https://www.kaggle.com/discussions/product-feedback/373153",
      "votes": 1
    },
    {
      "id": 2916777,
      "postDate": "2024-07-11T08:44:58.287Z",
      "content": "<p>Is there some code somewhere we can look at?</p>",
      "rawMarkdown": "Is there some code somewhere we can look at?",
      "votes": 1,
      "replies": [
        {
          "id": 2916874,
          "postDate": "2024-07-11T10:20:54.503Z",
          "content": "<p>Sure, I just need to recover it from Colab archives… Will do over the weekend…</p>",
          "rawMarkdown": "Sure, I just need to recover it from Colab archives… Will do over the weekend...",
          "votes": 1
        },
        {
          "id": 2931013,
          "postDate": "2024-07-21T15:27:44.730Z",
          "content": "<p>Just shared.</p>",
          "rawMarkdown": "Just shared.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2915611,
      "postDate": "2024-07-10T15:29:11.550Z",
      "content": "<p>Thanks for sharing this!</p>\n<p>I just wanted to know whether the key/value dimension that you used was per head or per attention layer?</p>",
      "rawMarkdown": "Thanks for sharing this!\n\nI just wanted to know whether the key/value dimension that you used was per head or per attention layer?",
      "votes": 1,
      "replies": [
        {
          "id": 2931014,
          "postDate": "2024-07-21T15:29:09.953Z",
          "content": "<p>Hi! Just published the code, hope it addresses the question…</p>",
          "rawMarkdown": "Hi! Just published the code, hope it addresses the question...",
          "votes": 1,
          "replies": [
            {
              "id": 2931452,
              "postDate": "2024-07-22T04:00:34.820Z",
              "content": "<p>What license is your code under?</p>",
              "rawMarkdown": "What license is your code under?"
            },
            {
              "id": 2933317,
              "postDate": "2024-07-23T15:57:39.623Z",
              "content": "<p>I don't know… I guess there's some general rules on Kaggle… Anyway, there's nothing really special about the code - not even close to \"Attention Is All You Need\" or \"ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction\" :)</p>",
              "rawMarkdown": "I don't know... I guess there's some general rules on Kaggle... Anyway, there's nothing really special about the code - not even close to \"Attention Is All You Need\" or \"ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction\" :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2914627,
      "postDate": "2024-07-10T04:44:08.717Z",
      "content": "<p>Congrats on the win!</p>",
      "rawMarkdown": "Congrats on the win!",
      "votes": 1
    },
    {
      "id": 2914380,
      "postDate": "2024-07-09T22:21:00.547Z",
      "content": "<p>Great job！I learnt a lot</p>",
      "rawMarkdown": "Great job！I learnt a lot",
      "votes": 1
    },
    {
      "id": 2913544,
      "postDate": "2024-07-09T14:18:55.953Z",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!",
      "votes": 1
    },
    {
      "id": 2913397,
      "postDate": "2024-07-09T13:08:10.303Z",
      "content": "<p>Great work! <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> !! Just, wow.</p>",
      "rawMarkdown": "Great work! @victorshlepov !! Just, wow.",
      "votes": 1
    },
    {
      "id": 2913277,
      "postDate": "2024-07-09T11:30:56.997Z",
      "content": "<p>WOw! thanks for sharing and congrats 🥳! <br>\ninteresting approach :) could you please share some code?</p>\n<p>I also wonder what part of the test was most successfully predicted? (at least which one of shared, non-shared triazines, and non-triazines) <br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518967\" target=\"_blank\">here</a> I have posted ours scores for non-triazines</p>",
      "rawMarkdown": "WOw! thanks for sharing and congrats 🥳! \ninteresting approach :) could you please share some code?\n\nI also wonder what part of the test was most successfully predicted? (at least which one of shared, non-shared triazines, and non-triazines) \n[here](https://www.kaggle.com/competitions/leash-BELKA/discussion/518967) I have posted ours scores for non-triazines",
      "votes": 1,
      "replies": [
        {
          "id": 2913287,
          "postDate": "2024-07-09T11:38:16.583Z",
          "content": "<p>Sure, will do - as soon as I find it in the Colab archives… :)</p>",
          "rawMarkdown": "Sure, will do - as soon as I find it in the Colab archives... :)",
          "votes": 2,
          "replies": [
            {
              "id": 2992794,
              "postDate": "2024-09-19T04:46:10.933Z",
              "content": "<p>we still highly waiting 🫠</p>",
              "rawMarkdown": "we still highly waiting 🫠"
            }
          ]
        },
        {
          "id": 2931016,
          "postDate": "2024-07-21T15:31:21.240Z",
          "content": "<p>Frankly, I have not checked the model performance by segments yet. At the end of the day, I'm not a real data-scientist :)</p>",
          "rawMarkdown": "Frankly, I have not checked the model performance by segments yet. At the end of the day, I'm not a real data-scientist :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2914580,
      "postDate": "2024-07-10T03:54:50.357Z",
      "content": "<p>Congratulations on winning this competition. Thanks for sharing details of your approach. </p>",
      "rawMarkdown": "Congratulations on winning this competition. Thanks for sharing details of your approach. ",
      "votes": 2
    },
    {
      "id": 2914468,
      "postDate": "2024-07-10T00:38:06.353Z",
      "content": "<p>Congratulations on winning the competition! Your innovative approach and persistence have paid off. Well done!</p>",
      "rawMarkdown": "Congratulations on winning the competition! Your innovative approach and persistence have paid off. Well done!",
      "votes": 2
    },
    {
      "id": 2913593,
      "postDate": "2024-07-09T15:08:59.210Z",
      "content": "<p>So you pretrained the model to predict ECFP first, and then with same inputs (smiles in both case) that took the pretrained final output layer as input to the transformer?</p>\n<p>Or something like that?</p>\n<p>Smart! As much randomness as there was, honestly something like that is what I would've expected to win, something that avoids just learning the character strings matched to \"good BBs\". </p>",
      "rawMarkdown": "So you pretrained the model to predict ECFP first, and then with same inputs (smiles in both case) that took the pretrained final output layer as input to the transformer?\n\nOr something like that?\n\nSmart! As much randomness as there was, honestly something like that is what I would've expected to win, something that avoids just learning the character strings matched to \"good BBs\". ",
      "votes": 2,
      "replies": [
        {
          "id": 2913653,
          "postDate": "2024-07-09T15:36:08.037Z",
          "content": "<p>Almost like that. I've pre-trained transformer on MLM/ECFP tasks (input - SMILES, outputs - SMILES and ECFP) and then just changed the head (dense layer with 3 units and sigmoid activation). </p>\n<p>Well, you never know, but pre-training on a larger task with low overfitting risks is usually my first move…</p>",
          "rawMarkdown": "Almost like that. I've pre-trained transformer on MLM/ECFP tasks (input - SMILES, outputs - SMILES and ECFP) and then just changed the head (dense layer with 3 units and sigmoid activation). \n\nWell, you never know, but pre-training on a larger task with low overfitting risks is usually my first move...",
          "votes": 6,
          "replies": [
            {
              "id": 2918029,
              "postDate": "2024-07-12T01:00:39.283Z",
              "content": "<p>What a brilliant approach, congratulations on your win! Would you mind sharing how much data you used and how many epochs you spent on pre-training and post-training, respectively?</p>",
              "rawMarkdown": "What a brilliant approach, congratulations on your win! Would you mind sharing how much data you used and how many epochs you spent on pre-training and post-training, respectively?",
              "votes": 1
            },
            {
              "id": 2930876,
              "postDate": "2024-07-21T13:42:28.600Z",
              "content": "<p>Done :) Let me know if missed some of the questions, OK?</p>",
              "rawMarkdown": "Done :) Let me know if missed some of the questions, OK?",
              "votes": 1
            },
            {
              "id": 2931394,
              "postDate": "2024-07-22T02:19:29.280Z",
              "content": "<p>Thank you! No further questions for now. Congratulations again!</p>",
              "rawMarkdown": "Thank you! No further questions for now. Congratulations again!"
            }
          ]
        }
      ]
    },
    {
      "id": 2992793,
      "postDate": "2024-09-19T04:44:51.130Z",
      "content": "<p>can we have a check to the winning notbook ? i'll appriciat it </p>",
      "rawMarkdown": "can we have a check to the winning notbook ? i'll appriciat it "
    },
    {
      "id": 2942176,
      "postDate": "2024-07-31T15:57:46.637Z",
      "content": "<p>Congratulations and thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations and thanks for sharing :)"
    },
    {
      "id": 2941267,
      "postDate": "2024-07-30T19:47:14.150Z",
      "content": "<p>Congratulations on winning this competition. Thanks for sharing your approach.</p>",
      "rawMarkdown": "Congratulations on winning this competition. Thanks for sharing your approach."
    },
    {
      "id": 2937294,
      "postDate": "2024-07-26T23:07:09.620Z",
      "content": "<p>In the additional dataset, how did you decide on 1 and 0 for your labels (binding to sEH)? Did you process read_count somehow?</p>",
      "rawMarkdown": "In the additional dataset, how did you decide on 1 and 0 for your labels (binding to sEH)? Did you process read_count somehow?",
      "replies": [
        {
          "id": 2938238,
          "postDate": "2024-07-27T21:27:21.880Z",
          "content": "<p>Just any read_count above zero…</p>",
          "rawMarkdown": "Just any read_count above zero...",
          "votes": 1
        }
      ]
    },
    {
      "id": 2937292,
      "postDate": "2024-07-26T23:06:04.137Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2914635,
      "postDate": "2024-07-10T04:59:08.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2914112,
      "postDate": "2024-07-09T18:22:51.077Z",
      "content": "<p>Nice! Awesome job, thanks for sharing!</p>",
      "rawMarkdown": "Nice! Awesome job, thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2932523,
      "postDate": "2024-07-23T04:42:23.580Z",
      "content": "<p>Congratulations!, thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations!, thanks for sharing :)"
    }
  ],
  "comments": [
    {
      "id": 2931387,
      "author_name": "Francois Berenger",
      "author_url": "",
      "post_date": "2024-07-22T01:57:14.293000",
      "content": "<p>I think the extracted tokens (and so your use of atomInSMILES) are correct. I know SMILES pretty well and I looked at your tokens dictionary.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2930971,
      "author_name": "Luigi Brancati",
      "author_url": "",
      "post_date": "2024-07-21T14:56:06.050000",
      "content": "<p>Quite informative post, congrats on the win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2928252,
      "author_name": "Rishit Jakharia",
      "author_url": "",
      "post_date": "2024-07-19T05:00:26.220000",
      "content": "<p>Congratulations!</p>\n<p>And thank you, the post sure is insightful </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2926996,
      "author_name": "amulya_incorrigible",
      "author_url": "",
      "post_date": "2024-07-18T07:35:21.537000",
      "content": "<p>Woww! Simple approach and yet powerful….can you share your notebook if possible?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2931010,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-21T15:25:53.887000",
          "content": "<p>Sure, just published it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2919764,
      "author_name": "Gitanshu Vaghasiya",
      "author_url": "",
      "post_date": "2024-07-13T07:27:30.923000",
      "content": "<p>Congratulations!!! Great work!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2919433,
      "author_name": "Kaavya Mahajan",
      "author_url": "",
      "post_date": "2024-07-12T21:16:46.620000",
      "content": "<p>Congratulations on the win! Thank you for sharing the approach as well as the things that didn't work, it's truly informative.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2919397,
      "author_name": "Vlad Vinogradov",
      "author_url": "",
      "post_date": "2024-07-12T20:37:53.230000",
      "content": "<p>Congratulations on the winning solution! Please, link it as your team's solution using this instruction: <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/373153</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2916777,
      "author_name": "Francois Berenger",
      "author_url": "",
      "post_date": "2024-07-11T08:44:58.287000",
      "content": "<p>Is there some code somewhere we can look at?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2916874,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-11T10:20:54.503000",
          "content": "<p>Sure, I just need to recover it from Colab archives… Will do over the weekend…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2931013,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-21T15:27:44.730000",
          "content": "<p>Just shared.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2915611,
      "author_name": "Andrius Bernatavicius",
      "author_url": "",
      "post_date": "2024-07-10T15:29:11.550000",
      "content": "<p>Thanks for sharing this!</p>\n<p>I just wanted to know whether the key/value dimension that you used was per head or per attention layer?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2931014,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-21T15:29:09.953000",
          "content": "<p>Hi! Just published the code, hope it addresses the question…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2931452,
              "author_name": "Francois Berenger",
              "author_url": "",
              "post_date": "2024-07-22T04:00:34.820000",
              "content": "<p>What license is your code under?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2933317,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-07-23T15:57:39.623000",
              "content": "<p>I don't know… I guess there's some general rules on Kaggle… Anyway, there's nothing really special about the code - not even close to \"Attention Is All You Need\" or \"ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction\" :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914627,
      "author_name": "Aaron Ma",
      "author_url": "",
      "post_date": "2024-07-10T04:44:08.717000",
      "content": "<p>Congrats on the win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2914380,
      "author_name": "qtx202",
      "author_url": "",
      "post_date": "2024-07-09T22:21:00.547000",
      "content": "<p>Great job！I learnt a lot</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2913544,
      "author_name": "kostopr4v",
      "author_url": "",
      "post_date": "2024-07-09T14:18:55.953000",
      "content": "<p>Congratulations!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2913397,
      "author_name": "William Zebrowski",
      "author_url": "",
      "post_date": "2024-07-09T13:08:10.303000",
      "content": "<p>Great work! <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> !! Just, wow.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2913277,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-07-09T11:30:56.997000",
      "content": "<p>WOw! thanks for sharing and congrats 🥳! <br>\ninteresting approach :) could you please share some code?</p>\n<p>I also wonder what part of the test was most successfully predicted? (at least which one of shared, non-shared triazines, and non-triazines) <br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/518967\" target=\"_blank\">here</a> I have posted ours scores for non-triazines</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2913287,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-09T11:38:16.583000",
          "content": "<p>Sure, will do - as soon as I find it in the Colab archives… :)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2992794,
              "author_name": "David khaldi",
              "author_url": "",
              "post_date": "2024-09-19T04:46:10.933000",
              "content": "<p>we still highly waiting 🫠</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2931016,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-21T15:31:21.240000",
          "content": "<p>Frankly, I have not checked the model performance by segments yet. At the end of the day, I'm not a real data-scientist :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2914580,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-07-10T03:54:50.357000",
      "content": "<p>Congratulations on winning this competition. Thanks for sharing details of your approach. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2914468,
      "author_name": "Jack Li",
      "author_url": "",
      "post_date": "2024-07-10T00:38:06.353000",
      "content": "<p>Congratulations on winning the competition! Your innovative approach and persistence have paid off. Well done!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2913593,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-07-09T15:08:59.210000",
      "content": "<p>So you pretrained the model to predict ECFP first, and then with same inputs (smiles in both case) that took the pretrained final output layer as input to the transformer?</p>\n<p>Or something like that?</p>\n<p>Smart! As much randomness as there was, honestly something like that is what I would've expected to win, something that avoids just learning the character strings matched to \"good BBs\". </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2913653,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-09T15:36:08.037000",
          "content": "<p>Almost like that. I've pre-trained transformer on MLM/ECFP tasks (input - SMILES, outputs - SMILES and ECFP) and then just changed the head (dense layer with 3 units and sigmoid activation). </p>\n<p>Well, you never know, but pre-training on a larger task with low overfitting risks is usually my first move…</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2918029,
              "author_name": "Frenio Redeker",
              "author_url": "",
              "post_date": "2024-07-12T01:00:39.283000",
              "content": "<p>What a brilliant approach, congratulations on your win! Would you mind sharing how much data you used and how many epochs you spent on pre-training and post-training, respectively?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2930876,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-07-21T13:42:28.600000",
              "content": "<p>Done :) Let me know if missed some of the questions, OK?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2931394,
              "author_name": "Frenio Redeker",
              "author_url": "",
              "post_date": "2024-07-22T02:19:29.280000",
              "content": "<p>Thank you! No further questions for now. Congratulations again!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2992793,
      "author_name": "David khaldi",
      "author_url": "",
      "post_date": "2024-09-19T04:44:51.130000",
      "content": "<p>can we have a check to the winning notbook ? i'll appriciat it </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2942176,
      "author_name": "Akshhat Shethia",
      "author_url": "",
      "post_date": "2024-07-31T15:57:46.637000",
      "content": "<p>Congratulations and thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2941267,
      "author_name": "yoruoo",
      "author_url": "",
      "post_date": "2024-07-30T19:47:14.150000",
      "content": "<p>Congratulations on winning this competition. Thanks for sharing your approach.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2937294,
      "author_name": "Arjan Hada",
      "author_url": "",
      "post_date": "2024-07-26T23:07:09.620000",
      "content": "<p>In the additional dataset, how did you decide on 1 and 0 for your labels (binding to sEH)? Did you process read_count somehow?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2938238,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-07-27T21:27:21.880000",
          "content": "<p>Just any read_count above zero…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2937292,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-26T23:06:04.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2914635,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-10T04:59:08.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2914112,
      "author_name": "Gcdelgado",
      "author_url": "",
      "post_date": "2024-07-09T18:22:51.077000",
      "content": "<p>Nice! Awesome job, thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2932523,
      "author_name": "rhubain mageswaran",
      "author_url": "",
      "post_date": "2024-07-23T04:42:23.580000",
      "content": "<p>Congratulations!, thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2913246": "[UPDATE]\n\na) Here is the [dataset](https://www.kaggle.com/datasets/victorshlepov/belka2024-finalsolution) with the (i) code  - model and data processing utils (ii) SMILES encoder vocabulary. I will upload processed training data too (just to save ones time). Once I recover the the trained weights I will add them as well - I have only the light version of the model left (tf.keras.export), so I will likely retrain it from scratch. For those of you who prefer Kaggle notebooks - here is [one](https://www.kaggle.com/code/victorshlepov/belka2024-1stplacesolution/edit) (but I would not really process data here - it should take close to infinity). \n\nb) The architecture is fairly simple and model is very flat - just 4 encoder layers with 8 heads. With a vocabulary size of just 43 tokens I end-up with a fixed dimensionality of 32. I've tried 64 and 16 too - they do not perform.\n\nc) I guess I used [atomInSmiles](https://pypi.org/project/atomInSmiles/) in a somewhat incorrect way and end up with a schema where separate tokens are either, atom (C, H, S, etc)  or digits, or anything in square brackets, like [C@@] are distinct tokens. I leave it to chemistry practitioners to decide what this mess really means :)\n\nd) Pre-training. I pre-trained model from scratch in two stages:\n\n- MLM - the standard prediction of masked tokens (15% of which 80% are masked, 10% are replaced with random token and 10% are retained - classics from \"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding\"). I used dynamic mini-batch alpha weights for CategoricalFocalCrossEntropy, but it was just for fun. Not sure it contributed much. II've trained for about 100 epochs 10K steps each with 2028 samples per batch - the model processed dataset for about 20 times. Note that I've combined all the data here - train, test and external data (reference is in the original text below).\n\n- SMILES-to-ECFP (size=2048, include_chirality=True). Same model, just a different head (Dense layer with sigmoid activation) and locked embeddings. Some 20-50 epochs, if I recall correctly. The model did not performed great (there's many papers saying that SMILES encoders generally have a difficult time to predict topological fingerprints, and it was exactly the case) with MAP around 0.4, however, I guess that's exactly where it learned some useful representations.\n\n- My motivation here was to train model on some general task without taking a major overfitting risk. I picked ECFP for 2 reasons (a) performance - they were fast enough to compute, especially with [scikit-fingerprints](https://scikit-fingerprints.github.io/scikit-fingerprints/index.html) library, and (b) predicting fingerprints with no predefined meanings for each bit position (unlike MACCS or PubChem) is a challenging task for SMILES transformer, which is good - \"no pain - no gain\", as they say...\n\n And, yes, I still wonder what \"chirality\" is :)\n\ne) Training. Combined BELKA train set and external data. Masked loss and metrics since external data has labels just for sEH protein.\n\nf) Validation. I put aside 3% of blocks from train set so my validation set included molecules with one ore more non-shared blocks - some 9 million of samples.\n\ng) Tech - A100 on google collab.\n\nThat's it. As being said - no magic, just a pure luck and randomness...\n\n----\n\nFrankly speaking, the final LB results came as a bit of a surprise to me. The winning model is a very basic encoder: Self-Attention -> FeedForward with 4 layers and 8 heads per layer and  key/value dimension of 32. The classics from the Transformers chapter of Tensorflow tutorials :)\n\nI used the [atomInSmiles](https://pypi.org/project/atomInSmiles/) tokenizer, but I did it incorrectly, so my tokenization scheme was almost character-based. I have not used any pre-trained models like ChemBERTa or similar.\n\nThe difference might have come from a two-stage pre-training schedule: (a) MLM - with 15% masking rate (b) SMILES to ECFP prediction. I'm not good at chemistry, to tell the truth, but I guess the second stage is where the encoder \"learned\" to extract some meaningful results from the SMILES.\n\nOh, last but not least: I used the data provided by the competition host and the dataset from [\"Building Block-Based Binding Predictions for DNA-Encoded Libraries\"](https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57), sited by @hengck23 and preprocessed by @chemdatafarmer early in the competition.\n\nNow, the list of things that didn't work out as expected:\n1) Complex tokenization schemes: bi- and tri-grams, atomInSmiles\n2) Any model with a depth above 32 and more than 6 encoder layers\n3) Multi-input models (SMILES + fingerprints)\n4) Pre-training on a larger dataset—I spent about a month experimenting with ZINC...\n5) Custom loss functions—BinaryFocusLoss was just fine\n6) Gated fusion of building blocks\n7) And many more—I will update the list.",
    "2931387": "I think the extracted tokens (and so your use of atomInSMILES) are correct. I know SMILES pretty well and I looked at your tokens dictionary.",
    "2930971": "Quite informative post, congrats on the win!",
    "2928252": "Congratulations!\n\nAnd thank you, the post sure is insightful ",
    "2926996": "Woww! Simple approach and yet powerful....can you share your notebook if possible?",
    "2919764": "Congratulations!!! Great work!!!",
    "2919433": "Congratulations on the win! Thank you for sharing the approach as well as the things that didn't work, it's truly informative.",
    "2919397": "Congratulations on the winning solution! Please, link it as your team's solution using this instruction: https://www.kaggle.com/discussions/product-feedback/373153",
    "2916777": "Is there some code somewhere we can look at?",
    "2915611": "Thanks for sharing this!\n\nI just wanted to know whether the key/value dimension that you used was per head or per attention layer?",
    "2914627": "Congrats on the win!",
    "2914380": "Great job！I learnt a lot",
    "2913544": "Congratulations!!!",
    "2913397": "Great work! @victorshlepov !! Just, wow.",
    "2913277": "WOw! thanks for sharing and congrats 🥳! \ninteresting approach :) could you please share some code?\n\nI also wonder what part of the test was most successfully predicted? (at least which one of shared, non-shared triazines, and non-triazines) \n[here](https://www.kaggle.com/competitions/leash-BELKA/discussion/518967) I have posted ours scores for non-triazines",
    "2914580": "Congratulations on winning this competition. Thanks for sharing details of your approach. ",
    "2914468": "Congratulations on winning the competition! Your innovative approach and persistence have paid off. Well done!",
    "2913593": "So you pretrained the model to predict ECFP first, and then with same inputs (smiles in both case) that took the pretrained final output layer as input to the transformer?\n\nOr something like that?\n\nSmart! As much randomness as there was, honestly something like that is what I would've expected to win, something that avoids just learning the character strings matched to \"good BBs\". ",
    "2992793": "can we have a check to the winning notbook ? i'll appriciat it ",
    "2942176": "Congratulations and thanks for sharing :)",
    "2941267": "Congratulations on winning this competition. Thanks for sharing your approach.",
    "2937294": "In the additional dataset, how did you decide on 1 and 0 for your labels (binding to sEH)? Did you process read_count somehow?",
    "2937292": "",
    "2914635": "",
    "2914112": "Nice! Awesome job, thanks for sharing!",
    "2932523": "Congratulations!, thanks for sharing :)"
  }
}