{
  "id": 441550,
  "title": "ChemBERTa v2 Embeddings for smiles",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/441550",
  "author_name": "Aleksey Trepetsky",
  "post_date": "2023-09-19T08:28:28.144000",
  "votes": 24,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hello everyone! I have generated embeddings for smiles using ChemBERTa v2.<br>\nIt was trained on 77M smiles from the PubChem database and fine-tuned on BACE, Clearance, Delaney, Lipophilicity, BBBP, ClinTox, HIV, Delaney, and Tox21 datasets.</p>\n<p><a href=\"https://arxiv.org/abs/2209.01712\" target=\"_blank\">arxiv</a><br>\n<a href=\"https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr\" target=\"_blank\">Dataset</a><br>\n<a href=\"https://www.kaggle.com/code/alekseytrepetsky/create-chemberta-embed\" target=\"_blank\">Notebook</a></p>",
  "messages": [
    {
      "id": 2446075,
      "postDate": "2023-09-19T08:28:28.143Z",
      "content": "<p>Hello everyone! I have generated embeddings for smiles using ChemBERTa v2.<br>\nIt was trained on 77M smiles from the PubChem database and fine-tuned on BACE, Clearance, Delaney, Lipophilicity, BBBP, ClinTox, HIV, Delaney, and Tox21 datasets.</p>\n<p><a href=\"https://arxiv.org/abs/2209.01712\" target=\"_blank\">arxiv</a><br>\n<a href=\"https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr\" target=\"_blank\">Dataset</a><br>\n<a href=\"https://www.kaggle.com/code/alekseytrepetsky/create-chemberta-embed\" target=\"_blank\">Notebook</a></p>",
      "rawMarkdown": "Hello everyone! I have generated embeddings for smiles using ChemBERTa v2.\nIt was trained on 77M smiles from the PubChem database and fine-tuned on BACE, Clearance, Delaney, Lipophilicity, BBBP, ClinTox, HIV, Delaney, and Tox21 datasets.\n\n[arxiv](https://arxiv.org/abs/2209.01712)\n[Dataset](https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr)\n[Notebook](https://www.kaggle.com/code/alekseytrepetsky/create-chemberta-embed)",
      "votes": 24
    },
    {
      "id": 2464126,
      "postDate": "2023-10-02T01:20:23.450Z",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a>, thanks for sharing. It was very helpful. I have a suggestion that it would be better not to use the lm_head for extracting features from SMILES, as I believe the pretrained checkpoint doesn't include the corresponding weights. Indeed, I received a warning message indicating the lm_head was randomly initialized when I loaded the pretrained model. In practice, you can skip the lm_head by doing the following;<br>\n<code>chemberta._modules[\"lm_head\"] = nn.Identity()</code><br>\nThen we can expect 384-dim features instead. I appreciate that you find this helpful.</p>",
      "rawMarkdown": "@alekseytrepetsky, thanks for sharing. It was very helpful. I have a suggestion that it would be better not to use the lm_head for extracting features from SMILES, as I believe the pretrained checkpoint doesn't include the corresponding weights. Indeed, I received a warning message indicating the lm_head was randomly initialized when I loaded the pretrained model. In practice, you can skip the lm_head by doing the following;\n`chemberta._modules[\"lm_head\"] = nn.Identity()`\nThen we can expect 384-dim features instead. I appreciate that you find this helpful.",
      "votes": 10
    },
    {
      "id": 2509138,
      "postDate": "2023-11-02T07:39:12.717Z",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> , Thanks for your great job. But I'm confused when I apply this code on my drug data.<br>\nFor example, I tried to get this smiles embedding <code>COC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl</code>, But the tokenizer incorrectly consider the last 3 <code>Cl</code> as <code>C</code>. Also I checked the token vocab, there indeed has <code>Cl</code> it dosen't make sense.</p>\n<p>The splited seq: </p>\n<pre><code>COC1=(C=(=C1)CCN=C2C3=(=(C=C3)Cl)Cl)Cl\n\n\n</code></pre>",
      "rawMarkdown": "@alekseytrepetsky , Thanks for your great job. But I'm confused when I apply this code on my drug data.\nFor example, I tried to get this smiles embedding `COC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl`, But the tokenizer incorrectly consider the last 3 `Cl` as `C`. Also I checked the token vocab, there indeed has `Cl` it dosen't make sense.\n\nThe splited seq: \n```\nCOC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl\n\n['C', 'O', 'C', '1', '=', 'C', '(', 'C', '=', 'C', '2', 'C', '(', '=', 'C', '1', ')', 'C', 'C', 'N', '=', 'C', '2', 'C', '3', '=', 'C', 'C', '(', '=', 'C', '(', 'C', '=', 'C', '3', ')', 'C', ')', 'C', ')', 'C']\n```\n\n\n",
      "votes": 3
    },
    {
      "id": 2457674,
      "postDate": "2023-09-27T05:51:00.013Z",
      "content": "<p>Thank you very much for sharing! <br>\nCould you please explain the difference between \"pad_true\" and \"pad_false\" and also between \"cls_pad\" and \"mean_pad\" files?</p>",
      "rawMarkdown": "Thank you very much for sharing! \nCould you please explain the difference between \"pad_true\" and \"pad_false\" and also between \"cls_pad\" and \"mean_pad\" files?",
      "votes": 3,
      "replies": [
        {
          "id": 2460108,
          "postDate": "2023-09-28T15:22:03.030Z",
          "content": "<p>Antonina, regarding cls and mean. Before feeding smiles into Chemberta2, they are tokenized, meaning they are converted into numbers. Additionally, a special token CLS is added at the beginning. After passing through Chemberta2, embeddings are obtained. One embedding is for the cls token, which contains information about the entire sequence and has a dimension of (1, 600). The embeddings for each token individually also exist. As a result, we have a matrix of size (len_sequence + 1, 600). However, it would be more convenient to work with a vector rather than a matrix, and there are two ways to handle this.</p>\n<p>One way is to simply take the embedding of the cls token. The other way is to aggregate the matrix into a vector, for example, by averaging.</p>\n<p>I would like to note that in protein embeddings for the CAFA5 competition, the mean approach yielded slightly better results than cls.</p>\n<p>Regarding pad, it refers to padding. It involves adding zeros to the sequences to make them of equal length. This is useful when you want to use batches for computing embeddings. However, in this case, I did not use batches and computed the embeddings one by one. Therefore, padding was not actually applied, and the embeddings in both folders are the same. This was my mistake, and I apologize for confusing you. Thank you for asking the question. I will fix it today by deleting the pad_true folder.</p>",
          "rawMarkdown": "Antonina, regarding cls and mean. Before feeding smiles into Chemberta2, they are tokenized, meaning they are converted into numbers. Additionally, a special token CLS is added at the beginning. After passing through Chemberta2, embeddings are obtained. One embedding is for the cls token, which contains information about the entire sequence and has a dimension of (1, 600). The embeddings for each token individually also exist. As a result, we have a matrix of size (len_sequence + 1, 600). However, it would be more convenient to work with a vector rather than a matrix, and there are two ways to handle this.\n\nOne way is to simply take the embedding of the cls token. The other way is to aggregate the matrix into a vector, for example, by averaging.\n\nI would like to note that in protein embeddings for the CAFA5 competition, the mean approach yielded slightly better results than cls.\n\nRegarding pad, it refers to padding. It involves adding zeros to the sequences to make them of equal length. This is useful when you want to use batches for computing embeddings. However, in this case, I did not use batches and computed the embeddings one by one. Therefore, padding was not actually applied, and the embeddings in both folders are the same. This was my mistake, and I apologize for confusing you. Thank you for asking the question. I will fix it today by deleting the pad_true folder.",
          "votes": 7,
          "replies": [
            {
              "id": 2461127,
              "postDate": "2023-09-29T09:50:28.597Z",
              "content": "<p>thank you for such a detailed answer! Very helpful!</p>",
              "rawMarkdown": "thank you for such a detailed answer! Very helpful!",
              "votes": 1
            },
            {
              "id": 2461823,
              "postDate": "2023-09-29T19:31:06.353Z",
              "content": "<p>Can you elaborate a little more on the CLS token? What exactly is it? What is its purpose? How come it encodes the information of the entire SMILE sequence in it?</p>\n<p>Is it possible to obtain an interpretation of these embeddings into more intuitive features like, types and number of atoms, bond, chirality etc?</p>",
              "rawMarkdown": "Can you elaborate a little more on the CLS token? What exactly is it? What is its purpose? How come it encodes the information of the entire SMILE sequence in it?\n\nIs it possible to obtain an interpretation of these embeddings into more intuitive features like, types and number of atoms, bond, chirality etc?",
              "votes": 1
            },
            {
              "id": 2488053,
              "postDate": "2023-10-19T01:42:16.180Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> , I want to get a drug's emb matrix, but I'm confused about the matrix shape. The smiles' length is 51, but the <strong>model_output[0]</strong> shape is <strong>(53, 600)</strong>. As you said above, there is a cls token and the shape should be <strong>(52,600)</strong>.</p>",
              "rawMarkdown": "Hi @alekseytrepetsky , I want to get a drug's emb matrix, but I'm confused about the matrix shape. The smiles' length is 51, but the **model_output[0]** shape is **(53, 600)**. As you said above, there is a cls token and the shape should be **(52,600)**."
            }
          ]
        }
      ]
    },
    {
      "id": 2446151,
      "postDate": "2023-09-19T09:23:22.020Z",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> a lot of effort thanks for sharing </p>",
      "rawMarkdown": "@alekseytrepetsky a lot of effort thanks for sharing ",
      "votes": 2
    },
    {
      "id": 2488882,
      "postDate": "2023-10-19T15:05:50.313Z",
      "content": "<p>How did you train it? cause it doesn't seems like it was trained or did you just loaded the pre-trained model?<br>\ncause if you do it with \"torch.no_grad()\" pytorch won't take that loss into account no?</p>",
      "rawMarkdown": "How did you train it? cause it doesn't seems like it was trained or did you just loaded the pre-trained model?\ncause if you do it with \"torch.no_grad()\" pytorch won't take that loss into account no?"
    },
    {
      "id": 2447776,
      "postDate": "2023-09-20T08:43:25.607Z",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> I think your data has a lot of redundance.</p>\n<p>There are only 147 SMILES, which are the same in train and test. But in your dataset, there are 614 for train and 255 for test.<br>\nYou are duplicate the data (614+255)/147 ≈ 6 times!</p>",
      "rawMarkdown": "@alekseytrepetsky I think your data has a lot of redundance.\n\nThere are only 147 SMILES, which are the same in train and test. But in your dataset, there are 614 for train and 255 for test.\nYou are duplicate the data (614+255)/147 ≈ 6 times!",
      "replies": [
        {
          "id": 2449983,
          "postDate": "2023-09-21T14:55:11.253Z",
          "content": "<p>Thank you for your observation. Yes, there is indeed duplication of data, and it was done intentionally. The embeddings were created in the same sequence as the chemical compounds in the train and test sets. This way, you can simply replace the sm_name in the datasets with the embeddings and use them. I thought it would be more convenient this way. Also, we have a small amount of data, so duplicating it is not critical.<br>\nHowever, if you find it more convenient to work with non-duplicated embeddings, please let me know, and I can add data without duplication to the dataset.</p>",
          "rawMarkdown": "Thank you for your observation. Yes, there is indeed duplication of data, and it was done intentionally. The embeddings were created in the same sequence as the chemical compounds in the train and test sets. This way, you can simply replace the sm_name in the datasets with the embeddings and use them. I thought it would be more convenient this way. Also, we have a small amount of data, so duplicating it is not critical.\nHowever, if you find it more convenient to work with non-duplicated embeddings, please let me know, and I can add data without duplication to the dataset.",
          "votes": 3,
          "replies": [
            {
              "id": 2451071,
              "postDate": "2023-09-22T09:45:12.860Z",
              "content": "<p>ok, that make sense.</p>",
              "rawMarkdown": "ok, that make sense."
            }
          ]
        }
      ]
    },
    {
      "id": 2485877,
      "postDate": "2023-10-17T14:08:58.713Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2462376,
      "postDate": "2023-09-30T09:57:21.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2447508,
      "postDate": "2023-09-20T05:57:50.707Z",
      "content": "<p>Great jobs and thanks for sharing !!!</p>",
      "rawMarkdown": "Great jobs and thanks for sharing !!!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2464126,
      "author_name": "Mt.Panda",
      "author_url": "",
      "post_date": "2023-10-02T01:20:23.450000",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a>, thanks for sharing. It was very helpful. I have a suggestion that it would be better not to use the lm_head for extracting features from SMILES, as I believe the pretrained checkpoint doesn't include the corresponding weights. Indeed, I received a warning message indicating the lm_head was randomly initialized when I loaded the pretrained model. In practice, you can skip the lm_head by doing the following;<br>\n<code>chemberta._modules[\"lm_head\"] = nn.Identity()</code><br>\nThen we can expect 384-dim features instead. I appreciate that you find this helpful.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 2509138,
      "author_name": "Chris Tang 0002",
      "author_url": "",
      "post_date": "2023-11-02T07:39:12.717000",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> , Thanks for your great job. But I'm confused when I apply this code on my drug data.<br>\nFor example, I tried to get this smiles embedding <code>COC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl</code>, But the tokenizer incorrectly consider the last 3 <code>Cl</code> as <code>C</code>. Also I checked the token vocab, there indeed has <code>Cl</code> it dosen't make sense.</p>\n<p>The splited seq: </p>\n<pre><code>COC1=(C=(=C1)CCN=C2C3=(=(C=C3)Cl)Cl)Cl\n\n\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2457674,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2023-09-27T05:51:00.013000",
      "content": "<p>Thank you very much for sharing! <br>\nCould you please explain the difference between \"pad_true\" and \"pad_false\" and also between \"cls_pad\" and \"mean_pad\" files?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2460108,
          "author_name": "Aleksey Trepetsky",
          "author_url": "",
          "post_date": "2023-09-28T15:22:03.030000",
          "content": "<p>Antonina, regarding cls and mean. Before feeding smiles into Chemberta2, they are tokenized, meaning they are converted into numbers. Additionally, a special token CLS is added at the beginning. After passing through Chemberta2, embeddings are obtained. One embedding is for the cls token, which contains information about the entire sequence and has a dimension of (1, 600). The embeddings for each token individually also exist. As a result, we have a matrix of size (len_sequence + 1, 600). However, it would be more convenient to work with a vector rather than a matrix, and there are two ways to handle this.</p>\n<p>One way is to simply take the embedding of the cls token. The other way is to aggregate the matrix into a vector, for example, by averaging.</p>\n<p>I would like to note that in protein embeddings for the CAFA5 competition, the mean approach yielded slightly better results than cls.</p>\n<p>Regarding pad, it refers to padding. It involves adding zeros to the sequences to make them of equal length. This is useful when you want to use batches for computing embeddings. However, in this case, I did not use batches and computed the embeddings one by one. Therefore, padding was not actually applied, and the embeddings in both folders are the same. This was my mistake, and I apologize for confusing you. Thank you for asking the question. I will fix it today by deleting the pad_true folder.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2461127,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2023-09-29T09:50:28.597000",
              "content": "<p>thank you for such a detailed answer! Very helpful!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2461823,
              "author_name": "arpitamhow",
              "author_url": "",
              "post_date": "2023-09-29T19:31:06.353000",
              "content": "<p>Can you elaborate a little more on the CLS token? What exactly is it? What is its purpose? How come it encodes the information of the entire SMILE sequence in it?</p>\n<p>Is it possible to obtain an interpretation of these embeddings into more intuitive features like, types and number of atoms, bond, chirality etc?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2488053,
              "author_name": "Chris Tang 0002",
              "author_url": "",
              "post_date": "2023-10-19T01:42:16.180000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> , I want to get a drug's emb matrix, but I'm confused about the matrix shape. The smiles' length is 51, but the <strong>model_output[0]</strong> shape is <strong>(53, 600)</strong>. As you said above, there is a cls token and the shape should be <strong>(52,600)</strong>.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2446151,
      "author_name": "𝔄ℌ𝔐𝔈𝔇 𝔄𝔖ℌℜ𝔄𝔉",
      "author_url": "",
      "post_date": "2023-09-19T09:23:22.020000",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> a lot of effort thanks for sharing </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2488882,
      "author_name": "plsdontkillme",
      "author_url": "",
      "post_date": "2023-10-19T15:05:50.313000",
      "content": "<p>How did you train it? cause it doesn't seems like it was trained or did you just loaded the pre-trained model?<br>\ncause if you do it with \"torch.no_grad()\" pytorch won't take that loss into account no?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2447776,
      "author_name": "hoho",
      "author_url": "",
      "post_date": "2023-09-20T08:43:25.607000",
      "content": "<p><a href=\"https://www.kaggle.com/alekseytrepetsky\" target=\"_blank\">@alekseytrepetsky</a> I think your data has a lot of redundance.</p>\n<p>There are only 147 SMILES, which are the same in train and test. But in your dataset, there are 614 for train and 255 for test.<br>\nYou are duplicate the data (614+255)/147 ≈ 6 times!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2449983,
          "author_name": "Aleksey Trepetsky",
          "author_url": "",
          "post_date": "2023-09-21T14:55:11.253000",
          "content": "<p>Thank you for your observation. Yes, there is indeed duplication of data, and it was done intentionally. The embeddings were created in the same sequence as the chemical compounds in the train and test sets. This way, you can simply replace the sm_name in the datasets with the embeddings and use them. I thought it would be more convenient this way. Also, we have a small amount of data, so duplicating it is not critical.<br>\nHowever, if you find it more convenient to work with non-duplicated embeddings, please let me know, and I can add data without duplication to the dataset.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2451071,
              "author_name": "hoho",
              "author_url": "",
              "post_date": "2023-09-22T09:45:12.860000",
              "content": "<p>ok, that make sense.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2485877,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-17T14:08:58.713000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2462376,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-30T09:57:21.567000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2447508,
      "author_name": "BItoAI",
      "author_url": "",
      "post_date": "2023-09-20T05:57:50.707000",
      "content": "<p>Great jobs and thanks for sharing !!!</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2446075": "Hello everyone! I have generated embeddings for smiles using ChemBERTa v2.\nIt was trained on 77M smiles from the PubChem database and fine-tuned on BACE, Clearance, Delaney, Lipophilicity, BBBP, ClinTox, HIV, Delaney, and Tox21 datasets.\n\n[arxiv](https://arxiv.org/abs/2209.01712)\n[Dataset](https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr)\n[Notebook](https://www.kaggle.com/code/alekseytrepetsky/create-chemberta-embed)",
    "2464126": "@alekseytrepetsky, thanks for sharing. It was very helpful. I have a suggestion that it would be better not to use the lm_head for extracting features from SMILES, as I believe the pretrained checkpoint doesn't include the corresponding weights. Indeed, I received a warning message indicating the lm_head was randomly initialized when I loaded the pretrained model. In practice, you can skip the lm_head by doing the following;\n`chemberta._modules[\"lm_head\"] = nn.Identity()`\nThen we can expect 384-dim features instead. I appreciate that you find this helpful.",
    "2509138": "@alekseytrepetsky , Thanks for your great job. But I'm confused when I apply this code on my drug data.\nFor example, I tried to get this smiles embedding `COC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl`, But the tokenizer incorrectly consider the last 3 `Cl` as `C`. Also I checked the token vocab, there indeed has `Cl` it dosen't make sense.\n\nThe splited seq: \n```\nCOC1=C(C=C2C(=C1)CCN=C2C3=CC(=C(C=C3)Cl)Cl)Cl\n\n['C', 'O', 'C', '1', '=', 'C', '(', 'C', '=', 'C', '2', 'C', '(', '=', 'C', '1', ')', 'C', 'C', 'N', '=', 'C', '2', 'C', '3', '=', 'C', 'C', '(', '=', 'C', '(', 'C', '=', 'C', '3', ')', 'C', ')', 'C', ')', 'C']\n```\n\n\n",
    "2457674": "Thank you very much for sharing! \nCould you please explain the difference between \"pad_true\" and \"pad_false\" and also between \"cls_pad\" and \"mean_pad\" files?",
    "2446151": "@alekseytrepetsky a lot of effort thanks for sharing ",
    "2488882": "How did you train it? cause it doesn't seems like it was trained or did you just loaded the pre-trained model?\ncause if you do it with \"torch.no_grad()\" pytorch won't take that loss into account no?",
    "2447776": "@alekseytrepetsky I think your data has a lot of redundance.\n\nThere are only 147 SMILES, which are the same in train and test. But in your dataset, there are 614 for train and 255 for test.\nYou are duplicate the data (614+255)/147 ≈ 6 times!",
    "2485877": "",
    "2462376": "",
    "2447508": "Great jobs and thanks for sharing !!!"
  }
}