{
  "id": 491324,
  "title": "Should Models be symmetric under the exchange of buildingblock2_smiles and buildingblock3_smiles?",
  "url": "/competitions/leash-BELKA/discussion/491324",
  "author_name": "",
  "post_date": "2024-04-05T12:40:25.781501500Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>According to the definition of the DEL, it seems that the exchange of buildingblock2_smiles (The structure, in SMILES, of the second building block) and buildingblock3_smiles (that of the third one) does not change the molecule_smiles (that of the fully assembled molecule).</p>\n<p>So, if we include information about buildingblock2_smiles and buildingblock3_smiles in our predictive model, do we need to restrict the model to be invariant with respect to the exchange of buildingblock2_smiles and buildingblock3_smiles? </p>\n<p>The simplest way to implement would be to make the dataset symmetric under the exchange of buildingblock2_smiles and buildingblock3_smiles, i.e. if there is a row with buildingblock2_smiles=A and buildingblock3_smiles=B, add another row with buildingblock2_smiles=B and buildingblock3_smiles=A.</p>",
  "messages": [
    {
      "id": "2736854",
      "postDate": "04/05/2024 12:40:25",
      "content": "<p>According to the definition of the DEL, it seems that the exchange of buildingblock2_smiles (The structure, in SMILES, of the second building block) and buildingblock3_smiles (that of the third one) does not change the molecule_smiles (that of the fully assembled molecule).</p>\n<p>So, if we include information about buildingblock2_smiles and buildingblock3_smiles in our predictive model, do we need to restrict the model to be invariant with respect to the exchange of buildingblock2_smiles and buildingblock3_smiles? </p>\n<p>The simplest way to implement would be to make the dataset symmetric under the exchange of buildingblock2_smiles and buildingblock3_smiles, i.e. if there is a row with buildingblock2_smiles=A and buildingblock3_smiles=B, add another row with buildingblock2_smiles=B and buildingblock3_smiles=A.</p>",
      "rawMarkdown": "According to the definition of the DEL, it seems that the exchange of buildingblock2_smiles (The structure, in SMILES, of the second building block) and buildingblock3_smiles (that of the third one) does not change the molecule_smiles (that of the fully assembled molecule).\n\nSo, if we include information about buildingblock2_smiles and buildingblock3_smiles in our predictive model, do we need to restrict the model to be invariant with respect to the exchange of buildingblock2_smiles and buildingblock3_smiles? \n\nThe simplest way to implement would be to make the dataset symmetric under the exchange of buildingblock2_smiles and buildingblock3_smiles, i.e. if there is a row with buildingblock2_smiles=A and buildingblock3_smiles=B, add another row with buildingblock2_smiles=B and buildingblock3_smiles=A.",
      "votes": null
    },
    {
      "id": "2736983",
      "postDate": "04/05/2024 14:11:44",
      "content": "<p>This seems like an interesting line of inquiry, maybe there are gains by understanding these symmetries.</p>\n<p>but remember, this symmetry is true for this training set, but it's not something you can guarantee for all molecules. </p>",
      "rawMarkdown": "This seems like an interesting line of inquiry, maybe there are gains by understanding these symmetries.\n\nbut remember, this symmetry is true for this training set, but it's not something you can guarantee for all molecules.",
      "votes": null
    },
    {
      "id": "2737038",
      "postDate": "04/05/2024 14:33:33",
      "content": "<p>I would recommend not to include BB data in the model, since the test set is stratified by BB from the train set as per the description: ' (e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets)'.</p>\n<p>Edit: Turn out <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">only part of the test set is stratified</a>. Including BB structures may be a good idea for the non-stratified part.</p>",
      "rawMarkdown": "I would recommend not to include BB data in the model, since the test set is stratified by BB from the train set as per the description: ' (e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets)'.\n\nEdit: Turn out [only part of the test set is stratified](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103). Including BB structures may be a good idea for the non-stratified part.",
      "votes": null
    },
    {
      "id": "2749171",
      "postDate": "04/12/2024 22:00:20",
      "content": "<p>Thanks for your reply. <br>\nI mistakenly thought building blocks are always connected to symmetric triazine. Of course, this discussion about symmetry does not hold true because test dataset contains non triazine molecules.</p>",
      "rawMarkdown": "Thanks for your reply. \nI mistakenly thought building blocks are always connected to symmetric triazine. Of course, this discussion about symmetry does not hold true because test dataset contains non triazine molecules.",
      "votes": null
    },
    {
      "id": "2749176",
      "postDate": "04/12/2024 22:02:16",
      "content": "<p>You're right. BBs should be removed for general prediction of test dataset. </p>",
      "rawMarkdown": "You're right. BBs should be removed for general prediction of test dataset.",
      "votes": null
    },
    {
      "id": "2749260",
      "postDate": "04/13/2024 00:32:23",
      "content": "<p>I think I'm asking the same question when I ask, re the below quote, whether Mickey zipper ears are interchangeable with Mickey velcro ears? In the example below, 10x10x10 = 1000. The question is whether the data we're given treats building_block_B as a \"connects to zipper ear\" building block, and you'll never find a building_block_B that matches any building_block_C. (Because no B block contains the right \"velcro\" to connect to C). Or are A, B, C arbitrary, and they're interchangeable? So if you have 30 Mickey building blocks, it's not 10x10x10, but actually 30x29x28 for all permutations (or 30x29x28/(3x2) for all combinations?</p>\n<blockquote>\n  <p>We could purchase ten different Mickey Mouse faces, ten different zipper ears, and ten different velcro ears, and use them to construct our small molecule library. By creating every combination of these three, we’ll have 1,000 small molecules, but we only needed thirty building blocks (faces and ears) to make them. This combinatorial approach is what allows DELs to have so many members: the library in this competition is composed of 133M small molecules. The 133M small molecule library used here, AMA014, was provided by AlphaMa. It has a triazine core and superficially resembles the DELs described here.</p>\n</blockquote>",
      "rawMarkdown": "I think I'm asking the same question when I ask, re the below quote, whether Mickey zipper ears are interchangeable with Mickey velcro ears? In the example below, 10x10x10 = 1000. The question is whether the data we're given treats building_block_B as a \"connects to zipper ear\" building block, and you'll never find a building_block_B that matches any building_block_C. (Because no B block contains the right \"velcro\" to connect to C). Or are A, B, C arbitrary, and they're interchangeable? So if you have 30 Mickey building blocks, it's not 10x10x10, but actually 30x29x28 for all permutations (or 30x29x28/(3x2) for all combinations?\n\n> We could purchase ten different Mickey Mouse faces, ten different zipper ears, and ten different velcro ears, and use them to construct our small molecule library. By creating every combination of these three, we’ll have 1,000 small molecules, but we only needed thirty building blocks (faces and ears) to make them. This combinatorial approach is what allows DELs to have so many members: the library in this competition is composed of 133M small molecules. The 133M small molecule library used here, AMA014, was provided by AlphaMa. It has a triazine core and superficially resembles the DELs described here.",
      "votes": null
    },
    {
      "id": "2749313",
      "postDate": "04/13/2024 02:03:02",
      "content": "<p>The analogy is a little tricky here. You can make the same bond with different reactions. i.e. a zipper ear and a velcro ear may be able to make the same connection between Mickey's head and the ear. While a zipper and velcro are different, there is nothing about the type of reaction used that is retained in the product of a chemical reaction. That said, the type of reactions that you can use are dependent on the product you want to make. Also, you can make the same molecule many different ways, with different building blocks and/or different reactions. Finally, you can make many different products from the same building blocks. Example, if you have three building blocks A, B, and C, you can make a molecule ABC, CBA, ACB, etc even if you're iteratively running the same kind of reaction, because the order can matter.</p>\n<p>For our train data, things are a bit simple. Because the first building block is the one that has DNA on it, all of those are unique and each one will provide a different product. However, the order with which you add building blocks 2 and 3 is meaningless in this case because triazines have a plane of symmetry and you end up making the same product. </p>\n<p>For our test data, it's more complex. Without significant test set probing, which I'm currently avoiding beyond identifying that there are non-triazines in the test set, it's difficult to say what symmetries the final molecules will have.</p>",
      "rawMarkdown": "The analogy is a little tricky here. You can make the same bond with different reactions. i.e. a zipper ear and a velcro ear may be able to make the same connection between Mickey's head and the ear. While a zipper and velcro are different, there is nothing about the type of reaction used that is retained in the product of a chemical reaction. That said, the type of reactions that you can use are dependent on the product you want to make. Also, you can make the same molecule many different ways, with different building blocks and/or different reactions. Finally, you can make many different products from the same building blocks. Example, if you have three building blocks A, B, and C, you can make a molecule ABC, CBA, ACB, etc even if you're iteratively running the same kind of reaction, because the order can matter.\n\nFor our train data, things are a bit simple. Because the first building block is the one that has DNA on it, all of those are unique and each one will provide a different product. However, the order with which you add building blocks 2 and 3 is meaningless in this case because triazines have a plane of symmetry and you end up making the same product. \n\nFor our test data, it's more complex. Without significant test set probing, which I'm currently avoiding beyond identifying that there are non-triazines in the test set, it's difficult to say what symmetries the final molecules will have.",
      "votes": null
    },
    {
      "id": "2749336",
      "postDate": "04/13/2024 02:34:17",
      "content": "<p>That's super helpful, thanks!</p>\n<p>Even though it's different for (55% of) the test set. I guess I'll think of it as A, B1, B2, and treat B1 and B2 as interchangeable</p>",
      "rawMarkdown": "That's super helpful, thanks!\n\nEven though it's different for (55% of) the test set. I guess I'll think of it as A, B1, B2, and treat B1 and B2 as interchangeable",
      "votes": null
    },
    {
      "id": "2749360",
      "postDate": "04/13/2024 02:46:43",
      "content": "<p>I'd say if you're using the building blocks to make train/test splits, that's perfectly reasonable since we only have data on triazines. However, if you're using that information directly for inference, I'd be a bit more cautious.</p>\n<p>I actually haven't checked whether or not BB2 and BB3 share building blocks yet since I'm not currently using them. I'm (perhaps naively naively) using Murcko decompositions (done in RDKit) for my train/test splits.</p>",
      "rawMarkdown": "I'd say if you're using the building blocks to make train/test splits, that's perfectly reasonable since we only have data on triazines. However, if you're using that information directly for inference, I'd be a bit more cautious.\n\nI actually haven't checked whether or not BB2 and BB3 share building blocks yet since I'm not currently using them. I'm (perhaps naively naively) using Murcko decompositions (done in RDKit) for my train/test splits.",
      "votes": null
    },
    {
      "id": "2749484",
      "postDate": "04/13/2024 04:50:29",
      "content": "<p>I saw that test has BB2 and BB3 overlap. Didn't yet run the same check on train but kinda expect the same result. </p>\n<p>Pretty simple check using just the dictionaries in the shrunk dataset, no need to even load the actual molecules data at all to check this. </p>",
      "rawMarkdown": "I saw that test has BB2 and BB3 overlap. Didn't yet run the same check on train but kinda expect the same result. \n\nPretty simple check using just the dictionaries in the shrunk dataset, no need to even load the actual molecules data at all to check this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2736983,
      "author_name": "andrewdblevins",
      "author_url": "",
      "post_date": "04/05/2024 14:11:44",
      "content": "<p>This seems like an interesting line of inquiry, maybe there are gains by understanding these symmetries.</p>\n<p>but remember, this symmetry is true for this training set, but it's not something you can guarantee for all molecules. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2749171,
          "author_name": "atsuno",
          "author_url": "",
          "post_date": "04/12/2024 22:00:20",
          "content": "<p>Thanks for your reply. <br>\nI mistakenly thought building blocks are always connected to symmetric triazine. Of course, this discussion about symmetry does not hold true because test dataset contains non triazine molecules.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2737038,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "04/05/2024 14:33:33",
      "content": "<p>I would recommend not to include BB data in the model, since the test set is stratified by BB from the train set as per the description: ' (e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets)'.</p>\n<p>Edit: Turn out <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103\" target=\"_blank\">only part of the test set is stratified</a>. Including BB structures may be a good idea for the non-stratified part.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2749176,
          "author_name": "atsuno",
          "author_url": "",
          "post_date": "04/12/2024 22:02:16",
          "content": "<p>You're right. BBs should be removed for general prediction of test dataset. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2749260,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/13/2024 00:32:23",
      "content": "<p>I think I'm asking the same question when I ask, re the below quote, whether Mickey zipper ears are interchangeable with Mickey velcro ears? In the example below, 10x10x10 = 1000. The question is whether the data we're given treats building_block_B as a \"connects to zipper ear\" building block, and you'll never find a building_block_B that matches any building_block_C. (Because no B block contains the right \"velcro\" to connect to C). Or are A, B, C arbitrary, and they're interchangeable? So if you have 30 Mickey building blocks, it's not 10x10x10, but actually 30x29x28 for all permutations (or 30x29x28/(3x2) for all combinations?</p>\n<blockquote>\n  <p>We could purchase ten different Mickey Mouse faces, ten different zipper ears, and ten different velcro ears, and use them to construct our small molecule library. By creating every combination of these three, we’ll have 1,000 small molecules, but we only needed thirty building blocks (faces and ears) to make them. This combinatorial approach is what allows DELs to have so many members: the library in this competition is composed of 133M small molecules. The 133M small molecule library used here, AMA014, was provided by AlphaMa. It has a triazine core and superficially resembles the DELs described here.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2749313,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "04/13/2024 02:03:02",
          "content": "<p>The analogy is a little tricky here. You can make the same bond with different reactions. i.e. a zipper ear and a velcro ear may be able to make the same connection between Mickey's head and the ear. While a zipper and velcro are different, there is nothing about the type of reaction used that is retained in the product of a chemical reaction. That said, the type of reactions that you can use are dependent on the product you want to make. Also, you can make the same molecule many different ways, with different building blocks and/or different reactions. Finally, you can make many different products from the same building blocks. Example, if you have three building blocks A, B, and C, you can make a molecule ABC, CBA, ACB, etc even if you're iteratively running the same kind of reaction, because the order can matter.</p>\n<p>For our train data, things are a bit simple. Because the first building block is the one that has DNA on it, all of those are unique and each one will provide a different product. However, the order with which you add building blocks 2 and 3 is meaningless in this case because triazines have a plane of symmetry and you end up making the same product. </p>\n<p>For our test data, it's more complex. Without significant test set probing, which I'm currently avoiding beyond identifying that there are non-triazines in the test set, it's difficult to say what symmetries the final molecules will have.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2749336,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "04/13/2024 02:34:17",
              "content": "<p>That's super helpful, thanks!</p>\n<p>Even though it's different for (55% of) the test set. I guess I'll think of it as A, B1, B2, and treat B1 and B2 as interchangeable</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2749360,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "04/13/2024 02:46:43",
                  "content": "<p>I'd say if you're using the building blocks to make train/test splits, that's perfectly reasonable since we only have data on triazines. However, if you're using that information directly for inference, I'd be a bit more cautious.</p>\n<p>I actually haven't checked whether or not BB2 and BB3 share building blocks yet since I'm not currently using them. I'm (perhaps naively naively) using Murcko decompositions (done in RDKit) for my train/test splits.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2749484,
                      "author_name": "roberthatch",
                      "author_url": "",
                      "post_date": "04/13/2024 04:50:29",
                      "content": "<p>I saw that test has BB2 and BB3 overlap. Didn't yet run the same check on train but kinda expect the same result. </p>\n<p>Pretty simple check using just the dictionaries in the shrunk dataset, no need to even load the actual molecules data at all to check this. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2736854": "According to the definition of the DEL, it seems that the exchange of buildingblock2_smiles (The structure, in SMILES, of the second building block) and buildingblock3_smiles (that of the third one) does not change the molecule_smiles (that of the fully assembled molecule).\n\nSo, if we include information about buildingblock2_smiles and buildingblock3_smiles in our predictive model, do we need to restrict the model to be invariant with respect to the exchange of buildingblock2_smiles and buildingblock3_smiles? \n\nThe simplest way to implement would be to make the dataset symmetric under the exchange of buildingblock2_smiles and buildingblock3_smiles, i.e. if there is a row with buildingblock2_smiles=A and buildingblock3_smiles=B, add another row with buildingblock2_smiles=B and buildingblock3_smiles=A.",
    "2736983": "This seems like an interesting line of inquiry, maybe there are gains by understanding these symmetries.\n\nbut remember, this symmetry is true for this training set, but it's not something you can guarantee for all molecules.",
    "2737038": "I would recommend not to include BB data in the model, since the test set is stratified by BB from the train set as per the description: ' (e.g., for a given building block in the test set, all molecules containing that building block must be removed from the training and validation sets)'.\n\nEdit: Turn out [only part of the test set is stratified](https://www.kaggle.com/competitions/leash-BELKA/discussion/491362#2737103). Including BB structures may be a good idea for the non-stratified part.",
    "2749171": "Thanks for your reply. \nI mistakenly thought building blocks are always connected to symmetric triazine. Of course, this discussion about symmetry does not hold true because test dataset contains non triazine molecules.",
    "2749176": "You're right. BBs should be removed for general prediction of test dataset.",
    "2749260": "I think I'm asking the same question when I ask, re the below quote, whether Mickey zipper ears are interchangeable with Mickey velcro ears? In the example below, 10x10x10 = 1000. The question is whether the data we're given treats building_block_B as a \"connects to zipper ear\" building block, and you'll never find a building_block_B that matches any building_block_C. (Because no B block contains the right \"velcro\" to connect to C). Or are A, B, C arbitrary, and they're interchangeable? So if you have 30 Mickey building blocks, it's not 10x10x10, but actually 30x29x28 for all permutations (or 30x29x28/(3x2) for all combinations?\n\n> We could purchase ten different Mickey Mouse faces, ten different zipper ears, and ten different velcro ears, and use them to construct our small molecule library. By creating every combination of these three, we’ll have 1,000 small molecules, but we only needed thirty building blocks (faces and ears) to make them. This combinatorial approach is what allows DELs to have so many members: the library in this competition is composed of 133M small molecules. The 133M small molecule library used here, AMA014, was provided by AlphaMa. It has a triazine core and superficially resembles the DELs described here.",
    "2749313": "The analogy is a little tricky here. You can make the same bond with different reactions. i.e. a zipper ear and a velcro ear may be able to make the same connection between Mickey's head and the ear. While a zipper and velcro are different, there is nothing about the type of reaction used that is retained in the product of a chemical reaction. That said, the type of reactions that you can use are dependent on the product you want to make. Also, you can make the same molecule many different ways, with different building blocks and/or different reactions. Finally, you can make many different products from the same building blocks. Example, if you have three building blocks A, B, and C, you can make a molecule ABC, CBA, ACB, etc even if you're iteratively running the same kind of reaction, because the order can matter.\n\nFor our train data, things are a bit simple. Because the first building block is the one that has DNA on it, all of those are unique and each one will provide a different product. However, the order with which you add building blocks 2 and 3 is meaningless in this case because triazines have a plane of symmetry and you end up making the same product. \n\nFor our test data, it's more complex. Without significant test set probing, which I'm currently avoiding beyond identifying that there are non-triazines in the test set, it's difficult to say what symmetries the final molecules will have.",
    "2749336": "That's super helpful, thanks!\n\nEven though it's different for (55% of) the test set. I guess I'll think of it as A, B1, B2, and treat B1 and B2 as interchangeable",
    "2749360": "I'd say if you're using the building blocks to make train/test splits, that's perfectly reasonable since we only have data on triazines. However, if you're using that information directly for inference, I'd be a bit more cautious.\n\nI actually haven't checked whether or not BB2 and BB3 share building blocks yet since I'm not currently using them. I'm (perhaps naively naively) using Murcko decompositions (done in RDKit) for my train/test splits.",
    "2749484": "I saw that test has BB2 and BB3 overlap. Didn't yet run the same check on train but kinda expect the same result. \n\nPretty simple check using just the dictionaries in the shrunk dataset, no need to even load the actual molecules data at all to check this."
  },
  "source": "meta"
}