{
  "id": 499870,
  "title": "molecule_smiles' - The structure of the fully assembled molecule",
  "url": "/competitions/leash-BELKA/discussion/499870",
  "author_name": "",
  "post_date": "2024-05-03T10:53:22.499548500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As described in the Data tab of this competition, “molecule_smiles' represent structure of the fully assembled molecule, in SMILES. This includes the three building blocks” . Is it true that this   feature is sufficient to consider for predicting the target variable or we should focus all features. It means, we can ignore other three columns (i.e. buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles) given in the dataset</p>",
  "messages": [
    {
      "id": "2790816",
      "postDate": "05/03/2024 10:53:22",
      "content": "<p>As described in the Data tab of this competition, “molecule_smiles' represent structure of the fully assembled molecule, in SMILES. This includes the three building blocks” . Is it true that this   feature is sufficient to consider for predicting the target variable or we should focus all features. It means, we can ignore other three columns (i.e. buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles) given in the dataset</p>",
      "rawMarkdown": "As described in the Data tab of this competition, “molecule_smiles' represent structure of the fully assembled molecule, in SMILES. This includes the three building blocks” . Is it true that this   feature is sufficient to consider for predicting the target variable or we should focus all features. It means, we can ignore other three columns (i.e. buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles) given in the dataset",
      "votes": null
    },
    {
      "id": "2790831",
      "postDate": "05/03/2024 10:58:05",
      "content": "<p>At the end of the day, you will need experiments to answer that question.</p>\n<p>In theory, you should only need the SMILES of the final molecule because in a perfect world, every reaction will go to 100% completion and the only compound available to bind will be the final one.</p>\n<p>In practice, using building block information <em>may</em> allow the model to learn when certain building blocks did not react well and did not make the final product. However, you will need to experiment with this to determine the answer.</p>\n<p>As a reference, none of my models currently account for building blocks and only use the final molecule.</p>",
      "rawMarkdown": "At the end of the day, you will need experiments to answer that question.\n\nIn theory, you should only need the SMILES of the final molecule because in a perfect world, every reaction will go to 100% completion and the only compound available to bind will be the final one.\n\nIn practice, using building block information *may* allow the model to learn when certain building blocks did not react well and did not make the final product. However, you will need to experiment with this to determine the answer.\n\nAs a reference, none of my models currently account for building blocks and only use the final molecule.",
      "votes": null
    },
    {
      "id": "2791193",
      "postDate": "05/03/2024 14:19:00",
      "content": "<p>i have some thoughts:<br>\nbuild a building block classifier:</p>\n<p>model(test bb+augment) = predict test block id</p>\n<p>then we have feature space of test bb (take the latent vector before classifier)<br>\ni wonder if we now project the train bb onto the test bb feature space, what do we get.</p>",
      "rawMarkdown": "i have some thoughts:\nbuild a building block classifier:\n\nmodel(test bb+augment) = predict test block id\n\nthen we have feature space of test bb (take the latent vector before classifier)\ni wonder if we now project the train bb onto the test bb feature space, what do we get.",
      "votes": null
    },
    {
      "id": "2791259",
      "postDate": "05/03/2024 15:08:11",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> </p>",
      "rawMarkdown": "Thank you very much @chemdatafarmer",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2790831,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "05/03/2024 10:58:05",
      "content": "<p>At the end of the day, you will need experiments to answer that question.</p>\n<p>In theory, you should only need the SMILES of the final molecule because in a perfect world, every reaction will go to 100% completion and the only compound available to bind will be the final one.</p>\n<p>In practice, using building block information <em>may</em> allow the model to learn when certain building blocks did not react well and did not make the final product. However, you will need to experiment with this to determine the answer.</p>\n<p>As a reference, none of my models currently account for building blocks and only use the final molecule.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2791193,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/03/2024 14:19:00",
          "content": "<p>i have some thoughts:<br>\nbuild a building block classifier:</p>\n<p>model(test bb+augment) = predict test block id</p>\n<p>then we have feature space of test bb (take the latent vector before classifier)<br>\ni wonder if we now project the train bb onto the test bb feature space, what do we get.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2791259,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "05/03/2024 15:08:11",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2790816": "As described in the Data tab of this competition, “molecule_smiles' represent structure of the fully assembled molecule, in SMILES. This includes the three building blocks” . Is it true that this   feature is sufficient to consider for predicting the target variable or we should focus all features. It means, we can ignore other three columns (i.e. buildingblock1_smiles, buildingblock2_smiles, buildingblock3_smiles) given in the dataset",
    "2790831": "At the end of the day, you will need experiments to answer that question.\n\nIn theory, you should only need the SMILES of the final molecule because in a perfect world, every reaction will go to 100% completion and the only compound available to bind will be the final one.\n\nIn practice, using building block information *may* allow the model to learn when certain building blocks did not react well and did not make the final product. However, you will need to experiment with this to determine the answer.\n\nAs a reference, none of my models currently account for building blocks and only use the final molecule.",
    "2791193": "i have some thoughts:\nbuild a building block classifier:\n\nmodel(test bb+augment) = predict test block id\n\nthen we have feature space of test bb (take the latent vector before classifier)\ni wonder if we now project the train bb onto the test bb feature space, what do we get.",
    "2791259": "Thank you very much @chemdatafarmer"
  },
  "source": "meta"
}