{
  "id": 519136,
  "title": "Ideas thread",
  "url": "/competitions/leash-BELKA/discussion/519136",
  "author_name": "",
  "post_date": "2024-07-09T21:03:21.029196100Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>For me, especially as I was unable to give this sufficient time (and didn't have sufficient compute power),  this part of the competition always sparks my interest and brainstorming \"what if\" ideas. </p>\n<p>What are future directions you would try? What have you read in other solutions that you would want to expand upon? Post you ideas here!</p>",
  "messages": [
    {
      "id": "2914311",
      "postDate": "07/09/2024 21:03:21",
      "content": "<p>For me, especially as I was unable to give this sufficient time (and didn't have sufficient compute power),  this part of the competition always sparks my interest and brainstorming \"what if\" ideas. </p>\n<p>What are future directions you would try? What have you read in other solutions that you would want to expand upon? Post you ideas here!</p>",
      "rawMarkdown": "For me, especially as I was unable to give this sufficient time (and didn't have sufficient compute power),  this part of the competition always sparks my interest and brainstorming \"what if\" ideas. \n\nWhat are future directions you would try? What have you read in other solutions that you would want to expand upon? Post you ideas here!",
      "votes": null
    },
    {
      "id": "2914319",
      "postDate": "07/09/2024 21:13:29",
      "content": "<p>For me, one idea I worked on a bit and always wanted to continue down is 1dCNN and/or insert favorite smiles NN architecture, but with disynthon smiles. That is, take the final smiles, load it in rdkit, then subtract 1 building block. Do that for every building block pair and you get about 600k disynthon smiles in train data. </p>\n<p>Now train your 1dCNN that used to be shape (batch, max_token_length) on input of (batch, 3,  smaller_max_disynthon_token_len). Keep the CNN to process each disynthon smiles the same, then have later layer that concats or pools the layers before whatever additional layers you run on the merged results. </p>\n<p>This concept could work really well for a transformer head as well, that can choose how much attention to give to each part of each building block. For the 1d CNN simple case, if you using a pooling layer, then it successfully makes bb1_2 disynthon eqiivarient with bb1_3 disynthon. </p>",
      "rawMarkdown": "For me, one idea I worked on a bit and always wanted to continue down is 1dCNN and/or insert favorite smiles NN architecture, but with disynthon smiles. That is, take the final smiles, load it in rdkit, then subtract 1 building block. Do that for every building block pair and you get about 600k disynthon smiles in train data. \n\nNow train your 1dCNN that used to be shape (batch, max_token_length) on input of (batch, 3,  smaller_max_disynthon_token_len). Keep the CNN to process each disynthon smiles the same, then have later layer that concats or pools the layers before whatever additional layers you run on the merged results. \n\nThis concept could work really well for a transformer head as well, that can choose how much attention to give to each part of each building block. For the 1d CNN simple case, if you using a pooling layer, then it successfully makes bb1_2 disynthon eqiivarient with bb1_3 disynthon.",
      "votes": null
    },
    {
      "id": "2914332",
      "postDate": "07/09/2024 21:24:53",
      "content": "<p>A new idea, inspired by the pretraining the winning team and some successful teams did:</p>\n<p>Using the concept of NAMs (neural additive models), have independent model architectures that are added at the end. Have \"good\" model architecture(s) and \"bad\" model architecture(s). Good is a pretrained SMILES model that was pretrained to predict chem features and/or ecfp, and maybe MLM as well. Bad is just numbered building blocks model as entire input. Perhaps one NAM that's [a, b, c], and other subset NAMs ([a, b], [a, c], etc). The idea is you train this model on the training data, and monitor on shared validation data. But also monitor on non-shared BBs, but for the non-shared validation data, you remove the building block NAMs and only use the residual predictions of the \"good\" architecture(s)!</p>\n<p>No idea if it would work, but the intent is to further teach the model to ignore the \"easy\" way to fit the training data, and keep track of the residual ability to generalize based on chemical or small subsequence properties</p>",
      "rawMarkdown": "A new idea, inspired by the pretraining the winning team and some successful teams did:\n\nUsing the concept of NAMs (neural additive models), have independent model architectures that are added at the end. Have \"good\" model architecture(s) and \"bad\" model architecture(s). Good is a pretrained SMILES model that was pretrained to predict chem features and/or ecfp, and maybe MLM as well. Bad is just numbered building blocks model as entire input. Perhaps one NAM that's [a, b, c], and other subset NAMs ([a, b], [a, c], etc). The idea is you train this model on the training data, and monitor on shared validation data. But also monitor on non-shared BBs, but for the non-shared validation data, you remove the building block NAMs and only use the residual predictions of the \"good\" architecture(s)!\n\nNo idea if it would work, but the intent is to further teach the model to ignore the \"easy\" way to fit the training data, and keep track of the residual ability to generalize based on chemical or small subsequence properties",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2914319,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/09/2024 21:13:29",
      "content": "<p>For me, one idea I worked on a bit and always wanted to continue down is 1dCNN and/or insert favorite smiles NN architecture, but with disynthon smiles. That is, take the final smiles, load it in rdkit, then subtract 1 building block. Do that for every building block pair and you get about 600k disynthon smiles in train data. </p>\n<p>Now train your 1dCNN that used to be shape (batch, max_token_length) on input of (batch, 3,  smaller_max_disynthon_token_len). Keep the CNN to process each disynthon smiles the same, then have later layer that concats or pools the layers before whatever additional layers you run on the merged results. </p>\n<p>This concept could work really well for a transformer head as well, that can choose how much attention to give to each part of each building block. For the 1d CNN simple case, if you using a pooling layer, then it successfully makes bb1_2 disynthon eqiivarient with bb1_3 disynthon. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914332,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/09/2024 21:24:53",
      "content": "<p>A new idea, inspired by the pretraining the winning team and some successful teams did:</p>\n<p>Using the concept of NAMs (neural additive models), have independent model architectures that are added at the end. Have \"good\" model architecture(s) and \"bad\" model architecture(s). Good is a pretrained SMILES model that was pretrained to predict chem features and/or ecfp, and maybe MLM as well. Bad is just numbered building blocks model as entire input. Perhaps one NAM that's [a, b, c], and other subset NAMs ([a, b], [a, c], etc). The idea is you train this model on the training data, and monitor on shared validation data. But also monitor on non-shared BBs, but for the non-shared validation data, you remove the building block NAMs and only use the residual predictions of the \"good\" architecture(s)!</p>\n<p>No idea if it would work, but the intent is to further teach the model to ignore the \"easy\" way to fit the training data, and keep track of the residual ability to generalize based on chemical or small subsequence properties</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2914311": "For me, especially as I was unable to give this sufficient time (and didn't have sufficient compute power),  this part of the competition always sparks my interest and brainstorming \"what if\" ideas. \n\nWhat are future directions you would try? What have you read in other solutions that you would want to expand upon? Post you ideas here!",
    "2914319": "For me, one idea I worked on a bit and always wanted to continue down is 1dCNN and/or insert favorite smiles NN architecture, but with disynthon smiles. That is, take the final smiles, load it in rdkit, then subtract 1 building block. Do that for every building block pair and you get about 600k disynthon smiles in train data. \n\nNow train your 1dCNN that used to be shape (batch, max_token_length) on input of (batch, 3,  smaller_max_disynthon_token_len). Keep the CNN to process each disynthon smiles the same, then have later layer that concats or pools the layers before whatever additional layers you run on the merged results. \n\nThis concept could work really well for a transformer head as well, that can choose how much attention to give to each part of each building block. For the 1d CNN simple case, if you using a pooling layer, then it successfully makes bb1_2 disynthon eqiivarient with bb1_3 disynthon.",
    "2914332": "A new idea, inspired by the pretraining the winning team and some successful teams did:\n\nUsing the concept of NAMs (neural additive models), have independent model architectures that are added at the end. Have \"good\" model architecture(s) and \"bad\" model architecture(s). Good is a pretrained SMILES model that was pretrained to predict chem features and/or ecfp, and maybe MLM as well. Bad is just numbered building blocks model as entire input. Perhaps one NAM that's [a, b, c], and other subset NAMs ([a, b], [a, c], etc). The idea is you train this model on the training data, and monitor on shared validation data. But also monitor on non-shared BBs, but for the non-shared validation data, you remove the building block NAMs and only use the residual predictions of the \"good\" architecture(s)!\n\nNo idea if it would work, but the intent is to further teach the model to ignore the \"easy\" way to fit the training data, and keep track of the residual ability to generalize based on chemical or small subsequence properties"
  },
  "source": "meta"
}