{
  "id": 519152,
  "title": "Took a gamble.  Didn't pay off.  1122th place.  😕",
  "url": "/competitions/leash-BELKA/discussion/519152",
  "author_name": "KirkDCO",
  "post_date": "2024-07-09T22:34:53.143000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congratulations to everyone and especially the medalists.  This was a very interesting competition, but I'm sadly disappointed in my performance.</p>\n<p>I took a very different approach than I've seen in the write-ups thus far.  Specifically I see many descriptions of 1D-CNNs trained on tokenized SMILES and use of Morgan (ECFP) fingerprints or RDKit (path-based) fingerprints. </p>\n<p>I was concerned that using the SMILES representations and either of these fingerprints would lead to an inability to generalize to the non-shared and non-triazine test sets.  Instead, I used 2D-pharmacophore fingerprints in hopes of having a representation that could transfer outside the training set, and typically saw good CV performance (see CV approach, below).  I was hopeful I would benefit from the shakeup, alas, it wasn't to be. </p>\n<p>I'm still very curious about performances across the 9 subsets - shared, non-shared, and non-triazine across the three protein targets.  I'm concerned that these two fingerprints and use of the training set SMILES directly are too specific, and a more abstract representation is needed for good generalization.  Did these methods just memorize the training set BBs and perform well on non-shared BBs that were structurally very similar, or did they actually move into new chemical space?</p>\n<p>My approach: </p>\n<p>Each protein was modeled independently. </p>\n<p>CV - I did random CV and also CV by building block, stratified for both.  For CV by Bb, I split the BBs such that the validation set contained no BBs present on the training set.  CV performance wasn't as good as random but it wasn't bad, suggesting generalization to new BBs outside those used for training is possible.</p>\n<p>BB fingerprints - for the shared BBs, I developed an 1145-bit fingerprint that had bits set for the BBs present in a molecule.  This boosted my performance when used for shared BB predictions, but was obviously worthless for non-shared.  I included predictions using this fingerprint only for the shared BBs and saw improvement in both public and private LB.</p>\n<p>BB only representations - to address the non-triazine set, I created 2D-pharmacphore fingerprints for each BB and not the whole molecule.  Looking at protein models with bound ligands, it looks like only one BB at a time will fit into a binding pocket.  I thought representing each building block, independently would allow for single or multiple BBs in a given binding site.</p>\n<p>I also tried a representation with an 1145-element vector consisting on the similarity of the BB to the 1145 training set BBs.  These were concatenated into a 3435 element vector (1145 for each BB) to represent the whole molecule.</p>\n<p>XGBoost - Since I focused on fingerprints, I used XGB rather than any NN model.  </p>",
  "messages": [
    {
      "id": 2914396,
      "postDate": "2024-07-09T22:34:53.143Z",
      "content": "<p>Congratulations to everyone and especially the medalists.  This was a very interesting competition, but I'm sadly disappointed in my performance.</p>\n<p>I took a very different approach than I've seen in the write-ups thus far.  Specifically I see many descriptions of 1D-CNNs trained on tokenized SMILES and use of Morgan (ECFP) fingerprints or RDKit (path-based) fingerprints. </p>\n<p>I was concerned that using the SMILES representations and either of these fingerprints would lead to an inability to generalize to the non-shared and non-triazine test sets.  Instead, I used 2D-pharmacophore fingerprints in hopes of having a representation that could transfer outside the training set, and typically saw good CV performance (see CV approach, below).  I was hopeful I would benefit from the shakeup, alas, it wasn't to be. </p>\n<p>I'm still very curious about performances across the 9 subsets - shared, non-shared, and non-triazine across the three protein targets.  I'm concerned that these two fingerprints and use of the training set SMILES directly are too specific, and a more abstract representation is needed for good generalization.  Did these methods just memorize the training set BBs and perform well on non-shared BBs that were structurally very similar, or did they actually move into new chemical space?</p>\n<p>My approach: </p>\n<p>Each protein was modeled independently. </p>\n<p>CV - I did random CV and also CV by building block, stratified for both.  For CV by Bb, I split the BBs such that the validation set contained no BBs present on the training set.  CV performance wasn't as good as random but it wasn't bad, suggesting generalization to new BBs outside those used for training is possible.</p>\n<p>BB fingerprints - for the shared BBs, I developed an 1145-bit fingerprint that had bits set for the BBs present in a molecule.  This boosted my performance when used for shared BB predictions, but was obviously worthless for non-shared.  I included predictions using this fingerprint only for the shared BBs and saw improvement in both public and private LB.</p>\n<p>BB only representations - to address the non-triazine set, I created 2D-pharmacphore fingerprints for each BB and not the whole molecule.  Looking at protein models with bound ligands, it looks like only one BB at a time will fit into a binding pocket.  I thought representing each building block, independently would allow for single or multiple BBs in a given binding site.</p>\n<p>I also tried a representation with an 1145-element vector consisting on the similarity of the BB to the 1145 training set BBs.  These were concatenated into a 3435 element vector (1145 for each BB) to represent the whole molecule.</p>\n<p>XGBoost - Since I focused on fingerprints, I used XGB rather than any NN model.  </p>",
      "rawMarkdown": "Congratulations to everyone and especially the medalists.  This was a very interesting competition, but I'm sadly disappointed in my performance.\n\nI took a very different approach than I've seen in the write-ups thus far.  Specifically I see many descriptions of 1D-CNNs trained on tokenized SMILES and use of Morgan (ECFP) fingerprints or RDKit (path-based) fingerprints. \n\nI was concerned that using the SMILES representations and either of these fingerprints would lead to an inability to generalize to the non-shared and non-triazine test sets.  Instead, I used 2D-pharmacophore fingerprints in hopes of having a representation that could transfer outside the training set, and typically saw good CV performance (see CV approach, below).  I was hopeful I would benefit from the shakeup, alas, it wasn't to be. \n\nI'm still very curious about performances across the 9 subsets - shared, non-shared, and non-triazine across the three protein targets.  I'm concerned that these two fingerprints and use of the training set SMILES directly are too specific, and a more abstract representation is needed for good generalization.  Did these methods just memorize the training set BBs and perform well on non-shared BBs that were structurally very similar, or did they actually move into new chemical space?\n\nMy approach: \n\nEach protein was modeled independently. \n\nCV - I did random CV and also CV by building block, stratified for both.  For CV by Bb, I split the BBs such that the validation set contained no BBs present on the training set.  CV performance wasn't as good as random but it wasn't bad, suggesting generalization to new BBs outside those used for training is possible.\n\nBB fingerprints - for the shared BBs, I developed an 1145-bit fingerprint that had bits set for the BBs present in a molecule.  This boosted my performance when used for shared BB predictions, but was obviously worthless for non-shared.  I included predictions using this fingerprint only for the shared BBs and saw improvement in both public and private LB.\n\nBB only representations - to address the non-triazine set, I created 2D-pharmacphore fingerprints for each BB and not the whole molecule.  Looking at protein models with bound ligands, it looks like only one BB at a time will fit into a binding pocket.  I thought representing each building block, independently would allow for single or multiple BBs in a given binding site.\n\nI also tried a representation with an 1145-element vector consisting on the similarity of the BB to the 1145 training set BBs.  These were concatenated into a 3435 element vector (1145 for each BB) to represent the whole molecule.\n\nXGBoost - Since I focused on fingerprints, I used XGB rather than any NN model.  ",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2914396": "Congratulations to everyone and especially the medalists.  This was a very interesting competition, but I'm sadly disappointed in my performance.\n\nI took a very different approach than I've seen in the write-ups thus far.  Specifically I see many descriptions of 1D-CNNs trained on tokenized SMILES and use of Morgan (ECFP) fingerprints or RDKit (path-based) fingerprints. \n\nI was concerned that using the SMILES representations and either of these fingerprints would lead to an inability to generalize to the non-shared and non-triazine test sets.  Instead, I used 2D-pharmacophore fingerprints in hopes of having a representation that could transfer outside the training set, and typically saw good CV performance (see CV approach, below).  I was hopeful I would benefit from the shakeup, alas, it wasn't to be. \n\nI'm still very curious about performances across the 9 subsets - shared, non-shared, and non-triazine across the three protein targets.  I'm concerned that these two fingerprints and use of the training set SMILES directly are too specific, and a more abstract representation is needed for good generalization.  Did these methods just memorize the training set BBs and perform well on non-shared BBs that were structurally very similar, or did they actually move into new chemical space?\n\nMy approach: \n\nEach protein was modeled independently. \n\nCV - I did random CV and also CV by building block, stratified for both.  For CV by Bb, I split the BBs such that the validation set contained no BBs present on the training set.  CV performance wasn't as good as random but it wasn't bad, suggesting generalization to new BBs outside those used for training is possible.\n\nBB fingerprints - for the shared BBs, I developed an 1145-bit fingerprint that had bits set for the BBs present in a molecule.  This boosted my performance when used for shared BB predictions, but was obviously worthless for non-shared.  I included predictions using this fingerprint only for the shared BBs and saw improvement in both public and private LB.\n\nBB only representations - to address the non-triazine set, I created 2D-pharmacphore fingerprints for each BB and not the whole molecule.  Looking at protein models with bound ligands, it looks like only one BB at a time will fit into a binding pocket.  I thought representing each building block, independently would allow for single or multiple BBs in a given binding site.\n\nI also tried a representation with an 1145-element vector consisting on the similarity of the BB to the 1145 training set BBs.  These were concatenated into a 3435 element vector (1145 for each BB) to represent the whole molecule.\n\nXGBoost - Since I focused on fingerprints, I used XGB rather than any NN model.  "
  }
}