{
  "id": 518938,
  "title": "Public 44th Private 898th Solution",
  "url": "/competitions/leash-BELKA/discussion/518938",
  "author_name": "2g",
  "post_date": "2024-07-09T00:12:11.785000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Overview</h2>\n<ul>\n<li>Main: Create features for each building block and pass them through FC layers.</li>\n<li>Occasionally: use embedding features of Tokenized SMILES.</li>\n</ul>\n<p>models created by various conditions were used for ensemble.</p>\n<h3><strong>Data</strong></h3>\n<ul>\n<li>Molecules that bind to any protein: <strong>use all</strong></li>\n<li>Molecules that do not bind to any protein: <strong>randomly sample 1/5 of them</strong><br>\n(Using all records did not improve the score, so only part of them was used.<br>\nUsed different random states for sampling as ensemble models)</li>\n</ul>\n<h3>Preprocessing</h3>\n<ul>\n<li><p>Cap the binding sites of bb1, bb2, bb3 with specific structures (e.g., fluorene)<br>\nBecause…</p>\n<ul>\n<li>To distinguish binding sites</li>\n<li>To align binding sites of new library bb with those in train data</li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2F1e411c70cc6cd68b6c66d19247296463%2Fpreprocess.png?generation=1720483827105452&amp;alt=media\"> </p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fda8391dcfe0f3ca19834d42b204a426f%2FUntitled.png?generation=1720483916092332&amp;alt=media\"></p>\n<h3><strong>Features</strong></h3>\n<ul>\n<li>RDkit descriptors: Represent chemical properties</li>\n<li>ECFP4 (Fingerprint, radius=2, nBit=1024): Represent the presence of substructures with 0 or 1</li>\n<li>Others (used instead of or concatenated with the above two)<ul>\n<li>MACCSkey, AvalonFP, original descriptors, ECFP6, mordred descriptor</li></ul></li>\n<li>Used scaffold features of bb1 (bb1 scaffold)</li>\n</ul>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fd4cdb2bf6a7253c9f733288294979074%2Farchitechture.png?generation=1720483787166354&amp;alt=media\"></p>\n<h3><strong>Augmentation</strong></h3>\n<ul>\n<li>Swap bb2 and bb3</li>\n</ul>\n<h3><strong>Loss</strong></h3>\n<ul>\n<li>nn.BCEWithLogitsLoss()</li>\n</ul>\n<h3><strong>CV Strategy</strong></h3>\n<ol>\n<li>Extract Murcko scaffolds of bb1 and calculate ECFP4</li>\n<li>Based on ECFP4 values, molecules were divided into 5 groups by k-NN.</li>\n</ol>\n<h3><strong>Post-process</strong></h3>\n<ul>\n<li>In train data, set the prediction value to 0 if no combination of building blocks binds to any protein.</li>\n</ul>\n<h3><strong>Ensemble</strong></h3>\n<ul>\n<li>Searched for the combination with the highest CV using oof (using optuna)</li>\n<li>Did not fine-tune weights; optimized whether to use each submodel (0 or 1)</li>\n</ul>",
  "messages": [
    {
      "id": 2912500,
      "postDate": "2024-07-09T00:12:11.787Z",
      "content": "<h2>Overview</h2>\n<ul>\n<li>Main: Create features for each building block and pass them through FC layers.</li>\n<li>Occasionally: use embedding features of Tokenized SMILES.</li>\n</ul>\n<p>models created by various conditions were used for ensemble.</p>\n<h3><strong>Data</strong></h3>\n<ul>\n<li>Molecules that bind to any protein: <strong>use all</strong></li>\n<li>Molecules that do not bind to any protein: <strong>randomly sample 1/5 of them</strong><br>\n(Using all records did not improve the score, so only part of them was used.<br>\nUsed different random states for sampling as ensemble models)</li>\n</ul>\n<h3>Preprocessing</h3>\n<ul>\n<li><p>Cap the binding sites of bb1, bb2, bb3 with specific structures (e.g., fluorene)<br>\nBecause…</p>\n<ul>\n<li>To distinguish binding sites</li>\n<li>To align binding sites of new library bb with those in train data</li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2F1e411c70cc6cd68b6c66d19247296463%2Fpreprocess.png?generation=1720483827105452&amp;alt=media\"> </p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fda8391dcfe0f3ca19834d42b204a426f%2FUntitled.png?generation=1720483916092332&amp;alt=media\"></p>\n<h3><strong>Features</strong></h3>\n<ul>\n<li>RDkit descriptors: Represent chemical properties</li>\n<li>ECFP4 (Fingerprint, radius=2, nBit=1024): Represent the presence of substructures with 0 or 1</li>\n<li>Others (used instead of or concatenated with the above two)<ul>\n<li>MACCSkey, AvalonFP, original descriptors, ECFP6, mordred descriptor</li></ul></li>\n<li>Used scaffold features of bb1 (bb1 scaffold)</li>\n</ul>\n<h3>Architecture</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fd4cdb2bf6a7253c9f733288294979074%2Farchitechture.png?generation=1720483787166354&amp;alt=media\"></p>\n<h3><strong>Augmentation</strong></h3>\n<ul>\n<li>Swap bb2 and bb3</li>\n</ul>\n<h3><strong>Loss</strong></h3>\n<ul>\n<li>nn.BCEWithLogitsLoss()</li>\n</ul>\n<h3><strong>CV Strategy</strong></h3>\n<ol>\n<li>Extract Murcko scaffolds of bb1 and calculate ECFP4</li>\n<li>Based on ECFP4 values, molecules were divided into 5 groups by k-NN.</li>\n</ol>\n<h3><strong>Post-process</strong></h3>\n<ul>\n<li>In train data, set the prediction value to 0 if no combination of building blocks binds to any protein.</li>\n</ul>\n<h3><strong>Ensemble</strong></h3>\n<ul>\n<li>Searched for the combination with the highest CV using oof (using optuna)</li>\n<li>Did not fine-tune weights; optimized whether to use each submodel (0 or 1)</li>\n</ul>",
      "rawMarkdown": "## Overview\n\n- Main: Create features for each building block and pass them through FC layers.\n- Occasionally: use embedding features of Tokenized SMILES.\n\nmodels created by various conditions were used for ensemble.\n\n### **Data**\n\n- Molecules that bind to any protein: **use all**\n- Molecules that do not bind to any protein: **randomly sample 1/5 of them**\n(Using all records did not improve the score, so only part of them was used.\nUsed different random states for sampling as ensemble models)\n\n### Preprocessing\n\n- Cap the binding sites of bb1, bb2, bb3 with specific structures (e.g., fluorene)\nBecause…\n    - To distinguish binding sites\n    - To align binding sites of new library bb with those in train data\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2F1e411c70cc6cd68b6c66d19247296463%2Fpreprocess.png?generation=1720483827105452&alt=media) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fda8391dcfe0f3ca19834d42b204a426f%2FUntitled.png?generation=1720483916092332&alt=media)\n\n\n\n### **Features**\n\n- RDkit descriptors: Represent chemical properties\n- ECFP4 (Fingerprint, radius=2, nBit=1024): Represent the presence of substructures with 0 or 1\n- Others (used instead of or concatenated with the above two)\n    - MACCSkey, AvalonFP, original descriptors, ECFP6, mordred descriptor\n- Used scaffold features of bb1 (bb1 scaffold)\n\n### Architecture\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fd4cdb2bf6a7253c9f733288294979074%2Farchitechture.png?generation=1720483787166354&alt=media)\n\n### **Augmentation**\n\n- Swap bb2 and bb3\n\n### **Loss**\n\n- nn.BCEWithLogitsLoss()\n\n### **CV Strategy**\n\n1. Extract Murcko scaffolds of bb1 and calculate ECFP4\n2. Based on ECFP4 values, molecules were divided into 5 groups by k-NN.\n\n### **Post-process**\n\n- In train data, set the prediction value to 0 if no combination of building blocks binds to any protein.\n\n### **Ensemble**\n\n- Searched for the combination with the highest CV using oof (using optuna)\n- Did not fine-tune weights; optimized whether to use each submodel (0 or 1)",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2912500": "## Overview\n\n- Main: Create features for each building block and pass them through FC layers.\n- Occasionally: use embedding features of Tokenized SMILES.\n\nmodels created by various conditions were used for ensemble.\n\n### **Data**\n\n- Molecules that bind to any protein: **use all**\n- Molecules that do not bind to any protein: **randomly sample 1/5 of them**\n(Using all records did not improve the score, so only part of them was used.\nUsed different random states for sampling as ensemble models)\n\n### Preprocessing\n\n- Cap the binding sites of bb1, bb2, bb3 with specific structures (e.g., fluorene)\nBecause…\n    - To distinguish binding sites\n    - To align binding sites of new library bb with those in train data\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2F1e411c70cc6cd68b6c66d19247296463%2Fpreprocess.png?generation=1720483827105452&alt=media) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fda8391dcfe0f3ca19834d42b204a426f%2FUntitled.png?generation=1720483916092332&alt=media)\n\n\n\n### **Features**\n\n- RDkit descriptors: Represent chemical properties\n- ECFP4 (Fingerprint, radius=2, nBit=1024): Represent the presence of substructures with 0 or 1\n- Others (used instead of or concatenated with the above two)\n    - MACCSkey, AvalonFP, original descriptors, ECFP6, mordred descriptor\n- Used scaffold features of bb1 (bb1 scaffold)\n\n### Architecture\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2913095%2Fd4cdb2bf6a7253c9f733288294979074%2Farchitechture.png?generation=1720483787166354&alt=media)\n\n### **Augmentation**\n\n- Swap bb2 and bb3\n\n### **Loss**\n\n- nn.BCEWithLogitsLoss()\n\n### **CV Strategy**\n\n1. Extract Murcko scaffolds of bb1 and calculate ECFP4\n2. Based on ECFP4 values, molecules were divided into 5 groups by k-NN.\n\n### **Post-process**\n\n- In train data, set the prediction value to 0 if no combination of building blocks binds to any protein.\n\n### **Ensemble**\n\n- Searched for the combination with the highest CV using oof (using optuna)\n- Did not fine-tune weights; optimized whether to use each submodel (0 or 1)"
  }
}