{
  "id": 506730,
  "title": "Faster and more stable training solution",
  "url": "/competitions/leash-BELKA/discussion/506730",
  "author_name": "",
  "post_date": "2024-05-23T03:13:08.886092800Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<h1>Advancing Molecular Representation: Techniques and Strategies for Effective Multimodal Learning</h1>\n<p>I recently started learning about molecular representation and hope my ideas can help you.</p>\n<h2>Introduction to Molecular Representation</h2>\n<p>Molecular representation is a critical aspect of computational chemistry and drug discovery, encompassing the transformation of molecular structures into formats that can be processed by machine learning models. These representations can be classified into three primary categories: 1D, 2D, and 3D.</p>\n<h3>Categories of Molecular Representations</h3>\n<p><strong>1D Representations:</strong><br>\n1D representations include molecular descriptors and SMILES strings:</p>\n<ul>\n<li><strong>Molecular Descriptors:</strong> These are numerical values that describe various molecular properties, including fingerprints (such as ECFP, MACCS), and other statistical features. They are useful for capturing detailed molecular characteristics.</li>\n<li><strong>String Descriptions:</strong> SMILES strings are textual representations of molecules, ideal for sequence models and spatial models in NLP.</li>\n</ul>\n<p><strong>2D Representations:</strong><br>\n2D representations involve the topological structure of molecules:</p>\n<ul>\n<li><strong>Topological Graph Structures:</strong> These structures capture the connectivity between atoms in a molecule, providing essential information for graph-based neural networks.</li>\n</ul>\n<p><strong>3D Representations:</strong><br>\n3D representations consider the geometric spatial structure of molecules:</p>\n<ul>\n<li><strong>Geometric Structures:</strong> These include the 3D coordinates of atoms in a molecule, which are crucial for models that consider spatial configurations.</li>\n</ul>\n<h3>Applications and Current Focus</h3>\n<p>Modern research focuses on efficiently extracting and leveraging multimodal information and performing efficient pre-training from rich datasets. This includes 2D and 3D mutual reconstruction, and contrastive learning between NLP and Graph-based methods. Effective pre-training techniques from rich datasets are critical for enhancing model performance.</p>\n<h3>Model Architectures</h3>\n<p>Our model architectures can be categorized as follows:</p>\n<p><strong>Single-modal Models:</strong></p>\n<ul>\n<li><strong>NLP Models:</strong> Examples include ChemBERTa and SmilesLLM.</li>\n<li><strong>Graph Models:</strong> Examples include ECFP (using MLP), 2DGNNs (like MPNN, PAINN), and 3DGNNs (such as SchNet, DimeNet).</li>\n</ul>\n<p><strong>Multimodal Models:</strong></p>\n<ul>\n<li>Examples include UniMol, TransformerM, and GraphMVP, which integrate information from multiple modalities.</li>\n</ul>\n<h3>Challenges and Solutions</h3>\n<p>Given the vast amount of data, designing a data-efficient architecture is crucial. Key considerations include:</p>\n<ol>\n<li><strong>Maximizing GPU Utilization:</strong> Efficient IO operations to increase training efficiency.</li>\n<li><strong>Preventing Memory Leaks:</strong> Avoid setting an excessively high number of workers in DataLoader.</li>\n<li><strong>Pre-storing Graph Features:</strong> Pre-compute graph features to save computation time and storage space.</li>\n</ol>\n<h3>Data Structure and Storage Solutions</h3>\n<p>To address these challenges, we propose the following data structure and storage solutions:</p>\n<ol>\n<li><strong>Feature Storage:</strong> Use a dictionary to store 1D, 2D, and 3D features. Compress binary matrices using <code>np.packbits</code> and convert numerical features to <code>np.float16</code> for storage.</li>\n<li><strong>Database Storage:</strong> Read features from a database(LMDB is all you need) in the dataset to minimize memory usage and implement parallel IO operations to reduce memory leaks.</li>\n<li><strong>Batch Organization:</strong> Define the batch organization strategy in the DataLoader's <code>collate_fn</code>.</li>\n<li><strong>Compression Algorithms:</strong> Use multi-thread friendly, high compression ratio algorithms like <code>blosc2</code> and <code>lz4</code> to compress pre-generated data into the database.</li>\n</ol>",
  "messages": [
    {
      "id": "2830140",
      "postDate": "05/23/2024 03:13:08",
      "content": "<h1>Advancing Molecular Representation: Techniques and Strategies for Effective Multimodal Learning</h1>\n<p>I recently started learning about molecular representation and hope my ideas can help you.</p>\n<h2>Introduction to Molecular Representation</h2>\n<p>Molecular representation is a critical aspect of computational chemistry and drug discovery, encompassing the transformation of molecular structures into formats that can be processed by machine learning models. These representations can be classified into three primary categories: 1D, 2D, and 3D.</p>\n<h3>Categories of Molecular Representations</h3>\n<p><strong>1D Representations:</strong><br>\n1D representations include molecular descriptors and SMILES strings:</p>\n<ul>\n<li><strong>Molecular Descriptors:</strong> These are numerical values that describe various molecular properties, including fingerprints (such as ECFP, MACCS), and other statistical features. They are useful for capturing detailed molecular characteristics.</li>\n<li><strong>String Descriptions:</strong> SMILES strings are textual representations of molecules, ideal for sequence models and spatial models in NLP.</li>\n</ul>\n<p><strong>2D Representations:</strong><br>\n2D representations involve the topological structure of molecules:</p>\n<ul>\n<li><strong>Topological Graph Structures:</strong> These structures capture the connectivity between atoms in a molecule, providing essential information for graph-based neural networks.</li>\n</ul>\n<p><strong>3D Representations:</strong><br>\n3D representations consider the geometric spatial structure of molecules:</p>\n<ul>\n<li><strong>Geometric Structures:</strong> These include the 3D coordinates of atoms in a molecule, which are crucial for models that consider spatial configurations.</li>\n</ul>\n<h3>Applications and Current Focus</h3>\n<p>Modern research focuses on efficiently extracting and leveraging multimodal information and performing efficient pre-training from rich datasets. This includes 2D and 3D mutual reconstruction, and contrastive learning between NLP and Graph-based methods. Effective pre-training techniques from rich datasets are critical for enhancing model performance.</p>\n<h3>Model Architectures</h3>\n<p>Our model architectures can be categorized as follows:</p>\n<p><strong>Single-modal Models:</strong></p>\n<ul>\n<li><strong>NLP Models:</strong> Examples include ChemBERTa and SmilesLLM.</li>\n<li><strong>Graph Models:</strong> Examples include ECFP (using MLP), 2DGNNs (like MPNN, PAINN), and 3DGNNs (such as SchNet, DimeNet).</li>\n</ul>\n<p><strong>Multimodal Models:</strong></p>\n<ul>\n<li>Examples include UniMol, TransformerM, and GraphMVP, which integrate information from multiple modalities.</li>\n</ul>\n<h3>Challenges and Solutions</h3>\n<p>Given the vast amount of data, designing a data-efficient architecture is crucial. Key considerations include:</p>\n<ol>\n<li><strong>Maximizing GPU Utilization:</strong> Efficient IO operations to increase training efficiency.</li>\n<li><strong>Preventing Memory Leaks:</strong> Avoid setting an excessively high number of workers in DataLoader.</li>\n<li><strong>Pre-storing Graph Features:</strong> Pre-compute graph features to save computation time and storage space.</li>\n</ol>\n<h3>Data Structure and Storage Solutions</h3>\n<p>To address these challenges, we propose the following data structure and storage solutions:</p>\n<ol>\n<li><strong>Feature Storage:</strong> Use a dictionary to store 1D, 2D, and 3D features. Compress binary matrices using <code>np.packbits</code> and convert numerical features to <code>np.float16</code> for storage.</li>\n<li><strong>Database Storage:</strong> Read features from a database(LMDB is all you need) in the dataset to minimize memory usage and implement parallel IO operations to reduce memory leaks.</li>\n<li><strong>Batch Organization:</strong> Define the batch organization strategy in the DataLoader's <code>collate_fn</code>.</li>\n<li><strong>Compression Algorithms:</strong> Use multi-thread friendly, high compression ratio algorithms like <code>blosc2</code> and <code>lz4</code> to compress pre-generated data into the database.</li>\n</ol>",
      "rawMarkdown": "# Advancing Molecular Representation: Techniques and Strategies for Effective Multimodal Learning\n\nI recently started learning about molecular representation and hope my ideas can help you.\n\n## Introduction to Molecular Representation\n\nMolecular representation is a critical aspect of computational chemistry and drug discovery, encompassing the transformation of molecular structures into formats that can be processed by machine learning models. These representations can be classified into three primary categories: 1D, 2D, and 3D.\n\n### Categories of Molecular Representations\n\n**1D Representations:**\n1D representations include molecular descriptors and SMILES strings:\n- **Molecular Descriptors:** These are numerical values that describe various molecular properties, including fingerprints (such as ECFP, MACCS), and other statistical features. They are useful for capturing detailed molecular characteristics.\n- **String Descriptions:** SMILES strings are textual representations of molecules, ideal for sequence models and spatial models in NLP.\n\n**2D Representations:**\n2D representations involve the topological structure of molecules:\n- **Topological Graph Structures:** These structures capture the connectivity between atoms in a molecule, providing essential information for graph-based neural networks.\n\n**3D Representations:**\n3D representations consider the geometric spatial structure of molecules:\n- **Geometric Structures:** These include the 3D coordinates of atoms in a molecule, which are crucial for models that consider spatial configurations.\n\n### Applications and Current Focus\n\nModern research focuses on efficiently extracting and leveraging multimodal information and performing efficient pre-training from rich datasets. This includes 2D and 3D mutual reconstruction, and contrastive learning between NLP and Graph-based methods. Effective pre-training techniques from rich datasets are critical for enhancing model performance.\n\n### Model Architectures\n\nOur model architectures can be categorized as follows:\n\n**Single-modal Models:**\n- **NLP Models:** Examples include ChemBERTa and SmilesLLM.\n- **Graph Models:** Examples include ECFP (using MLP), 2DGNNs (like MPNN, PAINN), and 3DGNNs (such as SchNet, DimeNet).\n\n**Multimodal Models:**\n- Examples include UniMol, TransformerM, and GraphMVP, which integrate information from multiple modalities.\n\n### Challenges and Solutions\n\nGiven the vast amount of data, designing a data-efficient architecture is crucial. Key considerations include:\n1. **Maximizing GPU Utilization:** Efficient IO operations to increase training efficiency.\n2. **Preventing Memory Leaks:** Avoid setting an excessively high number of workers in DataLoader.\n3. **Pre-storing Graph Features:** Pre-compute graph features to save computation time and storage space.\n\n### Data Structure and Storage Solutions\n\nTo address these challenges, we propose the following data structure and storage solutions:\n1. **Feature Storage:** Use a dictionary to store 1D, 2D, and 3D features. Compress binary matrices using `np.packbits` and convert numerical features to `np.float16` for storage.\n2. **Database Storage:** Read features from a database(LMDB is all you need) in the dataset to minimize memory usage and implement parallel IO operations to reduce memory leaks.\n3. **Batch Organization:** Define the batch organization strategy in the DataLoader's `collate_fn`.\n4. **Compression Algorithms:** Use multi-thread friendly, high compression ratio algorithms like `blosc2` and `lz4` to compress pre-generated data into the database.",
      "votes": null
    },
    {
      "id": "2834741",
      "postDate": "05/25/2024 01:34:55",
      "content": "<p>let me contribute one magic (?),</p>\n<ul>\n<li>for this dataset, it is special. it has a very long tail (in most feature space)</li>\n<li>hence there is no difference if you</li>\n</ul>\n<ol>\n<li>train with all samples</li>\n<li>train initially with sufficient subset, then finetune with increasing more samples until using the full dataset</li>\n</ol>\n<p>this speedup experiments and training</p>",
      "rawMarkdown": "let me contribute one magic (?),\n- for this dataset, it is special. it has a very long tail (in most feature space)\n- hence there is no difference if you\n1. train with all samples\n2. train initially with sufficient subset, then finetune with increasing more samples until using the full dataset\n\nthis speedup experiments and training",
      "votes": null
    },
    {
      "id": "2835401",
      "postDate": "05/25/2024 09:50:01",
      "content": "<p>I find you great work for 3D-graph modeling, that maybe we should addHs in building conformer, and removeHs when calcuate position distance</p>",
      "rawMarkdown": "I find you great work for 3D-graph modeling, that maybe we should addHs in building conformer, and removeHs when calcuate position distance",
      "votes": null
    },
    {
      "id": "2835517",
      "postDate": "05/25/2024 10:59:37",
      "content": "<p>if you have results and code for using addHs, i can release conformers with H.</p>",
      "rawMarkdown": "if you have results and code for using addHs, i can release conformers with H.",
      "votes": null
    },
    {
      "id": "2837294",
      "postDate": "05/26/2024 12:16:23",
      "content": "<p>Before generating conformer objects, use the following line in the molecuke object.</p>\n<p>Chem.AddHs(mol)</p>\n<p>This is important for getting realistic conformers.</p>",
      "rawMarkdown": "Before generating conformer objects, use the following line in the molecuke object.\n\nChem.AddHs(mol)\n\nThis is important for getting realistic conformers.",
      "votes": null
    },
    {
      "id": "2837789",
      "postDate": "05/26/2024 17:11:55",
      "content": "<pre><code>        mol = AllChem.\n        params = AllChem.\n        params.useBasicKnowledge = True\n        params.useRandomCoords = True\n        params.randomSeed = \n        AllChem.\n        AllChem.\n        mol = AllChem.\n        conf = mol.\n        z = np.(, dtype=np.uint8)\n        positions = conf..astype(np.float16)\n</code></pre>",
      "rawMarkdown": "```\n        mol = AllChem.AddHs(mol)\n        params = AllChem.ETKDGv3()\n        params.useBasicKnowledge = True\n        params.useRandomCoords = True\n        params.randomSeed = 0xF00D\n        AllChem.EmbedMolecule(mol, params)\n        AllChem.MMFFOptimizeMolecule(mol)\n        mol = AllChem.RemoveHs(mol)\n        conf = mol.GetConformer()\n        z = np.array([atom.GetAtomicNum() for atom in mol.GetAtoms()], dtype=np.uint8)\n        positions = conf.GetPositions().astype(np.float16)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2834741,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/25/2024 01:34:55",
      "content": "<p>let me contribute one magic (?),</p>\n<ul>\n<li>for this dataset, it is special. it has a very long tail (in most feature space)</li>\n<li>hence there is no difference if you</li>\n</ul>\n<ol>\n<li>train with all samples</li>\n<li>train initially with sufficient subset, then finetune with increasing more samples until using the full dataset</li>\n</ol>\n<p>this speedup experiments and training</p>",
      "votes": null,
      "replies": [
        {
          "id": 2835401,
          "author_name": "lblhandsome",
          "author_url": "",
          "post_date": "05/25/2024 09:50:01",
          "content": "<p>I find you great work for 3D-graph modeling, that maybe we should addHs in building conformer, and removeHs when calcuate position distance</p>",
          "votes": null,
          "replies": [
            {
              "id": 2835517,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/25/2024 10:59:37",
              "content": "<p>if you have results and code for using addHs, i can release conformers with H.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2837294,
                  "author_name": "chemdatafarmer",
                  "author_url": "",
                  "post_date": "05/26/2024 12:16:23",
                  "content": "<p>Before generating conformer objects, use the following line in the molecuke object.</p>\n<p>Chem.AddHs(mol)</p>\n<p>This is important for getting realistic conformers.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2837789,
                  "author_name": "lblhandsome",
                  "author_url": "",
                  "post_date": "05/26/2024 17:11:55",
                  "content": "<pre><code>        mol = AllChem.\n        params = AllChem.\n        params.useBasicKnowledge = True\n        params.useRandomCoords = True\n        params.randomSeed = \n        AllChem.\n        AllChem.\n        mol = AllChem.\n        conf = mol.\n        z = np.(, dtype=np.uint8)\n        positions = conf..astype(np.float16)\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2830140": "# Advancing Molecular Representation: Techniques and Strategies for Effective Multimodal Learning\n\nI recently started learning about molecular representation and hope my ideas can help you.\n\n## Introduction to Molecular Representation\n\nMolecular representation is a critical aspect of computational chemistry and drug discovery, encompassing the transformation of molecular structures into formats that can be processed by machine learning models. These representations can be classified into three primary categories: 1D, 2D, and 3D.\n\n### Categories of Molecular Representations\n\n**1D Representations:**\n1D representations include molecular descriptors and SMILES strings:\n- **Molecular Descriptors:** These are numerical values that describe various molecular properties, including fingerprints (such as ECFP, MACCS), and other statistical features. They are useful for capturing detailed molecular characteristics.\n- **String Descriptions:** SMILES strings are textual representations of molecules, ideal for sequence models and spatial models in NLP.\n\n**2D Representations:**\n2D representations involve the topological structure of molecules:\n- **Topological Graph Structures:** These structures capture the connectivity between atoms in a molecule, providing essential information for graph-based neural networks.\n\n**3D Representations:**\n3D representations consider the geometric spatial structure of molecules:\n- **Geometric Structures:** These include the 3D coordinates of atoms in a molecule, which are crucial for models that consider spatial configurations.\n\n### Applications and Current Focus\n\nModern research focuses on efficiently extracting and leveraging multimodal information and performing efficient pre-training from rich datasets. This includes 2D and 3D mutual reconstruction, and contrastive learning between NLP and Graph-based methods. Effective pre-training techniques from rich datasets are critical for enhancing model performance.\n\n### Model Architectures\n\nOur model architectures can be categorized as follows:\n\n**Single-modal Models:**\n- **NLP Models:** Examples include ChemBERTa and SmilesLLM.\n- **Graph Models:** Examples include ECFP (using MLP), 2DGNNs (like MPNN, PAINN), and 3DGNNs (such as SchNet, DimeNet).\n\n**Multimodal Models:**\n- Examples include UniMol, TransformerM, and GraphMVP, which integrate information from multiple modalities.\n\n### Challenges and Solutions\n\nGiven the vast amount of data, designing a data-efficient architecture is crucial. Key considerations include:\n1. **Maximizing GPU Utilization:** Efficient IO operations to increase training efficiency.\n2. **Preventing Memory Leaks:** Avoid setting an excessively high number of workers in DataLoader.\n3. **Pre-storing Graph Features:** Pre-compute graph features to save computation time and storage space.\n\n### Data Structure and Storage Solutions\n\nTo address these challenges, we propose the following data structure and storage solutions:\n1. **Feature Storage:** Use a dictionary to store 1D, 2D, and 3D features. Compress binary matrices using `np.packbits` and convert numerical features to `np.float16` for storage.\n2. **Database Storage:** Read features from a database(LMDB is all you need) in the dataset to minimize memory usage and implement parallel IO operations to reduce memory leaks.\n3. **Batch Organization:** Define the batch organization strategy in the DataLoader's `collate_fn`.\n4. **Compression Algorithms:** Use multi-thread friendly, high compression ratio algorithms like `blosc2` and `lz4` to compress pre-generated data into the database.",
    "2834741": "let me contribute one magic (?),\n- for this dataset, it is special. it has a very long tail (in most feature space)\n- hence there is no difference if you\n1. train with all samples\n2. train initially with sufficient subset, then finetune with increasing more samples until using the full dataset\n\nthis speedup experiments and training",
    "2835401": "I find you great work for 3D-graph modeling, that maybe we should addHs in building conformer, and removeHs when calcuate position distance",
    "2835517": "if you have results and code for using addHs, i can release conformers with H.",
    "2837294": "Before generating conformer objects, use the following line in the molecuke object.\n\nChem.AddHs(mol)\n\nThis is important for getting realistic conformers.",
    "2837789": "```\n        mol = AllChem.AddHs(mol)\n        params = AllChem.ETKDGv3()\n        params.useBasicKnowledge = True\n        params.useRandomCoords = True\n        params.randomSeed = 0xF00D\n        AllChem.EmbedMolecule(mol, params)\n        AllChem.MMFFOptimizeMolecule(mol)\n        mol = AllChem.RemoveHs(mol)\n        conf = mol.GetConformer()\n        z = np.array([atom.GetAtomicNum() for atom in mol.GetAtoms()], dtype=np.uint8)\n        positions = conf.GetPositions().astype(np.float16)\n```"
  },
  "source": "meta"
}