{
  "id": 493607,
  "title": " A script to extract a refined small subset",
  "url": "/competitions/leash-BELKA/discussion/493607",
  "author_name": "Wang-Lin-Boop",
  "post_date": "2024-04-14T05:39:51.550000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Training a model on a dataset of 98 million molecules is a formidable task. Fortunately, the organizers have provided us with building blocks that built the compounds. We can extract all the positive compounds and a subset of the negative compounds based on the building blocks of the positive compounds. This will allow us to create a smaller subset (&lt;5M) for accelerated training, facilitating the adjustment of model architecture and the evaluation of their generalization capabilities.</p>\n<p>Here, we will process the <code>train.parquet</code> file provided by the organizers in five steps:</p>\n<ol>\n<li>Remove [Dy] from the SMILES, which will result in a naked molecule that can be recognized by chemical informatics programs such as RDKit.</li>\n<li>Reshape the Dataframe so that each molecule can correspond to three labels simultaneously.</li>\n<li>Extract all the molecules that interact with at least one target and sample negative data three times the number of positive samples from the remaining molecules.</li>\n<li>Split the dataset based on building blocks to ensure that the training and testing sets do not contain the same building blocks.</li>\n<li>Take a portion of the remaining molecules as the validation set. In the training process, it has been observed that HSA is more challenging to generalize compared to the other two targets. Therefore, we will sample twice the number of positive HSA molecules for the validation set.</li>\n</ol>\n<p>Please note that this script will generate a <code>test.parquet</code> file, which may overwrite any existing file in the specified output path, such as the provided test set by the organizers.</p>\n<pre><code>     = sys.argv[]   \n     = sys.argv[]  \n     =   \n     = [, , ]  \n     = \n</code></pre>",
  "messages": [
    {
      "id": 2751236,
      "postDate": "2024-04-14T05:39:51.550Z",
      "content": "<p>Training a model on a dataset of 98 million molecules is a formidable task. Fortunately, the organizers have provided us with building blocks that built the compounds. We can extract all the positive compounds and a subset of the negative compounds based on the building blocks of the positive compounds. This will allow us to create a smaller subset (&lt;5M) for accelerated training, facilitating the adjustment of model architecture and the evaluation of their generalization capabilities.</p>\n<p>Here, we will process the <code>train.parquet</code> file provided by the organizers in five steps:</p>\n<ol>\n<li>Remove [Dy] from the SMILES, which will result in a naked molecule that can be recognized by chemical informatics programs such as RDKit.</li>\n<li>Reshape the Dataframe so that each molecule can correspond to three labels simultaneously.</li>\n<li>Extract all the molecules that interact with at least one target and sample negative data three times the number of positive samples from the remaining molecules.</li>\n<li>Split the dataset based on building blocks to ensure that the training and testing sets do not contain the same building blocks.</li>\n<li>Take a portion of the remaining molecules as the validation set. In the training process, it has been observed that HSA is more challenging to generalize compared to the other two targets. Therefore, we will sample twice the number of positive HSA molecules for the validation set.</li>\n</ol>\n<p>Please note that this script will generate a <code>test.parquet</code> file, which may overwrite any existing file in the specified output path, such as the provided test set by the organizers.</p>\n<pre><code>     = sys.argv[]   \n     = sys.argv[]  \n     =   \n     = [, , ]  \n     = \n</code></pre>",
      "rawMarkdown": "Training a model on a dataset of 98 million molecules is a formidable task. Fortunately, the organizers have provided us with building blocks that built the compounds. We can extract all the positive compounds and a subset of the negative compounds based on the building blocks of the positive compounds. This will allow us to create a smaller subset (<5M) for accelerated training, facilitating the adjustment of model architecture and the evaluation of their generalization capabilities.\n\nHere, we will process the `train.parquet` file provided by the organizers in five steps:\n\n1.  Remove [Dy] from the SMILES, which will result in a naked molecule that can be recognized by chemical informatics programs such as RDKit.\n2.  Reshape the Dataframe so that each molecule can correspond to three labels simultaneously.\n3.  Extract all the molecules that interact with at least one target and sample negative data three times the number of positive samples from the remaining molecules.\n4.  Split the dataset based on building blocks to ensure that the training and testing sets do not contain the same building blocks.\n5.  Take a portion of the remaining molecules as the validation set. In the training process, it has been observed that HSA is more challenging to generalize compared to the other two targets. Therefore, we will sample twice the number of positive HSA molecules for the validation set.\n\nPlease note that this script will generate a `test.parquet` file, which may overwrite any existing file in the specified output path, such as the provided test set by the organizers.\n\n```\n    data_path = sys.argv[1]   # the path to orignal datasets\n    output_path = sys.argv[2]  # output path\n    pos_neg_ratio = 3  # ratio of P/N\n    buildingblock_ratio = [16, 32, 48]  # sample ratios for buildingblock 1, 2, and 3\n    random_seed = 1207\n```",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2751236": "Training a model on a dataset of 98 million molecules is a formidable task. Fortunately, the organizers have provided us with building blocks that built the compounds. We can extract all the positive compounds and a subset of the negative compounds based on the building blocks of the positive compounds. This will allow us to create a smaller subset (<5M) for accelerated training, facilitating the adjustment of model architecture and the evaluation of their generalization capabilities.\n\nHere, we will process the `train.parquet` file provided by the organizers in five steps:\n\n1.  Remove [Dy] from the SMILES, which will result in a naked molecule that can be recognized by chemical informatics programs such as RDKit.\n2.  Reshape the Dataframe so that each molecule can correspond to three labels simultaneously.\n3.  Extract all the molecules that interact with at least one target and sample negative data three times the number of positive samples from the remaining molecules.\n4.  Split the dataset based on building blocks to ensure that the training and testing sets do not contain the same building blocks.\n5.  Take a portion of the remaining molecules as the validation set. In the training process, it has been observed that HSA is more challenging to generalize compared to the other two targets. Therefore, we will sample twice the number of positive HSA molecules for the validation set.\n\nPlease note that this script will generate a `test.parquet` file, which may overwrite any existing file in the specified output path, such as the provided test set by the organizers.\n\n```\n    data_path = sys.argv[1]   # the path to orignal datasets\n    output_path = sys.argv[2]  # output path\n    pos_neg_ratio = 3  # ratio of P/N\n    buildingblock_ratio = [16, 32, 48]  # sample ratios for buildingblock 1, 2, and 3\n    random_seed = 1207\n```"
  }
}