{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"},{"sourceId":8336706,"sourceType":"datasetVersion","datasetId":4951137},{"sourceId":170595844,"sourceType":"kernelVersion"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BELKA: Random Forest with RDKit Descriptors Demo\n\nAuthor: \n- Stephen Lee ([LinkedIn](https://www.linkedin.com/in/stephendongsoolee/))\n\n# 📖 Background\n---\n\nMolecular descriptors are quantitative properties of a chemical structure, and can be used for machine learning. Descriptors are one of the simplest feature types that can be generated from SMILES and have the advantage of being readily interpretable. \n\nThis notebook aims to provide demo code to\n1. generate RDKit descriptors from SMILES \n2. preprocess descriptors and apply feature selection\n3. train Random Forest classifiers (one per target, without tuning) using selected descriptors\n4. conduct inference and submit predictions","metadata":{}},{"cell_type":"markdown","source":"# 🛠️ Environment Setup\n---","metadata":{}},{"cell_type":"code","source":"!pip install -q rdkit-pypi","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:23:09.632292Z","iopub.execute_input":"2024-05-06T12:23:09.633228Z","iopub.status.idle":"2024-05-06T12:23:23.333523Z","shell.execute_reply.started":"2024-05-06T12:23:09.633114Z","shell.execute_reply":"2024-05-06T12:23:23.331983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \n\nfrom joblib import Parallel, delayed\n\nfrom rdkit import Chem\nfrom rdkit.Chem import Descriptors, AllChem\nfrom rdkit.ML.Descriptors import MoleculeDescriptors\n\nfrom sklearn.feature_selection import VarianceThreshold\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import precision_score","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-06T12:23:23.336255Z","iopub.execute_input":"2024-05-06T12:23:23.336901Z","iopub.status.idle":"2024-05-06T12:23:26.688085Z","shell.execute_reply.started":"2024-05-06T12:23:23.336848Z","shell.execute_reply":"2024-05-06T12:23:26.687021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# define config variables\nclass CFG:\n    train_path = '/kaggle/input/belka-shrinking-the-dataset/train.csv'\n    test_path = '/kaggle/input/leash-BELKA/test.parquet'\n    n_samples_per_label = 5000\n    low_var_cutoff = 0.01\n    corr_threshold = 0.7","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:23:26.689384Z","iopub.execute_input":"2024-05-06T12:23:26.689862Z","iopub.status.idle":"2024-05-06T12:23:26.694832Z","shell.execute_reply.started":"2024-05-06T12:23:26.689829Z","shell.execute_reply":"2024-05-06T12:23:26.693816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🗃️ Load Data\n---","metadata":{}},{"cell_type":"markdown","source":"The original training dataset is too large to be loaded in Kaggle. Credit to shloromon for creating and making available a [shrunken version of the training dataset](https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset).","metadata":{}},{"cell_type":"code","source":"%%time\n# load train dataset\n# adapted from https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset\ndtypes = {'binds_BRD4': np.byte, 'binds_HSA': np.byte, 'binds_sEH': np.byte}\n\ntrain_df = pd.read_csv(\n    CFG.train_path, \n    usecols=['molecule_smiles', 'binds_BRD4', 'binds_HSA', 'binds_sEH'], \n    dtype=dtypes\n)\nprint(len(train_df))\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:23:26.697890Z","iopub.execute_input":"2024-05-06T12:23:26.698342Z","iopub.status.idle":"2024-05-06T12:27:52.942223Z","shell.execute_reply.started":"2024-05-06T12:23:26.698302Z","shell.execute_reply":"2024-05-06T12:27:52.941088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For this demo, a small subset of the training dataset is sampled for each target with an equal abundance of positive and negative labels.","metadata":{}},{"cell_type":"code","source":"%%time\n# randomly sample an equal number of negative and positive samples for each target\ndef sample_data(df, target):\n    col = f'binds_{target}'\n    pos_samples = df[df[col]==True].sample(CFG.n_samples_per_label, random_state=42)\n    neg_samples = df[df[col]==False].sample(CFG.n_samples_per_label, random_state=42)\n\n    return pd.concat([pos_samples, neg_samples])\n\nbrd4_df = sample_data(train_df, 'BRD4')\nhsa_df = sample_data(train_df, 'HSA')\nseh_df = sample_data(train_df, 'sEH')\n\nprint(len(brd4_df))\nprint(brd4_df['binds_BRD4'].value_counts())\n\n# delete train_df for free up memory\ndel train_df","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:27:52.944661Z","iopub.execute_input":"2024-05-06T12:27:52.944993Z","iopub.status.idle":"2024-05-06T12:28:23.960364Z","shell.execute_reply.started":"2024-05-06T12:27:52.944965Z","shell.execute_reply":"2024-05-06T12:28:23.959170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following sections uses the BRD4 subset for demonstrating the feature generation, selection and model training steps.","metadata":{}},{"cell_type":"markdown","source":"# 🧪 Generate Molecular Descriptors\n---","metadata":{}},{"cell_type":"markdown","source":"RDKit can generate 200 unique descriptors out-of-the-box. While domain knowledge can be used to preemptively select descriptors, the approach employed in this notebook is to generate all RDKit descriptors and to use data-driven selection methods to narrow them down.","metadata":{}},{"cell_type":"code","source":"# generate RDKit Mol objects from SMILES\nmols = brd4_df['molecule_smiles'].apply(Chem.MolFromSmiles)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:28:23.961699Z","iopub.execute_input":"2024-05-06T12:28:23.962028Z","iopub.status.idle":"2024-05-06T12:28:27.001096Z","shell.execute_reply.started":"2024-05-06T12:28:23.961999Z","shell.execute_reply":"2024-05-06T12:28:26.999791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we are dealing with a large dataset, the following optimizations were included in the following generate_descriptors function:\n- casting smaller data types to conserve memory\n- parallel processing using joblib\n\nNote that the desc_names parameter allows for the generation of specified descriptors - this will be relevant at inference time since only the selected descriptors need to be generated for the test set.","metadata":{}},{"cell_type":"code","source":"%%time\ndef generate_descriptors(mols, desc_names=None):\n    if desc_names is not None:\n        desc_list = [x for x in Descriptors._descList if x[0] in desc_names]\n        calc = MoleculeDescriptors.MolecularDescriptorCalculator([x[0] for x in desc_list])\n    else:  # all RDKit descriptors\n        calc = MoleculeDescriptors.MolecularDescriptorCalculator([x[0] for x in Descriptors._descList])\n\n    desc_names = calc.GetDescriptorNames()\n    \n    def process_molecule(mol):\n        rdkit_descriptors = calc.CalcDescriptors(mol)\n        descriptors_converted = np.array(rdkit_descriptors, dtype=np.float32)\n\n        # identify indices for binary and integer descriptors\n        binary_indices = [i for i, desc in enumerate(desc_names) if 'fr_' in desc]\n        integer_indices = [i for i, val in enumerate(rdkit_descriptors) if isinstance(val, int) and i not in binary_indices]\n\n        # convert binary and integer descriptors to smaller dtype\n        descriptors_converted[binary_indices] = descriptors_converted[binary_indices].astype(np.byte)\n        descriptors_converted[integer_indices] = descriptors_converted[integer_indices].astype(np.int32)\n\n        return descriptors_converted\n\n    # use joblib to parallelize the molecule processing\n    mol_descriptors = Parallel(n_jobs=-1)(delayed(process_molecule)(mol) for mol in mols)\n\n    return pd.DataFrame(mol_descriptors, columns=desc_names)\n    \nX = generate_descriptors(mols)\nprint(X.shape)\nX.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:28:27.002862Z","iopub.execute_input":"2024-05-06T12:28:27.003295Z","iopub.status.idle":"2024-05-06T12:30:05.874803Z","shell.execute_reply.started":"2024-05-06T12:28:27.003235Z","shell.execute_reply":"2024-05-06T12:30:05.873496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, the data is split into a train and validation set prior to feature selection to avoid data leakage.","metadata":{}},{"cell_type":"code","source":"y = brd4_df['binds_BRD4']\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\nprint(len(X_train))\nprint(len(X_val))","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:05.876791Z","iopub.execute_input":"2024-05-06T12:30:05.877256Z","iopub.status.idle":"2024-05-06T12:30:05.892485Z","shell.execute_reply.started":"2024-05-06T12:30:05.877213Z","shell.execute_reply":"2024-05-06T12:30:05.891246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ✅ Feature Selection\n---","metadata":{}},{"cell_type":"markdown","source":"Descriptors should be examined for undesired artifacts such as constant features, infinite values and missing values, and processed accordingly. Optionally, features can be further selected to exclude those that are unlikely to contribute to model training such as features with low variance or highly correlated features.\n\nFirst we check and drop any features that have a constant value across all samples since they do not contribute any useful information for training.","metadata":{}},{"cell_type":"code","source":"# drop constant features\nprint(f'Starting with {len(X_train.columns)} features')\n\ndef drop_constant_feats(df):\n    constant_feats = [col for col in df.columns if len(df[col].unique()) == 1]\n    output_df = df.drop(constant_feats, axis=1)\n    print(f'-- Dropped {len(constant_feats)} constant features, {len(output_df.columns)} remaining')\n    return output_df\n\nX_train = drop_constant_feats(X_train)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:05.893969Z","iopub.execute_input":"2024-05-06T12:30:05.894421Z","iopub.status.idle":"2024-05-06T12:30:05.950103Z","shell.execute_reply.started":"2024-05-06T12:30:05.894381Z","shell.execute_reply":"2024-05-06T12:30:05.949040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Certain RDKit descriptors can sometimes produce infinite values - the data is checked for such descriptors below.","metadata":{}},{"cell_type":"code","source":"# check for infinite values\ninf_counts = X_train.isin([np.inf, -np.inf]).sum()\ninf_cols_and_counts = inf_counts[inf_counts > 0]\nprint(inf_cols_and_counts)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:05.954984Z","iopub.execute_input":"2024-05-06T12:30:05.955363Z","iopub.status.idle":"2024-05-06T12:30:06.060032Z","shell.execute_reply.started":"2024-05-06T12:30:05.955331Z","shell.execute_reply":"2024-05-06T12:30:06.058808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For simplicity, we drop the columns containing infinite values. Another option would be to replace them with some imputed value (e.g. mean, KNN imputed value).","metadata":{}},{"cell_type":"code","source":"# drop features containing infinite values\nX_train.drop(inf_cols_and_counts.index, axis=1, inplace=True)\nprint(f'-- {len(X_train.columns)} features remaining')","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:06.061597Z","iopub.execute_input":"2024-05-06T12:30:06.061936Z","iopub.status.idle":"2024-05-06T12:30:06.069581Z","shell.execute_reply.started":"2024-05-06T12:30:06.061895Z","shell.execute_reply":"2024-05-06T12:30:06.068438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for features containing missing values\nmissing_counts = X_train.isna().sum()\nmissing_counts[missing_counts > 0]","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:06.071131Z","iopub.execute_input":"2024-05-06T12:30:06.071968Z","iopub.status.idle":"2024-05-06T12:30:06.083540Z","shell.execute_reply.started":"2024-05-06T12:30:06.071929Z","shell.execute_reply":"2024-05-06T12:30:06.082543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we apply a filter to exclude any features that exhibit a low variance. The rationale here is that features with very low variance are unlikely to provide discriminative information for the model.","metadata":{}},{"cell_type":"code","source":"def filter_by_low_var(df, threshold):\n    selector = VarianceThreshold(threshold)\n    selector.fit(df)\n    output_df = df[df.columns[selector.get_support(indices=True)]]\n    print(f'-- Dropped features with variance lower than {threshold}, {len(output_df.columns)} features remaining')\n    return output_df\n\nX_train = filter_by_low_var(X_train, threshold=CFG.low_var_cutoff)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:06.085009Z","iopub.execute_input":"2024-05-06T12:30:06.085408Z","iopub.status.idle":"2024-05-06T12:30:06.120478Z","shell.execute_reply.started":"2024-05-06T12:30:06.085379Z","shell.execute_reply":"2024-05-06T12:30:06.119301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we apply a filter to drop highly correlated features. The rationale here is that features that exhibit high correlation with each other are likely redundant (multicolinearity). Below is an implementation of the methodology for removing highly correlated descriptors described in this [paper](https://www.sciencedirect.com/science/article/pii/S0016236122006950).\n\nTo summarize the method:\n- The process begins by selecting the feature with the highest number of correlations\n- It then identifies all features that are highly correlated with this selected feature\n- These correlated features, including the selected feature itself, are then removed from the pool of all features\n- The process repeats until no features are left to select from","metadata":{}},{"cell_type":"code","source":"def filter_by_correlation(df, threshold=0.7):\n    # generate correlation matrix\n    corr_matrix = df.corr(method='spearman')\n\n    # rank each feature based on how many others are highly correlated with it (as per threshold)\n    correlated_counts = corr_matrix.apply(lambda x: (x >= threshold).sum())\n    ranked_features = correlated_counts.sort_values(ascending=False)\n\n    # iteratively select feature with the most number of other correlated features\n    selected_features = []\n    while not ranked_features.empty:\n        # select the highest reanked feature\n        selected_feature = ranked_features.index[0]\n        selected_features.append(selected_feature)\n\n        # find features correlated to the selected feature\n        correlated_features = corr_matrix[selected_feature][corr_matrix[selected_feature] >= threshold].index\n\n        # remove the selected feature and its correlated features from the correlation matrix\n        corr_matrix = corr_matrix.drop(index=correlated_features, columns=correlated_features)\n\n        # update the ranked_features for the next iteration\n        correlated_counts = corr_matrix.apply(lambda x: (x >= threshold).sum())\n        ranked_features = correlated_counts.sort_values(ascending=False)\n\n    # drop features that are not selected\n    output_df = df[selected_features]\n    print(f'-- {len(output_df.columns)} remaining features')\n\n    return output_df\n\nX_train = filter_by_correlation(X_train, threshold=CFG.corr_threshold)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:06.121883Z","iopub.execute_input":"2024-05-06T12:30:06.122270Z","iopub.status.idle":"2024-05-06T12:30:07.756030Z","shell.execute_reply.started":"2024-05-06T12:30:06.122233Z","shell.execute_reply":"2024-05-06T12:30:07.754938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that feature selection is finished, X_val is updated to contain the same features as X_train.","metadata":{}},{"cell_type":"code","source":"X_val = X_val[X_train.columns]\nprint(len(X_val.columns))","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:07.757442Z","iopub.execute_input":"2024-05-06T12:30:07.757786Z","iopub.status.idle":"2024-05-06T12:30:07.764590Z","shell.execute_reply.started":"2024-05-06T12:30:07.757756Z","shell.execute_reply":"2024-05-06T12:30:07.763261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 💪 Train Model\n---\n\nDescriptors tend to have skewed distributions and widely variable scales. This makes tree-based algorithms preferable as they are not susceptible to such characteristics. A simple Random Forest classifier is fitted on the training data and evaluated against the validation set using the precision score. ","metadata":{}},{"cell_type":"code","source":"brd4_model = RandomForestClassifier(n_estimators=100, random_state=42)\nbrd4_model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:07.765945Z","iopub.execute_input":"2024-05-06T12:30:07.766417Z","iopub.status.idle":"2024-05-06T12:30:11.322661Z","shell.execute_reply.started":"2024-05-06T12:30:07.766387Z","shell.execute_reply":"2024-05-06T12:30:11.321531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = brd4_model.predict(X_val)\nprint(f'{precision_score(y_val, y_pred):.2f}')","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:11.323962Z","iopub.execute_input":"2024-05-06T12:30:11.324361Z","iopub.status.idle":"2024-05-06T12:30:11.389508Z","shell.execute_reply.started":"2024-05-06T12:30:11.324329Z","shell.execute_reply":"2024-05-06T12:30:11.388418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, all of the steps are encapsulated in a function and applied to derive models for HSA and sEH. Note that a different set of descriptors may have been selected for different targets - the function below also returns the selected descriptors to be used at inference time.","metadata":{}},{"cell_type":"code","source":"def run_all_steps(df, target):\n    print('Generating Mol objects')\n    mols = df['molecule_smiles'].apply(Chem.MolFromSmiles)\n    \n    print('Generating RDKit descriptors')\n    X = generate_descriptors(mols)\n    y = df[f'binds_{target}']\n    \n    X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\n    \n    print('Selecting Features')\n    X_train = drop_constant_feats(X_train)\n    \n    inf_counts = X_train.isin([np.inf, -np.inf]).sum()\n    inf_cols_and_counts = inf_counts[inf_counts > 0]\n    if len(inf_cols_and_counts) > 0:\n        X_train.drop(inf_cols_and_counts.index, axis=1, inplace=True)\n        print(f'-- Dropped features with inf values, {len(X_train.columns)} features remaining')\n\n    missing_counts = X_train.isna().sum()\n    missing_cols_and_counts = missing_counts[missing_counts > 0]\n    if len(missing_cols_and_counts) > 0:\n        X_train.drop(missing_cols_and_counts.index, axis=1, inplace=True)\n        print(f'-- Dropped features with missing values, {len(X_train.columns)} features remaining')\n        \n    X_train = filter_by_low_var(X_train, threshold=CFG.low_var_cutoff)\n    X_train = filter_by_correlation(X_train, threshold=CFG.corr_threshold)\n    \n    # get filtered descriptor names for inference\n    desc_names = X_train.columns\n    X_val = X_val[desc_names]\n    \n    print('Training Model')\n    rf_classifier = RandomForestClassifier(n_estimators=100, random_state=42)\n    rf_classifier.fit(X_train, y_train)\n    \n    y_pred = rf_classifier.predict(X_val)\n    print(f'-- Precision Score: {precision_score(y_val, y_pred):.2f}')\n    \n    return rf_classifier, desc_names","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:11.390706Z","iopub.execute_input":"2024-05-06T12:30:11.391070Z","iopub.status.idle":"2024-05-06T12:30:11.401333Z","shell.execute_reply.started":"2024-05-06T12:30:11.391041Z","shell.execute_reply":"2024-05-06T12:30:11.400252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nhsa_model, hsa_descriptors = run_all_steps(hsa_df, 'HSA')","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:30:11.402584Z","iopub.execute_input":"2024-05-06T12:30:11.402923Z","iopub.status.idle":"2024-05-06T12:31:56.475320Z","shell.execute_reply.started":"2024-05-06T12:30:11.402894Z","shell.execute_reply":"2024-05-06T12:31:56.474296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nseh_model, seh_descriptors = run_all_steps(seh_df, 'sEH')","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:31:56.476706Z","iopub.execute_input":"2024-05-06T12:31:56.477126Z","iopub.status.idle":"2024-05-06T12:33:40.903606Z","shell.execute_reply.started":"2024-05-06T12:31:56.477096Z","shell.execute_reply":"2024-05-06T12:33:40.902576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🎯 Inference\n---","metadata":{}},{"cell_type":"markdown","source":"To generate predictions for the three targets, the following steps are taken:\n- load test dataset\n- generate all descriptors across the three sets of selected descriptors\n- for each target, subset test data accordingly (using protein_name and selected descriptors) and generate probabilities for the positive class\n- write predictions to file for submission","metadata":{}},{"cell_type":"code","source":"# load test dataset\ntest_df = pd.read_parquet('/kaggle/input/leash-BELKA/test.parquet')\ntest_df.drop(columns=['buildingblock1_smiles', 'buildingblock2_smiles', 'buildingblock3_smiles'], inplace=True)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:33:40.905035Z","iopub.execute_input":"2024-05-06T12:33:40.905389Z","iopub.status.idle":"2024-05-06T12:33:42.515387Z","shell.execute_reply.started":"2024-05-06T12:33:40.905361Z","shell.execute_reply":"2024-05-06T12:33:42.514315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"brd4_descriptors = X_train.columns\nprint(len(brd4_descriptors))\n\n# combine all unique descriptors across the three targets\nall_descriptors = set(brd4_descriptors.append(hsa_descriptors).append(seh_descriptors))\nprint(len(all_descriptors))","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:33:42.517001Z","iopub.execute_input":"2024-05-06T12:33:42.517353Z","iopub.status.idle":"2024-05-06T12:33:42.522851Z","shell.execute_reply.started":"2024-05-06T12:33:42.517323Z","shell.execute_reply":"2024-05-06T12:33:42.521752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndef generate_features(df, feature_names):\n    # get all unique smiles from df\n    smiles = list(set(df['molecule_smiles']))\n    mols = [Chem.MolFromSmiles(smi) for smi in smiles]\n    \n    desc_df = generate_descriptors(mols, feature_names)\n    desc_df.insert(0, 'molecule_smiles', smiles)\n    \n    X_test = pd.merge(df, desc_df, on='molecule_smiles', how='left')\n    return X_test","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:33:42.536036Z","iopub.execute_input":"2024-05-06T12:33:42.536748Z","iopub.status.idle":"2024-05-06T12:33:42.552497Z","shell.execute_reply.started":"2024-05-06T12:33:42.536713Z","shell.execute_reply":"2024-05-06T12:33:42.551038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generation of descriptors using the entire test set can lead to OOM, hence the following cells are commented. For this notebook's submission, I ran the code on Google Colab and uploaded the resulting submission.csv as a dataset.","metadata":{}},{"cell_type":"code","source":"# %%time\n# test_feature_df = generate_features(test_df, all_descriptors)\n# print(test_feature_df.shape)\n# print(test_feature_df.head())","metadata":{"execution":{"iopub.status.busy":"2024-05-06T12:33:42.553899Z","iopub.execute_input":"2024-05-06T12:33:42.554307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def predict(df, target, descriptors, model):\n#     output_df = df[df['protein_name']==target]\n    \n#     X_test = output_df[descriptors.tolist()]\n\n#     output_df = output_df[['id']]\n#     output_df['binds'] = model.predict_proba(X_test)[:, 1]\n#     return output_df","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# brd_pred_df = predict(test_feature_df, 'BRD4', brd4_descriptors, brd4_model)\n# hsa_pred_df = predict(test_feature_df, 'HSA', hsa_descriptors, hsa_model)\n# seh_pred_df = predict(test_feature_df, 'sEH', seh_descriptors, seh_model)\n# submission_df = pd.concat([brd_pred_df, hsa_pred_df, seh_pred_df])\n# submission_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.read_csv('/kaggle/input/belka-sdlee94-submission/submission_df.csv')\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Have any feedback or questions? Comment below!","metadata":{}}]}