{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"},{"sourceId":8042988,"sourceType":"datasetVersion","datasetId":4740586}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-29T04:23:52.910589Z","iopub.execute_input":"2024-05-29T04:23:52.911021Z","iopub.status.idle":"2024-05-29T04:23:54.342409Z","shell.execute_reply.started":"2024-05-29T04:23:52.910989Z","shell.execute_reply":"2024-05-29T04:23:54.341175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a Kaggle Post Dataset\nThis notebook demonstrates the process of creating a shrunken dataset from the BELKA competition train set. The shrunken dataset significantly reduces the size of the original train set while retaining the essential information. This work was facilitated using ChatGPT-4 and referenced the following dataset:\n\n## Reference Dataset:\n* BELKA: Shrunken train set by graysnow, available at: [Kaggle Dataset Link](https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set)\n\n*Note*: This optimized version is created with the help of ChatGPT-4 and follows a similar shrinking strategy as the referenced dataset.","metadata":{}},{"cell_type":"markdown","source":"## About Dataset\nThe dataset contains the following optimizations:\n\n* No ID column.\n* binds columns saved in bytes.\n* Building block SMILES columns saved as int16, with encoded indices.\n* Protein/label columns transformed into three columns of labels per protein, reducing the dataset length by a third.\n","metadata":{}},{"cell_type":"markdown","source":"## Loading and Preprocessing the Train Data\nFirst, we load and preprocess the train data. This involves mapping the building block SMILES strings to integer indices and transforming the labels.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ndef load_and_preprocess_train_data(filepath, chunksize=10000):\n    # Load dictionaries for building blocks\n    bb1_dict = pd.read_pickle('/kaggle/input/belka-shrunken-train-set/test_dicts/BBs_dict_1_test.p')\n    bb2_dict = pd.read_pickle('/kaggle/input/belka-shrunken-train-set/test_dicts/BBs_dict_2_test.p')\n    bb3_dict = pd.read_pickle('/kaggle/input/belka-shrunken-train-set/test_dicts/BBs_dict_3_test.p')\n\n    chunks = []\n    for chunk in pd.read_csv(filepath, chunksize=chunksize):\n        chunk['buildingblock1_smiles'] = chunk['buildingblock1_smiles'].map(bb1_dict).fillna(-1).astype('int16')\n        chunk['buildingblock2_smiles'] = chunk['buildingblock2_smiles'].map(bb2_dict).fillna(-1).astype('int16')\n        chunk['buildingblock3_smiles'] = chunk['buildingblock3_smiles'].map(bb3_dict).fillna(-1).astype('int16')\n        chunk = transform_labels(chunk)\n        chunks.append(chunk)\n\n    return pd.concat(chunks, ignore_index=True), bb1_dict, bb2_dict, bb3_dict\n\ndef transform_labels(df):\n    min_length = len(df) - len(df) % 3\n    df_transformed = pd.DataFrame({\n        'buildingblock1_smiles': df['buildingblock1_smiles'][:min_length:3].values,\n        'buildingblock2_smiles': df['buildingblock2_smiles'][:min_length:3].values,\n        'buildingblock3_smiles': df['buildingblock3_smiles'][:min_length:3].values,\n        'molecule_smiles': df['molecule_smiles'][:min_length:3].values,\n        'binds_BRD4': df['binds'][:min_length:3].values.astype('int'),\n        'binds_HSA': df['binds'][1:min_length:3].values.astype('int'),\n        'binds_sEH': df['binds'][2:min_length:3].values.astype('int'),\n    })\n    return df_transformed\n\n# Load and preprocess the train data\ndf_train, bb1_dict, bb2_dict, bb3_dict = load_and_preprocess_train_data('/kaggle/input/leash-BELKA/train.csv')","metadata":{"execution":{"iopub.status.busy":"2024-05-29T05:40:24.080924Z","iopub.execute_input":"2024-05-29T05:40:24.081506Z","iopub.status.idle":"2024-05-29T06:09:13.713091Z","shell.execute_reply.started":"2024-05-29T05:40:24.081459Z","shell.execute_reply":"2024-05-29T06:09:13.711165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_and_preprocess_test_data(filepath, chunksize=10000):\n    chunks = []\n    for chunk in pd.read_csv(filepath, chunksize=chunksize):\n        chunk['buildingblock1_smiles'] = chunk['buildingblock1_smiles'].map(bb1_dict).fillna(-1).astype('int16')\n        chunk['buildingblock2_smiles'] = chunk['buildingblock2_smiles'].map(bb2_dict).fillna(-1).astype('int16')\n        chunk['buildingblock3_smiles'] = chunk['buildingblock3_smiles'].map(bb3_dict).fillna(-1).astype('int16')\n        chunks.append(chunk)\n    return pd.concat(chunks, ignore_index=True)\n\ndf_test_transformed = load_and_preprocess_test_data('/kaggle/input/leash-BELKA/test.csv')\n\n# Save the transformed test data\ndf_test_transformed.to_csv('transformed_test.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-05-29T06:09:13.716497Z","iopub.execute_input":"2024-05-29T06:09:13.717191Z","iopub.status.idle":"2024-05-29T06:09:34.254022Z","shell.execute_reply.started":"2024-05-29T06:09:13.717145Z","shell.execute_reply":"2024-05-29T06:09:34.252774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Columns\n### Before Transformation:\n\n* id: Unique identifier for each row.\n* buildingblock1_smiles: SMILES string for the first building block.\n* buildingblock2_smiles: SMILES string for the second building block.\n* buildingblock3_smiles: SMILES string for the third building block.\n* molecule_smiles: SMILES string for the molecule.\n* protein_name: Name of the protein.\n* binds: Binary indicator (0 or 1) indicating if the molecule binds to the protein.\n\n### After Transformation:\n\n* buildingblock1_smiles: Integer index for the first building block (after mapping).\n* buildingblock2_smiles: Integer index for the second building block (after mapping).\n* buildingblock3_smiles: Integer index for the third building block (after mapping).\n* molecule_smiles: SMILES string for the molecule.\n* binds_BRD4: Binary indicator (0 or 1) for binding to BRD4 protein.\n* binds_HSA: Binary indicator (0 or 1) for binding to HSA protein.\n* binds_sEH: Binary indicator (0 or 1) for binding to sEH protein.","metadata":{}},{"cell_type":"markdown","source":"## Transformation Example (conceptual)\nThe transform_labels function consolidates the repeated rows for each protein into a single row with separate columns for each binding indicator.\n\n### Before Transformation:","metadata":{}},{"cell_type":"markdown","source":"| id   | buildingblock1_smiles | buildingblock2_smiles | buildingblock3_smiles | molecule_smiles       | protein_name | binds |\n|------|-----------------------|-----------------------|-----------------------|-----------------------|--------------|-------|\n| 1    | CCCCCC                | O=C=O                 | CCN(CC)CC             | CCCCCCOCC             | BRD4         | 1     |\n| 2    | CCCCCC                | O=C=O                 | CCN(CC)CC             | CCCCCCOCC             | HSA          | 0     |\n| 3    | CCCCCC                | O=C=O                 | CCN(CC)CC             | CCCCCCOCC             | sEH          | 1     |\n| ...  | ...                   | ...                   | ...                   | ...                   | ...          | ...   |\n| 4    | CCN(CC)CC             | O=C=O                 | CCCCCC                | CCN(CC)CCO=C=O        | BRD4         | 0     |\n| 5    | CCN(CC)CC             | O=C=O                 | CCCCCC                | CCN(CC)CCO=C=O        | HSA          | 1     |\n| 6    | CCN(CC)CC             | O=C=O                 | CCCCCC                | CCN(CC)CCO=C=O        | sEH          | 0     |","metadata":{}},{"cell_type":"markdown","source":"### After Transformation:\n| buildingblock1_smiles | buildingblock2_smiles | buildingblock3_smiles | molecule_smiles       | binds_BRD4 | binds_HSA | binds_sEH |\n|-----------------------|-----------------------|-----------------------|-----------------------|------------|-----------|-----------|\n| CCCCCC                | O=C=O                 | CCN(CC)CC             | CCCCCCOCC             | 1          | 0         | 1         |\n| CCN(CC)CC             | O=C=O                 | CCCCCC                | CCN(CC)CCO=C=O        | 0          | 1         | 0         |","metadata":{}},{"cell_type":"markdown","source":"The transformation consolidates the repeated rows for each protein into a single row with separate columns for each binding indicator.\n\n## Saving the Preprocessed Dataset\nAfter preprocessing the data, save it to disk as a Parquet file for efficient storage and retrieval.","metadata":{}},{"cell_type":"code","source":"df_train.to_parquet('train_shrunken.parquet', index=False)\ndf_train.to_csv('train_shrunken.csv', index=False)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-29T06:09:34.255625Z","iopub.execute_input":"2024-05-29T06:09:34.256069Z","iopub.status.idle":"2024-05-29T06:21:57.837479Z","shell.execute_reply.started":"2024-05-29T06:09:34.256037Z","shell.execute_reply":"2024-05-29T06:21:57.835796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test_transformed.to_parquet('test_shrunken.parquet', index=False)\ndf_test_transformed.to_csv('test_shrunken.csv', index=False)\ndf_test_transformed.head()","metadata":{"execution":{"iopub.status.busy":"2024-05-29T06:28:52.714871Z","iopub.execute_input":"2024-05-29T06:28:52.715340Z","iopub.status.idle":"2024-05-29T06:29:06.105396Z","shell.execute_reply.started":"2024-05-29T06:28:52.715305Z","shell.execute_reply":"2024-05-29T06:29:06.103955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Summary\nThis notebook outlines the process of creating a shrunken dataset from the BELKA competition train set. The transformation includes mapping building block SMILES strings to integer indices and restructuring the data to consolidate repeated rows into a single row with separate columns for each protein binding indicator. This work was performed using ChatGPT-4. The transformed dataset can now be used for further analysis and modeling.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}