{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook's purpose is to shrink the size of the dataset for the [BELKA competition](https://www.kaggle.com/competitions/leash-BELKA/discussion?sort=published) .   \nShrinking strategy:  \n1. No ID column.  \n2. binds columns saved in bytes.  \n3. buildingblock1_smiles/buildingblock2_smiles/buildingblock3_smiles columns saved as int16, with encoded indices of the building blobks. I saved the building blocks and their indices in separate dictionaries.  \n4. I transformed the protein/label columns into three columns of labels per protein, shrinking the dataset length by three. (The other columns have identical values for each three consecutive rows).\n\nNOTE: TPU is not intended for EDA and data manipulation. Using TPU notebooks for the RAM capacity is considered a misuse of the TPU resource by Kaggle rules. Also, as an avid user of TPU, it is in my interest that people don't misuse it. In creating this notebook, I tried to make minimal use of TPU, developing and debugging on a regular notebook, and I published this notebook only because I feel that on this specific occasion, it is in the community's best interest to get a normal-size dataset instead of the given bloated one. Please don't fork/rerun this notebook on TPU, and please don't make similar use of the TPU notebook resource as I did here. I did it once, for the community, so that we all have a dataset that we can work with it. Thank you.  \n\nIf I see too many forks, I will turn this notebook to private (the dataset would still be public so don't worry). I hope I don't need to do this, so please don't fork. ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom pyarrow.parquet import ParquetFile\nimport pickle\nimport os","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-06T06:47:00.036683Z","iopub.execute_input":"2024-04-06T06:47:00.037786Z","iopub.status.idle":"2024-04-06T06:47:01.250496Z","shell.execute_reply.started":"2024-04-06T06:47:00.037746Z","shell.execute_reply":"2024-04-06T06:47:01.249308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ParquetFile('/kaggle/input/leash-BELKA/train.parquet').metadata","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:01.252291Z","iopub.execute_input":"2024-04-06T06:47:01.252748Z","iopub.status.idle":"2024-04-06T06:47:01.289186Z","shell.execute_reply.started":"2024-04-06T06:47:01.252720Z","shell.execute_reply":"2024-04-06T06:47:01.287967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ParquetFile('/kaggle/input/leash-BELKA/train.parquet').schema","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:01.291124Z","iopub.execute_input":"2024-04-06T06:47:01.291806Z","iopub.status.idle":"2024-04-06T06:47:01.308071Z","shell.execute_reply.started":"2024-04-06T06:47:01.291763Z","shell.execute_reply":"2024-04-06T06:47:01.306424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DEBUG = False\nif DEBUG:\n    NUM_ROWS = 30000000\nelse:\n    NUM_ROWS = 295246830","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:01.311609Z","iopub.execute_input":"2024-04-06T06:47:01.311943Z","iopub.status.idle":"2024-04-06T06:47:01.317432Z","shell.execute_reply.started":"2024-04-06T06:47:01.311918Z","shell.execute_reply":"2024-04-06T06:47:01.316228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_path = '/kaggle/input/leash-BELKA/train.parquet'","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:01.319857Z","iopub.execute_input":"2024-04-06T06:47:01.320225Z","iopub.status.idle":"2024-04-06T06:47:01.330792Z","shell.execute_reply.started":"2024-04-06T06:47:01.320182Z","shell.execute_reply":"2024-04-06T06:47:01.329646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### A quick verifications that id column is what we expect it to be","metadata":{}},{"cell_type":"code","source":"def id_eda(dataset_path):\n    id_arr = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=['id']).id.to_numpy()\n    id_arr_2 = range(295246830)\n    print(np.mean(id_arr == id_arr_2))\nid_eda(dataset_path)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:01.332020Z","iopub.execute_input":"2024-04-06T06:47:01.332630Z","iopub.status.idle":"2024-04-06T06:47:49.395218Z","shell.execute_reply.started":"2024-04-06T06:47:01.332598Z","shell.execute_reply":"2024-04-06T06:47:49.394303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The dataset consists of three rows of the same small molecule with binding labels to the three different proteins, followed by three rows of the next small molecule, etc. We will verify it for each relevant column along the way.","metadata":{}},{"cell_type":"code","source":"def protein_name_eda(dataset_path):\n    protein_name =  pd.read_parquet(dataset_path, engine = 'pyarrow', columns=['protein_name']).protein_name.to_numpy()\n    protein_name_reshaped = np.reshape(protein_name, [-1, 3])\n    print(np.mean(protein_name_reshaped[:, 0] == 'BRD4'))\n    print(np.mean(protein_name_reshaped[:, 1] == 'HSA'))\n    print(np.mean(protein_name_reshaped[:, 2] == 'sEH'))\n    \nprotein_name_eda(dataset_path)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:47:49.396517Z","iopub.execute_input":"2024-04-06T06:47:49.396820Z","iopub.status.idle":"2024-04-06T06:48:10.914995Z","shell.execute_reply.started":"2024-04-06T06:47:49.396795Z","shell.execute_reply":"2024-04-06T06:48:10.914173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_binds(dataset_path):\n    binds =  pd.read_parquet(dataset_path, engine = 'pyarrow', columns=['binds']).binds.to_numpy()\n    binds = binds[:NUM_ROWS]\n    return np.reshape(binds.astype('byte'), [-1, 3])\n\nbinds = get_binds(dataset_path)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:48:10.915954Z","iopub.execute_input":"2024-04-06T06:48:10.916283Z","iopub.status.idle":"2024-04-06T06:48:14.166781Z","shell.execute_reply.started":"2024-04-06T06:48:10.916253Z","shell.execute_reply":"2024-04-06T06:48:14.165608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef get_unique_BB(dataset_path, col):\n    BBs = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=[col])\n    BBs = BBs[:NUM_ROWS]\n    BBs = BBs.to_numpy()[:, 0]\n    BBs_reshaped = np.reshape(BBs, [-1, 3])\n    \n    if np.mean(BBs_reshaped[:, 0] == BBs_reshaped[:, 1]) != 1:\n        print('ERROR')\n    if np.mean(BBs_reshaped[:, 0] == BBs_reshaped[:, 2]) != 1:\n        print('ERROR')\n    \n    BBs_unique = np.unique(BBs_reshaped[:, 0])\n    BBs_unique = list(BBs_unique)\n    BBs_dict = {BBs_unique[i]:i for i in range(len(BBs_unique))}\n    BBs_dict_reverse = {i:BBs_unique[i] for i in range(len(BBs_unique))}\n    return BBs_dict, BBs_dict_reverse\n\nBBs_dict_1, BBs_dict_reverse_1 = get_unique_BB(dataset_path, 'buildingblock1_smiles')\nprint(len(BBs_dict_1))\nBBs_dict_2, BBs_dict_reverse_2 = get_unique_BB(dataset_path, 'buildingblock2_smiles')\nprint(len(BBs_dict_2))\nBBs_dict_3, BBs_dict_reverse_3 = get_unique_BB(dataset_path, 'buildingblock3_smiles')\nprint(len(BBs_dict_3))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:48:14.168193Z","iopub.execute_input":"2024-04-06T06:48:14.168603Z","iopub.status.idle":"2024-04-06T06:49:43.963425Z","shell.execute_reply.started":"2024-04-06T06:48:14.168569Z","shell.execute_reply":"2024-04-06T06:49:43.962284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef get_encoded(dataset_path, col, BBs_dict):\n    BBs = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=[col])\n    BBs = BBs[:NUM_ROWS]\n    BBs = BBs[col].to_numpy()\n    BBs_reshaped = np.reshape(BBs, [-1, 3])\n    BBs = BBs_reshaped[:, 0]\n    encoded_BBs = [BBs_dict[x] for x in BBs]\n    encoded_BBs = np.asarray(encoded_BBs, dtype = np.int16)\n    return encoded_BBs\n\nencoded_BBs_1 = get_encoded(dataset_path, 'buildingblock1_smiles', BBs_dict_1)\nencoded_BBs_2 = get_encoded(dataset_path, 'buildingblock2_smiles', BBs_dict_2)\nencoded_BBs_3 = get_encoded(dataset_path, 'buildingblock3_smiles', BBs_dict_3)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:49:43.964587Z","iopub.execute_input":"2024-04-06T06:49:43.964866Z","iopub.status.idle":"2024-04-06T06:50:49.345539Z","shell.execute_reply.started":"2024-04-06T06:49:43.964843Z","shell.execute_reply":"2024-04-06T06:50:49.344358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_molecule_smiles(dataset_path):\n    if DEBUG:\n        molecule_smiles = pd.read_csv(f'{dataset_path[:-7]}csv', usecols=['molecule_smiles'], nrows = NUM_ROWS)\n    else:\n        molecule_smiles = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=['molecule_smiles'])\n    molecule_smiles = molecule_smiles.molecule_smiles.to_numpy()\n    molecule_smiles = np.reshape(molecule_smiles, [-1, 3])\n    if np.mean(molecule_smiles[:, 0] == molecule_smiles[:, 1]) != 1:\n        print('ERROR')\n    if np.mean(molecule_smiles[:, 0] == molecule_smiles[:, 2]) != 1:\n        print('ERROR')\n    return molecule_smiles[:, 0]\n\nmolecule_smiles = get_molecule_smiles(dataset_path)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:50:49.346965Z","iopub.execute_input":"2024-04-06T06:50:49.347279Z","iopub.status.idle":"2024-04-06T06:52:30.655621Z","shell.execute_reply.started":"2024-04-06T06:50:49.347252Z","shell.execute_reply":"2024-04-06T06:52:30.654577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/leash-BELKA/train.csv', nrows = 2)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:52:30.656779Z","iopub.execute_input":"2024-04-06T06:52:30.657323Z","iopub.status.idle":"2024-04-06T06:52:30.679507Z","shell.execute_reply.started":"2024-04-06T06:52:30.657292Z","shell.execute_reply":"2024-04-06T06:52:30.678370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = {'buildingblock1_smiles':encoded_BBs_1, 'buildingblock2_smiles':encoded_BBs_2, 'buildingblock3_smiles':encoded_BBs_3,\n        'molecule_smiles':molecule_smiles, 'binds_BRD4':binds[:, 0], 'binds_HSA':binds[:, 1], 'binds_sEH':binds[:, 2]}\ndf = pd.DataFrame(data=data)\ndf.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:52:30.680743Z","iopub.execute_input":"2024-04-06T06:52:30.681058Z","iopub.status.idle":"2024-04-06T06:52:30.928335Z","shell.execute_reply.started":"2024-04-06T06:52:30.681032Z","shell.execute_reply":"2024-04-06T06:52:30.927285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_parquet('train.parquet', index = False)\ndf.to_csv('train.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:52:30.932354Z","iopub.execute_input":"2024-04-06T06:52:30.932684Z","iopub.status.idle":"2024-04-06T06:53:12.947367Z","shell.execute_reply.started":"2024-04-06T06:52:30.932657Z","shell.execute_reply":"2024-04-06T06:53:12.945881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    os.mkdir('train_dicts')\nexcept:\n    print('Folder exist')\n    \npickle.dump(BBs_dict_1, open('train_dicts/BBs_dict_1.p', 'bw'))\npickle.dump(BBs_dict_2, open('train_dicts/BBs_dict_2.p', 'bw'))\npickle.dump(BBs_dict_3, open('train_dicts/BBs_dict_3.p', 'bw'))\npickle.dump(BBs_dict_reverse_1, open('train_dicts/BBs_dict_reverse_1.p', 'bw'))\npickle.dump(BBs_dict_reverse_2, open('train_dicts/BBs_dict_reverse_2.p', 'bw'))\npickle.dump(BBs_dict_reverse_3, open('train_dicts/BBs_dict_reverse_3.p', 'bw'))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:12.948881Z","iopub.execute_input":"2024-04-06T06:53:12.950158Z","iopub.status.idle":"2024-04-06T06:53:12.959688Z","shell.execute_reply.started":"2024-04-06T06:53:12.950118Z","shell.execute_reply":"2024-04-06T06:53:12.958702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# For the test set","metadata":{}},{"cell_type":"code","source":"test_path = '/kaggle/input/leash-BELKA/test.parquet'\n\ndf = pd.read_csv('/kaggle/input/leash-BELKA/test.csv', nrows = 2)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:12.961001Z","iopub.execute_input":"2024-04-06T06:53:12.961430Z","iopub.status.idle":"2024-04-06T06:53:13.976815Z","shell.execute_reply.started":"2024-04-06T06:53:12.961393Z","shell.execute_reply":"2024-04-06T06:53:13.976026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ParquetFile('/kaggle/input/leash-BELKA/test.parquet').metadata","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:13.978194Z","iopub.execute_input":"2024-04-06T06:53:13.978764Z","iopub.status.idle":"2024-04-06T06:53:14.756921Z","shell.execute_reply.started":"2024-04-06T06:53:13.978735Z","shell.execute_reply":"2024-04-06T06:53:14.755917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def id_eda_test(dataset_path):\n    id_arr = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=['id']).id.to_numpy()\n    id_arr_2 = range(295246830, 295246830+1674896)\n    print(np.mean(id_arr == id_arr_2))\nid_eda_test(test_path)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:14.758046Z","iopub.execute_input":"2024-04-06T06:53:14.758375Z","iopub.status.idle":"2024-04-06T06:53:15.516892Z","shell.execute_reply.started":"2024-04-06T06:53:14.758348Z","shell.execute_reply":"2024-04-06T06:53:15.515693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The length of the test set is not dividable by 3. So, for some small molecules, we need to predict only one or two proteins.","metadata":{}},{"cell_type":"code","source":"molecule_smiles = pd.read_parquet(test_path, engine = 'pyarrow', columns=['molecule_smiles']).molecule_smiles.to_numpy()\nprotein_name = pd.read_parquet(test_path, engine = 'pyarrow', columns=['protein_name']).protein_name.to_numpy()\nfirst_unique_molecule_smiles_indices = []\nmolecule_smiles_unique = {}\nis_BRD4 = {}\nis_HSA = {}\nis_sEH = {}\nfor i,x in enumerate(molecule_smiles):\n    if x not in molecule_smiles_unique:\n        molecule_smiles_unique[x] = [i]\n        first_unique_molecule_smiles_indices.append(i)\n        is_BRD4[x] = False\n        is_HSA[x] = False\n        is_sEH[x] = False\n        if protein_name[i] == 'BRD4':\n            is_BRD4[x] = True\n        if protein_name[i] == 'HSA':\n            is_HSA[x] = True\n        if protein_name[i] == 'sEH':\n            is_sEH[x] = True\n    else:\n        molecule_smiles_unique[x].append(i)\n        if protein_name[i] == 'BRD4':\n            is_BRD4[x] = True\n        if protein_name[i] == 'HSA':\n            is_HSA[x] = True\n        if protein_name[i] == 'sEH':\n            is_sEH[x] = True\nfirst_unique_molecule_smiles_indices = np.asarray(first_unique_molecule_smiles_indices)\nprint(len(is_BRD4))\nprint(np.sum([is_BRD4[x] for x in is_BRD4]))\nprint(np.sum([is_HSA[x] for x in is_HSA]))\nprint(np.sum([is_sEH[x] for x in is_sEH]))\n\nmolecule_smiles_unique_arr = molecule_smiles[first_unique_molecule_smiles_indices]\nprint(len(np.unique(molecule_smiles_unique_arr)) == len(molecule_smiles_unique_arr))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:15.518291Z","iopub.execute_input":"2024-04-06T06:53:15.518717Z","iopub.status.idle":"2024-04-06T06:53:22.313039Z","shell.execute_reply.started":"2024-04-06T06:53:15.518688Z","shell.execute_reply":"2024-04-06T06:53:22.311359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"is_BRD4_arr = np.asarray([is_BRD4[x] for x in molecule_smiles_unique])\nis_HSA_arr = np.asarray([is_HSA[x] for x in molecule_smiles_unique])\nis_sEH_arr = np.asarray([is_sEH[x] for x in molecule_smiles_unique])\n\nprint(np.sum(is_BRD4_arr))\nprint(np.sum(is_HSA_arr))\nprint(np.sum(is_sEH_arr))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:22.315142Z","iopub.execute_input":"2024-04-06T06:53:22.315694Z","iopub.status.idle":"2024-04-06T06:53:23.329951Z","shell.execute_reply.started":"2024-04-06T06:53:22.315652Z","shell.execute_reply":"2024-04-06T06:53:23.328869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_unique_BB_test(dataset_path, col):\n    BBs = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=[col])\n    BBs = BBs[col].to_numpy()\n    BBs_unique = np.unique(BBs)\n    BBs_unique = list(BBs_unique)\n    BBs_dict = {BBs_unique[i]:i for i in range(len(BBs_unique))}\n    BBs_dict_reverse = {i:BBs_unique[i] for i in range(len(BBs_unique))}\n    return BBs_dict, BBs_dict_reverse\n\nBBs_dict_1_test, BBs_dict_reverse_1_test = get_unique_BB_test(test_path, 'buildingblock1_smiles')\nprint(len(BBs_dict_1_test))\nBBs_dict_2_test, BBs_dict_reverse_2_test = get_unique_BB_test(test_path, 'buildingblock2_smiles')\nprint(len(BBs_dict_2_test))\nBBs_dict_3_test, BBs_dict_reverse_3_test = get_unique_BB_test(test_path, 'buildingblock3_smiles')\nprint(len(BBs_dict_3_test))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:23.332137Z","iopub.execute_input":"2024-04-06T06:53:23.332593Z","iopub.status.idle":"2024-04-06T06:53:27.575920Z","shell.execute_reply.started":"2024-04-06T06:53:23.332553Z","shell.execute_reply":"2024-04-06T06:53:27.574815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_encoded_test(dataset_path, col, BBs_dict):\n    BBs = pd.read_parquet(dataset_path, engine = 'pyarrow', columns=[col])\n    BBs = BBs[col].to_numpy()\n    BBs = BBs[first_unique_molecule_smiles_indices]\n    encoded_BBs = [BBs_dict[x] for x in BBs]\n    encoded_BBs = np.asarray(encoded_BBs, dtype = np.int16)\n    return encoded_BBs\n\nencoded_BBs_1_test = get_encoded_test(test_path, 'buildingblock1_smiles', BBs_dict_1_test)\nencoded_BBs_2_test = get_encoded_test(test_path, 'buildingblock2_smiles', BBs_dict_2_test)\nencoded_BBs_3_test = get_encoded_test(test_path, 'buildingblock3_smiles', BBs_dict_3_test)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:27.577449Z","iopub.execute_input":"2024-04-06T06:53:27.577765Z","iopub.status.idle":"2024-04-06T06:53:28.365910Z","shell.execute_reply.started":"2024-04-06T06:53:27.577739Z","shell.execute_reply":"2024-04-06T06:53:28.364803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = {'buildingblock1_smiles':encoded_BBs_1_test, 'buildingblock2_smiles':encoded_BBs_2_test,\n        'buildingblock3_smiles':encoded_BBs_3_test,'molecule_smiles':molecule_smiles_unique_arr,\n        'is_BRD4':is_BRD4_arr, 'is_HSA':is_HSA_arr, 'is_sEH':is_sEH_arr}\ndf = pd.DataFrame(data=data)\ndf.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:28.367177Z","iopub.execute_input":"2024-04-06T06:53:28.367491Z","iopub.status.idle":"2024-04-06T06:53:28.564907Z","shell.execute_reply.started":"2024-04-06T06:53:28.367465Z","shell.execute_reply":"2024-04-06T06:53:28.563821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_parquet('test.parquet', index = False)\ndf.to_csv('test.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:28.566318Z","iopub.execute_input":"2024-04-06T06:53:28.567184Z","iopub.status.idle":"2024-04-06T06:53:32.523851Z","shell.execute_reply.started":"2024-04-06T06:53:28.567145Z","shell.execute_reply":"2024-04-06T06:53:32.522562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    os.mkdir('test_dicts')\nexcept:\n    print('Folder exist')\n    \npickle.dump(BBs_dict_1_test, open('test_dicts/BBs_dict_1_test.p', 'bw'))\npickle.dump(BBs_dict_2_test, open('test_dicts/BBs_dict_2_test.p', 'bw'))\npickle.dump(BBs_dict_3_test, open('test_dicts/BBs_dict_3_test.p', 'bw'))\npickle.dump(BBs_dict_reverse_1_test, open('test_dicts/BBs_dict_reverse_1_test.p', 'bw'))\npickle.dump(BBs_dict_reverse_2_test, open('test_dicts/BBs_dict_reverse_2_test.p', 'bw'))\npickle.dump(BBs_dict_reverse_3_test, open('test_dicts/BBs_dict_reverse_3_test.p', 'bw'))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:32.525192Z","iopub.execute_input":"2024-04-06T06:53:32.525533Z","iopub.status.idle":"2024-04-06T06:53:32.536070Z","shell.execute_reply.started":"2024-04-06T06:53:32.525508Z","shell.execute_reply":"2024-04-06T06:53:32.535007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pickle.dump(molecule_smiles_unique, open('test_dicts/molecule_smiles_unique.p', 'bw'))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T06:53:32.537581Z","iopub.execute_input":"2024-04-06T06:53:32.537902Z","iopub.status.idle":"2024-04-06T06:53:33.316921Z","shell.execute_reply.started":"2024-04-06T06:53:32.537877Z","shell.execute_reply":"2024-04-06T06:53:33.315942Z"},"trusted":true},"execution_count":null,"outputs":[]}]}