{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"🧬Open Problems – Single-Cell Perturbations\nPredict how small molecules change gene expression in different cell types\n\n\nFeature Augmentation & Fragments of SMILES\n\nTwo available features are \"cell_type\" and \"sm_name,\" and we want to add two new columns (two new features) to them.\n\nIf we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, we will see that for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.\n\nAlso, if we separate the cells based on 'sm_name,' we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.\n\nObviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1]).\"\n\nBy adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')\n#:::::::::::::::::::::::::::::::::::\nimport os\nimport gc\nimport glob\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n#:::::::::::::::::::::::::::::::::::\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*\n\nimport torch\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f'Using{device}')","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2025-01-18T06:49:30.729535Z","iopub.execute_input":"2025-01-18T06:49:30.729759Z","iopub.status.idle":"2025-01-18T06:49:36.941409Z","shell.execute_reply.started":"2025-01-18T06:49:30.729737Z","shell.execute_reply":"2025-01-18T06:49:36.940523Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n<div>\n    <h1 align=\"center\" style=\"color:gray;\">Competition Data (Eight files)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"(1) de_train.parquet","metadata":{}},{"cell_type":"code","source":"de_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet')\nde_train.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:36.943885Z","iopub.execute_input":"2025-01-18T06:49:36.944663Z","iopub.status.idle":"2025-01-18T06:49:39.328443Z","shell.execute_reply.started":"2025-01-18T06:49:36.944634Z","shell.execute_reply":"2025-01-18T06:49:39.327524Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"(7) id_map.csv","metadata":{}},{"cell_type":"code","source":"id_map = pd.read_csv('../input/open-problems-single-cell-perturbations/id_map.csv')\nid_map.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:39.329332Z","iopub.execute_input":"2025-01-18T06:49:39.329594Z","iopub.status.idle":"2025-01-18T06:49:39.344291Z","shell.execute_reply.started":"2025-01-18T06:49:39.329573Z","shell.execute_reply":"2025-01-18T06:49:39.343329Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"(8) sample_submission.csv","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('../input/open-problems-single-cell-perturbations/sample_submission.csv', index_col='id')\nsample_submission.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:39.345246Z","iopub.execute_input":"2025-01-18T06:49:39.345489Z","iopub.status.idle":"2025-01-18T06:49:41.754078Z","shell.execute_reply.started":"2025-01-18T06:49:39.345468Z","shell.execute_reply":"2025-01-18T06:49:41.753207Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"train | test | target","metadata":{}},{"cell_type":"code","source":"xlist  = ['cell_type','sm_name']\n_ylist = ['cell_type','sm_name','sm_lincs_id','SMILES','control']\n\ny = de_train.drop(columns=_ylist)\ny.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.7551Z","iopub.execute_input":"2025-01-18T06:49:41.755348Z","iopub.status.idle":"2025-01-18T06:49:41.794473Z","shell.execute_reply.started":"2025-01-18T06:49:41.755327Z","shell.execute_reply":"2025-01-18T06:49:41.793448Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"get_dummies (OneHotEncoder)","metadata":{}},{"cell_type":"code","source":"train = pd.get_dummies(de_train[xlist], columns=xlist)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.795895Z","iopub.execute_input":"2025-01-18T06:49:41.79618Z","iopub.status.idle":"2025-01-18T06:49:41.811816Z","shell.execute_reply.started":"2025-01-18T06:49:41.796157Z","shell.execute_reply":"2025-01-18T06:49:41.811031Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test1 = pd.get_dummies(id_map[xlist], columns=xlist)\ntest1.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.814386Z","iopub.execute_input":"2025-01-18T06:49:41.814613Z","iopub.status.idle":"2025-01-18T06:49:41.823563Z","shell.execute_reply.started":"2025-01-18T06:49:41.814593Z","shell.execute_reply":"2025-01-18T06:49:41.82291Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Uncommon deleted","metadata":{}},{"cell_type":"code","source":"uncommon = [f for f in train if f not in test1]\nlen(uncommon)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.824462Z","iopub.execute_input":"2025-01-18T06:49:41.824682Z","iopub.status.idle":"2025-01-18T06:49:41.834849Z","shell.execute_reply.started":"2025-01-18T06:49:41.824663Z","shell.execute_reply":"2025-01-18T06:49:41.834027Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X1 = train.drop(columns=uncommon)\nX1.shape[1], test1.shape[1]","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.835578Z","iopub.execute_input":"2025-01-18T06:49:41.835794Z","iopub.status.idle":"2025-01-18T06:49:41.847623Z","shell.execute_reply.started":"2025-01-18T06:49:41.835775Z","shell.execute_reply":"2025-01-18T06:49:41.846772Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"list(X1.columns) == list(test1.columns)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.848485Z","iopub.execute_input":"2025-01-18T06:49:41.849399Z","iopub.status.idle":"2025-01-18T06:49:41.856784Z","shell.execute_reply.started":"2025-01-18T06:49:41.849375Z","shell.execute_reply":"2025-01-18T06:49:41.856094Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Evaluation\n\nMean Rowwise Root Mean Squared Error (MRRMSE)","metadata":{}},{"cell_type":"code","source":"def mrrmse_pd(y_pred: pd.DataFrame, y_true: pd.DataFrame):\n    \n    return ((y_pred - y_true)**2).mean(axis=1).apply(np.sqrt).mean()","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.857763Z","iopub.execute_input":"2025-01-18T06:49:41.858312Z","iopub.status.idle":"2025-01-18T06:49:41.866357Z","shell.execute_reply.started":"2025-01-18T06:49:41.858289Z","shell.execute_reply":"2025-01-18T06:49:41.865609Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def mrrmse_np(y_pred, y_true):\n    \n    return np.sqrt(np.square(y_true - y_pred).mean(axis=1)).mean()\n","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.867409Z","iopub.execute_input":"2025-01-18T06:49:41.867892Z","iopub.status.idle":"2025-01-18T06:49:41.875415Z","shell.execute_reply.started":"2025-01-18T06:49:41.867862Z","shell.execute_reply":"2025-01-18T06:49:41.874555Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Feature Augmentation","metadata":{}},{"cell_type":"code","source":"de_cell_type = de_train.iloc[:, [0] + list(range(5, de_train.shape[1]))]\nde_sm_name = de_train.iloc[:, [1] + list(range(5, de_train.shape[1]))]\n\nde_cell_type.shape, de_sm_name.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.876471Z","iopub.execute_input":"2025-01-18T06:49:41.877221Z","iopub.status.idle":"2025-01-18T06:49:41.953955Z","shell.execute_reply.started":"2025-01-18T06:49:41.877189Z","shell.execute_reply":"2025-01-18T06:49:41.953082Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Calculate averages based on 'cell_type' and 'sm_name' for all columns\n\n","metadata":{}},{"cell_type":"code","source":"mean_cell_type = de_cell_type.groupby('cell_type').mean().reset_index()\nmean_sm_name = de_sm_name.groupby('sm_name').mean().reset_index()\n\ndisplay(mean_cell_type)\ndisplay(mean_sm_name)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:41.954962Z","iopub.execute_input":"2025-01-18T06:49:41.955251Z","iopub.status.idle":"2025-01-18T06:49:42.309605Z","shell.execute_reply.started":"2025-01-18T06:49:41.955227Z","shell.execute_reply":"2025-01-18T06:49:42.308793Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Paste the results into the train file","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in de_cell_type['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\ntr_cell_type = pd.concat(rows)\ntr_cell_type = tr_cell_type.reset_index(drop=True)\ntr_cell_type","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:42.310683Z","iopub.execute_input":"2025-01-18T06:49:42.311017Z","iopub.status.idle":"2025-01-18T06:49:43.242361Z","shell.execute_reply.started":"2025-01-18T06:49:42.310986Z","shell.execute_reply":"2025-01-18T06:49:43.241462Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rows = []\nfor name in de_sm_name['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\ntr_sm_name = pd.concat(rows)\ntr_sm_name = tr_sm_name.reset_index(drop=True)\ntr_sm_name","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:43.243353Z","iopub.execute_input":"2025-01-18T06:49:43.243592Z","iopub.status.idle":"2025-01-18T06:49:44.27546Z","shell.execute_reply.started":"2025-01-18T06:49:43.243571Z","shell.execute_reply":"2025-01-18T06:49:44.274565Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Paste the results into the test file","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in id_map['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\nte_cell_type = pd.concat(rows)\nte_cell_type = te_cell_type.reset_index(drop=True)\nte_cell_type","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:44.27641Z","iopub.execute_input":"2025-01-18T06:49:44.276648Z","iopub.status.idle":"2025-01-18T06:49:44.667081Z","shell.execute_reply.started":"2025-01-18T06:49:44.276627Z","shell.execute_reply":"2025-01-18T06:49:44.66621Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rows = []\nfor name in id_map['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\nte_sm_name = pd.concat(rows)\nte_sm_name = te_sm_name.reset_index(drop=True)\nte_sm_name","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:44.668005Z","iopub.execute_input":"2025-01-18T06:49:44.668266Z","iopub.status.idle":"2025-01-18T06:49:45.037465Z","shell.execute_reply.started":"2025-01-18T06:49:44.668244Z","shell.execute_reply":"2025-01-18T06:49:45.036643Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Sample: Column number zero - A1BG","metadata":{}},{"cell_type":"code","source":"y0 = y.iloc[:, 0].copy()\ny0","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.03865Z","iopub.execute_input":"2025-01-18T06:49:45.03891Z","iopub.status.idle":"2025-01-18T06:49:45.046398Z","shell.execute_reply.started":"2025-01-18T06:49:45.038887Z","shell.execute_reply":"2025-01-18T06:49:45.045477Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X0 = X1.join(tr_cell_type.iloc[:, 0+1]).copy()\nX0 = X0.join(tr_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\nX0","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.04772Z","iopub.execute_input":"2025-01-18T06:49:45.047931Z","iopub.status.idle":"2025-01-18T06:49:45.078855Z","shell.execute_reply.started":"2025-01-18T06:49:45.047912Z","shell.execute_reply":"2025-01-18T06:49:45.078109Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test0 = test1.join(te_cell_type.iloc[:, 0+1]).copy()\ntest0 = test0.join(te_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\ntest0","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.079755Z","iopub.execute_input":"2025-01-18T06:49:45.080436Z","iopub.status.idle":"2025-01-18T06:49:45.103991Z","shell.execute_reply.started":"2025-01-18T06:49:45.080412Z","shell.execute_reply":"2025-01-18T06:49:45.103275Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Correlation - Column #0</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"X0_corr = X0.copy()\nX0_corr['y0'] = y0\n\ncorr = X0_corr.iloc[: , 131:].corr(numeric_only=True).round(3)\ncorr.style.background_gradient(cmap='Pastel1')","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.108501Z","iopub.execute_input":"2025-01-18T06:49:45.108706Z","iopub.status.idle":"2025-01-18T06:49:45.154256Z","shell.execute_reply.started":"2025-01-18T06:49:45.108689Z","shell.execute_reply":"2025-01-18T06:49:45.153482Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cor_matrix = X0_corr.iloc[: , 131:].corr()\nfig = plt.figure(figsize=(6,6));\n\ncmap=sns.diverging_palette(240, 10, s=75, l=50, sep=1, n=6, center='light', as_cmap=False);\nsns.heatmap(cor_matrix, center=0, annot=True, cmap=cmap, linewidths=5);\nplt.suptitle('Train Set (Heatmap)', y=0.92, fontsize=16, c='darkred');\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.155192Z","iopub.execute_input":"2025-01-18T06:49:45.155496Z","iopub.status.idle":"2025-01-18T06:49:45.432599Z","shell.execute_reply.started":"2025-01-18T06:49:45.155468Z","shell.execute_reply":"2025-01-18T06:49:45.431706Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"LightGBM - Column #0","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\nfrom sklearn.svm import LinearSVR\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.multioutput import MultiOutputRegressor\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(X0, y0, test_size=0.20, random_state=421)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:45.433645Z","iopub.execute_input":"2025-01-18T06:49:45.433877Z","iopub.status.idle":"2025-01-18T06:49:48.007506Z","shell.execute_reply.started":"2025-01-18T06:49:45.433857Z","shell.execute_reply":"2025-01-18T06:49:48.00678Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = lgb.LGBMRegressor()\nmodel.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.008508Z","iopub.execute_input":"2025-01-18T06:49:48.008811Z","iopub.status.idle":"2025-01-18T06:49:48.076403Z","shell.execute_reply.started":"2025-01-18T06:49:48.008784Z","shell.execute_reply":"2025-01-18T06:49:48.075616Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"predict = model.predict(X_test) \n# mrrmse_pd(pd.DataFrame(predict), pd.DataFrame(y_test.values))","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.077475Z","iopub.execute_input":"2025-01-18T06:49:48.077774Z","iopub.status.idle":"2025-01-18T06:49:48.08418Z","shell.execute_reply.started":"2025-01-18T06:49:48.077751Z","shell.execute_reply":"2025-01-18T06:49:48.083378Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"N = 0\nplt.style.use('seaborn-whitegrid') \nplt.figure(figsize=(8, 4), facecolor='lightyellow')\nplt.title(f'Column:  #{N}', fontsize=12)\nplt.gca().set_facecolor('lightgray')\n\nsns.distplot(y_test.values-predict, bins=100, color='red')\nplt.legend(['y_true','y_pred'], loc=1)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.085169Z","iopub.execute_input":"2025-01-18T06:49:48.085417Z","iopub.status.idle":"2025-01-18T06:49:48.497259Z","shell.execute_reply.started":"2025-01-18T06:49:48.085397Z","shell.execute_reply":"2025-01-18T06:49:48.49638Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Feature Augmentation - Based on : 'SMILES'","metadata":{}},{"cell_type":"code","source":"sm_name_smiles_dict = de_train.set_index('sm_name')['SMILES'].to_dict()\n\nid_map['SMILES'] = id_map['sm_name'].map(sm_name_smiles_dict)\nid_map['SMILES']","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.498391Z","iopub.execute_input":"2025-01-18T06:49:48.498651Z","iopub.status.idle":"2025-01-18T06:49:48.545941Z","shell.execute_reply.started":"2025-01-18T06:49:48.498629Z","shell.execute_reply":"2025-01-18T06:49:48.545113Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def split_sign(text):\n    text = text.replace(')(', ' ')\n    text = text.replace('(' , ' ')\n    text = text.replace(')' , ' ')\n    return text.split(\" \")","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.546804Z","iopub.execute_input":"2025-01-18T06:49:48.547087Z","iopub.status.idle":"2025-01-18T06:49:48.552067Z","shell.execute_reply.started":"2025-01-18T06:49:48.54706Z","shell.execute_reply":"2025-01-18T06:49:48.551202Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"An example for the method of splitting each 'SMILES'","metadata":{}},{"cell_type":"code","source":"pip install rdkit","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:48.5531Z","iopub.execute_input":"2025-01-18T06:49:48.553776Z","iopub.status.idle":"2025-01-18T06:49:59.837232Z","shell.execute_reply.started":"2025-01-18T06:49:48.553746Z","shell.execute_reply":"2025-01-18T06:49:59.836181Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem import AllChem","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:49:59.838968Z","iopub.execute_input":"2025-01-18T06:49:59.839382Z","iopub.status.idle":"2025-01-18T06:50:00.253179Z","shell.execute_reply.started":"2025-01-18T06:49:59.839332Z","shell.execute_reply":"2025-01-18T06:50:00.252239Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import re\n# Thanks to: https://github.com/DocMinus\n\ndef element_count(data, N):\n    smiles = data['SMILES'].iloc[N]\n\n    pattern = \"Si|Ti|Al|Zn|Pd|Pt|Br?|Cl?|N|O|S|P|F|I|B|b|c|n|o|s|p\" \n    regex = re.compile(pattern)\n    elements = [token for token in regex.findall(smiles)]\n    ele_length = len(elements)\n    lowercase_pattern = \"b|c|n|o|s|p\"\n    regex_low = re.compile(lowercase_pattern)\n    for i in range(ele_length):\n        if regex_low.findall(elements[i]):\n            elements[i] = elements[i].upper()\n    element_count = [elements.count(ele) for ele in elements]\n    formula = dict(zip(elements, element_count)) \n\n    print('==> Element_Count: ', str(formula))","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.254445Z","iopub.execute_input":"2025-01-18T06:50:00.255105Z","iopub.status.idle":"2025-01-18T06:50:00.261306Z","shell.execute_reply.started":"2025-01-18T06:50:00.255072Z","shell.execute_reply":"2025-01-18T06:50:00.260537Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cols_sel = ['cell_type','sm_name','sm_lincs_id','SMILES']\n\nsmiles = de_train['SMILES'][0]\nmol = Chem.MolFromSmiles(smiles)\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)    \ndisplay(pd.DataFrame(de_train[cols_sel].iloc[0]))\n\nelement_count(de_train, 0)\nimg","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.262468Z","iopub.execute_input":"2025-01-18T06:50:00.262782Z","iopub.status.idle":"2025-01-18T06:50:00.320325Z","shell.execute_reply.started":"2025-01-18T06:50:00.262753Z","shell.execute_reply":"2025-01-18T06:50:00.319551Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"smiles0 = de_train['SMILES'][0]\nsmiles0","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.321264Z","iopub.execute_input":"2025-01-18T06:50:00.321469Z","iopub.status.idle":"2025-01-18T06:50:00.326551Z","shell.execute_reply.started":"2025-01-18T06:50:00.321451Z","shell.execute_reply":"2025-01-18T06:50:00.325744Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"split0 = split_sign(smiles0)\nsplit0","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.327482Z","iopub.execute_input":"2025-01-18T06:50:00.327699Z","iopub.status.idle":"2025-01-18T06:50:00.33843Z","shell.execute_reply.started":"2025-01-18T06:50:00.327679Z","shell.execute_reply":"2025-01-18T06:50:00.337685Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mol = Chem.MolFromSmiles(split0[0])\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)   \nimg","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.339328Z","iopub.execute_input":"2025-01-18T06:50:00.339704Z","iopub.status.idle":"2025-01-18T06:50:00.369053Z","shell.execute_reply.started":"2025-01-18T06:50:00.339682Z","shell.execute_reply":"2025-01-18T06:50:00.368269Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mol = Chem.MolFromSmiles(split0[1])\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)   \nimg","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.370076Z","iopub.execute_input":"2025-01-18T06:50:00.370371Z","iopub.status.idle":"2025-01-18T06:50:00.398933Z","shell.execute_reply.started":"2025-01-18T06:50:00.370343Z","shell.execute_reply":"2025-01-18T06:50:00.398181Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Feature Augmentation - Based on : 'SMILES' - train file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"de_train['_SMILES'] = [split_sign(text) for text in de_train['SMILES'].values]\nde_train['_SMILES']","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.399824Z","iopub.execute_input":"2025-01-18T06:50:00.400079Z","iopub.status.idle":"2025-01-18T06:50:00.411668Z","shell.execute_reply.started":"2025-01-18T06:50:00.400034Z","shell.execute_reply":"2025-01-18T06:50:00.410908Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sign = []\nfor row in de_train['_SMILES'].values:\n    for ele in row:\n        sign.append(ele)\n        \nde_train_sign_list = list(set(sign))\nlen(de_train_sign_list)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.412701Z","iopub.execute_input":"2025-01-18T06:50:00.412997Z","iopub.status.idle":"2025-01-18T06:50:00.425727Z","shell.execute_reply.started":"2025-01-18T06:50:00.412968Z","shell.execute_reply":"2025-01-18T06:50:00.425015Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data = np.zeros((len(de_train), len(de_train_sign_list)), dtype=int)\nde_train_sign = pd.DataFrame(data=data, columns=de_train_sign_list)\n\nfor sign in de_train_sign_list:\n    for i in range(len(de_train)):\n        row = de_train['_SMILES'].values[i]\n        \n        if (sign in row):\n            de_train_sign[sign].iloc[i] = 1\n            \nde_train_sign","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:00.426671Z","iopub.execute_input":"2025-01-18T06:50:00.426951Z","iopub.status.idle":"2025-01-18T06:50:02.531401Z","shell.execute_reply.started":"2025-01-18T06:50:00.426931Z","shell.execute_reply":"2025-01-18T06:50:02.530596Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Feature Augmentation - Based on : 'SMILES'","metadata":{}},{"cell_type":"code","source":"id_map['_SMILES'] = [split_sign(text) for text in id_map['SMILES'].values]\nid_map['_SMILES']","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:02.532757Z","iopub.execute_input":"2025-01-18T06:50:02.533146Z","iopub.status.idle":"2025-01-18T06:50:02.542355Z","shell.execute_reply.started":"2025-01-18T06:50:02.533113Z","shell.execute_reply":"2025-01-18T06:50:02.541526Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sign = []\nfor row in id_map['_SMILES'].values:\n    for ele in row:\n        sign.append(ele)\n        \nid_map_sign_list = list(set(sign))\nlen(id_map_sign_list)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:02.543398Z","iopub.execute_input":"2025-01-18T06:50:02.543993Z","iopub.status.idle":"2025-01-18T06:50:02.555267Z","shell.execute_reply.started":"2025-01-18T06:50:02.54397Z","shell.execute_reply":"2025-01-18T06:50:02.55434Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data = np.zeros((len(id_map), len(id_map_sign_list)), dtype=int)\nid_map_sign = pd.DataFrame(data=data, columns=id_map_sign_list)\n\nfor sign in id_map_sign_list:\n    for i in range(len(id_map)):\n        row = id_map['_SMILES'].values[i]\n        \n        if (sign in row):\n            id_map_sign[sign].iloc[i] = 1\n            \nid_map_sign","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:02.556261Z","iopub.execute_input":"2025-01-18T06:50:02.556566Z","iopub.status.idle":"2025-01-18T06:50:03.324859Z","shell.execute_reply.started":"2025-01-18T06:50:02.556536Z","shell.execute_reply":"2025-01-18T06:50:03.323865Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Uncommon Deleted","metadata":{}},{"cell_type":"code","source":"uncommon = [f for f in de_train_sign if f not in id_map_sign]\nlen(uncommon)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.325848Z","iopub.execute_input":"2025-01-18T06:50:03.326152Z","iopub.status.idle":"2025-01-18T06:50:03.332404Z","shell.execute_reply.started":"2025-01-18T06:50:03.32613Z","shell.execute_reply":"2025-01-18T06:50:03.331559Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sign = de_train_sign.drop(columns=uncommon)\ntrain_sign.shape, id_map_sign.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.333486Z","iopub.execute_input":"2025-01-18T06:50:03.333796Z","iopub.status.idle":"2025-01-18T06:50:03.346393Z","shell.execute_reply.started":"2025-01-18T06:50:03.333766Z","shell.execute_reply":"2025-01-18T06:50:03.345539Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"list(train_sign.columns) == list(id_map_sign.columns)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.347347Z","iopub.execute_input":"2025-01-18T06:50:03.347589Z","iopub.status.idle":"2025-01-18T06:50:03.357749Z","shell.execute_reply.started":"2025-01-18T06:50:03.34757Z","shell.execute_reply":"2025-01-18T06:50:03.356801Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_sign = train_sign.sort_index(axis = 1)\nid_map_sign = id_map_sign.sort_index(axis = 1)\n\nlist(train_sign.columns) == list(id_map_sign.columns)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.35874Z","iopub.execute_input":"2025-01-18T06:50:03.358979Z","iopub.status.idle":"2025-01-18T06:50:03.373833Z","shell.execute_reply.started":"2025-01-18T06:50:03.358958Z","shell.execute_reply":"2025-01-18T06:50:03.373103Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X2 = X1.join(train_sign).copy()\nX2.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.375009Z","iopub.execute_input":"2025-01-18T06:50:03.375333Z","iopub.status.idle":"2025-01-18T06:50:03.393666Z","shell.execute_reply.started":"2025-01-18T06:50:03.375311Z","shell.execute_reply":"2025-01-18T06:50:03.392733Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test2 = test1.join(id_map_sign).copy()\ntest2.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.394676Z","iopub.execute_input":"2025-01-18T06:50:03.394909Z","iopub.status.idle":"2025-01-18T06:50:03.402076Z","shell.execute_reply.started":"2025-01-18T06:50:03.394888Z","shell.execute_reply":"2025-01-18T06:50:03.401267Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Create a morgan fingerprint from SMILES","metadata":{}},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import AllChem\nradius = 2\nnBits = 2048\n\n# def get_frag(smiles_str):   \n#     mol = Chem.MolFromSmiles(smiles_str)\n#     frag = AllChem.GetMorganFingerprintAsBitVect(mol, radius, nBits=nBits)\n#     return frag\n\n\n\n# Function using GetMorganGenerator\ndef get_frag(smiles_str, radius=2, nBits=2048):\n    mol = Chem.MolFromSmiles(smiles_str)\n    if mol:\n        # Create a Morgan fingerprint generator\n        generator = AllChem.GetMorganGenerator(radius=radius, fpSize=nBits)\n        # Generate the fingerprint bit vector for the molecule\n        frag = generator.GetFingerprint(mol)\n        return list(frag)  # Converts the fingerprint to a list of bits\n    else:\n        raise ValueError(f\"Invalid SMILES string: {smiles_str}\")","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.402974Z","iopub.execute_input":"2025-01-18T06:50:03.403251Z","iopub.status.idle":"2025-01-18T06:50:03.408973Z","shell.execute_reply.started":"2025-01-18T06:50:03.40323Z","shell.execute_reply":"2025-01-18T06:50:03.408003Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"columns = np.arange(nBits).astype(str)\n\ndata1 = np.zeros((len(de_train), nBits), dtype=int)\ndata2 = np.zeros((len(id_map), nBits), dtype=int)\n\ndf_frag1 = pd.DataFrame(data=data1, columns=columns)\ndf_frag2 = pd.DataFrame(data=data2, columns=columns)\n\ndf_frag1.shape, df_frag2.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.410163Z","iopub.execute_input":"2025-01-18T06:50:03.410433Z","iopub.status.idle":"2025-01-18T06:50:03.423424Z","shell.execute_reply.started":"2025-01-18T06:50:03.410413Z","shell.execute_reply":"2025-01-18T06:50:03.422568Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for f in range(len(df_frag1)):\n    df_frag1.iloc[f] = get_frag(de_train['SMILES'][f])\n\ndf_frag1.shape\n","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:03.424405Z","iopub.execute_input":"2025-01-18T06:50:03.424634Z","iopub.status.idle":"2025-01-18T06:50:04.869788Z","shell.execute_reply.started":"2025-01-18T06:50:03.424615Z","shell.execute_reply":"2025-01-18T06:50:04.868913Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Ensure df_frag2 and id_map['SMILES'] have the same length\nfor f in range(len(id_map['SMILES'])):\n    # Get the fingerprint as a list\n    frag_list = get_frag(id_map['SMILES'][f])\n    \n    # Ensure that the list is of appropriate length to fit into the DataFrame\n    df_frag2.iloc[f] = frag_list\n\n# Check the shape of the DataFrame\ndf_frag2.shape\n","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:04.870904Z","iopub.execute_input":"2025-01-18T06:50:04.871181Z","iopub.status.idle":"2025-01-18T06:50:05.470566Z","shell.execute_reply.started":"2025-01-18T06:50:04.87116Z","shell.execute_reply":"2025-01-18T06:50:05.469669Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = X2.join(df_frag1).copy()\nX.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:05.471845Z","iopub.execute_input":"2025-01-18T06:50:05.47211Z","iopub.status.idle":"2025-01-18T06:50:05.508171Z","shell.execute_reply.started":"2025-01-18T06:50:05.472088Z","shell.execute_reply":"2025-01-18T06:50:05.507326Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test = test2.join(df_frag2).copy()\ntest.shape","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:05.509159Z","iopub.execute_input":"2025-01-18T06:50:05.509411Z","iopub.status.idle":"2025-01-18T06:50:05.522331Z","shell.execute_reply.started":"2025-01-18T06:50:05.50939Z","shell.execute_reply":"2025-01-18T06:50:05.521571Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"KNN & LinearSVR - Final model","metadata":{}},{"cell_type":"code","source":"model0 = lgb.LGBMRegressor()\nmodel1 = KNeighborsRegressor(n_neighbors=13)\nmodel2 = LinearSVR(max_iter= 2000, epsilon= 0.1)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:05.523282Z","iopub.execute_input":"2025-01-18T06:50:05.523572Z","iopub.status.idle":"2025-01-18T06:50:05.530428Z","shell.execute_reply.started":"2025-01-18T06:50:05.523551Z","shell.execute_reply":"2025-01-18T06:50:05.529686Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pred = []\nfor i in range(y.shape[1]):\n    \n    yi = y.iloc[:, i].copy()\n       \n    Xi = X.join(tr_cell_type.iloc[:, i+1], lsuffix='_', rsuffix='__').copy()\n    Xi = Xi.join(tr_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    testi = test.join(te_cell_type.iloc[:, i+1], lsuffix='_', rsuffix='__').copy()\n    testi = testi.join(te_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    model1.fit(Xi, yi)\n    model2.fit(Xi, yi)\n    \n    pred1 = model1.predict(testi)\n    pred2 = model2.predict(testi)    \n    pred.append((pred1 *0.40) + (pred2 *0.60))\n    \nlen(pred)","metadata":{"execution":{"iopub.status.busy":"2025-01-18T06:50:05.531458Z","iopub.execute_input":"2025-01-18T06:50:05.531688Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nimport numpy as np\nimport pandas as pd\n\n# Assuming X and y are your feature and target data respectively\n# Ensure X and y are DataFrames with appropriate shapes\ndef create_model(input_shape):\n    model = keras.Sequential()\n    model.add(keras.layers.Input(shape=input_shape))\n    model.add(keras.layers.Dense(64, activation='relu'))\n    model.add(keras.layers.Dense(32, activation='relu'))\n    model.add(keras.layers.Dense(1))  # Adjust output units as per your needs\n    model.compile(optimizer='adam', loss='mean_squared_error')\n    return model\n\npred = []\nfor i in range(y.shape[1]):\n    yi = y.iloc[:, i].copy()\n    \n    Xi = X.join(tr_cell_type.iloc[:, i + 1], lsuffix='_', rsuffix='__').copy()\n    Xi = Xi.join(tr_sm_name.iloc[:, i + 1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    testi = test.join(te_cell_type.iloc[:, i + 1], lsuffix='_', rsuffix='__').copy()\n    testi = testi.join(te_sm_name.iloc[:, i + 1], lsuffix='_cell_type', rsuffix='_sm_name')\n\n    # Convert boolean columns in Xi and testi to integers\n    for col in [*Xi.columns, *testi.columns]:\n        if Xi[col].dtype == 'bool':\n            Xi[col] = Xi[col].astype(int)\n        if testi[col].dtype == 'bool':\n            testi[col] = testi[col].astype(int)\n\n    # Ensure all data types are compatible\n    Xi = Xi.astype(float)\n    yi = yi.astype(float)\n\n    # Check and handle missing values\n    Xi = Xi.fillna(0)\n    yi = yi.fillna(0)\n    testi = testi.fillna(0)\n\n    # Create and fit the neural network model\n    model = create_model(input_shape=(Xi.shape[1],))\n    model.fit(Xi, yi, epochs=10, batch_size=32, verbose=0)\n    \n    # Predict using the neural network\n    pred1 = model.predict(testi).flatten()  # Flatten if necessary\n    pred.append(pred1)\n\nlen(pred)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"prediction.to_csv('prediction.csv')\n!ls","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}}]}