{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkcyan\"></p>\n\n<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkcyan;\">🧬Open Problems – Single-Cell Perturbations</h1> \n    <h3 align=\"center\" style=\"color:gray;\">Predict how small molecules change gene expression in different cell types</h3> \n    <h3 align=\"center\" style=\"color:gray;\">By: Somayyeh Gholami & Mehran Kazeminia</h3> \n</div>\n\n<p style=\"border-bottom: 5px solid darkcyan\"></p>\n\n# <div style=\"color:white;background-color:darkcyan;padding:1.5%;border-radius:15px 15px;font-size:1em;text-align:center\">Feature Augmentation & Fragments of SMILES</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Description for notebook number three :</p></div>\n\n- This notebook is the continuation of [notebook number one](https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain).\n\n- Two available features; are \"cell_type\" and \"sm_name\" and we want to add two new columns (two new features) to them.\n\n- If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, we will see that for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.\n\n- Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.\n\n- Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".\n\n- By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.\n\n- Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.\n\n- It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Adding \"fragments of SMILES\" as a feature, has been done since version 8.</p></div>\n\nWe speculated that \"fragments of SMILES\" may represent a special feature to counteract the addition of drugs. So we did the same thing and added \"fragments of SMILES\" as a feature to the calculations. Because the results have improved, it has been included in our notebook from version eighth onwards. Please note that accuracy in splitting \"SMILES\" can probably improve the results. Also, because it seemed a little complicated, we tried to explain the matter a little more for the first line.\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Adding morgan fingerprint from SMILES, has been done since version 10.</p></div>\n\n- Good luck.\n\n![](https://cdn-images-1.medium.com/max/1000/1*6lNZoZbkS_vBu54byPS3Yw.jpeg)\n\n[Image Reference](https://www.a-star.edu.sg/gis/our-science/spatial-and-single-cell-systems)","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')\n#:::::::::::::::::::::::::::::::::::\nimport os\nimport gc\nimport glob\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n#:::::::::::::::::::::::::::::::::::\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2023-11-12T18:50:02.299346Z","iopub.execute_input":"2023-11-12T18:50:02.301236Z","iopub.status.idle":"2023-11-12T18:50:03.476759Z","shell.execute_reply.started":"2023-11-12T18:50:02.301181Z","shell.execute_reply":"2023-11-12T18:50:03.474952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n<div>\n    <h1 align=\"center\" style=\"color:gray;\">Competition Data (Eight files)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(1) de_train.parquet</p></div>","metadata":{}},{"cell_type":"code","source":"de_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet')\nde_train.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:03.481726Z","iopub.execute_input":"2023-11-12T18:50:03.482392Z","iopub.status.idle":"2023-11-12T18:50:05.093719Z","shell.execute_reply.started":"2023-11-12T18:50:03.482336Z","shell.execute_reply":"2023-11-12T18:50:05.092053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(7) id_map.csv</p></div>","metadata":{}},{"cell_type":"code","source":"id_map = pd.read_csv('../input/open-problems-single-cell-perturbations/id_map.csv')\nid_map.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:05.095742Z","iopub.execute_input":"2023-11-12T18:50:05.096223Z","iopub.status.idle":"2023-11-12T18:50:05.127239Z","shell.execute_reply.started":"2023-11-12T18:50:05.096187Z","shell.execute_reply":"2023-11-12T18:50:05.125772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(8) sample_submission.csv</p></div>","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('../input/open-problems-single-cell-perturbations/sample_submission.csv', index_col='id')\nsample_submission.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:05.131015Z","iopub.execute_input":"2023-11-12T18:50:05.131388Z","iopub.status.idle":"2023-11-12T18:50:09.892097Z","shell.execute_reply.started":"2023-11-12T18:50:05.131359Z","shell.execute_reply":"2023-11-12T18:50:09.890654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>train | test | target</p></div>","metadata":{}},{"cell_type":"code","source":"xlist  = ['cell_type','sm_name']\n_ylist = ['cell_type','sm_name','sm_lincs_id','SMILES','control']\n\ny = de_train.drop(columns=_ylist)\ny.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:09.895433Z","iopub.execute_input":"2023-11-12T18:50:09.895813Z","iopub.status.idle":"2023-11-12T18:50:09.954097Z","shell.execute_reply.started":"2023-11-12T18:50:09.89578Z","shell.execute_reply":"2023-11-12T18:50:09.95272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">get_dummies (OneHotEncoder)</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"train = pd.get_dummies(de_train[xlist], columns=xlist)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:09.955912Z","iopub.execute_input":"2023-11-12T18:50:09.956242Z","iopub.status.idle":"2023-11-12T18:50:09.97063Z","shell.execute_reply.started":"2023-11-12T18:50:09.956215Z","shell.execute_reply":"2023-11-12T18:50:09.969185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test1 = pd.get_dummies(id_map[xlist], columns=xlist)\ntest1.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:09.972543Z","iopub.execute_input":"2023-11-12T18:50:09.972904Z","iopub.status.idle":"2023-11-12T18:50:09.986567Z","shell.execute_reply.started":"2023-11-12T18:50:09.972865Z","shell.execute_reply":"2023-11-12T18:50:09.985152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Uncommon deleted</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"uncommon = [f for f in train if f not in test1]\nlen(uncommon)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:09.988672Z","iopub.execute_input":"2023-11-12T18:50:09.989138Z","iopub.status.idle":"2023-11-12T18:50:09.997313Z","shell.execute_reply.started":"2023-11-12T18:50:09.989088Z","shell.execute_reply":"2023-11-12T18:50:09.996287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X1 = train.drop(columns=uncommon)\nX1.shape[1], test1.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:09.998791Z","iopub.execute_input":"2023-11-12T18:50:09.99916Z","iopub.status.idle":"2023-11-12T18:50:10.0133Z","shell.execute_reply.started":"2023-11-12T18:50:09.999123Z","shell.execute_reply":"2023-11-12T18:50:10.012351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(X1.columns) == list(test1.columns)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.014722Z","iopub.execute_input":"2023-11-12T18:50:10.015046Z","iopub.status.idle":"2023-11-12T18:50:10.02668Z","shell.execute_reply.started":"2023-11-12T18:50:10.015019Z","shell.execute_reply":"2023-11-12T18:50:10.025381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Evaluation</p></div>\n\n### <span style=\"color:navy;\">Mean Rowwise Root Mean Squared Error (MRRMSE)</span>","metadata":{}},{"cell_type":"code","source":"def mrrmse_pd(y_pred: pd.DataFrame, y_true: pd.DataFrame):\n    \n    return ((y_pred - y_true)**2).mean(axis=1).apply(np.sqrt).mean()","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.028782Z","iopub.execute_input":"2023-11-12T18:50:10.02929Z","iopub.status.idle":"2023-11-12T18:50:10.038396Z","shell.execute_reply.started":"2023-11-12T18:50:10.029249Z","shell.execute_reply":"2023-11-12T18:50:10.036931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mrrmse_np(y_pred, y_true):\n    \n    return np.sqrt(np.square(y_true - y_pred).mean(axis=1)).mean()","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.040048Z","iopub.execute_input":"2023-11-12T18:50:10.040356Z","iopub.status.idle":"2023-11-12T18:50:10.051808Z","shell.execute_reply.started":"2023-11-12T18:50:10.04033Z","shell.execute_reply":"2023-11-12T18:50:10.050347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:navy;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:lightgray;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Feature Augmentation</p></div>","metadata":{}},{"cell_type":"code","source":"de_cell_type = de_train.iloc[:, [0] + list(range(5, de_train.shape[1]))]\nde_sm_name = de_train.iloc[:, [1] + list(range(5, de_train.shape[1]))]\n\nde_cell_type.shape, de_sm_name.shape","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.053818Z","iopub.execute_input":"2023-11-12T18:50:10.054291Z","iopub.status.idle":"2023-11-12T18:50:10.146323Z","shell.execute_reply.started":"2023-11-12T18:50:10.05425Z","shell.execute_reply":"2023-11-12T18:50:10.145121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Calculate averages based on 'cell_type' and 'sm_name' for all columns</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"mean_cell_type = de_cell_type.groupby('cell_type').mean().reset_index()\nmean_sm_name = de_sm_name.groupby('sm_name').mean().reset_index()\n\ndisplay(mean_cell_type)\ndisplay(mean_sm_name)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.152671Z","iopub.execute_input":"2023-11-12T18:50:10.153248Z","iopub.status.idle":"2023-11-12T18:50:10.576988Z","shell.execute_reply.started":"2023-11-12T18:50:10.15321Z","shell.execute_reply":"2023-11-12T18:50:10.575777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Paste the results into the train file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in de_cell_type['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\ntr_cell_type = pd.concat(rows)\ntr_cell_type = tr_cell_type.reset_index(drop=True)\ntr_cell_type","metadata":{"execution":{"iopub.status.busy":"2023-11-12T18:50:10.579023Z","iopub.execute_input":"2023-11-12T18:50:10.579453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rows = []\nfor name in de_sm_name['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\ntr_sm_name = pd.concat(rows)\ntr_sm_name = tr_sm_name.reset_index(drop=True)\ntr_sm_name","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Paste the results into the test file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in id_map['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\nte_cell_type = pd.concat(rows)\nte_cell_type = te_cell_type.reset_index(drop=True)\nte_cell_type","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rows = []\nfor name in id_map['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\nte_sm_name = pd.concat(rows)\nte_sm_name = te_sm_name.reset_index(drop=True)\nte_sm_name","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:cyan;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Sample: Column number zero - A1BG</p></div>","metadata":{}},{"cell_type":"code","source":"y0 = y.iloc[:, 0].copy()\ny0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X0 = X1.join(tr_cell_type.iloc[:, 0+1]).copy()\nX0 = X0.join(tr_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\nX0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test0 = test1.join(te_cell_type.iloc[:, 0+1]).copy()\ntest0 = test0.join(te_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\ntest0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Correlation - Column #0</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"X0_corr = X0.copy()\nX0_corr['y0'] = y0\n\ncorr = X0_corr.iloc[: , 131:].corr(numeric_only=True).round(3)\ncorr.style.background_gradient(cmap='Pastel1')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cor_matrix = X0_corr.iloc[: , 131:].corr()\nfig = plt.figure(figsize=(6,6));\n\ncmap=sns.diverging_palette(240, 10, s=75, l=50, sep=1, n=6, center='light', as_cmap=False);\nsns.heatmap(cor_matrix, center=0, annot=True, cmap=cmap, linewidths=5);\nplt.suptitle('Train Set (Heatmap)', y=0.92, fontsize=16, c='darkred');\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">LightGBM - Column #0</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\nfrom sklearn.svm import LinearSVR\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.multioutput import MultiOutputRegressor\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(X0, y0, test_size=0.20, random_state=421)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = lgb.LGBMRegressor()\nmodel.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict = model.predict(X_test) \n# mrrmse_pd(pd.DataFrame(predict), pd.DataFrame(y_test.values))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N = 0\nplt.style.use('seaborn-whitegrid') \nplt.figure(figsize=(8, 4), facecolor='lightyellow')\nplt.title(f'Column:  #{N}', fontsize=12)\nplt.gca().set_facecolor('lightgray')\n\nsns.distplot(y_test.values-predict, bins=100, color='red')\nplt.legend(['y_true','y_pred'], loc=1)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Feature Augmentation - Based on : 'SMILES'</p></div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Add column : id_map['SMILES']</p></div>","metadata":{}},{"cell_type":"code","source":"sm_name_smiles_dict = de_train.set_index('sm_name')['SMILES'].to_dict()\n\nid_map['SMILES'] = id_map['sm_name'].map(sm_name_smiles_dict)\nid_map['SMILES']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def split_sign(text):\n    text = text.replace(')(', ' ')\n    text = text.replace('(' , ' ')\n    text = text.replace(')' , ' ')\n    return text.split(\" \")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">An example for the method of splitting each 'SMILES'</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"pip install rdkit","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem import AllChem","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import re\n# Thanks to: https://github.com/DocMinus\n\ndef element_count(data, N):\n    smiles = data['SMILES'].iloc[N]\n\n    pattern = \"Si|Ti|Al|Zn|Pd|Pt|Br?|Cl?|N|O|S|P|F|I|B|b|c|n|o|s|p\" \n    regex = re.compile(pattern)\n    elements = [token for token in regex.findall(smiles)]\n    ele_length = len(elements)\n    lowercase_pattern = \"b|c|n|o|s|p\"\n    regex_low = re.compile(lowercase_pattern)\n    for i in range(ele_length):\n        if regex_low.findall(elements[i]):\n            elements[i] = elements[i].upper()\n    element_count = [elements.count(ele) for ele in elements]\n    formula = dict(zip(elements, element_count)) \n\n    print('==> Element_Count: ', str(formula))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_sel = ['cell_type','sm_name','sm_lincs_id','SMILES']\n\nsmiles = de_train['SMILES'][0]\nmol = Chem.MolFromSmiles(smiles)\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)    \ndisplay(pd.DataFrame(de_train[cols_sel].iloc[0]))\n\nelement_count(de_train, 0)\nimg","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"smiles0 = de_train['SMILES'][0]\nsmiles0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"split0 = split_sign(smiles0)\nsplit0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mol = Chem.MolFromSmiles(split0[0])\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)   \nimg","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mol = Chem.MolFromSmiles(split0[1])\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)   \nimg","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Feature Augmentation - Based on : 'SMILES' - train file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"de_train['_SMILES'] = [split_sign(text) for text in de_train['SMILES'].values]\nde_train['_SMILES']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sign = []\nfor row in de_train['_SMILES'].values:\n    for ele in row:\n        sign.append(ele)\n        \nde_train_sign_list = list(set(sign))\nlen(de_train_sign_list)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = np.zeros((len(de_train), len(de_train_sign_list)), dtype=int)\nde_train_sign = pd.DataFrame(data=data, columns=de_train_sign_list)\n\nfor sign in de_train_sign_list:\n    for i in range(len(de_train)):\n        row = de_train['_SMILES'].values[i]\n        \n        if (sign in row):\n            de_train_sign[sign].iloc[i] = 1\n            \nde_train_sign","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Feature Augmentation - Based on : 'SMILES' - test file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"id_map['_SMILES'] = [split_sign(text) for text in id_map['SMILES'].values]\nid_map['_SMILES']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sign = []\nfor row in id_map['_SMILES'].values:\n    for ele in row:\n        sign.append(ele)\n        \nid_map_sign_list = list(set(sign))\nlen(id_map_sign_list)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = np.zeros((len(id_map), len(id_map_sign_list)), dtype=int)\nid_map_sign = pd.DataFrame(data=data, columns=id_map_sign_list)\n\nfor sign in id_map_sign_list:\n    for i in range(len(id_map)):\n        row = id_map['_SMILES'].values[i]\n        \n        if (sign in row):\n            id_map_sign[sign].iloc[i] = 1\n            \nid_map_sign","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Uncommon Deleted</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"uncommon = [f for f in de_train_sign if f not in id_map_sign]\nlen(uncommon)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sign = de_train_sign.drop(columns=uncommon)\ntrain_sign.shape, id_map_sign.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(train_sign.columns) == list(id_map_sign.columns)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sign = train_sign.sort_index(axis = 1)\nid_map_sign = id_map_sign.sort_index(axis = 1)\n\nlist(train_sign.columns) == list(id_map_sign.columns)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X2 = X1.join(train_sign).copy()\nX2.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test2 = test1.join(id_map_sign).copy()\ntest2.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Create a morgan fingerprint from SMILES</p></div>","metadata":{}},{"cell_type":"code","source":"radius = 2\nnBits = 2048\n\ndef get_frag(smiles_str):   \n    mol = Chem.MolFromSmiles(smiles_str)\n    frag = AllChem.GetMorganFingerprintAsBitVect(mol, radius, nBits=nBits)\n    return frag","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = np.arange(nBits).astype(str)\n\ndata1 = np.zeros((len(de_train), nBits), dtype=int)\ndata2 = np.zeros((len(id_map), nBits), dtype=int)\n\ndf_frag1 = pd.DataFrame(data=data1, columns=columns)\ndf_frag2 = pd.DataFrame(data=data2, columns=columns)\n\ndf_frag1.shape, df_frag2.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for f in range(len(df_frag1)):\n    df_frag1.iloc[f] = list(get_frag(de_train['SMILES'][f]))\n    \ndf_frag1.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for f in range(len(df_frag2)):\n    df_frag2.iloc[f] = list(get_frag(id_map['SMILES'][f]))\n    \ndf_frag2.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = X2.join(df_frag1).copy()\nX.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = test2.join(df_frag2).copy()\ntest.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:navy;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:lightgray;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>KNN & LinearSVR - Final model</p></div>","metadata":{}},{"cell_type":"code","source":"model0 = lgb.LGBMRegressor()\nmodel1 = KNeighborsRegressor(n_neighbors=13)\nmodel2 = LinearSVR(max_iter= 2000, epsilon= 0.1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = []\nfor i in range(y.shape[1]):\n    \n    yi = y.iloc[:, i].copy()\n       \n    Xi = X.join(tr_cell_type.iloc[:, i+1], lsuffix='_', rsuffix='__').copy()\n    Xi = Xi.join(tr_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    testi = test.join(te_cell_type.iloc[:, i+1], lsuffix='_', rsuffix='__').copy()\n    testi = testi.join(te_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    model1.fit(Xi, yi)\n    model2.fit(Xi, yi)\n    \n    pred1 = model1.predict(testi)\n    pred2 = model2.predict(testi)    \n    pred.append((pred1 *0.40) + (pred2 *0.60))\n    \nlen(pred)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"de_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet')\n\nprediction = pd.DataFrame(pred).T\nprediction.columns = de_train.columns[5:]\nprediction.index.name = 'id'\nprediction","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction.to_csv('prediction.csv')\n!ls","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>","metadata":{}}]}