{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkcyan\"></p>\n\n<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkcyan;\">🧬Open Problems – Single-Cell Perturbations</h1> \n    <h3 align=\"center\" style=\"color:gray;\">Predict how small molecules change gene expression in different cell types</h3> \n    <h3 align=\"center\" style=\"color:gray;\">By: Somayyeh Gholami & Mehran Kazeminia</h3> \n</div>\n\n<p style=\"border-bottom: 5px solid darkcyan\"></p>\n\n# <div style=\"color:white;background-color:darkcyan;padding:1.5%;border-radius:15px 15px;font-size:1em;text-align:center\">Feature Augmentation - LightGBM</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Description for notebook number three :</p></div>\n\n- This notebook is the continuation of [notebook number one](https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain).\n\n- Two available features; are \"cell_type\" and \"sm_name\" and we want to add two new columns (two new features) to them.\n\n- If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, we will see that for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.\n\n- Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.\n\n- Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".\n\n- By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.\n\n- Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.\n\n- It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.\n\n- Good luck.\n\n![](https://cdn-images-1.medium.com/max/1000/1*6lNZoZbkS_vBu54byPS3Yw.jpeg)\n\n[Image Reference](https://www.a-star.edu.sg/gis/our-science/spatial-and-single-cell-systems)","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')\n#:::::::::::::::::::::::::::::::::::\nimport os\nimport gc\nimport glob\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n#:::::::::::::::::::::::::::::::::::\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2023-10-30T18:33:17.55728Z","iopub.execute_input":"2023-10-30T18:33:17.557726Z","iopub.status.idle":"2023-10-30T18:33:18.702825Z","shell.execute_reply.started":"2023-10-30T18:33:17.557693Z","shell.execute_reply":"2023-10-30T18:33:18.701512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkgray;\">Competition Data (Eight files)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(1) de_train.parquet</p></div>","metadata":{}},{"cell_type":"code","source":"de_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet')\nde_train.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:18.705829Z","iopub.execute_input":"2023-10-30T18:33:18.706268Z","iopub.status.idle":"2023-10-30T18:33:20.218674Z","shell.execute_reply.started":"2023-10-30T18:33:18.706228Z","shell.execute_reply":"2023-10-30T18:33:20.217625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(7) id_map.csv</p></div>","metadata":{}},{"cell_type":"code","source":"id_map = pd.read_csv('../input/open-problems-single-cell-perturbations/id_map.csv')\nid_map.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:20.222674Z","iopub.execute_input":"2023-10-30T18:33:20.223037Z","iopub.status.idle":"2023-10-30T18:33:20.238851Z","shell.execute_reply.started":"2023-10-30T18:33:20.223008Z","shell.execute_reply":"2023-10-30T18:33:20.23759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>(8) sample_submission.csv</p></div>","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('../input/open-problems-single-cell-perturbations/sample_submission.csv', index_col='id')\nsample_submission.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:20.241745Z","iopub.execute_input":"2023-10-30T18:33:20.242144Z","iopub.status.idle":"2023-10-30T18:33:25.034198Z","shell.execute_reply.started":"2023-10-30T18:33:20.242112Z","shell.execute_reply":"2023-10-30T18:33:25.032914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>train | test | target</p></div>","metadata":{}},{"cell_type":"code","source":"xlist  = ['cell_type','sm_name']\n_ylist = ['cell_type','sm_name','sm_lincs_id','SMILES','control']\n\ny = de_train.drop(columns=_ylist)\ny.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.0388Z","iopub.execute_input":"2023-10-30T18:33:25.039359Z","iopub.status.idle":"2023-10-30T18:33:25.115692Z","shell.execute_reply.started":"2023-10-30T18:33:25.039307Z","shell.execute_reply":"2023-10-30T18:33:25.114297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">get_dummies (OneHotEncoder)</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"train = pd.get_dummies(de_train[xlist], columns=xlist)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.117227Z","iopub.execute_input":"2023-10-30T18:33:25.11759Z","iopub.status.idle":"2023-10-30T18:33:25.136953Z","shell.execute_reply.started":"2023-10-30T18:33:25.11756Z","shell.execute_reply":"2023-10-30T18:33:25.136004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.get_dummies(id_map[xlist], columns=xlist)\ntest.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.138529Z","iopub.execute_input":"2023-10-30T18:33:25.138895Z","iopub.status.idle":"2023-10-30T18:33:25.15567Z","shell.execute_reply.started":"2023-10-30T18:33:25.138859Z","shell.execute_reply":"2023-10-30T18:33:25.154239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Uncommon deleted</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"uncommon = [f for f in train if f not in test]\nlen(uncommon)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.15994Z","iopub.execute_input":"2023-10-30T18:33:25.160898Z","iopub.status.idle":"2023-10-30T18:33:25.175062Z","shell.execute_reply.started":"2023-10-30T18:33:25.160795Z","shell.execute_reply":"2023-10-30T18:33:25.173781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train.drop(columns=uncommon)\nX.shape[1], test.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.179193Z","iopub.execute_input":"2023-10-30T18:33:25.17967Z","iopub.status.idle":"2023-10-30T18:33:25.189207Z","shell.execute_reply.started":"2023-10-30T18:33:25.179638Z","shell.execute_reply":"2023-10-30T18:33:25.187798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(X.columns) == list(test.columns)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.19365Z","iopub.execute_input":"2023-10-30T18:33:25.194355Z","iopub.status.idle":"2023-10-30T18:33:25.205754Z","shell.execute_reply.started":"2023-10-30T18:33:25.19432Z","shell.execute_reply":"2023-10-30T18:33:25.204342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Evaluation</p></div>\n\n### <span style=\"color:navy;\">Mean Rowwise Root Mean Squared Error (MRRMSE)</span>","metadata":{}},{"cell_type":"code","source":"def mrrmse_pd(y_pred: pd.DataFrame, y_true: pd.DataFrame):\n    \n    return ((y_pred - y_true)**2).mean(axis=1).apply(np.sqrt).mean()","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.210108Z","iopub.execute_input":"2023-10-30T18:33:25.210499Z","iopub.status.idle":"2023-10-30T18:33:25.216486Z","shell.execute_reply.started":"2023-10-30T18:33:25.210469Z","shell.execute_reply":"2023-10-30T18:33:25.215567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mrrmse_np(y_pred, y_true):\n    \n    return np.sqrt(np.square(y_true - y_pred).mean(axis=1)).mean()","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.218026Z","iopub.execute_input":"2023-10-30T18:33:25.218376Z","iopub.status.idle":"2023-10-30T18:33:25.228666Z","shell.execute_reply.started":"2023-10-30T18:33:25.218339Z","shell.execute_reply":"2023-10-30T18:33:25.227512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:navy;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:lightgray;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Feature Augmentation</p></div>","metadata":{}},{"cell_type":"code","source":"de_cell_type = de_train.iloc[:, [0] + list(range(5, de_train.shape[1]))]\nde_sm_name = de_train.iloc[:, [1] + list(range(5, de_train.shape[1]))]\n\nde_cell_type.shape, de_sm_name.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.230275Z","iopub.execute_input":"2023-10-30T18:33:25.2307Z","iopub.status.idle":"2023-10-30T18:33:25.342235Z","shell.execute_reply.started":"2023-10-30T18:33:25.230659Z","shell.execute_reply":"2023-10-30T18:33:25.341043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Calculate averages based on 'cell_type' and 'sm_name' for all columns</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"mean_cell_type = de_cell_type.groupby('cell_type').mean().reset_index()\nmean_sm_name = de_sm_name.groupby('sm_name').mean().reset_index()\n\ndisplay(mean_cell_type)\ndisplay(mean_sm_name)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.344057Z","iopub.execute_input":"2023-10-30T18:33:25.344902Z","iopub.status.idle":"2023-10-30T18:33:25.880232Z","shell.execute_reply.started":"2023-10-30T18:33:25.344859Z","shell.execute_reply":"2023-10-30T18:33:25.878917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Paste the results into the train file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in de_cell_type['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\ntr_cell_type = pd.concat(rows)\ntr_cell_type = tr_cell_type.reset_index(drop=True)\ntr_cell_type","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:25.882374Z","iopub.execute_input":"2023-10-30T18:33:25.88288Z","iopub.status.idle":"2023-10-30T18:33:27.773006Z","shell.execute_reply.started":"2023-10-30T18:33:25.882838Z","shell.execute_reply":"2023-10-30T18:33:27.771702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rows = []\nfor name in de_sm_name['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\ntr_sm_name = pd.concat(rows)\ntr_sm_name = tr_sm_name.reset_index(drop=True)\ntr_sm_name","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:27.77462Z","iopub.execute_input":"2023-10-30T18:33:27.775082Z","iopub.status.idle":"2023-10-30T18:33:29.498987Z","shell.execute_reply.started":"2023-10-30T18:33:27.775045Z","shell.execute_reply":"2023-10-30T18:33:29.497801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Paste the results into the test file</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"rows = []\nfor name in id_map['cell_type']:\n    mean_rows = mean_cell_type[mean_cell_type['cell_type'] == name].copy()\n    rows.append(mean_rows)\n\nte_cell_type = pd.concat(rows)\nte_cell_type = te_cell_type.reset_index(drop=True)\nte_cell_type","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:29.500784Z","iopub.execute_input":"2023-10-30T18:33:29.50154Z","iopub.status.idle":"2023-10-30T18:33:30.195259Z","shell.execute_reply.started":"2023-10-30T18:33:29.501498Z","shell.execute_reply":"2023-10-30T18:33:30.194308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rows = []\nfor name in id_map['sm_name']:\n    mean_rows = mean_sm_name[mean_sm_name['sm_name'] == name].copy()\n    rows.append(mean_rows)\n\nte_sm_name = pd.concat(rows)\nte_sm_name = te_sm_name.reset_index(drop=True)\nte_sm_name","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:30.196693Z","iopub.execute_input":"2023-10-30T18:33:30.197287Z","iopub.status.idle":"2023-10-30T18:33:30.926984Z","shell.execute_reply.started":"2023-10-30T18:33:30.197256Z","shell.execute_reply":"2023-10-30T18:33:30.92579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:pink;font-block:cyan;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Sample: Column number zero - A1BG</p></div>","metadata":{}},{"cell_type":"code","source":"y0 = y.iloc[:, 0].copy()\ny0","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:30.928872Z","iopub.execute_input":"2023-10-30T18:33:30.929628Z","iopub.status.idle":"2023-10-30T18:33:30.94345Z","shell.execute_reply.started":"2023-10-30T18:33:30.929586Z","shell.execute_reply":"2023-10-30T18:33:30.941724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X0 = X.join(tr_cell_type.iloc[:, 0+1]).copy()\nX0 = X0.join(tr_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\nX0","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:30.945458Z","iopub.execute_input":"2023-10-30T18:33:30.94596Z","iopub.status.idle":"2023-10-30T18:33:30.997479Z","shell.execute_reply.started":"2023-10-30T18:33:30.945919Z","shell.execute_reply":"2023-10-30T18:33:30.99576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test0 = test.join(te_cell_type.iloc[:, 0+1]).copy()\ntest0 = test0.join(te_sm_name.iloc[:, 0+1], lsuffix='_cell_type', rsuffix='_sm_name')\ntest0","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:30.999972Z","iopub.execute_input":"2023-10-30T18:33:31.000488Z","iopub.status.idle":"2023-10-30T18:33:31.041029Z","shell.execute_reply.started":"2023-10-30T18:33:31.000443Z","shell.execute_reply":"2023-10-30T18:33:31.039828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Correlation - Column #0</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"X0_corr = X0.copy()\nX0_corr['y0'] = y0\n\ncorr = X0_corr.iloc[: , 131:].corr(numeric_only=True).round(3)\ncorr.style.background_gradient(cmap='Pastel1')","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.042597Z","iopub.execute_input":"2023-10-30T18:33:31.043323Z","iopub.status.idle":"2023-10-30T18:33:31.061652Z","shell.execute_reply.started":"2023-10-30T18:33:31.043288Z","shell.execute_reply":"2023-10-30T18:33:31.060447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cor_matrix = X0_corr.iloc[: , 131:].corr()\nfig = plt.figure(figsize=(6,6));\n\ncmap=sns.diverging_palette(240, 10, s=75, l=50, sep=1, n=6, center='light', as_cmap=False);\nsns.heatmap(cor_matrix, center=0, annot=True, cmap=cmap, linewidths=5);\nplt.suptitle('Train Set (Heatmap)', y=0.92, fontsize=16, c='darkred');\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.063382Z","iopub.execute_input":"2023-10-30T18:33:31.063753Z","iopub.status.idle":"2023-10-30T18:33:31.466964Z","shell.execute_reply.started":"2023-10-30T18:33:31.063723Z","shell.execute_reply":"2023-10-30T18:33:31.465903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">LightGBM - Column #0</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\nfrom sklearn.svm import LinearSVR\nfrom sklearn.neighbors import KNeighborsRegressor\nfrom sklearn.multioutput import MultiOutputRegressor\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(X0, y0, test_size=0.20, random_state=421)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.468673Z","iopub.execute_input":"2023-10-30T18:33:31.469368Z","iopub.status.idle":"2023-10-30T18:33:31.478624Z","shell.execute_reply.started":"2023-10-30T18:33:31.469327Z","shell.execute_reply":"2023-10-30T18:33:31.477303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = lgb.LGBMRegressor()\nmodel.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.482166Z","iopub.execute_input":"2023-10-30T18:33:31.482514Z","iopub.status.idle":"2023-10-30T18:33:31.80095Z","shell.execute_reply.started":"2023-10-30T18:33:31.482486Z","shell.execute_reply":"2023-10-30T18:33:31.799762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict = model.predict(X_test) \n# mrrmse_pd(pd.DataFrame(predict), pd.DataFrame(y_test.values))","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.802611Z","iopub.execute_input":"2023-10-30T18:33:31.803745Z","iopub.status.idle":"2023-10-30T18:33:31.814193Z","shell.execute_reply.started":"2023-10-30T18:33:31.803702Z","shell.execute_reply":"2023-10-30T18:33:31.813025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N = 0\nplt.style.use('seaborn-whitegrid') \nplt.figure(figsize=(8, 4), facecolor='lightyellow')\nplt.title(f'Column:  #{N}', fontsize=12)\nplt.gca().set_facecolor('lightgray')\n\nsns.distplot(y_test.values-predict, bins=100, color='red')\nplt.legend(['y_true','y_pred'], loc=1)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:31.815942Z","iopub.execute_input":"2023-10-30T18:33:31.816392Z","iopub.status.idle":"2023-10-30T18:33:32.399865Z","shell.execute_reply.started":"2023-10-30T18:33:31.816351Z","shell.execute_reply":"2023-10-30T18:33:32.398441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:navy;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:lightgray;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>LightGBM - Final mode</p></div>","metadata":{}},{"cell_type":"code","source":"model = lgb.LGBMRegressor()\n# model = KNeighborsRegressor(n_neighbors=13)\n# model = LinearSVR(max_iter= 2000, epsilon= 0.1)","metadata":{"execution":{"iopub.status.busy":"2023-10-30T18:33:39.459195Z","iopub.execute_input":"2023-10-30T18:33:39.460463Z","iopub.status.idle":"2023-10-30T18:33:39.465362Z","shell.execute_reply.started":"2023-10-30T18:33:39.460415Z","shell.execute_reply":"2023-10-30T18:33:39.464117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = []\nfor i in range(y.shape[1]):\n    \n    yi = y.iloc[:, i].copy()\n    \n    Xi = X.join(tr_cell_type.iloc[:, i+1]).copy()\n    Xi = Xi.join(tr_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    testi = test.join(te_cell_type.iloc[:, i+1]).copy()\n    testi = testi.join(te_sm_name.iloc[:, i+1], lsuffix='_cell_type', rsuffix='_sm_name')\n    \n    model.fit(Xi, yi)\n    pred.append(model.predict(testi))\n    \nlen(pred)","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:43:21.96189Z","iopub.execute_input":"2023-10-27T12:43:21.963142Z","iopub.status.idle":"2023-10-27T12:53:56.208375Z","shell.execute_reply.started":"2023-10-27T12:43:21.9631Z","shell.execute_reply":"2023-10-27T12:53:56.207024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction = pd.DataFrame(pred).T\nprediction.columns = de_train.columns[5:]\nprediction.index.name = 'id'\nprediction","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:53:56.210998Z","iopub.execute_input":"2023-10-27T12:53:56.211981Z","iopub.status.idle":"2023-10-27T12:54:00.339395Z","shell.execute_reply.started":"2023-10-27T12:53:56.21193Z","shell.execute_reply":"2023-10-27T12:54:00.337996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction.to_csv('prediction.csv')\n!ls","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:00.345757Z","iopub.execute_input":"2023-10-27T12:54:00.346165Z","iopub.status.idle":"2023-10-27T12:54:15.121671Z","shell.execute_reply.started":"2023-10-27T12:54:00.346135Z","shell.execute_reply":"2023-10-27T12:54:15.119795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:cyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Ensembling</p></div>","metadata":{}},{"cell_type":"markdown","source":"Several public notebooks presented different methods based on averaging, different guesses, etc. The following file is actually the optimization and ensembling of these results.\n\nThanks to: **@alexandervc**","metadata":{}},{"cell_type":"code","source":"import1 = pd.read_csv('../input/op2-603/op2_603.csv', index_col='id')\nimport1.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:15.124113Z","iopub.execute_input":"2023-10-27T12:54:15.124624Z","iopub.status.idle":"2023-10-27T12:54:21.817601Z","shell.execute_reply.started":"2023-10-27T12:54:15.124587Z","shell.execute_reply":"2023-10-27T12:54:21.816194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The file below is the result of a great notebook that uses the \"Autoencoder\" method.\n\nThanks to: **@vendekagonlabs**","metadata":{}},{"cell_type":"code","source":"import2 = pd.read_csv('../input/op2-720/op2_720.csv', index_col='id')\nimport2.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:21.819507Z","iopub.execute_input":"2023-10-27T12:54:21.821031Z","iopub.status.idle":"2023-10-27T12:54:28.678146Z","shell.execute_reply.started":"2023-10-27T12:54:21.820982Z","shell.execute_reply":"2023-10-27T12:54:28.676058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The file below is the result of a great notebook that uses the \"Neural Network\" method.\n\nThanks to: **@kishanvavdara**","metadata":{}},{"cell_type":"code","source":"import3 = pd.read_csv('../input/op2-604/submission_df.csv', index_col='id')\nimport3.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:28.680647Z","iopub.execute_input":"2023-10-27T12:54:28.6812Z","iopub.status.idle":"2023-10-27T12:54:34.878051Z","shell.execute_reply.started":"2023-10-27T12:54:28.681152Z","shell.execute_reply":"2023-10-27T12:54:34.876572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The file below is the result of a great notebook that uses the \"NLP\" method.\n\nThanks to: **@kishanvavdara**","metadata":{}},{"cell_type":"code","source":"import4 = pd.read_csv('../input/op2-607/OP2_607.csv', index_col='id')\nimport4.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:34.880351Z","iopub.execute_input":"2023-10-27T12:54:34.880885Z","iopub.status.idle":"2023-10-27T12:54:40.874128Z","shell.execute_reply.started":"2023-10-27T12:54:34.880838Z","shell.execute_reply":"2023-10-27T12:54:40.872765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = list(de_train.columns[5:])\nsubmission = sample_submission.copy()\n\nsubmission[col] = (import1[col] *0.24) + (import2[col] *0.16) + (import3[col] *0.2) + (import4[col] *0.24) + (prediction[col] *0.16)\nsubmission.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:40.875764Z","iopub.execute_input":"2023-10-27T12:54:40.876238Z","iopub.status.idle":"2023-10-27T12:54:51.611135Z","shell.execute_reply.started":"2023-10-27T12:54:40.876201Z","shell.execute_reply":"2023-10-27T12:54:51.609855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv')\n!ls","metadata":{"execution":{"iopub.status.busy":"2023-10-27T12:54:51.613409Z","iopub.execute_input":"2023-10-27T12:54:51.613894Z","iopub.status.idle":"2023-10-27T12:55:50.260884Z","shell.execute_reply.started":"2023-10-27T12:54:51.613855Z","shell.execute_reply":"2023-10-27T12:55:50.259023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:navy;\">Ensembling Histograms</span>\n\n<p style=\"border-bottom: 5px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"N = random.randrange(y.shape[1])\n\nprint(':' *40)\nprint('Column number :', N)\nprint('Column name :', list(y.columns)[N])\nprint(':' *40)\n\nhist_data = [submission.iloc[:, N], import1.iloc[:, N], import2.iloc[:, N], import3.iloc[:, N], import4.iloc[:, N], prediction.iloc[:, N]]\ngroup_labels = ['Submission', 'Mean & Paste', 'Autoencoder', 'Neural Network', 'NLP', 'Prediction']\n    \nfig = ff.create_distplot(hist_data, group_labels, bin_size=.2, show_hist=False, show_rug=False)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-27T13:06:20.549396Z","iopub.execute_input":"2023-10-27T13:06:20.549895Z","iopub.status.idle":"2023-10-27T13:06:20.666775Z","shell.execute_reply.started":"2023-10-27T13:06:20.549861Z","shell.execute_reply":"2023-10-27T13:06:20.665645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<p style=\"border-bottom: 5px solid darkgray\"></p>","metadata":{}}]}