{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":6567471,"sourceType":"datasetVersion","datasetId":3755626}],"dockerImageVersionId":30559,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nEducational notebook on the new breakthough gradient boosting package \"PYBOOST\". \n\n### PYBOOST - is the best boosting for MULTI-target tasks\n\nIt designed especially for MULTI-Target(!) tasks.  For such tasks it consitenly outperforms the other boostings,\nas well as incomparably faster. \n\n#### PYBOOST -> Gold on Kaggle\n\nPyboost is created by team lead by the Kaggle Grandmaster - Btbpanda - Anton Vakhrushev  https://www.kaggle.com/btbpanda\n\nHe and others got many gold medals on Kaggle: recent CAFA5 (https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/434064 ),\nlast year \"Open Problems\" top7-team \"Chromosom\", and public top2 team, etc.\n\nSome inpendent check - PyBoost outperaforming other boostings:  \"Table 2. Performance of Kinase Activity Prediction Modelsa\"\nhttps://pubs.acs.org/doi/full/10.1021/acs.jcim.3c00132\n\n#### PYBOOST - easy to use - similar interface and params as other boostings and sklearn models\n\n    model = GradientBoosting('mse')\n    model.fit(X,Y)\n    model.predict(X)\n    \n    Main params are similar to other boostings:\n    # ntrees - number of trees (iterations)\n    # lr  - learning rate\n    # max_depth - maximal depth for trees\n    # subsample - subsampling fraction\n    # colsample - subsampleling fraction for columns\n    # Etc \n    # Example:\n    model = GradientBoosting('mse', ntrees=ntrees, lr=lr,  max_depth=max_depth  , subsample=subsample, colsample=colsample, min_data_in_leaf=1, min_gain_to_split=0, verbose=100  )\n    \n\n#### OP2 - easily LB-score 0.604 -  welcome improve it \n\nThe present notebooks shares a Pyboost prediction with LB 0.604 sore - one of the best public solo-model solutions for the moment.\n\nWelcome to improve it with more tuning.  Notebook provides simple framework to tune - just change the intervals of the params - and run it.\nSubmissions will be saved to CSV files. Try: more iterations, tune colsample, other params, encoders, average several random seeds, etc...\nThe default tuning is done by \"AmbrosM\" validation scheme, you can change it to MT, Random or any other scheme. The LB - CV correspondence is poor,\nbut still not absent. \n\nThe LB604 file prediction can be found in Kaggle dataset: \nhttps://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv\nor uploaded from the present notebook version 19:\nhttps://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool/output?scriptVersionId=149987684&select=submission_Pyboost_max_depth10_ntrees1000_lr001_subsample1_colsample02_n_components50.csv\n\n\n#### PYBOOST - is for GPU\n\nCurrent version of Pyboost requires GPU.\n\n#### Organization of the notebook\n\n    Section: \" PY-BOOST Install (GPU only) \" \n    Shows how to install/import pyboost - do not forget to turn on GPU.\n\n    Section: \" Modeling \"\n    Contains the main modeling example.\n    One can see how simple is to use the pyboost.\n    \n    Also:\n    \n    a kind of grid-search loops to change some (not all) params are provided\n    as well as cross-validation scoring.\n    Scoring and other stastistics are saved to dataframe \"df_stat\" \n    It can be used to choose the best candidate for submission.\n    All the models generate submissions which are saved to csv\n    (To turn it off set: flag_save_each_model_submit = False) \n    \n    Other sections provide auxiliary code:\n    data load,     custom cross-validation scheme,     etc.\n    \n    The section \"  Outcomes of early versions \"\n    contains brief results of the tuning in the early versions of the notebook.\n    The main outcome seems to be - use target encoding instead of one-hot or ordinal.\n    We have not tuned much for target encoders - welcome to do it - improve the score. \n    \n#### Other tutorials\n\nTutorial_1_Basics for simple usage examples https://github.com/AILab-MLTools/Py-Boost/blob/master/tutorials/Tutorial_1_Basics.ipynb\n\nTutorial_2_Advanced_multioutput for advanced multioutput features https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb\n\nTutorial_3_Custom_features for examples of customization https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_3_Custom_features.ipynb\n\n#### Citation (NeurIPS 2022) and Github\n\nhttps://openreview.net/forum?id=WSxarC8t-T\n\nIosipoi, Leonid, and Anton Vakhrushev. \"SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems.\" Advances in Neural Information Processing Systems 35 (2022): 25422-25435.\n\n\nGithub: https://github.com/sb-ai-lab/Py-Boost\n\n","metadata":{}},{"cell_type":"markdown","source":"## Preliminaries","metadata":{}},{"cell_type":"code","source":"import time\nt0start = time.time() \n\nimport random \nimport joblib\n\nimport numpy as np\nimport pandas as pd \nimport tensorflow as tf\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport category_encoders as ce\n\nfrom sklearn.datasets import make_regression\nfrom sklearn.model_selection import KFold\n\nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-11-16T19:02:33.745782Z","iopub.execute_input":"2023-11-16T19:02:33.746038Z","iopub.status.idle":"2023-11-16T19:02:43.433579Z","shell.execute_reply.started":"2023-11-16T19:02:33.746012Z","shell.execute_reply":"2023-11-16T19:02:43.432747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PY-BOOST Install (GPU only)","metadata":{}},{"cell_type":"code","source":"# that works for CPU also\n!pip install py-boost","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:04:45.153395Z","iopub.execute_input":"2023-11-16T19:04:45.154114Z","iopub.status.idle":"2023-11-16T19:04:57.585815Z","shell.execute_reply.started":"2023-11-16T19:04:45.154082Z","shell.execute_reply":"2023-11-16T19:04:57.584694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import will work only for GPU","metadata":{}},{"cell_type":"code","source":"import os\n# Optional: set the device to run\nos.environ[\"CUDA_DEVICE_ORDER\"] = \"PCI_BUS_ID\"\nos.environ[\"CUDA_VISIBLE_DEVICES\"] = \"0\"\n\nos.makedirs('../data', exist_ok=True)\n\n# simple case - just one class is used\nfrom py_boost import GradientBoosting, TLPredictor, TLCompiledPredictor\nfrom py_boost.cv import CrossValidation","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:04:57.587677Z","iopub.execute_input":"2023-11-16T19:04:57.587964Z","iopub.status.idle":"2023-11-16T19:05:14.728769Z","shell.execute_reply.started":"2023-11-16T19:04:57.587940Z","shell.execute_reply":"2023-11-16T19:05:14.727577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check","metadata":{}},{"cell_type":"code","source":"%%time\nmodel = GradientBoosting('mse')\nmodel\n# model.fit(X,Y)\n# model.fit(X, y[:, 0], eval_sets=[{'X': X_test, 'y': y_test[:, 0]},])","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:14.733381Z","iopub.execute_input":"2023-11-16T19:05:14.734610Z","iopub.status.idle":"2023-11-16T19:05:14.745115Z","shell.execute_reply.started":"2023-11-16T19:05:14.734539Z","shell.execute_reply":"2023-11-16T19:05:14.744221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\ndisplay(df_de_train)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:14.747421Z","iopub.execute_input":"2023-11-16T19:05:14.747732Z","iopub.status.idle":"2023-11-16T19:05:16.835524Z","shell.execute_reply.started":"2023-11-16T19:05:14.747697Z","shell.execute_reply":"2023-11-16T19:05:16.834647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,4))\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs)\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20)\nplt.grid()\nplt.show()\ndisplay(v.head(15))","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:16.836603Z","iopub.execute_input":"2023-11-16T19:05:16.836908Z","iopub.status.idle":"2023-11-16T19:05:17.260950Z","shell.execute_reply.started":"2023-11-16T19:05:16.836884Z","shell.execute_reply":"2023-11-16T19:05:17.260115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\ndisplay(df_id_map)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:17.261999Z","iopub.execute_input":"2023-11-16T19:05:17.262268Z","iopub.status.idle":"2023-11-16T19:05:17.283062Z","shell.execute_reply.started":"2023-11-16T19:05:17.262244Z","shell.execute_reply":"2023-11-16T19:05:17.281976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\ndisplay(df_sample_submit)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:17.284356Z","iopub.execute_input":"2023-11-16T19:05:17.284742Z","iopub.status.idle":"2023-11-16T19:05:21.416729Z","shell.execute_reply.started":"2023-11-16T19:05:17.284690Z","shell.execute_reply":"2023-11-16T19:05:21.415913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nselected_features_columns = ['cell_type','sm_name']  \n\nX_submit_categorical = df_id_map[selected_features_columns] \nX_full_categorical = df_de_train[selected_features_columns]\nY_full = df_de_train.iloc[:,5:].values\n\nprint(f\"X_submit_categorical.shape: {X_submit_categorical.shape}\\n\", \n      f\"X_full_categorical.shape: {X_full_categorical.shape}\\n\",\n      f\"Y_full.shape: {Y_full.shape}\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.417922Z","iopub.execute_input":"2023-11-16T19:05:21.418207Z","iopub.status.idle":"2023-11-16T19:05:21.454971Z","shell.execute_reply.started":"2023-11-16T19:05:21.418183Z","shell.execute_reply":"2023-11-16T19:05:21.453901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load ChemBERT ","metadata":{}},{"cell_type":"code","source":"fn = '/kaggle/input/chemberta-v2-77-mtr/pad_false/train_ChemBERTa_v2_77MTR_mean_pad_False.npy'\nX_full_chembert = np.load(fn)\nprint(X_full_chembert.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.456424Z","iopub.execute_input":"2023-11-16T19:05:21.456820Z","iopub.status.idle":"2023-11-16T19:05:21.485589Z","shell.execute_reply.started":"2023-11-16T19:05:21.456786Z","shell.execute_reply":"2023-11-16T19:05:21.484609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn = '/kaggle/input/chemberta-v2-77-mtr/pad_false/test_ChemBERTa_v2_77MTR_mean_pad_False.npy'\nX_submit_chembert = np.load(fn)\nprint(X_submit_chembert.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.488341Z","iopub.execute_input":"2023-11-16T19:05:21.488615Z","iopub.status.idle":"2023-11-16T19:05:21.502513Z","shell.execute_reply.started":"2023-11-16T19:05:21.488592Z","shell.execute_reply":"2023-11-16T19:05:21.501661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes\n\nSee the notebook: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\n\nMore analysis of the CV schemes can be found in slides: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing and sheet: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n\n\nCV schemes (see discussion https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494 ):\n\nAmbrosM: https://www.kaggle.com/code/ambrosm/scp-quickstart\n\nMT: https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy\n\nKishan's notebook: https://www.kaggle.com/code/liudacheldieva/neural-network-regression/notebook#Predicting-on-test-data ","metadata":{}},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', \n                  'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', 'Belinostat', \n                  'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', \n                  'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021'\n                 ]\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \n\nfolds_index_data_AmbrosM = []\nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.503895Z","iopub.execute_input":"2023-11-16T19:05:21.504157Z","iopub.status.idle":"2023-11-16T19:05:21.515169Z","shell.execute_reply.started":"2023-11-16T19:05:21.504134Z","shell.execute_reply":"2023-11-16T19:05:21.514211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {\n    0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', \n        'O-Demethylated Adapalene'],\n    1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', \n        'Porcn Inhibitor III'],\n    2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  \n        'R428']}\n\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    folds_index_data_MT.append( [IX_train, IX_test ]) \n    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.516309Z","iopub.execute_input":"2023-11-16T19:05:21.516579Z","iopub.status.idle":"2023-11-16T19:05:21.527694Z","shell.execute_reply.started":"2023-11-16T19:05:21.516549Z","shell.execute_reply":"2023-11-16T19:05:21.526945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [[np.arange(0), np.arange(0), np.arange(0)]] # Example data\n        self.n_splits = 1 # Example data\n        \n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [[IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]]\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list(kf.split( np.arange(614)))\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [[IX_train_full, IX_test_full]]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))) , item[2] # Return train, test only \n            elif self.CV_scheme in ['AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.528934Z","iopub.execute_input":"2023-11-16T19:05:21.529230Z","iopub.status.idle":"2023-11-16T19:05:21.549498Z","shell.execute_reply.started":"2023-11-16T19:05:21.529207Z","shell.execute_reply":"2023-11-16T19:05:21.548802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \nprint(CV_scheme)\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3]\")\n\nfor i_fold,(IX_train,IX_valid, IX_test) in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train), len(IX_valid),len(IX_test), \n          IX_train[:3],IX_valid[:3], IX_test[:3])\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n\n    print(CV_scheme, 'valid_size', valid_size)\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3]\")\n    \n    for i_fold,(IX_train,IX_test) in enumerate( kf.split(None)):\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3])\n    print()\n    \n    valid_size = 0.2 \n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n\n    print(CV_scheme, 'valid_size', valid_size)\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3]\")\n    for i_fold,(IX_train, IX_valid, IX_test) in enumerate(kf.split(None)):\n        print(i_fold, len(IX_train), len(IX_valid), len(IX_test), \n              IX_train[:3],IX_valid[:3], IX_test[:3])\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.550576Z","iopub.execute_input":"2023-11-16T19:05:21.550850Z","iopub.status.idle":"2023-11-16T19:05:21.572536Z","shell.execute_reply.started":"2023-11-16T19:05:21.550818Z","shell.execute_reply":"2023-11-16T19:05:21.571472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Embeddings","metadata":{}},{"cell_type":"markdown","source":"Best combination for this moment:\n\n***OneHot cell embeddings + TargetEncoding drugs + ChemBERT -- 0.6 (without tuning), 0.597 (with tuning)***\n\nOther results:\n\n- TargetEncoding cell + drugs + ChemBERT -- 0.601\n- OneHot cell embeddings + TargetEncoding drugs -- 0.602\n- TargetEncoding cell + drugs -- 0.604\n- OneHot cell embeddings + ChemBERT -- 0.658 (seems TargetEncoding for drug is very important)","metadata":{}},{"cell_type":"code","source":"def create_embeddings(X_train_categorical, X_test_categorical, X_submit_categorical_loc, \n                      X_train_chembert, X_test_chembert, X_submit_chembert,\n                      Y_train_red, enc_cell, enc_drug):\n    ### -------------- Target Encoding -------------------------------------------------------------------\n    # Encoding is done by each of the reduced targets \n    # One can add other features here\n    # category encoders package can accept only 1-dimensional \"y\", so we need a loop and  encode using Y[:,i] one by one for all \"i\" \n\n    for i_target in range(Y_train_red.shape[1]):\n        if i_target == 0:\n            X_train_encoded = enc_drug.fit_transform(X_train_categorical[:, 1], Y_train_red[:,i_target])\n            X_test_encoded = enc_drug.transform(X_test_categorical[:, 1])\n            X_submit_encoded = enc_drug.transform(X_submit_categorical_loc[:, 1])\n        else:\n            X_encoded_tmp = enc_drug.fit_transform(X_train_categorical[:, 1], Y_train_red[:,i_target])\n            X_train_encoded = np.concatenate([X_train_encoded, X_encoded_tmp], axis = 1)\n            X_encoded_tmp = enc_drug.transform(X_test_categorical[:, 1])\n            X_test_encoded = np.concatenate([X_test_encoded, X_encoded_tmp], axis = 1)\n            X_encoded_tmp = enc_drug.transform(X_submit_categorical_loc[:, 1])\n            X_submit_encoded = np.concatenate([X_submit_encoded, X_encoded_tmp], axis = 1)\n\n    X_encoded_tmp = enc_cell.fit_transform(X_train_categorical[:, 0])\n    X_train_encoded = np.concatenate([X_train_encoded, X_encoded_tmp], axis = 1)\n    X_encoded_tmp = enc_cell.transform(X_test_categorical[:, 0])\n    X_test_encoded = np.concatenate([X_test_encoded, X_encoded_tmp], axis = 1)\n    X_encoded_tmp = enc_cell.transform(X_submit_categorical_loc[:, 0])\n    X_submit_encoded = np.concatenate([X_submit_encoded, X_encoded_tmp], axis = 1)\n\n#         X_train_encoded = enc_cell.fit_transform(X_train_categorical[:, 0])\n#         X_test_encoded = enc_cell.transform(X_test_categorical[:, 0])\n#         X_submit_encoded = enc_cell.transform(X_submit_categorical_loc[:, 0])\n\n    X_train_encoded = np.concatenate([X_train_encoded, X_train_chembert], axis = 1)\n    X_test_encoded = np.concatenate([X_test_encoded, X_test_chembert], axis = 1)\n    X_submit_encoded = np.concatenate([X_submit_encoded, X_submit_chembert], axis = 1)\n\n    return X_train_encoded, X_test_encoded, X_submit_encoded","metadata":{"execution":{"iopub.status.busy":"2023-11-16T19:05:21.574010Z","iopub.execute_input":"2023-11-16T19:05:21.574324Z","iopub.status.idle":"2023-11-16T19:05:21.585282Z","shell.execute_reply.started":"2023-11-16T19:05:21.574301Z","shell.execute_reply":"2023-11-16T19:05:21.584329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"markdown","source":"## Cross Validation","metadata":{}},{"cell_type":"code","source":"flag_save_each_model_submit = True\nverbose = 1000\n\nCV_scheme = 'AmbrosM' #  'Random'\nvalid_size = 0 \n# kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\ndf_stat = pd.DataFrame(); # report with scores will be here ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T11:57:28.842760Z","iopub.execute_input":"2023-11-16T11:57:28.843176Z","iopub.status.idle":"2023-11-16T11:57:28.849561Z","shell.execute_reply.started":"2023-11-16T11:57:28.843145Z","shell.execute_reply":"2023-11-16T11:57:28.848353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def save_submit(Y_submit_pred, str_fn_info,  dict_params = {} ):\n    '''\n    Save submit to csv - optionally include into filename the params of the model.\n    First column should be named \"id\" - do not forget.\n    Data - 255x18211: samples x genes\n    '''\n    \n    df_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n    df_submit.index.name = 'id'\n    print( df_submit.shape )\n    display(df_submit.head(2))\n    str_detailed_info = ''\n    for k in dict_params:\n        val = dict_params[k]\n        str_detailed_info += '_'+str(k)+str(val).replace('.','')\n    fn_save = 'submission_' + str_fn_info + str_detailed_info + '.csv' \n    print(fn_save)\n    df_submit.to_csv(fn_save)    \n    \ndef test_model(n_components, max_depth, ntrees, subsample, colsample, lr, enc_cell, enc_drug):\n    # Submission will be averaged submission from each fold\n    Y_submit_pred = np.zeros((255,  Y_full.shape[1])) \n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\n    i_blend_for_submit = 0 \n    list_mrrmse, list_r2, list_mae = [], [], []\n    t0_single_model_all_folds = time.time() \n\n    # Initialize CV-scheme generator:\n    kf = KFold_custom(CV_scheme, valid_size = valid_size) \n    \n    #  ------------------------------------------ Loop over folds  ------------------------------------------\n    for i_fold,(IX_train,IX_test) in enumerate( kf.split()) :\n\n        cd8_loc = df_de_train['cell_type'] == 'T cells CD8+'\n        cd8_idx = list(np.where(cd8_loc > 0)[0])\n        IX_train = [i for i in IX_train if i not in cd8_idx]\n\n        # ------------------- Get train, test. (Also technical: coversion to numpy if necessary) -------------------\n        if hasattr(X_full_categorical,'values'):\n            X_train_categorical,X_test_categorical = X_full_categorical.values[IX_train], X_full_categorical.values[IX_test]\n        else:\n            X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n        \n        if hasattr(X_submit_categorical,'values'):\n            X_submit_categorical_loc = X_submit_categorical.values\n        else:\n            X_submit_categorical_loc = X_submit_categorical\n\n        X_train_chembert, X_test_chembert = X_full_chembert[IX_train], X_full_chembert[IX_test]\n        #  Dimensional reduction for targets  \n        Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n\n        # ------------------------------------------ Initialize the model  ------------------------------------------\n        # Pay attention - it is important to make initialization of py-boost for each fold\n\n        # Use Ridge for simple/fast tests and debuging:\n        #model = Ridge(); str_model_id = 'Ridge'\n\n        str_model_id = 'Pyboost'; \n        model = GradientBoosting('mse', ntrees=ntrees, lr=lr,  max_depth=max_depth,\n                                 subsample=subsample, colsample=colsample, \n                                 min_data_in_leaf=1, min_gain_to_split=0, verbose=100)           \n        if verbose >= 100:\n            print('i_total, n_components, ntrees, lr, max_depth, colsample, subsample')\n            print(i_total, n_components,  ntrees, lr, max_depth, colsample, subsample  )\n        \n        X_train_encoded, X_test_encoded, X_submit_encoded = create_embeddings(\n            X_train_categorical, X_test_categorical, X_submit_categorical_loc, \n            X_train_chembert, X_test_chembert, X_submit_chembert,\n            Y_train_red, enc_cell, enc_drug\n        )\n        \n        if verbose >= 10:\n            print('fold:', i_fold)\n            print('X_train_encoded.shape, X_test_encoded.shape,  X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape')\n            print(X_train_encoded.shape, X_test_encoded.shape, X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape)\n\n        ### ---------------------------------- Train the model -----------------------------------------------------------------------\n        model.fit(X_train_encoded, Y_train_red)\n\n        ### ---------------------------------- Prediction for validation  -------------------------------------------------------------\n        Y_test_pred_red = model.predict(X_test_encoded)\n        \n        ### ---------------------------------- Inverse dimensional reduction  ---------------------------------------------------------\n        Y_test_pred = reducer.inverse_transform(Y_test_pred_red)\n\n        ### ---------------------------------- Prediction, Inverse-reducer  for Submit part---------------------------------------------\n        Y_submit_pred_from_one_fold = reducer.inverse_transform(model.predict(X_submit_encoded))\n        ### ---------------------------------- Average submit part from all folds          ---------------------------------------------\n        Y_submit_pred =  (Y_submit_pred * i_blend_for_submit + Y_submit_pred_from_one_fold )/(i_blend_for_submit + 1)\n        i_blend_for_submit += 1\n\n        ### ---------------------------------- Scoring computation for local test ---------------------------------------------\n        Y_test = Y_full[IX_test]\n        mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ); mae_1_fold = mean_absolute_error( Y_test , Y_test_pred  )\n        if verbose >= 10:\n            print(i_total, \n                  'fold:', i_fold, \n                  'mrrmse:', np.round(mrrmse,4), \n                  'r2:', np.round(r2,4),  \n                  'mae:', np.round(mae_1_fold,4))\n        list_mrrmse.append(mrrmse); list_r2.append(r2); list_mae.append(mae_1_fold)\n\n        # ------------------------------------ END of fold-loop -----------------------------------------------------------------------\n\n\n\n    # ------------------------------------------------ Optional. Save Scores, other statistics ----------------------------------------    \n    # Statistics/reporting part    \n    # Save scores,params to the report datataframe    \n    if verbose >= 1:\n        print(i_total, 'Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ),  'Average mae:', np.round(np.mean(list_r2 ),4 ))\n    IX = len(df_stat)+1\n    df_stat.loc[IX,'mrrmse'] =  np.mean(list_mrrmse )\n    df_stat.loc[IX,'r2'] =  np.mean(list_r2 ) \n    df_stat.loc[IX,'mae'] =  np.mean(list_mae ) \n    df_stat.loc[IX,'max_depth'] =  max_depth\n    df_stat.loc[IX,'ntrees'] = ntrees\n    df_stat.loc[IX,'colsample'] =  colsample\n    df_stat.loc[IX,'lr'] = lr\n    df_stat.loc[IX,'n_components'] = n_components\n    df_stat.loc[IX,'subsample'] =   subsample\n    for i_tmp in range(len(list_mrrmse)):\n        df_stat.loc[IX,'mrrmse ' + CV_scheme + ' ' + str(i_tmp)] = list_mrrmse[i_tmp]\n    for i_tmp in range(len(list_r2)):\n        df_stat.loc[IX,'r2 ' + CV_scheme + ' ' + str(i_tmp)] = list_r2[i_tmp]\n    for i_tmp in range(len(list_mae)):\n        df_stat.loc[IX,'mae ' + CV_scheme + ' ' + str(i_tmp)] = list_mae[i_tmp]\n    df_stat.loc[IX,'Time'] = np.round(time.time() - t0_single_model_all_folds,1)  \n    if verbose >= 1:\n        display(df_stat.tail(1))\n    if i_total % 20 == 0:\n        df_stat.sort_values(df_stat.columns[0] ).to_csv('df_stat.csv')\n\n\n    # ------------------------------------------------ Optional. Save Submission to CSV ----------------------------------------    \n    if flag_save_each_model_submit :\n        dict_params_loc = {'max_depth': max_depth, \n                           'ntrees': ntrees, \n                           'lr': lr, \n                           'subsample': subsample, \n                           'colsample': colsample, \n                           'n_components': n_components}\n        save_submit(Y_submit_pred, str_fn_info = str_model_id, dict_params = dict_params_loc )\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T11:57:29.021685Z","iopub.execute_input":"2023-11-16T11:57:29.022245Z","iopub.status.idle":"2023-11-16T11:57:29.051095Z","shell.execute_reply.started":"2023-11-16T11:57:29.022211Z","shell.execute_reply":"2023-11-16T11:57:29.050087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\ni_total = 0\n# Optional loops for params tuning - hand-made Grid-Search: \nfor n_components in [50]: # [30,50,100]:\n    # ------------------------------------------  Encoder   ------------------------------------------\n    # One can tune params for encoders also. In current version we will NOT do it.\n    #list_smoothing = [0,1,5,8,9,10,11,12,15, 20,50,100,200,500, 1000,10_000, np.inf] # For Target Encoder\n    #for smoothing in list_smoothing:\n        # enc = ce.TargetEncoder(smoothing = smoothing )\n        # print('smoothing:', smoothing, )\n        \n#     enc_drug = ce.LeaveOneOutEncoder()\n    enc_drug = ce.QuantileEncoder(quantile =.8)\n    enc_cell = ce.OneHotEncoder()\n        \n    for max_depth in [10]: #  \n        for ntrees in [5000]: # \n            for subsample in [1]: # From 0 to 1\n                for colsample in [0.35]: # From 0 to 1 \n                    for lr in [0.01]:#  From 0 to 1 \n                        test_model(n_components, max_depth, ntrees, subsample, colsample, lr, \n                                   enc_cell, enc_drug)\n                        i_total += 1\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T11:57:29.430964Z","iopub.execute_input":"2023-11-16T11:57:29.431830Z","iopub.status.idle":"2023-11-16T11:58:15.728887Z","shell.execute_reply.started":"2023-11-16T11:57:29.431796Z","shell.execute_reply":"2023-11-16T11:58:15.727907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Full Training","metadata":{}},{"cell_type":"code","source":"# n_components = 50\n# max_depth = 10 \n# ntrees = 1000 \n# subsample = 1 \n# colsample = 0.2 \n# lr = 0.01 \n\n# enc_cell = ce.OneHotEncoder()\n\n# Y_submit_pred = np.zeros((255,  Y_full.shape[1])) \n# reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\n# Y_train_red = reducer.fit_transform(Y_full) \n\n# str_model_id = 'Pyboost'; \n# model = GradientBoosting('mse', ntrees=ntrees, lr=lr, max_depth=max_depth,\n#                          subsample=subsample, colsample=colsample, min_data_in_leaf=1, \n#                          min_gain_to_split=0, verbose=100)    \n\n# X_train_encoded = enc_cell.fit_transform(X_full_categorical.values[:, 0])\n# X_submit_encoded = enc_cell.transform(X_submit_categorical.values[:, 0])\n\n# X_train_encoded = np.concatenate([X_train_encoded, X_full_chembert], axis = 1)\n# X_submit_encoded = np.concatenate([X_submit_encoded, X_submit_chembert], axis = 1)\n\n# model.fit(X_train_encoded, Y_train_red)\n# Y_submit_pred = reducer.inverse_transform(model.predict(X_submit_encoded))\n\n# df_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n# df_submit.index.name = 'id'\n# df_submit.to_csv('submission_.csv')    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T11:58:15.730775Z","iopub.execute_input":"2023-11-16T11:58:15.731069Z","iopub.status.idle":"2023-11-16T11:58:15.736247Z","shell.execute_reply.started":"2023-11-16T11:58:15.731044Z","shell.execute_reply":"2023-11-16T11:58:15.735218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_stat.sort_values(df_stat.columns[0]))\ndf_stat.sort_values(df_stat.columns[0]).to_csv('df_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-16T11:58:15.737423Z","iopub.execute_input":"2023-11-16T11:58:15.737710Z","iopub.status.idle":"2023-11-16T11:58:15.819733Z","shell.execute_reply.started":"2023-11-16T11:58:15.737687Z","shell.execute_reply":"2023-11-16T11:58:15.818443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show CV scores, etc","metadata":{}},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\ndisplay(df_stat.sort_values(df_stat.columns[0]))","metadata":{"execution":{"iopub.status.busy":"2023-11-16T10:07:33.485996Z","iopub.execute_input":"2023-11-16T10:07:33.486755Z","iopub.status.idle":"2023-11-16T10:07:33.511547Z","shell.execute_reply.started":"2023-11-16T10:07:33.486706Z","shell.execute_reply":"2023-11-16T10:07:33.510649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}