{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"}],"dockerImageVersionId":30559,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nEducational notebook on the new breakthough gradient boosting package \"PYBOOST\". \n\n### PYBOOST - is the best boosting for MULTI-target tasks\n\nIt designed especially for MULTI-Target(!) tasks.  For such tasks it consitenly outperforms the other boostings,\nas well as incomparably faster. \n\n#### PYBOOST -> Gold on Kaggle\n\nPyboost is created by team lead by the Kaggle Grandmaster - Btbpanda - Anton Vakhrushev  https://www.kaggle.com/btbpanda\n\nHe and others got many gold medals on Kaggle: recent CAFA5 (https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/434064 ),\nlast year \"Open Problems\" top7-team \"Chromosom\", and public top2 team, etc.\n\nSome inpendent check - PyBoost outperaforming other boostings:  \"Table 2. Performance of Kinase Activity Prediction Modelsa\"\nhttps://pubs.acs.org/doi/full/10.1021/acs.jcim.3c00132\n\n#### PYBOOST - easy to use - similar interface and params as other boostings and sklearn models\n\n    model = GradientBoosting('mse')\n    model.fit(X,Y)\n    model.predict(X)\n    \n    Main params are similar to other boostings:\n    # ntrees - number of trees (iterations)\n    # lr  - learning rate\n    # max_depth - maximal depth for trees\n    # subsample - subsampling fraction\n    # colsample - subsampleling fraction for columns\n    # Etc \n    # Example:\n    model = GradientBoosting('mse', ntrees=ntrees, lr=lr,  max_depth=max_depth  , subsample=subsample, colsample=colsample, min_data_in_leaf=1, min_gain_to_split=0, verbose=100  )\n    \n\n#### OP2 - easily LB-score 0.604 -  welcome improve it \n\nThe present notebooks shares a Pyboost prediction with LB 0.604 sore - one of the best public solo-model solutions for the moment.\n\nWelcome to improve it with more tuning.  Notebook provides simple framework to tune - just change the intervals of the params - and run it.\nSubmissions will be saved to CSV files. Try: more iterations, tune colsample, other params, encoders, average several random seeds, etc...\nThe default tuning is done by \"AmbrosM\" validation scheme, you can change it to MT, Random or any other scheme. The LB - CV correspondence is poor,\nbut still not absent. \n\nThe LB604 file prediction can be found in Kaggle dataset: \nhttps://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection?select=LB604_PyboostLeaveOneOutTsvd50maxdepth10ntrees1000lr001subsample1colsample02_Alex_nbV19.csv\nor uploaded from the present notebook version 19:\nhttps://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool/output?scriptVersionId=149987684&select=submission_Pyboost_max_depth10_ntrees1000_lr001_subsample1_colsample02_n_components50.csv\n\n\n#### PYBOOST - is for GPU\n\nCurrent version of Pyboost requires GPU.\n\n#### Organization of the notebook\n\n    Section: \" PY-BOOST Install (GPU only) \" \n    Shows how to install/import pyboost - do not forget to turn on GPU.\n\n    Section: \" Modeling \"\n    Contains the main modeling example.\n    One can see how simple is to use the pyboost.\n    \n    Also:\n    \n    a kind of grid-search loops to change some (not all) params are provided\n    as well as cross-validation scoring.\n    Scoring and other stastistics are saved to dataframe \"df_stat\" \n    It can be used to choose the best candidate for submission.\n    All the models generate submissions which are saved to csv\n    (To turn it off set: flag_save_each_model_submit = False) \n    \n    Other sections provide auxiliary code:\n    data load,     custom cross-validation scheme,     etc.\n    \n    The section \"  Outcomes of early versions \"\n    contains brief results of the tuning in the early versions of the notebook.\n    The main outcome seems to be - use target encoding instead of one-hot or ordinal.\n    We have not tuned much for target encoders - welcome to do it - improve the score. \n    \n#### Other tutorials\n\nTutorial_1_Basics for simple usage examples https://github.com/AILab-MLTools/Py-Boost/blob/master/tutorials/Tutorial_1_Basics.ipynb\n\nTutorial_2_Advanced_multioutput for advanced multioutput features https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb\n\nTutorial_3_Custom_features for examples of customization https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_3_Custom_features.ipynb\n\n#### Citation (NeurIPS 2022)\n\nhttps://openreview.net/forum?id=WSxarC8t-T\n\nIosipoi, Leonid, and Anton Vakhrushev. \"SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems.\" Advances in Neural Information Processing Systems 35 (2022): 25422-25435.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Preliminaries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-11-11T18:15:12.485626Z","iopub.execute_input":"2023-11-11T18:15:12.486029Z","iopub.status.idle":"2023-11-11T18:15:12.495361Z","shell.execute_reply.started":"2023-11-11T18:15:12.485997Z","shell.execute_reply":"2023-11-11T18:15:12.49451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PY-BOOST Install (GPU only)","metadata":{}},{"cell_type":"code","source":"# that works for CPU also\n!pip install py-boost","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:12.732674Z","iopub.execute_input":"2023-11-11T18:15:12.733096Z","iopub.status.idle":"2023-11-11T18:15:25.31533Z","shell.execute_reply.started":"2023-11-11T18:15:12.733065Z","shell.execute_reply":"2023-11-11T18:15:25.314197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import will work only for GPU","metadata":{}},{"cell_type":"code","source":"import os\n# Optional: set the device to run\nos.environ[\"CUDA_DEVICE_ORDER\"] = \"PCI_BUS_ID\"\nos.environ[\"CUDA_VISIBLE_DEVICES\"] = \"0\"\n\nos.makedirs('../data', exist_ok=True)\n\nimport joblib\nfrom sklearn.datasets import make_regression\nimport numpy as np\n\n# simple case - just one class is used\nfrom py_boost import GradientBoosting, TLPredictor, TLCompiledPredictor\nfrom py_boost.cv import CrossValidation","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:25.317626Z","iopub.execute_input":"2023-11-11T18:15:25.317969Z","iopub.status.idle":"2023-11-11T18:15:25.324862Z","shell.execute_reply.started":"2023-11-11T18:15:25.317937Z","shell.execute_reply":"2023-11-11T18:15:25.323892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check","metadata":{}},{"cell_type":"code","source":"%%time\nmodel = GradientBoosting('mse')\nmodel\n# model.fit(X,Y)\n# model.fit(X, y[:, 0], eval_sets=[{'X': X_test, 'y': y_test[:, 0]},])","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:25.325904Z","iopub.execute_input":"2023-11-11T18:15:25.326196Z","iopub.status.idle":"2023-11-11T18:15:25.342157Z","shell.execute_reply.started":"2023-11-11T18:15:25.32617Z","shell.execute_reply":"2023-11-11T18:15:25.341172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )\n\n\n#features_columns = ['cell_type','sm_name']\nselected_features_columns = ['cell_type','sm_name']  # [ 'sm_name'] # What features to consider - others not used \n    # For OP2 task we have two features - cell-type and compound - both categorical\n    # Some simple models (like Ridge) better work when one uses only one feature - 'sm_name' (=\"compound\") - so you can try that option. \n    # But seems pyboost is powerful enough to proper work with the both features.\n    # Table with results for Ridge can be found here: https://www.kaggle.com/code/alexandervc/op2-target-encoders?scriptVersionId=148576642&cellId=1\n    # Or here: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing (sheet \"Encoders\") and discussed here:\n    \nX_submit_categorical = df_id_map[selected_features_columns] # pd.DataFrame(df_id_map, columns= features_columns )\nX_full_categorical = df_de_train[ selected_features_columns ]\n\nY_full = df_de_train.iloc[:,5:].values\nprint('X_submit_categorical.shape, X_full_categorical.shape ,  Y_full.shape', X_submit_categorical.shape, X_full_categorical.shape ,  Y_full.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:25.343501Z","iopub.execute_input":"2023-11-11T18:15:25.343827Z","iopub.status.idle":"2023-11-11T18:15:31.132843Z","shell.execute_reply.started":"2023-11-11T18:15:25.343791Z","shell.execute_reply":"2023-11-11T18:15:31.131717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes\n\nSee the notebook: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\n\nMore analysis of the CV schemes can be found in slides: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing and sheet: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n\n\nCV schemes (see discussion https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494 ):\n\nAmbrosM: https://www.kaggle.com/code/ambrosm/scp-quickstart\n\nMT: https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy\n\nKishan's notebook: https://www.kaggle.com/code/liudacheldieva/neural-network-regression/notebook#Predicting-on-test-data ","metadata":{}},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ]) \n    \nimport random \nfrom sklearn.model_selection import KFold\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n        elif 'Full'.lower() in CV_scheme.lower():\n            # Just return the full set as both train and test - it useful to re-train the model on the entire data\n            IX_train_full = np.arange(614)\n            IX_test_full = np.arange(614)\n            self.n_splits = 1\n            self.folds_index_data = [   [IX_train_full, IX_test_full]  ]\n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['AmbrosM','MT', 'Random', 'Full']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n        \n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:31.135999Z","iopub.execute_input":"2023-11-11T18:15:31.136328Z","iopub.status.idle":"2023-11-11T18:15:31.197487Z","shell.execute_reply.started":"2023-11-11T18:15:31.136298Z","shell.execute_reply":"2023-11-11T18:15:31.196343Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Auxiliary function - save submit to csv","metadata":{}},{"cell_type":"code","source":"def save_submit(Y_submit_pred, str_fn_info,  dict_params = {} ):\n    '''\n    Save submit to csv - optionally include into filename the params of the model.\n    First column should be named \"id\" - do not forget.\n    Data - 255x18211: samples x genes\n    '''\n    \n    df_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\n    df_submit.index.name = 'id'\n    print( df_submit.shape )\n    display(df_submit.head(2))\n    str_detailed_info = ''\n    for k in dict_params:\n        val = dict_params[k]\n        str_detailed_info += '_'+str(k)+str(val).replace('.','')\n    fn_save = 'submission_' + str_fn_info + str_detailed_info + '.csv' \n    print(fn_save)\n    df_submit.to_csv(fn_save)    ","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:31.198892Z","iopub.execute_input":"2023-11-11T18:15:31.199307Z","iopub.status.idle":"2023-11-11T18:15:31.207206Z","shell.execute_reply.started":"2023-11-11T18:15:31.199269Z","shell.execute_reply":"2023-11-11T18:15:31.206104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_absolute_error\nimport category_encoders as ce\n\n\nflag_save_each_model_submit = True\n\nverbose = 1000\n\nCV_scheme = 'AmbrosM' #  'Random'\nprint('CV_scheme:', CV_scheme); print(); valid_size = 0 \n# kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\n\ndf_stat = pd.DataFrame(); # report with scores will be here \ni_total = 0\n# Optional loops for params tuning - hand-made Grid-Search: \nfor n_components in [50]: # [30,50,100]:\n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\n    # ------------------------------------------  Encoder   ------------------------------------------\n    # One can tune params for encoders also. In current version we will NOT do it.\n    smoothing = 1 # [0,1,5,8,9,10,]#11,12,15, 20,50,100,200,500, 1000,10_000, np.inf] # For Target Encoder\n    \n    enc = ce.QuantileEncoder(quantile =.8)\n    print('smoothing:', smoothing, )\n    #enc = ce.LeaveOneOutEncoder(  )\n        \n    for max_depth in [10]: #  \n        for ntrees in [5000]: # \n            for subsample in [1]: # From 0 to 1\n                for colsample in [0.35]: # From 0 to 1 \n                    for lr in [0.01]:#  From 0 to 1 \n                        i_total += 1\n\n                        # Submission will be averaged submission from each fold\n                        Y_submit_pred = np.zeros( (255,  Y_full.shape[1])   ); i_blend_for_submit = 0 \n\n                        # Fold-wise metrics will be stored here:\n                        list_mrrmse = [];list_r2 = []; list_mae = []  # Fold-wise metrics\n                        \n                        t0_single_model_all_folds = time.time() # Timer starts. All folds computation timing.\n\n                        # Initialize CV-scheme generator:\n                        kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n                        #  ------------------------------------------ Loop over folds  ------------------------------------------\n                        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n                            t0_one_fold = time.time()\n                            \n                            m_tmp1 = df_de_train['cell_type'] == 'T cells CD8+'\n                            l_tmp1 = list(np.where(m_tmp1>0)[0])\n                            IX_train = [i for i in IX_train if i not in l_tmp1]\n                            IX_train_basic = IX_train.copy() \n                            n_multiplex = 10\n                            for i in range(n_multiplex): \n                                IX_train += IX_train_basic\n                            IX_train = np.array( IX_train )\n                           \n                            \n   \n                            # ------------------- Get train, test. (Also technical: coversion to numpy if necessary) -------------------\n                            if hasattr(X_full_categorical,'values'):\n                                X_train_categorical,X_test_categorical = X_full_categorical.values[IX_train], X_full_categorical.values[IX_test]\n                            else:\n                                X_train_categorical,X_test_categorical = X_full_categorical[IX_train], X_full_categorical[IX_test]\n                            if hasattr(X_submit_categorical,'values'):\n                                X_submit_categorical_loc = X_submit_categorical.values\n                            else:\n                                X_submit_categorical_loc = X_submit_categorical\n\n\n                            #  Dimensional reduction for targets  \n                            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n\n\n                            # ------------------------------------------ Initialize the model  ------------------------------------------\n                            # Pay attention - it is important to make initialization of py-boost for each fold\n\n                            # Use Ridge for simple/fast tests and debuging:\n                            #model = Ridge(); str_model_id = 'Ridge'\n\n                            str_model_id = 'Pyboost'; \n                            model = GradientBoosting('mse', ntrees=ntrees, lr=lr,  max_depth=max_depth  , subsample=subsample, colsample=colsample, min_data_in_leaf=1, min_gain_to_split=0, verbose=100  )           \n                            if verbose >= 100:\n                                print('i_total, n_components, ntrees, lr, max_depth, colsample, subsample', i_total, n_components,  ntrees, lr, max_depth, colsample, subsample  )\n                            \n\n\n                            ### -------------- Target Encoding -------------------------------------------------------------------\n                            # Encoding is done by each of the reduced targets \n                            # One can add other features here\n                            # category encoders package can accept only 1-dimensional \"y\", so we need a loop and  encode using Y[:,i] one by one for all \"i\" \n                            for i_target in range(Y_train_red.shape[1]):\n                                if i_target == 0:\n                                    X_train_encoded = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                                    X_test_encoded = enc.transform(X_test_categorical)\n                                    X_submit_encoded = enc.transform(X_submit_categorical_loc)\n                                else:\n                                    X_encoded_tmp = enc.fit_transform(X_train_categorical, Y_train_red[:,i_target])\n                                    X_train_encoded = np.concatenate( [X_train_encoded, X_encoded_tmp], axis = 1)\n                                    X_encoded_tmp = enc.transform(X_test_categorical)\n                                    X_test_encoded = np.concatenate( [X_test_encoded, X_encoded_tmp], axis = 1)\n                                    X_encoded_tmp = enc.transform(X_submit_categorical_loc)\n                                    X_submit_encoded = np.concatenate( [X_submit_encoded, X_encoded_tmp], axis = 1)\n\n                            if verbose >= 10:\n                                print('fold:', i_fold, 'X_train_encoded.shape,X_test_encoded.shape,  X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape', X_train_encoded.shape,X_test_encoded.shape, X_train_categorical.shape, X_test_categorical.shape, Y_train_red.shape)\n\n\n                            ### ---------------------------------- Train the model -----------------------------------------------------------------------\n                            model.fit(X_train_encoded, Y_train_red)\n                            \n                            ### ---------------------------------- Prediction for validation  -------------------------------------------------------------\n                            Y_test_pred_red = model.predict( X_test_encoded )\n                            ### ---------------------------------- Inverse dimensional reduction  ---------------------------------------------------------\n                            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n                            \n                            ### ---------------------------------- Prediction, Inverse-reducer  for Submit part---------------------------------------------\n                            Y_submit_pred_from_one_fold =  reducer.inverse_transform( model.predict( X_submit_encoded ) )\n                            ### ---------------------------------- Average submit part from all folds          ---------------------------------------------\n                            Y_submit_pred =  (Y_submit_pred * i_blend_for_submit + Y_submit_pred_from_one_fold )/(i_blend_for_submit + 1)\n                            i_blend_for_submit += 1\n\n                            ### ----------------------------------      Scoring computation for local test     ---------------------------------------------\n                            Y_test = Y_full[IX_test]\n                            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ); mae_1_fold = mean_absolute_error( Y_test , Y_test_pred  )\n                            if verbose >= 10:\n                                print(i_total, 'fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  'mae:', np.round(mae_1_fold,4),  )\n                            list_mrrmse.append(mrrmse); list_r2.append(r2); list_mae.append(mae_1_fold)\n\n                            # ------------------------------------ END of fold-loop -----------------------------------------------------------------------\n                            \n\n                            \n                        # ------------------------------------------------ Optional. Save Scores, other statistics ----------------------------------------    \n                        # Statistics/reporting part    \n                        # Save scores,params to the report datataframe    \n                        if verbose >= 1:\n                            print(i_total, 'Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ),  'Average r2:', np.round(np.mean(list_r2 ),4 ),  'Average mae:', np.round(np.mean(list_r2 ),4 ))\n                        IX = len(df_stat)+1\n                        df_stat.loc[IX,'mrrmse'] =  np.mean(list_mrrmse )\n                        df_stat.loc[IX,'r2'] =  np.mean(list_r2 ) \n                        df_stat.loc[IX,'mae'] =  np.mean(list_mae ) \n                        df_stat.loc[IX,'max_depth'] =  max_depth\n                        df_stat.loc[IX,'ntrees'] = ntrees\n                        df_stat.loc[IX,'colsample'] =  colsample\n                        df_stat.loc[IX,'lr'] = lr\n                        df_stat.loc[IX,'n_components'] = n_components\n                        df_stat.loc[IX,'subsample'] =   subsample\n                        for i_tmp in range(len(list_mrrmse)):\n                            df_stat.loc[IX,'mrrmse '+CV_scheme + ' ' + str(i_tmp)] =  list_mrrmse[i_tmp]\n                        for i_tmp in range(len(list_r2)):\n                            df_stat.loc[IX,'r2 '+CV_scheme + ' ' + str(i_tmp)] =  list_r2[i_tmp]\n                        for i_tmp in range(len(list_mae)):\n                            df_stat.loc[IX,'mae '+CV_scheme + ' ' + str(i_tmp)] =  list_mae[i_tmp]\n                        df_stat.loc[IX,'Time'] = np.round(time.time() - t0_single_model_all_folds,1)  \n                        if verbose >= 1:\n                            display(df_stat.tail(1))\n                        if i_total % 20 == 0:\n                            df_stat.sort_values(df_stat.columns[0] ).to_csv('df_stat.csv')\n                                \n                                \n                        # ------------------------------------------------ Optional. Save Submission to CSV ----------------------------------------    \n                        if flag_save_each_model_submit :\n                            dict_params_loc = {'max_depth':max_depth, 'ntrees':ntrees, 'lr':lr, 'subsample':subsample, 'colsample':colsample, \n                                               'n_components':n_components  }\n                            save_submit(Y_submit_pred, str_fn_info = str_model_id,  dict_params = dict_params_loc )\n                                \n\ndisplay(df_stat.sort_values(df_stat.columns[0]  ) )\ndf_stat.sort_values(df_stat.columns[0] ).to_csv('df_stat.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-11-11T18:15:31.209106Z","iopub.execute_input":"2023-11-11T18:15:31.209505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show CV scores, etc","metadata":{}},{"cell_type":"code","source":"display(df_stat.sort_values(df_stat.columns[0]  ) )\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Outcomes of early versions","metadata":{}},{"cell_type":"markdown","source":"### Conclusions:\n\n    We considered one-hot, ordinal, target-encoding. In our framework TE is quite better, a bit surpsingly ordinal encoding is better than one-hot.\n    (Related facts discussed in literature: https://towardsdatascience.com/one-hot-encoding-is-making-your-tree-based-ensembles-worse-heres-why-d64b282b5769 )\n    For one-hot - one needs very deep trees, which will finally to enourmous execution time and RAM crashes.\n\n    Early versions contain tuning for ordinal and one-hot encoding for the Random CV-scheme with 5 folds.\n    \n    Versions 12,14 - tuning for Target Encoding - LeaveOneOutEncoder with Random and AmbrosM schemes respectively. Both schemes lead to similar params.\n    One can check other TargetEnciders and tune params for them,  it is not have been done yet for pyboost for that task.\n    We only choose Leaveoneout because it showed better results for Ridge and SVR, but it can different for the pyboost.\n    \n    \n\nRidge (Random 5 folds:)\n\n    fold: 0 mrrmse: 1.0286 r2: -0.0645\n    fold: 1 mrrmse: 1.5037 r2: 0.044\n    fold: 2 mrrmse: 1.1317 r2: -0.431\n    fold: 3 mrrmse: 1.501 r2: -0.0362\n    fold: 4 mrrmse: 1.0357 r2: 0.0714\n\n    Average mrrmse: 1.2401\n    Average r2: -0.0833\n    CPU times: user 16.8 s, sys: 11.4 s, total: 28.2 s\n    Wall time: 7.32 s\n\n\n    AmrbosM CV LOO-enc Version 14\n               mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\tn_components\tmrrmse AmbrosM\tr2 AmbrosM\n    35\t0.934482\t0.041665\t10.0\t1000.0\t0.2\t0.01\t50.0\t0.939138\t0.073975\n    36\t0.934675\t0.013133\t10.0\t1000.0\t0.5\t0.01\t50.0\t0.933020\t0.067515\n    53\t0.935853\t0.024890\t10.0\t1000.0\t0.2\t0.01\t100.0\t0.937071\t0.072993\n    16\t0.936676\t-0.014845\t10.0\t1000.0\t0.5\t0.01\t30.0\t0.935848\t0.081557\n    18\t0.937177\t-0.012405\t10.0\t1000.0\t0.5\t0.01\t30.0\t0.929801\t0.072386\n    15\t0.938012\t0.022985\t10.0\t1000.0\t0.2\t0.01\t30.0\t0.918119\t0.128154\n    \n    RandomCV version 12 \n                  mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\tn_components\n    35\t1.163426\t0.100887\t10.0\t1000.0\t0.2\t0.01\t50.0\n    23\t1.164302\t0.103402\t5.0\t1000.0\t0.2\t0.01\t50.0\n    29\t1.164691\t0.101295\t8.0\t1000.0\t0.2\t0.01\t50.0\n    17\t1.164854\t0.084335\t10.0\t1000.0\t0.2\t0.01\t30.0\n    11\t1.167825\t0.076399\t8.0\t1000.0\t0.2\t0.01\t30.0\n    33\t1.168018\t0.086795\t10.0\t1000.0\t0.2\t0.01\t50.0\n    47\t1.168054\t0.087553\t8.0\t1000.0\t0.2\t0.01\t100.0\n\n\n    Version 8: \n    selected_features_columns = ['cell_type','sm_name']\n    mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\n    19\t1.179872\t0.051471\t10.0\t1000.0\t0.5\t0.0100\n    46\t1.181120\t0.052247\t15.0\t1000.0\t0.5\t0.0100\n    22\t1.185480\t0.040529\t10.0\t1000.0\t0.9\t0.0100\n    49\t1.188173\t0.036996\t15.0\t1000.0\t0.9\t0.0100\n    52\t1.188414\t0.040948\t15.0\t1000.0\t1.0\t0.0100\n\n\n    \n    Version 9:\n    LOO-enc tsvd30 ONLY DRUG \n        mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\n    20\t1.216731\t0.064573\t10.0\t1000.0\t0.5\t0.0010\n    10\t1.218191\t0.055137\t10.0\t100.0\t0.5\t0.0100\n    13\t1.218365\t0.062862\t10.0\t100.0\t0.9\t0.0100\n    16\t1.219039\t0.066023\t10.0\t100.0\t1.0\t0.0100\n    26\t1.219111\t0.066407\t10.0\t1000.0\t1.0\t0.0010\n    23\t1.219372\t0.065715\t10.0\t1000.0\t0.9\t0.0010\n    47\t1.220931\t0.066633\t15.0\t1000.0\t0.5\t0.0010\n    53\t1.222609\t0.067975\t15.0\t1000.0\t1.0\t0.0010\n    37\t1.223011\t0.056021\t15.0\t100.0\t0.5\t0.0100\n    50\t1.223013\t0.067273\t15.0\t1000.0\t0.9\t0.0010\n\n    \n    LeaveOneOut both CT and Drug\n    Average mrrmse: 1.2232 Average r2: 0.0671\n\n    \n    LeaveOneOut drug only\n    Average mrrmse: 1.2201 Average r2: 0.0654\n\n    \n\n    Ordinal - drug only\n    \n        mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\n    23\t1.227360\t0.019551\t10.0\t1000.0\t0.9\t0.0010\n    20\t1.227360\t0.019549\t10.0\t1000.0\t0.5\t0.0010\n    26\t1.227597\t0.018875\t10.0\t1000.0\t1.0\t0.0010\n    10\t1.227980\t0.018493\t10.0\t100.0\t0.5\t0.0100\n    13\t1.227980\t0.018493\t10.0\t100.0\t0.9\t0.0100\n    16\t1.228215\t0.016203\t10.0\t100.0\t1.0\t0.0100\n    47\t1.229801\t0.017722\t15.0\t1000.0\t0.5\t0.0010\n    50\t1.229801\t0.017722\t15.0\t1000.0\t0.9\t0.0010\n    53\t1.230086\t0.016992\t15.0\t1000.0\t1.0\t0.0010\n    37\t1.230370\t0.016650\t15.0\t100.0\t0.5\t0.0100\n    40\t1.230370\t0.016650\t15.0\t100.0\t0.9\t0.0100\n    43\t1.230730\t0.014203\t15.0\t100.0\t1.0\t0.0100\n    7\t1.253304\t0.038348\t10.0\t50.0\t1.0\t0.0100\n    1\t1.253764\t0.039490\t10.0\t50.0\t0.5\t0.0100\n    4\t1.253764\t0.039490\t10.0\t50.0\t0.9\t0.0100\n\n\n    Ordinal - both - CT and Drug\n    \n        mrrmse\tr2\tmax_depth\tntrees\tcolsample\tlr\n    10\t1.237836\t0.051941\t10.0\t100.0\t0.5\t0.0100\n    13\t1.237836\t0.051941\t10.0\t100.0\t0.9\t0.0100\n    37\t1.238562\t0.051370\t15.0\t100.0\t0.5\t0.0100\n    40\t1.238562\t0.051370\t15.0\t100.0\t0.9\t0.0100\n    23\t1.239908\t0.054912\t10.0\t1000.0\t0.9\t0.0010\n    20\t1.239909\t0.054912\t10.0\t1000.0\t0.5\t0.0010\n    50\t1.240593\t0.054277\t15.0\t1000.0\t0.9\t0.0010\n    47\t1.240594\t0.054278\t15.0\t1000.0\t0.5\t0.0010\n    4\t1.268880\t0.036781\t10.0\t50.0\t0.9\t0.0100\n    1\t1.268880\t0.036781\t10.0\t50.0\t0.5\t0.0100\n    28\t1.269403\t0.036496\t15.0\t50.0\t0.5\t0.0100\n    31\t1.269408\t0.036493\t15.0\t50.0\t0.9\t0.0100\n    19\t1.305464\t-0.180885\t10.0\t1000.0\t0.5\t0.0100\n    22\t1.305504\t-0.180885\t10.0\t1000.0\t0.9\t0.0100\n    49\t1.316761\t-0.190236\t15.0\t1000.0\t0.9\t0.0100\n    46\t1.316767\t-0.190240\t15.0\t1000.0\t0.5\t0.0100\n    14\t1.319407\t-0.008435\t10.0\t100.0\t0.9\t0.0010\n    11\t1.319407\t-0.008435\t10.0\t100.0\t0.5\t0.0010\n    38\t1.319549\t-0.008504\t15.0\t100.0\t0.5\t0.0010\n    41\t1.319549\t-0.008504\t15.0\t100.0\t0.9\t0.0010\n    24\t1.319919\t-0.008704\t10.0\t1000.0\t0.9\t0.0001\n    21\t1.319919\t-0.008704\t10.0\t1000.0\t0.5\t0.0001\n    48\t1.320064\t-0.008783\t15.0\t1000.0\t0.5\t0.0001\n    51\t1.320064\t-0.008783\t15.0\t1000.0\t0.9\t0.0001\n    27\t1.327581\t-0.011636\t10.0\t1000.0\t1.0\t0.0001\n    5\t1.327597\t-0.016774\t10.0\t50.0\t0.9\t0.0010\n    2\t1.327597\t-0.016774\t10.0\t50.0\t0.5\t0.0010\n    32\t1.327681\t-0.016806\t15.0\t50.0\t0.9\t0.0010\n    29\t1.327682\t-0.016807\t15.0\t50.0\t0.5\t0.0010\n    17\t1.327981\t-0.012478\t10.0\t100.0\t1.0\t0.0010\n    54\t1.328842\t-0.012124\t15.0\t1000.0\t1.0\t0.0001\n    7\t1.328938\t-0.010233\t10.0\t50.0\t1.0\t0.0100\n    44\t1.329341\t-0.013038\t15.0\t100.0\t1.0\t0.0010\n    8\t1.331488\t-0.018432\t10.0\t50.0\t1.0\t0.0010\n    35\t1.332187\t-0.018708\t15.0\t50.0\t1.0\t0.0010\n    34\t1.333957\t-0.012900\t15.0\t50.0\t1.0\t0.0100\n    12\t1.334731\t-0.024247\t10.0\t100.0\t0.5\t0.0001\n    15\t1.334731\t-0.024247\t10.0\t100.0\t0.9\t0.0001\n    42\t1.334747\t-0.024254\t15.0\t100.0\t0.9\t0.0001\n    39\t1.334747\t-0.024254\t15.0\t100.0\t0.5\t0.0001\n    18\t1.335379\t-0.024462\t10.0\t100.0\t1.0\t0.0001\n    45\t1.335511\t-0.024514\t15.0\t100.0\t1.0\t0.0001\n    3\t1.335623\t-0.025185\t10.0\t50.0\t0.5\t0.0001\n    6\t1.335623\t-0.025185\t10.0\t50.0\t0.9\t0.0001\n    30\t1.335632\t-0.025188\t15.0\t50.0\t0.5\t0.0001\n    33\t1.335632\t-0.025188\t15.0\t50.0\t0.9\t0.0001\n    9\t1.335951\t-0.025298\t10.0\t50.0\t1.0\t0.0001\n    36\t1.336016\t-0.025323\t15.0\t50.0\t1.0\t0.0001\n    26\t1.354519\t-0.060866\t10.0\t1000.0\t1.0\t0.0010\n    16\t1.356253\t-0.064348\t10.0\t100.0\t1.0\t0.0100\n    53\t1.363119\t-0.066754\t15.0\t1000.0\t1.0\t0.0010\n    43\t1.365004\t-0.071763\t15.0\t100.0\t1.0\t0.0100\n    25\t1.537157\t-0.589301\t10.0\t1000.0\t1.0\t0.0100\n    52\t1.577464\t-0.666247\t15.0\t1000.0\t1.0\t0.0100\n\n\n\n        mrrmse\n    smoothing\t\n    100.0\t1.228143\n    50.0\t1.228178\n    200.0\t1.228248\n    10000.0\t1.228273\n    inf\t1.228283\n    500.0\t1.228316\n    1000.0\t1.228317\n    15.0\t1.228518\n    20.0\t1.228572\n    12.0\t1.228924\n    11.0\t1.229052\n    10.0\t1.229240\n    5.0\t1.229384\n    9.0\t1.229439\n    8.0\t1.229675\n    1.0\t1.231323\n    0.0\t1.336181\n\n\n    \n    max_depth 25 Average mrrmseL 1.2383 Average r2: 0.0268 - 5hours(!)\n    max_depth 24 Average mrrmse: 1.2387 Average r2: 0.0268\n    max_depth 22 Average mrrmse: 1.2394 Average r2: 0.0266\n    max_depth 21 Average mrrmse: 1.2398 Average r2: 0.0264\n    max_depth 20 Average mrrmse: 1.2403 Average r2: 0.0263\n    mrrmse\tr2\tmax_depth\n    10\t1.240853\t0.026067\t19.0\n    9\t1.241418\t0.025876\t18.0\n    8\t1.242047\t0.025684\t17.0\n    7\t1.242776\t0.025446\t16.0\n    6\t1.243554\t0.025254\t15.0\n    5\t1.244384\t0.024969\t14.0\n    4\t1.245430\t0.024568\t13.0\n    3\t1.246498\t0.024308\t12.0\n    2\t1.247458\t0.024128\t11.0\n    1\t1.248641\t0.023593\t10.0\n\n    \n    mrrmse\tr2\tmax_depth\n    9\t1.249972\t0.023016\t9.0\n    8\t1.251896\t0.021569\t8.0\n    7\t1.255513\t0.017397\t7.0\n    6\t1.260368\t0.011708\t6.0\n    5\t1.266247\t0.006205\t5.0\n    4\t1.273336\t0.000881\t4.0\n    3\t1.281519\t-0.002841\t3.0\n    2\t1.293277\t-0.006253\t2.0\n    1\t1.312721\t-0.011352\t1.0\n\n        # ---------------\n        # 10 maxdepth\n        #    \n\n        # n_trees: 5000 , lr = 0002\n        # mrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.241797\t0.011268\t10.0\t100.0\n        # Wall time: 25min 44s\n\n        # ntrees1000, lr 0.001\n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.241702\t0.010889\t10.0\t100.0\n\n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.242173\t0.008403\t10.0\t100.0\n\n        # 10 maxdepth\n\n        # ---------------\n\n        # ntrees = 1000, le = 0.001, colsample = 1  \n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.236224\t0.012712\t15.0\t100.0\n        # Same but colsample = 0.8 - quite worse \n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.242032\t0.025552\t15.0\t100.0\n        # That is with standard ntrees, lr: \n        # 1.243554    0.025254    15.0\n        # 100 tres with low LR   \n        #1\t1.314346\t-0.003038\t15.0\t100.0    \n\n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.241702\t0.010889\t10.0\t100.0    \n\n\n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.242173\t0.008403\t10.0\t100.0\n\n        # 1000 lr=1e-4\n        # mrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.314263\t-0.002907\t15.0\t100.0\n\n\n\n        # \tmrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.24221\t0.008955\t10.0\t100.0    \n\n        # mrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.248641\t0.023593\t10.0\t50.0\n        # 2\t1.248641\t0.023593\t10.0\t200.0\n        # 3\t1.248641\t0.023593\t10.0\t400.0\n\n\n        # mrrmse\tr2\tmax_depth\tntrees\n        # 1\t1.248641\t0.023593\t10.0\t100.0    \n","metadata":{}},{"cell_type":"markdown","source":"# Final timing","metadata":{"execution":{"iopub.status.busy":"2023-10-25T10:04:59.25663Z","iopub.execute_input":"2023-10-25T10:04:59.25698Z","iopub.status.idle":"2023-10-25T10:04:59.261681Z","shell.execute_reply.started":"2023-10-25T10:04:59.256955Z","shell.execute_reply":"2023-10-25T10:04:59.260221Z"}}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}