{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 12 - switch to big size of playground - half of full train - it takes long time - 1.5 hour\n\n    LightGBM - runs 0.6 seconds.\n    It still worse than Ridge, but difference is not big. \n    RF tuned - runs for 10 mintes  (500 interators)\n    Kernel Ridge - is very slow - 5 minutes - and again terrbible quality \n    XGBOost is 10 times slower with default params than Lgb, but quality is quite lower\n    Catboost default is second after the Ridge, but it takes 8 seconds, comparing to 0.6 lgb \n    \n#### 11 - cosmetic change\n    small error corrected - LGB Tuned2 params were not correct\n    Resulting table with statistcs - has been saved.  \n    \n#### 10 added parameter to control size of the play-ground\n\n    playground_fraction_of_train = 10# Should be Integer N - we create play-ground of the size Full_Train / N. \n\n#### 9 Blend and BIG SURPRISE(!) - how to explain ? ideas - welcome ! \n\n    added blending part for all solutions - and got suprise:\n    and the WORST/TERRIBLE (r2 negative, mse solution 10 times worse than others -   KernelRidge\n    enters the blends and improve it ! \n    \n    The only thing - that it is very uncorrelated with the other solutions. \n\nSee discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230\n\n\n#### 8: Added - statistics on all methods - collected in one table (dataframe)\n    \n    Still Ridge, SVR are the best. Cat\n\n#### 7: added XGBoost+Optuna, CatBoost, MLP\n    \n    CatBoost is quite good with default params\n\n#### 3,4,5,6 - search for optimal LightGBM params; added: models and param tuning for other models,\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n    Version 6 - seems around 40 PCA features is better for LightGBM  - current_best_r2 = 0.46792863038644295 # \n    \n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"tune_RF = False # True # It is slow - takes 4 minutes for 6 params  for play-ground of size 10% from original \nplayground_fraction_of_train = 1# Should be Integer N - we create play-ground of the size Full_Train / N. ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:36:26.445685Z","iopub.execute_input":"2022-11-07T20:36:26.446419Z","iopub.status.idle":"2022-11-07T20:36:26.482212Z","shell.execute_reply.started":"2022-11-07T20:36:26.446237Z","shell.execute_reply":"2022-11-07T20:36:26.481301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-07T20:36:26.48489Z","iopub.execute_input":"2022-11-07T20:36:26.485297Z","iopub.status.idle":"2022-11-07T20:36:26.528019Z","shell.execute_reply.started":"2022-11-07T20:36:26.485266Z","shell.execute_reply":"2022-11-07T20:36:26.526763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:36:26.529946Z","iopub.execute_input":"2022-11-07T20:36:26.530689Z","iopub.status.idle":"2022-11-07T20:36:54.876705Z","shell.execute_reply.started":"2022-11-07T20:36:26.530629Z","shell.execute_reply":"2022-11-07T20:36:54.875304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:36:54.88127Z","iopub.execute_input":"2022-11-07T20:36:54.881707Z","iopub.status.idle":"2022-11-07T20:37:19.242849Z","shell.execute_reply.started":"2022-11-07T20:36:54.881672Z","shell.execute_reply":"2022-11-07T20:37:19.241241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:19.244978Z","iopub.execute_input":"2022-11-07T20:37:19.245387Z","iopub.status.idle":"2022-11-07T20:37:43.861849Z","shell.execute_reply.started":"2022-11-07T20:37:19.245333Z","shell.execute_reply":"2022-11-07T20:37:43.860451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:43.863785Z","iopub.execute_input":"2022-11-07T20:37:43.864291Z","iopub.status.idle":"2022-11-07T20:37:44.88076Z","shell.execute_reply.started":"2022-11-07T20:37:43.864234Z","shell.execute_reply":"2022-11-07T20:37:44.879515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:44.883381Z","iopub.execute_input":"2022-11-07T20:37:44.883927Z","iopub.status.idle":"2022-11-07T20:37:46.791105Z","shell.execute_reply.started":"2022-11-07T20:37:44.883879Z","shell.execute_reply":"2022-11-07T20:37:46.789558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:46.792716Z","iopub.execute_input":"2022-11-07T20:37:46.793236Z","iopub.status.idle":"2022-11-07T20:37:47.344401Z","shell.execute_reply.started":"2022-11-07T20:37:46.793186Z","shell.execute_reply":"2022-11-07T20:37:47.342921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\"  (i.e. downsample)\n\nSmall 10% of data of data where we can make prelimanary experiments. \n\nit will be labeled by special column \"Playground\" in df_meta (meta data for CITE-seq train only)","metadata":{}},{"cell_type":"code","source":"# Prepare for creation of additional holdout folds with 10% of samples \n# We will use stratified Kfold to achieve that days, cell_types and donors are equally distributed \nimport numpy as np\nfrom sklearn.model_selection import StratifiedKFold\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\nflagged_column_name = 'Playground'\n\nif playground_fraction_of_train > 1:\n    # playground_fraction_of_train = 10\n    skf = StratifiedKFold(n_splits= playground_fraction_of_train,  shuffle=True, random_state=40)\n    skf.get_n_splits(df_meta, df_meta[scol] )\n\n\n\n    y = df_meta[scol] \n    for train_index, test_index in skf.split(df_meta, df_meta[scol]):\n        print(\"TRAIN:\", len(train_index), \"TEST:\", len(test_index) ); \n        break\n    print(test_index)\n    print(df_meta[scol].value_counts().head(5)   )\n    print(df_meta.iloc[test_index,:][scol].value_counts().head(5)    )\n\n\n    df_meta[flagged_column_name] = 0 \n    df_meta.loc[df_meta.index[test_index],flagged_column_name]  = 1\n\nelse:\n    df_meta[flagged_column_name] = 1 # Everything is playground \n    \ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.346373Z","iopub.execute_input":"2022-11-07T20:37:47.34671Z","iopub.status.idle":"2022-11-07T20:37:47.571612Z","shell.execute_reply.started":"2022-11-07T20:37:47.34668Z","shell.execute_reply":"2022-11-07T20:37:47.570447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"code","source":"selected_target = 'CD31'\nn_features = 100\n\nX= df_cite.iloc[:70988,:n_features][df_meta['Playground']==1]\ny= df_cite_train_y[df_meta['Playground']==1][selected_target]\nprint('X.shape, y.shape', X.shape, y.shape )\n\n# Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n# Private like test - new DAY, and donor, \n# While public like - only new donor (days are the same as in train):\n# Step 1: \nmask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \nX_train = df_cite.iloc[:70988,:n_features][mask_train]\ny_train = df_cite_train_y[mask_train][ selected_target ]\nprint('X_train.shape, y_train.shape', X_train.shape, y_train.shape )\n\n# Step 2:\nmask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\nX_test_private_like = df_cite.iloc[:70988,:n_features][ mask_test_private_like  ]\ny_test_private_like = df_cite_train_y[mask_test_private_like][ selected_target ]\nX_test = X_test_private_like\ny_test = y_test_private_like\nprint('X_test.shape, y_test.shape (\"private like\"  - with new day and donor )', X_test.shape, y_test.shape )\n\n\n# Step 3: \nmask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \nX_test_public_like = df_cite.iloc[:70988,:n_features][mask_test_public_like ]\ny_test_public_like = df_cite_train_y[mask_test_public_like][ selected_target ]\nX_test2 = X_test_public_like\ny_test2 = y_test_public_like\nprint('X_test2.shape, y_test2.shape (\"public like\" only new donor )', X_test2.shape, y_test2.shape )\n\n\nif playground_fraction_of_train > 1:\n    mask_out_of_playground = (df_meta['Playground']==0)\n    X_oop = df_cite.iloc[:70988,:n_features][mask_out_of_playground]\n    y_oop = df_cite_train_y[mask_out_of_playground][ selected_target ]\nelse:\n    # Playground is everything so we that should be empty, but not to crash we will take something - not important what in particular\n    X_oop = df_cite.iloc[:10,:n_features]\n    y_oop = df_cite_train_y.iloc[:10][ selected_target ]\n    \nprint('X_oop, y_oop', X_oop.shape, y_oop.shape )\n\n\nX.shape,y.shape, X_train.shape, X_test_private_like.shape, X_test_public_like.shape, y_train.shape, y_test_private_like.shape, y_test_public_like.shape","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.575353Z","iopub.execute_input":"2022-11-07T20:37:47.575718Z","iopub.status.idle":"2022-11-07T20:37:47.755797Z","shell.execute_reply.started":"2022-11-07T20:37:47.575687Z","shell.execute_reply":"2022-11-07T20:37:47.754623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling Preparations","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.757609Z","iopub.execute_input":"2022-11-07T20:37:47.758337Z","iopub.status.idle":"2022-11-07T20:37:47.764134Z","shell.execute_reply.started":"2022-11-07T20:37:47.758293Z","shell.execute_reply":"2022-11-07T20:37:47.763085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat = pd.DataFrame()# columns = [ 'r2_score','mse','Time',  'n_feat', 'Target' ])\n#IXmdl = 0\ndf_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.76608Z","iopub.execute_input":"2022-11-07T20:37:47.766618Z","iopub.status.idle":"2022-11-07T20:37:47.785955Z","shell.execute_reply.started":"2022-11-07T20:37:47.766572Z","shell.execute_reply":"2022-11-07T20:37:47.784441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict_save_predictions = {}\ndef update_models_stat(model, model_ID, t0 ):\n    \n    IXmdl = model_ID # df_models_stat.shape[0] + 1\n    #df_models_stat.loc[IXmdl,'Model'] = model_ID\n    y_pred_loc = model.predict(X_test)\n    df_models_stat.loc[IXmdl,'r2_score'] = r2_score(y_test, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse'] =  mean_squared_error(y_test, y_pred_loc)\n    df_models_stat.loc[IXmdl,'Time'] = np.round( time.time() - t0,3)\n    df_models_stat.loc[IXmdl,'Target'] = selected_target\n    df_models_stat.loc[IXmdl,'n_feat'] = X_train.shape[1]\n    df_models_stat.loc[IXmdl,'n_samples_train'] = X_train.shape[0]\n    \n    y_pred_loc = model.predict(X_test_public_like)\n    df_models_stat.loc[IXmdl,'r2_score Test2 PublLike'] = r2_score(y_test_public_like, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Test2 PublLike'] =  mean_squared_error(y_test_public_like, y_pred_loc)\n    \n    y_pred_loc = model.predict(X_oop)\n    df_models_stat.loc[IXmdl,'r2_score Out of Playgr'] = r2_score(y_oop, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Out of Playgr'] =  mean_squared_error(y_oop, y_pred_loc)\n\n    y_pred_loc = model.predict(X_train)\n    df_models_stat.loc[IXmdl,'r2_score Train'] = r2_score(y_train, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Train'] =  mean_squared_error(y_train, y_pred_loc)\n    \n    dict_save_predictions[model_ID] = (model.predict(X_train), model.predict(X_test), \n                                       model.predict(X_test_public_like), model.predict(X_oop)   )\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.787916Z","iopub.execute_input":"2022-11-07T20:37:47.788509Z","iopub.status.idle":"2022-11-07T20:37:47.804917Z","shell.execute_reply.started":"2022-11-07T20:37:47.788474Z","shell.execute_reply":"2022-11-07T20:37:47.803691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Lama\n","metadata":{}},{"cell_type":"code","source":"!pip install -U lightautoml","metadata":{"execution":{"iopub.status.busy":"2022-11-07T21:00:28.055222Z","iopub.execute_input":"2022-11-07T21:00:28.055753Z","iopub.status.idle":"2022-11-07T21:02:28.513929Z","shell.execute_reply.started":"2022-11-07T21:00:28.055709Z","shell.execute_reply":"2022-11-07T21:02:28.511353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_full_train = X_train\ndata_full_train['answer'] = y_train","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.806891Z","iopub.execute_input":"2022-11-07T20:37:47.807287Z","iopub.status.idle":"2022-11-07T20:37:47.822887Z","shell.execute_reply.started":"2022-11-07T20:37:47.80723Z","shell.execute_reply":"2022-11-07T20:37:47.821303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lightautoml.automl.presets.tabular_presets import TabularAutoML, TabularUtilizedAutoML\nfrom lightautoml.tasks import Task\n\ndef r2_metric(y_true, y_pred):\n    return f1_score(y_true, y_pred)\n\ntask = Task('reg', metric = r2_score)\n\nroles = {'target': 'answer'}\n\nautoml = TabularAutoML(task = task, \n                       timeout = 600*1, # 600 seconds = 10 minutes\n                       cpu_limit = 4, # Optimal for Kaggle kernels\n                       general_params = {'use_algos': [['linear_l2', \n                                         'lgb', 'lgb_tuned'], ['linear_l2', 'cb']]})\n\noof_pred = automl.fit_predict(data_full_train, roles = roles)\ny_pred = automl.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T22:01:45.415648Z","iopub.execute_input":"2022-11-07T22:01:45.416304Z","iopub.status.idle":"2022-11-07T22:14:50.663284Z","shell.execute_reply.started":"2022-11-07T22:01:45.416258Z","shell.execute_reply":"2022-11-07T22:14:50.661572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = automl.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T22:22:54.071702Z","iopub.execute_input":"2022-11-07T22:22:54.072174Z","iopub.status.idle":"2022-11-07T22:22:57.641285Z","shell.execute_reply.started":"2022-11-07T22:22:54.072137Z","shell.execute_reply":"2022-11-07T22:22:57.639615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'train score : {r2_score(y_train, oof_pred.data[:, 0])}')\nprint(f'test score : {r2_score(y_test, y_pred.data[:, 0])}')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T22:23:17.827653Z","iopub.execute_input":"2022-11-07T22:23:17.829287Z","iopub.status.idle":"2022-11-07T22:23:17.840598Z","shell.execute_reply.started":"2022-11-07T22:23:17.829226Z","shell.execute_reply":"2022-11-07T22:23:17.838846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Autogluon","metadata":{}},{"cell_type":"code","source":"-","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install autogluon","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.976474Z","iopub.status.idle":"2022-11-07T20:37:47.977088Z","shell.execute_reply.started":"2022-11-07T20:37:47.976865Z","shell.execute_reply":"2022-11-07T20:37:47.976887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from autogluon.tabular import TabularDataset, TabularPredictor","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.97919Z","iopub.status.idle":"2022-11-07T20:37:47.97968Z","shell.execute_reply.started":"2022-11-07T20:37:47.979461Z","shell.execute_reply":"2022-11-07T20:37:47.979482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import autogluon.core as ag\n\nnn_torch_options = {  # specifies non-default hyperparameter values for neural network models\n    'num_epochs': 10,  # number of training epochs (controls training time of NN models)\n    'learning_rate': ag.space.Real(4e-3, 1e-1, default=4e-3, log=True),  # learning rate used in training (real-valued hyperparameter searched on log-scale)\n    'activation': ag.space.Categorical('relu', 'softrelu', 'tanh'),  # activation function used in NN (categorical hyperparameter, default = first entry)\n    'dropout_prob': ag.space.Real(0.0, 0.5, default=0.1),  # dropout probability (real-valued hyperparameter)\n}\n\nnn_fast_options = {  # specifies non-default hyperparameter values for neural network models\n    'num_epochs': 10,  # number of training epochs (controls training time of NN models)\n    'learning_rate': ag.space.Real(4e-3, 1e-1, default=4e-3, log=True),  # learning rate used in training (real-valued hyperparameter searched on log-scale)\n    'activation': ag.space.Categorical('relu', 'softrelu', 'tanh'),  # activation function used in NN (categorical hyperparameter, default = first entry)\n    'dropout_prob': ag.space.Real(0.0, 0.5, default=0.1),  # dropout probability (real-valued hyperparameter)\n}\n\ngbm_options = {  # specifies non-default hyperparameter values for lightGBM gradient boosted trees\n    'num_boost_round': 100,  # number of boosting rounds (controls training time of GBM models)\n    'num_leaves': ag.space.Int(lower=40, upper=512, default=100),  # number of leaves in trees (integer hyperparameter)\n    'n_estimators': ag.space.Int(lower=10, upper=256),\n    'max_depth': ag.space.Int(lower=50, upper=500),\n    'min_child_samples': ag.space.Int(lower=1, upper=300),\n    'reg_alpha': ag.space.Real(0.0, 1.0),\n    'reg_lambda': ag.space.Real(0.0, 1.0),\n    \n}\n\nhyperparameters = {  # hyperparameters of each model type\n                   'GBM': gbm_options,\n                    'FAST_AI': nn_fast_options,\n                   'NN_TORCH': nn_torch_options,  # NOTE: comment this line out if you get errors on Mac OSX\n                  }  # When these keys are missing from hyperparameters dict, no models of that type are trained\n\ntime_limit = 30*60  # 10 min limit for model \nnum_trials = 8  # try at most 5 different hyperparameter configurations for each type of model\nsearch_strategy = 'auto'  # to tune hyperparameters using random search routine with a local scheduler\n\nmetric = 'r2'\nlabel = 'answer'\n\nhyperparameter_tune_kwargs = {  # HPO is not performed unless hyperparameter_tune_kwargs is specified\n    'num_trials': num_trials,\n    'scheduler' : 'local',\n    'searcher': search_strategy,\n}\n\nnum_bag_folds=10\nnum_bag_sets=60\nnum_stack_levels=10\n\npredictor = TabularPredictor(label=label, eval_metric=metric).fit(\n    data_full, time_limit=time_limit,\n    num_bag_folds=num_bag_folds, num_bag_sets=num_bag_sets, num_stack_levels=num_stack_levels\n)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.981881Z","iopub.status.idle":"2022-11-07T20:37:47.982327Z","shell.execute_reply.started":"2022-11-07T20:37:47.98212Z","shell.execute_reply":"2022-11-07T20:37:47.982139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = predictor.predict(X_test)\npredictor.evaluate_predictions(y_true=y_test, y_pred=y_pred, auxiliary_metrics=True)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.983743Z","iopub.status.idle":"2022-11-07T20:37:47.984499Z","shell.execute_reply.started":"2022-11-07T20:37:47.984234Z","shell.execute_reply":"2022-11-07T20:37:47.984263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = predictor.fit_summary()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.986298Z","iopub.status.idle":"2022-11-07T20:37:47.986755Z","shell.execute_reply.started":"2022-11-07T20:37:47.986557Z","shell.execute_reply":"2022-11-07T20:37:47.986577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor.info()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.987858Z","iopub.status.idle":"2022-11-07T20:37:47.988455Z","shell.execute_reply.started":"2022-11-07T20:37:47.988083Z","shell.execute_reply":"2022-11-07T20:37:47.988102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor.leaderboard(data_full, extra_metrics=['mean_squared_error', 'mean_absolute_error', 'r2', 'pearsonr', 'median_absolute_error'], silent=True)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.989722Z","iopub.status.idle":"2022-11-07T20:37:47.990122Z","shell.execute_reply.started":"2022-11-07T20:37:47.989926Z","shell.execute_reply":"2022-11-07T20:37:47.989945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#predictor.feature_importance(data_full)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.991716Z","iopub.status.idle":"2022-11-07T20:37:47.992119Z","shell.execute_reply.started":"2022-11-07T20:37:47.991921Z","shell.execute_reply":"2022-11-07T20:37:47.99194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# RVM","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install https://github.com/JamesRitchie/scikit-rvm/archive/master.zip","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.994154Z","iopub.status.idle":"2022-11-07T20:37:47.995194Z","shell.execute_reply.started":"2022-11-07T20:37:47.994918Z","shell.execute_reply":"2022-11-07T20:37:47.994959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from numpy import hstack\nfrom sklearn.datasets import make_classification\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.linear_model import LogisticRegression\nfrom catboost import CatBoostRegressor\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.svm import SVR\nfrom sklearn.linear_model import ElasticNet\nfrom sklearn.kernel_ridge import KernelRidge\nfrom xgboost.sklearn import XGBRegressor\nfrom sklearn.linear_model import Ridge\n\ndef get_models():\n    models = list()\n    models.append(('elastic', ElasticNet()))\n    models.append(('boosting', GradientBoostingRegressor()))\n    models.append(('kernel_ridge', KernelRidge(kernel='rbf')))\n    models.append(('catboost', CatBoostRegressor(verbose=0)))\n    models.append(('svr', SVR()))\n    models.append(('xgb', XGBRegressor()))\n    \n    return models\n \n# fit the blending ensemble\ndef fit_ensemble(models, X_train, X_val, y_train, y_val):\n    # fit all models on the training set and predict on hold out set\n    meta_X = list()\n    for name, model in models:\n        # fit in training set\n        model.fit(X_train, y_train)\n        # predict on hold out set\n        yhat = model.predict(X_val)\n        # reshape predictions into a matrix with one column\n        yhat = yhat.reshape(len(yhat), 1)\n        # store predictions as input for blending\n        meta_X.append(yhat)\n        # create 2d array from predictions, each set is an input feature\n    meta_X = hstack(meta_X)\n    # define blending model\n    blender = Ridge()\n    # fit on predictions from base models\n    blender.fit(meta_X, y_val)\n    return blender\n \n# make a prediction with the blending ensemble\ndef predict_ensemble(models, blender, X_test):\n    # make predictions with base models\n    meta_X = list()\n    for name, model in models:\n        # predict with base model\n        yhat = model.predict(X_test)\n        # reshape predictions into a matrix with one column\n        yhat = yhat.reshape(len(yhat), 1)\n        # store prediction\n        meta_X.append(yhat)\n    # create 2d array from predictions, each set is an input feature\n    meta_X = hstack(meta_X)\n    # predict\n    return blender.predict(meta_X)\n \n\nX_train_2, X_val, y_train_2, y_val = train_test_split(X_train, y_train, test_size=0.33, random_state=1)\nprint('Train: %s, Val: %s, Test: %s' % (X_train.shape, X_val.shape, X_test.shape))\n# create the base models\nmodels = get_models()\n# train the blending ensemble\nblender = fit_ensemble(models, X_train_2, X_val, y_train_2, y_val)\n# make predictions on test set\nyhat = predict_ensemble(models, blender, X_test)\n# evaluate predictions\nr2_score(y_test, yhat)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.997045Z","iopub.status.idle":"2022-11-07T20:37:47.997487Z","shell.execute_reply.started":"2022-11-07T20:37:47.997263Z","shell.execute_reply":"2022-11-07T20:37:47.997283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LightGBM model","metadata":{}},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm\nprint('Default Params')\nt0 = time.time()\nmodel = lgbm.LGBMRegressor(random_state = 0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\ntm0 = time.time() - t0\nupdate_models_stat(model, 'Lgbm Default', t0 )\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred), 'Secs: %.2f'%(time.time() - t0)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:47.998862Z","iopub.status.idle":"2022-11-07T20:37:47.999262Z","shell.execute_reply.started":"2022-11-07T20:37:47.999072Z","shell.execute_reply":"2022-11-07T20:37:47.999092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.000714Z","iopub.status.idle":"2022-11-07T20:37:48.001231Z","shell.execute_reply.started":"2022-11-07T20:37:48.000918Z","shell.execute_reply":"2022-11-07T20:37:48.000937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 40 features\n# 0.46792863038644295\n\n# 50 features\n# 0.4665950972767404\n# \n# 100 features:\n# 0.4471078544471576 # learning_rate= 0.07\n#\n# 500 features - tune by hands - one param after another: \n# (0.4354558552914086, 22.998228325953246) # (learning_rate= 0.067 ,num_leaves = 12, max_depth = 9, n_estimators= 150, subsample = 0.4, random_state = 0 )\n# (0.43090615704586654, 23.18357255464802) # learning_rate= 0.067 ,num_leaves = 12, max_depth = 9\n# (0.42965094120477576, 23.23470715728686) # learning_rate= 0.067 ,num_leaves = 12, max_depth = 8\n# (0.42768216014074323, 23.314910763787474) # learning_rate= 0.067 ,num_leaves = 12\n# (0.4234603331885415, 23.486898271070043) # learning_rate= 0.067,num_leaves = 19\n# (0.42228006014959485, 23.534979876540724) # learning_rate= 0.067\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.003049Z","iopub.status.idle":"2022-11-07T20:37:48.003494Z","shell.execute_reply.started":"2022-11-07T20:37:48.003264Z","shell.execute_reply":"2022-11-07T20:37:48.003283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape, y_train.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.005217Z","iopub.status.idle":"2022-11-07T20:37:48.005661Z","shell.execute_reply.started":"2022-11-07T20:37:48.005453Z","shell.execute_reply":"2022-11-07T20:37:48.005474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport time\nt0 = time.time()\ncurrent_best_r2 = 0.46792863038644295 #  0.4665950972767404 # 0.4471078544471576 # 0.4366682658232913 # 0.43590181064083766\nmodel = lgbm.LGBMRegressor(random_state = 0, \n    #extra_trees = True, # For 0.46792863038644295 it worsens - probably need to tune other params \n    learning_rate= 0.07 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n    num_leaves = 12, # hypersensitive - change to 11, worsens about 2%\n    max_depth = 9, # hypersensitive - change to 10 - worsens about 1%\n    n_estimators= 150, # senstitive - change by 10 - worsents about 0.5%\n    min_child_samples = 20,# hypersensitive - change to 21 - worsens about 3% \n    min_split_gain = 12.2,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n    reg_lambda = 0.0, # seems only makes worse\n    reg_alpha = 0.0, # seems only makes worse\n    subsample_for_bin = 10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n    colsample_bytree = 1, # only worsens\n    subsample = 1, # Does not seem to influence at all\n    other_rate = 1,# no influence ?  \n    min_child_weight =  0.1, # no influence ?\n    subsample_freq = 10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n              )# \nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'Lgbm Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred), (r2_score(y_test, y_pred) - current_best_r2 ) / current_best_r2 * 100","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.008218Z","iopub.status.idle":"2022-11-07T20:37:48.008687Z","shell.execute_reply.started":"2022-11-07T20:37:48.00847Z","shell.execute_reply":"2022-11-07T20:37:48.008492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nLGB_PARAMETERS = {\n    \"random_state\": 0, \n    \"learning_rate\": 0.09 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n    \"num_leaves\": 12, # hypersensitive - change to 11, worsens about 2%\n    \"max_depth\": 15, # hypersensitive - change to 10 - worsens about 1%\n    \"n_estimators\": 150, # senstitive - change by 10 - worsents about 0.5%\n    \"min_child_samples\": 50,# hypersensitive - change to 21 - worsens about 3% \n    \"min_split_gain\": 3.119890406724053,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n    \"reg_lambda\": 0.8575143843171859, # seems only makes worse\n    \"reg_alpha\": 0.0, # seems only makes worse\n    \"subsample_for_bin\": 10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n    \"colsample_bytree\": 1, # only worsens\n    \"subsample\": 1, # Does not seem to influence at all\n    \"other_rate\": 1,# no influence ?  \n    \"min_child_weight\":  0.1, # no influence ?\n    \"subsample_freq\": 10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n        }\n\n\nimport time\nt0 = time.time()\ncurrent_best_r2 = 0.46792863038644295 #  0.4665950972767404 # 0.4471078544471576 # 0.4366682658232913 # 0.43590181064083766\nmodel = lgbm.LGBMRegressor(**LGB_PARAMETERS) \nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'Lgbm Tuned2', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred), (r2_score(y_test, y_pred) - current_best_r2 ) / current_best_r2 * 100","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.012875Z","iopub.status.idle":"2022-11-07T20:37:48.013313Z","shell.execute_reply.started":"2022-11-07T20:37:48.013112Z","shell.execute_reply":"2022-11-07T20:37:48.013133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.014813Z","iopub.status.idle":"2022-11-07T20:37:48.015217Z","shell.execute_reply.started":"2022-11-07T20:37:48.015024Z","shell.execute_reply":"2022-11-07T20:37:48.015043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Params from some notebook: \n#         params ={\n#                         'task': 'train',\n#                         'boosting': 'goss',\n#                         'objective': 'regression',\n#                         'metric': 'rmse',\n#                         'learning_rate': 0.005,\n#                         'subsample': 0.9855232997390695,\n#                         'max_depth': 8,\n#                         'top_rate': 0.9064148448434349,\n#                         'num_leaves': 87,\n#                         'min_child_weight': 41.9612869171337,\n#                         'other_rate': 0.0721768246018207,\n#                         'reg_alpha': 9.677537745007898,\n#                         'colsample_bytree': 0.5665320670155495,\n#                         'min_split_gain': 9.820197773625843,\n#                         'reg_lambda': 8.2532317400459,\n#                         'min_data_in_leaf': 21,\n#                         'verbose': -1,\n#                         'seed':int(2**n_fold),\n#                         'bagging_seed':int(2**n_fold),\n#                         'drop_seed':int(2**n_fold)\n#                         }\n\ndir(model)\nmodel.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.016857Z","iopub.status.idle":"2022-11-07T20:37:48.01726Z","shell.execute_reply.started":"2022-11-07T20:37:48.017069Z","shell.execute_reply":"2022-11-07T20:37:48.017088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ridge","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.linear_model import Ridge\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.018712Z","iopub.status.idle":"2022-11-07T20:37:48.019121Z","shell.execute_reply.started":"2022-11-07T20:37:48.018915Z","shell.execute_reply":"2022-11-07T20:37:48.018933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nt0= time.time()\nmodel = Ridge(alpha=1) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'Ridge Alpha1', t0 ) # Updates df_models_stat\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.020316Z","iopub.status.idle":"2022-11-07T20:37:48.020762Z","shell.execute_reply.started":"2022-11-07T20:37:48.020563Z","shell.execute_reply":"2022-11-07T20:37:48.020583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.021928Z","iopub.status.idle":"2022-11-07T20:37:48.022431Z","shell.execute_reply.started":"2022-11-07T20:37:48.022197Z","shell.execute_reply":"2022-11-07T20:37:48.022217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nt0 = time.time()\nmodel = Ridge(alpha=1e4) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'Ridge Alpha1e4', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.024015Z","iopub.status.idle":"2022-11-07T20:37:48.024428Z","shell.execute_reply.started":"2022-11-07T20:37:48.02421Z","shell.execute_reply":"2022-11-07T20:37:48.024229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nt0 = time.time()\nmodel = Ridge(alpha=1e5) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'Ridge Alpha1e5', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.025468Z","iopub.status.idle":"2022-11-07T20:37:48.025844Z","shell.execute_reply.started":"2022-11-07T20:37:48.02566Z","shell.execute_reply":"2022-11-07T20:37:48.025679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.026842Z","iopub.status.idle":"2022-11-07T20:37:48.027226Z","shell.execute_reply.started":"2022-11-07T20:37:48.027041Z","shell.execute_reply":"2022-11-07T20:37:48.02706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Random Forest","metadata":{}},{"cell_type":"code","source":"%%time\nprint('Default params')\nt0 = time.time()\nfrom sklearn.ensemble import RandomForestRegressor\nmodel = RandomForestRegressor() # max_depth=20, n_estimators=1000 ,  random_state=0,\n     #min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'RandForest Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.029621Z","iopub.status.idle":"2022-11-07T20:37:48.030025Z","shell.execute_reply.started":"2022-11-07T20:37:48.029826Z","shell.execute_reply":"2022-11-07T20:37:48.029844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.031282Z","iopub.status.idle":"2022-11-07T20:37:48.032049Z","shell.execute_reply.started":"2022-11-07T20:37:48.031829Z","shell.execute_reply":"2022-11-07T20:37:48.031851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.ensemble import RandomForestRegressor\nmodel = RandomForestRegressor(max_depth=20, n_estimators=1000 ,  random_state=0,\n     min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'RandForest Tuned', t0 ) # Updates df_models_stat\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.033691Z","iopub.status.idle":"2022-11-07T20:37:48.03414Z","shell.execute_reply.started":"2022-11-07T20:37:48.033934Z","shell.execute_reply":"2022-11-07T20:37:48.033955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.037171Z","iopub.status.idle":"2022-11-07T20:37:48.037671Z","shell.execute_reply.started":"2022-11-07T20:37:48.037413Z","shell.execute_reply":"2022-11-07T20:37:48.037441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Results: \n# Best:  {'max_depth': 20, 'n_estimators': 500} r2: 0.433280 mse 23.086863 time overall 264.60\n# {'max_depth': None, 'n_estimators': 500} r2: 0.429443 mse 23.243160 time 31.37\n# {'max_depth': None, 'n_estimators': 1000} r2: 0.431275 mse 23.168528 time 62.94\n# {'max_depth': 10, 'n_estimators': 500} r2: 0.431057 mse 23.177438 time 25.30\n# {'max_depth': 10, 'n_estimators': 1000} r2: 0.429593 mse 23.237065 time 50.75\n# {'max_depth': 20, 'n_estimators': 500} r2: 0.433280 mse 23.086863 time 31.37\n# {'max_depth': 20, 'n_estimators': 1000} r2: 0.430806 mse 23.187661 time 62.88\n# CPU times: user 4min 24s, sys: 128 ms, total: 4min 24s\n# Wall time: 4min 24s\n    \n\nfrom sklearn.model_selection import ParameterGrid\nimport time\ngrid = {'max_depth': [None, 10,20], 'n_estimators': [500,1000]}\n\nif tune_RF: \n    verbose = 10; t00 = time.time()\n    r2_best_loc = 0; params_best = {}; rmse_best = np.inf;\n    for params in list(ParameterGrid(grid)):\n        t0 = time.time()\n        model = RandomForestRegressor(**params) # RandomForestRegressor(max_depth=2, random_state=0)\n        model.fit(X_train,y_train)\n        y_pred = model.predict(X_test); r2_score_loc = r2_score(y_test, y_pred); mse_loc = mean_squared_error(y_test, y_pred)\n        if verbose >= 10: # Print information \n            print(params, 'r2: %.6f'%(r2_score_loc), 'mse %.6f'%(  mse_loc ), 'time %.2f'%(time.time()-t0)    )\n        if r2_score_loc > r2_best_loc: # Save best results \n            r2_best_loc = r2_score_loc; params_best = params.copy(); mse_best = mse_loc\n\n    print()        \n    print('Best: ', params_best, 'r2: %.6f'%(r2_best_loc), 'mse %.6f'%(  mse_best ), 'time overall %.2f'%(time.time()-t00) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.039224Z","iopub.status.idle":"2022-11-07T20:37:48.040002Z","shell.execute_reply.started":"2022-11-07T20:37:48.039792Z","shell.execute_reply":"2022-11-07T20:37:48.039813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ExtraTreesRegressor","metadata":{}},{"cell_type":"code","source":"%%time\nprint('Default params')\nfrom sklearn.ensemble import ExtraTreesRegressor\nt0 = time.time()\nmodel = ExtraTreesRegressor() # max_depth=20, n_estimators=1000 ,  random_state=0,\n     #min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'ExtraTrees Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.041729Z","iopub.status.idle":"2022-11-07T20:37:48.042108Z","shell.execute_reply.started":"2022-11-07T20:37:48.041926Z","shell.execute_reply":"2022-11-07T20:37:48.041944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.04324Z","iopub.status.idle":"2022-11-07T20:37:48.043656Z","shell.execute_reply.started":"2022-11-07T20:37:48.04345Z","shell.execute_reply":"2022-11-07T20:37:48.043469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.ensemble import ExtraTreesRegressor\nt0 = time.time()\nmodel = ExtraTreesRegressor(max_depth = None, n_estimators = 500 ) # max_depth=20, n_estimators=1000 ,  random_state=0,\n     #min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'ExtraTrees Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n        ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.045445Z","iopub.status.idle":"2022-11-07T20:37:48.045841Z","shell.execute_reply.started":"2022-11-07T20:37:48.045652Z","shell.execute_reply":"2022-11-07T20:37:48.04567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.047256Z","iopub.status.idle":"2022-11-07T20:37:48.048096Z","shell.execute_reply.started":"2022-11-07T20:37:48.047873Z","shell.execute_reply":"2022-11-07T20:37:48.047896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# n_estimators=100, *, criterion='squared_error', max_depth=None, min_samples_split=2, min_samples_leaf=1, \n# min_weight_fraction_leaf=0.0, max_features=1.0, max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=False, \n# oob_score=False, n_jobs=None, random_state=None, verbose=0, warm_start=False, ccp_alpha=0.0, max_samples=None)\n\nfrom sklearn.model_selection import ParameterGrid\nimport time\ngrid = {'max_depth': [None, 10,20], 'n_estimators': [500,1000]}\n\nverbose = 10; t00 = time.time()\nr2_best_loc = 0; params_best = {}; rmse_best = np.inf;\nfor params in list(ParameterGrid(grid)):\n    t0 = time.time()\n    model = ExtraTreesRegressor(**params) # RandomForestRegressor(max_depth=2, random_state=0)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test); r2_score_loc = r2_score(y_test, y_pred); mse_loc = mean_squared_error(y_test, y_pred)\n    if verbose >= 10: # Print information \n        print(params, 'r2: %.6f'%(r2_score_loc), 'mse %.6f'%(  mse_loc ), 'time %.2f'%(time.time()-t0)    )\n    if r2_score_loc > r2_best_loc: # Save best results \n        r2_best_loc = r2_score_loc; params_best = params.copy(); mse_best = mse_loc\n        \nprint()        \nprint('Best: ', params_best, 'r2: %.6f'%(r2_best_loc), 'mse %.6f'%(  mse_best ), 'time overall %.2f'%(time.time()-t00) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.049337Z","iopub.status.idle":"2022-11-07T20:37:48.049772Z","shell.execute_reply.started":"2022-11-07T20:37:48.049586Z","shell.execute_reply":"2022-11-07T20:37:48.049605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SVR (support vector machine regression)","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.svm import SVR\nt0 = time.time()\nmodel = SVR()#C=1.9, epsilon=0.000) # RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'SVR Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.052098Z","iopub.status.idle":"2022-11-07T20:37:48.052849Z","shell.execute_reply.started":"2022-11-07T20:37:48.052643Z","shell.execute_reply":"2022-11-07T20:37:48.052664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.svm import SVR\nt0 = time.time()\nmodel = SVR(C=1.9, epsilon=0.000) # RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'SVR Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.054403Z","iopub.status.idle":"2022-11-07T20:37:48.054797Z","shell.execute_reply.started":"2022-11-07T20:37:48.054612Z","shell.execute_reply":"2022-11-07T20:37:48.054631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.055877Z","iopub.status.idle":"2022-11-07T20:37:48.056936Z","shell.execute_reply.started":"2022-11-07T20:37:48.056669Z","shell.execute_reply":"2022-11-07T20:37:48.056695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.model_selection import ParameterGrid\nimport time\ngrid = {'C': [0.5,1.,1.8,1.9,2,2.1,3], 'epsilon': [0,.5, 1,2]}\n\nverbose = 10; t00 = time.time()\nr2_best_loc = 0; params_best = {}; rmse_best = np.inf;\nfor params in list(ParameterGrid(grid)):\n    t0 = time.time()\n    model = SVR(**params) # RandomForestRegressor(max_depth=2, random_state=0)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test); r2_score_loc = r2_score(y_test, y_pred); mse_loc = mean_squared_error(y_test, y_pred)\n    if verbose >= 10: # Print information \n        print(params, 'r2: %.6f'%(r2_score_loc), 'mse %.6f'%(  mse_loc ), 'time %.2f'%(time.time()-t0)    )\n    if r2_score_loc > r2_best_loc: # Save best results \n        r2_best_loc = r2_score_loc; params_best = params.copy(); mse_best = mse_loc\n        \nprint()        \nprint('Best: ', params_best, 'r2: %.6f'%(r2_best_loc), 'mse %.6f'%(  r2_best_loc ), 'time overall %.2f'%(time.time()-t00) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.058349Z","iopub.status.idle":"2022-11-07T20:37:48.058793Z","shell.execute_reply.started":"2022-11-07T20:37:48.058604Z","shell.execute_reply":"2022-11-07T20:37:48.058623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KernelRidge","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.kernel_ridge import KernelRidge\nprint('Default params:')\nt0 = time.time()\nmodel = KernelRidge()# alpha=10000.01)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'KernelRidge Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.060115Z","iopub.status.idle":"2022-11-07T20:37:48.061035Z","shell.execute_reply.started":"2022-11-07T20:37:48.060803Z","shell.execute_reply":"2022-11-07T20:37:48.060825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.gaussian_process.kernels import RBF\nfrom sklearn.kernel_ridge import KernelRidge\nkernel = RBF(length_scale = 10)\nt0 = time.time()\nmodel = KernelRidge(alpha=1, kernel=kernel)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'KernelRidge Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.06277Z","iopub.status.idle":"2022-11-07T20:37:48.063165Z","shell.execute_reply.started":"2022-11-07T20:37:48.06298Z","shell.execute_reply":"2022-11-07T20:37:48.062998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.06486Z","iopub.status.idle":"2022-11-07T20:37:48.065228Z","shell.execute_reply.started":"2022-11-07T20:37:48.065049Z","shell.execute_reply":"2022-11-07T20:37:48.065066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.model_selection import ParameterGrid\nimport time\ngrid = {'length_scale': [1,10,100], 'alpha': [0,1,10,100] }#list(range(1,20))+[20,30,50,100,200]}\n\nverbose = 10; t00 = time.time()\nr2_best_loc = 0; params_best = {}; rmse_best = np.inf;\nfor params in list(ParameterGrid(grid)):\n    t0 = time.time()\n    kernel = RBF(length_scale = params['length_scale'])\n    model = KernelRidge(alpha = params['alpha'])# alpha=0.2, kernel=kernel)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test); r2_score_loc = r2_score(y_test, y_pred); mse_loc = mean_squared_error(y_test, y_pred)\n    if verbose >= 10: # Print information \n        print(params, 'r2: %.6f'%(r2_score_loc), 'mse %.6f'%(  mse_loc ), 'time %.2f'%(time.time()-t0)    )\n    if r2_score_loc > r2_best_loc: # Save best results \n        r2_best_loc = r2_score_loc; params_best = params.copy(); mse_best = mse_loc\n        \nprint()        \nprint('Best: ', params_best, 'r2: %.6f'%(r2_best_loc), 'mse %.6f'%(  r2_best_loc ), 'time overall %.2f'%(time.time()-t00) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.06657Z","iopub.status.idle":"2022-11-07T20:37:48.066948Z","shell.execute_reply.started":"2022-11-07T20:37:48.066767Z","shell.execute_reply":"2022-11-07T20:37:48.066785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KNeighborsRegressor","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.neighbors import KNeighborsRegressor\n\nmodel = KNeighborsRegressor(n_neighbors=15)\nt0 = time.time()\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'KNNRegres Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.068957Z","iopub.status.idle":"2022-11-07T20:37:48.069328Z","shell.execute_reply.started":"2022-11-07T20:37:48.069146Z","shell.execute_reply":"2022-11-07T20:37:48.069164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.07056Z","iopub.status.idle":"2022-11-07T20:37:48.070935Z","shell.execute_reply.started":"2022-11-07T20:37:48.070752Z","shell.execute_reply":"2022-11-07T20:37:48.07077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.model_selection import ParameterGrid\nimport time\ngrid = {'n_neighbors': list(range(1,20))+[20,30,50,100,200]}\n\nverbose = 10; t00 = time.time()\nr2_best_loc = 0; params_best = {}; rmse_best = np.inf;\nfor params in list(ParameterGrid(grid)):\n    t0 = time.time()\n    model = KNeighborsRegressor(**params) # RandomForestRegressor(max_depth=2, random_state=0)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test); r2_score_loc = r2_score(y_test, y_pred); mse_loc = mean_squared_error(y_test, y_pred)\n    if verbose >= 10: # Print information \n        print(params, 'r2: %.6f'%(r2_score_loc), 'mse %.6f'%(  mse_loc ), 'time %.2f'%(time.time()-t0)    )\n    if r2_score_loc > r2_best_loc: # Save best results \n        r2_best_loc = r2_score_loc; params_best = params.copy(); mse_best = mse_loc\n        \nprint()        \nprint('Best: ', params_best, 'r2: %.6f'%(r2_best_loc), 'mse %.6f'%(  r2_best_loc ), 'time overall %.2f'%(time.time()-t00) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.075914Z","iopub.status.idle":"2022-11-07T20:37:48.076389Z","shell.execute_reply.started":"2022-11-07T20:37:48.076168Z","shell.execute_reply":"2022-11-07T20:37:48.076189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# XGBoost , with Optuna","metadata":{}},{"cell_type":"code","source":"%%time\nimport xgboost as xgb\nprint('Default params')\nt0 = time.time()\nmodel = xgb.XGBRegressor()# tree_method=\"gpu_hist\")(n_neighbors=15)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'XGBoost Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.077812Z","iopub.status.idle":"2022-11-07T20:37:48.07822Z","shell.execute_reply.started":"2022-11-07T20:37:48.07802Z","shell.execute_reply":"2022-11-07T20:37:48.07804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.080448Z","iopub.status.idle":"2022-11-07T20:37:48.080889Z","shell.execute_reply.started":"2022-11-07T20:37:48.080696Z","shell.execute_reply":"2022-11-07T20:37:48.080716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.082677Z","iopub.status.idle":"2022-11-07T20:37:48.083121Z","shell.execute_reply.started":"2022-11-07T20:37:48.082927Z","shell.execute_reply":"2022-11-07T20:37:48.082947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Simplest example to use optuna: https://optuna.org/#code_examples\n# \n# import optuna\n# def objective(trial):\n#     x = trial.suggest_float('x', -10, 10)\n#     return (x - 2) ** 2\n# study = optuna.create_study()\n# study.optimize(objective, n_trials=100)\n# study.best_params  # E.g. {'x': 2.002108042}\n\n# Some example of Optuna with Lightgbm\n# https://www.kaggle.com/code/xiafire/lb0-830-lgbm-optuna-msci-citeseq\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.08524Z","iopub.status.idle":"2022-11-07T20:37:48.085693Z","shell.execute_reply.started":"2022-11-07T20:37:48.085487Z","shell.execute_reply":"2022-11-07T20:37:48.085507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna\n\ndef objective(trial):\n    params = {\n        'n_estimators' :  trial.suggest_categorical('n_estimators', [100, 200, 500]) ,\n       'max_depth': trial.suggest_categorical('max_depth', [6, 10,20,100]),\n#        'max_leaves' : trial.suggest_int('max_leaves', 0, 1000),\n#         'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n    }\n    \n    model = xgb.XGBRegressor(**params)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    mse = mean_squared_error(y_test, y_pred)\n    \n    return mse","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.087505Z","iopub.status.idle":"2022-11-07T20:37:48.087967Z","shell.execute_reply.started":"2022-11-07T20:37:48.087775Z","shell.execute_reply":"2022-11-07T20:37:48.087795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfind_params = True\nif find_params:\n#     study = optuna.create_study(\n#         direction='minimize', \n#         pruner=optuna.pruners.MedianPruner(n_warmup_steps=20),\n#         study_name='small')\n#     study.optimize(objective, n_trials=20)\n    study = optuna.create_study()\n    study.optimize(objective, n_trials=10)\n    \n    # Output for best found params: \n    print(); print('Best params:')\n    print(study.best_params)  # E.g. {'x': 2.002108042}\n    t0  = time.time()\n    model = xgb.XGBRegressor(**study.best_params)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    \n    update_models_stat(model, 'XGBoost Tuned', t0 ) # Updates df_models_stat\n\n    print('Best r2:%.6f'%r2_score(y_test, y_pred),'Mse:%.3f'%mean_squared_error(y_test, y_pred) )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.089785Z","iopub.status.idle":"2022-11-07T20:37:48.090234Z","shell.execute_reply.started":"2022-11-07T20:37:48.090038Z","shell.execute_reply":"2022-11-07T20:37:48.090058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Catboost","metadata":{}},{"cell_type":"code","source":"%%time\nfrom catboost import CatBoostRegressor\nprint('Default params')\nt0 = time.time()\nmodel = CatBoostRegressor(verbose = 0 ) # iterations=2, learning_rate=1, depth=2)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'CatBoost Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.091739Z","iopub.status.idle":"2022-11-07T20:37:48.092186Z","shell.execute_reply.started":"2022-11-07T20:37:48.091983Z","shell.execute_reply":"2022-11-07T20:37:48.092001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.093693Z","iopub.status.idle":"2022-11-07T20:37:48.094439Z","shell.execute_reply.started":"2022-11-07T20:37:48.0942Z","shell.execute_reply":"2022-11-07T20:37:48.094221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_all_params()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.095927Z","iopub.status.idle":"2022-11-07T20:37:48.09639Z","shell.execute_reply.started":"2022-11-07T20:37:48.096174Z","shell.execute_reply":"2022-11-07T20:37:48.096193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# MLPRegressor","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.neural_network import MLPRegressor\nprint('Default params')\nt0 = time.time()\nmodel = MLPRegressor(random_state=1, max_iter=500)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'MLP Default', t0 ) # Updates df_models_stat\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.09765Z","iopub.status.idle":"2022-11-07T20:37:48.098131Z","shell.execute_reply.started":"2022-11-07T20:37:48.097903Z","shell.execute_reply":"2022-11-07T20:37:48.097921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.100616Z","iopub.status.idle":"2022-11-07T20:37:48.101032Z","shell.execute_reply.started":"2022-11-07T20:37:48.100832Z","shell.execute_reply":"2022-11-07T20:37:48.10085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.neural_network import MLPRegressor\nt0 = time.time()\nprint('Aiguls params for 0.806') # https://www.kaggle.com/code/user327934/mmscel-crossvalidation-schemes?scriptVersionId=108617498&cellId=35\nmodel = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=42, \n                         hidden_layer_sizes=(300, 200))\n\nmodel.fit(X_train.values,y_train.values)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'MLP Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.102331Z","iopub.status.idle":"2022-11-07T20:37:48.102797Z","shell.execute_reply.started":"2022-11-07T20:37:48.1026Z","shell.execute_reply":"2022-11-07T20:37:48.102619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.105237Z","iopub.status.idle":"2022-11-07T20:37:48.10567Z","shell.execute_reply.started":"2022-11-07T20:37:48.105465Z","shell.execute_reply":"2022-11-07T20:37:48.105486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blend","metadata":{}},{"cell_type":"code","source":"predictions_test = np.zeros( (len(y_test), len(df_models_stat ))  )\npredictions_test2 = np.zeros( (len(y_test2), len(df_models_stat ))  )\npredictions_test3_oop = np.zeros( (len(y_oop), len(df_models_stat ))  )\n\nfor i,model_ID in enumerate( dict_save_predictions):\n    #print(i,model_ID)\n    tuple_preds = dict_save_predictions[model_ID]\n    predictions_test[:,i] = tuple_preds[1]\n    predictions_test2[:,i] = tuple_preds[2]\n    predictions_test3_oop[:,i] = tuple_preds[3]\npredictions_test\ny_blend = predictions_test[:,:4].mean(axis = 1 )\ny_pred = y_blend\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.106853Z","iopub.status.idle":"2022-11-07T20:37:48.107255Z","shell.execute_reply.started":"2022-11-07T20:37:48.107057Z","shell.execute_reply":"2022-11-07T20:37:48.107076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\ncm = pd.DataFrame( predictions_test , columns =  df_models_stat.index ).corr()\nplt.figure( figsize = (20,10) )\nsns.heatmap( cm )\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.108578Z","iopub.status.idle":"2022-11-07T20:37:48.109008Z","shell.execute_reply.started":"2022-11-07T20:37:48.108814Z","shell.execute_reply":"2022-11-07T20:37:48.108832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlist_models_ids = list( dict_save_predictions.keys() ) # or list(df_models_stat.index )  \n\n#sklearn.linear_model.LinearRegression\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import Lasso\nreg = LinearRegression()\nreg = Lasso(alpha = 0.01)\nreg = Lasso(alpha = 0.01, positive = True)\n\nreg.fit(predictions_test, y_test)\nprint(reg.coef_)\ny_pred = reg.predict(predictions_test)\nprint('R2:', r2_score(y_test, y_pred), 'MSE:', mean_squared_error(y_test, y_pred) )\ny_pred = reg.predict(predictions_test2)\nprint('Test2 R2:', r2_score(y_test2, y_pred), 'MSE:', mean_squared_error(y_test2, y_pred) )\ny_pred = reg.predict(predictions_test3_oop)\nprint('OOP R2:', r2_score(y_oop, y_pred), 'OOP MSE:', mean_squared_error(y_oop, y_pred) )\n\ndd = pd.DataFrame(index = df_models_stat.index, data = reg.coef_, columns = ['Blend Coef'] )\n#dd['r2_score'] = df_models_stat['r2_score']\ndd = dd.join(df_models_stat)\ndisplay(dd.sort_values('r2_score', ascending = False))\n#print(r2_score(y_test, y_pred), mean_squared_error(y_test, y_pred) )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.110631Z","iopub.status.idle":"2022-11-07T20:37:48.111214Z","shell.execute_reply.started":"2022-11-07T20:37:48.110844Z","shell.execute_reply":"2022-11-07T20:37:48.110863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprint('Optuna cannot find reasonable blend coefficients at least in 100 trials ')\n\nlist_models_ids = list( dict_save_predictions.keys() ) # or list(df_models_stat.index )  \n\noptuna.logging.set_verbosity(optuna.logging.WARNING)\ndef objective(trial):\n    \n    # 2. Suggest values of the hyperparameters using a trial object.\n    vec_coefs = np.zeros(predictions_test.shape[1])#    len(df_models_stat )\n\n    for i in range( len( vec_coefs )):\n        model_ID = list_models_ids[i]\n        vec_coefs = trial.suggest_float(model_ID,0, 1) #  1e-8, 10.0, log=True)\n    y_pred = (predictions_test*vec_coefs).mean(axis = 1)\n    mse = mean_squared_error(y_test, y_pred)\n    return mse\n\nstudy = optuna.create_study(direction='minimize')# 'maximize')\nstudy.optimize(objective, n_trials=100)#\n\nprint(); print('Best params:')\nprint(study.best_params)  # E.g. {'x': 2.002108042}\ny_pred = (predictions_test*np.array(list(study.best_params.values())) ).mean(axis = 1)\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.113088Z","iopub.status.idle":"2022-11-07T20:37:48.113546Z","shell.execute_reply.started":"2022-11-07T20:37:48.113311Z","shell.execute_reply":"2022-11-07T20:37:48.11333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show summary stat","metadata":{}},{"cell_type":"code","source":"df_models_stat.sort_values('r2_score',ascending = False)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.114877Z","iopub.status.idle":"2022-11-07T20:37:48.115298Z","shell.execute_reply.started":"2022-11-07T20:37:48.115108Z","shell.execute_reply":"2022-11-07T20:37:48.115127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat.sort_values('r2_score',ascending = False).to_csv('df_models_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.116533Z","iopub.status.idle":"2022-11-07T20:37:48.11691Z","shell.execute_reply.started":"2022-11-07T20:37:48.116723Z","shell.execute_reply":"2022-11-07T20:37:48.11674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:37:48.117883Z","iopub.status.idle":"2022-11-07T20:37:48.118288Z","shell.execute_reply.started":"2022-11-07T20:37:48.118099Z","shell.execute_reply":"2022-11-07T20:37:48.118117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}