{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 9 Blend and BIG SURPRISE(!) - how to explain ? ideas - welcome ! \n\n    added blending part for all solutions - and got suprise:\n    and the WORST/TERRIBLE (r2 negative, mse solution 10 times worse than others -   KernelRidge\n    enters the blends and improve it ! \n    \n    The only thing - that it is very uncorrelated with the other solutions. \n\n#### 8: Added - statistics on all methods - collected in one table (dataframe)\n    \n    Still Ridge, SVR are the best. Cat\n\n#### 7: added XGBoost+Optuna, CatBoost, MLP\n    \n    CatBoost is quite good with default params\n\n#### 3,4,5,6 - search for optimal LightGBM params; added: models and param tuning for other models,\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n    Version 6 - seems around 40 PCA features is better for LightGBM  - current_best_r2 = 0.46792863038644295 # \n    \n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"code","source":"tune_RF = False # True # Takes 4 minutes for 6 params ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:28:37.894226Z","iopub.execute_input":"2022-11-07T16:28:37.894617Z","iopub.status.idle":"2022-11-07T16:28:37.899344Z","shell.execute_reply.started":"2022-11-07T16:28:37.894583Z","shell.execute_reply":"2022-11-07T16:28:37.897931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-07T16:28:40.178677Z","iopub.execute_input":"2022-11-07T16:28:40.179041Z","iopub.status.idle":"2022-11-07T16:28:40.190191Z","shell.execute_reply.started":"2022-11-07T16:28:40.179010Z","shell.execute_reply":"2022-11-07T16:28:40.188954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:28:44.322993Z","iopub.execute_input":"2022-11-07T16:28:44.323785Z","iopub.status.idle":"2022-11-07T16:29:04.498216Z","shell.execute_reply.started":"2022-11-07T16:28:44.323743Z","shell.execute_reply":"2022-11-07T16:29:04.497076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:04.500856Z","iopub.execute_input":"2022-11-07T16:29:04.501231Z","iopub.status.idle":"2022-11-07T16:29:24.421752Z","shell.execute_reply.started":"2022-11-07T16:29:04.501194Z","shell.execute_reply":"2022-11-07T16:29:24.420600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:24.423765Z","iopub.execute_input":"2022-11-07T16:29:24.424149Z","iopub.status.idle":"2022-11-07T16:29:33.977614Z","shell.execute_reply.started":"2022-11-07T16:29:24.424110Z","shell.execute_reply":"2022-11-07T16:29:33.976365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.mean(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:33.980823Z","iopub.execute_input":"2022-11-07T16:29:33.981618Z","iopub.status.idle":"2022-11-07T16:29:34.147609Z","shell.execute_reply.started":"2022-11-07T16:29:33.981572Z","shell.execute_reply":"2022-11-07T16:29:34.146453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.std(axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:34.149136Z","iopub.execute_input":"2022-11-07T16:29:34.149902Z","iopub.status.idle":"2022-11-07T16:29:34.810300Z","shell.execute_reply.started":"2022-11-07T16:29:34.149845Z","shell.execute_reply":"2022-11-07T16:29:34.809141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:34.811888Z","iopub.execute_input":"2022-11-07T16:29:34.812502Z","iopub.status.idle":"2022-11-07T16:29:35.501359Z","shell.execute_reply.started":"2022-11-07T16:29:34.812459Z","shell.execute_reply":"2022-11-07T16:29:35.499741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.502717Z","iopub.execute_input":"2022-11-07T16:29:35.503173Z","iopub.status.idle":"2022-11-07T16:29:35.758726Z","shell.execute_reply.started":"2022-11-07T16:29:35.503136Z","shell.execute_reply":"2022-11-07T16:29:35.757691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\"  (i.e. downsample)\n\nSmall 10% of data of data where we can make prelimanary experiments. \n\nit will be labeled by special column \"Playground\" in df_meta (meta data for CITE-seq train only)","metadata":{}},{"cell_type":"code","source":"# Prepare for creation of additional holdout folds with 10% of samples \n# We will use stratified Kfold to achieve that days, cell_types and donors are equally distributed \nimport numpy as np\nfrom sklearn.model_selection import StratifiedKFold\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n\nskf = StratifiedKFold(n_splits=2,  shuffle=True, random_state=40)\nskf.get_n_splits(df_meta, df_meta[scol] )\n\n\n\ny = df_meta[scol] \nfor train_index, test_index in skf.split(df_meta, df_meta[scol]):\n    print(\"TRAIN:\", len(train_index), \"TEST:\", len(test_index) ); \n    break\nprint(test_index)\nprint(df_meta[scol].value_counts().head(5)   )\nprint(df_meta.iloc[test_index,:][scol].value_counts().head(5)    )\n\n\nflagged_column_name = 'Playground'\ndf_meta[flagged_column_name] = 0 \ndf_meta.loc[df_meta.index[test_index],flagged_column_name]  = 1\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.760396Z","iopub.execute_input":"2022-11-07T16:29:35.761084Z","iopub.status.idle":"2022-11-07T16:29:35.952980Z","shell.execute_reply.started":"2022-11-07T16:29:35.761047Z","shell.execute_reply":"2022-11-07T16:29:35.951940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score\n\ndf_models_stat = pd.DataFrame()# columns = [ 'r2_score','mse','Time',  'n_feat', 'Target' ])","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.954512Z","iopub.execute_input":"2022-11-07T16:29:35.955549Z","iopub.status.idle":"2022-11-07T16:29:35.961404Z","shell.execute_reply.started":"2022-11-07T16:29:35.955492Z","shell.execute_reply":"2022-11-07T16:29:35.960223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna\n\nimport lightgbm as lgbm\nimport xgboost as xgb","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.965213Z","iopub.execute_input":"2022-11-07T16:29:35.965610Z","iopub.status.idle":"2022-11-07T16:29:35.974121Z","shell.execute_reply.started":"2022-11-07T16:29:35.965574Z","shell.execute_reply":"2022-11-07T16:29:35.973039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def objective_lgbm(trial):\n    params = {\n        'boosting_type': trial.suggest_categorical('boosting_type', ['gbdt', 'dart', 'goss']),\n        'num_leaves': trial.suggest_categorical('num_leaves', list(range(20, 100))),\n        'max_depth': trial.suggest_categorical('max_depth', list(range(6, 100))),\n        'learning_rate': trial.suggest_float('learning_rate', 1e-8, 1e-1),\n        'n_estimators' :  trial.suggest_categorical('n_estimators', list(range(50, 501, 25))), \n        'class_weight' :  trial.suggest_categorical('class_weight', ['balanced', None]), \n        'reg_alpha': trial.suggest_float('reg_alpha', 0, 5),\n        'reg_lambda': trial.suggest_float('reg_lambda', 0, 5),\n    }\n    \n    model = lgbm.LGBMRegressor(**params)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    mse = mean_squared_error(y_test, y_pred)\n    \n    return mse","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.975618Z","iopub.execute_input":"2022-11-07T16:29:35.976051Z","iopub.status.idle":"2022-11-07T16:29:35.986225Z","shell.execute_reply.started":"2022-11-07T16:29:35.976012Z","shell.execute_reply":"2022-11-07T16:29:35.985255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def objective_xgb(trial):\n    params = {\n        'n_estimators' :  trial.suggest_categorical('n_estimators', list(range(50, 501, 25))), \n        'max_depth': trial.suggest_categorical('max_depth', list(range(6, 21))),\n        'learning_rate': trial.suggest_float('learning_rate', 1e-8, 1e-1),\n        'gamma': trial.suggest_float('gamma', 0, 5),\n        'base_score': trial.suggest_float('base_score', 0, 2),\n        'tree_method': trial.suggest_categorical('tree_method', ['gpu_hist']),\n        'predictor': trial.suggest_categorical('predictor', ['gpu_predictor']),\n        'max_leaves' : trial.suggest_int('max_leaves', 0, 1000)\n#         'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n    }\n    \n    model = xgb.XGBRegressor(**params)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    mse = mean_squared_error(y_test, y_pred)\n    \n    return mse","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:35.987807Z","iopub.execute_input":"2022-11-07T16:29:35.988223Z","iopub.status.idle":"2022-11-07T16:29:36.001975Z","shell.execute_reply.started":"2022-11-07T16:29:35.988149Z","shell.execute_reply":"2022-11-07T16:29:36.000927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict_save_predictions = {}\ndef update_models_stat(model, model_ID, t0 ):\n    \n    IXmdl = model_ID # df_models_stat.shape[0] + 1\n    #df_models_stat.loc[IXmdl,'Model'] = model_ID\n    y_pred_loc = model.predict(X_test)\n    df_models_stat.loc[IXmdl,'r2_score'] = r2_score(y_test, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse'] =  mean_squared_error(y_test, y_pred_loc)\n    df_models_stat.loc[IXmdl,'Time'] = np.round( time.time() - t0,3)\n    df_models_stat.loc[IXmdl,'Target'] = selected_target\n    df_models_stat.loc[IXmdl,'n_feat'] = X_train.shape[1]\n    df_models_stat.loc[IXmdl,'n_samples_train'] = X_train.shape[0]\n    \n    y_pred_loc = model.predict(X_test_public_like)\n    df_models_stat.loc[IXmdl,'r2_score Test2 PublLike'] = r2_score(y_test_public_like, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Test2 PublLike'] =  mean_squared_error(y_test_public_like, y_pred_loc)\n    \n    y_pred_loc = model.predict(X_oop)\n    df_models_stat.loc[IXmdl,'r2_score Out of Playgr'] = r2_score(y_oop, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Out of Playgr'] =  mean_squared_error(y_oop, y_pred_loc)\n\n    y_pred_loc = model.predict(X_train)\n    df_models_stat.loc[IXmdl,'r2_score Train'] = r2_score(y_train, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Train'] =  mean_squared_error(y_train, y_pred_loc)\n    \n    dict_save_predictions[model_ID] = (model.predict(X_train), model.predict(X_test), \n                                       model.predict(X_test_public_like), model.predict(X_oop)   )\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:36.003316Z","iopub.execute_input":"2022-11-07T16:29:36.005293Z","iopub.status.idle":"2022-11-07T16:29:36.015619Z","shell.execute_reply.started":"2022-11-07T16:29:36.005256Z","shell.execute_reply":"2022-11-07T16:29:36.014855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"code","source":"# all_targets = ['CD31', 'CD44', 'CD71', 'CD244', 'CD49d', 'CD36', 'CD29', 'CD41', 'CD11a', 'CD62L' ,'CD3' ,'CD196', \n#               'CD23', 'CD16', 'CD24', 'IgD', 'CD27', 'CD94', 'CD79b', 'Rat-IgG1', 'CD25', 'CD62P', 'CD45RO', 'CD161', 'CD20']\n\nall_targets = ['CD161', 'CD20']\nn_trails = 50\n\n# best_parameters_per_target = pd.DataFrame(columns=['Model', 'Target', 'r2', 'mse', 'best_params'])\n\n# best_params_lgbm = {}\n# best_params_xgb = {}\n\nfor selected_target in all_targets:\n    n_features = 40\n\n    X= df_cite.iloc[:70988,:n_features][df_meta['Playground']==1]\n    y= df_cite_train_y[df_meta['Playground']==1][selected_target]\n\n    # Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n    # Private like test - new DAY, and donor, \n    # While public like - only new donor (days are the same as in train):\n    # Step 1: \n    mask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \n    X_train = df_cite.iloc[:70988,:n_features][mask_train]\n    y_train = df_cite_train_y[mask_train][ selected_target ]\n    # Step 2:\n    mask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\n    X_test_private_like = df_cite.iloc[:70988,:n_features][ mask_test_private_like  ]\n    y_test_private_like = df_cite_train_y[mask_test_private_like][ selected_target ]\n    X_test = X_test_private_like\n    y_test = y_test_private_like\n    # Step 3: \n    mask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \n    X_test_public_like = df_cite.iloc[:70988,:n_features][mask_test_public_like ]\n    y_test_public_like = df_cite_train_y[mask_test_public_like][ selected_target ]\n    X_test2 = X_test_public_like\n    y_test2 = y_test_public_like\n\n    mask_out_of_playground = (df_meta['Playground']==0)\n    X_oop = df_cite.iloc[:70988,:n_features][mask_out_of_playground]\n    y_oop = df_cite_train_y[mask_out_of_playground][ selected_target ]\n\n    #optuna for LGBM\n    study = optuna.create_study()\n    study.optimize(objective_lgbm, n_trials=n_trails)\n    model = lgbm.LGBMRegressor(**study.best_params)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n\n    row = best_parameters_per_target.shape[0]\n    best_parameters_per_target.loc[row, 'Model'] = 'LGBM'\n    best_parameters_per_target.loc[row, 'Target'] = selected_target\n    best_parameters_per_target.loc[row, 'r2'] = r2_score(y_test, y_pred)\n    best_parameters_per_target.loc[row, 'mse'] = mean_squared_error(y_test, y_pred)\n    best_parameters_per_target.loc[row, 'best_params'] = str(study.best_params)\n    best_params_lgbm[selected_target] = study.best_params\n                       \n    #optuna XGB\n    study = optuna.create_study()\n    study.optimize(objective_xgb, n_trials=n_trails)\n    model = xgb.XGBRegressor(**study.best_params)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    \n    row = best_parameters_per_target.shape[0]\n    best_parameters_per_target.loc[row, 'Model'] = 'XGB'\n    best_parameters_per_target.loc[row, 'Target'] = selected_target\n    best_parameters_per_target.loc[row, 'r2'] = r2_score(y_test, y_pred)\n    best_parameters_per_target.loc[row, 'mse'] = mean_squared_error(y_test, y_pred)\n    best_parameters_per_target.loc[row, 'best_params'] = str(study.best_params)\n    best_params_xgb[selected_target] = study.best_params\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T20:48:03.276692Z","iopub.execute_input":"2022-11-07T20:48:03.277056Z","iopub.status.idle":"2022-11-07T21:01:38.692077Z","shell.execute_reply.started":"2022-11-07T20:48:03.277023Z","shell.execute_reply":"2022-11-07T21:01:38.691260Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params_xgb","metadata":{"execution":{"iopub.status.busy":"2022-11-07T21:01:38.694679Z","iopub.execute_input":"2022-11-07T21:01:38.695615Z","iopub.status.idle":"2022-11-07T21:01:38.711764Z","shell.execute_reply.started":"2022-11-07T21:01:38.695574Z","shell.execute_reply":"2022-11-07T21:01:38.710563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params_lgbm","metadata":{"execution":{"iopub.status.busy":"2022-11-07T21:01:38.713796Z","iopub.execute_input":"2022-11-07T21:01:38.714732Z","iopub.status.idle":"2022-11-07T21:01:38.731753Z","shell.execute_reply.started":"2022-11-07T21:01:38.714692Z","shell.execute_reply":"2022-11-07T21:01:38.730567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_parameters_per_target","metadata":{"execution":{"iopub.status.busy":"2022-11-07T21:01:38.734447Z","iopub.execute_input":"2022-11-07T21:01:38.735254Z","iopub.status.idle":"2022-11-07T21:01:38.757715Z","shell.execute_reply.started":"2022-11-07T21:01:38.735216Z","shell.execute_reply":"2022-11-07T21:01:38.756779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Catboost","metadata":{}},{"cell_type":"code","source":"%%time\nfrom catboost import CatBoostRegressor\nprint('Default params')\nt0 = time.time()\nmodel = CatBoostRegressor(verbose = 0 ) # iterations=2, learning_rate=1, depth=2)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'CatBoost Default', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T10:05:15.372117Z","iopub.execute_input":"2022-11-07T10:05:15.372468Z","iopub.status.idle":"2022-11-07T10:05:26.377535Z","shell.execute_reply.started":"2022-11-07T10:05:15.372438Z","shell.execute_reply":"2022-11-07T10:05:26.376400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_params()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def objective(trial):\n    params = {\n        'iterations' :  trial.suggest_categorical('iterations', list(range(50, 501, 25))), \n        'learning_rate': trial.suggest_float('learning_rate', 1e-8, 1e-1),\n        'max_depth' : trial.suggest_int('max_depth', 0, 16),\n#         'n_estimators' :  trial.suggest_categorical('n_estimators', list(range(50, 501, 25))), \n        'l2_leaf_reg': trial.suggest_float('l2_leaf_reg', 0, 10),\n        'early_stopping_rounds' :  trial.suggest_categorical('early_stopping_rounds', list(range(50, 200, 25))), \n\n    }\n    \n    model = CatBoostRegressor(**params, verbose=False)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    mse = mean_squared_error(y_test, y_pred)\n    \n    return mse","metadata":{"execution":{"iopub.status.busy":"2022-11-07T11:02:25.942984Z","iopub.execute_input":"2022-11-07T11:02:25.943346Z","iopub.status.idle":"2022-11-07T11:02:25.950857Z","shell.execute_reply.started":"2022-11-07T11:02:25.943317Z","shell.execute_reply":"2022-11-07T11:02:25.949909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfind_params = True\nif find_params:\n#     study = optuna.create_study(\n#         direction='minimize', \n#         pruner=optuna.pruners.MedianPruner(n_warmup_steps=20),\n#         study_name='small')\n#     study.optimize(objective, n_trials=20)\n    study = optuna.create_study()\n    study.optimize(objective, n_trials=10)\n    \n    # Output for best found params: \n    print(); print('Best params:')\n    print(study.best_params)  # E.g. {'x': 2.002108042}\n    best_xgb_params = study.best_params\n    t0  = time.time()\n    model = CatBoostRegressor(**study.best_params)# tree_method=\"gpu_hist\")(n_neighbors=15)\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    \n#     update_models_stat(model, 'XGBoost Tuned', t0 ) # Updates df_models_stat\n\n    print('Best r2:%.6f'%r2_score(y_test, y_pred),'Mse:%.3f'%mean_squared_error(y_test, y_pred) )","metadata":{"execution":{"iopub.status.busy":"2022-11-07T11:02:27.506578Z","iopub.execute_input":"2022-11-07T11:02:27.506964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-11-05T21:41:21.938812Z","iopub.execute_input":"2022-11-05T21:41:21.939161Z","iopub.status.idle":"2022-11-05T21:41:21.955183Z","shell.execute_reply.started":"2022-11-05T21:41:21.939135Z","shell.execute_reply":"2022-11-05T21:41:21.953693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_all_params()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:05.704138Z","iopub.execute_input":"2022-10-31T14:16:05.704774Z","iopub.status.idle":"2022-10-31T14:16:05.717203Z","shell.execute_reply.started":"2022-10-31T14:16:05.704735Z","shell.execute_reply":"2022-10-31T14:16:05.715789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# MLPRegressor","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.neural_network import MLPRegressor\nprint('Default params')\nt0 = time.time()\nmodel = MLPRegressor(random_state=1, max_iter=500)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'MLP Default', t0 ) # Updates df_models_stat\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:05.718917Z","iopub.execute_input":"2022-10-31T14:16:05.719638Z","iopub.status.idle":"2022-10-31T14:16:22.697528Z","shell.execute_reply.started":"2022-10-31T14:16:05.719604Z","shell.execute_reply":"2022-10-31T14:16:22.696294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:22.69908Z","iopub.execute_input":"2022-10-31T14:16:22.699889Z","iopub.status.idle":"2022-10-31T14:16:22.752821Z","shell.execute_reply.started":"2022-10-31T14:16:22.699842Z","shell.execute_reply":"2022-10-31T14:16:22.751644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.neural_network import MLPRegressor\nt0 = time.time()\nprint('Aiguls params for 0.806') # https://www.kaggle.com/code/user327934/mmscel-crossvalidation-schemes?scriptVersionId=108617498&cellId=35\nmodel = MLPRegressor(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-5, random_state=42, \n                         hidden_layer_sizes=(300, 200))\n\nmodel.fit(X_train.values,y_train.values)\ny_pred = model.predict(X_test)\nupdate_models_stat(model, 'MLP Tuned', t0 ) # Updates df_models_stat\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:22.754605Z","iopub.execute_input":"2022-10-31T14:16:22.755333Z","iopub.status.idle":"2022-10-31T14:16:33.842085Z","shell.execute_reply.started":"2022-10-31T14:16:22.755285Z","shell.execute_reply":"2022-10-31T14:16:33.839344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:33.851035Z","iopub.execute_input":"2022-10-31T14:16:33.85219Z","iopub.status.idle":"2022-10-31T14:16:33.920012Z","shell.execute_reply.started":"2022-10-31T14:16:33.852118Z","shell.execute_reply":"2022-10-31T14:16:33.917284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blend","metadata":{}},{"cell_type":"code","source":"predictions_test = np.zeros( (len(y_test), len(df_models_stat ))  )\npredictions_test2 = np.zeros( (len(y_test2), len(df_models_stat ))  )\npredictions_test3_oop = np.zeros( (len(y_oop), len(df_models_stat ))  )\n\nfor i,model_ID in enumerate( dict_save_predictions):\n    #print(i,model_ID)\n    tuple_preds = dict_save_predictions[model_ID]\n    predictions_test[:,i] = tuple_preds[1]\n    predictions_test2[:,i] = tuple_preds[2]\n    predictions_test3_oop[:,i] = tuple_preds[3]\npredictions_test\ny_blend = predictions_test[:,:4].mean(axis = 1 )\ny_pred = y_blend\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:56:30.501756Z","iopub.execute_input":"2022-10-31T14:56:30.503328Z","iopub.status.idle":"2022-10-31T14:56:30.552002Z","shell.execute_reply.started":"2022-10-31T14:56:30.503126Z","shell.execute_reply":"2022-10-31T14:56:30.550773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\ncm = pd.DataFrame( predictions_test , columns =  df_models_stat.index ).corr()\nplt.figure( figsize = (20,10) )\nsns.heatmap( cm )\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:56:31.278888Z","iopub.execute_input":"2022-10-31T14:56:31.279333Z","iopub.status.idle":"2022-10-31T14:56:31.838231Z","shell.execute_reply.started":"2022-10-31T14:56:31.2793Z","shell.execute_reply":"2022-10-31T14:56:31.836674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlist_models_ids = list( dict_save_predictions.keys() ) # or list(df_models_stat.index )  \n\n#sklearn.linear_model.LinearRegression\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import Lasso\nreg = LinearRegression()\nreg = Lasso(alpha = 0.01)\nreg = Lasso(alpha = 0.01, positive = True)\n\nreg.fit(predictions_test, y_test)\nprint(reg.coef_)\ny_pred = reg.predict(predictions_test)\nprint('R2:', r2_score(y_test, y_pred), 'MSE:', mean_squared_error(y_test, y_pred) )\ny_pred = reg.predict(predictions_test2)\nprint('Test2 R2:', r2_score(y_test2, y_pred), 'MSE:', mean_squared_error(y_test2, y_pred) )\ny_pred = reg.predict(predictions_test3_oop)\nprint('OOP R2:', r2_score(y_oop, y_pred), 'OOP MSE:', mean_squared_error(y_oop, y_pred) )\n\ndd = pd.DataFrame(index = df_models_stat.index, data = reg.coef_, columns = ['Blend Coef'] )\n#dd['r2_score'] = df_models_stat['r2_score']\ndd = dd.join(df_models_stat)\ndisplay(dd.sort_values('r2_score', ascending = False))\n#print(r2_score(y_test, y_pred), mean_squared_error(y_test, y_pred) )\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:59:26.550975Z","iopub.execute_input":"2022-10-31T14:59:26.551383Z","iopub.status.idle":"2022-10-31T14:59:26.707548Z","shell.execute_reply.started":"2022-10-31T14:59:26.551342Z","shell.execute_reply":"2022-10-31T14:59:26.706396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprint('Optuna cannot find reasonable blend coefficients at least in 100 trials ')\n\nlist_models_ids = list( dict_save_predictions.keys() ) # or list(df_models_stat.index )  \n\noptuna.logging.set_verbosity(optuna.logging.WARNING)\ndef objective(trial):\n    \n    # 2. Suggest values of the hyperparameters using a trial object.\n    vec_coefs = np.zeros(predictions_test.shape[1])#    len(df_models_stat )\n\n    for i in range( len( vec_coefs )):\n        model_ID = list_models_ids[i]\n        vec_coefs = trial.suggest_float(model_ID,0, 1) #  1e-8, 10.0, log=True)\n    y_pred = (predictions_test*vec_coefs).mean(axis = 1)\n    mse = mean_squared_error(y_test, y_pred)\n    return mse\n\nstudy = optuna.create_study(direction='minimize')# 'maximize')\nstudy.optimize(objective, n_trials=100)#\n\nprint(); print('Best params:')\nprint(study.best_params)  # E.g. {'x': 2.002108042}\ny_pred = (predictions_test*np.array(list(study.best_params.values())) ).mean(axis = 1)\n\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:34.573827Z","iopub.execute_input":"2022-10-31T14:16:34.575333Z","iopub.status.idle":"2022-10-31T14:16:41.31578Z","shell.execute_reply.started":"2022-10-31T14:16:34.575262Z","shell.execute_reply":"2022-10-31T14:16:41.313853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show summary stat","metadata":{}},{"cell_type":"code","source":"df_models_stat.sort_values('r2_score',ascending = False)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:41.317452Z","iopub.execute_input":"2022-10-31T14:16:41.317939Z","iopub.status.idle":"2022-10-31T14:16:41.349014Z","shell.execute_reply.started":"2022-10-31T14:16:41.31789Z","shell.execute_reply":"2022-10-31T14:16:41.347225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat.sort_values('r2_score',ascending = False).to_csv('df_models_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:41.350978Z","iopub.execute_input":"2022-10-31T14:16:41.351497Z","iopub.status.idle":"2022-10-31T14:16:41.362144Z","shell.execute_reply.started":"2022-10-31T14:16:41.351448Z","shell.execute_reply":"2022-10-31T14:16:41.360908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T14:16:41.364231Z","iopub.execute_input":"2022-10-31T14:16:41.364766Z","iopub.status.idle":"2022-10-31T14:16:41.375895Z","shell.execute_reply.started":"2022-10-31T14:16:41.364731Z","shell.execute_reply":"2022-10-31T14:16:41.374758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}