{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 12 - switch to big size of playground - half of full train - it takes long time - 1.5 hour\n\n    LightGBM - runs 0.6 seconds.\n    It still worse than Ridge, but difference is not big. \n    RF tuned - runs for 10 mintes  (500 interators)\n    Kernel Ridge - is very slow - 5 minutes - and again terrbible quality \n    XGBOost is 10 times slower with default params than Lgb, but quality is quite lower\n    Catboost default is second after the Ridge, but it takes 8 seconds, comparing to 0.6 lgb \n    \n#### 11 - cosmetic change\n    small error corrected - LGB Tuned2 params were not correct\n    Resulting table with statistcs - has been saved.  \n    \n#### 10 added parameter to control size of the play-ground\n\n    playground_fraction_of_train = 10# Should be Integer N - we create play-ground of the size Full_Train / N. \n\n#### 9 Blend and BIG SURPRISE(!) - how to explain ? ideas - welcome ! \n\n    added blending part for all solutions - and got suprise:\n    and the WORST/TERRIBLE (r2 negative, mse solution 10 times worse than others -   KernelRidge\n    enters the blends and improve it ! \n    \n    The only thing - that it is very uncorrelated with the other solutions. \n\nSee discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/363230\n\n\n#### 8: Added - statistics on all methods - collected in one table (dataframe)\n    \n    Still Ridge, SVR are the best. Cat\n\n#### 7: added XGBoost+Optuna, CatBoost, MLP\n    \n    CatBoost is quite good with default params\n\n#### 3,4,5,6 - search for optimal LightGBM params; added: models and param tuning for other models,\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n    Version 6 - seems around 40 PCA features is better for LightGBM  - current_best_r2 = 0.46792863038644295 # \n    \n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"tune_RF = False # True # It is slow - takes 4 minutes for 6 params  for play-ground of size 10% from original \nplayground_fraction_of_train = 2# Should be Integer N - we create play-ground of the size Full_Train / N. ","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:42:30.216870Z","iopub.execute_input":"2022-11-03T12:42:30.217268Z","iopub.status.idle":"2022-11-03T12:42:30.221991Z","shell.execute_reply.started":"2022-11-03T12:42:30.217237Z","shell.execute_reply":"2022-11-03T12:42:30.220778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nfrom sklearn.model_selection import StratifiedKFold\nimport optuna\n\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:42:30.228383Z","iopub.execute_input":"2022-11-03T12:42:30.229353Z","iopub.status.idle":"2022-11-03T12:42:51.381036Z","shell.execute_reply.started":"2022-11-03T12:42:30.229318Z","shell.execute_reply":"2022-11-03T12:42:51.379815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:42:51.383006Z","iopub.execute_input":"2022-11-03T12:42:51.383310Z","iopub.status.idle":"2022-11-03T12:43:11.931140Z","shell.execute_reply.started":"2022-11-03T12:42:51.383282Z","shell.execute_reply":"2022-11-03T12:43:11.929918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:11.933542Z","iopub.execute_input":"2022-11-03T12:43:11.933997Z","iopub.status.idle":"2022-11-03T12:43:22.855715Z","shell.execute_reply.started":"2022-11-03T12:43:11.933956Z","shell.execute_reply":"2022-11-03T12:43:22.854605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_cite.mean(axis = 0).head(3))\nprint('Pay attention - PCA features are ordered by magnitude:')\nprint(df_cite.std(axis=0).head(5))","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:22.857935Z","iopub.execute_input":"2022-11-03T12:43:22.858231Z","iopub.status.idle":"2022-11-03T12:43:23.646267Z","shell.execute_reply.started":"2022-11-03T12:43:22.858204Z","shell.execute_reply":"2022-11-03T12:43:23.645173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:23.647490Z","iopub.execute_input":"2022-11-03T12:43:23.647856Z","iopub.status.idle":"2022-11-03T12:43:24.385630Z","shell.execute_reply.started":"2022-11-03T12:43:23.647812Z","shell.execute_reply":"2022-11-03T12:43:24.384545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:24.387146Z","iopub.execute_input":"2022-11-03T12:43:24.387650Z","iopub.status.idle":"2022-11-03T12:43:24.651091Z","shell.execute_reply.started":"2022-11-03T12:43:24.387608Z","shell.execute_reply":"2022-11-03T12:43:24.650304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\"  (i.e. downsample)\n\nSmall 10% of data of data where we can make prelimanary experiments. \n\nit will be labeled by special column \"Playground\" in df_meta (meta data for CITE-seq train only)","metadata":{}},{"cell_type":"code","source":"# Prepare for creation of additional holdout folds with 10% of samples \n# We will use stratified Kfold to achieve that days, cell_types and donors are equally distributed \n\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n# playground_fraction_of_train = 10\nskf = StratifiedKFold(n_splits= playground_fraction_of_train,  shuffle=True, random_state=40)\nskf.get_n_splits(df_meta, df_meta[scol] )\n\n\n\ny = df_meta[scol] \nfor train_index, test_index in skf.split(df_meta, df_meta[scol]):\n    print(\"TRAIN:\", len(train_index), \"TEST:\", len(test_index) ); \n    break\nprint(test_index)\nprint(df_meta[scol].value_counts().head(5)   )\nprint(df_meta.iloc[test_index,:][scol].value_counts().head(5)    )\n\n\nflagged_column_name = 'Playground'\ndf_meta[flagged_column_name] = 0 \ndf_meta.loc[df_meta.index[test_index],flagged_column_name]  = 1\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:24.652356Z","iopub.execute_input":"2022-11-03T12:43:24.653304Z","iopub.status.idle":"2022-11-03T12:43:24.924230Z","shell.execute_reply.started":"2022-11-03T12:43:24.653272Z","shell.execute_reply":"2022-11-03T12:43:24.923331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"code","source":"selected_target = 'CD31'\nn_features = 500\n\nX= df_cite.iloc[:70988,:n_features][df_meta['Playground']==1]\ny= df_cite_train_y[df_meta['Playground']==1][selected_target]\nprint('X.shape, y.shape', X.shape, y.shape )\n\n# Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n# Private like test - new DAY, and donor, \n# While public like - only new donor (days are the same as in train):\n# Step 1: \nmask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \nX_train = df_cite.iloc[:70988,:n_features][mask_train]\ny_train = df_cite_train_y[mask_train][ selected_target ]\n# Step 2:\nmask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\nX_test_private_like = df_cite.iloc[:70988,:n_features][ mask_test_private_like  ]\ny_test_private_like = df_cite_train_y[mask_test_private_like][ selected_target ]\nX_test = X_test_private_like\ny_test = y_test_private_like\n# Step 3: \nmask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \nX_test_public_like = df_cite.iloc[:70988,:n_features][mask_test_public_like ]\ny_test_public_like = df_cite_train_y[mask_test_public_like][ selected_target ]\nX_test2 = X_test_public_like\ny_test2 = y_test_public_like\n\nmask_out_of_playground = (df_meta['Playground']==0)\nX_oop = df_cite.iloc[:70988,:n_features][mask_out_of_playground]\ny_oop = df_cite_train_y[mask_out_of_playground][ selected_target ]\nprint('X_oop, y_oop', X_oop.shape, y_oop.shape )\n\n\nX.shape,y.shape, X_train.shape, X_test_private_like.shape, X_test_public_like.shape, y_train.shape, y_test_private_like.shape, y_test_public_like.shape\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:24.925427Z","iopub.execute_input":"2022-11-03T12:43:24.926494Z","iopub.status.idle":"2022-11-03T12:43:25.298896Z","shell.execute_reply.started":"2022-11-03T12:43:24.926460Z","shell.execute_reply":"2022-11-03T12:43:25.297880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling Preparations","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:25.301611Z","iopub.execute_input":"2022-11-03T12:43:25.302947Z","iopub.status.idle":"2022-11-03T12:43:25.307580Z","shell.execute_reply.started":"2022-11-03T12:43:25.302913Z","shell.execute_reply":"2022-11-03T12:43:25.306618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat = pd.DataFrame()# columns = [ 'r2_score','mse','Time',  'n_feat', 'Target' ])\n#IXmdl = 0\ndf_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:25.310723Z","iopub.execute_input":"2022-11-03T12:43:25.311471Z","iopub.status.idle":"2022-11-03T12:43:25.322186Z","shell.execute_reply.started":"2022-11-03T12:43:25.311441Z","shell.execute_reply":"2022-11-03T12:43:25.321209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict_save_predictions = {}\ndef update_models_stat(model, model_ID, t0 ):\n    \n    IXmdl = model_ID # df_models_stat.shape[0] + 1\n    #df_models_stat.loc[IXmdl,'Model'] = model_ID\n    y_pred_loc = model.predict(X_test)\n    df_models_stat.loc[IXmdl,'r2_score'] = r2_score(y_test, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse'] =  mean_squared_error(y_test, y_pred_loc)\n    df_models_stat.loc[IXmdl,'Time'] = np.round( time.time() - t0,3)\n    df_models_stat.loc[IXmdl,'Target'] = selected_target\n    df_models_stat.loc[IXmdl,'n_feat'] = X_train.shape[1]\n    df_models_stat.loc[IXmdl,'n_samples_train'] = X_train.shape[0]\n    \n    y_pred_loc = model.predict(X_test_public_like)\n    df_models_stat.loc[IXmdl,'r2_score Test2 PublLike'] = r2_score(y_test_public_like, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Test2 PublLike'] =  mean_squared_error(y_test_public_like, y_pred_loc)\n    \n    y_pred_loc = model.predict(X_oop)\n    df_models_stat.loc[IXmdl,'r2_score Out of Playgr'] = r2_score(y_oop, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Out of Playgr'] =  mean_squared_error(y_oop, y_pred_loc)\n\n    y_pred_loc = model.predict(X_train)\n    df_models_stat.loc[IXmdl,'r2_score Train'] = r2_score(y_train, y_pred_loc) \n    df_models_stat.loc[IXmdl,'mse Train'] =  mean_squared_error(y_train, y_pred_loc)\n    \n    dict_save_predictions[model_ID] = (model.predict(X_train), model.predict(X_test), \n                                       model.predict(X_test_public_like), model.predict(X_oop)   )\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:25.323247Z","iopub.execute_input":"2022-11-03T12:43:25.323527Z","iopub.status.idle":"2022-11-03T12:43:25.334144Z","shell.execute_reply.started":"2022-11-03T12:43:25.323500Z","shell.execute_reply":"2022-11-03T12:43:25.333365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LightGBM model","metadata":{}},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm\nprint('Default Params')\nt0 = time.time()\nmodel = lgbm.LGBMRegressor(random_state = 0, boosting='dart')\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\ntm0 = time.time() - t0\nupdate_models_stat(model, 'Lgbm Default', t0 )\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred), 'Secs: %.2f'%(time.time() - t0)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:25.335409Z","iopub.execute_input":"2022-11-03T12:43:25.335670Z","iopub.status.idle":"2022-11-03T12:43:35.116783Z","shell.execute_reply.started":"2022-11-03T12:43:25.335646Z","shell.execute_reply":"2022-11-03T12:43:35.115838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:35.120888Z","iopub.execute_input":"2022-11-03T12:43:35.122077Z","iopub.status.idle":"2022-11-03T12:43:35.138157Z","shell.execute_reply.started":"2022-11-03T12:43:35.122041Z","shell.execute_reply":"2022-11-03T12:43:35.137142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape, y_train.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:35.141573Z","iopub.execute_input":"2022-11-03T12:43:35.141928Z","iopub.status.idle":"2022-11-03T12:43:35.151140Z","shell.execute_reply.started":"2022-11-03T12:43:35.141898Z","shell.execute_reply":"2022-11-03T12:43:35.150370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndef objective(trial):\n    params = {\n        'n_estimators' :  trial.suggest_int('n_estimators', 100, 1000) ,\n        'num_leaves': trial.suggest_int('num_leaves', 1, 1000),\n        'max_depth': trial.suggest_int('max_depth', 50, 500),\n        'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n        'min_split_gain': trial.suggest_int(\"min_split_gain\", 1, 100),\n        'subsample_for_bin': trial.suggest_int(\"subsample_for_bin\", 1, 100000, log=True),\n        'colsample_bytree': trial.suggest_float(\"colsample_bytree\", 0.0, 1.0),\n        'reg_alpha': trial.suggest_float(\"reg_alpha\", 0.0, 10),\n        'reg_lambda': trial.suggest_float(\"reg_lambda\", 0.0, 10)\n    }\n    \n    model = lgbm.LGBMRegressor(**params, boosting='dart',random_state = 0)#\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    mse = mean_squared_error(y_test, y_pred)\n    \n    return mse","metadata":{"execution":{"iopub.status.busy":"2022-11-03T12:43:35.152496Z","iopub.execute_input":"2022-11-03T12:43:35.153059Z","iopub.status.idle":"2022-11-03T12:43:35.282246Z","shell.execute_reply.started":"2022-11-03T12:43:35.153031Z","shell.execute_reply":"2022-11-03T12:43:35.281190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"find_params = True\nif find_params:\n#     study = optuna.create_study(\n#         direction='minimize', \n#         pruner=optuna.pruners.MedianPruner(n_warmup_steps=20),\n#         study_name='small')\n#     study.optimize(objective, n_trials=20)\n    study = optuna.create_study()\n    study.optimize(objective, n_trials=10)\n    \n    # Output for best found params: \n  # E.g. {'x': 2.002108042}\n    t0  = time.time()\n    model = lgbm.LGBMRegressor(**study.best_params, boosting='dart')#\n    model.fit(X_train,y_train)\n    y_pred = model.predict(X_test)\n    #r2_score(y_test, y_pred), \n    \n    update_models_stat(model, 'LGBM Tuned', t0 ) # Updates df_models_stat\n\n","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-11-03T13:01:29.076139Z","iopub.execute_input":"2022-11-03T13:01:29.076496Z","iopub.status.idle":"2022-11-03T13:07:12.342708Z","shell.execute_reply.started":"2022-11-03T13:01:29.076457Z","shell.execute_reply":"2022-11-03T13:07:12.341002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Best params:')\nprint(study.best_params)\nprint('Best r2:%.6f'%r2_score(y_test, y_pred),'Mse:%.3f'%mean_squared_error(y_test, y_pred) )\n\ndf_models_stat","metadata":{"execution":{"iopub.status.busy":"2022-11-03T13:07:20.274631Z","iopub.execute_input":"2022-11-03T13:07:20.275018Z","iopub.status.idle":"2022-11-03T13:07:20.294812Z","shell.execute_reply.started":"2022-11-03T13:07:20.274983Z","shell.execute_reply":"2022-11-03T13:07:20.293859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"losses=[x.values for x in study.trials]\ntimes=[(x.datetime_complete-x.datetime_start).total_seconds() for x in study.trials]\nplt.figure(figsize=(16,6))\nplt.subplot(1,2,1)\nplt.title('Loss')\nplt.xlabel('epoch')\nplt.ylabel('mse_loss')\nplt.plot(losses)\nplt.subplot(1,2,2)\nplt.title('Time')\nplt.xlabel('epoch')\nplt.ylabel('seconds')\nplt.plot(times)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T13:08:00.760635Z","iopub.execute_input":"2022-11-03T13:08:00.761056Z","iopub.status.idle":"2022-11-03T13:08:01.025549Z","shell.execute_reply.started":"2022-11-03T13:08:00.761020Z","shell.execute_reply":"2022-11-03T13:08:01.024539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for key in study.best_params.keys():\n    print(key)\n    value=[x.params[key] for x in study.trials]\n    plt.figure(figsize=(16,6))\n    plt.subplot(1,2,1)\n    plt.title('Loss/'+key)\n    plt.xlabel('mse_loss')\n    plt.ylabel(key)\n    plt.scatter(losses, value)\n    plt.subplot(1,2,2)\n    plt.title('Time/'+key)\n    plt.xlabel('seconds')\n    plt.ylabel(key)\n    plt.scatter(times, value)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T13:08:08.028756Z","iopub.execute_input":"2022-11-03T13:08:08.029893Z","iopub.status.idle":"2022-11-03T13:08:11.316112Z","shell.execute_reply.started":"2022-11-03T13:08:08.029852Z","shell.execute_reply":"2022-11-03T13:08:11.315088Z"},"trusted":true},"execution_count":null,"outputs":[]}]}