{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n\n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-01T16:33:24.798885Z","iopub.execute_input":"2022-11-01T16:33:24.799371Z","iopub.status.idle":"2022-11-01T16:33:24.885389Z","shell.execute_reply.started":"2022-11-01T16:33:24.799265Z","shell.execute_reply":"2022-11-01T16:33:24.884160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:33:25.378326Z","iopub.execute_input":"2022-11-01T16:33:25.378761Z","iopub.status.idle":"2022-11-01T16:33:54.193911Z","shell.execute_reply.started":"2022-11-01T16:33:25.378721Z","shell.execute_reply":"2022-11-01T16:33:54.192555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:33:54.196118Z","iopub.execute_input":"2022-11-01T16:33:54.196489Z","iopub.status.idle":"2022-11-01T16:34:16.741579Z","shell.execute_reply.started":"2022-11-01T16:33:54.196455Z","shell.execute_reply":"2022-11-01T16:34:16.740213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:16.743601Z","iopub.execute_input":"2022-11-01T16:34:16.744040Z","iopub.status.idle":"2022-11-01T16:34:37.146748Z","shell.execute_reply.started":"2022-11-01T16:34:16.743999Z","shell.execute_reply":"2022-11-01T16:34:37.145605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.mean(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:37.149167Z","iopub.execute_input":"2022-11-01T16:34:37.149540Z","iopub.status.idle":"2022-11-01T16:34:37.338605Z","shell.execute_reply.started":"2022-11-01T16:34:37.149505Z","shell.execute_reply":"2022-11-01T16:34:37.337298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.std(axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:37.340911Z","iopub.execute_input":"2022-11-01T16:34:37.341419Z","iopub.status.idle":"2022-11-01T16:34:38.162553Z","shell.execute_reply.started":"2022-11-01T16:34:37.341357Z","shell.execute_reply":"2022-11-01T16:34:38.161290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:38.164262Z","iopub.execute_input":"2022-11-01T16:34:38.164875Z","iopub.status.idle":"2022-11-01T16:34:38.986422Z","shell.execute_reply.started":"2022-11-01T16:34:38.164841Z","shell.execute_reply":"2022-11-01T16:34:38.985155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:38.987865Z","iopub.execute_input":"2022-11-01T16:34:38.988206Z","iopub.status.idle":"2022-11-01T16:34:39.385101Z","shell.execute_reply.started":"2022-11-01T16:34:38.988174Z","shell.execute_reply":"2022-11-01T16:34:39.383866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\"  (i.e. downsample)\n\nSmall 10% of data of data where we can make prelimanary experiments. \n\nit will be labeled by special column \"Playground\" in df_meta (meta data for CITE-seq train only)","metadata":{}},{"cell_type":"code","source":"# Prepare for creation of additional holdout folds with 10% of samples \n# We will use stratified Kfold to achieve that days, cell_types and donors are equally distributed \nimport numpy as np\nfrom sklearn.model_selection import StratifiedKFold\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n\nskf = StratifiedKFold(n_splits=10,  shuffle=True, random_state=40)\nskf.get_n_splits(df_meta, df_meta[scol] )\n\n\n\ny = df_meta[scol] \nfor train_index, test_index in skf.split(df_meta, df_meta[scol]):\n    print(\"TRAIN:\", len(train_index), \"TEST:\", len(test_index) ); \n    break\nprint(test_index)\nprint(df_meta[scol].value_counts().head(5)   )\nprint(df_meta.iloc[test_index,:][scol].value_counts().head(5)    )\n\n\nflagged_column_name = 'Playground'\ndf_meta[flagged_column_name] = 0 \ndf_meta.loc[df_meta.index[test_index],flagged_column_name]  = 1\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:39.386908Z","iopub.execute_input":"2022-11-01T16:34:39.387572Z","iopub.status.idle":"2022-11-01T16:34:39.751062Z","shell.execute_reply.started":"2022-11-01T16:34:39.387526Z","shell.execute_reply":"2022-11-01T16:34:39.749911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"code","source":"selected_target = 'CD31'\n\nX= df_cite.iloc[:70988,:][df_meta['Playground']==1]\ny= df_cite_train_y[df_meta['Playground']==1][selected_target]\nprint('X.shape, y.shape', X.shape, y.shape )\n\n# Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n# Private like test - new DAY, and donor, \n# While public like - only new donor (days are the same as in train):\n# Step 1: \nmask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \nX_train = df_cite.iloc[:70988,:][mask_train]\ny_train = df_cite_train_y[mask_train][ selected_target ]\n# Step 2:\nmask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\nX_test_private_like = df_cite.iloc[:70988,:][ mask_test_private_like  ]\ny_test_private_like = df_cite_train_y[mask_test_private_like][ selected_target ]\nX_test = X_test_private_like\ny_test = y_test_private_like\n# Step 3: \nmask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \nX_test_public_like = df_cite.iloc[:70988,:][mask_test_public_like ]\ny_test_public_like = df_cite_train_y[mask_test_public_like][ selected_target ]\n\nX.shape,y.shape, X_train.shape, X_test_private_like.shape, X_test_public_like.shape, y_train.shape, y_test_private_like.shape, y_test_public_like.shape\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:39.753027Z","iopub.execute_input":"2022-11-01T16:34:39.753497Z","iopub.status.idle":"2022-11-01T16:34:39.895305Z","shell.execute_reply.started":"2022-11-01T16:34:39.753451Z","shell.execute_reply":"2022-11-01T16:34:39.894109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:39.899360Z","iopub.execute_input":"2022-11-01T16:34:39.899743Z","iopub.status.idle":"2022-11-01T16:34:39.905376Z","shell.execute_reply.started":"2022-11-01T16:34:39.899710Z","shell.execute_reply":"2022-11-01T16:34:39.903620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LightGBM model","metadata":{}},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:34:39.907087Z","iopub.execute_input":"2022-11-01T16:34:39.907484Z","iopub.status.idle":"2022-11-01T16:34:40.908927Z","shell.execute_reply.started":"2022-11-01T16:34:39.907448Z","shell.execute_reply":"2022-11-01T16:34:40.907723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nparams = {\n    'task' : 'predict',\n    'application' : 'regression',\n#     'objective' : 'root_mean_squared_error',\n#     'boosting_type': \"gbdt\",\n# #     'num_iterations' : 2500,\n#     'learning_rate' : 0.05,\n#     'num_leaves' : 15,\n#     'tree_learner' : 'feature',\n#     'max_depth' : 10,\n#     'min_data_in_leaf' : 7,\n#     'bagging_fraction' : 1,\n#     'bagging_freq' : 100,\n#     'reg_sqrt' : 'True',\n#     'metric' : 'rmse',\n#     'feature_fraction' : 0.6,\n    'random_state' : 42\n}\n\nlgb_model = lgbm.LGBMRegressor()\n\n\nparameters = {\n    'task' : ['predict'],\n#     'boosting': ['gbdt' ],\n    'objective': ['root_mean_squared_error'],\n#     'num_iterations': [  1500, 2000,5000  ],\n    'learning_rate':[  0.05, 0.005 ],\n   'num_leaves':[ 7, 15, 31  ],\n   'max_depth' :[ 10,15,25],\n    'random_state' : [42],\n#    'min_data_in_leaf':[15,25 ],\n#   'feature_fraction': [ 0.6, 0.8,  0.9],\n#     'bagging_fraction': [  0.6, 0.8 ],\n#     'bagging_freq': [   100, 200, 400  ],\n     \n}","metadata":{"execution":{"iopub.status.busy":"2022-10-31T15:28:30.704847Z","iopub.status.idle":"2022-10-31T15:28:30.705431Z","shell.execute_reply.started":"2022-10-31T15:28:30.705114Z","shell.execute_reply":"2022-10-31T15:28:30.705141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\ngsearch_lgb = (lgb_model, param_grid = parameters, n_jobs=1, verbose=1)\ngsearch_lgb.fit(X_train,y_train)\n\nprint (gsearch_lgb.best_params_)\ny_pred = gsearch_lgb.predict(X_test)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T15:28:30.706835Z","iopub.status.idle":"2022-10-31T15:28:30.707451Z","shell.execute_reply.started":"2022-10-31T15:28:30.707100Z","shell.execute_reply":"2022-10-31T15:28:30.707126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\nscaler.fit(X_train)\nX_train_scaled = scaler.transform(X_train)\nX_test_scaled = scaler.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T15:28:30.709789Z","iopub.status.idle":"2022-10-31T15:28:30.710400Z","shell.execute_reply.started":"2022-10-31T15:28:30.710073Z","shell.execute_reply":"2022-10-31T15:28:30.710102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\nparameters = {\n# 'task' : ['predict'],\n# 'objective': ['root_mean_squared_error'],\n'learning_rate':[  0.05, 0.005 ],\n'num_leaves':[ 3, 7, 15, 31],\n'max_depth' :[ 10,15,25],\n}\nrand_search = RandomizedSearchCV(\n    estimator = lgb_model, param_distributions = parameters, \n    random_state = 42,\n    verbose = 2)\n\nrand_search.fit(X_train,y_train)\nprint (rand_search.best_params_)\ny_pred = rand_search.predict(X_test)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-10-31T15:28:30.712755Z","iopub.status.idle":"2022-10-31T15:28:30.713955Z","shell.execute_reply.started":"2022-10-31T15:28:30.713633Z","shell.execute_reply":"2022-10-31T15:28:30.713665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n\nlgb_model = lgbm.LGBMRegressor()\n\nfrom sklearn.experimental import enable_halving_search_cv\nfrom sklearn.model_selection import HalvingGridSearchCV\nparameters = {\n# 'task' : ['predict'],\n# 'objective': ['root_mean_squared_error'],\n'learning_rate':[0.01, 0.03, 0.04, 0.05, 0.075, 0.1],\n'boosting_type': ['gbdt', 'dart'],\n'num_leaves':[ 3, 7, 15, 20, 31 ],\n'max_depth' :[ 10,15,20, 25],\n'num_iterations' : [2500],\n}\n\nMAX_RESOURCE_DIVISOR = 20\nFACTOR = 3\nn_samples = len(X_train_scaled)\n                        \n\nhv_grid_search = HalvingGridSearchCV(\n    min_resources=n_samples // MAX_RESOURCE_DIVISOR,\n    estimator = lgb_model,\n    param_grid = parameters, \n    factor=FACTOR,\n    random_state = 42,\n    verbose = 1)\n\nhv_grid_search.fit(X_train_scaled,y_train)\nprint (hv_grid_search.best_params_)\ny_pred = hv_grid_search.predict(X_test_scaled)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\n","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-10-31T15:28:30.716237Z","iopub.status.idle":"2022-10-31T15:28:30.716836Z","shell.execute_reply.started":"2022-10-31T15:28:30.716538Z","shell.execute_reply":"2022-10-31T15:28:30.716567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"r2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T15:28:30.718404Z","iopub.status.idle":"2022-10-31T15:28:30.718981Z","shell.execute_reply.started":"2022-10-31T15:28:30.718676Z","shell.execute_reply":"2022-10-31T15:28:30.718705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlgb_model = lgbm.LGBMRegressor(**hv_grid_search.best_params_, num_leaves=31)\nlgb_model.fit(X_train,y_train)\ny_pred = lgb_model.predict(X_test)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T15:28:30.720715Z","iopub.status.idle":"2022-10-31T15:28:30.721317Z","shell.execute_reply.started":"2022-10-31T15:28:30.720991Z","shell.execute_reply":"2022-10-31T15:28:30.721018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.experimental import enable_halving_search_cv\nfrom sklearn.model_selection import HalvingRandomSearchCV\nparameters = {\n# 'task' : ['predict'],\n# 'objective': ['root_mean_squared_error'],\n'learning_rate':[0.005, 0.01, 0.03, 0.04, 0.05, 0.1, 0.5, 1],\n'num_leaves':[ 3, 7, 15, 20, 31 ],\n'max_depth' :[ 10,15,25],\n}\n\nhv_random_search = HalvingRandomSearchCV(\n    estimator = lgb_model, param_distributions = parameters, \n    random_state = 42,\n    verbose = 1)\n\nhv_random_search.fit(X_train,y_train)\nprint (hv_grid_search.best_params_)\ny_pred = hv_random_search.predict(X_test)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-10-31T15:28:30.723267Z","iopub.status.idle":"2022-10-31T15:28:30.723854Z","shell.execute_reply.started":"2022-10-31T15:28:30.723560Z","shell.execute_reply":"2022-10-31T15:28:30.723586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna\nfrom sklearn.multioutput import MultiOutputRegressor\n\nsmall_train_x, small_train_y =  X_train.iloc[:1000],y_train.iloc[:1000]\ndef objective(trial):\n#     params0 = {\n#         'metric': 'mae', \n#         'random_state': 42,\n#         'n_estimators': 2000,\n#         'reg_alpha': trial.suggest_float('reg_alpha', 1e-3, 10.0),\n#         'reg_lambda': trial.suggest_float('reg_lambda', 1e-3, 10.0),\n#         'colsample_bytree': trial.suggest_categorical('colsample_bytree', [0.3,0.4,0.5,0.6,0.7,0.8,0.9, 1.0]),\n#         'subsample': trial.suggest_categorical('subsample', [0.4,0.5,0.6,0.7,0.8,1.0]),\n#         'learning_rate': trial.suggest_float('learning_rate', 0.01, 0.1),\n#         'max_depth': trial.suggest_categorical('max_depth', [10,20,100]),\n#         'num_leaves' : trial.suggest_int('num_leaves', 1, 1000),\n#         'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n#         'cat_smooth' : trial.suggest_int('min_data_per_groups', 1, 100)\n#     }\n    params = {\n        'learning_rate': trial.suggest_float('learning_rate', 0.01, 0.1),\n        'boosting_type': trial.suggest_categorical('boosting_type', ['gbdt', 'dart']),\n        'num_leaves': trial.suggest_categorical('num_leaves', [ 3, 7, 15, 20, 31 ]),\n        'max_depth' : trial.suggest_categorical('max_depth', [ 10, 20, 31]),\n    }\n\n    model = lgbm.LGBMRegressor(**params)\n\n    model.fit(small_train_x, small_train_y)\n\n    y_va_pred = model.predict(X_test)\n    mse = mean_squared_error(y_test, y_va_pred)\n    \n    return mse\n\nfind_params = True\nif find_params:\n    study = optuna.create_study(\n        direction='minimize', \n        pruner=optuna.pruners.MedianPruner(n_warmup_steps=20),\n        study_name='small')\n    study.optimize(objective, n_trials=20)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-10-29T18:09:32.971333Z","iopub.execute_input":"2022-10-29T18:09:32.972647Z","iopub.status.idle":"2022-10-29T18:10:08.843104Z","shell.execute_reply.started":"2022-10-29T18:09:32.972545Z","shell.execute_reply":"2022-10-29T18:10:08.842091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params = {'reg_alpha': 0.013909414872153632, 'reg_lambda': 0.005755407865365625, 'colsample_bytree': 0.7, 'subsample': 0.4, 'learning_rate': 0.010454682205474793, 'max_depth': 20, 'num_leaves': 345, 'min_child_samples': 3, 'min_data_per_groups': 40}\nmodel = lgbm.LGBMRegressor(**best_params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\n\n# model.fit(X_train,y_train)\n# y_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T09:49:05.417773Z","iopub.execute_input":"2022-10-30T09:49:05.418197Z","iopub.status.idle":"2022-10-30T09:49:40.747934Z","shell.execute_reply.started":"2022-10-30T09:49:05.418164Z","shell.execute_reply":"2022-10-30T09:49:40.747001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna.integration.lightgbm as lgb\nfrom sklearn.model_selection import train_test_split\n\ntrain_x, val_x, train_y, val_y = train_test_split(X_train, y_train, test_size=0.25)\ndtrain = lgb.Dataset(train_x, label=train_y)\ndval = lgb.Dataset(val_x, label=val_y)\n\nparams = {\n\"objective\": \"root_mean_squared_error\",\n\"metric\": \"rmse\",\n\"verbosity\": -1,\n\"boosting_type\": \"gbdt\",\n}\n\ntuner = lgb.LightGBMTuner(\n     params, dtrain, verbose_eval=100, early_stopping_rounds=1000, valid_sets=[dtrain, dval],\n  model_dir= 'directory_to_save_boosters'\n)\n\ntuner.run()","metadata":{"execution":{"iopub.status.busy":"2022-10-29T18:18:43.984634Z","iopub.execute_input":"2022-10-29T18:18:43.985011Z","iopub.status.idle":"2022-10-29T18:28:44.084115Z","shell.execute_reply.started":"2022-10-29T18:18:43.984981Z","shell.execute_reply":"2022-10-29T18:28:44.082484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:53:07.827987Z","iopub.status.idle":"2022-10-29T14:53:07.828863Z","shell.execute_reply.started":"2022-10-29T14:53:07.828641Z","shell.execute_reply":"2022-10-29T14:53:07.828663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Optimizing hyperparameters","metadata":{}},{"cell_type":"code","source":"import optuna\nfrom sklearn.model_selection import train_test_split\ndef objective(trial,data=X_train,target=y_train):    \n    train_x, test_x, train_y, test_y = train_test_split(data, target, test_size=0.2,random_state=42)\n# def objective(trial,data=X_train,target=y_train, test_x = X_test, test_y=y_test):\n#     train_x, test_x, train_y, test_y = data, test_x, target, test_y \n#     param = {\n#         'metric': 'rmse', \n#         'random_state': 48,\n#         'n_estimators': 20000,\n#         'reg_alpha': trial.suggest_float('reg_alpha', 1e-3, 10.0),\n#         'reg_lambda': trial.suggest_float('reg_lambda', 1e-3, 10.0),\n#         'colsample_bytree': trial.suggest_categorical('colsample_bytree', [0.3,0.4,0.5,0.6,0.7,0.8,0.9, 1.0]),\n#         'subsample': trial.suggest_categorical('subsample', [0.4,0.5,0.6,0.7,0.8,1.0]),\n#         'learning_rate': trial.suggest_categorical('learning_rate', [0.006,0.008,0.01,0.014,0.017,0.02]),\n#         'max_depth': trial.suggest_categorical('max_depth', [10,20,100]),\n#         'num_leaves' : trial.suggest_int('num_leaves', 1, 1000),\n#         'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n#         'cat_smooth' : trial.suggest_int('min_data_per_groups', 1, 100)\n#         'boosting_type': trial.suggest_categorical('boosting_type', ['gbdt', 'dart']),\n#         'num_leaves':[ 3, 7, 15, 20, 31 ],\n#         'max_depth' :[ 10,15,20, 25],\n#         'num_iterations' : trial.suggest_categorical('num_iterations', [100]),\n#     }\n    param = {\n        'boosting_type': trial.suggest_categorical('boosting_type', ['gbdt', 'dart']),\n        'learning_rate': trial.suggest_float('learning_rate', 0.001, 0.2),\n#         'num_leaves': trial.suggest_categorical('num_leaves', [ 3, 7, 12, 20, 31 ]),!\n        'num_leaves': trial.suggest_categorical('num_leaves', [12]),\n        'n_estimators': trial.suggest_categorical('n_estimators', [150]),\n#         'max_depth' : trial.suggest_categorical('max_depth', [ 10, 20, 31, 50, 100]),!\n        'max_depth' : trial.suggest_categorical('max_depth', [50]),\n#        'subsample': trial.suggest_float('subsample', 0.3,0.5),!\n        'subsample': trial.suggest_categorical('subsample',[0.457]),\n        'min_split_gain': trial.suggest_float('min_split_gain', 0,10),\n        'bagging_fraction' : trial.suggest_float('bagging_fraction', 0,1), \n        'feature_fraction' : trial.suggest_float('feature_fraction', 0,1), \n        'random_state': trial.suggest_categorical('random_state', [0]),\n        'extra_trees' : trial.suggest_categorical('extra_trees', [True, False])\n#         'colsample_bytree': trial.suggest_categorical('colsample_bytree', [0.3,0.4,0.5,0.6,0.7,0.8,0.9, 1.0]),\n#         'min_child_samples': trial.suggest_int('min_child_samples', 1, 300),\n    }\n    model = lgbm.LGBMRegressor(**param) \n#     model.fit(train_x,train_y,eval_set=[(test_x,test_y)], callbacks=[lgbm.early_stopping(50)])\n    model.fit(train_x,train_y)\n    \n    preds = model.predict(test_x)\n    \n    rmse = mean_squared_error(test_y, preds,squared=False)\n    r2 = r2_score(test_y, preds)\n    \n#     return rmse\n    return r2\n\n# study = optuna.create_study(direction='minimize')\nstudy = optuna.create_study(direction='maximize')\nstudy.optimize(objective, n_trials=750)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T16:49:13.132305Z","iopub.execute_input":"2022-10-30T16:49:13.132756Z","iopub.status.idle":"2022-10-30T16:49:38.307458Z","shell.execute_reply.started":"2022-10-30T16:49:13.132692Z","shell.execute_reply":"2022-10-30T16:49:38.304766Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\n\n#https://www.kaggle.com/code/hamzaghanmi/lgbm-hyperparameter-tuning-using-optuna\noptuna.visualization.plot_parallel_coordinate(study)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T15:56:42.506157Z","iopub.execute_input":"2022-10-30T15:56:42.507189Z","iopub.status.idle":"2022-10-30T15:56:43.091048Z","shell.execute_reply.started":"2022-10-30T15:56:42.507139Z","shell.execute_reply":"2022-10-30T15:56:43.090234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"optuna.visualization.plot_param_importances(study)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T15:56:47.834358Z","iopub.execute_input":"2022-10-30T15:56:47.834762Z","iopub.status.idle":"2022-10-30T15:58:05.432105Z","shell.execute_reply.started":"2022-10-30T15:56:47.834702Z","shell.execute_reply":"2022-10-30T15:58:05.430791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#(0.4515336738095388, 22.34325502629381)\nparams = {'boosting_type': 'gbdt',\n          'learning_rate': 0.06574490834089135,\n          'num_leaves': 12,\n          'n_estimators': 150,\n          'max_depth': 50,\n          'subsample': 0.457,\n          'min_split_gain': 8.664344164634976,\n          'bagging_fraction': 0.47599085037377187,\n          'feature_fraction': 0.9537437332112573,\n          'random_state': 0,\n          'extra_trees': True}\n\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T15:58:58.774392Z","iopub.execute_input":"2022-10-30T15:58:58.774818Z","iopub.status.idle":"2022-10-30T15:58:59.997677Z","shell.execute_reply.started":"2022-10-30T15:58:58.774782Z","shell.execute_reply":"2022-10-30T15:58:59.996551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#(0.4392921253842067, 22.84194751719605) - warning, test was used to get these parameters!!\nparams = {'boosting_type': 'dart',\n          'learning_rate': 0.19292388778174555,\n          'num_leaves': 12,\n          'n_estimators': 150,\n          'max_depth': 50,\n          'subsample': 0.457,\n          'min_split_gain': 3.7874444939895326,\n          'bagging_fraction': 0.05206157146226077, \n          'feature_fraction': 0.8622272500714694,\n          'random_state': 0\n         }\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T15:21:09.530620Z","iopub.execute_input":"2022-10-30T15:21:09.531815Z","iopub.status.idle":"2022-10-30T15:21:12.402796Z","shell.execute_reply.started":"2022-10-30T15:21:09.531771Z","shell.execute_reply":"2022-10-30T15:21:12.401593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {'boosting_type': 'gbdt', 'learning_rate': 0.11416087231737981, 'num_leaves': 12, 'n_estimators': 150, 'max_depth': 50, 'subsample': 0.457, 'min_split_gain': 8.62497144495012, 'bagging_fraction': 0.6437210517259985, 'feature_fraction': 0.467623452206488, 'random_state': 0}\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T12:24:32.749758Z","iopub.execute_input":"2022-10-30T12:24:32.750328Z","iopub.status.idle":"2022-10-30T12:24:34.132740Z","shell.execute_reply.started":"2022-10-30T12:24:32.750280Z","shell.execute_reply":"2022-10-30T12:24:34.131684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#(0.4035033699064735, 24.29989899484972)\nparams =  {'boosting_type': 'gbdt', 'learning_rate': 0.09447661797639521, 'num_leaves': 12, 'n_estimators': 150, 'max_depth': 50, 'subsample': 0.457, 'min_split_gain': 7.280626081286166, 'bagging_fraction': 0.8994893813409378, 'feature_fraction': 0.4081022471202951, 'random_state': 0}\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:20:06.173979Z","iopub.execute_input":"2022-10-30T11:20:06.174400Z","iopub.status.idle":"2022-10-30T11:20:07.448638Z","shell.execute_reply.started":"2022-10-30T11:20:06.174367Z","shell.execute_reply":"2022-10-30T11:20:07.447606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# (0.42386386023586475, 23.470459508473667)\nparams = {'boosting_type': 'gbdt', 'learning_rate': 0.09077072260231679, 'num_leaves': 12, 'max_depth': 50, 'subsample': 0.4570609721468997, 'min_split_gain': 5.588206623830203, 'random_state' : 0, 'n_estimators':150}\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T10:53:45.615804Z","iopub.execute_input":"2022-10-30T10:53:45.617008Z","iopub.status.idle":"2022-10-30T10:53:48.525235Z","shell.execute_reply.started":"2022-10-30T10:53:45.616955Z","shell.execute_reply":"2022-10-30T10:53:48.524084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#(0.41817803960135835, 23.702086746836564)\nparams = {'boosting_type': 'gbdt', 'learning_rate': 0.0954386162837535, 'num_leaves': 12, 'max_depth': 31, 'subsample': 0.33610547409731384, 'random_state' : 0, 'n_estimators':150}\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T10:12:15.033417Z","iopub.execute_input":"2022-10-30T10:12:15.034286Z","iopub.status.idle":"2022-10-30T10:12:17.498168Z","shell.execute_reply.started":"2022-10-30T10:12:15.034240Z","shell.execute_reply":"2022-10-30T10:12:17.496651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#(0.4354558552914086, 22.998228325953246)\nparams = {'learning_rate': 0.067,\n          'num_leaves': 12,\n          'max_depth' : 9,\n          'n_estimators':150, \n          'subsample': 0.4,\n          'random_state' : 0}\nmodel = lgbm.LGBMRegressor(**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T09:55:29.711132Z","iopub.execute_input":"2022-10-30T09:55:29.711596Z","iopub.status.idle":"2022-10-30T09:55:32.274269Z","shell.execute_reply.started":"2022-10-30T09:55:29.711562Z","shell.execute_reply":"2022-10-30T09:55:32.273299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_trial = {'learning_rate': 0.09034439781047379,\n              'boosting_type': 'gbdt',\n              'num_leaves': 15,\n              'max_depth': 31,\n              'colsample_bytree': 0.5,\n              'min_child_samples': 119}\nmodel = lgbm.LGBMRegressor(**best_trial)#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:10:01.912416Z","iopub.execute_input":"2022-10-29T20:10:01.912885Z","iopub.status.idle":"2022-10-29T20:10:03.118720Z","shell.execute_reply.started":"2022-10-29T20:10:01.912849Z","shell.execute_reply":"2022-10-29T20:10:03.117843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T19:54:17.399080Z","iopub.execute_input":"2022-10-29T19:54:17.399471Z","iopub.status.idle":"2022-10-29T19:54:17.414593Z","shell.execute_reply.started":"2022-10-29T19:54:17.399441Z","shell.execute_reply":"2022-10-29T19:54:17.413339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_trial = {'learning_rate': 0.06411795521317262, 'boosting_type': 'gbdt', 'num_leaves': 31, 'max_depth': 10, 'colsample_bytree': 0.4, 'min_child_samples': 99}\nmodel = lgbm.LGBMRegressor(**best_trial, num_iterations=100)#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T20:03:34.551921Z","iopub.execute_input":"2022-10-29T20:03:34.552297Z","iopub.status.idle":"2022-10-29T20:03:35.590813Z","shell.execute_reply.started":"2022-10-29T20:03:34.552266Z","shell.execute_reply":"2022-10-29T20:03:35.589873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_trial = {'learning_rate': 0.06292206417901433, 'boosting_type': 'gbdt', 'num_leaves': 20, 'max_depth': 31, 'colsample_bytree': 0.6, 'min_child_samples': 44}\nmodel = lgbm.LGBMRegressor(**study.best_trial.params)#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T19:54:36.771399Z","iopub.execute_input":"2022-10-29T19:54:36.771796Z","iopub.status.idle":"2022-10-29T19:54:38.457800Z","shell.execute_reply.started":"2022-10-29T19:54:36.771764Z","shell.execute_reply":"2022-10-29T19:54:38.456877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_trial = {'learning_rate': 0.08747963388041238, 'boosting_type': 'dart', 'num_leaves': 15, 'max_depth': 20, 'num_iterations': 2500}\nmodel = lgbm.LGBMRegressor(**study.best_trial.params)#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T09:45:33.379484Z","iopub.execute_input":"2022-10-30T09:45:33.379959Z","iopub.status.idle":"2022-10-30T09:45:33.477232Z","shell.execute_reply.started":"2022-10-30T09:45:33.379863Z","shell.execute_reply":"2022-10-30T09:45:33.475677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_trial_no_limit = {'learning_rate': 0.08747963388041238, 'boosting_type': 'dart', 'num_leaves': 15, 'max_depth': 20,  'num_iterations': 25000, 'early_stopping_rounds': 20}\nmodel = lgbm.LGBMRegressor(**best_trial_no_limit)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T19:46:54.156428Z","iopub.execute_input":"2022-10-29T19:46:54.157507Z","iopub.status.idle":"2022-10-29T19:47:10.541452Z","shell.execute_reply.started":"2022-10-29T19:46:54.157463Z","shell.execute_reply":"2022-10-29T19:47:10.540385Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Catboost","metadata":{}},{"cell_type":"code","source":"%%time\nfrom catboost import CatBoostRegressor\n# (0.419706323373084, 23.63982798557151)\nmodel = CatBoostRegressor(verbose=0)\nmodel.fit(X_train, y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-11-01T16:37:13.656444Z","iopub.execute_input":"2022-11-01T16:37:13.657752Z","iopub.status.idle":"2022-11-01T16:38:20.387144Z","shell.execute_reply.started":"2022-11-01T16:37:13.657705Z","shell.execute_reply":"2022-11-01T16:38:20.385922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport optuna\nfrom sklearn.model_selection import train_test_split\n\ndef objective(trial,data=X_train,target=y_train, test_x = X_test, test_y=y_test):\n    train_x, test_x, train_y, test_y = data, test_x, target, test_y \n    parameters = {\n        'learning_rate': trial.suggest_float('learning_rate', 0.05, 0.2), \n        'depth': trial.suggest_int('depth', 1, 7), \n        'l2_leaf_reg': 2, \n        'loss_function': trial.suggest_categorical('loss_function', ['RMSE']), \n#         'task_type': trial.suggest_categorical('task_type', ['GPU']), \n#         'iterations': trial.suggest_categorical('iterations', [100, 300, 800]),\n        'od_type': trial.suggest_categorical('od_type', ['Iter']), \n        'boosting_type': trial.suggest_categorical('boosting_type', ['Plain']), \n        'bootstrap_type': trial.suggest_categorical('bootstrap_type',['Bayesian']),\n        'bagging_temperature': trial.suggest_float('bagging_temperature', 0.2, 0.4),\n        'allow_const_label': trial.suggest_categorical('allow_const_label', [True, False]), \n        'random_state': trial.suggest_categorical('random_state', [0]),\n        'verbose': trial.suggest_categorical('verbose', [0])\n    }\n    print(parameters)\n    model = CatBoostRegressor(**parameters) \n    model.fit(train_x,train_y)\n    preds = model.predict(test_x)\n    \n    rmse = mean_squared_error(test_y, preds,squared=False)\n    r2 = r2_score(test_y, preds)\n#     return rmse\n    return r2\n\n# study = optuna.create_study(direction='minimize')\nstudy = optuna.create_study(direction='maximize')\nstudy.optimize(objective, n_trials=250)\nprint('Number of finished trials:', len(study.trials))\nprint('Best trial:', study.best_trial.params)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-11-01T16:44:06.177327Z","iopub.execute_input":"2022-11-01T16:44:06.177791Z","iopub.status.idle":"2022-11-01T18:27:39.855592Z","shell.execute_reply.started":"2022-11-01T16:44:06.177757Z","shell.execute_reply":"2022-11-01T18:27:39.854305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom catboost import CatBoostRegressor\n#(0.43755657703583684, 22.91264975998903)\nparams_best_1 = {'learning_rate': 0.05221440462531609, 'depth': 3, 'loss_function': 'RMSE', 'od_type': 'Iter', 'boosting_type': 'Plain', 'bootstrap_type': 'Bayesian', 'bagging_temperature': 0.3063450856699467, 'allow_const_label': False, 'random_state': 0, 'verbose': 0}\n#(0.44362093406060543, 22.665601820845637)\nparams_best_2 = {'learning_rate': 0.053934133317539365, 'depth': 3, 'l2_leaf_reg': 2, 'loss_function': 'RMSE', 'od_type': 'Iter', 'boosting_type': 'Plain', 'bootstrap_type': 'Bayesian', 'bagging_temperature': 0.2643632532675586, 'allow_const_label': False, 'random_state': 0, 'verbose': 0}\n\nparams_best_3 = {'learning_rate': 0.05029554257639053, 'depth': 3, 'loss_function': 'RMSE', 'od_type': 'Iter', 'boosting_type': 'Plain', 'bootstrap_type': 'Bayesian', 'bagging_temperature': 0.34681598154977983, 'allow_const_label': True, 'random_state': 0, 'verbose': 0}\n\nmodel = CatBoostRegressor(**params_best_3)\nmodel.fit(X_train, y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ridge","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.linear_model import Ridge\n","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:27:53.659251Z","iopub.execute_input":"2022-10-29T15:27:53.659835Z","iopub.status.idle":"2022-10-29T15:27:53.667128Z","shell.execute_reply.started":"2022-10-29T15:27:53.659785Z","shell.execute_reply":"2022-10-29T15:27:53.665539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:27:53.818102Z","iopub.execute_input":"2022-10-29T15:27:53.818527Z","iopub.status.idle":"2022-10-29T15:27:53.967220Z","shell.execute_reply.started":"2022-10-29T15:27:53.818495Z","shell.execute_reply":"2022-10-29T15:27:53.965698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1e4) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-29T15:27:57.847338Z","iopub.execute_input":"2022-10-29T15:27:57.847809Z","iopub.status.idle":"2022-10-29T15:27:57.950789Z","shell.execute_reply.started":"2022-10-29T15:27:57.847767Z","shell.execute_reply":"2022-10-29T15:27:57.948791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1e5) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-29T17:42:54.027834Z","iopub.execute_input":"2022-10-29T17:42:54.028585Z","iopub.status.idle":"2022-10-29T17:42:54.118985Z","shell.execute_reply.started":"2022-10-29T17:42:54.028551Z","shell.execute_reply":"2022-10-29T17:42:54.117416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Random Forest","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.ensemble import RandomForestRegressor\nmodel = RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:53:08.706561Z","iopub.execute_input":"2022-10-29T14:53:08.707300Z","iopub.status.idle":"2022-10-29T14:53:08.866155Z","shell.execute_reply.started":"2022-10-29T14:53:08.707261Z","shell.execute_reply":"2022-10-29T14:53:08.864840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SVR (support vector machine regression)","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.svm import SVR\nmodel = SVR(C=1.0, epsilon=0.2) # RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:53:09.022397Z","iopub.execute_input":"2022-10-29T14:53:09.022854Z","iopub.status.idle":"2022-10-29T14:53:09.041819Z","shell.execute_reply.started":"2022-10-29T14:53:09.022815Z","shell.execute_reply":"2022-10-29T14:53:09.040813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KernelRidge","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.kernel_ridge import KernelRidge\nmodel = KernelRidge(alpha=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:53:09.343783Z","iopub.execute_input":"2022-10-29T14:53:09.344234Z","iopub.status.idle":"2022-10-29T14:53:09.367195Z","shell.execute_reply.started":"2022-10-29T14:53:09.344186Z","shell.execute_reply":"2022-10-29T14:53:09.365915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KNeighborsRegressor","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsRegressor\nmodel = KNeighborsRegressor(n_neighbors=2)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-29T14:53:09.701185Z","iopub.execute_input":"2022-10-29T14:53:09.701907Z","iopub.status.idle":"2022-10-29T14:53:09.721340Z","shell.execute_reply.started":"2022-10-29T14:53:09.701870Z","shell.execute_reply":"2022-10-29T14:53:09.719638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2022-10-29T12:55:24.383485Z","iopub.status.idle":"2022-10-29T12:55:24.384280Z","shell.execute_reply.started":"2022-10-29T12:55:24.383977Z","shell.execute_reply":"2022-10-29T12:55:24.384005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}