{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 9 Blend and BIG SURPRISE(!) - how to explain ? ideas - welcome ! \n\n    added blending part for all solutions - and got suprise:\n    and the WORST/TERRIBLE (r2 negative, mse solution 10 times worse than others -   KernelRidge\n    enters the blends and improve it ! \n    \n    The only thing - that it is very uncorrelated with the other solutions. \n\n#### 8: Added - statistics on all methods - collected in one table (dataframe)\n    \n    Still Ridge, SVR are the best. Cat\n\n#### 7: added XGBoost+Optuna, CatBoost, MLP\n    \n    CatBoost is quite good with default params\n\n#### 3,4,5,6 - search for optimal LightGBM params; added: models and param tuning for other models,\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n    Version 6 - seems around 40 PCA features is better for LightGBM  - current_best_r2 = 0.46792863038644295 # \n    \n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"code","source":"tune_RF = False # True # Takes 4 minutes for 6 params ","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:02.225804Z","iopub.execute_input":"2022-11-04T14:31:02.226292Z","iopub.status.idle":"2022-11-04T14:31:02.231788Z","shell.execute_reply.started":"2022-11-04T14:31:02.226249Z","shell.execute_reply":"2022-11-04T14:31:02.230548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:02.234127Z","iopub.execute_input":"2022-11-04T14:31:02.234494Z","iopub.status.idle":"2022-11-04T14:31:26.153246Z","shell.execute_reply.started":"2022-11-04T14:31:02.234447Z","shell.execute_reply":"2022-11-04T14:31:26.151883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score\nimport lightgbm as lgbm\nfrom sklearn.linear_model import Ridge, Lasso, ElasticNet, SGDRegressor, OrthogonalMatchingPursuit\nfrom sklearn.svm import NuSVR\nfrom sklearn.neural_network import MLPRegressor\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom typing import *\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:26.155596Z","iopub.execute_input":"2022-11-04T14:31:26.156691Z","iopub.status.idle":"2022-11-04T14:31:27.245910Z","shell.execute_reply.started":"2022-11-04T14:31:26.156638Z","shell.execute_reply":"2022-11-04T14:31:27.244801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:27.248172Z","iopub.execute_input":"2022-11-04T14:31:27.248599Z","iopub.status.idle":"2022-11-04T14:31:27.256337Z","shell.execute_reply.started":"2022-11-04T14:31:27.248567Z","shell.execute_reply":"2022-11-04T14:31:27.255177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:27.257610Z","iopub.execute_input":"2022-11-04T14:31:27.257953Z","iopub.status.idle":"2022-11-04T14:31:27.687950Z","shell.execute_reply.started":"2022-11-04T14:31:27.257921Z","shell.execute_reply":"2022-11-04T14:31:27.686893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:27.689313Z","iopub.execute_input":"2022-11-04T14:31:27.689676Z","iopub.status.idle":"2022-11-04T14:31:48.708888Z","shell.execute_reply.started":"2022-11-04T14:31:27.689648Z","shell.execute_reply":"2022-11-04T14:31:48.707768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.mean(axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:48.710295Z","iopub.execute_input":"2022-11-04T14:31:48.710616Z","iopub.status.idle":"2022-11-04T14:31:48.889167Z","shell.execute_reply.started":"2022-11-04T14:31:48.710579Z","shell.execute_reply":"2022-11-04T14:31:48.888056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.std(axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:48.890407Z","iopub.execute_input":"2022-11-04T14:31:48.890738Z","iopub.status.idle":"2022-11-04T14:31:50.385677Z","shell.execute_reply.started":"2022-11-04T14:31:48.890710Z","shell.execute_reply":"2022-11-04T14:31:50.384700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:50.388948Z","iopub.execute_input":"2022-11-04T14:31:50.389294Z","iopub.status.idle":"2022-11-04T14:31:51.241035Z","shell.execute_reply.started":"2022-11-04T14:31:50.389263Z","shell.execute_reply":"2022-11-04T14:31:51.239893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:51.243014Z","iopub.execute_input":"2022-11-04T14:31:51.243845Z","iopub.status.idle":"2022-11-04T14:31:51.637149Z","shell.execute_reply.started":"2022-11-04T14:31:51.243800Z","shell.execute_reply":"2022-11-04T14:31:51.636104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create \"Playground\"  (i.e. downsample)\n\nSmall 10% of data of data where we can make prelimanary experiments. \n\nit will be labeled by special column \"Playground\" in df_meta (meta data for CITE-seq train only)","metadata":{}},{"cell_type":"code","source":"# Prepare for creation of additional holdout folds with 10% of samples \n# We will use stratified Kfold to achieve that days, cell_types and donors are equally distributed \nimport numpy as np\nfrom sklearn.model_selection import StratifiedKFold\nscol = 'donor&day&CT'\ndf_meta[scol] =df_meta['donor'].apply(lambda x:str(x)+'_') + df_meta['day'].apply(lambda x:str(x)+'_') + df_meta['cell_type']\n\n\nskf = StratifiedKFold(n_splits=10,  shuffle=True, random_state=40)\nskf.get_n_splits(df_meta, df_meta[scol] )\n\n\n\ny = df_meta[scol] \nfor train_index, test_index in skf.split(df_meta, df_meta[scol]):\n    print(\"TRAIN:\", len(train_index), \"TEST:\", len(test_index) ); \n    break\nprint(test_index)\nprint(df_meta[scol].value_counts().head(5)   )\nprint(df_meta.iloc[test_index,:][scol].value_counts().head(5)    )\n\n\nflagged_column_name = 'Playground'\ndf_meta[flagged_column_name] = 0 \ndf_meta.loc[df_meta.index[test_index],flagged_column_name]  = 1\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:31:51.638555Z","iopub.execute_input":"2022-11-04T14:31:51.638869Z","iopub.status.idle":"2022-11-04T14:31:51.867253Z","shell.execute_reply.started":"2022-11-04T14:31:51.638841Z","shell.execute_reply":"2022-11-04T14:31:51.866017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"markdown","source":"### Preparations","metadata":{}},{"cell_type":"markdown","source":"If you want to add new model, just add new `model_name:[model_class, params_dict]` pair to **models_dict**.\n\nEach model must have *fit* method","metadata":{}},{"cell_type":"code","source":"models_dict = {\n    \"lgb_default\": [lgbm.LGBMRegressor, dict(random_state = 0)],\n    \"lgb_tuned\": [lgbm.LGBMRegressor, dict(random_state = 0, \n        learning_rate= 0.07 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n        num_leaves = 12, # hypersensitive - change to 11, worsens about 2%\n        max_depth = 9, # hypersensitive - change to 10 - worsens about 1%\n        n_estimators= 150, # senstitive - change by 10 - worsents about 0.5%\n        min_child_samples = 20,# hypersensitive - change to 21 - worsens about 3% \n        min_split_gain = 12.2,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n        reg_lambda = 0.0, # seems only makes worse\n        reg_alpha = 0.0, # seems only makes worse\n        subsample_for_bin = 10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n        colsample_bytree = 1, # only worsens\n        subsample = 1, # Does not seem to influence at all\n        other_rate = 1,# no influence ?  \n        min_child_weight =  0.1, # no influence ?\n        subsample_freq = 10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n                  )],\n    \"ridge_default\": [Ridge, dict(random_state=0)],\n#     \"ridge_tuned\": [Ridge, dict(random_state=0, alpha=1e-04)],\n    \"lasso_default\": [Lasso, dict(random_state=0)],\n    \"elasticnet_default\": [ElasticNet, dict(random_state=0)],\n#     \"nu_svr_default\": [NuSVR, dict()],\n#     \"nu_svr\": [NuSVR, dict(kernel='poly', degree=4, gamma='auto', nu=0.52, coef0=0.053)],\n    \"OMP_default\": [OrthogonalMatchingPursuit, dict(normalize=True)],\n#     \"SGD_default\": [SGDRegressor, dict(random_state=0)],\n    \"MLP_regressor\": [MLPRegressor, dict(max_iter=500, activation='logistic', early_stopping=True,\n                         solver='adam', alpha=1e-4, random_state=0, \n                         hidden_layer_sizes=(300, 200))]\n}","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:00:57.447296Z","iopub.execute_input":"2022-11-04T15:00:57.447734Z","iopub.status.idle":"2022-11-04T15:00:57.457632Z","shell.execute_reply.started":"2022-11-04T15:00:57.447697Z","shell.execute_reply":"2022-11-04T15:00:57.456636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"r2_dict = {model_name: [] for model_name in models_dict.keys()}\ncorr_dict = {model_name: [] for model_name in models_dict.keys()}\nmse_dict = {model_name: [] for model_name in models_dict.keys()}\ntime_dict = {model_name: [] for model_name in models_dict.keys()}","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:00:58.005781Z","iopub.execute_input":"2022-11-04T15:00:58.006207Z","iopub.status.idle":"2022-11-04T15:00:58.011938Z","shell.execute_reply.started":"2022-11-04T15:00:58.006172Z","shell.execute_reply":"2022-11-04T15:00:58.011183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# to see the whole dataset\npd.set_option('display.max_rows', None)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:01:00.781869Z","iopub.execute_input":"2022-11-04T15:01:00.782961Z","iopub.status.idle":"2022-11-04T15:01:00.788415Z","shell.execute_reply.started":"2022-11-04T15:01:00.782880Z","shell.execute_reply":"2022-11-04T15:01:00.787403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_columns = df_cite_train_y.columns","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:01:01.255171Z","iopub.execute_input":"2022-11-04T15:01:01.255612Z","iopub.status.idle":"2022-11-04T15:01:01.260270Z","shell.execute_reply.started":"2022-11-04T15:01:01.255576Z","shell.execute_reply":"2022-11-04T15:01:01.259125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fit_from_scratch(name: str, \n                     x: np.ndarray, \n                     y: np.ndarray,\n                     return_time: bool = True) -> Union[object, List[Union[object, float]]]: \n    \"\"\"\n    Create model from models_dict by given name. \n    Fit (x, y) into model.\n    Return trained model or trained model and time spent for training.\n    \"\"\"\n    params = models_dict[name][1]\n    model = models_dict[name][0](**params)\n    \n    s_t = time.time()\n    model.fit(x, y)\n    total_time = time.time() - s_t\n    \n    if return_time:\n        return model, total_time\n    else:\n        return model\n\n\ndef get_stats_from_model(model: object,\n                         x: np.ndarray, \n                         y: np.ndarray) -> List[float]:\n    \"\"\"\n    Get y_pred from model.predict(x).\n    Return r2 and mse of (y, y_pred)\n    \"\"\"\n    y_pred = model.predict(x)\n    \n    r2_ = r2_score(y, y_pred)\n    mse_ = mean_squared_error(y, y_pred)\n    corr_ = np.corrcoef(y, y_pred)[1, 0]\n    \n    return r2_, mse_, corr_","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:01:01.786530Z","iopub.execute_input":"2022-11-04T15:01:01.787262Z","iopub.status.idle":"2022-11-04T15:01:01.797281Z","shell.execute_reply.started":"2022-11-04T15:01:01.787213Z","shell.execute_reply":"2022-11-04T15:01:01.796215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:01:03.451275Z","iopub.execute_input":"2022-11-04T15:01:03.452441Z","iopub.status.idle":"2022-11-04T15:01:03.457317Z","shell.execute_reply.started":"2022-11-04T15:01:03.452390Z","shell.execute_reply":"2022-11-04T15:01:03.456132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for selected_target in tqdm(target_columns):\n#     print(\"Target: \", selected_target)\n    n_features = 40\n\n    X= df_cite.iloc[:70988,:n_features][df_meta['Playground']==1]\n    y= df_cite_train_y[df_meta['Playground']==1][selected_target]\n\n    # Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n    # Private like test - new DAY, and donor, \n    # While public like - only new donor (days are the same as in train):\n    # Step 1: \n    mask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \n    X_train = df_cite.iloc[:70988,:n_features][mask_train]\n    y_train = df_cite_train_y[mask_train][ selected_target ]\n    # Step 2:\n    mask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\n    X_test_private_like = df_cite.iloc[:70988,:n_features][ mask_test_private_like  ]\n    y_test_private_like = df_cite_train_y[mask_test_private_like][ selected_target ]\n    X_test = X_test_private_like\n    y_test = y_test_private_like\n    # Step 3: \n    mask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \n    X_test_public_like = df_cite.iloc[:70988,:n_features][mask_test_public_like ]\n    y_test_public_like = df_cite_train_y[mask_test_public_like][ selected_target ]\n    X_test2 = X_test_public_like\n    y_test2 = y_test_public_like\n\n    mask_out_of_playground = (df_meta['Playground']==0)\n    X_oop = df_cite.iloc[:70988,:n_features][mask_out_of_playground]\n    y_oop = df_cite_train_y[mask_out_of_playground][ selected_target ]\n    \n    # Fit models and update info\n    for model_name in models_dict.keys():\n        model, time_ = fit_from_scratch(model_name, \n                                        X_train, \n                                        y_train,\n                                        return_time=True)\n        r2_, mse_, corr_ = get_stats_from_model(model, X_test, y_test)\n        time_dict[model_name].append(time_)\n        r2_dict[model_name].append(r2_)\n        corr_dict[model_name].append(corr_)\n        mse_dict[model_name].append(mse_)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:01:03.956002Z","iopub.execute_input":"2022-11-04T15:01:03.956434Z","iopub.status.idle":"2022-11-04T15:10:46.309801Z","shell.execute_reply.started":"2022-11-04T15:01:03.956397Z","shell.execute_reply":"2022-11-04T15:10:46.308603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results","metadata":{}},{"cell_type":"code","source":"corr_df = pd.DataFrame(corr_dict, index=target_columns)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:10:46.312384Z","iopub.execute_input":"2022-11-04T15:10:46.313236Z","iopub.status.idle":"2022-11-04T15:10:46.324784Z","shell.execute_reply.started":"2022-11-04T15:10:46.313188Z","shell.execute_reply":"2022-11-04T15:10:46.322807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(corr_df.style.set_caption(\"Pearson corr table\"))","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:10:46.327345Z","iopub.execute_input":"2022-11-04T15:10:46.328377Z","iopub.status.idle":"2022-11-04T15:10:46.403965Z","shell.execute_reply.started":"2022-11-04T15:10:46.328329Z","shell.execute_reply":"2022-11-04T15:10:46.402739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_idxs, best_counts = np.unique(corr_df.values.argmax(1), return_counts=True)\nmodel_names = list(models_dict.keys())\nmean_abs_corrs = np.nanmean(np.abs(corr_df.values), 0)\n\nfor idx in model_idxs:\n    print(f\"{model_names[idx]} was best {best_counts[idx]} times with mean abs corr {mean_abs_corrs[idx]}\")","metadata":{"execution":{"iopub.status.busy":"2022-11-04T15:23:00.807667Z","iopub.execute_input":"2022-11-04T15:23:00.808080Z","iopub.status.idle":"2022-11-04T15:23:00.817004Z","shell.execute_reply.started":"2022-11-04T15:23:00.808034Z","shell.execute_reply":"2022-11-04T15:23:00.815822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"r2_df = pd.DataFrame(r2_dict, index=target_columns)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:47:43.793563Z","iopub.execute_input":"2022-11-04T14:47:43.794225Z","iopub.status.idle":"2022-11-04T14:47:43.800024Z","shell.execute_reply.started":"2022-11-04T14:47:43.794190Z","shell.execute_reply":"2022-11-04T14:47:43.798820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(r2_df.style.set_caption(\"R2 score table\"))","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:48:02.115581Z","iopub.execute_input":"2022-11-04T14:48:02.115968Z","iopub.status.idle":"2022-11-04T14:48:02.149453Z","shell.execute_reply.started":"2022-11-04T14:48:02.115936Z","shell.execute_reply":"2022-11-04T14:48:02.148327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mse_df = pd.DataFrame(mse_dict, index=target_columns)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:48:02.310927Z","iopub.execute_input":"2022-11-04T14:48:02.312104Z","iopub.status.idle":"2022-11-04T14:48:02.318761Z","shell.execute_reply.started":"2022-11-04T14:48:02.312020Z","shell.execute_reply":"2022-11-04T14:48:02.317441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(mse_df.style.set_caption(\"MSE score table\"))","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:48:02.815485Z","iopub.execute_input":"2022-11-04T14:48:02.816943Z","iopub.status.idle":"2022-11-04T14:48:02.848933Z","shell.execute_reply.started":"2022-11-04T14:48:02.816889Z","shell.execute_reply":"2022-11-04T14:48:02.847613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Averaged time spent for training\npd.DataFrame(time_dict).mean(0)","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:48:28.421677Z","iopub.execute_input":"2022-11-04T14:48:28.422160Z","iopub.status.idle":"2022-11-04T14:48:28.434853Z","shell.execute_reply.started":"2022-11-04T14:48:28.422118Z","shell.execute_reply":"2022-11-04T14:48:28.433608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2022-11-04T14:48:10.321342Z","iopub.execute_input":"2022-11-04T14:48:10.321763Z","iopub.status.idle":"2022-11-04T14:48:10.328462Z","shell.execute_reply.started":"2022-11-04T14:48:10.321725Z","shell.execute_reply":"2022-11-04T14:48:10.327210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}