{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n### Briefly - downsample and make quick experiments\n\nHere we will downsample the data. I.e. select 10% of data as a kind of \"Playground\" - for quick experiments. \nAnd will train/tune different models on that playground first.\n\n\n### About CV scheme\n\nPublic and private test sets are quite different. (Private - contains new DAY (and donor), while public ONLY new donor).\nSo one should be careful with the validation schemes.\nSome proposals are described here: \n\nNotebook: https://www.kaggle.com/code/alexandervc/mmscel-crossvalidation-schemes\nTopic: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\n\nWe will be based on them. \n\n### Start with selection of \"Playground\" - 10% part of data - to quickly test models, ideas\n\nData is quite big, and so training models, tuning params might take long time. That is not always affordable.\nIn the present script we first choose some 10% part of data - to make quick experiments.\nThat part is chosen to have SAME proporitions of key characteristics: donors, days, cell types as the initial data\n\nSo we can create CV scheme for that \"playground\" part.\nBut we can also simplify even further - for start - use  not 6-fold scheme but just splite by days. And have only 1 train subset for quick experiments. \n\nTo simplify even further we can start play with just one target only. \n\n\n\n### Versions\n\n#### 3,4,5,6 - search for optimal LightGBM params\n    Surprise - Ridge is better than LightGBM ! Even tuning of params of LightGBM does not help much !\n    \n    Version 5 - changed number of features to 100 - Ridge - quite improved, boosting - less. \n\n#### 1,2 Test several models - Ridge,LGB, SVR, RF etc...\n\n    Target chosen - only CD31\n    No params tuning - just fist look \n","metadata":{}},{"cell_type":"markdown","source":"# Install/import modules, load technical data\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","scrolled":true,"execution":{"iopub.status.busy":"2022-11-06T15:13:08.244855Z","iopub.execute_input":"2022-11-06T15:13:08.245176Z","iopub.status.idle":"2022-11-06T15:13:08.306620Z","shell.execute_reply.started":"2022-11-06T15:13:08.245104Z","shell.execute_reply":"2022-11-06T15:13:08.305768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:13:08.308260Z","iopub.execute_input":"2022-11-06T15:13:08.308500Z","iopub.status.idle":"2022-11-06T15:13:30.762074Z","shell.execute_reply.started":"2022-11-06T15:13:08.308477Z","shell.execute_reply":"2022-11-06T15:13:30.760718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:13:30.764952Z","iopub.execute_input":"2022-11-06T15:13:30.766197Z","iopub.status.idle":"2022-11-06T15:13:48.294719Z","shell.execute_reply.started":"2022-11-06T15:13:30.766162Z","shell.execute_reply":"2022-11-06T15:13:48.293281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load prepared Features for CITE-seq part of task","metadata":{}},{"cell_type":"code","source":"%%time\n\nprint('Load prepared features for CITE-seq')\n# These files contain both train and test parts .\n# For CITEseq part - first 70988 elements - train, and later 48663 - test. Overall 119651 samples.\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_TruncatedSVD200_niter7_rs42.csv'\nfn = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/citeseq_train_and_test_PCA500.csv'\ndf_cite = pd.read_csv(fn,index_col = 0)\ndisplay(df_cite)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:13:48.297212Z","iopub.execute_input":"2022-11-06T15:13:48.297564Z","iopub.status.idle":"2022-11-06T15:14:03.718320Z","shell.execute_reply.started":"2022-11-06T15:13:48.297535Z","shell.execute_reply":"2022-11-06T15:14:03.717420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.mean(axis = 0).plot()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:03.719731Z","iopub.execute_input":"2022-11-06T15:14:03.720293Z","iopub.status.idle":"2022-11-06T15:14:04.069952Z","shell.execute_reply.started":"2022-11-06T15:14:03.720259Z","shell.execute_reply":"2022-11-06T15:14:04.068916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite.std(axis=0).plot()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:04.071469Z","iopub.execute_input":"2022-11-06T15:14:04.071841Z","iopub.status.idle":"2022-11-06T15:14:04.569150Z","shell.execute_reply.started":"2022-11-06T15:14:04.071809Z","shell.execute_reply":"2022-11-06T15:14:04.568033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Targets for CITE-seq","metadata":{}},{"cell_type":"code","source":"%%time\n#if 1:\nprint('Load CITE-seq targets and ')\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndisplay(df_cite_train_y.sample(10))","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:04.570490Z","iopub.execute_input":"2022-11-06T15:14:04.571339Z","iopub.status.idle":"2022-11-06T15:14:05.274837Z","shell.execute_reply.started":"2022-11-06T15:14:04.571303Z","shell.execute_reply":"2022-11-06T15:14:05.273780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load and start prepare some metadata (cut only CITE-seq train part)","metadata":{}},{"cell_type":"code","source":"%%time\nfn2 = '/kaggle/input/feature-shop-for-multimodal-singlecell-competition/_citeseq_meta_all_text_also.csv'\ndf_meta_full = pd.read_csv(fn2,index_col = 0)\ndisplay(df_meta_full)\n#if 1:\ndf_meta = pd.DataFrame(index = df_cite_train_y.index) \ndf_meta = df_meta.join(df_cell.set_index('cell_id') )\ndisplay(df_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:05.275948Z","iopub.execute_input":"2022-11-06T15:14:05.276222Z","iopub.status.idle":"2022-11-06T15:14:05.591409Z","shell.execute_reply.started":"2022-11-06T15:14:05.276198Z","shell.execute_reply":"2022-11-06T15:14:05.589986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create X,y, X_train, y_train, etc - INSIDE \"Playground\"\n","metadata":{}},{"cell_type":"code","source":"df_cite.sample(5)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T17:22:09.518251Z","iopub.execute_input":"2022-11-03T17:22:09.519153Z","iopub.status.idle":"2022-11-03T17:22:09.550918Z","shell.execute_reply.started":"2022-11-03T17:22:09.519113Z","shell.execute_reply":"2022-11-03T17:22:09.549719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.sample(5)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T17:22:09.554480Z","iopub.execute_input":"2022-11-03T17:22:09.554873Z","iopub.status.idle":"2022-11-03T17:22:09.575980Z","shell.execute_reply.started":"2022-11-03T17:22:09.554838Z","shell.execute_reply":"2022-11-03T17:22:09.574616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import PredefinedSplit","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:29.164153Z","iopub.execute_input":"2022-11-02T13:19:29.164511Z","iopub.status.idle":"2022-11-02T13:19:29.170124Z","shell.execute_reply.started":"2022-11-02T13:19:29.164478Z","shell.execute_reply":"2022-11-02T13:19:29.168596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_splits = []\ndf_meta_modified = df_meta.copy()\ndf_meta_modified.reset_index(inplace=True)\nfor i, day2test in enumerate(df_meta_modified.day.unique()):\n    print('Day =', day2test)\n    df_meta_modified[f'predefined_{i}'] = 42\n    idx2choose = df_meta_modified.loc[df_meta_modified.day==day2test].index\n    df_meta_modified.loc[idx2choose, [f'predefined_{i}']] = -1 # just magic number, do not pay attention :)\n    for j, pat in enumerate(df_meta_modified.donor.unique()):\n        print('donor ->', pat)\n        df_meta_modified[f'predefined_{i}_{j}'] = df_meta_modified[f'predefined_{i}'].values\n        # exclude from this split iteration samples not with that donor AND that day\n        idx2drop = df_meta_modified.loc[(((df_meta_modified.donor == pat)&(df_meta_modified.day != day2test))|\\\n                            ((df_meta_modified.donor != pat)&(df_meta_modified.day == day2test)))].index\n        \n#         display(df_meta_modified[['day', 'donor', f'predefined_{i}_{j}']].value_counts().sort_index())\n        print(df_meta_modified.shape, df_meta_modified.drop(idx2drop, axis=0).shape, idx2drop.shape)\n        display(df_meta_modified.drop(idx2drop, axis=0)[[f'predefined_{i}_{j}', 'day', 'donor']].value_counts().sort_index())\n        data = df_meta_modified.drop(idx2drop, axis=0)\n        ps = PredefinedSplit(data[f'predefined_{i}_{j}'])\n        print(data[f'predefined_{i}_{j}'].unique(), '-> variants to split =', ps.get_n_splits())\n        \n        for test_index, train_index in ps.split():\n#             print(train_index)\n            print('\\tTrain:')\n            display(data.iloc[train_index][['day', 'donor']].value_counts().sort_index())\n            print()\n            print('\\tTest:')\n            display(data.iloc[test_index][['day', 'donor']].value_counts().sort_index())\n            print('-----------')\n            \n            train_test_splits.append((df_meta_modified.iloc[data.iloc[train_index].index].index,\n                                      df_meta_modified.iloc[data.iloc[test_index].index].index))","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:29.171760Z","iopub.execute_input":"2022-11-02T13:19:29.172098Z","iopub.status.idle":"2022-11-02T13:19:29.948101Z","shell.execute_reply.started":"2022-11-02T13:19:29.172070Z","shell.execute_reply":"2022-11-02T13:19:29.946952Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for train_ind, test_ind in train_test_splits:\n    display(df_meta.iloc[test_ind][['day', 'donor']].value_counts().sort_index())\n    print(max(train_ind), max(test_ind))","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:29.949963Z","iopub.execute_input":"2022-11-02T13:19:29.950399Z","iopub.status.idle":"2022-11-02T13:19:30.075825Z","shell.execute_reply.started":"2022-11-02T13:19:29.950327Z","shell.execute_reply":"2022-11-02T13:19:30.074216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def make_cv(list_train_test_pairs):\n    for train_ind, test_ind in list_train_test_pairs:\n        yield (train_ind, test_ind)","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:30.077401Z","iopub.execute_input":"2022-11-02T13:19:30.077799Z","iopub.status.idle":"2022-11-02T13:19:30.084189Z","shell.execute_reply.started":"2022-11-02T13:19:30.077761Z","shell.execute_reply":"2022-11-02T13:19:30.082498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# selected_target = 'CD31'\nn_features = 250\n\nX= df_cite.iloc[:70988, 1:n_features][df_meta['Playground']==1]\ny= df_cite_train_y[df_meta['Playground']==1] # [selected_target]\nprint('X.shape, y.shape', X.shape, y.shape )\n\n\ntrain_test_splits = []\ndf_meta_modified = df_meta.loc[X.index].copy()\ndf_meta_modified.reset_index(inplace=True)\nfor i, day2test in enumerate(df_meta_modified.day.unique()):\n    print('Day =', day2test)\n    df_meta_modified[f'predefined_{i}'] = 42\n    idx2choose = df_meta_modified.loc[df_meta_modified.day==day2test].index\n    df_meta_modified.loc[idx2choose, [f'predefined_{i}']] = -1 # just magic number, do not pay attention :)\n    for j, pat in enumerate(df_meta_modified.donor.unique()):\n        print('donor ->', pat)\n        df_meta_modified[f'predefined_{i}_{j}'] = df_meta_modified[f'predefined_{i}'].values\n        # exclude from this split iteration samples not with that donor AND that day\n        idx2drop = df_meta_modified.loc[(((df_meta_modified.donor == pat)&(df_meta_modified.day != day2test))|\\\n                            ((df_meta_modified.donor != pat)&(df_meta_modified.day == day2test)))].index\n        \n#         display(df_meta_modified[['day', 'donor', f'predefined_{i}_{j}']].value_counts().sort_index())\n        print(df_meta_modified.shape, df_meta_modified.drop(idx2drop, axis=0).shape, idx2drop.shape)\n        display(df_meta_modified.drop(idx2drop, axis=0)[[f'predefined_{i}_{j}', 'day', 'donor']].value_counts().sort_index())\n        data = df_meta_modified.drop(idx2drop, axis=0)\n        ps = PredefinedSplit(data[f'predefined_{i}_{j}'])\n        print(data[f'predefined_{i}_{j}'].unique(), '-> variants to split =', ps.get_n_splits())\n        \n        for test_index, train_index in ps.split():\n#             print(train_index)\n            print('\\tTrain:')\n            display(data.iloc[train_index][['day', 'donor']].value_counts().sort_index())\n            print()\n            print('\\tTest:')\n            display(data.iloc[test_index][['day', 'donor']].value_counts().sort_index())\n            print('-----------')\n            \n            train_test_splits.append((df_meta_modified.iloc[data.iloc[train_index].index].index,\n                                      df_meta_modified.iloc[data.iloc[test_index].index].index))\n\n# # Create simplfied validation scheme - like real test data - with two test-sets private-like, public-like:\n# # Private like test - new DAY, and donor, \n# # While public like - only new donor (days are the same as in train):\n# # Step 1: \n# mask_train = (df_meta['Playground']==1)&(df_meta['day']!=4)&(df_meta['donor']!=31800) \n# X_train = df_cite.iloc[:70988,:n_features][mask_train]\n# y_train = df_cite_train_y[mask_train] # [ selected_target ]\n# # Step 2:\n# mask_test_private_like = (df_meta['Playground']==1)&(df_meta['day']==4)\n# X_test_private_like = df_cite.iloc[:70988,:n_features][ mask_test_private_like  ]\n# y_test_private_like = df_cite_train_y[mask_test_private_like] # [ selected_target ]\n# X_test = X_test_private_like\n# y_test = y_test_private_like\n# # Step 3: \n# mask_test_public_like = (df_meta['Playground']==1)&(df_meta['day']!=4)  &(df_meta['donor']==31800) \n# X_test_public_like = df_cite.iloc[:70988,:n_features][mask_test_public_like ]\n# y_test_public_like = df_cite_train_y[mask_test_public_like] # [ selected_target ]\n\nX.shape,y.shape, # X_train.shape, X_test_private_like.shape, X_test_public_like.shape, y_train.shape, y_test_private_like.shape, y_test_public_like.shape\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:30.086191Z","iopub.execute_input":"2022-11-02T13:19:30.086545Z","iopub.status.idle":"2022-11-02T13:19:30.484553Z","shell.execute_reply.started":"2022-11-02T13:19:30.086517Z","shell.execute_reply":"2022-11-02T13:19:30.483242Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import r2_score","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:05.865424Z","iopub.execute_input":"2022-11-06T15:14:05.866188Z","iopub.status.idle":"2022-11-06T15:14:05.871489Z","shell.execute_reply.started":"2022-11-06T15:14:05.866160Z","shell.execute_reply":"2022-11-06T15:14:05.870056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV, RandomizedSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:05.872641Z","iopub.execute_input":"2022-11-06T15:14:05.872957Z","iopub.status.idle":"2022-11-06T15:14:05.882488Z","shell.execute_reply.started":"2022-11-06T15:14:05.872933Z","shell.execute_reply":"2022-11-06T15:14:05.881512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LightGBM model","metadata":{}},{"cell_type":"code","source":"%%time\nimport lightgbm as lgbm","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:05.883907Z","iopub.execute_input":"2022-11-06T15:14:05.884258Z","iopub.status.idle":"2022-11-06T15:14:06.674323Z","shell.execute_reply.started":"2022-11-06T15:14:05.884226Z","shell.execute_reply":"2022-11-06T15:14:06.673348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.multioutput import MultiOutputRegressor","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.675781Z","iopub.execute_input":"2022-11-06T15:14:06.676082Z","iopub.status.idle":"2022-11-06T15:14:06.681957Z","shell.execute_reply.started":"2022-11-06T15:14:06.676057Z","shell.execute_reply":"2022-11-06T15:14:06.680909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sklearn\nsklearn.metrics.SCORERS.keys()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.683179Z","iopub.execute_input":"2022-11-06T15:14:06.683997Z","iopub.status.idle":"2022-11-06T15:14:06.695342Z","shell.execute_reply.started":"2022-11-06T15:14:06.683958Z","shell.execute_reply":"2022-11-06T15:14:06.694165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"est = MultiOutputRegressor(estimator=lgbm.LGBMRegressor())","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.696831Z","iopub.execute_input":"2022-11-06T15:14:06.697644Z","iopub.status.idle":"2022-11-06T15:14:06.703538Z","shell.execute_reply.started":"2022-11-06T15:14:06.697591Z","shell.execute_reply":"2022-11-06T15:14:06.702790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"est.get_params().keys()","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:31.505757Z","iopub.execute_input":"2022-11-02T13:19:31.506204Z","iopub.status.idle":"2022-11-02T13:19:31.518631Z","shell.execute_reply.started":"2022-11-02T13:19:31.506167Z","shell.execute_reply":"2022-11-02T13:19:31.517536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'estimator__boosting_type': ['gbdt', 'dart', 'goss', 'rf'],\n    'estimator__objective': ['rmse'],\n    'estimator__n_estimators': np.arange(50, 200, 10),\n    'estimator__learning_rate': np.logspace(-5, -1, 10),\n   'estimator__num_leaves':np.arange(5, 150, 20),\n   'estimator__max_depth': np.arange(5, 10).tolist() + [-1],\n   'estimator__min_child_samples': np.arange(15, 25),\n  'estimator__colsample_bytree': np.arange(0.1, 0.9),\n    'estimator__reg_lambda': np.logspace(-3, 5, 9),\n    'estimator__reg_alpha': np.logspace(-3, 5, 9),\n    'estimator__random_state': [42],\n    'estimator__n_jobs': [-1],\n     \n}","metadata":{"execution":{"iopub.status.busy":"2022-11-02T13:19:31.531754Z","iopub.execute_input":"2022-11-02T13:19:31.532057Z","iopub.status.idle":"2022-11-02T13:19:31.543007Z","shell.execute_reply.started":"2022-11-02T13:19:31.532032Z","shell.execute_reply":"2022-11-02T13:19:31.541619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nParamSearch = RandomizedSearchCV(\n    estimator=MultiOutputRegressor(lgbm.LGBMRegressor()),\n    param_distributions = params,\n    n_iter=200, # 10**3\n    scoring='neg_mean_squared_error',\n    n_jobs=-1,\n    refit=True,\n    cv=make_cv(train_test_splits),\n    verbose=-1,\n    random_state=33,)\nParamSearch.fit(X, y)\n \nprint('best params')\nprint(ParamSearch.best_params_)\nprint('best score')\nprint (ParamSearch.best_score_)\n\n","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"scrolled":true,"execution":{"iopub.status.busy":"2022-11-02T13:34:18.746620Z","iopub.execute_input":"2022-11-02T13:34:18.746990Z","iopub.status.idle":"2022-11-02T22:40:31.199358Z","shell.execute_reply.started":"2022-11-02T13:34:18.746962Z","shell.execute_reply":"2022-11-02T22:40:31.198615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"best params\n{'estimator__reg_lambda': 100.0, 'estimator__reg_alpha': 1.0, 'estimator__random_state': 42, 'estimator__objective': 'rmse', 'estimator__num_leaves': 25, 'estimator__n_jobs': -1, 'estimator__n_estimators': 170, 'estimator__min_child_samples': 17, 'estimator__max_depth': 9, 'estimator__learning_rate': 0.1, 'estimator__colsample_bytree': 0.1, 'estimator__boosting_type': 'gbdt'}\nbest score\n-4.037030972299572\nCPU times: user 6min 18s, sys: 2.44 s, total: 6min 20s\nWall time: 9h 6min 12s","metadata":{}},{"cell_type":"code","source":"import joblib\n\njoblib.dump(ParamSearch.best_estimator_, 'MultiLGBMReg_model.joblib')","metadata":{"execution":{"iopub.status.busy":"2022-11-02T22:40:31.200869Z","iopub.execute_input":"2022-11-02T22:40:31.201215Z","iopub.status.idle":"2022-11-02T22:40:32.339388Z","shell.execute_reply.started":"2022-11-02T22:40:31.201181Z","shell.execute_reply":"2022-11-02T22:40:32.338544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_params = {'estimator__reg_lambda': 100.0, 'estimator__reg_alpha': 1.0, 'estimator__random_state': 42, 'estimator__objective': 'rmse', 'estimator__num_leaves': 25, 'estimator__n_jobs': -1, 'estimator__n_estimators': 170, 'estimator__min_child_samples': 17, 'estimator__max_depth': 9, 'estimator__learning_rate': 0.1, 'estimator__colsample_bytree': 0.1, 'estimator__boosting_type': 'gbdt'}","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.704892Z","iopub.execute_input":"2022-11-06T15:14:06.705407Z","iopub.status.idle":"2022-11-06T15:14:06.714609Z","shell.execute_reply.started":"2022-11-06T15:14:06.705374Z","shell.execute_reply":"2022-11-06T15:14:06.713938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = MultiOutputRegressor(lgbm.LGBMRegressor())\nmodel.set_params(**best_params)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.715979Z","iopub.execute_input":"2022-11-06T15:14:06.716292Z","iopub.status.idle":"2022-11-06T15:14:06.737068Z","shell.execute_reply.started":"2022-11-06T15:14:06.716260Z","shell.execute_reply":"2022-11-06T15:14:06.736083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_raw = df_cite.copy()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:06.738151Z","iopub.execute_input":"2022-11-06T15:14:06.738413Z","iopub.status.idle":"2022-11-06T15:14:06.812338Z","shell.execute_reply.started":"2022-11-06T15:14:06.738389Z","shell.execute_reply":"2022-11-06T15:14:06.811291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y = df_cite_train_y.copy()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:18.996475Z","iopub.execute_input":"2022-11-06T15:14:18.996845Z","iopub.status.idle":"2022-11-06T15:14:19.004995Z","shell.execute_reply.started":"2022-11-06T15:14:18.996816Z","shell.execute_reply":"2022-11-06T15:14:19.003892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel.fit(X_raw.iloc[:70988,:], Y)\nY_pred4submit = model.predict(X_raw.iloc[70988:,:])\nprint(Y_pred4submit.shape)\nprint(Y_pred4submit[:3,:3])","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:14:20.035729Z","iopub.execute_input":"2022-11-06T15:14:20.036105Z","iopub.status.idle":"2022-11-06T15:29:52.194358Z","shell.execute_reply.started":"2022-11-06T15:14:20.036076Z","shell.execute_reply":"2022-11-06T15:29:52.193684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmode_subm = 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' # 'put_zeros_to_multiome_part'\n\nif mode_subm == 'USE: sskknts MSCI CITEseq Keras Quickstart + Dropout' :\n    df_submission_full = pd.read_csv('../input/msci-citeseq-keras-quickstart-dropout/submission.csv',\n                             index_col='row_id', squeeze=True)\n    df_submission_full = df_submission_full.to_frame()\nelif mode_subm == 'put_zeros_to_multiome_part':\n    df_submission_full = pd.DataFrame(index = range(65_744_180), columns = ['target'], data = np.zeros(65_744_180) )\n    df_submission_full.index.name = 'row_id'\n    \ndisplay(df_submission_full.info() )\nprint()\ndisplay(df_submission_full)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:29:52.195678Z","iopub.execute_input":"2022-11-06T15:29:52.196109Z","iopub.status.idle":"2022-11-06T15:30:51.797627Z","shell.execute_reply.started":"2022-11-06T15:29:52.196083Z","shell.execute_reply":"2022-11-06T15:30:51.796917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_submission_full['target'].iloc[:6_812_820] = Y_pred4submit.ravel()\ndisplay(df_submission_full)\nsubmit_filename_postfix = 'RandSearch_MultiOutput_lightGBM_MBr'\ndf_submission_full.to_csv('submission_cite_seq_'+submit_filename_postfix +'.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-06T15:30:51.798691Z","iopub.execute_input":"2022-11-06T15:30:51.798989Z","iopub.status.idle":"2022-11-06T15:32:29.762386Z","shell.execute_reply.started":"2022-11-06T15:30:51.798964Z","shell.execute_reply":"2022-11-06T15:32:29.761225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#         params ={\n#                         'task': 'train',\n#                         'boosting': 'goss',\n#                         'objective': 'regression',\n#                         'metric': 'rmse',\n#                         'learning_rate': 0.005,\n#                         'subsample': 0.9855232997390695,\n#                         'max_depth': 8,\n#                         'top_rate': 0.9064148448434349,\n#                         'num_leaves': 87,\n#                         'min_child_weight': 41.9612869171337,\n#                         'other_rate': 0.0721768246018207,\n#                         'reg_alpha': 9.677537745007898,\n#                         'colsample_bytree': 0.5665320670155495,\n#                         'min_split_gain': 9.820197773625843,\n#                         'reg_lambda': 8.2532317400459,\n#                         'min_data_in_leaf': 21,\n#                         'verbose': -1,\n#                         'seed':int(2**n_fold),\n#                         'bagging_seed':int(2**n_fold),\n#                         'drop_seed':int(2**n_fold)\n#                         }\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T05:11:21.098538Z","iopub.execute_input":"2022-10-31T05:11:21.099054Z","iopub.status.idle":"2022-10-31T05:11:21.105936Z","shell.execute_reply.started":"2022-10-31T05:11:21.098998Z","shell.execute_reply":"2022-10-31T05:11:21.104335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# (0.4354558552914086, 22.998228325953246) # (learning_rate= 0.067 ,num_leaves = 12, max_depth = 9, n_estimators= 150, subsample = 0.4, random_state = 0 )\n# (0.43090615704586654, 23.18357255464802) # learning_rate= 0.067 ,num_leaves = 12, max_depth = 9\n# (0.42965094120477576, 23.23470715728686) # learning_rate= 0.067 ,num_leaves = 12, max_depth = 8\n# (0.42768216014074323, 23.314910763787474) # learning_rate= 0.067 ,num_leaves = 12\n# (0.4234603331885415, 23.486898271070043) # learning_rate= 0.067,num_leaves = 19\n# (0.42228006014959485, 23.534979876540724) # learning_rate= 0.067\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T05:11:21.352215Z","iopub.execute_input":"2022-10-31T05:11:21.353330Z","iopub.status.idle":"2022-10-31T05:11:21.358320Z","shell.execute_reply.started":"2022-10-31T05:11:21.353284Z","shell.execute_reply":"2022-10-31T05:11:21.357180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape, y_train.shape)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T05:11:21.627580Z","iopub.execute_input":"2022-10-31T05:11:21.628272Z","iopub.status.idle":"2022-10-31T05:11:21.634847Z","shell.execute_reply.started":"2022-10-31T05:11:21.628233Z","shell.execute_reply":"2022-10-31T05:11:21.633707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# \ncurrent_best_r2 = 0.4471078544471576 # 0.4366682658232913 # 0.43590181064083766\nmodel = lgbm.LGBMRegressor(random_state = 0, \n    learning_rate= 0.07 ,  # hypersensitive - change to 0.0671 - worsens about 2% from 0.43590181064083766\n    num_leaves = 12, # hypersensitive - change to 11, worsens about 2%\n    max_depth = 9, # hypersensitive - change to 10 - worsens about 1%\n    n_estimators= 150, # senstitive - change by 10 - worsents about 0.5%\n    min_child_samples = 20,# hypersensitive - change to 21 - worsens about 3% \n    min_split_gain = 12.2,# sometimes hypersenstive - change by 0.1 worsens by 0.9%, But for 100 features changes 11.8-12.2 - not change AT ALLL !!!  # Uplifts to  0.43590.. !!! \n    reg_lambda = 0, # seems only makes worse\n    reg_alpha = 0, # seems only makes worse\n    subsample_for_bin = 10000, # seems  less 10000 - worsens, but after 10 000 does not influence  \n    colsample_bytree = 1, # only worsens\n    subsample = 1, # Does not seem to influence at all\n    other_rate = 1,# no influence ?  \n    min_child_weight =  0.1, # no influence ?\n    subsample_freq = 10,# no influence ? # Integer; alias: bagging_freq; k means perform bagging at every k iteration\n              )# \nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred), (r2_score(y_test, y_pred) - current_best_r2 ) / current_best_r2 * 100","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:17:37.852181Z","iopub.execute_input":"2022-10-30T11:17:37.85263Z","iopub.status.idle":"2022-10-30T11:17:38.377989Z","shell.execute_reply.started":"2022-10-30T11:17:37.85259Z","shell.execute_reply":"2022-10-30T11:17:38.376829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dir(model)\nmodel.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:26.71301Z","iopub.execute_input":"2022-10-30T11:10:26.715555Z","iopub.status.idle":"2022-10-30T11:10:26.724255Z","shell.execute_reply.started":"2022-10-30T11:10:26.71551Z","shell.execute_reply":"2022-10-30T11:10:26.723025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ridge","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.linear_model import Ridge\n","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:26.726338Z","iopub.execute_input":"2022-10-30T11:10:26.726821Z","iopub.status.idle":"2022-10-30T11:10:26.80366Z","shell.execute_reply.started":"2022-10-30T11:10:26.726776Z","shell.execute_reply":"2022-10-30T11:10:26.802452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:26.805323Z","iopub.execute_input":"2022-10-30T11:10:26.809274Z","iopub.status.idle":"2022-10-30T11:10:26.87359Z","shell.execute_reply.started":"2022-10-30T11:10:26.809229Z","shell.execute_reply":"2022-10-30T11:10:26.87Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1e4) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:26.875776Z","iopub.execute_input":"2022-10-30T11:10:26.876743Z","iopub.status.idle":"2022-10-30T11:10:26.972765Z","shell.execute_reply.started":"2022-10-30T11:10:26.876684Z","shell.execute_reply":"2022-10-30T11:10:26.971093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodel = Ridge(alpha=1e5) # lgbm.LGBMRegressor()#**params)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:26.996872Z","iopub.execute_input":"2022-10-30T11:10:26.997901Z","iopub.status.idle":"2022-10-30T11:10:27.077168Z","shell.execute_reply.started":"2022-10-30T11:10:26.99784Z","shell.execute_reply":"2022-10-30T11:10:27.075455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Random Forest","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.ensemble import RandomForestRegressor\nmodel = RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:27.080675Z","iopub.execute_input":"2022-10-30T11:10:27.082991Z","iopub.status.idle":"2022-10-30T11:10:30.952277Z","shell.execute_reply.started":"2022-10-30T11:10:27.082911Z","shell.execute_reply":"2022-10-30T11:10:30.95115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SVR (support vector machine regression)","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.svm import SVR\nmodel = SVR(C=1.0, epsilon=0.2) # RandomForestRegressor(max_depth=2, random_state=0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:30.953652Z","iopub.execute_input":"2022-10-30T11:10:30.953982Z","iopub.status.idle":"2022-10-30T11:10:32.748015Z","shell.execute_reply.started":"2022-10-30T11:10:30.953939Z","shell.execute_reply":"2022-10-30T11:10:32.747001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KernelRidge","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.kernel_ridge import KernelRidge\nmodel = KernelRidge(alpha=1.0)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:32.752247Z","iopub.execute_input":"2022-10-30T11:10:32.752604Z","iopub.status.idle":"2022-10-30T11:10:33.388467Z","shell.execute_reply.started":"2022-10-30T11:10:32.75257Z","shell.execute_reply":"2022-10-30T11:10:33.387262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KNeighborsRegressor","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsRegressor\nmodel = KNeighborsRegressor(n_neighbors=2)\nmodel.fit(X_train,y_train)\ny_pred = model.predict(X_test)\nr2_score(y_test, y_pred), mean_squared_error(y_test, y_pred)","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:33.389964Z","iopub.execute_input":"2022-10-30T11:10:33.390741Z","iopub.status.idle":"2022-10-30T11:10:33.58399Z","shell.execute_reply.started":"2022-10-30T11:10:33.390695Z","shell.execute_reply":"2022-10-30T11:10:33.582895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2022-10-30T11:10:33.585524Z","iopub.execute_input":"2022-10-30T11:10:33.586766Z","iopub.status.idle":"2022-10-30T11:10:33.595267Z","shell.execute_reply.started":"2022-10-30T11:10:33.586719Z","shell.execute_reply":"2022-10-30T11:10:33.593643Z"},"trusted":true},"execution_count":null,"outputs":[]}]}