{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":6541963,"sourceType":"competition"},{"sourceId":6567471,"sourceType":"datasetVersion","datasetId":3755626},{"sourceId":6806175,"sourceType":"datasetVersion","datasetId":3765688}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n#### Brief\n\nWe use simple category encoders as well as some other features like: ChemBert, Morgan fingerprints, Molecula descriptors.\n\nCategory encoders considered:  'onehot', 'BackwardDifferenceEncoder', 'HelmertEncoder'  - these are contrast encoders - i.e. NOT using targets. \nTartget encoders planned to be considered in the other notebook. \n\n#### Conclusions\n\n    The simplest one-hot encoder works quite well and non of the encoders/features consitently outperform it in the present setup.\n    The Helmert encoder deserves attention - it is either the second or the first one - depending on CV-scheme\n\n#### More details \n\nFor each encoding we run the Ridge model with various alpha and several cross-validation schemes. Scores dependedce on regularization alpha is reported for each CV-scheme. \n\nOne can change n_components in tsvd as well as encoded features (compounds or cell-types, or both) - see section \"Key params\". \n\nDifferent versions of the notebook correspond to different encodings. Brief summary of the results is below. It is collected from the output files \"df_brief_report.csv\". More detailed scores collection - see https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing (sheet \"Encoders\"). \n\n                \t\t\t Alpha\tCV-MT\tCV-AmbrosM\tCV-Random\tAlpha\tCV-MT\tCV-AmbrosM\tCV-Random\tAlpha\tLB\t LB with priors\n    V9 onehot:\t\t\t\t 1.25\t2.6668\t0.99534\t1.2363\t\t1\t2.66032\t0.99785\t1.24013\t\t\t1\t0.613 \t 0.605\n    V3 BackwardDifference\t 0.2\t2.66246\t1.01891\t1.27087\t\t1\t2.6989\t1.03997\t1.27831\n    V21 HelmertEncoder\t \t 1500\t2.65864\t1.00846\t1.26319\t1\t2.67429\t1.02866\t1.29488\n    V5 ChembertClsEncoding \t0.6\t2.66555\t1.02795\t1.30795\t\t1\t2.66446\t1.03244\t1.31435\n    V6 ChembertMeanEncoding\t0.06\t2.66481\t1.02863\t1.30931\t\t1\t2.70268\t1.09262\t1.37668\n    V7 MolecularDescriptors\t0.0003\t2.66736\t1.03241\t1.33141\t1\t2.73042\t1.16143\t1.44951\n    V8 MorganFingerprints\t\t35\t2.6609\t1.01267\t1.26405\t\t1\t2.67261\t1.02744\t1.29396\n    \n    Versions above encode only compound (cell type is omitted - it works better for simple models like Ridge). \n    Here the first \"Alpha\" (Ridge regularization) is chosen as halfsum of optimal alphas for MT-scheme and AmbrosM scheme.\n    While the second \"Alpha\" is fixed to 1 for another comparaison. More detailed information can be found: \n    We copy these scores from the df_brief_report.csv file generated by the notebook. \n    \n    Encode Both CellType and Compound - apparantly even stronger regularization is required:\n                 \t\t\t Alpha\tCV-MT\tCV-AmbrosM\tCV-Random\tAlpha\tCV-MT\tCV-AmbrosM\tCV-Random\tAlpha\tLB\t LB with priors\n    V22 onehot\t\t\t \t 500001\t3.1558\t1.10573\t1.33653\t1\t2.88918\t1.50913\t1.28756\n    V25 BackwardDifference\t 2500.25\t3.15941\t1.10502\t1.33661\t1\t2.91915\t1.48635\t1.31733 \n    V27 HelmertEncoder\t \t 1500\t2.6107\t0.9987\t1.24782\t1\t2.9723\t1.59741\t1.35047\t\t\t\t1500\t?\t0.605\n\n    Encode only  CellType:\n    V20 HelmertEncoder\t \t501000\t3.15559\t1.10573\t1.33643\t1\t3.62095\t1.74043\t1.36914\n    V24 BackwardDifference\t1250\t3.15055\t1.10546\t1.33239\t1\t3.57319\t1.67329\t1.36381\n    V23 onehot\t\t\t\t500050\t3.1558\t1.10574\t1.33653\t1\t3.57055\t1.7143\t1.36555\n\n\nCollection of the CV/LB scores for other models can be found here: \nhttps://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing,  see Kaggle dataset:\nhttps://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\n\n\n#### PS\n\nWe custom CV-class supporting Ambros, MT, Random, etc cross-validation schemes proposed here: \nhttps://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes . It supports sklearn interace.\n\nChemBert features prepared by ALEKSEY TREPETSKY: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550 (please upvote)\n\nMorgan fingerprints, Molecula descriptors  by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nNotebook is based on the previous notebooks: CV-schemes: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes, \nNN analysis : https://www.kaggle.com/code/alexandervc/op2-kishan-s-nn-streamlined-and-blended","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"encoding = 'ChembertClsEncoding' # 'BackwardDifferenceEncoder' #    'onehot' # 'HelmertEncoder' # 'ChembertClsEncoding' #'ChembertMeanEncoding' #'MolecularDescriptors'# 'MorganFingerprints' # \n\n\nn_components = 30 # Reduction of targets is done by tsvd to n_components, after predicting components we take tsvd.inverse_transform\n\nlist_features_to_encode = ['cell_type','sm_name']# ['']#   [ 'sm_name'] # ['cell_type','sm_name'], ['cell_type'],  \n# For simple models like Ridge and that kind of features seems encoding only the compounds just forgetting about cell type at all gives sometimes better results. \n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:31.689509Z","iopub.execute_input":"2023-10-26T10:17:31.689940Z","iopub.status.idle":"2023-10-26T10:17:31.728025Z","shell.execute_reply.started":"2023-10-26T10:17:31.689907Z","shell.execute_reply":"2023-10-26T10:17:31.726754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create info string on params:\nstr_inf_cfg = encoding #+ ' '\nif 'sm_name' in list_features_to_encode: str_inf_cfg += 'Compound'\nif 'cell_type' in list_features_to_encode: str_inf_cfg += 'CellType'\nstr_inf_cfg += ' tsvd' + str(n_components)\nprint(str_inf_cfg)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:31.730532Z","iopub.execute_input":"2023-10-26T10:17:31.730961Z","iopub.status.idle":"2023-10-26T10:17:31.739423Z","shell.execute_reply.started":"2023-10-26T10:17:31.730927Z","shell.execute_reply":"2023-10-26T10:17:31.737918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preliminaries ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-26T10:17:31.741433Z","iopub.execute_input":"2023-10-26T10:17:31.741822Z","iopub.status.idle":"2023-10-26T10:17:44.293086Z","shell.execute_reply.started":"2023-10-26T10:17:31.741789Z","shell.execute_reply":"2023-10-26T10:17:44.291252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:44.297087Z","iopub.execute_input":"2023-10-26T10:17:44.298212Z","iopub.status.idle":"2023-10-26T10:17:53.359452Z","shell.execute_reply.started":"2023-10-26T10:17:44.298150Z","shell.execute_reply":"2023-10-26T10:17:53.358229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes\n\nSee the notebook: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\n\nMore analysis of the CV schemes can be found in slides: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing and sheet: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n\n\nCV schemes (see discussion https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494 ):\n\nAmbrosM: https://www.kaggle.com/code/ambrosm/scp-quickstart\n\nMT: https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy\n\nKishan's notebook: https://www.kaggle.com/code/liudacheldieva/neural-network-regression/notebook#Predicting-on-test-data ","metadata":{}},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ]) \n    \nimport random \nfrom sklearn.model_selection import KFold\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n            \n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['AmbrosM','MT', 'Random']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n        \n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:53.361784Z","iopub.execute_input":"2023-10-26T10:17:53.362577Z","iopub.status.idle":"2023-10-26T10:17:53.829137Z","shell.execute_reply.started":"2023-10-26T10:17:53.362514Z","shell.execute_reply":"2023-10-26T10:17:53.827939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding","metadata":{}},{"cell_type":"code","source":"%%time \nfeatures_columns = ['cell_type','sm_name']\nselected_features_columns =  list_features_to_encode #   [ 'sm_name']\nX_submit_categorical = df_id_map[selected_features_columns] # pd.DataFrame(df_id_map, columns= features_columns )\nX_full_categorical = df_de_train[ selected_features_columns ]\nY_full = df_de_train.iloc[:,5:].values\n","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:53.830726Z","iopub.execute_input":"2023-10-26T10:17:53.831049Z","iopub.status.idle":"2023-10-26T10:17:53.902308Z","shell.execute_reply.started":"2023-10-26T10:17:53.831022Z","shell.execute_reply":"2023-10-26T10:17:53.901272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One hot encoding","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.preprocessing import OneHotEncoder\n\nif encoding == 'onehot':\n\n    # Create an instance of the encoder\n    encoder = OneHotEncoder()\n\n    X_full = encoder.fit_transform(X_full_categorical).toarray()\n    X_submit = encoder.transform(X_submit_categorical).toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:53.903925Z","iopub.execute_input":"2023-10-26T10:17:53.905048Z","iopub.status.idle":"2023-10-26T10:17:53.920172Z","shell.execute_reply.started":"2023-10-26T10:17:53.905009Z","shell.execute_reply":"2023-10-26T10:17:53.918665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Backward difference contrast coding","metadata":{}},{"cell_type":"markdown","source":"### Demo example ","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Info on Backward difference contrast coding for encoding categorical variables. \n\nNO Tuneable params. \n\nDoes not depend on target \n\nWarning - returns dataframe even for one feature, (not just one column as for TEs)\n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nenc = ce.BackwardDifferenceEncoder( )\nd = enc.fit_transform(v)#, v_targets) \ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:53.922136Z","iopub.execute_input":"2023-10-26T10:17:53.922836Z","iopub.status.idle":"2023-10-26T10:17:54.454194Z","shell.execute_reply.started":"2023-10-26T10:17:53.922800Z","shell.execute_reply":"2023-10-26T10:17:54.453064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encoding BackwardDifferenceEncoder","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'BackwardDifferenceEncoder': # 'onehot'\n    encoder = ce.BackwardDifferenceEncoder( )\n    print(encoder)\n\n    X_full = encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.456153Z","iopub.execute_input":"2023-10-26T10:17:54.456496Z","iopub.status.idle":"2023-10-26T10:17:54.465501Z","shell.execute_reply.started":"2023-10-26T10:17:54.456466Z","shell.execute_reply":"2023-10-26T10:17:54.464141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Helmert contrast coding for encoding categorical features","metadata":{}},{"cell_type":"markdown","source":"### Demo example ","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Helmert contrast coding for encoding categorical features.\n\nNO Tuneable params. \n\nDoes not depend on target \n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nenc = ce.HelmertEncoder( )\nd = enc.fit_transform(v)#, v_targets) \ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.472199Z","iopub.execute_input":"2023-10-26T10:17:54.472695Z","iopub.status.idle":"2023-10-26T10:17:54.540048Z","shell.execute_reply.started":"2023-10-26T10:17:54.472658Z","shell.execute_reply":"2023-10-26T10:17:54.538556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encoding Helmert  ","metadata":{}},{"cell_type":"code","source":"%%time\n\nif  encoding == 'HelmertEncoder': # 'onehot'\n    encoder = ce.HelmertEncoder( )\n    print(encoder)\n\n    X_full = encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.542416Z","iopub.execute_input":"2023-10-26T10:17:54.542937Z","iopub.status.idle":"2023-10-26T10:17:54.552166Z","shell.execute_reply.started":"2023-10-26T10:17:54.542893Z","shell.execute_reply":"2023-10-26T10:17:54.551014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## ChemBert Encoding\n\nChemBERT embeddings for compounds created by ALEKSEY TREPETSKY: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550 (please upvote)\n\nDataset https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/code/alexandervc/look-on-chembert-data\n","metadata":{}},{"cell_type":"code","source":"%%time\nX_fulltrain_chembert_cls = np.load( '/kaggle/input/chemberta-v2-77-mtr/pad_false/train_ChemBERTa_v2_77MTR_cls_pad_False.npy')\nX_submit_chembert_cls = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/test_ChemBERTa_v2_77MTR_cls_pad_False.npy')\nX_fulltrain_chembert_mean = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/train_ChemBERTa_v2_77MTR_mean_pad_False.npy')\nX_submit_chembert_mean = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/test_ChemBERTa_v2_77MTR_mean_pad_False.npy')\nprint('chemBerta shapes:', X_fulltrain_chembert_cls.shape, X_submit_chembert_cls.shape, X_fulltrain_chembert_mean.shape, X_submit_chembert_mean.shape   ) \n\n# 'ChembertClsEncoding' 'ChembertMeanEncoding'\nif  encoding == 'ChembertClsEncoding': \n    print(encoding)\n\n    X_full = X_fulltrain_chembert_cls # encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = X_submit_chembert_cls # encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    \nif  encoding == 'ChembertMeanEncoding': \n    print(encoding)\n\n    X_full = X_fulltrain_chembert_mean # encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = X_submit_chembert_mean # encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.553216Z","iopub.execute_input":"2023-10-26T10:17:54.553581Z","iopub.status.idle":"2023-10-26T10:17:54.681169Z","shell.execute_reply.started":"2023-10-26T10:17:54.553533Z","shell.execute_reply":"2023-10-26T10:17:54.679455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## MorganFingerprints\n\n\nMorgan fingerprints, Molecula descriptors  by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/code/alexandervc/eda-morgan-fingerprint-features","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'MorganFingerprints': \n    fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_test.csv'\n    fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_train.csv'\n    df_train = pd.read_csv(fn_train)\n    df_test = pd.read_csv(fn_test)\n    \n    X_full = df_train.values\n    X_submit = df_test.values\n    \n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.683511Z","iopub.execute_input":"2023-10-26T10:17:54.684050Z","iopub.status.idle":"2023-10-26T10:17:54.692362Z","shell.execute_reply.started":"2023-10-26T10:17:54.684001Z","shell.execute_reply":"2023-10-26T10:17:54.691206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Molecular descriptors \n\nMorgan fingerprints, Molecula descriptors by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/alexandervc/eda-molecular-descriptors-features","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'MolecularDescriptors': \n    #fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_test.csv'\n    #fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_train.csv'\n    fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/molecula_descriptors/mol_descriptors_features_train.csv'\n    fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/molecula_descriptors/mol_descriptors_features_test.csv'\n    \n    df_train = pd.read_csv(fn_train)\n    df_test = pd.read_csv(fn_test)\n    \n    X_full = df_train.values\n    X_submit = df_test.values\n    \n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.693378Z","iopub.execute_input":"2023-10-26T10:17:54.693702Z","iopub.status.idle":"2023-10-26T10:17:54.716821Z","shell.execute_reply.started":"2023-10-26T10:17:54.693673Z","shell.execute_reply":"2023-10-26T10:17:54.715776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Micro EDA for encoded features","metadata":{}},{"cell_type":"code","source":"%%time\nprint('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \nprint(X_full[:3,:5])\nprint(X_submit[:3,:5])\n\nplt.figure(figsize = (12,4 ) )\nplt.hist(X_full.ravel() , bins = 100, label = 'train' )\nplt.hist(X_submit.ravel() , bins = 100 , label = 'test' )\nplt.title(encoding, fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.grid()\nplt.show()\nd1 = pd.Series( X_full.ravel() ).value_counts().head(5).to_frame().reset_index()\nd1.columns = ['value train', 'count train']\nd2 = pd.Series( X_submit.ravel() ).value_counts().head(5).to_frame().reset_index()\nd2.columns = ['value test', 'count test']\nd = pd.concat( [d1, d2], axis = 1)\nprint('Top frequent values')\ndisplay(d)\n\nv1 = np.mean(X_full, axis = 0 )\nv2 = np.mean(X_submit, axis = 0 )\nplt.figure(figsize = (12,4 ) )\nplt.plot(v1, label = 'train' )\nplt.plot(v2, label = 'test')\nplt.title(encoding + ' mean values', fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.xlabel( 'feature index', fontsize = 20 )\nplt.grid()\nplt.show()\n\nv1 = np.max(X_full, axis = 0 )\nv2 = np.max(X_submit, axis = 0 )\nplt.figure(figsize = (12,4 ) )\nplt.plot(v1, label = 'train' )\nplt.plot(v2, label = 'test')\nplt.title(encoding + ' max values', fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.xlabel( 'feature index', fontsize = 20 )\nplt.grid()\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:54.718717Z","iopub.execute_input":"2023-10-26T10:17:54.719631Z","iopub.status.idle":"2023-10-26T10:17:56.268928Z","shell.execute_reply.started":"2023-10-26T10:17:54.719578Z","shell.execute_reply":"2023-10-26T10:17:56.267391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling. Simple","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\nverbose = 1\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 1\nmodel = Ridge(alpha=alpha)\n\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme)\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\nlist_mrrmse = [];list_r2 = []\nfor i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n    X_train,X_test = X_full[IX_train], X_full[IX_test]\n    \n    Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n    if verbose >= 10:\n        print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n    \n    model.fit(X_train, Y_train_red)\n    Y_test_pred_red = model.predict( X_test )\n    Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n    \n    Y_test = Y_full[IX_test]\n    mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n    print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n    list_mrrmse.append(mrrmse); list_r2.append(r2)\nprint()\nprint('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ))\nprint('Average r2:', np.round(np.mean(list_r2 ),4 ))","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:17:56.270901Z","iopub.execute_input":"2023-10-26T10:17:56.271226Z","iopub.status.idle":"2023-10-26T10:18:03.931480Z","shell.execute_reply.started":"2023-10-26T10:17:56.271198Z","shell.execute_reply":"2023-10-26T10:18:03.929859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wrapper CV-computatation function","metadata":{"execution":{"iopub.status.busy":"2023-10-25T15:29:05.45711Z","iopub.execute_input":"2023-10-25T15:29:05.460841Z","iopub.status.idle":"2023-10-25T15:29:05.470101Z","shell.execute_reply.started":"2023-10-25T15:29:05.460777Z","shell.execute_reply":"2023-10-25T15:29:05.468659Z"}}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\ndef compute_cv( model, list_CV_scheme , verbose = 0 ):\n    '''\n    Compute cross validation(s) for model.\n    Returns dictionary with scores for each CV-scheme and fold\n    \n    dict_res = compute_cv( Ridge(alpha=1), ['MT','AmbrosM'] , verbose = 0 )\n    '''\n\n    dict_res = {}\n    \n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n    #alpha = 1\n    #model = Ridge(alpha=alpha)\n    for CV_scheme in list_CV_scheme:\n        dict_res[CV_scheme] = {}\n        #CV_scheme = 'MT'\n        if verbose >= 1:\n            print('CV_scheme:', CV_scheme)\n         \n        kf = KFold_custom(CV_scheme)# , valid_size = valid_size)  # valid_size = 0 \n\n        list_mrrmse = [];list_r2 = []\n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n            X_train,X_test = X_full[IX_train], X_full[IX_test]\n\n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n            model.fit(X_train, Y_train_red)\n            Y_test_pred_red = model.predict( X_test )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            if verbose >= 1:\n                print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4), CV_scheme )\n            list_mrrmse.append(mrrmse); list_r2.append(r2)\n        if verbose >= 1:\n            print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ), CV_scheme)\n            print('Average r2:', np.round(np.mean(list_r2 ),4 ), CV_scheme)\n            print()\n        dict_res[CV_scheme]['mrrmse'] = list_mrrmse\n        dict_res[CV_scheme]['r2'] = list_mrrmse\n    \n    \n    return dict_res\n\n\ndict_res = compute_cv( Ridge(alpha=1), ['MT','AmbrosM'] , verbose = 100 )\ndict_res","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:18:03.939782Z","iopub.execute_input":"2023-10-26T10:18:03.940952Z","iopub.status.idle":"2023-10-26T10:18:21.154694Z","shell.execute_reply.started":"2023-10-26T10:18:03.940863Z","shell.execute_reply":"2023-10-26T10:18:21.152439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compare several CV. Computation \n\nAbout 7-10 mins","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport warnings\nwarnings.filterwarnings(\"ignore\") \n\nverbose = 1\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 1\nmodel = Ridge(alpha=alpha)\n\n# CV_scheme = 'MT'\n# print('CV_scheme:', CV_scheme)\nvalid_size = 0 \ndf_stat = pd.DataFrame();\n\nlist_alpha = [0.0001,0.0002,0.0005,0.001, 0.002, 0.005, 0.01,0.02, 0.05,0.1,0.2,0.5,1,2,5,10,20,50 ,1e2,2e2,5e2,1e3,2e3,5e3,1e4,1e5,1e6]\n# list_alpha = [0.001,1,2] # 0.0001,0.0002,0.0005,0.001, 0.002, 0.005, 0.01,0.02, 0.05,0.1,0.2,0.5,1,2,5,10,20,50 ,1e2,2e2,5e2] # 1e3,1e4,1e5,1e6]:\nlist_CV_scheme =  ['AmbrosM','MT','Random']\n\nfor CV_scheme in  list_CV_scheme:\n    print();print('CV_scheme', CV_scheme)\n    list_mrrmse = []\n    IX_stat = 0;\n    for alpha in list_alpha:\n        if verbose > 0:\n            print('alpha:',alpha)\n        model = Ridge(alpha=alpha)\n        list_mrrmse_folds = []; list_r2_folds = []\n        kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n            \n            m_tmp1 = df_de_train['cell_type'] ==  'T cells CD8+' \n            l_tmp1 = list(      np.where(m_tmp1>0)[0] )\n            IX_train = [i for i in IX_train if i not in l_tmp1]\n            \n            X_train,X_test = X_full[IX_train], X_full[IX_test]\n\n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n            \n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n            model.fit(X_train, Y_train_red)\n            Y_test_pred_red = model.predict( X_test )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n            list_mrrmse_folds.append( mrrmse ); list_r2_folds.append(r2)\n\n            df_stat.loc[IX_stat,'alpha' ] = alpha\n            df_stat.loc[IX_stat,'mrrmse fold '+str(i_fold) +  ' ' + CV_scheme] = mrrmse\n            df_stat.loc[IX_stat,'r2 fold '+str(i_fold) + ' ' + CV_scheme] = r2\n        IX_stat += 1\n\n        list_mrrmse.append(np.mean( list_mrrmse_folds ))\n        print('Average:', 'mrrmse:', np.round(np.mean(list_mrrmse_folds),4), 'r2:',np.round(np.mean(list_r2_folds),4)  )\n        \n        \nwarnings.filterwarnings(\"default\")         \ndf_stat    ","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:18:21.156346Z","iopub.execute_input":"2023-10-26T10:18:21.157122Z","iopub.status.idle":"2023-10-26T10:30:28.636258Z","shell.execute_reply.started":"2023-10-26T10:18:21.157082Z","shell.execute_reply":"2023-10-26T10:30:28.635401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_stat.to_csv('df_stat.csv')\ndf_stat_save = df_stat.copy()\n\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nprint( df_stat.shape )\ndisplay( df_stat )\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:28.637717Z","iopub.execute_input":"2023-10-26T10:30:28.638504Z","iopub.status.idle":"2023-10-26T10:30:28.708901Z","shell.execute_reply.started":"2023-10-26T10:30:28.638469Z","shell.execute_reply":"2023-10-26T10:30:28.707398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plots  scores/alpha  ","metadata":{}},{"cell_type":"code","source":"df_stat = df_stat_save.copy()\nstr_model_inf = 'Ridge; '+ str_inf_cfg\nfor score_name in ['mrrmse','r2']:\n    print();print(); print('Score:', score_name ); print()\n    for CV_scheme in list_CV_scheme:\n        list_selected_cols =[ col for col in df_stat.columns if (score_name in col) and (CV_scheme in col) ]\n        col_average_name = score_name + ' ' + CV_scheme   + ' ' +  'average'\n        df_stat[col_average_name] = df_stat[ list_selected_cols ].mean(axis = 1)\n        list_selected_cols = [ col_average_name ] + list_selected_cols\n        plt.figure(figsize = (20,4))\n        \n        str_inf = 'Best alphas: '\n        for col in list_selected_cols:\n            v = df_stat[col]\n            plt.plot(np.log10(list_alpha), df_stat[col] ,'*-' , label = col)\n            #print('Best alpha:', df_stat.sort_values(col)['alpha'].iat[0], 'for ', col )\n            if score_name in ['mrrmse']:\n                str_score = str( np.round(df_stat.sort_values(col)[col].iat[0],4) )\n                str_inf += str(df_stat.sort_values(col)['alpha'].iat[0]) + ' for ' + col + ' ' + str_score +' ; '\n            else:\n                str_score = str( np.round(df_stat.sort_values(col,ascending = False)[col].iat[0],4) )\n                str_inf += str(df_stat.sort_values(col,ascending = False)['alpha'].iat[0]) + ' for ' + col  + ' ' + str_score + ' ; '\n                \n        print(str_inf)\n        plt.xlabel('LOG10 alpha',fontsize = 20 )\n        plt.legend(fontsize = 12 )\n        plt.title('scores: '  + score_name + ', CV: ' + CV_scheme + ', Model: ' + str_model_inf, fontsize = 20 )\n        plt.grid()\n        plt.show()\n\n\n# df_stat","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:28.710827Z","iopub.execute_input":"2023-10-26T10:30:28.711323Z","iopub.status.idle":"2023-10-26T10:30:31.449289Z","shell.execute_reply.started":"2023-10-26T10:30:28.711276Z","shell.execute_reply":"2023-10-26T10:30:31.447745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nprint( df_stat.shape )\nlist_cols1 = [t for t in df_stat.columns if ('average' in t) or ('alpha' == t) ]\nlist_cols2 = [t for t in df_stat.columns if ('average' not in t)  ]\ndf_stat = df_stat[list_cols1+list_cols2]\ndisplay( df_stat) # [list_cols1+list_cols2] )\ndf_stat.to_csv('df_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:31.451187Z","iopub.execute_input":"2023-10-26T10:30:31.451572Z","iopub.status.idle":"2023-10-26T10:30:31.539279Z","shell.execute_reply.started":"2023-10-26T10:30:31.451528Z","shell.execute_reply":"2023-10-26T10:30:31.537765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m = (df_stat['alpha'].iloc[:,0] == 0.1) | (df_stat['alpha'].iloc[:,0] == 0.2) | ( (df_stat['alpha'].iloc[:,0] == 1) )  | ( (df_stat['alpha'].iloc[:,0] == .5) )   | ( (df_stat['alpha'].iloc[:,0] == 1) )\ndf_stat[m].round(4)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:31.540793Z","iopub.execute_input":"2023-10-26T10:30:31.541143Z","iopub.status.idle":"2023-10-26T10:30:31.593200Z","shell.execute_reply.started":"2023-10-26T10:30:31.541115Z","shell.execute_reply":"2023-10-26T10:30:31.591678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', 100)\npd.set_option('display.max_rows', 50)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:31.595127Z","iopub.execute_input":"2023-10-26T10:30:31.595468Z","iopub.status.idle":"2023-10-26T10:30:31.601878Z","shell.execute_reply.started":"2023-10-26T10:30:31.595440Z","shell.execute_reply":"2023-10-26T10:30:31.600389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Brief final report ","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_brief_report = pd.DataFrame()\nlist_alpha = []\nfor col in ['mrrmse AmbrosM average', 'mrrmse MT average', 'mrrmse Random average',  ]:\n    alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0,0]\n    score_tmp = df_stat.sort_values(col)[col].iat[0]\n    list_alpha.append(alpha_tmp)\n    CV_scheme_loc = col.split(' ')[1]\n    str_score_inf = col.split(' ')[0]\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha opt '+ CV_scheme_loc    ] = alpha_tmp\n    \n    dict_res = compute_cv( Ridge(alpha=alpha_tmp), ['MT','AmbrosM','Random'] , verbose = 0 )\n    for CV2 in ['MT','AmbrosM','Random']:\n        df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha: ' + str(alpha_tmp) + ' opt for '+ CV_scheme_loc   ] = np.mean( dict_res[CV2][ str_score_inf ] )\n    \nalpha_average =    np.round( np.mean(list_alpha[:2]), 4)\ndf_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha average '   ] = alpha_average\ndict_res = compute_cv( Ridge(alpha=alpha_average), ['MT','AmbrosM','Random'] , verbose = 0 )\nfor CV2 in ['MT','AmbrosM','Random']:\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha(mean): ' + str(alpha_average)    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n\nalpha_selected = 1 #   \ndf_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha selected '   ] = alpha_selected\ndict_res = compute_cv( Ridge(alpha=alpha_selected), ['MT','AmbrosM','Random'] , verbose = 0 )\nfor CV2 in ['MT','AmbrosM','Random']:\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha(selected): ' + str(alpha_selected)    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n    \nprint( list_alpha     )\n\ndf_brief_report.round(5).to_csv('df_brief_report.csv')\ndf_brief_report.round(4)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:30:31.603954Z","iopub.execute_input":"2023-10-26T10:30:31.604407Z","iopub.status.idle":"2023-10-26T10:32:43.390524Z","shell.execute_reply.started":"2023-10-26T10:30:31.604377Z","shell.execute_reply":"2023-10-26T10:32:43.389392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Choose optimal alpha as average between optimial for selected CV-schemes","metadata":{}},{"cell_type":"code","source":"list_alpha = []\nfor col in ['mrrmse AmbrosM average', 'mrrmse MT average']:\n    alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0,0]\n    list_alpha.append(alpha_tmp)\n    \nalpha_average  = np.round( np.mean(list_alpha), 4)\nprint(alpha_average )\nprint( list_alpha )\n\ndict_res = compute_cv( Ridge(alpha=alpha_average), ['MT','AmbrosM','Random'] , verbose = 100 )\ndict_res","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:32:43.392228Z","iopub.execute_input":"2023-10-26T10:32:43.393667Z","iopub.status.idle":"2023-10-26T10:33:09.997450Z","shell.execute_reply.started":"2023-10-26T10:32:43.393610Z","shell.execute_reply":"2023-10-26T10:33:09.996366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submit \n\nWith alpha as average of optimal alphas for selected CV-schemes","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\nverbose = 1\n\nmask = df_de_train['cell_type'] != 'T cells CD8+'\n\nX_filtered = X_full[mask]\nY_filtered = Y_full[mask]\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nmodel = Ridge(alpha=alpha_average)\n\n\n\nY_filtered_red = reducer.fit_transform(Y_filtered)\nmodel.fit(X_filtered, Y_filtered_red)\n\nY_submit_pred_red = model.predict(X_submit)\nY_submit_pred = reducer.inverse_transform(Y_submit_pred_red)\n    \n    \ndf_submit = pd.DataFrame(Y_submit_pred, columns=df_de_train.columns[5:])\ndf_submit.index.name = 'id'\nprint(df_submit.shape)\ndisplay(df_submit)\nfn_save = 'submission_' + 'RidgeAlpha'+str(alpha_average).replace('.','')+'_'+str_inf_cfg.replace(' ','_')+'.csv' \nprint(fn_save)\ndf_submit.to_csv(fn_save)","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:33:09.999397Z","iopub.execute_input":"2023-10-26T10:33:10.000601Z","iopub.status.idle":"2023-10-26T10:33:26.036845Z","shell.execute_reply.started":"2023-10-26T10:33:10.000534Z","shell.execute_reply":"2023-10-26T10:33:26.035388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blend with \"priors\"\n\nThe trick typically uplifts the score about 0.01\n\nTrick to add \"priors\" (aggregates over the cell type and compound - which typically improves scores by 0.01 ): ( df_submit = (1-w_compound - w_cell_type)(df_submit) + w_compound df_submit_aggr_compound + w_cell_type * df_submit_aggr_cell_type )\n\nLIUDA CHELDIEVA: https://www.kaggle.com/code/liudacheldieva/streamlined-baseline-approach\n\nZXMKCD : https://www.kaggle.com/code/zmcxjt/streamlined-baseline-approach","metadata":{}},{"cell_type":"code","source":"%%time\n#group by drug and take the mean\ndf_tmp = df_de_train.iloc[:, [1] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"sm_name\"\ndf_aggr = df_tmp.groupby('sm_name').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_compound = pd.merge( df_id_map.reset_index(),  df_aggr, on='sm_name', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_compound.shape)\ndisplay(df_submit_aggr_compound.head(3))\n\nprint( )\n\ndf_tmp = df_de_train.iloc[:, [0] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"cell_type\"\ndf_aggr = df_tmp.groupby('cell_type').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_cell_type = pd.merge( df_id_map.reset_index(),  df_aggr, on='cell_type', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_cell_type.shape)\ndisplay(df_submit_aggr_cell_type.head(3))\n\n\nw_cell_type = 0.1\nw_compound = 0.45\ndf_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\n# fn_save2 = 'submission_' + 'RidgeAlpha'+str(alpha)+'_tsvd'+str(n_components)+'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \nfn_save2 = 'submission_' + 'RidgeAlpha'+str(alpha_average).replace('.','')+'_'+str_inf_cfg.replace(' ','_')+'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \nprint(fn_save2)\ndf_submit_2.to_csv(fn_save2)\nprint(df_submit_2.shape)\ndisplay(df_submit_2.head(3) )","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:33:26.038927Z","iopub.execute_input":"2023-10-26T10:33:26.039992Z","iopub.status.idle":"2023-10-26T10:33:40.668651Z","shell.execute_reply.started":"2023-10-26T10:33:26.039942Z","shell.execute_reply":"2023-10-26T10:33:40.667392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{"execution":{"iopub.status.busy":"2023-10-25T10:04:59.25663Z","iopub.execute_input":"2023-10-25T10:04:59.25698Z","iopub.status.idle":"2023-10-25T10:04:59.261681Z","shell.execute_reply.started":"2023-10-25T10:04:59.256955Z","shell.execute_reply":"2023-10-25T10:04:59.260221Z"}}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-10-26T10:33:40.674059Z","iopub.execute_input":"2023-10-26T10:33:40.674462Z","iopub.status.idle":"2023-10-26T10:33:40.681872Z","shell.execute_reply.started":"2023-10-26T10:33:40.674429Z","shell.execute_reply":"2023-10-26T10:33:40.680361Z"},"trusted":true},"execution_count":null,"outputs":[]}]}