{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":6541963,"sourceType":"competition"},{"sourceId":6567471,"sourceType":"datasetVersion","datasetId":3755626},{"sourceId":6774281,"sourceType":"datasetVersion","datasetId":3765688}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\n#### Brief\n\nWe use simple category encoders as well as some other features like: ChemBert, Morgan fingerprints, Molecula descriptors.\n\nCategory encoders considered:  'onehot', 'BackwardDifferenceEncoder', 'HelmertEncoder'  - these are contrast encoders - i.e. NOT using targets. \nTartget encoders planned to be considered in the other notebook. \n\n#### Conclusions\n\n    The simplest one-hot encoder works quite well and non of the encoders/features consitently outperform it in the present setup.\n    The Helmert encoder deserves attention - it is either the second or the first one - depending on CV-scheme\n\n#### More details \n\nFor each encoding we run the Ridge model with various alpha and several cross-validation schemes. Scores dependedce on regularization alpha is reported for each CV-scheme. \n\nOne can change n_components in tsvd as well as encoded features (compounds or cell-types, or both) - see section \"Key params\". \n\nDifferent versions of the notebook correspond to different encodings. Brief summary of the results is below. It is collected from the output files \"df_brief_report.csv\". More detailed scores collection - see https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing (sheet \"Encoders\"). \n\n    \n|  Notebook Version, Encoding       | Alpha | CV-MT  | CV-AmbrosM | CV-Random | Alpha | CV-MT  | CV-AmbrosM | CV-Random | Alpha | LB    | LB with priors |\n|------------------------|-------|--------|------------|-----------|-------|--------|------------|-----------|-------|-------|----------------|\n| V9 onehot               | 1.25  | 2.6668 | 0.99534    | 1.2363    | 1     | 2.66032| 0.99785    | 1.24013   | 1     | 0.613 | 0.605          |\n| V3 BackwardDifference   | 0.2   | 2.66246| 1.01891    | 1.27087   | 1     | 2.6989 | 1.03997    | 1.27831   |       |       |                |\n| V21 HelmertEncoder      | 1500  | 2.65864| 1.00846    | 1.26319   | 1     | 2.67429| 1.02866    | 1.29488   |       |       |                |\n| V5 ChembertClsEncoding  | 0.6   | 2.66555| 1.02795    | 1.30795   | 1     | 2.66446| 1.03244    | 1.31435   |       |       |                |\n| V6 ChembertMeanEncoding | 0.06  | 2.66481| 1.02863    | 1.30931   | 1     | 2.70268| 1.09262    | 1.37668   |       |       |                |\n| V7 MolecularDescriptors | 0.0003| 2.66736| 1.03241    | 1.33141   | 1     | 2.73042| 1.16143    | 1.44951   |       |       |                |\n| V8 MorganFingerprints   | 35    | 2.6609 | 1.01267    | 1.26405   | 1     | 2.67261| 1.02744    | 1.29396   |       |       |                |\n\n    \n    Versions above encode only compound (cell type is omitted - it works better for simple models like Ridge). \n    Here the first \"Alpha\" (Ridge regularization) is chosen as halfsum of optimal alphas for MT-scheme and AmbrosM scheme.\n    While the second \"Alpha\" is fixed to 1 for another comparaison. More detailed information can be found: \n    We copy these scores from the df_brief_report.csv file generated by the notebook. \n    \n    Encode Both CellType and Compound - apparantly even stronger regularization is required:\n    \n| Notebook Version, Encoding                       | Alpha   | CV-MT   | CV-AmbrosM | CV-Random | Alpha | CV-MT   | CV-AmbrosM | CV-Random | Alpha | LB      | LB with priors |\n|------------------------|---------|---------|------------|-----------|-------|---------|------------|-----------|-------|---------|----------------|\n| V22 onehot             | 500001  | 3.1558  | 1.10573    | 1.33653   | 1     | 2.88918 | 1.50913    | 1.28756   |       |         |                |\n| V25 BackwardDifference | 2500.25 | 3.15941 | 1.10502    | 1.33661   | 1     | 2.91915 | 1.48635    | 1.31733   |       |         |                |\n| V27 HelmertEncoder     | 1500    | 2.6107  | 0.9987     | 1.24782   | 1     | 2.9723  | 1.59741    | 1.35047   | 1500  | ?       | 0.605          |\n\n    Encode only CellType (results are very bad):\n\n| Notebook Version, Encoding                       | Alpha   | CV-MT   | CV-AmbrosM | CV-Random | Alpha | CV-MT   | CV-AmbrosM | CV-Random | Alpha | LB      | LB with priors |\n|------------------------|---------|---------|------------|-----------|-------|---------|------------|-----------|-------|---------|----------------|\n| V20 HelmertEncoder     | 501000  | 3.15559 | 1.10573    | 1.33643   | 1     | 3.62095 | 1.74043    | 1.36914   |       |         |                |\n| V24 BackwardDifference | 1250    | 3.15055 | 1.10546    | 1.33239   | 1     | 3.57319 | 1.67329    | 1.36381   |       |         |                |\n| V23 onehot             | 500050  | 3.1558  | 1.10574    | 1.33653   | 1     | 3.57055 | 1.7143     | 1.36555   |       |         |                |\n\n\n\nCollection of the CV/LB scores for other models can be found here: \nhttps://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing,  see Kaggle dataset:\nhttps://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc\n\n\n#### PS\n\nWe custom CV-class supporting Ambros, MT, Random, etc cross-validation schemes proposed here: \nhttps://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes . It supports sklearn interace.\n\nChemBert features prepared by ALEKSEY TREPETSKY: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550 (please upvote)\n\nMorgan fingerprints, Molecula descriptors  by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nNotebook is based on the previous notebooks: CV-schemes: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes, \nNN analysis : https://www.kaggle.com/code/alexandervc/op2-kishan-s-nn-streamlined-and-blended","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"encoding = 'onehot' # 'HelmertEncoder' # 'BackwardDifferenceEncoder' #     'HelmertEncoder' # 'ChembertClsEncoding' #'ChembertMeanEncoding' #'MolecularDescriptors'# 'MorganFingerprints' # \n\n\nn_components = 30 # Reduction of targets is done by tsvd to n_components, after predicting components we take tsvd.inverse_transform\n\nlist_features_to_encode = ['sm_name']# ['']#   [ 'sm_name'] # ['cell_type','sm_name'], ['cell_type'],  \n# For simple models like Ridge and that kind of features seems encoding only the compounds just forgetting about cell type at all gives sometimes better results. \n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:08.000253Z","iopub.execute_input":"2023-11-16T13:31:08.000624Z","iopub.status.idle":"2023-11-16T13:31:08.036400Z","shell.execute_reply.started":"2023-11-16T13:31:08.000596Z","shell.execute_reply":"2023-11-16T13:31:08.035126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create info string on params:\nstr_inf_cfg = encoding #+ ' '\nif 'sm_name' in list_features_to_encode: str_inf_cfg += 'Compound'\nif 'cell_type' in list_features_to_encode: str_inf_cfg += 'CellType'\nstr_inf_cfg += ' tsvd' + str(n_components)\nprint(str_inf_cfg)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:08.038909Z","iopub.execute_input":"2023-11-16T13:31:08.039651Z","iopub.status.idle":"2023-11-16T13:31:08.047561Z","shell.execute_reply.started":"2023-11-16T13:31:08.039605Z","shell.execute_reply":"2023-11-16T13:31:08.046162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preliminaries ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-11-16T13:31:08.049373Z","iopub.execute_input":"2023-11-16T13:31:08.050566Z","iopub.status.idle":"2023-11-16T13:31:20.296689Z","shell.execute_reply.started":"2023-11-16T13:31:08.050522Z","shell.execute_reply":"2023-11-16T13:31:20.295793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndisplay(df_de_train )\n\nplt.figure(figsize = (20,4) )\nv = df_de_train.iloc[:,5:].max(axis = 0 ).sort_values(ascending = False, key = abs )\nplt.plot(v.values,'*-')\nplt.title('Max-abs DE for genes',fontsize = 20 )\nplt.grid()\nplt.show()\ndisplay(v.head(15))\n\n# %%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn,index_col = 0)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf_sample_submit = pd.read_csv(fn, index_col = 0)\nprint(df_sample_submit.shape)\ndisplay( df_sample_submit )","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:20.299683Z","iopub.execute_input":"2023-11-16T13:31:20.300532Z","iopub.status.idle":"2023-11-16T13:31:27.556063Z","shell.execute_reply.started":"2023-11-16T13:31:20.300494Z","shell.execute_reply":"2023-11-16T13:31:27.555089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom CV-schemes\n\nSee the notebook: https://www.kaggle.com/code/alexandervc/op2-class-for-custom-cv-schemes#Ridge-CV-scoring.-Simple-example\n\nMore analysis of the CV schemes can be found in slides: https://docs.google.com/presentation/d/1wiz0Wmt4D54pqMMsIOyJHuQYMZ3hTBZQQnjbLzwoGYY/edit?usp=sharing and sheet: https://docs.google.com/spreadsheets/d/1APN63PMaWZygVjYimK9Ivt0RvifdAU5JRYkxiDn4szw/edit?usp=sharing\n\n\nCV schemes (see discussion https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/444494 ):\n\nAmbrosM: https://www.kaggle.com/code/ambrosm/scp-quickstart\n\nMT: https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy\n\nKishan's notebook: https://www.kaggle.com/code/liudacheldieva/neural-network-regression/notebook#Predicting-on-test-data ","metadata":{}},{"cell_type":"code","source":"# Prepare indexes for AmbrosM scheme:\n# For AmbrosM scheme: \nfolds_index_data_AmbrosM = [ ]\ntrain_sm_names = ['Idelalisib', 'Crizotinib', 'Linagliptin', 'Palbociclib', 'Dabrafenib', 'Alvocidib', 'LDN 193189', 'R428', 'Porcn Inhibitor III', \n  'Belinostat', 'Foretinib', 'MLN 2238', 'Penfluridol', 'Dactolisib', 'O-Demethylated Adapalene', 'Oprozomib (ONX 0912)', 'CHIR-99021']\nlist_fold_ids =  ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells'] \nfor fold_id in list_fold_ids:\n    mask_va = (df_de_train.cell_type == fold_id) & ~df_de_train.sm_name.isin(train_sm_names)\n    mask_tr = ~mask_va # 485 or 487 training rows\n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_AmbrosM.append( [IX_train, IX_test ])\n\n# For MT CV scheme:\nfolds_index_data_MT = []\nfold_to_compounds = {0: ['Alvocidib', 'Belinostat', 'Foretinib', 'LDN 193189',  'Linagliptin', 'O-Demethylated Adapalene'],\n 1: ['Dabrafenib', 'Dactolisib', 'Idelalisib', 'MLN 2238', 'Palbociclib', 'Porcn Inhibitor III'],\n 2: ['CHIR-99021', 'Crizotinib', 'Oprozomib (ONX 0912)', 'Penfluridol',  'R428']}\nfor fold_id in [0,1,2]:\n    mask_va = df_de_train['cell_type'].isin(['Myeloid cells', 'B cells']) & df_de_train['sm_name'].isin(fold_to_compounds[fold_id])\n    mask_tr = ~mask_va    \n    IX_train = np.where( mask_tr > 0 )[0]\n    IX_test = np.where( mask_va > 0 )[0]\n    #print(fold_id,  len(IX_test), type(IX_test), IX_test[:3], len(IX_train), type(IX_train), IX_train[:3]  )\n    folds_index_data_MT.append( [IX_train, IX_test ]) \n    \nimport random \nfrom sklearn.model_selection import KFold\n\nclass KFold_custom:\n    '''\n    Class with similar to sklearn \"KFold\" class, interface.\n    Support of CV schemes proposed by AmbrosM, MT, Kishan and standard sklearn KFold\n    Supports options to return IX_train,IX_valid,IX_test - triplet - if valid_size > 0\n    Examples:\n    kf = KFold_custom('AmbrosM'):     \n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n        pass\n    kf = KFold_custom('AmbrosM', valid_size = 0.1 )     \n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split()) :\n        pass\n        \n    kf = KFold_custom(CV_scheme = 'Random') # wrapper for sklearn Kfolds with shuffle = True\n    kf = KFold_custom(CV_scheme = CV_scheme, random_state = 42 ) # fixing random_state ensures reprodicibility of folds \n    '''\n    def __init__(self, CV_scheme,  valid_size = 0, n_splits=5, random_state = 42, verbose = 0 ):\n        self.CV_scheme = CV_scheme\n        self.CV_scheme_inf =  CV_scheme\n        self.valid_size = np.clip(valid_size,0,1)\n        self.random_state = random_state\n\n        self.folds_index_data = [ [np.arange(0), np.arange(0), np.arange(0)] ] # Example data\n        self.n_splits = 1 # Example data\n        if CV_scheme == 'Kishan1':\n            # One fold scheme which contains train, valid, test parts \n            IX_train_Kishan1 = np.arange(429)\n            IX_val_Kishan1 = np.arange(429,521)\n            IX_test_Kishan1 = np.arange(521,614)\n            self.n_splits = 1\n            self.folds_index_data =  [   [IX_train_Kishan1, IX_val_Kishan1, IX_test_Kishan1]  ]\n        elif CV_scheme == 'AmbrosM':\n            self.n_splits = 4\n            self.folds_index_data = folds_index_data_AmbrosM\n        elif CV_scheme == 'MT':\n            self.n_splits = 3\n            self.folds_index_data = folds_index_data_MT\n        elif 'Random'.lower() in CV_scheme.lower():\n            self.n_splits = n_splits\n            self.CV_scheme_inf = 'Random_'+str(n_splits) +'_'+str(random_state )\n            kf = KFold(n_splits=n_splits, random_state = random_state, shuffle=True )#,\n            self.folds_index_data = list( kf.split( np.arange(614) ) )\n            \n        else:\n            s = 'Uncrecognized CV_scheme ' + str(CV_scheme) \n            raise ValueError(s)\n\n        self.verbose = verbose \n        if verbose >= 10:\n            print(self.CV_scheme, self.n_splits, len(self.folds_index_data) )\n            #print(self.folds_index_data)\n            \n    def split(self, X=None):\n        '''\n        X - NOT used, just for compatibility with sklearn \n        '''\n        for item in self.folds_index_data:\n            if self.CV_scheme in ['Kishan1']:\n                if self.valid_size > 0:\n                    yield item[0],item[1],item[2]\n                else :\n                    yield np.array( list(set(item[0])|set(item[1]))  ) , item[2] # Return train, test only \n            elif self.CV_scheme in ['AmbrosM','MT', 'Random']:\n                if self.valid_size == 0:\n                    yield item[0],item[1]\n                else:\n                    # Split \"full-train\"->( real-train, valid ) \n                    index_for_real_train_part = int(  (1-self.valid_size) * len(item[0]) )\n                    if self.random_state is None:\n                        IX_train =  item[0][:index_for_real_train_part]\n                        IX_valid =  item[0][index_for_real_train_part:]\n                    elif self.random_state == -1:\n                        p = np.random.permutation(len(item[0])) \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    else:\n                        np.random.seed(self.random_state) # Set temporary random seed\n                        p = np.random.permutation(len(item[0])) \n                        np.random.seed(random.randint(0,30000)) # Randomize seed again - use Python random, not numpy       \n                        IX_train =  item[0][p][:index_for_real_train_part]\n                        IX_valid =  item[0][p][index_for_real_train_part:]\n                    yield IX_train, IX_valid, item[1]\n                    \n        \n    def get_list_main_CV_schemes(self):\n        return ['AmbrosM','MT','Kishan1','Random']\n    def get_n_splits(self, X=np.array([])):\n        return self.n_splits\n        \n    \nprint();print();    \nprint('--------------------------------Examples-----------------------------------------------')\nprint();print();    \n        \nCV_scheme = 'Kishan1'\nprint(CV_scheme,  )\nkf = KFold_custom('Kishan1', valid_size = 0.18)     \n#list( kf.split() )\nprint(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\nfor i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n    print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\nprint()\n\nfor CV_scheme in ['AmbrosM','MT','Random']:\n    valid_size = 0\n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size)     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train),type(IX_test)  \" )\n    for i_fold,(IX_train,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train), len(IX_test), IX_train[:3], IX_test[:3],  type(IX_train), type(IX_test) )\n    \n    print()\n    valid_size = 0.2 \n    print(CV_scheme, 'valid_size', valid_size )\n    kf = KFold_custom(CV_scheme, valid_size = valid_size )     \n    #list( kf.split() )\n    print(\" i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test)  \" )\n    for i_fold,(IX_train ,IX_valid,IX_test)   in enumerate( kf.split(None)) :\n        print(i_fold, len(IX_train),len(IX_valid),len(IX_test), IX_train[:3],IX_valid[:3], IX_test[:3],  type(IX_train),type(IX_valid),type(IX_test) )\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:27.557391Z","iopub.execute_input":"2023-11-16T13:31:27.557872Z","iopub.status.idle":"2023-11-16T13:31:28.012085Z","shell.execute_reply.started":"2023-11-16T13:31:27.557840Z","shell.execute_reply":"2023-11-16T13:31:28.010767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoding","metadata":{}},{"cell_type":"code","source":"%%time \nfeatures_columns = ['cell_type','sm_name']\nselected_features_columns =  list_features_to_encode #   [ 'sm_name']\nX_submit_categorical = df_id_map[selected_features_columns] # pd.DataFrame(df_id_map, columns= features_columns )\nX_full_categorical = df_de_train[ selected_features_columns ]\nY_full = df_de_train.iloc[:,5:].values\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.015457Z","iopub.execute_input":"2023-11-16T13:31:28.015909Z","iopub.status.idle":"2023-11-16T13:31:28.107715Z","shell.execute_reply.started":"2023-11-16T13:31:28.015872Z","shell.execute_reply":"2023-11-16T13:31:28.106500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One hot encoding","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.preprocessing import OneHotEncoder\n\nif encoding == 'onehot':\n\n    # Create an instance of the encoder\n    encoder = OneHotEncoder()\n\n    X_full = encoder.fit_transform(X_full_categorical).toarray()\n    X_submit = encoder.transform(X_submit_categorical).toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.109696Z","iopub.execute_input":"2023-11-16T13:31:28.110437Z","iopub.status.idle":"2023-11-16T13:31:28.130390Z","shell.execute_reply.started":"2023-11-16T13:31:28.110390Z","shell.execute_reply":"2023-11-16T13:31:28.129021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Backward difference contrast coding","metadata":{}},{"cell_type":"markdown","source":"### Demo example ","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Info on Backward difference contrast coding for encoding categorical variables. \n\nNO Tuneable params. \n\nDoes not depend on target \n\nWarning - returns dataframe even for one feature, (not just one column as for TEs)\n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nenc = ce.BackwardDifferenceEncoder( )\nd = enc.fit_transform(v)#, v_targets) \ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.131679Z","iopub.execute_input":"2023-11-16T13:31:28.132029Z","iopub.status.idle":"2023-11-16T13:31:28.640766Z","shell.execute_reply.started":"2023-11-16T13:31:28.131998Z","shell.execute_reply":"2023-11-16T13:31:28.639589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encoding BackwardDifferenceEncoder","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'BackwardDifferenceEncoder': # 'onehot'\n    encoder = ce.BackwardDifferenceEncoder( )\n    print(encoder)\n\n    X_full = encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.642358Z","iopub.execute_input":"2023-11-16T13:31:28.642706Z","iopub.status.idle":"2023-11-16T13:31:28.649791Z","shell.execute_reply.started":"2023-11-16T13:31:28.642676Z","shell.execute_reply":"2023-11-16T13:31:28.648621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Helmert contrast coding for encoding categorical features","metadata":{}},{"cell_type":"markdown","source":"### Demo example ","metadata":{}},{"cell_type":"code","source":"%%time\nimport category_encoders as ce\n\nprint(''' Helmert contrast coding for encoding categorical features.\n\nNO Tuneable params. \n\nDoes not depend on target \n''')\nv =  np.random.randint(0, 10, size=100)\nv_targets = np.random.randint(0, 10, size=100)\nv = [str(t) for t in v ]\nd = pd.DataFrame(); d.index.name = 'category'\nenc = ce.HelmertEncoder( )\nd = enc.fit_transform(v)#, v_targets) \ndisplay( d.head(10).round(3) )\nenc.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.654160Z","iopub.execute_input":"2023-11-16T13:31:28.654511Z","iopub.status.idle":"2023-11-16T13:31:28.710989Z","shell.execute_reply.started":"2023-11-16T13:31:28.654482Z","shell.execute_reply":"2023-11-16T13:31:28.710085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encoding Helmert  ","metadata":{}},{"cell_type":"code","source":"%%time\n\nif  encoding == 'HelmertEncoder': # 'onehot'\n    encoder = ce.HelmertEncoder( )\n    print(encoder)\n\n    X_full = encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.712459Z","iopub.execute_input":"2023-11-16T13:31:28.713228Z","iopub.status.idle":"2023-11-16T13:31:28.721500Z","shell.execute_reply.started":"2023-11-16T13:31:28.713192Z","shell.execute_reply":"2023-11-16T13:31:28.719483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## ChemBert Encoding\n\nChemBERT embeddings for compounds created by ALEKSEY TREPETSKY: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550 (please upvote)\n\nDataset https://www.kaggle.com/datasets/alekseytrepetsky/chemberta-v2-77-mtr (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/code/alexandervc/look-on-chembert-data\n","metadata":{}},{"cell_type":"code","source":"%%time\nX_fulltrain_chembert_cls = np.load( '/kaggle/input/chemberta-v2-77-mtr/pad_false/train_ChemBERTa_v2_77MTR_cls_pad_False.npy')\nX_submit_chembert_cls = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/test_ChemBERTa_v2_77MTR_cls_pad_False.npy')\nX_fulltrain_chembert_mean = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/train_ChemBERTa_v2_77MTR_mean_pad_False.npy')\nX_submit_chembert_mean = np.load('/kaggle/input/chemberta-v2-77-mtr/pad_false/test_ChemBERTa_v2_77MTR_mean_pad_False.npy')\nprint('chemBerta shapes:', X_fulltrain_chembert_cls.shape, X_submit_chembert_cls.shape, X_fulltrain_chembert_mean.shape, X_submit_chembert_mean.shape   ) \n\n# 'ChembertClsEncoding' 'ChembertMeanEncoding'\nif  encoding == 'ChembertClsEncoding': \n    print(encoding)\n\n    X_full = X_fulltrain_chembert_cls # encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = X_submit_chembert_cls # encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)\n    \nif  encoding == 'ChembertMeanEncoding': \n    print(encoding)\n\n    X_full = X_fulltrain_chembert_mean # encoder.fit_transform(X_full_categorical).values#.toarray()\n    X_submit = X_submit_chembert_mean # encoder.transform(X_submit_categorical).values#.toarray()\n\n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.722888Z","iopub.execute_input":"2023-11-16T13:31:28.723268Z","iopub.status.idle":"2023-11-16T13:31:28.814746Z","shell.execute_reply.started":"2023-11-16T13:31:28.723237Z","shell.execute_reply":"2023-11-16T13:31:28.813533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## MorganFingerprints\n\n\nMorgan fingerprints, Molecula descriptors  by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/code/alexandervc/eda-morgan-fingerprint-features","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'MorganFingerprints': \n    fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_test.csv'\n    fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_train.csv'\n    df_train = pd.read_csv(fn_train)\n    df_test = pd.read_csv(fn_test)\n    \n    X_full = df_train.values\n    X_submit = df_test.values\n    \n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.816782Z","iopub.execute_input":"2023-11-16T13:31:28.817949Z","iopub.status.idle":"2023-11-16T13:31:28.826358Z","shell.execute_reply.started":"2023-11-16T13:31:28.817901Z","shell.execute_reply":"2023-11-16T13:31:28.825119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Molecular descriptors \n\nMorgan fingerprints, Molecula descriptors by Antonina Dolgorukova https://www.kaggle.com/datasets/antoninadolgorukova/op2-supplementary-calcs-for-ml/ (please upvote)\n\nSome EDA can be found here: https://www.kaggle.com/alexandervc/eda-molecular-descriptors-features","metadata":{}},{"cell_type":"code","source":"%%time\n\nif encoding == 'MolecularDescriptors': \n    #fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_test.csv'\n    #fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/morgan_firgerptints_features/finger_features_train.csv'\n    fn_train = '/kaggle/input/op2-supplementary-calcs-for-ml/molecula_descriptors/mol_descriptors_features_train.csv'\n    fn_test = '/kaggle/input/op2-supplementary-calcs-for-ml/molecula_descriptors/mol_descriptors_features_test.csv'\n    \n    df_train = pd.read_csv(fn_train)\n    df_test = pd.read_csv(fn_test)\n    \n    X_full = df_train.values\n    X_submit = df_test.values\n    \n    print('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \n    print(X_full[:3,:5])\n    print(X_submit[:3,:5])    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.828214Z","iopub.execute_input":"2023-11-16T13:31:28.828971Z","iopub.status.idle":"2023-11-16T13:31:28.844600Z","shell.execute_reply.started":"2023-11-16T13:31:28.828922Z","shell.execute_reply":"2023-11-16T13:31:28.842939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Micro EDA for encoded features","metadata":{}},{"cell_type":"code","source":"%%time\nprint('X_full.shape, X_submit.shape ,  Y_full.shape', X_full.shape, X_submit.shape ,  Y_full.shape)  \nprint(X_full[:3,:5])\nprint(X_submit[:3,:5])\n\nplt.figure(figsize = (12,4 ) )\nplt.hist(X_full.ravel() , bins = 100, label = 'train' )\nplt.hist(X_submit.ravel() , bins = 100 , label = 'test' )\nplt.title(encoding, fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.grid()\nplt.show()\nd1 = pd.Series( X_full.ravel() ).value_counts().head(5).to_frame().reset_index()\nd1.columns = ['value train', 'count train']\nd2 = pd.Series( X_submit.ravel() ).value_counts().head(5).to_frame().reset_index()\nd2.columns = ['value test', 'count test']\nd = pd.concat( [d1, d2], axis = 1)\nprint('Top frequent values')\ndisplay(d)\n\nv1 = np.mean(X_full, axis = 0 )\nv2 = np.mean(X_submit, axis = 0 )\nplt.figure(figsize = (12,4 ) )\nplt.plot(v1, label = 'train' )\nplt.plot(v2, label = 'test')\nplt.title(encoding + ' mean values', fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.xlabel( 'feature index', fontsize = 20 )\nplt.grid()\nplt.show()\n\nv1 = np.max(X_full, axis = 0 )\nv2 = np.max(X_submit, axis = 0 )\nplt.figure(figsize = (12,4 ) )\nplt.plot(v1, label = 'train' )\nplt.plot(v2, label = 'test')\nplt.title(encoding + ' max values', fontsize = 20 )\nplt.legend(fontsize = 12)\nplt.xlabel( 'feature index', fontsize = 20 )\nplt.grid()\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:28.845966Z","iopub.execute_input":"2023-11-16T13:31:28.846351Z","iopub.status.idle":"2023-11-16T13:31:30.106553Z","shell.execute_reply.started":"2023-11-16T13:31:28.846315Z","shell.execute_reply":"2023-11-16T13:31:30.105432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling. Simple","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\nverbose = 1\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 1\nmodel = Ridge(alpha=alpha)\n\nCV_scheme = 'MT'\nprint('CV_scheme:', CV_scheme)\nvalid_size = 0 \nkf = KFold_custom(CV_scheme, valid_size = valid_size)   \n\nlist_mrrmse = [];list_r2 = []\nfor i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n    X_train,X_test = X_full[IX_train], X_full[IX_test]\n    \n    Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n    if verbose >= 10:\n        print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n    \n    model.fit(X_train, Y_train_red)\n    Y_test_pred_red = model.predict( X_test )\n    Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n    \n    Y_test = Y_full[IX_test]\n    mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n    print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n    list_mrrmse.append(mrrmse); list_r2.append(r2)\nprint()\nprint('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ))\nprint('Average r2:', np.round(np.mean(list_r2 ),4 ))","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:30.107862Z","iopub.execute_input":"2023-11-16T13:31:30.108785Z","iopub.status.idle":"2023-11-16T13:31:35.529898Z","shell.execute_reply.started":"2023-11-16T13:31:30.108751Z","shell.execute_reply":"2023-11-16T13:31:35.528426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wrapper CV-computatation function","metadata":{"execution":{"iopub.status.busy":"2023-10-25T15:29:05.457110Z","iopub.execute_input":"2023-10-25T15:29:05.460841Z","iopub.status.idle":"2023-10-25T15:29:05.470101Z","shell.execute_reply.started":"2023-10-25T15:29:05.460777Z","shell.execute_reply":"2023-10-25T15:29:05.468659Z"}}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\ndef compute_cv( model, list_CV_scheme , verbose = 0 ):\n    '''\n    Compute cross validation(s) for model.\n    Returns dictionary with scores for each CV-scheme and fold\n    \n    dict_res = compute_cv( Ridge(alpha=1), ['MT','AmbrosM'] , verbose = 0 )\n    '''\n\n    dict_res = {}\n    \n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n    #alpha = 1\n    #model = Ridge(alpha=alpha)\n    for CV_scheme in list_CV_scheme:\n        dict_res[CV_scheme] = {}\n        #CV_scheme = 'MT'\n        if verbose >= 1:\n            print('CV_scheme:', CV_scheme)\n         \n        kf = KFold_custom(CV_scheme)# , valid_size = valid_size)  # valid_size = 0 \n\n        list_mrrmse = [];list_r2 = []\n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n            X_train,X_test = X_full[IX_train], X_full[IX_test]\n\n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n            model.fit(X_train, Y_train_red)\n            Y_test_pred_red = model.predict( X_test )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            if verbose >= 1:\n                print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4), CV_scheme )\n            list_mrrmse.append(mrrmse); list_r2.append(r2)\n        if verbose >= 1:\n            print('Average mrrmse:', np.round(np.mean(list_mrrmse ),4 ), CV_scheme)\n            print('Average r2:', np.round(np.mean(list_r2 ),4 ), CV_scheme)\n            print()\n        dict_res[CV_scheme]['mrrmse'] = list_mrrmse\n        dict_res[CV_scheme]['r2'] = list_mrrmse\n    \n    \n    return dict_res\n\n\ndict_res = compute_cv( Ridge(alpha=1), ['MT','AmbrosM'] , verbose = 100 )\ndict_res","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:35.531989Z","iopub.execute_input":"2023-11-16T13:31:35.532492Z","iopub.status.idle":"2023-11-16T13:31:47.552680Z","shell.execute_reply.started":"2023-11-16T13:31:35.532447Z","shell.execute_reply":"2023-11-16T13:31:47.551436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compare several CV. Computation \n\nAbout 7-10 mins","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\nimport warnings\nwarnings.filterwarnings(\"ignore\") \n\nverbose = 1\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nalpha = 1\nmodel = Ridge(alpha=alpha)\n\n# CV_scheme = 'MT'\n# print('CV_scheme:', CV_scheme)\nvalid_size = 0 \ndf_stat = pd.DataFrame();\n\nlist_alpha = [0.0001,0.0002,0.0005,0.001, 0.002, 0.005, 0.01,0.02, 0.05,0.1,0.2,0.5,1,2,5,10,20,50 ,1e2,2e2,5e2,1e3,2e3,5e3,1e4,1e5,1e6]\n# list_alpha = [0.001,1,2] # 0.0001,0.0002,0.0005,0.001, 0.002, 0.005, 0.01,0.02, 0.05,0.1,0.2,0.5,1,2,5,10,20,50 ,1e2,2e2,5e2] # 1e3,1e4,1e5,1e6]:\nlist_CV_scheme =  ['AmbrosM','MT','Random']\n\nfor CV_scheme in  list_CV_scheme:\n    print();print('CV_scheme', CV_scheme)\n    list_mrrmse = []\n    IX_stat = 0;\n    for alpha in list_alpha:\n        if verbose > 0:\n            print('alpha:',alpha)\n        model = Ridge(alpha=alpha)\n        list_mrrmse_folds = []; list_r2_folds = []\n        kf = KFold_custom(CV_scheme, valid_size = valid_size)   \n        for i_fold,(IX_train,IX_test)   in enumerate( kf.split()) :\n            X_train,X_test = X_full[IX_train], X_full[IX_test]\n\n            Y_train_red = reducer.fit_transform(Y_full[IX_train] ) \n            if verbose >= 10:\n                print('fold:', i_fold, 'X_train.shape, X_test.shape, Y_train_red.shape', X_train.shape, X_test.shape, Y_train_red.shape)\n\n            model.fit(X_train, Y_train_red)\n            Y_test_pred_red = model.predict( X_test )\n            Y_test_pred = reducer.inverse_transform( Y_test_pred_red )\n\n            Y_test = Y_full[IX_test]\n            mrrmse = np.sqrt(np.square(Y_test - Y_test_pred).mean(axis=1)).mean();  r2 = r2_score( Y_test , Y_test_pred ) \n            print('fold:', i_fold, 'mrrmse:', np.round(mrrmse,4), 'r2:', np.round(r2,4),  )\n            list_mrrmse_folds.append( mrrmse ); list_r2_folds.append(r2)\n\n            df_stat.loc[IX_stat,'alpha' ] = alpha\n            df_stat.loc[IX_stat,'mrrmse fold '+str(i_fold) +  ' ' + CV_scheme] = mrrmse\n            df_stat.loc[IX_stat,'r2 fold '+str(i_fold) + ' ' + CV_scheme] = r2\n        IX_stat += 1\n\n        list_mrrmse.append(np.mean( list_mrrmse_folds ))\n        print('Average:', 'mrrmse:', np.round(np.mean(list_mrrmse_folds),4), 'r2:',np.round(np.mean(list_r2_folds),4)  )\n        \n        \nwarnings.filterwarnings(\"default\")         \ndf_stat    ","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:31:47.558626Z","iopub.execute_input":"2023-11-16T13:31:47.561754Z","iopub.status.idle":"2023-11-16T13:40:43.422955Z","shell.execute_reply.started":"2023-11-16T13:31:47.561695Z","shell.execute_reply":"2023-11-16T13:40:43.421004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_stat.to_csv('df_stat.csv')\ndf_stat_save = df_stat.copy()\n\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nprint( df_stat.shape )\ndisplay( df_stat )\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:43.424630Z","iopub.execute_input":"2023-11-16T13:40:43.424992Z","iopub.status.idle":"2023-11-16T13:40:43.478129Z","shell.execute_reply.started":"2023-11-16T13:40:43.424958Z","shell.execute_reply":"2023-11-16T13:40:43.476987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plots  scores/alpha  ","metadata":{}},{"cell_type":"code","source":"df_stat = df_stat_save.copy()\nstr_model_inf = 'Ridge; '+ str_inf_cfg\nfor score_name in ['mrrmse','r2']:\n    print();print(); print('Score:', score_name ); print()\n    for CV_scheme in list_CV_scheme:\n        list_selected_cols =[ col for col in df_stat.columns if (score_name in col) and (CV_scheme in col) ]\n        col_average_name = score_name + ' ' + CV_scheme   + ' ' +  'average'\n        df_stat[col_average_name] = df_stat[ list_selected_cols ].mean(axis = 1)\n        list_selected_cols = [ col_average_name ] + list_selected_cols\n        plt.figure(figsize = (20,4))\n        \n        str_inf = 'Best alphas: '\n        for col in list_selected_cols:\n            v = df_stat[col]\n            plt.plot(np.log10(list_alpha), df_stat[col] ,'*-' , label = col)\n            #print('Best alpha:', df_stat.sort_values(col)['alpha'].iat[0], 'for ', col )\n            if score_name in ['mrrmse']:\n                str_score = str( np.round(df_stat.sort_values(col)[col].iat[0],4) )\n                str_inf += str(df_stat.sort_values(col)['alpha'].iat[0]) + ' for ' + col + ' ' + str_score +' ; '\n            else:\n                str_score = str( np.round(df_stat.sort_values(col,ascending = False)[col].iat[0],4) )\n                str_inf += str(df_stat.sort_values(col,ascending = False)['alpha'].iat[0]) + ' for ' + col  + ' ' + str_score + ' ; '\n                \n        print(str_inf)\n        plt.xlabel('LOG10 alpha',fontsize = 20 )\n        plt.legend(fontsize = 12 )\n        plt.title('scores: '  + score_name + ', CV: ' + CV_scheme + ', Model: ' + str_model_inf, fontsize = 20 )\n        plt.grid()\n        plt.show()\n\n\n# df_stat","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:43.479683Z","iopub.execute_input":"2023-11-16T13:40:43.480022Z","iopub.status.idle":"2023-11-16T13:40:45.682598Z","shell.execute_reply.started":"2023-11-16T13:40:43.479993Z","shell.execute_reply":"2023-11-16T13:40:45.681337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\n\nprint( df_stat.shape )\nlist_cols1 = [t for t in df_stat.columns if ('average' in t) or ('alpha' == t) ]\nlist_cols2 = [t for t in df_stat.columns if ('average' not in t)  ]\ndf_stat = df_stat[list_cols1+list_cols2]\ndisplay( df_stat) # [list_cols1+list_cols2] )\ndf_stat.to_csv('df_stat.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:45.684569Z","iopub.execute_input":"2023-11-16T13:40:45.684991Z","iopub.status.idle":"2023-11-16T13:40:45.752465Z","shell.execute_reply.started":"2023-11-16T13:40:45.684951Z","shell.execute_reply":"2023-11-16T13:40:45.751405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m = (df_stat['alpha'].iloc[:,0] == 0.1) | (df_stat['alpha'].iloc[:,0] == 0.2) | ( (df_stat['alpha'].iloc[:,0] == 1) )  | ( (df_stat['alpha'].iloc[:,0] == .5) )   | ( (df_stat['alpha'].iloc[:,0] == 1) )\ndf_stat[m].round(4)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:45.753852Z","iopub.execute_input":"2023-11-16T13:40:45.754285Z","iopub.status.idle":"2023-11-16T13:40:45.798677Z","shell.execute_reply.started":"2023-11-16T13:40:45.754245Z","shell.execute_reply":"2023-11-16T13:40:45.797460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', 100)\npd.set_option('display.max_rows', 50)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:45.800349Z","iopub.execute_input":"2023-11-16T13:40:45.801129Z","iopub.status.idle":"2023-11-16T13:40:45.806265Z","shell.execute_reply.started":"2023-11-16T13:40:45.801083Z","shell.execute_reply":"2023-11-16T13:40:45.805126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Brief final report ","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_brief_report = pd.DataFrame()\nlist_alpha = []\nfor col in ['mrrmse AmbrosM average', 'mrrmse MT average', 'mrrmse Random average',  ]:\n    alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0,0]\n    score_tmp = df_stat.sort_values(col)[col].iat[0]\n    list_alpha.append(alpha_tmp)\n    CV_scheme_loc = col.split(' ')[1]\n    str_score_inf = col.split(' ')[0]\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha opt '+ CV_scheme_loc    ] = alpha_tmp\n    \n    dict_res = compute_cv( Ridge(alpha=alpha_tmp), ['MT','AmbrosM','Random'] , verbose = 0 )\n    for CV2 in ['MT','AmbrosM','Random']:\n        df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha: ' + str(alpha_tmp) + ' opt for '+ CV_scheme_loc   ] = np.mean( dict_res[CV2][ str_score_inf ] )\n    \nalpha_average =    np.round( np.mean(list_alpha[:2]), 4)\ndf_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha average '   ] = alpha_average\ndict_res = compute_cv( Ridge(alpha=alpha_average), ['MT','AmbrosM','Random'] , verbose = 0 )\nfor CV2 in ['MT','AmbrosM','Random']:\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha(mean): ' + str(alpha_average)    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n\nalpha_selected = 1 #   \ndf_brief_report.loc[ 'Ridge '+str_inf_cfg   , 'alpha selected '   ] = alpha_selected\ndict_res = compute_cv( Ridge(alpha=alpha_selected), ['MT','AmbrosM','Random'] , verbose = 0 )\nfor CV2 in ['MT','AmbrosM','Random']:\n    df_brief_report.loc[ 'Ridge '+str_inf_cfg   ,  str_score_inf  +  CV2 + ' alpha(selected): ' + str(alpha_selected)    ] = np.mean( dict_res[CV2][ str_score_inf ] )\n    \nprint( list_alpha     )\n\ndf_brief_report.round(5).to_csv('df_brief_report.csv')\ndf_brief_report.round(4)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:40:45.807846Z","iopub.execute_input":"2023-11-16T13:40:45.808216Z","iopub.status.idle":"2023-11-16T13:42:22.439222Z","shell.execute_reply.started":"2023-11-16T13:40:45.808186Z","shell.execute_reply":"2023-11-16T13:42:22.438326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Choose optimal alpha as average between optimial for selected CV-schemes","metadata":{}},{"cell_type":"code","source":"list_alpha = []\nfor col in ['mrrmse AmbrosM average', 'mrrmse MT average']:\n    alpha_tmp = df_stat.sort_values(col)['alpha'].iloc[0,0]\n    list_alpha.append(alpha_tmp)\n    \nalpha_average  = np.round( np.mean(list_alpha), 4)\nprint(alpha_average )\nprint( list_alpha )\n\ndict_res = compute_cv( Ridge(alpha=alpha_average), ['MT','AmbrosM','Random'] , verbose = 100 )\ndict_res","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:42:22.441153Z","iopub.execute_input":"2023-11-16T13:42:22.441970Z","iopub.status.idle":"2023-11-16T13:42:42.183546Z","shell.execute_reply.started":"2023-11-16T13:42:22.441925Z","shell.execute_reply":"2023-11-16T13:42:42.182355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submit \n\nWith alpha as average of optimal alphas for selected CV-schemes","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.linear_model import Ridge\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.metrics import r2_score\n\nverbose = 1\n\nreducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\n\nmodel = Ridge(alpha=alpha_average)\n\nY_full_red = reducer.fit_transform(Y_full ) \nmodel.fit(X_full, Y_full_red)\nY_submit_pred_red = model.predict( X_submit )\nY_submit_pred = reducer.inverse_transform( Y_submit_pred_red )\n    \n    \ndf_submit = pd.DataFrame(Y_submit_pred, columns = df_de_train.columns[5:])\ndf_submit.index.name = 'id'\nprint( df_submit.shape )\ndisplay(df_submit)\nfn_save = 'submission_' + 'RidgeAlpha'+str(alpha_average).replace('.','')+'_'+str_inf_cfg.replace(' ','_')+'.csv' \nprint(fn_save)\ndf_submit.to_csv(fn_save)","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:42:42.185332Z","iopub.execute_input":"2023-11-16T13:42:42.186389Z","iopub.status.idle":"2023-11-16T13:42:53.234138Z","shell.execute_reply.started":"2023-11-16T13:42:42.186340Z","shell.execute_reply":"2023-11-16T13:42:53.232885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blend with \"priors\"\n\nThe trick typically uplifts the score about 0.01\n\nTrick to add \"priors\" (aggregates over the cell type and compound - which typically improves scores by 0.01 ): ( df_submit = (1-w_compound - w_cell_type)(df_submit) + w_compound df_submit_aggr_compound + w_cell_type * df_submit_aggr_cell_type )\n\nLIUDA CHELDIEVA: https://www.kaggle.com/code/liudacheldieva/streamlined-baseline-approach\n\nZXMKCD : https://www.kaggle.com/code/zmcxjt/streamlined-baseline-approach","metadata":{}},{"cell_type":"code","source":"%%time\n#group by drug and take the mean\ndf_tmp = df_de_train.iloc[:, [1] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"sm_name\"\ndf_aggr = df_tmp.groupby('sm_name').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_compound = pd.merge( df_id_map.reset_index(),  df_aggr, on='sm_name', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_compound.shape)\ndisplay(df_submit_aggr_compound.head(3))\n\nprint( )\n\ndf_tmp = df_de_train.iloc[:, [0] + list(range(5, df_de_train.shape[1]))] # Take only numeric columns and \"cell_type\"\ndf_aggr = df_tmp.groupby('cell_type').mean().reset_index()\nprint(df_aggr.shape)\ndf_submit_aggr_cell_type = pd.merge( df_id_map.reset_index(),  df_aggr, on='cell_type', how = 'left' ).sort_values('id').drop(columns = ['cell_type', 'sm_name']).set_index('id')\nprint(df_submit_aggr_cell_type.shape)\ndisplay(df_submit_aggr_cell_type.head(3))\n\n\nw_cell_type = 0.1\nw_compound = 0.45\ndf_submit_2 = (1-w_compound - w_cell_type)*(df_submit) + w_compound * df_submit_aggr_compound +  w_cell_type * df_submit_aggr_cell_type\n# fn_save2 = 'submission_' + 'RidgeAlpha'+str(alpha)+'_tsvd'+str(n_components)+'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \nfn_save2 = 'submission_' + 'RidgeAlpha'+str(alpha_average).replace('.','')+'_'+str_inf_cfg.replace(' ','_')+'_blendWithPriors_'+str(w_compound)+'_'+str( w_cell_type )+'.csv' \nprint(fn_save2)\ndf_submit_2.to_csv(fn_save2)\nprint(df_submit_2.shape)\ndisplay(df_submit_2.head(3) )","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:42:53.235623Z","iopub.execute_input":"2023-11-16T13:42:53.236003Z","iopub.status.idle":"2023-11-16T13:43:03.761110Z","shell.execute_reply.started":"2023-11-16T13:42:53.235971Z","shell.execute_reply":"2023-11-16T13:43:03.760309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final timing","metadata":{"execution":{"iopub.status.busy":"2023-10-25T10:04:59.25663Z","iopub.execute_input":"2023-10-25T10:04:59.25698Z","iopub.status.idle":"2023-10-25T10:04:59.261681Z","shell.execute_reply.started":"2023-10-25T10:04:59.256955Z","shell.execute_reply":"2023-10-25T10:04:59.260221Z"}}},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-11-16T13:43:03.765733Z","iopub.execute_input":"2023-11-16T13:43:03.766367Z","iopub.status.idle":"2023-11-16T13:43:03.771703Z","shell.execute_reply.started":"2023-11-16T13:43:03.766330Z","shell.execute_reply":"2023-11-16T13:43:03.770681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}