{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# DRAFT UNDER CONSTRUCTION\n\n# What is about ?   (planned to be)\n\nThe first brief look on the \"Open Problems – Single-Cell Perturbations\" data.\n\nAnd some baselines\n\n\n###  Notes \n\n- number of train samples - 614, test 255 - not much ! ML for such low number of samples - hmmm.... very unclearly how successful it can be ...   \n\n- 18 211 targets (genes)  \n\n- in some sense we have ONLY TWO features - cell type and compound. Both of them are categorical. Cell types - 6, compounds - 146 categories. Thus it is natural to use target encodings. Alternative way - one might think in direction of recsys approach, but with the modification that we have 18211 targets. \n\n\n\n    V32 Catboost depth = 2\n    V30 LB 0.905 ( terrible :)  - Catboost first draft\n\n    Versions 23-27 - first modeling approach - target encode cell type and compound and simple Ridge model - results currently worse than naive approach used before - without models. \n        V27 LB 0.659 LB Ridge nCT1 nCD25 Al10 TSVD35 - even more relaxed : alpha and encoded_compound size \n        V26 LB 0.668 try to relax a model a bit: Ridge nCT1 nCD10 Al100 TSVD35 - results are better near mean/zero submission\n        V25 LB 0.677 - stronger constraining the model: Ridge nCT1 nCD5 Al100 TSVD35, but still we are worse than even predict by mean \n        V24 LB 0.702 - try avoid overfit Ridge nCT3 nCD10 Al100 TSVD35 - better results but still bad \n        V23 LB 0.747 Ridge model in target encoded features Ridge nCT10 nCD35 Al1 TSVD35 - modeling gives worse results than simple approach, might be bug or overfit\n    \n    V22 - cell cycle genes brief analysis \n    V19,20,21 - clustering 10000, 15000, 18211(all) genes by sns.clustermap - 15min,47min,RAMcrash - see two clear clusters in genes - what are they ? \n    V18 LB 0.623 TSVD-35 denoising  quantile 0.54, tsvd is better than pca/ica - similar to previous OP\n    V15 - bug -  LB 0.626 - NO it was not: TSVD-35 denoising  quantile 0.54 - so worse than pca,ICA, at least for these params\n    V14 LB 0.624 - ICA-35 denoising   quantile 0.54 - so ICA is not better than just pca at least for these the same params\n    Fork: LB 0.624 - pca-35 denoising, quantile 0.54 \n    V13 LB 0.626 pca25 denoising \n    V12 LB 0.627 same with pca100 denoising\n            That means: first reduce data to pca100, and  take pca inverse transfrom (denoising)\n            and then apply same as in V9: groupby by compound and quantile(0.6)         \n            so idea is - pca100 reductions hopefully kills some noise\n            In the previous challenge it worked okay, but the number of samples was not 614, but near 100 000\n    V11 - bug - pay attention that column names in submit file should correspond to genes\n    V9 LB 0.638 groupby by compound and quantile(0.6)\n    V6 LB 0.666 quantile(0.7)\n    V5 LB 0.657 quantile(0.6) \n            results are again better, that probably indicates some shift between train and public data \n    V4 LB 0.659 - median instead of mean - results are a bit better, \n            it might mean either a bit of presense of outliers, or  public is somewhat different from train - next experiments suggests second is true \n    V1 LB 0.664 - submission of train means - the simplest baseline ","metadata":{}},{"cell_type":"markdown","source":"# Preliminaries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-09-16T17:32:18.941638Z","iopub.execute_input":"2023-09-16T17:32:18.942100Z","iopub.status.idle":"2023-09-16T17:32:18.957779Z","shell.execute_reply.started":"2023-09-16T17:32:18.942058Z","shell.execute_reply":"2023-09-16T17:32:18.956105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndf_de_train","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:18.960753Z","iopub.execute_input":"2023-09-16T17:32:18.961234Z","iopub.status.idle":"2023-09-16T17:32:20.789425Z","shell.execute_reply.started":"2023-09-16T17:32:18.961200Z","shell.execute_reply":"2023-09-16T17:32:20.788078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_de_train.iloc[:,5:].head(1)","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:20.791801Z","iopub.execute_input":"2023-09-16T17:32:20.792185Z","iopub.status.idle":"2023-09-16T17:32:20.868655Z","shell.execute_reply.started":"2023-09-16T17:32:20.792157Z","shell.execute_reply":"2023-09-16T17:32:20.867318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dimensional reductions (pca, umap,...), visualizations, clustering train target data ","metadata":{}},{"cell_type":"code","source":"X = df_de_train.iloc[:,5:]\nprint(X.shape)","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:20.870517Z","iopub.execute_input":"2023-09-16T17:32:20.870907Z","iopub.status.idle":"2023-09-16T17:32:20.946547Z","shell.execute_reply.started":"2023-09-16T17:32:20.870874Z","shell.execute_reply":"2023-09-16T17:32:20.944664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.decomposition import PCA\n\nv1_color = df_de_train[  'cell_type']\nv2_color = df_de_train[  'sm_name'].copy()\nv3_color = df_de_train[  'sm_name'].copy()\nl = [t for t in df_de_train[  'sm_name'] if t.endswith('nib') ]\nm = v2_color.isin( l)\nv2_color[~m] = 'non -nib'\nv3_color[m] = '*nib'\n\nv4_color = df_de_train[  'control']#.copy()\n\nlist_top_drugs = ['MLN 2238', 'Resminostat', 'CEP-18770 (Delanzomib)', 'Oprozomib (ONX 0912)', 'Belinostat', 'Vorinostat', 'Ganetespib (STA-9090)', 'Scriptaid', 'Proscillaridin A;Proscillaridin-A', 'Alvocidib', 'IN1451']\nm = df_de_train[  'sm_name'].isin( list_top_drugs)\nv5_color = df_de_train[  'sm_name'].copy()\nv5_color[~m] = np.nan\n\n\n\nlist_cfg = [ ['cell type',v1_color], ['control' , v4_color ] , ['top compounds',v5_color] ]\n#     str_inf1 = ''\n    #X = np.clip(df.iloc[N0:N1,33:137].fillna(0),0, 1)\nstr_inf = 'PCA' \nreducer = PCA(n_components=100 )\nXr = reducer.fit_transform(X)\nfor i,j in [[0,1],[0,2],[1,2],[3,4],[5,6],[7,8]]:\n    plt.figure(figsize = (20,10)); ic=0\n    for str_inf1, v_for_color in list_cfg: # , ['*nib compounds ',v2_color], ['non *nib compounds',v3_color ] ]:\n        ic+=1; plt.subplot(1,len(list_cfg),ic)\n        sns.scatterplot(x= Xr[:,i], y = Xr[:,j], hue =  v_for_color ,s = 100) # df['reads'])\n        plt.xlabel(str_inf+str(i+1), fontsize = 20)\n        plt.ylabel(str_inf+str(j+1), fontsize = 20)\n        plt.title(str_inf1 + ' ', fontsize = 20 )\n\n    plt.show()\n    \n\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:20.950137Z","iopub.execute_input":"2023-09-16T17:32:20.950563Z","iopub.status.idle":"2023-09-16T17:32:33.443934Z","shell.execute_reply.started":"2023-09-16T17:32:20.950531Z","shell.execute_reply":"2023-09-16T17:32:33.442398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d = df_de_train.iloc[:,:5]\nd['PCA1'] = Xr[:,0]\nd['PCA2'] = Xr[:,1]\nd['PCA3'] = Xr[:,2]\nlist_top_drugs = []\ndisplay( d.sort_values('PCA1', ascending = False ).head(8) )\nlist_top_drugs += d.sort_values('PCA1', ascending = False ).head(8)['sm_name'].to_list()\nprint(list_top_drugs)\ndisplay( d.sort_values('PCA2', ascending = False ).head(8) )\nlist_top_drugs += d.sort_values('PCA2', ascending = False ).head(8)['sm_name'].to_list()\ndisplay( d.sort_values('PCA3', ascending = False ).head(8) )\nlist_top_drugs += d.sort_values('PCA3', ascending = False ).head(8)['sm_name'].to_list()\nprint(list(set(list_top_drugs)))","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:33.446257Z","iopub.execute_input":"2023-09-16T17:32:33.447404Z","iopub.status.idle":"2023-09-16T17:32:33.512576Z","shell.execute_reply.started":"2023-09-16T17:32:33.447353Z","shell.execute_reply":"2023-09-16T17:32:33.511088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Clustering  Cell Types","metadata":{}},{"cell_type":"code","source":"%%time\nN = df_de_train.shape[1]# 5000\nprint(N)\nX = df_de_train[ ['cell_type'] + list(df_de_train.columns[5:N]) ].groupby('cell_type').median()\nprint(X.shape)\ncm = np.corrcoef(X)\nprint(cm[:3,:2])\ncm = np.abs(cm)\nl = list(X.index)# [df_de_train['sm_name'].iat[i] +' '+ df_de_train['cell_type'].iat[i]  for i in range(len(df_de_train))] # .columns[5:N]\ncm = pd.DataFrame(cm, index =l , columns = l )\nprint(cm.shape)\nsns.clustermap(cm,  annot=True, fmt=\".2f\", cmap=\"coolwarm\" )\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:33.515349Z","iopub.execute_input":"2023-09-16T17:32:33.516194Z","iopub.status.idle":"2023-09-16T17:32:34.575262Z","shell.execute_reply.started":"2023-09-16T17:32:33.516143Z","shell.execute_reply":"2023-09-16T17:32:34.573361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Clustering compounds","metadata":{}},{"cell_type":"code","source":"%%time\nN = df_de_train.shape[1]# 5000\nprint(N)\nX = df_de_train[ ['sm_name'] + list(df_de_train.columns[5:N]) ].groupby('sm_name').median()\nprint(X.shape)\ncm = np.corrcoef(X)\nprint(cm[:3,:2])\ncm = np.abs(cm)\nl = list(X.index)# [df_de_train['sm_name'].iat[i] +' '+ df_de_train['cell_type'].iat[i]  for i in range(len(df_de_train))] # .columns[5:N]\nl = [t[:20] for t in l] # cut long names\ncm = pd.DataFrame(cm, index =l , columns = l )\nprint(cm.shape)\nclustergrid = sns.clustermap(cm,cmap=\"coolwarm\" )# ,  annot=True, fmt=\".2f\", \nplt.show()\nreordered_columns = clustergrid.dendrogram_col.reordered_ind\nreordered_rows = clustergrid.dendrogram_row.reordered_ind\nprint(len(reordered_rows), len(reordered_columns) )\nprint( list(cm.index[reordered_rows]) )\n# print( list(X.columns[reordered_columns]) )\n\nsns.clustermap(cm,  annot=True, fmt=\".2f\", cmap=\"coolwarm\" )\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:32:34.577678Z","iopub.execute_input":"2023-09-16T17:32:34.578154Z","iopub.status.idle":"2023-09-16T17:33:29.150818Z","shell.execute_reply.started":"2023-09-16T17:32:34.578115Z","shell.execute_reply":"2023-09-16T17:33:29.149231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Clustering samples (i.e. pairs cell + compound)","metadata":{}},{"cell_type":"code","source":"%%time\nN = df_de_train.shape[1]# 5000\nX = df_de_train.iloc[:,5:N]\nprint(X.shape)\ncm = np.corrcoef(X)\nprint(cm[:3,:2])\ncm = np.abs(cm)\nl = [df_de_train['sm_name'].iat[i] +' '+ df_de_train['cell_type'].iat[i]  for i in range(len(df_de_train))] # .columns[5:N]\ncm = pd.DataFrame(cm, index =l , columns = l )\nprint(cm.shape)\nclustergrid = sns.clustermap(cm)\nplt.show()\nreordered_columns = clustergrid.dendrogram_col.reordered_ind\nreordered_rows = clustergrid.dendrogram_row.reordered_ind\nprint(len(reordered_rows), len(reordered_columns) )\nprint( list(cm.index[reordered_rows]) )\n# print( list(cm.columns[reordered_columns]) )\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:33:29.152947Z","iopub.execute_input":"2023-09-16T17:33:29.153998Z","iopub.status.idle":"2023-09-16T17:33:31.625430Z","shell.execute_reply.started":"2023-09-16T17:33:29.153951Z","shell.execute_reply":"2023-09-16T17:33:31.623626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Look in genes space ","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.decomposition import PCA\n\nX = df_de_train.iloc[:,5:].T\nprint(X.shape)\n\nv1_color = pd.Series(range(len(X)), name = 'index') # df_de_train[  'cell_type']\n\n\n\nlist_cfg = [ ['Genes',v1_color]]# , ['control' , v4_color ] , ['top compounds',v5_color] ]\n#     str_inf1 = ''\n    #X = np.clip(df.iloc[N0:N1,33:137].fillna(0),0, 1)\nstr_inf = 'PCA' \nreducer = PCA(n_components=10 )\nXr = reducer.fit_transform(X)\nfor i,j in [[0,1],[0,2],[1,2],[3,4],[5,6],[7,8]]:\n    plt.figure(figsize = (20,10)); ic=0\n    for str_inf1, v_for_color in list_cfg: # , ['*nib compounds ',v2_color], ['non *nib compounds',v3_color ] ]:\n        ic+=1; plt.subplot(1,len(list_cfg),ic)\n        sns.scatterplot(x= Xr[:,i], y = Xr[:,j], hue =  v_for_color ,s = 100) # df['reads'])\n        plt.xlabel(str_inf+str(i+1), fontsize = 20)\n        plt.ylabel(str_inf+str(j+1), fontsize = 20)\n        plt.title(str_inf1 + ' ', fontsize = 20 )\n\n    plt.show()\n    \n\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:33:31.627754Z","iopub.execute_input":"2023-09-16T17:33:31.628659Z","iopub.status.idle":"2023-09-16T17:33:44.770780Z","shell.execute_reply.started":"2023-09-16T17:33:31.628603Z","shell.execute_reply":"2023-09-16T17:33:44.769258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.decomposition import PCA\nimport umap \n\nX = df_de_train.iloc[:,5:].T\nprint(X.shape)\n\nv1_color = pd.Series(range(len(X)), name = 'index') # df_de_train[  'cell_type']\n\n\n\nlist_cfg = [ ['Genes',v1_color]]# , ['control' , v4_color ] , ['top compounds',v5_color] ]\n# str_inf = 'PCA' \n# reducer = PCA(n_components=10 )\nstr_inf = 'UMAP' \nreducer = umap.UMAP()# (n_components=10 )\n\nXr = reducer.fit_transform(X)\nfor i,j in [[0,1] ]:# ,[0,2],[1,2],[3,4],[5,6],[7,8]]:\n    plt.figure(figsize = (20,10)); ic=0\n    for str_inf1, v_for_color in list_cfg: # , ['*nib compounds ',v2_color], ['non *nib compounds',v3_color ] ]:\n        ic+=1; plt.subplot(1,len(list_cfg),ic)\n        sns.scatterplot(x= Xr[:,i], y = Xr[:,j], hue =  v_for_color ,s = 100) # df['reads'])\n        plt.xlabel(str_inf+str(i+1), fontsize = 20)\n        plt.ylabel(str_inf+str(j+1), fontsize = 20)\n        plt.title(str_inf1 + ' ', fontsize = 20 )\n\n    plt.show()\n    \n\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:33:44.777074Z","iopub.execute_input":"2023-09-16T17:33:44.777622Z","iopub.status.idle":"2023-09-16T17:34:10.503111Z","shell.execute_reply.started":"2023-09-16T17:33:44.777578Z","shell.execute_reply":"2023-09-16T17:34:10.500805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Genes clustering ","metadata":{}},{"cell_type":"code","source":"%%time\nN = 1000#  18211 #  15_000# 10000 #\nX = df_de_train.iloc[:,5:N].T\nprint(X.shape)\ncm = np.corrcoef(X)\nprint(cm[:3,:2])\ncm = np.abs(cm)\ncm = pd.DataFrame(cm, index = df_de_train.columns[5:N], columns = df_de_train.columns[5:N] )\nprint(cm.shape)\nsns.clustermap(cm)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:10.505439Z","iopub.execute_input":"2023-09-16T17:34:10.506498Z","iopub.status.idle":"2023-09-16T17:34:13.911636Z","shell.execute_reply.started":"2023-09-16T17:34:10.506449Z","shell.execute_reply":"2023-09-16T17:34:13.910012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Genes related to proliferation (cell cycle genes )\n\nProliferation cycle or cell cycle - key process for many cells. Many drugs here are anti-cancer - so should affect the ability of cells to proliferate. So might be interesting to look on drugs affect on these particular genes.\n\nTirosh et.al. proposed list of about 100 cell-cycle genes which are the most effectively seen by single cell data. It is good starting point.\n\nIn general there are much more genes related to the cell cycle, with various degree of \"relation\". Some list are cell cycle genes may contain thousands genes, but in fact in such huge lists most of the genes are related to cell cycle very weakly or these genes not good captured by single cell sequencing, while Tirosh list contains genes very strongly related to cell cycle and well captured by single cell technology. Moreover, many genes are associated to various biological processes in the cell, but Tirosh genes are mostly associated with cell cycle, not with the other processes - another argument why they are convenient to work with. \n\nOne may look at Computational challenges of cell cycle analysis using single cell transcriptomics\nAlexander Chervov, Andrei Zinovyev https://arxiv.org/abs/2208.05229\n\nAnd hundreds Kaggle notebooks/datasets related to that work e.g. that discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350314 , or that notebook: https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03b-daybydaychange-allcelltypes/notebook\n\nPS\n\nOther cell cycle genes sets e.g. by Tom Freeman\n\nSee some comparison e.g. here: https://www.kaggle.com/code/alexandervc/tabmuris-cell-cycle-1-data-one-by-one#Cell-cycle\nSo list is bigger but kind of more \"dirty\", that means containing more non-cell cycle effects, and G1S - G2M split is less prominent. \n","metadata":{}},{"cell_type":"code","source":"G1S_genes_Tirosh = ['MCM5', 'PCNA', 'TYMS', 'FEN1', 'MCM2', 'MCM4', 'RRM1', 'UNG', 'GINS2', 'MCM6', 'CDCA7', 'DTL', 'PRIM1', 'UHRF1', 'MLF1IP', 'HELLS', 'RFC2', 'RPA2', 'NASP', 'RAD51AP1', 'GMNN', 'WDR76', 'SLBP', 'CCNE2', 'UBR7', 'POLD3', 'MSH2', 'ATAD2', 'RAD51', 'RRM2', 'CDC45', 'CDC6', 'EXO1', 'TIPIN', 'DSCC1', 'BLM', 'CASP8AP2', 'USP1', 'CLSPN', 'POLA1', 'CHAF1B', 'BRIP1', 'E2F8']\nG2M_genes_Tirosh = ['HMGB2', 'CDK1', 'NUSAP1', 'UBE2C', 'BIRC5', 'TPX2', 'TOP2A', 'NDC80', 'CKS2', 'NUF2', 'CKS1B', 'MKI67', 'TMPO', 'CENPF', 'TACC3', 'FAM64A', 'SMC4', 'CCNB2', 'CKAP2L', 'CKAP2', 'AURKB', 'BUB1', 'KIF11', 'ANP32E', 'TUBB4B', 'GTSE1', 'KIF20B', 'HJURP', 'CDCA3', 'HN1', 'CDC20', 'TTK', 'CDC25C', 'KIF2C', 'RANGAP1', 'NCAPD2', 'DLGAP5', 'CDCA2', 'CDCA8', 'ECT2', 'KIF23', 'HMMR', 'AURKA', 'PSRC1', 'ANLN', 'LBR', 'CKAP5', 'CENPE', 'CTCF', 'NEK2', 'G2E3', 'GAS2L3', 'CBX5', 'CENPA']\ngenes_Tirosh = G1S_genes_Tirosh + G2M_genes_Tirosh\n\n# Subset of Tirosh genes to capture \"fast\" cell cycle pattern = see https://arxiv.org/abs/2208.05229\nlist_genes_fastCCsign = ['CDK1', 'UBE2C', 'TOP2A', 'TMPO', 'HJURP', 'RRM1', 'RAD51AP1', 'RRM2', 'CDC45', 'BLM', 'BRIP1', 'E2F8', 'HIST2H2AC']\n\nG1S_genes_Freeman = ['ADAMTS1', 'ASF1B', 'ATAD2', 'BARD1', 'BLM', 'BRCA1', 'BRIP1', 'C17orf75', 'C9orf40', 'CACYBP', 'CASP8AP2', 'CCDC15', 'CCNE1', 'CCNE2', 'CCP110', 'CDC25A', 'CDC45', 'CDC6', 'CDC7', 'CDK2', 'CDT1', 'CENPJ', 'CENPQ', 'CENPU', 'CEP57', 'CHAF1A', 'CHAF1B', 'CHEK1', 'CLSPN', 'CREBZF', 'CRYL1', 'CSE1L', 'DCLRE1B', 'DCTPP1', 'DEK', 'DERA', 'DHFR', 'DNA2', 'DNAJC9', 'DNMT1', 'DONSON', 'DSCC1', 'DSN1', 'DTL', 'E2F8', 'EED', 'EFCAB11', 'ENDOD1', 'ETAA1', 'EXO1', 'EYA2', 'EZH2', 'FAM111A', 'FANCE', 'FANCG', 'FANCI', 'FANCL', 'FBXO5', 'FEN1', 'GGH', 'GINS1', 'GINS2', 'GINS3', 'GLMN', 'GMNN', 'GMPS', 'GPD2', 'HADH', 'HELLS', 'HSF2', 'ITGB3BP', 'KIAA0101', 'KNTC1', 'LIG1', 'MCM10', 'MCM2', 'MCM3', 'MCM4', 'MCM5', 'MCM6', 'MCM7', 'MCMBP', 'METTL9', 'MMD', 'MNS1', 'MPP1', 'MRE11A', 'MSH2', 'MSH6', 'MYO19', 'NASP', 'NPAT', 'NSMCE4A', 'ORC1', 'OSGEPL1', 'PAK1', 'PAQR4', 'PARP2', 'PASK', 'PAXIP1', 'PBX3', 'PCNA', 'PKMYT1', 'PMS1', 'POLA1', 'POLA2', 'POLD3', 'POLE2', 'PRIM1', 'PRPS2', 'PSMC3IP', 'RAB23', 'RAD51', 'RAD51AP1', 'RAD54L', 'RBBP8', 'RBL1', 'RDX', 'RFC2', 'RFC3', 'RFC4', 'RMI1', 'RNASEH2A', 'RPA1', 'RRM1', 'RRM2', 'SLBP', 'SLC25A40', 'SMC2', 'SMC3', 'SSX2IP', 'SUPT16H', 'TEX30', 'TFDP1', 'THAP10', 'THEM6', 'TIMELESS', 'TIPIN', 'TMEM106C', 'TMEM38B', 'TRIM45', 'TRIP13', 'TSPYL4', 'TTI1', 'TUBGCP5', 'TYMS', 'UBR7', 'UNG', 'USP1', 'WDHD1', 'WDR76', 'WRB', 'YEATS4', 'ZBTB14', 'ZWINT']\nG2M_genes_Freeman = ['ADGRE5', 'ARHGAP11A', 'ARHGDIB', 'ARL6IP1', 'ASPM', 'AURKA', 'AURKB', 'BIRC5', 'BORA', 'BRD8', 'BUB1', 'BUB1B', 'BUB3', 'CCNA2', 'CCNB1', 'CCNB2', 'CCNF', 'CDC20', 'CDC25B', 'CDC25C', 'CDC27', 'CDCA3', 'CDCA8', 'CDK1', 'CDKN1B', 'CDKN3', 'CENPE', 'CENPF', 'CENPI', 'CENPN', 'CEP55', 'CEP70', 'CEP85', 'CKAP2', 'CKAP5', 'CKS1B', 'CKS2', 'CTCF', 'DBF4', 'DBF4B', 'DCAF7', 'DEPDC1', 'DLGAP5', 'ECT2', 'ERCC6L', 'ESPL1', 'FAM64A', 'FOXM1', 'FZD2', 'FZD7', 'FZR1', 'GPSM2', 'GTF2E1', 'GTSE1', 'H2AFX', 'HJURP', 'HMGB2', 'HMGB3', 'HMMR', 'HN1', 'INCENP', 'JADE2', 'KIF11', 'KIF14', 'KIF15', 'KIF18A', 'KIF18B', 'KIF20A', 'KIF20B', 'KIF22', 'KIF23', 'KIF2C', 'KIF4A', 'KIF5B', 'KIFC1', 'KPNA2', 'LBR', 'LMNB2', 'MAD2L1', 'MELK', 'MET', 'METTL4', 'MIS18BP1', 'MKI67', 'MPHOSPH9', 'MTMR6', 'NCAPD2', 'NCAPG', 'NCAPG2', 'NCAPH', 'NDC1', 'NDC80', 'NDE1', 'NEIL3', 'NEK2', 'NRF1', 'NUSAP1', 'OIP5', 'PAFAH2', 'PARPBP', 'PBK', 'PLEKHG3', 'PLK1', 'PLK4', 'PRC1', 'PRR11', 'PSRC1', 'PTTG1', 'PTTG3P', 'RACGAP1', 'RAD21', 'RASSF1', 'REEP4', 'SAP30', 'SHCBP1', 'SKA1', 'SLCO1B3', 'SOGA1', 'SPA17', 'SPAG5', 'SPC25', 'SPDL1', 'STIL', 'STK17B', 'TACC3', 'TAF5', 'TBC1D2', 'TBC1D31', 'TMPO', 'TOP2A', 'TPX2', 'TROAP', 'TTF2', 'TTK', 'TUBB4B', 'TUBD1', 'UBE2C', 'UBE2S', 'VANGL1', 'WEE1', 'WHSC1', 'XPO1', 'ZMYM1']\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:13.914080Z","iopub.execute_input":"2023-09-16T17:34:13.915015Z","iopub.status.idle":"2023-09-16T17:34:13.944470Z","shell.execute_reply.started":"2023-09-16T17:34:13.914972Z","shell.execute_reply":"2023-09-16T17:34:13.942716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nN = 1000#  18211 #  15_000# 10000 #\n\nfor l,str_inf in [ [G1S_genes_Tirosh, 'G1S Tirosh'], [G2M_genes_Tirosh, 'G2M Tirosh'],  [G1S_genes_Tirosh + G2M_genes_Tirosh, 'All Tirosh'],\n                  [list_genes_fastCCsign, 'FastCC Signature'],\n                  [ G1S_genes_Freeman, 'G1S Freeman' ],  [ G2M_genes_Freeman, 'G2M Freeman' ], [ G1S_genes_Freeman + G2M_genes_Freeman, 'All Freeman' ],   ]: \n    ll = set(l) & set(df_de_train.columns) \n    ll = list(ll)\n    print(len(ll), str_inf )\n    X = df_de_train[ll].T # .iloc[:,5:N].T\n    print(X.shape)\n    cm = np.corrcoef(X)\n    print(cm[:3,:2])\n    cm = np.abs(cm)\n    cm = pd.DataFrame(cm, index = ll, columns = ll )\n    print(cm.shape)\n    clustergrid = sns.clustermap(cm)\n    plt.title(str_inf, fontsize = 20 )\n    plt.show()\n    reordered_columns = clustergrid.dendrogram_col.reordered_ind\n    reordered_rows = clustergrid.dendrogram_row.reordered_ind\n    print(len(reordered_rows), len(reordered_columns) )\n    print( list(cm.index[reordered_rows]) )\n#     print( list(cm.columns[reordered_columns]) )    \n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:13.946114Z","iopub.execute_input":"2023-09-16T17:34:13.946494Z","iopub.status.idle":"2023-09-16T17:34:22.192921Z","shell.execute_reply.started":"2023-09-16T17:34:13.946466Z","shell.execute_reply":"2023-09-16T17:34:22.189544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_de_train","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:22.195110Z","iopub.execute_input":"2023-09-16T17:34:22.195571Z","iopub.status.idle":"2023-09-16T17:34:22.240546Z","shell.execute_reply.started":"2023-09-16T17:34:22.195534Z","shell.execute_reply":"2023-09-16T17:34:22.238977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# for l,str_inf in [ [G1S_genes_Tirosh, 'G1S Tirosh'], [G2M_genes_Tirosh, 'G2M Tirosh'],  [G1S_genes_Tirosh + G2M_genes_Tirosh, 'All Tirosh'],\n#                   [list_genes_fastCCsign, 'FastCC Signature'],\n#                   [ G1S_genes_Freeman, 'G1S Freeman' ],  [ G2M_genes_Freeman, 'G2M Freeman' ], [ G1S_genes_Freeman + G2M_genes_Freeman, 'All Freeman' ],   ]: \nfor l,str_inf in [  [G1S_genes_Tirosh + G2M_genes_Tirosh, 'All Tirosh']   ]: \n    ll = set(l) & set(df_de_train.columns) \n    ll = list(ll)\n    print(len(ll), str_inf )\n    X = df_de_train[ ['cell_type'] + list(df_de_train.columns[5:]) ].groupby('cell_type').median()\n    X = X[ll]\n    print(X.shape)\n    clustergrid = sns.clustermap(X)# ,  annot=True, fmt=\".2f\", cmap=\"coolwarm\" )\n    plt.title(str_inf, fontsize = 20 )\n    plt.show()\n    reordered_columns = clustergrid.dendrogram_col.reordered_ind\n    reordered_rows = clustergrid.dendrogram_row.reordered_ind\n    print(len(reordered_rows), len(reordered_columns) )\n    print( list(X.index[reordered_rows]) )\n    print( list(X.columns[reordered_columns]) )\n    \ncol = 'sm_name'    \nfor l,str_inf in [  [G1S_genes_Tirosh + G2M_genes_Tirosh, 'All Tirosh']   ]: \n    ll = set(l) & set(df_de_train.columns) \n    ll = list(ll)\n    print(len(ll), str_inf )\n    X = df_de_train[ [col] + list(df_de_train.columns[5:]) ].groupby(col).median()\n    X = X[ll]\n    print(X.shape)\n    X.index = [t[:20] for t in X.index] # cut too long names\n    clustergrid = sns.clustermap(X)# ,  annot=True, fmt=\".2f\", cmap=\"coolwarm\" )\n    plt.title(str_inf, fontsize = 20 )\n    plt.show()    \n    reordered_columns = clustergrid.dendrogram_col.reordered_ind\n    reordered_rows = clustergrid.dendrogram_row.reordered_ind\n    print(len(reordered_rows), len(reordered_columns) )\n    print( list(X.index[reordered_rows]) )\n    print( list(X.columns[reordered_columns]) )","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:22.242832Z","iopub.execute_input":"2023-09-16T17:34:22.243257Z","iopub.status.idle":"2023-09-16T17:34:24.999757Z","shell.execute_reply.started":"2023-09-16T17:34:22.243223Z","shell.execute_reply":"2023-09-16T17:34:24.998132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Look on compounds ( count = 146  )\n\n15 compounds - 6 times data - only in train \n","metadata":{}},{"cell_type":"code","source":"d = df_de_train[['sm_name','sm_lincs_id','SMILES']].drop_duplicates()\nprint(d.shape)\nd.to_csv('compounds.csv')\ndisplay( d.head(10) )\n\nprint(list(df_de_train['sm_name'].unique() ) )","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:25.001860Z","iopub.execute_input":"2023-09-16T17:34:25.002293Z","iopub.status.idle":"2023-09-16T17:34:25.031723Z","shell.execute_reply.started":"2023-09-16T17:34:25.002258Z","shell.execute_reply":"2023-09-16T17:34:25.030243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = [len(s) for s in df_de_train['SMILES']]\nnp.sort(list(set(l)) )","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:25.033727Z","iopub.execute_input":"2023-09-16T17:34:25.034427Z","iopub.status.idle":"2023-09-16T17:34:25.046243Z","shell.execute_reply.started":"2023-09-16T17:34:25.034371Z","shell.execute_reply":"2023-09-16T17:34:25.044842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display( df_de_train['sm_name'].value_counts().head(20) )\ndisplay( df_de_train['sm_name'].value_counts().tail(10) )\ndf_de_train['sm_name'].value_counts().value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:25.048114Z","iopub.execute_input":"2023-09-16T17:34:25.048485Z","iopub.status.idle":"2023-09-16T17:34:25.072319Z","shell.execute_reply.started":"2023-09-16T17:34:25.048452Z","shell.execute_reply":"2023-09-16T17:34:25.071281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aggregations by compounds, cell_types \n\nIt is used for prediction in early versions of the notebook","metadata":{}},{"cell_type":"code","source":"%%time\ntrain_aggregate_mean_or_median_or_whatever = df_de_train.iloc[:,5:].quantile(0.7)# median()\ntrain_aggregate_mean_or_median_or_whatever\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:25.074235Z","iopub.execute_input":"2023-09-16T17:34:25.074555Z","iopub.status.idle":"2023-09-16T17:34:25.550925Z","shell.execute_reply.started":"2023-09-16T17:34:25.074528Z","shell.execute_reply":"2023-09-16T17:34:25.549771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nd = train_aggregate_mean_or_median_or_whatever\nplt.figure(figsize = (20,4) )\nplt.plot(d.values)\nplt.show()\nplt.figure(figsize = (10,4) )\nplt.hist(d.values, bins = 100)\nplt.show()\n\ndisplay( d.describe() )","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:25.552589Z","iopub.execute_input":"2023-09-16T17:34:25.552989Z","iopub.status.idle":"2023-09-16T17:34:26.384805Z","shell.execute_reply.started":"2023-09-16T17:34:25.552947Z","shell.execute_reply":"2023-09-16T17:34:26.383155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission data","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf = pd.read_csv(fn, index_col = 0)\nprint(df.shape)\ndf","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:26.386563Z","iopub.execute_input":"2023-09-16T17:34:26.386912Z","iopub.status.idle":"2023-09-16T17:34:32.421596Z","shell.execute_reply.started":"2023-09-16T17:34:26.386884Z","shell.execute_reply":"2023-09-16T17:34:32.419914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submission by aggregated-target","metadata":{}},{"cell_type":"markdown","source":"## Key params","metadata":{}},{"cell_type":"code","source":"# predict_method = 'train_aggregation_by_compounds'\n# predict_method = 'train_aggregation_by_compounds_with_denoising_pca'\n# predict_method = 'train_aggregation_by_compounds_with_denoising_ICA'\npredict_method = 'train_aggregation_by_compounds_with_denoising_TSVD'\nquantile = 0.54\nn_components = 35\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:32.423657Z","iopub.execute_input":"2023-09-16T17:34:32.425138Z","iopub.status.idle":"2023-09-16T17:34:32.431490Z","shell.execute_reply.started":"2023-09-16T17:34:32.425082Z","shell.execute_reply":"2023-09-16T17:34:32.430122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare predictions by direct aggregation of targets by compounds","metadata":{}},{"cell_type":"code","source":"%%time\ntrain_aggr_direct = df_de_train[ ['sm_name'] + list(df_de_train.columns[5:])  ].groupby('sm_name' ).quantile(quantile)# median()\ntrain_aggr_direct\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:32.433222Z","iopub.execute_input":"2023-09-16T17:34:32.433662Z","iopub.status.idle":"2023-09-16T17:34:34.401463Z","shell.execute_reply.started":"2023-09-16T17:34:32.433620Z","shell.execute_reply":"2023-09-16T17:34:34.399884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepapre predictions by aggregation of PCA-reduced targets by compounds","metadata":{}},{"cell_type":"code","source":"%%time \nfrom sklearn.decomposition import PCA\nfrom sklearn.decomposition import FastICA\nfrom sklearn.decomposition import TruncatedSVD\n\nY = df_de_train.iloc[:,5:]\nprint(X.shape)\nif '_pca' in predict_method:\n    str_inf_target_dimred = 'PCA' \n    reducer = PCA(n_components=n_components )\nelif '_ICA' in predict_method:\n    str_inf_target_dimred = 'ICA' \n#     reducer = PCA(n_components=n_components )\n    reducer = FastICA(n_components=n_components, random_state=0, whiten='unit-variance')\nelif '_TSVD' in predict_method:\n    str_inf_target_dimred = 'TSVD' \n#     reducer = PCA(n_components=n_components )\n#     reducer = FastICA(n_components=n_components, random_state=0, whiten='unit-variance')\n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\nelse:\n    str_inf_target_dimred = ''\n    \nprint(str_inf_target_dimred , reducer)\nYr = reducer.fit_transform(Y)\nYr_inv_trans = reducer.inverse_transform(Yr)\ndf_red_inv_trans = pd.DataFrame(Yr_inv_trans, columns = df_de_train.columns[5:])\ndf_red_inv_trans['sm_name'] = df_de_train['sm_name']\n\ntrain_aggr_denoised = df_red_inv_trans.groupby('sm_name' ).quantile(quantile)# median()\ntrain_aggr_denoised\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:34.403613Z","iopub.execute_input":"2023-09-16T17:34:34.404115Z","iopub.status.idle":"2023-09-16T17:34:39.406708Z","shell.execute_reply.started":"2023-09-16T17:34:34.404071Z","shell.execute_reply":"2023-09-16T17:34:39.404971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nif predict_method == 'train_aggregation_by_compounds':\n    df = df_id_map.merge(train_aggr, how = 'left', on = 'sm_name')\n    df = df.set_index('id').iloc[:,2:]\nelif predict_method.startswith('train_aggregation_by_compounds_with_denoising_'):\n    df = df_id_map.merge(train_aggr_denoised, how = 'left', on = 'sm_name')\n    df = df.set_index('id').iloc[:,2:]\nelse:\n    # consant for each target submission:\n    for i,col in enumerate( df.columns ):\n        df[col] = train_aggregate_mean_or_median_or_whatever[col]\n        if (i%1000) == 0: print(i,col)\n    \ndf    ","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:39.408586Z","iopub.execute_input":"2023-09-16T17:34:39.409013Z","iopub.status.idle":"2023-09-16T17:34:39.501177Z","shell.execute_reply.started":"2023-09-16T17:34:39.408981Z","shell.execute_reply":"2023-09-16T17:34:39.499497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:39.503081Z","iopub.execute_input":"2023-09-16T17:34:39.503421Z","iopub.status.idle":"2023-09-16T17:34:39.544724Z","shell.execute_reply.started":"2023-09-16T17:34:39.503393Z","shell.execute_reply":"2023-09-16T17:34:39.543251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:39.546993Z","iopub.execute_input":"2023-09-16T17:34:39.547388Z","iopub.status.idle":"2023-09-16T17:34:54.716112Z","shell.execute_reply.started":"2023-09-16T17:34:39.547357Z","shell.execute_reply":"2023-09-16T17:34:54.715147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Towards modeling","metadata":{}},{"cell_type":"markdown","source":"## Key params","metadata":{}},{"cell_type":"code","source":"n_components_for_cell_type_encoding = 1\nn_components_for_compound_encoding = 25\nalpha_regularization_for_linear_models = 10\n\n# predict_method\n\nmodel_type = 'Ridge'\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:54.722239Z","iopub.execute_input":"2023-09-16T17:34:54.722803Z","iopub.status.idle":"2023-09-16T17:34:54.726883Z","shell.execute_reply.started":"2023-09-16T17:34:54.722772Z","shell.execute_reply":"2023-09-16T17:34:54.726119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"str_model_id = model_type\nstr_model_id += ' nCT'+ str(n_components_for_cell_type_encoding)\nstr_model_id += ' nCD'+ str(n_components_for_compound_encoding)\nstr_model_id += ' Al'+ str(alpha_regularization_for_linear_models)\nstr_model_id += ' ' +str_inf_target_dimred+str( n_components )\n\nprint( str_model_id )\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:54.728182Z","iopub.execute_input":"2023-09-16T17:34:54.728782Z","iopub.status.idle":"2023-09-16T17:34:54.752489Z","shell.execute_reply.started":"2023-09-16T17:34:54.728751Z","shell.execute_reply":"2023-09-16T17:34:54.750653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target encoded features ","metadata":{"execution":{"iopub.status.busy":"2023-09-15T09:29:29.58894Z","iopub.execute_input":"2023-09-15T09:29:29.589366Z","iopub.status.idle":"2023-09-15T09:29:29.623736Z","shell.execute_reply.started":"2023-09-15T09:29:29.589335Z","shell.execute_reply":"2023-09-15T09:29:29.623004Z"}}},{"cell_type":"code","source":"%%time\n# Yr = reducer.fit_transform(X)\n# n_components_for_cell_type_encoding = 10\ndf_tmp = pd.DataFrame(Yr[:, :n_components_for_cell_type_encoding  ], index = df_de_train.index  )\ndf_tmp['column for aggregation'] = df_de_train['cell_type']\ndf_cell_type_encoded = df_tmp.groupby('column for aggregation').quantile( quantile )\nprint('df_cell_type_encoded.shape', df_cell_type_encoded.shape )\ndisplay( df_cell_type_encoded )\n\n\n# n_components_for_compound_encoding = 10\ndf_tmp = pd.DataFrame(Yr[:, :n_components_for_compound_encoding  ], index = df_de_train.index  )\ndf_tmp['column for aggregation'] = df_de_train['sm_name']\ndf_compound_encoded = df_tmp.groupby('column for aggregation').quantile( quantile )\nprint('df_compound_encoded.shape', df_compound_encoded.shape )\ndisplay( df_compound_encoded )\n\n","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:54.754782Z","iopub.execute_input":"2023-09-16T17:34:54.755890Z","iopub.status.idle":"2023-09-16T17:34:54.823351Z","shell.execute_reply.started":"2023-09-16T17:34:54.755836Z","shell.execute_reply":"2023-09-16T17:34:54.821697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare X_train, X_submit - target encoded cell type and compound features","metadata":{}},{"cell_type":"code","source":"%%time\nX_train = np.zeros( (len( df_de_train ) , n_components_for_cell_type_encoding + n_components_for_compound_encoding ))\n\nfor i in range(len( X_train )):\n    cell_type = df_de_train['cell_type'].iat[i] \n    X_train[i,:n_components_for_cell_type_encoding] = df_cell_type_encoded.loc[cell_type,:].values  \n    compound = df_de_train['sm_name'].iat[i] \n    X_train[i,n_components_for_cell_type_encoding:] = df_compound_encoded.loc[ compound, : ].values\nprint( X_train.shape)     \nprint( X_train)     \n    \n\nX_submit = np.zeros( (len( df_id_map ) , n_components_for_cell_type_encoding + n_components_for_compound_encoding ))\nfor i in range(len( X_submit )):\n    cell_type = df_id_map['cell_type'].iat[i] \n    X_submit[i,:n_components_for_cell_type_encoding] = df_cell_type_encoded.loc[cell_type,:].values  \n    compound = df_id_map['sm_name'].iat[i] \n    X_submit[i,n_components_for_cell_type_encoding:] = df_compound_encoded.loc[ compound, : ].values\n    \n    \nprint( X_submit.shape)     \nprint( X_submit)     \n    ","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:54.825694Z","iopub.execute_input":"2023-09-16T17:34:54.826268Z","iopub.status.idle":"2023-09-16T17:34:55.019022Z","shell.execute_reply.started":"2023-09-16T17:34:54.826219Z","shell.execute_reply":"2023-09-16T17:34:55.017512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.linear_model import Ridge\n\nmodel = Ridge(alpha=alpha_regularization_for_linear_models)\nprint(model)\nmodel.fit(X_train, Yr)\n\nY_submit = reducer.inverse_transform(   model.predict(X_submit) )\nprint(Y_submit.shape)\nY_submit","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:55.021590Z","iopub.execute_input":"2023-09-16T17:34:55.022094Z","iopub.status.idle":"2023-09-16T17:34:55.085334Z","shell.execute_reply.started":"2023-09-16T17:34:55.022026Z","shell.execute_reply":"2023-09-16T17:34:55.080934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save submission CSV","metadata":{}},{"cell_type":"code","source":"%%time\ndf_submit = pd.DataFrame(Y_submit, columns = df_de_train.columns[5:])\ndf_submit.index.name = 'id'\nprint( df_submit.shape )\ndisplay(df_submit)\ndf_submit.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:34:55.087744Z","iopub.execute_input":"2023-09-16T17:34:55.089060Z","iopub.status.idle":"2023-09-16T17:35:09.501517Z","shell.execute_reply.started":"2023-09-16T17:34:55.088996Z","shell.execute_reply":"2023-09-16T17:35:09.500107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Catboost - way 3","metadata":{}},{"cell_type":"code","source":"import catboost\nfrom catboost import CatBoostRegressor, Pool","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.503330Z","iopub.execute_input":"2023-09-16T17:35:09.503925Z","iopub.status.idle":"2023-09-16T17:35:09.509377Z","shell.execute_reply.started":"2023-09-16T17:35:09.503893Z","shell.execute_reply":"2023-09-16T17:35:09.508108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_features = ['cell_type','sm_name']\ndf_train = df_de_train[['cell_type','sm_name']]\nY = df_de_train['A1BG']\n\ntrain_data = Pool(data=df_train, \n                  label=Y,\n                  cat_features=categorical_features)\n\nmodel = CatBoostRegressor(iterations=100,  # Number of boosting iterations\n                          depth=2,        # Depth of the trees\n                          learning_rate=0.1,  # Learning rate\n                          loss_function='RMSE',  # Specify your loss function (e.g., RMSE for regression)\n                          cat_features=categorical_features,  # Categorical features\n                          verbose=0)  # Set verbose to 0 to suppress output\n\n# Train the model\nmodel.fit(train_data)\n\n\n# model = CatBoostClassifier(iterations=100,  # Number of boosting iterations\n#                            depth=6,        # Depth of the trees\n#                            learning_rate=0.1,  # Learning rate\n#                            loss_function='Logloss',  # Specify your loss function\n#                            cat_features=categorical_features,  # Categorical features\n#                            verbose=0)  # Set verbose to 0 to suppress output\n\n# # Train the model\n# model.fit(train_data)","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.511163Z","iopub.execute_input":"2023-09-16T17:35:09.511501Z","iopub.status.idle":"2023-09-16T17:35:09.895709Z","shell.execute_reply.started":"2023-09-16T17:35:09.511468Z","shell.execute_reply":"2023-09-16T17:35:09.894357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # For classification models (CatBoostClassifier)\n# predicted_class = model.predict(input_data)\n\n# For regression models (CatBoostRegressor)\npredicted_value = model.predict(train_data)\npredicted_value.shape","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.897875Z","iopub.execute_input":"2023-09-16T17:35:09.898332Z","iopub.status.idle":"2023-09-16T17:35:09.905947Z","shell.execute_reply.started":"2023-09-16T17:35:09.898301Z","shell.execute_reply":"2023-09-16T17:35:09.905116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = Pool(data=df_id_map[['cell_type','sm_name']], \n                  cat_features=categorical_features)","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.907755Z","iopub.execute_input":"2023-09-16T17:35:09.908094Z","iopub.status.idle":"2023-09-16T17:35:09.927451Z","shell.execute_reply.started":"2023-09-16T17:35:09.908066Z","shell.execute_reply":"2023-09-16T17:35:09.926273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_id_map[['cell_type','sm_name']]\npredicted_value = model.predict(test_data)\npredicted_value.shape","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.928894Z","iopub.execute_input":"2023-09-16T17:35:09.929477Z","iopub.status.idle":"2023-09-16T17:35:09.946464Z","shell.execute_reply.started":"2023-09-16T17:35:09.929447Z","shell.execute_reply":"2023-09-16T17:35:09.945525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Yr.shape","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.947891Z","iopub.execute_input":"2023-09-16T17:35:09.948553Z","iopub.status.idle":"2023-09-16T17:35:09.959686Z","shell.execute_reply.started":"2023-09-16T17:35:09.948509Z","shell.execute_reply":"2023-09-16T17:35:09.958166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nY_reduced_submit = np.zeros( (len(df_id_map) , Yr.shape[1] )   )\nfor k in range(Yr.shape[1]):\n    train_data = Pool(data=df_train, \n                      label=Yr[:,k],\n                      cat_features=categorical_features)\n    test_data = Pool(data=df_id_map[['cell_type','sm_name']], \n                      cat_features=categorical_features)\n    model.fit(train_data)    \n    predicted_value = model.predict(test_data)\n    Y_reduced_submit[:,k] = predicted_value\n    \n    ","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:09.963208Z","iopub.execute_input":"2023-09-16T17:35:09.963542Z","iopub.status.idle":"2023-09-16T17:35:12.813727Z","shell.execute_reply.started":"2023-09-16T17:35:09.963515Z","shell.execute_reply":"2023-09-16T17:35:12.812304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nY_submit = reducer.inverse_transform(  Y_reduced_submit )\nprint(Y_submit.shape)\nY_submit","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:12.815361Z","iopub.execute_input":"2023-09-16T17:35:12.815714Z","iopub.status.idle":"2023-09-16T17:35:12.857490Z","shell.execute_reply.started":"2023-09-16T17:35:12.815683Z","shell.execute_reply":"2023-09-16T17:35:12.855756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_submit = pd.DataFrame(Y_submit, columns = df_de_train.columns[5:])\ndf_submit.index.name = 'id'\nprint( df_submit.shape )\ndisplay(df_submit)\ndf_submit.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:12.860134Z","iopub.execute_input":"2023-09-16T17:35:12.860741Z","iopub.status.idle":"2023-09-16T17:35:27.120228Z","shell.execute_reply.started":"2023-09-16T17:35:12.860686Z","shell.execute_reply":"2023-09-16T17:35:27.119067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\nprint('%.1f minutes passed total '%( (time.time()-t0start)/60)  )\nprint('%.2f hours passed total '%( (time.time()-t0start)/3600)  )","metadata":{"execution":{"iopub.status.busy":"2023-09-16T17:35:27.121731Z","iopub.execute_input":"2023-09-16T17:35:27.123341Z","iopub.status.idle":"2023-09-16T17:35:27.131717Z","shell.execute_reply.started":"2023-09-16T17:35:27.123300Z","shell.execute_reply":"2023-09-16T17:35:27.130462Z"},"trusted":true},"execution_count":null,"outputs":[]}]}