{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":7092391,"sourceType":"datasetVersion","datasetId":3834323},{"sourceId":7122895,"sourceType":"datasetVersion","datasetId":3948965},{"sourceId":7131543,"sourceType":"datasetVersion","datasetId":3944109}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nModels diversity analysis U900 team for Kaggle challenge [Open Problems – Single-Cell Perturbations](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations)\n\nDRAFT REPORT VERSION TO BE UPDATED. COMMENTS ARE WELCOME.\n\nIt is  supplementary to:  Explaining the blend scheme: https://www.kaggle.com/code/alexandervc/op2-u900-team-blend (see it for more information).\n\nMain write-up: TODO - LINK\n\n\n### Context/problem/finding\n\nThe standard idea: more diversity - better blend. However quantification of the diversity itself and the correpodence - what diversity gives what uplift in blend - is not an easy task, especially for  multi-target tasks. The standard route - choose weights for blend by CV-scoring is not working here - because CV-LB correspondence is poor.\n\nOur **finding** - the simple quantification works well as guidence for blend - i.e.  compute correlations for each target and average over targets, then  blend of two models with  correlation score 0.8-0.9 typically improves the score about 0.01 - 0.006  .  \n\n\n### Content \n\nHere we provide code and examples how we estimated the diversity of models. \n\nThree sets  of models are presented - 4 models in the core blend scheme; 11 models/variations in the final blend submit; many (52) models/variations.\n\nThe attached Kaggle datasets contain many of our submissions  and with the clustermap plot we can see  general  diversity landscape for large number of models.\nIt helps to choose the models which are diverse enough.\n\nSimilar analysis of public submits has been shared with community during the challenge: https://www.kaggle.com/code/alexandervc/op2-submits-correlations-and-analysis\n\n### Details:\n\nCorrelation score between two models is computed as follows:\nConsider two models. For each target take prediction of model1 and  prediction of model2  and compute their Pearson correlation.\nAnd take average of these correlations over all 18211 targets.\n\nWe take that averaged correlation as our measure of diversity - the lower - the better.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Preliminaries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pandas as pd\nimport tensorflow as tf\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\ncc = 0\nprint('Show first 15 files:')\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        cc += 1\n        if cc <= 15:\n            print(os.path.join(dirname, filename))\nprint()            \nprint('Total files count:', cc)\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:06:10.890166Z","iopub.execute_input":"2023-12-07T10:06:10.890532Z","iopub.status.idle":"2023-12-07T10:06:10.906003Z","shell.execute_reply.started":"2023-12-07T10:06:10.890505Z","shell.execute_reply":"2023-12-07T10:06:10.905274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlation analysis for key submits ","metadata":{}},{"cell_type":"code","source":"%%time\n\n##################################################################################\n# Specify list of submit files to be processed and short aliases for their names \n##################################################################################\n\n\nlist_fn = []\nlist_ids = []\n\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Catboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Pyboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\nlist_fn.append( fn )\nlist_ids.append('LB574_NN_NLP')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\nlist_fn.append( fn )\nlist_ids.append('LB566_NN_TargetEnc_Ensemble')\n\n                \nprint(len(list_fn))    \nprint(len(list_ids))    \nprint(list_ids)\nprint(list_fn)\nprint()\n\n\n\n##################################################################################\n# Load stored submits specified by the list_fn \n##################################################################################\n\nprint('Start load submit files:')\n\n#print(list_ids)\nlist_df = []\n\ni_blend = 0\nfor i,fn in enumerate(list_fn): \n#     print(i, fn)\n    df = pd.read_csv(fn, index_col = 'id')\n    print(i,df.shape, fn , np.round( [df.iloc[0,0],  df.iloc[0,1],  df.iloc[1,0] ,  df.iloc[1,1] ], 4) , df.columns[0],   df.columns[1]   )\n    #display(df.head(2) )\n    list_df.append(df)\n    if i > 0: # not blend the first one\n        if i_blend == 0:\n            df_blend = df.copy()\n        else:\n            df_blend = (df_blend * i_blend  + df )/ (i_blend + 1)\n        i_blend += 1\n        \nflag_include_average_blend_of_all = False        \nif flag_include_average_blend_of_all:        \n    list_ids.append('Blend')  \n    list_df.append(df_blend )    \n\n\n##################################################################################\n# Prepare fast correlations computation - standard scaler each submit dataframe\n##################################################################################\n\n# %%time\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nlist_np = []\nif 1:\n    for k in range(len(list_df)): # dict_df:\n        df = list_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\nelse:\n    for k in dict_df:\n        df = dict_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\n    \nprint(len(list_np)) \n\n\n##################################################################################\n# Compute correlations\n# I.e. average column-wise correlations\n# Due to standard scaling on the previous step: correlation = produce + average\n##################################################################################\n\n\nfig = plt.figure(figsize = (20,10))\ndf_stat = pd.DataFrame()\ndf_corr_averages_for_submits = pd.DataFrame()\ni_total = 0\nfor i0,d0 in  enumerate(list_np):\n    for i1,d1 in  enumerate(list_np):\n        i_total += 1\n        v = np.mean( d0*d1, axis = 0 )\n        if i0<i1:\n            plt.hist(v, label = str(i0)+ ' vs ' + str(i1))\n        dt = pd.Series(v).describe().to_frame()\n        dt.columns = [str(i0)+ ' vs  ' + str(i1)]\n        df_stat = pd.concat( [df_stat, dt] , axis = 1)\n        #print(  )\n        # c = np.mean( np.abs(v) )\n        c = np.mean( (v) )\n        if i_total<3:\n            print(c)\n        df_corr_averages_for_submits.loc[i0,i1] = c\n        \ndisplay(df_stat)        \n# plt.legend(fontsize = 12 )\nplt.grid()\nplt.title('Distribution of correlations (over genes) for each prediction pair ', fontsize = 20)\nplt.show()        \ndf_corr_averages_for_submits.columns = [t.replace('_',' ') for t in  list_ids]\ndf_corr_averages_for_submits.index = [t.replace('_',' ') for t in  list_ids]\n\nprint('Show 3x3 part of the correlation matrix:')\ndisplay(df_corr_averages_for_submits.iloc[:5,:5].round(2) )  \n\n##################################################################################\n# Show clustermap\n##################################################################################\n\nsns.clustermap( df_corr_averages_for_submits,figsize = (20,15),  annot=True, cmap='coolwarm')\nplt.show()\n\n# sns.clustermap( df_corr_averages_for_submits,figsize = (20,12), annot=False, cmap='coolwarm')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:06:10.907916Z","iopub.execute_input":"2023-12-07T10:06:10.908381Z","iopub.status.idle":"2023-12-07T10:06:31.167213Z","shell.execute_reply.started":"2023-12-07T10:06:10.908357Z","shell.execute_reply":"2023-12-07T10:06:31.166067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_corr_averages_for_submits.round(2)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:06:31.169442Z","iopub.execute_input":"2023-12-07T10:06:31.169933Z","iopub.status.idle":"2023-12-07T10:06:31.186562Z","shell.execute_reply.started":"2023-12-07T10:06:31.169899Z","shell.execute_reply":"2023-12-07T10:06:31.184318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlations of other submits included in the final submit","metadata":{}},{"cell_type":"code","source":"%%time\n\n##################################################################################\n# Specify list of submit files to be processed and short aliases for their names \n##################################################################################\n\n\nlist_fn = []\nlist_ids = []\n\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Catboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Pyboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\nlist_fn.append( fn )\nlist_ids.append('LB574_NN_NLP')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\nlist_fn.append( fn )\nlist_ids.append('LB566_NN_TargetEnc_Ensemble')\n\n\n# Additional Pyboosts and NNs:\n\n\nfn = '/kaggle/input/open-problems-2-submits-collection/LB574_Pyboostmaxdepth12ntrees2000lr001subsample1colsample035ncomponents50T8T8b7t17_MadrisMillerBasedAlex_nbV8.csv'\nlist_fn.append( fn )\nlist_ids.append('LB574_Pyboost')\n\nfn = '/kaggle/input/open-problems-2-submits-collection/LB577_PyboostCVRandom5_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50_NikolenkoPubl.csv'\nlist_fn.append( fn )\nlist_ids.append('LB577_Pyboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPpseudo50ctstratmrrmsetfsmilesv_nbV2.csv'\nlist_fn.append( fn )\nlist_ids.append('LB587_NN1')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPblendown.csv'\nlist_fn.append( fn )\nlist_ids.append('LB587_NN2')\n\n\n# Additional Target Encoded NN :\nfn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_T8_Treg_3kmeans_Antonina.csv'\nlist_fn.append( fn )\nlist_ids.append('LB570_NN_TargetEnc_Additional1')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\nlist_fn.append( fn )\nlist_ids.append('LB569_NN_TargetEnc_Additional2')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB572_ave_blend_T4_T8_20ep_augm50_s1_0.1.csv'\nlist_fn.append( fn )\nlist_ids.append('LB569_NN_TargetEnc_Additional3')\n\n                \nprint(len(list_fn))    \nprint(len(list_ids))    \nprint(list_ids)\nprint(list_fn)\nprint()\n\n\n\n##################################################################################\n# Load stored submits specified by the list_fn \n##################################################################################\n\nprint('Start load submit files:')\n\n#print(list_ids)\nlist_df = []\n\ni_blend = 0\nfor i,fn in enumerate(list_fn): \n#     print(i, fn)\n    df = pd.read_csv(fn, index_col = 'id')\n    print(i,df.shape, fn , np.round( [df.iloc[0,0],  df.iloc[0,1],  df.iloc[1,0] ,  df.iloc[1,1] ], 4) , df.columns[0],   df.columns[1]   )\n    #display(df.head(2) )\n    list_df.append(df)\n    if i > 0: # not blend the first one\n        if i_blend == 0:\n            df_blend = df.copy()\n        else:\n            df_blend = (df_blend * i_blend  + df )/ (i_blend + 1)\n        i_blend += 1\n        \nflag_include_average_blend_of_all = False        \nif flag_include_average_blend_of_all:        \n    list_ids.append('Blend')  \n    list_df.append(df_blend )    \n\n\n##################################################################################\n# Prepare fast correlations computation - standard scaler each submit dataframe\n##################################################################################\n\n# %%time\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nlist_np = []\nif 1:\n    for k in range(len(list_df)): # dict_df:\n        df = list_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\nelse:\n    for k in dict_df:\n        df = dict_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\n    \nprint(len(list_np)) \n\n\n##################################################################################\n# Compute correlations\n# I.e. average column-wise correlations\n# Due to standard scaling on the previous step: correlation = produce + average\n##################################################################################\n\n\nfig = plt.figure(figsize = (20,10))\ndf_stat = pd.DataFrame()\ndf_corr_averages_for_submits = pd.DataFrame()\ni_total = 0\nfor i0,d0 in  enumerate(list_np):\n    for i1,d1 in  enumerate(list_np):\n        i_total += 1\n        v = np.mean( d0*d1, axis = 0 )\n        if i0<i1:\n            plt.hist(v, label = str(i0)+ ' vs ' + str(i1))\n        dt = pd.Series(v).describe().to_frame()\n        dt.columns = [str(i0)+ ' vs  ' + str(i1)]\n        df_stat = pd.concat( [df_stat, dt] , axis = 1)\n        #print(  )\n        # c = np.mean( np.abs(v) )\n        c = np.mean( (v) )\n        if i_total<3:\n            print(c)\n        df_corr_averages_for_submits.loc[i0,i1] = c\n        \ndisplay(df_stat)        \n# plt.legend(fontsize = 12 )\nplt.grid()\nplt.title('Distribution of correlations (over genes) for each prediction pair ', fontsize = 20)\nplt.show()        \ndf_corr_averages_for_submits.columns = [t.replace('_',' ') for t in  list_ids]\ndf_corr_averages_for_submits.index = [t.replace('_',' ') for t in  list_ids]\n\nprint('Show 3x3 part of the correlation matrix:')\ndisplay(df_corr_averages_for_submits.iloc[:5,:5].round(2) )  \n\n##################################################################################\n# Show clustermap\n##################################################################################\n\nsns.clustermap( df_corr_averages_for_submits,figsize = (20,15),  annot=True, cmap='coolwarm')\nplt.show()\n\n# sns.clustermap( df_corr_averages_for_submits,figsize = (20,12), annot=False, cmap='coolwarm')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:06:31.190194Z","iopub.execute_input":"2023-12-07T10:06:31.190702Z","iopub.status.idle":"2023-12-07T10:07:21.791822Z","shell.execute_reply.started":"2023-12-07T10:06:31.190665Z","shell.execute_reply":"2023-12-07T10:07:21.790479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_corr_averages_for_submits.round(2)","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:07:21.793283Z","iopub.execute_input":"2023-12-07T10:07:21.793575Z","iopub.status.idle":"2023-12-07T10:07:21.817164Z","shell.execute_reply.started":"2023-12-07T10:07:21.793552Z","shell.execute_reply":"2023-12-07T10:07:21.815973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlation analysis for most of the other submits","metadata":{}},{"cell_type":"code","source":"list_fn = []\nlist_ids = []\nlist_fn.append('/kaggle/input/open-problems-2-submits-etc/LB0571_blendA0574Publ0583w0703_Alex_OP2Blends_nbV6.csv')\nlist_ids.append('LB571Ant575publ583')\n\n# fn  = '/kaggle/input/open-problems-2-submits-collection/LB720_JAXautoencoder_VENDEKAGONLABS_nbV4.csv'\n# list_fn.append(fn)\n# list_ids.append('LB720JAX')\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        # if 'kaggle/input/op2-submissions' in dirname : #  1: # ('LB0615' in dirname):\n        if ('kaggle/input/open-problems-2-submits-etc' in dirname)  and (not filename.startswith('df_stat_') ):\n            #if  ('simple' in filename) and ('blend' not in filename):\n            #if  ('Random' in filename) and ('blend' not in filename) and ('.csv' in filename):\n            if '.csv' in filename : #   ('Random' in filename) and ('priors' in filename) and ('.csv' in filename):\n                f = os.path.join(dirname, filename)\n                list_fn.append( f )\n                list_ids.append(f.split('/')[-1][:20]) \n                \nprint(len(list_fn))    \nprint(len(list_ids))    \nprint(list_ids)\nprint(list_fn )\nprint()\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:07:21.818364Z","iopub.execute_input":"2023-12-07T10:07:21.819136Z","iopub.status.idle":"2023-12-07T10:07:21.839684Z","shell.execute_reply.started":"2023-12-07T10:07:21.819107Z","shell.execute_reply":"2023-12-07T10:07:21.837854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprint(list_ids)\nlist_df = []\n\ni_blend = 0\nfor i,fn in enumerate(list_fn): \n#     print(i, fn)\n    df = pd.read_csv(fn, index_col = 'id')\n    print(i,df.shape, fn , np.round( [df.iloc[0,0],  df.iloc[0,1],  df.iloc[1,0] ,  df.iloc[1,1] ], 4) , df.columns[0],   df.columns[1]   )\n    #display(df.head(2) )\n    list_df.append(df)\n    if i > 0: # not blend the first one\n        if i_blend == 0:\n            df_blend = df.copy()\n        else:\n            df_blend = (df_blend * i_blend  + df )/ (i_blend + 1)\n        i_blend += 1\nlist_ids.append('Blend')  \nlist_df.append(df_blend )    \n\n\n# %%time\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nlist_np = []\nif 1:\n    for k in range(len(list_df)): # dict_df:\n        df = list_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\nelse:\n    for k in dict_df:\n        df = dict_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\n    \nprint(len(list_np)) ","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:07:21.841042Z","iopub.execute_input":"2023-12-07T10:07:21.841401Z","iopub.status.idle":"2023-12-07T10:11:20.143335Z","shell.execute_reply.started":"2023-12-07T10:07:21.841375Z","shell.execute_reply":"2023-12-07T10:11:20.141765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute correlations","metadata":{}},{"cell_type":"code","source":"%%time\n\nfig = plt.figure(figsize = (20,10))\ndf_stat = pd.DataFrame()\ndf_corr_averages_for_submits = pd.DataFrame()\ni_total = 0\nfor i0,d0 in  enumerate(list_np):\n    for i1,d1 in  enumerate(list_np):\n        i_total += 1\n        v = np.mean( d0*d1, axis = 0 )\n        if i0<i1:\n            plt.hist(v, label = str(i0)+ ' vs ' + str(i1))\n        dt = pd.Series(v).describe().to_frame()\n        dt.columns = [str(i0)+ ' vs  ' + str(i1)]\n        df_stat = pd.concat( [df_stat, dt] , axis = 1)\n        #print(  )\n        # c = np.mean( np.abs(v) )\n        c = np.mean( (v) )\n        if i_total<3:\n            print(c)\n        df_corr_averages_for_submits.loc[i0,i1] = c\n        \ndisplay(df_stat)        \n# plt.legend(fontsize = 12 )\nplt.grid()\nplt.title('Distribution of correlations (over genes) for each prediction pair ', fontsize = 20)\nplt.show()        \ndf_corr_averages_for_submits.columns = [t.replace('_',' ') for t in  list_ids]\ndf_corr_averages_for_submits.index = [t.replace('_',' ') for t in  list_ids]\n\nprint('Show 3x3 part of the correlation matrix:')\ndisplay(df_corr_averages_for_submits.iloc[:5,:5].round(2) )      ","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:11:20.144834Z","iopub.execute_input":"2023-12-07T10:11:20.145232Z","iopub.status.idle":"2023-12-07T10:13:37.759749Z","shell.execute_reply.started":"2023-12-07T10:11:20.145201Z","shell.execute_reply":"2023-12-07T10:13:37.758029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Clustermap","metadata":{}},{"cell_type":"code","source":"sns.clustermap( df_corr_averages_for_submits,figsize = (20,15),  annot=True, cmap='coolwarm')\nplt.show()\n\n# sns.clustermap( df_corr_averages_for_submits,figsize = (20,12), annot=False, cmap='coolwarm')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-07T10:13:37.762759Z","iopub.execute_input":"2023-12-07T10:13:37.763174Z","iopub.status.idle":"2023-12-07T10:13:44.319070Z","shell.execute_reply.started":"2023-12-07T10:13:37.763139Z","shell.execute_reply":"2023-12-07T10:13:44.318226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}