{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":7068220,"sourceType":"datasetVersion","datasetId":3955392},{"sourceId":7122895,"sourceType":"datasetVersion","datasetId":3948965},{"sourceId":7131543,"sourceType":"datasetVersion","datasetId":3944109}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nEnsembling (blending)  by U900 team for Kaggle challenge [Open Problems – Single-Cell Perturbations](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations).\n\n\nHere we comment  on logic and problems with blends for that challenge. \nPyboost was a main innovative tool for us ([breakthrough gradient boosting](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/454700) developped by team member Anton Vakhrushev), but in fear of  shake-up we decided to rely more on diversity rather than on strengh of individual models.\nOur ensemble includes several Pyboost models, one Catboost, and several NN ensembles . \n\nThe detailed write-up is available here: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/460858\n\nThe reasons to expect the shake-up (not happend for us) and in particular  poor CV - LB correspondence discussed here: TODO - LINK\n    \n## Highlights:\n\n* Main problem - absence of good CV - LB correspondence - forbids the usual strategy to choose weights by CV \n* Solution relies on: how to increase and control divesification;  how to avoid overfit-proning choice of  weights - a scheme with the only weight = 0.5.\n    * Measure of diversity - average target-wise correlation of predictions\n    * Check by various experiments:  models with correlation score  0.8-0.9 - consistenly give substantial uplift in blend (about +0.01 - 0.006)\n    * Avoid overfit-proning question: how to choose weights of the models with DIFFERENT scores, by the following scheme:\n    * Core scheme: sequential blend of the models with EQUAL (almost) scores, giving them  EQUAL blend weight (=0.5):\n    * step1: LB score 0.575 = 0.5 * Pyboost(0.584) + 0.5 * Catboost(0.584) TODO LINK\n    * step2: LB score 0.566 = 0.5 * step1 (0.575 )  + 0.5 * NN-NLP  (0.574) TODO LINK\n    * step3: LB score 0.559 = 0.5 * step2 (0.566 )  + 0.5 * [MLP_TargetEncEnsemble](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution) (0.566)\n    * final polishing: 0.558  - blend more Pyboost (0.574,0.577) and NN (0.569,0.570,0.572,0.587) models:\n\n#### Other blend ideas tried:\n\n* Different weights for B-cells and Myeloid cells (partially successful)\n* Tried, but had not enougth time to succeeded:\n    * Estimate variance and correlations for each target (or each row) of predictions  and choose blend weights according to modifications of the classical stastical formula - weights are proportianal to variance - bigger variance - less confidence - lower weight in blend\n\n#### Future plans:\n\nWe are sure that blend schemes for such multi-target tasks can be improved, we plan to try the following:\n\n* Incorpoate blend scheme (target-wise) proposed by [Mehran Kazeminia](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459001#2547289). Setup: blend a strong solution with a weak one ; method:  choose for blend only those targets where correlation is positive - and otherwise select just the strong solution (no blend for that target)\n* Improve esimates of the variances and correlations for columns/rows and try again the classical statistical approach (see details below)\n\n\n#### Sharing with community during and after the challenge: \n\nFor the sake of the community we have shared our ideas on blend and method to measure the diversity during the competition: https://www.kaggle.com/code/alexandervc/op2-submits-correlations-and-analysis , and created and made publicly available the Kaggle datasets with partitial collection of our own submissions and oof predictions: https://www.kaggle.com/datasets/alexandervc/open-problems-single-cell-perturbations-submitsetc , as well as collected some by the others: https://www.kaggle.com/datasets/alexandervc/open-problems-2-blends, https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-collection\n\n\nAfter the challenge - most of the submits by our team are stored in the Kaggle datasets: https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-etc/ and models by Antonina Dolgorukova:  https://www.kaggle.com/datasets/antoninadolgorukova/op2-submits-and-yoof , https://www.kaggle.com/datasets/antoninadolgorukova/op2-submissions . Hope it can be useful for researchers who work on that important topic.\n\nSelected final submit (0.558 public, 0.745 private) is the notebook version 15 here or can be downloaded from the dataset: \nhttps://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-etc?select=LB558_745_U900final_blend_558_02netsAntoninaOnFirst128Bcells.csv\n\n\n## Details : \n\n### Poor CV-LB correspondence - problem for blend\n\n    Due to absence of good CV - LB correspondence:  blend is not an easy task - becuase standard strategy to choose weights looking on CV suffers since CV does not match LB.\n\n    CV-LB correspondence was not completely absent for individual models,\n    and all individual models were chosen such that they increase both CV and LB scores (at least to what extent it is possible).\n\n    But comparing CV from models of different classes - like Pyboost and neural networks - seemed to us hopeless.\n    We observed quite different CV scores for the same LB scores.\n\n    More details on poor CV-LB correspondence here: TODO - LINK\n    \n### Solution ingredient 1: controlling high diversification       \n    \n     The standard idea: more diversity - better blend\n     With clear rationale - independent flucations around ground truth better cancel each other\n\n     Thus we aimed to include in blend the models which are diverse enough.\n     Even if CV and public LB scores are misleading for private LB - the diversification will still work for private LB, because it does not rely on scores !\n     And thus we hoped  ensuring high diversification we can be more stable for private LB.\n\n     It is not so clear what is good measure of the diversity for the MULTI-target task.\n     We have chosen the simple one:\n     Consider two models. For each target take prediction of model1 and  prediction of model2  and compute their Pearson correlation.\n     And take average of these correlations over all 18211 targets.\n     We take that averaged correlation as our measure of diversity - the lower - the better.\n\n     We spent submits to understand how good is that measure.\n     Conclusion: models with correlation score  0.8-0.9 - consistenly give substantial uplift in blend (about +0.01 - 0.006)\n     We have seen that for phenomena for quite different models.\n     It gave us confidence that we can rely on that. \n     So strategy - blending models with correlation score 0.8-0.9 should hopefully work quite well even on private LB.\n     \n     \n### Solution ingredient 2: compare models of different origin by public LB, not by CV       \n     \n    Despite the usual data science mantra \"trust your CV\" - it is not always the case.\n    Our (not-easy) decision was to trust LB, not CV when we compared models of different nature - like NN vs boostings.\n    Aftermath shows - it was correct decision: the correlation between public and private LB scores is 0.98 for our submits\n    See analyais: \n    https://www.kaggle.com/code/alexandervc/op2-cv-vs-lb-analysis\n    \n    The motivation for that decision was the following:\n    First it seemed quite hopeless to match CV scores for models of diverse origin.\n    Second gradually it became clear that the key problem is that  B-cells and Myeloid cells (leaderbord) are quite different from train cell types,\n    locallay we see good model for NK-cells would not be good for T-cells CD8+ and so on. \n    And so if key difference is in cell types - public LB should quite correspond private LB,\n    because both of them contain the same cell types  (only drugs are different).\n    As said above - it happened indeed to be true - public LB quite mateches private LB - 0.98 correlation.\n\n\n### Still there is an issue - how to choose weights \n\n    So we decided to blend looking mainly on LB score.\n    But still there is an isuse: having models with different scores, what weight to give to one model, what to another ? \n    Making decision based too much on public LB, may lead to overfit and poor performance on private LB and moreover it requires many submits which we did not have. \n\n    The lucky case - two or more submits have approximately same score - so we can just average them with the equal weight. \n    We had a luck to find core blend scheme where we  only needed to blend models with almost the same LB scores:\n\n\n### Solution ingredient 3: blend with the only weight = 0.5 and  SAME scored models \n\n    Our keys models: several versions of pyboost (0.574,0.577,0.584), catboost (0.584), NLP-NN (0.574), TargetEnc NN Ensemble (0.566) and some other models.\n    All these models are diverse enough to expect good uplift from the blend.\n    The description can be found here: TODO - LINK\n    It is not immediate how to organize blend of these models since we canot rely on CV.\n    \n    During the competition gods did not send us luck, we relied only on hard work, but there was just one exception:\n    The core blend scheme - \"gods send us a bit of luck\" -  was organized as follows :\n\n    Pyboost 0.584, and Catboost 0.584 -- average --> 0.575 - almost the same as our NLP NN and very diverse from it (submit - nb V6 here )\n    0.575  + 0.574 NLP-NN -- average --> 0.566  -- exactly the same as our TargetEnc NNs ensemble and very diverse from it  (submit - nb V7 here )\n    0.566  + 0.566 TE-NN  -- average --> 0.559  (submit - nb V8 here )\n    That forms a basis for further (quite small) improvements.\n    \n    So we were lucky that we found:\n    submits such they have same LB score - and so there was no question how to choose weights.\n    \n    On the other hand models are quite diverse - so it gave us hope to some stability in case of shake-up, which we were very afraid of.\n    \n\n### Further small improvements/experiments: \n\n    We improved 0.559 a bit to 0.558 (0.744/0.745 private) blending with some more Pyboost and NN models.\n    Though these imporovements might be inessential.\n    \n    Weighted blend: 0.8 * 0.559 + 0.2 (others pyboosts (0.574,0.577) + nerual net 0.587) -> 0.558 ( 0.744 on private - actually the best private). See nb V11 here.\n\n    Based on some previous analysis it was natural to think that pyboost predictions \n    work better for Myeloid cells, and NN for B-cells.\n    We tried to increase weight of pyboosts on myeloid cells - but public score goes a bit down to 0.559.\n    We blended with 3 more out TargetEnc NN (only on B-cells):\n    0.8 * 0.558 + 0.2 ( three more TargetEnc NN (0.569-0.571) ) -> 0.558 (better)  (and 0.745 - worse private score) (nb V 15 here).\n    \n    \n    So finally selected : 0.558 (public) 0.745 (private) - notebook version 15 here \n    That submission was selected for final evaluation. It was not the best one on private, but almost the best one.\n    It incorporated many diverse models - so we hoped it can be stable in case of shake-up.\n    But there was practically no shake-up for reasonable models.\n    Aftermath analysis shows 0.98 correlation between public and private LB scores,\n    see: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939\n    \n### Links, references\n\n    TODO: \n    Pyboost models:\n    Catboost models:\n    NLP-NN model\n    TargetEnc NN models: \n\n\n## Further remarks and ideas\n\n\n#### Blend with weights by variance and correlation\n\n\n    It is well-known from classical statistics that optimal weights for blend of two independent variable with the same mean can be be given by the formulas:\n        weight1 = (variance2 )/(variance1 + variance2 )\n        weight2 = (variance1 )/(variance1 + variance2 )\n    And for correllated case:\n        weight1 = (variance2 - C)/(variance1 + variance2 - 2C)\n        weight2 = (variance1 - C)/(variance1 + variance2 - 2C)\n        \n    Which have simple rationale - higher variance - means less confidence, and so the weight should be lower.\n    See e.g. https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459001#2547400\n    \n    The theorerical assumptions for these formulas not met in practice like current challenge,\n    still one can try to modify these formulas and use them to find blend weights.\n    \n    We have made a couple of attempts -  row-wisely and global-wisely, but not quite successful.\n    Nevertheless we might hope to find appropriate modifications in future.\n    \n#### Possible modificatiosn\n\n    TODO  - describe approach\n    \n","metadata":{}},{"cell_type":"markdown","source":"# Preliminaries","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\ncc = 0\nprint('Show first 15 files:')\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        cc += 1\n        if cc <= 15:\n            print(os.path.join(dirname, filename))\nprint()            \nprint('Total files count:', cc)\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-06T17:02:29.935101Z","iopub.execute_input":"2023-12-06T17:02:29.936280Z","iopub.status.idle":"2023-12-06T17:02:29.950873Z","shell.execute_reply.started":"2023-12-06T17:02:29.936238Z","shell.execute_reply":"2023-12-06T17:02:29.948987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Diversity analysis of the models in the core blend\n\nHere is an example to analyse the diveristy of the models - compute average target-wise correlation.\nThe general pracitice - more diversified models we have - the better blend results.\n\nThe proposed diversity score (average target-wise correlation) between the  four core models  models is at most 0.85. \n\nVarious experiments during the competition suggest that it is low enough to provide good uplift (0.01-0.006) in blend.\n\nThus it gave us some confidence in our model - and hope it can survive the shake-up - which we were quite afraid of.\n\nMore detailed diversity analysis can be found in: \nhttps://www.kaggle.com/alexandervc/op2-correlations-submits-u900-team\n\nhttps://www.kaggle.com/code/alexandervc/op2-submits-correlations-and-analysis\n\n","metadata":{}},{"cell_type":"code","source":"%%time\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n##################################################################################\n# Specify list of submit files to be processed and short aliases for their names \n##################################################################################\n\n\nlist_fn = []\nlist_ids = []\n\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Catboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\nlist_fn.append( fn )\nlist_ids.append('LB584_Pyboost')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\nlist_fn.append( fn )\nlist_ids.append('LB574_NN_NLP')\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\nlist_fn.append( fn )\nlist_ids.append('LB566_NN_TargetEnc_Ensemble')\n\n                \nprint(len(list_fn))    \nprint(len(list_ids))    \nprint(list_ids)\nprint(list_fn)\nprint()\n\n\n\n##################################################################################\n# Load stored submits specified by the list_fn \n##################################################################################\n\nprint('Start load submit files:')\n\n#print(list_ids)\nlist_df = []\n\ni_blend = 0\nfor i,fn in enumerate(list_fn): \n#     print(i, fn)\n    df = pd.read_csv(fn, index_col = 'id')\n    print(i,df.shape, fn , np.round( [df.iloc[0,0],  df.iloc[0,1],  df.iloc[1,0] ,  df.iloc[1,1] ], 4) , df.columns[0],   df.columns[1]   )\n    #display(df.head(2) )\n    list_df.append(df)\n    if i > 0: # not blend the first one\n        if i_blend == 0:\n            df_blend = df.copy()\n        else:\n            df_blend = (df_blend * i_blend  + df )/ (i_blend + 1)\n        i_blend += 1\n        \nflag_include_average_blend_of_all = False        \nif flag_include_average_blend_of_all:        \n    list_ids.append('Blend')  \n    list_df.append(df_blend )    \n\n\n##################################################################################\n# Prepare fast correlations computation - standard scaler each submit dataframe\n##################################################################################\n\n# %%time\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nlist_np = []\nif 1:\n    for k in range(len(list_df)): # dict_df:\n        df = list_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\nelse:\n    for k in dict_df:\n        df = dict_df[k] \n        d = scaler.fit_transform(df)\n        list_np.append(d)\n    \nprint(len(list_np)) \n\n\n##################################################################################\n# Compute correlations\n# I.e. average column-wise correlations\n# Due to standard scaling on the previous step: correlation = produce + average\n##################################################################################\n\nflag_plot_hists = False\n\nif flag_plot_hists:\n    fig = plt.figure(figsize = (20,10))\n    \ndf_stat = pd.DataFrame()\ndf_corr_averages_for_submits = pd.DataFrame()\ni_total = 0\nfor i0,d0 in  enumerate(list_np):\n    for i1,d1 in  enumerate(list_np):\n        i_total += 1\n        v = np.mean( d0*d1, axis = 0 )\n        if i0<i1:\n            if flag_plot_hists:\n                plt.hist(v, label = str(i0)+ ' vs ' + str(i1))\n        dt = pd.Series(v).describe().to_frame()\n        dt.columns = [str(i0)+ ' vs  ' + str(i1)]\n        df_stat = pd.concat( [df_stat, dt] , axis = 1)\n        #print(  )\n        # c = np.mean( np.abs(v) )\n        c = np.mean( (v) )\n        if i_total<3:\n            print(c)\n        df_corr_averages_for_submits.loc[i0,i1] = c\n        \ndisplay(df_stat)        \n# plt.legend(fontsize = 12 )\nif flag_plot_hists:\n    plt.grid()\n    plt.title('Distribution of correlations (over genes) for each prediction pair ', fontsize = 20)\n    plt.show()        \n    \ndf_corr_averages_for_submits.columns = [t.replace('_',' ') for t in  list_ids]\ndf_corr_averages_for_submits.index = [t.replace('_',' ') for t in  list_ids]\n\nprint('Show 3x3 part of the correlation matrix:')\ndisplay(df_corr_averages_for_submits.iloc[:5,:5].round(2) )  \n\n##################################################################################\n# Show clustermap\n##################################################################################\n\nsns.clustermap( df_corr_averages_for_submits,figsize = (20,15),  annot=True, cmap='coolwarm')\nplt.show()\n\n# sns.clustermap( df_corr_averages_for_submits,figsize = (20,12), annot=False, cmap='coolwarm')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-06T15:51:13.519163Z","iopub.execute_input":"2023-12-06T15:51:13.520738Z","iopub.status.idle":"2023-12-06T15:51:36.445309Z","shell.execute_reply.started":"2023-12-06T15:51:13.520683Z","shell.execute_reply":"2023-12-06T15:51:36.443952Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blends","metadata":{}},{"cell_type":"markdown","source":"# Core blend - 0.559 on public LB ( 0.746 private LB)\n\n    Notebook version 8 here \n\n    The main blend - \"Gods send us a bit of luck\" was organized as follows :\n\n    Pyboost 0.584, and Catboost 0.584 -- average --> 0.575 - almost the same as our NLP NN (nb V6 here )\n    0.575  + 0.574 NLP-NN -- average --> 0.566  -- exactly the same as our TargetEnc NNs ensemble   (nb V7 here )\n    0.566  + 0.566 TE-NN  -- average --> 0.559  (nb V7 here )\n    That forms a basis for further (quite small) improvements.\n    \n    So we were lucky that we found:\n    submits such they have same LB score - and so there was no question how to choose weights.\n    \n    On the other hand models are quite diverse - so it gave us hope to some stability in case of shake-up, which we were very afraid of.\n    \nThe submit  is the notebook version 8 here or available here:\nhttps://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-etc/data?select=LB559_blendBasic_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv    \n","metadata":{}},{"cell_type":"code","source":"list_df = []\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf1 = df.copy()\nlist_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\nfn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf2 = df.copy()\nlist_df.append(df)\n\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf3 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf4 = df.copy()\nlist_df.append(df)\n\n\n#df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n#df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n# df = list_df[0]\n# n = len(list_df)\n# print( n )\n# for k in range(1,len(list_df)):\n#     df += list_df[k]\n# df /= n    \n\ndf = ( ((df1+df2)/2 + df3)/2 +df4)/2\n\ndisplay(df.head(2) )\n\n#df.to_csv('submission_blend_Tonya_569_570_572.csv')\n#df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n# df.to_csv('submission_catboost_and_pyboost_584_average.csv')\n# df.to_csv('submission_NLP574_and_575_catboost_pyboost_both584.csv')\ndf.to_csv('submission_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv')","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:27:02.706904Z","iopub.execute_input":"2023-12-05T18:27:02.707452Z","iopub.status.idle":"2023-12-05T18:27:40.043519Z","shell.execute_reply.started":"2023-12-05T18:27:02.707410Z","shell.execute_reply":"2023-12-05T18:27:40.042096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blend 0.558, 0.744 - best on private - gold zone  (but not selected) \n\n    \n    We improved 0.559 a bit to 0.558, 0.744 blending with some more Pyboost and NN models.\n    \n    Weighted blend: 0.8 * 0.559 + 0.2 (others pyboosts (0.574,0.577) + nerual net 0.587) -> 0.558 ( 0.744 on private - actually the best private).\n\n\nThe submit is in version 11 of the present notebook or can be downloaded here:     https://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-etc/data?select=LB558_blend_559_02newpyboostAnd587.csv\n    \n","metadata":{}},{"cell_type":"code","source":"%%time\nlist_df = []\nfn = '/kaggle/input/open-problems-2-submits-etc/LB559_blendBasic_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf559 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-collection/LB574_Pyboostmaxdepth12ntrees2000lr001subsample1colsample035ncomponents50T8T8b7t17_MadrisMillerBasedAlex_nbV8.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf1 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-collection/LB577_PyboostCVRandom5_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50_NikolenkoPubl.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf2 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPpseudo50ctstratmrrmsetfsmilesv_nbV2.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf3 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPblendown.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf4 = df.copy()\nlist_df.append(df)\n\nprint();print()\nprint('--------------------------------------------------------------------------------------------')\nprint(); print();\n\ndf = 0.8*df559 + 0.2*( (df1+df2)/2*0.9+0.1*(df3+df4)  )\n\ndisplay(df.head(2) )\ndf.to_csv('submission_blend_559_02newpyboostAnd587.csv')","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:27:40.046231Z","iopub.execute_input":"2023-12-05T18:27:40.046667Z","iopub.status.idle":"2023-12-05T18:28:24.251881Z","shell.execute_reply.started":"2023-12-05T18:27:40.046625Z","shell.execute_reply":"2023-12-05T18:28:24.250671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 0.558 , 0.745 - selected submit for final evaluation (almost best on private)\n\n    Based on some previous analysis it was natural to think that pyboost predictions \n    work better for Myeloid cells, and NN for B-cells.\n    We tried to increase weight of pyboosts on myeloid cells - but public score goes a bit down to 0.559.\n    We blended with 3 more out TargetEnc NN (only on B-cells):\n    0.8 * 0.558 + 0.2 ( three more TargetEnc NN (0.569-0.571) ) -> 0.558 (better)  (and 0.745 - worse private score) (nb V 15 here).\n    \n    So finally selected : 0.558 (public) 0.745 (private) - notebook version 15 here \n    That submission was selected for final evaluation. It was not the best one on private, but almost the best one.\n    It incorporated many diverse models - so we hoped it can be stable in case of shake-up.\n    But there was practically no shake-up for reasonable model.\n    Aftermath analysis shows 0.98 correlation between public and private LB scores,\n    see: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/458939\n\nSelected final submit (0.558 public, 0.745 private) is the notebook version 15 here or can be downloaded from the dataset: \nhttps://www.kaggle.com/datasets/alexandervc/open-problems-2-submits-etc?select=LB558_745_U900final_blend_558_02netsAntoninaOnFirst128Bcells.csv\n\n","metadata":{}},{"cell_type":"code","source":"%%time\nlist_df = []\nfn = '/kaggle/input/open-problems-2-submits-etc/LB558_blend_559_02newpyboostAnd587.csv'\nprint(fn)\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndf558 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_T8_Treg_3kmeans_Antonina.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf1 = df.copy()\nlist_df.append(df)\n\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf2 = df.copy()\nlist_df.append(df)\n\nfn = '/kaggle/input/open-problems-2-submits-etc/LB572_ave_blend_T4_T8_20ep_augm50_s1_0.1.csv'\ndf = pd.read_csv(fn, index_col = 'id')\nprint(df.shape)\ndisplay(df.head(2))\ndf3 = df.copy()\nlist_df.append(df)\n\nprint();print()\nprint('--------------------------------------------------------------------------------------------')\nprint(); print();\n\ndf_nets = (df1+df2+df3)/3\n\ndf_blend = df558.copy()\ndf_blend[:128] = ( df_blend[:128]*0.8 + df_nets[:128]*0.2) \n\ndisplay(df_blend.head(2) )\ndf_blend.to_csv('submission_blend_558_02netsAntoninaOnFirst128Bcells.csv')","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:54:31.837580Z","iopub.execute_input":"2023-12-05T18:54:31.838975Z","iopub.status.idle":"2023-12-05T18:55:09.483870Z","shell.execute_reply.started":"2023-12-05T18:54:31.838934Z","shell.execute_reply":"2023-12-05T18:55:09.482764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n\n# public_ids = [\n#     'LSM-43216', 'LSM-1050', 'LSM-45849', 'LSM-42800', 'LSM-1131', 'LSM-6335', 'LSM-1211',\n#     'LSM-45239', 'LSM-1130', 'LSM-45786', 'LSM-5199', 'LSM-45281',\n#     'LSM-6324', # 'ACY-1215' -> 'Ricolinostat'\n#     'LSM-3309', 'LSM-1056', 'LSM-45591', 'LSM-46203', 'LSM-5662',\n#     'LSM-47134',  # 'SB-2342' -> '5-(9-Isopropyl-8-methyl-2-morpholino-9H-purin-6-yl)pyrimidin-2-amine  '\n#     'LSM-45637', 'LSM-1127', 'LSM-46971', 'LSM-1172', 'LSM-46042', 'LSM-1101', 'LSM-45758',\n#     'LSM-5218', 'LSM-2287', 'LSM-1014',\n#     'LSM-1040', #  'fostamatinib' -> 'Tamatinib'\n#     'LSM-1476;LSM-5290',\n#     'LSM-45680',  # 'basimglurant' -> 'RG7090'\n#     'LSM-4349',  # '5-iodotubercidin' -> 'IN1451'\n#     'LSM-3425', 'LSM-45806',\n#     'LSM-45616',  # 'SB-683698' -> 'TR-14035'\n#     'LSM-1055',\n#     'LSM-43281',  # 'C-646' -> 'STK219801'\n#     'LSM-5690', 'LSM-1155', 'LSM-2499',\n#     'LSM-2382',  # 'JTC-801' -> 'UNII-BXU45ZH6LI'\n#     'LSM-45220', 'LSM-1037', 'LSM-1005', 'LSM-1180', 'LSM-36812',\n#     'LSM-45924',  # 'filgotinib' -> 'GLPG0634'\n#     'LSM-2013',  # 'TL-HRAS-61' -> TL_HRAS26'\n#     'LSM-4738'\n# ]\n\n# fn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\n# df_de_train = pd.read_parquet(fn)# , index_col = 0)\n# print(df_de_train.shape)\n# display(df_de_train.head(2) )\n\n\n# # %%time\n# fn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\n# df_id_map = pd.read_csv(fn,index_col = 0)\n# print(df_id_map.shape)\n# display(df_id_map.head(2))\n\n# lst = list( df_de_train['sm_name'][ df_de_train['sm_lincs_id'].isin(public_ids ) ].unique() )\n# mask_pub_samples = df_id_map['sm_name'].isin(lst )\n# pub_ix = list( df_id_map[mask_pub_samples].index )\n# print( pub_ix )\n# print(len(pub_ix))\n\n# fn = '/kaggle/input/open-problems-2-blends/LB531_Push_nbV27.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df531 = df.copy()\n\n# # df2 = df.copy()\n# # for k in pub_ix:\n# #     df2.iloc[k,:] = (df531.iloc[k,:] +  df.iloc[k,:])/2  \n# # df2    ","metadata":{"execution":{"iopub.status.busy":"2023-11-29T19:56:30.391512Z","iopub.execute_input":"2023-11-29T19:56:30.391849Z","iopub.status.idle":"2023-11-29T19:56:36.841125Z","shell.execute_reply.started":"2023-11-29T19:56:30.391822Z","shell.execute_reply":"2023-11-29T19:56:36.840203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# list_df = []\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB559_blendBasic_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv'\n# print(fn)\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df559 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-collection/LB574_Pyboostmaxdepth12ntrees2000lr001subsample1colsample035ncomponents50T8T8b7t17_MadrisMillerBasedAlex_nbV8.csv'\n# print(fn)\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-collection/LB577_PyboostCVRandom5_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50_NikolenkoPubl.csv'\n# print(fn)\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPpseudo50ctstratmrrmsetfsmilesv_nbV2.csv'\n# print(fn)\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPblendown.csv'\n# print(fn)\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df4 = df.copy()\n# list_df.append(df)\n\n# print()\n# print()\n# print('--------------------------------------------------------------------------------------------')\n# print(); print();\n\n# df_boost =  ( (df1+df2)/2*0.9+0.1*(df3+df4)  )\n# df_blend = 0.8*df559 + 0.2*df_boost \n\n# df_blend[128:] = 0.5*df559[128: ] + 0.5* df_boost[128:] \n# # ( (df1+df2)/2*0.9+0.1*(df3+df4)  )\n\n\n# display(df_blend.head(2) )\n# df_blend.to_csv('submission_blend_559_02newpyboostAnd587_BoostUp05Mielod.csv')\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-29T19:56:57.093832Z","iopub.execute_input":"2023-11-29T19:56:57.094160Z","iopub.status.idle":"2023-11-29T19:57:23.638206Z","shell.execute_reply.started":"2023-11-29T19:56:57.094133Z","shell.execute_reply":"2023-11-29T19:57:23.637180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df2 = df_blend.copy()\n# for k in pub_ix:\n#     df2.iloc[k,:] = (df531.iloc[k,:]*0.8 +  0.2*df_blend.iloc[k,:])\n# display(df2.head(2) )\n# df2.to_csv('submission_public531_blend_559_02newpyboostAnd587_BoostUp05Mielod.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-11-29T19:58:03.275066Z","iopub.execute_input":"2023-11-29T19:58:03.275383Z","iopub.status.idle":"2023-11-29T19:58:09.592718Z","shell.execute_reply.started":"2023-11-29T19:58:03.275359Z","shell.execute_reply":"2023-11-29T19:58:09.591823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list_df = []\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_NNohe8985_E100_OHE_Dr_CT_SB1_Tonya_nbV107.csv'\n# fn = '/kaggle/input/open-problems-2-submits-collection/LB574_Pyboostmaxdepth12ntrees2000lr001subsample1colsample035ncomponents50T8T8b7t17_MadrisMillerBasedAlex_nbV8.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_NNohe.8954_E100_OHE_Dr_CT_SB1_Tonya_nbV106.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB577_PyboostCVRandom5_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50_NikolenkoPubl.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n\n# fn = '/kaggle/input/open-problems-2-submits-collection/LB720_JAXautoencoder_VENDEKAGONLABS_nbV4.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB559_blendBasic_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df4 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPpseudo50ctstratmrrmsetfsmilesv_nbV2.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df5 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB587_ErshovSCPblendown.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df6 = df.copy()\n# list_df.append(df)\n\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# # df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n\n# df = list_df[0]\n# n = len(list_df)\n# print( n )\n# for k in range(1,len(list_df)):\n#     df += list_df[k]\n# df /= n    \n\n# # df = ( ((df1+df2)/2 + df3)/2 +df4)/2\n\n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# # df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n# # df.to_csv('submission_catboost_and_pyboost_584_average.csv')\n# # df.to_csv('submission_NLP574_and_575_catboost_pyboost_both584.csv')\n# df.to_csv('submission_NNohe_two_averaged_v106_107.csv')\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list_df = []\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_NNohe8985_E100_OHE_Dr_CT_SB1_Tonya_nbV107.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_NNohe.8954_E100_OHE_Dr_CT_SB1_Tonya_nbV106.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\n# # df = pd.read_csv(fn, index_col = 'id')\n# # print(df.shape)\n# # display(df.head(2))\n# # df3 = df.copy()\n# # list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\n# # df = pd.read_csv(fn, index_col = 'id')\n# # print(df.shape)\n# # display(df.head(2))\n# # df4 = df.copy()\n# # list_df.append(df)\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# # df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n\n# df = list_df[0]\n# n = len(list_df)\n# print( n )\n# for k in range(1,len(list_df)):\n#     df += list_df[k]\n# df /= n    \n\n# # df = ( ((df1+df2)/2 + df3)/2 +df4)/2\n\n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# # df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n# # df.to_csv('submission_catboost_and_pyboost_584_average.csv')\n# # df.to_csv('submission_NLP574_and_575_catboost_pyboost_both584.csv')\n# df.to_csv('submission_NNohe_two_averaged_v106_107.csv')\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list_df = []\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n# list_df.append(df)\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB566_blend2_corr_cl_plus_scores_Antonina.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df4 = df.copy()\n# list_df.append(df)\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# # df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n# # df = list_df[0]\n# # n = len(list_df)\n# # print( n )\n# # for k in range(1,len(list_df)):\n# #     df += list_df[k]\n# # df /= n    \n\n# df = ( ((df1+df2)/2 + df3)/2 +df4)/2\n\n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# # df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n# # df.to_csv('submission_catboost_and_pyboost_584_average.csv')\n# # df.to_csv('submission_NLP574_and_575_catboost_pyboost_both584.csv')\n# df.to_csv('submission_Tonya566and566NLPandBoosts574_and_575_catboost_pyboost_both584.csv')\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list_df = []\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB584_CATBnoCD8WO3RandSamplesFull_tsvd30_modelCATB_NI250_MD6_LR0.03_SS1_CS0.5_encQuantileEncoder_quantile0.8_Dr_CT_SB1_Tonya.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# # fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB584_PyboostStanPublic_max_depth10_ntrees5000_lr001_subsample1_colsample035_n_components50.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# # df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n# df = list_df[0]\n# n = len(list_df)\n# print( n )\n# for k in range(1,len(list_df)):\n#     df += list_df[k]\n# df /= n    \n    \n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# # df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n# df.to_csv('submission_catboost_and_pyboost_584_average.csv')\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list_df = []\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md7_Y_submit_tsvd70_CATB_QuantileEncoderCompoundCellType_Dr_CT_SB1_Tonya.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n# list_df.append(df)\n\n# # fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# fn = '/kaggle/input/open-problems-2-submits-etc/CATB_iter5000_md6_Y_submit_CATBtsvd70_QuantileEncoder80_Dr_CT_SB1_Tonya.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n# list_df.append(df)\n\n\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# # df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n# df = list_df[0]\n# n = len(list_df)\n# print( n )\n# for k in range(1,len(list_df)):\n#     df += list_df[k]\n# df /= n    \n    \n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# df.to_csv('submission_two_catboost5000_md67_averaged.csv')\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-11-26T20:53:31.988365Z","iopub.execute_input":"2023-11-26T20:53:31.988750Z","iopub.status.idle":"2023-11-26T20:53:57.380361Z","shell.execute_reply.started":"2023-11-26T20:53:31.988721Z","shell.execute_reply":"2023-11-26T20:53:57.378877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fn = '/kaggle/input/open-problems-2-submits-etc/LB570_ave_blend_0.571_T4_T8_20ep_augm50_s1_0.1.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB569_ave_blend_T4_T8_3kmeans_sep_f_20ep_augm50_s1_0.1.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB572_blend_v39_095_v58_0603_005_Antonina.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df4 = df.copy()\n\n\n# fn = '/kaggle/input/open-problems-2-blends/LB547_Angelova_SCPEnsemblingSubmissions_nbV15.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df5 = df.copy()\n\n\n\n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# #df = ( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3 \n# df = df5*0.7+ 0.3*(( ( df1 + df2)/2 *0.7 + df3*0.3)*0.7 + df4*0.3) \n# display(df.head(2) )\n\n# #df.to_csv('submission_blend_Tonya_569_570_572.csv')\n# #df.to_csv('submission_blend_Tonya_569_570_572_574DimaAlex.csv')\n# df.to_csv('submission_blend_pub547_Tonya_569_570_572_574DimaAlex.csv')\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_sameprm574v1NLPregr_Rudenko_nbV122.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_sameprm574v2NLPregr_Rudenko_nbV123.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n\n# df = ( df1 + df2 + df3 ) /3 \n# display(df.head(2) )\n\n# df.to_csv('submission_574NLPregressor_and_two_more_runs_with_same_prms.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-25T21:14:48.692197Z","iopub.execute_input":"2023-11-25T21:14:48.692565Z","iopub.status.idle":"2023-11-25T21:14:58.000896Z","shell.execute_reply.started":"2023-11-25T21:14:48.692534Z","shell.execute_reply":"2023-11-25T21:14:57.999489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_sameprm574v1NLPregr_Rudenko_nbV122.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df1 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LBXXX_sameprm574v2NLPregr_Rudenko_nbV123.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df2 = df.copy()\n\n# fn = '/kaggle/input/open-problems-2-submits-etc/LB574_NLPmultiplex25_DimaAlex.csv'\n# df = pd.read_csv(fn, index_col = 'id')\n# print(df.shape)\n# display(df.head(2))\n# df3 = df.copy()\n\n# df = ( df1 + df2 + df3 ) /3 \n# display(df.head(2) )\n\n# df.to_csv('submission_574NLPregressor_and_two_more_runs_with_same_prms.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-25T21:15:48.390761Z","iopub.execute_input":"2023-11-25T21:15:48.391153Z","iopub.status.idle":"2023-11-25T21:15:55.136083Z","shell.execute_reply.started":"2023-11-25T21:15:48.391123Z","shell.execute_reply":"2023-11-25T21:15:55.135166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}