{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nWe will try to find(analyse) good predictors for CD** protein levels - based in biological knowledge.\nI.e. we search for genes which RNA levels (i.e. features of CITEseq task)\nare correlated/good predictors for the CD** proteins (targets of the CITEseq task).\n\nWe widely use Nature paper recommended by orgs - see table on Figure 3a. \nhttps://www.nature.com/articles/ncb3493    \nHuman haematopoietic stem cell lineage commitment is a continuous process  \nPublished: 20 March 2017\n\n\n## Versions\n\n#### 15'CD14', 'LY96' and 'TLR4' - have no correlation\n\nProposed by Alexander Romanishin\n\n#### 14 'CD56', 'CD94' and RORC - conlusion RORC is absent, CD56, CD94 proteins - have some correlation 0.16 \n\nProposed by Alexander Romanishin\nBased on: \nhttps://www.frontiersin.org/articles/10.3389/fimmu.2019.02078/full\n\n#### 13 CD69 vs PRKCA - conclusion - expression is to small to be confident (correlation near zero) \n\nProtein kinase C alpha (PKCα) is an enzyme \n\nhttps://en.wikipedia.org/wiki/PKC_alpha\n\nProposed by Alexander Romanishin\n\n#### 12 CD44 with ERG gene - correlation - 0.2 - not bad \n\n    Not completely bio motivated. \n    \n\n    \n\n#### 11 Check correlation of CD44 with HAS1, HAS2, HAS3 -  Conclusion - HAS1,2 - not in data, HAS3 is low expressed, so probably no reliable conclusions (correlation is almost zero). \n\n    In addition, the interaction of HA with the leukocyte receptor **CD44**  \n    is important in tissue-specific homing by leukocytes, and overexpression of HA receptors has been correlated with tumor metastasis. HAS1 is a member of the newly identified vertebrate gene family encoding putative hyaluronan synthases, and its amino acid sequence shows significant homology to the hasA gene product of Streptococcus pyogenes, a glycosaminoglycan synthetase (DG42) from Xenopus laevis, and a recently described murine hyaluronan synthase.[6]\n\nhttps://en.wikipedia.org/wiki/HAS1\n\nProposed by Alexander Romanishin \n\n\n\n#### 10 cosmetic chages\n\n#### 9 for CD45RA from immature cells part of table - MPO SPINK2 ENO1 ATP884 SELL are correlated around 0.3  \n\n#### 8 for CD45RA from neutrophils part of table - CEBPA CEPBD CSTG are correlated around 0.3\n\n\n#### 7 CD45RA is ANTI-correllated (0.2) with HDC and GATA2 - \n    \n    which is in correpondence that CD45RA(-) is marker for cell type Eosinophil/basophil/mast cell progenitors\n    where these genes are expressed\n    (according to paper)\n    \n    \n#### 6 - mild changes\n    \n    'CD45RA, CD45, CD45RO , which are coded by the SAME gene CD45 = ENSG00000081237 = PTPRC (HGNC Symbol)\n    It seems only main CD45 is correlated wtih it its gene and the other two proteins\n    \n\n#### 5 - continue to use the Nature paper - continue with table on Figure 3a.\n\n    Conclusion - for CD45RA, we have IRF8 and TCF4 correlated around 0.3 but all the other correlations are quite lower. \n\n    Note we have 3 target proteins:  'CD45RA, CD45, CD45RO , which are coded by the SAME gene CD45 = ENSG00000081237 = PTPRC (HGNC Symbol)\n\n    Monocyte/dendritic cell progenitors, 'CD45RA' and related CD45, CD45RO , Immunophenotype: CD10–CD45RAhiCD135hi\n\n    Genes: SCT, TGFBI, LGMN TFs: IRF8, IRF7, TCF4\n    Immunophenotype: CD10–CD45RAhiCD135hi\n    \n    CD135 - not in included in CITEseq targets dataset, so skip it \n    \n    \n\n#### 4 use Nature paper recomenned by orgs - for some targets like CD71 - see quite (0.5) correlated genes \n\nhttps://www.nature.com/articles/ncb3493   \nHuman haematopoietic stem cell lineage commitment is a continuous process  \nPublished: 20 March 2017\n\n\n\n#### 2,3 - cosmetic changes\n\n#### 1 - gene \"SET\" not that much bad for CD47\n\n    Proposed by Alexander Romanishin, [10/8/2022 4:39 PM]\n\n    Paper:\n    Berkovits BD, Mayr C. Alternative 3' UTRs act as scaffolds to regulate membrane protein\n    localization. Nature. 2015 Jun 18;522(7556):363-7.\n\n    Idea:\n    HuR и SET genes are responsible for transfer and binding of the CD47 proteins - so one may look on them.\n\n    Conclusion:\n    At least SET gene is not worse correlated with CD47 protein than CD47 RNA itself with protein\n    Correlation is not that much big - about 0.1 but nevertheless.\n    For other gene - HuR - correlation is low. But HuR and SET are correlated among themselves - level about 0.25\n    \n    ","metadata":{}},{"cell_type":"markdown","source":"# Install/Import ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-07T12:16:25.813415Z","iopub.execute_input":"2022-11-07T12:16:25.813828Z","iopub.status.idle":"2022-11-07T12:16:25.825865Z","shell.execute_reply.started":"2022-11-07T12:16:25.813794Z","shell.execute_reply":"2022-11-07T12:16:25.824888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install scanpy\nimport scanpy as sc\nimport anndata\n\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:16:25.827594Z","iopub.execute_input":"2022-11-07T12:16:25.828152Z","iopub.status.idle":"2022-11-07T12:17:12.026570Z","shell.execute_reply.started":"2022-11-07T12:16:25.828119Z","shell.execute_reply":"2022-11-07T12:17:12.025124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\nprint('8.2G RAM consumed')\ndisplay( df_cite_train_x.head() )\n\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndf_cite_train_y.head()\nprint('8.3G RAM consumed in addtion to 8.2 for features part ')\n\nif 0:\n    df_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\n    display( df_cite_test_x.head() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:17:12.029285Z","iopub.execute_input":"2022-11-07T12:17:12.029682Z","iopub.status.idle":"2022-11-07T12:18:07.938340Z","shell.execute_reply.started":"2022-11-07T12:17:12.029617Z","shell.execute_reply":"2022-11-07T12:18:07.937080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"%%time\nlist_genes_names = [t.split('_')[1] for t in df_cite_train_x.columns ]\nprint(len(list_genes_names), list_genes_names[:10])\nlist_genes_ids = [t.split('_')[0] for t in df_cite_train_x.columns ]\nprint(len(list_genes_ids), list_genes_ids[:10])\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:07.940312Z","iopub.execute_input":"2022-11-07T12:18:07.941167Z","iopub.status.idle":"2022-11-07T12:18:07.968133Z","shell.execute_reply.started":"2022-11-07T12:18:07.941111Z","shell.execute_reply":"2022-11-07T12:18:07.966726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'ENSG00000066044' in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:07.970877Z","iopub.execute_input":"2022-11-07T12:18:07.971275Z","iopub.status.idle":"2022-11-07T12:18:07.981945Z","shell.execute_reply.started":"2022-11-07T12:18:07.971237Z","shell.execute_reply":"2022-11-07T12:18:07.979992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculations of the correlations","metadata":{"execution":{"iopub.status.busy":"2022-10-08T15:02:29.460937Z","iopub.execute_input":"2022-10-08T15:02:29.461385Z","iopub.status.idle":"2022-10-08T15:02:29.481832Z","shell.execute_reply.started":"2022-10-08T15:02:29.461348Z","shell.execute_reply":"2022-10-08T15:02:29.480641Z"}}},{"cell_type":"code","source":"df_new = pd.DataFrame()\ndf_new['CD44_protein'] = df_cite_train_y['CD44']\n\nlist_ids = ['ENSG00000196776', 'ENSG00000066044', 'ENSG00000119335' ]\nlist_names = ['CD44', 'LY96', 'TLR4' ]\nfor i,id1 in enumerate(list_ids): \n    gene_name1 = list_names[i]\n    IX = list_genes_ids.index(id1)\n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    df_new[gene_name1 + '_RNA'] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:07.983855Z","iopub.execute_input":"2022-11-07T12:18:07.984311Z","iopub.status.idle":"2022-11-07T12:18:08.096200Z","shell.execute_reply.started":"2022-11-07T12:18:07.984279Z","shell.execute_reply":"2022-11-07T12:18:08.094952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new.corr()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.098029Z","iopub.execute_input":"2022-11-07T12:18:08.098512Z","iopub.status.idle":"2022-11-07T12:18:08.118771Z","shell.execute_reply.started":"2022-11-07T12:18:08.098468Z","shell.execute_reply":"2022-11-07T12:18:08.117602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new.corr(method = 'spearman')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.120482Z","iopub.execute_input":"2022-11-07T12:18:08.121292Z","iopub.status.idle":"2022-11-07T12:18:08.183975Z","shell.execute_reply.started":"2022-11-07T12:18:08.121257Z","shell.execute_reply":"2022-11-07T12:18:08.182718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new.corr(method = 'kendall')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.185848Z","iopub.execute_input":"2022-11-07T12:18:08.186598Z","iopub.status.idle":"2022-11-07T12:18:08.336286Z","shell.execute_reply.started":"2022-11-07T12:18:08.186547Z","shell.execute_reply":"2022-11-07T12:18:08.334972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('At least SET gene is not worse correlated with CD47 protein than CD47 RNA itself with protein ')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.337892Z","iopub.execute_input":"2022-11-07T12:18:08.338339Z","iopub.status.idle":"2022-11-07T12:18:08.344261Z","shell.execute_reply.started":"2022-11-07T12:18:08.338294Z","shell.execute_reply":"2022-11-07T12:18:08.342799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Nature paper recommended by orgs:\n\nHuman haematopoietic stem cell lineage commitment is a continuous process\nPublished: 20 March 2017\nhttps://www.nature.com/articles/ncb3493\n\n\n##  Mk/Ery progenitor,  Immunophenotype: CD10– CD135– CD71+        \n\nME-2: Mk/Ery progenitor, low cell cycle activity       \nGenes: CSF1, TFR2, CNRIP1 TFs: KLF1, MYC, GATA1\n    \nME-1: Mk/Ery progenitor, high cell cycle activity    \nGenes: TK1, RRM2, ITGA2B TFs: MYBL2, KLF1, GATA1\n    \n## E-1: Erythroid progenitor  ,  Immunophenotype: CD10– CD135– CD71+ KEL+       \nGenes: CA1, HBB, LMNA TFs: KLF1, GATA1, NFIA  \n\n## E-2: Erythroid progenitor Immunophenotype: CD10–CD135–CD71+ KEL+\nGenes: HBB, AHSP, CA1 TFs: KLF1,HES6, E2F4\n\n## Mk: Megakaryocyte progenitor , Immunophenotype: CD34midCD38midCD10 – CD135–CD71+KEL-\n\nGenes: GP1BB, PLEK, ITGA2B TFs: GATA2, PBX1, MEIS1\n\n","metadata":{}},{"cell_type":"code","source":"# CD10 - as protein is absent in CITEseq targets \n# \tMME, CALLA, CD10, NEP, SFE, membrane metallo-endopeptidase, membrane metalloendopeptidase, CMT2T, SCA43\n# 'ENSG00000196549' - gene CD10\n\n# CD135 - also - no such protein in dataset targets\n\n# KEL (CD238) - also absent\n# Kell Metallo-Endopeptidase (Kell Blood Group) 2 3 5\n# CD238 2 3 5\n# ECE3 2 3 5\n#Kell Blood Group Glycoprotein 3 4\n#Kell Blood Group, Metallo-Endopeptidase 3\n\n# anti-human CD71\tCY1G4\tCCGTGTTCCTCATTA\tENSG00000072274\n\n# GP1BB, PLEK, ITGA2B TFs: GATA2, PBX1, MEIS1\n    \n\nprint( 'CD10' in list_genes_names, 'ENSG00000196549' in list_genes_ids, 'CD10' in df_cite_train_y.columns )\n\nprint( 'CD71' in list_genes_names, 'CY1G4' in list_genes_names , 'ENSG00000072274' in list_genes_ids, 'CD71' in df_cite_train_y.columns )\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.350076Z","iopub.execute_input":"2022-11-07T12:18:08.350738Z","iopub.status.idle":"2022-11-07T12:18:08.365569Z","shell.execute_reply.started":"2022-11-07T12:18:08.350681Z","shell.execute_reply":"2022-11-07T12:18:08.364169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = pd.DataFrame()\ndf_new['CD14_prot'] = df_cite_train_y['CD14'] # CD71 or TRFC \n\n#  CNRIP1 TFs: KLF1, MYC, GATA1\n#     ME-1: Mk/Ery progenitor, high cell cycle activity    \n# Genes: TK1, RRM2, ITGA2B TFs: MYBL2, KLF1, GATA1\n        \n\nlist_ids = ['ENSG00000196776', 'ENSG00000066044', 'ENSG00000119335']\nlist_names = ['CD14', 'LY96', 'TLR4']\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.367309Z","iopub.execute_input":"2022-11-07T12:18:08.368018Z","iopub.status.idle":"2022-11-07T12:18:08.467753Z","shell.execute_reply.started":"2022-11-07T12:18:08.367969Z","shell.execute_reply":"2022-11-07T12:18:08.466483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cms = df_new.corr()\nfor i in range(len(cms)):\n    cms.iloc[i,i] = np.nan\nplt.figure(figsize=(20,16) )\n#sns.heatmap(cms)\nax = sns.heatmap( cms[ ((cms >= 0) | (cms <= -0))  ], \n            cmap='rainbow', vmax=cms.max().max(), vmin=-cms.min().min(), linewidths=0.1,\n            annot=True, annot_kws={\"size\": 12}, square=True);\nax.xaxis.set_tick_params(labelsize=11)\nax.yaxis.set_tick_params(labelsize=11)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:08.469886Z","iopub.execute_input":"2022-11-07T12:18:08.470743Z","iopub.status.idle":"2022-11-07T12:18:09.069388Z","shell.execute_reply.started":"2022-11-07T12:18:08.470685Z","shell.execute_reply":"2022-11-07T12:18:09.068074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for method_corr in ['pearson', 'spearman','kendall']:\n    cm = df_new.corr(method = method_corr)\n    # print(cm)\n    plt.figure(figsize  = (20,5))\n    plt.bar(x = cm.index, height = cm.iloc[:,0], color = 'r')\n    plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n    plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n\n    plt.title(method_corr+  ' correlations with ' + str(cm.index[0]), fontsize = 20 )\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:09.070904Z","iopub.execute_input":"2022-11-07T12:18:09.071272Z","iopub.status.idle":"2022-11-07T12:18:09.909572Z","shell.execute_reply.started":"2022-11-07T12:18:09.071237Z","shell.execute_reply":"2022-11-07T12:18:09.908197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Pre-B cells , but CD10 - not in included in targets - so we skip it. \n\n    sB-2: small pre-B-cell     Immunophenotype: CD10+ FSC-Alow\n    Genes: DNTT, VPREB1, HHIP TFs: EBF1, ID3, ATF3\n    \n    sB-1: small pre-B-cell     Immunophenotype: CD10+ FSC-Alow\n    Genes: JCHAIN, DNTT, CD79A TFs: EBF1, SATB1, SP140\n","metadata":{}},{"cell_type":"markdown","source":"# Monocyte/dendritic cell progenitors, 'CD45RA' and related CD45, CD45RO , Immunophenotype: CD10–CD45RAhiCD135hi\n\n    Genes: SCT, TGFBI, LGMN TFs: IRF8, IRF7, TCF4\n    Immunophenotype: CD10–CD45RAhiCD135hi\n    \n    CD135 - not in included in CITEseq targets dataset, so skip it \n    \n    \n# Conclusion - for CD45RA, we have IRF8 and TCF4 correlated around 0.3 but all the other correlations are quite lower. ","metadata":{}},{"cell_type":"code","source":"l_prot = ['CD45RA', 'CD45', 'CD45RO'] # - all these from the same gene \"ENSG00000081237\" PTPRC (HGNC Symbol)\nfor p in l_prot:\n    print(p, p in df_cite_train_y.columns )\n    \n'ENSG00000081237' in list_genes_ids \n# print( 'CD45' in list_genes_names, 'ENSG00000081237' in list_genes_ids , 'CD45' in df_cite_train_y.columns )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:09.911223Z","iopub.execute_input":"2022-11-07T12:18:09.911722Z","iopub.status.idle":"2022-11-07T12:18:09.924003Z","shell.execute_reply.started":"2022-11-07T12:18:09.911676Z","shell.execute_reply":"2022-11-07T12:18:09.922507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nprint( 'CD71' in list_genes_names, 'CY1G4' in list_genes_names , 'ENSG00000072274' in list_genes_ids, 'CD71' in df_cite_train_y.columns )\n\ndf_new = pd.DataFrame()\nfor p in l_prot:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\nlist_ids = ['ENSG00000081237', 'SCT', 'TGFBI', 'LGMN', 'IRF8', 'IRF7', 'TCF4' ]\nlist_names = ['CD45', 'SCT', 'TGFBI', 'LGMN', 'IRF8', 'IRF7', 'TCF4' ]\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:09.925844Z","iopub.execute_input":"2022-11-07T12:18:09.926173Z","iopub.status.idle":"2022-11-07T12:18:10.142388Z","shell.execute_reply.started":"2022-11-07T12:18:09.926139Z","shell.execute_reply":"2022-11-07T12:18:10.141020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for IX in [0,1,2]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.grid()\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:10.143821Z","iopub.execute_input":"2022-11-07T12:18:10.144688Z","iopub.status.idle":"2022-11-07T12:18:10.845148Z","shell.execute_reply.started":"2022-11-07T12:18:10.144622Z","shell.execute_reply":"2022-11-07T12:18:10.843728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for method_corr in ['pearson','spearman','kendall']:\n    print(method_corr)\n    display( df_new.corr(method = method_corr ).iloc[:4,:4] )\n    print()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:10.846540Z","iopub.execute_input":"2022-11-07T12:18:10.847037Z","iopub.status.idle":"2022-11-07T12:18:11.549858Z","shell.execute_reply.started":"2022-11-07T12:18:10.846989Z","shell.execute_reply":"2022-11-07T12:18:11.548575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cms = df_new.corr().iloc[:4,:4]\ndisplay(cms)\nfor i in range(len(cms)):\n    cms.iloc[i,i] = np.nan\nplt.figure(figsize=(5,5) )\n#sns.heatmap(cms)\nax = sns.heatmap( cms[ ((cms >= 0.1) | (cms <= -0.2))  ], \n            cmap='rainbow', vmax=cms.max().max(), vmin=-cms.min().min(), linewidths=0.1,\n            annot=True, annot_kws={\"size\": 12}, square=True);\nax.xaxis.set_tick_params(labelsize=11)\nax.yaxis.set_tick_params(labelsize=11)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:11.551392Z","iopub.execute_input":"2022-11-07T12:18:11.551833Z","iopub.status.idle":"2022-11-07T12:18:11.901709Z","shell.execute_reply.started":"2022-11-07T12:18:11.551801Z","shell.execute_reply":"2022-11-07T12:18:11.900375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Negative correlations for CD45RA - indeed there are some !  \n\nGenes: CLC, HDC, PRG2 TFs: LMO4, GATA2, NR4A3\nImmunophenotype: CD10–CD45RA–CD135mid\n\n\nSo these genes should be not or anti correlated with CD45RA\n\nSince we see \"-\" after CD45RA\n\n","metadata":{}},{"cell_type":"code","source":"\ndf_new = pd.DataFrame()\nfor p in ['CD45RA']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\nlist_ids = ['CLC', 'HDC', 'PRG2', 'LMO4', 'GATA2', 'NR4A3' ]\nlist_names = list_ids \n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:11.902859Z","iopub.execute_input":"2022-11-07T12:18:11.903181Z","iopub.status.idle":"2022-11-07T12:18:12.072547Z","shell.execute_reply.started":"2022-11-07T12:18:11.903152Z","shell.execute_reply":"2022-11-07T12:18:12.071298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for IX in [0]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n        plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n        ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:12.074468Z","iopub.execute_input":"2022-11-07T12:18:12.074824Z","iopub.status.idle":"2022-11-07T12:18:12.358975Z","shell.execute_reply.started":"2022-11-07T12:18:12.074786Z","shell.execute_reply":"2022-11-07T12:18:12.357685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CD45RA  and  neutrophil progenitors\n\nN-1: intermediate-stage neutrophil progenitor\nGenes: CTSG, AZU1, ELANE TFs: CEBPD, CEBPA, CEBPE\n        \nN-3: late-stage neutrophil progenitor\nGenes: LYZ, CSTA, LGALS1 TFs: EGR1, CEBPD, FOSB\n        \nN-2: intermediate-stage neutrophil progenitor\nGenes: AZU1, ELANE, PRTN3 TFs: CEBPA, CEBPD, CEBPE\n        ","metadata":{}},{"cell_type":"code","source":"        \ndf_new = pd.DataFrame()\nfor p in ['CD45RA']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\nlist_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n           'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n           'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\nlist_ids = list(set(list_ids ))\nlist_names = list_ids \n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n        ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:12.360804Z","iopub.execute_input":"2022-11-07T12:18:12.361156Z","iopub.status.idle":"2022-11-07T12:18:12.963037Z","shell.execute_reply.started":"2022-11-07T12:18:12.361123Z","shell.execute_reply":"2022-11-07T12:18:12.961581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Markers of \"Immature populations\" - all include CD45RA\n\nIm-2: immature myeloid progenitor, high cell cycle activity\nGenes: TK1, TMEM92, SPINK2 TFs: MYBL2, BATF3, TCF19\nImmunophenotype: CD10–CD45RA–CD135\n    \nIm-1: immature myeloid progenitor\nGenes: SELL, SPINK2, CD52 TFs: HOXB5, GATA3, ENO1\nImmunophenotype: CD10–CD45RA–CD135mid\n\nN-0: early-stage neutrophil progenitor\nGenes: MPO, RNASE2, ATP8B4 TFs: SPI1, RUNX1, MAFK\nImmunophenotype: CD10–CD45RAmidCD135mid\n","metadata":{}},{"cell_type":"code","source":"\n    \n        \ndf_new = pd.DataFrame()\nfor p in ['CD45RA']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\n# list_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n#            'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n#            'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\nlist_ids = ['TK1', 'TMEM92', 'SPINK2', 'MYBL2', 'BATF3', 'TCF19',\n'SELL', 'SPINK2', 'CD52', 'HOXB5', 'GATA3', 'ENO1',\n'MPO', 'RNASE2', 'ATP8B4', 'SPI1', 'RUNX1', 'MAFK',]\nlist_ids = list(set(list_ids ))\nlist_names = list_ids \n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n            ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:12.964678Z","iopub.execute_input":"2022-11-07T12:18:12.965142Z","iopub.status.idle":"2022-11-07T12:18:13.759320Z","shell.execute_reply.started":"2022-11-07T12:18:12.965109Z","shell.execute_reply":"2022-11-07T12:18:13.758373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check correlation of CD44 with HAS1, HAS2, HAS3\n\nIn addition, the interaction of HA with the leukocyte receptor **CD44**  \nis important in tissue-specific homing by leukocytes, and overexpression of HA receptors has been correlated with tumor metastasis. HAS1 is a member of the newly identified vertebrate gene family encoding putative hyaluronan synthases, and its amino acid sequence shows significant homology to the hasA gene product of Streptococcus pyogenes, a glycosaminoglycan synthetase (DG42) from Xenopus laevis, and a recently described murine hyaluronan synthase.[6]\n\nhttps://en.wikipedia.org/wiki/HAS1\n\nProposed by Alexander Romanishin \n\n# Conclusion - HAS1,2 - not in data, HAS3 is low expressed, so probably no reliable conclusions (correlation is almost zero). \n","metadata":{}},{"cell_type":"code","source":"\n    \n        \ndf_new = pd.DataFrame()\nfor p in ['CD95']: # 'CD45RA', 'CD45','CD45RO']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\n# list_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n#            'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n#            'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\n# list_ids = ['TK1', 'TMEM92', 'SPINK2', 'MYBL2', 'BATF3', 'TCF19',\n# 'SELL', 'SPINK2', 'CD52', 'HOXB5', 'GATA3', 'ENO1',\n# 'MPO', 'RNASE2', 'ATP8B4', 'SPI1', 'RUNX1', 'MAFK',]\nlist_ids = ['ENSG00000002330', 'ENSG00000030110', 'ENSG00000026103', 'ENSG00000087088', 'ENSG00000171791', 'ENSG00000142867', 'ENSG00000105327', 'ENSG00000100290', 'ENSG00000165233', 'ENSG00000163098', 'ENSG00000105483', 'ENSG00000137752', 'ENSG00000106144', 'ENSG00000164305', 'ENSG00000196954', 'ENSG00000137757', 'ENSG00000138794', 'ENSG00000165806', 'ENSG00000064012', 'ENSG00000132906', 'ENSG00000003400', 'ENSG00000204403', 'ENSG00000003402', 'ENSG00000169372', 'ENSG00000184047', 'ENSG00000168040', 'ENSG00000125746', 'ENSG00000117560', 'ENSG00000106511', 'ENSG00000145649', 'ENSG00000100453', 'ENSG00000148346', 'ENSG00000186868', 'ENSG00000116688', 'ENSG00000168404', 'ENSG00000249437', 'ENSG00000165030', 'ENSG00000141682', 'ENSG00000123240', 'ENSG00000132646', 'ENSG00000150593', 'ENSG00000156709', 'ENSG00000177595', 'ENSG00000180644', 'ENSG00000111679', 'ENSG00000105327', 'ENSG00000104332', 'ENSG00000105329', 'ENSG00000092969', 'ENSG00000102871', 'ENSG00000121858', 'ENSG00000101966', 'ENSG00000156709', 'ENSG00000142208', 'ENSG00000105221', 'ENSG00000117020', 'ENSG00000120868', 'ENSG00000149311', 'ENSG00000171552', 'ENSG00000187446', 'ENSG00000166869', 'ENSG00000213341', 'ENSG00000100368', 'ENSG00000172115', 'ENSG00000160049', 'ENSG00000169598', 'ENSG00000149218', 'ENSG00000167136', 'ENSG00000157036', 'ENSG00000026103', 'ENSG00000117560', 'ENSG00000104365', 'ENSG00000269335', 'ENSG00000115008', 'ENSG00000125538', 'ENSG00000115594', 'ENSG00000196083', 'ENSG00000164399', 'ENSG00000185291', 'ENSG00000184216', 'ENSG00000134070', 'ENSG00000090376', 'ENSG00000198001', 'ENSG00000006062', 'ENSG00000172936', 'ENSG00000109320', 'ENSG00000100906', 'ENSG00000134259', 'ENSG00000198400', 'ENSG00000121879', 'ENSG00000051382', 'ENSG00000171608', 'ENSG00000105851','ENSG00000145675', 'ENSG00000105647', 'ENSG00000117461', 'ENSG00000141506', 'ENSG00000138814', 'ENSG00000107758', 'ENSG00000120910', 'ENSG00000221823', 'ENSG00000188386', 'ENSG00000072062', 'ENSG00000142875', 'ENSG00000165059', 'ENSG00000108946', 'ENSG00000188191', 'ENSG00000188191', 'ENSG00000005249', 'ENSG00000183943', 'ENSG00000173039', 'ENSG00000137275', 'ENSG00000232810', 'ENSG00000104689', 'ENSG00000120889', 'ENSG00000173535', 'ENSG00000173530', 'ENSG00000067182', 'ENSG00000121858', 'ENSG00000141510', 'ENSG00000127191']\n\n# list_ids = list(set(list_ids ))\n# list_names = list_ids \nlist_names = ['CD95', 'BID',  'BIRC2',  'BIRC3',  'CAPN1' ,  'CAPN2',  'CHP1',  'CHP2',  'CHUK',  'CSF2RB',  'CYCS',  'DFFA',  'DFFB',  'ENDOD1' ,  'ENDOG',  'EXOG',  'FAS',  'FASLG',  'IKBKB',  'IKBKG',  'IL1A',  'IL1B',  'IL1R1',  'IL1RAP',  'IL3',  'IL3RA',  'IRAK1',  'IRAK2',  'IRAK3',  'IRAK4',  'MAP3K14',  'MYD88',  'NFKB1',  'NFKBIA',  'NGF',  'NTRK1',  'PIK3CA',  'PIK3CB' , 'PIK3CD',  'PIK3CG' , 'PIK3R1',  'PIK3R2',  'PIK3R3', 'PIK3R5',  'PPP3CA',  'PPP3CB',  'PPP3CC',  'PPP3R1',  'PPP3R2',  'PRKACA', 'PRKACB',  'PRKACG',  'PRKAR1A',  'PRKAR1B',  'PRKAR2A',  'PRKAR2B',  'PRKX' , 'RELA',  'RIPK1',  'TNF' , 'TNFRSF10A',  'TNFRSF10B',  'TNFRSF10C',  'TNFRSF10D',  'TNFRSF1A',  'TNFSF10',  'TP53',  'TRAF2',  'BAD',  'BAK1',  'BAX',  'BCL2',  'BCL10',  'BclXL',  'BIK',  'BINCARD',  'BIRC8',  'CARD8',  'CASP1',  'CASP2',  'CASP3',  'CASP4',  'CASP5',  'CASP6',  'CASP7',  'CASP8',  'CASP9',  'CASP10', 'CASP12',  'CFLAR',  'CRADD',  'DIABLO',  'FADD' , 'EMAP2',  'FASL',  'GAX/MEOX2',  'GZMA',  'GZMB',  'LCN2',  'MAPT',  'MFN2',  'MLKL',  'NAIP1',  'NFIL3',  'NOXA/PMAIP1',  'OPTN' ,'PCNA', 'PDCD4' ,'PDCD8/AIFM1',  'PIDD' , 'PRF1',  'PTPN6',  'PUMA/BBC3',  'SARP2',  'TGFB1',  'TGFB2',  'TRADD',  'TRAIL/TNFSF10',  'XIAP',  'AIFM1' , 'AKT1',  'AKT2',  'AKT3',  'APAF1',  'ATM',  'BCL2L1']\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0]:\n    for method_corr in ['spearman']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        cm = cm.sort_values('CD95_prot', key=abs)[-12:]\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n            ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:55:33.247237Z","iopub.execute_input":"2022-11-07T12:55:33.247674Z","iopub.status.idle":"2022-11-07T12:55:42.266275Z","shell.execute_reply.started":"2022-11-07T12:55:33.247623Z","shell.execute_reply":"2022-11-07T12:55:42.264929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CD44 vs ERG gene  - correlation 0.2 - not bad ! \n\nProposed by Alexander Romanishin  - based on previous finding of the correlation by Anton Kostin - based on semantic clustering of genes descriptions. \nSo it is not completely bio motivated. \n\nThis protein is required for platelet adhesion to the subendothelium, inducing vascular cell remodeling. It also regulates hematopoesis, and the differentiation and maturation of megakaryocytic cells.\n","metadata":{}},{"cell_type":"code","source":"\n    \n        \ndf_new = pd.DataFrame()\nfor p in ['CD95']: # 'CD45RA', 'CD45','CD45RO']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'CD95_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\n# list_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n#            'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n#            'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\n# list_ids = ['TK1', 'TMEM92', 'SPINK2', 'MYBL2', 'BATF3', 'TCF19',\n# 'SELL', 'SPINK2', 'CD52', 'HOXB5', 'GATA3', 'ENO1',\n# 'MPO', 'RNASE2', 'ATP8B4', 'SPI1', 'RUNX1', 'MAFK',]\nlist_ids = ['ENSG00000170961','HAS3'] #  , 'ENSG00000170961','HAS3']\n\n# list_ids = list(set(list_ids ))\n# list_names = list_ids \nlist_names = ['HAS1', 'HAS2', 'HAS3' ]#,  'HAS1', 'HAS2', 'HAS3']\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n            ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:25.363275Z","iopub.execute_input":"2022-11-07T12:18:25.363956Z","iopub.status.idle":"2022-11-07T12:18:25.683318Z","shell.execute_reply.started":"2022-11-07T12:18:25.363917Z","shell.execute_reply":"2022-11-07T12:18:25.682129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 13 CD69 vs PRKCA - conclusion - expression is to small to be confident (correlation near zero) \n\nProtein kinase C alpha (PKCα) is an enzyme \n\nhttps://en.wikipedia.org/wiki/PKC_alpha\n\nProposed by Alexander Romanishin\n\n","metadata":{}},{"cell_type":"code","source":"\n    \n        \ndf_new = pd.DataFrame()\nfor p in ['CD69']: # 'CD45RA', 'CD45','CD45RO']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\n# list_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n#            'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n#            'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\n# list_ids = ['TK1', 'TMEM92', 'SPINK2', 'MYBL2', 'BATF3', 'TCF19',\n# 'SELL', 'SPINK2', 'CD52', 'HOXB5', 'GATA3', 'ENO1',\n# 'MPO', 'RNASE2', 'ATP8B4', 'SPI1', 'RUNX1', 'MAFK',]\nlist_ids = ['ENSG00000110848','ENSG00000154229'] #  , 'ENSG00000170961','HAS3']\n\n# list_ids = list(set(list_ids ))\n# list_names = list_ids \nlist_names = ['CD69', 'PRKCA' ]#,  'HAS1', 'HAS2', 'HAS3']\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n            ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:25.685137Z","iopub.execute_input":"2022-11-07T12:18:25.686007Z","iopub.status.idle":"2022-11-07T12:18:26.219048Z","shell.execute_reply.started":"2022-11-07T12:18:25.685852Z","shell.execute_reply":"2022-11-07T12:18:26.217797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ","metadata":{}},{"cell_type":"markdown","source":"# 'CD56', 'CD94' and RORC - conlusion RORC is absent, CD56, CD94 proteins - have some correlation 0.16 \n\nProposed by Alexander Romanishin\nBased on: \nhttps://www.frontiersin.org/articles/10.3389/fimmu.2019.02078/full\n\n\n","metadata":{}},{"cell_type":"code","source":"\n    \n        \ndf_new = pd.DataFrame()\nfor p in ['CD56', 'CD94']: ##'CD69']: # 'CD45RA', 'CD45','CD45RO']:\n    if p in df_cite_train_y.columns:\n        df_new[p+'_prot'] = df_cite_train_y[p] # CD71 or TRFC \n\n\n# list_ids = ['CTSG', 'AZU1', 'ELANE', 'CEBPD', 'CEBPA', 'CEBPE',\n#            'LYZ', 'CSTA', 'LGALS1', 'EGR1', 'CEBPD', 'FOSB',\n#            'AZU1', 'ELANE', 'PRTN3', 'CEBPA', 'CEBPD', 'CEBPE']\n# list_ids = ['TK1', 'TMEM92', 'SPINK2', 'MYBL2', 'BATF3', 'TCF19',\n# 'SELL', 'SPINK2', 'CD52', 'HOXB5', 'GATA3', 'ENO1',\n# 'MPO', 'RNASE2', 'ATP8B4', 'SPI1', 'RUNX1', 'MAFK',]\nlist_ids = ['ENSG00000149294','ENSG00000134539', 'ENSG00000143365'] #  , 'ENSG00000170961','HAS3']\n\n# list_ids = list(set(list_ids ))\n# list_names = list_ids \nlist_names = ['CD56', 'CD94','RORC'  ]#,  'HAS1', 'HAS2', 'HAS3']\n              \nfor i,id1 in enumerate(list_ids): \n    if id1 in list_genes_ids:\n        gene_name1 = list_names[i]\n        IX = list_genes_ids.index(id1)\n    elif  id1 in list_genes_names: \n        gene_name1 = list_names[i]\n        IX = list_genes_names.index(id1)\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \n    print(IX, list_genes_ids[IX], list_genes_names[IX], df_cite_train_x.columns[IX] )\n    v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n    df_new[gene_name1 ] = v\n\n\ndisplay(df_new)\ndisplay(df_new.describe() )\ndisplay(df_new.corr() )\n\nfor IX in [0,1]:\n    for method_corr in ['pearson']: #, 'spearman','kendall']:\n        cm = df_new.corr(method = method_corr)\n        # print(cm)\n        plt.figure(figsize  = (20,5))\n        v = cm.iloc[:,IX]\n        v2 = list(v.iloc[:IX]) + list(v.iloc[(IX+1):])\n        plt.bar(x = cm.index, height = v, color = 'r')\n        plt.scatter(x = cm.index, y = 0.5*np.ones(len(cm)),  color = 'b')\n        plt.text(cm.index[-1],0.5, '0.5', fontsize = 20)\n        plt.scatter(x = cm.index, y = 0.3*np.ones(len(cm)),  color = 'g')\n        plt.text(cm.index[-1],0.3, '0.3', fontsize = 20)\n#         plt.scatter(x = cm.index, y = -0.2*np.ones(len(cm)),  color = 'g')\n#         plt.text(cm.index[-1],-0.2, '-0.2', fontsize = 20)\n\n        plt.grid()# which='both')\n        plt.title(method_corr+  ' correlations with ' + str(cm.index[IX]) + ' Mean = ' +'%.2f'%np.mean(v2), fontsize = 20 )\n        plt.show()\n        \n            ","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:26.220895Z","iopub.execute_input":"2022-11-07T12:18:26.221993Z","iopub.status.idle":"2022-11-07T12:18:26.795794Z","shell.execute_reply.started":"2022-11-07T12:18:26.221956Z","shell.execute_reply":"2022-11-07T12:18:26.794676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-07T12:18:26.797754Z","iopub.execute_input":"2022-11-07T12:18:26.798230Z","iopub.status.idle":"2022-11-07T12:18:26.805287Z","shell.execute_reply.started":"2022-11-07T12:18:26.798174Z","shell.execute_reply":"2022-11-07T12:18:26.803995Z"},"trusted":true},"execution_count":null,"outputs":[]}]}