{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nWe will try to find(analyse) good predictors for CD** protein levels - based in biological knowledge.\nI.e. we search for genes which RNA levels (i.e. features of CITEseq task)\nare correlated/good predictors for the CD** proteins (targets of the CITEseq task).\n\nWe widely use Nature paper recommended by orgs - see table on Figure 3a. \nhttps://www.nature.com/articles/ncb3493    \nHuman haematopoietic stem cell lineage commitment is a continuous process  \nPublished: 20 March 2017\n\n\n## Versions\n\n#### 15'CD14', 'LY96' and 'TLR4' - have no correlation\n\nProposed by Alexander Romanishin\n\n#### 14 'CD56', 'CD94' and RORC - conlusion RORC is absent, CD56, CD94 proteins - have some correlation 0.16 \n\nProposed by Alexander Romanishin\nBased on: \nhttps://www.frontiersin.org/articles/10.3389/fimmu.2019.02078/full\n\n#### 13 CD69 vs PRKCA - conclusion - expression is to small to be confident (correlation near zero) \n\nProtein kinase C alpha (PKCα) is an enzyme \n\nhttps://en.wikipedia.org/wiki/PKC_alpha\n\nProposed by Alexander Romanishin\n\n#### 12 CD44 with ERG gene - correlation - 0.2 - not bad \n\n    Not completely bio motivated. \n    \n\n    \n\n#### 11 Check correlation of CD44 with HAS1, HAS2, HAS3 -  Conclusion - HAS1,2 - not in data, HAS3 is low expressed, so probably no reliable conclusions (correlation is almost zero). \n\n    In addition, the interaction of HA with the leukocyte receptor **CD44**  \n    is important in tissue-specific homing by leukocytes, and overexpression of HA receptors has been correlated with tumor metastasis. HAS1 is a member of the newly identified vertebrate gene family encoding putative hyaluronan synthases, and its amino acid sequence shows significant homology to the hasA gene product of Streptococcus pyogenes, a glycosaminoglycan synthetase (DG42) from Xenopus laevis, and a recently described murine hyaluronan synthase.[6]\n\nhttps://en.wikipedia.org/wiki/HAS1\n\nProposed by Alexander Romanishin \n\n\n\n#### 10 cosmetic chages\n\n#### 9 for CD45RA from immature cells part of table - MPO SPINK2 ENO1 ATP884 SELL are correlated around 0.3  \n\n#### 8 for CD45RA from neutrophils part of table - CEBPA CEPBD CSTG are correlated around 0.3\n\n\n#### 7 CD45RA is ANTI-correllated (0.2) with HDC and GATA2 - \n    \n    which is in correpondence that CD45RA(-) is marker for cell type Eosinophil/basophil/mast cell progenitors\n    where these genes are expressed\n    (according to paper)\n    \n    \n#### 6 - mild changes\n    \n    'CD45RA, CD45, CD45RO , which are coded by the SAME gene CD45 = ENSG00000081237 = PTPRC (HGNC Symbol)\n    It seems only main CD45 is correlated wtih it its gene and the other two proteins\n    \n\n#### 5 - continue to use the Nature paper - continue with table on Figure 3a.\n\n    Conclusion - for CD45RA, we have IRF8 and TCF4 correlated around 0.3 but all the other correlations are quite lower. \n\n    Note we have 3 target proteins:  'CD45RA, CD45, CD45RO , which are coded by the SAME gene CD45 = ENSG00000081237 = PTPRC (HGNC Symbol)\n\n    Monocyte/dendritic cell progenitors, 'CD45RA' and related CD45, CD45RO , Immunophenotype: CD10–CD45RAhiCD135hi\n\n    Genes: SCT, TGFBI, LGMN TFs: IRF8, IRF7, TCF4\n    Immunophenotype: CD10–CD45RAhiCD135hi\n    \n    CD135 - not in included in CITEseq targets dataset, so skip it \n    \n    \n\n#### 4 use Nature paper recomenned by orgs - for some targets like CD71 - see quite (0.5) correlated genes \n\nhttps://www.nature.com/articles/ncb3493   \nHuman haematopoietic stem cell lineage commitment is a continuous process  \nPublished: 20 March 2017\n\n\n\n#### 2,3 - cosmetic changes\n\n#### 1 - gene \"SET\" not that much bad for CD47\n\n    Proposed by Alexander Romanishin, [10/8/2022 4:39 PM]\n\n    Paper:\n    Berkovits BD, Mayr C. Alternative 3' UTRs act as scaffolds to regulate membrane protein\n    localization. Nature. 2015 Jun 18;522(7556):363-7.\n\n    Idea:\n    HuR и SET genes are responsible for transfer and binding of the CD47 proteins - so one may look on them.\n\n    Conclusion:\n    At least SET gene is not worse correlated with CD47 protein than CD47 RNA itself with protein\n    Correlation is not that much big - about 0.1 but nevertheless.\n    For other gene - HuR - correlation is low. But HuR and SET are correlated among themselves - level about 0.25\n    \n    ","metadata":{}},{"cell_type":"markdown","source":"# Install/Import ","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-14T16:22:49.988870Z","iopub.execute_input":"2022-11-14T16:22:49.989359Z","iopub.status.idle":"2022-11-14T16:22:50.014895Z","shell.execute_reply.started":"2022-11-14T16:22:49.989322Z","shell.execute_reply":"2022-11-14T16:22:50.013454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install scanpy\nimport scanpy as sc\nimport anndata\n\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:22:54.137341Z","iopub.execute_input":"2022-11-14T16:22:54.137873Z","iopub.status.idle":"2022-11-14T16:23:32.285952Z","shell.execute_reply.started":"2022-11-14T16:22:54.137834Z","shell.execute_reply":"2022-11-14T16:23:32.284589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"%%time\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\nprint('8.2G RAM consumed')\ndisplay( df_cite_train_x.head() )\n\ndf_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_TARGETS)\ndf_cite_train_y.head()\nprint('8.3G RAM consumed in addtion to 8.2 for features part ')\n\nif 0:\n    df_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\n    display( df_cite_test_x.head() )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:23:32.288317Z","iopub.execute_input":"2022-11-14T16:23:32.288684Z","iopub.status.idle":"2022-11-14T16:24:29.623528Z","shell.execute_reply.started":"2022-11-14T16:23:32.288651Z","shell.execute_reply":"2022-11-14T16:24:29.622147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"%%time\nlist_genes_names = [t.split('_')[1] for t in df_cite_train_x.columns ]\nprint(len(list_genes_names), list_genes_names[:10])\nlist_genes_ids = [t.split('_')[0] for t in df_cite_train_x.columns ]\nprint(len(list_genes_ids), list_genes_ids[:10])\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:27:05.174181Z","iopub.execute_input":"2022-11-14T16:27:05.175239Z","iopub.status.idle":"2022-11-14T16:27:05.203997Z","shell.execute_reply.started":"2022-11-14T16:27:05.175186Z","shell.execute_reply":"2022-11-14T16:27:05.202865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'ENSG00000066044' in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:11:34.843143Z","iopub.execute_input":"2022-11-14T16:11:34.843812Z","iopub.status.idle":"2022-11-14T16:11:34.854016Z","shell.execute_reply.started":"2022-11-14T16:11:34.843777Z","shell.execute_reply":"2022-11-14T16:11:34.852895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculations of the correlations (rna)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T15:02:29.460937Z","iopub.execute_input":"2022-10-08T15:02:29.461385Z","iopub.status.idle":"2022-10-08T15:02:29.481832Z","shell.execute_reply.started":"2022-10-08T15:02:29.461348Z","shell.execute_reply":"2022-10-08T15:02:29.480641Z"}}},{"cell_type":"markdown","source":"Соответствия РНК и белков:","metadata":{}},{"cell_type":"code","source":"rna=pd.read_excel(\"../input/proteinrna/TotalSeq_B_Universal_Cocktail_v1_140_Antibodies_399904_Barcodes.xlsx\")","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:27:11.717669Z","iopub.execute_input":"2022-11-14T16:27:11.718396Z","iopub.status.idle":"2022-11-14T16:27:11.766317Z","shell.execute_reply.started":"2022-11-14T16:27:11.718352Z","shell.execute_reply":"2022-11-14T16:27:11.765292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:11:35.362455Z","iopub.execute_input":"2022-11-14T16:11:35.363183Z","iopub.status.idle":"2022-11-14T16:11:35.380752Z","shell.execute_reply.started":"2022-11-14T16:11:35.363139Z","shell.execute_reply":"2022-11-14T16:11:35.379820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = pd.DataFrame()\nens=rna[\"Ensemble ID\"].dropna().to_numpy()\nnf=[]\n\n#list_names = rna[\"Description\"].to_numpy()\n              \nfor i,id1 in enumerate(ens): \n    if id1 in list_genes_ids:\n        gene_name1 = ens[i]+\"_rna\"\n        IX = list_genes_ids.index(id1)\n        v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n        df_new[gene_name1 ] = v\n    else:\n        print(id1, 'cannot be found')\n        nf.append(id1)\n        continue \n        \nmatr=df_new.corr()","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:27:17.696204Z","iopub.execute_input":"2022-11-14T16:27:17.697417Z","iopub.status.idle":"2022-11-14T16:27:20.913084Z","shell.execute_reply.started":"2022-11-14T16:27:17.697371Z","shell.execute_reply":"2022-11-14T16:27:20.912053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ens_rna=[]\nfor elem in df_new.columns:\n    ens_rna.append(elem[:-4])","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:51:03.751791Z","iopub.execute_input":"2022-11-14T16:51:03.752275Z","iopub.status.idle":"2022-11-14T16:51:03.760094Z","shell.execute_reply.started":"2022-11-14T16:51:03.752234Z","shell.execute_reply":"2022-11-14T16:51:03.758511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import Ridge\ny_predicted=[]\nfor elem in ens_rna:\n    IX = list_genes_ids.index(elem)\n    X=df_cite_train_x.drop(columns=[df_cite_train_x.columns[IX]])\n    y =  df_cite_train_x.iloc[:, IX]\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)\n    r=Ridge()\n    r.fit(X_train, y_train)\n    y_predicted.append(r.predict(X_test))\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:51:16.548366Z","iopub.execute_input":"2022-11-14T16:51:16.548840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import Ridge\ny_predicted=[]\nfor elem in ens:\n    if elem in list_genes_ids:\n        IX = list_genes_ids.index(elem)\n        X=df_cite_train_x.drop(columns=[df_cite_train_x.columns[IX]])\n        y =  df_cite_train_x.iloc[:, IX]\n        X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)\n        r=Ridge()\n        r.fit(X_train, y_train)\n        y_predicted.append(r.predict(X_test))\n    else:\n        continue","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:32:33.374134Z","iopub.execute_input":"2022-11-14T16:32:33.374758Z","iopub.status.idle":"2022-11-14T16:49:07.027643Z","shell.execute_reply.started":"2022-11-14T16:32:33.374697Z","shell.execute_reply":"2022-11-14T16:49:07.025477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame(y_predicted, index=df_new.columns).T.to_csv('./ridge.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:32.988268Z","iopub.status.idle":"2022-11-14T16:19:32.990448Z","shell.execute_reply.started":"2022-11-14T16:19:32.989651Z","shell.execute_reply":"2022-11-14T16:19:32.989749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"v = matr.values[np.triu_indices(len(matr),k=1)]\nt =  np.min([30, len(v)])\nthreshold4top_corr = np.sort(v[~np.isnan(v)])[-t ]\nthreshold4bottom_corr = np.sort(v[~np.isnan(v)])[ t ]","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:32.993976Z","iopub.status.idle":"2022-11-14T16:19:32.994798Z","shell.execute_reply.started":"2022-11-14T16:19:32.994447Z","shell.execute_reply":"2022-11-14T16:19:32.994483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"z =    np.where( np.triu(matr.values,1) >= threshold4top_corr )\ndf_corrs = pd.DataFrame()\nIX = 0\nfor (i,j) in zip( z[0], z[1]):\n    if i>=j : continue \n    #print(i,j, list_genes_ids[i], list_genes_ids[j], cm.iloc[i,j])\n    IX += 1\n    df_corrs.loc[IX, 'I1'] = i\n    df_corrs.loc[IX, 'I2'] = j\n    df_corrs.loc[IX, 'Id1'] = matr.columns[i]\n    df_corrs.loc[IX, 'Id2'] = matr.columns[j]\n    df_corrs.loc[IX, 'Correlation'] = matr.iloc[i,j]\n    \ndf_corrs['I1'] = df_corrs['I1'].astype(int) \ndf_corrs['I2'] = df_corrs['I2'].astype(int) \n\nprint(df_corrs.shape)\n\ndisplay ( df_corrs.sort_values('Correlation' , ascending = False) )","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:32.996762Z","iopub.status.idle":"2022-11-14T16:19:32.997428Z","shell.execute_reply.started":"2022-11-14T16:19:32.997135Z","shell.execute_reply":"2022-11-14T16:19:32.997164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hlp=[]\nfor elem in ens:\n    if elem not in hlp:\n        hlp.append(elem)\n    else:\n        print(elem)\nlen(np.unique(ens))","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:32.999140Z","iopub.status.idle":"2022-11-14T16:19:32.999738Z","shell.execute_reply.started":"2022-11-14T16:19:32.999444Z","shell.execute_reply":"2022-11-14T16:19:32.999473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna1=[]\nfor elem in matr.columns:\n    rna1.append(elem[:-4])\nrna1","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.003272Z","iopub.status.idle":"2022-11-14T16:19:33.003870Z","shell.execute_reply.started":"2022-11-14T16:19:33.003563Z","shell.execute_reply":"2022-11-14T16:19:33.003592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna_nan=[]\nfor i in range(len(matr.columns)):\n    for j in range(i,len(matr.columns)):\n        if np.isnan(matr[matr.columns[i]][matr.columns[j]]):\n            rna_nan.append(matr.columns[i]+\" \"+matr.columns[j])\nrna_nan","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.006394Z","iopub.status.idle":"2022-11-14T16:19:33.007197Z","shell.execute_reply.started":"2022-11-14T16:19:33.006840Z","shell.execute_reply":"2022-11-14T16:19:33.006874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.clustermap(np.abs(matr.fillna(0)), figsize=(25, 25), xticklabels=matr.columns, yticklabels=matr.columns, cmap='vlag')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.009439Z","iopub.status.idle":"2022-11-14T16:19:33.010171Z","shell.execute_reply.started":"2022-11-14T16:19:33.009797Z","shell.execute_reply":"2022-11-14T16:19:33.009840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1 кластер:","metadata":{}},{"cell_type":"code","source":"clst1=[]\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000117091\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000166825\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000026508\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000188404\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000135218\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000143226\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000179639\"][\"Description\"])\nclst1.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000005961\"][\"Description\"])\nclst1 #41,13, 32 +","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.012570Z","iopub.status.idle":"2022-11-14T16:19:33.013982Z","shell.execute_reply.started":"2022-11-14T16:19:33.013579Z","shell.execute_reply":"2022-11-14T16:19:33.013615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"2 кластер:","metadata":{}},{"cell_type":"code","source":"clst2=[]\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000121807\"][\"Description\"])\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000182578\"][\"Description\"])\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000099250\"][\"Description\"])\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000168685\"][\"Description\"])\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000170458\"][\"Description\"])\nclst2.append(rna[rna[\"Ensemble ID\"]==\"ENSG00000177575\"][\"Description\"])\nclst2 #192, 304, 127 все в отдельных","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.017455Z","iopub.status.idle":"2022-11-14T16:19:33.018070Z","shell.execute_reply.started":"2022-11-14T16:19:33.017762Z","shell.execute_reply":"2022-11-14T16:19:33.017789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Итого, в РНК хорошо выделяются 2 кластера, тогда как для белков выделяются как минимум 5 кластеров.\n\nВ 1 кластере по РНК корреллируют гены такие как CD48,13, 44, 62L, 36, 32, 41; при этом корреляция CD 41,13, 32 наблюдалась также в белках.\n\nВо 2 кластере по РНК коррелируются гены CD192, 115, 304, 127, 14, 163; из них CD192, 304, 127 присутствовали в корреляции белков, но каждый в своём кластере\n\nТакже не все РНК из таблицы TotalSeq_B_Universal_Cocktail_v1_140_Antibodies_399904_Barcodes.xlsx удалось найти в датасете соревнования, а также для большого количества РНК отсутствовала корреляция (см. таблицу выше).","metadata":{}},{"cell_type":"code","source":"citeseq_target=pd.read_csv('../input/target-genes-denoised/citeseq_target_genes.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.020674Z","iopub.status.idle":"2022-11-14T16:19:33.021869Z","shell.execute_reply.started":"2022-11-14T16:19:33.021657Z","shell.execute_reply":"2022-11-14T16:19:33.021679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ensg=[]\nfor elem in citeseq_target.columns.drop(\"cell_id\"):\n    ensg.append(elem[:15])\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.022994Z","iopub.status.idle":"2022-11-14T16:19:33.023831Z","shell.execute_reply.started":"2022-11-14T16:19:33.023612Z","shell.execute_reply":"2022-11-14T16:19:33.023634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = pd.DataFrame()\nens=citeseq_target.columns.drop(\"cell_id\").to_numpy()\n\n#list_names = rna[\"Description\"].to_numpy()\n              \nfor i,id1 in enumerate(ensg): \n    if id1 in list_genes_ids:\n        gene_name1 = ensg[i]+\"_rna\"\n        IX = list_genes_ids.index(id1)\n        v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n        df_new[gene_name1 ] = v\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \nmatr=df_new.corr()\nmatr","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.025246Z","iopub.status.idle":"2022-11-14T16:19:33.026062Z","shell.execute_reply.started":"2022-11-14T16:19:33.025788Z","shell.execute_reply":"2022-11-14T16:19:33.025811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"v = matr.values[np.triu_indices(len(matr),k=1)]\nt =  np.min([30, len(v)])\nthreshold4top_corr = np.sort(v[~np.isnan(v)])[-t ]\nthreshold4bottom_corr = np.sort(v[~np.isnan(v)])[ t ]","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.028037Z","iopub.status.idle":"2022-11-14T16:19:33.028950Z","shell.execute_reply.started":"2022-11-14T16:19:33.028704Z","shell.execute_reply":"2022-11-14T16:19:33.028734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"z =    np.where( np.triu(matr.values,1) >= threshold4top_corr )\ndf_corrs = pd.DataFrame()\nIX = 0\nfor (i,j) in zip( z[0], z[1]):\n    if i>=j : continue \n    #print(i,j, list_genes_ids[i], list_genes_ids[j], cm.iloc[i,j])\n    IX += 1\n    df_corrs.loc[IX, 'I1'] = i\n    df_corrs.loc[IX, 'I2'] = j\n    df_corrs.loc[IX, 'Id1'] = matr.columns[i]\n    df_corrs.loc[IX, 'Id2'] = matr.columns[j]\n    df_corrs.loc[IX, 'Correlation'] = matr.iloc[i,j]\n    \ndf_corrs['I1'] = df_corrs['I1'].astype(int) \ndf_corrs['I2'] = df_corrs['I2'].astype(int) \n\nprint(df_corrs.shape)\n\ndisplay ( df_corrs.sort_values('Correlation' , ascending = False) )","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.030124Z","iopub.status.idle":"2022-11-14T16:19:33.030801Z","shell.execute_reply.started":"2022-11-14T16:19:33.030593Z","shell.execute_reply":"2022-11-14T16:19:33.030614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matr.values","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.031971Z","iopub.status.idle":"2022-11-14T16:19:33.032755Z","shell.execute_reply.started":"2022-11-14T16:19:33.032543Z","shell.execute_reply":"2022-11-14T16:19:33.032565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"z","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.034512Z","iopub.status.idle":"2022-11-14T16:19:33.035409Z","shell.execute_reply.started":"2022-11-14T16:19:33.035082Z","shell.execute_reply":"2022-11-14T16:19:33.035114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matr","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.037051Z","iopub.status.idle":"2022-11-14T16:19:33.037866Z","shell.execute_reply.started":"2022-11-14T16:19:33.037536Z","shell.execute_reply":"2022-11-14T16:19:33.037564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for elem in ensg:\n    if elem not in rna1:\n        print(elem)","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.039501Z","iopub.status.idle":"2022-11-14T16:19:33.040358Z","shell.execute_reply.started":"2022-11-14T16:19:33.040036Z","shell.execute_reply":"2022-11-14T16:19:33.040064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna_nan=[]\nfor i in range(len(matr.columns)):\n    for j in range(i,len(matr.columns)):\n        if np.isnan(matr[matr.columns[i]][matr.columns[j]]):\n            rna_nan.append(matr.columns[i]+\" \"+matr.columns[j])\nlen(rna_nan)","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.042054Z","iopub.status.idle":"2022-11-14T16:19:33.042728Z","shell.execute_reply.started":"2022-11-14T16:19:33.042496Z","shell.execute_reply":"2022-11-14T16:19:33.042518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.clustermap(np.abs(matr.fillna(0)), figsize=(25, 25), xticklabels=matr.columns, yticklabels=matr.columns, cmap='vlag')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.043864Z","iopub.status.idle":"2022-11-14T16:19:33.044630Z","shell.execute_reply.started":"2022-11-14T16:19:33.044397Z","shell.execute_reply":"2022-11-14T16:19:33.044419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"citeseq_target_dca=pd.read_csv('../input/target-genes-denoised/citeseq_target_genes_denoised_dca.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.046325Z","iopub.status.idle":"2022-11-14T16:19:33.047288Z","shell.execute_reply.started":"2022-11-14T16:19:33.047064Z","shell.execute_reply":"2022-11-14T16:19:33.047092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = pd.DataFrame()\n\nensg=[]\nfor elem in citeseq_target_dca.columns.drop(\"cell_id\"):\n    ensg.append(elem[:15])\n\n#list_names = rna[\"Description\"].to_numpy()\n              \nfor i,id1 in enumerate(ensg): \n    if id1 in list_genes_ids:\n        gene_name1 = ensg[i]+\"_rna\"\n        IX = list_genes_ids.index(id1)\n        v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n        df_new[gene_name1 ] = v\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \nmatr=df_new.corr()\nmatr","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.048954Z","iopub.status.idle":"2022-11-14T16:19:33.049939Z","shell.execute_reply.started":"2022-11-14T16:19:33.049572Z","shell.execute_reply":"2022-11-14T16:19:33.049606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.clustermap(np.abs(matr.fillna(0)), figsize=(25, 25), xticklabels=matr.columns, yticklabels=matr.columns, cmap='vlag')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.051739Z","iopub.status.idle":"2022-11-14T16:19:33.052884Z","shell.execute_reply.started":"2022-11-14T16:19:33.052473Z","shell.execute_reply":"2022-11-14T16:19:33.052508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna_nan=[]\nfor i in range(len(matr.columns)):\n    for j in range(i,len(matr.columns)):\n        if np.isnan(matr[matr.columns[i]][matr.columns[j]]):\n            rna_nan.append(matr.columns[i]+\" \"+matr.columns[j])\nlen(rna_nan)","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.055006Z","iopub.status.idle":"2022-11-14T16:19:33.055885Z","shell.execute_reply.started":"2022-11-14T16:19:33.055553Z","shell.execute_reply":"2022-11-14T16:19:33.055590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"citeseq_target_magic=pd.read_csv('../input/target-genes-denoised/citeseq_target_genes_denoised_magic.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.057384Z","iopub.status.idle":"2022-11-14T16:19:33.058212Z","shell.execute_reply.started":"2022-11-14T16:19:33.057951Z","shell.execute_reply":"2022-11-14T16:19:33.057981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_new = pd.DataFrame()\n\nensg=[]\nfor elem in citeseq_target_magic.columns.drop(\"cell_id\"):\n    ensg.append(elem[:15])\n\n#list_names = rna[\"Description\"].to_numpy()\n              \nfor i,id1 in enumerate(ensg): \n    if id1 in list_genes_ids:\n        gene_name1 = ensg[i]+\"_rna\"\n        IX = list_genes_ids.index(id1)\n        v =  df_cite_train_x.iloc[:, IX]\n    #df_new[gene_name1 + '_RNA'] = v\n        df_new[gene_name1 ] = v\n    else:\n        print(id1, 'cannot be found')\n        continue \n        \nmatr=df_new.corr()\nmatr","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.059511Z","iopub.status.idle":"2022-11-14T16:19:33.060276Z","shell.execute_reply.started":"2022-11-14T16:19:33.059978Z","shell.execute_reply":"2022-11-14T16:19:33.060000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.clustermap(np.abs(matr.fillna(0)), figsize=(25, 25), xticklabels=matr.columns, yticklabels=matr.columns, cmap='vlag')","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.061417Z","iopub.status.idle":"2022-11-14T16:19:33.062842Z","shell.execute_reply.started":"2022-11-14T16:19:33.062465Z","shell.execute_reply":"2022-11-14T16:19:33.062505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rna_nan=[]\nfor i in range(len(matr.columns)):\n    for j in range(i,len(matr.columns)):\n        if np.isnan(matr[matr.columns[i]][matr.columns[j]]):\n            rna_nan.append(matr.columns[i]+\" \"+matr.columns[j])\nlen(rna_nan)","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.064330Z","iopub.status.idle":"2022-11-14T16:19:33.064874Z","shell.execute_reply.started":"2022-11-14T16:19:33.064592Z","shell.execute_reply":"2022-11-14T16:19:33.064624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'ENSG00000227993' in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.067229Z","iopub.status.idle":"2022-11-14T16:19:33.068364Z","shell.execute_reply.started":"2022-11-14T16:19:33.068142Z","shell.execute_reply":"2022-11-14T16:19:33.068167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"ENSG00000230726\" in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.069907Z","iopub.status.idle":"2022-11-14T16:19:33.071117Z","shell.execute_reply.started":"2022-11-14T16:19:33.070839Z","shell.execute_reply":"2022-11-14T16:19:33.070861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"ENSG00000228987\" in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.072380Z","iopub.status.idle":"2022-11-14T16:19:33.073846Z","shell.execute_reply.started":"2022-11-14T16:19:33.073456Z","shell.execute_reply":"2022-11-14T16:19:33.073485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"ENSG00000204287\" in list_genes_ids","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.075521Z","iopub.status.idle":"2022-11-14T16:19:33.076199Z","shell.execute_reply.started":"2022-11-14T16:19:33.075983Z","shell.execute_reply":"2022-11-14T16:19:33.076011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\n","metadata":{"execution":{"iopub.status.busy":"2022-11-14T16:19:33.078007Z","iopub.status.idle":"2022-11-14T16:19:33.078901Z","shell.execute_reply.started":"2022-11-14T16:19:33.078684Z","shell.execute_reply":"2022-11-14T16:19:33.078706Z"},"trusted":true},"execution_count":null,"outputs":[]}]}