{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p style=\"background-color:#2D735F;font-family:Verdana;color:white;font-size:210%;text-align:center;border-radius: 15px;\">Differentiation of possible protein and RNA regulation pathways. Part 2</p> ","metadata":{}},{"cell_type":"markdown","source":"## Chapter: CD45 isoforms","metadata":{}},{"cell_type":"markdown","source":"<h3>What we do here:</h3>\n<div style=\"line-height:24px; font-size:16px\">\n    <ul style=\"list-style:circle\">\n<li>GO enrichment analysis for top300 lists of RNA that correlate positively and negativey with CD45 isoforms and/or their RNA (lists were prepared in the <a href= \"https://www.kaggle.com/code/antoninadolgorukova/citeseq23-difference-in-prot-rna-correlations?scriptVersionId=121495403\">Part 1 notebook</a>) \n    </ul>\n     <a href= \"https://yulab-smu.top/biomedical-knowledge-mining-book/index.html\">Reference tutorial</a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<hr>\n\n<h3>Data</h3> \n\n<div style=\"line-height:24px; font-size:14px\">\n    <ol>\n<li> RDS files with sparse matrices of <a href=\"http://www.kaggle.com/datasets/stautxie/sparse-measurement-data-open-problems-multimodal\">normalised counts data</a> for <a href=\"http://www.kaggle.com/competitions/open-problems-multimodal\">Open Problems - Multimodal Single-Cell Integration</a> (CITEseq2022). The dataset for this competition comprises single-cell multiomics data (n = 70988 cells) collected from mobilized peripheral CD34+ hematopoietic stem and progenitor cells (HSPCs) isolated from healthy human donors (we use the train dataset: 3 days, 3 donors, 7 cell types).\n<li> RDS files with top300 lists of RNA that correlate positively or negatively with proteins and/or RNA of 112 protein-RNA pairs (lists were prepared in the <a href= \"https://www.kaggle.com/code/antoninadolgorukova/mmscel-difference-in-prot-rna-correlations-p1\">Part 1 notebook</a>).\n\n</div>\n<hr>","metadata":{}},{"cell_type":"code","source":"#function for figure size adjusment\nfig <- function(width, heigth) {\n    options(repr.plot.width = width, repr.plot.height = heigth)\n}\n\nsuppressPackageStartupMessages({\n    library(Matrix)\n    library(stringr)\n    library(ggplot2)\n    library(tidyr)\n\n    library(patchwork) #arrange plots\n\n    BiocManager::install('org.Hs.eg.db')\n    library(org.Hs.eg.db)\n\n    BiocManager::install('clusterProfiler')\n    library(clusterProfiler)\n\n    BiocManager::install(\"ggnewscale\")\n    library(ggnewscale) \n    library(enrichplot)\n\n    BiocManager::install(\"ReactomePA\")\n    library(\"ReactomePA\")\n    \n    BiocManager::install(\"GOSim\")\n    library(GOSim)\n    \n})","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-21T20:30:41.111608Z","iopub.execute_input":"2023-03-21T20:30:41.113255Z","iopub.status.idle":"2023-03-21T20:30:41.125761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Load data**","metadata":{}},{"cell_type":"code","source":"path <- \"/kaggle/input/sparse-measurement-data-open-problems-multimodal/sp_train_cite_inputs.rds\"\nmat_RNA <- readRDS(path)\n\ncat(\"\\nRNA data:\", ncol(mat_RNA),\n    \"genes in columns and\", nrow(mat_RNA), \"cells in rows\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:30:41.129182Z","iopub.execute_input":"2023-03-21T20:30:41.130546Z","iopub.status.idle":"2023-03-21T20:31:22.314759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"n = 300\ndat <- read.csv(paste0(\"/kaggle/input/mmscel-difference-in-prot-rna-correlations-p1/top\", n, \"_pos_norib_mit_EIF_corr.csv\"))\n\ncat(\"\\n\\nCorrelation coefficients for pairs with CD45 isoforms\")\npos_dat <- dat[grep(\"CD45R\", dat$Pair), ]\npos_dat %>% pivot_wider(names_from = \"RNA_list\", id_cols = -RNA,\n                                 values_from = \"corr.coef\",\n                                 values_fn = function(x) paste(round(min(x), 3), \"-\", round(max(x), 3), \", n = \", length(x)) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.318002Z","iopub.execute_input":"2023-03-21T20:31:22.319276Z","iopub.status.idle":"2023-03-21T20:31:22.553253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Select correlations above 0.1**","metadata":{}},{"cell_type":"code","source":"pos_dat <- pos_dat[pos_dat$corr.coef >= 0.1, ]\n\ncat(\"Top\", n, \"lists of RNA that correlate positively with proteins and/or RNA of protein-RNA pairs were matched and separated by:\n\\nCommon RNA - present in both top\", n, \"lists (for the protein and the RNA of the pair)\n\\nUnique for protein - present only in the top\", n, \"list for the  protein\n\\nUnique for RNA - present only in the top\", n, \"list for the  RNA\")\n\npos_dat %>%\n    pivot_wider(id_cols = !corr.coef, names_from = \"RNA_list\",\n                values_from = \"RNA\", values_fn = list)\n\ncat(\"Correlation  coefficients and sample sizes\")\npos_dat %>% pivot_wider(names_from = \"RNA_list\", id_cols = -RNA,\n                                 values_from = \"corr.coef\",\n                                 values_fn = function(x) paste(round(min(x), 3), \"-\", round(max(x), 3), \", n = \", length(x)) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.556539Z","iopub.execute_input":"2023-03-21T20:31:22.557800Z","iopub.status.idle":"2023-03-21T20:31:22.617851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"n = 300\ndat <- read.csv(paste0(\"/kaggle/input/mmscel-difference-in-prot-rna-correlations-p1/top\", n, \"_neg_norib_mit_EIF_corr.csv\"))\n\ncat(\"\\n\\nCorrelation coefficients for pairs with CD45 isoforms\")\nneg_dat <- dat[grep(\"CD45R\", dat$Pair), ]\nneg_dat <- neg_dat[order(neg_dat$RNA_list, -neg_dat$corr.coef), ]\nneg_dat %>% pivot_wider(names_from = \"RNA_list\", id_cols = -RNA,\n                                 values_from = \"corr.coef\",\n                                 values_fn = function(x) paste(round(min(x), 3), \"to\", round(max(x), 3), \", n = \", length(x)) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.620962Z","iopub.execute_input":"2023-03-21T20:31:22.622108Z","iopub.status.idle":"2023-03-21T20:31:22.876823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Select correlations below 0.1**","metadata":{}},{"cell_type":"code","source":"neg_dat <- neg_dat[neg_dat$corr.coef <= -0.1, ]\n\ncat(\"Top\", n, \"lists of RNA that correlate negatively with proteins and/or RNA of protein-RNA pairs were matched and separated by:\n\\nCommon RNA - present in both top\", n, \"lists (for the protein and the RNA of the pair)\n\\nUnique for protein - present only in the top\", n, \"list for the  protein\n\\nUnique for RNA - present only in the top\", n, \"list for the  RNA\")\nneg_dat %>%\n    pivot_wider(id_cols = !corr.coef, names_from = \"RNA_list\",\n                values_from = \"RNA\", values_fn = list)\n\ncat(\"Correlation  coefficients and sample sizes\")\nneg_dat %>% pivot_wider(names_from = \"RNA_list\", id_cols = -RNA,\n                                 values_from = \"corr.coef\",\n                                 values_fn = function(x) paste(round(min(x), 3), \"-\", round(max(x), 3), \", n = \", length(x)) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.879996Z","iopub.execute_input":"2023-03-21T20:31:22.881337Z","iopub.status.idle":"2023-03-21T20:31:22.940169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#For convenience we remove full name of the RNA for the analysis.\npos_dat$Pair <- gsub(\" -.*\", \"\", pos_dat$Pair) #%>% head\nneg_dat$Pair <- gsub(\" -.*\", \"\", neg_dat$Pair) #%>% head","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.943097Z","iopub.execute_input":"2023-03-21T20:31:22.944366Z","iopub.status.idle":"2023-03-21T20:31:22.959094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Make a list af all genes in the dataset with Entrez IDs**","metadata":{}},{"cell_type":"code","source":"geneUniverse <- gsub(\"_.*\",\"\", colnames(mat_RNA)) #ENSEMBL\n#geneUniverse <- gsub(\".*_\",\"\", colnames(mat_RNA)) #SYMBOL\n\n#get Entrez IDs of genes\ngeneUniverse <- AnnotationDbi::select(org.Hs.eg.db, keys = geneUniverse,\n                                  columns = c('ENTREZID'), keytype = 'ENSEMBL') #SYMBOL\ncat(\"All RNA (genes collection): For\", sum(is.na(geneUniverse$ENTREZID)),\"out of\", nrow(geneUniverse), \"genes the Entrez ID was not found\")\n#geneUniverse[is.na(geneUniverse$ENTREZID), ]$ENTREZID\n\ngeneUniverse %>% head","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:22.962203Z","iopub.execute_input":"2023-03-21T20:31:22.963401Z","iopub.status.idle":"2023-03-21T20:31:23.252931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Add missing Gene IDs manually**","metadata":{}},{"cell_type":"code","source":"gen_id <- read.csv(\"/kaggle/input/research-project-01-around-multimodal-singlecell/convers_TXNAME_GENEID_SYMBOL_ENSG.csv\")\ncat(\"The avalilable data of the data for conversion\")\ngen_id <- gen_id[gen_id$ENSG %in% geneUniverse[is.na(geneUniverse$ENTREZID), ]$ENSEMBL, ]\ngen_id %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:23.256097Z","iopub.execute_input":"2023-03-21T20:31:23.257311Z","iopub.status.idle":"2023-03-21T20:31:23.589550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for(s in gen_id$ENSEMBL) {\n    \n    geneUniverse[geneUniverse$ENSEMBL == s, ]$ENTREZID <-\n        unique(gen_id[gen_id$ENSG == s, ]$GENEID)\n}\ncat(\"the remaining missimg gene IDs\")\ngeneUniverse[is.na(geneUniverse$ENTREZID), ] %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:23.592618Z","iopub.execute_input":"2023-03-21T20:31:23.593868Z","iopub.status.idle":"2023-03-21T20:31:23.620050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"background-color:#1C588C;font-family:Verdana;color:white;font-size:80%;text-align:left;border-radius: 15px;padding:10px 15px\">Common RNAs for CD45RO and CD45RA proteins and their RNA</p>","metadata":{}},{"cell_type":"markdown","source":"**Select the set of genes for enrichment analysis**","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\nsetRO_pos <- pos_dat[pos_dat$Pair == \"CD45RO\" & pos_dat$RNA_list == \"Common for prot and RNA\", \"RNA\"]\ncat(\"Common RNA for CD45RO and PTPRC:\", length(setRO_pos))\nsetRO_pos\nsetRA_pos <- pos_dat[pos_dat$Pair == \"CD45RA\" & pos_dat$RNA_list == \"Common for prot and RNA\", \"RNA\"]\ncat(\"Common RNA for CD45RA and PTPRC:\", length(setRA_pos))\nsetRA_pos\n\n# check if there are intersectoins\ncat(\"Number of intersections:\")\nintersect(setRO_pos, setRA_pos) %>% length\n\ncat(\"Total genes for enrichment analysis:\")\nc(setRO_pos, setRA_pos) %>% length\n\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\nsetRO_neg <- neg_dat[neg_dat$Pair == \"CD45RO\" & neg_dat$RNA_list == \"Common for prot and RNA\", \"RNA\"]\ncat(\"Common RNA for CD45RO and PTPRC:\", length(setRO_neg))\nsetRO_neg\nsetRA_neg <- neg_dat[neg_dat$Pair == \"CD45RA\" & neg_dat$RNA_list == \"Common for prot and RNA\", \"RNA\"]\ncat(\"Common RNA for CD45RA and PTPRC:\", length(setRA_neg))\nsetRA_neg\n\n# check if there are intersectoins\ncat(\"Number of intersections:\")\nintersect(setRO_neg, setRA_neg) %>% length\n\ncat(\"Total genes for enrichment analysis:\")\nc(setRO_neg, setRA_neg) %>% length","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:23.623029Z","iopub.execute_input":"2023-03-21T20:31:23.624284Z","iopub.status.idle":"2023-03-21T20:31:23.717898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Get Entrez IDs of genes**","metadata":{}},{"cell_type":"code","source":"#positive\ngene_set_pos <- c(setRO_pos, setRA_pos)\n\n#get Entrez IDs of genes\ngene_set_pos <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_set_pos,\n                                  columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"For\", sum(is.na(gene_set_pos$ENTREZID)),\"out of\", nrow(gene_set_pos), \"genes the Entrez ID was not found\")\ngene_set_pos[is.na(gene_set_pos$ENTREZID), ]\n\n#negative\ngene_set_neg <- c(setRO_neg, setRA_neg)\n\n#get Entrez IDs of genes\ngene_set_neg <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_set_neg,\n                                      columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"For\", sum(is.na(gene_set_neg$ENTREZID)),\"out of\", nrow(gene_set_neg), \"genes the Entrez ID was not found\")\ngene_set_neg[is.na(gene_set_neg$ENTREZID), ]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:23.721035Z","iopub.execute_input":"2023-03-21T20:31:23.722283Z","iopub.status.idle":"2023-03-21T20:31:24.026870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Add missing Gene IDs manually or omit\n#gene_set_pos <- na.omit(gene_set_pos)\n#gene_set_neg <- na.omit(gene_set_neg)\n\ngeneUniverse <- na.omit(geneUniverse)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:24.030174Z","iopub.execute_input":"2023-03-21T20:31:24.031581Z","iopub.status.idle":"2023-03-21T20:31:24.047021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">GO enrichment over-representation test (biological processes)</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"cat(\"Common RNA for both top\", n, \"lists (for the protein and the RNA of the protein-RNA pairs) with CD45:\", nrow(gene_set_pos) )\ngene_set_pos$SYMBOL\n\ngo_pos <- enrichGO(gene = gene_set_pos$ENTREZID, ont = \"BP\",\n                   OrgDb =\"org.Hs.eg.db\",\n                   universe = geneUniverse$ENTREZID,\n                   pvalueCutoff = 0.05,\n                  readable = TRUE)\ndf_go_pos <- go_pos %>% as.data.frame\n\ncat(\"GO enrichment analysis:\\n\")\ncat(nrow(df_go_pos), \"enriched terms found, top 10 are shown\")\n\ndf_go_pos$geneID <- str_wrap(df_go_pos$geneID, width=30)\ndf_go_pos %>% head(10)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:24.050201Z","iopub.execute_input":"2023-03-21T20:31:24.051555Z","iopub.status.idle":"2023-03-21T20:31:34.556934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[From the tutorial:](http://yulab-smu.top/biomedical-knowledge-mining-book/useful-utilities.html) GO is organized in parent-child structure, thus a parent term can be overlapped with a large proportion with all its child terms. This can result in redundant findings. To solve this issue, clusterProfiler implements simplify method to reduce redundant GO terms from the outputs of enrichGO() and gseGO(). The function internally called GOSemSim (Yu et al. 2010) to calculate semantic similarity among GO terms and remove those highly similar terms by keeping one representative term. The simplify() method apply select_fun (which can be a user defined function) to feature by to select one representative term from redundant terms (which have similarity higher than cutoff).","metadata":{}},{"cell_type":"code","source":"goCom_pos <- pairwise_termsim(go_pos)\ngoCom_pos <- simplify(goCom_pos, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoCom_pos %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:34.560110Z","iopub.execute_input":"2023-03-21T20:31:34.561517Z","iopub.status.idle":"2023-03-21T20:31:35.444753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"cat(\"Common RNA for both top\", n, \"lists (for the protein and the RNA of the protein-RNA pairs) with CD45:\", nrow(gene_set_neg) )\ngene_set_neg$SYMBOL\n\ngo_neg <- enrichGO(gene = gene_set_neg$ENTREZID, ont = \"BP\",\n                   OrgDb =\"org.Hs.eg.db\",\n                   universe = geneUniverse$ENTREZID,\n                   pvalueCutoff = 0.05,\n                   readable = TRUE)\ndf_go_neg <- go_neg %>% as.data.frame\n\ncat(\"GO enrichment analysis:\\n\")\ncat(nrow(df_go_neg), \"enriched terms found\")\n\ndf_go_neg$geneID <- str_wrap(df_go_neg$geneID, width=30)\n#df_go_neg %>% head(10)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:35.447926Z","iopub.execute_input":"2023-03-21T20:31:35.450024Z","iopub.status.idle":"2023-03-21T20:31:44.891417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Visualize enriched GO terms**","metadata":{}},{"cell_type":"code","source":"fig(30,25)\nn = goCom_pos %>% as.data.frame %>% nrow\nplt1 <- dotplot(goCom_pos, showCategory = n, orderBy=\"GeneRatio\") + \n    scale_y_discrete(labels=function(x) str_wrap(x, width=70)) + \n    theme_bw(base_size = 20) \nplt2 <- barplot(goCom_pos, showCategory = n) + \n    scale_y_discrete(labels=function(x) str_wrap(x, width=70)) + \n    theme_bw(base_size = 20)\n                     \nplt1 + plt2 + plot_annotation(\n    title = paste0(\"Common RNAs for CD45RO and CD45RA proteins and their RNA\\nPositive correlations\\nEnriched terms shown: \",n, \" of \", goCom_pos %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:44.894542Z","iopub.execute_input":"2023-03-21T20:31:44.895794Z","iopub.status.idle":"2023-03-21T20:31:46.194913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(30,25)\nplt1 <- goplot(goCom_pos, showCategory=n) + \n    theme_void(base_size = 22) \nplt2 <- emapplot(goCom_pos) + \n    theme_void(base_size = 22)\n\nplt1 + plt2 + plot_annotation(\n    title = paste0(\"Common RNAs for CD45RO and CD45RA proteins and their RNA\\nPositive correlations\\nEnriched terms shown: \", n, \" of \", goCom_pos %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:46.197952Z","iopub.execute_input":"2023-03-21T20:31:46.199876Z","iopub.status.idle":"2023-03-21T20:31:53.267147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\np1 <- cnetplot(DOSE::setReadable(goCom_pos, \"org.Hs.eg.db\", \"ENSEMBL\"),\n         categorySize = \"pvalue\",\n         showCategory = n) + \n    theme_void(base_size = 22) \np1 + plot_annotation(\n    title = paste0(\"Common RNAs for CD45RO and CD45RA proteins and their RNA\\nPositive correlations\\nEnriched terms shown: \",n, \" of \", goCom_pos %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:53.270492Z","iopub.execute_input":"2023-03-21T20:31:53.272277Z","iopub.status.idle":"2023-03-21T20:31:58.979197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,28) \np1 <- heatplot(goCom_pos, showCategory = n) +\n    scale_y_discrete(labels=function(x) str_wrap(x, width=70)) + \n    theme_bw(base_size = 20) +\n    theme(axis.text.x = element_text(angle = 45, hjust=1)) \np1 + plot_annotation(\n    title = paste0(\"Common RNAs for CD45RO and CD45RA proteins and their RNA\\nPositive correlations\\nEnriched terms shown: \",n, \" of \", goCom_pos %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) )                  ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:31:58.982333Z","iopub.execute_input":"2023-03-21T20:31:58.984246Z","iopub.status.idle":"2023-03-21T20:32:00.310806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">GO Gene Set Enrichment Analysis</p>","metadata":{}},{"cell_type":"markdown","source":"**Prepare geneList**","metadata":{}},{"cell_type":"code","source":"setRO <- pos_dat[pos_dat$Pair == \"CD45RO\" & pos_dat$RNA_list == \"Common for prot and RNA\", ]\nsetRO <- select(setRO, c(\"RNA\", \"corr.coef\"))\n\nsetRA <- pos_dat[pos_dat$Pair == \"CD45RA\" & pos_dat$RNA_list == \"Common for prot and RNA\", ]\nsetRA <- select(setRA, c(\"RNA\", \"corr.coef\"))\n\nsetRO_RA <- rbind(setRO, setRA)\nsetRO_RA$RNA <- gene_set_pos$ENTREZID\n\ngeneList <- setRO_RA[,2]\nnames(geneList) = as.character(setRO_RA[,1])\ngeneList = sort(geneList, decreasing = TRUE)\ngeneList %>% head","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:00.314287Z","iopub.execute_input":"2023-03-21T20:32:00.322965Z","iopub.status.idle":"2023-03-21T20:32:00.391265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"go_set <- gseGO(geneList    = geneList,\n              OrgDb        = org.Hs.eg.db,\n              ont          = \"BP\",\n              scoreType = \"pos\",\n              minGSSize    = 100,\n              maxGSSize    = 500,\n              pvalueCutoff = 0.05,\n              verbose      = FALSE)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:00.396016Z","iopub.execute_input":"2023-03-21T20:32:00.399298Z","iopub.status.idle":"2023-03-21T20:32:10.213897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is not enough data for GO Gene Set Enrichment Analysis","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">KEGG Gene Set Enrichment Analysis</p>","metadata":{}},{"cell_type":"markdown","source":"Check if KEGG Gene Set Enrichment Analysis is possible","metadata":{}},{"cell_type":"code","source":"bitr_kegg(gene_set_pos$ENTREZID, fromType = \"kegg\", toType = \"Path\", organism = \"hsa\")\nbitr_kegg(gene_set_pos$ENTREZID, fromType = \"kegg\", toType = \"Module\", organism = \"hsa\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:10.217031Z","iopub.execute_input":"2023-03-21T20:32:10.218248Z","iopub.status.idle":"2023-03-21T20:32:17.790249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is not enough data for KEGG Gene Set Enrichment Analysis","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">Reactome pathway over-representation analysis</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"epath_pos <- enrichPathway(gene = gene_set_pos$ENTREZID, pvalueCutoff = 0.05, readable=TRUE)\n\ndf_epath_pos <- epath_pos %>% as.data.frame\n\ncat(\"Reactome pathway over-representation analysis:\\n\")\ncat(nrow(df_epath_pos), \"enriched terms found\")\ndf_epath_pos","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:17.793347Z","iopub.execute_input":"2023-03-21T20:32:17.794758Z","iopub.status.idle":"2023-03-21T20:32:24.490333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"epath_neg <- enrichPathway(gene = gene_set_neg$ENTREZID, pvalueCutoff = 0.05, readable=TRUE)\n\ndf_epath_neg <- epath_neg %>% as.data.frame\n\ncat(\"Reactome pathway over-representation analysis:\\n\")\ncat(nrow(df_epath_neg), \"enriched terms found\")\ndf_epath_neg","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:24.493344Z","iopub.execute_input":"2023-03-21T20:32:24.494652Z","iopub.status.idle":"2023-03-21T20:32:32.348430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">Reactome pathway gene set enrichment analysis</p>","metadata":{}},{"cell_type":"code","source":"gse_path_pos <- gsePathway(geneList, \n                pvalueCutoff = 0.2,\n                scoreType = \"pos\",\n                pAdjustMethod = \"BH\", \n                verbose = FALSE)\n\ndf_gse_path_pos <- gse_path_pos %>% as.data.frame\n\ncat(\"Reactome pathway gene set enrichment analysis:\\n\")\ncat(nrow(df_gse_path_pos), \"enriched terms found\")\ndf_gse_path_pos","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:32.351526Z","iopub.execute_input":"2023-03-21T20:32:32.352860Z","iopub.status.idle":"2023-03-21T20:32:39.326181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">Pathway Visualization</p>","metadata":{}},{"cell_type":"code","source":"# viewPathway(\"Alternative splicing\", \n#              readable = TRUE, \n#              foldChange = geneList) #","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:39.329407Z","iopub.execute_input":"2023-03-21T20:32:39.330687Z","iopub.status.idle":"2023-03-21T20:32:39.341330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is not enough data","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"background-color:#1C588C;font-family:Verdana;color:white;font-size:80%;text-align:left;border-radius: 15px;padding:10px 15px\">Unique RNA for CD45RO and CD45RA proteins and RNAs</p>","metadata":{}},{"cell_type":"markdown","source":"**Select the set of genes for enrichment analysis**","metadata":{}},{"cell_type":"markdown","source":"We use RNA lists from the main table above with unique RNA for CD45 proteins and pooled list with RNA for CD45 RNA to see if they form different enrichment terms. This may help us to differentiate regulation of:\n- CD45 proteins (RO vs RA)\n- CD45RO (protein vs RNA PTPRC)\n- CD45RA (protein vs RNA PTPRC)","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\n#unique RNA for proteins CD45RO and CD45RA\nsetRO_pos <- pos_dat[pos_dat$Pair == \"CD45RO\" & pos_dat$RNA_list == \"Unique for protein\", \"RNA\"]\nsetRA_pos <- pos_dat[pos_dat$Pair == \"CD45RA\" & pos_dat$RNA_list == \"Unique for protein\", \"RNA\"]\n\n# check if there are intersectoins\ncat(\"Number of intersections in RNA subsets for CD45 proteins:\")\nintersect(setRO_pos, setRA_pos) %>% length\nintersect(setRO_pos, setRA_pos)\n\n#Pooled for RNA of the CD45RO and CD45RA\nsetROrna_pos <- pos_dat[pos_dat$Pair == \"CD45RO\" & pos_dat$RNA_list == \"Unique for RNA\", \"RNA\"]\nsetRArna_pos <- pos_dat[pos_dat$Pair == \"CD45RA\" & pos_dat$RNA_list == \"Unique for RNA\", \"RNA\"]\n\n# check if there are intersectoins\ncat(\"Number of intersections in RNA subsets for PTPRC:\")\nintersect(setROrna_pos, setRArna_pos) %>% length\n\ncat(\"The pooled subset of RNA for PTPRC:\")\nunique(c(setROrna_pos, setRArna_pos))  %>% length\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\n#unique RNA for proteins CD45RO and CD45RA\nsetRO_neg <- neg_dat[neg_dat$Pair == \"CD45RO\" & neg_dat$RNA_list == \"Unique for protein\", \"RNA\"]\nsetRA_neg <- neg_dat[neg_dat$Pair == \"CD45RA\" & neg_dat$RNA_list == \"Unique for protein\", \"RNA\"]\n\n# check if there are intersectoins\ncat(\"Number of intersections in RNA subsets for CD45 proteins:\")\nintersect(setRO_neg, setRA_neg) %>% length\nintersect(setRO_neg, setRA_neg)\n\n#Pooled for RNA of the CD45RO and CD45RA\nsetROrna_neg <- neg_dat[neg_dat$Pair == \"CD45RO\" & neg_dat$RNA_list == \"Unique for RNA\", \"RNA\"]\nsetRArna_neg <- neg_dat[neg_dat$Pair == \"CD45RA\" & neg_dat$RNA_list == \"Unique for RNA\", \"RNA\"]\n\n# check if there are intersectoins\ncat(\"Number of intersections in RNA subsets for PTPRC:\")\nintersect(setROrna_neg, setRArna_neg) %>% length\n\ncat(\"The pooled subset of RNA for PTPRC:\")\nunique(c(setROrna_neg, setRArna_neg))  %>% length","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:39.344533Z","iopub.execute_input":"2023-03-21T20:32:39.345849Z","iopub.status.idle":"2023-03-21T20:32:39.435662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Get Entrez IDs of genes**","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\n# subset unique RNAs\ngene_setRO_pos <- c(setRO_pos)\ngene_setRA_pos <- c(setRA_pos)\n\n#c(setdiff(setROrna, setRArna))\n#c(setdiff(setRArna, setROrna))\ngene_setRNA_pos <- unique(c(setROrna_pos, setRArna_pos))\n\n#get Entrez IDs of genes\n\n#RO prot\ngene_setRO_pos <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRO_pos,\n                                    columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RO protein: For\", sum(is.na(gene_setRO_pos$ENTREZID)),\"out of\", nrow(gene_setRO_pos), \"genes the Entrez ID is not found\")\ngene_setRO_pos[is.na(gene_setRO_pos$ENTREZID), ]\n\n#RA prot\ngene_setRA_pos <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRA_pos,\n                                    columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RA protein: For\", sum(is.na(gene_setRA_pos$ENTREZID)),\"out of\", nrow(gene_setRA_pos), \"genes the Entrez ID is not found\")\ngene_setRA_pos[is.na(gene_setRA_pos$ENTREZID), ]\n\n#RNA\ngene_setRNA_pos <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRNA_pos,\n                                          columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RO - CD45RA RNA: For\", sum(is.na(gene_setRNA_pos$ENTREZID)),\"out of\", nrow(gene_setRNA_pos), \"genes the Entrez ID is not found\")\ngene_setRNA_pos[is.na(gene_setRNA_pos$ENTREZID), ]\n\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\n# subset unique RNAs\ngene_setRO_neg <- c(setRO_neg)\ngene_setRA_neg <- c(setRA_neg)\n\n#c(setdiff(setROrna, setRArna))\n#c(setdiff(setRArna, setROrna))\ngene_setRNA_neg <- unique(c(setROrna_neg, setRArna_neg))\n\n#get Entrez IDs of genes\n\n#RO prot\ngene_setRO_neg <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRO_neg,\n                                        columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RO protein: For\", sum(is.na(gene_setRO_neg$ENTREZID)),\"out of\", nrow(gene_setRO_neg), \"genes the Entrez ID is not found\")\ngene_setRO_neg[is.na(gene_setRO_neg$ENTREZID), ]\n\n#RA prot\ngene_setRA_neg <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRA_neg,\n                                        columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RA protein: For\", sum(is.na(gene_setRA_neg$ENTREZID)),\"out of\", nrow(gene_setRA_neg), \"genes the Entrez ID is not found\")\ngene_setRA_neg[is.na(gene_setRA_neg$ENTREZID), ]\n\n#RNA\ngene_setRNA_neg <- AnnotationDbi::select(org.Hs.eg.db, keys = gene_setRNA_neg,\n                                         columns = c('ENTREZID'), keytype = 'SYMBOL')\ncat(\"Subset CD45RO - CD45RA RNA: For\", sum(is.na(gene_setRNA_neg$ENTREZID)),\"out of\", nrow(gene_setRNA_neg), \"genes the Entrez ID is not found\")\ngene_setRNA_neg[is.na(gene_setRNA_neg$ENTREZID), ]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:39.438855Z","iopub.execute_input":"2023-03-21T20:32:39.440133Z","iopub.status.idle":"2023-03-21T20:32:40.316316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Add missing Gene IDs manually or omit**","metadata":{}},{"cell_type":"code","source":"gene_setRO_pos <- na.omit(gene_setRO_pos)\ngene_setRA_pos <- na.omit(gene_setRA_pos)\ngene_setRNA_pos <- na.omit(gene_setRNA_pos)\n\ngene_setRO_neg <- na.omit(gene_setRO_neg)\ngene_setRA_neg <- na.omit(gene_setRA_neg)\ngene_setRNA_neg <- na.omit(gene_setRNA_neg)\n\n#gene_set[gene_set$SYMBOL == \"ATP5E\", ]$ENTREZID <- \"514\"\n#gene_set[gene_set$SYMBOL == \"C10orf54\", ]$ENTREZID <- \"64115\"\n#gene_set[gene_set$SYMBOL == \"TCEB2\", ]$ENTREZID <- \"6923\"\n#gene_set[gene_set$SYMBOL == \"GLTSCR2\", ]$ENTREZID <- \"29997\"\n#gene_set[gene_set$SYMBOL == \"FAM65B\", ]$ENTREZID <- \"6923\"\n\n#geneUniverse[geneUniverse$SYMBOL == \"ATP5E\", ]$ENTREZID <- \"514\"\n#geneUniverse[geneUniverse$SYMBOL == \"C10orf54\", ]$ENTREZID <- \"64115\"\n#geneUniverse[geneUniverse$SYMBOL == \"TCEB2\", ]$ENTREZID <- \"9750\"\n#geneUniverse[geneUniverse$SYMBOL == \"GLTSCR2\", ]$ENTREZID <- \"29997\"\n#geneUniverse[geneUniverse$SYMBOL == \"FAM65B\", ]$ENTREZID <- \"9750\"","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:40.319442Z","iopub.execute_input":"2023-03-21T20:32:40.320692Z","iopub.status.idle":"2023-03-21T20:32:40.341742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">GO enrichment analysis (biological processes) for CD45RO</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRO_pos\n\ncat(\"Unique RNA for CD45RO protein:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRO_pos <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                   OrgDb =\"org.Hs.eg.db\",\n                   universe = geneUniverse$ENTREZID,\n                   pvalueCutoff = 0.05,\n                  readable = TRUE)\ndf_go <- goRO_pos %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for CD45RO:\\n\")\ncat(nrow(df_go), \"enriched terms found\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:40.344721Z","iopub.execute_input":"2023-03-21T20:32:40.345957Z","iopub.status.idle":"2023-03-21T20:32:50.405343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRO_pos <- pairwise_termsim(goRO_pos)\ngoRO_pos <- simplify(goRO_pos, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRO_pos %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:50.408488Z","iopub.execute_input":"2023-03-21T20:32:50.409770Z","iopub.status.idle":"2023-03-21T20:32:50.442675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRO_neg\n\ncat(\"Unique RNA for CD45RO protein:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRO_neg <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                     OrgDb =\"org.Hs.eg.db\",\n                     universe = geneUniverse$ENTREZID,\n                     pvalueCutoff = 0.05,\n                     readable = TRUE)\ndf_go <- goRO_neg %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for CD45RO:\\n\")\ncat(nrow(df_go), \"enriched terms found, top 10 are shown\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% head(10) %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:32:50.445492Z","iopub.execute_input":"2023-03-21T20:32:50.446711Z","iopub.status.idle":"2023-03-21T20:33:01.065653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRO_neg <- pairwise_termsim(goRO_neg)\ngoRO_neg <- simplify(goRO_neg, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRO_neg %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:01.069210Z","iopub.execute_input":"2023-03-21T20:33:01.070561Z","iopub.status.idle":"2023-03-21T20:33:01.237574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">GO enrichment analysis (biological processes) for CD45RA</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRA_pos\n\ncat(\"Unique RNA for CD45RA protein:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRA_pos <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                     OrgDb =\"org.Hs.eg.db\",\n                     universe = geneUniverse$ENTREZID,\n                     pvalueCutoff = 0.05,\n                     readable = TRUE)\ndf_go <- goRA_pos %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for CD45RA:\\n\")\ncat(nrow(df_go), \"enriched terms found, top 10 are shown\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% head(10) %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:01.240751Z","iopub.execute_input":"2023-03-21T20:33:01.241963Z","iopub.status.idle":"2023-03-21T20:33:14.243180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRA_pos <- pairwise_termsim(goRA_pos)\ngoRA_pos <- simplify(goRA_pos, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRA_pos %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:14.246326Z","iopub.execute_input":"2023-03-21T20:33:14.247646Z","iopub.status.idle":"2023-03-21T20:33:16.232061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRA_neg\n\ncat(\"Unique RNA for CD45RA protein:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRA_neg <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                     OrgDb =\"org.Hs.eg.db\",\n                     universe = geneUniverse$ENTREZID,\n                     pvalueCutoff = 0.05,\n                     readable = TRUE)\ndf_go <- goRA_neg %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for CD45RA:\\n\")\ncat(nrow(df_go), \"enriched terms found, top 10 are shown\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% head(10) %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:16.235159Z","iopub.execute_input":"2023-03-21T20:33:16.236396Z","iopub.status.idle":"2023-03-21T20:33:28.046522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRA_neg <- pairwise_termsim(goRA_neg)\ngoRA_neg <- simplify(goRA_neg, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRA_neg %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:28.049473Z","iopub.execute_input":"2023-03-21T20:33:28.050873Z","iopub.status.idle":"2023-03-21T20:33:28.317163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">GO enrichment analysis (biological processes) for PTPRC (CD45 RNA)</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRNA_pos\n\ncat(\"Pooled RNA for PTPRC:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRNA_pos <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                     OrgDb =\"org.Hs.eg.db\",\n                     universe = geneUniverse$ENTREZID,\n                     pvalueCutoff = 0.05,\n                     readable = TRUE)\ndf_go <- goRNA_pos %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for PTPRC:\\n\")\ncat(nrow(df_go), \"enriched terms found, top 10 are shown\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% head(10) %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:28.320292Z","iopub.execute_input":"2023-03-21T20:33:28.321585Z","iopub.status.idle":"2023-03-21T20:33:42.232718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRNA_pos <- pairwise_termsim(goRNA_pos)\ngoRNA_pos <- simplify(goRNA_pos, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRNA_pos %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:42.236070Z","iopub.execute_input":"2023-03-21T20:33:42.237459Z","iopub.status.idle":"2023-03-21T20:33:52.770348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"gene_set <- gene_setRNA_neg\n\ncat(\"Pooled RNA for PTPRC:\", nrow(gene_set) )\ngene_set$SYMBOL\n\ngoRNA_neg <- enrichGO(gene = gene_set$ENTREZID, ont = \"BP\",\n                      OrgDb =\"org.Hs.eg.db\",\n                      universe = geneUniverse$ENTREZID,\n                      pvalueCutoff = 0.05,\n                      readable = TRUE)\ndf_go <- goRNA_neg %>% as.data.frame\nrownames(df_go) <- NULL\n\ncat(\"GO enrichment analysis for PTPRC:\\n\")\ncat(nrow(df_go), \"enriched terms found, top 10 are shown\")\n\ndf_go$geneID <- str_wrap(df_go$geneID, width=30)\ndf_go %>% head(10) %>% select(1,2,3,4,5,8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:33:52.773562Z","iopub.execute_input":"2023-03-21T20:33:52.774829Z","iopub.status.idle":"2023-03-21T20:34:05.727029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"goRNA_neg <- pairwise_termsim(goRNA_neg)\ngoRNA_neg <- simplify(goRNA_neg, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ngoRNA_neg %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T20:34:05.731973Z","iopub.execute_input":"2023-03-21T20:34:05.735565Z","iopub.status.idle":"2023-03-21T20:34:20.331867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"background-color:#1C588C;font-family:Verdana;color:white;font-size:80%;text-align:left;border-radius: 15px;padding:10px 15px\">Compare enriched GO terms</p>","metadata":{}},{"cell_type":"markdown","source":"**Positive correlations**","metadata":{}},{"cell_type":"code","source":"fig(25, 22)\n\nn = nrow(as.data.frame(goRO_pos))\nplt1 <- emapplot(goRO_pos, showCategory = n) + \n    theme_void(base_size = 22) +\n    ggtitle(paste0(\"CD45RO\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRO_pos))))\n\nn = 20\nplt2 <- emapplot(goRA_pos, showCategory = n) + \n    theme_void(base_size = 22) +\n    ggtitle(paste0(\"CD45RA\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRA_pos))))\n\nplt3 <- emapplot(goRNA_pos, showCategory = n) + \n    theme_void(base_size = 22) +\n    ggtitle(paste0(\"CD45 RNA\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRNA_pos))))\n\nplt3 | (plt2 / plt1 ) + plot_annotation(\n    title = paste0(\"Positively correlated RNA\")) & \ntheme(text = element_text(size = 22) )                  ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:38:36.385212Z","iopub.execute_input":"2023-03-21T22:38:36.386733Z","iopub.status.idle":"2023-03-21T22:38:39.620615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Negative correlations**","metadata":{}},{"cell_type":"code","source":"fig(25, 22)\nn =  nrow(as.data.frame(goRO_neg)) \n\nplt1 <- emapplot(goRO_neg, showCategory = n) + \n  theme_void(base_size = 22) +\n  ggtitle(paste0(\"CD45RO\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRO_neg))))\n\nn = nrow(as.data.frame(goRA_neg))\n\nplt2 <- emapplot(goRA_neg, showCategory = n) + \n  theme_void(base_size = 22) +\n  ggtitle(paste0(\"CD45RA\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRA_neg)))) \n\nn = 50\n\nplt3 <- emapplot(goRNA_neg, showCategory = n) + \n  theme_void(base_size = 22) +\n  ggtitle(paste0(\"CD45 RNA\\nEnriched terms shown: \",n, \" of \", nrow(as.data.frame(goRNA_neg))))\n\nplt3 | (plt2 / plt1 ) + plot_annotation(\n  title = paste0(\"Negatively correlated RNA\")) & \n  theme(text = element_text(size = 22) )  ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:38:39.623705Z","iopub.execute_input":"2023-03-21T22:38:39.625784Z","iopub.status.idle":"2023-03-21T22:38:46.081766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">1. CD45 proteins (RO vs RA)</p>","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RA and CD45RO proteins\")\ndata.frame(\n    \"Common for CD45RA and CD45RO proteins\" = paste(intersect(goRO_pos$Description, goRA_pos$Description), collapse = \";\\n\"),\n    \"Unique for CD45RO\" = paste(setdiff(goRO_pos$Description, goRA_pos$Description), collapse = \";\\n\"),\n    \"Unique for CD45RA\" = paste(setdiff(goRA_pos$Description, goRO_pos$Description), collapse = \";\\n\"),\n    check.names = FALSE\n)\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RA and CD45RO proteins\")\ndata.frame(\n  \"Common for CD45RA and CD45RO proteins\" = paste(intersect(goRO_neg$Description, goRA_neg$Description), collapse = \";\\n\"),\n  \"Unique for CD45RO\" = paste(setdiff(goRO_neg$Description, goRA_neg$Description), collapse = \";\\n\"),\n  \"Unique for CD45RA\" = paste(setdiff(goRA_neg$Description, goRO_neg$Description), collapse = \";\\n\"),\n  check.names = FALSE\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:38:46.084974Z","iopub.execute_input":"2023-03-21T22:38:46.086699Z","iopub.status.idle":"2023-03-21T22:38:46.159802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Comparing multiple gene lists**","metadata":{}},{"cell_type":"code","source":"proteins <- list(\"CD45RO_pos\" = gene_setRO_pos$ENTREZID, \n                \"CD45RA_pos\" = gene_setRA_pos$ENTREZID,\n                \"CD45RO_neg\" = gene_setRO_neg$ENTREZID, \n                \"CD45RA_neg\" = gene_setRA_neg$ENTREZID)\nstr(proteins)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:38:46.162904Z","iopub.execute_input":"2023-03-21T22:38:46.164182Z","iopub.status.idle":"2023-03-21T22:38:46.182332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichGO**","metadata":{}},{"cell_type":"code","source":"comp_lst <- compareCluster(geneCluster = proteins, OrgDb = org.Hs.eg.db, fun = enrichGO, ont = \"BP\")\ncomp_lst <- setReadable(comp_lst, OrgDb = org.Hs.eg.db, keyType=\"ENTREZID\")\n\ndf_go_prot <- comp_lst %>% as.data.frame\nrownames(df_go_prot) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RO_pos:\", nrow(df_go_prot[df_go_prot$Cluster == \"CD45RO_pos\",]), \"enriched terms found\",\n   \"\\nCD45RA_pos:\", nrow(df_go_prot[df_go_prot$Cluster == \"CD45RA_pos\",]), \"enriched terms found\",\n   \"\\nCD45RO_neg:\", nrow(df_go_prot[df_go_prot$Cluster == \"CD45RO_neg\",]), \"enriched terms found\",\n   \"\\nCD45RA_neg:\", nrow(df_go_prot[df_go_prot$Cluster == \"CD45RA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_go_prot))\n\ndf_go_prot$geneID <- str_wrap(df_go_prot$geneID, width=30)\ndf_go_prot %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:38:46.185651Z","iopub.execute_input":"2023-03-21T22:38:46.186966Z","iopub.status.idle":"2023-03-21T22:39:35.365607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp_lst <- simplify(comp_lst, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ncomp_lst %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:39:35.368550Z","iopub.execute_input":"2023-03-21T22:39:35.369742Z","iopub.status.idle":"2023-03-21T22:39:37.295758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,30)\nn = 20\ndotplot(comp_lst, showCategory = n) + \n    scale_y_discrete(labels=function(x) str_wrap(x, width=70)) + \n    theme_bw(base_size = 18)  + plot_annotation(\n    title = paste0(\"enrichGO\\nCD45 proteins (RO vs RA)\\nEnriched terms shown (per cluster): \",n, \" of \", comp_lst %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 20) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:39:37.298700Z","iopub.execute_input":"2023-03-21T22:39:37.299884Z","iopub.status.idle":"2023-03-21T22:39:38.320194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\nn = 30 #nrow(as.data.frame(comp_lst))\ncnetplot(comp_lst, showCategory = n) + \n    theme_void(base_size = 22)  + plot_annotation(\n    title = paste0(\"enrichGO\\nCD45 proteins (RO vs RA)\\nEnriched terms shown: \",n, \" of \", comp_lst %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:39:38.323295Z","iopub.execute_input":"2023-03-21T22:39:38.325381Z","iopub.status.idle":"2023-03-21T22:39:44.453410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichPathway**","metadata":{}},{"cell_type":"code","source":"comp_lst <- compareCluster(geneCluster = proteins, fun = enrichPathway, readable = TRUE)\n\ndf_path_prot <- comp_lst %>% as.data.frame\nrownames(df_path_prot) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RO_pos:\", nrow(df_path_prot[df_path_prot$Cluster == \"CD45RO_pos\",]), \"enriched terms found\",\n    \"\\nCD45RA_pos:\", nrow(df_path_prot[df_path_prot$Cluster == \"CD45RA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RO_neg:\", nrow(df_path_prot[df_path_prot$Cluster == \"CD45RO_neg\",]), \"enriched terms found\",\n    \"\\nCD45RA_neg:\", nrow(df_path_prot[df_path_prot$Cluster == \"CD45RA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_path_prot)\n)\n\ndf_path_prot$geneID <- str_wrap(df_path_prot$geneID, width=30)\ndf_path_prot %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:39:44.456648Z","iopub.execute_input":"2023-03-21T22:39:44.458489Z","iopub.status.idle":"2023-03-21T22:40:11.416225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\nn = nrow(df_path_prot)\ndotplot(comp_lst, showCategory=n) + \n    scale_y_discrete(labels=function(x) str_wrap(x, width=70)) + \n    theme_bw(base_size = 18)  + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45 proteins (RO vs RA)\\nEnriched terms shown: \",n, \" of \", nrow(df_path_prot))) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:40:11.419361Z","iopub.execute_input":"2023-03-21T22:40:11.420558Z","iopub.status.idle":"2023-03-21T22:40:12.230930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,15)\nn = nrow(df_path_prot)\ncnetplot(comp_lst, showCategory = n) + \n    theme_void(base_size = 22)  + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45 proteins (RO vs RA)\\nEnriched terms shown: \",n, \" of \", nrow(df_path_prot))) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:40:12.234226Z","iopub.execute_input":"2023-03-21T22:40:12.236049Z","iopub.status.idle":"2023-03-21T22:40:18.846488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">2. CD45RO vs its RNA</p>","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RO and its RNA\")\ndata.frame(\n    \"Common for CD45RO and its RNA\" = paste(intersect(goRO_pos$Description, goRNA_pos$Description), collapse = \";\\n\"),\n    \"Unique for CD45RO\" = paste(setdiff(goRO_pos$Description, goRNA_pos$Description), collapse = \";\\n\"),\n    \"Unique for RNA\" = paste(setdiff(goRNA_pos$Description, goRO_pos$Description), collapse = \";\\n\"),\n    check.names = FALSE\n)\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RO and its RNA\")\ndata.frame(\n  \"Common for CD45RO and its RNA\" = paste(intersect(goRO_neg$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for CD45RO\" = paste(setdiff(goRO_neg$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for RNA\" = paste(setdiff(goRNA_pos$Description, goRO_neg$Description), collapse = \";\\n\"),\n  check.names = FALSE\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:40:18.849766Z","iopub.execute_input":"2023-03-21T22:40:18.851736Z","iopub.status.idle":"2023-03-21T22:40:18.918037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Comparing multiple gene lists**","metadata":{}},{"cell_type":"code","source":"prot_RNA <- list(\"CD45RO_pos\" = gene_setRO_pos$ENTREZID, \n                \"CD45RNA_pos\" = gene_setRNA_pos$ENTREZID,\n                \"CD45RO_neg\" = gene_setRO_neg$ENTREZID, \n                \"CD45RNA_neg\" = gene_setRNA_neg$ENTREZID)\nstr(prot_RNA)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:40:18.921024Z","iopub.execute_input":"2023-03-21T22:40:18.922396Z","iopub.status.idle":"2023-03-21T22:40:18.939996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichGO**","metadata":{}},{"cell_type":"code","source":"comp_go_ROrna <- compareCluster(geneCluster = prot_RNA,OrgDb = org.Hs.eg.db, fun = enrichGO, ont = \"BP\")\ncomp_go_ROrna <- setReadable(comp_go_ROrna, OrgDb = org.Hs.eg.db, keyType=\"ENTREZID\")\n\ndf_go_ROrna <- comp_go_ROrna %>% as.data.frame\nrownames(df_go_ROrna) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RO_pos:\", nrow(df_go_ROrna[df_go_ROrna$Cluster == \"CD45RO_pos\",]), \"enriched terms found\",\n    \"\\nCD45RNA_pos:\", nrow(df_go_ROrna[df_go_ROrna$Cluster == \"CD45RNA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RO_neg:\", nrow(df_go_ROrna[df_go_ROrna$Cluster == \"CD45RO_neg\",]), \"enriched terms found\",\n    \"\\nCD45RNA_neg:\", nrow(df_go_ROrna[df_go_ROrna$Cluster == \"CD45RNA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_go_ROrna)\n)\n\ndf_go_ROrna$geneID <- str_wrap(df_go_ROrna$geneID, width=30)\ndf_go_ROrna %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:40:18.942860Z","iopub.execute_input":"2023-03-21T22:40:18.944117Z","iopub.status.idle":"2023-03-21T22:41:07.511649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp_go_ROrna <- simplify(comp_go_ROrna, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ncomp_go_ROrna %>% as.data.frame %>% nrow","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:07.514770Z","iopub.execute_input":"2023-03-21T22:41:07.515954Z","iopub.status.idle":"2023-03-21T22:41:21.499343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,22)\nn = 20\n\ndotplot(comp_go_ROrna, showCategory = n) + \n    scale_y_discrete(labels = function(x) str_wrap(x, width = 70)) + \n    theme_bw(base_size = 18) + plot_annotation(\n    title = paste0(\"enrichGo\\nCD45RO vs its RNA\\nEnriched terms shown (per cluster): \",n, \" of \", comp_go_ROrna %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 20) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:21.502421Z","iopub.execute_input":"2023-03-21T22:41:21.503770Z","iopub.status.idle":"2023-03-21T22:41:22.379985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\nn = 30\n\ncnetplot(comp_go_ROrna, showCategory = n) + \n    theme_void(base_size = 22) + plot_annotation(\n    title = paste0(\"enrichGo\\nCD45RO vs its RNA\\nEnriched terms shown: \",n, \" of \", comp_go_ROrna %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:22.383059Z","iopub.execute_input":"2023-03-21T22:41:22.384994Z","iopub.status.idle":"2023-03-21T22:41:29.155505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichPathway**","metadata":{}},{"cell_type":"code","source":"comp_path_ROrna <- compareCluster(geneCluster = prot_RNA, fun = enrichPathway, readable = TRUE)\n\ndf_path_ROrna <- comp_path_ROrna %>% as.data.frame\nrownames(df_path_ROrna) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RO_pos:\", nrow(df_path_ROrna[df_path_ROrna$Cluster == \"CD45RO_pos\",]), \"enriched terms found\",\n    \"\\nCD45RNA_pos:\", nrow(df_path_ROrna[df_path_ROrna$Cluster == \"CD45RNA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RO_neg:\", nrow(df_path_ROrna[df_path_ROrna$Cluster == \"CD45RO_neg\",]), \"enriched terms found\",\n    \"\\nCD45RNA_neg:\", nrow(df_path_ROrna[df_path_ROrna$Cluster == \"CD45RNA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_path_ROrna)\n)\n\ndf_path_ROrna$geneID <- str_wrap(df_path_ROrna$geneID, width=30)\ndf_path_ROrna %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:29.158555Z","iopub.execute_input":"2023-03-21T22:41:29.160227Z","iopub.status.idle":"2023-03-21T22:41:55.813062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\nn = 20\n\ndotplot(comp_path_ROrna, showCategory = n) + \n    scale_y_discrete(labels = function(x) str_wrap(x, width = 70)) + \n    theme_bw(base_size = 18)  + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45RO vs its RNA\\nEnriched terms shown (per cluster): \",n, \" of \", nrow(df_path_ROrna))) & \ntheme(text = element_text(size = 22) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:55.816181Z","iopub.execute_input":"2023-03-21T22:41:55.817420Z","iopub.status.idle":"2023-03-21T22:41:57.025860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,15)\nn = 30\n\ncnetplot(comp_path_ROrna, showCategory = n) + \n    theme_void(base_size = 22)  + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45RO vs its RNA\\nEnriched terms shown: \",n, \" of \", nrow(df_path_ROrna))) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:41:57.029216Z","iopub.execute_input":"2023-03-21T22:41:57.037945Z","iopub.status.idle":"2023-03-21T22:42:02.422622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">3. CD45RA vs its RNA</p>","metadata":{}},{"cell_type":"code","source":"cat(\"------Positively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RA and its RNA\")\ndata.frame(\n  \"Common for CD45RA and its RNA\" = paste(intersect(goRA_pos$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for CD45RA\" = paste(setdiff(goRA_pos$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for RNA\" = paste(setdiff(goRNA_pos$Description, goRA_pos$Description), collapse = \";\\n\"),\n  check.names = FALSE\n)\n\ncat(\"------Negatively correlated RNA------\\n\\n\")\n\ncat(\"Common and unique GO terms for CD45RA and its RNA\")\ndata.frame(\n  \"Common for CD45RA and its RNA\" = paste(intersect(goRA_neg$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for CD45RA\" = paste(setdiff(goRA_neg$Description, goRNA_pos$Description), collapse = \";\\n\"),\n  \"Unique for RNA\" = paste(setdiff(goRNA_pos$Description, goRA_neg$Description), collapse = \";\\n\"),\n  check.names = FALSE\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:42:02.425657Z","iopub.execute_input":"2023-03-21T22:42:02.427658Z","iopub.status.idle":"2023-03-21T22:42:02.505817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Comparing multiple gene lists**","metadata":{}},{"cell_type":"code","source":"prot_RNA <- list(\"CD45RA_pos\" = gene_setRA_pos$ENTREZID, \n                 \"CD45RNA_pos\" = gene_setRNA_pos$ENTREZID,\n                 \"CD45RA_neg\" = gene_setRA_neg$ENTREZID, \n                 \"CD45RNA_neg\" = gene_setRNA_neg$ENTREZID)\nstr(prot_RNA)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:42:02.508879Z","iopub.execute_input":"2023-03-21T22:42:02.510043Z","iopub.status.idle":"2023-03-21T22:42:02.528898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichGO**","metadata":{}},{"cell_type":"code","source":"comp_go_RArna <- compareCluster(geneCluster = prot_RNA, OrgDb = org.Hs.eg.db, fun = enrichGO, ont = \"BP\")\ncomp_go_RArna <- setReadable(comp_go_RArna, OrgDb = org.Hs.eg.db, keyType=\"ENTREZID\")\n\ndf_go_RArna <- comp_go_RArna %>% as.data.frame\nrownames(df_go_RArna) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RA_pos:\", nrow(df_go_RArna[df_go_RArna$Cluster == \"CD45RA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RNA_pos:\", nrow(df_go_RArna[df_go_RArna$Cluster == \"CD45RNA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RA_neg:\", nrow(df_go_RArna[df_go_RArna$Cluster == \"CD45RA_neg\",]), \"enriched terms found\",\n    \"\\nCD45RNA_neg:\", nrow(df_go_RArna[df_go_RArna$Cluster == \"CD45RNA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_go_RArna)\n)\n\ndf_go_RArna$geneID <- str_wrap(df_go_RArna$geneID, width=30)\ndf_go_RArna %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:42:02.531776Z","iopub.execute_input":"2023-03-21T22:42:02.533036Z","iopub.status.idle":"2023-03-21T22:42:53.803118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"comp_go_RArna <- simplify(comp_go_RArna, cutoff = 0.7, by = \"p.adjust\", select_fun = min)\ncat(\"The clusterProfiler simplify method reduced GO terms to:\")\ncomp_go_RArna %>% as.data.frame %>% nrow","metadata":{"execution":{"iopub.status.busy":"2023-03-21T22:42:53.806379Z","iopub.execute_input":"2023-03-21T22:42:53.807615Z","iopub.status.idle":"2023-03-21T22:43:08.793871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,30)\nn = 20\n\ndotplot(comp_go_RArna, showCategory=n) + \n    scale_y_discrete(labels = function(x) str_wrap(x, width = 70)) + \n    theme_bw(base_size = 18) + plot_annotation(\n    title = paste0(\"enrichGo\\nCD45RA vs its RNA\\nEnriched terms shown (per cluster): \",n, \" of \", comp_go_RArna %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 20) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:08.796902Z","iopub.execute_input":"2023-03-21T22:43:08.798117Z","iopub.status.idle":"2023-03-21T22:43:09.993800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,20)\nn = 30 #nrow(as.data.frame(comp_lst))\n\ncnetplot(comp_go_RArna, showCategory = n) + \n    theme_void(base_size = 22) + plot_annotation(\n    title = paste0(\"enrichGo\\nCD45RA vs its RNA\\nEnriched terms shown: \",n, \" of \", comp_go_RArna %>% as.data.frame %>% nrow)) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:09.997199Z","iopub.execute_input":"2023-03-21T22:43:09.999537Z","iopub.status.idle":"2023-03-21T22:43:19.411895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**enrichPathway**","metadata":{}},{"cell_type":"code","source":"comp_path_RArna <- compareCluster(geneCluster = prot_RNA, fun = enrichPathway, readable = TRUE)\n\ndf_path_RArna <- comp_path_RArna %>% as.data.frame\nrownames(df_path_RArna) <- NULL\n\ncat(\"GO enrichment analysis:\\n\")\ncat(\"\\nCD45RA_pos:\", nrow(df_path_RArna[df_path_RArna$Cluster == \"CD45RA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RNA_pos:\", nrow(df_path_RArna[df_path_RArna$Cluster == \"CD45RNA_pos\",]), \"enriched terms found\",\n    \"\\nCD45RA_neg:\", nrow(df_path_RArna[df_path_RArna$Cluster == \"CD45RA_neg\",]), \"enriched terms found\",\n    \"\\nCD45RNA_neg:\", nrow(df_path_RArna[df_path_RArna$Cluster == \"CD45RNA_neg\",]), \"enriched terms found\",\n   \"\\nTotal:\", nrow(df_path_RArna)\n)\n\ndf_path_RArna$geneID <- str_wrap(df_path_RArna$geneID, width=30)\ndf_path_RArna %>% head(10) %>% select(1,3,4,5,7,9)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:19.415037Z","iopub.execute_input":"2023-03-21T22:43:19.416676Z","iopub.status.idle":"2023-03-21T22:43:47.613869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,25)\nn = 20\n\ndotplot(comp_path_RArna, showCategory = n) + \n    scale_y_discrete(labels = function(x) str_wrap(x, width = 70)) + \n    theme_bw(base_size = 18) + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45RA vs its RNA\\nEnriched terms shown (per cluster): \",n, \" of \", nrow(df_path_RArna))) & \ntheme(text = element_text(size = 22) )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:47.617128Z","iopub.execute_input":"2023-03-21T22:43:47.618333Z","iopub.status.idle":"2023-03-21T22:43:48.555304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,25)\nn = 40 #nrow(df_path_RArna)\n\ncnetplot(comp_path_RArna, showCategory = n) + \n    theme_void(base_size = 22) + plot_annotation(\n    title = paste0(\"enrichPathway\\nCD45RA vs its RNA\\nEnriched terms shown: \",n, \" of \", nrow(df_path_RArna))) & \ntheme(text = element_text(size = 22) ) ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:48.558599Z","iopub.execute_input":"2023-03-21T22:43:48.560493Z","iopub.status.idle":"2023-03-21T22:43:56.181435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:Verdana; font-weight:normal; letter-spacing: 1px; color:#207d06; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid #207d06;\">4. Translation, transcription, splicing</p>","metadata":{}},{"cell_type":"markdown","source":"Here we find and visualise enriched terms related to translation, transcription, splicing","metadata":{}},{"cell_type":"code","source":"dat <- rbind(df_go_prot,\n             df_go_ROrna,\n             df_go_RArna,\n             df_path_prot,\n             df_path_ROrna,\n             df_path_ROrna)\ndat[grepl(\"splicing|trancrition|translation\", dat$Description, ignore.case = TRUE), ] %>% unique","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:56.184636Z","iopub.execute_input":"2023-03-21T22:43:56.186496Z","iopub.status.idle":"2023-03-21T22:43:56.256548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,13) \n\nplt <- filter(comp_go_RArna, grepl(\"splicing\", Description, ignore.case = TRUE))\n\nn = nrow(plt)\n                    \np1 <- cnetplot(plt, showCategory = n) + \n    theme_void(base_size = 22) + plot_annotation(\n    title = paste0(\"Genes, negatively correlated with CD45 RNA and involved in the enriched terms containing 'splicing'\\nEnriched terms shown: \",n)) & \ntheme(text = element_text(size = 22) )  \np1","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:56.259498Z","iopub.execute_input":"2023-03-21T22:43:56.260713Z","iopub.status.idle":"2023-03-21T22:43:58.326718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"genes <- paste0(df_path_ROrna[grepl(\"splicing\", df_path_ROrna$Description, ignore.case = TRUE), ]$geneID,\"/\",\n           df_go_ROrna[grepl(\"splicing\", df_go_ROrna$Description, ignore.case = TRUE), ]$geneID)\ngenes <- gsub(\"\\\\\\n\", \"\", genes)\ngenes <- str_split(genes, \"/\")\ngenes <- genes %>% unlist %>% unique\n\ncat(\"Genes:\\n\", \"'\", paste(genes, collapse = \"', '\"),\"'\", sep = \"\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:58.329720Z","iopub.execute_input":"2023-03-21T22:43:58.331638Z","iopub.status.idle":"2023-03-21T22:43:58.359273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig(25,13) \n\nplt <- filter(comp_go_RArna, grepl(\"trancrition|translation\", Description, ignore.case = TRUE))\n\nn = nrow(plt)\n                    \np1 <- cnetplot(plt, showCategory = n) + \n    theme_void(base_size = 22) + plot_annotation(\n    title = paste0(\"Genes, negatively correlated with CD45 RNA and involved in the enriched terms containing 'trancrition|translation'\\nEnriched terms shown: \",n)) & \ntheme(text = element_text(size = 22) )  \np1","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:58.362175Z","iopub.execute_input":"2023-03-21T22:43:58.363377Z","iopub.status.idle":"2023-03-21T22:43:59.943503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"genes <- paste0(df_path_ROrna[grepl(\"trancrition|translation\", df_path_ROrna$Description, ignore.case = TRUE), ]$geneID,\"/\",\n           df_go_ROrna[grepl(\"splicing\", df_go_ROrna$Description, ignore.case = TRUE), ]$geneID)\ngenes <- gsub(\"\\\\\\n\", \"\", genes)\ngenes <- str_split(genes, \"/\")\ngenes <- genes %>% unlist %>% unique\ncat(\"Genes:\\n\", \"'\", paste(genes, collapse = \"', '\"),\"'\", sep = \"\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:59.947072Z","iopub.execute_input":"2023-03-21T22:43:59.948760Z","iopub.status.idle":"2023-03-21T22:43:59.972771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#The corresponding information is directly extracted from the \"GO\" library. The result depends on the currently set ontology (\"BP\",\"MF\",\"CC\"), i.e. only GO terms within the actual ontology are considered. The shown GO information refers to the actually installed GO library.\n\n#getGOInfo(\"92906\") #HNRNPLL\n# genes <- c('SNRPD3','POLR2F','SNRPG','YBX1',\n#                  'SF3B5','POLR2L','SNU13','SNRPB','SRSF7',\n#                  'SRSF2','SNRPD1','UBL5','HSPA8','C1QBP')\n# genes <- AnnotationDbi::select(org.Hs.eg.db, keys = genes,\n#                                   columns = c('ENTREZID'), keytype = 'SYMBOL') #SYMBOL\n# getGOInfo(genes$ENTREZID)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:59.975864Z","iopub.execute_input":"2023-03-21T22:43:59.977221Z","iopub.status.idle":"2023-03-21T22:43:59.987546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-21T22:43:59.990585Z","iopub.execute_input":"2023-03-21T22:43:59.991886Z","iopub.status.idle":"2023-03-21T22:44:00.125036Z"},"trusted":true},"execution_count":null,"outputs":[]}]}