{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### What is about ?\n\nOne of the first things people look analysing the single cell RNA sequencing data (scRNA-seq) - number of non zeros genes for each cell. Here we make certain observations on that feature. \n\n**Findings Briefly:** \n\n    1) it decreases with days \n    2) it is correlated with proliferation activity (measured by standard genes signatures e.g. used in Scanpy/Seurat - Tirosh \n    genes).\n    3) considering just the sum over genes - gives the same effect - that is a bit unexpected since sum of exponents of them normalized to 1e6. That sum is correlated 0.99 to count of non-zero elements. \n    4) consideration of sums over chromosome - shows some increase of mitochondrial genes at days 3 and 7 - probabably more \"bad cells\". By Y chromosome we see that there was probably one woman donor. Sum over other chromosomes - are quite correlated. \n\n**Remark:** see also discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353629\n\n\n**Context:** For the current competition - scRNA-seq datasets play two roles: 1) one datasets consists of targets of \"Multiome\" task, 2) another gives features of \"CITEseq\" task.  But both datasets were produced from the same types of cells, but slighly different technologies. \n\nThe observed dependences are similar for all cases, but there are some strange effects for some donors and days. \n\n\n**Remark:** genes sets and even numbers of genes for Multiome and CITEseq task are slighly different, thus the numbers of non-zero genes slighly diverage. But neverthless there is no problem to compare  the daily trends and other dependences - which basically agree, modula some exceptions. It is probably worth to look more carefully on those exceptions. \n\n\n**Remark2:** See paper: https://arxiv.org/abs/2208.05229 for cell cycle analysis. \"Computational challenges of cell cycle analysis using single cell transcriptomics\" Alexander Chervov, Andrei Zinovyev. All the work for that paper has been done on Kaggle - see datasets and code therein: https://www.kaggle.com/alexandervc/datasets?scroll=true https://www.kaggle.com/andreizinovyev Hundreds of single cell RNA seq datasets were analyzed. See also discussion https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350314, and notebooks: https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03b-daybydaychange-allcelltypes, https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03b-daybydaychange-allcelltypes\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-multiome-targets-01-cell-cycle\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Install/import","metadata":{"execution":{"iopub.status.busy":"2022-09-17T09:37:16.789122Z","iopub.execute_input":"2022-09-17T09:37:16.791624Z","iopub.status.idle":"2022-09-17T09:37:16.800788Z","shell.execute_reply.started":"2022-09-17T09:37:16.791532Z","shell.execute_reply":"2022-09-17T09:37:16.799362Z"}}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-19T11:33:10.981430Z","iopub.execute_input":"2022-09-19T11:33:10.981850Z","iopub.status.idle":"2022-09-19T11:33:10.998477Z","shell.execute_reply.started":"2022-09-19T11:33:10.981814Z","shell.execute_reply":"2022-09-19T11:33:10.996330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install scanpy\nimport scanpy as sc\nimport anndata\n\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:33:12.150588Z","iopub.execute_input":"2022-09-19T11:33:12.151322Z","iopub.status.idle":"2022-09-19T11:33:52.912540Z","shell.execute_reply.started":"2022-09-19T11:33:12.151278Z","shell.execute_reply":"2022-09-19T11:33:52.911025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data","metadata":{}},{"cell_type":"code","source":"str_data_inf = ' scRNA-seq - Train Part of \"Multiome\" '","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:33:52.916750Z","iopub.execute_input":"2022-09-19T11:33:52.917243Z","iopub.status.idle":"2022-09-19T11:33:52.925029Z","shell.execute_reply.started":"2022-09-19T11:33:52.917191Z","shell.execute_reply":"2022-09-19T11:33:52.923520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/scrnaseq-competition-open-problems-adata-fmt/train_multi_targets_sparse.h5ad'\nadata = sc.read(fn)\nprint(adata)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:33:52.927560Z","iopub.execute_input":"2022-09-19T11:33:52.927945Z","iopub.status.idle":"2022-09-19T11:34:15.443772Z","shell.execute_reply.started":"2022-09-19T11:33:52.927912Z","shell.execute_reply":"2022-09-19T11:34:15.442469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('4.8G RAM consumed')","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:15.446665Z","iopub.execute_input":"2022-09-19T11:34:15.447028Z","iopub.status.idle":"2022-09-19T11:34:15.453297Z","shell.execute_reply.started":"2022-09-19T11:34:15.446996Z","shell.execute_reply":"2022-09-19T11:34:15.452113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#adata.var['Index Original'] = adata.var.index\n#adata.var.set_index('symbol', inplace = True) \n#adata.var \n\n# Causes error:\n# Cannot setitem on a Categorical with a new category, set the categories first\n# adata.var_names_make_unique()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:15.454982Z","iopub.execute_input":"2022-09-19T11:34:15.455417Z","iopub.status.idle":"2022-09-19T11:34:15.467835Z","shell.execute_reply.started":"2022-09-19T11:34:15.455382Z","shell.execute_reply":"2022-09-19T11:34:15.466803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Count of nonzero genes per cell shows clear daily dependent pattern. Might be due to decrease of proliferation","metadata":{}},{"cell_type":"code","source":"df_meta = pd.DataFrame()\ndf_meta['n_nonzeros'] = adata.X.indptr[1:] - adata.X.indptr[:-1] # Number of nonzero elements in sparse matrix row by row. Borrowed from SO: \ndf_meta['day']= adata.obs['day'].values\ndf_meta['donor']= adata.obs['donor'].values","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:15.469529Z","iopub.execute_input":"2022-09-19T11:34:15.470069Z","iopub.status.idle":"2022-09-19T11:34:15.504421Z","shell.execute_reply.started":"2022-09-19T11:34:15.470015Z","shell.execute_reply":"2022-09-19T11:34:15.503444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['n_nonzeros'].values, label = 'n_nonzeros' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*1000, label = 'Day')\nplt.plot( adata.obs['donor'].values/10, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('Count of nonzero genes per cell shows clear daily dependent pattern. Might be due to decrease of proliferation ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:15.506173Z","iopub.execute_input":"2022-09-19T11:34:15.506656Z","iopub.status.idle":"2022-09-19T11:34:17.799570Z","shell.execute_reply.started":"2022-09-19T11:34:15.506609Z","shell.execute_reply":"2022-09-19T11:34:17.798009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.groupby('day')['n_nonzeros'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:17.801166Z","iopub.execute_input":"2022-09-19T11:34:17.801520Z","iopub.status.idle":"2022-09-19T11:34:17.819220Z","shell.execute_reply.started":"2022-09-19T11:34:17.801488Z","shell.execute_reply":"2022-09-19T11:34:17.817987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.groupby(['donor', 'day'] )['n_nonzeros'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:17.821541Z","iopub.execute_input":"2022-09-19T11:34:17.822626Z","iopub.status.idle":"2022-09-19T11:34:17.846253Z","shell.execute_reply.started":"2022-09-19T11:34:17.822568Z","shell.execute_reply":"2022-09-19T11:34:17.844785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.groupby('day')['n_nonzeros'].median()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:17.851728Z","iopub.execute_input":"2022-09-19T11:34:17.852227Z","iopub.status.idle":"2022-09-19T11:34:17.872510Z","shell.execute_reply.started":"2022-09-19T11:34:17.852183Z","shell.execute_reply":"2022-09-19T11:34:17.870892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \"Correlation\" of proliferation activity and number of non zeros genes  (Multiome targets)","metadata":{}},{"cell_type":"code","source":"G1S_genes_Tirosh = ['ENSMUSG00000005410', 'ENSMUSG00000027342', 'ENSMUSG00000025747', 'ENSMUSG00000024742', 'ENSMUSG00000002870', 'ENSMUSG00000022673', 'ENSMUSG00000030978', 'ENSMUSG00000029591', 'ENSMUSG00000031821', 'ENSMUSG00000026355', 'ENSMUSG00000055612', 'ENSMUSG00000037474', 'ENSMUSG00000025395', 'ENSMUSG00000001228', 'ENSMUSG00000031629', 'ENSMUSG00000025001', 'ENSMUSG00000023104', 'ENSMUSG00000028884', 'ENSMUSG00000028693', 'ENSMUSG00000030346', 'ENSMUSG00000006715', 'ENSMUSG00000027242', 'ENSMUSG00000004642', 'ENSMUSG00000028212', 'ENSMUSG00000041712', 'ENSMUSG00000030726', 'ENSMUSG00000024151', 'ENSMUSG00000022360', 'ENSMUSG00000027323', 'ENSMUSG00000020649', 'ENSMUSG00000000028', 'ENSMUSG00000017499', 'ENSMUSG00000039748', 'ENSMUSG00000032397', 'ENSMUSG00000022422', 'ENSMUSG00000030528', 'ENSMUSG00000028282', 'ENSMUSG00000028560', 'ENSMUSG00000042489', 'ENSMUSG00000006678', 'ENSMUSG00000022945', 'ENSMUSG00000034329', 'ENSMUSG00000046179']\nG2M_genes_Tirosh = ['ENSMUSG00000054717', 'ENSMUSG00000019942', 'ENSMUSG00000027306', 'ENSMUSG00000001403', 'ENSMUSG00000017716', 'ENSMUSG00000027469', 'ENSMUSG00000020914', 'ENSMUSG00000024056', 'ENSMUSG00000062248', 'ENSMUSG00000026683', 'ENSMUSG00000028044', 'ENSMUSG00000031004', 'ENSMUSG00000019961', 'ENSMUSG00000026605', 'ENSMUSG00000037313', 'ENSMUSG00000020808', 'ENSMUSG00000034349', 'ENSMUSG00000032218', 'ENSMUSG00000048327', 'ENSMUSG00000037725', 'ENSMUSG00000020897', 'ENSMUSG00000027379', 'ENSMUSG00000012443', 'ENSMUSG00000015749', 'ENSMUSG00000036752', 'ENSMUSG00000022385', 'ENSMUSG00000024795', 'ENSMUSG00000044783', 'ENSMUSG00000023505', 'ENSMUSG00000020737', 'ENSMUSG00000006398', 'ENSMUSG00000038379', 'ENSMUSG00000044201', 'ENSMUSG00000028678', 'ENSMUSG00000022391', 'ENSMUSG00000038252', 'ENSMUSG00000037544', 'ENSMUSG00000048922', 'ENSMUSG00000028873', 'ENSMUSG00000027699', 'ENSMUSG00000032254', 'ENSMUSG00000020330', 'ENSMUSG00000027496', 'ENSMUSG00000068744', 'ENSMUSG00000036777', 'ENSMUSG00000004880', 'ENSMUSG00000040549', 'ENSMUSG00000045328', 'ENSMUSG00000005698', 'ENSMUSG00000026622', 'ENSMUSG00000035293', 'ENSMUSG00000074802', 'ENSMUSG00000009575', 'ENSMUSG00000029177']\nlist_genes_fastCCsign = ['ENSG00000170312', 'ENSG00000175063', 'ENSG00000131747', 'ENSG00000120802', 'ENSG00000123485', 'ENSG00000167325', 'ENSG00000111247', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000197299', 'ENSG00000136492', 'ENSG00000129173', 'ENSG00000184260']\n\nG1S_genes_Tirosh = ['MCM5', 'PCNA', 'TYMS', 'FEN1', 'MCM2', 'MCM4', 'RRM1', 'UNG', 'GINS2', 'MCM6', 'CDCA7', 'DTL', 'PRIM1', 'UHRF1', 'MLF1IP', 'HELLS', 'RFC2', 'RPA2', 'NASP', 'RAD51AP1', 'GMNN', 'WDR76', 'SLBP', 'CCNE2', 'UBR7', 'POLD3', 'MSH2', 'ATAD2', 'RAD51', 'RRM2', 'CDC45', 'CDC6', 'EXO1', 'TIPIN', 'DSCC1', 'BLM', 'CASP8AP2', 'USP1', 'CLSPN', 'POLA1', 'CHAF1B', 'BRIP1', 'E2F8']\nG2M_genes_Tirosh = ['HMGB2', 'CDK1', 'NUSAP1', 'UBE2C', 'BIRC5', 'TPX2', 'TOP2A', 'NDC80', 'CKS2', 'NUF2', 'CKS1B', 'MKI67', 'TMPO', 'CENPF', 'TACC3', 'FAM64A', 'SMC4', 'CCNB2', 'CKAP2L', 'CKAP2', 'AURKB', 'BUB1', 'KIF11', 'ANP32E', 'TUBB4B', 'GTSE1', 'KIF20B', 'HJURP', 'CDCA3', 'HN1', 'CDC20', 'TTK', 'CDC25C', 'KIF2C', 'RANGAP1', 'NCAPD2', 'DLGAP5', 'CDCA2', 'CDCA8', 'ECT2', 'KIF23', 'HMMR', 'AURKA', 'PSRC1', 'ANLN', 'LBR', 'CKAP5', 'CENPE', 'CTCF', 'NEK2', 'G2E3', 'GAS2L3', 'CBX5', 'CENPA']\nlist_genes_fastCCsign = ['CDK1', 'UBE2C', 'TOP2A', 'TMPO', 'HJURP', 'RRM1', 'RAD51AP1', 'RRM2', 'CDC45', 'BLM', 'BRIP1', 'E2F8', 'HIST2H2AC']\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:17.874973Z","iopub.execute_input":"2022-09-19T11:34:17.875539Z","iopub.status.idle":"2022-09-19T11:34:18.119391Z","shell.execute_reply.started":"2022-09-19T11:34:17.875486Z","shell.execute_reply":"2022-09-19T11:34:18.118137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = G1S_genes_Tirosh\nIX = np.where( adata.var['symbol'].isin(l) )[0]\nv1 = np.asarray(adata.X[:,IX].sum(axis=1) ).ravel()\nl = G2M_genes_Tirosh\nIX = np.where( adata.var['symbol'].isin(l) )[0]\nv2 = np.asarray(adata.X[:,IX].sum(axis=1) ).ravel()\nplt.figure(figsize= (20,12))\nax = sns.scatterplot(x=v1,y=v2, hue = adata.obs['day'], palette = 'rainbow')\nplt.setp(ax.get_legend().get_texts(), fontsize = 20) # for legend text\nplt.setp(ax.get_legend().get_title(), fontsize = 20) # for legend title  \nplt.title('Proliferation activity clearly decrease with days',fontsize = 20)\nplt.xlabel('G1S score (proliferation (cell cycle) activity)', fontsize = 20 )\nplt.ylabel('G2M score (proliferation (cell cycle) activity)', fontsize = 20)\nplt.show()\nplt.figure(figsize= (20,12))\nax = sns.scatterplot(x=v1,y=v2, hue = df_meta['n_nonzeros'], palette = 'rainbow')\nplt.setp(ax.get_legend().get_texts(), fontsize = 20) # for legend text\nplt.setp(ax.get_legend().get_title(), fontsize = 20) # for legend title  \nplt.title('Number of non zero genes agrees with G1S+G2M scores (proliferation activity) ',fontsize = 20)\nplt.xlabel('G1S score (proliferation (cell cycle) activity)', fontsize = 20 )\nplt.ylabel('G2M score (proliferation (cell cycle) activity)', fontsize = 20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:18.120860Z","iopub.execute_input":"2022-09-19T11:34:18.121262Z","iopub.status.idle":"2022-09-19T11:34:26.978905Z","shell.execute_reply.started":"2022-09-19T11:34:18.121226Z","shell.execute_reply":"2022-09-19T11:34:26.977252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta['G1S score'] = v1\ndf_meta['G2M score'] = v2\ndf_meta['G1S+G2M cell cycle score'] = v1 + v2\n\ndf_meta.corr()\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:26.981683Z","iopub.execute_input":"2022-09-19T11:34:26.982554Z","iopub.status.idle":"2022-09-19T11:34:27.030626Z","shell.execute_reply.started":"2022-09-19T11:34:26.982504Z","shell.execute_reply":"2022-09-19T11:34:27.029636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plt.plot(df_meta['G1S+G2M cell cycle score'])\nplt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['G1S+G2M cell cycle score'].values, label = 'G1S+G2M score' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*100, label = 'Day')\nplt.plot( adata.obs['donor'].values/100, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('G1S+G2M score (proliferation score) shows similar pattern to  count of nonzero genes per cell shows clear daily dependent pattern ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:27.031865Z","iopub.execute_input":"2022-09-19T11:34:27.032971Z","iopub.status.idle":"2022-09-19T11:34:28.079331Z","shell.execute_reply.started":"2022-09-19T11:34:27.032924Z","shell.execute_reply":"2022-09-19T11:34:28.077946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta['cell_id'] =  adata.obs.index\ndf_meta.head(1)","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:28.080822Z","iopub.execute_input":"2022-09-19T11:34:28.081294Z","iopub.status.idle":"2022-09-19T11:34:28.101436Z","shell.execute_reply.started":"2022-09-19T11:34:28.081254Z","shell.execute_reply":"2022-09-19T11:34:28.100118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = list_genes_fastCCsign #  l =  G1S_genes_Tirosh # \nIX = np.where( adata.var['symbol'].isin(l) )[0]\nv3 = np.asarray(adata.X[:,IX].sum(axis=1) ).ravel()\nplt.plot(v3)\nv3[:20]","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:28.103051Z","iopub.execute_input":"2022-09-19T11:34:28.103615Z","iopub.status.idle":"2022-09-19T11:34:29.715776Z","shell.execute_reply.started":"2022-09-19T11:34:28.103562Z","shell.execute_reply":"2022-09-19T11:34:29.714193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta['FAST score'] = v3# adata.obs.index","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:29.718213Z","iopub.execute_input":"2022-09-19T11:34:29.718768Z","iopub.status.idle":"2022-09-19T11:34:29.726354Z","shell.execute_reply.started":"2022-09-19T11:34:29.718716Z","shell.execute_reply":"2022-09-19T11:34:29.724947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.asarray(adata.X.sum(axis = 1)).ravel()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:29.728324Z","iopub.execute_input":"2022-09-19T11:34:29.729160Z","iopub.status.idle":"2022-09-19T11:34:29.991699Z","shell.execute_reply.started":"2022-09-19T11:34:29.729106Z","shell.execute_reply":"2022-09-19T11:34:29.990085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sum of  elements shows similar behaviour to count non-zeros. And is 0.99 correlated. That a bit unexpected since sum of exponenets of these values are normalized to the same constant 1e6","metadata":{}},{"cell_type":"code","source":"df_meta['sum GEX'] = np.asarray(adata.X.sum(axis = 1)).ravel()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:29.993700Z","iopub.execute_input":"2022-09-19T11:34:29.994587Z","iopub.status.idle":"2022-09-19T11:34:30.204241Z","shell.execute_reply.started":"2022-09-19T11:34:29.994533Z","shell.execute_reply":"2022-09-19T11:34:30.202750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plt.plot(df_meta['G1S+G2M cell cycle score'])\nplt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['sum GEX'].values, label = 'sum GEX' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*1000, label = 'Day')\nplt.plot( adata.obs['donor'].values/10, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('Sum of GEX values shows similar pattern to  count of nonzero genes per cell shows clear daily dependent pattern ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:30.206367Z","iopub.execute_input":"2022-09-19T11:34:30.206923Z","iopub.status.idle":"2022-09-19T11:34:32.305662Z","shell.execute_reply.started":"2022-09-19T11:34:30.206869Z","shell.execute_reply":"2022-09-19T11:34:32.304027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.corr()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:32.307592Z","iopub.execute_input":"2022-09-19T11:34:32.307983Z","iopub.status.idle":"2022-09-19T11:34:32.354127Z","shell.execute_reply.started":"2022-09-19T11:34:32.307938Z","shell.execute_reply":"2022-09-19T11:34:32.352205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis sums with respect along to chromosomes. Basically all such somes are quite correlated except very small \"Y\" and \"MT\". (Expression of mitochrondrial genes - typical quality control - high expression cell is near dead). ","metadata":{}},{"cell_type":"code","source":"adata.var['chr'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:34:32.355599Z","iopub.execute_input":"2022-09-19T11:34:32.355978Z","iopub.status.idle":"2022-09-19T11:34:32.370350Z","shell.execute_reply.started":"2022-09-19T11:34:32.355945Z","shell.execute_reply":"2022-09-19T11:34:32.368597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = adata.var['chr'].value_counts()\nfor c in s.index:\n    if s[c]> 10:\n        print(c, s[c])\n        m = (adata.var['chr'] == c)\n        t = np.asarray( adata[:,m].X.sum(axis = 1) ) .ravel() \n        df_meta['Sum chr'+c] = t\ndf_meta.head(2)        ","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:40:25.469876Z","iopub.execute_input":"2022-09-19T11:40:25.470432Z","iopub.status.idle":"2022-09-19T11:40:57.345237Z","shell.execute_reply.started":"2022-09-19T11:40:25.470394Z","shell.execute_reply":"2022-09-19T11:40:57.344139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.corr()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:44:21.165815Z","iopub.execute_input":"2022-09-19T11:44:21.166238Z","iopub.status.idle":"2022-09-19T11:44:21.573426Z","shell.execute_reply.started":"2022-09-19T11:44:21.166202Z","shell.execute_reply":"2022-09-19T11:44:21.572166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_meta.columns","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:49:19.656328Z","iopub.execute_input":"2022-09-19T11:49:19.657360Z","iopub.status.idle":"2022-09-19T11:49:19.666908Z","shell.execute_reply.started":"2022-09-19T11:49:19.657302Z","shell.execute_reply":"2022-09-19T11:49:19.665379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plt.plot(df_meta['G1S+G2M cell cycle score'])\nplt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['Sum chrMT'].values, label = 'sum GEX - 13 mitochondrial genes' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*10, label = 'Day')\nplt.plot( adata.obs['donor'].values/1000, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('Sum of 13 mitochodrial genes - shows increase at days 3 and 7  ( \"bad cells ?? \")  ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:56:37.299829Z","iopub.execute_input":"2022-09-19T11:56:37.301149Z","iopub.status.idle":"2022-09-19T11:56:39.106270Z","shell.execute_reply.started":"2022-09-19T11:56:37.301092Z","shell.execute_reply":"2022-09-19T11:56:39.104910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plt.plot(df_meta['G1S+G2M cell cycle score'])\nplt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['Sum chrY'].values, label = 'sum GEX chromosome Y' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*10, label = 'Day')\nplt.plot( adata.obs['donor'].values/1000, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('Sum of Y chromosome - one donor woman - no Y chromosome.  Some decrease at day 7 for other donors ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T11:55:57.232233Z","iopub.execute_input":"2022-09-19T11:55:57.232809Z","iopub.status.idle":"2022-09-19T11:55:59.250826Z","shell.execute_reply.started":"2022-09-19T11:55:57.232757Z","shell.execute_reply":"2022-09-19T11:55:59.249251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plt.plot(df_meta['G1S+G2M cell cycle score'])\nplt.figure(figsize = (20,5))\n#df_meta['n_nonzeros'].plot()\nplt.plot( df_meta['Sum chr4'].values, label = 'sum GEX - chromosome4' )\nv = adata.obs['day']\nplt.plot(v.values.ravel()*100, label = 'Day')\nplt.plot( adata.obs['donor'].values/100, label = 'donor')\nplt.legend(fontsize = 14)\nplt.suptitle(str_data_inf , fontsize = 20)\nplt.title('Sum of GEX chromosome 4 - the same decrease at day 7 as for overall sum ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T12:07:16.720076Z","iopub.execute_input":"2022-09-19T12:07:16.720569Z","iopub.status.idle":"2022-09-19T12:07:18.972540Z","shell.execute_reply.started":"2022-09-19T12:07:16.720531Z","shell.execute_reply":"2022-09-19T12:07:18.971074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp = df_meta.copy()\ntmp.columns = ['n_nonzeros GEX']+   list(df_meta.columns[1:])\ntmp.head(1)","metadata":{"execution":{"iopub.status.busy":"2022-09-19T12:07:33.475563Z","iopub.execute_input":"2022-09-19T12:07:33.476005Z","iopub.status.idle":"2022-09-19T12:07:33.519733Z","shell.execute_reply.started":"2022-09-19T12:07:33.475962Z","shell.execute_reply":"2022-09-19T12:07:33.517903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp.to_csv('multiome_meta_and_scores_etc.csv')","metadata":{"execution":{"iopub.status.busy":"2022-09-19T12:07:36.317460Z","iopub.execute_input":"2022-09-19T12:07:36.318212Z","iopub.status.idle":"2022-09-19T12:07:39.535114Z","shell.execute_reply.started":"2022-09-19T12:07:36.318175Z","shell.execute_reply":"2022-09-19T12:07:39.533688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del adata","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:12:38.388183Z","iopub.execute_input":"2022-09-19T08:12:38.389156Z","iopub.status.idle":"2022-09-19T08:12:38.394175Z","shell.execute_reply.started":"2022-09-19T08:12:38.389116Z","shell.execute_reply":"2022-09-19T08:12:38.392997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:12:39.290679Z","iopub.execute_input":"2022-09-19T08:12:39.291120Z","iopub.status.idle":"2022-09-19T08:12:40.061207Z","shell.execute_reply.started":"2022-09-19T08:12:39.291088Z","shell.execute_reply":"2022-09-19T08:12:40.059849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Look on CITE-seq part of data","metadata":{}},{"cell_type":"code","source":"str_data_inf = 'scRNA-seq Train part of \"CITE-seq\" '","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:33:21.507688Z","iopub.execute_input":"2022-09-19T08:33:21.508272Z","iopub.status.idle":"2022-09-19T08:33:21.514277Z","shell.execute_reply.started":"2022-09-19T08:33:21.508234Z","shell.execute_reply":"2022-09-19T08:33:21.513121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:14:17.704935Z","iopub.execute_input":"2022-09-19T08:14:17.705318Z","iopub.status.idle":"2022-09-19T08:14:17.713611Z","shell.execute_reply.started":"2022-09-19T08:14:17.705288Z","shell.execute_reply":"2022-09-19T08:14:17.712288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:14:35.248317Z","iopub.execute_input":"2022-09-19T08:14:35.248812Z","iopub.status.idle":"2022-09-19T08:14:35.695014Z","shell.execute_reply.started":"2022-09-19T08:14:35.248776Z","shell.execute_reply":"2022-09-19T08:14:35.692836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\n#df_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\ndisplay(df_cite_train_x.head())","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:15:25.292594Z","iopub.execute_input":"2022-09-19T08:15:25.293063Z","iopub.status.idle":"2022-09-19T08:16:43.639017Z","shell.execute_reply.started":"2022-09-19T08:15:25.293024Z","shell.execute_reply":"2022-09-19T08:16:43.637812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('9.2G RAM consumed, during loading max was 14.2 RAM - near to 16G limit ')","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:17:23.002190Z","iopub.execute_input":"2022-09-19T08:17:23.002604Z","iopub.status.idle":"2022-09-19T08:17:23.009066Z","shell.execute_reply.started":"2022-09-19T08:17:23.002571Z","shell.execute_reply":"2022-09-19T08:17:23.007634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nv = np.zeros(len(df_cite_train_x))\nfor i in range(len(df_cite_train_x)):\n    v[i] = df_cite_train_x.iloc[i,:].astype(bool).sum()\n    ","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:19:22.089968Z","iopub.execute_input":"2022-09-19T08:19:22.090367Z","iopub.status.idle":"2022-09-19T08:20:31.605479Z","shell.execute_reply.started":"2022-09-19T08:19:22.090336Z","shell.execute_reply":"2022-09-19T08:20:31.604313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite_train_x","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:24:47.003393Z","iopub.execute_input":"2022-09-19T08:24:47.003866Z","iopub.status.idle":"2022-09-19T08:24:47.066297Z","shell.execute_reply.started":"2022-09-19T08:24:47.003828Z","shell.execute_reply":"2022-09-19T08:24:47.065104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmp = pd.DataFrame(index = df_cite_train_x.index, data = v, columns = ['n_nonzeros'])\ndf_2_meta = tmp.join(df_cell.set_index('cell_id'), how = 'left',  ) #, on = 'cell_id'\ndf_2_meta","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:27:41.697124Z","iopub.execute_input":"2022-09-19T08:27:41.697556Z","iopub.status.idle":"2022-09-19T08:27:41.804681Z","shell.execute_reply.started":"2022-09-19T08:27:41.697521Z","shell.execute_reply":"2022-09-19T08:27:41.803428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Count non-zero elements for CITE-seq data has similar but less pronounced pattern - decrease with days and it is more donor dependent","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\nplt.plot(df_2_meta['n_nonzeros'].values, label = 'n_nonzeros' )# .plot()\nv = df_2_meta['day'].values\nplt.plot(v.ravel()*1000, label = 'Day')\n\nplt.plot(df_2_meta['donor'].values/10, label = 'donor')\n\nplt.legend()\nplt.suptitle(str_data_inf, fontsize = 20)\nplt.title('Count of nonzero genes per cell shows daily dependent pattern. Might be due to decrease of proliferation ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:36:10.563590Z","iopub.execute_input":"2022-09-19T08:36:10.564984Z","iopub.status.idle":"2022-09-19T08:36:12.502007Z","shell.execute_reply.started":"2022-09-19T08:36:10.564938Z","shell.execute_reply":"2022-09-19T08:36:12.501098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Where is RAM ?","metadata":{}},{"cell_type":"code","source":"del df_cite_train_x\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:42:46.239340Z","iopub.execute_input":"2022-09-19T08:42:46.240089Z","iopub.status.idle":"2022-09-19T08:42:46.245744Z","shell.execute_reply.started":"2022-09-19T08:42:46.240047Z","shell.execute_reply":"2022-09-19T08:42:46.244359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('9.2G consumed - but all big objects were deleted - where is the RAM ? ')","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:43:24.360590Z","iopub.execute_input":"2022-09-19T08:43:24.361041Z","iopub.status.idle":"2022-09-19T08:43:24.366817Z","shell.execute_reply.started":"2022-09-19T08:43:24.361007Z","shell.execute_reply":"2022-09-19T08:43:24.365831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CITE seq - TEST part - need to restart notebook due to RAM crash","metadata":{}},{"cell_type":"code","source":"# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:48:42.546488Z","iopub.execute_input":"2022-09-19T08:48:42.546931Z","iopub.status.idle":"2022-09-19T08:49:08.864541Z","shell.execute_reply.started":"2022-09-19T08:48:42.546893Z","shell.execute_reply":"2022-09-19T08:49:08.863259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:49:08.867789Z","iopub.execute_input":"2022-09-19T08:49:08.868290Z","iopub.status.idle":"2022-09-19T08:49:08.877866Z","shell.execute_reply.started":"2022-09-19T08:49:08.868252Z","shell.execute_reply":"2022-09-19T08:49:08.876154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cell = pd.read_csv(FP_CELL_METADATA)\ndf_cell","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:50:53.033634Z","iopub.execute_input":"2022-09-19T08:50:53.034078Z","iopub.status.idle":"2022-09-19T08:50:53.443248Z","shell.execute_reply.started":"2022-09-19T08:50:53.034026Z","shell.execute_reply":"2022-09-19T08:50:53.442113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"str_data_inf = 'scRNA-seq Test part of \"CITE-seq\" '","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:50:53.616471Z","iopub.execute_input":"2022-09-19T08:50:53.617562Z","iopub.status.idle":"2022-09-19T08:50:53.622789Z","shell.execute_reply.started":"2022-09-19T08:50:53.617505Z","shell.execute_reply":"2022-09-19T08:50:53.621726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#df_cite_train_x = pd.read_hdf(FP_CITE_TRAIN_INPUTS)\ndf_cite_test_x = pd.read_hdf(FP_CITE_TEST_INPUTS)\ndisplay(df_cite_test_x.head())","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:49:08.892400Z","iopub.execute_input":"2022-09-19T08:49:08.892768Z","iopub.status.idle":"2022-09-19T08:49:46.281347Z","shell.execute_reply.started":"2022-09-19T08:49:08.892738Z","shell.execute_reply":"2022-09-19T08:49:46.280217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nv = np.zeros(len(df_cite_test_x))\nfor i in range(len(df_cite_test_x)):\n    v[i] = df_cite_test_x.iloc[i,:].astype(bool).sum()\n    \n\ntmp = pd.DataFrame(index = df_cite_test_x.index, data = v, columns = ['n_nonzeros'])\ndf_2_meta = tmp.join(df_cell.set_index('cell_id'), how = 'left',  ) #, on = 'cell_id'\ndisplay(df_2_meta)","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:51:07.222349Z","iopub.execute_input":"2022-09-19T08:51:07.222763Z","iopub.status.idle":"2022-09-19T08:51:33.163026Z","shell.execute_reply.started":"2022-09-19T08:51:07.222731Z","shell.execute_reply":"2022-09-19T08:51:33.161977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\nplt.plot(df_2_meta['n_nonzeros'].values, label = 'n_nonzeros' )# .plot()\nv = df_2_meta['day'].values\nplt.plot(v.ravel()*1000, label = 'Day')\n\nplt.plot(df_2_meta['donor'].values/10, label = 'donor')\n\nplt.legend()\nplt.suptitle(str_data_inf, fontsize = 20)\nplt.title('Count of nonzero genes daily decrease for one of donors. But there is also donor dependence for day 7 ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:55:33.987943Z","iopub.execute_input":"2022-09-19T08:55:33.990101Z","iopub.status.idle":"2022-09-19T08:55:34.860283Z","shell.execute_reply.started":"2022-09-19T08:55:33.990022Z","shell.execute_reply":"2022-09-19T08:55:34.859293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_2_meta.groupby(['day','donor'])['n_nonzeros'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T08:59:21.142780Z","iopub.execute_input":"2022-09-19T08:59:21.143240Z","iopub.status.idle":"2022-09-19T08:59:21.163549Z","shell.execute_reply.started":"2022-09-19T08:59:21.143198Z","shell.execute_reply":"2022-09-19T08:59:21.162695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \"Correlation\" of proliferation activity and number of non zeros genes  (CITE-seq test features)","metadata":{}},{"cell_type":"code","source":"\nG1S_genes_Tirosh = ['MCM5', 'PCNA', 'TYMS', 'FEN1', 'MCM2', 'MCM4', 'RRM1', 'UNG', 'GINS2', 'MCM6', 'CDCA7', 'DTL', 'PRIM1', 'UHRF1', 'MLF1IP', 'HELLS', 'RFC2', 'RPA2', 'NASP', 'RAD51AP1', 'GMNN', 'WDR76', 'SLBP', 'CCNE2', 'UBR7', 'POLD3', 'MSH2', 'ATAD2', 'RAD51', 'RRM2', 'CDC45', 'CDC6', 'EXO1', 'TIPIN', 'DSCC1', 'BLM', 'CASP8AP2', 'USP1', 'CLSPN', 'POLA1', 'CHAF1B', 'BRIP1', 'E2F8']\nG2M_genes_Tirosh = ['HMGB2', 'CDK1', 'NUSAP1', 'UBE2C', 'BIRC5', 'TPX2', 'TOP2A', 'NDC80', 'CKS2', 'NUF2', 'CKS1B', 'MKI67', 'TMPO', 'CENPF', 'TACC3', 'FAM64A', 'SMC4', 'CCNB2', 'CKAP2L', 'CKAP2', 'AURKB', 'BUB1', 'KIF11', 'ANP32E', 'TUBB4B', 'GTSE1', 'KIF20B', 'HJURP', 'CDCA3', 'HN1', 'CDC20', 'TTK', 'CDC25C', 'KIF2C', 'RANGAP1', 'NCAPD2', 'DLGAP5', 'CDCA2', 'CDCA8', 'ECT2', 'KIF23', 'HMMR', 'AURKA', 'PSRC1', 'ANLN', 'LBR', 'CKAP5', 'CENPE', 'CTCF', 'NEK2', 'G2E3', 'GAS2L3', 'CBX5', 'CENPA']\nlist_genes_fastCCsign = ['CDK1', 'UBE2C', 'TOP2A', 'TMPO', 'HJURP', 'RRM1', 'RAD51AP1', 'RRM2', 'CDC45', 'BLM', 'BRIP1', 'E2F8', 'HIST2H2AC']\n\nG1S_genes_Tirosh = ['ENSMUSG00000005410', 'ENSMUSG00000027342', 'ENSMUSG00000025747', 'ENSMUSG00000024742', 'ENSMUSG00000002870', 'ENSMUSG00000022673', 'ENSMUSG00000030978', 'ENSMUSG00000029591', 'ENSMUSG00000031821', 'ENSMUSG00000026355', 'ENSMUSG00000055612', 'ENSMUSG00000037474', 'ENSMUSG00000025395', 'ENSMUSG00000001228', 'ENSMUSG00000031629', 'ENSMUSG00000025001', 'ENSMUSG00000023104', 'ENSMUSG00000028884', 'ENSMUSG00000028693', 'ENSMUSG00000030346', 'ENSMUSG00000006715', 'ENSMUSG00000027242', 'ENSMUSG00000004642', 'ENSMUSG00000028212', 'ENSMUSG00000041712', 'ENSMUSG00000030726', 'ENSMUSG00000024151', 'ENSMUSG00000022360', 'ENSMUSG00000027323', 'ENSMUSG00000020649', 'ENSMUSG00000000028', 'ENSMUSG00000017499', 'ENSMUSG00000039748', 'ENSMUSG00000032397', 'ENSMUSG00000022422', 'ENSMUSG00000030528', 'ENSMUSG00000028282', 'ENSMUSG00000028560', 'ENSMUSG00000042489', 'ENSMUSG00000006678', 'ENSMUSG00000022945', 'ENSMUSG00000034329', 'ENSMUSG00000046179']\nG2M_genes_Tirosh = ['ENSMUSG00000054717', 'ENSMUSG00000019942', 'ENSMUSG00000027306', 'ENSMUSG00000001403', 'ENSMUSG00000017716', 'ENSMUSG00000027469', 'ENSMUSG00000020914', 'ENSMUSG00000024056', 'ENSMUSG00000062248', 'ENSMUSG00000026683', 'ENSMUSG00000028044', 'ENSMUSG00000031004', 'ENSMUSG00000019961', 'ENSMUSG00000026605', 'ENSMUSG00000037313', 'ENSMUSG00000020808', 'ENSMUSG00000034349', 'ENSMUSG00000032218', 'ENSMUSG00000048327', 'ENSMUSG00000037725', 'ENSMUSG00000020897', 'ENSMUSG00000027379', 'ENSMUSG00000012443', 'ENSMUSG00000015749', 'ENSMUSG00000036752', 'ENSMUSG00000022385', 'ENSMUSG00000024795', 'ENSMUSG00000044783', 'ENSMUSG00000023505', 'ENSMUSG00000020737', 'ENSMUSG00000006398', 'ENSMUSG00000038379', 'ENSMUSG00000044201', 'ENSMUSG00000028678', 'ENSMUSG00000022391', 'ENSMUSG00000038252', 'ENSMUSG00000037544', 'ENSMUSG00000048922', 'ENSMUSG00000028873', 'ENSMUSG00000027699', 'ENSMUSG00000032254', 'ENSMUSG00000020330', 'ENSMUSG00000027496', 'ENSMUSG00000068744', 'ENSMUSG00000036777', 'ENSMUSG00000004880', 'ENSMUSG00000040549', 'ENSMUSG00000045328', 'ENSMUSG00000005698', 'ENSMUSG00000026622', 'ENSMUSG00000035293', 'ENSMUSG00000074802', 'ENSMUSG00000009575', 'ENSMUSG00000029177']\nlist_genes_fastCCsign = ['ENSG00000170312', 'ENSG00000175063', 'ENSG00000131747', 'ENSG00000120802', 'ENSG00000123485', 'ENSG00000167325', 'ENSG00000111247', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000197299', 'ENSG00000136492', 'ENSG00000129173', 'ENSG00000184260']\n\nG1S_genes_Tirosh = ['ENSG00000100297', 'ENSG00000132646', 'ENSG00000176890', 'ENSG00000168496', 'ENSG00000073111', 'ENSG00000104738', 'ENSG00000167325', 'ENSG00000076248', 'ENSG00000131153', 'ENSG00000076003', 'ENSG00000144354', 'ENSG00000143476', 'ENSG00000198056', 'ENSG00000276043', 'ENSG00000151725', 'ENSG00000119969', 'ENSG00000049541', 'ENSG00000117748', 'ENSG00000132780', 'ENSG00000111247', 'ENSG00000112312', 'ENSG00000092470', 'ENSG00000163950', 'ENSG00000175305', 'ENSG00000012963', 'ENSG00000077514', 'ENSG00000095002', 'ENSG00000156802', 'ENSG00000051180', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000094804', 'ENSG00000174371', 'ENSG00000075131', 'ENSG00000136982', 'ENSG00000197299', 'ENSG00000118412', 'ENSG00000162607', 'ENSG00000092853', 'ENSG00000101868', 'ENSG00000159259', 'ENSG00000136492', 'ENSG00000129173']\nG2M_genes_Tirosh = ['ENSG00000164104', 'ENSG00000170312', 'ENSG00000137804', 'ENSG00000175063', 'ENSG00000089685', 'ENSG00000088325', 'ENSG00000131747', 'ENSG00000080986', 'ENSG00000123975', 'ENSG00000143228', 'ENSG00000173207', 'ENSG00000148773', 'ENSG00000120802', 'ENSG00000117724', 'ENSG00000013810', 'ENSG00000129195', 'ENSG00000113810', 'ENSG00000157456', 'ENSG00000169607', 'ENSG00000136108', 'ENSG00000178999', 'ENSG00000169679', 'ENSG00000138160', 'ENSG00000143401', 'ENSG00000188229', 'ENSG00000075218', 'ENSG00000138182', 'ENSG00000123485', 'ENSG00000111665', 'ENSG00000189159', 'ENSG00000117399', 'ENSG00000112742', 'ENSG00000158402', 'ENSG00000142945', 'ENSG00000100401', 'ENSG00000010292', 'ENSG00000126787', 'ENSG00000184661', 'ENSG00000134690', 'ENSG00000114346', 'ENSG00000137807', 'ENSG00000072571', 'ENSG00000087586', 'ENSG00000134222', 'ENSG00000011426', 'ENSG00000143815', 'ENSG00000175216', 'ENSG00000138778', 'ENSG00000102974', 'ENSG00000117650', 'ENSG00000092140', 'ENSG00000139354', 'ENSG00000094916', 'ENSG00000115163']\nlist_genes_fastCCsign = ['ENSG00000170312', 'ENSG00000175063', 'ENSG00000131747', 'ENSG00000120802', 'ENSG00000123485', 'ENSG00000167325', 'ENSG00000111247', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000197299', 'ENSG00000136492', 'ENSG00000129173', 'ENSG00000184260']\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T09:48:13.258355Z","iopub.execute_input":"2022-09-19T09:48:13.258835Z","iopub.status.idle":"2022-09-19T09:48:13.283044Z","shell.execute_reply.started":"2022-09-19T09:48:13.258800Z","shell.execute_reply":"2022-09-19T09:48:13.282223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_genes = [ t.split('_')[0] for t in df_cite_test_x.columns]\nprint(len(list_genes), df_cite_test_x.shape, list_genes[:10])","metadata":{"execution":{"iopub.status.busy":"2022-09-19T09:48:15.410556Z","iopub.execute_input":"2022-09-19T09:48:15.411650Z","iopub.status.idle":"2022-09-19T09:48:15.428171Z","shell.execute_reply.started":"2022-09-19T09:48:15.411602Z","shell.execute_reply":"2022-09-19T09:48:15.426750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = G1S_genes_Tirosh\nIX = np.where( pd.Series(list_genes).isin(l) )[0]\nv1 = np.asarray(df_cite_test_x.iloc[:,IX].sum(axis=1) ).ravel()\nl = G2M_genes_Tirosh\nIX = np.where( pd.Series(list_genes).isin(l) )[0]\nv2 = np.asarray(df_cite_test_x.iloc[:,IX].sum(axis=1) ).ravel()\nplt.figure(figsize= (20,12))\nax = sns.scatterplot(x=v1,y=v2, hue = df_2_meta['n_nonzeros'], palette = 'rainbow')\nplt.setp(ax.get_legend().get_texts(), fontsize = 20) # for legend text\nplt.setp(ax.get_legend().get_title(), fontsize = 20) # for legend title  \nplt.title( 'Number of non zero genes agrees with G1S+G2M scores (proliferation activity)',fontsize = 20)\nplt.xlabel('G1S score (proliferation (cell cycle) activity)', fontsize = 20 )\nplt.ylabel('G2M score (proliferation (cell cycle) activity)', fontsize = 20)\nplt.show()\n\nplt.figure(figsize= (20,12))\nax = sns.scatterplot(x=v1,y=v2, hue = df_2_meta['day'], palette = 'rainbow', alpha = 0.1)\nplt.setp(ax.get_legend().get_texts(), fontsize = 20) # for legend text\nplt.setp(ax.get_legend().get_title(), fontsize = 20) # for legend title  \nplt.title('Day 7 - less proliferation ',fontsize = 20)\nplt.xlabel('G1S score (proliferation (cell cycle) activity)', fontsize = 20 )\nplt.ylabel('G2M score (proliferation (cell cycle) activity)', fontsize = 20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T10:15:44.274820Z","iopub.execute_input":"2022-09-19T10:15:44.275345Z","iopub.status.idle":"2022-09-19T10:15:48.087735Z","shell.execute_reply.started":"2022-09-19T10:15:44.275302Z","shell.execute_reply":"2022-09-19T10:15:48.086332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_2_meta['G1S score'] = v1\ndf_2_meta['G2M score'] = v2\ndf_2_meta['G1S+G2M score'] = v1+v2\n\ndf_2_meta.corr()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T09:50:58.901387Z","iopub.execute_input":"2022-09-19T09:50:58.901815Z","iopub.status.idle":"2022-09-19T09:50:58.929777Z","shell.execute_reply.started":"2022-09-19T09:50:58.901782Z","shell.execute_reply":"2022-09-19T09:50:58.928639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\nplt.plot(df_2_meta['G1S+G2M score'].values, label = 'G1S+G2M score (proliferation)' )# .plot()\nv = df_2_meta['day'].values\nplt.plot(v.ravel()*100, label = 'Day')\n\nplt.plot(df_2_meta['donor'].values/100, label = 'donor')\n\nplt.legend()\nplt.suptitle(str_data_inf, fontsize = 20)\nplt.title('Proliferation score falls down at day 3, more or less stable other days and day 7 for other donors ')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T10:45:44.087992Z","iopub.execute_input":"2022-09-19T10:45:44.088531Z","iopub.status.idle":"2022-09-19T10:45:45.569778Z","shell.execute_reply.started":"2022-09-19T10:45:44.088493Z","shell.execute_reply":"2022-09-19T10:45:45.568140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{},"execution_count":null,"outputs":[]}]}