{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ? \n\n**Briefly:** Example how to work with huge .h5 files without loading them in RAM. \nSee also example in https://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk \n\n**Data** from the competition , but denoising preprocessing by MAGIC package (https://github.com/KrishnaswamyLab/MAGIC) has been done by Liza Geraseva. See discussion: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\n\n**Work principle:** The main data can be accessed as f['0']['block0_values'] - without loading to RAM, but  when one takes a slice: f['0']['block0_values'][:,IX] - it is converted to numpy array (and loaded into memory?). \nThere are some details: one can do constant slices like  f['0']['block0_values'][:10,:10], but to make a  variable slice one need to process in two steps: f['0']['block0_values'][:,IX1][IX0]. \n\n**Details:** We load the data. And as an exercise and sanity check we produce plots for cell proliferation cycle visualization - so-called G1/S-G2/M plots. The results are as expected - the \"cyclic\" structes is better seen comparing as to original files (see version 3 of the present notebook). Remark: one can compare with another (more simple denoising (\"pooling\")) here: https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte#G1/S-G2/M-plots-to-analyse-the-cell-cycle\n\n**Context:** See paper: https://arxiv.org/abs/2208.05229 for cell cycle analysis. \"Computational challenges of cell cycle analysis using single cell transcriptomics\" Alexander Chervov, Andrei Zinovyev. All the work for that paper has been done on Kaggle - see datasets and code therein: https://www.kaggle.com/alexandervc/datasets?scroll=true https://www.kaggle.com/andreizinovyev Hundreds of single cell RNA seq datasets were analyzed.\n\n\n\n### Remarks: \n    \n    ( in Vesrstion 1- 6 we look mainly on Megakaryocyte Progenitors - just almost smallest and representative (from previous analysis ) \n\n### Versions:\n\n#### 11 with MAGIC again \n\n#### 10 - for comparaison - same data as 9, but NO MAGIC \n\n\n#### 9 - added plots all cell types\n    \n    Same MAGIC CITEseq TEST\n\n#### 7,8 cosmetic changes\n\n\n#### 6 - return to MAGIC - same as version 4 - \"test\" part of CITE-seq (features)\n        \n\n#### 5 - for comparaison  NO(!) MAGIC - difference is STRIKING\n\n    For days 2,4 - we cannot see clear \"cyclic\" structures without denoising.\n    Compare with denoised versions - where \"cyclic\" structure can be clearly seen ! (E.g. version 4) \n    \"test\" part of CITE-seq (features)\n\n\n#### 4 - MAGIC on \"test\" part of CITE-seq (features)\n\n    \"Cyclic\" structure is seen VERY well for many days - may be because only one donor, but for \"train\" part we have 3 donors and they mix and create less clear picture\n\n#### 3 - for comparaison look on original input file\n\n    The \"cyclic\" structure is less clearly seen for original files, comparing to denoised - as expected. \n\n#### 1,2 - work with MAGIC denoised data. \n\n","metadata":{}},{"cell_type":"markdown","source":"# Install/import","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-31T16:43:00.355528Z","iopub.execute_input":"2022-10-31T16:43:00.356261Z","iopub.status.idle":"2022-10-31T16:43:00.375389Z","shell.execute_reply.started":"2022-10-31T16:43:00.356206Z","shell.execute_reply":"2022-10-31T16:43:00.373616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nimport time\nt0start = time.time()\n\nimport pandas as pd\nimport numpy as np\nimport os\nimport sys\n\nimport matplotlib.pyplot as plt\n#plt.style.use('dark_background')\nimport seaborn as sns\n\n#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n\n\nimport h5py\n!pip install hdf5plugin~=2.0 # https://forum.hdfgroup.org/t/cant-open-directory-usr-local-hdf5-lib-plugin/9738/4\nimport hdf5plugin\n\n# !pip install scanpy\n# import scanpy as sc\n# import anndata\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")\n\ndf_cell = pd.read_csv(FP_CELL_METADATA)\n#df_cell","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:00.378471Z","iopub.execute_input":"2022-10-31T16:43:00.379506Z","iopub.status.idle":"2022-10-31T16:43:09.342007Z","shell.execute_reply.started":"2022-10-31T16:43:00.379457Z","shell.execute_reply":"2022-10-31T16:43:09.340382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# \"Load\" data and work with it ","metadata":{}},{"cell_type":"code","source":"print('Look on Features: ')\nfilename = '/kaggle/input/citeseq-denoised/test_cite_inputs_denoised.h5'\nfilename = '/kaggle/input/citeseq-denoised/train_cite_inputs_denoised.h5'\n#filename = '/kaggle/input/open-problems-multimodal/train_multi_inputs.h5'\n#filename = '/kaggle/input/open-problems-multimodal/train_multi_inputs.h5'\nfilename = '/kaggle/input/open-problems-multimodal/train_cite_inputs.h5'\n\n#filename = '/kaggle/input/citeseq-denoised/test_cite_inputs_denoised.h5'\n\n#filename = '/kaggle/input/open-problems-multimodal/test_cite_inputs.h5'\n\nf2 = h5py.File(filename,'r')#, mode)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.343880Z","iopub.execute_input":"2022-10-31T16:43:09.344280Z","iopub.status.idle":"2022-10-31T16:43:09.358926Z","shell.execute_reply.started":"2022-10-31T16:43:09.344246Z","shell.execute_reply":"2022-10-31T16:43:09.357586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f2.keys() )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.362340Z","iopub.execute_input":"2022-10-31T16:43:09.363222Z","iopub.status.idle":"2022-10-31T16:43:09.375173Z","shell.execute_reply.started":"2022-10-31T16:43:09.363161Z","shell.execute_reply":"2022-10-31T16:43:09.373987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"str_data_inf = ''\nif filename == '/kaggle/input/open-problems-multimodal/train_cite_inputs.h5':\n    first_key = 'train_cite_inputs'\n    str_data_inf = ' CITEseq Train '\n# elif filename == '/kaggle/input/open-problems-multimodal/train_multi_inputs.h5':\n#     first_key = 'train_multi_inputs'\n#     str_data_inf = ' Multi Train Features '\nelif filename == '/kaggle/input/open-problems-multimodal/test_cite_inputs.h5':\n    first_key = 'test_cite_inputs'#'train_multi_inputs' \n    str_data_inf = ' CITEseq Test '\nelif filename == '/kaggle/input/citeseq-denoised/test_cite_inputs_denoised.h5':\n    first_key = '0'\n    str_data_inf = ' MAGIC CITEseq Test '\nelif filename == '/kaggle/input/citeseq-denoised/train_cite_inputs_denoised.h5':\n    first_key = '0'\n    str_data_inf = ' MAGIC CITEseq Train '\n# else:\n#     first_key = '0'#'train_multi_inputs' \n    \nprint( f2[first_key].keys() )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.376500Z","iopub.execute_input":"2022-10-31T16:43:09.376880Z","iopub.status.idle":"2022-10-31T16:43:09.390443Z","shell.execute_reply.started":"2022-10-31T16:43:09.376849Z","shell.execute_reply":"2022-10-31T16:43:09.389110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for k in f2[first_key].keys():\n    print(type(f2[first_key][k]), f2[first_key][k].shape)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.392337Z","iopub.execute_input":"2022-10-31T16:43:09.392845Z","iopub.status.idle":"2022-10-31T16:43:09.412841Z","shell.execute_reply.started":"2022-10-31T16:43:09.392792Z","shell.execute_reply":"2022-10-31T16:43:09.411715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f2[first_key]['axis0'][0], f2[first_key]['axis1'][0], f2[first_key]['block0_items'][0],  f2[first_key]['block0_values'][0,0], ","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.414653Z","iopub.execute_input":"2022-10-31T16:43:09.415151Z","iopub.status.idle":"2022-10-31T16:43:09.450965Z","shell.execute_reply.started":"2022-10-31T16:43:09.415105Z","shell.execute_reply":"2022-10-31T16:43:09.449372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(f2[first_key]['block0_values'][:10,:10])","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.453257Z","iopub.execute_input":"2022-10-31T16:43:09.453755Z","iopub.status.idle":"2022-10-31T16:43:09.470589Z","shell.execute_reply.started":"2022-10-31T16:43:09.453714Z","shell.execute_reply":"2022-10-31T16:43:09.469430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( dir(f2[first_key]['block0_values'][:10,:10]) )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.472489Z","iopub.execute_input":"2022-10-31T16:43:09.472916Z","iopub.status.idle":"2022-10-31T16:43:09.483171Z","shell.execute_reply.started":"2022-10-31T16:43:09.472883Z","shell.execute_reply":"2022-10-31T16:43:09.481870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_genes_ensembl_ids = [t.decode().split('_')[0] for t in f2[first_key]['axis0'] ]\nprint( len(list_genes_ensembl_ids), list_genes_ensembl_ids[:10] )\nlist_genes_symbols = [t.decode().split('_')[1] for t in f2[first_key]['axis0'] ]\nprint( len(list_genes_symbols), list_genes_symbols[:10] )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:09.488627Z","iopub.execute_input":"2022-10-31T16:43:09.489495Z","iopub.status.idle":"2022-10-31T16:43:13.540370Z","shell.execute_reply.started":"2022-10-31T16:43:09.489454Z","shell.execute_reply":"2022-10-31T16:43:13.539008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_cells_barcodes = [t.decode() for t in f2[first_key]['axis1'] ]\nprint(len(list_cells_barcodes) , list_cells_barcodes[:10] )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:13.542216Z","iopub.execute_input":"2022-10-31T16:43:13.542610Z","iopub.status.idle":"2022-10-31T16:43:20.013753Z","shell.execute_reply.started":"2022-10-31T16:43:13.542576Z","shell.execute_reply":"2022-10-31T16:43:20.012304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( type(f2[first_key]['block0_values']), dir(f2[first_key]['block0_values'] ) )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:20.015454Z","iopub.execute_input":"2022-10-31T16:43:20.016234Z","iopub.status.idle":"2022-10-31T16:43:20.025035Z","shell.execute_reply.started":"2022-10-31T16:43:20.016183Z","shell.execute_reply":"2022-10-31T16:43:20.023664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M cell proliferation cycle plot for all cells \n\nIt will not be very clear - we need to separate days and cell types - see next subsections.","metadata":{}},{"cell_type":"code","source":"\nG1S_genes_Tirosh = ['ENSMUSG00000005410', 'ENSMUSG00000027342', 'ENSMUSG00000025747', 'ENSMUSG00000024742', 'ENSMUSG00000002870', 'ENSMUSG00000022673', 'ENSMUSG00000030978', 'ENSMUSG00000029591', 'ENSMUSG00000031821', 'ENSMUSG00000026355', 'ENSMUSG00000055612', 'ENSMUSG00000037474', 'ENSMUSG00000025395', 'ENSMUSG00000001228', 'ENSMUSG00000031629', 'ENSMUSG00000025001', 'ENSMUSG00000023104', 'ENSMUSG00000028884', 'ENSMUSG00000028693', 'ENSMUSG00000030346', 'ENSMUSG00000006715', 'ENSMUSG00000027242', 'ENSMUSG00000004642', 'ENSMUSG00000028212', 'ENSMUSG00000041712', 'ENSMUSG00000030726', 'ENSMUSG00000024151', 'ENSMUSG00000022360', 'ENSMUSG00000027323', 'ENSMUSG00000020649', 'ENSMUSG00000000028', 'ENSMUSG00000017499', 'ENSMUSG00000039748', 'ENSMUSG00000032397', 'ENSMUSG00000022422', 'ENSMUSG00000030528', 'ENSMUSG00000028282', 'ENSMUSG00000028560', 'ENSMUSG00000042489', 'ENSMUSG00000006678', 'ENSMUSG00000022945', 'ENSMUSG00000034329', 'ENSMUSG00000046179']\nG2M_genes_Tirosh = ['ENSMUSG00000054717', 'ENSMUSG00000019942', 'ENSMUSG00000027306', 'ENSMUSG00000001403', 'ENSMUSG00000017716', 'ENSMUSG00000027469', 'ENSMUSG00000020914', 'ENSMUSG00000024056', 'ENSMUSG00000062248', 'ENSMUSG00000026683', 'ENSMUSG00000028044', 'ENSMUSG00000031004', 'ENSMUSG00000019961', 'ENSMUSG00000026605', 'ENSMUSG00000037313', 'ENSMUSG00000020808', 'ENSMUSG00000034349', 'ENSMUSG00000032218', 'ENSMUSG00000048327', 'ENSMUSG00000037725', 'ENSMUSG00000020897', 'ENSMUSG00000027379', 'ENSMUSG00000012443', 'ENSMUSG00000015749', 'ENSMUSG00000036752', 'ENSMUSG00000022385', 'ENSMUSG00000024795', 'ENSMUSG00000044783', 'ENSMUSG00000023505', 'ENSMUSG00000020737', 'ENSMUSG00000006398', 'ENSMUSG00000038379', 'ENSMUSG00000044201', 'ENSMUSG00000028678', 'ENSMUSG00000022391', 'ENSMUSG00000038252', 'ENSMUSG00000037544', 'ENSMUSG00000048922', 'ENSMUSG00000028873', 'ENSMUSG00000027699', 'ENSMUSG00000032254', 'ENSMUSG00000020330', 'ENSMUSG00000027496', 'ENSMUSG00000068744', 'ENSMUSG00000036777', 'ENSMUSG00000004880', 'ENSMUSG00000040549', 'ENSMUSG00000045328', 'ENSMUSG00000005698', 'ENSMUSG00000026622', 'ENSMUSG00000035293', 'ENSMUSG00000074802', 'ENSMUSG00000009575', 'ENSMUSG00000029177']\n\nG1S_genes_Tirosh = ['MCM5', 'PCNA', 'TYMS', 'FEN1', 'MCM2', 'MCM4', 'RRM1', 'UNG', 'GINS2', 'MCM6', 'CDCA7', 'DTL', 'PRIM1', 'UHRF1', 'MLF1IP', 'HELLS', 'RFC2', 'RPA2', 'NASP', 'RAD51AP1', 'GMNN', 'WDR76', 'SLBP', 'CCNE2', 'UBR7', 'POLD3', 'MSH2', 'ATAD2', 'RAD51', 'RRM2', 'CDC45', 'CDC6', 'EXO1', 'TIPIN', 'DSCC1', 'BLM', 'CASP8AP2', 'USP1', 'CLSPN', 'POLA1', 'CHAF1B', 'BRIP1', 'E2F8']\nG2M_genes_Tirosh = ['HMGB2', 'CDK1', 'NUSAP1', 'UBE2C', 'BIRC5', 'TPX2', 'TOP2A', 'NDC80', 'CKS2', 'NUF2', 'CKS1B', 'MKI67', 'TMPO', 'CENPF', 'TACC3', 'FAM64A', 'SMC4', 'CCNB2', 'CKAP2L', 'CKAP2', 'AURKB', 'BUB1', 'KIF11', 'ANP32E', 'TUBB4B', 'GTSE1', 'KIF20B', 'HJURP', 'CDCA3', 'HN1', 'CDC20', 'TTK', 'CDC25C', 'KIF2C', 'RANGAP1', 'NCAPD2', 'DLGAP5', 'CDCA2', 'CDCA8', 'ECT2', 'KIF23', 'HMMR', 'AURKA', 'PSRC1', 'ANLN', 'LBR', 'CKAP5', 'CENPE', 'CTCF', 'NEK2', 'G2E3', 'GAS2L3', 'CBX5', 'CENPA']\nlist_genes_fastCCsign = ['CDK1', 'UBE2C', 'TOP2A', 'TMPO', 'HJURP', 'RRM1', 'RAD51AP1', 'RRM2', 'CDC45', 'BLM', 'BRIP1', 'E2F8', 'HIST2H2AC']\n\nG1S_genes_Tirosh = ['ENSG00000100297', 'ENSG00000132646', 'ENSG00000176890', 'ENSG00000168496', 'ENSG00000073111', 'ENSG00000104738', 'ENSG00000167325', 'ENSG00000076248', 'ENSG00000131153', 'ENSG00000076003', 'ENSG00000144354', 'ENSG00000143476', 'ENSG00000198056', 'ENSG00000276043', 'ENSG00000151725', 'ENSG00000119969', 'ENSG00000049541', 'ENSG00000117748', 'ENSG00000132780', 'ENSG00000111247', 'ENSG00000112312', 'ENSG00000092470', 'ENSG00000163950', 'ENSG00000175305', 'ENSG00000012963', 'ENSG00000077514', 'ENSG00000095002', 'ENSG00000156802', 'ENSG00000051180', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000094804', 'ENSG00000174371', 'ENSG00000075131', 'ENSG00000136982', 'ENSG00000197299', 'ENSG00000118412', 'ENSG00000162607', 'ENSG00000092853', 'ENSG00000101868', 'ENSG00000159259', 'ENSG00000136492', 'ENSG00000129173']\nG2M_genes_Tirosh = ['ENSG00000164104', 'ENSG00000170312', 'ENSG00000137804', 'ENSG00000175063', 'ENSG00000089685', 'ENSG00000088325', 'ENSG00000131747', 'ENSG00000080986', 'ENSG00000123975', 'ENSG00000143228', 'ENSG00000173207', 'ENSG00000148773', 'ENSG00000120802', 'ENSG00000117724', 'ENSG00000013810', 'ENSG00000129195', 'ENSG00000113810', 'ENSG00000157456', 'ENSG00000169607', 'ENSG00000136108', 'ENSG00000178999', 'ENSG00000169679', 'ENSG00000138160', 'ENSG00000143401', 'ENSG00000188229', 'ENSG00000075218', 'ENSG00000138182', 'ENSG00000123485', 'ENSG00000111665', 'ENSG00000189159', 'ENSG00000117399', 'ENSG00000112742', 'ENSG00000158402', 'ENSG00000142945', 'ENSG00000100401', 'ENSG00000010292', 'ENSG00000126787', 'ENSG00000184661', 'ENSG00000134690', 'ENSG00000114346', 'ENSG00000137807', 'ENSG00000072571', 'ENSG00000087586', 'ENSG00000134222', 'ENSG00000011426', 'ENSG00000143815', 'ENSG00000175216', 'ENSG00000138778', 'ENSG00000102974', 'ENSG00000117650', 'ENSG00000092140', 'ENSG00000139354', 'ENSG00000094916', 'ENSG00000115163']\nlist_genes_fastCCsign = ['ENSG00000170312', 'ENSG00000175063', 'ENSG00000131747', 'ENSG00000120802', 'ENSG00000123485', 'ENSG00000167325', 'ENSG00000111247', 'ENSG00000171848', 'ENSG00000093009', 'ENSG00000197299', 'ENSG00000136492', 'ENSG00000129173', 'ENSG00000184260']\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:20.026525Z","iopub.execute_input":"2022-10-31T16:43:20.026901Z","iopub.status.idle":"2022-10-31T16:43:20.047954Z","shell.execute_reply.started":"2022-10-31T16:43:20.026870Z","shell.execute_reply":"2022-10-31T16:43:20.047020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nlen(IX), len( G1S_genes_Tirosh)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:20.048927Z","iopub.execute_input":"2022-10-31T16:43:20.049245Z","iopub.status.idle":"2022-10-31T16:43:20.076041Z","shell.execute_reply.started":"2022-10-31T16:43:20.049217Z","shell.execute_reply":"2022-10-31T16:43:20.074565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nv1 = f2[first_key]['block0_values'][:,IX].sum(axis = 1 )","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:43:20.078187Z","iopub.execute_input":"2022-10-31T16:43:20.078748Z","iopub.status.idle":"2022-10-31T16:44:02.128651Z","shell.execute_reply.started":"2022-10-31T16:43:20.078709Z","shell.execute_reply":"2022-10-31T16:44:02.127266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nlen(IX), len( G2M_genes_Tirosh)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:02.130854Z","iopub.execute_input":"2022-10-31T16:44:02.131212Z","iopub.status.idle":"2022-10-31T16:44:02.142640Z","shell.execute_reply.started":"2022-10-31T16:44:02.131183Z","shell.execute_reply":"2022-10-31T16:44:02.141426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nv2 = f2[first_key]['block0_values'][:,IX].sum(axis = 1 )\nprint( len(v2), type(v2), v2.shape) ","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:02.144155Z","iopub.execute_input":"2022-10-31T16:44:02.144672Z","iopub.status.idle":"2022-10-31T16:44:21.018438Z","shell.execute_reply.started":"2022-10-31T16:44:02.144627Z","shell.execute_reply":"2022-10-31T16:44:21.017321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(x = v1, y=v2)\nplt.title(str_data_inf)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:21.020009Z","iopub.execute_input":"2022-10-31T16:44:21.020367Z","iopub.status.idle":"2022-10-31T16:44:21.420460Z","shell.execute_reply.started":"2022-10-31T16:44:21.020338Z","shell.execute_reply":"2022-10-31T16:44:21.419522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=1, cell_type = MkP (Megakaryocyte Progenitors)","metadata":{}},{"cell_type":"code","source":"df_cite_train_y = pd.read_hdf('../input/open-problems-multimodal/train_cite_targets.h5')\ngenes = list(df_cite_train_y.keys())\ndf_cell_hue = df_cell.merge(df_cite_train_y, on = 'cell_id')\ndisplay(df_cell_hue.head(5))","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:21.421827Z","iopub.execute_input":"2022-10-31T16:44:21.422905Z","iopub.status.idle":"2022-10-31T16:44:22.508992Z","shell.execute_reply.started":"2022-10-31T16:44:21.422863Z","shell.execute_reply":"2022-10-31T16:44:22.507647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 1\nday = 7\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in genes: # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:22.512707Z","iopub.execute_input":"2022-10-31T16:44:22.513794Z","iopub.status.idle":"2022-10-31T16:44:41.256106Z","shell.execute_reply.started":"2022-10-31T16:44:22.513742Z","shell.execute_reply":"2022-10-31T16:44:41.253570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=2, cell_type = MkP (Megakaryocyte Progenitors)\n\nCompare with: https://www.kaggle.com/code/alexandervc/mmscel-cell-cycle-03-daybydaychange-megakaryocyte?scriptVersionId=105307517&cellId=12\n\n","metadata":{}},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 2\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in genes: # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.257701Z","iopub.status.idle":"2022-10-31T16:44:41.258577Z","shell.execute_reply.started":"2022-10-31T16:44:41.258235Z","shell.execute_reply":"2022-10-31T16:44:41.258264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=3, cell_type = MkP (Megakaryocyte Progenitors)","metadata":{}},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 3\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in genes: # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.260849Z","iopub.status.idle":"2022-10-31T16:44:41.261461Z","shell.execute_reply.started":"2022-10-31T16:44:41.261167Z","shell.execute_reply":"2022-10-31T16:44:41.261195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=4, cell_type = MkP (Megakaryocyte Progenitors)","metadata":{}},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 4\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in genes: # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.263037Z","iopub.status.idle":"2022-10-31T16:44:41.264078Z","shell.execute_reply.started":"2022-10-31T16:44:41.263762Z","shell.execute_reply":"2022-10-31T16:44:41.263801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=2, cell_type = MkP (Megakaryocyte Progenitors), Input data for hue\n","metadata":{}},{"cell_type":"code","source":"cite_cols_important={'CD86': ['ENSG00000114013_CD86'],\n'HLA-A-B-C': ['ENSG00000204525_HLA-C'],\n'CD274': ['ENSG00000120217_CD274'],\n'CD270': ['ENSG00000157873_TNFRSF14'],\n'CD155': ['ENSG00000073008_PVR'],\n'CD112': ['ENSG00000130202_NECTIN2'],\n'CD47': ['ENSG00000196776_CD47'],\n'CD48': ['ENSG00000117091_CD48'],\n'CD40': ['ENSG00000101017_CD40'],\n'CD154': ['ENSG00000102245_CD40LG'],\n'CD52': ['ENSG00000169442_CD52'],\n'CD3': ['ENSG00000167286_CD3D'],\n'CD8': [],\n'CD56': ['ENSG00000149294_NCAM1'],\n'CD19': ['ENSG00000177455_CD19'],\n'CD33': ['ENSG00000105383_CD33'],\n'CD11c': ['ENSG00000140678_ITGAX'],\n'CD45RA': ['ENSG00000081237_PTPRC'],\n'CD123': ['ENSG00000185291_IL3RA'],\n'CD7': ['ENSG00000173762_CD7'],\n'CD105': ['ENSG00000106991_ENG'],\n'CD49f': ['ENSG00000091409_ITGA6'],\n'CD194': ['ENSG00000183813_CCR4'],\n'CD4': ['ENSG00000010610_CD4'],\n'CD44': ['ENSG00000026508_CD44'],\n'CD14': ['ENSG00000170458_CD14'],\n'CD16': [],\n'CD25': ['ENSG00000134460_IL2RA'],\n'CD45RO': ['ENSG00000081237_PTPRC'],\n'CD279': [],\n'TIGIT': [],\n'Mouse-IgG1': [],\n'Mouse-IgG2a': [],\n'Mouse-IgG2b': [],\n'Rat-IgG2b': [],\n'CD20': ['ENSG00000156738_MS4A1'],\n'CD335': ['ENSG00000189430_NCR1'],\n'CD31': ['ENSG00000261371_PECAM1'],\n'Podoplanin': [],\n'CD146': ['ENSG00000076706_MCAM'],\n'IgM': ['ENSG00000211899_IGHM'],\n'CD5': [],\n'CD195': ['ENSG00000160791_CCR5'],\n'CD32': ['ENSG00000143226_FCGR2A'],\n'CD196': [],\n'CD185': ['ENSG00000160683_CXCR5'],\n'CD103': ['ENSG00000083457_ITGAE'],\n'CD69': ['ENSG00000110848_CD69'],\n'CD62L': ['ENSG00000188404_SELL'],\n'CD161': ['ENSG00000111796_KLRB1'],\n'CD152': [],\n'CD223': ['ENSG00000089692_LAG3'],\n'KLRG1': ['ENSG00000139187_KLRG1'],\n'CD27': ['ENSG00000139193_CD27'],\n'CD107a': ['ENSG00000185896_LAMP1'],\n'CD95': ['ENSG00000026103_FAS'],\n'CD134': ['ENSG00000186827_TNFRSF4'],\n'HLA-DR': ['ENSG00000204287_HLA-DRA'],\n'CD1c': ['ENSG00000158481_CD1C'],\n'CD11b': ['ENSG00000169896_ITGAM'],\n'CD64': ['ENSG00000150337_FCGR1A'],\n'CD141': ['ENSG00000178726_THBD'],\n'CD1d': ['ENSG00000158473_CD1D'],\n'CD314': [],\n'CD35': ['ENSG00000203710_CR1'],\n'CD57': [],\n'CD272': [],\n'CD278': ['ENSG00000163600_ICOS'],\n'CD58': ['ENSG00000116815_CD58'],\n'CD39': ['ENSG00000138185_ENTPD1'],\n'CX3CR1': ['ENSG00000168329_CX3CR1'],\n'CD24': ['ENSG00000272398_CD24'],\n'CD21': ['ENSG00000117322_CR2'],\n'CD11a': ['ENSG00000005844_ITGAL'],\n'CD79b': ['ENSG00000007312_CD79B'],\n'CD244': ['ENSG00000122223_CD244'],\n'CD169': [],\n'integrinB7': ['ENSG00000139626_ITGB7'],\n'CD268': ['ENSG00000159958_TNFRSF13C'],\n'CD42b': ['ENSG00000185245_GP1BA'],\n'CD54': ['ENSG00000090339_ICAM1'],\n'CD62P': ['ENSG00000174175_SELP'],\n'CD119': ['ENSG00000027697_IFNGR1'],\n'TCR': [],\n'Rat-IgG1': [],\n'Rat-IgG2a': [],\n'CD192': ['ENSG00000121807_CCR2'],\n'CD122': ['ENSG00000100385_IL2RB'],\n'FceRIa': ['ENSG00000179639_FCER1A'],\n'CD41': ['ENSG00000005961_ITGA2B'],\n'CD137': ['ENSG00000049249_TNFRSF9'],\n'CD163': ['ENSG00000177575_CD163'],\n'CD83': ['ENSG00000112149_CD83'],\n'CD124': ['ENSG00000077238_IL4R'],\n'CD13': ['ENSG00000166825_ANPEP'],\n'CD2': ['ENSG00000116824_CD2'],\n'CD226': ['ENSG00000150637_CD226'],\n'CD29': ['ENSG00000150093_ITGB1'],\n'CD303': ['ENSG00000198178_CLEC4C'],\n'CD49b': ['ENSG00000164171_ITGA2'],\n'CD81': ['ENSG00000110651_CD81'],\n'IgD': ['ENSG00000211898_IGHD'],\n'CD18': ['ENSG00000160255_ITGB2'],\n'CD28': [],\n'CD38': ['ENSG00000004468_CD38'],\n'CD127': ['ENSG00000168685_IL7R'],\n'CD45': ['ENSG00000081237_PTPRC'],\n'CD22': ['ENSG00000012124_CD22'],\n'CD71': ['ENSG00000072274_TFRC'],\n'CD26': ['ENSG00000197635_DPP4'],\n'CD115': ['ENSG00000182578_CSF1R'],\n'CD63': ['ENSG00000135404_CD63'],\n'CD304': ['ENSG00000099250_NRP1'],\n'CD36': ['ENSG00000135218_CD36'],\n'CD172a': ['ENSG00000198053_SIRPA'],\n'CD72': ['ENSG00000137101_CD72'],\n'CD158': [],\n'CD93': ['ENSG00000125810_CD93'],\n'CD49a': ['ENSG00000213949_ITGA1'],\n'CD49d': ['ENSG00000115232_ITGA4'],\n'CD73': [],\n'CD9': ['ENSG00000010278_CD9'],\n'TCRVa7.2': [],\n'TCRVd2': [],\n'LOX-1': ['ENSG00000173391_OLR1'],\n'CD158b': [],\n'CD158e1': [],\n'CD142': ['ENSG00000117525_F3'],\n'CD319': ['ENSG00000026751_SLAMF7'],\n'CD352': ['ENSG00000162739_SLAMF6'],\n'CD94': ['ENSG00000134539_KLRD1'],\n'CD162': ['ENSG00000110876_SELPLG'],\n'CD85j': ['ENSG00000104972_LILRB1'],\n'CD23': ['ENSG00000104921_FCER2'],\n'CD328': ['ENSG00000168995_SIGLEC7'],\n'HLA-E': ['ENSG00000204592_HLA-E'],\n'CD82': ['ENSG00000085117_CD82'],\n'CD101': ['ENSG00000134256_CD101'],\n'CD88': ['ENSG00000197405_C5AR1'],\n'CD224': ['ENSG00000100031_GGT1']}","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:54:39.744291Z","iopub.execute_input":"2022-10-31T16:54:39.744783Z","iopub.status.idle":"2022-10-31T16:54:39.770519Z","shell.execute_reply.started":"2022-10-31T16:54:39.744743Z","shell.execute_reply":"2022-10-31T16:54:39.769291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_dict = dict([(value[0], key) for key, value in cite_cols_important.items() if value!=[]])\nprint(new_dict.values())\nkeys = list(new_dict.keys())+['cell_id','day','donor','cell_type','technology', 'gene_id']\nprint(keys)","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:54:51.266702Z","iopub.execute_input":"2022-10-31T16:54:51.267194Z","iopub.status.idle":"2022-10-31T16:54:51.274572Z","shell.execute_reply.started":"2022-10-31T16:54:51.267159Z","shell.execute_reply":"2022-10-31T16:54:51.273306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite_train_y = pd.read_hdf(FP_CITE_TRAIN_INPUTS, axis = ['cell_id','gene_id'], usecols=keys)\ndf_cell_hue = df_cell.merge(df_cite_train_y, on = 'cell_id')\ndisplay(df_cell_hue.head(5))","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:54:54.542196Z","iopub.execute_input":"2022-10-31T16:54:54.542720Z","iopub.status.idle":"2022-10-31T16:56:02.396345Z","shell.execute_reply.started":"2022-10-31T16:54:54.542677Z","shell.execute_reply":"2022-10-31T16:56:02.395231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cite_train_y.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:52:29.571657Z","iopub.execute_input":"2022-10-31T16:52:29.572163Z","iopub.status.idle":"2022-10-31T16:52:29.610228Z","shell.execute_reply.started":"2022-10-31T16:52:29.572129Z","shell.execute_reply":"2022-10-31T16:52:29.608909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_to_be_deleted = [col for col in df_cell_hue.columns if col not in keys]\nprint(len(cols_to_be_deleted))\nif len(cols_to_be_deleted)!=22050:\n    df_cell_hue.drop(cols_to_be_deleted, axis=1, inplace=True)\ndf_cell_hue.rename(columns = new_dict, inplace = True)\ndf_cell_hue.columns","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:56:39.531652Z","iopub.execute_input":"2022-10-31T16:56:39.532100Z","iopub.status.idle":"2022-10-31T16:56:39.716315Z","shell.execute_reply.started":"2022-10-31T16:56:39.532067Z","shell.execute_reply":"2022-10-31T16:56:39.714907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(new_dict.values()))","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:56:44.394923Z","iopub.execute_input":"2022-10-31T16:56:44.395506Z","iopub.status.idle":"2022-10-31T16:56:44.401815Z","shell.execute_reply.started":"2022-10-31T16:56:44.395461Z","shell.execute_reply":"2022-10-31T16:56:44.400723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 2\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in new_dict.values(): # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:57:28.382723Z","iopub.execute_input":"2022-10-31T16:57:28.383183Z","iopub.status.idle":"2022-10-31T16:58:53.116407Z","shell.execute_reply.started":"2022-10-31T16:57:28.383152Z","shell.execute_reply":"2022-10-31T16:58:53.115260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=3, cell_type = MkP (Megakaryocyte Progenitors), Input data for hue\n","metadata":{}},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 3\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in new_dict.values(): # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T17:00:08.264772Z","iopub.execute_input":"2022-10-31T17:00:08.265150Z","iopub.status.idle":"2022-10-31T17:01:24.809492Z","shell.execute_reply.started":"2022-10-31T17:00:08.265119Z","shell.execute_reply":"2022-10-31T17:01:24.807987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# G1/S-G2/M plot for day=4, cell_type = MkP (Megakaryocyte Progenitors), Input data for hue\n","metadata":{}},{"cell_type":"code","source":"palette = 'rainbow'#'viridis'\nn_x_subplots = 4\nday = 4\n\nmask = (df_cell_hue['day'] == day) & (df_cell_hue['cell_type'] == 'MkP' )                      \ncell_ids = list(df_cell_hue['cell_id'][mask])\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G1S_genes_Tirosh) >0) [0]\nv1 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n# Remark: Error - if write [IX0,IX] - \"TypeError: Only one indexing vector or array is currently allowed for fancy indexing\"\n# so we proceed as above - in two steps\n\nIX = np.where(pd.Series(list_genes_ensembl_ids ).isin(G2M_genes_Tirosh) >0) [0]\nv2 = (f2[first_key]['block0_values'][:,IX]).sum(axis = 1 )\n\nIX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n\n\nc = 0\nfor gene_id in new_dict.values(): # ['day', 'donor', 'cell_type']:# , 'technology']:\n    if c % n_x_subplots == 0:\n        if c > 0:\n            plt.show()\n        fig = plt.figure(figsize = (20,5) ); c = 0\n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n\n    c += 1\n    fig.add_subplot(1,n_x_subplots ,c)    \n\n    v = df_cell_hue[gene_id]\n\n    ax = sns.scatterplot(x=v1[IX0], y=v2[IX0], hue=v[IX0], palette = palette )\n    plt.setp(ax.get_legend().get_texts(), fontsize=10) # for legend text\n    plt.setp(ax.get_legend().get_title(), fontsize=10) # for legend title    \n    plt.title('Colored by '+gene_id, fontsize = 10)\n\n#plt.title(str_data_inf + 'Day 4, MkP (Megakaryocyte Progenitors) ', fontsize = 20)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T17:02:44.486319Z","iopub.execute_input":"2022-10-31T17:02:44.486713Z","iopub.status.idle":"2022-10-31T17:04:05.044303Z","shell.execute_reply.started":"2022-10-31T17:02:44.486681Z","shell.execute_reply":"2022-10-31T17:04:05.043116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loop over all cell types ","metadata":{}},{"cell_type":"code","source":"dict_cell_types = { 'MasP': 'Mast Cell Progenitor',\n'MkP': 'Megakaryocyte Progenitor',\n'NeuP': 'Neutrophil Progenitor',\n'MoP': 'Monocyte Progenitor',\n'EryP': 'Erythrocyte Progenitor',\n'HSC': 'Hematoploetic Stem Cell',\n'BP': 'B-Cell Progenitor'}","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.284933Z","iopub.status.idle":"2022-10-31T16:44:41.285887Z","shell.execute_reply.started":"2022-10-31T16:44:41.285551Z","shell.execute_reply":"2022-10-31T16:44:41.285581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_cell.head(1)\nd = pd.DataFrame(index =list_cells_barcodes,  )\nd = d.join(df_cell.set_index('cell_id'), how = 'left')\nd['str_donor'] = d['donor'].apply(lambda x: str(x))\nd","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.287520Z","iopub.status.idle":"2022-10-31T16:44:41.288527Z","shell.execute_reply.started":"2022-10-31T16:44:41.288308Z","shell.execute_reply":"2022-10-31T16:44:41.288335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for cell_type in ['MkP','HSC',  'EryP', 'NeuP', 'MasP' ,  'MoP' , 'BP' ]:\n    print();print();print()\n    print(cell_type)\n    \n    for day in [2,3,4,7]:\n\n        mask = (df_cell['day'] == day) & (df_cell['cell_type'] == cell_type )                      \n        #print('Day',day, mask.sum() )\n        cell_ids = list(df_cell['cell_id'][mask])\n        #print(len(cell_ids) , cell_ids[:10])\n\n        IX0 = np.where( pd.Series(list_cells_barcodes).isin(cell_ids) > 0)[0]\n        print('Day', day,mask.sum(), len(IX0))# , IX0[:10]) # , len(cell_ids ))\n\n        plt.figure(figsize = (20,10))\n        if d['str_donor'].iloc[IX0].nunique() > 1:\n            sns.scatterplot(x = v1[IX0], y=v2[IX0] , hue = d['str_donor'].iloc[IX0], palette = 'rainbow')#  ['blue'])\n        else:\n            sns.scatterplot(x = v1[IX0], y=v2[IX0] , hue = d['str_donor'].iloc[IX0], palette =  ['blue'])\n            \n        plt.xlabel('G1/S score', fontsize = 20 )\n        plt.ylabel('G2/M score', fontsize = 20 )\n        plt.title(str_data_inf + '  Day '+str(day) + ', '+ cell_type +  ' ' + dict_cell_types[cell_type] + ' n_cells=%d'%len(IX0) , fontsize = 20)\n\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.289948Z","iopub.status.idle":"2022-10-31T16:44:41.290800Z","shell.execute_reply.started":"2022-10-31T16:44:41.290582Z","shell.execute_reply":"2022-10-31T16:44:41.290609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )\n","metadata":{"execution":{"iopub.status.busy":"2022-10-31T16:44:41.292527Z","iopub.status.idle":"2022-10-31T16:44:41.293644Z","shell.execute_reply.started":"2022-10-31T16:44:41.293400Z","shell.execute_reply":"2022-10-31T16:44:41.293424Z"},"trusted":true},"execution_count":null,"outputs":[]}]}