{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### In this notebook, we examine if [scgpt](https://github.com/bowang-lab/scGPT/tree/main) could be useful for this competition. Specifically we check how many genes are overlapped between scgpt and the competition's dataset","metadata":{}},{"cell_type":"markdown","source":"### 1. Get the competition's gene list","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport os","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"PATH = '/kaggle/input/open-problems-single-cell-perturbations'\nos.listdir(PATH)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"de_train = pd.read_parquet(f'{PATH}/de_train.parquet')\nde_train.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The gene names start from column 5.","metadata":{}},{"cell_type":"code","source":"genes = de_train.columns[5:].values.tolist()\nlen(genes)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 14211 genes in the competition dataset.","metadata":{}},{"cell_type":"markdown","source":"### 2. Get scgpt's gene list","metadata":{}},{"cell_type":"code","source":"!git clone https://github.com/bowang-lab/scGPT.git","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\ngene_json_path = './scGPT/scgpt/tokenizer/default_gene_vocab.json'\nwith open(gene_json_path) as f:\n    genes_scgpt = json.load(f)\nlen(genes_scgpt), type(genes_scgpt)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(genes_scgpt.items())[:5]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3. Get the genes in both competition data and scgpt's vocab","metadata":{}},{"cell_type":"code","source":"genes_in_both = [g for g in genes if g in genes_scgpt]\nlen(genes_in_both), len(genes_in_both)/len(genes)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 80% genes in the competition also exist in scgpt's vocab. So it is promissing to use scgpt for this competition.","metadata":{}},{"cell_type":"markdown","source":"### 4. Compare the expressed genes in `adata_train` and the overlapped genes `genes_in_both`","metadata":{}},{"cell_type":"code","source":"adata_train = pd.read_parquet(f'{PATH}/adata_train.parquet')\nprint(adata_train.shape)\nadata_train.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmask = adata_train.gene.isin(genes_scgpt)\nmask.sum()/adata_train.shape[0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusion: `95%` of the gene expressions in `adata_train` can be found in the vocab of `scgpt`.","metadata":{}}]}