{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"About this notebook:\n- the original notebook name: kagglenb003e-meta-and-eval\n- the goal of this notebook: EDA to see the sampling result of Multiome using metadata.csv and evaluation_ids.csv","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mf = pd.read_csv('../input/open-problems-multimodal/metadata.csv')\nef = pd.read_csv('../input/open-problems-multimodal/evaluation_ids.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:32:28.418882Z","iopub.execute_input":"2022-10-05T10:32:28.419304Z","iopub.status.idle":"2022-10-05T10:33:46.275815Z","shell.execute_reply.started":"2022-10-05T10:32:28.419269Z","shell.execute_reply":"2022-10-05T10:33:46.274375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to https://www.kaggle.com/competitions/open-problems-multimodal/data,<br><br>\nYour task is to predict the labels corresponding to the inputs in the test set. To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\n\n- **evaluation_ids.csv** - Identifies the labels from the test set to be evaluated. It provides a join key from the cell_id / gene_id identifiers of the label matrix to the row_id needed for the submission file.","metadata":{}},{"cell_type":"markdown","source":"- **metadata.csv**\n\n    -  cell_id - A unique identifier for each observed cell.\n    - donor - An identifier for the four cell donors.\n    - day - The day of the experiment the observation was made.\n    - technology - Either citeseq or multiome.\n    - cell_type - One of the above cell types or else hidden.","metadata":{}},{"cell_type":"markdown","source":"# Hypothesis\nevaluation_ids.csvにmultiomeのサンプル後の行が記載されてる?<br>\n[English] Does evaluation_ids.csv have which rows are to be evaluated for Multiome?","metadata":{}},{"cell_type":"code","source":"mf = mf[mf['technology']=='multiome'].reset_index(drop=True)\ndisplay(mf)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:33:46.277737Z","iopub.execute_input":"2022-10-05T10:33:46.278168Z","iopub.status.idle":"2022-10-05T10:33:46.384740Z","shell.execute_reply.started":"2022-10-05T10:33:46.278131Z","shell.execute_reply":"2022-10-05T10:33:46.383533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"assert len(mf) == mf['cell_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:33:46.387118Z","iopub.execute_input":"2022-10-05T10:33:46.387499Z","iopub.status.idle":"2022-10-05T10:33:46.455360Z","shell.execute_reply.started":"2022-10-05T10:33:46.387464Z","shell.execute_reply":"2022-10-05T10:33:46.453584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# cell_ids for Multiome\ncell_ids = mf['cell_id'].unique()\nassert len(mf) == len(cell_ids)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:33:46.457624Z","iopub.execute_input":"2022-10-05T10:33:46.458003Z","iopub.status.idle":"2022-10-05T10:33:46.493681Z","shell.execute_reply.started":"2022-10-05T10:33:46.457973Z","shell.execute_reply":"2022-10-05T10:33:46.492225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ef_m = ef[ef['cell_id'].isin(cell_ids)].reset_index(drop=True)\nef_c = ef[~ef['cell_id'].isin(cell_ids)].reset_index(drop=True)\nassert len(ef) == len(ef_m) + len(ef_c)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:33:46.495647Z","iopub.execute_input":"2022-10-05T10:33:46.496245Z","iopub.status.idle":"2022-10-05T10:34:01.253610Z","shell.execute_reply.started":"2022-10-05T10:33:46.496193Z","shell.execute_reply":"2022-10-05T10:34:01.251955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Multiome has only 'ENSG' for gene_id's suffix\nprint(ef_m['gene_id'].apply(lambda x: x[:4]).unique())","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:35:00.499363Z","iopub.execute_input":"2022-10-05T10:35:00.499783Z","iopub.status.idle":"2022-10-05T10:35:27.046664Z","shell.execute_reply.started":"2022-10-05T10:35:00.499751Z","shell.execute_reply":"2022-10-05T10:35:27.045473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cite has 'CD', 'HL', 'TI', 'Mo', 'Ra', 'Po', 'Ig', 'KL', 'CX', 'in', 'TC', 'Fc', 'LO' for gene_id's suffix\nprint(ef_c['gene_id'].apply(lambda x: x[:2]).unique())","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:36:40.431302Z","iopub.execute_input":"2022-10-05T10:36:40.433380Z","iopub.status.idle":"2022-10-05T10:36:43.728576Z","shell.execute_reply.started":"2022-10-05T10:36:40.433338Z","shell.execute_reply":"2022-10-05T10:36:43.726669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of unique cell_id for Multiome which is said to be sampled from entire cell_id for Multiome\nnum_cell_sample = ef_m['cell_id'].nunique()\nprint(num_cell_sample)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:43:06.556285Z","iopub.execute_input":"2022-10-05T10:43:06.556887Z","iopub.status.idle":"2022-10-05T10:43:11.242048Z","shell.execute_reply.started":"2022-10-05T10:43:06.556851Z","shell.execute_reply":"2022-10-05T10:43:11.240577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of entire cell_id for Multiome\nnum_cell_entire = mf['cell_id'].nunique()\nprint(num_cell_entire)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:43:24.900276Z","iopub.execute_input":"2022-10-05T10:43:24.900841Z","iopub.status.idle":"2022-10-05T10:43:24.965640Z","shell.execute_reply.started":"2022-10-05T10:43:24.900799Z","shell.execute_reply":"2022-10-05T10:43:24.963944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sampling ratio\nprint(num_cell_sample / num_cell_entire)","metadata":{"execution":{"iopub.status.busy":"2022-10-05T10:43:47.356197Z","iopub.execute_input":"2022-10-05T10:43:47.356747Z","iopub.status.idle":"2022-10-05T10:43:47.364623Z","shell.execute_reply.started":"2022-10-05T10:43:47.356703Z","shell.execute_reply":"2022-10-05T10:43:47.362931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Again,according to https://www.kaggle.com/competitions/open-problems-multimodal/data,<br><br>\nYour task is to predict the labels corresponding to the inputs in the test set. To facilitate submission scoring, we only require predictions on a subset of the Multiome data. This subset was created by **sampling 30% of the Multiome rows**, and for each row, 15% of the columns.<br><br>\n\n# Conclusion\nしかしcell_idで見ると, Multiomeは全体の約10%しかevaluationに提供されないようだ.<br>\n[English] However, in terms of cell_id, approximately only 10% of the entire cell_id seem to be served for evaluation.<br><br>\n~そしてここに現れるcell_idはMultiomeのCVを作る際に含めるべきだろう.~<br>\n~[English] And cell_ids listed here should be included in CV data for Multiome.~<br>\n\n# Next step\nHow many do cell_ids listed in evaluation_ids.csv appear in train data?","metadata":{}}]}