{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Preview\nWe cannot see a preview of the .h5 format files on the competition page, so check the first few rows at first.\n\nThe following notebook was used for data loading.\n\nコンペページでは、.h5ファイルのプレビューが見れないので、ひとまず先頭行を確認します。\n\nデータの読み込みについては、下記のnotebookを参考にしました。\n\n[Getting Started - Data Loading created @ Peter Holderrieth](https://www.kaggle.com/code/peterholderrieth/getting-started-data-loading)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T21:21:49.972077Z","iopub.execute_input":"2022-08-11T21:21:49.972474Z","iopub.status.idle":"2022-08-11T21:21:49.977181Z","shell.execute_reply.started":"2022-08-11T21:21:49.972443Z","shell.execute_reply":"2022-08-11T21:21:49.975948Z"}}},{"cell_type":"markdown","source":"# Data files\n9 files are provided.\n\n9つのファイルが提供されています。\n1. metadata.csv\n>*     cell_id - A unique identifier for each observed cell. / 観測された各セルに一意な識別子。\n>*     donor - An identifier for the four cell donors. / 4人の細胞提供者の識別子。\n>*     day - The day of the experiment the observation was made. / 実験観察が行われた日付。\n>*     technology - Either citeseq or multiome. / \"citeseq\"か\"multiome\"のいずれか。\n>*     cell_type - One of the  following cell types or else hidden. / 下記のセルタイプのいずれか、またはそれ以外が\"hidden\"。\n>>* MasP = Mast Cell Progenitor\n>>* MkP = Megakaryocyte Progenitor\n>>* NeuP = Neutrophil Progenitor\n>>* MoP = Monocyte Progenitor\n>>* EryP = Erythrocyte Progenitor\n>>* HSC = Hematoploetic Stem Cell\n>>* BP = B-Cell ProgenitorMasP = Mast Cell Progenitor\n>>* MkP = Megakaryocyte Progenitor\n>>* NeuP = Neutrophil Progenitor\n>>* MoP = Monocyte Progenitor\n>>* EryP = Erythrocyte Progenitor\n>>* HSC = Hematoploetic Stem Cell\n>>* BP = B-Cell Progenitor\n    \n2. train_multi_inputs.h5\n3. test_multi_inputs.h5\n>* train/test_multi_inputs.h5 - ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF) output (chromatin accessibility), with rows corresponding to cells and columns corresponding to the location of the genome whose level of accessibility is measured, here identified by the genomic coordinates on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).\n>* train/test_multi_inputs.h5 - ATAC-seqのピーク数をデフォルトのlog(TF) * log(IDF)出力でTF-IDF変換したもの（クロマチンアクセス性）。行は細胞、列はアクセス性のレベルが測定されたゲノムの位置に対応し、ここでは10x References - 2020-A (July 7, 2020) で提供された参照ゲノムGRCh38のゲノム座標で特定されています。\n\n4. train_multi_targets.h5\n>* train_multi_labels.h5 - RNA gene expression levels as library-size normalized and log1p transformed counts for the same cells.\n>* train_multi_labels.h5 - RNA遺伝子の発現量は、同じ細胞のライブラリーサイズで正規化し、log1p変換したカウント値。\n\n5. train_cite_inputs.h5\n6. test_cite_inputs.h5\n>* train/test_cite_inputs.h5 - RNA library-size normalized and log1p transformed counts (gene expression levels), with rows corresponding to cells and columns corresponding to genes given by {gene_name}_{gene_ensemble-ids}.\n>* train/test_cite_inputs.h5 - RNAライブラリーサイズを正規化し、log1p変換したカウント（遺伝子発現量）。行は細胞、列は{gene_name}_{gene_ensemble-ids}で指定した遺伝子に対応する。\n7. train_cite_targets.h5\n>* train_cite_labels.h5 - Surface protein levels for the same cells that have been dsb normalized.\n>* train_cite_labels.h5 - dsbで正規化された同じ細胞の表面タンパク質レベル。\n\n8. evaluation_ids.csv\n>* Identifies the labels from the test set to be evaluated. It provides a join key from the cell_id / gene_id identifiers of the label matrix to the row_id needed for the submission file.\n>* 評価対象のテストセットからラベルを特定する。ラベルマトリックスのcell_id / gene_id 識別子から、提出ファイルに必要なrow_idへの結合キーを提供します。\n\n9. sample_submission.csv\n>* A sample submission file in the correct format. See the Evaluation page for more information.\n>* 正しい形式の投稿ファイルのサンプルです。詳しくは「評価」のページをご覧ください。\n\nwww.DeepL.com/Translator（無料版）で翻訳しました。","metadata":{}},{"cell_type":"markdown","source":"# Data Preview\n\nOnly the first few lines are read, as memory will be insufficient.\n\nメモリが不足するため、先頭の数行のみ読み込みます。","metadata":{}},{"cell_type":"code","source":"num_rows = 5","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:34.692196Z","iopub.execute_input":"2022-08-18T07:54:34.692744Z","iopub.status.idle":"2022-08-18T07:54:34.704506Z","shell.execute_reply.started":"2022-08-18T07:54:34.692693Z","shell.execute_reply":"2022-08-18T07:54:34.703500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:34.706530Z","iopub.execute_input":"2022-08-18T07:54:34.707928Z","iopub.status.idle":"2022-08-18T07:54:51.862805Z","shell.execute_reply.started":"2022-08-18T07:54:34.707874Z","shell.execute_reply":"2022-08-18T07:54:51.860934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nDATA_DIR = \"/kaggle/input/open-problems-multimodal/\"\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")  # 1\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")  # 2\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")  # 4\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")  # 3\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")  # 5\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")  # 7\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")  # 6\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")  # 8\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")  # 9","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:51.866392Z","iopub.execute_input":"2022-08-18T07:54:51.866859Z","iopub.status.idle":"2022-08-18T07:54:52.565614Z","shell.execute_reply.started":"2022-08-18T07:54:51.866819Z","shell.execute_reply":"2022-08-18T07:54:52.564438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_1_metadata = pd.read_csv(FP_CELL_METADATA, nrows=num_rows)\ndf_1_metadata.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:52.568224Z","iopub.execute_input":"2022-08-18T07:54:52.568662Z","iopub.status.idle":"2022-08-18T07:54:52.608578Z","shell.execute_reply.started":"2022-08-18T07:54:52.568625Z","shell.execute_reply":"2022-08-18T07:54:52.607326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_2_train_multi_inputs = pd.read_hdf(FP_MULTIOME_TRAIN_INPUTS, stop=num_rows)\ndf_2_train_multi_inputs.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:52.612275Z","iopub.execute_input":"2022-08-18T07:54:52.612656Z","iopub.status.idle":"2022-08-18T07:54:53.444097Z","shell.execute_reply.started":"2022-08-18T07:54:52.612623Z","shell.execute_reply":"2022-08-18T07:54:53.442954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_3_test_multi_inputs = pd.read_hdf(FP_MULTIOME_TEST_INPUTS, stop=num_rows)\ndf_3_test_multi_inputs.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:53.446149Z","iopub.execute_input":"2022-08-18T07:54:53.446629Z","iopub.status.idle":"2022-08-18T07:54:54.252322Z","shell.execute_reply.started":"2022-08-18T07:54:53.446574Z","shell.execute_reply":"2022-08-18T07:54:54.250903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_4_train_multi_targets = pd.read_hdf(FP_MULTIOME_TRAIN_TARGETS, stop=num_rows)\ndf_4_train_multi_targets.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.254024Z","iopub.execute_input":"2022-08-18T07:54:54.254898Z","iopub.status.idle":"2022-08-18T07:54:54.377980Z","shell.execute_reply.started":"2022-08-18T07:54:54.254861Z","shell.execute_reply":"2022-08-18T07:54:54.377020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_5_train_cite_inputs = pd.read_hdf(FP_CITE_TRAIN_INPUTS, stop=num_rows)\ndf_5_train_cite_inputs.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.379124Z","iopub.execute_input":"2022-08-18T07:54:54.379795Z","iopub.status.idle":"2022-08-18T07:54:54.509184Z","shell.execute_reply.started":"2022-08-18T07:54:54.379753Z","shell.execute_reply":"2022-08-18T07:54:54.507875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_6_test_cite_inputs = pd.read_hdf(FP_CITE_TEST_INPUTS, stop=num_rows)\ndf_6_test_cite_inputs.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.510683Z","iopub.execute_input":"2022-08-18T07:54:54.511088Z","iopub.status.idle":"2022-08-18T07:54:54.637022Z","shell.execute_reply.started":"2022-08-18T07:54:54.511052Z","shell.execute_reply":"2022-08-18T07:54:54.635571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_7_train_cite_targets = pd.read_hdf(FP_CITE_TRAIN_TARGETS, stop=num_rows)\ndf_7_train_cite_targets.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.640458Z","iopub.execute_input":"2022-08-18T07:54:54.642118Z","iopub.status.idle":"2022-08-18T07:54:54.699561Z","shell.execute_reply.started":"2022-08-18T07:54:54.642021Z","shell.execute_reply":"2022-08-18T07:54:54.698447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_8_sample_submission = pd.read_csv(FP_SUBMISSION, nrows=num_rows)\ndf_8_sample_submission.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.701033Z","iopub.execute_input":"2022-08-18T07:54:54.702219Z","iopub.status.idle":"2022-08-18T07:54:54.722208Z","shell.execute_reply.started":"2022-08-18T07:54:54.702168Z","shell.execute_reply":"2022-08-18T07:54:54.721249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_9_evaluation_ids = pd.read_csv(FP_EVALUATION_IDS, nrows=num_rows)\ndf_9_evaluation_ids.head(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T07:54:54.723753Z","iopub.execute_input":"2022-08-18T07:54:54.724186Z","iopub.status.idle":"2022-08-18T07:54:54.742769Z","shell.execute_reply.started":"2022-08-18T07:54:54.724146Z","shell.execute_reply":"2022-08-18T07:54:54.741686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Column Name Patterns\n### Check for patterns in column names by replacing numbers in column names with \"@\" symbols\n### カラム名の数字を「@」記号に置き換えることにより、カラム名のパターンを確認","metadata":{}},{"cell_type":"markdown","source":"For Multiome, the number of columns is very large, but the pattern is limited.\n\nMultiomeについて、カラム数は非常に多いが、パターンは限られている。","metadata":{}},{"cell_type":"code","source":"print(len(df_2_train_multi_inputs.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:21:07.196798Z","iopub.execute_input":"2022-08-18T08:21:07.197264Z","iopub.status.idle":"2022-08-18T08:21:07.202814Z","shell.execute_reply.started":"2022-08-18T08:21:07.197228Z","shell.execute_reply":"2022-08-18T08:21:07.201859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import re\ncols_replace_digit = []\nfor col in df_2_train_multi_inputs.columns:\n    cols_replace_digit.append(re.sub(r'\\d', \"@\", col))\nset(cols_replace_digit)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:19:18.650021Z","iopub.execute_input":"2022-08-18T08:19:18.650499Z","iopub.status.idle":"2022-08-18T08:19:19.759424Z","shell.execute_reply.started":"2022-08-18T08:19:18.650460Z","shell.execute_reply":"2022-08-18T08:19:19.758039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(df_3_test_multi_inputs.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:25:24.992884Z","iopub.execute_input":"2022-08-18T08:25:24.993375Z","iopub.status.idle":"2022-08-18T08:25:25.002349Z","shell.execute_reply.started":"2022-08-18T08:25:24.993337Z","shell.execute_reply":"2022-08-18T08:25:25.001072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import re\ncols_replace_digit = []\nfor col in df_3_test_multi_inputs.columns:\n    cols_replace_digit.append(re.sub(r'\\d', \"@\", col))\nset(cols_replace_digit)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:25:35.563810Z","iopub.execute_input":"2022-08-18T08:25:35.564415Z","iopub.status.idle":"2022-08-18T08:25:36.693289Z","shell.execute_reply.started":"2022-08-18T08:25:35.564364Z","shell.execute_reply":"2022-08-18T08:25:36.691941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(df_4_train_multi_targets.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:21:41.833033Z","iopub.execute_input":"2022-08-18T08:21:41.833470Z","iopub.status.idle":"2022-08-18T08:21:41.839228Z","shell.execute_reply.started":"2022-08-18T08:21:41.833433Z","shell.execute_reply":"2022-08-18T08:21:41.838243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_replace_digit = []\nfor col in df_4_train_multi_targets.columns:\n    cols_replace_digit.append(re.sub(r'\\d', \"@\", col))\nset(cols_replace_digit)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:22:08.234498Z","iopub.execute_input":"2022-08-18T08:22:08.234929Z","iopub.status.idle":"2022-08-18T08:22:08.327785Z","shell.execute_reply.started":"2022-08-18T08:22:08.234893Z","shell.execute_reply":"2022-08-18T08:22:08.326418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"CITEseq has many variations of column name patterns\n\nCITEseqは、カラム名のパターンのバリエーションが多い","metadata":{}},{"cell_type":"code","source":"print(len(df_5_train_cite_inputs.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:22:28.054927Z","iopub.execute_input":"2022-08-18T08:22:28.055350Z","iopub.status.idle":"2022-08-18T08:22:28.063664Z","shell.execute_reply.started":"2022-08-18T08:22:28.055318Z","shell.execute_reply":"2022-08-18T08:22:28.062059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_replace_digit = []\nfor col in df_5_train_cite_inputs.columns:\n    cols_replace_digit.append(re.sub(r'\\d', \"@\", col))\nset(cols_replace_digit)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:30:36.192602Z","iopub.execute_input":"2022-08-18T08:30:36.193035Z","iopub.status.idle":"2022-08-18T08:30:36.312054Z","shell.execute_reply.started":"2022-08-18T08:30:36.192999Z","shell.execute_reply":"2022-08-18T08:30:36.310812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(set(cols_replace_digit))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:30:36.908514Z","iopub.execute_input":"2022-08-18T08:30:36.909393Z","iopub.status.idle":"2022-08-18T08:30:36.918224Z","shell.execute_reply.started":"2022-08-18T08:30:36.909349Z","shell.execute_reply":"2022-08-18T08:30:36.916788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(df_7_train_cite_targets.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:26:45.307192Z","iopub.execute_input":"2022-08-18T08:26:45.307724Z","iopub.status.idle":"2022-08-18T08:26:45.318026Z","shell.execute_reply.started":"2022-08-18T08:26:45.307684Z","shell.execute_reply":"2022-08-18T08:26:45.315741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_replace_digit = []\nfor col in df_7_train_cite_targets.columns:\n    cols_replace_digit.append(re.sub(r'\\d', \"@\", col))\nset(cols_replace_digit)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T08:26:53.025507Z","iopub.execute_input":"2022-08-18T08:26:53.025920Z","iopub.status.idle":"2022-08-18T08:26:53.036230Z","shell.execute_reply.started":"2022-08-18T08:26:53.025885Z","shell.execute_reply":"2022-08-18T08:26:53.034634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ","metadata":{}}]}