{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Multimodal Single-Cell Integration: Creating a Sparse Matrix Dataset","metadata":{}},{"cell_type":"markdown","source":"This notebook goal is to generate a more efficient version of the dataset of the competition \"Open Problems: Multimodal Single-Cell Integration\", using scipy's sparse matrices.\n\nUsing sparse matrices will lead to:\n- smaller file sizes\n- faster loading\n- much smaller memory footprint (so that you can actually keep all data in memory)\n\n(Especially for the multiome data, which is very big and very sparse; not so much for the citeseq data which is smaller and only ~75% sparse).\n\nThe downside is that we cannot use the nice pandas DataFrame anymore. Instead, each \"*xxx.h5*\" file is converted into two files:\n- One \"*xxx_values.sparse*\" file that can be loaded with `scipy.sparse.load_npz` and contains all the values of the corresponding dataframe (i.e. the result of `df.values` in a sparse format)\n- One \"*xxx_idxcol.npz*\" file that can be loaded with `np.load` and contains the values of the index and the columns of the corresponding dataframe (i.e the results of `df.index` and `df.columns`)\n\nFor convenience, the csv files are also converted into a more efficient parquet version.\n\nThe generated dataset is available [here](https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix).","metadata":{}},{"cell_type":"code","source":"!conda install pytables -y","metadata":{"execution":{"iopub.status.busy":"2022-08-27T00:48:04.241238Z","iopub.execute_input":"2022-08-27T00:48:04.242673Z","iopub.status.idle":"2022-08-27T00:49:11.665804Z","shell.execute_reply.started":"2022-08-27T00:48:04.242624Z","shell.execute_reply":"2022-08-27T00:49:11.664489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport scipy.sparse","metadata":{"execution":{"iopub.status.busy":"2022-08-27T00:50:24.922027Z","iopub.execute_input":"2022-08-27T00:50:24.922684Z","iopub.status.idle":"2022-08-27T00:50:24.929579Z","shell.execute_reply.started":"2022-08-27T00:50:24.922632Z","shell.execute_reply":"2022-08-27T00:50:24.928066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conversion Functions","metadata":{}},{"cell_type":"code","source":"def convert_to_parquet(filename, out_filename):\n    df = pd.read_csv(filename)\n    df.to_parquet(out_filename + \".parquet\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import scipy\ndef convert_h5_to_sparse_csr(filename, out_filename, chunksize=2500):\n    start = 0\n    total_rows = 0\n\n    sparse_chunks_data_list = []\n    chunks_index_list = []\n    columns_name = None\n    while True:\n        df_chunk = pd.read_hdf(filename, start=start, stop=start+chunksize)\n        if len(df_chunk) == 0:\n            break\n        chunk_data_as_sparse = scipy.sparse.csr_matrix(df_chunk.to_numpy())\n        sparse_chunks_data_list.append(chunk_data_as_sparse)\n        chunks_index_list.append(df_chunk.index.to_numpy())\n\n        if columns_name is None:\n            columns_name = df_chunk.columns.to_numpy()\n        else:\n            assert np.all(columns_name == df_chunk.columns.to_numpy())\n\n        total_rows += len(df_chunk)\n        print(total_rows)\n        if len(df_chunk) < chunksize: \n            del df_chunk\n            break\n        del df_chunk\n        start += chunksize\n        \n    all_data_sparse = scipy.sparse.vstack(sparse_chunks_data_list)\n    del sparse_chunks_data_list\n    \n    all_indices = np.hstack(chunks_index_list)\n    \n    scipy.sparse.save_npz(out_filename+\"_values.sparse\", all_data_sparse)\n    np.savez(out_filename+\"_idxcol.npz\", index=all_indices, columns =columns_name)\n    \n    \n","metadata":{"execution":{"iopub.status.busy":"2022-08-27T00:35:30.991692Z","iopub.execute_input":"2022-08-27T00:35:30.992275Z","iopub.status.idle":"2022-08-27T00:35:31.004102Z","shell.execute_reply.started":"2022-08-27T00:35:30.992227Z","shell.execute_reply":"2022-08-27T00:35:31.003112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# H5 -> Sparse Conversion","metadata":{}},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/train_multi_targets.h5\", \"train_multi_targets\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/train_multi_inputs.h5\", \"train_multi_inputs\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/train_cite_targets.h5\", \"train_cite_targets\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/train_cite_inputs.h5\", \"train_cite_inputs\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/test_multi_inputs.h5\", \"test_multi_inputs\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_h5_to_sparse_csr(\"../input/open-problems-multimodal/test_cite_inputs.h5\", \"test_cite_inputs\")","metadata":{"execution":{"iopub.status.busy":"2022-08-27T00:36:07.103674Z","iopub.execute_input":"2022-08-27T00:36:07.104119Z","iopub.status.idle":"2022-08-27T00:39:51.734351Z","shell.execute_reply.started":"2022-08-27T00:36:07.104084Z","shell.execute_reply":"2022-08-27T00:39:51.732947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CSV -> PARQUET conversion\nFor convenience, let us also convert the other files from CSV to the more efficient parquet format.\n\n(Then just replace pd.read_csv(xxx.csv) by pd.read_parquet(xxx.parquet) to read the file into a pandas DataFrame)","metadata":{}},{"cell_type":"code","source":"convert_to_parquet(\"../input/open-problems-multimodal/metadata.csv\", \"metadata\")","metadata":{"execution":{"iopub.status.busy":"2022-08-26T23:14:58.866585Z","iopub.execute_input":"2022-08-26T23:14:58.867085Z","iopub.status.idle":"2022-08-26T23:14:58.873794Z","shell.execute_reply.started":"2022-08-26T23:14:58.867036Z","shell.execute_reply":"2022-08-26T23:14:58.87243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_to_parquet(\"../input/open-problems-multimodal/evaluation_ids.csv\", \"evaluation\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"convert_to_parquet(\"../input/open-problems-multimodal/sample_submission.csv\", \"sample_submission\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -lh","metadata":{"execution":{"iopub.status.busy":"2022-08-27T00:39:51.737483Z","iopub.execute_input":"2022-08-27T00:39:51.738052Z","iopub.status.idle":"2022-08-27T00:39:52.975702Z","shell.execute_reply.started":"2022-08-27T00:39:51.737997Z","shell.execute_reply":"2022-08-27T00:39:52.974278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}