{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Lower Memory Footprint via Python Sparse Matrix","metadata":{"_uuid":"3991f95f-75cc-4b46-8251-81376b065a7b","_cell_guid":"2090cd09-d905-455d-ade5-f3580379f988","jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"It has been a month since [the Kaggle competition on Open Problem Multimodal](https://www.kaggle.com/competitions/open-problems-multimodal) started and there have been a few implementations on using sparse matrix to save RAM, e.g. [this one by SBUZINI](https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices) and [this one by Fabiean Crom](https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data/). It is becoming increasingly clear that sparse matrix is critical in solving this large-scale problem as dense matrix format will cost more than 100 GB in memory and few computers have this capacity. \n\nThese existing approaches resort to the underlying sparse matrix data structure, i.e., triplets of (row, column, value) or (row, row_pointer, value), to convert dense matrices to sparse counterpart. While technically correct, it is preferred to insulate the end users from such technical details if at all possible so that the end users will have easier time to understand the implementation, which in turn improves transparency and adoption.\n\nIn this code, we present a python implementation of [our previous work in R](https://www.kaggle.com/code/stautxie/reduce-memory-footprint-of-competition-data). This approach takes a divide-and-conquer approach in that chunk of rows from the large matrix is converted into sparse matrix, which are then concatenated into a full sparse matrix.","metadata":{}},{"cell_type":"markdown","source":"UPDATE: As [the raw data](https://www.kaggle.com/datasets/ryanholbrook/open-problems-raw-counts) has been released, we will revise the notebook to generate data for these raw data set as well so that we can test if these raw data file makes difference to the model quality.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nfrom scipy.sparse import csr_matrix, csc_matrix, vstack, save_npz, load_npz\nimport gc","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.279475Z","iopub.execute_input":"2022-10-14T04:14:37.280800Z","iopub.status.idle":"2022-10-14T04:14:37.286913Z","shell.execute_reply.started":"2022-10-14T04:14:37.280754Z","shell.execute_reply":"2022-10-14T04:14:37.286012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    import tables as tb\nexcept:\n    !pip install tables\n    import tables as tb","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.288640Z","iopub.execute_input":"2022-10-14T04:14:37.289514Z","iopub.status.idle":"2022-10-14T04:14:37.300852Z","shell.execute_reply.started":"2022-10-14T04:14:37.289483Z","shell.execute_reply":"2022-10-14T04:14:37.300163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DAT_DIR = \"../input/open-problems-multimodal\"\nRAW_DIR = \"../input/open-problems-raw-counts\"\nOUT_DIR = \".\"","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.302388Z","iopub.execute_input":"2022-10-14T04:14:37.302742Z","iopub.status.idle":"2022-10-14T04:14:37.319853Z","shell.execute_reply.started":"2022-10-14T04:14:37.302713Z","shell.execute_reply":"2022-10-14T04:14:37.317560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This function mat_h5_to_sparse() if the workhorse of this script. We are feeding a handle to the dense matrix stored in the h5 file as well as a parameter for number of rows to read in a time.  The function iteratively calculates the starting and ending rows for each iteration and converts the dense submatrix into a sparse matrix (compressed sparse row matrix in scipy.sparse module) through a single function call of csr_matrix(). Finally, we stack these sparse submatrix vertically to return the final sparse matrix again through a single call of numpy's vstack function.","metadata":{}},{"cell_type":"code","source":"def mat_h5_to_sparse(h5_array, chunk_size):\n    num_chunks = h5_array.shape[0] // chunk_size\n    \n    sp_matrix_list = []\n    for i in range(num_chunks+1):\n        row_start = i * chunk_size\n        row_end = (i+1) * chunk_size \n        print(f'Iteration = {i}, row_start = {row_start}, row_end = {row_end}')\n        mat = h5_array[row_start:row_end, :]\n        sp_mat = csr_matrix(mat)\n        del mat; gc.collect()\n        sp_matrix_list.append(sp_mat)\n        del sp_mat; gc.collect()\n        \n    mat_sp = vstack(sp_matrix_list)\n    del sp_matrix_list; gc.collect()  \n    \n    return mat_sp\n","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.322366Z","iopub.execute_input":"2022-10-14T04:14:37.324193Z","iopub.status.idle":"2022-10-14T04:14:37.333717Z","shell.execute_reply.started":"2022-10-14T04:14:37.324130Z","shell.execute_reply":"2022-10-14T04:14:37.332583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We also define the following function.","metadata":{}},{"cell_type":"code","source":"def gen_npz(input_data_stem, input_dir, output_dir):\n    \"\"\"read input data from h5 file and save the result into npz file in the output.\"\"\"\n    inputs_h5 = tb.open_file(os.path.join(input_dir, f\"{input_data_stem}.h5\"), \"r\")\n\n    inputs_block0_values = inputs_h5.root[f\"{input_data_stem}\"].block0_values\n    print(f'inputs_block0_values.shape = {inputs_block0_values}')\n\n    inputs_sp = mat_h5_to_sparse(h5_array = inputs_block0_values, chunk_size = 5000)\n    # pickle.dump(tr_mu_inputs_sp, open(os.path.join(OUT_DIR, \"tr_mu_inputs_sp.pkl\"), \"wb\"))\n    save_npz(os.path.join(OUT_DIR, f\"{input_data_stem}_sp.npz\"), inputs_sp)\n    del inputs_sp; gc.collect()\n    del inputs_block0_values; gc.collect()\n    del inputs_h5; gc.collect()  \n    return","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.335777Z","iopub.execute_input":"2022-10-14T04:14:37.336741Z","iopub.status.idle":"2022-10-14T04:14:37.354137Z","shell.execute_reply.started":"2022-10-14T04:14:37.336709Z","shell.execute_reply":"2022-10-14T04:14:37.352087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_input_data_stems = [\"test_cite_inputs\", \"test_multi_inputs\", \n                        \"train_cite_inputs\", \"train_multi_inputs\", \n                        \"train_cite_targets\", \"train_multi_targets\"]\nall_raw_data_stems = [c + \"_raw\" for c in all_input_data_stems]\nprint(f'all_input_data_stems = {all_input_data_stems}')\nprint(f'all_raw_data_stems = {all_raw_data_stems}')","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:15:18.061710Z","iopub.execute_input":"2022-10-14T04:15:18.062131Z","iopub.status.idle":"2022-10-14T04:15:18.069248Z","shell.execute_reply.started":"2022-10-14T04:15:18.062102Z","shell.execute_reply":"2022-10-14T04:15:18.068468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generate sparse matrix files for the original competition data files","metadata":{}},{"cell_type":"code","source":"for data_stem in all_input_data_stems:\n    gen_npz(input_data_stem = data_stem, input_dir = DAT_DIR, output_dir=OUT_DIR)","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.375511Z","iopub.execute_input":"2022-10-14T04:14:37.375948Z","iopub.status.idle":"2022-10-14T04:14:37.421311Z","shell.execute_reply.started":"2022-10-14T04:14:37.375912Z","shell.execute_reply":"2022-10-14T04:14:37.418747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generate sparse matrix files for the raw data files.","metadata":{}},{"cell_type":"code","source":"for data_stem in all_raw_data_stems:\n    gen_npz(input_data_stem = data_stem, input_dir = RAW_DIR, output_dir=OUT_DIR)","metadata":{"execution":{"iopub.status.busy":"2022-10-14T04:14:37.423647Z","iopub.status.idle":"2022-10-14T04:14:37.424174Z","shell.execute_reply.started":"2022-10-14T04:14:37.423952Z","shell.execute_reply":"2022-10-14T04:14:37.423973Z"},"trusted":true},"execution_count":null,"outputs":[]}]}