{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Querying specific subsets of the data without loading all of it into memory","metadata":{"execution":{"iopub.execute_input":"2022-08-11T21:21:49.972474Z","iopub.status.busy":"2022-08-11T21:21:49.972077Z","iopub.status.idle":"2022-08-11T21:21:49.977181Z","shell.execute_reply":"2022-08-11T21:21:49.975948Z","shell.execute_reply.started":"2022-08-11T21:21:49.972443Z"}}},{"cell_type":"markdown","source":"I have implemented a `DataReader` class that allows to query a given h5 file. The user can, e.g., ask for the training multiome inputs of a specific donor in a specific day, without the need to load all of the `train_multi_inputs.h5` file into memory.","metadata":{}},{"cell_type":"code","source":"#If you see a urllib warning running this cell, go to \"Settings\" on the right hand side, \n#and turn on internet. Note, you need to be phone verified.\n!pip install --quiet tables\n!pip install hdf5plugin","metadata":{"execution":{"iopub.status.busy":"2022-08-23T19:59:47.653481Z","iopub.execute_input":"2022-08-23T19:59:47.654328Z","iopub.status.idle":"2022-08-23T20:00:06.065017Z","shell.execute_reply.started":"2022-08-23T19:59:47.654146Z","shell.execute_reply":"2022-08-23T20:00:06.063928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport h5py\nimport hdf5plugin\nimport tables\nimport psutil\nimport time\nimport gc","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:06.067137Z","iopub.execute_input":"2022-08-23T20:00:06.06758Z","iopub.status.idle":"2022-08-23T20:00:06.621112Z","shell.execute_reply.started":"2022-08-23T20:00:06.067541Z","shell.execute_reply":"2022-08-23T20:00:06.620281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/open-problems-multimodal\"\n\nFP_CELL_METADATA = os.path.join(DATA_DIR,\"metadata.csv\")\n\nFP_CITE_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_cite_inputs.h5\")\nFP_CITE_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_cite_targets.h5\")\nFP_CITE_TEST_INPUTS = os.path.join(DATA_DIR,\"test_cite_inputs.h5\")\n\nFP_MULTIOME_TRAIN_INPUTS = os.path.join(DATA_DIR,\"train_multi_inputs.h5\")\nFP_MULTIOME_TRAIN_TARGETS = os.path.join(DATA_DIR,\"train_multi_targets.h5\")\nFP_MULTIOME_TEST_INPUTS = os.path.join(DATA_DIR,\"test_multi_inputs.h5\")\n\nFP_SUBMISSION = os.path.join(DATA_DIR,\"sample_submission.csv\")\nFP_EVALUATION_IDS = os.path.join(DATA_DIR,\"evaluation_ids.csv\")","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2022-08-23T20:00:06.622303Z","iopub.execute_input":"2022-08-23T20:00:06.622581Z","iopub.status.idle":"2022-08-23T20:00:06.630458Z","shell.execute_reply.started":"2022-08-23T20:00:06.622555Z","shell.execute_reply":"2022-08-23T20:00:06.629338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I keep a copy of the metadata dataframe for convenience:","metadata":{}},{"cell_type":"code","source":"md = pd.read_csv(FP_CELL_METADATA)\nmd.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:06.631582Z","iopub.execute_input":"2022-08-23T20:00:06.631859Z","iopub.status.idle":"2022-08-23T20:00:06.852483Z","shell.execute_reply.started":"2022-08-23T20:00:06.631833Z","shell.execute_reply":"2022-08-23T20:00:06.851187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The one below is the class definition. The `query_data` method is the most important. It uses `h5py` library (upon suggestion by @alexandervc).","metadata":{}},{"cell_type":"code","source":"class DataReader:\n    \"\"\"\n    Class to access subsets of data in one of the *.h5 files, without loading all of the data into memory.\n    E.g., if you are working in a kaggle notebook, and want to read multiome training inputs, initialize a \n    data reader like:\n    dr = DataReader(data_dir = '/kaggle/input',\n                    filename = 'training_multi_inputs.h5',\n                    metadata_file_name = 'metadata.csv')\n                    \n    Use `DataReader.query_data` method to query the data.\n    \n    GOOD TO KNOW: if you want to query data from, say, donor 31800, you can use donor_id = 2. Example:\n    dr.query_data(donor_id = 2) --> returns all data available from donor = 31800. \n    The mapping between donor and donor_id is given below:\n    \n    donor | donor_id\n    ----------------\n    13176 | 1\n    31800 | 2\n    32606 | 3\n    27678 | 4\n    \n    Note: the donor_id identifier is set to 4 for donor 27678 (instead of 2) in order to be in agreement with\n    Data documentation provided by the organizers (donor 27678 is informally called \"Donor 4\" in the Data tab \n    of the competition's main page, and all around).\n    \n    \n    \"\"\"\n    \n    def __init__(self, data_dir, filename, metadata_file_name):\n        self.filename = filename\n        self.prefix = filename.replace('.h5','')\n        self.file_path = os.path.join(data_dir, filename)\n        self.md = pd.read_csv(os.path.join(data_dir, metadata_file_name))\n        self.add_donor_id(self.md)\n        \n    def add_donor_id(self, df):\n\n        donor_dict = {13176:1, 31800:2,32606:3,27678:4}\n        df['donor_id'] = df.donor.map(donor_dict)\n        \n        \n    def get_query_metadata(self, **kwargs):\n        \"\"\"\n        Filters metadata according to variable query args (kwargs).\n        \"\"\"\n        qmd = self.md.copy()\n        for arg in kwargs:\n            qmd = qmd.loc[qmd[arg] == kwargs[arg]].copy()\n            \n        return qmd\n    \n    @property\n    def shape(self):\n        with h5py.File(self.file_path,'r') as f:\n            shape = f[self.prefix]['block0_values'].shape\n        return shape\n    \n    def query_data(self, col_sample = None, **kwargs):\n        \"\"\"\n        This function uses h5py to query the file. It returns a dataframe (which might well be huge). \n        Depending on your query, you might still run out of RAM and your session crash.\n        \n        Inputs:\n            - col_sample: int | list of string's | list of int's\n                if col_sample is None (i.e., simply omitted), then returns all columns.\n                if int: a random sample of `col_sample` columns will be chosen for the output\n                if list of int's: these int's are interpreted as the desired columns positions.\n                if list of string's: these must be column names.\n                \n                Warning: If you ask for a number larger than the number of columns, or you choose \n                         (string) column names that are not present in the actual columns, this \n                         will break, and not in a nice way.\n                Remark: The output dataframe will have its columns ordered coherently with the actual\n                        column ordering in the files. E.g., if you pass in col_sample = [5000, 14],\n                        the first output column will be column nb. 14, and the second will be the 5000-th\n                        columns.\n                        \n        kwargs: additional query args. These args must be chosen among the column names of the\n                metadata.csv file.\n                NOTE: there are four donors with four 5-digit codes. If you want to get data from \n                donor 13176, but you don't remember all these 5-digit codes, you can filter by donor_id.\n                The donor <--> donor_id mapping is as follows:\n                \n                donor | donor_id\n                ----------------\n                13176 | 1\n                31800 | 2\n                32606 | 3\n                27678 | 4                \n                \n                \n        Output:\n            - out: pandas.DataFrame with the desired rows (filtered according to kwargs) and columns.\n            \n        Examples:\n        self.query_data(col_sample = 400, day = 2, donor_id = 1)\n        Returns all rows for donor_id = 1 (i.e., donor = 13176), for day = 2, and 400 randomly sampled columns.\n        self.query_data(col_sample = [3, 5, 1], day = 3, donor_id = 2, cell_type = 'BP')\n        Returns all rows for donor_id = 2 (i.e., donor = 31800), for day = 3, and the second, fourth and sixth columns\n        (in that order).\n        \n        \"\"\"\n        \n        f = h5py.File(self.file_path, 'r')\n        \n        query_df = self.get_query_metadata(**kwargs)\n\n        rows = f[self.prefix]['axis1'][:]\n\n        rows = list(map(lambda x: x.decode(), rows))\n\n        rows = pd.DataFrame(rows, columns = ['cell_id']).reset_index().\\\n        rename(columns = {'index':'index_col'})\n\n        rows = pd.merge(query_df[['cell_id']]\n                      ,rows\n                      ,on = 'cell_id').sort_values(by = 'index_col')\n        \n        if col_sample == None: \n            N_cols = f[self.prefix]['axis0'].shape[0]\n            col_sample_idx = np.arange(N_cols)\n            columns = [col.decode() for col in f[self.prefix]['axis0'][:]]\n        elif type(col_sample) == int: \n            N_cols = f[self.prefix]['axis0'].shape[0]\n            col_sample_idx = np.random.choice(N_cols, size = col_sample, replace = False)\n            col_sample_idx = sorted(list(col_sample_idx))\n            columns = [col.decode() for col in f[self.prefix]['axis0'][col_sample_idx]]\n        elif type(col_sample) == list and type(col_sample[0]) == int:\n            col_sample_idx = sorted(col_sample)\n            columns = [col.decode() for col in f[self.prefix]['axis0'][col_sample_idx]]\n        elif type(col_sample) == list and type(col_sample[0]) == str:\n            all_cols = np.array([col.decode() for col in f[self.prefix]['axis0'][:]])\n            all_cols_sorted = np.sort(all_cols)\n            all_cols_sorted_idx = np.argsort(all_cols)\n\n            col_sample_pseudo_idx = np.searchsorted(all_cols_sorted, col_sample)\n            col_sample_idx = all_cols_sorted_idx[col_sample_pseudo_idx]\n            col_sample_idx = np.sort(col_sample_idx)\n            columns = all_cols[col_sample_idx]\n            \n        out = f[self.prefix]['block0_values'][rows.index_col.values,:][:,col_sample_idx]\n        \n        out = pd.DataFrame(out, index = rows.cell_id, columns = columns)\n        \n        f.close()\n\n        return out\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:06.854813Z","iopub.execute_input":"2022-08-23T20:00:06.855353Z","iopub.status.idle":"2022-08-23T20:00:06.877225Z","shell.execute_reply.started":"2022-08-23T20:00:06.855317Z","shell.execute_reply":"2022-08-23T20:00:06.876261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### A couple of tests\nLet's first check how much RAM is under use:","metadata":{}},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:06.878304Z","iopub.execute_input":"2022-08-23T20:00:06.87948Z","iopub.status.idle":"2022-08-23T20:00:06.896324Z","shell.execute_reply.started":"2022-08-23T20:00:06.879444Z","shell.execute_reply":"2022-08-23T20:00:06.895229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Set up a DataReader for Training Multiome inputs and targets.","metadata":{}},{"cell_type":"code","source":"dr_tr_mi = DataReader(data_dir = DATA_DIR,\n                      filename = 'train_multi_inputs.h5',\n                      metadata_file_name = 'metadata.csv')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:08.872808Z","iopub.execute_input":"2022-08-23T20:00:08.873806Z","iopub.status.idle":"2022-08-23T20:00:09.070958Z","shell.execute_reply.started":"2022-08-23T20:00:08.87377Z","shell.execute_reply":"2022-08-23T20:00:09.07Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dr_tr_mt = DataReader(data_dir = DATA_DIR,\n                      filename = 'train_multi_targets.h5',\n                      metadata_file_name = 'metadata.csv')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:12.183472Z","iopub.execute_input":"2022-08-23T20:00:12.183903Z","iopub.status.idle":"2022-08-23T20:00:12.379926Z","shell.execute_reply.started":"2022-08-23T20:00:12.183866Z","shell.execute_reply":"2022-08-23T20:00:12.37887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Multiome training input data has {dr_tr_mi.shape[0]} rows (cells) and {dr_tr_mi.shape[1]} columns (i.e., loci in the DNA, as far as I understand).\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:00:14.569628Z","iopub.execute_input":"2022-08-23T20:00:14.57024Z","iopub.status.idle":"2022-08-23T20:00:14.653618Z","shell.execute_reply.started":"2022-08-23T20:00:14.570207Z","shell.execute_reply":"2022-08-23T20:00:14.652744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Multiome training target data has {dr_tr_mt.shape[0]} rows (cells) and {dr_tr_mt.shape[1]} columns (i.e., mRNA transcriptions, as far as I understand).\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:02:51.977967Z","iopub.execute_input":"2022-08-23T20:02:51.97835Z","iopub.status.idle":"2022-08-23T20:02:52.021431Z","shell.execute_reply.started":"2022-08-23T20:02:51.978321Z","shell.execute_reply":"2022-08-23T20:02:52.020584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's query the data for donor `13176`, for day nb. 7:","metadata":{}},{"cell_type":"code","source":"%%time\ninputs = dr_tr_mi.query_data(day = 7, donor = 13176)\nprint(f\"Shape: {inputs.shape}.\")\ninputs.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:03:05.739544Z","iopub.execute_input":"2022-08-23T20:03:05.739932Z","iopub.status.idle":"2022-08-23T20:03:43.478124Z","shell.execute_reply.started":"2022-08-23T20:03:05.739901Z","shell.execute_reply":"2022-08-23T20:03:43.477144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let' choose the same cells, but just a random sample of the columns, of size 1000:","metadata":{}},{"cell_type":"code","source":"%%time\n# donor_id = 1 corresponds to donor 13176. See the docs of the DataReader class\ninputs = dr_tr_mi.query_data(day = 7, donor_id = 1, col_sample = 1000) \nprint(f\"Shape: {inputs.shape}.\")\ninputs.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:03:52.831399Z","iopub.execute_input":"2022-08-23T20:03:52.83183Z","iopub.status.idle":"2022-08-23T20:04:19.187609Z","shell.execute_reply.started":"2022-08-23T20:03:52.831794Z","shell.execute_reply":"2022-08-23T20:04:19.186664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:04:19.188926Z","iopub.execute_input":"2022-08-23T20:04:19.190005Z","iopub.status.idle":"2022-08-23T20:04:19.196144Z","shell.execute_reply.started":"2022-08-23T20:04:19.189978Z","shell.execute_reply":"2022-08-23T20:04:19.195173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's read Multiome training *targets*, with the same query parameters:","metadata":{}},{"cell_type":"code","source":"%%time\ntargets = dr_tr_mt.query_data(day = 7, donor = 13176)\nprint(f\"Shape: {targets.shape}.\")\ntargets.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:04:31.939464Z","iopub.execute_input":"2022-08-23T20:04:31.939873Z","iopub.status.idle":"2022-08-23T20:04:41.098109Z","shell.execute_reply.started":"2022-08-23T20:04:31.939838Z","shell.execute_reply":"2022-08-23T20:04:41.097055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:04:54.324751Z","iopub.execute_input":"2022-08-23T20:04:54.325259Z","iopub.status.idle":"2022-08-23T20:04:54.330508Z","shell.execute_reply.started":"2022-08-23T20:04:54.325219Z","shell.execute_reply":"2022-08-23T20:04:54.329753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Inputs' and targets' rows (cells) are ordered in the same way. See below.**\n\nThe rows of `DataReader.query_data`'s output are ordered in the same way as in the file from which they were read.","metadata":{}},{"cell_type":"code","source":"# cell_id's are the same row-wise: \n(inputs.index == targets.index).sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:29.793065Z","iopub.execute_input":"2022-08-23T20:05:29.794018Z","iopub.status.idle":"2022-08-23T20:05:29.801346Z","shell.execute_reply.started":"2022-08-23T20:05:29.79398Z","shell.execute_reply":"2022-08-23T20:05:29.799949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice that the number of rows (cells) in the metadata corresponding to the query parameters is the same as the number of rows in the above results:","metadata":{}},{"cell_type":"code","source":"md[(md.day == 7) &\n  (md.donor == 13176) &\n  (md.technology == 'multiome')].shape","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:44.834613Z","iopub.execute_input":"2022-08-23T20:05:44.835027Z","iopub.status.idle":"2022-08-23T20:05:44.864046Z","shell.execute_reply.started":"2022-08-23T20:05:44.834991Z","shell.execute_reply":"2022-08-23T20:05:44.863013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del inputs, targets\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:46.217929Z","iopub.execute_input":"2022-08-23T20:05:46.218687Z","iopub.status.idle":"2022-08-23T20:05:46.350292Z","shell.execute_reply.started":"2022-08-23T20:05:46.218646Z","shell.execute_reply":"2022-08-23T20:05:46.349246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's try with another file and another donor and day:","metadata":{}},{"cell_type":"code","source":"dr_tr_ci = DataReader(data_dir = DATA_DIR,\n                      filename = 'train_cite_inputs.h5',\n                      metadata_file_name = 'metadata.csv')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:52.124206Z","iopub.execute_input":"2022-08-23T20:05:52.124586Z","iopub.status.idle":"2022-08-23T20:05:52.33199Z","shell.execute_reply.started":"2022-08-23T20:05:52.124557Z","shell.execute_reply":"2022-08-23T20:05:52.330997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dr_tr_ct = DataReader(data_dir = DATA_DIR,\n                      filename = 'train_cite_targets.h5',\n                      metadata_file_name = 'metadata.csv')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:53.743399Z","iopub.execute_input":"2022-08-23T20:05:53.74426Z","iopub.status.idle":"2022-08-23T20:05:53.94302Z","shell.execute_reply.started":"2022-08-23T20:05:53.744225Z","shell.execute_reply":"2022-08-23T20:05:53.942212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's query the data for donor `32606`, for day nb. 4:","metadata":{}},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:05:58.136202Z","iopub.execute_input":"2022-08-23T20:05:58.136564Z","iopub.status.idle":"2022-08-23T20:05:58.143195Z","shell.execute_reply.started":"2022-08-23T20:05:58.136535Z","shell.execute_reply":"2022-08-23T20:05:58.141674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ninputs = dr_tr_ci.query_data(day = 4, donor_id = 3)\nprint(f\"Shape: {inputs.shape}.\")\ninputs.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:01.045592Z","iopub.execute_input":"2022-08-23T20:06:01.045987Z","iopub.status.idle":"2022-08-23T20:06:13.902841Z","shell.execute_reply.started":"2022-08-23T20:06:01.045954Z","shell.execute_reply":"2022-08-23T20:06:13.901847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:13.904343Z","iopub.execute_input":"2022-08-23T20:06:13.904611Z","iopub.status.idle":"2022-08-23T20:06:13.909322Z","shell.execute_reply.started":"2022-08-23T20:06:13.904586Z","shell.execute_reply":"2022-08-23T20:06:13.90855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntargets = dr_tr_ct.query_data(day = 4, donor_id = 3)\nprint(f\"Shape: {targets.shape}.\")\ntargets.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:13.910328Z","iopub.execute_input":"2022-08-23T20:06:13.91057Z","iopub.status.idle":"2022-08-23T20:06:15.020139Z","shell.execute_reply.started":"2022-08-23T20:06:13.910546Z","shell.execute_reply":"2022-08-23T20:06:15.019352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:15.021906Z","iopub.execute_input":"2022-08-23T20:06:15.02244Z","iopub.status.idle":"2022-08-23T20:06:15.027916Z","shell.execute_reply.started":"2022-08-23T20:06:15.022404Z","shell.execute_reply":"2022-08-23T20:06:15.027184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As before, the cell_id's are the same for each pair of rows:","metadata":{}},{"cell_type":"code","source":"# cell_id's are the same row-wise: \n(inputs.index == targets.index).sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:15.029421Z","iopub.execute_input":"2022-08-23T20:06:15.030271Z","iopub.status.idle":"2022-08-23T20:06:15.040827Z","shell.execute_reply.started":"2022-08-23T20:06:15.030217Z","shell.execute_reply":"2022-08-23T20:06:15.04012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**However, if your query returns too many rows, you might run out of RAM:**","metadata":{}},{"cell_type":"code","source":"# I try to free up some RAM. I'm not a Computer Scientist: this is supposed to work, sometimes it does, sometimes it doesn't:\ndel inputs, targets\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:19.309048Z","iopub.execute_input":"2022-08-23T20:06:19.30979Z","iopub.status.idle":"2022-08-23T20:06:19.433281Z","shell.execute_reply.started":"2022-08-23T20:06:19.309753Z","shell.execute_reply":"2022-08-23T20:06:19.432352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_gbs = round(0.01 * psutil.virtual_memory().percent * 16, 1)\nprint(f\"{used_gbs} Gb are being used.\")","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:06:20.605576Z","iopub.execute_input":"2022-08-23T20:06:20.606271Z","iopub.status.idle":"2022-08-23T20:06:20.611164Z","shell.execute_reply.started":"2022-08-23T20:06:20.606238Z","shell.execute_reply":"2022-08-23T20:06:20.610491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The cell below caused my session to crash, after having run the cells above. Even after restarting the session and not running the above queries, attempting to read all of donor 13176 multiome data into RAM causes a session crash. In the following days, I might try to adjust the `query_data` method to read the file in chunks, and convert those chunks in sparse matrices.**","metadata":{}},{"cell_type":"code","source":"f = %%time\ninputs = dr_tr_mi.query_data(donor = 13176)\nprint(f\"Shape: {inputs.shape}.\")\ninputs.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}