{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Efficient dataframe loading with Datatable\n\nFirst of all, lets start by quoting what [Sohier Dane](https://www.kaggle.com/sohier) said about the training dataset in the competition's starter [notebook](https://www.kaggle.com/sohier/competition-api-detailed-introduction).\n\n> It's larger than will fit in memory with default settings, so we'll specify more efficient datatypes and only load a subset of the data for now.\n\nAfter that, an instruction is given on how to efficiently load the dataset using specific data types. However, if you try to load the entire dataset that way using pandas, your RAM memmory limit will be likely reached.\n\nInspired by [Vopani](https://www.kaggle.com/rohanrao)'s excelent [notebook](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid), we'll see how to load heavy .csv data using the [**Python datatable**](https://datatable.readthedocs.io/en/latest/index.html) package."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Install the datatable package.\n!pip install datatable","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport datatable as dt\nimport gc\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Loading with datatable\n\nLoading .csv data and converting it to a Pandas Dataframe with Datatable is straightforward."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"%%time\ntrain_df = dt.fread(\"/kaggle/input/riiid-test-answer-prediction/train.csv\").to_pandas()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see thath the entire dataset was loaded in hoghly 44 seconds, which is nice given the dataset size.\n\nNow, let's check the dataframe's information."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see, **datatable** has automatically infered some columns types, in contrast with the rather conservative Pandas's data loading.\n\nThe entire dataset fits nicely in 4.6 GB without effort. But, as stated in the starter notebook, the data types can be further tweaked in order to improve memmory consumption."},{"metadata":{"trusted":true},"cell_type":"code","source":"dtype={\n    'row_id': np.int64, 'timestamp': np.int64, 'user_id': np.int32, 'content_id': np.int16, 'content_type_id': np.int8,\n    'task_container_id': np.int16, 'user_answer': np.int8, 'answered_correctly': np.int8, 'prior_question_elapsed_time': np.float32, \n    'prior_question_had_explanation': np.bool,\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in dtype.keys():\n    train_df[col] = train_df[col].astype(dtype[col])\ntrain_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, we've gained approximately 1.6 GB of extra RAM to be used in something really useful."},{"metadata":{},"cell_type":"markdown","source":"## Saving in binary format\n\nAs an additional step, well check the benefits of saving the processed data into a binary format.\n\nDatatable uses .jay format, which makes reading our dataset a breeze."},{"metadata":{"trusted":true},"cell_type":"code","source":"dt.Frame(train_df).to_jay(\"train_df.jay\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del train_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"gc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, we'll load our data back."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntrain_df = dt.fread(\"train_df.jay\").to_pandas()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Cool! The entire dataset was loaded in amzing 4.84 seconds (~ 4 times faster). As a bonus, our data types were preserved."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That's all folks!\n\nThis is the first of a series of short notebook that I'm planning to make. The goal is to build the critical phases of an end-to-end project, step by step. So, stay tuned for the other kernels.\n\nHave you found something useful? Please, give a an upvote!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}