{"cells":[{"metadata":{},"cell_type":"markdown","source":"In this competition, many participants face difficulties in handling large size `train.csv`. We have to use Kaggle Notebook env and its RAM is 16GB.\n\nIn this notebook, I introduce `pandas.DataFrame.memory_usage()` and how to filter columns when loading."},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\n\nimport riiideducation","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# You can only call make_env() once, so don't lose it!\nenv = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv',\n                       low_memory=False,\n                       nrows=10**6,\n                       dtype={'row_id': 'int64',\n                              'timestamp': 'int64',\n                              'user_id': 'int32',\n                              'content_id': 'int16',\n                              'content_type_id': 'int8',\n                              'task_container_id': 'int16',\n                              'user_answer': 'int8',\n                              'answered_correctly': 'int8',\n                              'prior_question_elapsed_time': 'float32',\n                              'prior_question_had_explanation': 'boolean',\n                             }\n                      )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"At the end of the log, we can see that memory usage is 31.5MB.\n\n>memory usage: 31.5 MB"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see the results for each column by `momory_usage()`. `row_id` and `timestamp` use much more memory than others."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.memory_usage()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.memory_usage().plot.barh()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"If you only use specific columns like notebook I published, you can filter columns by `usecols`.\n\nhttps://www.kaggle.com/sishihara/riiid-answered-correctly-benchmark"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv',\n                       low_memory=False,\n                       nrows=10**6,\n                       usecols=['content_id', 'answered_correctly'],\n                       dtype={'content_id': 'int16', 'answered_correctly': 'int8'}\n                      )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The momory usage was reduced to 2.9M. This means that you can increase the size of `nrows`.\n\nAnd another options is use proper type in `dtype` which we can see from the results of `memory_usage()`."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}