{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-09T11:43:03.527598Z","iopub.execute_input":"2022-09-09T11:43:03.528132Z","iopub.status.idle":"2022-09-09T11:43:03.564242Z","shell.execute_reply.started":"2022-09-09T11:43:03.528027Z","shell.execute_reply":"2022-09-09T11:43:03.563136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = r'/kaggle/input/amex-default-prediction/test_data.csv'","metadata":{"execution":{"iopub.status.busy":"2022-09-09T11:43:03.566006Z","iopub.execute_input":"2022-09-09T11:43:03.566922Z","iopub.status.idle":"2022-09-09T11:43:03.571365Z","shell.execute_reply.started":"2022-09-09T11:43:03.566883Z","shell.execute_reply":"2022-09-09T11:43:03.570408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport gc\nimport os\nimport glob","metadata":{"execution":{"iopub.status.busy":"2022-09-09T11:43:03.572898Z","iopub.execute_input":"2022-09-09T11:43:03.573563Z","iopub.status.idle":"2022-09-09T11:43:03.586082Z","shell.execute_reply.started":"2022-09-09T11:43:03.573529Z","shell.execute_reply":"2022-09-09T11:43:03.584502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Split and Save the Big Dataset into Chunks","metadata":{}},{"cell_type":"markdown","source":"- This code only for the explanation purpose and I am restricting the long iterations\n- Here, we are restricting upto 5 iteration due to the memory limit in Kaggle","metadata":{}},{"cell_type":"code","source":"# Let's load dataset by every 5,00,000 rows\nchunk_size=500000\nnum=1\nfor chunk in pd.read_csv(test_data,chunksize=chunk_size):\n    \n    # This is to break the loop after 5 iteration to avoid memory error in kaggle\n    # In local desktop yo can remove this if conditions\n    if num > 5:\n        break\n        \n    chunk.to_csv('chunk'+str(num)+'.csv',index=False)\n    gc.collect()\n    num+=1","metadata":{"execution":{"iopub.status.busy":"2022-09-09T11:43:03.589567Z","iopub.execute_input":"2022-09-09T11:43:03.590611Z","iopub.status.idle":"2022-09-09T11:56:53.939896Z","shell.execute_reply.started":"2022-09-09T11:43:03.590560Z","shell.execute_reply":"2022-09-09T11:56:53.938680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load the splitted chunk file and Optimize the Memory then concatenate","metadata":{}},{"cell_type":"code","source":"path = r'./'\nall_files = glob.glob(os.path.join(path, \"*.csv\"))\n\ngrouped_files = []\n\n\nfor filename in all_files:\n    chunk = pd.read_csv(filename, index_col=None, header=0)\n    for col in chunk.columns:\n        \n        # Let's convert all float64 features into float16 feature & int64 features into int 8 feature\n        if chunk[col].dtype == 'float64':\n            chunk[col] = chunk[col].astype('float16')\n        if chunk[col].dtype == 'int64':\n            chunk[col] = chunk[col].astype('int8')\n            \n            \n        \n    gc.collect()\n    grouped_files.append(chunk)\n\n    df = pd.concat(grouped_files, axis=0, ignore_index=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-09T11:56:53.942551Z","iopub.execute_input":"2022-09-09T11:56:53.943310Z","iopub.status.idle":"2022-09-09T12:01:54.037349Z","shell.execute_reply.started":"2022-09-09T11:56:53.943270Z","shell.execute_reply":"2022-09-09T12:01:54.035979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b style='color:blue'>NOTE</b>: Remember, this code only for explanation purpose, still there are other memory optimizatio steps available such as converting the object columns to str or categories.","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-09T12:01:54.039404Z","iopub.execute_input":"2022-09-09T12:01:54.040232Z","iopub.status.idle":"2022-09-09T12:01:54.086812Z","shell.execute_reply.started":"2022-09-09T12:01:54.040178Z","shell.execute_reply":"2022-09-09T12:01:54.085646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Now, we can load all the 25,00,000 rows in a dataframe","metadata":{}},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-09T12:01:54.088428Z","iopub.execute_input":"2022-09-09T12:01:54.088861Z","iopub.status.idle":"2022-09-09T12:01:54.098913Z","shell.execute_reply.started":"2022-09-09T12:01:54.088826Z","shell.execute_reply":"2022-09-09T12:01:54.097371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## converting to feather format\n\nThere are many formats available to store big data, here feather is one of my favorite for lazy load","metadata":{}},{"cell_type":"code","source":"df.to_feather(\"optimised_amex.feather\")","metadata":{"execution":{"iopub.status.busy":"2022-09-09T12:01:54.128058Z","iopub.execute_input":"2022-09-09T12:01:54.128421Z","iopub.status.idle":"2022-09-09T12:01:57.769764Z","shell.execute_reply.started":"2022-09-09T12:01:54.128386Z","shell.execute_reply":"2022-09-09T12:01:57.768438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = pd.read_feather('optimised_amex.feather')\ndf1.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-09T12:01:57.773018Z","iopub.execute_input":"2022-09-09T12:01:57.773415Z","iopub.status.idle":"2022-09-09T12:01:59.882196Z","shell.execute_reply.started":"2022-09-09T12:01:57.773381Z","shell.execute_reply":"2022-09-09T12:01:59.880945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h4>Summary</h4>\n\nWhat Next! Even though we have loaded the data in feather format, we could not achieve all the functionalities like Pandas DataFrame.However, there are other tools such as dask and data table will do the magic. Let me discuss those methods on upcoming notes.","metadata":{}},{"cell_type":"markdown","source":"If you like to download complete working code, please check out my GitHub repo.\n<a href='https://github.com/kumarnarun/Amex-Default-Prediciton'>Github repo</a>","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}