{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-05-26T05:29:52.333021Z","iopub.execute_input":"2022-05-26T05:29:52.333890Z","iopub.status.idle":"2022-05-26T05:29:52.342214Z","shell.execute_reply.started":"2022-05-26T05:29:52.333843Z","shell.execute_reply":"2022-05-26T05:29:52.341189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## load data using Pandas","metadata":{}},{"cell_type":"code","source":"import json\nimport time","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:29:53.888246Z","iopub.execute_input":"2022-05-26T05:29:53.889059Z","iopub.status.idle":"2022-05-26T05:29:53.893106Z","shell.execute_reply.started":"2022-05-26T05:29:53.889003Z","shell.execute_reply":"2022-05-26T05:29:53.892343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nDo Not Run the CODE:- Pandas is not able to load this much size of data\n\nError i got:-\nYour notebook tried to allocate more memory than is available. It has restarted.\n\"\"\"\n\n# start = time.time()\n# df = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\")\n# end = time.time()\n# print(f\"{end - start}\")\n\n","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:29:54.948032Z","iopub.execute_input":"2022-05-26T05:29:54.948838Z","iopub.status.idle":"2022-05-26T05:29:54.955670Z","shell.execute_reply.started":"2022-05-26T05:29:54.948794Z","shell.execute_reply":"2022-05-26T05:29:54.954589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## load data using DASK","metadata":{}},{"cell_type":"markdown","source":"**What is DASK?**\n\nDask is a parallel computing library that scales the existing Python ecosystem. This tutorial will introduce Dask and parallel data analysis more generally.\n\n\nDask provides multi-core and distributed parallel execution on larger-than-memory datasets.\n\n- **High-level collections:** Dask provides high-level Array, Bag, and DataFrame collections that mimic NumPy, lists, and Pandas but can operate in parallel on datasets that don’t fit into memory. Dask’s high-level collections are alternatives to NumPy and Pandas for large datasets.\n\n- **Low-level schedulers:** Dask provides dynamic task schedulers that execute task graphs in parallel. These execution engines power the high-level collections mentioned above but can also power custom, user-defined workloads. These schedulers are low-latency (around 1ms) and work hard to run computations in a small memory footprint. Dask’s schedulers are an alternative to direct use of threading or multiprocessing libraries in complex cases or other task scheduling systems like Luigi or IPython parallel.","metadata":{}},{"cell_type":"code","source":"!pip install dask","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:29:59.288267Z","iopub.execute_input":"2022-05-26T05:29:59.288742Z","iopub.status.idle":"2022-05-26T05:30:12.281077Z","shell.execute_reply.started":"2022-05-26T05:29:59.288698Z","shell.execute_reply":"2022-05-26T05:30:12.279822Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### loading Train data","metadata":{}},{"cell_type":"code","source":"from dask import dataframe as dd\n\nstart = time.time()\ndask_df = dd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\")\nend = time.time()\nprint(f\"Total time Disk take to read the data is :- {end - start} sec\")","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:30:12.283629Z","iopub.execute_input":"2022-05-26T05:30:12.284403Z","iopub.status.idle":"2022-05-26T05:30:13.593452Z","shell.execute_reply.started":"2022-05-26T05:30:12.284331Z","shell.execute_reply":"2022-05-26T05:30:13.592597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndask_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:40:55.907817Z","iopub.execute_input":"2022-05-26T05:40:55.908278Z","iopub.status.idle":"2022-05-26T05:40:57.031388Z","shell.execute_reply.started":"2022-05-26T05:40:55.908240Z","shell.execute_reply":"2022-05-26T05:40:57.030237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndask_df.tail()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:41:00.987970Z","iopub.execute_input":"2022-05-26T05:41:00.988445Z","iopub.status.idle":"2022-05-26T05:41:02.022698Z","shell.execute_reply.started":"2022-05-26T05:41:00.988406Z","shell.execute_reply":"2022-05-26T05:41:02.021539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dask_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:39:42.272963Z","iopub.execute_input":"2022-05-26T05:39:42.273451Z","iopub.status.idle":"2022-05-26T05:39:42.300805Z","shell.execute_reply.started":"2022-05-26T05:39:42.273416Z","shell.execute_reply":"2022-05-26T05:39:42.299696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dask_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:40:18.427454Z","iopub.execute_input":"2022-05-26T05:40:18.427899Z","iopub.status.idle":"2022-05-26T05:40:18.435332Z","shell.execute_reply.started":"2022-05-26T05:40:18.427862Z","shell.execute_reply":"2022-05-26T05:40:18.434324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndask_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:35:51.807934Z","iopub.execute_input":"2022-05-26T05:35:51.808907Z","iopub.status.idle":"2022-05-26T05:35:53.115664Z","shell.execute_reply.started":"2022-05-26T05:35:51.808862Z","shell.execute_reply":"2022-05-26T05:35:53.114657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### loading Test data","metadata":{}},{"cell_type":"code","source":"%%time\n\nstart = time.time()\ntest_df = dd.read_csv(\"/kaggle/input/amex-default-prediction/test_data.csv\")\nend = time.time()\nprint(f\"Total time Disk take to read the data is :- {end - start} sec\")","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:42:54.928685Z","iopub.execute_input":"2022-05-26T05:42:54.929324Z","iopub.status.idle":"2022-05-26T05:42:54.997451Z","shell.execute_reply.started":"2022-05-26T05:42:54.929282Z","shell.execute_reply":"2022-05-26T05:42:54.996285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T05:43:03.757910Z","iopub.execute_input":"2022-05-26T05:43:03.758703Z","iopub.status.idle":"2022-05-26T05:43:05.982373Z","shell.execute_reply.started":"2022-05-26T05:43:03.758663Z","shell.execute_reply":"2022-05-26T05:43:05.981350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### NEXT:- I'm going to work on the prediction model stay tuned for that....","metadata":{},"execution_count":null,"outputs":[]}]}