{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-23T07:03:09.337495Z","iopub.execute_input":"2023-03-23T07:03:09.337963Z","iopub.status.idle":"2023-03-23T07:03:09.376483Z","shell.execute_reply.started":"2023-03-23T07:03:09.337911Z","shell.execute_reply":"2023-03-23T07:03:09.375301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](https://docs.dask.org/en/stable/_images/dask-dataframe.svg)","metadata":{}},{"cell_type":"markdown","source":"# Reading the Training-Data with the help of **Dask library**\n- **Panda** is great library but it is **not** always `computationally efficient`, especially when there are **GBs** of **data to manipulate**.So what can you do to get around this obstacle?\n\n- This is where **Dask weaves** its magic ! It works with pandas dataframes and Numpy data structures to help you perform **data wrangling** and model building **using large datasets** on not-so-powerful machines.\n\n- When it comes to working with large datasets using these `python libraries(numpy, pandas, sklearn)`, the `run time` can become very high due to **memory constraints**.\n\n- Dask can efficiently **perform parallel computations** on a single machine using multi-core CPUs. For example, if you have a quad core processor, Dask can effectively use all 4 cores of your system simultaneously for processing. In order to use lesser memory during computations, Dask stores the complete data on the disk, and uses chunks of data (smaller parts, rather than the whole data) from the disk for processing. During the processing, the intermediate values generated (if any) are discarded as soon as possible, to save the memory consumption.\n\n- It took around **52-min** to **reduce training data** from **4GB to 1GB** by changing data types of data.","metadata":{}},{"cell_type":"code","source":"! pip install \"dask[complete]\"\n","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-03-23T07:03:13.607888Z","iopub.execute_input":"2023-03-23T07:03:13.608372Z","iopub.status.idle":"2023-03-23T07:03:27.892134Z","shell.execute_reply.started":"2023-03-23T07:03:13.608330Z","shell.execute_reply":"2023-03-23T07:03:27.890782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import dask.dataframe as dd","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:03:27.895968Z","iopub.execute_input":"2023-03-23T07:03:27.896414Z","iopub.status.idle":"2023-03-23T07:03:29.213947Z","shell.execute_reply.started":"2023-03-23T07:03:27.896371Z","shell.execute_reply":"2023-03-23T07:03:29.212558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf = dd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:03:29.216108Z","iopub.execute_input":"2023-03-23T07:03:29.216977Z","iopub.status.idle":"2023-03-23T07:03:29.284497Z","shell.execute_reply.started":"2023-03-23T07:03:29.216922Z","shell.execute_reply":"2023-03-23T07:03:29.282755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:03:44.991289Z","iopub.execute_input":"2023-03-23T07:03:44.991753Z","iopub.status.idle":"2023-03-23T07:03:46.828265Z","shell.execute_reply.started":"2023-03-23T07:03:44.991709Z","shell.execute_reply":"2023-03-23T07:03:46.827180Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.visualize()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:03:52.080285Z","iopub.execute_input":"2023-03-23T07:03:52.081291Z","iopub.status.idle":"2023-03-23T07:03:52.446511Z","shell.execute_reply.started":"2023-03-23T07:03:52.081218Z","shell.execute_reply":"2023-03-23T07:03:52.442789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \ndf.shape[0].compute()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:04:01.519687Z","iopub.execute_input":"2023-03-23T07:04:01.520141Z","iopub.status.idle":"2023-03-23T07:05:16.771753Z","shell.execute_reply.started":"2023-03-23T07:04:01.520101Z","shell.execute_reply":"2023-03-23T07:05:16.770652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlen(df)","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:05:16.773773Z","iopub.execute_input":"2023-03-23T07:05:16.777588Z","iopub.status.idle":"2023-03-23T07:05:36.896471Z","shell.execute_reply.started":"2023-03-23T07:05:16.777529Z","shell.execute_reply":"2023-03-23T07:05:36.895336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isnull().sum().compute()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:06:01.601126Z","iopub.execute_input":"2023-03-23T07:06:01.601615Z","iopub.status.idle":"2023-03-23T07:07:05.563700Z","shell.execute_reply.started":"2023-03-23T07:06:01.601573Z","shell.execute_reply":"2023-03-23T07:07:05.562533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:07:05.569275Z","iopub.execute_input":"2023-03-23T07:07:05.569723Z","iopub.status.idle":"2023-03-23T07:07:05.582347Z","shell.execute_reply.started":"2023-03-23T07:07:05.569685Z","shell.execute_reply":"2023-03-23T07:07:05.581293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npercent_missing = (df.isnull().sum().compute() * 100) / len(df)","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:07:05.584051Z","iopub.execute_input":"2023-03-23T07:07:05.584964Z","iopub.status.idle":"2023-03-23T07:08:30.577988Z","shell.execute_reply.started":"2023-03-23T07:07:05.584915Z","shell.execute_reply":"2023-03-23T07:08:30.576772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"percent_missing","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:08:30.583195Z","iopub.execute_input":"2023-03-23T07:08:30.583682Z","iopub.status.idle":"2023-03-23T07:08:30.599595Z","shell.execute_reply.started":"2023-03-23T07:08:30.583638Z","shell.execute_reply":"2023-03-23T07:08:30.598299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:08:30.601128Z","iopub.execute_input":"2023-03-23T07:08:30.602351Z","iopub.status.idle":"2023-03-23T07:08:30.611675Z","shell.execute_reply.started":"2023-03-23T07:08:30.602298Z","shell.execute_reply":"2023-03-23T07:08:30.610504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    start_mem = df.memory_usage().sum().compute() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        \n        if col_type != object:\n            c_min = df[col].min().compute()\n            c_max = df[col].max().compute()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            df[col] = df[col].astype('category')\n    end_mem = df.memory_usage().sum().compute() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:08:30.613324Z","iopub.execute_input":"2023-03-23T07:08:30.613784Z","iopub.status.idle":"2023-03-23T07:08:30.631061Z","shell.execute_reply.started":"2023-03-23T07:08:30.613747Z","shell.execute_reply":"2023-03-23T07:08:30.630013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf =reduce_mem_usage(df)","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:08:30.632367Z","iopub.execute_input":"2023-03-23T07:08:30.633530Z","iopub.status.idle":"2023-03-23T07:42:09.867080Z","shell.execute_reply.started":"2023-03-23T07:08:30.633481Z","shell.execute_reply":"2023-03-23T07:42:09.866080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.memory_usage().sum().compute()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:43:36.707562Z","iopub.execute_input":"2023-03-23T07:43:36.708225Z","iopub.status.idle":"2023-03-23T07:45:09.561611Z","shell.execute_reply.started":"2023-03-23T07:43:36.708148Z","shell.execute_reply":"2023-03-23T07:45:09.560607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:53:52.799651Z","iopub.execute_input":"2023-03-23T07:53:52.800107Z","iopub.status.idle":"2023-03-23T07:53:52.811019Z","shell.execute_reply.started":"2023-03-23T07:53:52.800070Z","shell.execute_reply":"2023-03-23T07:53:52.809912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:54:15.470351Z","iopub.execute_input":"2023-03-23T07:54:15.470819Z","iopub.status.idle":"2023-03-23T07:54:17.474227Z","shell.execute_reply.started":"2023-03-23T07:54:15.470779Z","shell.execute_reply":"2023-03-23T07:54:17.473099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"percent_missing = (df.isnull().sum().compute() * 100) / len(df)\npercent_missing","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:50:04.057847Z","iopub.execute_input":"2023-03-23T07:50:04.058513Z","iopub.status.idle":"2023-03-23T07:53:08.962129Z","shell.execute_reply.started":"2023-03-23T07:50:04.058445Z","shell.execute_reply":"2023-03-23T07:53:08.961021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}