{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45533,"databundleVersionId":5748852,"sourceType":"competition"}],"dockerImageVersionId":30615,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Predict Student Performance from Game Play - Save Data**\n\n### **Ho Chi Minh City University of Science**\n\n### **Faculty of Information Technology**\n\n#### **Class: 20KHDL**\n\n#### **Lecturers:**\n- **Dr. Nguyễn Tiến Huy**\n- **Mr. Nguyễn Trần Duy Minh**","metadata":{}},{"cell_type":"markdown","source":"# **Special thank**","metadata":{}},{"cell_type":"markdown","source":"- First of all, we would like to thank to [**Jack (Japan)**](https://www.kaggle.com/rsakata) for his notebooks relating to the competition.\n- For further information about [**Jack (Japan)'s**](https://www.kaggle.com/rsakata) works, please refer to [this discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119).","metadata":{}},{"cell_type":"markdown","source":"# **EDA process**","metadata":{}},{"cell_type":"markdown","source":"> For more information about the EDA process, please refer to the original notebook: https://www.kaggle.com/code/goldencheem/hcmus-fit-2023-dynamind","metadata":{}},{"cell_type":"markdown","source":"# **Outline**","metadata":{}},{"cell_type":"markdown","source":"This file is used to read the original `train.csv` file, separate it into 3 chunks and then save them into parquet files in order to reduce the memory storage. Furthermore, this file reads `train_label.csv` file, then extracts and adds some extra information, and finally saves it into parquet file.","metadata":{}},{"cell_type":"markdown","source":"# Libraries used","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-13T09:18:53.690485Z","iopub.execute_input":"2023-12-13T09:18:53.691054Z","iopub.status.idle":"2023-12-13T09:18:53.698748Z","shell.execute_reply.started":"2023-12-13T09:18:53.691014Z","shell.execute_reply":"2023-12-13T09:18:53.697748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes = {\n    'session_id': np.int64, \n    'elapsed_time': np.int64,\n    'event_name': str,\n    'name': str,\n    'level': np.int8,\n    'page': str,\n    'room_coor_x': np.float32,\n    'room_coor_y': np.float32,\n    'screen_coor_x': np.float32,\n    'screen_coor_y': np.float32,\n    'hover_duration': np.float32,\n    'text': str,\n    'fqid': str,\n    'room_fqid': str,\n    'text_fqid': str,\n    'fullscreen': np.int8,\n    'hq': np.int8,\n    'music': np.int8,\n    'level_group': str\n}","metadata":{"execution":{"iopub.status.busy":"2023-12-13T09:18:53.701145Z","iopub.execute_input":"2023-12-13T09:18:53.701934Z","iopub.status.idle":"2023-12-13T09:18:53.711487Z","shell.execute_reply.started":"2023-12-13T09:18:53.701893Z","shell.execute_reply":"2023-12-13T09:18:53.710450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Read the original data into 3 chunks with size of 10^7. For each chunk, convert level_group from string to integer and then save it to parquet file.","metadata":{}},{"cell_type":"code","source":"reader = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\", dtype=dtypes, chunksize=1e7)\nfor i, df in enumerate(reader):\n    df[\"level_group\"] = df[\"level_group\"].map({\"0-4\": 0, \"5-12\": 1, \"13-22\": 2}).astype(np.int8)\n    df.to_parquet(f\"train_{i}.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-12-13T09:18:53.713565Z","iopub.execute_input":"2023-12-13T09:18:53.714329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Create a question mapping dictionary. Questions 1-3, 5-13, 14-18 will be corresponding to level group 0 (0-4), 1 (5-12) and 2 (13-22), respectively.","metadata":{}},{"cell_type":"code","source":"q_map = {\n    1: 0,\n    2: 0,\n    3: 0,\n    4: 1,\n    5: 1,\n    6: 1,\n    7: 1,\n    8: 1,\n    9: 1,\n    10: 1,\n    11: 1,\n    12: 1,\n    13: 1,\n    14: 2,\n    15: 2,\n    16: 2,\n    17: 2,\n    18: 2,\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- First, read `train_labels.csv` file. Then create `q` and `level_group` columns which is questions and their corresponding level groups. Finally, extract session_ids from the original values of this column and save data into parquet file.","metadata":{}},{"cell_type":"code","source":"df_label = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\ndf_label[\"q\"] = df_label[\"session_id\"].apply(lambda x: int(x.split(\"_\")[1][1:])).astype(np.int8)\ndf_label[\"level_group\"] = df_label[\"q\"].map(q_map).astype(np.int8)\ndf_label[\"session_id\"] = df_label[\"session_id\"].apply(lambda x: x.split(\"_\")[0]).astype(np.int64)\ndf_label = df_label[[\"session_id\", \"level_group\", \"q\", \"correct\"]].sort_values([\"session_id\", \"q\"]).reset_index(drop=True)\ndf_label.to_parquet(\"labels.parquet\")\ndisplay(df_label)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}