{"metadata":{"kernelspec":{"display_name":"ida-env","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.8"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45533,"databundleVersionId":5748852,"sourceType":"competition"},{"sourceId":10135951,"sourceType":"datasetVersion","datasetId":6187360}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #0A6EBD; font-family: 'Trebuchet MS', Arial, sans-serif; color: white; padding: 20px; font-size: 40px; font-weight: bold; border-radius: 0 0 0 0; box-shadow: 0px 6px 8px rgba(0, 0, 0, 0.2);\">\n\tProject - Intelligent Data Analysis @ FIT-HCMUS, VNU-HCM 📌\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #5A96E3; font-family: 'Trebuchet MS', Arial, sans-serif; color: white; padding: 20px; font-size: 40px; font-weight: bold; border-radius: 0 0 0 0; box-shadow: 0px 6px 8px rgba(0, 0, 0, 0.2);\">\n\tPredict Student Performance from Game Play 📌\n</div>\n\n#\n---","metadata":{}},{"cell_type":"markdown","source":"**Sinh viên thực hiện:**\n\nHọ và tên: Vũ Minh Phát\n\nMSSV: 21127739\n\nLớp: Phân tích dữ liệu thông minh - 21KHDL\n\n**Giảng viên hướng dẫn**:\n- Nguyễn Tiến Huy\n- Nguyễn Trần Duy Minh\n- Lê Thanh Tùng\n\n#\n---","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #5A96E3; font-family: 'Trebuchet MS', Arial, sans-serif; color: white; padding: 20px; font-size: 40px; font-weight: bold; border-radius: 0 0 0 0; box-shadow: 0px 6px 8px rgba(0, 0, 0, 0.2);\">\n  Exploratory data analysis process\n</div>\n\n#\n---\n\nĐể biết thêm thông tin chi tiết về quy trình Phân tích Khám phá Dữ liệu (Exploratory Data Analysis - EDA) trên bộ dữ liệu của cuộc thi \"**Predict Student Performance from Game Play**\", vui lòng xem tại file notebook gốc [Project_Predict_Student_Performance_21127739.ipynb](Project_Predict_Student_Performance_21127739.ipynb) được đính kèm trong folder này.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #5A96E3; font-family: 'Trebuchet MS', Arial, sans-serif; color: white; padding: 20px; font-size: 40px; font-weight: bold; border-radius: 0 0 0 0; box-shadow: 0px 6px 8px rgba(0, 0, 0, 0.2);\">\n\tSetup and Import necessary Python modules 📚\n</div>\n\n#\n---","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os\nimport pickle\nfrom collections import Counter","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Lấy đường dẫn đến tất cả các file đầu vào trong thư mục \"/kaggle/input\"\nfile_paths = []\nfor dirname, _, filenames in os.walk(\"/kaggle/input\"):\n    for filename in filenames:\n        file_paths.append(os.path.join(dirname, filename))\n\n# Sắp xếp file theo thứ tự tên file\nfile_paths = sorted(file_paths)\n\n# Hiển thị tên các file\nfile_paths","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"text-align: center; background-color: #5A96E3; font-family: 'Trebuchet MS', Arial, sans-serif; color: white; padding: 20px; font-size: 40px; font-weight: bold; border-radius: 0 0 0 0; box-shadow: 0px 6px 8px rgba(0, 0, 0, 0.2);\">\n\tStage 5.4 - Testing data\n</div>\n\n#\n---","metadata":{}},{"cell_type":"markdown","source":"## 📌 Tổng quan\n\nFile notebook này được dùng để load các mô hình LightGBM đã được huấn luyện ở giai đoạn trước và sau đó sử dụng mô hình đã huấn luyện để dự đoán kết quả của người chơi trên tập dữ liệu kiểm tra thực sự.","metadata":{}},{"cell_type":"markdown","source":"## 📌 Quá trình chuẩn bị dữ liệu","metadata":{}},{"cell_type":"markdown","source":"Định nghĩa đường dẫn đến các tập tin chứa dữ liệu đã được xử lý trong các giai đoạn trước.","metadata":{}},{"cell_type":"code","source":"# Path of trained models\nMODEL_DIR = \"/kaggle/input/21127739-models/\"","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Định nghĩa lại lớp \"`FeatureMaker`\".","metadata":{}},{"cell_type":"code","source":"class FeatureMaker():\n    def __init__(self):\n        self.map_key = None\n        self.valid_keys = None\n        self.list_text_seq = None\n        self.map_text_seq = None\n        self.feature_names = []\n                \n    def prepare(self, df: pd.DataFrame, threshold: float):\n        # Concatenation of variables\n        keys = df[[\"level\", \"name\", \"event_name\", \"room_fqid\", \"fqid\", \"text\"]].values.tolist()\n        keys = [str(values[0]).zfill(2) + \"_\" + \"_\".join([\"None\" if type(v) != str else v for v in values[1:]])\n                for values in keys]\n        count_keys = Counter(keys)\n        valid_keys = [key for key, value in count_keys.items() if value >= threshold]\n        self.map_key = {key: i for i, key in enumerate(valid_keys)}\n        self.valid_keys = set(valid_keys)\n        \n        # Text sequence of important events (notification_click)\n        tmp_df = df.query(\"event_name == 'notification_click'\")[[\"session_id\", \"text\"]].fillna(\"\")\n        tmp_df[\"text_prev\"] = tmp_df.groupby(\"session_id\")[\"text\"].shift()\n        tmp_df = tmp_df.dropna()\n        agg_df = tmp_df.groupby([\"text_prev\", \"text\"], as_index=False).size().sort_values(\"size\", ascending=False)\n        self.list_text_seq = [(text1, text2) for text1, text2 in zip(agg_df[\"text_prev\"], agg_df[\"text\"])]\n        self.map_text_seq = {value: i for i, value in enumerate(self.list_text_seq)}\n        \n        self.feature_names += [f\"count_{key}\" for key in self.map_key.keys()]\n        self.feature_names += [f\"bdiff_{key}\" for key in self.map_key.keys()]\n        self.feature_names += [f\"fdiff_{key}\" for key in self.map_key.keys()]\n        self.feature_names += [f\"time_between_{text1}_and_{text2}\" for text1, text2 in self.list_text_seq]\n        self.feature_names += [\"last_time\", \"diff_level_group_0to1\", \"diff_level_group_1to2\"]\n    \n    def make_feature(self, features, session_df: pd.DataFrame, session_id, level_group):\n        keys = session_df[[\"level\", \"name\", \"event_name\", \"room_fqid\", \"fqid\", \"text\"]].values.tolist()\n        keys = [str(values[0]).zfill(2) + \"_\" + \"_\".join([\"None\" if type(v) != str else v for v in values[1:]])\n                for values in keys]\n        \n        values_time = session_df[\"elapsed_time\"].tolist()\n        values_event = session_df[\"event_name\"].tolist()\n        values_text = session_df[\"text\"].tolist()\n        \n        # Sort by index (dealing with api issue)\n        argsort = np.argsort(session_df[\"index\"].values).tolist()\n        keys = list(map(keys.__getitem__, argsort))\n        values_time = list(map(values_time.__getitem__, argsort))\n        values_event = list(map(values_event.__getitem__, argsort))\n        values_text = list(map(values_text.__getitem__, argsort))\n        \n        # Time diff between level_group\n        if level_group == 1:\n            features[-2] = values_time[0] - features[-3]\n        elif level_group == 2:\n            features[-1] = values_time[0] - features[-3]\n        # Last time\n        features[-3] = values_time[-1]\n\n        text_prev = \"\"\n        time_prev = 0\n        for i in range(len(session_df)):\n            if keys[i] in self.valid_keys:\n                # count\n                feature_idx = self.map_key[keys[i]]\n                if level_group <= 1:\n                    features[feature_idx] += 1\n                # bdiff\n                feature_idx += len(self.map_key)\n                if level_group <= 1 and i > 0:\n                    features[feature_idx] += values_time[i] - values_time[i-1]\n                # fdiff\n                feature_idx += len(self.map_key)\n                if i < len(session_df) - 1:\n                    features[feature_idx] += values_time[i+1] - values_time[i]\n            # Time between important events\n            if values_event[i] == \"notification_click\":\n                if (text_prev, values_text[i]) in self.map_text_seq:\n                    feature_idx = len(self.map_key)*3 + self.map_text_seq[(text_prev, values_text[i])]\n                    features[feature_idx] += values_time[i] - time_prev\n                text_prev = values_text[i]\n                time_prev = values_time[i]\n\n        return features","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Load dữ liệu của đối tượng \"`FeatureMaker`\" từ file pickle.","metadata":{}},{"cell_type":"code","source":"# Load initial feature maker\nwith open(MODEL_DIR + \"feature_maker.pickle\", \"rb\") as f:\n    fm = pickle.load(f)\nfeature_index_dict = {f: i for i, f in enumerate(fm.feature_names)}","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Load các mô hình đã được huấn luyện và các features quan trọng.","metadata":{}},{"cell_type":"code","source":"# Load models and important features\nmodels = []\nfeatures_idx = [[], [], []]\nvalid_keys = set()\nfor level_group in range(3):\n    # Load each model\n    with open(MODEL_DIR + f\"model{level_group}.pickle\", \"rb\") as f:\n        model = pickle.load(f)\n        \n    # Add each model to list of models\n    models.append(model)\n    \n    # Get number of model's features\n    n_features = len(model.feature_name())\n    \n    # Get the important features\n    importance_df = pd.read_csv(MODEL_DIR + f\"importance{level_group}.csv\")\n    importance_df = importance_df.groupby(\"feature\")[\"importance\"].mean().sort_values(ascending=False)\n    features = importance_df.index[1:n_features-1].tolist()\n    features_idx[level_group] = [feature_index_dict[f] for f in features]\n    \n    # Get valid keys\n    valid_keys |= set([f[6:] for f in features if f[:5] in [\"count\", \"bdiff\", \"fdiff\"]])","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Xác định danh sách các feature hợp lệ cho model.","metadata":{}},{"cell_type":"code","source":"# Set valid keys for model\nfm.valid_keys = valid_keys","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Khởi tạo mảng lưu trữ số câu hỏi của từng nhóm cấp độ và các câu hỏi tương ứng.","metadata":{}},{"cell_type":"code","source":"# Initialize arrays storing number of questions of each level group and corresponding questions\nn_questions = [3, 10, 5]\nquestions = [\n    np.array([1, 2, 3], dtype=np.float32),\n    np.array([4, 5, 6, 7, 8, 9, 10, 11, 12, 13], dtype=np.float32),\n    np.array([14, 15, 16, 17, 18], dtype=np.float32)\n]","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Thiết lập giá trị ngưỡng để mô hình ra quyết định cho kết quả dự đoán.","metadata":{}},{"cell_type":"code","source":"# Set threshold to 0.620\nthreshold = 0.620","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📌 Sử dụng mô hình đã huấn luyện để dự đoán kết quả trên tập dữ liệu kiểm tra thực sự","metadata":{}},{"cell_type":"markdown","source":"Lấy môi trường kiểm tra ẩn.","metadata":{}},{"cell_type":"code","source":"# Get hidden test environment\ntry:\n\timport jo_wilder\nexcept ModuleNotFoundError:\n\timport jo_wilder_310 as jo_wilder\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Bắt đầu quá trình dự đoán.","metadata":{}},{"cell_type":"code","source":"features_dict = dict()\nfor session_df, pred_df in iter_test:\n    # Get session id of the current test session dataframe\n    session_id = session_df.iloc[0, 0]\n    \n    # If length of pred_df is 3, 10 or 5 then assign level group to 0, 1 or 2, respectively\n    if len(pred_df) == 3:\n        level_group = 0\n    elif len(pred_df) == 10:\n        level_group = 1\n    elif len(pred_df) == 5:\n        level_group = 2\n    \n    if level_group == 0:\n        # Initialize value of `features_dict` at key `session_id`\n        features_dict[session_id] = [0] * len(fm.feature_names)\n    \n    # Update value of `features_dict` at key `session_id` \n    features_dict[session_id] = fm.make_feature(features_dict[session_id], session_df, session_id, level_group)\n    \n    # Get test set and use model to predict\n    test_X = list(map(features_dict[session_id].__getitem__, features_idx[level_group]))\n    test_X = np.array([0] + [value if value != 0 else np.nan for value in test_X] + [22], dtype=np.float32)  # append \"q\" and \"level_max\"\n    test_X = np.tile(test_X, (n_questions[level_group], 1))\n    test_X[:, 0] = questions[level_group]\n    preds = (models[level_group].predict(test_X) > threshold).astype(int)\n    \n    # Overwrite the session_id values (dealing with api issue)\n    pred_df[\"session_id\"] = [f\"{session_id}_q{int(q)}\" for q in questions[level_group]]\n    pred_df[\"correct\"] = preds\n    env.predict(pred_df)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📌 Kết quả dự đoán của mô hình","metadata":{}},{"cell_type":"markdown","source":"In một vài dòng đầu tiên của kết quả dự đoán được lưu trữ trong file \"`submission.csv`\".","metadata":{}},{"cell_type":"code","source":"# Print results in `submission.csv`\n! head submission.csv","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Tạo một biểu đồ cột đứng để thể hiện số lượng câu trả lời đúng và sai được dự đoán.","metadata":{}},{"cell_type":"code","source":"# Check the number of correct and incorrect answers\npred_df = pd.read_csv(\"submission.csv\")\npred_df[\"correct\"].value_counts()\n\nunique, count = np.unique(pred_df[\"correct\"], return_counts=True)\nunique = unique.astype(\"int\")\n\nfig, ax = plt.subplots(figsize=(5,5))\nbars = ax.bar([0, 0.25], count, color=[\"b\",\"c\"], width=0.1)\nax.set_xticks([0, 0.25], unique)\nax.bar_label(bars, fmt=\"{:.0f}\")\n\nax.set_xlabel(\"Correctness\")\nax.set_ylabel(\"Count\")\nax.set_title(\"Number of correct and incorrect answers\");","metadata":{},"outputs":[],"execution_count":null}]}