{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45533,"databundleVersionId":5748852,"sourceType":"competition"}],"dockerImageVersionId":30458,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Predict Student Performance from Game Play**\n\n### **Vietnam National University HCMC**\n\n### **University of Science**\n\n### **Faculty of Information Technology**\n\n#### **Course: CSC17001 - Intelligent Data Analysis**\n\n#### **Lecturers:**\n- **PhD. Nguyễn Tiến Huy**\n- **PhD. Nguyễn Trần Duy Minh**\n- **PhD. Lê Thanh Tùng**\n\n#### **Athours:**\n- **21127731 - Nguyễn Trọng Tín**","metadata":{}},{"cell_type":"markdown","source":"# **Acknowledgment of References:**","metadata":{}},{"cell_type":"markdown","source":"- This notebook is the final project of my Inteliigent Data Analysis course at HCMUS\n- Special thanks to [**Gusthema**](https://www.kaggle.com/gusthema), which provided the inspiration and foundation for the ideas and work that helped complete this project.\n- Please refer to [**this notebook**](https://www.kaggle.com/code/gusthema/student-performance-w-tensorflow-decision-forests) if you need further information about the Data Modeling & Experiments.\n- For my work, I will focus on data preprocessing, data visualization, and the data discovery process, aiming to enhance understanding and insights about the model and its behavior. Feel free to reach out to me for more information.","metadata":{}},{"cell_type":"markdown","source":"# **General Introduction**","metadata":{}},{"cell_type":"markdown","source":"# Import the Required Libraries","metadata":{"id":"zAXHC6-Tn2O5"}},{"cell_type":"code","source":"import tensorflow as tf\nimport tensorflow_addons as tfa\nimport tensorflow_decision_forests as tfdf\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"id":"IanlX-Eqn2O5","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:54:34.378037Z","iopub.execute_input":"2024-12-07T18:54:34.378502Z","iopub.status.idle":"2024-12-07T18:54:45.483670Z","shell.execute_reply.started":"2024-12-07T18:54:34.378465Z","shell.execute_reply":"2024-12-07T18:54:45.482377Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Version of Tensorflow Library\nprint(\"TensorFlow Decision Forests v\" + tfdf.__version__)\nprint(\"TensorFlow Addons v\" + tfa.__version__)\nprint(\"TensorFlow v\" + tf.__version__)","metadata":{"id":"gLpK2yAen2O7","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:54:45.488001Z","iopub.execute_input":"2024-12-07T18:54:45.488649Z","iopub.status.idle":"2024-12-07T18:54:45.495220Z","shell.execute_reply.started":"2024-12-07T18:54:45.488613Z","shell.execute_reply":"2024-12-07T18:54:45.494187Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Load the Dataset","metadata":{}},{"cell_type":"markdown","source":"To efficiently work with large datasets, we use a `dictionary` (dtypes) to explicitly define the data types for each column when loading the dataset. This approach helps in several ways:\n\n- Memory Optimization: By downcasting numerical columns to smaller types (e.g., int32, float32), we reduce memory consumption compared to the default int64 and float64 types, which are larger than necessary for many columns.\n\n- Faster Loading: Specifying data types prevents Pandas from having to infer them, which speeds up the reading process.\n\n- Efficient Handling of Categorical Data: Columns with repetitive string values (like `names` or `categories`) are better stored as category types, which saves memory by internally encoding strings as numeric codes.","metadata":{"id":"22DpLVFLn2O7"}},{"cell_type":"code","source":"dtypes={\n    'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n    'text':'category',\n    'fqid':'category',\n    'room_fqid':'category',\n    'text_fqid':'category',\n    'fullscreen':'category',\n    'hq':'category',\n    'music':'category',\n    'level_group':'category'}\n\ndataset_df = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', dtype=dtypes)\nprint(\"Train dataset shape: {}\".format(dataset_df.shape))","metadata":{"id":"_XItl24kn2O7","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:54:45.496210Z","iopub.execute_input":"2024-12-07T18:54:45.496515Z","iopub.status.idle":"2024-12-07T18:56:42.346717Z","shell.execute_reply.started":"2024-12-07T18:54:45.496485Z","shell.execute_reply":"2024-12-07T18:56:42.345621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display the first 5 examples\ndataset_df.head(5)","metadata":{"id":"-RTRVRiWn2O8","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:42.348717Z","iopub.execute_input":"2024-12-07T18:56:42.349040Z","iopub.status.idle":"2024-12-07T18:56:42.383793Z","shell.execute_reply.started":"2024-12-07T18:56:42.349001Z","shell.execute_reply":"2024-12-07T18:56:42.382468Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Data Discovery**","metadata":{}},{"cell_type":"code","source":"# 1. Memory Usage and Dataset Size\ndef mem_usage(pandas_obj):\n    if isinstance(pandas_obj, pd.DataFrame):\n        usage_bytes = pandas_obj.memory_usage(deep=True).sum()\n    else:  # we assume it's a Series\n        usage_bytes = pandas_obj.memory_usage(deep=True)\n    usage_mb = usage_bytes / 1024 ** 2  # convert to megabytes\n    return f\"{usage_mb:.2f} MB\"\n\nprint(f\"Dataset Memory Usage: {mem_usage(dataset_df)}\")\nprint(f\"Total Rows: {len(dataset_df)}\")\nprint(f\"Total Columns: {len(dataset_df.columns)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:42.385064Z","iopub.execute_input":"2024-12-07T18:56:42.385435Z","iopub.status.idle":"2024-12-07T18:56:42.402530Z","shell.execute_reply.started":"2024-12-07T18:56:42.385401Z","shell.execute_reply":"2024-12-07T18:56:42.401345Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 2. Categorical Variable Distribution\ndef plot_categorical_distribution(df, column, top_n=10):\n    plt.figure(figsize=(6, 3))\n    df[column].value_counts().head(top_n).plot(kind='bar')\n    plt.title(f'Top {top_n} {column} Distribution')\n    plt.xlabel(column)\n    plt.ylabel('Count')\n    plt.xticks(rotation=45)\n    plt.tight_layout()\n    plt.show()\n\ncategorical_columns = ['event_name', 'name', 'level_group', 'fullscreen', 'hq', 'music']\nfor col in categorical_columns:\n    print(f\"\\n{col} Distribution:\")\n    print(dataset_df[col].value_counts(normalize=True))\n    plot_categorical_distribution(dataset_df, col)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:42.403857Z","iopub.execute_input":"2024-12-07T18:56:42.404162Z","iopub.status.idle":"2024-12-07T18:56:44.963207Z","shell.execute_reply.started":"2024-12-07T18:56:42.404134Z","shell.execute_reply":"2024-12-07T18:56:44.962341Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Temporal Analysis\nprint(\"\\nTime-based Analysis:\")\ndataset_df['elapsed_time_minutes'] = dataset_df['elapsed_time'] / 60000  # Convert ms to minutes\ntime_by_level = dataset_df.groupby('level_group')['elapsed_time_minutes'].agg(['mean', 'median', 'max'])\nprint(time_by_level)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:44.964550Z","iopub.execute_input":"2024-12-07T18:56:44.965425Z","iopub.status.idle":"2024-12-07T18:56:46.168231Z","shell.execute_reply.started":"2024-12-07T18:56:44.965386Z","shell.execute_reply":"2024-12-07T18:56:46.167159Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Correlation of Coordinate Features\ncoordinate_columns = ['room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y']\ncorrelation_matrix = dataset_df[coordinate_columns].corr()\nplt.figure(figsize=(5, 4))\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', center=0)\nplt.title('Coordinate Features Correlation')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:46.169477Z","iopub.execute_input":"2024-12-07T18:56:46.169774Z","iopub.status.idle":"2024-12-07T18:56:49.754459Z","shell.execute_reply.started":"2024-12-07T18:56:46.169746Z","shell.execute_reply":"2024-12-07T18:56:49.753230Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 5. Level-based Analysis\nlevel_analysis = dataset_df.groupby('level')['session_id'].nunique()\nplt.figure(figsize=(7.5, 3))\nlevel_analysis.plot(kind='bar')\nplt.title('Number of Unique Sessions per Level')\nplt.xlabel('Level')\nplt.ylabel('Unique Sessions')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:49.755935Z","iopub.execute_input":"2024-12-07T18:56:49.756283Z","iopub.status.idle":"2024-12-07T18:56:52.430866Z","shell.execute_reply.started":"2024-12-07T18:56:49.756231Z","shell.execute_reply":"2024-12-07T18:56:52.429729Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 6. Event Type Analysis\nevent_level_crosstab = pd.crosstab(dataset_df['event_name'], dataset_df['level_group'], normalize='index')\nplt.figure(figsize=(6, 4))\nevent_level_crosstab.plot(kind='bar', stacked=True)\nplt.title('Event Types Across Level Groups')\nplt.xlabel('Event Name')\nplt.ylabel('Proportion')\nplt.legend(title='Level Group', bbox_to_anchor=(1.05, 1), loc='upper left')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:52.435298Z","iopub.execute_input":"2024-12-07T18:56:52.435636Z","iopub.status.idle":"2024-12-07T18:56:54.620171Z","shell.execute_reply.started":"2024-12-07T18:56:52.435605Z","shell.execute_reply":"2024-12-07T18:56:54.619019Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Load the labels\n\nThe labels for the training dataset are stored in the `train_labels.csv`. It consists of the information on whether the user in a particular session answered each question correctly. Load the labels data by running the following code. `","metadata":{"id":"ReNY-i3bn2O8"}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')","metadata":{"id":"KD4uayl2n2O9","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:54.621590Z","iopub.execute_input":"2024-12-07T18:56:54.622434Z","iopub.status.idle":"2024-12-07T18:56:54.977990Z","shell.execute_reply.started":"2024-12-07T18:56:54.622394Z","shell.execute_reply":"2024-12-07T18:56:54.976870Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Each value in the column, `session_id` is a combination of both the session and the question number. \nWe will split these into individual columns for ease of use.","metadata":{"id":"SMKh2KAPn2O9"}},{"cell_type":"code","source":"labels['session'] = labels.session_id.apply(lambda x: int(x.split('_')[0]) )\nlabels['q'] = labels.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )","metadata":{"id":"Kva8_Dbqn2O9","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:54.981269Z","iopub.execute_input":"2024-12-07T18:56:54.981859Z","iopub.status.idle":"2024-12-07T18:56:55.737995Z","shell.execute_reply.started":"2024-12-07T18:56:54.981817Z","shell.execute_reply":"2024-12-07T18:56:55.736579Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":" Let us take a look at the first 5 entries of `labels` using the following code:","metadata":{"id":"0EQFN3vin2O9"}},{"cell_type":"code","source":"# Display the first 5 examples\nlabels.head(5)","metadata":{"id":"0eD-KZMvn2O-","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:55.739822Z","iopub.execute_input":"2024-12-07T18:56:55.740287Z","iopub.status.idle":"2024-12-07T18:56:55.752756Z","shell.execute_reply.started":"2024-12-07T18:56:55.740210Z","shell.execute_reply":"2024-12-07T18:56:55.751697Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Our goal is to train models for each question to predict the label `correct` for any input user session. ","metadata":{"id":"7rQSOYAYqcZ2"}},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"- The competition dataset introduces a unique structure where game data is organized into three distinct level segments: `0-4`, `5-12`, and `13-22`.These segments are represented by the `level_group` column. Our primary objective is to predict question correctness for each of these segments using the available game interaction data.\n- To tackle this challenge, we'll develop a feature engineering strategy that transforms the raw event data into meaningful predictive features. While the provided columns offer a rich source of information, we'll be strategic in our feature selection.\n\nKey Considerations for Feature Engineering:\n- Some columns like `fullscreen`, `hq`, and `music` will be excluded as they don't directly contribute to predicting question correctness\n- The goal is to extract meaningful signals from the game interaction data that can help predict student performance","metadata":{"id":"y5fK05dsn2O_"}},{"cell_type":"code","source":"CATEGORICAL = ['event_name', 'name','fqid', 'room_fqid', 'text_fqid']\nNUMERICAL = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']","metadata":{"id":"cCZWGiL_n2PA","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:55.753885Z","iopub.execute_input":"2024-12-07T18:56:55.754194Z","iopub.status.idle":"2024-12-07T18:56:55.763676Z","shell.execute_reply.started":"2024-12-07T18:56:55.754164Z","shell.execute_reply":"2024-12-07T18:56:55.762554Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We'll create new features through two main steps:\n\n- Categorical Columns:\n\nGroup data by `session_id` and `level_group`\nCount unique elements for each categorical column\nStore these counts in temporary dataframes\n\n\n- Numerical Columns:\n\nGroup data by `session_id` and `level_group`\nCalculate mean and standard deviation for each numerical column\nStore statistical summaries in temporary dataframes\n\n\n\nFinal step: Concatenate all temporary dataframes to generate a comprehensive feature-engineered dataset.","metadata":{"id":"Us9sScDSn2PA"}},{"cell_type":"code","source":"def feature_engineer(dataset_df):\n    dfs = []\n    for c in CATEGORICAL:\n        tmp = dataset_df.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n    for c in NUMERICAL:\n        tmp = dataset_df.groupby(['session_id','level_group'])[c].agg('mean')\n        dfs.append(tmp)\n    for c in NUMERICAL:\n        tmp = dataset_df.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    dataset_df = pd.concat(dfs,axis=1)\n    dataset_df = dataset_df.fillna(-1)\n    dataset_df = dataset_df.reset_index()\n    dataset_df = dataset_df.set_index('session_id')\n    return dataset_df","metadata":{"id":"nHWhAOtTn2PA","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:55.765559Z","iopub.execute_input":"2024-12-07T18:56:55.766300Z","iopub.status.idle":"2024-12-07T18:56:55.779085Z","shell.execute_reply.started":"2024-12-07T18:56:55.766229Z","shell.execute_reply":"2024-12-07T18:56:55.777579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset_df = feature_engineer(dataset_df)\nprint(\"Prepared dataset shape is {}\".format(dataset_df.shape))","metadata":{"id":"JKcoPoemn2PA","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:56:55.781156Z","iopub.execute_input":"2024-12-07T18:56:55.781659Z","iopub.status.idle":"2024-12-07T18:57:36.837558Z","shell.execute_reply.started":"2024-12-07T18:56:55.781611Z","shell.execute_reply":"2024-12-07T18:57:36.836374Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploration of the dataset","metadata":{"id":"Ij7TT3x-n2PB"}},{"cell_type":"markdown","source":"Let us print out the first 5 entries using the following code:","metadata":{"id":"s1c59fMAn2PB"}},{"cell_type":"code","source":"# Display the first 5 examples\ndataset_df.head(5)","metadata":{"id":"mvQEsdV1n2PB","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:36.838591Z","iopub.execute_input":"2024-12-07T18:57:36.838903Z","iopub.status.idle":"2024-12-07T18:57:36.871814Z","shell.execute_reply.started":"2024-12-07T18:57:36.838872Z","shell.execute_reply":"2024-12-07T18:57:36.870351Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dataset_df.describe()","metadata":{"id":"DRusg-N1n2PB","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:36.873165Z","iopub.execute_input":"2024-12-07T18:57:36.873513Z","iopub.status.idle":"2024-12-07T18:57:37.021681Z","shell.execute_reply.started":"2024-12-07T18:57:36.873483Z","shell.execute_reply":"2024-12-07T18:57:37.020558Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Numerical data distribution¶","metadata":{"id":"xboIq9oDn2PB"}},{"cell_type":"code","source":"figure, axis = plt.subplots(3, 2, figsize=(10, 10))\n\nfor name, data in dataset_df.groupby('level_group'):\n    axis[0, 0].plot(range(1, len(data['room_coor_x_std'])+1), data['room_coor_x_std'], label=name)\n    axis[0, 1].plot(range(1, len(data['room_coor_y_std'])+1), data['room_coor_y_std'], label=name)\n    axis[1, 0].plot(range(1, len(data['screen_coor_x_std'])+1), data['screen_coor_x_std'], label=name)\n    axis[1, 1].plot(range(1, len(data['screen_coor_y_std'])+1), data['screen_coor_y_std'], label=name)\n    axis[2, 0].plot(range(1, len(data['hover_duration'])+1), data['hover_duration_std'], label=name)\n    axis[2, 1].plot(range(1, len(data['elapsed_time_std'])+1), data['elapsed_time_std'], label=name)\n    \n\naxis[0, 0].set_title('room_coor_x')\naxis[0, 1].set_title('room_coor_y')\naxis[1, 0].set_title('screen_coor_x')\naxis[1, 1].set_title('screen_coor_y')\naxis[2, 0].set_title('hover_duration')\naxis[2, 1].set_title('elapsed_time_std')\n\nfor i in range(3):\n    axis[i, 0].legend()\n    axis[i, 1].legend()\n\nplt.show()","metadata":{"id":"mXIiaq_bn2PC","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:37.023192Z","iopub.execute_input":"2024-12-07T18:57:37.023774Z","iopub.status.idle":"2024-12-07T18:57:39.672805Z","shell.execute_reply.started":"2024-12-07T18:57:37.023737Z","shell.execute_reply":"2024-12-07T18:57:39.671887Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Split Dataset into training and testing\ndef split_dataset(dataset, test_ratio=0.20):\n    USER_LIST = dataset.index.unique()\n    split = int(len(USER_LIST) * (1 - 0.20))\n    return dataset.loc[USER_LIST[:split]], dataset.loc[USER_LIST[split:]]\n\ntrain_x, valid_x = split_dataset(dataset_df)\nprint(\"{} examples in Training, {} examples in Testing.\".format(\n    len(train_x), len(valid_x)))","metadata":{"id":"OZfTcCJfn2PC","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:39.673786Z","iopub.execute_input":"2024-12-07T18:57:39.674227Z","iopub.status.idle":"2024-12-07T18:57:39.744022Z","shell.execute_reply.started":"2024-12-07T18:57:39.674174Z","shell.execute_reply":"2024-12-07T18:57:39.742822Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Building","metadata":{}},{"cell_type":"markdown","source":"\nThere are several tree-based models wr can use:\n\n- RandomForestModel\n- GradientBoostedTreesModel\n- CartModel\n- DistributedGradientBoostedTreesModel\n\nBut in this implementation, we are going to use `Random Forest Model` to train for each questions.","metadata":{"id":"ZnVfKZfzn2PE"}},{"cell_type":"code","source":"tfdf.keras.get_all_models()","metadata":{"id":"KZBdcVU1n2PE","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:39.745207Z","iopub.execute_input":"2024-12-07T18:57:39.745760Z","iopub.status.idle":"2024-12-07T18:57:39.752379Z","shell.execute_reply.started":"2024-12-07T18:57:39.745725Z","shell.execute_reply":"2024-12-07T18:57:39.751437Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Training","metadata":{"id":"UdibIrM-XP5-"}},{"cell_type":"markdown","source":"\nWe'll develop a tailored approach for predicting question correctness:\n\n- Total Questions: 18\n- Training Methodology: Individual model for each question\n- Key Data Structures Needed: Storage for trained models - Validation set predictions - Performance evaluation scores","metadata":{"id":"BtkKMa7sXd65"}},{"cell_type":"code","source":"VALID_USER_LIST = valid_x.index.unique()\nprediction_df = pd.DataFrame(data=np.zeros((len(VALID_USER_LIST),18)), index=VALID_USER_LIST)\n\nmodels = {}\n\nevaluation_dict ={}","metadata":{"id":"7Brds67Wn2PD","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:39.753526Z","iopub.execute_input":"2024-12-07T18:57:39.753844Z","iopub.status.idle":"2024-12-07T18:57:39.765530Z","shell.execute_reply.started":"2024-12-07T18:57:39.753813Z","shell.execute_reply":"2024-12-07T18:57:39.764281Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Before training the data we have to understand how `level_groups` and `questions` are connected\n\nIn this game the first quiz checkpoint(i.e., questions 1 to 3) comes after finishing levels 0 to 4\n\nSo for training questions 1 to 3 we will use data from the `level_group` 0-4\n\nSimilarly, we will use data from the `level_group` 5-12 to train questions from 4 to 13 and data from the `level_group` 13-22 to train questions from 14 to 18.\n\nTraining Approach: Create a `unique model` for each question, storing them in a dictionary for precise performance prediction across different game stages.","metadata":{"id":"p1GoOMRHn2PF"}},{"cell_type":"code","source":"# Iterate through questions 1 to 18\nfor q_no in range(1,19):\n    # Determine the level group based on question number\n    if q_no<=3: grp = '0-4'\n    elif q_no<=13: grp = '5-12'\n    elif q_no<=22: grp = '13-22'\n    print(\"### Question Number\", q_no, \"Group\", grp)\n    \n    # Filter training data for the specific level group\n    train_df = train_x.loc[train_x.level_group == grp]\n    train_users = train_df.index.values\n    \n    # Filter validation data for the specific level group\n    valid_df = valid_x.loc[valid_x.level_group == grp]\n    valid_users = valid_df.index.values\n    \n    # Extract labels for the current question for training users\n    train_labels = labels.loc[labels.q==q_no].set_index('session').loc[train_users]\n    \n    # Extract labels for the current question for validation users\n    valid_labels = labels.loc[labels.q==q_no].set_index('session').loc[valid_users]\n    \n    # Add 'correct' column to training and validation dataframes\n    train_df[\"correct\"] = train_labels[\"correct\"]\n    valid_df[\"correct\"] = valid_labels[\"correct\"]\n    \n    # Convert training dataframe to TensorFlow dataset, excluding 'level_group'\n    train_ds = tfdf.keras.pd_dataframe_to_tf_dataset(train_df.loc[:, train_df.columns != 'level_group'], label=\"correct\")\n    \n    # Convert validation dataframe to TensorFlow dataset, excluding 'level_group'\n    valid_ds = tfdf.keras.pd_dataframe_to_tf_dataset(valid_df.loc[:, valid_df.columns != 'level_group'], label=\"correct\")\n    \n    # Create a Random Forest model\n    gbtm = tfdf.keras.RandomForestModel(verbose=0)\n    \n    # Compile the model with accuracy metric\n    gbtm.compile(metrics=[\"accuracy\"])\n    \n    # Train the model on the training dataset\n    gbtm.fit(x=train_ds)\n    \n    # Store the trained model in the models dictionary\n    models[f'{grp}_{q_no}'] = gbtm\n    \n    # Create an inspector to evaluate the model\n    inspector = gbtm.make_inspector()\n    inspector.evaluation()\n    \n    # Evaluate the model on the validation dataset and store accuracy\n    evaluation = gbtm.evaluate(x=valid_ds,return_dict=True)\n    evaluation_dict[q_no] = evaluation[\"accuracy\"]\n    \n    # Make predictions on the validation dataset\n    predict = gbtm.predict(x=valid_ds)\n    \n    # Store predictions in the prediction dataframe\n    prediction_df.loc[valid_users, q_no-1] = predict.flatten() ","metadata":{"id":"VBO3VCOJn2PF","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T18:57:39.767531Z","iopub.execute_input":"2024-12-07T18:57:39.767961Z","iopub.status.idle":"2024-12-07T19:03:41.580159Z","shell.execute_reply.started":"2024-12-07T18:57:39.767922Z","shell.execute_reply":"2024-12-07T19:03:41.579068Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Accuracy & Achieved ","metadata":{}},{"cell_type":"code","source":"for name, value in evaluation_dict.items():\n  print(f\"Question {name}: accuracy {value:.4f}\")\n\nprint(\"\\nAverage accuracy\", sum(evaluation_dict.values())/18)","metadata":{"id":"qPOfPkm7n2PG","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:41.582964Z","iopub.execute_input":"2024-12-07T19:03:41.583337Z","iopub.status.idle":"2024-12-07T19:03:41.588924Z","shell.execute_reply.started":"2024-12-07T19:03:41.583302Z","shell.execute_reply":"2024-12-07T19:03:41.587979Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Visualize the model","metadata":{"id":"qtiCtfNCn2PG"}},{"cell_type":"code","source":"tfdf.model_plotter.plot_model_in_colab(models['0-4_1'], tree_idx=0, max_depth=3)","metadata":{"id":"od__6uAan2PG","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:41.589811Z","iopub.execute_input":"2024-12-07T19:03:41.590100Z","iopub.status.idle":"2024-12-07T19:03:41.676439Z","shell.execute_reply.started":"2024-12-07T19:03:41.590071Z","shell.execute_reply":"2024-12-07T19:03:41.675315Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Finding Key Features","metadata":{}},{"cell_type":"code","source":"inspector = models['0-4_1'].make_inspector()\n\nprint(f\"Available variable importances:\")\nfor importance in inspector.variable_importances().keys():\n  print(\"\\t\", importance)","metadata":{"id":"d_gvL9nbn2PH","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:41.677799Z","iopub.execute_input":"2024-12-07T19:03:41.678148Z","iopub.status.idle":"2024-12-07T19:03:41.689508Z","shell.execute_reply.started":"2024-12-07T19:03:41.678113Z","shell.execute_reply":"2024-12-07T19:03:41.688025Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"inspector.variable_importances()[\"NUM_AS_ROOT\"]","metadata":{"id":"ZDsxqRrwn2PH","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:41.691750Z","iopub.execute_input":"2024-12-07T19:03:41.692101Z","iopub.status.idle":"2024-12-07T19:03:41.702709Z","shell.execute_reply.started":"2024-12-07T19:03:41.692068Z","shell.execute_reply":"2024-12-07T19:03:41.701315Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# More Aspects to consider","metadata":{"id":"PIK5aUH-n2PH"}},{"cell_type":"code","source":"# Create a DataFrame to store true labels for validation users\n# Initialize with zeros, with length equal to number of validation users\ntrue_df = pd.DataFrame(data=np.zeros((len(VALID_USER_LIST),18)), index=VALID_USER_LIST)\n\n# Iterate through 18 questions\nfor i in range(18):\n    # Extract true labels for each question for validation users\n    # Set index to session, filter for specific validation users\n    tmp = labels.loc[labels.q == i+1].set_index('session').loc[VALID_USER_LIST]\n    \n    # Store true labels in the corresponding column of true_df\n    true_df[i] = tmp.correct.values\n\n# Initialize variables to track best threshold\nmax_score = 0\nbest_threshold = 0\n\n# Iterate through threshold values from 0.4 to 0.8 with 0.01 step\nfor threshold in np.arange(0.4,0.8,0.01):\n    # Create F1 Score metric with macro averaging\n    metric = tfa.metrics.F1Score(num_classes=2,average=\"macro\",threshold=threshold)\n    \n    # Convert true labels to one-hot encoded format\n    y_true = tf.one_hot(true_df.values.reshape((-1)), depth=2)\n    \n    # Convert predictions to binary labels based on current threshold\n    # Reshape predictions and apply threshold\n    y_pred = tf.one_hot((prediction_df.values.reshape((-1))>threshold).astype('int'), depth=2)\n    \n    # Update metric state with true and predicted labels\n    metric.update_state(y_true, y_pred)\n    \n    # Get F1 score\n    f1_score = metric.result().numpy()\n    \n    # Update best threshold if current F1 score is higher\n    if f1_score > max_score:\n        max_score = f1_score\n        best_threshold = threshold\n\n# Print the best threshold and corresponding F1 score\nprint(\"Best threshold \", best_threshold, \"\\tF1 score \", max_score)","metadata":{"id":"2wptRs3In2PH","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:41.708286Z","iopub.execute_input":"2024-12-07T19:03:41.708886Z","iopub.status.idle":"2024-12-07T19:03:42.751528Z","shell.execute_reply.started":"2024-12-07T19:03:41.708849Z","shell.execute_reply":"2024-12-07T19:03:42.750598Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Submission","metadata":{"id":"ezA40GQ4n2PH"}},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    test_df = feature_engineer(test)\n    grp = test_df.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        gbtm = models[f'{grp}_{t}']\n        test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df.loc[:, test_df.columns != 'level_group'])\n        predictions = gbtm.predict(test_ds)\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        n_predictions = (predictions > best_threshold).astype(int)\n        sample_submission.loc[mask,'correct'] = n_predictions.flatten()\n    \n    env.predict(sample_submission)","metadata":{"id":"gHiXTnTVn2PI","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:42.752451Z","iopub.execute_input":"2024-12-07T19:03:42.752735Z","iopub.status.idle":"2024-12-07T19:03:48.045569Z","shell.execute_reply.started":"2024-12-07T19:03:42.752706Z","shell.execute_reply":"2024-12-07T19:03:48.043809Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"! head submission.csv","metadata":{"id":"iYBXokAyn2PI","trusted":true,"execution":{"iopub.status.busy":"2024-12-07T19:03:48.047567Z","iopub.execute_input":"2024-12-07T19:03:48.047997Z","iopub.status.idle":"2024-12-07T19:03:49.370961Z","shell.execute_reply.started":"2024-12-07T19:03:48.047956Z","shell.execute_reply":"2024-12-07T19:03:49.369163Z"}},"outputs":[],"execution_count":null}]}