{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction \n\nThis notebook illustrates the use of a **XGBoost model** for the Predict Student Performance from Game Play challenge. The model assumes no correlation across rows (i.e., ignore the time-series nature of the problem) and uses a **minimal set of training features**. The XGBoost model is optimised with a 5-fold cross training procedure looking for the optimal maximum depth of the model. A simple estimation of the accuracy on a partitioned dataset built from the training data yields to about 70% accuracy score.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-26T20:50:19.719401Z","iopub.execute_input":"2023-02-26T20:50:19.720113Z","iopub.status.idle":"2023-02-26T20:50:19.750983Z","shell.execute_reply.started":"2023-02-26T20:50:19.720011Z","shell.execute_reply":"2023-02-26T20:50:19.750233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Import libraries to be used. This notebook will use XGBoost.  ","metadata":{}},{"cell_type":"code","source":"from pandas import read_csv\n\n# Basic plots\nimport matplotlib.pyplot as plt\nimport seaborn as sn\n\nfrom xgboost import XGBClassifier\n\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:50:19.752688Z","iopub.execute_input":"2023-02-26T20:50:19.753260Z","iopub.status.idle":"2023-02-26T20:50:20.927982Z","shell.execute_reply.started":"2023-02-26T20:50:19.753228Z","shell.execute_reply":"2023-02-26T20:50:20.926784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Define functionalities (options and plotting function to be used later).\nThe definition of the features to be used later is based on having a small initial set with **low rate of NaNs**. This is clearly a sub-optimal approach.  ","metadata":{}},{"cell_type":"code","source":"dummy_optimisation = True\nfeature_importance_plot = True \n\ndef conf_matrix_plot(matrix, plot_name):\n    fig = plt.figure(figsize = (10, 10))\n    plt.title('Correlation matrix')\n    mask = np.triu(matrix)\n    ax = sn.heatmap(matrix, annot=True, fmt='.2f', vmin=-1, vmax=1, center= 0, cmap= 'vlag') # mask=mask)\n    plt.tight_layout()\n    plt.savefig('heatmap_'+plot_name+'.png')    \n\n#features_to_keep = ['elapsed_time', 'event_name', 'name', 'level', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'level_group', 'correct']\n#features_to_keep_test = ['elapsed_time', 'event_name', 'name', 'level', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'level_group']\n\n\nfeatures_to_keep = ['elapsed_time', 'event_name', 'level_group', 'correct']\nfeatures_to_keep_test = ['elapsed_time', 'event_name', 'level_group']\n\n#categorical_data = ['event_name', 'name', 'level_group']\ncategorical_data = ['event_name', 'level_group']","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:50:20.932066Z","iopub.execute_input":"2023-02-26T20:50:20.932391Z","iopub.status.idle":"2023-02-26T20:50:20.942317Z","shell.execute_reply.started":"2023-02-26T20:50:20.932362Z","shell.execute_reply":"2023-02-26T20:50:20.941549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Load the data !** Also define a new column in the training dataset containing the values for the target label.  ","metadata":{}},{"cell_type":"code","source":"data = read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv')\ndata_test = read_csv('/kaggle/input/predict-student-performance-from-game-play/test.csv')\ndata_sub = read_csv('/kaggle/input/predict-student-performance-from-game-play/test.csv')\ntrain_labels = read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ndata['correct'] = train_labels['correct']","metadata":{"execution":{"iopub.status.busy":"2023-02-26T20:50:20.947322Z","iopub.execute_input":"2023-02-26T20:50:20.949417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Check the input data.**   ","metadata":{}},{"cell_type":"code","source":"print(data.head())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# NaNs are our enemies !\n\nLooking now for NaNs in the dataset, will make a plot of their distribution in the training dataset.   ","metadata":{}},{"cell_type":"code","source":"# ---------- Check for Nans\n\nnans_fractions = []\n# Count how many nans there are in each column\nprint('\\n Looking at distribution of NaNs')\nfor column in data.columns:\n    num_of_nans = data[column].isna().sum()\n    nans_fractions.append(num_of_nans/len(data))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now look at the **distribution of the NaNs per feature**. I will remove features that have large amount of NaNs. Here the goal is to find a very basic solution with **minimal features extraction and imputation.**   ","metadata":{}},{"cell_type":"code","source":"ax = sn.barplot(x=data.columns, y=nans_fractions, color='lightblue', edgecolor='black')\nax.set(xlabel='Feature', ylabel='Fraction of NaNs [%]')\nplt.xticks(rotation=90)\nplt.grid()\nplt.tight_layout()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's remove all features with large fraction of NaNs, keep things as simple as possible. As a reminder the features to consider are defined at the beginning of the notebook (['elapsed_time', 'event_name', 'name', 'level', 'room_coor_x', 'room_coor_y', 'screen_coor_x', 'screen_coor_y', 'level_group', 'correct'])","metadata":{}},{"cell_type":"code","source":"data = data[features_to_keep]\ndata = data.dropna()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Categorical data to numeric aka XGBoost speaks only numbers !\n\nNow let's look at a selection of categorical data ! ","metadata":{}},{"cell_type":"code","source":"# Look at categorical data and their unique values\nfor element in categorical_data:\n    print('Unique values for categorical data '+str(element)+' ='+str(data[element].nunique()))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Each categorical feature has a limited set of values, so it should be well usable in a simple ML model.** ","metadata":{}},{"cell_type":"code","source":"# Convert other categorical data objects into numbers\ncategorical_data_to_convert = categorical_data\nfor cat_data_to_cnv in categorical_data_to_convert:\n  data[cat_data_to_cnv] = pd.Categorical(data[cat_data_to_cnv]).codes\n  data_test[cat_data_to_cnv] = pd.Categorical(data_test[cat_data_to_cnv]).codes","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training summary\n\nLet's see if all the data is in numerical format now. ","metadata":{}},{"cell_type":"code","source":"print(data.head())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's have a look at the **correlation between features** and the target to see if we can get a glimpse of what features will matter in the end. ","metadata":{}},{"cell_type":"code","source":"conf_matrix_plot(data.corr(), 'correlation_features')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training time !\n\nNow let's split the data in multiple subsets (training/testing/validation). The used fractions of data are 60% training, 20% validation, 20% testing.  ","metadata":{}},{"cell_type":"code","source":"X = data.drop('correct', axis = 1)\ny = data['correct'] \n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.4, random_state=42)\nX_val, X_test, y_val, y_test = train_test_split(X_test, y_test, test_size=0.5, random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Define and fit the model.**","metadata":{}},{"cell_type":"code","source":"model = XGBClassifier()\n\nfrom sklearn.model_selection import GridSearchCV\n\nparameters = {}\n\nif dummy_optimisation:\n    parameters = {\n        'max_depth': [5,6,7] # Default is 6 \n    }\nelse:\n    parameters = {\n        'eta': [0.05*i for i in range(2, 8)], # Default is 0.3\n        'gamma': [0.05*i for i in range(0, 3)], # Default is 0 \n        'max_depth': [4,5,6,7,8], # Default is 6 \n        'max_leaves': [0,1,2] # Default is 0 \n    }    \n    \ncv = GridSearchCV(model, parameters, cv=5)\ncv.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's do a plot of the features importance now.** ","metadata":{}},{"cell_type":"code","source":"if feature_importance_plot: \n    fig = plt.figure(figsize = (4, 8))\n    feat_imp = cv.best_estimator_.feature_importances_\n    indices = np.argsort(feat_imp)\n    plt.yticks(range(len(indices)), [X_train.columns[i] for i in indices])\n    plt.barh(range(len(indices)), feat_imp[indices], color='lightgreen', align='center', edgecolor='black')\n    plt.grid()\n    plt.tight_layout()    \n    plt.savefig('feature_importance.png')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Testing time \n\nLet's now try to get an estimation of the accuracy on the 'testing data' prepared above. This testing data is a partition of the training data. The accuracy is estimated **without refitting the model to the full training data** so it is likely under-estimated here. ","metadata":{}},{"cell_type":"code","source":"y_pred = cv.best_estimator_.predict(X_test)\npredictions = [value for value in y_pred]\nfrom sklearn.metrics import accuracy_score\naccuracy = accuracy_score(y_test, predictions)\nprint(\"Accuracy: %.2f%%\" % (accuracy * 100.0))\nprint('And the best HP setting is: ... ')\nprint(cv.best_params_)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's talk to Kaggle now\n\nOK all seems set, now we are left with preparing the submission sample. \n**This part is just a draft**. \n\nFirst load the right libraries.  ","metadata":{}},{"cell_type":"code","source":"# For the submission in Kaggle \nimport jo_wilder\nenv = jo_wilder.make_env()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"iter_test = env.iter_test()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And the relevant function needed for the scoring ","metadata":{}},{"cell_type":"code","source":"counter = 0\n# The API will deliver two dataframes in this specific order,\n# for every session+level grouping (one group per session for each checkpoint)\nprint(counter)\nprint(iter_test)\nfor (sample_submission, test) in iter_test:\n    if counter == 0:\n        print(sample_submission.head())\n        print(test.head())\n        print(test.shape)\n\n    print('In test API')\n    print(test.columns)\n    test = test.fillna(-1)\n    for cat_data_to_cnv in categorical_data_to_convert:\n      print (\"Handling now data category \"+cat_data_to_cnv)\n      test[cat_data_to_cnv] = pd.Categorical(test[cat_data_to_cnv]).codes\n\n    test = test.dropna()  \n    y_pred_final = cv.best_estimator_.predict(test[features_to_keep_test])\n    predictions = [1 if value== 1 else 0 for value in y_pred_final]\n\n    ## users make predictions here using the test data\n    sample_submission['correct'] = pd.Series(predictions)\n    \n    ## env.predict appends the session+level sample_submission to the overall\n    ## submission\n    env.predict(sample_submission)\n    counter += 1","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}