{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Predicting student performance with datatable LinearModel\n\n[datatable](https://datatable.readthedocs.io/) is a Python package for manipulating data frames. It is close in spirit to pandas, however, specific emphasis is put on speed and big data support. A limited number of machine learning models and tools are also available in the package.\n\nIn this notebook we will focus on the newly developed datatable [LinearModel](https://datatable.readthedocs.io/en/latest/api/models/linear_model.html) to see how well it can predict student performance from game play. We will also use datatable to read the competition data, do data munging and feature engineering.\n\nThere are already many great baselines published for this competition, that involve much more sophisticated models. From the other hand, there are pretty tough restrictions on the computational resources. This makes us to believe that `LinearModel` may also have some potential, as it is fast, lightweight and has high interpretability. So let's give it a try!\n\n*Spoiler: after just six seconds of fitting, the model achieves the same F1 score as much more sophisticated baselines.*\n\n## Contents\n - [Install datatable and read in the competition data](#install) \n - [Feature engineering](#features)\n - [Fitting the model](#fit)\n - [Determining the optimal threshold](#threshold)\n - [Predicting and submitting](#predict)\n - [Conclusions](#conclusions)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"install\"></a>\n## Install datatable and read in the competition data\n\nFirst, we install the [latest datatable wheels](https://www.kaggle.com/datasets/kononenko/datatable), so that the package could be used in a no-internet code competition. \n\nSecond, we read the training data and the labels with fully parallel [datatable.fread()](https://datatable.readthedocs.io/en/latest/api/dt/fread.html) function. Unlike with pandas, there is no need to manually adjust any of the column types. The resulting data frames fit into 8Gb of RAM out of the box.","metadata":{}},{"cell_type":"code","source":"%%time\n# Install the latest datatable wheel\n!pip install /kaggle/input/datatable/datatable-1.1.0a2240-cp310-cp310-manylinux_2_17_x86_64.whl\nfrom datatable import dt, by, f, g\n\n# Read the training data\nDT_train_full = dt.fread(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\")\n\n# Read labels, then parse session ids and question numbers\nDT_labels = dt.Frame(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\nDT_labels = DT_labels[:, {\n                          \"session_id\" : f.session_id[:17].as_type(dt.int64),\n                          \"q\" : f.session_id[19:].as_type(dt.int8),\n                          \"correct\" : f.correct\n                     }]","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:45:30.483138Z","iopub.execute_input":"2023-06-06T07:45:30.483924Z","iopub.status.idle":"2023-06-06T07:46:55.636933Z","shell.execute_reply.started":"2023-06-06T07:45:30.483890Z","shell.execute_reply":"2023-06-06T07:46:55.634931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DT_train_full","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:46:55.639142Z","iopub.execute_input":"2023-06-06T07:46:55.640040Z","iopub.status.idle":"2023-06-06T07:46:55.649636Z","shell.execute_reply.started":"2023-06-06T07:46:55.639986Z","shell.execute_reply":"2023-06-06T07:46:55.648364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DT_labels","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:46:55.651519Z","iopub.execute_input":"2023-06-06T07:46:55.652683Z","iopub.status.idle":"2023-06-06T07:46:55.667682Z","shell.execute_reply.started":"2023-06-06T07:46:55.652645Z","shell.execute_reply":"2023-06-06T07:46:55.666363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"features\"></a>\n## Feature engineering\n\nIn this notebook we will stick to a limited number of features, namely\n- number of events per a session;\n- session time;\n- number of unique elements in each of the categorical columns.","metadata":{}},{"cell_type":"code","source":"# Feature engineering on `DT` columns\ndef get_features(DT): \n    # From the numeric columns we create the following two features: \n    # - number of events per session, i.e. `session_id.count()`;\n    # - session time, i.e. `elapsed_time.max()`.\n    # To reduce skewness, we also apply log transformation to these two features.\n    features = {\n                \"session_id_log_count\" : dt.math.log(1 + f.session_id.count()),\n                \"elapsed_time_log_max\" : dt.math.log(1 + f.elapsed_time.max()),\n               }\n\n    # From the categorical columns we create the corresponding \"nunique\" features\n    CAT_COLS = (\"event_name\", \"name\", \"text\", \"fqid\", \"room_fqid\", \"text_fqid\")\n    for col in CAT_COLS:\n        features[col + \"_nunique\"] = f[col].nunique()\n    \n    # The above aggregations are done by `session_id` and `level_group` columns   \n    return DT[: , features, by(\"session_id\", \"level_group\")]\n\n%time DT_train = get_features(DT_train_full)\n\n# Make sure the number of sessions per a level group in the training data\n# is the same as the number of sessions per question in labels.\nprint(\"Training data consistent:\", DT_train.nrows//3 == DT_labels.nrows//18)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-06-06T07:46:55.670895Z","iopub.execute_input":"2023-06-06T07:46:55.671670Z","iopub.status.idle":"2023-06-06T07:47:22.960967Z","shell.execute_reply.started":"2023-06-06T07:46:55.671631Z","shell.execute_reply":"2023-06-06T07:47:22.959613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DT_train","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:47:22.962975Z","iopub.execute_input":"2023-06-06T07:47:22.964268Z","iopub.status.idle":"2023-06-06T07:47:22.981040Z","shell.execute_reply.started":"2023-06-06T07:47:22.964228Z","shell.execute_reply":"2023-06-06T07:47:22.980114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"fit\"></a>\n## Fitting the model\n\nFor each question we create one `LinearModel`, i.e. `18` models in total. We also standardize all the training data to improve performance of these models.","metadata":{}},{"cell_type":"code","source":"%%time\nfrom datatable.models import LinearModel\n\n# Correspondance of level groups and question numbers.\nlevel_groups = {\"0-4\"   : (1, 4), \n                \"5-12\"  : (4, 14), \n                \"13-22\" : (14, 19)}\n\n# Get a level group the q-th question corresponds to\ndef get_level_group(q):\n    for level_group, (q_from, q_to) in level_groups.items():\n        if q >= q_from and q < q_to:\n            return level_group\n    \n# Model parameters\neta0 = 0.0005 # Learning rate\nnepochs = 10  # Number of training epochs \n\nmodels = []\ntrain_stats = []\nDT_train_qs = []\nDT_labels_qs = []\n\n# Fit one linear model per question\nfor q in range(1, 19):\n    print(\"Fitting model for question %d...\" % q)\n    model = LinearModel(eta0=eta0, nepochs=nepochs)\n    \n    # Preparing training data.\n    # For question `q` we determine the corresponding `level_group`,\n    # and then use training data for this `level_group` only.\n    level_group_q = get_level_group(q)\n    DT_train_q = DT_train[f.level_group == level_group_q, :]\n\n    # Preparing labels.\n    # First, we select labels for question `q`. To make sure\n    # these labels go in the same order as the training data,\n    # we join `DT_train_q` and `DT_labels_q` on `session_id`.\n    DT_labels_q = DT_labels[f.q == q, :]\n    DT_labels_q.key = [\"session_id\"]\n    DT_labels_q = DT_train_q[:, g.correct, dt.join(DT_labels_q)]\n    DT_labels_qs.append(DT_labels_q)\n    \n    # Finally, we can remove first two columns, i.e. `session_id` \n    # and `level_group`, from the training data.\n    DT_train_q = DT_train_q[:, 2:] \n\n    # Calculate mean and standard deviation for each feature.\n    # Then, standardize training data for each question to improve\n    # performance of the linear model.\n    train_stats.append({\"mean\" : DT_train_q.mean(), \n                        \"sd\"   : DT_train_q.sd()})\n    f_stand = (f[:] - train_stats[q - 1][\"mean\"]) / train_stats[q - 1][\"sd\"]\n    DT_train_qs.append(DT_train_q[:, f_stand])\n    \n    # Fit the q-th model and store it\n    model.fit(DT_train_qs[q - 1], DT_labels_qs[q - 1])\n    models.append(model)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:47:22.982578Z","iopub.execute_input":"2023-06-06T07:47:22.983353Z","iopub.status.idle":"2023-06-06T07:47:29.148746Z","shell.execute_reply.started":"2023-06-06T07:47:22.983322Z","shell.execute_reply":"2023-06-06T07:47:29.147475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"threshold\"></a>\n## Determining the optimal threshold\n\nWe are going to use the fitted models to make predictions. In order to convert predicted probabilities into labels, we need to determine the optimal threshold, i.e. the one which results in the largest F1 score on the training data.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nfrom sklearn.metrics import f1_score\n\n# Predict on the training data for each question\ny_true = []\nDT_preds = dt.Frame()\nfor q in range(1, 19):\n    DT_pred = models[q - 1].predict(DT_train_qs[q - 1])\n    DT_preds.rbind(DT_pred)\n    y_true.extend(DT_labels_qs[q - 1].to_list()[0])\n\n# Go through the reasonable thresholds to determine the one,\n# that gives the largest F1 score\nf1_max = 0\nthreshold_max = 0\nfor threshold in np.arange(0.6, 0.7, 0.01):\n    y_pred = DT_preds[:, f[\"True\"] > threshold].to_list()[0]\n    f1 = f1_score(y_true, y_pred, average=\"macro\")\n    print(\"Threshold: %f -> F1 score: %f\" % (threshold, f1))\n    if f1 > f1_max:\n        f1_max = f1\n        threshold_max = threshold\n        \nprint(\"Optimal threshold: %f -> F1 score: %f\" % (threshold_max, f1_max))","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:47:29.150145Z","iopub.execute_input":"2023-06-06T07:47:29.150464Z","iopub.status.idle":"2023-06-06T07:47:38.603886Z","shell.execute_reply.started":"2023-06-06T07:47:29.150437Z","shell.execute_reply":"2023-06-06T07:47:38.602708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"predict\"></a>\n## Predicting and submitting\n\nTo make predictions, we standardize the test data with means and standard deviations saved at the fitting step. To convert predicted probabilities into classes, we use `threshold_max`, that demonstrated the largest F1 score on the training data.","metadata":{}},{"cell_type":"code","source":"%%time\nimport jo_wilder_310\n\njo_wilder_310.make_env.__called__ = False\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nfor (PD_test, PD_submission) in iter_test:\n    # Create datatable frame from pandas\n    DT_test = dt.Frame(PD_test)\n    \n    # Do feature engineering for the test data\n    DT_test = get_features(DT_test)\n    \n    # Get the level group information. We are supposed to predict\n    # for one level group at a time.\n    level_group = DT_test[0, \"level_group\"]\n    (q_from, q_to) = level_groups[level_group]\n    \n    # Remove `session_id` and `level_group` columns from the test data\n    DT_test = DT_test[:, 2:]\n\n    print(\"\\nPredicting for questions %d-%d\" % (q_from, q_to - 1))\n    for q in range(q_from, q_to):\n        # Predict for the standardized test data\n        f_stand = (f[:] - train_stats[q - 1][\"mean\"]) / train_stats[q - 1][\"sd\"]\n        DT_p = models[q - 1].predict(DT_test[:, f_stand])\n        \n        # Convert probabilities to labels by using F1-optimal threshold\n        DT_p = DT_p[:, (f[\"True\"] > threshold_max).as_type(dt.int8)]\n        \n        # Set up predictions for the q-th question in the submission frame\n        mask_q = PD_submission.session_id.str.contains(f'q{q}')\n        PD_submission.loc[mask_q, 'correct'] = DT_p.to_list()\n    \n    print(PD_submission)\n    env.predict(PD_submission)     \n","metadata":{"execution":{"iopub.status.busy":"2023-06-06T07:47:38.605632Z","iopub.execute_input":"2023-06-06T07:47:38.606148Z","iopub.status.idle":"2023-06-06T07:47:38.809070Z","shell.execute_reply.started":"2023-06-06T07:47:38.606109Z","shell.execute_reply":"2023-06-06T07:47:38.808055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"conclusions\"></a>\n## Conclusions\n\nDespite less than ten seconds of training time, `LinearModel` achieves F1 score of `0.676`, that is similar to much more sophisticated baselines ([TensorFlow Decision Forests](https://www.kaggle.com/code/gusthema/student-performance-w-tensorflow-decision-forests) `0.672`, [XGBoost Baseline](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680) `0.677`, etc.). \n\nClearly, the model could be further improved by exploring other numeric features and paying more attention to categoricals. However, it definitely has a lot of potential when it comes to efficiency and limited computational resources.","metadata":{}}]}