{"cells":[{"metadata":{},"cell_type":"markdown","source":"<img src='https://www.riiid.co/assets/opengraph.png' width='700'>\n\n<h1><center><strong>A Deeper Dive Into Riiid!: EDA|Models|Blending </strong></center><h1>\n    \n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:green; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Problem Statement</center></h2>\n    \n<div align='left'><font size=\"3\" color=\"#000000\">In this competition, your challenge is to create algorithms for \"Knowledge Tracing,\" the modeling of student knowledge over time. The goal is to accurately predict how students will perform on future interactions. You will pair your machine learning skills using Riiid’s EdNet data.</font></div>\n<hr>    \n<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong> AIM: </strong> </font></div></h1>\n<div align='left'><font size=\"3\" color=\"#000000\">1) To create algorithms for \"Knowledge Tracing,\" the modeling of student knowledge over time.</font></div>     \n<div align='left'><font size=\"3\" color=\"#000000\">2) To accurately predict how students will perform on future interactions.</font></div> \n<hr>\n<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong> Purpose: </strong> </font></div></h1>\n<div align='left'><font size=\"3\" color=\"#000000\">1) Help tackle global challenges in education.</font></div>     \n<div align='left'><font size=\"3\" color=\"#000000\">2) Using good ML model,it is possible that any student with an Internet connection can enjoy the benefits of a personalized learning experience.</font></div> \n\n\n"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Evaluation Metrics: ROC-AUC Score</center></h2> \n<div align='left'><font size=\"3\" color=\"#000000\">Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target.</font></div> \n<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong> What is AUC-ROC Curve? </strong> </font></div></h1>\n<div align='left'><font size=\"3\" color=\"#000000\">The Receiver Operator Characteristic (ROC) curve is an evaluation metric for binary classification problems. It is a probability curve that plots the TPR against FPR at various threshold values and essentially separates the ‘signal’ from the ‘noise’. The Area Under the Curve (AUC) is the measure of the ability of a classifier to distinguish between classes and is used as a summary of the ROC curve.\n<strong>ROC-AUC Score is the best measure for binary classification and imbalance class problem. </strong>This iste gives more information about AUC ROC curve<a href=\"https://www.analyticsvidhya.com/blog/2020/06/auc-roc-curve-machine-learning/\" target=\"_blank\">AUC-ROC-Curve.</a> </font></div> \n<img src='https://classeval.files.wordpress.com/2015/06/roc-balanced-imbalanced.png?w=768&h=302' width='700'>\n\n"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:black; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Table Contents</center></h2> \n    \n<div align='left'><font size=\"3\" color=\"#000000\">1. Exploratory Data Analysis</font></div>\n<div align='left'><font size=\"3\" color=\"#000000\">2. Building Baseline Models</font></div>\n<div align='left'><font size=\"3\" color=\"#000000\">3. Ensembling</font></div>\n<div align='left'><font size=\"3\" color=\"#000000\">4. Conclusion</font></div>\n    "},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:black; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>1. Exploratory Data Analysis</center></h2> \n\n    \n    \n<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong>Importing Necessary Libraries</strong> </font></div>    \n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nfrom collections import Counter\nimport pandas_profiling as pp\nfrom sklearn.model_selection import StratifiedKFold\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\nfrom sklearn.metrics import roc_auc_score\nimport warnings\nwarnings.simplefilter('ignore')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import riiideducation\nenv = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" <div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong>Data Description</strong> </font></div>    \n<hr>\n<div align='left'><font size=\"3\" color=\"#000000\"><strong>1. row_id:</strong>ID code for the row.</font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>2. timestamp: </strong>time between this user interaction and the first event.</font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>3. user_id: </strong>ID code for the user</font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>4. content_id:</strong>ID code for the user interaction</font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>5. content_type_id: </strong>0 if question else 1 for Lecture.</font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>6. user_answer: </strong>the user's answer to the question, Read -1 as null, for lectures. </font></div>\n<div align='left'><font size=\"3\" color=\"#000000\"><strong>7. user_answer: </strong>IF the user responded correctly. Read -1 as null, for lectures.</font></div>\n<div align='left'><font size=\"3\" color=\"#000000\"><strong>8. prior_question_elapsed_time: </strong>How long it took a user to answer their previous question bundle, ignoring any lectures in between. It is the total time a user took to solve all questions in the previous bundle. </font></div>\n<div align='left'><font size=\"3\" color=\"#000000\"><strong>9. prior_question_had_explanation: </strong>Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between.</font></div>\n\n"},{"metadata":{},"cell_type":"markdown","source":"<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong>Loading Dataset</strong> </font></div> "},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', low_memory=False,nrows=10**5,  dtype={'row_id': 'int64',\n    'timestamp': 'int64','user_id': 'int32','content_id': 'int16', 'content_type_id': 'int8','task_container_id': 'int16',\n    'user_answer': 'int8','answered_correctly': 'int8','prior_question_elapsed_time': 'float32', 'prior_question_had_explanation': 'boolean'} )    \ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div align='left'><font size=\"4\" color=\"#000000\"><h1 style=\"text-transform: uppercase\"><strong>Data Cleaning</strong> </font></div>    \n<div align='left'><font size=\"3\" color=\"#000000\"><strong>Removing Unwanted columns</strong> </font></div>    \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.drop('user_answer',axis=1)\ntrain = train.query('answered_correctly != -1').reset_index(drop=True)\ntrain.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div align='left'><font size=\"3\" color=\"#000000\"><strong>Missing Value Detection </strong> </font></div>    \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"def missing(df):\n    total = df.isnull().sum().sort_values(ascending = False)\n    total = total[total>0]\n    percent = df.isnull().sum().sort_values(ascending = False)/len(df)*100\n    percent = percent[percent>0]\n    return pd.concat([total, percent], axis=1, keys=['Total','Percentage'])\nmissing(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div align='left'><font size=\"3\" color=\"#000000\"><strong>Missing Value Treatment</strong> </font></div>    \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"train['prior_question_had_explanation'] = train['prior_question_had_explanation'].fillna(train['prior_question_had_explanation'].mode()[0])\ntrain['prior_question_elapsed_time'] = train['prior_question_elapsed_time'].fillna(train['prior_question_elapsed_time'].mean())\nmissing(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Counter(train['answered_correctly'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:black; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>2. Building Baseline Models</center></h2> \n\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install -U pycaret","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['prior_question_had_explanation'] = train['prior_question_had_explanation'].astype(float)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from pycaret.classification import *\nclf1 = setup(train, target = 'answered_correctly',session_id = 786,use_gpu=True,silent = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"compare_models()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lightgbm = create_model('lightgbm')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# !pip install lightgbm ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Tune model using scikit-learn auto"},{"metadata":{"trusted":true},"cell_type":"code","source":"tuned_lightgbm = tune_model(lightgbm,optimize='AUC',n_iter=20)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Tune model using scikit-optimize"},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install scikit-optimize","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tuned_lightgbm_auto = tune_model(lightgbm,optimize='AUC',n_iter=25,search_library=\"scikit-optimize\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Ensemble Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"bagged_lightgbm = ensemble_model(tuned_lightgbm_auto)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"boosted_lightgbm = ensemble_model(tuned_lightgbm_auto ,method = 'Boosting')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Blend Models"},{"metadata":{"trusted":true},"cell_type":"code","source":"blender = blend_models(estimator_list = [tuned_lightgbm,tuned_lightgbm_auto], method = 'soft')\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Stack model "},{"metadata":{"trusted":true},"cell_type":"code","source":"# stacker = stack_models(estimator_list = [tuned_lightgbm,tuned_lightgbm_auto], meta_model=lightgbm)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analyze Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Best Model : tuned_lightgbm : AUC : 66\n\nplot_model(tuned_lightgbm)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:black; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Conclusion </center></h2> \n\n<div align='left'><font size=\"3\" color=\"#000000\">LGBMClassifier perform better compare to other model.</font></div>\n"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='font-size:30px;background:black; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Things to be taken care before submission </center></h2> \n<div align='left'><font size=\"3\" color=\"#000000\">This competition is different from most of other Kaggle Competitions.You will loop through a series of batches of questions. Once you make that prediction, you can move on to the next batch, you will receive test set data and make predictions with Kaggle's time-series API. So it is good if you refer these kerenel before submission</font></div> \n<hr> \n<div align='left'><font size=\"3\"><a href=\"https://www.kaggle.com/sohier/competition-api-detailed-introduction\" target=\"_blank\">1. Competition API Detailed Introduction</a></div>\n<div align='left'><font size=\"3\"><a href=\"https://www.kaggle.com/sohier/quick-sample-submission\" target=\"_blank\">2. Quick Sample Submission</a></div>\n<hr>   "}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}