{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Riiid! Answer Correctness Prediction. Data Analysis and visualization and Modeling"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport plotly.express as px \n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#Number of records"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!wc -l ../input/riiid-test-answer-prediction/train.csv","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Train.csv**\n* row_id: (int64) ID code for the row.\n\n* timestamp: (int64) the time between this user interaction and the first event from that user.\n\n* user_id: (int32) ID code for the user.\n\n* content_id: (int16) ID code for the user interaction\n\n* content_type_id: (int8) 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture.\n\n* task_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id. Monotonically increasing for each user.\n\n* user_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n\n* answered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.\n\n* prior_question_elapsed_time: (float32) How long it took a user to answer their previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Note that the time is the total time a user took to solve all the questions in the previous bundle.\n\n* prior_question_had_explanation: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nimport riiideducation\n#tr = pd.read_csv(\"../input/riiid-test-answer-prediction/train.csv\", low_memory=False, nrows=10**7)\n#tr = pd.read_csv(\"../input/riiid-test-answer-prediction/train.csv\")\n#tr.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n%%time\n\ndtypes = {\n    \"row_id\": \"int64\",\n    \"timestamp\": \"int64\",\n    \"user_id\": \"int32\",\n    \"content_id\": \"int16\",\n    \"content_type_id\": \"boolean\",\n    \"task_container_id\": \"int16\",\n    \"user_answer\": \"int8\",\n    \"answered_correctly\": \"int8\",\n    \"prior_question_elapsed_time\": \"float32\", \n    \"prior_question_had_explanation\": \"boolean\"\n}\n\ntr = pd.read_csv(\"../input/riiid-test-answer-prediction/train.csv\", dtype=dtypes, low_memory=False, nrows=30000000)\n\nprint(\"Train size:\", tr.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Print missing value\nprint('Part of missing values for every column')\nprint(tr.isnull().sum() / len(tr))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nprint(tr.answered_correctly.value_counts())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"CATS = [\"content_id\",  \"task_container_id\",\"prior_question_had_explanation\"]\nNUM = [\"prior_question_elapsed_time\"]\nASIDE = [\"user_id\", \"content_type_id\",\"user_answer\"]\nTARGET = \"answered_correctly\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for col in CATS:\n    print(col, tr[col].nunique(), tr[col].max())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TIME_MEAN = tr.prior_question_elapsed_time.median()\nTIME_MIN = tr.prior_question_elapsed_time.min()\nTIME_MAX = tr.prior_question_elapsed_time.max()\nprint(TIME_MEAN,TIME_MAX, TIME_MIN)\nmap_prior = {True:1, False:0}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\npiv1 = (tr.loc[tr.answered_correctly!=-1, [\"content_id\", \"answered_correctly\"]].\n        groupby(\"content_id\").agg([\"mean\", \"sum\"])).reset_index()\npiv1.columns = [\"content_id\", \"content_emb\", \"content_sum\"]\npiv2 = (tr.loc[tr.answered_correctly!=-1, [\"task_container_id\", \"answered_correctly\"]].\n        groupby(\"task_container_id\").agg([\"mean\", \"sum\"])).reset_index()\npiv2.columns = [\"task_container_id\", \"task_container_emb\", \"task_container_sum\"]\npiv3 = (tr.loc[tr.answered_correctly!=-1, [\"user_id\", \"answered_correctly\"]].\n        groupby(\"user_id\").agg([\"mean\", \"sum\"])).reset_index()\npiv3.columns = [\"user_id\", \"user_emb\", \"user_sum\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nfor col, df in zip([\"content_sum\", \"task_container_sum\", \"user_sum\"], [piv1, piv2, piv3]):\n    df[col] = (df[col] - df[col].min()) / (df[col].max() - df[col].min())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"m1 = piv1[\"content_sum\"].median()\nm2 = piv2[\"task_container_sum\"].median()\nm3 = piv3[\"user_sum\"].median()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def preprocess(df):\n    df = df.merge(piv1, how=\"left\", on=\"content_id\")\n    df[\"content_emb\"] = df[\"content_emb\"].fillna(0.5)\n    df[\"content_sum\"] = df[\"content_sum\"].fillna(m1)\n    df = df.merge(piv2, how=\"left\", on=\"task_container_id\")\n    df[\"task_container_emb\"] = df[\"task_container_emb\"].fillna(0.5)\n    df[\"task_container_sum\"] = df[\"task_container_sum\"].fillna(m2)\n    df = df.merge(piv3, how=\"left\", on=\"user_id\")\n    df[\"user_emb\"] = df[\"user_emb\"].fillna(0.5)\n    df[\"user_sum\"] = df[\"user_sum\"].fillna(m3)\n    df[\"prior_question_elapsed_time\"] = df[\"prior_question_elapsed_time\"].fillna(TIME_MEAN)\n    df[\"duration\"] = (df[\"prior_question_elapsed_time\"] - TIME_MIN) / (TIME_MAX - TIME_MIN)\n    df[\"prior_answer\"] = df[\"prior_question_had_explanation\"].map(map_prior)\n    df[\"prior_answer\"] = df[\"prior_answer\"].fillna(0.5)\n    #df = df.fillna(-1)\n    epsilon = 1e-6\n    df[\"score\"] = 2*df[\"content_emb\"]*df[\"user_emb\"] / (df[\"content_emb\"]+ df[\"user_emb\"] + epsilon)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntr = preprocess(tr)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tr.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"FE = [\"content_emb\",\"content_sum\" ,\"task_container_emb\", \"task_container_sum\",\n      \"user_emb\", \"user_sum\",\"duration\", \"prior_answer\",\"score\"]\nx = tr.loc[tr.answered_correctly!=-1, FE].values\ny = tr.loc[tr.answered_correctly!=-1, TARGET].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Free memory\ndel tr","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n#X_train, X_test, y_train, y_test = train_test_split(x, y, test_size = 0.3, random_state = 0)\nX_train, X_valid, y_train, y_valid = train_test_split(x,y,\n                                                      test_size=0.3, \n                                                      stratify = y,\n                                                      random_state=7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from xgboost import XGBClassifier","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"params = {              'n_estimators': 100,\n                        'seed': 44,\n                        'colsample_bytree': 0.8,\n                        'subsample': 0.7,\n                        'learning_rate': 0.01,\n                        'objective': 'binary:logistic',\n                        'max_depth': 5,\n                        'num_parallel_tree': 1000,\n                        'min_child_weight': 20,\n                        'eval_metric':'auc',\n                        'gamma':0.1,\n                        'tree_method':'gpu_hist'}\n\nmodel = XGBClassifier(**params)\neval_set = [(X_train, y_train), (X_valid, y_valid)]\nmodel.fit(X_train, y_train, early_stopping_rounds=5, eval_metric=\"auc\", eval_set=eval_set, verbose=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#PREDICTION","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"env = riiideducation.make_env()\niter_test = env.iter_test()\n#it = 0\nfor test_df, sample_prediction_df in iter_test:\n    #it += 1\n    #if it % 100 == 0:\n    #    print(it)\n    test_df = preprocess(test_df)\n    x_te = test_df[FE].values\n    test_df['answered_correctly'] = model.predict(x_te)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}