{"cells":[{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"%%HTML\n\n<style type=\"text/css\">\ndiv.h1 {\n    background-color:#339933; \n    color: white; \n    padding: 8px; \n    padding-right: 300px; \n    font-size: 35px; \n    max-width: 1500px; \n    margin: auto; \n    margin-top: 50px;\n}\n\ndiv.h2 {\n    background-color:#83ccd2; \n    color: white; \n    padding: 8px; \n    padding-right: 300px; \n    font-size: 35px; \n    max-width: 1500px; \n    margin: auto; \n    margin-top: 50px;\n}\n</style>","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# <div class=h1>About this notebook</div>"},{"metadata":{},"cell_type":"markdown","source":"In this notebook, I'll load train data and make sample plot to get good insight for data.\n\nIn version1, to begin with, the data was too large to read and visualize, so I aimed to read and visualize it first. This notebook is going to update.\n\nIn this competition, your challenge is to create algorithms for \"Knowledge Tracing,\" the modeling of student knowledge over time. The goal is to accurately predict how students will perform on future interactions.\n\nOur innovative algorithms will help tackle global challenges in education. If successful, it’s possible that any student with an Internet connection can enjoy the benefits of a personalized learning experience, regardless of where they live. "},{"metadata":{},"cell_type":"markdown","source":"\n\n<img src=\"https://storage.googleapis.com/kaggle-media/competitions/Riiid/Graphic%20or%20image%20within%20description%20(min%20size%20350x350).png\" width=\"300\">"},{"metadata":{},"cell_type":"markdown","source":"# <div class=h2> visualization</div>"},{"metadata":{},"cell_type":"markdown","source":"### load library."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import math\nimport matplotlib.pyplot as plt\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!ls ../input/riiid-test-answer-prediction/","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Load data\n\nI'll load data. Especially, train.csv is so big (5.45 GB !!!), we have to get creative.\n\nFirst, I specified a partial dtype. And second, I used reduce_mem_usage function. This is knoladge from \"ASHRAE - Great Energy Predictor III\" comp."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/riiid-test-answer-prediction/train.csv\",\n                    dtype = {\"content_type_id\":\"int8\", \"task_container_id\":\"int32\",\n                             \"user_answer\":\"int8\", \"answered_correctly\":\"int8\"})\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#From https://www.kaggle.com/rohanrao/ashrae-half-and-half\n\nfrom pandas.api.types import is_datetime64_any_dtype as is_datetime\nfrom pandas.api.types import is_categorical_dtype\n\ndef reduce_mem_usage(df, use_float16=False):\n    \"\"\"\n    Iterate through all the columns of a dataframe and modify the data type to reduce memory usage.        \n    \"\"\"\n    \n    start_mem = df.memory_usage().sum() / 1024**2\n    print(\"Memory usage of dataframe is {:.2f} MB\".format(start_mem))\n    \n    for col in df.columns:\n        if is_datetime(df[col]) or is_categorical_dtype(df[col]):\n            continue\n        col_type = df[col].dtype\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == \"int\":\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if use_float16 and c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            df[col] = df[col].astype(\"category\")\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print(\"Memory usage after optimization is: {:.2f} MB\".format(end_mem))\n    print(\"Decreased by {:.1f}%\".format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = reduce_mem_usage(train, use_float16=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions = pd.read_csv(\"../input/riiid-test-answer-prediction/questions.csv\")\nlectures = pd.read_csv(\"../input/riiid-test-answer-prediction/lectures.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.dtypes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.dtypes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.dtypes","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Simple visualization\n\nI'll check some features of dat by seaborn."},{"metadata":{},"cell_type":"markdown","source":"### train\n\nThere are too much data, so I focus on user_id == 115."},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t1 = sns.distplot(train[train[\"user_id\"]==115][\"timestamp\"])\ng_t1.set_title(\"distplot of timestamp of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 5))\ng_t2 = sns.countplot(data=train[train[\"user_id\"]==115],x=\"content_id\")\n#g_t2 = sns.distplot(train[train[\"user_id\"]==115][\"content_id\"])\ng_t2.set_title(\"distplot of content_id of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t3 = sns.countplot(data=train[train[\"user_id\"]==115],x=\"content_type_id\")\ng_t3.set_title(\"distplot of content_type_id of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t4 = sns.countplot(data=train[train[\"user_id\"]==115],x=\"user_answer\")\ng_t4.set_title(\"distplot of user_answer of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t5 = sns.countplot(data=train[train[\"user_id\"]==115],x=\"answered_correctly\")\ng_t5.set_title(\"distplot of answered_correctly of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t6 = sns.distplot(train[train[\"user_id\"]==115][\"prior_question_elapsed_time\"])\ng_t6.set_title(\"distplot of prior_question_elapsed_time of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_t7 = sns.countplot(data=train[train[\"user_id\"]==115],x=\"prior_question_had_explanation\")\ng_t7.set_title(\"distplot of prior_question_had_explanation of user_id=115\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#it takes so much time and memory...\n#g2 = sns.countplot(data=train, x=\"user_id\")\n#g2.set_title(\"countplot of user_id\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### questions"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 5))\ng_q1 = sns.countplot(data=questions, x=\"question_id\")\ng_q1.set_title(\"countplot of question_id\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_q2 = sns.countplot(data=questions, x=\"bundle_id\")\ng_q2.set_title(\"countplot of bundle_id\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_q3 = sns.countplot(data=questions, x=\"correct_answer\")\ng_q3.set_title(\"countplot of correct_answer\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_q4 = sns.countplot(data=questions, x=\"part\")\ng_q4.set_title(\"countplot of part\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 5))\ng_q5 = sns.countplot(data=questions, x=\"tags\")\ng_q5.set_title(\"countplot of tags\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### lectures"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 5))\ng_l1 = sns.countplot(data=lectures, x=\"lecture_id\")\ng_l1.set_title(\"countplot of lecture_id\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_l3 = sns.countplot(data=lectures, x=\"part\")\ng_l3.set_title(\"countplot of part\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_l4 = sns.distplot(lectures[\"tag\"],bins=50)\ng_l4.set_title(\"countplot of tag\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"g_l5 = sns.countplot(data=lectures, x=\"type_of\")\ng_l5.set_title(\"countplot of type_of\")","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}