{"cells":[{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# Introduction\nIn this competition you need to predict how well each student will answer the questions. This is a timeseries classification problem using ROC AUC metric.\n\n\nThe custom riiideducation Python module is provided to use for inference. It will ensure that the test set data will be made avaialble as a series of batches including questions or lectures. After predicting for the given batch, you will get the next batch. Find more details [here](https://www.kaggle.com/sohier/competition-api-detailed-introduction)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)import os\nfrom pathlib import Path\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"class config:\n    PATH = Path(\"/kaggle/input/riiid-test-answer-prediction\")\n    \n    dtype = {'row_id': 'int64', \n             'timestamp': 'int64', \n             'user_id': 'int32', \n             'content_id': 'int16', \n             'content_type_id': 'int8',\n             'task_container_id': 'int16',\n             'user_answer': 'int8', \n             'answered_correctly': 'int8', \n             'prior_question_elapsed_time': 'float32', \n             'prior_question_had_explanation': 'boolean',\n             }\n    \n    LINE_WIDTH = 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## What you need to predict?"},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = pd.read_csv(config.PATH/\"example_sample_submission.csv\")\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load training data"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(config.PATH/'train.csv', low_memory=False, nrows=10**5, \n                       dtype=config.dtype\n                      )\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"f\"Number of Unique students in {train_df.shape[0]} samples are {train_df['user_id'].nunique()}\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def apply_plot_layout(fig, feature, annot=\"\", annot_size=60, y_title=\"\", title=\"\", tickangle=-90, unified=True):\n    fig.update_layout(\n        hovermode='x unified' if unified else 'x',\n        title=title,\n        xaxis= {\"tickangle\":tickangle,\n                \"showgrid\":False,\n                \"showline\":False,\n                \"gridwidth\":.1,\n                \"zeroline\":False,\n                },\n        yaxis= {\"showline\":False,\n                \"gridcolor\":'rgba(203, 210, 211,.3)',\n                \"gridwidth\":.1,\n                \"zeroline\":False,\n                \"title\":y_title\n                },\n        #xaxis_title=\"Toggle the legends to show/hide corresponding curve\",\n        plot_bgcolor='#ffffff',\n        paper_bgcolor='#ffffff',\n    ) \n    return fig","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# user_id\n: (int32) ID code for the user."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'user_id'\nagg_fun = 'count'\ndf = train_df.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace1 = go.Bar(x=df[feature].astype(str) + \"-\",\n                y=df[agg_fun],\n                hovertext=['{} : {},\\n{} : {:,d}'.format(feature, id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"Number of interactions by each user\")\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# content_id\n: (int16) ID code for the user interaction"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"feature = 'content_id'\nagg_fun = 'count'\ndf = train_df.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace1 = go.Bar(x=df[feature].astype(str) + \"-\",\n                y=df[agg_fun],\n                hovertext=['{} : {},\\n{} : {:,d}'.format(feature, id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"Number of interactions for each content\")\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# content_type_id\n: (int8) 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"labels = [\"solved question\", 'watched lecture']\nvalues = train_df.content_type_id.value_counts().values\n\nfig = go.Figure(data=[go.Pie(labels=labels, values=values, textinfo='label+percent',\n                             insidetextorientation='radial'\n                            )])\n\nfig.update_layout(\n    title_text=\"Questions vs Lectures\",\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# task_container_id\n: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'task_container_id'\nagg_fun = 'count'\ndf = train_df.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace1 = go.Bar(x=df[feature].astype(str) + \"-\",\n                y=df[agg_fun],\n                hovertext=['{} : {},\\n{} : {:,d}'.format(\"Batch id\", id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\nfig.update_traces(marker_color='rgb(158,202,225)', marker_line_color='rgb(8,48,107)',\n                  marker_line_width=1.5, opacity=0.6)\n\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"Number of interactions per batch\")\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# user_answer\n: (int8) the user's answer to the question, if any. Read -1 as null, for lectures."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'user_answer'\nagg_fun = 'count'\ndf = train_df.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace1 = go.Bar(x=df[feature],\n                y=df[agg_fun],\n                hovertext=['{} : {:,d},\\n{} : {:,d}'.format(\"User's answer\", id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\nfig.update_traces(marker_color='rgb(158,202,225)', marker_line_color='rgb(8,48,107)',\n                  marker_line_width=1.5, opacity=0.6)\n\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"User's answer to the question\", tickangle=0, unified=False)\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# answered_correctly\n: (int8) if the user responded correctly. Read -1 as null, for lectures."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'answered_correctly'\nagg_fun = 'count'\ndf = train_df[train_df[feature] != -1].groupby([feature])[feature].agg([agg_fun]).reset_index()\n\nis_correct = {-1: \"Watching lecture\", 0:'Wrong', 1:'Correct'}\n\ntrace1 = go.Bar(x=df[feature],\n                y=df[agg_fun],\n                hovertext=['{} answer,\\n{} : {:,d}'.format(is_correct[id],\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\n\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"Answered correctly ?\", tickangle=0, unified=False)\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'user_answer'\nagg_fun = 'count'\nwrongly_answered = train_df.query('user_answer != -1 and answered_correctly == 0')\ndf = wrongly_answered.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace1 = go.Bar(x=df[feature],\n                y=df[agg_fun],\n                name=\"Wrong\",\n                hovertext=['{} : {:,d},\\n{} : {:,d}'.format(\"User's answer\", id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\ncorrectly_answered = train_df.query('user_answer != -1 and answered_correctly == 1')\ndf = correctly_answered.groupby([feature])[feature].agg([agg_fun]).reset_index()\n\ntrace2 = go.Bar(x=df[feature],\n                y=df[agg_fun],\n                name=\"Correct\",\n                hovertext=['{} : {:,d},\\n{} : {:,d}'.format(\"User's answer\", id,\n                            agg_fun, c) for id, c in zip(df[feature], df[agg_fun])],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace2, trace1])\n#fig.update_traces(marker_color='rgb(158,202,225)', marker_line_color='rgb(8,48,107)',\n#                  marker_line_width=1.5, opacity=0.6)\n\nfig = apply_plot_layout(fig, feature=feature, annot_size=60, y_title=agg_fun, title=\"Correct vs Wrong proportion for each User's answer\", tickangle=0, unified=False)\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# prior_question_elapsed_time\n: (float32) The average time it took a user to answer each question in the previous question bundle, ignoring any lectures in between. Is null for a user's first question bundle or lecture. Note that the time is the average time a user took to solve each question in the previous bundle."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"feature = 'user_id'\nagg_fun = 'unique'\ndf = train_df.groupby(['user_id'])['timestamp', 'task_container_id', 'prior_question_elapsed_time'].agg([agg_fun]).reset_index()\n\nall_traces = []\nTOP = None\nfor i in range(len(df)):\n    user_id = str(df.iloc[i]['user_id'].values[0])\n    x = df.iloc[i]['timestamp']['unique']\n    y = df.iloc[i]['prior_question_elapsed_time']['unique']\n    trace = go.Scatter(x=x,\n                    y=y,\n                    text=x,\n                    name=f'user {user_id}',\n                    mode='markers+lines',\n                    line={\"width\":config.LINE_WIDTH, \"shape\":'spline'},\n                    hovertext=['user {} took time : {}'.format(user_id, tk) for tk in y],\n                    hovertemplate='%{hovertext}' +\n                                '<extra></extra>'\n\n    )\n    all_traces.append(trace)\n    if TOP and i > TOP:\n        break\n    \nfig = go.Figure(data=all_traces)\nfig.update_layout(\n    \n    hovermode='x unified',\n    title='Average time taken for solving prior question bundle',\n    xaxis= {\n                #\"tickangle\":tickangle,\n                \"showgrid\":False,\n                \"showline\":False,\n                \"gridwidth\":.1,\n                \"zeroline\":False,\n                \"title\":\"Timestamp\"\n                },\n    yaxis= {    \"showline\":False, #linecolor='#272e3e',\n                \"gridcolor\":'rgba(203, 210, 211,.3)',\n                \"gridwidth\":.1,\n                \"zeroline\":False,\n                \"title\":'prior question elapsed time'\n                },\n    #xaxis_title=\"Toggle the legends to show/hide corresponding curve\",\n    plot_bgcolor='#ffffff',\n    paper_bgcolor='#ffffff',\n) \n\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# prior_question_had_explanation\n: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df = train_df[train_df['answered_correctly'] != -1].prior_question_had_explanation.value_counts().reset_index()\n\ntrace1 = go.Bar(x=df[\"index\"],\n                y=df.prior_question_had_explanation,\n                hovertext=['{} : {:,d}'.format(id, c) for id, c in zip(df[\"index\"], df.prior_question_had_explanation)],\n                hovertemplate='%{hovertext}' +\n                            '<extra></extra>'\n)\n\nfig = go.Figure(data=[trace1])\n#fig.update_traces(marker_color='rgb(158,202,225)', marker_line_color='rgb(8,48,107)',\n#                  marker_line_width=1.5, opacity=0.6)\n\nfig = apply_plot_layout(fig, feature='prior_question_had_explanation', annot_size=60, y_title='Count', \n                        title=\"Seen explanation for prior question (ignoring lectures)?\", tickangle=0,\n                        unified = False)\nfig.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Thanks for reading. Work in progess ..."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}