{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Riiid classification using Naive Bayes"},{"metadata":{},"cell_type":"markdown","source":"The initial purpose of this notebook was to make a very simple baseline for my own usage to better understand the data and the submission process. As I got a fairly decent score (given the simplicity of the model) I decided to share the notebook in case it can be helpful to someone else. This model could be improved by adding more features."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Dense\nfrom numpy import loadtxt\nimport tensorflow as tf\nfrom sklearn.model_selection import train_test_split\n\nimport pandas as pd\nimport numpy as np\nimport riiideducation\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Read the data\n<p>I read the data in feather format from this notebook https://www.kaggle.com/aralai/riiid-feather-dataset. It's much faster! \nIf you don't want to use feather, you can just replace the following lines with the csv format reading."},{"metadata":{"trusted":true},"cell_type":"code","source":"#%%time\ntrain = pd.read_feather('../input/riiid-feather-dataset/train.feather')\nquestions = pd.read_feather('../input/riiid-feather-dataset/questions.feather')\nlectures = pd.read_feather('../input/riiid-feather-dataset/lectures.feather')\nexample_test = pd.read_feather('../input/riiid-feather-dataset/example_test.feather')\nexample_sample_submission = pd.read_feather('../input/riiid-feather-dataset/example_sample_submission.feather')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#df = train[\"user_id\",\"content_id\",\"answered_correctly\"]\n\n# Try to use/make different features user_id and content_id are probably not super informative to models.\ndf = train[['user_id','content_id','answered_correctly']]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nX = df.iloc[:,0:2]\n\ny = df.iloc[:,2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.01, random_state = 0)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential()\nmodel.add(Dense(12, input_dim=2, activation='relu'))\nmodel.add(Dense(8, activation='relu'))\nmodel.add(Dense(1, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[tf.keras.metrics.AUC()]) #ERFAN: added auc since that's what we are optimizing for - loss seems static.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.fit(X_train, y_train, epochs=2, batch_size=50) #Erfan: AUC seems static, loss decreases when batch_size < 100. Why do we want a bigger batchsize?","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions = model.predict_classes(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Submission<p>\n<code>example_test</code> contains just a few dummy rows that can be used for development. The real submission must be made by using <code>riiideducation</code> package. To do that, we must create an environment and then loop through all the batches provided by <code>iter_test</code>."},{"metadata":{"trusted":true},"cell_type":"code","source":"env = riiideducation.make_env()\niter_test = env.iter_test()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"For each batch in <code>test_df</code>, we will predict the probability of answering correctly (<code>nb.predict</code>) and then we will send the resulting data back to the environment. This last part is done in <code>env.predict</code>. Notice that we must not create any <code>submission.csv</code> file, this is done automatically by <code>env.predict</code>."},{"metadata":{"trusted":true},"cell_type":"code","source":"for (test_df, sample_prediction_df) in iter_test:\n    test_questions = test_df['content_id']\n    test_users = test_df['user_id']\n    answered_correctly = nb.predict({'question':test_questions, 'user':test_users})\n    test_df['answered_correctly'] = answered_correctly\n    env.predict(test_df.loc[test_df['content_type_id']==0,['row_id','answered_correctly']])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}