{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# Riiid! Answer Correctness Prediction"},{"metadata":{},"cell_type":"markdown","source":"The main target of this notebook is giving the base understanding of our data and some useful features.\n\nFirst of all you can find here:\n\n>Comprehensive description of our data\n\n>Feature Engineering\n\n>Baseline with LightGBM without serious tuning model's parameters"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"### **Important: I use here just a part of the full dataset because of the limited RAM**"},{"metadata":{"trusted":true},"cell_type":"code","source":"import warnings\nwarnings.simplefilter('ignore')\n\nimport pandas as pd\nfrom pandas.plotting import scatter_matrix\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.metrics import roc_auc_score, confusion_matrix\nfrom sklearn.model_selection import train_test_split, GridSearchCV, RandomizedSearchCV, learning_curve\nfrom sklearn.utils import shuffle\nimport lightgbm as lgb\nfrom lightgbm import LGBMClassifier\nimport eli5\n\nimport riiideducation\n\n%matplotlib inline\n# for heatmap and other plots\ncolorMap1 = sns.color_palette(\"RdBu_r\")\n# for countplot and others plots\ncolorMap2 = 'Blues_r'\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input/riiid-test-answer-prediction'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### PATHS"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_path = \"../input/riiid-train-pickled-data/train.pkl.gzip\"\nquestions_path = \"../input/riiid-test-answer-prediction/questions.csv\"\nlectures_path = \"../input/riiid-test-answer-prediction/lectures.csv\"\n\ntest = \"../input/riiid-test-answer-prediction/example_test.csv\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"1\"></a>\n\n<div style=\"background:write; border:0; color:black; width: 100%; height: 50px\">\n    <div style=\"vertical-align: middle; text-align:center\"><h1>DATA EXPLORATION & EDA</h1></div>\n</div>"},{"metadata":{},"cell_type":"markdown","source":"### TRAIN"},{"metadata":{},"cell_type":"markdown","source":">**row_id**: (int64) ID code for the row.\n\n>**timestamp**: (int64) the time in milliseconds between this user interaction and the first event completion from that user.\n\n>**user_id**: (int32) ID code for the user.\n\n>**content_id**: (int16) ID code for the user interaction\n\n>**content_type_id**: (int8) 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture.\n\n>**task_container_id**: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.\n\n>**user_answer**: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n\n>**answered_correctly**: (int8) if the user responded correctly. Read -1 as null, for lectures.\n\n>**prior_question_elapsed_time**: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. Is null for a user's first question bundle or lecture. Note that the time is the average time a user took to solve each question in the previous bundle.\n\n>**prior_question_had_explanation**: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback.\n"},{"metadata":{},"cell_type":"markdown","source":"I'll use a pickle file prepared by Rohan Rao in this kernel: https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets/ \n\nThanks a lot!"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\ndf = pd.read_pickle(train_path)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"Train shape: {df.shape}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.memory_usage()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Drop **row_id** and **timestamp** because they look like useless and take a lot of memory"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.drop(['row_id', 'timestamp'], axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.describe().style.background_gradient(cmap='Blues')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f'Number of unique users: {len(np.unique(df.user_id))}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look at the missing data"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(df.isnull().sum() / len(df))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also let's check a correlation matrix to get more information between the columns"},{"metadata":{"trusted":true},"cell_type":"code","source":"corr_matrix = df.corr()\ncorr_matrix[\"answered_correctly\"].sort_values(ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(13, 10))\nsns.heatmap(corr_matrix, annot=True, \n            linewidths=.5, cmap=colorMap1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check the time distribution for **prior_question_elapsed_time**"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 10))\nax = sns.countplot(x=\"prior_question_elapsed_time\", \n                   data=df[df['prior_question_elapsed_time'].notnull()],\n                   palette=colorMap2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check any connection between our target value and a frequency of answering questions"},{"metadata":{"trusted":true},"cell_type":"code","source":"freq_answered_tasks = df['task_container_id'].value_counts().reset_index()\nfreq_answered_tasks.columns = [\n    'task_container_id', \n    'freq'\n]\n\ndf['freq_task_id'] = ''\ndf.loc[df['task_container_id'].isin(freq_answered_tasks[freq_answered_tasks['freq'] < 10000]['task_container_id'].values), 'freq_task_id'] = 'very rare answered'\ndf.loc[df['task_container_id'].isin(freq_answered_tasks[freq_answered_tasks['freq'] >= 10000]['task_container_id'].values), 'freq_task_id'] = 'rare answered'\ndf.loc[df['task_container_id'].isin(freq_answered_tasks[freq_answered_tasks['freq'] >= 50000]['task_container_id'].values), 'freq_task_id'] = 'normal answered'\ndf.loc[df['task_container_id'].isin(freq_answered_tasks[freq_answered_tasks['freq'] >= 200000]['task_container_id'].values), 'freq_task_id'] = 'often answered'\ndf.loc[df['task_container_id'].isin(freq_answered_tasks[freq_answered_tasks['freq'] >= 400000]['task_container_id'].values), 'freq_task_id'] = 'very often answered'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.sample(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 10))\nsns.countplot(x='freq_task_id', hue='answered_correctly', data=df, palette=colorMap2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**prior_question_had_explanation** has a medium correlation with the target value. So let's look at his distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 11))\nax = sns.countplot(x=\"prior_question_had_explanation\", hue=\"answered_correctly\", data=df[df['prior_question_had_explanation'].notnull()], palette=colorMap2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The answer's showing increases the probability of a successful answering. Let's go further.\n\nCheck the most active **user_id**"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 30 # number of users\n\nuser_freq = df['user_id'].value_counts().reset_index()\nuser_freq.columns = [\n    'user_id', \n    'count'\n]\n\n# Add ' - ' to convert user_id to str and not sort\nuser_freq['user_id'] = user_freq['user_id'].astype(str) + ' - '\nuser_freq = user_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 15))\nsns.barplot(x='count', y='user_id', data=user_freq, orient='h', palette=colorMap2)\nplt.title(f'Top {N} the most active users', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And the most useful **content_id**"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 30 # number of users\n\ncontent_id_freq = df['content_id'].value_counts().reset_index()\ncontent_id_freq.columns = [\n    'content_id', \n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\ncontent_id_freq['content_id'] = content_id_freq['content_id'].astype(str) + ' - '\ncontent_id_freq = content_id_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 15))\nsns.barplot(x='count', y='content_id', data=content_id_freq, orient='h', palette=colorMap2)\nplt.title(f'Top {N} the most useful content_id', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also we need to check distribution in **content_type_id**: number of video lectures and questions"},{"metadata":{"trusted":true},"cell_type":"code","source":"content_type_freq = df['content_type_id'].value_counts().reset_index()\ncontent_type_freq.columns = ['content_type_id',\n                             'share']\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='content_type_id', y='share', data=content_type_freq,\n            palette=colorMap2)\nplt.title('Share of the Questions (0) and Video lectures (1)', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Exclude lectures with -1 in **answered_correctly** col"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df[df['answered_correctly'] != -1].reset_index(drop=True, inplace=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look where is the most part of the incorrect answers"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.groupby(['content_type_id', 'answered_correctly']).agg({'answered_correctly': 'count'})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So you can see we don't have any correct answers for the lectures.\n\nNext check the distribution of the **task_container_id**"},{"metadata":{"trusted":true},"cell_type":"code","source":"task_ids_freq = df['task_container_id'].value_counts().reset_index()\ntask_ids_freq.columns = ['task_container_id', 'count']\n\nfig, ax = plt.subplots(figsize=(15, 10))\n\nsns.pointplot(x='task_container_id', y='count', data=task_ids_freq, palette=colorMap2)\nxticks_range = range(min(task_ids_freq['task_container_id']), \n                     max(task_ids_freq['task_container_id']),\n                     1000)\nplt.xticks(list(xticks_range), list(xticks_range))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look at the correct answers distribution in the user's answers"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15, 11))\nax = sns.countplot(x=\"user_answer\", hue=\"answered_correctly\", data=df, palette=colorMap2)\n\nfor i, p in enumerate(ax.patches):\n    x = p.get_bbox().get_points()[:,0]\n    y = p.get_bbox().get_points()[1,1]\n    if i > 3:\n        i -= 4\n    ax.annotate('{:.1f}%'.format(100.*y/len(np.where(df['user_answer'] == i)[0])), (x.mean(), y), \n                ha='center', va='bottom') # set the alignment of the text","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### QUESTIONS.CSV"},{"metadata":{},"cell_type":"markdown","source":">**question_id**: foreign key for the train/test content_id column, when the content type is question (0).\n\n>**bundle_id**: code for which questions are served together.\n\n>**correct_answer**: the answer to the question. Can be compared with the train user_answer column to check if the user was right.\n\n>**part**: top level category code for the question.\n\n>**tags**: one or more detailed tag codes for the question. The meaning of the tags will not be provided, but these codes are sufficient for clustering the questions together."},{"metadata":{"trusted":true},"cell_type":"code","source":"questions = pd.read_csv(questions_path)\nquestions.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.describe().style.background_gradient(cmap='Blues')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(questions.isnull().sum() / len(questions))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check the parts distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"part_freq = questions['part'].value_counts().reset_index()\npart_freq.columns = [\n    'part', \n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\npart_freq['part'] = part_freq['part'].astype(str) + ' - '\npart_freq = part_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='count', y='part', data=part_freq, orient='h', palette=colorMap2)\nplt.title(f'The most frequent parts', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And next check the tags distribution (without splitting the group of the tags)"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 30\n\ntags_freq = questions['tags'].value_counts().reset_index()\ntags_freq.columns = [\n    'tag',\n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\ntags_freq['tag'] = tags_freq['tag'].astype(str) + ' - '\ntags_freq = tags_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='count', y='tag', data=tags_freq, orient='h', palette=colorMap2)\nplt.title(f'Top {N} the most frequent tags', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And with splitting the group of the tags"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 30\n\ntags = questions['tags'].str.split(' ').explode('tags').reset_index()\ntags_freq = tags['tags'].value_counts().reset_index()\ntags_freq.columns = [\n    'tag',\n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\ntags_freq['tag'] = tags_freq['tag'].astype(str) + ' - '\ntags_freq = tags_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='count', y='tag', data=tags_freq, orient='h', palette=colorMap2)\nplt.title(f'Top {N} the most frequent tags', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### LECTURES.CSV"},{"metadata":{},"cell_type":"markdown","source":"Metadata for the lectures watched by users as they progress in their education.\n\n>**lecture_id**: foreign key for the train/test content_id column, when the content type is lecture (1).\n\n>**part**: top level category code for the lecture.\n\n>**tag**: one tag codes for the lecture. The meaning of the tags will not be provided, but these codes are sufficient for clustering the lectures together.\n\n>**type_of**: brief description of the core purpose of the lecture"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures = pd.read_csv(lectures_path)\nlectures.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also let's check the part distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"part_freq = lectures['part'].value_counts().reset_index()\npart_freq.columns = [\n    'part', \n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\npart_freq['part'] = part_freq['part'].astype(str) + ' - '\npart_freq = part_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='count', y='part', data=part_freq, orient='h', palette=colorMap2)\nplt.title(f'The most frequent parts', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The most frequent tags of lectures"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 30\n\ntags_freq = lectures['tag'].value_counts().reset_index()\ntags_freq.columns = [\n    'tag',\n    'count'\n]\n\n# Add ' - ' to convert content_id to str and not sort\ntags_freq['tag'] = tags_freq['tag'].astype(str) + ' - '\ntags_freq = tags_freq.sort_values(['count'], ascending=False).head(N)\n\nplt.figure(figsize=(15, 10))\nsns.barplot(x='count', y='tag', data=tags_freq, orient='h', palette=colorMap2)\nplt.title(f'Top {N} the most frequent tags', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And finally check the **type_of** column"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures['type_of'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### EXAMPLE-TEST.CSV"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_example = pd.read_csv(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_example.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"2\"></a>\n\n<div style=\"background:write; border:0; color:black; width: 100%; height: 50px\">\n    <div style=\"vertical-align: middle; text-align:center\"><h1>FEATURE ENGINEERING</h1></div>\n</div>"},{"metadata":{},"cell_type":"markdown","source":"I'll give just some part from our data bacause of the RAM limit on Kaggle kernel"},{"metadata":{"trusted":true},"cell_type":"code","source":"n = int(df.shape[0] * 0.1)\ntrain = df.sample(n=n, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del questions\ndel lectures","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_characteristics = df.groupby('user_id').agg({'answered_correctly':\n                                                  ['mean', 'median', 'std', 'skew', 'count']})\nuser_characteristics.columns = [\n    'mean_user_acc',\n    'median_user_acc',\n    'std_user_acc',\n    'skew_user_acc',\n    'number_of_answered_q'\n]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_characteristics.head(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We saw earlier some dependencies between **answered_correctly** and the frequency of **task_container_id**. Therefore I want to add some features for the **task_container_id**"},{"metadata":{"trusted":true},"cell_type":"code","source":"task_container_characteristics = df.groupby('task_container_id').agg({'answered_correctly':\n                                                                      ['mean', 'median', 'std', 'skew', 'count']})\ntask_container_characteristics.columns = [\n    'mean_task_acc',\n    'median_task_acc',\n    'std_task_acc',\n    'skew_task_acc',\n    'number_of_asked_task_containers'\n]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"task_container_characteristics.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_characteristics = df.groupby('content_id').agg({'answered_correctly':\n                                                        ['mean', 'median', 'std', 'skew', 'count']})\ncontent_characteristics.columns = [\n    'mean_acc',\n    'median_acc',\n    'std_acc',\n    'skew_acc',\n    'number_of_asked_q'\n]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_characteristics.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train.copy()\ndel train","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Merge all of our data"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.merge(user_characteristics, how='left', on='user_id')\ndf = df.merge(task_container_characteristics, how='left', on='task_container_id')\ndf = df.merge(content_characteristics, how='left', on='content_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"features = [\n    'prior_question_elapsed_time', \n    'prior_question_had_explanation',\n    'mean_user_acc',\n    'median_user_acc',\n    'std_user_acc',\n    'skew_user_acc',\n    'number_of_answered_q',\n    'mean_task_acc',\n    'median_task_acc',\n    'std_task_acc',\n    'skew_task_acc',\n    'number_of_asked_task_containers',\n    'mean_acc',\n    'median_acc',\n    'std_acc',\n    'skew_acc',\n    'number_of_asked_q'\n]\n\ntarget = 'answered_correctly'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Drop features that we are not going to use in our model"},{"metadata":{"trusted":true},"cell_type":"code","source":"col_to_drop = set(df.columns.values.tolist()).difference(features + [target])\nfor col in col_to_drop:\n    del df[col]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df['prior_question_had_explanation'] = df['prior_question_had_explanation'].fillna(value=False).astype(bool)\ndf = df.fillna(value=0.5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = df.replace([np.inf, -np.inf], np.nan)\ndf = df.fillna(0.5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"3\"></a>\n\n<div style=\"background:write; border:0; color:black; width: 100%; height: 50px\">\n    <div style=\"vertical-align: middle; text-align:center\"><h1>MODELING</h1></div>\n</div>"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df, test_df, y_train, y_test = train_test_split(df[features], df[target], random_state=777, test_size=0.2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here I'm not going to tune LGB parameters.\nI want just choose the most relevant from the most useful"},{"metadata":{"trusted":true},"cell_type":"code","source":"# clf = LGBMClassifier(random_state=777)\n\n# params = {\n#     'n_estimators': [50, 150, 300],\n#     'max_depth': [3, 5, 10],\n#     'num_leaves': [5, 15, 30],\n#     'min_data_in_leaf': [5, 50, 100],\n#     'feature_fraction': [0.1, 0.5, 1.],\n#     'lambda': [0., 0.5, 1.],\n# }\n\n# cv = RandomizedSearchCV(clf, param_distributions=params, cv=5, n_iter=50, verbose=2)\n# cv.fit(df[features], df[target])\n\n# print(cv.best_params_)\n# print(cv.best_score_)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It was the best params. Therefore I will use them"},{"metadata":{"trusted":true},"cell_type":"code","source":"params = {\n    'num_leaves': 30, \n    'n_estimators': 300, \n    'min_data_in_leaf': 100, \n    'max_depth': 5, \n    'lambda': 0.0, \n    'feature_fraction': 1.0\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = LGBMClassifier(**params)\nmodel.fit(train_df, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('LGB ROC-AUC score: ', roc_auc_score(y_test.values, model.predict_proba(test_df)[:, 1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"eli5.show_weights(model, top=20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lgb.plot_importance(model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"4\"></a>\n\n<div style=\"background:write; border:0; color:black; width: 100%; height: 50px\">\n    <div style=\"vertical-align: middle; text-align:center\"><h1>SUBMISSION PREPARATION</h1></div>\n</div>"},{"metadata":{"trusted":true},"cell_type":"code","source":"env = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"iter_test = env.iter_test()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for (test_df, sample_prediction_df) in iter_test:\n    # merge\n    test_df = test_df.merge(user_characteristics, on = \"user_id\", how = \"left\")\n    test_df = test_df.merge(task_container_characteristics, on = \"task_container_id\", how = \"left\")\n    test_df = test_df.merge(content_characteristics, on = \"content_id\", how = \"left\")\n    \n    # type transformation\n    test_df['prior_question_had_explanation'] = test_df['prior_question_had_explanation'].fillna(value=False).astype(bool)\n    test_df.fillna(value = 0.5, inplace = True)\n    test_df = test_df.replace([np.inf, -np.inf], np.nan)\n    test_df = test_df.fillna(0.5)\n    \n    # preds\n    test_df['answered_correctly'] = model.predict_proba(test_df[features])[:, 1]\n    cols_to_submission = ['row_id', 'answered_correctly', 'group_num']\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}