{"cells":[{"metadata":{},"cell_type":"markdown","source":"# ***Riiid AIEd Challenge 2020***\n\nBefore describing the project I would like to mention why i chose the topic. I loved teaching since forever, I have 6+ years of volunteering experience in teaching underprivileged children in India. I started this in my undergrad first year and continued till the time in came US. The topic really struck a old string and I picked this.\n\nIn 2018, 260 million children weren't attending school. \nAt the same time, more than half of these young students didn't meet minimum reading and math standards. \nEducation was already in a tough place when COVID-19 forced most countries to temporarily close schools.\nThis further delayed learning opportunities and intellectual development. \nThe equity gaps in every country could grow wider. \nWe need to re-think the current education system in terms of attendance, engagement, and individualized attention.\nWith a strong belief in equal opportunity in education, Riiid launched an AI tutor based on deep-learning algorithms in 2017 that attracted more than one million South Korean students.\nThis year, the company released EdNet, the world’s largest open database for AI education containing more than 130 million student interactions coming from over 780,000 students.\n\n\nTailoring education to a student's ability level is one of the many valuable things an AI tutor can do.We will predict whether students are able to answer their next questions correctly.We will be provided with the same sorts of information a complete education app would have: that student's historic performance, the performance of other students on the same question, metadata about the question itself, and more.\n\n\nIn this competition, our  challenge is to create algorithms for \"Knowledge Tracing,\" the modeling of student knowledge over time. The goal is to accurately predict how students will perform on future interactions. We will pair our machine learning skills using the world’s largest education dataset, EdNet, which consists of more than 130 million interactions coming from over 780,000 students.\n\nSubmissions are evaluated on area under the ROC curve between the predicted probability and the observed target."},{"metadata":{"trusted":true},"cell_type":"code","source":"pip install datatable","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Datatable (heavily inspired by R's data.table) can read large datasets fairly quickly and is often faster than pandas. It is specifically meant for data processing of tabular datasets with emphasis on speed and support for large sized data.\n\nDocumentation: https://datatable.readthedocs.io/en/latest/index.html"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Import Statements\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport datatable as dt\n\nimport plotly.express as px\nfrom plotly.subplots import make_subplots\nimport plotly.graph_objs as go\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.style as style\nstyle.use('fivethirtyeight')\nimport seaborn as sns\nimport os\nfrom matplotlib.ticker import FuncFormatter\n\n# import riiideducation\n# env = riiideducation.make_env()\n\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.model_selection import train_test_split\nfrom lightgbm import LGBMClassifier\n\nimport optuna\nfrom optuna.samplers import TPESampler\n\nfrom sklearn.feature_selection import RFE\nfrom sklearn.tree import DecisionTreeClassifier","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sampler = TPESampler(\n    seed=666\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# To check all file paths\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Loading Dataset Files using DataTable\nlectures = dt.fread(\"../input/riiid-test-answer-prediction/lectures.csv\").to_pandas()\nquestions = dt.fread(\"../input/riiid-test-answer-prediction/questions.csv\").to_pandas()\nsample_sub = dt.fread(\"../input/riiid-test-answer-prediction/example_sample_submission.csv\").to_pandas()\nexample_test = dt.fread(\"../input/riiid-test-answer-prediction/example_test.csv\").to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Dataset size overview\nprint(\"Lectures Data Shape\",lectures.shape)\nprint(\"Questions Data Shape\",questions.shape)\nprint(\"Sample Submission Data Shape\",sample_sub.shape)\nprint(\"Example Test Data Shape\",example_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Data Overview\nprint(lectures.head())\nprint(questions.head())\nprint(sample_sub.head())\nprint(example_test.head())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Loading Training Dataset using Datatable\ntrain = dt.fread(\"../input/riiid-test-answer-prediction/train.csv\").to_pandas()\n\n# Saving Training model in .jay format for easier reload\n# train = dt.fread(\"../input/riiid-test-answer-prediction/train.csv\").to_jay(\"train.jay\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# types = {\n#         'row_id': 'int64', \n#         'timestamp': 'int64', \n#         'user_id': 'int32', \n#         'content_id': 'int16', \n#         'content_type_id': 'int8',\n#         'task_container_id': 'int16', \n#         'user_answer': 'int8', \n#         'answered_correctly': 'int8', \n#         'prior_question_elapsed_time': 'float32', \n#         'prior_question_had_explanation': 'boolean'\n# }\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# train = pd.read_csv(\n#     '/kaggle/input/riiid-test-answer-prediction/train.csv', \n#     low_memory=False, \n#     nrows=10**6, \n#     dtype=types\n# )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Loading Training Dataset using .jay file\n# train = dt.fread(\"./train.jay\").to_pandas()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**train.csv**\n\nrow_id: (int64) ID code for the row.\n\ntimestamp: (int64) the time between this user interaction and the first event from that user.\n\nuser_id: (int32) ID code for the user.\n\ncontent_id: (int16) ID code for the user interaction\n\ncontent_type_id: (int8) 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture.\n\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id. Monotonically increasing for each user.\n\nuser_answer: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n\nanswered_correctly: (int8) if the user responded correctly. Read -1 as null, for lectures.\n\nprior_question_elapsed_time: (float32) How long it took a user to answer their previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Note that the time is the total time a user took to solve all the questions in the previous bundle.\n\nprior_question_had_explanation: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Training Data Overview\nprint(\"Training Data Shape\",train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Training Data Overview\ntrain.head().style.applymap(lambda x:\"background-color:lightblue\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Part of missing values for every column')\nprint(100* train.isnull().sum() / len(train))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Data Structure and its analysis\n**\n\nWe have been given three csv (Comma Seperated Values) files: Train, Questions, Lectures. All the three has been given with different features and related to each other\n\nTrain: This File Contains the ID's related to Questions, User and Content. Along with Timestamps of Interaction of User with Content.Answers Provided by the user and their overall effeciency.Now Content is Divided into two parts: As this Data has been collected from an Educational Application, they are divided into Questions and Lectures. Metadata related to Questions and Lectures has been provided in different files.\n\nQuestions: This File contain question ID's, their correct answers and tags to which these questions are related to.\n\nLectures: This File contain Lecture ID along with Summary to what part is covered by this particular lecture."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.memory_usage(deep=True)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that 'prior_question_had_explanation' is object and taking a lot of memory, while it is supposed to be boolean. Let's fix this before continuing."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['prior_question_had_explanation'] = train['prior_question_had_explanation'].astype('boolean')\ntrain.memory_usage(deep=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f'We have {train.user_id.nunique()} unique users in our train set')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Questions or Lectures\n# Content_type_id = False means that a question was asked. True means that the user was watching a lecture.\ntrain.content_type_id.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Content_id is question_id if content_type is question\n# Content_id is a code for the user interaction. Basically, these are the questions if content_type is question (question_id: foreign key for the train/test content_id column, when the content type is question).\nprint(f'We have {train.content_id.nunique()} content ids in our train set, of which {train[train.content_type_id == False].content_id.nunique()} are questions.')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cids = train.content_id.value_counts()[:30]\n\nfig = plt.figure(figsize=(12,6))\nax = cids.plot.bar()\nplt.title(\"Thirty most used content id's\")\nplt.xticks(rotation=90)\nax.get_yaxis().set_major_formatter(FuncFormatter(lambda x, p: format(int(x), ','))) #add thousands separator\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Top 40 users by number of actions\nds = train['user_id'].value_counts().reset_index()\n\nds.columns = ['user_id','count']\n\nds['user_id'] = ds['user_id'].astype(str) + '-'\nds = ds.sort_values(['count']).tail(40)\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='user_id', \n    orientation='h', \n    title='Top 40 users by number of actions', \n    width=800,\n    height=900 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#User action distribution\nds = train['user_id'].value_counts().reset_index()\n\nds.columns = ['user_id', 'count']\n\nds = ds.sort_values('user_id')\n\nfig = px.line(\n    ds, \n    x='user_id', \n    y='count', \n    title='User action distribution', \n    width=800,\n    height=600 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Top 40 most useful content_ids\nds = train['content_id'].value_counts().reset_index()\n\nds.columns = [\n    'content_id', \n    'count'\n]\n\nds['content_id'] = ds['content_id'].astype(str) + '-'\nds = ds.sort_values(['count']).tail(40)\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='content_id', \n    orientation='h', \n    title='Top 40 most useful content_ids',  \n    width=800,\n    height=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Lecures & questions\nds = train['content_type_id'].value_counts().reset_index()\n\nds.columns = [\n    'content_type_id', \n    'percent'\n]\n\nds['percent'] /= len(train)\n\nfig = px.pie(\n    ds, \n    names='content_type_id', \n    values='percent', \n    title='Lecures & questions', \n    width=800,\n    height=500 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Top 40 most useful task_container_ids\nds = train['task_container_id'].value_counts().reset_index()\n\nds.columns = [\n    'task_container_id', \n    'count'\n]\n\nds['task_container_id'] = ds['task_container_id'].astype(str) + '-'\nds = ds.sort_values(['count']).tail(40)\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='task_container_id', \n    orientation='h', \n    title='Top 40 most useful task_container_ids', \n    width=800,\n    height=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Percent of user answers for every option\nds = train['user_answer'].value_counts().reset_index()\n\nds.columns = [\n    'user_answer', \n    'percent_of_answers'\n]\n\nds['percent_of_answers'] /= len(train)\nds = ds.sort_values(['percent_of_answers'])\n\nfig = px.bar(\n    ds, \n    x='user_answer', \n    y='percent_of_answers', \n    orientation='v', \n    title='Percent of user answers for every option', \n    width=800,\n    height=400 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Percent of correct answers\nds = train['answered_correctly'].value_counts().reset_index()\n\nds.columns = [\n    'answered_correctly', \n    'percent_of_answers'\n]\n\nds['percent_of_answers'] /= len(train)\nds = ds.sort_values(['percent_of_answers'])\n\nfig = px.pie(\n    ds, \n    names='answered_correctly', \n    values='percent_of_answers', \n    title='Percent of correct answers', \n    width=800,\n    height=500 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Percent of correct answers for every option\nfig = make_subplots(rows=3, cols=2)\n\ntraces = [\n    go.Bar(\n        x=[\n            -1, 0, 1\n        ], \n        y=[\n            len(train[(train['user_answer'] == item) & (train['answered_correctly'] == -1)]),\n            len(train[(train['user_answer'] == item) & (train['answered_correctly'] == 0)]),\n            len(train[(train['user_answer'] == item) & (train['answered_correctly'] == 1)])\n        ], \n        name='Option: ' + str(item),\n        text = [\n            str(round(100 * len(train[(train['user_answer'] == item) & (train['answered_correctly'] == -1)]) / len(train[(train['user_answer'] == item)]), 2)) + '%',\n            str(round(100 * len(train[(train['user_answer'] == item) & (train['answered_correctly'] == -0)]) / len(train[(train['user_answer'] == item)]), 2)) + '%',\n            str(round(100 * len(train[(train['user_answer'] == item) & (train['answered_correctly'] == 1)]) / len(train[(train['user_answer'] == item)]), 2)) + '%',\n        ],\n        textposition='auto'\n    ) for item in train['user_answer'].unique().tolist()\n]\n\nfor i in range(len(traces)):\n    fig.append_trace(\n        traces[i], \n        (i // 2) + 1, \n        (i % 2)  + 1\n    )\n\nfig.update_layout(\n    title_text='Percent of correct answers for every option',\n    height=900,\n    width=800\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# #prior_question_elapsed_time distribution\n# fig = px.histogram(\n#     train, \n#     x=\"prior_question_elapsed_time\",\n#     nbins=50,\n#     title='prior_question_elapsed_time distribution',\n#     width=800,\n#     height=500\n# )\n\n# fig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**questions.csv:** metadata for the questions posed to users.\n\nquestion_id: foreign key for the train/test content_id column, when the content type is question (0).\n\nbundle_id: code for which questions are served together.\n\ncorrect_answer: the answer to the question. Can be compared with the train user_answer column to check if the user was right.\n\npart: top level category code for the question.\n\ntags: one or more detailed tag codes for the question. The meaning of the tags will not be provided, but these codes are sufficient for clustering the questions together."},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Part of missing values for every column')\nprint(questions.isnull().sum() / len(questions))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Number of correct answers per group\nds = questions['correct_answer'].value_counts().reset_index()\n\nds.columns = [\n    'correct_answer', \n    'number_of_answers'\n]\n\nds['correct_answer'] = ds['correct_answer'].astype(str) + '-'\nds = ds.sort_values(['number_of_answers'])\n\nfig = px.bar(\n    ds, \n    x='number_of_answers', \n    y='correct_answer', \n    orientation='h', \n    title='Number of correct answers per group', \n    width=800,\n    height=300\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Parts distribution\nds = questions['part'].value_counts().reset_index()\n\nds.columns = [\n    'part', \n    'count'\n]\n\nds['part'] = ds['part'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='part', \n    orientation='h', \n    title='Parts distribution',\n    width=800,\n    height=400\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions['tag'] = questions['tags'].str.split(' ')\nquestions = questions.explode('tag')\nquestions = pd.merge(\n    questions, \n    questions.groupby('question_id')['tag'].count().reset_index(), \n    on='question_id'\n)\n\nquestions = questions.drop(['tag_x'], axis=1)\n\nquestions.columns = [\n    'question_id', \n    'bundle_id', \n    'correct_answer', \n    'part', \n    'tags', \n    'tags_number'\n]\n\nquestions = questions.drop_duplicates()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Number tags distribution\nds = questions['tags_number'].value_counts().reset_index()\n\nds.columns = [\n    'tags_number', \n    'count'\n]\n\nds['tags_number'] = ds['tags_number'].astype(str) + '-'\nds = ds.sort_values(['tags_number'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='tags_number', \n    orientation='h', \n    title='Number tags distribution', \n    width=800,\n    height=400 \n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**lectures.csv**: metadata for the lectures watched by users as they progress in their education.\n\nlecture_id: foreign key for the train/test content_id column, when the content type is lecture (1).\n\npart: top level category code for the lecture.\n\ntag: one tag codes for the lecture. The meaning of the tags will not be provided, but these codes are sufficient for clustering the lectures together.\n\ntype_of: brief description of the core purpose of the lecture"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Part of missing values for every column')\nprint(lectures.isnull().sum() / len(lectures))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#type_of column distribution\nds = lectures['type_of'].value_counts().reset_index()\n\nds.columns = [\n    'type_of', \n    'count'\n]\n\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='type_of', \n    orientation='h', \n    title='type_of column distribution', \n    height=300, \n    width=800\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Modeling**"},{"metadata":{"trusted":true},"cell_type":"code","source":"used_data_types_dict = {\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float16',\n    'prior_question_had_explanation': 'boolean'\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# train_df = pd.read_csv(\n#     '/kaggle/input/riiid-test-answer-prediction/train.csv',\n#     usecols = used_data_types_dict.keys(),\n#     dtype=used_data_types_dict, \n#     index_col = 0\n# )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"features_df = train_df.iloc[:int(9/10 * len(train_df))]\ntrain_df = train_df.iloc[int(9/10 * len(train_df)):]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_questions_only_df = features_df[features_df['answered_correctly']!=-1]\ngrouped_by_user_df = train_questions_only_df.groupby('user_id')\n\nuser_answers_df = grouped_by_user_df.agg(\n    {\n        'answered_correctly': [\n            'mean', \n            'count', \n            'std', \n            'median', \n            'skew'\n        ]\n    }\n).copy()\n\nuser_answers_df.columns = [\n    'mean_user_accuracy', \n    'questions_answered', \n    'std_user_accuracy', \n    'median_user_accuracy', \n    'skew_user_accuracy'\n]\n\nuser_answers_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"grouped_by_content_df = train_questions_only_df.groupby('content_id')\ncontent_answers_df = grouped_by_content_df.agg(\n    {\n        'answered_correctly': [\n            'mean', \n            'count', \n            'std', \n            'median', \n            'skew'\n        ]\n    }\n).copy()\n\ncontent_answers_df.columns = [\n    'mean_accuracy', \n    'question_asked', \n    'std_accuracy', \n    'median_accuracy', \n    'skew_accuracy'\n]\n\ncontent_answers_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del features_df\ndel grouped_by_user_df\ndel grouped_by_content_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"features = [\n    'mean_user_accuracy', \n    'questions_answered',\n    'std_user_accuracy', \n    'median_user_accuracy',\n    'skew_user_accuracy',\n    'mean_accuracy', \n    'question_asked',\n    'std_accuracy', \n    'median_accuracy',\n    'prior_question_elapsed_time', \n    'prior_question_had_explanation',\n    'skew_accuracy'\n]\n\ntarget = 'answered_correctly'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = train_df[train_df[target] != -1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = train_df.merge(user_answers_df, how='left', on='user_id')\ntrain_df = train_df.merge(content_answers_df, how='left', on='content_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['prior_question_had_explanation'] = train_df['prior_question_had_explanation'].fillna(value=False).astype(bool)\ntrain_df = train_df.fillna(value=0.5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = train_df[features + [target]]\ntrain_df = train_df.replace([np.inf, -np.inf], np.nan)\ntrain_df = train_df.fillna(0.5)\n\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df, test_df = train_test_split(train_df, random_state=666, test_size=0.2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rfe = RFE(\n    estimator=DecisionTreeClassifier(\n        random_state=666\n    ), \n    n_features_to_select=8\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rfe.fit(train_df[features], train_df[target])\nX_transformed = rfe.transform(train_df[features])\nX_transformed = pd.DataFrame(X_transformed)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_transformed.columns = ['col_1', 'col_2', 'col_3', 'col_4', 'col_5', 'col_6', 'col_7', 'col_8']\nX_transformed","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_transformed_test = rfe.transform(test_df[features])\nX_transformed_test = pd.DataFrame(X_transformed_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_transformed_test.columns = ['col_1', 'col_2', 'col_3', 'col_4', 'col_5', 'col_6', 'col_7', 'col_8']\nX_transformed_test","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"X_transformed_test['col_1'] = X_transformed_test['col_1'].astype(np.float32)\nX_transformed_test['col_2'] = X_transformed_test['col_2'].astype(np.float32)\nX_transformed_test['col_3'] = X_transformed_test['col_3'].astype(np.int32)\nX_transformed_test['col_4'] = X_transformed_test['col_4'].astype(np.float32)\nX_transformed_test['col_5'] = X_transformed_test['col_5'].astype(np.int32)\nX_transformed_test['col_6'] = X_transformed_test['col_6'].astype(np.int32)\nX_transformed_test['col_7'] = X_transformed_test['col_7'].astype(np.int32)\nX_transformed_test['col_8'] = X_transformed_test['col_8'].astype(np.float32)\n\nX_transformed['col_1'] = X_transformed['col_1'].astype(np.float32)\nX_transformed['col_2'] = X_transformed['col_2'].astype(np.float32)\nX_transformed['col_3'] = X_transformed['col_3'].astype(np.int32)\nX_transformed['col_4'] = X_transformed['col_4'].astype(np.float32)\nX_transformed['col_5'] = X_transformed['col_5'].astype(np.int32)\nX_transformed['col_6'] = X_transformed['col_6'].astype(np.int32)\nX_transformed['col_7'] = X_transformed['col_7'].astype(np.int32)\nX_transformed['col_8'] = X_transformed['col_8'].astype(np.float32)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def create_model(trial):\n    num_leaves = trial.suggest_int(\"num_leaves\", 2, 31)\n    n_estimators = trial.suggest_int(\"n_estimators\", 50, 300)\n    max_depth = trial.suggest_int('max_depth', 3, 8)\n    min_child_samples = trial.suggest_int('min_child_samples', 100, 1200)\n    learning_rate = trial.suggest_uniform('learning_rate', 0.0001, 0.99)\n    min_data_in_leaf = trial.suggest_int('min_data_in_leaf', 5, 90)\n    bagging_fraction = trial.suggest_uniform('bagging_fraction', 0.0001, 1.0)\n    feature_fraction = trial.suggest_uniform('feature_fraction', 0.0001, 1.0)\n    \n    model = LGBMClassifier(\n        num_leaves=num_leaves,\n        n_estimators=n_estimators, \n        max_depth=max_depth, \n        min_child_samples=min_child_samples, \n        min_data_in_leaf=min_data_in_leaf,\n        learning_rate=learning_rate,\n        feature_fraction=feature_fraction,\n        random_state=666\n    )\n    return model\n\ndef objective(trial):\n    model = create_model(trial)\n    model.fit(X_transformed, train_df[target])\n    score = roc_auc_score(\n        test_df[target].values, \n        model.predict_proba(X_transformed_test)[:,1]\n    )\n    return score\n\n# uncomment to use optuna\n# final params is in study.best_params\nstudy = optuna.create_study(direction=\"maximize\", sampler=sampler)\nstudy.optimize(objective, n_trials=70)\nparams = study.best_params\nparams['random_state'] = 666\n\n\nprint(params)\n\n# params = {\n#     'bagging_fraction': 0.5817242323514327,\n#     'feature_fraction': 0.6884588361650144,\n#     'learning_rate': 0.42887924851375825, \n#     'max_depth': 6,\n#     'min_child_samples': 946, \n#     'min_data_in_leaf': 47, \n#     'n_estimators': 169,\n#     'num_leaves': 29,\n#     'random_state': 666\n# }\n\nmodel = LGBMClassifier(\n    **params\n)\n\nmodel.fit(X_transformed, train_df[target])\nprint('LGB score: ', roc_auc_score(test_df[target].values, model.predict_proba(X_transformed_test)[:,1]))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}