{
  "id": 209597,
  "title": "18th place solution",
  "url": "/competitions/riiid-test-answer-prediction/writeups/westq-18th-place-solution",
  "author_name": "",
  "post_date": "2021-01-08T03:52:21.213Z",
  "votes": 98,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Congrats to all medal teams and new Grandmasters,Masters,Experts.Thanks to Organizers and kaggle for such a good competition,it shows that kaggle competition is not just a game but also can be a useful machine learning project.</p>\n<h1>Team</h1>\n<ul>\n<li>At first we have three  teams individually, tomoyo and me, ethan and qyxs , wrb0312.We focus on feature engineering and optimization before wrb0312 joined us,I think we made many good features but neural network dominated this competition.wrb0312 did a great job even he use transformer for the first time.After wrb0312 joined us there are only ten days left,we focus on ensembling our models for inference,and also improved transformer very much.Our team members are from china and japan, it's very interesting to see we use chinese,japanese,english mixed-language to communicate.Greate job everyone!</li>\n</ul>\n<h1>Optimization</h1>\n<ul>\n<li><p>For GBM features, rather than using many dictionary to save features' data, we developed a nubma-based framework to speed up feature engineering process and online calculation. Firstly, the data are sorted by ['user_id', 'timestamp', 'content_id'] and split into different arrays. Then we created features in different array via self-designed rolling function or self-designed cumlative function. Actually, it provides us a very flexible way to create features and test it. In 10m data, the feature engineering process needs only 5 minutes to finish it.</p></li>\n<li><p>Some examples are listed as below. </p></li>\n</ul>\n<pre><code>from tqdm import tqdm\nfrom numba import jit,njit\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\nimport gc\nfrom multiprocessing import Process, Manager,Pool\nfrom functools import partial\nfrom numba import prange\nimport numpy as np\nimport pandas as pd\nfrom numba import types\nfrom numba.typed import Dict\nimport functools, time\nfrom numba.typed import List\n\n\ndef timeit(f):\n    def wrap(*args, **kwargs):\n        time1 = time.time()\n        ret = f(*args, **kwargs)\n        time2 = time.time()\n        print('{:s} function took {:.3f} s'.format(f.__name__, np.round(time2-time1, 2)))\n\n        return ret\n    return wrap\n\n\ndef rolling_feat_group(train, col_used):\n    a = train[col_used].values\n    ind = np.lexsort((a[:,2],a[:,1],a[:,0]))\n    a = a[ind]\n    g = np.split(a, np.unique(a[:, 0], return_index=True)[1][1:])\n    return g, ind, col_used\n\n@jit(nopython = True, fastmath = True)\ndef rolling_cal(arr, step, window = 5, shift_ = 1):\n    m = 2\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    beg = window\n    for i in step: \n        tmp = arr_[beg-window:beg]\n        ret[beg - window:(beg - window + i), 0] = np.nanmean(tmp)\n        ret[beg - window:(beg - window + i), 1] = np.nansum(tmp)\n        beg += i\n    return ret\n\n\n@jit(nopython = True, fastmath = True)\ndef rolling_time_cal(arr, window = 5, shift_ = 1):\n    m = 1\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    for i in range(0,arr.shape[0], 1): \n        tmp = arr_[i:i+window+1]\n        ret[i, 0] = np.nanmean(tmp)\n    return ret\n\ndef rolling_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    step = np.unique(tmp_g[:, 1], return_counts=True)[1]\n    for window_size in shift_period:\n        tmp = rolling_cal(tmp_g[:, 2], step, window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\ndef rolling_time_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    for window_size in shift_period:\n        tmp = rolling_time_cal(tmp_g[:, 2], window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\n\ndef rolling_feat_cal(tmp_g, name_dict, global_period):\n    answer_idx = name_dict.index('answered_correctly')\n    prior_idx = name_dict.index('prior_question_elapsed_time')\n    item_mean_idx = name_dict.index('item_mean')\n    task_set_idx = name_dict.index('task_set_distance')\n    tmp_res1 = rolling_cal_wrap(tmp_g[:,[0,1, answer_idx]], global_period)\n    tmp_res2 = rolling_time_cal_wrap(tmp_g[:,[0,1, prior_idx]], global_period)\n    tmp_res3 = rolling_time_cal_wrap(tmp_g[:,[0,1, item_mean_idx]], global_period)\n    tmp_res4 = rolling_time_cal_wrap(tmp_g[:,[0,1, task_set_idx]], global_period)\n    tmp_res = np.concatenate([tmp_res1, tmp_res2, tmp_res3, tmp_res4], axis = 1)\n    return tmp_res\n</code></pre>\n<ul>\n<li>If anyone interested in how to create features via numba-framework, Tomoyo publiced his full GBM pipeline in github(<a href=\"url\" target=\"_blank\">https://github.com/ZiwenYeee/Riiid-numba-framework</a>)</li>\n</ul>\n<h1>Catboost(LB 0.807)</h1>\n<h3>summary</h3>\n<ul>\n<li>We created 183 features for final catboost model,including some original features,global statistics(item base),cumulative and rolling statistics(user base),tfidf-svd(base on question's user list),word2vec(base on user's question list, wrong and correct tag list ),timedelta from many perspective,last same part groups features.</li>\n</ul>\n<h3>gbm benchmark</h3>\n<ul>\n<li>We compared lightgbm ,xgboost,catboost,catboost is the best for the training and inference speed,and memory consuming.When train the full data,lightgbm need over 100 hours with my AMD Ryzen ThreadRipper 3970X,xgboost always have out of memory error even using dask with 4 RTX 3090.</li>\n</ul>\n<h3>strong features and interesting finding by qyxs</h3>\n<ul>\n<li>1.  the history difficulty statistics features of user who had correct/wrong answers, boost almost 0.003</li>\n</ul>\n<pre><code>tmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n</code></pre>\n<ul>\n<li>2.  focus on the records about the current part of user connect with last same part, generate the features include answer correct ratio, time diff, frequency etc, boost almost 0.002</li>\n</ul>\n<pre><code>for_question_df['rank_part'] = for_question_df.groupby(['user_id', 'part'])['timestamp'].rank(method='first')\nfor_question_df['rank_user'] = for_question_df.groupby(['user_id'])['timestamp'].rank(method='first')\nfor_question_df['rank_diff'] = for_question_df['rank_user'] - for_question_df['rank_part']\nfor_question_df['part_times'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense')\nfor_question_df['rank_diff'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense', ascending=False)\n\nlast_part = for_question_df[for_question_df['rank_diff']==1]\npart_times = for_question_df.groupby(['user_id', 'part'])['part_times'].agg([['part_times', 'max']]).reset_index()\n\nlast_part_df = last_part.groupby(['user_id', 'part'])['answered_correctly'].agg([['last_continue_part_ratio', 'mean'], ['last_continue_part_cnt', 'count']]).reset_index()\nlast_part_time = last_part.groupby(['user_id', 'part'])['timestamp'].agg([['last_continue_part_time_start', 'min'], ['last_continue_part_time_end', 'max']]).reset_index()\nlast_part_df = last_part_df.merge(last_part_time, on=['user_id', 'part'], how='left')\nlast_part_df = last_part_df.merge(part_times, on=['user_id', 'part'], how='left')\nlast_part_df['part_time_diff'] = last_part_df['last_continue_part_time_end'] - last_part_df['last_continue_part_time_start']\nlast_part_df['part_time_freq'] = last_part_df['last_continue_part_cnt']/last_part_df['part_time_diff']\n\nfor_train = for_train.merge(last_part_df, on=['user_id', 'part'], how='left')\nfor_train['last_continue_part_time_start'] = for_train['timestamp'] - for_train['last_continue_part_time_start']\nfor_train['last_continue_part_time_end'] = for_train['timestamp'] - for_train['last_continue_part_time_end']\n</code></pre>\n<ul>\n<li>3.  the answer correctly ratio of each question under differenct user abilititys (split for 11 bins), boost almost 0.001</li>\n</ul>\n<pre><code>for_question_df['user_ability'] = for_question_df.groupby('user_id')['answered_correctly'].transform('mean').round(1)\ntmp_df = for_question_df.pivot_table(index='content_id', columns='user_ability', values='answered_correctly', aggfunc='mean').reset_index()\ntmp_df.columns = ['content_id'] + [f'c_mean_{i}_ratio' for i in range(11)]\nfor_train = for_train.merge(tmp_df, on='content_id', how='left')\n</code></pre>\n<ul>\n<li>Some interseting points:<ul>\n<li>1.they would watch lecture after users had wrong answers, so we could generated some features from this. LB is not improved caused by the lectures info in next group maybe.</li>\n<li>2.the content_id such as 0-195， 7851-7984 etc, then are all same in one part and continuous with each other，we could build a new bundle to generate features</li></ul></li>\n</ul>\n<h3>strong features and interesting finding by ethan</h3>\n<ul>\n<li>1.  user's behavior in last 1,5,…,60 minutes, 0.001 boost</li>\n</ul>\n<pre><code>for w in [1, 5, 10, 15, 30, 45, 60]:\n    print(w)\n    tmp = q_logs[q_logs['timestamp']&gt;=(q_logs['end_time']-w*60*1000)].copy()\n    group_df = tmp.groupby(['user_id'])['content_id'].agg([['user_content_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['part'].agg([['user_part_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['answered_correctly'].agg([['user_correct_raito_in_last{}mins'.format(w), 'mean']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n</code></pre>\n<ul>\n<li>2. \"users' ablility\" statistics in each question, seperately by \"answered_correctly\"(0/1), 0.002 boost</li>\n</ul>\n<pre><code>cc = q_logs.groupby(['user_id'])['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ngg = q_logs[['user_id', 'content_id', 'answered_correctly']].merge(cc, on=['user_id'], how='left')\n\ngroup_df1 = gg[gg['answered_correctly']==1].groupby(['content_id'])['corr_ratio'].agg([['question_correct_user_ablility_min', 'min'], \n                                                                                       ['question_correct_user_ablility_max', 'max'], \n                                                                                       ['question_correct_user_ablility_mean', 'mean'], \n                                                                                       ['question_correct_user_ablility_skew', 'skew'],\n                                                                                       ['question_correct_user_ablility_med', 'median'],\n                                                                                       ['question_correct_user_ablility_std', 'std']]).reset_index()\ngroup_df2 = gg[gg['answered_correctly']==0].groupby(['content_id'])['corr_ratio'].agg([['question_wrong_user_ablility_min','min'], \n                                                                                       ['question_wrong_user_ablility_max','max'],\n                                                                                       ['question_wrong_user_ablility_mean','mean'],\n                                                                                       ['question_wrong_user_ablility_skew','skew'],\n                                                                                       ['question_wrong_user_ablility_med','median'],\n                                                                                       ['question_wrong_user_ablility_std','std']]).reset_index()\n</code></pre>\n<ul>\n<li>3. \"lagtime\" statistics in each question, seperately by \"answered_correctly\"(0/1), means the distribution of users' preprare time for answering this question correctly, about 0.001 boost</li>\n</ul>\n<pre><code>user_task_timestamp = q_logs[['user_id', 'task_container_id', 'timestamp']].drop_duplicates()\nuser_task_timestamp['lag_time'] = user_task_timestamp['timestamp'] - user_task_timestamp.groupby(['user_id'])['timestamp'].shift(1)\ntmp = q_logs[['user_id', 'task_container_id', 'content_id', 'answered_correctly']].merge(user_task_timestamp.drop(['timestamp'], axis=1), on=['user_id', 'task_container_id'], how='left')\ngroup_df = tmp[tmp['answered_correctly']==1].groupby(['content_id'])['lag_time'].agg([['c_lag_time_mean', 'mean'],\n                                                                                     ['c_lag_time_std', 'std'],\n                                                                                     ['c_lag_time_max', 'max'],\n                                                                                     ['c_lag_time_min', 'min'],\n                                                                                     ['c_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n\ngroup_df = tmp[tmp['answered_correctly']==0].groupby(['content_id'])['lag_time'].agg([['w_lag_time_mean', 'mean'],\n                                                                                       ['w_lag_time_std', 'std'],\n                                                                                       ['w_lag_time_max', 'max'],\n                                                                                       ['w_lag_time_min', 'min'],\n                                                                                       ['w_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n</code></pre>\n<h3>feature list</h3>\n<pre><code>['content_id',\n 'prior_question_elapsed_time',\n 'prior_question_had_explanation',\n 'correct_answer',\n 'user_count',\n 'user_sum',\n 'user_mean',\n 'item_count',\n 'item_sum',\n 'item_mean',\n 'answer_ratio_0',\n 'answer_ratio_1',\n 'answer_ratio_2',\n 'bundle_id',\n 'part',\n 'le_tag',\n 'question_correct_user_ablility_mean',\n 'question_correct_user_ablility_median',\n 'question_wrong_user_ablility_mean',\n 'question_wrong_user_ablility_median',\n 'word2vec_0',\n 'word2vec_1',\n 'word2vec_2',\n 'word2vec_3',\n 'word2vec_4',\n 'svd_0',\n 'svd_1',\n 'svd_2',\n 'svd_3',\n 'svd_4',\n 'tags_w2v_correct_mean_0',\n 'tags_w2v_wrong_mean_0',\n 'tags_w2v_correct_mean_1',\n 'tags_w2v_wrong_mean_1',\n 'tags_w2v_correct_mean_2',\n 'tags_w2v_wrong_mean_2',\n 'tags_w2v_correct_mean_3',\n 'tags_w2v_wrong_mean_3',\n 'tags_w2v_correct_mean_4',\n 'tags_w2v_wrong_mean_4',\n 'real_time_wrong_mean',\n 'real_time_wrong_median',\n 'real_time_correct_mean',\n 'real_time_correct_median',\n 'task_set_distance_wrong_mean',\n 'task_set_distance_wrong_median',\n 'task_set_distance_correct_mean',\n 'task_set_distance_correct_median',\n 'mean_0_ratio',\n 'mean_1_ratio',\n 'mean_3_ratio',\n 'mean_4_ratio',\n 'mean_5_ratio',\n 'mean_6_ratio',\n 'mean_7_ratio',\n 'mean_8_ratio',\n 'mean_9_ratio',\n 'mean_10_ratio',\n 'user_d1',\n 'user_d2',\n 'task_set_distance',\n 'user_diff_mean',\n 'user_diff_std',\n 'user_diff_min',\n 'user_diff_max',\n 'task_set_item_mean',\n 'task_set_item_min',\n 'task_set_item_max',\n 'task_set_distance2',\n 'task_distance_shift',\n 'task_set_distance_diff',\n 'task_distance_diff_shift',\n 'container_mean_1',\n 'container_mean_5',\n 'container_std_5',\n 'container_mean_10',\n 'container_std_10',\n 'container_mean_20',\n 'container_std_20',\n 'container_mean_30',\n 'container_std_30',\n 'container_mean_40',\n 'container_std_40',\n 'prior_question_elapsed_time_mean_1',\n 'prior_question_elapsed_time_mean_5',\n 'prior_question_elapsed_time_mean_10',\n 'prior_question_elapsed_time_mean_20',\n 'prior_question_elapsed_time_mean_30',\n 'prior_question_elapsed_time_mean_40',\n 'item_mean_mean_30',\n 'item_mean_mean_40',\n 'task_set_distance_mean_1',\n 'task_set_distance_mean_5',\n 'task_set_distance_mean_10',\n 'task_set_distance_mean_20',\n 'task_set_distance_mean_30',\n 'begin_time_diff',\n 'end_time_diff',\n 'part_time_diff_mean',\n 'part_session_mean',\n 'part_session_sum',\n 'part_session_count',\n 'full_group0_item_mean_mean',\n 'full_group0_item_mean_median',\n 'full_group0_task_set_distance_median',\n 'full_group0_timestamp_mean',\n 'full_group0_timestamp_median',\n 'full_group1_item_mean_mean',\n 'full_group1_item_mean_median',\n 'full_group1_task_set_distance_median',\n 'full_group1_timestamp_median',\n 'part_sum',\n 'part_count',\n 'part_mean',\n 'part_sum_global_ratio',\n 'part_sum_1',\n 'part_sum_5',\n 'part_mean_5',\n 'part_sum_10',\n 'part_mean_10',\n 'cum_answer0_mean_item_mean',\n 'cum_answer0_median_item_mean',\n 'cum_answer0_median_task_set_distance',\n 'cum_answer1_mean_item_mean',\n 'cum_answer1_median_item_mean',\n 'cum_answer1_mean_task_set_distance',\n 'cum_answer1_median_task_set_distance',\n 'cum_answer0_time_diff',\n 'cum_answer1_time_diff',\n 'global_task_set_shift1',\n 'global_task_set_shift2',\n 'global_task_set_shift4',\n 'global_task_set_shift5',\n 'cum_answer0_mean_wrong_time_diff',\n 'cum_answer0_median_wrong_time_diff',\n 'cum_answer1_mean_right_time_diff',\n 'content_correct_mean',\n 'content_correct_sum',\n 'content_correct_count',\n 'hard_answer0_time',\n 'hard_answer1_time',\n 'full_bundle_item_mean_mean',\n 'full_bundle_item_mean_median',\n 'full_bundle_task_set_distance_mean',\n 'full_bundle_task_set_distance_median',\n 'full_bundle_timestamp_mean',\n 'full_bundle_timestamp_median',\n 'bundle_sum',\n 'bundle_mean',\n 'bundle_count',\n 'user_trend_mean',\n 'user_trend_median',\n 'user_trend_roll_user_ans_sum',\n 'user_trend_roll_user_ans_mean',\n 'user_trend_roll_user_ans_count',\n 'user_trend_roll_item_ans_mean',\n 'user_trend_roll_item_ans_count',\n 'div_ratio1',\n 'div_ratio2',\n 'div_ratio3',\n 'new_Feat0',\n 'new_Feat1',\n 'new_Feat2',\n 'new_Feat3',\n 'part_time_wrong_div',\n 'part_time_right_div',\n 'diff_lag_median_div',\n 'diff_item_median_div',\n 'diff_time_median_div',\n 'diff_item_mean_div',\n 'diff_task_set_mean_div',\n 'diff_timestamp_mean_div',\n 'last_20_frequent_answer',\n 'last_20_frequent_answer_count',\n 'last_20_frequent_answer_mean',\n 'last_20_frequent_answer_sum',\n 'last_user_same_answer_tf',\n 'last_item_same_answer_tf',\n 'last_right_time_diff',\n 'last_wrong_time_diff',\n 'last_5_part_time_div',\n 'last_10_part_time_div',\n 'last_20_part_time_div']\n</code></pre>\n<h1>Transformer(LB 0.808)</h1>\n<ul>\n<li><p>Used questions only. Lectures did't improve our validation score.</p></li>\n<li><p>In training window size 800 (the bigger,the better), in inference window size 300(the bigger,the better,but time consuming).</p></li>\n<li><p>Optimizer : Adam  with lr = 8e-4, beta1 = 0.9, beta2 = 0.999 with warmup steps to 4000.</p>\n<ul>\n<li>Without warmup steps, didn't converge.</li></ul></li>\n<li><p>Number of layers = 4, dimension of the model = 256, dimension of the FFN = 2048.</p></li>\n<li><p>Batch size=80</p>\n<ul>\n<li>With smaller size,  didn't converge.</li></ul></li>\n<li><p>Dropout = 0</p></li>\n<li><p>Postion encoding : Axial Positional Embedding</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.12180\" target=\"_blank\">https://arxiv.org/abs/1912.12180</a></li>\n<li><a href=\"https://github.com/lucidrains/axial-positional-embedding\" target=\"_blank\">https://github.com/lucidrains/axial-positional-embedding</a></li>\n<li>This encoding improved score significantly.</li></ul></li>\n<li><p>Augmentation : Mixup-Transformer</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2010.02394\" target=\"_blank\">https://arxiv.org/abs/2010.02394</a></li></ul></li>\n<li><p>Inputs</p>\n<ul>\n<li>question_id</li>\n<li>part</li>\n<li>prior_question_elapsed_time / 1000</li>\n<li>lagtime<ul>\n<li>log1p((timestamp_t - timestamp_(t-1)) /1000 / 60)</li>\n<li>This feature improved the score significantly</li></ul></li>\n<li>answered_correctly</li>\n<li>GBT feats (imortance top N)</li></ul></li>\n<li><p>model image<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F6743aa3ef1403f4084dd4fa3d8372702%2FScreen%20Shot%202021-01-08%20at%2011.34.11.png?generation=1610073332748004&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h1>Ensemble(LB 0.812)</h1>\n<ul>\n<li>Finally we use one catboost and two transformers for ensemble due to the inference time limitation.Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.</li>\n</ul>",
  "messages": [
    {
      "id": "1143618",
      "postDate": "01/08/2021 01:14:57",
      "content": "<p>Congrats to all medal teams and new Grandmasters,Masters,Experts.Thanks to Organizers and kaggle for such a good competition,it shows that kaggle competition is not just a game but also can be a useful machine learning project.</p>\n<h1>Team</h1>\n<ul>\n<li>At first we have three  teams individually, tomoyo and me, ethan and qyxs , wrb0312.We focus on feature engineering and optimization before wrb0312 joined us,I think we made many good features but neural network dominated this competition.wrb0312 did a great job even he use transformer for the first time.After wrb0312 joined us there are only ten days left,we focus on ensembling our models for inference,and also improved transformer very much.Our team members are from china and japan, it's very interesting to see we use chinese,japanese,english mixed-language to communicate.Greate job everyone!</li>\n</ul>\n<h1>Optimization</h1>\n<ul>\n<li><p>For GBM features, rather than using many dictionary to save features' data, we developed a nubma-based framework to speed up feature engineering process and online calculation. Firstly, the data are sorted by ['user_id', 'timestamp', 'content_id'] and split into different arrays. Then we created features in different array via self-designed rolling function or self-designed cumlative function. Actually, it provides us a very flexible way to create features and test it. In 10m data, the feature engineering process needs only 5 minutes to finish it.</p></li>\n<li><p>Some examples are listed as below. </p></li>\n</ul>\n<pre><code>from tqdm import tqdm\nfrom numba import jit,njit\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\nimport gc\nfrom multiprocessing import Process, Manager,Pool\nfrom functools import partial\nfrom numba import prange\nimport numpy as np\nimport pandas as pd\nfrom numba import types\nfrom numba.typed import Dict\nimport functools, time\nfrom numba.typed import List\n\n\ndef timeit(f):\n    def wrap(*args, **kwargs):\n        time1 = time.time()\n        ret = f(*args, **kwargs)\n        time2 = time.time()\n        print('{:s} function took {:.3f} s'.format(f.__name__, np.round(time2-time1, 2)))\n\n        return ret\n    return wrap\n\n\ndef rolling_feat_group(train, col_used):\n    a = train[col_used].values\n    ind = np.lexsort((a[:,2],a[:,1],a[:,0]))\n    a = a[ind]\n    g = np.split(a, np.unique(a[:, 0], return_index=True)[1][1:])\n    return g, ind, col_used\n\n@jit(nopython = True, fastmath = True)\ndef rolling_cal(arr, step, window = 5, shift_ = 1):\n    m = 2\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    beg = window\n    for i in step: \n        tmp = arr_[beg-window:beg]\n        ret[beg - window:(beg - window + i), 0] = np.nanmean(tmp)\n        ret[beg - window:(beg - window + i), 1] = np.nansum(tmp)\n        beg += i\n    return ret\n\n\n@jit(nopython = True, fastmath = True)\ndef rolling_time_cal(arr, window = 5, shift_ = 1):\n    m = 1\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    for i in range(0,arr.shape[0], 1): \n        tmp = arr_[i:i+window+1]\n        ret[i, 0] = np.nanmean(tmp)\n    return ret\n\ndef rolling_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    step = np.unique(tmp_g[:, 1], return_counts=True)[1]\n    for window_size in shift_period:\n        tmp = rolling_cal(tmp_g[:, 2], step, window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\ndef rolling_time_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    for window_size in shift_period:\n        tmp = rolling_time_cal(tmp_g[:, 2], window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\n\ndef rolling_feat_cal(tmp_g, name_dict, global_period):\n    answer_idx = name_dict.index('answered_correctly')\n    prior_idx = name_dict.index('prior_question_elapsed_time')\n    item_mean_idx = name_dict.index('item_mean')\n    task_set_idx = name_dict.index('task_set_distance')\n    tmp_res1 = rolling_cal_wrap(tmp_g[:,[0,1, answer_idx]], global_period)\n    tmp_res2 = rolling_time_cal_wrap(tmp_g[:,[0,1, prior_idx]], global_period)\n    tmp_res3 = rolling_time_cal_wrap(tmp_g[:,[0,1, item_mean_idx]], global_period)\n    tmp_res4 = rolling_time_cal_wrap(tmp_g[:,[0,1, task_set_idx]], global_period)\n    tmp_res = np.concatenate([tmp_res1, tmp_res2, tmp_res3, tmp_res4], axis = 1)\n    return tmp_res\n</code></pre>\n<ul>\n<li>If anyone interested in how to create features via numba-framework, Tomoyo publiced his full GBM pipeline in github(<a href=\"url\" target=\"_blank\">https://github.com/ZiwenYeee/Riiid-numba-framework</a>)</li>\n</ul>\n<h1>Catboost(LB 0.807)</h1>\n<h3>summary</h3>\n<ul>\n<li>We created 183 features for final catboost model,including some original features,global statistics(item base),cumulative and rolling statistics(user base),tfidf-svd(base on question's user list),word2vec(base on user's question list, wrong and correct tag list ),timedelta from many perspective,last same part groups features.</li>\n</ul>\n<h3>gbm benchmark</h3>\n<ul>\n<li>We compared lightgbm ,xgboost,catboost,catboost is the best for the training and inference speed,and memory consuming.When train the full data,lightgbm need over 100 hours with my AMD Ryzen ThreadRipper 3970X,xgboost always have out of memory error even using dask with 4 RTX 3090.</li>\n</ul>\n<h3>strong features and interesting finding by qyxs</h3>\n<ul>\n<li>1.  the history difficulty statistics features of user who had correct/wrong answers, boost almost 0.003</li>\n</ul>\n<pre><code>tmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n</code></pre>\n<ul>\n<li>2.  focus on the records about the current part of user connect with last same part, generate the features include answer correct ratio, time diff, frequency etc, boost almost 0.002</li>\n</ul>\n<pre><code>for_question_df['rank_part'] = for_question_df.groupby(['user_id', 'part'])['timestamp'].rank(method='first')\nfor_question_df['rank_user'] = for_question_df.groupby(['user_id'])['timestamp'].rank(method='first')\nfor_question_df['rank_diff'] = for_question_df['rank_user'] - for_question_df['rank_part']\nfor_question_df['part_times'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense')\nfor_question_df['rank_diff'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense', ascending=False)\n\nlast_part = for_question_df[for_question_df['rank_diff']==1]\npart_times = for_question_df.groupby(['user_id', 'part'])['part_times'].agg([['part_times', 'max']]).reset_index()\n\nlast_part_df = last_part.groupby(['user_id', 'part'])['answered_correctly'].agg([['last_continue_part_ratio', 'mean'], ['last_continue_part_cnt', 'count']]).reset_index()\nlast_part_time = last_part.groupby(['user_id', 'part'])['timestamp'].agg([['last_continue_part_time_start', 'min'], ['last_continue_part_time_end', 'max']]).reset_index()\nlast_part_df = last_part_df.merge(last_part_time, on=['user_id', 'part'], how='left')\nlast_part_df = last_part_df.merge(part_times, on=['user_id', 'part'], how='left')\nlast_part_df['part_time_diff'] = last_part_df['last_continue_part_time_end'] - last_part_df['last_continue_part_time_start']\nlast_part_df['part_time_freq'] = last_part_df['last_continue_part_cnt']/last_part_df['part_time_diff']\n\nfor_train = for_train.merge(last_part_df, on=['user_id', 'part'], how='left')\nfor_train['last_continue_part_time_start'] = for_train['timestamp'] - for_train['last_continue_part_time_start']\nfor_train['last_continue_part_time_end'] = for_train['timestamp'] - for_train['last_continue_part_time_end']\n</code></pre>\n<ul>\n<li>3.  the answer correctly ratio of each question under differenct user abilititys (split for 11 bins), boost almost 0.001</li>\n</ul>\n<pre><code>for_question_df['user_ability'] = for_question_df.groupby('user_id')['answered_correctly'].transform('mean').round(1)\ntmp_df = for_question_df.pivot_table(index='content_id', columns='user_ability', values='answered_correctly', aggfunc='mean').reset_index()\ntmp_df.columns = ['content_id'] + [f'c_mean_{i}_ratio' for i in range(11)]\nfor_train = for_train.merge(tmp_df, on='content_id', how='left')\n</code></pre>\n<ul>\n<li>Some interseting points:<ul>\n<li>1.they would watch lecture after users had wrong answers, so we could generated some features from this. LB is not improved caused by the lectures info in next group maybe.</li>\n<li>2.the content_id such as 0-195， 7851-7984 etc, then are all same in one part and continuous with each other，we could build a new bundle to generate features</li></ul></li>\n</ul>\n<h3>strong features and interesting finding by ethan</h3>\n<ul>\n<li>1.  user's behavior in last 1,5,…,60 minutes, 0.001 boost</li>\n</ul>\n<pre><code>for w in [1, 5, 10, 15, 30, 45, 60]:\n    print(w)\n    tmp = q_logs[q_logs['timestamp']&gt;=(q_logs['end_time']-w*60*1000)].copy()\n    group_df = tmp.groupby(['user_id'])['content_id'].agg([['user_content_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['part'].agg([['user_part_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['answered_correctly'].agg([['user_correct_raito_in_last{}mins'.format(w), 'mean']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n</code></pre>\n<ul>\n<li>2. \"users' ablility\" statistics in each question, seperately by \"answered_correctly\"(0/1), 0.002 boost</li>\n</ul>\n<pre><code>cc = q_logs.groupby(['user_id'])['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ngg = q_logs[['user_id', 'content_id', 'answered_correctly']].merge(cc, on=['user_id'], how='left')\n\ngroup_df1 = gg[gg['answered_correctly']==1].groupby(['content_id'])['corr_ratio'].agg([['question_correct_user_ablility_min', 'min'], \n                                                                                       ['question_correct_user_ablility_max', 'max'], \n                                                                                       ['question_correct_user_ablility_mean', 'mean'], \n                                                                                       ['question_correct_user_ablility_skew', 'skew'],\n                                                                                       ['question_correct_user_ablility_med', 'median'],\n                                                                                       ['question_correct_user_ablility_std', 'std']]).reset_index()\ngroup_df2 = gg[gg['answered_correctly']==0].groupby(['content_id'])['corr_ratio'].agg([['question_wrong_user_ablility_min','min'], \n                                                                                       ['question_wrong_user_ablility_max','max'],\n                                                                                       ['question_wrong_user_ablility_mean','mean'],\n                                                                                       ['question_wrong_user_ablility_skew','skew'],\n                                                                                       ['question_wrong_user_ablility_med','median'],\n                                                                                       ['question_wrong_user_ablility_std','std']]).reset_index()\n</code></pre>\n<ul>\n<li>3. \"lagtime\" statistics in each question, seperately by \"answered_correctly\"(0/1), means the distribution of users' preprare time for answering this question correctly, about 0.001 boost</li>\n</ul>\n<pre><code>user_task_timestamp = q_logs[['user_id', 'task_container_id', 'timestamp']].drop_duplicates()\nuser_task_timestamp['lag_time'] = user_task_timestamp['timestamp'] - user_task_timestamp.groupby(['user_id'])['timestamp'].shift(1)\ntmp = q_logs[['user_id', 'task_container_id', 'content_id', 'answered_correctly']].merge(user_task_timestamp.drop(['timestamp'], axis=1), on=['user_id', 'task_container_id'], how='left')\ngroup_df = tmp[tmp['answered_correctly']==1].groupby(['content_id'])['lag_time'].agg([['c_lag_time_mean', 'mean'],\n                                                                                     ['c_lag_time_std', 'std'],\n                                                                                     ['c_lag_time_max', 'max'],\n                                                                                     ['c_lag_time_min', 'min'],\n                                                                                     ['c_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n\ngroup_df = tmp[tmp['answered_correctly']==0].groupby(['content_id'])['lag_time'].agg([['w_lag_time_mean', 'mean'],\n                                                                                       ['w_lag_time_std', 'std'],\n                                                                                       ['w_lag_time_max', 'max'],\n                                                                                       ['w_lag_time_min', 'min'],\n                                                                                       ['w_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n</code></pre>\n<h3>feature list</h3>\n<pre><code>['content_id',\n 'prior_question_elapsed_time',\n 'prior_question_had_explanation',\n 'correct_answer',\n 'user_count',\n 'user_sum',\n 'user_mean',\n 'item_count',\n 'item_sum',\n 'item_mean',\n 'answer_ratio_0',\n 'answer_ratio_1',\n 'answer_ratio_2',\n 'bundle_id',\n 'part',\n 'le_tag',\n 'question_correct_user_ablility_mean',\n 'question_correct_user_ablility_median',\n 'question_wrong_user_ablility_mean',\n 'question_wrong_user_ablility_median',\n 'word2vec_0',\n 'word2vec_1',\n 'word2vec_2',\n 'word2vec_3',\n 'word2vec_4',\n 'svd_0',\n 'svd_1',\n 'svd_2',\n 'svd_3',\n 'svd_4',\n 'tags_w2v_correct_mean_0',\n 'tags_w2v_wrong_mean_0',\n 'tags_w2v_correct_mean_1',\n 'tags_w2v_wrong_mean_1',\n 'tags_w2v_correct_mean_2',\n 'tags_w2v_wrong_mean_2',\n 'tags_w2v_correct_mean_3',\n 'tags_w2v_wrong_mean_3',\n 'tags_w2v_correct_mean_4',\n 'tags_w2v_wrong_mean_4',\n 'real_time_wrong_mean',\n 'real_time_wrong_median',\n 'real_time_correct_mean',\n 'real_time_correct_median',\n 'task_set_distance_wrong_mean',\n 'task_set_distance_wrong_median',\n 'task_set_distance_correct_mean',\n 'task_set_distance_correct_median',\n 'mean_0_ratio',\n 'mean_1_ratio',\n 'mean_3_ratio',\n 'mean_4_ratio',\n 'mean_5_ratio',\n 'mean_6_ratio',\n 'mean_7_ratio',\n 'mean_8_ratio',\n 'mean_9_ratio',\n 'mean_10_ratio',\n 'user_d1',\n 'user_d2',\n 'task_set_distance',\n 'user_diff_mean',\n 'user_diff_std',\n 'user_diff_min',\n 'user_diff_max',\n 'task_set_item_mean',\n 'task_set_item_min',\n 'task_set_item_max',\n 'task_set_distance2',\n 'task_distance_shift',\n 'task_set_distance_diff',\n 'task_distance_diff_shift',\n 'container_mean_1',\n 'container_mean_5',\n 'container_std_5',\n 'container_mean_10',\n 'container_std_10',\n 'container_mean_20',\n 'container_std_20',\n 'container_mean_30',\n 'container_std_30',\n 'container_mean_40',\n 'container_std_40',\n 'prior_question_elapsed_time_mean_1',\n 'prior_question_elapsed_time_mean_5',\n 'prior_question_elapsed_time_mean_10',\n 'prior_question_elapsed_time_mean_20',\n 'prior_question_elapsed_time_mean_30',\n 'prior_question_elapsed_time_mean_40',\n 'item_mean_mean_30',\n 'item_mean_mean_40',\n 'task_set_distance_mean_1',\n 'task_set_distance_mean_5',\n 'task_set_distance_mean_10',\n 'task_set_distance_mean_20',\n 'task_set_distance_mean_30',\n 'begin_time_diff',\n 'end_time_diff',\n 'part_time_diff_mean',\n 'part_session_mean',\n 'part_session_sum',\n 'part_session_count',\n 'full_group0_item_mean_mean',\n 'full_group0_item_mean_median',\n 'full_group0_task_set_distance_median',\n 'full_group0_timestamp_mean',\n 'full_group0_timestamp_median',\n 'full_group1_item_mean_mean',\n 'full_group1_item_mean_median',\n 'full_group1_task_set_distance_median',\n 'full_group1_timestamp_median',\n 'part_sum',\n 'part_count',\n 'part_mean',\n 'part_sum_global_ratio',\n 'part_sum_1',\n 'part_sum_5',\n 'part_mean_5',\n 'part_sum_10',\n 'part_mean_10',\n 'cum_answer0_mean_item_mean',\n 'cum_answer0_median_item_mean',\n 'cum_answer0_median_task_set_distance',\n 'cum_answer1_mean_item_mean',\n 'cum_answer1_median_item_mean',\n 'cum_answer1_mean_task_set_distance',\n 'cum_answer1_median_task_set_distance',\n 'cum_answer0_time_diff',\n 'cum_answer1_time_diff',\n 'global_task_set_shift1',\n 'global_task_set_shift2',\n 'global_task_set_shift4',\n 'global_task_set_shift5',\n 'cum_answer0_mean_wrong_time_diff',\n 'cum_answer0_median_wrong_time_diff',\n 'cum_answer1_mean_right_time_diff',\n 'content_correct_mean',\n 'content_correct_sum',\n 'content_correct_count',\n 'hard_answer0_time',\n 'hard_answer1_time',\n 'full_bundle_item_mean_mean',\n 'full_bundle_item_mean_median',\n 'full_bundle_task_set_distance_mean',\n 'full_bundle_task_set_distance_median',\n 'full_bundle_timestamp_mean',\n 'full_bundle_timestamp_median',\n 'bundle_sum',\n 'bundle_mean',\n 'bundle_count',\n 'user_trend_mean',\n 'user_trend_median',\n 'user_trend_roll_user_ans_sum',\n 'user_trend_roll_user_ans_mean',\n 'user_trend_roll_user_ans_count',\n 'user_trend_roll_item_ans_mean',\n 'user_trend_roll_item_ans_count',\n 'div_ratio1',\n 'div_ratio2',\n 'div_ratio3',\n 'new_Feat0',\n 'new_Feat1',\n 'new_Feat2',\n 'new_Feat3',\n 'part_time_wrong_div',\n 'part_time_right_div',\n 'diff_lag_median_div',\n 'diff_item_median_div',\n 'diff_time_median_div',\n 'diff_item_mean_div',\n 'diff_task_set_mean_div',\n 'diff_timestamp_mean_div',\n 'last_20_frequent_answer',\n 'last_20_frequent_answer_count',\n 'last_20_frequent_answer_mean',\n 'last_20_frequent_answer_sum',\n 'last_user_same_answer_tf',\n 'last_item_same_answer_tf',\n 'last_right_time_diff',\n 'last_wrong_time_diff',\n 'last_5_part_time_div',\n 'last_10_part_time_div',\n 'last_20_part_time_div']\n</code></pre>\n<h1>Transformer(LB 0.808)</h1>\n<ul>\n<li><p>Used questions only. Lectures did't improve our validation score.</p></li>\n<li><p>In training window size 800 (the bigger,the better), in inference window size 300(the bigger,the better,but time consuming).</p></li>\n<li><p>Optimizer : Adam  with lr = 8e-4, beta1 = 0.9, beta2 = 0.999 with warmup steps to 4000.</p>\n<ul>\n<li>Without warmup steps, didn't converge.</li></ul></li>\n<li><p>Number of layers = 4, dimension of the model = 256, dimension of the FFN = 2048.</p></li>\n<li><p>Batch size=80</p>\n<ul>\n<li>With smaller size,  didn't converge.</li></ul></li>\n<li><p>Dropout = 0</p></li>\n<li><p>Postion encoding : Axial Positional Embedding</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.12180\" target=\"_blank\">https://arxiv.org/abs/1912.12180</a></li>\n<li><a href=\"https://github.com/lucidrains/axial-positional-embedding\" target=\"_blank\">https://github.com/lucidrains/axial-positional-embedding</a></li>\n<li>This encoding improved score significantly.</li></ul></li>\n<li><p>Augmentation : Mixup-Transformer</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2010.02394\" target=\"_blank\">https://arxiv.org/abs/2010.02394</a></li></ul></li>\n<li><p>Inputs</p>\n<ul>\n<li>question_id</li>\n<li>part</li>\n<li>prior_question_elapsed_time / 1000</li>\n<li>lagtime<ul>\n<li>log1p((timestamp_t - timestamp_(t-1)) /1000 / 60)</li>\n<li>This feature improved the score significantly</li></ul></li>\n<li>answered_correctly</li>\n<li>GBT feats (imortance top N)</li></ul></li>\n<li><p>model image<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F6743aa3ef1403f4084dd4fa3d8372702%2FScreen%20Shot%202021-01-08%20at%2011.34.11.png?generation=1610073332748004&amp;alt=media\" alt=\"\"></p></li>\n</ul>\n<h1>Ensemble(LB 0.812)</h1>\n<ul>\n<li>Finally we use one catboost and two transformers for ensemble due to the inference time limitation.Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.</li>\n</ul>",
      "rawMarkdown": "Congrats to all medal teams and new Grandmasters,Masters,Experts.Thanks to Organizers and kaggle for such a good competition,it shows that kaggle competition is not just a game but also can be a useful machine learning project.\n\n# Team\n* At first we have three  teams individually, tomoyo and me, ethan and qyxs , wrb0312.We focus on feature engineering and optimization before wrb0312 joined us,I think we made many good features but neural network dominated this competition.wrb0312 did a great job even he use transformer for the first time.After wrb0312 joined us there are only ten days left,we focus on ensembling our models for inference,and also improved transformer very much.Our team members are from china and japan, it's very interesting to see we use chinese,japanese,english mixed-language to communicate.Greate job everyone!\n\n# Optimization \n* For GBM features, rather than using many dictionary to save features' data, we developed a nubma-based framework to speed up feature engineering process and online calculation. Firstly, the data are sorted by ['user_id', 'timestamp', 'content_id'] and split into different arrays. Then we created features in different array via self-designed rolling function or self-designed cumlative function. Actually, it provides us a very flexible way to create features and test it. In 10m data, the feature engineering process needs only 5 minutes to finish it.\n\n* Some examples are listed as below. \n```\nfrom tqdm import tqdm\nfrom numba import jit,njit\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\nimport gc\nfrom multiprocessing import Process, Manager,Pool\nfrom functools import partial\nfrom numba import prange\nimport numpy as np\nimport pandas as pd\nfrom numba import types\nfrom numba.typed import Dict\nimport functools, time\nfrom numba.typed import List\n\n\ndef timeit(f):\n    def wrap(*args, **kwargs):\n        time1 = time.time()\n        ret = f(*args, **kwargs)\n        time2 = time.time()\n        print('{:s} function took {:.3f} s'.format(f.__name__, np.round(time2-time1, 2)))\n\n        return ret\n    return wrap\n\n\ndef rolling_feat_group(train, col_used):\n    a = train[col_used].values\n    ind = np.lexsort((a[:,2],a[:,1],a[:,0]))\n    a = a[ind]\n    g = np.split(a, np.unique(a[:, 0], return_index=True)[1][1:])\n    return g, ind, col_used\n\n@jit(nopython = True, fastmath = True)\ndef rolling_cal(arr, step, window = 5, shift_ = 1):\n    m = 2\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    beg = window\n    for i in step: \n        tmp = arr_[beg-window:beg]\n        ret[beg - window:(beg - window + i), 0] = np.nanmean(tmp)\n        ret[beg - window:(beg - window + i), 1] = np.nansum(tmp)\n        beg += i\n    return ret\n\n\n@jit(nopython = True, fastmath = True)\ndef rolling_time_cal(arr, window = 5, shift_ = 1):\n    m = 1\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    for i in range(0,arr.shape[0], 1): \n        tmp = arr_[i:i+window+1]\n        ret[i, 0] = np.nanmean(tmp)\n    return ret\n\ndef rolling_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    step = np.unique(tmp_g[:, 1], return_counts=True)[1]\n    for window_size in shift_period:\n        tmp = rolling_cal(tmp_g[:, 2], step, window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\ndef rolling_time_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    for window_size in shift_period:\n        tmp = rolling_time_cal(tmp_g[:, 2], window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\n\ndef rolling_feat_cal(tmp_g, name_dict, global_period):\n    answer_idx = name_dict.index('answered_correctly')\n    prior_idx = name_dict.index('prior_question_elapsed_time')\n    item_mean_idx = name_dict.index('item_mean')\n    task_set_idx = name_dict.index('task_set_distance')\n    tmp_res1 = rolling_cal_wrap(tmp_g[:,[0,1, answer_idx]], global_period)\n    tmp_res2 = rolling_time_cal_wrap(tmp_g[:,[0,1, prior_idx]], global_period)\n    tmp_res3 = rolling_time_cal_wrap(tmp_g[:,[0,1, item_mean_idx]], global_period)\n    tmp_res4 = rolling_time_cal_wrap(tmp_g[:,[0,1, task_set_idx]], global_period)\n    tmp_res = np.concatenate([tmp_res1, tmp_res2, tmp_res3, tmp_res4], axis = 1)\n    return tmp_res\n```\n\n* If anyone interested in how to create features via numba-framework, Tomoyo publiced his full GBM pipeline in github([https://github.com/ZiwenYeee/Riiid-numba-framework](url))\n\n# Catboost(LB 0.807)\n### summary\n* We created 183 features for final catboost model,including some original features,global statistics(item base),cumulative and rolling statistics(user base),tfidf-svd(base on question's user list),word2vec(base on user's question list, wrong and correct tag list ),timedelta from many perspective,last same part groups features.\n\n### gbm benchmark\n* We compared lightgbm ,xgboost,catboost,catboost is the best for the training and inference speed,and memory consuming.When train the full data,lightgbm need over 100 hours with my AMD Ryzen ThreadRipper 3970X,xgboost always have out of memory error even using dask with 4 RTX 3090.\n\n### strong features and interesting finding by qyxs\n* 1.  the history difficulty statistics features of user who had correct/wrong answers, boost almost 0.003\n```\ntmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n```\n\n* 2.  focus on the records about the current part of user connect with last same part, generate the features include answer correct ratio, time diff, frequency etc, boost almost 0.002\n```\nfor_question_df['rank_part'] = for_question_df.groupby(['user_id', 'part'])['timestamp'].rank(method='first')\nfor_question_df['rank_user'] = for_question_df.groupby(['user_id'])['timestamp'].rank(method='first')\nfor_question_df['rank_diff'] = for_question_df['rank_user'] - for_question_df['rank_part']\nfor_question_df['part_times'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense')\nfor_question_df['rank_diff'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense', ascending=False)\n\nlast_part = for_question_df[for_question_df['rank_diff']==1]\npart_times = for_question_df.groupby(['user_id', 'part'])['part_times'].agg([['part_times', 'max']]).reset_index()\n\nlast_part_df = last_part.groupby(['user_id', 'part'])['answered_correctly'].agg([['last_continue_part_ratio', 'mean'], ['last_continue_part_cnt', 'count']]).reset_index()\nlast_part_time = last_part.groupby(['user_id', 'part'])['timestamp'].agg([['last_continue_part_time_start', 'min'], ['last_continue_part_time_end', 'max']]).reset_index()\nlast_part_df = last_part_df.merge(last_part_time, on=['user_id', 'part'], how='left')\nlast_part_df = last_part_df.merge(part_times, on=['user_id', 'part'], how='left')\nlast_part_df['part_time_diff'] = last_part_df['last_continue_part_time_end'] - last_part_df['last_continue_part_time_start']\nlast_part_df['part_time_freq'] = last_part_df['last_continue_part_cnt']/last_part_df['part_time_diff']\n\nfor_train = for_train.merge(last_part_df, on=['user_id', 'part'], how='left')\nfor_train['last_continue_part_time_start'] = for_train['timestamp'] - for_train['last_continue_part_time_start']\nfor_train['last_continue_part_time_end'] = for_train['timestamp'] - for_train['last_continue_part_time_end']\n```\n\n* 3.  the answer correctly ratio of each question under differenct user abilititys (split for 11 bins), boost almost 0.001\n```\nfor_question_df['user_ability'] = for_question_df.groupby('user_id')['answered_correctly'].transform('mean').round(1)\ntmp_df = for_question_df.pivot_table(index='content_id', columns='user_ability', values='answered_correctly', aggfunc='mean').reset_index()\ntmp_df.columns = ['content_id'] + [f'c_mean_{i}_ratio' for i in range(11)]\nfor_train = for_train.merge(tmp_df, on='content_id', how='left')\n```\n\n* Some interseting points:\n * 1.they would watch lecture after users had wrong answers, so we could generated some features from this. LB is not improved caused by the lectures info in next group maybe.\n *  2.the content_id such as 0-195， 7851-7984 etc, then are all same in one part and continuous with each other，we could build a new bundle to generate features\n\n### strong features and interesting finding by ethan\n* 1.  user's behavior in last 1,5,...,60 minutes, 0.001 boost\n```\nfor w in [1, 5, 10, 15, 30, 45, 60]:\n    print(w)\n    tmp = q_logs[q_logs['timestamp']>=(q_logs['end_time']-w*60*1000)].copy()\n    group_df = tmp.groupby(['user_id'])['content_id'].agg([['user_content_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['part'].agg([['user_part_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['answered_correctly'].agg([['user_correct_raito_in_last{}mins'.format(w), 'mean']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n ```\n \n* 2. \"users' ablility\" statistics in each question, seperately by \"answered_correctly\"(0/1), 0.002 boost\n```\ncc = q_logs.groupby(['user_id'])['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ngg = q_logs[['user_id', 'content_id', 'answered_correctly']].merge(cc, on=['user_id'], how='left')\n\ngroup_df1 = gg[gg['answered_correctly']==1].groupby(['content_id'])['corr_ratio'].agg([['question_correct_user_ablility_min', 'min'], \n                                                                                       ['question_correct_user_ablility_max', 'max'], \n                                                                                       ['question_correct_user_ablility_mean', 'mean'], \n                                                                                       ['question_correct_user_ablility_skew', 'skew'],\n                                                                                       ['question_correct_user_ablility_med', 'median'],\n                                                                                       ['question_correct_user_ablility_std', 'std']]).reset_index()\ngroup_df2 = gg[gg['answered_correctly']==0].groupby(['content_id'])['corr_ratio'].agg([['question_wrong_user_ablility_min','min'], \n                                                                                       ['question_wrong_user_ablility_max','max'],\n                                                                                       ['question_wrong_user_ablility_mean','mean'],\n                                                                                       ['question_wrong_user_ablility_skew','skew'],\n                                                                                       ['question_wrong_user_ablility_med','median'],\n                                                                                       ['question_wrong_user_ablility_std','std']]).reset_index()\n```\n                        \n* 3. \"lagtime\" statistics in each question, seperately by \"answered_correctly\"(0/1), means the distribution of users' preprare time for answering this question correctly, about 0.001 boost\n```\nuser_task_timestamp = q_logs[['user_id', 'task_container_id', 'timestamp']].drop_duplicates()\nuser_task_timestamp['lag_time'] = user_task_timestamp['timestamp'] - user_task_timestamp.groupby(['user_id'])['timestamp'].shift(1)\ntmp = q_logs[['user_id', 'task_container_id', 'content_id', 'answered_correctly']].merge(user_task_timestamp.drop(['timestamp'], axis=1), on=['user_id', 'task_container_id'], how='left')\ngroup_df = tmp[tmp['answered_correctly']==1].groupby(['content_id'])['lag_time'].agg([['c_lag_time_mean', 'mean'],\n                                                                                     ['c_lag_time_std', 'std'],\n                                                                                     ['c_lag_time_max', 'max'],\n                                                                                     ['c_lag_time_min', 'min'],\n                                                                                     ['c_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n\ngroup_df = tmp[tmp['answered_correctly']==0].groupby(['content_id'])['lag_time'].agg([['w_lag_time_mean', 'mean'],\n                                                                                       ['w_lag_time_std', 'std'],\n                                                                                       ['w_lag_time_max', 'max'],\n                                                                                       ['w_lag_time_min', 'min'],\n                                                                                       ['w_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n```\n\n### feature list\n```\n['content_id',\n 'prior_question_elapsed_time',\n 'prior_question_had_explanation',\n 'correct_answer',\n 'user_count',\n 'user_sum',\n 'user_mean',\n 'item_count',\n 'item_sum',\n 'item_mean',\n 'answer_ratio_0',\n 'answer_ratio_1',\n 'answer_ratio_2',\n 'bundle_id',\n 'part',\n 'le_tag',\n 'question_correct_user_ablility_mean',\n 'question_correct_user_ablility_median',\n 'question_wrong_user_ablility_mean',\n 'question_wrong_user_ablility_median',\n 'word2vec_0',\n 'word2vec_1',\n 'word2vec_2',\n 'word2vec_3',\n 'word2vec_4',\n 'svd_0',\n 'svd_1',\n 'svd_2',\n 'svd_3',\n 'svd_4',\n 'tags_w2v_correct_mean_0',\n 'tags_w2v_wrong_mean_0',\n 'tags_w2v_correct_mean_1',\n 'tags_w2v_wrong_mean_1',\n 'tags_w2v_correct_mean_2',\n 'tags_w2v_wrong_mean_2',\n 'tags_w2v_correct_mean_3',\n 'tags_w2v_wrong_mean_3',\n 'tags_w2v_correct_mean_4',\n 'tags_w2v_wrong_mean_4',\n 'real_time_wrong_mean',\n 'real_time_wrong_median',\n 'real_time_correct_mean',\n 'real_time_correct_median',\n 'task_set_distance_wrong_mean',\n 'task_set_distance_wrong_median',\n 'task_set_distance_correct_mean',\n 'task_set_distance_correct_median',\n 'mean_0_ratio',\n 'mean_1_ratio',\n 'mean_3_ratio',\n 'mean_4_ratio',\n 'mean_5_ratio',\n 'mean_6_ratio',\n 'mean_7_ratio',\n 'mean_8_ratio',\n 'mean_9_ratio',\n 'mean_10_ratio',\n 'user_d1',\n 'user_d2',\n 'task_set_distance',\n 'user_diff_mean',\n 'user_diff_std',\n 'user_diff_min',\n 'user_diff_max',\n 'task_set_item_mean',\n 'task_set_item_min',\n 'task_set_item_max',\n 'task_set_distance2',\n 'task_distance_shift',\n 'task_set_distance_diff',\n 'task_distance_diff_shift',\n 'container_mean_1',\n 'container_mean_5',\n 'container_std_5',\n 'container_mean_10',\n 'container_std_10',\n 'container_mean_20',\n 'container_std_20',\n 'container_mean_30',\n 'container_std_30',\n 'container_mean_40',\n 'container_std_40',\n 'prior_question_elapsed_time_mean_1',\n 'prior_question_elapsed_time_mean_5',\n 'prior_question_elapsed_time_mean_10',\n 'prior_question_elapsed_time_mean_20',\n 'prior_question_elapsed_time_mean_30',\n 'prior_question_elapsed_time_mean_40',\n 'item_mean_mean_30',\n 'item_mean_mean_40',\n 'task_set_distance_mean_1',\n 'task_set_distance_mean_5',\n 'task_set_distance_mean_10',\n 'task_set_distance_mean_20',\n 'task_set_distance_mean_30',\n 'begin_time_diff',\n 'end_time_diff',\n 'part_time_diff_mean',\n 'part_session_mean',\n 'part_session_sum',\n 'part_session_count',\n 'full_group0_item_mean_mean',\n 'full_group0_item_mean_median',\n 'full_group0_task_set_distance_median',\n 'full_group0_timestamp_mean',\n 'full_group0_timestamp_median',\n 'full_group1_item_mean_mean',\n 'full_group1_item_mean_median',\n 'full_group1_task_set_distance_median',\n 'full_group1_timestamp_median',\n 'part_sum',\n 'part_count',\n 'part_mean',\n 'part_sum_global_ratio',\n 'part_sum_1',\n 'part_sum_5',\n 'part_mean_5',\n 'part_sum_10',\n 'part_mean_10',\n 'cum_answer0_mean_item_mean',\n 'cum_answer0_median_item_mean',\n 'cum_answer0_median_task_set_distance',\n 'cum_answer1_mean_item_mean',\n 'cum_answer1_median_item_mean',\n 'cum_answer1_mean_task_set_distance',\n 'cum_answer1_median_task_set_distance',\n 'cum_answer0_time_diff',\n 'cum_answer1_time_diff',\n 'global_task_set_shift1',\n 'global_task_set_shift2',\n 'global_task_set_shift4',\n 'global_task_set_shift5',\n 'cum_answer0_mean_wrong_time_diff',\n 'cum_answer0_median_wrong_time_diff',\n 'cum_answer1_mean_right_time_diff',\n 'content_correct_mean',\n 'content_correct_sum',\n 'content_correct_count',\n 'hard_answer0_time',\n 'hard_answer1_time',\n 'full_bundle_item_mean_mean',\n 'full_bundle_item_mean_median',\n 'full_bundle_task_set_distance_mean',\n 'full_bundle_task_set_distance_median',\n 'full_bundle_timestamp_mean',\n 'full_bundle_timestamp_median',\n 'bundle_sum',\n 'bundle_mean',\n 'bundle_count',\n 'user_trend_mean',\n 'user_trend_median',\n 'user_trend_roll_user_ans_sum',\n 'user_trend_roll_user_ans_mean',\n 'user_trend_roll_user_ans_count',\n 'user_trend_roll_item_ans_mean',\n 'user_trend_roll_item_ans_count',\n 'div_ratio1',\n 'div_ratio2',\n 'div_ratio3',\n 'new_Feat0',\n 'new_Feat1',\n 'new_Feat2',\n 'new_Feat3',\n 'part_time_wrong_div',\n 'part_time_right_div',\n 'diff_lag_median_div',\n 'diff_item_median_div',\n 'diff_time_median_div',\n 'diff_item_mean_div',\n 'diff_task_set_mean_div',\n 'diff_timestamp_mean_div',\n 'last_20_frequent_answer',\n 'last_20_frequent_answer_count',\n 'last_20_frequent_answer_mean',\n 'last_20_frequent_answer_sum',\n 'last_user_same_answer_tf',\n 'last_item_same_answer_tf',\n 'last_right_time_diff',\n 'last_wrong_time_diff',\n 'last_5_part_time_div',\n 'last_10_part_time_div',\n 'last_20_part_time_div']\n```\n# Transformer(LB 0.808)\n* Used questions only. Lectures did't improve our validation score.\n*  In training window size 800 (the bigger,the better), in inference window size 300(the bigger,the better,but time consuming).\n* Optimizer : Adam  with lr = 8e-4, beta1 = 0.9, beta2 = 0.999 with warmup steps to 4000.\n  * Without warmup steps, didn't converge.\n* Number of layers = 4, dimension of the model = 256, dimension of the FFN = 2048.\n* Batch size=80\n  * With smaller size,  didn't converge.\n* Dropout = 0\n* Postion encoding : Axial Positional Embedding\n  * https://arxiv.org/abs/1912.12180\n  * https://github.com/lucidrains/axial-positional-embedding\n  * This encoding improved score significantly.\n* Augmentation : Mixup-Transformer\n  * https://arxiv.org/abs/2010.02394\n\n* Inputs\n  * question_id\n  * part\n  * prior_question_elapsed_time / 1000\n  * lagtime\n      * log1p((timestamp_t - timestamp_(t-1)) /1000 / 60)\n      * This feature improved the score significantly\n  * answered_correctly\n  * GBT feats (imortance top N)\n* model image\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F6743aa3ef1403f4084dd4fa3d8372702%2FScreen%20Shot%202021-01-08%20at%2011.34.11.png?generation=1610073332748004&alt=media)\n\n# Ensemble(LB 0.812)\n* Finally we use one catboost and two transformers for ensemble due to the inference time limitation.Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.",
      "votes": null
    },
    {
      "id": "1143632",
      "postDate": "01/08/2021 01:22:39",
      "content": "<p>I've never used nubma. Really inspiring. Thank you for sharing.</p>",
      "rawMarkdown": "I've never used nubma. Really inspiring. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1143636",
      "postDate": "01/08/2021 01:24:39",
      "content": "<p>Thanks for your work and effort, bro. Looking forward to future cooperation!</p>\n<p>If anyone interested in my numba-based framework for GBM, leave a message below. I'm considering write a detailed explanation about it.</p>",
      "rawMarkdown": "Thanks for your work and effort, bro. Looking forward to future cooperation!\n\nIf anyone interested in my numba-based framework for GBM, leave a message below. I'm considering write a detailed explanation about it.",
      "votes": null
    },
    {
      "id": "1143637",
      "postDate": "01/08/2021 01:25:59",
      "content": "<p>I have a strong interest on it 👍</p>",
      "rawMarkdown": "I have a strong interest on it :+1:",
      "votes": null
    },
    {
      "id": "1143647",
      "postDate": "01/08/2021 01:38:16",
      "content": "<p>Congrats to all winners! Thanks for my teamates <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> <a href=\"https://www.kaggle.com/guziye\" target=\"_blank\">@guziye</a> <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/wrb0312\" target=\"_blank\">@wrb0312</a> ! Great teamwork!</p>",
      "rawMarkdown": "Congrats to all winners! Thanks for my teamates @senkin13 @guziye @juzqyxs @wrb0312 ! Great teamwork!",
      "votes": null
    },
    {
      "id": "1143653",
      "postDate": "01/08/2021 01:45:22",
      "content": "<p>Very interesting!</p>",
      "rawMarkdown": "Very interesting!",
      "votes": null
    },
    {
      "id": "1143665",
      "postDate": "01/08/2021 01:52:43",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> question for you - looks like catboost had a pretty strong performance based on features. Were you and your team not able to get a similar score with lightgbm? Also, does catboost have a similar type of feature importance plot? Just curious to know which of these features gave the greatest gains</p>",
      "rawMarkdown": "senkin13 question for you - looks like catboost had a pretty strong performance based on features. Were you and your team not able to get a similar score with lightgbm? Also, does catboost have a similar type of feature importance plot? Just curious to know which of these features gave the greatest gains",
      "votes": null
    },
    {
      "id": "1143669",
      "postDate": "01/08/2021 01:57:24",
      "content": "<p>Interested. Thanks.</p>",
      "rawMarkdown": "Interested. Thanks.",
      "votes": null
    },
    {
      "id": "1143670",
      "postDate": "01/08/2021 01:57:38",
      "content": "<p>I also interested in looking at it's src; Thanks a lot for sharing. I also used numba for computing the rolling FE's, the time dropped to &lt;3 secs for 100M rows as compared to more than 30 mins of Pandas! So very excited to explore numba!</p>\n<pre><code># &lt;3 secs for 100M rows on my lapi\nfrom numba import njit # no python basically\n\nwindow_width = 5\ndummy_window_5 = [np.nan]*4 # window_width-1\n\n@njit\ndef running_window_stats_window_width_5(vector):\n    # vector must be pre-padded with a zero\n    cumsum_vec = np.cumsum(vector, )\n    return (cumsum_vec[window_width:] - cumsum_vec[:-window_width]) / window_width\n</code></pre>",
      "rawMarkdown": "I also interested in looking at it's src; Thanks a lot for sharing. I also used numba for computing the rolling FE's, the time dropped to <3 secs for 100M rows as compared to more than 30 mins of Pandas! So very excited to explore numba!\n\n\n```\n# <3 secs for 100M rows on my lapi\nfrom numba import njit # no python basically\n\nwindow_width = 5\ndummy_window_5 = [np.nan]*4 # window_width-1\n\n@njit\ndef running_window_stats_window_width_5(vector):\n    # vector must be pre-padded with a zero\n    cumsum_vec = np.cumsum(vector, )\n    return (cumsum_vec[window_width:] - cumsum_vec[:-window_width]) / window_width\n```",
      "votes": null
    },
    {
      "id": "1143672",
      "postDate": "01/08/2021 02:00:15",
      "content": "<p>the reason we use catboost is just speed,I think lightgbm can get a similar score and feature gains</p>",
      "rawMarkdown": "the reason we use catboost is just speed,I think lightgbm can get a similar score and feature gains",
      "votes": null
    },
    {
      "id": "1143694",
      "postDate": "01/08/2021 02:23:11",
      "content": "<p><a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> Instead of \"lagtime\", our good features focus on \"user's ablilty\" and \"question's difficulty\". User's correctness in history represents \"user's ablility\" and Question's correctness(global) represents \"question's difficulty\". For example, \"quesions' difficulty\" statics(mean, max,min…) in each user's history and \"user's ablilty\" statics in each quesion. Especially create these features seperately by answered_correctly(0 or 1) will get great gains.</p>",
      "rawMarkdown": "rdizzl3 Instead of \"lagtime\", our good features focus on \"user's ablilty\" and \"question's difficulty\". User's correctness in history represents \"user's ablility\" and Question's correctness(global) represents \"question's difficulty\". For example, \"quesions' difficulty\" statics(mean, max,min...) in each user's history and \"user's ablilty\" statics in each quesion. Especially create these features seperately by answered_correctly(0 or 1) will get great gains.",
      "votes": null
    },
    {
      "id": "1143700",
      "postDate": "01/08/2021 02:28:24",
      "content": "<p>Great feature engineering work. Congrats on the results <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and team</p>",
      "rawMarkdown": "Great feature engineering work. Congrats on the results @senkin13 and team",
      "votes": null
    },
    {
      "id": "1143714",
      "postDate": "01/08/2021 02:55:01",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "1143789",
      "postDate": "01/08/2021 04:18:28",
      "content": "<p>Congratulations. Great models and strong finish! Your transformer is interesting. Did it do better than SAINT+? For reference, SAINT+ is pictured below! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8d4e260205a6a6599aefeee99bdf1b6a%2Fsaint2.png?generation=1610079431451477&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations. Great models and strong finish! Your transformer is interesting. Did it do better than SAINT+? For reference, SAINT+ is pictured below! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8d4e260205a6a6599aefeee99bdf1b6a%2Fsaint2.png?generation=1610079431451477&alt=media)",
      "votes": null
    },
    {
      "id": "1143822",
      "postDate": "01/08/2021 04:48:07",
      "content": "<p>Thanks!<br>\nI didn't try your picture's model architecture.<br>\nBut maybe we can get same score by SAINT+ architecture, with shorter inference time.<br>\nIn this competition I tried transformer first time, so I think it is possible to optimize the model structure more (it was hard for me).</p>",
      "rawMarkdown": "Thanks!\nI didn't try your picture's model architecture.\nBut maybe we can get same score by SAINT+ architecture, with shorter inference time.\nIn this competition I tried transformer first time, so I think it is possible to optimize the model structure more (it was hard for me).",
      "votes": null
    },
    {
      "id": "1144386",
      "postDate": "01/08/2021 12:38:33",
      "content": "<p>Congratz to all your team for the nice finish !</p>\n<blockquote>\n  <p>Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.</p>\n</blockquote>\n<p>That's heartbreaking, however do let people know if you manage to fix the bug &amp; improve your score</p>",
      "rawMarkdown": "Congratz to all your team for the nice finish !\n\n> Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.\n\nThat's heartbreaking, however do let people know if you manage to fix the bug & improve your score",
      "votes": null
    },
    {
      "id": "1144443",
      "postDate": "01/08/2021 13:17:22",
      "content": "<p>thanks for your care,we found the bug is from changing some features from float64 to float32 by mistake in inference notebook,if it is normal we still can not get gold medal,no regret! </p>",
      "rawMarkdown": "thanks for your care,we found the bug is from changing some features from float64 to float32 by mistake in inference notebook,if it is normal we still can not get gold medal,no regret!",
      "votes": null
    },
    {
      "id": "1144488",
      "postDate": "01/08/2021 13:38:37",
      "content": "<p>Good to hear, thanks !</p>",
      "rawMarkdown": "Good to hear, thanks !",
      "votes": null
    },
    {
      "id": "1144674",
      "postDate": "01/08/2021 15:50:38",
      "content": "<p>congratulations, a great innovative job! </p>",
      "rawMarkdown": "congratulations, a great innovative job!",
      "votes": null
    },
    {
      "id": "1148158",
      "postDate": "01/11/2021 00:21:38",
      "content": "<p>Congratulations for the great solution and nice finish! I have two questions:</p>\n<p>-1 In the example code, It seems like the corr_ratio are calculated by each user_id with forward looking. Model still had stable improvement between CV and LB?</p>\n<pre><code>tmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n</code></pre>\n<p>-2 How did you update features with numba functions in the test api? (didn't find numba functions for test updating in the gihub)</p>",
      "rawMarkdown": "Congratulations for the great solution and nice finish! I have two questions:\n\n-1 In the example code, It seems like the corr_ratio are calculated by each user_id with forward looking. Model still had stable improvement between CV and LB?\n\n```\ntmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n```\n\n-2 How did you update features with numba functions in the test api? (didn't find numba functions for test updating in the gihub)",
      "votes": null
    },
    {
      "id": "1148173",
      "postDate": "01/11/2021 00:52:02",
      "content": "<p>for question 2, We only update saving data and calculate with numba function rather than updating features. I loaded up one of our submit notebook and our test script in my github. <a href=\"https://github.com/ZiwenYeee/Riiid-numba-framework\" target=\"_blank\">https://github.com/ZiwenYeee/Riiid-numba-framework</a></p>",
      "rawMarkdown": "for question 2, We only update saving data and calculate with numba function rather than updating features. I loaded up one of our submit notebook and our test script in my github. https://github.com/ZiwenYeee/Riiid-numba-framework",
      "votes": null
    },
    {
      "id": "1149081",
      "postDate": "01/11/2021 15:20:19",
      "content": "<p>Congratulation for your amazing work ! Thanks for sharing your great ideas so much.</p>\n<p>Can you give additional details about w2v &amp; svd features ?</p>\n<p>1) While <code>tags_w2v_correct_mean_0 ... tags_w2v_wrong_mean_4</code> are self-explanatory and clear, I don't understand what <strong><code>word2vec_0 ... word2vec_4</code></strong> stand for. You said \"based on user's question list\" : is this feat the w2v of the current q_id ? Or the mean of w2v of each q_id in user's question list, just like tags ? <br>\n2) What about <strong><code>svd_0 ... svd_4</code></strong> ?</p>",
      "rawMarkdown": "Congratulation for your amazing work ! Thanks for sharing your great ideas so much.\n\nCan you give additional details about w2v & svd features ?\n\n1) While `tags_w2v_correct_mean_0 ... tags_w2v_wrong_mean_4` are self-explanatory and clear, I don't understand what **`word2vec_0 ... word2vec_4`** stand for. You said \"based on user's question list\" : is this feat the w2v of the current q_id ? Or the mean of w2v of each q_id in user's question list, just like tags ? \n2) What about **`svd_0 ... svd_4`** ?",
      "votes": null
    },
    {
      "id": "1150513",
      "postDate": "01/12/2021 16:29:50",
      "content": "<p>yeah,the improvement was stable.</p>",
      "rawMarkdown": "yeah,the improvement was stable.",
      "votes": null
    },
    {
      "id": "1150527",
      "postDate": "01/12/2021 16:35:15",
      "content": "<p>word2vec features are the w2v of current q_id, svd features get from the tfidf of users' history q_id.</p>",
      "rawMarkdown": "word2vec features are the w2v of current q_id, svd features get from the tfidf of users' history q_id.",
      "votes": null
    },
    {
      "id": "1150752",
      "postDate": "01/12/2021 20:07:40",
      "content": "<p>Makes sense. Many thanks, and congratz again !</p>",
      "rawMarkdown": "Makes sense. Many thanks, and congratz again !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143632,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 01:22:39",
      "content": "<p>I've never used nubma. Really inspiring. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143636,
      "author_name": "guziye",
      "author_url": "",
      "post_date": "01/08/2021 01:24:39",
      "content": "<p>Thanks for your work and effort, bro. Looking forward to future cooperation!</p>\n<p>If anyone interested in my numba-based framework for GBM, leave a message below. I'm considering write a detailed explanation about it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143637,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "01/08/2021 01:25:59",
          "content": "<p>I have a strong interest on it 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143669,
          "author_name": "scaomath",
          "author_url": "",
          "post_date": "01/08/2021 01:57:24",
          "content": "<p>Interested. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143670,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/08/2021 01:57:38",
          "content": "<p>I also interested in looking at it's src; Thanks a lot for sharing. I also used numba for computing the rolling FE's, the time dropped to &lt;3 secs for 100M rows as compared to more than 30 mins of Pandas! So very excited to explore numba!</p>\n<pre><code># &lt;3 secs for 100M rows on my lapi\nfrom numba import njit # no python basically\n\nwindow_width = 5\ndummy_window_5 = [np.nan]*4 # window_width-1\n\n@njit\ndef running_window_stats_window_width_5(vector):\n    # vector must be pre-padded with a zero\n    cumsum_vec = np.cumsum(vector, )\n    return (cumsum_vec[window_width:] - cumsum_vec[:-window_width]) / window_width\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143647,
      "author_name": "chenxin1991",
      "author_url": "",
      "post_date": "01/08/2021 01:38:16",
      "content": "<p>Congrats to all winners! Thanks for my teamates <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> <a href=\"https://www.kaggle.com/guziye\" target=\"_blank\">@guziye</a> <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> <a href=\"https://www.kaggle.com/wrb0312\" target=\"_blank\">@wrb0312</a> ! Great teamwork!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143653,
      "author_name": "fredegrec",
      "author_url": "",
      "post_date": "01/08/2021 01:45:22",
      "content": "<p>Very interesting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143665,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "01/08/2021 01:52:43",
      "content": "<p><a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> question for you - looks like catboost had a pretty strong performance based on features. Were you and your team not able to get a similar score with lightgbm? Also, does catboost have a similar type of feature importance plot? Just curious to know which of these features gave the greatest gains</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143672,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "01/08/2021 02:00:15",
          "content": "<p>the reason we use catboost is just speed,I think lightgbm can get a similar score and feature gains</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1143694,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "01/08/2021 02:23:11",
          "content": "<p><a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> Instead of \"lagtime\", our good features focus on \"user's ablilty\" and \"question's difficulty\". User's correctness in history represents \"user's ablility\" and Question's correctness(global) represents \"question's difficulty\". For example, \"quesions' difficulty\" statics(mean, max,min…) in each user's history and \"user's ablilty\" statics in each quesion. Especially create these features seperately by answered_correctly(0 or 1) will get great gains.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143700,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 02:28:24",
      "content": "<p>Great feature engineering work. Congrats on the results <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a> and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143714,
      "author_name": "a763337092",
      "author_url": "",
      "post_date": "01/08/2021 02:55:01",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1143789,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/08/2021 04:18:28",
      "content": "<p>Congratulations. Great models and strong finish! Your transformer is interesting. Did it do better than SAINT+? For reference, SAINT+ is pictured below! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8d4e260205a6a6599aefeee99bdf1b6a%2Fsaint2.png?generation=1610079431451477&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1143822,
          "author_name": "wrb0312",
          "author_url": "",
          "post_date": "01/08/2021 04:48:07",
          "content": "<p>Thanks!<br>\nI didn't try your picture's model architecture.<br>\nBut maybe we can get same score by SAINT+ architecture, with shorter inference time.<br>\nIn this competition I tried transformer first time, so I think it is possible to optimize the model structure more (it was hard for me).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144386,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "01/08/2021 12:38:33",
      "content": "<p>Congratz to all your team for the nice finish !</p>\n<blockquote>\n  <p>Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.</p>\n</blockquote>\n<p>That's heartbreaking, however do let people know if you manage to fix the bug &amp; improve your score</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144443,
          "author_name": "senkin13",
          "author_url": "",
          "post_date": "01/08/2021 13:17:22",
          "content": "<p>thanks for your care,we found the bug is from changing some features from float64 to float32 by mistake in inference notebook,if it is normal we still can not get gold medal,no regret! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144488,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "01/08/2021 13:38:37",
          "content": "<p>Good to hear, thanks !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144674,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/08/2021 15:50:38",
      "content": "<p>congratulations, a great innovative job! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148158,
      "author_name": "zonemercy",
      "author_url": "",
      "post_date": "01/11/2021 00:21:38",
      "content": "<p>Congratulations for the great solution and nice finish! I have two questions:</p>\n<p>-1 In the example code, It seems like the corr_ratio are calculated by each user_id with forward looking. Model still had stable improvement between CV and LB?</p>\n<pre><code>tmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n</code></pre>\n<p>-2 How did you update features with numba functions in the test api? (didn't find numba functions for test updating in the gihub)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148173,
          "author_name": "guziye",
          "author_url": "",
          "post_date": "01/11/2021 00:52:02",
          "content": "<p>for question 2, We only update saving data and calculate with numba function rather than updating features. I loaded up one of our submit notebook and our test script in my github. <a href=\"https://github.com/ZiwenYeee/Riiid-numba-framework\" target=\"_blank\">https://github.com/ZiwenYeee/Riiid-numba-framework</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150513,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "01/12/2021 16:29:50",
          "content": "<p>yeah,the improvement was stable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149081,
      "author_name": "johannhuber",
      "author_url": "",
      "post_date": "01/11/2021 15:20:19",
      "content": "<p>Congratulation for your amazing work ! Thanks for sharing your great ideas so much.</p>\n<p>Can you give additional details about w2v &amp; svd features ?</p>\n<p>1) While <code>tags_w2v_correct_mean_0 ... tags_w2v_wrong_mean_4</code> are self-explanatory and clear, I don't understand what <strong><code>word2vec_0 ... word2vec_4</code></strong> stand for. You said \"based on user's question list\" : is this feat the w2v of the current q_id ? Or the mean of w2v of each q_id in user's question list, just like tags ? <br>\n2) What about <strong><code>svd_0 ... svd_4</code></strong> ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1150527,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "01/12/2021 16:35:15",
          "content": "<p>word2vec features are the w2v of current q_id, svd features get from the tfidf of users' history q_id.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150752,
          "author_name": "johannhuber",
          "author_url": "",
          "post_date": "01/12/2021 20:07:40",
          "content": "<p>Makes sense. Many thanks, and congratz again !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143618": "Congrats to all medal teams and new Grandmasters,Masters,Experts.Thanks to Organizers and kaggle for such a good competition,it shows that kaggle competition is not just a game but also can be a useful machine learning project.\n\n# Team\n* At first we have three  teams individually, tomoyo and me, ethan and qyxs , wrb0312.We focus on feature engineering and optimization before wrb0312 joined us,I think we made many good features but neural network dominated this competition.wrb0312 did a great job even he use transformer for the first time.After wrb0312 joined us there are only ten days left,we focus on ensembling our models for inference,and also improved transformer very much.Our team members are from china and japan, it's very interesting to see we use chinese,japanese,english mixed-language to communicate.Greate job everyone!\n\n# Optimization \n* For GBM features, rather than using many dictionary to save features' data, we developed a nubma-based framework to speed up feature engineering process and online calculation. Firstly, the data are sorted by ['user_id', 'timestamp', 'content_id'] and split into different arrays. Then we created features in different array via self-designed rolling function or self-designed cumlative function. Actually, it provides us a very flexible way to create features and test it. In 10m data, the feature engineering process needs only 5 minutes to finish it.\n\n* Some examples are listed as below. \n```\nfrom tqdm import tqdm\nfrom numba import jit,njit\nfrom joblib import Parallel, delayed\nfrom tqdm import tqdm\nimport gc\nfrom multiprocessing import Process, Manager,Pool\nfrom functools import partial\nfrom numba import prange\nimport numpy as np\nimport pandas as pd\nfrom numba import types\nfrom numba.typed import Dict\nimport functools, time\nfrom numba.typed import List\n\n\ndef timeit(f):\n    def wrap(*args, **kwargs):\n        time1 = time.time()\n        ret = f(*args, **kwargs)\n        time2 = time.time()\n        print('{:s} function took {:.3f} s'.format(f.__name__, np.round(time2-time1, 2)))\n\n        return ret\n    return wrap\n\n\ndef rolling_feat_group(train, col_used):\n    a = train[col_used].values\n    ind = np.lexsort((a[:,2],a[:,1],a[:,0]))\n    a = a[ind]\n    g = np.split(a, np.unique(a[:, 0], return_index=True)[1][1:])\n    return g, ind, col_used\n\n@jit(nopython = True, fastmath = True)\ndef rolling_cal(arr, step, window = 5, shift_ = 1):\n    m = 2\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    beg = window\n    for i in step: \n        tmp = arr_[beg-window:beg]\n        ret[beg - window:(beg - window + i), 0] = np.nanmean(tmp)\n        ret[beg - window:(beg - window + i), 1] = np.nansum(tmp)\n        beg += i\n    return ret\n\n\n@jit(nopython = True, fastmath = True)\ndef rolling_time_cal(arr, window = 5, shift_ = 1):\n    m = 1\n    arr_ = np.concatenate((np.full((window, ), np.nan), arr))\n    ret = np.zeros((arr.shape[0], m))\n    for i in range(0,arr.shape[0], 1): \n        tmp = arr_[i:i+window+1]\n        ret[i, 0] = np.nanmean(tmp)\n    return ret\n\ndef rolling_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    step = np.unique(tmp_g[:, 1], return_counts=True)[1]\n    for window_size in shift_period:\n        tmp = rolling_cal(tmp_g[:, 2], step, window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\ndef rolling_time_cal_wrap(tmp_g, shift_period):\n    m = 2\n    tmp_res = []\n    for window_size in shift_period:\n        tmp = rolling_time_cal(tmp_g[:, 2], window_size)\n        tmp_res.append(tmp)\n    tmp_res = np.concatenate(tmp_res, axis = 1)\n    return tmp_res\n\n\ndef rolling_feat_cal(tmp_g, name_dict, global_period):\n    answer_idx = name_dict.index('answered_correctly')\n    prior_idx = name_dict.index('prior_question_elapsed_time')\n    item_mean_idx = name_dict.index('item_mean')\n    task_set_idx = name_dict.index('task_set_distance')\n    tmp_res1 = rolling_cal_wrap(tmp_g[:,[0,1, answer_idx]], global_period)\n    tmp_res2 = rolling_time_cal_wrap(tmp_g[:,[0,1, prior_idx]], global_period)\n    tmp_res3 = rolling_time_cal_wrap(tmp_g[:,[0,1, item_mean_idx]], global_period)\n    tmp_res4 = rolling_time_cal_wrap(tmp_g[:,[0,1, task_set_idx]], global_period)\n    tmp_res = np.concatenate([tmp_res1, tmp_res2, tmp_res3, tmp_res4], axis = 1)\n    return tmp_res\n```\n\n* If anyone interested in how to create features via numba-framework, Tomoyo publiced his full GBM pipeline in github([https://github.com/ZiwenYeee/Riiid-numba-framework](url))\n\n# Catboost(LB 0.807)\n### summary\n* We created 183 features for final catboost model,including some original features,global statistics(item base),cumulative and rolling statistics(user base),tfidf-svd(base on question's user list),word2vec(base on user's question list, wrong and correct tag list ),timedelta from many perspective,last same part groups features.\n\n### gbm benchmark\n* We compared lightgbm ,xgboost,catboost,catboost is the best for the training and inference speed,and memory consuming.When train the full data,lightgbm need over 100 hours with my AMD Ryzen ThreadRipper 3970X,xgboost always have out of memory error even using dask with 4 RTX 3090.\n\n### strong features and interesting finding by qyxs\n* 1.  the history difficulty statistics features of user who had correct/wrong answers, boost almost 0.003\n```\ntmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n```\n\n* 2.  focus on the records about the current part of user connect with last same part, generate the features include answer correct ratio, time diff, frequency etc, boost almost 0.002\n```\nfor_question_df['rank_part'] = for_question_df.groupby(['user_id', 'part'])['timestamp'].rank(method='first')\nfor_question_df['rank_user'] = for_question_df.groupby(['user_id'])['timestamp'].rank(method='first')\nfor_question_df['rank_diff'] = for_question_df['rank_user'] - for_question_df['rank_part']\nfor_question_df['part_times'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense')\nfor_question_df['rank_diff'] = for_question_df.groupby(['user_id', 'part'])['rank_diff'].rank(method='dense', ascending=False)\n\nlast_part = for_question_df[for_question_df['rank_diff']==1]\npart_times = for_question_df.groupby(['user_id', 'part'])['part_times'].agg([['part_times', 'max']]).reset_index()\n\nlast_part_df = last_part.groupby(['user_id', 'part'])['answered_correctly'].agg([['last_continue_part_ratio', 'mean'], ['last_continue_part_cnt', 'count']]).reset_index()\nlast_part_time = last_part.groupby(['user_id', 'part'])['timestamp'].agg([['last_continue_part_time_start', 'min'], ['last_continue_part_time_end', 'max']]).reset_index()\nlast_part_df = last_part_df.merge(last_part_time, on=['user_id', 'part'], how='left')\nlast_part_df = last_part_df.merge(part_times, on=['user_id', 'part'], how='left')\nlast_part_df['part_time_diff'] = last_part_df['last_continue_part_time_end'] - last_part_df['last_continue_part_time_start']\nlast_part_df['part_time_freq'] = last_part_df['last_continue_part_cnt']/last_part_df['part_time_diff']\n\nfor_train = for_train.merge(last_part_df, on=['user_id', 'part'], how='left')\nfor_train['last_continue_part_time_start'] = for_train['timestamp'] - for_train['last_continue_part_time_start']\nfor_train['last_continue_part_time_end'] = for_train['timestamp'] - for_train['last_continue_part_time_end']\n```\n\n* 3.  the answer correctly ratio of each question under differenct user abilititys (split for 11 bins), boost almost 0.001\n```\nfor_question_df['user_ability'] = for_question_df.groupby('user_id')['answered_correctly'].transform('mean').round(1)\ntmp_df = for_question_df.pivot_table(index='content_id', columns='user_ability', values='answered_correctly', aggfunc='mean').reset_index()\ntmp_df.columns = ['content_id'] + [f'c_mean_{i}_ratio' for i in range(11)]\nfor_train = for_train.merge(tmp_df, on='content_id', how='left')\n```\n\n* Some interseting points:\n * 1.they would watch lecture after users had wrong answers, so we could generated some features from this. LB is not improved caused by the lectures info in next group maybe.\n *  2.the content_id such as 0-195， 7851-7984 etc, then are all same in one part and continuous with each other，we could build a new bundle to generate features\n\n### strong features and interesting finding by ethan\n* 1.  user's behavior in last 1,5,...,60 minutes, 0.001 boost\n```\nfor w in [1, 5, 10, 15, 30, 45, 60]:\n    print(w)\n    tmp = q_logs[q_logs['timestamp']>=(q_logs['end_time']-w*60*1000)].copy()\n    group_df = tmp.groupby(['user_id'])['content_id'].agg([['user_content_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['part'].agg([['user_part_nunique_in_last{}mins'.format(w), 'nunique']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n    group_df = tmp.groupby(['user_id'])['answered_correctly'].agg([['user_correct_raito_in_last{}mins'.format(w), 'mean']]).reset_index()\n    train = train.merge(group_df, on=['user_id'], how='left')\n ```\n \n* 2. \"users' ablility\" statistics in each question, seperately by \"answered_correctly\"(0/1), 0.002 boost\n```\ncc = q_logs.groupby(['user_id'])['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ngg = q_logs[['user_id', 'content_id', 'answered_correctly']].merge(cc, on=['user_id'], how='left')\n\ngroup_df1 = gg[gg['answered_correctly']==1].groupby(['content_id'])['corr_ratio'].agg([['question_correct_user_ablility_min', 'min'], \n                                                                                       ['question_correct_user_ablility_max', 'max'], \n                                                                                       ['question_correct_user_ablility_mean', 'mean'], \n                                                                                       ['question_correct_user_ablility_skew', 'skew'],\n                                                                                       ['question_correct_user_ablility_med', 'median'],\n                                                                                       ['question_correct_user_ablility_std', 'std']]).reset_index()\ngroup_df2 = gg[gg['answered_correctly']==0].groupby(['content_id'])['corr_ratio'].agg([['question_wrong_user_ablility_min','min'], \n                                                                                       ['question_wrong_user_ablility_max','max'],\n                                                                                       ['question_wrong_user_ablility_mean','mean'],\n                                                                                       ['question_wrong_user_ablility_skew','skew'],\n                                                                                       ['question_wrong_user_ablility_med','median'],\n                                                                                       ['question_wrong_user_ablility_std','std']]).reset_index()\n```\n                        \n* 3. \"lagtime\" statistics in each question, seperately by \"answered_correctly\"(0/1), means the distribution of users' preprare time for answering this question correctly, about 0.001 boost\n```\nuser_task_timestamp = q_logs[['user_id', 'task_container_id', 'timestamp']].drop_duplicates()\nuser_task_timestamp['lag_time'] = user_task_timestamp['timestamp'] - user_task_timestamp.groupby(['user_id'])['timestamp'].shift(1)\ntmp = q_logs[['user_id', 'task_container_id', 'content_id', 'answered_correctly']].merge(user_task_timestamp.drop(['timestamp'], axis=1), on=['user_id', 'task_container_id'], how='left')\ngroup_df = tmp[tmp['answered_correctly']==1].groupby(['content_id'])['lag_time'].agg([['c_lag_time_mean', 'mean'],\n                                                                                     ['c_lag_time_std', 'std'],\n                                                                                     ['c_lag_time_max', 'max'],\n                                                                                     ['c_lag_time_min', 'min'],\n                                                                                     ['c_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n\ngroup_df = tmp[tmp['answered_correctly']==0].groupby(['content_id'])['lag_time'].agg([['w_lag_time_mean', 'mean'],\n                                                                                       ['w_lag_time_std', 'std'],\n                                                                                       ['w_lag_time_max', 'max'],\n                                                                                       ['w_lag_time_min', 'min'],\n                                                                                       ['w_lag_time_median', 'median']]).reset_index()\ntrain = train.merge(group_df, on=['content_id'], how='left')\n```\n\n### feature list\n```\n['content_id',\n 'prior_question_elapsed_time',\n 'prior_question_had_explanation',\n 'correct_answer',\n 'user_count',\n 'user_sum',\n 'user_mean',\n 'item_count',\n 'item_sum',\n 'item_mean',\n 'answer_ratio_0',\n 'answer_ratio_1',\n 'answer_ratio_2',\n 'bundle_id',\n 'part',\n 'le_tag',\n 'question_correct_user_ablility_mean',\n 'question_correct_user_ablility_median',\n 'question_wrong_user_ablility_mean',\n 'question_wrong_user_ablility_median',\n 'word2vec_0',\n 'word2vec_1',\n 'word2vec_2',\n 'word2vec_3',\n 'word2vec_4',\n 'svd_0',\n 'svd_1',\n 'svd_2',\n 'svd_3',\n 'svd_4',\n 'tags_w2v_correct_mean_0',\n 'tags_w2v_wrong_mean_0',\n 'tags_w2v_correct_mean_1',\n 'tags_w2v_wrong_mean_1',\n 'tags_w2v_correct_mean_2',\n 'tags_w2v_wrong_mean_2',\n 'tags_w2v_correct_mean_3',\n 'tags_w2v_wrong_mean_3',\n 'tags_w2v_correct_mean_4',\n 'tags_w2v_wrong_mean_4',\n 'real_time_wrong_mean',\n 'real_time_wrong_median',\n 'real_time_correct_mean',\n 'real_time_correct_median',\n 'task_set_distance_wrong_mean',\n 'task_set_distance_wrong_median',\n 'task_set_distance_correct_mean',\n 'task_set_distance_correct_median',\n 'mean_0_ratio',\n 'mean_1_ratio',\n 'mean_3_ratio',\n 'mean_4_ratio',\n 'mean_5_ratio',\n 'mean_6_ratio',\n 'mean_7_ratio',\n 'mean_8_ratio',\n 'mean_9_ratio',\n 'mean_10_ratio',\n 'user_d1',\n 'user_d2',\n 'task_set_distance',\n 'user_diff_mean',\n 'user_diff_std',\n 'user_diff_min',\n 'user_diff_max',\n 'task_set_item_mean',\n 'task_set_item_min',\n 'task_set_item_max',\n 'task_set_distance2',\n 'task_distance_shift',\n 'task_set_distance_diff',\n 'task_distance_diff_shift',\n 'container_mean_1',\n 'container_mean_5',\n 'container_std_5',\n 'container_mean_10',\n 'container_std_10',\n 'container_mean_20',\n 'container_std_20',\n 'container_mean_30',\n 'container_std_30',\n 'container_mean_40',\n 'container_std_40',\n 'prior_question_elapsed_time_mean_1',\n 'prior_question_elapsed_time_mean_5',\n 'prior_question_elapsed_time_mean_10',\n 'prior_question_elapsed_time_mean_20',\n 'prior_question_elapsed_time_mean_30',\n 'prior_question_elapsed_time_mean_40',\n 'item_mean_mean_30',\n 'item_mean_mean_40',\n 'task_set_distance_mean_1',\n 'task_set_distance_mean_5',\n 'task_set_distance_mean_10',\n 'task_set_distance_mean_20',\n 'task_set_distance_mean_30',\n 'begin_time_diff',\n 'end_time_diff',\n 'part_time_diff_mean',\n 'part_session_mean',\n 'part_session_sum',\n 'part_session_count',\n 'full_group0_item_mean_mean',\n 'full_group0_item_mean_median',\n 'full_group0_task_set_distance_median',\n 'full_group0_timestamp_mean',\n 'full_group0_timestamp_median',\n 'full_group1_item_mean_mean',\n 'full_group1_item_mean_median',\n 'full_group1_task_set_distance_median',\n 'full_group1_timestamp_median',\n 'part_sum',\n 'part_count',\n 'part_mean',\n 'part_sum_global_ratio',\n 'part_sum_1',\n 'part_sum_5',\n 'part_mean_5',\n 'part_sum_10',\n 'part_mean_10',\n 'cum_answer0_mean_item_mean',\n 'cum_answer0_median_item_mean',\n 'cum_answer0_median_task_set_distance',\n 'cum_answer1_mean_item_mean',\n 'cum_answer1_median_item_mean',\n 'cum_answer1_mean_task_set_distance',\n 'cum_answer1_median_task_set_distance',\n 'cum_answer0_time_diff',\n 'cum_answer1_time_diff',\n 'global_task_set_shift1',\n 'global_task_set_shift2',\n 'global_task_set_shift4',\n 'global_task_set_shift5',\n 'cum_answer0_mean_wrong_time_diff',\n 'cum_answer0_median_wrong_time_diff',\n 'cum_answer1_mean_right_time_diff',\n 'content_correct_mean',\n 'content_correct_sum',\n 'content_correct_count',\n 'hard_answer0_time',\n 'hard_answer1_time',\n 'full_bundle_item_mean_mean',\n 'full_bundle_item_mean_median',\n 'full_bundle_task_set_distance_mean',\n 'full_bundle_task_set_distance_median',\n 'full_bundle_timestamp_mean',\n 'full_bundle_timestamp_median',\n 'bundle_sum',\n 'bundle_mean',\n 'bundle_count',\n 'user_trend_mean',\n 'user_trend_median',\n 'user_trend_roll_user_ans_sum',\n 'user_trend_roll_user_ans_mean',\n 'user_trend_roll_user_ans_count',\n 'user_trend_roll_item_ans_mean',\n 'user_trend_roll_item_ans_count',\n 'div_ratio1',\n 'div_ratio2',\n 'div_ratio3',\n 'new_Feat0',\n 'new_Feat1',\n 'new_Feat2',\n 'new_Feat3',\n 'part_time_wrong_div',\n 'part_time_right_div',\n 'diff_lag_median_div',\n 'diff_item_median_div',\n 'diff_time_median_div',\n 'diff_item_mean_div',\n 'diff_task_set_mean_div',\n 'diff_timestamp_mean_div',\n 'last_20_frequent_answer',\n 'last_20_frequent_answer_count',\n 'last_20_frequent_answer_mean',\n 'last_20_frequent_answer_sum',\n 'last_user_same_answer_tf',\n 'last_item_same_answer_tf',\n 'last_right_time_diff',\n 'last_wrong_time_diff',\n 'last_5_part_time_div',\n 'last_10_part_time_div',\n 'last_20_part_time_div']\n```\n# Transformer(LB 0.808)\n* Used questions only. Lectures did't improve our validation score.\n*  In training window size 800 (the bigger,the better), in inference window size 300(the bigger,the better,but time consuming).\n* Optimizer : Adam  with lr = 8e-4, beta1 = 0.9, beta2 = 0.999 with warmup steps to 4000.\n  * Without warmup steps, didn't converge.\n* Number of layers = 4, dimension of the model = 256, dimension of the FFN = 2048.\n* Batch size=80\n  * With smaller size,  didn't converge.\n* Dropout = 0\n* Postion encoding : Axial Positional Embedding\n  * https://arxiv.org/abs/1912.12180\n  * https://github.com/lucidrains/axial-positional-embedding\n  * This encoding improved score significantly.\n* Augmentation : Mixup-Transformer\n  * https://arxiv.org/abs/2010.02394\n\n* Inputs\n  * question_id\n  * part\n  * prior_question_elapsed_time / 1000\n  * lagtime\n      * log1p((timestamp_t - timestamp_(t-1)) /1000 / 60)\n      * This feature improved the score significantly\n  * answered_correctly\n  * GBT feats (imortance top N)\n* model image\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F715257%2F6743aa3ef1403f4084dd4fa3d8372702%2FScreen%20Shot%202021-01-08%20at%2011.34.11.png?generation=1610073332748004&alt=media)\n\n# Ensemble(LB 0.812)\n* Finally we use one catboost and two transformers for ensemble due to the inference time limitation.Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.",
    "1143632": "I've never used nubma. Really inspiring. Thank you for sharing.",
    "1143636": "Thanks for your work and effort, bro. Looking forward to future cooperation!\n\nIf anyone interested in my numba-based framework for GBM, leave a message below. I'm considering write a detailed explanation about it.",
    "1143637": "I have a strong interest on it :+1:",
    "1143647": "Congrats to all winners! Thanks for my teamates @senkin13 @guziye @juzqyxs @wrb0312 ! Great teamwork!",
    "1143653": "Very interesting!",
    "1143665": "senkin13 question for you - looks like catboost had a pretty strong performance based on features. Were you and your team not able to get a similar score with lightgbm? Also, does catboost have a similar type of feature importance plot? Just curious to know which of these features gave the greatest gains",
    "1143669": "Interested. Thanks.",
    "1143670": "I also interested in looking at it's src; Thanks a lot for sharing. I also used numba for computing the rolling FE's, the time dropped to <3 secs for 100M rows as compared to more than 30 mins of Pandas! So very excited to explore numba!\n\n\n```\n# <3 secs for 100M rows on my lapi\nfrom numba import njit # no python basically\n\nwindow_width = 5\ndummy_window_5 = [np.nan]*4 # window_width-1\n\n@njit\ndef running_window_stats_window_width_5(vector):\n    # vector must be pre-padded with a zero\n    cumsum_vec = np.cumsum(vector, )\n    return (cumsum_vec[window_width:] - cumsum_vec[:-window_width]) / window_width\n```",
    "1143672": "the reason we use catboost is just speed,I think lightgbm can get a similar score and feature gains",
    "1143694": "rdizzl3 Instead of \"lagtime\", our good features focus on \"user's ablilty\" and \"question's difficulty\". User's correctness in history represents \"user's ablility\" and Question's correctness(global) represents \"question's difficulty\". For example, \"quesions' difficulty\" statics(mean, max,min...) in each user's history and \"user's ablilty\" statics in each quesion. Especially create these features seperately by answered_correctly(0 or 1) will get great gains.",
    "1143700": "Great feature engineering work. Congrats on the results @senkin13 and team",
    "1143714": "Congratulations!",
    "1143789": "Congratulations. Great models and strong finish! Your transformer is interesting. Did it do better than SAINT+? For reference, SAINT+ is pictured below! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8d4e260205a6a6599aefeee99bdf1b6a%2Fsaint2.png?generation=1610079431451477&alt=media)",
    "1143822": "Thanks!\nI didn't try your picture's model architecture.\nBut maybe we can get same score by SAINT+ architecture, with shorter inference time.\nIn this competition I tried transformer first time, so I think it is possible to optimize the model structure more (it was hard for me).",
    "1144386": "Congratz to all your team for the nice finish !\n\n> Unfortunately our inference notebook has bug although we improved our model about 0.0006 at last day.\n\nThat's heartbreaking, however do let people know if you manage to fix the bug & improve your score",
    "1144443": "thanks for your care,we found the bug is from changing some features from float64 to float32 by mistake in inference notebook,if it is normal we still can not get gold medal,no regret!",
    "1144488": "Good to hear, thanks !",
    "1144674": "congratulations, a great innovative job!",
    "1148158": "Congratulations for the great solution and nice finish! I have two questions:\n\n-1 In the example code, It seems like the corr_ratio are calculated by each user_id with forward looking. Model still had stable improvement between CV and LB?\n\n```\ntmp_df = for_question_df.groupby('content_id')['answered_correctly'].agg([['corr_ratio', 'mean']]).reset_index()\ntmp_fe = for_question_df[for_question_df['answered_correctly']==0].merge(tmp_df, on='content_id').groupby('user_id')['corr_ratio'].agg(['min', 'max', 'mean', 'std']).reset_index()\nfor_train = for_train.merge(tmp_fe, on='user_id', how='left')\n```\n\n-2 How did you update features with numba functions in the test api? (didn't find numba functions for test updating in the gihub)",
    "1148173": "for question 2, We only update saving data and calculate with numba function rather than updating features. I loaded up one of our submit notebook and our test script in my github. https://github.com/ZiwenYeee/Riiid-numba-framework",
    "1149081": "Congratulation for your amazing work ! Thanks for sharing your great ideas so much.\n\nCan you give additional details about w2v & svd features ?\n\n1) While `tags_w2v_correct_mean_0 ... tags_w2v_wrong_mean_4` are self-explanatory and clear, I don't understand what **`word2vec_0 ... word2vec_4`** stand for. You said \"based on user's question list\" : is this feat the w2v of the current q_id ? Or the mean of w2v of each q_id in user's question list, just like tags ? \n2) What about **`svd_0 ... svd_4`** ?",
    "1150513": "yeah,the improvement was stable.",
    "1150527": "word2vec features are the w2v of current q_id, svd features get from the tfidf of users' history q_id.",
    "1150752": "Makes sense. Many thanks, and congratz again !"
  },
  "source": "meta"
}