{
  "id": 194335,
  "title": "LightGBM takes too long time",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194335",
  "author_name": "",
  "post_date": "2020-11-01T08:58:31.695567300Z",
  "votes": 13,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Most of public notebooks use target encoding features.<br>\nIt seems to be working, but as is said <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437\" target=\"_blank\">here</a>, if you average the target using all the data, it will leak information from the future into your average.</p>\n<p>Then, I made this feature:</p>\n<p>it takes average of the past target values for <code>user_id</code>.</p>\n<pre><code>@pickle_cache\ndef user_id_target_encoded(train_df: pd.DataFrame) -&gt; pd.Series:\n    # target encode over timestamp\n    train_df = train_df[train_df['content_type_id'] == 0].reset_index(drop=True)\n    df = train_df[['user_id', 'timestamp', 'answered_correctly']]\n    encoded = (\n        df.sort_values('timestamp')\n        .groupby('user_id')['answered_correctly']\n        .expanding().mean()\n        .groupby('user_id')\n        .shift(1)\n    )\n    encoded = pd.Series(encoded.values,\n                        index=encoded.index.get_level_values(1),\n                        name='user_id_target_encoded').sort_index()\n\n    # deal with same user_id, same timestamp\n    df = pd.concat([df, encoded], axis=1, sort=False)\n    encoded = df.groupby(['user_id', 'timestamp'])['user_id_target_encoded'].transform('nth', n=0)\n\n    return encoded\n</code></pre>\n<p>However, if I use this feature, LightGBM takes more than 10 hours (&gt; 1800 rounds) for training, and fails submission (&gt; 9 hours for prediction).</p>\n<p>How do you deal with this? Any idea?</p>\n<p>My settings:</p>\n<pre><code>features = [\n    'timestamp',\n    'user_id_target_encoded',\n    'user_id_count_encoded',\n    'content_id_target_encoded',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation'\n]\n\nlgb_params = {\n    'boosting_type': 'gbdt',\n    'tree_learner': 'data',\n    'objective': 'binary',\n    'metric': 'auc',\n    'learning_rate': 0.1,\n    'num_leaves': 80,\n    'max_depth': 9,\n    'max_bin': 511,\n    'bagging_fraction': 0.8,\n    'bagging_freq': 1,\n    'bin_construct_sample_cnt': 1000000,\n    'seed': 1029,\n    'verbose': 0\n}\nnum_boost_round = 5000\nearly_stopping_rounds = 30\n</code></pre>\n<p>validation split:</p>\n<pre><code>def user_lastn_split(train_df: pd.DataFrame,\n                     max_n: int = 15,\n                     random_state: int = 0) -&gt; Tuple[np.ndarray, np.ndarray]:\n    set_seed(random_state)\n    tqdm.pandas(desc='create validation split')\n\n    def _lastn_pickup(x: pd.Series, n: int = 10) -&gt; pd.Series:\n        last_n_max = min(max_n, len(x))\n        n_pickup = np.random.choice(range(last_n_max))\n        return x.tail(n_pickup)\n\n    df = train_df.reset_index(drop=True)\n    valid_df = df.groupby('user_id').progress_apply(_lastn_pickup, n=max_n)\n    valid_idx = valid_df.index.get_level_values(1).values\n    train_idx = np.delete(range(len(df)), valid_idx)\n\n    return train_idx, valid_idx\n</code></pre>",
  "messages": [
    {
      "id": "1066056",
      "postDate": "11/01/2020 08:58:31",
      "content": "<p>Hi all,</p>\n<p>Most of public notebooks use target encoding features.<br>\nIt seems to be working, but as is said <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437\" target=\"_blank\">here</a>, if you average the target using all the data, it will leak information from the future into your average.</p>\n<p>Then, I made this feature:</p>\n<p>it takes average of the past target values for <code>user_id</code>.</p>\n<pre><code>@pickle_cache\ndef user_id_target_encoded(train_df: pd.DataFrame) -&gt; pd.Series:\n    # target encode over timestamp\n    train_df = train_df[train_df['content_type_id'] == 0].reset_index(drop=True)\n    df = train_df[['user_id', 'timestamp', 'answered_correctly']]\n    encoded = (\n        df.sort_values('timestamp')\n        .groupby('user_id')['answered_correctly']\n        .expanding().mean()\n        .groupby('user_id')\n        .shift(1)\n    )\n    encoded = pd.Series(encoded.values,\n                        index=encoded.index.get_level_values(1),\n                        name='user_id_target_encoded').sort_index()\n\n    # deal with same user_id, same timestamp\n    df = pd.concat([df, encoded], axis=1, sort=False)\n    encoded = df.groupby(['user_id', 'timestamp'])['user_id_target_encoded'].transform('nth', n=0)\n\n    return encoded\n</code></pre>\n<p>However, if I use this feature, LightGBM takes more than 10 hours (&gt; 1800 rounds) for training, and fails submission (&gt; 9 hours for prediction).</p>\n<p>How do you deal with this? Any idea?</p>\n<p>My settings:</p>\n<pre><code>features = [\n    'timestamp',\n    'user_id_target_encoded',\n    'user_id_count_encoded',\n    'content_id_target_encoded',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation'\n]\n\nlgb_params = {\n    'boosting_type': 'gbdt',\n    'tree_learner': 'data',\n    'objective': 'binary',\n    'metric': 'auc',\n    'learning_rate': 0.1,\n    'num_leaves': 80,\n    'max_depth': 9,\n    'max_bin': 511,\n    'bagging_fraction': 0.8,\n    'bagging_freq': 1,\n    'bin_construct_sample_cnt': 1000000,\n    'seed': 1029,\n    'verbose': 0\n}\nnum_boost_round = 5000\nearly_stopping_rounds = 30\n</code></pre>\n<p>validation split:</p>\n<pre><code>def user_lastn_split(train_df: pd.DataFrame,\n                     max_n: int = 15,\n                     random_state: int = 0) -&gt; Tuple[np.ndarray, np.ndarray]:\n    set_seed(random_state)\n    tqdm.pandas(desc='create validation split')\n\n    def _lastn_pickup(x: pd.Series, n: int = 10) -&gt; pd.Series:\n        last_n_max = min(max_n, len(x))\n        n_pickup = np.random.choice(range(last_n_max))\n        return x.tail(n_pickup)\n\n    df = train_df.reset_index(drop=True)\n    valid_df = df.groupby('user_id').progress_apply(_lastn_pickup, n=max_n)\n    valid_idx = valid_df.index.get_level_values(1).values\n    train_idx = np.delete(range(len(df)), valid_idx)\n\n    return train_idx, valid_idx\n</code></pre>",
      "rawMarkdown": "Hi all,\n\nMost of public notebooks use target encoding features.\nIt seems to be working, but as is said [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437), if you average the target using all the data, it will leak information from the future into your average.\n\nThen, I made this feature:\n\nit takes average of the past target values for `user_id`.\n\n```\n@pickle_cache\ndef user_id_target_encoded(train_df: pd.DataFrame) -> pd.Series:\n    # target encode over timestamp\n    train_df = train_df[train_df['content_type_id'] == 0].reset_index(drop=True)\n    df = train_df[['user_id', 'timestamp', 'answered_correctly']]\n    encoded = (\n        df.sort_values('timestamp')\n        .groupby('user_id')['answered_correctly']\n        .expanding().mean()\n        .groupby('user_id')\n        .shift(1)\n    )\n    encoded = pd.Series(encoded.values,\n                        index=encoded.index.get_level_values(1),\n                        name='user_id_target_encoded').sort_index()\n \n    # deal with same user_id, same timestamp\n    df = pd.concat([df, encoded], axis=1, sort=False)\n    encoded = df.groupby(['user_id', 'timestamp'])['user_id_target_encoded'].transform('nth', n=0)\n\n    return encoded\n```\n\nHowever, if I use this feature, LightGBM takes more than 10 hours (> 1800 rounds) for training, and fails submission (> 9 hours for prediction).\n\nHow do you deal with this? Any idea?\n\nMy settings:\n\n```\nfeatures = [\n    'timestamp',\n    'user_id_target_encoded',\n    'user_id_count_encoded',\n    'content_id_target_encoded',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation'\n]\n\nlgb_params = {\n    'boosting_type': 'gbdt',\n    'tree_learner': 'data',\n    'objective': 'binary',\n    'metric': 'auc',\n    'learning_rate': 0.1,\n    'num_leaves': 80,\n    'max_depth': 9,\n    'max_bin': 511,\n    'bagging_fraction': 0.8,\n    'bagging_freq': 1,\n    'bin_construct_sample_cnt': 1000000,\n    'seed': 1029,\n    'verbose': 0\n}\nnum_boost_round = 5000\nearly_stopping_rounds = 30\n```\n\nvalidation split:\n\n```\ndef user_lastn_split(train_df: pd.DataFrame,\n                     max_n: int = 15,\n                     random_state: int = 0) -> Tuple[np.ndarray, np.ndarray]:\n    set_seed(random_state)\n    tqdm.pandas(desc='create validation split')\n\n    def _lastn_pickup(x: pd.Series, n: int = 10) -> pd.Series:\n        last_n_max = min(max_n, len(x))\n        n_pickup = np.random.choice(range(last_n_max))\n        return x.tail(n_pickup)\n\n    df = train_df.reset_index(drop=True)\n    valid_df = df.groupby('user_id').progress_apply(_lastn_pickup, n=max_n)\n    valid_idx = valid_df.index.get_level_values(1).values\n    train_idx = np.delete(range(len(df)), valid_idx)\n\n    return train_idx, valid_idx\n```",
      "votes": null
    },
    {
      "id": "1066072",
      "postDate": "11/01/2020 09:29:40",
      "content": "<p>For me it was both time and RAM. My RAM initially exploded to 60GB+. Have you tried using a lower max_bin? Plus, a larger value for <code>bin_construct_sample_cnt</code> will also increase the data loading time. Are you doing a k fold Cv to do target-encoding? I am suing the below attached and it runs quite fast. (NB Not using any cat features explicitly, it's auto by default)</p>\n<p>Glad to see you here Sakami!</p>\n<pre><code>    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': \"binary\",\n        'metric': metrics,\n        'learning_rate': 0.1,\n        'num_leaves': 2**6,\n        'max_depth': 8,\n        'colsample_bytree': 0.7,\n        'min_child_samples': 100,\n        'subsample': 0.7, # ideally  it can be sqrt(number of features)/(number of features)\n        'num_threads': 8, # real_cores\n        'seed': 2020,\n        'first_metric_only': True,\n        'use_two_round_loading': True,\n        'max_bin':128,\n        'verbose': -1\n    }\n</code></pre>",
      "rawMarkdown": "For me it was both time and RAM. My RAM initially exploded to 60GB+. Have you tried using a lower max_bin? Plus, a larger value for `bin_construct_sample_cnt` will also increase the data loading time. Are you doing a k fold Cv to do target-encoding? I am suing the below attached and it runs quite fast. (NB Not using any cat features explicitly, it's auto by default)\n\nGlad to see you here Sakami!\n\n```\n    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': \"binary\",\n        'metric': metrics,\n        'learning_rate': 0.1,\n        'num_leaves': 2**6,\n        'max_depth': 8,\n        'colsample_bytree': 0.7,\n        'min_child_samples': 100,\n        'subsample': 0.7, # ideally  it can be sqrt(number of features)/(number of features)\n        'num_threads': 8, # real_cores\n        'seed': 2020,\n        'first_metric_only': True,\n        'use_two_round_loading': True,\n        'max_bin':128,\n        'verbose': -1\n    }\n```",
      "votes": null
    },
    {
      "id": "1066333",
      "postDate": "11/01/2020 16:19:24",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> for your kind comment.</p>\n<blockquote>\n  <p>Have you tried using a lower max_bin?</p>\n</blockquote>\n<p>Yes, I tried <code>max_bin = 255</code> but it didn't change drastically… :(</p>\n<blockquote>\n  <p>Are you doing a k fold Cv to do target-encoding?</p>\n</blockquote>\n<p>Yes, I'm using fold target encoding in <code>content_id_target_encoded</code>.</p>\n<blockquote>\n  <p>I am suing the below attached and it runs quite fast.</p>\n</blockquote>\n<p>Thank you, I'll tune my parameters with reference to yours.</p>",
      "rawMarkdown": "Thank you @adityaecdrid for your kind comment.\n\n> Have you tried using a lower max_bin?\n\nYes, I tried `max_bin = 255` but it didn't change drastically... :(\n\n> Are you doing a k fold Cv to do target-encoding?\n\nYes, I'm using fold target encoding in `content_id_target_encoded`.\n\n> I am suing the below attached and it runs quite fast.\n\nThank you, I'll tune my parameters with reference to yours.",
      "votes": null
    },
    {
      "id": "1066815",
      "postDate": "11/02/2020 04:17:15",
      "content": "<p>Wow, doing a k-fold CV when we have 100M rows is 🙏. And these features are definitely gonna help out as well;</p>",
      "rawMarkdown": "Wow, doing a k-fold CV when we have 100M rows is 🙏. And these features are definitely gonna help out as well;",
      "votes": null
    },
    {
      "id": "1069224",
      "postDate": "11/04/2020 08:16:40",
      "content": "<p>what about train the model in your own machine，and upload the model file as private dataset<br>\nI dont know the auc improvement of the rounds&gt;200,1000… , I guess it did not get much gain(may not need to do such many rounds?). my case below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5900557%2F2c8028c52cba45d5294ad5b0e0400153%2F_20201104160629.png?generation=1604477778718612&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "what about train the model in your own machine，and upload the model file as private dataset\nI dont know the auc improvement of the rounds>200,1000... , I guess it did not get much gain(may not need to do such many rounds?). my case below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5900557%2F2c8028c52cba45d5294ad5b0e0400153%2F_20201104160629.png?generation=1604477778718612&alt=media)",
      "votes": null
    },
    {
      "id": "1071972",
      "postDate": "11/07/2020 17:03:23",
      "content": "<p>Thank you for your advice!<br>\nIn my case, training a model in local environment and predicting with 1000 rounds also got \"submission scoring error\". 😓</p>",
      "rawMarkdown": "Thank you for your advice!\nIn my case, training a model in local environment and predicting with 1000 rounds also got \"submission scoring error\". 😓",
      "votes": null
    },
    {
      "id": "1071974",
      "postDate": "11/07/2020 17:06:58",
      "content": "<p>That's sad to hear but i am pretty sure that you will resolve the problem pretty soon :)</p>",
      "rawMarkdown": "That's sad to hear but i am pretty sure that you will resolve the problem pretty soon :)",
      "votes": null
    },
    {
      "id": "1071985",
      "postDate": "11/07/2020 17:23:43",
      "content": "<p>your code may cost too much time on merge features for provided samples(for lgbm model with 1000 rounds cost only about 0.03s to make prediction for provided 120 samples but more than 3s to merge features with 4 merge operations).check whole time the loop you make to make prediction for the provided 120 samples,It should be no more than 5s.</p>",
      "rawMarkdown": "your code may cost too much time on merge features for provided samples(for lgbm model with 1000 rounds cost only about 0.03s to make prediction for provided 120 samples but more than 3s to merge features with 4 merge operations).check whole time the loop you make to make prediction for the provided 120 samples,It should be no more than 5s.",
      "votes": null
    },
    {
      "id": "1072252",
      "postDate": "11/08/2020 01:23:52",
      "content": "<p>Sakami, How's the speed for you now? Yesterday i tried on 100M rows on MBP and even after 8 hours, it's just 1000 rounds.</p>",
      "rawMarkdown": "Sakami, How's the speed for you now? Yesterday i tried on 100M rows on MBP and even after 8 hours, it's just 1000 rounds.",
      "votes": null
    },
    {
      "id": "1082058",
      "postDate": "11/17/2020 14:52:18",
      "content": "<p>Legend is legend after all 🙏; Any tips ? Ty!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F7c2abedd70610dd42751582e2ffdb517%2FScreenshot%202020-11-17%20at%208.20.53%20PM.png?generation=1605624715514499&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Legend is legend after all 🙏; Any tips ? Ty!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F7c2abedd70610dd42751582e2ffdb517%2FScreenshot%202020-11-17%20at%208.20.53%20PM.png?generation=1605624715514499&alt=media)",
      "votes": null
    },
    {
      "id": "1083868",
      "postDate": "11/19/2020 13:29:59",
      "content": "<p>Thanks!</p>\n<p>I completely changed my approach. Our best score so far (LB 0.780) is by a SAINT model.</p>",
      "rawMarkdown": "Thanks!\n\nI completely changed my approach. Our best score so far (LB 0.780) is by a SAINT model.",
      "votes": null
    },
    {
      "id": "1083877",
      "postDate": "11/19/2020 13:40:24",
      "content": "<p>Oh wow! Any tips you wanna share in the SAINT discussion thread :🙏</p>",
      "rawMarkdown": "Oh wow! Any tips you wanna share in the SAINT discussion thread :🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1066072,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "11/01/2020 09:29:40",
      "content": "<p>For me it was both time and RAM. My RAM initially exploded to 60GB+. Have you tried using a lower max_bin? Plus, a larger value for <code>bin_construct_sample_cnt</code> will also increase the data loading time. Are you doing a k fold Cv to do target-encoding? I am suing the below attached and it runs quite fast. (NB Not using any cat features explicitly, it's auto by default)</p>\n<p>Glad to see you here Sakami!</p>\n<pre><code>    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': \"binary\",\n        'metric': metrics,\n        'learning_rate': 0.1,\n        'num_leaves': 2**6,\n        'max_depth': 8,\n        'colsample_bytree': 0.7,\n        'min_child_samples': 100,\n        'subsample': 0.7, # ideally  it can be sqrt(number of features)/(number of features)\n        'num_threads': 8, # real_cores\n        'seed': 2020,\n        'first_metric_only': True,\n        'use_two_round_loading': True,\n        'max_bin':128,\n        'verbose': -1\n    }\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1066333,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "11/01/2020 16:19:24",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> for your kind comment.</p>\n<blockquote>\n  <p>Have you tried using a lower max_bin?</p>\n</blockquote>\n<p>Yes, I tried <code>max_bin = 255</code> but it didn't change drastically… :(</p>\n<blockquote>\n  <p>Are you doing a k fold Cv to do target-encoding?</p>\n</blockquote>\n<p>Yes, I'm using fold target encoding in <code>content_id_target_encoded</code>.</p>\n<blockquote>\n  <p>I am suing the below attached and it runs quite fast.</p>\n</blockquote>\n<p>Thank you, I'll tune my parameters with reference to yours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066815,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/02/2020 04:17:15",
          "content": "<p>Wow, doing a k-fold CV when we have 100M rows is 🙏. And these features are definitely gonna help out as well;</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072252,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/08/2020 01:23:52",
          "content": "<p>Sakami, How's the speed for you now? Yesterday i tried on 100M rows on MBP and even after 8 hours, it's just 1000 rounds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1069224,
      "author_name": "huangtaogan",
      "author_url": "",
      "post_date": "11/04/2020 08:16:40",
      "content": "<p>what about train the model in your own machine，and upload the model file as private dataset<br>\nI dont know the auc improvement of the rounds&gt;200,1000… , I guess it did not get much gain(may not need to do such many rounds?). my case below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5900557%2F2c8028c52cba45d5294ad5b0e0400153%2F_20201104160629.png?generation=1604477778718612&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1071972,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "11/07/2020 17:03:23",
          "content": "<p>Thank you for your advice!<br>\nIn my case, training a model in local environment and predicting with 1000 rounds also got \"submission scoring error\". 😓</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1071974,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/07/2020 17:06:58",
          "content": "<p>That's sad to hear but i am pretty sure that you will resolve the problem pretty soon :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1071985,
          "author_name": "huangtaogan",
          "author_url": "",
          "post_date": "11/07/2020 17:23:43",
          "content": "<p>your code may cost too much time on merge features for provided samples(for lgbm model with 1000 rounds cost only about 0.03s to make prediction for provided 120 samples but more than 3s to merge features with 4 merge operations).check whole time the loop you make to make prediction for the provided 120 samples,It should be no more than 5s.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1082058,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "11/17/2020 14:52:18",
      "content": "<p>Legend is legend after all 🙏; Any tips ? Ty!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F7c2abedd70610dd42751582e2ffdb517%2FScreenshot%202020-11-17%20at%208.20.53%20PM.png?generation=1605624715514499&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1083868,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "11/19/2020 13:29:59",
          "content": "<p>Thanks!</p>\n<p>I completely changed my approach. Our best score so far (LB 0.780) is by a SAINT model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1083877,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/19/2020 13:40:24",
          "content": "<p>Oh wow! Any tips you wanna share in the SAINT discussion thread :🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1066056": "Hi all,\n\nMost of public notebooks use target encoding features.\nIt seems to be working, but as is said [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189437), if you average the target using all the data, it will leak information from the future into your average.\n\nThen, I made this feature:\n\nit takes average of the past target values for `user_id`.\n\n```\n@pickle_cache\ndef user_id_target_encoded(train_df: pd.DataFrame) -> pd.Series:\n    # target encode over timestamp\n    train_df = train_df[train_df['content_type_id'] == 0].reset_index(drop=True)\n    df = train_df[['user_id', 'timestamp', 'answered_correctly']]\n    encoded = (\n        df.sort_values('timestamp')\n        .groupby('user_id')['answered_correctly']\n        .expanding().mean()\n        .groupby('user_id')\n        .shift(1)\n    )\n    encoded = pd.Series(encoded.values,\n                        index=encoded.index.get_level_values(1),\n                        name='user_id_target_encoded').sort_index()\n \n    # deal with same user_id, same timestamp\n    df = pd.concat([df, encoded], axis=1, sort=False)\n    encoded = df.groupby(['user_id', 'timestamp'])['user_id_target_encoded'].transform('nth', n=0)\n\n    return encoded\n```\n\nHowever, if I use this feature, LightGBM takes more than 10 hours (> 1800 rounds) for training, and fails submission (> 9 hours for prediction).\n\nHow do you deal with this? Any idea?\n\nMy settings:\n\n```\nfeatures = [\n    'timestamp',\n    'user_id_target_encoded',\n    'user_id_count_encoded',\n    'content_id_target_encoded',\n    'prior_question_elapsed_time',\n    'prior_question_had_explanation'\n]\n\nlgb_params = {\n    'boosting_type': 'gbdt',\n    'tree_learner': 'data',\n    'objective': 'binary',\n    'metric': 'auc',\n    'learning_rate': 0.1,\n    'num_leaves': 80,\n    'max_depth': 9,\n    'max_bin': 511,\n    'bagging_fraction': 0.8,\n    'bagging_freq': 1,\n    'bin_construct_sample_cnt': 1000000,\n    'seed': 1029,\n    'verbose': 0\n}\nnum_boost_round = 5000\nearly_stopping_rounds = 30\n```\n\nvalidation split:\n\n```\ndef user_lastn_split(train_df: pd.DataFrame,\n                     max_n: int = 15,\n                     random_state: int = 0) -> Tuple[np.ndarray, np.ndarray]:\n    set_seed(random_state)\n    tqdm.pandas(desc='create validation split')\n\n    def _lastn_pickup(x: pd.Series, n: int = 10) -> pd.Series:\n        last_n_max = min(max_n, len(x))\n        n_pickup = np.random.choice(range(last_n_max))\n        return x.tail(n_pickup)\n\n    df = train_df.reset_index(drop=True)\n    valid_df = df.groupby('user_id').progress_apply(_lastn_pickup, n=max_n)\n    valid_idx = valid_df.index.get_level_values(1).values\n    train_idx = np.delete(range(len(df)), valid_idx)\n\n    return train_idx, valid_idx\n```",
    "1066072": "For me it was both time and RAM. My RAM initially exploded to 60GB+. Have you tried using a lower max_bin? Plus, a larger value for `bin_construct_sample_cnt` will also increase the data loading time. Are you doing a k fold Cv to do target-encoding? I am suing the below attached and it runs quite fast. (NB Not using any cat features explicitly, it's auto by default)\n\nGlad to see you here Sakami!\n\n```\n    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': \"binary\",\n        'metric': metrics,\n        'learning_rate': 0.1,\n        'num_leaves': 2**6,\n        'max_depth': 8,\n        'colsample_bytree': 0.7,\n        'min_child_samples': 100,\n        'subsample': 0.7, # ideally  it can be sqrt(number of features)/(number of features)\n        'num_threads': 8, # real_cores\n        'seed': 2020,\n        'first_metric_only': True,\n        'use_two_round_loading': True,\n        'max_bin':128,\n        'verbose': -1\n    }\n```",
    "1066333": "Thank you @adityaecdrid for your kind comment.\n\n> Have you tried using a lower max_bin?\n\nYes, I tried `max_bin = 255` but it didn't change drastically... :(\n\n> Are you doing a k fold Cv to do target-encoding?\n\nYes, I'm using fold target encoding in `content_id_target_encoded`.\n\n> I am suing the below attached and it runs quite fast.\n\nThank you, I'll tune my parameters with reference to yours.",
    "1066815": "Wow, doing a k-fold CV when we have 100M rows is 🙏. And these features are definitely gonna help out as well;",
    "1069224": "what about train the model in your own machine，and upload the model file as private dataset\nI dont know the auc improvement of the rounds>200,1000... , I guess it did not get much gain(may not need to do such many rounds?). my case below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5900557%2F2c8028c52cba45d5294ad5b0e0400153%2F_20201104160629.png?generation=1604477778718612&alt=media)",
    "1071972": "Thank you for your advice!\nIn my case, training a model in local environment and predicting with 1000 rounds also got \"submission scoring error\". 😓",
    "1071974": "That's sad to hear but i am pretty sure that you will resolve the problem pretty soon :)",
    "1071985": "your code may cost too much time on merge features for provided samples(for lgbm model with 1000 rounds cost only about 0.03s to make prediction for provided 120 samples but more than 3s to merge features with 4 merge operations).check whole time the loop you make to make prediction for the provided 120 samples,It should be no more than 5s.",
    "1072252": "Sakami, How's the speed for you now? Yesterday i tried on 100M rows on MBP and even after 8 hours, it's just 1000 rounds.",
    "1082058": "Legend is legend after all 🙏; Any tips ? Ty!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F7c2abedd70610dd42751582e2ffdb517%2FScreenshot%202020-11-17%20at%208.20.53%20PM.png?generation=1605624715514499&alt=media)",
    "1083868": "Thanks!\n\nI completely changed my approach. Our best score so far (LB 0.780) is by a SAINT model.",
    "1083877": "Oh wow! Any tips you wanna share in the SAINT discussion thread :🙏"
  },
  "source": "meta"
}