{
  "id": 206529,
  "title": "Tricks & Tips to Dramatically Speed Up Submission Running",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206529",
  "author_name": "william.wu",
  "post_date": "2020-12-25T05:11:20.818000",
  "votes": 46,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I saw many participants are struggling with submission running. Here are the tricks I used to speed up the submission running. I used 69 features with a single LGBM model and submission running is &lt; 40 mins. I also share some code snippets about the implementation here, hope it will be helpful to you.</p>\n<ul>\n<li>Store users' features and questions' features in <code>{u_id: UserFeats}</code> and <code>{q_id: QuesFeats}</code> format, this can increase the cache hit rate when looking up which means when you are looking up features of a given user, all his features will be in the cache.</li>\n<li>Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in <code>test_df</code>. (Refer to <code>_get_row_fvs</code> function in <code>FeatureEngineer</code> below.)</li>\n</ul>\n<p><strong>UserFeats</strong></p>\n<pre><code>class UserFeats(object):\n\n    def __init__(\n        self\n    ):\n        self._ans_cnt, self._ans_corr_cnt = 0, 0\n\n    def get_ans_cnt(self):\n        return self._ans_cnt\n\n    def get_ans_corr_cnt(self):\n        return self._ans_corr_cnt\n\n    def incr_ans_cnt(self, val):\n        self._ans_cnt += val\n\n    def incr_ans_corr_cnt(self, val):\n        self._ans_corr_cnt += val\n</code></pre>\n<p><strong>QuesFeats</strong></p>\n<pre><code>class QuesFeats(object):\n    def __init__(self, feats_tuple):\n        self.acc = feats_tuple[self.FEAT_KEY_TO_INDEX[self.ACC]]\n        self.cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CNT]]\n        self.corr_cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CORR_CNT]]\n</code></pre>\n<p><strong>FeatureEngineer for feature extraction during inference</strong></p>\n<pre><code>class FeatureEngineer(object):\n    def __init__(self, user_feats, ques_feats, elapsed_time_mean, feat_names):\n        self._user_feats = user_feats\n        self._ques_feats = ques_feats\n        self._elapsed_time_mean = elapsed_time_mean\n        self._init_feat_extact_fns(feat_names)\n\n    def _init_feat_extact_fns(self):\n        name_to_extract_fn = {\n            \"prior_question_elapsed_time\": lambda u_feats, q_feats, row, cache: row[\n                \"prior_question_elapsed_time\"\n            ]\n            if is_valid(row[\"prior_question_elapsed_time\"])\n            else self._elapsed_time_mean,\n            \"prior_question_had_explanation\": lambda u_feats, q_feats, row, cache: 1\n            if (row[\"prior_question_had_explanation\"] is True)\n            or (row[\"prior_question_had_explanation\"] == 1)\n            else 0,\n            \"user_ans_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_cnt(),\n            \"user_ans_corr_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_corr_cnt(),\n            \"ques_acc\": lambda u_feats, q_feats, row, cache: q_feats.acc,\n            \"ques_cnt\": lambda u_feats, q_feats, row, cache: q_feats.cnt,\n            \"ques_corr_cnt\": lambda u_feats, q_feats, row, cache: q_feats.corr_cnt,\n        }\n        self._feat_extract_fns = [name_to_extract_fn[name] for name in feat_names]\n\n\n    def get_feat_values_and_pre_labels(\n        self, df: DataFrame, update_per_row: bool = False\n    ) -&gt; Tuple[np.ndarray, List[int]]:\n        \"\"\"\n        Get feature values and previous labels based on test_df\n        \"\"\"\n        pre_labels = ... # extract previous labels\n        # update if necessary\n        fvs = []\n        for i, row in df.iterrows():\n            # skip lecture\n            u_feats = self._user_feats[row[\"user_id\"]]\n            q_feats = self._ques_feats[row[\"content_id\"]]\n            fvs.append(self._get_row_fvs(row, u_feats, q_feats))\n\n\n        return np.array(fvs), pre_labels\n\n\n    def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -&gt; List[float]:\n         cache = ... # extract stats that is used by multiple features\n         return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n</code></pre>",
  "messages": [
    {
      "id": 1125813,
      "postDate": "2020-12-25T05:11:20.820Z",
      "content": "<p>I saw many participants are struggling with submission running. Here are the tricks I used to speed up the submission running. I used 69 features with a single LGBM model and submission running is &lt; 40 mins. I also share some code snippets about the implementation here, hope it will be helpful to you.</p>\n<ul>\n<li>Store users' features and questions' features in <code>{u_id: UserFeats}</code> and <code>{q_id: QuesFeats}</code> format, this can increase the cache hit rate when looking up which means when you are looking up features of a given user, all his features will be in the cache.</li>\n<li>Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in <code>test_df</code>. (Refer to <code>_get_row_fvs</code> function in <code>FeatureEngineer</code> below.)</li>\n</ul>\n<p><strong>UserFeats</strong></p>\n<pre><code>class UserFeats(object):\n\n    def __init__(\n        self\n    ):\n        self._ans_cnt, self._ans_corr_cnt = 0, 0\n\n    def get_ans_cnt(self):\n        return self._ans_cnt\n\n    def get_ans_corr_cnt(self):\n        return self._ans_corr_cnt\n\n    def incr_ans_cnt(self, val):\n        self._ans_cnt += val\n\n    def incr_ans_corr_cnt(self, val):\n        self._ans_corr_cnt += val\n</code></pre>\n<p><strong>QuesFeats</strong></p>\n<pre><code>class QuesFeats(object):\n    def __init__(self, feats_tuple):\n        self.acc = feats_tuple[self.FEAT_KEY_TO_INDEX[self.ACC]]\n        self.cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CNT]]\n        self.corr_cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CORR_CNT]]\n</code></pre>\n<p><strong>FeatureEngineer for feature extraction during inference</strong></p>\n<pre><code>class FeatureEngineer(object):\n    def __init__(self, user_feats, ques_feats, elapsed_time_mean, feat_names):\n        self._user_feats = user_feats\n        self._ques_feats = ques_feats\n        self._elapsed_time_mean = elapsed_time_mean\n        self._init_feat_extact_fns(feat_names)\n\n    def _init_feat_extact_fns(self):\n        name_to_extract_fn = {\n            \"prior_question_elapsed_time\": lambda u_feats, q_feats, row, cache: row[\n                \"prior_question_elapsed_time\"\n            ]\n            if is_valid(row[\"prior_question_elapsed_time\"])\n            else self._elapsed_time_mean,\n            \"prior_question_had_explanation\": lambda u_feats, q_feats, row, cache: 1\n            if (row[\"prior_question_had_explanation\"] is True)\n            or (row[\"prior_question_had_explanation\"] == 1)\n            else 0,\n            \"user_ans_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_cnt(),\n            \"user_ans_corr_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_corr_cnt(),\n            \"ques_acc\": lambda u_feats, q_feats, row, cache: q_feats.acc,\n            \"ques_cnt\": lambda u_feats, q_feats, row, cache: q_feats.cnt,\n            \"ques_corr_cnt\": lambda u_feats, q_feats, row, cache: q_feats.corr_cnt,\n        }\n        self._feat_extract_fns = [name_to_extract_fn[name] for name in feat_names]\n\n\n    def get_feat_values_and_pre_labels(\n        self, df: DataFrame, update_per_row: bool = False\n    ) -&gt; Tuple[np.ndarray, List[int]]:\n        \"\"\"\n        Get feature values and previous labels based on test_df\n        \"\"\"\n        pre_labels = ... # extract previous labels\n        # update if necessary\n        fvs = []\n        for i, row in df.iterrows():\n            # skip lecture\n            u_feats = self._user_feats[row[\"user_id\"]]\n            q_feats = self._ques_feats[row[\"content_id\"]]\n            fvs.append(self._get_row_fvs(row, u_feats, q_feats))\n\n\n        return np.array(fvs), pre_labels\n\n\n    def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -&gt; List[float]:\n         cache = ... # extract stats that is used by multiple features\n         return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n</code></pre>",
      "rawMarkdown": "I saw many participants are struggling with submission running. Here are the tricks I used to speed up the submission running. I used 69 features with a single LGBM model and submission running is < 40 mins. I also share some code snippets about the implementation here, hope it will be helpful to you.\n* Store users' features and questions' features in `{u_id: UserFeats}` and `{q_id: QuesFeats}` format, this can increase the cache hit rate when looking up which means when you are looking up features of a given user, all his features will be in the cache.\n* Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in `test_df`. (Refer to `_get_row_fvs ` function in `FeatureEngineer` below.)\n\n**UserFeats**\n```Python\nclass UserFeats(object):\n\n    def __init__(\n        self\n    ):\n        self._ans_cnt, self._ans_corr_cnt = 0, 0\n\n    def get_ans_cnt(self):\n        return self._ans_cnt\n\n    def get_ans_corr_cnt(self):\n        return self._ans_corr_cnt\n\n    def incr_ans_cnt(self, val):\n        self._ans_cnt += val\n\n    def incr_ans_corr_cnt(self, val):\n        self._ans_corr_cnt += val\n```\n\n**QuesFeats**\n```Python\nclass QuesFeats(object):\n\tdef __init__(self, feats_tuple):\n        self.acc = feats_tuple[self.FEAT_KEY_TO_INDEX[self.ACC]]\n        self.cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CNT]]\n        self.corr_cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CORR_CNT]]\n```\n\n**FeatureEngineer for feature extraction during inference**\n```\nclass FeatureEngineer(object):\n\tdef __init__(self, user_feats, ques_feats, elapsed_time_mean, feat_names):\n\t\tself._user_feats = user_feats\n\t\tself._ques_feats = ques_feats\n\t\tself._elapsed_time_mean = elapsed_time_mean\n\t\tself._init_feat_extact_fns(feat_names)\n\n\tdef _init_feat_extact_fns(self):\n\t\tname_to_extract_fn = {\n\t\t\t\"prior_question_elapsed_time\": lambda u_feats, q_feats, row, cache: row[\n                \"prior_question_elapsed_time\"\n            ]\n            if is_valid(row[\"prior_question_elapsed_time\"])\n            else self._elapsed_time_mean,\n            \"prior_question_had_explanation\": lambda u_feats, q_feats, row, cache: 1\n            if (row[\"prior_question_had_explanation\"] is True)\n            or (row[\"prior_question_had_explanation\"] == 1)\n            else 0,\n            \"user_ans_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_cnt(),\n            \"user_ans_corr_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_corr_cnt(),\n            \"ques_acc\": lambda u_feats, q_feats, row, cache: q_feats.acc,\n            \"ques_cnt\": lambda u_feats, q_feats, row, cache: q_feats.cnt,\n            \"ques_corr_cnt\": lambda u_feats, q_feats, row, cache: q_feats.corr_cnt,\n\t\t}\n\t\tself._feat_extract_fns = [name_to_extract_fn[name] for name in feat_names]\n\n\n\tdef get_feat_values_and_pre_labels(\n        self, df: DataFrame, update_per_row: bool = False\n    ) -> Tuple[np.ndarray, List[int]]:\n    \t\"\"\"\n\t\tGet feature values and previous labels based on test_df\n    \t\"\"\"\n    \tpre_labels = ... # extract previous labels\n    \t# update if necessary\n    \tfvs = []\n    \tfor i, row in df.iterrows():\n    \t\t# skip lecture\n    \t\tu_feats = self._user_feats[row[\"user_id\"]]\n    \t\tq_feats = self._ques_feats[row[\"content_id\"]]\n    \t\tfvs.append(self._get_row_fvs(row, u_feats, q_feats))\n\n\n    \treturn np.array(fvs), pre_labels\n\n\n    def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -> List[float]:\n \t\tcache = ... # extract stats that is used by multiple features\n \t\treturn [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n```",
      "votes": 46
    },
    {
      "id": 1184380,
      "postDate": "2021-02-03T14:06:48.860Z",
      "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> Hi William. Congrats on your work in the competition. I was so impressed in this thread that I have started to apply your code to the small problem set. But still stucking on how to compare my own code and yours. Is there any plan to share some sort of real usage of your code? If so, it will be really helpful!</p>",
      "rawMarkdown": "@wuwenmin Hi William. Congrats on your work in the competition. I was so impressed in this thread that I have started to apply your code to the small problem set. But still stucking on how to compare my own code and yours. Is there any plan to share some sort of real usage of your code? If so, it will be really helpful!"
    },
    {
      "id": 1133157,
      "postDate": "2020-12-31T01:50:36.850Z",
      "content": "<p>I'd like to know why <code>this can increase the cache hit rate</code>!</p>",
      "rawMarkdown": "I'd like to know why `this can increase the cache hit rate`!",
      "replies": [
        {
          "id": 1133869,
          "postDate": "2020-12-31T15:38:21.973Z",
          "content": "<p>Comparing to store <code>user_cnts</code>, <code>user_corr_cnts</code>,  <code>user_part_cnts</code>, <code>user_part_corr_cnts</code> separately in <code>{u_id: cnt}</code>, <code>{u_id: corr_cnt}</code>, <code>{u_id: {part: cnt}}</code>, <code>{u_id: {part: corr_cnt}}</code> just like most public notebooks did. Let's suppose you are retrieving features for user <code>0</code>, because the limit size of cache (several KBs), some parts of <code>user_cnts</code>,   …, <code>user_part_corr_cnts</code> will be loaded into cache sequentially, which means <code>cache_hit_rate = 0</code>. In my solution, if you're retrieving features for user <code>0</code>, <code>UserFeats</code> of user <code>0</code> can be loaded into cache when retrieving the 1st feature, since the size of <code>UserFeats</code> is too small. So cache hit rate is <code>(n_user_feats - 1) / n_user_features</code>.</p>\n<p>I just checked the Kaggle kernel caches' sizes:</p>\n<pre><code>L1d cache:           32K\nL1i cache:           32K\nL2 cache:            256K\nL3 cache:            56320K\n</code></pre>\n<p><code>L1</code> cache can hold <code>1142</code> python <code>ints</code>, it's enough for most users.</p>",
          "rawMarkdown": "Comparing to store `user_cnts`, `user_corr_cnts`,  `user_part_cnts`, `user_part_corr_cnts` separately in `{u_id: cnt}`, `{u_id: corr_cnt}`, `{u_id: {part: cnt}}`, `{u_id: {part: corr_cnt}}` just like most public notebooks did. Let's suppose you are retrieving features for user `0`, because the limit size of cache (several KBs), some parts of `user_cnts`,   ..., `user_part_corr_cnts` will be loaded into cache sequentially, which means `cache_hit_rate = 0`. In my solution, if you're retrieving features for user `0`, `UserFeats` of user `0` can be loaded into cache when retrieving the 1st feature, since the size of `UserFeats` is too small. So cache hit rate is `(n_user_feats - 1) / n_user_features`.\n\nI just checked the Kaggle kernel caches' sizes:\n```\nL1d cache:           32K\nL1i cache:           32K\nL2 cache:            256K\nL3 cache:            56320K\n```\n`L1` cache can hold `1142` python `ints`, it's enough for most users.",
          "votes": 2
        },
        {
          "id": 1185900,
          "postDate": "2021-02-04T13:22:41.960Z",
          "content": "<p>But If u saved like <code>{(u_id, part): cnt}</code>, <code>{(u_id, part): corr_cnt}</code> instead of <code>{u_id: {part: cnt}}</code>, <code>{u_id: {part: corr_cnt}}</code>, then there would be no cache issue, wouldn't it? </p>",
          "rawMarkdown": "But If u saved like `{(u_id, part): cnt}`, `{(u_id, part): corr_cnt}` instead of `{u_id: {part: cnt}}`, `{u_id: {part: corr_cnt}}`, then there would be no cache issue, wouldn't it? "
        }
      ]
    },
    {
      "id": 1130441,
      "postDate": "2020-12-29T03:59:48.010Z",
      "content": "<p>Great tricks, thanks for sharing. One question:</p>\n<blockquote>\n  <p>Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in test_df. (Refer to _get_row_fvs function in FeatureEngineer below.)</p>\n</blockquote>\n<p>In the list of feature extraction functions, there is also if/else statements to extract features for each test_df row, only it`s in lambda functions, could you please explain why using if/else of lambda fuction is more efficient than direct if/else when looping test_df rows? Thanks.</p>",
      "rawMarkdown": "Great tricks, thanks for sharing. One question:\n\n> Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in test_df. (Refer to _get_row_fvs function in FeatureEngineer below.)\n\nIn the list of feature extraction functions, there is also if/else statements to extract features for each test_df row, only it`s in lambda functions, could you please explain why using if/else of lambda fuction is more efficient than direct if/else when looping test_df rows? Thanks.",
      "replies": [
        {
          "id": 1130743,
          "postDate": "2020-12-29T09:38:51.970Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a>, the <code>if/else</code> of lambda functions in <code>_init_feat_extact_fns</code> run only once during the initialization. Then extracting features for each row will call <code>_get_row_fvs</code>, there's no <code>if/else</code>. The complexity of <code>if/else</code> is O(N) depends on how many branches you have, whereas the complexity of the lambda function is <code>O(1)</code>. That's why a list of <code>lambda functions</code> is faster than do <code>if/else</code> judgment for each feature of each row.</p>",
          "rawMarkdown": "Hi @superchenhao, the `if/else` of lambda functions in `_init_feat_extact_fns ` run only once during the initialization. Then extracting features for each row will call `_get_row_fvs `, there's no `if/else`. The complexity of `if/else` is O(N) depends on how many branches you have, whereas the complexity of the lambda function is `O(1)`. That's why a list of `lambda functions` is faster than do `if/else` judgment for each feature of each row."
        },
        {
          "id": 1132297,
          "postDate": "2020-12-30T09:45:04.223Z",
          "content": "<p>Thanks for your explaination, but I just tried this approach and found my online inference speed still no change. I still need 4+ hours totally to run FE+training+Inference, similar speed as before. I<code>m wondering whether something still need improve in my code, if you don</code>t mind, could you please share how you handle prior_test_df for updating necessary values? My prior_test_df updating is common in other public kernel, it loop over each row of prior_test_df and update nessary values to {u_id: UserFeats} and {q_id: QuesFeats}, I doubt this part of code slows my inference down, any tricks in this part? Thanks.</p>",
          "rawMarkdown": "Thanks for your explaination, but I just tried this approach and found my online inference speed still no change. I still need 4+ hours totally to run FE+training+Inference, similar speed as before. I`m wondering whether something still need improve in my code, if you don`t mind, could you please share how you handle prior_test_df for updating necessary values? My prior_test_df updating is common in other public kernel, it loop over each row of prior_test_df and update nessary values to {u_id: UserFeats} and {q_id: QuesFeats}, I doubt this part of code slows my inference down, any tricks in this part? Thanks.",
          "votes": 1
        },
        {
          "id": 1133123,
          "postDate": "2020-12-31T00:45:32.533Z",
          "content": "<p>Yup, this part could also impact the speed, I don't use prior_test_df to update the states. Instead I store prior values in FeatureEngineer fields when get feature values for prior test df, such as user ids and question parts. When processing current test df, I extract prior targets using json.loads, then the states can be updated based on prior values and prior targets.</p>\n<p>BTW, I trained the model offline and upload the trained model and states for inference. So the 40mins is the only the inference time. I'm using 80 features now. How these trucks can speed up your inference mainly depends on how many features you have. You can try to do inference from a trained model to know the actual inference time. It can also help you when the training takes hours to finish.</p>",
          "rawMarkdown": "Yup, this part could also impact the speed, I don't use prior_test_df to update the states. Instead I store prior values in FeatureEngineer fields when get feature values for prior test df, such as user ids and question parts. When processing current test df, I extract prior targets using json.loads, then the states can be updated based on prior values and prior targets.\n\nBTW, I trained the model offline and upload the trained model and states for inference. So the 40mins is the only the inference time. I'm using 80 features now. How these trucks can speed up your inference mainly depends on how many features you have. You can try to do inference from a trained model to know the actual inference time. It can also help you when the training takes hours to finish.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1127513,
      "postDate": "2020-12-26T15:46:52.157Z",
      "content": "<p>谢谢大佬，很有帮助</p>",
      "rawMarkdown": "谢谢大佬，很有帮助\n"
    },
    {
      "id": 1126575,
      "postDate": "2020-12-25T17:49:41.970Z",
      "content": "<p>很棒，我也是这么做的。😝</p>",
      "rawMarkdown": "很棒，我也是这么做的。😝",
      "replies": [
        {
          "id": 1127280,
          "postDate": "2020-12-26T11:47:28.443Z",
          "content": "<p>da lao         </p>",
          "rawMarkdown": "da lao         "
        }
      ]
    },
    {
      "id": 1581056,
      "postDate": "2021-11-13T11:11:59.227Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1184380,
      "author_name": "jwc",
      "author_url": "",
      "post_date": "2021-02-03T14:06:48.860000",
      "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> Hi William. Congrats on your work in the competition. I was so impressed in this thread that I have started to apply your code to the small problem set. But still stucking on how to compare my own code and yours. Is there any plan to share some sort of real usage of your code? If so, it will be really helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1133157,
      "author_name": "jwc",
      "author_url": "",
      "post_date": "2020-12-31T01:50:36.850000",
      "content": "<p>I'd like to know why <code>this can increase the cache hit rate</code>!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1133869,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2020-12-31T15:38:21.973000",
          "content": "<p>Comparing to store <code>user_cnts</code>, <code>user_corr_cnts</code>,  <code>user_part_cnts</code>, <code>user_part_corr_cnts</code> separately in <code>{u_id: cnt}</code>, <code>{u_id: corr_cnt}</code>, <code>{u_id: {part: cnt}}</code>, <code>{u_id: {part: corr_cnt}}</code> just like most public notebooks did. Let's suppose you are retrieving features for user <code>0</code>, because the limit size of cache (several KBs), some parts of <code>user_cnts</code>,   …, <code>user_part_corr_cnts</code> will be loaded into cache sequentially, which means <code>cache_hit_rate = 0</code>. In my solution, if you're retrieving features for user <code>0</code>, <code>UserFeats</code> of user <code>0</code> can be loaded into cache when retrieving the 1st feature, since the size of <code>UserFeats</code> is too small. So cache hit rate is <code>(n_user_feats - 1) / n_user_features</code>.</p>\n<p>I just checked the Kaggle kernel caches' sizes:</p>\n<pre><code>L1d cache:           32K\nL1i cache:           32K\nL2 cache:            256K\nL3 cache:            56320K\n</code></pre>\n<p><code>L1</code> cache can hold <code>1142</code> python <code>ints</code>, it's enough for most users.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1185900,
          "author_name": "jwc",
          "author_url": "",
          "post_date": "2021-02-04T13:22:41.960000",
          "content": "<p>But If u saved like <code>{(u_id, part): cnt}</code>, <code>{(u_id, part): corr_cnt}</code> instead of <code>{u_id: {part: cnt}}</code>, <code>{u_id: {part: corr_cnt}}</code>, then there would be no cache issue, wouldn't it? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1130441,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2020-12-29T03:59:48.010000",
      "content": "<p>Great tricks, thanks for sharing. One question:</p>\n<blockquote>\n  <p>Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in test_df. (Refer to _get_row_fvs function in FeatureEngineer below.)</p>\n</blockquote>\n<p>In the list of feature extraction functions, there is also if/else statements to extract features for each test_df row, only it`s in lambda functions, could you please explain why using if/else of lambda fuction is more efficient than direct if/else when looping test_df rows? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1130743,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2020-12-29T09:38:51.970000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a>, the <code>if/else</code> of lambda functions in <code>_init_feat_extact_fns</code> run only once during the initialization. Then extracting features for each row will call <code>_get_row_fvs</code>, there's no <code>if/else</code>. The complexity of <code>if/else</code> is O(N) depends on how many branches you have, whereas the complexity of the lambda function is <code>O(1)</code>. That's why a list of <code>lambda functions</code> is faster than do <code>if/else</code> judgment for each feature of each row.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1132297,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2020-12-30T09:45:04.223000",
          "content": "<p>Thanks for your explaination, but I just tried this approach and found my online inference speed still no change. I still need 4+ hours totally to run FE+training+Inference, similar speed as before. I<code>m wondering whether something still need improve in my code, if you don</code>t mind, could you please share how you handle prior_test_df for updating necessary values? My prior_test_df updating is common in other public kernel, it loop over each row of prior_test_df and update nessary values to {u_id: UserFeats} and {q_id: QuesFeats}, I doubt this part of code slows my inference down, any tricks in this part? Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1133123,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2020-12-31T00:45:32.533000",
          "content": "<p>Yup, this part could also impact the speed, I don't use prior_test_df to update the states. Instead I store prior values in FeatureEngineer fields when get feature values for prior test df, such as user ids and question parts. When processing current test df, I extract prior targets using json.loads, then the states can be updated based on prior values and prior targets.</p>\n<p>BTW, I trained the model offline and upload the trained model and states for inference. So the 40mins is the only the inference time. I'm using 80 features now. How these trucks can speed up your inference mainly depends on how many features you have. You can try to do inference from a trained model to know the actual inference time. It can also help you when the training takes hours to finish.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1127513,
      "author_name": "Y.L",
      "author_url": "",
      "post_date": "2020-12-26T15:46:52.157000",
      "content": "<p>谢谢大佬，很有帮助</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1126575,
      "author_name": "林有夕",
      "author_url": "",
      "post_date": "2020-12-25T17:49:41.970000",
      "content": "<p>很棒，我也是这么做的。😝</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1127280,
          "author_name": "Zhenghan Chen",
          "author_url": "",
          "post_date": "2020-12-26T11:47:28.443000",
          "content": "<p>da lao         </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1581056,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-13T11:11:59.227000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125813": "I saw many participants are struggling with submission running. Here are the tricks I used to speed up the submission running. I used 69 features with a single LGBM model and submission running is < 40 mins. I also share some code snippets about the implementation here, hope it will be helpful to you.\n* Store users' features and questions' features in `{u_id: UserFeats}` and `{q_id: QuesFeats}` format, this can increase the cache hit rate when looking up which means when you are looking up features of a given user, all his features will be in the cache.\n* Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in `test_df`. (Refer to `_get_row_fvs ` function in `FeatureEngineer` below.)\n\n**UserFeats**\n```Python\nclass UserFeats(object):\n\n    def __init__(\n        self\n    ):\n        self._ans_cnt, self._ans_corr_cnt = 0, 0\n\n    def get_ans_cnt(self):\n        return self._ans_cnt\n\n    def get_ans_corr_cnt(self):\n        return self._ans_corr_cnt\n\n    def incr_ans_cnt(self, val):\n        self._ans_cnt += val\n\n    def incr_ans_corr_cnt(self, val):\n        self._ans_corr_cnt += val\n```\n\n**QuesFeats**\n```Python\nclass QuesFeats(object):\n\tdef __init__(self, feats_tuple):\n        self.acc = feats_tuple[self.FEAT_KEY_TO_INDEX[self.ACC]]\n        self.cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CNT]]\n        self.corr_cnt = feats_tuple[self.FEAT_KEY_TO_INDEX[self.CORR_CNT]]\n```\n\n**FeatureEngineer for feature extraction during inference**\n```\nclass FeatureEngineer(object):\n\tdef __init__(self, user_feats, ques_feats, elapsed_time_mean, feat_names):\n\t\tself._user_feats = user_feats\n\t\tself._ques_feats = ques_feats\n\t\tself._elapsed_time_mean = elapsed_time_mean\n\t\tself._init_feat_extact_fns(feat_names)\n\n\tdef _init_feat_extact_fns(self):\n\t\tname_to_extract_fn = {\n\t\t\t\"prior_question_elapsed_time\": lambda u_feats, q_feats, row, cache: row[\n                \"prior_question_elapsed_time\"\n            ]\n            if is_valid(row[\"prior_question_elapsed_time\"])\n            else self._elapsed_time_mean,\n            \"prior_question_had_explanation\": lambda u_feats, q_feats, row, cache: 1\n            if (row[\"prior_question_had_explanation\"] is True)\n            or (row[\"prior_question_had_explanation\"] == 1)\n            else 0,\n            \"user_ans_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_cnt(),\n            \"user_ans_corr_cnt\": lambda u_feats, q_feats, row, cache: u_feats.get_ans_corr_cnt(),\n            \"ques_acc\": lambda u_feats, q_feats, row, cache: q_feats.acc,\n            \"ques_cnt\": lambda u_feats, q_feats, row, cache: q_feats.cnt,\n            \"ques_corr_cnt\": lambda u_feats, q_feats, row, cache: q_feats.corr_cnt,\n\t\t}\n\t\tself._feat_extract_fns = [name_to_extract_fn[name] for name in feat_names]\n\n\n\tdef get_feat_values_and_pre_labels(\n        self, df: DataFrame, update_per_row: bool = False\n    ) -> Tuple[np.ndarray, List[int]]:\n    \t\"\"\"\n\t\tGet feature values and previous labels based on test_df\n    \t\"\"\"\n    \tpre_labels = ... # extract previous labels\n    \t# update if necessary\n    \tfvs = []\n    \tfor i, row in df.iterrows():\n    \t\t# skip lecture\n    \t\tu_feats = self._user_feats[row[\"user_id\"]]\n    \t\tq_feats = self._ques_feats[row[\"content_id\"]]\n    \t\tfvs.append(self._get_row_fvs(row, u_feats, q_feats))\n\n\n    \treturn np.array(fvs), pre_labels\n\n\n    def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -> List[float]:\n \t\tcache = ... # extract stats that is used by multiple features\n \t\treturn [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n```",
    "1184380": "@wuwenmin Hi William. Congrats on your work in the competition. I was so impressed in this thread that I have started to apply your code to the small problem set. But still stucking on how to compare my own code and yours. Is there any plan to share some sort of real usage of your code? If so, it will be really helpful!",
    "1133157": "I'd like to know why `this can increase the cache hit rate`!",
    "1130441": "Great tricks, thanks for sharing. One question:\n\n> Extract features with a list of feature extraction functions instead of stupid if/else statements for each row in test_df. (Refer to _get_row_fvs function in FeatureEngineer below.)\n\nIn the list of feature extraction functions, there is also if/else statements to extract features for each test_df row, only it`s in lambda functions, could you please explain why using if/else of lambda fuction is more efficient than direct if/else when looping test_df rows? Thanks.",
    "1127513": "谢谢大佬，很有帮助\n",
    "1126575": "很棒，我也是这么做的。😝",
    "1581056": ""
  }
}