{
  "id": 205199,
  "title": "Handy Class to Calculate Last N Stats in O(1) Time",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205199",
  "author_name": "",
  "post_date": "2020-12-19T01:27:58.466362400Z",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Since users' skills grow over time, so the recent performance is more important, also it's easy to calculate the moving average performance based on the last N stats. I wrote a class that calculates the last N Stats such as <code>correct_cnt</code> (calls <code>get_sum</code>) and <code>accuracy</code> (calls <code>get_mean</code>) in O(1) time. Hopes it's also helpful to you.</p>\n<h2>Class</h2>\n<pre><code>from collections import deque\n\nclass LastNStats(object):\n    def __init__(self, n, vals=None):\n        self.n = n\n        self.cum_sums = deque()\n        self.left_cum_sum = 0\n        if vals is not None:\n            self.init(vals)\n\n    def init(self, vals):\n        for val in vals:\n            self.append(val)\n\n    def append(self, val):\n        if len(self.cum_sums) == self.n:\n            self.left_cum_sum = self.cum_sums.popleft()\n        pre_sum = self.cum_sums[-1] if len(self.cum_sums) &gt; 0 else 0\n        self.cum_sums.append(pre_sum + val)\n\n    def get_sum(self):\n        if len(self.cum_sums) &lt; self.n:\n            return 0\n        return self.cum_sums[-1] - self.left_cum_sum\n\n    def get_mean(self):\n        sum_ = self.get_sum()\n        return None if len(self.cum_sums) &lt; self.n else sum_ / self.n\n</code></pre>\n<h2>Usage</h2>\n<pre><code>from collections import defaultdict\n\nALPHA = 0.8\n\nu_last_20_stats = defaultdict(lambda: LastNStats(n=20))\nu_moving_acc = {}\n\nfor i, row in pdf.iterrows():\n    u = row[\"user_id\"]\n    ans_res = row[\"answered_correctly\"]\n    last_20_acc = u_last_n_stats[u].get_mean()\n    # add last_20_acc to feature dataframe\n\n    mov_acc = u_moving_acc.get(u)\n    # add mov_acc to feature dataframe\n\n\n    # Update\n    u_last_n_stats[u].append(ans_res)\n    if mov_acc is not None:\n        mov_acc = (1 - ALPHA) * mov_acc + ALPHA * u_last_n_stats[u].get_mean()\n    else:\n        mov_acc = last_20_acc\n    u_moving_acc[u] = mov_acc\n</code></pre>",
  "messages": [
    {
      "id": "1118344",
      "postDate": "12/19/2020 01:27:58",
      "content": "<p>Since users' skills grow over time, so the recent performance is more important, also it's easy to calculate the moving average performance based on the last N stats. I wrote a class that calculates the last N Stats such as <code>correct_cnt</code> (calls <code>get_sum</code>) and <code>accuracy</code> (calls <code>get_mean</code>) in O(1) time. Hopes it's also helpful to you.</p>\n<h2>Class</h2>\n<pre><code>from collections import deque\n\nclass LastNStats(object):\n    def __init__(self, n, vals=None):\n        self.n = n\n        self.cum_sums = deque()\n        self.left_cum_sum = 0\n        if vals is not None:\n            self.init(vals)\n\n    def init(self, vals):\n        for val in vals:\n            self.append(val)\n\n    def append(self, val):\n        if len(self.cum_sums) == self.n:\n            self.left_cum_sum = self.cum_sums.popleft()\n        pre_sum = self.cum_sums[-1] if len(self.cum_sums) &gt; 0 else 0\n        self.cum_sums.append(pre_sum + val)\n\n    def get_sum(self):\n        if len(self.cum_sums) &lt; self.n:\n            return 0\n        return self.cum_sums[-1] - self.left_cum_sum\n\n    def get_mean(self):\n        sum_ = self.get_sum()\n        return None if len(self.cum_sums) &lt; self.n else sum_ / self.n\n</code></pre>\n<h2>Usage</h2>\n<pre><code>from collections import defaultdict\n\nALPHA = 0.8\n\nu_last_20_stats = defaultdict(lambda: LastNStats(n=20))\nu_moving_acc = {}\n\nfor i, row in pdf.iterrows():\n    u = row[\"user_id\"]\n    ans_res = row[\"answered_correctly\"]\n    last_20_acc = u_last_n_stats[u].get_mean()\n    # add last_20_acc to feature dataframe\n\n    mov_acc = u_moving_acc.get(u)\n    # add mov_acc to feature dataframe\n\n\n    # Update\n    u_last_n_stats[u].append(ans_res)\n    if mov_acc is not None:\n        mov_acc = (1 - ALPHA) * mov_acc + ALPHA * u_last_n_stats[u].get_mean()\n    else:\n        mov_acc = last_20_acc\n    u_moving_acc[u] = mov_acc\n</code></pre>",
      "rawMarkdown": "Since users' skills grow over time, so the recent performance is more important, also it's easy to calculate the moving average performance based on the last N stats. I wrote a class that calculates the last N Stats such as `correct_cnt` (calls `get_sum`) and `accuracy` (calls `get_mean`) in O(1) time. Hopes it's also helpful to you.\n\n## Class\n```Python\nfrom collections import deque\n\nclass LastNStats(object):\n    def __init__(self, n, vals=None):\n        self.n = n\n        self.cum_sums = deque()\n        self.left_cum_sum = 0\n        if vals is not None:\n            self.init(vals)\n    \n    def init(self, vals):\n        for val in vals:\n            self.append(val)\n            \n    def append(self, val):\n        if len(self.cum_sums) == self.n:\n            self.left_cum_sum = self.cum_sums.popleft()\n        pre_sum = self.cum_sums[-1] if len(self.cum_sums) > 0 else 0\n        self.cum_sums.append(pre_sum + val)\n            \n    def get_sum(self):\n        if len(self.cum_sums) < self.n:\n            return 0\n        return self.cum_sums[-1] - self.left_cum_sum\n    \n    def get_mean(self):\n        sum_ = self.get_sum()\n        return None if len(self.cum_sums) < self.n else sum_ / self.n\n```\n## Usage\n\n```Python\nfrom collections import defaultdict\n\nALPHA = 0.8\n\nu_last_20_stats = defaultdict(lambda: LastNStats(n=20))\nu_moving_acc = {}\n\nfor i, row in pdf.iterrows():\n\tu = row[\"user_id\"]\n\tans_res = row[\"answered_correctly\"]\n\tlast_20_acc = u_last_n_stats[u].get_mean()\n\t# add last_20_acc to feature dataframe\n\n\tmov_acc = u_moving_acc.get(u)\n\t# add mov_acc to feature dataframe\n\n\n\t# Update\n\tu_last_n_stats[u].append(ans_res)\n\tif mov_acc is not None:\n\t\tmov_acc = (1 - ALPHA) * mov_acc + ALPHA * u_last_n_stats[u].get_mean()\n\telse:\n\t\tmov_acc = last_20_acc\n\tu_moving_acc[u] = mov_acc\n\n```",
      "votes": null
    },
    {
      "id": "1118873",
      "postDate": "12/19/2020 13:40:57",
      "content": "<p>Hi William !</p>\n<p>If I may, you are iterating using the dataframe. You shall try considering iterating on the numpy array, it goes much faster</p>",
      "rawMarkdown": "Hi William !\n\nIf I may, you are iterating using the dataframe. You shall try considering iterating on the numpy array, it goes much faster",
      "votes": null
    },
    {
      "id": "1119333",
      "postDate": "12/19/2020 23:56:31",
      "content": "<p>Thanks for your advice, let me try.</p>",
      "rawMarkdown": "Thanks for your advice, let me try.",
      "votes": null
    },
    {
      "id": "1142333",
      "postDate": "01/07/2021 10:17:35",
      "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>  i think many people use np.append to add their  train data. <br>\nAny efficient method that can be there ?</p>",
      "rawMarkdown": "bowaka  i think many people use np.append to add their  train data. \nAny efficient method that can be there ?",
      "votes": null
    },
    {
      "id": "1142343",
      "postDate": "01/07/2021 10:26:38",
      "content": "<p>numpy.append is slow because numpy has to reallocate memory. A much faster way is to preallocate memory by creating an empty numpy array (with np.zeros) and populating that array using indices like x[idx] = value<br>\npython list doesn't have this problem (but takes a lot of memory because everything uses 64bits unlike numpy types). An alternative is too use the array module of python which both allocates \"intelligently\" and use types.</p>",
      "rawMarkdown": "numpy.append is slow because numpy has to reallocate memory. A much faster way is to preallocate memory by creating an empty numpy array (with np.zeros) and populating that array using indices like x[idx] = value\npython list doesn't have this problem (but takes a lot of memory because everything uses 64bits unlike numpy types). An alternative is too use the array module of python which both allocates \"intelligently\" and use types.",
      "votes": null
    },
    {
      "id": "1142817",
      "postDate": "01/07/2021 15:56:06",
      "content": "<p>hi <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> and <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a>, I didn't use <code>np.append</code> because the complexity of updating any size of data with <code>np.append</code> is O(N), where N is the total number of data you store for that user.  I'm using python deque instead. Just share the structure I used to store features of the user for the SAINT+ model. Since each time for each user, we only have few data points to update, the complexity of updating M data points with deque is O(M).</p>\n<pre><code>class _UserFeats(object):\n    __slots__ = [\n        \"last_ques\", # last questions\n        \"last_ets\",  # last elapsed_times\n        \"last_lts\",  # last lag times\n        \"last_ans\",  # last answers\n    ]\n\n    def __init__(self, max_len, last_seq=None):\n        self.last_ques = deque(maxlen=max_len)\n        self.last_ets = deque(maxlen=max_len)\n        self.last_lts = deque(maxlen=max_len)\n        self.last_ans = deque(maxlen=max_len)\n\n        if last_seq is not None:\n            self.init(last_seq)\n\n    def init(self, last_seq):\n        (\n            last_ques,\n            last_ets,\n            last_lts,\n            last_ans,\n        ) = last_seq\n        self.last_ques.extend(last_ques)\n        self.last_ets.extend(last_ets)\n        self.last_lts.extend(last_lts)\n        self.last_ans.extend(last_ans)\n\n    def add_ques(self, ques):\n        self.last_ques.append(ques)\n\n    def add_et(self, et):\n        self.last_ets.append(et)\n\n    def add_lt(self, lt):\n        self.last_lts.append(lt)\n\n    def add_ans(self, ans):\n        self.last_ans.append(ans)\n\n    def __repr__(self):\n        return \"_UserFeats: last_ques={}, last_ets={}, last_lts={}, last_ans={}\".format(\n            self.last_ques,\n            self.last_ets,\n            self.last_lts,\n            self.last_ans,\n        )\n</code></pre>",
      "rawMarkdown": "hi @jaideepvalani and @rodolphelampe, I didn't use `np.append` because the complexity of updating any size of data with `np.append` is O(N), where N is the total number of data you store for that user.  I'm using python deque instead. Just share the structure I used to store features of the user for the SAINT+ model. Since each time for each user, we only have few data points to update, the complexity of updating M data points with deque is O(M).\n```Python\nclass _UserFeats(object):\n    __slots__ = [\n        \"last_ques\", # last questions\n        \"last_ets\",  # last elapsed_times\n        \"last_lts\",  # last lag times\n        \"last_ans\",  # last answers\n    ]\n\n    def __init__(self, max_len, last_seq=None):\n        self.last_ques = deque(maxlen=max_len)\n        self.last_ets = deque(maxlen=max_len)\n        self.last_lts = deque(maxlen=max_len)\n        self.last_ans = deque(maxlen=max_len)\n\n        if last_seq is not None:\n            self.init(last_seq)\n\n    def init(self, last_seq):\n        (\n            last_ques,\n            last_ets,\n            last_lts,\n            last_ans,\n        ) = last_seq\n        self.last_ques.extend(last_ques)\n        self.last_ets.extend(last_ets)\n        self.last_lts.extend(last_lts)\n        self.last_ans.extend(last_ans)\n\n    def add_ques(self, ques):\n        self.last_ques.append(ques)\n\n    def add_et(self, et):\n        self.last_ets.append(et)\n\n    def add_lt(self, lt):\n        self.last_lts.append(lt)\n\n    def add_ans(self, ans):\n        self.last_ans.append(ans)\n\n    def __repr__(self):\n        return \"_UserFeats: last_ques={}, last_ets={}, last_lts={}, last_ans={}\".format(\n            self.last_ques,\n            self.last_ets,\n            self.last_lts,\n            self.last_ans,\n        )\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1118873,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/19/2020 13:40:57",
      "content": "<p>Hi William !</p>\n<p>If I may, you are iterating using the dataframe. You shall try considering iterating on the numpy array, it goes much faster</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119333,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/19/2020 23:56:31",
          "content": "<p>Thanks for your advice, let me try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142333,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/07/2021 10:17:35",
          "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>  i think many people use np.append to add their  train data. <br>\nAny efficient method that can be there ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142343,
          "author_name": "rodolphelampe",
          "author_url": "",
          "post_date": "01/07/2021 10:26:38",
          "content": "<p>numpy.append is slow because numpy has to reallocate memory. A much faster way is to preallocate memory by creating an empty numpy array (with np.zeros) and populating that array using indices like x[idx] = value<br>\npython list doesn't have this problem (but takes a lot of memory because everything uses 64bits unlike numpy types). An alternative is too use the array module of python which both allocates \"intelligently\" and use types.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142817,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/07/2021 15:56:06",
          "content": "<p>hi <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> and <a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a>, I didn't use <code>np.append</code> because the complexity of updating any size of data with <code>np.append</code> is O(N), where N is the total number of data you store for that user.  I'm using python deque instead. Just share the structure I used to store features of the user for the SAINT+ model. Since each time for each user, we only have few data points to update, the complexity of updating M data points with deque is O(M).</p>\n<pre><code>class _UserFeats(object):\n    __slots__ = [\n        \"last_ques\", # last questions\n        \"last_ets\",  # last elapsed_times\n        \"last_lts\",  # last lag times\n        \"last_ans\",  # last answers\n    ]\n\n    def __init__(self, max_len, last_seq=None):\n        self.last_ques = deque(maxlen=max_len)\n        self.last_ets = deque(maxlen=max_len)\n        self.last_lts = deque(maxlen=max_len)\n        self.last_ans = deque(maxlen=max_len)\n\n        if last_seq is not None:\n            self.init(last_seq)\n\n    def init(self, last_seq):\n        (\n            last_ques,\n            last_ets,\n            last_lts,\n            last_ans,\n        ) = last_seq\n        self.last_ques.extend(last_ques)\n        self.last_ets.extend(last_ets)\n        self.last_lts.extend(last_lts)\n        self.last_ans.extend(last_ans)\n\n    def add_ques(self, ques):\n        self.last_ques.append(ques)\n\n    def add_et(self, et):\n        self.last_ets.append(et)\n\n    def add_lt(self, lt):\n        self.last_lts.append(lt)\n\n    def add_ans(self, ans):\n        self.last_ans.append(ans)\n\n    def __repr__(self):\n        return \"_UserFeats: last_ques={}, last_ets={}, last_lts={}, last_ans={}\".format(\n            self.last_ques,\n            self.last_ets,\n            self.last_lts,\n            self.last_ans,\n        )\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1118344": "Since users' skills grow over time, so the recent performance is more important, also it's easy to calculate the moving average performance based on the last N stats. I wrote a class that calculates the last N Stats such as `correct_cnt` (calls `get_sum`) and `accuracy` (calls `get_mean`) in O(1) time. Hopes it's also helpful to you.\n\n## Class\n```Python\nfrom collections import deque\n\nclass LastNStats(object):\n    def __init__(self, n, vals=None):\n        self.n = n\n        self.cum_sums = deque()\n        self.left_cum_sum = 0\n        if vals is not None:\n            self.init(vals)\n    \n    def init(self, vals):\n        for val in vals:\n            self.append(val)\n            \n    def append(self, val):\n        if len(self.cum_sums) == self.n:\n            self.left_cum_sum = self.cum_sums.popleft()\n        pre_sum = self.cum_sums[-1] if len(self.cum_sums) > 0 else 0\n        self.cum_sums.append(pre_sum + val)\n            \n    def get_sum(self):\n        if len(self.cum_sums) < self.n:\n            return 0\n        return self.cum_sums[-1] - self.left_cum_sum\n    \n    def get_mean(self):\n        sum_ = self.get_sum()\n        return None if len(self.cum_sums) < self.n else sum_ / self.n\n```\n## Usage\n\n```Python\nfrom collections import defaultdict\n\nALPHA = 0.8\n\nu_last_20_stats = defaultdict(lambda: LastNStats(n=20))\nu_moving_acc = {}\n\nfor i, row in pdf.iterrows():\n\tu = row[\"user_id\"]\n\tans_res = row[\"answered_correctly\"]\n\tlast_20_acc = u_last_n_stats[u].get_mean()\n\t# add last_20_acc to feature dataframe\n\n\tmov_acc = u_moving_acc.get(u)\n\t# add mov_acc to feature dataframe\n\n\n\t# Update\n\tu_last_n_stats[u].append(ans_res)\n\tif mov_acc is not None:\n\t\tmov_acc = (1 - ALPHA) * mov_acc + ALPHA * u_last_n_stats[u].get_mean()\n\telse:\n\t\tmov_acc = last_20_acc\n\tu_moving_acc[u] = mov_acc\n\n```",
    "1118873": "Hi William !\n\nIf I may, you are iterating using the dataframe. You shall try considering iterating on the numpy array, it goes much faster",
    "1119333": "Thanks for your advice, let me try.",
    "1142333": "bowaka  i think many people use np.append to add their  train data. \nAny efficient method that can be there ?",
    "1142343": "numpy.append is slow because numpy has to reallocate memory. A much faster way is to preallocate memory by creating an empty numpy array (with np.zeros) and populating that array using indices like x[idx] = value\npython list doesn't have this problem (but takes a lot of memory because everything uses 64bits unlike numpy types). An alternative is too use the array module of python which both allocates \"intelligently\" and use types.",
    "1142817": "hi @jaideepvalani and @rodolphelampe, I didn't use `np.append` because the complexity of updating any size of data with `np.append` is O(N), where N is the total number of data you store for that user.  I'm using python deque instead. Just share the structure I used to store features of the user for the SAINT+ model. Since each time for each user, we only have few data points to update, the complexity of updating M data points with deque is O(M).\n```Python\nclass _UserFeats(object):\n    __slots__ = [\n        \"last_ques\", # last questions\n        \"last_ets\",  # last elapsed_times\n        \"last_lts\",  # last lag times\n        \"last_ans\",  # last answers\n    ]\n\n    def __init__(self, max_len, last_seq=None):\n        self.last_ques = deque(maxlen=max_len)\n        self.last_ets = deque(maxlen=max_len)\n        self.last_lts = deque(maxlen=max_len)\n        self.last_ans = deque(maxlen=max_len)\n\n        if last_seq is not None:\n            self.init(last_seq)\n\n    def init(self, last_seq):\n        (\n            last_ques,\n            last_ets,\n            last_lts,\n            last_ans,\n        ) = last_seq\n        self.last_ques.extend(last_ques)\n        self.last_ets.extend(last_ets)\n        self.last_lts.extend(last_lts)\n        self.last_ans.extend(last_ans)\n\n    def add_ques(self, ques):\n        self.last_ques.append(ques)\n\n    def add_et(self, et):\n        self.last_ets.append(et)\n\n    def add_lt(self, lt):\n        self.last_lts.append(lt)\n\n    def add_ans(self, ans):\n        self.last_ans.append(ans)\n\n    def __repr__(self):\n        return \"_UserFeats: last_ques={}, last_ets={}, last_lts={}, last_ans={}\".format(\n            self.last_ques,\n            self.last_ets,\n            self.last_lts,\n            self.last_ans,\n        )\n```"
  },
  "source": "meta"
}