{
  "id": 206065,
  "title": "Why does \"number of iteration per second\" getting smaller as time goes?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206065",
  "author_name": "",
  "post_date": "2020-12-23T04:45:58.927800100Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I used <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">tito's simulator</a> (Thanks again!) and found out that as the iteration went, \"the number of iterations per second\"  that <code>tqdm</code> showed was getting smaller: first 400~600 iter/s in the first few iterations and 100~150 iter/s in the last iterations (you can check this out if you used <code>tqdm</code> and <code>iter_test</code> together)</p>\n<p>I guess this is because the caching dictionaries, such as <code>sum_dict</code>, <code>count_dict</code> are getting larger as the new data comes, but with respect to the time complexity of dictionary(hash table), I think it supposed to be the same(or almost same) asymptotically because find operation is O(1).</p>\n<p>What you guys think?</p>",
  "messages": [
    {
      "id": "1123254",
      "postDate": "12/23/2020 04:45:58",
      "content": "<p>I used <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">tito's simulator</a> (Thanks again!) and found out that as the iteration went, \"the number of iterations per second\"  that <code>tqdm</code> showed was getting smaller: first 400~600 iter/s in the first few iterations and 100~150 iter/s in the last iterations (you can check this out if you used <code>tqdm</code> and <code>iter_test</code> together)</p>\n<p>I guess this is because the caching dictionaries, such as <code>sum_dict</code>, <code>count_dict</code> are getting larger as the new data comes, but with respect to the time complexity of dictionary(hash table), I think it supposed to be the same(or almost same) asymptotically because find operation is O(1).</p>\n<p>What you guys think?</p>",
      "rawMarkdown": "I used [tito's simulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) (Thanks again!) and found out that as the iteration went, \"the number of iterations per second\"  that `tqdm` showed was getting smaller: first 400~600 iter/s in the first few iterations and 100~150 iter/s in the last iterations (you can check this out if you used `tqdm` and `iter_test` together)\n\nI guess this is because the caching dictionaries, such as `sum_dict`, `count_dict` are getting larger as the new data comes, but with respect to the time complexity of dictionary(hash table), I think it supposed to be the same(or almost same) asymptotically because find operation is O(1).\n\nWhat you guys think?",
      "votes": null
    },
    {
      "id": "1123820",
      "postDate": "12/23/2020 14:23:39",
      "content": "<p>Yes this might be the reason.<br>\nRather than storing the full history, you shall consider computing your features on the way.<br>\nFor example, to calculate the rolling mean, you just need 2 parameters to update at each step: count, and good_answers that is easy to update.<br>\nIf you want to calculate mean or sum over 5 last elements or 10 last element, you can as well considering poping the first element of the list every time your list is bigger than a given size:</p>\n<pre><code>if len(list)&gt;10:\n    list.pop(0)\n</code></pre>",
      "rawMarkdown": "Yes this might be the reason.\nRather than storing the full history, you shall consider computing your features on the way.\nFor example, to calculate the rolling mean, you just need 2 parameters to update at each step: count, and good_answers that is easy to update.\nIf you want to calculate mean or sum over 5 last elements or 10 last element, you can as well considering poping the first element of the list every time your list is bigger than a given size:\n```\nif len(list)>10:\n    list.pop(0)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123820,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/23/2020 14:23:39",
      "content": "<p>Yes this might be the reason.<br>\nRather than storing the full history, you shall consider computing your features on the way.<br>\nFor example, to calculate the rolling mean, you just need 2 parameters to update at each step: count, and good_answers that is easy to update.<br>\nIf you want to calculate mean or sum over 5 last elements or 10 last element, you can as well considering poping the first element of the list every time your list is bigger than a given size:</p>\n<pre><code>if len(list)&gt;10:\n    list.pop(0)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123254": "I used [tito's simulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) (Thanks again!) and found out that as the iteration went, \"the number of iterations per second\"  that `tqdm` showed was getting smaller: first 400~600 iter/s in the first few iterations and 100~150 iter/s in the last iterations (you can check this out if you used `tqdm` and `iter_test` together)\n\nI guess this is because the caching dictionaries, such as `sum_dict`, `count_dict` are getting larger as the new data comes, but with respect to the time complexity of dictionary(hash table), I think it supposed to be the same(or almost same) asymptotically because find operation is O(1).\n\nWhat you guys think?",
    "1123820": "Yes this might be the reason.\nRather than storing the full history, you shall consider computing your features on the way.\nFor example, to calculate the rolling mean, you just need 2 parameters to update at each step: count, and good_answers that is easy to update.\nIf you want to calculate mean or sum over 5 last elements or 10 last element, you can as well considering poping the first element of the list every time your list is bigger than a given size:\n```\nif len(list)>10:\n    list.pop(0)\n```"
  },
  "source": "meta"
}