{
  "id": 202714,
  "title": "Has Anyone used Time Series Based Rolling and Lags Features ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202714",
  "author_name": "",
  "post_date": "2020-12-11T14:22:56.618427500Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Although there are very great kernels published on this competition , I have seen most of the people used average features based on content and user_id correctness ! Moreover as pointed out by a discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202086\" target=\"_blank\">here </a> , it is better to use loops for feature generation , Since Pandas is annoying for this tasks ! </p>\n<p>I wanted to ask has anyone considered using features like </p>\n<ol>\n<li>Average correctness based on Last 7 questions</li>\n<li>Lag Features Based on Answered_correctly </li>\n</ol>\n<p>Most of these features are extensively used in time series competitions , However using it requires using Pandas ! If anyone has used it please let me know in the comments !</p>",
  "messages": [
    {
      "id": "1109311",
      "postDate": "12/11/2020 14:22:56",
      "content": "<p>Although there are very great kernels published on this competition , I have seen most of the people used average features based on content and user_id correctness ! Moreover as pointed out by a discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202086\" target=\"_blank\">here </a> , it is better to use loops for feature generation , Since Pandas is annoying for this tasks ! </p>\n<p>I wanted to ask has anyone considered using features like </p>\n<ol>\n<li>Average correctness based on Last 7 questions</li>\n<li>Lag Features Based on Answered_correctly </li>\n</ol>\n<p>Most of these features are extensively used in time series competitions , However using it requires using Pandas ! If anyone has used it please let me know in the comments !</p>",
      "rawMarkdown": "Although there are very great kernels published on this competition , I have seen most of the people used average features based on content and user_id correctness ! Moreover as pointed out by a discussion [here ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202086) , it is better to use loops for feature generation , Since Pandas is annoying for this tasks ! \n\nI wanted to ask has anyone considered using features like \n1. Average correctness based on Last 7 questions\n2. Lag Features Based on Answered_correctly \n\nMost of these features are extensively used in time series competitions , However using it requires using Pandas ! If anyone has used it please let me know in the comments !",
      "votes": null
    },
    {
      "id": "1109373",
      "postDate": "12/11/2020 15:57:32",
      "content": "<p>I do use them, without pandas.</p>\n<p>The initial dataframe is sorted by user and timestamp. To limitate RAM usage, I have a dictionnary that cache the results I am computing for the current id.</p>\n<p>Among the features I have, I keep track in lists of the N lattest answers of the user. Once my list reach the size N, I pop at each iteration the first element etc… So I am sure my list takes only the N latest answers (or whatever over metric I want to track).</p>",
      "rawMarkdown": "I do use them, without pandas.\n\nThe initial dataframe is sorted by user and timestamp. To limitate RAM usage, I have a dictionnary that cache the results I am computing for the current id.\n\nAmong the features I have, I keep track in lists of the N lattest answers of the user. Once my list reach the size N, I pop at each iteration the first element etc... So I am sure my list takes only the N latest answers (or whatever over metric I want to track).",
      "votes": null
    },
    {
      "id": "1109390",
      "postDate": "12/11/2020 16:07:35",
      "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> that sounds interesting , I already improved my cv by good amount using One of Your Discussions ! Will Try this also ! </p>",
      "rawMarkdown": "bowaka that sounds interesting , I already improved my cv by good amount using One of Your Discussions ! Will Try this also !",
      "votes": null
    },
    {
      "id": "1109401",
      "postDate": "12/11/2020 16:23:43",
      "content": "<p>That's commonly referred to as a <a href=\"https://en.wikipedia.org/wiki/Circular_buffer\" target=\"_blank\">circular buffer</a>. In case you're using Python, you can use <a href=\"https://docs.python.org/3/library/collections.html#collections.deque\" target=\"_blank\"><code>collections.deque</code></a> for this.</p>",
      "rawMarkdown": "That's commonly referred to as a [circular buffer](https://en.wikipedia.org/wiki/Circular_buffer). In case you're using Python, you can use [`collections.deque`](https://docs.python.org/3/library/collections.html#collections.deque) for this.",
      "votes": null
    },
    {
      "id": "1110021",
      "postDate": "12/12/2020 10:48:07",
      "content": "<p>use lag questions mean correctness improve my cv score ,use dict and list  ,index = count%len(list).<br>\nuse deque may be better,I'll try</p>",
      "rawMarkdown": "use lag questions mean correctness improve my cv score ,use dict and list  ,index = count%len(list).\nuse deque may be better,I'll try",
      "votes": null
    },
    {
      "id": "1110023",
      "postDate": "12/12/2020 10:49:32",
      "content": "<p>You reminded me, thanks!</p>",
      "rawMarkdown": "You reminded me, thanks!",
      "votes": null
    },
    {
      "id": "1110066",
      "postDate": "12/12/2020 11:52:40",
      "content": "<p>Thanks will check that out !</p>",
      "rawMarkdown": "Thanks will check that out !",
      "votes": null
    },
    {
      "id": "1112106",
      "postDate": "12/14/2020 09:30:24",
      "content": "<p><a href=\"https://www.kaggle.com/qiaqia\" target=\"_blank\">@qiaqia</a> <br>\nare you talking about using this feature for LGBM or Transformers.<br>\nI doubt how  Transformer can  work for these feature during inference  .. although we do update test stats,but then Transformer is trained for fixed stats  for mean correctness..<br>\nSimilarly for Tree .. its branching is based on train average data but it changes during the test</p>",
      "rawMarkdown": "qiaqia \nare you talking about using this feature for LGBM or Transformers.\nI doubt how  Transformer can  work for these feature during inference  .. although we do update test stats,but then Transformer is trained for fixed stats  for mean correctness..\nSimilarly for Tree .. its branching is based on train average data but it changes during the test",
      "votes": null
    },
    {
      "id": "1112178",
      "postDate": "12/14/2020 10:54:33",
      "content": "<p>for LGBM,it's useful,update the lag_question dict during inference</p>",
      "rawMarkdown": "for LGBM,it's useful,update the lag_question dict during inference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1109373,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/11/2020 15:57:32",
      "content": "<p>I do use them, without pandas.</p>\n<p>The initial dataframe is sorted by user and timestamp. To limitate RAM usage, I have a dictionnary that cache the results I am computing for the current id.</p>\n<p>Among the features I have, I keep track in lists of the N lattest answers of the user. Once my list reach the size N, I pop at each iteration the first element etc… So I am sure my list takes only the N latest answers (or whatever over metric I want to track).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109401,
          "author_name": "christoffer",
          "author_url": "",
          "post_date": "12/11/2020 16:23:43",
          "content": "<p>That's commonly referred to as a <a href=\"https://en.wikipedia.org/wiki/Circular_buffer\" target=\"_blank\">circular buffer</a>. In case you're using Python, you can use <a href=\"https://docs.python.org/3/library/collections.html#collections.deque\" target=\"_blank\"><code>collections.deque</code></a> for this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1110023,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/12/2020 10:49:32",
          "content": "<p>You reminded me, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109390,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "12/11/2020 16:07:35",
      "content": "<p><a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> that sounds interesting , I already improved my cv by good amount using One of Your Discussions ! Will Try this also ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1110021,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/12/2020 10:48:07",
      "content": "<p>use lag questions mean correctness improve my cv score ,use dict and list  ,index = count%len(list).<br>\nuse deque may be better,I'll try</p>",
      "votes": null,
      "replies": [
        {
          "id": 1110066,
          "author_name": "sayedathar11",
          "author_url": "",
          "post_date": "12/12/2020 11:52:40",
          "content": "<p>Thanks will check that out !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112106,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/14/2020 09:30:24",
          "content": "<p><a href=\"https://www.kaggle.com/qiaqia\" target=\"_blank\">@qiaqia</a> <br>\nare you talking about using this feature for LGBM or Transformers.<br>\nI doubt how  Transformer can  work for these feature during inference  .. although we do update test stats,but then Transformer is trained for fixed stats  for mean correctness..<br>\nSimilarly for Tree .. its branching is based on train average data but it changes during the test</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1112178,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/14/2020 10:54:33",
          "content": "<p>for LGBM,it's useful,update the lag_question dict during inference</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1109311": "Although there are very great kernels published on this competition , I have seen most of the people used average features based on content and user_id correctness ! Moreover as pointed out by a discussion [here ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202086) , it is better to use loops for feature generation , Since Pandas is annoying for this tasks ! \n\nI wanted to ask has anyone considered using features like \n1. Average correctness based on Last 7 questions\n2. Lag Features Based on Answered_correctly \n\nMost of these features are extensively used in time series competitions , However using it requires using Pandas ! If anyone has used it please let me know in the comments !",
    "1109373": "I do use them, without pandas.\n\nThe initial dataframe is sorted by user and timestamp. To limitate RAM usage, I have a dictionnary that cache the results I am computing for the current id.\n\nAmong the features I have, I keep track in lists of the N lattest answers of the user. Once my list reach the size N, I pop at each iteration the first element etc... So I am sure my list takes only the N latest answers (or whatever over metric I want to track).",
    "1109390": "bowaka that sounds interesting , I already improved my cv by good amount using One of Your Discussions ! Will Try this also !",
    "1109401": "That's commonly referred to as a [circular buffer](https://en.wikipedia.org/wiki/Circular_buffer). In case you're using Python, you can use [`collections.deque`](https://docs.python.org/3/library/collections.html#collections.deque) for this.",
    "1110021": "use lag questions mean correctness improve my cv score ,use dict and list  ,index = count%len(list).\nuse deque may be better,I'll try",
    "1110023": "You reminded me, thanks!",
    "1110066": "Thanks will check that out !",
    "1112106": "qiaqia \nare you talking about using this feature for LGBM or Transformers.\nI doubt how  Transformer can  work for these feature during inference  .. although we do update test stats,but then Transformer is trained for fixed stats  for mean correctness..\nSimilarly for Tree .. its branching is based on train average data but it changes during the test",
    "1112178": "for LGBM,it's useful,update the lag_question dict during inference"
  },
  "source": "meta"
}