{
  "id": 194019,
  "title": "Allowance Calculating \"Current Elapsed Time\"",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194019",
  "author_name": "",
  "post_date": "2020-10-30T08:33:42.935608900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>I'm just curious whether it is allowed (and reasonable) to derive the time a user needs to answer a question from <code>prior_elapsed_time</code> (regardless of how someone would approach this via shifting, grouping by <code>user_id</code>/<code>bundle_id</code> etc.)</p>\n<p>I just think that this should not be allowed since let's say <code>current_elapsed_time</code> is a feature that is only observed posterior answering a question. Of course you can calculate it to use it for training your model, but at least the test batches should be organized in a way that someone can not calculate elapsed time for a question (e.g. by ensuring that each test batch only contains different users) and therefore you have to replace these missing values.</p>\n<p>Does anyone know if it's allowed to calculate this feature or has different thoughts on this?</p>",
  "messages": [
    {
      "id": "1064546",
      "postDate": "10/30/2020 08:33:42",
      "content": "<p>Hi everyone, </p>\n<p>I'm just curious whether it is allowed (and reasonable) to derive the time a user needs to answer a question from <code>prior_elapsed_time</code> (regardless of how someone would approach this via shifting, grouping by <code>user_id</code>/<code>bundle_id</code> etc.)</p>\n<p>I just think that this should not be allowed since let's say <code>current_elapsed_time</code> is a feature that is only observed posterior answering a question. Of course you can calculate it to use it for training your model, but at least the test batches should be organized in a way that someone can not calculate elapsed time for a question (e.g. by ensuring that each test batch only contains different users) and therefore you have to replace these missing values.</p>\n<p>Does anyone know if it's allowed to calculate this feature or has different thoughts on this?</p>",
      "rawMarkdown": "Hi everyone, \n\nI'm just curious whether it is allowed (and reasonable) to derive the time a user needs to answer a question from `prior_elapsed_time` (regardless of how someone would approach this via shifting, grouping by `user_id`/`bundle_id` etc.)\n\nI just think that this should not be allowed since let's say `current_elapsed_time` is a feature that is only observed posterior answering a question. Of course you can calculate it to use it for training your model, but at least the test batches should be organized in a way that someone can not calculate elapsed time for a question (e.g. by ensuring that each test batch only contains different users) and therefore you have to replace these missing values.\n\nDoes anyone know if it's allowed to calculate this feature or has different thoughts on this?",
      "votes": null
    },
    {
      "id": "1064555",
      "postDate": "10/30/2020 08:42:33",
      "content": "<p>Hi Marcus,</p>\n<p>Using  current_elapsed_time  as a feature makes sense in sequence models since the prediction will depend on past interactions. It should be ensured that there is no future peeking during training. Otherwise the model will just fail spectacularly on test data. Also, at predict time the current sample cannot contain this feature. </p>",
      "rawMarkdown": "Hi Marcus,\n\nUsing  current_elapsed_time  as a feature makes sense in sequence models since the prediction will depend on past interactions. It should be ensured that there is no future peeking during training. Otherwise the model will just fail spectacularly on test data. Also, at predict time the current sample cannot contain this feature.",
      "votes": null
    },
    {
      "id": "1064557",
      "postDate": "10/30/2020 08:46:19",
      "content": "<p>From the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">data description</a>:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but <strong>no more than one task_container_id of questions from any single user</strong>. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>So it will not be possible to calculate <code>current_elapsed_time</code> in the test data.</p>",
      "rawMarkdown": "From the [data description](https://www.kaggle.com/c/riiid-test-answer-prediction/data):\n> The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but **no more than one task_container_id of questions from any single user**. Each group has between 1 and 1000 users.\n\nSo it will not be possible to calculate `current_elapsed_time` in the test data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1064555,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/30/2020 08:42:33",
      "content": "<p>Hi Marcus,</p>\n<p>Using  current_elapsed_time  as a feature makes sense in sequence models since the prediction will depend on past interactions. It should be ensured that there is no future peeking during training. Otherwise the model will just fail spectacularly on test data. Also, at predict time the current sample cannot contain this feature. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1064557,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/30/2020 08:46:19",
      "content": "<p>From the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">data description</a>:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but <strong>no more than one task_container_id of questions from any single user</strong>. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>So it will not be possible to calculate <code>current_elapsed_time</code> in the test data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1064546": "Hi everyone, \n\nI'm just curious whether it is allowed (and reasonable) to derive the time a user needs to answer a question from `prior_elapsed_time` (regardless of how someone would approach this via shifting, grouping by `user_id`/`bundle_id` etc.)\n\nI just think that this should not be allowed since let's say `current_elapsed_time` is a feature that is only observed posterior answering a question. Of course you can calculate it to use it for training your model, but at least the test batches should be organized in a way that someone can not calculate elapsed time for a question (e.g. by ensuring that each test batch only contains different users) and therefore you have to replace these missing values.\n\nDoes anyone know if it's allowed to calculate this feature or has different thoughts on this?",
    "1064555": "Hi Marcus,\n\nUsing  current_elapsed_time  as a feature makes sense in sequence models since the prediction will depend on past interactions. It should be ensured that there is no future peeking during training. Otherwise the model will just fail spectacularly on test data. Also, at predict time the current sample cannot contain this feature.",
    "1064557": "From the [data description](https://www.kaggle.com/c/riiid-test-answer-prediction/data):\n> The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but **no more than one task_container_id of questions from any single user**. Each group has between 1 and 1000 users.\n\nSo it will not be possible to calculate `current_elapsed_time` in the test data."
  },
  "source": "meta"
}