{
  "id": 194710,
  "title": "Strategy on each test set iteration?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194710",
  "author_name": "",
  "post_date": "2020-11-02T12:03:43.986122300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I wonder whether I'm thinking in the right way or not:</p>\n<p>Let say I generated some mean encoded features in the train data set and build a model.</p>\n<p>When I submit the codes, it will iterate over the test set. Each chunk will be group0, group1, …</p>\n<p>Then, should I prepare the code like this?</p>\n<pre><code>// test set 0 (Group 0) comes out\n// generate encoded features using (train + test set 0)\n// train a new model based on them\n// Inference on test set 0\n\n// test set 1 (Group 1) comes out\n// generate encoded features using (train + test set 0 + test set 1)\n// train a new model based on them\n// Inference on test set 1\n\n....\n\n// test set k (Group k) comes out\n// generate encoded features using (train + test set 0 + test set 1 + .... + test set k)\n// train a new model based on them\n// Inference on test set k\n</code></pre>\n<p><strong>The reason I added the test set on total dataset gradually is that we can know the ground truth value(answer, response) from <code>prior_group_answerers_correct</code> and <code>prior_group_responses</code>.</strong></p>\n<p>Is it the right strategy? or will it be a very heavy load?</p>",
  "messages": [
    {
      "id": "1067225",
      "postDate": "11/02/2020 12:03:43",
      "content": "<p>I wonder whether I'm thinking in the right way or not:</p>\n<p>Let say I generated some mean encoded features in the train data set and build a model.</p>\n<p>When I submit the codes, it will iterate over the test set. Each chunk will be group0, group1, …</p>\n<p>Then, should I prepare the code like this?</p>\n<pre><code>// test set 0 (Group 0) comes out\n// generate encoded features using (train + test set 0)\n// train a new model based on them\n// Inference on test set 0\n\n// test set 1 (Group 1) comes out\n// generate encoded features using (train + test set 0 + test set 1)\n// train a new model based on them\n// Inference on test set 1\n\n....\n\n// test set k (Group k) comes out\n// generate encoded features using (train + test set 0 + test set 1 + .... + test set k)\n// train a new model based on them\n// Inference on test set k\n</code></pre>\n<p><strong>The reason I added the test set on total dataset gradually is that we can know the ground truth value(answer, response) from <code>prior_group_answerers_correct</code> and <code>prior_group_responses</code>.</strong></p>\n<p>Is it the right strategy? or will it be a very heavy load?</p>",
      "rawMarkdown": "I wonder whether I'm thinking in the right way or not:\n\nLet say I generated some mean encoded features in the train data set and build a model.\n\nWhen I submit the codes, it will iterate over the test set. Each chunk will be group0, group1, ...\n\nThen, should I prepare the code like this?\n\n```\n// test set 0 (Group 0) comes out\n// generate encoded features using (train + test set 0)\n// train a new model based on them\n// Inference on test set 0\n\n// test set 1 (Group 1) comes out\n// generate encoded features using (train + test set 0 + test set 1)\n// train a new model based on them\n// Inference on test set 1\n\n....\n\n// test set k (Group k) comes out\n// generate encoded features using (train + test set 0 + test set 1 + .... + test set k)\n// train a new model based on them\n// Inference on test set k\n\n```\n\n**The reason I added the test set on total dataset gradually is that we can know the ground truth value(answer, response) from `prior_group_answerers_correct` and `prior_group_responses`.**\n\nIs it the right strategy? or will it be a very heavy load?",
      "votes": null
    },
    {
      "id": "1067406",
      "postDate": "11/02/2020 14:13:45",
      "content": "<p>Indeed, you have to generate features again each turn because there could be new users and aggregating new users' information is important. But I don't think you need to retrain the model because it wouldn't change much from adding just a few test cases in each group. Also, it would make predictions very slow.</p>",
      "rawMarkdown": "Indeed, you have to generate features again each turn because there could be new users and aggregating new users' information is important. But I don't think you need to retrain the model because it wouldn't change much from adding just a few test cases in each group. Also, it would make predictions very slow.",
      "votes": null
    },
    {
      "id": "1067624",
      "postDate": "11/02/2020 15:53:34",
      "content": "<p>Maybe you can keep a buffer df of k groups and retrain every k iterations. Where k can be tuned. That maybe more efficient. </p>",
      "rawMarkdown": "Maybe you can keep a buffer df of k groups and retrain every k iterations. Where k can be tuned. That maybe more efficient.",
      "votes": null
    },
    {
      "id": "1067766",
      "postDate": "11/02/2020 17:40:30",
      "content": "<p>It's more efficient and it saves us the time as well. (I saw a difference of around ~0.05 if i don't do so) We can use <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">this</a> which i wrote, works like a charm for me.</p>",
      "rawMarkdown": "It's more efficient and it saves us the time as well. (I saw a difference of around ~0.05 if i don't do so) We can use [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856) which i wrote, works like a charm for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1067406,
      "author_name": "mateuscco",
      "author_url": "",
      "post_date": "11/02/2020 14:13:45",
      "content": "<p>Indeed, you have to generate features again each turn because there could be new users and aggregating new users' information is important. But I don't think you need to retrain the model because it wouldn't change much from adding just a few test cases in each group. Also, it would make predictions very slow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1067624,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/02/2020 15:53:34",
          "content": "<p>Maybe you can keep a buffer df of k groups and retrain every k iterations. Where k can be tuned. That maybe more efficient. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1067766,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/02/2020 17:40:30",
          "content": "<p>It's more efficient and it saves us the time as well. (I saw a difference of around ~0.05 if i don't do so) We can use <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">this</a> which i wrote, works like a charm for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1067225": "I wonder whether I'm thinking in the right way or not:\n\nLet say I generated some mean encoded features in the train data set and build a model.\n\nWhen I submit the codes, it will iterate over the test set. Each chunk will be group0, group1, ...\n\nThen, should I prepare the code like this?\n\n```\n// test set 0 (Group 0) comes out\n// generate encoded features using (train + test set 0)\n// train a new model based on them\n// Inference on test set 0\n\n// test set 1 (Group 1) comes out\n// generate encoded features using (train + test set 0 + test set 1)\n// train a new model based on them\n// Inference on test set 1\n\n....\n\n// test set k (Group k) comes out\n// generate encoded features using (train + test set 0 + test set 1 + .... + test set k)\n// train a new model based on them\n// Inference on test set k\n\n```\n\n**The reason I added the test set on total dataset gradually is that we can know the ground truth value(answer, response) from `prior_group_answerers_correct` and `prior_group_responses`.**\n\nIs it the right strategy? or will it be a very heavy load?",
    "1067406": "Indeed, you have to generate features again each turn because there could be new users and aggregating new users' information is important. But I don't think you need to retrain the model because it wouldn't change much from adding just a few test cases in each group. Also, it would make predictions very slow.",
    "1067624": "Maybe you can keep a buffer df of k groups and retrain every k iterations. Where k can be tuned. That maybe more efficient.",
    "1067766": "It's more efficient and it saves us the time as well. (I saw a difference of around ~0.05 if i don't do so) We can use [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856) which i wrote, works like a charm for me."
  },
  "source": "meta"
}