{
  "id": 208922,
  "title": "Which features of train_data are used by the test API for scoring my model",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208922",
  "author_name": "",
  "post_date": "2021-01-05T15:28:54.203903300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>If one is free to choose whatever set of features from the train data to train his/her model, then what happens during scoring of the model by the test API and the leaderboard? Do they use same features as the ones I trained the model on to make predictions? Because it seems to me the predictions would be inaccurate if my model is tested with different input features different from the one I chose to train the model on. Please I need clarification. Thanks!</p>",
  "messages": [
    {
      "id": "1139731",
      "postDate": "01/05/2021 15:28:54",
      "content": "<p>If one is free to choose whatever set of features from the train data to train his/her model, then what happens during scoring of the model by the test API and the leaderboard? Do they use same features as the ones I trained the model on to make predictions? Because it seems to me the predictions would be inaccurate if my model is tested with different input features different from the one I chose to train the model on. Please I need clarification. Thanks!</p>",
      "rawMarkdown": "If one is free to choose whatever set of features from the train data to train his/her model, then what happens during scoring of the model by the test API and the leaderboard? Do they use same features as the ones I trained the model on to make predictions? Because it seems to me the predictions would be inaccurate if my model is tested with different input features different from the one I chose to train the model on. Please I need clarification. Thanks!",
      "votes": null
    },
    {
      "id": "1139913",
      "postDate": "01/05/2021 17:38:27",
      "content": "<p>The same set of features are available to you during the scoring of your model. It is on you however to filter/manipulate them to get the data into the same format as to what you trained on.</p>",
      "rawMarkdown": "The same set of features are available to you during the scoring of your model. It is on you however to filter/manipulate them to get the data into the same format as to what you trained on.",
      "votes": null
    },
    {
      "id": "1140049",
      "postDate": "01/05/2021 19:13:13",
      "content": "<p>Thanks for your informative response. But how do I access the test data? Is it through the iter_test function provided in the submission guideline?</p>",
      "rawMarkdown": "Thanks for your informative response. But how do I access the test data? Is it through the iter_test function provided in the submission guideline?",
      "votes": null
    },
    {
      "id": "1140124",
      "postDate": "01/05/2021 20:15:14",
      "content": "<p>Yes you loop over it like below, the test_data is in test_df.<br>\nAfter you have made the predictions with your model you extract row_id and answered_correctly (your model produces this) then plug into env.predict()</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\n\n# You can only iterate through a result from `env.iter_test()` once\n# so be careful not to lose it once you start iterating.\niter_test = env.iter_test()\n\nprior_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n\n    pred_df = test_df[['row_id', 'answered_correctly']]\n    env.predict(pred_df)\n</code></pre>",
      "rawMarkdown": "Yes you loop over it like below, the test_data is in test_df.\nAfter you have made the predictions with your model you extract row_id and answered_correctly (your model produces this) then plug into env.predict()\n\n```\nimport riiideducation\nenv = riiideducation.make_env()\n\n# You can only iterate through a result from `env.iter_test()` once\n# so be careful not to lose it once you start iterating.\niter_test = env.iter_test()\n\nprior_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n\n    pred_df = test_df[['row_id', 'answered_correctly']]\n    env.predict(pred_df)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1139913,
      "author_name": "darren1515",
      "author_url": "",
      "post_date": "01/05/2021 17:38:27",
      "content": "<p>The same set of features are available to you during the scoring of your model. It is on you however to filter/manipulate them to get the data into the same format as to what you trained on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140049,
          "author_name": "nawasnaziru",
          "author_url": "",
          "post_date": "01/05/2021 19:13:13",
          "content": "<p>Thanks for your informative response. But how do I access the test data? Is it through the iter_test function provided in the submission guideline?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140124,
          "author_name": "darren1515",
          "author_url": "",
          "post_date": "01/05/2021 20:15:14",
          "content": "<p>Yes you loop over it like below, the test_data is in test_df.<br>\nAfter you have made the predictions with your model you extract row_id and answered_correctly (your model produces this) then plug into env.predict()</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\n\n# You can only iterate through a result from `env.iter_test()` once\n# so be careful not to lose it once you start iterating.\niter_test = env.iter_test()\n\nprior_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n\n    pred_df = test_df[['row_id', 'answered_correctly']]\n    env.predict(pred_df)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1139731": "If one is free to choose whatever set of features from the train data to train his/her model, then what happens during scoring of the model by the test API and the leaderboard? Do they use same features as the ones I trained the model on to make predictions? Because it seems to me the predictions would be inaccurate if my model is tested with different input features different from the one I chose to train the model on. Please I need clarification. Thanks!",
    "1139913": "The same set of features are available to you during the scoring of your model. It is on you however to filter/manipulate them to get the data into the same format as to what you trained on.",
    "1140049": "Thanks for your informative response. But how do I access the test data? Is it through the iter_test function provided in the submission guideline?",
    "1140124": "Yes you loop over it like below, the test_data is in test_df.\nAfter you have made the predictions with your model you extract row_id and answered_correctly (your model produces this) then plug into env.predict()\n\n```\nimport riiideducation\nenv = riiideducation.make_env()\n\n# You can only iterate through a result from `env.iter_test()` once\n# so be careful not to lose it once you start iterating.\niter_test = env.iter_test()\n\nprior_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n\n    pred_df = test_df[['row_id', 'answered_correctly']]\n    env.predict(pred_df)\n```"
  },
  "source": "meta"
}