{
  "id": 196367,
  "title": "Scaling data in env.predict loop",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196367",
  "author_name": "",
  "post_date": "2020-11-10T22:12:32.734754400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am currently using the sklearn StandardScaler with my training set.<br>\nHow would I scale the features with this in the env.predict() loop?</p>",
  "messages": [
    {
      "id": "1074636",
      "postDate": "11/10/2020 22:12:32",
      "content": "<p>I am currently using the sklearn StandardScaler with my training set.<br>\nHow would I scale the features with this in the env.predict() loop?</p>",
      "rawMarkdown": "I am currently using the sklearn StandardScaler with my training set.\nHow would I scale the features with this in the env.predict() loop?",
      "votes": null
    },
    {
      "id": "1074661",
      "postDate": "11/10/2020 22:55:30",
      "content": "<pre><code>my_encoder = StandardScaler() \nmy_encoder.fit(X_train)\nmy_encoder.transform(X_train)\n\n# train my great algorithm\n\nmy_encoder.transform(X_test)\npredict()\n</code></pre>\n<p>dump your encoder to pickle if needed. </p>",
      "rawMarkdown": "```\nmy_encoder = StandardScaler() \nmy_encoder.fit(X_train)\nmy_encoder.transform(X_train)\n\n# train my great algorithm\n\nmy_encoder.transform(X_test)\npredict()\n```\n\ndump your encoder to pickle if needed.",
      "votes": null
    },
    {
      "id": "1074676",
      "postDate": "11/10/2020 23:41:39",
      "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/jacquespeeters\" target=\"_blank\">@jacquespeeters</a>!</p>\n<p>I understand that this is how I would usually do the test set scaling, however I'm curious how this fits in with the:</p>\n<p>env.predict(test_df.loc[test_df['content_type_id'] == 0,<br>\n                           ['row_id', 'answered_correctly']])</p>\n<p>submission loop, should I wrap this with the transform call? such as:</p>\n<p>env.predict(scaler.transform(test_df.loc[test_df['content_type_id']… </p>",
      "rawMarkdown": "Thanks for the answer @jacquespeeters!\n\nI understand that this is how I would usually do the test set scaling, however I'm curious how this fits in with the:\n\nenv.predict(test_df.loc[test_df['content_type_id'] == 0,\n                           ['row_id', 'answered_correctly']])\n\nsubmission loop, should I wrap this with the transform call? such as:\n\nenv.predict(scaler.transform(test_df.loc[test_df['content_type_id']...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1074661,
      "author_name": "jacquespeeters",
      "author_url": "",
      "post_date": "11/10/2020 22:55:30",
      "content": "<pre><code>my_encoder = StandardScaler() \nmy_encoder.fit(X_train)\nmy_encoder.transform(X_train)\n\n# train my great algorithm\n\nmy_encoder.transform(X_test)\npredict()\n</code></pre>\n<p>dump your encoder to pickle if needed. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1074676,
      "author_name": "cshorten30",
      "author_url": "",
      "post_date": "11/10/2020 23:41:39",
      "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/jacquespeeters\" target=\"_blank\">@jacquespeeters</a>!</p>\n<p>I understand that this is how I would usually do the test set scaling, however I'm curious how this fits in with the:</p>\n<p>env.predict(test_df.loc[test_df['content_type_id'] == 0,<br>\n                           ['row_id', 'answered_correctly']])</p>\n<p>submission loop, should I wrap this with the transform call? such as:</p>\n<p>env.predict(scaler.transform(test_df.loc[test_df['content_type_id']… </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1074636": "I am currently using the sklearn StandardScaler with my training set.\nHow would I scale the features with this in the env.predict() loop?",
    "1074661": "```\nmy_encoder = StandardScaler() \nmy_encoder.fit(X_train)\nmy_encoder.transform(X_train)\n\n# train my great algorithm\n\nmy_encoder.transform(X_test)\npredict()\n```\n\ndump your encoder to pickle if needed.",
    "1074676": "Thanks for the answer @jacquespeeters!\n\nI understand that this is how I would usually do the test set scaling, however I'm curious how this fits in with the:\n\nenv.predict(test_df.loc[test_df['content_type_id'] == 0,\n                           ['row_id', 'answered_correctly']])\n\nsubmission loop, should I wrap this with the transform call? such as:\n\nenv.predict(scaler.transform(test_df.loc[test_df['content_type_id']..."
  },
  "source": "meta"
}