{
  "id": 194053,
  "title": "How do you deal with StandardScaler taking up so much memory and crashing the notebook?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194053",
  "author_name": "",
  "post_date": "2020-10-30T12:00:41.514653Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The memory usage in my notebook before standard scaling starts is 1.5GB, then it exceeds 13GB and makes my notebook crash. Is there any technique to overcome this ? Note that here I'm only using a subset of the data, this means it would be impossible to normalize all of the data without some sort of parallelization.</p>",
  "messages": [
    {
      "id": "1064678",
      "postDate": "10/30/2020 12:00:41",
      "content": "<p>The memory usage in my notebook before standard scaling starts is 1.5GB, then it exceeds 13GB and makes my notebook crash. Is there any technique to overcome this ? Note that here I'm only using a subset of the data, this means it would be impossible to normalize all of the data without some sort of parallelization.</p>",
      "rawMarkdown": "The memory usage in my notebook before standard scaling starts is 1.5GB, then it exceeds 13GB and makes my notebook crash. Is there any technique to overcome this ? Note that here I'm only using a subset of the data, this means it would be impossible to normalize all of the data without some sort of parallelization.",
      "votes": null
    },
    {
      "id": "1064712",
      "postDate": "10/30/2020 12:50:11",
      "content": "<p>Are you doing it feature by feature?</p>",
      "rawMarkdown": "Are you doing it feature by feature?",
      "votes": null
    },
    {
      "id": "1064750",
      "postDate": "10/30/2020 13:39:48",
      "content": "<p>First, are you sure you need to scale (what type of model are you using?) Neural nets and some more traditional models might need it, but it shouldn't make any difference for tree-based models. If you definitely need to scale, a few thoughts--</p>\n<ol>\n<li><p>If you are using <code>sklearn</code> <code>StandardScaler.fit_transform()</code>, in my experience that is much more RAM intensive than calling <code>fit</code> and <code>transform</code> separately. </p></li>\n<li><p>Consider trying to process the data in even smaller amounts, either with the <code>partial_fit</code> method from <code>StandardScaler</code> to process more data batch by batch, or by just reducing the size of your subset. Even a relatively small % of the data is probably representative enough that scaling parameters learned on it will work ok, since the dataset is so large. </p></li>\n</ol>",
      "rawMarkdown": "First, are you sure you need to scale (what type of model are you using?) Neural nets and some more traditional models might need it, but it shouldn't make any difference for tree-based models. If you definitely need to scale, a few thoughts--\n\n1. If you are using `sklearn` `StandardScaler.fit_transform()`, in my experience that is much more RAM intensive than calling `fit` and `transform` separately. \n\n2. Consider trying to process the data in even smaller amounts, either with the `partial_fit` method from `StandardScaler` to process more data batch by batch, or by just reducing the size of your subset. Even a relatively small % of the data is probably representative enough that scaling parameters learned on it will work ok, since the dataset is so large.",
      "votes": null
    },
    {
      "id": "1064866",
      "postDate": "10/30/2020 15:55:52",
      "content": "<p>I stumble upon the same problems as you when using GLM (since I want to use it as feature selection). To tackle this I make my own scaler, something like this</p>\n<pre><code>class MinMaxScaler :\n    def __init__(self):\n        self.statistics = {}\n    def fit(self,df) :\n        for col in df.columns :\n            self.statistics[col] = {'max': np.max(list(df[col])) ,'min':np.min(list(df[col]))}\n    def transform(self,df):\n        for col in df.columns :\n            df[col] = (df[col] - self.statistics[col]['min']) / (self.statistics[col]['max'] - self.statistics[col]['min'])\n        return df\n</code></pre>",
      "rawMarkdown": "I stumble upon the same problems as you when using GLM (since I want to use it as feature selection). To tackle this I make my own scaler, something like this\n\n```\nclass MinMaxScaler :\n    def __init__(self):\n        self.statistics = {}\n    def fit(self,df) :\n        for col in df.columns :\n            self.statistics[col] = {'max': np.max(list(df[col])) ,'min':np.min(list(df[col]))}\n    def transform(self,df):\n        for col in df.columns :\n            df[col] = (df[col] - self.statistics[col]['min']) / (self.statistics[col]['max'] - self.statistics[col]['min'])\n        return df\n```",
      "votes": null
    },
    {
      "id": "1064993",
      "postDate": "10/30/2020 18:20:41",
      "content": "<p>This is pretty helpful, thank you!</p>",
      "rawMarkdown": "This is pretty helpful, thank you!",
      "votes": null
    },
    {
      "id": "1064995",
      "postDate": "10/30/2020 18:21:14",
      "content": "<p>I will test that now, thank you for your response</p>",
      "rawMarkdown": "I will test that now, thank you for your response",
      "votes": null
    },
    {
      "id": "1064996",
      "postDate": "10/30/2020 18:21:42",
      "content": "<p>No, the whole dataframe at once. Probably it be less memory intensive if its feature by feature.</p>",
      "rawMarkdown": "No, the whole dataframe at once. Probably it be less memory intensive if its feature by feature.",
      "votes": null
    },
    {
      "id": "1065019",
      "postDate": "10/30/2020 18:46:01",
      "content": "<p>`</p>\n<pre><code>def normalize_features(df):\n\n    for column in df.columns:\n\n        df[column] = (df[column] - df[column].mean())/df[column].std()\n\n    return df\n</code></pre>\n<p>`</p>\n<p>Try this. It worked for me</p>",
      "rawMarkdown": "`\n\n    def normalize_features(df):\n\n        for column in df.columns:\n\n            df[column] = (df[column] - df[column].mean())/df[column].std()\n\n        return df\n`\n\nTry this. It worked for me",
      "votes": null
    },
    {
      "id": "1065728",
      "postDate": "10/31/2020 16:44:49",
      "content": "<p>I had this problem and ended up just using MInMaxScaler from sklearn. But if you're using  LightGBM you don't need it, I ended up not using it and the model's performance didn't change anything. Just use the normal lgb.Dataset</p>",
      "rawMarkdown": "I had this problem and ended up just using MInMaxScaler from sklearn. But if you're using  LightGBM you don't need it, I ended up not using it and the model's performance didn't change anything. Just use the normal lgb.Dataset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1064712,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/30/2020 12:50:11",
      "content": "<p>Are you doing it feature by feature?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1064996,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "10/30/2020 18:21:42",
          "content": "<p>No, the whole dataframe at once. Probably it be less memory intensive if its feature by feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065019,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/30/2020 18:46:01",
          "content": "<p>`</p>\n<pre><code>def normalize_features(df):\n\n    for column in df.columns:\n\n        df[column] = (df[column] - df[column].mean())/df[column].std()\n\n    return df\n</code></pre>\n<p>`</p>\n<p>Try this. It worked for me</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1064750,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/30/2020 13:39:48",
      "content": "<p>First, are you sure you need to scale (what type of model are you using?) Neural nets and some more traditional models might need it, but it shouldn't make any difference for tree-based models. If you definitely need to scale, a few thoughts--</p>\n<ol>\n<li><p>If you are using <code>sklearn</code> <code>StandardScaler.fit_transform()</code>, in my experience that is much more RAM intensive than calling <code>fit</code> and <code>transform</code> separately. </p></li>\n<li><p>Consider trying to process the data in even smaller amounts, either with the <code>partial_fit</code> method from <code>StandardScaler</code> to process more data batch by batch, or by just reducing the size of your subset. Even a relatively small % of the data is probably representative enough that scaling parameters learned on it will work ok, since the dataset is so large. </p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1064995,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "10/30/2020 18:21:14",
          "content": "<p>I will test that now, thank you for your response</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1064866,
      "author_name": "marcellosusanto",
      "author_url": "",
      "post_date": "10/30/2020 15:55:52",
      "content": "<p>I stumble upon the same problems as you when using GLM (since I want to use it as feature selection). To tackle this I make my own scaler, something like this</p>\n<pre><code>class MinMaxScaler :\n    def __init__(self):\n        self.statistics = {}\n    def fit(self,df) :\n        for col in df.columns :\n            self.statistics[col] = {'max': np.max(list(df[col])) ,'min':np.min(list(df[col]))}\n    def transform(self,df):\n        for col in df.columns :\n            df[col] = (df[col] - self.statistics[col]['min']) / (self.statistics[col]['max'] - self.statistics[col]['min'])\n        return df\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1064993,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "10/30/2020 18:20:41",
          "content": "<p>This is pretty helpful, thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065728,
      "author_name": "iuryck",
      "author_url": "",
      "post_date": "10/31/2020 16:44:49",
      "content": "<p>I had this problem and ended up just using MInMaxScaler from sklearn. But if you're using  LightGBM you don't need it, I ended up not using it and the model's performance didn't change anything. Just use the normal lgb.Dataset</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1064678": "The memory usage in my notebook before standard scaling starts is 1.5GB, then it exceeds 13GB and makes my notebook crash. Is there any technique to overcome this ? Note that here I'm only using a subset of the data, this means it would be impossible to normalize all of the data without some sort of parallelization.",
    "1064712": "Are you doing it feature by feature?",
    "1064750": "First, are you sure you need to scale (what type of model are you using?) Neural nets and some more traditional models might need it, but it shouldn't make any difference for tree-based models. If you definitely need to scale, a few thoughts--\n\n1. If you are using `sklearn` `StandardScaler.fit_transform()`, in my experience that is much more RAM intensive than calling `fit` and `transform` separately. \n\n2. Consider trying to process the data in even smaller amounts, either with the `partial_fit` method from `StandardScaler` to process more data batch by batch, or by just reducing the size of your subset. Even a relatively small % of the data is probably representative enough that scaling parameters learned on it will work ok, since the dataset is so large.",
    "1064866": "I stumble upon the same problems as you when using GLM (since I want to use it as feature selection). To tackle this I make my own scaler, something like this\n\n```\nclass MinMaxScaler :\n    def __init__(self):\n        self.statistics = {}\n    def fit(self,df) :\n        for col in df.columns :\n            self.statistics[col] = {'max': np.max(list(df[col])) ,'min':np.min(list(df[col]))}\n    def transform(self,df):\n        for col in df.columns :\n            df[col] = (df[col] - self.statistics[col]['min']) / (self.statistics[col]['max'] - self.statistics[col]['min'])\n        return df\n```",
    "1064993": "This is pretty helpful, thank you!",
    "1064995": "I will test that now, thank you for your response",
    "1064996": "No, the whole dataframe at once. Probably it be less memory intensive if its feature by feature.",
    "1065019": "`\n\n    def normalize_features(df):\n\n        for column in df.columns:\n\n            df[column] = (df[column] - df[column].mean())/df[column].std()\n\n        return df\n`\n\nTry this. It worked for me",
    "1065728": "I had this problem and ended up just using MInMaxScaler from sklearn. But if you're using  LightGBM you don't need it, I ended up not using it and the model's performance didn't change anything. Just use the normal lgb.Dataset"
  },
  "source": "meta"
}