{
  "id": 55379,
  "title": "Save more Memory in Pandas",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55379",
  "author_name": "Silverback",
  "post_date": "2018-04-25T21:24:11.071000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>We can save more memory in many of the public kernels by encoding <code>is_attributed</code> and <code>click_id</code> as <code>uint</code>. Here is an example where I have used the maximum values of their respective <code>uint</code> to encode missing or nan values.</p>\n\n<pre><code>train_df['is_attributed'] = train_df.is_attributed.fillna(2**8 - 1).astype('uint8')\ntrain_df['click_id'] = train_df.click_id.fillna(2**32 - 1).astype('uint32')\n</code></pre>\n\n<p>This trick saved me 0.6 gb on 16.0 gb RAM which was equivalent to 1.2 gb and 32.0 gb in total (swap + mem) that I need to run some of the leading public kernels.</p>\n\n<p><strong>Edit</strong>: I should mention, the filling of nans allows us to save memory, as, if a nan is present, Pandas has to assume that column is of dtype <code>float64</code>. By filling this with a large positive number, we can encode these nans AND save memory.</p>",
  "messages": [
    {
      "id": 319359,
      "postDate": "2018-04-25T21:24:11.070Z",
      "content": "<p>We can save more memory in many of the public kernels by encoding <code>is_attributed</code> and <code>click_id</code> as <code>uint</code>. Here is an example where I have used the maximum values of their respective <code>uint</code> to encode missing or nan values.</p>\n\n<pre><code>train_df['is_attributed'] = train_df.is_attributed.fillna(2**8 - 1).astype('uint8')\ntrain_df['click_id'] = train_df.click_id.fillna(2**32 - 1).astype('uint32')\n</code></pre>\n\n<p>This trick saved me 0.6 gb on 16.0 gb RAM which was equivalent to 1.2 gb and 32.0 gb in total (swap + mem) that I need to run some of the leading public kernels.</p>\n\n<p><strong>Edit</strong>: I should mention, the filling of nans allows us to save memory, as, if a nan is present, Pandas has to assume that column is of dtype <code>float64</code>. By filling this with a large positive number, we can encode these nans AND save memory.</p>",
      "rawMarkdown": "We can save more memory in many of the public kernels by encoding `is_attributed` and `click_id` as `uint`. Here is an example where I have used the maximum values of their respective `uint` to encode missing or nan values.\n\n    train_df['is_attributed'] = train_df.is_attributed.fillna(2**8 - 1).astype('uint8')\n    train_df['click_id'] = train_df.click_id.fillna(2**32 - 1).astype('uint32')\n\nThis trick saved me 0.6 gb on 16.0 gb RAM which was equivalent to 1.2 gb and 32.0 gb in total (swap + mem) that I need to run some of the leading public kernels.\n\n**Edit**: I should mention, the filling of nans allows us to save memory, as, if a nan is present, Pandas has to assume that column is of dtype `float64`. By filling this with a large positive number, we can encode these nans AND save memory.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "319359": "We can save more memory in many of the public kernels by encoding `is_attributed` and `click_id` as `uint`. Here is an example where I have used the maximum values of their respective `uint` to encode missing or nan values.\n\n    train_df['is_attributed'] = train_df.is_attributed.fillna(2**8 - 1).astype('uint8')\n    train_df['click_id'] = train_df.click_id.fillna(2**32 - 1).astype('uint32')\n\nThis trick saved me 0.6 gb on 16.0 gb RAM which was equivalent to 1.2 gb and 32.0 gb in total (swap + mem) that I need to run some of the leading public kernels.\n\n**Edit**: I should mention, the filling of nans allows us to save memory, as, if a nan is present, Pandas has to assume that column is of dtype `float64`. By filling this with a large positive number, we can encode these nans AND save memory."
  }
}