{
  "id": 332645,
  "title": "Data Size Reducing🏋 + feather 🪶format data",
  "url": "/competitions/amex-default-prediction/discussion/332645",
  "author_name": "",
  "post_date": "2022-06-22T17:20:00.116539400Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Everyone,<br>\nThis Competition is challenging because of the data size that is difficult to handle.<br>\nas mentioned <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in his very insightful topic here Before discussion about the format to apply to our dataset we have to optimize every column type so that it would take up as little memory space as possible.</p>\n<p>I propose this code sample that a friend (@mathurinache) gave me once:</p>\n<pre><code>def reduce_mem_usage(df, verbose=True):\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage before optimization is: {:.2f} MB'.format(start_mem))\n    for col in df.columns:\n        col_type = df[col].dtypes\n        if col_type in numerics:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min &gt; np.iinfo(np.int8).min and c_max &lt; np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min &gt; np.iinfo(np.int16).min and c_max &lt; np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min &gt; np.iinfo(np.int32).min and c_max &lt; np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min &gt; np.iinfo(np.int64).min and c_max &lt; np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min &gt; np.finfo(np.float16).min and c_max &lt; np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n    return df\n</code></pre>\n<p>I have created Feather format datasets for the initial datasets<br>\n(the final train dataset is merged with the labels)</p>\n<p>train_data : <strong>16.39 GB</strong> + 30.75Mb (train_labels.csv) --&gt; <strong>1.8 GB</strong><br>\ntest_data: <strong>33.82 GB</strong> --&gt; <strong>3.7 GB</strong></p>\n<p>The Overall : <strong>50.31 GB</strong>✨✨--&gt; ✨✨ <strong>5.5 GB</strong></p>\n<p>link to dataset --&gt; <a href=\"https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart</a></p>",
  "messages": [
    {
      "id": "1829546",
      "postDate": "06/22/2022 17:20:00",
      "content": "<p>Hello Everyone,<br>\nThis Competition is challenging because of the data size that is difficult to handle.<br>\nas mentioned <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in his very insightful topic here Before discussion about the format to apply to our dataset we have to optimize every column type so that it would take up as little memory space as possible.</p>\n<p>I propose this code sample that a friend (@mathurinache) gave me once:</p>\n<pre><code>def reduce_mem_usage(df, verbose=True):\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage before optimization is: {:.2f} MB'.format(start_mem))\n    for col in df.columns:\n        col_type = df[col].dtypes\n        if col_type in numerics:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min &gt; np.iinfo(np.int8).min and c_max &lt; np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min &gt; np.iinfo(np.int16).min and c_max &lt; np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min &gt; np.iinfo(np.int32).min and c_max &lt; np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min &gt; np.iinfo(np.int64).min and c_max &lt; np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min &gt; np.finfo(np.float16).min and c_max &lt; np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n    return df\n</code></pre>\n<p>I have created Feather format datasets for the initial datasets<br>\n(the final train dataset is merged with the labels)</p>\n<p>train_data : <strong>16.39 GB</strong> + 30.75Mb (train_labels.csv) --&gt; <strong>1.8 GB</strong><br>\ntest_data: <strong>33.82 GB</strong> --&gt; <strong>3.7 GB</strong></p>\n<p>The Overall : <strong>50.31 GB</strong>✨✨--&gt; ✨✨ <strong>5.5 GB</strong></p>\n<p>link to dataset --&gt; <a href=\"https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart</a></p>",
      "rawMarkdown": "Hello Everyone,\nThis Competition is challenging because of the data size that is difficult to handle.\nas mentioned @cdeotte in his very insightful topic here Before discussion about the format to apply to our dataset we have to optimize every column type so that it would take up as little memory space as possible.\n\nI propose this code sample that a friend (@mathurinache) gave me once:\n\n```\ndef reduce_mem_usage(df, verbose=True):\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage before optimization is: {:.2f} MB'.format(start_mem))\n    for col in df.columns:\n        col_type = df[col].dtypes\n        if col_type in numerics:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n    return df\n```\n\nI have created Feather format datasets for the initial datasets\n(the final train dataset is merged with the labels)\n\ntrain_data : **16.39 GB** + 30.75Mb (train_labels.csv) --> **1.8 GB**\ntest_data: **33.82 GB** --> **3.7 GB**\n\nThe Overall : **50.31 GB**✨✨--> ✨✨ **5.5 GB**\n\nlink to dataset --> [https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart](https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart)",
      "votes": null
    },
    {
      "id": "1898657",
      "postDate": "08/14/2022 17:53:27",
      "content": "<p>Thanks for the effort you put on reducing the datasize. It helped me a lot.</p>",
      "rawMarkdown": "Thanks for the effort you put on reducing the datasize. It helped me a lot.",
      "votes": null
    },
    {
      "id": "1899109",
      "postDate": "08/15/2022 04:22:06",
      "content": "<p>You're Welcome 😊        </p>",
      "rawMarkdown": "You're Welcome 😊",
      "votes": null
    },
    {
      "id": "1899503",
      "postDate": "08/15/2022 09:56:09",
      "content": "<p>Thanks for your topic</p>",
      "rawMarkdown": "Thanks for your topic",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1898657,
      "author_name": "gopiravindran",
      "author_url": "",
      "post_date": "08/14/2022 17:53:27",
      "content": "<p>Thanks for the effort you put on reducing the datasize. It helped me a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1899109,
          "author_name": "schopenhacker75",
          "author_url": "",
          "post_date": "08/15/2022 04:22:06",
          "content": "<p>You're Welcome 😊        </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1899503,
      "author_name": "perrypck",
      "author_url": "",
      "post_date": "08/15/2022 09:56:09",
      "content": "<p>Thanks for your topic</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1829546": "Hello Everyone,\nThis Competition is challenging because of the data size that is difficult to handle.\nas mentioned @cdeotte in his very insightful topic here Before discussion about the format to apply to our dataset we have to optimize every column type so that it would take up as little memory space as possible.\n\nI propose this code sample that a friend (@mathurinache) gave me once:\n\n```\ndef reduce_mem_usage(df, verbose=True):\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage before optimization is: {:.2f} MB'.format(start_mem))\n    for col in df.columns:\n        col_type = df[col].dtypes\n        if col_type in numerics:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)\n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n    return df\n```\n\nI have created Feather format datasets for the initial datasets\n(the final train dataset is merged with the labels)\n\ntrain_data : **16.39 GB** + 30.75Mb (train_labels.csv) --> **1.8 GB**\ntest_data: **33.82 GB** --> **3.7 GB**\n\nThe Overall : **50.31 GB**✨✨--> ✨✨ **5.5 GB**\n\nlink to dataset --> [https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart](https://www.kaggle.com/datasets/schopenhacker75/amex-optimizedfeather-formart)",
    "1898657": "Thanks for the effort you put on reducing the datasize. It helped me a lot.",
    "1899109": "You're Welcome 😊",
    "1899503": "Thanks for your topic"
  },
  "source": "meta"
}