{
  "id": 356540,
  "title": "Dtypes transformation (Reduce Memory Usage 75%)",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356540",
  "author_name": "SSS",
  "post_date": "2022-10-01T00:42:16.732000",
  "votes": 40,
  "comment_count": 20,
  "views": 0,
  "content": "<pre><code>def reduce_mem_usage(df, verbose=True):\n    start_mem = df.memory_usage().sum()/1024**2\n    numerics = ['int8', 'int16', 'int32', 'int64',\n                'float16', 'float32', 'float64']\n\n\n\n    for col in df.columns:\n        if col == 'team_scoring_next':\n            continue\n        col_type = df[col].dtypes\n        limit_int = abs(df[col]).max()\n        precision_float = df[col].apply(lambda x: len(str(x).split('.')[1])).max() \n\n        for tp in numerics:\n            cond1 = str(col_type)[0] == tp[0]\n            if tp[0] == 'i': cond2 = limit &lt;= np.iinfo(tp).max\n            else: cond2 = precision_float &lt; np.finfo(tp).precision\n\n            if cond1 and cond2:\n                df[col] = df[col].astype(tp)\n                break\n\n    end_mem = df.memory_usage().sum()/1024**2\n\n    reduction = (start_mem - end_mem)*100/start_mem\n    if verbose:\n        print(f'[INFO] Mem. usage decreased to {end_mem:.2f}'\n              f' MB {reduction:.2f}% reduction.')\n    return df\n</code></pre>\n<p><strong>One step:</strong><br>\n[INFO] Mem. usage decreased to 260.33 MB 73.98% reduction.</p>\n<p><strong>Overall:</strong><br>\n[INFO] Mem. usage decreased from 10027.15 MB to 3153.70 MB 68.55% reduction.</p>\n<p>P.s. <a href=\"https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated\" target=\"_blank\">my code here</a>.<br>\nP.P.s <a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet\" target=\"_blank\">dataset I use</a>.</p>\n<pre><code>train = pd.read_parquet('../input/tps2022octparquet/train.parquet.gzip')\ntest = pd.read_parquet('../input/tps2022octparquet/test.parquet.gzip')\n\ndtypes_dict_train = dict(pd.read_csv(../input/tps2022octparquet/dtypes_train.csv').values)\ndtypes_dict_test = dict(pd.read_csv('../input/tps2022octparquet/dtypes_test.csv').values)\n\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\ntest = test.astype(dtypes_dict_test)\n</code></pre>",
  "messages": [
    {
      "id": 1964881,
      "postDate": "2022-10-01T00:42:16.733Z",
      "content": "<pre><code>def reduce_mem_usage(df, verbose=True):\n    start_mem = df.memory_usage().sum()/1024**2\n    numerics = ['int8', 'int16', 'int32', 'int64',\n                'float16', 'float32', 'float64']\n\n\n\n    for col in df.columns:\n        if col == 'team_scoring_next':\n            continue\n        col_type = df[col].dtypes\n        limit_int = abs(df[col]).max()\n        precision_float = df[col].apply(lambda x: len(str(x).split('.')[1])).max() \n\n        for tp in numerics:\n            cond1 = str(col_type)[0] == tp[0]\n            if tp[0] == 'i': cond2 = limit &lt;= np.iinfo(tp).max\n            else: cond2 = precision_float &lt; np.finfo(tp).precision\n\n            if cond1 and cond2:\n                df[col] = df[col].astype(tp)\n                break\n\n    end_mem = df.memory_usage().sum()/1024**2\n\n    reduction = (start_mem - end_mem)*100/start_mem\n    if verbose:\n        print(f'[INFO] Mem. usage decreased to {end_mem:.2f}'\n              f' MB {reduction:.2f}% reduction.')\n    return df\n</code></pre>\n<p><strong>One step:</strong><br>\n[INFO] Mem. usage decreased to 260.33 MB 73.98% reduction.</p>\n<p><strong>Overall:</strong><br>\n[INFO] Mem. usage decreased from 10027.15 MB to 3153.70 MB 68.55% reduction.</p>\n<p>P.s. <a href=\"https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated\" target=\"_blank\">my code here</a>.<br>\nP.P.s <a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet\" target=\"_blank\">dataset I use</a>.</p>\n<pre><code>train = pd.read_parquet('../input/tps2022octparquet/train.parquet.gzip')\ntest = pd.read_parquet('../input/tps2022octparquet/test.parquet.gzip')\n\ndtypes_dict_train = dict(pd.read_csv(../input/tps2022octparquet/dtypes_train.csv').values)\ndtypes_dict_test = dict(pd.read_csv('../input/tps2022octparquet/dtypes_test.csv').values)\n\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\ntest = test.astype(dtypes_dict_test)\n</code></pre>",
      "rawMarkdown": "```\ndef reduce_mem_usage(df, verbose=True):\n    start_mem = df.memory_usage().sum()/1024**2\n    numerics = ['int8', 'int16', 'int32', 'int64',\n                'float16', 'float32', 'float64']\n    \n    \n    \n    for col in df.columns:\n        if col == 'team_scoring_next':\n            continue\n        col_type = df[col].dtypes\n        limit_int = abs(df[col]).max()\n        precision_float = df[col].apply(lambda x: len(str(x).split('.')[1])).max() \n\n        for tp in numerics:\n            cond1 = str(col_type)[0] == tp[0]\n            if tp[0] == 'i': cond2 = limit <= np.iinfo(tp).max\n            else: cond2 = precision_float < np.finfo(tp).precision\n\n            if cond1 and cond2:\n                df[col] = df[col].astype(tp)\n                break\n\n    end_mem = df.memory_usage().sum()/1024**2\n    \n    reduction = (start_mem - end_mem)*100/start_mem\n    if verbose:\n        print(f'[INFO] Mem. usage decreased to {end_mem:.2f}'\n              f' MB {reduction:.2f}% reduction.')\n    return df\n```\n**One step:**\n[INFO] Mem. usage decreased to 260.33 MB 73.98% reduction.\n\n**Overall:**\n[INFO] Mem. usage decreased from 10027.15 MB to 3153.70 MB 68.55% reduction.\n\nP.s. [my code here](https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated).\nP.P.s [dataset I use](https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet).\n```\ntrain = pd.read_parquet('../input/tps2022octparquet/train.parquet.gzip')\ntest = pd.read_parquet('../input/tps2022octparquet/test.parquet.gzip')\n\ndtypes_dict_train = dict(pd.read_csv(../input/tps2022octparquet/dtypes_train.csv').values)\ndtypes_dict_test = dict(pd.read_csv('../input/tps2022octparquet/dtypes_test.csv').values)\n    \n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\ntest = test.astype(dtypes_dict_test)\n```\n",
      "votes": 39
    },
    {
      "id": 1981893,
      "postDate": "2022-10-11T06:10:38.523Z",
      "content": "<p>How did you obtain the parquet files? Did you read the csv and saved it to parquet one by one?</p>",
      "rawMarkdown": "How did you obtain the parquet files? Did you read the csv and saved it to parquet one by one?",
      "votes": 1,
      "replies": [
        {
          "id": 1982016,
          "postDate": "2022-10-11T07:39:15.907Z",
          "content": "<p>Yes the author created it using the code above and saved it to parquet format.<br>\nYou can also check this <a href=\"https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather\" target=\"_blank\">notebook</a> here where I have done something similar and the data is also available.</p>",
          "rawMarkdown": "Yes the author created it using the code above and saved it to parquet format.\nYou can also check this [notebook](https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather) here where I have done something similar and the data is also available.",
          "votes": 1
        },
        {
          "id": 1983674,
          "postDate": "2022-10-12T07:03:12.740Z",
          "content": "<p>It is confusing a bit, since looking at this:</p>\n<p><code># Reduce memory usage by 70%.</code><br>\n<code>train = train.astype(dtypes_dict_train)</code><br>\n<code>test = test.astype(dtypes_dict_test)</code></p>\n<p>it looks like there have been some parquet files obtained somehow; then, they are read and the values converted to lower-precision types, which reduces the memory usage. </p>\n<p>It is not clear how the parquet files were obtained in the first place. </p>\n<p>Is it by converting the data  to lower-precision types + saving it as parquet? If so, then why do we need to convert the data once again e.g.: <code>train = train.astype(dtypes_dict_train)</code> ?</p>",
          "rawMarkdown": "It is confusing a bit, since looking at this:\n\n`# Reduce memory usage by 70%.`\n`train = train.astype(dtypes_dict_train)`\n`test = test.astype(dtypes_dict_test)`\n\nit looks like there have been some parquet files obtained somehow; then, they are read and the values converted to lower-precision types, which reduces the memory usage. \n\nIt is not clear how the parquet files were obtained in the first place. \n\nIs it by converting the data  to lower-precision types + saving it as parquet? If so, then why do we need to convert the data once again e.g.: `train = train.astype(dtypes_dict_train)` ?"
        },
        {
          "id": 1985564,
          "postDate": "2022-10-13T12:18:39.030Z",
          "content": "<p><code>dtypes_dict_train</code> has different data-type compared to what you get when you do <code>pd.read_csv(train_0.csv)</code> (which will by-default convert them to higher datatype, like float32 will be read as float64).</p>",
          "rawMarkdown": "`dtypes_dict_train` has different data-type compared to what you get when you do `pd.read_csv(train_0.csv)` (which will by-default convert them to higher datatype, like float32 will be read as float64)."
        }
      ]
    },
    {
      "id": 1980880,
      "postDate": "2022-10-10T12:43:40.503Z",
      "content": "<p>Did you check that the numbers did not change ? Changing the dtype like this can modify your value. It is best to be aware of : <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5652850%2F36473aa2e1dc62d689301aae07e6d678%2FCapture%20dcran%202022-10-10%20144250.png?generation=1665405790188479&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Did you check that the numbers did not change ? Changing the dtype like this can modify your value. It is best to be aware of : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5652850%2F36473aa2e1dc62d689301aae07e6d678%2FCapture%20dcran%202022-10-10%20144250.png?generation=1665405790188479&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1981236,
          "postDate": "2022-10-10T17:13:22.740Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1978148,
      "postDate": "2022-10-08T14:57:40.560Z",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>, how similar this is to the reduce _mem_usage that has been around for some time, it's the same or a superior version for this competition</p>",
      "rawMarkdown": "Thanks for sharing, @sergiosaharovskiy, how similar this is to the reduce _mem_usage that has been around for some time, it's the same or a superior version for this competition",
      "votes": 1,
      "replies": [
        {
          "id": 1979369,
          "postDate": "2022-10-09T11:12:48.283Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1979423,
          "postDate": "2022-10-09T12:29:43.767Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
          "rawMarkdown": "Thanks @sergiosaharovskiy ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1973089,
      "postDate": "2022-10-05T13:16:14.053Z",
      "content": "<p>Wow, it was nice and simple but very effective. Thanks for sharing.</p>",
      "rawMarkdown": "Wow, it was nice and simple but very effective. Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1970234,
      "postDate": "2022-10-04T00:40:04.500Z",
      "content": "<p>Fantastic!</p>",
      "rawMarkdown": "Fantastic!",
      "votes": 1
    },
    {
      "id": 1968376,
      "postDate": "2022-10-03T04:21:57.440Z",
      "content": "<p>Very cool!</p>",
      "rawMarkdown": "Very cool!",
      "votes": 1
    },
    {
      "id": 1966589,
      "postDate": "2022-10-02T04:54:12.210Z",
      "content": "<p>That's a clever move🙏</p>",
      "rawMarkdown": "That's a clever move🙏",
      "votes": 1
    },
    {
      "id": 1966240,
      "postDate": "2022-10-01T19:36:26.280Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. <br>\nI wish I had seen it sooner. I also explored similar <a href=\"https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather\" target=\"_blank\">approach</a>. Anyone interested feel free to download the output data shared.</p>",
      "rawMarkdown": "Thanks for sharing @sergiosaharovskiy. \nI wish I had seen it sooner. I also explored similar [approach](https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather). Anyone interested feel free to download the output data shared.",
      "votes": 1
    },
    {
      "id": 1964905,
      "postDate": "2022-10-01T01:45:07.520Z",
      "content": "<p>This brings me some memories from TPS Dec 21</p>",
      "rawMarkdown": "This brings me some memories from TPS Dec 21",
      "votes": 1
    },
    {
      "id": 1964893,
      "postDate": "2022-10-01T01:18:58.907Z",
      "content": "<p>Awesome, nice way to reduce some unnecessary memory usage and allocation!</p>",
      "rawMarkdown": "Awesome, nice way to reduce some unnecessary memory usage and allocation!",
      "votes": 1
    },
    {
      "id": 1983682,
      "postDate": "2022-10-12T07:10:55.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1971385,
      "postDate": "2022-10-04T15:12:47.240Z",
      "content": "<p>It's very useful for me.Upvoted! 👍</p>",
      "rawMarkdown": "It's very useful for me.Upvoted! 👍",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1965446,
      "postDate": "2022-10-01T11:17:57.140Z",
      "content": "<p>Great! Thanks for that!</p>",
      "rawMarkdown": "Great! Thanks for that!",
      "votes": 1
    },
    {
      "id": 1964894,
      "postDate": "2022-10-01T01:19:02.297Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "rawMarkdown": "Thanks for sharing @sergiosaharovskiy ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1981893,
      "author_name": "Ol S",
      "author_url": "",
      "post_date": "2022-10-11T06:10:38.523000",
      "content": "<p>How did you obtain the parquet files? Did you read the csv and saved it to parquet one by one?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1982016,
          "author_name": "Ayush Bihani",
          "author_url": "",
          "post_date": "2022-10-11T07:39:15.907000",
          "content": "<p>Yes the author created it using the code above and saved it to parquet format.<br>\nYou can also check this <a href=\"https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather\" target=\"_blank\">notebook</a> here where I have done something similar and the data is also available.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1983674,
          "author_name": "Ol S",
          "author_url": "",
          "post_date": "2022-10-12T07:03:12.740000",
          "content": "<p>It is confusing a bit, since looking at this:</p>\n<p><code># Reduce memory usage by 70%.</code><br>\n<code>train = train.astype(dtypes_dict_train)</code><br>\n<code>test = test.astype(dtypes_dict_test)</code></p>\n<p>it looks like there have been some parquet files obtained somehow; then, they are read and the values converted to lower-precision types, which reduces the memory usage. </p>\n<p>It is not clear how the parquet files were obtained in the first place. </p>\n<p>Is it by converting the data  to lower-precision types + saving it as parquet? If so, then why do we need to convert the data once again e.g.: <code>train = train.astype(dtypes_dict_train)</code> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1985564,
          "author_name": "Ayush Bihani",
          "author_url": "",
          "post_date": "2022-10-13T12:18:39.030000",
          "content": "<p><code>dtypes_dict_train</code> has different data-type compared to what you get when you do <code>pd.read_csv(train_0.csv)</code> (which will by-default convert them to higher datatype, like float32 will be read as float64).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1980880,
      "author_name": "MaximeTut",
      "author_url": "",
      "post_date": "2022-10-10T12:43:40.503000",
      "content": "<p>Did you check that the numbers did not change ? Changing the dtype like this can modify your value. It is best to be aware of : <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5652850%2F36473aa2e1dc62d689301aae07e6d678%2FCapture%20dcran%202022-10-10%20144250.png?generation=1665405790188479&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1981236,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-10T17:13:22.740000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1978148,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-10-08T14:57:40.560000",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>, how similar this is to the reduce _mem_usage that has been around for some time, it's the same or a superior version for this competition</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1979369,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-09T11:12:48.283000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1979423,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2022-10-09T12:29:43.767000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1973089,
      "author_name": "Shoaib Hossain",
      "author_url": "",
      "post_date": "2022-10-05T13:16:14.053000",
      "content": "<p>Wow, it was nice and simple but very effective. Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1970234,
      "author_name": "Godsent Abode",
      "author_url": "",
      "post_date": "2022-10-04T00:40:04.500000",
      "content": "<p>Fantastic!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1968376,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2022-10-03T04:21:57.440000",
      "content": "<p>Very cool!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1966589,
      "author_name": "Aklima Akter Rimi",
      "author_url": "",
      "post_date": "2022-10-02T04:54:12.210000",
      "content": "<p>That's a clever move🙏</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1966240,
      "author_name": "Ayush Bihani",
      "author_url": "",
      "post_date": "2022-10-01T19:36:26.280000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. <br>\nI wish I had seen it sooner. I also explored similar <a href=\"https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather\" target=\"_blank\">approach</a>. Anyone interested feel free to download the output data shared.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1964905,
      "author_name": "Jose Cáliz",
      "author_url": "",
      "post_date": "2022-10-01T01:45:07.520000",
      "content": "<p>This brings me some memories from TPS Dec 21</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1964893,
      "author_name": "Michael Gwinn",
      "author_url": "",
      "post_date": "2022-10-01T01:18:58.907000",
      "content": "<p>Awesome, nice way to reduce some unnecessary memory usage and allocation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1983682,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-12T07:10:55.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1971385,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-04T15:12:47.240000",
      "content": "<p>It's very useful for me.Upvoted! 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1965446,
      "author_name": "Daniil7191",
      "author_url": "",
      "post_date": "2022-10-01T11:17:57.140000",
      "content": "<p>Great! Thanks for that!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1964894,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-10-01T01:19:02.297000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1964881": "```\ndef reduce_mem_usage(df, verbose=True):\n    start_mem = df.memory_usage().sum()/1024**2\n    numerics = ['int8', 'int16', 'int32', 'int64',\n                'float16', 'float32', 'float64']\n    \n    \n    \n    for col in df.columns:\n        if col == 'team_scoring_next':\n            continue\n        col_type = df[col].dtypes\n        limit_int = abs(df[col]).max()\n        precision_float = df[col].apply(lambda x: len(str(x).split('.')[1])).max() \n\n        for tp in numerics:\n            cond1 = str(col_type)[0] == tp[0]\n            if tp[0] == 'i': cond2 = limit <= np.iinfo(tp).max\n            else: cond2 = precision_float < np.finfo(tp).precision\n\n            if cond1 and cond2:\n                df[col] = df[col].astype(tp)\n                break\n\n    end_mem = df.memory_usage().sum()/1024**2\n    \n    reduction = (start_mem - end_mem)*100/start_mem\n    if verbose:\n        print(f'[INFO] Mem. usage decreased to {end_mem:.2f}'\n              f' MB {reduction:.2f}% reduction.')\n    return df\n```\n**One step:**\n[INFO] Mem. usage decreased to 260.33 MB 73.98% reduction.\n\n**Overall:**\n[INFO] Mem. usage decreased from 10027.15 MB to 3153.70 MB 68.55% reduction.\n\nP.s. [my code here](https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated).\nP.P.s [dataset I use](https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet).\n```\ntrain = pd.read_parquet('../input/tps2022octparquet/train.parquet.gzip')\ntest = pd.read_parquet('../input/tps2022octparquet/test.parquet.gzip')\n\ndtypes_dict_train = dict(pd.read_csv(../input/tps2022octparquet/dtypes_train.csv').values)\ndtypes_dict_test = dict(pd.read_csv('../input/tps2022octparquet/dtypes_test.csv').values)\n    \n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\ntest = test.astype(dtypes_dict_test)\n```\n",
    "1981893": "How did you obtain the parquet files? Did you read the csv and saved it to parquet one by one?",
    "1980880": "Did you check that the numbers did not change ? Changing the dtype like this can modify your value. It is best to be aware of : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5652850%2F36473aa2e1dc62d689301aae07e6d678%2FCapture%20dcran%202022-10-10%20144250.png?generation=1665405790188479&alt=media)",
    "1978148": "Thanks for sharing, @sergiosaharovskiy, how similar this is to the reduce _mem_usage that has been around for some time, it's the same or a superior version for this competition",
    "1973089": "Wow, it was nice and simple but very effective. Thanks for sharing.",
    "1970234": "Fantastic!",
    "1968376": "Very cool!",
    "1966589": "That's a clever move🙏",
    "1966240": "Thanks for sharing @sergiosaharovskiy. \nI wish I had seen it sooner. I also explored similar [approach](https://www.kaggle.com/code/hsuyab/fast-loading-high-compression-with-feather). Anyone interested feel free to download the output data shared.",
    "1964905": "This brings me some memories from TPS Dec 21",
    "1964893": "Awesome, nice way to reduce some unnecessary memory usage and allocation!",
    "1983682": "",
    "1971385": "It's very useful for me.Upvoted! 👍",
    "1965446": "Great! Thanks for that!",
    "1964894": "Thanks for sharing @sergiosaharovskiy "
  }
}