{
  "id": 328054,
  "title": "How To Reduce Data Size",
  "url": "/competitions/amex-default-prediction/discussion/328054",
  "author_name": "Chris Deotte",
  "post_date": "2022-05-30T16:54:06.636000",
  "votes": 521,
  "comment_count": 137,
  "views": 0,
  "content": "<h1>How To Reduce Data Size</h1>\n<p>This competition's tabular data is 50GB! That's huge. To engineer features from this data and train models with this data, we need to reduce data size and efficiently use memory and disk first. Here are some tips</p>\n<h1>Step 1 - Reduce Data Types!</h1>\n<p>Many discussion topics discuss different file formats like Parquet, Feather, NumPy, Pickle, CSV, etc etc. This overlooks the most important point. The first step is reducing each column to the least data size possible. Afterward we can choose our file format and whether to save as multiple files or single file.</p>\n<h3>Column <code>customer_ID</code> - Reduce 64 bytes to 4 bytes!</h3>\n<p>This column is a string of length 64 which uses 64 bytes per row! That is too much! We can convert this <code>int32</code> or <code>int64</code> which only use 4 bytes or 8 bytes. My favorite technique is to take the last 16 letters of the hexadecimal string and convert that base16 number into base10 and save as <code>int64</code>. Discussion to explain this is <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a></p>\n<h3>Column <code>S_2</code> - Reduce 10 bytes to 3 bytes!</h3>\n<p>This column is a date with time. This column is provided as a string of length 10 which uses 10 bytes per row! This is too much! If we convert this column with <code>pd.to_datetime()</code> then it becomes only 4 bytes. Or we can save this column as three columns of <code>year_last_2_digits</code>, <code>month</code>, and <code>day</code> as <code>int8</code> each and only use 3 bytes per row.</p>\n<h3>11 Categorical Columns - Reduce 88 bytes to 11 bytes!</h3>\n<p>The 11 columns <code>['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']</code> are categorical with maximum 8 values. Therefore each of these columns can be converted into <code>int8</code> which is 1 byte per row. Originally they are each 8 bytes per row.</p>\n<h3>177 Numeric Columns - Reduce 1416 bytes to 353 bytes!</h3>\n<p>Lastly there are 177 numerical columns. These columns are <code>float64</code> with 8 bytes per row. At the bare minimum, we can convert <code>float64</code> to <code>float32</code> (4 bytes per row) without losing any important information. We are also discovering that we can convert these to <code>float16</code> which is 2 bytes per row since we suspect that Amex has added uniform noise. (And column <code>B_31</code> has only two values and can be converted to <code>int8</code> 1 byte per row)</p>\n<h1>Step 2 - Choose Your File Format</h1>\n<p>After making the above size reductions, we are now ready to save files to disk. The two important properties of the different options are compression ratio and save order. Some file formats like CSV save the data row by row. And some file formats like Parquet save the data column by column. This will affect reading the data later. If in the future we want to read a subset of rows. perhaps row order is better and faster. If in the future we want to read a subset of columns, perhaps column order is better and faster.</p>\n<p>The second property is compression. Above we talked about number of bytes per row. The train data has <code>5,531,451</code> rows. After the reductions above, we have <code>4 + 3 + 11 + 353 = 371</code> bytes per row. Therefore our uncompressed train data size is 2GB. And test data has <code>11,363,762</code> rows, therefore uncompressed test data is 4GB. Therefore uncompressed, all the competition data becomes 6GB instead of 50GB.  Wow!</p>\n<p>Certain file formats like Feather, Parquet, and Pickle compress the data. (And NumPy has an option <code>np.savez()</code> too). If we compress the data, then all the data can reduce to 4GB or 3GB, or 2GB. Wow! </p>\n<p>UPDATE: A third property is whether the file format remembers the dtype. When saving as parquet, if you downcast an <code>int64</code> to <code>int8</code>, then the file format remembers this and next time you read the file it loads as <code>int8</code>. However some formats like <code>CSV</code> do not remember this. And if you save as <code>int8</code>, it will still read as <code>int64</code> unless you specify dtype in your <code>read_csv()</code> command.</p>\n<h1>Step 3 - Choose Multiple Files or Not</h1>\n<p>Our last decision is whether to save the data as multiple files or one large file. If we have trouble processing the entire train and test dataset at once, then we can consider processing in chunks and saving the data to disk as separate files.</p>\n<h1>Step 4 - Read Raddar's Discussion</h1>\n<p>UPDATE: Many of the variables are actually (low cardinality) integers with noise added. Therefore for some columns, if we remove noise, we can further reduce the <code>float32</code> (4 bytes) into <code>int8</code> (1 byte) columns and reduce the data more. For more info, read Raddar's discussion post <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>\n<h1>Conclusion</h1>\n<p>In conclusion, there are many options to reduce the data and save it in a new file format. Many Kagglers have posted many discussion topics and many Kaggle datasets. Which is best depends on many factors and personal preference. </p>\n<p>We certainly do not want to use the original 50GB CSV files provided in this competition in our pipeline. We will certainly want to do <code>step-1</code> above and then we can pick our favorite <code>step-2</code> and <code>step-3</code>. </p>\n<p>I provide one example in my GRU starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>. For <code>step-3</code>, I choose to load the original train data in 10 chunks and process each chunk separately and save as 10 separate files. First i perform <code>step-1</code> above and for <code>step-2</code>, i choose to save the chunks as uncompressed NumPy files. I uploaded the result to a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 1805980,
      "postDate": "2022-05-30T16:54:06.637Z",
      "content": "<h1>How To Reduce Data Size</h1>\n<p>This competition's tabular data is 50GB! That's huge. To engineer features from this data and train models with this data, we need to reduce data size and efficiently use memory and disk first. Here are some tips</p>\n<h1>Step 1 - Reduce Data Types!</h1>\n<p>Many discussion topics discuss different file formats like Parquet, Feather, NumPy, Pickle, CSV, etc etc. This overlooks the most important point. The first step is reducing each column to the least data size possible. Afterward we can choose our file format and whether to save as multiple files or single file.</p>\n<h3>Column <code>customer_ID</code> - Reduce 64 bytes to 4 bytes!</h3>\n<p>This column is a string of length 64 which uses 64 bytes per row! That is too much! We can convert this <code>int32</code> or <code>int64</code> which only use 4 bytes or 8 bytes. My favorite technique is to take the last 16 letters of the hexadecimal string and convert that base16 number into base10 and save as <code>int64</code>. Discussion to explain this is <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a></p>\n<h3>Column <code>S_2</code> - Reduce 10 bytes to 3 bytes!</h3>\n<p>This column is a date with time. This column is provided as a string of length 10 which uses 10 bytes per row! This is too much! If we convert this column with <code>pd.to_datetime()</code> then it becomes only 4 bytes. Or we can save this column as three columns of <code>year_last_2_digits</code>, <code>month</code>, and <code>day</code> as <code>int8</code> each and only use 3 bytes per row.</p>\n<h3>11 Categorical Columns - Reduce 88 bytes to 11 bytes!</h3>\n<p>The 11 columns <code>['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']</code> are categorical with maximum 8 values. Therefore each of these columns can be converted into <code>int8</code> which is 1 byte per row. Originally they are each 8 bytes per row.</p>\n<h3>177 Numeric Columns - Reduce 1416 bytes to 353 bytes!</h3>\n<p>Lastly there are 177 numerical columns. These columns are <code>float64</code> with 8 bytes per row. At the bare minimum, we can convert <code>float64</code> to <code>float32</code> (4 bytes per row) without losing any important information. We are also discovering that we can convert these to <code>float16</code> which is 2 bytes per row since we suspect that Amex has added uniform noise. (And column <code>B_31</code> has only two values and can be converted to <code>int8</code> 1 byte per row)</p>\n<h1>Step 2 - Choose Your File Format</h1>\n<p>After making the above size reductions, we are now ready to save files to disk. The two important properties of the different options are compression ratio and save order. Some file formats like CSV save the data row by row. And some file formats like Parquet save the data column by column. This will affect reading the data later. If in the future we want to read a subset of rows. perhaps row order is better and faster. If in the future we want to read a subset of columns, perhaps column order is better and faster.</p>\n<p>The second property is compression. Above we talked about number of bytes per row. The train data has <code>5,531,451</code> rows. After the reductions above, we have <code>4 + 3 + 11 + 353 = 371</code> bytes per row. Therefore our uncompressed train data size is 2GB. And test data has <code>11,363,762</code> rows, therefore uncompressed test data is 4GB. Therefore uncompressed, all the competition data becomes 6GB instead of 50GB.  Wow!</p>\n<p>Certain file formats like Feather, Parquet, and Pickle compress the data. (And NumPy has an option <code>np.savez()</code> too). If we compress the data, then all the data can reduce to 4GB or 3GB, or 2GB. Wow! </p>\n<p>UPDATE: A third property is whether the file format remembers the dtype. When saving as parquet, if you downcast an <code>int64</code> to <code>int8</code>, then the file format remembers this and next time you read the file it loads as <code>int8</code>. However some formats like <code>CSV</code> do not remember this. And if you save as <code>int8</code>, it will still read as <code>int64</code> unless you specify dtype in your <code>read_csv()</code> command.</p>\n<h1>Step 3 - Choose Multiple Files or Not</h1>\n<p>Our last decision is whether to save the data as multiple files or one large file. If we have trouble processing the entire train and test dataset at once, then we can consider processing in chunks and saving the data to disk as separate files.</p>\n<h1>Step 4 - Read Raddar's Discussion</h1>\n<p>UPDATE: Many of the variables are actually (low cardinality) integers with noise added. Therefore for some columns, if we remove noise, we can further reduce the <code>float32</code> (4 bytes) into <code>int8</code> (1 byte) columns and reduce the data more. For more info, read Raddar's discussion post <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>\n<h1>Conclusion</h1>\n<p>In conclusion, there are many options to reduce the data and save it in a new file format. Many Kagglers have posted many discussion topics and many Kaggle datasets. Which is best depends on many factors and personal preference. </p>\n<p>We certainly do not want to use the original 50GB CSV files provided in this competition in our pipeline. We will certainly want to do <code>step-1</code> above and then we can pick our favorite <code>step-2</code> and <code>step-3</code>. </p>\n<p>I provide one example in my GRU starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>. For <code>step-3</code>, I choose to load the original train data in 10 chunks and process each chunk separately and save as 10 separate files. First i perform <code>step-1</code> above and for <code>step-2</code>, i choose to save the chunks as uncompressed NumPy files. I uploaded the result to a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "\n# How To Reduce Data Size\nThis competition's tabular data is 50GB! That's huge. To engineer features from this data and train models with this data, we need to reduce data size and efficiently use memory and disk first. Here are some tips\n\n# Step 1 - Reduce Data Types!\nMany discussion topics discuss different file formats like Parquet, Feather, NumPy, Pickle, CSV, etc etc. This overlooks the most important point. The first step is reducing each column to the least data size possible. Afterward we can choose our file format and whether to save as multiple files or single file.\n\n### Column `customer_ID` - Reduce 64 bytes to 4 bytes!\nThis column is a string of length 64 which uses 64 bytes per row! That is too much! We can convert this `int32` or `int64` which only use 4 bytes or 8 bytes. My favorite technique is to take the last 16 letters of the hexadecimal string and convert that base16 number into base10 and save as `int64`. Discussion to explain this is [here][1]\n\n### Column `S_2` - Reduce 10 bytes to 3 bytes!\nThis column is a date with time. This column is provided as a string of length 10 which uses 10 bytes per row! This is too much! If we convert this column with `pd.to_datetime()` then it becomes only 4 bytes. Or we can save this column as three columns of `year_last_2_digits`, `month`, and `day` as `int8` each and only use 3 bytes per row.\n\n### 11 Categorical Columns - Reduce 88 bytes to 11 bytes!\nThe 11 columns `['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']` are categorical with maximum 8 values. Therefore each of these columns can be converted into `int8` which is 1 byte per row. Originally they are each 8 bytes per row.\n\n### 177 Numeric Columns - Reduce 1416 bytes to 353 bytes!\nLastly there are 177 numerical columns. These columns are `float64` with 8 bytes per row. At the bare minimum, we can convert `float64` to `float32` (4 bytes per row) without losing any important information. We are also discovering that we can convert these to `float16` which is 2 bytes per row since we suspect that Amex has added uniform noise. (And column `B_31` has only two values and can be converted to `int8` 1 byte per row)\n\n# Step 2 - Choose Your File Format\nAfter making the above size reductions, we are now ready to save files to disk. The two important properties of the different options are compression ratio and save order. Some file formats like CSV save the data row by row. And some file formats like Parquet save the data column by column. This will affect reading the data later. If in the future we want to read a subset of rows. perhaps row order is better and faster. If in the future we want to read a subset of columns, perhaps column order is better and faster.\n  \nThe second property is compression. Above we talked about number of bytes per row. The train data has `5,531,451` rows. After the reductions above, we have `4 + 3 + 11 + 353 = 371` bytes per row. Therefore our uncompressed train data size is 2GB. And test data has `11,363,762` rows, therefore uncompressed test data is 4GB. Therefore uncompressed, all the competition data becomes 6GB instead of 50GB.  Wow!\n  \nCertain file formats like Feather, Parquet, and Pickle compress the data. (And NumPy has an option `np.savez()` too). If we compress the data, then all the data can reduce to 4GB or 3GB, or 2GB. Wow! \n  \nUPDATE: A third property is whether the file format remembers the dtype. When saving as parquet, if you downcast an `int64` to `int8`, then the file format remembers this and next time you read the file it loads as `int8`. However some formats like `CSV` do not remember this. And if you save as `int8`, it will still read as `int64` unless you specify dtype in your `read_csv()` command.\n  \n# Step 3 - Choose Multiple Files or Not\nOur last decision is whether to save the data as multiple files or one large file. If we have trouble processing the entire train and test dataset at once, then we can consider processing in chunks and saving the data to disk as separate files.\n\n# Step 4 - Read Raddar's Discussion\nUPDATE: Many of the variables are actually (low cardinality) integers with noise added. Therefore for some columns, if we remove noise, we can further reduce the `float32` (4 bytes) into `int8` (1 byte) columns and reduce the data more. For more info, read Raddar's discussion post [here][5]\n\n# Conclusion\nIn conclusion, there are many options to reduce the data and save it in a new file format. Many Kagglers have posted many discussion topics and many Kaggle datasets. Which is best depends on many factors and personal preference. \n  \nWe certainly do not want to use the original 50GB CSV files provided in this competition in our pipeline. We will certainly want to do `step-1` above and then we can pick our favorite `step-2` and `step-3`. \n   \nI provide one example in my GRU starter notebook [here][2]. For `step-3`, I choose to load the original train data in 10 chunks and process each chunk separately and save as 10 separate files. First i perform `step-1` above and for `step-2`, i choose to save the chunks as uncompressed NumPy files. I uploaded the result to a Kaggle dataset [here][3] and discussion [here][4]\n\n[1]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\n[3]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\n[4]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[5]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
      "votes": 519
    },
    {
      "id": 1806007,
      "postDate": "2022-05-30T17:34:48.377Z",
      "content": "<p>Great tips <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as always! I've actually made some youtube videos on this exact topic. Hope it's okay to post the links here:</p>\n<ul>\n<li><a href=\"https://youtu.be/u4_c2LDi4b8\" target=\"_blank\">This video explains how you can make your pandas dataframe more efficient</a> by aproprately casting the dtypes of the columns- exactly like how you discuss in this post. I also show some benchmarks showing how this can impact the dataframe size and speed.</li>\n<li><a href=\"https://youtu.be/u4rsA5ZiTls\" target=\"_blank\">This video discusses the differences in common file formats</a>, (parquet, feather, csv) each has it's unique positives and negatives. I also do some benchmarks to show the differences.</li>\n</ul>\n<p>I have other videos about pandas like this one about how to <a href=\"https://youtu.be/SAFmrTnEHLg\" target=\"_blank\">speed up pandas code by over 2500x</a>!</p>\n<p>Would love to hear any feedback and/or suggestions for future videos.</p>\n<p>Sorry to hijack your thread! If anyone finds these videos helpful please consider <a href=\"https://bit.ly/3N5ygGA\" target=\"_blank\">subscribing to my channel here</a>.</p>",
      "rawMarkdown": "Great tips @cdeotte as always! I've actually made some youtube videos on this exact topic. Hope it's okay to post the links here:\n\n- [This video explains how you can make your pandas dataframe more efficient](https://youtu.be/u4_c2LDi4b8) by aproprately casting the dtypes of the columns- exactly like how you discuss in this post. I also show some benchmarks showing how this can impact the dataframe size and speed.\n- [This video discusses the differences in common file formats](https://youtu.be/u4rsA5ZiTls), (parquet, feather, csv) each has it's unique positives and negatives. I also do some benchmarks to show the differences.\n\nI have other videos about pandas like this one about how to [speed up pandas code by over 2500x](https://youtu.be/SAFmrTnEHLg)!\n\nWould love to hear any feedback and/or suggestions for future videos.\n\nSorry to hijack your thread! If anyone finds these videos helpful please consider [subscribing to my channel here](https://bit.ly/3N5ygGA).",
      "votes": 21,
      "replies": [
        {
          "id": 1806025,
          "postDate": "2022-05-30T17:53:40.767Z",
          "content": "<p>Great videos <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> Thanks for sharing the links</p>",
          "rawMarkdown": "Great videos @robikscube Thanks for sharing the links",
          "votes": 5
        },
        {
          "id": 1820635,
          "postDate": "2022-06-14T19:49:03.680Z",
          "content": "<p>Thanks both of you <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "Thanks both of you @robikscube and @cdeotte ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1806495,
      "postDate": "2022-05-31T08:38:30.917Z",
      "content": "<p>I don't know much about hex. Is there any chance that last 16 characters of the customer_ID can cause integer overflow for 64 bit integers? If I understand correctly, 64 bit signed integer type is just enough for 16 characters of hex strings, 32 bit signed integer type is just enough for 8 characters of hex strings, and it goes on like that.</p>\n<p>I also found an elegant way to convert hex string into base 10, 64 bit integer for pandas.<br>\n<code>df_train['customer_ID'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)</code></p>\n<p>I guess we can even push it further and convert it to 32 bit integer by using last 8 characters. We have to check number of unique values to see whether the cardinality is affected or not.</p>\n<p>Edit: You should use int64/16 characters because cardinality is affected when int32/8 characters are used for training set. </p>\n<pre><code>&gt;&gt;&gt; df_train['customer_ID'].nunique()\n458913\n&gt;&gt;&gt; df_train['customer_ID_int'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n&gt;&gt;&gt; df_train['customer_ID_int'].nunique()\n458913\n&gt;&gt;&gt; df_train['customer_ID_int'] = df_train['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n&gt;&gt;&gt; df_train['customer_ID_int'].nunique()\n458884\n</code></pre>\n<pre><code>&gt;&gt;&gt; df_test['customer_ID'].nunique()\n924621\n&gt;&gt;&gt; df_test['customer_ID_int'] = df_test['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n&gt;&gt;&gt; df_test['customer_ID_int'].nunique()\n924621\n&gt;&gt;&gt; df_test['customer_ID_int'] = df_test['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n&gt;&gt;&gt; df_test['customer_ID_int'].nunique()\n924547\n</code></pre>",
      "rawMarkdown": "I don't know much about hex. Is there any chance that last 16 characters of the customer_ID can cause integer overflow for 64 bit integers? If I understand correctly, 64 bit signed integer type is just enough for 16 characters of hex strings, 32 bit signed integer type is just enough for 8 characters of hex strings, and it goes on like that.\n\nI also found an elegant way to convert hex string into base 10, 64 bit integer for pandas.\n`df_train['customer_ID'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)`\n\nI guess we can even push it further and convert it to 32 bit integer by using last 8 characters. We have to check number of unique values to see whether the cardinality is affected or not.\n\nEdit: You should use int64/16 characters because cardinality is affected when int32/8 characters are used for training set. \n\n```\n>>> df_train['customer_ID'].nunique()\n458913\n>>> df_train['customer_ID_int'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n>>> df_train['customer_ID_int'].nunique()\n458913\n>>> df_train['customer_ID_int'] = df_train['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n>>> df_train['customer_ID_int'].nunique()\n458884\n```\n\n```\n>>> df_test['customer_ID'].nunique()\n924621\n>>> df_test['customer_ID_int'] = df_test['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n>>> df_test['customer_ID_int'].nunique()\n924621\n>>> df_test['customer_ID_int'] = df_test['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n>>> df_test['customer_ID_int'].nunique()\n924547\n```",
      "votes": 11,
      "replies": [
        {
          "id": 1806591,
          "postDate": "2022-05-31T10:14:28.587Z",
          "content": "<p>You can save even more if you cast the column to <code>category</code> instead of <code>int64</code> by exploiting the duplication of <code>customer_ID</code>.</p>\n<pre><code>print(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('int64').memory_usage(deep=True))\nprint(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('category').memory_usage(deep=True))\n</code></pre>\n<pre><code>44251736\n42705564\n</code></pre>",
          "rawMarkdown": "You can save even more if you cast the column to `category` instead of `int64` by exploiting the duplication of `customer_ID`.\n```\nprint(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('int64').memory_usage(deep=True))\nprint(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('category').memory_usage(deep=True))\n```\n```\n44251736\n42705564\n```",
          "votes": 1
        },
        {
          "id": 1806813,
          "postDate": "2022-05-31T13:32:04.370Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Great analysis Gunes. Each hexadecimal digit ( i.e. one character ) represents 16 values or 4 bits. So yes, 16 characters are 64 bits and 8 characters are 32 bits. The way you checked unique customers is very important (and I did this also). This is called <code>hash collision</code> (wikipedia <a href=\"https://en.wikipedia.org/wiki/Hash_collision\" target=\"_blank\">here</a>). Using <code>int64</code> has <strong>no</strong> hash collision and performs perfect conversion for both train and test.</p>\n<p>If we want to compress the <code>customer_ID</code> into <code>int32</code>, we can do it, but we need to use a label encoder like <code>train['customer'],codes = train.customer.factorize()</code>. But even though this has greater compression, i don't like this because then we need to save the label encoder map to disk too and we need to concatenate train and test together before label encoding (in most competitions when there are ID overlaps). </p>\n<p>Using hexadecimal conversion is very elegant because we can open any file like <code>sample_submission.csv</code> and immediately know how to convert the <code>customer_ID</code> without needing to load an additional label encode map from disk.</p>",
          "rawMarkdown": "@gunesevitan Great analysis Gunes. Each hexadecimal digit ( i.e. one character ) represents 16 values or 4 bits. So yes, 16 characters are 64 bits and 8 characters are 32 bits. The way you checked unique customers is very important (and I did this also). This is called `hash collision` (wikipedia [here][1]). Using `int64` has **no** hash collision and performs perfect conversion for both train and test.\n\nIf we want to compress the `customer_ID` into `int32`, we can do it, but we need to use a label encoder like `train['customer'],codes = train.customer.factorize()`. But even though this has greater compression, i don't like this because then we need to save the label encoder map to disk too and we need to concatenate train and test together before label encoding (in most competitions when there are ID overlaps). \n\nUsing hexadecimal conversion is very elegant because we can open any file like `sample_submission.csv` and immediately know how to convert the `customer_ID` without needing to load an additional label encode map from disk.\n\n\n[1]: https://en.wikipedia.org/wiki/Hash_collision",
          "votes": 9
        },
        {
          "id": 1840363,
          "postDate": "2022-07-02T07:38:30.897Z",
          "content": "<p>Is it ok to use unsigned integer, for example, if using pandas, leaving out the <code>.astype(np.int64)</code>?  It looks like <code>np.uint64</code> has the same memory usage.</p>",
          "rawMarkdown": "Is it ok to use unsigned integer, for example, if using pandas, leaving out the `.astype(np.int64)`?  It looks like `np.uint64` has the same memory usage."
        }
      ]
    },
    {
      "id": 1833511,
      "postDate": "2022-06-26T04:51:51.367Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Hi!Thank you for this!It is pretty helpful!Could you please elaborate on the reduction for customer_id column?</p>",
      "rawMarkdown": "@cdeotte  Hi!Thank you for this!It is pretty helpful!Could you please elaborate on the reduction for customer_id column?",
      "votes": 3,
      "replies": [
        {
          "id": 1833575,
          "postDate": "2022-06-26T06:15:21.553Z",
          "content": "<p>Hi. The basic idea is that the dataset provides us with <code>customer_ID</code> that are strings of length 64. Each string takes 64 bytes per string. We note that there are only about 450k unique customers in train and 900k unique customers in test. So instead of calling each customer as a long string, we can call the first customer <code>int</code> value 1. We can call the second <code>int</code> value 2. The next <code>int</code> value 3 etc etc.</p>\n<p>Then we have renumbered all the train customers from int value 1 thru 450k. And we have numbered all the test customers as int val 450,001 thru 1,300,000. After renaming all the customers, we convert the column into <code>int32</code> with <code>df['customer_ID'] = df.customer_ID.map(CHANGE_CUSTOMER).astype('int32')</code>. That's the basic idea.</p>",
          "rawMarkdown": "Hi. The basic idea is that the dataset provides us with `customer_ID` that are strings of length 64. Each string takes 64 bytes per string. We note that there are only about 450k unique customers in train and 900k unique customers in test. So instead of calling each customer as a long string, we can call the first customer `int` value 1. We can call the second `int` value 2. The next `int` value 3 etc etc.\n\nThen we have renumbered all the train customers from int value 1 thru 450k. And we have numbered all the test customers as int val 450,001 thru 1,300,000. After renaming all the customers, we convert the column into `int32` with `df['customer_ID'] = df.customer_ID.map(CHANGE_CUSTOMER).astype('int32')`. That's the basic idea.",
          "votes": 4
        },
        {
          "id": 1862808,
          "postDate": "2022-07-20T02:56:09.617Z",
          "content": "<p>How would you then go back to the real customer_IDs in the Test dataframe prior to the submit?</p>",
          "rawMarkdown": "How would you then go back to the real customer_IDs in the Test dataframe prior to the submit?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1806935,
      "postDate": "2022-05-31T15:30:28.053Z",
      "content": "<p>All this is good, but how are you ever going to do all these transformations without first opening the data set?<br>\nWhich I can't open even with 16GB RAM?</p>",
      "rawMarkdown": "All this is good, but how are you ever going to do all these transformations without first opening the data set?\nWhich I can't open even with 16GB RAM?",
      "votes": 3,
      "replies": [
        {
          "id": 1806974,
          "postDate": "2022-05-31T16:08:20.943Z",
          "content": "<p>Pandas and cuDF allow you to read a subset of the data rows from the CSV. Therefore i suggest you read the train data as 10 chunks. And the test data as 20 chunks. Then each chunk only has around 500,000 rows and will easily fit in memory. For an example, see my notebook <a href=\"https://tinyurl.com/4pwp47p9\" target=\"_blank\">here</a>. The example is in hidden code cell 6. </p>\n<pre><code>train = cudf.read_csv('../input/amex-default-prediction/train_data.csv', \n    nrows=rows[k], skiprows=skip, header=None, names=T_COLS)\n</code></pre>\n<p>This same code works for both <code>cudf</code> and <code>pandas</code>.</p>",
          "rawMarkdown": "Pandas and cuDF allow you to read a subset of the data rows from the CSV. Therefore i suggest you read the train data as 10 chunks. And the test data as 20 chunks. Then each chunk only has around 500,000 rows and will easily fit in memory. For an example, see my notebook [here][1]. The example is in hidden code cell 6. \n\n    train = cudf.read_csv('../input/amex-default-prediction/train_data.csv', \n        nrows=rows[k], skiprows=skip, header=None, names=T_COLS)\n\nThis same code works for both `cudf` and `pandas`.\n\n[1]: https://tinyurl.com/4pwp47p9",
          "votes": 11
        },
        {
          "id": 1808621,
          "postDate": "2022-06-02T03:34:04.470Z",
          "content": "<p>You can use the <strong>chunksize</strong> parameter to iterate multiple chunks in pandas. In the example below I am using a dummy value</p>\n<pre><code># make data chunks\ndata_chunks  = pd.read_csv(\"file.csv\", chunksize=10000)\n\n# iterate over chunks\nfor chunk in data_chunks:\n     # chunk is of type df\n     # perform all the operations on this df slice here\n</code></pre>",
          "rawMarkdown": "You can use the **chunksize** parameter to iterate multiple chunks in pandas. In the example below I am using a dummy value\n\n```\n# make data chunks\ndata_chunks  = pd.read_csv(\"file.csv\", chunksize=10000)\n\n# iterate over chunks\nfor chunk in data_chunks:\n     # chunk is of type df\n     # perform all the operations on this df slice here\n```",
          "votes": 2
        },
        {
          "id": 1810937,
          "postDate": "2022-06-04T05:05:32.350Z",
          "content": "<p>Thank you, very helpful!</p>",
          "rawMarkdown": "Thank you, very helpful!"
        },
        {
          "id": 1814716,
          "postDate": "2022-06-08T07:38:46.497Z",
          "content": "<p>I have successfully loaded entire training data to the provided instance with 16GB with <code>dask</code> and while loading define the columns types as  following:<br>\n <code>df_train = dd.read_csv(\"../input/amex-default-prediction/train_data.csv\", dtype=col_dtypes)</code><br>\nsee kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash</a></p>",
          "rawMarkdown": "I have successfully loaded entire training data to the provided instance with 16GB with `dask` and while loading define the columns types as  following:\n `df_train = dd.read_csv(\"../input/amex-default-prediction/train_data.csv\", dtype=col_dtypes)`\nsee kernel: https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash"
        }
      ]
    },
    {
      "id": 1819567,
      "postDate": "2022-06-13T23:36:39.070Z",
      "content": "<p>Great tips, especially for those working at home on personal hardware 😅</p>",
      "rawMarkdown": "Great tips, especially for those working at home on personal hardware 😅",
      "votes": 4,
      "replies": [
        {
          "id": 1822033,
          "postDate": "2022-06-16T03:08:49.010Z",
          "content": "<p>haha,I work at home with 16G memory</p>",
          "rawMarkdown": "haha,I work at home with 16G memory",
          "votes": 1
        }
      ]
    },
    {
      "id": 1808486,
      "postDate": "2022-06-01T21:37:34.153Z",
      "content": "<p>train_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>",
      "rawMarkdown": "train_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)",
      "votes": 4,
      "replies": [
        {
          "id": 1808660,
          "postDate": "2022-06-02T04:40:12.060Z",
          "content": "<p>Great suggestions</p>",
          "rawMarkdown": "Great suggestions"
        }
      ]
    },
    {
      "id": 1906332,
      "postDate": "2022-08-19T19:15:40.017Z",
      "content": "<p>So let's say you do <code>df['customer_ID'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')</code>, then how do you get back to the original customer IDs?</p>\n<p>The original IDs are required for a submission. I could imagine one way would be to re-load all test IDs, then convert the int versions to base 16 strings (e.g. using numpy.base_repr(id_num, 16)) and then match those to the last 16 characters of the IDs? Seems a bit roundabout but I guess there's no other way.</p>",
      "rawMarkdown": "So let's say you do `df['customer_ID'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')`, then how do you get back to the original customer IDs?\n\nThe original IDs are required for a submission. I could imagine one way would be to re-load all test IDs, then convert the int versions to base 16 strings (e.g. using numpy.base_repr(id_num, 16)) and then match those to the last 16 characters of the IDs? Seems a bit roundabout but I guess there's no other way.",
      "votes": 1,
      "replies": [
        {
          "id": 1906337,
          "postDate": "2022-08-19T19:24:53.800Z",
          "content": "<p>You need to create a mapping. To do this, load the original CSV file and make a new column called `code', then export a dictionary map:</p>\n<pre><code>df = pd.read_csv('train_data.csv')\ndf['code'] = df['customer_ID']\\\n    .apply(lambda x: int(x[-16:],16) ).astype('int64')\ndf = df.set_index('code')\nMAPPING = df['customer_ID'].to_dict()\n</code></pre>\n<p>That creates a dictionary. Then you can convert a <code>code</code> column like the following</p>\n<pre><code>my_dataframe['customer_ID'] = my_dataframe['code'].map( MAPPING)\n</code></pre>",
          "rawMarkdown": "You need to create a mapping. To do this, load the original CSV file and make a new column called `code', then export a dictionary map:\n\n    df = pd.read_csv('train_data.csv')\n    df['code'] = df['customer_ID']\\\n        .apply(lambda x: int(x[-16:],16) ).astype('int64')\n    df = df.set_index('code')\n    MAPPING = df['customer_ID'].to_dict()\n\nThat creates a dictionary. Then you can convert a `code` column like the following\n\n    my_dataframe['customer_ID'] = my_dataframe['code'].map( MAPPING)",
          "votes": 4
        },
        {
          "id": 1906359,
          "postDate": "2022-08-19T19:54:39.677Z",
          "content": "<p>Aha, that's a better idea than what I proposed.</p>",
          "rawMarkdown": "Aha, that's a better idea than what I proposed.",
          "votes": 1
        },
        {
          "id": 1906364,
          "postDate": "2022-08-19T20:05:21.647Z",
          "content": "<p>The best way is what i do in my stater notebook. After making test predictions, put those test predictions into a new dataframe that has column <code>code</code> and column <code>prediction</code>. Then load Kaggle's submission.csv file. Then make a new column <code>code</code> in Kaggle's submission.csv. Finally merge (i.e. <code>sub = sub.merge(df, on='code')</code> ) the predictions from your dataframe onto the submission.csv dataframe using the column <code>code</code> for the merge.</p>",
          "rawMarkdown": "The best way is what i do in my stater notebook. After making test predictions, put those test predictions into a new dataframe that has column `code` and column `prediction`. Then load Kaggle's submission.csv file. Then make a new column `code` in Kaggle's submission.csv. Finally merge (i.e. `sub = sub.merge(df, on='code')` ) the predictions from your dataframe onto the submission.csv dataframe using the column `code` for the merge.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1897922,
      "postDate": "2022-08-14T06:43:33.190Z",
      "content": "<p>great discussion</p>",
      "rawMarkdown": "great discussion",
      "votes": 1
    },
    {
      "id": 1856525,
      "postDate": "2022-07-15T11:47:11.323Z",
      "content": "<p>Good tips!</p>",
      "rawMarkdown": "Good tips!",
      "votes": 1
    },
    {
      "id": 1853867,
      "postDate": "2022-07-13T08:23:20.347Z",
      "content": "<p>very friendly to new guys，I learned a lot from the share，thank you very much</p>",
      "rawMarkdown": "very friendly to new guys，I learned a lot from the share，thank you very much",
      "votes": 1
    },
    {
      "id": 1842847,
      "postDate": "2022-07-04T10:41:44.960Z",
      "content": "<p>Thanks for sharing. however, I believe it depends on case to case. one technique may be good in one situation but not ideal for the second.</p>",
      "rawMarkdown": "Thanks for sharing. however, I believe it depends on case to case. one technique may be good in one situation but not ideal for the second.",
      "votes": 1
    },
    {
      "id": 1835110,
      "postDate": "2022-06-27T13:29:45.053Z",
      "content": "<p>By implementing a hash, you can reduce it to 8 (or whatever resolution you want) bits per feature, at the cost of some computational speed.</p>",
      "rawMarkdown": "By implementing a hash, you can reduce it to 8 (or whatever resolution you want) bits per feature, at the cost of some computational speed.",
      "votes": 1
    },
    {
      "id": 1834873,
      "postDate": "2022-06-27T09:09:35.883Z",
      "content": "<p>A very nice tips. Thanks a lot to share this here</p>",
      "rawMarkdown": " A very nice tips. Thanks a lot to share this here",
      "votes": 1
    },
    {
      "id": 1833909,
      "postDate": "2022-06-26T12:52:45.357Z",
      "content": "<p>Thanks a lot for your great effort. what is the difference in practice between saving data in multiple files or in one file? can we say that one of them is better?</p>",
      "rawMarkdown": "Thanks a lot for your great effort. what is the difference in practice between saving data in multiple files or in one file? can we say that one of them is better?",
      "votes": 1
    },
    {
      "id": 1833510,
      "postDate": "2022-06-26T04:50:41.337Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi!Thank you so much!This is extremely helpful. Could you please elaborate on transforming the customer_id column?</p>",
      "rawMarkdown": "@cdeotte Hi!Thank you so much!This is extremely helpful. Could you please elaborate on transforming the customer_id column?",
      "votes": 1
    },
    {
      "id": 1830087,
      "postDate": "2022-06-23T06:44:53.087Z",
      "content": "<p>It's very helpful. Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 👍</p>",
      "rawMarkdown": "It's very helpful. Thanks for sharing @cdeotte 👍",
      "votes": 1
    },
    {
      "id": 1828483,
      "postDate": "2022-06-21T18:23:47.510Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> While converting the categorical features into \"int8\", it gives an error saying cannot convert NA values to int. How do we best approach this error? Thanks</p>",
      "rawMarkdown": "@cdeotte While converting the categorical features into \"int8\", it gives an error saying cannot convert NA values to int. How do we best approach this error? Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 1833564,
          "postDate": "2022-06-26T06:10:35.910Z",
          "content": "<p>hi <a href=\"https://www.kaggle.com/apurbapandey\" target=\"_blank\">@apurbapandey</a> . When using <code>Pandas</code>, dtype <code>int</code> cannot be <code>NA</code>. Therefore we must <code>fillna</code> before converting to dtype like <code>df[col] = df[col].fillna(-1).astype('int8')</code>. Note that when using RAPIDS cuDF, we the dtype <code>int</code> can contain NA.</p>",
          "rawMarkdown": "hi @apurbapandey . When using `Pandas`, dtype `int` cannot be `NA`. Therefore we must `fillna` before converting to dtype like `df[col] = df[col].fillna(-1).astype('int8')`. Note that when using RAPIDS cuDF, we the dtype `int` can contain NA.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1822035,
      "postDate": "2022-06-16T03:09:20.280Z",
      "content": "<p>so useful!👍</p>",
      "rawMarkdown": "so useful!👍",
      "votes": 1
    },
    {
      "id": 1820646,
      "postDate": "2022-06-14T20:06:13.763Z",
      "content": "<p>Very helpful,i will definetly utilise these tricks.thanks for posting.</p>",
      "rawMarkdown": "Very helpful,i will definetly utilise these tricks.thanks for posting.",
      "votes": 1
    },
    {
      "id": 1820276,
      "postDate": "2022-06-14T13:42:26.430Z",
      "content": "<p>This is quite helpful </p>",
      "rawMarkdown": "This is quite helpful ",
      "votes": 1
    },
    {
      "id": 1819380,
      "postDate": "2022-06-13T17:41:02.403Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Very useful informative post , Thanks for posting .</p>",
      "rawMarkdown": "@cdeotte Very useful informative post , Thanks for posting .",
      "votes": 1
    },
    {
      "id": 1818814,
      "postDate": "2022-06-13T07:05:28.117Z",
      "content": "<p>This is very informative. Thanks a lot for this.</p>",
      "rawMarkdown": "This is very informative. Thanks a lot for this.",
      "votes": 1
    },
    {
      "id": 1818635,
      "postDate": "2022-06-13T02:32:19.347Z",
      "content": "<p>Thanks a bunch for this informative post</p>",
      "rawMarkdown": "Thanks a bunch for this informative post",
      "votes": 1
    },
    {
      "id": 1818266,
      "postDate": "2022-06-12T12:55:56.947Z",
      "content": "<p>Thanks a lot for sharing this technique dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> .</p>",
      "rawMarkdown": "Thanks a lot for sharing this technique dear @cdeotte .",
      "votes": 1
    },
    {
      "id": 1817442,
      "postDate": "2022-06-11T10:51:24.343Z",
      "content": "<p>Thanks to the people who advised on this matter.</p>\n<p>But it's of no use for people who work with R, so here's something they might find useful:</p>\n<p>install.packages('sqldf')<br>\nlibrary('sqldf')</p>\n<p>Then use the command read.csv.sql to load the file in parts:</p>\n<p>your.data.set=read.csv.sql('train.csv', 'select [name of variable] from file where [whatever]')</p>\n<p>Note that the word \"file\" must not be changed; By \"file\" R understands the first argument, which is \"train.csv\".</p>\n<p>And good luck with it. I for one feel that it is not worth burning my computer to work with such huge files. And I'd like to know what American Express actually want from this competition; a few good algorithms, or handling huge data sets with home PCs?</p>",
      "rawMarkdown": "Thanks to the people who advised on this matter.\n\nBut it's of no use for people who work with R, so here's something they might find useful:\n\ninstall.packages('sqldf')\nlibrary('sqldf')\n\nThen use the command read.csv.sql to load the file in parts:\n\nyour.data.set=read.csv.sql('train.csv', 'select [name of variable] from file where [whatever]')\n\nNote that the word \"file\" must not be changed; By \"file\" R understands the first argument, which is \"train.csv\".\n\nAnd good luck with it. I for one feel that it is not worth burning my computer to work with such huge files. And I'd like to know what American Express actually want from this competition; a few good algorithms, or handling huge data sets with home PCs?",
      "votes": 1
    },
    {
      "id": 1816701,
      "postDate": "2022-06-10T13:23:56.443Z",
      "content": "<p>Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!</p>",
      "rawMarkdown": "Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!",
      "votes": 1
    },
    {
      "id": 1816605,
      "postDate": "2022-06-10T11:03:29.937Z",
      "content": "<p>This is great stuff Chris! I was able to reduce the size by adopting your advice! Also, I used Dask instead of Pandas. It helped me process the data fast :)</p>",
      "rawMarkdown": "This is great stuff Chris! I was able to reduce the size by adopting your advice! Also, I used Dask instead of Pandas. It helped me process the data fast :)",
      "votes": 1
    },
    {
      "id": 1815962,
      "postDate": "2022-06-09T17:04:18.990Z",
      "content": "<p>Excellent advices. Will give it a go.</p>",
      "rawMarkdown": "Excellent advices. Will give it a go.",
      "votes": 1
    },
    {
      "id": 1815868,
      "postDate": "2022-06-09T15:03:56.860Z",
      "content": "<p>Great insight <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. You have taught me a new way of looking at data. Thanks for sharing </p>",
      "rawMarkdown": "Great insight @cdeotte. You have taught me a new way of looking at data. Thanks for sharing ",
      "votes": 1
    },
    {
      "id": 1815644,
      "postDate": "2022-06-09T09:02:14.490Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!</p>",
      "rawMarkdown": "@cdeotte Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!",
      "votes": 1
    },
    {
      "id": 1814747,
      "postDate": "2022-06-08T08:39:43.533Z",
      "content": "<p>Thank you for your advices, very interesting !</p>",
      "rawMarkdown": "Thank you for your advices, very interesting !",
      "votes": 1
    },
    {
      "id": 1812571,
      "postDate": "2022-06-06T03:01:02.030Z",
      "content": "<p>great and useful insight</p>",
      "rawMarkdown": "great and useful insight",
      "votes": 1
    },
    {
      "id": 1810120,
      "postDate": "2022-06-03T09:33:50.820Z",
      "content": "<p>Very interesting!</p>",
      "rawMarkdown": "Very interesting!",
      "votes": 1
    },
    {
      "id": 1810072,
      "postDate": "2022-06-03T08:33:51.773Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for writing the detailed steps. Super helpful!</p>",
      "rawMarkdown": " @cdeotte Thanks for writing the detailed steps. Super helpful!",
      "votes": 1
    },
    {
      "id": 1809914,
      "postDate": "2022-06-03T06:03:24.047Z",
      "content": "<p>Thnks a lot!!!!</p>",
      "rawMarkdown": "Thnks a lot!!!!",
      "votes": 1
    },
    {
      "id": 1809359,
      "postDate": "2022-06-02T16:32:27.297Z",
      "content": "<p>Great insight to start. Thanks for sharing</p>",
      "rawMarkdown": "Great insight to start. Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1808946,
      "postDate": "2022-06-02T09:02:29.523Z",
      "content": "<p>This is an amazing post. Thanks a lot for putting this up.</p>",
      "rawMarkdown": "This is an amazing post. Thanks a lot for putting this up.",
      "votes": 1
    },
    {
      "id": 1808710,
      "postDate": "2022-06-02T05:22:44.367Z",
      "content": "<p>Quite helpful!!</p>",
      "rawMarkdown": "Quite helpful!!",
      "votes": 1
    },
    {
      "id": 1808619,
      "postDate": "2022-06-02T03:31:48.820Z",
      "content": "<p>This is so helpful that I will always do it, thanks!</p>",
      "rawMarkdown": "This is so helpful that I will always do it, thanks!",
      "votes": 1
    },
    {
      "id": 1808263,
      "postDate": "2022-06-01T16:59:22.650Z",
      "content": "<p>UPDATE: Added a <code>Step 4</code> about removing noise to further reduce data size of some columns. For more info, see Raddar's disccusion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "UPDATE: Added a `Step 4` about removing noise to further reduce data size of some columns. For more info, see Raddar's disccusion [here][1]\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
      "votes": 1
    },
    {
      "id": 1807697,
      "postDate": "2022-06-01T09:11:41.707Z",
      "content": "<p>Very useful and interesting.</p>",
      "rawMarkdown": "Very useful and interesting.\n",
      "votes": 1
    },
    {
      "id": 1806905,
      "postDate": "2022-05-31T14:57:17.143Z",
      "content": "<p>Very useful and interesting.<br>\nLove you. Thanks!👍</p>",
      "rawMarkdown": "Very useful and interesting.\nLove you. Thanks!👍",
      "votes": 1
    },
    {
      "id": 1806862,
      "postDate": "2022-05-31T14:03:02.320Z",
      "content": "<p>Very informative and useful. Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "rawMarkdown": "Very informative and useful. Thank you @cdeotte",
      "votes": 1
    },
    {
      "id": 1806377,
      "postDate": "2022-05-31T05:18:33.283Z",
      "content": "<p>Very informative and useful. Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Very informative and useful. Thank you @cdeotte ",
      "votes": 1
    },
    {
      "id": 1882231,
      "postDate": "2022-08-03T06:31:09.653Z",
      "content": "<p>Thanks for giving us your knowledge, it seems really good</p>",
      "rawMarkdown": "Thanks for giving us your knowledge, it seems really good",
      "votes": 2
    },
    {
      "id": 1820883,
      "postDate": "2022-06-15T04:41:22.390Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the amazing post. Really helps learn practical issues when dealing with industry scale datasets</p>",
      "rawMarkdown": "@cdeotte for the amazing post. Really helps learn practical issues when dealing with industry scale datasets",
      "votes": 2
    },
    {
      "id": 1818349,
      "postDate": "2022-06-12T14:51:35.590Z",
      "content": "<p>Thank you! I didn't know about these techniques! </p>",
      "rawMarkdown": "Thank you! I didn't know about these techniques! ",
      "votes": 2
    },
    {
      "id": 1814712,
      "postDate": "2022-06-08T07:33:05.503Z",
      "content": "<p>I have successfully loaded almost all columns as float16 so I could fit the entire training data to the provided instance with 16GB<br>\nsee kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash</a></p>",
      "rawMarkdown": "I have successfully loaded almost all columns as float16 so I could fit the entire training data to the provided instance with 16GB\nsee kernel: https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash",
      "votes": 2
    },
    {
      "id": 1808622,
      "postDate": "2022-06-02T03:35:14.087Z",
      "content": "<p>Thanks a lot for putting in the effort for writing this! The best part is that this technique is useful across almost any use case using tabular data.</p>",
      "rawMarkdown": "Thanks a lot for putting in the effort for writing this! The best part is that this technique is useful across almost any use case using tabular data.",
      "votes": 2,
      "replies": [
        {
          "id": 1808661,
          "postDate": "2022-06-02T04:40:39.627Z",
          "content": "<p>Absolutely. This is the first thing i do in every tabular data modeling project.</p>",
          "rawMarkdown": "Absolutely. This is the first thing i do in every tabular data modeling project.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1808396,
      "postDate": "2022-06-01T19:11:39.367Z",
      "content": "<p>Nice points <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Step 4 was bugging my mind for a while too. I wonder would it be better to convert some of the noise injected features into categorical ones, otherwise we're working with small data residing in huge noise which adds only an extra memory load. On the other hand could we use the patterns in the noise injections to our benefit if there's any…</p>",
      "rawMarkdown": "Nice points @cdeotte! Step 4 was bugging my mind for a while too. I wonder would it be better to convert some of the noise injected features into categorical ones, otherwise we're working with small data residing in huge noise which adds only an extra memory load. On the other hand could we use the patterns in the noise injections to our benefit if there's any...",
      "votes": 2
    },
    {
      "id": 1806220,
      "postDate": "2022-05-30T23:04:20.067Z",
      "content": "<p>Really nice!. I am struggling with the test data, this may help!</p>",
      "rawMarkdown": "Really nice!. I am struggling with the test data, this may help!",
      "votes": 2
    },
    {
      "id": 1806053,
      "postDate": "2022-05-30T18:29:13.287Z",
      "content": "<p>Thanks for sharing, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. This post is so informative and helpful; the size of these datasets makes everything more challenging in this competition, so the explanations here will help me to improve my strategy</p>",
      "rawMarkdown": "Thanks for sharing, @cdeotte. This post is so informative and helpful; the size of these datasets makes everything more challenging in this competition, so the explanations here will help me to improve my strategy",
      "votes": 2
    },
    {
      "id": 3462433,
      "postDate": "2026-05-23T09:13:42.617Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for this, it's a really valuable and informative, and i learned a lot from this share.</p>",
      "rawMarkdown": "Thanks @cdeotte for this, it's a really valuable and informative, and i learned a lot from this share."
    },
    {
      "id": 3107954,
      "postDate": "2025-01-27T06:56:02.757Z",
      "content": "<p>That was really helpful ! Thankyou!</p>",
      "rawMarkdown": "That was really helpful ! Thankyou!\n"
    },
    {
      "id": 2735011,
      "postDate": "2024-04-04T13:21:34.457Z",
      "content": "<p>Great tips, especially for those working at home on not expensive personal hardware.</p>",
      "rawMarkdown": "Great tips, especially for those working at home on not expensive personal hardware.\n\n\n"
    },
    {
      "id": 2515833,
      "postDate": "2023-11-07T09:15:12.753Z",
      "content": "<p>thanks a lot for your effort. learned a lot from your code and dataset :)</p>",
      "rawMarkdown": "thanks a lot for your effort. learned a lot from your code and dataset :)"
    },
    {
      "id": 2509919,
      "postDate": "2023-11-02T16:26:17.617Z",
      "content": "<p>Thank you Chris for your in-depth explanation on reducing the data size. It has been really helpful! </p>\n<p>Can you please elaborate why you only converted the last 16 characters of the hexadecimal string into int? What if the are multiple customer_ids that are different but have the same last 16 characters? </p>\n<p>Moreover, is saving to int64 needed - is it not enough to do train['customer_id'].apply(lambda x: int( x[-16:], 16))?</p>",
      "rawMarkdown": "Thank you Chris for your in-depth explanation on reducing the data size. It has been really helpful! \n\nCan you please elaborate why you only converted the last 16 characters of the hexadecimal string into int? What if the are multiple customer_ids that are different but have the same last 16 characters? \n\nMoreover, is saving to int64 needed - is it not enough to do train['customer_id'].apply(lambda x: int( x[-16:], 16))?"
    },
    {
      "id": 2150109,
      "postDate": "2023-02-19T00:29:19.787Z",
      "content": "<p>These suggestions are useful.Thank you for these good suggestions</p>",
      "rawMarkdown": "These suggestions are useful.Thank you for these good suggestions"
    },
    {
      "id": 1847949,
      "postDate": "2022-07-08T09:14:39.473Z",
      "content": "<p>Very Very detail and usefull,  learning a lot about \"Choose Your File Format\", Thank you CHRIS,  nice to see you in this competition!</p>",
      "rawMarkdown": "Very Very detail and usefull,  learning a lot about \"Choose Your File Format\", Thank you CHRIS,  nice to see you in this competition!"
    },
    {
      "id": 1846124,
      "postDate": "2022-07-06T21:03:01.670Z",
      "content": "<p>nice idea! ilike it!</p>",
      "rawMarkdown": "nice idea! ilike it!"
    },
    {
      "id": 1843194,
      "postDate": "2022-07-04T16:08:01.853Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Its simply lovely post. I have no other words. 💕💕💕</p>",
      "rawMarkdown": "@cdeotte Its simply lovely post. I have no other words. 💕💕💕"
    },
    {
      "id": 1807968,
      "postDate": "2022-06-01T13:05:38.330Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Reducing data size feels like … starting a journey of a thousand miles begins with a single step.</p>",
      "rawMarkdown": "Thanks for sharing @cdeotte , Reducing data size feels like ... starting a journey of a thousand miles begins with a single step."
    },
    {
      "id": 1903422,
      "postDate": "2022-08-17T12:26:21.027Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1905855,
          "postDate": "2022-08-19T11:16:12.657Z",
          "content": "<p>Yes, we must check this. We compute the <code>nunique()</code> before and after transformation and we observe that the number of unique customers is the same. So for this competition's dataset, using only the last 16 characters works. </p>\n<p>(I also checked whether we could use only the last 8 characters with <code>int32</code> but in that case the number of unique customers decreases which implies that two or more customers have the same last 8 characters).</p>",
          "rawMarkdown": "Yes, we must check this. We compute the `nunique()` before and after transformation and we observe that the number of unique customers is the same. So for this competition's dataset, using only the last 16 characters works. \n\n(I also checked whether we could use only the last 8 characters with `int32` but in that case the number of unique customers decreases which implies that two or more customers have the same last 8 characters).",
          "votes": 2
        },
        {
          "id": 1905963,
          "postDate": "2022-08-19T13:39:40.343Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1855668,
      "postDate": "2022-07-14T19:33:04.060Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1845153,
      "postDate": "2022-07-06T05:26:56.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1820963,
      "postDate": "2022-06-15T06:34:20.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1809388,
      "postDate": "2022-06-02T16:45:12.057Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2022069,
      "postDate": "2022-11-08T17:04:23.343Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": 1
    },
    {
      "id": 1905640,
      "postDate": "2022-08-19T07:43:22.240Z",
      "content": "<p>Thanks for your help.</p>",
      "rawMarkdown": "Thanks for your help.",
      "votes": 1
    },
    {
      "id": 1905636,
      "postDate": "2022-08-19T07:41:48.283Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1897383,
      "postDate": "2022-08-13T17:29:01.033Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1888146,
      "postDate": "2022-08-07T11:44:35.737Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1888066,
      "postDate": "2022-08-07T10:50:06.990Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1885252,
      "postDate": "2022-08-05T03:23:19.220Z",
      "content": "<p>thanks for very useful your knowledge</p>",
      "rawMarkdown": "thanks for very useful your knowledge",
      "votes": 1
    },
    {
      "id": 1882214,
      "postDate": "2022-08-03T06:17:44.413Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1879613,
      "postDate": "2022-08-01T06:57:21.877Z",
      "content": "<p>Well Explained .Thanks for sharing</p>",
      "rawMarkdown": "Well Explained .Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1875602,
      "postDate": "2022-07-29T06:28:07.747Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1860121,
      "postDate": "2022-07-18T06:45:54.147Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1857916,
      "postDate": "2022-07-16T14:02:53.100Z",
      "content": "<p>Thanks for sharing. Very valuable tips</p>",
      "rawMarkdown": "Thanks for sharing. Very valuable tips",
      "votes": 1
    },
    {
      "id": 1855841,
      "postDate": "2022-07-15T01:14:19.640Z",
      "content": "<p>Thanks for this post!</p>",
      "rawMarkdown": "Thanks for this post!",
      "votes": 1
    },
    {
      "id": 1854199,
      "postDate": "2022-07-13T13:50:46.067Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1852078,
      "postDate": "2022-07-11T18:31:47.993Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "votes": 1
    },
    {
      "id": 1850641,
      "postDate": "2022-07-10T15:09:56.440Z",
      "content": "<p>thank you so much!!</p>",
      "rawMarkdown": "thank you so much!!",
      "votes": 1
    },
    {
      "id": 1850257,
      "postDate": "2022-07-10T09:01:56.737Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1841725,
      "postDate": "2022-07-03T12:29:39.437Z",
      "content": "<p>This is good idea. Thanks so much!</p>",
      "rawMarkdown": "This is good idea. Thanks so much!\n\n",
      "votes": 1
    },
    {
      "id": 1837017,
      "postDate": "2022-06-29T09:18:16.883Z",
      "content": "<p>Thank you o much for sharing this </p>",
      "rawMarkdown": "Thank you o much for sharing this ",
      "votes": 1
    },
    {
      "id": 1835827,
      "postDate": "2022-06-28T06:04:07.780Z",
      "content": "<p>This is very efficient, thanks!</p>",
      "rawMarkdown": "This is very efficient, thanks!",
      "votes": 1
    },
    {
      "id": 1835743,
      "postDate": "2022-06-28T03:47:57.900Z",
      "content": "<p>It is very helpful, Thanks !!</p>",
      "rawMarkdown": "It is very helpful, Thanks !!",
      "votes": 1
    },
    {
      "id": 1828253,
      "postDate": "2022-06-21T15:27:47.847Z",
      "content": "<p>This is gold. Thanks so much :)</p>",
      "rawMarkdown": "This is gold. Thanks so much :)",
      "votes": 1
    },
    {
      "id": 1822297,
      "postDate": "2022-06-16T08:19:47.910Z",
      "content": "<p>Thanks for sharing these techniques</p>",
      "rawMarkdown": "Thanks for sharing these techniques",
      "votes": 1
    },
    {
      "id": 1819178,
      "postDate": "2022-06-13T14:02:29Z",
      "content": "<p>thanks for tips</p>",
      "rawMarkdown": "thanks for tips",
      "votes": 1
    },
    {
      "id": 1818120,
      "postDate": "2022-06-12T09:43:08.283Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1813554,
      "postDate": "2022-06-07T01:52:19.403Z",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot.",
      "votes": 1
    },
    {
      "id": 1811961,
      "postDate": "2022-06-05T10:49:23.323Z",
      "content": "<p>thanks a lot </p>",
      "rawMarkdown": "thanks a lot ",
      "votes": 1
    },
    {
      "id": 1811857,
      "postDate": "2022-06-05T08:09:15.117Z",
      "content": "<p><strong><em><em>Thanks, useful insight..!</em></em></strong></p>",
      "rawMarkdown": "****Thanks, useful insight..!****",
      "votes": 1
    },
    {
      "id": 1811656,
      "postDate": "2022-06-05T01:02:58.703Z",
      "content": "<p>Thanks, this was very helpful!</p>",
      "rawMarkdown": "Thanks, this was very helpful!",
      "votes": 1
    },
    {
      "id": 1811252,
      "postDate": "2022-06-04T13:41:58.773Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot",
      "votes": 1
    },
    {
      "id": 1811059,
      "postDate": "2022-06-04T08:44:38.013Z",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot.",
      "votes": 1
    },
    {
      "id": 1811015,
      "postDate": "2022-06-04T07:09:43.700Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1810637,
      "postDate": "2022-06-03T18:41:04.233Z",
      "content": "<p>Thanks, great start, keep adding more.</p>",
      "rawMarkdown": "Thanks, great start, keep adding more.",
      "votes": 1
    },
    {
      "id": 1810206,
      "postDate": "2022-06-03T10:49:42.653Z",
      "content": "<p>thanks a lot.</p>",
      "rawMarkdown": "thanks a lot.",
      "votes": 1
    },
    {
      "id": 1809832,
      "postDate": "2022-06-03T05:05:07.353Z",
      "content": "<p>Such a good way, thanks a lot.</p>",
      "rawMarkdown": "Such a good way, thanks a lot.",
      "votes": 1
    },
    {
      "id": 1809737,
      "postDate": "2022-06-03T02:33:31.897Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot",
      "votes": 1
    },
    {
      "id": 1809660,
      "postDate": "2022-06-02T23:26:42.043Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": 1
    },
    {
      "id": 1809296,
      "postDate": "2022-06-02T15:29:12.960Z",
      "content": "<p>Thanks a lot </p>",
      "rawMarkdown": "Thanks a lot ",
      "votes": 1
    },
    {
      "id": 1808951,
      "postDate": "2022-06-02T09:07:23.400Z",
      "content": "<p>Thanks a lot for this insight.</p>",
      "rawMarkdown": "Thanks a lot for this insight.",
      "votes": 1
    },
    {
      "id": 1808904,
      "postDate": "2022-06-02T08:30:41.003Z",
      "content": "<p>Thanks a lot &lt;</p>",
      "rawMarkdown": "Thanks a lot <",
      "votes": 1
    },
    {
      "id": 1807655,
      "postDate": "2022-06-01T08:34:00.990Z",
      "content": "<p>Thanks for sharing the info!👍</p>",
      "rawMarkdown": "Thanks for sharing the info!👍",
      "votes": 1
    },
    {
      "id": 1807542,
      "postDate": "2022-06-01T06:00:35.063Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot",
      "votes": 1
    },
    {
      "id": 1807419,
      "postDate": "2022-06-01T02:17:27.910Z",
      "content": "<p>very good!!<br>\nThanks a lot!!</p>",
      "rawMarkdown": "very good!!\nThanks a lot!!",
      "votes": 1
    },
    {
      "id": 1807161,
      "postDate": "2022-05-31T19:09:59.877Z",
      "content": "<p>Thanks a lot!!</p>",
      "rawMarkdown": "Thanks a lot!!",
      "votes": 1
    },
    {
      "id": 1807151,
      "postDate": "2022-05-31T18:57:06.113Z",
      "content": "<p>Thank you!! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thank you!! @cdeotte ",
      "votes": 1
    },
    {
      "id": 1891175,
      "postDate": "2022-08-09T09:24:56.383Z",
      "content": "<p>thanks for your share</p>",
      "rawMarkdown": "thanks for your share"
    },
    {
      "id": 1862886,
      "postDate": "2022-07-20T04:33:34.240Z",
      "content": "<p>Thanks for sharing!!!</p>",
      "rawMarkdown": "Thanks for sharing!!!",
      "votes": 2
    },
    {
      "id": 2987894,
      "postDate": "2024-09-13T06:22:15.680Z",
      "content": "<p>Very helpful, thank you!</p>",
      "rawMarkdown": "Very helpful, thank you!"
    },
    {
      "id": 2377631,
      "postDate": "2023-08-07T08:54:21.183Z",
      "content": "<p>Very Useful. Thanks</p>",
      "rawMarkdown": "Very Useful. Thanks"
    },
    {
      "id": 1842837,
      "postDate": "2022-07-04T10:31:25.860Z",
      "content": "<p>Thanks for sharing. very informative. </p>",
      "rawMarkdown": "Thanks for sharing. very informative. "
    }
  ],
  "comments": [
    {
      "id": 1806007,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2022-05-30T17:34:48.377000",
      "content": "<p>Great tips <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as always! I've actually made some youtube videos on this exact topic. Hope it's okay to post the links here:</p>\n<ul>\n<li><a href=\"https://youtu.be/u4_c2LDi4b8\" target=\"_blank\">This video explains how you can make your pandas dataframe more efficient</a> by aproprately casting the dtypes of the columns- exactly like how you discuss in this post. I also show some benchmarks showing how this can impact the dataframe size and speed.</li>\n<li><a href=\"https://youtu.be/u4rsA5ZiTls\" target=\"_blank\">This video discusses the differences in common file formats</a>, (parquet, feather, csv) each has it's unique positives and negatives. I also do some benchmarks to show the differences.</li>\n</ul>\n<p>I have other videos about pandas like this one about how to <a href=\"https://youtu.be/SAFmrTnEHLg\" target=\"_blank\">speed up pandas code by over 2500x</a>!</p>\n<p>Would love to hear any feedback and/or suggestions for future videos.</p>\n<p>Sorry to hijack your thread! If anyone finds these videos helpful please consider <a href=\"https://bit.ly/3N5ygGA\" target=\"_blank\">subscribing to my channel here</a>.</p>",
      "votes": 21,
      "replies": [
        {
          "id": 1806025,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-30T17:53:40.767000",
          "content": "<p>Great videos <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> Thanks for sharing the links</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1820635,
          "author_name": "Gaju Ahmed",
          "author_url": "",
          "post_date": "2022-06-14T19:49:03.680000",
          "content": "<p>Thanks both of you <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1806495,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-05-31T08:38:30.917000",
      "content": "<p>I don't know much about hex. Is there any chance that last 16 characters of the customer_ID can cause integer overflow for 64 bit integers? If I understand correctly, 64 bit signed integer type is just enough for 16 characters of hex strings, 32 bit signed integer type is just enough for 8 characters of hex strings, and it goes on like that.</p>\n<p>I also found an elegant way to convert hex string into base 10, 64 bit integer for pandas.<br>\n<code>df_train['customer_ID'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)</code></p>\n<p>I guess we can even push it further and convert it to 32 bit integer by using last 8 characters. We have to check number of unique values to see whether the cardinality is affected or not.</p>\n<p>Edit: You should use int64/16 characters because cardinality is affected when int32/8 characters are used for training set. </p>\n<pre><code>&gt;&gt;&gt; df_train['customer_ID'].nunique()\n458913\n&gt;&gt;&gt; df_train['customer_ID_int'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n&gt;&gt;&gt; df_train['customer_ID_int'].nunique()\n458913\n&gt;&gt;&gt; df_train['customer_ID_int'] = df_train['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n&gt;&gt;&gt; df_train['customer_ID_int'].nunique()\n458884\n</code></pre>\n<pre><code>&gt;&gt;&gt; df_test['customer_ID'].nunique()\n924621\n&gt;&gt;&gt; df_test['customer_ID_int'] = df_test['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n&gt;&gt;&gt; df_test['customer_ID_int'].nunique()\n924621\n&gt;&gt;&gt; df_test['customer_ID_int'] = df_test['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n&gt;&gt;&gt; df_test['customer_ID_int'].nunique()\n924547\n</code></pre>",
      "votes": 11,
      "replies": [
        {
          "id": 1806591,
          "author_name": "broccoli beef",
          "author_url": "",
          "post_date": "2022-05-31T10:14:28.587000",
          "content": "<p>You can save even more if you cast the column to <code>category</code> instead of <code>int64</code> by exploiting the duplication of <code>customer_ID</code>.</p>\n<pre><code>print(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('int64').memory_usage(deep=True))\nprint(train_data.customer_ID.apply(lambda x:int(x[-16:],16)).astype('category').memory_usage(deep=True))\n</code></pre>\n<pre><code>44251736\n42705564\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1806813,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-31T13:32:04.370000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Great analysis Gunes. Each hexadecimal digit ( i.e. one character ) represents 16 values or 4 bits. So yes, 16 characters are 64 bits and 8 characters are 32 bits. The way you checked unique customers is very important (and I did this also). This is called <code>hash collision</code> (wikipedia <a href=\"https://en.wikipedia.org/wiki/Hash_collision\" target=\"_blank\">here</a>). Using <code>int64</code> has <strong>no</strong> hash collision and performs perfect conversion for both train and test.</p>\n<p>If we want to compress the <code>customer_ID</code> into <code>int32</code>, we can do it, but we need to use a label encoder like <code>train['customer'],codes = train.customer.factorize()</code>. But even though this has greater compression, i don't like this because then we need to save the label encoder map to disk too and we need to concatenate train and test together before label encoding (in most competitions when there are ID overlaps). </p>\n<p>Using hexadecimal conversion is very elegant because we can open any file like <code>sample_submission.csv</code> and immediately know how to convert the <code>customer_ID</code> without needing to load an additional label encode map from disk.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1840363,
          "author_name": "wafflebufflo",
          "author_url": "",
          "post_date": "2022-07-02T07:38:30.897000",
          "content": "<p>Is it ok to use unsigned integer, for example, if using pandas, leaving out the <code>.astype(np.int64)</code>?  It looks like <code>np.uint64</code> has the same memory usage.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1833511,
      "author_name": "Shrinidhi Narasimhan",
      "author_url": "",
      "post_date": "2022-06-26T04:51:51.367000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  Hi!Thank you for this!It is pretty helpful!Could you please elaborate on the reduction for customer_id column?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1833575,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-26T06:15:21.553000",
          "content": "<p>Hi. The basic idea is that the dataset provides us with <code>customer_ID</code> that are strings of length 64. Each string takes 64 bytes per string. We note that there are only about 450k unique customers in train and 900k unique customers in test. So instead of calling each customer as a long string, we can call the first customer <code>int</code> value 1. We can call the second <code>int</code> value 2. The next <code>int</code> value 3 etc etc.</p>\n<p>Then we have renumbered all the train customers from int value 1 thru 450k. And we have numbered all the test customers as int val 450,001 thru 1,300,000. After renaming all the customers, we convert the column into <code>int32</code> with <code>df['customer_ID'] = df.customer_ID.map(CHANGE_CUSTOMER).astype('int32')</code>. That's the basic idea.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1862808,
          "author_name": "Adrian Faela",
          "author_url": "",
          "post_date": "2022-07-20T02:56:09.617000",
          "content": "<p>How would you then go back to the real customer_IDs in the Test dataframe prior to the submit?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1806935,
      "author_name": "querty550",
      "author_url": "",
      "post_date": "2022-05-31T15:30:28.053000",
      "content": "<p>All this is good, but how are you ever going to do all these transformations without first opening the data set?<br>\nWhich I can't open even with 16GB RAM?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1806974,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-31T16:08:20.943000",
          "content": "<p>Pandas and cuDF allow you to read a subset of the data rows from the CSV. Therefore i suggest you read the train data as 10 chunks. And the test data as 20 chunks. Then each chunk only has around 500,000 rows and will easily fit in memory. For an example, see my notebook <a href=\"https://tinyurl.com/4pwp47p9\" target=\"_blank\">here</a>. The example is in hidden code cell 6. </p>\n<pre><code>train = cudf.read_csv('../input/amex-default-prediction/train_data.csv', \n    nrows=rows[k], skiprows=skip, header=None, names=T_COLS)\n</code></pre>\n<p>This same code works for both <code>cudf</code> and <code>pandas</code>.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1808621,
          "author_name": "Anubhav Chhabra",
          "author_url": "",
          "post_date": "2022-06-02T03:34:04.470000",
          "content": "<p>You can use the <strong>chunksize</strong> parameter to iterate multiple chunks in pandas. In the example below I am using a dummy value</p>\n<pre><code># make data chunks\ndata_chunks  = pd.read_csv(\"file.csv\", chunksize=10000)\n\n# iterate over chunks\nfor chunk in data_chunks:\n     # chunk is of type df\n     # perform all the operations on this df slice here\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1810937,
          "author_name": "Aiden White",
          "author_url": "",
          "post_date": "2022-06-04T05:05:32.350000",
          "content": "<p>Thank you, very helpful!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1814716,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-06-08T07:38:46.497000",
          "content": "<p>I have successfully loaded entire training data to the provided instance with 16GB with <code>dask</code> and while loading define the columns types as  following:<br>\n <code>df_train = dd.read_csv(\"../input/amex-default-prediction/train_data.csv\", dtype=col_dtypes)</code><br>\nsee kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1819567,
      "author_name": "Arandeep Dhillon",
      "author_url": "",
      "post_date": "2022-06-13T23:36:39.070000",
      "content": "<p>Great tips, especially for those working at home on personal hardware 😅</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1822033,
          "author_name": "huan jun",
          "author_url": "",
          "post_date": "2022-06-16T03:08:49.010000",
          "content": "<p>haha,I work at home with 16G memory</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1808486,
      "author_name": "Mohsin",
      "author_url": "",
      "post_date": "2022-06-01T21:37:34.153000",
      "content": "<p>train_df['B_31'].unique()<br>\narray([1, 0], dtype=int64)<br>\ntrain_df['D_87'].unique()<br>\narray([nan,  1.], dtype=float16)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1808660,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-02T04:40:12.060000",
          "content": "<p>Great suggestions</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1906332,
      "author_name": "Nathan George",
      "author_url": "",
      "post_date": "2022-08-19T19:15:40.017000",
      "content": "<p>So let's say you do <code>df['customer_ID'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')</code>, then how do you get back to the original customer IDs?</p>\n<p>The original IDs are required for a submission. I could imagine one way would be to re-load all test IDs, then convert the int versions to base 16 strings (e.g. using numpy.base_repr(id_num, 16)) and then match those to the last 16 characters of the IDs? Seems a bit roundabout but I guess there's no other way.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1906337,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-19T19:24:53.800000",
          "content": "<p>You need to create a mapping. To do this, load the original CSV file and make a new column called `code', then export a dictionary map:</p>\n<pre><code>df = pd.read_csv('train_data.csv')\ndf['code'] = df['customer_ID']\\\n    .apply(lambda x: int(x[-16:],16) ).astype('int64')\ndf = df.set_index('code')\nMAPPING = df['customer_ID'].to_dict()\n</code></pre>\n<p>That creates a dictionary. Then you can convert a <code>code</code> column like the following</p>\n<pre><code>my_dataframe['customer_ID'] = my_dataframe['code'].map( MAPPING)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1906359,
          "author_name": "Nathan George",
          "author_url": "",
          "post_date": "2022-08-19T19:54:39.677000",
          "content": "<p>Aha, that's a better idea than what I proposed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1906364,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-19T20:05:21.647000",
          "content": "<p>The best way is what i do in my stater notebook. After making test predictions, put those test predictions into a new dataframe that has column <code>code</code> and column <code>prediction</code>. Then load Kaggle's submission.csv file. Then make a new column <code>code</code> in Kaggle's submission.csv. Finally merge (i.e. <code>sub = sub.merge(df, on='code')</code> ) the predictions from your dataframe onto the submission.csv dataframe using the column <code>code</code> for the merge.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1897922,
      "author_name": "llllllkkkkkk",
      "author_url": "",
      "post_date": "2022-08-14T06:43:33.190000",
      "content": "<p>great discussion</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1856525,
      "author_name": "Melvin",
      "author_url": "",
      "post_date": "2022-07-15T11:47:11.323000",
      "content": "<p>Good tips!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1853867,
      "author_name": "tianzizhao",
      "author_url": "",
      "post_date": "2022-07-13T08:23:20.347000",
      "content": "<p>very friendly to new guys，I learned a lot from the share，thank you very much</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842847,
      "author_name": "Sohail Ahmed",
      "author_url": "",
      "post_date": "2022-07-04T10:41:44.960000",
      "content": "<p>Thanks for sharing. however, I believe it depends on case to case. one technique may be good in one situation but not ideal for the second.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1835110,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2022-06-27T13:29:45.053000",
      "content": "<p>By implementing a hash, you can reduce it to 8 (or whatever resolution you want) bits per feature, at the cost of some computational speed.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1834873,
      "author_name": "Ronny Daniel Taroreh",
      "author_url": "",
      "post_date": "2022-06-27T09:09:35.883000",
      "content": "<p>A very nice tips. Thanks a lot to share this here</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1833909,
      "author_name": "Fatima HABIB",
      "author_url": "",
      "post_date": "2022-06-26T12:52:45.357000",
      "content": "<p>Thanks a lot for your great effort. what is the difference in practice between saving data in multiple files or in one file? can we say that one of them is better?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1833510,
      "author_name": "Shrinidhi Narasimhan",
      "author_url": "",
      "post_date": "2022-06-26T04:50:41.337000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi!Thank you so much!This is extremely helpful. Could you please elaborate on transforming the customer_id column?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1830087,
      "author_name": "Sakshi Jha",
      "author_url": "",
      "post_date": "2022-06-23T06:44:53.087000",
      "content": "<p>It's very helpful. Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1828483,
      "author_name": "Apurba Pandey",
      "author_url": "",
      "post_date": "2022-06-21T18:23:47.510000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> While converting the categorical features into \"int8\", it gives an error saying cannot convert NA values to int. How do we best approach this error? Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1833564,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-26T06:10:35.910000",
          "content": "<p>hi <a href=\"https://www.kaggle.com/apurbapandey\" target=\"_blank\">@apurbapandey</a> . When using <code>Pandas</code>, dtype <code>int</code> cannot be <code>NA</code>. Therefore we must <code>fillna</code> before converting to dtype like <code>df[col] = df[col].fillna(-1).astype('int8')</code>. Note that when using RAPIDS cuDF, we the dtype <code>int</code> can contain NA.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1822035,
      "author_name": "huan jun",
      "author_url": "",
      "post_date": "2022-06-16T03:09:20.280000",
      "content": "<p>so useful!👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1820646,
      "author_name": "Vandit",
      "author_url": "",
      "post_date": "2022-06-14T20:06:13.763000",
      "content": "<p>Very helpful,i will definetly utilise these tricks.thanks for posting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1820276,
      "author_name": "Aneruth Mohanasundaram",
      "author_url": "",
      "post_date": "2022-06-14T13:42:26.430000",
      "content": "<p>This is quite helpful </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1819380,
      "author_name": "Alaa Taha El Maria",
      "author_url": "",
      "post_date": "2022-06-13T17:41:02.403000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Very useful informative post , Thanks for posting .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818814,
      "author_name": "sampathsomayajula",
      "author_url": "",
      "post_date": "2022-06-13T07:05:28.117000",
      "content": "<p>This is very informative. Thanks a lot for this.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818635,
      "author_name": "Kanha patil",
      "author_url": "",
      "post_date": "2022-06-13T02:32:19.347000",
      "content": "<p>Thanks a bunch for this informative post</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818266,
      "author_name": "Saber",
      "author_url": "",
      "post_date": "2022-06-12T12:55:56.947000",
      "content": "<p>Thanks a lot for sharing this technique dear <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1817442,
      "author_name": "querty550",
      "author_url": "",
      "post_date": "2022-06-11T10:51:24.343000",
      "content": "<p>Thanks to the people who advised on this matter.</p>\n<p>But it's of no use for people who work with R, so here's something they might find useful:</p>\n<p>install.packages('sqldf')<br>\nlibrary('sqldf')</p>\n<p>Then use the command read.csv.sql to load the file in parts:</p>\n<p>your.data.set=read.csv.sql('train.csv', 'select [name of variable] from file where [whatever]')</p>\n<p>Note that the word \"file\" must not be changed; By \"file\" R understands the first argument, which is \"train.csv\".</p>\n<p>And good luck with it. I for one feel that it is not worth burning my computer to work with such huge files. And I'd like to know what American Express actually want from this competition; a few good algorithms, or handling huge data sets with home PCs?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1816701,
      "author_name": "wangyuan0619",
      "author_url": "",
      "post_date": "2022-06-10T13:23:56.443000",
      "content": "<p>Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1816605,
      "author_name": "Vivek Chowdhury",
      "author_url": "",
      "post_date": "2022-06-10T11:03:29.937000",
      "content": "<p>This is great stuff Chris! I was able to reduce the size by adopting your advice! Also, I used Dask instead of Pandas. It helped me process the data fast :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1815962,
      "author_name": "JiJung",
      "author_url": "",
      "post_date": "2022-06-09T17:04:18.990000",
      "content": "<p>Excellent advices. Will give it a go.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1815868,
      "author_name": "Dhamu",
      "author_url": "",
      "post_date": "2022-06-09T15:03:56.860000",
      "content": "<p>Great insight <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. You have taught me a new way of looking at data. Thanks for sharing </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1815644,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-09T09:02:14.490000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1814747,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-08T08:39:43.533000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1812571,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-06T03:01:02.030000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1810120,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T09:33:50.820000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1810072,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T08:33:51.773000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809914,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T06:03:24.047000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809359,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T16:32:27.297000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808946,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T09:02:29.523000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808710,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T05:22:44.367000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808619,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T03:31:48.820000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808263,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T16:59:22.650000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807697,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T09:11:41.707000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1806905,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-31T14:57:17.143000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1806862,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-31T14:03:02.320000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1806377,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-31T05:18:33.283000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1882231,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T06:31:09.653000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1820883,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-15T04:41:22.390000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1818349,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-12T14:51:35.590000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1814712,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-08T07:33:05.503000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1808622,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T03:35:14.087000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1808661,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-02T04:40:39.627000",
          "content": "",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1808396,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T19:11:39.367000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1806220,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-30T23:04:20.067000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1806053,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-30T18:29:13.287000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3462433,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-05-23T09:13:42.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3107954,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-27T06:56:02.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2735011,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-04T13:21:34.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2515833,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-07T09:15:12.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2509919,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-02T16:26:17.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2150109,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-19T00:29:19.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1847949,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-08T09:14:39.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1846124,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-06T21:03:01.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1843194,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-04T16:08:01.853000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1807968,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T13:05:38.330000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1903422,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T12:26:21.027000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1905855,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-19T11:16:12.657000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1905963,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-19T13:39:40.343000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1855668,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-14T19:33:04.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1845153,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-06T05:26:56.317000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1820963,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-15T06:34:20.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809388,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T16:45:12.057000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2022069,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-08T17:04:23.343000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1905640,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-19T07:43:22.240000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1905636,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-19T07:41:48.283000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1897383,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T17:29:01.033000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1888146,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-07T11:44:35.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1888066,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-07T10:50:06.990000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1885252,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T03:23:19.220000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1882214,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T06:17:44.413000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1879613,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-01T06:57:21.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1875602,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-29T06:28:07.747000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1860121,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-18T06:45:54.147000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1857916,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-16T14:02:53.100000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1855841,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-15T01:14:19.640000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1854199,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-13T13:50:46.067000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1852078,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-11T18:31:47.993000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1850641,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-10T15:09:56.440000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1850257,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-10T09:01:56.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1841725,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-03T12:29:39.437000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1837017,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-29T09:18:16.883000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1835827,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-28T06:04:07.780000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1835743,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-28T03:47:57.900000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1828253,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-21T15:27:47.847000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1822297,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-16T08:19:47.910000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1819178,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-13T14:02:29",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1818120,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-12T09:43:08.283000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1813554,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-07T01:52:19.403000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811961,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T10:49:23.323000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811857,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T08:09:15.117000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811656,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T01:02:58.703000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811252,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T13:41:58.773000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811059,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T08:44:38.013000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811015,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T07:09:43.700000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1810637,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T18:41:04.233000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1810206,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T10:49:42.653000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809832,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T05:05:07.353000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809737,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-03T02:33:31.897000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809660,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T23:26:42.043000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809296,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T15:29:12.960000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808951,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T09:07:23.400000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1808904,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T08:30:41.003000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807655,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T08:34:00.990000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807542,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T06:00:35.063000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807419,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-01T02:17:27.910000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807161,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-31T19:09:59.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807151,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-31T18:57:06.113000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1891175,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-09T09:24:56.383000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1862886,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-20T04:33:34.240000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2987894,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-13T06:22:15.680000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2377631,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-07T08:54:21.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1842837,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-04T10:31:25.860000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1805980": "\n# How To Reduce Data Size\nThis competition's tabular data is 50GB! That's huge. To engineer features from this data and train models with this data, we need to reduce data size and efficiently use memory and disk first. Here are some tips\n\n# Step 1 - Reduce Data Types!\nMany discussion topics discuss different file formats like Parquet, Feather, NumPy, Pickle, CSV, etc etc. This overlooks the most important point. The first step is reducing each column to the least data size possible. Afterward we can choose our file format and whether to save as multiple files or single file.\n\n### Column `customer_ID` - Reduce 64 bytes to 4 bytes!\nThis column is a string of length 64 which uses 64 bytes per row! That is too much! We can convert this `int32` or `int64` which only use 4 bytes or 8 bytes. My favorite technique is to take the last 16 letters of the hexadecimal string and convert that base16 number into base10 and save as `int64`. Discussion to explain this is [here][1]\n\n### Column `S_2` - Reduce 10 bytes to 3 bytes!\nThis column is a date with time. This column is provided as a string of length 10 which uses 10 bytes per row! This is too much! If we convert this column with `pd.to_datetime()` then it becomes only 4 bytes. Or we can save this column as three columns of `year_last_2_digits`, `month`, and `day` as `int8` each and only use 3 bytes per row.\n\n### 11 Categorical Columns - Reduce 88 bytes to 11 bytes!\nThe 11 columns `['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']` are categorical with maximum 8 values. Therefore each of these columns can be converted into `int8` which is 1 byte per row. Originally they are each 8 bytes per row.\n\n### 177 Numeric Columns - Reduce 1416 bytes to 353 bytes!\nLastly there are 177 numerical columns. These columns are `float64` with 8 bytes per row. At the bare minimum, we can convert `float64` to `float32` (4 bytes per row) without losing any important information. We are also discovering that we can convert these to `float16` which is 2 bytes per row since we suspect that Amex has added uniform noise. (And column `B_31` has only two values and can be converted to `int8` 1 byte per row)\n\n# Step 2 - Choose Your File Format\nAfter making the above size reductions, we are now ready to save files to disk. The two important properties of the different options are compression ratio and save order. Some file formats like CSV save the data row by row. And some file formats like Parquet save the data column by column. This will affect reading the data later. If in the future we want to read a subset of rows. perhaps row order is better and faster. If in the future we want to read a subset of columns, perhaps column order is better and faster.\n  \nThe second property is compression. Above we talked about number of bytes per row. The train data has `5,531,451` rows. After the reductions above, we have `4 + 3 + 11 + 353 = 371` bytes per row. Therefore our uncompressed train data size is 2GB. And test data has `11,363,762` rows, therefore uncompressed test data is 4GB. Therefore uncompressed, all the competition data becomes 6GB instead of 50GB.  Wow!\n  \nCertain file formats like Feather, Parquet, and Pickle compress the data. (And NumPy has an option `np.savez()` too). If we compress the data, then all the data can reduce to 4GB or 3GB, or 2GB. Wow! \n  \nUPDATE: A third property is whether the file format remembers the dtype. When saving as parquet, if you downcast an `int64` to `int8`, then the file format remembers this and next time you read the file it loads as `int8`. However some formats like `CSV` do not remember this. And if you save as `int8`, it will still read as `int64` unless you specify dtype in your `read_csv()` command.\n  \n# Step 3 - Choose Multiple Files or Not\nOur last decision is whether to save the data as multiple files or one large file. If we have trouble processing the entire train and test dataset at once, then we can consider processing in chunks and saving the data to disk as separate files.\n\n# Step 4 - Read Raddar's Discussion\nUPDATE: Many of the variables are actually (low cardinality) integers with noise added. Therefore for some columns, if we remove noise, we can further reduce the `float32` (4 bytes) into `int8` (1 byte) columns and reduce the data more. For more info, read Raddar's discussion post [here][5]\n\n# Conclusion\nIn conclusion, there are many options to reduce the data and save it in a new file format. Many Kagglers have posted many discussion topics and many Kaggle datasets. Which is best depends on many factors and personal preference. \n  \nWe certainly do not want to use the original 50GB CSV files provided in this competition in our pipeline. We will certainly want to do `step-1` above and then we can pick our favorite `step-2` and `step-3`. \n   \nI provide one example in my GRU starter notebook [here][2]. For `step-3`, I choose to load the original train data in 10 chunks and process each chunk separately and save as 10 separate files. First i perform `step-1` above and for `step-2`, i choose to save the chunks as uncompressed NumPy files. I uploaded the result to a Kaggle dataset [here][3] and discussion [here][4]\n\n[1]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\n[3]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\n[4]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[5]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
    "1806007": "Great tips @cdeotte as always! I've actually made some youtube videos on this exact topic. Hope it's okay to post the links here:\n\n- [This video explains how you can make your pandas dataframe more efficient](https://youtu.be/u4_c2LDi4b8) by aproprately casting the dtypes of the columns- exactly like how you discuss in this post. I also show some benchmarks showing how this can impact the dataframe size and speed.\n- [This video discusses the differences in common file formats](https://youtu.be/u4rsA5ZiTls), (parquet, feather, csv) each has it's unique positives and negatives. I also do some benchmarks to show the differences.\n\nI have other videos about pandas like this one about how to [speed up pandas code by over 2500x](https://youtu.be/SAFmrTnEHLg)!\n\nWould love to hear any feedback and/or suggestions for future videos.\n\nSorry to hijack your thread! If anyone finds these videos helpful please consider [subscribing to my channel here](https://bit.ly/3N5ygGA).",
    "1806495": "I don't know much about hex. Is there any chance that last 16 characters of the customer_ID can cause integer overflow for 64 bit integers? If I understand correctly, 64 bit signed integer type is just enough for 16 characters of hex strings, 32 bit signed integer type is just enough for 8 characters of hex strings, and it goes on like that.\n\nI also found an elegant way to convert hex string into base 10, 64 bit integer for pandas.\n`df_train['customer_ID'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)`\n\nI guess we can even push it further and convert it to 32 bit integer by using last 8 characters. We have to check number of unique values to see whether the cardinality is affected or not.\n\nEdit: You should use int64/16 characters because cardinality is affected when int32/8 characters are used for training set. \n\n```\n>>> df_train['customer_ID'].nunique()\n458913\n>>> df_train['customer_ID_int'] = df_train['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n>>> df_train['customer_ID_int'].nunique()\n458913\n>>> df_train['customer_ID_int'] = df_train['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n>>> df_train['customer_ID_int'].nunique()\n458884\n```\n\n```\n>>> df_test['customer_ID'].nunique()\n924621\n>>> df_test['customer_ID_int'] = df_test['customer_ID'].str[-16:].apply(int, base=16).astype(np.int64)\n>>> df_test['customer_ID_int'].nunique()\n924621\n>>> df_test['customer_ID_int'] = df_test['customer_ID'].str[-8:].apply(int, base=16).astype(np.int32)\n>>> df_test['customer_ID_int'].nunique()\n924547\n```",
    "1833511": "@cdeotte  Hi!Thank you for this!It is pretty helpful!Could you please elaborate on the reduction for customer_id column?",
    "1806935": "All this is good, but how are you ever going to do all these transformations without first opening the data set?\nWhich I can't open even with 16GB RAM?",
    "1819567": "Great tips, especially for those working at home on personal hardware 😅",
    "1808486": "train_df['B_31'].unique()\narray([1, 0], dtype=int64)\ntrain_df['D_87'].unique()\narray([nan,  1.], dtype=float16)",
    "1906332": "So let's say you do `df['customer_ID'] = df['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')`, then how do you get back to the original customer IDs?\n\nThe original IDs are required for a submission. I could imagine one way would be to re-load all test IDs, then convert the int versions to base 16 strings (e.g. using numpy.base_repr(id_num, 16)) and then match those to the last 16 characters of the IDs? Seems a bit roundabout but I guess there's no other way.",
    "1897922": "great discussion",
    "1856525": "Good tips!",
    "1853867": "very friendly to new guys，I learned a lot from the share，thank you very much",
    "1842847": "Thanks for sharing. however, I believe it depends on case to case. one technique may be good in one situation but not ideal for the second.",
    "1835110": "By implementing a hash, you can reduce it to 8 (or whatever resolution you want) bits per feature, at the cost of some computational speed.",
    "1834873": " A very nice tips. Thanks a lot to share this here",
    "1833909": "Thanks a lot for your great effort. what is the difference in practice between saving data in multiple files or in one file? can we say that one of them is better?",
    "1833510": "@cdeotte Hi!Thank you so much!This is extremely helpful. Could you please elaborate on transforming the customer_id column?",
    "1830087": "It's very helpful. Thanks for sharing @cdeotte 👍",
    "1828483": "@cdeotte While converting the categorical features into \"int8\", it gives an error saying cannot convert NA values to int. How do we best approach this error? Thanks",
    "1822035": "so useful!👍",
    "1820646": "Very helpful,i will definetly utilise these tricks.thanks for posting.",
    "1820276": "This is quite helpful ",
    "1819380": "@cdeotte Very useful informative post , Thanks for posting .",
    "1818814": "This is very informative. Thanks a lot for this.",
    "1818635": "Thanks a bunch for this informative post",
    "1818266": "Thanks a lot for sharing this technique dear @cdeotte .",
    "1817442": "Thanks to the people who advised on this matter.\n\nBut it's of no use for people who work with R, so here's something they might find useful:\n\ninstall.packages('sqldf')\nlibrary('sqldf')\n\nThen use the command read.csv.sql to load the file in parts:\n\nyour.data.set=read.csv.sql('train.csv', 'select [name of variable] from file where [whatever]')\n\nNote that the word \"file\" must not be changed; By \"file\" R understands the first argument, which is \"train.csv\".\n\nAnd good luck with it. I for one feel that it is not worth burning my computer to work with such huge files. And I'd like to know what American Express actually want from this competition; a few good algorithms, or handling huge data sets with home PCs?",
    "1816701": "Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!",
    "1816605": "This is great stuff Chris! I was able to reduce the size by adopting your advice! Also, I used Dask instead of Pandas. It helped me process the data fast :)",
    "1815962": "Excellent advices. Will give it a go.",
    "1815868": "Great insight @cdeotte. You have taught me a new way of looking at data. Thanks for sharing ",
    "1815644": "@cdeotte Thanks a lot for sharing this technique. It's great to learn such stuff that is useful not only for one competition but even in personal projects and professional work!",
    "1814747": "Thank you for your advices, very interesting !",
    "1812571": "great and useful insight",
    "1810120": "Very interesting!",
    "1810072": " @cdeotte Thanks for writing the detailed steps. Super helpful!",
    "1809914": "Thnks a lot!!!!",
    "1809359": "Great insight to start. Thanks for sharing",
    "1808946": "This is an amazing post. Thanks a lot for putting this up.",
    "1808710": "Quite helpful!!",
    "1808619": "This is so helpful that I will always do it, thanks!",
    "1808263": "UPDATE: Added a `Step 4` about removing noise to further reduce data size of some columns. For more info, see Raddar's disccusion [here][1]\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
    "1807697": "Very useful and interesting.\n",
    "1806905": "Very useful and interesting.\nLove you. Thanks!👍",
    "1806862": "Very informative and useful. Thank you @cdeotte",
    "1806377": "Very informative and useful. Thank you @cdeotte ",
    "1882231": "Thanks for giving us your knowledge, it seems really good",
    "1820883": "@cdeotte for the amazing post. Really helps learn practical issues when dealing with industry scale datasets",
    "1818349": "Thank you! I didn't know about these techniques! ",
    "1814712": "I have successfully loaded almost all columns as float16 so I could fit the entire training data to the provided instance with 16GB\nsee kernel: https://www.kaggle.com/code/jirkaborovec/amex-eda-baseline-lightning-flash",
    "1808622": "Thanks a lot for putting in the effort for writing this! The best part is that this technique is useful across almost any use case using tabular data.",
    "1808396": "Nice points @cdeotte! Step 4 was bugging my mind for a while too. I wonder would it be better to convert some of the noise injected features into categorical ones, otherwise we're working with small data residing in huge noise which adds only an extra memory load. On the other hand could we use the patterns in the noise injections to our benefit if there's any...",
    "1806220": "Really nice!. I am struggling with the test data, this may help!",
    "1806053": "Thanks for sharing, @cdeotte. This post is so informative and helpful; the size of these datasets makes everything more challenging in this competition, so the explanations here will help me to improve my strategy",
    "3462433": "Thanks @cdeotte for this, it's a really valuable and informative, and i learned a lot from this share.",
    "3107954": "That was really helpful ! Thankyou!\n",
    "2735011": "Great tips, especially for those working at home on not expensive personal hardware.\n\n\n",
    "2515833": "thanks a lot for your effort. learned a lot from your code and dataset :)",
    "2509919": "Thank you Chris for your in-depth explanation on reducing the data size. It has been really helpful! \n\nCan you please elaborate why you only converted the last 16 characters of the hexadecimal string into int? What if the are multiple customer_ids that are different but have the same last 16 characters? \n\nMoreover, is saving to int64 needed - is it not enough to do train['customer_id'].apply(lambda x: int( x[-16:], 16))?",
    "2150109": "These suggestions are useful.Thank you for these good suggestions",
    "1847949": "Very Very detail and usefull,  learning a lot about \"Choose Your File Format\", Thank you CHRIS,  nice to see you in this competition!",
    "1846124": "nice idea! ilike it!",
    "1843194": "@cdeotte Its simply lovely post. I have no other words. 💕💕💕",
    "1807968": "Thanks for sharing @cdeotte , Reducing data size feels like ... starting a journey of a thousand miles begins with a single step.",
    "1903422": "",
    "1855668": "",
    "1845153": "",
    "1820963": "",
    "1809388": "",
    "2022069": "Thanks for sharing!!",
    "1905640": "Thanks for your help.",
    "1905636": "Thanks for sharing.",
    "1897383": "Thanks for sharing.",
    "1888146": "Thanks for sharing.",
    "1888066": "Thanks for sharing",
    "1885252": "thanks for very useful your knowledge",
    "1882214": "Thanks for sharing",
    "1879613": "Well Explained .Thanks for sharing",
    "1875602": "Thanks for sharing.",
    "1860121": "Thanks for sharing!",
    "1857916": "Thanks for sharing. Very valuable tips",
    "1855841": "Thanks for this post!",
    "1854199": "Thanks for sharing!",
    "1852078": "Thanks for sharing!!",
    "1850641": "thank you so much!!",
    "1850257": "Thanks for sharing!",
    "1841725": "This is good idea. Thanks so much!\n\n",
    "1837017": "Thank you o much for sharing this ",
    "1835827": "This is very efficient, thanks!",
    "1835743": "It is very helpful, Thanks !!",
    "1828253": "This is gold. Thanks so much :)",
    "1822297": "Thanks for sharing these techniques",
    "1819178": "thanks for tips",
    "1818120": "Thanks for sharing",
    "1813554": "Thanks a lot.",
    "1811961": "thanks a lot ",
    "1811857": "****Thanks, useful insight..!****",
    "1811656": "Thanks, this was very helpful!",
    "1811252": "Thanks a lot",
    "1811059": "Thanks a lot.",
    "1811015": "Thanks for sharing",
    "1810637": "Thanks, great start, keep adding more.",
    "1810206": "thanks a lot.",
    "1809832": "Such a good way, thanks a lot.",
    "1809737": "Thanks a lot",
    "1809660": "Thank you.",
    "1809296": "Thanks a lot ",
    "1808951": "Thanks a lot for this insight.",
    "1808904": "Thanks a lot <",
    "1807655": "Thanks for sharing the info!👍",
    "1807542": "Thanks a lot",
    "1807419": "very good!!\nThanks a lot!!",
    "1807161": "Thanks a lot!!",
    "1807151": "Thank you!! @cdeotte ",
    "1891175": "thanks for your share",
    "1862886": "Thanks for sharing!!!",
    "2987894": "Very helpful, thank you!",
    "2377631": "Very Useful. Thanks",
    "1842837": "Thanks for sharing. very informative. "
  }
}