{
  "id": 308635,
  "title": "Memory Trick - Reduce Memory 8x or 16x!",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635",
  "author_name": "Chris Deotte",
  "post_date": "2022-02-19T14:59:54.288000",
  "votes": 337,
  "comment_count": 42,
  "views": 0,
  "content": "<p>I would like to share a memory trick that <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> taught me when we competed together in the RecSys-Twitter competition in 2020 and 2021. This memory trick will reduce the training data memory by 5x! (from 96 bytes per row to 20 bytes per row)</p>\n<h1>Reduce Customer_Id by 8x (or 16x)!</h1>\n<pre><code>import pandas as pd\ntrain = pd.read_csv('transactions_train.csv')\ntrain['customer_id'] =\\\n    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n</code></pre>\n<p>The <code>customer_id</code> is a length 64 string which uses 64 bytes. The code above coverts the column to int64 which only takes 8 bytes! Also the mapping is 1:1 which means each customer gets a unique <code>int64</code>. After you are done with your processing, i.e. <code>train.groupby('customer_id')</code> etc etc, just merge this onto the <code>sample_submission.csv</code> dataframe.</p>\n<p>When using <strong>RAPIDS cuDF</strong>, do this</p>\n<pre><code>import cudf\ntrain = cudf.read_csv('transactions_train.csv')\ntrain['customer_id'] =\\\n    train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n</code></pre>\n<h1>UPDATE: Achieve 16x Reduction</h1>\n<p>See comments below from Clear n' Simple <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> on how to achieve <strong>16x</strong> with a mapping.<br>\nSee comments below from Giba, to learn how to achieve <strong>16x</strong> memory reduction with <code>int32</code>!</p>\n<h1>Reduce Article_Id by 2.5x!</h1>\n<pre><code>train['article_id'] = train['article_id'].astype('int32')\n</code></pre>\n<p>Many people store the <code>article_id</code> as a string to keep the leading zero. However a string of length 10 uses 10 bytes. It is better to convert this column to <code>int32</code> which only uses 4 bytes. Then do all your processing, i.e. <code>train.groupby('article_id')</code> etc etc. Then before writing submission, just convert it back to string with zero</p>\n<pre><code>train['article_id'] = '0' + train.article_id.astype('str')\n</code></pre>\n<h1>Reduce Other Columns 3x!</h1>\n<p>Lastly you can reduce or remove the other columns</p>\n<pre><code>train.t_dat = cudf.to_datetime( train.t_dat )\ntrain['year'] = (train.t_dat.dt.year-2000).astype('int8')\ntrain['month'] = (train.t_dat.dt.month).astype('int8')\ntrain['day'] = (train.t_dat.dt.day).astype('int8')\ndel train['t_dat']\n</code></pre>\n<p>And</p>\n<pre><code>train['price'] = train['price'].astype('float32')\ntrain['sales_channel_id'] = train['sales_channel_id'].astype('int8')\n</code></pre>\n<h1>Example</h1>\n<p>I post an example of using <code>customer_id</code> and <code>article_id</code> memory reduction <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a>. The basic idea is this. After you make predictions for customers using the new int64 id, we merge it back with:</p>\n<pre><code>sub = cudf.read_csv('sample_submission.csv')[['customer_id']]\nsub['customer_id_2'] =\\\n    sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.merge(PREDS_DF.rename({'customer_id':'customer_id_2'},axis=1),\\\n    on='customer_id_2', how='left').fillna('')\ndel sub['customer_id_2']\nsub.to_csv('submission.csv',index=False)\n</code></pre>\n<h1>UPDATE -  Starter Notebook</h1>\n<p>I published a starter notebook <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> which uses the above memory reductions. Then the entire solution uses dataframe manipulations on the reduced dataframe. Finally the predictions are merged onto <code>sample_submission.csv</code> and submitted to Kaggle to achieve LB 0.021</p>",
  "messages": [
    {
      "id": 1697346,
      "postDate": "2022-02-19T14:59:54.287Z",
      "content": "<p>I would like to share a memory trick that <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> taught me when we competed together in the RecSys-Twitter competition in 2020 and 2021. This memory trick will reduce the training data memory by 5x! (from 96 bytes per row to 20 bytes per row)</p>\n<h1>Reduce Customer_Id by 8x (or 16x)!</h1>\n<pre><code>import pandas as pd\ntrain = pd.read_csv('transactions_train.csv')\ntrain['customer_id'] =\\\n    train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n</code></pre>\n<p>The <code>customer_id</code> is a length 64 string which uses 64 bytes. The code above coverts the column to int64 which only takes 8 bytes! Also the mapping is 1:1 which means each customer gets a unique <code>int64</code>. After you are done with your processing, i.e. <code>train.groupby('customer_id')</code> etc etc, just merge this onto the <code>sample_submission.csv</code> dataframe.</p>\n<p>When using <strong>RAPIDS cuDF</strong>, do this</p>\n<pre><code>import cudf\ntrain = cudf.read_csv('transactions_train.csv')\ntrain['customer_id'] =\\\n    train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n</code></pre>\n<h1>UPDATE: Achieve 16x Reduction</h1>\n<p>See comments below from Clear n' Simple <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> on how to achieve <strong>16x</strong> with a mapping.<br>\nSee comments below from Giba, to learn how to achieve <strong>16x</strong> memory reduction with <code>int32</code>!</p>\n<h1>Reduce Article_Id by 2.5x!</h1>\n<pre><code>train['article_id'] = train['article_id'].astype('int32')\n</code></pre>\n<p>Many people store the <code>article_id</code> as a string to keep the leading zero. However a string of length 10 uses 10 bytes. It is better to convert this column to <code>int32</code> which only uses 4 bytes. Then do all your processing, i.e. <code>train.groupby('article_id')</code> etc etc. Then before writing submission, just convert it back to string with zero</p>\n<pre><code>train['article_id'] = '0' + train.article_id.astype('str')\n</code></pre>\n<h1>Reduce Other Columns 3x!</h1>\n<p>Lastly you can reduce or remove the other columns</p>\n<pre><code>train.t_dat = cudf.to_datetime( train.t_dat )\ntrain['year'] = (train.t_dat.dt.year-2000).astype('int8')\ntrain['month'] = (train.t_dat.dt.month).astype('int8')\ntrain['day'] = (train.t_dat.dt.day).astype('int8')\ndel train['t_dat']\n</code></pre>\n<p>And</p>\n<pre><code>train['price'] = train['price'].astype('float32')\ntrain['sales_channel_id'] = train['sales_channel_id'].astype('int8')\n</code></pre>\n<h1>Example</h1>\n<p>I post an example of using <code>customer_id</code> and <code>article_id</code> memory reduction <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a>. The basic idea is this. After you make predictions for customers using the new int64 id, we merge it back with:</p>\n<pre><code>sub = cudf.read_csv('sample_submission.csv')[['customer_id']]\nsub['customer_id_2'] =\\\n    sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.merge(PREDS_DF.rename({'customer_id':'customer_id_2'},axis=1),\\\n    on='customer_id_2', how='left').fillna('')\ndel sub['customer_id_2']\nsub.to_csv('submission.csv',index=False)\n</code></pre>\n<h1>UPDATE -  Starter Notebook</h1>\n<p>I published a starter notebook <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> which uses the above memory reductions. Then the entire solution uses dataframe manipulations on the reduced dataframe. Finally the predictions are merged onto <code>sample_submission.csv</code> and submitted to Kaggle to achieve LB 0.021</p>",
      "rawMarkdown": "I would like to share a memory trick that @titericz taught me when we competed together in the RecSys-Twitter competition in 2020 and 2021. This memory trick will reduce the training data memory by 5x! (from 96 bytes per row to 20 bytes per row)\n\n# Reduce Customer_Id by 8x (or 16x)!\n\n    import pandas as pd\n    train = pd.read_csv('transactions_train.csv')\n    train['customer_id'] =\\\n        train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n\nThe `customer_id` is a length 64 string which uses 64 bytes. The code above coverts the column to int64 which only takes 8 bytes! Also the mapping is 1:1 which means each customer gets a unique `int64`. After you are done with your processing, i.e. `train.groupby('customer_id')` etc etc, just merge this onto the `sample_submission.csv` dataframe.\n\nWhen using **RAPIDS cuDF**, do this\n\n    import cudf\n    train = cudf.read_csv('transactions_train.csv')\n    train['customer_id'] =\\\n        train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n\n# UPDATE: Achieve 16x Reduction\nSee comments below from Clear n' Simple @jacob34 on how to achieve **16x** with a mapping.\nSee comments below from Giba, to learn how to achieve **16x** memory reduction with `int32`!\n\n# Reduce Article_Id by 2.5x!\n\n    train['article_id'] = train['article_id'].astype('int32')\n\nMany people store the `article_id` as a string to keep the leading zero. However a string of length 10 uses 10 bytes. It is better to convert this column to `int32` which only uses 4 bytes. Then do all your processing, i.e. `train.groupby('article_id')` etc etc. Then before writing submission, just convert it back to string with zero\n\n    train['article_id'] = '0' + train.article_id.astype('str')\n\n# Reduce Other Columns 3x!\nLastly you can reduce or remove the other columns\n\n    train.t_dat = cudf.to_datetime( train.t_dat )\n    train['year'] = (train.t_dat.dt.year-2000).astype('int8')\n    train['month'] = (train.t_dat.dt.month).astype('int8')\n    train['day'] = (train.t_dat.dt.day).astype('int8')\n    del train['t_dat']\n\nAnd\n\n    train['price'] = train['price'].astype('float32')\n    train['sales_channel_id'] = train['sales_channel_id'].astype('int8')\n\n# Example\nI post an example of using `customer_id` and `article_id` memory reduction [here][2] and [here][1]. The basic idea is this. After you make predictions for customers using the new int64 id, we merge it back with:\n\n    sub = cudf.read_csv('sample_submission.csv')[['customer_id']]\n    sub['customer_id_2'] =\\\n        sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n    sub = sub.merge(PREDS_DF.rename({'customer_id':'customer_id_2'},axis=1),\\\n        on='customer_id_2', how='left').fillna('')\n    del sub['customer_id_2']\n    sub.to_csv('submission.csv',index=False)\n\n# UPDATE -  Starter Notebook\nI published a starter notebook [here][2] which uses the above memory reductions. Then the entire solution uses dataframe manipulations on the reduced dataframe. Finally the predictions are merged onto `sample_submission.csv` and submitted to Kaggle to achieve LB 0.021\n\n[1]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[2]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021",
      "votes": 335
    },
    {
      "id": 1698725,
      "postDate": "2022-02-20T15:59:07.597Z",
      "content": "<p>Thank you!<br>\nThe customer_id code is very cool.</p>\n<p>Is there a reason not to use the index in the customers csv file?</p>\n<pre><code>id_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\n# for memory efficiency\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# for switching back for submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n</code></pre>",
      "rawMarkdown": "Thank you!\nThe customer_id code is very cool.\n\nIs there a reason not to use the index in the customers csv file?\n```\nid_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\n# for memory efficiency\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# for switching back for submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n```",
      "votes": 20,
      "replies": [
        {
          "id": 1698774,
          "postDate": "2022-02-20T16:40:17.540Z",
          "content": "<p>Great suggestion. Your code is the most efficient. It achieves <strong>16x</strong> because after</p>\n<pre><code>transactions[\"customer_id\"] =\\\n    transactions[\"customer_id\"].map(id_to_index_dict)\n</code></pre>\n<p>you can cast to <code>int32</code> and have unique ids for each customer. </p>\n<pre><code>transactions[\"customer_id\"] = transactions[\"customer_id\"].astype('int32')\n</code></pre>\n<p>The advantage of my method is that there is no need for a map (So it's simple. ). But the most important thing is to reduce memory size, so your method is best. </p>",
          "rawMarkdown": "Great suggestion. Your code is the most efficient. It achieves **16x** because after\n\n    transactions[\"customer_id\"] =\\\n        transactions[\"customer_id\"].map(id_to_index_dict)\n\nyou can cast to `int32` and have unique ids for each customer. \n\n    transactions[\"customer_id\"] = transactions[\"customer_id\"].astype('int32')\n\nThe advantage of my method is that there is no need for a map (So it's simple. ~~And it's easy to use with RAPIDS cuDF which doesn't implement mapping a column with a dictionary yet. With RAPIDS you would need to merge it on~~). But the most important thing is to reduce memory size, so your method is best. ",
          "votes": 10
        },
        {
          "id": 1713446,
          "postDate": "2022-03-06T03:20:27.893Z",
          "content": "<blockquote>\n  <p>RAPIDS cuDF doesn't implement mapping a column with a dictionary yet.</p>\n</blockquote>\n<p>Just tested it now, and it worked for me.<br>\nIs this something that just happened?</p>",
          "rawMarkdown": "> RAPIDS cuDF doesn't implement mapping a column with a dictionary yet.\n\nJust tested it now, and it worked for me.\nIs this something that just happened?",
          "votes": 2
        },
        {
          "id": 1713451,
          "postDate": "2022-03-06T03:27:17.880Z",
          "content": "<p>Awesome, you are right it works. Thanks for pointing this out.</p>",
          "rawMarkdown": "Awesome, you are right it works. Thanks for pointing this out.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1697752,
      "postDate": "2022-02-19T20:10:46.737Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Great post and I appreciate the mention. <br>\nNote that sometimes is possible to compress even more 'customer_id' by using int32 when it don't affect much the cardinality.</p>",
      "rawMarkdown": "@cdeotte Great post and I appreciate the mention. \nNote that sometimes is possible to compress even more 'customer_id' by using int32 when it don't affect much the cardinality.",
      "votes": 6,
      "replies": [
        {
          "id": 1697761,
          "postDate": "2022-02-19T20:20:32.813Z",
          "content": "<p>Yes, great point.</p>\n<p>What Giba is referring to is using more compression. We can reduce <code>customer_id</code> by <strong>16x</strong>! (instead of 8x) with the following code</p>\n<pre><code>train['customer_id'] =\\\n    train['customer_id'].apply(lambda x: int(x[-8:],16) ).astype('int32')\n</code></pre>\n<p>The above code converts 64 bytes of string into 4 bytes of <code>int32</code>. The original <code>customer_id</code> has <code>1,362,281</code> unique users. If we hash the string customer into <code>int32</code> then afterward there are <code>1,362,059</code> unique users, so we see that 222 users got mapped to the same hash value as another user. (These are called hash collisions). </p>\n<p>This is not many collisions (only 0.02%), so we can consider using <code>int32</code> and just ignoring these 222 collisions. And consequently this will reduce the memory by another 2x! That's huge, so we should consider this. (If we use <code>int64</code> we have zero collisions and 1:1 mapping).</p>",
          "rawMarkdown": "Yes, great point.\n\nWhat Giba is referring to is using more compression. We can reduce `customer_id` by **16x**! (instead of 8x) with the following code\n\n    train['customer_id'] =\\\n        train['customer_id'].apply(lambda x: int(x[-8:],16) ).astype('int32')\n\nThe above code converts 64 bytes of string into 4 bytes of `int32`. The original `customer_id` has `1,362,281` unique users. If we hash the string customer into `int32` then afterward there are `1,362,059` unique users, so we see that 222 users got mapped to the same hash value as another user. (These are called hash collisions). \n\nThis is not many collisions (only 0.02%), so we can consider using `int32` and just ignoring these 222 collisions. And consequently this will reduce the memory by another 2x! That's huge, so we should consider this. (If we use `int64` we have zero collisions and 1:1 mapping).",
          "votes": 11
        },
        {
          "id": 1698881,
          "postDate": "2022-02-20T18:29:08.557Z",
          "content": "<p>I believe it is cleaner to use a separate index (String -&gt; int32) for customers and don't risk any collision. It is enough to do it once when the data is loaded and then replace them for submission.</p>",
          "rawMarkdown": "I believe it is cleaner to use a separate index (String -> int32) for customers and don't risk any collision. It is enough to do it once when the data is loaded and then replace them for submission.",
          "votes": 9
        },
        {
          "id": 1699005,
          "postDate": "2022-02-20T20:26:48.690Z",
          "content": "<p>I like the simplicity of directly converting hexadecimal digits to integer. However, I agree it would be best to just use a one-to-one mapping <code>String -&gt; Int32</code> to avoid collision and achieve maximum reduction.</p>",
          "rawMarkdown": "I like the simplicity of directly converting hexadecimal digits to integer. However, I agree it would be best to just use a one-to-one mapping `String -> Int32` to avoid collision and achieve maximum reduction.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1699238,
      "postDate": "2022-02-21T03:35:56.027Z",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> which demonstrates the use of the above dataframe memory reduction. After memory reduction, the complete solution uses dataframe operations and then merges the predictions onto sample_submission.csv before submitting to Kaggle. The notebook achieves LB 0.021</p>",
      "rawMarkdown": "UPDATE: I posted a starter notebook [here][1] which demonstrates the use of the above dataframe memory reduction. After memory reduction, the complete solution uses dataframe operations and then merges the predictions onto sample_submission.csv before submitting to Kaggle. The notebook achieves LB 0.021\n\n[1]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021",
      "votes": 3
    },
    {
      "id": 1697378,
      "postDate": "2022-02-19T15:22:50.323Z",
      "content": "<p>Thanks so much for sharing this!</p>\n<p>I believe you had named the mapping trick as \"reverse target encoding\"? Sorry If I'm mixing the two concepts up.</p>\n<p>Here's a timestamped <a href=\"https://youtu.be/W3aWEXqIkWk?t=3797\" target=\"_blank\">link</a> where Chris spoke about this.</p>",
      "rawMarkdown": "Thanks so much for sharing this!\n\nI believe you had named the mapping trick as \"reverse target encoding\"? Sorry If I'm mixing the two concepts up.\n\nHere's a timestamped [link](https://youtu.be/W3aWEXqIkWk?t=3797) where Chris spoke about this.",
      "votes": 3,
      "replies": [
        {
          "id": 1697430,
          "postDate": "2022-02-19T16:09:24.197Z",
          "content": "<p>Thanks for posting a video to team Nvidia's 2021 RecSys winning solution. Reverse TE is a potential second step. This discussion's memory trick is just the first step.</p>\n<p>When we do TE or any other feature engineering, we will execute dataframe operations. Mainly they are <code>groupby</code> operations. Before we can do dataframe operations fast and efficiently, we need to reduce the dataframe memory size as much as possible. </p>",
          "rawMarkdown": "Thanks for posting a video to team Nvidia's 2021 RecSys winning solution. Reverse TE is a potential second step. This discussion's memory trick is just the first step.\n\nWhen we do TE or any other feature engineering, we will execute dataframe operations. Mainly they are `groupby` operations. Before we can do dataframe operations fast and efficiently, we need to reduce the dataframe memory size as much as possible. ",
          "votes": 3
        },
        {
          "id": 1697447,
          "postDate": "2022-02-19T16:17:52.850Z",
          "content": "<p>Sorry, I didnt run the lambda and skimmed assuming it was also applying grouping. This makes perfect sense! </p>\n<p>Thanks for answering my noob question 🙏</p>",
          "rawMarkdown": "Sorry, I didnt run the lambda and skimmed assuming it was also applying grouping. This makes perfect sense! \n\nThanks for answering my noob question 🙏",
          "votes": 2
        }
      ]
    },
    {
      "id": 1706719,
      "postDate": "2022-02-27T18:14:06.480Z",
      "content": "<p>Thanks for sharing. Really useful tricks !!</p>",
      "rawMarkdown": "Thanks for sharing. Really useful tricks !!",
      "votes": 1
    },
    {
      "id": 1704284,
      "postDate": "2022-02-25T11:27:18.607Z",
      "content": "<p>Thank you very much!, that has been helpful for me . </p>",
      "rawMarkdown": "Thank you very much!, that has been helpful for me . ",
      "votes": 1
    },
    {
      "id": 1702252,
      "postDate": "2022-02-23T13:16:03.393Z",
      "content": "<p>I didn't have this idea, I will use it for analysis!</p>",
      "rawMarkdown": "I didn't have this idea, I will use it for analysis!",
      "votes": 1
    },
    {
      "id": 1701288,
      "postDate": "2022-02-22T16:25:38.873Z",
      "content": "<p>wow i never thought about how you can use such tricks in order to increase the speed. Thanks for sharing !!<br>\nall of this falls back to old school way of handling data, when I was working for an application for TV set top box, I had to think so much for memory optimization, and the things never changes.. Memory optimization is always important </p>",
      "rawMarkdown": "wow i never thought about how you can use such tricks in order to increase the speed. Thanks for sharing !!\nall of this falls back to old school way of handling data, when I was working for an application for TV set top box, I had to think so much for memory optimization, and the things never changes.. Memory optimization is always important ",
      "votes": 1,
      "replies": [
        {
          "id": 1701294,
          "postDate": "2022-02-22T16:33:27.960Z",
          "content": "<p>Absolutely. Even if memory can handle large dataset, the first thing I do in every competition is reduce the data to the smallest memory types. Because this makes all subsequent computation faster.</p>",
          "rawMarkdown": "Absolutely. Even if memory can handle large dataset, the first thing I do in every competition is reduce the data to the smallest memory types. Because this makes all subsequent computation faster.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1701135,
      "postDate": "2022-02-22T14:39:22.613Z",
      "content": "<p>Very tricky, thanks for sharing!</p>",
      "rawMarkdown": "Very tricky, thanks for sharing!",
      "votes": 1,
      "replies": [
        {
          "id": 1701152,
          "postDate": "2022-02-22T14:50:25.770Z",
          "content": "<p>Exactly. That's why i like this trick. We could of course just label encode the column with <code>labels, codes = df.customer_id.factorize()</code> and save the labels and codes as a map. But the original customer_id are hexadecimal digits and only the last 16 digits are needed to identify unique customers. The beginning 48 digits are unnecessary.</p>",
          "rawMarkdown": "Exactly. That's why i like this trick. We could of course just label encode the column with `labels, codes = df.customer_id.factorize()` and save the labels and codes as a map. But the original customer_id are hexadecimal digits and only the last 16 digits are needed to identify unique customers. The beginning 48 digits are unnecessary.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1700678,
      "postDate": "2022-02-22T06:33:46.813Z",
      "content": "<p>Excellent work, as always 👏🏻 Thank you very much!</p>",
      "rawMarkdown": "Excellent work, as always 👏🏻 Thank you very much!",
      "votes": 1
    },
    {
      "id": 1699426,
      "postDate": "2022-02-21T07:10:09.697Z",
      "content": "<p>Thanks for sharing, it is really useful👍</p>",
      "rawMarkdown": "Thanks for sharing, it is really useful👍",
      "votes": 1
    },
    {
      "id": 1698516,
      "postDate": "2022-02-20T12:59:25.650Z",
      "content": "<p>Nice one! For further memory usage, I would also suggest using this trick:</p>\n<pre><code>pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\",\n            dtype={\"t_dat\": \"object\", \"customer_id\": \"object\", \"article_id\": \"object\", \"price\": float, \"sales_channel_id\": int}\n           )\n</code></pre>",
      "rawMarkdown": "Nice one! For further memory usage, I would also suggest using this trick:\n\n```python\npd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\",\n            dtype={\"t_dat\": \"object\", \"customer_id\": \"object\", \"article_id\": \"object\", \"price\": float, \"sales_channel_id\": int}\n           )\n```",
      "votes": 2
    },
    {
      "id": 1806684,
      "postDate": "2022-05-31T12:02:42.987Z",
      "content": "<p>Cool! Helps much!</p>",
      "rawMarkdown": "Cool! Helps much!"
    },
    {
      "id": 1773449,
      "postDate": "2022-05-01T07:11:26.353Z",
      "content": "<p>Great job.  It really helped me get past a lot of them hurdles.  Went from using 16g of ram to only 5.  Thanks.  </p>\n<p>Of course I did other things as well, but this was extra helpful.💪👍</p>",
      "rawMarkdown": "Great job.  It really helped me get past a lot of them hurdles.  Went from using 16g of ram to only 5.  Thanks.  \n\nOf course I did other things as well, but this was extra helpful.💪👍"
    },
    {
      "id": 1702073,
      "postDate": "2022-02-23T09:54:48.843Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1754767,
      "postDate": "2022-04-14T02:03:37.703Z",
      "content": "<p>a nice share. thanks</p>",
      "rawMarkdown": "a nice share. thanks\n",
      "votes": 1
    },
    {
      "id": 1728523,
      "postDate": "2022-03-19T01:51:16.260Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1720883,
      "postDate": "2022-03-13T08:35:58.803Z",
      "content": "<p>Thanks for sharing. Nice tip! :)</p>",
      "rawMarkdown": "Thanks for sharing. Nice tip! :)",
      "votes": 1
    },
    {
      "id": 1715414,
      "postDate": "2022-03-08T01:27:02.563Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1714838,
      "postDate": "2022-03-07T11:56:18.157Z",
      "content": "<p>Thanks for sharing！</p>",
      "rawMarkdown": "Thanks for sharing！",
      "votes": 1
    },
    {
      "id": 1711840,
      "postDate": "2022-03-04T11:11:38.983Z",
      "content": "<p>Thanks for sharing it really useful!</p>",
      "rawMarkdown": "Thanks for sharing it really useful!",
      "votes": 1
    },
    {
      "id": 1710583,
      "postDate": "2022-03-03T06:44:20.397Z",
      "content": "<p>Great Thank you for sharing..</p>",
      "rawMarkdown": "Great Thank you for sharing..",
      "votes": 1
    },
    {
      "id": 1708838,
      "postDate": "2022-03-01T18:29:11.257Z",
      "content": "<p>Thanks for sharing👍</p>",
      "rawMarkdown": "Thanks for sharing👍",
      "votes": 1
    },
    {
      "id": 1705536,
      "postDate": "2022-02-26T15:20:26.140Z",
      "content": "<p>Thanks alot!</p>",
      "rawMarkdown": "Thanks alot!",
      "votes": 1
    },
    {
      "id": 1701803,
      "postDate": "2022-02-23T04:40:59.067Z",
      "content": "<p>Awesome! Thanks for sharing</p>",
      "rawMarkdown": "Awesome! Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1700832,
      "postDate": "2022-02-22T09:35:22.373Z",
      "content": "<p>Thank you!! 👍</p>",
      "rawMarkdown": "Thank you!! 👍",
      "votes": 1
    },
    {
      "id": 1700467,
      "postDate": "2022-02-22T02:51:14.717Z",
      "content": "<p>Thank you for sharing this tip!</p>",
      "rawMarkdown": "Thank you for sharing this tip!",
      "votes": 1
    },
    {
      "id": 1699324,
      "postDate": "2022-02-21T05:27:28.810Z",
      "content": "<p>Thank you,really helpful.</p>",
      "rawMarkdown": "Thank you,really helpful.",
      "votes": 1
    },
    {
      "id": 1699135,
      "postDate": "2022-02-21T00:16:52.457Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": 2
    },
    {
      "id": 1697407,
      "postDate": "2022-02-19T15:57:02.137Z",
      "content": "<p>Nice Tricks ! 👋<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Nice Tricks ! 👋\nThanks for sharing.",
      "votes": 2
    },
    {
      "id": 1780918,
      "postDate": "2022-05-08T03:32:04.157Z",
      "content": "<p>Really nice. Thanks alot.</p>",
      "rawMarkdown": "Really nice. Thanks alot."
    },
    {
      "id": 1700687,
      "postDate": "2022-02-22T06:52:36.227Z",
      "content": "<p>Really helpful. Thanks for sharing</p>",
      "rawMarkdown": "Really helpful. Thanks for sharing",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1698725,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-02-20T15:59:07.597000",
      "content": "<p>Thank you!<br>\nThe customer_id code is very cool.</p>\n<p>Is there a reason not to use the index in the customers csv file?</p>\n<pre><code>id_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\n# for memory efficiency\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# for switching back for submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n</code></pre>",
      "votes": 20,
      "replies": [
        {
          "id": 1698774,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-20T16:40:17.540000",
          "content": "<p>Great suggestion. Your code is the most efficient. It achieves <strong>16x</strong> because after</p>\n<pre><code>transactions[\"customer_id\"] =\\\n    transactions[\"customer_id\"].map(id_to_index_dict)\n</code></pre>\n<p>you can cast to <code>int32</code> and have unique ids for each customer. </p>\n<pre><code>transactions[\"customer_id\"] = transactions[\"customer_id\"].astype('int32')\n</code></pre>\n<p>The advantage of my method is that there is no need for a map (So it's simple. ). But the most important thing is to reduce memory size, so your method is best. </p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1713446,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-03-06T03:20:27.893000",
          "content": "<blockquote>\n  <p>RAPIDS cuDF doesn't implement mapping a column with a dictionary yet.</p>\n</blockquote>\n<p>Just tested it now, and it worked for me.<br>\nIs this something that just happened?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1713451,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-06T03:27:17.880000",
          "content": "<p>Awesome, you are right it works. Thanks for pointing this out.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1697752,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2022-02-19T20:10:46.737000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Great post and I appreciate the mention. <br>\nNote that sometimes is possible to compress even more 'customer_id' by using int32 when it don't affect much the cardinality.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1697761,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-19T20:20:32.813000",
          "content": "<p>Yes, great point.</p>\n<p>What Giba is referring to is using more compression. We can reduce <code>customer_id</code> by <strong>16x</strong>! (instead of 8x) with the following code</p>\n<pre><code>train['customer_id'] =\\\n    train['customer_id'].apply(lambda x: int(x[-8:],16) ).astype('int32')\n</code></pre>\n<p>The above code converts 64 bytes of string into 4 bytes of <code>int32</code>. The original <code>customer_id</code> has <code>1,362,281</code> unique users. If we hash the string customer into <code>int32</code> then afterward there are <code>1,362,059</code> unique users, so we see that 222 users got mapped to the same hash value as another user. (These are called hash collisions). </p>\n<p>This is not many collisions (only 0.02%), so we can consider using <code>int32</code> and just ignoring these 222 collisions. And consequently this will reduce the memory by another 2x! That's huge, so we should consider this. (If we use <code>int64</code> we have zero collisions and 1:1 mapping).</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1698881,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-20T18:29:08.557000",
          "content": "<p>I believe it is cleaner to use a separate index (String -&gt; int32) for customers and don't risk any collision. It is enough to do it once when the data is loaded and then replace them for submission.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1699005,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-20T20:26:48.690000",
          "content": "<p>I like the simplicity of directly converting hexadecimal digits to integer. However, I agree it would be best to just use a one-to-one mapping <code>String -&gt; Int32</code> to avoid collision and achieve maximum reduction.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1699238,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-02-21T03:35:56.027000",
      "content": "<p>UPDATE: I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">here</a> which demonstrates the use of the above dataframe memory reduction. After memory reduction, the complete solution uses dataframe operations and then merges the predictions onto sample_submission.csv before submitting to Kaggle. The notebook achieves LB 0.021</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1697378,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2022-02-19T15:22:50.323000",
      "content": "<p>Thanks so much for sharing this!</p>\n<p>I believe you had named the mapping trick as \"reverse target encoding\"? Sorry If I'm mixing the two concepts up.</p>\n<p>Here's a timestamped <a href=\"https://youtu.be/W3aWEXqIkWk?t=3797\" target=\"_blank\">link</a> where Chris spoke about this.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1697430,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-19T16:09:24.197000",
          "content": "<p>Thanks for posting a video to team Nvidia's 2021 RecSys winning solution. Reverse TE is a potential second step. This discussion's memory trick is just the first step.</p>\n<p>When we do TE or any other feature engineering, we will execute dataframe operations. Mainly they are <code>groupby</code> operations. Before we can do dataframe operations fast and efficiently, we need to reduce the dataframe memory size as much as possible. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1697447,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-19T16:17:52.850000",
          "content": "<p>Sorry, I didnt run the lambda and skimmed assuming it was also applying grouping. This makes perfect sense! </p>\n<p>Thanks for answering my noob question 🙏</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1706719,
      "author_name": "Sanskriti Agrawal",
      "author_url": "",
      "post_date": "2022-02-27T18:14:06.480000",
      "content": "<p>Thanks for sharing. Really useful tricks !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1704284,
      "author_name": "Mohamed Elsayed",
      "author_url": "",
      "post_date": "2022-02-25T11:27:18.607000",
      "content": "<p>Thank you very much!, that has been helpful for me . </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1702252,
      "author_name": "HIRO",
      "author_url": "",
      "post_date": "2022-02-23T13:16:03.393000",
      "content": "<p>I didn't have this idea, I will use it for analysis!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1701288,
      "author_name": "Umashankar Somasekar",
      "author_url": "",
      "post_date": "2022-02-22T16:25:38.873000",
      "content": "<p>wow i never thought about how you can use such tricks in order to increase the speed. Thanks for sharing !!<br>\nall of this falls back to old school way of handling data, when I was working for an application for TV set top box, I had to think so much for memory optimization, and the things never changes.. Memory optimization is always important </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1701294,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-22T16:33:27.960000",
          "content": "<p>Absolutely. Even if memory can handle large dataset, the first thing I do in every competition is reduce the data to the smallest memory types. Because this makes all subsequent computation faster.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1701135,
      "author_name": "Neesham",
      "author_url": "",
      "post_date": "2022-02-22T14:39:22.613000",
      "content": "<p>Very tricky, thanks for sharing!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1701152,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-22T14:50:25.770000",
          "content": "<p>Exactly. That's why i like this trick. We could of course just label encode the column with <code>labels, codes = df.customer_id.factorize()</code> and save the labels and codes as a map. But the original customer_id are hexadecimal digits and only the last 16 digits are needed to identify unique customers. The beginning 48 digits are unnecessary.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1700678,
      "author_name": "FPiotro",
      "author_url": "",
      "post_date": "2022-02-22T06:33:46.813000",
      "content": "<p>Excellent work, as always 👏🏻 Thank you very much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1699426,
      "author_name": "Ravi_kr",
      "author_url": "",
      "post_date": "2022-02-21T07:10:09.697000",
      "content": "<p>Thanks for sharing, it is really useful👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1698516,
      "author_name": "Shion Honda",
      "author_url": "",
      "post_date": "2022-02-20T12:59:25.650000",
      "content": "<p>Nice one! For further memory usage, I would also suggest using this trick:</p>\n<pre><code>pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\",\n            dtype={\"t_dat\": \"object\", \"customer_id\": \"object\", \"article_id\": \"object\", \"price\": float, \"sales_channel_id\": int}\n           )\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1806684,
      "author_name": "Narendiranath",
      "author_url": "",
      "post_date": "2022-05-31T12:02:42.987000",
      "content": "<p>Cool! Helps much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1773449,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-01T07:11:26.353000",
      "content": "<p>Great job.  It really helped me get past a lot of them hurdles.  Went from using 16g of ram to only 5.  Thanks.  </p>\n<p>Of course I did other things as well, but this was extra helpful.💪👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1702073,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-23T09:54:48.843000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1754767,
      "author_name": "CunchiLv",
      "author_url": "",
      "post_date": "2022-04-14T02:03:37.703000",
      "content": "<p>a nice share. thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1728523,
      "author_name": "aop",
      "author_url": "",
      "post_date": "2022-03-19T01:51:16.260000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1720883,
      "author_name": "Prakarn Kidngun",
      "author_url": "",
      "post_date": "2022-03-13T08:35:58.803000",
      "content": "<p>Thanks for sharing. Nice tip! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1715414,
      "author_name": "Cafelatte1",
      "author_url": "",
      "post_date": "2022-03-08T01:27:02.563000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1714838,
      "author_name": "TapTap~",
      "author_url": "",
      "post_date": "2022-03-07T11:56:18.157000",
      "content": "<p>Thanks for sharing！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1711840,
      "author_name": "Artem Burenok",
      "author_url": "",
      "post_date": "2022-03-04T11:11:38.983000",
      "content": "<p>Thanks for sharing it really useful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1710583,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-03-03T06:44:20.397000",
      "content": "<p>Great Thank you for sharing..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1708838,
      "author_name": "Ishan Mehta115",
      "author_url": "",
      "post_date": "2022-03-01T18:29:11.257000",
      "content": "<p>Thanks for sharing👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1705536,
      "author_name": "highoncoffee",
      "author_url": "",
      "post_date": "2022-02-26T15:20:26.140000",
      "content": "<p>Thanks alot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1701803,
      "author_name": "Leonardo Berlatto",
      "author_url": "",
      "post_date": "2022-02-23T04:40:59.067000",
      "content": "<p>Awesome! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1700832,
      "author_name": "Patricio Javier Copado Pesce",
      "author_url": "",
      "post_date": "2022-02-22T09:35:22.373000",
      "content": "<p>Thank you!! 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1700467,
      "author_name": "Tayes Baldwin",
      "author_url": "",
      "post_date": "2022-02-22T02:51:14.717000",
      "content": "<p>Thank you for sharing this tip!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1699324,
      "author_name": "ABHIRUPDAS2",
      "author_url": "",
      "post_date": "2022-02-21T05:27:28.810000",
      "content": "<p>Thank you,really helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1699135,
      "author_name": "yuihmoo",
      "author_url": "",
      "post_date": "2022-02-21T00:16:52.457000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1697407,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-19T15:57:02.137000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1780918,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-08T03:32:04.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700687,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-22T06:52:36.227000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1697346": "I would like to share a memory trick that @titericz taught me when we competed together in the RecSys-Twitter competition in 2020 and 2021. This memory trick will reduce the training data memory by 5x! (from 96 bytes per row to 20 bytes per row)\n\n# Reduce Customer_Id by 8x (or 16x)!\n\n    import pandas as pd\n    train = pd.read_csv('transactions_train.csv')\n    train['customer_id'] =\\\n        train['customer_id'].apply(lambda x: int(x[-16:],16) ).astype('int64')\n\nThe `customer_id` is a length 64 string which uses 64 bytes. The code above coverts the column to int64 which only takes 8 bytes! Also the mapping is 1:1 which means each customer gets a unique `int64`. After you are done with your processing, i.e. `train.groupby('customer_id')` etc etc, just merge this onto the `sample_submission.csv` dataframe.\n\nWhen using **RAPIDS cuDF**, do this\n\n    import cudf\n    train = cudf.read_csv('transactions_train.csv')\n    train['customer_id'] =\\\n        train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n\n# UPDATE: Achieve 16x Reduction\nSee comments below from Clear n' Simple @jacob34 on how to achieve **16x** with a mapping.\nSee comments below from Giba, to learn how to achieve **16x** memory reduction with `int32`!\n\n# Reduce Article_Id by 2.5x!\n\n    train['article_id'] = train['article_id'].astype('int32')\n\nMany people store the `article_id` as a string to keep the leading zero. However a string of length 10 uses 10 bytes. It is better to convert this column to `int32` which only uses 4 bytes. Then do all your processing, i.e. `train.groupby('article_id')` etc etc. Then before writing submission, just convert it back to string with zero\n\n    train['article_id'] = '0' + train.article_id.astype('str')\n\n# Reduce Other Columns 3x!\nLastly you can reduce or remove the other columns\n\n    train.t_dat = cudf.to_datetime( train.t_dat )\n    train['year'] = (train.t_dat.dt.year-2000).astype('int8')\n    train['month'] = (train.t_dat.dt.month).astype('int8')\n    train['day'] = (train.t_dat.dt.day).astype('int8')\n    del train['t_dat']\n\nAnd\n\n    train['price'] = train['price'].astype('float32')\n    train['sales_channel_id'] = train['sales_channel_id'].astype('int8')\n\n# Example\nI post an example of using `customer_id` and `article_id` memory reduction [here][2] and [here][1]. The basic idea is this. After you make predictions for customers using the new int64 id, we merge it back with:\n\n    sub = cudf.read_csv('sample_submission.csv')[['customer_id']]\n    sub['customer_id_2'] =\\\n        sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n    sub = sub.merge(PREDS_DF.rename({'customer_id':'customer_id_2'},axis=1),\\\n        on='customer_id_2', how='left').fillna('')\n    del sub['customer_id_2']\n    sub.to_csv('submission.csv',index=False)\n\n# UPDATE -  Starter Notebook\nI published a starter notebook [here][2] which uses the above memory reductions. Then the entire solution uses dataframe manipulations on the reduced dataframe. Finally the predictions are merged onto `sample_submission.csv` and submitted to Kaggle to achieve LB 0.021\n\n[1]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[2]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021",
    "1698725": "Thank you!\nThe customer_id code is very cool.\n\nIs there a reason not to use the index in the customers csv file?\n```\nid_to_index_dict = dict(zip(customers[\"customer_id\"], customers.index))\nindex_to_id_dict = dict(zip(customers.index, customers[\"customer_id\"]))\n\n# for memory efficiency\ntransactions[\"customer_id\"] = transactions[\"customer_id\"].map(id_to_index_dict)\n\n# for switching back for submission\nsub[\"customer_id\"] = sub[\"customer_id\"].map(index_to_id_dict)\n```",
    "1697752": "@cdeotte Great post and I appreciate the mention. \nNote that sometimes is possible to compress even more 'customer_id' by using int32 when it don't affect much the cardinality.",
    "1699238": "UPDATE: I posted a starter notebook [here][1] which demonstrates the use of the above dataframe memory reduction. After memory reduction, the complete solution uses dataframe operations and then merges the predictions onto sample_submission.csv before submitting to Kaggle. The notebook achieves LB 0.021\n\n[1]: https://www.kaggle.com/cdeotte/recommend-items-purchased-together-0-021",
    "1697378": "Thanks so much for sharing this!\n\nI believe you had named the mapping trick as \"reverse target encoding\"? Sorry If I'm mixing the two concepts up.\n\nHere's a timestamped [link](https://youtu.be/W3aWEXqIkWk?t=3797) where Chris spoke about this.",
    "1706719": "Thanks for sharing. Really useful tricks !!",
    "1704284": "Thank you very much!, that has been helpful for me . ",
    "1702252": "I didn't have this idea, I will use it for analysis!",
    "1701288": "wow i never thought about how you can use such tricks in order to increase the speed. Thanks for sharing !!\nall of this falls back to old school way of handling data, when I was working for an application for TV set top box, I had to think so much for memory optimization, and the things never changes.. Memory optimization is always important ",
    "1701135": "Very tricky, thanks for sharing!",
    "1700678": "Excellent work, as always 👏🏻 Thank you very much!",
    "1699426": "Thanks for sharing, it is really useful👍",
    "1698516": "Nice one! For further memory usage, I would also suggest using this trick:\n\n```python\npd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\",\n            dtype={\"t_dat\": \"object\", \"customer_id\": \"object\", \"article_id\": \"object\", \"price\": float, \"sales_channel_id\": int}\n           )\n```",
    "1806684": "Cool! Helps much!",
    "1773449": "Great job.  It really helped me get past a lot of them hurdles.  Went from using 16g of ram to only 5.  Thanks.  \n\nOf course I did other things as well, but this was extra helpful.💪👍",
    "1702073": "",
    "1754767": "a nice share. thanks\n",
    "1728523": "Thanks for sharing !",
    "1720883": "Thanks for sharing. Nice tip! :)",
    "1715414": "Thanks for sharing !",
    "1714838": "Thanks for sharing！",
    "1711840": "Thanks for sharing it really useful!",
    "1710583": "Great Thank you for sharing..",
    "1708838": "Thanks for sharing👍",
    "1705536": "Thanks alot!",
    "1701803": "Awesome! Thanks for sharing",
    "1700832": "Thank you!! 👍",
    "1700467": "Thank you for sharing this tip!",
    "1699324": "Thank you,really helpful.",
    "1699135": "Thank you for sharing",
    "1697407": "Nice Tricks ! 👋\nThanks for sharing.",
    "1780918": "Really nice. Thanks alot.",
    "1700687": "Really helpful. Thanks for sharing"
  }
}