{
  "id": 311747,
  "title": "Feature Engineering Tricks | Thanks GM Chris and Giba",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/311747",
  "author_name": "",
  "post_date": "2022-03-08T15:45:32.161157700Z",
  "votes": 73,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi Everyone, </p>\n<p>GrandMaster <a href=\"http://kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> had shared this trick during an interview about their team's RecSys Winning Solution, I believe one of these was taught to him by GrandMaster <a href=\"http://kaggle.com/titericz\" target=\"_blank\">Giba</a> so thanks to both of them for sharing this. </p>\n<p>The interview is over an hour so I have made two 3 minute clips where Chris spoke about these:</p>\n<ul>\n<li>Target Encoding Trick <a href=\"https://youtu.be/OjxsAL45Lgg\" target=\"_blank\">Video</a></li>\n<li>\"Reverse Target &amp; Count Encoding\" Trick <a href=\"https://youtu.be/rX5Kj7iW9Ew\" target=\"_blank\">Video</a></li>\n</ul>\n<h3>Target Encoding Trick:</h3>\n<p><a href=\"https://www.kaggle.com/ryanholbrook/target-encoding\" target=\"_blank\">This tutorial</a> by Ryan explains really well why we apply encoding and applying smoothing to Target Encoding:</p>\n<p>Whenever working with categorical variables, we need to process numbers inside or models so we figure out a way to \"encode\" this information. There are various ways of doing this. The \"trick\" here is applying to smooth and creating folds to avoid overfitting. </p>\n<p>The NVTabular Package provides a nice function that allows this. You can find it <a href=\"https://nvidia-merlin.github.io/NVTabular/v0.7.1/api/ops/targetencoding.html\" target=\"_blank\">here</a> and leverage the same on GPUs without having to implement it :)</p>\n<p>The trick is to create k-random folds and apply smoothing using mean of the values or another method.</p>\n<h3>\"Reverse Target and Count Encoding\":</h3>\n<p>In the video above, we discussed the RecSys competition which had a somewhat similar split in data containing authors and readers in different columns. Visualised below:</p>\n<p>Ex: Table A has </p>\n<p><code>| UserID | ColA |..... | Col Z|</code></p>\n<p>Table B has: </p>\n<p><code>| UserID | Col A1 | ... | Col Z1|</code></p>\n<p>So here we can apply a Feature Engineering of Count Encoding and merge based on the same userid onto TableB. </p>\n<h3>Why is this powerful?</h3>\n<p>Chris has already shared a trick to reduce memory usage <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a></p>\n<p>On top of the same, in this competition case, for example-we can build features in one table of <code>articles</code> and <code>customers</code> and merge them together or onto <code>transactions.</code></p>\n<p>Once again, Thanks again Grandmaster Chris and Giba for sharing this with us, I have merely spent time understanding and then re-sharing it here. </p>\n<p>I was quite new to these tricks and I didn't see them mentioned in this competition so I decided to post these.</p>\n<p>Hope this helps! :)</p>",
  "messages": [
    {
      "id": "1716040",
      "postDate": "03/08/2022 15:45:32",
      "content": "<p>Hi Everyone, </p>\n<p>GrandMaster <a href=\"http://kaggle.com/cdeotte\" target=\"_blank\">Chris Deotte</a> had shared this trick during an interview about their team's RecSys Winning Solution, I believe one of these was taught to him by GrandMaster <a href=\"http://kaggle.com/titericz\" target=\"_blank\">Giba</a> so thanks to both of them for sharing this. </p>\n<p>The interview is over an hour so I have made two 3 minute clips where Chris spoke about these:</p>\n<ul>\n<li>Target Encoding Trick <a href=\"https://youtu.be/OjxsAL45Lgg\" target=\"_blank\">Video</a></li>\n<li>\"Reverse Target &amp; Count Encoding\" Trick <a href=\"https://youtu.be/rX5Kj7iW9Ew\" target=\"_blank\">Video</a></li>\n</ul>\n<h3>Target Encoding Trick:</h3>\n<p><a href=\"https://www.kaggle.com/ryanholbrook/target-encoding\" target=\"_blank\">This tutorial</a> by Ryan explains really well why we apply encoding and applying smoothing to Target Encoding:</p>\n<p>Whenever working with categorical variables, we need to process numbers inside or models so we figure out a way to \"encode\" this information. There are various ways of doing this. The \"trick\" here is applying to smooth and creating folds to avoid overfitting. </p>\n<p>The NVTabular Package provides a nice function that allows this. You can find it <a href=\"https://nvidia-merlin.github.io/NVTabular/v0.7.1/api/ops/targetencoding.html\" target=\"_blank\">here</a> and leverage the same on GPUs without having to implement it :)</p>\n<p>The trick is to create k-random folds and apply smoothing using mean of the values or another method.</p>\n<h3>\"Reverse Target and Count Encoding\":</h3>\n<p>In the video above, we discussed the RecSys competition which had a somewhat similar split in data containing authors and readers in different columns. Visualised below:</p>\n<p>Ex: Table A has </p>\n<p><code>| UserID | ColA |..... | Col Z|</code></p>\n<p>Table B has: </p>\n<p><code>| UserID | Col A1 | ... | Col Z1|</code></p>\n<p>So here we can apply a Feature Engineering of Count Encoding and merge based on the same userid onto TableB. </p>\n<h3>Why is this powerful?</h3>\n<p>Chris has already shared a trick to reduce memory usage <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a></p>\n<p>On top of the same, in this competition case, for example-we can build features in one table of <code>articles</code> and <code>customers</code> and merge them together or onto <code>transactions.</code></p>\n<p>Once again, Thanks again Grandmaster Chris and Giba for sharing this with us, I have merely spent time understanding and then re-sharing it here. </p>\n<p>I was quite new to these tricks and I didn't see them mentioned in this competition so I decided to post these.</p>\n<p>Hope this helps! :)</p>",
      "rawMarkdown": "Hi Everyone, \n\nGrandMaster [Chris Deotte](http://kaggle.com/cdeotte) had shared this trick during an interview about their team's RecSys Winning Solution, I believe one of these was taught to him by GrandMaster [Giba](http://kaggle.com/titericz) so thanks to both of them for sharing this. \n\nThe interview is over an hour so I have made two 3 minute clips where Chris spoke about these:\n\n- Target Encoding Trick [Video](https://youtu.be/OjxsAL45Lgg)\n- \"Reverse Target & Count Encoding\" Trick [Video](https://youtu.be/rX5Kj7iW9Ew)\n\n### Target Encoding Trick:\n\n[This tutorial](https://www.kaggle.com/ryanholbrook/target-encoding) by Ryan explains really well why we apply encoding and applying smoothing to Target Encoding:\n\nWhenever working with categorical variables, we need to process numbers inside or models so we figure out a way to \"encode\" this information. There are various ways of doing this. The \"trick\" here is applying to smooth and creating folds to avoid overfitting. \n\nThe NVTabular Package provides a nice function that allows this. You can find it [here](https://nvidia-merlin.github.io/NVTabular/v0.7.1/api/ops/targetencoding.html) and leverage the same on GPUs without having to implement it :)\n\nThe trick is to create k-random folds and apply smoothing using mean of the values or another method.\n\n### \"Reverse Target and Count Encoding\":\n\nIn the video above, we discussed the RecSys competition which had a somewhat similar split in data containing authors and readers in different columns. Visualised below:\n\nEx: Table A has \n\n` | UserID | ColA |..... | Col Z|`\n\nTable B has: \n\n` | UserID | Col A1 | ... | Col Z1|`\n\nSo here we can apply a Feature Engineering of Count Encoding and merge based on the same userid onto TableB. \n\n### Why is this powerful?\n\nChris has already shared a trick to reduce memory usage [here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635)\n\nOn top of the same, in this competition case, for example-we can build features in one table of `articles` and `customers` and merge them together or onto `transactions.`\n\nOnce again, Thanks again Grandmaster Chris and Giba for sharing this with us, I have merely spent time understanding and then re-sharing it here. \n\nI was quite new to these tricks and I didn't see them mentioned in this competition so I decided to post these.\n\nHope this helps! :)",
      "votes": null
    },
    {
      "id": "1716141",
      "postDate": "03/08/2022 17:21:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a>, thank you for sharing.</p>",
      "rawMarkdown": "Hi @init27, thank you for sharing.",
      "votes": null
    },
    {
      "id": "1716379",
      "postDate": "03/09/2022 00:15:32",
      "content": "<p>Thanks for sharing Sanyam. I will say that my current solution uses \"TE\" and/or \"reverse TE\" in the following sense. For each customer and/or item, we can create new features by using <code>groupby('customer_id')</code> or <code>groupby('article_id')</code> and then aggregating some statistic. Using this information we can either add new candidates to recommend or rerank recommendations that were found using a previous algorithm.</p>\n<p>Remember the general approach to this competition is (1) find 12 or 24 or 36 candidates to recommend each customer (2) rerank these recommendations arranging the most likely (to be purchased) first and least likely last. So there are two steps (1) find candidates (2) rearrange those candidates</p>",
      "rawMarkdown": "Thanks for sharing Sanyam. I will say that my current solution uses \"TE\" and/or \"reverse TE\" in the following sense. For each customer and/or item, we can create new features by using `groupby('customer_id')` or `groupby('article_id')` and then aggregating some statistic. Using this information we can either add new candidates to recommend or rerank recommendations that were found using a previous algorithm.\n\nRemember the general approach to this competition is (1) find 12 or 24 or 36 candidates to recommend each customer (2) rerank these recommendations arranging the most likely (to be purchased) first and least likely last. So there are two steps (1) find candidates (2) rearrange those candidates",
      "votes": null
    },
    {
      "id": "1716699",
      "postDate": "03/09/2022 09:58:47",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi Chris, how do we ensure the ranking as there are multiple entries of same article which also indicate a return of these articles.<br>\nor simply put <br>\nhow do we apply LTR(learning to rank) technique here </p>",
      "rawMarkdown": "cdeotte Hi Chris, how do we ensure the ranking as there are multiple entries of same article which also indicate a return of these articles.\nor simply put \nhow do we apply LTR(learning to rank) technique here",
      "votes": null
    },
    {
      "id": "1716936",
      "postDate": "03/09/2022 14:20:29",
      "content": "<p>There are many ways to rerank. We can use heuristics or we can train a reranking model. The model/heuristic can be general or specific to the customer and/or item.</p>\n<p>Here is one heuristic example. Imagine that we are simply recommending every customer the top 24 most popular items. Next we can loop through every customer. If the customer gender is male, we can move the female items lower in the recommendation rank (for that specific customer). If the customer gender is female, we can move the male items lower in the recommendation rank. </p>\n<p>We begin with 24 candidates per customer. Then we re-ranked them. Finally we keep only the top 12 and make a submission to Kaggle. Using just top 12 gets one score, but using top 24 then reranking then picking top 12 afterward gets a better score.</p>",
      "rawMarkdown": "There are many ways to rerank. We can use heuristics or we can train a reranking model. The model/heuristic can be general or specific to the customer and/or item.\n\nHere is one heuristic example. Imagine that we are simply recommending every customer the top 24 most popular items. Next we can loop through every customer. If the customer gender is male, we can move the female items lower in the recommendation rank (for that specific customer). If the customer gender is female, we can move the male items lower in the recommendation rank. \n\nWe begin with 24 candidates per customer. Then we re-ranked them. Finally we keep only the top 12 and make a submission to Kaggle. Using just top 12 gets one score, but using top 24 then reranking then picking top 12 afterward gets a better score.",
      "votes": null
    },
    {
      "id": "1716956",
      "postDate": "03/09/2022 14:35:31",
      "content": "<p>Bala, please thank the Grandmasters that taught me these tricks :)</p>",
      "rawMarkdown": "Bala, please thank the Grandmasters that taught me these tricks :)",
      "votes": null
    },
    {
      "id": "1716981",
      "postDate": "03/09/2022 14:55:07",
      "content": "<p>Pawel describes the procedure in more detail <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Pawel describes the procedure in more detail [here][1]\n\n[1]: https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288",
      "votes": null
    },
    {
      "id": "1716999",
      "postDate": "03/09/2022 15:17:58",
      "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. will try to implement them.</p>",
      "rawMarkdown": "Thanks alot @cdeotte. will try to implement them.",
      "votes": null
    },
    {
      "id": "1717515",
      "postDate": "03/10/2022 01:54:07",
      "content": "<p>(In response to top comment) </p>\n<p>Thank you for the reply, Chris! </p>\n<p>I will share my embarassing approach to understanding, incase its helpful for anyone else-I had asked the similar question to Chris around how this is applied in the Memory trick thread but then on paper after drawing and it didn't make sense to me still. </p>\n<p>I was still thinking if we merge all tables, maybe the trick won't apply here because then we have 1 final <code>df</code>. </p>\n<p>After wasting silly hours and then finally firing up a jupyter nb, I realised if we merge rows, there always isn't a 1:1 mapping and we can use <code>groupby()</code>. </p>\n<p>Lesson learned: Code first and draw on paper when stuck not the other way around to save time </p>",
      "rawMarkdown": "(In response to top comment) \n\nThank you for the reply, Chris! \n\nI will share my embarassing approach to understanding, incase its helpful for anyone else-I had asked the similar question to Chris around how this is applied in the Memory trick thread but then on paper after drawing and it didn't make sense to me still. \n\nI was still thinking if we merge all tables, maybe the trick won't apply here because then we have 1 final `df`. \n\nAfter wasting silly hours and then finally firing up a jupyter nb, I realised if we merge rows, there always isn't a 1:1 mapping and we can use `groupby()`. \n\nLesson learned: Code first and draw on paper when stuck not the other way around to save time",
      "votes": null
    },
    {
      "id": "1721578",
      "postDate": "03/13/2022 19:57:47",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I just want to make sure I undestand it correctly. If I want to make target encoding for each of the user combined with each of the products as a dynamic feature for ranking, without smoothing the target encoding would be number of times article appeared in a customer baskets and divided by the number of other articles that ever appeared in the customer basket right (in specified time range)? thank you for sharing your knowledge</p>",
      "rawMarkdown": "cdeotte I just want to make sure I undestand it correctly. If I want to make target encoding for each of the user combined with each of the products as a dynamic feature for ranking, without smoothing the target encoding would be number of times article appeared in a customer baskets and divided by the number of other articles that ever appeared in the customer basket right (in specified time range)? thank you for sharing your knowledge",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1716141,
      "author_name": "balabaskar",
      "author_url": "",
      "post_date": "03/08/2022 17:21:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a>, thank you for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1716956,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/09/2022 14:35:31",
          "content": "<p>Bala, please thank the Grandmasters that taught me these tricks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1716379,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/09/2022 00:15:32",
      "content": "<p>Thanks for sharing Sanyam. I will say that my current solution uses \"TE\" and/or \"reverse TE\" in the following sense. For each customer and/or item, we can create new features by using <code>groupby('customer_id')</code> or <code>groupby('article_id')</code> and then aggregating some statistic. Using this information we can either add new candidates to recommend or rerank recommendations that were found using a previous algorithm.</p>\n<p>Remember the general approach to this competition is (1) find 12 or 24 or 36 candidates to recommend each customer (2) rerank these recommendations arranging the most likely (to be purchased) first and least likely last. So there are two steps (1) find candidates (2) rearrange those candidates</p>",
      "votes": null,
      "replies": [
        {
          "id": 1716699,
          "author_name": "vjagannath786",
          "author_url": "",
          "post_date": "03/09/2022 09:58:47",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi Chris, how do we ensure the ranking as there are multiple entries of same article which also indicate a return of these articles.<br>\nor simply put <br>\nhow do we apply LTR(learning to rank) technique here </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1716936,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/09/2022 14:20:29",
          "content": "<p>There are many ways to rerank. We can use heuristics or we can train a reranking model. The model/heuristic can be general or specific to the customer and/or item.</p>\n<p>Here is one heuristic example. Imagine that we are simply recommending every customer the top 24 most popular items. Next we can loop through every customer. If the customer gender is male, we can move the female items lower in the recommendation rank (for that specific customer). If the customer gender is female, we can move the male items lower in the recommendation rank. </p>\n<p>We begin with 24 candidates per customer. Then we re-ranked them. Finally we keep only the top 12 and make a submission to Kaggle. Using just top 12 gets one score, but using top 24 then reranking then picking top 12 afterward gets a better score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1716981,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/09/2022 14:55:07",
          "content": "<p>Pawel describes the procedure in more detail <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1716999,
          "author_name": "vjagannath786",
          "author_url": "",
          "post_date": "03/09/2022 15:17:58",
          "content": "<p>Thanks alot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. will try to implement them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1717515,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/10/2022 01:54:07",
          "content": "<p>(In response to top comment) </p>\n<p>Thank you for the reply, Chris! </p>\n<p>I will share my embarassing approach to understanding, incase its helpful for anyone else-I had asked the similar question to Chris around how this is applied in the Memory trick thread but then on paper after drawing and it didn't make sense to me still. </p>\n<p>I was still thinking if we merge all tables, maybe the trick won't apply here because then we have 1 final <code>df</code>. </p>\n<p>After wasting silly hours and then finally firing up a jupyter nb, I realised if we merge rows, there always isn't a 1:1 mapping and we can use <code>groupby()</code>. </p>\n<p>Lesson learned: Code first and draw on paper when stuck not the other way around to save time </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721578,
          "author_name": "duuuscha",
          "author_url": "",
          "post_date": "03/13/2022 19:57:47",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I just want to make sure I undestand it correctly. If I want to make target encoding for each of the user combined with each of the products as a dynamic feature for ranking, without smoothing the target encoding would be number of times article appeared in a customer baskets and divided by the number of other articles that ever appeared in the customer basket right (in specified time range)? thank you for sharing your knowledge</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1716040": "Hi Everyone, \n\nGrandMaster [Chris Deotte](http://kaggle.com/cdeotte) had shared this trick during an interview about their team's RecSys Winning Solution, I believe one of these was taught to him by GrandMaster [Giba](http://kaggle.com/titericz) so thanks to both of them for sharing this. \n\nThe interview is over an hour so I have made two 3 minute clips where Chris spoke about these:\n\n- Target Encoding Trick [Video](https://youtu.be/OjxsAL45Lgg)\n- \"Reverse Target & Count Encoding\" Trick [Video](https://youtu.be/rX5Kj7iW9Ew)\n\n### Target Encoding Trick:\n\n[This tutorial](https://www.kaggle.com/ryanholbrook/target-encoding) by Ryan explains really well why we apply encoding and applying smoothing to Target Encoding:\n\nWhenever working with categorical variables, we need to process numbers inside or models so we figure out a way to \"encode\" this information. There are various ways of doing this. The \"trick\" here is applying to smooth and creating folds to avoid overfitting. \n\nThe NVTabular Package provides a nice function that allows this. You can find it [here](https://nvidia-merlin.github.io/NVTabular/v0.7.1/api/ops/targetencoding.html) and leverage the same on GPUs without having to implement it :)\n\nThe trick is to create k-random folds and apply smoothing using mean of the values or another method.\n\n### \"Reverse Target and Count Encoding\":\n\nIn the video above, we discussed the RecSys competition which had a somewhat similar split in data containing authors and readers in different columns. Visualised below:\n\nEx: Table A has \n\n` | UserID | ColA |..... | Col Z|`\n\nTable B has: \n\n` | UserID | Col A1 | ... | Col Z1|`\n\nSo here we can apply a Feature Engineering of Count Encoding and merge based on the same userid onto TableB. \n\n### Why is this powerful?\n\nChris has already shared a trick to reduce memory usage [here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635)\n\nOn top of the same, in this competition case, for example-we can build features in one table of `articles` and `customers` and merge them together or onto `transactions.`\n\nOnce again, Thanks again Grandmaster Chris and Giba for sharing this with us, I have merely spent time understanding and then re-sharing it here. \n\nI was quite new to these tricks and I didn't see them mentioned in this competition so I decided to post these.\n\nHope this helps! :)",
    "1716141": "Hi @init27, thank you for sharing.",
    "1716379": "Thanks for sharing Sanyam. I will say that my current solution uses \"TE\" and/or \"reverse TE\" in the following sense. For each customer and/or item, we can create new features by using `groupby('customer_id')` or `groupby('article_id')` and then aggregating some statistic. Using this information we can either add new candidates to recommend or rerank recommendations that were found using a previous algorithm.\n\nRemember the general approach to this competition is (1) find 12 or 24 or 36 candidates to recommend each customer (2) rerank these recommendations arranging the most likely (to be purchased) first and least likely last. So there are two steps (1) find candidates (2) rearrange those candidates",
    "1716699": "cdeotte Hi Chris, how do we ensure the ranking as there are multiple entries of same article which also indicate a return of these articles.\nor simply put \nhow do we apply LTR(learning to rank) technique here",
    "1716936": "There are many ways to rerank. We can use heuristics or we can train a reranking model. The model/heuristic can be general or specific to the customer and/or item.\n\nHere is one heuristic example. Imagine that we are simply recommending every customer the top 24 most popular items. Next we can loop through every customer. If the customer gender is male, we can move the female items lower in the recommendation rank (for that specific customer). If the customer gender is female, we can move the male items lower in the recommendation rank. \n\nWe begin with 24 candidates per customer. Then we re-ranked them. Finally we keep only the top 12 and make a submission to Kaggle. Using just top 12 gets one score, but using top 24 then reranking then picking top 12 afterward gets a better score.",
    "1716956": "Bala, please thank the Grandmasters that taught me these tricks :)",
    "1716981": "Pawel describes the procedure in more detail [here][1]\n\n[1]: https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288",
    "1716999": "Thanks alot @cdeotte. will try to implement them.",
    "1717515": "(In response to top comment) \n\nThank you for the reply, Chris! \n\nI will share my embarassing approach to understanding, incase its helpful for anyone else-I had asked the similar question to Chris around how this is applied in the Memory trick thread but then on paper after drawing and it didn't make sense to me still. \n\nI was still thinking if we merge all tables, maybe the trick won't apply here because then we have 1 final `df`. \n\nAfter wasting silly hours and then finally firing up a jupyter nb, I realised if we merge rows, there always isn't a 1:1 mapping and we can use `groupby()`. \n\nLesson learned: Code first and draw on paper when stuck not the other way around to save time",
    "1721578": "cdeotte I just want to make sure I undestand it correctly. If I want to make target encoding for each of the user combined with each of the products as a dynamic feature for ranking, without smoothing the target encoding would be number of times article appeared in a customer baskets and divided by the number of other articles that ever appeared in the customer basket right (in specified time range)? thank you for sharing your knowledge"
  },
  "source": "meta"
}