{
  "id": 314458,
  "title": "Care to share?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458",
  "author_name": "Clear n' Simple",
  "post_date": "2022-03-22T17:51:02.457000",
  "votes": 60,
  "comment_count": 28,
  "views": 0,
  "content": "<p>I'm sharing some details of my setup and what I've found helpful so far (<code>2022-03-22</code>)<br>\nLikely not very interesting for top teams, but hopefully helpful for other people like me.</p>\n<p>For those of you who've gotten decent results and would be willing to share what you've found helpful, I would very much appreciate it!</p>\n<ul>\n<li>Am working entirely with single Kaggle notebook.</li>\n<li>Abstracting as much as possible to helper functions, and storing in modules as kaggle dataset, to keep main notebook uncluttered and readable.</li>\n<li>Am using google cloud for storing outputs/inputs of different versions of notebooks.</li>\n<li>Using gpu with cudf allows quick iteration and testing - once data is loaded and prepared, experiments on full dataset with 3 cvs taking 1/2 a minute each (haven't started modeling yet)</li>\n<li>CV set up similar to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">here</a></li>\n<li>I've created a <code>week_number</code> column, with week numbers 0 to 104, so I can deal with weeks instead of days for cv (and other purposes)</li>\n<li>Like many people, I've been using <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>'s  <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">thread</a> for guidance.</li>\n<li>So far, have been focusing on step #1 - generating candidates. </li>\n<li>To evaluate how \"good\" my generated candidates are, in cv, I look at:    <ol>\n<li><code>recall</code> - what percentage of unique true customer_id/article_id purchases I have in the candidates.</li>\n<li><code>multiple factor</code> - how many times more total candidates I have than true candidates (like the inverse of precision)</li></ol></li>\n<li>For example (real example), if evaluation week of last cv has ~214k unique customer_id/article_id purchases, and I generated ~29,970k candidates, ~16k of which were actual purchases:<ol>\n<li><code>recall</code> = 16k / 214k = ~7.5%</li>\n<li><code>multiple factor</code> = 29,970k / 16k = 1,878</li></ol></li>\n<li>As I create candidates, I also create/store the relevant features on the basis of which candidates were chosen. Currently, I'm only using those features to manually sort/choose from among the candidates, but in next step, they can be features for a ranking model</li>\n<li>With gpu/cudf, only part that's taking longer than a minute is a version of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pair generating code, which is taking about 30-45 minutes for the entire dataset.</li>\n<li>Important to take into account article availability - if an item isn't available, doesn't really matter if it's likely that a customer would buy it.</li>\n</ul>\n<p>Update <code>2022-03-24</code>:<br>\nImplemented draft version of using LGBMRanker</p>\n<ul>\n<li>Memory starting to become an issue</li>\n<li>For training/prediction, converted from cudf to pandas (<code>to_pandas()</code>)</li>\n<li>Without any parameter tuning (except for n_estimators), got a +.0005 boost</li>\n</ul>\n<p>Update <code>2022-04-06</code>:</p>\n<ul>\n<li>wrote/setup modeling code properly</li>\n<li>Implemented <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">proper cv pipeline for modeling</a></li>\n<li>Got memory issues under control by:<ul>\n<li><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316193\" target=\"_blank\">using <code>gc.collect()</code> more often</a></li>\n<li>changing int columns with nulls to float32 before calling to <code>to_pandas()</code></li>\n<li><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314829\" target=\"_blank\">pickling items while they're not needed</a></li>\n<li>changing dfs in-place, instead of returning new df each time</li></ul></li>\n<li>able to improve LB, but can't get it to correlate to cv. Working on that now.</li>\n</ul>\n<p>Update <code>2022-04-27</code>:</p>\n<ul>\n<li>after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.</li>\n<li>In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.</li>\n<li>Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.</li>\n</ul>\n<p>Update <code>2022-05-05</code>:<br>\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage…<br>\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.<br>\nWe'll see what I can do in a few days…</p>",
  "messages": [
    {
      "id": 1731808,
      "postDate": "2022-03-22T17:51:02.457Z",
      "content": "<p>I'm sharing some details of my setup and what I've found helpful so far (<code>2022-03-22</code>)<br>\nLikely not very interesting for top teams, but hopefully helpful for other people like me.</p>\n<p>For those of you who've gotten decent results and would be willing to share what you've found helpful, I would very much appreciate it!</p>\n<ul>\n<li>Am working entirely with single Kaggle notebook.</li>\n<li>Abstracting as much as possible to helper functions, and storing in modules as kaggle dataset, to keep main notebook uncluttered and readable.</li>\n<li>Am using google cloud for storing outputs/inputs of different versions of notebooks.</li>\n<li>Using gpu with cudf allows quick iteration and testing - once data is loaded and prepared, experiments on full dataset with 3 cvs taking 1/2 a minute each (haven't started modeling yet)</li>\n<li>CV set up similar to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">here</a></li>\n<li>I've created a <code>week_number</code> column, with week numbers 0 to 104, so I can deal with weeks instead of days for cv (and other purposes)</li>\n<li>Like many people, I've been using <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>'s  <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">thread</a> for guidance.</li>\n<li>So far, have been focusing on step #1 - generating candidates. </li>\n<li>To evaluate how \"good\" my generated candidates are, in cv, I look at:    <ol>\n<li><code>recall</code> - what percentage of unique true customer_id/article_id purchases I have in the candidates.</li>\n<li><code>multiple factor</code> - how many times more total candidates I have than true candidates (like the inverse of precision)</li></ol></li>\n<li>For example (real example), if evaluation week of last cv has ~214k unique customer_id/article_id purchases, and I generated ~29,970k candidates, ~16k of which were actual purchases:<ol>\n<li><code>recall</code> = 16k / 214k = ~7.5%</li>\n<li><code>multiple factor</code> = 29,970k / 16k = 1,878</li></ol></li>\n<li>As I create candidates, I also create/store the relevant features on the basis of which candidates were chosen. Currently, I'm only using those features to manually sort/choose from among the candidates, but in next step, they can be features for a ranking model</li>\n<li>With gpu/cudf, only part that's taking longer than a minute is a version of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pair generating code, which is taking about 30-45 minutes for the entire dataset.</li>\n<li>Important to take into account article availability - if an item isn't available, doesn't really matter if it's likely that a customer would buy it.</li>\n</ul>\n<p>Update <code>2022-03-24</code>:<br>\nImplemented draft version of using LGBMRanker</p>\n<ul>\n<li>Memory starting to become an issue</li>\n<li>For training/prediction, converted from cudf to pandas (<code>to_pandas()</code>)</li>\n<li>Without any parameter tuning (except for n_estimators), got a +.0005 boost</li>\n</ul>\n<p>Update <code>2022-04-06</code>:</p>\n<ul>\n<li>wrote/setup modeling code properly</li>\n<li>Implemented <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">proper cv pipeline for modeling</a></li>\n<li>Got memory issues under control by:<ul>\n<li><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316193\" target=\"_blank\">using <code>gc.collect()</code> more often</a></li>\n<li>changing int columns with nulls to float32 before calling to <code>to_pandas()</code></li>\n<li><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314829\" target=\"_blank\">pickling items while they're not needed</a></li>\n<li>changing dfs in-place, instead of returning new df each time</li></ul></li>\n<li>able to improve LB, but can't get it to correlate to cv. Working on that now.</li>\n</ul>\n<p>Update <code>2022-04-27</code>:</p>\n<ul>\n<li>after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.</li>\n<li>In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.</li>\n<li>Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.</li>\n</ul>\n<p>Update <code>2022-05-05</code>:<br>\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage…<br>\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.<br>\nWe'll see what I can do in a few days…</p>",
      "rawMarkdown": "I'm sharing some details of my setup and what I've found helpful so far (`2022-03-22`)\nLikely not very interesting for top teams, but hopefully helpful for other people like me.\n\nFor those of you who've gotten decent results and would be willing to share what you've found helpful, I would very much appreciate it!\n\n- Am working entirely with single Kaggle notebook.\n- Abstracting as much as possible to helper functions, and storing in modules as kaggle dataset, to keep main notebook uncluttered and readable.\n- Am using google cloud for storing outputs/inputs of different versions of notebooks.\n- Using gpu with cudf allows quick iteration and testing - once data is loaded and prepared, experiments on full dataset with 3 cvs taking 1/2 a minute each (haven't started modeling yet)\n- CV set up similar to @cdeotte's [here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308919)\n- I've created a `week_number` column, with week numbers 0 to 104, so I can deal with weeks instead of days for cv (and other purposes)\n- Like many people, I've been using @paweljankiewicz's  [thread](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288) for guidance.\n- So far, have been focusing on step #1 - generating candidates. \n- To evaluate how \"good\" my generated candidates are, in cv, I look at:    \n    1. `recall` - what percentage of unique true customer_id/article_id purchases I have in the candidates.\n    2. `multiple factor` - how many times more total candidates I have than true candidates (like the inverse of precision)\n- For example (real example), if evaluation week of last cv has ~214k unique customer_id/article_id purchases, and I generated ~29,970k candidates, ~16k of which were actual purchases:\n    1. `recall` = 16k / 214k = ~7.5%\n    2. `multiple factor` = 29,970k / 16k = 1,878\n- As I create candidates, I also create/store the relevant features on the basis of which candidates were chosen. Currently, I'm only using those features to manually sort/choose from among the candidates, but in next step, they can be features for a ranking model\n- With gpu/cudf, only part that's taking longer than a minute is a version of @cdeotte pair generating code, which is taking about 30-45 minutes for the entire dataset.\n- Important to take into account article availability - if an item isn't available, doesn't really matter if it's likely that a customer would buy it.\n\nUpdate `2022-03-24`:\nImplemented draft version of using LGBMRanker\n- Memory starting to become an issue\n- For training/prediction, converted from cudf to pandas (`to_pandas()`)\n- Without any parameter tuning (except for n_estimators), got a +.0005 boost\n\nUpdate `2022-04-06`:\n- wrote/setup modeling code properly\n- Implemented [proper cv pipeline for modeling](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386)\n- Got memory issues under control by:\n    - [using `gc.collect()` more often](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316193)\n    - changing int columns with nulls to float32 before calling to `to_pandas()`\n    - [pickling items while they're not needed](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314829)\n    - changing dfs in-place, instead of returning new df each time\n- able to improve LB, but can't get it to correlate to cv. Working on that now.\n\nUpdate `2022-04-27`:\n- after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.\n- In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.\n- Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.\n\nUpdate `2022-05-05`:\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage...\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.\nWe'll see what I can do in a few days...",
      "votes": 58
    },
    {
      "id": 1754550,
      "postDate": "2022-04-13T18:54:06.817Z",
      "content": "<p>I've just noticed this thread. I admire you that you want to control the memory and use Kaggle notebook. I did almost the same till around half of the competition but used 128 GB RAM which I considered a concession when working with this problem - I'm blessed with a home PC with this amount of memory and 32 threads. Only since about 2 weeks I switched to the cloud and 256 GB RAM because I was having too many issues and I suspect by the time it is over I will probably train the last models on 512 GB RAM. I'm writing this just to give you a heads up what would you need to do if you wanted to \"quickly\" improve on the LB. Just generate more candidates/train bigger models.</p>",
      "rawMarkdown": "I've just noticed this thread. I admire you that you want to control the memory and use Kaggle notebook. I did almost the same till around half of the competition but used 128 GB RAM which I considered a concession when working with this problem - I'm blessed with a home PC with this amount of memory and 32 threads. Only since about 2 weeks I switched to the cloud and 256 GB RAM because I was having too many issues and I suspect by the time it is over I will probably train the last models on 512 GB RAM. I'm writing this just to give you a heads up what would you need to do if you wanted to \"quickly\" improve on the LB. Just generate more candidates/train bigger models.",
      "votes": 11
    },
    {
      "id": 1733846,
      "postDate": "2022-03-24T16:47:47.057Z",
      "content": "<p>I have been also having trouble with creating good recommendations.</p>\n<p>recall of about ~7.5% seems small. But I haven't got much better results than you. I am now trying to cluster customers and articles. The goal is to find a reasonable number of articles that are mostly bought by fewer customers, so it is easy to build recommendations for them.</p>\n<p>I've been focusing with Males and distinguishing sales_channel_id. 80% of articles bought in a week come from previously items bought in the past. So I am working with about 20.000 customers and 10.000 articles.</p>\n<p>And then I can build proper recommendations to rank them with LGBMRanker.</p>\n<p>I don't know if this approach will work. And if it would be feasible for other \"segments\" like females (most customers buy female clothes). </p>\n<p>The main goal to segment is simply to reduce the amount of articles each set of customers buy, so it is easy to build recommendations (I think this is the most difficult problem of the competition. As ranking with good recommendations is easy). </p>",
      "rawMarkdown": "I have been also having trouble with creating good recommendations.\n\nrecall of about ~7.5% seems small. But I haven't got much better results than you. I am now trying to cluster customers and articles. The goal is to find a reasonable number of articles that are mostly bought by fewer customers, so it is easy to build recommendations for them.\n\nI've been focusing with Males and distinguishing sales_channel_id. 80% of articles bought in a week come from previously items bought in the past. So I am working with about 20.000 customers and 10.000 articles.\n\nAnd then I can build proper recommendations to rank them with LGBMRanker.\n\nI don't know if this approach will work. And if it would be feasible for other \"segments\" like females (most customers buy female clothes). \n\nThe main goal to segment is simply to reduce the amount of articles each set of customers buy, so it is easy to build recommendations (I think this is the most difficult problem of the competition. As ranking with good recommendations is easy). ",
      "votes": 4,
      "replies": [
        {
          "id": 1754844,
          "postDate": "2022-04-14T03:49:19.273Z",
          "content": "<p>I would agree with this approach, though not sure what value that may add. The challenge would be to carry it to completion so that you may be able to get a reasonable public leaderboard score for validation. Basically, there is a lot of work before validation is possible. I am in the same boat.</p>",
          "rawMarkdown": "I would agree with this approach, though not sure what value that may add. The challenge would be to carry it to completion so that you may be able to get a reasonable public leaderboard score for validation. Basically, there is a lot of work before validation is possible. I am in the same boat."
        }
      ]
    },
    {
      "id": 1769938,
      "postDate": "2022-04-27T19:02:21.080Z",
      "content": "<p>Update <code>2022-04-27</code>:</p>\n<ul>\n<li>after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.</li>\n<li>In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.</li>\n<li>Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.</li>\n</ul>",
      "rawMarkdown": "Update `2022-04-27`:\n- after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.\n- In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.\n- Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.\n",
      "votes": 1
    },
    {
      "id": 1760747,
      "postDate": "2022-04-19T13:52:25.763Z",
      "content": "<p>Thanks for sharing, that's very useful</p>",
      "rawMarkdown": "Thanks for sharing, that's very useful",
      "votes": 1,
      "replies": [
        {
          "id": 1760941,
          "postDate": "2022-04-19T15:47:24.590Z",
          "content": "<p>Pleasure.<br>\nAppreciate the feedback!</p>",
          "rawMarkdown": "Pleasure.\nAppreciate the feedback!"
        }
      ]
    },
    {
      "id": 1742244,
      "postDate": "2022-04-01T16:01:20.217Z",
      "content": "<p>Very useful information, thanks for sharing.</p>\n<p>I have two doubts:</p>\n<p>1) You say that you use chris's code to generate products that are frequently purchased together.<br>\nThis code uses the entire transaction dataframe regardless of time. Therefore if we include this information directly in a model I think we can overfit the model. I'm curious about how you have implemented this.</p>\n<p>2) You say that it is important to take into account the availability of the products. Do you have any more details about how you get to know this availability.</p>\n<p>Thanks in advance</p>",
      "rawMarkdown": "Very useful information, thanks for sharing.\n\nI have two doubts:\n\n1) You say that you use chris's code to generate products that are frequently purchased together.\nThis code uses the entire transaction dataframe regardless of time. Therefore if we include this information directly in a model I think we can overfit the model. I'm curious about how you have implemented this.\n\n2) You say that it is important to take into account the availability of the products. Do you have any more details about how you get to know this availability.\n\nThanks in advance",
      "votes": 1,
      "replies": [
        {
          "id": 1742273,
          "postDate": "2022-04-01T16:33:55.670Z",
          "content": "<p>You're welcome!</p>\n<ol>\n<li>You're 100% right. I actually run my code separately for each cv fold.</li>\n<li>I don't have it worked out perfectly. But if an item sold in the past, and hasn't sold recently, it's likely no longer being sold. </li>\n</ol>",
          "rawMarkdown": "You're welcome!\n\n1. You're 100% right. I actually run my code separately for each cv fold.\n2. I don't have it worked out perfectly. But if an item sold in the past, and hasn't sold recently, it's likely no longer being sold. ",
          "votes": 1
        },
        {
          "id": 1742306,
          "postDate": "2022-04-01T17:06:50.060Z",
          "content": "<p>Okay, I got it; thanks :)</p>",
          "rawMarkdown": "Okay, I got it; thanks :)"
        }
      ]
    },
    {
      "id": 1733665,
      "postDate": "2022-03-24T13:51:51.840Z",
      "content": "<p>Updated my post with initial LGBMRanker info</p>",
      "rawMarkdown": "Updated my post with initial LGBMRanker info",
      "votes": 1
    },
    {
      "id": 1732090,
      "postDate": "2022-03-23T01:54:28.980Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Thanks for sharing! </p>\n<p>I would like to ask, the ~29,970k candidates are from:</p>\n<ul>\n<li>~69k customer in last week cv only or </li>\n<li>~1,372k customer in submission file</li>\n</ul>",
      "rawMarkdown": "Hi @jacob34 Thanks for sharing! \n\nI would like to ask, the ~29,970k candidates are from:\n- ~69k customer in last week cv only or \n- ~1,372k customer in submission file\n",
      "votes": 1,
      "replies": [
        {
          "id": 1732094,
          "postDate": "2022-03-23T02:02:27.583Z",
          "content": "<p>All customers in submission file.</p>",
          "rawMarkdown": "All customers in submission file."
        },
        {
          "id": 1732098,
          "postDate": "2022-03-23T02:08:40.773Z",
          "content": "<p>awesome! <br>\nThanks Jacob </p>",
          "rawMarkdown": "awesome! \nThanks Jacob ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1777728,
      "postDate": "2022-05-04T18:33:49.023Z",
      "content": "<p>Update <code>2022-05-05</code>:</p>\n<p>Finally figured out what's going on with my lgbmranker - there was a nasty, silent leakage…<br>\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.<br>\nWe'll see what I can do in a few days…</p>",
      "rawMarkdown": "Update `2022-05-05`:\n\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage...\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.\nWe'll see what I can do in a few days..."
    },
    {
      "id": 1774862,
      "postDate": "2022-05-02T14:16:10.017Z",
      "content": "<p>How do you deal with the prediction for billion of rows ?<br>\nI cross-product the users and items. Where one user require 1 second to predict.<br>\nEven I got 100x speedup I would have to wait for many hours.</p>",
      "rawMarkdown": "How do you deal with the prediction for billion of rows ?\nI cross-product the users and items. Where one user require 1 second to predict.\nEven I got 100x speedup I would have to wait for many hours."
    },
    {
      "id": 1765329,
      "postDate": "2022-04-23T10:54:32.180Z",
      "content": "<p>How can I install cudf in kaggle kernel?<br>\npip install cudf ocurr with error<br>\nThank you</p>",
      "rawMarkdown": "How can I install cudf in kaggle kernel?\npip install cudf ocurr with error\nThank you",
      "replies": [
        {
          "id": 1765362,
          "postDate": "2022-04-23T11:40:39.707Z",
          "content": "<p>it is already installed if you choose The GPU</p>",
          "rawMarkdown": "it is already installed if you choose The GPU"
        },
        {
          "id": 1765423,
          "postDate": "2022-04-23T13:35:48.357Z",
          "content": "<p>Thanks a lot</p>",
          "rawMarkdown": "Thanks a lot"
        }
      ]
    },
    {
      "id": 1762235,
      "postDate": "2022-04-20T14:31:03Z",
      "content": "<p>What dose CV mean? Thanks </p>",
      "rawMarkdown": "What dose CV mean? Thanks ",
      "replies": [
        {
          "id": 1762351,
          "postDate": "2022-04-20T16:09:17.550Z",
          "content": "<p>\"cross validation\",  but it's used more loosely to refer to the local validation you're doing on your models, as opposed to the result you get on the leader board.</p>\n<p>Check out these discussions in regards to CV for this competition:<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">discussion #1</a><br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">discussion #2</a></p>",
          "rawMarkdown": "\"cross validation\",  but it's used more loosely to refer to the local validation you're doing on your models, as opposed to the result you get on the leader board.\n\nCheck out these discussions in regards to CV for this competition:\n[discussion #1](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919)\n[discussion #2](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386)",
          "votes": 1
        },
        {
          "id": 1762828,
          "postDate": "2022-04-21T01:43:55.347Z",
          "content": "<p>Thank you for your patience！</p>",
          "rawMarkdown": "Thank you for your patience！"
        }
      ]
    },
    {
      "id": 1744086,
      "postDate": "2022-04-03T15:30:45.683Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Thanks for sharing!<br>\nI would like to ask <br>\nfor now, how many different recall methods are used when you generating candidates?</p>",
      "rawMarkdown": "Hi @jacob34 Thanks for sharing!\nI would like to ask \nfor now, how many different recall methods are used when you generating candidates?",
      "replies": [
        {
          "id": 1744137,
          "postDate": "2022-04-03T16:22:17.167Z",
          "content": "<p>The basic recall methods so far are just from cdeotte's notebook - <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021</a></p>\n<p>But I've made several changes.</p>",
          "rawMarkdown": "The basic recall methods so far are just from cdeotte's notebook - https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\n\nBut I've made several changes.",
          "votes": 1
        },
        {
          "id": 1749009,
          "postDate": "2022-04-08T07:06:20.293Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1762839,
      "postDate": "2022-04-21T02:18:15.370Z",
      "content": "<p>Thanks for sharing! Great job!</p>",
      "rawMarkdown": "Thanks for sharing! Great job!",
      "votes": 3
    },
    {
      "id": 1753555,
      "postDate": "2022-04-13T00:22:21.410Z",
      "content": "<p>Thank you very much for sharing</p>",
      "rawMarkdown": "Thank you very much for sharing",
      "votes": 1
    },
    {
      "id": 1732018,
      "postDate": "2022-03-22T23:57:59.470Z",
      "content": "<p>Thanks for share </p>",
      "rawMarkdown": "Thanks for share ",
      "votes": 1
    },
    {
      "id": 1776443,
      "postDate": "2022-05-04T01:39:13.523Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 1754550,
      "author_name": "Paweł Jankiewicz",
      "author_url": "",
      "post_date": "2022-04-13T18:54:06.817000",
      "content": "<p>I've just noticed this thread. I admire you that you want to control the memory and use Kaggle notebook. I did almost the same till around half of the competition but used 128 GB RAM which I considered a concession when working with this problem - I'm blessed with a home PC with this amount of memory and 32 threads. Only since about 2 weeks I switched to the cloud and 256 GB RAM because I was having too many issues and I suspect by the time it is over I will probably train the last models on 512 GB RAM. I'm writing this just to give you a heads up what would you need to do if you wanted to \"quickly\" improve on the LB. Just generate more candidates/train bigger models.</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 1733846,
      "author_name": "Carlos Pérez Ricardo",
      "author_url": "",
      "post_date": "2022-03-24T16:47:47.057000",
      "content": "<p>I have been also having trouble with creating good recommendations.</p>\n<p>recall of about ~7.5% seems small. But I haven't got much better results than you. I am now trying to cluster customers and articles. The goal is to find a reasonable number of articles that are mostly bought by fewer customers, so it is easy to build recommendations for them.</p>\n<p>I've been focusing with Males and distinguishing sales_channel_id. 80% of articles bought in a week come from previously items bought in the past. So I am working with about 20.000 customers and 10.000 articles.</p>\n<p>And then I can build proper recommendations to rank them with LGBMRanker.</p>\n<p>I don't know if this approach will work. And if it would be feasible for other \"segments\" like females (most customers buy female clothes). </p>\n<p>The main goal to segment is simply to reduce the amount of articles each set of customers buy, so it is easy to build recommendations (I think this is the most difficult problem of the competition. As ranking with good recommendations is easy). </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1754844,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-04-14T03:49:19.273000",
          "content": "<p>I would agree with this approach, though not sure what value that may add. The challenge would be to carry it to completion so that you may be able to get a reasonable public leaderboard score for validation. Basically, there is a lot of work before validation is possible. I am in the same boat.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1769938,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-04-27T19:02:21.080000",
      "content": "<p>Update <code>2022-04-27</code>:</p>\n<ul>\n<li>after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.</li>\n<li>In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.</li>\n<li>Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1760747,
      "author_name": "jiuda",
      "author_url": "",
      "post_date": "2022-04-19T13:52:25.763000",
      "content": "<p>Thanks for sharing, that's very useful</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1760941,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-19T15:47:24.590000",
          "content": "<p>Pleasure.<br>\nAppreciate the feedback!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1742244,
      "author_name": "Fnoa",
      "author_url": "",
      "post_date": "2022-04-01T16:01:20.217000",
      "content": "<p>Very useful information, thanks for sharing.</p>\n<p>I have two doubts:</p>\n<p>1) You say that you use chris's code to generate products that are frequently purchased together.<br>\nThis code uses the entire transaction dataframe regardless of time. Therefore if we include this information directly in a model I think we can overfit the model. I'm curious about how you have implemented this.</p>\n<p>2) You say that it is important to take into account the availability of the products. Do you have any more details about how you get to know this availability.</p>\n<p>Thanks in advance</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1742273,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-01T16:33:55.670000",
          "content": "<p>You're welcome!</p>\n<ol>\n<li>You're 100% right. I actually run my code separately for each cv fold.</li>\n<li>I don't have it worked out perfectly. But if an item sold in the past, and hasn't sold recently, it's likely no longer being sold. </li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1742306,
          "author_name": "Fnoa",
          "author_url": "",
          "post_date": "2022-04-01T17:06:50.060000",
          "content": "<p>Okay, I got it; thanks :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1733665,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-03-24T13:51:51.840000",
      "content": "<p>Updated my post with initial LGBMRanker info</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1732090,
      "author_name": "Hervind Philipe",
      "author_url": "",
      "post_date": "2022-03-23T01:54:28.980000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Thanks for sharing! </p>\n<p>I would like to ask, the ~29,970k candidates are from:</p>\n<ul>\n<li>~69k customer in last week cv only or </li>\n<li>~1,372k customer in submission file</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1732094,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-03-23T02:02:27.583000",
          "content": "<p>All customers in submission file.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1732098,
          "author_name": "Hervind Philipe",
          "author_url": "",
          "post_date": "2022-03-23T02:08:40.773000",
          "content": "<p>awesome! <br>\nThanks Jacob </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1777728,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-04T18:33:49.023000",
      "content": "<p>Update <code>2022-05-05</code>:</p>\n<p>Finally figured out what's going on with my lgbmranker - there was a nasty, silent leakage…<br>\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.<br>\nWe'll see what I can do in a few days…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1774862,
      "author_name": "WeiChing Lin",
      "author_url": "",
      "post_date": "2022-05-02T14:16:10.017000",
      "content": "<p>How do you deal with the prediction for billion of rows ?<br>\nI cross-product the users and items. Where one user require 1 second to predict.<br>\nEven I got 100x speedup I would have to wait for many hours.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1765329,
      "author_name": "Try Harder",
      "author_url": "",
      "post_date": "2022-04-23T10:54:32.180000",
      "content": "<p>How can I install cudf in kaggle kernel?<br>\npip install cudf ocurr with error<br>\nThank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1765362,
          "author_name": "Hervind Philipe",
          "author_url": "",
          "post_date": "2022-04-23T11:40:39.707000",
          "content": "<p>it is already installed if you choose The GPU</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1765423,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-04-23T13:35:48.357000",
          "content": "<p>Thanks a lot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1762235,
      "author_name": "Try Harder",
      "author_url": "",
      "post_date": "2022-04-20T14:31:03",
      "content": "<p>What dose CV mean? Thanks </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1762351,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-20T16:09:17.550000",
          "content": "<p>\"cross validation\",  but it's used more loosely to refer to the local validation you're doing on your models, as opposed to the result you get on the leader board.</p>\n<p>Check out these discussions in regards to CV for this competition:<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">discussion #1</a><br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">discussion #2</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1762828,
          "author_name": "Try Harder",
          "author_url": "",
          "post_date": "2022-04-21T01:43:55.347000",
          "content": "<p>Thank you for your patience！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1744086,
      "author_name": "Pan Adam",
      "author_url": "",
      "post_date": "2022-04-03T15:30:45.683000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Thanks for sharing!<br>\nI would like to ask <br>\nfor now, how many different recall methods are used when you generating candidates?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1744137,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-03T16:22:17.167000",
          "content": "<p>The basic recall methods so far are just from cdeotte's notebook - <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021</a></p>\n<p>But I've made several changes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1749009,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-08T07:06:20.293000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1762839,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T02:18:15.370000",
      "content": "<p>Thanks for sharing! Great job!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1753555,
      "author_name": "Zi Hao",
      "author_url": "",
      "post_date": "2022-04-13T00:22:21.410000",
      "content": "<p>Thank you very much for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1732018,
      "author_name": "stallone",
      "author_url": "",
      "post_date": "2022-03-22T23:57:59.470000",
      "content": "<p>Thanks for share </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1776443,
      "author_name": "Hao Yin J",
      "author_url": "",
      "post_date": "2022-05-04T01:39:13.523000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1731808": "I'm sharing some details of my setup and what I've found helpful so far (`2022-03-22`)\nLikely not very interesting for top teams, but hopefully helpful for other people like me.\n\nFor those of you who've gotten decent results and would be willing to share what you've found helpful, I would very much appreciate it!\n\n- Am working entirely with single Kaggle notebook.\n- Abstracting as much as possible to helper functions, and storing in modules as kaggle dataset, to keep main notebook uncluttered and readable.\n- Am using google cloud for storing outputs/inputs of different versions of notebooks.\n- Using gpu with cudf allows quick iteration and testing - once data is loaded and prepared, experiments on full dataset with 3 cvs taking 1/2 a minute each (haven't started modeling yet)\n- CV set up similar to @cdeotte's [here](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308919)\n- I've created a `week_number` column, with week numbers 0 to 104, so I can deal with weeks instead of days for cv (and other purposes)\n- Like many people, I've been using @paweljankiewicz's  [thread](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/307288) for guidance.\n- So far, have been focusing on step #1 - generating candidates. \n- To evaluate how \"good\" my generated candidates are, in cv, I look at:    \n    1. `recall` - what percentage of unique true customer_id/article_id purchases I have in the candidates.\n    2. `multiple factor` - how many times more total candidates I have than true candidates (like the inverse of precision)\n- For example (real example), if evaluation week of last cv has ~214k unique customer_id/article_id purchases, and I generated ~29,970k candidates, ~16k of which were actual purchases:\n    1. `recall` = 16k / 214k = ~7.5%\n    2. `multiple factor` = 29,970k / 16k = 1,878\n- As I create candidates, I also create/store the relevant features on the basis of which candidates were chosen. Currently, I'm only using those features to manually sort/choose from among the candidates, but in next step, they can be features for a ranking model\n- With gpu/cudf, only part that's taking longer than a minute is a version of @cdeotte pair generating code, which is taking about 30-45 minutes for the entire dataset.\n- Important to take into account article availability - if an item isn't available, doesn't really matter if it's likely that a customer would buy it.\n\nUpdate `2022-03-24`:\nImplemented draft version of using LGBMRanker\n- Memory starting to become an issue\n- For training/prediction, converted from cudf to pandas (`to_pandas()`)\n- Without any parameter tuning (except for n_estimators), got a +.0005 boost\n\nUpdate `2022-04-06`:\n- wrote/setup modeling code properly\n- Implemented [proper cv pipeline for modeling](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386)\n- Got memory issues under control by:\n    - [using `gc.collect()` more often](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316193)\n    - changing int columns with nulls to float32 before calling to `to_pandas()`\n    - [pickling items while they're not needed](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314829)\n    - changing dfs in-place, instead of returning new df each time\n- able to improve LB, but can't get it to correlate to cv. Working on that now.\n\nUpdate `2022-04-27`:\n- after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.\n- In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.\n- Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.\n\nUpdate `2022-05-05`:\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage...\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.\nWe'll see what I can do in a few days...",
    "1754550": "I've just noticed this thread. I admire you that you want to control the memory and use Kaggle notebook. I did almost the same till around half of the competition but used 128 GB RAM which I considered a concession when working with this problem - I'm blessed with a home PC with this amount of memory and 32 threads. Only since about 2 weeks I switched to the cloud and 256 GB RAM because I was having too many issues and I suspect by the time it is over I will probably train the last models on 512 GB RAM. I'm writing this just to give you a heads up what would you need to do if you wanted to \"quickly\" improve on the LB. Just generate more candidates/train bigger models.",
    "1733846": "I have been also having trouble with creating good recommendations.\n\nrecall of about ~7.5% seems small. But I haven't got much better results than you. I am now trying to cluster customers and articles. The goal is to find a reasonable number of articles that are mostly bought by fewer customers, so it is easy to build recommendations for them.\n\nI've been focusing with Males and distinguishing sales_channel_id. 80% of articles bought in a week come from previously items bought in the past. So I am working with about 20.000 customers and 10.000 articles.\n\nAnd then I can build proper recommendations to rank them with LGBMRanker.\n\nI don't know if this approach will work. And if it would be feasible for other \"segments\" like females (most customers buy female clothes). \n\nThe main goal to segment is simply to reduce the amount of articles each set of customers buy, so it is easy to build recommendations (I think this is the most difficult problem of the competition. As ranking with good recommendations is easy). ",
    "1769938": "Update `2022-04-27`:\n- after improving memory usage, I'm able to add more features and use more than one period's candidates for training. Not helping much so far.\n- In general, finding that adding candidates by casting a wider net is not improving my results - can't get lgbmranker to do well on higher-recall/lower-precision candidates.\n- Still struggling with issue posted here (lgbmranker doesn't improve loss, even on training set, and that's breaking my cv/lb correlation as well). It's a big problem, and I'm stuck on it.\n",
    "1760747": "Thanks for sharing, that's very useful",
    "1742244": "Very useful information, thanks for sharing.\n\nI have two doubts:\n\n1) You say that you use chris's code to generate products that are frequently purchased together.\nThis code uses the entire transaction dataframe regardless of time. Therefore if we include this information directly in a model I think we can overfit the model. I'm curious about how you have implemented this.\n\n2) You say that it is important to take into account the availability of the products. Do you have any more details about how you get to know this availability.\n\nThanks in advance",
    "1733665": "Updated my post with initial LGBMRanker info",
    "1732090": "Hi @jacob34 Thanks for sharing! \n\nI would like to ask, the ~29,970k candidates are from:\n- ~69k customer in last week cv only or \n- ~1,372k customer in submission file\n",
    "1777728": "Update `2022-05-05`:\n\nFinally figured out what's going on with my lgbmranker - there was a nasty, silent leakage...\nTook me a long time to track down, but very happy that I can now train my lgbmranker properly with a proper cv.\nWe'll see what I can do in a few days...",
    "1774862": "How do you deal with the prediction for billion of rows ?\nI cross-product the users and items. Where one user require 1 second to predict.\nEven I got 100x speedup I would have to wait for many hours.",
    "1765329": "How can I install cudf in kaggle kernel?\npip install cudf ocurr with error\nThank you",
    "1762235": "What dose CV mean? Thanks ",
    "1744086": "Hi @jacob34 Thanks for sharing!\nI would like to ask \nfor now, how many different recall methods are used when you generating candidates?",
    "1762839": "Thanks for sharing! Great job!",
    "1753555": "Thank you very much for sharing",
    "1732018": "Thanks for share ",
    "1776443": "Thanks for sharing."
  }
}