{
  "id": 308810,
  "title": "Best way to sample - Model construction",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/308810",
  "author_name": "",
  "post_date": "2022-02-20T12:45:20.446110500Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello, I have prepared the customer and articles (static and dynamic) attributes <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> suggested and I have the set of recommendations (the positive and negative registers) that we will enter to the LGBMRanker. </p>\n<p>Results have been quite good locally (with a sample of 20,000 customers). I have considered 8 weeks of training and 1 week of test. </p>\n<p>However, when I try to increment the number of customers the process cuts down due to RAM issues or lasts ages (when doing any merge or when training the model). </p>\n<p><strong>How would you approach this issue?</strong></p>\n<p>I was thinking about implementing as many models considering 20.000 samples of customers (we will have many models however). But I don't like this approach.  </p>\n<p>Thanks for any input! </p>",
  "messages": [
    {
      "id": "1698504",
      "postDate": "02/20/2022 12:45:20",
      "content": "<p>Hello, I have prepared the customer and articles (static and dynamic) attributes <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> suggested and I have the set of recommendations (the positive and negative registers) that we will enter to the LGBMRanker. </p>\n<p>Results have been quite good locally (with a sample of 20,000 customers). I have considered 8 weeks of training and 1 week of test. </p>\n<p>However, when I try to increment the number of customers the process cuts down due to RAM issues or lasts ages (when doing any merge or when training the model). </p>\n<p><strong>How would you approach this issue?</strong></p>\n<p>I was thinking about implementing as many models considering 20.000 samples of customers (we will have many models however). But I don't like this approach.  </p>\n<p>Thanks for any input! </p>",
      "rawMarkdown": "Hello, I have prepared the customer and articles (static and dynamic) attributes @paweljankiewicz suggested and I have the set of recommendations (the positive and negative registers) that we will enter to the LGBMRanker. \n\nResults have been quite good locally (with a sample of 20,000 customers). I have considered 8 weeks of training and 1 week of test. \n\nHowever, when I try to increment the number of customers the process cuts down due to RAM issues or lasts ages (when doing any merge or when training the model). \n\n**How would you approach this issue?**\n\nI was thinking about implementing as many models considering 20.000 samples of customers (we will have many models however). But I don't like this approach.  \n\nThanks for any input!",
      "votes": null
    },
    {
      "id": "1698687",
      "postDate": "02/20/2022 15:25:17",
      "content": "<p>Haven't gotten to that point yet, but I expect I will soon, and it will be something everyone will struggle with sooner or later.</p>\n<p>Some ideas I've been thinking of:</p>\n<ul>\n<li>delete any python objects you don't need (make sure to do this within the notebook cell the object was created, or you may not get your memory back)</li>\n<li>reduce feature sizes (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">see this discussion</a>)</li>\n<li>when creating/choosing features, add them one at a time, so you can tell which are actually important</li>\n<li>apply logic to manually cut down set of recommendations (i.e. - women won't purchase men's clothing)</li>\n<li>try alternatives to using merge (i.e. using pd.Series.map for each feature) - merge is a real RAM killer</li>\n<li>Move to colab for LB training/prediction (if you experiment with sampling, and just do entire dataset for periodic LB submission, you won't need that much usage)</li>\n<li>get a very good local cv, and use that to convince someone with a powerful machine to team up with you 😉</li>\n</ul>\n<p>But honestly, you (and most of us) would probably benefit from responses from people who have dealt with issues like this already in the past.</p>",
      "rawMarkdown": "Haven't gotten to that point yet, but I expect I will soon, and it will be something everyone will struggle with sooner or later.\n\nSome ideas I've been thinking of:\n- delete any python objects you don't need (make sure to do this within the notebook cell the object was created, or you may not get your memory back)\n- reduce feature sizes ([see this discussion](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635))\n- when creating/choosing features, add them one at a time, so you can tell which are actually important\n- apply logic to manually cut down set of recommendations (i.e. - women won't purchase men's clothing)\n- try alternatives to using merge (i.e. using pd.Series.map for each feature) - merge is a real RAM killer\n- Move to colab for LB training/prediction (if you experiment with sampling, and just do entire dataset for periodic LB submission, you won't need that much usage)\n- get a very good local cv, and use that to convince someone with a powerful machine to team up with you 😉\n\nBut honestly, you (and most of us) would probably benefit from responses from people who have dealt with issues like this already in the past.",
      "votes": null
    },
    {
      "id": "1698719",
      "postDate": "02/20/2022 15:54:33",
      "content": "<p>You can try online learning with Vowpal Wabbit</p>",
      "rawMarkdown": "You can try online learning with Vowpal Wabbit",
      "votes": null
    },
    {
      "id": "1698882",
      "postDate": "02/20/2022 18:30:52",
      "content": "<p>There are a couple of possible answers:</p>\n<ul>\n<li>First of all ranking algorithms work in such a way that you gain nothing from \"queries\" without the positive observation. So if your candidates for the customer don't have a transaction in the next week then you can safely delete the candidates for the customer on this date.</li>\n<li>Chunk the data as much as possible. For me chunks are generated like this - for every observation date I have 32 to 128 chunks of customers (32 for normal days, 128 chunks for submission date 09-22). I use Rust to generate all the data and it has excellent parallelization that speeds up whole process. Rust has this advantage that I have a shared transactions index by date. If I was using Python I would probably generate each chunk in a separate process even at a cost of loading all the transactions.</li>\n<li>There is a Dask LGBMRanker which can come handy - haven't tried it yet but I need to do it soon - it can be an answer for the costly merge of chunks</li>\n</ul>\n<p>I have been testing this simple code for merging - which is slower than pd.concat but doesn't blow up the memory</p>\n<pre><code>X_train = pd.DataFrame()\nfor df in tqdm(Xs_train):\n     X_train = pd.concat([X_train, df])\n     del df\n     gc.collect()\n</code></pre>\n<ul>\n<li>One thing that I just tried and it doesn't worsen my results is use initial prescoring of candidates to limit the number of candidates that you consider for the training and prediction. </li>\n</ul>",
      "rawMarkdown": "There are a couple of possible answers:\n- First of all ranking algorithms work in such a way that you gain nothing from \"queries\" without the positive observation. So if your candidates for the customer don't have a transaction in the next week then you can safely delete the candidates for the customer on this date.\n- Chunk the data as much as possible. For me chunks are generated like this - for every observation date I have 32 to 128 chunks of customers (32 for normal days, 128 chunks for submission date 09-22). I use Rust to generate all the data and it has excellent parallelization that speeds up whole process. Rust has this advantage that I have a shared transactions index by date. If I was using Python I would probably generate each chunk in a separate process even at a cost of loading all the transactions.\n- There is a Dask LGBMRanker which can come handy - haven't tried it yet but I need to do it soon - it can be an answer for the costly merge of chunks\n\nI have been testing this simple code for merging - which is slower than pd.concat but doesn't blow up the memory\n```\nX_train = pd.DataFrame()\nfor df in tqdm(Xs_train):\n     X_train = pd.concat([X_train, df])\n     del df\n     gc.collect()\n```\n\n- One thing that I just tried and it doesn't worsen my results is use initial prescoring of candidates to limit the number of candidates that you consider for the training and prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1698687,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "02/20/2022 15:25:17",
      "content": "<p>Haven't gotten to that point yet, but I expect I will soon, and it will be something everyone will struggle with sooner or later.</p>\n<p>Some ideas I've been thinking of:</p>\n<ul>\n<li>delete any python objects you don't need (make sure to do this within the notebook cell the object was created, or you may not get your memory back)</li>\n<li>reduce feature sizes (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">see this discussion</a>)</li>\n<li>when creating/choosing features, add them one at a time, so you can tell which are actually important</li>\n<li>apply logic to manually cut down set of recommendations (i.e. - women won't purchase men's clothing)</li>\n<li>try alternatives to using merge (i.e. using pd.Series.map for each feature) - merge is a real RAM killer</li>\n<li>Move to colab for LB training/prediction (if you experiment with sampling, and just do entire dataset for periodic LB submission, you won't need that much usage)</li>\n<li>get a very good local cv, and use that to convince someone with a powerful machine to team up with you 😉</li>\n</ul>\n<p>But honestly, you (and most of us) would probably benefit from responses from people who have dealt with issues like this already in the past.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1698719,
      "author_name": "souamesannis",
      "author_url": "",
      "post_date": "02/20/2022 15:54:33",
      "content": "<p>You can try online learning with Vowpal Wabbit</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1698882,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "02/20/2022 18:30:52",
      "content": "<p>There are a couple of possible answers:</p>\n<ul>\n<li>First of all ranking algorithms work in such a way that you gain nothing from \"queries\" without the positive observation. So if your candidates for the customer don't have a transaction in the next week then you can safely delete the candidates for the customer on this date.</li>\n<li>Chunk the data as much as possible. For me chunks are generated like this - for every observation date I have 32 to 128 chunks of customers (32 for normal days, 128 chunks for submission date 09-22). I use Rust to generate all the data and it has excellent parallelization that speeds up whole process. Rust has this advantage that I have a shared transactions index by date. If I was using Python I would probably generate each chunk in a separate process even at a cost of loading all the transactions.</li>\n<li>There is a Dask LGBMRanker which can come handy - haven't tried it yet but I need to do it soon - it can be an answer for the costly merge of chunks</li>\n</ul>\n<p>I have been testing this simple code for merging - which is slower than pd.concat but doesn't blow up the memory</p>\n<pre><code>X_train = pd.DataFrame()\nfor df in tqdm(Xs_train):\n     X_train = pd.concat([X_train, df])\n     del df\n     gc.collect()\n</code></pre>\n<ul>\n<li>One thing that I just tried and it doesn't worsen my results is use initial prescoring of candidates to limit the number of candidates that you consider for the training and prediction. </li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1698504": "Hello, I have prepared the customer and articles (static and dynamic) attributes @paweljankiewicz suggested and I have the set of recommendations (the positive and negative registers) that we will enter to the LGBMRanker. \n\nResults have been quite good locally (with a sample of 20,000 customers). I have considered 8 weeks of training and 1 week of test. \n\nHowever, when I try to increment the number of customers the process cuts down due to RAM issues or lasts ages (when doing any merge or when training the model). \n\n**How would you approach this issue?**\n\nI was thinking about implementing as many models considering 20.000 samples of customers (we will have many models however). But I don't like this approach.  \n\nThanks for any input!",
    "1698687": "Haven't gotten to that point yet, but I expect I will soon, and it will be something everyone will struggle with sooner or later.\n\nSome ideas I've been thinking of:\n- delete any python objects you don't need (make sure to do this within the notebook cell the object was created, or you may not get your memory back)\n- reduce feature sizes ([see this discussion](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635))\n- when creating/choosing features, add them one at a time, so you can tell which are actually important\n- apply logic to manually cut down set of recommendations (i.e. - women won't purchase men's clothing)\n- try alternatives to using merge (i.e. using pd.Series.map for each feature) - merge is a real RAM killer\n- Move to colab for LB training/prediction (if you experiment with sampling, and just do entire dataset for periodic LB submission, you won't need that much usage)\n- get a very good local cv, and use that to convince someone with a powerful machine to team up with you 😉\n\nBut honestly, you (and most of us) would probably benefit from responses from people who have dealt with issues like this already in the past.",
    "1698719": "You can try online learning with Vowpal Wabbit",
    "1698882": "There are a couple of possible answers:\n- First of all ranking algorithms work in such a way that you gain nothing from \"queries\" without the positive observation. So if your candidates for the customer don't have a transaction in the next week then you can safely delete the candidates for the customer on this date.\n- Chunk the data as much as possible. For me chunks are generated like this - for every observation date I have 32 to 128 chunks of customers (32 for normal days, 128 chunks for submission date 09-22). I use Rust to generate all the data and it has excellent parallelization that speeds up whole process. Rust has this advantage that I have a shared transactions index by date. If I was using Python I would probably generate each chunk in a separate process even at a cost of loading all the transactions.\n- There is a Dask LGBMRanker which can come handy - haven't tried it yet but I need to do it soon - it can be an answer for the costly merge of chunks\n\nI have been testing this simple code for merging - which is slower than pd.concat but doesn't blow up the memory\n```\nX_train = pd.DataFrame()\nfor df in tqdm(Xs_train):\n     X_train = pd.concat([X_train, df])\n     del df\n     gc.collect()\n```\n\n- One thing that I just tried and it doesn't worsen my results is use initial prescoring of candidates to limit the number of candidates that you consider for the training and prediction."
  },
  "source": "meta"
}